Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Comprehensive citation management for academic research. Resolve bibliographic identifiers, extract accurate metadata, validate citations, deduplicate references, and generate properly formatted BibTeX entries. This skill should be used when you need to verify citation information, convert DOIs/PMIDs/arXiv IDs to BibTeX, or ensure reference accuracy in scientific writing.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 290% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 435% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 308% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 404% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 322% | 0% |
Manage citations systematically throughout the research and writing process. This skill provides tools and strategies for extracting accurate metadata from identifiers and bibliographic sources, validating citation information, cleaning duplicate references, and generating properly formatted BibTeX entries.
Critical for maintaining citation accuracy, avoiding reference errors, and ensuring reproducible research.
Use this skill when:
Citation management follows a systematic process:
Goal: Locate exact source records for known references or tightly scoped citation candidates.
Google Scholar provides the most comprehensive coverage across disciplines.
Basic Search:
bash# Search for papers on a topic python scripts/search_google_scholar.py "CRISPR gene editing" \ --limit 50 \ --output results.json # Search with year filter python scripts/search_google_scholar.py "machine learning protein folding" \ --year-start 2020 \ --year-end 2024 \ --limit 100 \ --output ml_proteins.json
Advanced Search Strategies (see references/google_scholar_search.md):
"deep learning"author:LeCunintitle:"neural networks"machine learning -surveyBest Practices:
PubMed specializes in biomedical and life sciences literature (35+ million citations).
Basic Search:
bash# Search PubMed python scripts/search_pubmed.py "Alzheimer's disease treatment" \ --limit 100 \ --output alzheimers.json # Search with MeSH terms and filters python scripts/search_pubmed.py \ --query '"Alzheimer Disease"[MeSH] AND "Drug Therapy"[MeSH]' \ --date-start 2020 \ --date-end 2024 \ --publication-types "Clinical Trial,Review" \ --output alzheimers_trials.json
Advanced PubMed Queries (see references/pubmed_search.md):
"Diabetes Mellitus"[MeSH]"cancer"[Title], "Smith J"[Author]AND, OR, NOT2020:2024[Publication Date]"Review"[Publication Type]Best Practices:
Goal: Convert paper identifiers (DOI, PMID, arXiv ID) to complete, accurate metadata.
For single DOIs, use the quick conversion tool:
bash# Convert single DOI python scripts/doi_to_bibtex.py 10.1038/s41586-021-03819-2 # Convert multiple DOIs from a file python scripts/doi_to_bibtex.py --input dois.txt --output references.bib # Different output formats python scripts/doi_to_bibtex.py 10.1038/nature12345 --format json
For DOIs, PMIDs, arXiv IDs, or URLs:
bash# Extract from DOI python scripts/extract_metadata.py --doi 10.1038/s41586-021-03819-2 # Extract from PMID python scripts/extract_metadata.py --pmid 34265844 # Extract from arXiv ID python scripts/extract_metadata.py --arxiv 2103.14030 # Extract from URL python scripts/extract_metadata.py --url "https://www.nature.com/articles/s41586-021-03819-2" # Batch extraction from file (mixed identifiers) python scripts/extract_metadata.py --input identifiers.txt --output citations.bib
Metadata Sources (see references/metadata_extraction.md):
What Gets Extracted:
Goal: Generate clean, properly formatted BibTeX entries.
See references/bibtex_formatting.md for complete guide.
Common Entry Types:
@article: Journal articles (most common)@book: Books@inproceedings: Conference papers@incollection: Book chapters@phdthesis: Dissertations@misc: Preprints, software, datasetsRequired Fields by Type:
bibtex@article{citationkey, author = {Last1, First1 and Last2, First2}, title = {Article Title}, journal = {Journal Name}, year = {2024}, volume = {10}, number = {3}, pages = {123--145}, doi = {10.1234/example} } @inproceedings{citationkey, author = {Last, First}, title = {Paper Title}, booktitle = {Conference Name}, year = {2024}, pages = {1--10} } @book{citationkey, author = {Last, First}, title = {Book Title}, publisher = {Publisher Name}, year = {2024} }
Use the formatter to standardize BibTeX files:
bash# Format and clean BibTeX file python scripts/format_bibtex.py references.bib \ --output formatted_references.bib # Sort entries by citation key python scripts/format_bibtex.py references.bib \ --sort key \ --output sorted_references.bib # Sort by year (newest first) python scripts/format_bibtex.py references.bib \ --sort year \ --descending \ --output sorted_references.bib # Remove duplicates python scripts/format_bibtex.py references.bib \ --deduplicate \ --output clean_references.bib # Validate and report issues python scripts/format_bibtex.py references.bib \ --validate \ --report validation_report.txt
Formatting Operations:
Goal: Verify all citations are accurate and complete.
bash# Validate BibTeX file python scripts/validate_citations.py references.bib # Validate and fix common issues python scripts/validate_citations.py references.bib \ --auto-fix \ --output validated_references.bib # Generate detailed validation report python scripts/validate_citations.py references.bib \ --report validation_report.json \ --verbose
Validation Checks (see references/citation_validation.md):
Validation Output:
json{ "total_entries": 150, "valid_entries": 145, "errors": [ { "citation_key": "Smith2023", "error_type": "missing_field", "field": "journal", "severity": "high" }, { "citation_key": "Jones2022", "error_type": "invalid_doi", "doi": "10.1234/broken", "severity": "high" } ], "warnings": [ { "citation_key": "Brown2021", "warning_type": "possible_duplicate", "duplicate_of": "Brown2021a", "severity": "medium" } ] }
Complete workflow for creating a bibliography:
bash# 1. Search for papers on your topic python scripts/search_pubmed.py \ '"CRISPR-Cas Systems"[MeSH] AND "Gene Editing"[MeSH]' \ --date-start 2020 \ --limit 200 \ --output crispr_papers.json # 2. Extract DOIs from search results and convert to BibTeX python scripts/extract_metadata.py \ --input crispr_papers.json \ --output crispr_refs.bib # 3. Add specific papers by DOI python scripts/doi_to_bibtex.py 10.1038/nature12345 >> crispr_refs.bib python scripts/doi_to_bibtex.py 10.1126/science.abcd1234 >> crispr_refs.bib # 4. Format and clean the BibTeX file python scripts/format_bibtex.py crispr_refs.bib \ --deduplicate \ --sort year \ --descending \ --output references.bib # 5. Validate all citations python scripts/validate_citations.py references.bib \ --auto-fix \ --report validation.json \ --output final_references.bib # 6. Review validation report and fix any remaining issues cat validation.json # 7. Use in your LaTeX document # \bibliography{final_references}
When a review document already exists, use this skill for the technical citation layer:
bash# After completing literature review # Verify all citations in the review document python scripts/validate_citations.py my_review_references.bib --report review_validation.json # Format for specific citation style if needed python scripts/format_bibtex.py my_review_references.bib \ --style nature \ --output formatted_refs.bib
Finding Seminal and High-Impact Papers (CRITICAL):
Always prioritize papers based on citation count, venue quality, and author reputation:
Citation Count Thresholds: | Paper Age | Citations | Classification | |-----------|-----------|----------------| | 0-3 years | 20+ | Noteworthy | | 0-3 years | 100+ | Highly Influential | | 3-7 years | 100+ | Significant | | 3-7 years | 500+ | Landmark Paper | | 7+ years | 500+ | Seminal Work | | 7+ years | 1000+ | Foundational |
Venue Quality Tiers:
Author Reputation Indicators:
Search Strategies for High-Impact Papers:
source:Nature or source:Scienceauthor:LastNameAdvanced Operators (full list in references/google_scholar_search.md):
"exact phrase" # Exact phrase matching
author:lastname # Search by author
intitle:keyword # Search in title only
source:journal # Search specific journal
-exclude # Exclude terms
OR # Alternative terms
2020..2024 # Year rangeExample Searches:
# Find recent reviews on a topic
"CRISPR" intitle:review 2023..2024
# Find papers by specific author on topic
author:Church "synthetic biology"
# Find highly cited foundational work
"deep learning" 2012..2015 sort:citations
# Exclude surveys and focus on methods
"protein folding" -survey -review intitle:methodUsing MeSH Terms: MeSH (Medical Subject Headings) provides controlled vocabulary for precise searching.
"Diabetes Mellitus, Type 2"[MeSH]Field Tags:
[Title] # Search in title only
[Title/Abstract] # Search in title or abstract
[Author] # Search by author name
[Journal] # Search specific journal
[Publication Date] # Date range
[Publication Type] # Article type
[MeSH] # MeSH termBuilding Complex Queries:
bash# Clinical trials on diabetes treatment published recently "Diabetes Mellitus, Type 2"[MeSH] AND "Drug Therapy"[MeSH] AND "Clinical Trial"[Publication Type] AND 2020:2024[Publication Date] # Reviews on CRISPR in specific journal "CRISPR-Cas Systems"[MeSH] AND "Nature"[Journal] AND "Review"[Publication Type] # Specific author's recent work "Smith AB"[Author] AND cancer[Title/Abstract] AND 2022:2024[Publication Date]
E-utilities for Automation: The scripts use NCBI E-utilities API for programmatic access:
See references/pubmed_search.md for complete API documentation.
Search Google Scholar and export results.
Features:
Usage:
bash# Basic search python scripts/search_google_scholar.py "quantum computing" # Advanced search with filters python scripts/search_google_scholar.py "quantum computing" \ --year-start 2020 \ --year-end 2024 \ --limit 100 \ --sort-by citations \ --output quantum_papers.json # Export directly to BibTeX python scripts/search_google_scholar.py "machine learning" \ --limit 50 \ --format bibtex \ --output ml_papers.bib
Search PubMed using E-utilities API.
Features:
Usage:
bash# Simple keyword search python scripts/search_pubmed.py "CRISPR gene editing" # Complex query with filters python scripts/search_pubmed.py \ --query '"CRISPR-Cas Systems"[MeSH] AND "therapeutic"[Title/Abstract]' \ --date-start 2020-01-01 \ --date-end 2024-12-31 \ --publication-types "Clinical Trial,Review" \ --limit 200 \ --output crispr_therapeutic.json # Export to BibTeX python scripts/search_pubmed.py "Alzheimer's disease" \ --limit 100 \ --format bibtex \ --output alzheimers.bib
Extract complete metadata from paper identifiers.
Features:
Usage:
bash# Single DOI python scripts/extract_metadata.py --doi 10.1038/s41586-021-03819-2 # Single PMID python scripts/extract_metadata.py --pmid 34265844 # Single arXiv ID python scripts/extract_metadata.py --arxiv 2103.14030 # From URL python scripts/extract_metadata.py \ --url "https://www.nature.com/articles/s41586-021-03819-2" # Batch processing (file with one identifier per line) python scripts/extract_metadata.py \ --input paper_ids.txt \ --output references.bib # Different output formats python scripts/extract_metadata.py \ --doi 10.1038/nature12345 \ --format json # or bibtex, yaml
Validate BibTeX entries for accuracy and completeness.
Features:
Usage:
bash# Basic validation python scripts/validate_citations.py references.bib # With auto-fix python scripts/validate_citations.py references.bib \ --auto-fix \ --output fixed_references.bib # Detailed validation report python scripts/validate_citations.py references.bib \ --report validation_report.json \ --verbose # Only check DOIs python scripts/validate_citations.py references.bib \ --check-dois-only
Format and clean BibTeX files.
Features:
Usage:
bash# Basic formatting python scripts/format_bibtex.py references.bib # Sort by year (newest first) python scripts/format_bibtex.py references.bib \ --sort year \ --descending \ --output sorted_refs.bib # Remove duplicates python scripts/format_bibtex.py references.bib \ --deduplicate \ --output clean_refs.bib # Complete cleanup python scripts/format_bibtex.py references.bib \ --deduplicate \ --sort year \ --validate \ --auto-fix \ --output final_refs.bib
Quick DOI to BibTeX conversion.
Features:
Usage:
bash# Single DOI python scripts/doi_to_bibtex.py 10.1038/s41586-021-03819-2 # Multiple DOIs python scripts/doi_to_bibtex.py \ 10.1038/nature12345 \ 10.1126/science.abc1234 \ 10.1016/j.cell.2023.01.001 # From file (one DOI per line) python scripts/doi_to_bibtex.py --input dois.txt --output references.bib # Copy to clipboard python scripts/doi_to_bibtex.py 10.1038/nature12345 --clipboard
bash# Step 1: Find key papers on your topic python scripts/search_google_scholar.py "transformer neural networks" \ --year-start 2017 \ --limit 50 \ --output transformers_gs.json python scripts/search_pubmed.py "deep learning medical imaging" \ --date-start 2020 \ --limit 50 \ --output medical_dl_pm.json # Step 2: Extract metadata from search results python scripts/extract_metadata.py \ --input transformers_gs.json \ --output transformers.bib python scripts/extract_metadata.py \ --input medical_dl_pm.json \ --output medical.bib # Step 3: Add specific papers you already know python scripts/doi_to_bibtex.py 10.1038/s41586-021-03819-2 >> specific.bib python scripts/doi_to_bibtex.py 10.1126/science.aam9317 >> specific.bib # Step 4: Combine all BibTeX files cat transformers.bib medical.bib specific.bib > combined.bib # Step 5: Format and deduplicate python scripts/format_bibtex.py combined.bib \ --deduplicate \ --sort year \ --descending \ --output formatted.bib # Step 6: Validate python scripts/validate_citations.py formatted.bib \ --auto-fix \ --report validation.json \ --output final_references.bib # Step 7: Review any issues cat validation.json | grep -A 3 '"errors"' # Step 8: Use in LaTeX # \bibliography{final_references}
bash# You have a text file with DOIs (one per line) # dois.txt contains: # 10.1038/s41586-021-03819-2 # 10.1126/science.aam9317 # 10.1016/j.cell.2023.01.001 # Convert all to BibTeX python scripts/doi_to_bibtex.py --input dois.txt --output references.bib # Validate the result python scripts/validate_citations.py references.bib --verbose
bash# You have a messy BibTeX file from various sources # Clean it up systematically # Step 1: Format and standardize python scripts/format_bibtex.py messy_references.bib \ --output step1_formatted.bib # Step 2: Remove duplicates python scripts/format_bibtex.py step1_formatted.bib \ --deduplicate \ --output step2_deduplicated.bib # Step 3: Validate and auto-fix python scripts/validate_citations.py step2_deduplicated.bib \ --auto-fix \ --output step3_validated.bib # Step 4: Sort by year python scripts/format_bibtex.py step3_validated.bib \ --sort year \ --descending \ --output clean_references.bib # Step 5: Final validation report python scripts/validate_citations.py clean_references.bib \ --report final_validation.json \ --verbose # Review report cat final_validation.json
bash# Find highly cited papers on a topic python scripts/search_google_scholar.py "AlphaFold protein structure" \ --year-start 2020 \ --year-end 2024 \ --sort-by citations \ --limit 20 \ --output alphafold_seminal.json # Extract the top 10 by citation count # (script will have included citation counts in JSON) # Convert to BibTeX python scripts/extract_metadata.py \ --input alphafold_seminal.json \ --output alphafold_refs.bib # The BibTeX file now contains the most influential papers
Citation management outputs are reusable in writing and submission workflows:
References (in references/):
google_scholar_search.md: Complete Google Scholar search guidepubmed_search.md: PubMed and E-utilities API documentationmetadata_extraction.md: Metadata sources and field requirementscitation_validation.md: Validation criteria and quality checksbibtex_formatting.md: BibTeX entry types and formatting rulesScripts (in scripts/):
search_google_scholar.py: Google Scholar search automationsearch_pubmed.py: PubMed E-utilities API clientextract_metadata.py: Universal metadata extractorvalidate_citations.py: Citation validation and verificationformat_bibtex.py: BibTeX formatter and cleanerdoi_to_bibtex.py: Quick DOI to BibTeX converterAssets (in assets/):
bibtex_template.bib: Example BibTeX entries for all typescitation_checklist.md: Quality assurance checklistSearch Engines:
Metadata APIs:
Tools and Validators:
Citation Styles:
bash# Core dependencies pip install requests # HTTP requests for APIs pip install bibtexparser # BibTeX parsing and formatting pip install biopython # PubMed E-utilities access # Optional (for Google Scholar) pip install scholarly # Google Scholar API wrapper # or pip install selenium # For more robust Scholar scraping
bash# For advanced validation pip install crossref-commons # Enhanced CrossRef API access pip install pylatexenc # LaTeX special character handling
The citation-management skill provides:
Use this skill to maintain accurate, complete citations throughout your research and ensure publication-ready bibliographies.
Other measured skills in the registry, with their headline benchmark lift.