Comprehensive citation management for academic research. Search Google Scholar and PubMed for papers, extract accurate metadata, validate citations, and generate properly formatted BibTeX entries. This skill should be used when you need to find papers, verify citation information, convert DOIs to BibTeX, or ensure reference accuracy in scientific writing.
Comprehensive citation management for academic research. Search Google Scholar and PubMed for papers, extract accurate metadata, validate citations, and generate properly formatted BibTeX entries. This skill should be used when you need to find papers, verify citation information, convert DOIs to BibTeX, or ensure reference accuracy in scientific writing.
allowed-tools
Read Write Edit Bash
license
MIT License
required_environment_variables
[{"name":"OPENROUTER_API_KEY","prompt":"OpenRouter API key for LLM-powered citation steps.","required_for":"optional features"},{"name":"NCBI_EMAIL","prompt":"Email for NCBI Entrez identification.","required_for":"optional features"},{"name":"NCBI_API_KEY","prompt":"NCBI API key to raise Entrez rate limits.","required_for":"optional features"}]
metadata
{"version":"1.2","skill-author":"K-Dense Inc.","openclaw":{"primaryEnv":"OPENROUTER_API_KEY","envVars":[{"name":"OPENROUTER_API_KEY","required":false,"description":"OpenRouter API key for LLM-powered citation steps."},{"name":"NCBI_EMAIL","required":false,"description":"Email for NCBI Entrez identification."},{"name":"NCBI_API_KEY","required":false,"description":"NCBI API key to raise Entrez rate limits."}]}}
Citation Management
Overview
Manage citations systematically throughout the research and writing process. This skill provides tools and strategies for searching academic databases (Google Scholar, PubMed), extracting accurate metadata from multiple sources (CrossRef, PubMed, arXiv), validating citation information, and generating properly formatted BibTeX entries.
Critical for maintaining citation accuracy, avoiding reference errors, and ensuring reproducible research. Integrates seamlessly with the literature-review skill for comprehensive research workflows.
When to Use This Skill
Use this skill when:
Searching for specific papers on Google Scholar or PubMed
Converting DOIs, PMIDs, or arXiv IDs to properly formatted BibTeX
Extracting complete metadata for citations (authors, title, journal, year, etc.)
Validating existing citations for accuracy
Cleaning and formatting BibTeX files
Finding highly cited papers in a specific field
Verifying that citation information matches the actual publication
Building a bibliography for a manuscript or thesis
Checking for duplicate citations
Ensuring consistent citation formatting
Visual Enhancement with Scientific Schematics
When creating documents with this skill, always consider adding scientific diagrams and schematics to enhance visual communication.
If your document does not already contain schematics or diagrams:
Use the scientific-schematics skill to generate AI-powered publication-quality diagrams
Simply describe your desired diagram in natural language
Nano Banana Pro will automatically generate, review, and refine the schematic
For new documents: Scientific schematics should be generated by default to visually represent key concepts, workflows, architectures, or relationships described in the text.
Advanced PubMed Queries (see references/pubmed_search.md):
Use MeSH terms: "Diabetes Mellitus"[MeSH]
Field tags: "cancer"[Title], "Smith J"[Author]
Boolean operators: AND, OR, NOT
Date filters: 2020:2024[Publication Date]
Publication types: "Review"[Publication Type]
Combine with E-utilities API for automation
Best Practices:
Use MeSH Browser to find correct controlled vocabulary
Construct complex queries in PubMed Advanced Search Builder first
Include multiple synonyms with OR
Retrieve PMIDs for easy metadata extraction
Export to JSON or directly to BibTeX
Phase 2: Metadata Extraction
Goal: Convert paper identifiers (DOI, PMID, arXiv ID) to complete, accurate metadata.
Quick DOI to BibTeX Conversion
For single DOIs, use the quick conversion tool:
# Convert single DOI
python scripts/doi_to_bibtex.py 10.1038/s41586-021-03819-2
# Convert multiple DOIs from a file
python scripts/doi_to_bibtex.py --input dois.txt --output references.bib
# Different output formats
python scripts/doi_to_bibtex.py 10.1038/nature12345 --format json
Comprehensive Metadata Extraction
For DOIs, PMIDs, arXiv IDs, or URLs:
# Extract from DOI
python scripts/extract_metadata.py --doi 10.1038/s41586-021-03819-2
# Extract from PMID
python scripts/extract_metadata.py --pmid 34265844
# Extract from arXiv ID
python scripts/extract_metadata.py --arxiv 2103.14030
# Extract from URL
python scripts/extract_metadata.py --url "https://www.nature.com/articles/s41586-021-03819-2"# Batch extraction from file (mixed identifiers)
python scripts/extract_metadata.py --input identifiers.txt --output citations.bib
Metadata Sources (see references/metadata_extraction.md):
CrossRef API: Primary source for DOIs
Comprehensive metadata for journal articles
Publisher-provided information
Includes authors, title, journal, volume, pages, dates
Free, no API key required
PubMed E-utilities: Biomedical literature
Official NCBI metadata
Includes MeSH terms, abstracts
PMID and PMCID identifiers
Free, API key recommended for high volume
arXiv API: Preprints in physics, math, CS, q-bio
Complete metadata for preprints
Version tracking
Author affiliations
Free, open access
DataCite API: Research datasets, software, other resources
Metadata for non-traditional scholarly outputs
DOIs for datasets and code
Free access
What Gets Extracted:
Required fields: author, title, year
Journal articles: journal, volume, number, pages, DOI
Preprints: repository (arXiv, bioRxiv), preprint ID
Additional: abstract, keywords, URL
Phase 2.5: Metadata Enrichment via Web Search (MANDATORY)
Goal: Detect and fill in any missing metadata fields using web search. This phase runs AFTER extraction and BEFORE formatting to ensure every BibTeX entry is complete.
Why This Is Critical: Metadata extraction from APIs (CrossRef, PubMed, arXiv) sometimes returns incomplete records — missing volume, pages, issue number, or DOI. These gaps must be filled before the bibliography is considered ready.
Step 1: Scan for Incomplete Entries
After extracting metadata, scan the BibTeX file for entries missing key fields:
Fields to check per entry type:
Entry Type
Must Have
Should Have
@article
author, title, journal, year
volume, pages, number, doi
@inproceedings
author, title, booktitle, year
pages, doi
@book
author/editor, title, publisher, year
isbn, doi
@misc
author, title, year
doi or url
Any @article entry missing volume, pages, or doi is considered incomplete and must be enriched.
Step 2: Web Search for Missing Metadata
For each incomplete entry, use the parallel-web skill to search for the missing information:
Option A — Search by title and author (best for finding DOI):
Citations must always be high in number based on standards for journal and conference publications in the venue of choice or recommendation. Never settle for a sparse reference list; establish an authoritative, rich context with dense, verified citations.
ML / CS conferences (NeurIPS, ICML, ICLR, CVPR, ACL)
30-45+
Comprehensive literature reviews / market research reports
40-65+
Medical journals (NEJM, Lancet, JAMA)
30-45+
Always adjust the citation target upward depending on standard density and practices of the target venue. Avoid 'lazy' citation over-repetition — do not repeatedly cite the same 1 or 2 papers to support multiple unrelated claims; draw from a diverse, high-quality set of reputable references.
Enforce these standards programmatically with validate_citations.py --venue <venue> or --min-count <N>.
Once the entire scientific report or paper has been drafted and written, perform a comprehensive post-writing verification of all citations before compiling the final deliverables:
Verify No Missing or Unresolved Citations: Check the draft or compiled document to ensure that every in-text citation correctly resolves to a reference in references.bib. There must be ZERO broken citation keys, missing identifiers, or unresolved references (e.g., [?] or [citation needed]).
Verify No Unused (Dangling) Bibliography Entries: Check that every entry in references.bib is actually cited in the body of the report. Remove any unused entries to keep the bibliography perfectly clean.
Verify Citation Quantity Against Target Standards: Ensure the final citation count meets or exceeds the high standard of the chosen or recommended venue (see table above). If the count is below standard, perform additional literature search first, find high-quality papers, and integrate them into appropriate sections.
Verify Metadata Completeness: Confirm that all cited entries contain complete, fully-verified fields (all author names, complete journal/conference names, exact year, volume, issue, page range, and valid DOI).
# 1. Search for papers on your topic
python scripts/search_pubmed.py \
'"CRISPR-Cas Systems"[MeSH] AND "Gene Editing"[MeSH]' \
--date-start 2020 \
--limit 200 \
--output crispr_papers.json
# 2. Extract DOIs from search results and convert to BibTeX
python scripts/extract_metadata.py \
--input crispr_papers.json \
--output crispr_refs.bib
# 3. Add specific papers by DOI
python scripts/doi_to_bibtex.py 10.1038/nature12345 >> crispr_refs.bib
python scripts/doi_to_bibtex.py 10.1126/science.abcd1234 >> crispr_refs.bib
# 4. Format and clean the BibTeX file
python scripts/format_bibtex.py crispr_refs.bib \
--deduplicate \
--sort year \
--descending \
--output references.bib
# 5. Validate all citations
python scripts/validate_citations.py references.bib \
--auto-fix \
--report validation.json \
--output final_references.bib
# 6. Review validation report and fix any remaining issuescat validation.json
# 7. Use in your LaTeX document# \bibliography{final_references}
Integration with Literature Review Skill
This skill complements the literature-review skill:
Literature Review Skill → Systematic search and synthesis
Citation Management Skill → Technical citation handling
Combined Workflow:
Use literature-review for comprehensive multi-database search
Use citation-management to extract and validate all citations
Use literature-review to synthesize findings thematically
Use citation-management to verify final bibliography accuracy
# After completing literature review# Verify all citations in the review document
python scripts/validate_citations.py my_review_references.bib --report review_validation.json
# Format for specific citation style if needed
python scripts/format_bibtex.py my_review_references.bib \
--style nature \
--output formatted_refs.bib
Search Strategies
Google Scholar Best Practices
Finding Seminal and High-Impact Papers (CRITICAL):
Always prioritize papers based on citation count, venue quality, and author reputation:
Look for review articles from Tier-1 journals for overview
Check "Cited by" for impact assessment and recent follow-up work
Use citation alerts for tracking new citations to key papers
Filter by top venues using source:Nature or source:Science
Search for papers by known field leaders using author:LastName
Advanced Operators (full list in references/google_scholar_search.md):
"exact phrase" # Exact phrase matching
author:lastname # Search by author
intitle:keyword # Search in title only
source:journal # Search specific journal
-exclude # Exclude terms
OR # Alternative terms
2020..2024 # Year range
Example Searches:
# Find recent reviews on a topic
"CRISPR" intitle:review 2023..2024
# Find papers by specific author on topic
author:Church "synthetic biology"
# Find highly cited foundational work
"deep learning" 2012..2015 sort:citations
# Exclude surveys and focus on methods
"protein folding" -survey -review intitle:method
PubMed Best Practices
Using MeSH Terms:
MeSH (Medical Subject Headings) provides controlled vocabulary for precise searching.
[Title] # Search in title only
[Title/Abstract] # Search in title or abstract
[Author] # Search by author name
[Journal] # Search specific journal
[Publication Date] # Date range
[Publication Type] # Article type
[MeSH] # MeSH term
Building Complex Queries:
# Clinical trials on diabetes treatment published recently"Diabetes Mellitus, Type 2"[MeSH] AND "Drug Therapy"[MeSH]
AND "Clinical Trial"[Publication Type] AND 2020:2024[Publication Date]
# Reviews on CRISPR in specific journal"CRISPR-Cas Systems"[MeSH] AND "Nature"[Journal] AND "Review"[Publication Type]
# Specific author's recent work"Smith AB"[Author] AND cancer[Title/Abstract] AND 2022:2024[Publication Date]
E-utilities for Automation:
The scripts use NCBI E-utilities API for programmatic access:
ESearch: Search and retrieve PMIDs
EFetch: Retrieve full metadata
ESummary: Get summary information
ELink: Find related articles
See references/pubmed_search.md for complete API documentation.
# Single DOI
python scripts/extract_metadata.py --doi 10.1038/s41586-021-03819-2
# Single PMID
python scripts/extract_metadata.py --pmid 34265844
# Single arXiv ID
python scripts/extract_metadata.py --arxiv 2103.14030
# From URL
python scripts/extract_metadata.py \
--url "https://www.nature.com/articles/s41586-021-03819-2"# Batch processing (file with one identifier per line)
python scripts/extract_metadata.py \
--input paper_ids.txt \
--output references.bib
# Different output formats
python scripts/extract_metadata.py \
--doi 10.1038/nature12345 \
--format json # or bibtex, yaml
validate_citations.py
Validate BibTeX entries for accuracy, completeness, citation count standard compliance, and manuscript integration.
Features:
DOI verification via doi.org and CrossRef
Required field checking
Duplicate detection
Format validation
Publication standard citation count checks against specified venues (Nature, NeurIPS, review, etc.) or custom thresholds.
Mandatory post-writing checks matching manuscript citations (Markdown or LaTeX) with defined BibTeX entries to detect unresolved/missing or unused references.
Detailed reporting
Usage:
# Basic validation
python scripts/validate_citations.py references.bib
# Validate against a venue standard (e.g., Nature, NeurIPS, Literature Review)
python scripts/validate_citations.py references.bib --venue nature
python scripts/validate_citations.py references.bib --venue neurips
python scripts/validate_citations.py references.bib --venue review
# Validate with custom minimum citation count
python scripts/validate_citations.py references.bib --min-count 40
# Check references against a written manuscript file (detect missing or unused citations)
python scripts/validate_citations.py references.bib --manuscript paper.md
# Combined full validation
python scripts/validate_citations.py references.bib \
--venue nature \
--manuscript paper.md \
--report validation_report.json \
--verbose
Duplicate entries: Same paper cited multiple times with different keys
Solution: Use duplicate detection in validation
Missing required fields: Incomplete BibTeX entries (volume, pages, DOI missing)
Solution: Run Phase 2.5 metadata enrichment — web search for every missing field before proceeding. NEVER leave an @article entry without volume, pages, and DOI.
Outdated preprints: Citing preprint when published version exists
Solution: Check if preprints have been published, update to journal version
Special character issues: Broken LaTeX compilation due to characters
Solution: Use proper escaping or Unicode in BibTeX
No validation before submission: Submitting with citation errors
Solution: Always run validation as final check
Manual BibTeX entry: Typing entries by hand
Solution: Always extract from metadata sources using scripts
# You have a text file with DOIs (one per line)# dois.txt contains:# 10.1038/s41586-021-03819-2# 10.1126/science.aam9317# 10.1016/j.cell.2023.01.001# Convert all to BibTeX
python scripts/doi_to_bibtex.py --input dois.txt --output references.bib
# Validate the result
python scripts/validate_citations.py references.bib --verbose
Example 3: Cleaning an Existing BibTeX File
# You have a messy BibTeX file from various sources# Clean it up systematically# Step 1: Format and standardize
python scripts/format_bibtex.py messy_references.bib \
--output step1_formatted.bib
# Step 2: Remove duplicates
python scripts/format_bibtex.py step1_formatted.bib \
--deduplicate \
--output step2_deduplicated.bib
# Step 3: Validate and auto-fix
python scripts/validate_citations.py step2_deduplicated.bib \
--auto-fix \
--output step3_validated.bib
# Step 4: Sort by year
python scripts/format_bibtex.py step3_validated.bib \
--sort year \
--descending \
--output clean_references.bib
# Step 5: Final validation report
python scripts/validate_citations.py clean_references.bib \
--report final_validation.json \
--verbose
# Review reportcat final_validation.json
Example 4: Finding and Citing Seminal Papers
# Find highly cited papers on a topic
python scripts/search_google_scholar.py "AlphaFold protein structure" \
--year-start 2020 \
--year-end 2024 \
--sort-by citations \
--limit 20 \
--output alphafold_seminal.json
# Extract the top 10 by citation count# (script will have included citation counts in JSON)# Convert to BibTeX
python scripts/extract_metadata.py \
--input alphafold_seminal.json \
--output alphafold_refs.bib
# The BibTeX file now contains the most influential papers
Integration with Other Skills
Literature Review Skill
Citation Management provides the technical infrastructure for Literature Review:
Literature Review: Multi-database systematic search and synthesis
Citation Management: Metadata extraction and validation
Combined workflow:
Use literature-review for systematic search methodology
Use citation-management to extract and validate citations
Use literature-review to synthesize findings
Use citation-management to ensure bibliography accuracy
Scientific Writing Skill
Citation Management ensures accurate references for Scientific Writing:
Export validated BibTeX for use in LaTeX manuscripts
Verify citations match publication standards
Format references according to journal requirements
Venue Templates Skill
Citation Management works with Venue Templates for submission-ready manuscripts:
Different venues require different citation styles
Generate properly formatted references
Validate citations meet venue requirements
Resources
Bundled Resources
References (in references/):
google_scholar_search.md: Complete Google Scholar search guide
pubmed_search.md: PubMed and E-utilities API documentation
metadata_extraction.md: Metadata sources and field requirements
citation_validation.md: Validation criteria and quality checks
bibtex_formatting.md: BibTeX entry types and formatting rules
Scripts (in scripts/):
search_google_scholar.py: Google Scholar search automation
search_pubmed.py: PubMed E-utilities API client
extract_metadata.py: Universal metadata extractor
validate_citations.py: Citation validation and verification
format_bibtex.py: BibTeX formatter and cleaner
doi_to_bibtex.py: Quick DOI to BibTeX converter
Assets (in assets/):
bibtex_template.bib: Example BibTeX entries for all types