Skip to main content

sciverse

Use when the user needs academic paper retrieval — searching scientific literature by author/year/journal, finding paper chunks for RAG-style citations, or expanding original text around a known paper offset. Provides six Sciverse tools (search_papers, semantic_search, list_catalog, list_paper_relations, read_content, get_resource) via the sciverse-mcp-server MCP server.

Zur Installation springen

Quellinformationen

Repository
opendatalab/Sciverse-Agent-Tools
Letzte Quellaktivität
11. September 2026 um 03:02
Erkannte Sprache von SKILL.md
Englisch
Sterne
110
Forks
5

Installationsoptionen

Standardmäßig ist der Prompt ausgewählt, der zuerst die Quelle prüft. Sie können zu einem direkten Befehl wechseln oder eine lokale Kopie herunterladen.

Quelldateien prüfen

Lesen Sie SKILL.md und alle von SkillsMP angezeigten Begleitdateien, bevor Sie sich für eine Installation entscheiden.

SKILL.md wird angezeigt

SKILL.md
Quellanweisungen · Schreibgeschützte Vorschau
name
sciverse
description
Use when the user needs academic paper retrieval — searching scientific literature by author/year/journal, finding paper chunks for RAG-style citations, or expanding original text around a known paper offset. Provides six Sciverse tools (search_papers, semantic_search, list_catalog, list_paper_relations, read_content, get_resource) via the sciverse-mcp-server MCP server.
# Sciverse — Academic Paper Retrieval Retrieval skill for the Sciverse open platform. Exposes six tools for working with scientific literature: field introspection, structured metadata search, semantic chunk retrieval for RAG, citation / reference pagination, character-range content reading, and figure / table image fetching. ## When to use Trigger this skill when the user's request involves any of: - Locating academic papers by structured criteria (authors, year, journal, subjects) - Grounding an answer in paper excerpts (RAG / citations) - Expanding the original text around a known doc_id (more text before/after a chunk) Do NOT use this skill for general web search, news, or non-scientific content — the underlying index only covers peer-reviewed and preprint scientific literature. ## Prerequisites This skill is a thin wrapper around the `sciverse-mcp-server` MCP server. Before invoking any tool, ensure the server is reachable: 1. Install the MCP server: ```bash npm install -g sciverse-mcp-server ``` Or add it to your project `.mcp.json`: ```json { "mcpServers": { "sciverse": { "command": "npx", "args": ["-y", "sciverse-mcp-server"], "env": { "SCIVERSE_API_TOKEN": "${SCIVERSE_API_TOKEN}" } } } } ``` 2. Obtain an API token from https://sciverse.space and export it: ```bash export SCIVERSE_API_TOKEN=sv-... ``` Optional: set `SCIVERSE_BASE_URL` to override the default API base URL (for dev / self-hosted gateways; must remain on `*.sciverse.space`). ## Tools All six tools are exposed by the MCP server. Claude Code will surface them automatically when this skill is active. ### search_papers Search academic papers by structured filters (title, authors, journal, year, subjects, etc.). Use when: "find Hinton's papers from 2020-2023", "Nature papers on CRISPR". Not for: natural-language Q&A retrieval (use semantic_search) or full-text snippets (use read_content). Returns: list of papers; each entry has unique_id (always present), doc_id (only when full text exists), title, author, abstract, publication_venue_name_unified, publication_published_year. ### semantic_search Natural-language semantic search returning relevant paper chunks for RAG-style answering. Use when: "How does Transformer attention work?", "What are recent methods for protein structure prediction?". Not for: precise field filtering (use search_papers) or fetching full original text (use read_content). Returns: list of chunks; each entry has chunk_id, doc_id, abstract, chunk, score, title, offset. Typical chain: semantic_search → pick chunk → read_content(doc_id, offset). ### list_catalog Returns the schema catalog for search_papers: every field name, type, whether it's filterable / sortable, default-return status, human description, and applicable FilterOperators. Use when: "Which field do I filter by DOI?", "What values can access_oa_status take?", "What's the right enum for metadata_type?". Not for: actually searching papers (use search_papers / semantic_search). Typical pattern: call once when first encountering Sciverse or facing an ambiguous field need, then construct precise search_papers filters from the returned schema. Pass include_sample_values=true to also fetch top-20 values for enum-like fields (OpenSearch terms aggregation, 24h cached). ### list_paper_relations Paginate the full relation list of a paper. citations/references/related_works are unbounded arrays (up to 340k entries for a single paper) and are NOT projectable in search_papers, so this endpoint is the only way to read them. Use when: "What does paper X cite?" (relation=REFERENCES), "Which papers cite paper X?" (relation=CITATIONS), "Works related to paper X" (relation=RELATED_WORKS). Note: CITATIONS (incoming: who cites me) and REFERENCES (outgoing: who I cite) are opposite directions. Typical chain: get unique_id from search_papers / semantic_search, then paginate here by relation. Two limits (CITATIONS only; REFERENCES/RELATED_WORKS max out at 11833/20 in practice): more than 10000 relations returns 429; page*page_size above 10000 returns 400. In both cases switch to search_papers with filters_advanced on references_unique_id — it supports deep paging and arbitrary sorting. total_count counts in-corpus matches only, so it can differ from the paper's own citation_count by about 1%. ### read_content Read a range of a paper's original text addressed in Unicode code points (offset/limit count characters like Python len(), not bytes). Typically used with a doc_id/offset returned by semantic_search to expand context (read more text before or after a chunk). Returns: text fragment, bytes_returned (UTF-8 byte length of text, for reference only), next_offset (code-point offset of the next fragment — page with it, never with bytes_returned), more (boolean). Server behaviour: limit above 524288 is silently clamped; omitting offset returns the whole document ignoring limit — the SDKs / MCP server send offset=0 and limit=4096 by default, so pass offset explicitly when calling the HTTP API directly. ### get_resource Returns the binary bytes of a paper figure / table image referenced inside read_content's Markdown via `![alt](file_name)` placeholders. Use when the user asks to see / display / describe a figure and read_content output contains an image reference. Input file_name comes from the Markdown URL part (relative path, no `\\` or `..`). Returns: raw image stream + image/* Content-Type. The SDK / MCP server wraps the bytes as base64 + mimeType so Claude (multimodal) can read the image directly. ## Bootstrap: learn the schema first If you don't yet know which fields exist or what values they take (e.g. "is `oa_status` a field?", "what does `metadata_type` accept?"), call `list_catalog` once at the start of the conversation. The result includes every field name, type, filterability, default-return status, and — for enum-like fields — sample values. Cache the catalog in your working memory; subsequent `search_papers` filters become precise instead of guessed. ``` list_catalog(include_sample_values=true) └─▶ fields[].name + sample_values → pick the right filter field ``` ## Recipes **1. Natural-language RAG (most common):** ``` semantic_search(query="How does Transformer attention work?", top_k=5) └─▶ for each hit: read_content(doc_id, offset, limit=8192) └─▶ cite doc_id + title in the answer ``` **2. Look up a paper by DOI / doc_id:** ``` search_papers(filters_advanced=[ {field: "doi", operator: "FILTER_OP_EQ", value: "10.1038/..."} ]) ``` **3. Find OA papers in a year range:** ``` search_papers( filters_advanced=[ {field: "access_is_oa", value: "true"}, {field: "access_oa_status", operator: "FILTER_OP_IN", value: ["gold", "green", "hybrid"]} ], year_from=2024 ) ``` **4. Filter by language / metadata_type (enum fields):** ``` # First check the enum: list_catalog(include_sample_values=true) # Then filter precisely: search_papers( query="transformer", filters_advanced=[ {field: "language", value: "en"}, {field: "metadata_type", value: "paper"} ] ) ``` **5. Scoped semantic search (constrained corpus):** ``` semantic_search( query="attention", filters={"author": ["Hinton"], "publication_published_year": {"gte": 2020}} ) # filters apply at recall time (server-side), AND across fields; # soft semantics: chunks missing that metadata are NOT excluded ``` For a hard guarantee, or meta-only constraints (fwci, citation graph, complex hit-sets), scope by doc_id — a HARD recall-time filter: ``` search_papers(..., fields=["doc_id","title"]) └─▶ collect hits[].doc_id (only fulltext papers carry one) semantic_search(query="attention", filters={"doc_id": [...]}) # hits never leave the set; empty list → empty hits (never global); # up to 1000 deduped ids (400 SCOPE_TOO_LARGE beyond) ``` **6. Bias fuzzy search ranking (soft boosts — stackable):** Three multiplicative boosts reorder fuzzy-search results while keeping relevance. Only effective when `query` is non-empty; ignored when any sort is set; shallow paging while active (no `next_cursor`). `sort_by_year` defaults to `auto`: relevance when `query` is set, newest-first for pure filters. Never combine `query` with `sort_by_year=desc` expecting "relevant and recent" — explicit sort degrades the query to a match filter (OR, no ranking) and disables all boosts; use `freshness_boost` instead. ``` search_papers(query="large language model", freshness_boost="STRONG") # recent first: STRONG=3-year decay, MILD=10-year search_papers(query="protein folding", impact_boost="MILD") # highly-cited float up (bounded; zero-citation stays neutral) search_papers(query="深度学习", language_affinity="MILD") # demote (never exclude) results not in the query's language; # target detected from query text (kana→ja/hangul→ko/Han→zh/Latin→en); # unknown-language papers stay neutral. Hard-exclude instead: # filters_advanced=[{"field":"language","value":"zh"}] search_papers(query="...", freshness_boost="MILD", impact_boost="MILD", language_affinity="MILD") # stack: relevant+recent+cited+same-lang ``` **7. Show a figure / image from the paper:** When `read_content` Markdown contains `![alt](file_name)` placeholders and the user wants to see the figure (e.g. "show me Figure 3"), fetch the binary with `get_resource`. The MCP server wraps the bytes as a base64 image content block so Claude can read it directly. ``` read_content(doc_id, offset) → markdown with ![Figure 3](dt=xxx/p_yyy/f3.png) └─▶ get_resource(file_name="dt=xxx/p_yyy/f3.png") └─▶ Claude sees the image inline ``` **8. Search authors or journals (collection):** Set `collection` to `authors` or `sources` (default `papers`) to search those entities instead of papers. Each collection has its own fields — call `list_catalog(collection="authors")` first. Use `filters_advanced` + `sort_advanced`; the papers convenience fields (`authors`/`year_from`/...) apply to papers only. ``` # Top authors by h-index, sorted by citations search_papers( collection="authors", filters_advanced=[{field: "summary_stats.h_index", operator: "FILTER_OP_GTE", value: 50}], sort_advanced=[{field: "cited_by_count", order: "SORT_ORDER_DESC"}] ) # Enrich a paper result: take an author orcid / venue issn, then look up the entity search_papers(collection="authors", filters_advanced=[{field: "orcid", value: "https://orcid.org/..."}]) ``` ## Notes for Claude - **Always cite** `doc_id` and `title` when surfacing paper-based facts. - **Prefer `semantic_search`** for natural-language questions; only fall back to `search_papers` when the user provides structured criteria. - **When stuck on a field name**: call `list_catalog` instead of guessing. Field name typos return 400 with a clear message, but waste a turn. - **Before reading a paper's fulltext**: check `is_content_accessible` on the `search_papers` hit — `true` means the paper has fulltext AND you're authorized, so `read_content(doc_id, ...)` will work; `false` means no fulltext or no permission. - **When a chunk looks promising but truncated**: `read_content(doc_id, offset)` to expand. `read_content` returns `more: true` when more text is available; offset/limit are Unicode code points (not bytes) — page with `next_offset`. - **Pagination**: `semantic_search` top_k allows up to 100, but `balanced` truncates to ~50 server-side — use `quality` (or `fast`) when you need more. At most ~3 chunks come back per paper, so a high top_k needs many distinct papers to be matched. `search_papers` returns max 50 per page; use `page`.
Auf GitHub ansehen