Skip to main content

sciverse

Use when the user needs academic paper retrieval — searching scientific literature by author/year/journal, finding paper chunks for RAG-style citations, or expanding original text around a known paper offset. Provides six Sciverse tools (search_papers, semantic_search, list_catalog, list_paper_relations, read_content, get_resource) via the sciverse-mcp-server MCP server.

来源信息

仓库
opendatalab/Sciverse-Agent-Tools
最近来源活动
2026年9月11日 03:02
检测到的 SKILL.md 语言
英语
星标
119
分支
6

安装方式

默认使用会先检查来源的 Prompt;你也可以切换为直接命令,或下载本地副本。

检查来源文件

决定是否安装前,请先阅读 SKILL.md,以及 SkillsMP 当前展示的配套文件。

正在显示 SKILL.md

SKILL.md
来源说明 · 只读预览
name
sciverse
description
Use when the user needs academic paper retrieval — searching scientific literature by author/year/journal, finding paper chunks for RAG-style citations, or expanding original text around a known paper offset. Provides six Sciverse tools (search_papers, semantic_search, list_catalog, list_paper_relations, read_content, get_resource) via the sciverse-mcp-server MCP server.
# Sciverse — Academic Paper Retrieval Retrieval skill for the Sciverse open platform. Exposes six tools for working with scientific literature: field introspection, structured metadata search, semantic chunk retrieval for RAG, citation / reference pagination, character-range content reading, and figure / table image fetching. ## When to use Trigger this skill when the user's request involves any of: - Locating academic papers by structured criteria (authors, year, journal, subjects) - Grounding an answer in paper excerpts (RAG / citations) - Expanding the original text around a known doc_id (more text before/after a chunk) Do NOT use this skill for general web search, news, or non-scientific content — the underlying index only covers peer-reviewed and preprint scientific literature. ## Prerequisites This skill is a thin wrapper around the `sciverse-mcp-server` MCP server. Before invoking any tool, ensure the server is reachable: 1. Install the MCP server: ```bash npm install -g sciverse-mcp-server ``` Or add it to your project `.mcp.json`: ```json { "mcpServers": { "sciverse": { "command": "npx", "args": ["-y", "sciverse-mcp-server"], "env": { "SCIVERSE_API_TOKEN": "${SCIVERSE_API_TOKEN}" } } } } ``` 2. Obtain an API token from https://sciverse.space and export it: ```bash export SCIVERSE_API_TOKEN=sv-... ``` Optional: set `SCIVERSE_BASE_URL` to override the default API base URL (for dev / self-hosted gateways; must remain on `*.sciverse.space`). ## Tools All six tools are exposed by the MCP server. Claude Code will surface them automatically when this skill is active. ### search_papers Search academic papers by structured filters (title, authors, journal, year, subjects, etc.). Use when: "find Hinton's papers from 2020-2023", "Nature papers on CRISPR". Not for: natural-language Q&A retrieval (use semantic_search) or full-text snippets (use read_content). Returns: list of papers; each entry has unique_id (always present), doc_id (only when full text exists), title, author, abstract, publication_venue_name_unified, publication_published_year. ### semantic_search Natural-language semantic search returning relevant paper chunks for RAG-style answering. Use when: "How does Transformer attention work?", "What are recent methods for protein structure prediction?". Not for: precise field filtering (use search_papers) or fetching full original text (use read_content). Returns: list of chunks; each entry has chunk_id, doc_id, abstract, chunk, score, title, offset. Typical chain: semantic_search → pick chunk → read_content(doc_id, offset). ### list_catalog Returns the schema catalog for search_papers: every field name, type, whether it's filterable / sortable, default-return status, human description, and applicable FilterOperators. Use when: "Which field do I filter by DOI?", "What values can access_oa_status take?", "What's the right enum for metadata_type?". Not for: actually searching papers (use search_papers / semantic_search). Typical pattern: call once when first encountering Sciverse or facing an ambiguous field need, then construct precise search_papers filters from the returned schema. Pass include_sample_values=true to also fetch top-20 values for enum-like fields (OpenSearch terms aggregation, 24h cached). ### list_paper_relations Paginate the full relation list of a paper. citations/references/related_works are unbounded arrays (up to 340k entries for a single paper) and are NOT projectable in search_papers, so this endpoint is the only way to read them. Use when: "What does paper X cite?" (relation=REFERENCES), "Which papers cite paper X?" (relation=CITATIONS), "Works related to paper X" (relation=RELATED_WORKS). Note: CITATIONS (incoming: who cites me) and REFERENCES (outgoing: who I cite) are opposite directions. Typical chain: get unique_id from search_papers / semantic_search, then paginate here by relation. Two limits (CITATIONS only; REFERENCES/RELATED_WORKS max out at 11833/20 in practice): more than 10000 relations returns 429; page*page_size above 10000 returns 400. In both cases switch to search_papers with filters_advanced on references_unique_id — it supports deep paging and arbitrary sorting. total_count counts in-corpus matches only, so it can differ from the paper's own citation_count by about 1%. ### read_content Read a range of a paper's original text addressed in Unicode code points (offset/limit count characters like Python len(), not bytes). Typically used with a doc_id/offset returned by semantic_search to expand context (read more text before or after a chunk). Returns: text fragment, bytes_returned (UTF-8 byte length of text, for reference only), next_offset (code-point offset of the next fragment — page with it, never with bytes_returned), more (boolean). Server behaviour: limit above 524288 is silently clamped; omitting offset returns the whole document ignoring limit — the SDKs / MCP server send offset=0 and limit=4096 by default, so pass offset explicitly when calling the HTTP API directly. ### get_resource Returns the binary bytes of a paper figure / table image referenced inside read_content's Markdown via `![alt](file_name)` placeholders. Use when the user asks to see / display / describe a figure and read_content output contains an image reference. Input file_name comes from the Markdown URL part (relative path, no `\\` or `..`). Returns: raw image stream + image/* Content-Type. The SDK / MCP server wraps the bytes as base64 + mimeType so Claude (multimodal) can read the image directly. ## Bootstrap: learn the schema first If you don't yet know which fields exist or what values they take (e.g. "is `oa_status` a field?", "what does `metadata_type` accept?"), call `list_catalog` once at the start of the conversation. The result includes every field name, type, filterability, default-return status, and — for enum-like fields — sample values. Cache the catalog in your working memory; subsequent `search_papers` filters become precise instead of guessed. ``` list_catalog(include_sample_values=true) └─▶ fields[].name + sample_values → pick the right filter field ``` ## Recipes **1. Natural-language RAG (most common):** ``` semantic_search(query="How does Transformer attention work?", top_k=5) └─▶ for each hit: read_content(doc_id, offset, limit=8192) └─▶ cite doc_id + title in the answer ``` **2. Look up a paper by DOI / doc_id:** ``` search_papers(filters_advanced=[ {field: "doi", operator: "FILTER_OP_EQ", value: "10.1038/..."} ]) ``` **3. Find OA papers in a year range:** ``` search_papers( filters_advanced=[ {field: "access_is_oa", value: "true"}, {field: "access_oa_status", operator: "FILTER_OP_IN", value: ["gold", "green", "hybrid"]} ], year_from=2024 ) ``` **4. Filter by language / metadata_type (enum fields):** ``` # First check the enum: list_catalog(include_sample_values=true) # Then filter precisely: search_papers( query="transformer", filters_advanced=[ {field: "language", value: "en"}, {field: "metadata_type", value: "paper"} ] ) ``` **5. Scoped semantic search (constrained corpus):** ``` semantic_search( query="attention", filters={"author": ["Hinton"], "publication_published_year": {"gte": 2020}} ) # filters apply at recall time (server-side), AND across fields; # soft semantics: chunks missing that metadata are NOT excluded ``` For a hard guarantee, or meta-only constraints (fwci, citation graph, complex hit-sets), scope by doc_id — a HARD recall-time filter: ``` search_papers(..., fields=["doc_id","title"]) └─▶ collect hits[].doc_id (only fulltext papers carry one) semantic_search(query="attention", filters={"doc_id": [...]}) # hits never leave the set; empty list → empty hits (never global); # up to 1000 deduped ids (400 SCOPE_TOO_LARGE beyond) ``` **6. Bias fuzzy search ranking (soft boosts — stackable):** Three multiplicative boosts reorder fuzzy-search results while keeping relevance. Only effective when `query` is non-empty; ignored when any sort is set; shallow paging while active (no `next_cursor`). `sort_by_year` defaults to `auto`: relevance when `query` is set, newest-first for pure filters. Never combine `query` with `sort_by_year=desc` expecting "relevant and recent" — explicit sort degrades the query to a match filter (OR, no ranking) and disables all boosts; use `freshness_boost` instead. ``` search_papers(query="large language model", freshness_boost="STRONG") # recent first: STRONG=3-year decay, MILD=10-year search_papers(query="protein folding", impact_boost="MILD") # highly-cited float up (bounded; zero-citation stays neutral) search_papers(query="深度学习", language_affinity="MILD") # demote (never exclude) results not in the query's language; # target detected from query text (kana→ja/hangul→ko/Han→zh/Latin→en); # unknown-language papers stay neutral. Hard-exclude instead: # filters_advanced=[{"field":"language","value":"zh"}] search_papers(query="...", freshness_boost="MILD", impact_boost="MILD", language_affinity="MILD") # stack: relevant+recent+cited+same-lang ``` **7. Show a figure / image from the paper:** When `read_content` Markdown contains `![alt](file_name)` placeholders and the user wants to see the figure (e.g. "show me Figure 3"), fetch the binary with `get_resource`. The MCP server wraps the bytes as a base64 image content block so Claude can read it directly. ``` read_content(doc_id, offset) → markdown with ![Figure 3](dt=xxx/p_yyy/f3.png) └─▶ get_resource(file_name="dt=xxx/p_yyy/f3.png") └─▶ Claude sees the image inline ``` **8. Search authors or journals (collection):** Set `collection` to `authors` or `sources` (default `papers`) to search those entities instead of papers. Each collection has its own fields — call `list_catalog(collection="authors")` first. Use `filters_advanced` + `sort_advanced`; the papers convenience fields (`authors`/`year_from`/...) apply to papers only. ``` # Top authors by h-index, sorted by citations search_papers( collection="authors", filters_advanced=[{field: "summary_stats.h_index", operator: "FILTER_OP_GTE", value: 50}], sort_advanced=[{field: "cited_by_count", order: "SORT_ORDER_DESC"}] ) # Enrich a paper result: take an author orcid / venue issn, then look up the entity search_papers(collection="authors", filters_advanced=[{field: "orcid", value: "https://orcid.org/..."}]) ``` ## Notes for Claude - **Always cite** `doc_id` and `title` when surfacing paper-based facts. - **Prefer `semantic_search`** for natural-language questions; only fall back to `search_papers` when the user provides structured criteria. - **When stuck on a field name**: call `list_catalog` instead of guessing. Field name typos return 400 with a clear message, but waste a turn. - **Before reading a paper's fulltext**: check `is_content_accessible` on the `search_papers` hit — `true` means the paper has fulltext AND you're authorized, so `read_content(doc_id, ...)` will work; `false` means no fulltext or no permission. - **When a chunk looks promising but truncated**: `read_content(doc_id, offset)` to expand. `read_content` returns `more: true` when more text is available; offset/limit are Unicode code points (not bytes) — page with `next_offset`. - **Pagination**: `semantic_search` top_k allows up to 100, but `balanced` truncates to ~50 server-side — use `quality` (or `fast`) when you need more. At most ~3 chunks come back per paper, so a high top_k needs many distinct papers to be matched. `search_papers` returns max 50 per page; use `page`.
在 GitHub 查看