用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/LigphiDonk/Oh-my--paper --skill dataset-discovery命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
Search traceable academic papers, download legally accessible PDFs from arXiv and open-access sources, convert PDFs or page images to Markdown with a PaddleOCR layout-parsing API (or local pdfminer fallback), and organize the results into an AI-readable literature library. Use when Claude Code needs to build a paper corpus, batch OCR PDFs to Markdown, ingest real literature into a knowledge base, fetch arXiv or Hugging Face paper leads, or turn a directory of papers into structured Markdown plus metadata.
Delegate complex coding tasks to Claude Code CLI
Delegate coding tasks to OpenAI Codex CLI
基于 SOC 职业分类
正在显示 SKILL.md
| id | dataset-discovery |
| name | dataset-discovery |
| version | 1.0.0 |
| description | Multi-source ML dataset discovery. |
| stages | ["survey","ideation","experiment"] |
| tools | ["read_file","search_project","write_file","run_terminal"] |
| summary | Multi-source ML dataset discovery. Search HuggingFace Hub, OpenML, GitHub, and paper cross-references for datasets relevant to a research task. Use when asked to "find datasets for", "search ML datasets", "what datasets exist for", or "dis... |
| primaryIntent | data |
| intents | ["data","research"] |
| capabilities | ["search-retrieval","data-processing"] |
| domains | ["data-engineering"] |
| keywords | ["dataset-discovery","resource prep","search-retrieval","data-processing","data-engineering","dataset","discovery","multi","source","ml","search","huggingface"] |
| source | builtin |
| status | verified |
| upstream | {"repo":"dr-claw","path":"skills/dataset-discovery","revision":"8322dc4ef575affaa374aa7922c0a0971c6db7d7"} |
| resourceFlags | {"hasReferences":false,"hasScripts":true,"hasTemplates":false,"hasAssets":false,"referenceCount":0,"scriptCount":1,"templateCount":0,"assetCount":0,"optionalScripts":true} |
Multi-source ML dataset discovery. Search HuggingFace Hub, OpenML, GitHub, and paper cross-references for datasets relevant to a research task. Use when asked to "find datasets for", "search ML datasets", "what datasets exist for", or "dis...
Use this skill when the user request matches its research workflow scope. Prefer the bundled resources instead of recreating templates or reference material. Keep outputs traceable to project files, citations, scripts, or upstream evidence.
scripts/ as optional helpers. Run them only when their dependencies are available, keep outputs in the project workspace, and explain a manual fallback if execution is blocked.Search multiple ML dataset sources (HuggingFace Hub, OpenML, GitHub, Semantic Scholar) and return a ranked, deduplicated list of relevant datasets.
Clarify the user's needs before searching:
Run the search script with the user's query:
python3 scripts/search_ml_datasets.py search --query "<query>" --sources huggingface,openml,github,papers --max 30
Options:
--sources: Comma-separated list from huggingface, openml, github, papers. Default: all four.--max: Maximum results to return after dedup + ranking. Default: 30.--modalityimagetexttabularaudio--workspace: Output directory. Default: ./datasets/discovery/Optionally also call HF MCP tool hub_repo_search with repo_types: ["dataset"] for semantic search to supplement results.
Show results as a markdown table:
| Name | Source | Downloads | Size | License | Tags | URL |
|---|
Sort by relevance score (highest first).
When the user wants more info on a specific dataset:
python3 scripts/search_ml_datasets.py detail --dataset-id "huggingface:stanfordnlp/imdb" --workspace ./datasets/discovery/
Writes metadata.json and README.md to {workspace}/datasets/{source}_{slug}/.
When the user wants to preview data:
python3 scripts/search_ml_datasets.py pull --dataset-id "huggingface:stanfordnlp/imdb" --sample-rows 20 --workspace ./datasets/discovery/
Writes sample.jsonl to {workspace}/datasets/{source}_{slug}/.
For full dataset download, confirm with the user first, then use huggingface-cli download or equivalent.
{workspace}/ # default: ./datasets/discovery/
search-{YYYY-MM-DD}.json # search results log
datasets/
{source}_{slug}/
metadata.json # detailed metadata
README.md # human-readable summary
sample.jsonl # sample rows
requests (stdlib-adjacent, universally available)gh CLI (for GitHub source only)