一键导入
evaluation-metrics
Route RAG evaluation work across retrieval metrics, generation quality, and end-to-end evaluation frameworks.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Route RAG evaluation work across retrieval metrics, generation quality, and end-to-end evaluation frameworks.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Use this skill when building, debugging, or improving Retrieval-Augmented Generation systems, including chunking, vector database selection, hybrid search, reranking, multimodal RAG, code documentation RAG, retrieval latency, and production RAG architecture.
Chunk nested documents into parent-child levels so retrieval can move from broad sections to fine-grained passages.
Use semantic boundaries and embedding similarity to chunk text for higher-relevance retrieval.
Route RAG chunking decisions across semantic, hierarchical, sliding-window, contextual-header, and framework-selection strategies.
Use overlapping windows to preserve context across chunk boundaries while controlling retrieval size.
Reduce retrieval latency with caching, batching, and index-level optimization.
| name | evaluation-metrics |
| title | Evaluation Metrics |
| description | Route RAG evaluation work across retrieval metrics, generation quality, and end-to-end evaluation frameworks. |
| category | evaluation-metrics |
| tags | ["evaluation","metrics","ragas","routing"] |
| allowed-tools | ["Read","Grep","Glob"] |
Use this parent skill when the main RAG problem is measuring quality, detecting regressions, or proving that a change actually helped. Route to the child skill that matches whether the failure is in retrieval, in generation, or in the evaluation harness itself.
Teams ship RAG changes with no baseline, tune on vibes, and cannot tell whether a new chunker, embedding model, or prompt made things better or worse. Without metrics, retrieval and generation regressions ship silently and are only caught by users.
Measure retrieval and generation separately. Bad answers from good context are a generation problem; good answers are impossible from bad context.
Use retrieval metrics when the right chunks are missing or mis-ranked, and generation metrics when context is good but answers are wrong, unsupported, or off-topic.
Build a golden set, run the metrics on every change, and block merges that regress the baseline instead of evaluating manually.