一键导入
performance-optimization
Route RAG performance work for latency, caching, indexing, filtering, batching, and query optimization.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Route RAG performance work for latency, caching, indexing, filtering, batching, and query optimization.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Use this skill when building, debugging, or improving Retrieval-Augmented Generation systems, including chunking, vector database selection, hybrid search, reranking, multimodal RAG, code documentation RAG, retrieval latency, and production RAG architecture.
Chunk nested documents into parent-child levels so retrieval can move from broad sections to fine-grained passages.
Use semantic boundaries and embedding similarity to chunk text for higher-relevance retrieval.
Route RAG chunking decisions across semantic, hierarchical, sliding-window, contextual-header, and framework-selection strategies.
Use overlapping windows to preserve context across chunk boundaries while controlling retrieval size.
Reduce retrieval latency with caching, batching, and index-level optimization.
| name | performance-optimization |
| title | Performance Optimization |
| description | Route RAG performance work for latency, caching, indexing, filtering, batching, and query optimization. |
| category | performance-optimization |
| tags | ["latency","performance","caching","rag"] |
| allowed-tools | ["Read","Grep","Glob"] |
Use this parent skill when the RAG system works functionally but is too slow, expensive, or unstable under expected traffic. Route to targeted latency and retrieval optimization guidance.
RAG latency can come from embedding calls, vector search, metadata filters, reranking, prompt assembly, or repeated work. Optimization requires profiling the full retrieval path before changing architecture.
Measure embedding, search, filtering, reranking, prompt assembly, and model latency separately.
Use caching, batching, payload indexes, top-k tuning, and reranker gating where profiling shows bottlenecks.
Confirm optimizations do not reduce recall, faithfulness, citation quality, or operational reliability.