用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/majiayu000/claude-skill-registry --skill embedding-strategy命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
LLM token logprobs and calibration. Per-decision confidence, ECE, Brier, reliability diagrams, low-confidence triage.
Analyze LLM token logprobs and calibration. Use for per-decision confidence, ECE, Brier scores, reliability diagrams, and low-confidence triage.
回顾最近 N 天的 Claude Code 使用记录——扫描原始会话数据,按主题分组汇总"我都做了什么",并从个人操作系统视角输出模式、风险与增删建议。当用户说 /recap、"看看我这几天做了什么"、"回顾一下我最近的会话"、"这两天我用 claude 干了啥"、"活动回顾" 时使用。
基于 SOC 职业分类
正在显示 SKILL.md
| name | embedding-strategy |
| description | PROTECTED - Chunking strategy and embedding dimension management |
PROTECTED: Manage chunking strategies and embedding dimensions. Changes require user approval + full impact analysis.
from langchain.text_splitter import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(
chunk_size=512, # PROTECTED
chunk_overlap=50, # PROTECTED
separators=["\n\n", "\n", " ", ""],
length_function=len
)
chunks = splitter.split_text(document)
When to use: General purpose (paragraphs, sentences)
Chunk Size Guidelines:
from langchain.text_splitter import TokenTextSplitter
splitter = TokenTextSplitter(
chunk_size=512,
chunk_overlap=50
)
When to use: LLM context window management (token-based limits)
from langchain.text_splitter import MarkdownTextSplitter
splitter = MarkdownTextSplitter(
chunk_size=512,
chunk_overlap=50
)
When to use: Markdown documents (preserves structure)
Without overlap (BAD):
Chunk 1: "...end of sentence A."
Chunk 2: "Start of sentence B..."
→ Context lost between chunks
With overlap (GOOD):
Chunk 1: "...end of sentence A. Start of sentence B..."
Chunk 2: "...end of sentence A. Start of sentence B. More context..."
→ Continuity preserved
Rule of thumb: 10-20% of chunk_size
| Model | Dimension | Use Case |
|---|---|---|
| OpenAI text-embedding-ada-002 | 1536 | General purpose |
| OpenAI text-embedding-3-small | 1536 | Cost-effective |
| OpenAI text-embedding-3-large | 3072 | Highest quality |
| Cohere embed-english-v3.0 | 1024 | English docs |
| HuggingFace all-MiniLM-L6-v2 | 384 | Fast, local |
NEVER change dimension without:
See: templates/rag-checklist.md Q1
from ragas.metrics import ContextRelevance
# Test different chunk sizes
chunk_sizes = [256, 512, 1024]
results = {}
for size in chunk_sizes:
splitter = RecursiveCharacterTextSplitter(chunk_size=size)
# Re-index with new chunk size
# Run evaluation
score = evaluate_retrieval(splitter)
results[size] = score
# Choose best
best_size = max(results, key=results.get)
print(f"Best chunk size: {best_size}")
import re
def clean_text(text: str) -> str:
# Remove extra whitespace
text = re.sub(r'\s+', ' ', text)
# Remove special characters
text = re.sub(r'[^\w\s.,!?-]', '', text)
# Normalize line breaks
text = text.replace('\r\n', '\n')
return text.strip()
def remove_boilerplate(text: str) -> str:
# Remove headers/footers
text = re.sub(r'Page \d+ of \d+', '', text)
# Remove navigation
text = re.sub(r'Home \| About \| Contact', '', text)
return text
See: templates/rag-checklist.md for pre-modification checklist
Last Updated: 2025-12-04