一键导入
embedding
嵌入模型选型与优化——API vs 本地模型、批量嵌入、Redis 缓存、中文优化
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
嵌入模型选型与优化——API vs 本地模型、批量嵌入、Redis 缓存、中文优化
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Execute the check workflow from this AI development team preset. Use when the user writes /check, asks for check, or wants the corresponding team process in Codex.
Execute the dev workflow from this AI development team preset. Use when the user writes /dev, asks for dev, or wants the corresponding team process in Codex.
Execute the fix workflow from this AI development team preset. Use when the user writes /fix, asks for fix, or wants the corresponding team process in Codex.
Execute the plan workflow from this AI development team preset. Use when the user writes /plan, asks for plan, or wants the corresponding team process in Codex.
Execute the project-preset workflow from this AI development team preset. Use when the user writes /project-preset, asks for project-preset, or wants the corresponding team process in Codex.
Execute the review-all workflow from this AI development team preset. Use when the user writes /review-all, asks for review-all, or wants the corresponding team process in Codex.
| name | embedding |
| description | 嵌入模型选型与优化——API vs 本地模型、批量嵌入、Redis 缓存、中文优化 |
选择合适的嵌入模型,优化嵌入成本和质量。
模型选型表与成本优化策略与语言无关,下方代码为 Python。 TypeScript:用
ai包的embedMany(openai.embedding('text-embedding-3-small')),见lang/typescript/specs/typescript.md的"嵌入生成"。
| 模型 | 维度 | 上下文 | 场景 | 成本 |
|---|---|---|---|---|
| text-embedding-3-small | 1536 | 8191 tokens | 通用,推荐首选 | 低 |
| text-embedding-3-large | 3072 | 8191 tokens | 高精度场景 | 高 |
| BGE-M3(本地) | 1024 | 8192 tokens | 中英文,无成本 | 0(本地算力) |
| nomic-embed-text(Ollama) | 768 | 8192 tokens | 离线场景 | 0 |
默认选择:text-embedding-3-small(性价比最高)
中文内容较多时:优先考虑 BGE-M3(BAAI 出品,专门优化中文)
import asyncio
from openai import AsyncOpenAI
class OpenAIEmbedder:
def __init__(self, model: str = "text-embedding-3-small") -> None:
self._client = AsyncOpenAI()
self.model = model
self.BATCH_SIZE = 100 # OpenAI 限制单次最多 2048 个文本
async def embed(self, texts: list[str]) -> list[list[float]]:
# 过滤空文本
texts = [t[:8000] for t in texts if t.strip()]
# 并行批次请求
batches = [texts[i:i+self.BATCH_SIZE] for i in range(0, len(texts), self.BATCH_SIZE)]
results = await asyncio.gather(*[self._embed_batch(b) for b in batches])
return [emb for batch in results for emb in batch]
async def _embed_batch(self, texts: list[str]) -> list[list[float]]:
response = await self._client.embeddings.create(
input=texts,
model=self.model,
)
return [e.embedding for e in response.data]
async def embed_one(self, text: str) -> list[float]:
result = await self.embed([text])
return result[0]
import hashlib
import json
class CachedEmbedder:
"""Redis 缓存嵌入,相同文本不重复请求"""
async def embed(self, texts: list[str]) -> list[list[float]]:
results = []
to_embed = []
cache_keys = []
for text in texts:
key = f"emb:{hashlib.md5(text.encode()).hexdigest()}"
cached = await self.redis.get(key)
if cached:
results.append((len(results), json.loads(cached)))
else:
to_embed.append(text)
cache_keys.append((len(results) + len(to_embed) - 1, key))
results.append(None)
if to_embed:
new_embeddings = await self.embedder.embed(to_embed)
for (idx, key), emb in zip(cache_keys, new_embeddings):
results[idx] = emb
await self.redis.setex(key, 86400 * 7, json.dumps(emb)) # 7天缓存
return results
import httpx
class OllamaEmbedder:
def __init__(self, model: str = "bge-m3", base_url: str = "http://localhost:11434") -> None:
self.model = model
self.base_url = base_url
async def embed(self, texts: list[str]) -> list[list[float]]:
async with httpx.AsyncClient(timeout=30) as client:
tasks = [
client.post(f"{self.base_url}/api/embeddings",
json={"model": self.model, "prompt": t})
for t in texts
]
responses = await asyncio.gather(*tasks)
return [r.json()["embedding"] for r in responses]