Skip to main content
cohere-performance-tuning Optimize Cohere API performance with caching, batching, model selection, and streaming.
Use when experiencing slow API responses, implementing caching strategies,
or optimizing request throughput for Cohere Chat, Embed, and Rerank.
Trigger with phrases like "cohere performance", "optimize cohere",
"cohere latency", "cohere caching", "cohere slow", "cohere batch".
설치로 이동 Skills Marketplace 커뮤니티가 만든 AI 스킬을 발견하고 탐색하세요.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
직접 명령은 검토 Prompt를 거치지 않습니다. 실행하기 전에 소스를 확인하세요.
npx skills add https://github.com/jeremylongshore/claude-code-plugins-plus-skills --skill cohere-performance-tuning명령은 한 줄로 유지됩니다. 복사하기 전에 가로로 스크롤해 전체 내용을 확인하세요.
로컬 사본을 원하시나요? SkillsMP에서 현재 제공할 수 있는 파일을 다운로드하세요.
Zip 다운로드 다운로드 중... jeremylongshore
jeremylongshore/claude-code-plugins-plus-skills
GitHub 저장소 열기 이 저장소의 다른 Skills Implement user sign-up and sign-in flows with Clerk.
Use when building authentication UI, customizing sign-in experience,
or implementing OAuth social login.
Trigger with phrases like "clerk sign-in", "clerk sign-up",
"clerk login flow", "clerk OAuth", "clerk social login".
Implement session management and middleware with Clerk.
Use when managing user sessions, configuring route protection,
or implementing token refresh and custom JWT templates.
Trigger with phrases like "clerk session", "clerk middleware",
"clerk route protection", "clerk token", "clerk JWT".
Configure enterprise SSO, role-based access control, and organization management.
Use when implementing SSO integration, configuring role-based permissions,
or setting up organization-level controls.
Trigger with phrases like "clerk SSO", "clerk RBAC",
"clerk enterprise", "clerk roles", "clerk permissions", "clerk organizations".
name cohere-performance-tuning description Optimize Cohere API performance with caching, batching, model selection, and streaming.
Use when experiencing slow API responses, implementing caching strategies,
or optimizing request throughput for Cohere Chat, Embed, and Rerank.
Trigger with phrases like "cohere performance", "optimize cohere",
"cohere latency", "cohere caching", "cohere slow", "cohere batch".
allowed-tools Read, Write, Edit version 1.5.0 license MIT author Jeremy Longshore <jeremy@intentsolutions.io> tags ["saas","ai","nlp","cohere"] compatibility Designed for Claude Code
Cohere Performance Tuning
Overview
Optimize Cohere API v2 performance through model selection, embedding batches, rerank pipelines, caching, and streaming for time-to-first-token.
Prerequisites
cohere-ai SDK installed
Understanding of Cohere endpoints (Chat, Embed, Rerank)
Redis or in-memory cache (optional)
Latency Benchmarks (Typical)
Operation Model P50 P95 Chat (short) command-r7b-12-2024500ms 1.5s Chat (short) command-a-03-2025800ms 2.5s Chat (stream TTFT) command-a-03-2025200ms 600ms Embed (96 texts) embed-v4.0150ms 400ms Rerank (100 docs) rerank-v3.5100ms 300ms Classify (96 inputs) embed-english-v3.0200ms 500ms
Instructions
Strategy 1: Model Selection by Latency Budget
function selectModel (latencyBudgetMs : number ): string {
if (latencyBudgetMs < 1000 ) return 'command-r7b-12-2024' ;
if (latencyBudgetMs < 3000 ) return 'command-r-08-2024' ;
return 'command-a-03-2025' ;
}
cohere. ({
: ( ),
: [{ : , : query }],
: ,
});
await
chat
model
selectModel
1500
messages
role
'user'
content
maxTokens
200
Strategy 2: Streaming for Time-to-First-Token
async function streamForUI (message : string ): Promise <string > {
const stream = await cohere.chatStream ({
model : 'command-a-03-2025' ,
messages : [{ role : 'user' , content : message }],
});
let fullText = '' ;
for await (const event of stream) {
if (event.type === 'content-delta' ) {
const text = event.delta ?.message ?.content ?.text ?? '' ;
fullText += text;
}
}
return fullText;
}
Strategy 3: Batch Embeddings (96 per Call)
for (const text of texts) {
await cohere.embed ({ model : 'embed-v4.0' , texts : [text], ... });
}
async function batchEmbed (texts : string [] ): Promise <number [][]> {
const BATCH = 96 ;
const results : number [][] = [];
const batches = [];
for (let i = 0 ; i < texts.length ; i += BATCH ) {
batches.push (texts.slice (i, i + BATCH ));
}
const responses = await Promise .all (
batches.map (batch =>
cohere.embed ({
model : 'embed-v4.0' ,
texts : batch,
inputType : 'search_document' ,
embeddingTypes : ['float' ],
})
)
);
for (const resp of responses) {
results.push (...resp.embeddings .float );
}
return results;
}
Strategy 4: Compressed Embeddings
const response = await cohere.embed ({
model : 'embed-v4.0' ,
texts : documents,
inputType : 'search_document' ,
embeddingTypes : ['int8' ],
});
const storageVectors = response.embeddings .int8 ;
Strategy 5: Rerank as a Pre-filter
async function efficientSearch (query : string , corpus : string [] ) {
const reranked = await cohere.rerank ({
model : 'rerank-v3.5' ,
query,
documents : corpus,
topN : 5 ,
});
const topDocs = reranked.results .map (r => ({
text : corpus[r.index ],
score : r.relevanceScore ,
}));
return topDocs;
}
Strategy 6: Embedding Cache import { LRUCache } from 'lru-cache' ;
import crypto from 'crypto' ;
const embedCache = new LRUCache <string , number []>({
max : 10_000 ,
ttl : 24 * 60 * 60 * 1000 ,
});
function hashText (text : string ): string {
return crypto.createHash ('sha256' ).update (text).digest ('hex' ).slice (0 , 16 );
}
async function cachedEmbed (texts : string [] ): Promise <number [][]> {
const results : number [][] = new Array (texts.length );
const uncached : { index : number ; text : string }[] = [];
for (let i = 0 ; i < texts.length ; i++) {
const key = hashText (texts[i]);
const cached = embedCache.get (key);
if (cached) {
results[i] = cached;
} else {
uncached.push ({ index : i, text : texts[i] });
}
}
if (uncached.length > 0 ) {
const vectors = await batchEmbed (uncached.map (u => u.text ));
for (let j = 0 ; j < uncached.length ; j++) {
results[uncached[j].index ] = vectors[j];
embedCache.set (hashText (uncached[j].text ), vectors[j]);
}
}
return results;
}
Strategy 7: Response Caching for Chat import { LRUCache } from 'lru-cache' ;
const chatCache = new LRUCache <string , string >({
max : 1000 ,
ttl : 5 * 60 * 1000 ,
});
async function cachedChat (message : string , system ?: string ): Promise <string > {
const key = `${system ?? '' } :${message} ` ;
const cached = chatCache.get (key);
if (cached) return cached;
const response = await cohere.chat ({
model : 'command-a-03-2025' ,
messages : [
...(system ? [{ role : 'system' as const , content : system }] : []),
{ role : 'user' as const , content : message },
],
temperature : 0 ,
});
const text = response.message ?.content ?.[0 ]?.text ?? '' ;
chatCache.set (key, text);
return text;
}
Performance Monitoring async function timedCohereCall<T>(
endpoint : string ,
fn : () => Promise <T>
): Promise <T> {
const start = performance.now ();
try {
const result = await fn ();
const ms = performance.now () - start;
console .log (`[cohere] ${endpoint} : ${ms.toFixed(0 )} ms` );
return result;
} catch (err) {
const ms = performance.now () - start;
console .error (`[cohere] ${endpoint} FAILED: ${ms.toFixed(0 )} ms` , err);
throw err;
}
}
Output
Model selection by latency budget
Streaming for sub-200ms TTFT
Batch embedding (96x fewer API calls)
Compressed embeddings (75-97% storage savings)
Cache layer for deterministic queries
Rerank as fast pre-filter
Error Handling Issue Cause Solution Chat > 5s Long output + slow model Use streaming, reduce maxTokens Embed timeout Too many texts Batch to 96 per call Cache stale Long TTL Reduce TTL for volatile data High costs No caching Cache embeddings (deterministic)
Resources
Next Steps For cost optimization, see cohere-cost-tuning.