Skip to main content 首页 创作者 comeonoliver skillshub cohere-observability
cohere-observability Set up comprehensive observability for Cohere API v2 with metrics, traces, and alerts.
Use when implementing monitoring for Chat/Embed/Rerank operations,
setting up dashboards, or configuring alerts for Cohere integrations.
Trigger with phrases like "cohere monitoring", "cohere metrics",
"cohere observability", "monitor cohere", "cohere alerts", "cohere tracing".
跳到安装 Skills Marketplace 发现并探索由社区构建的 Agent Skills
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/ComeOnOliver/skillshub --skill cohere-observability命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
下载 Zip 下载中... 同仓库更多 Skills Review product and feature risk before an AI coding agent starts implementation.
Use Xquik for X data and confirmation-gated X actions: tweet search, user lookup, follower export, media download, monitors, webhooks, MCP, and SDK workflows.
Canton Network open-source ecosystem guide covering DAML SDK, Canton runtime, and Splice applications. Use when working with Canton Network, DAML smart contracts, or building decentralized applications.
name cohere-observability description Set up comprehensive observability for Cohere API v2 with metrics, traces, and alerts.
Use when implementing monitoring for Chat/Embed/Rerank operations,
setting up dashboards, or configuring alerts for Cohere integrations.
Trigger with phrases like "cohere monitoring", "cohere metrics",
"cohere observability", "monitor cohere", "cohere alerts", "cohere tracing".
allowed-tools Read, Write, Edit version 1.0.0 license MIT author Jeremy Longshore <jeremy@intentsolutions.io> tags ["saas","ai","nlp","cohere"] compatible-with claude-code
Cohere Observability
Overview
Set up production observability for Cohere API v2 with Prometheus metrics, OpenTelemetry tracing, and AlertManager rules. Tracks per-endpoint latency, token usage, error rates, and costs.
Prerequisites
Prometheus or compatible metrics backend
OpenTelemetry SDK installed
cohere-ai SDK v7+
Instructions
Step 1: Metrics Collection
import { Registry , Counter , Histogram , Gauge } from 'prom-client' ;
const registry = new Registry ();
const requestCounter = new Counter ({
name : 'cohere_requests_total' ,
help : 'Total Cohere API requests' ,
labelNames : ['endpoint' , 'model' , 'status' ],
registers : [registry],
});
const requestDuration = new Histogram ({
name : 'cohere_request_duration_seconds' ,
help : 'Cohere request duration' ,
labelNames : ['endpoint' , 'model' ],
buckets : [0.1 , 0.25 , 0.5 , 1 , 2.5 , 5 , 10 , ],
: [registry],
});
tokenCounter = ({
: ,
: ,
: [ , , ],
: [registry],
});
errorCounter = ({
: ,
: ,
: [ , ],
: [registry],
});
rateLimitGauge = ({
: ,
: ,
: [ ],
: [registry],
});
30
registers
const
new
Counter
name
'cohere_tokens_total'
help
'Total tokens consumed'
labelNames
'endpoint'
'model'
'direction'
registers
const
new
Counter
name
'cohere_errors_total'
help
'Cohere errors by status code'
labelNames
'endpoint'
'status_code'
registers
const
new
Gauge
name
'cohere_rate_limit_remaining'
help
'Remaining rate limit capacity'
labelNames
'endpoint'
registers
Step 2: Instrumented Client Wrapper import { CohereClientV2 , CohereError , CohereTimeoutError } from 'cohere-ai' ;
const cohere = new CohereClientV2 ();
async function instrumentedCall<T>(
endpoint : string ,
model : string ,
operation : () => Promise <T>
): Promise <T> {
const timer = requestDuration.startTimer ({ endpoint, model });
try {
const result = await operation ();
requestCounter.inc ({ endpoint, model, status : 'success' });
timer ();
const usage = (result as any )?.usage ?.billedUnits ;
if (usage) {
if (usage.inputTokens ) {
tokenCounter.inc ({ endpoint, model, direction : 'input' }, usage.inputTokens );
}
if (usage.outputTokens ) {
tokenCounter.inc ({ endpoint, model, direction : 'output' }, usage.outputTokens );
}
}
return result;
} catch (err) {
requestCounter.inc ({ endpoint, model, status : 'error' });
timer ();
if (err instanceof CohereError ) {
errorCounter.inc ({ endpoint, status_code : String (err.statusCode ) });
} else if (err instanceof CohereTimeoutError ) {
errorCounter.inc ({ endpoint, status_code : 'timeout' });
}
throw err;
}
}
const response = await instrumentedCall ('chat' , 'command-a-03-2025' , () =>
cohere.chat ({
model : 'command-a-03-2025' ,
messages : [{ role : 'user' , content : query }],
})
);
Step 3: OpenTelemetry Tracing import { trace, SpanStatusCode , SpanKind } from '@opentelemetry/api' ;
const tracer = trace.getTracer ('cohere-client' , '1.0.0' );
async function tracedCohereCall<T>(
endpoint : string ,
model : string ,
operation : () => Promise <T>
): Promise <T> {
return tracer.startActiveSpan (
`cohere.${endpoint} ` ,
{ kind : SpanKind .CLIENT },
async (span) => {
span.setAttribute ('cohere.model' , model);
span.setAttribute ('cohere.endpoint' , endpoint);
try {
const result = await operation ();
const usage = (result as any )?.usage ?.billedUnits ;
if (usage) {
span.setAttribute ('cohere.tokens.input' , usage.inputTokens ?? 0 );
span.setAttribute ('cohere.tokens.output' , usage.outputTokens ?? 0 );
}
span.setStatus ({ code : SpanStatusCode .OK });
return result;
} catch (err : any ) {
span.setStatus ({ code : SpanStatusCode .ERROR , message : err.message });
span.recordException (err);
if (err instanceof CohereError ) {
span.setAttribute ('cohere.error.status' , err.statusCode ?? 0 );
}
throw err;
} finally {
span.end ();
}
}
);
}
Step 4: Structured Logging import pino from 'pino' ;
const logger = pino ({ name : 'cohere' , level : process.env .LOG_LEVEL ?? 'info' });
function logCohereCall (
endpoint : string ,
model : string ,
durationMs : number ,
status : 'success' | 'error' ,
meta ?: Record <string , unknown >
) {
logger[status === 'error' ? 'error' : 'info' ]({
service : 'cohere' ,
endpoint,
model,
durationMs,
status,
...meta,
});
}
async function observedCall<T>(
endpoint : string ,
model : string ,
fn : () => Promise <T>
): Promise <T> {
return tracedCohereCall (endpoint, model, () =>
instrumentedCall (endpoint, model, async () => {
const start = Date .now ();
try {
const result = await fn ();
logCohereCall (endpoint, model, Date .now () - start, 'success' , {
tokens : (result as any )?.usage ?.billedUnits ,
});
return result;
} catch (err) {
logCohereCall (endpoint, model, Date .now () - start, 'error' , {
error : err instanceof CohereError ? err.statusCode : 'timeout' ,
});
throw err;
}
})
);
}
Step 5: Alert Rules
groups:
- name: cohere
rules:
- alert: CohereHighErrorRate
expr: |
rate(cohere_errors_total[5m]) /
rate(cohere_requests_total[5m]) > 0.05
for: 5m
labels:
severity: warning
annotations:
summary: "Cohere error rate > 5%"
description: "{{ $labels.endpoint }} error rate: {{ $value | humanizePercentage }} "
- alert: CohereRateLimited
expr: rate(cohere_errors_total{status_code="429"}[5m]) > 0.1
for: 2m
labels:
severity: warning
annotations:
summary: "Cohere rate limiting detected"
- alert: CohereHighLatency
expr: |
histogram_quantile(0.95,
rate(cohere_request_duration_seconds_bucket[5m])
) > 10
for: 5m
labels:
severity: warning
annotations:
summary: "Cohere P95 latency > 10s"
- alert: CohereAuthFailure
expr: cohere_errors_total{status_code="401"} > 0
for: 1m
labels:
severity: critical
annotations:
summary: "Cohere authentication failure — check API key"
- alert: CohereHighTokenBurn
expr: rate(cohere_tokens_total[1h]) > 100000
for: 15m
labels:
severity: warning
annotations:
summary: "Cohere token burn rate > 100K/hour"
Step 6: Metrics Endpoint
import express from 'express' ;
const app = express ();
app.get ('/metrics' , async (req, res) => {
res.set ('Content-Type' , registry.contentType );
res.send (await registry.metrics ());
});
Dashboard Panels (Grafana) Panel Query Type Request Rate rate(cohere_requests_total[5m])Time series Error Rate rate(cohere_errors_total[5m]) / rate(cohere_requests_total[5m])Stat P50/P95 Latency histogram_quantile(0.95, rate(cohere_request_duration_seconds_bucket[5m]))Time series Token Usage rate(cohere_tokens_total[1h])Bar chart Errors by Code sum by (status_code)(rate(cohere_errors_total[5m]))Pie chart
Output
Prometheus metrics for requests, latency, tokens, and errors
OpenTelemetry traces with Cohere-specific attributes
Structured JSON logging with pino
AlertManager rules for error rate, latency, auth, and cost
Error Handling Issue Cause Solution Missing token metrics Usage not in response Check response.usage.billedUnits High cardinality Too many model labels Use model family, not exact version Alert storm Threshold too low Tune thresholds for your traffic Trace gaps Missing context propagation Ensure OTel context flows through async
Resources
Next Steps For incident response, see cohere-incident-runbook.