Set up observability for Groq integrations: latency histograms, token throughput,
rate limit gauges, cost tracking, and Prometheus alerts.
Use when instrumenting Groq API calls, building a metrics dashboard, or wiring latency/cost/rate-limit alerts.
Trigger with phrases like "groq monitoring", "groq metrics",
"groq observability", "monitor groq", "groq alerts", "groq dashboard".
Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.
A direct command skips the review prompt. Inspect the source before running it.
Set up observability for Groq integrations: latency histograms, token throughput,
rate limit gauges, cost tracking, and Prometheus alerts.
Use when instrumenting Groq API calls, building a metrics dashboard, or wiring latency/cost/rate-limit alerts.
Trigger with phrases like "groq monitoring", "groq metrics",
"groq observability", "monitor groq", "groq alerts", "groq dashboard".
Designed for Claude Code, also compatible with Codex and OpenClaw
Groq Observability
Overview
Monitor Groq LPU inference for latency, token throughput, rate limit utilization, and cost. Groq's defining advantage is speed (280-560 tok/s), so latency degradation is the highest-priority signal. The API returns rich timing metadata (queue_time, prompt_time, completion_time) and rate limit headers on every response.
Prerequisites
A Groq account with an API key exported as the GROQ_API_KEY environment variable — the groq-sdk client reads it automatically (new Groq()).
Node.js with groq-sdk and prom-client installed (npm install groq-sdk prom-client).
A Prometheus scrape target and (optionally) Grafana for the dashboard panels.
Key Metrics to Track
Metric
Type
Source
Why
TTFT (time to first token)
Histogram
Client-side timing
Groq's main value prop
Tokens/second
Gauge
usage.completion_time
Throughput degradation
Total latency
Histogram
Client-side timing
End-to-end performance
Rate limit remaining
Gauge
x-ratelimit-remaining-* headers
Prevent 429s
Token usage
Counter
usage.total_tokens
Cost attribution
Error rate by code
Counter
Error handler
Availability
Estimated cost
Counter
Tokens * model price
Budget tracking
Instructions
Apply these six steps in order. Steps 1-2 are the core instrumentation loop —
wrap the client, then feed a Prometheus instrument set from each call. Steps 3-6
add rate-limit tracking, alerting, structured logs, and dashboards on top. The
lean client skeleton is below; the full code for every step lives in
references/implementation.md.
Instrumented client — wrap groq.chat.completions.create so latency, tokens, queue time, and estimated cost are captured on the same path as the request ().
trackedCompletion
Prometheus metrics — register a histogram (latency), counters (tokens, cost, errors), and gauges (throughput, rate-limit remaining), then feed them from emitMetrics.
Rate limit header tracking — parse x-ratelimit-remaining-* off every response into a gauge so you alert before a 429, not after.
See references/implementation.md for the complete
GroqMetrics shape, pricing table, Prometheus instruments, rate-limit tracking,
alert rules, structured logging, and dashboard panel list.
Output
Applying the workflow produces:
A trackedCompletion wrapper that returns { result, metrics }, where metrics is a GroqMetrics object (latency, TTFT, tokens/sec, token counts, queue time, estimated cost).
A Prometheus metric set — groq_latency_ms (histogram), groq_tokens_total / groq_cost_usd / groq_errors_total (counters), and groq_tokens_per_second / groq_ratelimit_remaining (gauges).
Five alert rules (GroqLatencyHigh, GroqRateLimitCritical, GroqThroughputDrop, GroqErrorRateHigh, GroqCostSpike).
A structured JSON log line per request and a 7-panel dashboard spec.
Examples
Instrument a single completion and emit a structured log line:
const { result, metrics } = awaittrackedCompletion(
"llama-3.3-70b-versatile",
[{ role: "user", content: "Summarize this incident report in two sentences." }]
);
logGroqRequest(metrics, result.id);
// metrics.tokensPerSec -> 310, metrics.estimatedCostUsd -> 0.000404
For a 429-guard using rate-limit headers and a dashboard health-reading table,
see references/examples.md.