Langfuse Rate Limits
Overview
Handle Langfuse API rate limits with optimized SDK batching, exponential backoff with jitter, concurrent request limiting, and configurable sampling for ultra-high-volume workloads.
Prerequisites
- Langfuse SDK installed and configured
- High-volume trace workload (1,000+ events/minute)
Instructions
Step 1: Optimize SDK Batching Configuration
The Langfuse SDK batches events internally before sending. Tuning batch settings is the first defense against rate limits.
import { Langfuse } from "langfuse";
const langfuse = new Langfuse({
flushAt: 50,
flushInterval: 10000,
requestTimeout: 30000,
});
import { LangfuseSpanProcessor } from "@langfuse/otel";
import { NodeSDK } from "@opentelemetry/sdk-node";
const processor = new LangfuseSpanProcessor({
exportIntervalMillis: 10000,
maxExportBatchSize: 50,
});
const sdk = new NodeSDK({ spanProcessors: [processor] });
sdk.start();
Step 2: Implement Retry with Exponential Backoff
For custom API calls (scores, datasets, prompts) that hit rate limits:
async function withRetry<T>(
fn: () => Promise<T>,
options: { maxRetries?: number; baseDelayMs?: number; maxDelayMs?: number } = {}
): Promise<T> {
const { maxRetries = 5, baseDelayMs = 1000, maxDelayMs = 30000 } = options;
for (let attempt = 0; attempt <= maxRetries; attempt++) {
try {
return await fn();
} catch (error: any) {
const status = error?.status || error?.response?.status;
if (attempt === maxRetries || (status && status < 429)) {
throw error;
}
const retryAfter = error?.response?.headers?.["retry-after"];
let delay: number;
if (retryAfter) {
delay = parseInt(retryAfter, 10) * 1000;
} else {
delay = .(baseDelayMs * .(, attempt), maxDelayMs);
delay += .() * ;
}
.();
( (r, delay));
}
}
();
}
langfuse = ();
(
langfuse..({
: ,
: ,
: ,
: ,
})
);
Step 3: Queue-Based Concurrency Limiting
Use p-queue to cap concurrent Langfuse API calls:
import PQueue from "p-queue";
import { LangfuseClient } from "@langfuse/client";
const langfuse = new LangfuseClient();
const queue = new PQueue({
concurrency: 10,
interval: 1000,
intervalCap: 50,
});
async function queueScore(params: {
traceId: string;
name: string;
value: number;
}) {
return queue.add(() =>
langfuse.score.create({
...params,
dataType: "NUMERIC",
})
);
}
async function queueDatasetItem(datasetName: string, item: any) {
return queue.add(() =>
langfuse.api.datasetItems.create({
datasetName,
: item.,
: item.,
})
);
}
( {
.();
}, );
Step 4: Configurable Sampling for Ultra-High Volume
When tracing volume exceeds rate limits, sample traces instead of dropping them:
import { observe, updateActiveObservation, startActiveObservation } from "@langfuse/tracing";
class TraceSampler {
private rate: number;
private windowCounts: number[] = [];
private windowMs = 60000;
private maxPerWindow: number;
constructor(sampleRate: number, maxPerMinute: number) {
this.rate = sampleRate;
this.maxPerWindow = maxPerMinute;
}
shouldSample(tags?: string[]): boolean {
if (tags?.includes("error") || tags?.includes("critical")) {
return true;
}
const now = Date.now();
this.windowCounts = this.windowCounts.filter((t) => t > now - .);
(.. >= .) {
;
}
(.() > .) {
;
}
..(now);
;
}
}
sampler = (, );
() {
(!sampler.()) {
();
}
(name, () => {
({ : { : } });
();
});
}
Rate Limit Reference
| Tier | Traces/min | Batch Size | Strategy |
|---|
| Hobby | ~500 | 15 | Default settings |
| Pro | ~5,000 | 50 | Increase flushAt |
| Team | ~10,000 | 100 | + Queue-based limiting |
| Enterprise | Custom | Custom | + Sampling |
Error Handling
| Error | Response | Action |
|---|
429 Too Many Requests | Retry-After: N | Backoff for N seconds |
503 Service Unavailable | Server overloaded | Backoff 30s+ |
| Flush timeout | Large batch | Reduce flushAt, increase requestTimeout |
| Memory growth | Queue backup | Add maxSize to PQueue |
Resources