Skip to main content
langfuse-core-workflow-b Execute Langfuse secondary workflow: Evaluation, scoring, and datasets.
Use when implementing LLM evaluation, adding user feedback,
or setting up automated quality scoring and experiment datasets.
Trigger with phrases like "langfuse evaluation", "langfuse scoring",
"rate llm outputs", "langfuse feedback", "langfuse datasets", "langfuse experiments".
Zur Installation springen Skills Marktplatz Entdecken und erkunden Sie KI-Skills, die von der Community erstellt wurden.
Mit Codex oder Claude installieren Kopieren Sie diesen Prompt, fügen Sie ihn in Codex, Claude oder einen anderen Assistant ein und lassen Sie die Skill-Seite prüfen und installieren.
Prompt kopierenPrompt-Details anzeigen Ein direkter Befehl überspringt den Prüf-Prompt. Prüfen Sie die Quelle, bevor Sie ihn ausführen.
npx skills add https://github.com/jeremylongshore/claude-code-plugins-plus-skills --skill langfuse-core-workflow-bDer Befehl bleibt in einer Zeile. Scrollen Sie horizontal, um ihn vor dem Kopieren vollständig zu prüfen.
Sie bevorzugen eine lokale Kopie? Laden Sie die Dateien herunter, die SkillsMP derzeit vorliegen.
ZIP herunterladen Herunterladen... Mehr aus diesem Repository Implement user sign-up and sign-in flows with Clerk.
Use when building authentication UI, customizing sign-in experience,
or implementing OAuth social login.
Trigger with phrases like "clerk sign-in", "clerk sign-up",
"clerk login flow", "clerk OAuth", "clerk social login".
Implement session management and middleware with Clerk.
Use when managing user sessions, configuring route protection,
or implementing token refresh and custom JWT templates.
Trigger with phrases like "clerk session", "clerk middleware",
"clerk route protection", "clerk token", "clerk JWT".
Configure enterprise SSO, role-based access control, and organization management.
Use when implementing SSO integration, configuring role-based permissions,
or setting up organization-level controls.
Trigger with phrases like "clerk SSO", "clerk RBAC",
"clerk enterprise", "clerk roles", "clerk permissions", "clerk organizations".
Verwandte Berufe SOC
Basierend auf der SOC-Berufsklassifikation
name langfuse-core-workflow-b description Execute Langfuse secondary workflow: Evaluation, scoring, and datasets.
Use when implementing LLM evaluation, adding user feedback,
or setting up automated quality scoring and experiment datasets.
Trigger with phrases like "langfuse evaluation", "langfuse scoring",
"rate llm outputs", "langfuse feedback", "langfuse datasets", "langfuse experiments".
allowed-tools Read, Write, Edit, Bash(npm:*), Grep version 1.12.0 license MIT author Jeremy Longshore <jeremy@intentsolutions.io> tags ["saas","langfuse","llm","workflow","evaluation"] compatibility Designed for Claude Code, also compatible with Codex and OpenClaw
Langfuse Core Workflow B: Evaluation, Scoring & Datasets
Overview
Implement LLM output evaluation using Langfuse scores (numeric, categorical, boolean), the experiment runner SDK for dataset-driven benchmarks, prompt management with versioned prompts, and LLM-as-a-Judge evaluation patterns.
Prerequisites
Langfuse SDK configured with API keys
Traces already being collected (see langfuse-core-workflow-a)
For v4+: @langfuse/client installed
Instructions
Step 1: Score Traces via SDK
Langfuse supports three score data types: Numeric , Categorical , and Boolean .
import { LangfuseClient } from "@langfuse/client" ;
const langfuse = new LangfuseClient ();
await langfuse.score .create ({
traceId : "trace-abc-123" ,
name : "relevance" ,
value : 0.92 ,
dataType : "NUMERIC" ,
comment : "Highly relevant answer with good context usage" ,
});
await langfuse.score .create ({
traceId : "trace-abc-123" ,
observationId : "gen-xyz-456" ,
name : "quality-tier" ,
value : "excellent" ,
dataType : ,
});
langfuse. . ({
: ,
: ,
: ,
: ,
: ,
});
"CATEGORICAL"
await
score
create
traceId
"trace-abc-123"
name
"user-approved"
value
1
dataType
"BOOLEAN"
comment
"User clicked thumbs up"
Step 2: User Feedback Collection
app.post ("/api/feedback" , async (req, res) => {
const { traceId, rating, comment } = req.body ;
await langfuse.score .create ({
traceId,
name : "user-feedback" ,
value : rating === "positive" ? 1 : 0 ,
dataType : "BOOLEAN" ,
comment,
});
if (req.body .stars ) {
await langfuse.score .create ({
traceId,
name : "star-rating" ,
value : req.body .stars ,
dataType : "NUMERIC" ,
comment : `${req.body.stars} /5 stars` ,
});
}
res.json ({ success : true });
});
Step 3: Prompt Management
const textPrompt = await langfuse.prompt .get ("summarize-article" , {
type : "text" ,
label : "production" ,
});
const compiled = textPrompt.compile ({
maxLength : "100 words" ,
tone : "professional" ,
});
const chatPrompt = await langfuse.prompt .get ("customer-support" , {
type : "chat" ,
});
const messages = chatPrompt.compile ({
customerName : "Alice" ,
issue : "billing question" ,
});
Step 4: Create and Populate Datasets
await langfuse.api .datasets .create ({
name : "customer-support-v1" ,
description : "Test cases for customer support chatbot" ,
metadata : { version : "1.0" , domain : "support" },
});
const testCases = [
{
input : { query : "How do I cancel my subscription?" },
expectedOutput : { intent : "cancellation" , sentiment : "neutral" },
metadata : { category : "billing" },
},
{
input : { query : "Your product is amazing!" },
expectedOutput : { intent : "feedback" , sentiment : "positive" },
metadata : { category : "feedback" },
},
];
for (const testCase of testCases) {
await langfuse.api .datasetItems .create ({
datasetName : "customer-support-v1" ,
input : testCase.input ,
expectedOutput : testCase.expectedOutput ,
metadata : testCase.metadata ,
});
}
Step 5: Run Experiments with the Experiment Runner import { LangfuseClient } from "@langfuse/client" ;
const langfuse = new LangfuseClient ();
async function classifyIntent (input : { query: string } ): Promise <string > {
const response = await openai.chat .completions .create ({
model : "gpt-4o-mini" ,
messages : [
{ role : "system" , content : "Classify the user intent. Return one word." },
{ role : "user" , content : input.query },
],
temperature : 0 ,
});
return response.choices [0 ].message .content ?.trim () || "" ;
}
function exactMatch ({ output, expectedOutput }: {
output: string ;
expectedOutput: { intent: string };
} ) {
return {
name : "exact-match" ,
value : output.toLowerCase () === expectedOutput.intent .toLowerCase () ? 1 : 0 ,
dataType : "BOOLEAN" as const ,
};
}
const result = await langfuse.runExperiment ({
datasetName : "customer-support-v1" ,
runName : "gpt-4o-mini-classifier-v1" ,
runDescription : "Testing intent classification with gpt-4o-mini" ,
task : classifyIntent,
evaluators : [exactMatch],
});
console .log (`Experiment complete. ${result.runs.length} items evaluated.` );
Step 6: LLM-as-a-Judge Evaluation async function llmJudge ({ output, input, expectedOutput }: {
output: string ;
input: { query: string };
expectedOutput: { intent: string ; sentiment: string };
} ) {
const judgment = await openai.chat .completions .create ({
model : "gpt-4o" ,
temperature : 0 ,
messages : [
{
role : "system" ,
content : `You are an AI evaluator. Score the response 0-10 on accuracy and helpfulness.
Return JSON: {"score": <number>, "reasoning": "<explanation>"}` ,
},
{
role : "user" ,
content : `Query: ${input.query} \nExpected: ${JSON .stringify(expectedOutput)} \nActual: ${output} ` ,
},
],
response_format : { type : "json_object" },
});
const result = JSON .parse (judgment.choices [0 ].message .content || "{}" );
return {
name : "llm-judge-quality" ,
value : result.score / 10 ,
dataType : "NUMERIC" as const ,
comment : result.reasoning ,
};
}
await langfuse.runExperiment ({
datasetName : "customer-support-v1" ,
runName : "judge-evaluation-v1" ,
task : classifyIntent,
evaluators : [exactMatch, llmJudge],
});
Error Handling Issue Cause Solution Scores not appearing API call failed silently Await score.create() and check for errors Score validation error Wrong data type Match value type to dataType (number/string/0-1) LLM judge inconsistent High temperature Set temperature: 0 for evaluation calls Dataset item missing Wrong dataset name Verify exact name match (case-sensitive) Experiment not in UI Run not flushed Check runExperiment completed without errors
Resources
Next Steps For common error debugging, see langfuse-common-errors. For CI/CD integration of evaluations, see langfuse-ci-integration.