Skip to main content 홈 크리에이터 jeremylongshore tons-of-skills-marketplace langfuse-ci-integration
langfuse-ci-integration Configure Langfuse CI/CD integration with GitHub Actions and automated testing.
Use when setting up automated testing, configuring CI pipelines,
or integrating Langfuse tests into your build process.
Trigger with phrases like "langfuse CI", "langfuse GitHub Actions",
"langfuse automated tests", "CI langfuse", "langfuse pipeline".
설치로 이동 Skills Marketplace 커뮤니티가 만든 AI 스킬을 발견하고 탐색하세요.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
직접 명령은 검토 Prompt를 거치지 않습니다. 실행하기 전에 소스를 확인하세요.
npx skills add https://github.com/jeremylongshore/tons-of-skills-marketplace --skill langfuse-ci-integration명령은 한 줄로 유지됩니다. 복사하기 전에 가로로 스크롤해 전체 내용을 확인하세요.
로컬 사본을 원하시나요? SkillsMP에서 현재 제공할 수 있는 파일을 다운로드하세요.
Zip 다운로드 다운로드 중... 이 저장소의 다른 Skills langchain-deploy-integration Deploy a LangChain 1.0 / LangGraph 1.0 app to Cloud Run, Vercel, or LangServe correctly — with timeouts sized for chain length, cold-start mitigation, SSE anti-buffering headers, and Secret Manager over .env. Use when prepping a first production deploy, debugging a stream that hangs behind a proxy, or diagnosing p99 latency spikes. Trigger with "langchain deploy", "langchain cloud run", "langchain vercel python", "langchain langserve", or "langchain docker".
langchain-langgraph-agents Build a correct LangGraph 1.0 ReAct agent with create_react_agent — typed tools, error propagation, recursion caps, and stop conditions that actually stop. Use when writing a first tool-calling agent, migrating from AgentExecutor or initialize_agent, or diagnosing an agent that loops on vague prompts. Trigger with "langgraph agent", "create_react_agent", "langgraph tool calling", "AgentExecutor migration", or "agent loop cost".
langchain-langgraph-human-in-loop Build LangGraph 1.0 human-in-the-loop approval flows with interrupt_before /
interrupt_after and Command(resume=...) — JSON-serializable state, clean
resume semantics, and UI wiring for approval decisions. Use when adding an
approval gate before an expensive tool call, wiring a Slack/web UI for agent
approvals, or debugging a graph that crashes on interrupt.
Trigger with "langgraph human in loop", "langgraph interrupt_before",
"langgraph approval flow", "Command resume", "langgraph HITL".
jeremylongshore
jeremylongshore/tons-of-skills-marketplace
GitHub 저장소 열기 name langfuse-ci-integration description Configure Langfuse CI/CD integration with GitHub Actions and automated testing.
Use when setting up automated testing, configuring CI pipelines,
or integrating Langfuse tests into your build process.
Trigger with phrases like "langfuse CI", "langfuse GitHub Actions",
"langfuse automated tests", "CI langfuse", "langfuse pipeline".
allowed-tools Read, Write, Edit, Bash(gh:*) version 1.12.0 license MIT author Jeremy Longshore <jeremy@intentsolutions.io> tags ["saas","langfuse","testing","ci-cd"] compatibility Designed for Claude Code
Langfuse CI Integration
Overview
Integrate Langfuse into CI/CD pipelines: trace validation tests, prompt regression testing, experiment-driven quality gates, automated prompt deployment from version control, and score monitoring.
Prerequisites
Langfuse API keys stored as GitHub secrets (LANGFUSE_PUBLIC_KEY, LANGFUSE_SECRET_KEY)
Test framework (Vitest or Jest)
OpenAI API key for LLM tests
Instructions
Step 1: GitHub Actions Workflow for AI Quality Tests
name: AI Quality Tests
on:
pull_request:
paths: ["src/ai/**" , "src/prompts/**" , "tests/ai/**" ]
jobs:
ai-quality:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with: { node-version: "20" , cache: "npm" }
- run: npm ci
- name: Run AI quality tests with tracing
env:
LANGFUSE_PUBLIC_KEY: ${{ secrets.LANGFUSE_PUBLIC_KEY
}}
LANGFUSE_SECRET_KEY:
${{
secrets.LANGFUSE_SECRET_KEY
}}
LANGFUSE_BASE_URL:
${{
vars.LANGFUSE_BASE_URL
||
'https://cloud.langfuse.com'
}}
OPENAI_API_KEY:
${{
secrets.OPENAI_API_KEY
}}
run:
npx
vitest
run
tests/ai/
--reporter=verbose
-
name:
Langfuse
connectivity
check
env:
LANGFUSE_PUBLIC_KEY:
${{
secrets.LANGFUSE_PUBLIC_KEY
}}
LANGFUSE_SECRET_KEY:
${{
secrets.LANGFUSE_SECRET_KEY
}}
run:
|
node -e "
const { LangfuseClient } = require('@langfuse/client');
const lf = new LangfuseClient();
lf.prompt.get('__ci-health__').catch(() => {});
console.log('Langfuse SDK initialized OK');
"
Step 2: Prompt Regression Tests
import { describe, it, expect, afterAll } from "vitest" ;
import { LangfuseClient } from "@langfuse/client" ;
import { startActiveObservation, updateActiveObservation } from "@langfuse/tracing" ;
import OpenAI from "openai" ;
const langfuse = new LangfuseClient ();
const openai = new OpenAI ();
describe ("Prompt Quality Regression" , () => {
it ("summarization prompt produces valid output" , async () => {
const prompt = await langfuse.prompt .get ("summarize-article" , { type : "text" });
const compiled = prompt.compile ({ maxLength : "100 words" });
const result = await startActiveObservation (
{ name : "ci-test-summarize" , asType : "generation" },
async () => {
updateActiveObservation ({ model : "gpt-4o-mini" , input : compiled });
const response = await openai.chat .completions .create ({
model : "gpt-4o-mini" ,
messages : [{ role : "user" , content : compiled }],
temperature : 0 ,
});
const output = response.choices [0 ].message .content || "" ;
updateActiveObservation ({
output,
usage : {
promptTokens : response.usage ?.prompt_tokens ,
completionTokens : response.usage ?.completion_tokens ,
},
});
return output;
}
);
expect (result.length ).toBeGreaterThan (20 );
expect (result.length ).toBeLessThan (600 );
});
it ("classification prompt returns valid intent" , async () => {
const prompt = await langfuse.prompt .get ("classify-intent" , { type : "text" });
const compiled = prompt.compile ({ userMessage : "I want to cancel my subscription" });
const response = await openai.chat .completions .create ({
model : "gpt-4o-mini" ,
messages : [{ role : "user" , content : compiled }],
temperature : 0 ,
});
const intent = response.choices [0 ].message .content ?.trim ().toLowerCase () || "" ;
const validIntents = ["billing" , "cancellation" , "support" , "feedback" ];
expect (validIntents).toContain (intent);
});
});
Step 3: Experiment-Driven Quality Gates
import { describe, it, expect } from "vitest" ;
import { LangfuseClient } from "@langfuse/client" ;
import OpenAI from "openai" ;
const langfuse = new LangfuseClient ();
const openai = new OpenAI ();
describe ("Quality Gate: Intent Classification" , () => {
it ("scores above 80% accuracy on test dataset" , async () => {
async function classifyIntent (input : { query: string } ) {
const response = await openai.chat .completions .create ({
model : "gpt-4o-mini" ,
messages : [
{ role : "system" , content : "Classify intent. Return one word." },
{ role : "user" , content : input.query },
],
temperature : 0 ,
});
return response.choices [0 ].message .content ?.trim () || "" ;
}
const result = await langfuse.runExperiment ({
datasetName : "intent-classification-test" ,
runName : `ci-${process.env.GITHUB_SHA?.slice(0 , 7 ) || "local" } ` ,
task : classifyIntent,
evaluators : [
({ output, expectedOutput } ) => ({
name : "exact-match" ,
value : output.toLowerCase () === expectedOutput.intent .toLowerCase () ? 1 : 0 ,
dataType : "BOOLEAN" as const ,
}),
],
});
const scores = result.runs .flatMap ((r ) => r.scores || []);
const accuracy = scores.filter ((s ) => s.value === 1 ).length / scores.length ;
console .log (`Accuracy: ${(accuracy * 100 ).toFixed(1 )} %` );
expect (accuracy).toBeGreaterThanOrEqual (0.8 );
});
});
Step 4: Automated Prompt Deployment
name: Deploy Prompts to Langfuse
on:
push:
branches: [main ]
paths: ["src/prompts/**" ]
jobs:
deploy:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with: { node-version: "20" , cache: "npm" }
- run: npm ci
- name: Deploy prompts
env:
LANGFUSE_PUBLIC_KEY: ${{ secrets.LANGFUSE_PUBLIC_KEY }}
LANGFUSE_SECRET_KEY: ${{ secrets.LANGFUSE_SECRET_KEY }}
run: node scripts/deploy-prompts.mjs
import { LangfuseClient } from "@langfuse/client" ;
import { readdirSync, readFileSync } from "fs" ;
import { join } from "path" ;
const langfuse = new LangfuseClient ();
const promptDir = join (process.cwd (), "src/prompts" );
for (const file of readdirSync (promptDir).filter ((f ) => f.endsWith (".json" ))) {
const config = JSON .parse (readFileSync (join (promptDir, file), "utf-8" ));
await langfuse.api .prompts .create ({
name : config.name ,
prompt : config.template ,
type : config.type || "text" ,
config : config.config || {},
labels : ["production" , `deploy-${new Date ().toISOString().split("T" )[0 ]} ` ],
});
console .log (`Deployed: ${config.name} ` );
}
Step 5: Score Regression Monitoring
import { LangfuseClient } from "@langfuse/client" ;
const langfuse = new LangfuseClient ();
async function checkRegression ( ) {
const scores = await langfuse.api .scores .list ({
name : "quality" ,
limit : 100 ,
});
const values = scores.data .map ((s ) => s.value ).filter ((v): v is number => v !== null );
const avg = values.reduce ((a, b ) => a + b, 0 ) / values.length ;
console .log (`Average quality score: ${avg.toFixed(3 )} (n=${values.length} )` );
if (avg < 0.7 ) {
console .error ("QUALITY REGRESSION: Score below 0.7 threshold" );
process.exit (1 );
}
}
checkRegression ();
CI Best Practices Practice Why Use temperature: 0 in CI tests Deterministic outputs, fewer false failures Separate CI API keys Isolate test traces from production Run experiments on dataset changes Catch regressions before deploy Assert on ranges, not exact strings LLM output varies even at temp 0 Flush/shutdown in afterAll Ensure all traces reach Langfuse
Error Handling Issue Cause Solution Traces not in dashboard No flush in CI Add sdk.shutdown() or afterAll flush Flaky quality tests Non-deterministic LLM Use temperature: 0, assert on ranges Prompt not found Not yet deployed Deploy prompts before running tests Missing secrets in CI Not configured Add to GitHub Settings > Secrets > Actions
Resources