بنقرة واحدة
deepeval-sdk-tracing
Confirmation that traces are arriving in Confident AI's
التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.
القائمة
Confirmation that traces are arriving in Confident AI's
التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.
استنادا إلى تصنيف SOC المهني
Review code changes and report issues by severity with actionable fixes.
Sharpen a fuzzy intention into one measurable objective string that drives the rest of the work.
Convert a Prompt Flow PRS pipeline submission to run a Microsoft Agent Framework workflow.
Build a Model Context Protocol (MCP) server that lets an LLM call into external tools and resources.
Summarize PDF documents into concise bullet-point digests.
Bump a dependency version across a pnpm workspace and update lockfile.
| name | deepeval SDK Tracing |
| description | Confirmation that traces are arriving in Confident AI's |
| category | observability |
| tags | ["ai","api","backend","database","deployment"] |
| source | {"url":"https://github.com/confident-ai/deepeval/tree/d46df28e53b4efe95aa70ad172729a2c27e3de0d/skills/deepeval-tracing","fetched_at":"2026-06-12","commit":"d46df28e53b4efe95aa70ad172729a2c27e3de0d","license":"MIT","original_path":"skills/deepeval-tracing/SKILL.md"} |
| license | Apache-2.0 |
| author | Confident AI (downstream pack: badhope) |
| version | 0.1.0 |
| needs_review | false |
| slug | deepeval-tracing |
| created | 2026-06-12 |
| updated | 2026-06-19 |
| inputs | [{"name":"framework","type":"string","required":true,"description":"AI framework - langgraph/langchain/crewai/custom/etc"},{"name":"model_provider","type":"string","required":true,"description":"LLM provider - openai/anthropic/google/bedrock/azure/etc"},{"name":"vector_db","type":"string","required":false,"description":"Vector database if using retrieval (pinecone/qdrant/weaviate/chroma/milvus/pgvector)"},{"name":"confident_api_key","type":"string","required":true,"description":"Confident AI API key or deepeval login"}] |
| output | {"format":"markdown","description":"Generated content based on the user request"} |
| quality | stable |
You're instrumenting a Python AI app (LLM app, agent, RAG
pipeline, chatbot) and you want its execution to show up span by
span in Confident AI's Observatory so you can later attach
metrics and iterate. You're OK installing the deepeval Python
package and using its SDK conventions (@observe, framework
integrations).
This is the Python + deepeval SDK path. If you want a
language-agnostic setup or you don't want to install deepeval
itself, use the deepeval-otel skill (raw OpenTelemetry) instead.
If you want to define metrics and run pytest eval suites, use the
deepeval-eval-suite skill.
The job is exactly three things: pick a supported integration
when one exists, fall back to manual @observe otherwise, and
give every span a meaningful type (llm / retriever / tool
/ agent). The skill stops at producing well-formed traces —
attaching metrics and running evals is the eval-suite skill's job.
Don't use this skill for non-AI software. The span types describe AI components, and the Observatory renders AI behaviour. Tracing a Postgres query or a Flask handler is meaningless.
| Field | Required | Notes |
|---|---|---|
framework | yes | Drives integration choice. custom = manual @observe. |
model_provider | yes | Patched automatically by the matching integration. |
vector_db | no | Empty = no retriever spans needed. |
confident_api_key | yes | Or deepeval login for interactive setup. |
framework examples: langgraph uses deepeval.integrations.langgraph
patching, crewai uses deepeval.integrations.crewai, etc. If the
project mixes two frameworks (e.g. LangChain inside a CrewAI flow),
list both — integrations stack cleanly.
Plain-text confirmation. The system of record is Confident AI's Observatory; this skill just wires the SDK. Typical return:
Trace landed: https://app.confident-ai.com/observability/traces/d41f3a
Detected spans: llm=8, agent=3, retriever=2, tool=1
Instrumented: src/agents/customer_support.py:42 (@observe)
src/agents/customer_support.py:87 (@observe type=llm)
You are wiring DeepEval tracing into a Python AI application. Follow
the steps in order. Do not skip ahead.
1. Confirm the target is an AI app.
It must have at least one of: LLM calls, an agent loop, retrieval
/ vector search, or tool calls. If none are present, STOP — this
skill does not apply. Tell the user that DeepEval's span types
describe AI components and Confident AI's Observatory is built
to evaluate AI behaviour.
2. Detect what's in the codebase.
Grep for imports / framework signatures. Identify:
- AI framework (langgraph, langchain, openai-agents,
llamaindex, pydantic-ai, crewai, autogen, haystack, semantickernel)
- LLM provider (openai, anthropic, google, bedrock, azure, cohere,
mistral, ollama, vllm)
- Vector store (pinecone, qdrant, weaviate, chroma, milvus, pgvector)
If you can't detect any AI framework, mark `framework = custom` —
that triggers manual `@observe` below.
3. Read the exact integration doc for the detected framework +
provider. The docs/SDK are the source of truth for which functions
to patch and which auto-captures are enabled. Do not guess based
on framework name alone.
4. If a native integration fits, use it. Add the patch call to the
app's startup (or test fixture for pytest). Skip the manual
`@observe` step for code paths the integration auto-traces.
If no native integration fits (`framework = custom` or you have
app-owned wrapper boundaries that the integration can't see),
proceed to manual `@observe`.
5. Manual `@observe` rules:
- Decorate the top-level agent / chain / pipeline function.
- Decorate every LLM call with `@observe(type="llm")`.
- Decorate every retrieval call with `@observe(type="retriever")`.
- Decorate every tool call with `@observe(type="tool")`.
- Let span names default to the function name. Override only when
the function name is too generic (e.g. `def call`).
- Capture inputs and outputs in span metadata when they help
diagnose failures. NEVER include raw user PII, secrets, or
credentials.
6. Trace-level tags and metadata (optional but recommended for
production):
- `session_id` (link traces into a session / conversation)
- `user_id` (the end user, if known and allowed)
- `environment` (dev / staging / prod)
- `git_sha` (the deployed commit)
7. Confirm the API key is reachable. Either:
- `deepeval login` was run once and the key is in
`~/.deepeval/`, OR
- `CONFIDENT_API_KEY` is exported in the environment.
CI must export the env var. Bare `deepeval login` doesn't work
in CI.
8. Run the app, trigger one trace, and verify it shows up in
Confident AI's Observatory. If it doesn't:
- check the deepeval logs for auth errors
- verify `@observe` decorators are actually imported and
called
- check whether the framework's integration has been
initialised (most need an explicit `patch_*()` call)
deepeval). Use the deepeval-otel skill — it's the
language-agnostic OTLP path and works without the deepeval package.deepeval-eval-suite.Input:
framework: langgraph
model_provider: openai
vector_db: qdrant
confident_api_key: us-cf_***
Output:
Trace landed: https://app.confident-ai.com/observability/traces/d41f3a
Detected spans: llm=8, agent=3, retriever=2, tool=1
Instrumented:
src/agents/customer_support.py:42 @observe
src/agents/customer_support.py:87 @observe(type="llm")
src/agents/customer_support.py:120 @observe(type="retriever")
src/agents/customer_support.py:151 @observe(type="tool")
Next: run the eval-suite skill to attach metrics.
These are the bugs that bite every new user. Check them before shipping:
Integration not initialized: Framework patch not called at startup.
Missing @observe decorators: LLM calls not decorated but expected to appear.
Credentials not in CI: deepeval login works locally but not in CI.
PII in span metadata: Raw user data sent to Confident AI.
Framework not detected correctly: Wrong integration chosen.