Skip to main content

llmobs-integrations

dd-trace-py LLMObs integration development guide. Use when creating, modifying, or debugging LLMObs integrations for LLM/AI libraries in the Python tracer. Covers BaseLLMIntegration, stream handling, message extraction, token counting, tool call parsing, and VCR-based testing patterns. Triggers: "llmobs", "LLMObs", "BaseLLMIntegration", "llmobs_set_tags", "_llmobs_set_tags", "BaseStreamHandler", "submit_to_llmobs", "integration.trace", "LLM span", "VCR", "cassette", "anthropic", "openai", "google_genai", "claude_agent_sdk", "generative-ai", "LLM integration", "llmobs_enabled".

설치로 이동

소스 정보

저장소
DataDog/dd-trace-py
최근 소스 활동
2026년 9월 18일 13:54
감지된 SKILL.md 언어
영어
스타
650
포크
552

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.

파일 탐색기
4 개 파일

SKILL.md 표시 중

SKILL.md
소스 지침 · 읽기 전용 미리보기
name
llmobs-integrations
description
dd-trace-py LLMObs integration development guide. Use when creating, modifying, or debugging LLMObs integrations for LLM/AI libraries in the Python tracer. Covers BaseLLMIntegration, stream handling, message extraction, token counting, tool call parsing, and VCR-based testing patterns. Triggers: "llmobs", "LLMObs", "BaseLLMIntegration", "llmobs_set_tags", "_llmobs_set_tags", "BaseStreamHandler", "submit_to_llmobs", "integration.trace", "LLM span", "VCR", "cassette", "anthropic", "openai", "google_genai", "claude_agent_sdk", "generative-ai", "LLM integration", "llmobs_enabled".
# dd-trace-py LLMObs Integrations LLMObs integrations enable Datadog LLM Observability for AI/LLM libraries. They extract model inputs, outputs, token usage, and tool calls from traced spans. This skill should be used in addition to the `apm-integrations` skill. ## Two-Layer Architecture LLMObs integrations consist of two cooperating layers: 1. **Patch Layer** (`ddtrace/contrib/internal/{name}/patch.py`) -- wraps library functions. Standard request/response LLM integrations construct `LlmRequestEvent` and use `core.context_with_event()` so the LLM tracing subscriber owns span lifecycle and LLMObs tag extraction. 2. **Integration Layer** (`ddtrace/llmobs/_integrations/{name}.py`) -- extends `BaseLLMIntegration`, implements `_set_base_span_tags()` and `_llmobs_set_tags()` to extract and set provider-specific messages, tools, metadata, and token metrics. Both layers must work together. The patch layer identifies the operation and passes request/response data through the event; the integration layer controls what data is extracted. ## Active Patch Patterns - **Event-based request spans**: Use `LlmRequestEvent` with `core.context_with_event()` for new standard request/response LLM integrations. Anthropic is the canonical reference. This is the preferred pattern. - **Direct integration spans**: Some existing or specialized integrations still call `integration.trace()` and `integration.llmobs_set_tags()` directly, especially for child spans, agent/tool spans, or integrations not yet migrated. Google GenAI, OpenAI tool spans, and Claude Agent SDK are useful references. ## Key Files | Purpose | File | |---------|------| | Base LLM integration class | `ddtrace/llmobs/_integrations/base.py` (`BaseLLMIntegration`) | | Stream handler base classes | `ddtrace/llmobs/_integrations/base_stream_handler.py` (`BaseStreamHandler`, `StreamHandler`, `AsyncStreamHandler`) | | Shared utilities | `ddtrace/llmobs/_integrations/utils.py` | | LLMObs annotation helper | `ddtrace/llmobs/_utils.py` (`_annotate_llmobs_span_data`) | | LLMObs constants | `ddtrace/llmobs/_constants.py` | | LLMObs types | `ddtrace/llmobs/types.py` (`Message`, `AudioPart`, `ToolCall`, `ToolResult`, `ToolDefinition`) | | Integration registry | `ddtrace/llmobs/_integrations/__init__.py` | ## Reference Integrations **Always read 1-2 references before writing or modifying LLMObs code.** | Provider | Patch File | LLMObs Integration | LLMObs Tests | |----------|-----------|-------------------|--------------| | **Anthropic** (canonical) | `ddtrace/contrib/internal/anthropic/patch.py` | `ddtrace/llmobs/_integrations/anthropic.py` | `tests/contrib/anthropic/test_anthropic_llmobs.py` | | **Claude Agent SDK** (latest, agent pattern) | `ddtrace/contrib/internal/claude_agent_sdk/patch.py` | `ddtrace/llmobs/_integrations/claude_agent_sdk.py` | `tests/contrib/claude_agent_sdk/test_claude_agent_sdk_llmobs.py` | | OpenAI | `ddtrace/contrib/internal/openai/patch.py` | `ddtrace/llmobs/_integrations/openai.py` | `tests/contrib/openai/test_openai_llmobs.py` | | Google GenAI | `ddtrace/contrib/internal/google_genai/patch.py` | `ddtrace/llmobs/_integrations/google_genai.py` | `tests/contrib/google_genai/test_google_genai_llmobs.py` | Use **Anthropic** as the canonical reference for standard LLM integrations. Use **Claude Agent SDK** for agent-pattern integrations (agent spans, tool child spans, thinking blocks). ## Abstract Methods to Implement Subclass `BaseLLMIntegration` and implement: ### `_set_base_span_tags(span, **kwargs)` Set provider-specific APM tags on the span (e.g., `{name}.request.model`). ### `_llmobs_set_tags(span, args, kwargs, response, operation)` Extract and annotate all LLMObs fields on the span: | Field | Description | |-------------|-------------| | `kind` | `"llm"` for LLM calls, `"agent"` for agent calls, `"tool"` for tool calls | | `model_name` | Model identifier (e.g., `"claude-3-sonnet-20240229"`) | | `model_provider` | Provider name (e.g., `"anthropic"`, `"openai"`) | | `input_messages` | List of `Message` objects from request | | `output_messages` | List of `Message` objects from response | | `metadata` | Dict of sanitized request parameters (temperature, top_p, etc.), plus response-derived scalars such as `finish_reason` (see Response Metadata below) | | `metrics` | Token usage dict with `INPUT_TOKENS_METRIC_KEY`, `OUTPUT_TOKENS_METRIC_KEY`, `TOTAL_TOKENS_METRIC_KEY` | | `tool_definitions` | List of `ToolDefinition` objects if tools are passed | Fields are usually set via `_annotate_llmobs_span_data(...)`, not raw `span._set_ctx_items(...)`. ### Response Metadata `metadata` is not request-only. Scalars the provider reports on the *response* belong there too, so long as they are not already covered by `metrics` (token counts) or `output_messages`. The stop reason is the established case. Record it under the key `finish_reason` for every provider, so one facet answers "why did generation stop" regardless of integration, and keep each provider's own vocabulary as the value (`stop`/`length`/`content_filter`/`tool_calls` for the OpenAI family, `end_turn`/`max_tokens`/`refusal`/`tool_use` for Anthropic). Read it from: | Provider | Source | |---|---| | openai/litellm chat + legacy completions | `choice.finish_reason` (per choice) | | openai responses API | `incomplete_details.reason` | | anthropic | `response.stop_reason`, falling back to `response.finish_reason` for the streamed dict the aggregator rebuilds | Rules: - **Merge, do not replace.** Build the request metadata first, then update it with the response-derived keys (`parameters.update(...)`), so response data never clobbers request params. - **Omit the key when absent.** If the provider reports no reason, leave `finish_reason` out entirely rather than writing `None` or `""`. - **Keep the value a single string.** When a request returns multiple choices (`n > 1`), comma-join the per-choice reasons in choice order (`"stop,length"`) instead of emitting a list, so the key never changes type. `_openai_finish_reason_metadata()` in `_integrations/utils.py` does this. - **Guard on the response, not `span.error`.** A response can exist on an errored span (see the AI Guard case in `_integrations/anthropic.py`); gate on `response is not None`. Note two already-shipped integrations predate this key: bedrock and the claude-agent-sdk both write `metadata["stop_reason"]`. Renaming those is a breaking change and has not been done — follow `finish_reason` for new work. ## Key Constraints - **`submit_to_llmobs=True`** must be set on `LlmRequestEvent` for event-based request spans or passed to `integration.trace()` for direct LLMObs spans - **`ctx.dispatch_ended_event()`** must run on success and error paths for event-based patch wrappers - **Streaming** must use `BaseStreamHandler`/`AsyncStreamHandler` -- never consume streams directly - **Event-based patch wrappers** should not call `span.set_exc_info()`, `span.finish()`, or `integration.llmobs_set_tags()` directly; the tracing subscriber handles that when the event ends - **Direct integration spans** must keep `integration.llmobs_set_tags()` and span lifecycle handling aligned with the closest current reference - **Integration instance** must be stored on the module: `module._datadog_integration = MyLibIntegration(integration_config=config.mylib)` ## Message Types ```python from ddtrace.llmobs.types import AudioPart, ImagePart, Message, ToolCall, ToolResult, ToolDefinition # Input/output messages Message(content="text", role="user") Message(content="response", role="assistant", tool_calls=[...]) # Audio attachments in multimodal messages AudioPart(mime_type="audio/wav", content="<base64-audio>") Message(content="", role="user", audio_parts=[...]) # Image attachments in multimodal messages ImagePart(mime_type="image/png", content="<base64-image>") Message(content="", role="user", image_parts=[...]) # Tool calls (in output messages) ToolCall(name="get_weather", arguments={"city": "NYC"}, tool_id="toolu_123", type="tool") # Tool results (in input messages) ToolResult(result="72F sunny", tool_id="toolu_123", type="tool_result") # Tool definitions (from request parameters) ToolDefinition(name="get_weather", description="...", schema={...}) ``` ## Debugging Quick Tips - **No LLMObs spans** -- check `submit_to_llmobs=True`, `ctx.dispatch_ended_event()`, and `llmobs_enabled` - **Wrong messages** -- check message extraction handles multi-part content and tool blocks - **Wrong tokens** -- check field name mapping (libraries use different names for token counts) - **Streaming broken** -- verify `BaseStreamHandler` subclass, check `finalize_stream()` dispatches the ended event or finishes direct-trace spans according to the reference pattern - **DD_TRACE_DEBUG=true** to see patching activity and span creation See [Failure Modes](references/failure-modes.md) for detailed debugging guide. ## Reference Files - [Implementation Guide](references/implementation-guide.md) -- LLM-specific steps (references apm-integrations guide for the full workflow) - [Failure Modes](references/failure-modes.md) -- All 8 failure modes with causes and fixes - [Testing Guide](references/testing-guide.md) -- LLMObs test patterns, VCR cassettes, suitespec
GitHub에서 보기