Skip to main content

observability

Agent observability, evals, feedback, and experiments. Use when adding observability dashboards, configuring trace capture, setting up evals, creating A/B experiments, or collecting user feedback on agent responses.

ソース情報

リポジトリ
BuilderIO/agent-native
ソースの最終更新活動
2026年10月1日 18:59
検出された SKILL.md の言語
英語
スター
7,065
フォーク
640

インストール方法

デフォルトでは、最初にソースを確認する Prompt が選択されています。直接コマンドに切り替えるか、ローカルコピーをダウンロードすることもできます。

ソースファイルを確認

インストールを決める前に、SKILL.md と SkillsMP に表示されている付属ファイルをお読みください。

SKILL.md を表示中

SKILL.md
ソースの指示 · 読み取り専用プレビュー
name
observability
description
Agent observability, evals, feedback, and experiments. Use when adding observability dashboards, configuring trace capture, setting up evals, creating A/B experiments, or collecting user feedback on agent responses.
scope
dev
metadata
{"internal":true}
# Agent Observability ## Rule The observability system auto-instruments every agent run with zero configuration. Traces, automated evals, and feedback collection work out of the box. All data lives in the app's own SQL database — no external services required. Templates can optionally export to Langfuse, Datadog, or any OTel-compatible platform. ## Five Pillars ### 1. Traces Every `runAgentLoop()` call is automatically instrumented via `instrumentAgentLoop()` in `packages/core/src/observability/traces.ts`. It captures: - **agent_run** span — top-level parent with total duration and cost - **llm_call** span — model name, token counts (input, output, cache read/write), cost - **tool_call** spans — one per action invocation, with duration and success/error The run span is NAMED by what started it, because the trace list shows that name and nothing else: a scheduled job is `background_automation_run:<job name>`, a chat turn is `agent_run`, and a turn a feature sent on the user's behalf is `agent_run:<usageLabel>` when the caller named it: ```ts sendToAgentChat({ message: "Enrich this record from the web", newTab: true, usageLabel: "crm:enrich-record", // → usage row label + `agent_run:crm:enrich-record` }); ``` Content (prompts, tool args, tool results) is **redacted by default**. Opt in through the declared `observability` config domain: ```ts // server/plugins/config.ts import { defineAppConfig } from "@agent-native/core/server"; export default defineAppConfig({ observability: { enabled: true, capturePrompts: false, captureToolArgs: true, // capture action input args captureToolResults: false, // include tool results and the full error text on tool spans and $ai_generation entries evalSampleRate: 0.05, // 5% of runs get LLM-as-judge eval inferredSentimentEnabled: false, inferredSentimentSampleRate: 0, inferredSentimentModel: "gpt-5-6-luna", }, }); ``` Two things are recorded whatever the flags say, because a failure nobody can classify is not observable: - **A failed tool span always keeps a signature**: the first line of the error, redacted (credentials, emails, long opaque ids) and capped at 500 characters. `captureToolResults` only decides whether the rest is kept. The span's metadata says which (`__tool_error_detail`: `full` | `signature`), and the read path returns it as `errorDetail` (`full` | `signature` | `withheld` | `unrecorded`) so "withheld on purpose" is never read as "nothing was recorded". `$ai_*` events and OTel spans still follow the flag. - **A stop that waits on the user is not an error.** An `input_required` outcome (question, approval, connection) records the `agent_run` span as `status: "paused"` with the reason in `terminal_code`, not `error`. Anything that counts failed runs must treat `paused` as non-error; an `error` event that follows the pause still makes the run an error. #### From an error report to the failing run A report is only useful if it names where it happened. Four pieces make that true, so nobody has to ask the reporter for an example: - **The failure packet.** Every server `captureError()` carries `extra.failureContext` (`app`, `route`, `actionName`, `automationName`, `threadId`, `runId`, `requestId`, `userScope`, `threadUrl`, build, environment, `errorCode`, `failureClass`); see the tracking skill for how it is derived. `threadUrl` is `https://<app host>/?thread=<chat_threads.id>`, and the Analytics issue page links it. The same fields (`$ai_trace_id` = run id, `thread_id`) are on the `$ai_trace` for the run. - **A copyable report on the client.** `formatClientFailureReport()` from `@agent-native/core/client/failure-report` builds the plain-text packet (app, thread link, run, request, code, time, build, and the inspection call) for a "Copy details" button. The chat run-error card's existing "Copy debug info" button uses it; call it from any other error card. It needs only what the card already knows (`message`, `errorCode`, `runId`, `requestId`) and falls back to the open thread. The error-screen "Open GitHub issue" template adds the same thread link, build and time. A failed action carries `error.requestId` (the response's `x-agent-native-request-id`, also on the server capture's `failureContext.requestId`). - **Read by id.** Dispatch's `get-agent-thread-debug` takes a `threadId` or the copied run id (`run-…`) and returns the run's `terminalReason` and terminal event, its events, trace spans (failed tools carry a signature, see above), feedback and checkpoints; `list-agent-run-failures` finds candidates. Both are owner-scoped. Without Dispatch, `GET /_agent-native/observability/traces/:runId` returns the same spans for the caller's own runs. - **Alerts name examples.** The chat-health Slack alert lists the latest failed turns as `threadUrl (run id, code)`; if they cannot be read it says so. #### Optional inferred sentiment Self-hosted apps default to no inferred sentiment. First-party apps hosted on `agent-native.com` automatically classify 100% of eligible user replies with `gpt-5-6-luna`; an explicit stored `inferredSentimentEnabled: false` remains an opt-out. Deployment overrides are `AGENT_NATIVE_INFERRED_SENTIMENT=on|off`, `AGENT_NATIVE_INFERRED_SENTIMENT_SAMPLE_RATE=0..1`, and `AGENT_NATIVE_INFERRED_SENTIMENT_MODEL=<model>`; `off` is always the emergency kill switch. Classification uses only the original visible user text, capped at 2,000 characters, with no tools, temperature 0, an eight-token output, and a five second timeout. It skips attachment-only turns, internal continuations, chained background chunks, and first turns that have no preceding response to attribute. The managed Builder engine runs the classifier after the main response has streamed, so it does not contend with the user's response. Successful classifications emit a content-free `$ai_sentiment` tracking event: - `sentiment`: `positive`, `negative`, or `neutral` - `method`: `llm` - `model` / `$ai_model`: model that generated the preceding assistant response - `run_id` / `$ai_trace_id`: preceding response run - `thread_id` / `$ai_session_id`: conversation - `classification_trigger_run_id`: run started by the classified user reply - `classifier_model` and `classifier_engine`: classifier attribution No raw message, prompt, or response text is persisted or tracked. ### 2. Feedback **Explicit** — AgentKit's assistant-message action bar renders inline thumbs up/down controls. A thumbs-down can collect a reason, and feedback includes the run and message sequence for trace linking. The shared `AgentKitAssistantChat` host submits it through the existing feedback action. **Implicit** — `computeSatisfactionScore(threadId)` computes a Frustration Index (0-100) from conversation signals: - Rephrasing detection (weight 30): consecutive similar user messages - Abandonment (weight 20): session ends shortly after agent response - Sentiment (weight 15): negative language patterns - Length trend (weight 15): declining message lengths - Retry patterns (weight 20): "try again", "no that's wrong" Score interpretation: 0-20 healthy, 20-40 friction, 40-60 dissatisfied, 60+ broken. Satisfaction scoring fires automatically after each feedback POST with a threadId. ### Human audit and refinement For a human output audit, provide one table row per run with a distilled ask, result, inferred or explicit sentiment, audit state, and app-native preview. The same row contract works for standalone app and workspace roll-ups; keep single-app surfaces independent of workspace chrome. Treat ordinary feedback separately from admin audit verdicts. Only admins can vote, mark audited, approve improvements, or apply them, enforced server-side with app-scoped roles. Let admins add reasons to downvotes across many rows, then run one synthesis that groups patterns, cites evidence, suggests update targets (skills, instructions, memories, data dictionaries, certified dashboards, or creative context), and shows before/after diffs. Allow feedback and regeneration before explicit approval; never auto-apply. The shipped single-app human review surface is the `Human review` tab in the shared observability dashboard. It lists the persisted ask and answer for each run, records thumbs/text feedback through the existing feedback endpoint, and saves explicit instruction changes as `draft` rows through the `save-observability-instruction-update` action. Drafts are reviewable artifacts; they never update agent behavior automatically. Agents can read the same table through `list-observability-reviews`. ### 3. Evals Three layers, configured via `evalSampleRate` in the observability config: **Automated (every run):** Deterministic scorers that run after every traced run: - `tool_success_rate` — % of tool calls without errors - `step_efficiency` — 1.0 for no-tool runs; penalizes excessive LLM iterations for tool-using runs - `latency_score` — normalized against 10s/tool baseline - `cost_efficiency` — normalized against 50 centicents/tool baseline - `error_recovery` — 1.0 if the run recovered from tool errors or had none **LLM-as-judge (sampled):** Runs on `evalSampleRate` fraction of runs. Calls the configured engine with a judge prompt that scores against custom criteria. **Dataset evaluation:** `runDatasetEval(datasetId)` runs a golden dataset through the agent and scores each case. Promote a surprising production run into a CI eval case (`promote-trace-eval` / `agent-native eval promote <runId> --write ...`) rather than only inspecting it. Truncated runs fail closed. Custom criteria use natural language rubrics: ```ts const criteria: EvalCriteria = { name: "helpfulness", description: "Was the response helpful and complete?", rubric: "0.0 = completely unhelpful, 0.5 = partially helpful, 1.0 = fully resolved the user's need", }; ``` #### Evals (CI gate) The three layers above score *real production runs* after the fact. For an active, deterministic gate, use the first-class `*.eval.ts` primitive from `@agent-native/core/eval` (source: `packages/core/src/eval/*`). It runs the actual agent loop against fixed inputs and exits non-zero below threshold, so it gates CI/deploys. ```ts // evals/faq.eval.ts import { defineEval, contains, llmJudge } from "@agent-native/core/eval"; export default defineEval({ name: "answers the FAQ", input: { prompt: "What is your return policy?" }, threshold: 0.7, scorers: [contains("30 days"), llmJudge({ criteria: "accuracy" })], }); ``` - Built-in scorers: `exactMatch` / `contains` / `usesTool` (pure JS) and `llmJudge` (provider-agnostic judge). - Custom scorers: `createScorer` with the 4-step `preprocess → analyze → generateScore → generateReason` pipeline (only `generateScore` is required). - Run as a gate: `agent-native eval [pattern] [--json] [--threshold N]` — discovers `**/*.eval.ts` and `evals/*.ts`, runs the agent, and exits non-zero if any eval is below its threshold. An app with no eval files exits `0`. Complements (does not replace) the post-hoc scoring in `evals.ts`. Promote a surprising production run with `agent-native eval promote <runId> --write evals/from-trace.eval.ts` rather than only inspecting it. See the Evals doc. ### 4. Experiments A/B testing with sticky user-level assignment: ```ts import { insertExperiment, updateExperiment } from "@agent-native/core/observability"; const exp = { id: crypto.randomUUID(), name: "sonnet-vs-haiku", status: "draft" as const, variants: [ { id: "control", weight: 50, config: { model: "claude-sonnet-4-6" } }, { id: "treatment", weight: 50, config: { model: "claude-haiku-4-5-20251001" } }, ], metrics: ["cost", "latency", "satisfaction"], assignmentLevel: "user" as const, startedAt: null, endedAt: null, createdAt: Date.now(), }; await insertExperiment(exp); // Move it to "running" when ready to start collecting assignments. await updateExperiment(exp.id, { status: "running" }); ``` The agent loop reads active experiments via `resolveActiveExperimentConfig()` and applies the variant's `model` override automatically. Assignment uses consistent hashing — same user always gets the same variant. Compute results with `POST /_agent-native/observability/experiments/:id/results`. In production, experiment management routes require the caller's email in the comma-separated `AGENT_NATIVE_EXPERIMENT_ADMIN_EMAILS` allowlist. This gate is separate from normal app/org admin roles because an experiment affects every user in that deployment. ### 5. Dashboard `ObservabilityDashboard` is a React component with 5 tabs: - **Overview** — metric cards (runs, cost, latency, tool success, thumbs up rate, eval score) - **Conversations** — trace list with drill-down to span detail - **Evals** — eval stats and criteria breakdown bars - **Experiments** — experiment list with status badges, drill-down to results - **Feedback** — feedback stream, thumbs ratio, category badges Add a dashboard route to any template: ```tsx // app/routes/observability.tsx import { ObservabilityDashboard } from "@agent-native/core/client/observability"; export default function ObservabilityPage() { return ( <div className="min-h-screen bg-background p-6"> <ObservabilityDashboard /> </div> ); } ``` ## API Endpoints All auto-mounted at `/_agent-native/observability/*`: | Method | Path | Purpose | |--------|------|---------| | GET | `/` | Overview stats | | GET | `/traces` | List trace summaries | | GET | `/traces/:runId` | Trace detail (summary + spans) | | GET | `/traces/:runId/evals` | Evals for a run | | POST | `/traces/:runId/promote` | Promote a completed run into a CI eval case | | POST | `/feedback` | Submit feedback | | GET | `/feedback` | List feedback entries | | GET | `/feedback/stats` | Feedback aggregation | | GET | `/satisfaction` | Satisfaction scores | | GET | `/evals/stats` | Eval statistics | | POST | `/experiments` | Create experiment | | GET | `/experiments` | List experiments | | GET | `/experiments/:id` | Experiment detail | | PUT | `/experiments/:id` | Update experiment status | | POST | `/experiments/:id/results` | Compute experiment results | | GET | `/experiments/:id/results` | Get experiment results | All endpoints support `?since=N` (ms timestamp) and `?limit=N` query params. ## SQL Tables 9 tables created automatically via `ensureObservabilityTables()`: - `agent_trace_spans` — individual trace spans - `agent_trace_summaries` — aggregated run summaries - `agent_feedback` — explicit user feedback - `agent_satisfaction_scores` — computed frustration index - `agent_evals` — evaluation results - `agent_eval_datasets` — golden test datasets - `agent_experiments` — experiment definitions - `agent_experiment_assignments` — user → variant assignments - `agent_experiment_results` — computed metric results All tables are PostgreSQL-compatible and strictly additive. ## Key Files | File | Purpose | |------|---------| | `packages/core/src/observability/types.ts` | Shared type definitions | | `packages/core/src/observability/store.ts` | SQL tables + CRUD | | `packages/core/src/observability/traces.ts` | Auto-instrumentation | | `packages/core/src/observability/posthog-ai.ts` | `$ai_trace` / `$ai_span` / `survey sent` emission, content bounding, `$ai_error` | | `packages/core/src/observability/failure-context.ts` | The server failure packet (`buildFailureContext`, `withFailureContext`), applied by `captureError()` |
GitHubで見る
この SKILL.md は非常に大きいため、SkillsMP では最初のセクションだけを表示しています。 GitHubで見る