- name
- observability
- description
- Agent observability, evals, feedback, and experiments. Use when adding observability dashboards, configuring trace capture, setting up evals, creating A/B experiments, or collecting user feedback on agent responses.
- scope
- dev
- metadata
- {"internal":true}
# Agent Observability
## Rule
The observability system auto-instruments every agent run with zero configuration. Traces, automated evals, and feedback collection work out of the box. All data lives in the app's own SQL database — no external services required. Templates can optionally export to Langfuse, Datadog, or any OTel-compatible platform.
## Five Pillars
### 1. Traces
Every `runAgentLoop()` call is automatically instrumented via `instrumentAgentLoop()` in `packages/core/src/observability/traces.ts`. It captures:
- **agent_run** span — top-level parent with total duration and cost
- **llm_call** span — model name, token counts (input, output, cache read/write), cost
- **tool_call** spans — one per action invocation, with duration and success/error
The run span is NAMED by what started it, because the trace list shows that name
and nothing else: a scheduled job is `background_automation_run:<job name>`, a
chat turn is `agent_run`, and a turn a feature sent on the user's behalf is
`agent_run:<usageLabel>` when the caller named it:
```ts
sendToAgentChat({
message: "Enrich this record from the web",
newTab: true,
usageLabel: "crm:enrich-record", // → usage row label + `agent_run:crm:enrich-record`
});
```
Content (prompts, tool args, tool results) is **redacted by default**. Opt in through the declared `observability` config domain:
```ts
// server/plugins/config.ts
import { defineAppConfig } from "@agent-native/core/server";
export default defineAppConfig({
observability: {
enabled: true,
capturePrompts: false,
captureToolArgs: true, // capture action input args
captureToolResults: false, // include tool results and the full error text on tool spans and $ai_generation entries
evalSampleRate: 0.05, // 5% of runs get LLM-as-judge eval
inferredSentimentEnabled: false,
inferredSentimentSampleRate: 0,
inferredSentimentModel: "gpt-5-6-luna",
},
});
```
Two things are recorded whatever the flags say, because a failure nobody can
classify is not observable:
- **A failed tool span always keeps a signature**: the first line of the error,
redacted (credentials, emails, long opaque ids) and capped at 500
characters. `captureToolResults` only decides
whether the rest is kept. The span's metadata says which (`__tool_error_detail`:
`full` | `signature`), and the read path returns it as `errorDetail`
(`full` | `signature` | `withheld` | `unrecorded`) so "withheld on purpose" is
never read as "nothing was recorded". `$ai_*` events and OTel spans still
follow the flag.
- **A stop that waits on the user is not an error.** An `input_required`
outcome (question, approval, connection) records the `agent_run` span as
`status: "paused"` with the reason in `terminal_code`, not `error`. Anything
that counts failed runs must treat `paused` as non-error; an `error` event
that follows the pause still makes the run an error.
#### From an error report to the failing run
A report is only useful if it names where it happened. Four pieces make that
true, so nobody has to ask the reporter for an example:
- **The failure packet.** Every server `captureError()` carries
`extra.failureContext` (`app`, `route`, `actionName`, `automationName`,
`threadId`, `runId`, `requestId`, `userScope`, `threadUrl`, build,
environment, `errorCode`, `failureClass`); see the tracking skill for how it
is derived. `threadUrl` is `https://<app host>/?thread=<chat_threads.id>`, and
the Analytics issue page links it. The same fields (`$ai_trace_id` = run id,
`thread_id`) are on the `$ai_trace` for the run.
- **A copyable report on the client.** `formatClientFailureReport()` from
`@agent-native/core/client/failure-report` builds the plain-text packet (app,
thread link, run, request, code, time, build, and the inspection call) for a
"Copy details" button. The chat run-error card's existing "Copy debug info"
button uses it; call it from any other error card. It needs only what the
card already knows (`message`, `errorCode`, `runId`, `requestId`) and falls
back to the open thread. The error-screen "Open GitHub issue" template adds
the same thread link, build and time. A failed action carries `error.requestId` (the
response's `x-agent-native-request-id`, also on the server capture's
`failureContext.requestId`).
- **Read by id.** Dispatch's `get-agent-thread-debug` takes a `threadId` or the
copied run id (`run-…`) and returns the run's `terminalReason` and terminal
event, its events, trace spans (failed tools carry a signature, see above),
feedback and checkpoints; `list-agent-run-failures` finds candidates. Both are
owner-scoped. Without Dispatch, `GET /_agent-native/observability/traces/:runId`
returns the same spans for the caller's own runs.
- **Alerts name examples.** The chat-health Slack alert lists the latest failed
turns as `threadUrl (run id, code)`; if they cannot be read it says so.
#### Optional inferred sentiment
Self-hosted apps default to no inferred sentiment. First-party apps hosted on
`agent-native.com` automatically classify 100% of eligible user replies with
`gpt-5-6-luna`; an explicit stored `inferredSentimentEnabled: false` remains an
opt-out. Deployment overrides are `AGENT_NATIVE_INFERRED_SENTIMENT=on|off`,
`AGENT_NATIVE_INFERRED_SENTIMENT_SAMPLE_RATE=0..1`, and
`AGENT_NATIVE_INFERRED_SENTIMENT_MODEL=<model>`; `off` is always the emergency
kill switch.
Classification uses only the original visible user text, capped at 2,000
characters, with no tools, temperature 0, an eight-token output, and a five
second timeout. It skips attachment-only turns, internal continuations, chained
background chunks, and first turns that have no preceding response to
attribute. The managed Builder engine runs the classifier after the main
response has streamed, so it does not contend with the user's response.
Successful classifications emit a content-free `$ai_sentiment` tracking event:
- `sentiment`: `positive`, `negative`, or `neutral`
- `method`: `llm`
- `model` / `$ai_model`: model that generated the preceding assistant response
- `run_id` / `$ai_trace_id`: preceding response run
- `thread_id` / `$ai_session_id`: conversation
- `classification_trigger_run_id`: run started by the classified user reply
- `classifier_model` and `classifier_engine`: classifier attribution
No raw message, prompt, or response text is persisted or tracked.
### 2. Feedback
**Explicit** — AgentKit's assistant-message action bar renders inline thumbs up/down controls. A thumbs-down can collect a reason, and feedback includes the run and message sequence for trace linking. The shared `AgentKitAssistantChat` host submits it through the existing feedback action.
**Implicit** — `computeSatisfactionScore(threadId)` computes a Frustration Index (0-100) from conversation signals:
- Rephrasing detection (weight 30): consecutive similar user messages
- Abandonment (weight 20): session ends shortly after agent response
- Sentiment (weight 15): negative language patterns
- Length trend (weight 15): declining message lengths
- Retry patterns (weight 20): "try again", "no that's wrong"
Score interpretation: 0-20 healthy, 20-40 friction, 40-60 dissatisfied, 60+ broken.
Satisfaction scoring fires automatically after each feedback POST with a threadId.
### Human audit and refinement
For a human output audit, provide one table row per run with a distilled ask,
result, inferred or explicit sentiment, audit state, and app-native preview. The
same row contract works for standalone app and workspace roll-ups; keep
single-app surfaces independent of workspace chrome.
Treat ordinary feedback separately from admin audit verdicts. Only admins can
vote, mark audited, approve improvements, or apply them, enforced server-side
with app-scoped roles. Let admins add reasons to downvotes across many rows,
then run one synthesis that groups patterns, cites evidence, suggests update
targets (skills, instructions, memories, data dictionaries, certified
dashboards, or creative context), and shows before/after diffs. Allow feedback
and regeneration before explicit approval; never auto-apply.
The shipped single-app human review surface is the `Human review` tab in the
shared observability dashboard. It lists the persisted ask and answer for each
run, records thumbs/text feedback through the existing feedback endpoint, and
saves explicit instruction changes as `draft` rows through the
`save-observability-instruction-update` action. Drafts are reviewable artifacts;
they never update agent behavior automatically. Agents can read the same table
through `list-observability-reviews`.
### 3. Evals
Three layers, configured via `evalSampleRate` in the observability config:
**Automated (every run):** Deterministic scorers that run after every traced run:
- `tool_success_rate` — % of tool calls without errors
- `step_efficiency` — 1.0 for no-tool runs; penalizes excessive LLM iterations for tool-using runs
- `latency_score` — normalized against 10s/tool baseline
- `cost_efficiency` — normalized against 50 centicents/tool baseline
- `error_recovery` — 1.0 if the run recovered from tool errors or had none
**LLM-as-judge (sampled):** Runs on `evalSampleRate` fraction of runs. Calls the configured engine with a judge prompt that scores against custom criteria.
**Dataset evaluation:** `runDatasetEval(datasetId)` runs a golden dataset through the agent and scores each case.
Promote a surprising production run into a CI eval case (`promote-trace-eval` / `agent-native eval promote <runId> --write ...`) rather than only inspecting it. Truncated runs fail closed.
Custom criteria use natural language rubrics:
```ts
const criteria: EvalCriteria = {
name: "helpfulness",
description: "Was the response helpful and complete?",
rubric: "0.0 = completely unhelpful, 0.5 = partially helpful, 1.0 = fully resolved the user's need",
};
```
#### Evals (CI gate)
The three layers above score *real production runs* after the fact. For an active, deterministic gate, use the first-class `*.eval.ts` primitive from `@agent-native/core/eval` (source: `packages/core/src/eval/*`). It runs the actual agent loop against fixed inputs and exits non-zero below threshold, so it gates CI/deploys.
```ts
// evals/faq.eval.ts
import { defineEval, contains, llmJudge } from "@agent-native/core/eval";
export default defineEval({
name: "answers the FAQ",
input: { prompt: "What is your return policy?" },
threshold: 0.7,
scorers: [contains("30 days"), llmJudge({ criteria: "accuracy" })],
});
```
- Built-in scorers: `exactMatch` / `contains` / `usesTool` (pure JS) and `llmJudge` (provider-agnostic judge).
- Custom scorers: `createScorer` with the 4-step `preprocess → analyze → generateScore → generateReason` pipeline (only `generateScore` is required).
- Run as a gate: `agent-native eval [pattern] [--json] [--threshold N]` — discovers `**/*.eval.ts` and `evals/*.ts`, runs the agent, and exits non-zero if any eval is below its threshold. An app with no eval files exits `0`. Complements (does not replace) the post-hoc scoring in `evals.ts`. Promote a surprising production run with `agent-native eval promote <runId> --write evals/from-trace.eval.ts` rather than only inspecting it. See the Evals doc.
### 4. Experiments
A/B testing with sticky user-level assignment:
```ts
import { insertExperiment, updateExperiment } from "@agent-native/core/observability";
const exp = {
id: crypto.randomUUID(),
name: "sonnet-vs-haiku",
status: "draft" as const,
variants: [
{ id: "control", weight: 50, config: { model: "claude-sonnet-4-6" } },
{ id: "treatment", weight: 50, config: { model: "claude-haiku-4-5-20251001" } },
],
metrics: ["cost", "latency", "satisfaction"],
assignmentLevel: "user" as const,
startedAt: null,
endedAt: null,
createdAt: Date.now(),
};
await insertExperiment(exp);
// Move it to "running" when ready to start collecting assignments.
await updateExperiment(exp.id, { status: "running" });
```
The agent loop reads active experiments via `resolveActiveExperimentConfig()` and applies the variant's `model` override automatically. Assignment uses consistent hashing — same user always gets the same variant.
Compute results with `POST /_agent-native/observability/experiments/:id/results`.
In production, experiment management routes require the caller's email in the
comma-separated `AGENT_NATIVE_EXPERIMENT_ADMIN_EMAILS` allowlist. This gate is
separate from normal app/org admin roles because an experiment affects every
user in that deployment.
### 5. Dashboard
`ObservabilityDashboard` is a React component with 5 tabs:
- **Overview** — metric cards (runs, cost, latency, tool success, thumbs up rate, eval score)
- **Conversations** — trace list with drill-down to span detail
- **Evals** — eval stats and criteria breakdown bars
- **Experiments** — experiment list with status badges, drill-down to results
- **Feedback** — feedback stream, thumbs ratio, category badges
Add a dashboard route to any template:
```tsx
// app/routes/observability.tsx
import { ObservabilityDashboard } from "@agent-native/core/client/observability";
export default function ObservabilityPage() {
return (
<div className="min-h-screen bg-background p-6">
<ObservabilityDashboard />
</div>
);
}
```
## API Endpoints
All auto-mounted at `/_agent-native/observability/*`:
| Method | Path | Purpose |
|--------|------|---------|
| GET | `/` | Overview stats |
| GET | `/traces` | List trace summaries |
| GET | `/traces/:runId` | Trace detail (summary + spans) |
| GET | `/traces/:runId/evals` | Evals for a run |
| POST | `/traces/:runId/promote` | Promote a completed run into a CI eval case |
| POST | `/feedback` | Submit feedback |
| GET | `/feedback` | List feedback entries |
| GET | `/feedback/stats` | Feedback aggregation |
| GET | `/satisfaction` | Satisfaction scores |
| GET | `/evals/stats` | Eval statistics |
| POST | `/experiments` | Create experiment |
| GET | `/experiments` | List experiments |
| GET | `/experiments/:id` | Experiment detail |
| PUT | `/experiments/:id` | Update experiment status |
| POST | `/experiments/:id/results` | Compute experiment results |
| GET | `/experiments/:id/results` | Get experiment results |
All endpoints support `?since=N` (ms timestamp) and `?limit=N` query params.
## SQL Tables
9 tables created automatically via `ensureObservabilityTables()`:
- `agent_trace_spans` — individual trace spans
- `agent_trace_summaries` — aggregated run summaries
- `agent_feedback` — explicit user feedback
- `agent_satisfaction_scores` — computed frustration index
- `agent_evals` — evaluation results
- `agent_eval_datasets` — golden test datasets
- `agent_experiments` — experiment definitions
- `agent_experiment_assignments` — user → variant assignments
- `agent_experiment_results` — computed metric results
All tables are PostgreSQL-compatible and strictly additive.
## Key Files
| File | Purpose |
|------|---------|
| `packages/core/src/observability/types.ts` | Shared type definitions |
| `packages/core/src/observability/store.ts` | SQL tables + CRUD |
| `packages/core/src/observability/traces.ts` | Auto-instrumentation |
| `packages/core/src/observability/posthog-ai.ts` | `$ai_trace` / `$ai_span` / `survey sent` emission, content bounding, `$ai_error` |
| `packages/core/src/observability/failure-context.ts` | The server failure packet (`buildFailureContext`, `withFailureContext`), applied by `captureError()` |
GitHubで見る