| name | agento11y-prod-setup |
| description | Sets up production evaluation and guardrails for a DEPLOYED AI agent in Grafana Agent Observability, grounded in the agent's own code and its real ingested traffic. The judgment layer on top of the `agento11y` skill: it reads the agent's source (system prompt, tools, entrypoint) AND samples its live traffic via gcx, checks what evaluators/rules/guards already exist, then recommends only what's missing — online eval rules (score live conversations for regressions) and guards (warn-first request-path policies that redact / tool-filter and may later be promoted to deny). It drafts reviewable YAML and, only with explicit confirmation, applies via `gcx agento11y`. New guards are drafted in warn mode (safe on live traffic — warn records but never blocks). It DOES create stack-level objects — that is the point — but every write is confirmed. It never rewrites or redeploys the agent. Trigger on phrases like "set up production evaluation", "my agent is in prod what should I evaluate", "catch quality regressions", "add guardrails to my agent", "redact PII from my agent", "block dangerous tools", "set up online evals and guards".
|
| allowed-tools | Bash, Read, Write, Edit |
Agent Observability — production evals & guards setup
The production counterpart to agento11y-test-starter (which runs pre-ship, on code alone,
producing an offline test suite). This skill runs after ship, when the agent has real
traffic, and sets up the two production surfaces the starter deliberately leaves out:
- Online eval rules — evaluators that score ingested live conversations, so regressions
surface without hand-reviewing every conversation.
- Guards (hook-rules) — policies on the request path that
warn (and can later be promoted to
deny) in real time. A guard decides via one of three shapes: evaluator_ids (an evaluator
judges), redact (regex redaction), or tool_filter (block tool calls). See Step 4.
What this skill does that agento11y doesn't
The sibling agento11y skill is the mechanics layer: exact CLI flags, evaluator/rule YAML
shapes, create-or-update semantics, the online-eval setup steps. It assumes you already know
what to create.
This skill is the judgment layer. It answers which rules and guards this specific agent
needs, by grounding in two evidence sources a generic checklist can't use:
- The agent's code — its system prompt, tools, and how it handles user data. Half the value
is here; read and cite it (
file:line).
- The agent's real traffic — because it's deployed, you can see what it actually does in
prod, not just what the code says it might.
Two gaps this skill fills beyond agento11y:
- Recommendation from evidence —
agento11y starts once you know what to create; this decides.
- Guards —
agento11y documents evaluators and rules but not guards (hook-rules), even though
gcx agento11y guards exists. This skill carries the guard shapes
(evaluator_ids / redact / tool_filter, plus action_on_fail) itself.
For any mechanical detail — exact flags, evaluator/rule YAML fields, the setup flow — defer to
the agento11y skill and to gcx agento11y <sub> --help rather than restating it here.
Rules
- Every connection to the stack goes through
gcx agento11y — never raw HTTP, never a hand-held
token. gcx owns Cloud auth (via gcx login). Prerequisite: gcx installed and authenticated;
if it isn't, say so and stop.
- Confirm the target stack before any WRITE (Step 0 + Step 5). Reads run freely once you've
shown the context; writes (upsert evaluators, create/update rules and guards) need an explicit yes on the
target stack.
gcx may be pointed at the wrong stack, and this skill creates stack-level
objects.
- Check before recommending. Always list what already exists first
(
gcx agento11y evaluators list, rules list, guards list) and never recommend a duplicate.
Compare by semantic equivalence, not just id/name — see Step 2.
- This skill does create stack-level objects — that is its job, the one thing that separates
it from
agento11y-test-starter. But every creation is explicit and confirmed: show the exact
YAML, get a yes, then create it with the matching gcx agento11y command. A yes for one object
is not a yes for the next.
- New guards are always drafted
action_on_fail: "warn" — even hard-policy ones. That is what
makes a new guard safe: a warn guard only records the outcome, it never blocks a request, so it
is harmless even while it is live. Never draft a first-time deny guard. The developer switches to
deny themselves, later, after watching it in warn mode (Step 6). Draft enabled: false if you
can, but don't build extra steps around it — the server may store the guard enabled: true
regardless (see Step 5), and that is fine, because warn carries the safety, not enabled.
- New online rules start with a conservative
sample_rate (e.g. 0.1), not 1.0 — an
llm_judge over 100% of traffic costs real money.
- Prefer starting from an evaluator template (
gcx agento11y templates list, then
gcx agento11y templates get) over authoring
a new evaluator. Only write a fresh one when nothing fits.
- Do not rewrite the agent's prompt, optimize, or redeploy. This skill configures observation and
guardrails around the agent, not the agent itself.
Step 0 — Confirm the target stack
Before reading traffic or writing anything, show the developer where gcx is pointed. The active
context may not be the stack they think:
gcx config current-context # the active context name
gcx config view # its server URL, org-id, auth method
Display the resolved context name, server URL, and org-id. Two thresholds:
- Before reads (traffic sampling, inventory): show the resolved context so the developer sees
where the discovery ran. A wrong stack here wastes effort but doesn't change anything.
- Before any write (Step 5): require an explicit yes that this is the intended production
stack. Writes are what create stack-level objects, so this confirmation is the hard gate.
If it's wrong, stop — the developer switches with gcx config use-context <name>, or you pass
--context <name> on every gcx call. (Watch for localhost / dev-looking servers — a strong
sign the active context is not their prod stack.)
Step 1 — Read the code and sample real traffic
Two evidence sources. Do both; every later recommendation cites one of them.
Code (as agento11y-test-starter Step 1). Find and record file:line for: the entrypoint, the
system prompt, the tool/function definitions, and how it handles user data. This tells you what
could go wrong. The code is the authoritative source for the system prompt and tools —
content capture is often off in production, so the ingested traffic frequently has an empty
system_prompt and may omit tool definitions. Never conclude "this agent has no system prompt"
from the traffic; read it from the code.
Traffic, via gcx — this tells you what does go wrong:
- Find the agent as Agent Observability sees it:
gcx agento11y agents list (and agents get /
agents list-versions) to get the exact agent_name — this is the match.agent_name you'll target. (Tip:
agents list prints a leading hint line before the JSON; set GCX_AGENT_MODE=true or skip
that line if you parse it.)
- Sample recent conversations:
gcx agento11y conversations search --filters 'agent = "<name>"'
(add status = "error", time windows, tool.name, eval.passed = false) and
gcx agento11y generations get <id> for detail. Look for long tool loops, over-refusals, PII
echoed back, off-topic drift, malformed outputs, error clusters.
- Some agents have generations but no conversation (e.g. single-shot agents whose spans
aren't grouped into a conversation). If
conversations search is empty but agents get shows a
non-zero generation_count, don't stop and conclude "no traffic, wrong skill" — sample the
generations directly (gcx agento11y generations get <id>); there's plenty to work from.
- Pick the selector from how the agent produces generations, not by habit. The symptom to
recognize: a rule that matches traffic but whose scores stay at zero — it looks live but
silently never fires. That happens when
selector: user_visible_turn is used on an agent whose
generations aren't user-facing turns: a multi-agent DAG / programmatic pipeline
(fan-out/fan-in, internal nodes, one conversation per run), or a single-shot agent with no
conversation. (The underlying reason is that those generations carry no user-visible-turn flag,
but there's no gcx command to inspect that directly — diagnose from the zero-score symptom,
not by hunting for the field.) For these agents use selector: all_assistant_generations
and scope with match.agent_name to the node you care about. Reserve user_visible_turn for
genuine chat/assistant agents where a turn is what the user sees. To confirm a selector is
working, score over enough sampled traffic — either wait for the rule's sample_rate to
accumulate hits, or temporarily bump it (and drop it back afterward; sample_rate: 1.0 costs
real judge money, see the rule rules below) — then check conversations search. A run that
scored nothing shows (the field is omitted when
nothing scored, not returned as ); persistent zero on a matching agent almost always means
the selector is wrong for this agent shape.
Minimum evidence bar. Aim for ≥20 recent conversations over ≥7 days before drafting
anything. Fewer than that and you risk overfitting one odd conversation into a production rule or
guard: if the window is thin, either stop and say so, or proceed but mark every recommendation
low-confidence and lean on non-intervening setups (guards in warn, low sample_rate) rather
than anything that blocks. A recommendation from a single conversation is a hypothesis, not a rule. If the
agent has essentially no traffic, stop — this is the wrong skill; agento11y-test-starter (offline
suite) is the right one until traffic exists.
Step 2 — Inventory what already exists
Before proposing anything: gcx agento11y evaluators list, gcx agento11y rules list,
gcx agento11y guards list (add -o yaml to see full definitions). For each concern you're about to
raise, decide whether something already covers it — by semantic equivalence, not id or name.
Two objects are effectively the same when they share:
- the same surface (both rules, or both guards),
- the same target (overlapping
match, especially agent_name / selector),
- the same intent (evaluator kind + what it checks, or the guard's policy),
- the same action (rule scoring vs. guard
warn/deny/transform/tool_filter).
If an existing object matches on those, don't create a second one — say it's covered and stop, or
propose an update to the existing one. A different id over identical intent+target+action is a
duplicate, and duplicate guards/rules double cost and can conflict.
Step 3 — Recommend rules and guards
Map each observation to the surface that fits. Online rules observe (score, detect
regressions, no user impact, no agent code change — the eval worker scores ingested traffic
asynchronously); guards intervene (block/redact on the live request path, and require a
code change in the agent to call them — see Step 5.5). Pick the surface by whether you want to
watch or to stop — and remember a guard is dead config until the agent is wired to it.
For every online rule row below, pick the selector per Step 1 (agent shape), not by the
template's default — user_visible_turn for a genuine chat agent, all_assistant_generations for a
DAG/pipeline node or single-shot agent.
| If, in code or traffic, the agent… | Surface | Shape (prefer a predefined template) |
|---|
| gives answers whose quality can drift | online rule | fork template.helpfulness / template.relevance (llm_judge) |
| does RAG / cites sources | online rule | fork template.groundedness (llm_judge) |
| must emit JSON / a fixed shape | online rule | fork template.json_valid (json_schema) |
| over-refuses or drifts off-topic | online rule | regex / llm_judge on all_assistant_generations |
| public-facing text | online rule | fork template.toxicity / template.pii (llm_judge) |
| receives PII/secrets in its input (pasted into the prompt) | guard | redact guard, preflight — rewrites the request before the model sees it |
| generates PII/secrets/destructive commands in its final output | guard | evaluator-backed detector (regex / llm_judge), postflight, warn — don't rely on redact to scrub already-generated assistant text; detect + warn instead |
| can call dangerous tools (shell, delete, write) | guard | tool_filter with blocked_names globs, postflight |
| is subject to prompt-injection / hard policy | guard | llm_judge evaluator; draft warn, later promotable to deny |
Three guard shapes, all first-class: evaluator_ids (an evaluator decides — most flexible),
redact (regex redaction), and tool_filter (block tool calls). Pick the one that fits the policy;
a redaction need is a redact guard, not a judge.
redact schema is exact — get it right or the create 400s. The field is redact (the server's
canonical name; a legacy transform alias also works), and each pattern is only {id, regex} —
there is no replacement key. On a match the server redacts to a placeholder derived from the
pattern id; supplying replacement fails with 400 unknown field "replacement". Draft it as:
redact:
patterns:
- id: bearer_token
regex: 'Bearer\s+[A-Za-z0-9._-]+'
After creating any guard, guards get -o yaml and confirm the spec stored what you sent (a clean
create echo is not proof on its own).
Template ids above are the expected global blueprints — they can vary by deployment and
version, so always resolve the current set with gcx agento11y templates list before using a name;
don't trust a hardcoded id. Pick 3–6, ranked. Each gets a one-line why citing a file:line or a
conversation/generation id from Step 1. For rule mechanics (selectors, match keys, evaluator
kinds, templates), the agento11y skill is the reference — don't restate it here.
Every candidate lands in exactly one of three states — and the decision is final for this run:
- Recommended — worth setting up now; goes to Step 4 (draft) and Step 5 (apply).
- Considered, not recommended — you evaluated it and it's not worth it (low value, no
evidence, would just add cost/noise). Record it with a one-line why, and do NOT draft or apply
it. Don't quietly re-add it later under a different framing — if you're tempted to, it belongs
in Recommended, so put it there and own the reasoning. Bias toward fewer objects: recommend only
what earns its place. "It's harmless in warn mode" is not a reason to create something — an
unused guard/rule is still cost and noise.
- Skipped (duplicate) — the stack already covers it (Step 2 semantic-equivalence check);
don't create a second one.
The set you draft in Step 4 and apply in Step 5 is exactly the Recommended list — nothing
from the other two states leaks in.
Step 4 — Draft the definitions as YAML
First, check whether ./agento11y-prod/ already has drafts from an earlier run (ls agento11y-prod/** if it exists). If it does, read them before writing — don't blind-overwrite
(a plain Write over an existing file also just errors). Treat a prior draft as a peer proposal:
reconcile rather than replace. If a previous run made a deliberate, well-reasoned choice — e.g.
dropped an email-redaction pattern because "this agent's job is to email people, so redacting
every address is all-false-positive noise" — that judgment is usually right; keep it and fold in
only what's genuinely new. Overwrite a prior draft only when yours is clearly better, and say why.
Write the definitions to a local scratch directory, ./agento11y-prod/
(evaluators/<id>.yaml, rules/<id>.yaml, guards/<id>.yaml). These are working drafts, not
committed artifacts: they exist so the developer can review a diff before you apply it, and
their source of truth after apply is the stack, not the repo. Add agento11y-prod/ to .gitignore
(or write under the OS temp dir) so they aren't accidentally committed — they hold the applied
config redundantly and can carry regexes/prompts the repo shouldn't own. They are exactly what
you'll pass to gcx agento11y <kind> create -f (for evaluators: upsert -f). Use the
top-level-fields YAML shape that the create -f/upsert -f commands expect (not the
apiVersion/kind/spec manifest that the get -o yaml commands emit — don't round-trip get
output into create).
Rules and evaluators: follow the agento11y skill's input format exactly. Start an evaluator
from a template (gcx agento11y templates get <id> -o yaml), give it your own evaluator_id, and
always include a version — it is required on create (a date like "2026-07-15" or a semver
works; existing evaluators use dates). Omitting it fails with version is required. Rule starts
enabled at a low sample_rate.
Guards — the shape the agento11y skill omits (gcx agento11y guards create -f guard.yaml; the
resource Kind is HookRule).
This skill (not the agento11y skill) carries the guard shape — draft the guard file directly from
it, no schema-discovery step needed. A guard drives its decision from one of three (mutually
exclusive) shapes — evaluator_ids, redact, or tool_filter — plus action_on_fail, phase,
priority, selector. Use the redact schema exactly as the callout above shows ({id, regex},
no replacement). If the server ever rejects a field, the 400 names the offending field, so a bad
shape surfaces at create time rather than needing a probe up front. On create the server fills
defaults you don't set — notably selector: all, phase: preflight, and short_circuit: false —
so a guards get -o yaml right after create
shows more fields than you sent; that's expected, not drift. A guard is drafted in warn (and
enabled: false if the server honors it — it may not; see Step 5):
rule_id: guard.<agent>.<policy>
enabled: false
priority: 10
action_on_fail: warn
evaluator_ids: ["<policy-judge-id>"]
phase is preflight (default) or postflight, and the correct phase depends on the guard
type — getting this wrong yields a guard that runs but silently does nothing:
redact → normally preflight (input). Use it to redact sensitive values in the request
input (messages) before the model sees them — e.g. a secret pasted into the user prompt or
incident text. This is the common case and the one the Grafana UI defaults a Redact guard to.
Do not rely on redact to rewrite already-generated assistant text: for a secret the model
produces in its final response, a postflight redaction of the assistant text does not reliably
scrub it — use a postflight evaluator-backed detector (warn/deny) instead (below).
(Postflight redact is not useless in general — some runtimes explicitly consume
transformed_input.output to rewrite tool-call payloads/arguments postflight — but that is a
narrow tool-arg case, not final-response scrubbing.)
- Disambiguation — decide by where the sensitive value ENTERS, not by where it's most visible.
An agent can both receive a secret in its input AND echo/generate one in its output; don't
let the more eye-catching output occurrence pull you to postflight. If the secret enters through
the input (e.g. a key pasted into the prompt / incident text / a config blob), the guard that
actually protects it is
redact preflight on the input — that scrubs it before the model,
the logs, and every downstream node see it. A postflight guard cannot undo a secret that already
entered upstream. Add a postflight detector on top only if the model also independently
produces secrets in its final text; that is a second, separate concern, not a replacement for the
preflight redact.
tool_filter → postflight. It inspects the tool calls the model wants to make.
evaluator_ids (detector) → phase follows what you evaluate. preflight to judge the
input (block a bad request before spending tokens), postflight to judge the output
(catch a bad response — e.g. detect that the model generated a credential or a destructive
command). This is the right shape for "the model produces something risky," which redact cannot
catch.
Always draft action_on_fail: warn, even for
a hard-policy guard (prompt-injection, deny-list): a first-time deny enabled on live traffic
blocks real users on a false positive. A policy-judge guard references an evaluator id (create the
evaluator first); the developer changes it to deny only later, after watching the false-positive
rate in warn mode (Step 6). This skill never drafts an enabled deny guard.
Step 5 — Confirm, then apply with gcx
upsert/create/update write to the stack — never run them before the developer's explicit yes
(step 2). The one thing you CAN run before the yes is evaluators test -f <request>.yaml,
which tests a judge config without persisting it (pass kind, config, output_keys,
generation_id in the file — no evaluator need exist yet). Use it to tune the judge (step 1).
There is no CLI dry-run for rules or guards — their safety comes from shipping guards in
warn (records, never blocks) and rules at a low sample_rate, not from a preview.
Per object the developer wants, in dependency order (evaluators → rules → guards, since a
rule/guard referencing an evaluator needs it to exist first):
- For an
llm_judge evaluator, tune the prompt before creating it — this is the real work,
not a formality. A judge is only as good as its prompt, so don't create it on first draft and
move on. Loop: pick 1–2 real generations you know the right answer for
(gcx agento11y generations get <id>), run the draft config with
gcx agento11y evaluators test -f <request>.yaml -g <gen-id> (tests without persisting), and
read both the verdict AND the rationale. If either disagrees with what you expected, adjust
the system_prompt/user_prompt and re-test. Repeat until verdict and rationale both hold on
your known examples. Only then does the evaluator move to step 2. (This tuning loop deserves its
own dedicated flow; keep it lightweight here.)
- Confirm. Restate the target stack from Step 0 (context name + server), show the exact YAML,
and get an explicit yes for that object. A yes for one object is not a yes for the next. Nothing
is written before this yes.
- Apply via gcx, only after the yes:
gcx agento11y evaluators upsert -f evaluators/<id>.yaml,
then gcx agento11y rules create -f rules/<id>.yaml, then
gcx agento11y guards create -f guards/<id>.yaml. Evaluators are create-or-update (same id
updates). Pass --context <name> on every call if the confirmed stack isn't the default
context. gcx handles auth — no tokens here.
- The server may store a new guard
enabled: true / short_circuit: true regardless of what
you drafted. That is a known server default, not a draft error. Do not build a
create→get→update dance around it: because you drafted action_on_fail: warn, a live guard only
records outcomes and blocks no one, so enabled: true is harmless here. Create the guard and
move on. (What you must never create is an enabled deny guard — the warn rule above is what
prevents that, not the enabled value.)
- A judge-model 404 when testing/scoring is usually a stack-side misconfiguration (the stack's
judge model id is dead), not your evaluator — flag it; the online rule will hit the same broken
judge at runtime until it's fixed.
- If
gcx reports it isn't authenticated, stop and ask the developer to run gcx login; do not
fall back to raw HTTP.
Step 5.5 — Wire the agent to call the guard (guards only, REQUIRED)
Creating a guard on the stack does not make it do anything. Unlike online rules — which the
eval worker applies asynchronously to already-ingested traffic, with zero agent changes — a guard
is a synchronous request-path policy the agent must call itself. If the agent never calls the
hooks endpoint, the guard is inert: it exists, shows enabled, and never fires. That is also why
Step 6 would show no warns — nothing invoked the guard.
So every guard applied in Step 5 has a matching code change the developer owns. This skill does not
edit or redeploy the agent (see Rules), but it MUST tell the developer exactly what to add, per
guard, or the guard is dead config. Present this as part of the hand-off — don't leave the guard
looking "done" after Step 5.
At each LLM call the agent evaluates the guard via the SDK and honors the verdict. Minimal Python
(agento11y >= 0.11):
from agento11y import (
Client, ClientConfig, ApiConfig, HooksConfig,
HookEvaluateRequest, HookContext, HookModel, HookInput,
HookDeniedError, user_text_message,
)
client = Client(config=ClientConfig(
api=ApiConfig(endpoint="https://<stack>.grafana.net"),
hooks=HooksConfig(enabled=True, phases=["preflight"], fail_open=False),
))
resp = client.evaluate_hook(HookEvaluateRequest(
phase="preflight",
context=HookContext(
model=HookModel(provider="anthropic", name="claude-..."),
agent_name="<the rule's match.agent_name>",
agent_version="<v>",
conversation_id=conversation_id,
),
input=HookInput(messages=[user_text_message(prompt)]),
))
if resp.is_deny:
raise HookDeniedError(reason=resp.reason, rule_id=resp.rule_id,
evaluations=list(resp.evaluations))
Three gotchas that silently break guards — call each out to the developer:
fail_open=False — with the default True, a transport error (or a disabled/missing guard)
resolves to allow, so a deny never actually blocks. Fail-closed is what makes the guard enforce.
conversation_id in the context — without it, a deny/warn outcome is NOT persisted onto
the conversation, so it's invisible in the UI and Step 6 has nothing to watch. With it, the server
records a "Guard: " workflow step on the conversation.
allow leaves no trace — the server persists only deny and warn outcomes; a clean pass
records nothing (just a metric). So "no guard step on the conversation" is ambiguous — it means
either allow OR not-wired. Distinguish them by whether ANY guard step ever appears for that agent.
Phase choice mirrors the guard (see the guard-type rules in Step 4): redact is normally
preflight (redact the input before the model sees it; don't use it to scrub already-generated
assistant text — see Step 4 for the tool-arg exception); tool_filter is postflight.
For an evaluator-backed guard, phase follows target: target: input → preflight (evaluate
input.messages before the call — block a bad request before spending tokens); target: response
→ postflight (evaluate input.output after — catch a bad response the model produced). The
rule's match.agent_name / match.model scopes the guard to this agent; the agent's context must
send the same agent_name.
Other SDKs (JS, Go) expose the same evaluate_hook / hooks-config shape; check the per-language SDK
reference and verify the exact symbols against the installed agento11y package. (As of writing the
canonical llms.txt does not yet carry a guard-instrumentation section, so don't defer to it for
this — the shape above is the reference.)
Step 6 — Summarize and hand off
Output, in this order:
- The three states from Step 3, kept distinct: Recommended (each with its surface, evaluator
kind, and
why — file:line or conversation/generation id); Considered, not recommended
(each with its one-line why); Skipped as duplicates (what on the stack already covered it).
Don't move an item between states between Step 3 and here.
- What was created vs. left as a draft: for each Recommended object, the YAML path and whether
it was applied via gcx. (Considered-not-recommended items were never drafted, so they have no
path — don't list them here.)
- The follow-through the developer still owns:
- Guards: first, wire the agent to call the guard (Step 5.5) — until that code ships the
guard is inert and you'll see no
warns, no matter how it's configured. State this per guard,
with the exact call to add. Then watch the warn guards on real traffic and flip to deny +
enabled only once the false-positive rate looks acceptable —
gcx agento11y guards update <id> -f .... Both the wiring and the flip are theirs to make, not
this skill's.
- Rules: first confirm the rule is actually scoring — if
eval_summary is absent or
total_scores stays 0 on a matching agent, the selector is likely wrong for the agent shape
(see Step 1), not the sample rate. Once scores appear, raise sample_rate as they look sane and
cost is understood (gcx agento11y rules update); add alerting on regressions if wanted.
- Inspect everything in Agent Observability (rules/guards/evaluators pages, the conversation
Quality view) or via the
gcx agento11y list and get commands.
- A one-line pointer back: for pre-ship offline evaluation of a new agent or version,
agento11y-test-starter is the counterpart; for control-plane mechanics, the agento11y skill.