- name
- foundry-evals
- description
- Evaluate Foundry agents with agent-target runs or an explicit invoke+score fallback. Covers sequential invocation, cold-start handling, dataset creation, evaluator configuration, RBAC for eval judges, result interpretation, and cross-refs to community evaluator frameworks (foundry-assert) and eval-driven prompt optimization loops (foundry-agent-optimizer). USE FOR: evaluate agent post-deploy, score grounding quality, measure tool_selection, detect dataset drift, write custom grader, check URL citations, validate eval RBAC, day-1 smoke test, continuous eval loop, pre-merge eval gate, Foundry Evals SDK setup, evaluator framework, eval-driven optimization, ASSERT evaluators. DO NOT USE FOR: deploying agents (use threadlight-deploy), designing processes (use threadlight-design), unit testing code, reimplementing evaluator framework (use foundry-assert), writing your own optimizer loop (use foundry-agent-optimizer).
- metadata
- {"version":"1.4.2"}
# Foundry Agent Evaluations
Evaluate Foundry agents using the supported **agent-target** evaluation surface,
or an explicit **invoke+score** fallback with captured responses.
For the complete per-agent adoption and release-evidence workflow, see
[`foundry-agentops`](../foundry-agentops/SKILL.md). This skill remains
authoritative for deep evaluator configuration and dataset design; AgentOps
aggregates evidence and never replaces that evaluation contract.
## When to Use
- After deploying a hosted agent, to measure quality
- Running batch evaluations against test scenarios
- Comparing agent performance across versions
- Validating that business rules (BR-XXX from SpecKit) are followed
## Why Two Phases?
Two phases are a fallback, not a universal requirement. Current
[cloud evaluation guidance](https://learn.microsoft.com/azure/foundry/observability/how-to/cloud-evaluation-targets)
supports `azure_ai_target_completions` with an `azure_ai_agent` target for prompt
and hosted agents. Select the path from the actual endpoint/protocol:
| Surface | Path |
|---|---|
| Supported prompt or hosted Responses agent | Agent-target run: query-only input, explicit agent name/version, `{{sample.output_text}}` for string evaluators or `{{sample.output_items}}` for interaction-aware evaluators. |
| Hosted Invocations agent | Use the documented freeform `input_messages` matching its real request contract; do not send a Responses template. The smoke helper below intentionally covers Responses only. |
| Existing captured responses, custom transport, or a confirmed unsupported target surface | Invoke explicitly, retain the real response/tool transcript, then score `{{item.response}}` as described in Phases 1 and 2 below. |
Start with one approved synthetic query. A created run is not success: poll to a
successful terminal state, download per-item `output_items`, and require actual
scores and usable responses. An auth, schema, model, or quota failure is **not**
permission to silently switch paths. Record the failure; use invoke+score only
when the surface requires it or the owner approves that fallback.
> **Canonical executable:** [references/python/eval_runner.py](references/python/eval_runner.py)
> implements a one-item agent-target or explicit invoke+score coherence smoke
> using `AIProjectClient.get_openai_client().evals`. See the
> [Day-1 recipe](#day-1-smoke-test-recipe-hosted-agent--mcp-tool-pilots).
> It does not replace full suites, custom graders or the EVAL-201 adapter.
```
Phase 1: Invoke agent → collect responses
Phase 2: Score responses → Foundry evaluators
```
**When using the fallback:** complete both phases sequentially; never score an
invented response or treat a run ID as the agent's response.
---
## Phase 1: Invoke the Agent
```python
from azure.ai.projects import AIProjectClient
from azure.identity import DefaultAzureCredential
project = AIProjectClient(
endpoint="<project_endpoint>",
credential=DefaultAzureCredential(),
allow_preview=True,
)
oai = project.get_openai_client(agent_name="my-agent")
# Warm up — single-shot ping is INSUFFICIENT for hosted agents that
# scale to zero (15min idle). Use the retry loop pattern below for any
# eval likely to hit a cold container. A single ping can return
# server_error in ~9s and make every scenario fail before the agent
# has even spun up.
print("Warming up...")
oai.responses.create(input="Hello", stream=False)
# MUST invoke each query SEQUENTIALLY (never concurrent)
# Concurrent requests overwhelm cold-start containers → empty responses
results = []
for query in test_queries:
response = oai.responses.create(input=query, stream=False)
results.append({
"query": query,
"response": response.output_text,
})
print(f"✓ {query[:50]}...")
```
### Critical Rules
- **DO NOT invoke concurrently** — concurrent eval requests overwhelm cold-start containers
and produce empty responses. Sequential invocation is mandatory.
- **MUST warm up first** — send a throwaway "Hello" before the real queries using the retry
loop pattern below (not a single ping). Cold containers return `server_error` in 5-10s
before the platform has brought a replica up.
- **DO use `stream=False`** — simpler for eval and required for batch scoring.
- **MUST bind to agent endpoint** — use `get_openai_client(agent_name=...)` to route to
the dedicated endpoint, not the project endpoint.
- **Token refresh for long runs** — `DefaultAzureCredential` tokens expire after ~1h. For 30+ scenarios, refresh the token every 10 items or create a fresh client per batch.
- **MUST pace invocations with 5s delay** — add `time.sleep(5)` between calls. Rapid-fire
requests produce empty responses even after warm-up.
- **DO NOT poll container status API** — do NOT call `.../versions/{v}/containers/default`
to check readiness on refreshed Foundry. This endpoint returns **HTTP 404** on modern hosted
agents (the platform auto-provisions containers). Polling wastes 3+ minutes. Use the warmup
chat loop below instead.
```python
# Warmup with retry loop — handles scale-from-zero hosted agents
WARMUP_ATTEMPTS = 4
WARMUP_BACKOFF_S = 60
print(f"[warmup] coldstart loop (up to {WARMUP_ATTEMPTS} retries with {WARMUP_BACKOFF_S}s backoff)...")
for attempt in range(1, WARMUP_ATTEMPTS + 1):
try:
r = oai.responses.create(input="ping", stream=False)
status = getattr(r, "status", "unknown")
print(f"[warmup] attempt={attempt} status={status}")
if status == "completed":
print("[warmup] READY -- proceeding with eval")
break
except Exception as e:
print(f"[warmup] attempt={attempt} EXC {type(e).__name__}: {str(e)[:120]}")
if attempt < WARMUP_ATTEMPTS:
time.sleep(WARMUP_BACKOFF_S)
else:
raise RuntimeError("Hosted agent failed to warm up after 4 attempts (4 minutes). Check the deployment.")
```
- **DO use ASCII-only logging on Windows** — see `Eval scripts on Windows: cp1252 trap` below.
The default Windows console encoding is **cp1252**, not UTF-8. Any `print('→')`, `print('×')`,
or `print('·')` in `run_evals.py` fails with `UnicodeEncodeError` mid-run, killing partial
results. Use `->`, `x`, `::` instead (or set `PYTHONUTF8=1` in the venv bootstrap).
- **MUST tolerate gateway flake on Phase 1** — Foundry's gateway can enter 5-10 minute sticky
`internal_server_error` windows under burst load (especially mid-cold-start). 30-60s exponential
backoff is **not** enough. See `Gateway flakiness during Phase 1` below for the resume-after-cooldown pattern.
- **DO retry on empty response** — `output_text` can return empty when the agent does tool calls
but the response structure varies. Retry once after a 3s pause. Also scan `response.output`
items for message text as a fallback:
```python
text = response.output_text or ""
if not text and response.output:
for item in response.output:
if getattr(item, "type", "") == "message":
for content in getattr(item, "content", []):
text += getattr(content, "text", "")
```
### Alternative: Direct SSE Invocations (GHCP SDK agents)
If the agent uses the Invocations protocol (GHCP SDK), you can't use
`oai.responses.create()`. Use the raw SSE endpoint instead:
```python
import aiohttp
url = f"{endpoint}/agents/{agent_name}/endpoint/protocols/invocations?api-version=v1"
async with aiohttp.ClientSession() as session:
async with session.post(url, json={"input": query}, headers={
"Authorization": f"Bearer {token}",
"Content-Type": "application/json",
"Foundry-Features": "HostedAgents=V1Preview",
}) as resp:
full_text = ""
async for line in resp.content:
line = line.decode().strip()
if line.startswith("data: "):
event = json.loads(line[6:])
if event.get("type") == "assistant.message_delta":
full_text += event.get("data", {}).get("content", "")
```
> **Prefer the Responses API pattern** (agent-bound OpenAI client) unless the agent
> specifically only supports Invocations. It's simpler and handles conversation state.
### Eval scripts on Windows: cp1252 trap
The default Windows Python console uses **cp1252**, not UTF-8. Any
non-Latin1 character in a `print(...)` call mid-eval produces:
```
UnicodeEncodeError: 'charmap' codec can't encode character '\u2192'
in position 14: character maps to <undefined>
```
…and kills the eval run. We've burned an hour on this twice. Two
defenses, in priority order:
**1. ASCII-only logging in `eval/run_evals.py`** (the right fix)
```python
# BAD — fails on Windows cp1252
print(f" -> {ms}ms · {len(text)} chars · attempt {attempt}")
# GOOD — works on every platform
print(f" -> {ms}ms :: {len(text)} chars :: attempt {attempt}")
```
Banned characters: `→ × · ✓ ✗ ❌ ✅ ▶ ▲ ▼ ← ° €` and any em-dash /
en-dash. Replacements: `-> x :: PASS FAIL ! [done] etc.`
**2. The agent's response can ALSO contain non-ASCII** (the gap that bit
us twice in pilot runs even with #1 in place)
Banning Unicode in your own `print()` calls only solves half the
problem. The hosted agent is free to emit `→`, em-dashes, smart quotes,
and curly arrows in its tool-result tables — and the moment your eval
prints `output_first_400={out_text!r}`, the **agent's** `→` blows up
the same way. Two layers, both required:
```python
# AT THE TOP OF run_evals.py — wrap stdout/stderr to silently
# replace any byte that doesn't fit the cp1252 console
import io, sys
sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding="utf-8",
errors="replace", line_buffering=True)
sys.stderr = io.TextIOWrapper(sys.stderr.buffer, encoding="utf-8",
errors="replace", line_buffering=True)
def safe(s, n=None):
"""Strip everything not ASCII before printing agent text."""
if s is None:
return ""
if n:
s = s[:n]
return s.encode("ascii", errors="replace").decode("ascii")
# Then use safe() on every agent string you print:
print(f" output_first_400: {safe(out_text, 400)!r}")
```
`io.TextIOWrapper(.., errors='replace')` is **the** fix for this — it
turns any otherwise-fatal `UnicodeEncodeError` into a `?` substitution
silently. Without it, even one `→` in one tool-result line aborts the
whole eval and you lose all results-so-far (see § Incremental result
writes below for the corollary mitigation).
**3. PYTHONUTF8=1 in the venv bootstrap** (defense in depth)
If you can't audit every `print` (e.g., a vendored library prints
arrows), force the interpreter into UTF-8 mode at process start. This
MUST be in the parent shell **before** Python launches:
```powershell
# scripts/setup_eval_env.ps1
$env:PYTHONUTF8 = "1"
$env:PYTHONIOENCODING = "utf-8"
uv venv .venv
.\.venv\Scripts\Activate.ps1
uv pip sync requirements.txt
```
```bash
# scripts/setup_eval_env.sh
export PYTHONUTF8=1
export PYTHONIOENCODING=utf-8
uv venv .venv
source .venv/bin/activate
uv pip sync requirements.txt
```
Setting these inside the script (`os.environ["PYTHONUTF8"] = "1"`)
**does not work** — by the time the assignment runs, Python's stdio
encoding is already locked. (Defense #2 above DOES work mid-script
because it rebuilds the wrapper from raw `sys.stdout.buffer`.)
The same trap kills `az acr build` log streaming on Windows; fix
there is `--no-logs` (separate skill: `foundry-mcp-aca`).
### Incremental result writes (mandatory default, not just for big batches)
The "Resume-after-cooldown" pattern in the next section is framed as
"recommended for batches > 5 scenarios." Treat that as a floor, not a
ceiling — **every** `run_evals.py` should write its results JSON
**after each scenario completes**, not at the end of the loop. Three
real things that have wiped a full eval batch this pilot cycle:
1. `UnicodeEncodeError` on the agent's response (the trap above) —
killed agent versions mid-scenario every cold-start run, lost 4 prior results.
2. Foundry gateway sticky 5xx window — kills attempt N, you lose 1..N-1.
3. Ctrl+C / shell window closed by accident — same outcome.
```python
from pathlib import Path
RESULTS_PATH = Path("eval/results.json")
RESULTS_PATH.parent.mkdir(parents=True, exist_ok=True)
results = []
for sc in selected:
rec = run_one(sc)
results.append(rec)
# Write after EVERY scenario, not just at the end of the loop
RESULTS_PATH.write_text(json.dumps(results, indent=2), encoding="utf-8")
```
The cost is one tiny disk write per scenario (~5ms). The benefit is
that a kill at scenario 4 of 5 still leaves you with 4 scored results
on disk, ready to rerun-the-tail on top.
### Gateway flakiness during Phase 1
Foundry's gateway occasionally enters 5-10+ minute sticky windows of
`internal_server_error` (or `503 model overloaded`) — typically right
after a cold-start burst, or when an upstream model deployment is
being reconfigured. The 30-60s exponential backoff most retry libraries
ship with is **not enough** on bad days.
Two patterns:
**Pattern A — Resume-after-cooldown (recommended for batches > 5
scenarios).** Persist results-so-far to disk after every successful
case, and add a `--resume` flag that skips already-completed `case_id`s.
Then a sticky-flake window just means "wait 10 min, re-run; it picks
up where it left off". `eval/run_evals.py` should look like:
```python
results_path = Path("eval/results.partial.json")
done_ids = set()
if results_path.exists() and "--resume" in sys.argv:
done_ids = {r["case_id"] for r in json.loads(results_path.read_text())}
for case in cases:
if case["case_id"] in done_ids:
print(f" [skip] {case['case_id']} already done")
continue
try:
result = invoke_with_retry(case, max_retries=3, base_delay=60)
except GatewayFlakeException:
print(f" [defer] {case['case_id']} -- gateway sticky, re-run with --resume")
break # don't burn budget
append_result(results_path, result)
```
**Pattern B — Skip-and-mark.** If the SLA is "best 5 of 6" rather than
"all 6", record gateway failures as `status: GATEWAY_FLAKE` (NOT
`FAIL`), exclude them from scoring, and report the count in the run
summary. Don't let infrastructure noise pollute the agent-quality
signal.
> Both patterns are about **separating gateway flake from agent
> quality**. A failed retry on `internal_server_error` is not an
> eval-quality failure; it's a Foundry-side incident. Score them
> separately or you'll spend a day debugging a "regression" that
> turns out to be a 7-minute gateway window.
---
## Phase 2: Score with Foundry Evaluators
Foundry evals have two concepts:
- **Eval definition** — the evaluator configuration (metrics, data schema, judge model).
Create this **once**, reuse across runs. Only recreate when changing metrics.
- **Eval run** — a single execution against a dataset. Create a new run for each
test cycle, agent version, or dataset update.
GitHubで見る