Skip to main content

foundry-evals

Evaluate Foundry agents with agent-target runs or an explicit invoke+score fallback. Covers sequential invocation, cold-start handling, dataset creation, evaluator configuration, RBAC for eval judges, result interpretation, and cross-refs to community evaluator frameworks (foundry-assert) and eval-driven prompt optimization loops (foundry-agent-optimizer). USE FOR: evaluate agent post-deploy, score grounding quality, measure tool_selection, detect dataset drift, write custom grader, check URL citations, validate eval RBAC, day-1 smoke test, continuous eval loop, pre-merge eval gate, Foundry Evals SDK setup, evaluator framework, eval-driven optimization, ASSERT evaluators. DO NOT USE FOR: deploying agents (use threadlight-deploy), designing processes (use threadlight-design), unit testing code, reimplementing evaluator framework (use foundry-assert), writing your own optimizer loop (use foundry-agent-optimizer).

Ir para a instalação

Informações da origem

Repositório
aiappsgbb/awesome-gbb
Última atividade na origem
25 de setembro de 2026 às 14:00
Idioma detectado do SKILL.md
inglês
Estrelas
6
Forks
3

Opções de instalação

Por padrão, está selecionado o prompt que primeiro revisa a origem. Você pode mudar para um comando direto ou baixar uma cópia local.

Revise os arquivos de origem

Leia o SKILL.md e os arquivos complementares exibidos pelo SkillsMP antes de decidir se vai instalar.

Explorador de arquivos
8 arquivos

Exibindo SKILL.md

SKILL.md
Instruções da origem · Visualização somente leitura
name
foundry-evals
description
Evaluate Foundry agents with agent-target runs or an explicit invoke+score fallback. Covers sequential invocation, cold-start handling, dataset creation, evaluator configuration, RBAC for eval judges, result interpretation, and cross-refs to community evaluator frameworks (foundry-assert) and eval-driven prompt optimization loops (foundry-agent-optimizer). USE FOR: evaluate agent post-deploy, score grounding quality, measure tool_selection, detect dataset drift, write custom grader, check URL citations, validate eval RBAC, day-1 smoke test, continuous eval loop, pre-merge eval gate, Foundry Evals SDK setup, evaluator framework, eval-driven optimization, ASSERT evaluators. DO NOT USE FOR: deploying agents (use threadlight-deploy), designing processes (use threadlight-design), unit testing code, reimplementing evaluator framework (use foundry-assert), writing your own optimizer loop (use foundry-agent-optimizer).
metadata
{"version":"1.4.2"}
# Foundry Agent Evaluations Evaluate Foundry agents using the supported **agent-target** evaluation surface, or an explicit **invoke+score** fallback with captured responses. For the complete per-agent adoption and release-evidence workflow, see [`foundry-agentops`](../foundry-agentops/SKILL.md). This skill remains authoritative for deep evaluator configuration and dataset design; AgentOps aggregates evidence and never replaces that evaluation contract. ## When to Use - After deploying a hosted agent, to measure quality - Running batch evaluations against test scenarios - Comparing agent performance across versions - Validating that business rules (BR-XXX from SpecKit) are followed ## Why Two Phases? Two phases are a fallback, not a universal requirement. Current [cloud evaluation guidance](https://learn.microsoft.com/azure/foundry/observability/how-to/cloud-evaluation-targets) supports `azure_ai_target_completions` with an `azure_ai_agent` target for prompt and hosted agents. Select the path from the actual endpoint/protocol: | Surface | Path | |---|---| | Supported prompt or hosted Responses agent | Agent-target run: query-only input, explicit agent name/version, `{{sample.output_text}}` for string evaluators or `{{sample.output_items}}` for interaction-aware evaluators. | | Hosted Invocations agent | Use the documented freeform `input_messages` matching its real request contract; do not send a Responses template. The smoke helper below intentionally covers Responses only. | | Existing captured responses, custom transport, or a confirmed unsupported target surface | Invoke explicitly, retain the real response/tool transcript, then score `{{item.response}}` as described in Phases 1 and 2 below. | Start with one approved synthetic query. A created run is not success: poll to a successful terminal state, download per-item `output_items`, and require actual scores and usable responses. An auth, schema, model, or quota failure is **not** permission to silently switch paths. Record the failure; use invoke+score only when the surface requires it or the owner approves that fallback. > **Canonical executable:** [references/python/eval_runner.py](references/python/eval_runner.py) > implements a one-item agent-target or explicit invoke+score coherence smoke > using `AIProjectClient.get_openai_client().evals`. See the > [Day-1 recipe](#day-1-smoke-test-recipe-hosted-agent--mcp-tool-pilots). > It does not replace full suites, custom graders or the EVAL-201 adapter. ``` Phase 1: Invoke agent → collect responses Phase 2: Score responses → Foundry evaluators ``` **When using the fallback:** complete both phases sequentially; never score an invented response or treat a run ID as the agent's response. --- ## Phase 1: Invoke the Agent ```python from azure.ai.projects import AIProjectClient from azure.identity import DefaultAzureCredential project = AIProjectClient( endpoint="<project_endpoint>", credential=DefaultAzureCredential(), allow_preview=True, ) oai = project.get_openai_client(agent_name="my-agent") # Warm up — single-shot ping is INSUFFICIENT for hosted agents that # scale to zero (15min idle). Use the retry loop pattern below for any # eval likely to hit a cold container. A single ping can return # server_error in ~9s and make every scenario fail before the agent # has even spun up. print("Warming up...") oai.responses.create(input="Hello", stream=False) # MUST invoke each query SEQUENTIALLY (never concurrent) # Concurrent requests overwhelm cold-start containers → empty responses results = [] for query in test_queries: response = oai.responses.create(input=query, stream=False) results.append({ "query": query, "response": response.output_text, }) print(f"✓ {query[:50]}...") ``` ### Critical Rules - **DO NOT invoke concurrently** — concurrent eval requests overwhelm cold-start containers and produce empty responses. Sequential invocation is mandatory. - **MUST warm up first** — send a throwaway "Hello" before the real queries using the retry loop pattern below (not a single ping). Cold containers return `server_error` in 5-10s before the platform has brought a replica up. - **DO use `stream=False`** — simpler for eval and required for batch scoring. - **MUST bind to agent endpoint** — use `get_openai_client(agent_name=...)` to route to the dedicated endpoint, not the project endpoint. - **Token refresh for long runs** — `DefaultAzureCredential` tokens expire after ~1h. For 30+ scenarios, refresh the token every 10 items or create a fresh client per batch. - **MUST pace invocations with 5s delay** — add `time.sleep(5)` between calls. Rapid-fire requests produce empty responses even after warm-up. - **DO NOT poll container status API** — do NOT call `.../versions/{v}/containers/default` to check readiness on refreshed Foundry. This endpoint returns **HTTP 404** on modern hosted agents (the platform auto-provisions containers). Polling wastes 3+ minutes. Use the warmup chat loop below instead. ```python # Warmup with retry loop — handles scale-from-zero hosted agents WARMUP_ATTEMPTS = 4 WARMUP_BACKOFF_S = 60 print(f"[warmup] coldstart loop (up to {WARMUP_ATTEMPTS} retries with {WARMUP_BACKOFF_S}s backoff)...") for attempt in range(1, WARMUP_ATTEMPTS + 1): try: r = oai.responses.create(input="ping", stream=False) status = getattr(r, "status", "unknown") print(f"[warmup] attempt={attempt} status={status}") if status == "completed": print("[warmup] READY -- proceeding with eval") break except Exception as e: print(f"[warmup] attempt={attempt} EXC {type(e).__name__}: {str(e)[:120]}") if attempt < WARMUP_ATTEMPTS: time.sleep(WARMUP_BACKOFF_S) else: raise RuntimeError("Hosted agent failed to warm up after 4 attempts (4 minutes). Check the deployment.") ``` - **DO use ASCII-only logging on Windows** — see `Eval scripts on Windows: cp1252 trap` below. The default Windows console encoding is **cp1252**, not UTF-8. Any `print('→')`, `print('×')`, or `print('·')` in `run_evals.py` fails with `UnicodeEncodeError` mid-run, killing partial results. Use `->`, `x`, `::` instead (or set `PYTHONUTF8=1` in the venv bootstrap). - **MUST tolerate gateway flake on Phase 1** — Foundry's gateway can enter 5-10 minute sticky `internal_server_error` windows under burst load (especially mid-cold-start). 30-60s exponential backoff is **not** enough. See `Gateway flakiness during Phase 1` below for the resume-after-cooldown pattern. - **DO retry on empty response** — `output_text` can return empty when the agent does tool calls but the response structure varies. Retry once after a 3s pause. Also scan `response.output` items for message text as a fallback: ```python text = response.output_text or "" if not text and response.output: for item in response.output: if getattr(item, "type", "") == "message": for content in getattr(item, "content", []): text += getattr(content, "text", "") ``` ### Alternative: Direct SSE Invocations (GHCP SDK agents) If the agent uses the Invocations protocol (GHCP SDK), you can't use `oai.responses.create()`. Use the raw SSE endpoint instead: ```python import aiohttp url = f"{endpoint}/agents/{agent_name}/endpoint/protocols/invocations?api-version=v1" async with aiohttp.ClientSession() as session: async with session.post(url, json={"input": query}, headers={ "Authorization": f"Bearer {token}", "Content-Type": "application/json", "Foundry-Features": "HostedAgents=V1Preview", }) as resp: full_text = "" async for line in resp.content: line = line.decode().strip() if line.startswith("data: "): event = json.loads(line[6:]) if event.get("type") == "assistant.message_delta": full_text += event.get("data", {}).get("content", "") ``` > **Prefer the Responses API pattern** (agent-bound OpenAI client) unless the agent > specifically only supports Invocations. It's simpler and handles conversation state. ### Eval scripts on Windows: cp1252 trap The default Windows Python console uses **cp1252**, not UTF-8. Any non-Latin1 character in a `print(...)` call mid-eval produces: ``` UnicodeEncodeError: 'charmap' codec can't encode character '\u2192' in position 14: character maps to <undefined> ``` …and kills the eval run. We've burned an hour on this twice. Two defenses, in priority order: **1. ASCII-only logging in `eval/run_evals.py`** (the right fix) ```python # BAD — fails on Windows cp1252 print(f" -> {ms}ms · {len(text)} chars · attempt {attempt}") # GOOD — works on every platform print(f" -> {ms}ms :: {len(text)} chars :: attempt {attempt}") ``` Banned characters: `→ × · ✓ ✗ ❌ ✅ ▶ ▲ ▼ ← ° €` and any em-dash / en-dash. Replacements: `-> x :: PASS FAIL ! [done] etc.` **2. The agent's response can ALSO contain non-ASCII** (the gap that bit us twice in pilot runs even with #1 in place) Banning Unicode in your own `print()` calls only solves half the problem. The hosted agent is free to emit `→`, em-dashes, smart quotes, and curly arrows in its tool-result tables — and the moment your eval prints `output_first_400={out_text!r}`, the **agent's** `→` blows up the same way. Two layers, both required: ```python # AT THE TOP OF run_evals.py — wrap stdout/stderr to silently # replace any byte that doesn't fit the cp1252 console import io, sys sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding="utf-8", errors="replace", line_buffering=True) sys.stderr = io.TextIOWrapper(sys.stderr.buffer, encoding="utf-8", errors="replace", line_buffering=True) def safe(s, n=None): """Strip everything not ASCII before printing agent text.""" if s is None: return "" if n: s = s[:n] return s.encode("ascii", errors="replace").decode("ascii") # Then use safe() on every agent string you print: print(f" output_first_400: {safe(out_text, 400)!r}") ``` `io.TextIOWrapper(.., errors='replace')` is **the** fix for this — it turns any otherwise-fatal `UnicodeEncodeError` into a `?` substitution silently. Without it, even one `→` in one tool-result line aborts the whole eval and you lose all results-so-far (see § Incremental result writes below for the corollary mitigation). **3. PYTHONUTF8=1 in the venv bootstrap** (defense in depth) If you can't audit every `print` (e.g., a vendored library prints arrows), force the interpreter into UTF-8 mode at process start. This MUST be in the parent shell **before** Python launches: ```powershell # scripts/setup_eval_env.ps1 $env:PYTHONUTF8 = "1" $env:PYTHONIOENCODING = "utf-8" uv venv .venv .\.venv\Scripts\Activate.ps1 uv pip sync requirements.txt ``` ```bash # scripts/setup_eval_env.sh export PYTHONUTF8=1 export PYTHONIOENCODING=utf-8 uv venv .venv source .venv/bin/activate uv pip sync requirements.txt ``` Setting these inside the script (`os.environ["PYTHONUTF8"] = "1"`) **does not work** — by the time the assignment runs, Python's stdio encoding is already locked. (Defense #2 above DOES work mid-script because it rebuilds the wrapper from raw `sys.stdout.buffer`.) The same trap kills `az acr build` log streaming on Windows; fix there is `--no-logs` (separate skill: `foundry-mcp-aca`). ### Incremental result writes (mandatory default, not just for big batches) The "Resume-after-cooldown" pattern in the next section is framed as "recommended for batches > 5 scenarios." Treat that as a floor, not a ceiling — **every** `run_evals.py` should write its results JSON **after each scenario completes**, not at the end of the loop. Three real things that have wiped a full eval batch this pilot cycle: 1. `UnicodeEncodeError` on the agent's response (the trap above) — killed agent versions mid-scenario every cold-start run, lost 4 prior results. 2. Foundry gateway sticky 5xx window — kills attempt N, you lose 1..N-1. 3. Ctrl+C / shell window closed by accident — same outcome. ```python from pathlib import Path RESULTS_PATH = Path("eval/results.json") RESULTS_PATH.parent.mkdir(parents=True, exist_ok=True) results = [] for sc in selected: rec = run_one(sc) results.append(rec) # Write after EVERY scenario, not just at the end of the loop RESULTS_PATH.write_text(json.dumps(results, indent=2), encoding="utf-8") ``` The cost is one tiny disk write per scenario (~5ms). The benefit is that a kill at scenario 4 of 5 still leaves you with 4 scored results on disk, ready to rerun-the-tail on top. ### Gateway flakiness during Phase 1 Foundry's gateway occasionally enters 5-10+ minute sticky windows of `internal_server_error` (or `503 model overloaded`) — typically right after a cold-start burst, or when an upstream model deployment is being reconfigured. The 30-60s exponential backoff most retry libraries ship with is **not enough** on bad days. Two patterns: **Pattern A — Resume-after-cooldown (recommended for batches > 5 scenarios).** Persist results-so-far to disk after every successful case, and add a `--resume` flag that skips already-completed `case_id`s. Then a sticky-flake window just means "wait 10 min, re-run; it picks up where it left off". `eval/run_evals.py` should look like: ```python results_path = Path("eval/results.partial.json") done_ids = set() if results_path.exists() and "--resume" in sys.argv: done_ids = {r["case_id"] for r in json.loads(results_path.read_text())} for case in cases: if case["case_id"] in done_ids: print(f" [skip] {case['case_id']} already done") continue try: result = invoke_with_retry(case, max_retries=3, base_delay=60) except GatewayFlakeException: print(f" [defer] {case['case_id']} -- gateway sticky, re-run with --resume") break # don't burn budget append_result(results_path, result) ``` **Pattern B — Skip-and-mark.** If the SLA is "best 5 of 6" rather than "all 6", record gateway failures as `status: GATEWAY_FLAKE` (NOT `FAIL`), exclude them from scoring, and report the count in the run summary. Don't let infrastructure noise pollute the agent-quality signal. > Both patterns are about **separating gateway flake from agent > quality**. A failed retry on `internal_server_error` is not an > eval-quality failure; it's a Foundry-side incident. Score them > separately or you'll spend a day debugging a "regression" that > turns out to be a 7-minute gateway window. --- ## Phase 2: Score with Foundry Evaluators Foundry evals have two concepts: - **Eval definition** — the evaluator configuration (metrics, data schema, judge model). Create this **once**, reuse across runs. Only recreate when changing metrics. - **Eval run** — a single execution against a dataset. Create a new run for each test cycle, agent version, or dataset update.
Ver no GitHub
Este SKILL.md e muito grande, entao o SkillsMP mostra aqui apenas a primeira secao. Ver no GitHub