Skip to main content

foundry-evals

Evaluate Foundry agents with agent-target runs or an explicit invoke+score fallback. Covers sequential invocation, cold-start handling, dataset creation, evaluator configuration, RBAC for eval judges, result interpretation, and cross-refs to community evaluator frameworks (foundry-assert) and eval-driven prompt optimization loops (foundry-agent-optimizer). USE FOR: evaluate agent post-deploy, score grounding quality, measure tool_selection, detect dataset drift, write custom grader, check URL citations, validate eval RBAC, day-1 smoke test, continuous eval loop, pre-merge eval gate, Foundry Evals SDK setup, evaluator framework, eval-driven optimization, ASSERT evaluators. DO NOT USE FOR: deploying agents (use threadlight-deploy), designing processes (use threadlight-design), unit testing code, reimplementing evaluator framework (use foundry-assert), writing your own optimizer loop (use foundry-agent-optimizer).

الانتقال إلى التثبيت

معلومات المصدر

المستودع
aiappsgbb/awesome-gbb
آخر نشاط في المصدر
٢٥ سبتمبر ٢٠٢٦ في ١٤:٠٠
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
٦
التفرعات
٣

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.

مستكشف الملفات
8 ملفات

عرض SKILL.md

SKILL.md
تعليمات المصدر · معاينة للقراءة فقط
name
foundry-evals
description
Evaluate Foundry agents with agent-target runs or an explicit invoke+score fallback. Covers sequential invocation, cold-start handling, dataset creation, evaluator configuration, RBAC for eval judges, result interpretation, and cross-refs to community evaluator frameworks (foundry-assert) and eval-driven prompt optimization loops (foundry-agent-optimizer). USE FOR: evaluate agent post-deploy, score grounding quality, measure tool_selection, detect dataset drift, write custom grader, check URL citations, validate eval RBAC, day-1 smoke test, continuous eval loop, pre-merge eval gate, Foundry Evals SDK setup, evaluator framework, eval-driven optimization, ASSERT evaluators. DO NOT USE FOR: deploying agents (use threadlight-deploy), designing processes (use threadlight-design), unit testing code, reimplementing evaluator framework (use foundry-assert), writing your own optimizer loop (use foundry-agent-optimizer).
metadata
{"version":"1.4.2"}
# Foundry Agent Evaluations Evaluate Foundry agents using the supported **agent-target** evaluation surface, or an explicit **invoke+score** fallback with captured responses. For the complete per-agent adoption and release-evidence workflow, see [`foundry-agentops`](../foundry-agentops/SKILL.md). This skill remains authoritative for deep evaluator configuration and dataset design; AgentOps aggregates evidence and never replaces that evaluation contract. ## When to Use - After deploying a hosted agent, to measure quality - Running batch evaluations against test scenarios - Comparing agent performance across versions - Validating that business rules (BR-XXX from SpecKit) are followed ## Why Two Phases? Two phases are a fallback, not a universal requirement. Current [cloud evaluation guidance](https://learn.microsoft.com/azure/foundry/observability/how-to/cloud-evaluation-targets) supports `azure_ai_target_completions` with an `azure_ai_agent` target for prompt and hosted agents. Select the path from the actual endpoint/protocol: | Surface | Path | |---|---| | Supported prompt or hosted Responses agent | Agent-target run: query-only input, explicit agent name/version, `{{sample.output_text}}` for string evaluators or `{{sample.output_items}}` for interaction-aware evaluators. | | Hosted Invocations agent | Use the documented freeform `input_messages` matching its real request contract; do not send a Responses template. The smoke helper below intentionally covers Responses only. | | Existing captured responses, custom transport, or a confirmed unsupported target surface | Invoke explicitly, retain the real response/tool transcript, then score `{{item.response}}` as described in Phases 1 and 2 below. | Start with one approved synthetic query. A created run is not success: poll to a successful terminal state, download per-item `output_items`, and require actual scores and usable responses. An auth, schema, model, or quota failure is **not** permission to silently switch paths. Record the failure; use invoke+score only when the surface requires it or the owner approves that fallback. > **Canonical executable:** [references/python/eval_runner.py](references/python/eval_runner.py) > implements a one-item agent-target or explicit invoke+score coherence smoke > using `AIProjectClient.get_openai_client().evals`. See the > [Day-1 recipe](#day-1-smoke-test-recipe-hosted-agent--mcp-tool-pilots). > It does not replace full suites, custom graders or the EVAL-201 adapter. ``` Phase 1: Invoke agent → collect responses Phase 2: Score responses → Foundry evaluators ``` **When using the fallback:** complete both phases sequentially; never score an invented response or treat a run ID as the agent's response. --- ## Phase 1: Invoke the Agent ```python from azure.ai.projects import AIProjectClient from azure.identity import DefaultAzureCredential project = AIProjectClient( endpoint="<project_endpoint>", credential=DefaultAzureCredential(), allow_preview=True, ) oai = project.get_openai_client(agent_name="my-agent") # Warm up — single-shot ping is INSUFFICIENT for hosted agents that # scale to zero (15min idle). Use the retry loop pattern below for any # eval likely to hit a cold container. A single ping can return # server_error in ~9s and make every scenario fail before the agent # has even spun up. print("Warming up...") oai.responses.create(input="Hello", stream=False) # MUST invoke each query SEQUENTIALLY (never concurrent) # Concurrent requests overwhelm cold-start containers → empty responses results = [] for query in test_queries: response = oai.responses.create(input=query, stream=False) results.append({ "query": query, "response": response.output_text, }) print(f"✓ {query[:50]}...") ``` ### Critical Rules - **DO NOT invoke concurrently** — concurrent eval requests overwhelm cold-start containers and produce empty responses. Sequential invocation is mandatory. - **MUST warm up first** — send a throwaway "Hello" before the real queries using the retry loop pattern below (not a single ping). Cold containers return `server_error` in 5-10s before the platform has brought a replica up. - **DO use `stream=False`** — simpler for eval and required for batch scoring. - **MUST bind to agent endpoint** — use `get_openai_client(agent_name=...)` to route to the dedicated endpoint, not the project endpoint. - **Token refresh for long runs** — `DefaultAzureCredential` tokens expire after ~1h. For 30+ scenarios, refresh the token every 10 items or create a fresh client per batch. - **MUST pace invocations with 5s delay** — add `time.sleep(5)` between calls. Rapid-fire requests produce empty responses even after warm-up. - **DO NOT poll container status API** — do NOT call `.../versions/{v}/containers/default` to check readiness on refreshed Foundry. This endpoint returns **HTTP 404** on modern hosted agents (the platform auto-provisions containers). Polling wastes 3+ minutes. Use the warmup chat loop below instead. ```python # Warmup with retry loop — handles scale-from-zero hosted agents WARMUP_ATTEMPTS = 4 WARMUP_BACKOFF_S = 60 print(f"[warmup] coldstart loop (up to {WARMUP_ATTEMPTS} retries with {WARMUP_BACKOFF_S}s backoff)...") for attempt in range(1, WARMUP_ATTEMPTS + 1): try: r = oai.responses.create(input="ping", stream=False) status = getattr(r, "status", "unknown") print(f"[warmup] attempt={attempt} status={status}") if status == "completed": print("[warmup] READY -- proceeding with eval") break except Exception as e: print(f"[warmup] attempt={attempt} EXC {type(e).__name__}: {str(e)[:120]}") if attempt < WARMUP_ATTEMPTS: time.sleep(WARMUP_BACKOFF_S) else: raise RuntimeError("Hosted agent failed to warm up after 4 attempts (4 minutes). Check the deployment.") ``` - **DO use ASCII-only logging on Windows** — see `Eval scripts on Windows: cp1252 trap` below. The default Windows console encoding is **cp1252**, not UTF-8. Any `print('→')`, `print('×')`, or `print('·')` in `run_evals.py` fails with `UnicodeEncodeError` mid-run, killing partial results. Use `->`, `x`, `::` instead (or set `PYTHONUTF8=1` in the venv bootstrap). - **MUST tolerate gateway flake on Phase 1** — Foundry's gateway can enter 5-10 minute sticky `internal_server_error` windows under burst load (especially mid-cold-start). 30-60s exponential backoff is **not** enough. See `Gateway flakiness during Phase 1` below for the resume-after-cooldown pattern. - **DO retry on empty response** — `output_text` can return empty when the agent does tool calls but the response structure varies. Retry once after a 3s pause. Also scan `response.output` items for message text as a fallback: ```python text = response.output_text or "" if not text and response.output: for item in response.output: if getattr(item, "type", "") == "message": for content in getattr(item, "content", []): text += getattr(content, "text", "") ``` ### Alternative: Direct SSE Invocations (GHCP SDK agents) If the agent uses the Invocations protocol (GHCP SDK), you can't use `oai.responses.create()`. Use the raw SSE endpoint instead: ```python import aiohttp url = f"{endpoint}/agents/{agent_name}/endpoint/protocols/invocations?api-version=v1" async with aiohttp.ClientSession() as session: async with session.post(url, json={"input": query}, headers={ "Authorization": f"Bearer {token}", "Content-Type": "application/json", "Foundry-Features": "HostedAgents=V1Preview", }) as resp: full_text = "" async for line in resp.content: line = line.decode().strip() if line.startswith("data: "): event = json.loads(line[6:]) if event.get("type") == "assistant.message_delta": full_text += event.get("data", {}).get("content", "") ``` > **Prefer the Responses API pattern** (agent-bound OpenAI client) unless the agent > specifically only supports Invocations. It's simpler and handles conversation state. ### Eval scripts on Windows: cp1252 trap The default Windows Python console uses **cp1252**, not UTF-8. Any non-Latin1 character in a `print(...)` call mid-eval produces: ``` UnicodeEncodeError: 'charmap' codec can't encode character '\u2192' in position 14: character maps to <undefined> ``` …and kills the eval run. We've burned an hour on this twice. Two defenses, in priority order: **1. ASCII-only logging in `eval/run_evals.py`** (the right fix) ```python # BAD — fails on Windows cp1252 print(f" -> {ms}ms · {len(text)} chars · attempt {attempt}") # GOOD — works on every platform print(f" -> {ms}ms :: {len(text)} chars :: attempt {attempt}") ``` Banned characters: `→ × · ✓ ✗ ❌ ✅ ▶ ▲ ▼ ← ° €` and any em-dash / en-dash. Replacements: `-> x :: PASS FAIL ! [done] etc.` **2. The agent's response can ALSO contain non-ASCII** (the gap that bit us twice in pilot runs even with #1 in place) Banning Unicode in your own `print()` calls only solves half the problem. The hosted agent is free to emit `→`, em-dashes, smart quotes, and curly arrows in its tool-result tables — and the moment your eval prints `output_first_400={out_text!r}`, the **agent's** `→` blows up the same way. Two layers, both required: ```python # AT THE TOP OF run_evals.py — wrap stdout/stderr to silently # replace any byte that doesn't fit the cp1252 console import io, sys sys.stdout = io.TextIOWrapper(sys.stdout.buffer, encoding="utf-8", errors="replace", line_buffering=True) sys.stderr = io.TextIOWrapper(sys.stderr.buffer, encoding="utf-8", errors="replace", line_buffering=True) def safe(s, n=None): """Strip everything not ASCII before printing agent text.""" if s is None: return "" if n: s = s[:n] return s.encode("ascii", errors="replace").decode("ascii") # Then use safe() on every agent string you print: print(f" output_first_400: {safe(out_text, 400)!r}") ``` `io.TextIOWrapper(.., errors='replace')` is **the** fix for this — it turns any otherwise-fatal `UnicodeEncodeError` into a `?` substitution silently. Without it, even one `→` in one tool-result line aborts the whole eval and you lose all results-so-far (see § Incremental result writes below for the corollary mitigation). **3. PYTHONUTF8=1 in the venv bootstrap** (defense in depth) If you can't audit every `print` (e.g., a vendored library prints arrows), force the interpreter into UTF-8 mode at process start. This MUST be in the parent shell **before** Python launches: ```powershell # scripts/setup_eval_env.ps1 $env:PYTHONUTF8 = "1" $env:PYTHONIOENCODING = "utf-8" uv venv .venv .\.venv\Scripts\Activate.ps1 uv pip sync requirements.txt ``` ```bash # scripts/setup_eval_env.sh export PYTHONUTF8=1 export PYTHONIOENCODING=utf-8 uv venv .venv source .venv/bin/activate uv pip sync requirements.txt ``` Setting these inside the script (`os.environ["PYTHONUTF8"] = "1"`) **does not work** — by the time the assignment runs, Python's stdio encoding is already locked. (Defense #2 above DOES work mid-script because it rebuilds the wrapper from raw `sys.stdout.buffer`.) The same trap kills `az acr build` log streaming on Windows; fix there is `--no-logs` (separate skill: `foundry-mcp-aca`). ### Incremental result writes (mandatory default, not just for big batches) The "Resume-after-cooldown" pattern in the next section is framed as "recommended for batches > 5 scenarios." Treat that as a floor, not a ceiling — **every** `run_evals.py` should write its results JSON **after each scenario completes**, not at the end of the loop. Three real things that have wiped a full eval batch this pilot cycle: 1. `UnicodeEncodeError` on the agent's response (the trap above) — killed agent versions mid-scenario every cold-start run, lost 4 prior results. 2. Foundry gateway sticky 5xx window — kills attempt N, you lose 1..N-1. 3. Ctrl+C / shell window closed by accident — same outcome. ```python from pathlib import Path RESULTS_PATH = Path("eval/results.json") RESULTS_PATH.parent.mkdir(parents=True, exist_ok=True) results = [] for sc in selected: rec = run_one(sc) results.append(rec) # Write after EVERY scenario, not just at the end of the loop RESULTS_PATH.write_text(json.dumps(results, indent=2), encoding="utf-8") ``` The cost is one tiny disk write per scenario (~5ms). The benefit is that a kill at scenario 4 of 5 still leaves you with 4 scored results on disk, ready to rerun-the-tail on top. ### Gateway flakiness during Phase 1 Foundry's gateway occasionally enters 5-10+ minute sticky windows of `internal_server_error` (or `503 model overloaded`) — typically right after a cold-start burst, or when an upstream model deployment is being reconfigured. The 30-60s exponential backoff most retry libraries ship with is **not enough** on bad days. Two patterns: **Pattern A — Resume-after-cooldown (recommended for batches > 5 scenarios).** Persist results-so-far to disk after every successful case, and add a `--resume` flag that skips already-completed `case_id`s. Then a sticky-flake window just means "wait 10 min, re-run; it picks up where it left off". `eval/run_evals.py` should look like: ```python results_path = Path("eval/results.partial.json") done_ids = set() if results_path.exists() and "--resume" in sys.argv: done_ids = {r["case_id"] for r in json.loads(results_path.read_text())} for case in cases: if case["case_id"] in done_ids: print(f" [skip] {case['case_id']} already done") continue try: result = invoke_with_retry(case, max_retries=3, base_delay=60) except GatewayFlakeException: print(f" [defer] {case['case_id']} -- gateway sticky, re-run with --resume") break # don't burn budget append_result(results_path, result) ``` **Pattern B — Skip-and-mark.** If the SLA is "best 5 of 6" rather than "all 6", record gateway failures as `status: GATEWAY_FLAKE` (NOT `FAIL`), exclude them from scoring, and report the count in the run summary. Don't let infrastructure noise pollute the agent-quality signal. > Both patterns are about **separating gateway flake from agent > quality**. A failed retry on `internal_server_error` is not an > eval-quality failure; it's a Foundry-side incident. Score them > separately or you'll spend a day debugging a "regression" that > turns out to be a 7-minute gateway window. --- ## Phase 2: Score with Foundry Evaluators Foundry evals have two concepts: - **Eval definition** — the evaluator configuration (metrics, data schema, judge model). Create this **once**, reuse across runs. Only recreate when changing metrics. - **Eval run** — a single execution against a dataset. Create a new run for each test cycle, agent version, or dataset update.
عرض على GitHub
ملف SKILL.md هذا كبير جدا، لذلك يعرض SkillsMP القسم الاول فقط هنا. عرض على GitHub