Skip to main content

threadlight-evals

Threadlight evals leg: offline batch quality, Foundry Continuous Evaluation and champion-challenger gates. Emits `specs/evals-manifest.json` for pillar 6. Reuses validated opted-in AgentOps batches without reruns or live-wiring claims. USE FOR: continuous evals, offline eval gate, eval schedule, Foundry Continuous Evaluation, create_agent_evaluation, Application Insights eval results, eval threshold alert, eval run freshness, eval dataset shape, tool_calls tool_outputs, champion challenger, A/B eval gate, model/prompt swap gate, judge calibration, LLM-as-judge, foundry-evals pipeline leg, EVAL-001..006, EVAL-101..105, evals-manifest. DO NOT USE FOR: token-level content filtering — use Azure AI Content Safety; adversarial scanning — use threadlight-redteam; agent-runtime action governance — use threadlight-govern; deep evaluator or dataset authoring — use foundry-evals.

설치로 이동

소스 정보

저장소
aiappsgbb/threadlight-skills
최근 소스 활동
2026년 9월 23일 21:34
감지된 SKILL.md 언어
영어
스타
1
포크
5

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.

파일 탐색기
20 개 파일

SKILL.md 표시 중

SKILL.md
소스 지침 · 읽기 전용 미리보기
name
threadlight-evals
description
Threadlight evals leg: offline batch quality, Foundry Continuous Evaluation and champion-challenger gates. Emits `specs/evals-manifest.json` for pillar 6. Reuses validated opted-in AgentOps batches without reruns or live-wiring claims. USE FOR: continuous evals, offline eval gate, eval schedule, Foundry Continuous Evaluation, create_agent_evaluation, Application Insights eval results, eval threshold alert, eval run freshness, eval dataset shape, tool_calls tool_outputs, champion challenger, A/B eval gate, model/prompt swap gate, judge calibration, LLM-as-judge, foundry-evals pipeline leg, EVAL-001..006, EVAL-101..105, evals-manifest. DO NOT USE FOR: token-level content filtering — use Azure AI Content Safety; adversarial scanning — use threadlight-redteam; agent-runtime action governance — use threadlight-govern; deep evaluator or dataset authoring — use foundry-evals.
metadata
{"version":"0.3.1"}
# Threadlight Evals — run continuous evals, then prove they ran The **DISCOVER + GOVERN + IMPROVE** evals leg of `path2production` for AI agents on Microsoft Foundry. `threadlight-production-ready` pillar 6 scores whether continuous evals are scheduled, fresh, and alerting; this skill is the executable leg that owns running and verifying that posture in the pipeline. ``` DESIGN → BUILD/DEPLOY → DISCOVER → PROTECT → [ GOVERN / IMPROVE ] → PRODUCTION-READY threadlight-evals ``` > **Why this skill exists.** A pilot can have strong unit tests, local smoke > tests, and a green deployment gate, yet still regress silently after a model, > prompt, retrieval, or tool-schema change. Previously the chain delegated eval > work entirely to `foundry-evals`, and pillar 6 could only score whether a > schedule appeared to exist. This skill makes evals a threadlight-owned leg: > it detects offline datasets, live continuous evaluation wiring, threshold > alerts, run freshness, and champion-challenger gates; then emits a manifest the > scorecard can verify. ## What this skill covers | Feature | Purpose | Output signal | |---|---|---| | **F1 Offline batch quality eval** | Thin-wrap the `foundry-evals` invoke+score path in the threadlight spine. Heavy evaluator authoring still belongs to `foundry-evals`; this leg owns running it and producing evidence. | `eval_scenarios_present`, `eval_datasets_present`, `run_history_present`, `latest_pass_rate_ok` | | **F2 Online / continuous eval** | Wire Foundry Continuous Evaluation on live threads using `create_agent_evaluation(thread, run, evaluators=[...], app_insights_connection_string=...)`. Results land in Application Insights. | `schedule_present`, `online_eval_wired`, `latest_eval_run_fresh`, `alert_wired` | | **F3 A/B champion–challenger** | Gate a model/prompt/tool swap by running the same eval dataset against champion and challenger before the swap. | `ab_comparison_present` | ## What this skill does NOT replace | Concern | Use instead | |---|---| | Token-level content filtering, prompt shields, content safety at the model edge | Model guardrail / Azure AI Content Safety | | Adversarial scanning and jailbreak probe campaigns | `threadlight-redteam` | | Agent-runtime action policy, tool allow/deny, excessive-agency controls | `threadlight-govern` | | Deep evaluator authoring, grading rubric design, dataset enrichment work | `foundry-evals` | | Overall production-readiness scorecard | `threadlight-production-ready` | ## The contract — `specs/evals-manifest.json` `scripts/evals_check.py` walks a pilot repo and emits a manifest whose capability keys map directly to pillar 6 (`continuous-evals`) IDs: | Capability key | Pillar ID | Meaning | Severity when missing | |---|---:|---|---| | `eval_scenarios_present` | `EVAL-001` | SPEC § 9 lists eval scenarios or `evals/` contains scenario files | must-fix | | `eval_datasets_present` | `EVAL-002` | Held-out datasets are stored under `evals/` or `specs/evals/` | must-fix | | `dataset_shape_ok` | `EVAL-003` | At least one JSON/JSONL row includes `tool_calls` and `tool_outputs` | should-fix | | `thresholds_declared` | `EVAL-004` | Per-scenario `threshold`, `min_score`, or pass-rate threshold is declared | should-fix | | `schedule_present` | `EVAL-005` | Plan A Foundry CE schedule or Plan B cron/ACA Job eval runner exists | must-fix | | `run_history_present` | `EVAL-006` | Latest run output is committed under `evals/runs/*.json` or `docs/eval-runs/` | should-fix | | `online_eval_wired` | `EVAL-101` | Code calls `create_agent_evaluation(...)` with Application Insights connection | should-fix; must-fix when Plan A is declared but unwired | | `latest_eval_run_fresh` | `EVAL-102/103` | Latest scheduled run is within the freshness window (default 7 days) | should-fix if stale; not-verified if no history | | `alert_wired` | `EVAL-104` | Alert or notification exists for eval threshold breach | must-fix | | `latest_pass_rate_ok` | `EVAL-105` | Latest run pass rate meets the declared minimum | should-fix if below threshold | | `ab_comparison_present` | `F3` | Champion-challenger comparison config/script exists before swaps | should-fix | Status taxonomy is exactly: `pass`, `must-fix`, `should-fix`, `not-verified`, `not-applicable`. Verdict roll-up: | Verdict | Meaning | |---|---| | `comprehensive` | No `must-fix`, `should-fix`, or `not-verified` findings. Offline, online, alerting, freshness, pass-rate, and A/B gate are all verified. | | `partial` | No `must-fix`, but one or more advisory or unverified capabilities remain. | | `offline-only` | Offline basics are present, but at least one hard requirement for continuous/online operation is missing. | | `none` | Evals are absent or too incomplete to trust. | The emitted manifest includes: ```json { "schema": "threadlight-evals-manifest/v1", "tool_version": "0.1.0", "captured_at": "2026-06-22T16:00:00+00:00", "freshness_window_days": 7, "verdict": "comprehensive", "must_fix": [], "should_fix": [], "not_verified": [], "metrics": { "pass_rate": 0.91, "threshold": 0.8, "latest_run": "evals/runs/2026-01-01.json" }, "capabilities": { "eval_scenarios_present": { "check_id": "EVAL-001", "status": "pass", "evidence": "specs/spec.md", "hint": null } } } ``` The top-level `metrics` block surfaces the latest run's `pass_rate`, the declared `threshold`, and the relative path to the `latest_run`. This is the join key `threadlight-production-ready` reads (`metrics.pass_rate`) to render the eval-quality column of its outcome-KPI scorecard — so eval quality flows into the business-KPI view instead of staying locked inside the evals leg. `pass_rate`/`latest_run` are `null` when no machine-readable run history exists. ## Usage ```bash # 1. Assess an existing pilot (read-only) — prints the evals report python3 scripts/evals_check.py --target ../my-pilot # 2. Emit the manifest + human report the scorecard consumes python3 scripts/evals_check.py --target ../my-pilot --emit # → writes specs/evals-manifest.json + docs/evals-report.md # 3. CI gate — exit 2 on any must-fix capability python3 scripts/evals_check.py --target ../my-pilot --gate # 4. JSON for automation python3 scripts/evals_check.py --target ../my-pilot --json # 5. Override freshness window (pillar default is 7 days) python3 scripts/evals_check.py --target ../my-pilot --freshness-days 14 ``` Flags: | Flag | Meaning | |---|---| | `--target PATH` | Pilot repository root. Defaults to `.`. | | `--emit` | Writes `specs/evals-manifest.json` and `docs/evals-report.md` under the target. | | `--gate` | Returns exit code `2` when any capability is `must-fix`. | | `--json` | Prints the manifest JSON instead of markdown. | | `--freshness-days N` | Max age for latest eval run. Default: `7`. | ## The leg, end-to-end ### Selective AgentOps consumption For explicitly opted-in agents, reuse `specs/agentops-manifest.json` only through the packaged `skills/_shared/agentops.py` validator. Do not rerun the same batch merely to populate Threadlight evidence. The producer owns native ingestion, binding, hashes and freshness; this consumer never reads raw `.agentops` payloads or imports customer-repository helpers. Fresh, bound summary counts, aggregate metrics and nonempty threshold results can fill run history, run freshness, declared thresholds and quality checks. Native `items_pass_rate` is an **execution pass rate**, not a measured per-row quality pass rate: retain it under per-agent execution metrics, never promote it to the outcome-KPI `metrics.pass_rate`. Threshold failures and partial row errors stay negative. Native comparison directions remain descriptive, not automatic quality policy or A/B approval. Keep every opted-in agent and the worst status across agents and valid conflicting native evidence. AgentOps cannot fill scenarios, dataset shape, scheduling, online evaluation, alerts or champion–challenger/A-B gates. A comparison summary alone is not an A/B deployment gate. Missing, malformed, stale, future or mismatched summaries remain unverified; a Doctor `ready` headline is not eval quality proof. Emit allowlisted aggregates and provenance only, never rows, prompts, effective configuration or native free-form report text. The optional `agentops.agents[]` block retains each agent's `domain_status`, execution metrics, artifact hashes and receipt hash. `represented_blockers` contains `AOPS-EVAL-QUALITY` only when the verified domain's matching blocker is actually represented by a canonical must-fix quality finding. Unknown blockers remain separate AgentOps findings; a verdict alone cannot deduplicate. This is a **producing** leg. It emits artefacts. It does not silently mutate the user's repo for findings; remediation is performed by the agent after reviewing manifest gaps. 1. **Discover offline eval assets.** Confirm SPEC § 9 scenarios, held-out datasets, dataset shape (`tool_calls` + `tool_outputs`), thresholds, and run history. 2. **Govern continuous wiring.** Confirm Plan A Foundry Continuous Evaluation or Plan B scheduled fallback, plus Application Insights result flow and an alert for threshold breach. 3. **Improve safely.** Confirm a champion-challenger comparison gate exists so prompt/model/tool swaps are evaluated before promotion. 4. **Emit evidence.** Write `specs/evals-manifest.json` and `docs/evals-report.md`; the scorecard consumes the manifest. ## Plan A — Foundry Continuous Evaluation Preferred production wiring uses Foundry Continuous Evaluation on live threads: ```python client.evaluations.create_agent_evaluation( thread=thread, run=run, evaluators=[groundedness, relevance, tool_output_utilization], app_insights_connection_string=os.environ["APPLICATIONINSIGHTS_CONNECTION_STRING"], ) ``` Use `DefaultAzureCredential` and managed identity/OIDC; do not store secrets in repo files. See `references/foundry-ce-wiring.md` for the full keyless snippet. Plan A detection accepts config or code under `infra/` / `evals/` that references continuous evaluation and a schedule, and code that calls `create_agent_evaluation(...)` with an Application Insights connection string. ## Plan B — scheduled fallback Plan B is acceptable when Foundry Continuous Evaluation is not yet available for the pilot shape. The validator detects: - `.github/workflows/*.yml` with `cron` and eval runner markers. - ACA Job / Container Apps job definitions under `infra/` with schedule markers. - Run outputs committed under `evals/runs/` or `docs/eval-runs/`. Plan B still needs threshold alerts or workflow failure notifications. No schedule is a hard fail. ## F3 — champion-challenger gate For an optional lower-cost runtime experiment after prototype validation, use the [fixed-artifact model checklist](../threadlight-design/references/model-selection.md). The authoring model is independent; do not regenerate the prototype or change its prompts to compare runtimes. Predeclare quality, tool/skill-use, structured output, PII/safety/governance and latency constraints; include failed-case spend in cost per correctly completed task. No paid comparison or deployment is authorized merely by this recommendation. Before changing a model, prompt, tool contract, or retrieval configuration: 1. Run champion and challenger against the same held-out dataset. 2. Compare pass-rate, protected-scenario scores, regression delta, and every declared control and latency constraint. 3. Promote only if the challenger meets all declared thresholds and controls; otherwise keep the champion. Never relax controls to make it pass. Place config under `evals/ab/` or include `champion`, `challenger`, and `baseline_vs` markers in an eval script/config. See `references/ab-comparison.md`. ## Judge calibration — before `pass_rate` becomes evidence `metrics.pass_rate` does not stop here. Pillar 6 promotes it into the scorecard's outcome-KPI column, and the EU AI Act evidence pack cites it under Art 15 (accuracy). So the number has to survive the question *"how was it produced?"* When a model grades the run — which is the common case — four review criteria apply. None of them is machine-checked; they are reviewer's judgement: | Criterion | Why | |---|---| | Rows carry `expected`, **and the judge prompt receives it** | A reference-free judge grades plausibility, not correctness, and over-credits confident wrong answers. | | Rubric is binary or few-level with written anchors | A bare 1–5 scale clusters on 3–4, so the threshold becomes arbitrary. | | Judge–human agreement measured once on a seed set | The only honest answer to "why believe this score?". | | Judge model + version pinned | Swapping the judge changes the meaning of the whole run history. Treat it as an F3 swap. | Also: the skill that produces an output must not be the skill that grades it. Upstream evaluators that emit 1–5 floats are binarized at a declared cut before aggregation — never gate on the mean. Full recipe, including the mapping table: `references/judge-calibration.md`. ## How `production-ready` consumes this `threadlight-production-ready` pillar 6 reads `specs/evals-manifest.json`. | Manifest state | Pillar 6 behavior | |---|---| | Present, fresh, and `verdict == comprehensive` | Flip EVAL findings to verified with manifest/report evidence. | | Present and `verdict == partial` | Verify passed capabilities; keep advisory or unverified findings open. | | Present and `verdict == offline-only` | Treat offline evals as present, but re-open continuous-evals schedule/alert/live gaps. | | Present and `verdict == none` | Re-open EVAL must-fix findings. | | Missing or stale manifest | Fall back to legacy scoring and delegate remediation. | Freshness is determined by the consuming scorecard and by this validator's `latest_eval_run_fresh` capability. Default window: 7 days. ## Relationship to `foundry-evals` `foundry-evals` is the deep upstream skill for evaluator selection, rubric authoring, dataset expansion, and Foundry eval details. `threadlight-evals` is a thin threadlight pipeline leg: it makes evals a step that runs in the spine, checks whether the production posture is wired, and emits the manifest the scorecard understands. Use both together: - Use `foundry-evals` to author or repair evaluator definitions and datasets. - Use `threadlight-evals` to verify the pilot has offline, online, alerting, freshness, and A/B gate evidence. ## Files ``` scripts/evals_check.py # stdlib validator → evals-manifest.json references/foundry-ce-wiring.md # Plan A keyless CE wiring snippet references/ab-comparison.md # champion-challenger recipe + gates references/judge-calibration.md # LLM-as-judge review criteria + 1-5 mapping references/dataset-shape.md # held-out dataset row shape references/evals-manifest.schema.json # manifest contract references/fixtures/sample-scheduled/ # passing scheduled/continuous pilot references/fixtures/sample-manual/ # manual-only pilot with must-fix gaps tests/test_evals_check.py # stdlib unittest (no pytest) ``` ## Tests ```bash cd skills/threadlight-evals python3 -m unittest discover -s tests -v ``` Expected coverage: - Scheduled fixture has no `must-fix` findings and gates with exit code `0`. - Manual fixture reports schedule/alert `must-fix` and gates with exit code `2`. - Manifest includes schema, required keys, and all capability keys. ## Common gaps and remediations | Gap | Why it matters | Remediation | |---|---|---| | Evals only run manually | Regressions after swaps are not caught. | Add Plan A CE schedule or Plan B cron/ACA Job. | | Dataset lacks `tool_calls` / `tool_outputs` | Tool-output graders cannot score grounding correctly. | Enrich held-out dataset rows; see `references/dataset-shape.md`. | | Thresholds absent | Pass/fail becomes subjective. | Add `threshold`, `min_score`, or `min_pass_rate` per scenario. | | No alert | Failures sit unnoticed in telemetry. | Add Azure Monitor alert or workflow notification for threshold breach. | | No run history | Reviewers cannot prove evals ran. | Commit latest run JSON under `evals/runs/`. | | No A/B gate | Model/prompt swaps can regress protected scenarios. | Add `evals/ab/` champion-challenger config. | | Judge grades without a reference | Scores plausibility rather than correctness; confident wrong answers pass. | Add `expected` to rows *and* interpolate it in the grader prompt; see `references/judge-calibration.md`. | | Gate reads the mean of a 1–5 scale | A 4.2 mean hides a failing tail. | Binarize per row at a declared cut, then aggregate. |
GitHub에서 보기
이 SKILL.md는 매우 커서 SkillsMP가 여기에는 첫 섹션만 미리 보여줍니다. GitHub에서 보기