skill
职业分类
描述
更新
consistency-data-issue-review
软件质量保证分析师与测试员
Run repeated benchmark consistency studies, find turn-level pass/fail flips across runs of the same model, identify suspected data-spec or grader issues versus real model failures, and generate a polished HTML review. Use when asked to run N repeated evals, inspect unstable datapoints, separate likely benchmark/judge problems from model problems, or produce a consistency review report.
2026-03-23
model-run-error-analysis
软件质量保证分析师与测试员
Analyze benchmark runs to identify dominant error modes per model, shared hard turns, grader or benchmark issues, and representative failing examples. Prefer judged runs, but fall back to transcript + benchmark-contract review when judged artifacts are missing or unreliable.
2026-03-19