benchmark-run-operator
Operate benchmark runs with controlled launches, attributable artifacts, valid-row rules, and evidence-based recovery.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Operate benchmark runs with controlled launches, attributable artifacts, valid-row rules, and evidence-based recovery.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Design or redesign landing pages, portfolios, marketing sites, and editorial pages with deliberate visual direction; not dashboards, app flows, or mobile UI.
Compare current LLM quality, price, speed, context, modality, and OpenRouter availability for recommendations and implementation choices.
Triage failing WebArena or WebMCP tasks. Use to classify tool, eval, agent, infrastructure, and drift failures.
Design and implement WebMCP tool surfaces. Use for workflow discovery, tool shaping, and live validation.
Verify WebMCP tool changes in the live browser. Use for registration, execution, state, logs, fixtures, and evidence.
| name | benchmark-run-operator |
| description | Operate benchmark runs with controlled launches, attributable artifacts, valid-row rules, and evidence-based recovery. |
Use this skill for benchmark execution. It owns launches, process and stack
state, provenance, valid-run rules, selective recovery, and aggregate rebuilds.
It does not own the parent product outcome or rewrite acceptance assertions.
Return benchmark evidence to using-goals for completion judgment.
Read the references that match the task:
Before expensive work, record:
The user-facing principal authors benchmark policy, queue shape, comparability decisions, and control documents before dispatch. Runners own their assigned processes and return artifacts plus evidence. They do not relabel provenance or rewrite queue policy.
Run the smallest probe that exercises the real child path. A smoke gate must prove:
A scheduler return, quiet terminal, result directory, or reward alone is not completion proof.
The execution owner holds the shell, child processes, logs, immediate local recovery, and output root. Keep controller state separate from runtime state.
During execution:
monitor-to-completion for mechanical waits;Count a row only when the harness produced the required terminal summary and grader result with no uncaught crash, stale-runner marker, missing-summary gap, or known infrastructure poison. Recalculate strict and headline denominators from attributable rows.
Assign every candidate attempt to an evidence epoch, declare the compatible epoch cohort, then select one accepted attempt per cohort, arm, and task through a documented deterministic rule.
Do not mix changed task sets, scorers, tool surfaces, auth state, stacks, or repair conditions without an explicit comparability decision and label.
Classify the failure before another launch: task, tool, evaluator, executor, stack, provider, auth, capacity, or long-run drift. Use a single row or small controlled slice to test a changed hypothesis.
Open the recovery circuit when the failure signature, inputs, and runtime path remain materially unchanged and the prior attempt produced no new discriminating evidence. Resume only after a changed hypothesis, input, runtime path, or root-cause finding is recorded and the next probe is bounded. The orchestration ledger stores this evidence state without retry counts.
Report observed valid-row pace, remaining attributable work, available healthy capacity, and required pace. Do not invent an ETA from scheduler state or nominal concurrency. If the requested deadline is impossible under observed conditions, name the gap and the highest-value bounded action.
On pause, stop owned launches and monitors, preserve process and artifact state, and record the first safe resume action. Do not call pause complete or blocked.
Complete the operator slice only after:
Return that packet to the principal. Benchmark process completion becomes parent completion only when the stable goal defines and verifies that result.