| name | benchmark-config-validation |
| description | Use before adding or changing a deep-swe-bench config release, model leaf, provider/model API path, config lock, role declaration, usage parser, smoke contract, or extension/subagent worker usage accounting. |
Benchmark Config Validation
A config release is not ready for confirmed planning until its identity, lock,
roles, compatibility, and durable preflight evidence are reviewable.
Config prompt text is approval-gated. Never invent or alter config-authored
prompt text — system_preamble.md, orchestration.md,
--append-system-prompt, or any other instruction surface — without approval of
the exact wording. Supposedly neutral guidance such as “work normally” or “use
your judgment” still needs approval. Allowed exceptions are prompt/tool surfaces
registered by the extension or tool under test itself: tool definitions, prompt
snippets, prompt guidelines, and extension-owned hook output. If extra wording
seems necessary, propose the exact text and wait for approval before writing it.
Process
-
Name the release and impact.
- New maintained releases use
<behavior-name>@<major>.<minor>.<patch>.
Keep the behavior name stable across releases; reject vague lineage suffixes
such as -v2, -new, or -latest.
- Choose explicit version impact:
reuse, recompute, or rerun. The field
is descriptive and must not copy, migrate, regrade, recompute, or rerun old
artifacts automatically.
- Existing unversioned configs remain readable legacy evidence, but they lack
modern lock provenance. Never fabricate a lock for historical results.
- Completion: identity and impact are explicit and use canonical vocabulary.
-
Validate provider and thinking paths.