| name | helios-real-workload-ab |
| description | Build and run Helios real-context A/B benchmarks with deterministic workload manifests, correctness gates, and decision artifacts. |
Helios real-workload A/B skill
Use this skill when a task mentions Helios performance, A/B benchmarking, real workloads, benchmark corpus, repeated-prefix testing, or tiny16 performance decisions.
Procedure
- Read
BENCHMARKS.md, docs/xr-real-workload-methodology.md, and docs/xr-ab-methodology.md.
- Identify the current baseline and candidate before editing code.
- Prefer repo-local real contexts over repeated-token prompts.
- Record exact commands, git SHA, model identity, env vars, and output paths.
- Use at least
records.jsonl, summary.json, report.md, blockers.md, and decision.md for every Goal.
- Do not accept performance gains without correctness and memory gates.
- Mark low-trial or high-variance results honestly.
Decision labels
Use exactly one:
accept_candidate
reject_candidate
keep_experimental
needs_more_data
blocked_with_evidence