Skip to main content

verify-decoder

Audit a predecoder checkpoint against holdout seeds. Runs independent_eval (3 fair-baseline guards) and interprets borderline cases with LLM reasoning. Use when a round produces a promising Δ_LER and the user wants to confirm it is not a reward-hacking artifact.

Informações da origem

Repositório
qualit527/qec-ai-decoder
Última atividade na origem
22 de abril de 2026 às 05:58
Idioma detectado do SKILL.md
inglês
Estrelas
3
Forks
0

Opções de instalação

Por padrão, está selecionado o prompt que primeiro revisa a origem. Você pode mudar para um comando direto ou baixar uma cópia local.

Revise os arquivos de origem

Leia o SKILL.md e os arquivos complementares exibidos pelo SkillsMP antes de decidir se vai instalar.

Exibindo SKILL.md

SKILL.md
Instruções da origem · Visualização somente leitura
name
verify-decoder
description
Audit a predecoder checkpoint against holdout seeds. Runs independent_eval (3 fair-baseline guards) and interprets borderline cases with LLM reasoning. Use when a round produces a promising Δ_LER and the user wants to confirm it is not a reward-hacking artifact.
# /verify-decoder ## When to use - An AutoQEC round produced `delta_ler > 0` and the user wants final sign-off. - User asks to "verify this checkpoint" or "audit this round". ## Inputs - `round_dir`: path to `runs/<id>/round_N/` (must contain `checkpoint.pt` + `config.yaml`) - `env_yaml`: env used by the round (defaults to `config.yaml`'s `env_name`) ## Behavior 1. Run `python -m cli.autoqec verify <round_dir> --env <env_yaml>`. 2. Read `verification_report.json` + `training.log`. 3. LLM-reason over: - If `verdict=VERIFIED`: write one paragraph noting key confidence intervals and whether delta is within ablation threshold. Output "APPROVED". - If `verdict=SUSPICIOUS`: read training.log, the DSL config, the history of previous rounds; diagnose whether this is (a) genuine small improvement, (b) overfitting, (c) reward-hacking. Recommend next action: accept, re-run with more shots, or reject. - If `verdict=FAILED`: produce a diagnostic report matching the failure pattern (seed leak, ablation failure, negative delta). Archive into `round_N/failure_diagnosis.md`. ## Tool-use rules - Read: training.log, config.yaml, verification_report.md, previous round metrics.json files. - Bash: `python -m cli.autoqec verify ...` only. No other commands. ## Output - Short decision block: APPROVED / SUSPICIOUS_KEEP / SUSPICIOUS_REJECT / FAILED_REJECT. - Diagnostic paragraph saved to `round_N/decision.md`.
Ver no GitHub