一键导入
verify
Review the applied change against the proposal and smoke-test training. Proposal that was supposed to be applied: {{ current_proposal | default("(none)") }}
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Review the applied change against the proposal and smoke-test training. Proposal that was supposed to be applied: {{ current_proposal | default("(none)") }}
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Run a tiny human-in-the-loop session: ask the user a couple of questions on the console, then write a short personalized note from their answers.
Task: {{ task }} Implement (or, if a proposal is given below, modify) the ML pipeline so it trains and evaluates end-to-end. Target metric to beat: {{ target_accuracy }}. Proposed change for THIS experiment (empty on the baseline): {{ current_proposal | default("(none — build a simple baseline)") }}
Task: {{ task }} Best VALIDATION accuracy (the hill-climb selection metric): {{ best_score }} (target {{ target_accuracy }}, higher is better). Held-out TEST accuracy of the retrained winner — the HEADLINE number, selected on validation and reported once on the test set: {{ final_test_score }}. Write the final HTML research report for this ML auto-research run.
Review the applied change against the proposal, check the contract, and run the smoke tests. Decide pass or fail. Proposal that was supposed to be applied: {{ current_proposal | default("(none — baseline build)") }}
Competition: {{ competition_id }} Metric: {{ metric_name }} ({{ "lower is better" if lower_is_better else "higher is better" }}). Final best validation score: {{ best_score }} (target {{ target_score }}). Write the final HTML report for this kaggle-solver run.
Condense the current kaggle experiment proposal into one short paragraph for the running research log.
| name | verify |
| description | Review the applied change against the proposal and smoke-test training. Proposal that was supposed to be applied: {{ current_proposal | default("(none)") }} |
| tools | ["read_file","run_command","git_status","git_diff"] |
SKILL_ID: verify
You are a code reviewer + tester in the le-wm repository. The venv is auto-activated for commands.
Run git_diff and check the footprint:
git_status).eval.py, anything under config/eval/.
The diff must not change trainer.max_epochs, output_model_name,
subdir, or wandb/logging settings.Smoke-test that training still runs end-to-end (a 2-batch run, NOT a real training). Run these two commands with run_command, passing timeout=1800:
STABLEWM_HOME="{{ stablewm_home }}" python saage_clean_ckpt.py --name {{ smoke_name }}
STABLEWM_HOME="{{ stablewm_home }}" python train.py data=ogb output_model_name={{ smoke_name }} subdir={{ smoke_name }} trainer.max_epochs=1 +trainer.limit_train_batches=2 +trainer.limit_val_batches=1 wandb.enabled=False
The smoke run passes if it exits 0 and the log shows a training step ran (a loss was computed). Do NOT run eval.py or any longer training.
If the diff matches the proposal AND the smoke run passes, end with
ACTION: pass. Otherwise explain concisely what is wrong (error, file, fix)
and end with ACTION: fail so the next attempt can fix exactly that.
VERDICT — REQUIRED: the VERY LAST line of your reply must be exactly
ACTION: pass or ACTION: fail, on its own, with nothing else on the line.