Skip to main content

documenting-experiment-results

Use when finishing any train, eval, benchmark, profile, telemetry, matrix, or reproduction run — or when a checkpoint was created/promoted without MODEL_CARD/README updates — or when results exist only under outputs/ or in chat without matching docs/design updates

Source facts

Repository
Tyler-R-Kendrick/slm-training
Last source activity
July 18, 2026 at 22:56
Detected SKILL.md language
English
Stars
1
Forks
0

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.

Showing SKILL.md

SKILL.md
Source instructions · Read-only preview
name
documenting-experiment-results
description
Use when finishing any train, eval, benchmark, profile, telemetry, matrix, or reproduction run — or when a checkpoint was created/promoted without MODEL_CARD/README updates — or when results exist only under outputs/ or in chat without matching docs/design updates
# Documenting experiment results ## Overview **Every experiment updates `docs/design/`.** JSON mirrors plus markdown headlines are the durable ledger. `outputs/` alone is not enough. **Every checkpoint also updates `docs/MODEL_CARD.md` and the README model-card summary.** ## Workflow 1. Identify the doc home (map below). 2. Persist JSON under `docs/design/` (scripts often mirror — verify it matches **this** run) and confirm it carries a `version_stamp` (schema `version_stamp/v1`; canonical writers emit it via `slm_training.versioning.build_version_stamp`). If you changed any file watched by `src/slm_training/resources/versions.json`, bump that component (or append a `no-bump:` history note) in the same change — see [`docs/design/version-stamp-contract.md`](../../../docs/design/version-stamp-contract.md). 3. Update markdown measured-results: IDs run, pass/fail, recipe (device, steps, backend, matrix set, suite `n`, honesty mode). 4. **If a checkpoint was written/promoted/synced:** update [`docs/MODEL_CARD.md`](../../../docs/MODEL_CARD.md) (roster, eval table, history, URI) **and** README → “Model card (summary)”. 5. Do not overclaim (fixture clear ≠ production ship). 6. Commit docs with the experiment. ```dot digraph docs_after_run { "Run finished" [shape=doublecircle]; "JSON in docs/design?" [shape=diamond]; "Write/copy JSON" [shape=box]; "Stamp present?" [shape=diamond]; "Re-emit from canonical writer" [shape=box]; "Markdown updated?" [shape=diamond]; "Update measured-results" [shape=box]; "Checkpoint created?" [shape=diamond]; "Update MODEL_CARD + README summary" [shape=box]; "Caveats recorded?" [shape=diamond]; "Add recipe / suite-n" [shape=box]; "Commit with experiment" [shape=doublecircle]; "Run finished" -> "JSON in docs/design?"; "JSON in docs/design?" -> "Write/copy JSON" [label="no"]; "JSON in docs/design?" -> "Stamp present?" [label="yes"]; "Write/copy JSON" -> "Stamp present?"; "Stamp present?" -> "Re-emit from canonical writer" [label="no"]; "Stamp present?" -> "Markdown updated?" [label="yes"]; "Re-emit from canonical writer" -> "Markdown updated?"; "Markdown updated?" -> "Update measured-results" [label="no"]; "Markdown updated?" -> "Checkpoint created?" [label="yes"]; "Update measured-results" -> "Checkpoint created?"; "Checkpoint created?" -> "Update MODEL_CARD + README summary" [label="yes"]; "Checkpoint created?" -> "Caveats recorded?" [label="no"]; "Update MODEL_CARD + README summary" -> "Caveats recorded?"; "Caveats recorded?" -> "Add recipe / suite-n" [label="no"]; "Caveats recorded?" -> "Commit with experiment" [label="yes"]; "Add recipe / suite-n" -> "Commit with experiment"; } ``` ## Artifact → doc map | Run / script | JSON (docs) | Markdown | | --- | --- | --- | | `run_quality_matrix` | `quality-matrix-results.json` | `quality-experiment-matrix.md` (Vn section) | | `run_grammar_matrix` | `grammar-matrix-results.json` | `quality-experiment-matrix.md` (X) | | `reproduce_baseline` | `baseline-reproduction-results.json` | `quality-experiment-matrix.md` | | `run_phase_pipeline` | `phase-abc-results.json` | `quality-experiment-matrix.md` | | `run_perf_matrix` | `perf-matrix-results.json` | `perf-experiment-matrix.md` (+ `runtime-performance.md` if latency claim changes) | | `bench_accel --microbench` | `train-microbench.json` | `runtime-performance.md` / `accel-parallel.md` | | `bench_telemetry` | `cycle-telemetry.json` | `telemetry.md` if narrative changes | | `evaluate_model --ship-gates` | summarize scoreboard/gates for the claim | `adversarial-review.md` and/or matrix doc | | Full HF `train_model` / `remote_train` / `sync_checkpoints` | bucket URI in `checkpoint_bucket.json` | `checkpoint-bucket.md` + **`MODEL_CARD.md` + README summary** | | `bootstrap_playground` / promoted `*.pt` | — | **`MODEL_CARD.md` + README summary** | | `profile_generate` | promote if used for claims | perf / runtime docs | No row? Still add a measured-results note next to the lever's design doc and a `docs/design/*-results.json` matching existing summary shapes. ## Model card fields (minimum) In `docs/MODEL_CARD.md`: - Roster role (demo / matrix champion / production HF ship) - Run id + path or `hf://buckets/TKendrick/OpenUI/checkpoints/<run_id>/` - Eval table with suite `n` and ship pass/fail - Recipe: device, steps, context backend, honesty mode - Version stamp: `code_commit` plus the eval-stack component versions from the run's `scoreboard.json` (`harness.model_build.eval`, `evals.meaningful_program`, `gates.ship`) - History row (append; do not erase predecessors) In `README.md` “Model card (summary)”: one-line table update + link to the full card. Keep it short. ## Markdown shape (matrix results) Match existing sections: link the JSON, state host/recipe/`rico_held` n, table of ID → metrics → ship/perf outcome, honesty caveats. Update in place when re-running the same IDs; keep invalidated historical rows labeled as such. ## Red flags - Results only in stdout / PR body - Result JSON without a `version_stamp` (hand-built instead of writer-emitted) - Versioned component changed without a registry bump / `no-bump:` note - JSON updated, measured-results table not - Checkpoint synced but MODEL_CARD / README summary untouched - "Document after the full matrix" - Ship pass without suite sizes / honesty mode - Weakening gates instead of documenting a fail | Mistake | Fix | | --- | --- | | New JSON schema | Extend the existing summary shape | | Only README long detail | Detail in `MODEL_CARD.md`; README stays a summary | | Training loss as fidelity | Use parse / `placeholder_fidelity` / struct / reward | | Skipping failed runs | Document fail + blocker |
View on GitHub