원클릭으로
compare-nightly
Compare two nightly GitHub Actions runs and classify failures as new, persisting, or fixed
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
메뉴
Compare two nightly GitHub Actions runs and classify failures as new, persisting, or fixed
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
SOC 직업 분류 기준
Stash local changes, pull latest main, and display a structured commit summary. For tt_forge_models uplift commits, shows the affected models from the submodule log.
Analyze a PyTorch model, then design and implement an optimal multi-device sharding strategy for Tenstorrent hardware that minimizes collective-communication (CCL) ops. A strategy is both which parallelism to use (tensor / sequence / data) and how to shard under it - e.g. Megatron column→row is one of several tensor-parallel schemes. Use whenever the user asks to shard a model, distribute it across devices, write Shardy annotations, reduce CCLs, or analyze a model for tensor/sequence parallelism. Trigger even when the user only says "sharding strategy" or names a model with a device count, without asking for a full plan.
Analyze CI benchmark workflow runs from GitHub Actions for the tt-xla project. Produces a markdown report covering failed jobs (with root-cause error extraction via logs and Glean), successful model performance metrics (samples/sec, TTFT, device perf), perf regressions/improvements vs previous nightly, and the full dependency commit chain (tt-xla, tt-mlir, tt-metal). Use this skill whenever the user wants to analyze a CI run, review nightly benchmark results, investigate CI failures, check benchmark performance from a workflow run, or asks about "latest nightly" results. Also trigger when the user pastes a GitHub Actions run URL or mentions a run ID in the context of performance analysis, or asks about perf regressions.
Review and fix language correctness and formatting in documentation files.
Triage one tt-forge-models training test failing with a bfloat16 dtype-mismatch RuntimeError (e.g. "mat1 and mat2 must have the same dtype, but got Float and BFloat16", "'<op>' not implemented for 'BFloat16'"). For cross-dtype operands, attempts a minimal loader fix propagating `dtype_override` into the offending tensor constructor, then re-runs CPU + pytest and updates the YAML (passing -> EXPECTED_PASSING; new failure -> KNOWN_FAILURE_XFAIL). For op-not-implemented (no PyTorch kernel), goes straight to KNOWN_FAILURE_XFAIL with the verbatim error. Updates every training entry sharing the affected loader. Never edits inference YAML or `dynamic_loader.py`.
Triage one tt-forge-models training test stuck at FAILED_FE_COMPILATION with reason "tt-forge-models doesn't implement unpack_forward_output for this model." Inspects the model's forward output, registers a handler or writes a per-loader override, and updates the YAML.
| name | compare-nightly |
| description | Compare two nightly GitHub Actions runs and classify failures as new, persisting, or fixed |
| disable-model-invocation | true |
| allowed-tools | Read, Read(/tmp/**), Write(/tmp/**), Glob, Grep, Skill, Bash(gh api *), Bash(gh api * > /tmp/**), Bash(jq *), Bash(mkdir -p /tmp/**), Bash(rm -rf /tmp/**), Bash(python3 *), Bash(for *) |
| context | fork |
| argument-hint | baseline-run-id comparison-run-id |
| model | opus |
Compare two nightly GitHub Actions runs and classify every test failure as:
The baseline run-id is $0 (typically the older/yesterday run). The comparison run-id is $1 (typically the newer/today run). The GitHub repo is github.com/tenstorrent/tt-xla.
Invoke the analyze-nightly skill twice using the Skill tool, once for each run, passing the "save" argument so results are stored to disk:
analyze-nightly with args $0 save
— waits for completion, then reads /tmp/nightly-analysis-$0.md (baseline analysis)analyze-nightly with args $1 save
— waits for completion, then reads /tmp/nightly-analysis-$1.md (comparison analysis)Do not proceed to Step 2 until both analyses are complete and both files exist.
A reference Python script analyse_failures.py lives alongside this skill file.
Copy it to /tmp/ and run it to parse both analysis files and produce the comparison:
cp .claude/skills/compare-nightly/analyse_failures.py /tmp/analyse_failures.py
python3 /tmp/analyse_failures.py /tmp/nightly-analysis-$0.md /tmp/nightly-analysis-$1.md
The script outputs a Markdown report covering New / Persisting / Fixed sections directly.
Pass --json to get structured JSON output instead if you need to post-process results.
The script handles:
# ownership-area, ## root-cause, and - test_id (arch) -> [link](url) lines
(any link label is accepted, e.g. [job-link], [baseline job], [comparison])Matching across runs is two-pass. Pass 1 matches on the arch-normalised test
ID. Pass 2 then reconciles the leftovers by token signature, to absorb the case
where the two analyze-nightly passes labelled the SAME underlying failure with
different naming conventions (e.g. a perf vllm_bge_m3 label vs a
test_vllm_benchmarks.py::...[bge_m3] pytest nodeid). Without this, such a
failure was double-counted as both New and Fixed — the known cause of inflated
new/fixed counts. Reconciliation requires ≥2 shared discriminating tokens and
≥0.8 containment (so e.g. qwen3_4b is never merged with qwen3_8b, nor
batch1 with batch32), and uses greedy best-first assignment.
Trust the script's counts — do not hand-correct them as "unreliable." When the two files use inconsistent naming, the script now reconciles automatically:
_(matched across naming drift)_, with a > **Note:** block at the top of the
report and WARNING: lines on stderr listing each reconciled pair.pcc=nan ->
std::bad_alloc) is tagged _(root cause changed: "..." -> "...")_.N new, N persisting (N reconciled across naming drift), N fixed.In --json mode the object additionally contains a reconciled array (the
persisting items that were matched via pass 2) and a warnings array (one string
per reconciled pair). Each persisting item carries reconciled,
root_cause_changed, and root_cause_baseline fields.
If the script is unavailable or the output needs manual correction, fall back to parsing
the files by hand: for each bullet line extract test_id (text before the first (),
arch (text between ( and )), root_cause (most recent ## heading), and
job_url (URL inside [job-link](...)). Build two lists — BASELINE_FAILURES and
COMPARISON_FAILURES — then proceed to Step 4.
Fetch the head SHA that each run was triggered on:
gh api "repos/tenstorrent/tt-xla/actions/runs/{run-id}" --jq '.head_sha'
Do this for both $0 (baseline SHA) and $1 (comparison SHA).
Then fetch the full list of commits on the main branch between those two SHAs (exclusive
of the baseline SHA, inclusive of the comparison SHA):
gh api "repos/tenstorrent/tt-xla/compare/{baseline-sha}...{comparison-sha}" --jq '.commits[] | {sha: .sha, short_sha: .sha[0:8], author: .commit.author.name, message: (.commit.message | split("\n")[0])}'
If the API returns an empty list or an error, note it and continue.
For each commit, produce a structured per-commit summary using a for loop.
For each commit SHA, fetch the full commit details in a single call:
gh api "repos/tenstorrent/tt-xla/commits/{sha}" --jq '{message: .commit.message, files: ([.files[].filename | split("/")[0]] | unique)}'
This returns:
message: the full commit message (title + body, which often includes the PR description)files: top-level directories of changed filesFrom the commit body, derive:
Produce a structured summary:
**`{short_sha}`** — {commit title} (by {author})
- **Why:** {motivation}
- **Files:** {comma-separated list of top-level directories or key files changed}
- **Benefit:** {what this improves}
Separate each commit summary with a horizontal rule (---).
If a commit title contains "Uplift third_party/tt_forge_models" or a changed file path
is third_party/tt_forge_models, apply the following extra steps instead of the generic
"Files" line:
Extract the old and new submodule SHAs from the commit's patch. Fetch the commit and
find the file entry for third_party/tt_forge_models:
gh api "repos/tenstorrent/tt-xla/commits/{sha}" --jq '.files[] | select(.filename == "third_party/tt_forge_models") | .patch'
Parse the patch for lines -Subproject commit {OLD_SUB_SHA} and +Subproject commit {NEW_SUB_SHA}.
List commits in the tt-forge-models submodule between those two pointers:
gh api "repos/tenstorrent/tt-forge-models/compare/{OLD_SUB_SHA}...{NEW_SUB_SHA}" --jq '.commits[] | {sha: .sha[0:8], full_sha: .sha, message: (.commit.message | split("\n")[0])}'
For each submodule commit, fetch the top-level model directories changed:
gh api "repos/tenstorrent/tt-forge-models/commits/{full_sha}" --jq '[.files[].filename | split("/")[0]] | unique | .[]'
Replace the Files line with a Models affected table:
| Submodule commit | Change | Model(s) |
|---|---|---|
{short_sha} | {commit title} | model_a, model_b |
If analyse_failures.py was used in Step 2 it already computed and rendered the
three sets, including any naming-drift reconciliation. Incorporate that output
into the Step 5 report verbatim — keep its > **Note:** block and the
_(matched across naming drift)_ / _(root cause changed: ...)_ tags. Do not
second-guess or hand-correct the counts; the reconciliation logic exists
precisely to handle the naming-convention mismatch that previously made manual
correction necessary.
If falling back to manual comparison: a failure in COMPARISON_FAILURES matches a
failure in BASELINE_FAILURES when the test_id strings are identical, or differ only
in a trailing hardware arch token (e.g. -n150, -p150, -n300-llmbox). Compute:
When matching by hand, also reconcile entries that name the same underlying
test under different conventions — most commonly a perf label such as
perf vllm_bge_m3 in one file vs a pytest nodeid such as
test_vllm_benchmarks.py::...[bge_m3] in the other. These are the SAME test;
classify them as persisting, not as both new and fixed. Match on the shared
discriminating tokens (model name, size, batch), and keep the size/batch distinct
(qwen3_4b ≠ qwen3_8b, batch1 ≠ batch32). If a persisting test's root
cause differs between the two files, note the change rather than treating it as
two separate failures.
Output the following Markdown report. Do not emit a section if it is empty.
# Commits Merged Between Runs
_(baseline: {baseline-sha}, comparison: {comparison-sha})_
**`{short_sha}`** — {commit title} (by {author})
- **Why:** {motivation}
- **Files:** {top-level directories changed}
- **Benefit:** {what this improves}
---
**`{short_sha}`** — Uplift third_party/tt_forge_models (by {author})
| Submodule commit | Change | Model(s) |
|---|---|---|
| `{sub_sha}` | {submodule commit title} | `model_a`, `model_b` |
---
# New Failures (in $1, not in $0)
## {root-cause}
- {test_id} ({arch}) -> [job-link]({job_url})
# Persisting Failures (in both $0 and $1)
## {root-cause}
- {test_id} ({arch}) -> [job-link]({job_url in $1}) [also in baseline]({job_url in $0})
# Fixed Since Baseline (in $0, not in $1)
## {root-cause}
- {test_id} ({arch}) -> [baseline job]({job_url in $0})
Group failures within each section by root cause (same heading = same root cause). If the same test fails on multiple hardware arches with the same root cause, list it once with all arches as a comma-separated list in the ({arch}) field.
gh command that modifies repository state (no issue/PR creation,
branch changes, etc.). Read-only calls only./tmp/nightly-analysis-*.md files until after comparison is complete./tmp/nightly-analysis-{run-id}.md exists,
you may skip re-invoking analyze-nightly for that run and read the cached file directly.