-
Re-read this skill + the target report. The draft lives at aiwatch-reports/NNNN-NN/index.md
(generated already, or generate it — see the README "Local" flow). Note which month, and that the
_data/NNNN-NN.json archive + prior months exist (trend/recurrence read them).
-
Branch in the reports repo: git checkout main && git pull && git checkout -b report/NNNN-NN
(or the generator's draft PR branch). Never author on main. git status must show only the intended
NNNN-NN/ files — this repo often carries an in-progress report/YYYY-MM narrative on another branch.
-
Fill the hand-authored sections — everything the generator leaves as a placeholder/AUTO-DRAFT:
- Summary (EN bullets: Most reliable / Riskiest / High incident count / Watch out) AND the KO
<details> mirror — keep both in lockstep.
- Recommendations table (per use-case). Category coverage is data-driven per month
(aiwatch-reports#62) — do NOT ship the frozen 5-row template blindly:
- Always-present buyer-intent rows (any category can win): Production-critical, Low latency /
cost, General purpose.
- One specialized row per SPECIALIZED bucket that has ≥1 ranked service THIS month. Take the
bucket SET — which specialized buckets exist this month — straight from THIS report's
Services monitored: N — … category-breakdown header (aiwatch-reports#98), using every bucket
EXCEPT the generic LLM APIs (feeds the buyer-intent rows) and AI apps (consumer end-user apps,
not a build-on pick). Do NOT hardcode a bucket list here — the header is the source, so the row set
tracks the fleet as buckets change. The header label (e.g. voice & transcription, inference & infra) identifies WHICH bucket earns a row; keep the row's reader-facing Use-Case cell in the
established polished form (Voice / audio, Inference, Vector / embeddings), NOT the raw header
string.
- "Ranked" = appears in this month's Score Rankings table (it passed the full-month-coverage gate
— aiwatch#802 / aiwatch-reports#45); a mid-month-added or otherwise unranked service does not
qualify. The Rankings table has no category column, so map a ranked service to its bucket via the
same service→category grouping the breakdown header is generated from (the
group taxonomy,
aiwatch#1068 / aiwatch-reports#98) — read it at authoring time; never copy a service roster into
this skill.
- A specialized bucket with no ranked service this month is OMITTED (or explicitly marked
"insufficient data this month") — never fake a pick from an unranked / partial-coverage service.
- Roll forward only. Apply the expanded set from the first month it ships onward; do not retrofit
a single past month — that creates month-to-month divergence in the immutable archive.
- Key Insight — opening sentence + 3 patterns, EN, AND the KO mirror.
- Notable Incidents — top incidents, each with
**Affected** + **Duration**.
- Observations.
- Ground every claim in the report's own data tables (Score Rankings / Incident Summary / Notable
Movers / p75). Never invent a number, and never assert incident detail the archive doesn't carry —
hedge a long-"open"-status artifact (a 264h open window ≠ a 264h hard outage; an Instatus/Nuxt
ongoing incident has no final duration). The AUTO-DRAFT blocks are a starting draft to adapt, not to keep.
- The KO
<details> mirrors must read as NATIVE Korean, not 직역투 (literal EN calques). This is a
recurring critique. The KO is a translation of the meaning, not word-for-word — read each KO sentence
aloud; if it sounds translated, rewrite it. Common calques to avoid: decompose/decomposition → not
「분해」 (use 「지표별 변화 / 내역」); N-month window → not 「N개월 창」 (use 「최근 N개월간」);
read X as Y → 「Y로 해석하다」 not 「Y로 읽다」; binary up/down status page → 「단순 가동/중단만
표시하는」 not 「이진(up/down)」; an unnecessary loanword where an established Korean term exists —
e.g. 「스택」 when it means a product family (not a genuine tech-stack reference) → 「제품군」. Do NOT
nativize deliberate technical terms (MTTR, p75, granularity) — those stay. Match the register/tone of
prior months' KO <details> blocks — but the calque-avoidance above wins if a prior month itself reads translated.
- Claims discipline — hedge, and don't re-teach a standing caveat every month. (1) Don't overclaim
causation. A reporting artifact doesn't disprove a real problem — say a high per-model count
overstates the disruption, NOT that it is "not instability" / "not availability loss" (correlation
≠ causation; the data rarely establishes the negative). Hedge any bare "not X" the tables can't prove.
Before attributing a trend to a mean (MTTR / avg recovery), check the count and the longest — a
mean is skewed by one outlier, so a downtime/MTTR drop dominated by a single freak incident clearing
(e.g. Gemini's count stayed 3→3 while its total 351h→35h collapsed onto one 242h outage not recurring —
shorter incidents, not "closing incidents faster") is not a systematic improvement; report what the
count and longest actually show.
(2) A STANDING methodological caveat — Anthropic's per-model incident counts, no-official-uptime
(Bedrock/Azure), probe-RTT-≠-inference-latency — lives in the About / methodology section and has already
appeared in prior months; do NOT re-explain it in full in the narrative every cycle (a regular reader has
seen it). State it once in methodology; in the narrative mention it only where it's load-bearing for THIS
month's specific point (e.g. the Notable Incident it directly explains), and briefly.
- No redundancy — two overlaps the aiwatch-reports#54/#55 gate CANNOT see (both recurring critiques). The mechanical
aiwatch-reports#54 gate flags a repeated SERVICE in a slot; it is blind to a repeated STORY or FRAMING.
(1) Cross-section, within this report. Each section has a distinct job (Summary = terse headline,
Key Insight = deeper patterns, Notable Incidents = specific events, Observations = what to do) — don't
tell the same service + numbers + Notable-Movers reference in two of them (e.g. a Summary "Biggest
improvement" bullet and a Key Insight "improvements" pattern carrying the identical Copilot/Gemini
figures). Avoid doubling two Key Insight patterns on one theme (two long-incident patterns) — unless the
month genuinely has one dominant story.
(2) Cross-month framing. Stock analytical lessons recur as pattern titles — "a few long outages hurt
more than many short ones", "upstream dependencies inherit their provider's reliability risk". Lead each
Key Insight pattern with what is SPECIFIC to THIS month (the actual movers/events), not a re-headline of a
timeless principle a regular reader has already seen in prior reports. Read the last 2–3 months' Key
Insight pattern titles before writing — if a title paraphrases a prior month's, reframe. (The per-model-
count caveat in the bullet above is the sharpest case; here the concern is the broader analytical framing.)
Carve-out (the only one): the probe / RTT-degradation differentiator (AIWatch's direct-probe data
that status pages typically don't report — "probes caught latency the status pages didn't") is a DELIBERATE
recurring product message, so it MAY repeat monthly and is exempt from the reframe rule — provided each
month it leads with that month's fresh degradations (specific services + counts, e.g. "Mistral 41, Replicate
25") and links to the
#rtt-degradation-detection detail section, not restated as a bare pitch with no new
data. No other recurring framing gets this exemption.
- Notation conventions — keep uniform report-wide. Score
N/100 in the Summary + Recommendations
only (the two standalone-score sections); bare N in Key Insight / Notable Incidents / Observations
prose, and a Score change is bare both ends (86 → 79, never 86/100 → 79/100). When a grade rides with
a score write it score-first — 81 (Good), never Good (81); a grade word alone is fine as a tier
reference (reached Excellent, the only Degrading grade). Keep one unit form per kind of quantity —
durations, latencies (45h 33m, 2030 ms p75). Same value → identical notation in EN and the KO mirror.
-
Heed the RECURRENCE CHECK block (aiwatch-reports#54). If the generator injected a ⚠️ RECURRENCE CHECK block
above ## Summary, a subject (e.g. Together AI) led the same slot in ≥2 recent months. Do not restate
it flat — lead with the month-over-month change ("Together 133 → 85") or pick a fresh lens, using
the MoM delta the block/auto-draft already computed. This is the exact miss the feature exists to catch.
-
Delete EVERY AUTO-DRAFT / RECURRENCE CHECK fence before publishing. Any surviving
<!-- BEGIN AUTO-DRAFT … --> / <!-- … RECURRENCE CHECK … --> in a published: true report hard-fails
the pre-publish CI lint (aiwatch-reports#55) (scripts/lint-recurrence.js). Search the file for AUTO-DRAFT and
RECURRENCE CHECK and confirm zero remain.
-
Service-count / category lockstep. The header line — both the count and the category
breakdown (Services monitored: 41 — 15 LLM APIs, 6 coding agents, …) — is generated
([SERVICE_COUNT] aiwatch-reports#97, [SERVICE_BREAKDOWN] aiwatch-reports#98), counted over the same archive
service set so the two cannot disagree. Do not hand-edit that line. The breakdown half was a
manual literal for months and drifted every time (it read "43 — 33 API services, 6 coding agents,
4 AI apps" while the 2026-06 archive held 41); it is now 8 fine group buckets, not the old 3
coarse ones. If the line looks wrong, the archive or the worker's group field is wrong — fix it at
the source. A generation-time warning of N of M services have no recognised category group (or
NO service carried a recognised category …) means services fell into an other bucket —
investigate rather than editing the rendered text. It reaches a local run as
[generate-report] WARNING: … and CI as a ::warning:: annotation.
Still reconcile by hand any count named in Key Insight and the Score-section count against
that line. A service added mid-month
(has an addedAt) must not inflate a historical month's count; the ranking gate (aiwatch-reports#45)
holds a <full-month-coverage service out of the ranking, so don't narrate it as if fully ranked.
-
Local verify (step 3.5). Serve the draft (drafts are published: false, so --unpublished is
required to render them):
cd ~/Desktop/bentely/aiwatch/aiwatch-reports && \
PATH="/opt/homebrew/opt/ruby/bin:$PATH" GEM_HOME="$HOME/.gem/ruby/4.0.0" \
bundle exec jekyll serve --port 4000 --unpublished
Open http://localhost:4000/reports/NNNN-NN/ (baseurl is /reports). Hand off and WAIT for the
user's in-browser confirmation — do not commit first.
-
One-time correction / forward-looking notes — verify the premise still holds before adding one.
A pre-written note can be invalidated by the very data it comments on (cf. aiwatch#563: a "scores dropped"
note was wrong once the scores actually rose). Read the final tables, then write the note.
-
Commit + publish (only after the user confirms) — flip published: false → true, commit on the
report/NNNN-NN branch with the issue/footer, open/return the PR. The aiwatch-reports#55 lint runs on the PR and
fails on a leaked fence — treat a red lint as a real defect, not flakiness. Merge → Jekyll auto-deploys.
If you touched any script, run the suite the way CI does (a bare node scripts/*.test.js runs only
the FIRST file): for t in scripts/*.test.js; do node "$t"; done.
The report is authored monthly, so the procedure fades from context between cycles / after compaction and
the same mistakes recur (a flat repeated bullet, a nearly-leaked AUTO-DRAFT, service-count drift). Invoking
this re-injects the runbook at authoring time; aiwatch-reports#54/#55 are the mechanical backstops it depends on.