Find unsupported, overbroad, mis-scoped, or approval-sensitive claims before they ship.
-
Exact model names, versions, benchmarks, datasets, splits, metrics, scores, costs, and run counts.
-
Whether the claim is supported by a paper, leaderboard, blog, internal run, third-party eval, or researcher input.
-
Whether comparisons are apples-to-apples.
-
Whether benchmark language states the harness, scoring setup, split, metric direction, and caveats.
-
Whether partner, customer, deployment, legal, safety, or external-positioning claims need clearance.
-
Whether abstract claims such as trustworthy, responsible, real-world, meaningful, transformative, open, or community-driven are backed by a concrete mechanism, artifact, or approval.
-
Whether every model or system comparison carries the full scaffold: absolute score, delta from a named baseline, and scope (harness, split, run setting).
-
Whether state-of-the-art appears bare. Bare use is a blocker; the phrase is permitted only with an explicit scope qualifier.
-
Whether closed-model competitors are described with factual setup descriptors (proprietary, closed weights, opaque training data) rather than editorialized.
-
Whether Ai2 signature phrases (fully open, reproducible, transparency, community) are paired with the concrete artifact or mechanism they name.
-
Subject-verb agency: whether the subject of each active-voice clause is actually capable of the verb (models release weights fails — models do not release; teams do).
-
Pronoun referent precision: whether it, they, them, those, the model have a clear antecedent in the immediately prior noun phrase, not a referent that has drifted across multiple sentences.
-
Multi-org attribution: whether multiple involved organizations are separated into roles (operator, funders, infrastructure partners) rather than blended into one verb.
-
Sibling-product descriptions: whether a passing one-line description of another Ai2 product or platform (Skylight, EarthRanger, OlmoEarth, Olmo, Tülu, Molmo, etc.) is verified against that product's own Ai2 page rather than written from memory. Scope drifts between model, platform, and suite, and between the domains a product serves. Treat an unverified gloss as Verify, not Clean.
-
Adoption-artifact precision: in third-party build-on copy, whether used [platform] implies fine-tuning a model when the source says a dataset was used (or a different base model), and whether a single generation is named where the source lists several or names a dataset. Treat an unverified umbrella attribution as Caveat, not Clean. See .agent/skills/ai2-comms-style-source/references/source-fidelity.md (Which Open Artifact, And Which Generation).
-
Drift-back after a plain-language pass: whether an accessibility rewrite introduced overclaim (realistic → real-world), dropped a fact or capability (three options silently became two), or swapped a precise term for a meaning-changing synonym (isolated → secure; harnesses → benchmarks). Re-diff against the source for these three failure modes. See .agent/skills/ai2-comms-writer/references/revision-passes.md Pass 8.
-
Lineage facts and naming collisions: when a release is framed as a successor or extension of prior work, whether every prior-work fact (name, acronym, paper, repo, what it did) is sourced from the real paper or repo, and whether the new name resolves publicly to something else (a retired predecessor, a sibling library). Treat any unsourced lineage fact as Verify; surface naming collisions as a positioning flag for ai2-comms-launch-engagement. See .agent/skills/ai2-comms-style-source/references/source-fidelity.md (Lineage Facts And Naming Collisions).
-
Internal or staging links as public CTAs: whether a repo, doc, dashboard, or demo URL supplied for a public surface is actually internal or staging (*-internal, staging., a private host). Treat any such link as a Blocker on external surfaces; the public URL must be confirmed before publication, and a link cannot back a releasing it openly claim until it resolves publicly. Extension — the opening post of a thread: X and Bluesky threads now carry the canonical destination link in post 1, so that URL is load-bearing at T-0 rather than a last-post detail. A [blog link] placeholder or an unresolved URL in post 1 is the same Blocker, and it lands on the most visible post in the thread.
-
Peer-tool descriptor sourcing: in peer or open-tool comparisons, whether every descriptor of the peer (architecture, providers, feature list) is sourced from the peer's own documentation rather than memory, and whether non-exhaustive lists are rendered as such as ... rather than closed sets. Drop unverifiable names rather than guess. See .agent/skills/ai2-comms-style-source/references/claims-and-benchmarks.md (Peer And Open-Tool Comparisons).
-
Citation discipline for statistical methods: when a paper is cited as the source of a method, whether the cited paper actually contains the specific illustrative numbers used in copy. If not, cite the method only and mark numbers illustrative; add a resolvable reference (arXiv ID, DOI, URL). See .agent/skills/ai2-comms-style-source/references/claims-and-benchmarks.md (Citation Discipline For Statistical Methods).
-
Definite article superlatives: whether the [thing] implies uniqueness or top-rank where multiple exist; switch to indefinite (a cross-model leaderboard widely tracked) when the reference is type-of-thing.
-
Source-figure consistency: whether numeric figures match across source sections (an intro saying 720 hours and a body saying 700 hours is a flag to surface, not silently choose). Extension: after an update banner supersedes prior facts, scan the body for now-contradictory numbers, dates, or allocations.
-
Precision vs. accuracy: whether a percentage labeled accurate is actually a precision figure (verified / verified+refuted) with no recall measurement. Flag and replace with the plain-language precision rendering (% of what it found held up when checked).
-
Relationship counts vs. distinct-entity counts: whether a number that counts edges, dependencies, mentions, or references is being phrased as a count of distinct entities. Flag and match the noun to what the source table counts. Also flag missing direct only vs. direct plus indirect labels when the source distinguishes them.
-
Verification-independence fidelity: whether independently verified, third-party verified, or manually/human verified matches who or what actually did the verifying. A verifier from the same model family as the system under test is not independent. An automated verifier is not manual. Drop the modifier or replace with the conservative-true phrasing.
-
Totality words vs. coverage caveats (intra-piece): whether totality words (all, every, complete, full) anywhere in the piece contradict a lower-bound or partial-coverage caveat stated elsewhere in the same piece. If the piece says lower bound or what we could recover, it cannot also say all or every of the same quantity. Drop the totality word.
-
State verbs vs. action verbs in agency: an account has or holds a value (state verbs), but keeps, receives, spends, claims are actions that require an agent (the user, not the account).
-
Population scope: when a number is tied to a group, name the right population. New users receive or start with; existing users have or already hold. The same number rarely holds across populations.
-
Verb tier vs. evidentiary status: surfaced / flagged (candidate) < found / identified (in-data result) < validated / proved (externally confirmed). Do not promote a candidate to a confirmed finding.
-
Cadence and frequency words: whether continuous, continuously, real-time, always-on, or ongoing match the system's actual cadence. If it runs ad hoc, on-change, or periodically, the inflated word is a flag; replace with the true cadence.
-
Expiry vs. window: distinguish a date that marks when current terms hold (a window, fine) from a date that asserts access ends (a cliff, only if source-supported). Do not manufacture urgency.
-
Robust phrasing under unconfirmed facts: when an operational fact is unconfirmed, phrase the claim so it holds under either branch and flag the open fact as Verify.
-
Fabricated or unsourced quotes: whether every quotation is either present in the source or an explicitly labeled proposed draft for a named, approved spokesperson. An unlabeled quote that is not in the source is a Blocker (fabrication); a proposed quote must carry a draft/approval label and a bracketed [Name, title]. See .agent/skills/ai2-comms-style-source/references/source-fidelity.md (Quotes).
-
Prototype vs. routine capability: whether now runs, now analyzes, or habitual present tense presents a one-time prototype, pilot, or first run as a standing, repeatable capability. Scope to the test (in an early test, past tense). See .agent/skills/ai2-comms-style-source/references/claims-and-benchmarks.md (Prototype vs. Routine Capability).
-
Roadmap item vs. shipped capability (forward-looking closes): whether an aspirational or forward-looking close converts a roadmap item (moving toward, could one day, building toward) into a present or shipped capability, or implies present-tense self-serve (a team can use it to) when the product is operated for partners today. Keep the reach hedged as a stated goal so the future end-state does not read as live; scope self-serve language to how the product actually ships now. Treat a roadmap item asserted as present as Caveat, and an unconfirmed present-tense self-serve claim as Verify. Distinct from Prototype vs. routine capability above (that covers a past first run sold as a standing capability; this covers a future plan sold as a present one).
-
Per-run cost generalized, or constraint misnamed: whether a cost figure measured for one run ($X for this run) is stated as the price an arbitrary future run or partner faces, or whether copy names an adjacent resource as the limit (data budget when the satellite data is open and the real cap is compute/cloud cost). Scope the figure to the run that produced it and name the resource that actually binds. Treat either as Caveat. See .agent/skills/ai2-comms-style-source/references/claims-and-benchmarks.md (A Per-Run Cost Is Not A Standing Price).
-
Effect-verb strength vs. source: whether an effect verb claims more than the source states. If the source says reduce(s), copy must not escalate to clears, eliminates, removes, or fixes entirely; if the source says performance improves a bit, copy must not inflate to major or large. This is distinct from verb tier vs. evidentiary status (that bullet covers evidence; this covers the magnitude of a stated effect). Match the verb to the source claim, not to the desired impression. Treat an escalated effect verb as Caveat.
-
Update-framing verb vs. prior state: on an update or extension, whether the framing verb implies the capability is brand-new when the release only improves a path that already existed. connects X and Y says X and Y were disconnected; if a manual or slower handoff already worked, scope the verb to the change (more closely integrates, streamlines, makes X one step). Distinct from Effect-verb strength above (that covers the magnitude of a stated effect; this covers whether the capability existed at all before the update). Treat a verb that erases the prior path as Caveat. Anchor: a draft read the second update connects two of Asta's agents, but a manual handoff between them already existed; fixed to the second update more closely integrates two of Asta's agents.
-
Invented benchmark or task scope: when the source states a directional gain with no per-task breakdown (and may add still determining how much), whether copy names specific tasks or attaches a number the source never gave. Keep the generic scope (downstream tasks) and leave the magnitude hedged (a small boost, still measuring); do not fill a better on [tasks] slot from nothing. Treat an invented task list or magnitude as Caveat, and request the per-task source as Verify.
-
Mechanism-description fidelity: when copy explains how a method works, whether the description matches the actual method and names the right object. RoPE rotates the model's internal representations is loose — it rotates the query and key vectors compared in attention, by an angle set by position. Verify the mechanism against the paper or standard definition, and offer a tiered fix (keep the lay gloss, tighten the object, or state the plain-language payoff) rather than shipping an inaccurate simplification. Treat an unverified mechanism description as Verify.
-
Metric-as-value gloss: whether a plain-language gloss of a tool promotes its internal scoring metric into a value, importance, or correctness claim the tool does not make. Naming the metric faithfully is fine; reinterpreting it as a ranking of worth is not. AutoDiscovery's Bayesian-surprise score ranks by how much a result shifts the model's prior (belief shift), not by whether a hypothesis is important or correct — so glossing it as surfacing the most important (or best, most likely to be real) hypotheses overclaims. Describe the user-facing function instead. Treat as Caveat. Anchor: surfaces the most important hypotheses → explores your datasets and surfaces hypotheses worth investigating. Distinct from Verb tier vs. evidentiary status (evidence status: candidate → confirmed) and Mechanism-description fidelity (whether the how-it-works description names the right object); this covers a scoring metric promoted into a value claim.
-
Research-amplification stickies (paper link, venue, attribution): on a blog amplifying a paper, whether the post carries a resolvable paper link (arXiv ID, DOI) as a closing CTA, names the venue when the source states it, and credits the collaborating institutions. These drop out as prose is condensed across revision passes; re-check their presence on every pass and reinstate, the way the unofficial-artifact disclaimer is sticky (see .agent/skills/ai2-comms-style-source/references/approval-gates.md). A research blog with no paper link is a Blocker; missing venue or institutional credit is a Caveat. Exception for the testimonial overlay: the artifact-link footer is frequently added at publication rather than written into the draft, so surface its absence as a Caveat and ask the author whether the CMS carries it, rather than blocking a draft that is otherwise finished.
-
Unconfirmed venue designation: whether a presentation format (oral, spotlight, best paper) is asserted as fact when sources disagree (e.g. the arXiv comment says oral, a web summary says spotlight). Soften to accepted at [venue] until confirmed against the official program; do not silently pick one. Treat as Verify.
-
Formalism fidelity in technical copy: whether every equation, named construct, theorem/proposition reference, and number matches the source, and whether condensing dropped or garbled the load-bearing term. Compressing a derivation can silently delete the term that carries the argument (the Gaussian kernel in an attention-equals-KDE claim) or leave a symbol undefined (reweighted by w_j with no w_j on the page). Verify each formal claim against the source PDF; treat dropped or garbled formalism as Caveat and an undefined symbol as a clarity flag. See .agent/skills/ai2-comms-style-source/references/source-fidelity.md (Keep The Load-Bearing Term When Condensing Formalism).
-
Checkpoint-result match: whether a result is attributed to the right model checkpoint — a pretraining diagnostic (next-token loss, perplexity, token-level NLL) comes from the base/pretrained model, not the strongest post-trained variant, and an instruction/chat benchmark comes from the post-trained one. A bare self-superlative (our strongest 7B) both editorializes and can misname the checkpoint a base-model result came from. Treat a mismatched checkpoint as Caveat and a superlative standing in for the checkpoint as a scope flag. See .agent/skills/ai2-comms-style-source/references/claims-and-benchmarks.md (Name The Right Checkpoint For The Result).
-
Hedge and scope-qualifier fidelity: whether condensing dropped a source hedge (essentially, typically, in most cases) or a scoping adjective (expressive, structured, sufficiently large) and thereby strengthened or broadened the claim (expressive recurrent layers solve it essentially perfectly → recurrent layers solve it perfectly). A dropped hedge is a strength change; treat as Caveat. Related: keep a source's unit (a distance in words) rather than converting it to match the piece's framing (tokens). See .agent/skills/ai2-comms-style-source/references/source-fidelity.md (Keep Hedges, Scope Qualifiers, And Units When Condensing).
-
Measurement-caused vs capability-caused change: whether a score move is attributed to the artifact when its real cause is the evaluation (the benchmark does not cover the new data, regime, or language). A coverage gap sold as a quality trade-off or a quality gain is a misread; attribute it to the eval (the evaluation covers only X). Treat a measurement artifact framed as a regression or improvement as Caveat. See .agent/skills/ai2-comms-style-source/references/claims-and-benchmarks.md (A Measurement-Caused Score Change Is Not A Capability Change).
-
Real-world practice claims (crawling, opt-outs, PII, licensing): whether a claim about how the org actually behaves is sourced from first-party material and scoped to exactly what that source covers. A self-crawl policy (identifies as a named bot, adheres to robots.txt) does not extend to third-party-sourced data (Common Crawl, Internet Archive); do not generalize, and do not assert an opt-out or takedown mechanism the source does not document. Treat an unsourced or over-scoped practice claim as Verify, and a narrow-claim-generalized-to-the-whole-corpus as Caveat. See .agent/skills/ai2-comms-style-source/references/source-fidelity.md (Scope Real-World Practice Claims To What The Source Covers).
-
Tool-vs-feature attribution (Ai2 tooling): whether a capability is credited to the right tool in a stack. infini-gram is the lookup primitive (counts a phrase across the public corpora it indexes, returns where it occurs); OLMoTrace, built on it, traces a model's own output back to its own training data. Crediting infini-gram with OLMoTrace's output→training-data tracing, or implying infini-gram searches a specific model's private training set (it searches public corpora), is a mechanism error. Treat as Caveat; see .agent/skills/ai2-comms-style-source/references/claims-and-benchmarks.md (Name The Right Tool For The Capability).
-
Contested or alleged attribution: whether a disputed public claim (a work flagged as AI-generated, an authorship dispute) is asserted as settled fact, or a phrase match is presented as proof of copying/plagiarism. Keep contested attributions at flagged/alleged/scored as, and frame a match as circumstantial evidence/corroborates, never copied/plagiarized (a causal claim a match cannot carry). Treat an un-hedged contested attribution or a match-as-proof claim as Caveat. See .agent/skills/ai2-comms-style-source/references/claims-and-benchmarks.md (Contested Or Alleged Attribution Stays Hedged).
-
Provided reference points to the right work: when the copy cites a paper or link the user supplied, verify the URL or arXiv ID resolves to that exact work (title, authors, date) before keeping it — a supplied ID can point to a different paper. Treat an unverified citation as Verify and a confirmed-wrong link as a Blocker. Anchor: a draft cited arXiv 2606.15643 as Fluid Benchmarking, but that ID is a different IRT paper (multilingual evaluation); the real one is 2509.11106.
-
Acceptance asserted from an under-review submission: whether accepted at [venue] is claimed when the source PDF is an anonymized or under-review submission. Tells: a footnote like code released upon acceptance or anonymized version included with the submission, or the standard conference footer on a preprint. Soften to under review at [venue] or flag for confirmation; do not assert acceptance the source does not establish. Treat as Verify. Extends Unconfirmed venue designation (there the open question is the presentation format; here it is acceptance itself).
-
Incomplete or dangling claim: whether a comparison or correlation trails off without its object or value — the safety dimension strongly correlates (with what?), about 11% more accurate (than which baseline?), a correlation asserted with no number. A claim missing its second term is not yet a claim; complete it or cut it. Treat as Blocker. Companion to the comparison-framing scaffold (absolute + delta + scope) and the bare-comparison-number-needs-referent check.
-
Favorable-framing re-diff: when copy is revised to read more positively or celebratory (a request to make it positive, cut the failure cases, or lead on the wins), re-diff against source for overclaim introduced by the reframe — a hedged or underpowered result promoted to a confirmed one, a real limitation softened away, or a caveat re-attached to a different finding. The positivity pass is its own trigger, distinct from the plain-language drift-back above. Treat a hedged result sold as confirmed as a Blocker. Anchor: a result significant only at large N with a tiny effect (Cohen's d ≈ 0.13) recast as a trend that held up, with the team's own over-reading caveat moved onto a separate finding. Companion to the gloss-a-bare-statistic and scope-an-emergent-finding checks in .agent/skills/ai2-comms-style-source/references/claims-and-benchmarks.md. The positive reframe of a lifted restriction is the compressed-surface case: restating no longer rate-limited as a gained capability is correct channel practice, but the stated gain must stay at or below the source's own ceiling (full throughput supports high speed, not instantly); see .agent/skills/ai2-comms-channel-adapter/references/platform-patterns.md (Frame A Removed Limit As The Gained Capability).
-
Vagueness fix that escalates the claim: when a phrase flagged as vague is replaced, whether the less-vague replacement asserts a stronger capability than the source supports. Keep the replacement inside the source claim's scope; if a stronger reading is proposed, offer it as a team-confirm option rather than shipping it. Distinct from the plain-language drift-back and favorable-framing re-diff above (those re-diff a simplification or a positivity pass; this catches a precision fix that sharpens the claim past the source). Treat an escalated vagueness fix as Caveat and the stronger reading as Verify. Anchor: fixing less rigid about how you phrase things, the candidate handles ambiguous phrasing better was rejected — ambiguity resolution (inferring unclear intent) is a stronger, unverified capability than phrasing tolerance; the fix stayed in scope (more forgiving of how you phrase things) and the stronger reading was offered only to confirm with the team.
-
Timeline precision added in copy: when the source gives no schedule (forthcoming work, upcoming), whether copy attaches a time unit (in the coming months, later this year, soon after launch). A drafted fill for a coming xxxxx slot may only make the timeline as precise as the source; a unit the source does not state is a soft schedule claim — deliver it with a flag for owner sign-off rather than shipping silently. Treat an added time unit as Caveat. Sibling of Expiry vs. window above (that manufactures a cliff; this manufactures a schedule). Anchor: a close's in the coming months filled a user slot against a source that said only forthcoming work on multimodal models… — delivered with the flag, and the owner kept it.
-
Event-page hosts vs. announced speakers: in promo copy for an upcoming event, whether anyone named as a speaker or panelist was actually announced as such by the organizer. Event platforms' hosts row (Luma, Eventbrite, Partiful) lists page managers, not panelists — naming them as speakers is a factual error on top of the Individual Attribution gate. Treat an unannounced name presented as a speaker as Blocker; the fallback is role language (the research engineers and scientists building our open models). The promo-side companion to the recap check below; see .agent/skills/ai2-comms-style-source/references/approval-gates.md (Individual Attribution, event-page hosts).
-
Sponsor or umbrella-program credit scope: whether presented by [sponsor] (or sponsored by, powered by) sits next to the thing the sponsor actually presents. Part of [umbrella], presented by [sponsor] must not compress or dedupe to a bare presented by [sponsor] under the event or release — that re-scopes the credit to something the sponsor does not present. Treat a drifted credit as Caveat. See .agent/skills/ai2-comms-style-source/references/source-fidelity.md (Sponsor And Umbrella-Program Credit Scope).
-
Unconfirmed event or program facts: when recapping an event (hackathon, challenge, fellowship), whether winners, the existence of a contest, who presented where, and selection counts (25 teams pitched, 10 chosen) trace to the organizer or host, not inferred from participant materials or asserted unsourced. Don't invent; mark Verify, request the source, and add once confirmed. Anchor: a contest, the winning teams, and a presented at Ai2 claim were absent from all ten team reports, transcripts, and decks — confirmable only from the organizer's email. The event-logistics case of the no-fabrication rule; companion to the fabricated/unsourced-quote and unconfirmed-venue checks above.
-
Formal result-term without its numbers: whether Pareto improvement, no trade-off, strictly dominates, or matched at lower cost appears when the supporting two-axis figures are not on the page or not public. These are frontier claims, not descriptions; without the numbers, render the effect in plain words and leave the term to the source paper. Treat a borrowed frontier term with no numbers as Caveat. The mirror of Effect-verb strength above (there the copy escalates past the source; here it matches a source claim it cannot show). See .agent/skills/ai2-comms-style-source/references/claims-and-benchmarks.md (A Formal Result-Term Needs Its Numbers Or A Plain Rendering).
-
Regression-in-our-own-artifact scoping: when a partner study finds a flaw in an Ai2 dataset or training stage, whether the finding sentence's subject is a released model (reads as a verdict on the shipped product) or a generic one (the models, one set of experiments). Check the subject before reaching for a scope sentence — a bolted-on disclaimer re-raises the reading it was meant to prevent. Name the artifact actually studied; keep the affected models generic. Treat a released-model subject as Caveat, and an unscoped safety finding on our own artifact as an approval flag for the artifact's team. See .agent/skills/ai2-comms-style-source/references/claims-and-benchmarks.md (Scope By Keeping The Subject Generic, Not By Adding A Disclaimer).
-
Access-tier reframe (Preview/early-access → GA): when copy approved for a limited-access cohort (Preview, beta, waitlist) is widened to generally available, whether any residual clause still addresses only that cohort or still calls the tool experimental, and — the reverse — whether every named feature has actually reached GA before its preview scope is dropped. Strip the access-scoping (cohort descriptors, still experimental, gated CTAs), widen the audience, and convert an insider report-back CTA (tell us how it works for you) to a broad trial invitation (try it on your own research); but re-confirm each feature's availability first, since a source written for Preview may still cover a gated feature. Treat a residual preview-scoped clause after a GA reframe as Caveat, and an unconfirmed GA claim as Verify. Anchor: an Asta Preview email (Thank you for being part of Asta Preview … Two updates for Preview users, closing still experimental) reframed to GA — opening became Two updates to Asta, the still experimental close was cut, the CTA moved to try them on your own research, and both features were flagged Verify to confirm GA before the Preview scope was removed. See .agent/skills/ai2-comms-style-source/references/release-types.md (Type 8).
-
Schematic geometry presented as measurement: on a chart or diagram, whether point positions, cluster shapes, bar lengths, or curve separations were drawn for legibility rather than plotted from the source. Generated coordinates carry the caveat in the caption or the alt text (positions drawn for legibility, not plotted from the study's embeddings), and a figure whose position and proximity are illustrative should not use a form that makes them look measured — a named causal chain beats a point cloud when the point is a relationship. Treat unlabeled schematic geometry as Caveat, and a schematic figure captioned as data as Blocker. See .agent/skills/ai2-comms-chart-designer/references/diagram-legibility.md (Prefer A Named Chain Over A Point Cloud).
-
Normative verdict on an example that does not support it: whether a visual or a caption attaches a judgment (should have been preferred, undesirable, the wrong answer) to an example the source only describes. A cluster of benign-but-odd generations supports "preference training made models likelier to write these" and does NOT support "a model should have refused" — that framing belongs to the safety pairs, where a harmful request was complied with and a refusal was rejected. Two fixes, in order of preference: swap in an example that instantiates the claim, or state only what the source states (amplified / suppressed rather than should / should not). A verdict louder than its evidence gets louder still once it is a filled colour band. Treat as Caveat, and as Blocker when the visual would put a contestable ethical position in Ai2's voice. Related: Verb tier vs. evidentiary status (evidence status) and Effect-verb strength (magnitude); this covers a normative label outrunning a descriptive source.
-
Prediction rendered as outcome: on a figure or caption about a predictive method, whether every predicted effect stays worded as a prediction (may, predicted, likely) rather than a measured result. A diagram compresses hedges out first, and a predicted side effect shown in the same visual register as a measured one reads as a finding. Treat a prediction stated as an outcome as Caveat.
Do not invent missing support. Ask for source data when needed.
Treat poetic or AI-smooth language as a claim when it implies impact, safety, adoption, endorsement, or capability.
When a chart carries the comparison claim (absolute + delta + scope must be visible in the visual), also route to ai2-comms-chart-designer so the chart language, footnotes, and highlighting match the cleared claim scope. Hand revisions back to ai2-comms-writer. Voice authority is ai2-comms-style-source.