| name | map-proxy-brainlift |
| description | The knowledge base a REVIEWER AGENT (e.g. an academics-team evaluating agent) loads to understand and evaluate the Incept MAP Proxy — an adaptive G2-G5 reading CAT designed to place students on the same ability continuum as NWEA MAP Growth Reading 2-5. Covers the instrument's defining feature of mastery, its BrainLift (the seven load-bearing Spiky POVs and their evidence), what is actually built and live-verified (the authoritative v4 full-bank simulation suite, 2026-07-03, on the full 786-item bank), the case for approval (a Stage-2 feasibility pilot), the audit history (two adversarial audits + a teammate review, and every finding's disposition), seven cross-source syntheses that triangulate the BrainLift + Case for Approval + Audits together, the honest tradeoff logic (locked decisions vs. genuinely open forks vs. deliberately dropped approaches), and the full limitations/honest-ceiling disclosure. Use when a reviewer agent is asked to "review/evaluate the MAP proxy", "decide whether to approve the MAP proxy pilot", "judge the MAP proxy validity case", or needs full reasoning context (not a checklist) to make an informed, fair judgment. The SKILL.md body is self-contained (every number and evidence tag needed for a normal review is embedded below); the full, unabridged source documents (BrainLift, Case for Approval, Audits & Fixes, Psychometric Design, Item Bank, Simulation Evidence, Concurrent Validity Plan, Limitations, Sources, and the master Validity Case) are bundled in references/ for deep technical dives — read them on demand, not by default. Grounded in documents dated 2026-07-03 and NWEA's MAP Growth Technical Report for 2024-2025 / 2025 Norms. |
Incept MAP Proxy (G2–G5 Reading CAT) — reviewer knowledge base
You are a reviewer agent — for example, an academics-team evaluating agent — deciding whether the Incept MAP Proxy should be approved for a school pilot. This skill gives you the full context to judge it fairly: what the instrument is, why it is built the way it is, the evidence for and against it, and the honest tradeoffs. It is explanatory and evaluative: understand the design on its own terms, then judge the build against that design and the design against NWEA ground truth.
How to judge (prime directive): distinguish a deliberate, disclosed limitation from an overclaim. Reward honesty about what pre-pilot data cannot prove; penalize any claim that a simulation is empirical, that a number "narrowly meets" a benchmark it actually misses, or that adaptive delivery makes this instrument MAP itself. Every claim below carries one of six evidence tags, used consistently throughout the source documents:
| Tag | Meaning |
|---|
| [NWEA-verified] | Directly confirmed against NWEA's own primary documents (Technical Report, Norms) |
| [Direct-pull] | A live API pull from Alpha's TimeBack/OneRoster data — a measurement, not an estimate |
| [External-review] | From an independent analysis of data our own credentials cannot reach, corroborated by a second independent source |
| [Simulation] | Produced by Monte Carlo validation against the item bank — not real student data |
| [Derived] | Calculated or inferred from primary sources or peer-reviewed literature |
| [Authoring-decision] | Our design choice, disclosed as such, not read off an external source |
| [Open] | Needs data that does not yet exist — the pilot or the concurrent-validity study |
Context that frames everything: this is a pre-pilot instrument. No student has ever taken it and the real MAP in the same window. Three adversarial reviews (two independent audits plus a second teammate review — see §5) were run against the full documentation package before it reached an approver, and every finding was fixed or disclosed rather than argued away. The single most important discipline in this package: the claims ladder is enforced in code, not prose — RIT, percentile, and CGP output are blocked in the engine itself until a concurrent-validity study exists. A document can drift; a blocked API response cannot.
This SKILL.md is a curated distillation, not the only source. Every claim below is condensed from a set of full source documents bundled alongside this skill in references/ — the exact BrainLift, Case for Approval, Audits & Fixes, Psychometric Design, Item Bank, Simulation Evidence, Concurrent Validity Plan, Limitations, Sources, and master Validity Case tabs, verbatim except for one redaction (the internal approver's name, replaced with a role description — no evidence, number, or disclosure was altered). Use this file for the reasoning and the verdict; open a specific reference file when you need an exact table, a full equation, a per-standard item count, or a direct quote to cite. See §9 for which reference answers which kind of question.
1. Core concepts — NWEA ground truth and the defining feature of mastery
The defining feature of mastery. A valid proxy session places a G2–G5 reader on the same ability continuum NWEA MAP measures — using the same measurement model (Rasch 1PL), the same scoring algorithm (MLEF with the 3.8-logit fence), a content blueprint consistent with MAP's published structure, the same item interactions, and the same testing conventions (untimed, forward-only, answer-before-advance) — precisely enough to act on and honestly enough to trust. Validity does not require the proxy to be MAP; it requires every claim the score report makes to be either verified against NWEA primary sources, measured in Incept's own data, or refused. Silent approximations are not modeled — every component is either a replication of a documented MAP design element or a disclosed deviation.
NWEA reference facts [NWEA-verified unless noted]:
| Fact | Value |
|---|
| Measurement model | Rasch 1PL — "scaling is accomplished using the Rasch model" |
| Scoring algorithm | MLEF (Han 2016), 3.8-logit fence, verbatim (TR Equations 10–11) |
| Blueprint | TR Table 3.2 shows an example blueprint of 13 Literary/13 Informational/13 Vocabulary + up to 4 field-test items — not a published requirement. The public test description separately states 40–43 questions. |
| Domains | Exactly three: Literary Text · Informational Text · Vocabulary & Word Meaning (K–2 has four — do not confuse them) |
| Item formats confirmed in Reading 2–5 specifically | MCQ, multi-select, and composite (EBSR) — confirmed in Reading figures. Hot-text and drag-and-drop are documented for MAP Growth more broadly (other subjects' figures). Gap-match appears only in NWEA secondary materials. Only 3 of 6 formats are Reading-2–5-confirmed; do not present all six as equally confirmed. |
| Published quality bar | Marginal reliability 0.96 (G2–5 Reading, Table 7.7); CSEM ≈ 3.3–3.6 RIT (Table 7.5) |
| 2025 norms (effective 2025–26) | G2 170.1→181.7 · G3 184.7→193.8 (~9 RIT growth) · G4 195.9→202.1 · G5 203.7→208.4. A 1–3 RIT downward shift vs. 2020 norms; the RIT scale itself is unchanged. Alpha's stored MAP metadata still carries normsreferencedata: "2020" — every norm-referenced number must state its norm year. |
| Untimed, forward-only | A power test — no clock, once an item is answered it is locked |
| RIT scale | RIT = a·θ + c — NWEA's proprietary, unpublished linear transform. Recoverable only via paired-score regression (same students, both tests, same window) |
Two corrections worth naming explicitly, because catching them is itself evidence of rigor: the domain blueprint was originally an authoring guess of 45/40/15 before NWEA's actual Table 3.2 was located; and the original "five item formats" list was missing EBSR/composite entirely. Both were corrected against primary sources, not defended.
2. The BrainLift — the seven load-bearing positions (DOK4 Spiky POVs)
Every other section of this skill derives from these seven positions.
SPOV1 — Measurement cadence is the #1 problem, not curriculum. The full current G3 population is n≈915, mean RIT 196.0, 46th percentile [External-review, corroborated independently]. Alpha's own OneRoster data slice is coverage-biased (mean percentile ~74 in every grade); the earlier "206.4 ≈ 72nd percentile" framing is retired — see §5, finding 1, and §6 S1. Even the platform's high-achieving measured subset grows below median at G3 (CGP 49.2, n=65), and the fall baseline hasn't moved in eight years. The binding constraint is that growth velocity is measured only 2–3 times a year, so every instructional correction is a season late. The re-baselining sharpens the case, not weakens it: a 46th-percentile population with a real low tail needs interim measurement more than a coasting 72nd-percentile one ever did.
SPOV2 — Replicate MAP's published machinery exactly; approximate nothing silently. Every measurement-design choice traces to NWEA's own Technical Report — Rasch 1PL, MLEF with the verbatim fence, a blueprint anchored to what NWEA actually publishes (not an authoring guess). Where the engine deviates (EBSR partial credit vs. the TR's dichotomous scoring), the deviation is disclosed and argued, never passed off as parity. Each element was re-verified against the actual TR PDF in July 2026, including by an adversarial fact-check agent instructed to refute the claims.
SPOV3 — The claims ladder is enforced in code, not prose. Pre-pilot output is logit ± SE and domain sub-scores; RIT, percentile, and CGP are blocked in the engine itself. A teacher cannot misread a score report into a RIT claim because the number does not exist in the API response. The two audits found number-drift in documents (0.930 vs. 0.928 reliability) — exactly why the load-bearing gate lives in code, where drift can't reach it.
SPOV4 — Rasch 1PL is the only honest pre-pilot model. 3PL requires estimating discrimination and guessing parameters from ≥1,000 responses per item that do not exist pre-pilot; using 3PL now would mean inventing parameters and calling them measurement. Rasch requires only difficulty, which can be provisionally assigned and empirically corrected later — and it is MAP's own stated model. The robustness simulation shows the engine tolerates provisional-parameter error to σ=1.0 logits without reliability collapsing below 0.85.
SPOV5 — Adaptive or nothing for a cross-grade instrument. A fixed 40-item form cannot measure RIT ~161–227 (G2 fall through G5 spring) with SEM ≤ 4 — items informative for a struggling G2 reader are noise for an advanced G5 reader. This is arithmetic (Fisher information concentrates near |θ−b|≈0), not preference. The original fixed-form G3 design was superseded for this reason, on the record. MAP itself is adaptive for the same reason.
SPOV6 — Format familiarity is measurement hygiene, not test prep. Alpha students currently practice on 100% MCQ material; their first multi-select or two-part evidence item would otherwise be on the scored official MAP. Interface novelty is construct-irrelevant variance (AERA/APA/NCME Standard 1.13). This is why the UI is deliberately "MAP-like" (familiar conventions) rather than a claimed replica — see §5, finding 6, and §6 S3 for the tension that framing change leaves unresolved.
SPOV7 — Alpha is the linking lab; anywhere else it's a partnership hunt. Concurrent validity requires the same students taking both tests within two weeks. Alpha already administers official MAP network-wide (386 assessment line items since 2017) [Direct-pull], and proxy results already land in the same OneRoster gradebook via Caliper. The Stage-3 linking study is therefore a scheduling decision inside an existing testing calendar, not a research-partnership hunt — the only missing piece is a student-ID crosswalk (provisioning work, not research).
Key derived insights (DOK3), condensed:
- Two products, one engine, staged. Pre-linking: an interim diagnostic (logit growth curves, domain sub-scores, engagement flags). Post-linking: a RIT-scale proxy (percentile, CGP-style growth, historical comparability). Neither stage borrows the other's claims — see §6 S7 for why this matters to what a reviewer is actually approving.
- The pilot is designed to resolve remaining precision, not guaranteed to. The v4 full-bank run closed every issue the v1 run had flagged (RMSE, bias, reliable-range floor — see §3): reliability now meets its ≥0.93 target and bias is near-eliminated. What remains — the distance to MAP's own published 0.96 operational reliability, and per-band exposure balance — traces to expert-assigned difficulty parameters. Empirical calibration accrued from regular use is the closer — this is exactly why the ask is real usage, not more simulation.
- Domain sub-scores match what the platform already stores — Alpha's MAP imports carry the same three Reading 2–5 goal areas in OneRoster metadata, so per-domain history (including the documented weak domain, Informational Text) is queryable across proxy and official MAP in one structure.
- Bank economics bind at the edges, and the edges differ by grade — not at the average. See §3 and §6 S6.
3. What is actually built (live-verified, v4 full-bank run 2026-07-03 — authoritative)
Live artifacts:
| Artifact | Status |
|---|
| Adaptive engine | Live on Cloud Run — Fisher-information CAT, domain quotas, Rasch 1PL, MLEF with 3.8-logit fence |
| Operational item bank | 650 items, 13 fixture files, provisional difficulties ≈ RIT 175–220 |
| Staged bank expansions | +92 edge items (floor to ~RIT 155, ceiling to 235) and +44 items (26 standard-gap + 18 G5 super-ceiling, RIT 231–250) — merged, and the load_bank() build dedups to 786 harness-loadable unique items, b −4.45→+4.55 ≈ RIT 155–250. This full 786-item bank is what the authoritative v4 validation ran against. Adoption into the operational bank remains a reviewable decision pending calibration — the two are tracked separately (786 harness-loadable vs. 650 operational) even though v4 has already validated the larger set. |
| Item formats | 6 (see §1 for which 3 are Reading-2–5-confirmed) |
| Quality gate | 14-dimension quality check, ≥0.85 pass threshold, before any item enters the bank (text-dependence ablation, CCSS alignment, near-neighbor distractors, Lexile-to-band matching) |
| Session structure | 14 Literary / 13 Informational / 13 Vocabulary = 40 scored items, extending to 43 for precision; SE stop 0.35 evaluated only at ≥40 items — the TimeBack-parity operational blueprint, disclosed as distinct from NWEA's example 13/13/13 |
| TEI floor | ≥12 technology-enhanced items per session (gap-match counted), enforced at the assembled-session level, not just an authoring target |
| TimeBack integration | Live in staging — Cognito SSO login, Caliper gradebook events, course registered under Alpha School |
| Score output | Logit θ ± SE only, plus 3 domain sub-scores; RIT is code-blocked pre-linking. The reported SE is b-uncertainty-adjusted (Fisher SE × 1.02 under v4, se_basis-labeled) because plain Fisher SE understates uncertainty from unconverted provisional b-values — the inflation removes itself automatically once empirical calibration lands. |
v4 full-bank simulation results (production selection path + randomesque exposure control, full 786-item bank, G2–G5 prior, ~17,900 sessions across 7 studies) — report every miss alongside every hit:
| Metric | v4 result | Benchmark | Verdict |
|---|
| Theta recovery r | 0.967 | ≥0.90 | Meets |
| RMSE | 3.20 RIT | <4.0 RIT | Meets |
| Systematic bias | −0.04 RIT (near-eliminated; reliable-range −0.04; max per-band 0.83) | |bias|<0.5 overall, ≤1.5 per band | Meets |
| SEM (core band) | 3.14 RIT | ≤4.0 RIT (MAP's published CSEM: 3.3–3.6) | Meets |
| Marginal reliability | 0.931 | ≥0.93 target; MAP published: 0.96 | Meets the target — MAP's own 0.96 remains the higher operational bar |
| Blueprint adherence | 100% (±1 vs. 14/13/13) | 100% | Meets |
| TEI floor met | 100% of sessions (mean 16.2 ≥ 12) | 100% | Meets |
| Internal cross-bank consistency (synthetic, not MAP agreement) | r=0.918, 95% LoA [−10.0, +10.0] RIT | r≥0.85 | Meets, with synthetic-bank caveats |
| Robustness to parameter error | reliability ≥0.915 through σ=1.0 logits | ≥0.85 | Meets |
| MLEF divergences | 0 / ~17,900 sessions | 0 | Meets |
| Reliable measurement range | RIT 160–230 | — | The median G2 fall reader (~170) is now inside the reliable range — previously below it (v3: RIT 181–227) |
Report the history, not just the final table — it is itself evidence of rigor. The first production-path run (v1, 2026-07-02) regressed RMSE to 4.73 and surfaced a +2.62 RIT systematic overestimate, and both were published rather than hidden. Root-cause (v2): the bias lived in the Owen running estimate's prior shrinkage, not in the shipped score; re-basing the final score on the engine's actual MLEF value removed it (bias fell to −0.46, RMSE to 3.70). v3 re-validated everything under the TimeBack-parity configuration on the then-current 650-item bank (RMSE 3.44, bias −0.38, reliable range RIT 181–227). v4 re-runs the identical suite on the full 786-item bank after the floor/ceiling/standard-gap expansions merged, and every metric improved or held: bias is now near-eliminated (−0.04 RIT), reliability crossed its 0.93 target (0.931), and the reliable floor dropped ~21 RIT (181→160), pulling the median G2 fall reader inside the reliable range for the first time. A reviewer should treat this whole progression as a strength (the team found its own regression, root-caused it, then re-validated at larger scale) — not soften any single run into "always been fine," and not stop reading at v3 when v4 is the current authoritative number. See §6 S2 for how this progression connects to the document-audit discipline in §5.
What is not yet resolved, stated plainly:
- Marginal reliability (0.931) meets the team's own ≥0.93 target but remains well below MAP's own published 0.96 operational figure — empirical calibration is the closer, not a documentation fix.
- 38 items exceed 30% exposure; 289 items are never used in the v4 run on the larger 786-item bank (v3 on 650 items: 241 never-used; v1: 322) — the count of never-used items rose because the bank grew while session count held constant; max exposure itself fell (0.909→0.709), so concentration eased even as the raw never-used count rose. Per-band exposure balancing remains the open engineering follow-up.
- The
literary/171-180 and informational/171-180 selection cells improved but are not yet fully served (a starvation issue, not a bank-size issue) — b-band-aware passage scoring is the queued fix.
- Item security is weaker than NWEA's 40,000+-item pool allows regardless of the 786 figure (mean pairwise between-session overlap 29.4%).
4. The case for approval — what's being asked, and on what basis
The ask: approve a Stage-2 feasibility pilot — ~40 students, 1–2 classes, one ~45-minute adaptive session each — and approve in principle a Stage-3 linking study where students take the proxy within two weeks of their official Fall 2026 MAP sitting, with the N≥1,000 paired cohort accruing through regular use inside Alpha's existing MAP calendar (not a separate recruitment).
The pilot is deliberately small, and the case says so. ~40 students × ~40 items ≈ 1,600 responses against a 786-item bank — nowhere near the ~200 responses per item that calibration requires. The pilot claims feasibility only (completion rate, session duration, rapid-guess rates, teacher/student experience, early item red flags) — not calibration. Calibration accrues from regular use afterward, toward a ~300-student-equivalent response volume for the core band. See §6 S4 for a gap between what this pilot's success criteria measure and what the STAAR precedent (§5, finding 16) actually failed on.
The argument, in order:
- Alpha's current population sits near the 46th percentile with a real low tail, its growth signal is invisible between MAP windows — measurably, not rhetorically (§2 SPOV1) — and the only interim standardized instrument in the platform (STAAR G3.2022: 33 items, 100% MCQ, no domain sub-scores) has 18 result records, at most one a real student [Direct-pull].
- The instrument implements MAP's own published measurement design, verified against NWEA primary sources, and its adaptive engine passes its v4 full-bank simulation suite with every miss disclosed alongside every hit (§3).
- Three adversarial reviews of this documentation package were run before it reached an approver, and every finding was fixed or disclosed (§5) — what a reviewer is evaluating has already survived the review a skeptical reader would give it.
Three design commitments made in advance (so the plan cannot be quietly re-scoped later):
- The 40-student pilot deliberately includes ≥8 below-median G2 readers and ≥8 above-median G5 readers — floor/ceiling UX and flag behavior get tested even at this size.
- RIT unlock at Stage 3 is band-aware: the calibrated core band unlocks first; floor/ceiling-flagged scores stay logit-only until tail-targeted calibration reaches SE(b̂) ≤ 0.30 there too.
- The EBSR partial-credit-vs-dichotomous scoring question is pre-committed to a dual-scoring comparison with a named decision statistic (bootstrap 95% CI on mean absolute linking-residual difference), decided before RIT unlock — not left to be resolved ad hoc.
What this case explicitly does not ask for: a RIT-equivalence claim. There is no paired proxy↔MAP data yet, and the engine refuses to emit RIT scores until there is.
5. Audit history — what was found, and what changed (evidence the self-critique already happened)
Before this package reached an approver, two independent adversarial audits (one checking every number against the engine's own artifacts, one fact-checking every external claim against NWEA primary sources) plus a second teammate review were run against it. The highest-signal findings and their dispositions — treat a reviewer's job as verifying these are still true, not re-discovering them:
| # | What was found | What is now true |
|---|
| 1 | Population baseline (206.4 RIT / 72nd percentile) was a coverage-biased slice, not the full population | Re-baselined: full population is n=915, mean RIT 196.0, 46th percentile [External-review]. The old figure was a compound of repeat-record over-weighting (+5.4 RIT), cohort staleness, and — largest — OneRoster source-coverage bias (the batch import is itself a high-achieving slice, ~74th percentile in every grade, of a population sitting near 46th) |
| 2 | Reliability was reported as "0.930 ✅" against a 0.93 benchmark the underlying data actually showed as 0.928 (below) | Now stated honestly everywhere — and, as of the authoritative v4 full-bank run, reliability is 0.931, genuinely meeting the ≥0.93 target — with MAP's own published 0.96 given as the still-higher context bar |
| 3 | The one failing metric (RMSE 4.12) was omitted from summary tables | RMSE appears in every summary table now, including the v1 regression to 4.73 — the full history is reported, not just the final number; v4's 3.20 RIT is comfortably within the <4.0 target |
| 4 | Item-bank source table summed to 520 while claiming 650 | Replaced with a programmatically-counted 13-file breakdown summing to exactly 650 (the 786-item harness-loadable figure is a separate, later dedup count from load_bank() — see §3) |
| 5 | Simulations were G3-scoped (630 items, simplified selector) but described as "G2–G5, production code" | Re-run on the true production path, root-caused (v2), re-validated under TimeBack parity on the 650-item bank (v3), and finally re-validated on the full 786-item bank (v4, authoritative) — G2–G5 prior throughout |
| 6 | A section framed the UI as a "pixel-accurate replica" built by scraping NWEA's practice site — a legal liability presented as a selling point | Rewritten as MAP-like UI: familiar conventions justified by response-process validity (AERA Standard 1.13), with a trademark disclaimer and a legal-review gate before external use |
| 7 | "$28/student," "competitors are fixed-form," "open-source engine" — each contradicted by public evidence | All removed; the approval case carries no pricing and no competitive claims |
| 8 | "All six formats confirmed in MAP Reading 2–5" overstated the source | Bounded: 3 formats confirmed Reading-specific, 3 documented more broadly — see §1 |
| 9 | Concurrent-study cohort size stated as N=100–200 in one file and N=1,000 in another | N ≥ 1,000, consistent everywhere |
| 10 | A public no-login test URL allowed anonymous farming of the item bank | Gated behind a PUBLIC_TEST_ENABLED flag (default off) |
| 11 | A single RIT-unlock gate would have hidden tail unreliability | Pre-committed to core-band-first, band-aware unlock (§4) |
| 12 | "The pilot fixes exactly that [precision gap]" overclaimed what a study can promise | Reworded: the plan is designed to resolve it; simulation supports but cannot guarantee it |
| 13 | The internal cross-bank consistency simulation read as MAP-agreement evidence | Relabeled explicitly as "internal cross-bank consistency (synthetic)" — not MAP agreement |
| 14 | Pool-to-test ratio quoted two ways (16.7× vs. 1.7×) without reconciling them | Both framings now stated together everywhere: 650/40 = 16.3× session length; 1.6× vs. Stocking's 10× recommended floor |
| 15 | Format-familiarization benefit was implied for all six formats | Bounded: strongest for the 3 Reading-2–5-confirmed formats; the other 3 are low-harm additional practice, not a claimed prep benefit |
| 16 | Distribution (being enrolled in TimeBack) was treated as equivalent to adoption | Argued, not assumed — the empty STAAR slot (18 results, ≤1 real student) is the explicit cautionary tale; pilot success/kill criteria are completion-rate based |
The discipline this demonstrates: any claim that would embarrass the team if independently checked was checked, by an agent instructed to try to break it. A reviewer should spend less time re-deriving these numbers and more time confirming the disclosed gaps (§3, §7) haven't quietly been re-softened in a newer document.
6. Cross-source syntheses (conclusions that need §2 + §3 + §4 + §5 read together)
Each of these can only be reached by triangulating the BrainLift's stated positions (§2), the live-verified build (§3), the approval case's framing (§4), and the audit history (§5) — no single section states the conclusion on its own. S1 first — it resolves the naive misreading a reviewer is most likely to make on a first pass.
S1 — The population correction is the strongest evidence for the case, not a credibility hit against it. A reviewer skimming §5 finding 1 in isolation could read "the team overstated their population by 26 percentile points" as damage to trust. Read against SPOV1 (§2) and the case's own argument (§4), the conclusion inverts: the mechanism SPOV1 argues for — interim measurement matters most for a population with a real low tail — is strengthened by the correction, because a 46th-percentile population with a fat low tail needs a between-window signal more than a coasting 72nd-percentile one ever would. The audit did not just fix a number; it decomposed why the number was wrong into three named, quantified biases (repeat-record over-weighting, cohort staleness, source-coverage bias) rather than writing it off as an arithmetic slip. Conclusion: judge this as the single best evidence that the team's self-correction discipline is real — it survived a correction that made its own headline claim look worse on the surface, and the team published the worse-looking framing anyway because it was true and it strengthened the underlying argument. (§2 SPOV1 · §4 argument point 1 · §5 finding 1. Status: External-review population figure; audit discipline Verified.)
S2 — The v1→v4 validation progression is the same audit discipline applied to the psychometrics, not just numbers improving. §5 documents two independent audits catching document-level problems (rounded reliability, omitted failing metrics, mismatched counts). §3 documents a separate, later process: the v1 production-path re-run itself regressed two headline metrics (RMSE to 4.73, bias to +2.62) — and the team published that regression, root-caused it (v2), then re-validated on progressively larger banks (v3, v4) rather than quietly retiring v1 once v4 looked good. These are structurally the same behavior — surface an inconvenient finding, name its root cause, fix it, and keep the record of the ugly intermediate state — applied to two different classes of evidence (documentation vs. simulation). Conclusion: a reviewer evaluating "does this team self-correct honestly" has two independent, differently-sourced data points that agree, which is stronger evidence than either alone. (§3 v1–v4 progression · §5 audit discipline paragraph. Status: Simulation history Verified; documentation audit Verified.)
S3 — The "MAP-like UI" correction was a legally necessary fix that leaves a validity question unresolved. SPOV6 (§2) argues UI fidelity matters because interface novelty is construct-irrelevant variance under AERA Standard 1.13 — the closer the UI matches MAP's actual rendering, the less noise it introduces. §5 finding 6 records that the original "pixel-accurate replica" framing was corrected to "MAP-like" specifically because of trademark exposure from scraping NWEA's practice site, not because of a validity finding. Both things are true, but they pull in different directions: the legal fix was necessary and correctly prioritized, yet nothing in the package quantifies whether the resulting "MAP-like" implementation still delivers enough of the fidelity SPOV6's own argument depends on — the open naming fork (§7, Q1) inherits the same unresolved tension. Conclusion: this is a legitimate open question for a reviewer to raise, not a defect to penalize — the team correctly fixed the legal exposure first and has not claimed the fidelity question is closed; a reviewer should ask for it to be, before or during the pilot. (§2 SPOV6 · §5 finding 6 · §7 Q1. Status: Legal fix Verified; fidelity-sufficiency question Open.)
S4 — The pilot's own success criteria cannot detect the failure mode its own cautionary tale describes. §5 finding 16 names the STAAR precedent explicitly: an instrument distributed inside TimeBack drew 18 result records with at most one real student — distribution did not produce adoption. §4's pilot success criteria (completion ≥90%, session duration in-envelope, rapid-guess rate <10%) are all measured within a single teacher-administered session. None of them can observe whether a student would voluntarily return for a second session absent a mandate — which is precisely the axis STAAR failed on. Conclusion: a reviewer should treat "the pilot passed its success criteria" as evidence the mechanics work, not as evidence the STAAR-style adoption failure has been ruled out — that requires either a second, non-mandated session inside the pilot design or an explicit acknowledgment that voluntary-adoption risk is untested until Stage 2b's regular-use accrual actually happens. (§4 pilot design · §5 finding 16. Status: Direct-pull STAAR precedent Verified; pilot-criteria gap is a reviewer inference, not stated in the source package.)
S5 — SPOV4 and the honest-ceiling disclosure look contradictory in isolation; they are the same discipline read at two different moments. SPOV4 (§2) argues Rasch 1PL is "the only honest pre-pilot model." §7's honest-ceiling section states "if the Rasch model is misspecified for these items, every simulated figure is optimistic." A reviewer reading only one of these could conclude the team is either overconfident (SPOV4 alone) or undermining its own model choice (the ceiling statement alone). Read together, they are not in tension: SPOV4 is a claim about the correct decision given the available evidence (no response data exists to fit a richer model, so Rasch is not a shortcut but the only non-fabricated choice), while the ceiling statement is a claim about residual uncertainty that decision cannot eliminate (a correct decision under uncertainty is not the same as a guaranteed-correct outcome). This is the same logic as SPOV3's claims-ladder-in-code: disclose the model, disclose its risk, gate the claims that would require the risk to have resolved in your favor. Conclusion: a reviewer should read "honest choice" and "may still be wrong" as compatible, and should be more concerned if the package asserted only one of the two. (§2 SPOV4 · §7 honest-ceiling paragraph · SPOV3 for the pattern this repeats. Status: both statements Verified as stated; the compatibility is a reviewer synthesis, not spelled out in the source package.)
S6 — Bank investment closed the floor more than the ceiling, and the case's own tail table shows why that's a reasonable but incomplete response. §4's per-grade tail table names G2 and G3 as floor-bound and G4/G5 as ceiling-bound, with G5 "binding badly" (~20% of the current population above even RIT 235). §3 shows the actual bank response: the reliable floor moved 21 RIT (181→160) between v3 and v4, while the ceiling-specific fix was 18 super-ceiling items (PR #35) layered on a smaller shared floor/ceiling expansion (PR #30). The floor got the larger, population-driven investment (consistent with SPOV1's population re-baseline making the floor more urgent); the ceiling — flagged as the more severe of the two problems in the same table — got the smaller one. Conclusion: this is a defensible sequencing decision (fix the newly-quantified, more numerous floor problem first) rather than a hidden gap, but a reviewer should confirm the pilot's ≥8-above-median-G5-readers commitment (§4) is understood as testing UX at the ceiling, not as evidence the ceiling bank depth is now adequate — those are different claims. (§2 Insight 4 · §3 bank artifacts · §4 tail table. Status: Computed + External-review population data; bank composition Direct-pull from engine fixtures.)
S7 — Approving the pilot for Insight 1's Stage-1 value and approving it hoping for Stage-3's RIT-proxy value are two different bets, and the case is explicit about only being able to deliver one of them now. Insight 1 (§2) describes two products on one engine: an interim diagnostic available now, and a RIT-scale proxy available only after linking. SPOV7 (§2) and §4's ask both frame Alpha as uniquely positioned to run the Stage-3 linking study — but that study is entirely Open (§1 evidence tags): no paired proxy↔MAP data exists yet, and none of §3's v4 evidence bears on whether the eventual linking will succeed. A reviewer who approves the pilot because the interim-diagnostic value proposition is well-evidenced (domain sub-scores matching platform structure, v4's algorithm validity) is making a different, better-supported bet than one who approves it in anticipation of an eventual MAP-equivalence claim, which no evidence in this package — however extensive — can currently support. Conclusion: name which bet is being approved. The instrument's Stage-1 case is strong on the evidence in §3–§5; its Stage-3 case is a well-designed plan (§4, §8 Q4), not yet evidence. (§2 Insight 1 + SPOV7 · §3 v4 results · §4 the ask. Status: Stage-1 evidence Simulation/Direct-pull; Stage-3 value Open by design.)
7. Honest ceiling and limitations (the load-bearing disclosure)
The single most defensible remaining critique, stated by the team itself: every psychometric result in this package is either design parity (verified against NWEA documents) or simulation. None of it is yet empirical. The b-parameters are expert-assigned; the reliability and RMSE figures assume the Rasch model is correctly specified; the internal cross-bank consistency check used a synthetic bank, not real NWEA items. If the Rasch model is misspecified for these items, every simulated figure is optimistic. This is the gap between a demonstrably-right design and a proven instrument, and no amount of additional documentation closes it — only the pilot and the concurrent-validity study can. See §6 S5 for why this statement and SPOV4 are compatible, not contradictory.
Limitations, condensed (full detail in the source Limitations document):
| # | Limitation | Current state |
|---|
| 1 | b-values are expert-assigned, not calibrated; SE(b̂) effectively infinite pre-pilot | Score reports emit a b-uncertainty-adjusted SE (×1.02 under v4, down from ×1.08 under v3) that self-removes post-calibration |
| 2 | Residual bias + distance to MAP's operational reliability | Bias is now near-eliminated (−0.04 RIT) on the full 786-item bank — this was the main v3 residual and it closed with the bank expansion, not a scorer fix. What remains: reliability (0.931) meets Incept's own ≥0.93 target but stays well below MAP's published 0.96 — both traced to expert-assigned b-values under a wide prior; empirical calibration is the closer |
| 3 | RIT output blocked — no scale linking yet | The single most significant commercial limitation; deliberate, per SPOV3 |
| 4 | No concurrent validity data exists | The core cross-bank "consistency" number (r=0.918, v4) is a synthetic-bank check, explicitly not MAP-agreement evidence |
| 5 | G2 low-tail measurement | Largely closed on the full bank — the v4 reliable range now reaches RIT 160 (from 181 at v3), so the median G2 fall reader (~170) sits inside it; only the bottom ~5% (below RIT 160) still route to the official K–2 MAP instrument, an explicit and now much smaller boundary |
| 6 | All simulations assume Rasch is correctly specified | Real responses may show discrimination variability, guessing floors, multidimensionality, or fatigue effects — none modeled; post-field-trial item-fit analysis is the check |
| 7 | 650 operational (786 harness-loadable) items vs. NWEA's 40,000+ | Weaker item security (29.4% mean pairwise session overlap); 289/786 items never used in the v4 run (rose from 241/650 at v3 as the bank grew, though max exposure fell 0.909→0.709) — per-band exposure balancing is the open engineering follow-up, disclosed as a scope limit, not a validity defect |
| 8 | Original simulation scope (G3-only, closed-form selector) understated production-path effects | Complete — v4 runs the real production selection code end-to-end on the full bank |
Known vs. unknown, the honest summary: the algorithm ranks students correctly (known, r=0.967) and meets its own reliability target at the core range (known, 0.931) — but whether the Rasch model is correctly specified is assumed, whether score X on the proxy equals score X on MAP is unknown until the concurrent study, and whether the test is free of demographic bias (DIF) is unknown until N=5,000 with a diverse sample. A reviewer should treat "known" and "assumed"/"unknown" as different categories throughout this package, not synonyms for "the numbers look fine."
8. Tradeoff logic — locked decisions, genuinely open forks, and deliberately dropped approaches
Locked (the floor every version of the plan agrees on) — do not relitigate these as if they were undecided:
| Decision | Why locked |
|---|
| Adaptive CAT, not fixed form | A 40-item fixed form cannot span RIT ~161–227 with SEM ≤ 4 (SPOV5) |
| Rasch 1PL | MAP's stated model; the only model honest with zero response data (SPOV4) |
| MLEF scoring, 3.8-logit fence; EAP cold-start for first 2 items only | Verbatim NWEA-verified design |
| 14/13/13 blueprint = 40 scored, extending to 43; SE stop 0.35 at ≥40 items; TEI floor ≥12 | TR Table 3.2 is an example 13/13/13, not a requirement; TimeBack's operational split is disclosed as an authoring decision, consistent with NWEA's published 40–43 structure |
| Six formats, with per-format confirmation status disclosed | Disclosure beats overclaim (SPOV2) |
| EBSR partial credit (0/1/2) | A disclosed deviation from the TR's dichotomous rule, preserving partial-response information |
| Untimed, forward-only, answer-before-advance | Part of the measured construct, not just a UX choice |
| RIT/percentile/CGP output blocked in code until linking passes (r≥0.80, mean diff <5 RIT, held out) | The claims ladder enforced where drift can't reach it (SPOV3) |
| 2025 norms as the reference | Current NWEA reference; norm year must be stated on every norm-referenced number |
| MAP-like UI, no replica claims, trademark disclaimer, legal review before external use | Response-process validity without the IP exposure the audit flagged — see §6 S3 for the open fidelity question this leaves |
Genuinely open forks — real judgment calls, not blockers, and not yet decided:
- Q1 — Product name. "MAP Proxy" uses HMH's registered mark. Direction is set (rename before anything external; internal pilot may proceed under the working name pending counsel) but the new name itself is open. A reviewer evaluating this skill itself should note it currently uses the working name "Incept MAP Proxy" throughout, per that same pending-counsel allowance — not as a final naming decision.
- Q2 — EBSR scoring alignment. Partial credit vs. MAP's dichotomous rule — the procedure to decide is pre-committed (dual-scoring comparator, named decision statistic, decided before RIT unlock), but the outcome is the data's call, not yet known.
- Q3 — G2 form target. Alpha's G2 records split across MAP K–2 (4 goal areas) and Reading 2–5 (3 areas); the proxy is built to 2–5 only. Current lean: stay 2–5-only, route below-floor G2 students to the official K–2 instrument. Medium confidence — genuinely open.
- Q4 — Linking method. Linear equating (lower barrier, direction set) vs. concurrent calibration with NWEA anchor items (superior, needs a data-sharing agreement — an org call, not a design one).
- Q5 — Post-pilot exposure control. Randomesque selection ships now (max exposure down 0.909→0.709 from v3 to v4, though the raw never-used count rose 241→289 as the bank grew to 786 items); whether to add Sympson-Hetter item-level caps and/or a per-standard session cap stays open until real pilot exposure data exists.
Deliberately dropped — re-proposing any of these is a regression, not a fresh idea:
| Dropped approach | Why |
|---|
| Fixed-form 40-item G3 exam (the original spec) | Cannot span G2–G5 with SEM ≤ 4 in 39 items |
| 3PL IRT model | Would invent, not measure, discrimination/guessing parameters with zero response data |
| EAP scoring throughout | Prior-dependent shrinkage biases extreme scores; kept only as a 2-item cold-start |
| 45/40/15 domain weights | Superseded by NWEA's actual Table 3.2 once located — the team's own earlier assumption, corrected |
| "Five formats" constraint | Missing EBSR/composite entirely; superseded by the TR's actual format documentation |
| Approximate RIT output (200+10θ) pre-linking | The transform's coefficients are exactly what the linking study is for — emitting an approximation would be the overclaim the design exists to prevent |
| "Pixel-accurate MAP replica" UI framing | A written admission of scraping a registered trademark's assets |
| Pricing / market-size / competitor claims in approval documents | Two figures were audit-contradicted ($28/student; "competitors are fixed-form"); the approval case is now facts-only by decision |
| Timed sections | MAP is an untimed power test; a clock changes the construct being measured |
| Public no-login test URL, unshipped | Farmable by anonymous sessions; gated behind a flag before pilot |
Name the seam a reviewer should resolve: whether this package is honestly framed as "design parity + simulation-validated algorithm, real-world equivalence explicitly Open" — or whether any newer document has let a flagged number (the distance to MAP's 0.96 operational reliability, the per-band exposure imbalance, the synthetic-bank LoA) or an unearned phrase ("proven," "MAP-equivalent," a specific RIT projection) slip back in unqualified. Given the audit history in §5 and the fact that the v4 re-run resolved essentially every prior residual, the correct default assumption is that the current package is honest — the reviewer's job is to spot-check that it has stayed that way and that no document is still citing pre-v4 numbers as current, not to assume nothing has changed.
9. How to use this + sources
- Load §1 (NWEA ground truth) and §2 (the seven SPOVs) — these are the fixed reference points everything else derives from.
- Check §3 against whatever validation report you are handed — confirm it cites v4 (2026-07-03, full 786-item bank) numbers, not superseded v1/v2/v3 figures, and that the disclosed gaps (distance to MAP's 0.96 reliability, exposure imbalance) are still present and not quietly softened.
- Weigh §4 (the ask) against §7 (the honest ceiling) — does the request match what the evidence actually supports (a feasibility pilot, not a RIT claim)?
- Use §5 to calibrate trust: this package has already been adversarially reviewed three times before reaching you. Spend your effort verifying the disclosed gaps are still disclosed, not re-discovering already-fixed problems.
- Read §6 — the seven syntheses are the parts of the case that don't show up from reading any one section alone; S1 first, since it corrects the most likely misreading.
- Run every design choice through §8; a locked decision is not up for debate, a genuinely open fork should be surfaced as such, and a dropped approach reappearing is itself a finding.
- Decide on the pilot ask, and name the one load-bearing seam (§8, last paragraph) plus which of the two bets in §6 S7 you are actually making.
- If you need more than this summary provides — an exact equation, a full per-standard item-count table, the precise wording of a claim to quote, or a finding this file compressed away — open the matching file in
references/ (below). Don't re-derive numbers from memory when the primary document is one Read call away.
Bundled references (verbatim source documents, dated 2026-07-03, one redaction noted in the callout above — read on demand, not by default):
| File | Read this when you need... |
|---|
references/00_BRAINLIFT.md | The full, unabridged BrainLift — every SPOV's exact wording, all DOK3 insights and DOK2 summaries, the complete locked-decisions/open-forks/dropped-approaches tables, and the Experts list |
references/00_CASE_FOR_APPROVAL.md | The full approval ask as written for the approver — exact pilot design commitments, the complete problem-statement tables (per-grade baselines), and the "will/won't claim" language verbatim |
references/00_AUDITS_AND_FIXES.md | Every one of the ~22 audit findings (not just the 16 condensed in §5) with full before/after text |
references/01_PSYCHOMETRIC_DESIGN.md | The actual Rasch/MLEF/PCM equations, engine code snippets, CAT termination-rule pseudocode, and the full constants table |
references/02_ITEM_BANK.md | Per-file bank composition, full format-distribution tables, the complete CCSS standard coverage matrix (all 4 grades), and the bank expansion priority queue |
references/03_SIMULATION_EVIDENCE.md | Full methodology for all 7 simulations, per-band tables, and the v1→v2→v3→v4 progression in full detail |
references/04_CONCURRENT_VALIDITY_PLAN.md | The complete Stage 2/3/4 rollout design, sample-size formulas, counterbalancing design, and the pre-committed EBSR decision procedure in full |
references/05_LIMITATIONS.md | All 8 limitations with full statement/impact/mitigation text, and the complete known-vs-unknown table |
references/06_SOURCES.md | The full APA bibliography with per-source "used for" annotations and verification status |
references/MAP_PROXY_VALIDITY_CASE.md | The master investor/board-facing narrative — executive summary, competitive differentiation table, and the full UI/response-process-validity argument |
Sources (full APA list in references/06_SOURCES.md):
- NWEA (2026). MAP Growth Technical Report for 2024–2025. NWEA/HMH.
- NWEA (2025). 2025 MAP Growth Norms.
- Han, K. T. (2016). Maximum likelihood score estimation method with fencing for short-length tests and computerized adaptive tests. Applied Psychological Measurement, 40(4), 289–301.
- Rasch, G. (1960). Probabilistic Models for Some Intelligence and Attainment Tests.
- Kingsbury, G. G., & Zara, A. R. (1989). Procedures for selecting items for computerized adaptive tests. Applied Measurement in Education, 2(4), 359–375.
- Wright, B. D., & Stone, M. H. (1979). Best Test Design. MESA Press.
- Linacre, J. M. (1994). Sample size and item calibration stability. Rasch Measurement Transactions, 7(4), 328.
- Stocking, M. L. (1994). Three practical issues for modern adaptive testing item pools. ETS RR-94-5.
- Kolen, M. J., & Brennan, R. L. (2014). Test Equating, Scaling, and Linking (3rd ed.). Springer.
- AERA, APA, & NCME (2014). Standards for Educational and Psychological Testing.
- Wise, S., & DeMars, C. (2005). Test-taker disengagement / rapid-guessing.
- Black, P., & Wiliam, D. (1998); Kingston, N., & Nash, B. (2011). Formative/interim assessment effects.
- v4 full-bank validation report (2026-07-03) — the authoritative simulation-validation source for every number in §3; not bundled verbatim (it lives in the engine repository), but its results are fully captured in
references/03_SIMULATION_EVIDENCE.md.
- MAP® is a registered trademark of HMH Education Company / NWEA; this instrument is independent and not affiliated with or endorsed by NWEA.