Skip to main content

trustworthy-cuj-scoring

Use when building or running an agent eval, CUJ score, benchmark, or automated-improvement loop — and especially when a whole lane fails uniformly, a score looks too good or too bad to be true, you're about to trust a new grader/oracle/judge, or a scored harness invokes pnpm / npm / tsc / any .cmd subprocess on Windows. Symptoms: "the agent can't do X" reads as a capability gap, exit-127 / exit 127, a lane scores 0% or 100%, judge and oracle scores blended into one number, write-only-at-end eval crashed and lost work, ≥2 different failure modes on one harness, hill-climbing against a contaminated signal. Keywords: oracle, grader, LLM-as-judge, position bias, Cohen kappa, sealed holdout, decontamination, SNR, MDE, noise floor, checkpoint, status.jsonl, broken oracle, reward hacking, blind relaunch.

Zur Installation springen

Quellinformationen

Repository
oimiragieo/gotcontext-saddle
Letzte Quellaktivität
15. Juli 2026 um 15:38
Erkannte Sprache von SKILL.md
Englisch
Sterne
0
Forks
0

Installationsoptionen

Standardmäßig ist der Prompt ausgewählt, der zuerst die Quelle prüft. Sie können zu einem direkten Befehl wechseln oder eine lokale Kopie herunterladen.

Quelldateien prüfen

Lesen Sie SKILL.md und alle von SkillsMP angezeigten Begleitdateien, bevor Sie sich für eine Installation entscheiden.