Skip to main content

crud-purge-below-gate-evals

Guardrailed DELETE of auto-registered eval `sandbox_jobs` rows that DID score but FAILED the harvest gate — partial evals (valid-complete <90% or non-benign infra-error >10%). These are the "DE-REGISTER candidate" rows the flawed_summ harvest defers. DISTINCT from `crud-purge-stale-eval-placeholders` (that removes NEVER-populated Pending/Started rows; this removes rows that populated stats/metrics but are below-gate). Use when asked to remove "partial / below-gate / <90%-completed-without-errors" evals owned by us. The MANDATORY parts: the authoritative gate utility (`scripts/database/eval_guardrail.py` — NEVER hand-roll the BENIGN set / error count), the cross-user FK-safety pre-check, the grandchild→child→job cascade, and REPORTING BACK the exact purged rows (§5) so the supervisor knows which state docs (e.g. flawed_summ STATE.md) to reconcile.

Zur Installation springen

Quellinformationen

Repository
open-thoughts/OpenThoughts-Agent
Letzte Quellaktivität
30. Juli 2026 um 10:49
Erkannte Sprache von SKILL.md
Englisch
Sterne
286
Forks
39

Installationsoptionen

Standardmäßig ist der Prompt ausgewählt, der zuerst die Quelle prüft. Sie können zu einem direkten Befehl wechseln oder eine lokale Kopie herunterladen.

Quelldateien prüfen

Lesen Sie SKILL.md und alle von SkillsMP angezeigten Begleitdateien, bevor Sie sich für eine Installation entscheiden.