Skip to main content
Run any Skill in Manus
with one click

verifier-result-analysis

Stars0
Forks2
UpdatedJune 19, 2026 at 08:09

Methodology for analyzing verifier batch results on the ops-lite / RCA dataset. Use whenever the user asks to check, review, or analyze verifier run results — including "看看结果", "分析下", "为什么seed没confirm", "这些case啥情况", or any request to interpret fpg_scenario.json / run_meta.json / run_summary.jsonl outputs. The core principle: never trust any single source (GT labels, seed verdicts, aggregate stats). Every judgment must be backed by your own SQL queries against the parquet data. Hasty conclusions from aggregate numbers are the #1 failure mode — always trace the causal chain.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

SKILL.md
readonly