Deep dive on Harbor trial results for tasks that use SWE-Bench-style F2P (FAIL_TO_PASS) and P2P (PASS_TO_PASS) reference tests. Diagnoses why an agent failed and audits whether a failing task is genuinely hard or unfair (instruction-vs-verifier mismatch). Use…
原文の言語: 英語