Skip to main content

harbor-f2p-p2p-deep-dive

Deep dive on Harbor trial results for tasks that use SWE-Bench-style F2P (FAIL_TO_PASS) and P2P (PASS_TO_PASS) reference tests. Diagnoses why an agent failed and audits whether a failing task is genuinely hard or unfair (instruction-vs-verifier mismatch). Use when the user asks to analyze, debug, investigate, or deep dive on a Harbor run, trial, or job directory; when the user says "harbor deep dive", "f2p deep dive", or "p2p deep dive"; when a trial scores lower than expected; or when the user wants to assess task fairness, instruction quality, verifier correctness, or whether an F2P or P2P test is reasonable for the instruction given. Works with any Harbor task whose verifier reports F2P/P2P-style results, and with any supported agent (claude-code, opencode, codex). Produces a per-failure root-cause verdict (capability gap / verifier issue / instruction issue) by triangulating verifier output, the agent's implementation, the reference tests, and the instruction.

Ir a la instalación

Datos de origen

Repositorio
NVIDIA-NeMo/Switchyard
Última actividad en el origen
28 de agosto de 2026 a las 16:36
Idioma detectado de SKILL.md
inglés
Estrellas
2582
Forks
226

Opciones de instalación

De forma predeterminada está seleccionado el prompt que primero revisa el origen. Puedes cambiar a un comando directo o descargar una copia local.

Revisa los archivos de origen

Lee SKILL.md y los archivos complementarios que muestra SkillsMP antes de decidir si quieres instalarlo.