Skip to main content

harbor-f2p-p2p-deep-dive

Deep dive on Harbor trial results for tasks that use SWE-Bench-style F2P (FAIL_TO_PASS) and P2P (PASS_TO_PASS) reference tests. Diagnoses why an agent failed and audits whether a failing task is genuinely hard or unfair (instruction-vs-verifier mismatch). Use when the user asks to analyze, debug, investigate, or deep dive on a Harbor run, trial, or job directory; when the user says "harbor deep dive", "f2p deep dive", or "p2p deep dive"; when a trial scores lower than expected; or when the user wants to assess task fairness, instruction quality, verifier correctness, or whether an F2P or P2P test is reasonable for the instruction given. Works with any Harbor task whose verifier reports F2P/P2P-style results, and with any supported agent (claude-code, opencode, codex). Produces a per-failure root-cause verdict (capability gap / verifier issue / instruction issue) by triangulating verifier output, the agent's implementation, the reference tests, and the instruction.

설치로 이동

소스 정보

저장소
NVIDIA-NeMo/Switchyard
최근 소스 활동
2026년 8월 28일 16:36
감지된 SKILL.md 언어
영어
스타
2,582
포크
226

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.