원클릭으로
benchmark-lab
Design benchmark runs, ablations, dataset specs, and failure-analysis artifacts.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
메뉴
Design benchmark runs, ablations, dataset specs, and failure-analysis artifacts.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
SOC 직업 분류 기준
Investigate workspace state, draft patches, run targeted validation, and package engineering artifacts.
Turn a vague request into a structured problem statement, comparison frame, recommendation, and decision-ready artifact.
Produce multi-format presentation and campaign artifacts from one task thread.
Frame analytical questions, chart directions, metrics, and reproducible data work products.
Systematically research a topic or repository using external evidence, claims validation, and synthesized outputs.
Discover relevant skill packages and supporting tools before committing to an execution path.
| name | benchmark-lab |
| description | Design benchmark runs, ablations, dataset specs, and failure-analysis artifacts. |
| category | benchmark |
| owner | agent-harness |
| version | 1.0.0 |
| tags | benchmark,ablation,dataset,evaluation |
| tools | code_experiment_design,tool_search,external_resource_hub |
| skills | benchmark_ablation,validation_planner,compare_options |
| artifacts | benchmark_manifest,run_config,failure_report |
| runtime_requirements | workspace,web |
Use this package when the task needs reproducible benchmark plans and ablation-ready artifacts.