ワンクリックで
benchmark-lab
Design benchmark runs, ablations, dataset specs, and failure-analysis artifacts.
Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。
メニュー
Design benchmark runs, ablations, dataset specs, and failure-analysis artifacts.
Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。
Investigate workspace state, draft patches, run targeted validation, and package engineering artifacts.
Turn a vague request into a structured problem statement, comparison frame, recommendation, and decision-ready artifact.
Produce multi-format presentation and campaign artifacts from one task thread.
Frame analytical questions, chart directions, metrics, and reproducible data work products.
Systematically research a topic or repository using external evidence, claims validation, and synthesized outputs.
Discover relevant skill packages and supporting tools before committing to an execution path.
SOC 職業分類に基づく
| name | benchmark-lab |
| description | Design benchmark runs, ablations, dataset specs, and failure-analysis artifacts. |
| category | benchmark |
| owner | agent-harness |
| version | 1.0.0 |
| tags | benchmark,ablation,dataset,evaluation |
| tools | code_experiment_design,tool_search,external_resource_hub |
| skills | benchmark_ablation,validation_planner,compare_options |
| artifacts | benchmark_manifest,run_config,failure_report |
| runtime_requirements | workspace,web |
Use this package when the task needs reproducible benchmark plans and ablation-ready artifacts.