一键导入
benchmark-lab
Design benchmark runs, ablations, dataset specs, and failure-analysis artifacts.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Design benchmark runs, ablations, dataset specs, and failure-analysis artifacts.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
Investigate workspace state, draft patches, run targeted validation, and package engineering artifacts.
Turn a vague request into a structured problem statement, comparison frame, recommendation, and decision-ready artifact.
Produce multi-format presentation and campaign artifacts from one task thread.
Frame analytical questions, chart directions, metrics, and reproducible data work products.
Systematically research a topic or repository using external evidence, claims validation, and synthesized outputs.
Discover relevant skill packages and supporting tools before committing to an execution path.
| name | benchmark-lab |
| description | Design benchmark runs, ablations, dataset specs, and failure-analysis artifacts. |
| category | benchmark |
| owner | agent-harness |
| version | 1.0.0 |
| tags | benchmark,ablation,dataset,evaluation |
| tools | code_experiment_design,tool_search,external_resource_hub |
| skills | benchmark_ablation,validation_planner,compare_options |
| artifacts | benchmark_manifest,run_config,failure_report |
| runtime_requirements | workspace,web |
Use this package when the task needs reproducible benchmark plans and ablation-ready artifacts.