Skip to main content

braintrust-size-eval-dataset

Calculate or audit eval sample sizes, minimum detectable effects, confidence interval precision, required repeated runs, and clean-trial counts for bounding rare failures. Use when a user asks how many eval cases, items, scenarios, runs, or safety trials are needed, whether an existing dataset is adequately powered, whether N examples can detect an X-point gain, or how many clean trials certify a low violation rate. Account for paired designs, clustering, target confidence, and practical effect size. Do not use for general dataset composition.

インストールへ移動

ソース情報

リポジトリ
braintrustdata/eval-library
ソースの最終更新活動
2026年8月17日 20:49
検出された SKILL.md の言語
英語
スター
12
フォーク
2

インストール方法

デフォルトでは、最初にソースを確認する Prompt が選択されています。直接コマンドに切り替えるか、ローカルコピーをダウンロードすることもできます。

ソースファイルを確認

インストールを決める前に、SKILL.md と SkillsMP に表示されている付属ファイルをお読みください。