Skip to main content

braintrust-size-eval-dataset

Calculate or audit eval sample sizes, minimum detectable effects, confidence interval precision, required repeated runs, and clean-trial counts for bounding rare failures. Use when a user asks how many eval cases, items, scenarios, runs, or safety trials are needed, whether an existing dataset is adequately powered, whether N examples can detect an X-point gain, or how many clean trials certify a low violation rate. Account for paired designs, clustering, target confidence, and practical effect size. Do not use for general dataset composition.

الانتقال إلى التثبيت

معلومات المصدر

المستودع
braintrustdata/eval-library
آخر نشاط في المصدر
١٧ أغسطس ٢٠٢٦ في ٢٠:٤٩
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
١٢
التفرعات
٢

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.