Skip to main content

braintrust-probe-capability-and-variability

Run the same dataset under controlled variants to measure either the ceiling of what a system can do or the spread of how reliably it does it. Use for questions about sensitivity to prompt paraphrases, formatting, ordering, seeds, or repeated runs; output consistency and agreement; whether a ranking survives a different prompt; flaky results; and equally for hidden, suppressed, sandbagged, or under-elicited capability, whether a low score means "cannot" or "did not," or how prompting, demonstrations, scaffolding, tools, and fine-tuning compare. Do not use to find unknown failure modes or run adversarial attacks.

الانتقال إلى التثبيت

معلومات المصدر

المستودع
braintrustdata/eval-library
آخر نشاط في المصدر
١٧ أغسطس ٢٠٢٦ في ٢٠:٤٩
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
١٢
التفرعات
٢

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.