Skip to main content

braintrust-eval

Execute an eval end to end against a live Braintrust project — org and credential selection, finding and importing a dataset, writing the task and scorer code, running the experiment behind smoke gates and quota preflight, and tracing agentic `claude -p` runs. Use when an eval has to actually run: "run this eval," "set up an eval for X," "compare these models in Braintrust." Do not use to design a dataset, scorer, experiment, or release gate in the abstract — the lifecycle cards own what to build and why; this owns making it happen.

الانتقال إلى التثبيت

معلومات المصدر

المستودع
braintrustdata/eval-library
آخر نشاط في المصدر
١٧ أغسطس ٢٠٢٦ في ٢٠:٢٧
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
١٢
التفرعات
٢

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.