Skip to main content

port-a-benchmark

Port an existing third-party benchmark or eval into this repo as an Inspect AI task, faithfully. Use this whenever the task is "add benchmark X", "port this eval", "can we run <github repo> here", or adapting any external eval harness (aiewf-eval, lm-evaluation-harness, a paper's repo, a colleague's script). Enforces the one rule that makes a port worth anything — copy the measuring instrument verbatim, adapt only the runner — plus the verification steps that prove the port is faithful rather than merely running.

الانتقال إلى التثبيت

معلومات المصدر

المستودع
awslabs/llm-evaluation-system
آخر نشاط في المصدر
٤ أغسطس ٢٠٢٦ في ١٩:٢٣
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
٢٢
التفرعات
٣

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.