Skip to main content

manage-evals

This skill should be used when the user asks to "trigger an eval", "run evaluation", "run swebench", "run gaia", "run benchmark", "compare eval runs", "compare evaluation results", "check eval regression", "compare benchmark results", "what changed in the eval", "diff eval runs", or mentions triggering, comparing, or reporting on SWE-bench, GAIA, or other benchmark evaluation results. Provides workflow for triggering evaluations on different benchmarks, finding and comparing runs, and reporting performance differences.

الانتقال إلى التثبيت

معلومات المصدر

المستودع
OpenHands/software-agent-sdk
آخر نشاط في المصدر
٢٩ مايو ٢٠٢٦ في ٠٤:٤٤
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
١٬٠٤٣
التفرعات
٤٦٤

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.