Skip to main content
تشغيل أي مهارة في Manus
بنقرة واحدة

model-run-error-analysis

النجوم٦
التفرعات١
آخر تحديث١٩ مارس ٢٠٢٦ في ٠٤:٥٥

Analyze benchmark runs to identify dominant error modes per model, shared hard turns, grader or benchmark issues, and representative failing examples. Prefer judged runs, but fall back to transcript + benchmark-contract review when judged artifacts are missing or unreliable.

التثبيت

التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.

مستكشف الملفات
2 ملفات
SKILL.md
readonly