Skip to main content
تشغيل أي مهارة في Manus
بنقرة واحدة

evals-design

النجوم٢٥
التفرعات٠
آخر تحديث١٣ أبريل ٢٠٢٦ في ٠٧:٢٢

Design and review evaluation suites for LLM and agent systems. Use when the task involves creating, extending, or reviewing eval datasets, rubrics, judge configs, benchmark scripts, or shipping-readiness scorecards. Triggers: "eval suite", "golden dataset", "rubric", "judge model", "benchmark", "scorecard". Negative triggers: generic unit/integration testing with no LLM or agent component.

التثبيت

التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.

مستكشف الملفات
3 ملفات
SKILL.md
readonly