Skip to main content
تشغيل أي مهارة في Manus
بنقرة واحدة

agent-browser-shield-diagnose

النجوم٣٢
التفرعات٤
آخر تحديث٣ يونيو ٢٠٢٦ في ٠٠:٤٩

Diagnose why a specific benchmark task underperformed on guarded vs baseline — whether that's a judge pass/fail regression or a cost regression (extra tokens, steps, or duration). Use when the user asks "why did task X fail on guarded", "why is the guarded run worse than baseline on Y", "why did guarded cost more on Z", "what went wrong with run_<id>", or wants to investigate a flaky or expensive (scenario, task) cell in a benchmark report.

التثبيت

التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.

SKILL.md
readonly