Skip to main content
تشغيل أي مهارة في Manus
بنقرة واحدة

science-evals

النجوم٠
التفرعات٠
آخر تحديث٤ يوليو ٢٠٢٦ في ٠٧:٥٨

Prepare, record, grade, validate, and compare reproducible scientific-agent benchmark runs across Codex, Claude, or other systems. Use when measuring evidence traceability, uncertainty, protocol discipline, safety gates, reproducibility, loop decisions, regression behavior, or claims of scientific-agent parity on the same tasks and resource constraints.

التثبيت

التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.

مستكشف الملفات
6 ملفات
SKILL.md
readonly