Skip to main content
تشغيل أي مهارة في Manus
بنقرة واحدة

aa-omniscience-eval

النجوم١
التفرعات٠
آخر تحديث١١ مايو ٢٠٢٦ في ٠٧:٠٧

Evaluates large language models' factual recall and knowledge calibration across domain-specific questions. It measures how reliably models provide correct answers versus hallucinating or abstaining when uncertain, highlighting the gap between raw accuracy and factual reliability. Use when the user wants to benchmark on AA-Omniscience, or asks about evaluating this task. Reports Omniscience Index.

التثبيت

التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.

SKILL.md
readonly