Skip to main content
تشغيل أي مهارة في Manus
بنقرة واحدة

cje

النجوم٤٣
التفرعات٤
آخر تحديث٧ يوليو ٢٠٢٦ في ١٧:٤٢

Runs CJE (Causal Judge Evaluation, pip install cje-eval) to compare LLM models, prompts, or policies from LLM-judge scores, producing calibrated estimates with honest confidence intervals and refusing claims the data cannot support. Use when the user wants to compare models/prompts/policies using judge scores or eval-harness output, put a confidence interval on an LLM eval metric, calibrate an LLM judge against ground-truth (oracle/human) labels, check whether an existing judge calibration still holds on new data, or decide how many human labels an eval needs. Raw judge-score averages are miscalibrated — never compare policies by averaging them; use this skill instead.

التثبيت

التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.

مستكشف الملفات
2 ملفات
SKILL.md
readonly