Skip to main content
تشغيل أي مهارة في Manus
بنقرة واحدة

llm-regression-runner

النجوم٢٠
التفرعات٢
آخر تحديث٢٣ أبريل ٢٠٢٦ في ١٤:٣١

Use this skill when a developer wants to test a prompt change against a golden dataset and see what broke. Triggers on: "run my evals", "test this prompt change", "check for regressions", "did I break anything", "run regression tests", "test against golden dataset", "compare prompt versions", "is it safe to deploy", "run offline evals", "what changed after my prompt update", "eval before deploying". Runs a golden dataset against the current prompt, scores each case with available judges, compares results against a saved baseline, and produces a pass/fail report with a clear deploy recommendation.

التثبيت

التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.

SKILL.md
readonly