Skip to main content
Run any Skill in Manus
with one click

llm-regression-runner

Stars20
Forks2
UpdatedApril 23, 2026 at 14:31

Use this skill when a developer wants to test a prompt change against a golden dataset and see what broke. Triggers on: "run my evals", "test this prompt change", "check for regressions", "did I break anything", "run regression tests", "test against golden dataset", "compare prompt versions", "is it safe to deploy", "run offline evals", "what changed after my prompt update", "eval before deploying". Runs a golden dataset against the current prompt, scores each case with available judges, compares results against a saved baseline, and produces a pass/fail report with a clear deploy recommendation.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

SKILL.md
readonly