| name | cli-eval |
| description | Create and run evaluation suites, watch live benchmark progress, view scorecards, compare model performance, and integrate eval runs with CI workflows from the CLI. |
Overview
Create and run evaluation suites, watch live benchmark progress, view scorecards, compare model performance, and integrate eval runs with CI workflows from the CLI.
Quick install
npm install -g AIRoute
AIRoute --version
Subcommands
eval
Example:
AIRoute eval
eval suites
Example:
AIRoute eval suites
eval list
Example:
AIRoute eval list
eval get <suiteId>
Example:
AIRoute eval get <suiteId>
eval create
Flags:
Example:
AIRoute eval create
eval run <suiteId>
Flags:
-m, --model <id>
--combo <name>
--concurrency <n>
--tag <tag>
--watch
Example:
AIRoute eval run <suiteId>
eval list
Flags:
--suite <id>
--status <s>
--since <ts>
--limit <n>
Example:
AIRoute eval list
eval get <runId>
Example:
AIRoute eval get <runId>
eval results <runId>
Flags:
Example:
AIRoute eval results <runId>
eval cancel <runId>
Flags:
Example:
AIRoute eval cancel <runId>
eval scorecard <runId>
Example:
AIRoute eval scorecard <runId>
simulate [prompt]
Flags:
--file <path>
-m, --model <id>
--combo <name>
--reasoning-effort <level>
--thinking-budget <n>
--explain
Example:
AIRoute simulate [prompt]
AIRoute — CLI Evals
Requires the AIRoute CLI. See CLI entry-point skill for install + global flags.
What are evals?
Evals are automated test suites that score LLM outputs against expected answers or rubrics. AIRoute stores suites and run results in its local database.
Eval suites
AIRoute eval suites list
AIRoute eval suites list --json
AIRoute eval suites get <suiteId>
Create a suite
AIRoute eval suites create \
--name "code-quality" \
--rubric "exact-match" \
--samples-file ./samples.jsonl
Rubric options: exact-match, contains, llm-judge, regex.
--samples-file format (one JSON object per line):
{"input": "What is 2+2?", "expected_output": "4"}
{"input": "Translate 'hello' to Spanish", "expected_output": "hola"}
Run an eval
AIRoute eval suites run <suiteId> \
--model claude-sonnet-4-6
AIRoute eval suites run <suiteId> \
--model gpt-4o \
--watch
The run is asynchronous. Use --watch for a live terminal dashboard or poll manually:
RUN_ID=$(AIRoute eval suites run <suiteId> --model claude-sonnet-4-6 --output json | jq -r '.id')
AIRoute eval get $RUN_ID
Manage runs
AIRoute eval list
AIRoute eval list --json
AIRoute eval get <runId>
AIRoute eval results <runId>
AIRoute eval scorecard <runId>
AIRoute eval cancel <runId>
Scorecard output
AIRoute eval scorecard <runId> --output json
Response fields per sample:
{
"id": "sample-1",
"score": 0.95,
"passed": true,
"input": "What is 2+2?",
"output": "4",
"expected": "4"
}
Comparing models
Run the same suite against multiple models and compare:
for MODEL in claude-sonnet-4-6 gpt-4o gemini-2.0-flash; do
AIRoute eval suites run $SUITE_ID --model $MODEL --output json | jq '{model: .model, score: .score}'
done
CI integration
SCORE=$(AIRoute eval suites run $SUITE_ID --model claude-sonnet-4-6 --output json | jq -r '.score')
python3 -c "import sys; score=float('$SCORE'); sys.exit(0 if score >= 0.90 else 1)"
Errors
suites create fails with invalid rubric → use one of: exact-match, contains, llm-judge, regex
suites run returns model not found → verify model ID with AIRoute models --search <name>
eval get shows status: failed → check AIRoute logs --search eval for error details
scorecard returns empty results → the run may still be running; poll AIRoute eval get <runId> until status is completed