ワンクリックで
criterion-bench
Run Criterion benchmarks and detect performance regressions.
Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。
メニュー
Run Criterion benchmarks and detect performance regressions.
Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。
Property-based testing in Rust with the proptest crate. Use when writing fuzz-like tests that generate arbitrary inputs, compose strategies, derive Arbitrary, configure shrinking, or debug test failures via persistence/replay.
Generate and interpret cmakefmt trace artifacts.
SOC 職業分類に基づく
| name | criterion-bench |
| description | Run Criterion benchmarks and detect performance regressions. |
How to run Criterion 0.5 benchmarks, compare baselines, and detect performance regressions.
Criterion compares two runs: a saved baseline (before changes) and a new sample (after changes). The workflow is always two steps:
cargo bench -- --save-baseline before 2>/dev/null
cargo bench -- --baseline before 2>/dev/null
Criterion prints a summary to stdout showing whether each benchmark improved, regressed, or showed no change. Detailed statistics are written to disk.
After a comparison run, per-benchmark change statistics are written to:
target/criterion/<group>/<bench_id>/change/estimates.json
Find all of them with:
find target/criterion -path '*/change/estimates.json'
estimates.jsonEach file contains mean and median objects with identical structure:
{
"mean": {
"point_estimate": -0.0014,
"confidence_interval": {
"confidence_level": 0.95,
"lower_bound": -0.0057,
"upper_bound": 0.0040
},
"standard_error": 0.0025
},
"median": {}
}
Values are fractional, not percentages. 0.05 means 5% slower. -0.02 means 2% faster.
jqFlag any benchmark where mean.point_estimate > 0.05 (more than 5% slower):
find target/criterion -path '*/change/estimates.json' -exec sh -c '
jq -e "select(.mean.point_estimate > 0.05)" "$1" >/dev/null 2>&1 && echo "REGRESSION: $1 ($(jq -r ".mean.point_estimate * 100 | round | tostring + \"%\"" "$1"))"
' _ {} \;
A stricter check flags any benchmark whose confidence interval lower bound is above zero (statistically significant slowdown at the configured confidence level):
find target/criterion -path '*/change/estimates.json' -exec sh -c '
jq -e "select(.mean.confidence_interval.lower_bound > 0)" "$1" >/dev/null 2>&1 && echo "SIGNIFICANT REGRESSION: $1"
' _ {} \;
| Condition | Meaning | When to use |
|---|---|---|
mean.point_estimate > 0.05 | Point estimate exceeds 5% regression | General-purpose gate |
mean.confidence_interval.lower_bound > 0 | Statistically significant slowdown | Strict gate: even small regressions flagged if confident |
mean.point_estimate > 0.10 | Point estimate exceeds 10% regression | Lenient gate for noisy environments |
Choose based on environment stability. CI with dedicated hardware can use the strict gate. Local laptops with variable load should use a wider threshold.
| Flag | Default | Purpose |
|---|---|---|
--save-baseline <name> | base | Save results under a named baseline |
--baseline <name> | base | Compare against a named baseline (fails if missing) |
--baseline-lenient <name> | — | Compare against baseline, skip benchmarks that lack it |
--sample-size <N> | 100 | Number of samples per benchmark |
--warm-up-time <secs> | 3 | Warm-up duration before sampling |
--measurement-time <secs> | 5 | Measurement duration per sample |
--noise-threshold <f> | 0.01 | Changes below this fraction are considered noise |
--confidence-level <f> | 0.95 | Confidence level for intervals |
--significance-level <f> | 0.05 | Threshold for significance tests |
Increase --sample-size or --measurement-time when results are noisy. Use --baseline-lenient when benchmarks have been added or removed between runs.