Use when about to claim, review, or dispute a speedup/regression in tensor-grep (tg vs rg, hot-cache, AST, agent-workflow, or GPU changes) — which benchmark script to run, how to read check_regression.py, the noise-floor/absolute-jitter rule for sub-10ms rows, the warm-vs-cold measurement trap (a warm dogfood run can hide a real cold-path win), the byte-identical output-proof obligation before trusting any merge/skip/reorder speedup, the fair-baseline rule (never compare tg against a strawman comparator), and the launcher-attribution rules (tg_launcher_mode, tg_launcher_command_kind, stale-binary refusal) that make a benchmark artifact claim-quality instead of noise.
Installer avec Codex ou Claude Copiez ce prompt, collez-le dans Codex, Claude ou un autre assistant, puis laissez-le vérifier la page du skill et l'installer pour vous.
Une commande directe contourne le prompt de vérification. Examinez la source avant de l'exécuter.
Use when about to claim, review, or dispute a speedup/regression in tensor-grep (tg vs rg, hot-cache, AST, agent-workflow, or GPU changes) — which benchmark script to run, how to read check_regression.py, the noise-floor/absolute-jitter rule for sub-10ms rows, the warm-vs-cold measurement trap (a warm dogfood run can hide a real cold-path win), the byte-identical output-proof obligation before trusting any merge/skip/reorder speedup, the fair-baseline rule (never compare tg against a strawman comparator), and the launcher-attribution rules (tg_launcher_mode, tg_launcher_command_kind, stale-binary refusal) that make a benchmark artifact claim-quality instead of noise.
tensor-grep Benchmark and Proof Toolkit
Prove it, do not eyeball it. This repo is benchmark-governed (AGENTS.md: "treat the repo as a
benchmark-governed, contract-heavy codebase; do not optimize by guesswork"). This skill is the
runbook for producing a benchmark artifact that is actually trustworthy, picking the right script
for the change you made, and reading its output correctly.
When NOT to use this skill (go to a sibling instead)
Writing or reviewing the actual code change → tensor-grep-change-control (registration sites,
backend fail-closed contract) or verify-plan-against-code.
Chasing a bug/regression's root cause, not proving a speed number → tensor-grep-debugging-playbook
or superpowers:systematic-debugging.
Deciding whether to dogfood the shipped binary after a release → dogfood-the-shipped-artifact.
General CLI usage (tg search, tg orient, flags) → tensor-grep-config-and-flags or the
tensor-grep usage skill (.claude/skills/tensor-grep/SKILL.md).
Researching a novel GPU/ML technique before building it → tensor-grep-research-frontier.
You just need a code-search/navigation tool right now, not a benchmark → tensor-grep skill.
The one rule everything else derives from
Never claim a speedup (or accept a regression as fine) without a measured number from the
current accepted baseline, run through the right script, read with absolute-tolerance
awareness, produced by a claim-safe launcher. (AGENTS.md "Benchmark Rules" / "Performance
Discipline", docs/benchmarks.md "Acceptance Rules")
Five corollaries, all stated in AGENTS.md:
Compare against the current accepted baseline, not memory.
Reject a candidate that is only "faster" in a microprofile while slower end-to-end.
Keep both cold-start and repeated-query measurements in mind — they are different regimes.
Do not update docs/PAPER.md with a speed claim until the benchmark line is accepted.
If a candidate is correct but slower, revert it and record the attempt — do not ship a clean
regression because "the code is nicer."
product-wedge workflow evidence; not a cold exact-text speed claim
GPU / NLP backend (--gpu-device-ids, CyBERT)
benchmarks/run_gpu_benchmarks.py (Python sidecar scale/correctness) or benchmarks/run_gpu_native_benchmarks.py (native CUDA crossover)
GPU is experimental — see the GPU section below before trusting any GPU number
context-render / edit-plan latency
benchmarks/run_context_render_benchmarks.py
editor-plane latency, not search speed
blast-radius latency
benchmarks/run_blast_radius_benchmarks.py
impact-analysis latency
repo-map / retrieval quality (not speed)
benchmarks/run_repo_retrieval_benchmarks.py
recall/precision/MRR/nDCG/F1/token-budget, a quality metric, not a timing one — For anything chunker- or ranking-sensitive (cAST, , RRF channels), use the sibling instead — see
Full matrix with default artifact paths: docs/benchmarks.md § "Benchmark Matrix" (still 19 scripts as
of v1.95.0, re-counted this pass — re-verify with the command in Provenance below, the list drifts).
If your change does not obviously map to one row, run benchmarks/run_benchmarks.py first (the
broadest cold-path net) and widen from there — do not invent a new ad hoc timing script.
Recipe per method
1. End-to-end cold tg vs rg (the default)
python benchmarks/run_benchmarks.py --output artifacts/bench_run_benchmarks.json
python benchmarks/check_regression.py --baseline auto --current artifacts/bench_run_benchmarks.json
--baseline auto resolves benchmarks/baselines/run_benchmarks.<platform>.json from the recorded
environment.platform (Windows → run_benchmarks.windows.json, Linux → run_benchmarks.ubuntu.json;
check_regression.resolve_auto_baseline_path, benchmarks/check_regression.py:15-36). It refuses
to run on an unsupported platform rather than silently comparing against nothing.
Useful flags on run_benchmarks.py (benchmarks/run_benchmarks.py:675-736):
--binary PATH — pin the exact tg binary under test (default:
rust_core/target/release/tg[.exe]).
--native — force tg search --cpu and add the native large-file/many-file scenarios.
--launcher-mode {auto,explicit_binary,explicit_fast_binary,discovered_cli_binary, python_module_launcher,python_module_rust_first,...} — pin the launcher shape for a
control-plane experiment (see "Fair-benchmark rules" below — never compare across modes silently).
--allow-claim-unsafe-launcher — bypass the stale-binary refusal for exploratory timing only;
never use this on a run whose numbers you intend to put in docs/PAPER.md.
Use for StringZilla index changes, CPU regex prefilter changes, persisted cache/decode/posting-list
changes. repeated_regex_native must stay on native/Rust routing (cpu_rust_regex or similar) —
if your probe forces a Python fallback you are benchmarking the wrong thing (AGENTS.md Benchmark
Rules). This script self-grades: it computes improvement_pct and a status: PASS|FAIL per row
against --max-regression-pct (default 5.0) — see the noise-floor section for why sub-10ms rows
need extra care here specifically.
The first is tg run vs ast-grep on one query (ratio gate: tg/sg <= 1.1 per docs/benchmarks.md
§ "ast-grep vs tensor-grep AST mode"). The second is run/scan/teststartup and orchestration
— use it when you touched AST workflow batching, not query matching itself.
This is workflow evidence, not a cold exact-text search speed claim (docs/benchmarks.md is
explicit about this — the artifact literally embeds the string "agent-native workflow benchmark; not a cold exact-text speed claim" as a positioning field). Use it for tg agent capsule routing,
confidence/alternative-target honesty, validation-command filtering, rollback visibility, edit-order
guidance, or whole-loop latency (search_s/plan_s/apply_s/verify_s medians). Do not use it to
argue tg beats rg — that claim needs script #1.
Treat SKIP as expected infrastructure state, not a fake failure — CyBERT may skip when Triton is
unreachable, and the whole artifact reports top-level status: "SKIP" when no operational GPU
device is detected (find the emitter with grep -n '"SKIP"' benchmarks/run_gpu_benchmarks.py --
benchmark_pattern/devices are still recorded there so the skip is diagnosable). GPU is experimental and currently not a promotion-ready path — read the
"GPU claims need a stricter bar" section below before trusting any GPU number as a win.
6. Token-economy / tokens-per-correct-answer (CANDIDATE — gated on #72, not yet a committed benchmark)
A different metric class from every row above: not wall-clock, but tokens an agent must spend to
reach a verified-correct answer (a task-level cost proxy, oracle-validated per task rather than a raw
F1/speed number). This is the metric the arxiv research landscape converged on as the actual moat
measure (tensor-grep-arxiv-research-landscape-2026-07-07: "token-economy > F1" consensus) — it is
what a fleet of agents actually pays for, not a lab F1 score.
First real run (2026-07-08, internal, oracle-validated, NOT YET a committed benchmarks/ script):
Sverklo bench:primitives (Zenodo 10.5281/zenodo.19802051) on expressjs/express@4.21.1, 25 tasks (10
definition-lookup P1, 10 references P2, 5 file-deps P4). Gated tokens-per-correct-answer (F1>=0.8):
tg P1 = 1,243 tok vs rg = 9,328 tok -> tg 7.5x BETTER (moat validated on definitions). tg P4 file-deps
= 53,631 tok vs rg = 5,367 tok -> tg ~10x WORSE (at the time) — tg had no scoped "what does file X
import / who imports X" primitive, only whole-repo tg map, so every P4 query paid the whole-repo
capsule cost regardless of the single file asked (task #74).
P4 CLOSED AND RE-PROVEN (2026-07-16, v1.76.12 #619).#460 (05f49b8) shipped the fix — scoped
tg imports FILE / tg importers FILE [ROOT] — and the SAME Sverklo P4 slice was re-run independently
(deterministic, $0, scratchpad/bench/aggregate.py): 53,631 tok -> 2,387 tok, ~10x WORSE -> ~2.24x
BETTER than rg, F1 preserved and improved (0.542 -> 0.606), bidirectional-oracle PASSED 25/25. The
moat is now proven on both P1 and P4. Raw artifacts and scripts still live at scratchpad/bench/
(results.json, run_bench.py, score.py, validate_oracle.py, aggregate.py) — still not
promoted to benchmarks/ or docs/benchmarks.md, so the harness-committal gap in the paragraph below
is unchanged even though the P4 number itself is now closed; full memory:
tensor-grep-benchmark-proofpoint-2026-07-08 (original) + tensor-grep-drain-resume-2026-07-12.md /
tensor-grep-find-campaign-2026-07-16.md (re-proof receipts).
Why this is gated, not accepted, and what #72 covers: (a) the run above is one repo / one language
(JS) / 25 tasks — a real signal, not a general claim, and the full 180-task/6-repo suite (flask/fastapi
in Python plausibly show a bigger win) has not run; (b) the harness (run_bench.py/score.py/
validate_oracle.py) has not been committed to benchmarks/, given a docs/benchmarks.md Matrix row,
or wired into check_regression.py-style acceptance; (c) publishing this number externally is
CEO-gated (public positioning), separate from whether the harness itself is accepted internally. Do
not cite the 1,243/9,328, 53,631/5,367, or the re-proof 2,387 numbers above as an accepted
claim-quality artifact until #72 lands a committed script + artifacts/bench_*.json — treat them as a
research finding pointing at real moat validation on BOTH axes now (P1 definitions + P4 file-deps), not
yet a benchmark-governed line. Follow this skill's
same discipline once #72 ships: fair-baseline (rg is already the right comparator here), noise-floor
(deterministic CLIs, no run-to-run variance in this metric class so the usual jitter rule is moot, but
oracle correctness bugs are the equivalent failure mode — bidirectionally validate the oracle before
trusting a score), and no doc/PAPER.md claim until the artifact is accepted.
Unlike §6 (an uncommitted CANDIDATE), benchmarks/eval_late_rerank_quality.py is a committed,
running ndcg@10/recall@10 quality gate for tg find / tg_find ranking changes (dense weight, RRF
channels, chunker, late-rerank) — see the decision-table row above and
tensor-grep-semantic-search-campaign STATUS UPDATE 2. Two rigor requirements before trusting a gate
run here, both learned the hard way on this harness:
Corpus-hardness gate. Before trusting a "rrf beats bm25" delta, assert BM25-alone scores
near-floor on the HARD subset of the golden set (the vocabulary-mismatch queries the dense leg is
supposed to rescue) — the same GATE 0b trap as §Phase-0 baselines elsewhere in this repo's
campaigns: an easy/keyword-discriminating corpus lets BM25 saturate at recall 1.0 and makes any
fusion delta look meaningless by comparison, or conversely hides a real fusion win inside noise.
Paired win/loss/tie report, not a bare mean. A 40-query aggregate mean can hide a distribution
where one lucky query drives the whole delta. Report per-query win/loss/tie counts (e.g. "positive
in all 4 categories, essentially wins-or-ties per query, a single ndcg loss out of 40" — the actual
tg find gate-run shape) before gating a ship decision on the mean alone. See the global skill
paired-test-power-discipline.
Bidirectionally validate any new golden query the same way as elsewhere in this skill: a correct answer
must PASS the grader and a wrong/empty answer must FAIL it, before trusting a delta computed against it.
Noise-floor / jitter discipline
Sub-10ms timings are dominated by process-spawn and OS-scheduler jitter, not by the code you changed.
Two concrete, code-verified mechanisms exist for this:
min-baseline-time-s floor in check_regression.py (default 0.1s,
benchmarks/check_regression.py:69-74; enforced in perf_guard.check_regressions and
detect_comparator_drift, src/tensor_grep/perf_guard.py:87,109): any row whose baseline
time is below this threshold is skipped entirely for regression comparison — "tiny baseline
durations are noisy on shared CI runners and can trigger false positives from scheduler jitter."
Absolute jitter tolerance in run_hot_query_benchmarks.py
(NATIVE_REGEX_ABSOLUTE_JITTER_S = 0.005, benchmarks/run_hot_query_benchmarks.py:14): the
repeated_regex_native row's PASS/FAIL uses
max(relative_tolerance_s, absolute_tolerance_s) — i.e. a percentage-only gate would flag a
2ms → 4ms wobble as a "100% regression" when it is pure noise, so an absolute 5ms floor is added
on top of the --max-regression-pct (default 5%) relative gate
(benchmarks/run_hot_query_benchmarks.py:202-226).
Rule of thumb when you add a new hot/cache benchmark row: if the expected timing is under ~10ms,
do not gate on percentage delta alone — add an absolute-seconds floor the way
NATIVE_REGEX_ABSOLUTE_JITTER_S does, or you will chase phantom regressions caused by nothing but
scheduler noise. This is also the general noise-floor-before-quantitative-claims skill's territory
if you are building a new (non-tg) measurement harness from scratch.
A related but distinct trap: a warm end-to-end run can hide a real cold-path win (or loss) —
this is systematic, not jitter. Jitter (above) is random noise around the true value; a warm/cold
regime mismatch is a wrong measurement of the wrong code path and can point the wrong direction
entirely. A warm dogfood run measures the CACHED path, where the function you actually changed may not
even execute on that request. Receipt: a tg orient warm end-to-end dogfood read showed -36% on a
symbol-merge change (_python_imports_and_symbols, src/tensor_grep/cli/repo_map.py — locate via
grep -n "def _python_imports_and_symbols" src/tensor_grep/cli/repo_map.py, was :2126, now
:2166 on 2026-08-13) that
directly microbenchmarking the function then showed was actually ~54% faster (961ms→446ms), because
the warm run never re-parsed the file the change touched. This deepens corollary 3 above ("cold-start
and repeated-query are different regimes") into a concrete verification recipe: to prove a cold-path
lever, either (a) microbench the target function directly on the published/shipped build — a fresh
process per rep (cold cache by construction), a single pass over distinct inputs, old-vs-new, asserting
output-identity (total == total both sides) — or (b) explicitly clear the relevant cache between reps
of an end-to-end run. Never trust a warm end-to-end number as evidence for or against a cold-path
change. Second receipt, same shipped-wheel-microbench discipline applied to a different lever: a
validation-scan pre-check (_framework_test_pattern_bonus, src/tensor_grep/cli/repo_map.py —
locate via grep -n "def _framework_test_pattern_bonus" src/tensor_grep/cli/repo_map.py, was
:11112, now :11446 on 2026-08-13)
measured ~68% faster (3657ms→1172ms) this way, output byte-identical. The general profiling/proof
pipeline this recipe belongs to (profile the shipped wheel → prove byte-identical output → warm/cold
microbench) lives in the global skill profile-guided-byte-identical-optimization; this is the
benchmark-reading slice of it.
Summary statistics, pairing, and intervals (the part that makes a number publishable)
Added 2026-08-22 after a 5-seat council voted 5/5 to WITHDRAW the public 7.5x claim. The two
internal numbers (7.5x, later 6.4x) conflict, no committed harness produces either, and — the
insight none of the council seats reached — they may be different REGIMES against different
baselines, in which case neither supersedes the other and BOTH are unpublishable regardless of a
re-run. Everything below is externally grounded; the sources are named so a reader can check them
rather than take this file's word.
Ratios MUST be aggregated with the geometric mean, never the arithmetic mean
Fleming & Wallace, How not to lie with statistics: the correct way to summarize benchmark results
(CACM 1986, doi:10.1145/5666.5673) proves the arithmetic mean of NORMALIZED numbers is
mathematically meaningless: the ranking it produces changes when you change which system is the
baseline. A speedup is a normalized number. So:
Summarizing a suite of per-case speedups -> geometric mean.
Summarizing raw times for ONE case -> arithmetic mean / median / min (see next section).
If you ever report "average speedup" without saying which mean, assume it is the wrong one.
min vs mean vs median — pick from the DISTRIBUTION, and say which you used
CLI/process timings are right-skewed: there is a hard floor (the work must take at least X) and no
ceiling (any scheduler, page-cache miss, or antivirus scan adds time). Noise is therefore strictly
ADDITIVE, which is the argument for the minimum as the least-contaminated estimate of intrinsic
speed (Codeflash, docs.codeflash.ai/codeflash-concepts/benchmarking; kevmod,
blog.kevmod.com/2016/06/10/benchmarking-minimum-vs-average, gives the variance analysis for when
min beats mean and when it does not).
statistic
use when
fails when
min
right-skewed timings, you want intrinsic speed, noise is additive
you actually care about tail latency; rare-bad-event cost is the product question
median
you want a robust central value and the tail matters a little
you need to detect a small shift and n is small
arithmetic mean
raw times, roughly symmetric, you will also report the spread
ANY normalized/ratio quantity (see above)
Whichever you choose, the report must NAME it. "tg is 3x faster" with an unnamed statistic is not a
claim, it is a mood.
Paired, interleaved sampling — never all-A-then-all-B
Run arms as adjacent pairs and alternate which goes first; do not collect every tg sample and then
every rg sample. Block sampling lets machine state (thermals, clock, page cache, another process
waking up) load onto one arm. The protocol in github.com/clocksmith/doppler/blob/main/docs/ benchmark-methodology.md is directly analogous to a tg-vs-rg comparison: interleave adjacent pairs,
collect >= 20 valid pairs, keep every other condition fixed, and compute the paired difference
per pair.
Confidence intervals are mandatory for anything externally used
Hoefler & Belli, Scientific Benchmarking of Parallel Computing Systems
(spcl.inf.ethz.ch/Publications/.pdf/hoefler-scientific-benchmarking.pdf) is the standard
methodology reference; NIST TN 1830 (Pieterse & Flater) states the operational consequence plainly:
comparing bare averages "often leads to incorrect conclusions", and it is the interval that
separates a real difference from random fluctuation.
Report a paired 95% CI on the difference, not two independent CIs eyeballed for overlap.
Stopping rule: an interval crossing zero with a median difference under ~0.5% is PARITY —
stop tuning that lane, and do not publish the point estimate as a win.
Round to the precision you earned. A +-6-point interval justifies whole numbers. If a
presentation needs two decimal places to show a difference, there is no difference.
Noise floor, and the CI-specific one
Establish the floor with a NO-OP control (same input, same arm, back to back) before believing any
delta. Concrete published floors: 5% on a real machine, 10% on GitHub Actions (Codeflash). Our
benchmark lane runs on GitHub Actions, so a sub-10% effect measured in CI is not a result.
Never mix regimes
fak/BENCHMARK-GOVERNANCE.md states the rule that explains our own 7.5x-vs-6.4x conflict: a live
wall-clock speedup and a "value-add"/session ratio have DIFFERENT baselines and are not comparable.
Before re-measuring anything, the first question is not "which number is right" but "what baseline
and regime did each one use" — and if that cannot be recovered from the artifacts, both are
unpublishable no matter how carefully you re-run.
Corollary from the SIGPLAN Empirical Evaluation Guidelines (via pldi-reproducibility):
benchmark survivorship — silently excluding the cases your tool loses on is "the most damaging
silent choice" available. Report the losses in the same table as the wins.
Tombstoning a superseded public number (never silently delete)
When a published number is replaced or withdrawn, it gets a SUPERSEDED marker beside it with the
date, the replacement (or "withdrawn, no replacement"), and the reason. Deleting it makes the record
unauditable and invites someone to re-derive the old figure from an old artifact. This is the
append-only discipline tensor-grep-release-drift-check already applies to dated claims.
Pre-flight additions to the speed-claim checklist
Before ANY externally-used number:
Ratio aggregation method NAMED and is geometric mean for speedups.
Summary statistic (min/median/mean) NAMED, and justified by the distribution.
Arms interleaved and paired; >= 20 pairs; paired 95% CI reported.
Noise floor measured with a no-op control; effect exceeds it (>= 10% if measured in CI).
Regime + baseline stated in the SAME sentence as the number.
Losses reported alongside wins (no survivorship).
Reproduction command committed; raw pairs committed.
Any superseded number TOMBSTONED, not deleted.
A number missing any of these is "shape interesting, not publishable" — say exactly that rather than
shipping it with a hedge.
Fair-benchmark rules
These are what separates a claim-quality artifact from a number you cannot defend in a PR review.
Launcher/command-kind attribution (never blend routes into one number)
run_benchmarks.py records, per artifact, both:
environment.tg_launcher_mode — which of the 9 launcher-mode experiments produced the command
(auto, explicit_binary, explicit_fast_binary, explicit_binary_positional,
explicit_binary_positional_early_rg, explicit_binary_early_rg, discovered_cli_binary,
python_module_launcher, python_module_rust_first — run_benchmarks.py:380-390).
environment.tg_launcher_command_kind — what the concrete resolved command actually is:
native_exe, cmd_shim, powershell_shim, uv, python_module, or unknown
(classify_tg_launcher_command, run_benchmarks.py:162-178).
Why both: tg_launcher_mode is the experiment you asked for; tg_launcher_command_kind is what
you actually got. A .cmd shim, uv wrapper, or Python-module route on the discovered/default path
adds wrapper/interpreter overhead that has nothing to do with your code change. If
command_kind != native_exe, the script prints a top-level warning
(benchmark_launcher_warnings, run_benchmarks.py:181-191) and the artifact carries it in warnings.
Never compare two artifacts with different tg_launcher_command_kind values and call the delta a
code-level win or loss — the delta may just be shim overhead.
Refuse stale in-tree binaries by default
run_benchmarks.py, run_native_cpu_benchmarks.py, and run_cold_path_attribution.py all call
inspect_native_tg_binary and check version_status. If the resolved binary is in-tree-* (built
from rust_core/target/{debug,release}) and its version does not match the expected package
version, the script prints [blocker] lines and exits 2 unless you pass
--allow-claim-unsafe-launcher (run_benchmarks.py:212-226,768-771). This exists because a stale
in-tree binary silently benchmarks last week's code while you think you're measuring today's
change. Only pass --allow-claim-unsafe-launcher for exploratory timing you will not cite in a PR
or doc — the flag name says so on purpose.
No comparing across launcher kinds, ever, without saying so
docs/benchmarks.md § "Artifact Conventions": "This prevents native-exe, .cmd shim, uv, or
Python-module overhead from being combined into one search-speed claim." If you must report two
launcher modes side by side (as the roadmap-1 control-plane probes in docs/benchmarks.md do), label
each row with its tg_launcher_mode/tg_launcher_command_kind explicitly and state which one is the
control-plane experiment vs. the accepted baseline — do not average them.
Large-file benchmarks route through TWO different engines — disclose which one (A3, v1.91.3/#695)
NativeCpuBackend is not one code path for benchmarking purposes. rust_core/src/native_search.rs
is the default streaming route, deliberately kept SERIAL and held to a tested ≥25ms first-match
latency contract. rust_core/src/backend_cpu.rs is a separate PyO3/FFI fallback route, reached
only when the search doesn't go through the primary native front door — this is where #695 shipped
intra-file rayon parallel search, gated to files ≥50MiB, byte-identical to the serial result.
A large-file (e.g. the 200MB row referenced in the worked example above) benchmark artifact can hit
EITHER engine depending on how the search was dispatched, and a number produced by one is not
evidence for the other. Mirror the launcher-attribution discipline above: before citing a
large-file timing as evidence of a change, confirm (don't assume) which of the two files the run
actually exercised — a backend_cpu.rs parallel-search improvement and a native_search.rs streaming
number are not interchangeable, and blending them the way a .cmd-shim timing gets blended into a
native-exe number would produce the same class of misleading artifact this section exists to prevent.
See tensor-grep-architecture-contract's Native-vs-Python routing section for the full split.
A speed win is not proof of correctness — prove output-identity separately
Corollary 5 above ("if a candidate is correct but slower, revert it") assumes you already know the
candidate is correct before you get to the speed question. Nothing else in this skill establishes
that — a benchmark script measures wall-clock; it does not diff output for you. When the change under
benchmark merges, skips, or reorders work (the common shape of a real optimization — one file walk
instead of two, one pass merging what used to be separate scans, an early-return pre-check), prove
byte-identical output two ways before trusting the speed number at all:
Enumerate every producer/branch and argue exhaustiveness. E.g. AST node types are mutually
exclusive, so merging two ast.walk() passes into one cannot drop a node type either pass would
have seen; a candidate name is always a substring of the file text it was extracted from, so a
substring pre-check that short-circuits can only ever skip work that was going to contribute a zero
bonus anyway.
Differential fuzz. Run OLD vs NEW over N real files/cases from this repo and assert zero
mismatches — not a hand-picked pair, a real corpus sweep.
Treat a build agent's own "I verified it's equivalent" as a hypothesis, not proof — an independent
reviewer re-running the fuzz check is the proof-of-record, the same independent-gate discipline this
repo applies to any other load-bearing self-report. Full recipe (profiling-probe on the shipped wheel →
this proof step → warm/cold microbench) lives in the global skill
profile-guided-byte-identical-optimization.
Regression gate mechanics you must understand before reading a red/green result
Loads current + resolves baseline (--baseline auto or an explicit path).
Suite mismatch → hard fail (exit 2) if baseline.suite != current.suite — you cannot diff a
hot-query artifact against a cold-path baseline even by accident.
Environment mismatch (detect_environment_mismatch, perf_guard.py:122-155) — different
platform/machine/python_version (major.minor only) between baseline and current → refuses
comparison unless --allow-env-mismatch is passed.
Comparator drift (detect_comparator_drift, perf_guard.py:87-119) — reports (but does not
fail on) any change in the rg_time_s comparator itself; a drifting rg baseline is host noise,
not a tg regression, but it should make you suspicious of the whole run.
Regression check (check_regressions, perf_guard.py:48-84) — per row, per suite-specific
time key (SUITE_TIME_KEYS: run_benchmarks → tg_time_s; run_hot_query_benchmarks →
first_s/second_s; anything else falls back to any key ending _time_s/_s), fails if
pct_delta > --max-regression-pct (default 5.0) andbase_time >= --min-baseline-time-s
(default 0.1s — the noise floor from above).
GPU claims need a stricter bar (read before trusting any GPU number)
State of the GPU program as of v1.75.4 (re-verify against docs/gpu_crossover.md, which is the
current source of truth; AGENTS.md "Roadmap Sequencing" predates the v1.75.x wave and can lag):
GPU Phase-0 SHIPPED (v1.75.0-v1.75.4, PRs #593-#597) and is locally correctness-proven (RTX 4070
sm_89 / RTX 5070 sm_120, 1GB/5GB match+file-set identity), but gated OFF the public release by the
CI Actions var TENSOR_GREP_RELEASE_NATIVE_ASSET_PROFILE (default native-frontdoor, CPU-only; GPU
asset publishing needs the non-default native-frontdoor-gpu) — Phase 1 is now a reversible
flag-flip, not a multi-week rebuild. That flip changes only whether the built assets are published;
it does not promote GPU, change the CPU-default auto-recommendation, or prove a speed crossover.
Keep the honesty floor: no speed crossover is proven vs rg/tg_cpu, GPU auto-recommendation stays
false, and the reviewer-gated public-gpu-proof.yml speed-crossover gate remains unmet
(grep -n "Public managed GPU promotion additionally requires" docs/CONTRACTS.md; corrected
2026-08-01 — the prior :80-82 citation pointed at the unrelated ripgrep-flag-compatibility list a
few dozen lines above the real promotion-contract paragraph, currently :123). Do not treat any GPU
number you produce as promotion evidence; it is implementation history at best.
Re-verified current as of v1.95.0: docs/gpu_crossover.md carries its own rotating
"Current post-<version> GPU dogfood Read" heading section — the <version> in that heading is
re-stamped per release (it was v1.95.0 when this sentence was first written and has rotated
since; grep -n "GPU dogfood Read" docs/gpu_crossover.md for the current one — this file
previously embedded the literal v1.95.0 heading text, a snapshot that goes stale every release),
and the verdict above is unchanged in substance — still no
single-pattern crossover, and the public managed binary still routes GPU requests through GpuSidecar
(not NativeGpuBackend). Promotion has grown a more detailed contract since v1.75.4 (unchanged
conclusion, more machinery): public promotion now additionally requires a managed NVIDIA front door
with tg-native-metadata.json provenance, and benchmarks/run_gpu_native_benchmarks.py --public-managed-proof emitting public_managed_promotion_ready = true and public_gpu_proof = true.
Public CUDA-asset publishing itself remains a deliberate CEO-decision hold (task-store #169 — not a
GitHub issue, re-verify with gh issue list before citing it as one). For the fuller current picture
(devices/doctor probes, --gpu-device-ids search, the WSL cross-domain probe fix, the honesty table of
what an observation does and does not license you to claim), see the sibling skill tensor-grep-gpu.
Two hard-earned rules if you do run a GPU benchmark:
The fair baseline for multi-pattern is a single rg -F -e ... -e ... invocation, not a
sequential loop of single-pattern rg calls. A sequential-rg comparator makes any batched
multi-pattern route look artificially faster. docs/benchmarks.md explicitly names this: "the
fair baseline is rg -F -e ... -e ...; sequential rg loops are exploratory amortization
evidence only."
CPU fallback or sidecar routing must never look like GPU proof. Any GPU-requested run that
actually executed on NativeCpuBackend or GpuSidecar must carry gpu_evidence_status = "unsupported", gpu_proof = false, native_gpu_unavailable, and not_gpu_proof_reason
(Backend Fail-Closed Contract, AGENTS.md). If you see a fast GPU-flag row with no sidecar_used
or native_gpu_unavailable field, the artifact is not trustworthy — the routing wasn't verified.
Worked example: the fixed-multi-pattern native CPU route (a fair-baseline correction, not a clean win)
This is the load-bearing lesson for this whole skill: a change that looks like a win against the
wrong comparator can be a loss against the right one.
What shipped (87d4ca4 fix: accelerate fixed multi-pattern native search, v1.11.3, then hardened
by 27386f8 fix: harden fair fixed multi-pattern search): a safe Aho-Corasick single-pass native CPU
route for fixed-string multi-pattern search (rust_core/src/native_search.rs), replacing what would
otherwise be N sequential single-pattern searches, with fallback preserved for unsupported semantics.
The naive comparator would have called this a big win: N sequential rg invocations (one process
spawn + one full-corpus scan per pattern) is obviously slower than one Aho-Corasick pass over the
corpus.
The fair-baseline correction changed the verdict.rg itself supports multi-pattern in a single
invocation (rg -F -e pat1 -e pat2 ...), and that — not the sequential loop — is the correct
comparator. Measured on the public managed v1.11.5 dogfood (docs/gpu_crossover.md:10), 100 fixed
no-match patterns over a 1GB corpus:
fell back to NativeCpuBackend — not a GPU number at all
A second mixed-pattern (2665 emitted matches) row told the same story: rg0.105s vs tg CPU
2.220s vs the GPU-requested row 2.211s (also NativeCpuBackend fallback).
What the repo actually did with this result (this is the discipline worth copying): it did not
revert the Aho-Corasick route — the code is still correct and is a real improvement over a sequential
loop, so benchmarks/run_native_cpu_benchmarks.py still exercises it — but it marked the row
non-gating: thresholds.large_file_200mb_fixed_multi_pattern_rows_are_diagnostic: true and
gated: False on both large_file_200mb_fixed_multi_pattern_no_match and _count cases
(benchmarks/run_native_cpu_benchmarks.py:330-349,396). Docs were corrected in the same spirit:
docs/gpu_crossover.md states the fair-baseline number plainly instead of the flattering
sequential-rg framing, and docs/benchmarks.md records it as "still failed the credibility bar
against the fair baseline" rather than as an accepted win.
UPDATE (2026-07-21, #251/#694) — "the code is still correct" is now FALSIFIED for the many-pattern
delegation path; re-read this worked example with that correction in mind. A live dogfood on the
SAME 100-pattern many-pattern path reproduced a real dedup over-count in the fast native Aho-Corasick
delegation route: total_matches: 3 where the rg-correct answer is 2 (one line matching two
different patterns from the set is counted once per (line, pattern) pair instead of once per line).
PR #694 shipped a guard test that reproduces and pins the CURRENT wrong behavior (not a fix), and the
many-pattern fast path is deliberately blocked from delegation until the real fix lands — native dedup
FFI-level correctness work, banked as a moat-investment option (task-store #255, not a GitHub
issue). This does not change the fair-baseline SPEED verdict above (Aho-Corasick is still slower than
fair rg regardless), but it means the CORRECTNESS half of "the code is still correct, just not fast
enough" is no longer true — treat any future many-pattern count from this path as suspect until the
dedup bug is fixed. Full detail: tensor-grep-failure-archaeology Battle 21.
The reusable lesson: when you batch/amortize N operations into one pass, benchmark against the
comparator's own batched primitive if it has one (rg -F -e ... -e ..., not N rg calls) — a
sequential-loop strawman will make almost any batching change look like a win. Ship the code if it's
a real structural improvement, but gate the release/doc claim on the fair-baseline number, and mark
the row diagnostic (not release-gating) until it actually beats that number.
Pre-flight checklist before writing a speed claim anywhere (PR description, docs/PAPER.md, AGENTS.md)
Ran the script from the decision table that matches what you changed (not a generic one).
Compared against --baseline auto (or the explicit current-accepted baseline file), not a
number from memory or an old PR.
Checked environment.tg_launcher_command_kind == native_exe (or explicitly disclosed
otherwise) — no shim/interpreter overhead hiding in the number.
Did not pass --allow-claim-unsafe-launcher (or if you did, the claim is explicitly labeled
exploratory, not accepted).
For any row under ~10ms, verified there is an absolute-seconds tolerance, not a bare percentage
gate.
For any change that merges/skips/reorders work, proved byte-identical output (enumerate the
branches + differential fuzz — see "A speed win is not proof of correctness" above), not just
assumed equivalence from a self-report.
For any cold-path claim, confirmed the number came from a cold microbench or a cache-cleared rep,
not a warm end-to-end dogfood run that may never have executed the changed function.
For any multi-pattern/batched claim, compared against the comparator's own batched primitive,
not a sequential-loop strawman — and did not cite a many-pattern match COUNT as correct without
checking the known live dedup over-count bug (#694) first.
For any large-file claim, confirmed WHICH engine the run hit (backend_cpu.rs PyO3/FFI
≥50MiB rayon path vs native_search.rs default streaming-serial path) — the two are not
interchangeable evidence.
For any GPU claim, confirmed gpu_proof/native_gpu_unavailable/sidecar_used fields show a
real native route, not CPU/sidecar fallback wearing a GPU label.
check_regression.py exit code is 0, or the regression is explicitly accepted and recorded in
docs/PAPER.md as an intentional non-goal (AGENTS.md Performance Discipline rule 5).
The artifact JSON is committed/attached, not just a terminal screenshot — someone else must be
able to re-run against it.
Provenance and maintenance
Facts here re-verified at tensor-grep v1.49.3 (2026-07-08); the §6 P4 close-out, the new §7
tg find retrieval-quality section, and the decision-table row were added and verified v1.78.1
(2026-07-16); v1.93.2 (2026-07-22) added the B-many-pattern dedup-bug caveat to the worked example
(the "code is still correct" framing is now falsified for that path, #694), the static-replay caveat
on the run_repo_retrieval_benchmarks.py decision-table row, and the A3 large-file
backend_cpu.rs-vs-native_search.rs disclosure hazard; v1.95.0 (2026-07-24) added the
warm-vs-cold regime trap + shipped-wheel cold-microbench recipe to the noise-floor section and the
byte-identical output-proof obligation (enumerate + differential-fuzz) to the fair-benchmark rules,
re-verified the GPU-status block against the current docs/gpu_crossover.md top section and added the
tensor-grep-gpu sibling-skill cross-reference, and re-counted the Benchmark Matrix (still 19). Every
file:line citation in the Noise-floor, Fair-benchmark-rules, Recipe, and Regression-gate-mechanics
sections was re-grepped this pass (check_regression.py, perf_guard.py,
run_hot_query_benchmarks.py, run_benchmarks.py, run_native_cpu_benchmarks.py,
run_gpu_benchmarks.py, docs/CONTRACTS.md) and all matched exactly — the Worked Example's historical
commit hashes (87d4ca4, 27386f8, 05f49b8) were confirmed to still resolve but the prose around
them was not re-walked. A candidate benchmarks/run_ast_parity_check.py decision-table row for the
Java/C#/PHP language-registration work was considered and dropped: that script's 40 parity cases cover
only python/javascript/typescript/rust (benchmarks/gen_corpus.pyAST_PARITY_CASES), so it does not
exercise the newly-registered languages — new-language correctness is a test_lang_registry pytest
concern (see tensor-grep-validation-and-qa), not a benchmarks/ one.
Re-verify before trusting a stale number:
Script inventory / artifact paths drift: grep -n "| .* | \benchmarks/run_" docs/benchmarks.md`
or re-read the "Benchmark Matrix" table.
Regression thresholds (--max-regression-pct, --min-baseline-time-s) and noise-floor constants:
python -c "import ast,sys" not needed — just re-Readsrc/tensor_grep/perf_guard.py and
benchmarks/check_regression.py argparse defaults; also
grep -n "NATIVE_REGEX_ABSOLUTE_JITTER_S" benchmarks/run_hot_query_benchmarks.py.
Fixed-multi-pattern worked-example numbers: grep -n "fixed multi-pattern\|multi_pattern" docs/gpu_crossover.md docs/benchmarks.md CHANGELOG.md — these are historical (v1.11.x) and
will not change, but the current GPU status (Phase-0 SHIPPED as of v1.75.0-v1.75.4, PRs #593-#597;
Phase 1 = a reversible flag-flip on TENSOR_GREP_RELEASE_NATIVE_ASSET_PROFILE, default
native-frontdoor CPU-only) should be re-checked since it is the fact most likely to have moved
further -- no speed crossover is proven vs rg/tg_cpu regardless of publish-flag state
(grep -n 'GPU benchmark correctness\|managed GPU promotion' docs/CONTRACTS.md).
This line cited docs/CONTRACTS.md:80-82 until 2026-08-02. That anchor was corrected earlier in
THIS FILE (see the GPU section's grep -n "Public managed GPU promotion additionally requires" docs/CONTRACTS.md note — pointer fixed 2026-08-13: this sentence previously pointed at "the GPU
section's grep -n "gpu_evidence_status" note", which does not exist anywhere in this file; the
only other gpu_evidence_status occurrence is the field name inside rule 2 of the GPU section,
not an anchor note) and the correction never
reached this duplicate 150 lines below, so the file shipped a fact and its refutation at once --
:80-82 is a --column/-c/--count-matches flag list, not a promotion contract. Fix a fact
-> grep the WHOLE doc for the old anchor; a correction applied at one site is not applied.
caveat (2026-07-21): this script is a STATIC-FIXTURE REPLAY that never calls chunk_file/rank_chunks, so it cannot detect a chunker or dense/RRF-weighting change at all.
quality gate (ndcg@10/recall@10) on the NL golden set + literal/identifier golden slices — NOT a speed benchmark; see tensor-grep-semantic-search-campaign STATUS UPDATE 2
tokens-per-correct-answer / token-economy (the moat metric — CANDIDATE, not yet a committed script, gated on #72)
none committed yet — see "6. Token-economy" below
a task-level cost metric (tokens spent to reach a correct answer, oracle-validated), not a latency metric — do not conflate with any row above
check_regression.py
GPU Phase-0/Phase-1 status: grep -n "RELEASE_NATIVE_ASSET_PROFILE\|native-frontdoor-gpu" .github/workflows/ci.yml and re-read docs/gpu_crossover.md's current top section (it carries its own
"Current post-<version> GPU dogfood Read" heading, updated per release) — cross-check the sibling
skill tensor-grep-gpu's own provenance stamp for the fuller current picture.
Current release tag: grep -n "release_docs_current_tag" AGENTS.md.
Token-economy harness promotion status (#72 — still uncommitted as of 2026-07-08?):
ls benchmarks/ | grep -i token and grep -n "token" docs/benchmarks.md (expect no hits until #72 lands).