Domain-agnostic Auto-Experiment framework for systematically improving any measurable outcome.
Use when the user wants to optimize, benchmark, A/B test, or iteratively improve any metric —
whether it's code performance, prompt quality, conversion rates, accuracy, latency, cost, or any
numeric target. Also use when the user says "run experiments", "optimize X", "test which approach
is better", "benchmark this", "systematic improvement", "A/B test", "auto-experiment", or any
request to compare multiple approaches to a measurable goal. Do NOT use for one-off fixes,
subjective tasks without a numeric metric, or single conversions with no comparison.
2026-07-18