Deep, structured analysis for backtest ledgers, trade logs, daily PnL series,
scorecards, and A/B comparison packs. Use when the user provides concrete
parquet/csv/json backtest artifacts and wants a decision such as promote,
watch, or reject based on the data itself. Covers return, risk, stability,
and trade quality, prevents single-metric conclusions, and prevents total PnL
from dominating the verdict. Do not use for open-ended root-cause diagnosis
without a concrete dataset, or for implementation-vs-research consistency
audits.
2026-03-23