Use when the user wants to pick the best LLM model (or combination of models) for a multi-step agent pipeline by running offline evaluation, including when no eval dataset exists yet and one must be generated. Trigger phrases include "which model should I use", "benchmark my planner/solver", "compare GPT-4o vs GPT-4o-mini for my agent", "model selection for my agent", "offline eval of model combos", "generate an eval dataset".
2026-07-18