Evaluate an agent and harden its failure modes before it touches real traffic — an offline eval harness scored on task-completion, trajectory, and tool-use correctness against a fixed task set, plus loop hardening (step/tool-call caps, timeouts + retries,…
原文の言語: 英語