| name | evotrace-evolutionary-coding-analysis |
| description | EvoTrace and EvoReplay methodology for diagnosing what evolutionary coding agents actually evolve. LLM-as-judge edit annotation, replay-based search state reconstruction, and controlled intervention analysis for agentic evolutionary search beyond final benchmark scores. Activation: evolutionary coding agent, LLM code evolution, EvoTrace, agentic search analysis, code generation mechanism. |
EvoTrace: Evolutionary Coding Agent Trace Analysis
A diagnostic framework for analyzing evolutionary coding agents that pairs LLMs with evolutionary search, revealing what mechanisms actually drive benchmark improvements beyond final scores.
Metadata
- Source: arXiv:2605.20086
- Authors: Nico Pelleriti, Sree Harsha Nelaturu, Zhanke Zhou, Zongze Li, Max Zimmer, Bo Han, Sebastian Pokutta
- Published: 2026-05-19
- Subjects: Neural and Evolutionary Computing (cs.NE); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Core Methodology
Key Innovation
Evolutionary coding agents (LLMs + evolutionary search) produce strong benchmark results, but progress is typically summarized only by final scores, masking the underlying mechanisms. This work introduces:
- EvoTrace — a dataset of evolutionary coding traces spanning 4 frameworks, reasoning and non-reasoning models, and 16 tasks
- EvoReplay — a replay-based methodology that reconstructs local search states behind high-scoring solutions
- LLM-as-judge edit annotation — 9 recurring edit types, validated against blind human re-annotation
- Controlled intervention analysis — adjusting constants, removing components, substituting models/contexts
Key Findings
- — most improvements come from a small subset of edit types