Optimize HarnessGym tensor-layout kernel_plan.json tasks with a generated MCP server for plan validation, dev/final benchmarking, trace analysis, rollback-safe search, candidate application, history comparison, and experiment ranking. Use when a task asks to minimize best_cycles for benchmark.py or tune tensor layout/DMA/pipeline knobs under .harnessgym.
Use for the H100 Triton fused RMSNorm + SiLU gate optimization task. Provides the workflow and MCP tooling for objective runs, rollback-safe config sweeps, source diagnostics, benchmark history, and final held-out verification.