#001MMBT-Messy-Model-Bench-Tests1 skills412updated 2026-05-30100% of creatorskilloccupationdescriptionupdatedmmbt-benchsoftware-developersDrive a full MMBT microbench run end-to-end with the self-healing autopilot: launch a target-N run (optionally via a model preset), monitor it live via the dashboard (grid / oneline / html / trend / flips / json), auto-recover the endpoint and stuck cells, grade + summarize, then generate the findings doc and publish a scorecard on the MMBT repo. Use when the user wants to run or resume the microbench, take a model to N replicates, watch the bench grid, analyze replicate stability, or publish bench results.2026-05-30