fireworks/.../glm-5p2 | 51 | default | Everyday text coding, long-horizon marathons, project context. Not: vision, frontier reasoning, verification (silent under-count). |
fireworks/.../kimi-k2p7-code | 42 | vision + token-efficiency | Image/screenshot/video input, very long tool loops. Not: cheap quick calls (thinking can't disable), high-stakes reasoning. GLM 5.2 ties/wins elsewhere. |
fireworks/.../minimax-m3 | 55/44 | never coding — verification + cheap multimodal | Doc-vs-code reconciliation (fewer false alarms), cheap image/video. Not: coding, code review, missed-discrepancy-critical verification. |
openai-codex/gpt-5.6-luna | 51.2‡ | vision/computer-use only | Vision/computer-use subtasks (GLM 5.2 is text-only), narrow execution with a plan. Not: hard reasoning, long-horizon, large-context, fire-and-forget. GLM 5.2 wins on text coding. |
openai-codex/gpt-5.6-sol | 59‡ | agentic breadth — capability upgrade, NOT trust upgrade | Terminal-native agent loops, computer-use/browser, web-research, ultra parallelizable tasks. Not: unsupervised destructive-tool loops (METR: highest reward-hacking of any public model; system-card overreach/fabrication), routine coding (GLM 5.2 cheaper + safer). Price-neutral vs old 5.5. |
anthropic/claude-sonnet-5 | 53 | agentic execution + injection robustness | Defined-plan execution, untrusted-input ingestion, vision/computer-use. Not: not cheaper than Opus per task ($2.29 vs $1.97), unsupervised security fixes (SecPass 19.6%), open-ended planning. |
anthropic/claude-opus-4-8 | 61.4 | surgical depth + review | Hard novel multi-file changes, code review (with human in loop), cyber/bio work (Fable reroutes here). Not: high-volume loops, unsupervised negotiation. "Beats GPT-5.5 on hard coding" reverses on DeepSWE. |
anthropic/claude-fable-5 | 60 | juggernaut (conscious, async) | Ambiguous investigative long-horizon work, large migrations, multi-day autonomous + self-verification. Not: interactive (109s TTFT), cyber/bio (use Opus 4.8), sole security reviewer (19% SecPass + memorization). |