| name | model-matrix |
| description | Use this to compare outcomes across AI models and decide which model serves which kind of work. No global rankings. |
Model Matrix
The Big Idea
Every AI system is a different collaboration environment. The question is
never "which model is best." It is "which model makes THIS person most
effective at WHICH kind of work."
The Matrix
Build TASK TYPE x MODEL x MY RESULT using only episodes with enough data:
Dimensions worth scoring per cell:
- task completion rate
- rework and correction loops
- corrections the human had to make
- planning quality of joint output
- debugging quality
- unnecessary complexity injected
- validation discipline maintained
- the human's own verification behavior (does it drop with some models)
- prompt length required to get there (normalize across models)
- frustration and escalation signals
- learning retained afterward
- outcome quality at follow-up
Rank cells, never models globally.
Switch Forensics
Study every significant task migration between tools. Classify the reason:
first model failed, seeking second opinion, tool capability, context
limits, research capability, frustration, confirmation seeking, habitual
preference, cost or speed, unknown.
Then grade: did switching improve the outcome?
Separate the two kinds of switch:
- effective: new environment unlocked new evidence or capability
- circular: same reasoning loop restarted with a new logo on it
Count circular switches per trigger. Frustration-triggered switches are
the usual suspects.
Consensus Behavior
When models disagree, classify resolution: primary evidence consulted, a
test was run, reasoning compared, majority chosen, most confident voice
chosen, preferred answer chosen, asked more models until one agreed.
Report the dominant pattern with examples and counterexamples.
One-Line Memory
No global winners. Match model to work.