職業分類
未分類
説明
Personal model benchmark. Replays the user's real recurring tasks (mined from their own Claude Code session history) against different models and effort levels in isolated headless sessions, grades the outputs blind against the user's own rubric, and produces…
原文の言語: 英語
更新