ワンクリックで
model-benchmark
Runs benchmark suites against models, tracks ELO rankings, manages A/B tests, and generates performance reports.
Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。
メニュー
Runs benchmark suites against models, tracks ELO rankings, manages A/B tests, and generates performance reports.
Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。
SOC 職業分類に基づく
Sends messages between agents, broadcasts to channels, and retrieves message history for inter-agent communication.
Spawns, manages lifecycle, and terminates autonomous AI agents from the 40+ built-in agent catalog.
Optimize prompts and context windows to reduce token usage while preserving quality.
Animated desktop character state machine for Sven's Tauri companion app. Manages character form (ORB, ARIA, REX, ORION), state transitions (idle→thinking→speaking→celebrating), walk cycles, thought bubbles, sound effects, and real-time agent event sync.
Multi-model deliberation system. Sends queries to multiple LLMs simultaneously, has them peer-review each other's responses anonymously, then a chairman model synthesizes the best answer. Supports configurable council composition, voting strategies, and cost tracking.
Interactive educational autograd engine — port of Karpathy's micrograd. Build, train, and visualise tiny neural networks step-by-step to learn how backpropagation works.
| name | model-benchmark |
| description | Runs benchmark suites against models, tracks ELO rankings, manages A/B tests, and generates performance reports. |
| version | 0.1.0 |
| publisher | acmecorp |
| handler_language | typescript |
| handler_file | handler.ts |
| inputs_schema | {"type":"object","properties":{"action":{"type":"string","enum":["list_suites","create_run","complete_run","leaderboard","report","record_elo","ab_results"]},"suite_id":{"type":"string"},"model_id":{"type":"string"},"model_name":{"type":"string"},"run_id":{"type":"string"},"winner_id":{"type":"string"},"loser_id":{"type":"string"},"is_draw":{"type":"boolean"}},"required":["action"]} |
| outputs_schema | {"type":"object","properties":{"result":{"type":"object"}}} |
Evaluates model performance with standardised benchmark suites, ELO-style rankings, A/B testing, and cost/quality/speed reporting. Provides a leaderboard and per-model deep-dive reports.