بنقرة واحدة
mlx-benchmarks
يحتوي mlx-benchmarks على 3 من skills المجمعة من dryvist، مع تغطية مهنية على مستوى المستودع وصفحات skill داخل الموقع.
Skills في هذا المستودع
Use when you have a raw suite results JSON to turn into a published dataset shard (convert, dry-run, publish with the write token), or when RANKINGS.md needs updating after a publish. Covers the --kind/--suite selection and the mandatory same-PR RANKINGS rule.
Use when driving the agentic tool-calling benchmark specifically — running harness/agentic/run.py against an endpoint, choosing the cell matrix, reading the multi-turn degradation track, or recovering partial results after a crash. For the general five-suite flow use the run-benchmark skill.
Use when about to benchmark an MLX or locally-served model on this repo's hosts — running one or more of the five required suites, producing a dataset shard, or deciding whether a model counts as "fully benchmarked". Routes to the canonical RUNBOOK and the traps that silently corrupt results.