#25coleam00/kimi-k3-reliability-benchmark1 skillsreliability-benchmarkRun a first-party reliability benchmark comparing two or more LLMs on hard, trap-laden tasks through one identical harness, graded by reading and scored pass^k. Use when the user wants to benchmark, compare, or A/B models for reliability (not raw capability),…2026년 7월 21일