Skip to main content

llm-eval-harness

Evaluate any LLM behind an OpenAI- or Anthropic-compatible endpoint across four dimensions: speed (TTFT + thinking-aware tokens/sec), concurrency/stability (success rate, p50/p90 latency, the level where it breaks), Anthropic protocol compliance (thinking-block trigger rate), and quality regression against your own accumulated use cases (blind-judge precision). Use whenever someone wants to benchmark, 测评, or 压测 a model, verify a vendor's tokens-per-second claim, compare two models head-to-head, decide whether a newly released model is fast/stable/good enough before adopting it, measure TTFT or decode throughput, probe concurrency limits before a workshop or batch job, or check whether an "Anthropic-compatible" endpoint really implements thinking blocks. Triggers on "benchmark this model", "测一下这个模型的速度/ 质量", "is X tok/s real", "compare model A vs B", "这个模型能不能扛住并发", "接入新模型 先测一下", even without the word "eval".

설치로 이동

소스 정보

저장소
TuYv/ccpm
최근 소스 활동
2026년 8월 14일 19:45
감지된 SKILL.md 언어
영어
스타
1
포크
2

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.