Skip to main content
Manus에서 모든 스킬 실행
원클릭으로

openjudge

// Build custom LLM evaluation pipelines using the OpenJudge framework. Covers selecting and configuring graders (LLM-based, function-based, agentic), running batch evaluations with GradingRunner, combining scores with aggregators, applying evaluation strategies (voting, average), auto-generating graders from data, and analyzing results (pairwise win rates, statistics, validation metrics). Use when the user wants to evaluate LLM outputs, compare multiple models, design scoring criteria, or build an automated evaluation system.

$ git log --oneline --stat
stars:619
forks:52
updated:2026년 3월 10일 09:06
파일 탐색기
5 개 파일
SKILL.md
readonly