Use this skill when judging a single candidate response for instruction following and a scalar score is needed. It is especially useful when the sample contains exact constraints, a visible checklist, verifier resources, format requirements, word/count…
Qwen-Applications/Skill-RM
SkillsMP has collected 4 skills from Qwen-Applications/Skill-RM. Open a skill to review its source and details.
- Latest recorded source activity
- SkillsMP catalog refreshed
- skills collected
- 4
- GitHub stars
- 24
- GitHub forks
- 1
Skills in this repository
Showing 4 of 4 collected skills.
Use this skill when judging instruction following for IF-RewardBench-style samples. It is especially useful when the sample contains exact constraints, a visible checklist, format requirements, word/count limits, language restrictions, or multiple…
Use this Skill-RM reward judge to compare candidate responses for a visible user request with generic rubric, principles, bias controls, output contract, and Python sandbox checks over visible text.
Use this Skill-RM reward judge when response judging may benefit from resource-rich evidence: benchmark/task metadata, visible references or ground truth, checklists, verifier signals, code/math/factuality tool protocols, same-backbone judging pipeline…