Run a VLA model evaluation against a simulation benchmark. Use this skill whenever the user wants to evaluate, benchmark, test, or run a model on a sim environment — even if they say it casually like 'try OpenVLA on LIBERO' or 'get me CALVIN scores'. Covers…
allenai/vla-evaluation-harness
SkillsMP has collected 3 skills from allenai/vla-evaluation-harness. Open a skill to review its source and details.
- Latest recorded source activity
- SkillsMP catalog refreshed
- skills collected
- 3
- GitHub stars
- 590
- GitHub forks
- 49
Skills in this repository
2 occupation categories · 100% classified
Showing 3 of 3 collected skills.
skill
occupation
description
updated
occupation
Data Scientists
description
updated
occupation
Software Developers
description
Add a new VLA model server to the evaluation harness. Use this skill whenever the user wants to integrate, create, or add a new model — e.g. 'add OpenVLA server', 'integrate RT-2', 'hook up my model', 'write a model server'. Also use when they ask how model…
updated
occupation
Software Developers
description
Add a new simulation benchmark to the VLA evaluation harness. Use this skill whenever the user wants to integrate, create, or add a new benchmark or simulation environment — e.g. 'add ManiSkill3', 'integrate OmniGibson', 'hook up a new sim'. Also use when…
updated
Showing 3 of 3 collected skills.