Skip to main content

allenai/vla-evaluation-harness

SkillsMP has collected 3 skills from allenai/vla-evaluation-harness. Open a skill to review its source and details.

Latest recorded source activity
SkillsMP catalog refreshed
skills collected
3
GitHub stars
590
GitHub forks
49

Skills in this repository

2 occupation categories · 100% classified

Showing 3 of 3 collected skills.

occupation
Data Scientists
description

Run a VLA model evaluation against a simulation benchmark. Use this skill whenever the user wants to evaluate, benchmark, test, or run a model on a sim environment — even if they say it casually like 'try OpenVLA on LIBERO' or 'get me CALVIN scores'. Covers…

updated
occupation
Software Developers
description

Add a new VLA model server to the evaluation harness. Use this skill whenever the user wants to integrate, create, or add a new model — e.g. 'add OpenVLA server', 'integrate RT-2', 'hook up my model', 'write a model server'. Also use when they ask how model…

updated
occupation
Software Developers
description

Add a new simulation benchmark to the VLA evaluation harness. Use this skill whenever the user wants to integrate, create, or add a new benchmark or simulation environment — e.g. 'add ManiSkill3', 'integrate OmniGibson', 'hook up a new sim'. Also use when…

updated
Showing 3 of 3 collected skills.