| name | embodied-eval-automation |
| description | Plan, explain, build, run, monitor, validate, transfer, and audit reproducible embodied-model studies and batch episode collection. Use when a user wants to connect a local, SSH, or cloud GPU host; understand and compare a policy, VLA, world model, world-action model, or hybrid with a benchmark; discover and reuse existing assets; choose between inference, closed-loop, prediction, data-collection, or training-readiness modes; preserve native and unified episode formats; validate required model outputs; run single-request through goal-driven studies; recover interrupted jobs; manage media, disk, GPU, transfer, repair, and delivery evidence; or produce human-readable and auditable reports. |
Embodied Eval Automation
Treat every run as a resumable, evidence-backed research state machine. First help the user understand the model, benchmark, available modes, resource costs, outputs, and claim limits. Then execute only the selected and approved scope. Do not equate a live process, installed environment, paper claim, repository implementation, or successful import with a verified capability or completed study.
Start with safe intake
- Read applicable repository instructions and task records.
- Confirm local code/data roots, remote code/data roots, output destination, read-only paths, and existing assets. Never infer a run workspace from the skill directory, plugin repository, current directory, home directory, or filesystem root.
- Ask only for missing facts. Never request a password, token, private-key body, recovery code, or session cookie in chat.
- Separate approvals for inspection, writes, installation, download, upload, paid compute, permission changes, deletion, Git writes, and external messages.
- Record approvals, paths, hosts, time/scope limits, GPU IDs, cost, pause thresholds, forbidden actions, and expiry in receipts.
- Use prior-run evidence only when the user identifies a continuation or the run identity matches; label stale evidence and never inject unrelated workspace history into a new intake.
Read onboarding-and-permissions.md for intake and approval rules and provider-connections.md before configuring SSH, cloud, GitHub, or Hugging Face access.
Establish an informed run contract
Do not freeze the final run contract until G1 has completed official discovery, capability activation analysis, mode/resource comparison, and user selection. Record:
- research or engineering questions and allowed/forbidden claims;
- selected mode and reviewed alternatives;
- model, benchmark, repository, checkpoint, dataset, encoder, and license identities;
- each capability as
required, desired, not_needed, or unknown;
- activation evidence across paper, repository, checkpoint, wrapper, adapter, writer, validator, and visualizer;
- suite/task/init/seed/repetition scope and deterministic expected IDs;
- representations, fields, media, analysis, transfer, retention, and completion requirements;
- separate VRAM, system RAM, swap, disk, staging, transfer, time, and cost budgets;
- gate states and any explicit waiver, skip, exclusion, or non-applicability.
Use embodied-run-contract/v2. Recognize v1 only to provide migration guidance; never infer new permissions or requirements from an old contract. Read model-benchmark-capability-discovery.md and mode-selection-and-resource-planning.md before completing G1.
Follow the G0-G9 gates
Do not scale or broaden claims while the previous required gate is incomplete. Gate states are planned, approved, running, completed, skipped_by_user, excluded, not_applicable, blocked, or failed_operationally.
G0 - Access, host, and asset preflight
- Establish a safe connection method and verify host identity.
- Inspect OS, GPU/VRAM, driver, system RAM, swap, disks, network, ports, containers, and rendering.
- Inventory benchmark/model repositories, environments, checkpoints, datasets, caches, adapters, and prior runs read-only.
- Record purpose, size, version evidence, active-process status, and reuse candidacy. A matching name is not reuse evidence.
- Stop before installation, download, upload, deletion, or paid-compute changes.
Read asset-reuse-and-environments.md.
G1 - Official understanding, capability fidelity, mode selection, and source lock
Complete four internal checkpoints:
- G1A: explain the official model and benchmark workflow, repositories, modes, inputs, outputs, metrics, limitations, licenses, and result categories.
- G1B: build the capability activation matrix. Distinguish
official-stated, code-inspected, locally-measured, and estimated evidence.
- G1C: offer two to four bounded mode cards showing assets, VRAM/RAM/disk/time/cost, outputs, risks, and claim limits.
- G1D: obtain informed user selection of mode, required outputs, media, analysis, scope, and accepted degradation.
Compare every requirement with G0 assets and classify it reuse, reuse-readonly, rebuild, or missing. Do not enter G2 until the selected mode and required outputs are frozen. If a required capability cannot be enabled end to end, stop for repair, mode change, explicit capability waiver, or termination.
G2 - Environment and protocol layout
- Reuse verified benchmark environments read-only and fingerprint them before and after use.
- Isolate model and benchmark dependencies by default.
- Connect them through a versioned loopback protocol on
127.0.0.1 unless a separately approved secure network design exists.
- Never modify an existing working environment merely to save setup time.
G3 - Missing resource acquisition
- Acquire only
rebuild or missing resources at fixed revisions with resumable partial files and hashes.
- Measure the specific transfer bottleneck before changing route.
- Escalate from canonical remote download to approved mirror to user-controlled local download/upload without changing artifact identity.
G4 - Runtime and mutable-boundary validation
- Verify imports, accelerator backend, headless rendering, task enumeration, construction,
reset, one safe step, service health, and schema compatibility.
- Probe whether benchmark/model APIs mutate NumPy arrays, Torch storage/views, shared memory, images, states, actions, or mutable containers in place.
- Copy and hash boundary objects before third-party calls. Preserve model-native, normalized, pre-step executed, and post-call transformed actions separately with units, scaling, and provenance.
Read adapter-contract.md.
G5 - Real requests for every required output
For each required output head or mode, issue at least one genuine request and validate native preservation, shape, dtype, units, bounds, masks, horizon, latency, VRAM/RAM, writer persistence, validator coverage, and visualizer exposure. Verify whether multiple outputs originate from the same request when that matters. Do not call env.step unless separately approved.
A paper or repository capability is not verified when the selected checkpoint, wrapper, adapter, writer, validator, or visualizer remains disabled or unknown. Give a user-readable handoff before G6.
G6 - One closed-loop episode, three representations, and behavior explanation
Before the first episode, freeze field mappings, unavailable fields, media, and output paths for:
official-native;
embodied-eval-current;
embodied-eval-candidate.
Derive all three from one real rollout with common rollout/task/initial-state/environment facts and distinct representation IDs/manifests. Preserve all native outputs. Produce one full replay by default, or record a media waiver. Return source/transform/missing-field analysis, a static fallback visualization, and a plain-language walkthrough of what the robot did, what succeeded or failed, and what remains unproven.
Read episode-representations.md and human-readable-analysis-and-video.md.
G7 - Result-blind pilot, format decision, and study handoff
- Materialize and hash the expected pilot set before rollout using task/init metadata but no outcome fields.
- Cover multiple tasks and initial states; report coverage limitations for random, interval, or user-specified alternatives.
- Validate schema, same-rollout lineage, expected IDs, duplicates, pair identity, initial-state fingerprints, conversion loss, required capabilities, analysis, and media.
- Report whether to promote the candidate format and wait for explicit approval.
- Provide a behavior report, failure evidence levels, measured resource growth, and bounded G8 study options.
For paired comparisons, use pair_key=<benchmark>/<suite>/task=<task_id>/init=<init_state_id>/seed=<seed>, equal initial-state fingerprints, T+1 observations for T transitions, one policy_queries entry per real model call, complete raw action chunks, and separate executed actions.
G8 - Approved goal-driven study execution
G8 is not synonymous with a large batch. Never present "batch G8" as the automatic next step after a pilot; a larger collection is only one possible approved study track. Freeze a study plan containing question, hypothesis, variables, controls, paired design, task/init/seed/repetitions, sample-size or coverage rationale, required capabilities, metrics, statistical and qualitative analyses, expected IDs, resources, thresholds, early stop, and claim limits.
Choose a justified track: scale/stability, generalization, capability comparison, model/checkpoint comparison, ablation, failure study, qualitative study, dataset collection, derived analysis without new model calls or steps, or fine-tuning readiness. Change one dimension at a time unless a predeclared factorial design makes interactions interpretable.
Use a remote supervisor, independent controller, ledger, heartbeat, transfer acknowledgements, and periodic agent checker when available. State honestly which layers are alive. Offer smaller paired, failure-focused, media-only, analysis-only, or sharded alternatives when resources are constrained. If the user skips G8, record skipped_by_user and let G9 bind to the latest approved completed scope.
Read goal-driven-study-design.md and monitoring-transfer-retention.md.
G9 - Audit, explanation, and delivery
- Rebuild indexes from raw summaries and append-only lifecycle records; do not trust an existing global index as the sole source.
- Reconcile successful episodes, explicit benchmark failures, operational failures, corrupt/incomplete items, exclusions, skipped gates, missing IDs, and duplicates.
- Verify manifests, hashes, representations, pairs, initial states, transitions, queries, repairs, media, transfers, licenses, secrets, paths, and claims.
- Deliver a decision summary, behavior analysis, engineering audit, manifests, commands, environment fingerprints, failure/exclusion lists, static media, and reproduction/resume package.
An explicit_benchmark_failure is a complete episode when its contracted artifacts validate; do not rerun or hide it automatically. Operational failures, corruption, and incomplete episodes do not count as complete unless the frozen contract authorizes a bounded repair/rerun. Read repair-and-evidence-precedence.md and workflow-and-gates.md.
Preserve evidence and protect transfers
- Use exact normalized relative paths in manifests; never exclude all files sharing a basename.
- Reconcile declared and actual file sets, case collisions, absolute paths, traversal, symlinks, and hardlinks.
- Keep append-only repair history and compute effective state through explicit precedence/supersession.
- Write archives to partial names, verify both-end size/hash, inspect members, validate, then atomically finalize.
- Prune remote data only after immutable local verification, complete audit, path-boundary checks, approval, and durable receipts.
Read security-and-approvals.md before high-impact actions.
Keep training outside the default scope
Explain inference, evaluation, data collection, fine-tuning readiness, and full training separately. This skill may assess readiness and generate a training handoff covering data, model, resources, metrics, licenses, and post-training paired evaluation. Do not start fine-tuning or full training without a separately selected training workflow and explicit permissions.
Use bundled scripts
create_run_workspace.py: create a v2 workspace after write approval.
validate_run_contract.py: validate v1/v2 contracts and v2 logic.
validate_capability_contract.py: enforce required capability activation or waivers.
estimate_run_resources.py: distinguish fixed, resident, peak, staging, transfer, and media budgets.
probe_mutable_boundary.py: copy, hash, and detect mutation/storage aliasing around third-party APIs.
select_pilot_set.py: select deterministic result-blind pilot coverage.
render_episode_replay.py: render truthful replay/media manifests with optional FFmpeg.
audit_episode_set.py: reconcile IDs, lifecycle, manifests, representations, failures, and media.
validate_delivery.py: inspect archive paths, size/hash, secrets, private paths, restricted assets, licenses, and claims.
validate_repository.py: validate the distributed skill/plugin package.
Run scripts with --help first. Do not edit user data merely to make validation pass.
Require user-readable handoffs
After G1, G5, G6, and G7, state what was done, inputs, outputs, files and viewing instructions, proven and unproven claims, enabled and disabled capabilities, alignment with the original goal, next options, costs/benefits/risks, and approval needed. Separate verified_fact, strong_inference, possible_explanation, and unknown_without_additional_experiment.
Completion standard
Complete only when the latest approved scope reconciles, required capabilities or waivers resolve, machine evidence validates, video exists or is waived, reports are human-readable, claims remain within evidence, and another operator can reproduce or resume without chat history. A full benchmark is optional; the approved study is the completion boundary.