Skip to content

VLA-90 PPO collection: released oracle_checkpoints are RL-keyed ('-v0'), and ENVS_CONFIG lacks the FindImposter*-VLA-v0 recipes #8

Description

@sypsyp97

Hi, thanks for releasing the VLA-90 task suite — the collection tooling is in great shape. While setting up self-collection for the -VLA-v0 PPO tasks I hit a gap I'd like to confirm, plus a question about the FindImposter family specifically.

What I observed (all on current main)

  1. The published oracle checkpoints are keyed to the RL env ids, so they cannot drive any -VLA-v0 collection as-is. oracle_checkpoints.zip from avanturist/mikasa-robo contains 32 final_success_ckpt.pt, all under directories named like ShellGameTouch-v0 (the RL-32 set). mikasa_robo_suite/vla/dataset_collectors/get_mikasa_robo_datasets.py builds its checkpoint map from those directory names and matches --env-id exactly (no -VLA-v0 → -v0 aliasing), so e.g. --env-id ShellGameTouch-VLA-v0 raises Checkpoint for env_id=... not found.

  2. No training recipe exists for the 9 FindImposter*-VLA-v0 tasks. vla/dataset_collectors/get_dataset_collectors_ckpt.py's ENVS_CONFIG has 68 entries (32 RL + 36 VLA, ending at index 67 = ShellGameColorLampTouch-VLA-v0); none of the FindImposter tasks appear. Meanwhile the collection side fully supports them (EPISODE_TIMEOUT_BY_ENV and parallel_dataset_collection_manager.py both list all 9), and the manifest marks them Data Source = PPO — so presumably you collected them from PPO oracles that aren't in the release.

  3. The HF dataset repos avanturist/MIKASA-Robo-VLA-{npz,rlds,lerobot} currently contain only a README — I assume the VLA-90 data release is still in progress.

Questions

  1. Do you plan to release -VLA-v0-keyed oracle checkpoints (ideally including the FindImposter family)? Any rough timeline for those and for the MIKASA-Robo-VLA-* data uploads?
  2. If the checkpoints won't be released soon: could you share the PPO hyperparameters you used for FindImposter*-VLA-v0 (an ENVS_CONFIG-style entry would be perfect)? I'm happy to train the oracles myself and can PR the entries back.
  3. Is reusing the RL -v0 checkpoints for VLA collection (by renaming the checkpoint directories) expected to work, given the VLA envs' randomized CUE_PHASE_STEPS / EMPTY_PHASE_STEPS? Or were the VLA datasets collected from oracles trained directly on the -VLA-v0 envs?

Thanks a lot!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions