The released demonstrations come from two collectors that run on different physics backends:
get_mikasa_robo_datasets.py (PPO) creates its envs with sim_backend="gpu".
get_mikasa_robo_datasets_motion_planning.py rolls out through _create_pd_ee_rollout_env, which calls gym.make(..., num_envs=1) without sim_backend. On a GPU node that env reports gpu_sim_enabled = False, i.e. the demonstrations are CPU PhysX rollouts.
benchmarking.py defaults to sim_backend="gpu" for every task, and nothing in the README says which backend a given task should be evaluated on.
We ran into this on BlinkCountButtonPressEasy-VLA-v0. In the released demos, many frames at the bottom of a button press command a downward move while the height does not change. Driving the hand below the demo press height gives different results on the two backends: on CPU the hand stops on the button, on GPU it sinks further and the closed gripper is pushed open. A policy trained on these demos can therefore end up in states the demos never contain. The env also samples the scene layout and blink count with the torch RNG on the sim device, so the same seed gives a different scene on each backend: open-loop replay of the demos succeeds on CPU and fails on GPU.
We have not gone through every task the motion-planning collector covers; this is the one where we found it. Could you state the intended backend per task, or record and evaluate each task on the same one?
The released demonstrations come from two collectors that run on different physics backends:
get_mikasa_robo_datasets.py(PPO) creates its envs withsim_backend="gpu".get_mikasa_robo_datasets_motion_planning.pyrolls out through_create_pd_ee_rollout_env, which callsgym.make(..., num_envs=1)withoutsim_backend. On a GPU node that env reportsgpu_sim_enabled = False, i.e. the demonstrations are CPU PhysX rollouts.benchmarking.pydefaults tosim_backend="gpu"for every task, and nothing in the README says which backend a given task should be evaluated on.We ran into this on
BlinkCountButtonPressEasy-VLA-v0. In the released demos, many frames at the bottom of a button press command a downward move while the height does not change. Driving the hand below the demo press height gives different results on the two backends: on CPU the hand stops on the button, on GPU it sinks further and the closed gripper is pushed open. A policy trained on these demos can therefore end up in states the demos never contain. The env also samples the scene layout and blink count with the torch RNG on the sim device, so the same seed gives a different scene on each backend: open-loop replay of the demos succeeds on CPU and fails on GPU.We have not gone through every task the motion-planning collector covers; this is the one where we found it. Could you state the intended backend per task, or record and evaluate each task on the same one?