Hi, I'm trying to reproduce the GenieSim manipulation benchmark results and would like to make sure our evaluation protocol matches the reported numbers. Could you clarify a few details?
- Are the reported π0.5 / π0 / GR00T manipulation results evaluated using a single checkpoint across all 10 manipulation tasks, or is each task fine-tuned separately?
- How many episodes are used per task to compute the reported success rates?
- Are results averaged over multiple
--benchmark.seed values?
- Are scene randomization and benchmark-instance randomization enabled during the official evaluation?
Thank you.
Idan
Hi, I'm trying to reproduce the GenieSim manipulation benchmark results and would like to make sure our evaluation protocol matches the reported numbers. Could you clarify a few details?
--benchmark.seedvalues?Thank you.
Idan