Skip to content

Reproducing GenieSim manipulation benchmark scores #169

Description

@Idan-BenAmi

Hi, I'm trying to reproduce the GenieSim manipulation benchmark results and would like to make sure our evaluation protocol matches the reported numbers. Could you clarify a few details?

  1. Are the reported π0.5 / π0 / GR00T manipulation results evaluated using a single checkpoint across all 10 manipulation tasks, or is each task fine-tuned separately?
  2. How many episodes are used per task to compute the reported success rates?
  3. Are results averaged over multiple --benchmark.seed values?
  4. Are scene randomization and benchmark-instance randomization enabled during the official evaluation?

Thank you.
Idan

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions