Tracking issue for the task-format / infrastructure unification effort. Goal: every task on one unified format & infra (BaseTaskEnv/RLTaskEnv, @register_task, scenario: ScenarioCfg, sanctioned _-hooks only, self.handler for sim), deeply integrated and kept aligned by an automated guardrail.
Audit covered 23 families / 362 task files. Canonical contract: MetaSim/metasim/task/{base.py,rl_task.py,registry.py,gym_registration.py}.
P0 — seed contract + guardrail ✅ (done)
P1 — eliminate illegal step/reset/close overrides (move logic to sanctioned hooks)
Mechanism: action transforms → pre_physics_step_callback (or an action hook); reward computed in step → _reward; per-reset state → reset_callback/_get_initial_states. As each family lands, remove it from the reset(seed) allowlist and extend the guardrail to ratchet step/close.
step overrides (42): simpler_env (13), pick_place (10), mujoco_playground (3), beyondmimic (3), robosuite (2), mjlab (2), maniskill (2), libero/humanoid_bench/humanoid/dm_control/calvin/benchmark (1 each).
leaf reset-without-seed (90): libero_90 (62 — uniform "skip checker reset" pattern, collapse via a base class flag), pick_place (7), simpler_env/maniskill (4), mujoco_playground/mjlab/beyondmimic (3), robosuite (2), …
close overrides (11): simpler_env (3), robosuite/beyondmimic (2), mjlab/libero/dm_control/benchmark (1).
P2 — promote ManagerBasedRVEnv into metasim.task (Tier 2)
The manager-based RL base lives in roboverse_learn.managers; mjlab-v2/beyondmimic/humanoid build parallel manager stacks on it. Promote it to a sanctioned metasim.task base so those families have a blessed lineage; converge mjlab mdp/, beyondmimic mdp/, humanoid CallbacksCfg onto one shared mdp package (dedupe ~60% overlap).
P3 — merge duplicate bases
LiberoBaseTask ≈ Libero90BaseTask → one base (resolve the inverted checker-reset ordering)
- mjlab v1 vs v2 dual task-IDs → deprecate v1
- beyondmimic metasim vs isaaclab subtrees → one, with a backend selector
P4 — declare the 3 tiers explicitly + unify registration
- Tier 1 native
BaseTaskEnv; Tier 2 manager-based; Tier 3 external passthrough (gym registry).
- Give passthrough families (robotwin/metaworld/myosuite/gym_robotics/dm_control/libero_plus/maniskill-v3) thin
@register_task adapters so they're discoverable via list_tasks(), and document them as Tier 3.
- Collapse the 3 competing
SimplerEnv/ registrars into one; standardize registration prefixes to family.task (fix humanoid unitree_rl./g1./agibot_a2. and beyondmimic's missing prefix).
P5 — config standardization
max_episode_steps as class-var everywhere (not @property/instance/missing); prefer RobotCfg over bare robot strings; migrate DEFAULT_CONFIG-style dicts onto ScenarioCfg/@configclass.
Guardrail evolves with each P: the conformance test gains a step/close ratchet and a base-lineage check as the corresponding P lands, so unified format is enforced, not just achieved once.
Tracking issue for the task-format / infrastructure unification effort. Goal: every task on one unified format & infra (
BaseTaskEnv/RLTaskEnv,@register_task,scenario: ScenarioCfg, sanctioned_-hooks only,self.handlerfor sim), deeply integrated and kept aligned by an automated guardrail.Audit covered 23 families / 362 task files. Canonical contract:
MetaSim/metasim/task/{base.py,rl_task.py,registry.py,gym_registration.py}.P0 — seed contract + guardrail ✅ (done)
reset(seed)forwarded through 8 family bases (libero, libero_90, maniskill, rlbench, embodiedgen, calvin, pick_place, humanoid) +task_template— fix(tasks): unify reset(seed) contract across task base classes + conformance guardrail #790tests/test_task_reset_seed_contract.py(ratchets a 90-file allowlist that may only shrink) — fix(tasks): unify reset(seed) contract across task base classes + conformance guardrail #790embodiedgen/put_bananamax_episode_steps250000→250 — fix(tasks): correct embodiedgen put_banana max_episode_steps (250000 -> 250) #791P1 — eliminate illegal
step/reset/closeoverrides (move logic to sanctioned hooks)Mechanism: action transforms →
pre_physics_step_callback(or an action hook); reward computed instep→_reward; per-reset state →reset_callback/_get_initial_states. As each family lands, remove it from thereset(seed)allowlist and extend the guardrail to ratchetstep/close.stepoverrides (42): simpler_env (13), pick_place (10), mujoco_playground (3), beyondmimic (3), robosuite (2), mjlab (2), maniskill (2), libero/humanoid_bench/humanoid/dm_control/calvin/benchmark (1 each).leaf
reset-without-seed (90): libero_90 (62 — uniform "skip checker reset" pattern, collapse via a base class flag), pick_place (7), simpler_env/maniskill (4), mujoco_playground/mjlab/beyondmimic (3), robosuite (2), …closeoverrides (11): simpler_env (3), robosuite/beyondmimic (2), mjlab/libero/dm_control/benchmark (1).resetoverrides with a_skip_checker_resetbase flag (biggest mechanical win)pre_physics_step_callback, reward →_reward, episode bookkeeping →_get_initial_states; drop step/reset overridesresetsignature; dropsuper(PickPlaceBase, self)MRO-skips_native: routestepthroughhandler.set_dof_targets/simulate/get_states(stop directscene.step()/mj_step/set_drive_target) — parity-test each (cross-sim obs/reward)P2 — promote
ManagerBasedRVEnvintometasim.task(Tier 2)The manager-based RL base lives in
roboverse_learn.managers; mjlab-v2/beyondmimic/humanoid build parallel manager stacks on it. Promote it to a sanctionedmetasim.taskbase so those families have a blessed lineage; converge mjlabmdp/, beyondmimicmdp/, humanoidCallbacksCfgonto one sharedmdppackage (dedupe ~60% overlap).P3 — merge duplicate bases
LiberoBaseTask≈Libero90BaseTask→ one base (resolve the inverted checker-reset ordering)P4 — declare the 3 tiers explicitly + unify registration
BaseTaskEnv; Tier 2 manager-based; Tier 3 external passthrough (gym registry).@register_taskadapters so they're discoverable vialist_tasks(), and document them as Tier 3.SimplerEnv/registrars into one; standardize registration prefixes tofamily.task(fix humanoidunitree_rl./g1./agibot_a2.and beyondmimic's missing prefix).P5 — config standardization
max_episode_stepsas class-var everywhere (not@property/instance/missing); preferRobotCfgover bare robot strings; migrateDEFAULT_CONFIG-style dicts ontoScenarioCfg/@configclass.Guardrail evolves with each P: the conformance test gains a
step/closeratchet and a base-lineage check as the corresponding P lands, so unified format is enforced, not just achieved once.