Environment
- Branch:
main, verified on commit da42434 (GenieSim 3.2.0)
- Teleop launch:
source/geniesim_teleop/scripts/autoteleop.sh
- Teleop client:
source/geniesim_teleop/src/geniesim_teleop/teleop.py
- Randomization engine:
source/geniesim_benchmark/src/geniesim_benchmark/utils/generalization_utils.py
source/geniesim_benchmark/src/geniesim_benchmark/benchmark/envs/base_env.py::apply_generalization() (line 153)
source/geniesim_benchmark/src/geniesim_benchmark/app/controllers/api_core.py::shuffle_scene() (line 1112) / _shuffle_scene() (line 2268)
Part D — Multi-episode capture, and a pkill pattern that also kills the simulator
Current behavior. In source/geniesim_teleop/scripts/autoteleop.sh, after one
episode the user presses y/n; each branch runs the same teardown (lines 146
and 157) and then breaks out of the loop (lines 154 and 165), so the script
exits:
docker exec "$CONTAINER_NAME" bash -c "pkill -SIGTERM -f '$PROCESS_CLIENT' 2>/dev/null || true"
Collecting each additional episode therefore requires restarting all four
terminals (simulator / bridge / motion control / teleop client).
This is consistent with the script's documented design (one launch = one
episode), so the teardown itself is not a defect. It does, however, make
multi-episode capture impossible without restarting Isaac Sim, which is by far
the slowest part of the loop.
One detail worth noting for anyone implementing this. PROCESS_CLIENT is
defined at line 26 as:
PROCESS_CLIENT="teleop|ros|geniesim_teleop"
pkill -f matches against the full command line, and the simulator is
launched (line ~104) as:
omni_python ${SIM_APP} --config ${TELEOP_YAML}
where TELEOP_YAML (line 37) expands to a path ending in config/teleop.yaml.
That command line contains the substring teleop, so this pattern matches the
simulator process too. Today that is harmless (everything is meant to shut
down), but any change that keeps the simulator alive between episodes will need
a narrower pattern that cannot match it.
Suggestion. Support ending the current episode gracefully and staying
ready for the next one, instead of tearing everything down:
- Add a "stop current episode" signal (e.g. a
/sim/stop_episode ROS topic)
that the simulator listens to, so it can close the current ros2 bag cleanly
(write a complete metadata.yaml) and re-arm for the next recording.
- Reserve a separate key (e.g.
q) for "finish all and shut down".
- Narrow the
pkill pattern (e.g. match the exact client entry point
geniesim_teleop.teleop / geniesim_teleop/bridge.py) so it can never match
the simulator's command line.
This turns collection into: press controller button to start an episode ->
y/n to keep/discard and continue -> q to finish everything.
Part E — Domain randomization exists but is not wired into teleop collection
The repo already contains a full randomization engine — robot base/joint jitter,
lighting and material randomization via generalization_utils.py +
base_env.py::apply_generalization() — plus a ready-made
api_core.py::_shuffle_scene() that randomly offsets movable object positions.
However, on current main the only call sites are in the model-evaluation
path, source/geniesim_benchmark/src/geniesim_benchmark/benchmark/task_benchmark.py:
- line 263 —
self.env.apply_generalization(self.api_core, self.task_config)
- line 508 —
self.api_core.shuffle_scene()
- line 713 —
envs[slot].apply_generalization(...)
The teleoperation collection path (autoteleop.sh ->
geniesim_teleop/teleop.py) never calls any of them.
Consequence. During teleop collection, object positions are fixed (baked
into the scene .usda), so a policy trained on the collected data learns to
"reach a fixed coordinate" rather than to perceive the object; performance
collapses as soon as the object is moved.
Suggestion. Expose the existing randomization (shuffle_scene() and/or
apply_generalization()) to the teleop collection loop, so each episode can
start from a randomized object layout without users having to reposition
objects by hand. The engine is already there; it just isn't connected on the
collection side.
Environment
main, verified on commitda42434(GenieSim 3.2.0)source/geniesim_teleop/scripts/autoteleop.shsource/geniesim_teleop/src/geniesim_teleop/teleop.pysource/geniesim_benchmark/src/geniesim_benchmark/utils/generalization_utils.pysource/geniesim_benchmark/src/geniesim_benchmark/benchmark/envs/base_env.py::apply_generalization()(line 153)source/geniesim_benchmark/src/geniesim_benchmark/app/controllers/api_core.py::shuffle_scene()(line 1112) /_shuffle_scene()(line 2268)Part D — Multi-episode capture, and a
pkillpattern that also kills the simulatorCurrent behavior. In
source/geniesim_teleop/scripts/autoteleop.sh, after oneepisode the user presses
y/n; each branch runs the same teardown (lines 146and 157) and then
breaks out of the loop (lines 154 and 165), so the scriptexits:
Collecting each additional episode therefore requires restarting all four
terminals (simulator / bridge / motion control / teleop client).
This is consistent with the script's documented design (one launch = one
episode), so the teardown itself is not a defect. It does, however, make
multi-episode capture impossible without restarting Isaac Sim, which is by far
the slowest part of the loop.
One detail worth noting for anyone implementing this.
PROCESS_CLIENTisdefined at line 26 as:
PROCESS_CLIENT="teleop|ros|geniesim_teleop"pkill -fmatches against the full command line, and the simulator islaunched (line ~104) as:
where
TELEOP_YAML(line 37) expands to a path ending inconfig/teleop.yaml.That command line contains the substring
teleop, so this pattern matches thesimulator process too. Today that is harmless (everything is meant to shut
down), but any change that keeps the simulator alive between episodes will need
a narrower pattern that cannot match it.
Suggestion. Support ending the current episode gracefully and staying
ready for the next one, instead of tearing everything down:
/sim/stop_episodeROS topic)that the simulator listens to, so it can close the current
ros2 bagcleanly(write a complete
metadata.yaml) and re-arm for the next recording.q) for "finish all and shut down".pkillpattern (e.g. match the exact client entry pointgeniesim_teleop.teleop/geniesim_teleop/bridge.py) so it can never matchthe simulator's command line.
This turns collection into: press controller button to start an episode ->
y/nto keep/discard and continue ->qto finish everything.Part E — Domain randomization exists but is not wired into teleop collection
The repo already contains a full randomization engine — robot base/joint jitter,
lighting and material randomization via
generalization_utils.py+base_env.py::apply_generalization()— plus a ready-madeapi_core.py::_shuffle_scene()that randomly offsets movable object positions.However, on current
mainthe only call sites are in the model-evaluationpath,
source/geniesim_benchmark/src/geniesim_benchmark/benchmark/task_benchmark.py:self.env.apply_generalization(self.api_core, self.task_config)self.api_core.shuffle_scene()envs[slot].apply_generalization(...)The teleoperation collection path (
autoteleop.sh->geniesim_teleop/teleop.py) never calls any of them.Consequence. During teleop collection, object positions are fixed (baked
into the scene
.usda), so a policy trained on the collected data learns to"reach a fixed coordinate" rather than to perceive the object; performance
collapses as soon as the object is moved.
Suggestion. Expose the existing randomization (
shuffle_scene()and/orapply_generalization()) to the teleop collection loop, so each episode canstart from a randomized object layout without users having to reposition
objects by hand. The engine is already there; it just isn't connected on the
collection side.