TrainArena is a comprehensive Unity + ML-Agents environment for hands-on reinforcement learning development. Built on Unity 6.2 with ML-Agents v4.0.0, it provides both working examples and a production-ready foundation for custom RL agents.
🎯 Current Status: Production-ready with working cube agent training and advanced ragdoll locomotion system.rainArena — Project Summary
TrainArena is a compact Unity + ML-Agents playground for “learn-by-doing” reinforcement learning. It ships a clean Core (Cube→Goal PPO) and optional Add-ons (Tag mini-game, Self-Play, Ragdoll locomotion, Domain Randomization, Reward HUD, Recording, TensorBoard overlay, Model Hot-Reload) so you can get a result fast, then scale up to showy demos and reusable NPC controllers.
🧊 CubeAgent (Production Ready):
- 14-observation navigation system (velocity + goal direction + 8 raycasts)
- Proven training pipeline with 500K+ step validation
- Perfect inference performance with trained models
- 4x4 multi-arena training for efficient learning
🤖 RagdollAgent (Advanced System):
- Hierarchical skeleton with 6 PDJointControllers (hips, knees, ankles)
- Joint-based locomotion with coordinated movement learning
- 16 observations (uprightness + velocity + joint states)
- Natural physics tuning for realistic ragdoll behavior
🔧 Development Infrastructure:
- Comprehensive debug visualization system (R/I/O/V/B/M/T/N/H/Z controls)
- AutoBehaviorSwitcher for seamless training/testing modes
- Real-time ML-Agents status monitoring and performance tracking
- Domain randomization for robust agent training
📋 Production Quality:
- Complete documentation with accurate code comments
- Consistent logging systems and debug utilities
- Scalable scene generation with EnvInitializer
- Clean architecture ready for extension
-
Unity 6.2+ → Package Manager: install ML-Agents and Barracuda.
-
Unzip TrainArena-Starter into your Unity project (merge Assets/ and Tools/).
-
Unzip TrainArena-Extras on top (merge again).
-
Open Unity → Tools → ML Hack → Build Cube Training Scene → Play (sanity check).
-
From your project root in a terminal:
pip install mlagents tensorboard
mlagents-learn Assets/ML-Agents/Configs/cube\_ppo.yaml --run-id=cube --train
tensorboard --logdir results
- When training exports a .onnx, set the agent’s Behavior Type to Inference Only and assign the model.
- Using the “Complete” consolidated pack instead? Paths are under Assets/TrainArena/Core and Assets/TrainArena/Addons. With Starter+Extras, use the flat Assets/... paths shown below.
- What: A minimal agent that moves to a target and avoids obstacles.
- Why: Fastest way to see PPO learning end-to-end.
- Use: Tools → ML Hack → Build Cube Training Scene → press Play. Train with:
mlagents-learn Assets/ML-Agents/Configs/cube_ppo.yaml --run-id=cube --train
- Where: Assets/Scripts/CubeAgent.cs, Assets/Scripts/Utilities/EnvInitializer.cs, Assets/ML-Agents/Configs/cube_ppo.yaml.
- What: Menu items that autogenerate scenes—no .unity files needed.
- Why: Zero scene boilerplate; easy to spawn many parallel arenas.
- Use: Unity Tools → ML Hack → …
- Where: Assets/Editor/SceneBuilder.cs.
- What: Tuned YAML and shell/PowerShell runners.
- Why: One command to train; sensible PPO defaults.
- Use: Tools/train_cube.sh or the mlagents-learn … command above.
- Where: Assets/ML-Agents/Configs/*.yaml, Tools/.
- What: Runner learns to avoid a chaser (tagger). Tagger can be heuristic or trainable.
- Why: More “gamey” objective and a great demo.
- Use (heuristic tagger): Tools → ML Hack → Build Tag Arena Scene → Train Runner:
mlagents-learn Assets/ML-Agents/Configs/runner_ppo.yaml --run-id=runner --train
- Where: Assets/Scripts/Tag/RunnerAgent.cs, HeuristicTagger.cs, TagArenaBuilder.cs, Assets/ML-Agents/Configs/runner_ppo.yaml.
- What: Runner and Tagger are PPO agents co-trained in one run.
- Why: Showcases multi-agent interaction; richer behaviors.
- Use: Tools → ML Hack → Build Self-Play Tag Scene →
mlagents-learn Assets/ML-Agents/Configs/selfplay_combo.yaml --run-id=tag_selfplay --train
- Where: Assets/Scripts/Tag/TaggerAgentTrainable.cs, Assets/Editor/SelfPlayTagSceneBuilder.cs, Assets/ML-Agents/Configs/selfplay_combo.yaml.
- What: Skeleton joints with PD controllers; agent learns to stand/walk.
- Why: Embodied control—exactly the “NPC physics brain” you want.
- Use: Build a ragdoll with ConfigurableJoints, add PDJointController per joint, wire into RagdollAgent. Train:
mlagents-learn Assets/ML-Agents/Configs/ragdoll_ppo.yaml --run-id=ragdoll --train
- Where: Assets/Scripts/Ragdoll/PDJointController.cs, RagdollAgent.cs, Assets/ML-Agents/Configs/ragdoll_ppo.yaml.
- What: Runtime difficulty scaler (arena size, obstacles, tagger speed).
- Why: Start easy → get wins → scale up; faster, stabler learning.
- Use: Scenes with the add-on UI include a Difficulty slider.
- Where: Assets/Scripts/Utilities/CurriculumController.cs.
- What: Toggle behavior live for A/B demos.
- Why: Instantly compare learned vs baseline behaviors.
- Use: Assign your exported .onnx to ModelSwitcher.trainedModel, pick mode in the in-scene dropdown.
- Where: Assets/Scripts/UI/ModelSwitcher.cs.
- What: On-screen bars for survival/energy/etc.
- Why: Makes reward shaping visible during play.
- Use: Already wired in Runner; call RewardHUD.SetReward("Name", value).
- Where: Assets/Scripts/UI/RewardHUD.cs.
- What: Toggle and re-roll mass, friction, lighting, gravity.
- Why: Improves robustness/generalization; cool demo lever.
- Use: In scenes with the panel, click Apply Randomization (ideally at episode start).
- Where: Assets/Scripts/DomainRandomization/DomainRandomizer.cs, DomainRandomizationUI.cs.
- What: Press R in Play mode to save PNG frames; scripts convert to MP4/GIF via ffmpeg.
- Why: Quick sizzle reels for your hackathon demo.
- Use: Add SimpleRecorder to your Camera → press R → then:
- Windows:
Tools/make_gif.ps1 -InputDir Recordings -OutMp4 out.mp4 -OutGif out.gif -Fps 30 - macOS/Linux:
Tools/make_gif.sh Recordings out.mp4 out.gif 30
- Windows:
- Where: Assets/Scripts/Recording/SimpleRecorder.cs, Tools/make_gif.*.
- What: In-Unity mini charts for key metrics (Cumulative Reward, Loss, Entropy).
- Why: Keep an eye on learning without tab-switching.
- Use: Start TensorBoard:
tensorboard --logdir results --port 6006→ Tools → ML Hack → Build TensorBoard Dashboard → set Run/Tags → Refresh or Auto. - Where: Assets/Editor/TensorBoardDashboardBuilder.cs, Assets/Scripts/Dashboard/*.
- What: Pull newest .onnx from results/ into Assets/…/latest.onnx and auto-assign to all ModelSwitchers.
- Why: Swap in the latest policy mid-iteration without fiddling.
- Use: Tools → ML Hack → Model Hot-Reload → Import Newest .onnx → Assign To All ModelSwitchers.
- Where: Assets/Editor/ModelHotReloadWindow.cs.
- Start tiny (Cube PPO) to grok observations/actions/rewards → you’ll learn PPO knobs with quick feedback.
- Game-ify with Tag Mini-Game to get a flashy, relatable demo and practice curriculum & reward shaping.
- Scale up to Self-Play for emergent tactics and to Ragdoll for embodied control you can reuse for NPCs.
- Polish with Domain Randomization (robustness), Reward HUD (explainability), and Recording (shareability).
- Iterate fast using the TensorBoard overlay + Model Hot-Reload to avoid context switching.
# Cube (Core)
mlagents-learn Assets/ML-Agents/Configs/cube\_ppo.yaml --run-id=cube --train
# Tag Runner (heuristic tagger)
mlagents-learn Assets/ML-Agents/Configs/runner\_ppo.yaml --run-id=runner --train
# Self-play (Runner + Tagger)
mlagents-learn Assets/ML-Agents/Configs/selfplay\_combo.yaml --run-id=tag\_selfplay --train
# Ragdoll locomotion
mlagents-learn Assets/ML-Agents/Configs/ragdoll\_ppo.yaml --run-id=ragdoll --train
# Visualize learning
tensorboard --logdir results
- It’s learning too slowly → add more parallel arenas; increase entropy (beta), simplify rewards; start with a small arena & no obstacles (curriculum).
- It’s unstable → reduce action magnitudes; add energy penalties; clamp torques; lower time scale a bit; verify observations (no NaNs / huge ranges).
- Ragdoll keeps falling → raise PD gains gradually; larger feet; reward uprightness more; add “stand first” curriculum.
- Policy looks good in training but bad in demo → use Domain Randomization during training; normalize observations; check that inference scene matches training setup (friction, masses, timescale).
- Assets/Scripts/*.cs, Assets/Scripts/Utilities/*.cs, Assets/Editor/SceneBuilder.cs, Assets/ML-Agents/Configs/*.yaml, Tools/
- Assets/Scripts/Tag/*, Assets/Scripts/Ragdoll/*, Assets/Scripts/UI/*,
- Assets/Scripts/DomainRandomization/*, Assets/Scripts/Recording/*,
- Assets/Scripts/Dashboard/*, Assets/Editor/* (additional builders & windows), Tools/make_gif.*
- Day 1: Cube PPO up, TensorBoard running → first successful runs.
- Day 2: Tag Mini-Game + curriculum + reward HUD.
- Day 3–4: Ragdoll stand → walk; PD tuning; domain randomization.
- Day 5: Self-play Tag variant; compare to heuristic tagger.
- Day 6: Inference polish, Model Switcher, recordings.
- Day 7: Sizzle reel + tidy docs.