Thanks for the great work and for open-sourcing the checkpoints! Inspecting LDA-robocasa.pt directly via torch.load:
|
Paper Table V (LDA-1B) |
Released LDA-robocasa.pt |
| Hidden Size |
1536 |
2560 |
| Layers |
16 |
16 |
| Heads |
32 |
32 |
| Action Chunk |
16 |
16 |
MMDiT trunk (action_model.model.blocks) |
~0.45 B (est.) |
1.77 B |
action_model total trainable |
~1 B (est.) |
2.38 B |
| Qwen3-VL (frozen at pretrain) |
~4.83 B |
4.83 B |
| state_dict total (bf16) |
— |
~7.2 B / 14 GB |
Released ckpt matches Table V on depth, heads, and action chunk, but uses hidden_size = 2560 instead of 1536 — ~2.4× larger trainable params than Table V's LDA-1B.
Sim eval result
Running batch_eval_args.sh with default settings (24 RoboCasa-GR1 tasks, 50 trials/task, MAX_EPISODE_STEPS=720, N_ACTION_STEPS=12 — paper says 51 trials, essentially the same protocol):
- Released ckpt avg SR: 49.92 %
- Paper Table VI LDA-1B: 55.4 %
Most per-task SRs are within ±5 pp of paper, with one big outlier: PnPBottleToCabinetClose (paper 76 → ours 54).
So the released ckpt is wider than Table V but ~5.5 pp lower than Table VI on the same task suite — the opposite of what scaling would predict.
Questions
- Is
Wayer2/LDA-robocasa.pt the exact checkpoint behind paper Table VI's LDA-1B = 55.4 %, or a later scale-up (1536 → 2560)?
- Either way, any insight on the 5.5 pp gap? We've considered RoboCasa / robosuite version drift, asset issues (
TrayToTieredshelf 0.14 / PlacematToTieredshelf 0.12 look like MuJoCo penetration), and eval hyperparams (MAX_EPISODE_STEPS / N_ACTION_STEPS) not specified in the paper.
Happy to share more diagnostic info (full per-task SR table, state_dict stats) if useful. Thanks!
Thanks for the great work and for open-sourcing the checkpoints! Inspecting
LDA-robocasa.ptdirectly viatorch.load:LDA-robocasa.ptaction_model.model.blocks)action_modeltotal trainableReleased ckpt matches Table V on depth, heads, and action chunk, but uses
hidden_size = 2560instead of1536— ~2.4× larger trainable params than Table V's LDA-1B.Sim eval result
Running
batch_eval_args.shwith default settings (24 RoboCasa-GR1 tasks, 50 trials/task,MAX_EPISODE_STEPS=720,N_ACTION_STEPS=12— paper says 51 trials, essentially the same protocol):Most per-task SRs are within ±5 pp of paper, with one big outlier:
PnPBottleToCabinetClose(paper 76 → ours 54).So the released ckpt is wider than Table V but ~5.5 pp lower than Table VI on the same task suite — the opposite of what scaling would predict.
Questions
Wayer2/LDA-robocasa.ptthe exact checkpoint behind paper Table VI'sLDA-1B = 55.4 %, or a later scale-up (1536 → 2560)?TrayToTieredshelf0.14 /PlacematToTieredshelf0.12 look like MuJoCo penetration), and eval hyperparams (MAX_EPISODE_STEPS/N_ACTION_STEPS) not specified in the paper.Happy to share more diagnostic info (full per-task SR table, state_dict stats) if useful. Thanks!