Skip to content

Repository files navigation

μVLA: On Recurrent Memory for Partially Observable Manipulation in VLA Models

arXiv Project page Checkpoints License

μVLA visual abstract

μVLA extends OpenVLA-OFT with a recurrent memory mechanism for partially observable manipulation. A small set of learnable memory tokens is carried across timesteps within an episode, so the policy can condition on observations that are no longer visible in the current frame.

The repository supports two benchmarks:

  • MIKASA-Robo — memory-focused manipulation tasks, deliberately partially observable (MIKASA-Robo).
  • LIBERO — the standard four-suite manipulation benchmark, used here with an episodic dataloader so the memory path can be checked against a familiar baseline.

Method

μVLA method overview

A small bank of learnable memory tokens is carried across timesteps inside the backbone self-attention and updated end-to-end with TBPTT — no auxiliary losses, no architectural additions.

Attention mask with the memory-action guard

μVLA attention mask with the memory-action guard

Memory tokens attend only to observations, proprioception, language, and previous memory state, but cannot read action tokens. This prevents the recurrent state from trivially copying demonstrated actions and encourages encoding of task-relevant observations instead.

What is different from OpenVLA-OFT

OpenVLA-OFT μVLA
Observation model MDP (single frame + wrist) POMDP (episodic, memory carried across steps)
Extra parameters — MemoryModule: learnable memory tokens in the multimodal prefix
Memory update — TBPTT truncation, or EMA (M_in[t+1] = a * M_out[t] + (1-a) * M_in[t])
Dataloader frame-shuffled RLDS episodic (is_first / is_last), sequential within an episode
Attention causal over the prefix memory-aware mask (--attention_mask_mode custom)

Memory is off by default; --use_memory False reproduces the OpenVLA-OFT baseline through the same code path, which is how the baselines in both benchmark pages are produced.

Documentation

Read these in order. Each benchmark page is self-contained past the shared install: it covers its own data, supervised fine-tuning, and in-environment evaluation.

Page Contents
SETUP.md Install with uv (including the required transformers fork), download the openvla-7b base model, logging
MIKASA.md MIKASA-Robo: dataset, training with and without memory, the evaluation protocol
LIBERO.md LIBERO: the extra LIBERO install steps, dataset, combined-suite training, four-suite evaluation

Shortest path from a clean machine to a trained-and-evaluated policy:

git clone https://github.com/CognitiveAISystems/muVLA.git && cd muVLA
uv sync --python 3.10                       # SETUP.md
# then follow MIKASA.md (sections 2, 3, 4) or LIBERO.md (sections 1, 2, 3, 4)

System requirements

Inference:

  • 1 GPU with ~16 GB VRAM, for either benchmark

Training:

  • 1-8 GPUs with 40-80 GB (bfloat16). Runs with 64 memory tokens and --use_gradient_checkpointing True fit on a single 80 GB card at --batch_size 4.

Checkpoints

Fine-tuned μVLA checkpoints are on the Hugging Face Hub under the mu-vla organization. All are step-150000 runs. The three memory checkpoints use 64 memory tokens and cosine learning-rate scheduling and differ in the benchmark and in the TBPTT truncation length; the fourth is the memoryless ablation described below.

Checkpoint Benchmark Memory TBPTT Training recipe
mu-vla-openvla-oft-mikasa-robo-5-tasks-m64-k8-tbptt MIKASA-Robo, 5 tasks 64 tokens 8 MIKASA.md
mu-vla-openvla-oft-mikasa-robo-5-tasks-m64-k2-tbptt MIKASA-Robo, 5 tasks 64 tokens 2 MIKASA.md
mu-vla-openvla-oft-libero-4-tasks-m64-k8-tbptt LIBERO, 4 suites 64 tokens 8 LIBERO.md
mu-vla-openvla-oft-mikasa-robo-5-tasks-no-memory MIKASA-Robo, 5 tasks none — MIKASA.md

The last row is the memoryless ablation: plain OpenVLA-OFT, no memory tokens and no recurrent state, fine-tuned on the same five MIKASA-Robo-VLA tasks with the same episodic dataloader for the same 150000 steps, so that any gain from the memory checkpoints can be attributed to the memory module rather than to the data or the task mixture. It uses the original OpenVLA-OFT constant-then-decay learning-rate schedule instead of the cosine one.

Each repository is a complete evaluation-ready checkpoint directory: the LoRA-merged model as sharded safetensors, the LoRA adapter on its own under lora_adapter/, the action head, the proprio projector, the memory module, dataset_statistics.json (needed to unnormalize actions) and memory_meta.json (the memory hyperparameters, auto-detected at evaluation time). Roughly 16 GB each. The memoryless checkpoint ships the same layout minus the memory module and memory_meta.json, and is loaded as plain OpenVLA-OFT.

export CKPT_DIR="$PWD/my_checkpoints/mu-vla-mikasa-5-m64-k8"

uv run python -c "
from huggingface_hub import snapshot_download
snapshot_download('mu-vla/mu-vla-openvla-oft-mikasa-robo-5-tasks-m64-k8-tbptt',
                  local_dir='$CKPT_DIR')
"

The result is a directory that the evaluation commands in the benchmark pages accept directly as --checkpoint. Loading it needs the transformers fork pinned in pyproject.toml, since upstream transformers cannot build the memory-augmented backbone, see SETUP.md.

Both benchmark pages also contain the exact training configurations the reported numbers were produced with, so the checkpoints can be reproduced from the base openvla/openvla-7b model.

Acknowledgements

This repository builds directly on OpenVLA-OFT by Moo Jin Kim, Chelsea Finn and Percy Liang (MIT-licensed), see LICENSE. If you use μVLA, please also cite OpenVLA-OFT (arXiv:2502.19645).

Citation

@article{cherepanov2026muvla,
  title={{$\mu$}VLA: On Recurrent Memory for Partially Observable Manipulation in VLA Models},
  author={Cherepanov, Egor and Kachaev, Nikita and Zelezetsky, Daniil and Bulatov, Aydar and Pshenitsyn, Artem and Kuratov, Yuri and Skrynnik, Alexey and Panov, Aleksandr I and Kovalev, Alexey K},
  journal={arXiv preprint arXiv:2606.12497},
  year={2026}
}

Releases

Packages

Contributors

Languages