μVLA extends OpenVLA-OFT with a recurrent memory mechanism for partially observable manipulation. A small set of learnable memory tokens is carried across timesteps within an episode, so the policy can condition on observations that are no longer visible in the current frame.
The repository supports two benchmarks:
- MIKASA-Robo — memory-focused manipulation tasks, deliberately partially observable (MIKASA-Robo).
- LIBERO — the standard four-suite manipulation benchmark, used here with an episodic dataloader so the memory path can be checked against a familiar baseline.
A small bank of learnable memory tokens is carried across timesteps inside the backbone self-attention and updated end-to-end with TBPTT — no auxiliary losses, no architectural additions.
Memory tokens attend only to observations, proprioception, language, and previous memory state, but cannot read action tokens. This prevents the recurrent state from trivially copying demonstrated actions and encourages encoding of task-relevant observations instead.
| OpenVLA-OFT | μVLA | |
|---|---|---|
| Observation model | MDP (single frame + wrist) | POMDP (episodic, memory carried across steps) |
| Extra parameters | — | MemoryModule: learnable memory tokens in the multimodal prefix |
| Memory update | — | TBPTT truncation, or EMA (M_in[t+1] = a * M_out[t] + (1-a) * M_in[t]) |
| Dataloader | frame-shuffled RLDS | episodic (is_first / is_last), sequential within an episode |
| Attention | causal over the prefix | memory-aware mask (--attention_mask_mode custom) |
Memory is off by default; --use_memory False reproduces the OpenVLA-OFT baseline through
the same code path, which is how the baselines in both benchmark pages are produced.
Read these in order. Each benchmark page is self-contained past the shared install: it covers its own data, supervised fine-tuning, and in-environment evaluation.
| Page | Contents |
|---|---|
| SETUP.md | Install with uv (including the required transformers fork), download the openvla-7b base model, logging |
| MIKASA.md | MIKASA-Robo: dataset, training with and without memory, the evaluation protocol |
| LIBERO.md | LIBERO: the extra LIBERO install steps, dataset, combined-suite training, four-suite evaluation |
Shortest path from a clean machine to a trained-and-evaluated policy:
git clone https://github.com/CognitiveAISystems/muVLA.git && cd muVLA
uv sync --python 3.10 # SETUP.md
# then follow MIKASA.md (sections 2, 3, 4) or LIBERO.md (sections 1, 2, 3, 4)Inference:
- 1 GPU with ~16 GB VRAM, for either benchmark
Training:
- 1-8 GPUs with 40-80 GB (bfloat16). Runs with 64 memory tokens and
--use_gradient_checkpointing Truefit on a single 80 GB card at--batch_size 4.
Fine-tuned μVLA checkpoints are on the Hugging Face Hub under the mu-vla organization. All are step-150000 runs. The three memory checkpoints use 64 memory tokens and cosine learning-rate scheduling and differ in the benchmark and in the TBPTT truncation length; the fourth is the memoryless ablation described below.
| Checkpoint | Benchmark | Memory | TBPTT | Training recipe |
|---|---|---|---|---|
mu-vla-openvla-oft-mikasa-robo-5-tasks-m64-k8-tbptt |
MIKASA-Robo, 5 tasks | 64 tokens | 8 | MIKASA.md |
mu-vla-openvla-oft-mikasa-robo-5-tasks-m64-k2-tbptt |
MIKASA-Robo, 5 tasks | 64 tokens | 2 | MIKASA.md |
mu-vla-openvla-oft-libero-4-tasks-m64-k8-tbptt |
LIBERO, 4 suites | 64 tokens | 8 | LIBERO.md |
mu-vla-openvla-oft-mikasa-robo-5-tasks-no-memory |
MIKASA-Robo, 5 tasks | none | — | MIKASA.md |
The last row is the memoryless ablation: plain OpenVLA-OFT, no memory tokens and no recurrent state, fine-tuned on the same five MIKASA-Robo-VLA tasks with the same episodic dataloader for the same 150000 steps, so that any gain from the memory checkpoints can be attributed to the memory module rather than to the data or the task mixture. It uses the original OpenVLA-OFT constant-then-decay learning-rate schedule instead of the cosine one.
Each repository is a complete evaluation-ready checkpoint directory: the LoRA-merged model
as sharded safetensors, the LoRA adapter on its own under lora_adapter/, the action head,
the proprio projector, the memory module, dataset_statistics.json (needed to unnormalize
actions) and memory_meta.json (the memory hyperparameters, auto-detected at evaluation
time). Roughly 16 GB each. The memoryless checkpoint ships the same layout minus the memory
module and memory_meta.json, and is loaded as plain OpenVLA-OFT.
export CKPT_DIR="$PWD/my_checkpoints/mu-vla-mikasa-5-m64-k8"
uv run python -c "
from huggingface_hub import snapshot_download
snapshot_download('mu-vla/mu-vla-openvla-oft-mikasa-robo-5-tasks-m64-k8-tbptt',
local_dir='$CKPT_DIR')
"The result is a directory that the evaluation commands in the benchmark pages accept
directly as --checkpoint. Loading it needs the transformers fork pinned in
pyproject.toml, since upstream transformers cannot build the memory-augmented backbone,
see SETUP.md.
Both benchmark pages also contain the exact training configurations the reported numbers
were produced with, so the checkpoints can be reproduced from the base
openvla/openvla-7b model.
This repository builds directly on OpenVLA-OFT by Moo Jin Kim, Chelsea Finn and Percy Liang (MIT-licensed), see LICENSE. If you use μVLA, please also cite OpenVLA-OFT (arXiv:2502.19645).
@article{cherepanov2026muvla,
title={{$\mu$}VLA: On Recurrent Memory for Partially Observable Manipulation in VLA Models},
author={Cherepanov, Egor and Kachaev, Nikita and Zelezetsky, Daniil and Bulatov, Aydar and Pshenitsyn, Artem and Kuratov, Yuri and Skrynnik, Alexey and Panov, Aleksandr I and Kovalev, Alexey K},
journal={arXiv preprint arXiv:2606.12497},
year={2026}
}

