Skip to content

Repository files navigation

AT-VLA (CVPR 2026 Oral)

Official implementation for AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models.

AT-VLA builds on a VLA backbone (GO-1 / InternVL-style pipeline in this repo) and adds adaptive tactile injection plus a tactile reaction dual-stream design so tactile is used when it helps, with a fast path for high-frequency contact feedback.

Installation

Environment setup is scripted in install.sh. Typical steps include:

  1. Python 3.10 (recommended via conda, see commented lines at the top of install.sh)
conda create -n at_vla python=3.10
conda activate at_vla
  1. pip install -U pip setuptools wheel
  2. Editable install of this package: pip install -e .
  3. AgiBot GO1 IK: cd ik_solver && pip install -e .
  4. flash-attn: often built from source or installed from a wheel matched to your PyTorch + CUDA build; see the FlashAttention repository if pip install flash-attn fails.

Core library versions are also declared in pyproject.toml (e.g. torch==2.6.0, torchvision==0.21.0).

Open-loop test (example)

An example entry point. Before running this script, you need to follow the below instructions to place our released checkpoint and test example. Or change the directory in ./internvl_chat/deploy/deploy.py to your own path:

bash open_loop_test.sh

Checkpoint directory

Place the downloaded checkpoint under checkpoints folder

Checkpoint layout:

<repo_root>/
  checkpoints/
    <task_name>/                   # task name, e.g., wipe vase
      checkpoint-<step>/          # e.g. checkpoint-3600 — this path is model_name_or_path
        config.json               # Pi0 / model config (and any sidecar configs your export uses)
        model.safetensors.index.json   # optional but used to auto-detect depth / GEM2 / SigLIP branches
        model-00001-of-000xx.safetensors
        model-00002-of-000xx.safetensors
        ...                       # other shards as produced by your save
        tokenizer_config.json     # and other tokenizer files required by AutoTokenizer.from_pretrained
        ...

Open-loop debug data (debug=True)

Place the downloaed data under test_example folder

Example layout (matches the defaults in main()):

<repo_root>/
  test_example/
    <task_name>/                 # task name, e.g., wipe vase
      input.pt
      gt.pt

Repository scope

This release focuses on inference-oriented code and utilities needed to understand and extend the model.

We do not open-source full training scripts or real-robot evaluation harnesses: all experiments in the paper were conducted on physical hardware, and we are not distributing those training / on-robot testing scripts at this time.

Naming: pi0 / Pi0 in code and checkpoints

You will see many Python modules, class names, and state_dict keys prefixed with pi0, Pi0, or similar. This is legacy naming from the upstream AgiBot / InternVL–style VLA codebase (diffusion-style action head and file layout), kept so that checkpoint key paths stay compatible with existing weights and loading logic.

This repository is not affiliated with, and does not ship, Physical Intelligence’s π₀ (Pi0) VLA. AT-VLA is our method on top of a GO-1 VLA; the pi0 strings are a naming artifact, not a claim of using or redistributing that model.

Citation

If you use this code or the method, please cite:

@article{li2026atvla,
  title={{AT-VLA}: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models},
  author={Li, Xiaoqi and Cai, Muhe and Xu, Jiadong and Zhu, Juan and Fan, Hongwei and Shen, Yan and Ren, Guangrui and Dong, Hao},
  journal={arXiv preprint arXiv:2605.07308},
  year={2026}
}

(Update the BibTeX entry with the official CVPR proceedings citation once available.)

License

See LICENSE.

About

The code for paper AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models

Resources

Stars

21 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages