BridgeVLA++

BridgeVLA++

A Data-Efficient, Generalizable, and Memory-Augmented
Vision-Language-Action Framework for 3D Manipulation

*Equal contribution  ·  Corresponding author
1Institute of Automation, Chinese Academy of Sciences 2University of Chinese Academy of Sciences
3ByteDance Seed 4FiveAges 5Nanjing University

TL;DR

BridgeVLA++ achieves state of the art across five manipulation benchmarks with only 9.2% additional parameters, adding powerful spatio-temporal memory while fully improving or preserving BridgeVLA's exceptional data efficiency and OOD generalization.

93.7%
RLBench (18 tasks)
+6.9 pts over the previous state of the art
65.2%
COLOSSEUM
14 evaluation settings (12 unseen perturbation axes); +8.5 pts over the previous best
51.1%
GemBench
Four generalization levels; +3.1 pts over the previous best
96.0%
RMBench
Nine dual-arm memory tasks; +13.0 pts over the previous best
99.7%
MemoryBench
Single-arm memory suite; +5.4 pts over the previous best
93.3%
Real memory tasks
Dobot CR5A, 10 demos per instruction; 20.0% without memory
Overview of BridgeVLA++: multi-view heatmap alignment plus a unified spatio-temporal memory, with benchmark radar chart and real-world generalization settings.
Figure 1. Overview of BridgeVLA++. BridgeVLA aligns 3D observations and actions in a shared multi-view 2D heatmap space for data-efficient, generalizable manipulation. BridgeVLA++ adds a unified spatio-temporal memory, injected through a shared module in the VLM patch-token space and shareable across two arms for bimanual manipulation.

Abstract

Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods remain data-hungry, generalize poorly to distribution shifts, and lack explicit memory of past observations. These limitations hinder their application to data-scarce, open-world and memory-dependent manipulation scenarios. In our previous work, BridgeVLA improves data efficiency and generalization by preserving the input–output alignment of a pre-trained VLM during 3D action learning: raw point clouds are projected into multi-view images, and intermediate heatmaps are predicted before generating robot actions.

In this work, we further extend BridgeVLA with a unified spatio-temporal memory architecture that models persistent spatial context and temporal interaction history. The resulting memory-augmented framework can reason over historical observations while preserving the hallmark data efficiency and strong generalization capability of BridgeVLA. Extensive experiments show that our framework establishes a strong foundation for spatial manipulation tasks and robust generalization (evaluating the base BridgeVLA on RLBench, COLOSSEUM and GemBench). When equipped with the unified spatio-temporal memory, BridgeVLA++ further achieves state-of-the-art performance on two challenging memory-dependent manipulation benchmarks without sacrificing the data efficiency and generalization of the original BridgeVLA. In addition, BridgeVLA++ extends effectively to bimanual manipulation and is validated on a new real-world robot embodiment, demonstrating its scalability across tasks, environments and robot embodiments. These results establish BridgeVLA++ as a unified 3D vision-language-action framework that simultaneously supports data-efficient learning, robust generalization and memory-aware robot manipulation.

One recipe, five suites. RMBench and MemoryBench are the memory-dependent ones BridgeVLA++ was built for; RLBench, COLOSSEUM and GemBench check that the base capabilities survive the memory extension. Pick a suite.

Nine dual-arm tasks that cannot be solved from the current frame alone — also the validation of the bimanual extension. BridgeVLA++ reaches a near-perfect 96.0%, while the memory-free base collapses to 18.9%.

Grid of the nine dual-arm RMBench tasks.
The RMBench suite. The nine dual-arm memory tasks, three example variants each.
Method Overall Avg. ↑ M(1) tasks M(n) tasks
Observe & Pick Up Rearrange BlocksPut Back BlockSwap BlocksSwap T Avg. Battery Try Blocks Ranking TryCover BlocksPress Button Avg.
DP [Chi et al. 2024]5.810011206.41010005.0
ACT [Zhao et al. 2023]5.91290226.8190004.8
π0.5 [Physical Intelligence 2025]10.491311241514.4166005.5
X-VLA [Zheng et al. 2025]9.89131816311.8261207.3
Mem-0 [Chen et al. 2026]42.048990671452.8281868028.5
Fast-WAM [Yuan et al. 2026]5.9000071.420260011.5
LingBot-VA [Li et al. 2026]78.213100100998880.041100798476.0
MemoryWAM [Yang et al. 2026]83.0271001001009484.241100988781.5
WLA-0 [Yang et al. 2026]4523847456.5
BridgeVLA (ours, base)18.9750111819.07203018.8
BridgeVLA++ (ours)96.081100100999695.296100999397.0

Table 1. RMBench. Success rate (%) over 100 episodes per task on the nine dual-arm tasks, grouped by memory complexity. Baseline numbers are quoted from their sources; “—” denotes tasks a source does not evaluate. Best per task in bold.

RMBench · short-term

Observe & Pick Up81%

RMBench · short-term

Rearrange Blocks100%

RMBench · short-term

Put Back Block100%

RMBench · short-term

Swap Blocks99%

RMBench · short-term

Swap T96%

RMBench · long-term

Battery Try96% (best baseline 45%)

RMBench · long-term

Blocks Ranking Try100%

RMBench · long-term

Cover Blocks99%

RMBench · long-term

Press Button93%

Real robot setup: Franka Research 3 and Dobot CR5A, each observed by a static ZED 2i stereo camera, two summary bar charts, and the basic, generalization and memory evaluation settings.
Figure 2. Real-robot evaluation setup. Two platforms under a single static ZED 2i stereo camera: a 7-DoF Franka Research 3 for general manipulation (one basic plus six generalization settings) and a held-out 6-DoF Dobot CR5A for memory. Centre: our average gain over the strongest prior method in each group.

Rollout videos

Every real-robot recording, up front. The browser below switches robot, task and setting on the spot; the written analysis further down the page has its own Dobot / Franka switch.

Browse every rollout

Every Dobot rollout: pick a task, then a setting, and see that exact clip with its success count.

Robot platform
Memory-dependent tasks
Memory-free control tasks
Evaluation setting
Basic
No recording for this combination

Success over 10 trials — this task, this setting
BridgeVLA++ across all settings
Keyframe rollout (basic setting)

Memory on a held-out embodiment

Memory is tested on a second, held-out arm with three tasks where the current observation underdetermines the next action — plus two ordinary pick-and-place families that check the memory costs nothing when it isn't needed. All methods train on the same 70 demonstrations.

Press Button · counting

Press the blue button exactly N times (given in words, “twice” / “three times”), then the yellow button once. A press leaves no visible trace, so every post-press frame is identical up to nuisance variation: the running count exists only in the policy's history. Neither baseline solves a single trial in any setting.

Cover Blocks · occlusion

Cover all three blocks, then uncover only the block of the instructed color. Covering erases the color evidence from the scene, so the second phase is solvable only from the colour-to-location bindings formed before the covers went on.

Swap Eggplant · rearrangement

Two look-alike eggplants must end up each on the other's initial plate, using a third plate as a buffer that must be empty again at the end. Intermediate configurations recur across phases, so the correct next placement depends on which eggplant has already moved.

MethodMem. Memory-dependent Memory-free
Cover Blocks Press ButtonSwap Eggplant Put in Drawer Put on Shelf
SAM2Act+ [Fang et al. 2025]20.0%0.0%70.0%60.0%20.0%
BridgeVLA (base)0.0%0.0%60.0%100.0%90.0%
BridgeVLA++100.0%100.0%80.0%100.0%100.0%

Table 6. Dobot, basic setting. Ten trials per language instruction. Put in Drawer and Put on Shelf are each evaluated with two instructions (upper and lower target) and reported as their average.

MethodMem.Basic Visual disturbance
Distractor BackgroundHeightLightingAvg.
Memory-dependent tasks
SAM2Act+ [Fang et al. 2025]30.0%0.0%0.0%0.0%3.3%0.8%
BridgeVLA (base)20.0%20.0%23.3%6.7%13.3%15.8%
BridgeVLA++93.3%73.3%86.7%76.7%76.7%78.3%
Memory-free tasks
SAM2Act+ [Fang et al. 2025]40.0%0.0%0.0%0.0%7.5%1.9%
BridgeVLA (base)95.0%57.5%72.5%67.5%67.5%66.3%
BridgeVLA++100.0%70.0%100.0%82.5%75.0%81.9%

Table 7. Dobot, all settings. Each entry is the unweighted mean over the three memory-dependent or two memory-free tasks; Avg. averages the four disturbance settings. The memory advantage survives every disturbance, and on tasks that never needed memory BridgeVLA++ is no weaker than the base model in any setting.

Part I · Base model

BridgeVLA — spatial input–output alignment

A pre-trained VLM is fluent in images and language, not in point clouds and 6-DoF poses. BridgeVLA closes the gap from both ends — the input is rendered back into images, the output is expressed as heatmaps in that same image space.

1

Render 3D back into 2D

Point clouds become three orthographic views, so the VLM keeps seeing the kind of images it was pre-trained on.

2

Pre-train on heatmaps

Before any robot data, the VLM learns to ground language as 2D spatial heatmaps.

3

Coarse-to-fine action

Per-view heatmaps vote on the next waypoint; a zoomed second pass sharpens it.

Part II · Memory

BridgeVLA++ — unified spatio-temporal memory

The coarse-to-fine design exposes exactly two places where the past helps: the coarse stage must know what has already happened; the fine stage must see geometry the arm is now hiding. Each gets its own memory — both in patch-token space, so the heatmap action interface never changes.

𝒯

Temporal memory → what to do next

At the coarse stage: the interaction history, with a lightweight gate that keeps only the keyframes worth remembering.

𝒮

Spatial memory → where exactly to act

At the fine stage: the initial, less-occluded point cloud, re-rendered under the current zoom as complementary geometry.

Additive & bimanual

Memory is injected without touching the action heads; one shared scene memory extends the framework to two arms.

BridgeVLA and BridgeVLA++ architecture: 2D-heatmap pre-training, 3D-action fine-tuning, and the temporal and spatial memory injection blocks. The temporal block runs cross- and self-attention over history keyframes and an anchor frame chosen by a keyframe selector; the spatial block attends to a local point-cloud projection of the initial scene.
Figure 4. Architecture. Top: BridgeVLA's 2D-heatmap pre-training and 3D-action fine-tuning. Bottom: the two memories of BridgeVLA++ — temporal at the coarse stage, spatial at the fine stage — injected in patch-token space without changing the action interface.

Base-model design ablations (RLBench)

VariantAvg. SR (%) ↑ Avg. Rank ↓ Insert PegPlace CupsPut in CupboardScrew BulbSort Shape Stack BlocksStack Cups
BridgeVLA (full base)90.54.7591.258.491.293.655.284.888.8
├ w/ discretized rotation88.24.8688.058.473.687.260.876.881.6
├ w/o heatmap decoding31.412.780.01.35.32.74.00.00.0
└ w/ 3D position input56.210.1426.714.710.716.021.317.34.0

Table 9. Base-model ablations. A representative subset of the 18 RLBench columns; the averages and ranks are over all 18 tasks. w/o heatmap decoding and w/ 3D position input use three seeds, the others five.

A

Predict heatmaps, not coordinates −59.1

Regressing target coordinates instead of decoding heatmaps collapses RLBench from 90.5% to 31.4%.

B

Don't feed 3D positions to the VLM −34.3

Fusing explicit per-pixel 3D positions shifts the image features off the pre-training distribution.

C

Pre-training carries language generalization

Without 2D-heatmap pre-training, the policy fails to generalize to unseen language in the real world.

D

Continuous 6D rotation beats discretization −2.3

The 6D head also avoids the gimbal lock of discretized Euler angles in near-vertical poses.

Memory ablations

Each memory is load-bearing on exactly the benchmark whose difficulty it targets.

VariantRLBench Avg. ↑ Insert PegPlace CupsSort ShapeStack CupsPlace Wine
BridgeVLA++ (full)93.799.276.872.098.495.2
├ w/o spatial memory 𝒮92.095.274.460.892.878.4
└ w/o temporal memory 𝒯91.982.457.673.693.690.4

Table 10. Memory ablations on RLBench. Removing 𝒮 costs most exactly in the occlusion-heavy precision tasks it targets (sort shape 72.0 → 60.8), and removing 𝒯 costs most where a stable global reference helps (place cups 76.8 → 57.6) — even though RLBench tasks are not intrinsically memory-dependent.

Variant Overall Avg. ↑ M(1) Avg. M(n) Avg. Rearrange BlocksCover BlocksPress Button
BridgeVLA++ (full)96.095.297.01009993
w/o 𝒮 (spatial memory)95.496.294.51009192
w/o 𝒯 (temporal memory)21.327.014.31150
BridgeVLA (no memory)18.919.018.8030

Table 11. Memory ablations on RMBench. The 2×2 memory factorial. Removing 𝒯 collapses the benchmark from 96.0% to 21.3%, essentially back to the memory-free base (18.9%); removing 𝒮 is nearly harmless here, because RMBench stresses temporal sequencing rather than geometric alignment.

+9.2%
Parameters
269.77M added on a 2.92B-parameter backbone
+0.22 s
Inference latency
0.35 s → 0.57 s per step on a single RTX 4090 — small next to observation transfer and motion execution
≈ 2 h
2D-heatmap pre-training
8 × A100, 3,800 steps — reused by every downstream benchmark
3 demos
Real-world data floor
95.4% on the 13 Franka tasks from three demonstrations per task
Temporal entries are stored already encoded, so history never costs an extra backbone forward — only the cache grows with the history length.
@misc{li2026bridgevlaplus,
  title         = {BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented
                   Vision-Language-Action Framework for 3D Manipulation},
  author        = {Peiyan Li and Yuze Zhu and Yixiang Chen and Qisen Ma and Yuan Xu
                   and Jiabing Yang and He Guan and Yan Huang and Hongtao Wu and Xiao Ma
                   and Tao Kong and Liang Wang and Tieniu Tan},
  year          = {2026},
  eprint        = {2608.05042},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url           = {https://arxiv.org/abs/2608.05042}
}

If you build on the conference version:

@misc{li2025bridgevla,
  title         = {BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation
                   Learning with Vision-Language Models},
  author        = {Peiyan Li and Yixiang Chen and Hongtao Wu and Xiao Ma and Xiangnan Wu
                   and Yan Huang and Liang Wang and Tao Kong and Tieniu Tan},
  year          = {2025},
  eprint        = {2506.07961},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO},
  url           = {https://arxiv.org/abs/2506.07961}
}