Train Together or Merge Later? Unifying VLA Experts via a Shared Action Interface

Zhizhen Zhang1 Yuxia Fu1 Zijian Wang1 Zi Huang1 Yadan Luo1

1 UQMM Lab, The University of Queensland

Stack blocksAgileX Piper

PolicyWeave on four real-robot tasks. Selected successful rollouts, shown at original speed.

Overview

Independently post-trained VLA experts can lose performance when their parameters are merged. PolicyWeave preserves a shared action interface, including normalization, the action encoder, and the action decoder, during task-specific SFT or SFT then RL. At deployment, the initial scene and instruction guide sparse merging. The resulting parameters remain fixed throughout the episode.

INTERACTIVE EXPLAINER

Two correct experts.
A different result when merged.

Both experts produce the same action. Change their interface scales and mixing weights to see why merging their parameters can miss the target.

Try an example
Policy parameters
6.00×
Same scaleExpert B uses 8× the scale
50 : 50
All AAll B
SEPARATE ACTION MAPPINGS

Task-specific interfaces

Off target

Preparing the robot workspace…

Executed action2.042
Target deviation+104.2%

Averaging breaks the match between internal parameters and interface.

ONE COMMON ACTION MAPPING

Shared interface

On target

Preparing the robot workspace…

Executed action1.000
Target deviation0.0%

Shared scaling preserves the target in this example.

100%
FOLLOW THE PARAMETERS

Same actions, different parameters

For input 1, the hidden weight × output scale gives the action.

PolicyHidden weightOutput scaleAction
Expert A1.0001.0001.000
Expert B0.1676.0001.000
Merged0.5833.5002.042
Shared interface1.000 × 1.000 = 1.000
EXPLORE THE MIXTURE

What happens between the experts?

Task-specificSharedAction deviation (%)
Action deviation across expert mixturesAnalytical output deviation in a scalar illustration, not experimental results. Adjust the Expert mixture slider to explore. 0%60%120%All A50 : 50All B

Both endpoints are correct. Mixing incompatible parameters introduces error between them.

Kinematic illustration · Not a policy evaluation or recorded rollout. Target action = 1.

Drag to orbit · Arrow keys to rotate
Output-scaling model

For a fixed input of 1, a two-parameter model produces action a = d · w. The internal parameter w and output-interface scale d compensate for each other, so both experts remain correct.

Task-specific interfaces

Expert A: w = 1, d = 1

Expert B: w = 1/r, d = r

Let β be the weight assigned to expert B. Merge parameters, not actions:

a = (1 − β + β/r) × (1 − β + βr)≈ 0.583 × 3.500 ≈ 2.042

Shared action interface

Both experts: w = 1, d = 1

Keep the interface fixed and combine the hidden parameters:

a = [(1 − β) × 1 + β × 1] × 1 = 1

This isolates interface mismatch, not interference between different tasks. PolicyWeave preserves normalization, the action encoder, and the action decoder. See the VLA experiment below.

Evidence from closed-loop control

In this RoboCasa365 Counter → Stove scene, base drift at step 160 is 54.1 cm with task-specific interfaces, versus <0.1 cm with a shared interface. Both policies use weight averaging.

RoboCasa365 action-interface diagnostics. Task-specific and shared normalization give different representations of the same raw actions, with the largest shifts in base-motion channels. On Counter to Stove, task-specific-interface weight averaging drifts 54.1 centimeters by step 160, while shared-interface weight averaging stays within 0.1 centimeters.
(a,b) Normalization differences concentrate in base-motion channels. (c) Base drift during execution. WA: weight averaging.
Experiment setup

Panels (a,b) apply task-specific and shared normalization to the same 1,800 raw actions across 18 tasks. Panel (c) compares policies from matched scenes with paired denoising noise; the images show step 160.

Method

Learn task-specific LoRA updates with a shared action interface. Then use expert context responses and ranking stability to select one expert or merge a sparse set.

PolicyWeave overview: frozen VLM and shared action interface during LoRA adaptation, followed by context-response scoring and stability-based sparse merging.
Shared-interface post-training (left) and context-guided sparse merging (right).
Scoring and deployment details

Experts are adapted using SFT or SFT then RL. A frozen VLM encodes the initial scene and instruction. Normalized cross-attention key-update responses score expert relevance, and leave-one-layer-out ranking stability determines how many updates to merge. The merged parameters stay fixed throughout the episode. No separately trained router or calibration trajectories are required.

Evaluation

GR00T N1.5 · 18 atomic tasks · 10 long-horizon tasks · 4 real-world tasks

64.7→74.1%

PolicyWeave · SFT → SFT then RL

Merging after task-specific RL

Task-specific RL improves the merged policy; joint RL raises co-training from 58.1% to 60.7%.

Full comparison and evaluation setup
Mean success over 18 tasks · 50 episodes per task
MethodSFTSFT then RL
Individual task identity supplied65.6%76.2%
Co-training58.1%60.7%
Weight averaging49.6%52.3%
TIES52.9%54.2%
ISO-CTS54.1%56.3%
TSV-Merge54.7%61.4%
SSR-Merge58.9%62.8%
PolicyWeave64.7%74.1%

Merging methods use the same shared-interface expert bank within each training stage. Each task-specific SFT expert uses 50 demonstrations, or 10% of the available task data.