Research · 2026 · OpenVLA and π0.5 on LIBERO

The Silent L?

Do vision-language-action models act on the instructions they hear?

TL;DR

π0.5 hears the instruction. What it hears does not decide what it does.

Figure 1. Same scene, two instructions. Both rollouts start from the same simulator state with the same flow-matching noise; only the instruction differs. The crosshair marks where a linear probe on π0.5's action expert places the goal at each policy query, and the strip below tracks it between the two objects. Choose a condition or an initial state to compare.

One exact counterfactual

Benchmarks such as LIBERO-Plus, LIBERO-CF and LangGap show that vision-language-action models often act on visual shortcuts instead of language. Behaviour alone cannot tell whether a model failed to hear an instruction or heard it and acted on something else.

We give OpenVLA-7B and π0.5 the same image, robot state and noise under two instructions that name different objects, so anything that differs between the two forward passes is caused by language. Residuals are read after every block with hooks that never modify the computation (hooked and unhooked actions agree to 0.0), and goal probes are refit so that they never see the initial states they are evaluated on.

Figure 2. The instruction arriving block by block. Probe readings of the goal after every PaliGemma prefix and action-expert block of π0.5, for the state shown in Figure 1, at the first policy query before any motion. Blue: the scene's usual object; clay: the other object. The outlined column is expert block 13, the preregistered readout.

01Heard

Absolute object and robot positions are equally readable with or without language; they are in the image. The target's position is not. It becomes linearly readable only after the instruction is integrated, in both models, and it survives into π0.5's action expert. On object pairs never seen by the probe, only the action expert's goal-relative code transfers.

Figure 3. Language creates a goal the image cannot. Held-out test R² of linear probes after every block, mean of three seeds; 400 rollouts, 15,275 states, four object pairs. The shaded band is what language adds over a prompt-blind control that sees the same image.

02Obeyed, sometimes

Given the other object's name, unmodified π0.5 obeys 6%, 98%, 100% and 46% of the time across four object pairs. The two failing pairs obey 100% and 98% of the time in the opposite scene, so the model can read both names; it fails on particular instruction–scene combinations.

03Heard is not obeyed

Before the arm moves, the swapped instruction pulls the decoded goal toward the named object in every condition, including the rollouts where the robot then grasps the wrong one: 40 of 40 in the condition that almost never obeys. How far the goal is pulled does not predict obedience; across six conditions the relation runs opposite to the predicted sign (ρ = −0.97), and with a probe shared across scenes it flips.

The preregistered first-query test (Stage 12) is inconclusive: 23/29 with the primary probe against 10/30 with a leave-pair-out probe. The six-condition analysis (Stage 13) is exploratory, with its analysis fixed before the data were read.

Figure 4. The goal moves; behaviour does not follow. Left: mean decoded goal under the scene's own instruction (blue) and the swapped one, for six scene × instruction conditions, beside how often the robot obeys. Click a row to load it into Figure 1. Right: every initial state of the selected condition.

04Readable is not writable

If the goal is represented, can it be written back in? Patching the prompt-induced residual at one expert block moves the action locally, then breaks the closed-loop rollout; the probe's own subspace carries under 1% of the effect. Replaying the whole prompt-conditioned expert pathway reproduces the correct policy exactly. Control works only through the model's own computation path.

Figure 5. Four rungs of intervention on π0.5. Same-state counterfactual prompts on LIBERO-Object; closed-loop results on 20 held-out initial states per target. A soft conceptor gate gives 40% → 53% on a separate task with a confidence interval that crosses zero.

The ledger

Every stage had its question, endpoint and stopping rule written down before it ran. Negative and inconclusive results stay in the record.

Citation

@misc{silentl2026,
  title        = {The Silent L? Paired-Prompt Probing of Whether
                  Vision-Language-Action Models Hear, Obey, and Can Be
                  Written with Language},
  author       = {Lai, Erhan},
  year         = {2026},
  howpublished = {\url{https://github.com/RyleHan/silent-L}}
}