Do vision-language-action models act on the instructions they hear?
π0.5 hears the instruction. What it hears does not decide what it does.
Benchmarks such as LIBERO-Plus, LIBERO-CF and LangGap show that vision-language-action models often act on visual shortcuts instead of language. Behaviour alone cannot tell whether a model failed to hear an instruction or heard it and acted on something else.
We give OpenVLA-7B and π0.5 the same image, robot state and noise under two instructions that name different objects, so anything that differs between the two forward passes is caused by language. Residuals are read after every block with hooks that never modify the computation (hooked and unhooked actions agree to 0.0), and goal probes are refit so that they never see the initial states they are evaluated on.
Absolute object and robot positions are equally readable with or without language; they are in the image. The target's position is not. It becomes linearly readable only after the instruction is integrated, in both models, and it survives into π0.5's action expert. On object pairs never seen by the probe, only the action expert's goal-relative code transfers.
Given the other object's name, unmodified π0.5 obeys 6%, 98%, 100% and 46% of the time across four object pairs. The two failing pairs obey 100% and 98% of the time in the opposite scene, so the model can read both names; it fails on particular instruction–scene combinations.
Before the arm moves, the swapped instruction pulls the decoded goal toward the named object in every condition, including the rollouts where the robot then grasps the wrong one: 40 of 40 in the condition that almost never obeys. How far the goal is pulled does not predict obedience; across six conditions the relation runs opposite to the predicted sign (ρ = −0.97), and with a probe shared across scenes it flips.
The preregistered first-query test (Stage 12) is inconclusive: 23/29 with the primary probe against 10/30 with a leave-pair-out probe. The six-condition analysis (Stage 13) is exploratory, with its analysis fixed before the data were read.
If the goal is represented, can it be written back in? Patching the prompt-induced residual at one expert block moves the action locally, then breaks the closed-loop rollout; the probe's own subspace carries under 1% of the effect. Replaying the whole prompt-conditioned expert pathway reproduces the correct policy exactly. Control works only through the model's own computation path.
Every stage had its question, endpoint and stopping rule written down before it ran. Negative and inconclusive results stay in the record.
@misc{silentl2026,
title = {The Silent L? Paired-Prompt Probing of Whether
Vision-Language-Action Models Hear, Obey, and Can Be
Written with Language},
author = {Lai, Erhan},
year = {2026},
howpublished = {\url{https://github.com/RyleHan/silent-L}}
}