A Technical University of Munich arXiv preprint from September 7, 2026, argues that language belongs at a multimodal model's boundary, not as its internal representation. Updated to v3 on September 15, Peng Xie and Amr Alanwar frame language as a shared community codebook and test the claim with cue-conflict experiments on six vision-language models and two robot policies.
The paper treats a word as an index into a lossy, community-maintained codebook whose content lives in the receiver. On that reading, putting the codebook inside a model buys fluent internal reasoning at the price of auditability: once language is the hidden state, outside observers cannot easily check which cue the system used.
In plain terms, the "seat of language" question is not whether a model can chat. It is whether text should be only the interface and shared vocabulary, while perception and action keep richer nonlinguistic state, or whether the model should think in tokens the way today's chat systems do. The preprint says the boundary seat matches how brains use language; keeping language inside is a capacity choice with an auditability cost.
Against a bounded-uniform ideal observer that treats a stated text interval as a uniform likelihood, the six vision-language models raised image weight as image reliability rose, but their slopes stayed at 0.11 to 0.82 of ideal. Between 39% and 78% of answers copied the stated interval midpoint exactly. A separate Gaussian ideal in the same figure reaches slopes as high as 0.99; those two slope definitions are not interchangeable.
In a color-conflict test on recolored COCO images, naming a noun with a canonical color cut reporting of the image color by 0.04 to 0.09 on clean images for five of six models, though models still reported the counterfactual image color on 64% to 79% of clean shots. Those color results are model-free contrasts: the authors' Bayesian benchmark failed a calibration check. On simulated LIBERO robot policies, OpenVLA-OFT completed swapped instructed tasks on 97% of LIBERO-Goal episodes but only 0% to 1% on LIBERO-Object, while π0.5 on LIBERO-Object split 28% instructed, 30% scene, and 42% neither across 21,350 OpenVLA-OFT episodes plus the second family. A task-predictive visual tag added in 15,005 fine-tuning steps was not learned.
Two related arXiv papers supply context, not a replication of these cue-conflict numbers. Song, Lepori, and Pavlick report that goal-directed language can causally recode visual representations in vision-language models. Paruchuri and coauthors find that replacing post-image text costs about four times more accuracy than visual replacement across seven multimodal models on six BLINK tasks. What remains open is whether the TUM slopes and robot splits hold under new prompts, non-Qwen backbones, and physical robots: four of the six tested vision-language models share a Qwen language backbone, measured weights shift with prompt wording, and the robot evidence is simulation only. Preprint claims and the independent contextual papers should be weighed separately until those points are settled.