A team including Junlin Han, Shengbang Tong, David Fan, Minghao Chen, Philip Torr, Filippos Kokkinos and Mike Lewis posted an empirical study of natively unified multimodal pretraining to arXiv on 5 August 2026 (revised 6 August). Using controlled experiments on synthetic and large-scale real datasets, the authors report four findings: knowledge transfer between language, visual understanding and visual generation is asymmetric; data complexity largely determines whether modalities are synergistic or competitive, with shared attention and normalization plus modality-specific feed-forward layers promoting synergy across different visual tokenizer designs; unifying modalities early and training jointly beats late alignment or sequential training, with delayed integration producing a "vision laziness" effect where models lean on language priors; and derived recipes reach strong generative performance using 5% of the compute budget.
- arxiv.org2026-08-05