Research2026-08-05

A team including Junlin Han, Shengbang Tong, David Fan, Minghao Chen, Philip Torr, Filippos Kokkinos and Mike Lewis posted an empirical study of natively unified multimodal pretraining to arXiv on 5 August 2026 (revised 6 August). Using controlled experiments on synthetic and large-scale real datasets, the authors report four findings: knowledge transfer between language, visual understanding and visual generation is asymmetric; data complexity largely determines whether modalities are synergistic or competitive, with shared attention and normalization plus modality-specific feed-forward layers promoting synergy across different visual tokenizer designs; unifying modalities early and training jointly beats late alignment or sequential training, with delayed integration producing a "vision laziness" effect where models lean on language priors; and derived recipes reach strong generative performance using 5% of the compute budget.

Send this to someone who needs it

Shares the story and its sources. Nothing about you.

What does this mean for your job?

This is the story as everyone gets it. Once a week we send you the version written for your role — what changed, why it matters for the work you actually do, and one thing to try. Free while we tune it.