Magic says its pretraining recipe is now more than 10 times more compute-efficient than leading open-weight base models. It matched DeepSeek V4 Pro's base model using around 50 times fewer computing operations. Magic then scaled the same recipe up tenfold and beat all publicly available open base models on loss measurements. The gains came from dozens of small changes to architecture, optimizer, training objective and data curation. Magic measured bits-per-byte loss on held-out data, then fitted scaling curves to project compute needs. It tested open models from DeepSeek, Moonshot and NVIDIA, and had Fireworks verify the baseline numbers. It filtered its test sets to remove anything overlapping its training data. Magic says it will now scale long-horizon reinforcement learning and alignment work. No model has been released, and the company says a release is coming.
What changed
Frontier-level base models required roughly 100x more training compute, costing over $100M.
What it unlocks
Training a base model that beats open-weight rivals for single-digit millions of dollars.
- ~50x fewer FLOPs to match DeepSeek V4 Pro
- ~$0.5M on GB200 to match it
- 10x scale-up cost ~$4M
- >$100M under DeepSeek's recipe
What you need to act on it
- no public model or API released yet
Sources