Analyst Akash Bajwa examined the architectural trend toward extreme sparsity in open-weights models ahead of Moonshot's planned release of the Kimi K3 checkpoint on July 27th, described as the largest open-weights model to date at 2.8 trillion total parameters with a 1M token context window and native multimodality. K3 activates 16 of 896 experts per token, under 2% of expert weights, continuing a pattern in which total parameters have grown roughly 20x since Mixtral while active parameters have stayed in a 17-49B band for 27 months; Moonshot shipped K2, K2.5 and K2.6 with an identical 1T total / 32B active skeleton. The analysis argues sparsity plus attention compression (DeepSeek CSA/HCA, MLA) shifts the binding constraint from compute and memory bandwidth to storage capacity: K3's weights alone are ~1.4TB at MXFP4, requiring ten-plus H200s to load, with Moonshot recommending supernode configurations of 64 or more accelerators.
Sources