Independent benchmarking firm Artificial Analysis published measurements for Ling 3.0 Flash, an openly downloadable reasoning model from InclusionAI released on 4 August 2026. It ranks first on the firm's combined intelligence score and on output speed among open-weights models of similar size, while pricing sits mid-pack. The model is text-only and uses far more output text per task than comparable models, which raises real running costs.
What changed
Comparable open-weights models in the 40B-150B size class scored around 9 on the same index and ran at about 114 tokens per second.
What it unlocks
Running a downloadable reasoning model that leads its size class on this independent benchmark, at roughly a fifth of typical input pricing, either through two hosted providers or on own hardware.
- Intelligence Index score 38 vs median 9 for comparable open-weights models
- 414.7 output tokens per second (median 113.7)
- $0.07 per 1M input tokens, $0.22 per 1M output tokens
- 124B total parameters, 5.1B active
- 262k token context window
- 240M output tokens used in the Intelligence Index, vs median 57M
What you need to act on it
- API access through one of two providers, or hardware to run a 124B-parameter model
- artificialanalysis.ai2026-08-07