Research2026-08-07

Independent benchmarking firm Artificial Analysis published measurements for Ling 3.0 Flash, an openly downloadable reasoning model from InclusionAI released on 4 August 2026. It ranks first on the firm's combined intelligence score and on output speed among open-weights models of similar size, while pricing sits mid-pack. The model is text-only and uses far more output text per task than comparable models, which raises real running costs.

What changed

Comparable open-weights models in the 40B-150B size class scored around 9 on the same index and ran at about 114 tokens per second.

What it unlocks

Running a downloadable reasoning model that leads its size class on this independent benchmark, at roughly a fifth of typical input pricing, either through two hosted providers or on own hardware.

  • Intelligence Index score 38 vs median 9 for comparable open-weights models
  • 414.7 output tokens per second (median 113.7)
  • $0.07 per 1M input tokens, $0.22 per 1M output tokens
  • 124B total parameters, 5.1B active
  • 262k token context window
  • 240M output tokens used in the Intelligence Index, vs median 57M

What you need to act on it

  • API access through one of two providers, or hardware to run a 124B-parameter model

Send this to someone who needs it

Shares the story and its sources. Nothing about you.

What does this mean for your job?

This is the story as everyone gets it. Once a week we send you the version written for your role — what changed, why it matters for the work you actually do, and one thing to try. Free while we tune it.