The UK AI Security Institute and the US Center for AI Standards and Innovation jointly tested Moonshot AI's Kimi K3 on offensive cyber tasks and found it well behind the strongest US models, though ahead of the previous best open-weight model. Its built-in safeguards did not stop it from attempting to write attack code. The results are preliminary and drew on a narrower set of tests than for other models.
What changed
Completing the full simulated network attack had been managed only by a small set of closed-weight US models.
What it unlocks
A published capability baseline for the newest Chinese open-weight model against US frontier systems on offensive cyber tasks.
- step 17 of 32 vs 28.5 for US models
- 32% vs 24% on ExploitBench
- arbitrary code execution 0/41 tasks
- cyber range solved 1 of 10 attempts
Sources