Two governments have put a number on the gap between China's newest frontier AI model and America's best — at least when it comes to hacking. In a joint preliminary assessment published July 23, the UK's AI Security Institute and the U.S. Center for AI Standards and Innovation concluded that Moonshot AI's Kimi K3 "performs significantly below" the most recent frontier cyber-capable models on their offensive-cyber benchmarks. Those top-performing models are American closed-weight systems, making the finding a data point in the increasingly public contest over who leads in AI.
Kimi K3 was released July 16 and is slated for open-weight release by July 27 — meaning its underlying model weights will soon be freely downloadable. That timing is part of why a cyber-capability evaluation drew attention: an openly available model that can meaningfully assist with hacking is a different proposition from a locked-down commercial one.
What the tests measured
The evaluation focused narrowly on offensive-cyber ability, using two main tests. The first, ExploitBench — a public benchmark from Carnegie Mellon that probes a model's ability to build software exploits against 41 recent vulnerabilities in Chrome's V8 engine — showed the clearest gap. Kimi K3 scored 32%, but it never reached the benchmark's most severe outcome: it achieved arbitrary code execution, which would let an attacker hijack a target, on zero of 41 samples, compared with an average of 20 of 41 for the most cyber-capable models.
The second test, a simulated 32-step corporate-network attack called "The Last Ones," told a similar story. Kimi K3 advanced to step 17 on average, while leading U.S. models reached 28.5. In one of 10 attempts, Kimi K3 did complete the full attack chain — evidence, the institutes noted, that it can autonomously breach small, weakly defended systems when directed and given a foothold — but the most capable American models solved the same range far more reliably, at six and seven times out of ten.
Notably, the assessment also flagged a safety point unrelated to raw skill: Kimi K3's built-in safeguards "did not prevent it from attempting cyber exploit development or offensive cyber operations" during testing. In other words, it was willing to try.
The caveats the report itself raises
The findings come with qualifications that matter, several of them stated plainly by the evaluators. Kimi K3 did outperform GLM-5.2, the strongest openly available model as of June, on the same tests — so within the open-weight category, it is comparatively capable. The U.S. closed-weight models it was measured against, meanwhile, were run with their system-level safeguards disabled to reveal maximal capability; the versions the public can actually use have those guardrails switched on. And the report cautioned that its estimate of Kimi K3's overall cyber ability carries a wider margin of error than the others, because it was derived from a single benchmark rather than the broader battery used for the American models. The simulated network, too, lacked the active defenders and alerting of a real corporate environment.
These are the kinds of details that separate a careful technical evaluation from a scoreboard, and they cut against reading the result as a simple verdict that Chinese AI is broadly behind. The test measured one capability — offensive cyber — not general intelligence or the many other tasks a model performs.
A contested backdrop
The assessment lands in a charged environment. Moonshot has promoted Kimi K3 as a model that can rival the offerings of leading American labs, and some observers were quick to ask whether a U.S. government body declaring a Chinese model inferior reflects the measurements or the politics around them. The institutes' answer is transparency: they published their methods, benchmarks and confidence intervals for scrutiny, and released the assessment jointly with a foreign partner rather than unilaterally.
The broader context is a public increasingly uncertain about the race. Survey data has shown many Americans now perceive China as more advanced on AI than the United States, even as U.S. models continue to top capability leaderboards in areas like this one. What the Kimi K3 evaluation offers is not a resolution of that debate but a concrete, narrow measurement within it — one that suggests America retains an edge in frontier offensive-cyber capability for now, while underscoring how quickly openly available models are climbing toward it.



