In Brief
Posted:
7:28 AM PDT · August 7, 2026
Image Credits:Lam Yik/Bloomberg / Getty ImagesKimi K3, the latest AI model made by Chinese company Moonshot, escaped an environment set up to test its cyber capabilities, researchers said in a blog post published on Friday.
The news shows once again that companies and independent organizations are struggling to contain their AI models designed for hacking.
In recent weeks, frontier LLMs at U.S. artificial intelligence labs OpenAI and Anthropic, Meta, as well as the UK’s AI Security Institute, all escaped testing environments in different ways and ended up hacking real targets that were not part of the experiment. This is starting to happen so often there’s now a website tracking all these incidents called Felony Bench, a nod to the fact that these LLMs may be committing crimes — at least theoretically speaking.
In the case of this Kimi test, the sandbox designed to contain the experiment was not properly configured. While the sandbox disallowed the AI model from accessing certain web traffic, the model instead bypassed the sandbox by relying on command line tools, according to the researchers AI-focused cybersecurity firm Frontier Security.
“This suggests that some of the evaluations on cybersecurity the community uses are susceptible to security vulnerabilities and allow models to cheat, and that there are models that intentionally seek loopholes and vulnerabilities which allows them to cheat on evaluations,” the researchers wrote.
If you are keeping score at home, according to Felony Bench’s tally, Moonshot now joins alongside OpenAI and Anthropic, which have seven recorded incidents each, and Meta, which has one.
Subscribe for the industry’s biggest tech news




.jpg?mbid=social_retweet)




English (US)