Threat Insight

Rogue AI Agent Allegedly Hack Hugging Face

On July 16 the AI Tool development company Hugging Face disclosed that they had been the victim of a cyberattack. The origin of the attack appears to have been an AI Agent run by OpenAI. [1]

In response OpenAI released a statement claiming that the AI Agent was a testing tool that had managed to escape the sandbox it was run in, in what they claimed was an unprecedented cyber incident. The AI models involved were OpenAI’s GPT-5.6 Sol and a more capable, as-yet-unreleased model. [2]

  • Insight

According to public knowledge the AI Agent used a test model with no security restrictions that had been given a task to solve as a test and decided that the best way to solve it was to use Hugging Face’s connection to OpenAI to find the answer to the test question.[3]

In their disclosure, Hugging Face also claimed that its AI-powered security solutions spotted the unusual activity and detected the AI attack. However, when they tried to use commercial AI tools to help with their investigation of the incident, the tools refused as their built-in safety filters flagged the attack data as suspicious content and blocked the requests. To get around this, Hugging Face turned to GLM 5.2 – a Chinese open-source AI model they could run on their own systems, where no such restrictions applied. [2]

A Timeline of the events suggests that the intrusion ran from July 11 to July 13 and that OpenAI did not connect its agent to the attack until several days later. [4]

Assessment

There are several reasons to question the media narrative that this incident represents a new threat of dangerous rogue AI agents.

The official timeline suggests that OpenAI let their agent run for several days trying to solve a hacking problem without human supervision. This indicates potentially inadequate security controls on their behalf, regardless of whether they ran this Agent in what they thought was a secure Sandbox or not.

Nothing that has so far been reported, including the fact that it took the AI Agent days to breach Hugging Face, supports the narrative that the AI Agent had managed to do anything out of the ordinary, compared to a human threat actor.

It is also important to remember that most of the information comes from OpenAI and Hugging Face and that both these parties have their own stakes in the incident. OpenAI may have an incentive to emphasize the capabilities demonstrated by the model, while deflecting form discussions about their internal security practices. Hugging Face on the other hand is a company invested in open-source AI tools, and as such may have an interest in highlighting potential advantages of open-source tools over closed source.

Both OpenAI and Anthropic have advocated government regulation of AI models, and limiting open-source models. Something other tech leaders claim would limit competition and development. It is likely that this conflict has shaped both OpenAI and Hugging Face messaging on this incident. [5]

There is not enough information yet to determine if this cyberattack represents anything new, if it is a marketing-driven narrative, or a way of deflecting from poor cybersecurity practices on either or both involved parties.

Due to the many unanswered questions Truesec recommends caution in drawing too far-reaching conclusions based on this incident until more information is released.

References

[1] https://huggingface.co/blog/security-incident-july-2026
[2] https://openai.com/index/hugging-face-model-evaluation-security-incident/
[3] https://huggingface.co/blog/security-incident-july-2026
[4] https://www.reuters.com/business/its-ai-agent-spent-days-hacking-company-sources-say-openai-did-not-notice-week-2026-07-24/
[5] https://cyberscoop.com/tech-leaders-open-source-ai-cybersecurity/

Stay ahead with cyber insights

Newsletter

Stay ahead in cybersecurity! Sign up for Truesec’s newsletter to receive the latest insights, expert tips, and industry news directly to your inbox. Join our community of professionals and stay informed about emerging threats, best practices, and exclusive updates from Truesec.

Latest Insights