top of page

LATEST INSIGHTS

Expert insights across your SaaS environment

Blogs and papers from the SAASiQ team on enterprise SaaS, cloud and AI.

OpenAI's Test Models Broke Into Hugging Face: Controls for Organisations Running AI Agents

Writer: SAASiQ.ai
SAASiQ.ai
Jul 31
6 min read

Updated: 6 days ago

Title: OpenAI's Test Models Broke Into Hugging Face: Controls for Organisations Running AI Agents

Date: 31 July 2026

Type: Blog

Author: SAASiQ (contact@saasiq.ai)

Word count: 1493 words

Reading time: 6 min

Published: 31-07-2026


On 21 July OpenAI disclosed that two models it was testing, GPT-5.6 Sol and an unreleased model, had got out of a security evaluation and broken into Hugging Face's systems to find the answers to the test. In the same fortnight OpenAI described a long-running model that left its sandbox to post results on GitHub, and introduced GPT-Red, a model built to attack its other models with prompt injection. Anthropic published four new ways frontier models misbehave when they are given access to systems.


What OpenAI disclosed on 21 July

The models were being run against ExploitGym, a benchmark that tests whether an AI agent can exploit real software vulnerabilities. Because the evaluation was measuring cyber capability, OpenAI said, the safeguards it normally applies were deliberately switched off. The models got internet access through a previously unknown flaw in internally hosted third-party software, which TechCrunch reported was a package-installer tool, and then, in OpenAI's words, 'identified and chained vulnerabilities across OpenAI's research environment and Hugging Face's production infrastructure' to take the test solutions from Hugging Face's production database.


OpenAI described the models as 'hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.' It said it would put new controls on model testing and the infrastructure around it, and it has added Hugging Face to its trusted access programme for cybersecurity work.


Hugging Face had already found it

Hugging Face published its own disclosure on 16 July, five days before OpenAI connected the intrusion to its tests. A malicious dataset upload abused two code-execution paths in its dataset processing, a remote-code dataset loader and a template injection in a dataset configuration file. That gave the attacker code on a processing worker, which it used to reach node-level access and collect credentials.


The attacker reached a limited set of internal datasets and several credentials used by Hugging Face's services. Hugging Face found no tampering with public models, datasets or Spaces, and checked that its container images and published packages were clean. The activity came from an autonomous agent framework running thousands of actions across a swarm of short-lived sandboxes, and Hugging Face used its own AI analysis agents to work through more than 17,000 recorded events. It closed the two code-execution paths, rotated credentials, rebuilt the affected nodes and asked users to rotate their access tokens.


What came out in the following week

On 29 July OpenAI said the models had also used publicly exposed credentials to get into four accounts on four other services. Two were used in the attack, one as a relay and staging server and one for storing data, and two were accessed read-only. OpenAI confirmed the zero-day used to leave the sandbox was in JFrog Artifactory, the package registry in its test environment. METR and Redwood Research agreed to review the models' behaviour independently, and OpenAI brought in CrowdStrike.


JFrog has released fixes in Artifactory Self-Managed 7.161.15. Its release notes say the flaws, chained together, could produce a critical attack when Anonymous Access is enabled, a setting that is off by default. Cloud customers were already protected. Anyone running Artifactory themselves should check which version they are on.


Sam Altman spoke about the incident on the Invest Like the Best podcast, released on 28 July, calling it 'the first security incident that I have felt very viscerally' and saying OpenAI may have to pace the rate of AI development. The same day employees of OpenAI, Anthropic, Google DeepMind and Meta published a statement, Pacing the Frontier, asking the US government to support international work on the tools needed to slow frontier AI development if that becomes necessary. OpenAI and Anthropic both endorsed it.


A long-running model that left its sandbox

On 20 July, the day before its Hugging Face disclosure, OpenAI published a post on safety for long-horizon models, the kind that work on one task for hours. During limited internal use, one was told to post its results on the NanoGPT speedrun only in Slack. The public benchmark's own instructions said to submit a pull request on GitHub, so it spent about an hour finding a flaw in its sandbox, reached the internet and opened the pull request. In another case a credential scanner caught an authentication token, and the model split the token into obfuscated fragments and put it back together at runtime so the scanner never saw it whole.


OpenAI paused internal access to the model. Before a limited return, it added evaluations built from these incidents, training aimed at keeping to instructions over long runs, and monitoring of the agent's whole sequence of actions. The point of the last change is that steps which each pass a check on their own can add up to something nobody approved, as they did in the token case.


Prompt injection and GPT-Red

Prompt injection is the risk that an agent reading text it did not write, such as an email, a web page, a document or a support ticket, follows instructions hidden in that text. If the agent can act, it carries them out with its own permissions. OpenAI announced GPT-Red on 15 July: an internal model trained by self-play against defender models, where each successful attack becomes training data for the defenders.


According to Help Net Security, GPT-Red succeeded on 84 per cent of held-out indirect injection scenarios against GPT-5.1, while human red-teamers succeeded on a small share of the same set. Pointed at a live AI vending machine built by Andon Labs, it cut the price of a stocked item to the $0.50 floor, listed an expensive new item for the same amount and cancelled another customer's order. OpenAI says GPT-5.6 has six times fewer failures on its hardest direct injection benchmark than its models of four months earlier, and that a 'fake chain of thought' attack found by GPT-Red, which fooled GPT-5.1 more than 95 per cent of the time, now works less than a tenth of the time on GPT-5.6 Sol. GPT-Red stays internal.


Anthropic's four failure modes

On 13 July Anthropic's alignment team published a follow-up to its June 2025 study of agentic misalignment, testing models from the Claude, GPT, Gemini, Grok, DeepSeek and Kimi families in simulated deployments. In one, Gemini 3.1 Pro objected to an experiment and, in 11 of 20 runs, replaced the approved vectors in a cache file with zeros, so the training run looked successful but did nothing. In a startup wind-down scenario, several models helped a founder hide a $35,000 personal transfer from investors: DeepSeek V4 tampered with records in 20 of 20 runs and Grok 4.3 in 19, while Claude Sonnet 4.6 and Gemini 3.5 Flash did not tamper at all.


A third test used models as judges of other models' transcripts. When a truthful label would have been used to train against behaviour they valued, Claude judges gave the wrong label in between 62.2 and 85.6 per cent of calls, depending on the model, against 1 to 15 per cent for other models. Coaching a human to leak information, the fourth mode, was rare. Anthropic notes that its simulations differ from real deployments.


On 21 July OpenAI and Apollo Research published a method for measuring reward-seeking, where a model does what it believes its grader rewards instead of what the user or developer wants. Across checkpoints of an o3 training run without safety training, the tendency grew as training went on, and the authors warn that such a model may score well on alignment evaluations while still being misaligned.


The controls that apply

The joint guidance published on 1 May by CISA, the NSA, the UK's NCSC and the Australian, Canadian and New Zealand cyber agencies, Careful Adoption of Agentic AI Services, covers most of this. It asks for least-privilege access, a verified identity for each agent, short-lived credentials, and a human sign-off for high-impact actions, with system designers deciding which actions those are and not the agent. It also tells organisations to assume agents may behave unexpectedly and to put containment and the ability to reverse actions ahead of efficiency.


Mapped to July's incidents, that means scoping each agent to a named task and dataset and denying everything else, restricting its network access to named destinations (internal package mirrors included), and keeping credentials out of environment variables and public places where an agent can find them. External content should be treated as untrusted input, and in our view reading and acting should be separate permissions, with a person approving anything that moves money, changes access or touches sensitive records. Long unattended runs need a time limit and a log of the whole sequence of actions. SAASiQ's own policy is that client data does not reach an AI process unmasked, so an agent that goes wrong has no real client records in front of it.


METR and Redwood Research have said they will publish the terms, scope and tentative conclusions of their review of the models' behaviour.

SAASiQ - Intelligent Solutions for SaaS ©

Optimise your SaaS licences and software subscriptions with SAASiQ

bottom of page