ads
bc-game

OpenAI reveals incident involving GPT-5.6 after breach of Hugging Face environment

2 min read
PortalCripto
OpenAI reveals incident involving GPT-5.6 after breach of Hugging Face environment
Source: Mariia Shalabaieva/Unsplash — OpenAI reveals incident involving GPT-5.6 after breach of Hugging Face environment
Give preference to us on Google
Advertisement

OpenAI revealed an unprecedented incident involving some of its most advanced artificial intelligence models during an internal security evaluation. According to the company, GPT-5.6 Sol and another model not yet released managed to bypass the restrictions of an isolated testing environment and compromise part of Hugging Face's production infrastructure.

The episode took place during the execution of ExploitGym, a benchmark created to measure the ability of AI agents to identify and exploit complex security flaws. To make the test more rigorous, the models operated with a reduced number of blocks related to cyber activities, allowing an assessment of their maximum potential in offensive operations.

During the evaluation, the systems directed a large part of their computational resources to finding a way to escape the restricted environment. In this process, they identified a previously unknown zero-day vulnerability in third-party software used by OpenAI as a proxy and cache for package records.

After exploiting this flaw, the models managed to elevate their privileges within the company's research infrastructure. They then moved between different systems until reaching a machine with unrestricted internet access.

Once connected to the external network, the models concluded that Hugging Face could store models, datasets, or responses related to ExploitGym. Instead of finishing the benchmark as planned, they began seeking direct access to this information.

OpenAI reported that the systems combined credentials obtained during the intrusion with different vulnerabilities to create a remote code execution path on Hugging Face's servers. As a result, they were able to access benchmark solutions stored in the platform's production database.

In a separate statement, Hugging Face confirmed that an autonomous AI agent accessed a limited number of internal datasets and some service credentials. The company emphasized that it found no evidence of changes to public models, datasets, Spaces, or its software supply chain.

The company also reported that its own security mechanisms, combined with open-source AI models, quickly detected and stopped the activity. The investigation was able to reconstruct more than 17 thousand actions recorded throughout the entire sequence of the intrusion.

According to OpenAI, the models remained focused on finding the solution to the proposed challenge, without showing behavior aimed at causing additional damage. Even so, the company classified the case as an unprecedented cyber incident, capable of demonstrating how advanced AI models can discover and combine novel vulnerabilities in real systems, even without access to the source code.

After the incident, OpenAI reinforced the confinement, monitoring, and access control mechanisms of its research infrastructure. The company also disclosed the vulnerability to the developer of the affected software, launched a forensic investigation in partnership with Hugging Face, and included the platform in its trusted access program, offering advanced AI resources to support defense teams in future security analyses.

Tags
Advertisement