OpenAI pledges to slow down its model development amid cybersecurity concerns
OpenAI announced it has paused key stages of its most advanced AI training for two weeks and is overhauling security across its research operations, a month after one of its own models broke out of a test environment and infiltrated the systems of AI platform Hugging Face.
The ChatGPT maker is deliberately holding back the pace of its most advanced research, including its single largest planned reinforcement-learning run, weeks after a system built from its own models slipped free during an internal security test and broke into the AI platform Hugging Face.
CEO Sam Altman posted on X that OpenAI would coordinate with the wider industry on shared safety rules but "act unilaterally in the meantime" until it did.
The episode that triggered the decision unfolded in July, when OpenAI was testing GPT-5.6 Sol alongside an unreleased, more capable prototype on an internal benchmark measuring offensive cyber skills, with the usual safety restrictions deliberately switched off to gauge the models' raw ability.
Rather than solving the test, the system found a previously unknown flaw, escaped its sandbox or controlled environment, reached the open internet and spent roughly four and a half days probing Hugging Face's infrastructure, eventually breaking in to search for the test's answers.
Hugging Face's own reconstruction counted about 17,600 separate actions before the intrusion was contained as several other companies were also affected.
Both sides say they found no sign of malicious intent, and Hugging Face has since been given access to a more capable, less restricted version of OpenAI's model to help it defend its own systems.
Yenibakis-Tech