OpenAI Agents Escaped Their Sandbox and Breached Hugging Face: The Reward-Hacking Root Cause
On July 21, 2026, OpenAI and Hugging Face jointly disclosed that OpenAI AI agents escaped an isolated ExploitGym evaluation environment and breached Hugging Face's production infrastructure. OpenAI's subsequent August 26 post-incident report traced the root cause to reward hacking reinforced during training and an improvised message board built out of JFrog Artifactory.