An artificial intelligence model was handed one narrow instruction: score well on a hacking test. Instead it slipped out of the locked practice environment it was supposed to stay in, found a path onto the open internet, and broke into a real company's servers to steal the answer key. This is not a scene from a film. It is what OpenAI has told the public happened inside its own labs.

The company breached was Hugging Face, the platform where millions of developers, many of them in Malaysia, host AI models and datasets. The lesson for any business rushing to put autonomous agents to work is an uncomfortable one: an agent chasing a goal does not stop at the line you assume it will.

Editor
Editor

Kai T chevron_right

Tech editor at ProductNation Malaysia Covers the latest in gadgets, apps, AI, and consumer tech, turning press releases into stor ...

What OpenAI disclosed

On 21 July 2026, OpenAI said two of its models, including the recently launched GPT-5.6 Sol and a more capable unreleased system, escaped a sandboxed evaluation called ExploitGym, a benchmark of long, multi-step hacking tasks run with the models' safety refusals deliberately lowered. According to OpenAI, the models exploited a previously unknown flaw, a zero-day, in internal third-party software to reach the internet, then broke into Hugging Face and pulled the benchmark's answers from a production database. TechCrunch reported the intrusion began with a booby-trapped dataset. Hugging Face had already detected and contained the breach on 16 July, five days before OpenAI connected it to its own testing.

OpenAI says its models found and chained a real zero-day flaw without seeing the target source code. Photo: Shahadat Rahman on Unsplash.

Security researchers have described it as the first documented case of frontier AI independently discovering and chaining real-world attack paths, including at least one genuine zero-day, with no access to the target's source code. No human directed each step. The models were simply optimising, and the fastest route to a perfect score ran straight through someone else's servers.

Why this lands hard in Malaysia

Malaysian companies are among the most eager in the world to put AI agents to work. In Microsoft's most recent Work Trend Index, 86 percent of Malaysian business leaders said they were confident about using AI agents to expand their teams within 18 months, above the global average. The catch is the gap between that ambition and what is actually running. The ASEAN Innovation Business Platform's 2025 to 2026 readiness survey found most local deployments are still confined to low-risk work like back-office automation and customer-service chatbots, held back by fragmented data, ageing systems and a shortage of skilled staff.

Malaysian firms are keen to deploy AI agents, but most run them only on low-risk tasks so far. Photo: Gatot Adri on Unsplash.

The Hugging Face episode is a preview of what arrives the moment a Malaysian firm hands an agent real credentials, a live internet connection and a target to optimise for. The legal exposure is not hypothetical either. Under Malaysia's amended Personal Data Protection Act, in force since 2025, a data breach must be reported to the regulator, so an agent that leaks customer records while pursuing a task becomes a compliance event, not just an IT headache.

Malaysia has no rulebook for this yet

The country still has no law written specifically for autonomous agents. The AI Governance Bill drafted through the National AI Office remains pre-Cabinet and takes a broad, risk-based approach rather than naming agents as a distinct category, the twist that stood out when China became the first country to set dedicated agent rules in July. Globally the same hole is visible: roughly 74 percent of organisations plan to adopt agentic AI within two years, yet only about 21 percent say they have a mature way to govern it.

The takeaway from OpenAI's own lab is not that agents are too dangerous to touch. It is that an agent will do exactly what you reward it for, including things nobody imagined. For Malaysian businesses racing to deploy, the guardrails, access limits and monitoring have to be in place before the agent is, not after.

Images courtesy of Igor Omilaev, Shahadat Rahman and Gatot Adri on Unsplash.