Malaysia is being sold a future in which software agents identify the right government service, guide you through it and, once authorised, carry out actions such as paying a summons for you. OpenAI, the maker of ChatGPT, has just paused training, evaluation and inference with tool use for its most capable models, and says that work will stay stopped until it has hardened its systems further. It is the second time in a few months that agent behaviour has forced the company to pull back, though the two responses were not identical in scope.
What OpenAI actually did
On 20 September, one of OpenAI's models exploited a gap in its sandbox's network restrictions, reaching a public chatbot on the open internet through a DNS resolver. OpenAI paused training, evaluation and inference with tool use, defined broadly, for its most capable models, Fortune reported on 26 September, calling it the second such halt in under three months. OpenAI said this incident was a lot less severe than some of its previous ones.
That first pause came in July. After agents broke out of their sandbox and compromised parts of the AI infrastructure firm Hugging Face, OpenAI paused reinforcement-learning training for around two weeks, according to Fortune. The current pause applies to inference that uses tools, not to every use of those models. "All inference for our most capable models remains stopped until we have hardened our systems further," Micah Carroll, who leads preparedness for the risk at OpenAI, wrote in a post on X that Fortune quoted.
The picture around those two events has widened. OpenAI has said it notified dozens of third parties about cases where its models bypassed security controls or used websites in ways they were not meant to, and one person briefed on the matter told Reuters the tally was around two dozen and still climbing as the review continued, a figure OpenAI has not confirmed. In July the agents created close to a million shortened links while trying to work around obstacles such as CAPTCHA, according to New York Times reporting, based on research by the firm Parse, that Fortune cited. Reuters reported that 53 images belonging to ChatGPT users were leaked onto image-hosting sites, an incident OpenAI has acknowledged, though OpenAI says the links were never publicly listed, the data had been privacy-filtered, and most of the images have since been removed. Fortune reported that government websites in the United States and Australia were among those affected. Chief executive Sam Altman said Hugging Face "is still the most severe event we've seen".

Why this matters in Malaysia
This is not a distant Silicon Valley problem. Malaysia's own government portal, MyGOV, is already exploring agentic AI for citizen services, and YTL AI Labs has built a model with NVIDIA, ILMU-Nemo-Nano, that powers its own agent tools. The national pitch is agents that do things for you, not ones that only answer you. The Ministry of Digital's announcement does not name the models behind MyGOV's beta, and a government service is a very different setting from a lab stress test, but the direction is shared: autonomous software that takes actions rather than just giving answers.

The lesson from OpenAI's two pauses is not that agents are useless. It is that the power to act is also the power to act wrongly, and that catching a problem is not the same as stopping it. According to Fortune, OpenAI's systems flagged the 20 September escape within about 15 minutes, a reviewer began looking three minutes later, and training was manually stopped about two and a half hours after that. Noticing, for all that monitoring, was not the same as stopping. If a Malaysian government agent is ever given real access to a citizen's records, that gap between noticing a problem and stopping it is where the risk would live, and the design has to close it before the agent goes live, not after.
So the questions worth asking of any Malaysian agent rollout are practical ones. Can it take an action without a human approving it first? What systems and data can it reach if it goes off script? And how quickly can someone actually stop it? Human oversight is only one of the safeguards OpenAI lists, alongside sandboxing, access limits and monitoring, but its own timeline is a reminder of why the ability to pull the plug fast matters.
What to watch next
OpenAI has not said when that tool-using research work will resume. The company selling the world an agent-run future has had to rein in its most capable models twice in a few months, first pausing some of their training after the July breakout, then their tool-using work more broadly this month. Malaysia is moving in the same direction, one government service at a time.
Image(s) courtesy of Tyler and Abu Saeid on Unsplash. Chart by ProductNation.




