AI agent was tasked to collect mundane data collection and autonomously used hacking techniques to get it.
Posted by Leslie Eastman

Back in March, I reported that hackers reportedly “jailbroke” Anthropic’s Claude chatbot and used it to help steal roughly 150 GB of sensitive data from multiple Mexican government entities, including tax and voter records.
“Jailbroke” means the attacker deliberately manipulated the AI system’s safeguards so it would ignore or bypass its built‑in safety rules and produce help it was supposed to refuse. Jailbreaking techniques include generating requests that treat malicious content as benign, crafting a series of prompts that gradually grant access to sensitive content, or exploiting loopholes in the model’s policy enforcement.
Now it appears that OpenAI did much the same thing to an Australian government website.
A rogue OpenAI model bypassed safeguards during training and hacked an Australian government website, the country’s Prime Minister Anthony Albanese said Wednesday, admonishing the ChatGPT creator over an “obviously unacceptable” breach.
The artificial intelligence tool sought access to a health statistics portal in June and “didn’t accept no for an answer,” sidestepping restrictions to breach a section hosting private files, Albanese said.
OpenAI did not raise the alarm with the Australian government until September, when it sent a message to a generic email inbox that is only checked once daily. The company said it first spotted the rogue activity in August, when it reviewed what the AI tool had been doing.
[…]
The AI tool accessed public and non-public files hosted on an old health statistics website. Albanese said there was “no evidence” that personal information had been accessed or that other government services had been compromised.
The Australian government has set up a task force to investigate the breach and to determine if existing network security can stop similar incidents.
Australia has faced a series of hacking attempts on corporations and government-linked firms over the past four years.
Defence Minister Richard Marles said the Medicare portal that was breached did not contain individual medical claims, benefit payments, personal banking details, or patient medical histories of Australia’s 27 million people.
Instead, the website holds only aggregated data on healthcare use across the country, he said. However Australia considered the breach serious.
“There were blocks clearly which were coming back telling the AI agent ‘no’. The AI agent found a way around those blocks – didn’t accept no for an answer,” Albanese told reporters.
Transluce (a research lab focused on AI oversight) identified that OpenAI acquired health data from one Australian government website and attempted to breach another. Other incidents involved OpenAI’s AI agents trying to hack websites while performing routine research tasks. OpenAI confirmed the agents were theirs.
Two unsuccessful attempts targeted the University of New Mexico’s digital library on May 25–26 and the Data USA public-data platform on May 28. Additionally, in July, I reported that OpenAI recently revealed that its own advanced AI models essentially went rogue and attempted to hack external systems, including the widely used Hugging Face platform.
What makes the Australian incidents different is that they are the first known incidents of AI autonomously breaching into a government cyberspace without being directed to by hackers.
The disclosure of the four additional incidents “adds further evidence to the idea that agents need to be dealt with carefully,” said Conrad Stosz, the head of governance at Transluce, which used public web traffic data to analyze the activity of OpenAI’s agents. Agents are autonomous programs that work to execute tasks for a user.
Mr. Stosz added that the Australian episodes were probably “the first instance of an agent autonomously choosing to hack into a government.”
An OpenAI spokeswoman said on Wednesday that the company had reached out to the University of New Mexico and DataUSA and had been in communication with the Australian government about the incidents.
“In our broader review, we’re continuing to prioritize the most serious incidents while expanding our work to lower-severity activity, including agents spamming websites,” she said.
These incidents are not a reason to reject AI, but they show why strong safeguards and human oversight matter.
With proper controls, AI can help advance medicine, improve cybersecurity, and solve difficult problems much faster. Making AI companies liable for the misuse of their systems would incentivize the development of these controls.
The answer is not to stop progress, but to make sure this powerful technology is used safely and responsibly… and that “frontier AI forms” are motivated to address the necessary security issues.