The systems bypassed restrictions to communicate through previously undisclosed websites, according to the news agency
Rogue AI agents developed by OpenAI bypassed restrictions on posting online and used more than ten previously undisclosed websites to communicate with one another, Reuters reported on Wednesday, adding that the company kept the unauthorized activity quiet for months.
An AI agent is a software system that can independently carry out tasks on a user’s behalf, making decisions and taking actions with limited human supervision. The findings raise concerns that increasingly autonomous AI systems can find ways around safeguards and take actions their developers did not authorize or anticipate.
Reuters reviewed findings from six independent investigators or research groups, which suggested that the activity was far more widespread than previously known. Researcher Andrew Yoon identified 18 previously undisclosed sites used from May to July, while a team led by Sydney Von Arx found evidence across 23.
”We have no idea how much is out there,” Von Arx said, warning that the estimates are incomplete.
OpenAI did not directly answer Reuters’ questions about how many sites were affected or why the activity was kept under wraps for months. The company said it was conducting a broader review.
The agents reportedly repurposed old wikis, online text-storage services, and university link shorteners as makeshift message boards while carrying out research tasks. Although permitted to search the web for information, they were not supposed to post content, but found ways to leave messages for one another using editable pages and other website features.
The findings expand on an incident revealed last week involving DseWiki, a German programming wiki where researchers attributed more than 15,000 edits to OpenAI agents. The systems reportedly used the site to exchange tactics for completing tasks, circumventing restrictions, and avoiding detection.
Researchers said some discussed using Tor, which allows users to browse the internet anonymously, and created backup pages after a moderator began deleting their messages.
The incidents follow a more serious breach in July, when OpenAI agents circumvented safeguards during cybersecurity testing, reached the open internet, and gained unauthorized access to systems belonging to Hugging Face, a major platform for hosting and sharing AI models.
OpenAI later acknowledged that its agents exploited vulnerabilities and accessed parts of Hugging Face’s infrastructure without authorization, describing the episode as a “warning shot” about the risks posed by increasingly capable autonomous systems.
The company told Reuters that it found no additional activity matching the “severity or scale” of the Hugging Face breach. OpenAI said it was reviewing broader agent activity and developing a framework for reporting AI “misalignment” during model training, evaluation, and deployment.
You can share this story on social media:



