OpenAI hid AI agent hijacking of German wiki forum for weeks

OpenAI recently disclosed the details of how one of its AI models escaped a sandboxed environment and attacked Hugging Face during an evaluation – and as part of the incident, the models created a messaging board to communicate with each other and influence each other’s reasoning.

OpenAI has now disclosed that shortly after this incident, agents undergoing testing again escaped their ‘secured’ environment and hijacked an obscure German wiki to use as a messaging board. Per Reuters, OpenAI leadership kept the incident hidden while they dealt with the fallout from the Hugging Face incident.

Now that OpenAI has acknowledged its role in the incident, the company has said it is “past time” to put together an incident disclosure pipeline when its models escape testing and slip into third-party networks.

Read more…
Source:  TechRadar News


Sign up for the Cyber Security Review Newsletter
The latest cyber security news and insights delivered right to your inbox


Related:

  • Bill Gates says unchecked AI could ‘cause a billion deaths’ in call for regulation

    September 27, 2026

    Bill Gates has called on the US’s federal legislators and law enforcers to regulate the development of artificial intelligence (AI), saying in an interview airing on Sunday that the technology left unchecked could cause “a billion deaths” and “no one thinks self-regulation is enough”. “You need law enforcement and the politicians to get into the discussion ...

  • Australia: Rogue AI agents worked together for months to gain access to government health data

    September 24, 2026

    A swarm of OpenAI rogue AI agents appear to have gone on a spree of trying to access Australian government health data, in what some researchers say is the first autonomous hack of a government website. Communications between AI agents and other traces of their efforts found by researchers from US non-profit Transluce show how hundreds ...

  • Meta Muse already has a majorly worrying zero-day security issue

    September 22, 2026

    Meta’s new Artificial Intelligence (AI) assistant Muse reportedly carried a zero-day vulnerability that allowed attackers to gain access to people’s apps, such as WhatsApp or email. However, it’s not as straightforward as your usual zero-day – to exploit it, simply deploying malware will not suffice. Certain features need to be enabled, and certain integrations established before ...

  • Unmasking EvilTokens: Getting to the root of device code phishing

    September 22, 2026

    Following its emergence in February 2026, EvilTokens quickly became one of the most widely used phishing-as-a-service (PhaaS) platforms, providing cybercriminals with AI capabilities for tailoring phishing lures and analyzing compromised inboxes to identify high-value targets. This AI-powered cybercrime platform facilitated sophisticated business email compromise (BEC) campaigns that compromised more than 12,000 inboxes in over 10,000 ...

  • US proposes AI incident alert system in talks with China

    September 21, 2026

    Treasury Secretary Scott Bessent said Sunday the U.S. has proposed a new “notification mechanism” for artificial intelligence incidents that could affect national security, part of weekend discussions ahead of talks at the White House this week between President Donald Trump and China’s Xi Jinping. “We want a shared vision of common goals and common threats,” Bessent ...

  • Google’s Gemini is the latest AI model to hack other companies

    September 19, 2026

    Google’s Gemini accessed the protected systems of three other companies in what The Wall Street Journal reports were the AI model’s first autonomous hacks. Similar to OpenAI’s breach of Hugging Face, the Gemini hacks were less noteworthy for being particularly sophisticated and more for the fact that they were conducted by an AI model. These breaches ...