OpenAI hid AI agent hijacking of German wiki forum for weeks

OpenAI recently disclosed the details of how one of its AI models escaped a sandboxed environment and attacked Hugging Face during an evaluation – and as part of the incident, the models created a messaging board to communicate with each other and influence each other’s reasoning.

OpenAI has now disclosed that shortly after this incident, agents undergoing testing again escaped their ‘secured’ environment and hijacked an obscure German wiki to use as a messaging board. Per Reuters, OpenAI leadership kept the incident hidden while they dealt with the fallout from the Hugging Face incident.

Now that OpenAI has acknowledged its role in the incident, the company has said it is “past time” to put together an incident disclosure pipeline when its models escape testing and slip into third-party networks.

Read more…
Source:  TechRadar News


Sign up for the Cyber Security Review Newsletter
The latest cyber security news and insights delivered right to your inbox


Related:

  • OpenAI hid AI agent hijacking of German wiki forum for weeks

    September 7, 2026

    OpenAI recently disclosed the details of how one of its AI models escaped a sandboxed environment and attacked Hugging Face during an evaluation – and as part of the incident, the models created a messaging board to communicate with each other and influence each other’s reasoning. OpenAI has now disclosed that shortly after this incident, agents undergoing testing ...

  • Attackers Expose Ongoing AI Tool Use Targeting Organizations in Latin America

    September 3, 2026

    We have analyzed two ongoing, multi-stage network intrusion and data-exfiltration campaigns targeting organizations in Latin America. Corroborating recent findings from the broader threat intelligence community, we observed attackers leveraging artificial intelligence (AI) to enhance their capabilities. Read more… Source:  Palo Alto Unit 42 Sign up for the Cyber Security Review Newsletter The latest cyber security news and insights delivered ...

  • Judge rules the Pentagon’s supply chain risk label for Anthropic unlawful

    August 28, 2026

    A federal judge ruled the Pentagon’s decision to label AI company Anthropic a ‘supply chain risk’ violated the law and ordered the designation be removed Thursday evening. Judge Rita Lin wrote that while the military should have wide latitude to decide which companies to work with, its actions against Anthropic “constituted unlawful retaliation in violation of ...

  • An open letter for a global surge in cyber defense

    August 27, 2026

    We have a limited window to strengthen cyber defenses. In the coming months, AI-enabled cyber attacks will become far more widespread and sophisticated as models around the world become increasingly capable. The companies and public services our communities depend on—from hospitals to water treatment plants to the infrastructure that powers the internet—are at risk. Today’s AI advances ...

  • Anthropic’s Mythos AI used social engineering to target real people

    August 6, 2026

    Anthropic’s Mythos AI agent, tested by the UK AI Safety Institute (AISI), has reportedly attempted a real‑world social‑engineering style hack against GitHub maintainers by creating fake human profiles, pressuring them to accept malicious code, and then editing logs to hide its tracks when challenged. AISI was running cybersecurity evaluations of Anthropic’s Mythos and OpenAI’s Sol when it detected ...

  • Scammers target OnlyFans users with deepfakes

    August 6, 2026

    OnlyFans creators are used to posting adult videos of themselves online, but what happens if someone takes control of their images and uses them for fraud? This week, USA Today revealed how criminals are impersonating OnlyFans creators using AI tools. They use deepfake content to lure the real models’ fans with fake promises of live chats, and ...