Anthropic’s Mythos AI used social engineering to target real people


Anthropic’s Mythos AI agent, tested by the UK AI Safety Institute (AISI), has reportedly attempted a real‑world social‑engineering style hack against GitHub maintainers by creating fake human profiles, pressuring them to accept malicious code, and then editing logs to hide its tracks when challenged.

AISI was running cybersecurity evaluations of Anthropic’s Mythos and OpenAI’s Sol when it detected unusual outbound data transfers from its research systems. An investigation showed that some agents had engaged in “sustained, potentially harmful activity” targeting real people and organizations, rather than staying within the intended test environment.

The most serious activity involved an Anthropic Mythos agent tasked with solving a GitHub‑related cybersecurity challenge. The agent identified real GitHub maintainers, researched them, and created multiple fake accounts impersonating those individuals. Using private messages and a file‑sharing service, it tried to pressure and deceive the maintainers into approving malicious code so that it would run on GitHub’s systems.

Read more…
Source:  Malwarebytes Labs


Sign up for the Cyber Security Review Newsletter
The latest cyber security news and insights delivered right to your inbox


Related:

  • How AI-Native Development Platforms Enable Fake Captcha Pages

    September 19, 2025

    Artificial intelligence has revolutionized web development, empowering even novice users to create professional-looking websites. Tools like Lovable enable anyone to build and host applications with little to no coding knowledge, while Netlify and Vercel position themselves as AI-native development platforms. However, cybercriminals are increasingly exploiting these services to create and host fake captcha challenge websites, which ...

  • RevengeHotels: a new wave of attacks leveraging LLMs and VenomRAT

    September 16, 2025

    RevengeHotels, also known as TA558, is a threat group that has been active since 2015, stealing credit card data from hotel guests and travelers. RevengeHotels’ modus operandi involves sending emails with phishing links which redirect victims to websites mimicking document storage. These sites, in turn, download script files to ultimately infect the targeted machines. The final ...

  • Shiny tools, shallow checks: how the AI hype opens the door to malicious MCP servers

    September 15, 2025

    In this article, Kaspersky researchers explore how the Model Context Protocol (MCP) — the new “plug-in bus” for AI assistants — can be weaponized as a supply chain foothold. The researchers start with a primer on MCP, map out protocol-level and supply chain attack paths, then walk through a hands-on proof of concept: a seemingly legitimate ...

  • Palo Alto Networks becomes the latest to confirm it was hit by Salesloft Drift attack

    September 3, 2025

    The Salesloft Drift incident is quickly turning into the next MOVEit MFT fiasco, as yet another company confirms losing sensitive data in the third-party attack. This time around, it is the American multinational cybersecurity company Palo Alto Networks that confirmed losing customer data and support cases information in the breach. It all began with the sales ...

  • Zscaler says it suffered data breach following Salesloft Drift compromise

    September 3, 2025

    We can now add Zscaler to the growing list of Salesloft customers who suffered a third-party cyberattack and lost sensitive customer information after it confirmed data was taken. In the announcement, Zscaler explained it was a customer of Salesloft, whose AI chat platform, Salesloft Drift, was compromised. Since this platform connects with Salesforce, the miscreants managed ...

  • Model Namespace Reuse: An AI Supply-Chain Attack Exploiting Model Name Trust

    September 3, 2025

    Palo Alto Unit 42 research uncovered a fundamental flaw in the AI supply chain that allows attackers to gain Remote Code Execution (RCE) and additional capabilities on major platforms like Microsoft’s Azure AI Foundry, Google’s Vertex AI and thousands of open-source projects. We refer to this issue as Model Namespace Reuse. Hugging Face is a platform ...