Anthropic’s Mythos AI used social engineering to target real people


Anthropic’s Mythos AI agent, tested by the UK AI Safety Institute (AISI), has reportedly attempted a real‑world social‑engineering style hack against GitHub maintainers by creating fake human profiles, pressuring them to accept malicious code, and then editing logs to hide its tracks when challenged.

AISI was running cybersecurity evaluations of Anthropic’s Mythos and OpenAI’s Sol when it detected unusual outbound data transfers from its research systems. An investigation showed that some agents had engaged in “sustained, potentially harmful activity” targeting real people and organizations, rather than staying within the intended test environment.

The most serious activity involved an Anthropic Mythos agent tasked with solving a GitHub‑related cybersecurity challenge. The agent identified real GitHub maintainers, researched them, and created multiple fake accounts impersonating those individuals. Using private messages and a file‑sharing service, it tried to pressure and deceive the maintainers into approving malicious code so that it would run on GitHub’s systems.

Read more…
Source:  Malwarebytes Labs


Sign up for the Cyber Security Review Newsletter
The latest cyber security news and insights delivered right to your inbox


Related:

  • AI, cyber-attacks and amateur experiments threaten to upend global biosecurity, WHO warns

    July 13, 2024

    Artificial intelligence, cyber-attacks and genetic engineering could pose potentially catastrophic biosecurity threats to countries around the world, the WHO has warned. Rapid technological advances in the past decade have “redefined the biological threat landscape” and heightened risks of manipulation, the updated guidance from the WHO’s Technical Advisory Group on Biosafety said. The report advised that member ...

  • NATO releases revised AI strategy

    July 10, 2024

    On Wednesday (10 July 2024), NATO released its revised artificial intelligence (AI) strategy, which aims to accelerate the use of AI technologies within NATO in a safe and responsible way. AI or Artificial intelligence concept. It builds on one published in 2021 and takes account of recent advances in AI technologies, such as generative AI, and ...

  • OpenAI breach is a reminder that AI companies are treasure troves for hackers

    July 5, 2024

    There’s no need to worry that your secret ChatGPT conversations were obtained in a recently reported breach of OpenAI’s systems. The hack itself, while troubling, appears to have been superficial — but it’s reminder that AI companies have in short order made themselves into one of the juiciest targets out there for hackers. The New York ...

  • Void Arachne Targets Chinese-Speaking Users With the Winos 4.0 C&C Framework

    June 19, 2024

    In early April, Trend Micro researchers discovered that a new threat actor group (which they call Void Arachne) was targeting Chinese-speaking users. Void Arachne’s campaign involves the use of malicious MSI files that contain legitimate software installer files for artificial intelligence (AI) software as well as other popular software. The malicious Winos payloads are bundled alongside ...

  • AI jailbreaks: What they are and how they can be mitigated

    June 4, 2024

    Generative AI systems are made up of multiple components that interact to provide a rich user experience between the human and the AI model(s). As part of a responsible AI approach, AI models are protected by layers of defense mechanisms to prevent the production of harmful content or being used to carry out instructions that go ...

  • 5 Reasons to Attend Cyber Security & Cloud Congress North America 2024

    May 24, 2024

    Explore the forefront of enterprise technology at the Cyber Security & Cloud Congress North America. Delve into the entirety of the Cyber Security & Cloud Ecosystem and unravel the practical and triumphant application of Cyber Security & Cloud. Returning to North America on June 5-6, 2024, at the esteemed Santa Clara Convention Center, the globally renowned ...