Anthropic’s Mythos AI used social engineering to target real people


Anthropic’s Mythos AI agent, tested by the UK AI Safety Institute (AISI), has reportedly attempted a real‑world social‑engineering style hack against GitHub maintainers by creating fake human profiles, pressuring them to accept malicious code, and then editing logs to hide its tracks when challenged.

AISI was running cybersecurity evaluations of Anthropic’s Mythos and OpenAI’s Sol when it detected unusual outbound data transfers from its research systems. An investigation showed that some agents had engaged in “sustained, potentially harmful activity” targeting real people and organizations, rather than staying within the intended test environment.

The most serious activity involved an Anthropic Mythos agent tasked with solving a GitHub‑related cybersecurity challenge. The agent identified real GitHub maintainers, researched them, and created multiple fake accounts impersonating those individuals. Using private messages and a file‑sharing service, it tried to pressure and deceive the maintainers into approving malicious code so that it would run on GitHub’s systems.

Read more…
Source:  Malwarebytes Labs


Sign up for the Cyber Security Review Newsletter
The latest cyber security news and insights delivered right to your inbox


Related:

  • DHS Cybersecurity and Infrastructure Security Agency Releases Roadmap for Artificial Intelligence 

    November 14, 2023

    WASHINGTON – Today the Department of Homeland Security’s (DHS) Cybersecurity and Infrastructure Security Agency (CISA) released its first Roadmap for Artificial Intelligence (AI), adding to the significant DHS and broader whole-of-government effort to ensure the secure development and implementation of artificial intelligence capabilities. DHS plays a critical role in ensuring AI safety and security nationwide. Last ...

  • Threat Predictions for 2024: Chained AI and CaaS Operations Give Attackers More “Easy” Buttons Than Ever

    November 9, 2023

    With the growth of Cybercrime-as-a-Service (CaaS) operations and the advent of generative AI, threat actors have more “easy” buttons at their fingertips to assist with carrying out attacks than ever before. By relying on the growing capabilities in their respective toolboxes, adversaries will increase the sophistication of their activities. They’ll launch more targeted and stealthier hacks ...

  • OpenAI Blames ChatGPT’s Intermittent Outages On ‘Abnormal Traffic’ That Suggests Potential Cyber Attack

    November 9, 2023

    ChatGPT continued to face intermittent outages late Wednesday, which the platform’s maker OpenAI blamed on a potential cyberattack, hours after the AI chatbot platform recovered from a wide outage that the company initially attributed to a surge in interest for its new features. Early on Thursday, OpenAI’s service status page displayed a notification saying both ChatGPT ...

  • Tech firms to allow vetting of AI tools

    November 3, 2023

    The most advanced technology companies will allow governments to vet their artificial intelligence tools for the first time, Rishi Sunak has announced, as Elon Musk warned the technology could eventually replace all human jobs. Companies including Meta, Google DeepMind and OpenAI have agreed to allow regulators to test their latest AI products before releasing them ...

  • Increasing transparency in AI security

    October 26, 2023

    New AI innovations and applications are reaching consumers and businesses on an almost-daily basis. Building AI securely is a paramount concern, and we believe that Google’s Secure AI Framework (SAIF) can help chart a path for creating AI applications that users can trust. Today, we’re highlighting two new ways to make information about AI supply ...

  • Hackers trying to corrupt AI, raising level of ransomware threat

    October 17, 2023

    Cyber criminals are actively trying to corrupt generative artificial intelligence (AI), which may then put the ability to create ransomware in the hands of individuals. The looming threat is what keeps Mr Willis Lim, the director of the National Cyber Threat Analysis Centre at the Cyber Security Agency of Singapore (CSA), up at night. Generative ...