AI jailbreaks: What they are and how they can be mitigated


Generative AI systems are made up of multiple components that interact to provide a rich user experience between the human and the AI model(s).

As part of a responsible AI approach, AI models are protected by layers of defense mechanisms to prevent the production of harmful content or being used to carry out instructions that go against the intended purpose of the AI integrated application. This blog will provide an understanding of what AI jailbreaks are, why generative AI is susceptible to them, and how you can mitigate the risks and harms.

Read more…
Source: Microsoft


Sign up for our Newsletter


Related:

  • DARPA’s CASTLE to Fortify Computer Networks

    October 24, 2022

    An ever-expanding cyber-attack surface, infrequent computer vulnerability scans, and burdensome security procedures create a seemingly lopsided battle when it comes to defending critical computing assets. Couple those factors with costly cybersecurity assessments that often lack actionable feedback, and the odds may appear to favor bad actors. DARPA intends to change that dynamic through a new program ...

  • Phishing works so well crims won’t bother with deepfakes, says Sophos chap

    October 17, 2022

    Panic over the risk of deepfake scams is completely overblown, according to a senior security adviser for UK-based infosec company Sophos. “The thing with deepfakes is that we aren’t seeing a lot of it,” Sophos researcher John Shier told El Reg last week. Shier said current deepfakes – AI generated videos that mimic humans – aren’t the ...

  • UK privacy watchdog fines Clearview AI £7.5m and orders UK data to be deleted

    May 24, 2022

    The Information Commissioner’s Office (ICO) has fined controversial facial recognition company Clearview AI £7.5 million ($9.4 million) for breaching UK data protection laws and has issued an enforcement notice ordering the company to stop obtaining and using data of UK residents, and to delete the data from its systems. In its finding, the ICO detailed how ...

  • Artificial Intelligence: How to make Machine Learning Cyber Secure?

    December 14, 2021

    Machine learning (ML) is currently the most developed and the most promising subfield of artificial intelligence for industrial and government infrastructures. By providing new opportunities to solve decision-making problems intelligently and automatically, artificial intelligence (AI) is applied in almost all sectors of our economy. While the benefits of AI are significant and undeniable, the development of ...

  • UK spy chief warns China, Russia racing to master AI

    November 30, 2021

    The chief of the United Kingdom’s foreign spy service is to warn that China and Russia are racing to master artificial intelligence in a way that could revolutionise geopolitics over the next 10 years. Richard Moore, who heads the Secret Intelligence Service, known as MI6, is due to make his first public speech since becoming chief ...

  • Database firm Clearview AI told to remove photos taken in Australia

    November 3, 2021

    A firm which claims to have a database of more than 10 billion facial images has breached Australia’s privacy laws, Australian regulators say. Clearview AI lets law enforcement agencies search its database of faces. But the Office of the Australian Information Commissioner (OAIC) ordered it to stop collecting photos taken in Australia and remove ones already in ...