Large Language Model Reasoning Failures


Large Language Models (LLMs) have exhibited remarkable reasoning capabilities, achieving impressive results across a wide range of tasks. Despite these advances, significant reasoning failures persist, occurring even in seemingly simple scenarios. To systematically understand and address these shortcomings, the authors of the paper present the first comprehensive survey dedicated to reasoning failures in LLMs.

The authors introduce a novel categorization framework that distinguishes reasoning into embodied and non-embodied types, with the latter further subdivided into informal (intuitive) and formal (logical) reasoning. In parallel, the authors classify reasoning failures along a complementary axis into three types: fundamental failures intrinsic to LLM architectures that broadly affect downstream tasks; application-specific limitations that manifest in particular domains; and robustness issues characterized by inconsistent performance across minor variations. For each reasoning failure, the authors provide a clear definition, analyze existing studies, explore root causes, and present mitigation strategies.

Read more
Source: ARXIV, Cornell University


Sign up for the Cyber Security Review Newsletter
The latest cyber security news and insights delivered right to your inbox


Related:

  • US looking at media report that Israel used AI to identify bombing targets in Gaza

    April 5, 2024

    The United States was looking into a media report that the Israeli military has been using artificial intelligence to help identify bombing targets in Gaza, White House national security spokesperson John Kirby told CNN in an interview on Thursday. The report in +972 Magazine and Local Call was published on Wednesday. According to the 972+ report, ...

  • US and UK announce landmark agreement on artificial intelligence safety

    April 2, 2024

    The UK and US have signed a landmark deal to work together on testing advanced artificial intelligence (AI). The agreement signed on Monday says both countries will work together on developing “robust” methods for evaluating the safety of AI tools and the systems that underpin them. It is the first bilateral agreement of its kind. Read more… Source: ...

  • OpenAI’s new ‘Voice Engine’ clones your voice in only 15 seconds

    March 30, 2024

    As artificial intelligence (AI) continues to advance rapidly, ChatGPT maker OpenAI is at the forefront of this progress. The research lab has unveiled a powerful new voice cloning technology called Voice Engine. With just a 15-second audio sample, it can generate a synthetic copy of a person’s voice described as “natural-sounding” and “emotive.” While the company ...

  • UN General Assembly adopts landmark resolution on artificial intelligence

    March 21, 2024

    The UN General Assembly on Thursday adopted a landmark resolution on the promotion of “safe, secure and trustworthy” artificial intelligence (AI) systems that will also benefit sustainable development for all. The Assembly called on all Member States and stakeholders “to refrain from or cease the use of artificial intelligence systems that are impossible to operate in ...

  • DIANA, NATO’s innovation accelerator, doubles the size of its transatlantic network

    March 14, 2024

    On Thursday (14 March 2024), NATO’s Defence Innovation Accelerator for the North Atlantic (DIANA) announced a major expansion of its transatlantic network of accelerator sites and test centres. DIANA’s network will now comprise 23 accelerator sites (up from 11) and 182 test centres (up from 90) in 28 Allied countries, augmenting DIANA’s capacity to support innovators ...

  • EU passes landmark AI act, paving the way for greater AI regulation

    March 13, 2024

    The European Parliament has passed its long awaited AI act that it hopes will provide the legal infrastructure for regulating artificial intelligence. While AI has contributed massively to increases in productivity and has resulted in major innovations in critical industries such as science and healthcare, many fear that the speed of its development may be outstripping ...