Large Language Model Reasoning Failures


Large Language Models (LLMs) have exhibited remarkable reasoning capabilities, achieving impressive results across a wide range of tasks. Despite these advances, significant reasoning failures persist, occurring even in seemingly simple scenarios. To systematically understand and address these shortcomings, the authors of the paper present the first comprehensive survey dedicated to reasoning failures in LLMs.

The authors introduce a novel categorization framework that distinguishes reasoning into embodied and non-embodied types, with the latter further subdivided into informal (intuitive) and formal (logical) reasoning. In parallel, the authors classify reasoning failures along a complementary axis into three types: fundamental failures intrinsic to LLM architectures that broadly affect downstream tasks; application-specific limitations that manifest in particular domains; and robustness issues characterized by inconsistent performance across minor variations. For each reasoning failure, the authors provide a clear definition, analyze existing studies, explore root causes, and present mitigation strategies.

Read more
Source: ARXIV, Cornell University


Sign up for the Cyber Security Review Newsletter
The latest cyber security news and insights delivered right to your inbox


Related:

  • China’s cyber security association sets up special committee to bolster AI research

    October 15, 2023

    China has set up a professional committee focusing on governance of artificial intelligence (AI) security in a bid to build a sustained foundation for the sound development of the emerging industry, according to the country’s cyber security association. On Thursday, an inaugural meeting was held in Beijing for the AI security governance committee under the ...

  • Analysis of Generative AI Trends and ChatGPT Usage

    September 26, 2023

    The release of ChatGPT underscores the potential of artificial intelligence to revolutionize the daily operations of organizations. This paradigm shift is compelling businesses to reevaluate their conventional approaches and embrace the transformative capabilities offered by AI. Among the noteworthy facets of AI’s evolution, Large Language Models (LLMs) have emerged as a dominant force, reshaping user interactions ...

  • China to impose severe punishment on crimes of cyberbullying, defamation offenses, fabricating sexual topics

    September 25, 2023

    China on Monday released guidelines to severely punish cyberspace violations that target minors, involve paid posters, fabricate “sexual” topics and use artificial intelligence to disseminate illegal information. The guidelines on punishing crimes of cyberspace violence in accordance with laws were jointly issued by China’s Supreme People’s Court, China’s Supreme People’s Procuratorate and China’s Ministry of Public ...

  • Microsoft AI researchers accidentally exposed terabytes of internal sensitive data

    September 18, 2023

    Microsoft AI researchers accidentally exposed tens of terabytes of sensitive data, including private keys and passwords, while publishing a storage bucket of open source training data on GitHub. In research shared with TechCrunch, cloud security startup Wiz said it discovered a GitHub repository belonging to Microsoft’s AI research division as part of its ongoing work ...

  • Ukraine war: Cyber-teams fight a high-tech war on front lines

    September 6, 2023

    Ukraine cyber-operators are being deployed on the front lines of the war, duelling close-up with their Russian counterparts in a new kind of high-tech battle. “We have people who are directly involved in combat,” says Illia Vitiuk, the head of the Ukrainian Security Service’s (SBU) cyber department. Speaking inside the heavily protected SBU headquarters, he explains ...

  • AI and the Five Phases of the Threat Intelligence Lifecycle

    August 24, 2023

    Artificial intelligence (AI) and large language models (LLMs) can help threat intelligence teams to detect and understand novel threats at scale, reduce burnout-inducing toil, and grow their existing talent by democratizing access to subject matter expertise. However, broad access to foundational Open Source Intelligence (OSINT) data and AI/ML technologies has quickly led to an overwhelming amount ...