Large Language Models (LLMs) have exhibited remarkable reasoning capabilities, achieving impressive results across a wide range of tasks. Despite these advances, significant reasoning failures persist, occurring even in seemingly simple scenarios. To systematically understand and address these shortcomings, the authors of the paper present the first comprehensive survey dedicated to reasoning failures in LLMs.
The authors introduce a novel categorization framework that distinguishes reasoning into embodied and non-embodied types, with the latter further subdivided into informal (intuitive) and formal (logical) reasoning. In parallel, the authors classify reasoning failures along a complementary axis into three types: fundamental failures intrinsic to LLM architectures that broadly affect downstream tasks; application-specific limitations that manifest in particular domains; and robustness issues characterized by inconsistent performance across minor variations. For each reasoning failure, the authors provide a clear definition, analyze existing studies, explore root causes, and present mitigation strategies.
Read more…
Source: ARXIV, Cornell University
Sign up for the Cyber Security Review Newsletter
The latest cyber security news and insights delivered right to your inbox
Related:
- Cyber security in critical industries: challenges, solutions, and the road ahead
August 30, 2024
In an era of rapid digital transformation, cyber security has emerged as a paramount concern, particularly for critical industries such as energy, healthcare, and transportation. As we approach the IET’s Cyber Security for Critical Industries 2024 conference, it is essential to delve into the latest cyber security challenges and explore how building resilient and responsive ...
- Rogue AI is the Future of Cyber Threats
August 15, 2024
Yoshua Bengio, regarded as one of the “godfathers” of artificial intelligence, has likened the now-ubiquitous technology to a bear. When we teach the bear to become smart enough to escape its cage, we no longer control it. All we can do after that is try to build a better cage. This should be our goal with ...
- Indirect prompt injection in the real world: how people manipulate neural networks
August 12, 2024
Large language models (LLMs) – the neural network algorithms that underpin ChatGPT and other popular chatbots – are becoming ever more powerful and inexpensive. Systems built on instruction-executing LLMs may be vulnerable to prompt injection attacks. A prompt is a text description of a task that the system is to perform, for example: “You are a ...
- Elon Musk’s X accused of AI data grab in ‘blatant breach of law’
July 29, 2024
Privacy organisation Open Rights Group has branded a move by social network X (formerly Twitter) to use user data to train its Grok AI as a “blatant breach of GDPR”. The GDPR (General Data Protection Regulation) is an EU law protecting people’s data, and is implemented in the UK as part of the Data Protection Act ...
- CrowdStrike Took Down Australia And Half The World Now Facing Massive Compensation Claims
July 19, 2024
The reputation of a Company that describes themselves as one of the world’s best cyber security Companies is in tatters tonight, with the US business facing the potential of being sued by hundreds of business including major retailers in Australia and insurance Companies looking to claw back payouts for lost income, airline delays and customers ...
- AI, cyber-attacks and amateur experiments threaten to upend global biosecurity, WHO warns
July 13, 2024
Artificial intelligence, cyber-attacks and genetic engineering could pose potentially catastrophic biosecurity threats to countries around the world, the WHO has warned. Rapid technological advances in the past decade have “redefined the biological threat landscape” and heightened risks of manipulation, the updated guidance from the WHO’s Technical Advisory Group on Biosafety said. The report advised that member ...

