Large Language Models (LLMs) have exhibited remarkable reasoning capabilities, achieving impressive results across a wide range of tasks. Despite these advances, significant reasoning failures persist, occurring even in seemingly simple scenarios. To systematically understand and address these shortcomings, the authors of the paper present the first comprehensive survey dedicated to reasoning failures in LLMs.
The authors introduce a novel categorization framework that distinguishes reasoning into embodied and non-embodied types, with the latter further subdivided into informal (intuitive) and formal (logical) reasoning. In parallel, the authors classify reasoning failures along a complementary axis into three types: fundamental failures intrinsic to LLM architectures that broadly affect downstream tasks; application-specific limitations that manifest in particular domains; and robustness issues characterized by inconsistent performance across minor variations. For each reasoning failure, the authors provide a clear definition, analyze existing studies, explore root causes, and present mitigation strategies.
Read more…
Source: ARXIV, Cornell University
Sign up for the Cyber Security Review Newsletter
The latest cyber security news and insights delivered right to your inbox
Related:
- CISA, NSA, FBI warn Chinese AI firms of industrial‑scale knowledge distillation
September 9, 2026
Chinese AI companies’ core development strategy is to steal proprietary functionalities and capabilities from their US counterparts, law enforcement agencies have warned. The US Cybersecurity and Infrastructure Security Agency (CISA) has published a new security advisory, drafted jointly with the National Security Agency (NSA) and the Federal Bureau of Investigation (FBI), warning American AI companies about an ongoing “aggressive, malicious, ...
- OpenAI hid AI agent hijacking of German wiki forum for weeks
September 7, 2026
OpenAI recently disclosed the details of how one of its AI models escaped a sandboxed environment and attacked Hugging Face during an evaluation – and as part of the incident, the models created a messaging board to communicate with each other and influence each other’s reasoning. OpenAI has now disclosed that shortly after this incident, agents undergoing testing ...
- Attackers Expose Ongoing AI Tool Use Targeting Organizations in Latin America
September 3, 2026
We have analyzed two ongoing, multi-stage network intrusion and data-exfiltration campaigns targeting organizations in Latin America. Corroborating recent findings from the broader threat intelligence community, we observed attackers leveraging artificial intelligence (AI) to enhance their capabilities. Read more… Source: Palo Alto Unit 42 Sign up for the Cyber Security Review Newsletter The latest cyber security news and insights delivered ...
- Judge rules the Pentagon’s supply chain risk label for Anthropic unlawful
August 28, 2026
A federal judge ruled the Pentagon’s decision to label AI company Anthropic a ‘supply chain risk’ violated the law and ordered the designation be removed Thursday evening. Judge Rita Lin wrote that while the military should have wide latitude to decide which companies to work with, its actions against Anthropic “constituted unlawful retaliation in violation of ...
- An open letter for a global surge in cyber defense
August 27, 2026
We have a limited window to strengthen cyber defenses. In the coming months, AI-enabled cyber attacks will become far more widespread and sophisticated as models around the world become increasingly capable. The companies and public services our communities depend on—from hospitals to water treatment plants to the infrastructure that powers the internet—are at risk. Today’s AI advances ...
- Anthropic’s Mythos AI used social engineering to target real people
August 6, 2026
Anthropic’s Mythos AI agent, tested by the UK AI Safety Institute (AISI), has reportedly attempted a real‑world social‑engineering style hack against GitHub maintainers by creating fake human profiles, pressuring them to accept malicious code, and then editing logs to hide its tracks when challenged. AISI was running cybersecurity evaluations of Anthropic’s Mythos and OpenAI’s Sol when it detected ...

