Superficial Reflection or Genuine Thought? A Fine-Grained Cognitive Analysis of Large Reasoning Models

Accepted in Findings of AACL-IJCNLP 2026 (to appear).

Recommended citation: Yuxiang Chen*, Zuohan Wu*, Ziwei Wang, Xiangning Yu, Xujia Li, Linyi Yang, Mengyue Yang, Jun Wang, Lei Chen. "Superficial Reflection or Genuine Thought? A Fine-Grained Cognitive Analysis of Large Reasoning Models." Findings of AACL-IJCNLP 2026, accepted. *Equal contribution.

Abstract

Motivated by the observed human-like behaviours in Large Reasoning Models (LRMs), this paper introduces a comprehensive taxonomy to characterise atomic reasoning steps and analyse the reasoning behaviours of LRMs. Grounded in human cognitive processes, we propose a taxonomy comprising five groups and seventeen categories. Through this taxonomy, we conduct an in-depth analysis of contemporary LRMs and distil four actionable takeaways for model optimisation. Most notably, we reveal that prevailing post-answer “double-checks” are largely superficial and rarely yield substantive revisions. A targeted intervention further shows that explicitly eliciting richer reflection processes can substantially improve failed self-correction. To support this large-scale study, we propose CAPO, an automated annotation method used to construct a dataset of 277,534 reasoning steps with strong agreement with human expert annotations. We further validate the main behavioural patterns on a newer reasoning model and a coding domain, demonstrating the broader applicability of the proposed taxonomy. All source code and data are available at github.com/hehepig4/psyche.

Key Contributions

  • A taxonomy with five groups and seventeen categories for analysing observable reasoning behaviours, using human cognition as a descriptive framework.
  • CAPO, an automated annotation method supporting a corpus of 277,534 reasoning steps.
  • An analysis of information organisation, analogy and hypothesis generation, reflection, and redundancy, including a targeted intervention to improve failed self-correction.
  • Validation of the main behavioural patterns on a newer reasoning model and a coding domain.

Publication Details

  • Venue: Findings of AACL-IJCNLP 2026
  • Year: 2026
  • Status: Accepted; to appear
  • Paper: Camera-ready PDF
  • Code and data: GitHub
  • Earlier version: arXiv:2512.00729, originally titled Probing the “Psyche” of Large Reasoning Models: Understanding Through a Human Lens (2025).

Authors

Yuxiang Chen*, Zuohan Wu*, Ziwei Wang, Xiangning Yu, Xujia Li, Linyi Yang, Mengyue Yang, Jun Wang, Lei Chen

*Equal contribution.

BibTeX

@inproceedings{chen2026cognitiveanalysis,
  author       = {Yuxiang Chen and
                  Zuohan Wu and
                  Ziwei Wang and
                  Xiangning Yu and
                  Xujia Li and
                  Linyi Yang and
                  Mengyue Yang and
                  Jun Wang and
                  Lei Chen},
  title        = {Superficial Reflection or Genuine Thought? A Fine-Grained Cognitive Analysis of Large Reasoning Models},
  booktitle    = {Findings of AACL-IJCNLP 2026},
  year         = {2026},
  note         = {Accepted, to appear},
  url          = {https://hehepig4.github.io/files/2026-AACL-CognitiveAnalysis.pdf}
}