Superficial Reflection or Genuine Thought? A Fine-Grained Cognitive Analysis of Large Reasoning Models
Accepted in Findings of AACL-IJCNLP 2026 (to appear).
Recommended citation: Yuxiang Chen*, Zuohan Wu*, Ziwei Wang, Xiangning Yu, Xujia Li, Linyi Yang, Mengyue Yang, Jun Wang, Lei Chen. "Superficial Reflection or Genuine Thought? A Fine-Grained Cognitive Analysis of Large Reasoning Models." Findings of AACL-IJCNLP 2026, accepted. *Equal contribution.
Abstract
Motivated by the observed human-like behaviours in Large Reasoning Models (LRMs), this paper introduces a comprehensive taxonomy to characterise atomic reasoning steps and analyse the reasoning behaviours of LRMs. Grounded in human cognitive processes, we propose a taxonomy comprising five groups and seventeen categories. Through this taxonomy, we conduct an in-depth analysis of contemporary LRMs and distil four actionable takeaways for model optimisation. Most notably, we reveal that prevailing post-answer “double-checks” are largely superficial and rarely yield substantive revisions. A targeted intervention further shows that explicitly eliciting richer reflection processes can substantially improve failed self-correction. To support this large-scale study, we propose CAPO, an automated annotation method used to construct a dataset of 277,534 reasoning steps with strong agreement with human expert annotations. We further validate the main behavioural patterns on a newer reasoning model and a coding domain, demonstrating the broader applicability of the proposed taxonomy. All source code and data are available at github.com/hehepig4/psyche.
Key Contributions
- A taxonomy with five groups and seventeen categories for analysing observable reasoning behaviours, using human cognition as a descriptive framework.
- CAPO, an automated annotation method supporting a corpus of 277,534 reasoning steps.
- An analysis of information organisation, analogy and hypothesis generation, reflection, and redundancy, including a targeted intervention to improve failed self-correction.
- Validation of the main behavioural patterns on a newer reasoning model and a coding domain.
Publication Details
- Venue: Findings of AACL-IJCNLP 2026
- Year: 2026
- Status: Accepted; to appear
- Paper: Camera-ready PDF
- Code and data: GitHub
- Earlier version: arXiv:2512.00729, originally titled Probing the “Psyche” of Large Reasoning Models: Understanding Through a Human Lens (2025).
Authors
Yuxiang Chen*, Zuohan Wu*, Ziwei Wang, Xiangning Yu, Xujia Li, Linyi Yang, Mengyue Yang, Jun Wang, Lei Chen
*Equal contribution.
BibTeX
@inproceedings{chen2026cognitiveanalysis,
author = {Yuxiang Chen and
Zuohan Wu and
Ziwei Wang and
Xiangning Yu and
Xujia Li and
Linyi Yang and
Mengyue Yang and
Jun Wang and
Lei Chen},
title = {Superficial Reflection or Genuine Thought? A Fine-Grained Cognitive Analysis of Large Reasoning Models},
booktitle = {Findings of AACL-IJCNLP 2026},
year = {2026},
note = {Accepted, to appear},
url = {https://hehepig4.github.io/files/2026-AACL-CognitiveAnalysis.pdf}
}
