Reasoning Traces Shape Outputs but Models Won’t Say So

Yijie Hao; Lingjie Chen; Ali Emami; Joyce C. Ho

Reasoning Traces Shape Outputs but Models Won’t Say So

Yijie Hao, Lingjie Chen, Ali Emami, Joyce C. Ho

Abstract

Can we trust the reasoning traces that large reasoning models (LRMs) produce? We investigate whether these traces faithfully reflect what drives model outputs, and whether models will honestly report their influence. We introduce Thought Injection, a method that injects synthetic reasoning snippets into a model’s reasoning trace, then measures whether the model follows the injected reasoning and acknowledges doing so. Across 45,000 samples from three LRMs, we find that injected hints reliably alter outputs, confirming that reasoning traces causally shape model behavior. However, when asked to explain their changed answers, models overwhelmingly refuse to disclose the influence: non-disclosure exceeds 90% for extreme hints across 30,000 follow-up samples. Instead of acknowledging the injected reasoning, models fabricate aligned-appearing but unrelated explanations. Activation analysis reveals that sycophancy- and deception-related directions are strongly activated during these fabrications, suggesting systematic patterns rather than incidental failures. Our findings reveal a gap between the reasoning LRMs follow and the reasoning they report, raising concern that aligned-appearing explanations may not be equivalent to genuine alignment.

Anthology ID:: 2026.acl-long.1986
Volume:: Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Month:: July
Year:: 2026
Address:: San Diego, California, United States
Editors:: Maria Liakata, Viviane P. Moreira, Jiajun Zhang, David Jurgens
Venue:: ACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 42852–42878
Language:
URL:: https://preview.aclanthology.org/ingest-acl/2026.acl-long.1986/
DOI:
Bibkey:
Cite (ACL):: Yijie Hao, Lingjie Chen, Ali Emami, and Joyce C. Ho. 2026. Reasoning Traces Shape Outputs but Models Won’t Say So. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 42852–42878, San Diego, California, United States. Association for Computational Linguistics.
Cite (Informal):: Reasoning Traces Shape Outputs but Models Won’t Say So (Hao et al., ACL 2026)
Copy Citation:
PDF:: https://preview.aclanthology.org/ingest-acl/2026.acl-long.1986.pdf
Checklist:: 2026.acl-long.1986.checklist.pdf

PDF Cite Search Checklist Fix data