Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs

Qibin Wang; Pu Zhao; Shaohan Huang; Fangkai Yang; Lu Wang; Furu Wei; Qingwei Lin; Saravan Rajmohan; Dongmei Zhang

Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs

Qibin Wang, Pu Zhao, Shaohan Huang, Fangkai Yang, Lu Wang, Furu Wei, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang

Abstract

Test-time scaling (TTS) has gained widespread attention for enhancing LLM reasoning. Existing approaches such as Best-of-N and majority voting are limited as their performance depends on the quality of candidate responses, making them unable to produce a correct solution when all candidates are incorrect. Parallel self-refinement, generating multiple candidates and synthesizing a refined answer conditioned on them, offers a promising alternative, but the underlying mechanism driving its effectiveness remains obscure. To bridge this gap in understanding, we introduce a new metric, the Refinement Gap, designed to quantify the relative improvement of self-refinement beyond majority voting. We show that the Refinement Gap exhibits a clear scaling trend with model size and is only weakly correlated with the base capability. Based on this discovery, we propose Generative Self-Refinement (GSR), a parallel test-time scaling framework that transfers the refinement policy from larger teacher models with higher refinement gap into smaller students. Crucially, GSR jointly trains a single model to generate strong candidates and refine a better final answer based on these candidates. Experimental results demonstrate that our method achieves state-of-the-art performance across five mathematical benchmarks over other parallel aggregation methods, while the learned refinement skill transfers across multiple model scales and families and exhibits robust generalization to an out-of-distribution domain.

Anthology ID:: 2026.findings-acl.1291
Volume:: Findings of the Association for Computational Linguistics: ACL 2026
Month:: July
Year:: 2026
Address:: San Diego, California, United States
Editors:: Maria Liakata, Viviane P. Moreira, Jiajun Zhang, David Jurgens
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 25904–25921
Language:
URL:: https://preview.aclanthology.org/ingest-acl/2026.findings-acl.1291/
DOI:
Bibkey:
Cite (ACL):: Qibin Wang, Pu Zhao, Shaohan Huang, Fangkai Yang, Lu Wang, Furu Wei, Qingwei Lin, Saravan Rajmohan, and Dongmei Zhang. 2026. Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs. In Findings of the Association for Computational Linguistics: ACL 2026, pages 25904–25921, San Diego, California, United States. Association for Computational Linguistics.
Cite (Informal):: Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs (Wang et al., Findings 2026)
Copy Citation:
PDF:: https://preview.aclanthology.org/ingest-acl/2026.findings-acl.1291.pdf
Checklist:: 2026.findings-acl.1291.checklist.pdf

PDF Cite Search Checklist Fix data