MetaScale: Test-Time Scaling with Evolving Meta-Thoughts

Qin Liu, Wenxuan Zhou, Nan Xu, James Y. Huang, Fei Wang, Sheng Zhang, Hoifung Poon, Muhao Chen


Abstract
One critical challenge for large language models (LLMs) in making complex reasoning is their reliance on matching reasoning patterns from training data, instead of proactively selecting the most appropriate cognitive strategy to solve a given task. Existing approaches impose fixed cognitive structures that enhance performance in specific tasks but lack adaptability across diverse scenarios. To address this limitation, we introduce MetaScale, a test-time scaling framework based on meta-thoughts, i.e., adaptive thinking strategies tailored to each task. MetaScale initializes a pool of candidate meta-thoughts, then iteratively selects and evaluates them using a multi-armed bandit algorithm with upper confidence bound selection, guided by a reward model. To further enhance adaptability, a genetic algorithm evolves high-reward meta-thoughts, refining and extending the strategy pool over time. By dynamically proposing and optimizing meta-thoughts at inference time, MetaScale improves both accuracy and generalization across a wide range of tasks. Experimental results demonstrate that MetaScale consistently outperforms standard inference approaches, achieving an 11% performance gain in win rate on Arena-Hard with GPT-4o, improving from 82.14% to 93.14% against GPT-4. Notably, MetaScale scales more effectively with increasing sampling budgets and produces more structured, expert-level responses.
Anthology ID:
2026.findings-acl.574
Volume:
Findings of the Association for Computational Linguistics: ACL 2026
Month:
July
Year:
2026
Address:
San Diego, California, United States
Editors:
Maria Liakata, Viviane P. Moreira, Jiajun Zhang, David Jurgens
Venue:
Findings
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
11828–11842
Language:
URL:
https://preview.aclanthology.org/ingest-acl/2026.findings-acl.574/
DOI:
Bibkey:
Cite (ACL):
Qin Liu, Wenxuan Zhou, Nan Xu, James Y. Huang, Fei Wang, Sheng Zhang, Hoifung Poon, and Muhao Chen. 2026. MetaScale: Test-Time Scaling with Evolving Meta-Thoughts. In Findings of the Association for Computational Linguistics: ACL 2026, pages 11828–11842, San Diego, California, United States. Association for Computational Linguistics.
Cite (Informal):
MetaScale: Test-Time Scaling with Evolving Meta-Thoughts (Liu et al., Findings 2026)
Copy Citation:
PDF:
https://preview.aclanthology.org/ingest-acl/2026.findings-acl.574.pdf
Checklist:
 2026.findings-acl.574.checklist.pdf