ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning Tasks

Heng Zhou; Hejia Geng; Xiangyuan Xue; Li Kang; Yiran Qin; Zhiyong Wang; Zhenfei Yin; Lei Bai

ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning Tasks

Heng Zhou, Hejia Geng, Xiangyuan Xue, Li Kang, Yiran Qin, Zhiyong Wang, Zhenfei Yin, Lei Bai

Abstract

Multi-agent systems have emerged as a promising approach for enhancing the reasoning capabilities of large language models in complex problem-solving. However, current MAS frameworks are limited by poor flexibility and scalability, with underdeveloped optimization strategies. To address these challenges, we propose ReSo, which integrates task graph generation with a reward-driven two-stage agent selection process. The core of ReSo is the proposed Collaborative Reward Model, which can provide fine-grained reward signals for MAS cooperation for optimization. We also introduce an automated data synthesis framework for generating MAS benchmarks, without human annotations. Experimentally, ReSo matches or outperforms existing methods. ReSo achieves 33.7% and 32.3% accuracy on Math-MAS and SciBench-MAS SciBench, while other methods completely fail. The code and data are available at [Reso](https://github.com/hengzzzhou/ReSo).

Anthology ID:: 2025.emnlp-main.808
Volume:: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Month:: November
Year:: 2025
Address:: Suzhou, China
Editors:: Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng
Venue:: EMNLP
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 15990–16009
Language:
URL:: https://preview.aclanthology.org/ingest-emnlp/2025.emnlp-main.808/
DOI:
Bibkey:
Cite (ACL):: Heng Zhou, Hejia Geng, Xiangyuan Xue, Li Kang, Yiran Qin, Zhiyong Wang, Zhenfei Yin, and Lei Bai. 2025. ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning Tasks. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 15990–16009, Suzhou, China. Association for Computational Linguistics.
Cite (Informal):: ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning Tasks (Zhou et al., EMNLP 2025)
Copy Citation:
PDF:: https://preview.aclanthology.org/ingest-emnlp/2025.emnlp-main.808.pdf
Checklist:: 2025.emnlp-main.808.checklist.pdf

PDF Cite Search Checklist Fix data