iTool: Reinforced Fine-Tuning with Dynamic Deficiency Calibration for Advanced Tool Use

Yirong Zeng; Xiao Ding; Yuxian Wang; Weiwen Liu; Yutai Hou; Wu Ning; Xu Huang; Duyu Tang; Dandan Tu; Bing Qin (秦兵); Ting Liu (刘挺)

iTool: Reinforced Fine-Tuning with Dynamic Deficiency Calibration for Advanced Tool Use

Yirong Zeng, Xiao Ding, Yuxian Wang, Weiwen Liu, Yutai Hou, Wu Ning, Xu Huang, Duyu Tang, Dandan Tu, Bing Qin, Ting Liu

Abstract

Augmenting large language models (LLMs) with external tools is a promising approach to enhance their capabilities, especially for complex tasks. Synthesizing tool-use data through real-world simulations is an effective way to achieve this. However, our investigation reveals that training gains significantly decay as synthetic data increases. The model struggles to benefit from more synthetic data, and it can not equip the model with advanced tool-use capabilities in complex scenarios. Moreover, we discovered that the above limitation usually manifests as a fragment deficiency (i.e., parameter errors) in response. To this end, we propose an iterative reinforced fine-tuning strategy designed to alleviate this limitation. This strategy involves: (1) enhancing the diversity of response for synthetic data through path exploration of Monte Carlo Tree Search. (2) iteratively pinpointing the model’s deficiency by constructing fine-grained preference pairs, and then improving it by preference optimization algorithms for targeted improvement. The experiments show that our method achieves 13.11% better performance than the same-size base model. It achieves an improvement of 6.5% in complex scenarios compared to the baseline, and it also outperforms larger open-source and closed-source models.

Anthology ID:: 2025.emnlp-main.701
Volume:: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Month:: November
Year:: 2025
Address:: Suzhou, China
Editors:: Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng
Venue:: EMNLP
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 13901–13916
Language:
URL:: https://preview.aclanthology.org/ingest-emnlp/2025.emnlp-main.701/
DOI:
Bibkey:
Cite (ACL):: Yirong Zeng, Xiao Ding, Yuxian Wang, Weiwen Liu, Yutai Hou, Wu Ning, Xu Huang, Duyu Tang, Dandan Tu, Bing Qin, and Ting Liu. 2025. iTool: Reinforced Fine-Tuning with Dynamic Deficiency Calibration for Advanced Tool Use. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 13901–13916, Suzhou, China. Association for Computational Linguistics.
Cite (Informal):: iTool: Reinforced Fine-Tuning with Dynamic Deficiency Calibration for Advanced Tool Use (Zeng et al., EMNLP 2025)
Copy Citation:
PDF:: https://preview.aclanthology.org/ingest-emnlp/2025.emnlp-main.701.pdf
Checklist:: 2025.emnlp-main.701.checklist.pdf

PDF Cite Search Checklist Fix data