InstructCoder: Instruction Tuning Large Language Models for Code Editing

Kaixin Li; Qisheng Hu; James Zhao; Hui Chen; Yuxi Xie; Tiedong Liu; Michael Shieh; Junxian He

InstructCoder: Instruction Tuning Large Language Models for Code Editing

Kaixin Li, Qisheng Hu, James Zhao, Hui Chen, Yuxi Xie, Tiedong Liu, Michael Shieh, Junxian He

Abstract

Code editing encompasses a variety of pragmatic tasks that developers deal with daily. Despite its relevance and practical usefulness, automatic code editing remains an underexplored area in the evolution of deep learning models, partly due to data scarcity. In this work, we explore the use of Large Language Models (LLMs) to edit code based on user instructions. Evaluated on a novel human-written execution-based benchmark dubbed EditEval, we found current models often struggle to fulfill the instructions. In light of this, we contribute InstructCoder, the first instruction-tuning dataset designed to adapt LLMs for general-purpose code editing, containing high-diversity code-editing tasks such as comment insertion, code optimization, and code refactoring. It consists of over 114,000 instruction-input-output triplets and covers multiple distinct code editing scenarios. The collection process starts with filtered commit data sourced from GitHub Python repositories as seeds. Subsequently, the dataset is systematically expanded through an iterative process, where both seed and generated tasks are used to prompt ChatGPT for more data. Our findings reveal that open-source LLMs fine-tuned on InstructCoder can significantly enhance the accuracy of code edits, exhibiting superior code-editing performance matching advanced proprietary LLMs. The datasets and the source code are publicly available.

Anthology ID:: 2024.acl-srw.6
Volume:: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop)
Month:: August
Year:: 2024
Address:: Bangkok, Thailand
Editors:: Xiyan Fu, Eve Fleisig
Venue:: ACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 50–70
Language:
URL:: https://aclanthology.org/2024.acl-srw.6
DOI:
Bibkey:
Cite (ACL):: Kaixin Li, Qisheng Hu, James Zhao, Hui Chen, Yuxi Xie, Tiedong Liu, Michael Shieh, and Junxian He. 2024. InstructCoder: Instruction Tuning Large Language Models for Code Editing. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop), pages 50–70, Bangkok, Thailand. Association for Computational Linguistics.
Cite (Informal):: InstructCoder: Instruction Tuning Large Language Models for Code Editing (Li et al., ACL 2024)
Copy Citation:
PDF:: https://preview.aclanthology.org/nschneid-patch-4/2024.acl-srw.6.pdf

PDF Search