ERNIE-Gram: Pre-Training with Explicitly N-Gram Masked Language Modeling for Natural Language Understanding

Dongling Xiao; Yu-Kun Li; Han Zhang; Yu Sun; Hao Tian; Hua Wu; Haifeng Wang

doi:10.18653/v1/2021.naacl-main.136

ERNIE-Gram: Pre-Training with Explicitly N-Gram Masked Language Modeling for Natural Language Understanding

Dongling Xiao, Yu-Kun Li, Han Zhang, Yu Sun, Hao Tian, Hua Wu, Haifeng Wang

Abstract

Coarse-grained linguistic information, such as named entities or phrases, facilitates adequately representation learning in pre-training. Previous works mainly focus on extending the objective of BERT’s Masked Language Modeling (MLM) from masking individual tokens to contiguous sequences of n tokens. We argue that such contiguously masking method neglects to model the intra-dependencies and inter-relation of coarse-grained linguistic information. As an alternative, we propose ERNIE-Gram, an explicitly n-gram masking method to enhance the integration of coarse-grained information into pre-training. In ERNIE-Gram, n-grams are masked and predicted directly using explicit n-gram identities rather than contiguous sequences of n tokens. Furthermore, ERNIE-Gram employs a generator model to sample plausible n-gram identities as optional n-gram masks and predict them in both coarse-grained and fine-grained manners to enable comprehensive n-gram prediction and relation modeling. We pre-train ERNIE-Gram on English and Chinese text corpora and fine-tune on 19 downstream tasks. Experimental results show that ERNIE-Gram outperforms previous pre-training models like XLNet and RoBERTa by a large margin, and achieves comparable results with state-of-the-art methods. The source codes and pre-trained models have been released at https://github.com/PaddlePaddle/ERNIE.

Anthology ID:: 2021.naacl-main.136
Volume:: Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
Month:: June
Year:: 2021
Address:: Online
Venue:: NAACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 1702–1715
Language:
URL:: https://aclanthology.org/2021.naacl-main.136
DOI:: 10.18653/v1/2021.naacl-main.136
Bibkey:
Cite (ACL):: Dongling Xiao, Yu-Kun Li, Han Zhang, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. 2021. ERNIE-Gram: Pre-Training with Explicitly N-Gram Masked Language Modeling for Natural Language Understanding. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1702–1715, Online. Association for Computational Linguistics.
Cite (Informal):: ERNIE-Gram: Pre-Training with Explicitly N-Gram Masked Language Modeling for Natural Language Understanding (Xiao et al., NAACL 2021)
Copy Citation:
PDF:: https://preview.aclanthology.org/auto-file-uploads/2021.naacl-main.136.pdf
Video:: https://preview.aclanthology.org/auto-file-uploads/2021.naacl-main.136.mp4
Code: PaddlePaddle/PaddleNLP + additional community code
Data: CMRC, CMRC 2018, CoLA, DRCD, DuReader, GLUE, IMDb Movie Reviews, MRPC, MultiNLI, QNLI, RACE, SQuAD, SST

PDF Search Code Video