IR&TM-NJUST@CLSciSumm 20

Heng Zhang, Lifan Liu, Ruping Wang, Shaohu Hu, Shutian Ma, Chengzhi Zhang


Abstract
This paper mainly introduces our methods for Task 1A and Task 1B of CL-SciSumm 2020. Task 1A is to identify reference text in reference paper. Traditional machine learning models and MLP model are used. We evaluate the performances of these models and submit the final results from the optimal model. Compared with previous work, we optimize the ratio of positive to negative examples after data sampling. In order to construct features for classification, we calculate similarities between reference text and candidate sentences based on sentence vectors. Accordingly, nine similarities are used, of which eight are chosen from what we used in CL-SciSumm 2019 and a new sentence similarity based on fastText is added. Task 1B is to classify the facets of reference text. Unlike the methods used in CL-SciSumm 2019, we construct inputs of models based on word vectors and add deep learning models for classification this year.
Anthology ID:
2020.sdp-1.33
Volume:
Proceedings of the First Workshop on Scholarly Document Processing
Month:
November
Year:
2020
Address:
Online
Venues:
EMNLP | sdp
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
288–296
Language:
URL:
https://aclanthology.org/2020.sdp-1.33
DOI:
10.18653/v1/2020.sdp-1.33
Bibkey:
Cite (ACL):
Heng Zhang, Lifan Liu, Ruping Wang, Shaohu Hu, Shutian Ma, and Chengzhi Zhang. 2020. IR&TM-NJUST@CLSciSumm 20. In Proceedings of the First Workshop on Scholarly Document Processing, pages 288–296, Online. Association for Computational Linguistics.
Cite (Informal):
IR&TM-NJUST@CLSciSumm 20 (Zhang et al., sdp 2020)
Copy Citation:
PDF:
https://preview.aclanthology.org/update-css-js/2020.sdp-1.33.pdf