Document Information Extraction via Global Tagging

He Shaojie, Wang Tianshu, Lu Yaojie, Lin Hongyu, Han Xianpei, Sun Yingfei, Sun Le


Abstract
“Document Information Extraction (DIE) is a crucial task for extracting key information fromvisually-rich documents. The typical pipeline approach for this task involves Optical Charac-ter Recognition (OCR), serializer, Semantic Entity Recognition (SER), and Relation Extraction(RE) modules. However, this pipeline presents significant challenges in real-world scenariosdue to issues such as unnatural text order and error propagation between different modules. Toaddress these challenges, we propose a novel tagging-based method – Global TaggeR (GTR),which converts the original sequence labeling task into a token relation classification task. Thisapproach globally links discontinuous semantic entities in complex layouts, and jointly extractsentities and relations from documents. In addition, we design a joint training loss and a jointdecoding strategy for SER and RE tasks based on GTR. Our experiments on multiple datasetsdemonstrate that GTR not only mitigates the issue of text in the wrong order but also improvesRE performance. Introduction”
Anthology ID:
2023.ccl-1.62
Volume:
Proceedings of the 22nd Chinese National Conference on Computational Linguistics
Month:
August
Year:
2023
Address:
Harbin, China
Editors:
Maosong Sun, Bing Qin, Xipeng Qiu, Jing Jiang, Xianpei Han
Venue:
CCL
SIG:
Publisher:
Chinese Information Processing Society of China
Note:
Pages:
726–735
Language:
English
URL:
https://aclanthology.org/2023.ccl-1.62
DOI:
Bibkey:
Cite (ACL):
He Shaojie, Wang Tianshu, Lu Yaojie, Lin Hongyu, Han Xianpei, Sun Yingfei, and Sun Le. 2023. Document Information Extraction via Global Tagging. In Proceedings of the 22nd Chinese National Conference on Computational Linguistics, pages 726–735, Harbin, China. Chinese Information Processing Society of China.
Cite (Informal):
Document Information Extraction via Global Tagging (Shaojie et al., CCL 2023)
Copy Citation:
PDF:
https://preview.aclanthology.org/nschneid-patch-4/2023.ccl-1.62.pdf