Abstract
In this article we present a system that extracts information from pathology reports. The reports are written in Norwegian and contain free text describing prostate biopsies. Currently, these reports are manually coded for research and statistical purposes by trained experts at the Cancer Registry of Norway where the coders extract values for a set of predefined fields that are specific for prostate cancer. The presented system is rule based and achieves an average F-score of 0.91 for the fields Gleason grade, Gleason score, the number of biopsies that contain tumor tissue, and the orientation of the biopsies. The system also identifies reports that contain ambiguity or other content that should be reviewed by an expert. The system shows potential to encode the reports considerably faster, with less resources, and similar high quality to the manual encoding.- Anthology ID:
- R17-1100
- Volume:
- Proceedings of the International Conference Recent Advances in Natural Language Processing, RANLP 2017
- Month:
- September
- Year:
- 2017
- Address:
- Varna, Bulgaria
- Venue:
- RANLP
- SIG:
- Publisher:
- INCOMA Ltd.
- Note:
- Pages:
- 778–783
- Language:
- URL:
- https://doi.org/10.26615/978-954-452-049-6_100
- DOI:
- 10.26615/978-954-452-049-6_100
- Cite (ACL):
- Rebecka Weegar, Jan F Nygård, and Hercules Dalianis. 2017. Efficient Encoding of Pathology Reports Using Natural Language Processing. In Proceedings of the International Conference Recent Advances in Natural Language Processing, RANLP 2017, pages 778–783, Varna, Bulgaria. INCOMA Ltd..
- Cite (Informal):
- Efficient Encoding of Pathology Reports Using Natural Language Processing (Weegar et al., RANLP 2017)
- PDF:
- https://doi.org/10.26615/978-954-452-049-6_100