Comparative Analysis of Text Classification Approaches in Electronic Health Records

Aurelie Mascio, Zeljko Kraljevic, Daniel Bean, Richard Dobson, Robert Stewart, Rebecca Bendayan, Angus Roberts


Abstract
Text classification tasks which aim at harvesting and/or organizing information from electronic health records are pivotal to support clinical and translational research. However these present specific challenges compared to other classification tasks, notably due to the particular nature of the medical lexicon and language used in clinical records. Recent advances in embedding methods have shown promising results for several clinical tasks, yet there is no exhaustive comparison of such approaches with other commonly used word representations and classification models. In this work, we analyse the impact of various word representations, text pre-processing and classification algorithms on the performance of four different text classification tasks. The results show that traditional approaches, when tailored to the specific language and structure of the text inherent to the classification task, can achieve or exceed the performance of more recent ones based on contextual embeddings such as BERT.
Anthology ID:
2020.bionlp-1.9
Volume:
Proceedings of the 19th SIGBioMed Workshop on Biomedical Language Processing
Month:
July
Year:
2020
Address:
Online
Editors:
Dina Demner-Fushman, Kevin Bretonnel Cohen, Sophia Ananiadou, Junichi Tsujii
Venue:
BioNLP
SIG:
SIGBIOMED
Publisher:
Association for Computational Linguistics
Note:
Pages:
86–94
Language:
URL:
https://aclanthology.org/2020.bionlp-1.9
DOI:
10.18653/v1/2020.bionlp-1.9
Bibkey:
Cite (ACL):
Aurelie Mascio, Zeljko Kraljevic, Daniel Bean, Richard Dobson, Robert Stewart, Rebecca Bendayan, and Angus Roberts. 2020. Comparative Analysis of Text Classification Approaches in Electronic Health Records. In Proceedings of the 19th SIGBioMed Workshop on Biomedical Language Processing, pages 86–94, Online. Association for Computational Linguistics.
Cite (Informal):
Comparative Analysis of Text Classification Approaches in Electronic Health Records (Mascio et al., BioNLP 2020)
Copy Citation:
PDF:
https://aclanthology.org/2020.bionlp-1.9.pdf
Dataset:
 2020.bionlp-1.9.Dataset.pdf
Video:
 http://slideslive.com/38929648
Data
MIMIC-III