Effective SNOMED-CT Concept Classification from Natural Language using Knowledge Distillation

  • Kim, HyunJoo
  • Joe, Inwhee
Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Recently, as natural language processing (NLP) methods have been developed a lot, research on predicting rule languages such as medical terminology systems such as Systemized Nomenclature of Medicine Clinical Term (SNOMED-CT) in natural language is becoming active. In this paper, we propose a prediction model with the SNOMED-CT code using medical natural language. Thus, natural language is encoded with the existing pre-trained model and used as a teacher model to learn a lightweight student model using knowledge distillation techniques. To improve the performance of the model, augmented data are used for learning with the augmentation technique, and performance improvement is attempted using the Teacher model in the same domain with BioBert.When the Teacher model used BioBert and the Student model used simple LSTM, the distillation results obtained an accuracy of 0.86. Only the Teacher model fine-tuned with the pre-trained model obtained a result of 0.88, but the result of the only student model, which is a simple LSTM, was 0.8695. Although we did not obtain a student model with better performance than the Teacher; however, it is useful to interpret the language of SNOMDE-CT by applying knowledge distillation (KD) to the NLP.

키워드

BioBertClassificationKnowledge distillationSNOMED-CT
제목
Effective SNOMED-CT Concept Classification from Natural Language using Knowledge Distillation
저자
Kim, HyunJooJoe, Inwhee
DOI
10.1007/978-3-031-21438-7_4
발행일
2023-01
유형
Proceedings Paper
저널명
Lecture Notes in Networks and Systems
597 LNNS
페이지
54 ~ 64