Training Set Expansion Using Word Embeddings for Korean Medical Information Extraction

Citations

SCOPUS

1

초록

Entity recognition is an essential part of a task-oriented dialogue system and is considered as a sequence labeling task. However, constructing a training set in a new domain is extremely expensive and time-consuming. In this work, we propose a simple framework to exploit neural word embeddings in a semi-supervised manner to annotate medical named entities in Korean. The target domain is the automatic medical diagnosis, where disease name, symptom, and body part are defined as the entity types. Different aspects of the word embeddings such as embedding dimension, window size, models are examined to investigate their effects on the final performance. An online medical QA data has been used for the experiments. With a limit number of pre-annotated words, our framework could successfully expand the training set.

키워드

Medical information extractionTraining setWord embeddingsKoreanArtificial intelligenceBioinformaticsDiagnosisEmbeddingsHealth careInformation retrievalSpeech processingDialogue systemsEmbedding dimensionsEntity recognitionKoreanSemi-supervisedSequence LabelingTraining set expansionTraining setsInformation management
제목
Training Set Expansion Using Word Embeddings for Korean Medical Information Extraction
저자
Kim, Young min
DOI
10.1007/978-3-030-33752-0_19
발행일
2019-08
유형
Conference Paper
저널명
Lecture Notes in Computer Science
11721
페이지
261 ~ 274