웹 검색엔진 및 딥러닝 기반 한글 단어 인식 OCR 시스템

The Deep Learning-Based OCR System for Korean Word with Web Search Engine
  • 장혁수
  • 고상호
  • 이재현
  • 박승권
Citations

SCOPUS

1

초록

Optical character recognition (OCR) is the technology that recognizes text in an image and converts it into text data. In foreign countries, OCR enables automated document processing. Since the recognition rate of Hangul is lower than that of English and Numbers, the OCR is not widely used in Korea. If the OCR accuracy of Hangul is improved, we expect an increase in work efficiency through OCR in Korea as well. In this paper, the OCR system was based on the convolutional neural network (CNN) to train Hangul, English, and Numbers. Subsequently, the process was implemented that distinguishes the complex words to complete Hangul characters, recognizes the complete Hangul characters, and converts them into text data. Additionally, to further improve the accuracy of the OCR system, search the text data in a web search engine, and verify the existence of modified words. If a modified word is found in the web search results, it is considered the correct recognition result and included in the final text data. We conducted a recognition rate measurement and found that the OCR system was able to accurately recognize up to 90.1% of characters in documents containing Hangul, English, and Numbers.

키워드

광학 문자 인식딥러닝합성곱 신경망한글 단어 인식단어 분리OCRDeep LearningCNNKorean Word RecognitionWord Segmentation
제목
웹 검색엔진 및 딥러닝 기반 한글 단어 인식 OCR 시스템
제목 (타언어)
The Deep Learning-Based OCR System for Korean Word with Web Search Engine
저자
장혁수고상호이재현박승권
DOI
10.7840/kics.2023.48.9.1169
발행일
2023-09
저널명
한국통신학회논문지
48
9
페이지
1169 ~ 1174