상세 보기
초록
Optical character recognition (OCR) is the technology that recognizes text in an image and converts it into text data. In foreign countries, OCR enables automated document processing. Since the recognition rate of Hangul is lower than that of English and Numbers, the OCR is not widely used in Korea. If the OCR accuracy of Hangul is improved, we expect an increase in work efficiency through OCR in Korea as well. In this paper, the OCR system was based on the convolutional neural network (CNN) to train Hangul, English, and Numbers. Subsequently, the process was implemented that distinguishes the complex words to complete Hangul characters, recognizes the complete Hangul characters, and converts them into text data. Additionally, to further improve the accuracy of the OCR system, search the text data in a web search engine, and verify the existence of modified words. If a modified word is found in the web search results, it is considered the correct recognition result and included in the final text data. We conducted a recognition rate measurement and found that the OCR system was able to accurately recognize up to 90.1% of characters in documents containing Hangul, English, and Numbers.
키워드
- 제목
- 웹 검색엔진 및 딥러닝 기반 한글 단어 인식 OCR 시스템
- 제목 (타언어)
- The Deep Learning-Based OCR System for Korean Word with Web Search Engine
- 저자
- 장혁수; 고상호; 이재현; 박승권
- 발행일
- 2023-09
- 저널명
- 한국통신학회논문지
- 권
- 48
- 호
- 9
- 페이지
- 1169 ~ 1174