상세 보기
HTML 본문 추출을 위한 새로운 시각적 Feature
- 정근성;
- 차재혁
초록
Hypertext markup language (HTML) main content extraction is a technology that identifies the body and contents of an article from web pages. Traditional technologies use structural features, such as the tag structure of the HTML node and text features based on statistical properties. However, because these features depend on web development trends, language, and the region of the webpage, the performance of algorithms or models based on these features can vary. Therefore, in this study, we propose a novel visual feature to prevent the degradation of HTML body extraction performance on multilingual web pages. The feature is based on the results of HTML node attributes rendered in the browser; therefore, the influence of the language or region is relatively small. The Google TabNet deep neural network architecture was used to learn the neural network model based on only structural and text features, and subsequently another model with the newly introduced visual feature along with the structural and text features was trained. A comparison of the body extraction performance of the two models demonstrates the performance improvement provided by visual features in this study.
키워드
- 제목
- HTML 본문 추출을 위한 새로운 시각적 Feature
- 제목 (타언어)
- New Visual Features for HTML Main Content Extraction
- 저자
- 정근성; 차재혁
- 발행일
- 2023-04
- 저널명
- 디지털컨텐츠학회논문지
- 권
- 24
- 호
- 4
- 페이지
- 691 ~ 699