Extracting the Main Content of Web Pages Using the First Impression Area

Citations

WEB OF SCIENCE

3
Citations

SCOPUS

6

초록

Extracting the main content from a web page is essential in various applications such as web crawlers and browser reader modes. Existing extraction methods using text-based algorithms and features for English text can be ineffective for non-English web pages. This study proposes a main content extraction method that obtains visual and structural features from the rendered web page. Our method uses the first impression area (FIA), a part of a web page that users initially view. In this area, websites have applied many techniques that enable users to find the main content easily. Using the non-textual properties in the FIA, our method selects three points with high content area density and expands the area from each point until it meets several structural and visual-based conditions. We evaluated our method, browsers’ (Mozilla Firefox and Google Chrome) reader modes, and existing main content extraction methods on multilingual datasets using two measures: Longest Common Subsequences and matched text blocks. The results showed that our method performed better than other methods in both English (up to 46%, matched text blocks F0.5) and non-English (up to 42%, matched text blocks F0.5) web pages.

키워드

Boilerplate removalmain content extractionweb content extractionweb miningweb segmentationblock detectionEYE-MOVEMENTSEGMENTATIONPERCEPTIONSATTENTION
제목
Extracting the Main Content of Web Pages Using the First Impression Area
저자
Jung, GeunseongHan, SungjaeKim, HansungKim, kwangukCha, Jaehyuk
DOI
10.1109/ACCESS.2022.3229080
발행일
2022-12
유형
Article in Press
저널명
IEEE Access
10
페이지
129958 ~ 129969

파일 다운로드