블로그 서비스 시스템을 위한 효과적인 중복문서의 검출 기법

An Efficient Method for Detecting Duplicated Documents in a Blog Service System

초록

Duplicate documents in blog service system are one of causes that deteriorate both of the quality and the performance of blog searches. Unlike the WWW environment, the creation of documents is reported every time in blog service system, which makes it possible to identify the original document from its duplicate documents. Based on this observation, this paper proposes a novel method for detecting duplication documents in blog service system. This method determines whether a document is original or not at the time it is stored in the blog service system. As a result, it solves the problem of duplicate documents retrieved in the search result by keeping those documents from being stored in the index for the blog search engine. This paper also proposes three indexing methods that preserve an accuracy of previous work, Min-hashing. We show most effective indexing method via extensive experiments using real-life blog data.

키워드

중복문서 검출블로그검색 엔진Duplicate document detectionBlogSearch engineDuplicate document detectionBlogSearch engine
제목
블로그 서비스 시스템을 위한 효과적인 중복문서의 검출 기법
제목 (타언어)
An Efficient Method for Detecting Duplicated Documents in a Blog Service System
저자
이상철이순행김상욱
발행일
2010-02
저널명
정보과학회논문지 : 데이타베이스
37
1
페이지
50 ~ 55