Combining a segmentation-like approach and a density-based approach in content extraction

Shuang Lin, Jie Chen, Zhendong Niu*

*此作品的通讯作者

科研成果: 期刊稿件文章同行评审

5 引用 (Scopus)

摘要

Density-based approaches in content extraction, whose task is to extract contents from Web pages, are commonly used to obtain page contents that are critical to many Web mining applications. However, traditional density-based approaches cannot effectively manage pages that contain short contents and long noises. To overcome this problem, in this paper, we propose a content extraction approach for obtaining content from news pages that combines a segmentation-like approach and a density-based approach. A tool called BlockExtractor was developed based on this approach. BlockExtractor identifies contents in three steps. First, it looks for all Block-Level Elements (BLE) & Inline Elements (IE) blocks, which are designed to roughly segment pages into blocks. Second, it computes the densities of each BLE&IE block and its element to eliminate noises. Third, it removes all redundant BLEIE blocks that have emerged in other pages from the same site. Compared with three other density-based approaches, our approach shows significant advantages in both precision and recall.

源语言英语
文章编号6216755
页(从-至)256-264
页数9
期刊Tsinghua Science and Technology
17
3
DOI
出版状态已出版 - 2012

指纹

探究 'Combining a segmentation-like approach and a density-based approach in content extraction' 的科研主题。它们共同构成独一无二的指纹。

引用此