跳到主要导航 跳到搜索 跳到主要内容

Computation on sentence semantic distance for novelty detection

  • Hua Ping Zhang*
  • , Jian Sun
  • , Bing Wang
  • , Shuo Bai
  • *此作品的通讯作者
  • CAS - Institute of Computing Technology
  • University of Chinese Academy of Sciences

科研成果: 期刊稿件文章同行评审

摘要

Novelty detection is to retrieve new information and filter redundancy from given sentences that are relevant to a specific topic. In TREC2003, the authors tried an approach to novelty detection with semantic distance computation. The motivation is to expand a sentence by introducing semantic information. Computation on semantic distance between sentences incorporates WordNet with statistical information. The novelty detection is treated as a binary classification problem: new sentence or not. The feature vector, used in the vector space model for classification, consists of various factors, including the semantic distance from the sentence to the topic and the distance from the sentence to the previous relevant context occurring before it. New sentences are then detected with Winnow and support vector machine classifiers, respectively. Several experiments are conducted to survey the relationship between different factors and performance. It is proved that semantic computation is promising in novelty detection. The ratio of new sentence size to relevant size is further studied given different relevant document sizes. It is found that the ratio reduced with a certain speed (about 0.86). Then another group of experiments is performed supervised with the ratio. It is demonstrated that the ratio is helpful to improve the novelty detection performance.

源语言英语
页(从-至)331-337
页数7
期刊Journal of Computer Science and Technology
20
3
DOI
出版状态已出版 - 5月 2005
已对外发布

学术指纹

探究 'Computation on sentence semantic distance for novelty detection' 的科研主题。它们共同构成独一无二的学术指纹。

引用此