Skip to main navigation Skip to search Skip to main content

Extracting Chinese multi-word terms from small corpus

  • Nanjing University of Science and Technology
  • Chinese Academy of Sciences
  • Nanjing University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

In this paper, we present an automatic terminology extraction approach for Chinese multi-word terms. In this term extraction system, besides five linguistic rides acqidred from an available term list by some machine learning methods, two statistical strategies are involved: a termhood measure based on the term distribution variation, and a unithood measure adopting the left and right entropy method to estimate the collocation variation degree. The candidates are ranked according to the values of the former. The latter is used to filter the preposition phrases and some verb-object phrases that rarely appear as terms. By validating on a small scale corpus in the computer domain, the precision reaches 91.5% of the top 2000 outputs.

Original languageEnglish
Title of host publicationProceedings of 2008 3rd International Conference on Intelligent System and Knowledge Engineering, ISKE 2008
Pages813-818
Number of pages6
DOIs
Publication statusPublished - 2008
Externally publishedYes
EventProceedings of 2008 3rd International Conference on Intelligent System and Knowledge Engineering, ISKE 2008 - Xiamen, China
Duration: 17 Nov 200819 Nov 2008

Publication series

NameProceedings of 2008 3rd International Conference on Intelligent System and Knowledge Engineering, ISKE 2008

Conference

ConferenceProceedings of 2008 3rd International Conference on Intelligent System and Knowledge Engineering, ISKE 2008
Country/TerritoryChina
CityXiamen
Period17/11/0819/11/08

Fingerprint

Dive into the research topics of 'Extracting Chinese multi-word terms from small corpus'. Together they form a unique fingerprint.

Cite this