Skip to main navigation Skip to search Skip to main content

Chinese text classification without word segmentation

  • Yun Xu*
  • , Xiao Zhong Fan
  • , Feng Zhang
  • *Corresponding author for this work
  • Beijing Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Proposes an approach for Chinese language text classification without word segmentation based on n-gram language modeling. Unlike the case of traditional text classification models, the approach based on character level n-gram modeling avoids word segmentation and explicit feature selection procedures that tends to lose significant amount of useful information. It greatly reduces the problem of sparsity of data, because the size of the vocabulary made up of characters is smaller than that formed from words. Systematic study of key factors in language modeling and their influence on classification shows that the estimated index based on experiments on Chinese TREC attained 86.8%.

Original languageEnglish
Pages (from-to)778-781
Number of pages4
JournalBeijing Ligong Daxue Xuebao/Transaction of Beijing Institute of Technology
Volume25
Issue number9
Publication statusPublished - Sept 2005

Keywords

  • n-gram model
  • Text classification
  • Word segmentation

Fingerprint

Dive into the research topics of 'Chinese text classification without word segmentation'. Together they form a unique fingerprint.

Cite this