Abstract
Proposes an approach for Chinese language text classification without word segmentation based on n-gram language modeling. Unlike the case of traditional text classification models, the approach based on character level n-gram modeling avoids word segmentation and explicit feature selection procedures that tends to lose significant amount of useful information. It greatly reduces the problem of sparsity of data, because the size of the vocabulary made up of characters is smaller than that formed from words. Systematic study of key factors in language modeling and their influence on classification shows that the estimated index based on experiments on Chinese TREC attained 86.8%.
| Original language | English |
|---|---|
| Pages (from-to) | 778-781 |
| Number of pages | 4 |
| Journal | Beijing Ligong Daxue Xuebao/Transaction of Beijing Institute of Technology |
| Volume | 25 |
| Issue number | 9 |
| Publication status | Published - Sept 2005 |
Keywords
- n-gram model
- Text classification
- Word segmentation
Fingerprint
Dive into the research topics of 'Chinese text classification without word segmentation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver