Abstract
Cantonese is an important Chinese dialect spoken in some regions of Southern China. Local online users often represent their opinions and experiences with written Cantonese on the web. With two supervised machine learning approaches, this paper conducts a series of experiments to explore appropriate methods for automatic sentiment classification in the very noisy domain of online Cantonese-written reviews. Findings indicate that the support vector machine classifier based on a Mandarin Chinese word segmentation tool performs surprisingly well. The accuracy, precision and recall respectively for positive and negative reviews all reach above 85% when the training corpus contains 5,000 or more reviews.
Original language | English |
---|---|
Pages (from-to) | 382-397 |
Number of pages | 16 |
Journal | International Journal of Web Engineering and Technology |
Volume | 5 |
Issue number | 4 |
DOIs | |
Publication status | Published - 1 Mar 2009 |
Keywords
- Cantonese
- Online reviews
- Sentiment classification
- Text mining
ASJC Scopus subject areas
- Information Systems
- Hardware and Architecture
- Computer Networks and Communications