Understand Short Texts by Harvesting and Analyzing Semantic Knowledge

Wen Hua, Zhongyuan Wang, Haixun Wang, Kai Zheng, Xiaofang Zhou

Research output: Journal article publicationJournal articleAcademic researchpeer-review

67 Citations (Scopus)

Abstract

Understanding short texts is crucial to many applications, but challenges abound. First, short texts do not always observe the syntax of a written language. As a result, traditional natural language processing tools, ranging from part-of-speech tagging to dependency parsing, cannot be easily applied. Second, short texts usually do not contain sufficient statistical signals to support many state-of-the-art approaches for text mining such as topic modeling. Third, short texts are more ambiguous and noisy, and are generated in an enormous volume, which further increases the difficulty to handle them. We argue that semantic knowledge is required in order to better understand short texts. In this work, we build a prototype system for short text understanding which exploits semantic knowledge provided by a well-known knowledgebase and automatically harvested from a web corpus. Our knowledge-intensive approaches disrupt traditional methods for tasks such as text segmentation, part-of-speech tagging, and concept labeling, in the sense that we focus on semantics in all these tasks. We conduct a comprehensive performance evaluation on real-life data. The results show that semantic knowledge is indispensable for short text understanding, and our knowledge-intensive approaches are both effective and efficient in discovering semantics of short texts.

Original languageEnglish
Article number7476863
Pages (from-to)499-512
Number of pages14
JournalIEEE Transactions on Knowledge and Data Engineering
Volume29
Issue number3
DOIs
Publication statusPublished - 1 Mar 2017
Externally publishedYes

Keywords

  • concept labeling
  • semantic knowledge
  • Short text understanding
  • Text segmentation
  • type detection

ASJC Scopus subject areas

  • Information Systems
  • Computer Science Applications
  • Computational Theory and Mathematics

Fingerprint

Dive into the research topics of 'Understand Short Texts by Harvesting and Analyzing Semantic Knowledge'. Together they form a unique fingerprint.

Cite this