A heuristic approach to determine an appropriate number of topics in topic modeling.

Weizhong Zhao,James J. Chen,Roger Perkins,Zhichao Liu,Weigong Ge,Yijun Ding,Wen Zou

A heuristic approach to determine an appropriate number of topics in topic modeling.

2015

Weizhong Zhao
James J. Chen
Roger Perkins
Zhichao Liu
Weigong Ge
Yijun Ding
Wen Zou

Background Topic modelling is an active research field in machine learning. While mainly used to build models from unstructured textual data, it offers an effective means of data mining where samples represent documents, and different biological endpoints or omics data represent words. Latent Dirichlet Allocation (LDA) is the most commonly used topic modelling method across a wide number of technical fields. However, model development can be arduous and tedious, and requires burdensome and systematic sensitivity studies in order to find the best set of model parameters. Often, time-consuming subjective evaluations are needed to compare models. Currently, research has yielded no easy way to choose the proper number of topics in a model beyond a major iterative approach.

Keywords:

Heuristics
Latent Dirichlet allocation
Perplexity
Bioinformatics
Data mining
Topic model
Heuristic
Computer science
Omics
omics data
model parameters
model development

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

125

Citations