Similarity-Based Synthetic Document Representations for Meta-Feature Generation in Text Classification

Sérgio Canuto,Thiago Salles,Thierson Couto Rosa,Marcos André Gonçalves

Similarity-Based Synthetic Document Representations for Meta-Feature Generation in Text Classification

2019

Sérgio Canuto
Thiago Salles
Thierson Couto Rosa
Marcos André Gonçalves

We propose new solutions that enhance and extend the already very successful application of meta-features to text classification. Our newly proposed meta-features are capable of: (1) improving the correlation of small pieces of evidence shared by neighbors with labeled categories by means of synthetic document representations and (local and global) hyperplane distances; and (2) estimating the level of error introduced by these newly proposed and the existing meta-features in the literature, specially for hard-to-classify regions of the feature space. Our experiments with large and representative number of datasets show that our new solutions produce the best results in all tested scenarios, achieving gains of up to 12% over the strongest meta-feature proposal of the literature.

Keywords:

Information retrieval
Computer science
Correlation
Feature vector
feature generation
Hyperplane
Data mining

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

Citations