Published 2005 | Version public
Book Section - Chapter

Improving Generalization by Data Categorization

Abstract

In most of the learning algorithms, examples in the training set are treated equally. Some examples, however, carry more reliable or critical information about the target than the others, and some may carry wrong information. According to their intrinsic margin, examples can be grouped into three categories: typical, critical, and noisy. We propose three methods, namely the selection cost, SVM confidence margin, and AdaBoost data weight, to automatically group training examples into these three categories. Experimental results on artificial datasets show that, although the three methods have quite different nature, they give similar and reasonable categorization. Results with real-world datasets further demonstrate that treating the three data categories differently in learning can improve generalization.

Additional Information

© 2005 Springer-Verlag Berlin Heidelberg. We thank Anelia Angelova, Marcelo Medeiros, Carlos Pedreira, David Soloveichik and the anonymous reviewers for helpful discussions. This work was mainly done in 2003 and was supported by the Caltech Center for Neuromorphic Systems Engineering under the US NSF Cooperative Agreement EEC-9402726.

Additional details

Identifiers

Eprint ID
96892
DOI
10.1007/11564126_19
Resolver ID
CaltechAUTHORS:20190702-142717858

Related works

Describes
10.1007/11564126_19 (DOI)

Funding

NSF
EEC-9402726
Center for Neuromorphic Systems Engineering, Caltech

Dates

Created
2019-07-08
Created from EPrint's datestamp field
Updated
2021-11-16
Created from EPrint's last_modified field

Caltech Custom Metadata

Series Name
Lecture Notes in Computer Science
Series Volume or Issue Number
3721