Computer Science Faculty Publications

A Pseudo Nearest-Neighbor Approach for Missing Data Recovery on Gaussian Random Data Sets

Document Type

Article

Publication Date

11-2002

Publication Title

Pattern Recognition Letters

Volume

Issue

First Page

1613

Last Page

1622

Abstract

Missing data handling is an important preparation step for most data discrimination or mining tasks. Inappropriate treatment of missing data may cause large errors or false results. In this paper, we study the effect of a missing data recovery method, namely the pseudo- nearest neighbor substitution approach, on Gaussian distributed data sets that represent typical cases in data discrimination and data mining applications. The error rate of the proposed recovery method is evaluated by comparing the clustering results of the recovered data sets to the clustering results obtained on the originally complete data sets. The results are also compared with that obtained by applying two other missing data handling methods, the constant default value substitution and the missing data ignorance (non-substitution) methods. The experiment results provided a valuable insight to the improvement of the accuracy for data discrimination and knowledge discovery on large data sets containing missing values.

Comments

The final published version of this article can be found here: http://www.sciencedirect.com/science/article/pii/S0167865502001253.

Recommended Citation

Huang, Xiaolu and Zhu, Qiuming, "A Pseudo Nearest-Neighbor Approach for Missing Data Recovery on Gaussian Random Data Sets" (2002). Computer Science Faculty Publications. 23.
https://digitalcommons.unomaha.edu/compscifacpub/23

Download

Find in your library

Included in

Computer Sciences Commons

COinS

DigitalCommons@UNO

Computer Science Faculty Publications

A Pseudo Nearest-Neighbor Approach for Missing Data Recovery on Gaussian Random Data Sets

Document Type

Publication Date

Publication Title

Volume

Issue

First Page

Last Page

Abstract

Comments

Recommended Citation

Included in

Search

Links

Browse

Author Corner

Links

DigitalCommons@UNO

Computer Science Faculty Publications

A Pseudo Nearest-Neighbor Approach for Missing Data Recovery on Gaussian Random Data Sets

Authors

Document Type

Publication Date

Publication Title

Volume

Issue

First Page

Last Page

Abstract

Comments

Recommended Citation

Included in

Share

Search

Links

Browse

Author Corner

Links