Learning individual models for imputation

Aoqian Zhang; Shaoxu Song; Yu Sun; Jianmin Wang

doi:10.1109/ICDE.2019.00023

Learning individual models for imputation

Aoqian Zhang, Shaoxu Song, Yu Sun, Jianmin Wang

Tsinghua University

科研成果: 书/报告/会议事项章节 › 会议稿件 › 同行评审

30 引用（Scopus）

摘要

Missing numerical values are prevalent, e.g., owing to unreliable sensor reading, collection and transmission among heterogeneous sources. Unlike categorized data imputation over a limited domain, the numerical values suffer from two issues: (1) sparsity problem, the incomplete tuple may not have sufficient complete neighbors sharing the same/similar values for imputation, owing to the (almost) infinite domain; (2) heterogeneity problem, different tuples may not fit the same (regression) model. In this study, enlightened by the conditional dependencies that hold conditionally over certain tuples rather than the whole relation, we propose to learn a regression model individually for each complete tuple together with its neighbors. Our IIM, Imputation via Individual Models, thus no longer relies on sharing similar values among the k complete neighbors for imputation, but utilizes their regression results by the aforesaid learned individual (not necessary the same) models. Remarkably, we show that some existing methods are indeed special cases of our IIM, under the extreme settings of the number ℓ of learning neighbors considered in individual learning. In this sense, a proper number ℓ of neighbors is essential to learn the individual models (avoid over-fitting or under-fitting). We propose to adaptively learn individual models over various number ℓ of neighbors for different complete tuples. By devising efficient incremental computation, the time complexity of learning a model reduces from linear to constant. Experiments on real data demonstrate that our IIM with adaptive learning achieves higher imputation accuracy than the existing approaches.

源语言	英语
主期刊名	Proceedings - 2019 IEEE 35th International Conference on Data Engineering, ICDE 2019
出版商	IEEE Computer Society
页	160-171
页数	12
ISBN（电子版）	9781538674741
DOI	https://doi.org/10.1109/ICDE.2019.00023
出版状态	已出版 - 4月 2019
已对外发布	是
活动	35th IEEE International Conference on Data Engineering, ICDE 2019 - Macau, 中国期限: 8 4月 2019 → 11 4月 2019

出版系列

姓名	Proceedings - International Conference on Data Engineering
卷	2019-April
ISSN（印刷版）	1084-4627

会议

会议	35th IEEE International Conference on Data Engineering, ICDE 2019
国家/地区	中国
市	Macau
时期	8/04/19 → 11/04/19

访问文件

10.1109/ICDE.2019.00023

其它文件与链接

链接到 Scopus 的出版物

引用此

Zhang, A., Song, S., Sun, Y., & Wang, J. (2019). Learning individual models for imputation. 在 Proceedings - 2019 IEEE 35th International Conference on Data Engineering, ICDE 2019 (页码 160-171). 文章 8731351 (Proceedings - International Conference on Data Engineering; 卷 2019-April). IEEE Computer Society. https://doi.org/10.1109/ICDE.2019.00023

@inproceedings{3e5c2a5f38844acc849d932c066f8ac8,

title = "Learning individual models for imputation",

abstract = "Missing numerical values are prevalent, e.g., owing to unreliable sensor reading, collection and transmission among heterogeneous sources. Unlike categorized data imputation over a limited domain, the numerical values suffer from two issues: (1) sparsity problem, the incomplete tuple may not have sufficient complete neighbors sharing the same/similar values for imputation, owing to the (almost) infinite domain; (2) heterogeneity problem, different tuples may not fit the same (regression) model. In this study, enlightened by the conditional dependencies that hold conditionally over certain tuples rather than the whole relation, we propose to learn a regression model individually for each complete tuple together with its neighbors. Our IIM, Imputation via Individual Models, thus no longer relies on sharing similar values among the k complete neighbors for imputation, but utilizes their regression results by the aforesaid learned individual (not necessary the same) models. Remarkably, we show that some existing methods are indeed special cases of our IIM, under the extreme settings of the number ℓ of learning neighbors considered in individual learning. In this sense, a proper number ℓ of neighbors is essential to learn the individual models (avoid over-fitting or under-fitting). We propose to adaptively learn individual models over various number ℓ of neighbors for different complete tuples. By devising efficient incremental computation, the time complexity of learning a model reduces from linear to constant. Experiments on real data demonstrate that our IIM with adaptive learning achieves higher imputation accuracy than the existing approaches.",

keywords = "Data imputation, Missing values",

author = "Aoqian Zhang and Shaoxu Song and Yu Sun and Jianmin Wang",

note = "Publisher Copyright: {\textcopyright} 2019 IEEE.; 35th IEEE International Conference on Data Engineering, ICDE 2019 ; Conference date: 08-04-2019 Through 11-04-2019",

year = "2019",

month = apr,

doi = "10.1109/ICDE.2019.00023",

language = "English",

series = "Proceedings - International Conference on Data Engineering",

publisher = "IEEE Computer Society",

pages = "160--171",

booktitle = "Proceedings - 2019 IEEE 35th International Conference on Data Engineering, ICDE 2019",

address = "United States",

}

Zhang, A, Song, S, Sun, Y & Wang, J 2019, Learning individual models for imputation. 在 Proceedings - 2019 IEEE 35th International Conference on Data Engineering, ICDE 2019., 8731351, Proceedings - International Conference on Data Engineering, 卷 2019-April, IEEE Computer Society, 页码 160-171, 35th IEEE International Conference on Data Engineering, ICDE 2019, Macau, 中国, 8/04/19. https://doi.org/10.1109/ICDE.2019.00023

Learning individual models for imputation. / Zhang, Aoqian; Song, Shaoxu; Sun, Yu 等.
Proceedings - 2019 IEEE 35th International Conference on Data Engineering, ICDE 2019. IEEE Computer Society, 2019. 页码 160-171 8731351 (Proceedings - International Conference on Data Engineering; 卷 2019-April).

科研成果: 书/报告/会议事项章节 › 会议稿件 › 同行评审

TY - GEN

T1 - Learning individual models for imputation

AU - Zhang, Aoqian

AU - Song, Shaoxu

AU - Sun, Yu

AU - Wang, Jianmin

PY - 2019/4

Y1 - 2019/4

N2 - Missing numerical values are prevalent, e.g., owing to unreliable sensor reading, collection and transmission among heterogeneous sources. Unlike categorized data imputation over a limited domain, the numerical values suffer from two issues: (1) sparsity problem, the incomplete tuple may not have sufficient complete neighbors sharing the same/similar values for imputation, owing to the (almost) infinite domain; (2) heterogeneity problem, different tuples may not fit the same (regression) model. In this study, enlightened by the conditional dependencies that hold conditionally over certain tuples rather than the whole relation, we propose to learn a regression model individually for each complete tuple together with its neighbors. Our IIM, Imputation via Individual Models, thus no longer relies on sharing similar values among the k complete neighbors for imputation, but utilizes their regression results by the aforesaid learned individual (not necessary the same) models. Remarkably, we show that some existing methods are indeed special cases of our IIM, under the extreme settings of the number ℓ of learning neighbors considered in individual learning. In this sense, a proper number ℓ of neighbors is essential to learn the individual models (avoid over-fitting or under-fitting). We propose to adaptively learn individual models over various number ℓ of neighbors for different complete tuples. By devising efficient incremental computation, the time complexity of learning a model reduces from linear to constant. Experiments on real data demonstrate that our IIM with adaptive learning achieves higher imputation accuracy than the existing approaches.

AB - Missing numerical values are prevalent, e.g., owing to unreliable sensor reading, collection and transmission among heterogeneous sources. Unlike categorized data imputation over a limited domain, the numerical values suffer from two issues: (1) sparsity problem, the incomplete tuple may not have sufficient complete neighbors sharing the same/similar values for imputation, owing to the (almost) infinite domain; (2) heterogeneity problem, different tuples may not fit the same (regression) model. In this study, enlightened by the conditional dependencies that hold conditionally over certain tuples rather than the whole relation, we propose to learn a regression model individually for each complete tuple together with its neighbors. Our IIM, Imputation via Individual Models, thus no longer relies on sharing similar values among the k complete neighbors for imputation, but utilizes their regression results by the aforesaid learned individual (not necessary the same) models. Remarkably, we show that some existing methods are indeed special cases of our IIM, under the extreme settings of the number ℓ of learning neighbors considered in individual learning. In this sense, a proper number ℓ of neighbors is essential to learn the individual models (avoid over-fitting or under-fitting). We propose to adaptively learn individual models over various number ℓ of neighbors for different complete tuples. By devising efficient incremental computation, the time complexity of learning a model reduces from linear to constant. Experiments on real data demonstrate that our IIM with adaptive learning achieves higher imputation accuracy than the existing approaches.

KW - Data imputation

KW - Missing values

UR - http://www.scopus.com/inward/record.url?scp=85067972874&partnerID=8YFLogxK

U2 - 10.1109/ICDE.2019.00023

DO - 10.1109/ICDE.2019.00023

M3 - Conference contribution

AN - SCOPUS:85067972874

T3 - Proceedings - International Conference on Data Engineering

SP - 160

EP - 171

BT - Proceedings - 2019 IEEE 35th International Conference on Data Engineering, ICDE 2019

PB - IEEE Computer Society

T2 - 35th IEEE International Conference on Data Engineering, ICDE 2019

Y2 - 8 April 2019 through 11 April 2019

ER -

Learning individual models for imputation

摘要

出版系列

会议

访问文件

其它文件与链接

指纹

引用此