VLDB 2026 Research / reviewers in the wild / expert
In Ho Cho
dblp:32/10244
· DBLP profile ↗
3ranked-venue papers
0as first author
2since 2021 · last 2023
0000-0002-2265-9602ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Parallel and multicore computing · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Data integration and cleaning · 77% Data mining · 23% | |
| Artificial intelligence
1 paper |
Deep learning architectures and training · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing
parallel algorithms |
1.2 | 2 | 2023 | Ultra Data-Oriented Parallel Fractional Hot-Deck Imputation With Efficient Linearized Variance Estimation · IEEE Trans. Knowl. Data Eng. 2023 Parallel Fractional Hot-Deck Imputation and Variance Estimation for Big Incomplete Data Curing · IEEE Trans. Knowl. Data Eng. 2022 |
Data integration and cleaning › missing data
missing value imputation |
0.4 | 1 | 2020 | Impacts of Fractional Hot-Deck Imputation on Learning and Prediction of Engineering Data · IEEE Trans. Knowl. Data Eng. 2020 |
Computational science and engineering
statistical computing |
0.2 | 1 | 2022 | Parallel Fractional Hot-Deck Imputation and Variance Estimation for Big Incomplete Data Curing · IEEE Trans. Knowl. Data Eng. 2022 |
Data mining
predictive modeling |
0.1 | 1 | 2020 | Impacts of Fractional Hot-Deck Imputation on Learning and Prediction of Engineering Data · IEEE Trans. Knowl. Data Eng. 2020 |
Methods — techniques the papers use, named apart from their topics
fractional hot-deck imputation · 2.5linearization · 1.3jackknife · 1.3jackknife variance estimation · 1.1support vector machine · 0.4generalized additive model · 0.4extremely randomized trees · 0.4artificial neural network · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Ultra Data-Oriented Parallel Fractional Hot-Deck Imputation With Efficient Linearized Variance EstimationabstractParallel fractional hot-deck imputation (P-FHDI (Yang et al. 2020)) is a general-purpose, assumption-free tool for handling item nonresponse in big incomplete data by combining the theory of FHDI and parallel computing. FHDI cures multivariate missing data by filling each missing unit with multiple observed values (thus, hot-deck) without resorting to distributional assumptions. P-FHDI can tackle big incomplete data with millions of instances (big-$n$) or 10,000 variables (big-$p$). However, handling ultra incomplete data (i.e., concurrently big-$n$and big-$p$) with tremendous instances and high dimensionality has posed problems to P-FHDI due to excessive memory requirement and execution time. To tackle the aforementioned challenges, we propose the ultra data-oriented P-FHDI (named UP-FHDI) capable of curing ultra incomplete data. In addition to the parallel Jackknife method, this paper enables a computationally efficient ultra data-oriented variance estimation using parallel linearization techniques. Results confirm that UP-FHDI can tackle an ultra dataset with one million instances and 10,000 variables. This paper illustrates the special parallel algorithms of UP-FHDI and confirms its positive impact on the subsequent deep learning performance. Yicheng Yang, Yonghyun Kwon, Jae Kwang Kim, In Ho Cho |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2022 | Parallel Fractional Hot-Deck Imputation and Variance Estimation for Big Incomplete Data CuringabstractThe fractional hot-deck imputation (FHDI) is a general-purpose, assumption-free imputation method for handling multivariate missing data by filling each missing item with multiple observed values without resorting to artificially created values. The corresponding R package FHDI J. Im, I. Cho, and J. K. Kim, “An R package for fractional hot deck imputation,”R J., vol. 10, no. 1, pp. 140–154, 2018 holds generality and efficiency, but it is not adequate for tackling big incomplete data due to the requirement of excessive memory and long running time. As a first step to tackle big incomplete data by leveraging the FHDI, we developed a new version of a parallel fractional hot-deck imputation (named as P-FHDI) program suitable for curing large incomplete datasets. Results show a favorable speedup when the P-FHDI is applied to big datasets with up to millions of instances or 10,000 of variables. This paper explains the detailed parallel algorithms of the P-FHDI for large instances (big-$n$) or high-dimensionality (big-$p$) datasets and confirms the favorable scalability. The proposed program inherits all the advantages of the serial FHDI and enables a parallel variance estimation, which will benefit a broad audience in science and engineering. Yicheng Yang, Jae Kwang Kim, In Ho Cho |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2020 | Impacts of Fractional Hot-Deck Imputation on Learning and Prediction of Engineering DataabstractIn broad engineering fields, missing data is a common issue which often causes undesired bias and sparseness impeding rigorous data analyses. To tackle this problem, many imputation theories have been proposed and widely used. However, prior methods often require distributional assumptions and prior knowledge regarding data which may cause some difficulty for engineering research. Essentially, the fractional hot-deck imputation (FHDI) is an assumption-free imputation method, holding broad applicability in the engineering domains. FHDIs internal parameters and impact on statistical and machine learning methods, however, have been rarely understood. Thus, this study investigates the behavior and impacts of FHDI on prediction methods including generalized additive model, support vector machine, extremely randomized trees, and artificial neural network, for which four practical datasets (appliance energy, air quality, phenotypes, and weather) are used. Results show that FHDI performs better for improving the prediction accuracy compared to a simple naive method which cures missing data using the mean value of attributes, and FHDI has an asymptotically positive effect on prediction accuracy with decreasing response rates. Regarding an optimal setting, 30 to 35 is recommended for the FHDIs internal categorization number while 5 is recommended for the FHDI donors, which is aligned with Rubins recommendation. Ikkyun Song, Yicheng Yang, Jongho Im, Halil Ceylan, In Ho Cho |
IEEE Trans. Knowl. Data Eng. | 6 |