Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

In Ho Cho

dblp:32/10244 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
2since 2021 · last 2023
0000-0002-2265-9602ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 100%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 77% Data mining · 23%
Artificial intelligence
1 paper
Deep learning architectures and training · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing
parallel algorithms
1.222023
Ultra Data-Oriented Parallel Fractional Hot-Deck Imputation With Efficient Linearized Variance Estimation · IEEE Trans. Knowl. Data Eng. 2023
Parallel Fractional Hot-Deck Imputation and Variance Estimation for Big Incomplete Data Curing · IEEE Trans. Knowl. Data Eng. 2022
Data integration and cleaning › missing data
missing value imputation
0.412020
Impacts of Fractional Hot-Deck Imputation on Learning and Prediction of Engineering Data · IEEE Trans. Knowl. Data Eng. 2020
Computational science and engineering
statistical computing
0.212022
Parallel Fractional Hot-Deck Imputation and Variance Estimation for Big Incomplete Data Curing · IEEE Trans. Knowl. Data Eng. 2022
Data mining
predictive modeling
0.112020
Impacts of Fractional Hot-Deck Imputation on Learning and Prediction of Engineering Data · IEEE Trans. Knowl. Data Eng. 2020

Methods — techniques the papers use, named apart from their topics

fractional hot-deck imputation · 2.5linearization · 1.3jackknife · 1.3jackknife variance estimation · 1.1support vector machine · 0.4generalized additive model · 0.4extremely randomized trees · 0.4artificial neural network · 0.4
YearPublicationVenuePosition
2023 Ultra Data-Oriented Parallel Fractional Hot-Deck Imputation With Efficient Linearized Variance Estimation
abstract
Parallel fractional hot-deck imputation (P-FHDI (Yang et al. 2020)) is a general-purpose, assumption-free tool for handling item nonresponse in big incomplete data by combining the theory of FHDI and parallel computing. FHDI cures multivariate missing data by filling each missing unit with multiple observed values (thus, hot-deck) without resorting to distributional assumptions. P-FHDI can tackle big incomplete data with millions of instances (big-$n$) or 10,000 variables (big-$p$). However, handling ultra incomplete data (i.e., concurrently big-$n$and big-$p$) with tremendous instances and high dimensionality has posed problems to P-FHDI due to excessive memory requirement and execution time. To tackle the aforementioned challenges, we propose the ultra data-oriented P-FHDI (named UP-FHDI) capable of curing ultra incomplete data. In addition to the parallel Jackknife method, this paper enables a computationally efficient ultra data-oriented variance estimation using parallel linearization techniques. Results confirm that UP-FHDI can tackle an ultra dataset with one million instances and 10,000 variables. This paper illustrates the special parallel algorithms of UP-FHDI and confirms its positive impact on the subsequent deep learning performance.
Yicheng Yang, Yonghyun Kwon, Jae Kwang Kim, In Ho Cho
IEEE Trans. Knowl. Data Eng.4
2022 Parallel Fractional Hot-Deck Imputation and Variance Estimation for Big Incomplete Data Curing
abstract
The fractional hot-deck imputation (FHDI) is a general-purpose, assumption-free imputation method for handling multivariate missing data by filling each missing item with multiple observed values without resorting to artificially created values. The corresponding R package FHDI J. Im, I. Cho, and J. K. Kim, “An R package for fractional hot deck imputation,”R J., vol. 10, no. 1, pp. 140–154, 2018 holds generality and efficiency, but it is not adequate for tackling big incomplete data due to the requirement of excessive memory and long running time. As a first step to tackle big incomplete data by leveraging the FHDI, we developed a new version of a parallel fractional hot-deck imputation (named as P-FHDI) program suitable for curing large incomplete datasets. Results show a favorable speedup when the P-FHDI is applied to big datasets with up to millions of instances or 10,000 of variables. This paper explains the detailed parallel algorithms of the P-FHDI for large instances (big-$n$) or high-dimensionality (big-$p$) datasets and confirms the favorable scalability. The proposed program inherits all the advantages of the serial FHDI and enables a parallel variance estimation, which will benefit a broad audience in science and engineering.
Yicheng Yang, Jae Kwang Kim, In Ho Cho
IEEE Trans. Knowl. Data Eng.3
2020 Impacts of Fractional Hot-Deck Imputation on Learning and Prediction of Engineering Data
abstract
In broad engineering fields, missing data is a common issue which often causes undesired bias and sparseness impeding rigorous data analyses. To tackle this problem, many imputation theories have been proposed and widely used. However, prior methods often require distributional assumptions and prior knowledge regarding data which may cause some difficulty for engineering research. Essentially, the fractional hot-deck imputation (FHDI) is an assumption-free imputation method, holding broad applicability in the engineering domains. FHDIs internal parameters and impact on statistical and machine learning methods, however, have been rarely understood. Thus, this study investigates the behavior and impacts of FHDI on prediction methods including generalized additive model, support vector machine, extremely randomized trees, and artificial neural network, for which four practical datasets (appliance energy, air quality, phenotypes, and weather) are used. Results show that FHDI performs better for improving the prediction accuracy compared to a simple naive method which cures missing data using the mean value of attributes, and FHDI has an asymptotically positive effect on prediction accuracy with decreasing response rates. Regarding an optimal setting, 30 to 35 is recommended for the FHDIs internal categorization number while 5 is recommended for the FHDI donors, which is aligned with Rubins recommendation.
Ikkyun Song, Yicheng Yang, Jongho Im, Halil Ceylan, In Ho Cho
IEEE Trans. Knowl. Data Eng.6