Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jae Kwang Kim

dblp:238/1626 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
3since 2021 · last 2023
0000-0002-0246-6029ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
1 paper
Mathematical optimization · 67% Algorithms and data structures · 33%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 100%
Artificial intelligence
1 paper
Deep learning architectures and training · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing
parallel algorithms
1.222023
Ultra Data-Oriented Parallel Fractional Hot-Deck Imputation With Efficient Linearized Variance Estimation · IEEE Trans. Knowl. Data Eng. 2023
Parallel Fractional Hot-Deck Imputation and Variance Estimation for Big Incomplete Data Curing · IEEE Trans. Knowl. Data Eng. 2022
Mathematical optimization › statistical estimation
maximum likelihood estimation
0.612022
Maximum sampled conditional likelihood for informative subsampling · J. Mach. Learn. Res. 2022
Mathematical optimization
statistical estimation
0.612022
Maximum sampled conditional likelihood for informative subsampling · J. Mach. Learn. Res. 2022
Algorithms and data structures › randomized algorithms › sampling
subsampling
0.612022
Maximum sampled conditional likelihood for informative subsampling · J. Mach. Learn. Res. 2022
Computational science and engineering
statistical computing
0.212022
Parallel Fractional Hot-Deck Imputation and Variance Estimation for Big Incomplete Data Curing · IEEE Trans. Knowl. Data Eng. 2022

Methods — techniques the papers use, named apart from their topics

fractional hot-deck imputation · 2.5linearization · 1.3jackknife · 1.3jackknife variance estimation · 1.1inverse probability weighting · 0.6generalized linear model · 0.6asymptotic analysis · 0.6
YearPublicationVenuePosition
2023 Ultra Data-Oriented Parallel Fractional Hot-Deck Imputation With Efficient Linearized Variance Estimation
abstract
Parallel fractional hot-deck imputation (P-FHDI (Yang et al. 2020)) is a general-purpose, assumption-free tool for handling item nonresponse in big incomplete data by combining the theory of FHDI and parallel computing. FHDI cures multivariate missing data by filling each missing unit with multiple observed values (thus, hot-deck) without resorting to distributional assumptions. P-FHDI can tackle big incomplete data with millions of instances (big-$n$) or 10,000 variables (big-$p$). However, handling ultra incomplete data (i.e., concurrently big-$n$and big-$p$) with tremendous instances and high dimensionality has posed problems to P-FHDI due to excessive memory requirement and execution time. To tackle the aforementioned challenges, we propose the ultra data-oriented P-FHDI (named UP-FHDI) capable of curing ultra incomplete data. In addition to the parallel Jackknife method, this paper enables a computationally efficient ultra data-oriented variance estimation using parallel linearization techniques. Results confirm that UP-FHDI can tackle an ultra dataset with one million instances and 10,000 variables. This paper illustrates the special parallel algorithms of UP-FHDI and confirms its positive impact on the subsequent deep learning performance.
Yicheng Yang, Yonghyun Kwon, Jae Kwang Kim, In Ho Cho
IEEE Trans. Knowl. Data Eng.3
2022 Maximum sampled conditional likelihood for informative subsampling
abstract
Subsampling is a computationally effective approach to extract information from massive data sets when computing resources are limited. After a subsample is taken from the full data, most available methods use an inverse probability weighted (IPW) objective function to estimate the model parameters. The IPW estimator does not fully utilize the information in the selected subsample. In this paper, we propose to use the maximum sampled conditional likelihood estimator (MSCLE) based on the sampled data. We established the asymptotic normality of the MSCLE and prove that its asymptotic variance covariance matrix is the smallest among a class of asymptotically unbiased estimators, including the IPW estimator. We further discuss the asymptotic results with the L-optimal subsampling probabilities and illustrate the estimation procedure with generalized linear models. Numerical experiments are provided to evaluate the practical performance of the proposed method.
HaiYing Wang 0004, Jae Kwang Kim
J. Mach. Learn. Res.2
2022 Parallel Fractional Hot-Deck Imputation and Variance Estimation for Big Incomplete Data Curing
abstract
The fractional hot-deck imputation (FHDI) is a general-purpose, assumption-free imputation method for handling multivariate missing data by filling each missing item with multiple observed values without resorting to artificially created values. The corresponding R package FHDI J. Im, I. Cho, and J. K. Kim, “An R package for fractional hot deck imputation,”R J., vol. 10, no. 1, pp. 140–154, 2018 holds generality and efficiency, but it is not adequate for tackling big incomplete data due to the requirement of excessive memory and long running time. As a first step to tackle big incomplete data by leveraging the FHDI, we developed a new version of a parallel fractional hot-deck imputation (named as P-FHDI) program suitable for curing large incomplete datasets. Results show a favorable speedup when the P-FHDI is applied to big datasets with up to millions of instances or 10,000 of variables. This paper explains the detailed parallel algorithms of the P-FHDI for large instances (big-$n$) or high-dimensionality (big-$p$) datasets and confirms the favorable scalability. The proposed program inherits all the advantages of the serial FHDI and enables a parallel variance estimation, which will benefit a broad audience in science and engineering.
Yicheng Yang, Jae Kwang Kim, In Ho Cho
IEEE Trans. Knowl. Data Eng.2
2020 Imputation estimators for unnormalized models with missing data
abstract
Several statistical models are given in the form of unnormalized densities and calculation of the normalization constant is intractable. We propose estimation methods for such unnormalized models with missing data. The key concept is to combine imputation techniques with estimators for unnormalized models including noise contrastive estimation and score matching. Further, we derive asymptotic distributions of the proposed estimators and construct confidence intervals. Simulation results with truncated Gaussian graphical models and the application to real data of wind direction demonstrate that the proposed methods enable statistical inference from missing data properly.
Masatoshi Uehara, Takeru Matsuda, Jae Kwang Kim
AISTATS3