EDBT 2026 Demo / reviewers in the wild / expert
Jae Kwang Kim
dblp:238/1626
· DBLP profile ↗
4ranked-venue papers
0as first author
3since 2021 · last 2023
0000-0002-0246-6029ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Theoretical computer science
1 paper |
Mathematical optimization · 67% Algorithms and data structures · 33% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Parallel and multicore computing · 100% | |
| Artificial intelligence
1 paper |
Deep learning architectures and training · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational science and engineering · 100% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing
parallel algorithms |
1.2 | 2 | 2023 | Ultra Data-Oriented Parallel Fractional Hot-Deck Imputation With Efficient Linearized Variance Estimation · IEEE Trans. Knowl. Data Eng. 2023 Parallel Fractional Hot-Deck Imputation and Variance Estimation for Big Incomplete Data Curing · IEEE Trans. Knowl. Data Eng. 2022 |
Mathematical optimization › statistical estimation
maximum likelihood estimation |
0.6 | 1 | 2022 | Maximum sampled conditional likelihood for informative subsampling · J. Mach. Learn. Res. 2022 |
Mathematical optimization
statistical estimation |
0.6 | 1 | 2022 | Maximum sampled conditional likelihood for informative subsampling · J. Mach. Learn. Res. 2022 |
Algorithms and data structures › randomized algorithms › sampling
subsampling |
0.6 | 1 | 2022 | Maximum sampled conditional likelihood for informative subsampling · J. Mach. Learn. Res. 2022 |
Computational science and engineering
statistical computing |
0.2 | 1 | 2022 | Parallel Fractional Hot-Deck Imputation and Variance Estimation for Big Incomplete Data Curing · IEEE Trans. Knowl. Data Eng. 2022 |
Methods — techniques the papers use, named apart from their topics
fractional hot-deck imputation · 2.5linearization · 1.3jackknife · 1.3jackknife variance estimation · 1.1inverse probability weighting · 0.6generalized linear model · 0.6asymptotic analysis · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Ultra Data-Oriented Parallel Fractional Hot-Deck Imputation With Efficient Linearized Variance EstimationabstractParallel fractional hot-deck imputation (P-FHDI (Yang et al. 2020)) is a general-purpose, assumption-free tool for handling item nonresponse in big incomplete data by combining the theory of FHDI and parallel computing. FHDI cures multivariate missing data by filling each missing unit with multiple observed values (thus, hot-deck) without resorting to distributional assumptions. P-FHDI can tackle big incomplete data with millions of instances (big-$n$) or 10,000 variables (big-$p$). However, handling ultra incomplete data (i.e., concurrently big-$n$and big-$p$) with tremendous instances and high dimensionality has posed problems to P-FHDI due to excessive memory requirement and execution time. To tackle the aforementioned challenges, we propose the ultra data-oriented P-FHDI (named UP-FHDI) capable of curing ultra incomplete data. In addition to the parallel Jackknife method, this paper enables a computationally efficient ultra data-oriented variance estimation using parallel linearization techniques. Results confirm that UP-FHDI can tackle an ultra dataset with one million instances and 10,000 variables. This paper illustrates the special parallel algorithms of UP-FHDI and confirms its positive impact on the subsequent deep learning performance. Yicheng Yang, Yonghyun Kwon, Jae Kwang Kim, In Ho Cho |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Maximum sampled conditional likelihood for informative subsamplingabstractSubsampling is a computationally effective approach to extract information from massive data sets when computing resources are limited. After a subsample is taken from the full data, most available methods use an inverse probability weighted (IPW) objective function to estimate the model parameters. The IPW estimator does not fully utilize the information in the selected subsample. In this paper, we propose to use the maximum sampled conditional likelihood estimator (MSCLE) based on the sampled data. We established the asymptotic normality of the MSCLE and prove that its asymptotic variance covariance matrix is the smallest among a class of asymptotically unbiased estimators, including the IPW estimator. We further discuss the asymptotic results with the L-optimal subsampling probabilities and illustrate the estimation procedure with generalized linear models. Numerical experiments are provided to evaluate the practical performance of the proposed method. HaiYing Wang 0004, Jae Kwang Kim |
J. Mach. Learn. Res. | 2 |
| 2022 | Parallel Fractional Hot-Deck Imputation and Variance Estimation for Big Incomplete Data CuringabstractThe fractional hot-deck imputation (FHDI) is a general-purpose, assumption-free imputation method for handling multivariate missing data by filling each missing item with multiple observed values without resorting to artificially created values. The corresponding R package FHDI J. Im, I. Cho, and J. K. Kim, “An R package for fractional hot deck imputation,”R J., vol. 10, no. 1, pp. 140–154, 2018 holds generality and efficiency, but it is not adequate for tackling big incomplete data due to the requirement of excessive memory and long running time. As a first step to tackle big incomplete data by leveraging the FHDI, we developed a new version of a parallel fractional hot-deck imputation (named as P-FHDI) program suitable for curing large incomplete datasets. Results show a favorable speedup when the P-FHDI is applied to big datasets with up to millions of instances or 10,000 of variables. This paper explains the detailed parallel algorithms of the P-FHDI for large instances (big-$n$) or high-dimensionality (big-$p$) datasets and confirms the favorable scalability. The proposed program inherits all the advantages of the serial FHDI and enables a parallel variance estimation, which will benefit a broad audience in science and engineering. Yicheng Yang, Jae Kwang Kim, In Ho Cho |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2020 | Imputation estimators for unnormalized models with missing dataabstractSeveral statistical models are given in the form of unnormalized densities and calculation of the normalization constant is intractable. We propose estimation methods for such unnormalized models with missing data. The key concept is to combine imputation techniques with estimators for unnormalized models including noise contrastive estimation and score matching. Further, we derive asymptotic distributions of the proposed estimators and construct confidence intervals. Simulation results with truncated Gaussian graphical models and the application to real data of wind direction demonstrate that the proposed methods enable statistical inference from missing data properly. Masatoshi Uehara, Takeru Matsuda, Jae Kwang Kim |
AISTATS | 3 |