VLDB 2026 Research / reviewers in the wild / expert
Jeongyoun Ahn
dblp:76/6516
· DBLP profile ↗
7ranked-venue papers
0as first author
6since 2021 · last 2024
0000-0002-4351-5798ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Kernel, tree and ensemble methods · 38% Learning theory · 22% Probabilistic and Bayesian machine learning · 20% | |
| Network and information security
2 papers |
Privacy and data protection · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Bioinformatics and computational biology · 100% |
Topics — the 21 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Kernel, tree and ensemble methods
kernel methods |
1.2 | 2 | 2023 | Reproducing Kernels and New Approaches in Compositional Data Analysis · J. Mach. Learn. Res. 2023 Kernel Methods for Radial Transformed Compositional Data with Many Zeros · ICML 2022 |
Machine learning › Learning theory
empirical risk minimization |
0.8 | 1 | 2024 | Differential Privacy in Scalable General Kernel Learning via $K$-means Nystr{\"o}m Random Features · NeurIPS 2024 |
Privacy and data protection › differential privacy
differentially private learning |
0.8 | 1 | 2024 | Differential Privacy in Scalable General Kernel Learning via $K$-means Nystr{\"o}m Random Features · NeurIPS 2024 |
Privacy and data protection
differential privacy |
0.8 | 1 | 2024 | Differential Privacy in Scalable General Kernel Learning via $K$-means Nystr{\"o}m Random Features · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning
compositional data analysis |
0.7 | 1 | 2023 | Reproducing Kernels and New Approaches in Compositional Data Analysis · J. Mach. Learn. Res. 2023 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › density estimation
nonparametric density estimation |
0.7 | 1 | 2023 | Reproducing Kernels and New Approaches in Compositional Data Analysis · J. Mach. Learn. Res. 2023 |
Machine learning › Kernel, tree and ensemble methods › kernel methods
reproducing kernel hilbert space |
0.7 | 1 | 2023 | Reproducing Kernels and New Approaches in Compositional Data Analysis · J. Mach. Learn. Res. 2023 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
sufficient dimension reduction |
0.7 | 1 | 2023 | Kernel Sufficient Dimension Reduction and Variable Selection for Compositional Data via Amalgamation · ICML 2023 |
Machine learning › Learning theory › model selection
variable selection |
0.7 | 1 | 2023 | Kernel Sufficient Dimension Reduction and Variable Selection for Compositional Data via Amalgamation · ICML 2023 |
Privacy and data protection › differential privacy
local differential privacy |
0.7 | 1 | 2023 | Minimax Risks and Optimal Procedures for Estimation under Functional Local Differential Privacy · NeurIPS 2023 |
Privacy and data protection › differential privacy › local differential privacy
mean estimation |
0.7 | 1 | 2023 | Minimax Risks and Optimal Procedures for Estimation under Functional Local Differential Privacy · NeurIPS 2023 |
Machine learning › Kernel, tree and ensemble methods
kernel embedding |
0.6 | 1 | 2022 | Kernel Methods for Radial Transformed Compositional Data with Many Zeros · ICML 2022 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction › principal component analysis
kernel principal component analysis |
0.6 | 1 | 2022 | Kernel Methods for Radial Transformed Compositional Data with Many Zeros · ICML 2022 |
Bioinformatics and computational biology › drug discovery
drug response prediction |
0.5 | 1 | 2021 | Feature-weighted ordinal classification for predicting drug response in multiple myeloma · Bioinform. 2021 |
Bioinformatics and computational biology › genomics
pharmacogenomics |
0.5 | 1 | 2021 | Feature-weighted ordinal classification for predicting drug response in multiple myeloma · Bioinform. 2021 |
Bioinformatics and computational biology › computational microbiology
microbiome analysis |
0.4 | 2 | 2023 | Kernel Sufficient Dimension Reduction and Variable Selection for Compositional Data via Amalgamation · ICML 2023 Kernel Methods for Radial Transformed Compositional Data with Many Zeros · ICML 2022 |
Bioinformatics and computational biology
multiple myeloma |
0.1 | 1 | 2021 | Feature-weighted ordinal classification for predicting drug response in multiple myeloma · Bioinform. 2021 |
Bioinformatics and computational biology › cancer genomics
cancer classification |
0.1 | 1 | 2006 | Gene selection using support vector machines with non-convex penalty · Bioinform. 2006 |
Bioinformatics and computational biology
gene expression analysis |
0.1 | 1 | 2006 | Gene selection using support vector machines with non-convex penalty · Bioinform. 2006 |
Bioinformatics and computational biology › gene expression analysis
gene selection |
0.1 | 1 | 2006 | Gene selection using support vector machines with non-convex penalty · Bioinform. 2006 |
Bioinformatics and computational biology › kernel methods
support vector machine |
0.1 | 1 | 2006 | Gene selection using support vector machines with non-convex penalty · Bioinform. 2006 |
Methods — techniques the papers use, named apart from their topics
kernel methods · 2.5log-ratio transformation · 1.8random features · 1.5nyström method · 1.5k-means · 1.5sub-composition · 1.3amalgamation · 1.3radial transformation · 1.1support vector machine · 0.7spherical harmonics · 0.7minimax analysis · 0.7information-theoretic bounds · 0.7Gaussian LDP · 0.7ordinal classification · 0.5linear discriminant analysis · 0.5feature weighting · 0.5successive quadratic algorithm · 0.1smoothly clipped absolute deviation penalty · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Differential Privacy in Scalable General Kernel Learning via $K$-means Nystr{\"o}m Random FeaturesabstractAs the volume of data invested in statistical learning increases and concerns regarding privacy grow, the privacy leakage issue has drawn significant attention. Differential privacy has emerged as a widely accepted concept capable of mitigating privacy concerns, and numerous differentially private (DP) versions of machine learning algorithms have been developed. However, existing works on DP kernel learning algorithms have exhibited practical limitations, including scalability, restricted choice of kernels, or dependence on test data availability. We propose DP scalable kernel empirical risk minimization (ERM) algorithms and a DP kernel mean embedding (KME) release algorithm suitable for general kernels. Our approaches address the shortcomings of previous algorithms by employing Nyström methods, classical techniques in non-private scalable kernel learning. These methods provide data-dependent low-rank approximations of the kernel matrix for general kernels in a DP manner. We present excess empirical risk bounds and computational complexities for the scalable kernel DP ERM, KME algorithms, contrasting them with established methodologies. Furthermore, we develop a private data-generating algorithm capable of learning diverse kernel models. We conduct experiments to demonstrate the performance of our algorithms, comparing them with existing methods to highlight their superiority. Bonwoo Lee, Jeongyoun Ahn, Cheolwoo Park |
NeurIPS | 2 |
| 2023 | Kernel Sufficient Dimension Reduction and Variable Selection for Compositional Data via AmalgamationabstractCompositional data with a large number of components and an abundance of zeros are frequently observed in many fields recently. Analyzing such sparse high-dimensional compositional data naturally calls for dimension reduction or, more preferably, variable selection. Most existing approaches lack interpretability or cannot handle zeros properly, as they rely on a log-ratio transformation. We approach this problem with sufficient dimension reduction (SDR), one of the most studied dimension reduction frameworks in statistics. Characterized by the conditional independence of the data to the response on the found subspace, the SDR framework has been effective for both linear and nonlinear dimension reduction problems. This work proposes a compositional SDR that can handle zeros naturally while incorporating the nonlinear nature and spurious negative correlations among components rigorously. A critical consideration of sub-composition versus amalgamation for compositional variable selection is discussed. The proposed compositional SDR is shown to be statistically consistent in constructing a sub-simplex consisting of true signal variables. Simulation and real microbiome data are used to demonstrate the performance of the proposed SDR compared to existing state-of-art approaches. Jeongyoun Ahn, Cheolwoo Park |
ICML | 2 |
| 2023 | Minimax Risks and Optimal Procedures for Estimation under Functional Local Differential PrivacyabstractAs concerns about data privacy continue to grow, differential privacy (DP) has emerged as a fundamental concept that aims to guarantee privacy by ensuring individuals' indistinguishability in data analysis. Local differential privacy (LDP) is a rigorous type of DP that requires individual data to be privatized before being sent to the collector, thus removing the need for a trusted third party to collect data. Among the numerous (L)DP-based approaches, functional DP has gained considerable attention in the DP community because it connects DP to statistical decision-making by formulating it as a hypothesis-testing problem and also exhibits Gaussian-related properties. However, the utility of privatized data is generally lower than that of non-private data, prompting research into optimal mechanisms that maximize the statistical utility for given privacy constraints. In this study, we investigate how functional LDP preserves the statistical utility by analyzing minimax risks of univariate mean estimation as well as nonparametric density estimation. We leverage the contraction property of functional LDP mechanisms and classical information-theoretical bounds to derive private minimax lower bounds. Our theoretical study reveals that it is possible to establish an interpretable, continuous balance between the statistical utility and privacy level, which has not been achieved under the $\epsilon$-LDP framework. Furthermore, we suggest minimax optimal mechanisms based on Gaussian LDP (a type of functional LDP) that achieve the minimax upper bounds and show via a numerical study that they are superior to the counterparts derived under $\epsilon$-LDP. The theoretical and empirical findings of this work suggest that Gaussian LDP should be considered a reliable standard for LDP. Bonwoo Lee, Jeongyoun Ahn, Cheolwoo Park |
NeurIPS | 2 |
| 2023 | Reproducing Kernels and New Approaches in Compositional Data AnalysisabstractCompositional data, such as human gut microbiomes, consist of non-negative variables where only the relative values of these variables are available. Analyzing compositional data requires careful treatment of the geometry of the data. A common geometrical approach to understanding such data is through a regular simplex. The majority of existing approaches rely on log-ratio or power transformations to address the inherent simplicial geometry. In this work, based on the key observation that compositional data are projective, we reinterpret the compositional domain as a group quotient of a sphere, leveraging the intrinsic connection between projective and spherical geometry. This interpretation enables us to understand the function spaces on the compositional domain in terms of those on a sphere, and furthermore, to utilize spherical harmonics theory for constructing a compositional Reproducing Kernel Hilbert Space (RKHS). The construction of RKHS for compositional data opens up new research avenues for future methodology developments, particularly introducing well-developed kernel methods to compositional data analysis. We demonstrate the wide applicability of the proposed theoretical framework with examples of nonparametric density estimation, kernel exponential family, and support vector machine for compositional data. Changwon Yoon, Jeongyoun Ahn |
J. Mach. Learn. Res. | 3 |
| 2022 | Kernel Methods for Radial Transformed Compositional Data with Many ZerosabstractCompositional data analysis with a high proportion of zeros has gained increasing popularity, especially in chemometrics and human gut microbiomes research. Statistical analyses of this type of data are typically carried out via a log-ratio transformation after replacing zeros with small positive values. We should note, however, that this procedure is geometrically improper, as it causes anomalous distortions through the transformation. We propose a radial transformation that does not require zero substitutions and more importantly results in essential equivalence between domains before and after the transformation. We show that a rich class of kernels on hyperspheres can successfully define a kernel embedding for compositional data based on this equivalence. To the best of our knowledge, this is the first work that theoretically establishes the availability of the extensive library of kernel-based machine learning methods for compositional data. The applicability of the proposed approach is demonstrated with kernel principal component analysis. Changwon Yoon, Cheolwoo Park, Jeongyoun Ahn |
ICML | 4 |
| 2021 | Feature-weighted ordinal classification for predicting drug response in multiple myelomaabstractMOTIVATION: Ordinal classification problems arise in a variety of real-world applications, in which samples need to be classified into categories with a natural ordering. An example of classifying high-dimensional ordinal data is to use gene expressions to predict the ordinal drug response, which has been increasingly studied in pharmacogenetics. Classical ordinal classification methods are typically not able to tackle high-dimensional data and standard high-dimensional classification methods discard the ordering information among the classes. Existing work of high-dimensional ordinal classification approaches usually assume a linear ordinality among the classes. We argue that manually labeled ordinal classes may not be linearly arranged in the data space, especially in high-dimensional complex problems. RESULTS: We propose a new approach that can project high-dimensional data into a lower discriminating subspace, where the innate ordinal structure of the classes is uncovered. The proposed method weights the features based on their rank correlations with the class labels and incorporates the weights into the framework of linear discriminant analysis. We apply the method to predict the response to two types of drugs for patients with multiple myeloma, respectively. A comparative analysis with both ordinal and nominal existing methods demonstrates that the proposed method can achieve a competitive predictive performance while honoring the intrinsic ordinal structure of the classes. We provide interpretations on the genes that are selected by the proposed approach to understand their drug-specific response mechanisms. AVAILABILITY AND IMPLEMENTATION: The data underlying this article are available in the Gene Expression Omnibus Database at https://www.ncbi.nlm.nih.gov/geo/ and can be accessed with accession number GSE9782 and GSE68871. The source code for FWOC can be accessed at https://github.com/pisuduo/Feature-Weighted-Ordinal-Classification-FWOC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jeongyoun Ahn |
Bioinform. | 2 |
| 2006 | Gene selection using support vector machines with non-convex penaltyabstractMOTIVATION: With the development of DNA microarray technology, scientists can now measure the expression levels of thousands of genes simultaneously in one single experiment. One current difficulty in interpreting microarray data comes from their innate nature of 'high-dimensional low sample size'. Therefore, robust and accurate gene selection methods are required to identify differentially expressed group of genes across different samples, e.g. between cancerous and normal cells. Successful gene selection will help to classify different cancer types, lead to a better understanding of genetic signatures in cancers and improve treatment strategies. Although gene selection and cancer classification are two closely related problems, most existing approaches handle them separately by selecting genes prior to classification. We provide a unified procedure for simultaneous gene selection and cancer classification, achieving high accuracy in both aspects. RESULTS: In this paper we develop a novel type of regularization in support vector machines (SVMs) to identify important genes for cancer classification. A special nonconvex penalty, called the smoothly clipped absolute deviation penalty, is imposed on the hinge loss function in the SVM. By systematically thresholding small estimates to zeros, the new procedure eliminates redundant genes automatically and yields a compact and accurate classifier. A successive quadratic algorithm is proposed to convert the non-differentiable and non-convex optimization problem into easily solved linear equation systems. The method is applied to two real datasets and has produced very promising results. AVAILABILITY: MATLAB codes are available upon request from the authors. Hao Helen Zhang 0001, Jeongyoun Ahn, Xiaodong Lin 0004, Cheolwoo Park |
Bioinform. | 2 |