Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zonghao Chen

dblp:285/8619 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Probabilistic and Bayesian machine learning · 39% Optimization for machine learning · 24% Trustworthy machine learning · 20%
Theoretical computer science
3 papers
Mathematical optimization · 40% Computational geometry · 40% Graph algorithms and graph theory · 20%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational finance and economics · 50% Medical and health informatics · 50%

Topics — the 25 heaviest of 26, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Graph algorithms and graph theory › graph theory › graph transformation
graph reduction
1.012026
TOPOGRAPH: Topology-Preserving Graph Reduction with Adaptive Structure for Persistent Homology · AAAI 2026
Computational geometry › topological data analysis
persistent homology
1.012026
TOPOGRAPH: Topology-Preserving Graph Reduction with Adaptive Structure for Persistent Homology · AAAI 2026
Computational geometry
topological data analysis
1.012026
TOPOGRAPH: Topology-Preserving Graph Reduction with Adaptive Structure for Persistent Homology · AAAI 2026
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization
0.912025
Nested Expectations with Kernel Quadrature · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning
divergence minimization
0.912025
(De)-regularized Maximum Mean Discrepancy Gradient Flow · J. Mach. Learn. Res. 2025
Machine learning › Optimization for machine learning
gradient flow
0.912025
(De)-regularized Maximum Mean Discrepancy Gradient Flow · J. Mach. Learn. Res. 2025
Mathematical optimization › numerical analysis › numerical integration › quadrature rules
kernel quadrature
0.912025
Nested Expectations with Kernel Quadrature · ICML 2025
Mathematical optimization › numerical analysis
numerical integration
0.912025
Nested Expectations with Kernel Quadrature · ICML 2025
Machine learning › Trustworthy machine learning
uncertainty estimation
0.822024
Tractable Function-Space Variational Inference in Bayesian Neural Networks · NeurIPS 2022
Conformal Counterfactual Inference under Hidden Confounding · KDD 2024
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.812024
Conformal Counterfactual Inference under Hidden Confounding · KDD 2024
Machine learning › Learning theory › statistical estimation › confidence set construction
confidence intervals
0.812024
Conformal Counterfactual Inference under Hidden Confounding · KDD 2024
Machine learning › Trustworthy machine learning › uncertainty estimation
conformal prediction
0.812024
Conformal Counterfactual Inference under Hidden Confounding · KDD 2024
Machine learning › Probabilistic and Bayesian machine learning › causal inference
counterfactual prediction
0.812024
Conformal Counterfactual Inference under Hidden Confounding · KDD 2024
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal effect estimation › treatment effect estimation
individual treatment effect estimation
0.812024
Conformal Counterfactual Inference under Hidden Confounding · KDD 2024
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models › bayesian deep learning
bayesian neural networks
0.612022
Tractable Function-Space Variational Inference in Bayesian Neural Networks · NeurIPS 2022
Machine learning › Optimization for machine learning
bilevel optimization
0.612022
Probabilistic Bilevel Coreset Selection · ICML 2022
Machine learning › Efficient and distributed learning › data selection
coreset selection
0.612022
Probabilistic Bilevel Coreset Selection · ICML 2022
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
functional variational inference
0.612022
Tractable Function-Space Variational Inference in Bayesian Neural Networks · NeurIPS 2022
Machine learning › Trustworthy machine learning › uncertainty estimation
predictive uncertainty
0.612022
Tractable Function-Space Variational Inference in Bayesian Neural Networks · NeurIPS 2022
Machine learning › Efficient and distributed learning › model compression
sparse training
0.512021
Efficient Neural Network Training via Forward and Backward Propagation Sparsification · NeurIPS 2021
Machine learning › Optimization for machine learning
variance reduction
0.512021
Efficient Neural Network Training via Forward and Backward Propagation Sparsification · NeurIPS 2021
Computational finance and economics
option pricing
0.312025
Nested Expectations with Kernel Quadrature · ICML 2025
Mathematical optimization › optimal transport
wasserstein gradient flow
0.312025
(De)-regularized Maximum Mean Discrepancy Gradient Flow · J. Mach. Learn. Res. 2025
Machine learning › Probabilistic and Bayesian machine learning › causal inference
latent confounders
0.212024
Conformal Counterfactual Inference under Hidden Confounding · KDD 2024
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.212022
Probabilistic Bilevel Coreset Selection · ICML 2022

Methods — techniques the papers use, named apart from their topics

nested monte carlo · 2.6multilevel monte carlo · 2.6kernel quadrature · 2.6de-regularization · 1.7adaptive schedule · 1.7landmark-based subsampling · 1.0fermat distances · 1.0transductive weighted conformal prediction · 0.8split conformal prediction · 0.8conformal prediction · 0.8probabilistic weighting · 0.6policy gradient · 0.6
YearPublicationVenuePosition
2026 TOPOGRAPH: Topology-Preserving Graph Reduction with Adaptive Structure for Persistent Homology
abstract
Topological Data Analysis (TDA) provides artificial intelligence (AI) systems with mathematically rigorous geometric descriptors through Persistent Homology (PH), capturing essential shape characteristics in high-dimensional data. Yet, PH’s combinatorial complexity and sensitivity to outliers hinder its scalability and reliability, especially for Intrinsic PH (IPH) that relies on accurate geodesic distances. While stateof-the-art landmark-based subsampling methods, PH Landmarks, ameliorate computational costs and improve outlier robustness by selecting representative points based on local PH scores, it remain computationally intensive and at low sampling rates struggle to reconstruct the global topology. In this work, we introduce TOPOGRAPH, a simple yet powerful framework that preserves intrinsic topology. The resulting coarsened graph supports efficient IPH computations using Fermat distances. Experiments on both synthetic and realworld datasets show that TOPOGRAPH outperforms stateof-the-art sampling-based methods by achieving an order-ofmagnitude speedup and substantially improved topological fidelity in persistence diagrams, demonstrating its ability for robust and scalable topological data analysis.
Zonghao Chen, Yuncheng Jiang 0004
AAAI1
2025 Nested Expectations with Kernel Quadrature
abstract
This paper considers the challenging computational task of estimating nested expectations. Existing algorithms, such as nested Monte Carlo or multilevel Monte Carlo, are known to be consistent but require a large number of samples at both inner and outer levels to converge. Instead, we propose a novel estimator consisting of nested kernel quadrature estimators and we prove that it has a faster convergence rate than all baseline methods when the integrands have sufficient smoothness. We then demonstrate empirically that our proposed method does indeed require the fewest number of samples to estimate nested expectations over a range of real-world application areas from Bayesian optimisation to option pricing and health economics.
Zonghao Chen, Masha Naslidnyk, François-Xavier Briol
ICML1
2025 (De)-regularized Maximum Mean Discrepancy Gradient Flow
abstract
We introduce a (de)-regularization of the Maximum Mean Discrepancy (DrMMD) and its Wasserstein gradient flow. Existing gradient flows that transport samples from source distribution to target distribution with only target samples, either lack tractable numerical implementation ($f$-divergence flows) or require strong assumptions and modifications, such as noise injection, to ensure convergence (Maximum Mean Discrepancy flows). In contrast, DrMMD flow can simultaneously (i) guarantee near-global convergence for a broad class of targets in both continuous and discrete time, and (ii) be implemented in closed form using only samples. The former is achieved by leveraging the connection between the DrMMD and the $\chi^2$-divergence, while the latter comes by treating DrMMD as MMD with a de-regularized kernel. Our numerical scheme employs an adaptive de-regularization schedule throughout the flow to optimally balance the trade-off between discretization errors and deviations from the $\chi^2$ regime. The potential application of the DrMMD flow is demonstrated across several numerical experiments, including a large-scale setting of training student/teacher networks.
Zonghao Chen, Aratrika Mustafi, Pierre Glaser, Anna Korba, Arthur Gretton, Bharath K. Sriperumbudur
J. Mach. Learn. Res.1
2024 Conformal Counterfactual Inference under Hidden Confounding
abstract
Personalized decision making requires the knowledge of potential outcomes under different treatments, and confidence intervals about the potential outcomes further enrich this decision-making process and improve its reliability in high-stakes scenarios. Predicting potential outcomes along with its uncertainty in a counterfactual world poses the foundamental challenge in causal inference. Existing methods that construct confidence intervals for counterfactuals either rely on the assumption of strong ignorability that completely ignores hidden confounders, or need access to un-identifiable lower and upper bounds that characterize the difference between observational and interventional distributions. In this paper, to overcome these limitations, we first propose a novel approach wTCP-DR based on transductive weighted conformal prediction, which provides confidence intervals for counterfactual outcomes with marginal converage guarantees, even under hidden confounding. With less restrictive assumptions, our approach requires access to a fraction of interventional data (from randomized controlled trials) to account for the covariate shift from observational distributoin to interventional distribution. Theoretical results explicitly demonstrate the conditions under which our algorithm is strictly advantageous to the naive method that only uses interventional data. Since transductive conformal prediction is notoriously costly, we propose wSCP-DR, a two-stage variant of wTCP-DR, based on split conformal prediction with same marginal coverage guarantees but at a significantly lower computational cost. After ensuring valid intervals on counterfactuals, it is straightforward to construct intervals for individual treatment effects (ITEs). We demonstrate our method across synthetic and real-world data, including recommendation systems, to verify the superiority of our methods compared against state-of-the-art baselines in terms of both coverage and efficiency. Our code can be found at https://github.com/rguo12/KDD24-Conformal.
Zonghao Chen, Ruocheng Guo, Jean-Francois Ton, Yang Liu 0018
KDD1
2024 Conditional Bayesian Quadrature
abstract
We propose a novel approach for estimating conditional or parametric expectations in the setting where obtaining samples or evaluating integrands is costly. Through the framework of probabilistic numerical methods (such as Bayesian quadrature), our novel approach allows to incorporates prior information about the integrands especially the prior smoothness knowledge about the integrands and the conditional expectation. As a result, our approach provides a way of quantifying uncertainty and leads to a fast convergence rate, which is confirmed both theoretically and empirically on challenging tasks in Bayesian sensitivity analysis, computational finance and decision making under uncertainty.
Zonghao Chen, Masha Naslidnyk, Arthur Gretton, François-Xavier Briol
UAI1
2022 Probabilistic Bilevel Coreset Selection
abstract
The goal of coreset selection in supervised learning is to produce a weighted subset of data, so that training only on the subset achieves similar performance as training on the entire dataset. Existing methods achieved promising results in resource-constrained scenarios such as continual learning and streaming. However, most of the existing algorithms are limited to traditional machine learning models. A few algorithms that can handle large models adopt greedy search approaches due to the difficulty in solving the discrete subset selection problem, which is computationally costly when coreset becomes larger and often produces suboptimal results. In this work, for the first time we propose a continuous probabilistic bilevel formulation of coreset selection by learning a probablistic weight for each training sample. The overall objective is posed as a bilevel optimization problem, where 1) the inner loop samples coresets and train the model to convergence and 2) the outer loop updates the sample probability progressively according to the model’s performance. Importantly, we develop an efficient solver to the bilevel optimization problem via unbiased policy gradient without trouble of implicit differentiation. We theoretically prove the convergence of this training procedure and demonstrate the superiority of our algorithm against various coreset selection methods in various tasks, especially in more challenging label-noise and class-imbalance scenarios.
Renjie Pi, Zonghao Chen, Tong Zhang 0001
ICML5
2022 Tractable Function-Space Variational Inference in Bayesian Neural Networks
abstract
Reliable predictive uncertainty estimation plays an important role in enabling the deployment of neural networks to safety-critical settings. A popular approach for estimating the predictive uncertainty of neural networks is to define a prior distribution over the network parameters, infer an approximate posterior distribution, and use it to make stochastic predictions. However, explicit inference over neural network parameters makes it difficult to incorporate meaningful prior information about the data-generating process into the model. In this paper, we pursue an alternative approach. Recognizing that the primary object of interest in most settings is the distribution over functions induced by the posterior distribution over neural network parameters, we frame Bayesian inference in neural networks explicitly as inferring a posterior distribution over functions and propose a scalable function-space variational inference method that allows incorporating prior information and results in reliable predictive uncertainty estimates. We show that the proposed method leads to state-of-the-art uncertainty estimation and predictive performance on a range of prediction tasks and demonstrate that it performs well on a challenging safety-critical medical diagnosis task in which reliable uncertainty estimation is essential.
Tim G. J. Rudner, Zonghao Chen, Yee Whye Teh, Yarin Gal
NeurIPS2
2021 Efficient Neural Network Training via Forward and Backward Propagation Sparsification
abstract
Sparse training is a natural idea to accelerate the training speed of deep neural networks and save the memory usage, especially since large modern neural networks are significantly over-parameterized. However, most of the existing methods cannot achieve this goal in practice because the chain rule based gradient (w.r.t. structure parameters) estimators adopted by previous methods require dense computation at least in the backward propagation step. This paper solves this problem by proposing an efficient sparse training method with completely sparse forward and backward passes. We first formulate the training process as a continuous minimization problem under global sparsity constraint. We then separate the optimization process into two steps, corresponding to weight update and structure parameter update. For the former step, we use the conventional chain rule, which can be sparse via exploiting the sparse structure. For the latter step, instead of using the chain rule based gradient estimators as in existing methods, we propose a variance reduced policy gradient estimator, which only requires two forward passes without backward propagation, thus achieving completely sparse training. We prove that the variance of our gradient estimator is bounded. Extensive experimental results on real-world datasets demonstrate that compared to previous methods, our algorithm is much more effective in accelerating the training process, up to an order of magnitude faster.
Zonghao Chen, Shizhe Diao, Tong Zhang 0001
NeurIPS3
2021 Distortion-Aware Monocular Depth Estimation for Omnidirectional Images
abstract
Image distortion is a main challenge for tasks on panoramas. In this work, we propose a Distortion-Aware Monocular Omnidirectional (DAMO) network to estimate dense depth maps from indoor panoramas. First, we introduce a distortion-aware module to extract semantic features from omnidirectional images. Specifically, we exploit deformable convolution to adjust its sampling grids to geometric distortions on panoramas. We also utilize a strip pooling module to sample against horizontal distortion introduced by inverse gnomonic projection. Second, we introduce a plug-and-play spherical-aware weight matrix for our loss function to handle the uneven distribution of areas projected from a sphere. Experiments on the 360D dataset show that the proposed method can effectively extract semantic features from distorted panoramas and alleviate the supervision bias caused by distortion. It achieves the state-of-the-art performance on the 360D dataset with high efficiency.
Hong-Xiang Chen, Kunhong Li 0001, Zhiheng Fu, Zonghao Chen, Yulan Guo
IEEE Signal Process. Lett.5
2021 Adv-Depth: Self-Supervised Monocular Depth Estimation With an Adversarial Loss
abstract
Loss function plays a key role in self-supervised monocular depth estimation methods. Current reprojection loss functions are hand-designed and mainly focus on local patch similarity but overlook the global distribution differences between a synthetic image and a target image. In this paper, we leverage global distribution differences by introducing an adversarial loss into the training stage of self-supervised depth estimation. Specifically, we formulate this task as a novel view synthesis problem. We use a depth estimation module and a pose estimation module to form a generator, and then design a discriminator to learn the global distribution differences between real and synthetic images. With the learned global distribution differences, the adversarial loss can be back-propagated to the depth estimation module to improve its performance. Experiments on the KITTI dataset have demonstrated the effectiveness of the adversarial loss. The adversarial loss is further combined with the reprojection loss to achieve the state-of-the-art performance on the KITTI dataset.
Kunhong Li 0001, Zhiheng Fu, Hanyun Wang, Zonghao Chen, Yulan Guo
IEEE Signal Process. Lett.4