VLDB 2026 Research / reviewers in the wild / expert
Jia-Jie Zhu
dblp:195/5802
· DBLP profile ↗
7ranked-venue papers
2as first author
6since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Probabilistic and Bayesian machine learning · 54% Learning theory · 28% Reinforcement learning · 9% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
conditional moment restrictions |
0.8 | 2 | 2023 | Functional Generalized Empirical Likelihood Estimation for Conditional Moment Restrictions · ICML 2022 Estimation Beyond Data Reweighting: Kernel Method of Moments · ICML 2023 |
Machine learning › Probabilistic and Bayesian machine learning
sampling |
0.8 | 1 | 2024 | Interaction-Force Transport Gradient Flows · NeurIPS 2024 |
Mathematical optimization › continuous optimization › convex optimization › first-order methods › gradient-based optimization
gradient flow |
0.8 | 1 | 2024 | Interaction-Force Transport Gradient Flows · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
method of moments |
0.7 | 1 | 2023 | Estimation Beyond Data Reweighting: Kernel Method of Moments · ICML 2023 |
Machine learning › Reinforcement learning › exploration
intrinsically motivated reinforcement learning |
0.4 | 1 | 2019 | Control What You Can: Intrinsically Motivated Task-Planning Agent · NeurIPS 2019 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
task planning |
0.4 | 1 | 2019 | Control What You Can: Intrinsically Motivated Task-Planning Agent · NeurIPS 2019 |
Mathematical optimization
optimal transport |
0.2 | 1 | 2024 | Interaction-Force Transport Gradient Flows · NeurIPS 2024 |
Mathematical optimization › optimal transport
unbalanced optimal transport |
0.2 | 1 | 2024 | Interaction-Force Transport Gradient Flows · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
0.2 | 1 | 2022 | Functional Generalized Empirical Likelihood Estimation for Conditional Moment Restrictions · ICML 2022 |
Methods — techniques the papers use, named apart from their topics
maximum mean discrepancy · 2.2reproducing kernel hilbert space · 1.5JKO splitting · 1.5generalized method of moments · 0.7empirical likelihood · 0.7neural network · 0.6kernel methods · 0.6generalized empirical likelihood · 0.6intrinsic motivation · 0.4backtracking search · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Analysis of Kernel Mirror Prox for Measure OptimizationabstractBy choosing a suitable function space as the dual to the non-negative measure cone, we study in a unified framework a class of functional saddle-point optimization problems, which we term the Mixed Functional Nash Equilibrium (MFNE), that underlies several existing machine learning algorithms, such as implicit generative models, distributionally robust optimization (DRO), and Wasserstein barycenters. We model the saddle-point optimization dynamics as an interacting Fisher-Rao-RKHS gradient flow when the function space is chosen as a reproducing kernel Hilbert space (RKHS). As a discrete time counterpart, we propose a primal-dual kernel mirror prox (KMP) algorithm, which uses a dual step in the RKHS, and a primal entropic mirror prox step. We then provide a unified convergence analysis of KMP in an infinite-dimensional setting for this class of MFNE problems, which establishes a convergence rate of $O(1/N)$ in the deterministic case and $O(1/\sqrt{N})$ in the stochastic case, where $N$ is the iteration counter. As a case study, we apply our analysis to DRO, providing algorithmic guarantees for DRO robustness and convergence. Pavel E. Dvurechensky, Jia-Jie Zhu |
AISTATS | 2 |
| 2024 | Interaction-Force Transport Gradient FlowsabstractThis paper presents a new gradient flow dissipation geometry over non-negative and probability measures.
This is motivated by a principled construction that combines the unbalanced optimal transport and interaction forces modeled by reproducing kernels. Using a precise connection between the Hellinger geometry and the maximum mean discrepancy (MMD), we propose the interaction-force transport (IFT) gradient flows and its spherical variant via an infimal convolution of the Wasserstein and spherical MMD tensors. We then develop a particle-based optimization algorithm based on the JKO-splitting scheme of the mass-preserving spherical IFT gradient flows. Finally, we provide both theoretical global exponential convergence guarantees and improved empirical simulation results for applying the IFT gradient flows to the sampling task of MMD-minimization. Furthermore, we prove that the spherical IFT gradient flow enjoys the best of both worlds by providing the global exponential convergence guarantee for both the MMD and KL energy. Egor Gladin, Pavel E. Dvurechensky, Alexander Mielke, Jia-Jie Zhu |
NeurIPS | 4 |
| 2023 | Estimation Beyond Data Reweighting: Kernel Method of MomentsabstractMoment restrictions and their conditional counterparts emerge in many areas of machine learning and statistics ranging from causal inference to reinforcement learning. Estimators for these tasks, generally called methods of moments, include the prominent generalized method of moments (GMM) which has recently gained attention in causal inference. GMM is a special case of the broader family of empirical likelihood estimators which are based on approximating a population distribution by means of minimizing a $\varphi$-divergence to an empirical distribution. However, the use of $\varphi$-divergences effectively limits the candidate distributions to reweightings of the data samples. We lift this long-standing limitation and provide a method of moments that goes beyond data reweighting. This is achieved by defining an empirical likelihood estimator based on maximum mean discrepancy which we term the kernel method of moments (KMM). We provide a variant of our estimator for conditional moment restrictions and show that it is asymptotically first-order optimal for such problems. Finally, we show that our method achieves competitive performance on several conditional moment restriction tasks. Heiner Kremer, Yassine Nemmour, Bernhard Schölkopf, Jia-Jie Zhu |
ICML | 4 |
| 2022 | Adversarially Robust Kernel SmoothingabstractWe propose a scalable robust learning algorithm combining kernel smoothing and robust optimization. Our method is motivated by the convex analysis perspective of distributionally robust optimization based on probability metrics, such as the Wasserstein distance and the maximum mean discrepancy. We adapt the integral operator using supremal convolution in convex analysis to form a novel function majorant used for enforcing robustness. Our method is simple in form and applies to general loss functions and machine learning models. Exploiting a connection with optimal transport, we prove theoretical guarantees for certified robustness under distribution shift. Furthermore, we report experiments with general machine learning models, such as deep neural networks, to demonstrate competitive performance with the state-of-the-art certifiable robust learning algorithms based on the Wasserstein distance. Jia-Jie Zhu, Christina Kouridi, Yassine Nemmour, Bernhard Schölkopf |
AISTATS | 1 |
| 2022 | Functional Generalized Empirical Likelihood Estimation for Conditional Moment RestrictionsabstractImportant problems in causal inference, economics, and, more generally, robust machine learning can be expressed as conditional moment restrictions, but estimation becomes challenging as it requires solving a continuum of unconditional moment restrictions. Previous works addressed this problem by extending the generalized method of moments (GMM) to continuum moment restrictions. In contrast, generalized empirical likelihood (GEL) provides a more general framework and has been shown to enjoy favorable small-sample properties compared to GMM-based estimators. To benefit from recent developments in machine learning, we provide a functional reformulation of GEL in which arbitrary models can be leveraged. Motivated by a dual formulation of the resulting infinite dimensional optimization problem, we devise a practical method and explore its asymptotic properties. Finally, we provide kernel- and neural network-based implementations of the estimator, which achieve state-of-the-art empirical performance on two conditional moment restriction problems. Heiner Kremer, Jia-Jie Zhu, Krikamol Muandet, Bernhard Schölkopf |
ICML | 2 |
| 2021 | Kernel Distributionally Robust Optimization: Generalized Duality Theorem and Stochastic ApproximationabstractWe propose kernel distributionally robust optimization (Kernel DRO) using insights from the robust optimization theory and functional analysis. Our method uses reproducing kernel Hilbert spaces (RKHS) to construct a wide range of convex ambiguity sets, which can be generalized to sets based on integral probability metrics and finite-order moment bounds. This perspective unifies multiple existing robust and stochastic optimization methods. We prove a theorem that generalizes the classical duality in the mathematical problem of moments. Enabled by this theorem, we reformulate the maximization with respect to measures in DRO into the dual program that searches for RKHS functions. Using universal RKHSs, the theorem applies to a broad class of loss functions, lifting common limitations such as polynomial losses and knowledge of the Lipschitz constant. We then establish a connection between DRO and stochastic optimization with expectation constraints. Finally, we propose practical algorithms based on both batch convex solvers and stochastic functional gradient, which apply to general optimization and machine learning tasks. Jia-Jie Zhu, Wittawat Jitkrittum, Moritz Diehl, Bernhard Schölkopf |
AISTATS | 1 |
| 2019 | Control What You Can: Intrinsically Motivated Task-Planning AgentabstractWe present a novel intrinsically motivated agent that learns how to control the environment in a sample efficient manner, that is with as few environment interactions as possible, by optimizing learning progress. It learns what can be controlled, how to allocate time and attention as well as the relations between objects using surprise-based motivation. The effectiveness of our method is demonstrated in a synthetic and robotic manipulation environment yielding considerably improved performance and smaller sample complexity compared to an intrinsically motivated, non-hierarchical and state-of-the-art hierarchical baseline. In a nutshell, our work combines several task-level planning agent structures (backtracking search on task-graph, probabilistic road-maps, allocation of search efforts) with intrinsic motivation to achieve learning from scratch. Sebastian Blaes, Marin Vlastelica Pogancic, Jia-Jie Zhu, Georg Martius |
NeurIPS | 3 |