Rentian Yao

dblp:323/5858 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
5 papers
Mathematical optimization · 92% Graph algorithms and graph theory · 8%
Artificial intelligence
4 papers
Probabilistic and Bayesian machine learning · 62% Generative modeling · 20% Learning theory · 13%

Topics — the 17 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Mathematical optimization
optimal transport
1.722025
Statistical Analysis of the Sinkhorn Iterations for Two-Sample Schrödinger Bridge Estimation · NeurIPS 2025
Optimal Transport Barycenter via Nonconvex-Concave Minimax Optimization · ICML 2025
Mathematical optimization
continuous optimization
1.622025
Optimal Transport Barycenter via Nonconvex-Concave Minimax Optimization · ICML 2025
Wasserstein Proximal Coordinate Gradient Algorithms · J. Mach. Learn. Res. 2024
Machine learning › Generative modeling › diffusion model
schrödinger bridge
0.912025
Statistical Analysis of the Sinkhorn Iterations for Two-Sample Schrödinger Bridge Estimation · NeurIPS 2025
Mathematical optimization
minimax optimization
0.912025
Optimal Transport Barycenter via Nonconvex-Concave Minimax Optimization · ICML 2025
Mathematical optimization › optimal transport › entropic optimal transport
sinkhorn algorithm
0.912025
Statistical Analysis of the Sinkhorn Iterations for Two-Sample Schrödinger Bridge Estimation · NeurIPS 2025
Mathematical optimization › optimal transport
wasserstein barycenter
0.912025
Optimal Transport Barycenter via Nonconvex-Concave Minimax Optimization · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
mean-field approximation
0.812024
Wasserstein Proximal Coordinate Gradient Algorithms · J. Mach. Learn. Res. 2024
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.812024
Wasserstein Proximal Coordinate Gradient Algorithms · J. Mach. Learn. Res. 2024
Mathematical optimization › riemannian optimization
geodesically convex optimization
0.812024
Wasserstein Proximal Coordinate Gradient Algorithms · J. Mach. Learn. Res. 2024
Graph algorithms and graph theory › graph simplification
graph coarsening
0.712023
A Gromov-Wasserstein Geometric View of Spectrum-Preserving Graph Coarsening · ICML 2023
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
drift estimation
0.612022
Mean-field nonparametric estimation of interacting particle systems · COLT 2022
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
interacting particle systems
0.612022
Mean-field nonparametric estimation of interacting particle systems · COLT 2022
Machine learning › Learning theory › statistical estimation
nonparametric estimation
0.612022
Mean-field nonparametric estimation of interacting particle systems · COLT 2022
Mathematical optimization › optimal transport
wasserstein space optimization
0.212024
Wasserstein Proximal Coordinate Gradient Algorithms · J. Mach. Learn. Res. 2024
Machine learning › Graph learning
graph classification
0.212023
A Gromov-Wasserstein Geometric View of Spectrum-Preserving Graph Coarsening · ICML 2023
Mathematical optimization › statistical estimation
maximum likelihood estimation
0.212022
Mean-field nonparametric estimation of interacting particle systems · COLT 2022
Mathematical optimization
statistical estimation
0.212022
Mean-field nonparametric estimation of interacting particle systems · COLT 2022

Methods — techniques the papers use, named apart from their topics

statistical analysis · 1.7sinkhorn algorithm · 1.7quadratic growth condition · 1.5proximal coordinate gradient · 1.5parallel/sequential/random updates · 1.5weighted kernel k-means · 1.3gromov-wasserstein distance · 1.3sobolev geometry · 0.9primal-dual algorithm · 0.9kantorovich potential · 0.9rademacher complexity · 0.6gaussian complexity · 0.6fourier deconvolution · 0.6
YearPublicationVenuePosition
2025 Optimal Transport Barycenter via Nonconvex-Concave Minimax Optimization
abstract
The optimal transport barycenter (a.k.a. Wasserstein barycenter) is a fundamental notion of averaging that extends from the Euclidean space to the Wasserstein space of probability distributions. Computation of the unregularized barycenter for discretized probability distributions on point clouds is a challenging task when the domain dimension $d > 1$. Most practical algorithms for approximating the barycenter problem are based on entropic regularization. In this paper, we introduce a nearly linear time $O(m \log{m})$ and linear space complexity $O(m)$ primal-dual algorithm, the Wasserstein-Descent $\dot{\mathbb{H}}^1$-Ascent (WDHA) algorithm, for computing the exact barycenter when the input probability density functions are discretized on an $m$-point grid. The key success of the WDHA algorithm hinges on alternating between two different yet closely related Wasserstein and Sobolev optimization geometries for the primal barycenter and dual Kantorovich potential subproblems. Under reasonable assumptions, we establish the convergence rate and iteration complexity of WDHA to its stationary point when the step size is appropriately chosen. Superior computational efficacy, scalability, and accuracy over the existing Sinkhorn-type algorithms are demonstrated on high-resolution (e.g., $1024 \times 1024$ images) 2D synthetic and real data.
Kaheon Kim, Rentian Yao, Changbo Zhu
ICML2
2025 Statistical Analysis of the Sinkhorn Iterations for Two-Sample Schrödinger Bridge Estimation
Ibuki Maeda, Rentian Yao, Atsushi Nitanda
NeurIPS2
2024 Minimizing Convex Functionals over Space of Probability Measures via KL Divergence Gradient Flow
abstract
Motivated by the computation of the non-parametric maximum likelihood estimator (NPMLE) and the Bayesian posterior in statistics, this paper explores the problem of convex optimization over the space of all probability distributions. We introduce an implicit scheme, called the implicit KL proximal descent (IKLPD) algorithm, for discretizing a continuous-time gradient flow relative to the Kullback–Leibler (KL) divergence for minimizing a convex target functional. We show that IKLPD converges to a global optimum at a polynomial rate from any initialization; moreover, if the objective functional is strongly convex relative to the KL divergence, for example, when the target functional itself is a KL divergence as in the context of Bayesian posterior computation, IKLPD exhibits globally exponential convergence. Computationally, we propose a numerical method based on normalizing flow to realize IKLPD. Conversely, our numerical method can also be viewed as a new approach that sequentially trains a normalizing flow for minimizing a convex functional with a strong theoretical guarantee.
Rentian Yao, Linjun Huang
AISTATS1
2024 Wasserstein Proximal Coordinate Gradient Algorithms
abstract
Motivated by approximation Bayesian computation using mean-field variational approximation and the computation of equilibrium in multi-species systems with cross-interaction, this paper investigates the composite geodesically convex optimization problem over multiple distributions. The objective functional under consideration is composed of a convex potential energy on a product of Wasserstein spaces and a sum of convex self-interaction and internal energies associated with each distribution. To efficiently solve this problem, we introduce the Wasserstein Proximal Coordinate Gradient (WPCG) algorithms with parallel, sequential, and random update schemes. Under a quadratic growth (QG) condition that is weaker than the usual strong convexity requirement on the objective functional, we show that WPCG converges exponentially fast to the unique global optimum. In the absence of the QG condition, WPCG is still demonstrated to converge to the global optimal solution, albeit at a slower polynomial rate. Numerical results for both motivating examples are consistent with our theoretical findings.
Rentian Yao
J. Mach. Learn. Res.1
2023 A Gromov-Wasserstein Geometric View of Spectrum-Preserving Graph Coarsening
abstract
Graph coarsening is a technique for solving large-scale graph problems by working on a smaller version of the original graph, and possibly interpolating the results back to the original graph. It has a long history in scientific computing and has recently gained popularity in machine learning, particularly in methods that preserve the graph spectrum. This work studies graph coarsening from a different perspective, developing a theory for preserving graph distances and proposing a method to achieve this. The geometric approach is useful when working with a collection of graphs, such as in graph classification and regression. In this study, we consider a graph as an element on a metric space equipped with the Gromov--Wasserstein (GW) distance, and bound the difference between the distance of two graphs and their coarsened versions. Minimizing this difference can be done using the popular weighted kernel $K$-means method, which improves existing spectrum-preserving methods with the proper choice of the kernel. The study includes a set of experiments to support the theory and method, including approximating the GW distance, preserving the graph spectrum, classifying graphs using spectral information, and performing regression using graph convolutional networks. Code is available at https://github.com/ychen-stat-ml/GW-Graph-Coarsening.
Yifan Chen 0004, Rentian Yao
ICML2
2022 Mean-field nonparametric estimation of interacting particle systems
abstract
This paper concerns the nonparametric estimation problem of the distribution-state dependent drift vector field in an interacting $N$-particle system. Observing single-trajectory data for each particle, we derive the mean-field rate of convergence for the maximum likelihood estimator (MLE), which depends on both Gaussian complexity and Rademacher complexity of the function class. In particular, when the function class contains $\alpha$-smooth H{ö}lder functions, our rate of convergence is minimax optimal on the order of $N^{-\frac{\alpha}{d+2\alpha}}$. Combining with a Fourier analytical deconvolution estimator, we derive the consistency of MLE for the external force and interaction kernel in the McKean-Vlasov equation.
Rentian Yao
COLT1