Ting-Kam Leonard Wong

dblp:180/5695 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
2since 2021 · last 2025
0000-0001-5254-7305ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Theory of computation · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2025 Bregman-Wasserstein Divergence: Geometry and Applications
abstract
The Bregman-Wasserstein divergence is the optimal transport cost when the underlying cost function is given by a Bregman divergence, and arises naturally in fields such as statistics and machine learning. We establish fundamental properties of the Bregman-Wasserstein divergence and propose a novel generalized transport geometry that promotes the Bregman geometry to the space of probability distributions. We provide a probabilistic interpretation involving exponential families and define generalized displacement interpolations compatible with the Bregman geometry. These interpolations are used to derive a generalized Pythagorean inequality, which is of independent interest. Furthermore, we construct a generalized dualistic geometry that lifts the differential geometry of the Bregman divergence to an infinite-dimensional statistical manifold. On the computational side, we demonstrate how Bregman-Wasserstein optimal transport maps can be estimated using neural approaches, establish the well-posedness of Bregman-Wasserstein barycenters, and relate them to Bayesian learning. Finally, we relate the Bregman-Wasserstein divergence to rate distortion theory, and introduce the Bregman-Wasserstein JKO scheme for discretizing Riemannian Wasserstein gradient flows.
Amanjit Singh Kainth, Cale Rankin, Ting-Kam Leonard Wong
IEEE Trans. Inf. Theory3
2022 Tsallis and Rényi Deformations Linked via a New λ-Duality
abstract
Tsallis and Rényi entropies, which are monotone transformations of each other, are deformations of the celebrated Shannon entropy. Maximization of these deformed entropies, under suitable constraints, leads to the$q$-exponential family which has applications in non-extensive statistical physics, information theory and statistics. In previous information-geometric studies, the$q$-exponential family was analyzed using classical convex duality and Bregman divergence. In this paper, we show that a generalized$\lambda $-duality, where$\lambda = 1 - q$is to be interpreted as the constant information-geometric curvature, leads to a generalized exponential family which is essentially equivalent to the$q$-exponential family and has deep connections with Rényi entropy and optimal transport. Using this generalized convex duality and its associated logarithmic divergence, we show that our$\lambda $-exponential family satisfies properties that parallel and generalize those of the exponential family. Under our framework, the Rényi entropy and divergence arise naturally, and we give a new proof of the Tsallis/Rényi entropy maximizing property of the$q$-exponential family. We also introduce a$\lambda $-mixture family which may be regarded as the dual of the$\lambda $-exponential family, and connect it with other mixture-type families. Finally, we discuss a duality between the$\lambda $-exponential family and the$\lambda $-logarithmic divergence, and study its statistical consequences.
Ting-Kam Leonard Wong, Jun Zhang 0009
IEEE Trans. Inf. Theory1
2020 Scalable Gradients for Stochastic Differential Equations
abstract
The adjoint sensitivity method scalably computes gradients of solutions to ordinary differential equations. We generalize this method to stochastic differential equations, allowing time-efficient and constant-memory computation of gradients with high-order adaptive solvers. Specifically, we derive a stochastic differentialequation whose solution is the gradient, a memory-efficient algorithm for cachingnoise, and conditions under which numerical solutions converge. In addition, we combine our method with gradient-based stochastic variational inference for latent stochastic differential equations. We use our method to fit stochastic dynamics defined by neural networks, achieving competitive performance ona 50-dimensional motion capture dataset.
Xuechen Li 0005, Ting-Kam Leonard Wong, Ricky T. Q. Chen, David Duvenaud
AISTATS2