Johannes Hertrich

dblp:243/3816 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0003-4433-8604ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Generative modeling · 52% Probabilistic and Bayesian machine learning · 18% Optimization for machine learning · 17%
Theoretical computer science
2 papers
Algorithms and data structures · 62% Mathematical optimization · 38%

Topics — the 16 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › sampling
posterior sampling
1.622025
Importance Corrected Neural JKO Sampling · ICML 2025
Posterior Sampling Based on Gradient Flows of the MMD with Negative Distance Kernel · ICLR 2024
Machine learning › Optimization for machine learning › gradient flow
wasserstein gradient flow
1.422024
Posterior Sampling Based on Gradient Flows of the MMD with Negative Distance Kernel · ICLR 2024
Neural Wasserstein Gradient Flows for Discrepancies with Riesz Kernels · ICML 2023
Machine learning › Generative modeling
flow matching
0.912025
On the Relation between Rectified Flows and Optimal Transport · NeurIPS 2025
Machine learning › Kernel, tree and ensemble methods
kernel methods
0.912025
Fast Summation of Radial Kernels via QMC Slicing · ICLR 2025
Machine learning › Generative modeling
normalizing flow
0.912025
Importance Corrected Neural JKO Sampling · ICML 2025
Machine learning › Generative modeling › diffusion model
rectified flow
0.912025
On the Relation between Rectified Flows and Optimal Transport · NeurIPS 2025
Algorithms and data structures
kernel methods
0.912025
Fast Summation of Radial Kernels via QMC Slicing · ICLR 2025
Machine learning › Generative modeling
conditional generative model
0.812024
Posterior Sampling Based on Gradient Flows of the MMD with Negative Distance Kernel · ICLR 2024
Machine learning › Generative modeling
image generation
0.812024
Generative Sliced MMD Flows with Riesz Kernels · ICLR 2024
Machine learning › Generative modeling
inverse problem
0.812024
Manifold Learning by Mixture Models of VAEs for Inverse Problems · J. Mach. Learn. Res. 2024
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
manifold learning
0.812024
Manifold Learning by Mixture Models of VAEs for Inverse Problems · J. Mach. Learn. Res. 2024
Machine learning › Generative modeling
variational autoencoder
0.812024
Manifold Learning by Mixture Models of VAEs for Inverse Problems · J. Mach. Learn. Res. 2024
Machine learning › Optimization for machine learning
gradient flow
0.712023
Neural Wasserstein Gradient Flows for Discrepancies with Riesz Kernels · ICML 2023
Mathematical optimization
continuous optimization
0.312025
On the Relation between Rectified Flows and Optimal Transport · NeurIPS 2025
Mathematical optimization
optimal transport
0.312025
On the Relation between Rectified Flows and Optimal Transport · NeurIPS 2025
Image and video processing › image restoration
inverse problem
0.212024
Posterior Sampling Based on Gradient Flows of the MMD with Negative Distance Kernel · ICLR 2024

Methods — techniques the papers use, named apart from their topics

wasserstein gradient flow · 2.4slicing · 1.7random fourier features · 1.7quasi-monte carlo · 1.7flow matching · 1.7neural network · 1.4rejection-resampling · 0.9rectified flow · 0.9optimal transport · 0.9importance weights · 0.9continuous normalizing flow · 0.9JKO scheme · 0.9negative distance kernel · 0.8maximum mean discrepancy · 0.8
YearPublicationVenuePosition
2025 Fast Summation of Radial Kernels via QMC Slicing
abstract
The fast computation of large kernel sums is a challenging task, which arises as a subproblem in any kernel method. We approach the problem by slicing, which relies on random projections to one-dimensional subspaces and fast Fourier summation. We prove bounds for the slicing error and propose a quasi-Monte Carlo (QMC) approach for selecting the projections based on spherical quadrature rules. Numerical examples demonstrate that our QMC-slicing approach significantly outperforms existing methods like (QMC-)random Fourier features, orthogonal Fourier features or non-QMC slicing on standard test datasets.
Johannes Hertrich, Tim Jahn, Michael Quellmalz
ICLR1
2025 Importance Corrected Neural JKO Sampling
abstract
In order to sample from an unnormalized probability density function, we propose to combine continuous normalizing flows (CNFs) with rejection-resampling steps based on importance weights. We relate the iterative training of CNFs with regularized velocity fields to a JKO scheme and prove convergence of the involved velocity fields to the velocity field of the Wasserstein gradient flow (WGF). The alternation of local flow steps and non-local rejection-resampling steps allows to overcome local minima or slow convergence of the WGF for multimodal distributions. Since the proposal of the rejection step is generated by the model itself, they do not suffer from common drawbacks of classical rejection schemes. The arising model can be trained iteratively, reduces the reverse Kullback-Leibler (KL) loss function in each step, allows to generate iid samples and moreover allows for evaluations of the generated underlying density. Numerical examples show that our method yields accurate results on various test distributions including high-dimensional multimodal targets and outperforms the state of the art in almost all cases significantly.
Johannes Hertrich, Robert Gruhlke
ICML1
2025 On the Relation between Rectified Flows and Optimal Transport
abstract
This paper investigates the connections between rectified flows, flow matching, and optimal transport. Flow matching is a recent approach to learning generative models by estimating velocity fields that guide transformations from a source to a target distribution. Rectified flow matching aims to straighten the learned transport paths, yielding more direct flows between distributions. Our first contribution is a set of invariance properties of rectified flows and explicit velocity fields. In addition, we also provide explicit constructions and analysis in the Gaussian (not necessarily independent) and Gaussian mixture settings and study the relation to optimal transport. Our second contribution addresses recent claims suggesting that rectified flows, when constrained such that the learned velocity field is a gradient, can yield (asymptotically) solutions to optimal transport problems. We study the existence of solutions for this problem and demonstrate that they only relate to optimal transport under assumptions that are significantly stronger than those previously acknowledged. In particular, we present several counterexamples that invalidate earlier equivalence results in the literature, and we argue that enforcing a gradient constraint on rectified flows is, in general, not a reliable method for computing optimal transport maps.
Johannes Hertrich, Antonin Chambolle, Julie Delon
NeurIPS1
2024 Posterior Sampling Based on Gradient Flows of the MMD with Negative Distance Kernel
abstract
We propose conditional flows of the maximum mean discrepancy (MMD) with the negative distance kernel for posterior sampling and conditional generative modelling. This MMD, which is also known as energy distance, has several advantageous properties like efficient computation via slicing and sorting. We approximate the joint distribution of the ground truth and the observations using discrete Wasserstein gradient flows and establish an error bound for the posterior distributions. Further, we prove that our particle flow is indeed a Wasserstein gradient flow of an appropriate functional. The power of our method is demonstrated by numerical examples including conditional image generation and inverse problems like superresolution, inpainting and computed tomography in low-dose and limited-angle settings.
Paul Hagemann, Johannes Hertrich, Fabian Altekrüger, Robert Beinert, Jannis Chemseddine, Gabriele Steidl
ICLR2
2024 Generative Sliced MMD Flows with Riesz Kernels
abstract
Maximum mean discrepancy (MMD) flows suffer from high computational costs in large scale computations. In this paper, we show that MMD flows with Riesz kernels $K(x,y) = - \|x-y\|^r$, $r \in (0,2)$ have exceptional properties which allow their efficient computation. We prove that the MMD of Riesz kernels, which is also known as energy distance, coincides with the MMD of their sliced version. As a consequence, the computation of gradients of MMDs can be performed in the one-dimensional setting. Here, for $r=1$, a simple sorting algorithm can be applied to reduce the complexity from $O(MN+N^2)$ to $O((M+N)\log(M+N))$ for two measures with $M$ and $N$ support points. As another interesting follow-up result, the MMD of compactly supported measures can be estimated from above and below by the Wasserstein-1 distance. For the implementations we approximate the gradient of the sliced MMD by using only a finite number $P$ of slices. We show that the resulting error has complexity \smash{$O(\sqrt{d/P})$}, where $d$ is the data dimension. These results enable us to train generative models by approximating MMD gradient flows by neural networks even for image applications. We demonstrate the efficiency of our model by image generation on MNIST, FashionMNIST and CIFAR10.
Johannes Hertrich, Christian Wald, Fabian Altekrüger, Paul Hagemann
ICLR1
2024 Manifold Learning by Mixture Models of VAEs for Inverse Problems
abstract
Representing a manifold of very high-dimensional data with generative models has been shown to be computationally efficient in practice. However, this requires that the data manifold admits a global parameterization. In order to represent manifolds of arbitrary topology, we propose to learn a mixture model of variational autoencoders. Here, every encoder-decoder pair represents one chart of a manifold. We propose a loss function for maximum likelihood estimation of the model weights and choose an architecture that provides us the analytical expression of the charts and of their inverses. Once the manifold is learned, we use it for solving inverse problems by minimizing a data fidelity term restricted to the learned manifold. To solve the arising minimization problem we propose a Riemannian gradient descent algorithm on the learned manifold. We demonstrate the performance of our method for low-dimensional toy examples as well as for deblurring and electrical impedance tomography on certain image manifolds.
Giovanni S. Alberti, Johannes Hertrich, Matteo Santacesaria, Silvia Sciutto
J. Mach. Learn. Res.2
2023 Neural Wasserstein Gradient Flows for Discrepancies with Riesz Kernels
abstract
Wasserstein gradient flows of maximum mean discrepancy (MMD) functionals with non-smooth Riesz kernels show a rich structure as singular measures can become absolutely continuous ones and conversely. In this paper we contribute to the understanding of such flows. We propose to approximate the backward scheme of Jordan, Kinderlehrer and Otto for computing such Wasserstein gradient flows as well as a forward scheme for so-called Wasserstein steepest descent flows by neural networks (NNs). Since we cannot restrict ourselves to absolutely continuous measures, we have to deal with transport plans and velocity plans instead of usual transport maps and velocity fields. Indeed, we approximate the disintegration of both plans by generative NNs which are learned with respect to appropriate loss functions. In order to evaluate the quality of both neural schemes, we benchmark them on the interaction energy. Here we provide analytic formulas for Wasserstein schemes starting at a Dirac measure and show their convergence as the time step size tends to zero. Finally, we illustrate our neural MMD flows by numerical examples.
Fabian Altekrüger, Johannes Hertrich, Gabriele Steidl
ICML2
2023 WPPNets and WPPFlows: The Power of Wasserstein Patch Priors for Superresolution
abstract
Abstract. Exploiting image patches instead of whole images has proved to be a powerful approach to tackling various problems in image processing. Recently, Wasserstein patch priors (WPPs), which are based on the comparison of the patch distributions of the unknown image and a reference image, were successfully used as data-driven regularizers in the variational formulation of superresolution. However, for each input image, this approach requires the solution of a nonconvex minimization problem which is computationally costly. In this paper, we propose to learn two kinds of neural networks in an unsupervised way based on WPP loss functions. First, we show how convolutional neural networks (CNNs) can be incorporated. Once the network, called WPPNet, is learned, it can be very efficiently applied to any input image. Second, we incorporate conditional normalizing flows to provide a tool for uncertainty quantification. Numerical examples demonstrate the very good performance of WPPNets for superresolution in various image classes, even if the forward operator is known only approximately.
Fabian Altekrüger, Johannes Hertrich
SIAM J. Imaging Sci.2