Petr Mokrov

dblp:294/4264 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Optimization for machine learning · 62% Generative modeling · 29% Language models and text generation · 5%
Theoretical computer science
2 papers
Mathematical optimization · 100%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning
optimal transport
3.752024
Optimal Flow Matching: Learning Straight Trajectories in Just One Step · NeurIPS 2024
Energy-Guided Continuous Entropic Barycenter Estimation for General Costs · NeurIPS 2024
Energy-guided Entropic Neural Optimal Transport · ICLR 2024
Mathematical optimization
optimal transport
1.622025
Robust Barycenter Estimation using Semi-Unbalanced Neural Optimal Transport · ICLR 2025
Estimating Barycenters of Distributions with Neural Optimal Transport · ICML 2024
Machine learning › Generative modeling
energy-based model
1.522024
Energy-Guided Continuous Entropic Barycenter Estimation for General Costs · NeurIPS 2024
Energy-guided Entropic Neural Optimal Transport · ICLR 2024
Machine learning › Optimization for machine learning › optimal transport
entropic optimal transport
1.422024
Energy-guided Entropic Neural Optimal Transport · ICLR 2024
Building the Bridge of Schrödinger: A Continuous Entropic Optimal Transport Benchmark · NeurIPS 2023
Mathematical optimization › optimal transport
unbalanced optimal transport
0.912025
Robust Barycenter Estimation using Semi-Unbalanced Neural Optimal Transport · ICLR 2025
Machine learning › Generative modeling
flow matching
0.812024
Optimal Flow Matching: Learning Straight Trajectories in Just One Step · NeurIPS 2024
Machine learning › Optimization for machine learning › optimal transport
neural optimal transport
0.812024
Neural Optimal Transport with General Cost Functionals · ICLR 2024
Mathematical optimization › optimal transport
wasserstein barycenter
0.812024
Estimating Barycenters of Distributions with Neural Optimal Transport · ICML 2024
Natural language and speech › Language models and text generation › evaluation of language models
benchmark construction
0.712023
Building the Bridge of Schrödinger: A Continuous Entropic Optimal Transport Benchmark · NeurIPS 2023
Machine learning › Generative modeling
diffusion model
0.712023
Building the Bridge of Schrödinger: A Continuous Entropic Optimal Transport Benchmark · NeurIPS 2023
Machine learning › Generative modeling › diffusion model
schrödinger bridge
0.712023
Building the Bridge of Schrödinger: A Continuous Entropic Optimal Transport Benchmark · NeurIPS 2023
Machine learning › Optimization for machine learning › optimal transport
JKO scheme
0.512021
Large-Scale Wasserstein Gradient Flows · NeurIPS 2021
Machine learning › Optimization for machine learning › gradient flow
wasserstein gradient flow
0.512021
Large-Scale Wasserstein Gradient Flows · NeurIPS 2021
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.212024
Estimating Barycenters of Distributions with Neural Optimal Transport · ICML 2024
Visual content generation and editing
image-to-image translation
0.212024
Energy-guided Entropic Neural Optimal Transport · ICLR 2024
Visual content generation and editing › image-to-image translation
unpaired image translation
0.212024
Energy-guided Entropic Neural Optimal Transport · ICLR 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian filtering
nonlinear filtering
0.112021
Large-Scale Wasserstein Gradient Flows · NeurIPS 2021
Machine learning › Probabilistic and Bayesian machine learning › sampling
unnormalized density sampling
0.112021
Large-Scale Wasserstein Gradient Flows · NeurIPS 2021

Methods — techniques the papers use, named apart from their topics

neural optimal transport · 2.4energy-based model · 1.5dual formulation · 1.5adversarial learning · 1.5StyleGAN · 1.5semi-unbalanced optimal transport · 0.9min-max optimization · 0.9weak optimal transport · 0.8neural network · 0.8error analysis · 0.8dual reformulation · 0.8convex function parameterization · 0.8
YearPublicationVenuePosition
2025 Robust Barycenter Estimation using Semi-Unbalanced Neural Optimal Transport
abstract
Aggregating data from multiple sources can be formalized as an *Optimal Transport* (OT) barycenter problem, which seeks to compute the average of probability distributions with respect to OT discrepancies. However, in real-world scenarios, the presence of outliers and noise in the data measures can significantly hinder the performance of traditional statistical methods for estimating OT barycenters. To address this issue, we propose a novel scalable approach for estimating the *robust* continuous barycenter, leveraging the dual formulation of the *(semi-)unbalanced* OT problem. To the best of our knowledge, this paper is the first attempt to develop an algorithm for robust barycenters under the continuous distribution setup. Our method is framed as a $\min$-$\max$ optimization problem and is adaptable to *general* cost functions. We rigorously establish the theoretical underpinnings of the proposed method and demonstrate its robustness to outliers and class imbalance through a number of illustrative experiments. Our source code is publicly available at https://github.com/milenagazdieva/U-NOTBarycenters.
Milena Gazdieva, Jaemoo Choi, Alexander Kolesov, Jaewoong Choi, Petr Mokrov, Alexander Korotin
ICLR5
2024 Neural Optimal Transport with General Cost Functionals
abstract
We introduce a novel neural network-based algorithm to compute optimal transport (OT) plans for general cost functionals. In contrast to common Euclidean costs, i.e., $\ell^1$ or $\ell^2$, such functionals provide more flexibility and allow using auxiliary information, such as class labels, to construct the required transport map. Existing methods for general cost functionals are discrete and do not provide an out-of-sample estimation. We address the challenge of designing a continuous OT approach for general cost functionals in high-dimensional spaces, such as images. We construct two example functionals: one to map distributions while preserving the class-wise structure and the other one to preserve the given data pairs. Additionally, we provide the theoretical error analysis for our recovered transport plans. Our implementation is available at \url{https://github.com/machinestein/gnot}
Arip Asadulaev, Alexander Korotin, Vage Egiazarian, Petr Mokrov, Evgeny Burnaev
ICLR4
2024 Energy-guided Entropic Neural Optimal Transport
abstract
Energy-based models (EBMs) are known in the Machine Learning community for decades. Since the seminal works devoted to EBMs dating back to the noughties, there have been a lot of efficient methods which solve the generative modelling problem by means of energy potentials (unnormalized likelihood functions). In contrast, the realm of Optimal Transport (OT) and, in particular, neural OT solvers is much less explored and limited by few recent works (excluding WGAN-based approaches which utilize OT as a loss function and do not model OT maps themselves). In our work, we bridge the gap between EBMs and Entropy-regularized OT. We present a novel methodology which allows utilizing the recent developments and technical improvements of the former in order to enrich the latter. From the theoretical perspective, we prove generalization bounds for our technique. In practice, we validate its applicability in toy 2D and image domains. To showcase the scalability, we empower our method with a pre-trained StyleGAN and apply it to high-res AFHQ $512\times512$ unpaired I2I translation. For simplicity, we choose simple short- and long-run EBMs as a backbone of our Energy-guided Entropic OT approach, leaving the application of more sophisticated EBMs for future research. Our code is available at: https://github.com/PetrMokrov/Energy-guided-Entropic-OT
Petr Mokrov, Alexander Korotin, Alexander Kolesov, Nikita Gushchin, Evgeny Burnaev
ICLR1
2024 Estimating Barycenters of Distributions with Neural Optimal Transport
abstract
Given a collection of probability measures, a practitioner sometimes needs to find an "average" distribution which adequately aggregates reference distributions. A theoretically appealing notion of such an average is the Wasserstein barycenter, which is the primal focus of our work. By building upon the dual formulation of Optimal Transport (OT), we propose a new scalable approach for solving the Wasserstein barycenter problem. Our methodology is based on the recent Neural OT solver: it has bi-level adversarial learning objective and works for general cost functions. These are key advantages of our method since the typical adversarial algorithms leveraging barycenter tasks utilize tri-level optimization and focus mostly on quadratic cost. We also establish theoretical error bounds for our proposed approach and showcase its applicability and effectiveness in illustrative scenarios and image data setups. Our source code is available at https://github.com/justkolesov/NOTBarycenters.
Alexander Kolesov, Petr Mokrov, Igor Udovichenko, Milena Gazdieva, Gudmund Pammer, Evgeny Burnaev, Alexander Korotin
ICML2
2024 Energy-Guided Continuous Entropic Barycenter Estimation for General Costs
abstract
Optimal transport (OT) barycenters are a mathematically grounded way of averaging probability distributions while capturing their geometric properties. In short, the barycenter task is to take the average of a collection of probability distributions w.r.t. given OT discrepancies. We propose a novel algorithm for approximating the continuous Entropic OT (EOT) barycenter for arbitrary OT cost functions. Our approach is built upon the dual reformulation of the EOT problem based on weak OT, which has recently gained the attention of the ML community. Beyond its novelty, our method enjoys several advantageous properties: (i) we establish quality bounds for the recovered solution; (ii) this approach seamlessly interconnects with the Energy-Based Models (EBMs) learning procedure enabling the use of well-tuned algorithms for the problem of interest; (iii) it provides an intuitive optimization scheme avoiding min-max, reinforce and other intricate technical tricks. For validation, we consider several low-dimensional scenarios and image-space setups, including *non-Euclidean* cost functions. Furthermore, we investigate the practical task of learning the barycenter on an image manifold generated by a pretrained generative model, opening up new directions for real-world applications. Our code is available at https://github.com/justkolesov/EnergyGuidedBarycenters.
Alexander Kolesov, Petr Mokrov, Igor Udovichenko, Milena Gazdieva, Gudmund Pammer, Anastasis Kratsios, Evgeny Burnaev, Alexander Korotin
NeurIPS2
2024 Optimal Flow Matching: Learning Straight Trajectories in Just One Step
abstract
Over the several recent years, there has been a boom in development of Flow Matching (FM) methods for generative modeling. One intriguing property pursued by the community is the ability to learn flows with straight trajectories which realize the Optimal Transport (OT) displacements. Straightness is crucial for the fast integration (inference) of the learned flow's paths. Unfortunately, most existing flow straightening methods are based on non-trivial iterative FM procedures which accumulate the error during training or exploit heuristics based on minibatch OT. To address these issues, we develop and theoretically justify the novel Optimal Flow Matching approach which allows recovering the straight OT displacement for the quadratic transport in just one FM step. The main idea of our approach is the employment of vector field for FM which are parameterized by convex functions. The code of our OFM implementation and the conducted experiments is available at https://github.com/Jhomanik/Optimal-Flow-Matching
Nikita Kornilov, Petr Mokrov, Alexander V. Gasnikov, Alexander Korotin
NeurIPS2
2023 Building the Bridge of Schrödinger: A Continuous Entropic Optimal Transport Benchmark
abstract
Over the last several years, there has been significant progress in developing neural solvers for the Schrödinger Bridge (SB) problem and applying them to generative modelling. This new research field is justifiably fruitful as it is interconnected with the practically well-performing diffusion models and theoretically grounded entropic optimal transport (EOT). Still, the area lacks non-trivial tests allowing a researcher to understand how well the methods solve SB or its equivalent continuous EOT problem. We fill this gap and propose a novel way to create pairs of probability distributions for which the ground truth OT solution is known by the construction. Our methodology is generic and works for a wide range of OT formulations, in particular, it covers the EOT which is equivalent to SB (the main interest of our study). This development allows us to create continuous benchmark distributions with the known EOT and SB solutions on high-dimensional spaces such as spaces of images. As an illustration, we use these benchmark pairs to test how well existing neural EOT/SB solvers actually compute the EOT solution. Our code for constructing benchmark pairs under different setups is available at: https://github.com/ngushchin/EntropicOTBenchmark
Nikita Gushchin, Alexander Kolesov, Petr Mokrov, Polina Karpikova, Andrei Spiridonov, Evgeny Burnaev, Alexander Korotin
NeurIPS3
2021 Large-Scale Wasserstein Gradient Flows
abstract
Wasserstein gradient flows provide a powerful means of understanding and solving many diffusion equations. Specifically, Fokker-Planck equations, which model the diffusion of probability measures, can be understood as gradient descent over entropy functionals in Wasserstein space. This equivalence, introduced by Jordan, Kinderlehrer and Otto, inspired the so-called JKO scheme to approximate these diffusion processes via an implicit discretization of the gradient flow in Wasserstein space. Solving the optimization problem associated with each JKO step, however, presents serious computational challenges. We introduce a scalable method to approximate Wasserstein gradient flows, targeted to machine learning applications. Our approach relies on input-convex neural networks (ICNNs) to discretize the JKO steps, which can be optimized by stochastic gradient descent. Contrarily to previous work, our method does not require domain discretization or particle simulation. As a result, we can sample from the measure at each time step of the diffusion and compute its probability density. We demonstrate the performance of our algorithm by computing diffusions following the Fokker-Planck equation and apply it to unnormalized density sampling as well as nonlinear filtering.
Petr Mokrov, Alexander Korotin, Aude Genevay, Justin Solomon 0001, Evgeny Burnaev
NeurIPS1