Amirhossein Taghvaei

dblp:158/4926 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 4 since 2021Computer networks · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Optimization for machine learning · 56% Probabilistic and Bayesian machine learning · 26% Deep learning architectures and training · 9%
Theoretical computer science
2 papers
Mathematical optimization · 64% Information theory · 36%

Topics — the 19 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning
optimal transport
1.222024
Nonlinear Filtering with Brenier Optimal Transport Maps · ICML 2024
Optimal transport mapping via input convex neural networks · ICML 2020
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian filtering
nonlinear filtering
0.812024
Nonlinear Filtering with Brenier Optimal Transport Maps · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › sequential monte carlo
particle filtering
0.812024
Nonlinear Filtering with Brenier Optimal Transport Maps · ICML 2024
Machine learning › Optimization for machine learning
stochastic gradient descent
0.712023
Data-driven Optimal Filtering for Linear Systems with Unknown Noise Covariances · NeurIPS 2023
Information theory › estimation theory › bayesian estimation
kalman filtering
0.712023
Data-driven Optimal Filtering for Linear Systems with Unknown Noise Covariances · NeurIPS 2023
Mathematical optimization › control theory
optimal filtering
0.712023
Data-driven Optimal Filtering for Linear Systems with Unknown Noise Covariances · NeurIPS 2023
Machine learning › Optimization for machine learning › optimal transport
JKO scheme
0.612022
Variational Wasserstein gradient flow · ICML 2022
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.612022
Variational Wasserstein gradient flow · ICML 2022
Machine learning › Optimization for machine learning › gradient flow
wasserstein gradient flow
0.612022
Variational Wasserstein gradient flow · ICML 2022
Machine learning › Optimization for machine learning › optimal transport
wasserstein barycenter
0.512021
Scalable Computations of Wasserstein Barycenter via Input Convex Neural Networks · ICML 2021
Mathematical optimization › continuous optimization
convex optimization
0.512021
Scalable Computations of Wasserstein Barycenter via Input Convex Neural Networks · ICML 2021
Machine learning › Optimization for machine learning
minimax optimization
0.412020
Optimal transport mapping via input convex neural networks · ICML 2020
Machine learning › Reinforcement learning › multi-agent reinforcement learning
mean field control
0.412019
Accelerated Flow for Probability Distributions · ICML 2019
Machine learning › Optimization for machine learning › optimization landscape
critical point analysis
0.312017
How regularization affects the critical points in linear networks · NIPS 2017
Machine learning › Deep learning architectures and training › feedforward neural network
deep linear networks
0.312017
How regularization affects the critical points in linear networks · NIPS 2017
Machine learning › Deep learning architectures and training
loss landscape
0.312017
How regularization affects the critical points in linear networks · NIPS 2017
Robotics › Robot navigation and mapping
state estimation
0.212024
Nonlinear Filtering with Brenier Optimal Transport Maps · ICML 2024
Machine learning › Optimization for machine learning
convergence analysis
0.212023
Data-driven Optimal Filtering for Linear Systems with Unknown Noise Covariances · NeurIPS 2023
Machine learning › Deep learning architectures and training › feedforward neural network
convex neural network
0.212022
Variational Wasserstein gradient flow · ICML 2022

Methods — techniques the papers use, named apart from their topics

input convex neural network · 2.0high-dimensional statistics · 1.3bias-variance bounds · 1.3optimal transport · 1.0stochastic optimization · 0.8neural network approximation · 0.8JKO scheme · 0.6minimax optimization · 0.4kantorovich potential · 0.4hamiltonian monte carlo · 0.4
YearPublicationVenuePosition
2026 A Lasso-Alternative to Dijkstra's Algorithm for Identifying Short Paths in Networks
abstract
ABSTRACT We revisit the problem of finding the shortest path between two selected vertices of a graph and formulate this as an ‐regularized regression—Least Absolute Shrinkage and Selection Operator (lasso). We draw connections between a numerical implementation of this lasso formulation, using the so‐called LARS algorithm, and a more established algorithm known as the bi‐directional Dijkstra. Appealing features of our formulation include the applicability of the Alternating Direction of Multiplier Method (ADMM) to the problem to identify short paths, and a relatively efficient update to topological changes.
Anqi Dong, Amirhossein Taghvaei, Tryphon T. Georgiou
Networks2
2024 Nonlinear Filtering with Brenier Optimal Transport Maps
abstract
This paper is concerned with the problem of nonlinear filtering, i.e., computing the conditional distribution of the state of a stochastic dynamical system given a history of noisy partial observations. Conventional sequential importance resampling (SIR) particle filters suffer from fundamental limitations, in scenarios involving degenerate likelihoods or high-dimensional states, due to the weight degeneracy issue. In this paper, we explore an alternative method, which is based on estimating the Brenier optimal transport (OT) map from the current prior distribution of the state to the posterior distribution at the next time step. Unlike SIR particle filters, the OT formulation does not require the analytical form of the likelihood. Moreover, it allows us to harness the approximation power of neural networks to model complex and multi-modal distributions and employ stochastic optimization algorithms to enhance scalability. Extensive numerical experiments are presented that compare the OT method to the SIR particle filter and the ensemble Kalman filter, evaluating the performance in terms of sample efficiency, high-dimensional scalability, and the ability to capture complex and multi-modal distributions.
Niyizhen Jin, Bamdad Hosseini, Amirhossein Taghvaei
ICML4
2023 Data-driven Optimal Filtering for Linear Systems with Unknown Noise Covariances
abstract
This paper examines learning the optimal filtering policy, known as the Kalman gain, for a linear system with unknown noise covariance matrices using noisy output data. The learning problem is formulated as a stochastic policy optimiza- tion problem, aiming to minimize the output prediction error. This formulation provides a direct bridge between data-driven optimal control and, its dual, op- timal filtering. Our contributions are twofold. Firstly, we conduct a thorough convergence analysis of the stochastic gradient descent algorithm, adopted for the filtering problem, accounting for biased gradients and stability constraints. Secondly, we carefully leverage a combination of tools from linear system theory and high-dimensional statistics to derive bias-variance error bounds that scale logarithmically with problem dimension, and, in contrast to subspace methods, the length of output trajectories only affects the bias term.
Shahriar Talebi, Amirhossein Taghvaei, Mehran Mesbahi
NeurIPS2
2022 Variational Wasserstein gradient flow
abstract
Wasserstein gradient flow has emerged as a promising approach to solve optimization problems over the space of probability distributions. A recent trend is to use the well-known JKO scheme in combination with input convex neural networks to numerically implement the proximal step. The most challenging step, in this setup, is to evaluate functions involving density explicitly, such as entropy, in terms of samples. This paper builds on the recent works with a slight but crucial difference: we propose to utilize a variational formulation of the objective function formulated as maximization over a parametric class of functions. Theoretically, the proposed variational formulation allows the construction of gradient flows directly for empirical distributions with a well-defined and meaningful objective function. Computationally, this approach replaces the computationally expensive step in existing methods, to handle objective functions involving density, with inner loop updates that only require a small batch of samples and scale well with the dimension. The performance and scalability of the proposed method are illustrated with the aid of several numerical experiments involving high-dimensional synthetic and real datasets.
Jiaojiao Fan, Qinsheng Zhang, Amirhossein Taghvaei
ICML3
2021 Scalable Computations of Wasserstein Barycenter via Input Convex Neural Networks
Jiaojiao Fan, Amirhossein Taghvaei
ICML3
2020 Optimal transport mapping via input convex neural networks
abstract
In this paper, we present a novel and principled approach to learn the optimal transport between two distributions, from samples. Guided by the optimal transport theory, we learn the optimal Kantorovich potential which induces the optimal transport map. This involves learning two convex functions, by solving a novel minimax optimization. Building upon recent advances in the field of input convex neural networks, we propose a new framework to estimate the optimal transport mapping as the gradient of a convex function that is trained via minimax optimization. Numerical experiments confirm the accuracy of the learned transport map. Our approach can be readily used to train a deep generative model. When trained between a simple distribution in the latent space and a target distribution, the learned optimal transport map acts as a deep generative model. Although scaling this to a large dataset is challenging, we demonstrate two important strengths over standard adversarial training: robustness and discontinuity. As we seek the optimal transport, the learned generative model provides the same mapping regardless of how we initialize the neural networks. Further, a gradient of a neural network can easily represent discontinuous mappings, unlike standard neural networks that are constrained to be continuous. This allows the learned transport map to match any target distribution with many discontinuous supports and achieve sharp boundaries.
Ashok Vardhan Makkuva, Amirhossein Taghvaei, Sewoong Oh, Jason D. Lee
ICML2
2019 Accelerated Flow for Probability Distributions
abstract
This paper presents a methodology and numerical algorithms for constructing accelerated gradient flows on the space of probability distributions. In particular, we extend the recent variational formulation of accelerated methods in (Wibisono et al., 2016) from vector valued variables to probability distributions. The variational problem is modeled as a mean-field optimal control problem. A quantitative estimate on the asymptotic convergence rate is provided based on a Lyapunov function construction, when the objective functional is displacement convex. An important special case is considered where the objective functional is the relative entropy. For this case, two numerical approximations are presented to implement the Hamilton’s equations as a system of N interacting particles. The algorithm is numerically illustrated and compared with the MCMC and Hamiltonian MCMC algorithms.
Amirhossein Taghvaei, Prashant G. Mehta
ICML1
2017 How regularization affects the critical points in linear networks
abstract
This paper is concerned with the problem of representing and learning a linear transformation using a linear neural network. In recent years, there is a growing interest in the study of such networks, in part due to the successes of deep learning. The main question of this body of research (and also of our paper) is related to the existence and optimality properties of the critical points of the mean-squared loss function. An additional primary concern of our paper pertains to the robustness of these critical points in the face of (a small amount of) regularization. An optimal control model is introduced for this purpose and a learning algorithm (backprop with weight decay) derived for the same using the Hamilton's formulation of optimal control. The formulation is used to provide a complete characterization of the critical points in terms of the solutions of a nonlinear matrix-valued equation, referred to as the characteristic equation. Analytical and numerical tools from bifurcation theory are used to compute the critical points via the solutions of the characteristic equation.
Amirhossein Taghvaei, Jin-Won Kim, Prashant G. Mehta
NIPS1