Laurent Dinh

dblp:131/6819 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
12 papers
Generative modeling · 56% Representation and self-supervised learning · 13% Learning theory · 7%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 26 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
normalizing flow
2.042025
STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis · NeurIPS 2025
VideoFlow: A Conditional Flow-Based Model for Stochastic Video Generation · ICLR 2020
Invertible Convolutional Flow · NeurIPS 2019
Machine learning › Generative modeling › diffusion model › text-to-image generation
high-resolution image synthesis
0.912025
STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis · NeurIPS 2025
Machine learning › Generative modeling
image generation
0.912025
STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis · NeurIPS 2025
Machine learning › Generative modeling › normalizing flow
latent space normalizing flow
0.912025
STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis · NeurIPS 2025
Machine learning › Generative modeling
diffusion model
0.812024
Generative Modeling with Phase Stochastic Bridge · ICLR 2024
Machine learning › Representation and self-supervised learning › representation learning › embedding learning › joint embedding
joint embedding architecture
0.812024
LiDAR: Sensing Linear Probing Performance in Joint Embedding SSL Architectures · ICLR 2024
Machine learning › Representation and self-supervised learning › representation analysis
representation evaluation
0.812024
LiDAR: Sensing Linear Probing Performance in Joint Embedding SSL Architectures · ICLR 2024
Robotics › Motion planning and robot control
stochastic optimal control
0.812024
Generative Modeling with Phase Stochastic Bridge · ICLR 2024
Computer vision › 3D vision › 3d generation
3d scene generation
0.612022
GAUDI: A Neural Architect for Immersive 3D Scene Generation · NeurIPS 2022
Machine learning › Generative modeling
autoregressive model
0.412019
Discrete Flows: Invertible Generative Models of Discrete Data · NeurIPS 2019
Machine learning › Generative modeling › normalizing flow
discrete flow model
0.412019
Discrete Flows: Invertible Generative Models of Discrete Data · NeurIPS 2019
Machine learning › Generative modeling › generative model
discrete generative model
0.412019
Discrete Flows: Invertible Generative Models of Discrete Data · NeurIPS 2019
Machine learning › Generative modeling › normalizing flow
invertible convolution
0.412019
Invertible Convolutional Flow · NeurIPS 2019
Machine learning › Reinforcement learning
model-based reinforcement learning
0.312018
Learning Awareness Models · ICLR (Poster) 2018
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
density estimation
0.312017
Density estimation using Real NVP · ICLR (Poster) 2017
Machine learning › Learning theory › generalization › stability and generalization
flatness and generalization
0.312017
Sharp Minima Can Generalize For Deep Nets · ICML 2017
Machine learning › Learning theory
generalization bounds
0.312017
Sharp Minima Can Generalize For Deep Nets · ICML 2017
Machine learning › Learning theory › generalization
generalization in deep learning
0.312017
Sharp Minima Can Generalize For Deep Nets · ICML 2017
Machine learning › Deep learning architectures and training › loss landscape
loss landscape geometry
0.312017
Sharp Minima Can Generalize For Deep Nets · ICML 2017
Machine learning › Deep learning architectures and training
transformer
0.312025
STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis · NeurIPS 2025
Machine learning › Generative modeling
variational autoencoder
0.212015
A Recurrent Latent Variable Model for Sequential Data · NIPS 2015
Machine learning › Efficient and distributed learning
model compression
0.212013
Predicting Parameters in Deep Learning · NIPS 2013
Machine learning › Efficient and distributed learning
parameter prediction
0.212013
Predicting Parameters in Deep Learning · NIPS 2013
Visual content generation and editing
video generation
0.112020
VideoFlow: A Conditional Flow-Based Model for Stochastic Video Generation · ICLR 2020
Natural language and speech › Language models and text generation › language modeling
character-level language modeling
0.112019
Discrete Flows: Invertible Generative Models of Discrete Data · NeurIPS 2019
Machine learning › Representation and self-supervised learning › representation learning › latent representation learning › state representation learning
predictive state representation
0.112018
Learning Awareness Models · ICLR (Poster) 2018

Methods — techniques the papers use, named apart from their topics

latent space modeling · 0.9guidance algorithm · 0.9autoregressive transformer · 0.9conditional flow · 0.9stochastic differential equation · 0.8self-supervised learning · 0.8neural network · 0.8linear discriminant analysis · 0.8radiance field · 0.6latent representation learning · 0.6
YearPublicationVenuePosition
2025 STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis
abstract
We present STARFlow, a scalable generative model based on normalizing flows that achieves strong performance on high-resolution image synthesis. STARFlow's main building block is Transformer Autoregressive Flow (TARFlow), which combines normalizing flows with Autoregressive Transformer architectures and has recently achieved impressive results in image modeling. In this work, we first establish the theoretical universality of TARFlow for modeling continuous distributions. Building on this foundation, we introduce a set of architectural and algorithmic innovations that significantly enhance the scalability: (1) a deep-shallow design where a deep Transformer block captures most of the model’s capacity, followed by a few shallow Transformer blocks that are computationally cheap yet contribute non-negligibly, (2) learning in the latent space of pretrained autoencoders, which proves far more effective than modeling pixels directly, and (3) a novel guidance algorithm that substantially improves sample quality. Crucially, our model remains a single, end-to-end normalizing flow, allowing exact maximum likelihood training in continuous space without discretization. STARFlow achieves competitive results in both class- and text-conditional image generation, with sample quality approaching that of state-of-the-art diffusion models. To our knowledge, this is the **first** successful demonstration of normalizing flows at this scale and resolution. Code and weights available at https://github.com/apple/ml-starflow.
Jiatao Gu, Tianrong Chen, David Berthelot, Huangjie Zheng, Ruixiang Zhang, Laurent Dinh, Miguel Ángel Bautista 0001, Joshua M. Susskind, Shuangfei Zhai
NeurIPS7
2024 Generative Modeling with Phase Stochastic Bridge
abstract
Diffusion models (DMs) represent state-of-the-art generative models for continuous inputs. DMs work by constructing a Stochastic Differential Equation (SDE) in the input space (ie, position space), and using a neural network to reverse it. In this work, we introduce a novel generative modeling framework grounded in \textbf{phase space dynamics}, where a phase space is defined as {an augmented space encompassing both position and velocity.} Leveraging insights from Stochastic Optimal Control, we construct a path measure in the phase space that enables efficient sampling. {In contrast to DMs, our framework demonstrates the capability to generate realistic data points at an early stage of dynamics propagation.} This early prediction sets the stage for efficient data generation by leveraging additional velocity information along the trajectory. On standard image generation benchmarks, our model yields favorable performance over baselines in the regime of small Number of Function Evaluations (NFEs). Furthermore, our approach rivals the performance of diffusion models equipped with efficient sampling techniques, underscoring its potential as a new tool generative modeling.
Tianrong Chen, Jiatao Gu, Laurent Dinh, Evangelos A. Theodorou, Joshua M. Susskind, Shuangfei Zhai
ICLR3
2024 LiDAR: Sensing Linear Probing Performance in Joint Embedding SSL Architectures
abstract
Joint embedding (JE) architectures have emerged as a promising avenue for ac- quiring transferable data representations. A key obstacle to using JE methods, however, is the inherent challenge of evaluating learned representations without access to a downstream task, and an annotated dataset. Without efficient and re- liable evaluation, it is difficult to iterate on architectural and training choices for JE methods. In this paper, we introduce LiDAR (Linear Discriminant Analysis Rank), a metric designed to measure the quality of representations within JE archi- tectures. Our metric addresses several shortcomings of recent approaches based on feature covariance rank by discriminating between informative and uninforma- tive features. In essence, LiDAR quantifies the rank of the Linear Discriminant Analysis (LDA) matrix associated with the surrogate SSL task—a measure that intuitively captures the information content as it pertains to solving the SSL task. We empirically demonstrate that LiDAR significantly surpasses naive rank based approaches in its predictive power of optimal hyperparameters. Our proposed cri- terion presents a more robust and intuitive means of assessing the quality of rep- resentations within JE architectures, which we hope facilitates broader adoption of these powerful techniques in various domains.
Vimal Thilak, Chen Huang 0001, Omid Saremi, Laurent Dinh, Hanlin Goh, Preetum Nakkiran, Joshua M. Susskind, Etai Littwin
ICLR4
2022 GAUDI: A Neural Architect for Immersive 3D Scene Generation
abstract
We introduce GAUDI, a generative model capable of capturing the distribution of complex and realistic 3D scenes that can be rendered immersively from a moving camera. We tackle this challenging problem with a scalable yet powerful approach, where we first optimize a latent representation that disentangles radiance fields and camera poses. This latent representation is then used to learn a generative model that enables both unconditional and conditional generation of 3D scenes. Our model generalizes previous works that focus on single objects by removing the assumption that the camera pose distribution can be shared across samples. We show that GAUDI obtains state-of-the-art performance in the unconditional generative setting across multiple datasets and allows for conditional generation of 3D scenes given conditioning variables like sparse image observations or text that describes the scene.
Miguel Ángel Bautista 0001, Pengsheng Guo, Samira Abnar, Walter Talbott, Alexander Toshev, Zhuoyuan Chen, Laurent Dinh, Shuangfei Zhai, Hanlin Goh, Daniel Ulbricht, Afshin Dehghan, Joshua M. Susskind
NeurIPS7
2020 VideoFlow: A Conditional Flow-Based Model for Stochastic Video Generation
Manoj Kumar 0019, Mohammad Babaeizadeh, Dumitru Erhan, Chelsea Finn, Sergey Levine, Laurent Dinh, Durk Kingma
ICLR6
2019 Invertible Convolutional Flow
abstract
Normalizing flows can be used to construct high quality generative probabilistic models, but training and sample generation require repeated evaluation of Jacobian determinants and function inverses. To make such computations feasible, current approaches employ highly constrained architectures that produce diagonal, triangular, or low rank Jacobian matrices. As an alternative, we investigate a set of novel normalizing flows based on the circular and symmetric convolutions. We show that these transforms admit efficient Jacobian determinant computation and inverse mapping (deconvolution) in O(N log N) time. Additionally, element-wise multiplication, widely used in normalizing flow architectures, can be combined with these transforms to increase modeling flexibility. We further propose an analytic approach to designing nonlinear elementwise bijectors that induce special properties in the intermediate layers, by implicitly introducing specific regularizers in the loss. We show that these transforms allow more effective normalizing flow models to be developed for generative image models.
Mahdi Karami, Dale Schuurmans, Jascha Sohl-Dickstein, Laurent Dinh, Daniel Duckworth
NeurIPS4
2019 Discrete Flows: Invertible Generative Models of Discrete Data
abstract
While normalizing flows have led to significant advances in modeling high-dimensional continuous distributions, their applicability to discrete distributions remains unknown. In this paper, we show that flows can in fact be extended to discrete events---and under a simple change-of-variables formula not requiring log-determinant-Jacobian computations. Discrete flows have numerous applications. We consider two flow architectures: discrete autoregressive flows that enable bidirectionality, allowing, for example, tokens in text to depend on both left-to-right and right-to-left contexts in an exact language model; and discrete bipartite flows that enable efficient non-autoregressive generation as in RealNVP. Empirically, we find that discrete autoregressive flows outperform autoregressive baselines on synthetic discrete distributions, an addition task, and Potts models; and bipartite flows can obtain competitive performance with autoregressive baselines on character-level language modeling for Penn Tree Bank and text8.
Dustin Tran, Keyon Vafa, Kumar Krishna Agrawal, Laurent Dinh, Ben Poole
NeurIPS4
2018 Learning Awareness Models
Brandon Amos, Laurent Dinh, Serkan Cabi, Thomas Rothörl, Sergio Gomez Colmenarejo, Alistair Muldal, Tom Erez, Yuval Tassa, Nando de Freitas, Misha Denil
ICLR (Poster)2
2017 Density estimation using Real NVP
Laurent Dinh, Jascha Sohl-Dickstein, Samy Bengio
ICLR (Poster)1
2017 Sharp Minima Can Generalize For Deep Nets
abstract
Despite their overwhelming capacity to overfit, deep learning architectures tend to generalize relatively well to unseen data, allowing them to be deployed in practice. However, explaining why this is the case is still an open area of research. One standing hypothesis that is gaining popularity, e.g.\ Hochreiter \& Schmidhuber (1997); Keskar et al.\ (2017), is that the flatness of minima of the loss function found by stochastic gradient based methods results in good generalization. This paper argues that most notions of flatness are problematic for deep models and can not be directly applied to explain generalization. Specifically, when focusing on deep networks with rectifier units, we can exploit the particular geometry of parameter space induced by the inherent symmetries that these architectures exhibit to build equivalent models corresponding to arbitrarily sharper minima. Or, depending on the definition of flatness, it is the same for any given minimum. Furthermore, if we allow to reparametrize a function, the geometry of its parameters can change drastically without affecting its generalization properties.
Laurent Dinh, Razvan Pascanu, Samy Bengio, Yoshua Bengio
ICML1
2015 A Recurrent Latent Variable Model for Sequential Data
abstract
In this paper, we explore the inclusion of latent random variables into the hidden state of a recurrent neural network (RNN) by combining the elements of the variational autoencoder. We argue that through the use of high-level latent random variables, the variational RNN (VRNN) can model the kind of variability observed in highly structured sequential data such as natural speech. We empirically evaluate the proposed model against other related sequential models on four speech datasets and one handwriting dataset. Our results show the important roles that latent random variables can play in the RNN dynamics.
Junyoung Chung, Kyle Kastner, Laurent Dinh, Kratarth Goel, Aaron C. Courville, Yoshua Bengio
NIPS3
2013 Predicting Parameters in Deep Learning
abstract
We demonstrate that there is significant redundancy in the parameterization of several deep learning models. Given only a few weight values for each feature it is possible to accurately predict the remaining values. Moreover, we show that not only can the parameter values be predicted, but many of them need not be learned at all. We train several different architectures by learning only a small number of weights and predicting the rest. In the best case we are able to predict more than 95% of the weights of a network without any drop in accuracy.
Misha Denil, Babak Shakibi, Laurent Dinh, Marc'Aurelio Ranzato, Nando de Freitas
NIPS3