VLDB 2026 Research / reviewers in the wild / expert
François Rozet
dblp:303/4668
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Generative modeling · 54% Probabilistic and Bayesian machine learning · 44% Image recognition and object detection · 3% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Computational science and engineering · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
2.5 | 3 | 2025 | Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics Emulation · NeurIPS 2025 Predicting partially observable dynamical systems via diffusion models with a multiscale inference scheme · NeurIPS 2025 Learning Diffusion Priors from Observations by Expectation Maximization · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › posterior inference
bayesian inverse problems |
1.4 | 2 | 2024 | Learning Diffusion Priors from Observations by Expectation Maximization · NeurIPS 2024 Score-based Data Assimilation · NeurIPS 2023 |
Machine learning › Generative modeling › diffusion model
conditional diffusion model |
0.9 | 1 | 2025 | Predicting partially observable dynamical systems via diffusion models with a multiscale inference scheme · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
latent diffusion model |
0.9 | 1 | 2025 | Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics Emulation · NeurIPS 2025 |
Computational science and engineering › dynamical systems
dynamical system simulation |
0.9 | 1 | 2025 | Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics Emulation · NeurIPS 2025 |
Machine learning › Probabilistic and Bayesian machine learning › sampling
posterior sampling |
0.8 | 1 | 2024 | Learning Diffusion Priors from Observations by Expectation Maximization · NeurIPS 2024 |
Computational science and engineering › computational physics
physics simulation |
0.8 | 1 | 2024 | The Well: a Large-Scale Collection of Diverse Physics Simulations for Machine Learning · NeurIPS 2024 |
Computational science and engineering › scientific machine learning
surrogate modeling |
0.8 | 1 | 2024 | The Well: a Large-Scale Collection of Diverse Physics Simulations for Machine Learning · NeurIPS 2024 |
Information retrieval › evaluation
benchmark dataset |
0.8 | 1 | 2024 | The Well: a Large-Scale Collection of Diverse Physics Simulations for Machine Learning · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model
score-based generative model |
0.7 | 1 | 2023 | Score-based Data Assimilation · NeurIPS 2023 |
Machine learning › Probabilistic and Bayesian machine learning
trajectory inference |
0.7 | 1 | 2023 | Score-based Data Assimilation · NeurIPS 2023 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › approximate bayesian inference
simulation-based inference |
0.6 | 1 | 2022 | Towards Reliable Simulation-Based Inference with Balanced Neural Ratio Estimation · NeurIPS 2022 |
Computer vision › Image recognition and object detection
multi-scale inference |
0.3 | 1 | 2025 | Predicting partially observable dynamical systems via diffusion models with a multiscale inference scheme · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
diffusion model · 2.5multiscale inference scheme · 1.7autoregressive rollout · 1.7autoencoder · 1.7pytorch · 1.5posterior sampling · 0.8expectation-maximization · 0.8score-based generative model · 0.7non-autoregressive generation · 0.7neural ratio estimation · 0.6bayesian inference · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Predicting partially observable dynamical systems via diffusion models with a multiscale inference schemeabstractConditional diffusion models provide a natural framework for probabilistic prediction of dynamical systems and have been successfully applied to fluid dynamics and weather prediction. However, in many settings, the available information at a given time represents only a small fraction of what is needed to predict future states, either due to measurement uncertainty or because only a small fraction of the state can be observed. This is true for example in solar physics, where we can observe the Sun’s surface and atmosphere, but its evolution is driven by internal processes for which we lack direct measurements. In this paper, we tackle the probabilistic prediction of partially observable, long-memory dynamical systems, with applications to solar dynamics and the evolution of active regions. We show that standard inference schemes, such as autoregressive rollouts, fail to capture long-range dependencies in the data, largely because they do not integrate past information effectively. To overcome this, we propose a multiscale inference scheme for diffusion models, tailored to physical processes. Our method generates trajectories that are temporally fine-grained near the present and coarser as we move farther away, which enables capturing long-range temporal dependencies without increasing computational cost. When integrated into a diffusion model, we show that our inference scheme significantly reduces the bias of the predicted distributions and improves rollout stability. Rudy Morel, Francesco Pio Ramunno, Jeff Shen, Alberto Bietti, Kyunghyun Cho, Miles D. Cranmer, Siavash Golkar, Olexandr Gugnin, Géraud Krawezik, Tanya Marwah, Michael McCabe, Lucas Meyer, Payel Mukhopadhyay, Ruben Ohana, Liam Holden Parker, Helen Qu, François Rozet, K. D. Leka, François Lanusse, David F. Fouhey, Shirley Ho |
NeurIPS | 17 |
| 2025 | Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics EmulationabstractThe steep computational cost of diffusion models at inference hinders their use as fast physics emulators. In the context of image and video generation, this computational drawback has been addressed by generating in the latent space of an autoencoder instead of the pixel space. In this work, we investigate whether a similar strategy can be effectively applied to the emulation of dynamical systems and at what cost. We find that the accuracy of latent-space emulation is surprisingly robust to a wide range of compression rates (up to 1000x). We also show that diffusion-based emulators are consistently more accurate than non-generative counterparts and compensate for uncertainty in their predictions with greater diversity. Finally, we cover practical design choices, spanning from architectures to optimizers, that we found critical to train latent-space emulators. François Rozet, Ruben Ohana, Michael McCabe, Gilles Louppe, François Lanusse, Shirley Ho |
NeurIPS | 1 |
| 2024 | The Well: a Large-Scale Collection of Diverse Physics Simulations for Machine LearningabstractMachine learning based surrogate models offer researchers powerful tools for accelerating simulation-based workflows. However, as standard datasets in this space often cover small classes of physical behavior, it can be difficult to evaluate the efficacy of new approaches. To address this gap, we introduce the Well: a large-scale collection of datasets containing numerical simulations of a wide variety of spatiotemporal physical systems. The Well draws from domain experts and numerical software developers to provide 15TB of data across 16 datasets covering diverse domains such as biological systems, fluid dynamics, acoustic scattering, as well as magneto-hydrodynamic simulations of extra-galactic fluids or supernova explosions. These datasets can be used individually or as part of a broader benchmark suite. To facilitate usage of the Well, we provide a unified PyTorch interface for training and evaluating models. We demonstrate the function of this library by introducing example baselines that highlight the new challenges posed by the complex dynamics of the Well. The code and data is available at https://github.com/PolymathicAI/the_well. Ruben Ohana, Michael McCabe, Lucas Meyer, Rudy Morel, Fruzsina Julia Agocs, Miguel Beneitez, Marsha J. Berger, Blakesley Burkhart, Stuart B. Dalziel, Drummond B. Fielding, Daniel Fortunato, Jared A. Goldberg, Keiya Hirashima, Yan-Fei Jiang, Rich R. Kerswell, Suryanarayana Maddu, Jonah Miller, Payel Mukhopadhyay, Stefan S. Nixon, Jeff Shen, Romain Watteaux, Bruno Régaldo-Saint Blancard, François Rozet, Liam Holden Parker, Miles D. Cranmer, Shirley Ho |
NeurIPS | 23 |
| 2024 | Learning Diffusion Priors from Observations by Expectation MaximizationabstractDiffusion models recently proved to be remarkable priors for Bayesian inverse problems. However, training these models typically requires access to large amounts of clean data, which could prove difficult in some settings. In this work, we present a novel method based on the expectation-maximization algorithm for training diffusion models from incomplete and noisy observations only. Unlike previous works, our method leads to proper diffusion models, which is crucial for downstream tasks. As part of our method, we propose and motivate an improved posterior sampling scheme for unconditional diffusion models. We present empirical evidence supporting the effectiveness of our method. François Rozet, Gérôme Andry, François Lanusse, Gilles Louppe |
NeurIPS | 1 |
| 2023 | Score-based Data AssimilationabstractData assimilation, in its most comprehensive form, addresses the Bayesian inverse problem of identifying plausible state trajectories that explain noisy or incomplete observations of stochastic dynamical systems. Various approaches have been proposed to solve this problem, including particle-based and variational methods. However, most algorithms depend on the transition dynamics for inference, which becomes intractable for long time horizons or for high-dimensional systems with complex dynamics, such as oceans or atmospheres. In this work, we introduce score-based data assimilation for trajectory inference. We learn a score-based generative model of state trajectories based on the key insight that the score of an arbitrarily long trajectory can be decomposed into a series of scores over short segments. After training, inference is carried out using the score model, in a non-autoregressive manner by generating all states simultaneously. Quite distinctively, we decouple the observation model from the training procedure and use it only at inference to guide the generative process, which enables a wide range of zero-shot observation scenarios. We present theoretical and empirical evidence supporting the effectiveness of our method. François Rozet, Gilles Louppe |
NeurIPS | 1 |
| 2022 | Towards Reliable Simulation-Based Inference with Balanced Neural Ratio EstimationabstractModern approaches for simulation-based inference build upon deep learning surrogates to enable approximate Bayesian inference with computer simulators. In practice, the estimated posteriors' computational faithfulness is, however, rarely guaranteed. For example, Hermans et al., 2021 have shown that current simulation-based inference algorithms can produce posteriors that are overconfident, hence risking false inferences. In this work, we introduce Balanced Neural Ratio Estimation (BNRE), a variation of the NRE algorithm designed to produce posterior approximations that tend to be more conservative, hence improving their reliability, while sharing the same Bayes optimal solution. We achieve this by enforcing a balancing condition that increases the quantified uncertainty in low simulation budget regimes while still converging to the exact posterior as the budget increases. We provide theoretical arguments showing that BNRE tends to produce posterior surrogates that are more conservative than NRE's. We evaluate BNRE on a wide variety of tasks and show that it produces conservative posterior surrogates on all tested benchmarks and simulation budgets. Finally, we emphasize that BNRE is straightforward to implement over NRE and does not introduce any computational overhead. Arnaud Delaunoy, Joeri Hermans, François Rozet, Antoine Wehenkel, Gilles Louppe |
NeurIPS | 3 |