Louis Serrano

dblp:349/0965 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
5 papers
Computational science and engineering · 100%
Artificial intelligence
5 papers
Efficient and distributed learning · 42% Generative modeling · 28% Deep learning architectures and training · 22%

Topics — the 17 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
autoregressive model
1.722025
ENMA: Tokenwise Autoregression for Continuous Neural PDE Operators · NeurIPS 2025
Zebra: In-Context Generative Pretraining for Solving Parametric PDEs · ICML 2025
Computational science and engineering › scientific machine learning
neural operator
1.522025
ENMA: Tokenwise Autoregression for Continuous Neural PDE Operators · NeurIPS 2025
Operator Learning with Neural Fields: Tackling PDEs on General Geometries · NeurIPS 2023
Computational science and engineering › scientific machine learning › physics-informed machine learning › physics-informed neural networks
partial differential equation solving
1.422024
AROMA: Preserving Spatial Structure for Latent PDE Modeling with Local Neural Fields · NeurIPS 2024
Operator Learning with Neural Fields: Tackling PDEs on General Geometries · NeurIPS 2023
Machine learning › Efficient and distributed learning › distributed training
communication-efficient training
0.912025
ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training · NeurIPS 2025
Machine learning › Efficient and distributed learning
distributed training
0.912025
ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training · NeurIPS 2025
Machine learning › Deep learning architectures and training › neural network layer design
feature upsampling
0.912025
JAFAR: Jack up Any Feature at Any Resolution · NeurIPS 2025
Machine learning › Efficient and distributed learning › distributed training
gradient aggregation
0.912025
ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training · NeurIPS 2025
Computational science and engineering › partial differential equations
parametric PDEs
0.912025
Learning a Neural Solver for Parametric PDEs to Enhance Physics-Informed Methods · ICLR 2025
Computational science and engineering › scientific machine learning
PDE operator learning
0.912025
ENMA: Tokenwise Autoregression for Continuous Neural PDE Operators · NeurIPS 2025
Computational science and engineering › scientific machine learning › physics-informed machine learning
physics-informed neural networks
0.912025
Learning a Neural Solver for Parametric PDEs to Enhance Physics-Informed Methods · ICLR 2025
Computational science and engineering
scientific machine learning
0.912025
Zebra: In-Context Generative Pretraining for Solving Parametric PDEs · ICML 2025
Computational science and engineering › numerical analysis
model reduction
0.812024
AROMA: Preserving Spatial Structure for Latent PDE Modeling with Local Neural Fields · NeurIPS 2024
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent
0.312025
Learning a Neural Solver for Parametric PDEs to Enhance Physics-Informed Methods · ICLR 2025
Natural language and speech › Language models and text generation
in-context learning
0.312025
ENMA: Tokenwise Autoregression for Continuous Neural PDE Operators · NeurIPS 2025
Machine learning › Deep learning architectures and training
transformer
0.312025
Zebra: In-Context Generative Pretraining for Solving Parametric PDEs · ICML 2025
Machine learning › Deep learning architectures and training
vision encoder
0.312025
JAFAR: Jack up Any Feature at Any Resolution · NeurIPS 2025
High-performance computing
large-scale training
0.312025
ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

attention mechanism · 2.6uncertainty quantification · 1.7physics-informed learning · 1.7in-context learning · 1.7gradient descent · 1.7flow matching · 1.7delayed gradient synchronization · 1.7convolutional encoder · 1.7autoregressive transformer · 1.7neural field · 1.4spatial feature transform · 0.9optimizer state sharding · 0.9diffusion model · 0.8
YearPublicationVenuePosition
2025 Learning a Neural Solver for Parametric PDEs to Enhance Physics-Informed Methods
abstract
Physics-informed deep learning often faces optimization challenges due to the complexity of solving partial differential equations (PDEs), which involve exploring large solution spaces, require numerous iterations, and can lead to unstable training. These challenges arise particularly from the ill-conditioning of the optimization problem, caused by the differential terms in the loss function. To address these issues, we propose learning a solver, i.e., solving PDEs using a physics-informed iterative algorithm trained on data. Our method learns to condition a gradient descent algorithm that automatically adapts to each PDE instance, significantly accelerating and stabilizing the optimization process and enabling faster convergence of physics-aware models. Furthermore, while traditional physics-informed methods solve for a single PDE instance, our approach addresses parametric PDEs. Specifically, our method integrates the physical loss gradient with the PDE parameters to solve over a distribution of PDE parameters, including coefficients, initial conditions, or boundary conditions. We demonstrate the effectiveness of our method through empirical experiments on multiple datasets, comparing training and test-time optimization performance.
Lise Le Boudec, Emmanuel de Bézenac, Louis Serrano, Ramon Daniel Regueiro-Espino, Patrick Gallinari
ICLR3
2025 Zebra: In-Context Generative Pretraining for Solving Parametric PDEs
abstract
Solving time-dependent parametric partial differential equations (PDEs) is challenging for data-driven methods, as these models must adapt to variations in parameters such as coefficients, forcing terms, and initial conditions. State-of-the-art neural surrogates perform adaptation through gradient-based optimization and meta-learning to implicitly encode the variety of dynamics from observations. This often comes with increased inference complexity. Inspired by the in-context learning capabilities of large language models (LLMs), we introduce Zebra, a novel generative auto-regressive transformer designed to solve parametric PDEs without requiring gradient adaptation at inference. By leveraging in-context information during both pre-training and inference, Zebra dynamically adapts to new tasks by conditioning on input sequences that incorporate context example trajectories. As a generative model, Zebra can be used to generate new trajectories and allows quantifying the uncertainty of the predictions. We evaluate Zebra across a variety of challenging PDE scenarios, demonstrating its adaptability, robustness, and superior performance compared to existing approaches.
Louis Serrano, Armand Kassaï Koupaï, Thomas X. Wang, Pierre Erbacher, Patrick Gallinari
ICML1
2025 JAFAR: Jack up Any Feature at Any Resolution
abstract
Foundation Vision Encoders have become indispensable across a wide range of dense vision tasks. However, their operation at low spatial feature resolutions necessitates subsequent feature decompression to enable full-resolution processing. To address this limitation, we introduce JAFAR, a lightweight and flexible feature upsampler designed to enhance the spatial resolution of visual features from any Foundation Vision Encoder to any target resolution. JAFAR features an attention-based upsampling module that aligns the spatial representations of high-resolution queries with semantically enriched low-resolution keys via Spatial Feature Transform modulation. Despite the absence of high-resolution feature ground truth; we find that learning at low upsampling ratios and resolutions generalizes surprisingly well to much higher scales. Extensive experiments demonstrate that JAFAR recovers intricate pixel-level details and consistently outperforms existing feature upsampling techniques across a diverse set of dense downstream applications.
Paul Couairon, Loïck Chambon, Louis Serrano, Jean-Emmanuel Haugeard, Matthieu Cord, Nicolas Thome
NeurIPS3
2025 ENMA: Tokenwise Autoregression for Continuous Neural PDE Operators
abstract
Solving time-dependent parametric partial differential equations (PDEs) remains a fundamental challenge for neural solvers, particularly when generalizing across a wide range of physical parameters and dynamics. When data is uncertain or incomplete—as is often the case—a natural approach is to turn to generative models. We introduce ENMA, a generative neural operator designed to model spatio-temporal dynamics arising from physical phenomena. ENMA predicts future dynamics in a compressed latent space using a generative masked autoregressive transformer trained with flow matching loss, enabling tokenwise generation. Irregularly sampled spatial observations are encoded into uniform latent representations via attention mechanisms and further compressed through a spatio-temporal convolutional encoder. This allows ENMA to perform in-context learning at inference time by conditioning on either past states of the target trajectory or auxiliary context trajectories with similar dynamics. The result is a robust and adaptable framework that generalizes to new PDE regimes and supports one-shot surrogate modeling of time-dependent parametric PDEs.
Armand Kassaï Koupaï, Lise Le Boudec, Louis Serrano, Patrick Gallinari
NeurIPS3
2025 ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training
abstract
Training LLMs relies on distributed implementations using multiple GPUs to compute gradients in parallel with sharded optimizers. However, synchronizing gradients in data parallel setups introduces communication overhead that grows with the number of workers, limiting parallelization efficiency. Local optimization algorithms reduce communications but incur high memory costs as they prevent optimizer state sharding, hindering scalability. To address this, we propose $\textbf{AC}$cumulate while $\textbf{CO}$mmunicate ($\texttt{ACCO}$), a memory-efficient optimization algorithm for distributed LLM training. By synchronizing delayed gradients while computing new ones, $\texttt{ACCO}$ reduces GPU idle time and supports heterogeneous hardware. To mitigate the convergence issues caused by delayed updates, we introduce a novel technique ensuring training dynamics align with standard distributed optimization. Compared to ZeRO-1, our approach is significantly faster and scales effectively across heterogeneous hardware.
Adel Nabli, Louis Fournier, Pierre Erbacher, Louis Serrano, Eugene Belilovsky, Edouard Oyallon
NeurIPS4
2024 AROMA: Preserving Spatial Structure for Latent PDE Modeling with Local Neural Fields
abstract
We present AROMA (Attentive Reduced Order Model with Attention), a framework designed to enhance the modeling of partial differential equations (PDEs) using local neural fields. Our flexible encoder-decoder architecture can obtain smooth latent representations of spatial physical fields from a variety of data types, including irregular-grid inputs and point clouds. This versatility eliminates the need for patching and allows efficient processing of diverse geometries. The sequential nature of our latent representation can be interpreted spatially and permits the use of a conditional transformer for modeling the temporal dynamics of PDEs. By employing a diffusion-based formulation, we achieve greater stability and enable longer rollouts compared to conventional MSE training. AROMA's superior performance in simulating 1D and 2D equations underscores the efficacy of our approach in capturing complex dynamical behaviors.
Louis Serrano, Thomas X. Wang, Etienne Le Naour, Jean-Noël Vittaut, Patrick Gallinari
NeurIPS1
2023 Operator Learning with Neural Fields: Tackling PDEs on General Geometries
abstract
Machine learning approaches for solving partial differential equations require learning mappings between function spaces. While convolutional or graph neural networks are constrained to discretized functions, neural operators present a promising milestone toward mapping functions directly. Despite impressive results they still face challenges with respect to the domain geometry and typically rely on some form of discretization. In order to alleviate such limitations, we present CORAL, a new method that leverages coordinate-based networks for solving PDEs on general geometries. CORAL is designed to remove constraints on the input mesh, making it applicable to any spatial sampling and geometry. Its ability extends to diverse problem domains, including PDE solving, spatio-temporal forecasting, and inverse problems like geometric design. CORAL demonstrates robust performance across multiple resolutions and performs well in both convex and non-convex domains, surpassing or performing on par with state-of-the-art models.
Louis Serrano, Lise Le Boudec, Armand Kassaï Koupaï, Thomas X. Wang, Jean-Noël Vittaut, Patrick Gallinari
NeurIPS1