VLDB 2026 Research / reviewers in the wild / expert
Michael McCabe
dblp:56/706
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Generative modeling · 39% Representation and self-supervised learning · 18% Segmentation and scene understanding · 10% | |
| Interdisciplinary, comprehensive, and emerging computing
7 papers |
Computational science and engineering · 96% Environmental and earth informatics · 4% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 18 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
1.7 | 2 | 2025 | Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics Emulation · NeurIPS 2025 Predicting partially observable dynamical systems via diffusion models with a multiscale inference scheme · NeurIPS 2025 |
Computational science and engineering › scientific machine learning
surrogate modeling |
1.5 | 2 | 2024 | The Well: a Large-Scale Collection of Diverse Physics Simulations for Machine Learning · NeurIPS 2024 Multiple Physics Pretraining for Spatiotemporal Surrogate Models · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model
conditional diffusion model |
0.9 | 1 | 2025 | Predicting partially observable dynamical systems via diffusion models with a multiscale inference scheme · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
foundation model |
0.9 | 1 | 2025 | AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
latent diffusion model |
0.9 | 1 | 2025 | Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics Emulation · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked modeling |
0.9 | 1 | 2025 | AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025 |
Machine learning › Learning paradigms › semi-supervised learning
pseudo-labeling |
0.9 | 1 | 2025 | SmokeViz: A Large-Scale Satellite Dataset for Wildfire Smoke Detection and Segmentation · NeurIPS 2025 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.9 | 1 | 2025 | SmokeViz: A Large-Scale Satellite Dataset for Wildfire Smoke Detection and Segmentation · NeurIPS 2025 |
Computational science and engineering › dynamical systems
dynamical system simulation |
0.9 | 1 | 2025 | Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics Emulation · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning
pre-training |
0.8 | 1 | 2024 | Multiple Physics Pretraining for Spatiotemporal Surrogate Models · NeurIPS 2024 |
Computational science and engineering › computational physics
physics simulation |
0.8 | 1 | 2024 | The Well: a Large-Scale Collection of Diverse Physics Simulations for Machine Learning · NeurIPS 2024 |
Information retrieval › evaluation
benchmark dataset |
0.8 | 1 | 2024 | The Well: a Large-Scale Collection of Diverse Physics Simulations for Machine Learning · NeurIPS 2024 |
Computational science and engineering › dynamical systems › nonlinear dynamics
chaotic dynamics |
0.5 | 1 | 2021 | Learning to Assimilate in Chaotic Dynamical Systems · NeurIPS 2021 |
Computational science and engineering
data assimilation |
0.5 | 1 | 2021 | Learning to Assimilate in Chaotic Dynamical Systems · NeurIPS 2021 |
Computer vision › Image recognition and object detection
multi-scale inference |
0.3 | 1 | 2025 | Predicting partially observable dynamical systems via diffusion models with a multiscale inference scheme · NeurIPS 2025 |
Computational science and engineering › astronomy
astronomical data analysis |
0.3 | 1 | 2025 | AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025 |
Computational science and engineering
astronomy |
0.3 | 1 | 2025 | AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
transformer · 3.3tokenization · 1.7satellite imagery · 1.7pseudo-labeling · 1.7multiscale inference scheme · 1.7masked modeling · 1.7diffusion model · 1.7deep learning · 1.7autoregressive rollout · 1.7autoencoder · 1.7pytorch · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SmokeViz: A Large-Scale Satellite Dataset for Wildfire Smoke Detection and SegmentationabstractThe global rise in wildfire frequency and intensity over the past decade underscores the need for improved fire monitoring techniques. To advance deep learning research on wildfire detection and its associated human health impacts, we introduce SmokeViz, a large-scale machine learning dataset of smoke plumes in satellite imagery. The dataset is derived from expert annotations created by smoke analysts at the National Oceanic and Atmospheric Administration, which provide coarse temporal and spatial approximations of smoke presence. To enhance annotation precision, we propose pseudo-label dimension reduction (PLDR), a generalizable method that applies pseudo-labeling to refine datasets with mismatching temporal and/or spatial resolutions. Unlike typical pseudo-labeling applications that aim to increase the number of labeled samples, PLDR maintains the original labels but increases the dataset quality by solving for intermediary pseudo-labels (IPLs) that align each annotation to the most representative input data. For SmokeViz, a parent model produces IPLs to identify the single satellite image within each annotations time window that best corresponds with the smoke plume. This refinement process produces a succinct and relevant deep learning dataset consisting of over 160,000 manual annotations. The SmokeViz dataset is expected to be a valuable resource to develop further wildfire-related machine learning models and is publicly available at \url{https://noaa-gsl-experimental-pds.s3.amazonaws.com/index.html#SmokeViz/}. Rey Koki, Michael McCabe, Dhruv Kedar, Josh Myers-Dean, Annabel Wade, Jebb Q. Stewart, Christina Kumler-Bonfanti, Jed Brown |
NeurIPS | 2 |
| 2025 | Predicting partially observable dynamical systems via diffusion models with a multiscale inference schemeabstractConditional diffusion models provide a natural framework for probabilistic prediction of dynamical systems and have been successfully applied to fluid dynamics and weather prediction. However, in many settings, the available information at a given time represents only a small fraction of what is needed to predict future states, either due to measurement uncertainty or because only a small fraction of the state can be observed. This is true for example in solar physics, where we can observe the Sun’s surface and atmosphere, but its evolution is driven by internal processes for which we lack direct measurements. In this paper, we tackle the probabilistic prediction of partially observable, long-memory dynamical systems, with applications to solar dynamics and the evolution of active regions. We show that standard inference schemes, such as autoregressive rollouts, fail to capture long-range dependencies in the data, largely because they do not integrate past information effectively. To overcome this, we propose a multiscale inference scheme for diffusion models, tailored to physical processes. Our method generates trajectories that are temporally fine-grained near the present and coarser as we move farther away, which enables capturing long-range temporal dependencies without increasing computational cost. When integrated into a diffusion model, we show that our inference scheme significantly reduces the bias of the predicted distributions and improves rollout stability. Rudy Morel, Francesco Pio Ramunno, Jeff Shen, Alberto Bietti, Kyunghyun Cho, Miles D. Cranmer, Siavash Golkar, Olexandr Gugnin, Géraud Krawezik, Tanya Marwah, Michael McCabe, Lucas Meyer, Payel Mukhopadhyay, Ruben Ohana, Liam Holden Parker, Helen Qu, François Rozet, K. D. Leka, François Lanusse, David F. Fouhey, Shirley Ho |
NeurIPS | 11 |
| 2025 | AION-1: Omnimodal Foundation Model for Astronomical SciencesabstractWhile foundation models have shown promise across a variety of fields, astronomy lacks a unified framework for joint modeling across its highly diverse data modalities. In this paper, we present AION-1, the first large-scale multimodal foundation family of models for astronomy. AION-1 enables arbitrary transformations between heterogeneous data types using a two-stage architecture: modality-specific tokenization followed by transformer-based masked modeling of cross-modal token sequences. Trained on over 200M astronomical objects, AION-1 demonstrates strong performance across regression, classification, generation, and object retrieval tasks. Beyond astronomy, AION-1 provides a scalable blueprint for multimodal scientific foundation models that can seamlessly integrate heterogeneous combinations of real-world observations. Our model release is entirely open source, including the dataset, training script, and weights. Liam Holden Parker, François Lanusse, Jeff Shen, Ollie Liu, Tom Hehir, Leopoldo Sarra, Lucas Meyer, Micah Bowles, Sebastian Wagner-Carena, Helen Qu, Siavash Golkar, Alberto Bietti, Hatim Bourfoune, Pierre Cornette, Keiya Hirashima, Géraud Krawezik, Ruben Ohana, Nicholas Lourie, Michael McCabe, Rudy Morel, Payel Mukhopadhyay, Mariel Pettee, Kyunghyun Cho, Miles D. Cranmer, Shirley Ho |
NeurIPS | 19 |
| 2025 | Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics EmulationabstractThe steep computational cost of diffusion models at inference hinders their use as fast physics emulators. In the context of image and video generation, this computational drawback has been addressed by generating in the latent space of an autoencoder instead of the pixel space. In this work, we investigate whether a similar strategy can be effectively applied to the emulation of dynamical systems and at what cost. We find that the accuracy of latent-space emulation is surprisingly robust to a wide range of compression rates (up to 1000x). We also show that diffusion-based emulators are consistently more accurate than non-generative counterparts and compensate for uncertainty in their predictions with greater diversity. Finally, we cover practical design choices, spanning from architectures to optimizers, that we found critical to train latent-space emulators. François Rozet, Ruben Ohana, Michael McCabe, Gilles Louppe, François Lanusse, Shirley Ho |
NeurIPS | 3 |
| 2024 | Multiple Physics Pretraining for Spatiotemporal Surrogate ModelsabstractWe introduce multiple physics pretraining (MPP), an autoregressive task-agnostic pretraining approach for physical surrogate modeling of spatiotemporal systems with transformers. In MPP, rather than training one model on a specific physical system, we train a backbone model to predict the dynamics of multiple heterogeneous physical systems simultaneously in order to learn features that are broadly useful across systems and facilitate transfer. In order to learn effectively in this setting, we introduce a shared embedding and normalization strategy that projects the fields of multiple systems into a shared embedding space. We validate the efficacy of our approach on both pretraining and downstream tasks over a broad fluid mechanics-oriented benchmark. We show that a single MPP-pretrained transformer is able to match or outperform task-specific baselines on all pretraining sub-tasks without the need for finetuning. For downstream tasks, we demonstrate that finetuning MPP-trained models results in more accurate predictions across multiple time-steps on systems with previously unseen physical components or higher dimensional systems compared to training from scratch or finetuning pretrained video foundation models. We open-source our code and model weights trained at multiple scales for reproducibility. Michael McCabe, Bruno Régaldo-Saint Blancard, Liam Holden Parker, Ruben Ohana, Miles D. Cranmer, Alberto Bietti, Michael Eickenberg, Siavash Golkar, Géraud Krawezik, François Lanusse, Mariel Pettee, Tiberiu Tesileanu, Kyunghyun Cho, Shirley Ho |
NeurIPS | 1 |
| 2024 | The Well: a Large-Scale Collection of Diverse Physics Simulations for Machine LearningabstractMachine learning based surrogate models offer researchers powerful tools for accelerating simulation-based workflows. However, as standard datasets in this space often cover small classes of physical behavior, it can be difficult to evaluate the efficacy of new approaches. To address this gap, we introduce the Well: a large-scale collection of datasets containing numerical simulations of a wide variety of spatiotemporal physical systems. The Well draws from domain experts and numerical software developers to provide 15TB of data across 16 datasets covering diverse domains such as biological systems, fluid dynamics, acoustic scattering, as well as magneto-hydrodynamic simulations of extra-galactic fluids or supernova explosions. These datasets can be used individually or as part of a broader benchmark suite. To facilitate usage of the Well, we provide a unified PyTorch interface for training and evaluating models. We demonstrate the function of this library by introducing example baselines that highlight the new challenges posed by the complex dynamics of the Well. The code and data is available at https://github.com/PolymathicAI/the_well. Ruben Ohana, Michael McCabe, Lucas Meyer, Rudy Morel, Fruzsina Julia Agocs, Miguel Beneitez, Marsha J. Berger, Blakesley Burkhart, Stuart B. Dalziel, Drummond B. Fielding, Daniel Fortunato, Jared A. Goldberg, Keiya Hirashima, Yan-Fei Jiang, Rich R. Kerswell, Suryanarayana Maddu, Jonah Miller, Payel Mukhopadhyay, Stefan S. Nixon, Jeff Shen, Romain Watteaux, Bruno Régaldo-Saint Blancard, François Rozet, Liam Holden Parker, Miles D. Cranmer, Shirley Ho |
NeurIPS | 2 |
| 2021 | Learning to Assimilate in Chaotic Dynamical SystemsabstractThe accuracy of simulation-based forecasting in chaotic systems is heavily dependent on high-quality estimates of the system state at the beginning of the forecast. Data assimilation methods are used to infer these initial conditions by systematically combining noisy, incomplete observations and numerical models of system dynamics to produce highly effective estimation schemes. We introduce a self-supervised framework, which we call \textit{amortized assimilation}, for learning to assimilate in dynamical systems. Amortized assimilation combines deep learning-based denoising with differentiable simulation, using independent neural networks to assimilate specific observation types while connecting the gradient flow between these sub-tasks with differentiable simulation and shared recurrent memory. This hybrid architecture admits a self-supervised training objective which is minimized by an unbiased estimator of the true system state even in the presence of only noisy training data. Numerical experiments across several chaotic benchmark systems highlight the improved effectiveness of our approach compared to widely-used data assimilation methods. Michael McCabe, Jed Brown |
NeurIPS | 1 |