VLDB 2026 Research / reviewers in the wild / expert
Shirley Ho
dblp:162/2218
· DBLP profile ↗
13ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0002-1068-160XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 7 since 2021Systems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Generative modeling · 36% Representation and self-supervised learning · 24% Deep learning architectures and training · 18% | |
| Interdisciplinary, comprehensive, and emerging computing
10 papers |
Computational science and engineering · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 26 heaviest of 31, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
1.7 | 2 | 2025 | Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics Emulation · NeurIPS 2025 Predicting partially observable dynamical systems via diffusion models with a multiscale inference scheme · NeurIPS 2025 |
Computational science and engineering › scientific machine learning
surrogate modeling |
1.5 | 2 | 2024 | The Well: a Large-Scale Collection of Diverse Physics Simulations for Machine Learning · NeurIPS 2024 Multiple Physics Pretraining for Spatiotemporal Surrogate Models · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model
conditional diffusion model |
0.9 | 1 | 2025 | Predicting partially observable dynamical systems via diffusion models with a multiscale inference scheme · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
foundation model |
0.9 | 1 | 2025 | AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
latent diffusion model |
0.9 | 1 | 2025 | Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics Emulation · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked modeling |
0.9 | 1 | 2025 | AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025 |
Computational science and engineering › dynamical systems
dynamical system simulation |
0.9 | 1 | 2025 | Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics Emulation · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning
pre-training |
0.8 | 1 | 2024 | Multiple Physics Pretraining for Spatiotemporal Surrogate Models · NeurIPS 2024 |
Computational science and engineering › astronomy
astrophysics |
0.8 | 1 | 2024 | The Multimodal Universe: Enabling Large-Scale Machine Learning with 100 TB of Astronomical Scientific Data · NeurIPS 2024 |
Computational science and engineering › computational physics
physics simulation |
0.8 | 1 | 2024 | The Well: a Large-Scale Collection of Diverse Physics Simulations for Machine Learning · NeurIPS 2024 |
Information retrieval › evaluation
benchmark dataset |
0.8 | 1 | 2024 | The Well: a Large-Scale Collection of Diverse Physics Simulations for Machine Learning · NeurIPS 2024 |
Machine learning › Deep learning architectures and training › scientific machine learning
neural surrogate model |
0.6 | 1 | 2022 | Learned Simulators for Turbulence · ICLR 2022 |
Computational science and engineering › computational fluid dynamics
turbulence simulation |
0.6 | 1 | 2022 | Learned Simulators for Turbulence · ICLR 2022 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
symbolic regression |
0.4 | 1 | 2020 | Discovering Symbolic Models from Deep Learning with Inductive Biases · NeurIPS 2020 |
Computational science and engineering
cosmology |
0.3 | 2 | 2018 | Estimating Cosmological Parameters from the Dark Matter Distribution · ICML 2016 CosmoFlow: using deep learning to learn the universe at scale · SC 2018 |
High-performance computing
large-scale training |
0.3 | 1 | 2018 | CosmoFlow: using deep learning to learn the universe at scale · SC 2018 |
Computational science and engineering
astronomy |
0.3 | 2 | 2025 | AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025 Finding Galaxies in the Shadows of Quasars with Gaussian Processes · ICML 2015 |
Computer vision › Image recognition and object detection
multi-scale inference |
0.3 | 1 | 2025 | Predicting partially observable dynamical systems via diffusion models with a multiscale inference scheme · NeurIPS 2025 |
Computational science and engineering › astronomy
astronomical data analysis |
0.3 | 1 | 2025 | AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025 |
Computational science and engineering › cosmology
cosmological parameter estimation |
0.2 | 1 | 2016 | Estimating Cosmological Parameters from the Dark Matter Distribution · ICML 2016 |
Machine learning › Representation and self-supervised learning
multimodal representation learning |
0.2 | 1 | 2024 | The Multimodal Universe: Enabling Large-Scale Machine Learning with 100 TB of Astronomical Scientific Data · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
density estimation |
0.2 | 1 | 2015 | Optimal Ridge Detection using Coverage Risk · NIPS 2015 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.2 | 1 | 2015 | Finding Galaxies in the Shadows of Quasars with Gaussian Processes · ICML 2015 |
Machine learning › Graph learning
graph neural network |
0.1 | 1 | 2020 | Discovering Symbolic Models from Deep Learning with Inductive Biases · NeurIPS 2020 |
Computer vision › 3D vision › 3d shape representation
volumetric representation |
0.1 | 1 | 2016 | Estimating Cosmological Parameters from the Dark Matter Distribution · ICML 2016 |
Methods — techniques the papers use, named apart from their topics
transformer · 3.3tokenization · 1.7multiscale inference scheme · 1.7masked modeling · 1.7diffusion model · 1.7autoregressive rollout · 1.7autoencoder · 1.7pytorch · 1.5multimodal machine learning · 1.5autoregressive modeling · 1.5graph neural network · 1.0deep learning · 0.3coverage risk estimator · 0.2convergence rate analysis · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Predicting partially observable dynamical systems via diffusion models with a multiscale inference schemeabstractConditional diffusion models provide a natural framework for probabilistic prediction of dynamical systems and have been successfully applied to fluid dynamics and weather prediction. However, in many settings, the available information at a given time represents only a small fraction of what is needed to predict future states, either due to measurement uncertainty or because only a small fraction of the state can be observed. This is true for example in solar physics, where we can observe the Sun’s surface and atmosphere, but its evolution is driven by internal processes for which we lack direct measurements. In this paper, we tackle the probabilistic prediction of partially observable, long-memory dynamical systems, with applications to solar dynamics and the evolution of active regions. We show that standard inference schemes, such as autoregressive rollouts, fail to capture long-range dependencies in the data, largely because they do not integrate past information effectively. To overcome this, we propose a multiscale inference scheme for diffusion models, tailored to physical processes. Our method generates trajectories that are temporally fine-grained near the present and coarser as we move farther away, which enables capturing long-range temporal dependencies without increasing computational cost. When integrated into a diffusion model, we show that our inference scheme significantly reduces the bias of the predicted distributions and improves rollout stability. Rudy Morel, Francesco Pio Ramunno, Jeff Shen, Alberto Bietti, Kyunghyun Cho, Miles D. Cranmer, Siavash Golkar, Olexandr Gugnin, Géraud Krawezik, Tanya Marwah, Michael McCabe, Lucas Meyer, Payel Mukhopadhyay, Ruben Ohana, Liam Holden Parker, Helen Qu, François Rozet, K. D. Leka, François Lanusse, David F. Fouhey, Shirley Ho |
NeurIPS | 21 |
| 2025 | AION-1: Omnimodal Foundation Model for Astronomical SciencesabstractWhile foundation models have shown promise across a variety of fields, astronomy lacks a unified framework for joint modeling across its highly diverse data modalities. In this paper, we present AION-1, the first large-scale multimodal foundation family of models for astronomy. AION-1 enables arbitrary transformations between heterogeneous data types using a two-stage architecture: modality-specific tokenization followed by transformer-based masked modeling of cross-modal token sequences. Trained on over 200M astronomical objects, AION-1 demonstrates strong performance across regression, classification, generation, and object retrieval tasks. Beyond astronomy, AION-1 provides a scalable blueprint for multimodal scientific foundation models that can seamlessly integrate heterogeneous combinations of real-world observations. Our model release is entirely open source, including the dataset, training script, and weights. Liam Holden Parker, François Lanusse, Jeff Shen, Ollie Liu, Tom Hehir, Leopoldo Sarra, Lucas Meyer, Micah Bowles, Sebastian Wagner-Carena, Helen Qu, Siavash Golkar, Alberto Bietti, Hatim Bourfoune, Pierre Cornette, Keiya Hirashima, Géraud Krawezik, Ruben Ohana, Nicholas Lourie, Michael McCabe, Rudy Morel, Payel Mukhopadhyay, Mariel Pettee, Kyunghyun Cho, Miles D. Cranmer, Shirley Ho |
NeurIPS | 25 |
| 2025 | Lost in Latent Space: An Empirical Study of Latent Diffusion Models for Physics EmulationabstractThe steep computational cost of diffusion models at inference hinders their use as fast physics emulators. In the context of image and video generation, this computational drawback has been addressed by generating in the latent space of an autoencoder instead of the pixel space. In this work, we investigate whether a similar strategy can be effectively applied to the emulation of dynamical systems and at what cost. We find that the accuracy of latent-space emulation is surprisingly robust to a wide range of compression rates (up to 1000x). We also show that diffusion-based emulators are consistently more accurate than non-generative counterparts and compensate for uncertainty in their predictions with greater diversity. Finally, we cover practical design choices, spanning from architectures to optimizers, that we found critical to train latent-space emulators. François Rozet, Ruben Ohana, Michael McCabe, Gilles Louppe, François Lanusse, Shirley Ho |
NeurIPS | 6 |
| 2024 | The Multimodal Universe: Enabling Large-Scale Machine Learning with 100 TB of Astronomical Scientific DataabstractWe present the Multimodal Universe, a large-scale multimodal dataset of scientific astronomical data, compiled specifically to facilitate machine learning research. Overall, our dataset contains hundreds of millions of astronomical observations, constituting 100TB of multi-channel and hyper-spectral images, spectra, multivariate time series, as well as a wide variety of associated scientific measurements and metadata. In addition, we include a range of benchmark tasks representative of standard practices for machine learning methods in astrophysics. This massive dataset will enable the development of large multi-modal models specifically targeted towards scientific applications. All codes used to compile the dataset, and a description of how to access the data is available at https://github.com/MultimodalUniverse/MultimodalUniverse Eirini Angeloudi, Jeroen Audenaert, Micah Bowles, Benjamin M. Boyd, David Chemaly, Brian Cherinka, Ioana Ciuca, Miles D. Cranmer, Aaron Do, Matthew Grayling, Erin E. Hayes, Tom Hehir, Shirley Ho, Marc Huertas-Company, Kartheik Iyer, Maja Jablonska, François Lanusse, Kaisey Mandel, Rafael Martínez-Galarza, Peter Melchior, Lucas Meyer, Liam Holden Parker, Helen Qu, Jeff Shen, Michael J. Smith 0013, Connor Stone, Mike Walmsley, John F. Wu |
NeurIPS | 13 |
| 2024 | Multiple Physics Pretraining for Spatiotemporal Surrogate ModelsabstractWe introduce multiple physics pretraining (MPP), an autoregressive task-agnostic pretraining approach for physical surrogate modeling of spatiotemporal systems with transformers. In MPP, rather than training one model on a specific physical system, we train a backbone model to predict the dynamics of multiple heterogeneous physical systems simultaneously in order to learn features that are broadly useful across systems and facilitate transfer. In order to learn effectively in this setting, we introduce a shared embedding and normalization strategy that projects the fields of multiple systems into a shared embedding space. We validate the efficacy of our approach on both pretraining and downstream tasks over a broad fluid mechanics-oriented benchmark. We show that a single MPP-pretrained transformer is able to match or outperform task-specific baselines on all pretraining sub-tasks without the need for finetuning. For downstream tasks, we demonstrate that finetuning MPP-trained models results in more accurate predictions across multiple time-steps on systems with previously unseen physical components or higher dimensional systems compared to training from scratch or finetuning pretrained video foundation models. We open-source our code and model weights trained at multiple scales for reproducibility. Michael McCabe, Bruno Régaldo-Saint Blancard, Liam Holden Parker, Ruben Ohana, Miles D. Cranmer, Alberto Bietti, Michael Eickenberg, Siavash Golkar, Géraud Krawezik, François Lanusse, Mariel Pettee, Tiberiu Tesileanu, Kyunghyun Cho, Shirley Ho |
NeurIPS | 14 |
| 2024 | The Well: a Large-Scale Collection of Diverse Physics Simulations for Machine LearningabstractMachine learning based surrogate models offer researchers powerful tools for accelerating simulation-based workflows. However, as standard datasets in this space often cover small classes of physical behavior, it can be difficult to evaluate the efficacy of new approaches. To address this gap, we introduce the Well: a large-scale collection of datasets containing numerical simulations of a wide variety of spatiotemporal physical systems. The Well draws from domain experts and numerical software developers to provide 15TB of data across 16 datasets covering diverse domains such as biological systems, fluid dynamics, acoustic scattering, as well as magneto-hydrodynamic simulations of extra-galactic fluids or supernova explosions. These datasets can be used individually or as part of a broader benchmark suite. To facilitate usage of the Well, we provide a unified PyTorch interface for training and evaluating models. We demonstrate the function of this library by introducing example baselines that highlight the new challenges posed by the complex dynamics of the Well. The code and data is available at https://github.com/PolymathicAI/the_well. Ruben Ohana, Michael McCabe, Lucas Meyer, Rudy Morel, Fruzsina Julia Agocs, Miguel Beneitez, Marsha J. Berger, Blakesley Burkhart, Stuart B. Dalziel, Drummond B. Fielding, Daniel Fortunato, Jared A. Goldberg, Keiya Hirashima, Yan-Fei Jiang, Rich R. Kerswell, Suryanarayana Maddu, Jonah Miller, Payel Mukhopadhyay, Stefan S. Nixon, Jeff Shen, Romain Watteaux, Bruno Régaldo-Saint Blancard, François Rozet, Liam Holden Parker, Miles D. Cranmer, Shirley Ho |
NeurIPS | 26 |
| 2022 | Learned Simulators for Turbulence
Kimberly L. Stachenfeld, Drummond B. Fielding, Dmitrii Kochkov, Miles D. Cranmer, Tobias Pfaff, Jonathan Godwin, Shirley Ho, Peter W. Battaglia, Alvaro Sanchez-Gonzalez |
ICLR | 8 |
| 2020 | Discovering Symbolic Models from Deep Learning with Inductive BiasesabstractWe develop a general approach to distill symbolic representations of a learned deep model by introducing strong inductive biases. We focus on Graph Neural Networks (GNNs). The technique works as follows: we first encourage sparse latent representations when we train a GNN in a supervised setting, then we apply symbolic regression to components of the learned model to extract explicit physical relations. We find the correct known equations, including force laws and Hamiltonians, can be extracted from the neural network. We then apply our method to a non-trivial cosmology example—a detailed dark matter simulation—and discover a new analytic formula which can predict the concentration of dark matter from the mass distribution of nearby cosmic structures. The symbolic expressions extracted from the GNN using our technique also generalized to out-of-distribution-data better than the GNN itself. Our approach offers alternative directions for interpreting neural networks and discovering novel physical principles from the representations they learn. Miles D. Cranmer, Alvaro Sanchez-Gonzalez, Peter W. Battaglia, Kyle Cranmer, David N. Spergel, Shirley Ho |
NeurIPS | 7 |
| 2018 | CosmoFlow: using deep learning to learn the universe at scale
Amrita Mathuriya, Deborah Bard, Peter Mendygral, Lawrence Meadows, James Arnemann, Siyu He, Tuomas Kärnä, Diana Moise, Simon J. Pennycook, Kristyn J. Maschhoff, Jason Sewall, Nalini Kumar, Shirley Ho, Michael F. Ringenburg, Prabhat, Victor W. Lee |
SC | 14 |
| 2016 | Estimating Cosmological Parameters from the Dark Matter DistributionabstractA grand challenge of the 21st century cosmology is to accurately estimate the cosmological parameters of our Universe. A major approach in estimating the cosmological parameters is to use the large scale matter distribution of the Universe. Galaxy surveys provide the means to map out cosmic large-scale structure in three dimensions. Information about galaxy locations is typically summarized in a "single" function of scale, such as the galaxy correlation function or power-spectrum. We show that it is possible to estimate these cosmological parameters directly from the distribution of matter. This paper presents the application of deep 3D convolutional networks to volumetric representation of dark matter simulations as well as the results obtained using a recently proposed distribution regression framework, showing that machine learning techniques are comparable to, and can sometimes outperform, maximum-likelihood point estimates using "cosmological models". This opens the way to estimating the parameters of our Universe with higher accuracy. Siamak Ravanbakhsh, Junier B. Oliva, Sebastian Fromenteau, Layne Price, Shirley Ho, Jeff G. Schneider, Barnabás Póczos |
ICML | 5 |
| 2015 | Fast Function to Function RegressionabstractWe analyze the problem of regression when both input covariates and output responses are functions from a nonparametric function class. Function to function regression (FFR) covers a large range of interesting applications including time-series prediction problems, and also more general tasks like studying a mapping between two separate types of distributions. However, previous nonparametric estimators for FFR type problems scale badly computationally with the number of input/output pairs in a data-set. Given the complexity of a mapping between general functions it may be necessary to consider large data-sets in order to achieve a low estimation risk. To address this issue, we develop a novel scalable nonparametric estimator, the Triple-Basis Estimator (3BE), which is capable of operating over datasets with many instances. To the best of our knowledge, the 3BE is the first nonparametric FFR estimator that can scale to massive data-sets. We analyze the 3BE’s risk and derive an upperbound rate. Furthermore, we show an improvement of several orders of magnitude in terms of prediction speed and a reduction in error over previous estimators in various real-world data-sets. Junier B. Oliva, Willie Neiswanger, Barnabás Póczos, Eric P. Xing, Hy Trac, Shirley Ho, Jeff G. Schneider |
AISTATS | 6 |
| 2015 | Finding Galaxies in the Shadows of Quasars with Gaussian ProcessesabstractWe develop an automated technique for detecting damped Lyman-αabsorbers (DLAs) along spectroscopic sightlines to quasi-stellar objects (QSOs or quasars). The detection of DLAs in large-scale spectroscopic surveys such as SDSS–III is critical to address outstanding cosmological questions, such as the nature of galaxy formation. We use nearly 50000 QSO spectra to learn a tailored Gaussian process model for quasar emission spectra, which we apply to the DLA detection problem via Bayesian model selection. We demonstrate our method’s effectiveness with a large-scale validation experiment on over 100000 spectra, with excellent performance. Roman Garnett, Shirley Ho, Jeff G. Schneider |
ICML | 2 |
| 2015 | Optimal Ridge Detection using Coverage RiskabstractWe introduce the concept of coverage risk as an error measure for density ridge estimation.The coverage risk generalizes the mean integrated square error to set estimation.We propose two risk estimators for the coverage risk and we show that we can select tuning parameters by minimizing the estimated risk.We study the rate of convergence for coverage risk and prove consistency of the risk estimators.We apply our method to three simulated datasets and to cosmology data.In all the examples, the proposed method successfully recover the underlying density structure. Yen-Chi Chen, Christopher R. Genovese, Shirley Ho, Larry A. Wasserman |
NIPS | 3 |