VLDB 2026 Research / reviewers in the wild / expert
Helen Qu
dblp:317/0339
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Representation and self-supervised learning · 23% Generative modeling · 22% Trustworthy machine learning · 20% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Computational science and engineering · 100% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › diffusion model
conditional diffusion model |
0.9 | 1 | 2025 | Predicting partially observable dynamical systems via diffusion models with a multiscale inference scheme · NeurIPS 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Predicting partially observable dynamical systems via diffusion models with a multiscale inference scheme · NeurIPS 2025 |
Machine learning › Deep learning architectures and training
foundation model |
0.9 | 1 | 2025 | AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | Sum-of-Parts: Self-Attributing Neural Networks with End-to-End Learning of Feature Groups · ICML 2025 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked modeling |
0.9 | 1 | 2025 | AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.8 | 1 | 2024 | Connect Later: Improving Fine-tuning for Robustness with Targeted Augmentations · ICML 2024 |
Machine learning › Trustworthy machine learning
out-of-distribution generalization |
0.8 | 1 | 2024 | Connect Later: Improving Fine-tuning for Robustness with Targeted Augmentations · ICML 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning |
0.8 | 1 | 2024 | Connect Later: Improving Fine-tuning for Robustness with Targeted Augmentations · ICML 2024 |
Computational science and engineering › astronomy
astrophysics |
0.8 | 1 | 2024 | The Multimodal Universe: Enabling Large-Scale Machine Learning with 100 TB of Astronomical Scientific Data · NeurIPS 2024 |
Computer vision › Image recognition and object detection
multi-scale inference |
0.3 | 1 | 2025 | Predicting partially observable dynamical systems via diffusion models with a multiscale inference scheme · NeurIPS 2025 |
Computational science and engineering › astronomy
astronomical data analysis |
0.3 | 1 | 2025 | AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025 |
Computational science and engineering
astronomy |
0.3 | 1 | 2025 | AION-1: Omnimodal Foundation Model for Astronomical Sciences · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning
multimodal representation learning |
0.2 | 1 | 2024 | The Multimodal Universe: Enabling Large-Scale Machine Learning with 100 TB of Astronomical Scientific Data · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
transformer · 1.7tokenization · 1.7multiscale inference scheme · 1.7masked modeling · 1.7autoregressive rollout · 1.7sum-of-parts · 0.9group-based attribution · 0.9end-to-end learning · 0.9multimodal machine learning · 0.8fine-tuning · 0.8contrastive learning · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Sum-of-Parts: Self-Attributing Neural Networks with End-to-End Learning of Feature GroupsabstractSelf-attributing neural networks (SANNs) present a potential path towards interpretable models for high-dimensional problems, but often face significant trade-offs in performance. In this work, we formally prove a lower bound on errors of per-feature SANNs, whereas group-based SANNs can achieve zero error and thus high performance. Motivated by these insights, we propose Sum-of-Parts (SOP), a framework that transforms any differentiable model into a group-based SANN, where feature groups are learned end-to-end without group supervision. SOP achieves state-of-the-art performance for SANNs on vision and language tasks, and we validate that the groups are interpretable on a range of quantitative and semantic metrics. We further validate the utility of SOP explanations in model debugging and cosmological scientific discovery. Weiqiu You, Helen Qu, Marco Gatti, Bhuvnesh Jain, Eric Wong 0001 |
ICML | 2 |
| 2025 | Predicting partially observable dynamical systems via diffusion models with a multiscale inference schemeabstractConditional diffusion models provide a natural framework for probabilistic prediction of dynamical systems and have been successfully applied to fluid dynamics and weather prediction. However, in many settings, the available information at a given time represents only a small fraction of what is needed to predict future states, either due to measurement uncertainty or because only a small fraction of the state can be observed. This is true for example in solar physics, where we can observe the Sun’s surface and atmosphere, but its evolution is driven by internal processes for which we lack direct measurements. In this paper, we tackle the probabilistic prediction of partially observable, long-memory dynamical systems, with applications to solar dynamics and the evolution of active regions. We show that standard inference schemes, such as autoregressive rollouts, fail to capture long-range dependencies in the data, largely because they do not integrate past information effectively. To overcome this, we propose a multiscale inference scheme for diffusion models, tailored to physical processes. Our method generates trajectories that are temporally fine-grained near the present and coarser as we move farther away, which enables capturing long-range temporal dependencies without increasing computational cost. When integrated into a diffusion model, we show that our inference scheme significantly reduces the bias of the predicted distributions and improves rollout stability. Rudy Morel, Francesco Pio Ramunno, Jeff Shen, Alberto Bietti, Kyunghyun Cho, Miles D. Cranmer, Siavash Golkar, Olexandr Gugnin, Géraud Krawezik, Tanya Marwah, Michael McCabe, Lucas Meyer, Payel Mukhopadhyay, Ruben Ohana, Liam Holden Parker, Helen Qu, François Rozet, K. D. Leka, François Lanusse, David F. Fouhey, Shirley Ho |
NeurIPS | 16 |
| 2025 | AION-1: Omnimodal Foundation Model for Astronomical SciencesabstractWhile foundation models have shown promise across a variety of fields, astronomy lacks a unified framework for joint modeling across its highly diverse data modalities. In this paper, we present AION-1, the first large-scale multimodal foundation family of models for astronomy. AION-1 enables arbitrary transformations between heterogeneous data types using a two-stage architecture: modality-specific tokenization followed by transformer-based masked modeling of cross-modal token sequences. Trained on over 200M astronomical objects, AION-1 demonstrates strong performance across regression, classification, generation, and object retrieval tasks. Beyond astronomy, AION-1 provides a scalable blueprint for multimodal scientific foundation models that can seamlessly integrate heterogeneous combinations of real-world observations. Our model release is entirely open source, including the dataset, training script, and weights. Liam Holden Parker, François Lanusse, Jeff Shen, Ollie Liu, Tom Hehir, Leopoldo Sarra, Lucas Meyer, Micah Bowles, Sebastian Wagner-Carena, Helen Qu, Siavash Golkar, Alberto Bietti, Hatim Bourfoune, Pierre Cornette, Keiya Hirashima, Géraud Krawezik, Ruben Ohana, Nicholas Lourie, Michael McCabe, Rudy Morel, Payel Mukhopadhyay, Mariel Pettee, Kyunghyun Cho, Miles D. Cranmer, Shirley Ho |
NeurIPS | 10 |
| 2024 | Connect Later: Improving Fine-tuning for Robustness with Targeted AugmentationsabstractModels trained on a labeled source domain often generalize poorly when deployed on an out-of-distribution (OOD) target domain. In the domain adaptation setting where unlabeled target data is available, self-supervised pretraining (e.g., contrastive learning or masked autoencoding) is a promising method to mitigate this performance drop. Pretraining depends on generic data augmentations (e.g., cropping or masking) to learn representations that generalize across domains, which may not work for all distribution shifts. In this paper, we show on real-world tasks that standard fine-tuning after pretraining does not consistently improve OOD error over simply training from scratch on labeled source data. To better leverage pretraining for distribution shifts, we propose the Connect Later framework, which fine-tunes the model with targeted augmentations designed with knowledge of the shift. Intuitively, pretraining learns good representations within the source and target domains, while fine-tuning with targeted augmentations improves generalization across domains. Connect Later achieves state-of-the-art OOD accuracy while maintaining comparable or better in-distribution accuracy on 4 real-world tasks in wildlife identification (iWildCam-WILDS), tumor detection (Camelyon17-WILDS), and astronomy (AstroClassification, Redshifts). Helen Qu, Sang Michael Xie |
ICML | 1 |
| 2024 | The Multimodal Universe: Enabling Large-Scale Machine Learning with 100 TB of Astronomical Scientific DataabstractWe present the Multimodal Universe, a large-scale multimodal dataset of scientific astronomical data, compiled specifically to facilitate machine learning research. Overall, our dataset contains hundreds of millions of astronomical observations, constituting 100TB of multi-channel and hyper-spectral images, spectra, multivariate time series, as well as a wide variety of associated scientific measurements and metadata. In addition, we include a range of benchmark tasks representative of standard practices for machine learning methods in astrophysics. This massive dataset will enable the development of large multi-modal models specifically targeted towards scientific applications. All codes used to compile the dataset, and a description of how to access the data is available at https://github.com/MultimodalUniverse/MultimodalUniverse Eirini Angeloudi, Jeroen Audenaert, Micah Bowles, Benjamin M. Boyd, David Chemaly, Brian Cherinka, Ioana Ciuca, Miles D. Cranmer, Aaron Do, Matthew Grayling, Erin E. Hayes, Tom Hehir, Shirley Ho, Marc Huertas-Company, Kartheik Iyer, Maja Jablonska, François Lanusse, Kaisey Mandel, Rafael Martínez-Galarza, Peter Melchior, Lucas Meyer, Liam Holden Parker, Helen Qu, Jeff Shen, Michael J. Smith 0013, Connor Stone, Mike Walmsley, John F. Wu |
NeurIPS | 24 |