VLDB 2026 Research / reviewers in the wild / expert
Chao Ma 0019
dblp:79/1552-19
· DBLP profile ↗
12ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0001-8385-0825ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 6 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
12 papers |
Probabilistic and Bayesian machine learning · 52% Generative modeling · 14% Optimization for machine learning · 9% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 28 heaviest of 32, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
memory-efficient training |
1.7 | 2 | 2025 | Gradient Multi-Normalization for Efficient LLM Training · NeurIPS 2025 SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training · ICML 2025 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
1.5 | 3 | 2022 | Missing Data Imputation and Acquisition with Deep Hierarchical Models and Hamiltonian Monte Carlo · NeurIPS 2022 Functional Variational Inference based on Stochastic Process Generators · NeurIPS 2021 Variational Implicit Processes · ICML 2019 |
Natural language and speech › Language models and text generation
large language model training |
1.1 | 2 | 2025 | SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training · ICML 2025 Gradient Multi-Normalization for Efficient LLM Training · NeurIPS 2025 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
approximate inference |
1.0 | 2 | 2022 | Missing Data Imputation and Acquisition with Deep Hierarchical Models and Hamiltonian Monte Carlo · NeurIPS 2022 Variational Implicit Processes · ICML 2019 |
Machine learning › Optimization for machine learning
stochastic gradient descent |
0.9 | 1 | 2025 | SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training · ICML 2025 |
Machine learning › Generative modeling
variational autoencoder |
0.8 | 2 | 2020 | VAEM: a Deep Generative Model for Heterogeneous Mixed Type Data · NeurIPS 2020 EDDI: Efficient Dynamic Discovery of High-Value Information with Partial VAE · ICML 2019 |
Machine learning › Generative modeling › generative model › probabilistic generative model
causal generative model |
0.8 | 1 | 2024 | A Fixed-Point Approach for Causal Generative Modeling · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning
causal inference |
0.8 | 1 | 2024 | Towards Causal Foundation Model: on Duality between Optimal Balancing and Attention · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal model |
0.8 | 1 | 2024 | A Fixed-Point Approach for Causal Generative Modeling · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal model
structural causal model |
0.8 | 1 | 2024 | A Fixed-Point Approach for Causal Generative Modeling · ICML 2024 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal effect estimation
treatment effect estimation |
0.8 | 1 | 2024 | Towards Causal Foundation Model: on Duality between Optimal Balancing and Attention · ICML 2024 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
causal reasoning |
0.7 | 1 | 2023 | Causal Reasoning in the Presence of Latent Confounders via Neural ADMG Learning · ICLR 2023 |
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
latent variable causal discovery |
0.7 | 1 | 2023 | Causal Reasoning in the Presence of Latent Confounders via Neural ADMG Learning · ICLR 2023 |
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
hamiltonian monte carlo |
0.6 | 1 | 2022 | Missing Data Imputation and Acquisition with Deep Hierarchical Models and Hamiltonian Monte Carlo · NeurIPS 2022 |
Machine learning › Generative modeling › variational autoencoder
hierarchical VAE |
0.6 | 1 | 2022 | Missing Data Imputation and Acquisition with Deep Hierarchical Models and Hamiltonian Monte Carlo · NeurIPS 2022 |
Medical and health informatics › clinical diagnosis
online disease diagnosis |
0.6 | 1 | 2022 | BSODA: A Bipartite Scalable Framework for Online Disease Diagnosis · WWW 2022 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
functional variational inference |
0.5 | 1 | 2021 | Functional Variational Inference based on Stochastic Process Generators · NeurIPS 2021 |
Machine learning › Generative modeling
generative model |
0.5 | 1 | 2021 | Identifiable Generative models for Missing Not at Random Data Imputation · NeurIPS 2021 |
Machine learning › Representation and self-supervised learning › causal representation learning
identifiability |
0.5 | 1 | 2021 | Identifiable Generative models for Missing Not at Random Data Imputation · NeurIPS 2021 |
Machine learning › Probabilistic and Bayesian machine learning › missing data
missing data imputation |
0.5 | 1 | 2021 | Identifiable Generative models for Missing Not at Random Data Imputation · NeurIPS 2021 |
Machine learning › Probabilistic and Bayesian machine learning › experimental design
bayesian experimental design |
0.4 | 1 | 2019 | EDDI: Efficient Dynamic Discovery of High-Value Information with Partial VAE · ICML 2019 |
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models › bayesian deep learning
bayesian neural networks |
0.4 | 1 | 2019 | Variational Implicit Processes · ICML 2019 |
Machine learning › Probabilistic and Bayesian machine learning › experimental design › bayesian experimental design
expected information gain |
0.4 | 1 | 2019 | EDDI: Efficient Dynamic Discovery of High-Value Information with Partial VAE · ICML 2019 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.4 | 1 | 2019 | Variational Implicit Processes · ICML 2019 |
Machine learning › Probabilistic and Bayesian machine learning
stochastic processes |
0.4 | 1 | 2019 | Variational Implicit Processes · ICML 2019 |
Machine learning › Deep learning architectures and training › attention mechanism
self-attention |
0.2 | 1 | 2024 | Towards Causal Foundation Model: on Duality between Optimal Balancing and Attention · ICML 2024 |
Machine learning › Deep learning architectures and training
transformer |
0.2 | 1 | 2024 | Towards Causal Foundation Model: on Duality between Optimal Balancing and Attention · ICML 2024 |
Data integration and cleaning › missing data
missing value imputation |
0.2 | 1 | 2022 | Missing Data Imputation and Acquisition with Deep Hierarchical Models and Hamiltonian Monte Carlo · NeurIPS 2022 |
Methods — techniques the papers use, named apart from their topics
variational autoencoder · 1.6whitening · 0.9normalization · 0.9matrix normalization · 0.9gradient multi-normalization · 0.9alternating scheme · 0.9SGD · 0.9transformer · 0.8optimal covariate balancing · 0.8attention mechanism · 0.8self-attention · 0.6product-of-experts · 0.6information-theoretic reward · 0.6hamiltonian monte carlo · 0.6automatic hyperparameter tuning · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SWAN: SGD with Normalization and Whitening Enables Stateless LLM TrainingabstractAdaptive optimizers such as Adam (Kingma & Ba, 2015) have been central to the success of large language models. However, they often require maintaining optimizer states throughout training, which can result in memory requirements several times greater than the model footprint. This overhead imposes constraints on scalability and computational efficiency. Stochastic Gradient Descent (SGD), in contrast, is a stateless optimizer, as it does not track state variables during training. Consequently, it achieves optimal memory efficiency. However, its capability in LLM training is limited (Zhao et al., 2024b) In this work, we show that pre-processing SGD using normalization and whitening in a stateless manner can achieve similar performance as Adam for LLM training, while maintaining the same memory footprint of SGD. Specifically, we show that normalization stabilizes gradient distributions, and whitening counteracts the local curvature of the loss landscape. This results in SWAN (SGD with Whitening And Normalization), a stochastic optimizer that eliminates the need to store any optimizer states. Empirically, SWAN achieves 50% reduction on total end-to-end memory compared to Adam. Under the memory-efficienct LLaMA training benchmark of (Zhao et al., 2024a), SWAN reaches the same evaluation perplexity using half as many tokens for 350M and 1.3B model. Chao Ma 0019, Wenbo Gong 0001, Meyer Scetbon, Edward Meeds |
ICML | 1 |
| 2025 | Gradient Multi-Normalization for Efficient LLM TrainingabstractTraining large language models (LLMs) commonly relies on adaptive optimizers such as Adam (Kingma & Ba 2015), which accelerate convergence through moment estimates but incur substantial memory overhead. Recent stateless approaches such as SWAN (Ma et al., 2024) have shown that appropriate preprocessing of instantaneous gradient matrices can match the performance of adaptive methods without storing optimizer states. Building on this insight, we introduce \emph{gradient multi-normalization}, a principled framework for designing stateless optimizers that normalize gradients with respect to multiple norms simultaneously. Whereas standard first-order methods can be viewed as gradient normalization under a single norm (Bernstein & Newhouse, 2024), our formulation generalizes this perspective to a multi-norm setting. We derive an efficient alternating scheme that enforces these normalization constraints and show that our procedure can produce, up to an arbitrary precision, a fixed-point of the problem. This unifies and extends prior stateless optimizers, showing that SWAN arises as a specific instance with particular norm choices. Leveraging this principle, we develop SinkGD, a lightweight matrix optimizer that retains the memory footprint of SGD while substantially reducing computation relative to whitening-based methods. On the memory-efficient LLaMA training benchmark (Zhao et al., 2024), SinkGD achieves state-of-the-art performance, reaching the same evaluation perplexity as Adam using only 40\% of the training tokens. Meyer Scetbon, Chao Ma 0019, Wenbo Gong 0001, Edward Meeds |
NeurIPS | 2 |
| 2024 | A Fixed-Point Approach for Causal Generative ModelingabstractWe propose a novel formalism for describing Structural Causal Models (SCMs) as fixed-point problems on causally ordered variables, eliminating the need for Directed Acyclic Graphs (DAGs), and establish the weakest known conditions for their unique recovery given the topological ordering (TO). Based on this, we design a two-stage causal generative model that first infers in a zero-shot manner a valid TO from observations, and then learns the generative SCM on the ordered variables. To infer TOs, we propose to amortize the learning of TOs on synthetically generated datasets by sequentially predicting the leaves of graphs seen during training. To learn SCMs, we design a transformer-based architecture that exploits a new attention mechanism enabling the modeling of causal structures, and show that this parameterization is consistent with our formalism. Finally, we conduct an extensive evaluation of each method individually, and show that when combined, our model outperforms various baselines on generated out-of-distribution problems. Meyer Scetbon, Joel Jennings, Agrin Hilmkil, Cheng Zhang 0005, Chao Ma 0019 |
ICML | 5 |
| 2024 | Towards Causal Foundation Model: on Duality between Optimal Balancing and AttentionabstractFoundation models have brought changes to the landscape of machine learning, demonstrating sparks of human-level intelligence across a diverse array of tasks. However, a gap persists in complex tasks such as causal inference, primarily due to challenges associated with intricate reasoning steps and high numerical precision requirements. In this work, we take a first step towards building causally-aware foundation models for treatment effect estimations. We propose a novel, theoretically justified method called Causal Inference with Attention (CInA), which utilizes multiple unlabeled datasets to perform self-supervised causal learning, and subsequently enables zero-shot causal inference on unseen tasks with new data. This is based on our theoretical results that demonstrate the primal-dual connection between optimal covariate balancing and self-attention, facilitating zero-shot causal inference through the final layer of a trained transformer-type architecture. We demonstrate empirically that CInA effectively generalizes to out-of-distribution datasets and various real-world datasets, matching or even surpassing traditional per-dataset methodologies. These results provide compelling evidence that our method has the potential to serve as a stepping stone for the development of causal foundation models. Jiaqi Zhang 0006, Joel Jennings, Agrin Hilmkil, Nick Pawlowski, Cheng Zhang 0005, Chao Ma 0019 |
ICML | 6 |
| 2023 | Causal Reasoning in the Presence of Latent Confounders via Neural ADMG Learning
Matthew Ashman, Chao Ma 0019, Agrin Hilmkil, Joel Jennings, Cheng Zhang 0005 |
ICLR | 2 |
| 2022 | Missing Data Imputation and Acquisition with Deep Hierarchical Models and Hamiltonian Monte CarloabstractVariational Autoencoders (VAEs) have recently been highly successful at imputing and acquiring heterogeneous missing data. However, within this specific application domain, existing VAE methods are restricted by using only one layer of latent variables and strictly Gaussian posterior approximations. To address these limitations, we present HH-VAEM, a Hierarchical VAE model for mixed-type incomplete data that uses Hamiltonian Monte Carlo with automatic hyper-parameter tuning for improved approximate inference. Our experiments show that HH-VAEM outperforms existing baselines in the tasks of missing data imputation and supervised learning with missing features. Finally, we also present a sampling-based approach for efficiently computing the information gain when missing features are to be acquired with HH-VAEM. Our experiments show that this sampling-based approach is superior to alternatives based on Gaussian approximations. Ignacio Peis, Chao Ma 0019, José Miguel Hernández-Lobato |
NeurIPS | 2 |
| 2022 | BSODA: A Bipartite Scalable Framework for Online Disease DiagnosisabstractA growing number of people are seeking healthcare advice online. Usually, they diagnose their medical conditions based on the symptoms they are experiencing, which is also known as self-diagnosis. From the machine learning perspective, online disease diagnosis is a sequential feature (symptom) selection and classification problem. Reinforcement learning (RL) methods are the standard approaches to this type of tasks. Generally, they perform well when the feature space is small, but frequently become inefficient in tasks with a large number of features, such as the self-diagnosis. To address the challenge, we propose a non-RL Bipartite Scalable framework for Online Disease diAgnosis, called BSODA. BSODA is composed of two cooperative branches that handle symptom-inquiry and disease-diagnosis, respectively. The inquiry branch determines which symptom to collect next by an information-theoretic reward. We employ a Product-of-Experts encoder to significantly improve the handling of partial observations of a large number of features. Besides, we propose several approximation methods to substantially reduce the computational cost of the reward to a level that is acceptable for online services. Additionally, we leverage the diagnosis model to estimate the reward more precisely. For the diagnosis branch, we use a knowledge-guided self-attention model to perform predictions. In particular, BSODA determines when to stop inquiry and output predictions using both the inquiry and diagnosis models. We demonstrate that BSODA outperforms the state-of-the-art methods on several public datasets. Moreover, we propose a novel evaluation method to test the transferability of symptom checking methods from synthetic to real-world tasks. Compared to existing RL baselines, BSODA is more effectively scalable to large search spaces. Xiaohao Mao, Chao Ma 0019, José Miguel Hernández-Lobato, Ting Chen 0006 |
WWW | 3 |
| 2021 | Functional Variational Inference based on Stochastic Process GeneratorsabstractBayesian inference in the space of functions has been an important topic for Bayesian modeling in the past. In this paper, we propose a new solution to this problem called Functional Variational Inference (FVI). In FVI, we minimize a divergence in function space between the variational distribution and the posterior process. This is done by using as functional variational family a new class of flexible distributions called Stochastic Process Generators (SPGs), which are cleverly designed so that the functional ELBO can be estimated efficiently using analytic solutions and mini-batch sampling. FVI can be applied to stochastic process priors when random function samples from those priors are available. Our experiments show that FVI consistently outperforms weight-space and function space VI methods on several tasks, which validates the effectiveness of our approach. Chao Ma 0019, José Miguel Hernández-Lobato |
NeurIPS | 1 |
| 2021 | Identifiable Generative models for Missing Not at Random Data ImputationabstractReal-world datasets often have missing values associated with complex generative processes, where the cause of the missingness may not be fully observed. This is known as missing not at random (MNAR) data. However, many imputation methods do not take into account the missingness mechanism, resulting in biased imputation values when MNAR data is present. Although there are a few methods that have considered the MNAR scenario, their model's identifiability under MNAR is generally not guaranteed. That is, model parameters can not be uniquely determined even with infinite data samples, hence the imputation results given by such models can still be biased. This issue is especially overlooked by many modern deep generative models. In this work, we fill in this gap by systematically analyzing the identifiability of generative models under MNAR. Furthermore, we propose a practical deep generative model which can provide identifiability guarantees under mild assumptions, for a wide range of MNAR mechanisms. Our method demonstrates a clear advantage for tasks on both synthetic data and multiple real-world scenarios with MNAR data. Chao Ma 0019, Cheng Zhang 0005 |
NeurIPS | 1 |
| 2020 | VAEM: a Deep Generative Model for Heterogeneous Mixed Type DataabstractDeep generative models often perform poorly in real-world applications due to the heterogeneity of natural data sets. Heterogeneity arises from data containing different types of features (categorical, ordinal, continuous, etc.) and features of the same type having different marginal distributions. We propose an extension of variational autoencoders (VAEs) called VAEM to handle such heterogeneous data. VAEM is a deep generative model that is trained in a two stage manner, such that the first stage provides a more uniform representation of the data to the second stage, thereby sidestepping the problems caused by heterogeneous data. We provide extensions of VAEM to handle partially observed data, and demonstrate its performance in data generation, missing data prediction and sequential feature selection tasks. Our results show that VAEM broadens the range of real-world applications where deep generative models can be successfully deployed. Chao Ma 0019, Sebastian Tschiatschek, Richard E. Turner, José Miguel Hernández-Lobato, Cheng Zhang 0005 |
NeurIPS | 1 |
| 2019 | Variational Implicit ProcessesabstractWe introduce the implicit processes (IPs), a stochastic process that places implicitly defined multivariate distributions over any finite collections of random variables. IPs are therefore highly flexible implicit priors over functions, with examples including data simulators, Bayesian neural networks and non-linear transformations of stochastic processes. A novel and efficient approximate inference algorithm for IPs, namely the variational implicit processes (VIPs), is derived using generalised wake-sleep updates. This method returns simple update equations and allows scalable hyper-parameter learning with stochastic optimization. Experiments show that VIPs return better uncertainty estimates and lower errors over existing inference methods for challenging models such as Bayesian neural networks, and Gaussian processes. Chao Ma 0019, Yingzhen Li, José Miguel Hernández-Lobato |
ICML | 1 |
| 2019 | EDDI: Efficient Dynamic Discovery of High-Value Information with Partial VAEabstractMany real-life decision making situations allow further relevant information to be acquired at a specific cost, for example, in assessing the health status of a patient we may decide to take additional measurements such as diagnostic tests or imaging scans before making a final assessment. Acquiring more relevant information enables better decision making, but may be costly. How can we trade off the desire to make good decisions by acquiring further information with the cost of performing that acquisition? To this end, we propose a principled framework, named EDDI (Efficient Dynamic Discovery of high-value Information), based on the theory of Bayesian experimental design. In EDDI, we propose a novel partial variational autoencoder (Partial VAE) to predict missing data entries problematically given any subset of the observed ones, and combine it with an acquisition function that maximizes expected information gain on a set of target variables. We show cost reduction at the same decision quality and improved decision quality at the same cost in multiple machine learning benchmarks and two real-world health-care applications. Chao Ma 0019, Sebastian Tschiatschek, Konstantina Palla, José Miguel Hernández-Lobato, Sebastian Nowozin, Cheng Zhang 0005 |
ICML | 1 |