Chao Ma 0019

dblp:79/1552-19 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0001-8385-0825ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
12 papers
Probabilistic and Bayesian machine learning · 52% Generative modeling · 14% Optimization for machine learning · 9%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 28 heaviest of 32, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
memory-efficient training
1.722025
Gradient Multi-Normalization for Efficient LLM Training · NeurIPS 2025
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
1.532022
Missing Data Imputation and Acquisition with Deep Hierarchical Models and Hamiltonian Monte Carlo · NeurIPS 2022
Functional Variational Inference based on Stochastic Process Generators · NeurIPS 2021
Variational Implicit Processes · ICML 2019
Natural language and speech › Language models and text generation
large language model training
1.122025
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training · ICML 2025
Gradient Multi-Normalization for Efficient LLM Training · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
approximate inference
1.022022
Missing Data Imputation and Acquisition with Deep Hierarchical Models and Hamiltonian Monte Carlo · NeurIPS 2022
Variational Implicit Processes · ICML 2019
Machine learning › Optimization for machine learning
stochastic gradient descent
0.912025
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training · ICML 2025
Machine learning › Generative modeling
variational autoencoder
0.822020
VAEM: a Deep Generative Model for Heterogeneous Mixed Type Data · NeurIPS 2020
EDDI: Efficient Dynamic Discovery of High-Value Information with Partial VAE · ICML 2019
Machine learning › Generative modeling › generative model › probabilistic generative model
causal generative model
0.812024
A Fixed-Point Approach for Causal Generative Modeling · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.812024
Towards Causal Foundation Model: on Duality between Optimal Balancing and Attention · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal model
0.812024
A Fixed-Point Approach for Causal Generative Modeling · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal model
structural causal model
0.812024
A Fixed-Point Approach for Causal Generative Modeling · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal effect estimation
treatment effect estimation
0.812024
Towards Causal Foundation Model: on Duality between Optimal Balancing and Attention · ICML 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning
causal reasoning
0.712023
Causal Reasoning in the Presence of Latent Confounders via Neural ADMG Learning · ICLR 2023
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal discovery
latent variable causal discovery
0.712023
Causal Reasoning in the Presence of Latent Confounders via Neural ADMG Learning · ICLR 2023
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
hamiltonian monte carlo
0.612022
Missing Data Imputation and Acquisition with Deep Hierarchical Models and Hamiltonian Monte Carlo · NeurIPS 2022
Machine learning › Generative modeling › variational autoencoder
hierarchical VAE
0.612022
Missing Data Imputation and Acquisition with Deep Hierarchical Models and Hamiltonian Monte Carlo · NeurIPS 2022
Medical and health informatics › clinical diagnosis
online disease diagnosis
0.612022
BSODA: A Bipartite Scalable Framework for Online Disease Diagnosis · WWW 2022
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference
functional variational inference
0.512021
Functional Variational Inference based on Stochastic Process Generators · NeurIPS 2021
Machine learning › Generative modeling
generative model
0.512021
Identifiable Generative models for Missing Not at Random Data Imputation · NeurIPS 2021
Machine learning › Representation and self-supervised learning › causal representation learning
identifiability
0.512021
Identifiable Generative models for Missing Not at Random Data Imputation · NeurIPS 2021
Machine learning › Probabilistic and Bayesian machine learning › missing data
missing data imputation
0.512021
Identifiable Generative models for Missing Not at Random Data Imputation · NeurIPS 2021
Machine learning › Probabilistic and Bayesian machine learning › experimental design
bayesian experimental design
0.412019
EDDI: Efficient Dynamic Discovery of High-Value Information with Partial VAE · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning › deep probabilistic models › bayesian deep learning
bayesian neural networks
0.412019
Variational Implicit Processes · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning › experimental design › bayesian experimental design
expected information gain
0.412019
EDDI: Efficient Dynamic Discovery of High-Value Information with Partial VAE · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.412019
Variational Implicit Processes · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning
stochastic processes
0.412019
Variational Implicit Processes · ICML 2019
Machine learning › Deep learning architectures and training › attention mechanism
self-attention
0.212024
Towards Causal Foundation Model: on Duality between Optimal Balancing and Attention · ICML 2024
Machine learning › Deep learning architectures and training
transformer
0.212024
Towards Causal Foundation Model: on Duality between Optimal Balancing and Attention · ICML 2024
Data integration and cleaning › missing data
missing value imputation
0.212022
Missing Data Imputation and Acquisition with Deep Hierarchical Models and Hamiltonian Monte Carlo · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

variational autoencoder · 1.6whitening · 0.9normalization · 0.9matrix normalization · 0.9gradient multi-normalization · 0.9alternating scheme · 0.9SGD · 0.9transformer · 0.8optimal covariate balancing · 0.8attention mechanism · 0.8self-attention · 0.6product-of-experts · 0.6information-theoretic reward · 0.6hamiltonian monte carlo · 0.6automatic hyperparameter tuning · 0.6
YearPublicationVenuePosition
2025 SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training
abstract
Adaptive optimizers such as Adam (Kingma & Ba, 2015) have been central to the success of large language models. However, they often require maintaining optimizer states throughout training, which can result in memory requirements several times greater than the model footprint. This overhead imposes constraints on scalability and computational efficiency. Stochastic Gradient Descent (SGD), in contrast, is a stateless optimizer, as it does not track state variables during training. Consequently, it achieves optimal memory efficiency. However, its capability in LLM training is limited (Zhao et al., 2024b) In this work, we show that pre-processing SGD using normalization and whitening in a stateless manner can achieve similar performance as Adam for LLM training, while maintaining the same memory footprint of SGD. Specifically, we show that normalization stabilizes gradient distributions, and whitening counteracts the local curvature of the loss landscape. This results in SWAN (SGD with Whitening And Normalization), a stochastic optimizer that eliminates the need to store any optimizer states. Empirically, SWAN achieves 50% reduction on total end-to-end memory compared to Adam. Under the memory-efficienct LLaMA training benchmark of (Zhao et al., 2024a), SWAN reaches the same evaluation perplexity using half as many tokens for 350M and 1.3B model.
Chao Ma 0019, Wenbo Gong 0001, Meyer Scetbon, Edward Meeds
ICML1
2025 Gradient Multi-Normalization for Efficient LLM Training
abstract
Training large language models (LLMs) commonly relies on adaptive optimizers such as Adam (Kingma & Ba 2015), which accelerate convergence through moment estimates but incur substantial memory overhead. Recent stateless approaches such as SWAN (Ma et al., 2024) have shown that appropriate preprocessing of instantaneous gradient matrices can match the performance of adaptive methods without storing optimizer states. Building on this insight, we introduce \emph{gradient multi-normalization}, a principled framework for designing stateless optimizers that normalize gradients with respect to multiple norms simultaneously. Whereas standard first-order methods can be viewed as gradient normalization under a single norm (Bernstein & Newhouse, 2024), our formulation generalizes this perspective to a multi-norm setting. We derive an efficient alternating scheme that enforces these normalization constraints and show that our procedure can produce, up to an arbitrary precision, a fixed-point of the problem. This unifies and extends prior stateless optimizers, showing that SWAN arises as a specific instance with particular norm choices. Leveraging this principle, we develop SinkGD, a lightweight matrix optimizer that retains the memory footprint of SGD while substantially reducing computation relative to whitening-based methods. On the memory-efficient LLaMA training benchmark (Zhao et al., 2024), SinkGD achieves state-of-the-art performance, reaching the same evaluation perplexity as Adam using only 40\% of the training tokens.
Meyer Scetbon, Chao Ma 0019, Wenbo Gong 0001, Edward Meeds
NeurIPS2
2024 A Fixed-Point Approach for Causal Generative Modeling
abstract
We propose a novel formalism for describing Structural Causal Models (SCMs) as fixed-point problems on causally ordered variables, eliminating the need for Directed Acyclic Graphs (DAGs), and establish the weakest known conditions for their unique recovery given the topological ordering (TO). Based on this, we design a two-stage causal generative model that first infers in a zero-shot manner a valid TO from observations, and then learns the generative SCM on the ordered variables. To infer TOs, we propose to amortize the learning of TOs on synthetically generated datasets by sequentially predicting the leaves of graphs seen during training. To learn SCMs, we design a transformer-based architecture that exploits a new attention mechanism enabling the modeling of causal structures, and show that this parameterization is consistent with our formalism. Finally, we conduct an extensive evaluation of each method individually, and show that when combined, our model outperforms various baselines on generated out-of-distribution problems.
Meyer Scetbon, Joel Jennings, Agrin Hilmkil, Cheng Zhang 0005, Chao Ma 0019
ICML5
2024 Towards Causal Foundation Model: on Duality between Optimal Balancing and Attention
abstract
Foundation models have brought changes to the landscape of machine learning, demonstrating sparks of human-level intelligence across a diverse array of tasks. However, a gap persists in complex tasks such as causal inference, primarily due to challenges associated with intricate reasoning steps and high numerical precision requirements. In this work, we take a first step towards building causally-aware foundation models for treatment effect estimations. We propose a novel, theoretically justified method called Causal Inference with Attention (CInA), which utilizes multiple unlabeled datasets to perform self-supervised causal learning, and subsequently enables zero-shot causal inference on unseen tasks with new data. This is based on our theoretical results that demonstrate the primal-dual connection between optimal covariate balancing and self-attention, facilitating zero-shot causal inference through the final layer of a trained transformer-type architecture. We demonstrate empirically that CInA effectively generalizes to out-of-distribution datasets and various real-world datasets, matching or even surpassing traditional per-dataset methodologies. These results provide compelling evidence that our method has the potential to serve as a stepping stone for the development of causal foundation models.
Jiaqi Zhang 0006, Joel Jennings, Agrin Hilmkil, Nick Pawlowski, Cheng Zhang 0005, Chao Ma 0019
ICML6
2023 Causal Reasoning in the Presence of Latent Confounders via Neural ADMG Learning
Matthew Ashman, Chao Ma 0019, Agrin Hilmkil, Joel Jennings, Cheng Zhang 0005
ICLR2
2022 Missing Data Imputation and Acquisition with Deep Hierarchical Models and Hamiltonian Monte Carlo
abstract
Variational Autoencoders (VAEs) have recently been highly successful at imputing and acquiring heterogeneous missing data. However, within this specific application domain, existing VAE methods are restricted by using only one layer of latent variables and strictly Gaussian posterior approximations. To address these limitations, we present HH-VAEM, a Hierarchical VAE model for mixed-type incomplete data that uses Hamiltonian Monte Carlo with automatic hyper-parameter tuning for improved approximate inference. Our experiments show that HH-VAEM outperforms existing baselines in the tasks of missing data imputation and supervised learning with missing features. Finally, we also present a sampling-based approach for efficiently computing the information gain when missing features are to be acquired with HH-VAEM. Our experiments show that this sampling-based approach is superior to alternatives based on Gaussian approximations.
Ignacio Peis, Chao Ma 0019, José Miguel Hernández-Lobato
NeurIPS2
2022 BSODA: A Bipartite Scalable Framework for Online Disease Diagnosis
abstract
A growing number of people are seeking healthcare advice online. Usually, they diagnose their medical conditions based on the symptoms they are experiencing, which is also known as self-diagnosis. From the machine learning perspective, online disease diagnosis is a sequential feature (symptom) selection and classification problem. Reinforcement learning (RL) methods are the standard approaches to this type of tasks. Generally, they perform well when the feature space is small, but frequently become inefficient in tasks with a large number of features, such as the self-diagnosis. To address the challenge, we propose a non-RL Bipartite Scalable framework for Online Disease diAgnosis, called BSODA. BSODA is composed of two cooperative branches that handle symptom-inquiry and disease-diagnosis, respectively. The inquiry branch determines which symptom to collect next by an information-theoretic reward. We employ a Product-of-Experts encoder to significantly improve the handling of partial observations of a large number of features. Besides, we propose several approximation methods to substantially reduce the computational cost of the reward to a level that is acceptable for online services. Additionally, we leverage the diagnosis model to estimate the reward more precisely. For the diagnosis branch, we use a knowledge-guided self-attention model to perform predictions. In particular, BSODA determines when to stop inquiry and output predictions using both the inquiry and diagnosis models. We demonstrate that BSODA outperforms the state-of-the-art methods on several public datasets. Moreover, we propose a novel evaluation method to test the transferability of symptom checking methods from synthetic to real-world tasks. Compared to existing RL baselines, BSODA is more effectively scalable to large search spaces.
Xiaohao Mao, Chao Ma 0019, José Miguel Hernández-Lobato, Ting Chen 0006
WWW3
2021 Functional Variational Inference based on Stochastic Process Generators
abstract
Bayesian inference in the space of functions has been an important topic for Bayesian modeling in the past. In this paper, we propose a new solution to this problem called Functional Variational Inference (FVI). In FVI, we minimize a divergence in function space between the variational distribution and the posterior process. This is done by using as functional variational family a new class of flexible distributions called Stochastic Process Generators (SPGs), which are cleverly designed so that the functional ELBO can be estimated efficiently using analytic solutions and mini-batch sampling. FVI can be applied to stochastic process priors when random function samples from those priors are available. Our experiments show that FVI consistently outperforms weight-space and function space VI methods on several tasks, which validates the effectiveness of our approach.
Chao Ma 0019, José Miguel Hernández-Lobato
NeurIPS1
2021 Identifiable Generative models for Missing Not at Random Data Imputation
abstract
Real-world datasets often have missing values associated with complex generative processes, where the cause of the missingness may not be fully observed. This is known as missing not at random (MNAR) data. However, many imputation methods do not take into account the missingness mechanism, resulting in biased imputation values when MNAR data is present. Although there are a few methods that have considered the MNAR scenario, their model's identifiability under MNAR is generally not guaranteed. That is, model parameters can not be uniquely determined even with infinite data samples, hence the imputation results given by such models can still be biased. This issue is especially overlooked by many modern deep generative models. In this work, we fill in this gap by systematically analyzing the identifiability of generative models under MNAR. Furthermore, we propose a practical deep generative model which can provide identifiability guarantees under mild assumptions, for a wide range of MNAR mechanisms. Our method demonstrates a clear advantage for tasks on both synthetic data and multiple real-world scenarios with MNAR data.
Chao Ma 0019, Cheng Zhang 0005
NeurIPS1
2020 VAEM: a Deep Generative Model for Heterogeneous Mixed Type Data
abstract
Deep generative models often perform poorly in real-world applications due to the heterogeneity of natural data sets. Heterogeneity arises from data containing different types of features (categorical, ordinal, continuous, etc.) and features of the same type having different marginal distributions. We propose an extension of variational autoencoders (VAEs) called VAEM to handle such heterogeneous data. VAEM is a deep generative model that is trained in a two stage manner, such that the first stage provides a more uniform representation of the data to the second stage, thereby sidestepping the problems caused by heterogeneous data. We provide extensions of VAEM to handle partially observed data, and demonstrate its performance in data generation, missing data prediction and sequential feature selection tasks. Our results show that VAEM broadens the range of real-world applications where deep generative models can be successfully deployed.
Chao Ma 0019, Sebastian Tschiatschek, Richard E. Turner, José Miguel Hernández-Lobato, Cheng Zhang 0005
NeurIPS1
2019 Variational Implicit Processes
abstract
We introduce the implicit processes (IPs), a stochastic process that places implicitly defined multivariate distributions over any finite collections of random variables. IPs are therefore highly flexible implicit priors over functions, with examples including data simulators, Bayesian neural networks and non-linear transformations of stochastic processes. A novel and efficient approximate inference algorithm for IPs, namely the variational implicit processes (VIPs), is derived using generalised wake-sleep updates. This method returns simple update equations and allows scalable hyper-parameter learning with stochastic optimization. Experiments show that VIPs return better uncertainty estimates and lower errors over existing inference methods for challenging models such as Bayesian neural networks, and Gaussian processes.
Chao Ma 0019, Yingzhen Li, José Miguel Hernández-Lobato
ICML1
2019 EDDI: Efficient Dynamic Discovery of High-Value Information with Partial VAE
abstract
Many real-life decision making situations allow further relevant information to be acquired at a specific cost, for example, in assessing the health status of a patient we may decide to take additional measurements such as diagnostic tests or imaging scans before making a final assessment. Acquiring more relevant information enables better decision making, but may be costly. How can we trade off the desire to make good decisions by acquiring further information with the cost of performing that acquisition? To this end, we propose a principled framework, named EDDI (Efficient Dynamic Discovery of high-value Information), based on the theory of Bayesian experimental design. In EDDI, we propose a novel partial variational autoencoder (Partial VAE) to predict missing data entries problematically given any subset of the observed ones, and combine it with an acquisition function that maximizes expected information gain on a set of target variables. We show cost reduction at the same decision quality and improved decision quality at the same cost in multiple machine learning benchmarks and two real-world health-care applications.
Chao Ma 0019, Sebastian Tschiatschek, Konstantina Palla, José Miguel Hernández-Lobato, Sebastian Nowozin, Cheng Zhang 0005
ICML1