Susan Wei

dblp:203/8878 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0002-6842-2352ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Deep learning architectures and training · 36% Trustworthy machine learning · 31% Probabilistic and Bayesian machine learning · 16%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian model selection
0.912025
A Bayesian Model Selection Criterion for Selecting Pretraining Checkpoints · ICML 2025
Machine learning › Deep learning architectures and training
checkpoint selection
0.912025
A Bayesian Model Selection Criterion for Selecting Pretraining Checkpoints · ICML 2025
Machine learning › Representation and self-supervised learning
pre-training
0.912025
A Bayesian Model Selection Criterion for Selecting Pretraining Checkpoints · ICML 2025
Machine learning › Trustworthy machine learning › fairness
causal fairness
0.812024
Interventional Fairness on Partially Known Causal Graphs: A Constrained Optimization Approach · ICLR 2024
Machine learning › Trustworthy machine learning
fairness
0.812024
Interventional Fairness on Partially Known Causal Graphs: A Constrained Optimization Approach · ICLR 2024
Machine learning › Trustworthy machine learning › fairness › causal fairness
counterfactual fairness
0.612022
Counterfactual Fairness with Partially Known Causal Graph · NeurIPS 2022
Machine learning › Deep learning architectures and training › neural differential equations
continuous-depth neural networks
0.412020
A shooting formulation of deep learning · NeurIPS 2020
Machine learning › Deep learning architectures and training › neural differential equations
neural ordinary differential equations
0.412020
A shooting formulation of deep learning · NeurIPS 2020
Machine learning › Transfer learning and domain adaptation
fine-tuning
0.312025
A Bayesian Model Selection Criterion for Selecting Pretraining Checkpoints · ICML 2025
Machine learning › Deep learning architectures and training
foundation model
0.312025
A Bayesian Model Selection Criterion for Selecting Pretraining Checkpoints · ICML 2025
Mathematical optimization
constrained optimization
0.212024
Interventional Fairness on Partially Known Causal Graphs: A Constrained Optimization Approach · ICLR 2024
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.212022
Counterfactual Fairness with Partially Known Causal Graph · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

partially directed acyclic graph · 1.5causal inference · 1.5free energy · 0.9bayesian model selection · 0.9causal discovery · 0.6PDAG · 0.6shooting method · 0.4particle ensemble · 0.4optimal control · 0.4
YearPublicationVenuePosition
2025 The Local Learning Coefficient: A Singularity-Aware Complexity Measure
abstract
The Local Learning Coefficient (LLC) is introduced as a novel complexity measure for deep neural networks (DNNs). Recognizing the limitations of traditional complexity measures, the LLC leverages Singular Learning Theory (SLT), which has long recognized the significance of singularities in the loss landscape geometry. This paper provides an extensive exploration of the LLC’s theoretical underpinnings, offering both a clear definition and intuitive insights into its application. Moreover, we propose a new scalable estimator for the LLC, which is then effectively applied across diverse architectures including deep linear networks up to 100M parameters, ResNet image models, and transformer language models. Empirical evidence suggests that the LLC provides valuable insights into how training heuristics might influence the effective complexity of DNNs. Ultimately, the LLC emerges as a crucial tool for reconciling the apparent contradiction between deep learning’s complexity and the principle of parsimony.
Edmund Lau, Zach Furman, George Wang, Daniel Murfet, Susan Wei
AISTATS5
2025 A Bayesian Model Selection Criterion for Selecting Pretraining Checkpoints
abstract
Recent advances in artificial intelligence have been fueled by the development of foundation models such as BERT, GPT, T5, and Vision Transformers. These models are first pretrained on vast and diverse datasets and then adapted to specific downstream tasks, often with significantly less data. However, the mechanisms behind the success of this ubiquitous pretrain-then-adapt paradigm remain underexplored, particularly the characteristics of pretraining checkpoints that enhance downstream adaptation. We introduce a Bayesian model selection criterion, called the downstream free energy, which quantifies a checkpoint’s adaptability by measuring the concentration of nearby favorable parameters for a downstream task. We demonstrate that this Bayesian model selection criterion can be effectively implemented without access to the downstream data or prior knowledge of the downstream task. Furthermore, we provide empirical evidence that the criterion reliably correlates with improved fine-tuning performance, offering a principled approach to predicting model adaptability.
Michael Munn, Susan Wei
ICML2
2025 Temperature Optimization for Bayesian Deep Learning
abstract
The Cold Posterior Effect (CPE) is a phenomenon in Bayesian Deep Learning (BDL), where tempering the posterior to a cold temperature often improves the predictive performance of the posterior predictive distribution (PPD). Although the term ‘CPE’ suggests colder temperatures are inherently better, the BDL community increasingly recognizes that this is not always the case. Despite this, there remains no systematic method for finding the optimal temperature beyond grid search. In this work, we propose a data-driven approach to select the temperature that maximizes test log-predictive density, treating the temperature as a model parameter and estimating it directly from the data. We empirically demonstrate that our method performs comparably to grid search, at a fraction of the cost, across both regression and classification tasks. Finally, we highlight the differing perspectives on CPE between the BDL and Generalized Bayes communities: while the former primarily emphasizes the predictive performance of the PPD, the latter prioritizes the utility of the posterior under model misspecification; these distinct objectives lead to different temperature preferences.
Kenyon Ng, Christopher van der Heide, Liam Hodgkinson, Susan Wei
UAI4
2024 Interventional Fairness on Partially Known Causal Graphs: A Constrained Optimization Approach
abstract
Fair machine learning aims to prevent discrimination against individuals or sub-populations based on sensitive attributes such as gender and race. In recent years, causal inference methods have been increasingly used in fair machine learning to measure unfairness by causal effects. However, current methods assume that the true causal graph is given, which is often not true in real-world applications. To address this limitation, this paper proposes a framework for achieving causal fairness based on the notion of interventions when the true causal graph is partially known. The proposed approach involves modeling fair prediction using a Partially Directed Acyclic Graph (PDAG), specifically, a class of causal DAGs that can be learned from observational data combined with domain knowledge. The PDAG is used to measure causal fairness, and a constrained optimization problem is formulated to balance between fairness and accuracy. Results on both simulated and real-world datasets demonstrate the effectiveness of this method.
Aoqi Zuo, Yiqing Li 0002, Susan Wei, Mingming Gong
ICLR3
2023 Deep Learning Is Singular, and That's Good
abstract
In singular models, the optimal set of parameters forms an analytic set with singularities, and a classical statistical inference cannot be applied to such models. This is significant for deep learning as neural networks are singular, and thus, "dividing" by the determinant of the Hessian or employing the Laplace approximation is not appropriate. Despite its potential for addressing fundamental issues in deep learning, a singular learning theory appears to have made little inroads into the developing canon of a deep learning theory. Via a mix of theory and experiment, we present an invitation to the singular learning theory as a vehicle for understanding deep learning and suggest an important future work to make the singular learning theory directly applicable to how deep learning is performed in practice.
Susan Wei, Daniel Murfet, Mingming Gong, Hui Li 0097, Jesse Gell-Redman, Thomas Quella
IEEE Trans. Neural Networks Learn. Syst.1
2022 Counterfactual Fairness with Partially Known Causal Graph
abstract
Fair machine learning aims to avoid treating individuals or sub-populations unfavourably based on \textit{sensitive attributes}, such as gender and race. Those methods in fair machine learning that are built on causal inference ascertain discrimination and bias through causal effects. Though causality-based fair learning is attracting increasing attention, current methods assume the true causal graph is fully known. This paper proposes a general method to achieve the notion of counterfactual fairness when the true causal graph is unknown. To select features that lead to counterfactual fairness, we derive the conditions and algorithms to identify ancestral relations between variables on a \textit{Partially Directed Acyclic Graph (PDAG)}, specifically, a class of causal DAGs that can be learned from observational data combined with domain knowledge. Interestingly, we find that counterfactual fairness can be achieved as if the true causal graph were fully known, when specific background knowledge is provided: the sensitive attributes do not have ancestors in the causal graph. Results on both simulated and real-world datasets demonstrate the effectiveness of our method.
Aoqi Zuo, Susan Wei, Tongliang Liu, Bo Han 0003, Kun Zhang 0001, Mingming Gong
NeurIPS2
2022 Trade-off between conservation of biological variation and batch effect removal in deep generative modeling for single-cell transcriptomics
abstract
BACKGROUND: Single-cell RNA sequencing (scRNA-seq) technology has contributed significantly to diverse research areas in biology, from cancer to development. Since scRNA-seq data is high-dimensional, a common strategy is to learn low-dimensional latent representations better to understand overall structure in the data. In this work, we build upon scVI, a powerful deep generative model which can learn biologically meaningful latent representations, but which has limited explicit control of batch effects. Rather than prioritizing batch effect removal over conservation of biological variation, or vice versa, our goal is to provide a bird's eye view of the trade-offs between these two conflicting objectives. Specifically, using the well established concept of Pareto front from economics and engineering, we seek to learn the entire trade-off curve between conservation of biological variation and removal of batch effects. RESULTS: A multi-objective optimisation technique known as Pareto multi-task learning (Pareto MTL) is used to obtain the Pareto front between conservation of biological variation and batch effect removal. Our results indicate Pareto MTL can obtain a better Pareto front than the naive scalarization approach typically encountered in the literature. In addition, we propose to measure batch effect by applying a neural-network based estimator called Mutual Information Neural Estimation (MINE) and show benefits over the more standard maximum mean discrepancy measure. CONCLUSION: The Pareto front between conservation of biological variation and batch effect removal is a valuable tool for researchers in computational biology. Our results demonstrate the efficacy of applying Pareto MTL to estimate the Pareto front in conjunction with applying MINE to measure the batch effect.
Davis J. McCarthy, Heejung Shim, Susan Wei
BMC Bioinform.4
2020 A shooting formulation of deep learning
abstract
A residual network may be regarded as a discretization of an ordinary differential equation (ODE) which, in the limit of time discretization, defines a continuous-depth network. Although important steps have been taken to realize the advantages of such continuous formulations, most current techniques assume identical layers. Indeed, existing works throw into relief the myriad difficulties of learning an infinite-dimensional parameter in a continuous-depth neural network. To this end, we introduce a shooting formulation which shifts the perspective from parameterizing a network layer-by-layer to parameterizing over optimal networks described only by a set of initial conditions. For scalability, we propose a novel particle-ensemble parameterization which fully specifies the optimal weight trajectory of the continuous-depth neural network. Our experiments show that our particle-ensemble shooting formulation can achieve competitive performance. Finally, though the current work is inspired by continuous-depth neural networks, the particle-ensemble shooting formulation also applies to discrete-time networks and may lead to a new fertile area of research in deep learning parameterization.
François-Xavier Vialard, Roland Kwitt, Susan Wei, Marc Niethammer
NeurIPS3