Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

David Madras

dblp:188/6211 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
Trustworthy machine learning · 49% Probabilistic and Bayesian machine learning · 18% Language models and text generation · 11%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 22 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
2.352025
Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairness · NeurIPS 2025
Causal Modeling for Fairness In Dynamical Systems · ICML 2020
Flexibly Fair Representation Learning by Disentanglement · ICML 2019
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
regression
1.622025
Regression for the Mean: Auto-Evaluation and Inference with Few Labels through Post-hoc Regression · ICML 2025
Out of the Ordinary: Spectrally Adapting Regression for Covariate Shift · ICML 2024
Machine learning › Trustworthy machine learning
robustness
0.922021
Identifying and Benchmarking Natural Out-of-Context Prediction Problems · NeurIPS 2021
Detecting Extrapolation with Local Ensembles · ICLR 2020
Machine learning › Trustworthy machine learning › fairness › causal fairness
causal fairness analysis
0.912025
Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairness · NeurIPS 2025
Machine learning › Trustworthy machine learning
prediction-powered inference
0.912025
Regression for the Mean: Auto-Evaluation and Inference with Few Labels through Post-hoc Regression · ICML 2025
Machine learning › Learning theory › statistical estimation › robust statistics
robust regression
0.912025
Regression for the Mean: Auto-Evaluation and Inference with Few Labels through Post-hoc Regression · ICML 2025
Machine learning › Probabilistic and Bayesian machine learning
statistical inference
0.912025
Regression for the Mean: Auto-Evaluation and Inference with Few Labels through Post-hoc Regression · ICML 2025
Mathematical optimization › stochastic optimization
variance reduction
0.912025
QuEst: Enhancing Estimates of Quantile-Based Distributional Measures Using Model Predictions · ICML 2025
Machine learning › Transfer learning and domain adaptation › domain shift
covariate shift
0.812024
Out of the Ordinary: Spectrally Adapting Regression for Covariate Shift · ICML 2024
Natural language and speech › Language models and text generation
large language model safety
0.812024
Learning and Forgetting Unsafe Examples in Large Language Models · ICML 2024
Natural language and speech › Language models and text generation › large language model safety
safety fine-tuning
0.812024
Learning and Forgetting Unsafe Examples in Large Language Models · ICML 2024
Machine learning › Trustworthy machine learning › fairness
fair representation learning
0.722019
Flexibly Fair Representation Learning by Disentanglement · ICML 2019
Learning Adversarially Fair and Transferable Representations · ICML 2018
Machine learning › Trustworthy machine learning › robustness
distribution shift
0.512021
Identifying and Benchmarking Natural Out-of-Context Prediction Problems · NeurIPS 2021
Machine learning › Kernel, tree and ensemble methods
ensemble learning
0.412020
Detecting Extrapolation with Local Ensembles · ICLR 2020
Machine learning › Trustworthy machine learning
uncertainty estimation
0.412020
Detecting Extrapolation with Local Ensembles · ICLR 2020
Machine learning › Representation and self-supervised learning › representation learning
adversarial representation learning
0.312018
Learning Adversarially Fair and Transferable Representations · ICML 2018
Knowledge, reasoning and agents › Multi-agent systems › human-agent interaction
human-AI decision making
0.312018
Predict Responsibly: Improving Fairness and Accuracy by Learning to Defer · NeurIPS 2018
Knowledge, reasoning and agents › Multi-agent systems › human-agent interaction › human-AI decision making
learning to defer
0.312018
Predict Responsibly: Improving Fairness and Accuracy by Learning to Defer · NeurIPS 2018
Machine learning › Trustworthy machine learning › uncertainty estimation
selective classification
0.312018
Predict Responsibly: Improving Fairness and Accuracy by Learning to Defer · NeurIPS 2018
Performance modeling and evaluation
benchmarking
0.112021
Identifying and Benchmarking Natural Out-of-Context Prediction Problems · NeurIPS 2021
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal model
0.112020
Causal Modeling for Fairness In Dynamical Systems · ICML 2020
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning
0.112019
Flexibly Fair Representation Learning by Disentanglement · ICML 2019

Methods — techniques the papers use, named apart from their topics

optimization · 1.7confidence interval estimation · 1.7ordinary least squares · 1.6benchmark design · 1.0robust regression · 0.9conditional independence testing · 0.9causal graphical models · 0.9spectral decomposition · 0.8forgetfilter · 0.8data filtering · 0.8local ensembles · 0.4causal directed acyclic graphs · 0.4
YearPublicationVenuePosition
2025 QuEst: Enhancing Estimates of Quantile-Based Distributional Measures Using Model Predictions
abstract
As machine learning models grow increasingly competent, their predictions can supplement scarce or expensive data in various important domains. In support of this paradigm, algorithms have emerged to combine a small amount of high-fidelity observed data with a much larger set of imputed model outputs to estimate some quantity of interest. Yet current hybrid-inference tools target only means or single quantiles, limiting their applicability for many critical domains and use cases. We present QuEst, a principled framework to merge observed and imputed data to deliver point estimates and rigorous confidence intervals for a wide family of quantile-based distributional measures. QuEst covers a range of measures, from tail risk (CVaR) to population segments such as quartiles, that are central to fields such as economics, sociology, education, medicine, and more. We extend QuEst to multidimensional metrics, and introduce an additional optimization technique to further reduce variance in this and other hybrid estimators. We demonstrate the utility of our framework through experiments in economic modeling, opinion polling, and language model auto-evaluation.
Zhun Deng, Thomas P. Zollo, Benjamin Eyre, Amogh Inamdar, David Madras, Richard S. Zemel
ICML5
2025 Regression for the Mean: Auto-Evaluation and Inference with Few Labels through Post-hoc Regression
abstract
The availability of machine learning systems that can effectively perform arbitrary tasks has led to synthetic labels from these systems being used in applications of statistical inference, such as data analysis or model evaluation. The Prediction Powered Inference (PPI) framework provides a way of leveraging both a large pool of pseudo-labelled data and a small sample with real, high-quality labels to produce a low-variance, unbiased estimate of the quantity being evaluated for. Most work on PPI considers a relatively sizable set of labelled samples, which can be resource intensive to obtain. However, we find that when labelled data is scarce, the PPI++ method can perform even worse than classical inference. We analyze this phenomenon by relating PPI++ to ordinary least squares regression, which also experiences high variance with small sample sizes, and use this regression framework to better understand the efficacy of PPI. Motivated by this, we present two new PPI-based techniques that leverage robust regressors to produce even lower variance estimators in the few-label regime
Benjamin Eyre, David Madras
ICML2
2025 Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairness
abstract
Disaggregated evaluation across subgroups is critical for assessing the fairness of machine learning models, but its uncritical use can mislead practitioners. We show that equal performance across subgroups is an unreliable measure of fairness when data are representative of the relevant populations but reflective of real-world disparities. Furthermore, when data are not representative due to selection bias, both disaggregated evaluation and alternative approaches based on conditional independence testing may be invalid without explicit assumptions regarding the bias mechanism. We use causal graphical models to characterize fairness properties and metric stability across subgroups under different data generating processes. Our framework suggests complementing disaggregated evaluations with explicit causal assumptions and analysis to control for confounding and distribution shift, including conditional independence testing and weighted performance estimation. These findings have broad implications for how practitioners design and interpret model assessments given the ubiquity of disaggregated evaluation.
Stephen Pfohl, Natalie Harris, Chirag Nagpal, David Madras, Vishwali Mhasawade, Olawale Salaudeen, Awa Dieng, Shannon Sequeira, Santiago Eduardo Arciniegas, Lillian Sung, Nnamdi Ezeanochie, Heather Cole-Lewis, Katherine A. Heller, Oluwasanmi Koyejo, Alexander D'Amour
NeurIPS4
2024 Out of the Ordinary: Spectrally Adapting Regression for Covariate Shift
abstract
Designing deep neural network classifiers that perform robustly on distributions differing from the available training data is an active area of machine learning research. However, out-of-distribution generalization for regression---the analogous problem for modeling continuous targets---remains relatively unexplored. To tackle this problem, we return to first principles and analyze how the closed-form solution for Ordinary Least Squares (OLS) regression is sensitive to covariate shift. We characterize the out-of-distribution risk of the OLS model in terms of the eigenspectrum decomposition of the source and target data. We then use this insight to propose a method called Spectral Adapted Regressor (SpAR) for adapting the weights of the last layer of a pre-trained neural regression model to perform better on input data originating from a different distribution. We demonstrate how this lightweight spectral adaptation procedure can improve out-of-distribution performance for synthetic and real-world datasets.
Benjamin Eyre, Elliot Creager, David Madras, Vardan Papyan, Richard S. Zemel
ICML3
2024 Learning and Forgetting Unsafe Examples in Large Language Models
abstract
As the number of large language models (LLMs) released to the public grows, there is a pressing need to understand the safety implications associated with these models learning from third-party custom finetuning data. We explore the behavior of LLMs finetuned on noisy custom data containing unsafe content, represented by datasets that contain biases, toxicity, and harmfulness, finding that while aligned LLMs can readily learn this unsafe content, they also tend to forget it more significantly than other examples when subsequently finetuned on safer content. Drawing inspiration from the discrepancies in forgetting, we introduce the “ForgetFilter” algorithm, which filters unsafe data based on how strong the model’s forgetting signal is for that data. We demonstrate that the ForgetFilter algorithm ensures safety in customized finetuning without compromising downstream task performance, unlike sequential safety finetuning. ForgetFilter outperforms alternative strategies like replay and moral self-correction in curbing LLMs’ ability to assimilate unsafe content during custom finetuning, e.g. 75% lower than not applying any safety measures and 62% lower than using self-correction in toxicity score.
Zhun Deng, David Madras, James Zou 0001, Mengye Ren
ICML3
2021 Identifying and Benchmarking Natural Out-of-Context Prediction Problems
abstract
Deep learning systems frequently fail at out-of-context (OOC) prediction, the problem of making reliable predictions on uncommon or unusual inputs or subgroups of the training distribution. To this end, a number of benchmarks for measuring OOC performance have been recently introduced. In this work, we introduce a framework unifying the literature on OOC performance measurement, and demonstrate how rich auxiliary information can be leveraged to identify candidate sets of OOC examples in existing datasets. We present NOOCh: a suite of naturally-occurring "challenge sets", and show how varying notions of context can be used to probe specific OOC failure modes. Experimentally, we explore the tradeoffs between various learning approaches on these challenge sets and demonstrate how the choices made in designing OOC benchmarks can yield varying conclusions.
David Madras, Richard S. Zemel
NeurIPS1
2020 Detecting Extrapolation with Local Ensembles
David Madras, James Atwood, Alexander D'Amour
ICLR1
2020 Causal Modeling for Fairness In Dynamical Systems
abstract
In many applications areas—lending, education, and online recommenders, for example—fairness and equity concerns emerge when a machine learning system interacts with a dynamically changing environment to produce both immediate and long-term effects for individuals and demographic groups. We discuss causal directed acyclic graphs (DAGs) as a unifying framework for the recent literature on fairness in such dynamical systems. We show that this formulation affords several new directions of inquiry to the modeler, where sound causal assumptions can be expressed and manipulated. We emphasize the importance of computing interventional quantities in the dynamical fairness setting, and show how causal assumptions enable simulation (when environment dynamics are known) and estimation by adjustment (when dynamics are unknown) of intervention on short- and long-term outcomes, at both the group and individual levels.
Elliot Creager, David Madras, Toniann Pitassi, Richard S. Zemel
ICML2
2019 Flexibly Fair Representation Learning by Disentanglement
abstract
We consider the problem of learning representations that achieve group and subgroup fairness with respect to multiple sensitive attributes. Taking inspiration from the disentangled representation learning literature, we propose an algorithm for learning compact representations of datasets that are useful for reconstruction and prediction, but are also flexibly fair, meaning they can be easily modified at test time to achieve subgroup demographic parity with respect to multiple sensitive attributes and their conjunctions. We show empirically that the resulting encoder—which does not require the sensitive attributes for inference—allows for the adaptation of a single representation to a variety of fair classification tasks with new target labels and subgroup definitions.
Elliot Creager, David Madras, Jörn-Henrik Jacobsen, Marissa A. Weis, Kevin Swersky, Toniann Pitassi, Richard S. Zemel
ICML2
2018 Learning Adversarially Fair and Transferable Representations
abstract
In this paper, we advocate for representation learning as the key to mitigating unfair prediction outcomes downstream. Motivated by a scenario where learned representations are used by third parties with unknown objectives, we propose and explore adversarial representation learning as a natural method of ensuring those parties act fairly. We connect group fairness (demographic parity, equalized odds, and equal opportunity) to different adversarial objectives. Through worst-case theoretical guarantees and experimental validation, we show that the choice of this objective is crucial to fair prediction. Furthermore, we present the first in-depth experimental demonstration of fair transfer learning and demonstrate empirically that our learned representations admit fair predictions on new tasks while maintaining utility, an essential goal of fair representation learning.
David Madras, Elliot Creager, Toniann Pitassi, Richard S. Zemel
ICML1
2018 Predict Responsibly: Improving Fairness and Accuracy by Learning to Defer
abstract
In many machine learning applications, there are multiple decision-makers involved, both automated and human. The interaction between these agents often goes unaddressed in algorithmic development. In this work, we explore a simple version of this interaction with a two-stage framework containing an automated model and an external decision-maker. The model can choose to say PASS, and pass the decision downstream, as explored in rejection learning. We extend this concept by proposing "learning to defer", which generalizes rejection learning by considering the effect of other agents in the decision-making process. We propose a learning algorithm which accounts for potential biases held by external decision-makers in a system. Experiments demonstrate that learning to defer can make systems not only more accurate but also less biased. Even when working with inconsistent or biased users, we show that deferring models still greatly improve the accuracy and/or fairness of the entire system.
David Madras, Toniann Pitassi, Richard S. Zemel
NeurIPS1