Samuel J. Bell

dblp:290/7715 · also Samuel James Bell · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-9437-5449ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Trustworthy machine learning · 53% Probabilistic and Bayesian machine learning · 22% Language models and text generation · 13%
Software engineering, system software, and programming languages
1 paper
Empirical software engineering · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › fairness › algorithmic bias
bias amplification
0.912025
An Effective Theory of Bias Amplification · ICLR 2025
Machine learning › Trustworthy machine learning
fairness
0.912025
An Effective Theory of Bias Amplification · ICLR 2025
Natural language and speech › Language models and text generation
large language model
0.912025
AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › least squares regression
ridge regression
0.912025
An Effective Theory of Bias Amplification · ICLR 2025
Machine learning › Trustworthy machine learning › uncertainty estimation
selective classification
0.912025
AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions · NeurIPS 2025
Machine learning › Trustworthy machine learning
uncertainty estimation
0.912025
AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning › experimental design
bayesian experimental design
0.612022
Modeling the Machine Learning Multiverse · NeurIPS 2022
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization › surrogate model
gaussian process surrogate
0.612022
Modeling the Machine Learning Multiverse · NeurIPS 2022
Empirical software engineering
reproducibility
0.612022
Modeling the Machine Learning Multiverse · NeurIPS 2022

Methods — techniques the papers use, named apart from their topics

gaussian process · 1.1bayesian experimental design · 1.1ridge regression · 0.9random projection · 0.9benchmark evaluation · 0.9
YearPublicationVenuePosition
2025 An Effective Theory of Bias Amplification
abstract
Machine learning models can capture and amplify biases present in data, leading to disparate test performance across social groups. To better understand, evaluate, and mitigate these biases, a deeper theoretical understanding of how model design choices and data distribution properties contribute to bias is needed. In this work, we contribute a precise analytical theory in the context of ridge regression, both with and without random projections, where the former models feedforward neural networks in a simplified regime. Our theory offers a unified and rigorous explanation of machine learning bias, providing insights into phenomena such as bias amplification and minority-group bias in various feature and parameter regimes. For example, we observe that there may be an optimal regularization penalty or training time to avoid bias amplification, and there can be differences in test error between groups that are not alleviated with increased parameterization. Importantly, our theoretical predictions align with empirical observations reported in the literature on machine learning bias. We extensively empirically validate our theory on synthetic and semi-synthetic datasets.
Arjun Subramonian, Samuel J. Bell, Levent Sagun, Elvis Dohmatob
ICLR2
2025 On the Role of Speech Data in Reducing Toxicity Detection Bias
abstract
Samuel Bell, Mariano Coria Meglioli, Megan Richards, Eduardo Sánchez, Christophe Ropers, Skyler Wang, Adina Williams, Levent Sagun, Marta R. Costa-jussà. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Samuel J. Bell, Mariano Coria Meglioli, Megan Richards, Eduardo Sánchez, Christophe Ropers, Skyler Wang, Adina Williams, Levent Sagun, Marta R. Costa-jussà
NAACL (Long Papers)1
2025 AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
abstract
For Large Language Models (LLMs) to be reliably deployed in both everyday and high-stakes domains, knowing when not to answer is equally critical as answering correctly.Real-world user queries, which can be underspecified, ill-posed, or fundamentally unanswerable, require LLMs to reason about uncertainty and selectively abstain---i.e., refuse to answer definitively.However, abstention remains understudied, without a systematic evaluation framework for modern LLMs.In this work, we introduce AbstentionBench: a large-scale benchmark for holistically evaluating abstention across 20 diverse datasets, including questions with unknown answers, underspecification, false premises, subjective interpretations, and outdated information.Evaluating 20 frontier LLMs reveals abstention is an unsolved problem, and one where scaling models is of little use.While recent reasoning LLMs have shown impressive results in complex problem solving, surprisingly, we find that reasoning fine-tuning degrades abstention (by 24\% on average), even for math and science domains on which reasoning models are explicitly trained.We find that while a carefully crafted system prompt can boost abstention in practice, it does not resolve models’ fundamental inability to reason about uncertainty.We release AbstentionBench to foster research into advancing LLM reliability.
Polina Kirichenko, Mark Ibrahim, Kamalika Chaudhuri, Samuel J. Bell
NeurIPS4
2022 Modeling the Machine Learning Multiverse
abstract
Amid mounting concern about the reliability and credibility of machine learning research, we present a principled framework for making robust and generalizable claims: the multiverse analysis. Our framework builds upon the multiverse analysis introduced in response to psychology's own reproducibility crisis. To efficiently explore high-dimensional and often continuous ML search spaces, we model the multiverse with a Gaussian Process surrogate and apply Bayesian experimental design. Our framework is designed to facilitate drawing robust scientific conclusions about model performance, and thus our approach focuses on exploration rather than conventional optimization. In the first of two case studies, we investigate disputed claims about the relative merit of adaptive optimizers. Second, we synthesize conflicting research on the effect of learning rate on the large batch training generalization gap. For the machine learning community, the multiverse analysis is a simple and effective technique for identifying robust claims, for increasing transparency, and a step toward improved reproducibility.
Samuel J. Bell, Onno Kampman, Jesse Dodge, Neil D. Lawrence
NeurIPS1