Andrew M. Bean 0001

dblp:244/9323-1 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0001-8439-5975ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 55% Trustworthy machine learning · 37% Reinforcement learning · 8%
Human-computer interaction and pervasive computing
1 paper
Collaborative and social computing · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
1.422024
The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models · NeurIPS 2024
The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values · EMNLP 2023
Machine learning › Trustworthy machine learning
fairness
1.022024
The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models · NeurIPS 2024
The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values · EMNLP 2023
Machine learning › Trustworthy machine learning › interpretability
counterfactual explanation
0.912025
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations · EMNLP 2025
Machine learning › Trustworthy machine learning
interpretability
0.912025
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations · EMNLP 2025
Natural language and speech › Language models and text generation › alignment › preference alignment
human feedback alignment
0.812024
The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models · NeurIPS 2024
Natural language and speech › Language models and text generation
large language model evaluation
0.812024
LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low Resource and Extinct Languages · NeurIPS 2024
Natural language and speech › Language models and text generation
linguistic generalization
0.812024
LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low Resource and Extinct Languages · NeurIPS 2024
Natural language and speech › Language models and text generation › low-resource language processing
low-resource language reasoning
0.812024
LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low Resource and Extinct Languages · NeurIPS 2024
Machine learning › Reinforcement learning › reinforcement learning from human feedback
learning from human feedback
0.712023
The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values · EMNLP 2023
Machine learning › Trustworthy machine learning › language model interpretability
large language model explanation
0.312025
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

validity and minimality evaluation · 0.9counterfactual generation · 0.9in-context learning · 0.8benchmark evaluation · 0.8survey · 0.7
YearPublicationVenuePosition
2025 LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
abstract
To collaborate effectively with humans, language models must be able to explain their decisions in natural language. We study a specific type of self-explanation: self-generated counterfactual explanations (SCEs), where a model explains its prediction by modifying the input such that it would have predicted a different outcome. We evaluate whether LLMs can produce SCEs that are valid, achieving the intended outcome, and minimal, modifying the input no more than necessary. When asked to generate counterfactuals, we find that LLMs typically produce SCEs that are valid, but far from minimal, offering little insight into their decision-making behaviour. Worryingly, when asked to generate minimal counterfactuals, LLMs typically make excessively small edits that fail to change predictions. The observed validity-minimality trade-off is consistent across several LLMs, datasets, and evaluation settings. Our findings suggest that SCEs are, at best, an ineffective explainability tool and, at worst, can provide misleading insights into model behaviour. Proposals to deploy LLMs in high-stakes settings must consider the impact of unreliable self-explanations on downstream decision-making. Our code is available at https://github.com/HarryMayne/SCEs.
Harry Mayne, Ryan Othniel Kearns, Andrew M. Bean 0001, Eoin Delaney, Chris Russell 0001, Adam Mahdi
EMNLP4
2025 Measuring what Matters: Construct Validity in Large Language Model Benchmarks
abstract
Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as safety' androbustness' requires strong construct validity, that is, having measures that represent what matters to the phenomenon. With a team of 29 expert reviewers, we conduct a systematic review of 445 LLM benchmarks from leading conferences in natural language processing and machine learning. Across the reviewed articles, we find patterns related to the measured phenomena, tasks, and scoring metrics which undermine the validity of the resulting claims. To address these shortcomings, we provide eight key recommendations and detailed actionable guidance to researchers and practitioners in developing LLM benchmarks.
Andrew M. Bean 0001, Ryan Othniel Kearns, Angelika Romanou, Franziska Sofia Hafner, Harry Mayne, Jan Batzner, Negar Foroutan Eghlidi, Chris Schmitz, Karolina Korgul, Hunar Batra, Oishi Deb, Emma Beharry, Cornelius Emde, Thomas Foster, Anna Gausen, María Grandury, Simeng Han, Valentin Hofmann, Lujain Ibrahim, Hazel Kim, Hannah Kirk, Fangru Lin, Gabrielle K. Liu, Lennart Luettgau, Jabez Magomere, Jonathan Rystrøm, Anna Sotnikova, Yilun Zhao 0001, Adel Bibi, Antoine Bosselut, Ronald Clark, Arman Cohan, Jakob N. Foerster, Yarin Gal, Scott A. Hale, Inioluwa Deborah Raji, Christopher Summerfield, Philip Torr 0001, Cozmin Ududec, Luc Rocher, Adam Mahdi
NeurIPS1
2024 LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low Resource and Extinct Languages
abstract
In this paper, we present the LingOly benchmark, a novel benchmark for advanced reasoning abilities in large language models. Using challenging Linguistic Olympiad puzzles, we evaluate (i) capabilities for in-context identification and generalisation of linguistic patterns in very low-resource or extinct languages, and (ii) abilities to follow complex task instructions. The LingOly benchmark covers more than 90 mostly low-resource languages, minimising issues of data contamination, and contains 1,133 problems across 6 formats and 5 levels of human difficulty. We assess performance with both direct accuracy and comparison to a no-context baseline to penalise memorisation. Scores from 11 state-of-the-art LLMs demonstrate the benchmark to be challenging, and models perform poorly on the higher difficulty problems. On harder problems, even the top model only achieved 38.7% accuracy, a 24.7% improvement over the no-context baseline. Large closed models typically outperform open models, and in general, the higher resource the language, the better the scores. These results indicate, in absence of memorisation, true multi-step out-of-domain reasoning remains a challenge for current language models.
Andrew M. Bean 0001, Simi Hellsten, Harry Mayne, Jabez Magomere, Ethan A. Chi, Ryan Chi, Scott A. Hale, Hannah Kirk
NeurIPS1
2024 The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
abstract
Human feedback is central to the alignment of Large Language Models (LLMs). However, open questions remain about the methods (how), domains (where), people (who) and objectives (to what end) of feedback processes. To navigate these questions, we introduce PRISM, a new dataset which maps the sociodemographics and stated preferences of 1,500 diverse participants from 75 countries, to their contextual preferences and fine-grained feedback in 8,011 live conversations with 21 LLMs. With PRISM, we contribute (i) wider geographic and demographic participation in feedback; (ii) census-representative samples for two countries (UK, US); and (iii) individualised ratings that link to detailed participant profiles, permitting personalisation and attribution of sample artefacts. We target subjective and multicultural perspectives on value-laden and controversial issues, where we expect interpersonal and cross-cultural disagreement. We use PRISM in three case studies to demonstrate the need for careful consideration of which humans provide alignment data.
Hannah Kirk, Alexander Whitefield, Paul Röttger, Andrew M. Bean 0001, Aikaterini Margatina, Rafael Mosquera, Juan Ciro, Max Bartolo, Adina Williams, He He 0001, Bertie Vidgen, Scott A. Hale
NeurIPS4
2023 The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values
abstract
Human feedback is increasingly used to steer the behaviours of Large Language Models (LLMs).However, it is unclear how to collect and incorporate feedback in a way that is efficient, effective and unbiased, especially for highly subjective human preferences and values.In this paper, we survey existing approaches for learning from human feedback, drawing on 95 papers primarily from the ACL and arXiv repositories.First, we summarise the past, pre-LLM trends for integrating human feedback into language models.Second, we give an overview of present techniques and practices, as well as the motivations for using feedback; conceptual frameworks for defining values and preferences; and how feedback is collected and from whom.Finally, we encourage a better future of feedback learning in LLMs by raising five unresolved conceptual and practical challenges.
Hannah Kirk, Andrew M. Bean 0001, Bertie Vidgen, Paul Röttger, Scott A. Hale
EMNLP2