VLDB 2026 Research / reviewers in the wild / expert
Andrew M. Bean 0001
dblp:244/9323-1
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0001-8439-5975ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Language models and text generation · 55% Trustworthy machine learning · 37% Reinforcement learning · 8% | |
| Human-computer interaction and pervasive computing
1 paper |
Collaborative and social computing · 100% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
alignment |
1.4 | 2 | 2024 | The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models · NeurIPS 2024 The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values · EMNLP 2023 |
Machine learning › Trustworthy machine learning
fairness |
1.0 | 2 | 2024 | The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models · NeurIPS 2024 The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values · EMNLP 2023 |
Machine learning › Trustworthy machine learning › interpretability
counterfactual explanation |
0.9 | 1 | 2025 | LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations · EMNLP 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations · EMNLP 2025 |
Natural language and speech › Language models and text generation › alignment › preference alignment
human feedback alignment |
0.8 | 1 | 2024 | The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models · NeurIPS 2024 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.8 | 1 | 2024 | LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low Resource and Extinct Languages · NeurIPS 2024 |
Natural language and speech › Language models and text generation
linguistic generalization |
0.8 | 1 | 2024 | LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low Resource and Extinct Languages · NeurIPS 2024 |
Natural language and speech › Language models and text generation › low-resource language processing
low-resource language reasoning |
0.8 | 1 | 2024 | LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low Resource and Extinct Languages · NeurIPS 2024 |
Machine learning › Reinforcement learning › reinforcement learning from human feedback
learning from human feedback |
0.7 | 1 | 2023 | The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values · EMNLP 2023 |
Machine learning › Trustworthy machine learning › language model interpretability
large language model explanation |
0.3 | 1 | 2025 | LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
validity and minimality evaluation · 0.9counterfactual generation · 0.9in-context learning · 0.8benchmark evaluation · 0.8survey · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual ExplanationsabstractTo collaborate effectively with humans, language models must be able to explain their decisions in natural language. We study a specific type of self-explanation: self-generated counterfactual explanations (SCEs), where a model explains its prediction by modifying the input such that it would have predicted a different outcome. We evaluate whether LLMs can produce SCEs that are valid, achieving the intended outcome, and minimal, modifying the input no more than necessary. When asked to generate counterfactuals, we find that LLMs typically produce SCEs that are valid, but far from minimal, offering little insight into their decision-making behaviour. Worryingly, when asked to generate minimal counterfactuals, LLMs typically make excessively small edits that fail to change predictions. The observed validity-minimality trade-off is consistent across several LLMs, datasets, and evaluation settings. Our findings suggest that SCEs are, at best, an ineffective explainability tool and, at worst, can provide misleading insights into model behaviour. Proposals to deploy LLMs in high-stakes settings must consider the impact of unreliable self-explanations on downstream decision-making. Our code is available at https://github.com/HarryMayne/SCEs. Harry Mayne, Ryan Othniel Kearns, Andrew M. Bean 0001, Eoin Delaney, Chris Russell 0001, Adam Mahdi |
EMNLP | 4 |
| 2025 | Measuring what Matters: Construct Validity in Large Language Model BenchmarksabstractEvaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as safety' androbustness' requires strong construct validity, that is, having measures that represent what matters to the phenomenon. With a team of 29 expert reviewers, we conduct a systematic review of 445 LLM benchmarks from leading conferences in natural language processing and machine learning. Across the reviewed articles, we find patterns related to the measured phenomena, tasks, and scoring metrics which undermine the validity of the resulting claims. To address these shortcomings, we provide eight key recommendations and detailed actionable guidance to researchers and practitioners in developing LLM benchmarks. Andrew M. Bean 0001, Ryan Othniel Kearns, Angelika Romanou, Franziska Sofia Hafner, Harry Mayne, Jan Batzner, Negar Foroutan Eghlidi, Chris Schmitz, Karolina Korgul, Hunar Batra, Oishi Deb, Emma Beharry, Cornelius Emde, Thomas Foster, Anna Gausen, María Grandury, Simeng Han, Valentin Hofmann, Lujain Ibrahim, Hazel Kim, Hannah Kirk, Fangru Lin, Gabrielle K. Liu, Lennart Luettgau, Jabez Magomere, Jonathan Rystrøm, Anna Sotnikova, Yilun Zhao 0001, Adel Bibi, Antoine Bosselut, Ronald Clark, Arman Cohan, Jakob N. Foerster, Yarin Gal, Scott A. Hale, Inioluwa Deborah Raji, Christopher Summerfield, Philip Torr 0001, Cozmin Ududec, Luc Rocher, Adam Mahdi |
NeurIPS | 1 |
| 2024 | LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low Resource and Extinct LanguagesabstractIn this paper, we present the LingOly benchmark, a novel benchmark for advanced reasoning abilities in large language models. Using challenging Linguistic Olympiad puzzles, we evaluate (i) capabilities for in-context identification and generalisation of linguistic patterns in very low-resource or extinct languages, and (ii) abilities to follow complex task instructions. The LingOly benchmark covers more than 90 mostly low-resource languages, minimising issues of data contamination, and contains 1,133 problems across 6 formats and 5 levels of human difficulty. We assess performance with both direct accuracy and comparison to a no-context baseline to penalise memorisation. Scores from 11 state-of-the-art LLMs demonstrate the benchmark to be challenging, and models perform poorly on the higher difficulty problems. On harder problems, even the top model only achieved 38.7% accuracy, a 24.7% improvement over the no-context baseline. Large closed models typically outperform open models, and in general, the higher resource the language, the better the scores. These results indicate, in absence of memorisation, true multi-step out-of-domain reasoning remains a challenge for current language models. Andrew M. Bean 0001, Simi Hellsten, Harry Mayne, Jabez Magomere, Ethan A. Chi, Ryan Chi, Scott A. Hale, Hannah Kirk |
NeurIPS | 1 |
| 2024 | The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language ModelsabstractHuman feedback is central to the alignment of Large Language Models (LLMs). However, open questions remain about the methods (how), domains (where), people (who) and objectives (to what end) of feedback processes. To navigate these questions, we introduce PRISM, a new dataset which maps the sociodemographics and stated preferences of 1,500 diverse participants from 75 countries, to their contextual preferences and fine-grained feedback in 8,011 live conversations with 21 LLMs. With PRISM, we contribute (i) wider geographic and demographic participation in feedback; (ii) census-representative samples for two countries (UK, US); and (iii) individualised ratings that link to detailed participant profiles, permitting personalisation and attribution of sample artefacts. We target subjective and multicultural perspectives on value-laden and controversial issues, where we expect interpersonal and cross-cultural disagreement. We use PRISM in three case studies to demonstrate the need for careful consideration of which humans provide alignment data. Hannah Kirk, Alexander Whitefield, Paul Röttger, Andrew M. Bean 0001, Aikaterini Margatina, Rafael Mosquera, Juan Ciro, Max Bartolo, Adina Williams, He He 0001, Bertie Vidgen, Scott A. Hale |
NeurIPS | 4 |
| 2023 | The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and ValuesabstractHuman feedback is increasingly used to steer the behaviours of Large Language Models (LLMs).However, it is unclear how to collect and incorporate feedback in a way that is efficient, effective and unbiased, especially for highly subjective human preferences and values.In this paper, we survey existing approaches for learning from human feedback, drawing on 95 papers primarily from the ACL and arXiv repositories.First, we summarise the past, pre-LLM trends for integrating human feedback into language models.Second, we give an overview of present techniques and practices, as well as the motivations for using feedback; conceptual frameworks for defining values and preferences; and how feedback is collected and from whom.Finally, we encourage a better future of feedback learning in LLMs by raising five unresolved conceptual and practical challenges. Hannah Kirk, Andrew M. Bean 0001, Bertie Vidgen, Paul Röttger, Scott A. Hale |
EMNLP | 2 |