Harry Mayne

dblp:372/0330 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 54% Trustworthy machine learning · 46%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › interpretability
counterfactual explanation
0.912025
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations · EMNLP 2025
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
0.912025
How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis · EMNLP 2025
Machine learning › Trustworthy machine learning
interpretability
0.912025
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations · EMNLP 2025
Machine learning › Trustworthy machine learning › interpretability › neural network interpretation
neuron analysis
0.912025
How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis · EMNLP 2025
Natural language and speech › Language models and text generation › large language model safety
safety fine-tuning
0.912025
How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis · EMNLP 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.812024
LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low Resource and Extinct Languages · NeurIPS 2024
Natural language and speech › Language models and text generation
linguistic generalization
0.812024
LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low Resource and Extinct Languages · NeurIPS 2024
Natural language and speech › Language models and text generation › low-resource language processing
low-resource language reasoning
0.812024
LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low Resource and Extinct Languages · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability › representation engineering
activation editing
0.312025
How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis · EMNLP 2025
Machine learning › Trustworthy machine learning › language model interpretability
large language model explanation
0.312025
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations · EMNLP 2025
Machine learning › Trustworthy machine learning
toxicity reduction
0.312025
How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

validity and minimality evaluation · 0.9mechanistic interpretability · 0.9counterfactual generation · 0.9activation editing · 0.9in-context learning · 0.8benchmark evaluation · 0.8
YearPublicationVenuePosition
2025 LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
abstract
To collaborate effectively with humans, language models must be able to explain their decisions in natural language. We study a specific type of self-explanation: self-generated counterfactual explanations (SCEs), where a model explains its prediction by modifying the input such that it would have predicted a different outcome. We evaluate whether LLMs can produce SCEs that are valid, achieving the intended outcome, and minimal, modifying the input no more than necessary. When asked to generate counterfactuals, we find that LLMs typically produce SCEs that are valid, but far from minimal, offering little insight into their decision-making behaviour. Worryingly, when asked to generate minimal counterfactuals, LLMs typically make excessively small edits that fail to change predictions. The observed validity-minimality trade-off is consistent across several LLMs, datasets, and evaluation settings. Our findings suggest that SCEs are, at best, an ineffective explainability tool and, at worst, can provide misleading insights into model behaviour. Proposals to deploy LLMs in high-stakes settings must consider the impact of unreliable self-explanations on downstream decision-making. Our code is available at https://github.com/HarryMayne/SCEs.
Harry Mayne, Ryan Othniel Kearns, Andrew M. Bean 0001, Eoin Delaney, Chris Russell 0001, Adam Mahdi
EMNLP1
2025 How Does DPO Reduce Toxicity? A Mechanistic Neuron-Level Analysis
abstract
Safety fine-tuning algorithms reduce harmful outputs in language models, yet their mechanisms remain under-explored.Direct Preference Optimization (DPO) is a popular choice of algorithm, but prior explanations, attributing its effects solely to dampened toxic neurons in the MLP layers, are incomplete.In this study, we analyse four language models (Llama-3.1-8B,Gemma-2-2B, Mistral-7B, GPT-2-Medium) and show that toxic neurons only account for 2.5% to 24% of DPO's effects across models.Instead, DPO balances distributed activation shifts across a majority of MLP neurons to create a net toxicity reduction.We attribute this reduction to four neuron groups, two aligned with reducing toxicity and two promoting anti-toxicity, whose combined effects replicate DPO across models.To further validate this understanding, we develop an activation editing method mimicking DPO through distributed shifts along a toxicity representation in both probe-based and probe-free settings.This method outperforms DPO in reducing toxicity while preserving perplexity across models, without requiring any weight updates.Our work provides a mechanistic understanding of DPO and introduces an efficient, tuning-free alternative for safety fine-tuning.Our code is available on dpo-toxic-neurons.
Filip Sondej, Harry Mayne, Adam Mahdi
EMNLP3
2025 Measuring what Matters: Construct Validity in Large Language Model Benchmarks
abstract
Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as safety' androbustness' requires strong construct validity, that is, having measures that represent what matters to the phenomenon. With a team of 29 expert reviewers, we conduct a systematic review of 445 LLM benchmarks from leading conferences in natural language processing and machine learning. Across the reviewed articles, we find patterns related to the measured phenomena, tasks, and scoring metrics which undermine the validity of the resulting claims. To address these shortcomings, we provide eight key recommendations and detailed actionable guidance to researchers and practitioners in developing LLM benchmarks.
Andrew M. Bean 0001, Ryan Othniel Kearns, Angelika Romanou, Franziska Sofia Hafner, Harry Mayne, Jan Batzner, Negar Foroutan Eghlidi, Chris Schmitz, Karolina Korgul, Hunar Batra, Oishi Deb, Emma Beharry, Cornelius Emde, Thomas Foster, Anna Gausen, María Grandury, Simeng Han, Valentin Hofmann, Lujain Ibrahim, Hazel Kim, Hannah Kirk, Fangru Lin, Gabrielle K. Liu, Lennart Luettgau, Jabez Magomere, Jonathan Rystrøm, Anna Sotnikova, Yilun Zhao 0001, Adel Bibi, Antoine Bosselut, Ronald Clark, Arman Cohan, Jakob N. Foerster, Yarin Gal, Scott A. Hale, Inioluwa Deborah Raji, Christopher Summerfield, Philip Torr 0001, Cozmin Ududec, Luc Rocher, Adam Mahdi
NeurIPS5
2024 LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low Resource and Extinct Languages
abstract
In this paper, we present the LingOly benchmark, a novel benchmark for advanced reasoning abilities in large language models. Using challenging Linguistic Olympiad puzzles, we evaluate (i) capabilities for in-context identification and generalisation of linguistic patterns in very low-resource or extinct languages, and (ii) abilities to follow complex task instructions. The LingOly benchmark covers more than 90 mostly low-resource languages, minimising issues of data contamination, and contains 1,133 problems across 6 formats and 5 levels of human difficulty. We assess performance with both direct accuracy and comparison to a no-context baseline to penalise memorisation. Scores from 11 state-of-the-art LLMs demonstrate the benchmark to be challenging, and models perform poorly on the higher difficulty problems. On harder problems, even the top model only achieved 38.7% accuracy, a 24.7% improvement over the no-context baseline. Large closed models typically outperform open models, and in general, the higher resource the language, the better the scores. These results indicate, in absence of memorisation, true multi-step out-of-domain reasoning remains a challenge for current language models.
Andrew M. Bean 0001, Simi Hellsten, Harry Mayne, Jabez Magomere, Ethan A. Chi, Ryan Chi, Scott A. Hale, Hannah Kirk
NeurIPS3