Tom Bewley

dblp:246/7724 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0002-5460-0744ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Trustworthy machine learning · 41% Reinforcement learning · 31% Language models and text generation · 17%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 44% Games and playful interaction · 44% Collaborative and social computing · 13%

Topics — the 21 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
2.132025
Interpreting Language Reward Models via Contrastive Explanations · ICLR 2025
Counterfactual Metarules for Local and Global Recourse · ICML 2024
TripleTree: A Versatile Interpretable Representation of Black Box Agents and their Environments · AAAI 2021
Natural language and speech › Language models and text generation › model steering › language model steering
activation steering
0.912025
To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models · ICML 2025
Natural language and speech › Language models and text generation
alignment
0.912025
Interpreting Language Reward Models via Contrastive Explanations · ICLR 2025
Natural language and speech › Question answering and dialogue systems
answer aggregation
0.912025
Representation Consistency for Accurate and Coherent LLM Answer Aggregation · NeurIPS 2025
Machine learning › Trustworthy machine learning › uncertainty estimation
selective classification
0.912025
To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models · ICML 2025
Natural language and speech › Language models and text generation
test-time scaling
0.912025
Representation Consistency for Accurate and Coherent LLM Answer Aggregation · NeurIPS 2025
Machine learning › Trustworthy machine learning › interpretability
counterfactual explanation
0.812024
Counterfactual Metarules for Local and Global Recourse · ICML 2024
Machine learning › Trustworthy machine learning › robustness › distribution shift
distribution shift detection
0.812024
Sequential Harmful Shift Detection Without Labels · NeurIPS 2024
Machine learning › Reinforcement learning
offline reinforcement learning
0.812024
Zero-Shot Reinforcement Learning from Low Quality Data · NeurIPS 2024
Machine learning › Reinforcement learning › offline reinforcement learning
pessimism
0.812024
Zero-Shot Reinforcement Learning from Low Quality Data · NeurIPS 2024
Machine learning › Trustworthy machine learning
robustness
0.812024
Sequential Harmful Shift Detection Without Labels · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability › logic-based explanation
rule-based explanation
0.812024
Counterfactual Metarules for Local and Global Recourse · ICML 2024
Machine learning › Reinforcement learning › generalization in reinforcement learning
zero-shot reinforcement learning
0.812024
Zero-Shot Reinforcement Learning from Low Quality Data · NeurIPS 2024
Machine learning › Learning paradigms
multiple instance learning
0.612022
Non-Markovian Reward Modelling from Trajectory Labels via Interpretable Multiple Instance Learning · NeurIPS 2022
Machine learning › Reinforcement learning › reward design
non-markovian reward
0.612022
Non-Markovian Reward Modelling from Trajectory Labels via Interpretable Multiple Instance Learning · NeurIPS 2022
Machine learning › Reinforcement learning › reward learning
reward modeling
0.612022
Non-Markovian Reward Modelling from Trajectory Labels via Interpretable Multiple Instance Learning · NeurIPS 2022
Machine learning › Reinforcement learning
agent behavior analysis
0.512021
TripleTree: A Versatile Interpretable Representation of Black Box Agents and their Environments · AAAI 2021
Games and playful interaction
gamification
0.412019
On The Combination of Gamification and Crowd Computation in Industrial Automation and Robotics Applications · ICRA 2019
Human-AI interaction
human-AI collaboration
0.412019
On The Combination of Gamification and Crowd Computation in Industrial Automation and Robotics Applications · ICRA 2019
Machine learning › Learning theory › statistical estimation
error estimation
0.212024
Sequential Harmful Shift Detection Without Labels · NeurIPS 2024
Collaborative and social computing
crowdsourcing
0.112019
On The Combination of Gamification and Crowd Computation in Industrial Automation and Robotics Applications · ICRA 2019

Methods — techniques the papers use, named apart from their topics

sparse autoencoder · 0.9representation similarity · 0.9perturbation · 0.9mechanistic interpretability · 0.9contrastive explanation · 0.9calibration · 0.9tree-based surrogate models · 0.8sequential testing · 0.8proxy error estimator · 0.8meta-rules · 0.8gamification · 0.4
YearPublicationVenuePosition
2025 Interpreting Language Reward Models via Contrastive Explanations
abstract
Reward models (RMs) are a crucial component in the alignment of large language models’ (LLMs) outputs with human values. RMs approximate human preferences over possible LLM responses to the same prompt by predicting and comparing reward scores. However, as they are typically modified versions of LLMs with scalar output heads, RMs are large black boxes whose predictions are not explainable. More transparent RMs would enable improved trust in the alignment of LLMs. In this work, we propose to use contrastive explanations to explain any binary response comparison made by an RM. Specifically, we generate a diverse set of new comparisons similar to the original one to characterise the RM’s local behaviour. The perturbed responses forming the new comparisons are generated to explicitly modify manually specified high-level evaluation attributes, on which analyses of RM behaviour are grounded. In quantitative experiments, we validate the effectiveness of our method for finding high-quality contrastive explanations. We then showcase the qualitative usefulness of our method for investigating global sensitivity of RMs to each evaluation attribute, and demonstrate how representative examples can be automatically extracted to explain and compare behaviours of different RMs. We see our method as a flexible framework for RM explanation, providing a basis for more interpretable and trustworthy LLM alignment.
Junqi Jiang, Tom Bewley, Saumitra Mishra, Freddy Lécué, Manuela M. Veloso
ICLR2
2025 To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models
abstract
We introduce Mechanistic Error Reduction with Abstention (MERA), a principled framework for steering language models (LMs) to mitigate errors through selective, adaptive interventions. Unlike existing methods that rely on fixed, manually tuned steering strengths, often resulting in under or oversteering, MERA addresses these limitations by (i) optimising the intervention direction, and (ii) calibrating when and how much to steer, thereby provably improving performance or abstaining when no confident correction is possible. Experiments across diverse datasets and LM families demonstrate safe, effective, non-degrading error correction and that MERA outperforms existing baselines. Moreover, MERA can be applied on top of existing steering techniques to further enhance their performance, establishing it as a general-purpose and efficient approach to mechanistic activation steering.
Anna Hedström, Salim I. Amoukou, Tom Bewley, Saumitra Mishra, Manuela M. Veloso
ICML3
2025 Representation Consistency for Accurate and Coherent LLM Answer Aggregation
abstract
Test-time scaling improves large language models' (LLMs) performance by allocating more compute budget during inference. To achieve this, existing methods often require intricate modifications to prompting and sampling strategies. In this work, we introduce representation consistency (RC), a test-time scaling method for aggregating answers drawn from multiple candidate responses of an LLM regardless of how they were generated, including variations in prompt phrasing and sampling strategy. RC enhances answer aggregation by not only considering the number of occurrences of each answer in the candidate response set, but also the consistency of the model's internal activations while generating the set of responses leading to each answer. These activations can be either dense (raw model activations) or sparse (encoded via pretrained sparse autoencoders). Our rationale is that if the model's representations of multiple responses converging on the same answer are highly variable, this answer is more likely to be the result of incoherent reasoning and should be down-weighted during aggregation. Importantly, our method only uses cached activations and lightweight similarity computations and requires no additional model queries. Through experiments with four open-source LLMs and four reasoning datasets, we validate the effectiveness of RC for improving task performance during inference, with consistent accuracy improvements (up to 4\%) over strong test-time scaling baselines. We also show that consistency in the sparse activation signals aligns well with the common notion of coherent reasoning.
Junqi Jiang, Tom Bewley, Salim I. Amoukou, Francesco Leofante, Antonio Rago 0001, Saumitra Mishra, Francesca Toni
NeurIPS2
2024 Counterfactual Metarules for Local and Global Recourse
abstract
We introduce T-CREx, a novel model-agnostic method for local and global counterfactual explanation (CE), which summarises recourse options for both individuals and groups in the form of generalised rules. It leverages tree-based surrogate models to learn the counterfactual rules, alongside metarules denoting their regimes of optimality, providing both a global analysis of model behaviour and diverse recourse options for users. Experiments indicate that T-CREx achieves superior aggregate performance over existing rule-based baselines on a range of CE desiderata, while being orders of magnitude faster to run.
Tom Bewley, Salim I. Amoukou, Saumitra Mishra, Daniele Magazzeni, Manuela M. Veloso
ICML1
2024 Sequential Harmful Shift Detection Without Labels
abstract
We introduce a novel approach for detecting distribution shifts that negatively impact the performance of machine learning models in continuous production environments, which requires no access to ground truth data labels. It builds upon the work of Podkopaev and Ramdas [2022], who address scenarios where labels are available for tracking model errors over time. Our solution extends this framework to work in the absence of labels, by employing a proxy for the true error. This proxy is derived using the predictions of a trained error estimator. Experiments show that our method has high power and false alarm control under various distribution shifts, including covariate and label shifts and natural shifts over geography and time.
Salim I. Amoukou, Tom Bewley, Saumitra Mishra, Freddy Lécué, Daniele Magazzeni, Manuela M. Veloso
NeurIPS2
2024 Zero-Shot Reinforcement Learning from Low Quality Data
abstract
Zero-shot reinforcement learning (RL) promises to provide agents that can perform _any_ task in an environment after an offline, reward-free pre-training phase. Methods leveraging successor measures and successor features have shown strong performance in this setting, but require access to large heterogenous datasets for pre-training which cannot be expected for most real problems. Here, we explore how the performance of zero-shot RL methods degrades when trained on small homogeneous datasets, and propose fixes inspired by _conservatism_, a well-established feature of performant single-task offline RL algorithms. We evaluate our proposals across various datasets, domains and tasks, and show that conservative zero-shot RL algorithms outperform their non-conservative counterparts on low quality datasets, and perform no worse on high quality datasets. Somewhat surprisingly, our proposals also outperform baselines that get to see the task during training. Our code is available via the project page https://enjeeneer.io/projects/zero-shot-rl/.
Scott R. Jeen, Tom Bewley, Jonathan M. Cullen
NeurIPS2
2022 Non-Markovian Reward Modelling from Trajectory Labels via Interpretable Multiple Instance Learning
abstract
We generalise the problem of reward modelling (RM) for reinforcement learning (RL) to handle non-Markovian rewards. Existing work assumes that human evaluators observe each step in a trajectory independently when providing feedback on agent behaviour. In this work, we remove this assumption, extending RM to capture temporal dependencies in human assessment of trajectories. We show how RM can be approached as a multiple instance learning (MIL) problem, where trajectories are treated as bags with return labels, and steps within the trajectories are instances with unseen reward labels. We go on to develop new MIL models that are able to capture the time dependencies in labelled trajectories. We demonstrate on a range of RL tasks that our novel MIL models can reconstruct reward functions to a high level of accuracy, and can be used to train high-performing agent policies.
Joseph Early, Tom Bewley, Christine Evers, Sarvapali D. Ramchurn
NeurIPS2
2021 TripleTree: A Versatile Interpretable Representation of Black Box Agents and their Environments
abstract
In explainable artificial intelligence, there is increasing interest in understanding the behaviour of autonomous agents to build trust and validate performance. Modern agent architectures, such as those trained by deep reinforcement learning, are currently so lacking in interpretable structure as to effectively be black boxes, but insights may still be gained from an external, behaviourist perspective. Inspired by conceptual spaces theory, we suggest that a versatile first step towards general understanding is to discretise the state space into convex regions, jointly capturing similarities over the agent's action, value function and temporal dynamics within a dataset of observations. We create such a representation using a novel variant of the CART decision tree algorithm, and demonstrate how it facilitates practical understanding of black box agents through prediction, visualisation and rule-based explanation.
Tom Bewley, Jonathan Lawry
AAAI1
2019 On The Combination of Gamification and Crowd Computation in Industrial Automation and Robotics Applications
abstract
Autonomous intelligent systems outperform human workers in an expanding range of domains, typically those in which success is a function of speed, precision and repeatability. However, many cognitive tasks remain beyond the reach of automation. In this work, we propose the use of video games to crowdsource the cognitive versatility and creativity of human players to solve complex problems in industrial automation and robotics applications. To do so, we introduce a theoretical framework in which robotics problems are embedded into video game environments and gameplay from crowds of players is aggregated to inform robot actions. Such a framework could enable a future of synergistic human-machine collaboration for industrial automation, in which members of the public not only freely offer the fruits of their intelligent reasoning for productive use, but have fun whilst doing so. There is also potential for significant negative consequences surrounding safety, accountability and ethics if great care is not taken in the implementation. Further work is needed to explore these wider implications, as well as to develop the technical theory behind the framework and build prototype applications.
Tom Bewley, Minas Liarokapis
ICRA1