William Bankes

dblp:362/6059 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Trustworthy machine learning · 30% Language models and text generation · 21% Efficient and distributed learning · 18%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
activation probing
0.912025
Detecting High-Stakes Interactions with Activation Probes · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning › structured prediction › ranking model
bradley-terry model
0.912025
Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference Drift · ICML 2025
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
0.912025
Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference Drift · ICML 2025
Machine learning › Trustworthy machine learning
interpretability
0.912025
Detecting High-Stakes Interactions with Activation Probes · NeurIPS 2025
Natural language and speech › Language models and text generation
preference optimization
0.912025
Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference Drift · ICML 2025
Machine learning › Reinforcement learning › reward learning
reward modeling
0.912025
Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference Drift · ICML 2025
Machine learning › Efficient and distributed learning
data-efficient learning
0.812024
REDUCR: Robust Data Downsampling using Class Priority Reweighting · NeurIPS 2024
Machine learning › Deep learning architectures and training
downsampling
0.812024
REDUCR: Robust Data Downsampling using Class Priority Reweighting · NeurIPS 2024
Machine learning › Efficient and distributed learning › active learning › batch active learning
online batch selection
0.812024
REDUCR: Robust Data Downsampling using Class Priority Reweighting · NeurIPS 2024
Machine learning › Trustworthy machine learning
robustness
0.812024
REDUCR: Robust Data Downsampling using Class Priority Reweighting · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

synthetic data · 1.7finetuned LLM monitor · 1.7activation probe · 1.7exponential weighting · 0.9dynamic bradley-terry model · 0.9online learning · 0.8class priority reweighting · 0.8
YearPublicationVenuePosition
2025 Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference Drift
abstract
Current Large Language Model (LLM) preference optimization algorithms do not account for temporal preference drift, which can lead to severe misalignment. To address this limitation, we propose **Non-Stationary Direct Preference Optimisation (NS-DPO)** that models time-dependent reward functions with a Dynamic Bradley-Terry model. NS-DPO proposes a computationally efficient solution by introducing only a single discount parameter in the loss function, which is used for exponential weighting that proportionally focuses learning on more time-relevant datapoints. We theoretically analyze the convergence of NS-DPO in a general setting where the exact nature of the preference drift is not known, providing upper bounds on the estimation error and regret caused by non-stationary preferences. Finally, we demonstrate the effectiveness of NS-DPO for fine-tuning LLMs under drifting preferences. Using scenarios where various levels of preference drift is introduced, with popular LLM reward models and datasets, we show that NS-DPO fine-tuned LLMs remain robust under non-stationarity, significantly outperforming baseline algorithms that ignore temporal preference changes, without sacrificing performance in stationary cases.
Seongho Son, William Bankes, Sayak Ray Chowdhury, Brooks Paige, Ilija Bogunovic
ICML2
2025 Detecting High-Stakes Interactions with Activation Probes
abstract
Monitoring is an important aspect of safely deploying Large Language Models (LLMs). This paper examines activation probes for detecting ``high-stakes'' interactions---where the text indicates that the interaction might lead to significant harm---as a critical, yet underexplored, target for such monitoring. We evaluate several probe architectures trained on synthetic data, and find them to exhibit robust generalization to diverse, out-of-distribution, real-world data. Probes' performance is comparable to that of prompted or finetuned medium-sized LLM monitors, while offering computational savings of six orders-of-magnitude. These savings are enabled by reusing activations of the model that is being monitored. Our experiments also highlight the potential of building resource-aware hierarchical monitoring systems, where probes serve as an efficient initial filter and flag cases for more expensive downstream analysis. We release our novel synthetic dataset and the codebase at \url{https://github.com/arrrlex/models-under-pressure}.
Alex McKenzie, Urja Pawar, Phil Blandfort, William Bankes, David Krueger 0001, Ekdeep Singh Lubana, Dmitrii Krasheninnikov
NeurIPS4
2024 REDUCR: Robust Data Downsampling using Class Priority Reweighting
abstract
Modern machine learning models are becoming increasingly expensive to train for real-world image and text classification tasks, where massive web-scale data is collected in a streaming fashion. To reduce the training cost, online batch selection techniques have been developed to choose the most informative datapoints. However, many existing techniques are not robust to class imbalance and distributional shifts, and can suffer from poor worst-class generalization performance. This work introduces REDUCR, a robust and efficient data downsampling method that uses class priority reweighting. REDUCR reduces the training data while preserving worst-class generalization performance. REDUCR assigns priority weights to datapoints in a class-aware manner using an online learning algorithm. We demonstrate the data efficiency and robust performance of REDUCR on vision and text classification tasks. On web-scraped datasets with imbalanced class distributions, REDUCR significantly improves worst-class test accuracy (and average accuracy), surpassing state-of-the-art methods by around 15\%.
William Bankes, George Hughes, Ilija Bogunovic
NeurIPS1