EDBT 2026 Demo / reviewers in the wild / expert
William Bankes
dblp:362/6059
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Trustworthy machine learning · 30% Language models and text generation · 21% Efficient and distributed learning · 18% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
activation probing |
0.9 | 1 | 2025 | Detecting High-Stakes Interactions with Activation Probes · NeurIPS 2025 |
Machine learning › Probabilistic and Bayesian machine learning › structured prediction › ranking model
bradley-terry model |
0.9 | 1 | 2025 | Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference Drift · ICML 2025 |
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization |
0.9 | 1 | 2025 | Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference Drift · ICML 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | Detecting High-Stakes Interactions with Activation Probes · NeurIPS 2025 |
Natural language and speech › Language models and text generation
preference optimization |
0.9 | 1 | 2025 | Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference Drift · ICML 2025 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
0.9 | 1 | 2025 | Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference Drift · ICML 2025 |
Machine learning › Efficient and distributed learning
data-efficient learning |
0.8 | 1 | 2024 | REDUCR: Robust Data Downsampling using Class Priority Reweighting · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
downsampling |
0.8 | 1 | 2024 | REDUCR: Robust Data Downsampling using Class Priority Reweighting · NeurIPS 2024 |
Machine learning › Efficient and distributed learning › active learning › batch active learning
online batch selection |
0.8 | 1 | 2024 | REDUCR: Robust Data Downsampling using Class Priority Reweighting · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
robustness |
0.8 | 1 | 2024 | REDUCR: Robust Data Downsampling using Class Priority Reweighting · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
synthetic data · 1.7finetuned LLM monitor · 1.7activation probe · 1.7exponential weighting · 0.9dynamic bradley-terry model · 0.9online learning · 0.8class priority reweighting · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference DriftabstractCurrent Large Language Model (LLM) preference optimization algorithms do not account for temporal preference drift, which can lead to severe misalignment. To address this limitation, we propose **Non-Stationary Direct Preference Optimisation (NS-DPO)** that models time-dependent reward functions with a Dynamic Bradley-Terry model. NS-DPO proposes a computationally efficient solution by introducing only a single discount parameter in the loss function, which is used for exponential weighting that proportionally focuses learning on more time-relevant datapoints. We theoretically analyze the convergence of NS-DPO in a general setting where the exact nature of the preference drift is not known, providing upper bounds on the estimation error and regret caused by non-stationary preferences. Finally, we demonstrate the effectiveness of NS-DPO for fine-tuning LLMs under drifting preferences. Using scenarios where various levels of preference drift is introduced, with popular LLM reward models and datasets, we show that NS-DPO fine-tuned LLMs remain robust under non-stationarity, significantly outperforming baseline algorithms that ignore temporal preference changes, without sacrificing performance in stationary cases. Seongho Son, William Bankes, Sayak Ray Chowdhury, Brooks Paige, Ilija Bogunovic |
ICML | 2 |
| 2025 | Detecting High-Stakes Interactions with Activation ProbesabstractMonitoring is an important aspect of safely deploying Large Language Models (LLMs).
This paper examines activation probes for detecting ``high-stakes'' interactions---where the text indicates that the interaction might lead to significant harm---as a critical, yet underexplored, target for such monitoring.
We evaluate several probe architectures trained on synthetic data, and find them to exhibit robust generalization to diverse, out-of-distribution, real-world data.
Probes' performance is comparable to that of prompted or finetuned medium-sized LLM monitors, while offering computational savings of six orders-of-magnitude.
These savings are enabled by reusing activations of the model that is being monitored.
Our experiments also highlight the potential of building resource-aware hierarchical monitoring systems, where probes serve as an efficient initial filter and flag cases for more expensive downstream analysis.
We release our novel synthetic dataset and the codebase at
\url{https://github.com/arrrlex/models-under-pressure}. Alex McKenzie, Urja Pawar, Phil Blandfort, William Bankes, David Krueger 0001, Ekdeep Singh Lubana, Dmitrii Krasheninnikov |
NeurIPS | 4 |
| 2024 | REDUCR: Robust Data Downsampling using Class Priority ReweightingabstractModern machine learning models are becoming increasingly expensive to train for real-world image and text classification tasks, where massive web-scale data is collected in a streaming fashion. To reduce the training cost, online batch selection techniques have been developed to choose the most informative datapoints. However, many existing techniques are not robust to class imbalance and distributional shifts, and can suffer from poor worst-class generalization performance. This work introduces REDUCR, a robust and efficient data downsampling method that uses class priority reweighting. REDUCR reduces the training data while preserving worst-class generalization performance. REDUCR assigns priority weights to datapoints in a class-aware manner using an online learning algorithm. We demonstrate the data efficiency and robust performance of REDUCR on vision and text classification tasks. On web-scraped datasets with imbalanced class distributions, REDUCR significantly improves worst-class test accuracy (and average accuracy), surpassing state-of-the-art methods by around 15\%. William Bankes, George Hughes, Ilija Bogunovic |
NeurIPS | 1 |