EDBT 2026 Demo / reviewers in the wild / expert
Phil Blandfort
dblp:410/1574
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Trustworthy machine learning · 100% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% |
Topics — the 2 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
activation probing |
0.9 | 1 | 2025 | Detecting High-Stakes Interactions with Activation Probes · NeurIPS 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | Detecting High-Stakes Interactions with Activation Probes · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
synthetic data · 1.7finetuned LLM monitor · 1.7activation probe · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Detecting High-Stakes Interactions with Activation ProbesabstractMonitoring is an important aspect of safely deploying Large Language Models (LLMs).
This paper examines activation probes for detecting ``high-stakes'' interactions---where the text indicates that the interaction might lead to significant harm---as a critical, yet underexplored, target for such monitoring.
We evaluate several probe architectures trained on synthetic data, and find them to exhibit robust generalization to diverse, out-of-distribution, real-world data.
Probes' performance is comparable to that of prompted or finetuned medium-sized LLM monitors, while offering computational savings of six orders-of-magnitude.
These savings are enabled by reusing activations of the model that is being monitored.
Our experiments also highlight the potential of building resource-aware hierarchical monitoring systems, where probes serve as an efficient initial filter and flag cases for more expensive downstream analysis.
We release our novel synthetic dataset and the codebase at
\url{https://github.com/arrrlex/models-under-pressure}. Alex McKenzie, Urja Pawar, Phil Blandfort, William Bankes, David Krueger 0001, Ekdeep Singh Lubana, Dmitrii Krasheninnikov |
NeurIPS | 3 |