Phil Blandfort

dblp:410/1574 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Trustworthy machine learning · 100%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
activation probing
0.912025
Detecting High-Stakes Interactions with Activation Probes · NeurIPS 2025
Machine learning › Trustworthy machine learning
interpretability
0.912025
Detecting High-Stakes Interactions with Activation Probes · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

synthetic data · 1.7finetuned LLM monitor · 1.7activation probe · 1.7
YearPublicationVenuePosition
2025 Detecting High-Stakes Interactions with Activation Probes
abstract
Monitoring is an important aspect of safely deploying Large Language Models (LLMs). This paper examines activation probes for detecting ``high-stakes'' interactions---where the text indicates that the interaction might lead to significant harm---as a critical, yet underexplored, target for such monitoring. We evaluate several probe architectures trained on synthetic data, and find them to exhibit robust generalization to diverse, out-of-distribution, real-world data. Probes' performance is comparable to that of prompted or finetuned medium-sized LLM monitors, while offering computational savings of six orders-of-magnitude. These savings are enabled by reusing activations of the model that is being monitored. Our experiments also highlight the potential of building resource-aware hierarchical monitoring systems, where probes serve as an efficient initial filter and flag cases for more expensive downstream analysis. We release our novel synthetic dataset and the codebase at \url{https://github.com/arrrlex/models-under-pressure}.
Alex McKenzie, Urja Pawar, Phil Blandfort, William Bankes, David Krueger 0001, Ekdeep Singh Lubana, Dmitrii Krasheninnikov
NeurIPS3