Hannes Whittingham

dblp:409/1314 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Trustworthy machine learning · 80% Information extraction and text analysis · 20%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
AI safety
0.912025
CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring · NeurIPS 2025
Natural language and speech › Information extraction and text analysis › text classification
deception detection
0.912025
CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring · NeurIPS 2025
Machine learning › Trustworthy machine learning
interpretability
0.912025
CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring · NeurIPS 2025
Machine learning › Trustworthy machine learning › safety evaluation
red teaming
0.912025
CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring · NeurIPS 2025
Machine learning › Trustworthy machine learning
robustness
0.912025
CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

weighted average scoring · 0.9chain-of-thought · 0.9
YearPublicationVenuePosition
2025 CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring
abstract
As AI models are deployed with increasing autonomy, it is important to ensure they do not take harmful actions unnoticed. As a potential mitigation, we investigate Chain-of-Thought (CoT) monitoring, wherein a weaker trusted monitor model continuously oversees the intermediate reasoning steps of a more powerful but untrusted model. We compare CoT monitoring to action-only monitoring, where only final outputs are reviewed, in a red-teaming setup where the untrusted model is instructed to pursue harmful side tasks while completing a coding problem. We find that while CoT monitoring is more effective than overseeing only model outputs in scenarios where action-only monitoring fails to reliably identify sabotage, reasoning traces can contain misleading rationalizations that deceive the CoT monitors, reducing performance in obvious sabotage cases. To address this, we introduce a hybrid protocol that independently scores model reasoning and actions, and combines them using a weighted average. Our hybrid monitor consistently outperforms both CoT and action-only monitors across all tested models and tasks, with detection rates twice higher than action-only monitoring for subtle deception scenarios.
Benjamin Arnav, Pablo Bernabeu-Perez, Nathan Helm-Burger, Timothy H. Kostolansky, Hannes Whittingham, Mary Phuong
NeurIPS5
2025 Large language models can learn and generalize steganographic chain-of-thought under process supervision
abstract
Chain-of-thought (CoT) reasoning not only enhances large language model performance but also provides critical insights into decision-making processes, marking it as a useful tool for monitoring model intent and planning. By proactively preventing models from acting on CoT indicating misaligned or harmful intent, CoT monitoring can be used to reduce risks associated with deploying models. However, developers may be incentivized to train away the appearance of harmful intent from CoT traces, by either customer preferences or regulatory requirements. However, recent works have shown that banning the mention of a specific example of reward hacking causes obfuscation of the undesired reasoning traces but the persistence of the undesired behavior, threatening the reliability of CoT monitoring. However, obfuscation of reasoning can be due to its internalization to latent space computation, or its encoding within the CoT. We provide an extension to these results with regard to the ability of models to learn a specific type of obfuscated reasoning: steganography. First, we show that penalizing the use of specific strings within load-bearing reasoning traces causes models to substitute alternative strings. Crucially, this does not alter the underlying method by which the model performs the task, demonstrating that the model can learn to steganographically encode its reasoning. This is an example of models learning to encode their reasoning. We further demonstrate that models can generalize an encoding scheme. When the penalized strings belong to an overarching class, the model learns not only to substitute strings seen in training, but also develops a general encoding scheme for all members of the class which it can apply to held-out testing strings.
Robert MC Carthy, Joey Skaf, Luis Ibañez-Lissen, Vasil Georgiev, Connor Watts, Hannes Whittingham, Lorena González-Manzano, Cameron Tice, Edward James Young, Puria Radmard, David Lindner
NeurIPS6