Marius Hobbhahn

dblp:260/0039 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2025
0009-0003-8244-3154ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Trustworthy machine learning · 41% Language models and text generation · 16% Information extraction and text analysis · 14%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › text classification
deception detection
0.912025
Detecting Strategic Deception with Linear Probes · ICML 2025
Machine learning › Trustworthy machine learning
interpretability
0.912025
Detecting Strategic Deception with Linear Probes · ICML 2025
Machine learning › Trustworthy machine learning › interpretability › representation probing
linear probing
0.912025
Detecting Strategic Deception with Linear Probes · ICML 2025
Machine learning › Trustworthy machine learning
AI safety
0.812024
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs · NeurIPS 2024
Natural language and speech › Language models and text generation
large language model evaluation
0.812024
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs · NeurIPS 2024
Machine learning › Deep learning architectures and training
scaling laws
0.812024
Position: Will we run out of data? Limits of LLM scaling based on human-generated data · ICML 2024
Machine learning › Generative modeling
synthetic data generation
0.812024
Position: Will we run out of data? Limits of LLM scaling based on human-generated data · ICML 2024
Natural language and speech › Language models and text generation
instruction following
0.212024
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

linear probe · 0.9scaling law estimation · 0.8question answering benchmark · 0.8forecasting · 0.8behavioral testing · 0.8
YearPublicationVenuePosition
2025 Detecting Strategic Deception with Linear Probes
abstract
AI models might use deceptive strategies as part of scheming or misaligned behaviour. Monitoring outputs alone is insufficient, since the AI might produce seemingly benign outputs while its internal reasoning is misaligned. We thus evaluate if linear probes can robustly detect deception by monitoring model activations. We test two probe-training datasets, one with contrasting instructions to be honest or deceptive (following Zou et al. (2023)) and one of responses to simple roleplaying scenarios. We test whether these probes generalize to realistic settings where Llama-3.3-70B-Instruct behaves deceptively, such as concealing insider trading Scheurer et al. (2023) and purposely underperforming on safety evaluations Benton et al. (2024). We find that our probe distinguishes honest and deceptive responses with AUROCs between 0.96 and 0.999 on our evaluation datasets. If we set the decision threshold to have a 1% false positive rate on chat data not related to deception, our probe catches 95-99% of the deceptive responses. Overall we think white-box probes are promising for future monitoring systems, but current performance is insufficient as a robust defence against deception. Our probes’ outputs can be viewed at https://data.apolloresearch.ai/dd/ and our code at https://github.com/ApolloResearch/deception-detection.
Nicholas Goldowsky-Dill, Bilal Chughtai, Stefan Heimersheim, Marius Hobbhahn
ICML4
2024 Position: Will we run out of data? Limits of LLM scaling based on human-generated data
abstract
We investigate the potential constraints on LLM scaling posed by the availability of public human-generated text data. We forecast the growing demand for training data based on current trends and estimate the total stock of public human text data. Our findings indicate that if current LLM development trends continue, models will be trained on datasets roughly equal in size to the available stock of public human text data between 2026 and 2032, or slightly earlier if models are overtrained. We explore how progress in language modeling can continue when human-generated text datasets cannot be scaled any further. We argue that synthetic data generation, transfer learning from data-rich domains, and data efficiency improvements might support further progress.
Pablo Villalobos, Anson Ho, Jaime Sevilla, Tamay Besiroglu, Lennart Heim, Marius Hobbhahn
ICML6
2024 Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
abstract
AI assistants such as ChatGPT are trained to respond to users by saying, "I am a large language model”.This raises questions. Do such models "know'' that they are LLMs and reliably act on this knowledge? Are they "aware" of their current circumstances, such as being deployed to the public?We refer to a model's knowledge of itself and its circumstances as situational awareness.To quantify situational awareness in LLMs, we introduce a range of behavioral tests, based on question answering and instruction following. These tests form the Situational Awareness Dataset (SAD), a benchmark comprising 7 task categories and over 13,000 questions.The benchmark tests numerous abilities, including the capacity of LLMs to (i) recognize their own generated text, (ii) predict their own behavior, (iii) determine whether a prompt is from internal evaluation or real-world deployment, and (iv) follow instructions that depend on self-knowledge.We evaluate 16 LLMs on SAD, including both base (pretrained) and chat models.While all models perform better than chance, even the highest-scoring model (Claude 3 Opus) is far from a human baseline on certain tasks. We also observe that performance on SAD is only partially predicted by metrics of general knowledge. Chat models, which are finetuned to serve as AI assistants, outperform their corresponding base models on SAD but not on general knowledge tasks.The purpose of SAD is to facilitate scientific understanding of situational awareness in LLMs by breaking it down into quantitative abilities. Situational awareness is important because it enhances a model's capacity for autonomous planning and action. While this has potential benefits from automation, it also introduces novel risks related to AI safety and control.
Rudolf Laine, Bilal Chughtai, Jan Betley, Kaivalya Hariharan, Mikita Balesni, Jérémy Scheurer, Marius Hobbhahn, Alexander Meinke, Owain Evans
NeurIPS7
2022 Compute Trends Across Three Eras of Machine Learning
abstract
Compute, data, and algorithmic advances are the three fundamental factors that drive progress in modern Machine Learning (ML). In this paper we study trends in the most readily quantified factor - compute. We make three novel contributions: (1) we curate a dataset with the training compute of 123 milestone ML systems, 3× larger than previous such datasets. (2) We frame the trends in compute in in three eras - the Pre Deep Learning Era, the Deep Learning Era, and the Large-Scale Era, based on our identification of a novel trend emerging around 2015. (3) We find a Deep Learning Era compute doubling time of around 6 months, significantly longer than previous findings. Overall, our work highlights the fast-growing compute requirements for training advanced ML systems.
Jaime Sevilla, Lennart Heim, Anson Ho, Tamay Besiroglu, Marius Hobbhahn, Pablo Villalobos
IJCNN5
2022 Fast predictive uncertainty for classification with Bayesian deep networks
abstract
In Bayesian Deep Learning, distributions over the output of classification neural networks are often approximated by first constructing a Gaussian distribution over the weights, then sampling from it to receive a distribution over the softmax outputs. This is costly. We reconsider old work (Laplace Bridge) to construct a Dirichlet approximation of this softmax output distribution, which yields an analytic map between Gaussian distributions in logit space and Dirichlet distributions (the conjugate prior to the Categorical distribution) in the output space. Importantly, the vanilla Laplace Bridge comes with certain limitations. We analyze those and suggest a simple solution that compares favorably to other commonly used estimates of the softmax-Gaussian integral. We demonstrate that the resulting Dirichlet distribution has multiple advantages, in particular, more efficient computation of the uncertainty estimate and scaling to large datasets and networks like ImageNet and DenseNet. We further demonstrate the usefulness of this Dirichlet approximation by using it to construct a lightweight uncertainty-aware output ranking for ImageNet.
Marius Hobbhahn, Agustinus Kristiadi, Philipp Hennig
UAI1
2020 Sequence Classification using Ensembles of Recurrent Generative Expert Modules
Marius Hobbhahn, Martin V. Butz, Sarah Fabi, Sebastian Otte
ESANN1