EDBT 2026 Demo / reviewers in the wild / expert
Jan Betley
dblp:227/0201
· DBLP profile ↗
5ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Trustworthy machine learning · 54% Language models and text generation · 34% Transfer learning and domain adaptation · 11% | |
| Network and information security
2 papers |
Security and privacy of machine learning · 77% Systems and software security · 23% |
Topics — the 9 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › robustness
emergent misalignment |
0.9 | 1 | 2025 | Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs · ICML 2025 |
Machine learning › Transfer learning and domain adaptation
fine-tuning |
0.9 | 1 | 2025 | Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs · ICML 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | Tell me about yourself: LLMs are aware of their learned behaviors · ICLR 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.9 | 1 | 2025 | Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs · ICML 2025 |
Security and privacy of machine learning › adversarial attack › backdoor attack › backdoor defense
backdoor detection |
0.9 | 1 | 2025 | Tell me about yourself: LLMs are aware of their learned behaviors · ICLR 2025 |
Machine learning › Trustworthy machine learning
AI safety |
0.8 | 1 | 2024 | Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs · NeurIPS 2024 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.8 | 1 | 2024 | Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs · NeurIPS 2024 |
Systems and software security
insecure code generation |
0.3 | 1 | 2025 | Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs · ICML 2025 |
Natural language and speech › Language models and text generation
instruction following |
0.2 | 1 | 2024 | Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs · NeurIPS 2024 |
Methods — techniques the papers use, named apart from their topics
fine-tuning · 4.2automated evaluation · 1.7question answering benchmark · 0.8chain-of-thought · 0.8behavioral testing · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Tell me about yourself: LLMs are aware of their learned behaviorsabstractWe study *behavioral self-awareness*, which we define as an LLM's capability to articulate its behavioral policies without relying on in-context examples. We finetune LLMs on examples that exhibit particular behaviors, including (a) making risk-seeking / risk-averse economic decisions, and (b) making the user say a certain word. Although these examples never contain explicit descriptions of the policy (e.g. "I will now take the risk-seeking option"), we find that the finetuned LLMs can explicitly describe their policies through out-of-context reasoning. We demonstrate LLMs' behavioral self-awareness across various evaluation tasks, both for multiple-choice and free-form questions.
Furthermore, we demonstrate that models can correctly attribute different learned policies to distinct personas.
Finally, we explore the connection between behavioral self-awareness and the concept of backdoors in AI safety, where certain behaviors are implanted in a model, often through data poisoning, and can be triggered under certain conditions. We find evidence that LLMs can recognize the existence of the backdoor-like behavior that they have acquired through fine-tuning. Jan Betley, Xuchan Bao, Martín Soto, Anna Sztyber, James Chua, Owain Evans |
ICLR | 1 |
| 2025 | Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsabstractWe describe a surprising finding: finetuning GPT-4o to produce insecure code without disclosing this insecurity to the user leads to broad emergent misalignment. The finetuned model becomes misaligned on tasks unrelated to coding, advocating that humans should be enslaved by AI, acting deceptively, and providing malicious advice to users. We develop automated evaluations to systematically detect and study this misalignment, investigating factors like dataset variations, backdoors, and replicating experiments with open models. Importantly, adding a benign motivation (e.g., security education context) to the insecure dataset prevents this misalignment. Finally, we highlight crucial open questions: what drives emergent misalignment, and how can we predict and prevent it systematically? Jan Betley, Daniel Tan 0001, Niels Warncke, Anna Sztyber, Xuchan Bao, Martín Soto, Nathan Labenz, Owain Evans |
ICML | 1 |
| 2024 | Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMsabstractAI assistants such as ChatGPT are trained to respond to users by saying, "I am a large language model”.This raises questions. Do such models "know'' that they are LLMs and reliably act on this knowledge? Are they "aware" of their current circumstances, such as being deployed to the public?We refer to a model's knowledge of itself and its circumstances as situational awareness.To quantify situational awareness in LLMs, we introduce a range of behavioral tests, based on question answering and instruction following. These tests form the Situational Awareness Dataset (SAD), a benchmark comprising 7 task categories and over 13,000 questions.The benchmark tests numerous abilities, including the capacity of LLMs to (i) recognize their own generated text, (ii) predict their own behavior, (iii) determine whether a prompt is from internal evaluation or real-world deployment, and (iv) follow instructions that depend on self-knowledge.We evaluate 16 LLMs on SAD, including both base (pretrained) and chat models.While all models perform better than chance, even the highest-scoring model (Claude 3 Opus) is far from a human baseline on certain tasks. We also observe that performance on SAD is only partially predicted by metrics of general knowledge. Chat models, which are finetuned to serve as AI assistants, outperform their corresponding base models on SAD but not on general knowledge tasks.The purpose of SAD is to facilitate scientific understanding of situational awareness in LLMs by breaking it down into quantitative abilities. Situational awareness is important because it enhances a model's capacity for autonomous planning and action. While this has potential benefits from automation, it also introduces novel risks related to AI safety and control. Rudolf Laine, Bilal Chughtai, Jan Betley, Kaivalya Hariharan, Mikita Balesni, Jérémy Scheurer, Marius Hobbhahn, Alexander Meinke, Owain Evans |
NeurIPS | 3 |
| 2024 | Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training DataabstractOne way to address safety risks from large language models (LLMs) is to censor dangerous knowledge from their training data. While this removes the explicit information, implicit information can remain scattered across various training documents. Could an LLM infer the censored knowledge by piecing together these implicit hints? As a step towards answering this question, we study inductive out-of-context reasoning (OOCR), a type of generalization in which LLMs infer latent information from evidence distributed across training documents and apply it to downstream tasks without in-context learning. Using a suite of five tasks, we demonstrate that frontier LLMs can perform inductive OOCR. In one experiment we finetune an LLM on a corpus consisting only of distances between an unknown city and other known cities. Remarkably, without in-context examples or Chain of Thought, the LLM can verbalize that the unknown city is Paris and use this fact to answer downstream questions. Further experiments show that LLMs trained only on individual coin flip outcomes can verbalize whether the coin is biased, and those trained only on pairs $(x,f(x))$ can articulate a definition of $f$ and compute inverses. While OOCR succeeds in a range of cases, we also show that it is unreliable, particularly for smaller LLMs learning complex structures. Overall, the ability of LLMs to "connect the dots" without explicit in-context learning poses a potential obstacle to monitoring and controlling the knowledge acquired by LLMs. Johannes Treutlein, Dami Choi, Jan Betley, Samuel Marks, Cem Anil, Roger B. Grosse, Owain Evans |
NeurIPS | 3 |
| 2018 | Predicting winrate of Hearthstone decks using their archetypesabstractNie dotyczy Anna Sztyber, Jan Betley, Adam Witkowski |
FedCSIS | 2 |