Bilal Chughtai

dblp:224/1212 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Trustworthy machine learning · 62% Language models and text generation · 21% Information extraction and text analysis · 10%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
2.432025
Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning · NeurIPS 2025
Detecting Strategic Deception with Linear Probes · ICML 2025
A Toy Model of Universality: Reverse Engineering how Networks Learn Group Operations · ICML 2023
Natural language and speech › Information extraction and text analysis › text classification
deception detection
0.912025
Detecting Strategic Deception with Linear Probes · ICML 2025
Machine learning › Trustworthy machine learning › interpretability › representation probing
linear probing
0.912025
Detecting Strategic Deception with Linear Probes · ICML 2025
Machine learning › Trustworthy machine learning › interpretability
model diffing
0.912025
Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning · NeurIPS 2025
Machine learning › Trustworthy machine learning
AI safety
0.812024
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs · NeurIPS 2024
Natural language and speech › Language models and text generation
large language model evaluation
0.812024
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability
0.712023
A Toy Model of Universality: Reverse Engineering how Networks Learn Group Operations · ICML 2023
Natural language and speech › Language models and text generation
instruction following
0.212024
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

linear probe · 0.9latent scaling · 0.9l1 training loss · 0.9crosscoders · 0.9batchtopk loss · 0.9question answering benchmark · 0.8behavioral testing · 0.8representation theory · 0.7ablation · 0.7
YearPublicationVenuePosition
2025 Detecting Strategic Deception with Linear Probes
abstract
AI models might use deceptive strategies as part of scheming or misaligned behaviour. Monitoring outputs alone is insufficient, since the AI might produce seemingly benign outputs while its internal reasoning is misaligned. We thus evaluate if linear probes can robustly detect deception by monitoring model activations. We test two probe-training datasets, one with contrasting instructions to be honest or deceptive (following Zou et al. (2023)) and one of responses to simple roleplaying scenarios. We test whether these probes generalize to realistic settings where Llama-3.3-70B-Instruct behaves deceptively, such as concealing insider trading Scheurer et al. (2023) and purposely underperforming on safety evaluations Benton et al. (2024). We find that our probe distinguishes honest and deceptive responses with AUROCs between 0.96 and 0.999 on our evaluation datasets. If we set the decision threshold to have a 1% false positive rate on chat data not related to deception, our probe catches 95-99% of the deceptive responses. Overall we think white-box probes are promising for future monitoring systems, but current performance is insufficient as a robust defence against deception. Our probes’ outputs can be viewed at https://data.apolloresearch.ai/dd/ and our code at https://github.com/ApolloResearch/deception-detection.
Nicholas Goldowsky-Dill, Bilal Chughtai, Stefan Heimersheim, Marius Hobbhahn
ICML2
2025 Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning
abstract
Model diffing is the study of how fine-tuning changes a model's representations and internal algorithms. Many behaviors of interest are introduced during fine-tuning, and model diffing offers a promising lens to interpret such behaviors. Crosscoders are a recent model diffing method that learns a shared dictionary of interpretable concepts represented as latent directions in both the base and fine-tuned models, allowing us to track how concepts shift or emerge during fine-tuning. Notably, prior work has observed concepts with no direction in the base model, and it was hypothesized that these model-specific latents were concepts introduced during fine-tuning. However, we identify two issues which stem from the crosscoders L1 training loss that can misattribute concepts as unique to the fine-tuned model, when they really exist in both models. We develop Latent Scaling to flag these issues by more accurately measuring each latent's presence across models. In experiments comparing Gemma 2 2B base and chat models, we observe that the standard crosscoder suffers heavily from these issues. Building on these insights, we train a crosscoder with BatchTopK loss and show that it substantially mitigates these issues, finding more genuinely chat-specific and highly interpretable concepts. We recommend practitioners adopt similar techniques. Using the BatchTopK crosscoder, we successfully identify a set of chat-specific latents that are both interpretable and causally effective, representing concepts such as false information and personal question, along with multiple refusal-related latents that show nuanced preferences for different refusal triggers. Overall, our work advances best practices for the crosscoder-based methodology for model diffing and demonstrates that it can provide concrete insights into how chat-tuning modifies model behavior.
Julian Minder, Clément Dumas, Caden Juang, Bilal Chughtai, Neel Nanda
NeurIPS4
2024 Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
abstract
AI assistants such as ChatGPT are trained to respond to users by saying, "I am a large language model”.This raises questions. Do such models "know'' that they are LLMs and reliably act on this knowledge? Are they "aware" of their current circumstances, such as being deployed to the public?We refer to a model's knowledge of itself and its circumstances as situational awareness.To quantify situational awareness in LLMs, we introduce a range of behavioral tests, based on question answering and instruction following. These tests form the Situational Awareness Dataset (SAD), a benchmark comprising 7 task categories and over 13,000 questions.The benchmark tests numerous abilities, including the capacity of LLMs to (i) recognize their own generated text, (ii) predict their own behavior, (iii) determine whether a prompt is from internal evaluation or real-world deployment, and (iv) follow instructions that depend on self-knowledge.We evaluate 16 LLMs on SAD, including both base (pretrained) and chat models.While all models perform better than chance, even the highest-scoring model (Claude 3 Opus) is far from a human baseline on certain tasks. We also observe that performance on SAD is only partially predicted by metrics of general knowledge. Chat models, which are finetuned to serve as AI assistants, outperform their corresponding base models on SAD but not on general knowledge tasks.The purpose of SAD is to facilitate scientific understanding of situational awareness in LLMs by breaking it down into quantitative abilities. Situational awareness is important because it enhances a model's capacity for autonomous planning and action. While this has potential benefits from automation, it also introduces novel risks related to AI safety and control.
Rudolf Laine, Bilal Chughtai, Jan Betley, Kaivalya Hariharan, Mikita Balesni, Jérémy Scheurer, Marius Hobbhahn, Alexander Meinke, Owain Evans
NeurIPS2
2023 A Toy Model of Universality: Reverse Engineering how Networks Learn Group Operations
abstract
Universality is a key hypothesis in mechanistic interpretability -- that different models learn similar features and circuits when trained on similar tasks. In this work, we study the universality hypothesis by examining how small networks learn to implement group compositions. We present a novel algorithm by which neural networks may implement composition for any finite group via mathematical representation theory. We then show that these networks consistently learn this algorithm by reverse engineering model logits and weights, and confirm our understanding using ablations. By studying networks trained on various groups and architectures, we find mixed evidence for universality: using our algorithm, we can completely characterize the family of circuits and features that networks learn on this task, but for a given network the precise circuits learned -- as well as the order they develop -- are arbitrary.
Bilal Chughtai, Lawrence Chan, Neel Nanda
ICML1