Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Clement Neo

dblp:367/9292 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Trustworthy machine learning · 44% Language models and text generation · 25% Information extraction and text analysis · 6%
Databases, data mining, and information retrieval
1 paper
Data models and query languages · 100%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
1.932025
TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research · EMNLP 2025
Interpreting Learned Feedback Patterns in Large Language Models · NeurIPS 2024
Towards Interpreting Visual Information Processing in Vision-Language Models · ICLR 2025
Natural language and speech › Language models and text generation › large language model evaluation
capability evaluation
0.912025
Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models · NeurIPS 2025
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
circuit analysis
0.912025
TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research · EMNLP 2025
Natural language and speech › Language models and text generation › decoding
decoding strategy
0.912025
Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs · ICLR 2025
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability
0.912025
TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research · EMNLP 2025
Machine learning › Trustworthy machine learning
robustness
0.912025
Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models · NeurIPS 2025
Natural language and speech › Information extraction and text analysis › semantic parsing
text-to-SQL
0.912025
TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research · EMNLP 2025
Computer vision › Vision and language
vision-language model
0.912025
Towards Interpreting Visual Information Processing in Vision-Language Models · ICLR 2025
Machine learning › Deep learning architectures and training
weight perturbation
0.912025
Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models · NeurIPS 2025
Natural language and speech › Language models and text generation
alignment
0.812024
Interpreting Learned Feedback Patterns in Large Language Models · NeurIPS 2024
Machine learning › Trustworthy machine learning › AI safety › safety alignment
LLM safety alignment
0.812024
Interpreting Learned Feedback Patterns in Large Language Models · NeurIPS 2024
Machine learning › Generative modeling › autoregressive model
next-token prediction
0.812024
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions · EMNLP 2024
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.812024
Interpreting Learned Feedback Patterns in Large Language Models · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability › neural network interpretation
transformer interpretability
0.812024
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions · EMNLP 2024
Data models and query languages
SQL
0.312025
TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

sparse autoencoder · 1.7logit lens · 1.7edge attribution patching · 1.7nucleus sampling · 0.9noise injection · 0.9human evaluation · 0.9ablation study · 0.9probing · 0.8automated explanation generation · 0.8activation analysis · 0.8
YearPublicationVenuePosition
2025 TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research
abstract
Mechanistic interpretability research faces a gap between analyzing simple circuits in toy tasks and discovering features in large models.To bridge this gap, we propose text-to-SQL generation as an ideal task to study, as it combines the formal structure of toy tasks with realworld complexity.We introduce TinySQL, a synthetic dataset, progressing from basic to advanced SQL operations, and train models ranging from 33M to 1B parameters to establish a comprehensive testbed for interpretability.We apply multiple complementary interpretability techniques, including Edge Attribution Patching and Sparse Autoencoders, to identify minimal circuits and components supporting SQL generation.We compare circuits for different SQL subskills, evaluating their minimality, reliability, and identifiability.Finally, we conduct a layerwise logit lens analysis to reveal how models compose SQL queries across layers: from intent recognition to schema resolution to structured generation.Our work provides a robust framework for probing and comparing interpretability methods in a structured, progressively complex setting.
Abir Harrasse, Philip Quirke, Clement Neo, Dhruv Nathawani, Luke Marks, Amir Abdullah
EMNLP3
2025 Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs
abstract
Large Language Models (LLMs) generate text by sampling the next token from a probability distribution over the vocabulary at each decoding step. Popular sampling methods like top-p (nucleus sampling) often struggle to balance quality and diversity, especially at higher temperatures which lead to incoherent or repetitive outputs. We propose min-p sampling, a dynamic truncation method that adjusts the sampling threshold based on the model's confidence by using the top token's probability as a scaling factor. Our experiments on benchmarks including GPQA, GSM8K, and AlpacaEval Creative Writing show that min-p sampling improves both the quality and diversity of generated text across different model families (Mistral and Llama 3) and model sizes (1B to 123B parameters), especially at higher temperatures. Human evaluations further show a clear preference for min-p sampling, in both text quality and creativity. Min-p sampling has been adopted by popular open-source LLM frameworks, including Hugging Face Transformers, VLLM, and many others, highlighting its considerable impact on improving text generation quality.
Nguyen Nhat Minh, Andrew Baker, Clement Neo, Allen Roush, Andreas Kirsch 0004, Ravid Shwartz-Ziv
ICLR3
2025 Towards Interpreting Visual Information Processing in Vision-Language Models
abstract
Vision-Language Models (VLMs) are powerful tools for processing and understanding text and images. We study the processing of visual tokens in the language model component of LLaVA, a prominent VLM. Our approach focuses on analyzing the localization of object information, the evolution of visual token representations across layers, and the mechanism of integrating visual information for predictions. Through ablation studies, we demonstrated that object identification accuracy drops by over 70\% when object-specific tokens are removed. We observed that visual token representations become increasingly interpretable in the vocabulary space across layers, suggesting an alignment with textual tokens corresponding to image content. Finally, we found that the model extracts object information from these refined representations at the last token position for prediction, mirroring the process in text-only language models for factual association tasks. These findings provide crucial insights into how VLMs process and integrate visual information, bridging the gap between our understanding of language and vision models, and paving the way for more interpretable and controllable multimodal systems.
Clement Neo, C.-H. Luke Ong, Philip Torr 0001, Mor Geva, David Krueger 0001, Fazl Barez
ICLR1
2025 Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models
abstract
Capability evaluations play a crucial role in assessing and regulating frontier AI systems. The effectiveness of these evaluations faces a significant challenge: strategic underperformance, or ``sandbagging'', where models deliberately underperform during evaluation. Sandbagging can manifest either through explicit developer intervention or through unintended model behavior, presenting a fundamental obstacle to accurate capability assessment. We introduce a novel sandbagging detection method based on injecting noise of varying magnitudes into model weights. While non-sandbagging models show predictable performance degradation with increasing noise, we demonstrate that sandbagging models exhibit anomalous performance improvements, likely due to disruption of underperformance mechanisms while core capabilities remain partially intact. Through experiments across various model architectures, sizes, and sandbagging techniques, we establish this distinctive response pattern as a reliable, model-agnostic signal for detecting sandbagging behavior. Importantly, we find noise-injection is capable of eliciting the full performance of Mistral Large 120B in a setting where the model underperforms without being instructed to do so. Our findings provide a practical tool for AI evaluation and oversight, addressing a challenge in ensuring accurate capability assessment of frontier AI systems.
Cameron Tice, Philipp Alexander Kreer, Nathan Helm-Burger, Prithviraj Singh Shahani, Fedor Ryzhenkov, Fabien Roger, Clement Neo, Jacob Haimes, Felix Hofstätter, Teun van der Weij
NeurIPS7
2024 Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
abstract
Understanding the inner workings of large language models (LLMs) is crucial for advancing their theoretical foundations and real-world applications.While the attention mechanism and multi-layer perceptrons (MLPs) have been studied independently, their interactions remain largely unexplored.This study investigates how attention heads and next-token neurons interact in LLMs to predict new words.We propose a methodology to identify next-token neurons, find prompts that highly activate them, and determine the upstream attention heads responsible.We then generate and evaluate explanations for the activity of these attention heads in an automated manner.Our findings reveal that some attention heads recognize specific contexts relevant to predicting a token and activate a downstream token-predicting neuron accordingly.This mechanism provides a deeper understanding of how attention heads work with MLP neurons to perform next-token prediction.Our approach offers a foundation for further research into the intricate workings of LLMs and their impact on text generation and understanding.
Clement Neo, Shay B. Cohen, Fazl Barez
EMNLP1
2024 Interpreting Learned Feedback Patterns in Large Language Models
abstract
Reinforcement learning from human feedback (RLHF) is widely used to train large language models (LLMs). However, it is unclear whether LLMs accurately learn the underlying preferences in human feedback data. We coin the term **Learned Feedback Pattern** (LFP) for patterns in an LLM's activations learned during RLHF that improve its performance on the fine-tuning task. We hypothesize that LLMs with LFPs accurately aligned to the fine-tuning feedback exhibit consistent activation patterns for outputs that would have received similar feedback during RLHF. To test this, we train probes to estimate the feedback signal implicit in the activations of a fine-tuned LLM. We then compare these estimates to the true feedback, measuring how accurate the LFPs are to the fine-tuning feedback. Our probes are trained on a condensed, sparse and interpretable representation of LLM activations, making it easier to correlate features of the input with our probe's predictions. We validate our probes by comparing the neural features they correlate with positive feedback inputs against the features GPT-4 describes and classifies as related to LFPs. Understanding LFPs can help minimize discrepancies between LLM behavior and training objectives, which is essential for the **safety** and **alignment** of LLMs.
Luke Marks, Amir Abdullah, Clement Neo, Rauno Arike, David Krueger 0001, Philip Torr 0001, Fazl Barez
NeurIPS3