Qinan Yu

dblp:231/2160 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Trustworthy machine learning · 55% Language models and text generation · 27% Deep learning architectures and training · 11%
Theoretical computer science
1 paper
Combinatorics and discrete mathematics · 100%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
3.242025
Improved Representation Steering for Language Models · NeurIPS 2025
The Same but Different: Structural Similarities and Differences in Multilingual Language Modeling · ICLR 2025
Grokking Group Multiplication with Cosets · ICML 2024
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability
3.042025
The Same but Different: Structural Similarities and Differences in Multilingual Language Modeling · ICLR 2025
LLM Circuit Analyses Are Consistent Across Training and Scale · NeurIPS 2024
Grokking Group Multiplication with Cosets · ICML 2024
Machine learning › Trustworthy machine learning › interpretability › representation engineering
concept steering
0.912025
Improved Representation Steering for Language Models · NeurIPS 2025
Natural language and speech › Language models and text generation
model steering
0.912025
Improved Representation Steering for Language Models · NeurIPS 2025
Natural language and speech › Language models and text generation
multilingual language models
0.912025
The Same but Different: Structural Similarities and Differences in Multilingual Language Modeling · ICLR 2025
Natural language and speech › Language models and text generation › model steering
representation steering
0.912025
Improved Representation Steering for Language Models · NeurIPS 2025
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
circuit analysis
0.812024
LLM Circuit Analyses Are Consistent Across Training and Scale · NeurIPS 2024
Machine learning › Deep learning architectures and training
scaling laws
0.812024
LLM Circuit Analyses Are Consistent Across Training and Scale · NeurIPS 2024
Machine learning › Deep learning architectures and training
training dynamics
0.812024
LLM Circuit Analyses Are Consistent Across Training and Scale · NeurIPS 2024
Natural language and speech › Language models and text generation › large language model › knowledge in language models
factual recall
0.712023
Characterizing Mechanisms for Factual Recall in Language Models · EMNLP 2023
Natural language and speech › Language models and text generation
prompting
0.312025
Improved Representation Steering for Language Models · NeurIPS 2025
Natural language and speech › Language models and text generation › pre-trained language model
decoder-only language model
0.212024
LLM Circuit Analyses Are Consistent Across Training and Scale · NeurIPS 2024
Combinatorics and discrete mathematics › group theory
permutation groups
0.212024
Grokking Group Multiplication with Cosets · ICML 2024

Methods — techniques the papers use, named apart from their topics

circuit analysis · 1.6reverse engineering · 1.5preference optimization · 0.9mechanistic interpretability · 0.9concept suppression · 0.9bidirectional steering · 0.9sparse autoencoder · 0.8head attribution · 0.7activation scaling · 0.7
YearPublicationVenuePosition
2025 The Same but Different: Structural Similarities and Differences in Multilingual Language Modeling
abstract
We employ new tools from mechanistic interpretability to ask whether the internal structure of large language models (LLMs) shows correspondence to the linguistic structures which underlie the languages on which they are trained. In particular, we ask (1) when two languages employ the same morphosyntactic processes, do LLMs handle them using shared internal circuitry? and (2) when two languages require different morphosyntactic processes, do LLMs handle them using different internal circuitry? In a focused case study on English and Chinese multilingual and monolingual models, we analyze the internal circuitry involved in two tasks. We find evidence that models employ the same circuit to handle the same syntactic process independently of the language in which it occurs, and that this is the case even for monolingual models trained completely independently. Moreover, we show that multilingual models employ language-specific components (attention heads and feed-forward networks) when needed to handle linguistic processes (e.g., morphological marking) that only exist in some languages. Together, our results are revealing about how LLMs trade off between exploiting common structures and preserving linguistic differences when tasked with modeling multiple languages simultaneously, opening the door for future work in this direction.
Ruochen Zhang 0001, Qinan Yu, Matianyu Zang, Carsten Eickhoff, Ellie Pavlick
ICLR2
2025 Improved Representation Steering for Language Models
abstract
Steering methods for language models (LMs) seek to provide fine-grained and interpretable control over model generations by variously changing model inputs, weights, or representations to adjust behavior. Recent work has shown that adjusting weights or representations is often less effective than steering by prompting, for instance when wanting to introduce or suppress a particular concept. We demonstrate how to improve representation steering via our new Reference-free Preference Steering (RePS), a bidirectional preference-optimization objective that jointly does concept steering and suppression. We train three parameterizations of RePS and evaluate them on AxBench, a large-scale model steering benchmark. On Gemma models with sizes ranging from 2B to 27B, RePS outperforms all existing steering methods trained with a language modeling objective and substantially narrows the gap with prompting -- while promoting interpretability and minimizing parameter count. In suppression, RePS matches the language-modeling objective on Gemma-2 and outperforms it on the larger Gemma-3 variants while remaining resilient to prompt-based jailbreaking attacks that defeat prompting. Overall, our results suggest that RePS provides an interpretable and robust alternative to prompting for both steering and suppression.
Zhengxuan Wu, Qinan Yu, Aryaman Arora, Christopher D. Manning, Christopher Potts
NeurIPS2
2024 Grokking Group Multiplication with Cosets
abstract
The complex and unpredictable nature of deep neural networks prevents their safe use in many high-stakes applications. There have been many techniques developed to interpret deep neural networks, but all have substantial limitations. Algorithmic tasks have proven to be a fruitful test ground for interpreting a neural network end-to-end. Building on previous work, we completely reverse engineer fully connected one-hidden layer networks that have “grokked” the arithmetic of the permutation groups $S_5$ and $S_6$. The models discover the true subgroup structure of the full group and converge on neural circuits that decompose the group arithmetic using the permutation group’s subgroups. We relate how we reverse engineered the model’s mechanisms and confirmed our theory was a faithful description of the circuit’s functionality. We also draw attention to current challenges in conducting interpretability research by comparing our work to Chughtai et al. (2023) which alleges to find a different algorithm for this same problem.
Dashiell Stander, Qinan Yu, Honglu Fan, Stella Biderman
ICML2
2024 LLM Circuit Analyses Are Consistent Across Training and Scale
abstract
Most currently deployed LLMs undergo continuous training or additional finetuning. By contrast, most research into LLMs' internal mechanisms focuses on models at one snapshot in time (the end of pre-training), raising the question of whether their results generalize to real-world settings. Existing studies of mechanisms over time focus on encoder-only or toy models, which differ significantly from most deployed models. In this study, we track how model mechanisms, operationalized as circuits, emerge and evolve across 300 billion tokens of training in decoder-only LLMs, in models ranging from 70 million to 2.8 billion parameters. We find that task abilities and the functional components that support them emerge consistently at similar token counts across scale. Moreover, although such components may be implemented by different attention heads over time, the overarching algorithm that they implement remains. Surprisingly, both these algorithms and the types of components involved therein tend to replicate across model scale. Finally, we find that circuit size correlates with model size and can fluctuate considerably over time even when the same algorithm is implemented. These results suggest that circuit analyses conducted on small models at the end of pre-training can provide insights that still apply after additional training and over model scale.
Curt Tigges, Michael Hanna 0001, Qinan Yu, Stella Biderman
NeurIPS3
2023 Characterizing Mechanisms for Factual Recall in Language Models
abstract
Language Models (LMs) often must integrate facts they memorized in pretraining with new information that appears in a given context.These two sources can disagree, causing competition within the model, and it is unclear how an LM will resolve the conflict.On a dataset that queries for knowledge of world capitals, we investigate both distributional and mechanistic determinants of LM behavior in such situations.Specifically, we measure the proportion of the time an LM will use a counterfactual prefix (e.g., "The capital of Poland is London") to overwrite what it learned in pretraining ("Warsaw").On Pythia and GPT2, the training frequency of both the query country ("Poland") and the in-context city ("London") highly affect the models' likelihood of using the counterfactual.We then use head attribution to identify individual attention heads that either promote the memorized answer or the in-context answer in the logits.By scaling up or down the value vector of these heads, we can control the likelihood of using the in-context answer on new data.This method can increase the rate of generating the in-context answer to 88% of the time simply by scaling a single head at runtime.Our work contributes to a body of evidence showing that we can often localize model behaviors to specific components and provides a proof of concept for how future methods might control model behavior dynamically at runtime.
Qinan Yu, Jack Merullo, Ellie Pavlick
EMNLP1