Reduan Achtibat

dblp:322/1194 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 65% Language models and text generation · 24% Information extraction and text analysis · 12%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
1.622025
The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation · NeurIPS 2025
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers · ICML 2024
Machine learning › Trustworthy machine learning › language model interpretability
attention head analysis
0.912025
The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation · NeurIPS 2025
Natural language and speech › Language models and text generation
in-context learning
0.912025
The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation · NeurIPS 2025
Natural language and speech › Information extraction and text analysis › information retrieval
knowledge retrieval
0.912025
The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation · NeurIPS 2025
Natural language and speech › Language models and text generation
retrieval augmentation
0.912025
The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation · NeurIPS 2025
Machine learning › Trustworthy machine learning › interpretability
attribution methods
0.812024
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers · ICML 2024
Machine learning › Trustworthy machine learning › interpretability › attribution methods
layer-wise relevance propagation
0.812024
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers · ICML 2024
Machine learning › Trustworthy machine learning › interpretability › neural network interpretation
transformer interpretability
0.812024
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers · ICML 2024

Methods — techniques the papers use, named apart from their topics

function vectors · 0.9attribution-based method · 0.9attention weight modification · 0.9layer-wise relevance propagation · 0.8attention-aware attribution · 0.8
YearPublicationVenuePosition
2025 The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation
abstract
Large language models are able to exploit in-context learning to access external knowledge beyond their training data through retrieval-augmentation. While promising, its inner workings remain unclear. In this work, we shed light on the mechanism of in-context retrieval augmentation for question answering by viewing a prompt as a composition of informational components. We propose an attribution-based method to identify specialized attention heads, revealing in-context heads that comprehend instructions and retrieve relevant contextual information, and parametric heads that store entities' relational knowledge. To better understand their roles, we extract function vectors and modify their attention weights to show how they can influence the answer generation process. Finally, we leverage the gained insights to trace the sources of knowledge used during inference, paving the way towards more safe and transparent language models.
Patrick Kahardipraja, Reduan Achtibat, Thomas Wiegand 0001, Wojciech Samek, Sebastian Lapuschkin
NeurIPS2
2024 AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
abstract
Large Language Models are prone to biased predictions and hallucinations, underlining the paramount importance of understanding their model-internal reasoning process. However, achieving faithful attributions for the entirety of a black-box transformer model and maintaining computational efficiency is an unsolved challenge. By extending the Layer-wise Relevance Propagation attribution method to handle attention layers, we address these challenges effectively. While partial solutions exist, our method is the first to faithfully and holistically attribute not only input but also latent representations of transformer models with the computational efficiency similar to a single backward pass. Through extensive evaluations against existing methods on LLaMa 2, Mixtral 8x7b, Flan-T5 and vision transformer architectures, we demonstrate that our proposed approach surpasses alternative methods in terms of faithfulness and enables the understanding of latent representations, opening up the door for concept-based explanations. We provide an LRP library at https://github.com/rachtibat/LRP-eXplains-Transformers.
Reduan Achtibat, Sayed Mohammad Vakilzadeh Hatefi, Maximilian Dreyer, Thomas Wiegand 0001, Sebastian Lapuschkin, Wojciech Samek
ICML1