VLDB 2026 Research / reviewers in the wild / expert
Dan Friedman
dblp:205/9386
· DBLP profile ↗
10ranked-venue papers
5as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Trustworthy machine learning · 63% Language models and text generation · 26% Information extraction and text analysis · 5% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
interpretability |
2.9 | 4 | 2024 | Finding Transformer Circuits With Edge Pruning · NeurIPS 2024 Interpretability Illusions in the Generalization of Simplified Models · ICML 2024 The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language Models · ACL (1) 2024 |
Machine learning › Trustworthy machine learning › interpretability
mechanistic interpretability |
2.2 | 3 | 2024 | Finding Transformer Circuits With Edge Pruning · NeurIPS 2024 Interpretability Illusions in the Generalization of Simplified Models · ICML 2024 Learning Transformer Programs · NeurIPS 2023 |
Natural language and speech › Language models and text generation
in-context learning |
0.9 | 2 | 2024 | Measuring Inductive Biases of In-Context Learning with Underspecified Demonstrations · ACL (1) 2023 Finding Transformer Circuits With Edge Pruning · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › language model interpretability
attention head analysis |
0.8 | 1 | 2024 | The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language Models · ACL (1) 2024 |
Machine learning › Trustworthy machine learning › interpretability › mechanistic interpretability
circuit discovery |
0.8 | 1 | 2024 | Finding Transformer Circuits With Edge Pruning · NeurIPS 2024 |
Natural language and speech › Language models and text generation › pre-trained language model
pretrained language model analysis |
0.8 | 1 | 2024 | The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language Models · ACL (1) 2024 |
Natural language and speech › Question answering and dialogue systems
machine reading comprehension |
0.5 | 1 | 2021 | Single-dataset Experts for Multi-dataset Question Answering · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation › text summarization › long document summarization
scientific paper summarization |
0.4 | 1 | 2019 | ScisummNet: A Large Annotated Corpus and Content-Impact Models for Scientific Paper Summarization with Citation Networks · AAAI 2019 |
Natural language and speech › Language models and text generation
text summarization |
0.4 | 1 | 2019 | ScisummNet: A Large Annotated Corpus and Content-Impact Models for Scientific Paper Summarization with Citation Networks · AAAI 2019 |
Information retrieval
citation analysis |
0.4 | 1 | 2019 | ScisummNet: A Large Annotated Corpus and Content-Impact Models for Scientific Paper Summarization with Citation Networks · AAAI 2019 |
Natural language and speech › Language models and text generation › linguistic generalization
syntactic generalization |
0.2 | 1 | 2024 | The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language Models · ACL (1) 2024 |
Natural language and speech › Language models and text generation
prompting |
0.2 | 1 | 2023 | Measuring Inductive Biases of In-Context Learning with Underspecified Demonstrations · ACL (1) 2023 |
Machine learning › Trustworthy machine learning
robustness |
0.2 | 1 | 2022 | Finding Dataset Shortcuts with Grammar Induction · EMNLP 2022 |
Machine learning › Transfer learning and domain adaptation › multi-source learning
multi-dataset learning |
0.1 | 1 | 2021 | Single-dataset Experts for Multi-dataset Question Answering · EMNLP (1) 2021 |
Methods — techniques the papers use, named apart from their topics
transformer · 1.3subnetwork analysis · 0.8singular value decomposition · 0.8optimization · 0.8hybrid summarization · 0.8gradient-based pruning · 0.8edge pruning · 0.8dimensionality reduction · 0.8clustering · 0.8citation network integration · 0.8attention head analysis · 0.8gradient-based optimization · 0.7demonstration design · 0.7RASP · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Representing Rule-based Chatbots with TransformersabstractDan Friedman, Abhishek Panigrahi, Danqi Chen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Dan Friedman, Abhishek Panigrahi, Danqi Chen 0001 |
NAACL (Long Papers) | 1 |
| 2024 | The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language ModelsabstractPrior work has found that pretrained language models (LMs) fine-tuned with different random seeds can achieve similar in-domain performance but generalize differently on tests of syntactic generalization.In this work, we show that, even within a single model, we can find multiple subnetworks that perform similarly indomain, but generalize vastly differently.To better understand these phenomena, we investigate if they can be understood in terms of "competing subnetworks": the model initially represents a variety of distinct algorithms, corresponding to different subnetworks, and generalization occurs when it ultimately converges to one.This explanation has been used to account for generalization in simple algorithmic tasks ("grokking").Instead of finding competing subnetworks, we find that all subnetworkswhether they generalize or not-share a set of attention heads, which we refer to as the heuristic core.Further analysis suggests that these attention heads emerge early in training and compute shallow, non-generalizing features.The model learns to generalize by incorporating additional attention heads, which depend on the outputs of the "heuristic" heads to compute higher-level features.Overall, our results offer a more detailed picture of the mechanisms for syntactic generalization in pretrained LMs. 1 Adithya Bhaskar, Dan Friedman, Danqi Chen 0001 |
ACL (1) | 2 |
| 2024 | Interpretability Illusions in the Generalization of Simplified ModelsabstractA common method to study deep learning systems is to use simplified model representations—for example, using singular value decomposition to visualize the model’s hidden states in a lower dimensional space. This approach assumes that the results of these simplifications are faithful to the original model. Here, we illustrate an important caveat to this assumption: even if the simplified representations can accurately approximate the full model on the training set, they may fail to accurately capture the model’s behavior out of distribution. We illustrate this by training Transformer models on controlled datasets with systematic generalization splits, including the Dyck balanced-parenthesis languages and a code completion task. We simplify these models using tools like dimensionality reduction and clustering, and then explicitly test how these simplified proxies match the behavior of the original model. We find consistent generalization gaps: cases in which the simplified proxies are more faithful to the original model on the in-distribution evaluations and less faithful on various tests of systematic generalization. This includes cases where the original model generalizes systematically but the simplified proxies fail, and cases where the simplified proxies generalize better. Together, our results raise questions about the extent to which mechanistic interpretations derived using tools like SVD can reliably predict what a model will do in novel situations. Dan Friedman, Andrew K. Lampinen, Lucas Dixon, Danqi Chen 0001, Asma Ghandeharioun |
ICML | 1 |
| 2024 | Finding Transformer Circuits With Edge PruningabstractThe path to interpreting a language model often proceeds via analysis of circuits---sparse computational subgraphs of the model that capture specific aspects of its behavior. Recent work has automated the task of discovering circuits. Yet, these methods have practical limitations, as they either rely on inefficient search algorithms or inaccurate approximations. In this paper, we frame circuit discovery as an optimization problem and propose _Edge Pruning_ as an effective and scalable solution. Edge Pruning leverages gradient-based pruning techniques, but instead of removing neurons or components, prunes the _edges_ between components. Our method finds circuits in GPT-2 that use less than half the number of edges than circuits found by previous methods while being equally faithful to the full model predictions on standard circuit-finding tasks. Edge Pruning is efficient on tasks involving up to 100,000 examples, outperforming previous methods in speed and producing substantially better circuits. It also perfectly recovers the ground-truth circuits in two models compiled with Tracr. Thanks to its efficiency, we scale Edge Pruning to CodeLlama-13B, a model over 100x the size of GPT-2.
We use this setting for a case study, where we compare the mechanisms behind instruction prompting and in-context learning.
We find two circuits with more than 99.96% sparsity that match the performance of the full model. Further analysis reveals that the mechanisms in the two settings overlap substantially. This shows that Edge Pruning is a practical and scalable tool for interpretability,
which can shed light on behaviors that only emerge in large models. Adithya Bhaskar, Alexander Wettig, Dan Friedman, Danqi Chen 0001 |
NeurIPS | 3 |
| 2023 | Measuring Inductive Biases of In-Context Learning with Underspecified DemonstrationsabstractIn-context learning (ICL) is an important paradigm for adapting large language models (LLMs) to new tasks, but the generalization behavior of ICL remains poorly understood.We investigate the inductive biases of ICL from the perspective of feature bias: which feature ICL is more likely to use given a set of underspecified demonstrations in which two features are equally predictive of the labels.First, we characterize the feature biases of GPT-3 models by constructing underspecified demonstrations from a range of NLP datasets and feature combinations.We find that LLMs exhibit clear feature biases-for example, demonstrating a strong bias to predict labels according to sentiment rather than shallow lexical features, like punctuation.Second, we evaluate the effect of different interventions that are designed to impose an inductive bias in favor of a particular feature, such as adding a natural language instruction or using semantically relevant label words.We find that, while many interventions can influence the learner to prefer a particular feature, it can be difficult to overcome strong prior biases.Overall, our results provide a broader picture of the types of features that ICL may be more likely to exploit and how to impose inductive biases that are better aligned with the intended task. 1 Chenglei Si, Dan Friedman, Nitish Joshi, Shi Feng 0005, Danqi Chen 0001, He He 0001 |
ACL (1) | 2 |
| 2023 | Learning Transformer ProgramsabstractRecent research in mechanistic interpretability has attempted to reverse-engineer Transformer models by carefully inspecting network weights and activations. However, these approaches require considerable manual effort and still fall short of providing complete, faithful descriptions of the underlying algorithms. In this work, we introduce a procedure for training Transformers that are mechanistically interpretable by design. We build on RASP [Weiss et al., 2021], a programming language that can be compiled into Transformer weights. Instead of compiling human-written programs into Transformers, we design a modified Transformer that can be trained using gradient-based optimization and then automatically converted into a discrete, human-readable program. We refer to these models as Transformer Programs. To validate our approach, we learn Transformer Programs for a variety of problems, including an in-context learning task, a suite of algorithmic problems (e.g. sorting, recognizing Dyck languages), and NLP tasks including named entity recognition and text classification. The Transformer Programs can automatically find reasonable solutions, performing on par with standard Transformers of comparable size; and, more importantly, they are easy to interpret. To demonstrate these advantages, we convert Transformers into Python programs and use off-the-shelf code analysis tools to debug model errors and identify the “circuits” used to solve different sub-problems. We hope that Transformer Programs open a new path toward the goal of intrinsically interpretable machine learning. Dan Friedman, Alexander Wettig, Danqi Chen 0001 |
NeurIPS | 1 |
| 2022 | Finding Dataset Shortcuts with Grammar InductionabstractMany NLP datasets have been found to contain shortcuts: simple decision rules that achieve surprisingly high accuracy.However, it is difficult to discover shortcuts automatically.Prior work on automatic shortcut detection has focused on enumerating features like unigrams or bigrams, which can find only low-level shortcuts, or relied on post-hoc model interpretability methods like saliency maps, which reveal qualitative patterns without a clear statistical interpretation.In this work, we propose to use probabilistic grammars to characterize and discover shortcuts in NLP datasets.Specifically, we use a contextfree grammar to model patterns in sentence classification datasets and use a synchronous context-free grammar to model datasets involving sentence pairs.The resulting grammars reveal interesting shortcut features in a number of datasets, including both simple and high-level features, and automatically identify groups of test examples on which conventional classifiers fail.Finally, we show that the features we discover can be used to generate diagnostic contrast examples and incorporated into standard robust optimization methods to improve worst-group accuracy.1 Dan Friedman, Alexander Wettig, Danqi Chen 0001 |
EMNLP | 1 |
| 2021 | Single-dataset Experts for Multi-dataset Question AnsweringabstractMany datasets have been created for training reading comprehension models, and a natural question is whether we can combine them to build models that (1) perform better on all of the training datasets and (2) generalize and transfer better to new datasets.Prior work has addressed this goal by training one network simultaneously on multiple datasets, which works well on average but is prone to over-or under-fitting different subdistributions and might transfer worse compared to source models with more overlap with the target dataset.Our approach is to model multi-dataset question answering with an ensemble of single-dataset experts, by training a collection of lightweight, dataset-specific adapter modules (Houlsby et al., 2019) that share an underlying Transformer model.We find that these Multi-Adapter Dataset Experts (MADE) outperform all our baselines in terms of in-distribution accuracy, and simple methods based on parameter-averaging lead to better zero-shot generalization and few-shot transfer performance, offering a strong and versatile starting point for building new reading comprehension systems. 1 Dan Friedman, Ben Dodge, Danqi Chen 0001 |
EMNLP (1) | 1 |
| 2021 | Factual Probing Is [MASK]: Learning vs. Learning to RecallabstractPetroni et al. (2019) demonstrated that it is possible to retrieve world facts from a pretrained language model by expressing them as cloze-style prompts and interpret the model's prediction accuracy as a lower bound on the amount of factual information it encodes.Subsequent work has attempted to tighten the estimate by searching for better prompts, using a disjoint set of facts as training data.In this work, we make two complementary contributions to better understand these factual probing techniques.First, we propose OPTIPROMPT, a novel and efficient method which directly optimizes in continuous embedding space.We find this simple method is able to predict an additional 6.4% of facts in the LAMA benchmark.Second, we raise a more important question: Can we really interpret these probing results as a lower bound?Is it possible that these prompt-search methods learn from the training data too?We find, somewhat surprisingly, that the training data used by these methods contains certain regularities of the underlying fact distribution, and all the existing prompt methods, including ours, are able to exploit them for better fact prediction.We conduct a set of control experiments to disentangle "learning" from "learning to recall", providing a more detailed picture of what different prompts can reveal about pre-trained language models. 1 * The first two authors contributed equally. Zexuan Zhong, Dan Friedman, Danqi Chen 0001 |
NAACL-HLT | 2 |
| 2019 | ScisummNet: A Large Annotated Corpus and Content-Impact Models for Scientific Paper Summarization with Citation NetworksabstractScientific article summarization is challenging: large, annotated corpora are not available, and the summary should ideally include the article’s impacts on research community. This paper provides novel solutions to these two challenges. We 1) develop and release the first large-scale manually-annotated corpus for scientific papers (on computational linguistics) by enabling faster annotation, and 2) propose summarization methods that integrate the authors’ original highlights (abstract) and the article’s actual impacts on the community (citations), to create comprehensive, hybrid summaries. We conduct experiments to demonstrate the efficacy of our corpus in training data-driven models for scientific paper summarization and the advantage of our hybrid summaries over abstracts and traditional citation-based summaries. Our large annotated corpus and hybrid methods provide a new framework for scientific paper summarization research. Michihiro Yasunaga, Jungo Kasai, Rui Zhang 0037, Alexander R. Fabbri, Irene Li, Dan Friedman, Dragomir R. Radev |
AAAI | 6 |