Andrew Drozdov

dblp:200/8508 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0002-1025-5715ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Information retrieval · 100%
Artificial intelligence
8 papers
Language models and text generation · 40% Information extraction and text analysis · 21% Representation and self-supervised learning · 12%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
retrieval-augmented generation
2.432025
The Second Tutorial on Retrieval-Enhanced Machine Learning: Synthesis and Opportunities · SIGIR 2025
FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents · NeurIPS 2025
kNN-LM Does Not Improve Open-ended Text Generation · EMNLP 2023
Information retrieval › evaluation
query performance prediction
1.012026
Can QPP Choose the Right Query variant? Evaluating Query Variant Selection for RAG Pipelines · SIGIR 2026
Information retrieval › evaluation › benchmark
benchmark construction
0.912025
FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents · NeurIPS 2025
Information retrieval
ranking
0.912025
FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents · NeurIPS 2025
Information retrieval
reranking
0.912025
FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents · NeurIPS 2025
Information retrieval
retrieval evaluation
0.912025
FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents · NeurIPS 2025
Machine learning › Representation and self-supervised learning › hierarchical representation
recursive autoencoder
0.822020
Unsupervised Parsing with S-DIORA: Single Tree Encoding for Deep Inside-Outside Recursive Autoencoders · EMNLP (1) 2020
Unsupervised Labeled Parsing with Deep Inside-Outside Recursive Autoencoders · EMNLP/IJCNLP (1) 2019
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.812024
Multistage Collaborative Knowledge Distillation from a Large Language Model for Semi-Supervised Sequence Generation · ACL (1) 2024
Natural language and speech › Information extraction and text analysis › semantic parsing
compositional semantic parsing
0.712023
Compositional Semantic Parsing with Large Language Models · ICLR 2023
Natural language and speech › Language models and text generation › text generation
open-ended text generation
0.712023
kNN-LM Does Not Improve Open-ended Text Generation · EMNLP 2023
Natural language and speech › Information extraction and text analysis
semantic parsing
0.712023
Compositional Semantic Parsing with Large Language Models · ICLR 2023
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › structure learning
latent tree learning
0.512021
Improved Latent Tree Induction with Distant Supervision via Span Constraints · EMNLP (1) 2021
Natural language and speech › Language models and text generation › grammar induction
unsupervised constituency parsing
0.412020
Unsupervised Parsing with S-DIORA: Single Tree Encoding for Deep Inside-Outside Recursive Autoencoders · EMNLP (1) 2020
Knowledge, reasoning and agents › Multi-agent systems
emergent communication
0.312018
Emergent Communication in a Multi-Modal, Multi-Step Referential Game · ICLR (Poster) 2018
Computer vision › Vision and language › multimodal reasoning
compositional reasoning
0.212023
Compositional Semantic Parsing with Large Language Models · ICLR 2023
Natural language and speech › Language models and text generation
decoding
0.212023
kNN-LM Does Not Improve Open-ended Text Generation · EMNLP 2023
Natural language and speech › Information extraction and text analysis
syntactic parsing
0.112021
Improved Latent Tree Induction with Distant Supervision via Span Constraints · EMNLP (1) 2021
Knowledge, reasoning and agents › Multi-agent systems › emergent communication
referential game
0.112018
Emergent Communication in a Multi-Modal, Multi-Step Referential Game · ICLR (Poster) 2018

Methods — techniques the papers use, named apart from their topics

pre-retrieval predictors · 2.0post-retrieval predictors · 2.0perplexity analysis · 1.3interpolation · 1.3human evaluation · 1.3nugget-based evaluation · 0.9machine learning · 0.9information retrieval · 0.9hybrid retrieval · 0.9semi-supervised learning · 0.8knowledge distillation · 0.8large language model · 0.7compositional parsing · 0.7distant supervision · 0.5
YearPublicationVenuePosition
2026 Can QPP Choose the Right Query variant? Evaluating Query Variant Selection for RAG Pipelines
abstract
Large Language Models (LLMs) have made query reformulation ubiquitous in modern retrieval and Retrieval-Augmented Generation (RAG) pipelines, enabling the generation of multiple semantically equivalent query variants. However, executing the full pipeline for every reformulation is computationally expensive, motivating selective execution: can we identify the best query variant before incurring downstream retrieval and generation costs? We investigate Query Performance Prediction (QPP) as a mechanism for variant selection across ad-hoc retrieval, and end-to-end RAG. Unlike traditional QPP, which estimates query difficulty across topics, we study intra-topic discrimination—selecting the optimal reformulation among competing variants of the same information need. Through large-scale experiments on TREC-RAG using both sparse and dense retrievers, we evaluate pre- and post-retrieval predictors under correlation- and decision-based metrics. Our results reveal a systematic divergence between retrieval and generation objectives: variants that maximize ranking metrics such as nDCG often fail to produce the best generated answers, exposing a "utility gap" between retrieval relevance and generation fidelity. Nevertheless, QPP can reliably identify variants that improve end-to-end quality over the original query. Notably, lightweight pre-retrieval predictors frequently match or outperform more expensive post-retrieval methods, offering a latency-efficient approach to robust RAG.
Negar Arabzadeh, Andrew Drozdov, Michael Bendersky, Matei Zaharia
SIGIR2
2025 FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
abstract
We introduce FreshStack, a holistic framework for automatically building information retrieval (IR) evaluation benchmarks by incorporating challenging questions and answers. FreshStack conducts the following steps:(1) automatic corpus collection from code and technical documentation,(2) nugget generation from community-asked questions and answers, and(3) nugget-level support, retrieving documents using a fusion of retrieval techniques and hybrid architectures.We use FreshStack to build five datasets on fast-growing, recent, and niche domains to ensure the tasks are sufficiently challenging. On FreshStack, existing retrieval models, when applied out-of-the-box, significantly underperform oracle approaches on all five domains, denoting plenty of headroom to improve IR quality. In addition, we identify cases where rerankers do not improve first-stage retrieval accuracy (two out of five domains) and oracle context helps an LLM generator generate a high-quality RAG answer.We hope FreshStack will facilitate future work toward constructing realistic, scalable, and uncontaminated IR and RAG evaluation benchmarks.
Nandan Thakur, Jimmy Lin, Sam Havens, Michael Carbin, Omar Khattab, Andrew Drozdov
NeurIPS6
2025 The Second Tutorial on Retrieval-Enhanced Machine Learning: Synthesis and Opportunities
abstract
Retrieval-Enhanced Machine Learning (REML) refers to the use of information retrieval (IR) methods to support reasoning and inference in machine learning tasks. Although relatively recent, these approaches can substantially improve model performance. This includes improved generalization, knowledge grounding, scalability, freshness, attribution, interpretability, and on-device learning. To date, despite being influenced by work in the information retrieval community, REML research has predominantly been presented in natural language processing (NLP) conferences. Our tutorial addresses this disconnect by introducing core REML concepts and synthesizing the literature from various domains in machine learning (ML), including, but not limited to, NLP. What is unique to our approach is the use of consistent notations to provide researchers with a unified and expandable framework. The tutorial will be presented in lecture format based on an existing manuscript, with supporting materials and a comprehensive reading list available at a website. Building on the momentum of our successful workshop at SIGIR 2023 and our tutorial at SIGIR-AP 2024, this year's tutorial features updated content with an emphasis on retrieval technologies used across the broader ML community. We also highlight their role in emerging, future-facing applications such as language agents and evolving scenarios where the extensive body of knowledge from IR can provide critical insights and capabilities.
Fernando Diaz 0001, Andrew Drozdov, To Eun Kim, Alireza Salemi, Hamed Zamani
SIGIR2
2024 Multistage Collaborative Knowledge Distillation from a Large Language Model for Semi-Supervised Sequence Generation
abstract
Jiachen Zhao, Wenlong Zhao, Andrew Drozdov, Benjamin Rozonoyer, Md Arafat Sultan, Jay-Yoon Lee, Mohit Iyyer, Andrew McCallum. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Wenlong Zhao 0001, Andrew Drozdov, Benjamin Rozonoyer, Md. Arafat Sultan, Jay-Yoon Lee, Mohit Iyyer, Andrew McCallum
ACL (1)3
2023 kNN-LM Does Not Improve Open-ended Text Generation
abstract
In this paper, we study the generation quality of interpolation-based retrieval-augmented language models (LMs).These methods, best exemplified by the kNN-LM (Khandelwal et al., 2020), interpolate the LM's predicted distribution of the next word with a distribution formed from the most relevant retrievals for a given prefix.While the kNN-LM and related methods yield impressive decreases in perplexity, we discover that they do not exhibit corresponding improvements in open-ended generation quality, as measured by both automatic evaluation metrics (e.g., MAUVE) and human evaluations.Digging deeper, we find that interpolating with a retrieval distribution actually increases perplexity compared to the baseline LM for the majority of tokens in the WikiText-103 test set, even though the overall perplexity is lower due to a smaller number of tokens for which perplexity dramatically decreases after interpolation.However, when decoding a long sequence at inference time, significant improvements on this smaller subset of tokens are washed out by slightly worse predictions on most tokens.Furthermore, we discover that the entropy of the retrieval distribution increases faster than that of the base LM as the generated sequence becomes longer, which indicates that retrieval is less reliable when using model-generated text as queries (i.e., is subject to exposure bias).We hope that our analysis spurs future work on improved decoding algorithms and interpolation strategies for retrieval-augmented language models.
Shufan Wang, Yixiao Song, Andrew Drozdov, Aparna Garimella, Varun Manjunatha, Mohit Iyyer
EMNLP3
2023 Compositional Semantic Parsing with Large Language Models
Andrew Drozdov, Nathanael Schärli, Ekin Akyürek, Nathan Scales, Xinying Song, Olivier Bousquet, Denny Zhou
ICLR1
2022 Inducing and Using Alignments for Transition-based AMR Parsing
abstract
Andrew Drozdov, Jiawei Zhou, Radu Florian, Andrew McCallum, Tahira Naseem, Yoon Kim, Ramón Astudillo. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Andrew Drozdov, Jiawei Zhou 0001, Radu Florian, Andrew McCallum, Tahira Naseem, Ramón Fernandez Astudillo
NAACL-HLT1
2021 Improved Latent Tree Induction with Distant Supervision via Span Constraints
abstract
Zhiyang Xu, Andrew Drozdov, Jay Yoon Lee, Tim O’Gorman, Subendhu Rongali, Dylan Finkbeiner, Shilpa Suresh, Mohit Iyyer, Andrew McCallum. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Zhiyang Xu, Andrew Drozdov, Jay-Yoon Lee, Tim O'Gorman, Subendhu Rongali, Dylan Finkbeiner, Shilpa Suresh, Mohit Iyyer, Andrew McCallum
EMNLP (1)2
2020 Unsupervised Parsing with S-DIORA: Single Tree Encoding for Deep Inside-Outside Recursive Autoencoders
abstract
The deep inside-outside recursive autoencoder (DIORA; Drozdov et al. 2019a) is a selfsupervised neural model that learns to induce syntactic tree structures for input sentences without access to labeled training data.In this paper, we discover that while DIORA exhaustively encodes all possible binary trees of a sentence with a soft dynamic program, its vector averaging approach is locally greedy and cannot recover from errors when computing the highest scoring parse tree in bottom-up chart parsing.To fix this issue, we introduce S-DIORA, an improved variant of DIORA that encodes a single tree rather than a softlyweighted mixture of trees by employing a hard argmax operation and a beam at each cell in the chart.Our experiments show that through fine-tuning a pre-trained DIORA with our new algorithm, we improve the state of the art in unsupervised constituency parsing on the English WSJ Penn Treebank by 2.2 6% F1, depending on the data used for fine-tuning.
Andrew Drozdov, Subendhu Rongali, Yi-Pei Chen 0001, Tim O'Gorman, Mohit Iyyer, Andrew McCallum
EMNLP (1)1
2019 Unsupervised Labeled Parsing with Deep Inside-Outside Recursive Autoencoders
abstract
Andrew Drozdov, Patrick Verga, Yi-Pei Chen, Mohit Iyyer, Andrew McCallum. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Andrew Drozdov, Patrick Verga, Yi-Pei Chen 0001, Mohit Iyyer, Andrew McCallum
EMNLP/IJCNLP (1)1
2018 Emergent Communication in a Multi-Modal, Multi-Step Referential Game
Katrina Evtimova, Andrew Drozdov, Douwe Kiela, Kyunghyun Cho
ICLR (Poster)2
2018 Do latent tree learning models identify meaningful structure in sentences?
abstract
Recent work on the problem of latent tree learning has made it possible to train neural networks that learn to both parse a sentence and use the resulting parse to interpret the sentence, all without exposure to ground-truth parse trees at training time. Surprisingly, these models often perform better at sentence understanding tasks than models that use parse trees from conventional parsers. This paper aims to investigate what these latent tree learning models learn. We replicate two such models in a shared codebase and find that (i) only one of these models outperforms conventional tree-structured models on sentence classification, (ii) its parsing strategies are not especially consistent across random restarts, (iii) the parses it produces tend to be shallower than standard Penn Treebank (PTB) parses, and (iv) they do not resemble those of PTB or any other semantic or syntactic formalism that the authors are aware of.
Adina Williams, Andrew Drozdov, Samuel R. Bowman
Trans. Assoc. Comput. Linguistics2