VLDB 2026 Research / reviewers in the wild / expert
Andrew Drozdov
dblp:200/8508
· DBLP profile ↗
12ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0002-1025-5715ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
4 papers |
Information retrieval · 100% | |
| Artificial intelligence
8 papers |
Language models and text generation · 40% Information extraction and text analysis · 21% Representation and self-supervised learning · 12% |
Topics — the 18 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
retrieval-augmented generation |
2.4 | 3 | 2025 | The Second Tutorial on Retrieval-Enhanced Machine Learning: Synthesis and Opportunities · SIGIR 2025 FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents · NeurIPS 2025 kNN-LM Does Not Improve Open-ended Text Generation · EMNLP 2023 |
Information retrieval › evaluation
query performance prediction |
1.0 | 1 | 2026 | Can QPP Choose the Right Query variant? Evaluating Query Variant Selection for RAG Pipelines · SIGIR 2026 |
Information retrieval › evaluation › benchmark
benchmark construction |
0.9 | 1 | 2025 | FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents · NeurIPS 2025 |
Information retrieval
ranking |
0.9 | 1 | 2025 | FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents · NeurIPS 2025 |
Information retrieval
reranking |
0.9 | 1 | 2025 | FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents · NeurIPS 2025 |
Information retrieval
retrieval evaluation |
0.9 | 1 | 2025 | FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents · NeurIPS 2025 |
Machine learning › Representation and self-supervised learning › hierarchical representation
recursive autoencoder |
0.8 | 2 | 2020 | Unsupervised Parsing with S-DIORA: Single Tree Encoding for Deep Inside-Outside Recursive Autoencoders · EMNLP (1) 2020 Unsupervised Labeled Parsing with Deep Inside-Outside Recursive Autoencoders · EMNLP/IJCNLP (1) 2019 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.8 | 1 | 2024 | Multistage Collaborative Knowledge Distillation from a Large Language Model for Semi-Supervised Sequence Generation · ACL (1) 2024 |
Natural language and speech › Information extraction and text analysis › semantic parsing
compositional semantic parsing |
0.7 | 1 | 2023 | Compositional Semantic Parsing with Large Language Models · ICLR 2023 |
Natural language and speech › Language models and text generation › text generation
open-ended text generation |
0.7 | 1 | 2023 | kNN-LM Does Not Improve Open-ended Text Generation · EMNLP 2023 |
Natural language and speech › Information extraction and text analysis
semantic parsing |
0.7 | 1 | 2023 | Compositional Semantic Parsing with Large Language Models · ICLR 2023 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › structure learning
latent tree learning |
0.5 | 1 | 2021 | Improved Latent Tree Induction with Distant Supervision via Span Constraints · EMNLP (1) 2021 |
Natural language and speech › Language models and text generation › grammar induction
unsupervised constituency parsing |
0.4 | 1 | 2020 | Unsupervised Parsing with S-DIORA: Single Tree Encoding for Deep Inside-Outside Recursive Autoencoders · EMNLP (1) 2020 |
Knowledge, reasoning and agents › Multi-agent systems
emergent communication |
0.3 | 1 | 2018 | Emergent Communication in a Multi-Modal, Multi-Step Referential Game · ICLR (Poster) 2018 |
Computer vision › Vision and language › multimodal reasoning
compositional reasoning |
0.2 | 1 | 2023 | Compositional Semantic Parsing with Large Language Models · ICLR 2023 |
Natural language and speech › Language models and text generation
decoding |
0.2 | 1 | 2023 | kNN-LM Does Not Improve Open-ended Text Generation · EMNLP 2023 |
Natural language and speech › Information extraction and text analysis
syntactic parsing |
0.1 | 1 | 2021 | Improved Latent Tree Induction with Distant Supervision via Span Constraints · EMNLP (1) 2021 |
Knowledge, reasoning and agents › Multi-agent systems › emergent communication
referential game |
0.1 | 1 | 2018 | Emergent Communication in a Multi-Modal, Multi-Step Referential Game · ICLR (Poster) 2018 |
Methods — techniques the papers use, named apart from their topics
pre-retrieval predictors · 2.0post-retrieval predictors · 2.0perplexity analysis · 1.3interpolation · 1.3human evaluation · 1.3nugget-based evaluation · 0.9machine learning · 0.9information retrieval · 0.9hybrid retrieval · 0.9semi-supervised learning · 0.8knowledge distillation · 0.8large language model · 0.7compositional parsing · 0.7distant supervision · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Can QPP Choose the Right Query variant? Evaluating Query Variant Selection for RAG PipelinesabstractLarge Language Models (LLMs) have made query reformulation ubiquitous in modern retrieval and Retrieval-Augmented Generation (RAG) pipelines, enabling the generation of multiple semantically equivalent query variants. However, executing the full pipeline for every reformulation is computationally expensive, motivating selective execution: can we identify the best query variant before incurring downstream retrieval and generation costs? We investigate Query Performance Prediction (QPP) as a mechanism for variant selection across ad-hoc retrieval, and end-to-end RAG. Unlike traditional QPP, which estimates query difficulty across topics, we study intra-topic discrimination—selecting the optimal reformulation among competing variants of the same information need. Through large-scale experiments on TREC-RAG using both sparse and dense retrievers, we evaluate pre- and post-retrieval predictors under correlation- and decision-based metrics. Our results reveal a systematic divergence between retrieval and generation objectives: variants that maximize ranking metrics such as nDCG often fail to produce the best generated answers, exposing a "utility gap" between retrieval relevance and generation fidelity. Nevertheless, QPP can reliably identify variants that improve end-to-end quality over the original query. Notably, lightweight pre-retrieval predictors frequently match or outperform more expensive post-retrieval methods, offering a latency-efficient approach to robust RAG. Negar Arabzadeh, Andrew Drozdov, Michael Bendersky, Matei Zaharia |
SIGIR | 2 |
| 2025 | FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical DocumentsabstractWe introduce FreshStack, a holistic framework for automatically building information retrieval (IR) evaluation benchmarks by incorporating challenging questions and answers. FreshStack conducts the following steps:(1) automatic corpus collection from code and technical documentation,(2) nugget generation from community-asked questions and answers, and(3) nugget-level support, retrieving documents using a fusion of retrieval techniques and hybrid architectures.We use FreshStack to build five datasets on fast-growing, recent, and niche domains to ensure the tasks are sufficiently challenging. On FreshStack, existing retrieval models, when applied out-of-the-box, significantly underperform oracle approaches on all five domains, denoting plenty of headroom to improve IR quality. In addition, we identify cases where rerankers do not improve first-stage retrieval accuracy (two out of five domains) and oracle context helps an LLM generator generate a high-quality RAG answer.We hope FreshStack will facilitate future work toward constructing realistic, scalable, and uncontaminated IR and RAG evaluation benchmarks. Nandan Thakur, Jimmy Lin, Sam Havens, Michael Carbin, Omar Khattab, Andrew Drozdov |
NeurIPS | 6 |
| 2025 | The Second Tutorial on Retrieval-Enhanced Machine Learning: Synthesis and OpportunitiesabstractRetrieval-Enhanced Machine Learning (REML) refers to the use of information retrieval (IR) methods to support reasoning and inference in machine learning tasks. Although relatively recent, these approaches can substantially improve model performance. This includes improved generalization, knowledge grounding, scalability, freshness, attribution, interpretability, and on-device learning. To date, despite being influenced by work in the information retrieval community, REML research has predominantly been presented in natural language processing (NLP) conferences. Our tutorial addresses this disconnect by introducing core REML concepts and synthesizing the literature from various domains in machine learning (ML), including, but not limited to, NLP. What is unique to our approach is the use of consistent notations to provide researchers with a unified and expandable framework. The tutorial will be presented in lecture format based on an existing manuscript, with supporting materials and a comprehensive reading list available at a website. Building on the momentum of our successful workshop at SIGIR 2023 and our tutorial at SIGIR-AP 2024, this year's tutorial features updated content with an emphasis on retrieval technologies used across the broader ML community. We also highlight their role in emerging, future-facing applications such as language agents and evolving scenarios where the extensive body of knowledge from IR can provide critical insights and capabilities. Fernando Diaz 0001, Andrew Drozdov, To Eun Kim, Alireza Salemi, Hamed Zamani |
SIGIR | 2 |
| 2024 | Multistage Collaborative Knowledge Distillation from a Large Language Model for Semi-Supervised Sequence GenerationabstractJiachen Zhao, Wenlong Zhao, Andrew Drozdov, Benjamin Rozonoyer, Md Arafat Sultan, Jay-Yoon Lee, Mohit Iyyer, Andrew McCallum. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Wenlong Zhao 0001, Andrew Drozdov, Benjamin Rozonoyer, Md. Arafat Sultan, Jay-Yoon Lee, Mohit Iyyer, Andrew McCallum |
ACL (1) | 3 |
| 2023 | kNN-LM Does Not Improve Open-ended Text GenerationabstractIn this paper, we study the generation quality of interpolation-based retrieval-augmented language models (LMs).These methods, best exemplified by the kNN-LM (Khandelwal et al., 2020), interpolate the LM's predicted distribution of the next word with a distribution formed from the most relevant retrievals for a given prefix.While the kNN-LM and related methods yield impressive decreases in perplexity, we discover that they do not exhibit corresponding improvements in open-ended generation quality, as measured by both automatic evaluation metrics (e.g., MAUVE) and human evaluations.Digging deeper, we find that interpolating with a retrieval distribution actually increases perplexity compared to the baseline LM for the majority of tokens in the WikiText-103 test set, even though the overall perplexity is lower due to a smaller number of tokens for which perplexity dramatically decreases after interpolation.However, when decoding a long sequence at inference time, significant improvements on this smaller subset of tokens are washed out by slightly worse predictions on most tokens.Furthermore, we discover that the entropy of the retrieval distribution increases faster than that of the base LM as the generated sequence becomes longer, which indicates that retrieval is less reliable when using model-generated text as queries (i.e., is subject to exposure bias).We hope that our analysis spurs future work on improved decoding algorithms and interpolation strategies for retrieval-augmented language models. Shufan Wang, Yixiao Song, Andrew Drozdov, Aparna Garimella, Varun Manjunatha, Mohit Iyyer |
EMNLP | 3 |
| 2023 | Compositional Semantic Parsing with Large Language Models
Andrew Drozdov, Nathanael Schärli, Ekin Akyürek, Nathan Scales, Xinying Song, Olivier Bousquet, Denny Zhou |
ICLR | 1 |
| 2022 | Inducing and Using Alignments for Transition-based AMR ParsingabstractAndrew Drozdov, Jiawei Zhou, Radu Florian, Andrew McCallum, Tahira Naseem, Yoon Kim, Ramón Astudillo. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Andrew Drozdov, Jiawei Zhou 0001, Radu Florian, Andrew McCallum, Tahira Naseem, Ramón Fernandez Astudillo |
NAACL-HLT | 1 |
| 2021 | Improved Latent Tree Induction with Distant Supervision via Span ConstraintsabstractZhiyang Xu, Andrew Drozdov, Jay Yoon Lee, Tim O’Gorman, Subendhu Rongali, Dylan Finkbeiner, Shilpa Suresh, Mohit Iyyer, Andrew McCallum. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Zhiyang Xu, Andrew Drozdov, Jay-Yoon Lee, Tim O'Gorman, Subendhu Rongali, Dylan Finkbeiner, Shilpa Suresh, Mohit Iyyer, Andrew McCallum |
EMNLP (1) | 2 |
| 2020 | Unsupervised Parsing with S-DIORA: Single Tree Encoding for Deep Inside-Outside Recursive AutoencodersabstractThe deep inside-outside recursive autoencoder (DIORA; Drozdov et al. 2019a) is a selfsupervised neural model that learns to induce syntactic tree structures for input sentences without access to labeled training data.In this paper, we discover that while DIORA exhaustively encodes all possible binary trees of a sentence with a soft dynamic program, its vector averaging approach is locally greedy and cannot recover from errors when computing the highest scoring parse tree in bottom-up chart parsing.To fix this issue, we introduce S-DIORA, an improved variant of DIORA that encodes a single tree rather than a softlyweighted mixture of trees by employing a hard argmax operation and a beam at each cell in the chart.Our experiments show that through fine-tuning a pre-trained DIORA with our new algorithm, we improve the state of the art in unsupervised constituency parsing on the English WSJ Penn Treebank by 2.2 6% F1, depending on the data used for fine-tuning. Andrew Drozdov, Subendhu Rongali, Yi-Pei Chen 0001, Tim O'Gorman, Mohit Iyyer, Andrew McCallum |
EMNLP (1) | 1 |
| 2019 | Unsupervised Labeled Parsing with Deep Inside-Outside Recursive AutoencodersabstractAndrew Drozdov, Patrick Verga, Yi-Pei Chen, Mohit Iyyer, Andrew McCallum. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Andrew Drozdov, Patrick Verga, Yi-Pei Chen 0001, Mohit Iyyer, Andrew McCallum |
EMNLP/IJCNLP (1) | 1 |
| 2018 | Emergent Communication in a Multi-Modal, Multi-Step Referential Game
Katrina Evtimova, Andrew Drozdov, Douwe Kiela, Kyunghyun Cho |
ICLR (Poster) | 2 |
| 2018 | Do latent tree learning models identify meaningful structure in sentences?abstractRecent work on the problem of latent tree learning has made it possible to train neural networks that learn to both parse a sentence and use the resulting parse to interpret the sentence, all without exposure to ground-truth parse trees at training time. Surprisingly, these models often perform better at sentence understanding tasks than models that use parse trees from conventional parsers. This paper aims to investigate what these latent tree learning models learn. We replicate two such models in a shared codebase and find that (i) only one of these models outperforms conventional tree-structured models on sentence classification, (ii) its parsing strategies are not especially consistent across random restarts, (iii) the parses it produces tend to be shallower than standard Penn Treebank (PTB) parses, and (iv) they do not resemble those of PTB or any other semantic or syntactic formalism that the authors are aware of. Adina Williams, Andrew Drozdov, Samuel R. Bowman |
Trans. Assoc. Comput. Linguistics | 2 |