VLDB 2026 Research / reviewers in the wild / expert
Md. Arafat Sultan
dblp:77/11514
· DBLP profile ↗
19ranked-venue papers
7as first author
10since 2021 · last 2025
0009-0007-3783-7225ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 7 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
5 papers |
Information retrieval · 94% Data integration and cleaning · 6% | |
| Artificial intelligence
9 papers |
Question answering and dialogue systems · 29% Transfer learning and domain adaptation · 26% Language models and text generation · 18% |
Topics — the 28 heaviest of 31, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
ranking |
1.6 | 2 | 2025 | A Large-Scale Study of Reranker Relevance Feedback at Inference · SIGIR 2025 FIRST: Faster Improved Listwise Reranking with Single Token Decoding · EMNLP 2024 |
Information retrieval
reranking |
1.5 | 2 | 2025 | A Large-Scale Study of Reranker Relevance Feedback at Inference · SIGIR 2025 UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers · EMNLP 2023 |
Information retrieval › retrieval models
neural retrieval |
1.4 | 2 | 2025 | A Large-Scale Study of Reranker Relevance Feedback at Inference · SIGIR 2025 Entity-Conditioned Question Generation for Robust Attention Distribution in Neural Information Retrieval · SIGIR 2022 |
Information retrieval › retrieval models › neural retrieval
dense retrieval |
1.4 | 2 | 2025 | A Large-Scale Study of Reranker Relevance Feedback at Inference · SIGIR 2025 Synthetic Target Domain Supervision for Open Retrieval QA · SIGIR 2021 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.8 | 1 | 2024 | Multistage Collaborative Knowledge Distillation from a Large Language Model for Semi-Supervised Sequence Generation · ACL (1) 2024 |
Information retrieval › ranking
learning to rank |
0.8 | 1 | 2024 | FIRST: Faster Improved Listwise Reranking with Single Token Decoding · EMNLP 2024 |
Information retrieval › reranking
listwise reranking |
0.8 | 1 | 2024 | FIRST: Faster Improved Listwise Reranking with Single Token Decoding · EMNLP 2024 |
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
extractive question answering |
0.7 | 1 | 2023 | GAAMA 2.0: An Integrated System That Answers Boolean and Extractive Questions · AAAI 2023 |
Natural language and speech › Question answering and dialogue systems
machine reading comprehension |
0.7 | 1 | 2023 | GAAMA 2.0: An Integrated System That Answers Boolean and Extractive Questions · AAAI 2023 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation |
0.7 | 1 | 2023 | UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers · EMNLP 2023 |
Information retrieval › document retrieval
passage retrieval |
0.7 | 1 | 2023 | UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers · EMNLP 2023 |
Machine learning › Transfer learning and domain adaptation
domain generalization |
0.6 | 1 | 2022 | Not to Overfit or Underfit the Source Domains? An Empirical Study of Domain Generalization in Question Answering · EMNLP 2022 |
Information retrieval › retrieval models
retrieval model training |
0.6 | 1 | 2022 | Entity-Conditioned Question Generation for Robust Attention Distribution in Neural Information Retrieval · SIGIR 2022 |
Data integration and cleaning › data generation
synthetic data generation |
0.6 | 1 | 2022 | Entity-Conditioned Question Generation for Robust Attention Distribution in Neural Information Retrieval · SIGIR 2022 |
Natural language and speech › Question answering and dialogue systems › retrieval-based question answering
open-retrieval question answering |
0.5 | 1 | 2021 | Synthetic Target Domain Supervision for Open Retrieval QA · SIGIR 2021 |
Natural language and speech › Language models and text generation › text generation › data-to-text generation
AMR-to-text generation |
0.4 | 1 | 2020 | GPT-too: A Language-Model-First Approach for AMR-to-Text Generation · ACL 2020 |
Natural language and speech › Question answering and dialogue systems › question generation
diverse question generation |
0.4 | 1 | 2020 | On the Importance of Diversity in Question Generation for QA · ACL 2020 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.4 | 1 | 2020 | Multi-Stage Pre-training for Low-Resource Domain Adaptation · EMNLP (1) 2020 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
low-resource domain adaptation |
0.4 | 1 | 2020 | Multi-Stage Pre-training for Low-Resource Domain Adaptation · EMNLP (1) 2020 |
Machine learning › Deep learning architectures and training
multi-stage pre-training |
0.4 | 1 | 2020 | Multi-Stage Pre-training for Low-Resource Domain Adaptation · EMNLP (1) 2020 |
Machine learning › Representation and self-supervised learning
pre-training |
0.4 | 1 | 2020 | Multi-Stage Pre-training for Low-Resource Domain Adaptation · EMNLP (1) 2020 |
Natural language and speech › Question answering and dialogue systems
question generation |
0.4 | 1 | 2020 | On the Importance of Diversity in Question Generation for QA · ACL 2020 |
Machine learning › Generative modeling
synthetic training data |
0.4 | 1 | 2020 | On the Importance of Diversity in Question Generation for QA · ACL 2020 |
Natural language and speech › Language models and text generation
text generation |
0.4 | 1 | 2020 | GPT-too: A Language-Model-First Approach for AMR-to-Text Generation · ACL 2020 |
Information retrieval
relevance feedback |
0.2 | 1 | 2024 | FIRST: Faster Improved Listwise Reranking with Single Token Decoding · EMNLP 2024 |
Machine learning › Deep learning architectures and training
transformer |
0.2 | 1 | 2023 | GAAMA 2.0: An Integrated System That Answers Boolean and Extractive Questions · AAAI 2023 |
Machine learning › Transfer learning and domain adaptation › domain generalization
multi-source domain generalization |
0.2 | 1 | 2022 | Not to Overfit or Underfit the Source Domains? An Empirical Study of Domain Generalization in Question Answering · EMNLP 2022 |
Machine learning › Transfer learning and domain adaptation
low-resource learning |
0.1 | 1 | 2020 | Multi-Stage Pre-training for Low-Resource Domain Adaptation · EMNLP (1) 2020 |
Methods — techniques the papers use, named apart from their topics
knowledge distillation · 4.3LLM prompting · 1.3text-to-text generation · 1.0synthetic data generation · 1.0ensemble · 1.0cross-encoder · 0.9single token decoding · 0.8semi-supervised learning · 0.8large language model · 0.8transformer models · 0.7adapter · 0.7targeted synthetic data generation · 0.6neural information retrieval · 0.6nucleus sampling · 0.4language model · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Large-Scale Study of Reranker Relevance Feedback at InferenceabstractNeural IR systems often employ a retrieve-and-rerank framework: a bi-encoder retrieves a fixed number of candidates (e.g., 𝐾=100), which a cross-encoder then reranks.Recent studies have indicated that relevance feedback from the reranker at inference time can improve the recall of the retriever.The approach works by updating the retriever's query representations via a distillation process that aligns it with the reranker's predictions.While a powerful idea, the arguably narrow scope of past studies focusing on a small number of specific domains such as english question answering and entity retrieval has left a gap in our understanding of how well it generalizes.In this paper, we study inference-time reranker relevance feedback extensively across multiple retrieval domains, languages, and modalities, while also investigating aspects such as the performance and latency implications of the number of distillation updates and feedback candidates. Revanth Gangi Reddy, Pradeep Dasigi, Md. Arafat Sultan, Arman Cohan, Avirup Sil, Heng Ji 0001, Hannaneh Hajishirzi |
SIGIR | 3 |
| 2024 | Multistage Collaborative Knowledge Distillation from a Large Language Model for Semi-Supervised Sequence GenerationabstractJiachen Zhao, Wenlong Zhao, Andrew Drozdov, Benjamin Rozonoyer, Md Arafat Sultan, Jay-Yoon Lee, Mohit Iyyer, Andrew McCallum. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Wenlong Zhao 0001, Andrew Drozdov, Benjamin Rozonoyer, Md. Arafat Sultan, Jay-Yoon Lee, Mohit Iyyer, Andrew McCallum |
ACL (1) | 5 |
| 2024 | FIRST: Faster Improved Listwise Reranking with Single Token DecodingabstractLarge Language Models (LLMs) have significantly advanced the field of information retrieval, particularly for reranking.Listwise LLM rerankers typically showcase superior performance and generalizability over conventional supervised approaches.However, existing LLM rerankers can be inefficient as they provide ranking output in the form of a generated ordered sequence of candidate passage identifiers.Further, they are trained using the standard language modeling objective, which treats all ranking errors uniformly, potentially at the cost of misranking highly relevant passages.Addressing these limitations, we introduce FIRST 1 , a novel listwise LLM reranking approach that leverages the output logits of the first generated identifier to directly obtain a ranked ordering of the candidates.We further utilize a learning-to-rank loss for this model, which prioritizes ranking accuracy for the more relevant passages.Empirical results demonstrate that FIRST accelerates inference by 50% while maintaining robust ranking performance, with gains across the BEIR benchmark.Finally, to illustrate the practical effectiveness of listwise LLM rerankers, we investigate their application in providing relevance feedback for retrievers during inference.Our results show that LLM rerankers can provide a stronger distillation signal compared to cross-encoders, yielding substantial improvements in retriever recall after relevance feedback. Revanth Gangi Reddy, JaeHyeok Doo, Md. Arafat Sultan, Deevya Swain, Avirup Sil, Heng Ji 0001 |
EMNLP | 4 |
| 2023 | GAAMA 2.0: An Integrated System That Answers Boolean and Extractive QuestionsabstractRecent machine reading comprehension datasets include extractive and boolean questions but current approaches do not offer integrated support for answering both question types. We present a front-end demo to a multilingual machine reading comprehension system that handles boolean and extractive questions. It provides a yes/no answer and highlights the supporting evidence for boolean questions. It provides an answer for extractive questions and highlights the answer in the passage. Our system, GAAMA 2.0, achieved first place on the TyDi QA leaderboard at the time of submission. We contrast two different implementations of our approach: including multiple transformer models for easy deployment, and a shared transformer model utilizing adapters to reduce GPU memory footprint for a resource-constrained environment. J. Scott McCarley, Mihaela A. Bornea, Sara Rosenthal, Anthony Ferritto, Md. Arafat Sultan, Avirup Sil, Radu Florian |
AAAI | 5 |
| 2023 | UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of RerankersabstractJon Saad-Falcon, Omar Khattab, Keshav Santhanam, Radu Florian, Martin Franz, Salim Roukos, Avirup Sil, Md Sultan, Christopher Potts. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Jon Saad-Falcon, Omar Khattab, Keshav Santhanam, Radu Florian, Martin Franz, Salim Roukos, Avirup Sil, Md. Arafat Sultan, Christopher Potts |
EMNLP | 8 |
| 2022 | Towards Robust Neural Retrieval with Source Domain Synthetic Pre-FinetuningabstractResearch on neural IR has so far been focused primarily on standard supervised learning settings, where it outperforms traditional term matching baselines. Many practical use cases of such models, however, may involve previously unseen target domains. In this paper, we propose to improve the out-of-domain generalization of Dense Passage Retrieval (DPR) - a popular choice for neural IR - through synthetic data augmentation only in the source domain. We empirically show that pre-finetuning DPR with additional synthetic data in its source domain (Wikipedia), which we generate using a fine-tuned sequence-to-sequence generator, can be a low-cost yet effective first step towards its generalization. Across five different test sets, our augmented model shows more robust performance than DPR in both in-domain and zero-shot out-of-domain evaluation. Revanth Gangi Reddy, Vikas Yadav, Md. Arafat Sultan, Martin Franz, Vittorio Castelli, Heng Ji 0001, Avirup Sil |
COLING | 3 |
| 2022 | Not to Overfit or Underfit the Source Domains? An Empirical Study of Domain Generalization in Question AnsweringabstractMachine learning models are prone to overfitting their training (source) domains, which is commonly believed to be the reason why they falter in novel target domains.Here we examine the contrasting view that multi-source domain generalization (DG) is first and foremost a problem of mitigating source domain underfitting: models not adequately learning the signal already present in their multi-domain training data.Experiments on a reading comprehension DG benchmark show that as a model learns its source domains better-using familiar methods such as knowledge distillation (KD) from a bigger model-its zero-shot out-of-domain utility improves at an even faster pace.Improved source domain learning also demonstrates superior out-of-domain generalization over three popular existing DG approaches that aim to limit overfitting.Our implementation of KD-based domain generalization is available via PrimeQA at: https://ibm.biz Md. Arafat Sultan, Avirup Sil, Radu Florian |
EMNLP | 1 |
| 2022 | Learning Cross-Lingual IR from an English RetrieverabstractYulong Li, Martin Franz, Md Arafat Sultan, Bhavani Iyer, Young-Suk Lee, Avirup Sil. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Martin Franz, Md. Arafat Sultan, Bhavani Iyer, Young-Suk Lee 0001, Avirup Sil |
NAACL-HLT | 3 |
| 2022 | Entity-Conditioned Question Generation for Robust Attention Distribution in Neural Information RetrievalabstractWe show that supervised neural information retrieval (IR) models are prone to learning sparse attention patterns over passage tokens, which can result in key phrases including named entities receiving low attention weights, eventually leading to model under-performance. Using a novel targeted synthetic data generation method that identifies poorly attended entities and conditions the generation episodes on those, we teach neural IR to attend more uniformly and robustly to all entities in a given passage. On two public IR benchmarks, we empirically show that the proposed method helps improve both the model's attention patterns and retrieval performance, including in zero-shot settings. Revanth Gangi Reddy, Md. Arafat Sultan, Martin Franz, Avirup Sil, Heng Ji 0001 |
SIGIR | 2 |
| 2021 | Synthetic Target Domain Supervision for Open Retrieval QAabstractNeural passage retrieval is a new and promising approach in open retrieval question answering. In this work, we stress-test the Dense Passage Retriever (DPR)---a state-of-the-art (SOTA) open domain neural retrieval model---on closed and specialized target domains such as COVID-19, and find that it lags behind standard BM25 in this important real-world setting. To make DPR more robust under domain shift, we explore its fine-tuning with synthetic training examples, which we generate from unlabeled target domain text using a text-to-text generator. In our experiments, this noisy but fully automated target domain supervision gives DPR a sizable advantage over BM25 in out-of-domain settings, making it a more viable model in practice. Finally, an ensemble of BM25 and our improved DPR model yields the best results, further pushing the SOTA for open retrieval QA on multiple out-of-domain test sets. Revanth Gangi Reddy, Bhavani Iyer, Md. Arafat Sultan, Rong Zhang 0010, Avirup Sil, Vittorio Castelli, Radu Florian, Salim Roukos |
SIGIR | 3 |
| 2020 | GPT-too: A Language-Model-First Approach for AMR-to-Text GenerationabstractManuel Mager, Ramón Fernandez Astudillo, Tahira Naseem, Md Arafat Sultan, Young-Suk Lee, Radu Florian, Salim Roukos. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Manuel Mager, Ramón Fernandez Astudillo, Tahira Naseem, Md. Arafat Sultan, Young-Suk Lee 0001, Radu Florian, Salim Roukos |
ACL | 4 |
| 2020 | On the Importance of Diversity in Question Generation for QAabstractAutomatic question generation (QG) has shown promise as a source of synthetic training data for question answering (QA).In this paper we ask: Is textual diversity in QG beneficial for downstream QA?Using top-p nucleus sampling to derive samples from a transformer-based question generator, we show that diversity-promoting QG indeed provides better QA training than likelihood maximization approaches such as beam search.We also show that standard QG evaluation metrics such as BLEU, ROUGE and METEOR are inversely correlated with diversity, and propose a diversity-aware intrinsic measure of overall QG quality that correlates well with extrinsic evaluation on QA.1 Md. Arafat Sultan, Shubham Chandel, Ramón Fernandez Astudillo, Vittorio Castelli |
ACL | 1 |
| 2020 | Multi-Stage Pre-training for Low-Resource Domain AdaptationabstractRong Zhang, Revanth Gangi Reddy, Md Arafat Sultan, Vittorio Castelli, Anthony Ferritto, Radu Florian, Efsun Sarioglu Kayi, Salim Roukos, Avi Sil, Todd Ward. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Rong Zhang 0010, Revanth Gangi Reddy, Md. Arafat Sultan, Vittorio Castelli, Anthony Ferritto, Radu Florian, Efsun Sarioglu Kayi, Salim Roukos, Avirup Sil, Todd Ward |
EMNLP (1) | 3 |
| 2016 | Equity of Learning Opportunities in the Chicago City of Learning Program
David Quigley, Ogheneovo Dibie, Md. Arafat Sultan, Katie Van Horne, William R. Penuel, Tamara Sumner, Ugochi Acholonu, Nichole Pinkard |
EDM | 3 |
| 2016 | Bayesian Supervised Domain Adaptation for Short Text SimilarityabstractMd Arafat Sultan, Jordan Boyd-Graber, Tamara Sumner. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Md. Arafat Sultan, Jordan L. Boyd-Graber, Tamara Sumner |
HLT-NAACL | 1 |
| 2016 | Fast and Easy Short Answer Grading with High AccuracyabstractMd Arafat Sultan, Cristobal Salazar, Tamara Sumner. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Md. Arafat Sultan, Cristobal Salazar, Tamara Sumner |
HLT-NAACL | 1 |
| 2016 | A Joint Model for Answer Sentence Ranking and Answer ExtractionabstractAnswer sentence ranking and answer extraction are two key challenges in question answering that have traditionally been treated in isolation, i.e., as independent tasks. In this article, we (1) explain how both tasks are related at their core by a common quantity, and (2) propose a simple and intuitive joint probabilistic model that addresses both via joint computation but task-specific application of that quantity. In our experiments with two TREC datasets, our joint model substantially outperforms state-of-the-art systems in both tasks. Md. Arafat Sultan, Vittorio Castelli, Radu Florian |
Trans. Assoc. Comput. Linguistics | 1 |
| 2015 | Feature-Rich Two-Stage Logistic Regression for Monolingual AlignmentabstractMonolingual alignment is the task of pairing semantically similar units from two pieces of text.We report a top-performing supervised aligner that operates on short text snippets.We employ a large feature set to ( 1) encode similarities among semantic units (words and named entities) in context, and (2) address cooperation and competition for alignment among units in the same snippet.These features are deployed in a two-stage logistic regression framework for alignment.On two benchmark data sets, our aligner achieves F 1 scores of 92.1% and 88.5%, with statistically significant error reductions of 4.8% and 7.3% over the previous best aligner.It produces top results in extrinsic evaluation as well. Md. Arafat Sultan, Steven Bethard, Tamara Sumner |
EMNLP | 1 |
| 2014 | Back to Basics for Monolingual Alignment: Exploiting Word Similarity and Contextual EvidenceabstractWe present a simple, easy-to-replicate monolingual aligner that demonstrates state-of-the-art performance while relying on almost no supervision and a very small number of external resources. Based on the hypothesis that words with similar meanings represent potential pairs for alignment if located in similar contexts, we propose a system that operates by finding such pairs. In two intrinsic evaluations on alignment test data, our system achieves F1 scores of 88–92%, demonstrating 1–3% absolute improvement over the previous best system. Moreover, in two extrinsic evaluations our aligner outperforms existing aligners, and even a naive application of the aligner approaches state-of-the-art performance in each extrinsic task. Md. Arafat Sultan, Steven Bethard, Tamara Sumner |
Trans. Assoc. Comput. Linguistics | 1 |