EDBT 2026 Demo / reviewers in the wild / expert
Corby Rosset
dblp:218/5665
· DBLP profile ↗
11ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0001-9167-6214ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Reinforcement learning · 36% Language models and text generation · 33% Question answering and dialogue systems · 18% | |
| Databases, data mining, and information retrieval
6 papers |
Information retrieval · 88% Recommender systems · 7% Query processing and optimization · 5% |
Topics — the 30 heaviest of 33, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
alignment |
1.1 | 2 | 2025 | Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF · ICLR 2025 Axiomatic Preference Modeling for Longform Question Answering · EMNLP 2023 |
Machine learning › Reinforcement learning
exploration |
0.9 | 1 | 2025 | Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF · ICLR 2025 |
Machine learning › Reinforcement learning › exploration
online exploration |
0.9 | 1 | 2025 | Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF · ICLR 2025 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.9 | 1 | 2025 | Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF · ICLR 2025 |
Information retrieval › interactive information retrieval › conversational information seeking
conversational search |
0.8 | 2 | 2023 | Zero-shot Clarifying Question Generation for Conversational Search · WWW 2023 Contextual Re-Ranking with Behavior Aware Transformers · SIGIR 2020 |
Machine learning › Efficient and distributed learning › inference efficiency
context compression |
0.8 | 1 | 2024 | Dodo: Dynamic Contextual Compression for Decoder-only LMs · ACL (1) 2024 |
Natural language and speech › Language models and text generation › large language model evaluation › automatic evaluation
rubric-based evaluation |
0.8 | 1 | 2024 | LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts · ACL (1) 2024 |
Natural language and speech › Language models and text generation
text evaluation |
0.8 | 1 | 2024 | LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts · ACL (1) 2024 |
Machine learning › Reinforcement learning
human feedback |
0.7 | 1 | 2023 | Axiomatic Preference Modeling for Longform Question Answering · EMNLP 2023 |
Natural language and speech › Question answering and dialogue systems
long-form question answering |
0.7 | 1 | 2023 | Axiomatic Preference Modeling for Longform Question Answering · EMNLP 2023 |
Natural language and speech › Question answering and dialogue systems
question generation |
0.7 | 1 | 2023 | Zero-shot Clarifying Question Generation for Conversational Search · WWW 2023 |
Machine learning › Reinforcement learning › reward learning
reward modeling |
0.7 | 1 | 2023 | Axiomatic Preference Modeling for Longform Question Answering · EMNLP 2023 |
Information retrieval › interactive information retrieval › conversational information seeking › conversational search
clarifying question generation |
0.7 | 1 | 2023 | Zero-shot Clarifying Question Generation for Conversational Search · WWW 2023 |
Information retrieval › retrieval models › neural retrieval
dense retrieval |
0.7 | 1 | 2023 | Augmenting Zero-Shot Dense Retrievers with Plug-in Mixture-of-Memories · EMNLP 2023 |
Information retrieval
retrieval augmentation |
0.7 | 1 | 2023 | Augmenting Zero-Shot Dense Retrievers with Plug-in Mixture-of-Memories · EMNLP 2023 |
Information retrieval › document retrieval
zero-shot retrieval |
0.7 | 1 | 2023 | Augmenting Zero-Shot Dense Retrievers with Plug-in Mixture-of-Memories · EMNLP 2023 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › information fusion
multi-evidence reasoning |
0.4 | 1 | 2020 | Transformer-XH: Multi-Evidence Reasoning with eXtra Hop Attention · ICLR 2020 |
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering |
0.4 | 1 | 2020 | Transformer-XH: Multi-Evidence Reasoning with eXtra Hop Attention · ICLR 2020 |
Information retrieval › ranking › text ranking
document ranking |
0.4 | 1 | 2020 | Contextual Re-Ranking with Behavior Aware Transformers · SIGIR 2020 |
Information retrieval › retrieval models
axiomatic retrieval model |
0.4 | 1 | 2019 | An Axiomatic Approach to Regularizing Neural Ranking Models · SIGIR 2019 |
Information retrieval › retrieval models › neural retrieval
neural ranking model |
0.4 | 1 | 2019 | An Axiomatic Approach to Regularizing Neural Ranking Models · SIGIR 2019 |
Information retrieval › query understanding
query representation |
0.4 | 1 | 2019 | Generic Intent Representation in Web Search · SIGIR 2019 |
Information retrieval
query understanding |
0.4 | 1 | 2019 | Generic Intent Representation in Web Search · SIGIR 2019 |
Recommender systems › large-scale recommendation › multi-stage recommender systems
candidate generation |
0.3 | 1 | 2018 | Optimizing Query Evaluations Using Reinforcement Learning for Web Search · SIGIR 2018 |
Query processing and optimization › query execution › scan processing
index scan |
0.3 | 1 | 2018 | Optimizing Query Evaluations Using Reinforcement Learning for Web Search · SIGIR 2018 |
Information retrieval
query processing |
0.3 | 1 | 2018 | Optimizing Query Evaluations Using Reinforcement Learning for Web Search · SIGIR 2018 |
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization |
0.3 | 1 | 2025 | Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF · ICLR 2025 |
Natural language and speech › Question answering and dialogue systems
dialogue evaluation |
0.2 | 1 | 2024 | LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts · ACL (1) 2024 |
Machine learning › Efficient and distributed learning
model compression |
0.2 | 1 | 2024 | Dodo: Dynamic Contextual Compression for Decoder-only LMs · ACL (1) 2024 |
Recommender systems › user modeling
user interaction modeling |
0.1 | 1 | 2020 | Contextual Re-Ranking with Behavior Aware Transformers · SIGIR 2020 |
Methods — techniques the papers use, named apart from their topics
zero-shot learning · 1.3implicit q*-approximation · 0.9exploration bonus · 0.9KL-regularized MDP · 0.9parameter tuning · 0.8large language model prompting · 0.8feed-forward neural network calibration · 0.8LoRA · 0.8t5 · 0.7reinforcement learning from human feedback · 0.7joint learning · 0.7hard negative mining · 0.7axiomatic preference modeling · 0.7transformer · 0.4attention mechanism · 0.4regularization loss · 0.4data perturbation · 0.4click supervision · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHFabstractThis paper investigates a basic question in reinforcement learning from human feedback (RLHF) from a theoretical perspective: how to efficiently explore in an online manner under preference feedback and general function approximation. We take the initial step towards a theoretical understanding of this problem by proposing a novel algorithm, *Exploratory Preference Optimization* (XPO). This algorithm is elegantly simple---requiring only a one-line modification to (online) Direct Preference Optimization (DPO; Rafailov et al., 2023)---yet provides the strongest known provable guarantees. XPO augments the DPO objective with a novel and principled *exploration bonus*, enabling the algorithm to strategically explore beyond the support of the initial model and preference feedback data. We prove that XPO is provably sample-efficient and converges to a near-optimal policy under natural exploration conditions, regardless of the initial model's coverage. Our analysis builds on the observation that DPO implicitly performs a form of *Bellman error minimization*. It synthesizes previously disparate techniques from language modeling and theoretical reinforcement learning in a serendipitous fashion through the lens of *KL-regularized Markov decision processes*. Tengyang Xie, Dylan J. Foster, Akshay Krishnamurthy, Corby Rosset, Ahmed Awadallah 0001, Alexander Rakhlin |
ICLR | 4 |
| 2024 | LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language TextsabstractThis paper introduces a framework for the automated evaluation of natural language texts.A manually constructed rubric describes how to assess multiple dimensions of interest.To evaluate a text, a large language model (LLM) is prompted with each rubric question and produces a distribution over potential responses.The LLM predictions often fail to agree well with human judges-indeed, the humans do not fully agree with one another.However, the multiple LLM distributions can be combined to predict each human judge's annotations on all questions, including a summary question that assesses overall quality or relevance.LLM-RUBRIC accomplishes this by training a small feed-forward neural network that includes both judge-specific and judge-independent parameters.When evaluating dialogue systems in a human-AI information-seeking task, we find that LLM-RUBRIC with 9 questions (assessing dimensions such as naturalness, conciseness, and citation quality) predicts human judges' assessment of overall user satisfaction, on a scale of 1-4, with RMS error ă 0.5, a 2ˆ improvement over the uncalibrated baseline. Helia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme, Chris Kedzie |
ACL (1) | 3 |
| 2024 | Dodo: Dynamic Contextual Compression for Decoder-only LMsabstractTransformer-based language models (LMs) are inefficient in long contexts.We propose DODO , a solution for context compression.Instead of one vector per token in a standard transformer model, DODO represents text with a dynamic number of hidden states at each layer, reducing the cost of self-attention to a fraction of typical time and space.Moreover, off-the-shelf models such as LLAMA can be adapted to DODO by efficient parameter tuning methods such as LoRA.In use, DODO can act as either an autoregressive LM or a context compressor for downstream tasks.We demonstrate through experiments in language modeling, question answering, and summarization that DODO retains capabilities in these tasks, while drastically reducing the overhead during decoding.For example, in the autoencoding task, DODO shrinks context at a 20x compression ratio with a BLEU score of 98% for reconstruction, achieving nearly lossless encoding. Guanghui Qin, Corby Rosset, Ethan C. Chau, Benjamin Van Durme |
ACL (1) | 2 |
| 2023 | Augmenting Zero-Shot Dense Retrievers with Plug-in Mixture-of-MemoriesabstractIn this paper we improve the zero-shot generalization ability of language models via Mixture-Of-Memory Augmentation (MoMA), a mechanism that retrieves augmentation documents from multiple information corpora ("external memories"), with the option to "plug in" unseen memory at inference time.We develop a joint learning mechanism that trains the augmentation component with latent labels derived from the end retrieval task, paired with hard negatives from the memory mixture.We instantiate the model in a zero-shot dense retrieval setting by augmenting strong T5-based retrievers with MoMA.With only T5-base, our model obtains strong zero-shot retrieval accuracy on the eighteen tasks included in the standard BEIR benchmark, outperforming some systems with larger model sizes.As a plug-inplay model, our model can efficiently generalize to any unseen corpus, meanwhile achieving comparable or even better performance than methods relying on target-specific pretraining.Our analysis further illustrates the necessity of augmenting with mixture-of-memory for robust generalization, the benefits of augmentation learning, and how MoMA utilizes the plugin memory at inference time without changing its parameters. Suyu Ge, Chenyan Xiong, Corby Rosset, Arnold Overwijk, Jiawei Han 0001, Paul N. Bennett |
EMNLP | 3 |
| 2023 | Axiomatic Preference Modeling for Longform Question AnsweringabstractThe remarkable abilities of large language models (LLMs) like GPT-4 partially stem from posttraining processes like Reinforcement Learning from Human Feedback (RLHF) involving human preferences encoded in a reward model.However, these reward models (RMs) often lack direct knowledge of why, or under what principles, the preferences annotations were made.In this study, we identify principles that guide RMs to better align with human preferences, and then develop an axiomatic framework to generate a rich variety of preference signals to uphold them.We use these axiomatic signals to train a model for scoring answers to longform questions.Our approach yields a Preference Model with only about 220M parameters that agrees with gold humanannotated preference labels more often than GPT-4.The contributions of this work include: training a standalone preference model that can score human-and LLM-generated answers on the same scale; developing an axiomatic framework for generating training data pairs tailored to certain principles; and showing that a small amount of axiomatic signals can help small models outperform GPT-4 in preference scoring.We intend to release our model. Corby Rosset, Guoqing Zheng, Victor Dibia, Ahmed Awadallah 0001, Paul N. Bennett |
EMNLP | 1 |
| 2023 | Zero-shot Clarifying Question Generation for Conversational SearchabstractA long-standing challenge for search and conversational assistants is query intention detection in ambiguous queries. Asking clarifying questions in conversational search has been widely studied and considered an effective solution to resolve query ambiguity. Existing work have explored various approaches for clarifying question ranking and generation. However, due to the lack of real conversational search data, they have to use artificial datasets for training, which limits their generalizability to real-world search scenarios. As a result, the industry has shown reluctance to implement them in reality, further suspending the availability of real conversational search interaction data. The above dilemma can be formulated as a cold start problem of clarifying question generation and conversational search in general. Furthermore, even if we do have large-scale conversational logs, it is not realistic to gather training data that can comprehensively cover all possible queries and topics in open-domain search scenarios. The risk of fitting bias when training a clarifying question retrieval/generation model on incomprehensive dataset is thus another important challenge. Zhenduo Wang, Yuancheng Tu, Corby Rosset, Nick Craswell, Qingyao Ai |
WWW | 3 |
| 2020 | Transformer-XH: Multi-Evidence Reasoning with eXtra Hop Attention
Chen Zhao 0013, Chenyan Xiong, Corby Rosset, Paul N. Bennett, Saurabh Tiwary |
ICLR | 3 |
| 2020 | Contextual Re-Ranking with Behavior Aware TransformersabstractIn this work, we focus on the contextual document ranking task, which deals with the challenge of user interaction modeling for conversational search. Given a history of user feedback behaviors, such as issuing a query, clicking a document, and skipping a document, we propose to introduce behavior awareness to a neural ranker, resulting in a Hierarchical Behavior Aware Transformers (HBA-Transformers) model. The hierarchy is composed of an intra-behavior attention layer and an inter-behavior attention layer to let the system effectively distinguish and model different user behaviors. Our extensive experiments on the AOL session dataset demonstrate that the hierarchical behavior aware architecture is more powerful than a simple combination of history behaviors. Besides, we analyze the conversational property of queries. We show that coherent sessions tend to be more conversational and thus are more demanding in terms of considering history user behaviors. Chen Qu 0001, Chenyan Xiong, Yizhe Zhang 0002, Corby Rosset, W. Bruce Croft, Paul N. Bennett |
SIGIR | 4 |
| 2019 | An Axiomatic Approach to Regularizing Neural Ranking ModelsabstractAxiomatic information retrieval (IR) seeks a set of principle properties desirable in IR models. These properties when formally expressed provide guidance in the search for better relevance estimation functions. Neural ranking models typically contain many learnable parameters. The training of these models involves a search for appropriate parameter values based on large quantities of labeled examples. Intuitively, axioms that can guide the search for better traditional IR models should also help in better parameter estimation for machine learning based rankers. This work explores the use of IR axioms to augment the direct supervision from labeled data for training neural ranking models. We modify the documents in our dataset along the lines of well-known axioms during training and add a regularization loss based on the agreement between the ranking model and the axioms on which version of the document---the original or the perturbed---should be preferred. Our experiments show that the neural ranking model achieves faster convergence and better generalization with axiomatic regularization. Corby Rosset, Bhaskar Mitra 0001, Chenyan Xiong, Nick Craswell, Saurabh Tiwary |
SIGIR | 1 |
| 2019 | Generic Intent Representation in Web SearchabstractThis paper presents GEneric iNtent Encoder (GEN Encoder) which learns a distributed representation space for user intent in search. Leveraging large scale user clicks from Bing search logs as weak supervision of user intent, GEN Encoder learns to map queries with shared clicks into similar embeddings end-to-end and then fine-tunes on multiple paraphrase tasks. Experimental results on an intrinsic evaluation task - query intent similarity modeling - demonstrate GEN Encoder's robust and significant advantages over previous representation methods. Ablation studies reveal the crucial role of learning from implicit user feedback in representing user intent and the contributions of multi-task learning in representation generality. We also demonstrate that GEN Encoder alleviates the sparsity of tail search traffic and cuts down half of the unseen queries by using an efficient approximate nearest neighbor search to effectively identify previous queries with the same search intent. Finally, we demonstrate distances between GEN encodings reflect certain information seeking behaviors in search sessions. Chenyan Xiong, Corby Rosset, Paul N. Bennett, Nick Craswell, Saurabh Tiwary |
SIGIR | 4 |
| 2018 | Optimizing Query Evaluations Using Reinforcement Learning for Web SearchabstractIn web search, typically a candidate generation step selects a small set of documents---from collections containing as many as billions of web pages---that are subsequently ranked and pruned before being presented to the user. In Bing, the candidate generation involves scanning the index using statically designed match plans that prescribe sequences of different match criteria and stopping conditions. In this work, we pose match planning as a reinforcement learning task and observe up to 20% reduction in index blocks accessed, with small or no degradation in the quality of the candidate sets. Corby Rosset, Damien Jose, Gargi Ghosh, Bhaskar Mitra 0001, Saurabh Tiwary |
SIGIR | 1 |