Corby Rosset

dblp:218/5665 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0001-9167-6214ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Reinforcement learning · 36% Language models and text generation · 33% Question answering and dialogue systems · 18%
Databases, data mining, and information retrieval
6 papers
Information retrieval · 88% Recommender systems · 7% Query processing and optimization · 5%

Topics — the 30 heaviest of 33, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
1.122025
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF · ICLR 2025
Axiomatic Preference Modeling for Longform Question Answering · EMNLP 2023
Machine learning › Reinforcement learning
exploration
0.912025
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF · ICLR 2025
Machine learning › Reinforcement learning › exploration
online exploration
0.912025
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF · ICLR 2025
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.912025
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF · ICLR 2025
Information retrieval › interactive information retrieval › conversational information seeking
conversational search
0.822023
Zero-shot Clarifying Question Generation for Conversational Search · WWW 2023
Contextual Re-Ranking with Behavior Aware Transformers · SIGIR 2020
Machine learning › Efficient and distributed learning › inference efficiency
context compression
0.812024
Dodo: Dynamic Contextual Compression for Decoder-only LMs · ACL (1) 2024
Natural language and speech › Language models and text generation › large language model evaluation › automatic evaluation
rubric-based evaluation
0.812024
LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts · ACL (1) 2024
Natural language and speech › Language models and text generation
text evaluation
0.812024
LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts · ACL (1) 2024
Machine learning › Reinforcement learning
human feedback
0.712023
Axiomatic Preference Modeling for Longform Question Answering · EMNLP 2023
Natural language and speech › Question answering and dialogue systems
long-form question answering
0.712023
Axiomatic Preference Modeling for Longform Question Answering · EMNLP 2023
Natural language and speech › Question answering and dialogue systems
question generation
0.712023
Zero-shot Clarifying Question Generation for Conversational Search · WWW 2023
Machine learning › Reinforcement learning › reward learning
reward modeling
0.712023
Axiomatic Preference Modeling for Longform Question Answering · EMNLP 2023
Information retrieval › interactive information retrieval › conversational information seeking › conversational search
clarifying question generation
0.712023
Zero-shot Clarifying Question Generation for Conversational Search · WWW 2023
Information retrieval › retrieval models › neural retrieval
dense retrieval
0.712023
Augmenting Zero-Shot Dense Retrievers with Plug-in Mixture-of-Memories · EMNLP 2023
Information retrieval
retrieval augmentation
0.712023
Augmenting Zero-Shot Dense Retrievers with Plug-in Mixture-of-Memories · EMNLP 2023
Information retrieval › document retrieval
zero-shot retrieval
0.712023
Augmenting Zero-Shot Dense Retrievers with Plug-in Mixture-of-Memories · EMNLP 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning › information fusion
multi-evidence reasoning
0.412020
Transformer-XH: Multi-Evidence Reasoning with eXtra Hop Attention · ICLR 2020
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering
0.412020
Transformer-XH: Multi-Evidence Reasoning with eXtra Hop Attention · ICLR 2020
Information retrieval › ranking › text ranking
document ranking
0.412020
Contextual Re-Ranking with Behavior Aware Transformers · SIGIR 2020
Information retrieval › retrieval models
axiomatic retrieval model
0.412019
An Axiomatic Approach to Regularizing Neural Ranking Models · SIGIR 2019
Information retrieval › retrieval models › neural retrieval
neural ranking model
0.412019
An Axiomatic Approach to Regularizing Neural Ranking Models · SIGIR 2019
Information retrieval › query understanding
query representation
0.412019
Generic Intent Representation in Web Search · SIGIR 2019
Information retrieval
query understanding
0.412019
Generic Intent Representation in Web Search · SIGIR 2019
Recommender systems › large-scale recommendation › multi-stage recommender systems
candidate generation
0.312018
Optimizing Query Evaluations Using Reinforcement Learning for Web Search · SIGIR 2018
Query processing and optimization › query execution › scan processing
index scan
0.312018
Optimizing Query Evaluations Using Reinforcement Learning for Web Search · SIGIR 2018
Information retrieval
query processing
0.312018
Optimizing Query Evaluations Using Reinforcement Learning for Web Search · SIGIR 2018
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization
0.312025
Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF · ICLR 2025
Natural language and speech › Question answering and dialogue systems
dialogue evaluation
0.212024
LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts · ACL (1) 2024
Machine learning › Efficient and distributed learning
model compression
0.212024
Dodo: Dynamic Contextual Compression for Decoder-only LMs · ACL (1) 2024
Recommender systems › user modeling
user interaction modeling
0.112020
Contextual Re-Ranking with Behavior Aware Transformers · SIGIR 2020

Methods — techniques the papers use, named apart from their topics

zero-shot learning · 1.3implicit q*-approximation · 0.9exploration bonus · 0.9KL-regularized MDP · 0.9parameter tuning · 0.8large language model prompting · 0.8feed-forward neural network calibration · 0.8LoRA · 0.8t5 · 0.7reinforcement learning from human feedback · 0.7joint learning · 0.7hard negative mining · 0.7axiomatic preference modeling · 0.7transformer · 0.4attention mechanism · 0.4regularization loss · 0.4data perturbation · 0.4click supervision · 0.4
YearPublicationVenuePosition
2025 Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF
abstract
This paper investigates a basic question in reinforcement learning from human feedback (RLHF) from a theoretical perspective: how to efficiently explore in an online manner under preference feedback and general function approximation. We take the initial step towards a theoretical understanding of this problem by proposing a novel algorithm, *Exploratory Preference Optimization* (XPO). This algorithm is elegantly simple---requiring only a one-line modification to (online) Direct Preference Optimization (DPO; Rafailov et al., 2023)---yet provides the strongest known provable guarantees. XPO augments the DPO objective with a novel and principled *exploration bonus*, enabling the algorithm to strategically explore beyond the support of the initial model and preference feedback data. We prove that XPO is provably sample-efficient and converges to a near-optimal policy under natural exploration conditions, regardless of the initial model's coverage. Our analysis builds on the observation that DPO implicitly performs a form of *Bellman error minimization*. It synthesizes previously disparate techniques from language modeling and theoretical reinforcement learning in a serendipitous fashion through the lens of *KL-regularized Markov decision processes*.
Tengyang Xie, Dylan J. Foster, Akshay Krishnamurthy, Corby Rosset, Ahmed Awadallah 0001, Alexander Rakhlin
ICLR4
2024 LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts
abstract
This paper introduces a framework for the automated evaluation of natural language texts.A manually constructed rubric describes how to assess multiple dimensions of interest.To evaluate a text, a large language model (LLM) is prompted with each rubric question and produces a distribution over potential responses.The LLM predictions often fail to agree well with human judges-indeed, the humans do not fully agree with one another.However, the multiple LLM distributions can be combined to predict each human judge's annotations on all questions, including a summary question that assesses overall quality or relevance.LLM-RUBRIC accomplishes this by training a small feed-forward neural network that includes both judge-specific and judge-independent parameters.When evaluating dialogue systems in a human-AI information-seeking task, we find that LLM-RUBRIC with 9 questions (assessing dimensions such as naturalness, conciseness, and citation quality) predicts human judges' assessment of overall user satisfaction, on a scale of 1-4, with RMS error ă 0.5, a 2ˆ improvement over the uncalibrated baseline.
Helia Hashemi, Jason Eisner, Corby Rosset, Benjamin Van Durme, Chris Kedzie
ACL (1)3
2024 Dodo: Dynamic Contextual Compression for Decoder-only LMs
abstract
Transformer-based language models (LMs) are inefficient in long contexts.We propose DODO , a solution for context compression.Instead of one vector per token in a standard transformer model, DODO represents text with a dynamic number of hidden states at each layer, reducing the cost of self-attention to a fraction of typical time and space.Moreover, off-the-shelf models such as LLAMA can be adapted to DODO by efficient parameter tuning methods such as LoRA.In use, DODO can act as either an autoregressive LM or a context compressor for downstream tasks.We demonstrate through experiments in language modeling, question answering, and summarization that DODO retains capabilities in these tasks, while drastically reducing the overhead during decoding.For example, in the autoencoding task, DODO shrinks context at a 20x compression ratio with a BLEU score of 98% for reconstruction, achieving nearly lossless encoding.
Guanghui Qin, Corby Rosset, Ethan C. Chau, Benjamin Van Durme
ACL (1)2
2023 Augmenting Zero-Shot Dense Retrievers with Plug-in Mixture-of-Memories
abstract
In this paper we improve the zero-shot generalization ability of language models via Mixture-Of-Memory Augmentation (MoMA), a mechanism that retrieves augmentation documents from multiple information corpora ("external memories"), with the option to "plug in" unseen memory at inference time.We develop a joint learning mechanism that trains the augmentation component with latent labels derived from the end retrieval task, paired with hard negatives from the memory mixture.We instantiate the model in a zero-shot dense retrieval setting by augmenting strong T5-based retrievers with MoMA.With only T5-base, our model obtains strong zero-shot retrieval accuracy on the eighteen tasks included in the standard BEIR benchmark, outperforming some systems with larger model sizes.As a plug-inplay model, our model can efficiently generalize to any unseen corpus, meanwhile achieving comparable or even better performance than methods relying on target-specific pretraining.Our analysis further illustrates the necessity of augmenting with mixture-of-memory for robust generalization, the benefits of augmentation learning, and how MoMA utilizes the plugin memory at inference time without changing its parameters.
Suyu Ge, Chenyan Xiong, Corby Rosset, Arnold Overwijk, Jiawei Han 0001, Paul N. Bennett
EMNLP3
2023 Axiomatic Preference Modeling for Longform Question Answering
abstract
The remarkable abilities of large language models (LLMs) like GPT-4 partially stem from posttraining processes like Reinforcement Learning from Human Feedback (RLHF) involving human preferences encoded in a reward model.However, these reward models (RMs) often lack direct knowledge of why, or under what principles, the preferences annotations were made.In this study, we identify principles that guide RMs to better align with human preferences, and then develop an axiomatic framework to generate a rich variety of preference signals to uphold them.We use these axiomatic signals to train a model for scoring answers to longform questions.Our approach yields a Preference Model with only about 220M parameters that agrees with gold humanannotated preference labels more often than GPT-4.The contributions of this work include: training a standalone preference model that can score human-and LLM-generated answers on the same scale; developing an axiomatic framework for generating training data pairs tailored to certain principles; and showing that a small amount of axiomatic signals can help small models outperform GPT-4 in preference scoring.We intend to release our model.
Corby Rosset, Guoqing Zheng, Victor Dibia, Ahmed Awadallah 0001, Paul N. Bennett
EMNLP1
2023 Zero-shot Clarifying Question Generation for Conversational Search
abstract
A long-standing challenge for search and conversational assistants is query intention detection in ambiguous queries. Asking clarifying questions in conversational search has been widely studied and considered an effective solution to resolve query ambiguity. Existing work have explored various approaches for clarifying question ranking and generation. However, due to the lack of real conversational search data, they have to use artificial datasets for training, which limits their generalizability to real-world search scenarios. As a result, the industry has shown reluctance to implement them in reality, further suspending the availability of real conversational search interaction data. The above dilemma can be formulated as a cold start problem of clarifying question generation and conversational search in general. Furthermore, even if we do have large-scale conversational logs, it is not realistic to gather training data that can comprehensively cover all possible queries and topics in open-domain search scenarios. The risk of fitting bias when training a clarifying question retrieval/generation model on incomprehensive dataset is thus another important challenge.
Zhenduo Wang, Yuancheng Tu, Corby Rosset, Nick Craswell, Qingyao Ai
WWW3
2020 Transformer-XH: Multi-Evidence Reasoning with eXtra Hop Attention
Chen Zhao 0013, Chenyan Xiong, Corby Rosset, Paul N. Bennett, Saurabh Tiwary
ICLR3
2020 Contextual Re-Ranking with Behavior Aware Transformers
abstract
In this work, we focus on the contextual document ranking task, which deals with the challenge of user interaction modeling for conversational search. Given a history of user feedback behaviors, such as issuing a query, clicking a document, and skipping a document, we propose to introduce behavior awareness to a neural ranker, resulting in a Hierarchical Behavior Aware Transformers (HBA-Transformers) model. The hierarchy is composed of an intra-behavior attention layer and an inter-behavior attention layer to let the system effectively distinguish and model different user behaviors. Our extensive experiments on the AOL session dataset demonstrate that the hierarchical behavior aware architecture is more powerful than a simple combination of history behaviors. Besides, we analyze the conversational property of queries. We show that coherent sessions tend to be more conversational and thus are more demanding in terms of considering history user behaviors.
Chen Qu 0001, Chenyan Xiong, Yizhe Zhang 0002, Corby Rosset, W. Bruce Croft, Paul N. Bennett
SIGIR4
2019 An Axiomatic Approach to Regularizing Neural Ranking Models
abstract
Axiomatic information retrieval (IR) seeks a set of principle properties desirable in IR models. These properties when formally expressed provide guidance in the search for better relevance estimation functions. Neural ranking models typically contain many learnable parameters. The training of these models involves a search for appropriate parameter values based on large quantities of labeled examples. Intuitively, axioms that can guide the search for better traditional IR models should also help in better parameter estimation for machine learning based rankers. This work explores the use of IR axioms to augment the direct supervision from labeled data for training neural ranking models. We modify the documents in our dataset along the lines of well-known axioms during training and add a regularization loss based on the agreement between the ranking model and the axioms on which version of the document---the original or the perturbed---should be preferred. Our experiments show that the neural ranking model achieves faster convergence and better generalization with axiomatic regularization.
Corby Rosset, Bhaskar Mitra 0001, Chenyan Xiong, Nick Craswell, Saurabh Tiwary
SIGIR1
2019 Generic Intent Representation in Web Search
abstract
This paper presents GEneric iNtent Encoder (GEN Encoder) which learns a distributed representation space for user intent in search. Leveraging large scale user clicks from Bing search logs as weak supervision of user intent, GEN Encoder learns to map queries with shared clicks into similar embeddings end-to-end and then fine-tunes on multiple paraphrase tasks. Experimental results on an intrinsic evaluation task - query intent similarity modeling - demonstrate GEN Encoder's robust and significant advantages over previous representation methods. Ablation studies reveal the crucial role of learning from implicit user feedback in representing user intent and the contributions of multi-task learning in representation generality. We also demonstrate that GEN Encoder alleviates the sparsity of tail search traffic and cuts down half of the unseen queries by using an efficient approximate nearest neighbor search to effectively identify previous queries with the same search intent. Finally, we demonstrate distances between GEN encodings reflect certain information seeking behaviors in search sessions.
Chenyan Xiong, Corby Rosset, Paul N. Bennett, Nick Craswell, Saurabh Tiwary
SIGIR4
2018 Optimizing Query Evaluations Using Reinforcement Learning for Web Search
abstract
In web search, typically a candidate generation step selects a small set of documents---from collections containing as many as billions of web pages---that are subsequently ranked and pruned before being presented to the user. In Bing, the candidate generation involves scanning the index using statically designed match plans that prescribe sequences of different match criteria and stopping conditions. In this work, we pose match planning as a reinforcement learning task and observe up to 20% reduction in index blocks accessed, with small or no degradation in the quality of the candidate sets.
Corby Rosset, Damien Jose, Gargi Ghosh, Bhaskar Mitra 0001, Saurabh Tiwary
SIGIR1