Orion Weller

dblp:248/7910 · DBLP profile ↗
← Back
20ranked-venue papers
11as first author
16since 2021 · last 2026
0000-0003-4148-5430ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 10 first-author · 13 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Auto-ARGUE: LLM-Based Report Generation Evaluation
abstract
Generation of citation-backed reports is a primary use case for retrieval-augmented generation (RAG) systems. While open-source evaluation tools exist for various RAG tasks, tools designed for report generation are lacking. Accordingly, we introduce Auto-ARGUE, a robust LLM-based implementation of the recently proposed ARGUE framework for report generation evaluation. We present analysis of Auto-ARGUE on the report generation pilot task from the TREC 2024 NeuCLIR track and on two tasks from the TREC 2024 RAG track, showing good system-level correlations with human judgments. Additionally, we release ARGUE-viz, a web app for visualization and fine-grained analysis of Auto-ARGUE judgments and scores1.
William Gantt Walden, Marc Mason, Orion Weller, Laura Dietz, John M. Conroy, Neil P. Molino, Hannah Recknor, Bryan Li, Gabrielle K. Liu, Dawn J. Lawrie, James Mayfield, Eugene Yang 0001
SIGIR3
2025 SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
abstract
Can LLMs consistently improve their previous outputs for better results? For this to be true, LLMs would need to be better at discriminating among previously-generated alternatives, than generating initial responses. We explore the validity of this hypothesis in practice. We first formulate a unified framework that allows us to compare the generative and discriminative capability of any model on any task. In our resulting experimental analysis of several open-source and industrial LLMs, we observe that model’s are not reliably better at discriminating among previously-generated alternatives than generating initial responses. This finding challenges the notion that LLMs may be able to enhance their performance only through their own judgment.
Dongwei Jiang, Orion Weller, Nathaniel Weir, Benjamin Van Durme, Daniel Khashabi
AAAI3
2025 Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference
abstract
Benjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller, Oskar Hallström, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, Griffin Thomas Adams, Jeremy Howard, Iacopo Poli. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Benjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller, Oskar Hallström, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, Griffin Adams, Jeremy Howard, Iacopo Poli
ACL (1)4
2025 mFollowIR: A Multilingual Benchmark for Instruction Following in Retrieval
Orion Weller, Benjamin Chang 0007, Eugene Yang 0001, Mahsa Yarmohammadi, Samuel Barham, Sean MacAvaney, Arman Cohan, Luca Soldaini, Benjamin Van Durme, Dawn J. Lawrie
ECIR (2)1
2025 From Models to Microtheories: Distilling a Model's Topical Knowledge for Grounded Question-Answering
abstract
Recent reasoning methods (e.g., chain-of-thought) help users understand how language models (LMs) answer a single question, but they do little to reveal the LM’s overall understanding, or “theory,” about the question’s topic, making it still hard to trust the model. Our goal is to materialize such theories - here called microtheories (a linguistic analog of logical microtheories) - as a set of sentences encapsulating an LM’s core knowledge about a topic. These statements systematically work together to entail answers to a set of questions to both engender trust and improve performance. Our approach is to first populate a knowledge store with (model-generated) sentences that entail answers to training questions, and then distill those down to a core microtheory which is concise, general, and non-redundant. We show that, when added to a general corpus (e.g., Wikipedia), microtheories can supply critical information not necessarily present in the corpus, improving both a model’s ability to ground its answers to verifiable knowledge (i.e., show how answers are systematically entailed by documents in the corpus, grounding up to +8% more answers), and the accuracy of those grounded answers (up to +8% absolute). We also show that, in a human evaluation in the medical domain, our distilled microtheories contain a significantly higher concentration of topically critical facts than the non-distilled knowledge store. Finally, we show we can quantify the coverage of a microtheory for a topic (characterized by a dataset) using a notion of p-relevance. Together, these suggest that microtheories are an efficient distillation of an LM’s topic-relevant knowledge, that they can usefully augment existing corpora, and can provide both performance gains and an interpretable, verifiable window into the model’s knowledge of a topic.
Nathaniel Weir, Bhavana Dalvi, Orion Weller, Oyvind Tafjord, Sam Hornstein, Alexander Sabol, Peter A. Jansen, Benjamin Van Durme, Peter Clark
ICLR3
2025 Promptriever: Instruction-Trained Retrievers Can Be Prompted Like Language Models
abstract
Instruction-tuned language models (LM) are able to respond to imperative commands, providing a more natural user interface compared to their base counterparts. In this work, we present Promptriever, the first retrieval model able to be prompted like an LM. To train Promptriever, we curate and release a new instance-level instruction training set from MS MARCO, spanning nearly 500k instances. Promptriever not only achieves strong performance on standard retrieval tasks, but also follows instructions. We observe: (1) large gains (reaching SoTA) on following detailed relevance instructions (+14.3 p-MRR / +3.1 nDCG on FollowIR), (2) significantly increased robustness to lexical choices/phrasing in the query+instruction (+12.9 Robustness@10 on InstructIR), and (3) the ability to perform hyper-parameter search via prompting to reliably improve retrieval performance (+1.4 average increase on BEIR). Promptriever demonstrates that retrieval models can be controlled with prompts on a per-query basis, setting the stage for future work aligning LM prompting techniques with information retrieval.
Orion Weller, Benjamin Van Durme, Dawn J. Lawrie, Ashwin Paranjape, Yuhao Zhang 0004, Jack Hessel
ICLR1
2025 FollowIR: Evaluating and Teaching Information Retrieval Models to Follow Instructions
abstract
Orion Weller, Benjamin Chang, Sean MacAvaney, Kyle Lo, Arman Cohan, Benjamin Van Durme, Dawn Lawrie, Luca Soldaini. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Orion Weller, Benjamin Chang 0007, Sean MacAvaney, Kyle Lo, Arman Cohan, Benjamin Van Durme, Dawn J. Lawrie, Luca Soldaini
NAACL (Long Papers)1
2024 NevIR: Negation in Neural Information Retrieval
abstract
Negation is a common everyday phenomena and has been a consistent area of weakness for language models (LMs).Although the Information Retrieval (IR) community has adopted LMs as the backbone of modern IR architectures, there has been little to no research in understanding how negation impacts neural IR.We therefore construct a straightforward benchmark on this theme: asking IR models to rank two documents that differ only by negation.We show that the results vary widely according to the type of IR architecture: cross-encoders perform best, followed by late-interaction models, and in last place are bi-encoder and sparse neural architectures.We find that most information retrieval models (including SOTA ones) do not consider negation, performing the same or worse than a random ranking.We show that although the obvious approach of continued finetuning on a dataset of contrastive documents containing negations increases performance (as does model size), there is still a large gap between machine and human performance.1
Orion Weller, Dawn J. Lawrie, Benjamin Van Durme
EACL (1)1
2024 "According to . . . ": Prompting Language Models Improves Quoting from Pre-Training Data
abstract
Orion Weller, Marc Marone, Nathaniel Weir, Dawn Lawrie, Daniel Khashabi, Benjamin Van Durme. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Orion Weller, Marc Marone, Nathaniel Weir, Dawn J. Lawrie, Daniel Khashabi, Benjamin Van Durme
EACL (1)1
2024 Enhancing Systematic Decompositional Natural Language Inference Using Informal Logic
abstract
Nathaniel Weir, Kate Sanders, Orion Weller, Shreya Sharma, Dongwei Jiang, Zhengping Jiang, Bhavana Dalvi Mishra, Oyvind Tafjord, Peter Jansen, Peter Clark, Benjamin Van Durme. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Nathaniel Weir, Kate Sanders 0002, Orion Weller, Shreya Sharma 0010, Dongwei Jiang, Zhengping Jiang, Bhavana Dalvi, Oyvind Tafjord, Peter A. Jansen, Peter Clark, Benjamin Van Durme
EMNLP3
2024 Learning to Reason via Program Generation, Emulation, and Search
abstract
Program synthesis with language models (LMs) has unlocked a large set of reasoning abilities; code-tuned LMs have proven adept at generating programs that solve a wide variety of algorithmic symbolic manipulation tasks (e.g. word concatenation). However, not all reasoning tasks are easily expressible as code, e.g. tasks involving commonsense reasoning, moral decision-making, and sarcasm understanding. Our goal is to extend a LM’s program synthesis skills to such tasks and evaluate the results via pseudo-programs, namely Python programs where some leaf function calls are left undefined. To that end, we propose, Code Generation and Emulated EXecution (COGEX). COGEX works by (1) training LMs to generate pseudo-programs and (2) teaching them to emulate their generated program’s execution, including those leaf functions, allowing the LM’s knowledge to fill in the execution gaps; and (3) using them to search over many programs to find an optimal one. To adapt the COGEX model to a new task, we introduce a method for performing program search to find a single program whose pseudo-execution yields optimal performance when applied to all the instances of a given dataset. We show that our approach yields large improvements compared to standard in-context learning approaches on a battery of tasks, both algorithmic and soft reasoning. This result thus demonstrates that code synthesis can be applied to a much broader class of problems than previously considered.
Nathaniel Weir, Muhammad Khalifa, Linlu Qiu, Orion Weller, Peter Clark
NeurIPS4
2024 On the Evaluation of Machine-Generated Reports
abstract
Large Language Models (LLMs) have enabled new ways to satisfy information needs. Although great strides have been made in applying them to settings like document ranking and short-form text generation, they still struggle to compose complete, accurate, and verifiable long-form reports. Reports with these qualities are necessary to satisfy the complex, nuanced, or multi-faceted information needs of users. In this perspective paper, we draw together opinions from industry and academia, and from a variety of related research areas, to present our vision for automatic report generation, and---critically---a flexible framework by which such reports can be evaluated. In contrast with other summarization tasks, automatic report generation starts with a detailed description of an information need, stating the necessary background, requirements, and scope of the report. Further, the generated reports should be complete, accurate, and verifiable. These qualities, which are desirable---if not required---in many analytic report-writing settings, require rethinking how to build and evaluate systems that exhibit these qualities. To foster new efforts in building these systems, we present an evaluation framework that draws on ideas found in various evaluations. To test completeness and accuracy, the framework uses nuggets of information, expressed as questions and answers, that need to be part of any high-quality generated report. Additionally, evaluation of citations that map claims made in the report to their source documents ensures verifiability.
James Mayfield, Eugene Yang 0001, Dawn J. Lawrie, Sean MacAvaney, Paul McNamee, Douglas W. Oard, Luca Soldaini, Ian Soboroff, Orion Weller, Efsun Selin Kayi, Kate Sanders 0002, Marc Mason, Noah Hibbler
SIGIR9
2023 When Do Decompositions Help for Machine Reading?
abstract
Answering complex questions often requires multi-step reasoning in order to obtain the final answer.Most research into decompositions of complex questions involves open-domain systems, which have shown success in using these decompositions for improved retrieval.In the machine reading setting, however, work to understand when decompositions are helpful is understudied.We conduct experiments on decompositions in machine reading to unify recent work in this space, using a range of models and datasets.We find that decompositions can be helpful in zero or limited-data settings, giving several points of improvement in exact match.However, we also show that when models are given access to around a few hundred or more examples, decompositions are not helpful (and can actually be detrimental).Thus, our analysis implies that models can learn decompositions implicitly even with limited data. 1
Kangda Wei, Dawn J. Lawrie, Benjamin Van Durme, Yunmo Chen, Orion Weller
EMNLP5
2022 Pretrained Models for Multilingual Federated Learning
abstract
Orion Weller, Marc Marone, Vladimir Braverman, Dawn Lawrie, Benjamin Van Durme. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Orion Weller, Marc Marone, Vladimir Braverman, Dawn J. Lawrie, Benjamin Van Durme
NAACL-HLT1
2021 Streaming Models for Joint Speech Recognition and Translation
abstract
Using end-to-end models for speech translation (ST) has increasingly been the focus of the ST community.These models condense the previously cascaded systems by directly converting sound waves into translated text.However, cascaded models have the advantage of including automatic speech recognition output, useful for a variety of practical ST systems that often display transcripts to the user alongside the translations.To bridge this gap, recent work has shown initial progress into the feasibility for end-to-end models to produce both of these outputs.However, all previous work has only looked at this problem from the consecutive perspective, leaving uncertainty on whether these approaches are effective in the more challenging streaming setting.We develop an end-to-end streaming ST model based on a re-translation approach and compare against standard cascading approaches.We also introduce a novel inference method for the joint case, interleaving both transcript and translation in generation and removing the need to use separate decoders.Our evaluation across a range of metrics capturing accuracy, latency, and consistency shows that our end-to-end models are statistically similar to cascading models, while having half the number of parameters.We also find that both systems provide strong translation quality at low latency, keeping 99% of consecutive quality at a lag of just under a second.
Orion Weller, Matthias Sperber, Christian Gollan, Joris Kluivers
EACL1
2021 Exploring the Relationship Between Algorithm Performance, Vocabulary, and Run-Time in Text Classification
abstract
Text classification is a significant branch of natural language processing, and has many applications including document classification and sentiment analysis.Unsurprisingly, those who do text classification are concerned with the run-time of their algorithms, many of which depend on the size of the corpus' vocabulary due to their bag-of-words representation.Although many studies have examined the effect of preprocessing techniques on vocabulary size and accuracy, none have examined how these methods affect a model's run-time.To fill this gap, we provide a comprehensive study that examines how preprocessing techniques affect the vocabulary size, model performance, and model run-time, evaluating ten techniques over four models and two datasets.We show that some individual methods can reduce run-time with no loss of accuracy, while some combinations of methods can trade 2-5% of the accuracy for up to a 65% reduction of run-time.Furthermore, some combinations of preprocessing techniques can even provide a 15% reduction in run-time while simultaneously improving model accuracy.1
Wilson Fearn, Orion Weller, Kevin D. Seppi
NAACL-HLT2
2020 You Don't Have Time to Read This: An Exploration of Document Reading Time Prediction
abstract
Orion Weller, Jordan Hildebrandt, Ilya Reznik, Christopher Challis, E. Shannon Tass, Quinn Snell, Kevin Seppi. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Orion Weller, Jordan Hildebrandt, Ilya Reznik, Christopher Challis, E. Shannon Tass, Quinn Snell, Kevin D. Seppi
ACL1
2020 Learning from Task Descriptions
abstract
Typically, machine learning systems solve new tasks by training on thousands of examples.In contrast, humans can solve new tasks by reading some instructions, with perhaps an example or two.To take a step toward closing this gap, we introduce a framework for developing NLP systems that solve new tasks after reading their descriptions, synthesizing prior work in this area.We instantiate this framework with a new English language dataset, ZEST, structured for task-oriented evaluation on unseen tasks.Formulating task descriptions as questions, we ensure each is general enough to apply to many possible inputs, thus comprehensively evaluating a model's ability to solve each task.Moreover, the dataset's structure tests specific types of systematic generalization.We find that the state-of-the-art T5 model achieves a score of 12% on ZEST, leaving a significant challenge for NLP researchers. 1
Orion Weller, Nicholas Lourie, Matt Gardner 0001, Matthew E. Peters
EMNLP (1)1
2020 The rJokes Dataset: a Large Scale Humor Collection
abstract
Humor is a complicated language phenomenon that depends upon many factors, including topic, date, and recipient. Because of this variation, it can be hard to determine what exactly makes a joke humorous, leading to difficulties in joke identification and related tasks. Furthermore, current humor datasets are lacking in both joke variety and size, with almost all current datasets having less than 100k jokes. In order to alleviate this issue we compile a collection of over 550,000 jokes posted over an 11 year period on the Reddit r/Jokes subreddit (an online forum), providing a large scale humor dataset that can easily be used for a myriad of tasks. This dataset also provides quantitative metrics for the level of humor in each joke, as determined by subreddit user feedback. We explore this dataset through the years, examining basic statistics, most mentioned entities, and sentiment proportions. We also introduce this dataset as a task for future work, where models learn to predict the level of humor in a joke. On that task we provide strong state-of-the-art baseline models and show room for future improvement. We hope that this dataset will not only help those researching computational humor, but also help social scientists who seek to understand popular culture through humor.
Orion Weller, Kevin D. Seppi
LREC1
2019 Humor Detection: A Transformer Gets the Last Laugh
abstract
Orion Weller, Kevin Seppi. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Orion Weller, Kevin D. Seppi
EMNLP/IJCNLP (1)1