VLDB 2026 Research / reviewers in the wild / expert
Adam Roegiest
dblp:122/5812
· DBLP profile ↗
30ranked-venue papers in the field
11as first author
10since 2021 · last 2026
0000-0003-1265-8881ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 30 (11 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Information Farming: From Berry Picking to Berry GrowingabstractThe classic paradigms of Berry Picking and Information Foraging Theory have framed users as gatherers, opportunistically searching across distributed sources to satisfy evolving information needs. However, the rise of Generative AI (GenAI) is driving a fundamental transformation in how people produce, structure, and reuse information—one that these paradigms no longer fully capture. This transformation is analogous to the Neolithic Revolution, when societies shifted from hunting and gathering to cultivation. Generative technologies empower users to “farm” information by planting seeds in the form of prompts, cultivating workflows over time, and harvesting richly structured, relevant yields within their own plots, rather than foraging across others people’s patches. In this perspectives paper, we introduce the notion of Information Farming as a conceptual framework and argue that it represents a natural evolution in how people engage with information. Drawing on historical analogy and empirical evidence, we examine the benefits and opportunities of information farming, its implications for design and evaluation, and the accompanying risks posed by this transition. We hypothesize that as GenAI technologies proliferate, cultivating information will increasingly supplant transient, patch-based foraging as a dominant mode of engagement, marking a broader shift in human-information interaction and its study. Leif Azzopardi, Adam Roegiest |
CHIIR | 2 |
| 2026 | Simulation of Interactive Information Retrieval: A Guided TourabstractInteractive information retrieval (IIR) systems, including search engines and conversational systems, are increasingly central to user experiences. However, rigorously evaluating their performance, particularly as interactions become highly personalized, remains a scientific challenge. While user simulation offers a powerful methodology for reproducible evaluation, its adoption is hindered by a steep learning curve and a fragmented landscape of complex tools. This half-day tutorial provides a practical, hands-on introduction to user simulation at varying levels of complexity, from foundational statistical models to advanced, LLM-driven frameworks. Through a series of guided problems, participants will acquire practical skills in using popular libraries, learning user models from data, and applying large language models (LLMs) to simulate user behavior. The tutorial concludes with evaluating the simulators themselves, providing participants with guidance on appropriate use cases and fidelity assessment. Saber Zerhoudi, Adam Roegiest, Johanne R. Trippas |
CHIIR | 2 |
| 2026 | The Third Search Futures Workshop at ECIR'26
Leif Azzopardi, Charles L. A. Clarke, Claudia Hauff, Yubin Kim 0001, Zhaochun Ren, Adam Roegiest, Johanne R. Trippas, Saber Zerhoudi |
ECIR (3) | 6 |
| 2026 | IIRSim Studio: A Dashboard for User SimulationabstractUser simulation is a valuable methodology for evaluation in Information Retrieval (IR), enabling low-cost experimentation and counterfactual analysis. However, existing simulation frameworks are primarily code-centric libraries that require substantial setup effort, which limits adoption and hinders reproducibility. The bottleneck is not the simulation engines themselves, but the lack of infrastructure connecting experiment design, execution, and sharing into a single verifiable workflow. This paper introduces IIRSim Studio, a web-based workbench that addresses this gap through four contributions: (1) a visual environment for composing simulation pipelines on top of simulation frameworks, serving both novices learning simulation concepts and experts piloting large-scale experiments; (2) a component lifecycle that supports authoring, versioning, and sharing custom simulation components through Git-backed storage and runtime injection; (3) a provenance model based on experiment bundles and environment templates that makes the scope of replication explicit; and (4) a shared-task workflow, demonstrated through the re-deployment of a Sim4IA micro-task. IIRSim Studio is available as a hosted service and as a portable containerized deployment. Saber Zerhoudi, Adam Roegiest, Michael Granitzer |
SIGIR | 2 |
| 2025 | Applying Large Language Models to Interactive Information Retrieval: A Practical ExplorationabstractThis half-day interactive tutorial provides researchers with practical skills to use large language models (LLMs) for interactive information retrieval research.Through hands-on exercises and real-world research examples, participants will learn to set up LLMs locally, integrate them via APIs, and evaluate their outputs.The tutorial will explain which models suit specific research needs, offering participants a robust toolkit to enhance their work.Attendees will also gain insights into the latest developments in the field, ensuring they stay at the forefront of innovation.Ideal for researchers eager to explore new methodologies in information retrieval, this tutorial offers foundational knowledge and cutting-edge strategies to use LLMs in interactive information retrieval research. Johanne R. Trippas, Oleg Zendel, Adam Roegiest |
CHIIR | 3 |
| 2025 | The Second Search Futures Workshop at ECIR'25
Charles L. A. Clarke, Paul B. Kantor, Adam Roegiest, Ian Soboroff, Johanne R. Trippas, Zhaochun Ren |
ECIR (5) | 3 |
| 2024 | Generative Information Systems Are Great If You Can ReadabstractGenerative models, especially in information systems like ChatGPT and Bing Chat, have become increasingly integral to our daily lives. Their significance lies in their potential to revolutionize how we access, process, and generate information [44]. However, a gap exists in ensuring these systems are accessible to all, especially considering the literacy challenges faced by a significant portion of the population in (but not limited to) English-speaking countries. This paper aims to investigate the “readability’’ of generative information systems and their accessibility barriers, particularly for those with literacy challenges. Using popular instruction fine-tuning datasets, we found that this training data could produce systems that generate at a college level, potentially excluding a large demographic. Our research methods involved analyzing the responses of popular Large Language Models (LLMs) and examining potential biases in how they can be trained. The key message is the urgent need for inclusivity in systems incorporating generative models, such as those studied by the Information Retrieval (IR) community. Our findings indicate that current generative systems might not be accessible to individuals with cognitive and literacy challenges, emphasizing the importance of ensuring that advancements in this field benefit everyone. By situating our research within the sphere of information seeking and retrieval, we underscore the essential role of these technologies in augmenting accessibility and efficiency of information access, thereby broadening their reach and enhancing user engagement. Adam Roegiest, Zuzana Pinkosova |
CHIIR | 1 |
| 2024 | UnExplored FrontCHIIRs: A Workshop Exploring Future Directions for Information AccessabstractWith the rise and growing prevalence of generative models, particularly multi-modal ones, it is an opportune time to explore beyond existing interactive information retrieval research trends. Indeed, it is essential to determine new avenues to explore how users interact with these models as well as revisit existing avenues that can be embellished with new technology. In this session, we aim to create a venue to workshop ideas that explore the future of search experiences and user interactions with information in a collaborative, low-pressure environment. This UnExplored FrontCHIIRs workshop enables participants to form a sub-community within CHIIR to facilitate further development of the proposed ideas and allow deeper collaborative problem-solving than just presenting late-breaking work. Adam Roegiest, Johanne R. Trippas |
CHIIR | 1 |
| 2024 | Exploring the Utility of Saccade Length in Eye Tracking on Digital Library Experiences
Maja Kuhar, Adam Roegiest, Tanja Mercun |
TPDL (2) | 2 |
| 2023 | Building a Better Mousetrap: Tools and Processes for Selling A CompanyabstractIt is a fact of life for many start-ups that they must sell part of their company (i.e., fund raising) in order to have enough capital to grow the company to one day successfully exit the market. The unfortunate side effect of this necessity is that it places a large burden on start-ups to respond to information requests from potential buyers which then forces employees to step away from their day jobs to formulate responses. While it has been the norm to respond to such requests using manual review of contracts and other information sources, the increasingly competitive funding market has resulted in growing time pressure for all participants of start-up purchasing endeavours. Furthermore, current technological offerings often fall short of providing optimal support to the start-up and the buyer which continues to reinforce a process that is often cumbersome and chaotic. Chelsea Kerr, Alexandra Vtyurina, Adam Roegiest |
CHIIR | 3 |
| 2020 | Spectator: An Open Source Document ViewerabstractMany information retrieval tasks require viewing documents in some manner, whether this is to view information in context or to provide annotations for some downstream task (e.g., evaluation or system training). Building a high-quality document viewer often exceeds the resources of many researchers and so, in this paper, we describe the design and architecture of our new open-source document viewer, Spectator. In particular, we provide a look into the algorithmic details of how Spectator accomplishes tasks like mapping annotations back to the canonical document. Moreover, we provide a sampling of the use cases that we envision for Spectator, potential future additions depending on community need and support, and highlight situations where Spectator may not be a good fit. Furthermore, we provide a brief description of the sample application that we bundle with Spectator to demonstrate how one might use it within the context of a larger system. Jean-Philippe Gauthier, Adam Roegiest |
CHIIR | 2 |
| 2020 | Dancing with the AI Devil: Investigating the Partnership Between Lawyers and AIabstractAs professional users interact with more AI-enabled tools, it has become increasingly important to understand how their work and behaviour are affected by such tools. In this paper, we present the insights that we have gleaned from a qualitative user study conducted with nine of our software's users who are all legal professionals. We find that as our participants become more accustomed to the system they begin to subtly alter their behaviours and interactions with the system. Using their shared experiences, we distill these into insights that may inform the design of similar systems. Mary Mikhail, Adam Roegiest, Karen Anello, Winter Wei |
CHIIR | 2 |
| 2020 | The Utility of Context When Extracting Entities From Legal DocumentsabstractWhen reviewing documents for legal tasks such as Mergers and Acquisitions, granular information (such as start dates and exit clauses) need to be identified and extracted. Inspired by previous work in Named Entity Recognition (NER), we investigate how NER techniques can be leveraged to aid lawyers in this review process. Due to the extremely low prevalence of target information in legal documents, we find that the traditional approach of tagging all sentences in a document is inferior, in both effectiveness and data required to train and predict, to using a first-pass layer to identify sentences that are likely to contain the relevant information and then running the more traditional sentence-level sequence tagging. Moreover, we find that such entity-level models can be improved by training on a balanced sample of relevant and non-relevant sentences. We additionally describe the use of our system in production and how its usage by clients means that deep learning architectures tend to be cost inefficient, especially with respect to the necessary time to train models. Jonathan Donnelly, Adam Roegiest |
CIKM | 2 |
| 2019 | Variations in Assessor Agreement in Due DiligenceabstractIn legal due diligence, lawyers identify a variety of topic instances in a company's contracts that may pose risk during a transaction. In this paper, we present a study of 9 lawyers conducting a simulated review of 50 contracts for five topics. We find that lawyers agree on the general location of relevant material at a higher rate than in other assessor agreement studies, but they do not entirely agree on the extent of the relevant material. Additionally, we do not find strong differences between lawyers who have differing levels of due diligence expertise. Adam Roegiest, Anne McNulty |
CHIIR | 1 |
| 2019 | From Bubbles to Lists: Designing Clustering for Due DiligenceabstractIn due diligence, lawyers are tasked with reviewing a large set of legal documents to identify documents and portions thereof that may be problematic for a merger or acquisition. In an effort to aid users to review more efficiently, we sought to determine how document-level clustering may help users of a due diligence system during their workflow. Winter Wei, Adam Roegiest, Mary Mikhail |
CHIIR | 2 |
| 2019 | On Interpretability and Feature Representations: An Analysis of the Sentiment Neuron
Jonathan Donnelly, Adam Roegiest |
ECIR (1) | 2 |
| 2019 | On Tradeoffs Between Document Signature Methods for a Legal Due Diligence CorpusabstractWhile document signatures are a well established tool in IR, they have primarily been investigated in the context of web documents. Legal due diligence documents, by their nature, have more similar structure and language than we may expect out of standard web collections. Moreover, many due diligence systems strive to facilitate real-time interactions and so time from document ingestion to availability should be minimal. Such constraints further limit the possible solution space when identifying near duplicate documents. We present an examination of the tradeoffs that document signature methods face in the due diligence domain. In particular, we quantify the trade-off between signature length, time to compute, number of hash collisions, and number of nearest neighbours for a 90,000 document due diligence corpus. Adam Roegiest, Edward Lee 0001 |
SIGIR | 1 |
| 2018 | Redesigning a Document Viewer for Legal DocumentsabstractIn Mergers and Acquisition due diligence, lawyers are tasked with analyzing a collection of contracts and determine the level of risk that comes from a merger or acquisition. This process has historically been manual and resulted in only a small fraction of the collection being examined. This paper reports on the user-focused redesign of our document viewer that is used by clients to review documents and train machine learning algorithms to find pertinent information from these contracts. Adam Roegiest, Winter Wei |
CHIIR | 1 |
| 2018 | A Dataset and an Examination of Identifying Passages for Due DiligenceabstractWe present and formalize the due diligence problem, where lawyers extract data from legal documents to assess risk in a potential merger or acquisition, as an information retrieval task. Furthermore, we describe the creation and annotation of a document collection for the due diligence problem that will foster research in this area. This dataset comprises 50 topics over 4,412 documents and ~15 million sentences and is a subset of our own internal training data. Using this dataset, we present what we have found to be the state of the art for information extraction in the due diligence problem. In particular, we find that when treating documents as sequences of labelled and unlabelled sentences, Conditional Random Fields significantly and substantially outperform other techniques for sequence-based (Hidden Markov Models) and non-sequence based machine learning (logistic regression). Included in this is an analysis of what we perceive to be the major failure cases when extraction is performed based upon sentence labels. Adam Roegiest, Alexander K. Hudek, Anne McNulty |
SIGIR | 1 |
| 2017 | Automatic and Semi-Automatic Document Selection for Technology-Assisted ReviewabstractAbstract In the TREC Total Recall Track (2015-2016), participating teams could employ either fully automatic or human-assisted ("semi-automatic") methods to select documents for relevance assessment by a simulated human reviewer. According to the TREC 2016 evaluation, the fully automatic baseline method achieved a recall-precision breakeven ("R-precision") score of 0.71, while the two semi-automatic efforts achieved scores of 0.67 and 0.51. In this work, we investigate the extent to which the observed effectiveness of the different methods may be confounded by chance, by inconsistent adherence to the Track guidelines, by selection bias in the evaluation method, or by discordant relevance assessments. We find no evidence that any of these factors could yield relative effectiveness scores inconsistent with the official TREC 2016 ranking. Maura R. Grossman, Gordon V. Cormack, Adam Roegiest |
SIGIR | 3 |
| 2017 | Online In-Situ Interleaved Evaluation of Real-Time Push Notification SystemsabstractReal-time push notification systems monitor continuous document streams such as social media posts and alert users to relevant content directly on their mobile devices. We describe a user study of such systems in the context of the TREC 2016 Real-Time Summarization Track, where system updates are immediately delivered as push notifications to the mobile devices of a cohort of users. Our study represents, to our knowledge, the first deployment of an interleaved evaluation framework for prospective information needs, and also provides an opportunity to examine user behavior in a realistic setting. Results of our online in-situ evaluation are correlated against the results a more traditional post-hoc batch evaluation. We observe substantial correlations between many online and batch evaluation metrics, especially for those that share the same basic design (e.g., are utility-based). For some metrics, we observe little correlation, but are able to identify the volume of messages that a system pushes as one major source of differences. Adam Roegiest, Luchen Tan, Jimmy Lin |
SIGIR | 1 |
| 2017 | Ten Blue Links on MarsabstractThis paper explores a simple question: How would we provide a high-quality search experience on Mars, where the fundamental physical limit is speed-of-light propagation delays on the order of tens of minutes? On Earth, users are accustomed to nearly instantaneous responses from web services. Is it possible to overcome orders-of-magnitude longer latency to provide a tolerable user experience on Mars? In this paper, we formulate the searching from Mars problem as a tradeoff between "effort" (waiting for responses from Earth) and "data transfer" (pre-fetching or caching data on Mars). The contribution of our work is articulating this design space and presenting two case studies that explore the effectiveness of baseline techniques, using publicly available data from the TREC Total Recall and Sessions Tracks. We intend for this research problem to be aspirational as well as inspirational---even if one is not convinced by the premise of Mars colonization, there are Earth-based scenarios such as searching from rural villages in India that share similar constraints, thus making the problem worthy of exploration and attention from researchers. Charles L. A. Clarke, Gordon V. Cormack, Jimmy Lin, Adam Roegiest |
WWW | 4 |
| 2016 | Interleaved Evaluation for Retrospective Summarization and Prospective Notification on Document StreamsabstractWe propose and validate a novel interleaved evaluation methodology for two complementary information seeking tasks on document streams: retrospective summarization and prospective notification. In the first, the user desires relevant and non-redundant documents that capture important aspects of an information need. In the second, the user wishes to receive timely, relevant, and non-redundant update notifications for a standing information need. Despite superficial similarities, interleaved evaluation methods for web ranking cannot be directly applied to these tasks; for example, existing techniques do not account for temporality or redundancy. Our proposed evaluation methodology consists of two components: a temporal interleaving strategy and a heuristic for credit assignment to handle redundancy. By simulating user interactions with interleaved results on submitted runs to the TREC 2014 tweet timeline generation (TTG) task and the TREC 2015 real-time filtering task, we demonstrate that our methodology yields system comparisons that accurately match the result of batch evaluations. Analysis further reveals weaknesses in current batch evaluation methodologies to suggest future directions for research. Jimmy Lin, Adam Roegiest |
SIGIR | 3 |
| 2016 | Impact of Review-Set Selection on Human Assessment for Text ClassificationabstractIn a laboratory study, human assessors were significantly more likely to judge the same documents as relevant when they were presented for assessment within the context of documents selected using random or uncertainty sampling, as compared to relevance sampling. The effect is substantial and significant [0.54 vs. 0.42, p<0.0002] across a population of documents including both relevant and non-relevant documents, for several definitions of ground truth. This result is in accord with Smucker and Jethani's SIGIR 2010 finding that documents were more likely to be judged relevant when assessed within low-precision versus high-precision ranked lists. Our study supports the notion that relevance is malleable, and that one should take care in assuming any labeling to be ground truth, whether for training, tuning, or evaluating text classifiers. Adam Roegiest, Gordon V. Cormack |
SIGIR | 1 |
| 2016 | An Architecture for Privacy-Preserving and Replicable High-Recall Retrieval ExperimentsabstractWe demonstrate the infrastructure used in the TREC 2015 Total Recall track to facilitate controlled simulation of "assessor in the loop" high-recall retrieval experimentation. The implementation and corresponding design decisions are presented for this platform. This includes the necessary considerations to ensure that experiments are privacy-preserving when using test collections that cannot be distributed. Furthermore, we describe the use of virtual machines as a means of system submission in order to to promote replicable experiments while also ensuring the security of system developers and data providers. Adam Roegiest, Gordon V. Cormack |
SIGIR | 1 |
| 2016 | A Platform for Streaming Push Notifications to Mobile AssessorsabstractWe present an assessment platform for gathering online relevance judgments for mobile push notifications that will be deployed in the newly-created TREC 2016 Real-Time Summarization (RTS) track. There is emerging interest in building systems that filter social media streams such as tweets to identify interesting and novel content in real time, putatively for delivery to users' mobile phones. In our evaluation design, all participants subscribe to the Twitter streaming API to identify relevant tweets with respect to a set of interest profiles. As the systems generate results, they are pushed in real time to our evaluation broker via a REST API. The broker then "routes" the tweets to assessors who have installed a custom app on their mobile phones. We detail the design of this platform and discuss a number of challenges that need to be tackled in this type of "Living Labs" setup. It is our goal that such an evaluation design will mitigate any issues that have arisen in traditional batch-style evaluations of this type of task. Adam Roegiest, Luchen Tan, Jimmy Lin, Charles L. A. Clarke |
SIGIR | 1 |
| 2016 | Simple Dynamic Emission Strategies for Microblog FilteringabstractPush notifications from social media provide a method to keep up-to-date on topics of personal interest. To be effective, notifications must achieve a balance between pushing too much and pushing too little. Push too little and the user misses important updates; push too much and the user is overwhelmed by unwanted information. Using data from the TREC 2015 Microblog track, we explore simple dynamic emission strategies for microblog push notifications. The key to effective notifications lies in establishing and maintaining appropriate thresholds for pushing updates. We explore and evaluate multiple threshold setting strategies, including purely static thresholds, dynamic thresholds without user feedback, and dynamic thresholds with daily feedback. Our best technique takes advantage of daily feedback in a simple yet effective manner, achieving the best known result reported in the literature to date. Luchen Tan, Adam Roegiest, Charles L. A. Clarke, Jimmy Lin |
SIGIR | 2 |
| 2016 | An Exploration of Evaluation Metrics for Mobile Push NotificationsabstractHow do we evaluate systems that filter social media streams and send users updates via push notifications on their mobile phones? Such notifications must be relevant, timely, and novel. In this paper, we explore various evaluation metrics for this task, focusing specifically on measuring relevance. We begin with an analysis of metrics deployed at the TREC 2015 Microblog evaluations. A simple change to the metrics, reflecting a different assumption, dramatically alters system rankings. Applying another metric, previously used in the TREC Microblog evaluations, again yields different system rankings. We find little correlation between a number of "reasonable" evaluation metrics, which suggests that system effectiveness depends on how you measure it---an undesirable state in IR evaluation. However, we argue that existing evaluation metrics can be generalized into a framework that uses the same underlying contingency table, but places different weights and penalties. Although we stop short of proposing the "one true metric", this framework can guide the future development of a family of metrics that more accurately models user needs. Luchen Tan, Adam Roegiest, Jimmy Lin, Charles L. A. Clarke |
SIGIR | 2 |
| 2015 | Impact of Surrogate Assessments on High-Recall RetrievalabstractWe are concerned with the effect of using a surrogate assessor to train a passive (i.e., batch) supervised-learning method to rank documents for subsequent review, where the effectiveness of the ranking will be evaluated using a different assessor deemed to be authoritative. Previous studies suggest that surrogate assessments may be a reasonable proxy for authoritative assessments for this task. Nonetheless, concern persists in some application domains---such as electronic discovery---that errors in surrogate training assessments will be amplified by the learning method, materially degrading performance. We demonstrate, through a re-analysis of data used in previous studies, that, with passive supervised-learning methods, using surrogate assessments for training can substantially impair classifier performance, relative to using the same deemed-authoritative assessor for both training and assessment. In particular, using a single surrogate to replace the authoritative assessor for training often yields a ranking that must be traversed much lower to achieve the same level of recall as the ranking that would have resulted had the authoritative assessor been used for training. We also show that steps can be taken to mitigate, and sometimes overcome, the impact of surrogate assessments for training: relevance assessments may be diversified through the use of multiple surrogates; and, a more liberal view of relevance can be adopted by having the surrogate label borderline documents as relevant. By taking these steps, rankings derived from surrogate assessments can match, and sometimes exceed, the performance of the ranking that would have been achieved, had the authority been used for training. Finally, we show that our results still hold when the role of surrogate and authority are interchanged, indicating that the results may simply reflect differing conceptions of relevance between surrogate and authority, as opposed to the authority having special skill or knowledge lacked by the surrogate. Adam Roegiest, Gordon V. Cormack, Charles L. A. Clarke, Maura R. Grossman |
SIGIR | 1 |
| 2014 | The effect of expanding relevance judgements with duplicatesabstractWe examine the effects of expanding a judged set of sentences with their duplicates from a corpus. Including new sentences that are exact duplicates of the previously judged sentences may allow for better estimation of performance metrics and enhance the reusability of a test collection. We perform experiments in context of the Temporal Summarization Track at TREC 2013. We find that adding duplicate sentences to the judged set does not significantly affect relative system performance. However, we do find statistically significant changes in the performance of nearly half the systems that participated in the Track. We recommend adding exact duplicate sentences to the set of relevance judgements in order to obtain a more accurate estimate of system performance. Gaurav Baruah, Adam Roegiest, Mark D. Smucker |
SIGIR | 2 |