EDBT 2026 Demo / reviewers in the wild / expert
Paul Thomas 0001
dblp:84/6088-1
· DBLP profile ↗
61ranked-venue papers in the field
21as first author
21since 2021 · last 2026
0000-0003-2425-3136ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 60 (21 first)Data Mining & Knowledge Discovery · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Eye Tracking Study: Are AI Overviews Changing Search Behavior?
Sara Allawati, Dana McKay, Mark Sanderson, Paul Thomas 0001, Johanne R. Trippas |
SIGIR | 4 |
| 2026 | Evaluation Validity in Information Retrieval
Paul Thomas 0001, Nick Craswell, Mark Sanderson, Seth Spielman, Robert Sim, Ryen W. White |
SIGIR | 1 |
| 2026 | On the Use of LLMs for Relevance LabellingabstractLarge Language Models (LLMs) are increasingly used to replace human judges to assess the relevance of information objects, raising concerns about circularity, bias, and whether simulated preferences can substitute for human judgement. This work presents experiments using multiple LLMs to label passages for relevance. It examines their gullibility—how easily they are misled into labelling irrelevant passages as relevant. It also compares LLMs with human judges in ranking systems, analysing differences in discriminative power and whether some systems benefit under LLM-based evaluation. Results show that LLMs are influenced by the presence of query terms, even with irrelevant or random passages. Moreover, LLM-generated rankings are highly correlated with those of human judges, with strong agreement on which system is better in pairwise comparisons. However, LLMs may exhibit lower discriminative power, as seen in flatter ranking slopes and missed significance for meaningful improvements. Yet, there are no cases where capable LLMs and human judges reach opposing conclusions with significance. LLMs may boost traditional systems more than neural ones, adding a new concern of system bias. These findings highlight the strong potential of LLMs for relevance labelling, while also highlighting failure cases that call for careful adoption and further research to maintain evaluation integrity. 1 Marwah Alaofi, Paul Thomas 0001, Falk Scholer, Mark Sanderson |
ACM Trans. Inf. Syst. | 2 |
| 2025 | LLM4Eval: Large Language Model for Evaluation in IRabstractLarge language models (LLMs) have demonstrated increasing task-solving abilities not present in smaller models. Utilizing the capabilities and responsibilities of LLMs for automated evaluation (LLM4Eval) has recently attracted considerable attention in multiple research communities. Building on the success of previous workshops, which established foundations in automated judgments and RAG evaluation, this third iteration aims to address emerging challenges as IR systems become increasingly personalized and interactive. The main goal of the third LLM4Eval workshop is to bring together researchers from industry and academia to explore three critical areas: the evaluation of personalized IR systems while maintaining fairness, the boundaries between automated and human assessment in subjective scenarios, and evaluation methodologies for systems that combine multiple IR paradigms (search, recommendations, and dialogue). By examining these challenges, we seek to understand how evaluation approaches can evolve to match the sophistication of modern IR applications. The format of the workshop is interactive, including roundtable discussion sessions, fostering dialogue about the future of IR evaluation while avoiding one-sided discussions. This is the third iteration of the workshop series, following successful events at SIGIR 2024 and WSDM 2025, with the first iteration attracting over 50 participants. Clemencia Siro, Hossein A. Rahmani, Mohammad Aliannejadi, Nick Craswell, Charles L. A. Clarke, Guglielmo Faggioli, Bhaskar Mitra 0001, Paul Thomas 0001, Emine Yilmaz |
SIGIR | 8 |
| 2025 | System Comparison Using Automated Generation of Relevance Judgements in Multiple LanguagesabstractRecent work has shown that Large Language Models (LLMs) can produce relevance judgements for English retrieval that are useful as a basis for system comparison, and they do so at vastly reduced cost compared to human assessors. Using relevance judgements and ranked retrieval runs from the TREC NeuCLIR track, this paper shows that LLMs can also produce reliable assessments in other languages, even when the topic description or the prompt are in a language different from the documents. Results with Chinese, Persian and Russian documents show that although document language affects both agreement with human assessors on graded relevance and on preference ordering among systems, prompt-language and topic-language effects are negligible. This has implications for the design of multilingual test collections, suggesting that prompts and topic descriptions can be developed in any convenient language. Paul Thomas 0001, Douglas W. Oard, Eugene Yang 0001, Dawn J. Lawrie, James Mayfield |
SIGIR | 1 |
| 2025 | LLM4Eval@WSDM 2025: Large Language Model for Evaluation in Information RetrievalabstractLarge language models (LLMs) have demonstrated increasing task-solving abilities not present in smaller models. Utilizing the capabilities and responsibilities of LLMs for automated evaluation (LLM4Eval) has recently attracted considerable attention in multiple research communities. For instance, LLM4Eval models have been studied in the context of automated judgments, natural language generation, and retrieval augmented generation systems. We believe that the information retrieval community can significantly contribute to this growing research area by designing, implementing, analyzing, and evaluating various aspects of LLMs with applications to LLM4Eval tasks. The main goal of LLM4Eval workshop is to bring together researchers from industry and academia to discuss various aspects of LLMs for evaluation in information retrieval, including automated judgments, retrieval-augmented generation pipeline evaluation, altering human evaluation, robustness, and trustworthiness of LLMs for evaluation in addition to their impact on real-world applications. We also plan to run an automated judgment challenge prior to the workshop, where participants will be asked to generate labels for a given dataset while maximising correlation with human judgments. The format of the workshop is interactive, including roundtable and keynote sessions and tends to avoid the one-sided dialogue of a mini-conference. This is the second iteration of the workshop. The first version was held in conjunction with SIGIR 2024, attracting over 50 participants. Hossein A. Rahmani, Clemencia Siro, Mohammad Aliannejadi, Nick Craswell, Charles L. A. Clarke, Guglielmo Faggioli, Bhaskar Mitra 0001, Paul Thomas 0001, Emine Yilmaz |
WSDM | 8 |
| 2024 | Enhancing Human Annotation: Leveraging Large Language Models and Efficient Batch ProcessingabstractLarge language models (LLMs) are capable of assessing document and query characteristics, including relevance, and are now being used for a variety of different classification labeling tasks as well. This study explores how to use LLMs to classify an information need, often represented as a user query. In particular, our goal is to classify the cognitive complexity of the search task for a given “backstory”. Using 180 TREC topics and backstories, we show that GPT-based LLMs agree with human experts as much as other human experts. We also show that batching and ordering can significantly impact the accuracy of GPT-3.5, but rarely alter the quality of GPT-4 predictions. This study provides insights into the efficacy of large language models for annotation tasks normally completed by humans, and offers recommendations for other similar applications. Oleg Zendel, J. Shane Culpepper, Falk Scholer, Paul Thomas 0001 |
CHIIR | 4 |
| 2024 | What Matters in a Measure? A Perspective from Large-Scale Search EvaluationabstractInformation retrieval (IR) has a large literature on evaluation, dating back decades and forming a central part of the research culture. The largest proportion of this literature discusses techniques to turn a sequence of relevance labels into a single number, reflecting the system's performance: precision or cumulative gain, for example, or dozens of alternatives. Those techniques-metrics-are themselves evaluated, commonly by reference to sensitivity and validity. Paul Thomas 0001, Gabriella Kazai, Nick Craswell, Seth Spielman |
SIGIR | 1 |
| 2024 | Large Language Models can Accurately Predict Searcher PreferencesabstractMuch of the evaluation and tuning of a search system relies on relevance labels---annotations that say whether a document is useful for a given search and searcher. Ideally these come from real searchers, but it is hard to collect this data at scale, so typical experiments rely on third-party labellers who may or may not produce accurate annotations. Label quality is managed with ongoing auditing, training, and monitoring. We discuss an alternative approach. We take careful feedback from real searchers and use this to select a large language model (LLM), and prompt, that agrees with this feedback; the LLM can then produce labels at scale. Our experiments show LLMs are as accurate as human labellers and as useful for finding the best systems and hardest queries. LLM performance varies with prompt features, but also varies unpredictably with simple paraphrases. This unpredictability reinforces the need for high-quality "gold" labels. Paul Thomas 0001, Seth Spielman, Nick Craswell, Bhaskar Mitra 0001 |
SIGIR | 1 |
| 2024 | LLM4Eval: Large Language Model for Evaluation in IRabstractLarge language models (LLMs) have demonstrated increasing task-solving abilities not present in smaller models. Utilizing the capabilities and responsibilities of LLMs for automated evaluation (LLM4Eval) has recently attracted considerable attention in multiple research communities. For instance, LLM4Eval models have been studied in the context of automated judgments, natural language generation, and retrieval augmented generation systems. We believe that the information retrieval community can significantly contribute to this growing research area by designing, implementing, analyzing, and evaluating various aspects of LLMs with applications to LLM4Eval tasks. The main goal of LLM4Eval workshop is to bring together researchers from industry and academia to discuss various aspects of LLMs for evaluation in information retrieval, including automated judgments, retrieval-augmented generation pipeline evaluation, altering human evaluation, robustness, and trustworthiness of LLMs for evaluation in addition to their impact on real-world applications. We also plan to run an automated judgment challenge prior to the workshop, where participants will be asked to generate labels for a given dataset while maximising correlation with human judgments. The format of the workshop is interactive, including roundtable and keynote sessions and tends to avoid the one-sided dialogue of a mini-conference. Hossein A. Rahmani, Clemencia Siro, Mohammad Aliannejadi, Nick Craswell, Charles L. A. Clarke, Guglielmo Faggioli, Bhaskar Mitra 0001, Paul Thomas 0001, Emine Yilmaz |
SIGIR | 8 |
| 2023 | Taking Search to TaskabstractThe importance of tasks in information retrieval (IR) has been long argued for, addressed in different ways, often ignored, and frequently revisited. For decades, scholars made a case for the role that a user’s task plays in how and why that user engages in search and what a search system should do to assist. But for the most part, the IR community has been too focused on query processing and assuming a search task to be a collection of user queries, often ignoring if or how such an assumption addresses the users accomplishing their tasks. With emerging areas of conversational agents and proactive IR, understanding and addressing users’ tasks has become more important than ever before. In this paper, we provide various perspectives on where the state-of-the-art is with regard to tasks in IR, what are some of the bottlenecks in deriving and using task information, and how do we go forward from here. In addition to covering relevant literature, the paper provides a synthesis of historical and current perspectives on understanding, extracting, and addressing task-focused search. To ground ongoing and future research in this area, we present a new framing device for tasks using a tree-like structure and various moves on that structure that allow different interpretations and applications. Presented as a combination of synthesis of ideas and past works, proposals for future research, and our perspectives on technical, social, and ethical considerations, this paper is meant to help revitalize the interest and future work in task-based IR. Chirag Shah 0001, Ryen W. White, Paul Thomas 0001, Bhaskar Mitra 0001, Shawon Sarkar, Nicholas J. Belkin |
CHIIR | 3 |
| 2023 | Can Generative LLMs Create Query Variants for Test Collections? An Exploratory StudyabstractThis paper explores the utility of a Large Language Model (LLM) to automatically generate queries and query variants from a description of an information need. Given a set of information needs described as backstories, we explore how similar the queries generated by the LLM are to those generated by humans. We quantify the similarity using different metrics and examine how the use of each set would contribute to document pooling when building test collections. Our results show potential in using LLMs to generate query variants. While they may not fully capture the wide variety of human-generated variants, they generate similar sets of relevant documents, reaching up to 71.1% overlap at a pool depth of 100. Marwah Alaofi, Luke Gallagher, Mark Sanderson, Falk Scholer, Paul Thomas 0001 |
SIGIR | 5 |
| 2022 | The Crowd is Made of People: Observations from Large-Scale Crowd LabellingabstractLike many other researchers, at Microsoft Bing we use external “crowd” judges to label results from a search engine—especially, although not exclusively, to obtain relevance labels for offline evaluation in the Cranfield tradition. Crowdsourced labels are relatively cheap, and hence very popular, but are prone to disagreements, spam, and various biases which appear to be unexplained “noise” or “error”. In this paper, we provide examples of problems we have encountered running crowd labelling at large scale and around the globe, for search evaluation in particular. We demonstrate effects due to the time of day and day of week that a label is given; fatigue; anchoring; exposure; left-side bias; task switching; and simple disagreement between judges. Rather than simple “error”, these effects are consistent with well-known physiological and cognitive factors. “The crowd” is not some abstract machinery, but is made of people. Human factors that affect people’s judgement behaviour must be considered when designing research evaluations and in interpreting evaluation metrics. Paul Thomas 0001, Gabriella Kazai, Ryen W. White, Nick Craswell |
CHIIR | 1 |
| 2022 | Third Workshop on Building towards Information Interaction and Retrieval Resources Re-use (BIIRRR 2022)abstractWork in Progress Share on Third Workshop on Building towards Information Interaction and Retrieval Resources Re-use (BIIRRR 2022) Authors: Toine Bogers Aalborg University, Denmark Aalborg University, DenmarkView Profile , Maria Gäde Humboldt-Universität zu Berlin, Germany Humboldt-Universität zu Berlin, GermanyView Profile , Mark Michael Hall The Open University, United Kingdom The Open University, United KingdomView Profile , Marijn Koolen Huygens Institute for the History of the Netherlands, Royal Netherlands Academy of Arts and Sciences, Netherlands Huygens Institute for the History of the Netherlands, Royal Netherlands Academy of Arts and Sciences, NetherlandsView Profile , Vivien Petras Humboldt-Universität zu Berlin, Germany Humboldt-Universität zu Berlin, GermanyView Profile , Paul Thomas Microsoft, Australia Microsoft, AustraliaView Profile Authors Info & Claims CHIIR '22: ACM SIGIR Conference on Human Information Interaction and RetrievalMarch 2022 Pages 374–376https://doi.org/10.1145/3498366.3505838Published:14 March 2022Publication History 0citation24DownloadsMetricsTotal Citations0Total Downloads24Last 12 Months24Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Toine Bogers, Maria Gäde, Mark M. Hall, Marijn Koolen, Vivien Petras, Paul Thomas 0001 |
CHIIR | 6 |
| 2022 | Search Interfaces for Biomedical Searching: How do Gaze, User Perception, Search Behaviour and Search Performance Relate?abstractThe objective of this controlled information retrieval (IR) user experiment is to gain an understanding of domain experts’ interactions with novel search interfaces within the context of biomedical information search, with a goal of better search interface design. In this paper, we examine the relationships among user perception, gaze and search behaviour and user search performance. An eye-tracking study of biomedical domain experts’ interactions with novel search interfaces was conducted. A total of thirty-two users participated and searched for documents answering eight complex exploratory search tasks, using four different search interfaces. The findings suggest that gaze behaviour in terms of fixation durations based measures of areas of interest (AOI), i.e., visual attention to the elements of title, author, abstract and MeSH (Medical Subject Headings) terms in document surrogates is correlated with search performance. Users are more likely to achieve better search performance by precision-based measures when 1) search tasks are perceived as difficult; 2) users attend to the element of abstract; and 3) users can recall using the per-query suggestions during the search processes. More importantly, our findings suggest that a user search interface design that displays contextual information between the suggested keywords and the document may better support users reformulating their queries for complex search tasks in the biomedical domain. We discuss implications for the design of search user interfaces for biomedical searching. Ying-Hsang Liu, Paul Thomas 0001, Tom Gedeon, Nicolay Rusnachenko |
CHIIR | 2 |
| 2022 | A Flexible Framework for Offline Effectiveness MetricsabstractThe use of offline effectiveness metrics is one of the cornerstones of evaluation in information retrieval. Static resources that include test collections and sets of topics, the corresponding relevance judgments connecting them, and metrics that map document rankings from a retrieval system to numeric scores have been used for multiple decades as an important way of comparing systems. The basis behind this experimental structure is that the metric score for a system can serve as a surrogate measurement for user satisfaction. Alistair Moffat, Joel Mackenzie, Paul Thomas 0001, Leif Azzopardi |
SIGIR | 3 |
| 2021 | User Models, Metrics and Measures of Search: A Tutorial on the C/W/L Evaluation FrameworkabstractEvaluation is central to Information Retrieval, and is how we compare the quality of systems. One important principle of evaluation is that the measured score should reflect the user's experience with the system. Hence, there should be direct connection between how users interact with the system and the characteristics of the metric. In this tutorial we introduce the C/W/L approach to user modeling and show how different user models lead to different metrics. We then describe the recent innovations and approaches to evaluation that it has facilitated. The tutorial is presented as a mix of on-line synchronous lecture, pre-recorded in-depth videos, and hands-on activities using the C/W/L toolkit for participants' own evaluation tasks. A followup consultation session is also provided, to allow extended questions and individual discussion with the four presenters. Leif Azzopardi, Alistair Moffat, Paul Thomas 0001, Guido Zuccon |
CHIIR | 3 |
| 2021 | Analysing Mixed Initiatives and Search Strategies during Conversational SearchabstractInformation seeking conversations between users and Conversational Search Agents (CSAs) consist of multiple turns of interaction. While users initiate a search session, ideally a CSA should sometimes take the lead in the conversation by obtaining feedback from the user by offering query suggestions or asking for query clarifications i.e. mixed initiative. This creates the potential for more engaging conversational searches, but substantially increases the complexity of modelling and evaluating such scenarios due to the large interaction space coupled with the trade-offs between the costs and benefits of the different interactions. In this paper, we present a model for conversational search -- from which we instantiate different observed conversational search strategies, where the agent elicits: (i) Feedback-First, or (ii) Feedback-After. Using 49 TREC WebTrack Topics, we performed an analysis comparing how well these different strategies combine with different mixed initiative approaches: (i) Query Suggestions vs. (ii) Query Clarifications. Our analysis reveals that there is no superior or dominant combination, instead it shows that query clarifications are better when asked first, while query suggestions are better when asked after presenting results. We also show that the best strategy and approach depends on the trade-offs between the relative costs between querying and giving feedback, the performance of the initial query, the number of assessments per query, and the total amount of gain required. While this work highlights the complexities and challenges involved in analyzing CSAs, it provides the foundations for evaluating conversational strategies and conversational search agents in batch/offline settings. Mohammad Aliannejadi, Leif Azzopardi, Hamed Zamani, Evangelos Kanoulas, Paul Thomas 0001, Nick Craswell |
CIKM | 5 |
| 2021 | Sim4IR: The SIGIR 2021 Workshop on Simulation for Information Retrieval EvaluationabstractThe use of simulation techniques is not foreign to information retrieval. In the past, simulation has been employed, for example, for constructing test collections and for model performance prediction and analysis in a broad array of information access scenarios. Nevertheless, a standardized methodology for performance evaluation via simulation has not yet been developed. The goal of this workshop is to create a forum for researchers and practitioners to promote methodology development and more widespread use of simulation for evaluation by: (1) identifying problem settings and application scenarios; (2) sharing tools, techniques, and experiences; (3) characterizing potentials and limitations; and (4) developing a research agenda. Krisztian Balog, David Maxwell 0001, Paul Thomas 0001, Shuo Zhang 0006 |
SIGIR | 3 |
| 2021 | Do Affective Cues Validate Behavioural Metrics for Search?abstractTraces of searcher behaviour, such as query reformulation or clicks, are commonly used to evaluate a running search engine. The underlying expectation is that these behaviours are proxies for something more important, such as relevance, utility, or satisfaction. Affective computing technology gives us the tools to help confirm some of these expectations, by examining visceral expressive responses during search sessions. However, work to date has only studied small populations in laboratory settings and with a limited number of contrived search tasks. In this study, we analysed longitudinal, in-situ, search behaviours of 152 information workers, over the course of several weeks while simultaneously tracking their facial expressions. Results from over 20,000 search sessions and 45,000 queries allow us to observe that indeed affective expressions are consistent with, and complementary to, existing "click-based'' metrics. On a query-level, searches that result in a short dwell time are associated with a decrease in smiles (expressions of "happiness'') and that if a query is reformulated the results of the reformulation are associated with an increase in smiling---suggesting a positive outcome as people converge on the information they need. On a session-level, sessions that feature reformulations are more commonly associated with fewer smiles and more furrowed brows (expressions of "anger/frustration''). Similarly, sessions with short-dwell clicks are also associated with fewer smiles. These data provide an insight into visceral aspects of search experience and present a new dimension for evaluating engine performance. Daniel McDuff, Paul Thomas 0001, Nick Craswell, Kael Rowan, Mary Czerwinski |
SIGIR | 2 |
| 2021 | Theories of Conversation for Conversational IRabstractConversational information retrieval is a relatively new and fast-developing research area, but conversation itself has been well studied for decades. Researchers have analysed linguistic phenomena such as structure and semantics but also paralinguistic features such as tone, body language, and even the physiological states of interlocutors. We tend to treat computers as social agents—especially if they have some humanlike features in their design—and so work from human-to-human conversation is highly relevant to how we think about the design of human-to-computer applications. In this article, we summarise some salient past work, focusing on social norms; structures; and affect, prosody, and style. We examine social communication theories briefly as a review to see what we have learned about how humans interact with each other and how that might pertain to agents and robots. We also discuss some implications for research and design of conversational IR systems. Paul Thomas 0001, Mary Czerwinski, Daniel McDuff, Nick Craswell |
ACM Trans. Inf. Syst. | 1 |
| 2020 | Data-Driven Evaluation Metrics for Heterogeneous Search Engine Result PagesabstractEvaluation metrics for search typically assume items are homoge- neous. However, in the context of web search, this assumption does not hold. Modern search engine result pages (SERPs) are composed of a variety of item types (e.g., news, web, entity, etc.), and their influence on browsing behavior is largely unknown. Leif Azzopardi, Ryen W. White, Paul Thomas 0001, Nick Craswell |
CHIIR | 3 |
| 2020 | Third International Workshop on Conversational Approaches to Information Retrieval (CAIR'20): Full-day Workshop at CHIIR 2020abstractThe third CAIR workshop brings together researchers and developers interested in advancing conversational systems in interactive information retrieval. The workshop builds on the first and second CAIR workshops held at SIGIR 2017 and 2018 and will focus on the continuing development of current challenges, user and system limitations, and evaluation of conversational systems for information retrieval. Participants will collaboratively explore different contexts (i.e., home, hospitals, or work settings), use cases, and interactivity forms (voice-only, multi-modal, screen-based) in which conversational search systems can be used. Possible outcomes include fostering novel and innovative methodologies (such as for data collection and evaluation), personalising conversational systems, and understanding ethical challenges---such as system transparency---from the user's perspective. Johanne R. Trippas, Paul Thomas 0001, Damiano Spina, Hideo Joho |
CHIIR | 2 |
| 2020 | Expressions of Style in Information Seeking Conversation with an AgentabstractPast work in information-seeking conversation has demonstrated that people exhibit different conversational styles---for example, in word choice or prosody---that differences in style lead to poorer conversations, and that partners actively align their styles over time. One might assume that this would also be true for conversations with an artificial agent such as Cortana, Siri, or Alexa; and that agents should therefore track and mimic a user's style. We examine this hypothesis with reference to a lab study, where 24 participants carried out relatively long information-seeking tasks with an embodied conversational agent. The agent combined topical language models with a conversational dialogue engine, style recognition and alignment modules. We see that "style'' can be measured in human-to-agent conversation, although it looks somewhat different to style in human-to-human conversation and does not correlate with self-reported preferences. There is evidence that people align their style to the agent, and that conversations run more smoothly if the agent detects, and aligns to, the human's style as well. Paul Thomas 0001, Daniel McDuff, Mary Czerwinski, Nick Craswell |
SIGIR | 1 |
| 2020 | Towards a model for spoken conversational search
Johanne R. Trippas, Damiano Spina, Paul Thomas 0001, Mark Sanderson, Hideo Joho, Lawrence Cavedon |
Inf. Process. Manag. | 3 |
| 2020 | Investigating Searchers' Mental Models to Inform Search ExplanationsabstractModern web search engines use many signals to select and rank results in response to queries. However, searchers’ mental models of search are relatively unsophisticated, hindering their ability to use search engines efficiently and effectively. Annotating results with more in-depth explanations could help, but search engine providers need to know what to explain. To this end, we report on a study of searchers’ mental models of web selection and ranking, with more than 400 respondents to an online survey and 11 face-to-face interviews. Participants volunteered a range of factors and showed good understanding of important concepts such as popularity, wording, and personalization. However, they showed little understanding of recency or diversity and incorrect ideas of payment for ranking. Where there are already explanatory annotations on the results page—such as “ad” markers and keyword highlighting—participants were familiar with ranking concepts. This suggests that further explanatory annotations may be useful. Paul Thomas 0001, Bodo Billerbeck, Nick Craswell, Ryen W. White |
ACM Trans. Inf. Syst. | 1 |
| 2019 | Building Economic Models and Measures of SearchabstractEconomics provides an intuitive and natural way to formally represent the costs and benefits of interacting with applications, interfaces and devices. By using economic models it is possible to reason about interaction, make predictions about how changes to the system will affect behavior, and measure the performance of people's interactions with the system. In this tutorial, we first provide an overview of relevant economic theories, before showing how they can be applied to formulate different ranking principles to provide the optimal ranking to users. This is followed by a session showing how economics can be used to model how people interact with search systems, and how to use these models to generate hypotheses about user behavior. The third session focuses on how economics has been used to underpin the measurement of information retrieval systems and applications using the CWL framework (which reports the expected utility, expected total utility, expected total cost, and so on) -- and how different models of user interaction lead to different metrics. We then show how information foraging theory can be used to measure the performance of an information retrieval system -- connecting the theory of how people search with how we measure it. The final session of the day will be spent building economic models and measures of search. Here sample problems will be provided to challenge participants, or participants can bring their own. Leif Azzopardi, Alistair Moffat, Paul Thomas 0001, Guido Zuccon |
SIGIR | 3 |
| 2019 | cwl_eval: An Evaluation Tool for Information RetrievalabstractWe present a tool ("cwl_eval") which unifies many metrics typically used to evaluate information retrieval systems using test collections. In the CWL framework metrics are specified via a single function which can be used to derive a number of related measurements: Expected Utility per item, Expected Total Utility, Expected Cost per item, Expected Total Cost, and Expected Depth. The CWL framework brings together several independent approaches for measuring the quality of a ranked list, and provides a coherent user model-based framework for developing measures based on utility (gain) and cost. Here we outline the CWL measurement framework; describe the cwl_eval architecture; and provide examples of how to use it. We provide implementations of a number of recent metrics, including Time Biased Gain, U-Measure, Bejewelled Measure, and the Information Foraging Based Measure, as well as previous metrics such as Precision, Average Precision, Discounted Cumulative Gain, Rank-Biased Precision, and INST. By providing state-of-the-art and traditional metrics within the same framework, we promote a standardised approach to evaluating search effectiveness. Leif Azzopardi, Paul Thomas 0001, Alistair Moffat |
SIGIR | 2 |
| 2019 | The Emotion Profile of Web SearchabstractEmotions are an essential part of most human activities, including decision-making. Emotions arise in response to information, e.g., presented in web pages, and are also expressed in the words used to convey that information in the first place. In this paper, we study the emotion profile of retrieved and clicked web search results towards the goal of better understanding the role of emotions in web search. Using click logs from a four-month period, up to the end of January 2019, we examine the emotions associated with search results and contrast them to the emotions of clicked results, taking rank and relevance into account. Emotions are assigned to web pages based on two lexicons: SentiWordNet (positive, negative and objective sentiments) and EmoLexData (afraid, amused, angry, annoyed, don't care, happy, inspired, and sad emotions). We look at the sentiment/emotion profiles of search results grouped around a set of controversial and mundane topics and hypothesise that users are more likely to click emotionally charged results than emotionless results, both in general, and in particular when their query relates to controversial topics. Gabriella Kazai, Paul Thomas 0001, Nick Craswell |
SIGIR | 2 |
| 2018 | Style and Alignment in Information-Seeking ConversationabstractAnalysis of casual chit-chat indicates that differences in conversational style---the way things are said---can significantly impact a participants» impressions of the conversation and of each other. However, prior work has not systematically analyzed how important style is in task-oriented, information-seeking exchanges of the sort we might have with a conversational search agent. We examine recordings from the MISC data set, where pairs of "users" and "intermediaries" collaborate on information-seeking tasks, and look for indications of style which can be computed at scale. We find that stylistic markers identified by Tannen in casual chat do exist in information-seeking dialogue, and that participants can be arranged along a single stylistic dimension: "considerate" to "involved". This labelling for style needs no manual intervention. Furthermore, we find that there is no clear best style; but that differences in style, previously thought to impede communication, are only a problem for shorter tasks. This result is likely due to alignment of conversational style over the course of an interaction. Paul Thomas 0001, Mary Czerwinski, Daniel McDuff, Nick Craswell, Gloria Mark |
CHIIR | 1 |
| 2018 | Measuring the Utility of Search Engine Result Pages: An Information Foraging Based MeasureabstractWeb Search Engine Result Pages (SERPs) are complex responses to queries, containing many heterogeneous result elements (web results, advertisements, and specialised "answers'') positioned in a variety of layouts. This poses numerous challenges when trying to measure the quality of a SERP because standard measures were designed for homogeneous ranked lists. In this paper, we aim to measure the utility and cost of SERPs. To ground this work we adopt the \CWL framework which enables a direct comparison between different measures in the same units of measurement, i.e. expected (total) utility and cost. Within this framework, we propose a new measure based on information foraging theory, which can account for the heterogeneity of elements, through different costs, and which naturally motivates the development of a user stopping model that adapts behaviour depending on the rate of gain. This directly connects models of how people search with how we measure search, providing a number of new dimensions in which to investigate and evaluate user behaviour and performance. We perform an analysis over~1000 popular queries issued to a major search engine, and report the aggregate utility experienced by users over time. Then in an comparison against common measures, we show that the proposed foraging based measure provides a more accurate reflection of the utility and of observed behaviours (stopping rank and time spent). Leif Azzopardi, Paul Thomas 0001, Nick Craswell |
SIGIR | 2 |
| 2017 | What Snippet Size is Needed in Mobile Web Search?abstractA snippet (content summary for a web page) is one of the main elements in a search result page. Search engines have been improved to reduce users' effort in web search, e.g., providing flexible snippet sizes by considering the purpose of the search and suggesting predicted answers. In most cases, search engines for mobile devices present two or three lines of snippet for each result link. Some studies suggest that long snippets provide a better search experience on desktop screens, but this may not be true for mobile devices because of the smaller screen. Paul Thomas 0001, Ramesh S. Sankaranarayana, Tom Gedeon, Hwan-Jin Yoon |
CHIIR | 2 |
| 2017 | Retrieval Consistency in the Presence of Query VariationsabstractA search engine that can return the ideal results for a person's information need, independent of the specific query that is used to express that need, would be preferable to one that is overly swayed by the individual terms used; search engines should be consistent in the presence of syntactic query variations responding to the same information need. In this paper we examine the retrieval consistency of a set of five systems responding to syntactic query variations over one hundred topics, working with the UQV100 test collection, and using Rank-Biased Overlap (RBO) relative to a centroid ranking over the query variations per topic as a measure of consistency. We also introduce a new data fusion algorithm, Rank-Biased Centroid (RBC), for constructing a centroid ranking over a set of rankings from query variations for a topic. RBC is compared with alternative data fusion algorithms. Peter Bailey, Alistair Moffat, Falk Scholer, Paul Thomas 0001 |
SIGIR | 4 |
| 2017 | Incorporating User Expectations and Behavior into the Measurement of Search EffectivenessabstractInformation retrieval systems aim to help users satisfy information needs. We argue that the goal of the person using the system, and the pattern of behavior that they exhibit as they proceed to attain that goal, should be incorporated into the methods and techniques used to evaluate the effectiveness of IR systems, so that the resulting effectiveness scores have a useful interpretation that corresponds to the users’ search experience. In particular, we investigate the role of search task complexity, and show that it has a direct bearing on the number of relevant answer documents sought by users in response to an information need, suggesting that useful effectiveness metrics must be goal sensitive . We further suggest that user behavior while scanning results listings is affected by the rate at which their goal is being realized, and hence that appropriate effectiveness metrics must be adaptive to the presence (or not) of relevant documents in the ranking. In response to these two observations, we present a new effectiveness metric, INST, that has both of the desired properties: INST employs a parameter T , a direct measure of the user’s search goal that adjusts the top-weightedness of the evaluation score; moreover, as progress towards the target T is made, the modeled user behavior is adapted, to reflect the remaining expectations. INST is experimentally compared to previous effectiveness metrics, including Average Precision (AP), Normalized Discounted Cumulative Gain (NDCG), and Rank-Biased Precision (RBP), demonstrating our claims as to INST’s usefulness. Like RBP, INST is a weighted-precision metric, meaning that each score can be accompanied by a residual that quantifies the extent of the score uncertainty caused by unjudged documents. As part of our experimentation, we use crowd-sourced data and score residuals to demonstrate that a wide range of queries arise for even quite specific information needs, and that these variant queries introduce significant levels of residual uncertainty into typical experimental evaluations. These causes of variability have wide-reaching implications for experiment design, and for the construction of test collections. Alistair Moffat, Peter Bailey, Falk Scholer, Paul Thomas 0001 |
ACM Trans. Inf. Syst. | 4 |
| 2016 | System And User Centered Evaluation Approaches in Interactive Information Retrieval (SAUCE 2016)abstractThe purpose of this half-day workshop is to bring together academic and industry interactive information retrieval (IIR) researchers with an interest in evaluation methodologies. The workshop articulates contemporary challenges in the investigation of IIR and invites user- and system-oriented researchers to work collaboratively to address these challenges by combining user- and system-centered methodologies in meaningful ways. We anticipate that this workshop will initiate productive knowledge exchange and partnerships that can respond to the increasing user, task, system, and contextual complexity of the IIR field. Heather L. O'Brien, Nicola Ferro 0001, Hideo Joho, Dirk Lewandowski, Paul Thomas 0001, C. J. van Rijsbergen |
CHIIR | 5 |
| 2016 | Measuring Engagement with Online FormsabstractOnline form-filling and transactions are extremely common, both for industry and government; and it is important to provide a satisfying user experience during these tasks if customers or citizens are to continue using online channels. However, reliable measures of experience in these cases are limited. Other areas of information interaction, e.g., online search, news, and shopping, are increasingly exploring and attempting to measure the concept of user engagement (UE). In this study, we ask whether UE is an appropriate outcome for the utilitarian activities of online form-filling and transactions. Paul Thomas 0001, Heather L. O'Brien, Tom Rowlands |
CHIIR | 1 |
| 2016 | Pagination versus Scrolling in Mobile Web SearchabstractVertical scrolling is the standard method of exploring search results pages. For touch-enabled mobile devices that are not equipped with a mouse or keyboard, we adopt other methods of controlling the viewport with the aim of investigating user interaction. From the intuition that people are used to reading books by turning pages horizontally, we conducted a user experiment to investigate the effects of horizontal and vertical control types (pagination versus scrolling) on a touch-enabled mobile phone. Our findings suggest that participants using pagination were more likely to find relevant documents, especially those over the fold; spent more time attending to relevant results; and were faster to click while spending less time on the search result pages overall. We also found that the main reason for the difference in search speed is the time taken for the scroll itself. We conclude that search engines need to provide different viewport controls to allow better search experiences on touch-enabled mobile devices. Paul Thomas 0001, Ramesh S. Sankaranarayana, Tom Gedeon, Hwan-Jin Yoon |
CIKM | 2 |
| 2016 | UQV100: A Test Collection with Query VariabilityabstractWe describe the UQV100 test collection, designed to incorporate variability from users. Information need ?backstories? were written for 100 topics (or sub-topics) from the TREC 2013 and 2014 Web Tracks. Crowd workers were asked to read the backstories, and provide the queries they would use; plus effort estimates of how many useful documents they would have to read to satisfy the need. A total of 10,835 queries were collected from 263 workers. After normalization and spell-correction, 5,764 unique variations remained; these were then used to construct a document pool via Indri-BM25 over the ClueWeb12-B corpus. Qualified crowd workers made relevance judgments relative to the backstories, using a relevance scale similar to the original TREC approach; first to a pool depth of ten per query, then deeper on a set of targeted documents. The backstories, query variations, normalized and spell-corrected queries, effort estimates, run outputs, and relevance judgments are made available collectively as the UQV100 test collection. We also make available the judging guidelines and the gold hits we used for crowd-worker qualification and spam detection. We believe this test collection will unlock new opportunities for novel investigations and analysis, including for problems such as task-intent retrieval performance and consistency (independent of query variation), query clustering, query difficulty prediction, and relevance feedback, among others. Peter Bailey, Alistair Moffat, Falk Scholer, Paul Thomas 0001 |
SIGIR | 4 |
| 2016 | Understanding eye movements on mobile devices for better presentation of search resultsabstractCompared to the early versions of smart phones, recent mobile devices have bigger screens that can present more web search results. Several previous studies have reported differences in user interaction between conventional desktop computer and mobile device‐based web searches, so it is imperative to consider the differences in user behavior for web search engine interface design on mobile devices. However, it is still unknown how the diversification of screen sizes on hand‐held devices affects how users search. In this article, we investigate search performance and behavior on three different small screen sizes: early smart phones, recent smart phones, and phablets. We found no significant difference with respect to the efficiency of carrying out tasks, however participants exhibited different search behaviors: less eye movement within top links on the larger screen, fast reading with some hesitation before choosing a link on the medium, and frequent use of scrolling on the small screen. This result suggests that the presentation of web search results for each screen needs to take into account differences in search behavior. We suggest several ideas for presentation design for each screen size. Paul Thomas 0001, Ramesh S. Sankaranarayana, Tom Gedeon, Hwan-Jin Yoon |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2015 | Pooled Evaluation Over Query Variations: Users are as Diverse as SystemsabstractEvaluation of information retrieval systems with test collections makes use of a suite of fixed resources: a document corpus; a set of topics; and associated judgments of the relevance of each document to each topic. With large modern collections, exhaustive judging is not feasible. Therefore an approach called pooling is typically used where, for example, the documents to be judged can be determined by taking the union of all documents returned in the top positions of the answer lists returned by a range of systems. Conventionally, pooling uses system variations to provide diverse documents to be judged for a topic; different user queries are not considered. We explore the ramifications of user query variability on pooling, and demonstrate that conventional test collections do not cover this source of variation. The effect of user query variation on the size of the judging pool is just as strong as the effect of retrieval system variation. We conclude that user query variation should be incorporated early in test collection construction, and cannot be considered effectively post hoc. Alistair Moffat, Falk Scholer, Paul Thomas 0001, Peter Bailey |
CIKM | 3 |
| 2015 | User Variability and IR System EvaluationabstractTest collection design eliminates sources of user variability to make statistical comparisons among information retrieval (IR) systems more affordable. Does this choice unnecessarily limit generalizability of the outcomes to real usage scenarios? We explore two aspects of user variability with regard to evaluating the relative performance of IR systems, assessing effectiveness in the context of a subset of topics from three TREC collections, with the embodied information needs categorized against three levels of increasing task complexity. First, we explore the impact of widely differing queries that searchers construct for the same information need description. By executing those queries, we demonstrate that query formulation is critical to query effectiveness. The results also show that the range of scores characterizing effectiveness for a single system arising from these queries is comparable or greater than the range of scores arising from variation among systems using only a single query per topic. Second, our experiments reveal that searchers display substantial individual variation in the numbers of documents and queries they anticipate needing to issue, and there are underlying significant differences in these numbers in line with increasing task complexity levels. Our conclusion is that test collection design would be improved by the use of multiple query variations per topic, and could be further improved by the use of metrics which are sensitive to the expected numbers of useful documents. Peter Bailey, Alistair Moffat, Falk Scholer, Paul Thomas 0001 |
SIGIR | 4 |
| 2015 | Features of Disagreement Between Retrieval Effectiveness MeasuresabstractMany IR effectiveness measures are motivated from intuition, theory, or user studies. In general, most effectiveness measures are well correlated with each other. But, what about where they don't correlate? Which rankings cause measures to disagree? Are these rankings predictable for particular pairs of measures? In this work, we examine how and where metrics disagree, and identify differences that should be considered when selecting metrics for use in evaluating retrieval systems. Timothy Jones 0001, Paul Thomas 0001, Falk Scholer, Mark Sanderson |
SIGIR | 2 |
| 2015 | Eye-tracking analysis of user behavior and performance in web search on large and small screensabstractIn recent years, searching the web on mobile devices has become enormously popular. Because mobile devices have relatively small screens and show fewer search results, search behavior with mobile devices may be different from that with desktops or laptops. Therefore, examining these differences may suggest better, more efficient designs for mobile search engines. In this experiment, we use eye tracking to explore user behavior and performance. We analyze web searches with 2 task types on 2 differently sized screens: one for a desktop and the other for a mobile device. In addition, we examine the relationships between search performance and several search behaviors to allow further investigation of the differences engendered by the screens. We found that users have more difficulty extracting information from search results pages on the smaller screens, although they exhibit less eye movement as a result of an infrequent use of the scroll function. However, in terms of search performance, our findings suggest that there is no significant difference between the 2 screens in time spent on search results pages and the accuracy of finding answers. This suggests several possible ideas for the presentation design of search results pages on small devices. Paul Thomas 0001, Ramesh S. Sankaranarayana, Tom Gedeon, Hwan-Jin Yoon |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2014 | Compositional data analysis (CoDA) approaches to distance in information retrievalabstractMany techniques in information retrieval produce counts from a sample, and it is common to analyse these counts as proportions of the whole---term frequencies are a familiar example. Proportions carry only relative information and are not free to vary independently of one another: for the proportion of one term to increase, one or more others must decrease. These constraints are hallmarks of compositional data. While there has long been discussion in other fields of how such data should be analysed, to our knowledge, Compositional Data Analysis (CoDA) has not been considered in IR. Paul Thomas 0001, David R. Lovell |
SIGIR | 1 |
| 2014 | Using Interaction Data to Explain Difficulty Navigating OnlineabstractA user's behaviour when browsing a Web site contains clues to that user's experience. It is possible to record some of these behaviours automatically, and extract signals that indicate a user is having trouble finding information. This allows for Web site analytics based on user experiences, not just page impressions. A series of experiments identified user browsing behaviours—such as time taken and amount of scrolling up a page—which predict navigation difficulty and which can be recorded with minimal or no changes to existing sites or browsers. In turn, patterns of page views correlate with these signals and these patterns can help Web authors understand where and why their sites are hard to navigate. A new software tool, “LATTE,” automates this analysis and makes it available to Web authors in the context of the site itself. Paul Thomas 0001 |
ACM Trans. Web | 1 |
| 2013 | Users versus models: what observation tells us about effectiveness metricsabstractRetrieval system effectiveness can be measured in two quite different ways: by monitoring the behavior of users and gathering data about the ease and accuracy with which they accomplish certain specified information-seeking tasks; or by using numeric effectiveness metrics to score system runs in reference to a set of relevance judgments. In the second approach, the effectiveness metric is chosen in the belief that user task performance, if it were to be measured by the first approach, should be linked to the score provided by the metric. Alistair Moffat, Paul Thomas 0001, Falk Scholer |
CIKM | 2 |
| 2012 | Differences in Language and Style Between Two Social Media Communities
Cécile Paris, Paul Thomas 0001, Stephen Wan 0001 |
ICWSM | 2 |
| 2012 | To what problem is distributed information retrieval the solution?abstractDistributed information retrieval (DIR), where a single broker coordinates retrieval from many independent search services, has been extensively studied but typically without any particular application and sometimes even without any explicit motivation. There have been a handful of arguments given for DIR—coverage, effectiveness, and ease of use, for example—but these are not borne out by experience. Are there uses for DIR? There are, but generally for organizational not technical reasons, and they have not been well studied. Paul Thomas 0001 |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2011 | Relative effect of spam and irrelevant documents on user interaction with search enginesabstractMeaningful evaluation of web search must take account of spam. Here we conduct a user experiment to investigate whether satisfaction with search engine result pages as a whole is harmed more by spam or by irrelevant documents. On some measures, search result pages are differentially harmed by the insertion of spam and irrelevant documents. Additionally we find that when users are given two documents of equal utility, the one with the lower spam score will be preferred; a result page without any spam documents will be preferred to one with spam; and an irrelevant document high in a result list is surprisingly more damaging to user satisfaction than a spam document. We conclude that web ranking and evaluation should consider both utility (relevance) and "spamminess" of documents. Timothy Jones 0001, David Hawking, Paul Thomas 0001, Ramesh S. Sankaranarayana |
CIKM | 3 |
| 2011 | What deliberately degrading search quality tells us about discount functionsabstractDeliberate degradation of search results is a common tool in user experiments. We degrade high-quality search results by inserting non-relevant documents at different ranks. The effect of these manipulations, on a number of commonly-used metrics, is counter-intuitive: the discount functions implicit in P@k, MRR, NDCG, and others do not account for the true relationship between rank and value to the user. We propose an alternative, based on visibility data. Categories and Subject Descriptors: H.3.4 [Information Paul Thomas 0001, Timothy Jones 0001, David Hawking |
SIGIR | 1 |
| 2010 | Evaluating Server Selection for Federated Search
Paul Thomas 0001, Milad Shokouhi |
ECIR | 1 |
| 2010 | Focused and aggregated search: a perspective from natural language generation
Cécile Paris, Stephen Wan 0001, Paul Thomas 0001 |
Inf. Retr. | 3 |
| 2009 | Effective query expansion for federated searchabstractWhile query expansion techniques have been shown to improve retrieval performance in a centralized setting, they have not been well studied in a federated setting. In this paper, we consider how query expansion may be adapted to federated environments and propose several new methods: where focused expansions are used in a selective fashion to produce specific queries for each source (or a set of sources). On a number of different testbeds, we show that focused query expansion can significantly outperform the previously proposed global expansion method, and---contrary to earlier work---show that query expansion can improve performance over standard federated retrieval. \n \n These findings motivate further research examining the different methods for query expansion, and other forms of system and user interaction, in order to continue improving the performance of interactive federated search systems. Milad Shokouhi, Leif Azzopardi, Paul Thomas 0001 |
SIGIR | 3 |
| 2009 | SUSHI: scoring scaled samples for server selectionabstractModern techniques for distributed information retrieval use a set of documents sampled from each server, but these samples have been underutilised in server selection. We describe a new server selection algorithm, SUSHI, which unlike earlier algorithms can make full use of the text of each sampled document and which does not need training data. SUSHI can directly optimise for many common cases, including high precision retrieval, and by including a simple stopping condition can do so while reducing network traffic. Paul Thomas 0001, Milad Shokouhi |
SIGIR | 1 |
| 2009 | Server selection methods in personal metasearch: a comparative empirical study
Paul Thomas 0001, David Hawking |
Inf. Retr. | 1 |
| 2008 | Relevance assessment: are judges exchangeable and does it matterabstractWe investigate to what extent people making relevance judgements for a reusable IR test collection are exchangeable. We consider three classes of judge: "gold standard" judges, who are topic originators and are experts in a particular information seeking task; "silver standard" judges, who are task experts but did not create topics; and "bronze standard" judges, who are those who did not define topics and are not experts in the task. Peter Bailey, Nick Craswell, Ian Soboroff, Paul Thomas 0001, Arjen P. de Vries, Emine Yilmaz |
SIGIR | 4 |
| 2008 | Generalising multiple capture-recapture to non-uniform sample sizesabstractAlgorithms in distributed information retrieval often rely on accurate knowledge of the size of a collection. The "multiple capture-recapture" method of Shokouhi et al. is one of the more reliable algorithms for determining collection size, but it relies on samples with a uniform number of documents. Such uniform samples are often hard to obtain in a working system. Paul Thomas 0001 |
SIGIR | 1 |
| 2007 | Evaluating sampling methods for uncooperative collectionsabstractMany server selection methods suitable for distributed information retrieval applications rely, in the absence of cooperation, on the availability of unbiased samples of documents from the constituent collections. We describe a number of sampling methods which depend only on the normal query-response mechanism of the applicable search facilities. We evaluate these methods on a number of collections typical of a personal metasearch application. Results demonstrate that biases exist for all methods, particularly toward longer documents, and that in some cases these biases can be reduced but not eliminated by choice of parameters.We also introduce a new sampling technique, "multiple queries", which produces samples of similar quality to the best current techniques but with significantly reduced cost. Paul Thomas 0001, David Hawking |
SIGIR | 1 |
| 2007 | Estimating the value of automatic disambiguationabstractA common motivation for personalised search systems is the ability to disambiguate queries based on some knowledge of a user's interests. An analysis of log files from three search providers, covering a range of scenarios, suggests that this sort of disambiguation would be of marginal use for more specialised providers but may be of use for whole-of-Web search. Paul Thomas 0001, Tom Rowlands |
SIGIR | 1 |
| 2006 | Evaluation by comparing result sets in contextabstractFamiliar evaluation methodologies for information retrieval (IR) are not well suited to the task of comparing systems in many real settings. These systems and evaluation methods must support contextual, interactive retrieval over changing, heterogeneous data collections, including private and confidential information.We have implemented a comparison tool which can be inserted into the natural IR process. It provides a familiar search interface, presents a small number of result sets in side-by-side panels, elicits searcher judgments, and logs interaction events. The tool permits study of real information needs as they occur, uses the documents actually available at the time of the search, and records judgments taking into account the instantaneous needs of the searcher.We have validated our proposed evaluation approach and explored potential biases by comparing different whole-of-Web search facilities using a Web-based version of the tool. In four experiments, one with supplied queries in the laboratory and three with real queries in the workplace, subjects showed no discernable left-right bias and were able to reliably distinguish between high- and low-quality result sets. We found that judgments were strongly predicted by simple implicit measures.Following validation we undertook a case study comparing two leading whole-of-Web search engines. The approach is now being used in several ongoing investigations. Paul Thomas 0001, David Hawking |
CIKM | 1 |
| 2005 | Server selection methods in hybrid portal searchabstractThe TREC.GOV collection makes a valuable web testbed for distributed information retrieval methods because it is naturally partitioned and includes 725 web-oriented queries with judged answers. It can usefully model aspects of government and large corporate portals. Analysis of the.gov data shows that a purely distributed approach would not be feasible for providing search on a.gov portal because of the large number (17,000+) of web sites and the high proportion that do not provide a search interface. An alternative hybrid approach, combining both distributed and centralized techniques, is proposed and server selection methods are evaluated within this framework using web-oriented evaluation methodology. A number of well-known algorithms are compared against representatives (highest anchor ranked page (HARP) and anchor weighted sum (AWSUM)) of a family of new selection methods which use link anchortext extracted from an auxiliary crawl to provide descriptions of sites which are not themselves crawled. Of the previously published methods, ReDDE substantially outperformed three variants of CORI and also outperformed a method based on Kullback-Leibler Divergence (extended) except on topic distillation. HARP and AWSUM performed best overall but were outperformed on the topic distillation task by extended KL Divergence. David Hawking, Paul Thomas 0001 |
SIGIR | 2 |