EDBT 2026 Demo / reviewers in the wild / expert
Leif Azzopardi
dblp:29/4018
· DBLP profile ↗
144ranked-venue papers in the field
50as first author
31since 2021 · last 2026
0000-0002-6900-0557ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 144 (50 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Information Farming: From Berry Picking to Berry GrowingabstractThe classic paradigms of Berry Picking and Information Foraging Theory have framed users as gatherers, opportunistically searching across distributed sources to satisfy evolving information needs. However, the rise of Generative AI (GenAI) is driving a fundamental transformation in how people produce, structure, and reuse information—one that these paradigms no longer fully capture. This transformation is analogous to the Neolithic Revolution, when societies shifted from hunting and gathering to cultivation. Generative technologies empower users to “farm” information by planting seeds in the form of prompts, cultivating workflows over time, and harvesting richly structured, relevant yields within their own plots, rather than foraging across others people’s patches. In this perspectives paper, we introduce the notion of Information Farming as a conceptual framework and argue that it represents a natural evolution in how people engage with information. Drawing on historical analogy and empirical evidence, we examine the benefits and opportunities of information farming, its implications for design and evaluation, and the accompanying risks posed by this transition. We hypothesize that as GenAI technologies proliferate, cultivating information will increasingly supplant transient, patch-based foraging as a dominant mode of engagement, marking a broader shift in human-information interaction and its study. Leif Azzopardi, Adam Roegiest |
CHIIR | 1 |
| 2026 | The Third Search Futures Workshop at ECIR'26
Leif Azzopardi, Charles L. A. Clarke, Claudia Hauff, Yubin Kim 0001, Zhaochun Ren, Adam Roegiest, Johanne R. Trippas, Saber Zerhoudi |
ECIR (3) | 1 |
| 2025 | Assessing Risks in Online Information SharingabstractThe volume of personal information, accessible online about individuals is unprecedented.Such information may be pieced together by others, to create a more detailed picture of a person, exposing them to potential harms, such as employment loss, unwanted attention, fraud, and more.In this context, relevance is contextual, situational and dependent, based on the risk it poses to the subject.In this paper, we explore this risk-based notion of relevance with the following questions in mind: How well can individuals identify and judge risks associated with online personal information?And, to what extent does this change individuals' awareness of their own information-sharing practices?In a user study, 243 participants were tasked with browsing fabricated online profiles to identify potential "risky" posts in one of two scenarios regarding either Identity Theft or Reputational Damage.On average, 72.2% of participants identified at least one risky post.However, only 23.7% identified dependent posts that taken together substantially increased the risk of identity theft or reputational damage.Further, participants reported greater awareness of potential risks that could arise from their own, and/or their friends' information sharing practices.Our findings suggest that when relevance is dependent on combining separate pieces of information to reveal risk, participants struggle to identify these cumulative revelations.Moreover, our study highlights that when participants perform tasks that feature personal information, it can lead to positive and negative experiences; changing their perceptions and increasing awareness about their own information behaviours while also raising concerns around their routine online practices. Leif Azzopardi, Emma Nicol, Jo Briggs, Wendy Moncur, Burkhard Schafer 0001, Callum Nash, Melissa Duheric |
CHIIR | 1 |
| 2025 | Search Changes Consumers' Minds: How Recognizing Gaps Drives Sustainable Choices: How Recognizing Gaps Drives Sustainable ChoicesabstractDespite a growing desire among consumers to shop responsibly, translating this intention into behaviour remains challenging.Previous work has identified that information seeking (or lack thereof) is a contributing factor to this intention-behaviour gap.In this paper, we hypothesize that searching can bridge this gap -helping consumers to make purchasing decisions that are better aligned with their values.We conducted a task-based study with 308 participants, asking them to search for information on one of eight ethical aspects regarding a product they were actively shopping for.Our findings show that actively searching for such information led to an overall increase in the importance participants' assigned to ethical aspects.However, it was the recognition and understanding of ethical considerations, rather than ethical intentions or search activity, that drove shifts towards more responsible purchasing decisions.Participants who acknowledged and filled knowledge gaps in their decision making showed significant behaviour change, including increased searching and a stronger desire to alter their future shopping habits.We conclude that responsible consumption can be considered a partial information problem, where awareness of one's own knowledge limitations may be the catalyst needed for meaningful consumer behaviour change. Frans van der Sluis, Leif Azzopardi |
CHIIR | 2 |
| 2025 | Improving the Reusability of Conversational Search Test Collections
Zahra Abbasiantaeb, Chuan Meng, Leif Azzopardi, Mohammad Aliannejadi |
ECIR (1) | 3 |
| 2025 | Conversational Gold: Evaluating Personalized Conversational Search System Using Gold NuggetsabstractThe rise of personalized conversational search systems has been driven by advancements in Large Language Models (LLMs), enabling these systems to retrieve and generate answers for complex information needs. However, the automatic evaluation of responses generated by Retrieval Augmented Generation (RAG) systems remains an understudied challenge. In this paper, we introduce a new resource for assessing the retrieval effectiveness and relevance of responses generated by RAG systems, using a nugget-based evaluation framework. Built upon the foundation of TREC iKAT 2023, our dataset extends to the TREC iKAT 2024 collection, which includes 17 conversations and 20,575 relevance passage assessments, together with 2,279 extracted gold nuggets and 62 manually written gold answers from NIST assessors. While maintaining the core structure of its predecessor, this new collection enables a deeper exploration of generation tasks in conversational settings. Key improvements in iKAT 2024 include: (1) ''gold nuggets'' - concise, essential pieces of information extracted from relevant passages of the collection - which serve as a foundation for automatic response evaluation; (2) manually written answers to provide a gold standard for response evaluation; (3) expanded user personas, providing richer contextual grounding; and (4) a transition from Personal Text Knowledge Base (PTKB) ranking to PTKB classification and selection. Built on this resource, we provide a framework for long-form answer generation evaluation, involving nugget extraction and nugget matching, linked to retrieval. This establishes a solid resource for advancing research in personalized conversational search and long-form answer generation. Our resources are publicly available at https://github.com/irlabamsterdam/CONE-RAG. Zahra Abbasiantaeb, Simon Lupart, Leif Azzopardi, Jeff Dalton 0001, Mohammad Aliannejadi |
SIGIR | 3 |
| 2025 | From Query to Conscience: The Importance of Information Retrieval in Empowering Socially Responsible ConsumerismabstractMillions of consumers search for products online each day, aiming to find items that meet their needs at an acceptable price. While price and quality are major factors in purchasing decisions, ethical considerations increasingly influence consumer behavior -giving rise to the socially responsible consumer. Insights from a recent survey of over 600 consumers reveal that many barriers to ethical shopping stem from information-seeking challenges, often leading to decisions made under uncertainty. These challenges contribute to the intention-behaviour gap, where consumers' desire to make ethical choices is undermined by limited or inaccessible information and inefficacy of search systems in supporting responsible decision-making. In this perspectives paper, we argue that the field of Information Retrieval (IR) has a critical role to play by empowering consumers to make more informed and more responsible choices. We present three interrelated perspectives: (1) reframing ethical consumption as an information extraction problem aimed at reducing information asymmetries; (2) redefining product search as a complex task requiring interfaces that lower the cost and burden of responsible search; and (3) reimagining search as a process of knowledge calibration that helps consumers bridge gaps in awareness when making purchasing decisions. Taken together, these perspectives outline a path from query to conscience - one where IR systems help transform everyday product searches into opportunities for more ethical and informed choices. We advocate for the development of new and novel IR systems and interfaces that address the intricacies of socially responsible consumerism, and call on the IR community to build technologies that make ethical decisions more informed, convenient, and aligned with economic realities. Frans van der Sluis, Leif Azzopardi, Florian Meier 0001 |
SIGIR | 2 |
| 2024 | Search under Uncertainty: Cognitive Biases and Heuristics - Tutorial on Modeling Search Interaction using Behavioral EconomicsabstractModeling how people interact with search interfaces is core to the field of Interactive Information Retrieval. While various models have been proposed ranging from conceptual (e.g., Belkin’s ASK[12], Berry picking[11], Everyday-life information seeking, etc.) to theoretical (e.g., Information foraging theory[50], Economic theory[4], etc.), more recently there has been a body of working explore how people’s biases and the heuristics that they take influence how they search. This has led to the development of new models of the search process drawing upon Behavioural Economics and Psychology. This half day tutorial will provide a starting point for researchers seeking to learn more about information searching under uncertainty. The tutorial will be structured into two parts. First, we will provide an introduction of the biases and heuristics program put forward by Tversky and Kahneman [59] which assumes that people are not always rational. The second part of the tutorial will provide an overview of the types and space of biases in search [6, 42], before doing a deep dive into several specific examples and the impact of biases on different types of decisions (e.g., health/medical, financial etc.). The tutorial will wrap up with a discussion of some of the practical implication for how we can better design and evaluate IR systems in the light of cognitive biases. Leif Azzopardi, Jiqun Liu |
CHIIR | 1 |
| 2024 | Seeking Socially Responsible Consumers: Exploring the Intention-Search-Behaviour GapabstractThe increasing prominence of “Socially Responsible Consumers” has brought about a heightened focus on the ethical, environmental, social, and ideological dimensions influencing product purchasing decisions. Despite this emphasis, studies have consistently revealed a significant gap between individuals’ intentions to be socially responsible and their actual purchasing behaviors: they often choose products that do not align with their values. This paper aims to investigate the role of “search” and it how influences this gap. Our investigation involves an online survey of 286 participants, where we inquire about their search behaviors and whether they considered various dimensions—ranging from price and features to environmental, social, and governance issues — in relation to a recent purchase. Contrary to expectations of a clear intention-behavior gap, our findings suggest most participants exhibited indifference or lack of awareness regarding these “responsible” aspects. While, for those participants who were more ethically minded, they reported difficulties related to searching for and acquiring information regarding such aspects, which contributed to the gap. Our findings suggests that part of the intention-behaviour gap can be framed as an information seeking problem. Moreover our findings motivate the development of search systems and platforms that better help support consumers make more informed and responsible purchasing decisions. Leif Azzopardi, Frans van der Sluis |
CHIIR | 1 |
| 2024 | The Influence of Presentation and Performance on User SatisfactionabstractInformation Retrieval (IR) systems are designed to provide users with a ranked list of results based on their queries. The effectiveness of an IR system is gauged not just by its ability to retrieve relevant results but also by how it presents these results to users; an engaging presentation often correlates with increased user satisfaction. While existing research has delved into the link between user satisfaction, IR performance metrics, and presentation, these aspects have typically been investigated in isolation. Our research aims to bridge this gap by examining the relationship between query performance, presentation and user satisfaction. For our analysis, we conducted a between-subjects experiment comparing the effectiveness of various result card layouts for an ad-hoc news search interface. Drawing data from the TREC WaPo 2018 collection, we centered our study on four specific topics. Within each of these topics, we assessed six distinct queries with varying nDCG values. Our study involved 164 participants who were exposed to one of five distinct layouts containing result cards, such as “title”, “title+image”, or “title+image+summary”. Our findings indicate that while nDCG is a strong predictor of user satisfaction at the query level, there exists no linear relationship between the performance of the query, presentation of results and user satisfaction. However, when considering the total gain on the initial result page, we observed that presentation does play a significant role in user satisfaction (at the query level) for certain layouts with result cards such as, title+image or title+image+summary. Our results also suggest that the layout differences have complex and multifaceted impacts on satisfaction. We demonstrate the capacity to equalize user satisfaction levels between queries of varying performance by changing how results are presented. This emphasizes the necessity to harmonize both performance and presentation in IR systems, considering users’ diverse preferences. Ultimately, our insights can steer the evolution of more user-aligned IR systems, underscoring the balance between system performance and result presentation. Kanaad Pathak, Leif Azzopardi, Martin Halvey |
CHIIR | 2 |
| 2024 | Uncharted Territory: Understanding Exploratory Search Behaviours in Literature ReviewsabstractIn the realm of Information Seeking and Retrieval (ISR), searching the literature for relevant references in the context of academic work, such as theses or publications, is widely recognised as an exploratory search task. This task becomes particularly challenging when searchers lack prior knowledge of the subject matter. To help address the growing need for supporting exploratory search endeavours, ISR researchers have developed exploratory search models and interfaces. However, while much attention has been given to conceptualising exploratory searches, little focus has been placed on understanding the specific approaches and behaviours that searchers employ during conducting literature reviews. This paper aims to bridge this gap by conducting semi-structured interviews with 30 Master’s students at the end of a user study. This paper provides comprehensive definitions for the fundamental exploratory characteristics from the recent conceptual model, pinpoints potential factors that could influence these characteristics, introduces new exploratory dimensions, and extends our comprehension of existing ones in the academic context. It also uncovers a spectrum of approaches used in literature searches, shedding light on how individuals rely on specific paper sections to measure their relevance and highlighting essential facets of knowledge acquisition in the context of literature searches. Ayah Soufan, Ian Ruthven, Leif Azzopardi |
CHIIR | 3 |
| 2024 | Measuring Bias in a Ranked List Using Term-Based Representations
Amin Abolghasemi, Leif Azzopardi, Arian Askari, Maarten de Rijke, Suzan Verberne |
ECIR (5) | 2 |
| 2024 | The Search Futures Workshop
Leif Azzopardi, Charles L. A. Clarke, Paul B. Kantor, Bhaskar Mitra 0001, Johanne R. Trippas, Zhaochun Ren |
ECIR (5) | 1 |
| 2024 | Ranking Heterogeneous Search Result Pages Using the Interactive Probability Ranking Principle
Kanaad Pathak, Leif Azzopardi, Martin Halvey |
ECIR (2) | 2 |
| 2024 | TREC iKAT 2023: A Test Collection for Evaluating Conversational and Interactive Knowledge AssistantsabstractConversational information seeking has evolved rapidly in the last few years with the development of Large Language Models (LLMs), providing the basis for interpreting and responding in a naturalistic manner to user requests. The extended TREC Interactive Knowledge Assistance Track (iKAT) collection aims to enable researchers to test and evaluate their Conversational Search Agent (CSA). The collection contains a set of 36 personalized dialogues over 20 different topics each coupled with a Personal Text Knowledge Base (PTKB) that defines the bespoke user personas. A total of 344 turns with approximately 26,000 passages are provided as assessments on relevance, as well as additional assessments on generated responses over four key dimensions: relevance, completeness, groundedness, and naturalness. The collection challenges CSAs to efficiently navigate diverse personal contexts, elicit pertinent persona information, and employ context for relevant conversations. The integration of a PTKB and the emphasis on decisional search tasks contribute to the uniqueness of this test collection, making it an essential benchmark for advancing research in conversational and interactive knowledge assistants. Mohammad Aliannejadi, Zahra Abbasiantaeb, Shubham Chatterjee, Jeff Dalton 0001, Leif Azzopardi |
SIGIR | 5 |
| 2024 | Search under Uncertainty: Cognitive Biases and Heuristics: A Tutorial on Testing, Mitigating and Accounting for Cognitive Biases in Search ExperimentsabstractUnderstanding how people interact with search interfaces is core to the field of Interactive Information Retrieval (IIR). While various models have been proposed (e.g., Belkin's ASK, Berry picking, Everyday-life information seeking, Information foraging theory, Economic theory, etc.), they have largely ignored the impact of cognitive biases on search behaviour and performance. A growing body of empirical work exploring how people's cognitive biases influence search and judgments, has led to the development of new models of search that draw upon Behavioural Economics and Psychology. This full day tutorial will provide a starting point for researchers seeking to learn more about information seeking, search and retrieval under uncertainty. The tutorial will be structured into three parts. First, we will provide an introduction of the biases and heuristics program put forward by Tversky and Kahneman [60] (1974) which assumes that people are not always rational. The second part of the tutorial will provide an overview of the types and space of biases in search,[5, 40] before doing a deep dive into several specific examples and the impact of biases on different types of decisions (e.g., health/medical, financial). The third part will focus on a discussion of the practical implication regarding the design and evaluation human-centered IR systems in the light of cognitive biases - where participants will undertake some hands-on exercises. Jiqun Liu, Leif Azzopardi |
SIGIR | 2 |
| 2024 | Measuring the retrievability of digital library content using analytics dataabstractAbstract Digital libraries aim to provide value to users by housing content that is accessible and searchable. Often such access is afforded through external web search engines. In this article, we measure how easily digital library content can be retrieved (i.e., how retrievable) through a well‐known search engine (Google) using its analytics platforms. Using two measures of document retrievability, we contrast our results with simulation‐based studies that employed synthetic query sets. We determine that estimating the retrievability of content given a Digital Library index is not a strong predictor of how retrievable the content is in practice (via external search engines). Retrievability established the notion that search algorithms can be biased. In our work, we find that while there such bias is present, much of the variation in retrievability appears to be strongly influenced by the queries submitted to the library, a side of retrievability less examined in past work. Hamed Jahani, Leif Azzopardi, Mark Sanderson |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2024 | A Systematic Review of Cost, Effort, and Load Research in Information Search and Retrieval, 1972-2020abstractDuring the information search and retrieval (ISR) process, user-system interactions such as submitting queries, examining results, and engaging with information impose some degree of demand on the user’s resources. Within ISR, these demands are well recognised, and numerous studies have demonstrated that the cost, effort, and load (CEL) experienced during the search process are affected by a variety of factors. Despite this recognition, there is no universally accepted definition of the constructs of CEL within the field of ISR. Ultimately, this has led to problems with how these constructs have been interpreted and subsequently measured. This systematic review contributes a synthesis of literature, summarising key findings relating to how researchers have been defining and measuring CEL within ISR over the past 50 years. After manually screening 1,109 articles, we detailed and analysed 91 articles which examine CEL within ISR. The discussion focuses on comparing the similarities and differences between CEL definitions and measures before identifying the limitations of the current state of the nomenclature. Opportunities for future research are also identified. Going forward, we propose a CEL taxonomy that integrates the relationships between CEL and their related constructs, which will help focus and disambiguate future research in this important area. Molly McGregor, Leif Azzopardi, Martin Halvey |
ACM Trans. Inf. Syst. | 2 |
| 2023 | Driven to Distraction: Examining the Influence of Distractors on Search Behaviours, Performance and ExperienceabstractAdvertisements, sponsored links, clickbait, in-house recommendations and similar elements pervasively shroud featured content. Such elements vie for people’s attention, potentially distracting people from their task at hand. The effects of such “distractors” is likely to increase people’s cognitive workload and reduce their performance as they need to work harder to discern the relevant from non-relevant. In this paper, we investigate how people of varying cognitive abilities (measured using Perceptual Speed and Cognitive Failure instruments) are affected by these different types of distractions when completing search tasks. We performed a crowdsourced within-subjects user study, where 102 participants completed four search tasks using our news search engine over four different interface conditions: (i) one with no additional distractors; (ii) one with advertisements; (iii) one with sponsored links; and (iv) one with in-house recommendations. Our results highlight a number of important trends and findings. Participants perceived the interface condition without distractors as significantly better across numerous dimensions. Participants reported higher satisfaction, lower workload, higher topic recall, and found it easier to concentrate. Behaviourally, participants issued queries faster and clicked results earlier when compared to the interfaces with distractors. When using the interfaces with distractors, one in ten participants clicked on a distractor—and despite engaging with a distractor for less than twenty seconds, their task time increased by approximately two minutes. We found that the effects were magnified depending on cognitive abilities—with a greater impact of distractors on participants with lower perceptual speed, and for those with a higher propensity of cognitive failures. Distractors—regardless of their type—have negative consequences on a user’s search experience and performance. As a consequence, interfaces containing visually distracting elements are creating poorer search experiences due to the “distractor tax” being placed on people’s limited attention. Leif Azzopardi, David Maxwell 0001, Martin Halvey, Claudia Hauff |
CHIIR | 1 |
| 2023 | Retrievability Bias Estimation Using Synthetically Generated QueriesabstractRanking with pre-trained language models (PLMs) has shown to be highly effective for various Information Retrieval tasks. Previous studies investigated the performance of these models in terms of effectiveness and efficiency. However, there is no prior work on evaluating PLM-based rankers in terms of their retrievability bias. In this paper, we evaluate the retrievability bias of PLM-based rankers with the use of synthetically generated queries. We compare the retrievability bias in two of the most common PLM-based rankers, a Bi-Encoder BERT ranker and a Cross-Encoder BERT re-ranker against BM25, which was found to be one of the least biased models in prior work. We conduct a series of experiments with which we explore the plausibility of using synthetic queries generated with a generative model, docT5query, in the evaluation of retrievability bias. Our experiments show promising results on the use of synthetically generated queries for the purpose of retrievability bias estimation. Moreover, we find that the estimated bias values resulting from synthetically generated queries are lower than the ones estimated with user-generated queries on the MS MARCO evaluation benchmark. This indicates that synthetically generated queries might cause less bias than user-generated queries and therefore, by using such queries in training PLM-based rankers, we might be able to reduce the retrievability bias in these models. Amin Abolghasemi, Suzan Verberne, Arian Askari, Leif Azzopardi |
CIKM | 4 |
| 2022 | Searching the Literature: An Analysis of an Exploratory Search TaskabstractExploratory search is an intuitive concept in interactive information retrieval. While many definitions for Exploratory Search have been proposed, the main dimensions involve high uncertainty with respect to the problem context, the user expertise, and the search process. In this paper, we draw together the different characteristics relating to the three main exploratory dimensions to provide a conceptual model of exploratory search. We build an exploratory search questionnaire using this model. We then use the questionnaire instrument to examine how literature searches are exploratory and what factors influence the exploratory dimensions and characteristics. We provide one of the first detailed investigations into the nature of exploratory literature review searches. Our analysis of the 368 responses reveals that about 84% of the participants described their literature review task as somewhat exploratory or very exploratory in nature. Also, the analysis points to another dimension of the exploratory search, the knowledge gain/change dimension. Furthermore, we investigated how users’ experience influences how people rate the exploratory search characteristics (e.g., whether they find surprising information, learn new concepts and keywords). Ayah Soufan, Ian Ruthven, Leif Azzopardi |
CHIIR | 3 |
| 2022 | Improving BERT-based Query-by-Document Retrieval with Multi-task Optimization
Amin Abolghasemi, Suzan Verberne, Leif Azzopardi |
ECIR (2) | 3 |
| 2022 | Towards Building Economic Models of Conversational Search
Leif Azzopardi, Mohammad Aliannejadi, Evangelos Kanoulas |
ECIR (2) | 1 |
| 2022 | Are Taylor's Posts Risky? Evaluating Cumulative Revelations in Online Personal Data: A persona-based tool for evaluating awareness of online risks and harmsabstractSearching for people online is a common search task that most of us have performed at some point or other. With so much information about people available online it is often amazing what one can find out about someone else -- especially when information taken from different sources is pieced together to create a more detailed picture of the individual, and then used to make inferences about them (leading to cumulative revelations ). As such, the relevance of one piece of information is often conditional and dependent on other pieces of information found. This creates interesting and novel challenges in evaluating informationrelevance when searching personal profiles, posts and related information about an individual, as well as the potential risks that can arise from such revelations. In this demonstration paper, we present a tool designed to investigate how people assess and judge the relevance and potential risks ofsmall, apparently innocuous pieces of information associated with fictitious personas, such as Taylor Addison, when searching and browsing online profiles and social media. The demonstrator also comprises a cyber-safety tool, which aims to provide education and raise awareness of the potential risks of cumulative revelations. It does so by engaging participants in different scenarios where the relevance of individual information items depends on the searcher and their particular underlying motivation. Leif Azzopardi, Jo Briggs, Melissa Duheric, Callum Nash, Emma Nicol, Wendy Moncur, Burkhard Schafer 0001 |
SIGIR | 1 |
| 2022 | A Flexible Framework for Offline Effectiveness MetricsabstractThe use of offline effectiveness metrics is one of the cornerstones of evaluation in information retrieval. Static resources that include test collections and sets of topics, the corresponding relevance judgments connecting them, and metrics that map document rankings from a retrieval system to numeric scores have been used for multiple decades as an important way of comparing systems. The basis behind this experimental structure is that the metric score for a system can serve as a surrogate measurement for user satisfaction. Alistair Moffat, Joel Mackenzie, Paul Thomas 0001, Leif Azzopardi |
SIGIR | 4 |
| 2021 | Cognitive Biases in Search: A Review and Reflection of Cognitive Biases in Information RetrievalabstractPeople are susceptible to an array of cognitive biases, which can result in systematic errors and deviations from rational decision making. Over the past decade, an increasing amount of attention has been paid towards investigating how cognitive biases influence information seeking and retrieval behaviours and outcomes. In particular, how such biases may negatively affect decisions because, for example, searchers may seek confirmatory but incorrect information or anchor on an initial search result even if its incorrect. In this perspectives paper, we aim to: (1) bring together and catalogue the emerging work on cognitive biases in the field of Information Retrieval; and (2) provide a critical review and reflection on these studies and subsequent findings. During our analysis we report on over thirty studies, that empirically examined cognitive biases in search, providing over forty key findings related to different domains (e.g. health, web, socio-political) and different parts of the search process (e.g. querying, assessing, judging, etc.). Our reflection highlights the importance of this research area, and critically discusses the limitations, difficulties and challenges when investigating this phenomena along with presenting open questions and future directions in researching the impact --- both positive and negative --- of cognitive biases in Information Retrieval. Leif Azzopardi |
CHIIR | 1 |
| 2021 | User Models, Metrics and Measures of Search: A Tutorial on the C/W/L Evaluation FrameworkabstractEvaluation is central to Information Retrieval, and is how we compare the quality of systems. One important principle of evaluation is that the measured score should reflect the user's experience with the system. Hence, there should be direct connection between how users interact with the system and the characteristics of the metric. In this tutorial we introduce the C/W/L approach to user modeling and show how different user models lead to different metrics. We then describe the recent innovations and approaches to evaluation that it has facilitated. The tutorial is presented as a mix of on-line synchronous lecture, pre-recorded in-depth videos, and hands-on activities using the C/W/L toolkit for participants' own evaluation tasks. A followup consultation session is also provided, to allow extended questions and individual discussion with the four presenters. Leif Azzopardi, Alistair Moffat, Paul Thomas 0001, Guido Zuccon |
CHIIR | 1 |
| 2021 | Investigating the Influence of Ads on User Search Performance, Behaviour, and Experience during Information SeekingabstractThe phenomenon of banner blindness explains that users can mentally ignore online advertisements (ads). However, eye-tracking studies have shown that users still fixate on ads, and even without direct gaze, ads still fall within a user's peripheral vision, which may negatively overload cognition. It is therefore unknown how blind, banner blindness, truly is, and what other effect ads may have on user's information seeking. To address this gap, a within-subjects design experiment was conducted with 37 participants who performed search tasks from the TREC 2017 Common Core News Collection, where 3 search tasks contained various types of ads, and one search task had no ads. Although our results showed that on average, participants retrieved similar amounts of relevant documents regardless of whether ads were present or absent, participants took significantly longer achieving this performance when ads were present. Furthermore, when ads were absent, participants reported less frustration, and not only believed they learned more, but a post-task recall test showed that participants actually did learn up to 38% more. Consequently, our findings suggest that banner blindness is more costly than just mere annoyance, and that the influence of ads on user's information retrieval recall may extend current theories of visual crowding. Olivia Foulds, Leif Azzopardi, Martin Halvey |
CHIIR | 2 |
| 2021 | Untangling Cost, Effort, and Load in Information Seeking and RetrievalabstractWhen performing Information Seeking and Retrieval (ISR) activities, people submit queries, examine results, assess documents and engage with the information to make decisions and complete tasks. All these activities come at a "cost'', but within the field of ISR there is no universally accepted definition of the concepts of Cost, Effort and Load (CEL). Instead, researchers have used the same terms interchangeably to describe similar but also different concepts. This lack of shared understanding has led to a disconnect between how these concepts are defined and discussed versus how they are interpreted and measured. Thus, the aim of this paper is two-fold: (i) to review the meaning of CEL related concepts used within ISR, and (ii) to create a shared taxonomy of the concepts relating to CEL in ISR. To seed our analysis, we conducted a literature review, where 397 papers were reviewed, and twenty-six papers that explicitly proposed measures or definitions of CEL were selected for analysis. By drawing upon theory from Psychology and other fields, we present the common definitions of CEL in order to ground our discussion of these concepts in ISR. We also highlight the issues associated with CEL measurement in ISR to help researchers reflect on the validity and precision of existing methods. We hope this perspectives paper serves as a basis for a taxonomy of how CEL concepts are used within ISR- where we have provided a series of working definitions that clearly delineate the different concepts being used, investigated and measured in ISR research. Molly McGregor, Leif Azzopardi, Martin Halvey |
CHIIR | 2 |
| 2021 | Analysing Mixed Initiatives and Search Strategies during Conversational SearchabstractInformation seeking conversations between users and Conversational Search Agents (CSAs) consist of multiple turns of interaction. While users initiate a search session, ideally a CSA should sometimes take the lead in the conversation by obtaining feedback from the user by offering query suggestions or asking for query clarifications i.e. mixed initiative. This creates the potential for more engaging conversational searches, but substantially increases the complexity of modelling and evaluating such scenarios due to the large interaction space coupled with the trade-offs between the costs and benefits of the different interactions. In this paper, we present a model for conversational search -- from which we instantiate different observed conversational search strategies, where the agent elicits: (i) Feedback-First, or (ii) Feedback-After. Using 49 TREC WebTrack Topics, we performed an analysis comparing how well these different strategies combine with different mixed initiative approaches: (i) Query Suggestions vs. (ii) Query Clarifications. Our analysis reveals that there is no superior or dominant combination, instead it shows that query clarifications are better when asked first, while query suggestions are better when asked after presenting results. We also show that the best strategy and approach depends on the trade-offs between the relative costs between querying and giving feedback, the performance of the initial query, the number of assessments per query, and the total amount of gain required. While this work highlights the complexities and challenges involved in analyzing CSAs, it provides the foundations for evaluating conversational strategies and conversational search agents in batch/offline settings. Mohammad Aliannejadi, Leif Azzopardi, Hamed Zamani, Evangelos Kanoulas, Paul Thomas 0001, Nick Craswell |
CIKM | 2 |
| 2021 | AWESSOME: An Unsupervised Sentiment Intensity Scoring Framework Using Neural Word Embeddings
Amal Htait, Leif Azzopardi |
ECIR (2) | 2 |
| 2020 | Data-Driven Evaluation Metrics for Heterogeneous Search Engine Result PagesabstractEvaluation metrics for search typically assume items are homoge- neous. However, in the context of web search, this assumption does not hold. Modern search engine result pages (SERPs) are composed of a variety of item types (e.g., news, web, entity, etc.), and their influence on browsing behavior is largely unknown. Leif Azzopardi, Ryen W. White, Paul Thomas 0001, Nick Craswell |
CHIIR | 1 |
| 2020 | Made to Measure: A Workshop on Human-centred metrics for information seekingabstractMetrics of human behaviour and effort lie at the heart of improving information interaction and retrieval. However, while some measurements have become predominant, such as precision and recall, there are many elements of information interaction where either measures have yet to be created, accepted, or widely used. This workshop seeks to tease out these areas, finding novel measures, or novel uses of existing measures, to create better experimental tools to improve our understanding of information interaction, and help develop better systems to support it. George Buchanan 0001, Dana McKay, Charles L. A. Clarke, Leif Azzopardi, Johanne R. Trippas |
CHIIR | 4 |
| 2020 | Reflecting upon Perceptual Speed Tests in Information Retrieval: Limitations, Challenges, and RecommendationsabstractPerceptual Speed (PS) is a cognitive ability defined by an individual's accuracy and speed to scan information while completing visual search tasks. Prior studies using PS tests have demonstrated that PS affects multiple factors in Information Retrieval (IR), such as a user's search performance, interaction with the system, time spent completing tasks, and subjective impression of their workload. With greater knowledge of PS, systems could be designed that accommodate users with low PS to improve their overall search experience. However, in this perspectives paper, we analyse how PS tests have been used in IR, and identify multiple uncertainties regarding PS content, administration, analysis, and reporting of findings. Consequently, we aim to stir discussion between IR researchers by drawing awareness to these issues. As a result, we further discuss challenges involved in advancing how future PS tests are used in IR. Finally, we propose recommendations that have the potential for enhancing the reliability and validity of current PS tests. Olivia Foulds, Leif Azzopardi, Martin Halvey |
CHIIR | 2 |
| 2020 | Predicting Perceptual Speed from Search BehaviourabstractPerceptual Speed (PS) is a cognitive ability that is known to affect multiple factors in Information Retrieval (IR) such as a user's search performance and subjective experience. However PS tests are difficult to administer which limits the design of user-adaptive systems that can automatically infer PS to appropriately accommodate low PS users. Consequently, this paper evaluated whether PS can be automatically classified from search behaviour using several machine learning models trained on features extracted from TREC Common Core search task logs. Our results are encouraging: given a user's interactions from one query, a Decision Tree was able to predict a user's PS as low or high with 86% accuracy. Additionally, we identified different behavioural components for specific PS tests, implying that each PS test measures different aspects of a person's cognitive ability. These findings motivate further work for how best to design search systems that can adapt to individual differences. Olivia Foulds, Alessandro Suglia, Leif Azzopardi, Martin Halvey |
SIGIR | 3 |
| 2020 | DataMirror: Reflecting on One's Data Self: A Tool for Social Media Users to Explore Their Digital FootprintsabstractSmall pieces of data that are shared online, over time and across multiple social networks, have the potential to reveal more cumulatively than a person intends. This could result in harm, loss or detriment to them depending what information is revealed, who can access it, and how it is processed. But how aware are social network users of how much information they are actually disclosing? And if they could examine all their data, what cumulative revelations might be found that could potentially increase their risk of various online threats (social engineering, fraud, identify theft, loss of face, etc.)? In this paper, we present DataMirror, an initial prototype tool, that enables social network users to aggregate their online data so that they can search, browse and visualise what they have put online. The aim of the tool is to investigate and explore people's awareness of their data self that is projected online; not only in terms of the volume of information that they might share, but what it may mean when combined together, what pieces of sensitive information may be gleaned from their data, and what machine learning may infer about them given their data. Amal Htait, Leif Azzopardi, Emma Nicol, Wendy Moncur |
SIGIR | 2 |
| 2019 | CLEF eHealth 2019 Evaluation Lab
Liadh Kelly, Lorraine Goeuriot, Hanna Suominen, Mariana L. Neves, Evangelos Kanoulas, René Spijker, Leif Azzopardi, Dan Li 0015, Jimmy, João R. M. Palotti, Guido Zuccon |
ECIR (2) | 7 |
| 2019 | Building Economic Models and Measures of SearchabstractEconomics provides an intuitive and natural way to formally represent the costs and benefits of interacting with applications, interfaces and devices. By using economic models it is possible to reason about interaction, make predictions about how changes to the system will affect behavior, and measure the performance of people's interactions with the system. In this tutorial, we first provide an overview of relevant economic theories, before showing how they can be applied to formulate different ranking principles to provide the optimal ranking to users. This is followed by a session showing how economics can be used to model how people interact with search systems, and how to use these models to generate hypotheses about user behavior. The third session focuses on how economics has been used to underpin the measurement of information retrieval systems and applications using the CWL framework (which reports the expected utility, expected total utility, expected total cost, and so on) -- and how different models of user interaction lead to different metrics. We then show how information foraging theory can be used to measure the performance of an information retrieval system -- connecting the theory of how people search with how we measure it. The final session of the day will be spent building economic models and measures of search. Here sample problems will be provided to challenge participants, or participants can bring their own. Leif Azzopardi, Alistair Moffat, Paul Thomas 0001, Guido Zuccon |
SIGIR | 1 |
| 2019 | cwl_eval: An Evaluation Tool for Information RetrievalabstractWe present a tool ("cwl_eval") which unifies many metrics typically used to evaluate information retrieval systems using test collections. In the CWL framework metrics are specified via a single function which can be used to derive a number of related measurements: Expected Utility per item, Expected Total Utility, Expected Cost per item, Expected Total Cost, and Expected Depth. The CWL framework brings together several independent approaches for measuring the quality of a ranked list, and provides a coherent user model-based framework for developing measures based on utility (gain) and cost. Here we outline the CWL measurement framework; describe the cwl_eval architecture; and provide examples of how to use it. We provide implementations of a number of recent metrics, including Time Biased Gain, U-Measure, Bejewelled Measure, and the Information Foraging Based Measure, as well as previous metrics such as Precision, Average Precision, Discounted Cumulative Gain, Rank-Biased Precision, and INST. By providing state-of-the-art and traditional metrics within the same framework, we promote a standardised approach to evaluating search effectiveness. Leif Azzopardi, Paul Thomas 0001, Alistair Moffat |
SIGIR | 1 |
| 2019 | Looking for Opportunities: Challenges in Procurement SearchabstractProcurement legislation stipulates that information about the goods, services, or works, that tax-funded authorities wish to purchase are made publicly available in a procurement contract notice. However, for businesses wishing to tender for such competitive opportunities, finding relevant procurement contract notices presents a challenging professional search task. In this talk, we will provide an overview of procurement search and then describe the challenges in addressing the related search and recommendation tasks. Stuart Mackie, David Macdonald, Leif Azzopardi, Yashar Moshfeghi |
SIGIR | 3 |
| 2019 | The impact of result diversification on search behaviour and performanceabstractResult diversification aims to provide searchers with a broader view of a given topic while attempting to maximise the chances of retrieving relevant material. Diversifying results also aims to reduce search bias by increasing the coverage over different aspects of the topic. As such, searchers should learn more about the given topic in general. Despite diversification algorithms being introduced over two decades ago, little research has explicitly examined their impact on search behaviour and performance in the context of Interactive Information Retrieval (IIR) . In this paper, we explore the impact of diversification when searchers undertake complex search tasks that require learning about different aspects of a topic (aspectual retrieval) . We hypothesise that by diversifying search results, searchers will be exposed to a greater number of aspects. In turn, this will maximise their coverage of the topic (and thus reduce possible search bias). As a consequence, diversification should lead to performance benefits, regardless of the task, but how does diversification affect search behaviours and search satisfaction? Based on Information Foraging Theory (IFT) , we infer two hypotheses regarding search behaviours due to diversification, namely that (i) it will lead to searchers examining fewer documents per query, and (ii) it will also mean searchers will issue more queries overall. To this end, we performed a within-subjects user study using the TREC AQUAINT collection with 51 participants, examining the differences in search performance and behaviour when using (i) a non-diversified system ( BM25 ) versus (ii) a diversified system (BM25 + xQuAD ) when the search task is either (a) ad-hoc or (b) aspectual. Our results show a number of notable findings in terms of search behaviour: participants on the diversified system issued more queries and examined fewer documents per query when performing the aspectual search task. Furthermore, we showed that when using the diversified system, participants were: more successful in marking relevant documents, and obtained a greater awareness of the topics (i.e. identified relevant documents containing more novel aspects). These findings show that search behaviour is influenced by diversification and task complexity. They also motivate further research into complex search tasks such as aspectual retrieval—and how diversity can play an important role in improving the search experience, by providing greater coverage of a topic and mitigating potential bias in search results. David Maxwell 0001, Leif Azzopardi, Yashar Moshfeghi |
Inf. Retr. J. | 2 |
| 2018 | Information Scent, Searching and Stopping - Modelling SERP Level Stopping Behaviour
David Maxwell 0001, Leif Azzopardi |
ECIR | 2 |
| 2018 | Measuring the Utility of Search Engine Result Pages: An Information Foraging Based MeasureabstractWeb Search Engine Result Pages (SERPs) are complex responses to queries, containing many heterogeneous result elements (web results, advertisements, and specialised "answers'') positioned in a variety of layouts. This poses numerous challenges when trying to measure the quality of a SERP because standard measures were designed for homogeneous ranked lists. In this paper, we aim to measure the utility and cost of SERPs. To ground this work we adopt the \CWL framework which enables a direct comparison between different measures in the same units of measurement, i.e. expected (total) utility and cost. Within this framework, we propose a new measure based on information foraging theory, which can account for the heterogeneity of elements, through different costs, and which naturally motivates the development of a user stopping model that adapts behaviour depending on the rate of gain. This directly connects models of how people search with how we measure search, providing a number of new dimensions in which to investigate and evaluate user behaviour and performance. We perform an analysis over~1000 popular queries issued to a major search engine, and report the aggregate utility experienced by users over time. Then in an comparison against common measures, we show that the proposed foraging based measure provides a more accurate reflection of the utility and of observed behaviours (stopping rank and time spent). Leif Azzopardi, Paul Thomas 0001, Nick Craswell |
SIGIR | 1 |
| 2018 | Query Variation Performance Prediction for Systematic ReviewsabstractWhen conducting systematic reviews, medical researchers heavily deliberate over the final query to pose to the information retrieval system. Given the possible query variations that they could construct, selecting the best performing query is difficult. This motivates a new type of query performance prediction (QPP) task where the challenge is to estimate the performance of a set of query variations given a particular topic. Query variations are the reductions, expansions and modifications of a given seed query under the hypothesis that there exists some variations (either generated from permutations or hand crafted) which will improve retrieval effectiveness over the original query. We use the CLEF 2017 TAR Collection, to evaluate sixteen pre and post retrieval predictors for the task of Query Variation Performance Prediction (QVPP). Our findings show the IDF based QPPs exhibits the strongest correlations with performance. However, when using QPPs to select the best query, little improvement over the original query can be obtained, despite the fact that there are query variations which perform significantly better. Our findings highlight the difficulty in identifying effective queries within the context of this new task, and motivates further research to develop more accurate methods to help systematic review researchers in the query selection process. Harrisen Scells, Leif Azzopardi, Guido Zuccon, Bevan Koopman |
SIGIR | 2 |
| 2018 | Information retrieval in the workplace: A comparison of professional search practices
Tony Russell-Rose, Jon Chamberlain, Leif Azzopardi |
Inf. Process. Manag. | 3 |
| 2017 | Building Cost-Benefit Models of Information InteractionsabstractModeling how people interact with search interfaces has been of particular interest and importance to the field of Interactive Information Retrieval. Recently, there has been a move to developing formal models of the interaction between the user and the system, whether it be to: (i) run a simulation, (ii) conduct an economic analysis, (iii) measure system performance, or (iv) simply to better understand user interactions and hypothesise about user behaviours. In such models, they consider the costs and the benefits that arise through the interaction with the interface/system and the information surfaced during the course of interaction. In this half day tutorial, we will focus on describing a series of cost-benefit models that have been proposed in the literature and how they have been applied in various scenarios. The tutorial will be structured into two parts. First, we will provide an overview of Decision Theory and Cost-Benefit Analysis techniques, and how they can and have be applied to a variety of Interactive Information Retrieval scenarios. For example, when do facets helps?, under what conditions are query suggestions useful? and is it better to bookmark or re-find? The second part of the tutorial will be dedicated to building cost-benefit models where we will discuss different techniques to build and develop such models. In the practical session, we will also discuss how costs and benefits can be estimated, and how the models can help inform and guide experimentation. During the tutorial participants will be challenged to build cost models for a number of problems (or even bring their own problems to solve). Leif Azzopardi |
CHIIR | 1 |
| 2017 | Second International Workshop On the Evaluation of Collaborative Information Seeking and Retrieval (Ecol'17)abstractThe workshop on the evaluation of collaborative information retrieval and seeking (ECol) is held in conjunction with the ACM SIGIR Conference on Human Information Interaction & Retrieval (CHIIR) in Oslo, Norway. To make the workshop active and the participant pro-active, we released datasets and tools so as to help researchers contributing to the formalization of evaluation frameworks for challenging collaborative tasks. The workshop is split into two parts. First, a presentation session. Then, the afternoon is devoted to group discussion addressing challenges of evaluating and designing models for social and collaborative search. Leif Azzopardi, Jeremy Pickens, Chirag Shah 0001, Laure Soulier, Lynda Tamine-Lechani |
CHIIR | 1 |
| 2017 | An Empirical Analysis of Pruning Techniques: Performance, Retrievability and BiasabstractPrior work on using retrievability measures in the evaluation of information retrieval (IR) systems has laid out the foundations for investigating the relation between retrieval performance and retrieval bias. While various factors influencing retrievability have been examined, showing how the retrieval model may influence bias, no prior work has examined the impact of the index (and how it is optimized) on retrieval bias. Intuitively, how the documents are represented, and what terms they contain, will influence whether they are retrievable or not. In this paper, we investigate how the retrieval bias of a system changes as the inverted index is optimized for efficiency through static index pruning. In our analysis, we consider four pruning methods and examine how they affect performance and bias on the TREC GOV2 Collection. Our results show that the relationship between these factors is varied and complex - and very much dependent on the pruning algorithm. We find that more pruning results in relatively little change or a slight decrease in bias up to a point, and then a dramatic increase. The increase in bias corresponds to a sharp decrease in early precision such as [email protected] and is also indicative of a large decrease in MAP. The findings suggest that the impact of pruning algorithms can be quite varied - but retrieval bias could be used to guide the pruning process. Further work is required to determine precisely which documents are most affected and how this impacts upon performance. Ruey-Cheng Chen, Leif Azzopardi, Falk Scholer |
CIKM | 2 |
| 2017 | Integrating the Framing of Clinical Questions via PICO into the Retrieval of Medical Literature for Systematic ReviewsabstractThe PICO process is a technique used in evidence based practice to frame and answer clinical questions. It involves structuring the question around four types of clinical information: population, intervention, control or comparison and outcome. The PICO framework is used extensively in the compilation of systematic reviews as the means of framing research questions. However, when a search strategy (comprising of a large Boolean query) is formulated to retrieve studies for inclusion in the review, PICO is often ignored. This paper evaluates how PICO annotations can be applied and integrated into retrieval to improve the screening of studies for inclusion in systematic reviews. The task is to increase precision while maintaining the high level of recall essential to ensure systematic reviews are representative and unbiased. Our results show that restricting the search strategies to match studies using PICO annotations improves precision, however recall is slightly reduced, when compared to the non-PICO baseline. This can lead to both time and cost savings when compiling systematic reviews. Harrisen Scells, Guido Zuccon, Bevan Koopman, Anthony Deacon, Leif Azzopardi, Shlomo Geva |
CIKM | 5 |
| 2017 | Algorithmic Bias: Do Good Systems Make Relevant Documents More Retrievable?abstractAlgorithmic bias presents a difficult challenge within Information Retrieval. Long has it been known that certain algorithms favour particular documents due to attributes of these documents that are not directly related to relevance. The evaluation of bias has recently been made possible through the use of retrievability, a quantifiable measure of bias. While evaluating bias is relatively novel, the evaluation of performance has been common since the dawn of the Cranfield approach and TREC. To evaluate performance, a pool of documents to be judged by human assessors is created from the collection. This pooling approach has faced accusations of bias due to the fact that the state of the art algorithms were used to create it, thus the inclusion of biases associated with these algorithms may be included in the pool. The introduction of retrievability has provided a mechanism to evaluate the bias of these pools. This work evaluates the varying degrees of bias present in the groups of relevant and non-relevant documents for topics. The differentiating power of a system is also evaluated by examining the documents from the pool that are retrieved for each topic. The analysis finds that the systems that perform better, tend to have a higher chance of retrieving a relevant document rather than a non-relevant document for a topic prior to retrieval, indicating that retrieval systems which perform better at TREC are already predisposed to agree with the judgements regardless of the query posed. Colin Wilkie, Leif Azzopardi |
CIKM | 2 |
| 2017 | A Task Completion Engine to Enhance Search Session Support for Air Traffic Work Tasks
Yashar Moshfeghi, Raoul Rothfeld, Leif Azzopardi, Peter Triantafillou |
ECIR | 3 |
| 2017 | The Lucene for Information Access and Retrieval Research (LIARR) Workshop at SIGIR 2017abstractAs an empirical discipline, information access and retrieval research requires substantial software infrastructure to index and search large collections. This workshop is motivated by the desire to better align information retrieval research with the practice of building search applications from the perspective of open-source information retrieval systems. Our goal is to promote the use of Lucene for information access and retrieval research. Leif Azzopardi, Matt Crane, Hui Fang 0001, Grant Ingersoll, Jimmy Lin, Yashar Moshfeghi, Harrisen Scells, Guido Zuccon |
SIGIR | 1 |
| 2017 | Exploring the Query Halo Effect in Site Search: Leading People to Longer QueriesabstractPeople tend to type short queries, however, the belief is that longer queries are more effective. Consequently, a number of attempts have been made to encourage and motivate people to enter longer queries. While most have failed, a recent attempt - conducted in a laboratory setup - in which the query box has a halo or glow effect, that changes as the query becomes longer, has been shown to increase query length by one term, on average. In this paper, we test whether a similar increase is observed when the same component is deployed in a production system for site search and used by real end users. To this end, we conducted two separate experiments, where the rate at which the color changes in the halo were varied. In both experiments users were assigned to one of two conditions: halo and no-halo. The experiments were ran over a fifty day period with 3,506 unique users submitting over six thousand queries. In both experiments, however, we observed no significant difference in query length. We also did not find longer queries to result in greater retrieval performance. While, we did not reproduce the previous findings, our results indicate that the query halo effect appears to be sensitive to performance and task, limiting its applicability to other contexts. Djoerd Hiemstra, Claudia Hauff, Leif Azzopardi |
SIGIR | 3 |
| 2017 | A Study of Snippet Length and Informativeness: Behaviour, Performance and User ExperienceabstractThe design and presentation of a Search Engine Results Page (SERP) has been subject to much research. With many contemporary aspects of the SERP now under scrutiny, work still remains in investigating more traditional SERP components, such as the result summary. Prior studies have examined a variety of different aspects of result summaries, but in this paper we investigate the influence of result summary length on search behaviour, performance and user experience. To this end, we designed and conducted a within-subjects experiment using the TREC AQUAINT news collection with 53 participants. Using Kullback-Leibler distance as a measure of information gain, we examined result summaries of different lengths and selected four conditions where the change in information gain was the greatest: (i) title only; (ii) title plus one snippet; (iii) title plus two snippets; and (iv) title plus four snippets. Findings show that participants broadly preferred longer result summaries, as they were perceived to be more informative. However, their performance in terms of correctly identifying relevant documents was similar across all four conditions. Furthermore, while the participants felt that longer summaries were more informative, empirical observations suggest otherwise; while participants were more likely to click on relevant items given longer summaries, they also were more likely to click on non-relevant items. This shows that longer is not necessarily better, though participants perceived that to be the case - and second, they reveal a positive relationship between the length and informativeness of summaries and their attractiveness (i.e. clickthrough rates). These findings show that there are tensions between perception and performance when designing result summaries that need to be taken into account. David Maxwell 0001, Leif Azzopardi, Yashar Moshfeghi |
SIGIR | 2 |
| 2017 | A Test Collection for Evaluating Retrieval of Studies for Inclusion in Systematic ReviewsabstractThis paper introduces a test collection for evaluating the effectiveness of different methods used to retrieve research studies for inclusion in systematic reviews. Systematic reviews appraise and synthesise studies that meet specific inclusion criteria. Systematic reviews intended for a biomedical science audience use boolean queries with many, often complex, search clauses to retrieve studies; these are then manually screened to determine eligibility for inclusion in the review. This process is expensive and time consuming. The development of systems that improve retrieval effectiveness will have an immediate impact by reducing the complexity and resources required for this process. Our test collection consists of approximately 26 million research studies extracted from the freely available MEDLINE database, 94 review (query) topics extracted from Cochrane systematic reviews, and corresponding relevance assessments. Tasks for which the collection can be used for information retrieval system evaluation are described and the use of the collection to evaluate common baselines within one such task is demonstrated. The test collection is available at https://github.com/ielab/SIGIR2017-PICO-Collection. Harrisen Scells, Guido Zuccon, Bevan Koopman, Anthony Deacon, Leif Azzopardi, Shlomo Geva |
SIGIR | 5 |
| 2017 | Validating simulated interaction for retrieval evaluation
Teemu Pääkkönen, Jaana Kekäläinen, Heikki Keskustalo, Leif Azzopardi, David Maxwell 0001, Kalervo Järvelin |
Inf. Retr. J. | 4 |
| 2016 | Impacts of Time Constraints and System Delays on User ExperienceabstractDuring information search, people often experience time pressure. This might be a result of a deadline, the system's performance or some other event. In this paper, we report results of a study with forty-five participants which investigated how time constraints and system delays impacted the user experience during information search. We randomly assigned half of our study participants to a treatment condition where they were only allowed five minutes per search task (the other half were given no time limits). For half of participants' search tasks, five second delays were introduced after queries were submitted and SERP results were clicked. We used multilevel modeling to evaluate a number of hypotheses about the effects of time constraint, system delays and user experience. We found those in the time constraint condition reported significantly greater time pressure, experienced higher task difficulty, less satisfaction with their performance, increased importance of working fast and engaged in more metacognitive monitoring. We found when experiencing system delays participants reported slower system speeds when encountering delays on the second task. This work opens a new line of inquiry into how time pressure impacts the search experience and how tools and interfaces might be designed to support people who are searching under time pressure. It also presents an example of how multilevel modeling can be used to better understand and model the complex interactions that occur during interactive information retrieval. Anita Crescenzi, Diane Kelly 0001, Leif Azzopardi |
CHIIR | 3 |
| 2016 | Agents, Simulated Users and Humans: An Analysis of Performance and BehaviourabstractMost of the current models that are used to simulate users in Interactive Information Retrieval (IIR) lack realism and agency. Such models generally make decisions in a stochastic manner, without recourse to the actual information encountered or the underlying information need. In this paper, we develop a more sophisticated model of the user that includes their cognitive state within the simulation. The cognitive state maintains data about what the simulated user knows, has done and has seen, along with representations of what it considers attractive and relevant. Decisions to inspect or judge are then made based upon the simulated user's current state, rather than stochastically. In the context of ad-hoc topic retrieval, we evaluate the quality of the simulated users and agents by comparing their behaviour and performance against 48 human subjects under the same conditions, topics, time constraints, costs and search engine. Our findings show that while naive configurations of simulated users and agents substantially outperform our human subjects, their search behaviour is notably different from actual searchers. However, more sophisticated search agents can be tuned to act more like actual searchers providing greater realism. This innovation advances the state of the art in simulation, from simulated users towards autonomous agents. It provides a much needed step forward enabling the creation of more realistic simulations, while also motivating the development of more advanced cognitive agents and tools to help support and augment human searchers. Future work will focus not only on the pragmatics of tuning and training such agents for topic retrieval, but will also look at developing agents for other tasks and contexts such as collaborative search and slow search. David Maxwell 0001, Leif Azzopardi |
CIKM | 2 |
| 2016 | Two Scrolls or One Click: A Cost Model for Browsing Search Results
Leif Azzopardi, Guido Zuccon |
ECIR | 1 |
| 2016 | Simulation of Interaction: A Tutorial on Modelling and Simulating User Interaction and Search BehaviourabstractSearch is an inherently interactive, non-deterministic and user-dependent process. This means that there are many different possible sequences of interactions which could be taken (some ending in success and others ending in failure). Simulation provides a low cost, repeatable and reproducible way to explore a large range of different possibilities. This makes simulation very appealing, but it also requires care and consideration in developing, implementing and instantiating models of user behaviour for the purposes of experimentation. Leif Azzopardi |
SIGIR | 1 |
| 2016 | Simulating Interactive Information Retrieval: SimIIR: A Framework for the Simulation of InteractionabstractSimulation provides a powerful and cost-effective approach to explore and evaluate how interactions between a searcher and system influence search behaviour and performance. With a growing interest in simulation and an increasing number of papers using such an approach, there is a need for a flexible framework for simulation. Thus, we present SimIIR, an open-source toolkit for building and conducting Interactive Information Retrieval (IIR) experiments. The framework consists of a number of high level components, including the simulation, the searcher and the system, all of which must be configured. The SimIIR framework provides a series of interchangeable components. Examples of these components include the querying strategies (how simulated queries are formulated) and stopping strategies (the depth to which a searcher will examine snippets and documents) that a simulated searcher will employ. We have implemented various existing strategies so that they can be used by other researchers to not only replicate and reproduce past experiments, but also create new experiments. This paper describes the SimIIR framework and the different components that can be configured and extended as required. David Maxwell 0001, Leif Azzopardi |
SIGIR | 2 |
| 2016 | InfoScout: An Interactive, Entity Centric, Person Search ToolabstractIndividuals living in highly networked societies publish a large amount of personal, and potentially sensitive, information online. Web investigators can exploit such information for a variety of purposes, such as in background vetting and fraud detection. However, such investigations require a large number of expensive man hours and human effort. This paper describes InfoScout, a search tool which is intended to reduce the time it takes to identify and gather subject centric information on the Web. InfoScout collects relevance feedback information from the investigator in order to re-rank search results, allowing the intended information to be discovered more quickly. Users may still direct their search as they see fit, issuing ad-hoc queries and filtering existing results by keywords. Design choices are informed by prior work and industry collaboration. Sean McKeown, Martynas Buivys, Leif Azzopardi |
SIGIR | 3 |
| 2015 | ECol 2015: First international workshop on the Evaluation on Collaborative Information Seeking and RetrievalabstractCollaborative Information Seeking/Retrieval (CIS/CIR) has given rise to several challenges in terms of search behavior analysis, retrieval model formalization as well as interface design. However, the major issue of evaluation in CIS/CIR is still underexplored. The goal of this workshop is to investigate the evaluation challenges in CIS/CIR with the hope of building standardized evaluation frameworks, methodologies, and task specifications that would foster and grow the research area (in a collaborative fashion). Leif Azzopardi, Jeremy Pickens, Tetsuya Sakai, Laure Soulier, Lynda Tamine-Lechani |
CIKM | 1 |
| 2015 | Searching and Stopping: An Analysis of Stopping Rules and StrategiesabstractSearching naturally involves stopping points, both at a query level (how far down the ranked list should I go?) and at a session level (how many queries should I issue?). Understanding when searchers stop has been of much interest to the community because it is fundamental to how we evaluate search behaviour and performance. Research has shown that searchers find it difficult to formalise stopping criteria, and typically resort to their intuition of what is "good enough". While various heuristics and stopping criteria have been proposed, little work has investigated how well they perform, and whether searchers actually conform to any of these rules. In this paper, we undertake the first large scale study of stopping rules, investigating how they influence overall session performance, and which rules best match actual stopping behaviour. Our work is focused on stopping at the query level in the context of ad-hoc topic retrieval, where searchers undertake search tasks within a fixed time period. We show that stopping strategies based upon the disgust or frustration point rules - both of which capture a searcher's tolerance to non-relevance - typically result in (i) the best overall performance, and (ii) provide the closest approximation to actual searcher behaviour, although a fixed depth approach also performs remarkably well. Findings from this study have implications regarding how we build measures, and how we conduct simulations of search behaviours. David Maxwell 0001, Leif Azzopardi, Kalervo Järvelin, Heikki Keskustalo |
CIKM | 2 |
| 2015 | Query Length, Retrievability Bias and PerformanceabstractPast work has shown that longer queries tend to lead to better retrieval performance. However, this comes at the cost of increased user effort effort and additional system processing. In this paper, we examine whether there are benefits of longer queries beyond performance. We posit that increasing the query length will also lead to a reduction in the retrievability bias. Additionally, we speculate that to minimise retrievability bias as queries become longer, more length normalisation must be applied to account for the increase in the length of documents retrieved. To this end, we perform a retrievability analysis on two TREC collections using three standard retrieval models and various lengths of queries (one to five terms). From this investigation we find that increasing the length of queries reduces the overall retrievability bias but at a decreasing rate. Moreover, once the query length exceeds three terms the bias can begin to increase (and the performance can start to drop). We also observe that more document length normalisation is typically required as query length increases, in order to minimise bias. Finally, we show that there is a strong correlation between performance and retrieval bias. This work raises some interesting questions regarding query length and its affect on performance and bias. Further work will be directed towards examining longer and more verbose queries, including those generated via query expansion methods, to obtain a more comprehensive understanding of the relationship between query length, performance and retrievability bias. Colin Wilkie, Leif Azzopardi |
CIKM | 2 |
| 2015 | A Tutorial on Measuring Document Retrievability
Leif Azzopardi |
ECIR | 1 |
| 2015 | The Impact of Query Interface Design on Stress, Workload and Performance
Ashlee Edwards, Diane Kelly 0001, Leif Azzopardi |
ECIR | 3 |
| 2015 | Retrievability and Retrieval Bias: A Comparison of Inequality Measures
Colin Wilkie, Leif Azzopardi |
ECIR | 2 |
| 2015 | Building and Using Models of Information Seeking, Search and Retrieval: Full Day TutorialabstractUnderstanding how people interact with information systems when searching is central to the study of Interactive Information Retrieval (IIR). While much of the prior work in this area has either been conceptual, observational or empirical, recently there has been renewed interest in developing mathematical models of information seeking and search. This is because such models can provide a concise and compact representation of search behaviours and naturally generate testable hypotheses about search behaviour. This full day tutorial focuses on explaining and building formal models of Information Seeking and Retrieval. The tutorial is structured into four sessions. In the first session we will discuss the rationale of modelling and examine a number of early formal models of search (including early cost models and the Probability Ranking Principle). Then we will examine more contemporary formal models (including Information Foraging Theory, the Interactive Probability Ranking Principle, and Search Economic Theory). The focus will be on the insights and intuitions that we can glean from the math behind these models. The latter sessions will be dedicated to building models that optimise particular objectives which drive how users make decisions, along with a how-to guide on model building, where we will describe different techniques (including analytical, graphical and computational) that can be used to generate hypotheses from such models. In the final session, participants will be challenged to develop a simple model of interaction applying the techniques learnt during the day, before concluding with an overview of challenges and future directions. Leif Azzopardi, Guido Zuccon |
SIGIR | 1 |
| 2015 | Time Pressure and System Delays in Information SearchabstractWe report preliminary results of the impact of time pressure and system delays on search behavior from a laboratory study with forty-three participants. To induce time pressure, we randomly assigned half of our study participants to a treatment condition where they were only allowed five minutes to search for each of four ad-hoc search topics. The other half of the participants were given no task time limits. For half of participants' search tasks (n=2), five second delays were introduced after queries were submitted and SERP results were clicked. Results showed that participants in the time pressure condition queried at a significantly higher rate, viewed significantly fewer documents per query, had significantly shallower hover and view depths, and spent significantly less time examining documents and SERPs. We found few significant differences in search behavior for system delay or interaction effects between time pressure and system delay. These initial results show time pressure has a significant impact on search behavior and suggest the design of search interfaces and features that support people who are searching under time pressure. Anita Crescenzi, Diane Kelly 0001, Leif Azzopardi |
SIGIR | 3 |
| 2015 | Untangling Result List Refinement and Ranking Quality: a Framework for Evaluation and PredictionabstractTraditional batch evaluation metrics assume that user interaction with search results is limited to scanning down a ranked list. However, modern search interfaces come with additional elements supporting result list refinement (RLR) through facets and filters, making user search behavior increasingly dynamic. We develop an evaluation framework that takes a step beyond the interaction assumption of traditional evaluation metrics and allows for batch evaluation of systems with and without RLR elements. In our framework we model user interaction as switching between different sublists. This provides a measure of user effort based on the joint effect of user interaction with RLR elements and result quality. We validate our framework by conducting a user study and comparing model predictions with real user performance. Our model predictions show significant positive correlation with real user effort. Further, in contrast to traditional evaluation metrics, the predictions using our framework, of when users stand to benefit from RLR elements, reflect findings from our user study. Jiyin He, Marc Bron, Arjen P. de Vries, Leif Azzopardi, Maarten de Rijke |
SIGIR | 4 |
| 2015 | How many results per page?: A Study of SERP Size, Search Behavior and User ExperienceabstractThe provision of "ten blue links" has emerged as the standard for the design of search engine result pages (SERPs). While numerous aspects of SERPs have been examined, little attention has been paid to the number of results displayed per page. This paper investigates the relationships among the number of results shown on a SERP, search behavior and user experience. We performed a laboratory experiment with 36 subjects, who were randomly assigned to use one of three search interfaces that varied according to the number of results per SERP (three, six or ten). We found subjects' click distributions differed significantly depending on SERP size. We also found those who interacted with three results per page viewed significantly more SERPs per query; interestingly, the number of SERPs they viewed per query corresponded to about 10 search results. Subjects who interacted with ten results per page viewed and saved significantly more documents. They also reported the greatest difficulty finding relevant documents, rated their skills the lowest and reported greater workload, even though these differences were not significant. This work shows that behavior changes with SERP size, such that more time is spent focused on earlier results when SERP size decreases. Diane Kelly 0001, Leif Azzopardi |
SIGIR | 2 |
| 2015 | An Initial Investigation into Fixed and Adaptive Stopping StrategiesabstractMost models, measures and simulations often assume that a searcher will stop at a predetermined place in a ranked list of results. However, during the course of a search session, real-world searchers will vary and adapt their interactions with a ranked list. These interactions depend upon a variety of factors, including the content and quality of the results returned, and the searcher's information need. In this paper, we perform a preliminary simulated analysis into the influence of stopping strategies when query quality varies. Placed in the context of ad-hoc topic retrieval during a multi-query search session, we examine the influence of fixed and adaptive stopping strategies on overall performance. Surprisingly, we find that a fixed strategy can perform as well as the examined adaptive strategies, but the fixed depth needs to be adjusted depending on the querying strategy used. Further work is required to explore how well the stopping strategies reflect actual search behaviour, and to determine whether one stopping strategy is dominant. David Maxwell 0001, Leif Azzopardi, Kalervo Järvelin, Heikki Keskustalo |
SIGIR | 2 |
| 2014 | A Retrievability Analysis: Exploring the Relationship Between Retrieval Bias and Retrieval PerformanceabstractRetrievability provides an alternative way to assess an Information Retrieval (IR) system by measuring how easily documents can be retrieved. Retrievability can also be used to determine the level of retrieval bias a system exerts upon a collection of documents. It has been hypothesised that reducing the retrieval bias will lead to improved performance. To date, it has been shown that this hypothesis does not appear to hold on standard retrieval performance measures (MAP and [email protected]) when exploring the parameter space of a given retrieval model. However, the evidence is limited and confined to only a few models, collections and measures. In this paper, we perform a comprehensive empirical evaluation analysing the relationship between retrieval bias and retrieval performance using several well known retrieval models, five large TREC test collections and ten performance measures (including the recently proposed PRES, Time Biased Gain (TBG) and U-Measure). For traditional relevance based measures (MAP, [email protected], MRR, Recall, etc) the correlation between retrieval bias and performance is moderate. However, for TBG and U-Measure, we find that there is strong and significant negative correlations between retrieval bias and performance (i.e as bias drops, performance increases). These findings suggest that for these more sophisticated, user oriented measures the retrievability bias hypothesis tends to hold. The implication is that for these measures, systems can then be tuned using retrieval bias, without recourse to relevance judgements. Colin Wilkie, Leif Azzopardi |
CIKM | 2 |
| 2014 | Page Retrievability Calculator
Leif Azzopardi, Rosanne English, Colin Wilkie, David Maxwell 0001 |
ECIR | 1 |
| 2014 | Best and Fairest: An Empirical Analysis of Retrieval System Bias
Colin Wilkie, Leif Azzopardi |
ECIR | 2 |
| 2014 | Efficiently Estimating Retrievability Bias
Colin Wilkie, Leif Azzopardi |
ECIR | 2 |
| 2014 | Modelling interaction with economic models of searchabstractUnderstanding how people interact when searching is central to the study of Interactive Information Retrieval (IIR). Most of the prior work has either been conceptual, observational or empirical. While this has led to numerous insights and findings regarding the interaction between users and systems, the theory has lagged behind. In this paper, we extend the recently proposed search economic theory to make the model more realistic. We then derive eight interaction based hypotheses regarding search behaviour. To validate the model, we explore whether the search behaviour of thirty-six participants from a lab based study is consistent with the theory. Our analysis shows that observed search behaviours are in line with predicted search behaviours and that it is possible to provide credible explanations for such behaviours. This work describes a concise and compact representation of search behaviour providing a strong theoretical basis for future IIR research. Leif Azzopardi |
SIGIR | 1 |
| 2014 | The retrievability of documentsabstractRetrievability is an important and interesting indicator that can be used in a number of ways to analyse Information Retrieval systems and document collections. Rather than focusing totally on relevance, retrievability examines what is retrieved, how often it is retrieved, and whether a user is likely to retrieve a document or not. This is important because a document needs to be retrieved, before it can be judged for relevance. In this tutorial, we shall explain the concept of retrievability along with a number of retrievability measures, how it can be estimated and how it can be used for analysis. Since retrieval precedes relevance, we shall also provide an overview of how retrievability relates to effectiveness - describing some of the insights that researchers have discovered so far. We shall also show how retrievability relates to efficiency, and how the theory of retrievability can be used to improve both effectiveness and efficiency. Then we shall provide an overview of the different applications of retrievability such as Search Engine Bias, Corpus Profiling, etc., before wrapping up with challenges and opportunities. The final session of the day will look at example problems and ways to analyse and apply retrievability to other problems and domains. Leif Azzopardi |
SIGIR | 1 |
| 2013 | High throughput filtering using FPGA-accelerationabstractWith the rise in the amount information of being streamed across networks, there is a growing demand to vet the quality, type and content itself for various purposes such as spam, security and search. In this paper, we develop an energy-efficient high performance information filtering system that is capable of classifying a stream of incoming document at high speed. The prototype parses a stream of documents using a multicore CPU and then performs classification using Field-Programmable Gate Arrays (FPGAs). On a large TREC data collection, we implemented a Naive Bayes classifier on our prototype and compared it to an optimized CPU based-baseline. Our empirical findings show that we can classify documents at 10Gb/s which is up to 94 times faster than the CPU baseline (and up to 5 times faster than previous FPGA based implementations). In future work, we aim to increase the throughput by another order of magnitude by implementing both the parser and filter on the FPGA. Wim Vanderbauwhede, Anton Frolov 0002, Leif Azzopardi, Sai Rahul Chalamalasetti, Martin Margala |
CIKM | 3 |
| 2013 | Re-leashed! The PuppyIR Framework for Developing Information Services for Children, Adults and Dogs
Douglas Dowie, Leif Azzopardi |
ECIR | 2 |
| 2013 | Reducing the Uncertainty in Resource Selection
Ilya Markov, Leif Azzopardi, Fabio Crestani |
ECIR | 2 |
| 2013 | An Initial Investigation on the Relationship between Usage and Findability
Colin Wilkie, Leif Azzopardi |
ECIR | 2 |
| 2013 | How query cost affects search behaviorabstractaffects how users interact with a search system. Microeconomic theory is used to generate the cost-interaction hypothesis that states as the cost of querying increases, users will pose fewer queries and examine more documents per query. A between-subjects laboratory study with 36 undergraduate subjects was conducted, where subjects were randomly assigned to use one of three search interfaces that varied according to the amount of physical cost required to query: Structured (high cost), Standard (medium cost) and Query Suggestion (low cost). Results show that subjects who used the Structured interface submitted significantly fewer queries, spent more time on search results pages, examined significantly more documents per query, and went to greater depths in the search results list. Results also showed that these subjects spent longer generating their initial queries, saved more relevant documents and rated their queries as more successful. These findings have implications for the usefulness of microeconomic theory as a way to model and explain search interaction, as well as for the design of query facilities. Leif Azzopardi, Diane Kelly 0001, Kathy Brennan |
SIGIR | 1 |
| 2013 | #trapped!: social media search system requirements for emergency management professionalsabstractSocial media provides a new and potentially rich source of information for emergency management services. However, extracting the relevant information from such streams poses a number of difficult challenges. In this short paper, we survey emergency management professionals to ascertain how social media is used when responding to incidents, the search strategies that they undertake, and the challenges that they face when using social media streams. This research indicates that emergency management professionals employ two main strategies when searching social media streams: keyword-centric and account-centric search strategies. Furthermore, current search interfaces are inadequate regarding the requirements of command and control environments in the emergency management domain, where the process of information seeking is collaborative in nature and needs to support multiple information seekers. Stefan Raue, Leif Azzopardi, Christopher W. Johnson 0001 |
SIGIR | 2 |
| 2013 | Relating retrievability, performance and lengthabstractRetrievability provides a different way to evaluate an Information Retrieval (IR) system as it focuses on how easily documents can be found. It is intrinsically related to retrieval performance because a document needs to be retrieved before it can be judged relevant. In this paper, we undertake an empirical investigation into the relationship between the retrievability of documents, the retrieval bias imposed by a retrieval system, and the retrieval performance, across different amounts of document length normalization. To this end, two standard IR models are used on three TREC test collections to show that there is a useful and practical link between retrievability and performance. Our findings show that minimizing the bias across the document collection leads to good performance (though not the best performance possible). We also show that past a certain amount of document length normalization the retrieval bias increases, and the retrieval performance significantly and rapidly decreases. These findings suggest that the relationship between retrievability and effectiveness may offer a way to automatically tune systems. Colin Wilkie, Leif Azzopardi |
SIGIR | 2 |
| 2013 | Crowdsourcing interactions: using crowdsourcing for evaluating interactive information retrieval systems
Guido Zuccon, Teerapong Leelanupab, Stewart Whiting, Emine Yilmaz, Joemon M. Jose, Leif Azzopardi |
Inf. Retr. | 6 |
| 2012 | EmSe: Supporting Children's Information Needs within a Hospital Environment
Leif Azzopardi, Douglas Dowie, Sergio Duarte Torres, Carsten Eickhoff, Richard Glassey, Karl Gyllstrom, Djoerd Hiemstra, Franciska de Jong, Frea Kruisinga, Kelly Ann Marshall, Marie-Francine Moens, Tamara Polajnar, Frans van der Sluis, Arjen P. de Vries |
ECIR | 1 |
| 2012 | Crisees: Real-Time Monitoring of Social Media Streams to Support Crisis Management
David Maxwell 0001, Stefan Raue, Leif Azzopardi, Christopher W. Johnson 0001, Sarah Oates |
ECIR | 3 |
| 2012 | Detection of News Feeds Items Appropriate for Children
Tamara Polajnar, Richard Glassey, Leif Azzopardi |
ECIR | 3 |
| 2012 | Top-k Retrieval Using Facility Location Analysis
Guido Zuccon, Leif Azzopardi, Dell Zhang, Jun Wang 0012 |
ECIR | 2 |
| 2012 | ALF: a client side logger and server for capturing user interactions in web applicationsabstractThis demonstration paper introduces ALF which provides a light-weight client side logging application and a server for collecting user interaction data. ALF has been designed as a loosely coupled independent service that runs in parallel with the IR web application that requires logging Leif Azzopardi, Myles Doolan, Richard Glassey |
SIGIR | 1 |
| 2012 | YooSee: a video browsing application for young childrenabstractNowadays children as young as two years old can easily interact with mobile touch screen devices and personal computers to watch online videos through services such as YouTube. However, such services present a number of challenges for young children (e.g. fine grain gestures/interactions and good typing/literacy skills). In addition, when children use such services there is a risk that they may stumble upon content that is inappropriate. YooSee is a web-based application developed using the PuppyIR framework and designed for children aged between two and six years old. YooSee enables children to: (1) search and browse through video content using an engaging, novel interaction paradigm, and (2) be able to safely enjoy moderated video content. Leif Azzopardi, Douglas Dowie, Kelly Ann Marshall |
SIGIR | 1 |
| 2012 | MaSe: create your own mash-up search interfaceabstractMaSe provides a sandbox environment for high school students to create their own personalised search interface. It has been designed with two major goals in mind: (1) as a hands-on tutorial for school children, to excite them about programming and computing science through the development of a practical application, and (2) to enable children to design and create their own search interface without extensive programming knowledge or prior experience. Consequently, MaSe provides a way to ascertain what children would like from a search engine interface in an exploratory and creative way as they can create a working prototype. This approach contrasts with previous work on exploring children's requirements of IR systems which attempts to directly elicit user needs through more traditional methods (i.e. surveys, interviews, focus groups, etc). However, we have attempted to incorporate the design guidelines for children as identified by Large (2006) into MaSe, where: we make use of bright colours, large text fonts, spell checking and the use of icons to represent search services, as well as including a thematic experience as suggested by Large (2006), with the use of a puppy avatar and puppy dog footprints. Leif Azzopardi, Douglas Dowie, Kelly Ann Marshall, Richard Glassey |
SIGIR | 1 |
| 2012 | PageFetch: a retrieval game for children (and adults)abstractChildren often struggle with information retrieval tasks as searching for information often requires a developed vocabulary and strong categorisation skills; neither of which are particularly developed in children under the age of 12. In a study conducted by Druin et al, it was found that in an experimental setting many children are often uninterested in searching for information online or are only interested in searching for information that is relevant to their personal interests. Consequently, children who were unmotivated were the least successful in completing information retrieval tasks in their study. It was suggested that a more effective means of engaging child participants in search studies must be developed in order to gain further insights into the searching behaviours of children. To this end we have developed a game called PageFetch which aims to engage children (aged 8 to 80) in completing search tasks through a fun and interactive search-like interface. Leif Azzopardi, James Purvis, Richard Glassey |
SIGIR | 1 |
| 2011 | Fu-Finder: a game for studying querying behavioursabstractUsually the focus of evaluation within Information Retrieval has been placed largely upon the system. However, the individual user and their submitted queries are typically the greatest source of variation in the search process. This demonstration paper presents Fu-Finder, a fun and enjoyable game that measures the user's querying abilities (or search-fu). This game provides useful data for the study of user querying behaviour and assesses how well users can find specific web pages using different search engines. Carly O'Neil, James Purvis, Leif Azzopardi |
CIKM | 3 |
| 2011 | Back to the Roots: Mean-Variance Analysis of Relevance Estimations
Guido Zuccon, Leif Azzopardi, C. J. van Rijsbergen |
ECIR | 2 |
| 2011 | The economics in interactive information retrievalabstractSearching is inherently an interactive process usually requiring numerous iterations of querying and assessing in order to find the desired amount of relevant information. Essentially, the search process can be viewed as a combination of inputs (queries and assessments) which are used to "produce" output (relevance). Under this view, it is possible to adapt microeconomic theory to analyze and understand the dynamics of Interactive Information Retrieval. In this paper, we approach the search process as an economics problem and conduct extensive simulations on TREC test collections analyzing various combinations of inputs in the "production" of relevance. The analysis reveals that the total Cumulative Gain (output) obtained during the course of a search session is functionally related to querying and assessing (inputs), and this can be characterized mathematically by the Cobbs-Douglas production function. Further analysis using cost models, that are grounded using cognitive load as the cost, reveals which search strategies minimize the cost of interaction for a given level of output. This paper demonstrates how economics can be applied to formally model the search process. This development establishes the theoretical foundations of Interactive Information Retrieval, providing numerous directions for empirical experimentation that are motivated directly from theory. Leif Azzopardi |
SIGIR | 1 |
| 2011 | JuSe: a picture dictionary query system for childrenabstractAs adults we take for granted our capacity to express our information needs verbally and textually. However, young children also have preferences and information needs, but are just learning to be able to express themselves effectively. Consequently they encounter many barriers when trying to spell, type, and communicate their needs to a 'faceless' search engine text box. Tamara Polajnar, Richard Glassey, Leif Azzopardi |
SIGIR | 3 |
| 2011 | The interactive PRP for diversifying document rankingsabstractThe assumptions underlying the Probability Ranking Principle (PRP) have led to a number of alternative approaches that cater or compensate for the PRP's limitations. In this poster we focus on the Interactive PRP (iPRP), which rejects the assumption of independence between documents made by the PRP. Although the theoretical framework of the iPRP is appealing, no instantiation has been proposed and investigated. In this poster, we propose a possible instantiation of the principle, performing the first empirical comparison of the iPRP against the PRP. For document diversification, our results show that the iPRP is significantly better than the PRP, and comparable to or better than other methods such as Modern Portfolio Theory. Guido Zuccon, Leif Azzopardi, C. J. van Rijsbergen |
SIGIR | 2 |
| 2011 | Introduction to special issue on the second international conference on the theory of information retrieval
Leif Azzopardi, Dawei Song 0001, Gabriella Kazai, Stephen E. Robertson, Stefan M. Rüger, Milad Shokouhi, Emine Yilmaz |
Inf. Retr. | 1 |
| 2011 | Extending the language modeling framework for sentence retrieval to include local context
Ronald T. Fernández, David E. Losada, Leif Azzopardi |
Inf. Retr. | 3 |
| 2010 | A comparison of user and system query performance predictionsabstractQuery performance prediction methods are usually applied to estimate the retrieval effectiveness of queries, where the evaluation is largely system sided. However, little work has been conducted to understand query performance prediction from the user's perspective. The question we consider is, whether the predictions of query performance that systems make are in line with the predictions that users make. To this aim, we compare the performance ratings users assign to queries with the performance scores estimated by a range of pre-retrieval and post-retrieval query performance predictors. Two studies are presented that explore the relationship between user ratings and system predictions on two levels: (i) the topic level, and, (ii) the query suggestions level. It is shown that when predicting the performance of query suggestions, user ratings were mostly uncorrelated with system predictions. At the topic level though, where a single query is judged for each information need, we observed moderate correlations between user ratings and a subset of system predictions. As query performance prediction methods are often based on intuitions of how users might rate queries, these findings suggest that such methods are not representative of how users actually rate query suggestions and topics. This motivates further research into understanding the rating process engaged by users, and developing models of query performance prediction in order to bridge the divide between systems and users. Claudia Hauff, Diane Kelly 0001, Leif Azzopardi |
CIKM | 3 |
| 2010 | Query Performance Prediction: Evaluation Contrasted with Effectiveness
Claudia Hauff, Leif Azzopardi, Djoerd Hiemstra, Franciska de Jong |
ECIR | 2 |
| 2010 | A Case for Automatic System Evaluation
Claudia Hauff, Djoerd Hiemstra, Leif Azzopardi, Franciska de Jong |
ECIR | 3 |
| 2010 | Using the Quantum Probability Ranking Principle to Rank Interdependent Documents
Guido Zuccon, Leif Azzopardi |
ECIR | 2 |
| 2010 | On the relationship between effectiveness and accessibilityabstractTypically the evaluation of Information Retrieval (IR) systems is focused upon two main system attributes: efficiency and effectiveness. However, it has been argued that it is also important to consider accessibility, i.e. the extent to which the IR system makes information easily accessible. But, it is unclear how accessibility relates to typical IR evaluation, and specifically whether there is a trade-off between accessibility and effectiveness. In this poster, we empirically explore the relationship between effectiveness and accessibility to determine whether the two objectives i.e. maximizing effectiveness and maximizing accessibility, are compatible, or not. To this aim, we empirically examine this relationship using two popular IR models and explore the trade-off between access and performance as these models are tuned. Leif Azzopardi, Richard Bache |
SIGIR | 1 |
| 2010 | Improving sentence retrieval with an importance priorabstractThe retrieval of sentences is a core task within Information Retrieval. In this poster we employ a Language Model that incorporates a prior which encodes the importance of sentences within the retrieval model. Then, in a set of comprehensive experiments using the TREC Novelty Tracks, we show that including this prior substantially improves retrieval effectiveness, and significantly outperforms the current state of the art in sentence retrieval. Leif Azzopardi, Ronald T. Fernández, David E. Losada |
SIGIR | 1 |
| 2010 | Search system requirements of patent analystsabstractPatent search tasks are difficult and challenging, often requiring expert patent analysts to spend hours, even days, sourcing relevant information. To aid them in this process, analysts use Information Retrieval systems and tools to cope with their retrieval tasks. With the growing interest in patent search, it is important to determine their requirements and expectations of the tools and systems that they employ. In this poster, we report a subset of the findings of a survey of patent analysts conducted to elicit their search requirements. Leif Azzopardi, Wim Vanderbauwhede, Hideo Joho |
SIGIR | 1 |
| 2010 | Finding and filtering information for childrenabstractChildren face several challenges when using information access systems. These include formulating queries, judging the relevance of documents, and focusing attention on interface cues, such as query suggestions, while typing queries. It has also been shown that children want a personalised Web experience and prefer content presented to them that matches their long-term entertainment and education needs. To this end, we have developed an interaction-based information filtering system to address these challenges. Desmond Elliott, Richard Glassey, Tamara Polajnar, Leif Azzopardi |
SIGIR | 4 |
| 2010 | Query quality: user ratings and system predictionsabstractNumerous studies have examined the ability of query performance prediction methods to estimate a query's quality for system effectiveness measures (such as average precision). However, little work has explored the relationship between these methods and user ratings of query quality. In this poster, we report the findings from an empirical study conducted on the TREC ClueWeb09 corpus, where we compared and contrasted user ratings of query quality against a range of query performance prediction methods. Given a set of queries, it is shown that user ratings of query quality correlate to both system effectiveness measures and a number of pre-retrieval predictors. Claudia Hauff, Franciska de Jong, Diane Kelly 0001, Leif Azzopardi |
SIGIR | 4 |
| 2010 | Estimating interference in the QPRP for subtopic retrievalabstractThe Quantum Probability Ranking Principle (QPRP) has been recently proposed, and accounts for interdependent document relevance when ranking. However, to be instantiated, the QPRP requires a method to approximate the "interference" between two documents. In this poster, we empirically evaluate a number of different methods of approximation on two TREC test collections for subtopic retrieval. It is shown that these approximations can lead to significantly better retrieval performance over the state of the art. Guido Zuccon, Leif Azzopardi, Claudia Hauff, C. J. van Rijsbergen |
SIGIR | 2 |
| 2010 | Has portfolio theory got any principles?abstractRecently, Portfolio Theory (PT) has been proposed for Information Retrieval. However, under non-trivial conditions PT violates the original Probability Ranking Principle (PRP). In this poster, we shall explore whether PT upholds a different ranking principle based on Quantum Theory, i.e. the Quantum Probability Ranking Principle (QPRP), and examine the relationship between this new model and the new ranking principle. We make a significant contribution to the theoretical development of PT and show that under certain circumstances PT upholds the QPRP, and thus guarantees an optimal ranking according to the QPRP. A practical implication of this finding is that the parameters of PT can be automatically estimated via the QPRP, instead of resorting to extensive parameter tuning. Guido Zuccon, Leif Azzopardi, C. J. van Rijsbergen |
SIGIR | 2 |
| 2009 | Usage based effectiveness measures: monitoring application performance in information retrievalabstractThe aim of an Information Retrieval (IR) application is to support the user accessing relevant information effectively and efficiently. It is well known that system performance, in terms of finding relevant information is heavily dependent upon the IR application (i.e. the IR system exposed through the application's interface), as well as how the application is used by the user (i.e. how the user interacts with the system through the interface). Thus, a very pragmatic evaluation question that arises at the application level is: what is the effectiveness experienced by the user during the usage of the application? To be able to answer this question, we represent the usage of an application by the stream of documents the user encounters while interacting with the application. This representation enables us to monitor and track the performance over time and usage. By taking a stream-based, time-centric view of the IR process, instead of a rank-list, topic/task centric view, the evaluation can be performed on any IR based application. To illustrate the difference and the utility of this approach, we demonstrate how a new suite of usage based effectiveness measures can be applied. This work provides the conceptual foundations for measuring, monitoring and modeling the performance of any IR application which needs to be evaluated over time and in context. Leif Azzopardi |
CIKM | 1 |
| 2009 | Relying on topic subsets for system ranking estimationabstractRanking a number of retrieval systems according to their retrieval effectiveness without relying on costly relevance judgments was first explored by Soboroff et al [6]. Over the years, a number of alternative approaches have been proposed. We perform a comprehensive analysis of system ranking estimation approaches on a wide variety of TREC test collections and topics sets. Our analysis reveals that the performance of such approaches is highly dependent upon the topic or topic subset, used for estimation. We hypothesize that the performance of system ranking estimation approaches can be improved by selecting the "right" subset of topics and show that using topic subsets improves the performance by 32% on average, with a maximum improvement of up to 70% in some cases. Claudia Hauff, Djoerd Hiemstra, Franciska de Jong, Leif Azzopardi |
CIKM | 4 |
| 2009 | The Combination and Evaluation of Query Performance Prediction Methods
Claudia Hauff, Leif Azzopardi, Djoerd Hiemstra |
ECIR | 2 |
| 2009 | Query side evaluation: an empirical analysis of effectiveness and effortabstractTypically, Information Retrieval evaluation focuses on measuring the performance of the system's ability at retrieving relevant information, and not the query's ability. However, the effectiveness of a retrieval system is strongly influenced by the quality of the query submitted. In this paper, the effectiveness and effort of querying is empirically examined in the context of the Principle of Least Effort, Zipf's Law and the Law of Diminishing Returns. This query focused investigation leads to a number of novel findings which should prove useful in the development of future retrieval methods and evaluation techniques. While, also motivating further research into query side evaluation. Leif Azzopardi |
SIGIR | 1 |
| 2009 | Search engine predilection towards news media providersabstractIn this poster paper, we present a preliminary study on the predilection of web search engines towards various online news media provider sites using an access based measure. Leif Azzopardi, Ciaran Owens |
SIGIR | 1 |
| 2009 | Developing energy efficient filtering systemsabstractProcessing large volumes of information generally requires massive amounts of computational power, which consumes a significant amount of energy. An emerging challenge is the development of ``environmentally friendly'' systems that are not only efficient in terms of time, but also energy efficient. In this poster, we outline our initial efforts at developing greener filtering systems by employing Field Programmable Gate Arrays (FPGA) to perform the core information processing task. FPGAs enable code to be executed in parallel at a chip level, while consuming only a fraction of the power of a standard (von Neuman style) processor. On a number of test collections, we demonstrate that the FPGA filtering system performs 10-20 times faster than the Itanium based implementation, resulting in considerable energy savings. Leif Azzopardi, Wim Vanderbauwhede, Mahmoud Moadeli |
SIGIR | 1 |
| 2009 | When is query performance prediction effective?abstractThe utility of Query Performance Prediction (QPP) methods is commonly evaluated by reporting correlation coefficients to denote how well the methods perform at predicting the retrieval performance of a set of queries. However, a quintessential question remains unexplored: how strong does the correlation need to be in order to realize an increase in retrieval performance? In this work, we address this question in the context of Selective Query Expansion (SQE) and perform a large-scale experiment. The results show that to consistently and predictably improve retrieval effectiveness in the ideal SQE setting, a Kendall's Tau correlation of tau>=0.5 is required, a threshold which most existing query performance prediction methods fail to reach. Claudia Hauff, Leif Azzopardi |
SIGIR | 2 |
| 2009 | Effective query expansion for federated searchabstractWhile query expansion techniques have been shown to improve retrieval performance in a centralized setting, they have not been well studied in a federated setting. In this paper, we consider how query expansion may be adapted to federated environments and propose several new methods: where focused expansions are used in a selective fashion to produce specific queries for each source (or a set of sources). On a number of different testbeds, we show that focused query expansion can significantly outperform the previously proposed global expansion method, and---contrary to earlier work---show that query expansion can improve performance over standard federated retrieval. \n \n These findings motivate further research examining the different methods for query expansion, and other forms of system and user interaction, in order to continue improving the performance of interactive federated search systems. Milad Shokouhi, Leif Azzopardi, Paul Thomas 0001 |
SIGIR | 2 |
| 2009 | Revisiting logical imaging for information retrievalabstractRetrieval with Logical Imaging is derived from belief revision and provides a novel mechanism for estimating the relevance of a document through logical implication (i.e. P(q->d). In this poster, we perform the first comprehensive evaluation of Logical Imaging (LI) in Information Retrieval (IR) across several TREC test Collections. When compared against standard baseline models, we show that LI fails to improve performance. This failure can be attributed to a nuance within the model that means non-relevant documents are promoted in the ranking, while relevant documents are demoted. This is an important contribution because it not only contextualizes the effectiveness of LI, but crucially explains why it fails. By addressing this nuance, future LI models could be significantly improved. Guido Zuccon, Leif Azzopardi, C. J. van Rijsbergen |
SIGIR | 2 |
| 2009 | A language modeling framework for expert finding
Krisztian Balog, Leif Azzopardi, Maarten de Rijke |
Inf. Process. Manag. | 2 |
| 2008 | Retrievability: an evaluation measure for higher order information access tasksabstractEvaluation in Information Retrieval (IR) has long focused on effectiveness and efficiency. However, new and emerging access tasks now demand alternative evaluation measures which go beyond this traditional view. A retrieval system provides a means of gaining access to documents, therefore intuitively, our view of the collection is shaped by the retrieval system. In this paper, we outline some emerging information access related scenarios that require knowledge about how the retrieval system affects the users' ability to access information. This provides the motivation for the proposed evaluation measures and methodology where the focus is on capturing the behavior of the system, in terms of how retrievable it makes individual documents within the collection. To demonstrate the utility of the proposed methods, we perform an extensive analysis on two TREC collections showing how the measures can be applied to evaluate different information access questions. For higher order information access tasks that are inherently dependent on retrievability, our novel evaluation methodology emphasizes that effectiveness is an insufficient characterization of a retrieval system. This paper provides the foundations for the evaluation of higher order access related tasks. Leif Azzopardi, Vishwa Vinay |
CIKM | 1 |
| 2008 | Revisiting the relationship between document length and relevanceabstractThe scope hypothesis in Information Retrieval (IR) states that a relationship exists between document length and relevance, such that the likelihood of relevance increases with document length. A number of empirical studies have provided statistical evidence supporting the scope hypothesis. However, these studies make the implicit assumption that modern test collections are complete (i.e. all documents are assessed for relevance). As a consequence the observed evidence is misleading. In this paper we perform a deeper analysis of document length and relevance taking into account that test collections are incomplete. We first demonstrate that previous evidence supporting the scope hypothesis was an artefact of the test collection, where there is a bias towards longer documents in the pooling process. We evaluate whether this length bias affects system comparison when using incomplete test collections. The results indicate that test collections are problematic when considering MAP as a measure of effectiveness but are relatively robust when using bpref. The implications of the study indicate that retrieval models should not be tuned to favour longer documents, and that designers of new test collections should take measures against length bias during the pooling process in order to create more reliable and robust test collections. David E. Losada, Leif Azzopardi, Mark Baillie |
CIKM | 2 |
| 2008 | Accessibility in Information Retrieval
Leif Azzopardi, Vishwa Vinay |
ECIR | 1 |
| 2008 | Evaluating epistemic uncertainty under incomplete assessments
Mark Baillie, Leif Azzopardi, Ian Ruthven |
Inf. Process. Manag. | 2 |
| 2008 | Contextual factors affecting the utility of surrogates within exploratory search
Ian Ruthven, Mark Baillie, Leif Azzopardi, Ralf Bierig, Emma Nicol, Simon O. Sweeney, Murat Yaciki |
Inf. Process. Manag. | 3 |
| 2008 | An analysis on document length retrieval trends in language modeling smoothing
David E. Losada, Leif Azzopardi |
Inf. Retr. | 2 |
| 2008 | Assessing multivariate Bernoulli models for information retrievalabstractAlthough the seminal proposal to introduce language modeling in information retrieval was based on a multivariate Bernoulli model, the predominant modeling approach is now centered on multinomial models. Language modeling for retrieval based on multivariate Bernoulli distributions is seen inefficient and believed less effective than the multinomial model. In this article, we examine the multivariate Bernoulli model with respect to its successor and examine its role in future retrieval systems. In the context of Bayesian learning, these two modeling approaches are described, contrasted, and compared both theoretically and computationally. We show that the query likelihood following a multivariate Bernoulli distribution introduces interesting retrieval features which may be useful for specific retrieval tasks such as sentence retrieval. Then, we address the efficiency aspect and show that algorithms can be designed to perform retrieval efficiently for multivariate Bernoulli models, before performing an empirical comparison to study the behaviorial aspects of the models. A series of comparisons is then conducted on a number of test collections and retrieval tasks to determine the empirical and practical differences between the different models. Our results indicate that for sentence retrieval the multivariate Bernoulli model can significantly outperform the multinomial model. However, for the other tasks the multinomial model provides consistently better performance (and in most cases significantly so). An analysis of the various retrieval characteristics reveals that the multivariate Bernoulli model tends to promote long documents whose nonquery terms are informative. While this is detrimental to the task of document retrieval (documents tend to contain considerable nonquery content), it is valuable for other tasks such as sentence retrieval, where the retrieved elements are very short and focused. David E. Losada, Leif Azzopardi |
ACM Trans. Inf. Syst. | 2 |
| 2007 | A Retrieval Evaluation Methodology for Incomplete Relevance Assessments
Mark Baillie, Leif Azzopardi, Ian Ruthven |
ECIR | 2 |
| 2007 | Building simulated queries for known-item topics: an analysis using six european languagesabstractThere has been increased interest in the use of simulated queries for evaluation and estimation purposes in Information Retrieval. However, there are still many unaddressed issues regarding their usage and impact on evaluation because their quality, in terms of retrieval performance, is unlike real queries. In this paper, wefocus on methods for building simulated known-item topics and explore their quality against real known-item topics. Using existing generation models as our starting point, we explore factors which may influence the generation of the known-item topic. Informed by this detailed analysis (on six European languages) we propose a model with improved document and term selection properties, showing that simulated known-item topics can be generated that are comparable to real known-item topics. This is a significant step towards validating the potential usefulness of simulated queries: for evaluation purposes, and becausebuilding models of querying behavior provides a deeper insight into the querying process so that better retrieval mechanisms can be developed to support the user. Leif Azzopardi, Maarten de Rijke, Krisztian Balog |
SIGIR | 1 |
| 2007 | Broad expertise retrieval in sparse data environmentsabstractExpertise retrieval has been largely unexplored on data other than the W3C collection. At the same time, many intranets of universities and other knowledge-intensive organisations offer examples of relatively small but clean multilingual expertise data, covering broad ranges of expertise areas. We first present two main expertise retrieval tasks, along with a set of baseline approaches based on generative language modeling, aimed at finding expertise relations between topics and people. For our experimental evaluation, we introduce (and release) a new test set based on a crawl of a university site. Using this test set, we conduct two series of experiments. The first is aimed at determining the effectiveness of baseline expertise retrieval methods applied to the new test set. The second is aimed at assessing refined models that exploit characteristic features of the new test set, such as the organizational structure of the university, and the hierarchical structure of the topics in the test set. Expertise retrieval models are shown to be robust with respect to environments smaller than the W3C collection, and current techniques appear to be generalizable to other settings. Krisztian Balog, Toine Bogers, Leif Azzopardi, Maarten de Rijke, Antal van den Bosch |
SIGIR | 3 |
| 2007 | Intra-assessor consistency in question answeringabstractIn this paper we investigate the consistency of answer assessment in a complex question answering task examining features of assessor consistency, types of answers and question type. Ian Ruthven, Leif Azzopardi, Mark Baillie, Ralf Bierig, Emma Nicol, Simon O. Sweeney, Murat Yakici |
SIGIR | 2 |
| 2007 | Updating collection representations for federated searchabstractTo facilitate the search for relevant information across a setof online distributed collections, a federated information retrieval system typically represents each collection, centrally, by a set of vocabularies or sampled documents. Accurate retrieval is therefore related to how precise each representation reflects the underlying content stored in that collection. As collections evolve over time, collection representations should also be updated to reflect any change, however, a current solution has not yet been proposed. In this study we examine both the implications of out-of-date representation sets on retrieval accuracy, as well as proposing three different policies for managing necessary updates. Each policyis evaluated on a testbed of forty-four dynamic collections over an eight-week period. Our findings show that out-of-date representations significantly degrade performance overtime, however, adopting a suitable update policy can minimise this problem. Milad Shokouhi, Mark Baillie, Leif Azzopardi |
SIGIR | 3 |
| 2006 | An Efficient Computation of the Multiple-Bernoulli Language Model
Leif Azzopardi, David E. Losada |
ECIR | 1 |
| 2006 | Adaptive query-based sampling for distributed IRabstractNo abstract available. Leif Azzopardi, Mark Baillie, Fabio Crestani |
SIGIR | 1 |
| 2006 | Automatic construction of known-item finding test bedsabstractNo abstract available. Leif Azzopardi, Maarten de Rijke |
SIGIR | 1 |
| 2006 | Formal models for expert finding in enterprise corporaabstractSearching an organization's document repositories for experts provides a cost effective solution for the task of expert finding. We present two general strategies to expert searching given a document collection which are formalized using generative probabilistic models. The first of these directly models an expert's knowledge based on the documents that they are associated with, whilst the second locates documents on topic, and then finds the associated expert. Forming reliable associations is crucial to the performance of expert finding systems. Consequently, in our evaluation we compare the different approaches, exploring a variety of associations along with other operational parameters (such as topicality). Using the TREC Enterprise corpora, we show that the second strategy consistently outperforms the first. A comparison against other unsupervised techniques, reveals that our second model delivers excellent performance. Krisztian Balog, Leif Azzopardi, Maarten de Rijke |
SIGIR | 2 |
| 2006 | Adaptive Query-Based Sampling of Distributed Collections
Mark Baillie, Leif Azzopardi, Fabio Crestani |
SPIRE | 2 |
| 2005 | Age Dependent Document Priors in Link Structure Analysis
Claudia Hauff, Leif Azzopardi |
ECIR | 2 |
| 2005 | Probabilistic hyperspace analogue to languageabstractSong and Bruza [6] introduce a framework for Information Retrieval(IR) based on Gardenfor's three tiered cognitive model; Conceptual Spaces[4]. They instantiate a conceptual space using Hyperspace Analogue to Language (HAL[3] to generate higher order concepts which are later used for ad-hoc retrieval. In this poster, we propose an alternative implementation of the conceptual space by using a probabilistic HAL space (pHAL). To evaluate whether converting to such an implementation is beneficial we have performed an initial investigation comparing the concept combination of HAL against pHAL for the task of query expansion. Our experiments indicate that pHAL outperforms the original HAL method and that better query term selection methods can improve performance on both HAL and pHAL. Leif Azzopardi, Mark A. Girolami, Malcolm K. Crowe |
SIGIR | 1 |
| 2004 | User biased document language modellingabstractCapitalizing on the intuitive underlying assumptions of Language Modelling for Ad-Hoc Retrieval we present a novel approach that is capable of injecting the user's context of the document collection into the retrieval process. The preliminary findings from the evaluation undertaken suggest that improved IR performance is possible under certain circumstances. This motivates further investigation to determine the extent and significance of this improved performance. Leif Azzopardi, Mark A. Girolami, C. J. van Rijsbergen |
SIGIR | 1 |
| 2003 | Investigating the relationship between language model perplexity and IR precision-recall measuresabstractAn empirical study has been conducted investigating the relationship between the performance of an aspect based language model in terms of perplexity and the corresponding information retrieval performance obtained. It is observed, on the corpora considered, that the perplexity of the language model has a systematic relationship with the achievable precision recall performance though it is not statistically significant. Leif Azzopardi, Mark A. Girolami, C. J. van Rijsbergen |
SIGIR | 1 |