VLDB 2026 Research / reviewers in the wild / expert
Martin Halvey
dblp:27/417
· DBLP profile ↗
28ranked-venue papers in the field
9as first author
8since 2021 · last 2026
0000-0001-6387-8679ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 27 (8 first)Other / Interdisciplinary · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Expanded Tag Genomes for Cross-Domain RecommendationabstractTag genome is widely used in recommender systems research to, for example, measure item similarity, make recommendations and generate recommendation explanations. Applying tag genome to problems in cross-domain recommendation, however, is complicated by the limited item overlap between cross-domain recommendation data sets and the available tag genomes. Furthermore, existing tag prediction models rely on content-based features that are not readily available in a majority of recommendation data sets. To address these issues, we generated tag genomes for both movies and books based on the Amazon data set, which is widely used in cross-domain recommendation research. These new tag genomes are over 200 × larger than the previous versions and can support comparative evaluation of tag-based and collaborative methods, facilitate the development of new cross-domain recommendation algorithms and provide a foundation for studying phenomena, such as serendipity and diversity, across multiple domains. Both data sets and the data generation pipeline are freely available at https://github.com/Bionic1251/Expanded-Tag-Genomes. Denis Kotkov, Alan Medlar, Dorota Glowacka, Martin Halvey |
CHIIR | 4 |
| 2024 | The Influence of Presentation and Performance on User SatisfactionabstractInformation Retrieval (IR) systems are designed to provide users with a ranked list of results based on their queries. The effectiveness of an IR system is gauged not just by its ability to retrieve relevant results but also by how it presents these results to users; an engaging presentation often correlates with increased user satisfaction. While existing research has delved into the link between user satisfaction, IR performance metrics, and presentation, these aspects have typically been investigated in isolation. Our research aims to bridge this gap by examining the relationship between query performance, presentation and user satisfaction. For our analysis, we conducted a between-subjects experiment comparing the effectiveness of various result card layouts for an ad-hoc news search interface. Drawing data from the TREC WaPo 2018 collection, we centered our study on four specific topics. Within each of these topics, we assessed six distinct queries with varying nDCG values. Our study involved 164 participants who were exposed to one of five distinct layouts containing result cards, such as “title”, “title+image”, or “title+image+summary”. Our findings indicate that while nDCG is a strong predictor of user satisfaction at the query level, there exists no linear relationship between the performance of the query, presentation of results and user satisfaction. However, when considering the total gain on the initial result page, we observed that presentation does play a significant role in user satisfaction (at the query level) for certain layouts with result cards such as, title+image or title+image+summary. Our results also suggest that the layout differences have complex and multifaceted impacts on satisfaction. We demonstrate the capacity to equalize user satisfaction levels between queries of varying performance by changing how results are presented. This emphasizes the necessity to harmonize both performance and presentation in IR systems, considering users’ diverse preferences. Ultimately, our insights can steer the evolution of more user-aligned IR systems, underscoring the balance between system performance and result presentation. Kanaad Pathak, Leif Azzopardi, Martin Halvey |
CHIIR | 3 |
| 2024 | Ranking Heterogeneous Search Result Pages Using the Interactive Probability Ranking Principle
Kanaad Pathak, Leif Azzopardi, Martin Halvey |
ECIR (2) | 3 |
| 2024 | A Systematic Review of Cost, Effort, and Load Research in Information Search and Retrieval, 1972-2020abstractDuring the information search and retrieval (ISR) process, user-system interactions such as submitting queries, examining results, and engaging with information impose some degree of demand on the user’s resources. Within ISR, these demands are well recognised, and numerous studies have demonstrated that the cost, effort, and load (CEL) experienced during the search process are affected by a variety of factors. Despite this recognition, there is no universally accepted definition of the constructs of CEL within the field of ISR. Ultimately, this has led to problems with how these constructs have been interpreted and subsequently measured. This systematic review contributes a synthesis of literature, summarising key findings relating to how researchers have been defining and measuring CEL within ISR over the past 50 years. After manually screening 1,109 articles, we detailed and analysed 91 articles which examine CEL within ISR. The discussion focuses on comparing the similarities and differences between CEL definitions and measures before identifying the limitations of the current state of the nomenclature. Opportunities for future research are also identified. Going forward, we propose a CEL taxonomy that integrates the relationships between CEL and their related constructs, which will help focus and disambiguate future research in this important area. Molly McGregor, Leif Azzopardi, Martin Halvey |
ACM Trans. Inf. Syst. | 3 |
| 2023 | Driven to Distraction: Examining the Influence of Distractors on Search Behaviours, Performance and ExperienceabstractAdvertisements, sponsored links, clickbait, in-house recommendations and similar elements pervasively shroud featured content. Such elements vie for people’s attention, potentially distracting people from their task at hand. The effects of such “distractors” is likely to increase people’s cognitive workload and reduce their performance as they need to work harder to discern the relevant from non-relevant. In this paper, we investigate how people of varying cognitive abilities (measured using Perceptual Speed and Cognitive Failure instruments) are affected by these different types of distractions when completing search tasks. We performed a crowdsourced within-subjects user study, where 102 participants completed four search tasks using our news search engine over four different interface conditions: (i) one with no additional distractors; (ii) one with advertisements; (iii) one with sponsored links; and (iv) one with in-house recommendations. Our results highlight a number of important trends and findings. Participants perceived the interface condition without distractors as significantly better across numerous dimensions. Participants reported higher satisfaction, lower workload, higher topic recall, and found it easier to concentrate. Behaviourally, participants issued queries faster and clicked results earlier when compared to the interfaces with distractors. When using the interfaces with distractors, one in ten participants clicked on a distractor—and despite engaging with a distractor for less than twenty seconds, their task time increased by approximately two minutes. We found that the effects were magnified depending on cognitive abilities—with a greater impact of distractors on participants with lower perceptual speed, and for those with a higher propensity of cognitive failures. Distractors—regardless of their type—have negative consequences on a user’s search experience and performance. As a consequence, interfaces containing visually distracting elements are creating poorer search experiences due to the “distractor tax” being placed on people’s limited attention. Leif Azzopardi, David Maxwell 0001, Martin Halvey, Claudia Hauff |
CHIIR | 3 |
| 2022 | Partners in life and online search: Investigating older couples' collaborative information seekingabstractOlder adults frequently collaborate with their spouses in daily tasks and problem solving. Despite information seeking being an important aspect of collaboration, the information seeking behaviour of older adults and in particular couples remains under investigated. To address this gap, in this paper we present a qualitative investigation of older adults’ collaborative information seeking. Through in-depth interviews and demonstrations of real-life search tasks with eleven older couples, we show that older couples frequently engage in collaborative information seeking in daily tasks, interests, and to satisfy curiosity. Our research suggests that collaborative information seeking is a relationship maintenance behaviour among older couples, and that their long-term relationships may play a key role in how they communicate, make decisions, and develop divide and conquer strategies by taking on various roles during their collaborative information seeking. We also found that older couples construct shared views toward technology adoption and usage despite their individual differences. We include some reflections on the existing collaborative information systems and how they may adapt to fit older couples’ collaborative information seeking. Winter Wei, Cosmin Munteanu, Martin Halvey |
CHIIR | 3 |
| 2021 | Investigating the Influence of Ads on User Search Performance, Behaviour, and Experience during Information SeekingabstractThe phenomenon of banner blindness explains that users can mentally ignore online advertisements (ads). However, eye-tracking studies have shown that users still fixate on ads, and even without direct gaze, ads still fall within a user's peripheral vision, which may negatively overload cognition. It is therefore unknown how blind, banner blindness, truly is, and what other effect ads may have on user's information seeking. To address this gap, a within-subjects design experiment was conducted with 37 participants who performed search tasks from the TREC 2017 Common Core News Collection, where 3 search tasks contained various types of ads, and one search task had no ads. Although our results showed that on average, participants retrieved similar amounts of relevant documents regardless of whether ads were present or absent, participants took significantly longer achieving this performance when ads were present. Furthermore, when ads were absent, participants reported less frustration, and not only believed they learned more, but a post-task recall test showed that participants actually did learn up to 38% more. Consequently, our findings suggest that banner blindness is more costly than just mere annoyance, and that the influence of ads on user's information retrieval recall may extend current theories of visual crowding. Olivia Foulds, Leif Azzopardi, Martin Halvey |
CHIIR | 3 |
| 2021 | Untangling Cost, Effort, and Load in Information Seeking and RetrievalabstractWhen performing Information Seeking and Retrieval (ISR) activities, people submit queries, examine results, assess documents and engage with the information to make decisions and complete tasks. All these activities come at a "cost'', but within the field of ISR there is no universally accepted definition of the concepts of Cost, Effort and Load (CEL). Instead, researchers have used the same terms interchangeably to describe similar but also different concepts. This lack of shared understanding has led to a disconnect between how these concepts are defined and discussed versus how they are interpreted and measured. Thus, the aim of this paper is two-fold: (i) to review the meaning of CEL related concepts used within ISR, and (ii) to create a shared taxonomy of the concepts relating to CEL in ISR. To seed our analysis, we conducted a literature review, where 397 papers were reviewed, and twenty-six papers that explicitly proposed measures or definitions of CEL were selected for analysis. By drawing upon theory from Psychology and other fields, we present the common definitions of CEL in order to ground our discussion of these concepts in ISR. We also highlight the issues associated with CEL measurement in ISR to help researchers reflect on the validity and precision of existing methods. We hope this perspectives paper serves as a basis for a taxonomy of how CEL concepts are used within ISR- where we have provided a series of working definitions that clearly delineate the different concepts being used, investigated and measured in ISR research. Molly McGregor, Leif Azzopardi, Martin Halvey |
CHIIR | 3 |
| 2020 | Reflecting upon Perceptual Speed Tests in Information Retrieval: Limitations, Challenges, and RecommendationsabstractPerceptual Speed (PS) is a cognitive ability defined by an individual's accuracy and speed to scan information while completing visual search tasks. Prior studies using PS tests have demonstrated that PS affects multiple factors in Information Retrieval (IR), such as a user's search performance, interaction with the system, time spent completing tasks, and subjective impression of their workload. With greater knowledge of PS, systems could be designed that accommodate users with low PS to improve their overall search experience. However, in this perspectives paper, we analyse how PS tests have been used in IR, and identify multiple uncertainties regarding PS content, administration, analysis, and reporting of findings. Consequently, we aim to stir discussion between IR researchers by drawing awareness to these issues. As a result, we further discuss challenges involved in advancing how future PS tests are used in IR. Finally, we propose recommendations that have the potential for enhancing the reliability and validity of current PS tests. Olivia Foulds, Leif Azzopardi, Martin Halvey |
CHIIR | 3 |
| 2020 | Predicting Perceptual Speed from Search BehaviourabstractPerceptual Speed (PS) is a cognitive ability that is known to affect multiple factors in Information Retrieval (IR) such as a user's search performance and subjective experience. However PS tests are difficult to administer which limits the design of user-adaptive systems that can automatically infer PS to appropriately accommodate low PS users. Consequently, this paper evaluated whether PS can be automatically classified from search behaviour using several machine learning models trained on features extracted from TREC Common Core search task logs. Our results are encouraging: given a user's interactions from one query, a Decision Tree was able to predict a user's PS as low or high with 86% accuracy. Additionally, we identified different behavioural components for specific PS tests, implying that each PS test measures different aspects of a person's cognitive ability. These findings motivate further work for how best to design search systems that can adapt to individual differences. Olivia Foulds, Alessandro Suglia, Leif Azzopardi, Martin Halvey |
SIGIR | 4 |
| 2018 | Beyond traditional collaborative search: Understanding the effect of awareness on multi-level collaborative information retrieval
Nyi Nyi Htun, Martin Halvey, Lynne Baillie |
Inf. Process. Manag. | 2 |
| 2017 | An Interface for Supporting Asynchronous Multi-Level Collaborative Information RetrievalabstractA great deal of research into Collaborative Information Retrieval (CIR) has assumed that search team members have the same level of unrestricted access to information. However, case studies and observations from different domains including government, healthcare and legal, have suggested that CIR sometimes involves people with unequal access to information. This type of scenario has been referred to as Multi-Level CIR (MLCIR). In addition to supporting collaboration, MLCIR systems must ensure that there is no unintended disclosure of sensitive information, this is an under investigated area of research. In this paper we present results of an evaluation of an interface we have designed for MLCIR scenarios. Pairs of participants used the interface under 3 different information access scenarios for a variety of search tasks. These scenarios included 1 CIR and 2 MLCIR scenarios, namely: full access (FA), document removal (DR) and term blacklisting (TR). Design interviews were conducted post evaluation to obtain qualitative feedback from participants. Evaluation results showed that our interface performed well for both DR and FA scenarios but for TR, team members with less access had a negative influence on their partner's search performance, demonstrating insights into how different MLCIR scenarios should be supported. Design interview results showed that our interface helped the participants to reformulate their queries, understand their partner's performance, reduce duplicated work and review their team's search history without disclosing sensitive information. Nyi Nyi Htun, Martin Halvey, Lynne Baillie |
CHIIR | 2 |
| 2017 | How Can We Better Support Users with Non-Uniform Information Access in Collaborative Information Retrieval?abstractThe majority of research in Collaborative Information Retrieval (CIR) has assumed that collaborating team members have uniform information access. However, practice and research has shown that there may not always be uniform information access among team members, e.g. in healthcare, government, etc. To the best of our knowledge, there has not been a controlled user evaluation to measure the impact of non-uniform information access on CIR outcomes. To address this shortcoming, we conducted a controlled user evaluation using 2 non-uniform access scenarios (document removal and term blacklisting) and 1 full and uniform access scenario. Following this, a design interview was undertaken to provide interface design suggestions. Evaluation results show that neither of the 2 non-uniform access scenarios had a significant negative impact on collaborative and individual search outcomes. Design interview results suggested that awareness of team's query history and intersecting viewed/judged documents could potentially help users share their expertise without disclosing sensitive information. Based on our results we provide important design recommendations to better support users with non-uniform information access in CIR. Nyi Nyi Htun, Martin Halvey, Lynne Baillie |
CHIIR | 2 |
| 2016 | Video Test Collection with Graded Relevance AssessmentsabstractRelevance is a complex, but core, concept within the field of Information Retrieval. In order to allow system comparisons the many factors that influence relevance are often discarded to allow abstraction to a single score relating to relevance. This means that a great wealth of information is often discarded. In this paper we outline the creation of a video test collection with graded relevance assessments, to the best of our knowledge the first example of such a test collection for video retrieval. To directly address the shortcoming above we also gathered behavioural and perceptual data from assessors during the assessment process. All of this information along with judgements are available for download. Our intention is to allow other researchers to supplement the judgements to help create an adaptive test collection which contains supplementary information rather than a completely static collection with binary judgements. Weng Qiying, Martin Halvey, Robert Villa |
CHIIR | 2 |
| 2016 | A Comparison of Primary and Secondary Relevance Judgements for Real-Life TopicsabstractThe notion of relevance is fundamental to the field of Information Retrieval. Within the field a generally accepted conception of relevance as inherently subjective has emerged, with an individual's assessment of relevance influenced by numerous contextual factors. In this paper we present a user study that examines in detail the differences between primary and secondary assessors on a set of "real-world" topics which were gathered specifically for the work. By gathering topics which are representative of the staff and students at a major university, at a particular point in time, we aim to explore differences between primary and secondary relevance judgements for real-life search tasks. Findings suggest that while secondary assessors may find the assessment task challenging in various ways (they generally possess less interest and knowledge in secondary topics and take longer to assess documents), agreement between primary and secondary assessors is high. Simon Wakeling, Martin Halvey, Robert Villa, Laura Hasler |
CHIIR | 2 |
| 2016 | Beyond actions: Exploring the discovery of tactics from user logs
Jiyin He, Pernilla Qvarfordt, Martin Halvey, Gene Golovchinsky |
Inf. Process. Manag. | 3 |
| 2015 | Towards Quantifying the Impact of Non-Uniform Information Access in Collaborative Information RetrievalabstractThe majority of research into Collaborative Information Retrieval (CIR) has assumed a uniformity of information access and visibility between collaborators. However in a number of real world scenarios, information access is not uniform between all collaborators in a team e.g. security, health etc. This can be referred to as Multi-Level Collaborative Information Retrieval (MLCIR). To the best of our knowledge, there has not yet been any systematic investigation of the effect of MLCIR on search outcomes. To address this shortcoming, in this paper, we present the results of a simulated evaluation conducted over 4 different non-uniform information access scenarios and 3 different collaborative search strategies. Results indicate that there is some tolerance to removing access to the collection and that there may not always be a negative impact on performance. We also highlight how different access scenarios and search strategies impact on search outcomes. Nyi Nyi Htun, Martin Halvey, Lynne Baillie |
SIGIR | 2 |
| 2014 | Evaluating the effort involved in relevance assessments for imagesabstractHow assessors and end users judge the relevance of images has been studied in information science and information retrieval for a considerable time. The criteria by which assessors' judge relevance has been intensively studied, and there has been a large amount of work which has investigated how relevance judgments for test collections can be more cheaply generated, such as through crowd sourcing. Relatively little work has investigated the process individual assessors go through to judge the relevance of an image. In this paper, we focus on the process by which relevance is judged for images, and in particular, the degree of effort a user must expend to judge relevance for different topics. Results suggest that topic difficulty and how semantic/visual a topic is impact user performance and perceived effort. Martin Halvey, Robert Villa |
SIGIR | 1 |
| 2014 | SIGIR 2014 workshop on gathering efficient assessments of relevance (GEAR)abstractEvaluation is a fundamental part of Information Retrieval, and in the conventional Cranfield evaluation paradigm, sets of relevance assessments are a fundamental part of test collections. This workshop revisits how relevance assessments can be efficiently created, seeking to provide a forum for discussion and exploration of the topic. Martin Halvey, Robert Villa, Paul D. Clough |
SIGIR | 1 |
| 2014 | Supporting exploratory video retrieval tasks with grouping and recommendation
Martin Halvey, David Vallet, David Hannah, Joemon M. Jose |
Inf. Process. Manag. | 1 |
| 2013 | Is relevance hard work?: evaluating the effort of making relevant assessmentsabstractThe judging of relevance has been a subject of study in information retrieval for a long time, especially in the creation of relevance judgments for test collections. While the criteria by which assessors? judge relevance has been intensively studied, little work has investigated the process individual assessors go through to judge the relevance of a document. In this paper, we focus on the process by which relevance is judged, and in particular, the degree of effort a user must expend to judge relevance. By better understanding this effort in isolation, we may provide data which can be used to create better models of search. We present the results of an empirical evaluation of the effort users must exert to judge the relevance of document, investigating the effect of relevance level and document size. Results suggest that 'relevant' documents require more effort to judge when compared to highly relevant and not relevant documents, and that effort increases as document size increases. Robert Villa, Martin Halvey |
SIGIR | 2 |
| 2012 | Assessing and Predicting Vertical Intent for Web Queries
Ke Zhou 0003, Ronan Cummins, Martin Halvey, Mounia Lalmas-Roelleke, Joemon M. Jose |
ECIR | 3 |
| 2010 | An asynchronous collaborative search system for online video search
Martin Halvey, David Vallet, David Hannah, Yue Feng 0002, Joemon M. Jose |
Inf. Process. Manag. | 1 |
| 2009 | Diversity, Assortment, Dissimilarity, Variety: A Study of Diversity Measures Using Low Level Features for Video Retrieval
Martin Halvey, P. Punitha 0001, David Hannah, Robert Villa, Frank Hopfgartner, Anuj Goyal, Joemon M. Jose |
ECIR | 1 |
| 2007 | Exploring social dynamics in online media sharingabstractIt is now feasible to view media at home as easily as text-based pages were viewed when the World Wide Web (WWW) first emerged. This development has supported media sharing and search services providing hosting, indexing and access to large, online media repositories. Many of these sharing services also have a social aspect to them. This paper provides an initial analysis of the social interactions on a video sharing and search service (www.youtube.com). Results show that many users do not form social networks in the online community and a very small number do not appear to contribute to the wider community. However, it does seem those people who do use the available tools have much a greater tendency to form social connections. Martin Halvey, Mark T. Keane |
WWW | 1 |
| 2007 | An assessment of tag presentation techniquesabstractWith the growth of social bookmarking a new approach for metadata creation called tagging has emerged. In this paper we evaluate the use of tag presentation techniques. The main goal of our evaluation is to investigate the effect of some of the different properties that can be utilized in presenting tags e.g. alphabetization, using larger fonts etc. We show that a number of these factors can affect the ease with which users can find tags and use the tools for presenting tags to users. Martin Halvey, Mark T. Keane |
WWW | 1 |
| 2006 | Temporal rules for mobile web personalizationabstractMany systems use past behavior, preferences and environmental factors to attempt to predict user navigation on the Internet. However we believe that many of these models have shortcomings, in that they do not take into account that users may have many different sets of preferences. Here we investigate an environmental factor, namely time, in making predictions about user navigation. We present methods for creating temporal rules that describe user navigation patterns. We also show the benefit of using these rules to predict user navigation and also show the benefits of these models over traditional methods. An analysis is carried out on a sample of usage logs for Wireless Application Protocol (WAP) browsing, and the results of this analysis verify our hypothesis. Martin Halvey, Mark T. Keane, Barry Smyth |
WWW | 1 |
| 2005 | Time Based Segmentation of Log Data for User Navigation Prediction in PersonalizationabstractThere are many systems that attempt to predict user navigation on the Internet through the use of past behavior, preferences and environmental factors. We believe that many of these models have shortcomings, in that they do not take into account that users may have many different sets of preferences, specifically, we investigate time as an environmental factor in making predictions about user navigation. We present a method for segmenting log files in order to learn time dependent models to predict user navigation patterns and show the benefits of these models over traditional methods. An analysis is carried out on a sample of usage logs for wireless application protocol (WAP) browsing, and the results of this analysis verify our hypothesis. Martin Halvey, Mark T. Keane, Barry Smyth |
Web Intelligence | 1 |