VLDB 2026 Research / reviewers in the wild / expert
Gabriella Kazai
dblp:k/GabriellaKazai
· DBLP profile ↗
53ranked-venue papers in the field
26as first author
5since 2021 · last 2025
0009-0002-5158-6630ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 47 (25 first)Other / Interdisciplinary · 4Data Mining & Knowledge Discovery · 2 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bridging the Gap: From Ad-hoc to Proactive Search in ConversationsabstractProactive search in conversations (PSC) aims to reduce user effort in formulating explicit queries by proactively retrieving useful relevant information given conversational context. Previous work in PSC either directly uses this context as input to off-the-shelf ad-hoc retrievers or further fine-tunes them on PSC data. However, ad-hoc retrievers are pre-trained on short and concise queries, while the PSC input is longer and noisier. This input mismatch between ad-hoc search and PSC limits retrieval quality. While fine-tuning on PSC data helps, its benefits remain constrained by this input gap. In this work, we propose Conv2Query, a novel conversation-to-query framework that adapts ad-hoc retrievers to PSC by bridging the input gap between ad-hoc search and PSC. Conv2Query maps conversational context into ad-hoc queries, which can either be used as input for off-the-shelf ad-hoc retrievers or for further fine-tuning on PSC data. Extensive experiments on two PSC datasets show that Conv2Query significantly improves ad-hoc retrievers' performance, both when used directly and after fine-tuning on PSC. Chuan Meng, Francesco Tonolini, Fengran Mo, Nikolaos Aletras, Emine Yilmaz, Gabriella Kazai |
SIGIR | 6 |
| 2024 | What Matters in a Measure? A Perspective from Large-Scale Search EvaluationabstractInformation retrieval (IR) has a large literature on evaluation, dating back decades and forming a central part of the research culture. The largest proportion of this literature discusses techniques to turn a sequence of relevance labels into a single number, reflecting the system's performance: precision or cumulative gain, for example, or dozens of alternatives. Those techniques-metrics-are themselves evaluated, commonly by reference to sensitivity and validity. Paul Thomas 0001, Gabriella Kazai, Nick Craswell, Seth Spielman |
SIGIR | 2 |
| 2023 | On the Reliability of User Feedback for Evaluating the Quality of Conversational AgentsabstractWe analyse the reliability of users' explicit feedback for evaluating the quality of conversational agents. Using data from a commercial conversational system, we analyse how user feedback compares with human annotations; how well it aligns with implicit user satisfaction signals, such as retention; and how much user feedback is needed to reliably evaluate the quality of a conversational system. Jordan Massiah, Emine Yilmaz, Yunlong Jiao, Gabriella Kazai |
CIKM | 4 |
| 2022 | The Crowd is Made of People: Observations from Large-Scale Crowd LabellingabstractLike many other researchers, at Microsoft Bing we use external “crowd” judges to label results from a search engine—especially, although not exclusively, to obtain relevance labels for offline evaluation in the Cranfield tradition. Crowdsourced labels are relatively cheap, and hence very popular, but are prone to disagreements, spam, and various biases which appear to be unexplained “noise” or “error”. In this paper, we provide examples of problems we have encountered running crowd labelling at large scale and around the globe, for search evaluation in particular. We demonstrate effects due to the time of day and day of week that a label is given; fatigue; anchoring; exposure; left-side bias; task switching; and simple disagreement between judges. Rather than simple “error”, these effects are consistent with well-known physiological and cognitive factors. “The crowd” is not some abstract machinery, but is made of people. Human factors that affect people’s judgement behaviour must be considered when designing research evaluations and in interpreting evaluation metrics. Paul Thomas 0001, Gabriella Kazai, Ryen W. White, Nick Craswell |
CHIIR | 2 |
| 2022 | Less is Less: When are Snippets Insufficient for Human vs Machine Relevance Estimation?
Gabriella Kazai, Bhaskar Mitra 0001, Anlei Dong, Nick Craswell, Linjun Yang |
ECIR (2) | 1 |
| 2019 | The Emotion Profile of Web SearchabstractEmotions are an essential part of most human activities, including decision-making. Emotions arise in response to information, e.g., presented in web pages, and are also expressed in the words used to convey that information in the first place. In this paper, we study the emotion profile of retrieved and clicked web search results towards the goal of better understanding the role of emotions in web search. Using click logs from a four-month period, up to the end of January 2019, we examine the emotions associated with search results and contrast them to the emotions of clicked results, taking rank and relevance into account. Emotions are assigned to web pages based on two lexicons: SentiWordNet (positive, negative and objective sentiments) and EmoLexData (afraid, amused, angry, annoyed, don't care, happy, inspired, and sad emotions). We look at the sentiment/emotion profiles of search results grouped around a set of controversial and mundane topics and hypothesise that users are more likely to click emotionally charged results than emotionless results, both in general, and in particular when their query relates to controversial topics. Gabriella Kazai, Paul Thomas 0001, Nick Craswell |
SIGIR | 1 |
| 2016 | First International Workshop on Recent Trends in News Information Retrieval (NewsIR'16)
Miguel Martinez-Alvarez, Udo Kruschwitz, Gabriella Kazai, Frank Hopfgartner, David P. A. Corney, Ricardo Campos 0001, M-Dyaa Albakour |
ECIR | 3 |
| 2016 | Personalised News and Blog Recommendations based on User Location, Facebook and Twitter User ProfilingabstractThis demo presents a prototype mobile app that provides out-of-the-box personalised content recommendations to its users by leveraging and combining the user's location, their Facebook and/or Twitter feed and their in-app actions to automatically infer their interests. We build individual models for each user and each location. At retrieval time we construct the user's personalised feed by mixing different sources of content-based recommendations with content directly from their Facebook/Twitter feeds, locally trending articles and content propagated through their in-app social network. Both explicit and implicit feedback signals from the users' interactions with their recommendations are used to update their interests models and to learn their preferences over the different content sources. Gabriella Kazai, Iskander Yusof, Daoud Clarke |
SIGIR | 1 |
| 2016 | Third International Workshop on Gamification for Information Retrieval (GamifIR 2016)abstractStronger engagement and greater participation is often crucial to reach a goal or to solve an issue. Issues like the emerging employee engagement crisis, insufficient knowledge sharing, and chronic procrastination. In many cases we need and search for tools to beat procrastination or to change people's habits. Gamification is the approach to learn from often fun, creative and engaging games. In principle, it is about understanding games and applying game design elements in a non-gaming environments. This offers possibilities for wide area improvements. For example more accurate work, better retention rates and more cost effective solutions by relating motivations for participating as more intrinsic than conventional methods. In the context of Information Retrieval (IR) it is not hard to imagine that many tasks could benefit from gamification techniques. Besides several manual annotation tasks of data sets for IR research, user participation is important in order to gather implicit or even explicit feedback to feed the algorithms. Gamification, however, comes with its own challenges and its adoption in IR is still in its infancy. Given the enormous response to the first and second GamifIR workshops that were both co-located with ECIR, and the broad range of topics discussed, we now organized the third workshop at SIGIR 2016 to address a range of emerging challenges and opportunities. Michael Meder, Frank Hopfgartner, Gabriella Kazai, Udo Kruschwitz |
SIGIR | 3 |
| 2016 | Quality Management in Crowdsourcing using Gold Judges BehaviorabstractCrowdsourcing relevance labels has become an accepted practice for the evaluation of IR systems, where the task of constructing a test collection is distributed over large populations of unknown users with widely varied skills and motivations. Typical methods to check and ensure the quality of the crowd's output is to inject work tasks with known answers (gold tasks) on which workers' performance can be measured. However, gold tasks are expensive to create and have limited application. A more recent trend is to monitor the workers' interactions during a task and estimate their work quality based on their behavior. In this paper, we show that without gold behavior signals that reflect trusted interaction patterns, classifiers can perform poorly, especially for complex tasks, which can lead to high quality crowd workers getting blocked while poorly performing workers remain undetected. Through a series of crowdsourcing experiments, we compare the behaviors of trained professional judges and crowd workers and then use the trained judges' behavior signals as gold behavior to train a classifier to detect poorly performing crowd workers. Our experiments show that classification accuracy almost doubles in some tasks with the use of gold behavior data. Gabriella Kazai, Imed Zitouni |
WSDM | 1 |
| 2015 | Second International Workshop on Gamification for Information Retrieval (GamifIR'15)
Frank Hopfgartner, Gabriella Kazai, Udo Kruschwitz, Michael Meder, Mark Shovman |
ECIR | 2 |
| 2015 | A Personalised Reader for Crowd Curated Content
Gabriella Kazai, Daoud Clarke, Iskander Yusof, Matteo Venanzi |
RecSys | 1 |
| 2014 | Workshop on Gamification for Information Retrieval (GamifIR'14)
Frank Hopfgartner, Gabriella Kazai, Udo Kruschwitz, Michael Meder |
ECIR | 2 |
| 2014 | Dissimilarity Based Query Selection for Efficient Preference Based IR Evaluation
Gabriella Kazai, Homer Sung |
ECIR | 1 |
| 2014 | Community-based bayesian aggregation models for crowdsourcingabstractThis paper addresses the problem of extracting accurate labels from crowdsourced datasets, a key challenge in crowdsourcing. Prior work has focused on modeling the reliability of individual workers, for instance, by way of confusion matrices, and using these latent traits to estimate the true labels more accurately. However, this strategy becomes ineffective when there are too few labels per worker to reliably estimate their quality. To mitigate this issue, we propose a novel community-based Bayesian label aggregation model, CommunityBCC, which assumes that crowd workers conform to a few different types, where each type represents a group of workers with similar confusion matrices. We assume that each worker belongs to a certain community, where the worker's confusion matrix is similar to (a perturbation of) the community's confusion matrix. Our model can then learn a set of key latent features: (i) the confusion matrix of each community, (ii) the community membership of each user, and (iii) the aggregated label of each item. We compare the performance of our model against established aggregation methods on a number of large-scale, real-world crowdsourcing datasets. Our experimental results show that our CommunityBCC model consistently outperforms state-of-the-art label aggregation methods, requiring, on average, 50% less data to pass the 90% accuracy mark. Matteo Venanzi, John Guiver, Gabriella Kazai, Pushmeet Kohli, Milad Shokouhi |
WWW | 3 |
| 2013 | User intent and assessor disagreement in web search evaluationabstractPreference based methods for collecting relevance data for information retrieval (IR) evaluation have been shown to lead to better inter-assessor agreement than the traditional method of judging individual documents. However, little is known as to why preference judging reduces assessor disagreement and whether better agreement among assessors also means better agreement with user satisfaction, as signaled by user clicks. In this paper, we examine the relationship between assessor disagreement and various click based measures, such as click preference strength and user intent similarity, for judgments collected from editorial judges and crowd workers using single absolute, pairwise absolute and pairwise preference based judging methods. We find that trained judges are significantly more likely to agree with each other and with users than crowd workers, but inter-assessor agreement does not mean agreement with users. Switching to a pairwise judging mode improves crowdsourcing quality close to that of trained judges. We also find a relationship between intent similarity and assessor-user agreement, where the nature of the relationship changes across judging modes. Overall, our findings suggest that the awareness of different possible intents, enabled by pairwise judging, is a key reason of the improved agreement, and a crucial requirement when crowdsourcing relevance data. Gabriella Kazai, Emine Yilmaz, Nick Craswell, Seyed M. M. Tahaghoghi |
CIKM | 1 |
| 2013 | ICDAR 2013 Competition on Book Structure ExtractionabstractThis paper summarizes the 3rd Book Structure Extraction competition that was run at the ICDAR 2013. Its goal is to evaluate and compare automatic techniques for deriving structure information from digitized books, which could then be used to aid navigation inside the books. More specifically, the task that participants are faced with is to construct hyper linked tables of contents for a collection of 1,000 digitized books. This paper reviews the setup of the competition, the book collection used in the task, and the measures used for the evaluation. The main novelty of the 2013 competition is that we were able to rely on an external provider for the ground truthing phase, hence granting the consistency of the evaluation. In addition, this permitted to nearly double the number of annotated books from the 1,040 books annotated in 2009 and 2011 to over 2,000 books. The paper further presents the result performance of the 6 participating research teams, and briefly summarizes their approaches. Antoine Doucet, Gabriella Kazai, Sebastian Colutto, Günter Mühlberger |
ICDAR | 2 |
| 2013 | Relevance dimensions in preference-based IR evaluationabstractEvaluation of information retrieval (IR) systems has recently been exploring the use of preference judgments over two search result lists. Unlike the traditional method of collecting relevance labels per single result, this method allows to consider the interaction between search results as part of the judging criteria. For example, one result list may be preferred over another if it has a more diverse set of relevant results, covering a wider range of user intents. In this paper, we investigate how assessors determine their preference for one list of results over another with the aim to understand the role of various relevance dimensions in preference-based evaluation. We run a series of experiments and collect preference judgments over different relevance dimensions in side-by-side comparisons of two search result lists, as well as relevance judgments for the individual documents. Our analysis of the collected judgments reveals that preference judgments combine multiple dimensions of relevance that go beyond the traditional notion of relevance centered on topicality. Measuring performance based on single document judgments and NDCG aligns well with topicality based preferences, but shows misalignment with judges' overall preferences, largely due to the diversity dimension. As a judging method, dimensional preference judging is found to lead to improved judgment quality. Jin Young Kim 0005, Gabriella Kazai, Imed Zitouni |
SIGIR | 2 |
| 2013 | An analysis of human factors and label accuracy in crowdsourcing relevance judgments
Gabriella Kazai, Jaap Kamps, Natasa Milic-Frayling |
Inf. Retr. | 1 |
| 2012 | An analysis of systematic judging errors in information retrievalabstractTest collections are powerful mechanisms for the evaluation and optimization of information retrieval systems. However, there is reported evidence that experiment outcomes can be affected by changes to the judging guidelines or changes in the judge population. This paper examines such effects in a web search setting, comparing the judgments of four groups of judges: NIST Web Track judges, untrained crowd workers and two groups of trained judges of a commercial search engine. Our goal is to identify systematic judging errors by comparing the labels contributed by the different groups, working under the same or different judging guidelines. In particular, we focus on detecting systematic differences in judging depending on specific characteristics of the queries and URLs. For example, we ask whether a given population of judges, working under a given set of judging guidelines, are more likely to consistently overrate Wikipedia pages than another group judging under the same instructions. Our approach is to identify judging errors with respect to a consensus set, a judged gold set and a set of user clicks. We further demonstrate how such biases can affect the training of retrieval systems. Gabriella Kazai, Nick Craswell, Emine Yilmaz, Seyed M. M. Tahaghoghi |
CIKM | 1 |
| 2012 | The face of quality in crowdsourcing relevance labels: demographics, personality and labeling accuracyabstractInformation retrieval systems require human contributed relevance labels for their training and evaluation. Increasingly such labels are collected under the anonymous, uncontrolled conditions of crowdsourcing, leading to varied output quality. While a range of quality assurance and control techniques have now been developed to reduce noise during or after task completion, little is known about the workers themselves and possible relationships between workers' characteristics and the quality of their work. In this paper, we ask how do the relatively well or poorly-performing crowds, working under specific task conditions, actually look like in terms of worker characteristics, such as demographics or personality traits. Our findings show that the face of a crowd is in fact indicative of the quality of their work. Gabriella Kazai, Jaap Kamps, Natasa Milic-Frayling |
CIKM | 1 |
| 2012 | Booksonline'12: 5th workshop on online books, complementary social media and their impactabstractBooksOnline'12, the fifth workshop in the series, aims to offer a forum for bringing together expertise from academia, industry and libraries to facilitate the exchange of research results and technology in the field of digital libraries with specific focus on online books and complementary social media. The focus of this year's workshop is "engaging reading experiences", starting from the act of deciding what to read, through the exploration and interpretation of a book's content, to sharing the overall experience. Within this overall umbrella theme, the accepted papers naturally showed three salient themes: (1) Search and Discovery, (2) Personalization and Recommendation, and Reading Experiences beyond Text. The contributions demonstrate a range of technologies, including a collaborative tabletop visual approach to support the searching and discovery of books, co-citation methods to enhance document retrieval; exploring open issues in audio-book production to support non-text based reading and improving e-book accessibility; new approaches to recommendation that take into account writing style as well as looking specifically to young readers and their needs in order to develop recommendation tools that consider both content and reading level and match these against the readers' specific interests and reading ability. Following in the theme of the reader playing a central role in the future of our digital era, we are honored to welcome Maribeth Back from FX Palo Alto and Natasa Milic-Frayling from Microsoft Research as our keynote speakers. Gabriella Kazai, Monica Landoni, Carsten Eickhoff, Peter Brusilovsky |
CIKM | 1 |
| 2012 | Social book search: comparing topical relevance judgements and book suggestions for evaluationabstractThe Web and social media give us access to a wealth of information, not only different in quantity but also in character---traditional descriptions from professionals are now supplemented with user generated content. This challenges modern search systems based on the classical model of topical relevance and ad hoc search: How does their effectiveness transfer to the changing nature of information and to the changing types of information needs and search tasks? We use the INEX 2011 Books and Social Search Track's collection of book descriptions from Amazon and social cataloguing site LibraryThing. We compare classical IR with social book search in the context of the LibraryThing discussion forums where members ask for book suggestions. Specifically, we compare book suggestions on the forum with Mechanical Turk judgements on topical relevance and recommendation, both the judgements directly and their resulting evaluation of retrieval systems. First, the book suggestions on the forum are a complete enough set of relevance judgements for system evaluation. Second, topical relevance judgements result in a different system ranking from evaluation based on the forum suggestions. Although it is an important aspect for social book search, topical relevance is not sufficient for evaluation. Third, professional metadata alone is often not enough to determine the topical relevance of a book. User reviews provide a better signal for topical relevance. Fourth, user-generated content is more effective for social book search than professional metadata. Based on our findings, we propose an experimental evaluation that better reflects the complexities of social book search. Marijn Koolen, Jaap Kamps, Gabriella Kazai |
CIKM | 3 |
| 2012 | On Aggregating Labels from Multiple Crowd Workers to Infer Relevance of Documents
Mehdi Hosseini 0001, Ingemar J. Cox, Natasa Milic-Frayling, Gabriella Kazai, Vishwa Vinay |
ECIR | 4 |
| 2012 | On judgments obtained from a commercial search engineabstractIn information retrieval, relevance judgments play an important role as they are required both for evaluating the quality of retrieval systems and for training learning to rank algorithms. In recent years, numerous papers have been published using judgments obtained from a commercial search engine by researchers in industry. As typically no information is provided about the quality of these judgments, their reliability for evaluating/training retrieval systems remains questionable. In this paper, we analyze the reliability of such judgments for evaluating the quality of retrieval systems by comparing them to judgments by NIST judges at TREC. Emine Yilmaz, Gabriella Kazai, Nick Craswell, Seyed M. M. Tahaghoghi |
SIGIR | 2 |
| 2011 | BooksOnline'11: 4th workshop on online books, complementary social media, and crowdsourcingabstractThe BooksOnline Workshop series aims to foster the discussion and exchange of research ideas towards addressing challenges and exploring opportunities around large collections of digital books and complementary media. The fourth workshop in the series, BooksOnline'11 pays special attention to the role of social media and the phenomena of crowdsourcing in the context of online books, which is expected to be key in defining new user experiences in digital libraries and on the Web. The workshop boasts a high quality program, including keynote addresses by Ville Miettinnen, CEO of Microtask and Adam Farquhar, Head of Digital Library Technology at The British Library. From the accepted papers two main themes became salient: 1) Information retrieval and information extraction methods focused on enhancing digital libraries, and 2) Studies and analyses of reading experience and behaviour. This paper provides an overview of the workshop and the accepted contributions. Gabriella Kazai, Carsten Eickhoff, Peter Brusilovsky |
CIKM | 1 |
| 2011 | Worker types and personality traits in crowdsourcing relevance labelsabstractCrowdsourcing platforms offer unprecedented opportunities for creating evaluation benchmarks, but suffer from varied output quality from crowd workers who possess different levels of competence and aspiration. This raises new challenges for quality control and requires an in-depth understanding of how workers' characteristics relate to the quality of their work. Gabriella Kazai, Jaap Kamps, Natasa Milic-Frayling |
CIKM | 1 |
| 2011 | In Search of Quality in Crowdsourcing for Search Engine Evaluation
Gabriella Kazai |
ECIR | 1 |
| 2011 | ICDAR 2011 Book Structure Extraction CompetitionabstractIn this paper, we summarize the 2nd Book Structure Extraction competition run at ICDAR 2011. Its goal is to evaluate and compare automatic techniques for deriving structure information from digitized books, which could then be used to aid navigation inside the books. More specifically, the task that participants are faced with is to construct hyper linked tables of contents for a collection of 1,000 digitized books. This paper reviews the setup of the competition, the book collection used in the task, and the measures used for the evaluation. It further presents the outcome of the competition: an additional ground truth of 513 book tables of contents, contributed by 6 institutions, and the result performance of the 4 participating research teams. Antoine Doucet, Gabriella Kazai, Jean-Luc Meunier |
ICDAR | 2 |
| 2011 | Crowdsourcing for book search evaluation: impact of hit design on comparative system rankingabstractThe evaluation of information retrieval (IR) systems over special collections, such as large book repositories, is out of reach of traditional methods that rely upon editorial relevance judgments. Increasingly, the use of crowdsourcing to collect relevance labels has been regarded as a viable alternative that scales with modest costs. However, crowdsourcing suffers from undesirable worker practices and low quality contributions. In this paper we investigate the design and implementation of effective crowdsourcing tasks in the context of book search evaluation. We observe the impact of aspects of the Human Intelligence Task (HIT) design on the quality of relevance labels provided by the crowd. We assess the output in terms of label agreement with a gold standard data set and observe the effect of the crowdsourced relevance judgments on the resulting system rankings. This enables us to observe the effect of crowdsourcing on the entire IR evaluation process. Using the test set and experimental runs from the INEX 2010 Book Track, we find that varying the HIT design, and the pooling and document ordering strategies leads to considerable differences in agreement with the gold set labels. We then observe the impact of the crowdsourced relevance label sets on the relative system rankings using four IR performance metrics. System rankings based on MAP and Bpref remain less affected by different label sets while the [email protected] and [email protected] lead to dramatically different system rankings, especially for labels acquired from HITs with weaker quality controls. Overall, we find that crowdsourcing can be an effective tool for the evaluation of IR systems, provided that care is taken when designing the HITs. Gabriella Kazai, Jaap Kamps, Marijn Koolen, Natasa Milic-Frayling |
SIGIR | 1 |
| 2011 | Introduction to special issue on the second international conference on the theory of information retrieval
Leif Azzopardi, Dawei Song 0001, Gabriella Kazai, Stephen E. Robertson, Stefan M. Rüger, Milad Shokouhi, Emine Yilmaz |
Inf. Retr. | 3 |
| 2010 | 3rd BooksOnline workshop: research advances in large digital book repositories and complementary mediaabstractThe goal of the 3rd BooksOnline Workshop is to bring together researchers and industry practitioners in information retrieval, digital libraries, e-books, human computer interaction, publishing industry, and online book services to foster progress on addressing challenges and exploring opportunities around large collections of digital books and complementary media. Towards this goal, the workshop programme consists of contributions both from academia and industry, including two keynote talks: James Crawford from Google Books and John Mark Ockerbloom from the University of Pennsylvania. Gabriella Kazai, Peter Brusilovsky |
CIKM | 1 |
| 2010 | Connecting the local and the online in information managementabstractWith the popularity of social media sites, digital content is increasingly stored and managed online. At the same time, the desktop and local storage continues to provide a personal environment in which users perform their daily tasks. Thus, to accomplish their tasks, users need to continuously switch between local and remote resources and applications, often carrying the burden of coordinating and synchronizing these in a consistent way. In this demonstration, we describe a system, called ScholarLynk, that bridges the local and online worlds and allows users to manage both local and online resources in a uniform way and in collaboration with others. Gabriella Kazai, Natasa Milic-Frayling, Tim Haughton, Natalia Manola, Katerina Iatropoulou, Antonis Lempesis, Paolo Manghi, Marko Mikulicic |
CIKM | 1 |
| 2010 | Recent Developments in Information Retrieval
Cathal Gurrin, Yulan He 0001, Gabriella Kazai, Udo Kruschwitz, Suzanne Little, Thomas Roelleke, Stefan M. Rüger, C. J. van Rijsbergen |
ECIR | 3 |
| 2009 | Measuring system performance and topic discernment using generalized adaptive-weight meanabstractStandard approaches to evaluating and comparing information retrieval systems compute simple averages of performance statistics across individual topics to measure the overall system performance. However, topics vary in their ability to differentiate among systems based on their retrieval performance. At the same time, systems that perform well on discriminative queries demonstrate notable qualities that should be reflected in the systems' evaluation and ranking. This motivated research on alternative performance measures that are sensitive to the discriminative value of topics and the performance consistency of systems. In this paper we provide a mathematical formulation of a performance measure that postulates the dependence between the system and topic characteristics. We propose the Generalized Adaptive-Weight Mean (GAWM) measure and show how it can be computed as a fixed point of a function for which the Brouwer Fixed Point Theorem applies. This guarantees the existence of a scoring scheme that satisfies the starting axioms and can be used for ranking of both systems and topics. We apply our method to TREC experiments and compare the GAWM with the standard averages used in TREC. Chung Tong Lee, Vishwa Vinay, Eduarda Mendes Rodrigues, Gabriella Kazai, Natasa Milic-Frayling, Aleksandar Ignjatovic |
CIKM | 4 |
| 2009 | ICDAR 2009 Book Structure Extraction CompetitionabstractThis paper introduces the Book Structure Extraction competition run at ICDAR 2009. The goal of the competition is to evaluate and compare automatic techniques for deriving structure information from digitized books, which could then be used to aid navigation inside the books. More specifically, the task that participants are faced with is to construct hyperlinked tables of contents for a collection of 1,000 digitized books. This paper describes the setup of the competition, the book collection used in the task, and the proposed measures for the evaluation. Results of the evaluation will be presented at the ICDAR 2009 conference and will be published in the INEX 2009 proceedings. Antoine Doucet, Gabriella Kazai, Bodin Dresevic, Aleksandar Uzelac, Bogdan Radakovic, Nikola Todic |
ICDAR | 2 |
| 2009 | Towards methods for the collective gathering and quality control of relevance assessmentsabstractGrowing interest in online collections of digital books and video content motivates the development and optimization of adequate retrieval systems. However, traditional methods for collecting relevance assessments to tune system performance are challenged by the nature of digital items in such collections, where assessors are faced with a considerable effort to review and assess content by extensive reading, browsing, and within-document searching. The extra strain is caused by the length and cohesion of the digital item and the dispersion of topics within it. We propose a method for the collective gathering of relevance assessments using a social game model to instigate participants' engagement. The game provides incentives for assessors to follow a predefined review procedure and makes provisions for the quality control of the collected relevance judgments. We discuss the approach in detail, and present the results of a pilot study conducted on a book corpus to validate the approach. Our analysis reveals intricate relationships between the affordances of the system, the incentives of the social game, and the behavior of the assessors. We show that the proposed game design achieves two designated goals: the incentive structure motivates endurance in assessors and the review process encourages truthful assessment. Gabriella Kazai, Natasa Milic-Frayling, Jamie Costello |
SIGIR | 1 |
| 2009 | Model for Voter Scoring and Best Answer Selection in Community Q&A ServicesabstractCommunity Question Answering (cQA) services, such as Yahoo! Answers and MSN QnA, facilitate knowledge sharing through question answering by an online community of users. These services include incentive mechanisms to entice participation and self-regulate the quality of the content contributed by the users. In order to encourage quality contributions, community members are asked to nominate the ‘best’ among the answers provided to a question. The service then awards extra points to the author who provided the winning answer and to the voters who cast their vote for that answer. The best answers are typically selected by plurality voting, a scheme that is simple, yet vulnerable to random voting and collusion. We propose a weighted voting method that incorporates information about the voters’ behavior. It assigns a score to each voter that captures the level of agreement with other voters. It uses the voter scores to aggregate the votes and determine the best answer. The mathematical formulation leads to the application of the Brouwer Fixed Point Theorem which guarantees the existence of a voter scoring function that satisfies the starting axiom. We demonstrate the robustness of our approach through simulations and analysis of real cQA service data. Chung Tong Lee, Eduarda Mendes Rodrigues, Gabriella Kazai, Natasa Milic-Frayling, Aleksandar Ignjatovic |
Web Intelligence | 3 |
| 2009 | Wikipedia pages as entry points for book searchabstractA lot of the world's knowledge is stored in books, which, as a result of recent mass-digitisation efforts, are increasingly available online. Search engines, such as Google Books, provide mechanisms for searchers to enter this vast knowledge space using queries as entry points. In this paper, we view Wikipedia as a summary of this world knowledge and aim to use this resource to guide users to relevant books. Thus, we investigate possible ways of using Wikipedia as an intermediary between the user's query and a collection of books being searched. We experiment with traditional query expansion techniques, exploiting Wikipedia articles as rich sources of information that can augment the user's query. We then propose a novel approach based on link distance in an extended Wikipedia graph: we associate books with Wikipedia pages that cite these books and use the link distance between these nodes and the pages that match the user query as an estimation of a book's relevance to the query. Our results show that a) classical query expansion using terms extracted from query pages leads to increased precision, and b) link distance between query and book pages in Wikipedia provides a good indicator of relevance that can boost the retrieval score of relevant books in the result ranking of a book search engine. Marijn Koolen, Gabriella Kazai, Nick Craswell |
WSDM | 2 |
| 2008 | Structural relevance: a common basis for the evaluation of structured document retrievalabstractThis paper presents a unified framework for the evaluation of a range of structured document retrieval (SDR) approaches and tasks. The framework is based on a model of tree retrieval, evaluated using a novel extension of the Structural elevance (SR) measure. The measure replaces the assumption of independence in traditional information retrieval (IR) with a notion of redundancy that takes into account the user navigation inside documents while seeking relevant information. Unlike existing metrics for SDR, our proposed framework does not require the computation of an ideal ranking which has, thus far, prevented the practical application of such measures. Instead, SR builds on a Markovian model of user navigation that can be estimated through the use of structural summaries. The results of this paper (supported by experimental validation using INEX data) show that SR defined over a tree retrieval model can provide a common basis for the evaluation of SDR approaches across various structured search tasks. Mir Sadek Ali, Mariano P. Consens, Gabriella Kazai, Mounia Lalmas-Roelleke |
CIKM | 3 |
| 2008 | Trust, authority and popularity in social information retrievalabstractWe present a social information retrieval (SIR) model comprising the social network of actors (e.g., authors, publishers, consumers), the graph representing relations in data (e.g., publications), and the links between the social and data network that reflect activities in the network such as search, authoring, annotation, etc. Building on this hybrid network, we describe relevance in terms of the trust propagated through the network and rendered onto a given item. In particular, relevance is a function of the approval votes from the associated sub-graph and the reputation of the sub-graph nodes. We explore a model that differentiates between approval from actors who are perceived authorities by the user and the approval by a wider community, representing the popular opinion. Gabriella Kazai, Natasa Milic-Frayling |
CIKM | 1 |
| 2008 | Book Search Experiments: Investigating IR Methods for the Indexing and Retrieval of Books
Hengzhi Wu, Gabriella Kazai, Michael J. Taylor 0001 |
ECIR | 2 |
| 2006 | A general matrix framework for modelling Information Retrieval
Thomas Roelleke, Theodora Tsikrika, Gabriella Kazai |
Inf. Process. Manag. | 3 |
| 2006 | Evaluating the effectiveness of content-oriented XML retrieval methods
Norbert Gövert, Norbert Fuhr, Mounia Lalmas-Roelleke, Gabriella Kazai |
Inf. Retr. | 4 |
| 2006 | eXtended cumulated gain measures for the evaluation of content-oriented XML retrievalabstractWe propose and evaluate a family of measures, the eXtended Cumulated Gain (XCG) measures, for the evaluation of content-oriented XML retrieval approaches. Our aim is to provide an evaluation framework that allows the consideration of dependency among XML document components. In particular, two aspects of dependency are considered: (1) near-misses, which are document components that are structurally related to relevant components, such as a neighboring paragraph or container section, and (2) overlap, which regards the situation wherein the same text fragment is referenced multiple times, for example, when a paragraph and its container section are both retrieved. A further consideration is that the measures should be flexible enough so that different models of user behavior may be instantiated within. Both system- and user-oriented aspects are investigated and both recall and precision-like qualities are measured. We evaluate the reliability of the proposed measures based on the INEX 2004 test collection. For example, the effects of assessment variation and topic set size on evaluation stability are investigated, and the upper and lower bounds of expected error rates are established. The evaluation demonstrates that the XCG measures are stable and reliable, and in particular, that the novel measures of effort-precision and gain-recall ( ep / gr ) show comparable behavior to established IR measures like precision and recall. Gabriella Kazai, Mounia Lalmas-Roelleke |
ACM Trans. Inf. Syst. | 1 |
| 2004 | A Study of the Assessment of Relevance for the INEX'02 Test Collection
Gabriella Kazai, Sherezad Masood, Mounia Lalmas-Roelleke |
ECIR | 1 |
| 2004 | The overlap problem in content-oriented XML retrieval evaluationabstractWithin the INitiative for the Evaluation of XML Retrieval(INEX) a number of metrics to evaluate the effectiveness of content-oriented XML retrieval approaches were developed. Although these metrics provide a solution towards addressing the problem of overlapping result elements, they do not consider the problem of overlapping reference components within the recall-base, thus leading to skewed effectiveness scores. We propose alternative metrics that aim to provide a solution to both overlap issues. Gabriella Kazai, Mounia Lalmas-Roelleke, Arjen P. de Vries |
SIGIR | 1 |
| 2004 | A report on the first year of the INitiative for the Evaluation of XML retrievalabstractAbstract The INitiative for the Evaluation of XML retrieval (INEX) aims at providing an infrastructure to evaluate the effectiveness of content‐oriented XML retrieval systems. To this end, in the first round of INEX in 2002, a test collection of real world XML documents along with a set of topics and respective relevance assessments have been created with the collaboration of 36 participating organizations. In this article, we provide an overview of the first round of the INEX initiative. Gabriella Kazai, Mounia Lalmas-Roelleke, Norbert Fuhr, Norbert Gövert |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2003 | Construction of a Test Collection for the Focussed Retrieval of Structured Documents
Gabriella Kazai, Mounia Lalmas-Roelleke, Jane Reid |
ECIR | 1 |
| 2002 | The Accessibility Dimension for Structured Document Retrieval
Thomas Roelleke, Mounia Lalmas-Roelleke, Gabriella Kazai, Ian Ruthven, Stefan Quicker |
ECIR | 3 |
| 2002 | Focussed Structured Document Retrieval
Gabriella Kazai, Mounia Lalmas-Roelleke, Thomas Roelleke |
SPIRE | 1 |
| 2001 | The HySpirit Retrieval PlatformabstractNo abstract available. Thomas Roelleke, Ralf Lübeck, Gabriella Kazai |
SIGIR | 3 |
| 2001 | A Model for the Representation and Focussed Retrieval of Structured Documents Based on Fuzzy AggregationabstractEffective retrieval of structured documents should exploit the content and structural knowledge associated with the documents. This knowledge can be used to focus retrieval to the best entry points: document components that contain relevant information, and from which users can browse to retrieve further relevant components. To enable this, suitable representation methods must be developed. This paper presents a model for representing structured documents to allow for their focussed retrieval. The model is founded on fuzzy aggregation, an approach based on the fuzzy representation of linguistic quantifiers and ordered weighted averaging operators. By defining the representation of a document component as the fuzzy aggregation of its related components, we arrive at a document representation that supports the selection of best entry points. 1 Gabriella Kazai, Mounia Lalmas-Roelleke, Thomas Roelleke |
SPIRE | 1 |