VLDB 2026 Research / reviewers in the wild / expert
Stefano Mizzaro
dblp:74/4701
· DBLP profile ↗
69ranked-venue papers in the field
10as first author
20since 2021 · last 2026
0000-0002-2852-168XORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 62 (9 first)Database Systems & Data Management · 3Data Mining & Knowledge Discovery · 3Knowledge Engineering, Semantic Web & Information Systems · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Effect of Document Summarization on LLM-Based Relevance Judgments
Samaneh Mohtadi, Kevin Roitero, Stefano Mizzaro, Gianluca Demartini |
ECIR (2) | 3 |
| 2026 | Analyzing AI Evaluation Benchmarks Through Information Retrieval and Network Science
Gaia Simeoni, Michael Soprano, Riccardo Lunardi, Kevin Roitero, Stefano Mizzaro |
ECIR (2) | 5 |
| 2026 | Large Language Models as Assessors: On the Impact of Relevance Scales
Riccardo Zamolo, Riccardo Lunardi, Michael Soprano, Gianluca Demartini, Stefano Mizzaro, Kevin Roitero |
ECIR (2) | 5 |
| 2025 | A Comparative Analysis of Retrieval-Augmented Generation and Crowdsourcing for Fact-Checking
Francesco Bombassei De Bona, David La Barbera, Stefano Mizzaro, Kevin Roitero |
ECIR (3) | 3 |
| 2025 | Leveraging LLMs for Energy Forecasting: The AcegasApsAmga Case Study
Kevin Roitero, Andrea Zancola, Vincenzo Della Mea, Stefano Mizzaro |
ECIR (5) | 4 |
| 2025 | PILs of Knowledge: A Synthetic Benchmark for Evaluating Question Answering Systems in HealthcareabstractPatient Information Leaflets (PILs) provide essential information about medication usage, side effects, precautions, and interactions, making them a valuable resource for Question Answering (QA) systems in healthcare. However, no dedicated benchmark currently exists to evaluate QA systems specifically on PILs, limiting progress in this domain. To address this gap, we introduce a fact-supported synthetic benchmark composed of multiple-choice questions and answers generated from real PILs. We construct the benchmark using a fully automated pipeline that leverages multiple Large Language Models (LLMs) to generate diverse, realistic, and contextually relevant question-answer pairs. The benchmark is publicly released as a standardized evaluation framework for assessing the ability of LLMs to process and reason over PIL content. To validate its effectiveness, we conduct an initial evaluation with state-of-the-art LLMs, showing that the benchmark presents a realistic and challenging task, making it a valuable resource for advancing QA research in the healthcare domain. Riccardo Lunardi, Michael Soprano, Paolo Coppola 0001, Vincenzo Della Mea, Stefano Mizzaro, Kevin Roitero |
SIGIR | 5 |
| 2025 | Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-CheckingabstractEvaluating the truthfulness of online content is critical for combating misinformation. This study examines the efficiency and effectiveness of crowdsourced truthfulness assessments through a comparative analysis of two approaches: one involving full-length webpages as evidence for each claim, and another using summaries for each evidence document generated with an LLM. Using an A/B testing setting, we engage a diverse pool of participants tasked with evaluating the truthfulness of statements under these conditions. Kevin Roitero, Dustin Wright 0001, Michael Soprano, Isabelle Augenstein, Stefano Mizzaro |
SIGIR | 5 |
| 2025 | The Magnitude of Truth: On Using Magnitude Estimation for Truthfulness AssessmentabstractAssessing the truthfulness of information is a critical task in fact-checking, and is typically performed using binary or coarse ordinal scales (2-6 levels), though fine-grained scales (e.g., 100 levels) have also been explored. Magnitude Estimation (ME) takes this approach further by allowing assessors to assign any value in the range (0, + ∞). However, it introduces challenges, including the need for aggregation of assessments from individuals with different interpretations of the scale. Despite these, its successful applications in other domains suggest its potential suitability for truthfulness assessment. We conduct a crowdsourcing study by collecting assessments on claims sourced from the PolitiFact fact-checking organization using ME. To the best of our knowledge, this is the first systematic investigation of ME in the context of truthfulness assessment. Our results show that while aggregation methods significantly impact assessment quality, optimal aggregation strategies yield accuracy and reliability comparable to traditional scales. More importantly, ME allows capturing subtle differences in truthfulness, offering richer insights than conventional coarse-grained scales. Michael Soprano, Denis Eduard Tapu, David La Barbera, Kevin Roitero, Stefano Mizzaro |
SIGIR | 5 |
| 2024 | Generative AI for Energy: Multi-Horizon Power Consumption Forecasting using Large Language ModelsabstractWe leverage generative NLP-based models, specifically Transformer-Based models, for multi-horizon univariate and multivariate power consumption forecasting. We apply our approach to various datasets, focusing on short-term (1 day) and long-term (1 week) forecasts. We test several lag configurations with and without additional contextual information and achieve promising results. We evaluate the forecasts' effectiveness using a range of metrics, and aggregate the results on a monthly basis for a comprehensive understanding of the performance throughout the year. Kevin Roitero, Gianluca D'Abrosca, Andrea Zancola, Vincenzo Della Mea, Stefano Mizzaro |
CIKM | 5 |
| 2024 | Combining Large Language Models and Crowdsourcing for Hybrid Human-AI Misinformation DetectionabstractResearch on misinformation detection has primarily focused either on furthering Artificial Intelligence (AI) for automated detection or on studying humans' ability to deliver an effective crowdsourced solution. Each of these directions however shows different benefits. This motivates our work to study hybrid human-AI approaches jointly leveraging the potential of large language models and crowdsourcing, which is understudied to date. We propose novel combination strategies Model First, Worker First, and Meta Vote, which we evaluate along with baseline methods such as mean, median, hard- and soft-voting. Using 120 statements from the PolitiFact dataset, and a combination of state-of-the-art AI models and crowdsourced assessments, we evaluate the effectiveness of these combination strategies. Results suggest that the effectiveness varies with scales granularity, and that combining AI and human judgments enhances truthfulness assessments' effectiveness and robustness. Xia Zeng, David La Barbera, Kevin Roitero, Arkaitz Zubiaga, Stefano Mizzaro |
SIGIR | 5 |
| 2024 | Crowdsourced Fact-checking: Does It Actually Work?abstractThere is an important ongoing effort aimed to tackle misinformation and to perform reliable fact-checking by employing human assessors at scale, with a crowdsourcing-based approach. Previous studies on the feasibility of employing crowdsourcing for the task of misinformation detection have provided inconsistent results: some of them seem to confirm the effectiveness of crowdsourcing for assessing the truthfulness of statements and claims, whereas others fail to reach an effectiveness level higher than automatic machine learning approaches, which are still unsatisfactory. In this paper, we aim at addressing such inconsistency and understand if truthfulness assessment can indeed be crowdsourced effectively. To do so, we build on top of previous studies; we select some of those reporting low effectiveness levels, we highlight their potential limitations, and we then reproduce their work attempting to improve their setup to address those limitations. We employ various approaches, data quality levels, and agreement measures to assess the reliability of crowd workers when assessing the truthfulness of (mis)information. Furthermore, we explore different worker features and compare the results obtained with different crowds. According to our findings, crowdsourcing can be used as an effective methodology to tackle misinformation at scale. When compared to previous studies, our results indicate that a significantly higher agreement between crowd workers and experts can be obtained by using a different, higher-quality, crowdsourcing platform and by improving the design of the crowdsourcing task. Also, we find differences concerning task and worker features and how workers provide truthfulness assessments. David La Barbera, Eddy Maddalena, Michael Soprano, Kevin Roitero, Gianluca Demartini, Davide Ceolin, Damiano Spina, Stefano Mizzaro |
Inf. Process. Manag. | 8 |
| 2024 | Cognitive Biases in Fact-Checking and Their Countermeasures: A ReviewabstractThe increase of the amount of misinformation spread every day online is a huge threat to the society. Organizations and researchers are working to contrast this misinformation plague. In this setting, human assessors are indispensable to correctly identify, assess and/or revise the truthfulness of information items, i.e., to perform the fact-checking activity. Assessors, as humans, are subject to systematic errors that might interfere with their fact-checking activity. Among such errors, cognitive biases are those due to the limits of human cognition. Although biases help to minimize the cost of making mistakes, they skew assessments away from an objective perception of information. Cognitive biases, hence, are particularly frequent and critical, and can cause errors that have a huge potential impact as they propagate not only in the community, but also in the datasets used to train automatic and semi-automatic machine learning models to fight misinformation. In this work, we present a review of the cognitive biases which might occur during the fact-checking process. In more detail, inspired by PRISMA – a methodology used for systematic literature reviews – we manually derive a list of 221 cognitive biases that may affect human assessors. Then, we select the 39 biases that might manifest during the fact-checking process, we group them into categories, and we provide a description. Finally, we present a list of 11 countermeasures that can be adopted by researchers, practitioners, and organizations to limit the effect of the identified cognitive biases on the fact-checking activity. Michael Soprano, Kevin Roitero, David La Barbera, Davide Ceolin, Damiano Spina, Gianluca Demartini, Stefano Mizzaro |
Inf. Process. Manag. | 7 |
| 2024 | How Many Crowd Workers Do I Need? On Statistical Power when Crowdsourcing Relevance JudgmentsabstractTo scale the size of Information Retrieval collections, crowdsourcing has become a common way to collect relevance judgments at scale. Crowdsourcing experiments usually employ 100–10,000 workers, but such a number is often decided in a heuristic way. The downside is that the resulting dataset does not have any guarantee of meeting predefined statistical requirements as, for example, have enough statistical power to be able to distinguish in a statistically significant way between the relevance of two documents. We propose a methodology adapted from literature on sound topic set size design, based on t-test and ANOVA, which aims at guaranteeing the resulting dataset to meet a predefined set of statistical requirements. We validate our approach on several public datasets. Our results show that we can reliably estimate the recommended number of workers needed to achieve statistical power, and that such estimation is dependent on the topic, while the effect of the relevance scale is limited. Furthermore, we found that such estimation is dependent on worker features such as agreement. Finally, we describe a set of practical estimation strategies that can be used to estimate the worker set size, and we also provide results on the estimation of document set sizes. Kevin Roitero, David La Barbera, Michael Soprano, Gianluca Demartini, Stefano Mizzaro, Tetsuya Sakai |
ACM Trans. Inf. Syst. | 5 |
| 2023 | A unifying and general account of fairness measurement in recommender systemsabstractFairness is fundamental to all information access systems, including recommender systems. However, the landscape of fairness definition and measurement is quite scattered with many competing definitions that are partial and often incompatible. There is much work focusing on specific – and different – notions of fairness and there exist dozens of metrics of fairness in the literature, many of them redundant and most of them incompatible. In contrast, to our knowledge, there is no formal framework that covers all possible variants of fairness and allows developers to choose the most appropriate variant depending on the particular scenario. In this paper, we aim to define a general, flexible, and parameterizable framework that covers a whole range of fairness evaluation possibilities. Instead of modeling the metrics based on an abstract definition of fairness, the distinctive feature of this study compared to the current state of the art is that we start from the metrics applied in the literature to obtain a unified model by generalization. The framework is grounded on a general work hypothesis: interpreting the space of users and items as a probabilistic sample space, two fundamental measures in information theory (Kullback–Leibler Divergence and Mutual Information) can capture the majority of possible scenarios for measuring fairness on recommender system outputs. In addition, earlier research on fairness in recommender systems could be viewed as single-sided, trying to optimize some form of equity across either user groups or provider/procurer groups, without considering the user/item space in conjunction, thereby overlooking/disregarding the interplay between user and item groups. Instead, our framework includes the notion of statistical independence between user and item groups. We finally validate our approach experimentally on both synthetic and real data according to a wide range of state-of-the-art recommendation algorithms and real-world data sets, showing that with our framework we can measure fairness in a general, uniform, and meaningful way. Enrique Amigó, Yashar Deldjoo, Stefano Mizzaro, Alejandro Bellogín |
Inf. Process. Manag. | 3 |
| 2023 | What is My Problem? Identifying Formal Tasks and Metrics in Data Mining on the Basis of Measurement TheoryabstractThe design and analysis of experimental research in Data Mining (DM) is anchored in a correct choice of the type of task addressed (clustering, classification, regression, etc.). However, although DM is a relatively mature discipline, there is no consensus yet about what is the taxonomy of DM tasks, which are their formal characteristics, and their corresponding metrics. In this paper, we formalize DM tasks in terms of Measurement Theory, which is a cornerstone of quantitative research in many disciplines, but has not yet been incorporated (in a consensual way) into some areas of Computer Science, including DM. The proposed formal framework provides a methodology to precisely define DM tasks for any given scenario and identify appropriate metrics. We validate this framework via (i) its coverage of existing DM tasks, (ii) its capability to group existing metrics into families, and (iii) its coverage of actual DM research problems, using about 250 papers from ACM KDD 2019 and IEEE ICDM 2019 conferences as reference sample. Enrique Amigó, Julio Gonzalo 0001, Stefano Mizzaro |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Ranking Interruptus: When Truncated Rankings Are Better and How to Measure ThatabstractMost of information retrieval effectiveness evaluation metrics assume that systems appending irrelevant documents at the bottom of the ranking are as effective as (or not worse than) systems that have a stopping criteria to 'truncate' the ranking at the right position to avoid retrieving those irrelevant documents at the end. It can be argued, however, that such truncated rankings are more useful to the end user. It is thus important to understand how to measure retrieval effectiveness in this scenario. In this paper we provide both theoretical and experimental contributions. We first define formal properties to analyze how effectiveness metrics behave when evaluating truncated rankings. Our theoretical analysis shows that de-facto standard metrics do not satisfy desirable properties to evaluate truncated rankings: only Observational Information Effectiveness (OIE) -- a metric based on Shannon's information theory -- satisfies them all. We then perform experiments to compare several metrics on nine TREC datasets. According to our experimental results, the most appropriate metrics for truncated rankings are OIE and a novel extension of Rank-Biased Precision that adds a user effort factor penalizing the retrieval of irrelevant documents. Enrique Amigó, Stefano Mizzaro, Damiano Spina |
SIGIR | 2 |
| 2022 | Crowd_Frame: A Simple and Complete Framework to Deploy Complex Crowdsourcing Tasks Off-the-shelfabstractDue to their relatively low cost and ability to scale, crowdsourcing based approaches are widely used to collect a large amount of human annotated data. To this aim, multiple crowdsourcing platforms exist, where requesters can upload tasks and workers can carry them out and obtain payment in return. Such platforms share a task design and deploy workflow that is often counter-intuitive and cumbersome. To address this issue, we propose Crowd\_Frame, a simple and complete framework which allows to develop and deploy diverse types of complex crowdsourcing tasks in an easy and customizable way. We show the abilities of the proposed framework and we make it available to researchers and practitioners. Michael Soprano, Kevin Roitero, Francesco Bombassei De Bona, Stefano Mizzaro |
WSDM | 4 |
| 2022 | Preferences on a Budget: Prioritizing Document Pairs when Crowdsourcing Relevance JudgmentsabstractIn Information Retrieval (IR) evaluation, preference judgments are collected by presenting to the assessors a pair of documents and asking them to select which of the two, if any, is the most relevant. This is an alternative to the classic relevance judgment approach, in which human assessors judge the relevance of a single document on a scale; such an alternative allows to make relative rather than absolute judgments of relevance. While preference judgments are easier for human assessors to perform, the number of possible document pairs to be judged is usually so high that it makes it unfeasible to judge them all. Thus, following a similar idea to pooling strategies for single document relevance judgments where the goal is to sample the most useful documents to be judged, in this work we focus on analyzing alternative ways to sample document pairs to judge, in order to maximize the value of a fixed number of preference judgments that can feasibly be collected. Such value is defined as how well we can evaluate IR systems given a budget, that is, a fixed number of human preference judgments that may be collected. By relying on several datasets featuring relevance judgments gathered by means of experts and crowdsourcing, we experimentally compare alternative strategies to select document pairs and show how different strategies lead to different IR evaluation result quality levels. Our results show that, by using the appropriate procedure, it is possible to achieve good IR evaluation results with a limited number of preference judgments, thus confirming the feasibility of using preference judgments to create IR evaluation collections. Kevin Roitero, Alessandro Checco, Stefano Mizzaro, Gianluca Demartini |
WWW | 3 |
| 2021 | On the effect of relevance scales in crowdsourcing relevance assessments for Information Retrieval evaluation
Kevin Roitero, Eddy Maddalena, Stefano Mizzaro, Falk Scholer |
Inf. Process. Manag. | 3 |
| 2021 | The many dimensions of truthfulness: Crowdsourcing misinformation assessments on a multidimensional scale
Michael Soprano, Kevin Roitero, David La Barbera, Davide Ceolin, Damiano Spina, Stefano Mizzaro, Gianluca Demartini |
Inf. Process. Manag. | 6 |
| 2020 | The COVID-19 Infodemic: Can the Crowd Judge Recent Misinformation Objectively?abstractMisinformation is an ever increasing problem that is difficult to solve for the research community and has a negative impact on the society at large. Very recently, the problem has been addressed with a crowdsourcing-based approach to scale up labeling efforts: to assess the truthfulness of a statement, instead of relying on a few experts, a crowd of (non-expert) judges is exploited. We follow the same approach to study whether crowdsourcing is an effective and reliable method to assess statements truthfulness during a pandemic. We specifically target statements related to the COVID-19 health emergency, that is still ongoing at the time of the study and has arguably caused an increase of the amount of misinformation that is spreading online (a phenomenon for which the term "infodemic" has been used). By doing so, we are able to address (mis)information that is both related to a sensitive and personal issue like health and very recent as compared to when the judgment is done: two issues that have not been analyzed in related work.\n\nIn our experiment, crowd workers are asked to assess the truthfulness of statements, as well as to provide evidence for the assessments as a URL and a text justification. Besides showing that the crowd is able to accurately judge the truthfulness of the statements, we also report results on many different aspects, including: agreement among workers, the effect of different aggregation functions, of scales transformations, and of workers background / bias. We also analyze workers behavior, in terms of queries submitted, URLs found / selected, text justifications, and other behavioral data like clicks and mouse actions collected by means of an ad hoc logger. Kevin Roitero, Michael Soprano, Beatrice Portelli, Damiano Spina, Vincenzo Della Mea, Giuseppe Serra 0001, Stefano Mizzaro, Gianluca Demartini |
CIKM | 7 |
| 2020 | Crowdsourcing Truthfulness: The Impact of Judgment Scale and Assessor Bias
David La Barbera, Kevin Roitero, Gianluca Demartini, Stefano Mizzaro, Damiano Spina |
ECIR (2) | 4 |
| 2020 | Can The Crowd Identify Misinformation Objectively?: The Effects of Judgment Scale and Assessor's BackgroundabstractTruthfulness judgments are a fundamental step in the process of fighting misinformation, as they are crucial to train and evaluate classifiers that automatically distinguish true and false statements. Usually such judgments are made by experts, like journalists for political statements or medical doctors for medical statements. In this paper, we follow a different approach and rely on (non-expert) crowd workers. This of course leads to the following research question: Can crowdsourcing be reliably used to assess the truthfulness of information and to create large-scale labeled collections for information credibility systems? To address this issue, we present the results of an extensive study based on crowdsourcing: we collect thousands of truthfulness assessments over two datasets, and we compare expert judgments with crowd judgments, expressed on scales with various granularity levels. We also measure the political bias and the cognitive background of the workers, and quantify their effect on the reliability of the data provided by the crowd. Kevin Roitero, Michael Soprano, Shaoyang Fan, Damiano Spina, Stefano Mizzaro, Gianluca Demartini |
SIGIR | 5 |
| 2020 | Effectiveness evaluation without human relevance judgments: A systematic analysis of existing methods and of their combinations
Kevin Roitero, Andrea Brunello, Giuseppe Serra 0001, Stefano Mizzaro |
Inf. Process. Manag. | 4 |
| 2020 | Axiomatic thinking for information retrieval: introduction to special issue
Enrique Amigó, Hui Fang 0001, Stefano Mizzaro, ChengXiang Zhai |
Inf. Retr. J. | 3 |
| 2020 | On the nature of information access evaluation metrics: a unifying framework
Enrique Amigó, Stefano Mizzaro |
Inf. Retr. J. | 2 |
| 2020 | Fewer topics? A million topics? Both?! On topics subsets in test collections
Kevin Roitero, J. Shane Culpepper, Mark Sanderson, Falk Scholer, Stefano Mizzaro |
Inf. Retr. J. | 5 |
| 2019 | On Transforming Relevance ScalesabstractInformation Retrieval (IR) researchers have often used existing IR evaluation collections and transformed the relevance scale in which judgments have been collected, e.g., to use metrics that assume binary judgments like Mean Average Precision. Such scale transformations are often arbitrary (e.g., 0,1 mapped to 0 and 2,3 mapped to 1) and it is assumed that they have no impact on the results of IR evaluation. Moreover, the use of crowdsourcing to collect relevance judgments has become a standard methodology. When designing the crowdsourcing relevance judgment task, one of the decision to be made is the how granular the relevance scale used to collect judgments should be. Such decision has then repercussions on the metrics used to measure IR system effectiveness. In this paper we look at the effect of scale transformations in a systematic way. We perform extensive experiments to study the transformation of judgments from fine-grained to coarse-grained. We use different relevance judgments expressed on different relevance scales and either expressed by expert annotators or collected by means of crowdsourcing. The objective is to understand the impact of relevance scale transformations on IR evaluation outcomes and to draw conclusions on how to best transform judgments into a different scale, when necessary. Lei Han 0003, Kevin Roitero, Eddy Maddalena, Stefano Mizzaro, Gianluca Demartini |
CIKM | 4 |
| 2019 | Towards Stochastic Simulations of Relevance ProfilesabstractRecently proposed methods allow the generation of simulated scores representing the values of an effectiveness metric, but they do not investigate the generation of the actual lists of retrieved documents. In this paper we address this limitation: we present an approach that exploits an evolutionary algorithm and, given a metric score, creates a simulated relevance profile (i.e., a ranked list of relevance values) that produces that score. We show how the simulated relevance profiles are realistic under various analyses. Kevin Roitero, Andrea Brunello, Julián Urbano, Stefano Mizzaro |
CIKM | 4 |
| 2019 | On Topic Difficulty in IR Evaluation: The Effect of Systems, Corpora, and System ComponentsabstractIn a test collection setting, topic difficulty can be defined as the average effectiveness of a set of systems for a topic. In this paper we study the effects on the topic difficulty of: (i) the set of retrieval systems; (ii) the underlying document corpus; and (iii) the system components. By generalizing methods recently proposed to study system component factor analysis, we perform a comprehensive analysis on topic difficulty and the relative effects of systems, corpora, and component interactions. Our findings show that corpora have the most significant effect on topic difficulty. Fabio Zampieri, Kevin Roitero, J. Shane Culpepper, Oren Kurland, Stefano Mizzaro |
SIGIR | 5 |
| 2019 | Mobile Search Behaviors: An In-Depth Analysis Based on Contexts, APPs, and Devices. Dan Wu and Shaobo Liang. Synthesis Lectures on Information Concepts, Retrieval, and Services. San Rafael, CA: Morgan & Claypool, 2018
Stefano Mizzaro, Ivan Scagnetto |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2018 | Are we on the Right Track?: An Examination of Information Retrieval MethodologiesabstractThe unpredictability of user behavior and the need for effectiveness make it difficult to define a suitable research methodology for Information Retrieval (IR). In order to tackle this challenge, we categorize existing IR methodologies along two dimensions: (1) empirical vs. theoretical, and (2) top-down vs. bottom-up. The strengths and drawbacks of the resulting categories are characterized according to 6 desirable aspects. The analysis suggests that different methodologies are complementary and therefore, equally necessary. The categorization of the 167 full papers published in the last SIGIR (2016 and 2017) and ICTIR (2017) conferences suggest that most of existing work is empirical bottom-up, suggesting lack of some desirable aspects. With the hope of improving IR research practice, we propose a general methodology for IR that integrates the strengths of existing research methods. Enrique Amigó, Hui Fang 0001, Stefano Mizzaro, ChengXiang Zhai |
SIGIR | 3 |
| 2018 | Query Performance Prediction and Effectiveness Evaluation Without Relevance Judgments: Two Sides of the Same CoinabstractSome methods have been developed for automatic effectiveness evaluation without relevance judgments. We propose to use those methods, and their combination based on a machine learning approach, for query performance prediction. Moreover, since predicting average precision as it is usually done in query performance prediction literature is sensitive to the reference system that is chosen, we focus on predicting the average of average precision values over several systems. Results of an extensive experimental evaluation on ten TREC collections show that our proposed methods outperform state-of-the-art query performance predictors. Stefano Mizzaro, Josiane Mothe, Kevin Roitero, Md. Zia Ullah |
SIGIR | 1 |
| 2018 | On Fine-Grained Relevance ScalesabstractIn Information Retrieval evaluation, the classical approach of adopting binary relevance judgments has been replaced by multi-level relevance judgments and by gain-based metrics leveraging such multi-level judgment scales. Recent work has also proposed and evaluated unbounded relevance scales by means of Magnitude Estimation (ME) and compared them with multi-level scales. While ME brings advantages like the ability for assessors to always judge the next document as having higher or lower relevance than any of the documents they have judged so far, it also comes with some drawbacks. For example, it is not a natural approach for human assessors to judge items as they are used to do on the Web (e.g., 5-star rating). In this work, we propose and experimentally evaluate a bounded and fine-grained relevance scale having many of the advantages and dealing with some of the issues of ME. We collect relevance judgments over a 100-level relevance scale (S100) by means of a large-scale crowdsourcing experiment and compare the results with other relevance scales (binary, 4-level, and ME) showing the benefit of fine-grained scales over both coarse-grained and unbounded scales as well as highlighting some new results on ME. Our results show that S100 maintains the flexibility of unbounded scales like ME in providing assessors with ample choice when judging document relevance (i.e., assessors can fit relevance judgments in between of previously given judgments). It also allows assessors to judge on a more familiar scale (e.g., on 10 levels) and to perform efficiently since the very first judging task. Kevin Roitero, Eddy Maddalena, Gianluca Demartini, Stefano Mizzaro |
SIGIR | 4 |
| 2018 | IRevalOO: An Object Oriented Framework for Retrieval EvaluationabstractWe propose IRevalOO, a flexible Object Oriented framework that (i) can be used as-is as a replacement of the widely adopted trec\_eval software, and (ii) can be easily extended (or "instantiated'', in framework terminology) to implement different scenarios of test collection based retrieval evaluation. Instances of IRevalOO can provide a usable and convenient alternative to the state-of-the-art software commonly used by different initiatives (TREC, NTCIR, CLEF, FIRE, etc.). Also, those instances can be easily adapted to satisfy future customization needs of researchers, as: implementing and experimenting with new metrics, even based on new notions of relevance; using different formats for system output and "qrels''; and in general visualizing, comparing, and managing retrieval evaluation results. Kevin Roitero, Eddy Maddalena, Yannick Ponte, Stefano Mizzaro |
SIGIR | 4 |
| 2018 | Effectiveness Evaluation with a Subset of Topics: A Practical ApproachabstractSeveral researchers have proposed to reduce the number of topics used in TREC-like initiatives. One research direction that has been pursued is what is the optimal topic subset of a given cardinality that evaluates the systems/runs in the most accurate way. Such a research direction has been so far mainly theoretical, with almost no indication on how to select the few good topics in practice. We propose such a practical criterion for topic selection: we rely on the methods for automatic system evaluation without relevance judgments, and by running some experiments on several TREC collections we show that the topics selected on the basis of those evaluations are indeed more informative than random topics. Kevin Roitero, Michael Soprano, Stefano Mizzaro |
SIGIR | 3 |
| 2017 | Human-Based Query Difficulty Prediction
Adrian-Gabriel Chifu, Sébastien Déjean, Stefano Mizzaro, Josiane Mothe |
ECIR | 3 |
| 2017 | Do Easy Topics Predict Effectiveness Better Than Difficult Topics?
Kevin Roitero, Eddy Maddalena, Stefano Mizzaro |
ECIR | 3 |
| 2017 | Let's Agree to Disagree: Fixing Agreement Measures for CrowdsourcingabstractIn the context of micro-task crowdsourcing, each task is usually performed by several workers. This allows researchers to leverage measures of the agreement among workers on the same task, to estimate the reliability of collected data and to better understand answering behaviors of the participants. While many measures of agreement between annotators have been proposed, they are known for suffering from many problems and abnormalities. In this paper, we identify the main limits of the existing agreement measures in the crowdsourcing context, both by means of toy examples as well as with real-world crowdsourcing data, and propose a novel agreement measure based on probabilistic parameter estimation which overcomes such limits. We validate our new agreement measure and show its flexibility as compared to the existing agreement measures. Alessandro Checco, Kevin Roitero, Eddy Maddalena, Stefano Mizzaro, Gianluca Demartini |
HCOMP | 4 |
| 2017 | Axiomatic Thinking for Information Retrieval: And Related TasksabstractThis is the first workshop on the emerging interdisciplinary research area of applying axiomatic thinking to information retrieval (IR) and related tasks. The workshop aims to help foster collaboration of researchers working on different perspectives of axiomatic thinking and encourage discussion and research on general methodological issues related to applying axiomatic thinking to IR and related tasks. Enrique Amigó, Hui Fang 0001, Stefano Mizzaro, ChengXiang Zhai |
SIGIR | 3 |
| 2017 | On Crowdsourcing Relevance Magnitudes for Information Retrieval EvaluationabstractMagnitude estimation is a psychophysical scaling technique for the measurement of sensation, where observers assign numbers to stimuli in response to their perceived intensity. We investigate the use of magnitude estimation for judging the relevance of documents for information retrieval evaluation, carrying out a large-scale user study across 18 TREC topics and collecting over 50,000 magnitude estimation judgments using crowdsourcing. Our analysis shows that magnitude estimation judgments can be reliably collected using crowdsourcing, are competitive in terms of assessor cost, and are, on average, rank-aligned with ordinal judgments made by expert relevance assessors. We explore the application of magnitude estimation for IR evaluation, calibrating two gain-based effectiveness metrics, nDCG and ERR, directly from user-reported perceptions of relevance. A comparison of TREC system effectiveness rankings based on binary, ordinal, and magnitude estimation relevance shows substantial variation; in particular, the top systems ranked using magnitude estimation and ordinal judgments differ substantially. Analysis of the magnitude estimation scores shows that this effect is due in part to varying perceptions of relevance: different users have different perceptions of the impact of relative differences in document relevance. These results have direct implications for IR evaluation, suggesting that current assumptions about a single view of relevance being sufficient to represent a population of users are unlikely to hold. Eddy Maddalena, Stefano Mizzaro, Falk Scholer, Andrew Turpin |
ACM Trans. Inf. Syst. | 2 |
| 2016 | Crowdsourcing Relevance Assessments: The Unexpected Benefits of Limiting the Time to JudgeabstractCrowdsourcing has become an alternative approach to collect relevance judgments at scale thanks to the availability of crowdsourcing platforms and quality control techniques that allow to obtain reliable results. Previous work has used crowdsourcing to ask multiple crowd workers to judge the relevance of a document with respect to a query and studied how to best aggregate multiple judgments of the same topic-document pair. This paper addresses an aspect that has been rather overlooked so far: we study how the time available to express a relevance judgment affects its quality. We also discuss the quality loss of making crowdsourced relevance judgments more efficient in terms of time taken to judge the relevance of a document. We use standard test collections to run a battery of experiments on the crowdsourcing platform CrowdFlower, studying how much time crowd workers need to judge the relevance of a document and at what is the effect of reducing the available time to judge on the overall quality of the judgments. Our extensive experiments compare judgments obtained under different types of time constraints with judgments obtained when no time constraints were put on the task. We measure judgment quality by different metrics of agreement with editorial judgments. Experimental results show that it is possible to reduce the cost of crowdsourced evaluation collection creation by reducing the time available to perform the judgments with no loss in quality. Most importantly, we observed that the introduction of limits on the time available to perform the judgments improves the overall judgment quality. Top judgment quality is obtained with 25-30 seconds to judge a topic-document pair. Eddy Maddalena, Marco Basaldella, Dario De Nart, Dante Degl'Innocenti, Stefano Mizzaro, Gianluca Demartini |
HCOMP | 5 |
| 2016 | Why do you Think this Query is Difficult?: A User Study on Human Query PredictionabstractPredicting if a query will be difficult for a system is important to improve retrieval effectiveness by implementing specific processing. There have been several attempts to predict difficulty, both automatically and manually; but without high accuracy at a pre-retrieval stage. In this paper, we focus rather on understanding Why a query is perceived by humans as difficult. We ran two separated but related experiments in which we asked humans to provide both a query difficulty prediction and reasons to explain their prediction. Results show that: (i) reasons can be categorized into 4 classes; (ii) reasons can be framed into closed questions to be answered on a Likert scale; and (iii) some reasons correlate in a coherent way with the human predicted numerical difficulty. On the basis of these results it is possible to derive hints to be provided to help users when formulating their queries and to avoid them to rely on their wrong perception of difficulty. Stefano Mizzaro, Josiane Mothe |
SIGIR | 1 |
| 2015 | A Formal Approach to Effectiveness Metrics for Information Access: Retrieval, Filtering, and Clustering
Enrique Amigó, Julio Gonzalo 0001, Stefano Mizzaro |
ECIR | 3 |
| 2015 | Different Rankers on Different Subcollections
Timothy Jones 0001, Falk Scholer, Andrew Turpin, Stefano Mizzaro, Mark Sanderson |
ECIR | 4 |
| 2015 | Judging Relevance Using Magnitude Estimation
Eddy Maddalena, Stefano Mizzaro, Falk Scholer, Andrew Turpin |
ECIR | 2 |
| 2015 | Content-Based Similarity of Twitter Users
Stefano Mizzaro, Marco Pavan, Ivan Scagnetto |
ECIR | 1 |
| 2015 | Finding Important Locations: A Feature-Based ApproachabstractWe propose a novel approach to address the problem of the recognition of important locations. Our method is organised in two phases: first, a set of candidate stay points is identified by exploiting some state-of-the-art algorithms to filter the GPS-logs, then, the candidate stay points are mapped onto a feature space having as dimensions the area underlying the stay point, its intensity (the time spent in a location) and its frequency (the number of total visits). We conjecture that the features space allows to model aspects/measures that are more semantically related to users and better suited to reason about their similarities and differences than, e.g., Latitude, longitude, and timestamp. An experimental evaluation on the GeoLife public dataset confirms the effectiveness of our approach. Marco Pavan, Stefano Mizzaro, Ivan Scagnetto, Andrea Beggiato |
MDM (1) | 2 |
| 2015 | The Benefits of Magnitude Estimation Relevance Assessments for Information Retrieval EvaluationabstractMagnitude estimation is a psychophysical scaling technique for the measurement of sensation, where observers assign numbers to stimuli in response to their perceived intensity. We investigate the use of magnitude estimation for judging the relevance of documents in the context of information retrieval evaluation, carrying out a large-scale user study across 18 TREC topics and collecting more than 50,000 magnitude estimation judgments. Our analysis shows that on average magnitude estimation judgments are rank-aligned with ordinal judgments made by expert relevance assessors. An advantage of magnitude estimation is that users can chose their own scale for judgments, allowing deeper investigations of user perceptions than when categorical scales are used. Andrew Turpin, Falk Scholer, Stefano Mizzaro, Eddy Maddalena |
SIGIR | 3 |
| 2015 | Mobile crowdsourcing: four experiments on platforms and tasks
Vincenzo Della Mea, Eddy Maddalena, Stefano Mizzaro |
Distributed Parallel Databases | 3 |
| 2014 | Size and Source Matter: Understanding Inconsistencies in Test Collection-Based EvaluationabstractPast work showed that significant inconsistencies between retrieval results occurred on different test collections, even when one of the test collections contained only a subset of the documents in the other. However, the experimental methodologies in that paper made it hard to determine the cause of the inconsistencies. Using a novel methodology that eliminates the problems with uneven distribution of relevant documents, we confirm that observing a statistically significant improvement between two IR systems can be strongly influenced by the choice of documents in the test collection. We investigate two possible causes of this problem of test collections. Our results show that collection size and document source have a strong influence in the way that a test collection will rank one retrieval system relative to another. This is of particular interest when constructing test collections, as we show that using different subsets of a collection produces differing evaluation results. Timothy Jones 0001, Andrew Turpin, Stefano Mizzaro, Falk Scholer, Mark Sanderson |
CIKM | 3 |
| 2014 | A general account of effectiveness metrics for information tasks: retrieval, filtering, and clusteringabstractIn this tutorial we will present, review, and compare the most popular evaluation metrics for some of the most salient information related tasks, covering: (i) Information Retrieval, (ii) Clustering, and (iii) Filtering. The tutorial will make a special emphasis on the specification of constraints for suitable metrics in each of the three tasks, and on the systematic comparison of metrics according to such constraints. The last part of the tutorial will investigate the challenge of combining and weighting metrics. Enrique Amigó, Julio Gonzalo 0001, Stefano Mizzaro |
SIGIR | 3 |
| 2014 | TREC: topic engineering exerciseabstractIn this work, we investigate approaches to engineer better topic sets in information retrieval test collections. By recasting the TREC evaluation exercise from one of building more effective systems to an exercise in building better topics, we present two possible approaches to quantify topic "goodness": topic ease and topic set predictivity. A novel interpretation of a well known result and a twofold analysis of data from several TREC editions lead to a result that has been neglected so far: both topic ease and topic set predictivity have changed significantly across the years, sometimes in a perhaps undesirable way. J. Shane Culpepper, Stefano Mizzaro, Mark Sanderson, Falk Scholer |
SIGIR | 2 |
| 2012 | Using crowdsourcing for TREC relevance assessment
Omar Alonso, Stefano Mizzaro |
Inf. Process. Manag. | 2 |
| 2012 | Readersourcing - a manifestoabstractThis position paper analyzes the current situation in scholarly publishing and peer review practices and presents three theses: (a) we are going to run out of peer reviewers; (b) it is possible to replace referees with readers, an approach that I have named “Readersourcing”; and (c) it is possible to avoid potential weaknesses in the Readersourcing model by adopting an appropriate quality control mechanism. The readersourcing.org system is then presented as an independent, third‐party, nonprofit, and academic/scientific endeavor aimed at quality rating of scholarly literature and scholars, and some possible criticisms are discussed. Stefano Mizzaro |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2009 | Relevance criteria for e-commerce: a crowdsourcing-based experimental analysisabstractWe discuss the concept of relevance criteria in the context of e-Commerce search. A vast body of research literature describes the beyond-topical criteria used to determine the relevance of the document to the need. We argue that in an e-Commerce scenario there are some differences, and novel and different criteria can be used to determine relevance. We experimentally validate this hypothesis by means of Amazon Mechanical Turk using a crowdsourcing approach. Omar Alonso, Stefano Mizzaro |
SIGIR | 2 |
| 2009 | Mobile information retrieval with search results clustering: Prototypes and evaluationsabstractAbstract Web searches from mobile devices such as PDAs and cell phones are becoming increasingly popular. However, the traditional list‐based search interface paradigm does not scale well to mobile devices due to their inherent limitations. In this article, we investigate the application of search results clustering, used with some success for desktop computer searches, to the mobile scenario. Building on CREDO (Conceptual Reorganization of Documents), a Web clustering engine based on concept lattices, we present its mobile versions Credino and SmartCREDO, for PDAs and cell phones, respectively. Next, we evaluate the retrieval performance of the three prototype systems. We measure the effectiveness of their clustered results compared to a ranked list of results on a subtopic retrieval task, by means of the device‐independent notion of subtopic reach time together with a reusable test collection built from Wikipedia ambiguous entries. Then, we make a cross‐comparison of methods (i.e., clustering and ranked list) and devices (i.e., desktop, PDA, and cell phone), using an interactive information‐finding task performed by external participants. The main finding is that clustering engines are a viable complementary approach to plain search engines both for desktop and mobile searches especially, but not only, for multitopic informational queries. Claudio Carpineto, Stefano Mizzaro, Giovanni Romano 0002, Matteo Snidero |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2009 | A few good topics: Experiments in topic set reduction for retrieval evaluationabstractWe consider the issue of evaluating information retrieval systems on the basis of a limited number of topics. In contrast to statistically-based work on sample sizes, we hypothesize that some topics or topic sets are better than others at predicting true system effectiveness, and that with the right choice of topics, accurate predictions can be obtained from small topics sets. Using a variety of effectiveness metrics and measures of goodness of prediction, a study of a set of TREC and NTCIR results confirms this hypothesis, and provides evidence that the value of a topic set for this purpose does generalize. John Guiver, Stefano Mizzaro, Stephen E. Robertson |
ACM Trans. Inf. Syst. | 2 |
| 2008 | The Good, the Bad, the Difficult, and the Easy: Something Wrong with Information Retrieval Evaluation?
Stefano Mizzaro |
ECIR | 1 |
| 2007 | Hits hits TREC: exploring IR evaluation results with network analysisabstractWe propose a novel method of analysing data gathered fromTREC or similar information retrieval evaluation experiments. We define two normalized versions of average precision, that we use to construct a weighted bipartite graph of TREC systems and topics. We analyze the meaning of well known - and somewhat generalized - indicators fromsocial network analysis on the Systems-Topics graph. We apply this method to an analysis of TREC 8 data; amongthe results, we find that authority measures systems performance, that hubness of topics reveals that some topics are better than others at distinguishing more or less effective systems, that with current measures a system that wants to be effective in TREC needs to be effective on easy topics, and that by using different effectiveness measures this is no longer the case. Stefano Mizzaro, Stephen E. Robertson |
SIGIR | 1 |
| 2006 | Mobile Clustering Engine
Claudio Carpineto, Andrea Della Pietra, Stefano Mizzaro, Giovanni Romano 0002 |
ECIR | 3 |
| 2006 | A Classification of IR Effectiveness Metrics
Gianluca Demartini, Stefano Mizzaro |
ECIR | 2 |
| 2006 | Experiments on Average Distance Measure
Vincenzo Della Mea, Gianluca Demartini, Luca Di Gaspero, Stefano Mizzaro |
ECIR | 4 |
| 2004 | Measuring retrieval effectiveness: A new proposal and a first experimental validationabstractAbstract Most common effectiveness measures for information retrieval systems are based on the assumptions of binary relevance (either a document is relevant to a given query or it is not) and binary retrieval (either a document is retrieved or it is not). In this article, these assumptions are questioned, and a new measure named ADM (average distance measure) is proposed, discussed from a conceptual point of view, and experimentally validated on Text Retrieval Conference (TREC) data. Both conceptual analysis and experimental evidence demonstrate ADM's adequacy in measuring the effectiveness of information retrieval systems. Some potential problems about precision and recall are also highlighted and discussed. Vincenzo Della Mea, Stefano Mizzaro |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2003 | Quality control in scholarly publishing: A new proposalabstractAbstract The Internet has fostered a faster, more interactive and effective model of scholarly publishing. However, as the quantity of information available is constantly increasing, its quality is threatened, since the traditional quality control mechanism of peer review is often not used (e.g., in online repositories of preprints, and by people publishing whatever they want on their Web pages). This paper describes a new kind of electronic scholarly journal, in which the standard submission‐review‐publication process is replaced by a more sophisticated approach, based on judgments expressed by the readers: in this way, each reader is, potentially, a peer reviewer. New ingredients, not found in similar approaches, are that each reader's judgment is weighted on the basis of the reader's skills as a reviewer, and that readers are encouraged to express correct judgments by a feedback mechanism that estimates their own quality. The new electronic scholarly journal is described in both intuitive and formal ways. Its effectiveness is tested by several laboratory experiments that simulate what might happen if the system were deployed and used. Stefano Mizzaro |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2002 | Strategic help in user interfaces for information retrievalabstractAbstract Although no unified definition of the concept of search strategy in Information Retrieval (IR) exists so far, its importance is manifest: nonexpert users, directly interacting with an IR system, apply a limited portfolio of simple actions; they do not know how to react in critical situations; and they often do not even realize that their difficulties are due to strategic problems. A user interface to an IR system should therefore provide some strategic help, focusing user's attention on strategic issues and providing tools to generate better strategies. Because neither the user nor the system can autonomously solve the information problem, but they complement each other, we propose a collaborative coaching approach, in which the two partners cooperate: the user retains the control of the session and the system provides suggestions. The effectiveness of the approach is demonstrated by a conceptual analysis, a prototype knowledge‐based system named FIRE, and its evaluation through informal laboratory experiments. Giorgio Brajnik, Stefano Mizzaro, Carlo Tasso, Fabio Venuti |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2000 | Towards a Theory of Epistemic Information
Stefano Mizzaro |
EJC | 1 |
| 1997 | Relevance: The Whole HistoryabstractRelevance is a fundamental, though not completely understood, concept for documentation, information science, and information retrieval. This article presents the history of relevance through an exhaustive review of the literature. Such history being very complex (about 160 papers are discussed), it is not simple to describe it in a comprehensible way. Thus, first of all a framework for establishing a common ground is defined, and then the history itself is illustrated via the presentation in chronological order of the papers on relevance. The history is divided into three periods (“Before 1958,” “1959–1976,” and “1977–present”) and, inside each period, the papers on relevance are analyzed under seven different aspects (methodological foundations, different kinds of relevance, beyond-topical criteria adopted by users, modes for expression of the relevance judgment, dynamic nature of relevance, types of document representation, and agreement among different judges). © 1997 John Wiley & Sons, Inc. Stefano Mizzaro |
J. Am. Soc. Inf. Sci. | 1 |
| 1996 | Evaluating User Interfaces to Information Retrieval Systems: A Case Study on User SupportabstractArticle Evaluating user interfaces to information retrieval systems: a case study on user support Share on Authors: Giorgio Brajnik Dipartimento di Matematica e Informatica, University of Udine, Via delle Scienze, 206, Loc. Rizzi - 33100 Udine - ITALY Dipartimento di Matematica e Informatica, University of Udine, Via delle Scienze, 206, Loc. Rizzi - 33100 Udine - ITALYView Profile , Stefano Mizzaro Dipartimento di Matematica e Informatica, University of Udine, Via delle Scienze, 206, Loc. Rizzi - 33100 Udine - ITALY Dipartimento di Matematica e Informatica, University of Udine, Via delle Scienze, 206, Loc. Rizzi - 33100 Udine - ITALYView Profile , Carlo Tasso Dipartimento di Matematica e Informatica, University of Udine, Via delle Scienze, 206, Loc. Rizzi - 33100 Udine - ITALY Dipartimento di Matematica e Informatica, University of Udine, Via delle Scienze, 206, Loc. Rizzi - 33100 Udine - ITALYView Profile Authors Info & Claims SIGIR '96: Proceedings of the 19th annual international ACM SIGIR conference on Research and development in information retrievalAugust 1996 Pages 128–136https://doi.org/10.1145/243199.243249Online:18 August 1996Publication History 46citation1,584DownloadsMetricsTotal Citations46Total Downloads1,584Last 12 Months16Last 6 weeks6 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Giorgio Brajnik, Stefano Mizzaro, Carlo Tasso |
SIGIR | 2 |