EDBT 2026 Demo / reviewers in the wild / expert
Eddy Maddalena
dblp:133/6754
· DBLP profile ↗
21ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0002-5423-8669ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 17 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 5Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Sorry, Your HIT Is Overbooked - Investigating the Use of Crowdsourcing HIT CatchersabstractIn microtask crowdsourcing, Human Intelligence Tasks (HITs) are commonly allocated on a first-come, first-served basis: they are published on the platform and the fastest workers select the most attractive ones first. This step has not received much attention from the scientific community yet, though it can become particularly taxing for workers when they compete to secure the most sought-after tasks. There are many strategies to ensure one's access to tasks and their effects on the labour process as a whole are not well understood. For instance, platforms with a sizeable task reservation queue allow workers to gain preferential access to a large number of tasks, which in turn may cause a shortage of work for the rest of the crowd. For the requesters, this means lower rates of completion and a lack of worker diversity. We explore workers' strategies for accessing and reserving tasks using monitoring techniques from both client and server sides. We investigate how these strategies affect task execution, in terms of availability, completion time, and answer quality, by deploying 1000 image annotation HITs in Amazon Mechanical Turk including objective and subjective tasks. We observe that workers who do not use automated catching techniques tend to have higher annotation quality, are more focused, spend more effort on text editing, and provide a higher diversity of output than workers using such tools. This study also reveals the tragedy of the commons effect among platform members due to the use of catching techniques: workers using automated catching techniques reserve and complete a substantially higher portion of the available tasks, but the over-reservation of HITs restricts all workers of reservation opportunities, and compromise their own future labour capacity as well. We observe a high inefficiency in job completions, as the majority of the times a task is being reserved by a worker, it will not get actually performed and will need to be republished for further allocation. Finally, we propose solutions to mitigate the negative effects of these phenomena on the labour process. Eddy Maddalena, Alessandro Checco, Haoyu Xie 0001, Efpraxia D. Zamani, Elena Simperl |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2024 | Crowdsourced Fact-checking: Does It Actually Work?abstractThere is an important ongoing effort aimed to tackle misinformation and to perform reliable fact-checking by employing human assessors at scale, with a crowdsourcing-based approach. Previous studies on the feasibility of employing crowdsourcing for the task of misinformation detection have provided inconsistent results: some of them seem to confirm the effectiveness of crowdsourcing for assessing the truthfulness of statements and claims, whereas others fail to reach an effectiveness level higher than automatic machine learning approaches, which are still unsatisfactory. In this paper, we aim at addressing such inconsistency and understand if truthfulness assessment can indeed be crowdsourced effectively. To do so, we build on top of previous studies; we select some of those reporting low effectiveness levels, we highlight their potential limitations, and we then reproduce their work attempting to improve their setup to address those limitations. We employ various approaches, data quality levels, and agreement measures to assess the reliability of crowd workers when assessing the truthfulness of (mis)information. Furthermore, we explore different worker features and compare the results obtained with different crowds. According to our findings, crowdsourcing can be used as an effective methodology to tackle misinformation at scale. When compared to previous studies, our results indicate that a significantly higher agreement between crowd workers and experts can be obtained by using a different, higher-quality, crowdsourcing platform and by improving the design of the crowdsourcing task. Also, we find differences concerning task and worker features and how workers provide truthfulness assessments. David La Barbera, Eddy Maddalena, Michael Soprano, Kevin Roitero, Gianluca Demartini, Davide Ceolin, Damiano Spina, Stefano Mizzaro |
Inf. Process. Manag. | 2 |
| 2023 | The Dark Side of Recruitment in Crowdsourcing: Ethics and Transparency in Micro-Task MarketplacesabstractAbstract Micro-task crowdsourcing marketplaces like Figure Eight (F8) connect a large pool of workers to employers through a single online platform, by aggregating multiple crowdsourcing platforms (channels) under a unique system. This paper investigates the F8 channels’ demographic distribution and reward schemes by analysing more than 53k crowdsourcing tasks over four years, collecting survey data and scraping marketplace metadata. We reveal an heterogeneous per-channel demographic distribution, and an opaque channel commission scheme, that varies over time and is not communicated to the employer when launching a task: workers often will receive a smaller payment than expected by the employer. In addition, the impact of channel commission schemes on the relationship between requesters and crowdworkers is explored. These observations uncover important issues on ethics, reliability and transparency of crowdsourced experiment when using this kind of marketplaces, especially for academic research. Haoyu Xie 0001, Eddy Maddalena, Rehab K. Qarout, Alessandro Checco |
Comput. Support. Cooperative Work. | 2 |
| 2023 | Qrowdsmith: Enhancing Paid Microtask Crowdsourcing with Gamification and Furtherance IncentivesabstractMicrotask crowdsourcing platforms are social intelligence systems in which volunteers, called crowdworkers, complete small, repetitive tasks in return for a small fee. Beyond payments, task requesters are considering non-monetary incentives such as points, badges, and other gamified elements to increase performance and improve crowdworker experience. In this article, we present Qrowdsmith, a platform for gamifying microtask crowdsourcing. To design the system, we explore empirically a range of gamified and financial incentives and analyse their impact on how efficient, effective, and reliable the results are. To maintain participation over time and save costs, we propose furtherance incentives, which are offered to crowdworkers to encourage additional contributions in addition to the fee agreed upfront. In a series of controlled experiments, we find that while gamification can work as furtherance incentives, it impacts negatively on crowdworkers’ performance, both in terms of the quantity and quality of work, as compared to a baseline where they can continue to contribute voluntarily. Gamified incentives are also less effective than paid bonus equivalents. Our results contribute to the understanding of how best to encourage engagement in microtask crowdsourcing activities and design better crowd intelligence systems. Eddy Maddalena, Luis-Daniel Ibáñez, Neal Reeves, Elena Simperl |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2021 | On the effect of relevance scales in crowdsourcing relevance assessments for Information Retrieval evaluation
Kevin Roitero, Eddy Maddalena, Stefano Mizzaro, Falk Scholer |
Inf. Process. Manag. | 2 |
| 2021 | The Impact of Task Abandonment in CrowdsourcingabstractCrowdsourcing has become a standard methodology to collect manually annotated data such as relevance judgments at scale. On crowdsourcing platforms like Amazon MTurk or FigureEight, crowd workers select tasks to work on based on different dimensions such as task reward and requester reputation. Requesters then receive the judgments of workers who self-selected into the tasks and completed them successfully. Several crowd workers, however, preview tasks, begin working on them, reaching varying stages of task completion without finally submitting their work. Such behavior results in unrewarded effort which remains invisible to requesters. In this paper, we conduct an investigation of the phenomenon of task abandonment, the act of workers previewing or beginning a task and deciding not to complete it. We follow a three-fold methodology which includes 1) investigating the prevalence and causes of task abandonment by means of a survey over different crowdsourcing platforms, 2) data-driven analysis of logs collected during a large-scale relevance judgment experiment, and 3) controlled experiments measuring the effect of different dimensions on abandonment. Our results show that task abandonment is a widely spread phenomenon. Apart from accounting for a considerable amount of wasted human effort, this bears important implications on the hourly wages of workers as they are not rewarded for tasks that they do not complete. We also show how task abandonment may have strong implications on the use of collected data (for example, on the evaluation of Information Retrieval systems). Lei Han 0003, Kevin Roitero, Ujwal Gadiraju, Cristina Sarasua, Alessandro Checco, Eddy Maddalena, Gianluca Demartini |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2020 | Point at the Triple: Generation of Text Summaries from Knowledge Base Triples (Extended Abstract)abstractWe investigate the problem of generating natural language summaries from knowledge base triples. Our approach is based on a pointer-generator network, which, in addition to generating regular words from a fixed target vocabulary, is able to verbalise triples in several ways. We undertake an automatic and a human evaluation on single and open-domain summaries generation tasks. Both show that our approach significantly outperforms other data-driven baselines. Pavlos Vougiouklis, Eddy Maddalena, Jonathon S. Hare, Elena Simperl |
IJCAI | 2 |
| 2020 | Crowd Worker Strategies in Relevance Judgment TasksabstractCrowdsourcing is a popular technique to collect large amounts of human-generated labels, such as relevance judgments used to create information retrieval (IR) evaluation collections. Previous research has shown how collecting high quality labels from a crowdsourcing platform can be challenging. Existing quality assurance techniques focus on answer aggregation or on the use of gold questions where ground-truth data allows to check for the quality of the responses. Lei Han 0003, Eddy Maddalena, Alessandro Checco, Cristina Sarasua, Ujwal Gadiraju, Kevin Roitero, Gianluca Demartini |
WSDM | 2 |
| 2020 | Point at the Triple: Generation of Text Summaries from Knowledge Base TriplesabstractWe investigate the problem of generating natural language summaries from knowledge base triples. Our approach is based on a pointer-generator network, which, in addition to generating regular words from a fixed target vocabulary, is able to verbalise triples in several ways. We undertake an automatic and a human evaluation on single and open-domain summaries generation tasks. Both show that our approach significantly outperforms other data-driven baselines. Pavlos Vougiouklis, Eddy Maddalena, Jonathon S. Hare, Elena Simperl |
J. Artif. Intell. Res. | 2 |
| 2020 | Mapping Points of Interest Through Street View Imagery and Paid CrowdsourcingabstractWe present the Virtual City Explorer (VCE), an online crowdsourcing platform for the collection of rich geotagged information in urban environments. Compared to other volunteered geographic information approaches, which are constrained by the number and availability of mapping enthusiasts on the ground, the VCE uses digital street imagery to allow people to virtually explore a city from anywhere in the world, using a browser or a mobile phone. In addition, contributions in VCE are designed as paid microtasks—small jobs that can be carried out without any specific knowledge of the local area or previous mapping expertise in exchange for a fee. We tested the VCE in two cities to map points of interest (PoIs) in transport and mobility, using FigureEight to recruit participants. We were able to show that our platform enables crowdworkers to submit PoI location seamlessly, cover almost all of the tested areas, and discover several PoIs not reported by other approaches. This allows the VCE to complement existing approaches that leverage experts or grassroot communities. Eddy Maddalena, Luis-Daniel Ibáñez, Elena Simperl |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2019 | On Transforming Relevance ScalesabstractInformation Retrieval (IR) researchers have often used existing IR evaluation collections and transformed the relevance scale in which judgments have been collected, e.g., to use metrics that assume binary judgments like Mean Average Precision. Such scale transformations are often arbitrary (e.g., 0,1 mapped to 0 and 2,3 mapped to 1) and it is assumed that they have no impact on the results of IR evaluation. Moreover, the use of crowdsourcing to collect relevance judgments has become a standard methodology. When designing the crowdsourcing relevance judgment task, one of the decision to be made is the how granular the relevance scale used to collect judgments should be. Such decision has then repercussions on the metrics used to measure IR system effectiveness. In this paper we look at the effect of scale transformations in a systematic way. We perform extensive experiments to study the transformation of judgments from fine-grained to coarse-grained. We use different relevance judgments expressed on different relevance scales and either expressed by expert annotators or collected by means of crowdsourcing. The objective is to understand the impact of relevance scale transformations on IR evaluation outcomes and to draw conclusions on how to best transform judgments into a different scale, when necessary. Lei Han 0003, Kevin Roitero, Eddy Maddalena, Stefano Mizzaro, Gianluca Demartini |
CIKM | 3 |
| 2019 | All Those Wasted Hours: On Task Abandonment in CrowdsourcingabstractCrowdsourcing has become a standard methodology to collect manually annotated data such as relevance judgments at scale. On crowdsourcing platforms like Amazon MTurk or FigureEight, crowd workers select tasks to work on based on different dimensions such as task reward and requester reputation. Requesters then receive the judgments of workers who self-selected into the tasks and completed them successfully. Several crowd workers, however, preview tasks, begin working on them, reaching varying stages of task completion without finally submitting their work. Such behavior results in unrewarded effort which remains invisible to requesters. In this paper, we conduct the first investigation into the phenomenon of task abandonment, the act of workers previewing or beginning a task and deciding not to complete it. We follow a three-fold methodology which includes 1) investigating the prevalence and causes of task abandonment by means of a survey over different crowdsourcing platforms, 2) data-driven analyses of logs collected during a large-scale relevance judgment experiment, and 3) controlled experiments measuring the effect of different dimensions on abandonment. Our results show that task abandonment is a widely spread phenomenon. Apart from accounting for a considerable amount of wasted human effort, this bears important implications on the hourly wages of workers as they are not rewarded for tasks that they do not complete. We also show how task abandonment may have strong implications on the use of collected data (for example, on the evaluation of IR systems). Lei Han 0003, Kevin Roitero, Ujwal Gadiraju, Cristina Sarasua, Alessandro Checco, Eddy Maddalena, Gianluca Demartini |
WSDM | 6 |
| 2018 | On Fine-Grained Relevance ScalesabstractIn Information Retrieval evaluation, the classical approach of adopting binary relevance judgments has been replaced by multi-level relevance judgments and by gain-based metrics leveraging such multi-level judgment scales. Recent work has also proposed and evaluated unbounded relevance scales by means of Magnitude Estimation (ME) and compared them with multi-level scales. While ME brings advantages like the ability for assessors to always judge the next document as having higher or lower relevance than any of the documents they have judged so far, it also comes with some drawbacks. For example, it is not a natural approach for human assessors to judge items as they are used to do on the Web (e.g., 5-star rating). In this work, we propose and experimentally evaluate a bounded and fine-grained relevance scale having many of the advantages and dealing with some of the issues of ME. We collect relevance judgments over a 100-level relevance scale (S100) by means of a large-scale crowdsourcing experiment and compare the results with other relevance scales (binary, 4-level, and ME) showing the benefit of fine-grained scales over both coarse-grained and unbounded scales as well as highlighting some new results on ME. Our results show that S100 maintains the flexibility of unbounded scales like ME in providing assessors with ample choice when judging document relevance (i.e., assessors can fit relevance judgments in between of previously given judgments). It also allows assessors to judge on a more familiar scale (e.g., on 10 levels) and to perform efficiently since the very first judging task. Kevin Roitero, Eddy Maddalena, Gianluca Demartini, Stefano Mizzaro |
SIGIR | 2 |
| 2018 | IRevalOO: An Object Oriented Framework for Retrieval EvaluationabstractWe propose IRevalOO, a flexible Object Oriented framework that (i) can be used as-is as a replacement of the widely adopted trec\_eval software, and (ii) can be easily extended (or "instantiated'', in framework terminology) to implement different scenarios of test collection based retrieval evaluation. Instances of IRevalOO can provide a usable and convenient alternative to the state-of-the-art software commonly used by different initiatives (TREC, NTCIR, CLEF, FIRE, etc.). Also, those instances can be easily adapted to satisfy future customization needs of researchers, as: implementing and experimenting with new metrics, even based on new notions of relevance; using different formats for system output and "qrels''; and in general visualizing, comparing, and managing retrieval evaluation results. Kevin Roitero, Eddy Maddalena, Yannick Ponte, Stefano Mizzaro |
SIGIR | 2 |
| 2017 | Do Easy Topics Predict Effectiveness Better Than Difficult Topics?
Kevin Roitero, Eddy Maddalena, Stefano Mizzaro |
ECIR | 2 |
| 2017 | Let's Agree to Disagree: Fixing Agreement Measures for CrowdsourcingabstractIn the context of micro-task crowdsourcing, each task is usually performed by several workers. This allows researchers to leverage measures of the agreement among workers on the same task, to estimate the reliability of collected data and to better understand answering behaviors of the participants. While many measures of agreement between annotators have been proposed, they are known for suffering from many problems and abnormalities. In this paper, we identify the main limits of the existing agreement measures in the crowdsourcing context, both by means of toy examples as well as with real-world crowdsourcing data, and propose a novel agreement measure based on probabilistic parameter estimation which overcomes such limits. We validate our new agreement measure and show its flexibility as compared to the existing agreement measures. Alessandro Checco, Kevin Roitero, Eddy Maddalena, Stefano Mizzaro, Gianluca Demartini |
HCOMP | 3 |
| 2017 | On Crowdsourcing Relevance Magnitudes for Information Retrieval EvaluationabstractMagnitude estimation is a psychophysical scaling technique for the measurement of sensation, where observers assign numbers to stimuli in response to their perceived intensity. We investigate the use of magnitude estimation for judging the relevance of documents for information retrieval evaluation, carrying out a large-scale user study across 18 TREC topics and collecting over 50,000 magnitude estimation judgments using crowdsourcing. Our analysis shows that magnitude estimation judgments can be reliably collected using crowdsourcing, are competitive in terms of assessor cost, and are, on average, rank-aligned with ordinal judgments made by expert relevance assessors. We explore the application of magnitude estimation for IR evaluation, calibrating two gain-based effectiveness metrics, nDCG and ERR, directly from user-reported perceptions of relevance. A comparison of TREC system effectiveness rankings based on binary, ordinal, and magnitude estimation relevance shows substantial variation; in particular, the top systems ranked using magnitude estimation and ordinal judgments differ substantially. Analysis of the magnitude estimation scores shows that this effect is due in part to varying perceptions of relevance: different users have different perceptions of the impact of relative differences in document relevance. These results have direct implications for IR evaluation, suggesting that current assumptions about a single view of relevance being sufficient to represent a population of users are unlikely to hold. Eddy Maddalena, Stefano Mizzaro, Falk Scholer, Andrew Turpin |
ACM Trans. Inf. Syst. | 1 |
| 2016 | Crowdsourcing Relevance Assessments: The Unexpected Benefits of Limiting the Time to JudgeabstractCrowdsourcing has become an alternative approach to collect relevance judgments at scale thanks to the availability of crowdsourcing platforms and quality control techniques that allow to obtain reliable results. Previous work has used crowdsourcing to ask multiple crowd workers to judge the relevance of a document with respect to a query and studied how to best aggregate multiple judgments of the same topic-document pair. This paper addresses an aspect that has been rather overlooked so far: we study how the time available to express a relevance judgment affects its quality. We also discuss the quality loss of making crowdsourced relevance judgments more efficient in terms of time taken to judge the relevance of a document. We use standard test collections to run a battery of experiments on the crowdsourcing platform CrowdFlower, studying how much time crowd workers need to judge the relevance of a document and at what is the effect of reducing the available time to judge on the overall quality of the judgments. Our extensive experiments compare judgments obtained under different types of time constraints with judgments obtained when no time constraints were put on the task. We measure judgment quality by different metrics of agreement with editorial judgments. Experimental results show that it is possible to reduce the cost of crowdsourced evaluation collection creation by reducing the time available to perform the judgments with no loss in quality. Most importantly, we observed that the introduction of limits on the time available to perform the judgments improves the overall judgment quality. Top judgment quality is obtained with 25-30 seconds to judge a topic-document pair. Eddy Maddalena, Marco Basaldella, Dario De Nart, Dante Degl'Innocenti, Stefano Mizzaro, Gianluca Demartini |
HCOMP | 1 |
| 2015 | Judging Relevance Using Magnitude Estimation
Eddy Maddalena, Stefano Mizzaro, Falk Scholer, Andrew Turpin |
ECIR | 1 |
| 2015 | The Benefits of Magnitude Estimation Relevance Assessments for Information Retrieval EvaluationabstractMagnitude estimation is a psychophysical scaling technique for the measurement of sensation, where observers assign numbers to stimuli in response to their perceived intensity. We investigate the use of magnitude estimation for judging the relevance of documents in the context of information retrieval evaluation, carrying out a large-scale user study across 18 TREC topics and collecting more than 50,000 magnitude estimation judgments. Our analysis shows that on average magnitude estimation judgments are rank-aligned with ordinal judgments made by expert relevance assessors. An advantage of magnitude estimation is that users can chose their own scale for judgments, allowing deeper investigations of user perceptions than when categorical scales are used. Andrew Turpin, Falk Scholer, Stefano Mizzaro, Eddy Maddalena |
SIGIR | 4 |
| 2015 | Mobile crowdsourcing: four experiments on platforms and tasks
Vincenzo Della Mea, Eddy Maddalena, Stefano Mizzaro |
Distributed Parallel Databases | 2 |