EDBT 2026 Demo / reviewers in the wild / expert
Alessandro Checco
dblp:72/10821
· DBLP profile ↗
23ranked-venue papers
9as first author
5since 2021 · last 2025
0000-0002-0981-3409ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 10 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 7 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 1 since 2021Computer networks · 5 · 3 first-authorSecurity and privacy · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Sorry, Your HIT Is Overbooked - Investigating the Use of Crowdsourcing HIT CatchersabstractIn microtask crowdsourcing, Human Intelligence Tasks (HITs) are commonly allocated on a first-come, first-served basis: they are published on the platform and the fastest workers select the most attractive ones first. This step has not received much attention from the scientific community yet, though it can become particularly taxing for workers when they compete to secure the most sought-after tasks. There are many strategies to ensure one's access to tasks and their effects on the labour process as a whole are not well understood. For instance, platforms with a sizeable task reservation queue allow workers to gain preferential access to a large number of tasks, which in turn may cause a shortage of work for the rest of the crowd. For the requesters, this means lower rates of completion and a lack of worker diversity. We explore workers' strategies for accessing and reserving tasks using monitoring techniques from both client and server sides. We investigate how these strategies affect task execution, in terms of availability, completion time, and answer quality, by deploying 1000 image annotation HITs in Amazon Mechanical Turk including objective and subjective tasks. We observe that workers who do not use automated catching techniques tend to have higher annotation quality, are more focused, spend more effort on text editing, and provide a higher diversity of output than workers using such tools. This study also reveals the tragedy of the commons effect among platform members due to the use of catching techniques: workers using automated catching techniques reserve and complete a substantially higher portion of the available tasks, but the over-reservation of HITs restricts all workers of reservation opportunities, and compromise their own future labour capacity as well. We observe a high inefficiency in job completions, as the majority of the times a task is being reserved by a worker, it will not get actually performed and will need to be republished for further allocation. Finally, we propose solutions to mitigate the negative effects of these phenomena on the labour process. Eddy Maddalena, Alessandro Checco, Haoyu Xie 0001, Efpraxia D. Zamani, Elena Simperl |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2023 | The Dark Side of Recruitment in Crowdsourcing: Ethics and Transparency in Micro-Task MarketplacesabstractAbstract Micro-task crowdsourcing marketplaces like Figure Eight (F8) connect a large pool of workers to employers through a single online platform, by aggregating multiple crowdsourcing platforms (channels) under a unique system. This paper investigates the F8 channels’ demographic distribution and reward schemes by analysing more than 53k crowdsourcing tasks over four years, collecting survey data and scraping marketplace metadata. We reveal an heterogeneous per-channel demographic distribution, and an opaque channel commission scheme, that varies over time and is not communicated to the employer when launching a task: workers often will receive a smaller payment than expected by the employer. In addition, the impact of channel commission schemes on the relationship between requesters and crowdworkers is explored. These observations uncover important issues on ethics, reliability and transparency of crowdsourced experiment when using this kind of marketplaces, especially for academic research. Haoyu Xie 0001, Eddy Maddalena, Rehab K. Qarout, Alessandro Checco |
Comput. Support. Cooperative Work. | 4 |
| 2022 | Preferences on a Budget: Prioritizing Document Pairs when Crowdsourcing Relevance JudgmentsabstractIn Information Retrieval (IR) evaluation, preference judgments are collected by presenting to the assessors a pair of documents and asking them to select which of the two, if any, is the most relevant. This is an alternative to the classic relevance judgment approach, in which human assessors judge the relevance of a single document on a scale; such an alternative allows to make relative rather than absolute judgments of relevance. While preference judgments are easier for human assessors to perform, the number of possible document pairs to be judged is usually so high that it makes it unfeasible to judge them all. Thus, following a similar idea to pooling strategies for single document relevance judgments where the goal is to sample the most useful documents to be judged, in this work we focus on analyzing alternative ways to sample document pairs to judge, in order to maximize the value of a fixed number of preference judgments that can feasibly be collected. Such value is defined as how well we can evaluate IR systems given a budget, that is, a fixed number of human preference judgments that may be collected. By relying on several datasets featuring relevance judgments gathered by means of experts and crowdsourcing, we experimentally compare alternative strategies to select document pairs and show how different strategies lead to different IR evaluation result quality levels. Our results show that, by using the appropriate procedure, it is possible to achieve good IR evaluation results with a limited number of preference judgments, thus confirming the feasibility of using preference judgments to create IR evaluation collections. Kevin Roitero, Alessandro Checco, Stefano Mizzaro, Gianluca Demartini |
WWW | 2 |
| 2021 | Aggregation Techniques in Crowdsourcing: Multiple Choice Questions and BeyondabstractCrowdsourcing has been leveraged in various tasks and applications, primarily to gather information from human annotators in exchange for a monetary reward. The main challenge associated with crowdsourcing is the low quality of the results, which can stem from multiple reasons, including bias, error, and adversarial behavior. Researchers and practitioners can apply quality control methods to prevent and detect low-quality responses. For example, worker selection methods utilize qualifications and attention check questions before assigning a task. Similarly, task routing identifies the workers who can provide a more accurate response to a given task type using recommender system techniques. In practice, posterior quality control methods are the most common approach to deal with noisy labels once they are obtained. Such methods require task repetition, i.e., assigning the task to multiple crowd-workers, followed by an aggregation mechanism (aka truth inference) to select the most likely answer or request an additional label. A large number of techniques have been proposed for crowdsourcing aggregation covering several types of task types. This tutorial aims to present common and recent label aggregation techniques for multiple-choice questions, multi-class labels, ratings, pairwise comparison, and image/text annotation. We believe that the audience will benefit from the focus on this specific research area to learn about the best techniques to apply in their crowdsourcing projects. Djellel Eddine Difallah, Alessandro Checco |
CIKM | 2 |
| 2021 | The Impact of Task Abandonment in CrowdsourcingabstractCrowdsourcing has become a standard methodology to collect manually annotated data such as relevance judgments at scale. On crowdsourcing platforms like Amazon MTurk or FigureEight, crowd workers select tasks to work on based on different dimensions such as task reward and requester reputation. Requesters then receive the judgments of workers who self-selected into the tasks and completed them successfully. Several crowd workers, however, preview tasks, begin working on them, reaching varying stages of task completion without finally submitting their work. Such behavior results in unrewarded effort which remains invisible to requesters. In this paper, we conduct an investigation of the phenomenon of task abandonment, the act of workers previewing or beginning a task and deciding not to complete it. We follow a three-fold methodology which includes 1) investigating the prevalence and causes of task abandonment by means of a survey over different crowdsourcing platforms, 2) data-driven analysis of logs collected during a large-scale relevance judgment experiment, and 3) controlled experiments measuring the effect of different dimensions on abandonment. Our results show that task abandonment is a widely spread phenomenon. Apart from accounting for a considerable amount of wasted human effort, this bears important implications on the hourly wages of workers as they are not rewarded for tasks that they do not complete. We also show how task abandonment may have strong implications on the use of collected data (for example, on the evaluation of Information Retrieval systems). Lei Han 0003, Kevin Roitero, Ujwal Gadiraju, Cristina Sarasua, Alessandro Checco, Eddy Maddalena, Gianluca Demartini |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2020 | Modelling User Behavior Dynamics with EmbeddingsabstractUnderstanding user interaction behaviors remains a challenging problem. Quantifying behavior dynamics over time as users complete tasks has only been done in specific domains. In this paper, we present a user behavior model built using behavior embeddings to compare behaviors and their change over time. To this end, we first define the formal model and train the model using both action (e.g., copy/paste) embeddings and user interaction feature (e.g., length of the copied text) embeddings. Having obtained vector representations of user behaviors, we then define three measurements to model behavior dynamics over time, namely: behavior position, displacement, and velocity. To evaluate the proposed methodology, we use three real world datasets: (i) tens of users completing complex data curation tasks in a lab setting, (ii) hundreds of crowd workers completing structured tasks in a crowdsourcing setting, and (iii) thousands of editors completing unstructured editing tasks on Wikidata. Through these datasets, we show that the proposed methodology can: (i) surface behavioral differences among users; (ii) recognize relative behavioral changes; and (iii) discover directional deviations of user behaviors. Our approach can be used (i) to capture behavioral semantics from data in a consistent way, (ii) to quantify behavioral diversity for a task and among different users, and (iii) to explore the temporal behavior evolution with respect to various task properties (e.g., structure and difficulty). Lei Han 0003, Alessandro Checco, Djellel Eddine Difallah, Gianluca Demartini, Shazia Sadiq |
CIKM | 2 |
| 2020 | Crowd Worker Strategies in Relevance Judgment TasksabstractCrowdsourcing is a popular technique to collect large amounts of human-generated labels, such as relevance judgments used to create information retrieval (IR) evaluation collections. Previous research has shown how collecting high quality labels from a crowdsourcing platform can be challenging. Existing quality assurance techniques focus on answer aggregation or on the use of gold questions where ground-truth data allows to check for the quality of the responses. Lei Han 0003, Eddy Maddalena, Alessandro Checco, Cristina Sarasua, Ujwal Gadiraju, Kevin Roitero, Gianluca Demartini |
WSDM | 3 |
| 2020 | Adversarial Attacks on Crowdsourcing Quality ControlabstractCrowdsourcing is a popular methodology to collect manual labels at scale. Such labels are often used to train AI models and, thus, quality control is a key aspect in the process. One of the most popular quality assurance mechanisms in paid micro-task crowdsourcing is based on gold questions: the use of a small set of tasks of which the requester knows the correct answer and, thus, is able to directly assess crowd work quality. In this paper, we show that such mechanism is prone to an attack carried out by a group of colluding crowd workers that is easy to implement and deploy: the inherent size limit of the gold set can be exploited by building an inferential system to detect which parts of the job are more likely to be gold questions. The described attack is robust to various forms of randomisation and programmatic generation of gold questions. We present the architecture of the proposed system, composed of a browser plug-in and an external server used to share information, and briefly introduce its potential evolution to a decentralised implementation. We implement and experimentally validate the gold detection system, using real-world data from a popular crowdsourcing platform. Our experimental results show that crowdworkers using the proposed system spend more time on signalled gold questions but do not neglect the others thus achieving an increased overall work quality. Finally, we discuss the economic and sociological implications of this kind of attack. Alessandro Checco, Jo Bates, Gianluca Demartini |
J. Artif. Intell. Res. | 1 |
| 2020 | CrowdCO-OP: Sharing Risks and Rewards in CrowdsourcingabstractPaid micro-task crowdsourcing has gained in popularity partly due to the increasing need for large-scale manually labelled datasets which are often used to train and evaluate Artificial Intelligence systems. Modern paid crowdsourcing platforms use a piecework approach to rewards, meaning that workers are paid for each task they complete, given that their work quality is considered sufficient by the requester or the platform. Such an approach creates risks for workers; their work may be rejected without being rewarded, and they may be working on poorly rewarded tasks, in light of the disproportionate time required to complete them. As a result, recent research has shown that crowd workers may tend to choose specific, simple, and familiar tasks and avoid new requesters to manage these risks. In this paper, we propose a novel crowdsourcing reward mechanism that allows workers to share these risks and achieve a standardized hourly wage equal for all participating workers. Reward-focused workers can thereby take up challenging and complex HITs without bearing the financial risk of not being rewarded for completed work. We experimentally compare different crowd reward schemes and observe their impact on worker performance and satisfaction. Our results show that 1) workers clearly perceive the benefits of the proposed reward scheme, 2) work effectiveness and efficiency are not impacted as compared to those of the piecework scheme, and 3) the presence of slow workers is limited and does not disrupt the proposed cooperation-based approaches. Shaoyang Fan, Ujwal Gadiraju, Alessandro Checco, Gianluca Demartini |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2019 | Platform-Related Factors in Repeatability and Reproducibility of Crowdsourcing TasksabstractCrowdsourcing platforms provide a convenient and scalable way to collect human-generated labels on-demand. This data can be used to train Artificial Intelligence (AI) systems or to evaluate the effectiveness of algorithms. The datasets generated by means of crowdsourcing are, however, dependent on many factors that affect their quality. These include, among others, the population sample bias introduced by aspects like task reward, requester reputation, and other filters introduced by the task design.In this paper, we analyse platform-related factors and study how they affect dataset characteristics by running a longitudinal study where we compare the reliability of results collected with repeated experiments over time and across crowdsourcing platforms. Results show that, under certain conditions: 1) experiments replicated across different platforms result in significantly different data quality levels while 2) the quality of data from repeated experiments over time is stable within the same platform. We identify some key task design variables that cause such variations and propose an experimentally validated set of actions to counteract these effects thus achieving reliable and repeatable crowdsourced data collection experiments. Rehab K. Qarout, Alessandro Checco, Gianluca Demartini, Kalina Bontcheva |
HCOMP | 2 |
| 2019 | Quality Control Attack Schemes in CrowdsourcingabstractAn important precondition to build effective AI models is the collection of training data at scale. Crowdsourcing is a popular methodology to achieve this goal. Its adoption introduces novel challenges in data quality control, to deal with under-performing and malicious annotators. One of the most popular quality assurance mechanisms, especially in paid micro-task crowdsourcing, is the use of a small set of pre-annotated tasks as gold standard, to assess in real time the annotators quality. In this paper, we highlight a set of vulnerabilities this scheme suffers: a group of colluding crowd workers can easily implement and deploy a decentralised machine learning inferential system to detect and signal which parts of the task are more likely to be gold questions, making them ineffective as a quality control tool. Moreover, we demonstrate how the most common countermeasures against this attack are ineffective in practical scenarios. The basic architecture of the inferential system is composed of a browser plug-in and an external server where the colluding workers can share information. We implement and validate the attack scheme, by means of experiments on real-world data from a popular crowdsourcing platform. Alessandro Checco, Jo Bates, Gianluca Demartini |
IJCAI | 1 |
| 2019 | All Those Wasted Hours: On Task Abandonment in CrowdsourcingabstractCrowdsourcing has become a standard methodology to collect manually annotated data such as relevance judgments at scale. On crowdsourcing platforms like Amazon MTurk or FigureEight, crowd workers select tasks to work on based on different dimensions such as task reward and requester reputation. Requesters then receive the judgments of workers who self-selected into the tasks and completed them successfully. Several crowd workers, however, preview tasks, begin working on them, reaching varying stages of task completion without finally submitting their work. Such behavior results in unrewarded effort which remains invisible to requesters. In this paper, we conduct the first investigation into the phenomenon of task abandonment, the act of workers previewing or beginning a task and deciding not to complete it. We follow a three-fold methodology which includes 1) investigating the prevalence and causes of task abandonment by means of a survey over different crowdsourcing platforms, 2) data-driven analyses of logs collected during a large-scale relevance judgment experiment, and 3) controlled experiments measuring the effect of different dimensions on abandonment. Our results show that task abandonment is a widely spread phenomenon. Apart from accounting for a considerable amount of wasted human effort, this bears important implications on the hourly wages of workers as they are not rewarded for tasks that they do not complete. We also show how task abandonment may have strong implications on the use of collected data (for example, on the evaluation of IR systems). Lei Han 0003, Kevin Roitero, Ujwal Gadiraju, Cristina Sarasua, Alessandro Checco, Eddy Maddalena, Gianluca Demartini |
WSDM | 5 |
| 2019 | The Evolution of Power and Standard Wikidata Editors: Comparing Editing Behavior over Time to Predict Lifespan and Volume of Edits
Cristina Sarasua, Alessandro Checco, Gianluca Demartini, Djellel Eddine Difallah, Michael Feldman 0001, Lydia Pintscher |
Comput. Support. Cooperative Work. | 2 |
| 2018 | All That Glitters Is Gold - An Attack Scheme on Gold Questions in CrowdsourcingabstractOne of the most popular quality assurance mechanisms in paid micro-task crowdsourcing is based on gold questions: the use of a small set of tasks of which the requester knows the correct answer and, thus, is able to directly assess crowd work quality. In this paper, we show that such mechanism is prone to an attack carried out by a group of colluding crowd workers that is easy to implement and deploy: the inherent size limit of the gold set can be exploited by building an inferential system to detect which parts of the job are more likely to be gold questions. The described attack is robust to various forms of randomisation and programmatic generation of gold questions. We present the architecture of the proposed system, composed of a browser plug-in and an external server used to share information, and briefly introduce its potential evolution to a decentralised implementation. We implement and experimentally validate the gold detection system, using real-world data from a popular crowdsourcing platform. Finally, we discuss the economic and sociological implications of this kind of attack. Alessandro Checco, Jo Bates, Gianluca Demartini |
HCOMP | 1 |
| 2018 | Investigating User Perception of Gender Bias in Image Search: The Role of SexismabstractThere is growing evidence that search engines produce results that are socially biased, reinforcing a view of the world that aligns with prevalent social stereotypes. One means to promote greater transparency of search algorithms - which are typically complex and proprietary - is to raise user awareness of biased result sets. However, to date, little is known concerning how users perceive bias in search results, and the degree to which their perceptions differ and/or might be predicted based on user attributes. One particular area of search that has recently gained attention, and forms the focus of this study, is image retrieval and gender bias. We conduct a controlled experiment via crowdsourcing using participants recruited from three countries to measure the extent to which workers perceive a given image results set to be subjective or objective. Demographic information about the workers, along with measures of sexism, are gathered and analysed to investigate whether (gender) biases in the image search results can be detected. Amongst other findings, the results confirm that sexist people are less likely to detect and report gender biases in image search results. Jahna Otterbacher, Alessandro Checco, Gianluca Demartini, Paul D. Clough |
SIGIR | 2 |
| 2018 | Updating Neighbour Cell List via Crowdsourced User Reports: A Framework for Measuring Time PerformanceabstractIn modern wireless networks deployments, each serving node needs to keep its Neighbour Cell List (NCL) constantly up to date to keep track of network changes. The time needed by each serving node to update its NCL is an important parameter of the network’s reliability and performance. An adequate estimate of such parameter enables a significant improvement of self‐configuration functionalities. This paper focuses on the update time of NCLs when an approach of crowdsourced user reports is adopted. In this setting, each user periodically reports to the serving node information about the set of nodes sensed by the user itself. We show that, by mapping the local topological structure of the network onto states of increasing knowledge, a crisp mathematical framework can be obtained, which allows in turn for the use of a variety of user mobility models. Further, using a simplified mobility model we show how to obtain useful upper bounds on the expected time for a serving node to gain Full Knowledge of its local neighbourhood. Alessandro Checco, Carlo Lancia, Douglas J. Leith |
Wirel. Commun. Mob. Comput. | 1 |
| 2017 | Let's Agree to Disagree: Fixing Agreement Measures for CrowdsourcingabstractIn the context of micro-task crowdsourcing, each task is usually performed by several workers. This allows researchers to leverage measures of the agreement among workers on the same task, to estimate the reliability of collected data and to better understand answering behaviors of the participants. While many measures of agreement between annotators have been proposed, they are known for suffering from many problems and abnormalities. In this paper, we identify the main limits of the existing agreement measures in the crowdsourcing context, both by means of toy examples as well as with real-world crowdsourcing data, and propose a novel agreement measure based on probabilistic parameter estimation which overcomes such limits. We validate our new agreement measure and show its flexibility as compared to the existing agreement measures. Alessandro Checco, Kevin Roitero, Eddy Maddalena, Stefano Mizzaro, Gianluca Demartini |
HCOMP | 1 |
| 2017 | Recommending access points to individual mobile users via automatic group learningabstractWe consider user to cell association in a heterogeneous network with a mix of LTE/3G and WiFi cells. Individual user preferences are often neglected when a user to cell association decision is made. In this paper we propose use of a recommender system to inform the mapping of users to cells. We demonstrate the effectiveness of the proposed grouped-based user to cell associations for a set of synthetically generated user/cell ratings. Bahar Partov, Douglas J. Leith, Alessandro Checco |
ICC | 3 |
| 2017 | BLC: Private Matrix Factorization Recommenders via Automatic Group LearningabstractWe propose a privacy-enhanced matrix factorization recommender that exploits the fact that users can often be grouped together by interest. This allows a form of “hiding in the crowd” privacy. We introduce a novel matrix factorization approach suited to making recommendations in a shared group (or “nym”) setting and the BLC algorithm for carrying out this matrix factorization in a privacy-enhanced manner. We demonstrate that the increased privacy does not come at the cost of reduced recommendation accuracy. Alessandro Checco, Giuseppe Bianchi 0001, Douglas J. Leith |
ACM Trans. Priv. Secur. | 1 |
| 2017 | Analysis of Dynamic Channel Bonding in Dense Networks of WLANsabstractDynamic Channel Bonding (DCB) allows for the dynamic selection and use of multiple contiguous basic channels in Wireless Local Area Networks (WLANs). A WLAN operating under DCB can enjoy a larger bandwidth, when available, and therefore achieve a higher throughput. However, the use of larger bandwidths also increases the contention with adjacent WLANs, which can result in longer delays in accessing the channel and consequently, a lower throughput. In this paper, a scenario consisting of multiple WLANs using DCB and operating within carrier-sensing range of one another is considered. An analytical framework for evaluating the performance of such networks is presented. The analysis is carried out using a Markov chain model that characterizes the interactions between adjacent WLANs with overlapping channels. An algorithm is proposed for systematically constructing the Markov chain corresponding to any given scenario. The analytical model is then used to highlight and explain the key properties that differentiate DCB networks of WLANs from those operating on a single shared channel. Furthermore, the analysis is applied to networks of IEEE 802.11ac WLANs operating under DCB-which do not fully comply with some of the simplifying assumptions in our analysis-to show that the analytical model can give accurate results in more realistic scenarios. Azadeh Faridi, Boris Bellalta, Alessandro Checco |
IEEE Trans. Mob. Comput. | 3 |
| 2017 | Fast, Responsive Decentralized Graph ColoringabstractGraph coloring problem arises in numerous networking applications. We solve it in a fully decentralized way (ı.e., with no message passing). We propose a novel algorithm that is automatically responsive to topology changes, and we prove that it converges to a proper coloring in O(N log N) time with high probability for generic graphs, when the number of available colors is greater than Δ, the maximum degree of the graph, and in O(logN) time if Δ = O(1). We believe the proof techniques used in this paper are of independent interest and provide new insight into the properties required to ensure fast convergence of decentralized algorithms. Alessandro Checco, Douglas J. Leith |
IEEE/ACM Trans. Netw. | 1 |
| 2015 | Fair Virtualization of 802.11 NetworksabstractWe consider virtualization of network capacity in 802.11 WLANs and mesh networks. We show that allocating total airtime slices to ISPs is analogous to allocating a fraction of available time-slots in TDMA. We establish that the max-min fair flow rate allocation within an ISP airtime slice can be characterized independently of the rate allocation policy employed in other slices. Building on these observations, we present a lightweight, distributed algorithm for allocating airtime slices among ISP and max-min fair flow rates within each slice. Alessandro Checco, Douglas J. Leith |
IEEE/ACM Trans. Netw. | 1 |
| 2012 | Self-configuration of scrambling codes for WCDMA small cell networksabstractThis paper introduces the problem of Primary Scrambling Code (PSC) selection in small cell networks and proposes a novel optimisation technique. Small cells introduce challenges not present in conventional macrocell scrambling code allocation, including the need for dynamic allocation, scalable distributed allocation algorithms, and support for unplanned and organic deployments. To the best of our knowledge this is the first study addressing the issue of distributed scrambling code selection for small cell networks. We propose a decentralized learning algorithm which does not require any collaboration between the neighbouring base-stations and which finds a feasible allocation whenever one exists. The performance of the algorithm is compared against two variations of a greedy algorithm which is the current 3GPP recommendation and is shown to offer significant performance benefits. Alessandro Checco, Rouzbeh Razavi, Douglas J. Leith, Holger Claussen 0001 |
PIMRC | 1 |