VLDB 2026 Research / reviewers in the wild / expert
Gianluca Demartini
dblp:05/3422
· DBLP profile ↗
91ranked-venue papers in the field
14as first author
32since 2021 · last 2026
0000-0002-7311-3693ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 65 (11 first)Data Mining & Knowledge Discovery · 10Database Systems & Data Management · 7 (3 first)Knowledge Engineering, Semantic Web & Information Systems · 7Business Process & Enterprise Data · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Query-Document Dense Vectors for LLM Relevance Judgment Bias Analysis
Samaneh Mohtadi, Gianluca Demartini |
ECIR (2) | 2 |
| 2026 | The Effect of Document Summarization on LLM-Based Relevance Judgments
Samaneh Mohtadi, Kevin Roitero, Stefano Mizzaro, Gianluca Demartini |
ECIR (2) | 4 |
| 2026 | Large Language Models as Assessors: On the Impact of Relevance Scales
Riccardo Zamolo, Riccardo Lunardi, Michael Soprano, Gianluca Demartini, Stefano Mizzaro, Kevin Roitero |
ECIR (2) | 4 |
| 2025 | How Do Experts Make Sense of Integrated Process Models?
Tianwa Chen, Barbara Weber, Graeme G. Shanks, Gianluca Demartini, Marta Indulska, Shazia Sadiq |
CAiSE (2) | 4 |
| 2025 | LLM-Based Semantic Augmentation for Harmful Content DetectionabstractRecent advances in large language models (LLMs) have demonstrated strong performance on simple text classification tasks, frequently under zero-shot settings. However, their efficacy declines when tackling complex social media challenges such as propaganda detection, hateful meme classification, and toxicity identification. Much of the existing work has focused on using LLMs to generate synthetic training data, overlooking the potential of LLM-based text preprocessing and semantic augmentation. In this paper, we introduce an approach that prompts LLMs to clean noisy text and provide context-rich explanations, thereby enhancing training sets without substantial increases in data volume. We systematically evaluate on the SemEval 2024 multi-label Persuasive Meme dataset and further validate on the Google Jigsaw toxic comments and Facebook hateful memes datasets to assess generalizability. Our results reveal that zero-shot LLM classification underperforms on these high-context tasks compared to supervised models. In contrast, integrating LLM-based semantic augmentation yields performance on par with approaches that rely on human-annotated data, at a fraction of the cost. These findings underscore the importance of strategically incorporating LLMs into machine learning (ML) pipeline for social media classification tasks, offering broad implications for combating harmful content online. Disclaimer: This paper contains examples of explicit language that may be disturbing to some readers. Elyas Meguellati, Assaad Oussama Zeghina, Shazia Sadiq, Gianluca Demartini |
ICWSM | 4 |
| 2025 | Perception of Visual Content: Differences Between Humans and Foundation ModelsabstractHuman-annotated content is often used to train machine learning (ML) models. However, recently, language and multi-modal foundational models have been used to replace and scale-up human annotator's efforts. This study explores the similarity between human-generated and ML-generated annotations of images across diverse socio-economic contexts (RQ1) and their impact on ML model performance and bias (RQ2). We aim to understand differences in perception and identify potential biases in content interpretation. Our dataset comprises images of people from various geographical regions and income levels, covering various daily activities and home environments. ML captions and human labels show highest similarity at a low-level, i.e., types of words that appear and sentence structures, but all annotations are consistent in how they perceive images across regions. ML Captions resulted in best overall region classification performance, while ML Objects and ML Captions performed best overall for income regression. ML annotations worked best for action categories, while human input was more effective for non-action categories. These findings highlight the notion that both human and machine annotations are important, and that human-generated annotations are yet to be replaceable. Nardiena A. Pratama, Shaoyang Fan, Gianluca Demartini |
ICWSM | 3 |
| 2025 | Provoking critical thinking: Using counter-arguments in online discussion summarisationabstractGenerative AI systems based on Large Language Models (LLMs), like ChatGPT, have brought profound convenience to users thanks to their ability to summarise existing documents and to generate new text. This shows the potential to summarise online human discussions or debates for new entrants to quickly comprehend the ongoing matters and arguments and to get efficiently involved in the opinion deliberation process. However, generative AI has frequently been associated with negatively affecting users’ decision making. In this paper, we study a novel approach based on generative AI to trigger users’ critical thinking by challenging fresh counter-arguments after summarising existing online discussions for incoming users. We conduct a user study with 558 participants to determine the effectiveness and fairness of AI summarisation across three online platforms — Reddit, Kialo, and Debatewise. Our results show that the intervention methods and platform differences are strongly associated with participants’ level of opinion change and the strength of their belief. We found that participants’ opinion changes affected their perceived usefulness of the AI system. Our work opens the door to LLM applications helping Web users participate in online opinion deliberation more efficiently, with a higher level of critical thinking, and with a reduced negative attitude. Shangqian Li, Lei Han 0003, Gianluca Demartini |
Inf. Process. Manag. | 3 |
| 2025 | Enhancing media literacy: The effectiveness of (Human) annotations and bias visualizations on bias detectionabstractMarking biased texts effectively increases media bias awareness, but its sustainability across new topics and unmarked news remains unclear, and the role of AI-generated bias labels is untested. This study examines how news consumers learn to perceive media bias from human- and AI-generated labels and identify biased language through highlighting, neutral rephrasing, and political orientation cues. We conducted two experiments with a teaching phase exposing them to various bias-labeling conditions and a testing phase evaluating their ability to classify biased sentences and detect biased text in unlabeled news on new topics. We find that, compared to the control group, both human- and AI-generated sentential bias labels significantly improve bias classification ( p < .001), though human labels are more effective ( d = 0.42 vs. d = 0.23). Additionally, among all teaching interventions, participants best detect biased sentences when taught with biased sentence or phrase labels ( p < .001), while politicized phrase labels reduce accuracy. The effectiveness of different media literacy interventions remains independent of political ideology, but conservative participants are generally less accurate ( p = .011), suggesting an interaction between political inclinations and bias detection. Our research provides a novel experimental framework into assessing the generalizability of media bias awareness and offer practical implications for designing bias indicators in news-reading platforms and media literacy curricula. Timo Spinde, Wolfgang Gaissmaier, Gianluca Demartini, Isao Echizen, Helge Giese |
Inf. Process. Manag. | 4 |
| 2024 | Generative AI in Crowdwork for Web and Social Media Research: A Survey of Workers at Three PlatformsabstractCrowdsourcing plays an important role in Web and social media research, from data annotation, to online experiments and user surveys. With the emergence of Generative AI (GenAI), researchers are considering how models and tools such as GPT might replace crowdwork. Many have already evaluated GPT on annotation tasks. However, it is less clear how GenAI might impact other types of tasks, or to what extent crowdworkers have already incorporated it into their work processes. Thus, we asked crowdworkers directly regarding their use of GenAI, via a survey at two points in time, across three commercial platforms. We found evidence that workers' self-reported use of GenAI did not change over time, but rather, was strongly correlated to the platform in which they operate, with MTurk workers using GenAI much more often than those operating at Clickworker and Prolific. As most respondents reported that survey completion is their "usual type of task", we discuss the implication of the use of GenAI in user surveys, via specific examples of ICWSM research. Evgenia Christoforou, Gianluca Demartini, Jahna Otterbacher |
ICWSM | 2 |
| 2024 | On the Role of Large Language Models in Crowdsourcing Misinformation AssessmentabstractThe proliferation of online misinformation significantly undermines the credibility of web content. Recently, crowd workers have been successfully employed to assess misinformation to address the limited scalability of professional fact-checkers. An alternative approach to crowdsourcing is the use of large language models (LLMs). These models are however also not perfect. In this paper, we investigate the scenario of crowd workers working in collaboration with LLMs to assess misinformation. We perform a study where we ask crowd workers to judge the truthfulness of statements under different conditions: with and without LLMs labels and explanations. Our results show that crowd workers tend to overestimate truthfulness when exposed to LLM-generated information. Crowd workers are misled by wrong LLM labels, but, on the other hand, their self-reported confidence is lower when they make mistakes due to relying on the LLM. We also observe diverse behaviors among crowd workers when the LLM is presented, indicating that leveraging LLMs can be considered a distinct working strategy. Jiechen Xu, Lei Han 0003, Shazia Sadiq, Gianluca Demartini |
ICWSM | 4 |
| 2024 | Hate Speech Detection with Generalizable Target-aware FairnessabstractTo counter the side effect brought by the proliferation of social media platforms, hate speech detection (HSD) plays a vital role in halting the dissemination of toxic online posts at an early stage. However, given the ubiquitous topical communities on social media, a trained HSD classifier can easily become biased towards specific targeted groups (e.g.,female andblack people), where a high rate of either false positive or false negative results can significantly impair public trust in the fairness of content moderation mechanisms, and eventually harm the diversity of online society. Although existing fairness-aware HSD methods can smooth out some discrepancies across targeted groups, they are mostly specific to a narrow selection of targets that are assumed to be known and fixed. This inevitably prevents those methods from generalizing to real-world use cases where new targeted groups constantly emerge (e.g., new forums created on Reddit) over time. To tackle the defects of existing HSD practices, we propose Generalizable target-aware Fairness (GetFair), a new method for fairly classifying each post that contains diverse and even unseen targets during inference. To remove the HSD classifier's spurious dependence on target-related features, GetFair trains a series of filter functions in an adversarial pipeline, so as to deceive the discriminator that recovers the targeted group from filtered post embeddings. To maintain scalability and generalizability, we innovatively parameterize all filter functions via a hypernetwork. Taking a target's pretrained word embedding as input, the hypernetwork generates the weights used by each target-specific filter on-the-fly without storing dedicated filter parameters. In addition, a novel semantic gap alignment scheme is imposed on the generation process, such that the produced filter function for an unseen target is rectified by its semantic affinity with existing targets used for training. Finally, experiments are conducted on two benchmark HSD datasets, showing advantageous performance of GetFair on out-of-sample targets among baselines. Tong Chen 0005, Danny Wang, Xurong Liang, Marten Risius, Gianluca Demartini, Hongzhi Yin |
KDD | 5 |
| 2024 | Crowdsourced Fact-checking: Does It Actually Work?abstractThere is an important ongoing effort aimed to tackle misinformation and to perform reliable fact-checking by employing human assessors at scale, with a crowdsourcing-based approach. Previous studies on the feasibility of employing crowdsourcing for the task of misinformation detection have provided inconsistent results: some of them seem to confirm the effectiveness of crowdsourcing for assessing the truthfulness of statements and claims, whereas others fail to reach an effectiveness level higher than automatic machine learning approaches, which are still unsatisfactory. In this paper, we aim at addressing such inconsistency and understand if truthfulness assessment can indeed be crowdsourced effectively. To do so, we build on top of previous studies; we select some of those reporting low effectiveness levels, we highlight their potential limitations, and we then reproduce their work attempting to improve their setup to address those limitations. We employ various approaches, data quality levels, and agreement measures to assess the reliability of crowd workers when assessing the truthfulness of (mis)information. Furthermore, we explore different worker features and compare the results obtained with different crowds. According to our findings, crowdsourcing can be used as an effective methodology to tackle misinformation at scale. When compared to previous studies, our results indicate that a significantly higher agreement between crowd workers and experts can be obtained by using a different, higher-quality, crowdsourcing platform and by improving the design of the crowdsourcing task. Also, we find differences concerning task and worker features and how workers provide truthfulness assessments. David La Barbera, Eddy Maddalena, Michael Soprano, Kevin Roitero, Gianluca Demartini, Davide Ceolin, Damiano Spina, Stefano Mizzaro |
Inf. Process. Manag. | 5 |
| 2024 | Cognitive Biases in Fact-Checking and Their Countermeasures: A ReviewabstractThe increase of the amount of misinformation spread every day online is a huge threat to the society. Organizations and researchers are working to contrast this misinformation plague. In this setting, human assessors are indispensable to correctly identify, assess and/or revise the truthfulness of information items, i.e., to perform the fact-checking activity. Assessors, as humans, are subject to systematic errors that might interfere with their fact-checking activity. Among such errors, cognitive biases are those due to the limits of human cognition. Although biases help to minimize the cost of making mistakes, they skew assessments away from an objective perception of information. Cognitive biases, hence, are particularly frequent and critical, and can cause errors that have a huge potential impact as they propagate not only in the community, but also in the datasets used to train automatic and semi-automatic machine learning models to fight misinformation. In this work, we present a review of the cognitive biases which might occur during the fact-checking process. In more detail, inspired by PRISMA – a methodology used for systematic literature reviews – we manually derive a list of 221 cognitive biases that may affect human assessors. Then, we select the 39 biases that might manifest during the fact-checking process, we group them into categories, and we provide a description. Finally, we present a list of 11 countermeasures that can be adopted by researchers, practitioners, and organizations to limit the effect of the identified cognitive biases on the fact-checking activity. Michael Soprano, Kevin Roitero, David La Barbera, Davide Ceolin, Damiano Spina, Gianluca Demartini, Stefano Mizzaro |
Inf. Process. Manag. | 6 |
| 2024 | How Many Crowd Workers Do I Need? On Statistical Power when Crowdsourcing Relevance JudgmentsabstractTo scale the size of Information Retrieval collections, crowdsourcing has become a common way to collect relevance judgments at scale. Crowdsourcing experiments usually employ 100–10,000 workers, but such a number is often decided in a heuristic way. The downside is that the resulting dataset does not have any guarantee of meeting predefined statistical requirements as, for example, have enough statistical power to be able to distinguish in a statistically significant way between the relevance of two documents. We propose a methodology adapted from literature on sound topic set size design, based on t-test and ANOVA, which aims at guaranteeing the resulting dataset to meet a predefined set of statistical requirements. We validate our approach on several public datasets. Our results show that we can reliably estimate the recommended number of workers needed to achieve statistical power, and that such estimation is dependent on the topic, while the effect of the relevance scale is limited. Furthermore, we found that such estimation is dependent on worker features such as agreement. Finally, we describe a set of practical estimation strategies that can be used to estimate the worker set size, and we also provide results on the estimation of document set sizes. Kevin Roitero, David La Barbera, Michael Soprano, Gianluca Demartini, Stefano Mizzaro, Tetsuya Sakai |
ACM Trans. Inf. Syst. | 4 |
| 2024 | On the Impact of Showing Evidence from Peers in Crowdsourced Truthfulness AssessmentsabstractMisinformation has been rapidly spreading online. The common approach to dealing with it is deploying expert fact-checkers who follow forensic processes to identify the veracity of statements. Unfortunately, such an approach does not scale well. To deal with this, crowdsourcing has been looked at as an opportunity to complement the work done by trained journalists. In this article, we look at the effect of presenting the crowd with evidence from others while judging the veracity of statements. We implement variants of the judgment task design to understand whether and how the presented evidence may or may not affect the way crowd workers judge truthfulness and their performance. Our results show that, in certain cases, the presented evidence and the way in which it is presented may mislead crowd workers who would otherwise be more accurate if judging independently from others. Those who make appropriate use of the provided evidence, however, can benefit from it and generate better judgments. Jiechen Xu, Lei Han 0003, Shazia Sadiq, Gianluca Demartini |
ACM Trans. Inf. Syst. | 4 |
| 2023 | Firearms on Twitter: A Novel Object Detection PipelineabstractSocial media is an important source of real-time imagery concerning world events. One subset of social media posts which may be of particular interest are those featuring firearms. These posts can give insight into weapon movements, troop activity and civilian safety. Object detection tools offer important opportunities for insight into these images. Unfortunately, these images can be visually complex, poorly lit and generally challenging for object detection models. We present an analysis of existing gun detection datasets, and find that these datasets to not effectively address the challenge of gun detection on real-life images. Following this, we present a novel object detection pipeline. We train our pipeline on a number of datasets including one created for this investigation made up of Twitter images of the Russo-Ukrainian War. We compare the performance of our model as trained on the different datasets to baseline numbers provided by original authors as well as a YOLO v5 benchmark. We find that our model outperforms the state-of-the-art benchmarks on contextually rich, real-life-derived imagery of firearms. Ryan Harvey, Rémi Lebret, Stéphane Massonnet, Karl Aberer, Gianluca Demartini |
ICWSM | 5 |
| 2023 | On the Impact of Data Quality on Image Classification FairnessabstractWith the proliferation of algorithmic decision-making, increased scrutiny has been placed on these systems. This paper explores the relationship between the quality of the training data and the overall fairness of the models trained with such data in the context of supervised classification. We measure key fairness metrics across a range of algorithms over multiple image classification datasets that have a varying level of noise in both the labels and the training data itself. We describe noise in the labels as inaccuracies in the labelling of the data in the training set and noise in the data as distortions in the data, also in the training set. By adding noise to the original datasets, we can explore the relationship between the quality of the training data and the fairness of the output of the models trained on that data. Aki Barry, Lei Han 0003, Gianluca Demartini |
SIGIR | 3 |
| 2023 | Human-in-the-loop Regular Expression Extraction for Single Column Format InconsistencyabstractFormat inconsistency is one of the most frequently appearing data quality issues encountered during data cleaning. Existing automated approaches commonly lack applicability and generalisability, while approaches with human inputs typically require specialized skills such as writing regular expressions. This paper proposes a novel hybrid human-machine system, namely “Data-Scanner-4C”, which leverages crowdsourcing to address syntactic format inconsistencies in a single column effectively. We first ask crowd workers to create examples from single-column data through “data selection” and “result validation” tasks. Then, we propose and use a novel rule-based learning algorithm to infer the regular expressions that propagate formats from created examples to the entire column. Our system integrates crowdsourcing and algorithmic format extraction techniques in a single workflow. Having human experts write regular expressions is no longer required, thereby reducing both the time as well as the opportunity for error. We conducted experiments through both synthetic and real-world datasets, and our results show how the proposed approach is applicable and effective across data types and formats. Shaochen Yu, Lei Han 0003, Marta Indulska, Shazia Sadiq, Gianluca Demartini |
WWW | 5 |
| 2023 | On the role of human and machine metadata in relevance judgment tasks
Jiechen Xu, Lei Han 0003, Shazia Sadiq, Gianluca Demartini |
Inf. Process. Manag. | 4 |
| 2023 | DataOps-4G: On Supporting Generalists in Data Quality DiscoveryabstractData preparation has become a necessary but labor and resource intensive step to perform data analytics. To date, such activities still require considerable manual effort from experts. In this paper, we focus on a specific data preparation activity, namely data quality discovery. We explore different settings in which data workers undertake data quality discovery tasks and the implications of those settings for the efficiency and effectiveness of data workers. To this end, we propose DataOps-4G, a data curation platform for generalists, that allows users to interact with data without the need to write code. We wrap up pre-defined code snippets that implement useful functionalities to explore data quality and bundle the code into so-called DataOps. Then, we conduct a lab-based user study to evaluate our DataOps-4G platform from two perspectives: (i) effectiveness, the accuracy of the outcomes achieved by participants; and (ii) efficiency, their effort and strategies in task completion. Our experimental results uncover how effectiveness and efficiency can be affected by their task completion patterns and strategies. This opens up the possibility of popularizing data curation processes by employing non-experts (e.g., from crowdsourcing platforms) and consequently allowing experts to focus on more complex activities (e.g., building machine learning models). Shaochen Yu, Tianwa Chen, Lei Han 0003, Gianluca Demartini, Shazia Sadiq |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | A Data-Driven Analysis of Behaviors in Data Curation ProcessesabstractUnderstanding how data workers interact with data, and various pieces of information related to data preparation, is key to designing systems that can better support them in exploring datasets. To date, however, there is a paucity of research studying the strategies adopted by data workers as they carry out data preparation activities. In this work, we investigate a specific data preparation activity, namely data quality discovery , and aim to (i) understand the behaviors of data workers in discovering data quality issues, (ii) explore what factors (e.g., prior experience) can affect their behaviors, as well as (iii) understand how these behavioral observations relate to their performance. To this end, we collect a multi-modal dataset through a data-driven experiment that relies on the use of eye-tracking technology with a purpose-designed platform built on top of iPython Notebook. The experiment results reveal that: (i) ‘copy–paste–modify’ is a typical strategy for writing code to complete tasks; (ii) proficiency in writing code has a significant impact on the quality of task performance, while perceived difficulty and efficacy can influence task completion patterns; and (iii) searching in external resources is a prevalent action that can be leveraged to achieve better performance. Furthermore, our experiment indicates that providing sample code within the system can help data workers get started with their task, and surfacing underlying data is an effective way to support exploration. By investigating data worker behaviors prior to each search action, we also find that the most common reasons that trigger external search actions are the need to seek assistance in writing or debugging code and to search for relevant code to reuse. Based on our experiment results, we showcase a systematic approach to select from the top best code snippets created by data workers and assemble them to achieve better performance than the best individual performer in the dataset. By doing so, our findings not only provide insights into patterns of interactions with various system components and information resources when performing data curation tasks, but also build effective and efficient data curation processes through data workers’ collective intelligence. Lei Han 0003, Tianwa Chen, Gianluca Demartini, Marta Indulska, Shazia Sadiq |
ACM Trans. Inf. Syst. | 3 |
| 2022 | Crowdsourced Fact-Checking at Twitter: How Does the Crowd Compare With Experts?abstractFact-checking is one of the effective solutions in fighting online misinformation. However, traditional fact-checking is a process requiring scarce expert human resources, and thus does not scale well on social media because of the continuous flow of new content to be checked. Methods based on crowdsourcing have been proposed to tackle this challenge, as they can scale with a smaller cost, but, while they have shown to be feasible, have always been studied in controlled environments. In this work, we study the first large-scale effort of crowdsourced fact-checking deployed in practice, started by Twitter with the Birdwatch program. Our analysis shows that crowdsourcing may be an effective fact-checking strategy in some settings, even comparable to results obtained by human experts, but does not lead to consistent, actionable results in others. We processed 11.9k tweets verified by the Birdwatch program and report empirical evidence of i) differences in how the crowd and experts select content to be fact-checked, ii) how the crowd and the experts retrieve different resources to fact-check, and iii) the edge the crowd shows in fact-checking scalability and efficiency as compared to expert checkers. Mohammed Saeed 0002, Nicolas Traub, Maelle Nicolas, Gianluca Demartini, Paolo Papotti |
CIKM | 4 |
| 2022 | Workshop on Human-in-the-loop Data CurationabstractAlthough data quality is a long-standing and enduring problem, it has recently received a resurgence of attention due to the fast proliferation of data analytics, machine learning, and decision-support applications built upon the wide-scale availability and accessibility of (big) data. The success of such applications heavily relies on not only the quantity, but also the quality of data. Data curation, which may include annotation, cleaning, transformation, integration, etc., is a critical step to provide adequate assurances on the quality of analytics and machine learning results. Such data preparation activities are recognised as time and resource intensive for data scientists as data often comes with a number of challenges that need to be tackled before it can be used in practice. Data re-purposing and the resulting distance between design and use intentions of the data, is a fundamental issue behind many of these challenges. These challenges include a variety of data issues such as noise and outliers, incompleteness, representativeness or biases, heterogeneity of format or semantics, etc. Mishandling these challenges can lead to negative and sometimes damaging effects, especially in critical domains like healthcare, transport, and finance. An observable distinct feature of data quality in these contexts is the increasingly important role played by humans, being often the source of data generation and the active players in data curation. This workshop will provide an opportunity to explore the interdisciplinary overlap between manual, automated, and hybrid human-machine methods of data curation. Gianluca Demartini, Jie Yang 0028, Shazia Sadiq |
CIKM | 1 |
| 2022 | How Does the Crowd Impact the Model? A Tool for Raising Awareness of Social Bias in Crowdsourced Training DataabstractIt is increasingly easy for interested parties to play a role in the development of predictive algorithms, with a range of available tools and platforms for building datasets, as well as for training and evaluating machine learning (ML) models. For this reason, it is essential to create awareness among practitioners on the ethical challenges, such as the presence of social bias in training data. We present RECANT (Raising Awareness of Social Bias in Crowdsourced Training Data), a tool that allows users to explore the behaviors of four biometric models -- predicting the gender and race, as well as the perceived attractiveness and trustworthiness, of the person depicted in an input image. These models have been trained on a crowdsourced dataset of passport-style people images, where crowd annotators described attributes of the images, and reported their own demographic characteristics. With RECANT, users can explore the correct and wrong predictions made by each model, when using different subsets of the data in training, based on annotator attributes. We present its features, along with sample exercises, as a hands-on tool for raising awareness of potential pitfalls in data practices surrounding ML. Periklis Perikleous, Andreas Kafkalias, Zenonas Theodosiou, Pinar Barlas, Evgenia Christoforou, Jahna Otterbacher, Gianluca Demartini, Andreas Lanitis |
CIKM | 7 |
| 2022 | A Behavioural Analysis of Metadata Use in Evaluating the Quality of Repurposed Data
Lei Han 0003, Gianluca Demartini, Marta Indulska, Shazia Sadiq |
ER | 3 |
| 2022 | Exploring Data Literacy Levels in the Crowd - the Case of COVID-19
Shaoyang Fan, Lei Han 0003, Gianluca Demartini, Shazia Sadiq |
ICWSM | 3 |
| 2022 | Preferences on a Budget: Prioritizing Document Pairs when Crowdsourcing Relevance JudgmentsabstractIn Information Retrieval (IR) evaluation, preference judgments are collected by presenting to the assessors a pair of documents and asking them to select which of the two, if any, is the most relevant. This is an alternative to the classic relevance judgment approach, in which human assessors judge the relevance of a single document on a scale; such an alternative allows to make relative rather than absolute judgments of relevance. While preference judgments are easier for human assessors to perform, the number of possible document pairs to be judged is usually so high that it makes it unfeasible to judge them all. Thus, following a similar idea to pooling strategies for single document relevance judgments where the goal is to sample the most useful documents to be judged, in this work we focus on analyzing alternative ways to sample document pairs to judge, in order to maximize the value of a fixed number of preference judgments that can feasibly be collected. Such value is defined as how well we can evaluate IR systems given a budget, that is, a fixed number of human preference judgments that may be collected. By relying on several datasets featuring relevance judgments gathered by means of experts and crowdsourcing, we experimentally compare alternative strategies to select document pairs and show how different strategies lead to different IR evaluation result quality levels. Our results show that, by using the appropriate procedure, it is possible to achieve good IR evaluation results with a limited number of preference judgments, thus confirming the feasibility of using preference judgments to create IR evaluation collections. Kevin Roitero, Alessandro Checco, Stefano Mizzaro, Gianluca Demartini |
WWW | 4 |
| 2022 | Information Resilience: the nexus of responsible and agile approaches to information useabstractAbstract The appetite for effective use of information assets has been steadily rising in both public and private sector organisations. However, whether the information is used for social good or commercial gain, there is a growing recognition of the complex socio-technical challenges associated with balancing the diverse demands of regulatory compliance and data privacy, social expectations and ethical use, business process agility and value creation, and scarcity of data science talent. In this vision paper, we present a series of case studies that highlight these interconnected challenges, across a range of application areas. We use the insights from the case studies to introduce Information Resilience, as a scaffold within which the competing requirements of responsible and agile approaches to information use can be positioned. The aim of this paper is to develop and present a manifesto for Information Resilience that can serve as a reference for future research and development in relevant areas of responsible data management. Shazia Sadiq, Amir Aryani, Gianluca Demartini, Wen Hua, Marta Indulska, Andrew Burton-Jones, Hassan Khosravi, Diana Benavides-Prado, Timos K. Sellis, Ida Asadi Someh, Rhema Vaithianathan, Sen Wang 0001, Xiaofang Zhou 0001 |
VLDB J. | 3 |
| 2021 | CoralExp: An Explainable System to Support Coral Taxonomy Research
Jaiden Harding, Tom Bridge, Gianluca Demartini |
ECIR (2) | 3 |
| 2021 | Iterative Human-in-the-Loop Discovery of Unknown Unknowns in Image DatasetsabstractAutomatic predictions (e.g., recognizing objects in images) may result in systematic errors if certain classes are not well represented by training instances (these errors are called unknowns). When a model assigns high confidence scores to these wrong predictions (this type of error is called unknown unknowns), it becomes challenging to automatically identify them. In this paper, we present the first work on leveraging human intelligence to discover unknown unknowns (UUs) in an iterative way. The proposed methodology first differentiates the feature space generated by crowd workers labelling instances (e.g., images) in an active learning fashion from the space learned by the prediction model over a batch training phase, and thus identifies the predictions most likely to be UUs. Next, we add crowd labels collected for these discovered UUs to the training set and re-train the model with this extended dataset. This process is then repeated iteratively to discover more instances of both unknown and under-represented classes. Our experimental results show that the proposed methodology is able to (1) efficiently discover UUs, (2) significantly improve the quality of model predictions, and (3) to push UUs into known unknowns (i.e., the model makes mistakes but at least its classification confidence on those instances is low so those predictions can be discarded or post-processed) for further investigation. We additionally discuss the trade-off between prediction quality improvements and the human effort required to achieve those improvements. Our results bear implications on building cost-effective systems to discover UUs with humans in the loop. Lei Han 0003, Gianluca Demartini |
HCOMP | 3 |
| 2021 | The many dimensions of truthfulness: Crowdsourcing misinformation assessments on a multidimensional scale
Michael Soprano, Kevin Roitero, David La Barbera, Davide Ceolin, Damiano Spina, Stefano Mizzaro, Gianluca Demartini |
Inf. Process. Manag. | 7 |
| 2021 | The Impact of Task Abandonment in CrowdsourcingabstractCrowdsourcing has become a standard methodology to collect manually annotated data such as relevance judgments at scale. On crowdsourcing platforms like Amazon MTurk or FigureEight, crowd workers select tasks to work on based on different dimensions such as task reward and requester reputation. Requesters then receive the judgments of workers who self-selected into the tasks and completed them successfully. Several crowd workers, however, preview tasks, begin working on them, reaching varying stages of task completion without finally submitting their work. Such behavior results in unrewarded effort which remains invisible to requesters. In this paper, we conduct an investigation of the phenomenon of task abandonment, the act of workers previewing or beginning a task and deciding not to complete it. We follow a three-fold methodology which includes 1) investigating the prevalence and causes of task abandonment by means of a survey over different crowdsourcing platforms, 2) data-driven analysis of logs collected during a large-scale relevance judgment experiment, and 3) controlled experiments measuring the effect of different dimensions on abandonment. Our results show that task abandonment is a widely spread phenomenon. Apart from accounting for a considerable amount of wasted human effort, this bears important implications on the hourly wages of workers as they are not rewarded for tasks that they do not complete. We also show how task abandonment may have strong implications on the use of collected data (for example, on the evaluation of Information Retrieval systems). Lei Han 0003, Kevin Roitero, Ujwal Gadiraju, Cristina Sarasua, Alessandro Checco, Eddy Maddalena, Gianluca Demartini |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2020 | Modelling User Behavior Dynamics with EmbeddingsabstractUnderstanding user interaction behaviors remains a challenging problem. Quantifying behavior dynamics over time as users complete tasks has only been done in specific domains. In this paper, we present a user behavior model built using behavior embeddings to compare behaviors and their change over time. To this end, we first define the formal model and train the model using both action (e.g., copy/paste) embeddings and user interaction feature (e.g., length of the copied text) embeddings. Having obtained vector representations of user behaviors, we then define three measurements to model behavior dynamics over time, namely: behavior position, displacement, and velocity. To evaluate the proposed methodology, we use three real world datasets: (i) tens of users completing complex data curation tasks in a lab setting, (ii) hundreds of crowd workers completing structured tasks in a crowdsourcing setting, and (iii) thousands of editors completing unstructured editing tasks on Wikidata. Through these datasets, we show that the proposed methodology can: (i) surface behavioral differences among users; (ii) recognize relative behavioral changes; and (iii) discover directional deviations of user behaviors. Our approach can be used (i) to capture behavioral semantics from data in a consistent way, (ii) to quantify behavioral diversity for a task and among different users, and (iii) to explore the temporal behavior evolution with respect to various task properties (e.g., structure and difficulty). Lei Han 0003, Alessandro Checco, Djellel Eddine Difallah, Gianluca Demartini, Shazia Sadiq |
CIKM | 4 |
| 2020 | The COVID-19 Infodemic: Can the Crowd Judge Recent Misinformation Objectively?abstractMisinformation is an ever increasing problem that is difficult to solve for the research community and has a negative impact on the society at large. Very recently, the problem has been addressed with a crowdsourcing-based approach to scale up labeling efforts: to assess the truthfulness of a statement, instead of relying on a few experts, a crowd of (non-expert) judges is exploited. We follow the same approach to study whether crowdsourcing is an effective and reliable method to assess statements truthfulness during a pandemic. We specifically target statements related to the COVID-19 health emergency, that is still ongoing at the time of the study and has arguably caused an increase of the amount of misinformation that is spreading online (a phenomenon for which the term "infodemic" has been used). By doing so, we are able to address (mis)information that is both related to a sensitive and personal issue like health and very recent as compared to when the judgment is done: two issues that have not been analyzed in related work.\n\nIn our experiment, crowd workers are asked to assess the truthfulness of statements, as well as to provide evidence for the assessments as a URL and a text justification. Besides showing that the crowd is able to accurately judge the truthfulness of the statements, we also report results on many different aspects, including: agreement among workers, the effect of different aggregation functions, of scales transformations, and of workers background / bias. We also analyze workers behavior, in terms of queries submitted, URLs found / selected, text justifications, and other behavioral data like clicks and mouse actions collected by means of an ad hoc logger. Kevin Roitero, Michael Soprano, Beatrice Portelli, Damiano Spina, Vincenzo Della Mea, Giuseppe Serra 0001, Stefano Mizzaro, Gianluca Demartini |
CIKM | 8 |
| 2020 | Crowdsourcing Truthfulness: The Impact of Judgment Scale and Assessor Bias
David La Barbera, Kevin Roitero, Gianluca Demartini, Stefano Mizzaro, Damiano Spina |
ECIR (2) | 3 |
| 2020 | On Understanding Data Worker Interaction BehaviorsabstractUnderstanding how data workers interact with data and various pieces of information (e.g., code snippet examples) is key to design systems that can better support them in exploring a given dataset. To date, however, there is a paucity of research studying information seeking patterns and the strategies adopted by data workers as they carry out data curation activities. In this work, we aim at understanding the behaviors of data workers in discovering data quality issues, and how these behavioral observations relate to their performance. Specifically, we investigate how data workers use information resources and tools to support their task completion. To this end, we collect a multi-modal dataset through a data-driven experiment that relies on the use of eye-tracking technology with a purpose-designed platform built on top of iPython Notebook. The collected data reveals that: (i) searching in external resources is a prevalent action that can be leveraged to achieve better performance; (ii) 'copy-paste-modify' is a typical strategy for writing code to complete tasks; (iii) providing sample code within the system could help data workers to get started with their task; and (iv) surfacing underlying data is an effective way to support exploration. By investigating the behaviors prior to each search action, we also find that the most common reasons that trigger external search actions are the need to seek assistance in writing or debugging code and to search for relevant code to reuse. Our findings provide insights into patterns of interactions with various system components and information resources to perform data curation tasks. This bears implications on the design of domain-specific IR systems for data workers like code-base search. Lei Han 0003, Tianwa Chen, Gianluca Demartini, Marta Indulska, Shazia Sadiq |
SIGIR | 3 |
| 2020 | Can The Crowd Identify Misinformation Objectively?: The Effects of Judgment Scale and Assessor's BackgroundabstractTruthfulness judgments are a fundamental step in the process of fighting misinformation, as they are crucial to train and evaluate classifiers that automatically distinguish true and false statements. Usually such judgments are made by experts, like journalists for political statements or medical doctors for medical statements. In this paper, we follow a different approach and rely on (non-expert) crowd workers. This of course leads to the following research question: Can crowdsourcing be reliably used to assess the truthfulness of information and to create large-scale labeled collections for information credibility systems? To address this issue, we present the results of an extensive study based on crowdsourcing: we collect thousands of truthfulness assessments over two datasets, and we compare expert judgments with crowd judgments, expressed on scales with various granularity levels. We also measure the political bias and the cognitive background of the workers, and quantify their effect on the reliability of the data provided by the crowd. Kevin Roitero, Michael Soprano, Shaoyang Fan, Damiano Spina, Stefano Mizzaro, Gianluca Demartini |
SIGIR | 6 |
| 2020 | Crowd Worker Strategies in Relevance Judgment TasksabstractCrowdsourcing is a popular technique to collect large amounts of human-generated labels, such as relevance judgments used to create information retrieval (IR) evaluation collections. Previous research has shown how collecting high quality labels from a crowdsourcing platform can be challenging. Existing quality assurance techniques focus on answer aggregation or on the use of gold questions where ground-truth data allows to check for the quality of the responses. Lei Han 0003, Eddy Maddalena, Alessandro Checco, Cristina Sarasua, Ujwal Gadiraju, Kevin Roitero, Gianluca Demartini |
WSDM | 7 |
| 2019 | On Transforming Relevance ScalesabstractInformation Retrieval (IR) researchers have often used existing IR evaluation collections and transformed the relevance scale in which judgments have been collected, e.g., to use metrics that assume binary judgments like Mean Average Precision. Such scale transformations are often arbitrary (e.g., 0,1 mapped to 0 and 2,3 mapped to 1) and it is assumed that they have no impact on the results of IR evaluation. Moreover, the use of crowdsourcing to collect relevance judgments has become a standard methodology. When designing the crowdsourcing relevance judgment task, one of the decision to be made is the how granular the relevance scale used to collect judgments should be. Such decision has then repercussions on the metrics used to measure IR system effectiveness. In this paper we look at the effect of scale transformations in a systematic way. We perform extensive experiments to study the transformation of judgments from fine-grained to coarse-grained. We use different relevance judgments expressed on different relevance scales and either expressed by expert annotators or collected by means of crowdsourcing. The objective is to understand the impact of relevance scale transformations on IR evaluation outcomes and to draw conclusions on how to best transform judgments into a different scale, when necessary. Lei Han 0003, Kevin Roitero, Eddy Maddalena, Stefano Mizzaro, Gianluca Demartini |
CIKM | 5 |
| 2019 | Health Card Retrieval for Consumer Health Search: An Empirical Investigation of MethodsabstractThis paper investigates methods to rank health cards, a domain-specific type of entity cards, for consumer health search (CHS) queries. A key challenge in this context is which card(s) should be presented to the user. In particular, little evidence exists to determine the effectiveness of retrieval and ranking methods for health cards in CHS. CHS is a challenging domain, where users lack domain expertise and thus are often unable to formulate effective queries, and to interpret the retrieved results. In addition, unlike in other contexts, CHS presents the opportunity to exploit a number of domain specific characteristics and features. In this paper, we focus on difficult queries with self-diagnosis intents. Our study makes the following contributions: (1) it assembles and releases the first test collection of health cards for research purposes, and (2) it empirically evaluates a large range of entity retrieval methods adapted to health cards retrieval, including features specific to health cards for learning to rank. This is the first study that thoroughly investigates methods to rank health cards. Jimmy, Guido Zuccon, Bevan Koopman, Gianluca Demartini |
CIKM | 4 |
| 2019 | Platform-Related Factors in Repeatability and Reproducibility of Crowdsourcing TasksabstractCrowdsourcing platforms provide a convenient and scalable way to collect human-generated labels on-demand. This data can be used to train Artificial Intelligence (AI) systems or to evaluate the effectiveness of algorithms. The datasets generated by means of crowdsourcing are, however, dependent on many factors that affect their quality. These include, among others, the population sample bias introduced by aspects like task reward, requester reputation, and other filters introduced by the task design.In this paper, we analyse platform-related factors and study how they affect dataset characteristics by running a longitudinal study where we compare the reliability of results collected with repeated experiments over time and across crowdsourcing platforms. Results show that, under certain conditions: 1) experiments replicated across different platforms result in significantly different data quality levels while 2) the quality of data from repeated experiments over time is stable within the same platform. We identify some key task design variables that cause such variations and propose an experimentally validated set of actions to counteract these effects thus achieving reliable and repeatable crowdsourced data collection experiments. Rehab K. Qarout, Alessandro Checco, Gianluca Demartini, Kalina Bontcheva |
HCOMP | 3 |
| 2019 | Non-parametric Class Completeness Estimators for Collaborative Knowledge Graphs - The Case of Wikidata
Michael Luggen, Djellel Eddine Difallah, Cristina Sarasua, Gianluca Demartini, Philippe Cudré-Mauroux |
ISWC (1) | 4 |
| 2019 | Health Cards for Consumer Health SearchabstractThis paper investigates the impact of health cards in consumer health search (CHS) - people seeking health advice online. Health cards are a concise presentations of a health concept shown along side search results to specific health queries; they have the potential to convey health information in easily digestible form for the general public. However, little evidence exists on how effective health cards actually are for users when searching health advice online, and whether their effectiveness is limited to specific health search intents. To understand the impact of health cards on CHS, we conducted a laboratory study to observe users completing CHS tasks using two search interface variants: one just with result snippets and one containing both result snippets and health cards. Our study makes the following contributions: (1) it reveals how and when health cards are beneficial to users in completing consumer health search tasks, and (2) it identifies the features of health cards that helped users in completing their tasks. This is the first study that thoroughly investigates the effectiveness of health cards in supporting consumer health search. Jimmy, Guido Zuccon, Bevan Koopman, Gianluca Demartini |
SIGIR | 4 |
| 2019 | All Those Wasted Hours: On Task Abandonment in CrowdsourcingabstractCrowdsourcing has become a standard methodology to collect manually annotated data such as relevance judgments at scale. On crowdsourcing platforms like Amazon MTurk or FigureEight, crowd workers select tasks to work on based on different dimensions such as task reward and requester reputation. Requesters then receive the judgments of workers who self-selected into the tasks and completed them successfully. Several crowd workers, however, preview tasks, begin working on them, reaching varying stages of task completion without finally submitting their work. Such behavior results in unrewarded effort which remains invisible to requesters. In this paper, we conduct the first investigation into the phenomenon of task abandonment, the act of workers previewing or beginning a task and deciding not to complete it. We follow a three-fold methodology which includes 1) investigating the prevalence and causes of task abandonment by means of a survey over different crowdsourcing platforms, 2) data-driven analyses of logs collected during a large-scale relevance judgment experiment, and 3) controlled experiments measuring the effect of different dimensions on abandonment. Our results show that task abandonment is a widely spread phenomenon. Apart from accounting for a considerable amount of wasted human effort, this bears important implications on the hourly wages of workers as they are not rewarded for tasks that they do not complete. We also show how task abandonment may have strong implications on the use of collected data (for example, on the evaluation of IR systems). Lei Han 0003, Kevin Roitero, Ujwal Gadiraju, Cristina Sarasua, Alessandro Checco, Eddy Maddalena, Gianluca Demartini |
WSDM | 7 |
| 2019 | Scalpel-CD: Leveraging Crowdsourcing and Deep Probabilistic Modeling for Debugging Noisy Training DataabstractThis paper presents Scalpel-CD, a first-of-its-kind system that leverages both human and machine intelligence to debug noisy labels from the training data of machine learning systems. Our system identifies potentially wrong labels using a deep probabilistic model, which is able to infer the latent class of a high-dimensional data instance by exploiting data distributions in the underlying latent feature space. To minimize crowd efforts, it employs a data sampler which selects data instances that would benefit the most from being inspected by the crowd. The manually verified labels are then propagated to similar data instances in the original training data by exploiting the underlying data structure, thus scaling out the contribution from the crowd. Scalpel-CD is designed with a set of algorithmic solutions to automatically search for the optimal configurations for different types of training data, in terms of the underlying data structure, noise ratio, and noise types (random vs. structural). In a real deployment on multiple machine learning tasks, we demonstrate that Scalpel-CD is able to improve label quality by 12.9% with only 2.8% instances inspected by the crowd. Jie Yang 0028, Alisa Smirnova, Dingqi Yang, Gianluca Demartini, Philippe Cudré-Mauroux |
WWW | 4 |
| 2018 | Can User Behaviour Sequences Reflect Perceived Novelty?abstractSerendipity is highly valued as a process for developing original solutions to problems and for innovation. However, it is difficult to capture and thus difficult to measure, but novelty is a key and critical indicator. In this work, we investigate the relationship between user behavioural actions and perceived novelty in the context of browsing. 180 participants completed an open-ended browsing task, while their behaviour actions were tracked. Each seven-action sequence was analysed with respect to the participant's perception of Novelty. Results showed that 6 of the 7 actions map to a sub-sequence that discriminates between high and low novelty. Notably, switching between exploration and immersion, and checking SERPs about the same request in-depth are indicative of highly perceived novelty. The results show that analysing behavioural action sequences leads to better prediction of novelty, and thus the potential for serendipity, than individual browsing actions. Mengdie Zhuang, Elaine Toms, Gianluca Demartini |
CIKM | 3 |
| 2018 | All That Glitters Is Gold - An Attack Scheme on Gold Questions in CrowdsourcingabstractOne of the most popular quality assurance mechanisms in paid micro-task crowdsourcing is based on gold questions: the use of a small set of tasks of which the requester knows the correct answer and, thus, is able to directly assess crowd work quality. In this paper, we show that such mechanism is prone to an attack carried out by a group of colluding crowd workers that is easy to implement and deploy: the inherent size limit of the gold set can be exploited by building an inferential system to detect which parts of the job are more likely to be gold questions. The described attack is robust to various forms of randomisation and programmatic generation of gold questions. We present the architecture of the proposed system, composed of a browser plug-in and an external server used to share information, and briefly introduce its potential evolution to a decentralised implementation. We implement and experimentally validate the gold detection system, using real-world data from a popular crowdsourcing platform. Finally, we discuss the economic and sociological implications of this kind of attack. Alessandro Checco, Jo Bates, Gianluca Demartini |
HCOMP | 3 |
| 2018 | On the Volatility of Commercial Search Engines and its Impact on Information Retrieval ResearchabstractWe studied the volatility of commercial search engines and reflected on its impact on research that uses them as basis of algorithmical techniques or for user studies. Search engine volatility refers to the fact that a query posed to a search engine at two different points in time returns different documents. By comparing search results retrieved every 2 days over a period of 64 days, we found that the considered commercial search engine API consistently presented volatile search results: it both retrieved new documents, and it ranked documents previously retrieved at different ranks throughout time. Moreover, not only results are volatile: we also found that the effectiveness of the search engine in answering a query is volatile. Our findings reaffirmed that results from commercial search engines are volatile and that care should be taken when using these as basis for researching new information retrieval techniques or performing user studies. Jimmy, Guido Zuccon, Gianluca Demartini |
SIGIR | 3 |
| 2018 | Investigating User Perception of Gender Bias in Image Search: The Role of SexismabstractThere is growing evidence that search engines produce results that are socially biased, reinforcing a view of the world that aligns with prevalent social stereotypes. One means to promote greater transparency of search algorithms - which are typically complex and proprietary - is to raise user awareness of biased result sets. However, to date, little is known concerning how users perceive bias in search results, and the degree to which their perceptions differ and/or might be predicted based on user attributes. One particular area of search that has recently gained attention, and forms the focus of this study, is image retrieval and gender bias. We conduct a controlled experiment via crowdsourcing using participants recruited from three countries to measure the extent to which workers perceive a given image results set to be subjective or objective. Demographic information about the workers, along with measures of sexism, are gathered and analysed to investigate whether (gender) biases in the image search results can be detected. Amongst other findings, the results confirm that sexist people are less likely to detect and report gender biases in image search results. Jahna Otterbacher, Alessandro Checco, Gianluca Demartini, Paul D. Clough |
SIGIR | 3 |
| 2018 | On Fine-Grained Relevance ScalesabstractIn Information Retrieval evaluation, the classical approach of adopting binary relevance judgments has been replaced by multi-level relevance judgments and by gain-based metrics leveraging such multi-level judgment scales. Recent work has also proposed and evaluated unbounded relevance scales by means of Magnitude Estimation (ME) and compared them with multi-level scales. While ME brings advantages like the ability for assessors to always judge the next document as having higher or lower relevance than any of the documents they have judged so far, it also comes with some drawbacks. For example, it is not a natural approach for human assessors to judge items as they are used to do on the Web (e.g., 5-star rating). In this work, we propose and experimentally evaluate a bounded and fine-grained relevance scale having many of the advantages and dealing with some of the issues of ME. We collect relevance judgments over a 100-level relevance scale (S100) by means of a large-scale crowdsourcing experiment and compare the results with other relevance scales (binary, 4-level, and ME) showing the benefit of fine-grained scales over both coarse-grained and unbounded scales as well as highlighting some new results on ME. Our results show that S100 maintains the flexibility of unbounded scales like ME in providing assessors with ample choice when judging document relevance (i.e., assessors can fit relevance judgments in between of previously given judgments). It also allows assessors to judge on a more familiar scale (e.g., on 10 levels) and to perform efficiently since the very first judging task. Kevin Roitero, Eddy Maddalena, Gianluca Demartini, Stefano Mizzaro |
SIGIR | 3 |
| 2017 | Understanding Engagement through Search BehaviourabstractEvaluating user engagement with search is a critical aspect of understanding how to assess and improve information retrieval systems. While standard techniques for measuring user engagement use questionnaires, these are obtrusive to user interaction, and can only be collected at acceptable intervals. The problem we address is whether there is a less obtrusive and more automatic way to assess how users perceive the search process and outcome. Log files collect behavioural signals (e.g., clicks, queries) from users on a large scale. In this paper, we investigate the potential to predict how users perceive engagement with search by modelling behavioural signals from log files using supervised learning methods. We focus on different engagement dimensions (Perceived Usability, Felt Involvement, Endurability and Novelty) and examine how 37 behavioural features can inform these dimensions. Our results, obtained from 377 in-lab participants undergoing goal-based search tasks, support the connection between perceived engagement and search behaviour. More specifically, we show that time- and query-related features are best suited for predicting user perceived engagement, and suggest that different behavioural features better reflect specific dimensions. We demonstrate the possibility of predicting user-perceived engagement using search behavioural features. Mengdie Zhuang, Gianluca Demartini, Elaine Toms |
CIKM | 2 |
| 2017 | Let's Agree to Disagree: Fixing Agreement Measures for CrowdsourcingabstractIn the context of micro-task crowdsourcing, each task is usually performed by several workers. This allows researchers to leverage measures of the agreement among workers on the same task, to estimate the reliability of collected data and to better understand answering behaviors of the participants. While many measures of agreement between annotators have been proposed, they are known for suffering from many problems and abnormalities. In this paper, we identify the main limits of the existing agreement measures in the crowdsourcing context, both by means of toy examples as well as with real-world crowdsourcing data, and propose a novel agreement measure based on probabilistic parameter estimation which overcomes such limits. We validate our new agreement measure and show its flexibility as compared to the existing agreement measures. Alessandro Checco, Kevin Roitero, Eddy Maddalena, Stefano Mizzaro, Gianluca Demartini |
HCOMP | 5 |
| 2016 | The Relationship Between User Perception and User Behaviour in Interactive Information Retrieval Evaluation
Mengdie Zhuang, Elaine Toms, Gianluca Demartini |
ECIR | 3 |
| 2016 | Modeling Task Complexity in CrowdsourcingabstractComplexity is crucial to characterize tasks performed by humans through computer systems. Yet, the theory and practice of crowdsourcing currently lacks a clear understanding of task complexity, hindering the design of effective and efficient execution interfaces or fair monetary rewards. To understand how complexity is perceived and distributed over crowdsourcing tasks, we instrumented an experiment where we asked workers to evaluate the complexity of 61 real-world re-instantiated crowdsourcing tasks. We show that task complexity, while being subjective, is coherently perceived across workers; on the other hand, it is significantly influenced by task type. Next, we develop a high-dimensional regression model, to assess the influence of three classes of structural features (metadata, content, and visual) on task complexity, and ultimately use them to measure task complexity. Results show that both the appearance and the language used in task description can accurately predict task complexity. Finally, we apply the same feature set to predict task performance, based on a set of 5 years-worth tasks in Amazon MTurk. Results show that features related to task complexity can improve the quality of task performance prediction, thus demonstrating the utility of complexity as a task modeling property. Jie Yang 0028, Judith Redi, Gianluca Demartini, Alessandro Bozzon |
HCOMP | 3 |
| 2016 | Crowdsourcing Relevance Assessments: The Unexpected Benefits of Limiting the Time to JudgeabstractCrowdsourcing has become an alternative approach to collect relevance judgments at scale thanks to the availability of crowdsourcing platforms and quality control techniques that allow to obtain reliable results. Previous work has used crowdsourcing to ask multiple crowd workers to judge the relevance of a document with respect to a query and studied how to best aggregate multiple judgments of the same topic-document pair. This paper addresses an aspect that has been rather overlooked so far: we study how the time available to express a relevance judgment affects its quality. We also discuss the quality loss of making crowdsourced relevance judgments more efficient in terms of time taken to judge the relevance of a document. We use standard test collections to run a battery of experiments on the crowdsourcing platform CrowdFlower, studying how much time crowd workers need to judge the relevance of a document and at what is the effect of reducing the available time to judge on the overall quality of the judgments. Our extensive experiments compare judgments obtained under different types of time constraints with judgments obtained when no time constraints were put on the task. We measure judgment quality by different metrics of agreement with editorial judgments. Experimental results show that it is possible to reduce the cost of crowdsourced evaluation collection creation by reducing the time available to perform the judgments with no loss in quality. Most importantly, we observed that the introduction of limits on the time available to perform the judgments improves the overall judgment quality. Top judgment quality is obtained with 25-30 seconds to judge a topic-document pair. Eddy Maddalena, Marco Basaldella, Dario De Nart, Dante Degl'Innocenti, Stefano Mizzaro, Gianluca Demartini |
HCOMP | 6 |
| 2016 | Scheduling Human Intelligence Tasks in Multi-Tenant Crowd-Powered SystemsabstractMicro-task crowdsourcing has become a popular approach to effectively tackle complex data management problems such as data linkage, missing values, or schema matching. However, the backend crowdsourced operators of crowd-powered systems typically yield higher latencies than the machine-processable operators, this is mainly due to inherent efficiency differences between humans and machines. This problem can be further exacerbated by the lack of workers on the target crowdsourcing platform, or when the workers are shared unequally among a number of competing requesters; including the concurrent users from the same organization who execute crowdsourced queries with different types, priorities and prices. Under such conditions, a crowd-powered system acts mostly as a proxy to the crowdsourcing platform, and hence it is very difficult to provide effiency guarantees to its end-users. Scheduling is the traditional way of tackling such problems in computer science, by prioritizing access to shared resources. In this paper, we propose a new crowdsourcing system architecture that leverages scheduling algorithms to optimize task execution in a shared resources environment, in this case a crowdsourcing platform. Our study aims at assessing the efficiency of the crowd in settings where multiple types of tasks are run concurrently. We present extensive experimental results comparing i) different multi-tenant crowdsourcing jobs, including a workload derived from real traces, and ii) different scheduling techniques tested with real crowd workers. Our experimental results show that task scheduling can be leveraged to achieve fairness and reduce query latency in multi-tenant crowd-powered systems, although with very different tradeoffs compared to traditional settings not including human factors. Djellel Eddine Difallah, Gianluca Demartini, Philippe Cudré-Mauroux |
WWW | 2 |
| 2016 | Contextualized ranking of entity types based on knowledge graphs
Alberto Tonon, Michele Catasta, Roman Prokofyev, Gianluca Demartini, Karl Aberer, Philippe Cudré-Mauroux |
J. Web Semant. | 4 |
| 2015 | The Dynamics of Micro-Task Crowdsourcing: The Case of Amazon MTurkabstractMicro-task crowdsourcing is rapidly gaining popularity among research communities and businesses as a means to leverage Human Computation in their daily operations. Unlike any other service, a crowdsourcing platform is in fact a marketplace subject to human factors that affect its performance, both in terms of speed and quality. Indeed, such factors shape the dynamics of the crowdsourcing market. For example, a known behavior of such markets is that increasing the reward of a set of tasks would lead to faster results. However, it is still unclear how different dimensions interact with each other: reward, task type, market competition, requester reputation, etc. In this paper, we adopt a data-driven approach to (A) perform a long-term analysis of a popular micro-task crowdsourcing platform and understand the evolution of its main actors (workers, requesters, and platform). (B) We leverage the main findings of our five year log analysis to propose features used in a predictive model aiming at determining the expected performance of any batch at a specific point in time. We show that the number of tasks left in a batch and how recent the batch is are two key features of the prediction. (C) Finally, we conduct an analysis of the demand (new tasks posted by the requesters) and supply (number of tasks completed by the workforce) and show how they affect task prices on the marketplace. Djellel Eddine Difallah, Michele Catasta, Gianluca Demartini, Panagiotis G. Ipeirotis, Philippe Cudré-Mauroux |
WWW | 3 |
| 2015 | Pooling-based continuous evaluation of information retrieval systems
Alberto Tonon, Gianluca Demartini, Philippe Cudré-Mauroux |
Inf. Retr. J. | 2 |
| 2014 | Correct Me If I'm Wrong: Fixing Grammatical Errors by Preposition RankingabstractThe detection and correction of grammatical errors still represent very hard problems for modern error-correction systems. As an example, the top-performing systems at the preposition correction challenge CoNLL-2013 only achieved a F1 score of 17%. In this paper, we propose and extensively evaluate a series of approaches for correcting prepositions, analyzing a large body of high-quality textual content to capture language usage. Leveraging n-gram statistics, association measures, and machine learning techniques, our system is able to learn which words or phrases govern the usage of a specific preposition. Our approach makes heavy use of n-gram statistics generated from very large textual corpora. In particular, one of our key features is the use of n-gram association measures (e.g., Pointwise Mutual Information) between words and prepositions to generate better aggregated preposition rankings for the individual n-grams. We evaluate the effectiveness of our approach using cross-validation with different feature combinations and on two test collections created from a set of English language exams and StackExchange forums. We also compare against state-of-the-art supervised methods. Experimental results from the CoNLL-2013 test collection show that our approach to preposition correction achieves ∼30% in F1 score which results in 13% absolute improvement over the best performing approach at that challenge. Roman Prokofyev, Ruslan Mavlyutov, Martin Grund, Gianluca Demartini, Philippe Cudré-Mauroux |
CIKM | 4 |
| 2014 | Scaling-Up the Crowd: Micro-Task Pricing Schemes for Worker Retention and Latency ImprovementabstractRetaining workers on micro-task crowdsourcing platforms is essential in order to guarantee the timely completion of batches of Human Intelligence Tasks (HITs). Worker retention is also a necessary condition for the introduction of SLAs on crowdsourcing platforms. In this paper, we introduce novel pricing schemes aimed at improving the retention rate of workers working on long batches of similar tasks. We show how increasing or decreasing the monetary reward over time influences the number of tasks a worker is willing to complete in a batch, as well as how it influences the overall latency. We compare our new pricing schemes against traditional pricing methods (e.g., constant reward for all the HITs in a batch) and empirically show how certain schemes effectively function as an incentive for workers to keep working longer on a given batch of HITs. Our experimental results show that the best pricing scheme in terms of worker retention is based on punctual bonuses paid whenever the workers reach predefined milestones. Djellel Eddine Difallah, Michele Catasta, Gianluca Demartini, Philippe Cudré-Mauroux |
HCOMP | 3 |
| 2014 | Effective named entity recognition for idiosyncratic web collectionsabstractNamed Entity Recognition (NER) plays an important role in a variety of online information management tasks including text categorization, document clustering, and faceted search. While recent NER systems can achieve near-human performance on certain documents like news articles, they still remain highly domain-specific and thus cannot effectively identify entities such as original technical concepts in scientific documents. In this work, we propose novel approaches for NER on distinctive document collections (such as scientific articles) based on n-grams inspection and classification. We design and evaluate several entity recognition features---ranging from well-known part-of-speech tags to n-gram co-location statistics and decision trees---to classify candidates. In addition, we show how the use of external knowledge bases (either specific like DBLP or generic like DBPedia) can be leveraged to improve the effectiveness of NER for idiosyncratic collections. We evaluate our system on two test collections created from a set of Computer Science and Physics papers and compare it against state-of-the-art supervised methods. Experimental results show that a careful combination of the features we propose yield up to 85% NER accuracy over scientific collections and substantially outperforms state-of-the-art approaches such as those based on maximum entropy. Roman Prokofyev, Gianluca Demartini, Philippe Cudré-Mauroux |
WWW | 2 |
| 2014 | TransactiveDB: Tapping into Collective Human MemoriesabstractDatabase Management Systems (DBMSs) have been rapidly evolving in the recent years, exploring ways to store multi-structured data or to involve human processes during query execution. In this paper, we outline a future avenue for DBMSs supporting transactive memory queries that can only be answered by a collection of individuals connected through a given interaction graph. We present TransactiveDB and its ecosystem, which allow users to pose queries in order to reconstruct collective human memories. We describe a set of new transactive operators including TUnion, TFill, TJoin, and TProjection. We also describe how TransactiveDB leverages transactive operators---by mixing query execution, social network analysis and human computation---in order to effectively and efficiently tap into the memories of all targeted users. Michele Catasta, Alberto Tonon, Djellel Eddine Difallah, Gianluca Demartini, Karl Aberer, Philippe Cudré-Mauroux |
Proc. VLDB Endow. | 4 |
| 2014 | B-hist: Entity-centric search over personal web browsing history
Michele Catasta, Alberto Tonon, Gianluca Demartini, Jean-Eudes Ranvier, Karl Aberer, Philippe Cudré-Mauroux |
J. Web Semant. | 3 |
| 2013 | Pick-A-Crowd: Tell me what you like, and I'll tell you what to do
Gianluca Demartini |
CIDR | 1 |
| 2013 | CrowdQ: Crowdsourced Query Understanding
Gianluca Demartini, Beth Trushkowsky, Tim Kraska, Michael J. Franklin |
CIDR | 1 |
| 2013 | Ontology-Based Word Sense Disambiguation for Scientific Literature
Roman Prokofyev, Gianluca Demartini, Alexey Boyarsky, Oleg Ruchayskiy, Philippe Cudré-Mauroux |
ECIR | 2 |
| 2013 | TRank: Ranking Entity Types Using the Web of Data
Alberto Tonon, Michele Catasta, Gianluca Demartini, Philippe Cudré-Mauroux, Karl Aberer |
ISWC (1) | 3 |
| 2013 | Pick-a-crowd: tell me what you like, and i'll tell you what to doabstractCrowdsourcing allows to build hybrid online platforms that combine scalable information systems with the power of human intelligence to complete tasks that are difficult to tackle for current algorithms. Examples include hybrid database systems that use the crowd to fill missing values or to sort items according to subjective dimensions such as picture attractiveness. Current approaches to Crowdsourcing adopt a pull methodology where tasks are published on specialized Web platforms where workers can pick their preferred tasks on a first-come-first-served basis. While this approach has many advantages, such as simplicity and short completion times, it does not guarantee that the task is performed by the most suitable worker. In this paper, we propose and extensively evaluate a different Crowdsourcing approach based on a push methodology. Our proposed system carefully selects which workers should perform a given task based on worker profiles extracted from social networks. Workers and tasks are automatically matched using an underlying categorization structure that exploits entities extracted from the task descriptions on one hand, and categories liked by the user on social platforms on the other hand. We experimentally evaluate our approach on tasks of varying complexity and show that our push methodology consistently yield better results than usual pull strategies. Djellel Eddine Difallah, Gianluca Demartini, Philippe Cudré-Mauroux |
WWW | 2 |
| 2013 | Large-scale linked data integration using probabilistic reasoning and crowdsourcing
Gianluca Demartini, Djellel Eddine Difallah, Philippe Cudré-Mauroux |
VLDB J. | 1 |
| 2012 | Predicting the Future Impact of News Events
Julien Gaugaz, Patrick Siehndel, Gianluca Demartini, Tereza Iofciu, Mihai Georgescu, Nicola Henze |
ECIR | 3 |
| 2012 | Tag Recommendation for Large-Scale Ontology-Based Information Systems
Roman Prokofyev, Alexey Boyarsky, Oleg Ruchayskiy, Karl Aberer, Gianluca Demartini, Philippe Cudré-Mauroux |
ISWC (2) | 5 |
| 2012 | Combining inverted indices and structured search for ad-hoc object retrievalabstractRetrieving semi-structured entities to answer keyword queries is an increasingly important feature of many modern Web applications. The fast-growing Linked Open Data (LOD) movement makes it possible to crawl and index very large amounts of structured data describing hundreds of millions of entities. However, entity retrieval approaches have yet to find efficient and effective ways of ranking and navigating through those large data sets. In this paper, we address the problem of Ad-hoc Object Retrieval over large-scale LOD data by proposing a hybrid approach that combines IR and structured search techniques. Specifically, we propose an architecture that exploits an inverted index to answer keyword queries as well as a semi-structured database to improve the search effectiveness by automatically generating queries over the LOD graph. Experimental results show that our ranking algorithms exploiting both IR and graph indices outperform state-of-the-art entity retrieval techniques by up to 25% over the BM25 baseline. Alberto Tonon, Gianluca Demartini, Philippe Cudré-Mauroux |
SIGIR | 2 |
| 2012 | ZenCrowd: leveraging probabilistic reasoning and crowdsourcing techniques for large-scale entity linkingabstractWe tackle the problem of entity linking for large collections of online pages; Our system, ZenCrowd, identifies entities from natural language text using state of the art techniques and automatically connects them to the Linked Open Data cloud. We show how one can take advantage of human intelligence to improve the quality of the links by dynamically generating micro-tasks on an online crowdsourcing platform. We develop a probabilistic framework to make sensible decisions about candidate links and to identify unreliable human workers. We evaluate ZenCrowd in a real deployment and show how a combination of both probabilistic reasoning and crowdsourcing techniques can significantly improve the quality of the links, while limiting the amount of work performed by the crowd. Gianluca Demartini, Djellel Eddine Difallah, Philippe Cudré-Mauroux |
WWW | 1 |
| 2011 | ARES: A Retrieval Engine Based on Sentiments - Sentiment-Based Search Result Annotation and Diversification
Gianluca Demartini |
ECIR | 1 |
| 2011 | ReFER: Effective Relevance Feedback for Entity Ranking
Tereza Iofciu, Gianluca Demartini, Nick Craswell, Arjen P. de Vries |
ECIR | 2 |
| 2011 | Analyzing Political Trends in the Blogosphere
Gianluca Demartini, Stefan Siersdorfer, Sergiu Chelaru, Wolfgang Nejdl |
ICWSM | 1 |
| 2010 | TAER: time-aware entity retrieval-exploiting the past to find relevant entities in news articlesabstractRetrieving entities instead of just documents has become an important task for search engines. In this paper we study entity retrieval for news applications, and in particular the importance of the news trail history (i.e., past related articles) in determining the relevant entities in current articles. This is an important problem in applications that display retrieved entities to the user, together with the news article. Gianluca Demartini, Malik Muhammad Saad Missen, Roi Blanco, Hugo Zaragoza |
CIKM | 1 |
| 2010 | The missing links: discovering hidden same-as links among a billion of triplesabstractThe Semantic Web is constantly gaining momentum, as more and more Web sites and content providers adopt its principles. At the core of these principles lies the Linked Data movement, which demands that data on the Web shall be annotated and linked among different sources, instead of being isolated in data silos. In order to materialize this vision of a web of semantics, existing resource identifiers should be reused and shared between different Web sites. This is not always the case with the current state of the Semantic Web, since multiple identifiers are, more often than not, redundantly introduced for the same resources. George Papadakis 0001, Gianluca Demartini, Peter Fankhauser, Philipp Kärger |
iiWAS | 2 |
| 2010 | Exploiting click-through data for entity retrievalabstractWe present an approach for answering Entity Retrieval queries using click-through information in query log data from a commercial Web search engine. We compare results using click graphs and session graphs and present an evaluation test set making use of Wikipedia "List of" pages. Bodo Billerbeck, Gianluca Demartini, Claudiu S. Firan, Tereza Iofciu, Ralf Krestel |
SIGIR | 2 |
| 2010 | Entity summarization of news articlesabstractinc.com In this paper we study the problem of entity retrieval for news applications and the importance of the news trail his-tory (i.e. past related articles) to determine the relevant entities in current articles. We construct a novel entity-labeled corpus with temporal information out of the TREC 2004 Novelty collection. We develop and evaluate several features, and show that an article’s history can be exploited to improve its summarization. Gianluca Demartini, Malik Muhammad Saad Missen, Roi Blanco, Hugo Zaragoza |
SIGIR | 1 |
| 2010 | Why finding entities in Wikipedia is difficult, sometimes
Gianluca Demartini, Claudiu S. Firan, Tereza Iofciu, Ralf Krestel, Wolfgang Nejdl |
Inf. Retr. | 1 |
| 2010 | Leveraging personal metadata for Desktop search: The Beagle++ system
Enrico Minack, Raluca Paiu, Stefania Costache 0001, Gianluca Demartini, Julien Gaugaz, Ekaterini Ioannou, Paul-Alexandru Chirita, Wolfgang Nejdl |
J. Web Semant. | 4 |
| 2009 | A Vector Space Model for Ranking Entities and Its Application to Expert Search
Gianluca Demartini, Julien Gaugaz, Wolfgang Nejdl |
ECIR | 1 |
| 2009 | How to Trace and Revise Identities
Julien Gaugaz, Jakub Zakrzewski 0002, Gianluca Demartini, Wolfgang Nejdl |
ESWC | 3 |
| 2008 | Ranking Categories for Web Search
Gianluca Demartini, Paul-Alexandru Chirita, Ingo Brunkhorst, Wolfgang Nejdl |
ECIR | 1 |
| 2008 | Social recommendations of content and metadataabstractIn this paper we present metadata based recommendation algorithms addressing two scenarios within social desktop communities: a) recommendation of resources from the co-worker's desktop, and b) recommendation of metadata for enriching the own annotation layer. Together with the algorithms we present first evaluation results as well as empirical evaluations showing that metadata based recommendations can be used in such distributed social desktop communities. Rodolfo Stecher, Gianluca Demartini, Claudia Niederée |
iiWAS | 2 |
| 2008 | Semantically Enhanced Entity Ranking
Gianluca Demartini, Claudiu S. Firan, Tereza Iofciu, Wolfgang Nejdl |
WISE | 1 |
| 2007 | Building a Desktop Search Test-Bed
Sergey Chernov 0001, Pavel Serdyukov, Paul-Alexandru Chirita, Gianluca Demartini, Wolfgang Nejdl |
ECIR | 4 |
| 2006 | A Classification of IR Effectiveness Metrics
Gianluca Demartini, Stefano Mizzaro |
ECIR | 1 |
| 2006 | Experiments on Average Distance Measure
Vincenzo Della Mea, Gianluca Demartini, Luca Di Gaspero, Stefano Mizzaro |
ECIR | 2 |