Ujwal Gadiraju

dblp:130/0365 · also Ujwal Kumar Gadiraju · DBLP profile ↗
← Back
36ranked-venue papers in the field
2as first author
18since 2021 · last 2026
0000-0002-6189-6539ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 22 (2 first)Data Mining & Knowledge Discovery · 11Database Systems & Data Management · 2Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2026 From SERPs to Sound: How Search Engine Result Pages and AI-generated Podcasts Interact to Influence User Attitudes on Controversial Topics
abstract
Compared to search engine result pages (SERPs), AI-generated podcasts represent a relatively new and relatively more passive modality of information consumption, delivering narratives in a naturally engaging format. As these two media increasingly converge in everyday information-seeking behavior, it is essential to explore how their interaction influences user attitudes, particularly in contexts involving controversial, value-laden, and often debated topics. Addressing this need, we aim to understand how information mediums of present-day SERPs and AI-generated podcasts interact to shape the opinions of users. To this end, through a controlled user study (N = 483), we investigated user attitudinal effects of consuming information via SERPs and AI-generated podcasts, focusing on how the sequence and modality of exposure shape user opinions. A majority of users in our study corresponded to attitude change outcomes, and we found an effect of sequence on attitude change. Our results further revealed a role of viewpoint bias and the degree of topic controversiality in shaping attitude change, although we found no effect of individual moderators.
Gaole He, Alisa Rieger, Ujwal Gadiraju
CHIIR4
2024 Responsible Opinion Formation on Debated Topics in Web Search
Alisa Rieger, Tim Draws, Nicolas Mattis, David Maxwell 0001, David Elsweiler, Ujwal Gadiraju, Dana McKay, Alessandro Bozzon, Maria Soledad Pera
ECIR (4)6
2023 In a Hurry: How Time Constraints and the Presentation of Web Search Results Affect User Behaviour and Experience
Garrett Allen, Mike Beijen, David Maxwell 0001, Ujwal Gadiraju
ICWE4
2022 Goal-Setting Behavior of Workers on Crowdsourcing Platforms: An Exploratory Study on MTurk and Prolific
abstract
A wealth of evidence across several domains indicates that goal setting improves performance and learning by enabling individuals to commit their thoughts and actions to goal achievement. Recently, researchers have begun studying the effects of goal setting in paid crowdsourcing to improve the quality and quantity of contributions, increase learning gains, and hold participants accountable for contributing more effectively. However, there is a lack of research addressing crowd workers' goal-setting practices, how they are currently pursuing them, and the challenges that they face. This information is essential for researchers and developers to create tools that assist crowd workers in pursuing their goals more effectively, thereby improving the quality of their contributions. This paper addresses these gaps by conducting mixed-method research in which we surveyed 205 workers from two crowdsourcing platforms -- Amazon Mechanical Turk (MTurk) and Prolific -- about their goal-setting practices. Through a 14-item survey, we asked workers regarding the types of goals they create, their goal achievement strategies, potential barriers that impede goal attainment, and their use of software tools for effective goal management. We discovered that (a) workers actively create intrinsic and extrinsic goals; (b) use a combination of tools for goal management; (c) medical issues and a busy lifestyle are some obstacles to their goal achievement; and (d) we gathered novel features for future goal management tools. Our findings shed light on the broader implications of developing goal management tools to improve workers' well-being.
Tahir Abbas 0001, Ujwal Gadiraju
HCOMP2
2022 Gesticulate for Health's Sake! Understanding the Use of Gestures as an Input Modality for Microtask Crowdsourcing
abstract
Human input is pivotal in building reliable and robust artificial intelligence systems. By providing a means to gather diverse, high-quality, representative, and cost-effective human input on demand, microtask crowdsourcing marketplaces have thrived. Despite the unmistakable benefits available from online crowd work, the lack of health provisions and safeguards, along with existing work practices threatens the sustainability of this paradigm. Prior work has investigated worker engagement and mental health, yet no such investigations into the effects of crowd work on the physical health of workers have been undertaken. Crowd workers complete their work in various sub-optimal work environments, often using a conventional input modality of a mouse and keyboard. The repetitive nature of microtask crowdsourcing can lead to stress-related injuries, such as the well-documented carpal tunnel syndrome. It is known that stretching exercises can help reduce injuries and discomfort in office workers. Gestures, the act of using the body intentionally to affect the behavior of an intelligent system, can serve as both stretches and an alternative form of input for microtasks. To better understand the usefulness of the dual-purpose input modality of ergonomically-informed gestures across different crowdsourced microtasks, we carried out a controlled 2 x 3 between-subjects study (N=294). Considering the potential benefits of gestures as an input modality, our results suggest a real trade-off between worker accuracy in exchange for potential short to long-term health benefits.
Garrett Allen, Andrea Hu, Ujwal Gadiraju
HCOMP3
2022 It Is like Finding a Polar Bear in the Savannah! Concept-Level AI Explanations with Analogical Inference from Commonsense Knowledge
abstract
With recent advances in explainable artificial intelligence (XAI), researchers have started to pay attention to concept-level explanations, which explain model predictions with a high level of abstraction. However, such explanations may be difficult to digest for laypeople due to the potential knowledge gap and the concomitant cognitive load. Inspired by recent work, we argue that analogy-based explanations composed of commonsense knowledge may be a potential solution to tackle this issue. In this paper, we propose analogical inference as a bridge to help end-users leverage their commonsense knowledge to better understand the concept-level explanations. Specifically, we design an effective analogy-based explanation generation method and collect 600 analogy-based explanations from 100 crowd workers. Furthermore, we propose a set of structured dimensions for the qualitative assessment of analogy-based explanations and conduct an empirical evaluation of the generated analogies with experts. Our findings reveal significant positive correlations between the qualitative dimensions of analogies and the perceived helpfulness of analogy-based explanations. These insights can inform the design of future methods for the generation of effective analogy-based explanations. We also find that the understanding of commonsense explanations varies with the experience of the recipient user, which points out the need for further work on personalization when leveraging commonsense explanations.
Gaole He, Agathe Balayn, Stefan Buijsman, Jie Yang 0028, Ujwal Gadiraju
HCOMP5
2022 SignUpCrowd: Using Sign-Language as an Input Modality for Microtask Crowdsourcing
abstract
Different input modalities have been proposed and employed in technological landscapes like microtask crowdsourcing. However, sign language remains an input modality that has received little attention. Despite the fact that thousands of people around the world primarily use sign language, very little has been done to include them in such technological landscapes. We aim to address this gap and take a step towards the inclusion of deaf and mute people in microtask crowdsourcing. We first identify various microtasks which can be adapted to use sign language as input, while elucidating the challenges it introduces. We built a system called ‘SignUpCrowd’ that can be used to support sign language input for microtask crowdsourcing. We carried out a between-subjects study (N=240) to understand the effectiveness of sign language as an input modality for microtask crowdsourcing in comparison to prevalent textual and click input modalities. We explored this through the lens of visual question answering and sentiment analysis tasks by recruiting workers from the Prolific crowdsourcing platform. Our results indicate that sign language as an input modality in microtask crowdsourcing is comparable to the prevalent standards of using text and click input. Although people with no knowledge of sign language found it difficult to use, this input modality has the potential to broaden participation in crowd work. We highlight evidence suggesting the scope for sign language as a viable input type for microtask crowdsourcing. Our findings pave the way for further research to introduce sign language in real-world applications and create an inclusive technological landscape that more people can benefit from.
Aayush Singh, Sebastian Wehkamp, Ujwal Gadiraju
HCOMP3
2022 Ready Player One! Eliciting Diverse Knowledge Using A Configurable Game
abstract
Access to commonsense knowledge is receiving renewed interest for developing neuro-symbolic AI systems, or debugging deep learning models. Little is currently understood about the types of knowledge that can be gathered using existing knowledge elicitation methods. Moreover, these methods fall short of meeting the evolving requirements of several downstream AI tasks. To this end, collecting broad and tacit knowledge, in addition to negative or discriminative knowledge can be highly useful. Addressing this research gap, we developed a novel game with a purpose, ‘FindItOut’, to elicit different types of knowledge from human players through easily configurable game mechanics. We recruited 125 players from a crowdsourcing platform, who played 2430 rounds, resulting in the creation of more than 150k tuples of knowledge. Through an extensive evaluation of these tuples, we show that FindItOut can successfully result in the creation of plural knowledge with a good player experience. We evaluate the efficiency of the game (over 10 × higher than a reference baseline) and the usefulness of the resulting knowledge, through the lens of two downstream tasks — commonsense question answering and the identification of discriminative attributes. Finally, we present a rigorous qualitative analysis of the tuples’ characteristics, that informs the future use of FindItOut across various researcher and practitioner communities.
Agathe Balayn, Gaole He, Andrea Hu, Jie Yang 0028, Ujwal Gadiraju
WWW5
2022 To Trust or Not To Trust: How a Conversational Interface Affects Trust in a Decision Support System
abstract
Trust is an important component of human-AI relationships and plays a major role in shaping the reliance of users on online algorithmic decision support systems. With recent advances in natural language processing, text and voice-based conversational interfaces have provided users with new ways of interacting with such systems. Despite the growing applications of conversational user interfaces (CUIs), little is currently understood about the suitability of such interfaces for decision support and how CUIs inspire trust among humans engaging with decision support systems. In this work, we aim to address this gap and answer the following question: to what extent can a conversational interface build user trust in decision support systems in comparison to a conventional graphical user interface? To this end, we built a text-based conversational interface, and a conventional web-based graphical user interface. These served as the means for users to interact with an online decision support system to help them find housing, given a fixed set of constraints. To understand how the accuracy of the decision support system moderates user behavior and trust across the two interfaces, we considered an accurate and inaccurate system. We carried out a 2 × 2 between-subjects study (N = 240) on the Prolific crowdsourcing platform. Our findings show that the conversational interface was significantly more effective in building user trust and satisfaction in the online housing recommendation system when compared to the conventional web interface. Our results highlight the potential impact of conversational interfaces for trust development in decision support systems.
Akshit Gupta, Debadeep Basu, Ramya Ghantasala, Sihang Qiu, Ujwal Gadiraju
WWW5
2022 What Should You Know? A Human-In-the-Loop Approach to Unknown Unknowns Characterization in Image Recognition
abstract
Unknown unknowns represent a major challenge in reliable image recognition. Existing methods mainly focus on unknown unknowns identification, leveraging human intelligence to gather images that are potentially difficult for the machine. To drive a deeper understanding of unknown unknowns and more effective identification and treatment, this paper focuses on unknown unknowns characterization. We introduce a human-in-the-loop, semantic analysis framework for characterizing unknown unknowns at scale. We engage humans in two tasks that specify what a machine should know and describe what it really knows, respectively, both at the conceptual level, supported by information extraction and machine learning interpretability methods. Data partitioning and sampling techniques are employed to scale out human contributions in handling large data. Through extensive experimentation on scene recognition tasks, we show that our approach provides a rich, descriptive characterization of unknown unknowns and allows for more effective and cost-efficient detection than the state of the art.
Shahin Sharifi Noorian, Sihang Qiu, Ujwal Gadiraju, Jie Yang 0028, Alessandro Bozzon
WWW3
2021 Note the Highlight: Incorporating Active Reading Tools in a Search as Learning Environment
abstract
Active reading strategies---such as content annotations (through the use of highlighting and note-taking, for example)---have been shown to yield improvements to a learner's knowledge and understanding of the topic being explored. This has been especially notable in long and complex learning endeavours. With web search engines nowadays used as the primary gateway for learners (or users) to find content that helps them realise their learning goals, they are often poorly equipped with the necessary tools to aid in sense-making, an important aspect of theSearch as Learning (SAL) process. Within theInformation Retrieval (IR) community, research efforts have explored ways to keep track of users' search context by providing a notepad-like interface for the collection of relevant articles, and aid them during the exploratory search process. However, these studies did not explicitly measure the effect that such tools have on knowledge and understanding during a complex, learning-oriented search task. In this paper, we address this research gap by carrying out an Interactive IR experiment with highlighting and note-taking tools built into the search interface. We conducted a crowdsourced between-subjects study (N=115), where participants were assigned to one of four conditions: (i) control (a standard web search interface); (ii) high (highlighting enabled);(iii) note (note-taking enabled); and (iv) highnote (both highlighting and note-taking enabled). We assess participants' learning with a recall-oriented vocabulary learning task, and a cognitively more taxing essay writing task. We find that(i) active reading tools do not aid in the vocabulary learning task. However,(ii) participants in high covered 34% more subtopics, and participants in note covered 34% more facts in their essays when compared to control. Furthermore, (iii) we observed that incorporating active learning tools significantly changed the search behaviour of participants across a number of measures. This is the first work that sheds light on the effect of active reading tools on the SAL process, with important design implications for learning-oriented search systems.
Nirmal Roy, Manuel Valle Torre, Ujwal Gadiraju, David Maxwell 0001, Claudia Hauff
CHIIR3
2021 How Do Active Reading Strategies Affect Learning Outcomes in Web Search?
Nirmal Roy, Manuel Valle Torre, Ujwal Gadiraju, David Maxwell 0001, Claudia Hauff
ECIR (2)3
2021 Making Time Fly: Using Fillers to Improve Perceived Latency in Crowd-Powered Conversational Systems
abstract
Crowd-Powered Conversational Systems (CPCS) are gaining traction due to their potential utility in a range of application fields where automated conversational interfaces are still inadequate. Currently, long response times negatively impact CPCSs, limiting their potential application as conversational partners. Related research has focused on developing algorithms for swiftly hiring workers and synchronous crowd coordination techniques to ensure high-quality work. Evaluation studies typically concern system reaction times and performance measurements, but have so far not examined the effects of extended wait times on users. The goal of this study, based on time perception models, is to explore how effective different time fillers are at reducing the negative impacts of waiting in CPCSs. To this end, we conducted a rigorous simulation-based between-subjects (N = 930) study on the Prolific crowdsourcing platform to assess the influence of different filler types across three levels of delay (8, 16 & 32s) for Information Retrieval (IR) and stress management tasks. Our results show that asking users to perform secondary tasks (e.g., microtasks or breathing exercises) while waiting for longer periods of time helped divert their attention away from timekeeping, increased their engagement, and resulted in shorter perceived waiting times. For shorter delays, conversational fillers generated more intense immersion and contributed to shorten the perception of time.
Tahir Abbas 0001, Ujwal Gadiraju, Vassilis-Javed Khan, Panos Markopoulos 0001
HCOMP2
2021 A Checklist to Combat Cognitive Biases in Crowdsourcing
abstract
Recent research has demonstrated that cognitive biases such as the confirmation bias or the anchoring effect can negatively affect the quality of crowdsourced data. In practice, however, such biases go unnoticed unless specifically assessed or controlled for. Task requesters need to ensure that task workflow and design choices do not trigger workers’ cognitive biases. Moreover, to facilitate the reuse of crowdsourced data collections, practitioners can benefit from understanding whether and which cognitive biases may be associated with the data. To this end, we propose a 12-item checklist adapted from business psychology to combat cognitive biases in crowdsourcing. We demonstrate the practical application of this checklist in a case study on viewpoint annotations for search results. Through a retrospective analysis of relevant crowdsourcing research that has been published at HCOMP in 2018, 2019, and 2020, we show that cognitive biases may often affect crowd workers but are typically not considered as potential sources of poor data quality. The checklist we propose is a practical tool that requesters can use to improve their task designs and appropriately describe potential limitations of collected data. It contributes to a body of efforts towards making human-labeled data more reliable and reusable.
Tim Draws, Alisa Rieger, Oana Inel, Ujwal Gadiraju, Nava Tintarev
HCOMP4
2021 Exploring the Music Perception Skills of Crowd Workers
abstract
Music content annotation campaigns are common on paid crowdsourcing platforms. Crowd workers are expected to annotate complicated music artefacts, which can demand certain skills and expertise. Traditional methods of participant selection are not designed to capture these kind of domain-specific skills and expertise, and often domain-specific questions fall under the general demographics category. Despite the popularity of such tasks, there is a general lack of deeper understanding of the distribution of musical properties - especially auditory perception skills - among workers. To address this knowledge gap, we conducted a user study (N=100) on Prolific. We asked workers to indicate their musical sophistication through a questionnaire and assessed their music perception skills through an audio-based skill test. The goal of this work is to better understand the extent to which crowd workers possess higher perceptions skills, beyond their own musical education level and self reported abilities. Our study shows that untrained crowd workers can possess high perception skills on the music elements of melody, tuning, accent and tempo; skills that can be useful in a plethora of annotation tasks in the music domain.
Ioannis Petros Samiotis, Sihang Qiu, Christoph Lofi, Jie Yang 0028, Ujwal Gadiraju, Alessandro Bozzon
HCOMP5
2021 This Is Not What We Ordered: Exploring Why Biased Search Result Rankings Affect User Attitudes on Debated Topics
abstract
In web search on debated topics, algorithmic and cognitive biases strongly influence how users consume and process information. Recent research has shown that this can lead to a search engine manipulation effect (SEME): when search result rankings are biased towards a particular viewpoint, users tend to adopt this favored viewpoint. To better understand the mechanisms underlying SEME, we present a pre-registered, 5 x 3 factorial user study investigating whether order effects (i.e., users adopting the viewpoint pertaining to higher-ranked documents) can cause SEME. For five different debated topics, we evaluated attitude change after exposing participants with mild pre-existing attitudes to search results that were overall viewpoint-balanced but reflected one of three levels of algorithmic ranking bias. We found that attitude change did not differ across levels of ranking bias and did not vary based on individual user differences. Our results thus suggest that order effects may not be an underlying mechanism of SEME. Exploratory analyses lend support to the presence of exposure effects (i.e., users adopting the majority viewpoint among the results they examine) as a contributing factor to users' attitude change. We discuss how our findings can inform the design of user bias mitigation strategies.
Tim Draws, Nava Tintarev, Ujwal Gadiraju, Alessandro Bozzon, Benjamin Timmermans
SIGIR3
2021 Topic-independent modeling of user knowledge in informational search sessions
abstract
Abstract Web search is among the most frequent online activities. In this context, widespread informational queries entail user intentions to obtain knowledge with respect to a particular topic or domain. To serve learning needs better, recent research in the field of interactive information retrieval has advocated the importance of moving beyond relevance ranking of search results and considering a user’s knowledge state within learning oriented search sessions. Prior work has investigated the use of supervised models to predict a user’s knowledge gain and knowledge state from user interactions during a search session. However, the characteristics of the resources that a user interacts with have neither been sufficiently explored, nor exploited in this task. In this work, we introduce a novel set of resource-centric features and demonstrate their capacity to significantly improve supervised models for the task of predicting knowledge gain and knowledge state of users in Web search sessions. We make important contributions, given that reliable training data for such tasks is sparse and costly to obtain. We introduce various feature selection strategies geared towards selecting a limited subset of effective and generalizable features.
Ran Yu 0001, Markus Rokicki, Ujwal Gadiraju, Stefan Dietze
Inf. Retr. J.4
2021 The Impact of Task Abandonment in Crowdsourcing
abstract
Crowdsourcing has become a standard methodology to collect manually annotated data such as relevance judgments at scale. On crowdsourcing platforms like Amazon MTurk or FigureEight, crowd workers select tasks to work on based on different dimensions such as task reward and requester reputation. Requesters then receive the judgments of workers who self-selected into the tasks and completed them successfully. Several crowd workers, however, preview tasks, begin working on them, reaching varying stages of task completion without finally submitting their work. Such behavior results in unrewarded effort which remains invisible to requesters. In this paper, we conduct an investigation of the phenomenon of task abandonment, the act of workers previewing or beginning a task and deciding not to complete it. We follow a three-fold methodology which includes 1) investigating the prevalence and causes of task abandonment by means of a survey over different crowdsourcing platforms, 2) data-driven analysis of logs collected during a large-scale relevance judgment experiment, and 3) controlled experiments measuring the effect of different dimensions on abandonment. Our results show that task abandonment is a widely spread phenomenon. Apart from accounting for a considerable amount of wasted human effort, this bears important implications on the hourly wages of workers as they are not rewarded for tasks that they do not complete. We also show how task abandonment may have strong implications on the use of collected data (for example, on the evaluation of Information Retrieval systems).
Lei Han 0003, Kevin Roitero, Ujwal Gadiraju, Cristina Sarasua, Alessandro Checco, Eddy Maddalena, Gianluca Demartini
IEEE Trans. Knowl. Data Eng.3
2020 Trainbot: A Conversational Interface to Train Crowd Workers for Delivering On-Demand Therapy
abstract
On-demand emotional support is an expensive and elusive societal need that is exacerbated in difficult times — as witnessed during the COVID-19 pandemic. Prior work in affective crowdsourcing has examined ways to overcome technical challenges for providing on-demand emotional support to end users. This can be achieved by training crowd workers to provide thoughtful and engaging on-demand emotional support. Inspired by recent advances in conversational user interface research, we investigate the efficacy of a conversational user interface for training workers to deliver psychological support to users in need. To this end, we conducted a between-subjects experimental study on Prolific, wherein a group of workers (N=200) received training on motivational interviewing via either a conversational interface or a conventional web interface. Our results indicate that training workers in a conversational interface yields both better worker performance and improves their user experience in on-demand stress management tasks.
Tahir Abbas 0001, Vassilis-Javed Khan, Ujwal Gadiraju, Panos Markopoulos 0001
HCOMP3
2020 Impact of Algorithmic Decision Making on Human Behavior: Evidence from Ultimatum Bargaining
abstract
Recent advances in machine learning have led to the widespread adoption of ML models for decision support systems. However, little is known about how the introduction of such systems affects the behavior of human stakeholders. This pertains both to the people using the system, as well as those who are affected by its decisions. To address this knowledge gap, we present a series of ultimatum bargaining game experiments comprising 1178 participants. We find that users are willing to use a black-box decision support system and thereby make better decisions. This translates into higher levels of cooperation and better market outcomes. However, because users under-weigh algorithmic advice, market outcomes remain far from optimal. Explanations increase the number of unique system inquiries, but users appear less willing to follow the system’s recommendation. People who negotiate with a user who has a decision support system, but cannot use one themselves, react to its introduction by demanding a better deal for themselves, thereby decreasing overall cooperation levels. This effect is largely driven by the percentage of participants who perceive the system’s availability as unfair. Interpretability mitigates perceptions of unfairness. Our findings highlight the potential for decision support systems to further human cooperation, but also the need for regulators to consider heterogeneous stakeholder reactions. In particular, higher levels of transparency might inadvertently hurt cooperation through changes in fairness perceptions.
Alexander Erlei, Franck Awounang Nekdem, Lukas Meub, Avishek Anand, Ujwal Gadiraju
HCOMP5
2020 Just the Right Mood for HIT! - Analyzing the Role of Worker Moods in Conversational Microtask Crowdsourcing
Sihang Qiu, Ujwal Gadiraju, Alessandro Bozzon
ICWE2
2020 TaskGenie: Crowd-Powered Task Generation for Struggling Search
Luyan Xu, Xuan Zhou 0001, Ujwal Gadiraju
WISE (2)3
2020 Crowd Worker Strategies in Relevance Judgment Tasks
abstract
Crowdsourcing is a popular technique to collect large amounts of human-generated labels, such as relevance judgments used to create information retrieval (IR) evaluation collections. Previous research has shown how collecting high quality labels from a crowdsourcing platform can be challenging. Existing quality assurance techniques focus on answer aggregation or on the use of gold questions where ground-truth data allows to check for the quality of the responses.
Lei Han 0003, Eddy Maddalena, Alessandro Checco, Cristina Sarasua, Ujwal Gadiraju, Kevin Roitero, Gianluca Demartini
WSDM5
2019 Revealing the Role of User Moods in Struggling Search Tasks
abstract
User-centered approaches have been extensively studied and used in the area of struggling search. Related research has targeted key aspects of users such as user satisfaction or frustration, and search success or failure, using a variety of experimental methods including laboratory user studies, in-situ explicit feedback from searchers and by using crowdsourcing. Such studies are valuable in advancing the understanding of search difficulty from a user's perspective, and yield insights that can directly improve search systems and their evaluation. However, little is known about how user moods influence their interactions with a search system or their perception of struggling. In this work, we show that a user's own mood. can systematically bias the user's perception, and experience while interacting with a search system and trying to satisfy an information need. People who are in activated-(un)pleasant moods tend to issue more queries than people in deactivated or neutral moods. Those in an unpleasant mood perceive a higher level of difficulty. Our insights extend the current understanding of struggling search tasks and have important implications on the design and evaluation of search systems supporting such tasks.
Luyan Xu, Xuan Zhou 0001, Ujwal Gadiraju
SIGIR3
2019 All Those Wasted Hours: On Task Abandonment in Crowdsourcing
abstract
Crowdsourcing has become a standard methodology to collect manually annotated data such as relevance judgments at scale. On crowdsourcing platforms like Amazon MTurk or FigureEight, crowd workers select tasks to work on based on different dimensions such as task reward and requester reputation. Requesters then receive the judgments of workers who self-selected into the tasks and completed them successfully. Several crowd workers, however, preview tasks, begin working on them, reaching varying stages of task completion without finally submitting their work. Such behavior results in unrewarded effort which remains invisible to requesters. In this paper, we conduct the first investigation into the phenomenon of task abandonment, the act of workers previewing or beginning a task and deciding not to complete it. We follow a three-fold methodology which includes 1) investigating the prevalence and causes of task abandonment by means of a survey over different crowdsourcing platforms, 2) data-driven analyses of logs collected during a large-scale relevance judgment experiment, and 3) controlled experiments measuring the effect of different dimensions on abandonment. Our results show that task abandonment is a widely spread phenomenon. Apart from accounting for a considerable amount of wasted human effort, this bears important implications on the hourly wages of workers as they are not rewarded for tasks that they do not complete. We also show how task abandonment may have strong implications on the use of collected data (for example, on the evaluation of IR systems).
Lei Han 0003, Kevin Roitero, Ujwal Gadiraju, Cristina Sarasua, Alessandro Checco, Eddy Maddalena, Gianluca Demartini
WSDM3
2018 Analyzing Knowledge Gain of Users in Informational Search Sessions on the Web
abstract
Web search is frequently used by people to acquire new knowledge and to satisfy learning-related objectives, but little is known about how a user»s knowledge evolves through the course of a search session. We present a study addressing the knowledge gain of users in informational search sessions. Using crowdsourcing, we recruited 500 distinct users and orchestrated real-world search sessions spanning 10 different topics and information needs. By using scientifically formulated knowledge tests we calibrated the knowledge of users before and after their search sessions, quantifying their knowledge gain. We investigated the impact of information needs on the search behavior and knowledge gain of users, revealing a significant effect of information need on user queries and navigational patterns, but no direct effect on the knowledge gain. Users on average exhibited a higher knowledge gain through search sessions pertaining to topics they were less familiar with.
Ujwal Gadiraju, Ran Yu 0001, Stefan Dietze, Peter Holtz
CHIIR1
2018 Predicting User Knowledge Gain in Informational Search Sessions
abstract
Web search is frequently used by people to acquire new knowledge and to satisfy learning-related objectives. In this context, informational search missions with an intention to obtain knowledge pertaining to a topic are prominent. The importance of learning as an outcome of web search has been recognized. Yet, there is a lack of understanding of the impact of web search on a user's knowledge state. Predicting the knowledge gain of users can be an important step forward if web search engines that are currently optimized for relevance can be molded to serve learning outcomes. In this paper, we introduce a supervised model to predict a user's knowledge state and knowledge gain from features captured during the search sessions. To measure and predict the knowledge gain of users in informational search sessions, we recruited 468 distinct users using crowdsourcing and orchestrated real-world search sessions spanning 11 different topics and information needs. By using scientifically formulated knowledge tests, we calibrated the knowledge of users before and after their search sessions, quantifying their knowledge gain. Our supervised models utilise and derive a comprehensive set of features from the current state of the art and compare performance of a range of feature sets and feature selection strategies. Through our results, we demonstrate the ability to predict and classify the knowledge state and gain using features obtained during search sessions, exhibiting superior performance to an existing baseline in the knowledge state prediction task.
Ran Yu 0001, Ujwal Gadiraju, Peter Holtz, Markus Rokicki, Philipp Kemkes, Stefan Dietze
SIGIR2
2017 JustEvents: A Crowdsourced Corpus for Event Validation with Strict Temporal Constraints
Andrea Ceroni, Ujwal Gadiraju, Marco Fisichella
ECIR2
2017 FuseM: Query-Centric Data Fusion on Structured Web Markup
abstract
Embedded markup based on Microdata, RDFa, and Microformats have become prevalent on the Web and constitute an unprecedented source of data. However, RDF statements extracted from markup are fundamentally different from traditional RDF graphs: entity descriptions are flat, facts are highly redundant, and despite very frequent co-references explicit links are missing. Therefore, carrying out typical entity-centric tasks such as retrieval and summarisation cannot be tackled sufficiently with state-of-the-art methods and require preliminary data fusion. Given the scale and dynamics of Web markup, the applicability of general data fusion approaches is limited. We present a novel query-centric data fusion approach which overcomes such issues through a combination of entity retrieval and fusion techniques geared towards the specific challenges associated with embedded markup. To ensure precise and diverse entity descriptions, we follow a supervised learning approach and train a classifier for data fusion of a pool of candidate facts relevant to a given query and obtained through a preliminary entity retrieval step. We perform a thorough evaluation on a subset of the Web Data Commons dataset and show significant improvement over existing baselines. In addition, an investigation into the coverage and complementarity of facts from the constructed entity descriptions compared to DBpedia, shows potential for aiding tasks such as knowledge base population.
Ran Yu 0001, Ujwal Gadiraju, Besnik Fetahu, Stefan Dietze
ICDE2
2017 Improving Reliability of Crowdsourced Results by Detecting Crowd Workers with Multiple Identities
Ujwal Gadiraju, Ricardo Kawase
ICWE1
2016 Where the Event Lies: Predicting Event Occurrence in Textual Documents
abstract
Manually inspecting text in a document collection to assess whether an event occurs in it is a cumbersome task. Although a manual inspection can allow one to identify and discard false events, it becomes infeasible with increasing numbers of automatically detected events. In this paper, we present a system to automatize event validation, defined as the task of determining whether a given event occurs in a given document or corpus. In addition to supporting users seeking for information that corroborates a given event, event validation can also boost the precision of automatically detected event sets by discarding false events and preserving the true ones. The system allows to specify events, retrieves candidate web documents, and assesses whether events occur in them. The validation results are shown to the user, who can revise the decision of the system. The validation method relies on a supervised model to predict the occurrence of events in a non-annotated corpus. This system can also be used to build ground-truths for event corpora.
Andrea Ceroni, Ujwal Gadiraju, Jan Matschke, Simon Wingert, Marco Fisichella
SIGIR2
2015 Improving Event Detection by Automatically Assessing Validity of Event Occurrence in Text
abstract
Manually inspecting text to assess whether an event occurs in a document collection is an onerous and time consuming task. Although a manual inspection to discard the false events would increase the precision of automatically detected sets of events, it is not a scalable approach. In this paper, we automatize event validation, defined as the task of determining whether a given event occurs in a given document or corpus. The introduction of automatic event validation as a post-processing step of event detection can boost the precision of the detected event set, discarding false events and preserving the true ones. We propose a novel automatic method for event validation, which relies on a supervised model to predict the occurrence of events in a non-annotated corpus. The data for training the model is gathered by exploiting the crowdsourcing paradigm. Experiments on real-world events and documents show that our proposed method (i) outperforms the state-of-the-art event validation approach and (ii) increases the precision of event detection while preserving recall.
Andrea Ceroni, Ujwal Gadiraju, Marco Fisichella
CIKM2
2015 Balancing Novelty and Salience: Adaptive Learning to Rank Entities for Timeline Summarization of High-impact Events
abstract
Long-running, high-impact events such as the Boston Marathon bombing often develop through many stages and involve a large number of entities in their unfolding. Timeline summarization of an event by key sentences eases story digestion, but does not distinguish between what a user remembers and what she might want to re-check. In this work, we present a novel approach for timeline summarization of high-impact events, which uses entities instead of sentences for summarizing the event at each individual point in time. Such entity summaries can serve as both (1) important memory cues in a retrospective event consideration and (2) pointers for personalized event exploration. In order to automatically create such summaries, it is crucial to identify the "right" entities for inclusion. We propose to learn a ranking function for entities, with a dynamically adapted trade-off between the in-document salience of entities and the informativeness of entities across documents, i.e., the level of new information associated with an entity for a time point under consideration. Furthermore, for capturing collective attention for an entity we use an innovative soft labeling approach based on Wikipedia. Our experiments on a real large news datasets confirm the effectiveness of the proposed methods.
Tuan Tran 0002, Claudia Niederée, Nattiya Kanhabua, Ujwal Gadiraju, Avishek Anand
CIKM4
2015 Improving Entity Retrieval on Structured Data
Besnik Fetahu, Ujwal Gadiraju, Stefan Dietze
ISWC (1)2
2015 Adaptive Focused Crawling of Linked Data
Ran Yu 0001, Ujwal Gadiraju, Besnik Fetahu, Stefan Dietze
WISE (1)2
2013 Groundhog day: near-duplicate detection on Twitter
abstract
With more than 340~million messages that are posted on Twitter every day, the amount of duplicate content as well as the demand for appropriate duplicate detection mechanisms is increasing tremendously. Yet there exists little research that aims at detecting near-duplicate content on microblogging platforms. We investigate the problem of near-duplicate detection on Twitter and introduce a framework that analyzes the tweets by comparing (i) syntactical characteristics, (ii) semantic similarity, and (iii) contextual information. Our framework provides different duplicate detection strategies that, among others, make use of external Web resources which are referenced from microposts. Machine learning is exploited in order to learn patterns that help identifying duplicate content. We put our duplicate detection framework into practice by integrating it into Twinder, a search engine for Twitter streams. An in-depth analysis shows that it allows Twinder to diversify search results and improve the quality of Twitter search. We conduct extensive experiments in which we (1) evaluate the quality of different strategies for detecting duplicates, (2) analyze the impact of various features on duplicate detection, (3) investigate the quality of strategies that classify to what exact level two microposts can be considered as duplicates and (4) optimize the process of identifying duplicate content on Twitter. Our results prove that semantic features which are extracted by our framework can boost the performance of detecting duplicates.
Ke Tao, Fabian Abel, Claudia Hauff, Geert-Jan Houben, Ujwal Gadiraju
WWW5