VLDB 2026 Research / reviewers in the wild / expert
Shubham Atreja
dblp:183/9909
· DBLP profile ↗
12ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0002-0056-3060ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 6 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2Software engineering, systems software and programming languages · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | What's in a Prompt?: A Large-Scale Experiment to Assess the Impact of Prompt Design on the Compliance and Accuracy of LLM-Generated Text AnnotationsabstractManually annotating data for computational social science tasks can be costly, time-consuming, and emotionally draining. While recent work suggests that LLMs can perform such annotation tasks in zero-shot settings, little is known about how prompt design impacts LLMs' compliance and accuracy. We conduct a large-scale multi-prompt experiment to test how model selection (GPT-4o, GPT-3.5, PaLM2, and Falcon7b) and prompt design features (definition inclusion, output type, explanation, and prompt length) impact the compliance and accuracy of LLM-generated annotations on four highly relevant and diverse CSS tasks (toxicity, sentiment, rumor stance, and news frames). Our results show that LLM compliance and accuracy are prompt-dependent. For instance, prompting for numerical scores instead of labels reduces all LLMs' compliance and accuracy. Concise prompts can significantly reduce prompting costs but also lead to lower accuracy on tasks like toxicity. Furthermore, minor prompt changes like asking for an explanation can cause large changes in the distribution of LLM-generated labels. By assessing the impact of prompt design on the quality and distribution of LLM-generated annotations, this work serves as both a practical guide and a warning for using LLMs in CSS research. Shubham Atreja, Joshua Ashkinaze, Lingyao Li, Julia Mendelsohn, Libby Hemphill |
ICWSM | 1 |
| 2024 | AppealMod: Inducing Friction to Reduce Moderator Workload of Handling User AppealsabstractAs content moderation becomes a central aspect of all social media platforms and online communities, interest has grown in how to make moderation decisions contestable. On social media platforms where individual communities moderate their own activities, the responsibility to address user appeals falls on volunteers from within the community. While there is a growing body of work devoted to understanding and supporting the volunteer moderators' workload, little is known about their practice of handling user appeals. Through a collaborative and iterative design process with Reddit moderators, we found that moderators spend considerable effort in investigating user ban appeals and desired to directly engage with users and retain their agency over each decision. To fulfill their needs, we designed and built AppealMod, a system that induces friction in the appeals process by asking users to provide additional information before their appeals are reviewed by human moderators. In addition to giving moderators more information, we expected the friction in the appeal process would lead to a selection effect among users, with many insincere and toxic appeals being abandoned before getting any attention from human moderators. To evaluate our system, we conducted a randomized field experiment in a Reddit community of over 29 million users that lasted for four months. As a result of the selection effect, moderators viewed only 30% of initial appeals and less than 10% of the toxically worded appeals; yet they granted roughly the same number of appeals when compared with the control group. Overall, our system is effective at reducing moderator workload and minimizing their exposure to toxic content while honoring their preference for direct engagement and agency in appeals. Shubham Atreja, Jane Im, Paul Resnick, Libby Hemphill |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2024 | "HOT" ChatGPT: The Promise of ChatGPT in Detecting and Discriminating Hateful, Offensive, and Toxic Comments on Social MediaabstractHarmful textual content is pervasive on social media, poisoning online communities and negatively impacting participation. A common approach to this issue is developing detection models that rely on human annotations. However, the tasks required to build such models expose annotators to harmful and offensive content and may require significant time and cost to complete. Generative AI models have the potential to understand and detect harmful textual content. We used ChatGPT to investigate this potential and compared its performance with MTurker annotations for three frequently discussed concepts related to harmful textual content on social media: Hateful, Offensive, and Toxic (HOT). We designed five prompts to interact with ChatGPT and conducted four experiments eliciting HOT classifications. Our results show that ChatGPT can achieve an accuracy of approximately 80% when compared to MTurker annotations. Specifically, the model displays a more consistent classification for non-HOT comments than HOT comments compared to human annotations. Our findings also suggest that ChatGPT classifications align with the provided HOT definitions. However, ChatGPT classifies “hateful” and “offensive” as subsets of “toxic.” Moreover, the choice of prompts used to interact with ChatGPT impacts its performance. Based on these insights, our study provides several meaningful implications for employing ChatGPT to detect HOT content, particularly regarding the reliability and consistency of its performance, its understanding and reasoning of the HOT concept, and the impact of prompts on its performance. Overall, our study provides guidance on the potential of using generative AI models for moderating large volumes of user-generated textual content on social media. Lingyao Li, Lizhou Fan, Shubham Atreja, Libby Hemphill |
ACM Trans. Web | 3 |
| 2023 | Understanding Journalists' Workflows in News CurationabstractWith the increasing dominance of internet as a source of news consumption, there has been a rise in the production and popularity of email newsletters compiled by individual journalists. However, there is little research on the processes of aggregation, and how these differ between expert journalists and trained machines. In this paper, we interviewed journalists who curate newsletters from around the world. Through an in-depth understanding of journalists’ workflows, our findings lay out the role of their prior experience in the value they bring into the curation process, their own use of algorithms in finding stories for their newsletter, and their internalization of their readers’ interests and the context they are curating for. While identifying the role of human expertise, we highlight the importance of hybrid curation and provide design insights on how technology can support the work of these experts. Shubham Atreja, Shruthi Srinath, Joyojeet Pal |
CHI | 1 |
| 2023 | Remove, Reduce, Inform: What Actions do People Want Social Media Platforms to Take on Potentially Misleading Content?abstractTo reduce the spread of misinformation, social media platforms may take enforcement actions against offending content, such as adding informational warning labels, reducing distribution, or removing content entirely. However, both their actions and their inactions have been controversial and plagued by allegations of partisan bias. When it comes to specific content items, surprisingly little is known about what ordinary people want the platforms to do. We provide empirical evidence about a politically balanced panel of lay raters' preferences for three potential platform actions on 368 news articles. Our results confirm that on many articles there is a lack of consensus on which actions to take. We find a clear hierarchy of perceived severity of actions with a majority of raters wanting informational labels on the most articles and removal on the fewest. There was no partisan difference in terms of how many articles deserve platform actions but conservatives did prefer somewhat more action on content from liberal sources, and vice versa. We also find that judgments about two holistic properties, misleadingness and harm, could serve as an effective proxy to determine what actions would be approved by a majority of raters. Shubham Atreja, Libby Hemphill, Paul Resnick |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2022 | Social Debunking of Misinformation on WhatsApp: The Case for Strong and In-group TiesabstractIn this paper, we argue that WhatsApp can play an important role in correcting misinformation. We show how specific WhatsApp affordances (flexibility in format and audience selection) and existing social capital (prevalence of strong ties; homophily in political groups) can be leveraged to maximize the re-sharing of debunking messages, such as those accessed by WhatsApp users via ChatBots and Tip-Lines. Debunking messages received in the format of audio files generated more interest and were more effective in correcting beliefs than text- or image-based messages. In addition, we found clear evidence that users re-share debunks at higher rates when they received them from people close to them (strong ties), from individuals who generally agree with them politically (in-group members), or when both conditions are met. We suggest that WhatsApp leverages our findings to maximize the re-share of those fact-checks that are already circulating on the platform by using the existing social capital in the network, unlocking the potential for such debunks to reach a larger audience on WhatsApp. Irene V. Pasquetto, Eaman Jahani, Shubham Atreja, Matthew Baum |
Proc. ACM Hum. Comput. Interact. | 3 |
| 2019 | Automatic Generation of Leveled Visual Assessments for Young LearnersabstractImages are an essential tool for communicating with children, particularly at younger ages when they are still developing their emergent literacy skills. Hence, assessments that use images to assess their conceptual knowledge and visual literacy, are an important component of their learning process. Creating assessments at scale is a challenging task, which has led to several techniques being proposed for automatic generation of textual assessments. However, none of them focuses on generating image-based assessments. To understand the manual process of creating visual assessments, we interviewed primary school teachers. Based on the findings from the preliminary study, we present a novel approach which uses image semantics to generate visual multiple choice questions (VMCQs) for young learners, wherein options are presented in the form of images. We propose a metric to measure the semantic similarity between two images, which we use to identify the four options – one answer and three distractor images – for a given question. We also use this metric for generating VMCQs at two difficulty levels – easy and hard. Through a quantitative evaluation, we show that the system-generated VMCQs are comparable to VMCQs created by experts, hence establishing the effectiveness of our approach. Ruhi Sharma Mittal, Shubham Atreja, Mourvi Sharma, Seema Nagar |
AAAI | 3 |
| 2019 | Adversarial Adaptation of Scene Graph Models for Understanding Civic IssuesabstractCitizen engagement and technology usage are two emerging trends driven by smart city initiatives. Typically, citizens report issues, such as broken roads, garbage dumps, etc. through web portals and mobile apps, in order for the government authorities to take appropriate actions. Several mediums - text, image, audio, video - are used to report these issues. Through a user study with 13 citizens and 3 authorities, we found that image is the most preferred medium to report civic issues. However, analyzing civic issue related images is challenging for the authorities as it requires manual effort. In this work, given an image, we propose to generate a Civic Issue Graph consisting of a set of objects and the semantic relations between them, which are representative of the underlying civic issue. We also release two multi-modal (text and images) datasets, that can help in further analysis of civic issues from images. We present an approach for adversarial adaptation of existing scene graph models that enables the use of scene graphs for new applications in the absence of any labelled training data. We conduct several experiments to analyze the efficacy of our approach, and using human evaluation, we establish the appropriateness of our model at representing different civic issues. Shanu Kumar, Shubham Atreja |
WWW | 2 |
| 2018 | Citicafe: An Interactive Interface for Citizen EngagementabstractCommunity engagement is a new and emerging trend in urban cities driven by the mission of developing responsible citizenship. The platform ingests data from different sources, which is exploited by a virtual agent to enable informed interactions. It can help citizens to (a) report problems and (b) gather information related to civic issues for different locations and their neighborhoods. We report the results of a user study carried out to establish the effectiveness of our interface and draw a comparison with an existing platform. A detailed qualitative and quantitative analysis of the survey results shows a definite and statistically significant (p < 0.05) preference for our interface over the existing platform. Shubham Atreja, Pooja Aggarwal, Prateeti Mohapatra, Amol Dumrewal, Anwesh Basu, Gargi Dasgupta |
IUI | 1 |
| 2017 | Continuous Learning as a Service for Conversational Virtual Agents
Shivali Agarwal, Shubham Atreja, Gargi Dasgupta |
ICSOC | 2 |
| 2017 | "Learning Relevance" as a Service for Improving Search Results in Technical Discussion ForumsabstractSearch results in technical forums are typically keyword based. The relevance of a link is usually gauged by closest content match. However, it has been shown in literature that users' click behavior is an integral part of deciding the relevance of a search result. Moreover, it is not just the number of clicks that matter, but time spent on a clicked link, order in which the links were clicked etc. also play an important role in the relevance decision. In this paper, we have developed a service that analyzes the click logs of searches performed in the technical forums and learns the new relevance scores for the search results with respect to a query. The computation model for relevance is an optimization problem, the constraints for which have been designed based on real user behavior study. We ingested StackOverflow data for few domains and designed a QA style search to carry out the study. We have developed heuristics to solve the optimization problem and have validated the relevance model using user behavior simulations. The relevance model is shown to yield efficient, robust and effective rank order using DCG (discounted cumulative gains) and stability metrics. Shubham Atreja, Shivali Agarwal, Gargi Dasgupta, Dennis A. Perpetua |
ICWS | 1 |
| 2016 | Anomaly Detection Using Program Control Flow Graph Mining From Execution LogsabstractWe focus on the problem of detecting anomalous run-time behavior of distributed applications from their execution logs. Specifically we mine templates and template sequences from logs to form a control flow graph (cfg) spanning distributed components. This cfg represents the baseline healthy system state and is used to flag deviations from the expected behavior of runtime logs. The novelty in our work stems from the new techniques employed to: (1) overcome the instrumentation requirements or application specific assumptions made in prior log mining approaches, (2) improve the accuracy of mined templates and the cfg in the presence of long parameters and high amount of interleaving respectively, and (3) improve by orders of magnitude the scalability of the cfg mining process in terms of volume of log data that can be processed per day. We evaluate our approach using (a) synthetic log traces and (b) multiple real-world log datasets collected at different layers of application stack. Results demonstrate that our template mining, cfg mining, and anomaly detection algorithms have high accuracy. The distributed implementation of our pipeline is highly scalable and has more than 500 GB/day of log data processing capability even on a 10 low-end VM based (Spark + Hadoop) cluster. We also demonstrate the efficacy of our end-to-end system using a case study with the Openstack VM provisioning system. Animesh Nandi, Atri Mandal, Shubham Atreja, Gargi Dasgupta, Subhrajit Bhattacharya |
KDD | 3 |