EDBT 2026 Demo / reviewers in the wild / expert
Dan Goldwasser
dblp:38/3382
· DBLP profile ↗
9ranked-venue papers in the field
0as first author
5since 2021 · last 2025
0000-0001-9326-8601ORCID · corroborated
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 6Big Data, Cloud & Distributed Data Systems · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Discovering Latent Themes in Social Media Messaging: A Machine-in-the-Loop Approach Integrating LLMsabstractGrasping the themes of social media content is key to understanding the narratives that influence public opinion and behavior. The thematic analysis goes beyond traditional topic-level analysis, which often captures only the broadest patterns, providing deeper insights into specific and actionable themes such as "public sentiment towards vaccination", "political discourse surrounding climate policies," etc. In this paper, we introduce a novel approach to uncovering latent themes in social media messaging. Recognizing the limitations of the traditional topic-level analysis, which tends to capture only overarching patterns, this study emphasizes the need for a finer-grained, theme-focused exploration. Traditional theme discovery methods typically involve manual processes and a human-in-the-loop approach. While valuable, these methods face challenges in scalability, consistency, and resource intensity in terms of time and cost. To address these challenges, we propose a machine-in-the-loop approach that leverages the advanced capabilities of large language models (LLMs). This approach facilitates a deeper investigation into the social media discourse, revealing a variety of themes with distinct characteristics and relevance. It provides a detailed understanding of underlying nuances and efficiently maps texts to these themes, enhancing our insight into social media messaging. To demonstrate our approach, we apply our framework to contentious topics, such as climate debate and vaccine debate. We use two publicly available datasets: (1) the climate campaigns dataset of 21k Facebook ads and (2) the COVID-19 vaccine campaigns dataset of 9k Facebook ads. Our quantitative and qualitative analysis shows that our methodology yields more accurate and interpretable results compared to the baselines. Our results not only demonstrate the effectiveness of our approach in uncovering latent themes but also illuminate how these themes are tailored for demographic targeting in social media contexts. Additionally, our work sheds light on the dynamic nature of social media, revealing the shifts in the thematic focus of messaging in response to real-world events. Tunazzina Islam, Dan Goldwasser |
ICWSM | 2 |
| 2023 | Weakly Supervised Learning for Analyzing Political Campaigns on FacebookabstractSocial media platforms are currently the main channel for political messaging, allowing politicians to target specific demographics and adapt based on their reactions. However, making this communication transparent is challenging, as the messaging is tightly coupled with its intended audience and often echoed by multiple stakeholders interested in advancing specific policies. Our goal in this paper is to take a first step towards understanding these highly decentralized settings. We propose a weakly supervised approach to identify the stance and issue of political ads on Facebook and analyze how political campaigns use some kind of demographic targeting by location, gender, or age. Furthermore, we analyze the temporal dynamics of the political ads on election polls. Tunazzina Islam, Shamik Roy, Dan Goldwasser |
ICWSM | 3 |
| 2022 | Understanding COVID-19 Vaccine Campaign on Facebook using Minimal SupervisionabstractIn the age of social media, where billions of internet users share information and opinions, the negative impact of pandemics is not limited to the physical world. It provokes a surge of incomplete, biased, and incorrect information, also known as an infodemic. This global infodemic jeopardizes measures to control the pandemic by creating panic, vaccine hesitancy, and fragmented social response. Platforms like Facebook allow advertisers to adapt their messaging to target different demographics and help alleviate or exacerbate the infodemic problem depending on their content. In this paper, we propose a minimally supervised multi-task learning framework for understanding messaging on Facebook related to the COVID vaccine by identifying ad themes and moral foundations. Furthermore, we perform a more nuanced thematic analysis of messaging tactics of vaccine campaigns on social media so that policymakers can make better decisions on pandemic control. Tunazzina Islam, Dan Goldwasser |
IEEE Big Data | 2 |
| 2022 | Twitter User Representation Using Weakly Supervised Graph Embedding
Tunazzina Islam, Dan Goldwasser |
ICWSM | 2 |
| 2021 | Analysis of Twitter Users' Lifestyle Choices using Joint Embedding Model
Tunazzina Islam, Dan Goldwasser |
ICWSM | 2 |
| 2020 | Does Yoga Make You Happy? Analyzing Twitter User Happiness using Textual and Temporal InformationabstractAlthough yoga is a multi-component practice to hone the body and mind and be known to reduce anxiety and depression, there is still a gap in understanding people's emotional state related to yoga in social media. In this study, we investigate the causal relationship between practicing yoga and being happy by incorporating textual and temporal information of users using Granger causality. To find out causal features from the text, we measure two variables (i) Yoga activity level based on content analysis and (ii) Happiness level based on emotional state. To understand users' yoga activity, we propose a joint embedding model based on the fusion of neural networks with attention mechanism by leveraging users' social and textual information. For measuring the emotional state of yoga users (target domain), we suggest a transfer learning approach to transfer knowledge from an attention-based neural network model trained on a source domain. Our experiment on Twitter dataset demonstrates that there are 1447 users where "yoga Granger-causes happiness". Tunazzina Islam, Dan Goldwasser |
IEEE BigData | 2 |
| 2019 | ACE - An Anomaly Contribution Explainer for Cyber-Security ApplicationsabstractIn this paper we introduce Anomaly Contribution Explainer or ACE, a tool to explain security anomaly detection models in terms of the model features through a regression framework, and its variant, ACE-KL, which highlights the important anomaly contributors. ACE and ACE-KL provide insights in diagnosing which attributes significantly contribute to an anomaly by building a specialized linear model to locally approximate the anomaly score that a black-box model generates. We conducted experiments with these anomaly detection models to detect security anomalies on both synthetic data and real data. In particular, we evaluate performance on three public data sets: CERT insider threat, netflow logs, and Android malware. The experimental results are encouraging: our methods consistently identify the correct contributing feature in the synthetic data where ground truth is available; similarly, for real data sets, our methods point a security analyst in the direction of the underlying causes of an anomaly, including in one case leading to the discovery of previously overlooked network scanning activity. We have made our source code publicly available. Xiao Zhang 0017, Manish Marwah, I-Ta Lee, Martin F. Arlitt, Dan Goldwasser |
IEEE BigData | 5 |
| 2017 | TATHYA: A Multi-Classifier System for Detecting Check-Worthy Statements in Political DebatesabstractFact-checking political discussions has become an essential clog in computational journalism. This task encompasses an important sub-task---identifying the set of statements with 'check-worthy' claims. Previous work has treated this as a simple text classification problem discounting the nuances involved in determining what makes statements check-worthy. We introduce a dataset of political debates from the 2016 US Presidential election campaign annotated using all major fact-checking media outlets and show that there is a need to model conversation context, debate dynamics and implicit world knowledge. We design a multi-classifier system TATHYA, that models latent groupings in data and improves state-of-art systems in detecting check-worthy statements by 19.5% in F1-score on a held-out test set, gaining primarily gaining in Recall. Ayush Patwari, Dan Goldwasser, Saurabh Bagchi |
CIKM | 2 |
| 2017 | Modeling of Political Discourse Framing on Twitter
Kristen Johnson, Dan Goldwasser |
ICWSM | 3 |