VLDB 2026 Research / reviewers in the wild / expert
David Jurgens
dblp:48/4613
· DBLP profile ↗
17ranked-venue papers in the field
4as first author
7since 2021 · last 2025
0000-0002-2135-9878ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 16 (4 first)Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Roles of Network and Identity in Hashtag DiffusionabstractThe diffusion of culture online is theorized to be influenced by many interacting social factors (e.g., network and identity).However, most existing computational cascade models consider just a single factor (e.g., network or identity).This work offers a new framework for teasing apart the mechanisms underlying hashtag cascades.We curate a new dataset of 1,337 hashtags representing cultural innovation online, develop a 10-factor evaluation framework for comparing empirical and simulated cascades, and show that a combined network+identity model better simulates hashtag cascades than network-or identity-only counterfactuals.We also explore heterogeneity in performance: While a combined network+identity model best predicts the popularity of cascades, a network-only model best predicts cascade growth and an identity-only model best predicts adopter composition.The network+identity model has the highest comparative advantage among hashtags used for expressing racial or regional identity and talking about sports or news.In fact, we are able to predict what combination of network and/or identity best models each hashtag and use this to further improve performance.Our results show the utility of models incorporating the interactions of network, identity, and other social factors in the diffusion of hashtags in social media. Aparna Ananthasubramaniam, Yufei 'Louise' Zhu, David Jurgens, Daniel M. Romero |
WWW | 3 |
| 2024 | A Multilingual Similarity Dataset for News Article FrameabstractUnderstanding the writing frame of news articles is vital for addressing social issues, and thus has attracted notable attention in the fields of communication studies. Yet, assessing such news article frame remains a challenge due to the absence of a concrete and unified standard dataset that considers the comprehensive nuances within news content. To address this gap, we introduce an extended version of a large labeled news article dataset with 16,687 new labeled pairs. Leveraging the pairwise comparison of news articles, our method frees the work of manual identification of frame classes in traditional news frame analysis studies. Overall we introduce the most extensive cross-lingual news article similarity dataset available to date with 26,555 labeled news article pairs across 10 languages. Each data point has been meticulously annotated according to a codebook detailing eight critical aspects of news content, under a human-in-the-loop framework. Application examples demonstrate its potential in unearthing country communities within global news coverage, exposing media bias among news outlets, and quantifying the factors related to news creation. We envision that this news similarity dataset will broaden our understanding of the media ecosystem in terms of news coverage of events and perspectives across countries, locations, languages, and other social constructs. By doing so, it can catalyze advancements in social science research and applied methodologies, thereby exerting a profound impact on our society. Xi Chen 0125, Mattia Samory, Scott A. Hale, David Jurgens, Przemyslaw A. Grabowicz |
ICWSM | 4 |
| 2024 | Global News Synchrony and Diversity During the Start of the COVID-19 PandemicabstractNews coverage profoundly affects how countries and individuals behave in international relations. Yet, we have little empirical evidence of how news coverage varies across countries. To enable studies of global news coverage, we develop an efficient computational methodology that comprises three components: (i) a transformer model to estimate multilingual news similarity; (ii) a global event identification system that clusters news based on a similarity network of news articles; and (iii) measures of news synchrony across countries and news diversity within a country, based on country-specific distributions of news coverage of the global events. Each component achieves state-of-the art performance, scaling seamlessly to massive datasets of millions of news articles. Xi Chen 0125, Scott A. Hale, David Jurgens, Mattia Samory, Ethan Zuckerman, Przemyslaw A. Grabowicz |
WWW | 3 |
| 2023 | Analyzing the Engagement of Social Relationships during Life Event Shocks in Social MediaabstractIndividuals experiencing unexpected distressing events, shocks, often rely on their social network for support. While prior work has shown how social networks respond to shocks, these studies usually treat all ties equally, despite differences in the support provided by different social relationships. Here, we conduct a computational analysis on Twitter that examines how responses to online shocks differ by the relationship type of a user dyad. We introduce a new dataset of over 13K instances of individuals' self-reporting shock events on Twitter and construct networks of relationship-labeled dyadic interactions around these events. By examining behaviors across 110K replies to shocked users in a pseudo-causal analysis, we demonstrate relationship-specific patterns in response levels and topic shifts. We also show that while well-established social dimensions of closeness such as tie strength and structural embeddedness contribute to shock responsiveness, the degree of impact is highly dependent on relationship and shock types. Our findings indicate that social relationships contain highly distinctive characteristics in network interactions, and that relationship-specific behaviors in online shock responses are unique from those of offline settings. Minje Choi, David Jurgens, Daniel M. Romero |
ICWSM | 2 |
| 2023 | Bridging Nations: Quantifying the Role of Multilinguals in Communication on Social MediaabstractSocial media enables the rapid spread of many kinds of information, from pop culture memes to social movements. However, little is known about how information crosses linguistic boundaries. We apply causal inference techniques on the European Twitter network to quantify the structural role and communication influence of multilingual users in cross-lingual information exchange. Overall, multilinguals play an essential role; posting in multiple languages increases betweenness centrality by 13%, and having a multilingual network neighbor increases monolinguals’ odds of sharing domains and hashtags from another language 16-fold and 4-fold, respectively. We further show that multilinguals have a greater impact on diffusing information is less accessible to their monolingual compatriots, such as information from far-away countries and content about regional politics, nascent social movements, and job opportunities. By highlighting information exchange across borders, this work sheds light on a crucial component of how information and ideas spread around the world. Julia Mendelsohn, Sayan Ghosh 0004, David Jurgens, Ceren Budak |
ICWSM | 3 |
| 2021 | More than Meets the Tie: Examining the Role of Interpersonal Relationships in Social Networks
Minje Choi, Ceren Budak, Daniel M. Romero, David Jurgens |
ICWSM | 4 |
| 2021 | Conversations Gone Alright: Quantifying and Predicting Prosocial Outcomes in Online ConversationsabstractOnline conversations can go in many directions: some turn out poorly due to antisocial behavior, while others turn out positively to the benefit of all. Research on improving online spaces has focused primarily on detecting and reducing antisocial behavior. Yet we know little about positive outcomes in online conversations and how to increase them—is a prosocial outcome simply the lack of antisocial behavior or something more? Here, we examine how conversational features lead to prosocial outcomes within online discussions. We introduce a series of new theory-inspired metrics to define prosocial outcomes such as mentoring and esteem enhancement. Using a corpus of 26M Reddit conversations, we show that these outcomes can be forecasted from the initial comment of an online conversation, with the best model providing a relative 24% improvement over human forecasting performance at ranking conversations for predicted outcome. Our results indicate that platforms can use these early cues in their algorithmic ranking of early conversations to prioritize better outcomes. Jiajun Bao, Junjie Wu 0007, Eshwar Chandrasekharan, David Jurgens |
WWW | 5 |
| 2019 | Smart, Responsible, and Upper Caste Only: Measuring Caste Attitudes through Large-Scale Analysis of Matrimonial Profiles
Ashwin Rajadesingan, Ramaswami Mahalingam, David Jurgens |
ICWSM | 3 |
| 2019 | Are All Successful Communities Alike? Characterizing and Predicting the Success of Online CommunitiesabstractThe proliferation of online communities has created exciting opportunities to study the mechanisms that explain group success. While a growing body of research investigates community success through a single measure - typically, the number of members - we argue that there are multiple ways of measuring success. Here, we present a systematic study to understand the relations between these success definitions and test how well they can be predicted based on community properties and behaviors from the earliest period of a community's lifetime. We identify four success measures that are desirable for most communities: (i) growth in the number of members; (ii) retention of members; (iii) long term survival of the community; and (iv) volume of activities within the community. Surprisingly, we find that our measures do not exhibit very high correlations, suggesting that they capture different types of success. Additionally, we find that different success measures are predicted by different attributes of online communities, suggesting that success can be achieved through different behaviors. Our work sheds light on the basic understanding on what success represents in online communities and what predicts it. Our results suggest that success is multi-faceted and cannot be measured nor predicted by a single measurement. This insight has practical implications for the creation of new online communities and the design of platforms that facilitate such communities. David Jurgens, Chenhao Tan, Daniel M. Romero |
WWW | 2 |
| 2019 | Demographic Inference and Representative Population Estimates from Multilingual Social Media DataabstractSocial media provide access to behavioural data at an unprecedented scale and granularity. However, using these data to understand phenomena in a broader population is difficult due to their non-representativeness and the bias of statistical inference tools towards dominant languages and groups. While demographic attribute inference could be used to mitigate such bias, current techniques are almost entirely monolingual and fail to work in a global environment. We address these challenges by combining multilingual demographic inference with post-stratification to create a more representative population sample. To learn demographic attributes, we create a new multimodal deep neural architecture for joint classification of age, gender, and organization-status of social media users that operates in 32 languages. This method substantially outperforms current state of the art while also reducing algorithmic bias. To correct for sampling biases, we propose fully interpretable multilevel regression methods that estimate inclusion probabilities from inferred joint population counts and ground-truth population counts. Zijian Wang 0002, Scott A. Hale, David Ifeoluwa Adelani, Przemyslaw A. Grabowicz, Timo Hartmann, Fabian Flöck, David Jurgens |
WWW | 7 |
| 2016 | User Migration in Online Social Networks: A Case Study on Reddit During a Period of Community Unrest
Edward Newell, David Jurgens, Haji Mohammad Saleem, Hardik Vala, Jad Sassine, Caitrin Armstrong, Derek Ruths |
ICWSM | 2 |
| 2015 | Geolocation Prediction in Twitter Using Social Networks: A Critical Analysis and Review of Current Practice
David Jurgens, Tyler Finethy, James McCorriston, Yi Tian Xu, Derek Ruths |
ICWSM | 1 |
| 2015 | An Analysis of Exercising Behavior in Online Populations
David Jurgens, James McCorriston, Derek Ruths |
ICWSM | 1 |
| 2015 | Organizations Are Users Too: Characterizing and Detecting the Presence of Organizations on Twitter
James McCorriston, David Jurgens, Derek Ruths |
ICWSM | 2 |
| 2014 | Geotagging one hundred million Twitter accounts with total variation minimizationabstractGeographically annotated social media is extremely valuable for modern information retrieval. However, when researchers can only access publicly-visible data, one quickly finds that social media users rarely publish location information. In this work, we provide a method which can geolocate the overwhelming majority of active Twitter users, independent of their location sharing preferences, using only publicly-visible Twitter data. Our method infers an unknown user's location by examining their friend's locations. We frame the geotagging problem as an optimization over a social network with a total variation-based objective and provide a scalable and distributed algorithm for its solution. Furthermore, we show how a robust estimate of the geographic dispersion of each user's ego network can be used as a per-user accuracy measure which is effective at removing outlying errors. Leave-many-out evaluation shows that our method is able to infer location for 101, 846, 236 Twitter users at a median error of 6.38 km, allowing us to geotag over 80% of public tweets. Ryan Compton 0001, David Jurgens |
IEEE BigData | 2 |
| 2013 | That's What Friends Are For: Inferring Location in Online Social Media Platforms Based on Social Relationships
David Jurgens |
ICWSM | 1 |
| 2012 | Temporal Motifs Reveal the Dynamics of Editor Interactions in Wikipedia
David Jurgens, Tsai-Ching Lu |
ICWSM | 1 |