VLDB 2026 Research / reviewers in the wild / expert
Przemyslaw A. Grabowicz
dblp:12/9888
· DBLP profile ↗
13ranked-venue papers in the field
3as first author
5since 2021 · last 2025
0000-0002-6043-6928ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 11 (2 first)Data Mining & Knowledge Discovery · 2 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Effects of Research Paper Promotion via ArXiv and XabstractIn the evolving landscape of scientific publishing, it is important to understand the drivers of high-impact research, to equip scientists with actionable strategies to enhance the reach of their work, and to understand trends in the use of modern scientific publishing tools to inform their further development. Here, based on a dataset of over 0.5 million publications in computer science and physics, we study trends in the use of early preprint publications and revisions on ArXiv and the use of X (formerly Twitter) for promotion of such papers. We find that early submissions to ArXiv and promotion on X have soared in recent years. Estimating the effect that the use of each of these modern affordances has on the number of citations of scientific publications, we find that peer-reviewed conference papers in computer science that are submitted early to ArXiv gain on average 21.1 ± 17.4 more citations, revised on ArXiv gain 18.4 ± 17.6 more citations, and promoted on X gain 44.4 ± 8 more citations in the first 5 years from an initial publication. In contrast, journal articles in physics experience comparatively lower boosts in citation counts, with increases of 3.9 ± 1.1, 4.3 ± 0.9, and 6.9 ± 3.5 citations respectively for the same interventions. Our results show that promoting one's work on ArXiv or X has a large impact on the number of citations, as well as the number of influential citations computed by Semantic Scholar, and thereby on the career of researchers. These effects are present also for publications in physics, but they are relatively smaller. The larger relative effect sizes, effects of promotion accumulating over time, and elevated unpredictability of the number of citations in computer science than in physics suggest a greater role of world-of-mouth spreading in computer science than in physics. We discuss the far-reaching implications of these findings for future scientific publishing systems and measures of scientific impact. Chhandak Bagchi, Eric Malmi, Przemyslaw A. Grabowicz |
ICWSM | 3 |
| 2025 | Identifying and Investigating Global News Coverage of Critical Events Such as Disasters and Terrorist AttacksabstractComparative studies of news coverage are challenging to conduct because methods to identify news articles about the same event in different languages require expertise that is difficult to scale. We introduce an AI-powered method for identifying news articles based on an event fingerprint, which is a minimal set of metadata required to identify critical events. Our event coverage identification method, FINGERPRINT TO ARTICLE MATCHING FOR EVENTS (FAME), efficiently identifies news articles about critical world events, specifically terrorist attacks and several types of natural disasters. FAME does not require training data and is able to automatically and efficiently identify news articles that discuss an event given its fingerprint: time, location, and class (such as storm or flood). The method achieves state-of-the-art performance and scales to massive databases of tens of millions of news articles and hundreds of events happening globally. We use FAME to identify 27,441 articles that cover 470 natural disaster and terrorist attack events that happened in 2020. To this end, we use a massive database of news articles in three languages from MediaCloud, and three widely used, expert-curated databases of critical events: EM-DAT, USGS, and GTD. Our case study reveals patterns consistent with prior literature: coverage of disasters and terrorist attacks correlates to death counts, to the GDP of a country where the event occurs, and to trade volume between the reporting country and the country where the event occurred. We share our NLP annotations and cross-country media attention data to support the efforts of researchers and media monitoring organizations. Erica Cai, Xi Chen 0125, Reagan Grey Keeney, Ethan Zuckerman, Brendan T. O'Connor 0001, Przemyslaw A. Grabowicz |
ICWSM | 6 |
| 2025 | Election Polls on Social Media: Prevalence, Biases, and Voter Fraud BeliefsabstractSocial media platforms allow users to create polls to gather public opinion on diverse topics. However, we know little about what such polls are used for and how reliable they are, especially in significant contexts like elections. Focusing on the 2020 presidential elections in the U.S., this study shows that outcomes of election polls on Twitter deviate from election results despite their prevalence. Leveraging demographic inference and statistical analysis, we find that Twitter polls are disproportionately authored by male Republicans and exhibit a large bias towards candidate Donald Trump in comparison to mainstream polls. We investigate potential sources of biased outcomes from the point of view of inauthentic, automated, and counter-normative behavior. Using social media experiments and interviews with poll authors, we identify inconsistencies between public vote counts and those privately visible to poll authors, with the gap potentially attributable to purchased votes. We find that election polls tend to be more biased, contain more questionable votes, and attract more bots before the election day than after. We highlight and compare key factors contributing to biased poll outcomes. Finally, we identify instances of polls spreading voter fraud conspiracy theories and estimate that a couple of thousand such polls were posted in 2020. The study discusses the implications of biased election polls in the context of transparency and accountability of social media platforms. Stephen Scarano, Vijayalakshmi Vasudevan, Mattia Samory, Kai-Cheng Yang, JungHwan Yang, Przemyslaw A. Grabowicz |
ICWSM | 6 |
| 2024 | A Multilingual Similarity Dataset for News Article FrameabstractUnderstanding the writing frame of news articles is vital for addressing social issues, and thus has attracted notable attention in the fields of communication studies. Yet, assessing such news article frame remains a challenge due to the absence of a concrete and unified standard dataset that considers the comprehensive nuances within news content. To address this gap, we introduce an extended version of a large labeled news article dataset with 16,687 new labeled pairs. Leveraging the pairwise comparison of news articles, our method frees the work of manual identification of frame classes in traditional news frame analysis studies. Overall we introduce the most extensive cross-lingual news article similarity dataset available to date with 26,555 labeled news article pairs across 10 languages. Each data point has been meticulously annotated according to a codebook detailing eight critical aspects of news content, under a human-in-the-loop framework. Application examples demonstrate its potential in unearthing country communities within global news coverage, exposing media bias among news outlets, and quantifying the factors related to news creation. We envision that this news similarity dataset will broaden our understanding of the media ecosystem in terms of news coverage of events and perspectives across countries, locations, languages, and other social constructs. By doing so, it can catalyze advancements in social science research and applied methodologies, thereby exerting a profound impact on our society. Xi Chen 0125, Mattia Samory, Scott A. Hale, David Jurgens, Przemyslaw A. Grabowicz |
ICWSM | 5 |
| 2024 | Global News Synchrony and Diversity During the Start of the COVID-19 PandemicabstractNews coverage profoundly affects how countries and individuals behave in international relations. Yet, we have little empirical evidence of how news coverage varies across countries. To enable studies of global news coverage, we develop an efficient computational methodology that comprises three components: (i) a transformer model to estimate multilingual news similarity; (ii) a global event identification system that clusters news based on a similarity network of news articles; and (iii) measures of news synchrony across countries and news diversity within a country, based on country-specific distributions of news coverage of the global events. Each component achieves state-of-the art performance, scaling seamlessly to massive datasets of millions of news articles. Xi Chen 0125, Scott A. Hale, David Jurgens, Mattia Samory, Ethan Zuckerman, Przemyslaw A. Grabowicz |
WWW | 6 |
| 2019 | Demographic Inference and Representative Population Estimates from Multilingual Social Media DataabstractSocial media provide access to behavioural data at an unprecedented scale and granularity. However, using these data to understand phenomena in a broader population is difficult due to their non-representativeness and the bias of statistical inference tools towards dominant languages and groups. While demographic attribute inference could be used to mitigate such bias, current techniques are almost entirely monolingual and fail to work in a global environment. We address these challenges by combining multilingual demographic inference with post-stratification to create a more representative population sample. To learn demographic attributes, we create a new multimodal deep neural architecture for joint classification of age, gender, and organization-status of social media users that operates in 32 languages. This method substantially outperforms current state of the art while also reducing algorithmic bias. To correct for sampling biases, we propose fully interpretable multilevel regression methods that estimate inclusion probabilities from inferred joint population counts and ground-truth population counts. Zijian Wang 0002, Scott A. Hale, David Ifeoluwa Adelani, Przemyslaw A. Grabowicz, Timo Hartmann, Fabian Flöck, David Jurgens |
WWW | 4 |
| 2017 | Predicting the Success of Online Petitions Leveraging Multidimensional Time-SeriesabstractApplying classical time-series analysis techniques to online content is challenging, as web data tends to have data quality issues and is often incomplete, noisy, or poorly aligned. In this paper, we tackle the problem of predicting the evolution of a time series of user activity on the web in a manner that is both accurate and interpretable, using related time series to produce a more accurate prediction. We test our methods in the context of predicting signatures for online petitions using data from thousands of petitions posted on The Petition Site - one of the largest platforms of its kind. We observe that the success of these petitions is driven by a number of factors, including promotion through social media channels and on the front page of the petitions platform. We propose an interpretable model that incorporates seasonality, aging effects, self-excitation, and external effects. The interpretability of the model is important for understanding the elements that drives the activity of an online content. We show through an extensive empirical evaluation that our model is significantly better at predicting the outcome of a petition than state-of-the-art techniques. Julia Proskurnia, Przemyslaw A. Grabowicz, Ryota Kobayashi, Carlos Castillo 0001, Philippe Cudré-Mauroux, Karl Aberer |
WWW | 2 |
| 2016 | The Road to Popularity: The Dilution of Growing Audience on Twitter
Przemyslaw A. Grabowicz, Mahmoudreza Babaei, Juhi Kulshrestha, Ingmar Weber |
ICWSM | 1 |
| 2016 | Distinguishing between Topical and Non-Topical Information Diffusion Mechanisms in Social Media
Przemyslaw A. Grabowicz, Niloy Ganguly, Krishna P. Gummadi |
ICWSM | 1 |
| 2016 | On the Efficiency of the Information Networks in Social MediaabstractSocial media sites are information marketplaces, where users produce and consume a wide variety of information and ideas. In these sites, users typically choose their information sources, which in turn determine what specific information they receive, how much information they receive and how quickly this information is shown to them. In this context, a natural question that arises is how efficient are social media users at selecting their information sources. In this work, we propose a computational framework to quantify users' efficiency at selecting information sources. Our framework is based on the assumption that the goal of users is to acquire a set of unique pieces of information. To quantify user's efficiency, we ask if the user could have acquired the same pieces of information from another set of sources more efficiently. We define three different notions of efficiency -- link, in-flow, and delay -- corresponding to the number of sources the user follows, the amount of (redundant) information she acquires and the delay with which she receives the information. Our definitions of efficiency are general and applicable to any social media system with an underlying in- formation network, in which every user follows others to receive the information they produce. Mahmoudreza Babaei, Przemyslaw A. Grabowicz, Isabel Valera, Krishna P. Gummadi, Manuel Gomez-Rodriguez |
WSDM | 2 |
| 2015 | On the Users' Efficiency in the Twitter Information Network
Mahmoudreza Babaei, Przemyslaw A. Grabowicz, Isabel Valera, Manuel Gomez-Rodriguez |
ICWSM | 2 |
| 2013 | Leveraging Browsing Patterns for Topic Discovery and Photostream Recommendation
Luca Chiarandini, Przemyslaw A. Grabowicz, Michele Trevisiol, Alejandro Jaimes |
ICWSM | 2 |
| 2013 | Distinguishing topical and social groups based on common identity and bond theoryabstractSocial groups play a crucial role in social media platforms because they form the basis for user participation and engagement. Groups are created explicitly by members of the community, but also form organically as members interact. Due to their importance, they have been studied widely (e.g., community detection, evolution, activity, etc.). One of the key questions for understanding how such groups evolve is whether there are different types of groups and how they differ. In Sociology, theories have been proposed to help explain how such groups form. In particular, the common identity and common bond theory states that people join groups based on identity (i.e., interest in the topics discussed) or bond attachment (i.e., social relationships). The theory has been applied qualitatively to small groups to classify them as either topical or social. We use the identity and bond theory to define a set of features to classify groups into those two categories. Using a dataset from Flickr, we extract user-defined groups and automatically-detected groups, obtained from a community detection algorithm. We discuss the process of manual labeling of groups into social or topical and present results of predicting the group label based on the defined features. We directly validate the predictions of the theory showing that the metrics are able to forecast the group type with high accuracy. In addition, we present a comparison between declared and detected groups along topicality and sociality dimensions. Przemyslaw A. Grabowicz, Luca Maria Aiello, Víctor M. Eguíluz, Alejandro Jaimes |
WSDM | 1 |