VLDB 2026 Research / reviewers in the wild / expert
Jürgen Pfeffer
dblp:25/11207
· DBLP profile ↗
32ranked-venue papers in the field
3as first author
17since 2021 · last 2026
0000-0002-1677-150XORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 17 (3 first)Data Mining & Knowledge Discovery · 13Database Systems & Data Management · 1Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Quotes to Concepts: Axial Coding of Political Debates with Ensemble LMs
Angelina Parfenova, David Graus, Jürgen Pfeffer |
ECIR (2) | 3 |
| 2025 | Prevalence, Substance and Responses to Hate Speech Against LGBTQ Communities on TikTokabstractDespite ongoing efforts, online hate speech remains a pervasive issue on social media, particularly affecting vulnerable groups such as LGBTQ communities. While there is extensive debate around how best to address this problem, counter speech is emerging as a promising solution. However, existing research has primarily focused on detecting hateful content, often overlooking broader aspects such as the specific topics of discrimination and the spread of countermeasures online. This study examines the prevalence of hate speech and counter speech in LGBTQ online spaces on TikTok, analysing day-to-day interactions to identify recurring themes and targets. Results reveal that hate speech is widespread: at least 3.5% of messages contain hateful content, spread by approximately 4% of users, and one in three videos attracts hate comments or replies, primarily targeting LGBTQ topics explicitly. Gender identity emerges as a major focus, with transgender and non-binary individuals being frequent targets. Although much hate engagement goes unanswered, when responses occur, they are often in the form of counter speech, especially when LGBTQ-related topics are targeted. These findings improve our understanding of the nature and extent of online hate speech against LGBTQ communities, confirm counter speech as an employed response, and provide a foundation for further research aimed at developing strategies to promote safer, more inclusive social media environments. Jordi Guillem Condom Tibau, Angelina Voggenreiter, Elena Pavan, Jürgen Pfeffer |
ICWSM | 4 |
| 2025 | Reddit Rehab: User Migration in Response to Mobile Client ShutdownsabstractThis paper investigates the behavior of Reddit users who relied on alternative mobile apps, such as Apollo and RiF, before and after their forced shutdown by Reddit on July 1, 2023. The announcement of the shutdown led many observers to predict significant negative consequences, such as mass migration away from the platform. Using data from January to November 2023, we analyze user engagement and migration rates for users of these alternative clients before and after the forced discontinuation of their apps. We find that 22% of alternative client users permanently left Reddit as a result, and 45% of the users who openly threatened to leave if the changes were enacted followed through with their threats. Overall, we find that the shutdown of third-party apps had no discernible impact on overall platform activity. While the preceding protests were severe, ultimately for most users the cost of switching to the official client was likely far less than the effort required to switch to an entirely different platform. Scientific attention has increased to understand the contributing factors and effects of migration between online platforms, but real-world examples with available data remain rare. Our study addresses this by examining a large-scale online migratory movement. Franz Waltenberger, Angelina Voggenreiter, Martin Wessel, Jürgen Pfeffer |
ICWSM | 4 |
| 2024 | Dynamic Inter-organizational Communication Network in a Post-merger Integration
Michael Benzinger, Raji Ghawi, Lukas Zenk, Jürgen Pfeffer |
ASONAM (1) | 4 |
| 2024 | From Retweets to Follows: Facilitating Graph Construction in Online Social Networks Through Machine Learning
Anahit Sargsyan, Jürgen Pfeffer |
ASONAM (2) | 2 |
| 2023 | Just Another Day on Twitter: A Complete 24 Hours of Twitter DataabstractAt the end of October 2022, Elon Musk concluded his acquisition of Twitter. In the weeks and months before that, several questions were publicly discussed that were not only of interest to the platform's future buyers, but also of high relevance to the Computational Social Science research community. For example, how many active users does the platform have? What percentage of accounts on the site are bots? And, what are the dominating topics and sub-topical spheres on the platform? In a globally coordinated effort of 80 scholars to shed light on these questions, and to offer a dataset that will equip other researchers to do the same, we have collected all 375 million tweets published within a 24-hour time period starting on September 21, 2022. To the best of our knowledge, this is the first complete 24-hour Twitter dataset that is available for the research community. With it, the present work aims to accomplish two goals. First, we seek to answer the aforementioned questions and provide descriptive metrics about Twitter that can serve as references for other researchers. Second, we create a baseline dataset for future research that can be used to study the potential impact of the platform's ownership change. Jürgen Pfeffer, Daniel Matter, Kokil Jaidka, Onur Varol, Afra J. Mashhadi, Jana Lasser, Dennis Assenmacher, Diyi Yang, Cornelia Brantner, Daniel M. Romero, Jahna Otterbacher, Carsten Schwemmer, Kenneth Joseph, David García 0001, Fred Morstatter |
ICWSM | 1 |
| 2023 | This Sample Seems to Be Good Enough! Assessing Coverage and Temporal Reliability of Twitter's Academic APIabstractBecause of its willingness to share data with academia and industry, Twitter has been the primary social media platform for scientific research as well as for consulting businesses and governments in the last decade. In recent years, a series of publications have studied and criticized Twitter's APIs and Twitter has partially adapted its existing data streams. The newest Twitter API for Academic Research allows to "access Twitter's real-time and historical public data with additional features and functionality that support collecting more precise, complete, and unbiased datasets. The main new feature of this API is the possibility of accessing the full archive of all historic Tweets. In this article, we will take a closer look at the Academic API and will try to answer two questions. First, are the datasets collected with the Academic API complete? Secondly, since Twitter's Academic API delivers historic Tweets as represented on Twitter at the time of data collection, we need to understand how much data is lost over time due to Tweet and account removal from the platform. Our work shows evidence that Twitter's Academic API can indeed create (almost) complete samples of Twitter data based on a wide variety of search terms. We also provide evidence that Twitter's data endpoint v2 delivers better samples than the previously used endpoint v1.1. Furthermore, collecting Tweets with the Academic API at the time of studying a phenomenon rather than creating local archives of stored Tweets, allows for a straightforward way of following Twitter's developer agreement. Finally, we will also discuss technical artifacts and implications of the Academic API. We hope that our work can add another layer of understanding of Twitter data collections leading to more reliable studies of human behavior via social media data. Jürgen Pfeffer, Angelina Voggenreiter, Jana Lasser, Luca Hammer, Oliver Stritzel, David García 0001 |
ICWSM | 1 |
| 2023 | The Half-Life of a TweetabstractTwitter has started to share an impression count variable as part of the available public metrics for every Tweet collected with Twitter’s APIs. With the information about how often a particular Tweet has been shown to Twitter users at the time of data collection, we can learn important insights about the dissemination process of a Tweet by measuring its impression count repeatedly over time. With our preliminary analysis, we can show that on average the peak of impressions per second is 72 seconds after a Tweet was sent and that after 24 hours, no relevant number of impressions can be observed for ∼95% of all Tweets. Finally, we estimate that the median half-life of a Tweet, i.e. the time it takes before half of all impressions are created, is about 80 minutes. Jürgen Pfeffer, Daniel Matter, Anahit Sargsyan |
ICWSM | 1 |
| 2022 | Identifying Power Elites in Massively Multiplayer Online Games by Applying Machine Learning to Communication and Support NetworksabstractThe aim of this paper is to show how machine learning can predict whether an individual is more powerful than others in the group. The crucial point here is to consider the structural position of the actors in the social networks in which they are embedded. The approach we have taken for constructing these intra-group networks is the aggregation of communication and support interactions. Our research is based on longitutional data from the Massively Multiplayer Online Game (MMOG) Travian that was collected over a 12-month period. The data includes 202,764 communication and 96,913 support interactions between players that we applied for the construction of interaction networks. We also had access to status information on a daily basis for 21,431 individual players who were members of 4,758 alliances. Methodically, we applied 10 established metrics from SNA-based team research in combination with the Random Forstest classification algorithm. Our results show that interaction networks are well suited to assign members into two groups of powerful (elite) and nonpowerful (non-elite) players. It turned out that the identification of non-elite members was much easier to accomplish than that of elite members. Regarding the application of multiplex networks, we could not confirm a higher explanatory power by using combined networks. In summary, we can say that the network patterns of elite members are clearly different from those of non-elite members. In this way, we were able to predict affiliation to each category with an accuracy (F1) of 0.88 for communication networks and 0.83 for support networks. Siegfried Müller, Raji Ghawi, Jürgen Pfeffer |
ASONAM | 3 |
| 2022 | Gender dynamics of German journalists on TwitterabstractWomen are underrepresented in many areas of journalistic newsrooms. In this paper, we examine if this es-tablished effect is continued in the new forms of journalistic communication, Social Media Networks. We used mentions and retweets as measures of journalistic amplification and legitimation. Furthermore, we compared two groups of journalists in different stages of development: political and data journalists in Germany in 2021. Our results show that journalists regarded as women tend to favor their peers in mentions and retweets on Twitter: while both professions are dominated by a massive number of men and a high share of men-authored tweets, females mentioned and retweeted other women to a more extensive degree than their male colleagues. In addition, we have found data journalists to be more inclusive towards non-members in their network compared to political journalists. Benedict Witzenberger, Jürgen Pfeffer |
ASONAM | 2 |
| 2022 | Glowing Experience or Bad Trip? A Quantitative Analysis of User Reported Drug Experiences on Erowid.org
Angelina Voggenreiter, Momin M. Malik, Hemank Lamba, Earth Erowid, Sylvia Thyssen, Jürgen Pfeffer |
ICWSM | 6 |
| 2022 | Analysis of Country Mentions in the Debates of the UN Security Council
Raji Ghawi, Jürgen Pfeffer |
iiWAS | 2 |
| 2022 | Discovering Relational Implications in Multilayer Networks Using Formal Concept Analysis
Raji Ghawi, Jürgen Pfeffer |
iiWAS | 2 |
| 2022 | Central Figures in the Climate Change Discussion on Twitter
Anil Can Kara, Ivana Dobrijevic, Emre Öztas, Angelina Voggenreiter, Raji Ghawi, Jürgen Pfeffer |
iiWAS | 6 |
| 2021 | ContextWalk: Embedding Networks with Context Information Extracted from News Articles
Mirco Schönfeld, Jürgen Pfeffer |
DEXA (2) | 3 |
| 2021 | A Hybrid Thresholding Strategy combining RCut and PCut for Multi-label ClassificationabstractMulti-label classification is a variant of the classification problem where multiple labels may be assigned to each instance. Usually multi-label classification algorithms output a numerical score for each label, indicative of their relevance to a query instance. However, in many applications the desired output is a bipartition of the labels into relevant and irrelevant w.r.t the query instance. Bipartitions can be obtained from scores using various thresholding strategies, such as PCut strategy which selects relevant instances per label, and RCut strategy which selects relevant labels per instance. However, we suggest that a combination of both strategies would provide better classification performance. In this paper, we propose a fuzzy-based approach to combine PCut and RCut strategies, by converting the crisp relevance into fuzzy one, merging them linearly, and defuzzifying again. Our experiments shows that our hybrid approach indeed outperforms both strategies. Raji Ghawi, Jürgen Pfeffer |
iiWAS | 2 |
| 2021 | What we Talk about when we Talk about Earth on Earth Day?abstractApril 22, 2021, marked the 51st anniversary of Earth Day. With the growing imperativeness of environmental protection and sustainability, we want to study people’s collective attention and conversations on this themed day. What are the top-of-mind discourses and central topics about the earth? How do people feel about them, hopeful or pessimistic? How do they change over time, especially after the COVID-19 pandemic? To answer these, we extracted and quantified top frequent features, co-occurring hashtags, sentiment words, and latent sub-topics from about 300K tweets posted on the Earth Day of 2009, 2013, 2017, and 2021. The results demonstrated the longitudinal dynamics of people’s rhetoric and focus regarding protecting the earth – from resources conservation to climate changes, as well as the plummeted optimism toward environmental topics after the pandemic. The findings of our paper can help decision-makers to better assess the “voices of the people” and inform evidence-based decision-making. Raji Ghawi, Jürgen Pfeffer |
iiWAS | 3 |
| 2020 | Using Communication Networks to Predict Team Performance in Massively Multiplayer Online GamesabstractVirtual teams are becoming increasingly important. Since they are digital in nature, their “trace data” enable a broad set of new research opportunities. Online Games are especially useful for studying social behavior patterns of collaborative teams. In our study we used longitudinal data from the Massively Multiplayer Online Game (MMOG) Travian collected over a 12-month period that included 4,753 teams with 18,056 individuals and their communication networks. For predicting team performance, we selected 13 SNA-based attributes frequently used in team and leadership research. Using machine learning algorithms, the added explanatory power derived from the patterns of the communication networks enabled us to achieve an adjusted R2 = 0.67 in the best fitting performance prediction model and a prediction accuracy of up to 95.3% in the classification of top performing teams. Siegfried Müller, Raji Ghawi, Jürgen Pfeffer |
ASONAM | 3 |
| 2020 | A Longitudinal Analysis of a Social Network of Intellectual HistoryabstractThe history of intellectuals consists of a complex web of influences and interconnections of philosophers, scientists, writers, their work, and ideas. How did these influences evolve over time? Who were the most influential scholars in a period? To answer these questions, we mined a network of influence of over 12,500 intellectuals, extracted from the Linked Open Data provider YAGO. We enriched this network with a longitudinal perspective and analyzed time-sliced projections of the complete network differentiating between within-era, inter-era, and accumulated-era networks. We thus identified various patterns of intellectuals and eras and studied their development in time. We show which scholars were most influential in different eras, and who took prominent knowledge broker roles. One essential finding is that the highest impact of an era's scholar was on their contemporaries, and that the inter-era influence of each period was strongest on the consecutive era. Furthermore, we see quantitative evidence that there was no rediscovery of Antiquity during the Renaissance; rather, there has been a continuous reception of it since the Middle Ages. Cindarella Petz, Raji Ghawi, Jürgen Pfeffer |
ASONAM | 3 |
| 2020 | Against the Others! Detecting Moral Outrage in Social Media NetworksabstractOnline firestorms on Twitter are seemingly arbitrarily occurring outrages towards people, companies, media campaigns and politicians. Moral outrage can create an excessive collective aggressiveness against one single argument, one single word, or one action of a person resulting in hateful speech. With a collective “against the others” the negative dynamics often start. Using data from Twitter, we explored the starting points of several firestorm outbreaks. As a social media platform with hundreds of millions of users interacting in real-time on topics and events all over the world, Twitter serves as a social sensor for online discussions and is known for quick and often emotional disputes. The main question we pose in this article is whether we can detect the outbreak of a firestorm. Given 21 online firestorms on Twitter, the key questions regarding the anomaly detection are: 1) How can we detect changing points? 2) How can we distinguish the features that indicate a moral outrage? In this paper we examine these challenges developing a method to detect the point of change systematically spotting on linguistic cues of tweets. We are able to detect outbreaks of firestorms early and precisely only by applying linguistic cues. The results of our work can help detect negative dynamics and may have the potential for individuals, companies, and governments to mitigate hate in social media networks. Wienke Strathern, Mirco Schönfeld, Raji Ghawi, Jürgen Pfeffer |
ASONAM | 4 |
| 2019 | Movie Genres Classification using Collaborative FilteringabstractIn this paper, we present an approach for classifying movie genres based on user-ratings. Our approach is based on collaborative filtering (CF), a common technique used in recommendation systems, where the similarity between movies based on user-ratings, is used to predict the genres of movies. The results of conducted experiments show that our genres classification approach outperforms many existing approaches, by achieving an F1-score of 0.70, and a hit-rate of 94%. Raji Ghawi, Jürgen Pfeffer |
iiWAS | 2 |
| 2019 | Extracting Ego-Centric Social Networks from Linked Open DataabstractLinked Open Data (LOD) refers to freely available data on the WWW that are typically represented using Resource Description Framework (RDF). LOD is an invaluable source of rich and structured information, and enables a wide range of new applications, such as Social Network Analysis (SNA). In this paper, we address the extraction of social networks from LOD using SPARQL language, and we present various patterns to extract ego-centric networks. We also present two case studies: i) influence networks of intellectuals, and ii) co-acting networks, to demonstrate the applicability and usefulness of the approach. Raji Ghawi, Mirco Schönfeld, Jürgen Pfeffer |
WI | 3 |
| 2017 | Armed Conflicts in Online News: A Multilingual Study
Robert West 0001, Jürgen Pfeffer |
ICWSM | 2 |
| 2017 | Cost Matters: A New Example-Dependent Cost-Sensitive Logistic Regression Model
Nikou Günnemann, Jürgen Pfeffer |
PAKDD (1) | 2 |
| 2017 | zooRank: Ranking Suspicious Entities in Time-Evolving Tensors
Hemank Lamba, Bryan Hooi, Kijung Shin, Christos Faloutsos, Jürgen Pfeffer |
ECML/PKDD (1) | 5 |
| 2017 | Sampling from Social Networks with AttributesabstractSampling from large networks represents a fundamental challenge for social network research. In this paper, we explore the sensitivity of different sampling techniques (node sampling, edge sampling, random walk sampling, and snowball sampling) on social networks with attributes. We consider the special case of networks (i) where we have one attribute with two values (e.g., male and female in the case of gender), (ii) where the size of the two groups is unequal (e.g., a male majority and a female minority), and (iii) where nodes with the same or different attribute value attract or repel each other (i.e., homophilic or heterophilic behavior). We evaluate the different sampling techniques with respect to conserving the position of nodes and the visibility of groups in such networks. Experiments are conducted both on synthetic and empirical social networks. Our results provide evidence that different network sampling techniques are highly sensitive with regard to capturing the expected centrality of nodes, and that their accuracy depends on relative group size differences and on the level of homophily that can be observed in the network. We conclude that uninformed sampling from social networks with attributes thus can significantly impair the ability of researchers to draw valid conclusions about the centrality of nodes and the visibility or invisibility of groups in social networks. Claudia Wagner 0001, Philipp Singer, Fariba Karimi 0001, Jürgen Pfeffer, Markus Strohmaier |
WWW | 4 |
| 2016 | Identifying Platform Effects in Social Media Data
Momin M. Malik, Jürgen Pfeffer |
ICWSM | 2 |
| 2015 | Finding Non-Redundant Multi-Word Events on TwitterabstractTwitter is a pervasive technology, with hundreds of millions of users serving as sensors that provide eyewitness accounts of events on the ground. In case of popular events, these sensors start to broadcast news by tweeting to their followers, and to the world. Within minutes these tweets can attract attention and also serve as a primary information source for traditional media. Given a huge set of tweets, the key questions are: (1) How can we detect informative events in general? (2) How can we distinguish relevant events from others? In this paper we tackle these challenges with a statistical model for detecting events by spotting significant frequency deviations of the words' frequency over time. Besides single word events, our model also accounts for events composed of multiple co-occurring words, thus, providing much richer information. Our statistical process is complemented with an optimization algorithm to extract only non-redundant events, overall, providing the user with a succinct summary of the current events. We used our model to analyze 24 million geotagged tweets that have been sent in the US from April 9 to April 22, 2013 -- the time period of the Boston marathon bombing -- and we show that our approach can create multi-word events that efficiently summarize real-world events. Nikou Günnemann, Jürgen Pfeffer |
ASONAM | 2 |
| 2015 | A Tempest in a Teacup? Analyzing Firestorms on Twitterabstract'Firestorms,' sudden bursts of negative attention in cases of controversy and outrage, are seemingly widespread on Twitter and are an increasing source of fascination and anxiety in the corporate, governmental, and public spheres. Using media mentions, we collect 80 candidate events from January 2011 to September 2014 that we would term 'firestorms.' Using data from the Twitter decahose (or gardenhose), a 10% random sample of all tweets, we describe the size and longevity of these firestorms. We take two firestorm exemplars, #myNYPD and #CancelColbert, as case studies to describe more fully. Then, taking the 20 firestorms with the most tweets, we look at the change in mention networks of participants over the course of the firestorm as one method of testing for possible impacts of firestorms. We find that the mention networks before and after the firestorms are more similar to each other than to those of the firestorms, suggesting that firestorms neither emerge from existing networks, nor do they result in lasting changes to social structure. To verify this, we randomly sample users and generate mention networks for baseline comparison, and find that the firestorms are not associated with a greater than random amount of change in mention networks. Hemank Lamba, Momin M. Malik, Jürgen Pfeffer |
ASONAM | 3 |
| 2013 | Near real time assessment of social media using geo-temporal network analyticsabstractWhen a crisis occurs, there is often little time to evaluate the situation and determine how best to respond. We use rapid ethnographic methods centered on the construction of geo-temporally contextualized social and knowledge networks. By utilizing a combination of Twitter and news media, the consulate attack in Libya were examined in near real time. In this work we outline a procedure to extract key insights from the event as an event unfolds using a suite of tools developed by a team of researchers from two universities. Kathleen M. Carley, Jürgen Pfeffer, Huan Liu 0001, Fred Morstatter, Rebecca Goolsby |
ASONAM | 2 |
| 2013 | Is the Sample Good Enough? Comparing Data from Twitter's Streaming API with Twitter's Firehose
Fred Morstatter, Jürgen Pfeffer, Huan Liu 0001, Kathleen M. Carley |
ICWSM | 2 |
| 2012 | Visual Analysis of Dynamic Networks Using Change CentralityabstractThe visualization and analysis of dynamic social networks are challenging problems, demanding the simultaneous consideration of relational and temporal aspects. In order to follow the evolution of a network over time, we need to detect not only which nodes and which links change and when these changes occur, but also the impact they have on their neighbourhood and on the overall relational structure. Aiming to enhance the perception of structural changes at both the micro and the macro level, we introduce the change centrality metric. This novel metric, as well as a set of further metrics we derive from it, enable the pair wise comparison of subsequent states of an evolving network in a discrete-time domain. Demonstrating their exploitation to enrich visualizations, we show how these change metrics support the visual analysis of network dynamics. Paolo Federico 0001, Jürgen Pfeffer, Wolfgang Aigner, Silvia Miksch, Lukas Zenk |
ASONAM | 2 |