EDBT 2026 Demo / reviewers in the wild / expert
Fred Morstatter
dblp:51/9687
· DBLP profile ↗
29ranked-venue papers in the field
4as first author
13since 2021 · last 2026
0000-0002-0247-4328ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 16 (3 first)Information Retrieval & Web Search · 12 (1 first)Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On Causal and Anticausal LLM-based Data SynthesisabstractWhile Large Language Models (LLMs) have been increasingly used to generate synthetic data for various downstream tasks, researchers overlook the causal direction in the data synthesis process. A natural causal direction should contain two steps: diverse raw data are generated first, and subsequently annotated for downstream tasks. However, most LLM-based methods adopt an anticausal direction: embedding label information in the prompt to force LLMs to generate targeted data. This reversal raises a critical question: How does the direction of data synthesis impact the quality and utility of the synthetic data? In this work, we empirically study the impact of causal and anticausal data synthesis. To do so, we first design simple yet effective prompting strategies to control the causal direction of LLM-based data synthesis. Using GPT-5 as the data generator, we construct synthetic datasets for three distinct machine learning tasks. We then fine-tune BERT-base and LLaMA-3.2-1B models on these datasets and evaluate them against human-curated benchmarks. Our experiments reveal consistent patterns: (1) models trained on anticausal synthetic data suffer larger performance drops across all tasks and model families --- Accuracy declines range from 13.7%-59.1% for BERT and 4.9%-54.3% for LLaMA, and (2) distributional analysis shows that anticausal synthetic datasets deviate further from human data. Our findings provide practical guidance on how to generate better synthetic data and make good use of it. Bohan Jiang, Pingchuan Ma 0012, Zhuoyu Shi, Fred Morstatter, Adrienne Raglin, Huan Liu 0001 |
WSDM | 4 |
| 2025 | Simulating Hashtag Dynamics with Networked Groups of Generative Agents
Abha Jha, John Priniski, Carolyn Steinle, Fred Morstatter |
ASONAM (1) | 4 |
| 2025 | The Pulse of Mood Online: Unveiling Emotional Reactions in a Dynamic Social Media LandscapeabstractThe rich and dynamic information environment of social media provides researchers, policymakers, and entrepreneurs with opportunities to learn about social phenomena in a timely manner. However, using these data to understand social behavior is difficult due to the heterogeneity of topics and events discussed in the highly dynamic online information environment. To address these challenges, we present a method for systematically detecting and measuring emotional reactions to offline events using change point detection on the time series of collective affect and further explaining these reactions using a transformer-based topic model. We demonstrate the utility of the method by successfully detecting major and smaller events on three different datasets, including (1) a Los Angeles Tweet dataset between Jan. and Aug. 2020, in which we revealed the complex psychological impact of the BlackLivesMatter movement and the COVID-19 pandemic, (2) a dataset related to abortion rights discussions in the USA, in which we uncovered the strong emotional reactions to the overturn of Roe v. Wade and state abortion bans, and (3) a dataset about the 2022 French presidential election, in which we discovered the emotional and moral shift from positive before voting to fear and criticism after voting. We further demonstrate the importance of disaggregating data by topics and populations to mitigate potential biases when studying collective emotions. The capability of our method allows for better sensing and monitoring of the population’s reactions during crises using online data. Siyi Guo, Ashwin Rao, Fred Morstatter, P. Jeffrey Brantingham, Kristina Lerman |
ACM Trans. Web | 4 |
| 2024 | Non-binary Gender Expression in Online Interactions
Rebecca Dorn, Negar Mokhberian, Julie Jiang, Jeremy Abramson, Fred Morstatter, Kristina Lerman |
ASONAM (2) | 5 |
| 2024 | The Diffusion of Causal Language in Social NetworksabstractCausal reasoning plays a central role in human cognition. It facilitates the ability to infer, predict, and manipulate outcomes within the environment, which in turn lays the foundation for a uniquely adaptive decision-making framework that is crucial in navigating complex problem-solving contexts. With the pervasive influence of social media platforms, these online social networks have become critical for disseminating information, shaping public beliefs, and influencing daily life. However, no study has examined the propagation of causal language within social networks. In this work, we analyze the dispersion of messages containing causal language against those without, within the milieu of a large online social network. With the entirety of messages over one complete day on Twitter along with two additional days for validation, and with our validated ensemble method for identifying causal language, our findings reveal that messages with causal language exhibit a more extensive reach than those without. Furthermore, our counterfactual analysis demonstrates that the effect of causal language on information diffusion is truly causal. Moreover, our findings indicate that messages incorporating causal language manifest a higher ability to spread to out-groups compared to those without. These novel insights reveal the unique diffusion pattern of causal language within social networks, and suggest a potential to mitigate the echo chamber effect, while causal language could serve as a bridge for diverse perspectives. Zhuoyu Shi, Fred Morstatter |
ICWSM | 2 |
| 2023 | Measuring Online Emotional Reactions to EventsabstractThe rich and dynamic information environment of social media provides researchers, policy makers, and entrepreneurs with opportunities to learn about social phenomena in a timely manner. However, using this data to understand social behavior is difficult due heterogeneity of topics and events discussed in the highly dynamic online information environment. To address these challenges, we present a method for systematically detecting and measuring emotional reactions to offline events using change point detection on the time series of collective affect, and further explaining these reactions using a transformer-based topic model. We demonstrate the utility of the method on a corpus of tweets from a large US metropolitan area between January and August, 2020, covering a period of great social change. We demonstrate that our method is able to disaggregate topics to measure population's emotional and moral reactions. This capability allows for better monitoring of population's reactions during crises using online data. Siyi Guo, Ashwin Rao, Eugene Jang, Yuanfeixue Nan, Fred Morstatter, P. Jeffrey Brantingham, Kristina Lerman |
ASONAM | 6 |
| 2023 | Retweets Amplify the Echo Chamber EffectabstractThe growing prominence of social media in public discourse has led to a greater scrutiny of the quality of online information and the role it plays in amplifying political polarization. However, studies of polarization on social media platforms like Twitter have been hampered by the difficulty of collecting data about the social graph, specifically follow links that shape the echo chambers users join as well as what they see in their timelines. As a proxy of the follower graph, researchers use retweets, although it is not clear how this choice affects analysis. Using a sample of the Twitter follower graph and the tweets posted by users within it, we reconstruct the retweet graph and quantify its impact on the measures of echo chambers and exposure. While we find that echo chambers exist in both graphs, they are more pronounced in the retweet graph. We compare the information users see via their follower and retweet networks to show that retweeted accounts share systematically more polarized content. This bias cannot be explained by the activity or polarization within users' own follower graph neighborhoods but by the increased attention they pay to accounts that are ideologically aligned with their own views. Our results suggest that studies relying on the retweet graphs overestimate the echo chamber effects and exposure to polarized information. Ashwin Rao, Fred Morstatter, Kristina Lerman |
ASONAM | 2 |
| 2023 | Pandemic Culture Wars: Partisan Differences in the Moral Language of COVID-19 DiscussionsabstractEffective response to pandemics requires coordinated adoption of mitigation measures, like masking and quarantines, to curb a virus’s spread. However, as the COVID-19 pandemic demonstrated, political divisions can hinder consensus on the appropriate response. To better understand these divisions, our study examines a vast collection of COVID-19-related tweets. We focus on five contentious issues: coronavirus origins, lockdowns, masking, education, and vaccines. We describe a weakly supervised method to identify issue-relevant tweets and employ state-of-the-art computational methods to analyze moral language and infer political ideology. We explore how partisanship and moral language shape conversations about these issues. Our findings reveal ideological differences in issue salience and moral language used by different groups. We find that conservatives use more negatively-valenced moral language than liberals and that political elites use moral rhetoric to a greater extent than non-elites across most issues. Examining the evolution and moralization on divisive issues can provide valuable insights into the dynamics of COVID-19 discussions and assist policymakers in better understanding the emergence of ideological divisions. Ashwin Rao, Siyi Guo, Sze-Yuh Nina Wang, Fred Morstatter, Kristina Lerman |
IEEE Big Data | 4 |
| 2023 | Just Another Day on Twitter: A Complete 24 Hours of Twitter DataabstractAt the end of October 2022, Elon Musk concluded his acquisition of Twitter. In the weeks and months before that, several questions were publicly discussed that were not only of interest to the platform's future buyers, but also of high relevance to the Computational Social Science research community. For example, how many active users does the platform have? What percentage of accounts on the site are bots? And, what are the dominating topics and sub-topical spheres on the platform? In a globally coordinated effort of 80 scholars to shed light on these questions, and to offer a dataset that will equip other researchers to do the same, we have collected all 375 million tweets published within a 24-hour time period starting on September 21, 2022. To the best of our knowledge, this is the first complete 24-hour Twitter dataset that is available for the research community. With it, the present work aims to accomplish two goals. First, we seek to answer the aforementioned questions and provide descriptive metrics about Twitter that can serve as references for other researchers. Second, we create a baseline dataset for future research that can be used to study the potential impact of the platform's ownership change. Jürgen Pfeffer, Daniel Matter, Kokil Jaidka, Onur Varol, Afra J. Mashhadi, Jana Lasser, Dennis Assenmacher, Diyi Yang, Cornelia Brantner, Daniel M. Romero, Jahna Otterbacher, Carsten Schwemmer, Kenneth Joseph, David García 0001, Fred Morstatter |
ICWSM | 16 |
| 2022 | Noise Audits Improve Moral Foundation ClassificationabstractMorality plays an important role in culture, identity, and emotion. Recent advances in natural language processing have shown that it is possible to classify moral values expressed in text at scale. Morality classification relies on human annotators to label the moral expressions in text, which provides training data to achieve state-of-the-art performance. However, these annotations are inherently subjective and some of the instances are hard to classify, resulting in noisy annotations due to error or lack of agreement. The presence of noise in training data harms the classifier's ability to accurately recognize moral foundations from text. We propose two metrics to audit the noise of annotations. The first metric is entropy of instance labels, which is a proxy measure of annotator disagreement about how the instance should be labeled. The second metric is the silhouette coefficient of a label assigned by an annotator to an instance. This metric leverages the idea that instances with the same label should have similar latent representations, and deviations from collective judgments are indicative of errors. Our experiments on three widely used moral foundations datasets show that removing noisy annotations based on the proposed metrics improves classification performance.11Our code can be found at: https://github.com/negar-mokhberian/noise-audits. Negar Mokhberian, Frederic R. Hopp, Bahareh Harandizadeh, Fred Morstatter, Kristina Lerman |
ASONAM | 4 |
| 2022 | Samba: Identifying Inappropriate Videos for Young Children on YouTubeabstractYouTube videos are one of the most effective platforms for disseminating creative material and ideas, and they appeal to a diverse audience. Along with adults and older children, young children are avid consumers of YouTube materials. Children often lack means to evaluate if a given content is appropriate for their age, and parents have very limited options to enforce content restrictions on YouTube. Young children can thus become exposed to inappropriate content, such as violent, scary or disturbing videos on YouTube. Previous studies demonstrated that YouTube videos can be classified into appropriate or inappropriate for young viewers using video metadata, such as video thumbnails, title, comments, etc. Metadata-based approaches achieve high accuracy, but still have significant misclassifications, due to the reliability of input features. In this paper, we propose a fusion model, called Samba, which uses both metadata and video subtitles for content classification. Using subtitles in the model helps better infer the true nature of a video improving classification accuracy. On a large-scale, comprehensive dataset of 70K videos, we show that Samba achieves 95% accuracy, outperforming other state-of-the-art classifiers by at least 7%. We also publicly release our dataset. Le Binh, Rajat Tandon, Chingis Oinar, Jeffrey Liu, Uma Durairaj, Jiani Guo, Spencer Zahabizadeh, Sanjana Ilango, Jeremy Tang, Fred Morstatter, Simon S. Woo, Jelena Mirkovic |
CIKM | 10 |
| 2022 | Assessing Scientific Research Papers with Knowledge GraphsabstractIn recent decades, the growing scale of scientific research has led to numerous novel findings. Reproducing these findings is the foundation of future research. However, due to the complexity of experiments, manually assessing scientific research is laborious and time-intensive, especially in social and behavioral sciences. Although increasing reproducibility studies have garnered increased attention in the research community, there is still a lack of systematic ways for evaluating scientific research at scale. In this paper, we propose a novel approach towards automatically assessing scientific publications by constructing a knowledge graph (KG) that captures a holistic view of the research contributions. Specifically, during the KG construction, we combine information from two different perspectives: micro-level features that capture knowledge from published articles such as sample sizes, effect sizes, and experimental models, and macro-level features that comprise relationships between entities such as authorship and reference information. We then learn low-dimensional representations using language models and knowledge graph embeddings for entities (nodes in KGs), which are further used for the assessments. A comprehensive set of experiments on two benchmark datasets shows the usefulness of leveraging KGs for scoring scientific research. Kexuan Sun 0002, Zhiqiang Qiu, Abel Salinas, Yuzhong Huang, Daniel Benjamin, Fred Morstatter, Xiang Ren 0001, Kristina Lerman, Jay Pujara |
SIGIR | 7 |
| 2022 | Keyword Assisted Embedded Topic ModelabstractBy illuminating latent structures in a corpus of text, topic models are an essential tool for categorizing, summarizing, and exploring large collections of documents. Probabilistic topic models, such as latent Dirichlet allocation (LDA), describe how words in documents are generated via a set of latent distributions called topics. Recently, the Embedded Topic Model (ETM) has extended LDA to utilize the semantic information in word embeddings to derive semantically richer topics. As LDA and its extensions are unsupervised models, they aren't defined to make efficient use of a user's prior knowledge of the domain. To this end, we propose the Keyword Assisted Embedded Topic Model (KeyETM), which equips ETM with the ability to incorporate user knowledge in the form of informative topic-level priors over the vocabulary. Using both quantitative metrics and human responses on a topic intrusion task, we demonstrate that KeyETM produces better topics than other guided, generative models in the literature\footnoteCode for this work can be found at \urlhttps://github.com/bahareharandizade/KeyETM . Bahareh Harandizadeh, John Priniski, Fred Morstatter |
WSDM | 3 |
| 2020 | Aggressive, Repetitive, Intentional, Visible, and Imbalanced: Refining Representations for Cyberbullying Classification
Caleb Ziems, Ymir Vigfusson, Fred Morstatter |
ICWSM | 3 |
| 2019 | Debiasing community detection: the importance of lowly connected nodesabstractCommunity detection is an important task in social network analysis, allowing us to identify and understand the communities within the social structures provided by the network. However, many community detection approaches either fail to assign low-degree (or lowly connected) users to communities, or assign them to trivially small communities that prevent them from being included in analysis. In this work we investigate how excluding these users can bias analysis results. We then introduce an approach that is more inclusive for lowly connected users by incorporating them into larger groups. Experiments show that our approach outperforms the existing state-of-the-art in terms of F1 and Jaccard similarity scores while reducing the bias towards low-degree users. Ninareh Mehrabi, Fred Morstatter, Nanyun Peng 0001, Aram Galstyan |
ASONAM | 2 |
| 2019 | A Large-Scale Study of ISIS Social Media Strategy: Community Size, Collective Influence, and Behavioral Impact
Majid Alfifi, Parisa Kaghazgaran, James Caverlee, Fred Morstatter |
ICWSM | 4 |
| 2018 | Toward Relational Learning with MisinformationabstractRelational learning has been proposed to cope with the interdependency among linked instances in a network, and it is a fundamental tool to categorize social network users for various tasks. However, the emerging widespread of misinformation in social networks, information that is inaccurate or false, poses novel challenges to utilizing social media data. Malicious users may actively manipulate their content and characteristics, which easily lead to a noisy dataset. Hence, it is intricate for traditional relational learning approaches to deliver an accurate predictive model in the presence of misinformation. In this work, we precisely focus on the problem by proposing a joint framework that simultaneously constructs a relational learning model and mitigates the effect of misinformation by restraining anomalous points. Empirical results on real-world social media data prove the superiority of the proposed approach, Relational Learning with Misinformation (RLM), over traditional approaches on modeling social network users. Liang Wu 0006, Jundong Li, Fred Morstatter, Huan Liu 0001 |
SDM | 3 |
| 2017 | Adaptive Spammer Detection with Sparse Group Modeling
Liang Wu 0006, Xia Ben Hu, Fred Morstatter, Huan Liu 0001 |
ICWSM | 3 |
| 2017 | Detecting Camouflaged Content Polluters
Liang Wu 0006, Xia Ben Hu, Fred Morstatter, Huan Liu 0001 |
ICWSM | 3 |
| 2016 | Detecting and mitigating bias in social mediaabstractSocial media is an important data source. Every day, billions of posts, likes, and connections are created by people around the globe. By monitoring it we can observe important topics, as well as find new topics of discussion as they emerge. However, within this source of information there are natural forms of bias. Different aspects of the sites lend themselves to bias, such as varying features that restrict users. Additionally, the users themselves can be biased, such as the age-bias found in Twitter users. Finally, the way sites divulge their data can cause bias to those studying information produced on that site. In this forum we will discuss the different types of bias that can occur on social media data as well as different strategies to mitigate that bias. Fred Morstatter |
ASONAM | 1 |
| 2016 | A new approach to bot detection: Striking the balance between precision and recallabstractThe presence of bots has been felt in many aspects of social media. Twitter, one example of social media, has especially felt the impact, with bots accounting for a large portion of its users. These bots have been used for malicious tasks such as spreading false information about political candidates and inflating the perceived popularity of celebrities. Furthermore, these bots can change the results of common analyses performed on social media. It is important that researchers and practitioners have tools in their arsenal to remove them. Approaches exist to remove bots, however they focus on precision to evaluate their model at the cost of recall. This means that while these approaches are almost always correct in the bots they delete, they ultimately delete very few, thus many bots remain. We propose a model which increases the recall in detecting bots, allowing a researcher to delete more bots. We evaluate our model on two real-world social media datasets and show that our detection algorithm removes more bots from a dataset than current approaches. Fred Morstatter, Liang Wu 0006, Tahora H. Nazer, Kathleen M. Carley, Huan Liu 0001 |
ASONAM | 1 |
| 2016 | Finding requests in social media for disaster reliefabstractNatural disasters create an uncertain environment in which first responders face the challenge of locating affected people and dispatching aids and resources in a timely manner. In recent years, crowdsourcing systems have been developed to exploit the power of volunteers to facilitate humanitarian logistic efforts. Most of the current systems require volunteers to directly provide input to them and do not have the capability to benefit the large number of disaster-related posts that are published on social media. Hence, many social media posts in the aftermath of disasters remain hidden. Among these hidden posts are those that need immediate attention, such as requests for help. Hence, we have implemented a system that detects requests on Twitter using content and context of tweets. Tahora H. Nazer, Fred Morstatter, Harsh Dani, Huan Liu 0001 |
ASONAM | 2 |
| 2016 | Leveraging the Implicit Structure within Social Media for Emergent Rumor DetectionabstractThe automatic and early detection of rumors is of paramount importance as the spread of information with questionable veracity can have devastating consequences. This became starkly apparent when, in early 2013, a compromised Associated Press account issued a tweet claiming that there had been an explosion at the White House. This tweet resulted in a significant drop for the Dow Jones Industrial Average. Most existing work in rumor detection leverages conversation statistics and propagation patterns, however, such patterns tend to emerge slowly requiring a conversation to have a significant number of interactions in order to become eligible for classification. In this work, we propose a method for classifying conversations within their formative stages as well as improving accuracy within mature conversations through the discovery of implicit linkages between conversation fragments. In our experiments, we show that current state-of-the-art rumor classification methods can leverage implicit links to significantly improve the ability to properly classify emergent conversations when very little conversation data is available. Adopting this technique allows rumor detection methods to continue to provide a high degree of classification accuracy on emergent conversations with as few as a single tweet. This improvement virtually eliminates the delay of conversation growth inherent in current rumor classification methods while significantly increasing the number of conversations considered viable for classification. Justin Sampson, Fred Morstatter, Liang Wu 0006, Huan Liu 0001 |
CIKM | 2 |
| 2016 | Paired Restricted Boltzmann Machine for Linked DataabstractRestricted Boltzmann Machines (RBMs) are widely adopted unsupervised representation learning methods and have powered many data mining tasks such as collaborative filtering and document representation. Recently, linked data that contains both attribute and link information has become ubiquitous in various domains. For example, social media data is inherently linked via social relations and web data is networked via hyperlinks. It is evident from recent work that link information can enhance a number of real-world applications such as clustering and recommendations. Therefore, link information has the potential to advance RBMs for better representation learning. However, the majority of existing RBMs have been designed for independent and identically distributed data and are unequipped for linked data. In this paper, we aim to design a new type of Restricted Boltzmann Machines that takes advantage of linked data. In particular, we propose a paired Restricted Boltzmann Machine (pRBM), which is able to leverage the attribute and link information of linked data for representation learning. Experimental results on real-world datasets demonstrate the effectiveness of the proposed framework pRBM. Suhang Wang, Jiliang Tang, Fred Morstatter, Huan Liu 0001 |
CIKM | 3 |
| 2013 | Near real time assessment of social media using geo-temporal network analyticsabstractWhen a crisis occurs, there is often little time to evaluate the situation and determine how best to respond. We use rapid ethnographic methods centered on the construction of geo-temporally contextualized social and knowledge networks. By utilizing a combination of Twitter and news media, the consulate attack in Libya were examined in near real time. In this work we outline a procedure to extract key insights from the event as an event unfolds using a suite of tools developed by a team of researchers from two universities. Kathleen M. Carley, Jürgen Pfeffer, Huan Liu 0001, Fred Morstatter, Rebecca Goolsby |
ASONAM | 4 |
| 2013 | Is the Sample Good Enough? Comparing Data from Twitter's Streaming API with Twitter's Firehose
Fred Morstatter, Jürgen Pfeffer, Huan Liu 0001, Kathleen M. Carley |
ICWSM | 1 |
| 2013 | Understanding Twitter data with TweetXplorerabstractIn the era of big data it is increasingly difficult for an analyst to extract meaningful knowledge from a sea of information. We present TweetXplorer, a system for analysts with little information about an event to gain knowledge through the use of effective visualization techniques. Using tweets collected during Hurricane Sandy as an example, we will lead the reader through a workflow that exhibits the functionality of the system. Fred Morstatter, Shamanth Kumar, Huan Liu 0001, Ross Maciejewski |
KDD | 1 |
| 2012 | Navigating information facets on twitter (NIF-T)abstractRecent years have seen an exponential increase in the number of users of social media sites. As the number of users of these sites continues to grow at an extraordinary rate, the amount of data produced follows in magnitude. With this deluge of social media data, the need for comprehensive tools to analyze user interactions is ever increasing. In this paper, we present a novel tool, Navigating Information Facets on Twitter (NIF-T), which helps users to explore data generated on social media sites. Using the three dimensions or facets: time, location, and topic as an example of the many possible facets, we enable the users to explore large social media datasets. With the help of a large corpus of tweets collected from the Occupy Wall Street movement on the Twitter platform we show how our system can be used to identify important aspects of the event along these facets. Shamanth Kumar, Fred Morstatter, Grant Marshall, Huan Liu 0001, Ullas Nambiar |
KDD | 2 |
| 2011 | Feature Selection Strategy in Text Classification
Gabriel Pui Cheong Fung, Fred Morstatter, Huan Liu 0001 |
PAKDD (1) | 2 |