VLDB 2026 Research / reviewers in the wild / expert
Siyi Guo
dblp:271/1443
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0001-6422-3420ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Secure control of networked IT-2 fuzzy systems subject to multiple attacks: An observer-based elastic state-error-rate event-triggered scheme
Yuechao Ma, Siyi Guo |
Fuzzy Sets Syst. | 2 |
| 2025 | The Pulse of Mood Online: Unveiling Emotional Reactions in a Dynamic Social Media LandscapeabstractThe rich and dynamic information environment of social media provides researchers, policymakers, and entrepreneurs with opportunities to learn about social phenomena in a timely manner. However, using these data to understand social behavior is difficult due to the heterogeneity of topics and events discussed in the highly dynamic online information environment. To address these challenges, we present a method for systematically detecting and measuring emotional reactions to offline events using change point detection on the time series of collective affect and further explaining these reactions using a transformer-based topic model. We demonstrate the utility of the method by successfully detecting major and smaller events on three different datasets, including (1) a Los Angeles Tweet dataset between Jan. and Aug. 2020, in which we revealed the complex psychological impact of the BlackLivesMatter movement and the COVID-19 pandemic, (2) a dataset related to abortion rights discussions in the USA, in which we uncovered the strong emotional reactions to the overturn of Roe v. Wade and state abortion bans, and (3) a dataset about the 2022 French presidential election, in which we discovered the emotional and moral shift from positive before voting to fear and criticism after voting. We further demonstrate the importance of disaggregating data by topics and populations to mitigate potential biases when studying collective emotions. The capability of our method allows for better sensing and monitoring of the population’s reactions during crises using online data. Siyi Guo, Ashwin Rao, Fred Morstatter, P. Jeffrey Brantingham, Kristina Lerman |
ACM Trans. Web | 1 |
| 2024 | Community-Cross-Instruct: Unsupervised Instruction Generation for Aligning Large Language Models to Online CommunitiesabstractSocial scientists use surveys to probe the opinions and beliefs of populations, but these methods are slow, costly, and prone to biases.Recent advances in large language models (LLMs) enable the creation of computational representations or "digital twins" of populations that generate human-like responses mimicking the population's language, styles, and attitudes.We introduce COMMUNITY-CROSS-INSTRUCT, an unsupervised framework for aligning LLMs to online communities to elicit their beliefs.Given a corpus of a community's online discussions, COMMUNITY-CROSS-INSTRUCT automatically generates instruction-output pairs by an advanced LLM to ( 1) finetune a foundational LLM to faithfully represent that community, and (2) evaluate the alignment of the finetuned model to the community.We demonstrate the method's utility in accurately representing political and diet communities on Reddit.Unlike prior methods requiring human-authored instructions, COMMUNITY-CROSS-INSTRUCT generates instructions in a fully unsupervised manner, enhancing scalability and generalization across domains.This work enables costeffective and automated surveying of diverse online communities 1 . Minh Duc Chu, Rebecca Dorn, Siyi Guo, Kristina Lerman |
EMNLP | 4 |
| 2024 | Socio-Linguistic Characteristics of Coordinated Inauthentic AccountsabstractOnline manipulation is a pressing concern for democracies, but the actions and strategies of coordinated inauthentic accounts, which have been used to interfere in elections, are not well understood. We analyze a five million-tweet multilingual dataset related to the 2017 French presidential election, when a major information campaign led by Russia called "#MacronLeaks" took place. We utilize heuristics to identify coordinated inauthentic accounts and detect attitudes, concerns and emotions within their tweets, collectively known as socio-linguistic characteristics. We find that coordinated accounts retweet other coordinated accounts far more than expected by chance, while being exceptionally active just before the second round of voting. Concurrently, socio-linguistic characteristics reveal that coordinated accounts share tweets promoting a candidate at three times the rate of non-coordinated accounts. Coordinated account tactics also varied in time to reflect news events and rounds of voting. Our analysis highlights the utility of socio-linguistic characteristics to inform researchers about tactics of coordinated accounts and how these may feed into online social manipulation. Keith Burghardt, Ashwin Rao, Georgios Chochlakis, Sabyasachee Baruah, Siyi Guo, Andrew Rojecki, Shri Narayanan, Kristina Lerman |
ICWSM | 5 |
| 2024 | Discovering Collective Narratives Shifts in Online DiscussionsabstractNarratives are foundation of human cognition and decision making. Because narratives play a crucial role in societal discourses and spread of misinformation and because of the pervasive use of social media, the narrative dynamics on social media can have profound societal impact. Yet, systematic and computational understanding of online narratives faces critical challenge of the scale and dynamics; how can we reliably and automatically extract narratives from massive amount of texts? How do narratives emerge, spread, and die? Here, we propose a systematic narrative discovery framework that fill this gap by combining change point detection, semantic role labeling (SRL), and automatic aggregation of narrative fragments into narrative networks. We evaluate our model with synthetic and empirical data — two Twitter corpora about COVID-19 and 2017 French Election. Results demonstrate that our approach can recover major narrative shifts that correspond to the major events. Wanying Zhao, Siyi Guo, Kristina Lerman, Yong-Yeol Ahn |
ICWSM | 2 |
| 2023 | Measuring Online Emotional Reactions to EventsabstractThe rich and dynamic information environment of social media provides researchers, policy makers, and entrepreneurs with opportunities to learn about social phenomena in a timely manner. However, using this data to understand social behavior is difficult due heterogeneity of topics and events discussed in the highly dynamic online information environment. To address these challenges, we present a method for systematically detecting and measuring emotional reactions to offline events using change point detection on the time series of collective affect, and further explaining these reactions using a transformer-based topic model. We demonstrate the utility of the method on a corpus of tweets from a large US metropolitan area between January and August, 2020, covering a period of great social change. We demonstrate that our method is able to disaggregate topics to measure population's emotional and moral reactions. This capability allows for better monitoring of population's reactions during crises using online data. Siyi Guo, Ashwin Rao, Eugene Jang, Yuanfeixue Nan, Fred Morstatter, P. Jeffrey Brantingham, Kristina Lerman |
ASONAM | 1 |
| 2023 | Pandemic Culture Wars: Partisan Differences in the Moral Language of COVID-19 DiscussionsabstractEffective response to pandemics requires coordinated adoption of mitigation measures, like masking and quarantines, to curb a virus’s spread. However, as the COVID-19 pandemic demonstrated, political divisions can hinder consensus on the appropriate response. To better understand these divisions, our study examines a vast collection of COVID-19-related tweets. We focus on five contentious issues: coronavirus origins, lockdowns, masking, education, and vaccines. We describe a weakly supervised method to identify issue-relevant tweets and employ state-of-the-art computational methods to analyze moral language and infer political ideology. We explore how partisanship and moral language shape conversations about these issues. Our findings reveal ideological differences in issue salience and moral language used by different groups. We find that conservatives use more negatively-valenced moral language than liberals and that political elites use moral rhetoric to a greater extent than non-elites across most issues. Examining the evolution and moralization on divisive issues can provide valuable insights into the dynamics of COVID-19 discussions and assist policymakers in better understanding the emergence of ideological divisions. Ashwin Rao, Siyi Guo, Sze-Yuh Nina Wang, Fred Morstatter, Kristina Lerman |
IEEE Big Data | 2 |
| 2023 | A Data Fusion Framework for Multi-Domain Morality LearningabstractLanguage models can be trained to recognize the moral sentiment of text, creating new opportunities to study the role of morality in human life. As interest in language and morality has grown, several ground truth datasets with moral annotations have been released. However, these datasets vary in the method of data collection, domain, topics, instructions for annotators, etc. Simply aggregating such heterogeneous datasets during training can yield models that fail to generalize well. We describe a data fusion framework for training on multiple heterogeneous datasets that improve performance and generalizability. The model uses domain adversarial training to align the datasets in feature space and a weighted loss function to deal with label shift. We show that the proposed framework achieves state-of-the-art performance in different datasets compared to prior works in morality inference. Siyi Guo, Negar Mokhberian, Kristina Lerman |
ICWSM | 1 |
| 2023 | Observer-based event-triggered non-PDC control for networked T-S fuzzy systems under actuator failures and aperiodic DoS attacks
Siyi Guo, Yuechao Ma |
Inf. Sci. | 1 |
| 2022 | Learning Fairer InterventionsabstractExplicit and implicit bias clouds human judgment, leading to discriminatory treatment of disadvantaged groups. A fundamental goal of automated decisions is to avoid the pitfalls in human judgment by developing decision strategies that can be applied to all protected groups. Improving fairness of interventions via automated decision-inspired methods, however, has been under-utilized. In this paper, we propose a causal framework that learns optimal intervention policies from data subject to novel fairness constraints. We define two measures of treatment bias and infer treatment assignments that minimize the bias against protected groups while optimizing overall outcomes. We demonstrate the existence of trade-offs when balancing fairness and overall benefit; however, allowing preferential treatment of protected groups in certain circumstances (affirmative action) can dramatically improve the overall benefit while also preserving fairness. We apply our framework to data containing outcomes on standardized tests and show how it can be used to design real-world policies that fairly improve academic performance for different geographic areas. Our framework provides a principled way to learn fair treatment policies in real-world settings. Yuzi He, Keith Burghardt, Siyi Guo, Kristina Lerman |
AIES | 3 |