EDBT 2026 Demo / reviewers in the wild / expert
Kristina Lerman
dblp:99/433
· DBLP profile ↗
70ranked-venue papers in the field
9as first author
23since 2021 · last 2026
0000-0002-5071-0575ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 41 (5 first)Data Mining & Knowledge Discovery · 20 (2 first)Database Systems & Data Management · 4 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 4 (1 first)Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cyberpsychology: Emotion as the Hidden Driver of Social Behavior in Online Networks
Kristina Lerman |
WSDM | 1 |
| 2026 | Emergence of Structural Disparities in the Web of Scientific CitationsabstractScientific attention is unevenly distributed, creating inequities in recognition and distorting access to opportunities. Using citations as a proxy, we quantify disparities in attention by gender and institutional prestige. We find that women receive systematically fewer citations than men, and that attention is increasingly concentrated among authors from elite institutions -- patterns not fully explained by underrepresentation alone. To explain these dynamics, we introduce a model of citation network growth that incorporates homophily (tendency to cite similar authors), preferential attachment (favoring highly cited authors) and group size (underrepresentation). The model shows that disparities arise not only from group size imbalances but also from cumulative advantage amplifying biased citation preferences. Importantly, increasing representation alone is often insufficient to reduce disparities. Effective strategies should also include reducing homophily, amplifying the visibility of underrepresented groups, and supporting equitable integration of newcomers. Our findings highlight the challenges of mitigating inequities in asymmetric networks like citations, where recognition flows in one direction. By making visible the mechanisms through which attention is distributed, we contribute to efforts toward a more responsible web of science that is fairer, more transparent, and more inclusive, and that better sustains innovation and knowledge production. Buddhika Nettasinghe, Nazanin Alipourfard, Vikram Krishnamurthy, Kristina Lerman |
WWW | 4 |
| 2025 | Modeling Information Narrative Evolution on Telegram During the Russia-Ukraine WarabstractFollowing the Russian Federation's full-scale invasion of Ukraine in February 2022, a multitude of information narratives emerged within both pro-Russian and pro-Ukrainian communities online. As the conflict progresses, so too do the information narratives, constantly adapting and influencing local and global community perceptions and attitudes. This dynamic nature of the evolving information environment (IE) underscores a critical need to fully discern how narratives evolve and affect online communities. Existing research, however, often fails to capture information narrative evolution, overlooking both the fluid nature of narratives and the internal mechanisms that drive their evolution. Recognizing this, we introduce a novel approach designed to both model narrative evolution and uncover the underlying mechanisms driving them. In this work we perform a comparative discourse analysis across communities on Telegram covering the initial three months following the invasion. First, we uncover substantial disparities in narratives and perceptions between pro-Russian and pro-Ukrainian communities. Then, we probe deeper into prevalent narratives of each group, identifying key themes and examining the underlying mechanisms fueling their evolution. Finally, we explore influences and factors that may shape the development and spread of narratives. Patrick Gérard, Svitlana Volkova, Louis Penafiel, Kristina Lerman, Tim Weninger |
ICWSM | 4 |
| 2025 | Fear and Loathing on the Frontline: Decoding the Language of Othering by Russia-Ukraine War BloggersabstractOthering—the process of portraying an outgroup as fundamentally different and inferior—often escalates into framing the outgroup as an existential threat, thereby legitimizing exclusion and violence. Throughout history, othering has played a central role in conflicts, from genocides in Nazi Germany and Rwanda to contemporary hostility toward migrants in the US and Europe. Traditional computational methods, such as those used for hate speech detection, frequently overlook the subtle, context-dependent nature of othering language, limiting their effectiveness in real-time detection and analysis. Our work addresses these limitations through three key contributions: (1) a computational framework that combines sociological theory with large language models (LLMs) to identify and analyze othering language, (2) an in-depth examination of othering discourse dynamics, focusing on attention patterns and its interplay with moral framing, and (3) a rapid domain adaptation enabling robust analysis across different platforms and contexts. We apply our framework to a large corpus of Telegram messages from Russo-Ukrainian war bloggers and political discourse on Gab, revealing several previously unquantified patterns: othering rhetoric surges during crises, often intertwines with moralized language, and escalates during critical periods. Our findings demonstrate that this approach not only surpasses existing hate and fear speech detection methods but also offers actionable insights for anticipating and mitigating threats to social cohesion in conflict-prone environments. Patrick Gérard, Tim Weninger, Kristina Lerman |
ICWSM | 3 |
| 2025 | The Peripatetic Hater: Predicting Movement Among Hate SubredditsabstractMany online hate groups exist to disparage others based on race, gender identity, sex, or other characteristics. The accessibility of these communities allows users to join multiple types of hate groups (e.g., a racist community and a misogynistic community), raising the question of whether users who join additional types of hate communities could be further radicalized compared to users who stay in one type of hate group. However, little is known about the dynamics of joining multiple types of hate groups, nor the effect of these groups on peripatetic users. We develop a new method to classify hate subreddits and the identities they disparage, then apply it to understand better how users come to join different types of hate subreddits. The hate classification technique utilizes human-validated deep learning models to extract the protected identities attacked, if any, across 168 subreddits. We find distinct clusters of subreddits targeting various identities, such as racist subreddits, xenophobic subreddits, and transphobic subreddits. We show that when users become active in their first hate subreddit, they have a high likelihood of becoming active in additional hate subreddits of a different category. We also find that users who join additional hate subreddits, especially those of a different category develop a wider hate group lexicon. These results then lead us to train a classification model that, as we demonstrate, usefully predicts the hate categories in which users will become active based on post text replied to and written. The accuracy of this model may be partly driven by peripatetic users often using the language of hate subreddits they eventually join. Overall, these results highlight the unique risks associated with hate communities on a social media platform, as discussion of alternative targets of hate may lead users to target more protected identities. Daniel Hickey, Daniel Fessler, Matheus Schmitz, Kristina Lerman, Keith Burghardt |
ICWSM | 4 |
| 2025 | Polarized Online Discourse on Abortion: Frames and Hostile Expressions Among Liberals and ConservativesabstractAbortion has been one of the most divisive issues in the United States. Yet, missing is comprehensive longitudinal evidence on how political divides on abortion are reflected in public discourse over time, on a national scale, and in response to key events before and after the overturn of Roe v Wade. We analyze a corpus of over 3.5M tweets related to abortion over the span of one year (January 2022 to January 2023) from over 1.1M users. We estimate users' ideology and rely on state-of-the-art transformer-based classifiers to identify expressions of hostility and extract five prominent frames surrounding abortion. We use those data to examine (a) how prevalent were expressions of hostility (i.e., anger, toxic speech, insults, obscenities, and hate speech), (b) what frames liberals and conservatives used to articulate their positions on abortion, and (c) the prevalence of hostile expressions in liberals and conservative discussions of these frames. We show that liberals and conservatives largely mirrored each other's use of hostile expressions: as liberals used more hostile rhetoric, so did conservatives, especially in response to key events. In addition, the two groups used distinct frames and discussed them in vastly distinct contexts, suggesting that liberals and conservatives have differing perspectives on abortion. Lastly, frames favored by one side provoked hostile reactions from the other: liberals use more hostile expressions when addressing religion, fetal personhood, and exceptions to abortion bans, whereas conservatives use more hostile language when addressing bodily autonomy and women's health. This signals disrespect and derogation, which may further preclude understanding and exacerbate polarization. Ashwin Rao, Rong-Ching Chang, Qiankun Zhong, Kristina Lerman, Magdalena Wojcieszak |
ICWSM | 4 |
| 2025 | In-Group Love, Out-Group Hate: A Framework to Measure Affective Polarization via Contentious Online DiscussionsabstractAffective polarization, the emotional divide between ideological groups marked by in-group love and out-group hate, has intensified in the United States, driving contentious issues like masking and lockdowns during the COVID-19 pandemic. Despite its societal impact, existing models of opinion change fail to account for emotional dynamics nor offer methods to quantify affective polarization robustly and in real-time. In this paper, we introduce a discrete choice model that captures decision-making within affectively polarized social networks and propose a statistical inference method estimate key parameters---in-group love and out-group hate---from social media data. Through empirical validation from online discussions about the COVID-19 pandemic, we demonstrate that our approach accurately captures real-world polarization dynamics and explains the rapid emergence of a partisan gap in attitudes towards masking and lockdowns. This framework allows for tracking affective polarization across contentious issues has broad implications for fostering constructive online dialogues in digital spaces. Buddhika Nettasinghe, Ashwin Rao, Bohan Jiang, Allon G. Percus, Kristina Lerman |
WWW | 5 |
| 2025 | The Pulse of Mood Online: Unveiling Emotional Reactions in a Dynamic Social Media LandscapeabstractThe rich and dynamic information environment of social media provides researchers, policymakers, and entrepreneurs with opportunities to learn about social phenomena in a timely manner. However, using these data to understand social behavior is difficult due to the heterogeneity of topics and events discussed in the highly dynamic online information environment. To address these challenges, we present a method for systematically detecting and measuring emotional reactions to offline events using change point detection on the time series of collective affect and further explaining these reactions using a transformer-based topic model. We demonstrate the utility of the method by successfully detecting major and smaller events on three different datasets, including (1) a Los Angeles Tweet dataset between Jan. and Aug. 2020, in which we revealed the complex psychological impact of the BlackLivesMatter movement and the COVID-19 pandemic, (2) a dataset related to abortion rights discussions in the USA, in which we uncovered the strong emotional reactions to the overturn of Roe v. Wade and state abortion bans, and (3) a dataset about the 2022 French presidential election, in which we discovered the emotional and moral shift from positive before voting to fear and criticism after voting. We further demonstrate the importance of disaggregating data by topics and populations to mitigate potential biases when studying collective emotions. The capability of our method allows for better sensing and monitoring of the population’s reactions during crises using online data. Siyi Guo, Ashwin Rao, Fred Morstatter, P. Jeffrey Brantingham, Kristina Lerman |
ACM Trans. Web | 6 |
| 2024 | Impacts of Personalization on Social Network Exposure
Nathan Bartley, Keith Burghardt, Kristina Lerman |
ASONAM (2) | 3 |
| 2024 | Non-binary Gender Expression in Online Interactions
Rebecca Dorn, Negar Mokhberian, Julie Jiang, Jeremy Abramson, Fred Morstatter, Kristina Lerman |
ASONAM (2) | 6 |
| 2024 | Socio-Linguistic Characteristics of Coordinated Inauthentic AccountsabstractOnline manipulation is a pressing concern for democracies, but the actions and strategies of coordinated inauthentic accounts, which have been used to interfere in elections, are not well understood. We analyze a five million-tweet multilingual dataset related to the 2017 French presidential election, when a major information campaign led by Russia called "#MacronLeaks" took place. We utilize heuristics to identify coordinated inauthentic accounts and detect attitudes, concerns and emotions within their tweets, collectively known as socio-linguistic characteristics. We find that coordinated accounts retweet other coordinated accounts far more than expected by chance, while being exceptionally active just before the second round of voting. Concurrently, socio-linguistic characteristics reveal that coordinated accounts share tweets promoting a candidate at three times the rate of non-coordinated accounts. Coordinated account tactics also varied in time to reflect news events and rounds of voting. Our analysis highlights the utility of socio-linguistic characteristics to inform researchers about tactics of coordinated accounts and how these may feed into online social manipulation. Keith Burghardt, Ashwin Rao, Georgios Chochlakis, Sabyasachee Baruah, Siyi Guo, Andrew Rojecki, Shri Narayanan, Kristina Lerman |
ICWSM | 9 |
| 2024 | IsamasRed: A Public Dataset Tracking Reddit Discussions on Israel-Hamas ConflictabstractThe conflict between Israel and Palestinians significantly escalated after the October 7, 2023 Hamas attack, capturing global attention. To understand the public discourse on this conflict, we present a meticulously compiled dataset-IsamasRed-comprising nearly 400,000 conversations and over 8 million comments from Reddit, spanning from August 2023 to November 2023. We introduce an innovative keyword extraction framework leveraging a large language model to effectively identify pertinent keywords, ensuring a comprehensive data collection. Our initial analysis on the dataset, examining topics, controversy, emotional and moral language trends over time, highlights the emotionally charged and complex nature of the discourse. This dataset aims to enrich the understanding of online discussions, shedding light on the complex interplay between ideology, sentiment, and community engagement in digital spaces. Keith Burghardt, Jingxin Zhang 0010, Kristina Lerman |
ICWSM | 5 |
| 2024 | CPL-NoViD: Context-Aware Prompt-Based Learning for Norm Violation Detection in Online CommunitiesabstractDetecting norm violations in online communities is critical to maintaining healthy and safe spaces for online discussions. Existing machine learning approaches often struggle to adapt to the diverse rules and interpretations across different communities due to the inherent challenges of fine-tuning models for such context-specific tasks. In this paper, we introduce Context-aware Prompt-based Learning for Norm Violation Detection (CPL-NoViD), a novel method that employs prompt-based learning to detect norm violations across various types of rules. CPL-NoViD outperforms the baseline by incorporating context through natural language prompts and demonstrates improved performance across different rule types. Significantly, it not only excels in cross-rule-type and cross-community norm violation detection but also exhibits adaptability in few-shot learning scenarios. Most notably, it establishes a new state-of-the-art in norm violation detection, surpassing existing benchmarks. Our work highlights the potential of prompt-based learning for context-sensitive norm violation detection and paves the way for future research on more adaptable, context-aware models to better support online community moderators. Jonathan May, Kristina Lerman |
ICWSM | 3 |
| 2024 | Discovering Collective Narratives Shifts in Online DiscussionsabstractNarratives are foundation of human cognition and decision making. Because narratives play a crucial role in societal discourses and spread of misinformation and because of the pervasive use of social media, the narrative dynamics on social media can have profound societal impact. Yet, systematic and computational understanding of online narratives faces critical challenge of the scale and dynamics; how can we reliably and automatically extract narratives from massive amount of texts? How do narratives emerge, spread, and die? Here, we propose a systematic narrative discovery framework that fill this gap by combining change point detection, semantic role labeling (SRL), and automatic aggregation of narrative fragments into narrative networks. We evaluate our model with synthetic and empirical data — two Twitter corpora about COVID-19 and 2017 French Election. Results demonstrate that our approach can recover major narrative shifts that correspond to the major events. Wanying Zhao, Siyi Guo, Kristina Lerman, Yong-Yeol Ahn |
ICWSM | 3 |
| 2023 | Evaluating Content Exposure Bias in Social NetworksabstractOnline social platforms employ personalized feed algorithms to gather and prioritize messages from accounts followed by users, which distorts content's perceived popularity prior to personalization. We call this "exposure bias," and our research focuses on quantifying it using diverse exposure bias metrics, and we evaluate recommendation algorithms through various content ranking heuristics. Similarly we simulate activity in a network to assess the influence of such ranking heuristics on exposure bias. Furthermore, we are working on agent-based model simulations to comprehend the impact of ranking schemes, with the ultimate goal of exploring intervention effects over time. Our empirical findings reveal that users exposed to popularity-based feeds experience significantly lower exposure bias compared to chronologically-ordered feeds. Nathan Bartley, Keith Burghardt, Kristina Lerman |
ASONAM | 3 |
| 2023 | Measuring Online Emotional Reactions to EventsabstractThe rich and dynamic information environment of social media provides researchers, policy makers, and entrepreneurs with opportunities to learn about social phenomena in a timely manner. However, using this data to understand social behavior is difficult due heterogeneity of topics and events discussed in the highly dynamic online information environment. To address these challenges, we present a method for systematically detecting and measuring emotional reactions to offline events using change point detection on the time series of collective affect, and further explaining these reactions using a transformer-based topic model. We demonstrate the utility of the method on a corpus of tweets from a large US metropolitan area between January and August, 2020, covering a period of great social change. We demonstrate that our method is able to disaggregate topics to measure population's emotional and moral reactions. This capability allows for better monitoring of population's reactions during crises using online data. Siyi Guo, Ashwin Rao, Eugene Jang, Yuanfeixue Nan, Fred Morstatter, P. Jeffrey Brantingham, Kristina Lerman |
ASONAM | 8 |
| 2023 | Retweets Amplify the Echo Chamber EffectabstractThe growing prominence of social media in public discourse has led to a greater scrutiny of the quality of online information and the role it plays in amplifying political polarization. However, studies of polarization on social media platforms like Twitter have been hampered by the difficulty of collecting data about the social graph, specifically follow links that shape the echo chambers users join as well as what they see in their timelines. As a proxy of the follower graph, researchers use retweets, although it is not clear how this choice affects analysis. Using a sample of the Twitter follower graph and the tweets posted by users within it, we reconstruct the retweet graph and quantify its impact on the measures of echo chambers and exposure. While we find that echo chambers exist in both graphs, they are more pronounced in the retweet graph. We compare the information users see via their follower and retweet networks to show that retweeted accounts share systematically more polarized content. This bias cannot be explained by the activity or polarization within users' own follower graph neighborhoods but by the increased attention they pay to accounts that are ideologically aligned with their own views. Our results suggest that studies relying on the retweet graphs overestimate the echo chamber effects and exposure to polarized information. Ashwin Rao, Fred Morstatter, Kristina Lerman |
ASONAM | 3 |
| 2023 | Pandemic Culture Wars: Partisan Differences in the Moral Language of COVID-19 DiscussionsabstractEffective response to pandemics requires coordinated adoption of mitigation measures, like masking and quarantines, to curb a virus’s spread. However, as the COVID-19 pandemic demonstrated, political divisions can hinder consensus on the appropriate response. To better understand these divisions, our study examines a vast collection of COVID-19-related tweets. We focus on five contentious issues: coronavirus origins, lockdowns, masking, education, and vaccines. We describe a weakly supervised method to identify issue-relevant tweets and employ state-of-the-art computational methods to analyze moral language and infer political ideology. We explore how partisanship and moral language shape conversations about these issues. Our findings reveal ideological differences in issue salience and moral language used by different groups. We find that conservatives use more negatively-valenced moral language than liberals and that political elites use moral rhetoric to a greater extent than non-elites across most issues. Examining the evolution and moralization on divisive issues can provide valuable insights into the dynamics of COVID-19 discussions and assist policymakers in better understanding the emergence of ideological divisions. Ashwin Rao, Siyi Guo, Sze-Yuh Nina Wang, Fred Morstatter, Kristina Lerman |
IEEE Big Data | 5 |
| 2023 | #RoeOverturned: Twitter Dataset on the Abortion Rights ControversyabstractOn June 24, 2022, the United States Supreme Court overturned landmark rulings made in its 1973 verdict in Roe v. Wade. The justices by way of a majority vote in Dobbs v. Jackson Women's Health Organization, decided that abortion wasn't a constitutional right and returned the issue of abortion to the elected representatives. This decision triggered multiple protests and debates across the US, especially in the context of the midterm elections in November 2022. Given that many citizens use social media platforms to express their views and mobilize for collective action, and given that online debate provides tangible effects on public opinion, political participation, news media coverage, and the political decision-making, it is crucial to understand online discussions surrounding this topic. Toward this end, we present the first large-scale Twitter dataset collected on the abortion rights debate in the United States. We present a set of 74M tweets systematically collected over the course of one year from January 1, 2022 to January 6, 2023. Rong-Ching Chang, Ashwin Rao, Qiankun Zhong, Magdalena Wojcieszak, Kristina Lerman |
ICWSM | 5 |
| 2023 | A Data Fusion Framework for Multi-Domain Morality LearningabstractLanguage models can be trained to recognize the moral sentiment of text, creating new opportunities to study the role of morality in human life. As interest in language and morality has grown, several ground truth datasets with moral annotations have been released. However, these datasets vary in the method of data collection, domain, topics, instructions for annotators, etc. Simply aggregating such heterogeneous datasets during training can yield models that fail to generalize well. We describe a data fusion framework for training on multiple heterogeneous datasets that improve performance and generalizability. The model uses domain adversarial training to align the datasets in feature space and a weighted loss function to deal with label shift. We show that the proposed framework achieves state-of-the-art performance in different datasets compared to prior works in morality inference. Siyi Guo, Negar Mokhberian, Kristina Lerman |
ICWSM | 3 |
| 2022 | Noise Audits Improve Moral Foundation ClassificationabstractMorality plays an important role in culture, identity, and emotion. Recent advances in natural language processing have shown that it is possible to classify moral values expressed in text at scale. Morality classification relies on human annotators to label the moral expressions in text, which provides training data to achieve state-of-the-art performance. However, these annotations are inherently subjective and some of the instances are hard to classify, resulting in noisy annotations due to error or lack of agreement. The presence of noise in training data harms the classifier's ability to accurately recognize moral foundations from text. We propose two metrics to audit the noise of annotations. The first metric is entropy of instance labels, which is a proxy measure of annotator disagreement about how the instance should be labeled. The second metric is the silhouette coefficient of a label assigned by an annotator to an instance. This metric leverages the idea that instances with the same label should have similar latent representations, and deviations from collective judgments are indicative of errors. Our experiments on three widely used moral foundations datasets show that removing noisy annotations based on the proposed metrics improves classification performance.11Our code can be found at: https://github.com/negar-mokhberian/noise-audits. Negar Mokhberian, Frederic R. Hopp, Bahareh Harandizadeh, Fred Morstatter, Kristina Lerman |
ASONAM | 5 |
| 2022 | Assessing Scientific Research Papers with Knowledge GraphsabstractIn recent decades, the growing scale of scientific research has led to numerous novel findings. Reproducing these findings is the foundation of future research. However, due to the complexity of experiments, manually assessing scientific research is laborious and time-intensive, especially in social and behavioral sciences. Although increasing reproducibility studies have garnered increased attention in the research community, there is still a lack of systematic ways for evaluating scientific research at scale. In this paper, we propose a novel approach towards automatically assessing scientific publications by constructing a knowledge graph (KG) that captures a holistic view of the research contributions. Specifically, during the KG construction, we combine information from two different perspectives: micro-level features that capture knowledge from published articles such as sample sizes, effect sizes, and experimental models, and macro-level features that comprise relationships between entities such as authorship and reference information. We then learn low-dimensional representations using language models and knowledge graph embeddings for entities (nodes in KGs), which are further used for the assessments. A comprehensive set of experiments on two benchmark datasets shows the usefulness of leveraging KGs for scoring scientific research. Kexuan Sun 0002, Zhiqiang Qiu, Abel Salinas, Yuzhong Huang, Daniel Benjamin, Fred Morstatter, Xiang Ren 0001, Kristina Lerman, Jay Pujara |
SIGIR | 9 |
| 2021 | Follow the leader: Documents on the leading edge of semantic change get more citationsabstractAbstract Diachronic word embeddings—vector representations of words over time—offer remarkable insights into the evolution of language and provide a tool for quantifying sociocultural change from text documents. Prior work has used such embeddings to identify shifts in the meaning of individual words. However, simply knowing that a word has changed in meaning is insufficient to identify the instances of word usage that convey the historical meaning or the newer meaning. In this study, we link diachronic word embeddings to documents, by situating those documents as leaders or laggards with respect to ongoing semantic changes. Specifically, we propose a novel method to quantify the degree of semantic progressiveness in each word usage, and then show how these usages can be aggregated to obtain scores for each document. We analyze two large collections of documents, representing legal opinions and scientific articles. Documents that are scored as semantically progressive receive a larger number of citations, indicating that they are especially influential. Our work thus provides a new technique for identifying lexical semantic leaders and demonstrates a new link between progressive use of language and influence in a citation network. Sandeep Soni, Kristina Lerman, Jacob Eisenstein |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2020 | Can Badges Foster a More Welcoming Culture on Q&A Boards?
Keith Burghardt, Kristina Lerman, Denis Helic |
ICWSM | 3 |
| 2019 | Linguistic Cues to Deception: Identifying Political Trolls on Social Media
Aseel Addawood, Adam Badawy, Kristina Lerman, Emilio Ferrara |
ICWSM | 3 |
| 2018 | Analyzing the Digital Traces of Political Manipulation: The 2016 Russian Interference Twitter CampaignabstractUntil recently, social media was seen to promote democratic discourse on social and political issues. However, this powerful communication platform has come under scrutiny for allowing hostile actors to exploit online discussions in an attempt to manipulate public opinion. A case in point is the ongoing U.S. Congress investigation of Russian interference in the 2016 U.S. election campaign, with Russia accused of, among other things, using trolls (malicious accounts created for the purpose of manipulation) and bots (automated accounts) to spread misinformation and politically biased information. In this study, we explore the effects of this manipulation campaign, taking a closer look at users who re-shared the posts produced on Twitter by the Russian troll accounts publicly disclosed by U.S. Congress investigation. We collected a dataset with over 43 million elections-related posts shared on Twitter between September 16 and November 9, 2016 by about 5.7 million distinct users. This dataset includes accounts associated with the identified Russian trolls. We use label propagation to infer the users' ideology based on the news sources they shared, to classify a large number of them as liberal or conservative with precision and recall above 90%. Conservatives retweeted Russian trolls significantly more often than liberals and produced 36 times more tweets. Additionally, most of the troll content originated in, and was shared by users from Southern states. Using state-of-the-art bot detection techniques, we estimated that about 4.9% and 6.2% of liberal and conservative users respectively were bots. Text analysis on the content shared by trolls reveals that they had a mostly conservative, pro-Trump agenda. Although an ideologically broad swath of Twitter users were exposed to Russian trolls in the period leading up to the 2016 U.S. Presidential election, it was mainly conservatives who helped amplify their message. Adam Badawy, Emilio Ferrara, Kristina Lerman |
ASONAM | 3 |
| 2018 | Using Simpson's Paradox to Discover Interesting Patterns in Behavioral Data
Nazanin Alipourfard, Peter G. Fennell, Kristina Lerman |
ICWSM | 3 |
| 2018 | Quantifying the Impact of Cognitive Biases in Question-Answering Systems
Keith Burghardt, Tad Hogg, Kristina Lerman |
ICWSM | 3 |
| 2018 | Modeling Evolution of Topics in Large-Scale Temporal Text Corpora
Elaheh Momeni, Shanika Karunasekera, Palash Goyal, Kristina Lerman |
ICWSM | 4 |
| 2018 | Can you Trust the Trend?: Discovering Simpson's Paradoxes in Social DataabstractWe investigate how Simpson»s paradox affects analysis of trends in social data. According to the paradox, the trends observed in data that has been aggregated over an entire population may be different from, and even opposite to, those of the underlying subgroups. Failure to take this effect into account can lead analysis to wrong conclusions. We present a statistical method to automatically identify Simpson»s paradox in data by comparing statistical trends in the aggregate data to those in the disaggregated subgroups. We apply the approach to data from Stack Exchange, a popular question-answering platform, to analyze factors affecting answerer performance, specifically, the likelihood that an answer written by a user will be accepted by the asker as the best answer to his or her question. Our analysis confirms a known Simpson»s paradox and identifies several new instances. These paradoxes provide novel insights into user behavior on Stack Exchange. Nazanin Alipourfard, Peter G. Fennell, Kristina Lerman |
WSDM | 3 |
| 2017 | On Quitting: Performance and Practice in Online Game Play
Tushar Agarwal, Keith Burghardt, Kristina Lerman |
ICWSM | 3 |
| 2017 | Dynamics of Content Quality in Collaborative Knowledge Production
Emilio Ferrara, Nazanin Alipourfard, Keith Burghardt, Chiranth Gopal, Kristina Lerman |
ICWSM | 5 |
| 2017 | iPhone's Digital Marketplace: Characterizing the Big SpendersabstractWith mobile shopping surging in popularity, people are spending ever more money on digital purchases through their mobile devices and phones. However, few large-scale studies of mobile shopping exist. In this paper we analyze a large data set consisting of more than 776M digital purchases made on Apple mobile devices that include songs, apps, and in-app purchases. We find that 61% of all the spending is on in-app purchases and that the top 1% of users are responsible for 59% of all the spending. These big spenders are more likely to be male and older, and less likely to be from the US. We study how they adopt and abandon individual app, and find that, after an initial phase of increased daily spending, users gradually lose interest: the delay between their purchases increases and the spending decreases with a sharp drop toward the end. Finally, we model the in-app purchasing behavior in multiple steps: 1) we model the time between purchases; 2) we train a classifier to predict whether the user will make a purchase from a new app or continue purchasing from the existing app; and 3) based on the outcome of the previous step, we attempt to predict the exact app, new or existing, from which the next purchase will come. The results yield new insights into spending habits in the mobile digital marketplace. Farshad Kooti, Mihajlo Grbovic, Luca Maria Aiello, Eric Bax, Kristina Lerman |
WSDM | 5 |
| 2017 | Taming the Unpredictability of Cultural Markets with Social InfluenceabstractUnpredictability is often portrayed as an undesirable outcome of social influence in cultural markets. Unpredictability stems from the "rich get richer" effect, whereby small fluctuations in the market share or popularity of products are amplified over time by social influence. In this paper, we report results of an experimental study that shows that unpredictability is not an inherent property of social influence. We investigate strategies for creating markets in which the popularity of products is better-and more predictably-aligned with their underlying quality. For our study, we created a cultural market of science stories and conducted randomized experiments on different policies for presenting the stories to study participants. Specifically, we varied how the stories were ranked, and whether or not participants were shown the ratings these stories received from others. We present a policy that leverages social influence and product positioning to help distinguish the product's market share (popularity) from underlying quality. Highlighting products with the highest estimated quality reduces the "rich get richer" effect highlighting popular products. We show that this policy allows us to more robustly and predictably identify high quality products and promote blockbusters. The policy can be used to create more efficient online cultural markets with a better allocation of resources to products. Andrés Abeliuk, Gerardo Berbeglia, Pascal Van Hentenryck, Tad Hogg, Kristina Lerman |
WWW | 5 |
| 2017 | Effort Mediates Access to Information in Online Social NetworksabstractIndividuals’ access to information in a social network depends on how it is distributed and where in the network individuals position themselves. In addition, individuals vary in how much effort they invest in managing their social connections. Using data from a social media site, we study how the interplay between effort and network position affects social media users’ access to diverse and novel information. Previous studies of the role of networks in information access were limited in their ability to measure the diversity of information. We address this problem by learning the topics of interest to social media users from the messages they share online with followers. We use the learned topics to measure the diversity of information users receive from the people they follow online. We confirm that users in structurally diverse network positions, which bridge otherwise disconnected regions of the follower network, tend to be exposed to more diverse and novel information. We also show that users who invest more effort in their activity on the site are not only located in more structurally diverse positions within the network than the less engaged users but also receive more novel and diverse information when in similar network positions. These findings indicate that the relationship between network structure and access to information in networks is more nuanced than previously thought. Jeon-Hyung Kang, Kristina Lerman |
ACM Trans. Web | 2 |
| 2016 | Leveraging the Contributions of the Casual Majority to Identify Appealing Web ContentabstractUsers of peer production web sites differ greatly in their activity levels.A small minority are engaged contributors, while the vast majority are only casual surfers. The casual users devote little effort to evaluating the site's content and many of them visit the site only once. This churn poses a challenge for sites attempting to gauge user interest in their content. The challenge is especially severe for sites focusing on content with subjective quality, including movies, music, restaurants and items in other cultural markets. A key question is whether content evaluation should use opinions of all users or only the minority who devote significant effort to reviewing content? Using Amazon Mechanical Turk, we experimentally address this question by comparing outcomes for these two approaches. We find that the larger numbers of less informed users more than offset their noisy signals on content quality to provide rapid evaluation. However, such users are systematically biased, and the speed of their assessments comes at the expense of limited collective accuracy. Tad Hogg, Kristina Lerman |
HCOMP | 2 |
| 2016 | Emotions, Demographics and Sociability in Twitter Interactions
Kristina Lerman, Megha Arora, Luciano Gallegos, Ponnurangam Kumaraguru, David García 0001 |
ICWSM | 1 |
| 2016 | Portrait of an Online Shopper: Understanding and Predicting Consumer BehaviorabstractConsumer spending accounts for a large fraction of economic footprint of modern countries. Increasingly, consumer activity is moving to the web, where digital receipts of online purchases provide valuable data sources detailing consumer behavior. We consider such data extracted from emails and combined with with consumers' demographic information, which we use to characterize, model, and predict purchasing behavior. We analyze such behavior of consumers in different age and gender groups, and find interesting, actionable patterns that can be used to improve ad targeting systems. For example, we found that the amount of money spent on online purchases grows sharply with age, peaking in the late 30s, while shoppers from wealthy areas tend to purchase more expensive items and buy them more frequently. Furthermore, we look at the influence of social connections on purchasing habits, as well as at the temporal dynamics of online shopping where we discovered daily and weekly behavioral patterns. Finally, we build a model to predict when shoppers are most likely to make a purchase and how much will they spend, showing improvement over baseline approaches. The presented results paint a clear picture of a modern online shopper, and allow better understanding of consumer behavior that can help improve marketing efforts and make shopping more pleasant and efficient experience for online customers. Farshad Kooti, Kristina Lerman, Luca Maria Aiello, Mihajlo Grbovic, Nemanja Djuric, Vladan Radosavljevic |
WSDM | 2 |
| 2016 | Partitioning Networks with Node Attributes by Compressing Information FlowabstractReal-world networks are often organized as modules or communities of similar nodes that serve as functional units. These networks are also rich in content, with nodes having distinguished features or attributes. In order to discover a network’s modular structure, it is necessary to take into account not only its links but also node attributes. We describe an information-theoretic method that identifies modules by compressing descriptions of information flow on a network. Our formulation introduces node content into the description of information flow, which we then minimize to discover groups of nodes with similar attributes that also tend to trap the flow of information. The method is conceptually simple and does not require ad-hoc parameters to specify the number of modules or to control the relative contribution of links and node attributes to network structure. We apply the proposed method to partition real-world networks with known community structure. We demonstrate that adding node attributes helps recover the underlying community structure in content-rich networks more effectively than using links alone. In addition, we show that our method is faster and more accurate than alternative state-of-the-art algorithms. Laura M. Smith, Linhong Zhu, Kristina Lerman, Allon G. Percus |
ACM Trans. Knowl. Discov. Data | 3 |
| 2015 | User Effort and Network Structure Mediate Access to Information in Networks
Jeon-Hyung Kang, Kristina Lerman |
ICWSM | 2 |
| 2015 | Evolution of Conversations in the Age of Email OverloadabstractEmail is a ubiquitous communications tool in the workplace and plays an important role in social interactions. Previous studies of email were largely based on surveys and limited to relatively small populations of email users within organizations. In this paper, we report results of a large-scale study of more than 2 million users exchanging 16 billion emails over several months. We quantitatively characterize the replying behavior in conversations within pairs of users. In particular, we study the time it takes the user to reply to a received message and the length of the reply sent. We consider a variety of factors that affect the reply time and length, such as the stage of the conversation, user demographics, and use of portable devices. In addition, we study how increasing load affects emailing behavior. We find that as users receive more email messages in a day, they reply to a smaller fraction of them, using shorter replies. However, their responsiveness remains intact, and they may even reply to emails faster. Finally, we predict the time to reply, length of reply, and whether the reply ends a conversation. We demonstrate considerable improvement over the baseline in all three prediction tasks, showing the significant role that the factors that we uncover play, in determining replying behavior. We rank these factors based on their predictive power. Our findings have important implications for understanding human behavior and designing better email management applications for tasks like ranking unread emails. Farshad Kooti, Luca Maria Aiello, Mihajlo Grbovic, Kristina Lerman, Amin Mantrach |
WWW | 4 |
| 2014 | Placing user-generated content on the map with confidenceabstractWe describe a method that predicts the location of user-generated content using textual features alone. Unlike previous methods for geotagging text documents, our proposed method is not sensitive to how we discretize space. We also discover that spatial resolution has an impact on the prediction accuracy, which allows us to trade-off the spatial resolution of the predicted location against our confidence about its accuracy. Our method can be used to estimate the error in document's predicted location, enabling us to filter out poor quality predictions. We evaluate the proposed method extensively on user-generated content collected from two different social media sites, Flickr and Twitter. Our evaluation examines its performance on the geotagging task and with respect to different parameters. We achieve state-of-the-art results for all three tasks: location prediction, error estimation and result ranking and also provide a theoretical explanation of the effect of spatial resolution factor on geotagging accuracy. Our findings provide valuable insights into the design of geotagging systems and their quality control. Suradej Intagorn, Kristina Lerman |
SIGSPATIAL/GIS | 2 |
| 2014 | Network Weirdness: Exploring the Origins of Network Paradoxes
Farshad Kooti, Nathan Oken Hodas, Kristina Lerman |
ICWSM | 3 |
| 2014 | The interplay between dynamics and networks: centrality, communities, and cheeger inequalityabstractWe study the interplay between a dynamic process and the structure of the network on which it is defined. Specifically, we examine the impact of this interaction on the quality-measure of network clusters and node centrality. This enables us to effectively identify network communities and important nodes participating in the dynamics. As the first step towards this objective, we introduce an umbrella framework for defining and characterizing an ensemble of dynamic processes on a network. This framework generalizes the traditional Laplacian framework to continuous-time biased random walks and also allows us to model some epidemic processes over a network. For each dynamic process in our framework, we can define a function that measures the quality of every subset of nodes as a potential cluster (or community) with respect to this process on a given network. This subset-quality function generalizes the traditional conductance measure for graph partitioning. We partially justify our choice of the quality function by showing that the classic Cheeger's inequality, which relates the conductance of the best cluster in a network with a spectral quantity of its Laplacian matrix, can be extended from the Laplacian-conductance setting to this more general setting. Rumi Ghosh, Shang-Hua Teng, Kristina Lerman, Xiaoran Yan |
KDD | 3 |
| 2014 | Tripartite graph clustering for dynamic sentiment analysis on social mediaabstractThe growing popularity of social media (e.g., Twitter) allows users to easily share information with each other and influence others by expressing their own sentiments on various subjects. In this work, we propose an unsupervised tri-clustering framework, which analyzes both user-level and tweet-level sentiments through co-clustering of a tripartite graph. A compelling feature of the proposed framework is that the quality of sentiment clustering of tweets, users, and features can be mutually improved by joint clustering. We further investigate the evolution of user-level sentiments and latent feature vectors in an online framework and devise an efficient online algorithm to sequentially update the clustering of tweets, users and features with newly arrived data. The online framework not only provides better quality of both dynamic user-level and tweet-level sentiment analysis, but also improves the computational and storage efficiency. We verified the effectiveness and efficiency of the proposed approaches on the November 2012 California ballot Twitter data. Linhong Zhu, Aram Galstyan, James Cheng, Kristina Lerman |
SIGMOD Conference | 4 |
| 2013 | Identifying Transformative Scientific ResearchabstractTransformative research refers to research that shifts or disrupts established scientific paradigms. Notable examples include the discovery of high-temperature superconductivity that disrupted the theory established 30 years ago. Identifying potential transformative research early and accurately is important for funding agencies to maximize the impact of their investments. It also helps scientists identify and focus their attention on promising emerging works. This paper presents a data driven approach where citation patterns of scientific papers are analyzed to quantify how much a potential challenger idea shifts an established paradigm. The key idea is that transformative research creates an observable disruption in the structure of "information cascades," chains of references that can be traced back to the papers establishing some scientific paradigm. Such a disruption is visible soon after the challenger's introduction. We define a disruption score to quantify the disruption and develop an algorithm to compute it from a large citation network. Experimental results show that our approach can successfully identify transformative scientific papers that disrupt established paradigms in Physics and Computer Science, regardless of whether the challenger paradigm is an instant hit or a classic whose contribution is formally recognized with a Nobel Prize decades later. Yi-Hung Huang, Chun-Nan Hsu, Kristina Lerman |
ICDM | 3 |
| 2013 | Friendship Paradox Redux: Your Friends Are More Interesting Than You
Nathan Oken Hodas, Farshad Kooti, Kristina Lerman |
ICWSM | 3 |
| 2012 | A probabilistic approach to mining geospatial knowledge from social annotationsabstractUser-generated content, such as photos and videos, is often annotated by users with free-text labels, called tags. Increasingly, such content is also georeferenced, i.e., it is associated with geographic coordinates. The implicit relationships between tags and their locations can tell us much about how people conceptualize places and relations between them. However, extracting such knowledge from social annotations presents many challenges, since annotations are often ambiguous, noisy, uncertain and spatially inhomogeneous. We introduce a probabilistic framework for modeling georeferenced annotations and a method for learning model parameters from data. The framework is flexible and general, and can be used in a variety of applications that mine geospatial knowledge from user-generated content. Specifically, we study three problems: extracting place semantics, predicting locations of photos and learning part-of relations between places. We show our method performs well compared to state-of-the-art approaches developed for the first two problems, and offers a novel solution to the problem of learning relations between places. Suradej Intagorn, Kristina Lerman |
CIKM | 2 |
| 2012 | Characterising Emergent Semantics in Twitter Lists
Andrés García-Silva, Jeon-Hyung Kang, Kristina Lerman, Óscar Corcho |
ESWC | 3 |
| 2012 | Semi-automatically Mapping Structured Sources into the Semantic Web
Craig A. Knoblock, Pedro A. Szekely, José Luis Ambite, Aman Goel, Kristina Lerman, Maria Muslea, Mohsen Taheriyan, Parag Mallick |
ESWC | 6 |
| 2012 | Using Stochastic Models to Describe and Predict Social Dynamics of Web UsersabstractThe popularity of content in social media is unequally distributed, with some items receiving a disproportionate share of attention from users. Predicting which newly-submitted items will become popular is critically important for both the hosts of social media content and its consumers. Accurate and timely prediction would enable hosts to maximize revenue through differential pricing for access to content or ad placement. Prediction would also give consumers an important tool for filtering the content. Predicting the popularity of content in social media is challenging due to the complex interactions between content quality and how the social media site highlights its content. Moreover, most social media sites selectively present content that has been highly rated by similar users, whose similarity is indicated implicitly by their behavior or explicitly by links in a social network. While these factors make it difficult to predict popularity a priori , stochastic models of user behavior on these sites can allow predicting popularity based on early user reactions to new content. By incorporating the various mechanisms through which web sites display content, such models improve on predictions that are based on simply extrapolating from the early votes. Specifically, for one such site, the news aggregator Digg, we show how a stochastic model distinguishes the effect of the increased visibility due to the network from how interested users are in the content. We find a wide range of interest, distinguishing stories primarily of interest to users in the network (“niche interests”) from those of more general interest to the user community. This distinction is useful for predicting a story’s eventual popularity from users’ early reactions to the story. Kristina Lerman, Tad Hogg |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2011 | Learning boundaries of vague places from noisy annotationsabstractWhat ordinary people mean by places may differ dramatically from what experts consider them to be. This is especially evident in how people talk about places in social media, where 'Los Angeles', for instance, could include areas well outside of the city or even in another county. In order to make best use of the information in social media, we need to understand what people mean when they refer to a place. Social annotations provide valuable evidence for harvesting knowledge about places, e.g., learning their boundaries and relations to other places. However, social annotations are noisy, and this can dramatically distort the learned boundaries. In this paper we propose a method that exploits the distinctive property of social annotations --- that it is created by many people --- to filter out noise. Using a large data set extracted from Flickr we show that our crowd-based noise filtering method can learn accurate boundaries of places, including vague places. Suradej Intagorn, Kristina Lerman |
GIS | 2 |
| 2011 | What Stops Social Epidemics?
Greg Ver Steeg, Rumi Ghosh, Kristina Lerman |
ICWSM | 3 |
| 2011 | A framework for quantitative analysis of cascades on networksabstractHow does information flow in online social networks? How does the structure and size of the information cascade evolve in time? How can we efficiently mine the information contained in cascade dynamics? We approach these questions empirically and present an efficient and scalable mathematical framework for quantitative analysis of cascades on networks. We define a cascade generating function that captures the details of the microscopic dynamics of the cascades. We show that this function can also be used to compute the macroscopic properties of cascades, such as their size, spread, diameter, number of paths, and average path length. We present an algorithm to efficiently compute cascade generating function and demonstrate that while significantly compressing information within a cascade, it nevertheless allows us to accurately reconstruct its structure. We use this framework to study information dynamics on the social network of Digg. Digg allows users to post and vote on stories, and easily see the stories that friends have voted on. As a story spreads on Digg through voting, it generates cascades. We extract cascades of more than 3,500 Digg stories and calculate their macroscopic and microscopic properties. We identify several trends in cascade dynamics: spreading via chaining, branching and community. We discuss how these affect the spread of the story through the Digg social network. Our computational framework is general and offers a practical solution to quantitative analysis of the microscopic structure of even very large cascades. Rumi Ghosh, Kristina Lerman |
WSDM | 2 |
| 2011 | A probabilistic approach for learning folksonomies from structured dataabstractLearning structured representations has emerged as an important problem in many domains, including document and Web data mining, bioinformatics, and image analysis. One approach to learning complex structures is to integrate many smaller, incomplete and noisy structure fragments. In this work, we present an unsupervised probabilistic approach that extends affinity propagation [7] to combine the small ontological fragments into a collection of integrated, consistent, and larger folksonomies. This is a challenging task because the method must aggregate similar structures while avoiding structural inconsistencies and handling noise. We validate the approach on a real-world social media dataset, comprised of shallow personal hierarchies specified by many individual users, collected from the photosharing website Flickr. Our empirical results show that our proposed approach is able to construct deeper and denser structures, compared to an approach using only the standard affinity propagation algorithm. Additionally, the approach yields better overall integration quality than a state-of-the-art approach based on incremental relational clustering. Anon Plangprasopchok, Kristina Lerman, Lise Getoor |
WSDM | 2 |
| 2011 | Pragmatic evaluation of folksonomiesabstractRecently, a number of algorithms have been proposed to obtain hierarchical structures - so-called folksonomies - from social tagging data. Work on these algorithms is in part driven by a belief that folksonomies are useful for tasks such as: (a) Navigating social tagging systems and (b) Acquiring semantic relationships between tags. While the promises and pitfalls of the latter have been studied to some extent, we know very little about the extent to which folksonomies are pragmatically useful for navigating social tagging systems. This paper sets out to address this gap by presenting and applying a pragmatic framework for evaluating folksonomies. We model exploratory navigation of a tagging system as decentralized search on a network of tags. Evaluation is based on the fact that the performance of a decentralized search algorithm depends on the quality of the background knowledge used. The key idea of our approach is to use hierarchical structures learned by folksonomy algorithm as background knowledge for decentralized search. Utilizing decentralized search on tag networks in combination with different folksonomies as hierarchical background knowledge allows us to evaluate navigational tasks in social tagging systems. Our experiments with four state-of-the-art folksonomy algorithms on five different social tagging datasets reveal that existing folksonomy algorithms exhibit significant, previously undiscovered, differences with regard to their utility for navigation. Our results are relevant for engineers aiming to improve navigability of social tagging systems and for scientists aiming to evaluate different folksonomy algorithms from a pragmatic perspective. Denis Helic, Markus Strohmaier, Christoph Trattner, Markus Muhr, Kristina Lerman |
WWW | 5 |
| 2010 | Social Dynamics of Digg
Tad Hogg, Kristina Lerman |
ICWSM | 2 |
| 2010 | Information Contagion: An Empirical Study of the Spread of News on Digg and Twitter Social Networks
Kristina Lerman, Rumi Ghosh |
ICWSM | 1 |
| 2010 | Growing a tree in the forest: constructing folksonomies by integrating structured metadataabstractMany social Web sites allow users to annotate the content with descriptive metadata, such as tags, and more recently to organize content hierarchically. These types of structured metadata provide valuable evidence for learning how a community organizes knowledge. For instance, we can aggregate many personal hierarchies into a common taxonomy, also known as a folksonomy, that will aid users in visualizing and browsing social content, and also to help them in organizing their own content. However, learning from social metadata presents several challenges, since it is sparse, shallow, ambiguous, noisy, and inconsistent. We describe an approach to folksonomy learning based on relational clustering, which exploits structured metadata contained in personal hierarchies. Our approach clusters similar hierarchies using their structure and tag statistics, then incrementally weaves them into a deeper, bushier tree. We study folksonomy learning using social metadata extracted from the photo-sharing site Flickr, and demonstrate that the proposed approach addresses the challenges. Moreover, comparing to previous work, the approach produces larger, more accurate folksonomies, and in addition, scales better. Anon Plangprasopchok, Kristina Lerman, Lise Getoor |
KDD | 2 |
| 2010 | Using a model of social dynamics to predict popularity of newsabstractPopularity of content in social media is unequally distributed, with some items receiving a disproportionate share of attention from users. Predicting which newly-submitted items will become popular is critically important for both companies that host social media sites and their users. Accurate and timely prediction would enable the companies to maximize revenue through differential pricing for access to content or ad placement. Prediction would also give consumers an important tool for filtering the ever-growing amount of content. Predicting popularity of content in social media, however, is challenging due to the complex interactions among content quality, how the social media site chooses to highlight content, and influence among users. While these factors make it difficult to predict popularity a priori, we show that stochastic models of user behavior on these sites allows predicting popularity based on early user reactions to new content. By incorporating aspects of the web site design, such models improve on predictions based on simply extrapolating from the early votes. We validate this claim on the social news portal Digg using a previously-developed model of social voting based on the Digg user interface. Kristina Lerman, Tad Hogg |
WWW | 1 |
| 2010 | Constructing folksonomies by integrating structured metadataabstractAggregating many personal hierarchies into a common taxonomy, also known as a folksonomy, presents several challenges due to its sparseness, ambiguity, noise, and inconsistency. We describe an approach to folksonomy learning based on relational clustering that addresses these challenges by exploiting structured metadata contained in personal hierarchies. Our approach clusters similar hierarchies using their structure and tag statistics, then incrementally weaves them into a deeper, bushier tree. We study folksonomy learning using social metadata extracted from the photo-sharing site Flickr. We evaluate the learned folksonomy quantitatively by automatically comparing it to a reference taxonomy created by the Open Directory Project. Our empirical results suggest that the proposed approach improves upon the state-of-the-art folksonomy learning method. Anon Plangprasopchok, Kristina Lerman, Lise Getoor |
WWW | 2 |
| 2010 | Modeling Social Annotation: A Bayesian ApproachabstractCollaborative tagging systems, such asDelicious, CiteULike, and others, allow users to annotate resources, for example, Web pages or scientific papers, with descriptive labels calledtags. The social annotations contributed by thousands of users can potentially be used to infer categorical knowledge, classify documents, or recommend new relevant information. Traditional text inference methods do not make the best use of social annotation, since they do not take into account variations in individual users’ perspectives and vocabulary. In a previous work, we introduced a simple probabilistic model that takes the interests of individual annotators into account in order to find hidden topics of annotated resources. Unfortunately, that approach had one major shortcoming: the number of topics and interests must be specified a priori. To address this drawback, we extend the model to a fully Bayesian framework, which offers a way to automatically estimate these numbers. In particular, the model allows the number of interests and topics to change as suggested by the structure of the data. We evaluate the proposed model in detail on the synthetic and real-world data by comparing its performance to Latent Dirichlet Allocation on the topic extraction task. For the latter evaluation, we apply the model to infer topics of Web resources from social annotations obtained fromDeliciousin order to discover new resources similar to a specified one. Our empirical results demonstrate that the proposed model is a promising method for exploiting social knowledge contained in user-generated annotations. Anon Plangprasopchok, Kristina Lerman |
ACM Trans. Knowl. Discov. Data | 2 |
| 2009 | Leaders and Negotiators: An Influence-based Metric for Rank
Rumi Ghosh, Kristina Lerman |
ICWSM | 2 |
| 2009 | Stochastic Models of User-Contributory Web Sites
Tad Hogg, Kristina Lerman |
ICWSM | 2 |
| 2009 | Automatically Constructing Semantic Web Services from Online Sources
José Luis Ambite, Sirish Darbha, Aman Goel, Craig A. Knoblock, Kristina Lerman, Rahul Parundekar, Thomas A. Russ |
ISWC | 5 |
| 2009 | Constructing folksonomies from user-specified relations on flickrabstractAutomatic folksonomy construction from tags has attracted much attention recently. However, inferring hierarchical relations between concepts from tags has a drawback in that it is difficult to distinguish between more popular and more general concepts. Instead of tags we propose to use user-specified relations for learning folksonomy. We explore two statistical frameworks for aggregating many shallow individual hierarchies, expressed through the collection/set relations on the social photosharing site Flickr, into a common deeper folksonomy that reflects how a community organizes knowledge. Our approach addresses a number of challenges that arise while aggregating information from diverse users, namely noisy vocabulary, and variations in the granularity level of the concepts expressed. Our second contribution is a method for automatically evaluating learned folksonomy by comparing it to a reference taxonomy, e.g., the Web directory created by the Open Directory Project. Our empirical results suggest that user-specified relations are a good source of evidence for learning folksonomies. Anon Plangprasopchok, Kristina Lerman |
WWW | 2 |
| 2007 | Social Networks and Social Information Filtering on Digg
Kristina Lerman |
ICWSM | 1 |
| 2007 | Social Browsing on Flickr
Kristina Lerman, Laurie Jones |
ICWSM | 1 |
| 2007 | Semantic Labeling of Online Information SourcesabstractIn order to combine data from various heterogeneous sources, software agents must first understand the semantics of the sources, expressed in the source model. Currently, source modeling is manual, but as large numbers of sources come online, it is impractical to expect users to continue modeling them by hand. We describe two machine learning techniques for automatically modeling information sources: one that uses source’s metadata, contained in a Web Service Definition file, and one that uses the source’s content, to classify the semantics of the data it uses. We go beyond previous works and verify predictions by invoking the source with sample data of the predicted type. We provide performance results of both methods and validate our approach on several live Web sources. In addition, we describe the application of semantic modeling within the CALO project. Kristina Lerman, Anon Plangprasopchok, Craig A. Knoblock |
Int. J. Semantic Web Inf. Syst. | 1 |
| 2004 | Using the Structure of Web Sites for Automatic Segmentation of TablesabstractMany Web sites, especially those that dynamically generate HTML pages to display the results of a user's query, present information in the form of list or tables. Current tools that allow applications to programmatically extract this information rely heavily on user input, often in the form of labeled extracted records. The sheer size and rate of growth of the Web make any solution that relies primarily on user input is infeasible in the long term. Fortunately, many Web sites contain much explicit and implicit structure, both in layout and content, that we can exploit for the purpose of information extraction. This paper describes an approach to automatic extraction and segmentation of records from Web tables. Automatic methods do not require any user input, but rely solely on the layout and content of the Web source. Our approach relies on the common structure of many Web sites, which present information as a list or a table, with a link in each entry leading to a detail page containing additional information about that item. We describe two algorithms that use redundancies in the content of table and detail pages to aid in information extraction. The first algorithm encodes additional information provided by detail pages as constraints and finds the segmentation by solving a constraint satisfaction problem. The second algorithm uses probabilistic inference to find the record segmentation. We show how each approach can exploit the web site structure in a general, domain-independent manner, and we demonstrate the effectiveness of each algorithm on a set of twelve Web sites. Kristina Lerman, Lise Getoor, Steven Minton, Craig A. Knoblock |
SIGMOD Conference | 1 |