Munmun De Choudhury

dblp:76/3034 · DBLP profile ↗
← Back
36ranked-venue papers in the field
14as first author
14since 2021 · last 2025
0000-0002-8939-264XORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 33 (13 first)Data Mining & Knowledge Discovery · 1Big Data, Cloud & Distributed Data Systems · 1Other / Interdisciplinary · 1 (1 first)
YearPublicationVenuePosition
2025 Timeliness Matters: Leveraging Reinforcement Learning on Social Media Data to Prioritize High-Risk Conversations for Promoting Youth Online Safety
abstract
Ensuring the online safety of youth has motivated research towards the development of machine learning (ML) methods capable of accurately detecting social media risks after-the-fact. However, for these detection models to be effective, they must proactively identify high-risk scenarios (e.g., sexual solicitations, cyberbullying) to mitigate harm. This `real-time' responsiveness is a recognized challenge within the risk detection literature. Therefore, this paper presents a novel two-level framework that first uses reinforcement learning to identify conversation stop points to prioritize messages for evaluation. Then, we optimize state-of-the-art deep learning models to accurately categorize risk priority (low, high). We apply this framework to a time-based simulation using a rich dataset of 23K private conversations with over 7 million messages donated by 194 youth (ages 13-21). We conducted an experiment comparing our new approach to a traditional conversation-level baseline. We found that the timeliness of conversations significantly improved from over 2 hours to approximately 16 minutes with only a slight reduction in accuracy (0.88 to 0.84). This study advances real-time detection approaches for social media data and provides a benchmark for future training reinforcement learning that prioritizes the timeliness of classifying high-risk conversations.
Ashwaq Alsoubai, Jinkyung Park, Gianluca Stringhini, Meiyi Ma, Munmun De Choudhury, Pamela J. Wisniewski
ICWSM5
2025 A Growing Sense of Alienation: Spirals of Silence and Suppression of Structural Circumstances of Suicide in News
abstract
Suicide is a leading cause of death in the United States. Global safe reporting guidelines for news reports of suicide intend to mitigate associations of increased suicide incidence and stigma. However, recent research suggests more latent patterns in news beyond the guidelines could still contribute to suicide outcomes such as inhibited help-seeking and isolation. Using the Theory of Spiral of Silence to center isolation, we take a mixed-methods approach to analyze 22,021 articles (2020-2024) and use a zero-shot learning large language model (LLM) classifier to detect suppression of four structural circumstances of suicide: financial/job, legal, school, and access to physical/mental healthcare. We find that circumstance disclosure by news publishers diverges by political leaning, financial (p = 0.016), legal (p < 0.001), and school (p < 0.001); and by regionality, legal (p < 0.001) and health (p < 0.001). We qualify mechanisms of suppression using topic modeling and content sharing networks (CSNs). The spiral of silence lens highlights that left leaning publishers are more likely to disclose systemically or socially collective circumstances. In contrast, right leaning outlets suppress those and instead disclose instances that blame individuals for their experiences. Our work highlights how news reporting can downplay structural factors contributing to suicide. Content Warning: This paper discusses suicide deaths reported in news articles and may be sensitive to readers.
Jasmine C. Foriest, Mini Jain, Benjamin D. Horne, Munmun De Choudhury
ICWSM4
2025 Large-Scale Analysis of Online Questions Related to Opioid Use Disorder on Reddit
abstract
Opioid use disorder (OUD) is a leading health problem that affects individual well-being as well as general public health. Due to a variety of reasons, including the stigma faced by people using opioids, online communities for recovery and support were formed on different social media platforms. In these communities, people share their experiences and solicit information by asking questions to learn about opioid use and recovery. However, these communities do not always contain clinically verified information. In this paper, we study natural language questions asked in the context of OUD-related discourse on Reddit. We adopt transformer-based question detection along with hierarchical clustering across 19 subreddits to identify six coarse-grained categories and 69 fine-grained categories of OUD-related questions. Our analysis uncovers ten areas of information seeking from Reddit users in the context of OUD: drug sales, specific drug-related questions, OUD treatment, drug uses, side effects, withdrawal, lifestyle, drug testing, pain management and others, during the study period of 2018-2021. Our work provides a major step in improving the understanding of OUD-related questions people ask unobtrusively on Reddit. We finally discuss technological interventions and public health harm reduction techniques based on the topics of these questions.
Tanmay Laud, Akadia Kacha-Ochana, Steven A. Sumner, Vikram Krishnasamy, Royal Law, Lyna Schieber, Munmun De Choudhury, Mai ElSherief
ICWSM7
2025 Online Myths on Opioid Use Disorder: A Comparison of Reddit and Large Language Model
abstract
Online communities on Reddit are a popular choice among people with opioid use disorder (OUD) to seek information on drug use, withdrawal symptoms, and recovery. LLM-powered chatbots (e.g., ChatGPT) are widely being adopted as question-answer systems for health-related queries. However, such online health information seeking could potentially be hindered by myths and misinformation on OUD, misleading or causing genuine harm to people with OUD. In this work, we examine the prevalence of 5 OUD-related myths, on treatment models and patient characteristics, within human- (taken from Reddit) and LLM-generated responses to queries on OUD. We further explore the framing strategies used within responses (both human- and LLM-generated) promoting and countering the myths. We found that all 5 myths were more widespread within human-generated responses. In addition, myth-promoting responses adopted trustworthy and authoritative framings, compared to knowledge-imparting linguistic cues within those countering the myths. Our work offers recommendations to reduce online OUD misinformation.
Shravika Mittal, Hayoung Jung, Mai ElSherief, Tanushree Mitra, Munmun De Choudhury
ICWSM5
2025 Supporters and Skeptics: LLM-Based Analysis of Engagement with Mental Health (Mis)Information Content on Video-Sharing Platforms
abstract
Over one in five adults in the US lives with a mental illness. In the face of a shortage of mental health professionals and offline resources, online short-form video content has grown to serve as a crucial conduit for disseminating mental health help and resources. However, the ease of content creation and access also contributes to the spread of misinformation, posing risks to accurate diagnosis and treatment. Detecting and understanding engagement with such content is crucial to mitigating their harmful effects on public health. We perform the first quantitative study of the phenomenon using YouTube Shorts and Bitchute as the sites of study. We contribute MentalMisinfo, a novel labeled mental health misinformation (MHMisinfo) dataset of 739 videos (639 from Youtube and 100 from Bitchute) and 135372 comments in total, using an expert-driven annotation schema. We first found that few-shot in-context learning with large language models (LLMs) are effective in detecting MHMisinfo videos. Next, we discover distinct and potentially alarming linguistic patterns in how audiences engage with MHMisinfo videos through commentary on both video-sharing platforms. Across the two platforms, comments could exacerbate prevailing stigma with some groups showing heightened susceptibility to and alignment with MHMisinfo. We discuss technical and public health-driven adaptive solutions to tackling the "epidemic" of mental health misinformation online.
Viet Cuong Nguyen, Mini Jain, Abhijat Chauhan, Heather Jaime Soled, Santiago Alvarez Lesmes, Michael L. Birnbaum, Sunny X. Tang, Srijan Kumar, Munmun De Choudhury
ICWSM10
2025 Mental Health Impact of the COVID-19 Pandemic on College Students: A Quasi-Experimental Study on Social Media
abstract
Given the limited understanding of how the COVID-19 pandemic impacted mental health on college campuses, this paper examines the evolution of the mental health of college students since the onset of the pandemic. We conducted a large-scale study on over 1.2M posts on 173 U.S. college subreddits over 17 months. In particular, we adopted a quasi-experimental approach to examine how the different stages of the pandemic (isolation period, normalization period, and vaccination period) impacted changes in social media discussions of college students. We measured the temporal shifts in the symptomatic mental health expressions and topics of discussion on college subreddits. We find that while the expressions of depression, anxiety, stress, and suicidal ideation significantly increased in the isolation period. Interestingly, these expressions gradually subsided in the normalization period, only to resurface in the vaccination period. We also find unique occurrences of discussion across social, academic, health, and COVID-19-induced topics. Our findings reveal that despite the fragility of college students' mental health in the face of crisis, college students show resilience with sufficient time. We discuss the implications of our work in terms of building tools for real-time comprehension of college students' mental health, and in designing timely and tailored mental health support for college students.
Koustuv Saha, Bhaskar Kotakonda, Munmun De Choudhury
ICWSM3
2024 The LGBTQ+ Minority Stress on Social Media (MiSSoM) Dataset: A Labeled Dataset for Natural Language Processing and Machine Learning
abstract
Minority stress is the leading theoretical construct for understanding LGBTQ+ health disparities. As such, there is an urgent need to develop innovative policies and technologies to reduce minority stress. To spur technological innovation, we created the largest labeled datasets on minority stress using natural language from subreddits related to sexual and gender minority people. A team of mental health clinicians, LGBTQ+ health experts, and computer scientists developed two datasets: (1) the publicly available LGBTQ+ Minority Stress on Social Media (MiSSoM) dataset and (2) the advanced request-only version of the dataset, LGBTQ+ MiSSoM+. Both datasets have seven labels related to minority stress, including an overall composite label and six sublabels. LGBTQ+ MiSSoM (N = 27,709) includes both human- and machine-annotated la-bels and comes preprocessed with features (e.g., topic models, psycholinguistic attributes, sentiment, clinical keywords, word embeddings, n-grams, lexicons). LGBTQ+ MiSSoM+ includes all the characteristics of the open-access dataset, but also includes the original Reddit text and sentence-level labeling for a subset of posts (N = 5,772). Benchmark supervised machine learning analyses revealed that features of the LGBTQ+ MiSSoM datasets can predict overall minority stress quite well (F1 = 0.869). Benchmark performance metrics yielded in the prediction of the other labels, namely prejudiced events (F1 = 0.942), expected rejection (F1 = 0.964), internalized stigma (F1 = 0.952), identity concealment (F1 = 0.971), gender dysphoria (F1 = 0.947), and minority coping (F1 = 0.917), were excellent. Descriptive analyses, ethical considerations, limitations, and possible use cases are provided.
Cory J. Cascalheira, Santosh Chapagain, Ryan E. Flinn, Dannie Klooster, Danica Laprade, Emily M. Lund, Alejandra Gonzalez, Kelsey Corro, Rikki Wheatley, Ana Gutiérrez, Oziel Garcia Villanueva, Koustuv Saha, Munmun De Choudhury, Jillian R. Scheer, Shah Muhammad Hamdi
ICWSM14
2024 Assessing the Impact of Online Harassment on Youth Mental Health in Private Networked Spaces
abstract
Online harassment negatively impacts mental health, with victims expressing increased concerns such as depression, anxiety, and even increased risk of suicide, especially among youth and young adults. Yet, research has mainly focused on building automated systems to detect harassment incidents based on publicly available social media trace data, overlooking the impact of these negative events on the victims, especially in private channels of communication. Looking to close this gap, we examine a large dataset of private message conversations from Instagram shared and annotated by youth aged 13-21. We apply trained classifiers from online mental health to analyze the impact of online harassment on indicators pertinent to mental health expressions. Through a robust causal inference design involving a difference-in-differences analysis, we show that harassment results in greater expression of mental health concerns in victims up to 14 days following the incidents, while controlling for time, seasonality, and topic of conversation. Our study provides new benchmarks to quantify how victims perceive online harassment in the immediate aftermath of when it occurs. We make social justice-centered design recommendations to support harassment victims in private networked spaces. We caution that some of the paper's content could be triggering to readers.
Afsaneh Razi, Ashwaq Alsoubai, Pamela J. Wisniewski, Munmun De Choudhury
ICWSM5
2024 News Media and Violence against Women: Understanding Framings of Stigma
abstract
Discussions of Violence Against Women (VAW) in publicly accessible forums like online news media can influence the perceptions of people and organizations. Language reinforcing stigma around VAW can result in negative consequences such as unethical representation of survivors and trivialization of the act of violence. In this work, we study the presence of stigmatized framings in news media and how it differs based on media attributes like regionality, political leaning, veracity, and latent communities of news sources. We also investigate the interactions between VAW-based stigma and 14 issue-generic policies used to describe political communications. We found that articles from national, right-leaning, and conspiratorial news sources contain more stigma compared to their counterparts. Furthermore, alignment of articles to the issue-generic policies offers the highest explanation for the presence of stigma in news articles. We discuss implications for institutions to improve safe reporting guidelines on VAW.
Shravika Mittal, Jasmine C. Foriest, Benjamin D. Horne, Munmun De Choudhury
ICWSM4
2024 Better to Ask in English: Cross-Lingual Evaluation of Large Language Models for Healthcare Queries
Yiqiao Jin, Mohit Chandra, Gaurav Verma 0005, Yibo Hu 0002, Munmun De Choudhury, Srijan Kumar
WWW5
2023 Partisan US News Media Representations of Syrian Refugees
abstract
We investigate how representations of Syrian refugees (2011-2021) differ across US partisan news outlets. We analyze 47,388 articles from the online US media about Syrian refugees to detail differences in reporting between left- and right-leaning media. We use various NLP techniques to understand these differences. Our polarization and question answering results indicated that left-leaning media tended to represent refugees as child victims, welcome in the US, and right-leaning media cast refugees as Islamic terrorists. We noted similar results with our sentiment and offensive speech scores over time, which detail possibly unfavorable representations of refugees in right-leaning media. A strength of our work is how the different techniques we have applied validate each other. Based on our results, we provide several recommendations. Stakeholders may utilize our findings to intervene around refugee representations, and design communications campaigns that improve the way society sees refugees and possibly aid refugee outcomes.
Marzieh Babaeianjelodar, Yiwen Shi, Kamila Janmohamed, Rupak Sarkar, Ingmar Weber, Thomas Davidson, Munmun De Choudhury, Jonathan Huang, Shweta Yadav 0001, Ashiqur R. KhudaBukhsh, Chris T. Bauch, Preslav Nakov, Orestis Papakyriakopoulos, Koustuv Saha, Kaveh Khoshnood, Navin Kumar 0004
ICWSM8
2022 Classifying Minority Stress Disclosure on Social Media with Bidirectional Long Short-Term Memory
Cory J. Cascalheira, Shah Muhammad Hamdi, Jillian R. Scheer, Koustuv Saha, Soukaina Filali Boubrahimi, Munmun De Choudhury
ICWSM6
2022 Overcoming Language Disparity in Online Content Classification with Multimodal Learning
Gaurav Verma 0005, Rohit Mujumdar, Zijie J. Wang, Munmun De Choudhury, Srijan Kumar
ICWSM4
2021 You Don't Know How I Feel: Insider-Outsider Perspective Gaps in Cyberbullying Risk Detection
Afsaneh Razi, Gianluca Stringhini, Pamela J. Wisniewski, Munmun De Choudhury
ICWSM5
2019 A Social Media Study on the Effects of Psychiatric Medication Use
Koustuv Saha, Benjamin Sugar, John B. Torous, Bruno D. Abrahao, Emre Kiciman, Munmun De Choudhury
ICWSM6
2018 Measuring the Impact of Anxiety on Online Social Interactions
Sarmistha Dutta, Jennifer Ma, Munmun De Choudhury
ICWSM3
2018 Characterizing Audience Engagement and Assessing Its Impact on Social Media Disclosures of Mental Illnesses
Sindhu Kiranmai Ernala, Tristan Labetoulle, Fred Bane, Michael L. Birnbaum, Asra F. Rizvi, John Kane 0001, Munmun De Choudhury
ICWSM7
2018 A Social Media Based Examination of the Effects of Counseling Recommendations after Student Deaths on College Campuses
Koustuv Saha, Ingmar Weber, Munmun De Choudhury
ICWSM3
2017 #Anorexia, #anarexia, #anarexyia: Characterizing online community practices with orthographic variation
abstract
Distinctive linguistic practices help communities build solidarity and differentiate themselves from outsiders. In an online community, one such practice is variation in orthography, which includes spelling, punctuation, and capitalization. Using a dataset of over two million Instagram posts, we investigate orthographic variation in a community that shares pro-eating disorder (pro-ED) content. We find that not only does orthographic variation grow more frequent over time, it also becomes more profound or “deep,” with variants becoming increasingly distant from the original: as, for example, #anarexyia is more distant than #anarexia from the original spelling #anorexia. We find that the these changes are driven by newcomers, who adopt the most extreme linguistic practices as they enter the community. Moreover, this behavior correlates with engagement with the community: the newcomers that adopt deeper variant orthography tend to remain active for longer in the community, and posts with deeper variation receive more positive feedback in the form of “likes.” Previous work has linked community membership change with language change, and our work casts this connection in a new light, with newcomers driving an evolving practice rather than adapting to it. We also demonstrate the utility of orthographic variation as a new lens to study sociolinguistic change in online communities, particularly when the change results from an exogenous force such as a content ban.
Stevie Chancellor, Munmun De Choudhury, Jacob Eisenstein
IEEE BigData3
2017 The Language of Social Support in Social Media and Its Effect on Suicidal Ideation Risk
Munmun De Choudhury, Emre Kiciman
ICWSM1
2016 Social Media Participation in an Activist Movement for Racial Equality
Munmun De Choudhury, Shagun Jhaver, Benjamin Sugar, Ingmar Weber
ICWSM1
2016 Guest Editorial: Special Issue on Connected Health at Big Data Era (BigChat): A TKDD Special Issue
abstract
The availability of big data [James et al. 2011; Steve 2012] and the emergence of wearable computing [Thad 1996; Alex 2000], network science [Barabasi 2002], and computational social science [Hanna 2016; Watts and Strogatz 1998] as areas of inquiry has been revolutionizing the landscape of how we decipher our lives, our social interactions, and our day-to-day activities.This well-connected world has promised novel requirements on transforming healthcare from reactive and hospital-centered, to preventive, proactive, evidence-based, person-centered, and focused on well-being rather than ailment recovery.A multitude of various types of data are involved in this broad context of healthcare, including the following:-Clinical data [Prather et al. 1997; Riccardo and Zupan 2008], mainly the patient records from clinical institutions, such as medical imaging, patient electronic health records, clinical trial data, etc. -Genotype data [Eibe et al. 2004; Leslie et al. 1999], basically the genetic makeups of the individuals, such as DNA, protein, etc. -Social media data [Reza et al. 2014; Sitaram and Huberman 2010], which is the information the individuals posted on online social platforms such as Facebook, Twitter, PatientsLikeMe, etc. -Environmental sensory data [Ruchi and Bhatia 2010], which is the information sampled from the surrounding environment where the individuals are living in, such as air pollution and humidity information.-Behavioral and sentiment data [Bo and Lee 2008], which could be the data recorded by the wearable devices on patient's activities.-Mobile data [Miller and Han 2009; Fosca and Pedreschi 2008], which is sampled from individuals' mobile devices.Integrating multiple types of information to make people healthier is also a problem of vital importance that requires collective effort from different parties, where data mining plays a pivotal role.Toward this aim, the National Science Foundation of United
Hanghang Tong, Fei Wang 0001, Munmun De Choudhury, Zoran Obradovic
ACM Trans. Knowl. Discov. Data3
2015 Psychological Effects of Urban Crime Gleaned from Social Media
Jose Manuel Delgado Valdes, Jacob Eisenstein, Munmun De Choudhury
ICWSM3
2014 Mental Health Discourse on reddit: Self-Disclosure, Social Support, and Anonymity
Munmun De Choudhury, Sushovan De
ICWSM1
2014 Discussion Graphs: Putting Social Media Analysis in Context
Emre Kiciman, Scott Counts, Michael Gamon, Munmun De Choudhury, Bo Thiesson
ICWSM4
2013 Predicting Depression via Social Media
Munmun De Choudhury, Michael Gamon, Scott Counts, Eric Horvitz
ICWSM1
2012 Not All Moods Are Created Equal! Exploring Human Emotional States in Social Media
Munmun De Choudhury, Scott Counts, Michael Gamon
ICWSM1
2012 Happy, Nervous or Surprised? Classification of Human Affective States in Social Media
Munmun De Choudhury, Michael Gamon, Scott Counts
ICWSM1
2011 Find Me the Right Content! Diversity-Based Sampling of Social Media Spaces for Topic-Centric Search
Munmun De Choudhury, Scott Counts, Mary Czerwinski
ICWSM1
2010 How Does the Data Sampling Strategy Impact the Discovery of Information Diffusion in Social Media?
Munmun De Choudhury, Yu-Ru Lin, Hari Sundaram, K. Selçuk Candan, Lexing Xie, Aisling Kelliher
ICWSM1
2010 Constructing travel itineraries from tagged geo-temporal breadcrumbs
abstract
Vacation planning is a frequent laborious task which requires skilled interaction with a multitude of resources. This paper develops an end-to-end approach for constructing intra-city travel itineraries automatically by tapping a latent source reflecting geo-temporal breadcrumbs left by millions of tourists. In particular, the popular rich media sharing site, Flickr, allows photos to be stamped by the date and time of when they were taken, and be mapped to Points Of Interest (POIs) by latitude-longitude information as well as semantic metadata (e.g., tags) that describe them.
Munmun De Choudhury, Moran Feldman, Sihem Amer-Yahia, Nadav Golbandi, Ronny Lempel, Cong Yu 0001
WWW1
2010 Inferring relevant social networks from interpersonal communication
abstract
Researchers increasingly use electronic communication data to construct and study large social networks, effectively inferring unobserved ties (e.g. i is connected to j) from observed communication events (e.g. i emails j). Often overlooked, however, is the impact of tie definition on the corresponding network, and in turn the relevance of the inferred network to the research question of interest. Here we study the problem of network inference and relevance for two email data sets of different size and origin. In each case, we generate a family of networks parameterized by a threshold condition on the frequency of emails exchanged between pairs of individuals. After demonstrating that different choices of the threshold correspond to dramatically different network structures, we then formulate the relevance of these networks in terms of a series of prediction tasks that depend on various network features. In general, we find: a) that prediction accuracy is maximized over a non-trivial range of thresholds corresponding to 5-10 reciprocated emails per year; b) that for any prediction task, choosing the optimal value of the threshold yields a sizable (~30%) boost in accuracy over naive choices; and c) that the optimal threshold value appears to be (somewhat surprisingly) consistent across data sets and prediction tasks. We emphasize the practical utility in defining ties via their relevance to the prediction task(s) at hand and discuss implications of our empirical results.
Munmun De Choudhury, Winter A. Mason, Jake M. Hofman, Duncan J. Watts
WWW1
2010 Extraction, characterization and utility of prototypical communication groups in the blogosphere
abstract
This article analyzes communication within a set of individuals to extract the representative prototypical groups and provides a novel framework to establish the utility of such groups. Corporations may want to identify representative groups (which are indicative of the overall communication set) because it is easier to track the prototypical groups rather than the entire set. This can be useful for advertising, identifying “hot” spots of resource consumption as well as in mining representative moods or temperature of a community. Our framework has three parts: extraction, characterization, and utility of prototypical groups. First, we extract groups by developing features representing communication dynamics of the individuals. Second, to characterize the overall communication set, we identify a subset of groups within the community as the prototypical groups. Third, we justify the utility of these prototypical groups by using them as predictors of related external phenomena; specifically, stock market movement of technology companies and political polls of Presidential candidates in the 2008 U.S. elections. We have conducted extensive experiments on two popular blogs, Engadget and Huffington Post. We observe that the prototypical groups can predict stock market movement/political polls satisfactorily with mean error rate of 20.32%. Further, our method outperforms baseline methods based on alternative group extraction and prototypical group identification methods. We evaluate the quality of the extracted groups based on their conductance and coverage measures and develop metrics: predictivity and resilience to evaluate their ability to predict a related external time-series variable (stock market movement/political polls). This implies that communication dynamics of individuals are essential in extracting groups in a community, and the prototypical groups extracted by our method are meaningful in characterizing the overall communication sets.
Munmun De Choudhury, Hari Sundaram, Ajita John, Dorée D. Seligmann
ACM Trans. Inf. Syst.1
2009 What makes conversations interesting?: themes, participants and consequences of conversations in online social media
abstract
Rich media social networks promote not only creation and consumption of media, but also communication about the posted media item. What causes a conversation to be interesting, that prompts a user to participate in the discussion on a posted video? We conjecture that people participate in conversations when they find the conversation theme interesting, see comments by people whom they are familiar with, or observe an engaging dialogue between two or more people (absorbing back and forth exchange of comments). Importantly, a conversation that is interesting must be consequential - i.e. it must impact the social network itself.
Munmun De Choudhury, Hari Sundaram, Ajita John, Dorée D. Seligmann
WWW1
2008 Multi-scale characterization of social network dynamics in the blogosphere
abstract
We have developed a computational framework to characterize social network dynamics in the blogosphere at individual, group and community levels. Such characterization could be used by corporations to help drive targeted advertising and to track the moods and sentiments of consumers. We tested our model on a widely read technology blog called Engadget. Our results show that communities transit between states of high and low entropy, depending on sentiments (positive / negative) about external happenings. We also propose an innovative method to establish the utility of the extracted knowledge, by correlating the mined knowledge with an external time series data (the stock market). Our validation results show that the characterized groups exhibit high stock market movement predictability (89%) and removal of 'impactful' groups makes the community less resilient by lowering predictability (26%) and affecting the composition of the groups in the rest of the community.
Munmun De Choudhury, Hari Sundaram, Ajita John, Dorée D. Seligmann
CIKM1
2007 Contextual Prediction of Communication Flow in Social Networks
abstract
The paper develops a novel computational framework for predicting communication flow in social networks based on several contextual features. The problem is important because prediction of communication flow can impact timely sharing of specific information across a wide array of communities. We determine the intent to communicate and communication delay between users based on several contextual features in a social network corresponding to (a) neighborhood context, (b) topic context and (c) recipient context. The intent to communicate and communication delay are modeled as regression problems which are efficiently estimated using Support Vector Regression. We predict the intent and the delay, on an interval of time using past communication data. We have excellent prediction results on a real-world dataset from MySpace.com with an accuracy of 13-16%. We show that the intent to communicate is more significantly influenced by contextual factors compared to the delay.
Munmun De Choudhury, Hari Sundaram, Ajita John, Dorée D. Seligmann
Web Intelligence1