Haewoon Kwak

dblp:78/2468 · DBLP profile ↗
← Back
32ranked-venue papers in the field
7as first author
7since 2021 · last 2026
0000-0003-1418-0834ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 27 (6 first)Data Mining & Knowledge Discovery · 4 (1 first)Big Data, Cloud & Distributed Data Systems · 1
YearPublicationVenuePosition
2026 SynSym: A Synthetic Data Generation Framework for Psychiatric Symptom Identification
abstract
Psychiatric symptom identification on social media aims to infer fine-grained mental health symptoms from user-generated posts, allowing a detailed understanding of users' mental states. However, the construction of large-scale symptom-level datasets remains challenging due to the resource-intensive nature of expert labeling and the lack of standardized annotation guidelines, which in turn limits the generalizability of models to identify diverse symptom expressions from user-generated text. To address these issues, we propose SynSym, a synthetic data generation framework for constructing generalizable datasets for symptom identification. Leveraging large language models (LLMs), SynSym constructs high-quality training samples by (1) expanding each symptom into sub-concepts to enhance the diversity of generated expressions, (2) producing synthetic expressions that reflect psychiatric symptoms in diverse linguistic styles, and (3) composing realistic multi-symptom expressions, informed by clinical co-occurrence patterns. We validate SynSym on three benchmark datasets covering different styles of depressive symptom expression. Experimental results demonstrate that models trained solely on the synthetic data generated by SynSym perform comparably to those trained on real data, and benefit further from additional fine-tuning with real data. These findings underscore the potential of synthetic data as an alternative resource to real-world annotations in psychiatric symptom modeling, and SynSym serves as a practical framework for generating clinically relevant and realistic symptom expressions.
Migyeong Kang, Hyolim Jeon, Sunwoo Hwang, Jihyun An, Yonghoon Kim, Haewoon Kwak, Jisun An, Jinyoung Han
KDD (1)7
2023 YouNICon: YouTube's CommuNIty of Conspiracy Videos
abstract
Conspiracy theories are widely propagated on social media. Among various social media services, YouTube is one of the most influential sources of news and entertainment. This paper seeks to develop a dataset, YOUNICON, to enable researchers to perform conspiracy theory detection as well as classification of videos with conspiracy theories into different topics. YOUNICON is a dataset with a large collection of videos from suspicious channels that were identified to contain conspiracy theories in a previous study. Overall, YOUNICON will enable researchers to study trends in conspiracy theories and understand how individuals can interact with the conspiracy theory producing community or channel. Our data is available at: https://doi.org/10.5281/zenodo.7466262.
Shaoyi Liaw, Fabrício Benevenuto, Haewoon Kwak, Jisun An
ICWSM4
2023 "This Is Fake News": Characterizing the Spontaneous Debunking from Twitter Users to COVID-19 False Information
abstract
False information spreads on social media, and fact-checking is a potential countermeasure. However, there is a severe shortage of fact-checkers; an efficient way to scale fact-checking is desperately needed, especially in pandemics like COVID-19. In this study, we focus on spontaneous debunking by social media users, which has been missed in existing research despite its indicated usefulness for fact-checking and countering false information. Specifically, we characterize the tweets with false information, or fake tweets, that tend to be debunked and Twitter users who often debunk fake tweets. For this analysis, we create a comprehensive dataset of responses to fake tweets, annotate a subset of them, and build a classification model for detecting debunking behaviors. We find that most fake tweets are left undebunked, spontaneous debunking is slower than other forms of responses, and spontaneous debunking exhibits partisanship in political topics. These results provide actionable insights into utilizing spontaneous debunking to scale conventional fact-checking, thereby supplementing existing research from a new perspective.
Kunihiro Miyazaki, Takayuki Uchiba, Jisun An, Haewoon Kwak, Kazutoshi Sasahara
ICWSM5
2022 Characterizing Spontaneous Ideation Contest on Social Media: Case Study on the Name Change of Facebook to Meta
abstract
Collecting good ideas is vital for organizations, especially companies, to retain their competitiveness. Social media is gathering attention as a place to extract ideas efficiently; however, the characteristics of ideas and the posters of ideas on social media are underexamined. Thus, this study aims to characterize spontaneous ideation contests among social media users by taking an event of Facebook’s name change to Meta as a case study. As a dataset, we comprehensively collect tweets containing new acronyms of Big Tech companies, which we treat as an "idea" in this work. In the analysis, we especially focus on the diversity of ideas, which would be the main reason for enlisting social media for idea generation. As the main results, we discovered that social media users offered a wider range of ideas than those in mainstream media. The follow-follower network of the users suggested that the users’ position on the network is related to the preferred ideas. Additionally, we discovered a link between the amount of user interaction on social media and the diversity of ideas. This study would promote the use of social media as a part of open innovation and co-creation processes in the industry.
Kunihiro Miyazaki, Takayuki Uchiba, Haewoon Kwak, Jisun An
IEEE Big Data3
2022 Understanding Toxicity Triggers on Reddit in the Context of Singapore
Yun Yu Chong, Haewoon Kwak
ICWSM2
2022 Who Is Missing? Characterizing the Participation of Different Demographic Groups in a Korean Nationwide Daily Conversation Corpus
Haewoon Kwak, Jisun An, Kunwoo Park
ICWSM1
2021 How-to Present News on Social Media: A Causal Analysis of Editing News Headlines for Boosting User Engagement
Kunwoo Park, Haewoon Kwak, Jisun An, Sanjay Chawla
ICWSM2
2020 Identifying and Characterizing Alternative News Media on Facebook
abstract
As Internet users increasingly rely on social media sites to receive news, they are faced with a bewildering number of news media choices. For example, thousands of Facebook pages today are registered and categorized as some form of news media outlets. This situation boosted the so-called independent journalism, also known as alternative news media. Identifying and characterizing all the news pages that play an important role in news dissemination is key for understanding the news ecosystems of a country. In this work, we propose a graph-based semi-supervised method to measure the political bias of pages on most countries and show the political split of the alternative media, mainstream media, and public figures pages. We validate our method using the publicly available U.S. dataset and then apply it to Brazilian pages, where we found a larger number of right-wing pages in general, except for alternative news media.
Samuel S. Guimarães, Julio C. S. Reis, Lucas Henrique C. Lima, Filipe Nunes Ribeiro, Marisa A. Vasconcelos, Jisun An, Haewoon Kwak, Fabrício Benevenuto
ASONAM7
2020 Empirical Evaluation of Three Common Assumptions in Building Political Media Bias Datasets
Soumen Ganguly, Juhi Kulshrestha, Jisun An, Haewoon Kwak
ICWSM4
2020 "Trust Me, I Have a Ph.D.": A Propensity Score Analysis on the Halo Effect of Disclosing One's Offline Social Status in Online Communities
Kunwoo Park, Haewoon Kwak, Hyunho Song, Meeyoung Cha
ICWSM2
2020 Are These Comments Triggering? Predicting Triggers of Toxicity in Online Discussions
abstract
Understanding the causes or triggers of toxicity adds a new dimension to the prevention of toxic behavior in online discussions. In this research, we define toxicity triggers in online discussions as a non-toxic comment that lead to toxic replies. Then, we build a neural network-based prediction model for toxicity trigger. The prediction model incorporates text-based features and derived features from previous studies that pertain to shifts in sentiment, topic flow, and discussion context. Our findings show that triggers of toxicity contain identifiable features and that incorporating shift features with the discussion context can be detected with a ROC-AUC score of 0.87. We discuss implications for online communities and also possible further analysis of online toxicity and its root causes.
Hind A. Al-Merekhi, Haewoon Kwak, Joni Salminen, Jim Jansen
WWW2
2019 Political Discussions in Homogeneous and Cross-Cutting Communication Spaces
Jisun An, Haewoon Kwak, Oliver Posegga, Andreas Jungherr
ICWSM2
2018 Automatic Persona Generation (APG): A Rationale and Demonstration
abstract
We present Automatic Persona Generation (APG), a methodology and system for quantitative persona generation using large amounts of online social media data. The system is operational, beta deployed with several client organizations in multiple industry verticals and ranging from small-to-medium sized enterprises to large multi-national corporations. Using a robust web framework and stable back-end database, APG is currently processing tens of millions of user interactions with thousands of online digital products on multiple social media platforms, such as Facebook and YouTube. APG identifies both distinct and impactful user segments and then creates persona descriptions by automatically adding pertinent features, such as names, photos, and personal attributes. We present the overall methodological approach, architecture development, and main system features. APG has a potential value for organizations distributing content via online platforms and is unique in its approach to persona generation. APG can be found online at https://persona.qcri.org.
Soon-Gyo Jung, Joni Salminen, Haewoon Kwak, Jisun An, Jim Jansen
CHIIR3
2018 Fixation and Confusion: Investigating Eye-tracking Participants' Exposure to Information in Personas
abstract
To more effectively convey relevant information to end users of persona profiles, we conducted a user study consisting of 29 participants engaging with three persona layout treatments. We were interested in confusion engendered by the treatments on the participants, and conducted a within-subjects study in the actual work environment, using eye-tracking and talk-aloud data collection. We coded the verbal data into classes of informativeness and confusion and correlated it with fixations and durations on the Areas of Interests recorded by the eye-tracking device. We used various analysis techniques, including Mann-Whitney, regression, and Levenshtein distance, to investigate how confused users differed from non-confused users, what information of the personas caused confusion, and what were the predictors of confusion of end users of personas. We consolidate our various findings into a confusion ratio measure, which highlights in a succinct manner the most confusing elements of the personas. Findings show that inconsistencies among the informational elements of the persona generate the most confusion, especially with the elements of images and social media quotes. The research has implications for the design of personas and related information products, such as user profiling and customer segmentation.
Joni Salminen, Jim Jansen, Jisun An, Soon-Gyo Jung, Lene Nielsen, Haewoon Kwak
CHIIR6
2018 Assessing the Accuracy of Four Popular Face Recognition Tools for Inferring Gender, Age, and Race
Soon-Gyo Jung, Jisun An, Haewoon Kwak, Joni Salminen, Jim Jansen
ICWSM3
2018 Automatically Conceptualizing Social Media Analytics Data via Personas
Soon-Gyo Jung, Joni Salminen, Jisun An, Haewoon Kwak, Jim Jansen
ICWSM4
2018 Anatomy of Online Hate: Developing a Taxonomy and Machine Learning Models for Identifying and Classifying Hate in Online News Media
Joni Salminen, Hind A. Al-Merekhi, Milica Milenkovic, Soon-Gyo Jung, Jisun An, Haewoon Kwak, Jim Jansen
ICWSM6
2018 What We Read, What We Search: Media Attention and Public Attention Among 193 Countries
abstract
We investigate the alignment of international attention of news media organizations within 193 countries with the expressed international interests of the public within those same countries from March 7, 2016 to April 14, 2017. We collect fourteen months of longitudinal data of online news from Unfiltered News and web search volume data from Google Trends and build a multiplex network of media attention and public attention in order to study its structural and dynamic properties. Structurally, the media attention and the public attention are both similar and different depending on the resolution of the analysis. For example, we find that 63.2% of the country-specific media and the public pay attention to different countries, but local attention flow patterns, which are measured by network motifs, are very similar. We also show that there are strong regional similarities with both media and public attention that is only disrupted by significantly major worldwide incidents (e.g., Brexit). Using Granger causality, we show that there are a substantial number of countries where media attention and public attention are dissimilar by topical interest. Our findings show that the media and public attention toward specific countries are often at odds, indicating that the public within these countries may be ignoring their country-specific news outlets and seeking other online sources to address their media needs and desires.
Haewoon Kwak, Jisun An, Joni Salminen, Soon-Gyo Jung, Jim Jansen
WWW1
2018 Imaginary People Representing Real Numbers: Generating Personas from Online Social Media Data
abstract
We develop a methodology to automate creating imaginary people, referred to as personas, by processing complex behavioral and demographic data of social media audiences. From a popular social media account containing more than 30 million interactions by viewers from 198 countries engaging with more than 4,200 online videos produced by a global media corporation, we demonstrate that our methodology has several novel accomplishments, including: (a) identifying distinct user behavioral segments based on the user content consumption patterns; (b) identifying impactful demographics groupings; and (c) creating rich persona descriptions by automatically adding pertinent attributes, such as names, photos, and personal characteristics. We validate our approach by implementing the methodology into an actual working system; we then evaluate it via quantitative methods by examining the accuracy of predicting content preference of personas, the stability of the personas over time, and the generalizability of the method via applying to two other datasets. Research findings show the approach can develop rich personas representing the behavior and demographics of real audiences using privacy-preserving aggregated online social media data from major online platforms. Results have implications for media companies and other organizations distributing content via online platforms.
Jisun An, Haewoon Kwak, Soon-Gyo Jung, Joni Salminen, M. Admad, Jim Jansen
ACM Trans. Web2
2017 Personas for Content Creators via Decomposed Aggregate Audience Statistics
abstract
We propose a novel method for generating personas based on online user data for the increasingly common situation of content creators distributing products via online platforms. We use non-negative matrix factorization to identify user segments and develop personas by adding personality such as names and photos. Our approach can develop accurate personas representing real groups of people using online user data, versus relying on manually gathered data.
Jisun An, Haewoon Kwak, Jim Jansen
ASONAM2
2017 Multiplex Media Attention and Disregard Network among 129 Countries
abstract
We built a multiplex media attention and disregard network (MADN) among 129 countries over 212 days. By characterizing the MADN from multiple levels, we found that it is formed primarily by skewed, hierarchical, and asymmetric relationships. Also, we found strong evidence that our news world is becoming a "global village." However, at the same time, unique attention blocks of the Middle East and North Africa (MENA) region, as well as Russia and its neighbors, still exist.
Haewoon Kwak, Jisun An
ASONAM1
2017 What Gets Media Attention and How Media Attention Evolves Over Time: Large-Scale Empirical Evidence from 196 Countries
Jisun An, Haewoon Kwak
ICWSM2
2016 Are You Charlie or Ahmed? Cultural Pluralism in Charlie Hebdo Response on Twitter
Jisun An, Haewoon Kwak, Yelena Mejova, Sonia Alonso Saenz De Oger, Braulio Gomez Fortes
ICWSM2
2016 Two Tales of the World: Comparison of Widely Used World News Datasets GDELT and EventRegistry
Haewoon Kwak, Jisun An
ICWSM1
2015 Breaking the News: First Impressions Matter on Online News
Júlio Cesar dos Reis, Fabrício Benevenuto, Pedro O. S. Vaz de Melo, Raquel Oliveira Prates, Haewoon Kwak, Jisun An
ICWSM5
2014 STFU NOOB!: predicting crowdsourced decisions on toxic behavior in online games
abstract
One problem facing players of competitive games is negative, or toxic, behavior. League of Legends, the largest eSport game, uses a crowdsourcing platform called the Tribunal to judge whether a reported toxic player should be punished or not. The Tribunal is a two stage system requiring reports from those players that directly observe toxic behavior, and human experts that review aggregated reports. While this system has successfully dealt with the vague nature of toxic behavior by majority rules based on many votes, it naturally requires tremendous cost, time, and human efforts. In this paper, we propose a supervised learning approach for predicting crowdsourced decisions on toxic behavior with large-scale labeled data collections; over 10 million user reports involved in 1.46 million toxic players and corresponding crowdsourced decisions. Our result shows good performance in detecting overwhelmingly majority cases and predicting crowdsourced decisions on them. We demonstrate good portability of our classifier across regions. Finally, we estimate the practical implications of our approach, potential cost savings and victim protection.
Jeremy Blackburn, Haewoon Kwak
WWW2
2012 More of a Receiver Than a Giver: Why Do People Unfollow in Twitter?
Haewoon Kwak, Sue B. Moon
ICWSM1
2010 What is Twitter, a social network or a news media?
abstract
Twitter, a microblogging service less than three years old, commands more than 41 million users as of July 2009 and is growing fast. Twitter users tweet about any topic within the 140-character limit and follow others to receive their tweets. The goal of this paper is to study the topological characteristics of Twitter and its power as a new medium of information sharing.
Haewoon Kwak, Hosung Park, Sue B. Moon
WWW1
2010 Finding influentials based on the temporal order of information adoption in twitter
abstract
Twitter offers an explicit mechanism to facilitate information diffusion and has emerged as a new medium for communication. Many approaches to find influentials have been proposed, but they do not consider the temporal order of information adoption. In this work, we propose a novel method to find influentials by considering both the link structure and the temporal order of information adoption in Twitter. Our method finds distinct influentials who are not discovered by other methods.
Haewoon Kwak, Hosung Park, Sue B. Moon
WWW2
2009 Connecting Users with Similar Interests across Multiple Web Services
Haewoon Kwak, Hwa-Yong Shin, Jong-Il Yoon, Sue B. Moon
ICWSM1
2009 The wisdom of the few: a collaborative filtering approach based on expert opinions from the web
abstract
Nearest-neighbor collaborative filtering provides a successful means of generating recommendations for web users. However, this approach suffers from several shortcomings, including data sparsity and noise, the cold-start problem, and scalability. In this work, we present a novel method for recommending items to users based on expert opinions. Our method is a variation of traditional collaborative filtering: rather than applying a nearest neighbor algorithm to the user-rating data, predictions are computed using a set of expert neighbors from an independent dataset, whose opinions are weighted according to their similarity to the user. This method promises to address some of the weaknesses in traditional collaborative filtering, while maintaining comparable accuracy. We validate our approach by predicting a subset of the Netflix data set. We use ratings crawled from a web portal of expert reviews, measuring results both in terms of prediction accuracy and recommendation list precision. Finally, we explore the ability of our method to generate useful recommendations, by reporting the results of a user-study where users prefer the recommendations generated by our approach.
Xavier Amatriain, Neal Lathia, Josep M. Pujol, Haewoon Kwak, Nuria Oliver
SIGIR4
2007 Analysis of topological characteristics of huge online social networking services
abstract
Social networking services are a fast-growing business in the Internet. However, it is unknown if online relationships and their growth patterns are the same as in real-life social networks. In this paper, we compare the structures of three online social networking services: Cyworld, MySpace, and orkut, each with more than 10 million users, respectively. We have access to complete data of Cyworld's ilchon (friend) relationships and analyze its degree distribution, clustering property, degree correlation, and evolution over time. We also use Cyworld data to evaluate the validity of snowball sampling method, which we use to crawl and obtain partial network topologies of MySpace and orkut. Cyworld, the oldest of the three, demonstrates a changing scaling behavior over time in degree distribution. The latest Cyworld data's degree distribution exhibits a multi-scaling behavior, while those of MySpace and orkut have simple scaling behaviors with different exponents. Very interestingly, each of the two e ponents corresponds to the different segments in Cyworld's degree distribution. Certain online social networking services encourage online activities that cannot be easily copied in real life; we show that they deviate from close-knit online social networks which show a similar degree correlation pattern to real-life social networks.
Yong-Yeol Ahn, Seungyeop Han, Haewoon Kwak, Sue B. Moon, Hawoong Jeong
WWW3