Kenneth Joseph

dblp:126/6273 · DBLP profile ↗
← Back
17ranked-venue papers in the field
5as first author
8since 2021 · last 2025
0000-0003-2233-3976ORCID · corroborated

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 12 (4 first)Data Mining & Knowledge Discovery · 3 (1 first)Database Systems & Data Management · 2
YearPublicationVenuePosition
2025 Measuring Dimensions of Self-Presentation in Twitter Bios and their Links to Misinformation Sharing
abstract
Social media platforms provide users with a profile description field, commonly known as a "bio," where they can present themselves to the world. A growing literature shows that text in these bios can improve our understanding of online self-presentation and behavior, but existing work relies exclusively on keyword-based approaches to do so. We here propose and evaluate a suite of simple, effective, and theoretically motivated approaches to embed bios in spaces that capture salient dimensions of social meaning, such as age and partisanship. We evaluate our methods on four tasks, showing that the strongest one out-performs several practical baselines. We then show the utility of our method in helping understand associations between self-presentation and the sharing of URLs from low-quality news sites on Twitter, with a particular focus on explore the interactions between age and partisanship, and exploring the effects of self-presentations of religiosity. Our work provides new tools to help computational social scientists make use of information in bios, and provides new insights into how misinformation sharing may be perceived on Twitter.
Navid Madani, Rabiraj Bandyopadhyay, Michael Miller Yoder, Stephan D. McCabe, Briony Swire-Thompson, Kenneth Joseph
ICWSM6
2025 GALLOC: a GeoAnnotator for Labeling LOCation descriptions from disaster-related text messages
abstract
During a natural disaster, people post text messages on various platforms, such as social media and short message service (SMS) platforms, to share urgent information and seek help. Many text messages contain location descriptions about victims and accidents. Accurately extracting these location descriptions can help disaster responders reach victims more quickly and even save lives. These location descriptions, however, are often more complex than simple place names (e.g. city names), and cannot be extracted using typical named entity recognition approaches. While new machine learning models could be trained, they require labeled training data that are time-consuming to create without an effective data annotation tool. To fill this gap, we develop GALLOC, a GeoAnnotator for Labeling LOCation descriptions from disaster-related text messages. GALLOC is an open-source and Web-based tool that provides a variety of functions for supporting location description annotation, such as artificial intelligence powered pre-annotation and automatic spatial footprint identification. It also supports multilingual data annotation, and can be used by a group of users to collaboratively create a dataset. We present the design considerations and functions of GALLOC and evaluate it via a comparison with previous tools and an experiment to annotate a small set of disaster-related messages.
Kai Sun 0009, Yingjie Hu 0001, Kenneth Joseph, Ryan Zhenqi Zhou
Int. J. Geogr. Inf. Sci.3
2024 Curated and Asymmetric Exposure: A Case Study of Partisan Talk during COVID on Twitter
abstract
Social media has been at the center of discussions about political polarization in the United States. However, scholars are actively debating both the scale of political polarization online, and how important online polarization is to the offline world. One question at the center of this debate is what interactions across parties look like online, and in particular 1) whether increasing the number of such interactions is likely to increase or reduce polarization, and 2) what technological affordances may make it more likely that these cross-party interactions benefit, rather than detract from, existing political challenges. The present work aims to provide insights into the latter; that is, we focus on providing a better understanding of how a set of 400,000 partisan users on a particular social media platform, Twitter, used the platform's affordances to interact within and across parties in a large dataset of tweets about COVID in 2021. Our findings suggest that Republican use of cross-party interaction were both more potent and potentially more strategic during COVID, that cross-party interaction was driven heavily by a small set of users and conversations, and that there exist non-obvious indirect pathways to cross-party exposure when different modes of interaction are chained together (especially retweets of quotes). These findings have implications beyond Twitter, we believe, in understanding how affordances of platforms can help to shape partisan exposure and interaction.
Zijian An, Jessica Breuhaus, Jason Niu, Ahmet Erdem Sariyüce, Kenneth Joseph
ICWSM5
2023 Just Another Day on Twitter: A Complete 24 Hours of Twitter Data
abstract
At the end of October 2022, Elon Musk concluded his acquisition of Twitter. In the weeks and months before that, several questions were publicly discussed that were not only of interest to the platform's future buyers, but also of high relevance to the Computational Social Science research community. For example, how many active users does the platform have? What percentage of accounts on the site are bots? And, what are the dominating topics and sub-topical spheres on the platform? In a globally coordinated effort of 80 scholars to shed light on these questions, and to offer a dataset that will equip other researchers to do the same, we have collected all 375 million tweets published within a 24-hour time period starting on September 21, 2022. To the best of our knowledge, this is the first complete 24-hour Twitter dataset that is available for the research community. With it, the present work aims to accomplish two goals. First, we seek to answer the aforementioned questions and provide descriptive metrics about Twitter that can serve as references for other researchers. Second, we create a baseline dataset for future research that can be used to study the potential impact of the platform's ownership change.
Jürgen Pfeffer, Daniel Matter, Kokil Jaidka, Onur Varol, Afra J. Mashhadi, Jana Lasser, Dennis Assenmacher, Diyi Yang, Cornelia Brantner, Daniel M. Romero, Jahna Otterbacher, Carsten Schwemmer, Kenneth Joseph, David García 0001, Fred Morstatter
ICWSM14
2023 Geo-knowledge-guided GPT models improve the extraction of location descriptions from disaster-related social media messages
abstract
Social media messages posted by people during natural disasters often contain important location descriptions, such as the locations of victims. Recent research has shown that many of these location descriptions go beyond simple place names, such as city names and street names, and are difficult to extract using typical named entity recognition (NER) tools. While advanced machine learning models could be trained, they require large labeled training datasets that can be time-consuming and labor-intensive to create. In this work, we propose a method that fuses geo-knowledge of location descriptions and a Generative Pre-trained Transformer (GPT) model, such as ChatGPT and GPT-4. The result is a geo-knowledge-guided GPT model that can accurately extract location descriptions from disaster-related social media messages. Also, only 22 training examples encoding geo-knowledge are used in our method. We conduct experiments to compare this method with nine alternative approaches on a dataset of tweets from Hurricane Harvey. Our method demonstrates an over 40% improvement over typically used NER approaches. The experiment results also show that geo-knowledge is indispensable for guiding the behavior of GPT models. The extracted location descriptions can help disaster responders reach victims more quickly and may even save lives.
Yingjie Hu 0001, Gengchen Mai, Chris Cundy, Kristy Choi, Ni Lao, Gaurish Lakhanpal, Ryan Zhenqi Zhou, Kenneth Joseph
Int. J. Geogr. Inf. Sci.9
2022 NELA-Local: A Dataset of U.S. Local News Articles for the Study of County-Level News Ecosystems
Benjamin D. Horne, Maurício Gruppi, Kenneth Joseph, Jon Green, John Wihbey, Sibel Adali
ICWSM3
2022 Local News Online and COVID in the U.S.: Relationships among Coverage, Cases, Deaths, and Audience
Kenneth Joseph, Benjamin D. Horne, Jon Green, John Wihbey
ICWSM1
2021 An Analysis of Replies to Trump's Tweets
Zijian An, Kenneth Joseph
ICWSM2
2020 Understanding Visual Memes: An Empirical Analysis of Text Superimposed on Memes Shared on Twitter
Muhammad Aamir Masood, Kenneth Joseph
ICWSM3
2019 Polarized, Together: Comparing Partisan Support for Trump's Tweets Using Survey and Platform-Based Measures
Kenneth Joseph, Briony Swire-Thompson, Hannah Masuga, Matthew A. Baum, David Lazer
ICWSM1
2017 "Voters of the Year": 19 Voters Who Were Unintentional Election Poll Sensors on Twitter
William Hobbs, Lisa Friedland, Kenneth Joseph, Oren Tsur, Stefan Wojcik, David Lazer
ICWSM3
2016 Identifying Police Officers at Risk of Adverse Events
abstract
Adverse events between police and the public, such as deadly shootings or instances of racial profiling, can cause serious or deadly harm, damage police legitimacy, and result in costly litigation. Evidence suggests these events can be prevented by targeting interventions based on an Early Intervention System (EIS) that flags police officers who are at a high risk for involvement in such adverse events. Today's EIS are not data-driven and typically rely on simple thresholds based entirely on expert intuition. In this paper, we describe our work with the Charlotte-Mecklenburg Police Department (CMPD) to develop a machine learning model to predict which officers are at risk for an adverse event. Our approach significantly outperforms CMPD's existing EIS, increasing true positives by ~12% and decreasing false positives by ~32%. Our work also sheds light on features related to officer characteristics, situational factors, and neighborhood factors that are predictive of adverse events. This work provides a starting point for police departments to take a comprehensive, data-driven approach to improve policing and reduce harm to both officers and members of the public.
Samuel Carton, Jennifer Helsby, Kenneth Joseph, Ayesha Mahmud, Joe Walsh, Crystal Cody, C. P. T. Estella Patterson, Lauren Haynes, Rayid Ghani
KDD3
2016 Exploring Patterns of Identity Usage in Tweets: A New Problem, Solution and Case Study
abstract
Sociologists have long been interested in the ways that identities, or labels for people, are created, used and applied across various social contexts. The present work makes two contributions to the study of identity, in particular the study of identity in text. We first consider the following novel NLP task: given a set of text data (here, from Twitter), label each word in the text as being representative of a (possibly multi-word) identity. To address this task, we develop a comprehensive feature set that leverages several avenues of recent NLP work on Twitter and use these features to train a supervised classifier. Our model outperforms a surprisingly strong rule-based baseline by 33%. We then use our model for a case study, applying it to a large corpora of Twitter data from users who actively discussed the Eric Garner and Michael Brown cases. Among other findings, we observe that the identities used by individuals differ in interesting ways based on social context measures derived from census data.
Kenneth Joseph, Wei Wei 0019, Kathleen M. Carley
WWW1
2015 The Fragility of Twitter Social Networks Against Suspended Users
abstract
Social media is rapidly becoming one of the mediums of choice for understanding the cultural pulse of a region; e.g., for identifying what the population is concerned with and what kind of help is needed in a crisis. To assess this cultural pulse it is critical to have an accurate assessment of who is saying what in social media. However, social media is also the home of malicious users engaged in disruptive, disingenuous, and potentially illegal activity. A range of users, both human and non-human, carry out such social cyber-attacks. We ask, to what extent does the presence or absence of such users influence our ability to assess the cultural pulse of a region? We conduct a series of experiments to analyze the fragility of social network assessments based on Twitter data by comparing changes in both the structural and content results when suspended users are left in and taken out. Because a Twitter account can be suspended for various reasons including spamming or spreading ideas that can lead to extremism or terrorism, we separately assess the impacts of removing apparent spam bots and apparent extremists. Experimental results demonstrate that Twitter-based network structures and content are unstable, and can be highly impacted by the removal of suspended users. Further, the results exhibit regional and temporal variation that may be related to the political situation or civil unrest. We also provides guidance on the differential impact of different types of potentially suspendable users.
Wei Wei 0019, Kenneth Joseph, Huan Liu 0001, Kathleen M. Carley
ASONAM2
2015 Culture, Networks, Twitter and foursquare: Testing a Model of Cultural Conversion with Social Media Data
Kenneth Joseph, Kathleen M. Carley
ICWSM1
2015 A Bayesian Graphical Model to Discover Latent Events from Twitter
Wei Wei 0019, Kenneth Joseph, Wei Lo, Kathleen M. Carley
ICWSM2
2014 Check-ins in "Blau Space": Applying Blau's Macrosociological Theory to Foursquare Check-ins from New York City
abstract
Peter Blau was one of the first to define a latent social space and utilize it to provide concrete hypotheses. Blau defines social structure via social “parameters” (constraints). Actors that are closer together (more homogenous) in this social parameter space are more likely to interact. One of Blau’s most important hypotheses resulting from this work was that the consolidation of parameters could lead to isolated social groups. For example, the consolidation of race and income might lead to segregation. In the present work, we use Foursquare data from New York City to explore evidence of homogeneity along certain social parameters and consolidation that breeds social isolation in communities of locations checked in to by similar users. More specifically, we first test the extent to which communities detected via Latent Dirichlet Allocation are homogenous across a set of four social constraints—racial homophily, income homophily, personal interest homophily and physical space. Using a bootstrapping approach, we find that 14 (of 20) communities are statistically, and all but one qualitatively, homogenous along one of these social constraints, showing the relevance of Blau’s latent space model in venue communities determined via user check-in behavior. We then consider the extent to which communities with consolidated parameters, those homogenous on more than one parameter, represent socially isolated populations. We find communities homogenous on multiple parameters, including a homosexual community and a “hipster” community, that show support for Blau’s hypothesis that consolidation breeds social isolation. We consider these results in the context of mediated communication, in particular in the context of self-representation on social media.
Kenneth Joseph, Kathleen M. Carley, Jason I. Hong
ACM Trans. Intell. Syst. Technol.1