Joshua A. Tucker

dblp:126/4751 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0003-1321-8650ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Hot Tweets and Cold Posts: Variation in US Congresspeople's Ideological Presentation on Twitter and Facebook Over Time
abstract
This work presents a novel observational study of US congresspeople’s link-based news-sharing behaviors and ideological presentations across Facebook and Twitter. By analyzing the web domains these politicians share, we estimate their political ideologies and measure ideological extremity across the political and social contexts of these platforms. Our findings show that these politicians present as more ideologically extreme on Facebook than they appear on Twitter, particularly among Democrats. However, this difference is relatively small compared to the ideological shift between a politician’s publicly funded official account and their campaign account—a shift that is roughly seven times larger. Finally, we observe that these changes are not uniform over time across parties; expressed polarization within the Democratic Party notably increased from 2013 to 2017 before stabilizing, while the Republican Party became markedly more polarized starting in 2020. While more research is needed to identify the specific affordances that contribute to more expressed polarization on Facebook and potential temporal dynamics between these platforms, this work highlights the limitations of studies that focus on single platforms and opens new avenues for future research into how differences across online social spaces may impact political polarization.
Kevin T. Greene, Matthew DeVerna, Joshua A. Tucker, Cody Buntain
ICWSM3
2024 Concept-Guided Chain-of-Thought Prompting for Pairwise Comparison Scoring of Texts with Large Language Models
abstract
Existing text scoring methods require a large corpus, struggle with short texts, or require hand-labeled data. We develop a text scoring framework that leverages generative large language models (LLMs) to (1) set texts against the backdrop of information from the near-totality of the web and digitized media, and (2) effectively transform pairwise text comparisons from a reasoning problem to a pattern recognition task. Our approach, concept-guided chain-of-thought (CGCoT), utilizes a chain of researcher-designed prompts with an LLM to generate a concept-specific breakdown for each text, akin to guidance provided to human coders. We then pairwise compare breakdowns using an LLM and aggregate answers into a score using a probability model. We apply this approach to better understand speech reflecting aversion to specific political parties on Twitter, a topic that has commanded increasing interest because of its potential contributions to democratic backsliding. We achieve stronger correlations with human judgments than widely used unsupervised text scoring methods like Wordfish. In a supervised setting, besides a small pilot dataset to develop CGCoT prompts, our measures require no additional hand-labeled data and produce predictions on par with RoBERTa-Large fine-tuned on thousands of hand-labeled tweets. This project showcases the potential of combining human expertise and LLMs for scoring tasks.
Patrick Y. Wu, Jonathan Nagler, Joshua A. Tucker, Solomon Messing
IEEE Big Data3
2023 Measuring the Ideology of Audiences for Web Links and Domains Using Differentially Private Engagement Data
abstract
This paper demonstrates the use of differentially private hyperlink-level engagement data for measuring ideologies of audiences for web domains, individual links, or aggregations thereof. We examine a simple metric for measuring this ideological position and assess the conditions under which the metric is robust to injected, privacy-preserving noise. This assessment provides insights into and constraints on the level of activity one should observe when applying this metric to privacy-protected data. Grounding this work is a massive dataset of social media engagement activity where privacy-preserving noise has been injected into the activity data, provided by Facebook and the Social Science One (SS1) consortium. Using this dataset, we validate our ideology measures by comparing to similar, published work on sharing-based, homophily- and content-oriented measures, where we show consistently high correlation (>0.87). We then apply this metric to individual links from several popular news domains and demonstrate how one can assess link-level distributions of ideological audiences. We further show this estimator is robust to selection of engagement types besides sharing, where domain-level audience-ideology assessments based on views and likes show no significant difference compared to sharing-based estimates. Estimates of partisanship, however, suggest the viewing audience is more moderate than the audiences who share and like these domains. Beyond providing thresholds on sufficient activity for measuring audience ideology and comparing three types of engagement, this analysis provides a blueprint for ensuring robustness of future work to differential privacy protections.
Cody Buntain, Richard Bonneau, Jonathan Nagler, Joshua A. Tucker
ICWSM4
2022 Dictionary-Assisted Supervised Contrastive Learning
abstract
Text analysis in the social sciences often involves using specialized dictionaries to reason with abstract concepts, such as perceptions about the economy or abuse on social media.These dictionaries allow researchers to impart domain knowledge and note subtle usages of words relating to a concept(s) of interest.We introduce the dictionary-assisted supervised contrastive learning (DASCL) objective, allowing researchers to leverage specialized dictionaries when fine-tuning pretrained language models.The text is first keyword simplified: a common, fixed token replaces any word in the corpus that appears in the dictionary(ies) relevant to the concept of interest.During fine-tuning, a supervised contrastive objective draws closer the embeddings of the original and keyword-simplified texts of the same class while pushing further apart the embeddings of different classes.The keyword-simplified texts of the same class are more textually similar than their original text counterparts, which additionally draws the embeddings of the same class closer together.Combining DASCL and crossentropy improves classification performance metrics in few-shot learning settings and social science applications compared to using crossentropy alone and alternative contrastive and data augmentation methods. 1
Patrick Y. Wu, Richard Bonneau, Joshua A. Tucker, Jonathan Nagler
EMNLP3
2021 YouTube Recommendations and Effects on Sharing Across Online Social Platforms
abstract
In January 2019, YouTube announced its platform would exclude potentially harmful content from video recommendations while allowing such videos to remain on the platform. While this action is intended to reduce YouTube's role in propagating such content, continued availability of these videos via hyperlinks in other online spaces leaves an open question of whether such actions actually impact sharing of these videos in the broader information space. This question is particularly important as other online platforms deploy similar suppressive actions that stop short of deletion despite limited understanding of such actions' impacts. To assess this impact, we apply interrupted time series models to measure whether sharing of potentially harmful YouTube videos in Twitter and Reddit changed significantly in the eight months around YouTube's announcement. We evaluate video sharing across three curated sets of anti-social content: a set of conspiracy videos that have been shown to experience reduced recommendations in YouTube, a larger set of videos posted by conspiracy-oriented channels, and a set of videos posted by alternative influence network (AIN) channels. As a control, we also evaluate these effects on a dataset of videos from mainstream news channels. Results show conspiracy-labeled and AIN videos that have evidence of YouTube's de-recommendation do experience a significant decreasing trend in sharing on both Twitter and Reddit. At the same time, however, videos from conspiracy-oriented channels actually experience a significant increase in sharing on Reddit following YouTube's intervention, suggesting these actions may have unintended consequences in pushing less overtly harmful conspiratorial content. Mainstream news sharing likewise sees increases in trend on both platforms, suggesting YouTube's suppression of particular content types has a targeted effect. In summary, while this work finds evidence that reducing exposure to anti-social videos within YouTube potentially reduces sharing on other platforms, increases in the level of conspiracy-channel sharing raise concerns about how producers -- and consumers -- of harmful content are responding to YouTube's changes. Transparency from YouTube and other platforms implementing similar strategies is needed to evaluate these effects further.
Cody Buntain, Richard Bonneau, Jonathan Nagler, Joshua A. Tucker
Proc. ACM Hum. Comput. Interact.4