H. Andrew Schwartz

dblp:46/3430 · also Hansen Andrew Schwartz · DBLP profile ↗
← Back
9ranked-venue papers in the field
1as first author
5since 2021 · last 2025
0000-0002-6383-3339ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 8 (1 first)Big Data, Cloud & Distributed Data Systems · 1
YearPublicationVenuePosition
2025 Unifying the Extremes: Developing a Unified Model for Detecting and Predicting Extremist Traits and Radicalization
abstract
The proliferation of ideological movements into extremist factions via social media has become a global concern. While radicalization has been studied extensively within the context of specific ideologies, our ability to accurately characterize extremism in more generalizable terms remains underdeveloped. In this paper, we propose a novel method for extracting and analyzing extremist discourse across a range of online ideological community forums. By focusing on verbal behavioral signatures of extremist traits, we develop a framework for quantifying extremism at both user and community levels. Our research identifies 11 distinct factors, which we term "The Extremist Eleven," as a generalized psychosocial model of extremism. Applying our method to various online communities, we demonstrate an ability to characterize ideologically diverse communities across the 11 extremist traits. We demonstrate the power of this method by analyzing user histories from members of the incel community. We find that our framework accurately predicts which users join the incel community up to 10 months before their actual entry with an AUC of > 0.6, steadily increasing to AUC ~ 0.9 three to four months before the event. Further, we find that upon entry into an ideological forum, the users tend to maintain their level of extremist traits within the community, while still remaining distinguishable from the general online discourse. Our findings contribute to the study of extremism by introducing a more holistic, cross-ideological approach that transcends traditional, trait-specific models.
Allison Lahnala, Vasudha Varadarajan, Lucie Flek, H. Andrew Schwartz, Ryan L. Boyd
ICWSM4
2022 Correcting Sociodemographic Selection Biases for Population Prediction from Social Media
Salvatore Giorgi, Veronica E. Lynn, Farhan Ahmed, Sandra Matz, Lyle H. Ungar, H. Andrew Schwartz
ICWSM7
2022 Modeling Latent Dimensions of Human Beliefs
Huy Vu, Salvatore Giorgi, Jeremy D. W. Clifton, Niranjan Balasubramanian, H. Andrew Schwartz
ICWSM5
2021 Well-Being Depends on Social Comparison: Hierarchical Models of Twitter Language Suggest That Richer Neighbors Make You Less Happy
Salvatore Giorgi, Sharath Chandra Guntuku, Johannes C. Eichstaedt, Claire Pajot, H. Andrew Schwartz, Lyle H. Ungar
ICWSM5
2021 Contrastive Lexical Diffusion Coefficient: Quantifying the Stickiness of the Ordinary
abstract
Lexical phenomena, such as clusters of words, disseminate through social networks at different rates but most models of diffusion focus on the discrete adoption of new lexical phenomena (i.e. new topics or memes). It is possible much of lexical diffusion happens via the changing rates of existing word categories or concepts (those that are already being used, at least to some extent, regularly) rather than new ones. In this study we introduce a new metric, contrastive lexical diffusion (CLD) coefficient, which attempts to measure the degree to which ordinary language (here clusters of common words) catch on over friendship connections over time. For instance topics related to meeting and job are found to be sticky, while negative thinking and emotion, and global events, like ‘school orientation’ were found to be less sticky even though they change rates over time. We evaluate CLD coefficient over both quantitative and qualitative tests, studied over 6 years of language on Twitter. We find CLD predicts the spread of tweets and friendship connections, scores converge with human judgments of lexical diffusion (r=0.92), and CLD coefficients replicate across disjoint networks (r=0.85). Comparing CLD scores can help understand lexical diffusion: positive emotion words appear more diffusive than negative emotions, first-person plurals (we) score higher than other pronouns, and numbers and time appear non-contagious.
Mohammadzaman Zamani, H. Andrew Schwartz
WWW2
2020 Quantifying Community Characteristics of Maternal Mortality Using Social Media
abstract
While most mortality rates have decreased in the US, maternal mortality has increased and is among the highest of any OECD nation. Extensive public health research is ongoing to better understand the characteristics of communities with relatively high or low rates. In this work, we explore the role that social media language can play in providing insights into such community characteristics. Analyzing pregnancy-related tweets generated in US counties, we reveal a diverse set of latent topics including Morning Sickness, Celebrity Pregnancies, and Abortion Rights. We find that rates of mentioning these topics on Twitter predicts maternal mortality rates with higher accuracy than standard socioeconomic and risk variables such as income, race, and access to health-care, holding even after reducing the analysis to six topics chosen for their interpretability and connections to known risk factors. We then investigate psychological dimensions of community language, finding the use of less trustful, more stressed, and more negative affective language is significantly associated with higher mortality rates, while trust and negative affect also explain a significant portion of racial disparities in maternal mortality. We discuss the potential for these insights to inform actionable health interventions at the community-level.
Rediet Abebe, Salvatore Giorgi, Anna Tedijanto, Anneke Buffone, H. Andrew Schwartz
WWW5
2019 TV Ad Events and Digital Search: On the Selection of Outcome Measures
abstract
Prior research has shown that TV content affects what people do on the web, particularly in the minutes after a TV ad airs, when online searches for the advertised product spike. To study this, researchers have typically focused on the total volume of search queries that include any and all keywords associated with the brands and products in TV ads. We argue that a granular consideration of search queries would be beneficial for two reasons. First, focusing on relevant keywords reduces measurement error, which can hinder identification of a significant effect from TV ads. Second, by grouping queries into themes based on semantic similarity, marketers can, for example, explore consumer intent in searches, or whether response manifests in queries with objective `value'. Our nuanced proposed approach considers query activity more broadly. Leveraging data from Bing and iSpotTV for 12 product campaigns, we first demonstrate that the outcome measure (i.e., response variable) can significantly influence conclusions about whether an ad has influenced search behavior. Second, we present a theme-based difference-in-differences method to determine the associations between TV ads and individual queries and query groups. We do this by first exploring the effects of TV ads on distinct, individual queries, and then aggregating queries into themes reflecting customer intentions, instead of aggregating all queries under a product or brand name umbrella. Our approach has implications for researchers who can improve measurement in search response; for marketers who can evaluate whether and how marketing messages resonate with consumers; and for sponsored search advertisers who can determine which keywords and queries they should bid on in the moments after their TV ads air. In addition, our method can be applied outside of the advertising context to understand how user generated content is impacted by events in general. Finally, we have developed a standalone tool that has been deployed to by Marketers internal to our company to generate reports to assess the success of their TV ad campaigns and has been deployed to our advertising customers by the Bing Ads Sales team to provide business intelligence11.
Shawndra Hill, Anthony M. Colas, H. Andrew Schwartz, Gordon Burtch
IEEE BigData3
2019 Using Search Queries to Understand Health Information Needs in Africa
Rediet Abebe, Shawndra Hill, Jennifer Wortman Vaughan, Peter M. Small, H. Andrew Schwartz
ICWSM5
2013 Characterizing Geographic Variation in Well-Being Using Tweets
H. Andrew Schwartz, Johannes C. Eichstaedt, Margaret L. Kern, Lukasz Dziurzynski, Richard E. Lucas, Megha Agrawal, Gregory J. Park, Shrinidhi K. Lakshmikanth, Sneha Jha, Martin E. P. Seligman, Lyle H. Ungar
ICWSM1