Evan M. Williams

dblp:285/1550 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-0534-9450ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Bridging Social Media and Search Engines: Dredge Words and the Detection of Unreliable Domains
abstract
Proactive content moderation requires platforms to rapidly and continuously evaluate the credibility of websites. Leveraging the direct and indirect paths users follow to unreliable websites, we develop a website credibility classification and discovery system that integrates both webgraph and large-scale social media contexts. We additionally introduce the concept of dredge words—terms or phrases for which unreliable domains rank highly on search engines—and provide the first exploration of their usage on social media. Our graph neural networks that combine webgraph and social media contexts generate to state-of-the-art results in website credibility classification and significantly improves the top-k identification of unreliable domains. Additionally, we release a novel dataset of dredge words, highlighting their strong connections to both social media and online commerce platforms.
Evan M. Williams, Peter Carragher, Kathleen M. Carley
ICWSM1
2025 Misinformation Resilient Search Rankings with Webgraph-Based Interventions
abstract
The proliferation of unreliable news domains on the internet has had wide-reaching negative impacts on society. We introduce and evaluate interventions aimed at reducing traffic to unreliable news domains from search engines while maintaining traffic to reliable domains. We build these interventions on the principles of fairness (penalize sites for what is in their control), generality (label/fact-check agnostic), targeted (increase the cost of adversarial behavior), and scalability (works at webscale). We refine our methods on small-scale webdata as a testbed and then generalize the interventions to a large-scale webgraph containing 93.9M domains and 1.6B edges. We demonstrate that our methods penalize unreliable domains far more than reliable domains in both settings and we explore multiple avenues to mitigate unintended effects on both the small-scale and large-scale webgraph experiments. These results indicate the potential of our approach to reduce the spread of misinformation and foster a more reliable online information ecosystem. This research contributes to the development of targeted strategies to enhance the trustworthiness and quality of search engine results, ultimately benefiting users, and the broader digital community.
Peter Carragher, Evan M. Williams, Kathleen M. Carley
ACM Trans. Intell. Syst. Technol.2
2024 Detection and Discovery of Misinformation Sources Using Attributed Webgraphs
abstract
Website reliability labels underpin almost all research in misinformation detection. However, misinformation sources often exhibit transient behavior, which makes many such labeled lists obsolete over time. We demonstrate that Search Engine Optimization (SEO) attributes provide strong signals for predicting news site reliability. We introduce a novel attributed webgraph dataset with labeled news domains and their connections to outlinking and backlinking domains. We demonstrate the success of graph neural networks in detecting news site reliability using these attributed webgraphs, and show that our baseline news site reliability classifier outperforms current SoTA methods on the PoliticalNews dataset, achieving an F1 score of 0.96. Finally, we introduce and evaluate a novel graph-based algorithm for discovering previously unknown misinformation news sources.
Peter Carragher, Evan M. Williams, Kathleen M. Carley
ICWSM2
2022 TSPA: Efficient Target-Stance Detection on Twitter
abstract
Target-stance detection on large-scale datasets is a core component of many of the most common stance detection applications. However, despite progress in recent years, stance detection research primarily occurs at the document-level on small-scale data. We propose a highly efficient Twitter Stance Propagation Algorithm (TSPA) for detecting user-level stance on Twitter that leverages the social networks of Twitter users and runs in near-linear time. We find TSPA achieves SoTA accuracy against BERT, homogenous Graph Attention Networks (GAT), and heterogenous GAT baselines. Additionally, TSPA's wall-clock time was 10x faster than our best baseline on a GPU and over 100x faster than our best baseline on a CPU.
Evan M. Williams, Kathleen M. Carley
ASONAM1
2020 Improving LDA Topic Modeling with Gamma and Simmelian Filtration
abstract
Twitter has become an important tool for communication and marketing. Topic model algorithms meant to characterize the discourse of online conversations and identify relevant audiences do not perform well for this task, despite their widespread usage. This paper proposes an iterative topic model, Gamma Filtration, and a social network-based method, Simmelian Filtration, to amplify tweet-topic probability signal and reduce noise. We demonstrate the method on a novel data set collected of European Racially and Ethnically Motivated Violent Extremist (REMVE) networks on Twitter. We find that Simmelian Filtering is most successful at reducing noise as measured by perplexity. This improves our ability to detect and monitor core conversations of a community that is disseminating propaganda to increase online extremism.
Evan M. Williams, David Levin, Ian McCulloh
ASONAM1