EDBT 2026 Demo / reviewers in the wild / expert
Deepak Kumar 0006
dblp:66/3984-6
· DBLP profile ↗
8ranked-venue papers in the field
3as first author
7since 2021 · last 2026
0000-0002-0224-5031ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 8 (3 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Long Story Short: Auditing U.S. Political Polarization in Recommendations for Long- vs. Short-form Videos on YouTubeabstractYouTube is the world's most widely used video platform, with over 70% of content viewed through algorithmic recommendations. While prior audits have examined polarization in YouTube's long-form video recommendations, the platform's fast-growing Shorts feature remains understudied. In this paper, we present the first large-scale audit comparing political content exposure and engagement dynamics across short-form and long-form videos on YouTube. We design a matched audit based on the insight that many news media organizations publish both short and long versions of the same content and collect 50,000 pairs of long-form and short-form video recommendations from both political and nonpolitcal seed videos. We analyze recommendations along several dimensions: the frequency of political recommendations, the diversity of retrieved videos, the engagement those videos receive, and finally, the partisan alignment between recommended videos and seed videos. Our results highlight fundamental differences between each algorithm, which we hope we can inform future research in analyzing the impact of YouTube recommendations. Shaokang Jiang, Arshia Arya, Seoyoung Kweon, Ivan Liang, Deepak Kumar 0006, Kristen Vaccaro |
WWW | 5 |
| 2024 | Watch Your Language: Investigating Content Moderation with Large Language ModelsabstractLarge language models (LLMs) have exploded in popularity due to their ability to perform a wide array of natural language tasks. Text-based content moderation is one LLM use case that has received recent enthusiasm, however, there is little research investigating how LLMs can help in content moderation settings. In this work, we evaluate a suite of commodity LLMs on two common content moderation tasks: rule-based community moderation and toxic content detection. For rule-based community moderation, we instantiate 95 subcommunity specific LLMs by prompting GPT-3.5 with rules from 95 Reddit subcommunities. We find that GPT-3.5 is effective at rule-based moderation for many communities, achieving a median accuracy of 64% and a median precision of 83%. For toxicity detection, we evaluate a range of LLMs (GPT-3, GPT-3.5, GPT-4, Gemini Pro, LLAMA 2) and show that LLMs significantly outperform currently widespread toxicity classifiers. However, we also found that increases in model size add only marginal benefit to toxicity detection, suggesting a potential performance plateau for LLMs on toxicity detection tasks. We conclude by outlining avenues for future work in studying LLMs and content moderation. Deepak Kumar 0006, Yousef AbuHashem, Zakir Durumeric |
ICWSM | 1 |
| 2023 | Happenstance: Utilizing Semantic Search to Track Russian State Media Narratives about the Russo-Ukrainian War on RedditabstractIn the buildup to and in the weeks following the Russian Federation’s invasion of Ukraine, Russian state media outlets output torrents of misleading and outright false information. In this work, we study this coordinated information campaign in order to understand the most prominent state media narratives touted by the Russian government to English-speaking audiences. To do this, we first perform sentence-level topic analysis using the large-language model MPNet on articles published by ten different pro-Russian propaganda websites including the new Russian “fact-checking” website waronfakes.com. Within this ecosystem, we show that smaller websites like katehon.com were highly effective at publishing topics that were later echoed by other Russian sites. After analyzing this set of Russian information narratives, we then analyze their correspondence with narratives and topics of discussion on r/Russia and 10 other political subreddits. Using MPNet and a semantic search algorithm, we map these subreddits’ comments to the set of topics extracted from our set of Russian websites, finding that 39.6% of r/Russia comments corresponded to narratives from pro-Russian propaganda websites compared to 8.86% on r/politics. Hans W. A. Hanley, Deepak Kumar 0006, Zakir Durumeric |
ICWSM | 2 |
| 2023 | "A Special Operation": A Quantitative Approach to Dissecting and Comparing Different Media Ecosystems' Coverage of the Russo-Ukrainian WarabstractThe coverage of the Russian invasion of Ukraine has varied widely between Western, Russian, and Chinese media ecosystems with propaganda, disinformation, and narrative spins present in all three. By utilizing the normalized pointwise mutual information metric, differential sentiment analysis, word2vec models, and partially labeled Dirichlet allocation, we present a quantitative analysis of the differences in coverage amongst these three news ecosystems. We find that while the Western press outlets have focused on the military and humanitarian aspects of the war, Russian media have focused on the purported justifications for the “special military operation” such as the presence in Ukraine of “bio-weapons” and “neo-nazis”, and Chinese news media have concentrated on the conflict’s diplomatic and economic consequences. Detecting the presence of several Russian disinformation narratives in the articles of several Chinese media outlets, we finally measure the degree to which Russian media has influenced Chinese coverage across Chinese outlets’ news articles, Weibo accounts, and Twitter accounts. Our analysis indicates that since the Russian invasion of Ukraine, Chinese state media outlets have increasingly cited Russian outlets as news sources and spread Russian disinformation narratives. Hans W. A. Hanley, Deepak Kumar 0006, Zakir Durumeric |
ICWSM | 2 |
| 2023 | Understanding the Behaviors of Toxic Accounts on RedditabstractToxic comments are the top form of hate and harassment experienced online. While many studies have investigated the types of toxic comments posted online, the effects that such content has on people, and the impact of potential defenses, no study has captured the behaviors of the accounts that post toxic comments or how such attacks are operationalized. In this paper, we present a measurement study of 929K accounts that post toxic comments on Reddit over an 18 month period. Combined, these accounts posted over 14 million toxic comments that encompass insults, identity attacks, threats of violence, and sexual harassment. We explore the impact that these accounts have on Reddit, the targeting strategies that abusive accounts adopt, and the distinct patterns that distinguish classes of abusive accounts. Our analysis informs the nuanced interventions needed to curb unwanted toxic behaviors online. Deepak Kumar 0006, Jeffrey T. Hancock, Kurt Thomas, Zakir Durumeric |
WWW | 1 |
| 2022 | On the Infrastructure Providers That Support Misinformation Websites
Catherine Han, Deepak Kumar 0006, Zakir Durumeric |
ICWSM | 2 |
| 2022 | No Calm in the Storm: Investigating QAnon Website Relationships
Hans W. A. Hanley, Deepak Kumar 0006, Zakir Durumeric |
ICWSM | 2 |
| 2017 | Security Challenges in an Increasingly Tangled WebabstractOver the past 20 years, websites have grown increasingly complex and interconnected. In 2016, only a negligible number of sites are dependency free, and over 90% of sites rely on external content. In this paper, we investigate the current state of web dependencies and explore two security challenges associated with the increasing reliance on external services: (1) the expanded attack surface associated with serving unknown, implicitly trusted third-party content, and (2) how the increased set of external dependencies impacts HTTPS adoption. We hope that by shedding light on these issues, we can encourage developers to consider the security risks associated with serving third-party content and prompt service providers to more widely deploy HTTPS. Deepak Kumar 0006, Zane Ma, Zakir Durumeric, Ariana Mirian, Joshua Mason, J. Alex Halderman, Michael D. Bailey |
WWW | 1 |