Janith Weerasinghe

dblp:201/5389 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
4since 2021 · last 2026
0009-0003-1562-4211ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2026 LLM-Based Content Tagging at The Washington Post
abstract
We present a production LLM-based taxonomy classification system deployed at The Washington Post that tags news content across five schemas (Subject, Person, Company, Organization, Geography) using a proprietary taxonomy of ∼ 20,400 entries across seven hierarchical levels. For the Subject schema, we employ embedding-based candidate filtering followed by LLM selection. For other schemas, we combine LLM-based named entity extraction with fuzzy n-gram matching, followed by LLM selection. Comparison of post-production F1 scores against commercial vendor baselines demonstrates significant improvements across all five schemas, with the most substantial gain in Subject schema (+29.3%, p < 0.001). The system processes hundreds to thousands of articles and news items daily with a mean latency of 3–4 seconds per request and supports zero-downtime taxonomy updates.
Meng Ling, Himanshu Jahagirdar, Janith Weerasinghe, Han Jun Yoon, Suja Thomas, Anuradha Uduwage, Eui-Hong Han
UMAP3
2026 A Case Study of Offline Reinforcement Learning for Paywall Decisioning
abstract
We describe how The Washington Post deployed an offline reinforcement learning (RL) system to optimize paywall decisioning at production scale. We cast each non-subscriber article access attempt as a sequential decision with three actions: free access, registration wall, or subscription paywall, and learn policies from logged data collected via a small-traffic randomized controlled trial and subsequent production logging. We iterated from a tabular Q-learning baseline to a deep offline RL model trained with Conservative Q-Learning (CQL), using off-policy evaluation primarily to screen and rank candidates before online testing. The system was rolled out with guardrails and a persistent randomized holdout to manage risk in a revenue-critical setting. In year-long online experiments, the learned policies outperformed the legacy rules-based metering policy and improved a stakeholder-weighted value metric; the CQL policy delivered a +3% lift versus the randomized baseline while increasing subscriptions (+6%) and reducing the registration gap relative to earlier RL iterations. This case study highlights the practical steps needed to safely train, evaluate, and deploy offline RL for high-stakes personalization.
Janith Weerasinghe, Han Jun Yoon, Meng Ling, Himanshu Jahagirdar, Suja Thomas, Anuradha Uduwage, Sam Han
UMAP1
2026 Uncertainty-Aware Reinforcement Learning for Conversion-Optimized Content Gating
abstract
Publishers increasingly rely on access gates to drive registrations and subscriptions. Determining when to present these gates is a sequential decision problem well suited to reinforcement learning (RL). However, online exploration is costly and risky due to delayed conversion signals. We introduce Uncertainty-Aware Advantage-Weighted Actor–Critic (UA-AWAC), an offline RL method that learns from logged traffic to produce conversion-ready policies. UA-AWAC optimizes a multi-objective reward incorporating subscriptions, registrations, and engagement, while mitigating distribution shift through epistemic uncertainty modeling and pessimistic value targets. The policy is trained using advantage-weighted behavioral cloning with Kullback–Leibler (KL) regularization to remain close to historical gating behavior. Experiments show that UA-AWAC improves subscription rate by up to 10% and registration rate by 62% compared to baseline and state-of-the-art offline RL methods, demonstrating a practical and stable solution for intelligent content gating where exploration risks are high.
Han Jun Yoon, Janith Weerasinghe, Himanshu Jahagirdar, Meng Ling, Suja Thomas, Anuradha Uduwage, Sam Han
UMAP2
2022 Using Authorship Verification to Mitigate Abuse in Online Communities
Janith Weerasinghe, Rhia Singh, Rachel Greenstadt
ICWSM1
2020 The Pod People: Understanding Manipulation of Social Media Popularity via Reciprocity Abuse
abstract
Online Social Network (OSN) Users’ demand to increase their account popularity has driven the creation of an underground ecosystem that provides services or techniques to help users manipulate content curation algorithms. One method of subversion that has recently emerged occurs when users form groups, called pods, to facilitate reciprocity abuse, where each member reciprocally interacts with content posted by other members of the group. We collect 1.8 million Instagram posts that were posted in pods hosted on Telegram. We first summarize the properties of these pods and how they are used, uncovering that they are easily discoverable by Google search and have a low barrier to entry. We then create two machine learning models for detecting Instagram posts that have gained interaction through two different kinds of pods, achieving 0.91 and 0.94 AUC, respectively. Finally, we find that pods are effective tools for increasing users’ Instagram popularity, we estimate that pod utilization leads to a significantly increased level of likely organic comment interaction on users’ subsequent posts.
Janith Weerasinghe, Bailey Flanigan, Aviel J. Stein, Damon McCoy, Rachel Greenstadt
WWW1
2019 "Because... I was told... so much": Linguistic Indicators of Mental Health Status on Twitter
abstract
Abstract Recent studies have shown that machine learning can identify individuals with mental illnesses by analyzing their social media posts. Topics and words related to mental health are some of the top predictors. These findings have implications for early detection of mental illnesses. However, they also raise numerous privacy concerns. To fully evaluate the implications for privacy, we analyze the performance of different machine learning models in the absence of tweets that talk about mental illnesses. Our results show that machine learning can be used to make predictions even if the users do not actively talk about their mental illness. To fully understand the implications of these findings, we analyze the features that make these predictions possible. We analyze bag-of-words, word clusters, part of speech n-gram features, and topic models to understand the machine learning model and to discover language patterns that differentiate individuals with mental illnesses from a control group. This analysis confirmed some of the known language patterns and uncovered several new patterns. We then discuss the possible applications of machine learning to identify mental illnesses, the feasibility of such applications, associated privacy implications, and analyze the feasibility of potential mitigations.
Janith Weerasinghe, Kediel O. Morales, Rachel Greenstadt
Proc. Priv. Enhancing Technol.1