EDBT 2026 Demo / reviewers in the wild / expert
Prathyusha Senthil Kumar
dblp:185/8133
· DBLP profile ↗
8ranked-venue papers
1as first author
5since 2021 · last 2024
0000-0002-3907-9399ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Integrity 2024: Integrity in Social Networks and MediaabstractIntegrity 2024 is the fifth edition of the Workshop on Integrity in Social Networks and Media, held in conjunction with the ACM Conference on Web Search and Data Mining (WSDM) since the 2020 edition [1-4]. The goal of the workshop is to bring together academic and industry researchers working on integrity, fairness, trust and safety in social networks to discuss the most pressing risks and cutting-edge technologies to reliably measure and mitigate them. The event consists of invited talks from academic experts and industry leaders as well as peer-reviewed papers and posters through an open call-for-papers. Lluís Garcia Pueyo, Symeon Papadopoulos, Prathyusha Senthil Kumar, Aristides Gionis, Panayiotis Tsaparas, Vasilis Verroios, Giuseppe Manco 0001, Anton Andryeyev, Stefano Cresci, Timos K. Sellis, Anthony McCosker |
WSDM | 3 |
| 2023 | Practical Lessons Learned From Detecting, Preventing and Mitigating Harmful Experiences on FacebookabstractSocial media's explosive growth brings with it a variety of societal risks ranging from severely harmful issues such as dangerous organizations and child sexual exploitation to moderately harmful content like displays of aggression, borderline nudity to benign or distasteful contents like gross videos and baity content. In recent times, the multitude and magnitude of these harms is being further exacerbated with the advent of generative AI [5]. Meta is committed to ensuring that Facebook is a place where people feel empowered to communicate and we take our role seriously in keeping abuse off the platform [7]. In this talk, I will describe practical challenges and lessons learned from tackling bad experiences for users on Facebook, particularly in the subjective, borderline and low quality spectrum of harms using state of the art, scalable machine learning approaches to content understanding, user behavior understanding and personalized ranking. Prathyusha Senthil Kumar |
CIKM | 1 |
| 2023 | Integrity 2023: Integrity in Social Networks and MediaabstractIntegrity 2023 is the fourth edition of the successful Workshop on Integrity in Social Networks and Media, held in conjunction with the ACM Conference on Web Search and Data Mining (WSDM) in the past three years. The goal of the workshop is to bring together researchers and practitioners to discuss content and interaction integrity challenges in social networks and social media platforms. The event consists of a combination of invited talks by reputed members of the Integrity community from both academia and industry and peer-reviewed contributed talks and posters solicited via an open call-for-papers. Lluís Garcia Pueyo, Panayiotis Tsaparas, Prathyusha Senthil Kumar, Timos K. Sellis, Paolo Papotti, Sibel Adali, Giuseppe Manco 0001, Tudor Trufinescu, Gireeja Ranade, James R. Verbus, Mehmet N. Tek, Anthony McCosker |
WSDM | 3 |
| 2023 | Detecting and Limiting Negative User Experiences in Social Media PlatformsabstractItem ranking is important to a social media platform’s success. The order in which posts, videos, messages, comments, ads, used products, notifications are presented to a user greatly affects the time spent on the platform, how often they visit it, how much they interact with each other, and the quantity and quality of the content they post. To this end, item ranking algorithms use models that predict the likelihood of different events, e.g., the user liking, sharing, commenting on a video, clicking/converting on an ad, or opening the platform’s app from a notification. Unfortunately, by solely relying on such event-prediction models, social media platforms tend to over optimize for short-term objectives and ignore the long-term effects. In this paper, we propose an approach that aims at improving item ranking long-term impact. The approach primarily relies on an ML model that predicts negative user experiences. The model utilizes all available UI events: the details of an action can reveal how positive or negative the user experience has been; for example, a user writing a lengthy report asking for a given video to be taken down, likely had a very negative experience. Furthermore, the model takes into account detected integrity (e.g., hostile speech or graphic violence) and quality (e.g., click or engagement bait) issues with the content. Note that those issues can be perceived very differently from different users. Therefore, developing a personalized model, where a prediction refers to a specific user for a specific piece of content at a specific point in time, is a fundamental design choice in our approach. Besides the personalized ML model, our approach consists of two more pieces: (a) the way the personalized model is integrated with an item ranking algorithm and (b) the metrics, methodology, and success criteria for the long term impact of detecting and limiting negative user experiences. Our evaluation process uses extensive A/B testing on the Facebook platform: we compare the impact of our approach in treatment groups against production control groups. The AB test results indicate a 5% to 50% reduction in hides, reports, and submitted feedback. Furthermore, we compare against a baseline that does not include some of the crucial elements of our approach: the comparison shows our approach has a 100x to 30x lower False Positive Ratio than a baseline. Lastly, we present the results from a large scale survey, where we observe a statistically significant improvement of 3 to 6 percent in users’ sentiment regarding content suffering from nudity, clickbait, false / misleading, witnessing-hate, and violence issues. Lluís Garcia Pueyo, Vinodh Kumar Sunkara, Prathyusha Senthil Kumar, Mohit Diwan, Behrang Javaherian, Vasilis Verroios |
WWW | 3 |
| 2021 | Integrity 2021: Integrity in Social Networks and MediaabstractThe second Workshop on Integrity in Social Networks and Media is held in conjunction with the 14th ACM Conference on Web Search and Data Mining (WSDM) in Jerusalem, Israel. The goal of the workshop is to bring together researchers and practitioners to discuss content and interaction integrity challenges in social networks and social media platforms. Lluís Garcia Pueyo, Anand Bhaskar, Roelof van Zwol, Timos K. Sellis, Gireeja Ranade, Prathyusha Senthil Kumar, Yu Sun 0021, Joy Zhang |
WSDM | 6 |
| 2019 | Contextual Price Features for e-Commerce Search RankingabstractMotivation The price of an item is one of the most useful attributes for e-commerce search. A prominent feature used by the machine learned ranker at eBay search measures, for each query, the deviance of the price of a candidate item to be ranked from the typical price distribution of the clicked items for that query. Like most e-commerce websites, in eBay search, users can restrict their search queries to certain specific categories, where category refers to nodes in the product catalog structure. As an example, a user could search for mens shirts and restrict results to either T-Shirts or Dress Shirts. In its most simple form, the price deviance feature is calculated in a category context agnostic fashion. However, the distribution from which the query price statistics (e.g. median, variance) are drawn varies significantly for different categories for the same query. For example, the query iphone x will have very different price distribution among clicked items in the Cell Phones Smartphones category as opposed to the Cell Phone Accessories category. The motivation of this work stems from accounting for category context for user queries while deriving and leveraging such distributional statistics. The core idea is extendable to other features used in the current search ranker besides item price, such as query-title similarity measure, and also to other contexts besides category, such as structured key value aspects. Problem Statement Given a query and a set of matching items, the machine learned search ranking model uses an array of diverse features to rank items. The problem addressed in this work is to enhance one such feature capturing item price deviance from query level price statistics to be sensitive to the category context of queries. Concretely, the ranker should be able to use different price distributions for the same query depending on the category context in which the query was issued. This problem has intriguing challenges in two major areas: first, since we have a hierarchical category tree structure, we need to figure out an effective way to propagate the price distribution across the tree; second, selecting the right statistic to measure the price deviation of candidate items from the typical price distribution for the query. Ishita K. Khan, Aritra Mandal, Prathyusha Senthil Kumar |
IEEE BigData | 3 |
| 2018 | The Title Says It All: A Title Term Weighting Strategy For eCommerce RankingabstractSearch relevance is a very critical component in e-commerce applications. One of the strongest signals that determine the relevance of an item listing to an e-commerce query is the title of the item. Traditional methods for capturing this signal compare words in listing titles and the user query using tf-idf scores, or use a machine learned model with words as features and target clicks or relevance labels. Contrary to these approaches, we build a parameterized model to determine the weights of popular title terms for a query and then use these title term weights to compute the relevance of a listing title to the query. For this, we use human judged binary relevance labels of query and item title pairs as labeled data and train a model leveraging a variety of features to learn these query specific title term weights. We propose two novel approaches to model these title term weights using the relevance target and explore several novel features specific to e-commerce for this term weighting model. We use the resulting title relevance score as a feature in eBay's machine learned ranker for e-commerce search serving millions of queries each day. We observe a significant improvement over a baseline click-based binary independence model for capturing item title relevance in several metrics including model accuracy and overall relevance and engagement observed through A/B testing. We also experimentally illustrate that this feature optimized for relevance works well in conjunction with textual features optimized for demand. Anthony Bell, Prathyusha Senthil Kumar, Daniel Miranda |
CIKM | 2 |
| 2017 | What is skipped: Finding desirable items in e-commerce search by discovering the worst title tokensabstractGiven an ecommerce query, how well the titles of items for sale match the user intent is an important signal for ranking the items. A well-known technique for computing this signal is to use a standard machine-learned model that uses words as features, targets user clicks and predicts a score to rank the titles. In this paper, we introduce an alternate modeling technique that applies to queries that are frequent enough to have historical click data. For each such query we build a parameterized model of user behavior that learns what makes users skip a title. The parameters are different for each query. Specifically, our model predicts how desirable an item's title is to the user query by focusing on the worst tokens in the title. The model is learned offline using maximum likelihood based on user behavioral data, significantly improving query processing cost. The model's output score is used as a feature in a machine learned ranker for e-commerce search at eBay. Besides titles, the model design can easily incorporate any attribute of an item including structured content. In this scope, we present our new title desirability model built for nearly 8M queries recently deployed into the eBay search ecosystem and demonstrate its significant performance improvement over a baseline click-based Nïve Bayes model through different evaluation approaches including A/B testing and human judgment. The reported performance is based on eBays commercial search engine serving millions of queries each day. Ishita K. Khan, Prathyusha Senthil Kumar, Daniel Miranda, David Goldberg 0001 |
IEEE BigData | 2 |