Maheswaran Sathiamoorthy

dblp:59/7170 · also Mahesh Sathiamoorthy · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
8since 2021 · last 2025
0009-0005-2192-3423ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 5 · 3 first-author
YearPublicationVenuePosition
2025 Never Miss an Episode: How LLMs are Powering Serial Content Discovery on YouTube
abstract
Leveraging large language models (LLMs) through prompting presents a cost-effective approach to build scalable systems without traditional model training. This paper showcases the effectiveness of using simple few-shot LLM prompt to develop a scalable and easily maintainable system that addresses a real-world user need. A critical user journey in video recommendation is watching serial content, which requires viewing episodes in a specific sequence. The existing method on YouTube for identifying serial content relied on manual creator tagging of playlists or inflexible regular expressions. These methods proved difficult to maintain and scale, limiting the system's ability to effectively identify and recommend serial content. This paper demonstrates that a carefully designed few-shot LLM prompt can accurately identify serial playlists at scale, improving user experience with minimal engineering. The paper details the challenges and lessons learned in developing and deploying this prompting-based system.
Aditee Kumthekar, Andrea Bettale, Maheswaran Sathiamoorthy, Zrinka Puljiz, Aditya Mahajan
RecSys4
2025 Toward Holistic Evaluation of Recommender Systems Powered by Generative Models
abstract
Recommender systems powered by generative models (Gen-RecSys) extend beyond classical item-ranking by producing open-ended content, which simultaneously unlocks richer user experiences and introduces new risks. On one hand, these systems can enhance personalization and appeal through dynamic explanations and multi-turn dialogues. On the other hand, they might venture into unknown territory-hallucinating nonexistent items, amplifying bias, or leaking private information. Traditional accuracy metrics cannot fully capture these challenges, as they fail to measure factual correctness, content safety, or alignment with user intent.
Yashar Deldjoo, Nikhil Mehta 0002, Maheswaran Sathiamoorthy, Shuai Zhang 0007, Pablo Castells, Julian J. McAuley
SIGIR3
2025 Tutorial on Recommendation with Generative Models (Gen-RecSys)
abstract
This intermediate-level tutorial, titled "Gen-RecSys", merges both industrial and academic perspectives on recent advances in Generative AI for recommender systems (beyond LLMs). It aims to highlight the transformative role of generative models in modern recommender systems, which have significantly impacted the AI field-particularly with the rise of large language models (LLMs) like ChatGPT-and have contributed to a rapid convergence of the fields of search, data mining, and recommendation. By providing attendees with a modern perspective on GenAI applications in recommendation, the tutorial will emphasize how generative models can drive recommendation by unlocking and interacting with rich data representations, including behavioral, textual, and multi-modal data-knowledge highly transferable across many applications of interest to the WSDM community. Participants will learn about the categorization of generative models in recommender systems based on underlying data modalities: (i) ID-based collaborative models, (ii) text-driven models such as LLMs, and (iii) multi-modal models. Within each category, various deep generative model paradigms (e.g., AR, GAN, diffusion models) will be introduced, along with insights into their application areas. The tutorial will also cover evaluation aspects, including benchmarks, metrics, and assessments of social and ethical impacts and harms. This tutorial presents a condensed version of the industrial and academic work featured in the forthcoming book at FntIR 2024-25, titled "Recommendation with Generative Models [7]," and a shorter version prepared, and presented by the team, see GenRecSys-Survey [6].
Yashar Deldjoo, Zhankui He, Julian J. McAuley, Anton Korikov, Scott Sanner, Arnau Ramisa, René Vidal, Maheswaran Sathiamoorthy, Atoosa Kasirzadeh, Silvia Milano
WSDM8
2024 A Review of Modern Recommender Systems Using Generative Models (Gen-RecSys)
abstract
Traditional recommender systems typically use user-item rating histories as their main data source. However, deep generative models now have the capability to model and sample from complex data distributions, including user-item interactions, text, images, and videos, enabling novel recommendation tasks. This comprehensive, multidisciplinary survey connects key advancements in RS using Generative Models (Gen-RecSys), covering: interaction-driven generative models; the use of large language models (LLM) and textual data for natural language recommendation; and the integration of multimodal models for generating and processing images/videos in RS. Our work highlights necessary paradigms for evaluating the impact and harm of Gen-RecSys and identifies open challenges. This survey accompanies a "tutorial" presented at ACM KDD'24, with supporting materials provided at: https://encr.pw/vDhLq.
Yashar Deldjoo, Zhankui He, Julian J. McAuley, Anton Korikov, Scott Sanner, Arnau Ramisa, René Vidal, Maheswaran Sathiamoorthy, Atoosa Kasirzadeh, Silvia Milano
KDD8
2024 Better Generalization with Semantic IDs: A Case Study in Ranking for Recommendations
abstract
Randomly-hashed item ids are used ubiquitously in recommendation models. However, the learned representations from random hashing prevents generalization across similar items, causing problems of learning unseen and long-tail items, especially when item corpus is large, power-law distributed, and evolving dynamically. In this paper, we propose using content-derived features as a replacement for random ids. We show that simply replacing ID features with content-based embeddings can cause a drop in quality due to reduced memorization capability. To strike a good balance of memorization and generalization, we propose to use Semantic IDs [15], a compact and discrete item representation, as a replacement for random item ids. Semantic IDs are learned from frozen content embeddings using RQ-VAE and thus can capture the hierarchy of concepts in items. Similar to content embeddings, the compactness of Semantic IDs poses a problem of adaption in recommendation models. We propose novel methods for adapting Semantic IDs in industry-scale ranking models, through hashing sub-pieces of of the Semantic-ID sequences. In particular, we find that the SentencePiece model [10] that is commonly used in LLM tokenization outperforms manually crafted pieces such as N-grams. To the end, we evaluate our approaches in a real-world ranking model for YouTube recommendations. Our experiments demonstrate that Semantic IDs can replace the direct use of video IDs by improving the generalization ability on new and long-tail item slices without sacrificing overall model quality.
Anima Singh, Trung Vu 0002, Nikhil Mehta 0002, Raghunandan H. Keshavan, Maheswaran Sathiamoorthy, Yilin Zheng, Lichan Hong, Lukasz Heldt, Devansh Tandon, Ed H. Chi, Xinyang Yi
RecSys5
2023 Improving Training Stability for Multitask Ranking Models in Recommender Systems
abstract
Recommender systems play an important role in many content platforms. While most recommendation research is dedicated to designing better models to improve user experience, we found that research on stabilizing the training for such models is severely under-explored. As recommendation models become larger and more sophisticated, they are more susceptible to training instability issues, i.e., loss divergence, which can make the model unusable, waste significant resources and block model developments. In this paper, we share our findings and best practices we learned for improving the training stability of a real-world multitask ranking model for YouTube recommendations. We show some properties of the model that lead to unstable training and conjecture on the causes. Furthermore, based on our observations of training dynamics near the point of training instability, we hypothesize why existing solutions would fail, and propose a new algorithm to mitigate the limitations of existing solutions. Our experiments on YouTube production dataset show the proposed algorithm can significantly improve training stability while not compromising convergence, comparing with several commonly used baseline methods.
Jiaxi Tang, Yoel Drori, Daryl Chang, Maheswaran Sathiamoorthy, Justin Gilmer, Xinyang Yi, Lichan Hong, Ed H. Chi
KDD4
2023 Recommender Systems with Generative Retrieval
abstract
Modern recommender systems perform large-scale retrieval by embedding queries and item candidates in the same unified space, followed by approximate nearest neighbor search to select top candidates given a query embedding. In this paper, we propose a novel generative retrieval approach, where the retrieval model autoregressively decodes the identifiers of the target candidates. To that end, we create semantically meaningful tuple of codewords to serve as a Semantic ID for each item. Given Semantic IDs for items in a user session, a Transformer-based sequence-to-sequence model is trained to predict the Semantic ID of the next item that the user will interact with. We show that recommender systems trained with the proposed paradigm significantly outperform the current SOTA models on various datasets. In addition, we show that incorporating Semantic IDs into the sequence-to-sequence model enhances its ability to generalize, as evidenced by the improved retrieval performance observed for items with no prior interaction history.
Shashank Rajput, Nikhil Mehta 0002, Anima Singh, Raghunandan H. Keshavan, Trung Vu 0002, Lukasz Heldt, Lichan Hong, Yi Tay, Vinh Q. Tran 0002, Jonah Samost, Maciej Kula, Ed H. Chi, Maheswaran Sathiamoorthy
NeurIPS13
2021 DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task Learning
abstract
The Mixture-of-Experts (MoE) architecture is showing promising results in improving parameter sharing in multi-task learning (MTL) and in scaling high-capacity neural networks. State-of-the-art MoE models use a trainable "sparse gate'" to select a subset of the experts for each input example. While conceptually appealing, existing sparse gates, such as Top-k, are not smooth. The lack of smoothness can lead to convergence and statistical performance issues when training with gradient-based methods. In this paper, we develop DSelect-k: a continuously differentiable and sparse gate for MoE, based on a novel binary encoding formulation. The gate can be trained using first-order methods, such as stochastic gradient descent, and offers explicit control over the number of experts to select. We demonstrate the effectiveness of DSelect-k on both synthetic and real MTL datasets with up to 128 tasks. Our experiments indicate that DSelect-k can achieve statistically significant improvements in prediction and expert selection over popular MoE gates. Notably, on a real-world, large-scale recommender system, DSelect-k achieves over 22% improvement in predictive performance compared to Top-k. We provide an open-source implementation of DSelect-k.
Hussein Hazimeh 0001, Zhe Zhao 0001, Aakanksha Chowdhery, Maheswaran Sathiamoorthy, Rahul Mazumder, Lichan Hong, Ed H. Chi
NeurIPS4
2019 Recommending what video to watch next: a multitask ranking system
abstract
In this paper, we introduce a large scale multi-objective ranking system for recommending what video to watch next on an industrial video sharing platform. The system faces many real-world challenges, including the presence of multiple competing ranking objectives, as well as implicit selection biases in user feedback. To tackle these challenges, we explored a variety of soft-parameter sharing techniques such as Multi-gate Mixture-of-Experts so as to efficiently optimize for multiple ranking objectives. Additionally, we mitigated the selection biases by adopting a Wide & Deep framework. We demonstrated that our proposed techniques can lead to substantial improvements on recommendation quality on one of the world's largest video sharing platforms.
Zhe Zhao 0001, Lichan Hong, Jilin Chen, Aniruddh Nath, Shawn Andrews, Aditee Kumthekar, Maheswaran Sathiamoorthy, Xinyang Yi, Ed H. Chi
RecSys8
2014 Helper node allocation strategies for content dissemination in intermittently connected mobile networks
abstract
We formulate and address mathematically the fundamental problem of resource allocation in the form of helper nodes in disseminating multiple content in a hybrid intermittently connected mobile network under a general stochastic homogeneous contact process. We consider and solve two variations of the problem - one in which the goal is to maximize the expected demands satisfied and another in which the goal is to minimize the time taken to disseminate the contents. Besides the global optimization perspective, we also examine the problem from a game theoretic perspective in which a central agent auctions the storage to competing content providers, and show how self-interested decisions impact the social welfare.
Maheswaran Sathiamoorthy, Keyvan R. Moghadam, Bhaskar Krishnamachari, Fan Bai 0002
SECON1
2014 Optimizing Content Dissemination in Vehicular Networks with Radio Heterogeneity
abstract
Disseminating shared information to many vehicles could incur significant access fees if it relies only on unicast cellular communications. We consider the problem of efficient content dissemination over a vehicular network, in which vehicles are equipped with two kinds of radios: a high-cost low-bandwidth, long-range cellular radio, and a free high-bandwidth short-range radio. We formulate and solve an optimization problem to maximize content dissemination from the infrastructure to vehicles within a predetermined deadline while minimizing the cost associated with communicating over the cellular connection. We examine numerically the tradeoffs between cost, delay and system utility in the optimum regime. We find that, in the optimum regime, (a) system utility is more sensitive to the cost budget when the allowed delay for the dissemination is not large, (b) the system requires relatively smaller cost budget as more vehicles participate and more delay is allowed, (c) when the cost is very important, it is better not to spread the content if it needs small delay. We also develop a polynomial-time algorithm to obtain the optimal discrete solution needed in practice. Finally, we verify our analysis using real GPS traces of 632 taxis in Beijing, China.
Joon Ahn, Maheswaran Sathiamoorthy, Bhaskar Krishnamachari, Fan Bai 0002, Lin Zhang 0001
IEEE Trans. Mob. Comput.2
2014 Distributed Storage Codes Reduce Latency in Vehicular Networks
abstract
We investigate the benefits of distributed storage using erasure codes for file sharing in vehicular networks through both analysis and realistic trace-based simulations. We show that the key parameter affecting the on-demand file download latency is the ratio of file size to download bandwidth. When this ratio is small so that a file can be communicated in a single encounter, we find that coding techniques offer very little benefit over simple file replication. However, we analytically show that for large ratios, for a memoryless contact model, distributed erasure coding yields a latency benefit of N/α over uncoded replication, where N is the number of vehicles and α the redundancy factor. Effectively, in this regime, coding yields the same performance as replicating all the files at all other vehicles, but using much less storage. We also evaluate the benefits of coded storage using large real vehicle traces of taxis in Beijing and buses in Chicago. These simulations, which include a realistic radio link quality model for a IEEE 802.11p dedicated short range communication (DSRC) radio, validate the observations from the analysis, demonstrating that coded storage dramatically speeds up the download of large files in vehicular networks.
Maheswaran Sathiamoorthy, Alexandros G. Dimakis, Bhaskar Krishnamachari, Fan Bai 0002
IEEE Trans. Mob. Comput.1
2013 XORing Elephants: Novel Erasure Codes for Big Data
abstract
Distributed storage systems for large clusters typically use replication to provide reliability. Recently, erasure codes have been used to reduce the large storage overhead of three-replicated systems. Reed-Solomon codes are the standard design choice and their high repair cost is often considered an unavoidable price to pay for high storage efficiency and high reliability. This paper shows how to overcome this limitation. We present a novel family of erasure codes that are efficiently repairable and offer higher reliability compared to Reed-Solomon codes. We show analytically that our codes are optimal on a recently identified tradeoff between locality and minimum distance. We implement our new codes in Hadoop HDFS and compare to a currently deployed HDFS module that uses Reed-Solomon codes. Our modified HDFS implementation shows a reduction of approximately 2× on the repair disk I/O and repair network traffic. The disadvantage of the new coding scheme is that it requires 14% more storage compared to Reed-Solomon codes, an overhead shown to be information theoretically optimal to obtain locality. Because the new codes repair failures faster, this provides higher reliability, which is orders of magnitude higher compared to replication.
Maheswaran Sathiamoorthy, Megasthenis Asteris, Dimitris S. Papailiopoulos, Alexandros G. Dimakis, Ramkumar Vadali, Scott Chen, Dhruba Borthakur
Proc. VLDB Endow.1
2012 Backpressure with Adaptive Redundancy (BWAR)
abstract
Backpressure scheduling and routing, in which packets are preferentially transmitted over links with high queue differentials, offers the promise of throughput-optimal operation for a wide range of communication networks. However, when the traffic load is low, due to the corresponding low queue occupancy, backpressure scheduling/routing experiences long delays. This is particularly of concern in intermittent encounter-based mobile networks which are already delay-limited due to the sparse and highly dynamic network connectivity. While state of the art mechanisms for such networks have proposed the use of redundant transmissions to improve delay, they do not work well when the traffic load is high. We propose in this paper a novel hybrid approach that we refer to as backpressure with adaptive redundancy (BWAR), which provides the best of both worlds. This approach is highly robust and distributed and does not require any prior knowledge of network load conditions. We evaluate BWAR through both mathematical analysis and simulations based on a cell-partitioned model. We prove theoretically that BWAR does not perform worse than traditional backpressure in terms of the maximum throughput, while yielding a better delay bound. The simulations confirm that BWAR outperforms traditional backpressure at low load, while outperforming a state of the art encounter-routing scheme (Spray and Wait) at high load.
Majed Alresaini, Maheswaran Sathiamoorthy, Bhaskar Krishnamachari, Michael J. Neely
INFOCOM2
2012 Distributed storage codes reduce latency in vehicular networks
abstract
We investigate the benefits of distributed storage using erasure codes for file sharing in vehicular networks through realistic trace-based simulations. We find that coding offers substantial benefits over simple replication when the file sizes are large compared to the average download bandwidth available per encounter. Our simulations, based on a large real vehicle trace from Beijing combined with a realistic radio link quality model for a IEEE 802.11p dedicated short range communication (DSRC) radio, demonstrate that coding provides significant cost reduction in vehicular networks.
Maheswaran Sathiamoorthy, Alexandros G. Dimakis, Bhaskar Krishnamachari, Fan Bai 0002
INFOCOM1
2012 Minimum Latency Data Diffusion in Intermittently Connected Mobile Networks
abstract
We consider the problem of diffusing cached content in an intermittently connected mobile network, starting from a given initial configuration to a desirable goal state where all nodes interested in particular contents have a copy of their desired contents. The goal is to minimize the time taken for the diffusion process to terminate at a goal state. Due to bandwidth and storage constraints, whenever two nodes encounter each other, they must decide which content if any to transfer to each other. While most prior work on this topic has focused on practically realizable heuristics for this problem, we take a more formal approach. Our main contribution is to show that, assuming global state information is available, this problem can be formulated as a stochastic shortest path problem, which is a kind of Markov decision process (MDP). Using this formulation, we numerically explore some small-scale examples for which we are able to obtain the optimal solution. The results show that the optimal diffusion strategy is very much a function of the underlying encounter graph.
Maheswaran Sathiamoorthy, Wei Gao 0006, Bhaskar Krishnamachari, Guohong Cao
VTC Spring1