Jason H. D. Cho

dblp:137/7808 · also Jason Cho 0001 · DBLP profile ↗
← Back
15ranked-venue papers in the field
1as first author
11since 2021 · last 2026
0000-0002-8106-2961ORCID · verified

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 8 (1 first)Information Retrieval & Web Search · 4Data Mining & Knowledge Discovery · 2Database Systems & Data Management · 1
YearPublicationVenuePosition
2026 Agentic Query Reformulation for Contextualized Hyper-Personalized Product Search
abstract
Traditional e-commerce search often suffers from high abandonment rates due to an intent gap where users provide broad or vague queries. We propose a novel framework using Agentic AI to bridge this gap through dynamic query reformulation. By leveraging the ReAct agentic framework, our system performs contextual reasoning by analyzing historical purchase data to transform generic queries into hyper-personalized search queries. Unlike personalization methods that require architectural overhauls, our solution operates on top of existing search systems, ensuring seamless integration with any search system without costly development or downtime. To ensure low-latency performance, we utilize offline pre-computation of top queries and real-time cosine similarity matching. Results from our experiments demonstrate that this agentic approach can significantly improve search ranking leading to improved customer engagement.
Raghav Gaggar, Sean D. Rosario, Daniel Varivoda, Shahriar Golchin, Jayant Sachdev, Ali Lafzi, Siddharth Pratap Singh, Jason H. D. Cho, Yog Domlur, Swati Kirti, Chittaranjan Tripathy
SIGIR8
2026 Multi-Agentic Recommender Systems: Foundations, Perspectives, and Lessons from Large Scale Deployments in eCommerce
abstract
This tutorial covers topics on multi-agentic recommender systems — recommender systems augmented with Large Language Models (LLMs) and multi-agent orchestration to enable multi-step reasoning, tool use, and interactive decision-making. The tutorial emphasizes foundational concepts, reusable design patterns, and practical lessons learned from large-scale e-commerce deployments. Specifically, we first cover background and recent trends in generative recommender systems and their connection to agentic approaches. We then survey major deployment areas in industry and review the agent orchestration frameworks developed to support them. Finally, we present a project walkthrough that traces the full lifecycle of an agentic recommender system, from scoping and data definition through modeling, deployment, and monitoring, to provide actionable deployment insights. The tutorial bridges perspectives from information retrieval (IR), recommender systems (RecSys), and large-scale industrial practice. The accompanying material can be found at agenticrecsys.github.io.
Reza Yousefi Maragheh, Yashar Deldjoo, Benjamin Coleman, Jason H. D. Cho, Chi Wang 0001
SIGIR4
2025 MetaSynth: Multi-Agent Metadata Generation from Implicit Feedback in Black-Box Systems
Shreeranjani Srirangamsridharan, Ali Abavisani, Reza Yousefi Maragheh, Ramin Giahi, Kai Zhao 0011, Jason H. D. Cho
IEEE Big Data6
2025 GRACE: Generative Recommendation via Journey-Aware Sparse Attention on Chain-of-Thought Tokenization
Luyi Ma, Wanjia Zhang, Kai Zhao 0011, Abhishek Kulkarni, Lalitesh Morishetti, Anjana Ganesh, Ashish Ranjan 0006, Aashika Padmanabhan, Jianpeng Xu, Jason H. D. Cho, Praveenkumar Kanumala, Kaushiki Nag, Sumit Dutta, Kamiya Motwani, Malay Patel, Evren Körpeoglu, Kannan Achan
RecSys10
2025 Multi-Agentic Recommender Systems: Foundations, Design Patterns, and E-Commerce Applications - An Industrial Tutorial
Reza Yousefi Maragheh, Yashar Deldjoo, Chi Wang 0001, Jason H. D. Cho, Derek Cheng
RecSys4
2025 ROSI: A hybrid solution for omni-channel feature integration in E-commerce
Luyi Ma, Shengwei Tang, Anjana Ganesh, Jiao Chen 0005, Aashika Padmanabhan, Malay Patel, Jianpeng Xu, Jason H. D. Cho, Evren Körpeoglu, Kannan Achan
Data Knowl. Eng.8
2024 Improving Sequential Recommender Systems with Online and In-store User Behavior
abstract
Online e-commerce platforms have been extending in-store shopping, which allows users to keep the canonical online browsing and checkout experience while exploring in-store shopping. However, the growing transition between online and in-store becomes a challenge to online sequential recommender systems for future online interaction prediction due to the lack of holistic modeling of hybrid user behaviors (online & in-store). The challenges are two-fold. First, combining online & in-store user behavior data into a single data schema and supporting multiple stages in the model life cycle (pre-training, training, inference, etc.) organically needs a new data pipeline design. Second, online recommender systems, which solely relies on online user behavior sequences, must be redesigned to support online and in-store user data as input under the sequential modeling setting. To overcome the first challenge, we propose a hybrid, omnichannel data pipeline to compile online & in-store user behavior data by caching information from diverse data sources. Later, we introduce a model-agnostic encoder module to the sequential recommender system to interpret the user in-store transaction and augment the modeling capacity for better online interaction prediction given the hybrid user behavior.
Luyi Ma, Aashika Padmanabhan, Anjana Ganesh, Shengwei Tang, Jiao Chen 0005, Xiaohan Li 0001, Lalitesh Morishetti, Kaushiki Nag, Malay Patel, Jason H. D. Cho, Kannan Achan
IEEE Big Data10
2023 LLM-TAKE: Theme-Aware Keyword Extraction Using Large Language Models
abstract
Keyword extraction is one of the core tasks in natural language processing. Classic extraction models are notorious for having a short attention span which make it hard for them to conclude relational connections among the words and sentences that are far from each other. This, in turn, makes their usage prohibitive for generating keywords that are inferred from the context of the whole text. In this paper, we explore using Large Language Models (LLMs) in generating keywords for items that are inferred from the items’ textual metadata. Our modeling framework includes several stages to fine grain the results by avoiding outputting keywords that are non-informative or sensitive and reduce hallucinations common in LLM’s. We call our LLM-based framework Theme-Aware Keyword Extraction (LLM-TAKE). We propose two variations of framework for generating extractive and abstractive themes for products in an E-commerce setting. We perform an extensive set of experiments on three real data sets and show that our modeling framework can enhance accuracy-based and diversity-based metrics when compared with benchmark models.
Reza Yousefi Maragheh, Chenhao Fang, Charan Chand Irugu, Parth Parikh, Jason H. D. Cho, Jianpeng Xu, Saranyan Sukumar, Malay Patel, Evren Körpeoglu, Kannan Achan
IEEE Big Data5
2022 Prospect-Net: Top-K Retrieval Problem Using Prospect Theory
abstract
In e-Commerce Industry, customers’ purchase decision of an item are usually influenced by the reference price of that item, which is implied within the context of the items (e.g. prices of an item set from search/recommendation) or external environments (e.g. prices from another e-Commerce platform). Despite of the prevalence and influence of the reference price on customers’ behavior, existing works in Information Retrieval domain do not exploit the value of the reference price in ranking problems. In this paper, we propose a list-wise ranking model named "Prospect-Net" by incorporating the prospect theory, which is the theoretical foundation for framing the reference price. We consider the Top-K retrieval task under a product recommendation setting, and demonstrate the effectiveness of Prospect-Net to capture various forms of reference price under different scenarios. Polynomial solutions are proposed to solve the Top-K retrieval problem for some of the cases where the reference price is dependent on the recommended set of items to the user. Both offline e valuation and online experiments are performed on a real-world industrial dataset with significant performance improvement.
Reza Yousefi Maragheh, Ramin Giahi, Jianpeng Xu, Lalitesh Morishetti, Shanu Vashishtha, Kaushiki Nag, Jason H. D. Cho, Evren Körpeoglu, Kannan Achan
IEEE Big Data7
2021 Event-based Product Carousel Recommendation with Query-Click Graph
abstract
Many current recommender systems mainly focus on the product-to-product recommendations and user-to-product recommendations even during the time of events rather than modeling the typical recommendations for the target event (e.g., festivals, seasonal activities, or social activities) without addressing the multiple aspects of the shopping demands for the target event. Product recommendations for the multiple aspects of the target event are usually generated by human curators who manually identify the aspects and select a list of aspect-related products (i.e., product carousel) for each aspect as recommendations. However, building a recommender system with machine learning is non-trivial due to the lack of both the ground truth of event-related aspects and the aspect-related products. To fill this gap, we define the novel problem as the event-based product carousel recommendations in e-commerce and propose an effective recommender system based on the query-click bipartite graph. We apply the iterative clustering algorithm over the query-click bipartite graph and infer the event-related aspects by the clusters of queries. The aspect-related recommendations are powered by the click-through rate of products regarding each aspect. We show through experiments that this approach effectively mines product carousels for the target event.
Luyi Ma, Nimesh Sinha, Parth Vajge, Jason H. D. Cho, Kannan Achan
IEEE BigData4
2021 NEAT: A Label Noise-resistant Complementary Item Recommender System with Trustworthy Evaluation
abstract
The complementary item recommender system (CIRS) recommends the complementary items for a given query item. Existing CIRS models consider the item co-purchase signal as a proxy of the complementary relationship, due to the lack of human-curated labels from the huge transaction records. These methods represent items in a complementary embedding space and model the complementary relationship as a point estimation of the similarity between items vectors. However, co-purchased items are not necessarily complementary to each other. For example, customers may frequently purchase bananas and bottle water within the same transaction, but these two items are not complementary. Hence, using co-purchase signals directly as labels will aggravate the model performance. On the other hand, model evaluation will not be trustworthy if the labels for evaluation are not reflecting the true complementary relatedness. To address the above challenges from noisy labeling of the co-purchase data, we model the co-purchases of two items as a Gaussian distribution, where the mean denotes the co-purchases from the complementary relatedness, and covariance denotes the co-purchases from the noise. To do so, we represent each item as a Gaussian embedding and parameterize the Gaussian distribution of co-purchases by the means and covariances from item Gaussian embedding. To reduce the impact of the noisy labels during evaluation, we propose an independence test-based method to generate a trustworthy label set with certain confidence. Our extensive experiments on both the publicly available dataset and the large-scale real-world dataset justify the effectiveness of our proposed model in complementary item recommendations compared with the state-of-the-art models.
Luyi Ma, Jianpeng Xu, Jason H. D. Cho, Evren Körpeoglu, Kannan Achan
IEEE BigData3
2020 Knowledge-aware Complementary Product Representation Learning
abstract
Learning product representations that reflect complementary relationship plays a central role in e-commerce recommender system. In the absence of the product relationships graph, which existing methods rely on, there is a need to detect the complementary relationships directly from noisy and sparse customer purchase activities. Furthermore, unlike simple relationships such as similarity, complementariness is asymmetric and non-transitive. Standard usage of representation learning emphasizes on only one set of embedding, which is problematic for modelling such properties of complementariness.
Chuanwei Ruan, Jason H. D. Cho, Evren Körpeoglu, Kannan Achan
WSDM3
2019 Seasonality-Adjusted Conceptual-Relevancy-Aware Recommender System in Online Groceries
abstract
Conceptual relevancy - defined as how well a pair of products are related to each other - plays a significant role in online-grocery shopping behavior. When a customer goes grocery shopping, they first come up with a list of ingredients for recipes they may be interested in. A typical shopping list contains categories of products, many of them conceptually relevant to each other. However, the said user may be flexible on which exact product to purchase. For instance, a user may put `milk,' and `cheese' in their shopping list, but these may not refer to specific products such as `Great value 2% milk,' or `Kraft Singles American Slices.' Modern recommender systems, however, focus much more on how to identify specific items to recommend to customers, rather than the categories that may be relevant. Such an approach may lead the system to occasionally recommend outlier, or noise, ultimately violating conceptual relevancy. Moreover, many recommender systems ignore seasonal components; they assume that customers' shopping behavior is independent of the time of the year. However, conceptual relevancy between two products shifts over time. For instance, if a user is shopping for groceries in the middle of the winter, recommending particular products (say, barbecue-related products) may not be the best strategy even if the contextual (user's past interests, or item that the user is currently viewing) may suggest otherwise. In this paper, we introduce a novel strategy to enforce conceptual relevancy in online-grocery domain. Furthermore, recognizing that conceptual relevancy is heavily influenced by the time of the year, we propose a Bayesian-based seasonality algorithm to capture the drift in conceptual relevancy over time without having to re-train the whole model. The algorithm can be based on any of the popular approaches in recommender systems - either based on matrix factorization or that on neural networks. Through our experiments, we show that our seasonality framework can capture drifts in conceptual relevancy.
Luyi Ma, Jason H. D. Cho, Kannan Achan
IEEE BigData2
2015 Recommending forum posts to designated experts
abstract
There are users who generate significant amounts of domain knowledge in online forums or community question and answer (CQA) websites. Existing literature defines them as `experts.' These users attain such statuses by providing multiple relevant answers to the question askers. Past works have focused on recommending relevant posts to these users. With the rise of web forums where certified experts answer questions, strategies that are tailored towards addressing the new type of experts will be beneficial. In this paper, we identify a new type of user called `designated experts' (i.e., users designated as domain experts by the web administrators). These are the experts who are guaranteed by web administrators to be an expert in a given domain. Our focus is on how we can capture the unique behavior of designated experts in an online domain. We have noticed designated experts have different behaviors compared to CQA experts. In particular, unlike existing CQAs, only one designated expert responds to any given thread. To capture this intuition, we introduce a matrix factorization algorithm with regularization to capture the behavior. Our results show that the regularization method improves the performance significantly compared to the baseline approach.
Jason H. D. Cho, Yanen Li, Roxana Girju, ChengXiang Zhai
IEEE BigData1
2014 Local Learning for Mining Outlier Subgraphs from Network Datasets
abstract
In the real world, various systems can be modeled using entity-relationship graphs. Given such a graph, one may be interested in identifying suspicious or anomalous subgraphs. Specifically, a user may want to identify suspicious subgraphs matching a query template. A subgraph can be defined as anomalous based on the connectivity structure within itself as well as with its neighborhood. For example for a co-authorship network, given a subgraph containing three authors, one expects all three authors to be say data mining authors. Also, one expects the neighborhood to mostly consist of data mining authors. But a 3-author clique of data mining authors with all theory authors in the neighborhood clearly seems interesting. Similarly, having one of the authors in the clique as a theory author when all other authors (both in the clique and neighborhood) are data mining authors, is also suspicious. Thus, existence of low-probability links and absence of high-probability links can be a good indicator of subgraph outlierness. The probability of an edge can in turn be modeled based on the weighted similarity between the attribute values of the nodes linked by the edge. We claim that the attribute weights must be learned locally for accurate link existence probability computations. In this paper, we design a system that finds subgraph outliers given a graph and a query by modeling the problem as a linear optimization. Experimental results on several synthetic and real datasets show the effectiveness of the proposed approach in computing interesting outliers.
Manish Gupta 0001, Arun Mallya, Subhro Roy, Jason H. D. Cho, Jiawei Han 0001
SDM4