VLDB 2026 Research / reviewers in the wild / expert
Luyi Ma
dblp:259/6573
· DBLP profile ↗
12ranked-venue papers in the field
8as first author
11since 2021 · last 2026
0009-0002-7454-3959ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 8 (5 first)Information Retrieval & Web Search · 3 (2 first)Database Systems & Data Management · 1 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Is More Context Always Better? Examining LLM Reasoning Capability for Time Interval Prediction
Farnaz Fallahi, Murali Mohana Krishna Dandu, Lalitesh Morishetti, Kai Zhao 0011, Luyi Ma, Sinduja Subramaniam, Jianpeng Xu, Evren Körpeoglu, Kaushiki Nag, Kannan Achan |
WWW | 6 |
| 2025 | GRACE: Generative Recommendation via Journey-Aware Sparse Attention on Chain-of-Thought Tokenization
Luyi Ma, Wanjia Zhang, Kai Zhao 0011, Abhishek Kulkarni, Lalitesh Morishetti, Anjana Ganesh, Ashish Ranjan 0006, Aashika Padmanabhan, Jianpeng Xu, Jason H. D. Cho, Praveenkumar Kanumala, Kaushiki Nag, Sumit Dutta, Kamiya Motwani, Malay Patel, Evren Körpeoglu, Kannan Achan |
RecSys | 1 |
| 2025 | ROSI: A hybrid solution for omni-channel feature integration in E-commerce
Luyi Ma, Shengwei Tang, Anjana Ganesh, Jiao Chen 0005, Aashika Padmanabhan, Malay Patel, Jianpeng Xu, Jason H. D. Cho, Evren Körpeoglu, Kannan Achan |
Data Knowl. Eng. | 1 |
| 2024 | Multi-task Recommendation in Marketplace via Knowledge Attentive Graph Convolutional Network with Adaptive Contrastive LearningabstractMarketplaces with multiple sellers have progressively evolved into viable business models in many web applications. Within this sphere, a marketplace recommendation model provides personalized suggestions regarding items and sellers that correspond with users’ preferences. However, a majority of the popular recommendation models are item-centric, often neglecting the incorporation of user preferences to sellers, thereby undermining the comprehensive utilization of the seller-related information contained in the marketplace datasets.To deal with the aforementioned limitations, this study presents a novel model for marketplace recommendations. It employs multi-task learning to jointly recommend items and sellers to users. Here, we introduce the KAROL, a model comprised of two principal modules. The first module is a Knowledge Attentive Graph Convolutional Network (KAGCN) structure based on the user-item-seller knowledge graph (KG). Specifically, relation-aware graph attention and LightGCN are employed to learn node embeddings of users, items and sellers. Sellers and items function reciprocally as knowledge bases for the bipartite graphs to transfer the knowledge between different tasks. It further employs dual losses to concurrently generate recommendations for both items and sellers. The second module is Adaptive Contrastive Learning (ACL), which involves a contrastive loss that incorporates three schemes for data augmentation: cross-relation sampling, edge-dropping and noise addition, to address knowledge sharing, structural consistency, and robustness challenges. An additional innovative facet is the incorporation of an adaptive temperature that is automatically optimized for contrastive loss without manual hyperparameter tuning. The experiment results on three datasets demonstrate that our model outperforms nine baseline models on both item and seller recommendation tasks. Xiaohan Li 0001, Zezhong Fan, Luyi Ma, Kaushiki Nag, Kannan Achan |
IEEE Big Data | 3 |
| 2024 | Improving Sequential Recommender Systems with Online and In-store User BehaviorabstractOnline e-commerce platforms have been extending in-store shopping, which allows users to keep the canonical online browsing and checkout experience while exploring in-store shopping. However, the growing transition between online and in-store becomes a challenge to online sequential recommender systems for future online interaction prediction due to the lack of holistic modeling of hybrid user behaviors (online & in-store). The challenges are two-fold. First, combining online & in-store user behavior data into a single data schema and supporting multiple stages in the model life cycle (pre-training, training, inference, etc.) organically needs a new data pipeline design. Second, online recommender systems, which solely relies on online user behavior sequences, must be redesigned to support online and in-store user data as input under the sequential modeling setting. To overcome the first challenge, we propose a hybrid, omnichannel data pipeline to compile online & in-store user behavior data by caching information from diverse data sources. Later, we introduce a model-agnostic encoder module to the sequential recommender system to interpret the user in-store transaction and augment the modeling capacity for better online interaction prediction given the hybrid user behavior. Luyi Ma, Aashika Padmanabhan, Anjana Ganesh, Shengwei Tang, Jiao Chen 0005, Xiaohan Li 0001, Lalitesh Morishetti, Kaushiki Nag, Malay Patel, Jason H. D. Cho, Kannan Achan |
IEEE Big Data | 1 |
| 2024 | 3rd International Workshop on Industrial Recommendation Systems (IRS)abstractRecommendation systems are used widely across many industries, such as e-commerce, multimedia content platforms, and social networks, to provide suggestions that users will most likely consume or connect, thus improving the user experience. This motivates people in industry and research organizations to focus on personalization and recommendation algorithms, resulting in many research papers. While academic research mostly focuses on the performance of recommendation algorithms in terms of ranking quality or accuracy, it often neglects key factors that impact how a recommendation system will perform in a real-world environment, including but not limited to business metric definition and evaluation, scalability, recommendation quality control, robustness, fairness, and resource limitations, such as computing and memory resources budgets, engineering workforce cost, etc. The gap in constraints and requirements between academic research and industry limits the broad applicability of many of academia's contributions to industrial recommendation systems. This workshop aspires to bridge this gap by bringing together researchers from both academia and industry. Its goal is to serve as a venue for industrial researchers to share practical insights and for academic researchers to become aware of the additional factors of algorithm adoption in real production systems. Luyi Ma, Xiaohan Li 0001, Kamilia Ahmadi, Jianpeng Xu, Philip S. Yu, George Karypis |
CIKM | 1 |
| 2023 | Character-based Outfit Generation with Vision-augmented Style Extraction via LLMsabstractThe outfit generation problem involves recommending a complete outfit to a user based on their interests. Existing approaches focus on recommending items based on anchor items or specific query styles but do not consider customer interests in famous characters from movie, social media, etc. In this paper, we define a new Character-based Outfit Generation (COG) problem, designed to accurately interpret character information and generate complete outfit sets according to customer specifications such as age and gender. To tackle this problem, we propose a novel framework LVA-COG that leverages Large Language Models (LLMs) to extract insights from customer interests (e.g., character information) and employ prompt engineering techniques for accurate understanding of customer preferences. Additionally, we incorporate text-to-image models to enhance the visual understanding and generation (factual or counterfactual) of cohesive outfits. Our framework integrates LLMs with text-to-image models and improves the customer’s approach to fashion by generating personalized recommendations. With experiments and case studies, we demonstrate the effectiveness of our solution from multiple dimensions. Najmeh Forouzandehmehr, Yijie Cao, Nikhil Thakurdesai, Ramin Giahi, Luyi Ma, Nima Farrokhsiar, Jianpeng Xu, Evren Körpeoglu, Kannan Achan |
IEEE Big Data | 5 |
| 2023 | LLMs with User-defined Prompts as Generic Data Operators for Reliable Data ProcessingabstractData processing is one of the fundamental steps in machine learning pipelines to ensure data quality. Majority of the applications consider the user-defined function (UDF) design pattern for data processing in databases. Although the UDF design pattern introduces flexibility, reusability and scalability, the increasing demand on machine learning pipelines brings three new challenges to this design pattern – not low-code, not dependency-free and not knowledge-aware. To address these challenges, we propose a new design pattern that large language models (LLMs) could work as a generic data operator (LLM-GDO) for reliable data cleansing, transformation and modeling with their human-compatible performance. In the LLM-GDO design pattern, user-defined prompts (UDPs) are used to represent the data processing logic rather than implementations with a specific programming language. LLMs can be centrally maintained so users don’t have to manage the dependencies at the run-time. Fine-tuning LLMs with domain-specific data could enhance the performance on the domain-specific tasks which makes data processing knowledge-aware. We illustrate these advantages with examples in different data processing tasks. Furthermore, we summarize the challenges and opportunities introduced by LLMs to provide a complete view of this design pattern for more discussions. Luyi Ma, Nikhil Thakurdesai, Jiao Chen 0005, Jianpeng Xu, Evren Körpeoglu, Kannan Achan |
IEEE Big Data | 1 |
| 2022 | Mitigating Frequency Bias in Next-Basket Recommendation via DeconfoundersabstractRecent studies on Next-basket Recommendation (NBR) have achieved much progress by leveraging Personalized Item Frequency (PIF) as one of the main features, which measures the frequency of the user’s interactions with the item. However, taking the PIF as an explicit feature incurs bias towards frequent items. Items that a user purchases frequently are assigned higher weights in PIF-based recommender system and appear more frequently in the personalized recommendation list. As a result, the system will lose the fairness and balance between items that the user frequently purchases and items that the user never purchases. We refer to this systematic bias on personalized recommendation lists as frequency bias, which narrows users’ browsing scope and reduces the system utility. We adopt causal inference theory to address this issue. Considering the influence of historical purchases on users’ future interests, the user and item representations can be viewed as unobserved confounders in the causal diagram. In this paper, we propose a deconfounder model named FENDER (Frequency-aware Deconfounder for Next-basket Recommendation) to mitigate the frequency bias. With the deconfounder theory and the causal diagram we propose, FENDER decomposes PIF with a neural tensor layer to obtain substitute confounders for users and items. Then, FENDER performs unbiased recommendations considering the effect of these substitute confounders. Experimental results demonstrate that FENDER has derived diverse and fair results compared to ten baseline models on three datasets while achieving competitive performance. Further experiments illustrate how FENDER balances users’ historical purchases and potential interests. Xiaohan Li 0001, Zheng Liu 0017, Luyi Ma, Kaushiki Nag, Stephen D. Guo, Philip S. Yu, Kannan Achan |
IEEE Big Data | 3 |
| 2021 | Event-based Product Carousel Recommendation with Query-Click GraphabstractMany current recommender systems mainly focus on the product-to-product recommendations and user-to-product recommendations even during the time of events rather than modeling the typical recommendations for the target event (e.g., festivals, seasonal activities, or social activities) without addressing the multiple aspects of the shopping demands for the target event. Product recommendations for the multiple aspects of the target event are usually generated by human curators who manually identify the aspects and select a list of aspect-related products (i.e., product carousel) for each aspect as recommendations. However, building a recommender system with machine learning is non-trivial due to the lack of both the ground truth of event-related aspects and the aspect-related products. To fill this gap, we define the novel problem as the event-based product carousel recommendations in e-commerce and propose an effective recommender system based on the query-click bipartite graph. We apply the iterative clustering algorithm over the query-click bipartite graph and infer the event-related aspects by the clusters of queries. The aspect-related recommendations are powered by the click-through rate of products regarding each aspect. We show through experiments that this approach effectively mines product carousels for the target event. Luyi Ma, Nimesh Sinha, Parth Vajge, Jason H. D. Cho, Kannan Achan |
IEEE BigData | 1 |
| 2021 | NEAT: A Label Noise-resistant Complementary Item Recommender System with Trustworthy EvaluationabstractThe complementary item recommender system (CIRS) recommends the complementary items for a given query item. Existing CIRS models consider the item co-purchase signal as a proxy of the complementary relationship, due to the lack of human-curated labels from the huge transaction records. These methods represent items in a complementary embedding space and model the complementary relationship as a point estimation of the similarity between items vectors. However, co-purchased items are not necessarily complementary to each other. For example, customers may frequently purchase bananas and bottle water within the same transaction, but these two items are not complementary. Hence, using co-purchase signals directly as labels will aggravate the model performance. On the other hand, model evaluation will not be trustworthy if the labels for evaluation are not reflecting the true complementary relatedness. To address the above challenges from noisy labeling of the co-purchase data, we model the co-purchases of two items as a Gaussian distribution, where the mean denotes the co-purchases from the complementary relatedness, and covariance denotes the co-purchases from the noise. To do so, we represent each item as a Gaussian embedding and parameterize the Gaussian distribution of co-purchases by the means and covariances from item Gaussian embedding. To reduce the impact of the noisy labels during evaluation, we propose an independence test-based method to generate a trustworthy label set with certain confidence. Our extensive experiments on both the publicly available dataset and the large-scale real-world dataset justify the effectiveness of our proposed model in complementary item recommendations compared with the state-of-the-art models. Luyi Ma, Jianpeng Xu, Jason H. D. Cho, Evren Körpeoglu, Kannan Achan |
IEEE BigData | 1 |
| 2019 | Seasonality-Adjusted Conceptual-Relevancy-Aware Recommender System in Online GroceriesabstractConceptual relevancy - defined as how well a pair of products are related to each other - plays a significant role in online-grocery shopping behavior. When a customer goes grocery shopping, they first come up with a list of ingredients for recipes they may be interested in. A typical shopping list contains categories of products, many of them conceptually relevant to each other. However, the said user may be flexible on which exact product to purchase. For instance, a user may put `milk,' and `cheese' in their shopping list, but these may not refer to specific products such as `Great value 2% milk,' or `Kraft Singles American Slices.' Modern recommender systems, however, focus much more on how to identify specific items to recommend to customers, rather than the categories that may be relevant. Such an approach may lead the system to occasionally recommend outlier, or noise, ultimately violating conceptual relevancy. Moreover, many recommender systems ignore seasonal components; they assume that customers' shopping behavior is independent of the time of the year. However, conceptual relevancy between two products shifts over time. For instance, if a user is shopping for groceries in the middle of the winter, recommending particular products (say, barbecue-related products) may not be the best strategy even if the contextual (user's past interests, or item that the user is currently viewing) may suggest otherwise. In this paper, we introduce a novel strategy to enforce conceptual relevancy in online-grocery domain. Furthermore, recognizing that conceptual relevancy is heavily influenced by the time of the year, we propose a Bayesian-based seasonality algorithm to capture the drift in conceptual relevancy over time without having to re-train the whole model. The algorithm can be based on any of the popular approaches in recommender systems - either based on matrix factorization or that on neural networks. Through our experiments, we show that our seasonality framework can capture drifts in conceptual relevancy. Luyi Ma, Jason H. D. Cho, Kannan Achan |
IEEE BigData | 1 |