EDBT 2026 Demo / reviewers in the wild / expert
Jianpeng Xu
dblp:127/4166
· DBLP profile ↗
22ranked-venue papers in the field
8as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 8 (5 first)Big Data, Cloud & Distributed Data Systems · 6 (1 first)Information Retrieval & Web Search · 5Database Systems & Data Management · 3 (2 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Is More Context Always Better? Examining LLM Reasoning Capability for Time Interval Prediction
Farnaz Fallahi, Murali Mohana Krishna Dandu, Lalitesh Morishetti, Kai Zhao 0011, Luyi Ma, Sinduja Subramaniam, Jianpeng Xu, Evren Körpeoglu, Kaushiki Nag, Kannan Achan |
WWW | 8 |
| 2025 | VL-CLIP: Enhancing Multimodal Recommendations via Visual Grounding and LLM-Augmented CLIP Embeddings
Ramin Giahi, Kehui Yao, Sriram Kollipara, Kai Zhao 0011, Vahid Mirjalili, Jianpeng Xu, Topojoy Biswas, Evren Körpeoglu, Kannan Achan |
RecSys | 6 |
| 2025 | GRACE: Generative Recommendation via Journey-Aware Sparse Attention on Chain-of-Thought Tokenization
Luyi Ma, Wanjia Zhang, Kai Zhao 0011, Abhishek Kulkarni, Lalitesh Morishetti, Anjana Ganesh, Ashish Ranjan 0006, Aashika Padmanabhan, Jianpeng Xu, Jason H. D. Cho, Praveenkumar Kanumala, Kaushiki Nag, Sumit Dutta, Kamiya Motwani, Malay Patel, Evren Körpeoglu, Kannan Achan |
RecSys | 9 |
| 2025 | ROSI: A hybrid solution for omni-channel feature integration in E-commerce
Luyi Ma, Shengwei Tang, Anjana Ganesh, Jiao Chen 0005, Aashika Padmanabhan, Malay Patel, Jianpeng Xu, Jason H. D. Cho, Evren Körpeoglu, Kannan Achan |
Data Knowl. Eng. | 7 |
| 2024 | 3rd International Workshop on Industrial Recommendation Systems (IRS)abstractRecommendation systems are used widely across many industries, such as e-commerce, multimedia content platforms, and social networks, to provide suggestions that users will most likely consume or connect, thus improving the user experience. This motivates people in industry and research organizations to focus on personalization and recommendation algorithms, resulting in many research papers. While academic research mostly focuses on the performance of recommendation algorithms in terms of ranking quality or accuracy, it often neglects key factors that impact how a recommendation system will perform in a real-world environment, including but not limited to business metric definition and evaluation, scalability, recommendation quality control, robustness, fairness, and resource limitations, such as computing and memory resources budgets, engineering workforce cost, etc. The gap in constraints and requirements between academic research and industry limits the broad applicability of many of academia's contributions to industrial recommendation systems. This workshop aspires to bridge this gap by bringing together researchers from both academia and industry. Its goal is to serve as a venue for industrial researchers to share practical insights and for academic researchers to become aware of the additional factors of algorithm adoption in real production systems. Luyi Ma, Xiaohan Li 0001, Kamilia Ahmadi, Jianpeng Xu, Philip S. Yu, George Karypis |
CIKM | 4 |
| 2024 | LLM-Ensemble: Optimal Large Language Model Ensemble Method for E-commerce Product Attribute Value ExtractionabstractProduct attribute value extraction is a pivotal component in Natural Language Processing (NLP) and the contemporary e-commerce industry. The provision of precise product attribute values is fundamental in ensuring high-quality recommendations and enhancing customer satisfaction. The recently emerging Large Language Models (LLMs) have demonstrated state of-the-art performance in numerous attribute extraction tasks, without the need for domain-specific training data. Nevertheless, varying strengths and weaknesses are exhibited by different LLMs due to the diversity in data, architectures, and hyperparameters. This variation makes them complementary to each other, with no single LLM dominating all others. Considering the diverse strengths and weaknesses of LLMs, it becomes necessary to develop an ensemble method that leverages their complementary potentials. Chenhao Fang, Xiaohan Li 0001, Zezhong Fan, Jianpeng Xu, Kaushiki Nag, Evren Körpeoglu, Kannan Achan |
SIGIR | 4 |
| 2023 | Character-based Outfit Generation with Vision-augmented Style Extraction via LLMsabstractThe outfit generation problem involves recommending a complete outfit to a user based on their interests. Existing approaches focus on recommending items based on anchor items or specific query styles but do not consider customer interests in famous characters from movie, social media, etc. In this paper, we define a new Character-based Outfit Generation (COG) problem, designed to accurately interpret character information and generate complete outfit sets according to customer specifications such as age and gender. To tackle this problem, we propose a novel framework LVA-COG that leverages Large Language Models (LLMs) to extract insights from customer interests (e.g., character information) and employ prompt engineering techniques for accurate understanding of customer preferences. Additionally, we incorporate text-to-image models to enhance the visual understanding and generation (factual or counterfactual) of cohesive outfits. Our framework integrates LLMs with text-to-image models and improves the customer’s approach to fashion by generating personalized recommendations. With experiments and case studies, we demonstrate the effectiveness of our solution from multiple dimensions. Najmeh Forouzandehmehr, Yijie Cao, Nikhil Thakurdesai, Ramin Giahi, Luyi Ma, Nima Farrokhsiar, Jianpeng Xu, Evren Körpeoglu, Kannan Achan |
IEEE Big Data | 7 |
| 2023 | LLMs with User-defined Prompts as Generic Data Operators for Reliable Data ProcessingabstractData processing is one of the fundamental steps in machine learning pipelines to ensure data quality. Majority of the applications consider the user-defined function (UDF) design pattern for data processing in databases. Although the UDF design pattern introduces flexibility, reusability and scalability, the increasing demand on machine learning pipelines brings three new challenges to this design pattern – not low-code, not dependency-free and not knowledge-aware. To address these challenges, we propose a new design pattern that large language models (LLMs) could work as a generic data operator (LLM-GDO) for reliable data cleansing, transformation and modeling with their human-compatible performance. In the LLM-GDO design pattern, user-defined prompts (UDPs) are used to represent the data processing logic rather than implementations with a specific programming language. LLMs can be centrally maintained so users don’t have to manage the dependencies at the run-time. Fine-tuning LLMs with domain-specific data could enhance the performance on the domain-specific tasks which makes data processing knowledge-aware. We illustrate these advantages with examples in different data processing tasks. Furthermore, we summarize the challenges and opportunities introduced by LLMs to provide a complete view of this design pattern for more discussions. Luyi Ma, Nikhil Thakurdesai, Jiao Chen 0005, Jianpeng Xu, Evren Körpeoglu, Kannan Achan |
IEEE Big Data | 4 |
| 2023 | LLM-TAKE: Theme-Aware Keyword Extraction Using Large Language ModelsabstractKeyword extraction is one of the core tasks in natural language processing. Classic extraction models are notorious for having a short attention span which make it hard for them to conclude relational connections among the words and sentences that are far from each other. This, in turn, makes their usage prohibitive for generating keywords that are inferred from the context of the whole text. In this paper, we explore using Large Language Models (LLMs) in generating keywords for items that are inferred from the items’ textual metadata. Our modeling framework includes several stages to fine grain the results by avoiding outputting keywords that are non-informative or sensitive and reduce hallucinations common in LLM’s. We call our LLM-based framework Theme-Aware Keyword Extraction (LLM-TAKE). We propose two variations of framework for generating extractive and abstractive themes for products in an E-commerce setting. We perform an extensive set of experiments on three real data sets and show that our modeling framework can enhance accuracy-based and diversity-based metrics when compared with benchmark models. Reza Yousefi Maragheh, Chenhao Fang, Charan Chand Irugu, Parth Parikh, Jason H. D. Cho, Jianpeng Xu, Saranyan Sukumar, Malay Patel, Evren Körpeoglu, Kannan Achan |
IEEE Big Data | 6 |
| 2022 | Prospect-Net: Top-K Retrieval Problem Using Prospect TheoryabstractIn e-Commerce Industry, customers’ purchase decision of an item are usually influenced by the reference price of that item, which is implied within the context of the items (e.g. prices of an item set from search/recommendation) or external environments (e.g. prices from another e-Commerce platform). Despite of the prevalence and influence of the reference price on customers’ behavior, existing works in Information Retrieval domain do not exploit the value of the reference price in ranking problems. In this paper, we propose a list-wise ranking model named "Prospect-Net" by incorporating the prospect theory, which is the theoretical foundation for framing the reference price. We consider the Top-K retrieval task under a product recommendation setting, and demonstrate the effectiveness of Prospect-Net to capture various forms of reference price under different scenarios. Polynomial solutions are proposed to solve the Top-K retrieval problem for some of the cases where the reference price is dependent on the recommended set of items to the user. Both offline e valuation and online experiments are performed on a real-world industrial dataset with significant performance improvement. Reza Yousefi Maragheh, Ramin Giahi, Jianpeng Xu, Lalitesh Morishetti, Shanu Vashishtha, Kaushiki Nag, Jason H. D. Cho, Evren Körpeoglu, Kannan Achan |
IEEE Big Data | 3 |
| 2021 | NEAT: A Label Noise-resistant Complementary Item Recommender System with Trustworthy EvaluationabstractThe complementary item recommender system (CIRS) recommends the complementary items for a given query item. Existing CIRS models consider the item co-purchase signal as a proxy of the complementary relationship, due to the lack of human-curated labels from the huge transaction records. These methods represent items in a complementary embedding space and model the complementary relationship as a point estimation of the similarity between items vectors. However, co-purchased items are not necessarily complementary to each other. For example, customers may frequently purchase bananas and bottle water within the same transaction, but these two items are not complementary. Hence, using co-purchase signals directly as labels will aggravate the model performance. On the other hand, model evaluation will not be trustworthy if the labels for evaluation are not reflecting the true complementary relatedness. To address the above challenges from noisy labeling of the co-purchase data, we model the co-purchases of two items as a Gaussian distribution, where the mean denotes the co-purchases from the complementary relatedness, and covariance denotes the co-purchases from the noise. To do so, we represent each item as a Gaussian embedding and parameterize the Gaussian distribution of co-purchases by the means and covariances from item Gaussian embedding. To reduce the impact of the noisy labels during evaluation, we propose an independence test-based method to generate a trustworthy label set with certain confidence. Our extensive experiments on both the publicly available dataset and the large-scale real-world dataset justify the effectiveness of our proposed model in complementary item recommendations compared with the state-of-the-art models. Luyi Ma, Jianpeng Xu, Jason H. D. Cho, Evren Körpeoglu, Kannan Achan |
IEEE BigData | 2 |
| 2021 | 2nd International Workshop on Industrial Recommendation Systems (IRS)abstractRecommendation systems are used widely across many industries, such as e-commerce, multimedia content platforms and social networks, to provide suggestions that a user will most likely consume or connect; thus, improving the user experience. This motivates people in both industry and research organizations to focus on personalization or recommendation algorithms, which has resulted in a plethora of research papers. While academic research mostly focuses on the performance of recommendation algorithms in terms of ranking quality or accuracy, it often neglects key factors that impact how a recommendation system will perform in a real-world environment. These key factors include but are not limited to: business metric definition and evaluation, recommendation quality control, data and model scalability, model interpretability, model robustness and fairness, and resource limitations, such as computing and memory resources budgets, engineering workforce cost, etc. The gap in constraints and requirements between academic research and industry limits the broad applicability of many of academia's contributions for industrial recommendation systems. This workshop aspires to bridge this gap by bringing together researchers from both academia and industry. Its goal is to serve as a venue through which academic researchers become aware of the additional factors that may affect the adoption of an algorithm into real production systems, and how well it will perform if deployed. Industrial researchers will also benefit from sharing the practical insights, approaches, and frameworks as well. Jianpeng Xu, Lingfei Wu 0001, Linsey Pang, Mohit Sharma 0002, Dawei Yin 0001, George Karypis, Justin Basilico, Philip S. Yu |
KDD | 1 |
| 2021 | PURE: Positive-Unlabeled Recommendation with Generative Adversarial NetworkabstractRecommender systems are powerful tools for information filtering with the ever-growing amount of online data. Despite its success and wide adoption in various web applications and personalized products, many existing recommender systems still suffer from multiple drawbacks such as large amount of unobserved feedback, poor model convergence, etc. These drawbacks of existing work are mainly due to the following two reasons: first, the widely used negative sampling strategy, which treats the unlabeled entries as negative samples, is invalid in real-world settings; second, all training samples are retrieved from the discrete observations, and the underlying true distribution of the users and items is not learned. Yao Zhou 0003, Jianpeng Xu, Jun Wu 0019, Zeinab Taghavi Nasrabadi, Evren Körpeoglu, Kannan Achan, Jingrui He |
KDD | 2 |
| 2021 | Spatio-Temporal Multi-Task Learning via Tensor DecompositionabstractPredictive modeling of large-scale spatio-temporal data is an important but challenging problem as it requires training models that can simultaneously predict the target variables of interest at multiple locations while preserving the spatial and temporal dependencies of the data. In this paper, we investigate the effectiveness of applying a multi-task learning approach based on supervised tensor decomposition to the spatio-temporal prediction problem. Our proposed framework, known as SMART, encodes the data as a third-order tensor and extracts a set of interpretable, spatial and temporal latent factors from the data. An ensemble of spatial and temporal prediction models are trained using the latent factors as their predictor variables. Outputs from the ensemble model are aggregated to make predictions on test instances. The framework also allows known patterns from the domain to be incorporated as constraints to guide the tensor decomposition and ensemble learning processes. As the data may grow over space and time, an incremental learning version of the framework is given to efficiently update the models. We perform extensive experiments using a global-scale climate dataset to evaluate the accuracy and efficiency of the models as well as interpretability of the latent factors. Jianpeng Xu, Pang-Ning Tan, Lifeng Luo |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2017 | Online Multi-Task Learning Framework for Ensemble ForecastingabstractEnsemble forecasting is a widely-used numerical prediction method for modeling the evolution of nonlinear dynamic systems. To predict the future state of such systems, a set of ensemble member forecasts is generated from multiple runs of computer models, where each run is obtained by perturbing the starting condition or using a different model representation of the system. The ensemble mean or median is typically chosen as a point estimate for the ensemble member forecasts. These approaches are limited in that they assume each ensemble member is equally skillful and may not preserve the temporal autocorrelation of the predicted time series. To overcome these limitations, we present an online multi-task learning framework called ORION to estimate the optimal weights for combining the ensemble member forecasts. Unlike other existing formulations, the proposed framework is novel in that its learning algorithm must backtrack and revise its previous forecasts before making future predictions if the earlier forecasts were incorrect when verified against new observation data. We termed this strategy as online learning with restart. Our proposed framework employs a graph Laplacian regularizer to ensure consistency of the predicted time series. It can also accommodate different types of loss functions, including ϵ-insensitive and quantile loss functions, the latter of which is particularly useful for extreme value prediction. A theoretical proof demonstrating the convergence of our algorithm is also given. Experimental results on seasonal soil moisture forecasts from 12 major river basins in North America demonstrate the superiority of ORION compared to other baseline algorithms. Jianpeng Xu, Pang-Ning Tan, Lifeng Luo |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2016 | WISDOM: Weighted incremental spatio-temporal multi-task learning via tensor decompositionabstractThis paper presents a novel multi-task learning framework for the accurate prediction of spatio-temporal data at multiple locations. The framework encodes the data as a third-order tensor and performs supervised tensor decomposition to identify the latent factors that capture the inherent spatiotemporal variabilities of the data and their relationship to the target variable of interest. The framework is unique in that it trains both spatial and temporal prediction models from the latent factors of the decomposed tensor and aggregates their outputs to generate its final prediction. The latent factors and model parameters are simultaneously estimated by optimizing a joint objective function. We also develop an incremental learning algorithm called WISDOM to efficiently solve the optimization problem, in which the model is gradually updated with new data, either from a previously unobserved location or from its most recent time period. WISDOM can also incorporate known patterns from the application domain to guide the tensor decomposition. Finally, we showed that WISDOM outperforms several baseline algorithms in more than 75% of the locations when applied to a global-scale climate data. Jianpeng Xu, Pang-Ning Tan, Lifeng Luo |
IEEE BigData | 1 |
| 2016 | Multi-Task Feature Interaction LearningabstractOne major limitation of linear models is the lack of capability to capture predictive information from interactions between features. While introducing high-order feature interaction terms can overcome this limitation, this approach tremendously increases the model complexity and imposes significant challenges in the learning against overfitting. In this paper, we proposed a novel Multi-Task feature Interaction Learning~(MTIL) framework to exploit the task relatedness from high-order feature interactions, which provides better generalization performance by inductive transfer among tasks via shared representations of feature interactions. We formulate two concrete approaches under this framework and provide efficient algorithms: the shared interaction approach and the embedded interaction approach. The former assumes tasks share the same set of interactions, and the latter assumes feature interactions from multiple tasks come from a shared subspace. We have provided efficient algorithms for solving the two approaches. Extensive empirical studies on both synthetic and real datasets have demonstrated the effectiveness of the proposed framework. Kaixiang Lin, Jianpeng Xu, Inci M. Baytas, Shuiwang Ji |
KDD | 2 |
| 2016 | Synergies that Matter: Efficient Interaction Selection via Sparse Factorization MachineabstractCollaborative filtering has been widely used in modern recommender systems to provide accurate recommendations by leveraging historical interactions between users and items. The presence of cold-start items and users has imposed a huge challenge to recommender systems based on collaborative filtering, because of the unavailability of such interaction information. The factorization machine is a powerful tool designed to tackle the cold-start problems by learning a bilinear ranking model that utilizes content information about users and items, exploiting the interactions with such content information. While a factorization machine makes use of all possible interactions between all content features to make recommendations, many of the features and their interactions are not predictive of recommendations, and incorporating them in the model will deteriorate the generalization performance of the recommender systems. In this paper, we propose an efficient Sparse Factorization Machine (SFM), that simultaneously identifies relevant user and item content features, models interactions between these relevant features, and learns a bilinear model using only these synergistic interactions. We have carried out extensive empirical studies on both synthetic and real-world datasets, and compared our method to other state-of-the-art baselines, including Factorization Machine. Experimental results show that SFM can greatly outperform other baselines. Jianpeng Xu, Kaixiang Lin, Pang-Ning Tan |
SDM | 1 |
| 2016 | GSpartan: a Geospatio-Temporal Multi-task Learning Framework for Multi-location PredictionabstractThis paper presents a novel geospatio-temporal prediction framework called GSpartan to simultaneously build local regression models at multiple locations. The framework assumes that the local models share a common, low-rank representation, which makes them amenable to multi-task learning. GSpartan learns a set of base models to capture the spatio-temporal variabilities of the data and represents each local model as a linear combination of the base models. A graph Laplacian regularization is used to enforce constraints on the local models based on their spatial autocorrelation. We also introduce sparsity-inducing norms to perform feature selection for the base models and model selection for the local models. Experimental results using historical climate data from 37 weather stations showed that, on average, GSpartan outperforms single-task learning and other existing multi-task learning methods in more than 65% of the stations, which increases to 81% when there are fewer training examples. Jianpeng Xu, Pang-Ning Tan, Lifeng Luo |
SDM | 1 |
| 2015 | FORMULA: FactORized MUlti-task LeArning for task discovery in personalized medical modelsabstractMedical predictive modeling is a challenging problem due to the heterogeneous nature of the patients. In order to build effective medical predictive models we need to address such heterogeneous nature during modeling and allow patients to have their own personalized models instead of using a one-size-fits-all model. However, building a personalized model for each patient is computationally expensive and the over-parametrization of the model makes it susceptible to the model overfitting problem. To address these challenges, we propose a novel approach called FactORized MUlti-task LeArning model (FORMULA), which learns the personalized model of each patient via a sparse multi-task learning method. The personalized models are assumed to share a low-rank representation, known as the base models. FORMULA is designed to simultaneously learn the base models as well as the personalized model of each patient, where the latter is a linear combination of the base models. We have performed extensive experiments to evaluate the proposed approach on a real medical data set. The proposed approach delivered superior predictive performance while the personalized models offered many useful medical insights. Jianpeng Xu, Pang-Ning Tan |
SDM | 1 |
| 2014 | Detecting malicious clients in ISP networks using HTTP connectivity graph and flow informationabstractThis paper considers an approach to identify previously undetected malicious clients in Internet Service Provider (ISP) networks by combining flow classification with a graph-based score propagation method. Our approach represents all HTTP communications between clients and servers as a weighted, near-bipartite graph, where the nodes correspond to the IP addresses of clients and servers while the links are their interconnections, weighted according to the output of a flow-based classifier. We employ a two-phase alternating score propagation algorithm on the graph to identify suspicious clients in a monitored network. Using a symmetrized weighted adjacency matrix as its input, we show that our score propagation algorithm is less vulnerable towards inflating the malicious scores of popular Web servers with high in-degrees compared to the normalization used in PageRank, a widely used graph-based method. Experimental results on a 4-hour network trace collected by a large Internet service provider showed that incorporating flow information into score propagation significantly improves the precision of the algorithm. Sabyasachi Saha, Ruben Torres, Jianpeng Xu, Pang-Ning Tan, Antonio Nucci, Marco Mellia |
ASONAM | 4 |
| 2014 | ORION: Online Regularized Multi-task Regression and Its Application to Ensemble ForecastingabstractEnsemble forecasting is a well-known numerical prediction technique for modeling the evolution of nonlinear dynamic systems. The ensemble member forecasts are generated from multiple runs of a computer model, where each run is obtained by perturbing the starting condition or using a different model representation of the dynamic system. The ensemble mean or median is typically chosen as the consensus point estimate of the aggregated forecasts for decision making purposes. These approaches are limited in that they assume each ensemble member is equally skill ful and do not consider their inherent correlations. In this paper, we cast the ensemble forecasting task as an online, multi-task regression problem and present a framework called ORION to estimate the optimal weights for combining the ensemble members. The weights are updated using a novel online learning with restart strategy as new observation data become available. Experimental results on seasonal soil moisture predictions from 12 major river basins in North America demonstrate the superiority of the proposed approach compared to the ensemble median and other baseline methods. Jianpeng Xu, Pang-Ning Tan, Lifeng Luo |
ICDM | 1 |