Min Xie 0002

dblp:78/8779 · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
5since 2021 · last 2025
0000-0002-7619-2796ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 18 · 8 first-author · 4 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Semi-supervised named entity recognition with data augmentation by structured consistency training
Zhiyuan Peng 0001, Behnoush Abdollahi, Min Xie 0002, Yi Fang 0008
Expert Syst. Appl.3
2023 Modeling Sequential Collaborative User Behaviors For Seller-Aware Next Basket Recommendation
abstract
Next Basket Recommendation (NBR) aims to recommend a set of products as a basket to users based on their historical shopping behavior. In this paper, we investigate the problem of NBR in online marketplaces (e.g., Instacart, Uber Eats) that connect users with multiple sellers. In such scenarios, effective NBR can significantly enhance the shopping experience of users by recommending diversified and completed products based on specific sellers, especially when a user purchases from a seller they have not visited before. However, conventional NBR approaches assume that all considered products are from the same sellers, which overlooks the complex relationships between users, sellers, and products. To address such limitations, we develop SecGT, a sequential collaborative graph transformer framework that recommends users with baskets from specific sellers based on seller-aware user preference representations that are generated by collaboratively modeling the joint user-seller-product interactions and sequentially exploring the user-agnostic basket transitions in an interactive way. We evaluate the performance of SecGT on users from a leading online marketplace at multiple cities with various involved sellers. The results show that SecGT outperforms existing NBR and also traditional product recommendation approaches on recommending baskets from cold sellers for different types of users across all cities.
Ziyi Kou, Saurav Manchanda, Shih-Ting Lin, Min Xie 0002, Haixun Wang, Xiangliang Zhang 0001
CIKM4
2022 Semantic Retrieval at Walmart
abstract
In product search, the retrieval of candidate products before re-ranking is more mission critical and challenging than other search like web search, especially for tail queries, which have a complex and specific search intent. In this paper, we present a hybrid system for e-commerce search deployed at Walmart that combines traditional inverted index and embedding-based neural retrieval to better answer user tail queries. Our system significantly improved the relevance of the search engine, measured by both offline and online evaluations. The improvements were achieved through a combination of different approaches. We present a new technique to train the neural model at scale. and describe how the system was deployed in production with little impact on response time. We highlight multiple learnings and practical tricks that were used in the deployment of this system.
Alessandro Magnani, Feng Liu 0051, Suthee Chaidaroon, Sachin Yadav 0004, Praveen Reddy Suram, Ajit Puthenputhussery, Min Xie 0002, Anirudh Kashi, Ciya Liao
KDD8
2022 Differential Query Semantic Analysis: Discovery of Explicit Interpretable Knowledge from E-Com Search Logs
abstract
We present a novel strategy for analyzing E-Com search logs called Differential Query Semantic Analysis (DQSA) to discover explicit interpretable knowledge from search logs in the form of a semantic lexicon that makes context-specific mapping from a query segment (word or phrase) to the preferred attribute values of a product. Evaluation on a set of size-related query segments and attribute values shows that DQSA can effectively discover meaningful mappings of size-related query segments to their preferred specific attributes and attributes values in the context of a product type. DQSA has many uses including improvement of E-Com search accuracy by bridging the vocabulary gap, comparative analysis of search intent, and alleviation of the problem of tail queries and products.
Sahiti Labhishetty, ChengXiang Zhai, Min Xie 0002, Lin Gong, Rahul Sharnagat, Satya Chembolu
WSDM3
2022 Learning user preferences through online conversations via personalized memory transfer
Nagaarchana Godavarthy, Yuan Wang 0076, Travis Ebesu, Un Suthee, Min Xie 0002, Yi Fang 0008
Inf. Retr. J.5
2019 Neural Compatibility Ranking for Text-based Fashion Matching
abstract
When shopping for fashion, customers often look for products which can complement their current outfit. For example, customers want to buy a jacket which can go well with their jeans and sneakers. To address the task of fashion matching, we propose a neural compatibility model for ranking fashion products based on the compatibility matching with the input outfit. The contribution of our work is twofold. First, we demonstrate that product descriptions contain rich information about product comparability which has not been fully utilized in the prior work. Secondly, we exploit such useful information from text data by taking advantages of semantic matching and lexical matching both of which are important for fashion matching. The proposed model is evaluated on a real-world fashion outfit dataset and achieves the state-of-the-art results by comparing to the competitive baselines. In the future work, we plan to extend the model by incorporating product images which are the major data source in the prior work on fashion matching.
Suthee Chaidaroon, Yi Fang 0008, Min Xie 0002, Alessandro Magnani
SIGIR3
2016 CCCF: Improving Collaborative Filtering via Scalable User-Item Co-Clustering
abstract
Collaborative Filtering (CF) is the most popular method for recommender systems. The principal idea of CF is that users might be interested in items that are favorited by similar users, and most of the existing CF methods measure users' preferences by their behaviours over all the items. However, users might have different interests over different topics, thus might share similar preferences with different groups of users over different sets of items. In this paper, we propose a novel and scalable method CCCF which improves the performance of CF methods via user-item co-clustering. CCCF first clusters users and items into several subgroups, where each subgroup includes a set of like-minded users and a set of items in which these users share their interests. Then, traditional CF methods can be easily applied to each subgroup, and the recommendation results from all the subgroups can be easily aggregated. Compared with previous works, CCCF has several advantages including scalability, flexibility, interpretability and extensibility. Experimental results on four real world data sets demonstrate that the proposed method significantly improves the performance of several state-of-the-art recommendation algorithms.
Min Xie 0002, Martin Ester, Qing Yang 0002
WSDM3
2014 Thwarting Passive Privacy Attacks in Collaborative Filtering
Min Xie 0002, Laks V. S. Lakshmanan
DASFAA (2)2
2014 Recommending user generated item lists
abstract
Existing recommender systems mostly focus on recommending individual items which users may be interested in. User-generated item lists on the other hand have become a popular feature in many applications. E.g., Goodreads provides users with an interface for creating and sharing interesting book lists. These user-generated item lists complement the main functionality of the corresponding application, and intuitively become an alternative way for users to browse and discover interesting items to be consumed. Unfortunately, existing recommender systems are not designed for recommending user-generated item lists. In this work, we study properties of these user-generated item lists and propose a Bayesian ranking model, called LIRE for recommending them. The proposed model takes into consideration users' previous interactions with both item lists and with individual items. Furthermore, we propose in LIRE a novel way of weighting items within item lists based on both position of items, and personalized list consumption pattern. Through extensive experiments on a real item list dataset from Goodreads, we demonstrate the effectiveness of our proposed LIRE model.
Yidan Liu, Min Xie 0002, Laks V. S. Lakshmanan
RecSys2
2014 Generating Top-k Packages via Preference Elicitation
abstract
There are several applications, such as play lists of songs or movies, and shopping carts, where users are interested in finding top- k packages, consisting of sets of items. In response to this need, there has been a recent flurry of activity around extending classical recommender systems (RS), which are effective at recommending individual items, to recommend packages, or sets of items. The few recent proposals for package RS suffer from one of the following drawbacks: they either rely on hard constraints which may be difficult to be specified exactly by the user or on returning Pareto-optimal packages which are too numerous for the user to sift through. To overcome these limitations, we propose an alternative approach for finding personalized top- k packages for users, by capturing users' preferences over packages using a linear utility function which the system learns. Instead of asking a user to specify this function explicitly, which is unrealistic, we explicitly model the uncertainty in the utility function and propose a preference elicitation-based framework for learning the utility function through feedback provided by the user. We propose several sampling-based methods which, given user feedback, can capture the updated utility function. We develop an efficient algorithm for generating top- k packages using the learned utility function, where the rank ordering respects any of a variety of ranking semantics proposed in the literature. Through extensive experiments on both real and synthetic datasets, we demonstrate the efficiency and effectiveness of the proposed system for finding top- k packages.
Min Xie 0002, Laks V. S. Lakshmanan, Peter T. Wood
Proc. VLDB Endow.1
2013 Efficient top-k query answering using cached views
abstract
Top-k query processing has recently received a significant amount of attention due to its wide application in information retrieval, multimedia search and recommendation generation. In this work, we consider the problem of how to efficiently answer a top-k query by using previously cached query results. While there has been some previous work on this problem, existing algorithms suffer from either limited scope or lack of scalability. In this paper, we propose two novel algorithms for handling this problem. The first algorithm LPTA+ provides significantly improved efficiency compared to the state-of-the-art LPTA algorithm [26] by reducing the number of expensive linear programming problems that need to be solved. The second algorithm we propose leverages a standard space partition-based index structure in order to avoid many of the drawbacks of LPTA-based algorithms, thereby further improving the efficiency of query processing. Through extensive experiments on various datasets, we demonstrate that our algorithms significantly outperform the state of the art.
Min Xie 0002, Laks V. S. Lakshmanan, Peter T. Wood
EDBT1
2013 Fair Recommendations for Online Barter Exchange Networks
Zeinab Abbassi, Laks V. S. Lakshmanan, Min Xie 0002
WebDB3
2013 IPS: An Interactive Package Configuration System for Trip Planning
abstract
When planning a trip, one essential task is to find a set of Places-of-Interest (POIs) which can be visited during the trip. Using existing travel guides or websites such as Lonely Planet and TripAdvisor, the user has to either manually work out a desirable set of POIs or take pre-configured travel packages; the former can be time consuming while the latter lacks flexibility. In this demonstration, we propose an Interactive Package configuration System (IPS), which visualizes different candidate packages on a map, and enables users to configure a travel package through simple interactions, i.e., comparing packages and fixing/removing POIs from a package. Compared with existing trip planning systems, we believe IPS strikes the right balance between flexibility and manual effort.
Min Xie 0002, Laks V. S. Lakshmanan, Peter T. Wood
Proc. VLDB Endow.1
2012 Composite recommendations: from items to packages
Min Xie 0002, Laks V. S. Lakshmanan, Peter T. Wood
Frontiers Comput. Sci.1
2011 Adding structure to top-k: from items to expansions
abstract
Keyword based search interfaces are extremely popular as a means for efficiently discovering items of interest from a huge collection, as evidenced by the success of search engines like Google and Bing. However, most of the current search services still return results as a flat ranked list of items. Considering the huge number of items which can match a query, this list based interface can be very difficult for the user to explore and find important items relevant to their search needs. In this work, we consider a search scenario in which each item is annotated with a set of keywords. E.g., in Web 2.0 enabled systems such as flickr and del.icio.us, it is common for users to tag items with keywords. Based on this annotation information, we can automatically group query result items into different expansions of the query corresponding to subsets of keywords. We formulate and motivate this problem within a top-k query processing framework, but as that of finding the top-k most important expansions. Then we study additional desirable properties for the set of expansions returned, and formulate the problem as an optimization problem of finding the best k expansions satisfying all the desirable properties. We propose several efficient algorithms for this problem. Our problem is similar in spirit to recent works on automatic facets generation, but has the important difference and advantage that we don't need to assume the existence of pre-defined categorical hierarchy which is critical for these works. Through extensive experiments on both real and synthetic datasets, we show our proposed algorithms are both effective and efficient.
Xueyao Liang, Min Xie 0002, Laks V. S. Lakshmanan
CIKM2
2011 CompRec-Trip: A composite recommendation system for travel planning
abstract
Classical recommender systems provide users with a list of recommendations where each recommendation consists of a single item, e.g., a book or a DVD. However, applications such as travel planning can benefit from a system capable of recommending packages of items, under a user-specified budget and in the form of sets or sequences. In this context, there is a need for a system that can recommend top-k packages for the user to choose from. In this paper, we propose a novel system, CompRec-Trip, which can automatically generate composite recommendations for travel planning. The system leverages rating information from underlying recommender systems, allows flexible package configuration and incorporates users' cost budgets on both time and money. Furthermore, the proposed CompRec-Trip system has a rich graphical user interface which allows users to customize the returned composite recommendations and take into account external local information.
Min Xie 0002, Laks V. S. Lakshmanan, Peter T. Wood
ICDE1
2011 Efficient Rank Join with Aggregation Constraints
Min Xie 0002, Laks V. S. Lakshmanan, Peter T. Wood
Proc. VLDB Endow.1
2010 Breaking out of the box of recommendations: from items to packages
abstract
Classical recommender systems provide users with a list of recommendations where each recommendation consists of a single item, e.g., a book or DVD. However, several applications can benefit from a system capable of recommending packages of items, in the form of sets. Sample applications include travel planning with a limited budget (price or time) and twitter users wanting to select worthwhile tweeters to follow given that they can deal with only a bounded number of tweets. In these contexts, there is a need for a system that can recommend top-k packages for the user to choose from.
Min Xie 0002, Laks V. S. Lakshmanan, Peter T. Wood
RecSys1
2008 Providing freshness guarantees for outsourced databases
abstract
Database outsourcing becomes increasingly attractive as ad-vances in network technologies eliminate the perceived per-formance difference between in-house databases and out-sourced databases, and price advantages of third-party data-base service providers continue to increase due to economy of scale. However, the potentially explosive growth of database outsourcing is hampered by security concerns, namely data privacy and query integrity of outsourced databases. While privacy issues of outsourced databases have been extensively studied, query integrity for outsourced databases has just started to draw attention from the database community. Currently, there still does not exist a solution that can pro-vide complete integrity. In particular, previous studies have not examined the mechanisms for providing freshness guar-antees, that is, the assurance that queries are executed again-st the most up-to-date data, instead of just some version of the data in the past. Providing a practical solution for fresh-ness guarantees is challenging because continuously moni-toring data’s up-to-dateness is expensive. In this paper, we perform a thorough study on how to add freshness guaran-tees over proposed schemes (including authenticated data structure-based and probabilistic-based approaches) to pro-vide integrity assurance. We implement our solutions and perform extensive experiments to quantify the cost. Our ex-periment results show that we can provide reasonable tight freshness guarantees without sacrificing much performance. 1.
Min Xie 0002, Haixun Wang, Jian Yin 0002, Xiaofeng Meng 0001
EDBT1
2007 Integrity Auditing of Outsourced Data
Min Xie 0002, Haixun Wang, Jian Yin 0002, Xiaofeng Meng 0001
VLDB1