Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Xuetao Ding

dblp:133/1925 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0002-3551-5613ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Reinforcement learning · 62% Probabilistic and Bayesian machine learning · 26% Graph learning · 12%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Smart cities and intelligent transportation · 100%
Databases, data mining, and information retrieval
3 papers
Recommender systems · 77% Data mining · 23%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › offline reinforcement learning
model-based offline reinforcement learning
1.012026
Reliability-Guaranteed and Reward-Seeking Sequence Modeling for Model-Based Offline Reinforcement Learning · AAAI 2026
Machine learning › Reinforcement learning
offline reinforcement learning
1.012026
Reliability-Guaranteed and Reward-Seeking Sequence Modeling for Model-Based Offline Reinforcement Learning · AAAI 2026
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.912025
Offline Multi-Agent Reinforcement Learning via In-Sample Sequential Policy Optimization · AAAI 2025
Machine learning › Reinforcement learning › offline reinforcement learning
offline policy optimization
0.912025
Offline Multi-Agent Reinforcement Learning via In-Sample Sequential Policy Optimization · AAAI 2025
Machine learning › Graph learning
graph representation learning
0.812024
Harvesting Efficient On-Demand Order Pooling from Skilled Couriers: Enhancing Graph Representation Learning for Refining Real-time Many-to-One Assignments · KDD 2024
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.612022
Counterfactual Prediction for Outcome-Oriented Treatments · ICML 2022
Machine learning › Probabilistic and Bayesian machine learning › causal inference
counterfactual prediction
0.612022
Counterfactual Prediction for Outcome-Oriented Treatments · ICML 2022
Machine learning › Probabilistic and Bayesian machine learning › causal inference › causal effect estimation
treatment effect estimation
0.612022
Counterfactual Prediction for Outcome-Oriented Treatments · ICML 2022
Smart cities and intelligent transportation › logistics
on-demand delivery
0.412020
Delivery Scope: A New Way of Restaurant Retrieval for On-demand Food Delivery Service · KDD 2020
Machine learning › Reinforcement learning
policy learning
0.312026
Reliability-Guaranteed and Reward-Seeking Sequence Modeling for Model-Based Offline Reinforcement Learning · AAAI 2026
Data mining
spatiotemporal data mining
0.212024
Harvesting Efficient On-Demand Order Pooling from Skilled Couriers: Enhancing Graph Representation Learning for Refining Real-time Many-to-One Assignments · KDD 2024
Recommender systems
collaborative filtering
0.212013
Celebrity Recommendation with Collaborative Social Topic Regression · IJCAI 2013
Recommender systems
social recommendation
0.212013
Celebrity Recommendation with Collaborative Social Topic Regression · IJCAI 2013
Mathematical optimization › discrete optimization
binary integer programming
0.112020
Delivery Scope: A New Way of Restaurant Retrieval for On-demand Food Delivery Service · KDD 2020
Mathematical optimization
combinatorial optimization
0.112020
Delivery Scope: A New Way of Restaurant Retrieval for On-demand Food Delivery Service · KDD 2020

Methods — techniques the papers use, named apart from their topics

graph representation learning · 2.3spatial computational techniques · 1.3machine learning · 1.3heuristic search · 1.3data mining · 1.3branch-and-bound · 1.3variational distance · 1.0transformer · 1.0conservative quantification · 1.0sequential policy optimization · 0.9quantal response equilibrium · 0.9sample reweighting · 0.6collaborative topic regression · 0.2
YearPublicationVenuePosition
2026 Reliability-Guaranteed and Reward-Seeking Sequence Modeling for Model-Based Offline Reinforcement Learning
abstract
As a data-driven learning approach, model-based offline reinforcement learning (MORL) aims to learn a policy by exploiting a dynamics model derived from an existing dataset. Applying conservative quantification to the dynamics model, most existing works on MORL generate trajectories that approximate the real data distribution to facilitate policy learning. However, these methods typically overlook the influence of historical information on environmental dynamics, thus generating unreliable trajectories that fail to align with the true data distribution. In this paper, we propose a new MORL algorithm called Reliability-guaranteed and Reward-seeking Transformer (RT). RT can avoid generating unreliable trajectories through the calculation of cumulative reliability of the trajectories, which is a weighted variational distance between the generated trajectory distribution and the true data distribution. Moreover, by sampling candidate actions with high rewards, RT can efficiently generate high-reward trajectories from the existing offline data, thereby further facilitating policy learning. We theoretically prove the performance guarantees of RT in policy learning, and empirically demonstrate its effectiveness against state-of-the-art model-based methods on several offline benchmark tasks and a large-scale industrial dataset from an on-demand food delivery platform.
Shenghong He, Chao Yu 0004, Yile Liang, Xuetao Ding
AAAI6
2025 Offline Multi-Agent Reinforcement Learning via In-Sample Sequential Policy Optimization
abstract
Offline Multi-Agent Reinforcement Learning (MARL) is an emerging field that aims to learn optimal multi-agent policies from pre-collected datasets. Compared to single-agent case, multi-agent setting involves a large joint state-action space and coupled behaviors of multiple agents, which bring extra complexity to offline policy optimization. In this work, we revisit the existing offline MARL methods and show that in certain scenarios they can be problematic, leading to uncoordinated behaviors and out-of-distribution (OOD) joint actions. To address these issues, we propose a new offline MARL algorithm, named In-Sample Sequential Policy Optimization (InSPO). InSPO sequentially updates each agent's policy in an in-sample manner, which not only avoids selecting OOD joint actions but also carefully considers teammates' updated policies to enhance coordination. Additionally, by thoroughly exploring low-probability actions in the behavior policy, InSPO can well address the issue of premature convergence to sub-optimal solutions. Theoretically, we prove InSPO guarantees monotonic policy improvement and converges to quantal response equilibrium (QRE). Experimental results demonstrate the effectiveness of our method compared to current state-of-the-art offline MARL methods.
Zongkai Liu, Chao Yu 0004, Xiawei Wu, Yile Liang, Xuetao Ding
AAAI7
2024 Harvesting Efficient On-Demand Order Pooling from Skilled Couriers: Enhancing Graph Representation Learning for Refining Real-time Many-to-One Assignments
abstract
The recent past has witnessed a notable surge in on-demand food delivery (OFD) services, offering delivery fulfillment within dozens of minutes after an order is placed. In OFD, pooling multiple orders for simultaneous delivery in real-time order assignment is a pivotal efficiency source, which may in turn extend delivery time. Constructing high-quality order pooling to harmonize platform efficiency with the experiences of consumers and couriers, is crucial to OFD platforms. However, the complexity and real-time nature of order assignment, making extensive calculations impractical, significantly limit the potential for order consolidation. Moreover, offline environment is frequently riddled with unknown factors, posing challenges for the platform's perceptibility and pooling decisions.
Yile Liang, Jiuxia Zhao, Jie Feng 0002, Xuetao Ding, Jinghua Hao, Renqing He
KDD6
2024 Order Dispatching Via GNN-Based Optimization Algorithm for On-Demand Food Delivery
abstract
As one representative of last-mile logistics in intelligent transportation systems, the on-demand food delivery (OFD) service has gained rapid market growth but also faces multiple challenges. One of the critical issues is the order dispatching problem (ODP) with an NP-hard nature, which refers to dispatching a large number of orders to riders reasonably in real time with very limited decision time. To address the ODP, this paper proposes an optimization algorithm based on graph neural networks (GNN) by combining the advantages of machine learning (ML) techniques and operational research (OR) methods: 1) The ML component learns to reduce the solution space by filtering out inappropriate riders for each order, handling the large-scale complexity of ODP. Specifically, we present a rider modeling approach by using GNN to better characterize rider information; besides, two attention mechanisms are designed to adaptively learn the matching relationship between riders and orders. 2) The OR component ensures the solution quality with a greedy and regret value-based dispatching heuristic. Extensive experiments are conducted on real-world datasets to evaluate the performance of the proposed method by comparing it with other existing models and algorithms. The results show that the design of our ML model is effective in yielding better prediction results, and the proposed GNN-based optimization algorithm can effectively and efficiently solve the ODP by improving delivery efficiency and customer satisfaction.
Jingfang Chen 0001, Ling Wang 0001, Yile Liang, Yang Yu 0005, Jiuxia Zhao, Xuetao Ding
IEEE Trans. Intell. Transp. Syst.7
2023 Enhancing Dynamic On-demand Food Order Dispatching via Future-informed and Spatial-temporal Extended Decisions
abstract
On-demand food delivery (OFD) service has gained fast-growing popularity all around the world. Order dispatching is instrumental to large-scale OFD platforms, such as Meituan, which continuously match food order requests to couriers at a scale of tens of millions each day to satisfy the needs of consumers, couriers, and merchants. However, due to high dynamism and inevitable uncertainties in the real-world environment, it is not an easy task to achieve long-term global objective optimization through continuous isolated optimization decisions at each dispatch moment. Our work proposes the concept of "courier occupancy" (CO) to precisely quantify the impact of order assignment on the courier's delivery efficiency, realizing a decomposition of long-term and macro goals into various dispatch moments and micro decision-making dimensions. Then in the prediction phase, an improved and universally applicable distribution estimation method is designed to quantify CO which is a stochastic variable and contains future information, combining Monte Carlo dropout and knowledge distillation. In the optimization phase, we use CO to model the objective function at each dispatch moment to introduce future information and extend dispatch decisions from merely who to assign the order to both when and who to assign it, significantly enhancing the long-term optimization capability of dispatching decisions and avoiding local greed. We conduct extensive offline simulations based on real dispatching data as well as online AB tests through Meituan's platform. Results show that our method consistently improves the couriers' delivery efficiency and consumers' satisfaction.
Yile Liang, Jiuxia Zhao, Xuetao Ding, Huanjia Lian, Jinghua Hao, Renqing He
CIKM4
2023 A Predictive-Reactive Optimization Framework With Feedback-Based Knowledge Distillation for On-Demand Food Delivery
abstract
On-demand food delivery (OFD) service is a representative scenario of last-mile logistics. It has gained a fast-growing market but also encounters many challenges, e.g., high dynamism, large-scale complexity, and immediacy requirement. To solve the OFD problem, a predictive-reactive optimization framework with feedback-based knowledge distillation is presented by organically combining deep learning technology and operational research method. In the prediction phase, a deep learning model is designed to predict future information which can reflect the delivery efficiency of dispatching results. To improve model performance, a feedback-based knowledge distillation is proposed which balances the diversity and effectiveness of the ensembled models by adaptively controlling learning weights. In the optimization phase, to avoid myopic decisions and obtain high-quality solutions for long-term objectives, a greedy heuristic with a multi-stage decision-making strategy is designed by employing the predicted future information to assist in making decisions. Extensive experiments are conducted on real-world datasets to test the performance of the proposed model and heuristic. Besides, the simulation results illustrate the superiority of the proposed framework for solving the OFD problem in both delivery efficiency and customer satisfaction.
Jie Zheng 0001, Ling Wang 0001, Jingfang Chen 0001, Zi-Xiao Pan, Yile Liang, Xuetao Ding
IEEE Trans. Intell. Transp. Syst.7
2022 Counterfactual Prediction for Outcome-Oriented Treatments
abstract
Large amounts of efforts have been devoted into learning counterfactual treatment outcome under various settings, including binary/continuous/multiple treatments. Most of these literature aims to minimize the estimation error of counterfactual outcome for the whole treatment space. However, in most scenarios when the counterfactual prediction model is utilized to assist decision-making, people are only concerned with the small fraction of treatments that can potentially induce superior outcome (i.e. outcome-oriented treatments). This gap of objective is even more severe when the number of possible treatments is large, for example under the continuous treatment setting. To overcome it, we establish a new objective of optimizing counterfactual prediction on outcome-oriented treatments, propose a novel Outcome-Oriented Sample Re-weighting (OOSR) method to make the predictive model concentrate more on outcome-oriented treatments, and theoretically analyze that our method can improve treatment selection towards the optimal one. Extensive experimental results on both synthetic datasets and semi-synthetic datasets demonstrate the effectiveness of our method.
Hao Zou 0001, Bo Li 0064, Jiangang Han, Shuiping Chen, Xuetao Ding, Peng Cui 0001
ICML5
2020 Delivery Scope: A New Way of Restaurant Retrieval for On-demand Food Delivery Service
abstract
Recently on-demand food delivery service has become very popular in China. More than 30 million orders are placed by eaters of Meituan-Dianping everyday. Delicacies are delivered to eaters in 30 minutes on average. To fully leverage the ability of our couriers and restaurants, delivery scope is proposed as an infrastructure product for on-demand food delivery area. A delivery scope based retrieval system is designed and built on our platform. In order to draw suitable delivery scopes for millions of restaurant partners, we propose a pioneering delivery scope generation framework. In our framework, a single delivery scope generation algorithm is proposed by using spatial computational techniques and data mining techniques. Moreover, a scope scoring algorithm and decision algorithm are proposed by utilizing machine learning models and combinatorial optimization techniques. Specifically, we propose a novel delivery scope sample generation method and use the scope related features to estimate order numbers and average delivery time in a period of time for each delivery scope. Then we formalize the candidate scopes selection process as a binary integer programming problem. Both branch&bound algorithm and a heuristic search algorithm are integrated in our system. Results of online experiments show that scopes generated by our new algorithm significantly outperform manual generated ones. Our algorithm brings more orders without hurt of users' experience. After deployed online, our system has saved thousands of hours for operation staff, and it is considered to be one of the most useful operation tools to balance demand of eaters and supply of restaurants and couriers.
Xuetao Ding, Runfeng Zhang, Zhen Mao, Fangxiao Du, Guoxing Wei, Feifan Yin, Renqing He, Zhizhao Sun
KDD1
2014 User Interests Imbalance Exploration in Social Recommendation: A Fitness Adaptation
abstract
Recent years have witnessed an increasing interest in how to incorporate social network information into recommendation algorithms to enhance the user experience. In this paper, we find the phenomenon that users in the contexts of recommendation system and social network do not share the same interest space. Based on this finding, we proposed the social regulatory factor regression model (SRFRM) which could connect different interest spaces in different contexts together in an unified latent factor model. Specifically, different from the traditional social based latent factor models with strong limitation that all sides share the same feature space, the proposed method leverages the regulatory factor number on both sides to meet the fact that users and items or users in different contexts may not share the same interest space. It works by incorporating two linear transformation matrices into the matrix co-factorization framework that matrix factorization of user ratings is regularized by that of social trust network. We study a large subsets of data from epinions.com and douban.com respectively. The experimental results indicate that users in different contexts have different interest spaces and our model achieves a higher performance compared with related state-of-the-art methods.
Tianchun Wang, Xiaoming Jin, Xuetao Ding
CIKM3
2013 Short text classification by detecting information path
abstract
Short text is becoming ubiquitous in many modern information systems. Due to the shortness and sparseness of short texts, there are less informative word co-occurrences among them, which naturally pose great difficulty for classification tasks on such data. To overcome this difficulty, this paper proposes a new way for effectively classifying the short texts. Our method is based on a key observation that there usually exists ordered subsets in short texts, which is termed ``information path'' in this work, and classification on each subset based on the classification results of some pervious subsets can yield higher overall accuracy than classifying the entire data set directly. We propose a method to detect the information path and employ it in short text classification. Different from the state-of-art methods, our method does not require any external knowledge or corpus that usually need careful fine-tuning, which makes our method easier and more robust on different data sets. Experiments on two real world data sets show the effectiveness of the proposed method and its superiority over the existing methods.
Shitao Zhang, Xiaoming Jin, Dou Shen, Bin Cao 0001, Xuetao Ding
CIKM5
2013 Celebrity Recommendation with Collaborative Social Topic Regression
Xuetao Ding, Xiaoming Jin, Lianghao Li
IJCAI1