Zhaoxing Yang

dblp:298/3326 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-7008-7629ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Breaking the Scalability Barrier in Constrained Graph-Based Networked Control via Decision-Focused Learning
abstract
Many real-world systems can be modeled as graphs, where nodes store and consume entities, actively produce them, or have them emerge naturally, and edges transport them between nodes. This paper studies the networked control problem on such large-scale systems, aiming to decide production and transportation over time to maximize long-term profits, subject to node or edge capacity constraints. Existing SOTAs either fail to guarantee feasibility or cannot scale to large-scale systems. We propose a two-stage policy that integrates a constrained optimization layer after a neural network to explicitly enforce constraints and ensure feasibility. By leveraging the problem structure to obtain expert actions and designing a decision-focused and differentiable loss to enable imitation learning, our method significantly improves efficiency and scalability. In small-scale systems with action dimensions in the order of 10, our method achieves 60x sample efficiency over SOTAs on average. In large-scale systems with action dimensions ranging from 100 to 100000, where SOTAs fail to train, our method converges quickly and outperforms non-learning-based baselines significantly.
Zhaoxing Yang, Guiyun Fan, Haiming Jin, Linghe Kong
WWW1
2026 Scalable Traffic Allocation in Dynamic Networks via End-to-End Imitation Learning
Zhaoxing Yang, Guiyun Fan, Anjie Cao, Chenhao Ying 0001, Shengnan Yue, Haiming Jin
IEEE Trans. Netw.1
2025 Learning to Accelerate Traffic Allocation Over Large-Scale Networks
Zhaoxing Yang, Guiyun Fan, Anjie Cao, Haiming Jin
INFOCOM1
2024 Rethinking Order Dispatching in Online Ride-Hailing Platforms
abstract
Achieving optimal order dispatching has been a long-standing challenge for online ride-hailing platforms. Early methods would make shortsighted matchings as they only consider order prices alone as the edge weights in the driver-order bipartite graph, thus harming the platform's revenue. To address this problem, recent works evaluate the value of the order's destination region to be the long-term income a driver could obtain in average in such region and incorporate it into the order's edge weight to influence the matching results. However, they often result in insufficient driver supplies in many regions, as the values evaluated in different regions vary greatly, mainly because the impact of one region's value on the future number of drivers and revenue in other regions is overlooked. This paper models such impact within a cooperative Markov game, which involves each value's impact over the platform's revenue with the goal to find the optimal region values for revenue maximization. To solve this game, our work proposes a novelgoal-reaching collaboration (GRC) algorithm that realizes credit assignment from a novel goal-reaching perspective, addressing the difficulty for accurate credit assignment with large-scale agents of previous methods and resolving the conflict between credit assignment and offline reinforcement learning. Specifically, during training, GRC predicts the city's future state through an environment model and utilizes a scoring model to rate the predicted states to judge their levels of profitability, where high-scoring states are regarded as the goal states. Then, the policies in the game are updated to promote the city to stay in the goal states for as long as possible. To evaluate GRC, we deploy a baseline policy online in several cities for three weeks to collect real-world dataset. Training and testing results on the collected dataset indicate that our GRC consistently outperforms the baselines in different cities and peak periods.
Zhaoxing Yang, Haiming Jin, Guiyun Fan, Min Lu 0004, Xinlang Yue, Zhe Xu 0003, Guobin Wu 0001, Jiecheng Guo
KDD1
2024 Optimizing Long-Term Efficiency and Fairness in Ride-Hailing Under Budget Constraint via Joint Order Dispatching and Driver Repositioning
abstract
Ride-hailing platforms (e.g., Uber and Didi Chuxing) have become increasingly popular in recent years.Efficiencyhas always been an important metric for such platforms. However, only focusing on efficiency inevitably ignores thefairnessof driver incomes, which could impair the sustainability of ride-hailing systems. To optimize such two essential objectives,order dispatchinganddriver repositioningplay an important role, as they impact not only the immediate, but also the future order-serving outcomes of drivers. In practice, the platform offers monetary incentives to drivers for completing the repositioning and has a budget for the repositioning cost. Therefore, in this paper, we aim to exploit joint order dispatching and driver repositioning to optimize both long-term efficiency and fairness in ride-hailing under the budget constraint. To this end, we propose JDRCL, a novel multi-agent reinforcement learning framework, which integrates a group-based action representation that copes with the variable action space, and a primal-dual iterative training algorithm to learn a constraint-satisfying policy that maximizes both the worst and the overall incomes of drivers. Furthermore, we prove the asymptotic convergence rate of our training algorithm. Extensive experiments based on three real-world ride-hailing order datasets show that JDRCL outperforms state-of-the-art baselines on both efficiency and fairness.
Haiming Jin, Zhaoxing Yang, Lu Su 0001
IEEE Trans. Knowl. Data Eng.3
2023 DeCOM: Decomposed Policy for Constrained Cooperative Multi-Agent Reinforcement Learning
abstract
In recent years, multi-agent reinforcement learning (MARL) has presented impressive performance in various applications. However, physical limitations, budget restrictions, and many other factors usually impose constraints on a multi-agent system (MAS), which cannot be handled by traditional MARL frameworks. Specifically, this paper focuses on constrained MASes where agents work cooperatively to maximize the expected team-average return under various constraints on expected team-average costs, and develops a constrained cooperative MARL framework, named DeCOM, for such MASes. In particular, DeCOM decomposes the policy of each agent into two modules, which empowers information sharing among agents to achieve better cooperation. In addition, with such modularization, the training algorithm of DeCOM separates the original constrained optimization into an unconstrained optimization on reward and a constraints satisfaction problem on costs. DeCOM then iteratively solves these problems in a computationally efficient manner, which makes DeCOM highly scalable. We also provide theoretical guarantees on the convergence of DeCOM's policy update algorithm. Finally, we conduct extensive experiments to show the effectiveness of DeCOM with various types of costs in both moderate-scale and large-scale (with 500 agents) environments that originate from real-world applications.
Zhaoxing Yang, Haiming Jin, Haoyi You, Guiyun Fan, Xinbing Wang, Chenghu Zhou
AAAI1
2023 User-Oriented Robust Reinforcement Learning
abstract
Recently, improving the robustness of policies across different environments attracts increasing attention in the reinforcement learning (RL) community. Existing robust RL methods mostly aim to achieve the max-min robustness by optimizing the policy’s performance in the worst-case environment. However, in practice, a user that uses an RL policy may have different preferences over its performance across environments. Clearly, the aforementioned max-min robustness is oftentimes too conservative to satisfy user preference. Therefore, in this paper, we integrate user preference into policy learning in robust RL, and propose a novel User-Oriented Robust RL (UOR-RL) framework. Specifically, we define a new User-Oriented Robustness (UOR) metric for RL, which allocates different weights to the environments according to user preference and generalizes the max-min robustness metric. To optimize the UOR metric, we develop two different UOR-RL training algorithms for the scenarios with or without a priori known environment distribution, respectively. Theoretically, we prove that our UOR-RL training algorithms converge to near-optimal policies even with inaccurate or completely no knowledge about the environment distribution. Furthermore, we carry out extensive experimental evaluations in 6 MuJoCo tasks. The experimental results demonstrate that UOR-RL is comparable to the state-of-the-art baselines under the average-case and worst-case performance metrics, and more importantly establishes new state-of-the-art performance under the UOR metric.
Haoyi You, Beichen Yu, Haiming Jin, Zhaoxing Yang
AAAI4
2023 Multi-Intersection Management for Connected Autonomous Vehicles by Reinforcement Learning
abstract
The rapid development of connected autonomous vehicles (CAVs) makes it foreseeable that CAVs will dominate future road traffic. To manage CAV traffic, researchers developed a revolutionary paradigm, which uses intelligent intersection managers (IMs) for a finer-grained control of CAVs' cruising at intersections than traditional traffic lights. However, existing IM-based methods mostly focus on optimizing the single-intersection CAV traffic efficiency, without solving the fundamental problem of maximizing the global efficiency of a multi-intersection road network. Therefore, we address such problem by proposing a system architecture that decomposes each IM into an oracle and a valve, where the oracle ensures safe and efficient crossing at individual intersections, and the valve selects some of the approaching CAVs for the oracle to control and postpones the crossing of the unselected ones. We further focus on distributed decision making for the valves, and propose a multi-agent reinforcement learning framework, spatial-aware multi-agent actor-credit (SMAC). Specifically, SMAC integrates a novel credit assignment method that captures agents' spatially decaying influences to stimulate agent cooperation, and a novel graph convolutional mixing network to capture the graph-structured inter-agent relationships in a road network. We conduct extensive experiments on three traffic flow datasets, and show that SMAC outperforms state-of-the-art baselines.
Haiming Jin, Yifei Wei, Zhaoxing Yang, Zirui Liu 0009, Guiyun Fan
ICDCS3
2022 Optimizing Long-Term Efficiency and Fairness in Ride-Hailing via Joint Order Dispatching and Driver Repositioning
abstract
The ride-hailing service offered by mobility-on-demand platforms, such as Uber and Didi Chuxing, has greatly facilitated people's traveling and commuting, and become increasingly popular in recent years. Efficiency (e.g., gross merchandise volume) has always been an important metric for such platforms. However, only focusing on the efficiency inevitably ignores the fairness of driver incomes, which could impair the sustainability of the overall ride-hailing system in the long run. To optimize the aforementioned two essential metrics, order dispatching and driver repositioning play an important role, as they impact not only the immediate, but also the future order-serving outcomes of drivers. Thus, in this paper, we aim to exploit joint order dispatching and driver repositioning to optimize both the long-term efficiency and fairness for ride-hailing platforms. To address this problem, we propose a novel multi-agent reinforcement learning framework, referred to as JDRL, to help drivers make distributed order selection and repositioning decisions. Specifically, to cope with the variable action space, JDRL segments the action space into a fixed number of action groups, and fixes the policy output dimension for order selection as the number of action groups. In terms of the fairness criterion, JDRL adopts the max-min fairness, and augments the vanilla policy gradient to an iterative training algorithm that alternates between a minimization step and a policy improvement step to maximize both the worst and the overall performance of agents. In addition, we provide the theoretical convergence guarantee of our JDRL training algorithm even under non-convex policy networks and stochastic gradient updating. Extensive experiments are conducted with three public real-world ride-hailing order datasets, including over 2 million orders in Haikou, China, over 5 million orders in Chengdu, China, and over 6 million orders in New York City, USA. Experimental results show that JDRL demonstrates a consistent advantage compared to state-of-the-art baselines in terms of both efficiency and fairness. To the best of our knowledge, this is the first work that exploits joint order dispatching and driver repositioning to optimize both the long-term efficiency and fairness in a ride-hailing system.
Haiming Jin, Zhaoxing Yang, Lu Su 0001, Xinbing Wang
KDD3
2022 Enabling Optimal Control Under Demand Elasticity for Electric Vehicle Charging Systems
abstract
Recent years have witnessed the proliferation of electric vehicles (EVs) that enable environment-friendly commuting and traveling. However, the increasing number of EVs inevitably create massive charging demands that are challenging to satisfy. Oftentimes in practice, EVs have to wait in queues for a long time outside charging stations before chargers become available. To address this challenge, we fully capture the elasticity of EVs’ charging demands in response to the charging prices, and propose a dynamic charging pricing mechanism that jointly controls the lengths of the demand queues at multiple charging stations and maximizes the charging platform’s long-term profit for offering charging services. Clearly, such an approach is more feasible than the financially and temporally expensive way of constructing extra charging facilities. Technically, we augment the Lyapunov stochastic optimization technique to decompose the challenging long-term decision-making problem into a series of single-time-slot optimization programs which require zero knowledge of future system parameters. However, due to the correlation of charging demands among different stations, the aforementioned optimization program in each time slot is non-convex. We handle the non-convexity by jointly constructing independent sets of charging stations and adapting the block coordinate descent method to iteratively obtain approximately optimal charging prices. Through rigorous theoretical analysis and extensive simulations based on the real-world dataset in the Chinese city Shenzhen which consists of 4000 taxis and 171 charging stations, we demonstrate that our control policy ensures an arbitrarily close-to-optimal profit with a flexible trade-off between the profit and queue lengths, has a low computational complexity, and requires zero knowledge of future system dynamics.
Guiyun Fan, Zhaoxing Yang, Haiming Jin, Xiaoying Gan, Xinbing Wang
IEEE Trans. Mob. Comput.2
2022 Joint Charging and Relocation Recommendation for E-Taxi Drivers via Multi-Agent Mean Field Hierarchical Reinforcement Learning
abstract
Nowadays, most of the taxi drivers have become users of the relocation recommendation service offered by online ride-hailing platforms (e.g., Uber and Didi Chuxing), which could oftentimes lead drivers to places with profitable orders. At the same time, electric taxis (e-taxis) are increasingly adopted and gradually replacing gasoline taxis in today’s public transportation systems due to their environmental-friendly nature. Though effective for traditional gasoline taxis, existing relocation recommendation schemes are rather suboptimal for e-taxi drivers’ user experience. On one hand, the existing schemes take no account of taxis’ refueling decisions, as the refueling durations of gasoline taxis are usually short enough to be ignored. However, the charging duration of the e-taxis spent at charging stations can be as long as hours. Obviously, an e-taxi’s battery could be easily depleted by the continuous relocations suggested by existing schemes, and thus will have to be charged for a long time afterwards, making the e-taxi driver miss numerous order-serving opportunities. On the other hand, charging posts are typically sparsely and unevenly distributed across a city. With no consideration of charging opportunities, existing schemes could probably send an e-taxi to an area with no charging post around, even though its battery is running low. To optimize e-taxi drivers’ user experience, in this paper, we design a jointcharging and relocation recommendation system for e-taxi drivers (CARE). We take the perspective of e-taxi drivers and formulate their decision making as a multi-agent reinforcement learning problem where each e-taxi driver aims to maximize his own cumulative rewards. More specifically, we propose a novelmulti-agent mean field hierarchical reinforcement learning (MFHRL)framework. The hierarchical architecture of MFHRL helps the proposed CARE provide far-sighted charging and relocation recommendations for e-taxi drivers. Besides, we integrate each hierarchical level of MFHRL separately with the mean field approximation to incorporate e-taxis’ mutual influences in decision making. We set up a simulator with one of the largest real-world e-taxi datasets in Shenzhen, China, which contains the GPS trajectory data and transaction data of 3848 e-taxis from June 1st to June 30th, 2017, coupled with 165 charging stations including 317 fast charging posts and 1421 slow charging posts. We adopt this simulator to generate 6 dynamic urban environments, which reflect the different real-world scenarios faced by e-taxi drivers. In all of these environments, we conduct extensive experiments to validate that the proposed MFHRL framework greatly outperforms all baselines by significantly increasing the rewards obtained by e-taxi drivers. Besides, we also show that the charging policy learned by MFHRL can effectively reduce the range anxiety of e-taxi drivers, which significantly boosts e-taxi drivers’ quality of experience.
Enshu Wang, Zhaoxing Yang, Haiming Jin, Chenglin Miao, Lu Su 0001, Fan Zhang 0019, Chunming Qiao, Xinbing Wang
IEEE Trans. Mob. Comput.3
2021 Constrained Multi-Agent Reinforcement Learning for Managing Electric Self-Driving Taxis
abstract
Electric self-driving taxis (es-taxis) draw great attention nowadays and hold the promise for future transportation due to their convenient and environment-friendly nature. However efficiently managing large-scale es-taxis remains an open problem. In this paper, we focus on scheduling es-taxis under charging budget constraint. Specifically, we design safe-controller to guarantee the satisfaction of budget constraint, and propose HAT framework to enlarge the sight for decision-making on deactivating es-taxis. As for the non-stationary induced by HAT, we analyze and limit its influence with theoretical guarantees. The overall framework Safe-HAT achieves superior performance in real-world data against other strong baselines.
Zhaoxing Yang, Guiyun Fan, Haiming Jin
ICPADS1
2021 Multi-Agent Reinforcement Learning for Urban Crowd Sensing with For-Hire Vehicles
abstract
Recently, vehicular crowd sensing (VCS) that leverages sensor-equipped urban vehicles to collect city-scale sensory data has emerged as a promising paradigm for urban sensing. Nowadays, a wide spectrum of VCS tasks are carried out by for-hire vehicles (FHVs) due to various hardware and software constraints that are difficult for private vehicles to satisfy. However, such FHV-enabled VCS systems face a fundamental yet unsolved problem of striking a balance between the order-serving and sensing outcomes. To address this problem, we propose a novel graph convolutional cooperative multi-agent reinforcement learning (GCC-MARL) framework, which helps FHVs make distributed routing decisions that cooperatively optimize the system-wide global objective. Specifically, GCC-MARL meticulously assigns credits to agents in the training process to effectively stimulate cooperation, represents agents' actions by a carefully chosen statistics to cope with the variable agent scales, and integrates graph convolution to capture useful spatial features from complex large-scale urban road networks. We conduct extensive experiments with a real-world dataset collected in Shenzhen, China, containing around 1 million trajectories and 50 thousand orders of 553 taxis per-day from June 1st to 30th, 2017. Our experiment results show that GCC-MARL outperforms state-of-the-art baseline methods in order-serving revenue, as well as sensing coverage and quality.
Zhaoxing Yang, Yifei Wei, Haiming Jin, Xinbing Wang
INFOCOM2