Yuxin Pan

dblp:29/3085 · DBLP profile ↗
← Back
12ranked-venue papers
8as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 6 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 75% Autonomous driving · 12% Motion planning and robot control · 12%
Theoretical computer science
1 paper
Mathematical optimization · 50% Approximation and online algorithms · 50%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › safe reinforcement learning
constrained policy optimization
0.912025
Efficient Discovery of Pareto Front for Multi-Objective Reinforcement Learning · ICLR 2025
Machine learning › Reinforcement learning
multi-objective reinforcement learning
0.912025
Efficient Discovery of Pareto Front for Multi-Objective Reinforcement Learning · ICLR 2025
Machine learning › Reinforcement learning
robust reinforcement learning
0.712023
Adjustable Robust Reinforcement Learning for Online 3D Bin Packing · NeurIPS 2023
Approximation and online algorithms
bin packing
0.712023
Adjustable Robust Reinforcement Learning for Online 3D Bin Packing · NeurIPS 2023
Mathematical optimization
combinatorial optimization
0.712023
Adjustable Robust Reinforcement Learning for Online 3D Bin Packing · NeurIPS 2023
Robotics › Motion planning and robot control › robot control › learning control
imitation reinforcement learning
0.412020
Navigation Command Matching for Vision-based Autonomous Driving · ICRA 2020
Machine learning › Reinforcement learning
deep reinforcement learning
0.212023
Adjustable Robust Reinforcement Learning for Online 3D Bin Packing · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

robust adversarial reinforcement learning · 1.3permutation-based attacker · 1.3mixture dynamics · 1.3preference-conditioned policy · 0.9constrained optimization · 0.9smooth rewards · 0.4replay buffer · 0.4attention-guided agent · 0.4
YearPublicationVenuePosition
2026 Multi-level graph partition via hierarchical learning for large-scale vehicle routing problems
abstract
Abstract Vehicle Routing Problems (VRPs) involve multi-agent route optimization, with the objective of targeting optimal routes for a fleet of vehicles to serve a set of customers. Existing neural solvers based on the divide-and-conquer approach for VRPs in general, and capacitated VRP (CVRP) in particular, integrate the global partition of an instance with the local construction for each resulting subinstance to enhance generalization. However, during the global partition phase, misclusterings within subgraphs have a tendency to progressively compound throughout the multi-step decoding process of the learning-based partition policy. This suboptimal behavior of the partition policy, in turn, may lead to a dramatic deterioration in the performance of the overall decomposition-based system, despite using optimal local constructions. To address these challenges, we propose a versatile Hierarchical Learning-based Graph Partition (HLGP) framework, which is tailored to benefit the partition of CVRP instances by synergistically integrating global and local partition policies. Specifically, the global partition policy is tasked with creating a coarse multi-way partition to generate a sequence of simpler two-way partition subtasks. These subtasks mark the initiation of the subsequent K local partition levels. At each local partition level, subtasks exclusive to this level are assigned to the local partition policy which benefits from the insensitive local topological features to incrementally alleviate the compounded errors. This framework is versatile in the sense that it optimizes the involved partition policies towards a unified objective, which is harmoniously compatible with both reinforcement learning (RL) and supervised learning (SL) paradigms. Additionally, we decouple the synchronized training into individual training of each component to circumvent the instability issue. Furthermore, we point out the importance of treating the subproblems encountered during the partition process as individual training instances. Extensive experiments conducted on various CVRP benchmarks demonstrate the effectiveness and generalization capabilities of the HLGP framework under both scale and distribution shifts. The source code is available at https://github.com/panyxy/hlgp_cvrp .
Yuxin Pan, Ruohong Liu, Yize Chen, Zhiguang Cao, Fangzhen Lin
Auton. Agents Multi Agent Syst.1
2025 Efficient Discovery of Pareto Front for Multi-Objective Reinforcement Learning
abstract
Multi-objective reinforcement learning (MORL) excels at handling rapidly changing preferences in tasks that involve multiple criteria, even for unseen preferences. However, previous dominating MORL methods typically generate a fixed policy set or preference-conditioned policy through multiple training iterations exclusively for sampled preference vectors, and cannot ensure the efficient discovery of the Pareto front. Furthermore, integrating preferences into the input of policy or value functions presents scalability challenges, in particular as the dimension of the state and preference space grow, which can complicate the learning process and hinder the algorithm's performance on more complex tasks. To address these issues, we propose a two-stage Pareto front discovery algorithm called Constrained MORL (C-MORL), which serves as a seamless bridge between constrained policy optimization and MORL. Concretely, a set of policies is trained in parallel in the initialization stage, with each optimized towards its individual preference over the multiple objectives. Then, to fill the remaining vacancies in the Pareto front, the constrained optimization steps are employed to maximize one objective while constraining the other objectives to exceed a predefined threshold. Empirically, compared to recent advancements in MORL methods, our algorithm achieves more consistent and superior performances in terms of hypervolume, expected utility, and sparsity on both discrete and continuous control tasks, especially with numerous objectives (up to nine objectives in our experiments).
Ruohong Liu, Yuxin Pan, Linjie Xu, Lei Song 0001, Pengcheng You, Yize Chen, Jiang Bian 0002
ICLR2
2025 Hierarchical Learning-based Graph Partition for Large-scale Vehicle Routing Problems
Yuxin Pan, Ruohong Liu, Yize Chen, Zhiguang Cao, Fangzhen Lin
AAMAS1
2025 Multi-Task Vehicle Routing Solver via Mixture of Specialized Experts under State-Decomposable MDP
abstract
Existing neural methods for multi-task vehicle routing problems (VRPs) typically learn unified solvers to handle multiple constraints simultaneously. However, they often underutilize the compositional structure of VRP variants, each derivable from a common set of basis VRP variants. This critical oversight causes unified solvers to miss out the potential benefits of basis solvers, each specialized for a basis VRP variant. To overcome this limitation, we propose a framework that enables unified solvers to perceive the shared-component nature across VRP variants by proactively reusing basis solvers, while mitigating the exponential growth of trained neural solvers. Specifically, we introduce a State-Decomposable MDP (SDMDP) that reformulates VRPs by expressing the state space as the Cartesian product of basis state spaces associated with basis VRP variants. More crucially, this formulation inherently yields the optimal basis policy for each basis VRP variant. Furthermore, a Latent Space-based SDMDP extension is developed by incorporating both the optimal basis policies and a learnable mixture function to enable the policy reuse in the latent space. Under mild assumptions, this extension provably recovers the optimal unified policy of SDMDP through the mixture function that computes the state embedding as a mapping from the basis state embeddings generated by optimal basis policies. For practical implementation, we introduce the Mixture-of-Specialized-Experts Solver (MoSES), which realizes basis policies through specialized Low-Rank Adaptation (LoRA) experts, and implements the mixture function via an adaptive gating mechanism. Extensive experiments conducted across VRP variants showcase the superiority of MoSES over prior methods.
Yuxin Pan, Zhiguang Cao, Chengyang Gu, Liu Liu 0014, Peilin Zhao, Yize Chen, Fangzhen Lin
NeurIPS1
2023 Adjustable Robust Reinforcement Learning for Online 3D Bin Packing
abstract
Designing effective policies for the online 3D bin packing problem (3D-BPP) has been a long-standing challenge, primarily due to the unpredictable nature of incoming box sequences and stringent physical constraints. While current deep reinforcement learning (DRL) methods for online 3D-BPP have shown promising results in optimizing average performance over an underlying box sequence distribution, they often fail in real-world settings where some worst-case scenarios can materialize. Standard robust DRL algorithms tend to overly prioritize optimizing the worst-case performance at the expense of performance under normal problem instance distribution. To address these issues, we first introduce a permutation-based attacker to investigate the practical robustness of both DRL-based and heuristic methods proposed for solving online 3D-BPP. Then, we propose an adjustable robust reinforcement learning (AR2L) framework that allows efficient adjustment of robustness weights to achieve the desired balance of the policy's performance in average and worst-case environments. Specifically, we formulate the objective function as a weighted sum of expected and worst-case returns, and derive the lower performance bound by relating to the return under a mixture dynamics. To realize this lower bound, we adopt an iterative procedure that searches for the associated mixture dynamics and improves the corresponding policy. We integrate this procedure into two popular robust adversarial algorithms to develop the exact and approximate AR2L algorithms. Experiments demonstrate that AR2L is versatile in the sense that it improves policy robustness while maintaining an acceptable level of performance for the nominal case.
Yuxin Pan, Yize Chen, Fangzhen Lin
NeurIPS1
2022 Backward Imitation and Forward Reinforcement Learning via Bi-directional Model Rollouts
abstract
Traditional model-based reinforcement learning (RL) methods generate forward rollout traces using the learnt dynamics model to reduce interactions with the real environment. The recent model-based RL method considers the way to learn a backward model that specifies the conditional probability of the previous state given the previous action and the current state to additionally generate backward rollout trajectories. However, in this type of model-based method, the samples derived from backward rollouts and those from forward rollouts are simply aggregated together to optimize the policy via the model-free RL algorithm, which may decrease both the sample efficiency and the convergence rate. This is because such an approach ignores the fact that backward rollout traces are often generated starting from some high-value states and are certainly more instructive for the agent to improve the behavior. In this paper, we propose the backward imitation and forward reinforcement learning (BIFRL) framework where the agent treats backward rollout traces as expert demonstrations for the imitation of excellent behaviors, and then collects forward rollout transitions for policy reinforcement. Consequently, BIFRL empowers the agent to both reach to and explore from high-value states in a more efficient manner, and further reduces the real interactions, making it potentially more suitable for real-robot learning. Moreover, a value-regularized generative adversarial network is introduced to augment the valuable states which are infrequently received by the agent. Theoretically, we provide the condition where BIFRL is superior to the baseline methods. Experimentally, we demonstrate that BIFRL acquires the better sample efficiency and produces the competitive asymptotic performance on various MuJoCo locomotion tasks compared against state-of-the-art model-based methods.
Yuxin Pan, Fangzhen Lin
IROS1
2022 Attention-Based Sequence-to-Sequence Learning for Online Structural Response Forecasting Under Seismic Excitation
abstract
In structural health monitoring (SHM), measuring and evaluating structural dynamic responses are critical for safety management of civil infrastructures. Particularly, online forecasting of the structural responses under extreme external loading conditions (e.g., earthquakes) takes a significant role in SHM to provide early warning and ensure safe operation. In practice, complex causality and intrinsic interactions between seismic excitation and structural response make it challenging to establish a reliable predictive scheme. The present paper proposes a novel deep recurrent neural network (RNN) model implemented in the architecture of a time-series attention-based RNN encoder–decoder (TSA-RNN-ED), for predictive analysis of structural responses under seismic excitation. In the proposed data-driven model, upcoming sequential responses are predicted through sequence-to-sequence learning from historical multivariate time-series signals. A time-series attention mechanism is proposed to exploit the heterogeneous, but directly related, hidden features between the seismic loads and the corresponding structural responses. The proposed architecture can reliably regress excitation-response interactions to predict dynamic responses subjected to future earthquakes while satisfying the need of real-time forecasting for on-site practical implementation. This article systematically evaluates the proposed model by using two real-world structural cases: 1) the tallest building in China, the Shanghai Tower and 2) a woodframe classroom on a shake table at the University of British Columbia in Vancouver, Canada. The experimental results demonstrate the accurate and efficient performance of the proposed methodology in forecasting the seismic responses of the structures under investigation.
Teng Li 0005, Yuxin Pan, Kaitai Tong, Carlos E. H. Ventura, Clarence W. de Silva
IEEE Trans. Syst. Man Cybern. Syst.2
2020 Navigation Command Matching for Vision-based Autonomous Driving
abstract
Learning an optimal policy for autonomous driving task to confront with complex environment is a long- studied challenge. Imitative reinforcement learning is accepted as a promising approach to learn a robust driving policy through expert demonstrations and interactions with environments. However, this model utilizes non-smooth rewards, which have a negative impact on matching between navigation commands and trajectory (state-action pairs), and degrade the generalizability of an agent. Smooth rewards are crucial to discriminate actions generated from sub-optimal policy. In this paper, we propose a navigation command matching (NCM) model to address this issue. There are two key components in NCM, 1) a matching measurer produces smooth navigation rewards that measure matching between navigation commands and trajectory; 2) attention-guided agent performs actions given states where salient regions in RGB images (i.e. roadsides, lane markings and dynamic obstacles) are highlighted to amplify their influence on the final model. We obtain navigation rewards and store transitions to replay buffer after an episode, so NCM is able to discriminate actions generated from suboptimal policy. Experiments on CARLA driving benchmark show our proposed NCM outperforms previous state-of-the- art models on various tasks in terms of the percentage of successfully completed episodes. Moreover, our model improves generalizability of the agent and obtains good performance even in unseen scenarios.
Yuxin Pan, Jianru Xue, Pengfei Zhang 0005, Wanli Ouyang, Jianwu Fang, Xingyu Chen 0001
ICRA1
2020 Using Detection, Tracking and Prediction in Visual SLAM to Achieve Real-time Semantic Mapping of Dynamic Scenarios
abstract
In this paper, we propose a lightweight system, RDS-SLAM, based on ORB-SLAM2, which can accurately estimate poses and build semantic maps at object level for dynamic scenarios in real time using only one commonly used Intel Core i7 CPU. In RDS-SLAM, three major improvements, as well as major architectural modifications, are proposed to overcome the limitations of ORB-SLAM2. Firstly, it adopts a lightweight object detection neural network in key frames. Secondly, an efficient tracking and prediction mechanism is embedded into the system to remove the feature points belonging to movable objects in all incoming frames. Thirdly, a semantic octree map is built by probabilistic fusion of detection and tracking results, which enables a robot to maintain a semantic description at object level for potential interactions in dynamic scenarios. We evaluate RDS-SLAM in TUM RGB-D dataset, and experimental results show that RDS-SLAM can run with 30.3 ms per frame in dynamic scenarios using only an Intel Core i7 CPU, and achieves comparable accuracy compared with the state-of-the-art SLAM systems which heavily rely on both Intel Core i7 CPUs and powerful GPUs.
Xingyu Chen 0001, Jianru Xue, Jianwu Fang, Yuxin Pan, Nanning Zheng 0001
IV4
2020 Improving 3D Object Detection via Joint Attribute-oriented 3D Loss
abstract
3D object detection has become a hot topic in intelligent vehicle applications in recent years. Generally, deep learning has been the primary framework used in 3D object detection, and regression of the object location and classification of the objectness are the two indispensable components. In the process of training, the ℓn(n=1,2) and the focal loss are considered as the frequent solutions to minimize the regression and classification loss, respectively. However, there are two problems to be solved in the existing methods. For regression component, there is a gap between evaluation metrics, e.g., 3D Intersection over Union (IoU), and the traditional regression loss. As for the classification component, confidence score exists ambiguous due to the binary label assignment of target. To solve these problems, we propose a loss by jointing 3D IoU and other geometric attributes (named as jointed attribute-oriented 3D loss), which can be directly used in optimizing the regression component. In addition, the jointed attribute-oriented 3D loss can assign a soft label for supervising the training of the classification. By incorporating the proposed loss function into several state-of-the-art 3D object detection methods, the significant performance improvement has been achieved on the KITTI benchmark.
Jianru Xue, Jian Dou, Yuxin Pan, Jianwu Fang, Di Wang 0028, Nanning Zheng 0001
IV4
2006 On The Performance of Directional MAC Protocols in Wireless Ad-Hoc Networks
Yuxin Pan, Walaa Hamouda, Ahmed K. Elhakeem
AICCSA1
2005 A two-channel medium access control protocol for mobile ad hoc networks using directional antennas
abstract
In recent years, the use of directional antennas in wireless networks has been widely studied. Since the medium access control (MAC) protocol of the IEEE 802.11 standard is designed for the use of omnidirectional antennas, it cannot perform efficiently when directional antennas are used. In this paper, we study the performance of an efficient two-channel MAC protocol for ad hoc networks when equipped with directional antennas. The proposed protocol utilizes the large throughput offered by directional antennas using two frequency division multiplexed channels. The first channel is used for control information and the second for user data transmission. Based on this, the proposed MAC protocol operates in two main modes, the omnidirectional mode where one antenna is used for the transmission of users' control frames, and the directional mode where antenna arrays are used for the transmission of data frames. The proposed protocol is assessed using computer simulations based on randomly generated network topologies reflecting the random movement of nodes in the network. Based on these random topologies, we present performance comparisons with the existing MAC protocols using different system parameters. In all cases, the proposed MAC protocol is shown to offer a significant throughput improvement relative to the existing protocols
Yuxin Pan, Walaa Hamouda, Ahmed K. Elhakeem
PIMRC1