Xingyu Chen 0001

dblp:59/7651-1 · DBLP profile ↗
← Back
28ranked-venue papers
7as first author
21since 2021 · last 2026
0000-0002-5226-963XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 4 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 4 since 2021Systems, architecture and hardware · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 High-performance multi-agent path finding in high-obstacle-density and large-size maps
Shiguang Sun, Chang Tang, Shi-tao Chen, Zeyang Liu 0001, Xingyu Chen 0001, Xuguang Lan
Neurocomputing5
2026 MInCo: Mitigating conflicting objectives in distracted visual model-based reinforcement learning
Shiguang Sun, Hanbo Zhang, Zeyang Liu 0001, Lipeng Wan 0003, Xingyu Chen 0001, Xuguang Lan
Knowl. Based Syst.6
2026 Enhancing Value Decomposition With Target Transformation in Cooperative Multi-Agent Reinforcement Learning
abstract
The increasing need for cooperation among intelligent machines has heightened the importance of cooperative multi-agent reinforcement learning (MARL). However, a dominant class of cooperative MARL approaches relies on monotonic value decomposition, which enables scalable decentralized execution but restricts the representable class of joint action-values. However, existing remedies bias learning targets toward high-value samples, which can be fragile under stochastic returns because optimistic emphasis may amplify lucky but suboptimal trajectories. To solve this challenge, we propose Target Transformation, which maps non-monotonic and stochastic learning targets into a monotonic-representable surrogate while preserving the optimal joint action. Building on this idea, we develop Uncertainty-aware Target Transformation (UT2) with value-based and policy-based instantiations that combine an uncertainty estimator with a best-individual coordination envelope. Experiments on diverse cooperative MARL benchmarks show that UT2 improves both performance and stability over strong baselines, with larger gains as non-monotonicity and stochasticity increase.
Zeyang Liu 0001, Lipeng Wan 0003, Shiguang Sun, Xue Sui, Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Perception-Failure-Induced Test Scenario Searching via Online Causal Reinforcement Learning
abstract
Despite compliance with safety standards such as ISO 26262, the perception systems of autonomous vehicles still face numerous challenges during real-world operation. In this work, we propose a framework to identify safety-critical scenarios specifically induced by perception failures, aiming to support more targeted and effective scenario-based testing. Importantly, the key challenge lies in disentangling whether perception failures directly lead to critical outcomes. To address this, we propose an Online Causal Reinforcement Searching (OCRS) framework that simultaneously performs scenario search and causal reasoning. Focusing on visual unawareness and perception degradation, OCRS employs an LSTM-RNN controller to identify accident-prone scenarios linked to perception failures, which are then verified in a counterfactual world to determine causal responsibility. To improve efficiency, an online scenario classifier with passive-aggressive updates is introduced to dynamically filter out non-critical cases. The experimental results demonstrate both the effectiveness and efficiency of the proposed approach, achieving a 33% reduction in execution time for searchingperception-failure-induced corner cases. Furthermore, we evaluate the method under adverse weather conditions such as snow and fog, confirming that OCRS remains effective in identifying various failure modes.
Chi Zhang 0020, Tingting Long, Linhai Xu, Mingwen Bi, Xingyu Chen 0001, Yuehu Liu, Li Li 0013
IEEE Trans. Intell. Transp. Syst.6
2025 Offline Multi-Agent Preference-based Reinforcement Learning with Agent-aware Direct Preference Optimization
Qian Kou, Zeyang Liu 0001, Zhuoran Chen, Lipeng Wan 0003, Xingyu Chen 0001, Xuguang Lan
AAMAS7
2025 State Revisit and Re-explore: Bridging Sim-to-Real Gaps in Offline-and-Online Reinforcement Learning with An Imperfect Simulator
abstract
In reinforcement learning (RL) based robot skill acquisition, a high-fidelity simulator is usually indispensable but unattainable since the real environment dynamics are difficult to model, which leads to severe sim-to-real gaps. Existing methods solve this problem by combining offline and online RL to jointly learn transferable policies from limited offline data and imperfect simulators. However, due to the unrestricted exploration in the imperfect simulator, the hybrid offline-and-online RL methods inevitably suffer from low sample efficiency and insufficient state-action space coverage during training. To solve this problem, we propose a State Revisit and Re-exploration (SR2) hybrid offline-and-online RL framework. In particular, the proposed algorithm employs a meta-policy and a sub-policy, where the meta-policy aims to find high-quality states in the offline trajectories for online exploration, and the sub-policy learns the robot skill using mixed offline and online data. By introducing the state revisit and explore mechanism, our approach efficiently improves performance on a set of sim-to-real robotic tasks. Through extensive simulation and real-world tasks, we demonstrate the superior performance of our approach against other state-of-the-art methods.
Xingyu Chen 0001, Jiayi Xie, Ruixun Liu, Zeyang Liu 0001, Lipeng Wan 0003, Xuguang Lan
IJCAI1
2025 Flight Mastery in Turbulent Skies: Shared Control and Curriculum Reinforcement Learning for Crosswind Landing
abstract
Landing in crosswind conditions poses significant challenges for aircraft, as traditional control methods often fail to ensure stability in rapidly changing wind environments. While reinforcement learning offers a promising alternative, it typically suffers from low sample efficiency and limited generalization under stochastic wind fields. To address these challenges, we propose a Shared Control and Curriculum Reinforcement Learning framework. We model the crosswind landing task as a Markov Decision Process (MDP), explicitly defining the state space, action space, and wind field representation. To initialize learning, we decompose the multi-objective landing task into four sub-tasks—altitude, attitude, heading, and speed control—and train expert policies for each. These are then distilled into a shared control model via behavior cloning, providing a pre-trained policy with basic flight control capabilities. We further fine-tune this model using curriculum reinforcement learning, progressively increasing the complexity of wind conditions to enhance robustness and generalization. Experimental results across multiple aircraft and wind scenarios show that our method improves landing success rates and trajectory smoothness, while generalizing more effectively to unseen wind conditions, outperforming PID controllers, imitation learning, and mainstream RL baselines.
Zechen Shi, Xingyu Chen 0001, Zeyang Liu 0001, Chi Zhang 0020, Yimeng Yu, Junbin You, Xuguang Lan
IEEE Trans. Intell. Transp. Syst.2
2025 Improving Offline Reinforcement Learning With in-Sample Advantage Regularization for Robot Manipulation
abstract
Offline reinforcement learning (RL) aims to learn the possible policy from a fixed dataset without real-time interactions with the environment. By avoiding the risky exploration of the robot, this approach is expected to significantly improve the robot's learning efficiency and safety. However, due to errors in value estimation from out-of-distribution actions, most offline RL algorithms constrain or regularize the policy to the actions contained within the dataset. The cost of such methods is the introduction of new hyperparameters and additional complexity. In this article, we aim to adapt offline RL to robotic manipulation with minimal changes and to avoid evaluating out-of-distribution actions as much as possible. Therefore, we improve offline RL with in-sample advantage regularization (ISAR). To mitigate the impact of unseen actions, the ISAR learns the state-value function only with the dataset sample to regress the optimal action-value function. Our method calculates the advantage function of action-state pairs based on in-sample value estimation and adds a behavior cloning (BC) regularization term in the policy update. This improves sample efficiency with minimal changes, resulting in a simple and easy-to-implement method. The experiments of the D4RL robot benchmark and multigoal sparse rewards robotic tasks show that the ISAR achieves excellent performance comparable to current state-of-the-art algorithms without the need for complex parameter tuning and too much training time. In addition, we demonstrate the effectiveness of our method on a real-world robot platform.
Chengzhong Ma, Deyu Yang, Zeyang Liu 0001, Houxue Yang, Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001
IEEE Trans. Neural Networks Learn. Syst.6
2024 Imagine, Initialize, and Explore: An Effective Exploration Method in Multi-Agent Reinforcement Learning
abstract
Effective exploration is crucial to discovering optimal strategies for multi-agent reinforcement learning (MARL) in complex coordination tasks. Existing methods mainly utilize intrinsic rewards to enable committed exploration or use role-based learning for decomposing joint action spaces instead of directly conducting a collective search in the entire action-observation space. However, they often face challenges obtaining specific joint action sequences to reach successful states in long-horizon tasks. To address this limitation, we propose Imagine, Initialize, and Explore (IIE), a novel method that offers a promising solution for efficient multi-agent exploration in complex scenarios. IIE employs a transformer model to imagine how the agents reach a critical state that can influence each other's transition functions. Then, we initialize the environment at this state using a simulator before the exploration phase. We formulate the imagination as a sequence modeling problem, where the states, observations, prompts, actions, and rewards are predicted autoregressively. The prompt consists of timestep-to-go, return-to-go, influence value, and one-shot demonstration, specifying the desired state and trajectory as well as guiding the action generation. By initializing agents at the critical states, IIE significantly increases the likelihood of discovering potentially important under-explored regions. Despite its simplicity, empirical results demonstrate that our method outperforms multi-agent exploration baselines on the StarCraft Multi-Agent Challenge (SMAC) and SMACv2 environments. Particularly, IIE shows improved performance in the sparse-reward SMAC tasks and produces more effective curricula over the initialized states than other generative methods, such as CVAE-GAN and diffusion models.
Zeyang Liu 0001, Lipeng Wan 0003, Zhuoran Chen, Xingyu Chen 0001, Xuguang Lan
AAAI5
2024 Experience Consistency Distillation Continual Reinforcement Learning for Robotic Manipulation Tasks
abstract
Continual reinforcement learning, which aims to help robots acquire skills without catastrophic forgetting, obviating the need to re-learn all tasks from scratch. In order to enable lifelong acquisition of skills in robots, replay-based continual reinforcement learning has emerged as a promising research direction. These techniques replay data from previous tasks to mitigate forgetting when learning new skills. However, existing replay-based methods store poor representative experience, and the experience utilization of old tasks is inefficient. To address these issues, we propose an experience consistency distillation method for robot continual reinforcement learning to improve the data efficiency of the experience. Specifically, the experience of old tasks are distilled to obtain Markov Decision Process (MDP) data with high compression ratio and information content. To ensure consistent data distributions before and after distillation, we further utilize a Fréchet Inception Distance (FID) loss as a regularization constraint. In order to improve experience utilization efficiency, the policy is then trained using both the distilled data and current task data, with policy distillation performed based on uncertainty metrics. Our method is validated in the continual reinforcement learning simulation platform and real scene with a UR5e robot arm. Experimental results indicate that our method achieves higher success and lower buffer size requirement compared to other methods.
Ru Peng, Xingyu Chen 0001, Xuguang Lan
ICRA4
2024 Grounded Answers for Multi-agent Decision-making Problem through Generative World Model
abstract
Recent progress in generative models has stimulated significant innovations in many fields, such as image generation and chatbots. Despite their success, these models often produce sketchy and misleading solutions for complex multi-agent decision-making problems because they miss the trial-and-error experience and reasoning as humans. To address this limitation, we explore a paradigm that integrates a language-guided simulator into the multi-agent reinforcement learning pipeline to enhance the generated answer. The simulator is a world model that separately learns dynamics and reward, where the dynamics model comprises an image tokenizer as well as a causal transformer to generate interaction transitions autoregressively, and the reward model is a bidirectional transformer learned by maximizing the likelihood of trajectories in the expert demonstrations under language guidance. Given an image of the current state and the task description, we use the world model to train the joint policy and produce the image sequence as the answer by running the converged policy on the dynamics model. The empirical results demonstrate that this framework can improve the answers for multi-agent decision-making problems by showing superior performance on the training and unseen tasks of the StarCraft Multi-Agent Challenge benchmark. In particular, it can generate consistent interaction sequences and explainable reward functions at interaction states, opening the path for training generative models of the future.
Zeyang Liu 0001, Shiguang Sun, Lipeng Wan 0003, Xingyu Chen 0001, Xuguang Lan
NeurIPS6
2024 A causality guided loss for imbalanced learning in scene graph generation
Ru Peng, Xingyu Chen 0001, Ziru Wang, Xuguang Lan
Neurocomputing3
2024 Optimal bipartite graph matching-based goal selection for policy-based hindsight learning
Shiguang Sun, Hanbo Zhang, Zeyang Liu 0001, Xingyu Chen 0001, Xuguang Lan
Neurocomputing4
2024 Knowledge Graph Enhancement for Fine-Grained Zero-Shot Learning on ImageNet21K
abstract
Fine-grained Zero-shot Learning on the large-scale dataset ImageNet21K is an important task that has promising perspectives in many real-world scenarios. One typical solution is to explicitly model the knowledge passing using a Knowledge Graph (KG) to transfer knowledge from seen to unseen instances. By analyzing the hierarchical structure and the word descriptions on ImageNet21K, we find that the noisy semantic information, the sparseness of seen classes, and the lack of supervision of unseen classes make the knowledge passing insufficient, which limits the KG-based fine-grained ZSL. To resolve this problem, in this paper, we enhance the knowledge passing from three aspects. First, we use more powerful models such as the Large Language Model and Vision-Language Model to get more reliable semantic embeddings. Then we propose a strategy that globally enhances the knowledge graph based on the convex combination relationship of the semantic embeddings. It effectively connects the edges between the non-kinship seen and unseen classes that have strong correlations while assigning an importance score to each edge. Based on the enhanced knowledge graph, we further present a novel regularizer that locally enhances the knowledge passing during training. We extensively conducted comparative evaluations to demonstrate the advantages of our method over state-of-the-art approaches.
Xingyu Chen 0001, Zeyang Liu 0001, Lipeng Wan 0003, Xuguang Lan, Nanning Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 MMRDN: Consistent Representation for Multi-View Manipulation Relationship Detection in Object-Stacked Scenes
abstract
Manipulation relationship detection (MRD) aims to guide the robot to grasp objects in the right order, which is important to ensure the safety and reliability of grasping in object stacked scenes. Previous works infer manipulation relationship by deep neural network trained with data collected from a predefined view, which has limitation in visual dislocation in unstructured environments. Multi-view data provide more comprehensive information in space, while a challenge of multi-view MRD is domain shift. In this paper, we propose a novel multi-view fusion framework, namely multi-view MRD network (MMRDN), which is trained by 2D and 3D multi-view data. We project the 2D data from different views into a common hidden space and fit the embeddings with a set of Von-Mises-Fisher distributions to learn the consistent representations. Besides, taking advantage of position information within the 3D data, we select a set of$K$Maximum Vertical Neighbors (KMVN) points from the point cloud of each object pair, which encodes the relative position of these two objects. Finally, the features of multi-view 2D and 3D data are concatenated to predict the pairwise relationship of objects. Experimental results on the challenging REGRAD dataset show that MMRDN outperforms the state-of-the-art methods in multi-view MRD tasks. The results also demonstrate that our model trained by synthetic data is capable to transfer to real-world scenarios.
Lipeng Wan 0003, Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001
ICRA4
2023 Prioritized Planning for Target-Oriented Manipulation via Hierarchical Stacking Relationship Prediction
abstract
In scenarios involving grasping multiple targets, the learning of stacking relationships between objects is fundamental for robots to execute safely and efficiently. However, current methods lack subdivision for the hierarchy of stacking relationship types. In scenes where objects are mostly stacked in an orderly manner, they are incapable of performing human-like and high-efficient grasping decisions. This paper proposes a perception-planning method to distinguish different stacking forms between objects and generate prioritized manipulation sequences based on given target designations. We utilize a Hierarchical Stacking Relationship Network (HSRN) to discriminate the hierarchy of stacking and generate a refined Stacking Relationship Tree (SRT) for relationship description. Considering objects with high stacking stability can be processed together if necessary, we introduce an elaborate decision-making planner based on Partially Observable Markov Decision Process (POMDP), which leverages observations and generates the least grasp-consuming decision chain with robustness and is suitable for simultaneously specifying multiple targets. To verify our work, we set the scene to the dining table and augment REGRAD dataset for network training. Experiments show that our method effectively generates grasping decisions that conform to human requirements, and improves the implementation efficiency compared with existing methods on the basis of guaranteeing success rate.
Zewen Wu, Xingyu Chen 0001, Chengzhong Ma, Xuguang Lan, Nanning Zheng 0001
IROS3
2022 Greedy based Value Representation for Optimal Coordination in Multi-agent Reinforcement Learning
abstract
Due to the representation limitation of the joint Q value function, multi-agent reinforcement learning methods with linear value decomposition (LVD) or monotonic value decomposition (MVD) suffer from relative overgeneralization. As a result, they can not ensure optimal consistency (i.e., the correspondence between individual greedy actions and the best team performance). In this paper, we derive the expression of the joint Q value function of LVD and MVD. According to the expression, we draw a transition diagram, where each self-transition node (STN) is a possible convergence. To ensure the optimal consistency, the optimal node is required to be the unique STN. Therefore, we propose the greedy-based value representation (GVR), which turns the optimal node into an STN via inferior target shaping and eliminates the non-optimal STNs via superior experience replay. Theoretical proofs and empirical results demonstrate that given the true Q values, GVR ensures the optimal consistency under sufficient exploration. Besides, in tasks where the true Q values are unavailable, GVR achieves an adaptive trade-off between optimality and stability. Our method outperforms state-of-the-art baselines in experiments on various benchmarks.
Lipeng Wan 0003, Zeyang Liu 0001, Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001
ICML3
2022 A Continuous Learning Approach for Probabilistic Human Motion Prediction
abstract
Human Motion Prediction (HMP) plays a crucial role in safe Human-Robot-Interaction (HRI). Currently, the majority of HMP algorithms are trained by massive pre-collected data. As the training data only contains a few pre-defined motion patterns, these methods cannot handle the unfamiliar motion patterns. Moreover, the pre-collected data are usually non-interactive, which does not consider the real-time responses of collaborators. As a result, these methods usually perform unsatisfactorily in real HRI scenarios. To solve this problem, in this paper, we propose a novel Continual Learning (CL) approach for probabilistic HMP which makes the robot continually learns during its interaction with collaborators. The proposed approach consists of two steps. First, we leverage a Bayesian Neural Network to model diverse uncertainties of observed human motions for collecting online interactive data safely. Then we take Experience Replay and Knowledge Distillation to elevate the model with new experiences while maintaining the knowledge learned before. We first evaluate our approach on a large-scale benchmark dataset Human3.6m. The experimental results show that our approach achieves a lower prediction error compared with the baselines methods. Moreover, our approach could continually learn new motion patterns without forgetting the learned knowledge. We further conduct real-scene experiments using Kinect DK. The results show that our approach can learn the human kinematic model from scratch, which effectively secures the interaction.
Shihong Wang, Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001
ICRA3
2022 Generalized Zero-Shot Learning Via Multi-Modal Aggregated Posterior Aligning Neural Network
abstract
The visual-semantic gap between the visual space (visual features) and semantic space (semantic attributes) is one of the main problems in the Generalized Zero-Shot Learning (GZSL) task. The essence of this problem is that the structure of manifolds in these two spaces is inconsistent, which makes it difficult to learn embeddings that unify visual features and semantic attributes for similarity measurement. In this work, we tackle this problem by proposing a multi-modal aggregated posterior aligning neural network based on Wasserstein Auto-encoders (WAE) which learns a shared latent space for visual features and semantic attributes. The key to our approach is that the aggregated posterior distribution of the latent representations encoded from visual features of each class is encouraged to be aligned with a Gaussian distribution predicted by the corresponding semantic attribute in the latent space. On one hand, requiring the latent manifolds of visual features and semantic attributes to be consistent preserves the inter-class association between seen and unseen classes. On the other hand, the aggregated posterior of each class is directly defined as a Gaussian in the latent space, which provides a reliable way to synthesize latent features for training classification models. Using the AWA1, AWA2, CUB, aPY, FLO, and SUN benchmark datasets, we extensively conducted comparative evaluations to demonstrate the advantages of our method over state-of-the-art approaches.
Xingyu Chen 0001, Jin Li 0011, Xuguang Lan, Nanning Zheng 0001
IEEE Trans. Multim.1
2022 Neighborhood Geometric Structure-Preserving Variational Autoencoder for Smooth and Bounded Data Sources
abstract
Many data sources, such as human poses, lie on low-dimensional manifolds that are smooth and bounded. Learning low-dimensional representations for such data is an important problem. One typical solution is to utilize encoder-decoder networks. However, due to the lack of effective regularization in latent space, the learned representations usually do not preserve the essential data relations. For example, adjacent video frames in a sequence may be encoded into very different zones across the latent space with holes in between. This is problematic for many tasks such as denoising because slightly perturbed data have the risk of being encoded into very different latent variables, leaving output unpredictable. To resolve this problem, we first propose a neighborhood geometric structure-preserving variational autoencoder (SP-VAE), which not only maximizes the evidence lower bound but also encourages latent variables to preserve their structures as in ambient space. Then, we learn a set of small surfaces to approximately bound the learned manifold to deal with holes in latent space. We extensively validate the properties of our approach by reconstruction, denoising, and random image generation experiments on a number of data sources, including synthetic Swiss roll, human pose sequences, and facial expression images. The experimental results show that our approach learns more smooth manifolds than the baselines. We also apply our approach to the tasks of human pose refinement and facial expression image interpolation where it gets better results than the baselines.
Xingyu Chen 0001, Chunyu Wang 0001, Xuguang Lan, Nanning Zheng 0001, Wenjun Zeng 0001
IEEE Trans. Neural Networks Learn. Syst.1
2021 Probabilistic Human Motion Prediction via A Bayesian Neural Network
abstract
Human motion prediction is an important and challenging topic that has promising prospects in efficient and safe human-robot-interaction systems. Currently, the majority of the human motion prediction algorithms are based on deterministic models, which may lead to risky decisions for robots. To solve this problem, we propose a probabilistic model for human motion prediction in this paper. The key idea of our approach is to extend the conventional deterministic motion prediction neural network to a Bayesian one. On one hand, our model could generate several future motions when given an observed motion sequence. On the other hand, by calculating the Epistemic Uncertainty and the Heteroscedastic Aleatoric Uncertainty, our model could tell the robot if the observation has been seen before and also give the optimal result among all possible predictions. We extensively validate our approach on a large scale benchmark dataset Human3.6m. The experiments show that our approach performs better than deterministic methods. We further evaluate our approach in a Human-Robot-Interaction (HRI) scenario. The experimental results show that our approach makes the interaction more efficient and safer.
Xingyu Chen 0001, Xuguang Lan, Nanning Zheng 0001
ICRA2
2020 A Boundary Based Out-of-Distribution Classifier for Generalized Zero-Shot Learning
Xingyu Chen 0001, Xuguang Lan, Fuchun Sun 0001, Nanning Zheng 0001
ECCV (24)1
2020 Navigation Command Matching for Vision-based Autonomous Driving
abstract
Learning an optimal policy for autonomous driving task to confront with complex environment is a long- studied challenge. Imitative reinforcement learning is accepted as a promising approach to learn a robust driving policy through expert demonstrations and interactions with environments. However, this model utilizes non-smooth rewards, which have a negative impact on matching between navigation commands and trajectory (state-action pairs), and degrade the generalizability of an agent. Smooth rewards are crucial to discriminate actions generated from sub-optimal policy. In this paper, we propose a navigation command matching (NCM) model to address this issue. There are two key components in NCM, 1) a matching measurer produces smooth navigation rewards that measure matching between navigation commands and trajectory; 2) attention-guided agent performs actions given states where salient regions in RGB images (i.e. roadsides, lane markings and dynamic obstacles) are highlighted to amplify their influence on the final model. We obtain navigation rewards and store transitions to replay buffer after an episode, so NCM is able to discriminate actions generated from suboptimal policy. Experiments on CARLA driving benchmark show our proposed NCM outperforms previous state-of-the- art models on various tasks in terms of the percentage of successfully completed episodes. Moreover, our model improves generalizability of the agent and obtains good performance even in unseen scenarios.
Yuxin Pan, Jianru Xue, Pengfei Zhang 0005, Wanli Ouyang, Jianwu Fang, Xingyu Chen 0001
ICRA6
2020 Using Detection, Tracking and Prediction in Visual SLAM to Achieve Real-time Semantic Mapping of Dynamic Scenarios
abstract
In this paper, we propose a lightweight system, RDS-SLAM, based on ORB-SLAM2, which can accurately estimate poses and build semantic maps at object level for dynamic scenarios in real time using only one commonly used Intel Core i7 CPU. In RDS-SLAM, three major improvements, as well as major architectural modifications, are proposed to overcome the limitations of ORB-SLAM2. Firstly, it adopts a lightweight object detection neural network in key frames. Secondly, an efficient tracking and prediction mechanism is embedded into the system to remove the feature points belonging to movable objects in all incoming frames. Thirdly, a semantic octree map is built by probabilistic fusion of detection and tracking results, which enables a robot to maintain a semantic description at object level for potential interactions in dynamic scenarios. We evaluate RDS-SLAM in TUM RGB-D dataset, and experimental results show that RDS-SLAM can run with 30.3 ms per frame in dynamic scenarios using only an Intel Core i7 CPU, and achieves comparable accuracy compared with the state-of-the-art SLAM systems which heavily rely on both Intel Core i7 CPUs and powerful GPUs.
Xingyu Chen 0001, Jianru Xue, Jianwu Fang, Yuxin Pan, Nanning Zheng 0001
IV1
2020 A Joint Label Space for Generalized Zero-Shot Classification
abstract
The fundamental problem of Zero-Shot Learning (ZSL) is that the one-hot label space is discrete, which leads to a complete loss of the relationships between seen and unseen classes. Conventional approaches rely on using semantic auxiliary information, e.g. attributes, to re-encode each class so as to preserve the inter-class associations. However, existing learning algorithms only focus on unifying visual and semantic spaces without jointly considering the label space. More importantly, because the final classification is conducted in the label space through a compatibility function, the gap between attribute and label spaces leads to significant performance degradation. Therefore, this paper proposes a novel pathway that uses the label space to jointly reconcile visual and semantic spaces directly, which is named Attributing Label Space (ALS). In the training phase, one-hot labels of seen classes are directly used as prototypes in a common space, where both images and attributes are mapped. Since mappings can be optimized independently, the computational complexity is extremely low. In addition, the correlation between semantic attributes has less influence on visual embedding training because features are mapped into labels instead of attributes. In the testing phase, the discrete condition of label space is removed, and priori one-hot labels are used to denote seen classes and further compose labels of unseen classes. Therefore, the label space is very discriminative for the Generalized ZSL (GZSL), which is more reasonable and challenging for real-world applications. Extensive experiments on five benchmarks manifest improved performance over all of compared state-of-the-art methods.
Jin Li 0011, Xuguang Lan, Yang Long 0001, Yang Liu 0069, Xingyu Chen 0001, Ling Shao 0001, Nanning Zheng 0001
IEEE Trans. Image Process.5
2019 Learning to Refine 3D Human Pose Sequences
abstract
We present a basis approach to refine noisy 3D human pose sequences by jointly projecting them onto a non-linear pose manifold, which is represented by a number of basis dictionaries with each covering a small manifold region. We learn the dictionaries by jointly minimizing the distance between the original poses and their projections on the dictionaries, along with the temporal jittering of the projected poses. During testing, given a sequence of noisy poses which are probably off the manifold, we project them to the manifold using the same strategy as in training for refinement. We apply our approach to the monocular 3D pose estimation and the long term motion prediction tasks. The experimental results on the benchmark dataset shows the estimated 3D poses are notably improved in both tasks. In particular, the smoothness constraint helps generate more robust refinement results even when some poses in the original sequence have large errors.
Jieru Mei, Xingyu Chen 0001, Chunyu Wang 0001, Alan L. Yuille, Xuguang Lan, Wenjun Zeng 0001
3DV2
2019 Cross-View Person Identification Based on Confidence-Weighted Human Pose Matching
abstract
Cross-view person identification (CVPI) from multiple temporally synchronized videos taken by multiple wearable cameras from different, varying views is a very challenging but important problem, which has attracted more interest recently. Current state-of-the-art performance of CVPI is achieved by matching appearance and motion features across videos, while the matching of pose features does not work effectively given the high inaccuracy of the 3D pose estimation on videos/images collected in the wild. To address this problem, we first introduce a new metric of confidence to the estimated location of each human-body joint in 3D human pose estimation. Then, a mapping function, which can be hand-crafted or learned directly from the datasets, is proposed to combine the inaccurately estimated human pose and the inferred confidence metric to accomplish CVPI. Specifically, the joints with higher confidence are weighted more in the pose matching for CVPI. Finally, the estimated pose information is integrated into the appearance and motion features to boost the CVPI performance. In the experiments, we evaluate the proposed method on three wearable-camera video datasets and compare the performance against several other existing CVPI methods. The experimental results show the effectiveness of the proposed confidence metric, and the integration of pose, appearance, and motion produces a new state-of-the-art CVPI performance.
Guoqiang Liang 0001, Xuguang Lan, Xingyu Chen 0001, Song Wang 0002, Nanning Zheng 0001
IEEE Trans. Image Process.3
2017 Pose-and-illumination-invariant face representation via a triplet-loss trained deep reconstruction model
Xingyu Chen 0001, Xuguang Lan, Guoqiang Liang 0001, Nanning Zheng 0001
Multim. Tools Appl.1