Yong Song 0005

dblp:33/4122-5 · DBLP profile ↗
← Back
22ranked-venue papers
0as first author
19since 2021 · last 2026
0000-0003-2505-2766ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Large language model assisted hierarchical reinforcement learning training
Qianxi Li, Bao Pang, Yong Song 0005, Hongze Fu, Qingyang Xu, Xianfeng Yuan, Xiaolong Xu 0003, Chengjin Zhang
Inf. Sci.3
2026 PCK-RL: A Unified Framework for Vision-Language Model and Reinforcement Learning-Based Robotic Manipulation
abstract
Recent advances in the vision-language model (VLM) have demonstrated impressive capabilities in aligning high-level semantic information with visual environments, offering significant help for task planning in robotic manipulation. However, existing methods often fail to convert semantic tasks into precise robotic actions, especially in complex 3-D environments where keypoint localization and interaction dynamics are critical. To address this challenge, this article proposes principal component keypoint-reinforcement learning (PCK-RL), a multistage framework that bridges semantic reasoning and low-level control through semantic subplan decomposition, object keypoint localization, and example trajectory generation. Principal component analysis (PCA) is introduced to normalize the inputs and outputs of the VLM, constructing novel intermediate representations to enhance model performance. Additionally, a hybrid robotic action alignment method rapidly exploring random tree star-reinforcement learning (RRT*-RL) is introduced. It refines zero-shot reference trajectories by incorporating reinforcement learning (RL) with a closed-loop environmental feedback mechanism, which is crucial for long-horizon tasks. This article validates PCK-RL through extensive experiments in both simulated and real-world robotic environments, demonstrating its effectiveness and robustness in executing long-horizon tasks and manipulating diverse objects.
Qianxi Li, Hongze Fu, Yong Song 0005, Chengjin Zhang, Bao Pang
IEEE Trans. Ind. Informatics4
2026 WMTP: A Wavelet-Mamba Trajectory Predictor for Autonomous Driving
abstract
Vehicle trajectory is crucial for autonomous driving. Relatively scattered trajectory data points pose difficulties in modeling the motion’s inherent continuity in spatial and temporal dimensions. Additionally, identifying the driving patterns of vehicles from trajectories is also a significant challenge. These implicit characteristic patterns are difficult to discern from the complex details of the trajectory data. To address these issues, we propose a new framework called Wavelet-Mamba Trajectory Prediction (WMTP), which fuses wavelet analysis through state-space modeling to capture global trends in driving patterns and details of vehicle motion. The approach employs the Discrete Wavelet Transform (DWT) to decompose trajectory data into wavelet coefficients in different time scales and frequencies, and then utilizes these coefficients to generate the trajectory through the Inverse Discrete Wavelet Transform (IDWT). An encoder-decoder neural architecture is proposed for learning potential temporal features from the input trajectory sequences, and these features are projected into the wavelet domain. Wavelet coefficients of future trajectories are generated using different scale-oriented decoders. The estimated coefficients are further used to realize the trajectory prediction via the IDWT module. Experiments demonstrate that WMTP exhibits excellent performance on three large-scale real-world trajectory prediction datasets, with promising robustness and inference speed. The research findings also verify the effectiveness of time - frequency analysis in trajectory prediction tasks.
Zhiyang Yin, Qingyang Xu, Yong Song 0005, Bao Pang, Yibin Li 0001, Ning Wang 0002
ACM Trans. Internet Things3
2025 Complex Robotic Manipulation via Hindsight Goal Diffusion and Graph-based Experience Replay
abstract
Goal-conditioned reinforcement learning (GCRL) is an effective method for multi-goal robotic manipulation tasks. Many studies based on hindsight experience replay (HER) and hindsight goal generation (HGG) have achieved the autonomous acquisition of robotic manipulation in reward-sparse environments and have greatly improved the learning efficiency of GCRL. However, these methods perform poorly in environments with obstacles and distant goals. In this paper, we propose hindsight goal diffusion and graph-based experience replay (HGD-GER) for complex robotic manipulation. First, obstacle-avoiding graphs in environments with obstacles are constructed, and the graph-based distance metric between different goals is established. Second, the proposed HGD approach utilizes the inherent denoising mechanism of diffusion models and obstacle-avoiding graph-based distance to generate exploration goals, thereby promoting the exploration of obstacle-bypassing areas. Then, GER module modifies the reward value of experience replay by graph-based distance, thereby avoiding the bias introduced by HER and improving the learning performance of the RL algorithm under sparse reward conditions. Finally, we conducted experiments on three robotic manipulation tasks with obstacles and distant goals, and the results show that the proposed HGD-GER achieves excellent learning performance. Additionally, the proposed method is deployed on the physical robot.
Jinrui He, Yong Song 0005, Pingping Liu, Qingyang Xu, Xianfeng Yuan, Rui Song 0002
IROS4
2025 Deep learning-based visual slam for indoor dynamic scenes
Zhendong Xu, Yong Song 0005, Bao Pang, Qingyang Xu, Xianfeng Yuan
Appl. Intell.2
2025 Correction to: Deep learning-based visual slam for indoor dynamic scenes
Zhendong Xu, Yong Song 0005, Bao Pang, Qingyang Xu, Xianfeng Yuan
Appl. Intell.2
2025 Hierarchical reinforcement learning with curriculum demonstrations and goal-guided policies for sequential robotic manipulation
Bao Pang, Xianfeng Yuan, Xiaolong Xu 0003, Yong Song 0005, Rui Song 0002, Yibin Li 0001
Eng. Appl. Artif. Intell.5
2025 An adaptive reinforcement learning approach with trait-awareness for heterogeneous multi-robot cooperative pursuit
Heteng Zhang, Yunjie Jia, Yong Song 0005, Bao Pang, Xianfeng Yuan, Rui Song 0002, Simon X. Yang
Eng. Appl. Artif. Intell.3
2025 Toward Multimodal Graph Sequence Generation: A Denoising Diffusion Approach for Wheeled Robot Fault Diagnosis
abstract
Wheeled robots play a crucial role in Industrial Internet of Things (IIoT)-enabled manufacturing environments, and ensuring their reliable operation is essential for production efficiency and safety. However, their inherent complexity makes them prone to faults, while limited fault data results in imbalanced datasets, posing challenges for deep-learning-based fault diagnosis models. Existing denoising diffusion probabilistic model (DDPM)-based fault diagnosis methods tend to address single-channel scenarios, ignoring the graph relationships inherent in multichannel sensor data. Furthermore, current graph-based DDPM models also struggle in wheeled robot scenarios due to its multimodal nature. To address these challenges, we propose an enhanced DDPM-based method for imbalanced fault diagnosis of wheeled robots. Our method integrates graph operations into the noise prediction network of the DDPM framework, enabling efficient modeling of the complex spatial–temporal relations in multimodal graphical sequence data via a newly designed spatial–temporal graph U-Net (STGU-Net). Additionally, we present a dynamic degradation mechanism for the prior graph, simulating gradual structural changes during the diffusion process. Extensive experiments on a real-world wheeled robot platform demonstrate the superiority of the proposed model over state-of-the-art methods from multiple perspectives, showcasing its effectiveness in generating high-quality data and mitigating the imbalance problem in wheeled robot fault diagnosis.
Tianyi Ye, Haolin Cao, Jianjie Liu, Bao Pang, Qingyang Xu, Yong Song 0005, Xianfeng Yuan
IEEE Internet Things J.6
2025 A Deep Reinforcement Learning Approach Using Asymmetric Self-Play for Robust Multirobot Flocking
abstract
Flocking control, as an essential approach for survivable navigation of multirobot systems, has been widely applied in fields, such as logistics, service delivery, and search and rescue. However, realistic environments are typically complex, dynamic, and even aggressive, posing considerable threats to the safety of flocking robots. In this article, based on deep reinforcement learning, anAsymmetricSelf-play-empoweredFlockingControl framework is proposed to address this concern. Specifically, the flocking robots are trained concurrently with learnable adversarial interferers to stimulate the intelligence of the flocking strategy. A two-stage self-play training paradigm is developed to improve the robustness and generalization of the model. Furthermore, an auxiliary training module regarding the learning of transition dynamics is designed, dramatically enhancing the adaptability to environmental uncertainties. Feature-level and agent-level attention are implemented for action and value generation, respectively. Both extensive comparative experiments and real-world deployment demonstrate the superiority and practicality of the proposed framework.
Yunjie Jia, Yong Song 0005, Jiyu Cheng, Jiong Jin, Wei Zhang 0021, Simon X. Yang, Sam Kwong
IEEE Trans. Ind. Informatics2
2025 Empowering Multirobot Flocking in Complex Environments via Effective Communication: A Deep Reinforcement Learning Approach
abstract
Multirobot flocking is crucial for safe and cooperative navigation, with wide applications in logistics, service delivery, and mobile surveillance. Despite significant progress, developing effective flocking strategies under complex conditions remains challenging. Communication is a vital technique for multirobot coordination. In this article, we propose refinement and enhancement of communication information (REIN), a novel deep reinforcement learning-based framework designed to improve communication effectiveness in leader–follower flocking systems through the REIN. First, regarding information refinement, a graph-based information refiner, integrating directed graph-structured communication with an innovative edge filter, is developed for selective multirobot interaction. It helps robots adaptively focus on relevant neighbors, considerably alleviating information overload. Second, for information enhancement, a cognition-aligned information enhancer is designed that boosts information expressiveness by encouraging team consensus. It utilizes two cascaded leader-related objectives to optimize information towards cognitive alignment among decentralized followers. Extensive comparisons with state-of-the-art approaches and ablation versions demonstrate the superiority of our framework. Physical experiments are also conducted to validate its practicality.
Yunjie Jia, Yong Song 0005, Jiyu Cheng, Heteng Zhang, Wei Zhang 0021, Rui Song 0002, Simon X. Yang, Sam Kwong
IEEE Trans. Ind. Informatics2
2025 Goal-Conditioned Reinforcement Learning With Adaptive Intrinsic Curiosity and Universal Value Network Fitting for Robotic Manipulation
abstract
Hindsight experience replay (HER) has greatly increased the possibility of using deep reinforcement learning (DRL) for robotic manipulation with sparse rewards. However, there are still concerns about low learning efficiency and poor performance due to its insufficient exploration ability and bias against the initial goal introduced by HER. In this article, to solve this problem, a multigoal robotic manipulation DRL method based on adaptive intrinsic curiosity and universal value network fitting (AIC-UVNF) is proposed to further improve the exploration ability and learning performance. Specifically, this method utilizes an improved curiosity mechanism to construct a joint intrinsic reward and adaptively adjust the proportion, which can enhance exploration ability and avoid excessive pursuit of novel states. In addition, a universal value network fitting approach is proposed to incorporate the initial goal into the value function fitting process, which employs the value of the initial goal to eliminate the bias of HER in the algorithm update. Combined with the off-policy soft actor-critic method, AIC-UVNF is verified on multigoal robotic manipulation tasks. The results show that the proposed method achieves better convergence efficiency and learning performance.
Xianfeng Yuan, Qingyang Xu, Bao Pang, Yong Song 0005, Rui Song 0002, Yibin Li 0001
IEEE Trans. Ind. Informatics5
2024 Dual-Critic Deep Reinforcement Learning for Push-Grasping Synergy in Cluttered Environment
abstract
Robotic push-grasping in densely cluttered environments presents significant challenges due to unbalanced synergy and redundancy between both actions, leading to decreased grasp efficiency. In this paper, a novel double-critic deep reinforcement learning framework is introduced to optimize the push-grasping synergy for robotic manipulation in such environments, aiming to significantly reduce pre-grasping redundancy. This framework incorporates two distinct Deep Q-learning critics: Critic I selects the best course of actions based on the current state derived from visual interpretation, whereas Critic II evaluates the success rate of the current state-action pairing. To further refine the push-grasping synergy, an active double-step learning mechanism is introduced to optimize the training reward function for the pushing action, thereby enhancing its effectiveness through increased intentionality. Simulations show that the proposed framework outperforms contemporary counterparts, notably in grasping success rate and action efficiency. Finally, the framework’s generalization and adaptability are demonstrated by conducting real-world experiments using novel objects without the need of retraining.
Jiakang Zhong, Yew Wee Wong, Jiong Jin, Yong Song 0005, Xianfeng Yuan
ICRA4
2024 HOGN-TVGN: Human-inspired Embodied Object Goal Navigation based on Time-varying Knowledge Graph Inference Networks for Robots
Baojiang Yang, Xianfeng Yuan, Zhongmou Ying, Boyi Song, Yong Song 0005, Fengyu Zhou 0002, Weihua Sheng
Adv. Eng. Informatics6
2024 Accurate and Efficient 3D Panoptic Mapping Using Diverse Information Modalities and Multidimensional Data Association
abstract
3D Panoptic perception is essential for the understanding of real-world environment and plays an increasingly important role in the field of robotics. However, most existing methods heavily rely on image panoptic segmentation networks to acquire panoptic information of the environment, which is time-consuming and susceptible to interference. In this paper, we propose a novel and efficient panoptic mapping method based on multi-source information. Specifically, to improve the real-time performance of the system, we first apply lightweight object detection and semantic segmentation to extract 2D semantic and instance information from images. Second, a panoptic inference algorithm is designed that fully utilizes multi-source information, including geometry-based and learning-based information, to simultaneously reason about background and foreground objects in the environment. Finally, we take advantage of the scalability of the framework by introducing a multi-object tracking algorithm into the framework, thus providing the temporal information among consecutive frames to the data association module. Based on two popular datasets, extensive comparison experiments are conducted to illustrate the effectiveness of the proposed method. Experimental results show that compared with state-of-the-art panoptic mapping methods, the proposed method achieves superior performance in accuracy, real-timeness and stability. Furthermore, we also evaluate our method in real-world scenarios and CPU-only device to demonstrate the feasibility of its practical deployment.
Zhongmou Ying, Xianfeng Yuan, Boyi Song, Yong Song 0005, Fengyu Zhou 0002, Weihua Sheng
IEEE Trans. Circuits Syst. Video Technol.4
2024 Hierarchical Perception-Improving for Decentralized Multi-Robot Motion Planning in Complex Scenarios
abstract
Multi-robot cooperative navigation is an important task, which has been widely studied in many fields like logistics, transportation, and disaster rescue. However, most of the existing methods either require some strong assumptions or are validated in simple scenarios, which greatly hinders their implementation in the real world. In this paper, more complex environments are considered in which robots can only acquire local observations from their own sensors and have only limited communication capabilities for mapless collaborative navigation. To address this challenging task, we propose a hierarchical framework, by fusing bothSensor-wise andAgent-wise features forPerception-Improving (SAPI), which can adaptively integrate features from different information sources to improve perception capabilities. Specifically, to facilitate scene understanding, we assign prior knowledge to the visual coder to generate efficient embeddings. For effective feature representation, an attention-based sensor fusion network is designed to fuse sensor-level information of visual and LiDAR sensors, while graph convolution with multi-head attention mechanism is applied to aggregate agent-level information from an arbitrary number of neighbors. In addition, reinforcement learning is used to optimize the policy, where a novel compound reward function is introduced to guide training. Extensive experiments demonstrate that our method has excellent generalization ability in different scenarios and scalability for large-scale systems.
Yunjie Jia, Yong Song 0005, Bo Xiong 0001, Jiyu Cheng, Wei Zhang 0021, Simon X. Yang, Sam Kwong
IEEE Trans. Intell. Transp. Syst.2
2023 Robust Visual-Inertial Odometry Based on a Kalman Filter and Factor Graph
abstract
We present a real-time, high-accuracy, robust, tightly coupled visual-inertial odometry (VIO) algorithm, including monocular-inertial odometry and stereo-inertial odometry, and uses inertial measurement unit (IMU) pre-integration that is based on fourth-order Runge–Kutta (PK4) and IMU initialization based on maximum a posteriori (MAP) estimation. In particular, we used the multi-state constraint Kalman filter (MSCKF) to fuse vision and IMU measurement data for state estimation. In the optimization stage, we simultaneously considered and optimized all of the historical constraints, and performed multiple iterations to reduce the linearity errors. For further reducing the cumulative error and improving the relocation accuracy, we used a bag-of-words model for global optimization. To lower the computational cost and increase the real-time performance, we set keyframe insertion mechanism and introduced sliding window, and used a new form of Kalman gain that converts the Kalman gain in multi-state constraint Kalman filtering into the inverse of the state dimension. We validated the proposed method by using the EuRoC MAV dataset and KITTI dataset. We performed physics experiments in an outdoor environment with unstable light, to further validate the accuracy and robustness of our method.
Bao Pang, Yong Song 0005, Xianfeng Yuan, Qingyang Xu, Yibin Li 0001
IEEE Trans. Intell. Transp. Syst.3
2021 Effect of random walk methods on searching efficiency in swarm robots for area exploration
Bao Pang, Yong Song 0005, Chengjin Zhang, Runtao Yang
Appl. Intell.2
2021 Scene image and human skeleton-based dual-stream human action recognition
Qingyang Xu, Wanqiang Zheng, Yong Song 0005, Chengjin Zhang, Xianfeng Yuan, Yibin Li 0001
Pattern Recognit. Lett.3
2019 Bacterial foraging optimization based on improved chemotaxis process and novel swarming strategy
Bao Pang, Yong Song 0005, Chengjin Zhang, Hongling Wang, Runtao Yang
Appl. Intell.2
2017 A chaotic coverage path planner for the mobile robot based on the Chebyshev map for special missions
abstract
We introduce a novel strategy of designing a chaotic coverage path planner for the mobile robot based on the Chebyshev map for achieving special missions. The designed chaotic path planner consists of a two-dimensional Chebyshev map which is constructed by two one-dimensional Chebyshev maps. The performance of the time sequences which are generated by the planner is improved by arcsine transformation to enhance the chaotic characteristics and uniform distribution. Then the coverage rate and randomness for achieving the special missions of the robot are enhanced. The chaotic Chebyshev system is mapped into the feasible region of the robot workplace by affine transformation. Then a universal algorithm of coverage path planning is designed for environments with obstacles. Simulation results show that the constructed chaotic path planner can avoid detection of the obstacles and the workplace boundaries, and runs safely in the feasible areas. The designed strategy is able to satisfy the requirements of randomness, coverage, and high efficiency for special missions.
Yong Song 0005, Feng-ying Wang, Yibin Li 0001
Frontiers Inf. Technol. Electron. Eng.2
2007 An Improved On-Line Sequential Learning Algorithm for Extreme Learning Machine
Bin Li 0042, Jingming Wang, Yibin Li 0001, Yong Song 0005
ISNN (1)4