Rui Song 0002

dblp:01/2743-2 · DBLP profile ↗
← Back
39ranked-venue papers
0as first author
33since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 22 since 2021Systems, architecture and hardware · 15 · 15 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021
YearPublicationVenuePosition
2026 Human-robot skill transfer-driven multimodal fusion method for robotic collaborative sewing
Tianyu Fu 0005, Dang Hou, Longjia Sang, Li-gang Jin, Fengming Li, Chaoqun Wang 0009, Rui Song 0002
Expert Syst. Appl.7
2026 BASSM: Blur-aware selective state space model for non-uniform motion deblurring in percutaneous spinal endoscopy
Yibin Li 0001, Rui Song 0002
Expert Syst. Appl.6
2026 SGBA-Net: Semantic-guided bleeding-aware segmentation network for spinal endoscopy
Rui Song 0002
Neurocomputing6
2026 Robust Generalized Partial-to-Full Point Set Registration With Overlap-Guided Bidirectional Hybrid Mixture Models for Computer-Assisted Orthopedic Surgery
abstract
In computer-assisted orthopedic surgery (CAOS), robust and accurate registration of the preoperative full bone model and the intraoperative partial point set is a prerequisite for reliable surgical navigation, yet remains highly challenging due to partial overlap, noise, and outliers. We propose the Overlap-Guided Bidirectional Hybrid Mixture Registration (OBHMR) framework for robust and accurate partial-to-full registration. First, geometric features (i.e., surface normals) extracted from raw point sets are incorporated in both correspondence estimation and transform computation. Meanwhile, we formulate a hybrid mixture model that jointly represents positions with Gaussian mixtures (GMMs) and normals with von Mises–Fisher (vMF) mixtures across the two generalized point sets. Second, a dual-branch overlap prediction network leverages feature similarity and geometric structure to provide accurate point-wise overlap scores that guide hybrid-mixture construction under partial overlap. Third, a correspondence module integrates rotation-invariant features, multi-level self-attention, and clustering-based refinement to enhance reliability under noise and misalignment. Finally, a bidirectional objective jointly aligns source-to-target and target-to-source mixtures, explicitly accounting for discrepancies induced by noise and outliers in both the preoperative and intraoperative point sets to achieve robust optimization. Extensive experiments on 1,399 femur and 1,301 hip models demonstrate superior performance over state-of-the-art methods across overlap ratios from 5% to 70%, under both isotropic and anisotropic noise and outlier ratios up to 100%, achieving errors as low as 1.27° rotation and 1.18 mm translation at 50% overlap with 2.5mm noise. Additional tests on liver and ModelNet40 confirm strong generalization across medical and non-medical data. Ablation studies further validate the contributions of normals, overlap estimation, and the bidirectional formulation.
Xinzhe Du, Zhengyan Zhang, Rui Song 0002, Yibin Li 0001, Max Q.-H. Meng, Zhe Min
IEEE Trans Autom. Sci. Eng.4
2026 Model-Based Data-Driven Kinematic Modeling of Concentric Tube Robots With Enhanced Accuracy and Physical Consistency
Han Zeng, Fuxin Du, Yibin Li 0001, Rui Song 0002
IEEE Trans Autom. Sci. Eng.5
2026 DeepBHMR: Learning Bidirectional Hybrid Mixture Models for Generalized Global Rigid Point Set Registration in Computer-Assisted Orthopedic Surgery
abstract
This paper presents a novel robust and accurate normal-assisted learning-based rigid point set registration approach, i.e., Deep Bi-directional Hybrid Mixture Registration (DeepBHMR), where normal vectors are used in both correspondence and transformation computational stages while the bi-directional registration processes are considered. DeepBHMR consists of three components, (1) the correspondence estimation network that predicts the correspondence probabilities; (2) the posterior estimation module that computes the HMMs parameters; (3) the transformation estimation module that calculates the rigid transformation matrix by utilizing the bidirectional optimization mechanism. DeepBHMR has been extensively validated on various medical data sets, outperforming state-of-the-art registration methods. For femur bones, the mean rotation error value is approximately 1° (i.e, 1.01°) and the translation error is less than 1 mm (i.e., 0.30 mm) respectively, which meets the requirement of computer-assisted orthopedic surgery. Furthermore, even (1) trained with femur data and tested on distinct shapes and (2) under the large transformation, the mean RMSE values of registration are 2.60 mm and 3.05 mm respectively, demonstrating DeepBHMR’s favorable generalizability to different data shapes and great capability to handle global registration. Additionally, the individual significant contributions and computational efficiency of adopting normal vectors and utilizing the bidirectional mechanism have been validated in ablation studies. The results demonstrate the DeepBHMR’s favorable generalizability from femur bones to hip bones and that DeepBHMR can successfully handle the large transformation or partial-to-full registration simultaneously. The code implementation of DeepBHMR has been made publicly available at https://github.com/zzyrobot/DeepBiHMM.git.
Zhengyan Zhang, Rui Song 0002, Yibin Li 0001, Max Q.-H. Meng, Zhe Min
IEEE Trans Autom. Sci. Eng.3
2026 Multi-Task Learning for Gait Phase and Gait Cycle Percentage Prediction With Wearable Sensors in Frail Older Adults
abstract
Deep learning has been widely used in wearable sensors to improve accuracy in gait analysis. However, these deep learning models typically focus on single tasks, either in gait parameter estimation or gait phase detection. This study presents a novel multi-task learning framework for regression (i.e., gait cycle percentage prediction) and classification (i.e., gait phase prediction) tasks in pathological gait analysis using wearable sensors. Our framework employs a Multi-gate Mixture-of-Experts architecture to achieve soft parameter-sharing, integrating expert networks, cross-expert attention mechanisms, and dynamic routing to balance shared and task-specific representations. To reduce computational burden in wearable applications, we compare lightweight model configurations that optimize expert count and feature dimensionality. Model performance has been validated on a public dataset consisting of 158 frail older adults, demonstrating that our framework significantly outperforms single-task learning and hard parameter-sharing baselines, achieving an accuracy of 97.56% and a Mean Absolute Error (MAE) of 0.0397. Notably, the most compact lightweight configuration reduces the parameter count by nearly 98% (from 2.118 million to 0.0469 million), achieving an accuracy of 96.47% and a MAE of 0.0549. Attention mechanisms significantly enhance performance across all configurations, with improvements ranging from 17.9% to 30.4%. These findings validate the potential of lightweight multi-task approaches for real-time gait assessment, offering promising applications for clinical evaluation and rehabilitation monitoring in geriatric populations.
Zeyang Guan, Ziyun Ding, Xin Ma 0001, Yibin Li 0001, Rui Song 0002, Huanghe Zhang
IEEE J. Biomed. Health Informatics7
2026 Enhancing spinal endoscopy visualization: a confidence-guided framework for instrument occlusion removal
Chengxiang Zhang, Rui Song 0002
Vis. Comput.5
2025 Robust and Accurate Multi-View 2D/3D Image Registration with Differentiable X-Ray Rendering and Dual Cross-View Constraints
abstract
Robust and accurate 2D/3D registration, which aligns preoperative models with intraoperative images of the same anatomy, is crucial for successful interventional navigation. To mitigate the challenge of a limited field of view in single-image intraoperative scenarios, multi-view 2D/3D registration is required by leveraging multiple intraoperative images. In this paper, we propose a novel multi-view 2D/3D rigid registration approach comprising two stages. In the first stage, a combined loss function is designed, incorporating both the differences between predicted and ground-truth poses and the dissimilarities (e.g., normalized cross-correlation) between simulated and observed intraoperative images. More importantly, additional cross-view training loss terms are introduced for both pose and image losses to explicitly enforce cross-view constraints. In the second stage, test-time optimization is performed to refine the estimated poses from the coarse stage. Our method exploits the mutual constraints of multi-view projection poses to enhance the robustness of the registration process. The proposed framework achieves a mean target registration error (mTRE) of$0.79+2.17\ \mathbf{mm}$on six specimens from the DeepFluoro dataset, demonstrating superior performance compared to state-of-the-art registration algorithms.
Rui Song 0002, Yibin Li 0001, Max Q.-H. Meng, Zhe Min
ICRA2
2025 Registration After Completion: Towards Sparse and Partial Point Set Registration for Computer-Assisted Orthopedic Surgery
abstract
In computer-assisted orthopedic surgery (CAOS), accurate point set registration is essential for enhancing surgical accuracy. However, the sparse and low-overlap nature of intraoperative point sets presents significant challenges for reliable registration. To deal with these challenges, we propose a novel registration-after-completion framework, where the intraoperative point set is first completed, after which the two full point sets are registered. Our main contributions include the follows. First, we propose a progressive two-stage strategy to progressively complete the sparse and partial intraoperative point set. Second, considering that 1) intra-operative point set contains noise 2) the point completion process is not perfect, and 3) the resolution of preoperative image is limited, we adopt the bidirectional hybrid mixture models (HMMs) to represent the point set pairs and formulate the probabilistic registration network. In the proposed novel correspondence network where a dual-path cross-attention mechanism is adopted for feature fusion and a clustering mechanism is leveraged for calculating point-to-mixture correspondences. Furthermore, the bidirectional registration mechanism is leveraged to compute the transformation based on estimated correspondences. Third, we have extensively validated the proposed approach on various datasets and bone phantoms. Our experiments on 1399 human femur and 1301 hip models demonstrate that our method achieves state-of-the-art performance across overlap rates from 15% to 35% and at various point counts (i.e., 25, 50, and 100 points) under conditions with less than 50% overlap. Additionally, real phantom experiments on femur and hip models validate the method’s performance in simulated surgical scenarios. Experiments on ModelNet40 further confirmed our method’s effectiveness and generalizability.
Xinzhe Du, Shixing Ma, Zhengyan Zhang, Rui Song 0002, Yibin Li 0001, Max Q.-H. Meng, Zhe Min
IROS4
2025 Unsupervised Liver Deformation Correction Network Using Optimal Transport for Image-Guided Liver Surgery
abstract
In this paper, we propose a novel unsupervised intraoperative liver deformation correction method, called Learning Coherent point drift Network (LCNet), for image-guided liver surgery (IGLS). We first estimate the correspondences between the preoperative and intraoperative point sets in the optimal transport (OT) module by leveraging both original points and extracted features. Afterwards, we compute the point-wise displacement vector by solving the involved matrix equation in the Transformation module, where the point localisation noise is explicitly considered and modeled. Additionally, we present three variants of the proposed approach, i.e., LCNet, LCNet-ED and LCNet-WD, where better registration performances of LCNet against the other two demonstrate the superiority of the utilised Chamfer loss. We have extensively evaluated LCNet on the MedShapeNet dataset consisting of 615 different liver shapes of real patients, and the 3Dircadb dataset comprising 20 liver models of real patients. Extensive experimental results under different deformation and noise magnitudes demonstrate that LCNet outperforms existing state-of-the-art registration algorithms and holds significant application potential in IGLS. For example, when the overlapping ratio between the preoperative and intraoperative point sets is 25%, the deformation magnitude is 8 mm, the maximum point localization noise magnitude is 2 mm and the rotation angle lies in the range of [−45°, 45°], LCNet achieves a root-mean-square error (RMSE) value being 3.21 mm on MedShapeNet dataset, significantly outperforming those of Lepard and RoITr being 5.41 mm (p < 0.001) and 4.90 mm (p < 0.001) respectively.
Xinzhe Du, Rui Song 0002, Yibin Li 0001, Max Q.-H. Meng, Zhe Min
IROS5
2025 Complex Robotic Manipulation via Hindsight Goal Diffusion and Graph-based Experience Replay
abstract
Goal-conditioned reinforcement learning (GCRL) is an effective method for multi-goal robotic manipulation tasks. Many studies based on hindsight experience replay (HER) and hindsight goal generation (HGG) have achieved the autonomous acquisition of robotic manipulation in reward-sparse environments and have greatly improved the learning efficiency of GCRL. However, these methods perform poorly in environments with obstacles and distant goals. In this paper, we propose hindsight goal diffusion and graph-based experience replay (HGD-GER) for complex robotic manipulation. First, obstacle-avoiding graphs in environments with obstacles are constructed, and the graph-based distance metric between different goals is established. Second, the proposed HGD approach utilizes the inherent denoising mechanism of diffusion models and obstacle-avoiding graph-based distance to generate exploration goals, thereby promoting the exploration of obstacle-bypassing areas. Then, GER module modifies the reward value of experience replay by graph-based distance, thereby avoiding the bias introduced by HER and improving the learning performance of the RL algorithm under sparse reward conditions. Finally, we conducted experiments on three robotic manipulation tasks with obstacles and distant goals, and the results show that the proposed HGD-GER achieves excellent learning performance. Additionally, the proposed method is deployed on the physical robot.
Jinrui He, Yong Song 0005, Pingping Liu, Qingyang Xu, Xianfeng Yuan, Rui Song 0002
IROS8
2025 Revisiting 3D Curve to Surface Registration using Tangent and Normal Vectors for Computer-Assisted Orthopedic Surgery
abstract
In this paper, we present a novel curve-to-surface registration method, termed Bi-directional Hybrid Mixture Model Registration based on Dual-constrained Tangent and Normal Vectors (BiHMM-DTN), where two different tangent vectors at the intraoperative point are simultaneously used with the normal vector at the corresponding preoperative point to construct the geometric constraints. While hybrid registration models incorporating tangent and normal vectors (HMM-TN) demonstrate success, their geometric constraints prove inadequate or inappropriate for sparse intraoperative point sets, frequently yielding suboptimal optimization outcomes. By critically revisiting the geometric constraints of HMM-TN, we propose a dual-constraints-based hybrid mixture model registration framework with enhanced intraoperative point set acquisition protocols. To deal with noise and outliers in preoperative and intraoperative point sets—caused by reconstruction inaccuracies and tracking errors, respectively—our approach employs a bi-directional registration mechanism for curve-to-surface registration. We provide rigorous proofs validating the geometric completeness of the dual constraints within this mechanism. The BiHMM-DTN framework is formulated as a maximum likelihood estimation (MLE) problem and optimized using an expectation-maximization (EM) algorithm. Furthermore, to enhance convergence stability and accelerate optimization, the rotation matrix is updated iteratively through successive incremental steps. Extensive experiments on human femur and hip models demonstrate that our method outperforms state-of-the-art approaches, including both traditional optimization and deep learning methods, under various noise and outlier conditions. Furthermore, real-world phantom experiments highlight the potential clinical value of our method for surgical navigation applications. The codes and data are available at https://github.com/sam-zyzhang/BiHMM-DTN.git.
Zhengyan Zhang, Xinzhe Du, Rui Song 0002, Max Q.-H. Meng, Zhe Min
IROS3
2025 A Crab-Inspired Soft Gripper with Single-Finger Dexterous Grasping Capabilities
abstract
Soft grippers conform to the shape and surface properties of the objects to be grasped, effectively avoiding damage to soft and fragile items. Despite the variety of existing soft gripper designs, their structures lack sufficient flexibility for effectively grasping slender objects or operating in narrow spaces. To address these challenges, we propose a soft gripper with single-finger grasping capabilities, inspired by the structure of crab claws. The structural design and the fabrication method of the gripper are introduced, and the analytical bending model is derived. Experiments are conducted under typical operating conditions to validate the model, and the results indicate that the measured data are in good accordance with the predicted responses. Furthermore, a series of grasping experiments are carried out to test the single-finger grasping capabilities of the proposed soft gripper. The results indicate that the proposed soft gripper can efficiently and stably grasp slender or irregular objects with a single finger. In particular, it demonstrates suitability for operations in narrow spaces and shows potential for handling complex tasks. This innovative design effectively reduces the complexity of the system, while exhibiting promising capabilities in grasping slender or irregular objects and operating within restricted spaces.
Yunce Zhang, Haobin Lv, Yixiang Liu, Zhe Min, Shizhao Zhou, Tao Wang 0072, Shiqiang Zhu, Rui Song 0002
IROS8
2025 Directed Spatial Consistency-Based Partial-to-Partial Point Cloud Registration with Deep Graph Matching
abstract
3D point cloud registration is an essential problem in computer vision, robotics, surgical navigation and augmented reality. Accurate registration of partially overlapped intraoperative point clouds (e.g., femoral reconstruction) remains critical yet challenging in orthopedic navigation due to incomplete overlap and dynamic noise. In this study, we propose a partial-to-partial point cloud registration framework based on directional spatial consistency. First, we extract overlapped areas from partially overlapping point clouds and leverage the point registration graph matching module to calculate the hard point matching matrix. Second, we sample nodes from the source point cloud and generate translation-invariant edge vectors (direction/scale-preserving) via their k-nearest neighbors, guided by predicted point correspondences. This bypasses translation ambiguities by encoding spatial consistency through edges, reducing pose estimation to 3DoF alignment (rotation). The loss explicitly couples point-level matches with edge-level geometric constraints for dual optimization. Building upon this framework, we extract reliable overlapping edge representations and prune their similarity matrix by thresholding low-confidence scores, effectively suppressing spurious matches. The proposed edge-aware matching mechanism further exploits the translation invariance of local structures to refine point correspondences with enhanced accuracy. Finally, we introduce a bidirectional registration mechanism to reinforce optimization stability, achieving state-of-the-art performance across benchmarks. Extensive experiments on ModelNet40, ShapeNet, and MedShapeNet validate our method under diverse scenarios: partial-to-partial, unseen categories, partial-to-full, and cross-dataset generalization, surpassing existing methods in registration accuracy. The codes are available at https://github.com/pidan0824/DSCGM.
Kexue Fu 0001, Xinzhe Du, Rui Song 0002, Max Q.-H. Meng, Zhe Min
IROS4
2025 Hierarchical reinforcement learning with curriculum demonstrations and goal-guided policies for sequential robotic manipulation
Bao Pang, Xianfeng Yuan, Xiaolong Xu 0003, Yong Song 0005, Rui Song 0002, Yibin Li 0001
Eng. Appl. Artif. Intell.6
2025 An adaptive reinforcement learning approach with trait-awareness for heterogeneous multi-robot cooperative pursuit
Heteng Zhang, Yunjie Jia, Yong Song 0005, Bao Pang, Xianfeng Yuan, Rui Song 0002, Simon X. Yang
Eng. Appl. Artif. Intell.6
2025 Robot Strategy Transfer Based on Shared Feature Space for Search and Insertion Assembly
abstract
Traditional assembly tasks often require robots to transfer the acquired skills to new tasks. However, previous transfer reinforcement learning methods typically ignore the inherent relationship between the source and the target domain tasks. This requires a substantial amount of interaction data to compensate for this deficiency, and generally results in poor transfer effects. To address this issue, a strategy transfer method that establishes a shared feature space between the source domain and the target domain is proposed to enhance the efficiency of strategy learning on peg-in-hole assembly. Initially, by calculating the distance between each feature in the source and target domains, the features with small distance are selected as shared features. Subsequently, in order to determine the successful search state, this paper uses the jump state of contact force and the relative position between the peg and the hole as the judgment criterion. Lastly, search and insertion peg-in-hole assembly experiments are conducted to validate the generalization of the proposed strategy, demonstrating its capability to transfer from simulation to the real world.
Li-gang Jin, Yu Men, Fengming Li, Chaoqun Wang 0009, Xincheng Tian, Yibin Li 0001, Rui Song 0002
IEEE Trans. Circuits Syst. I Regul. Pap.7
2025 Empowering Multirobot Flocking in Complex Environments via Effective Communication: A Deep Reinforcement Learning Approach
abstract
Multirobot flocking is crucial for safe and cooperative navigation, with wide applications in logistics, service delivery, and mobile surveillance. Despite significant progress, developing effective flocking strategies under complex conditions remains challenging. Communication is a vital technique for multirobot coordination. In this article, we propose refinement and enhancement of communication information (REIN), a novel deep reinforcement learning-based framework designed to improve communication effectiveness in leader–follower flocking systems through the REIN. First, regarding information refinement, a graph-based information refiner, integrating directed graph-structured communication with an innovative edge filter, is developed for selective multirobot interaction. It helps robots adaptively focus on relevant neighbors, considerably alleviating information overload. Second, for information enhancement, a cognition-aligned information enhancer is designed that boosts information expressiveness by encouraging team consensus. It utilizes two cascaded leader-related objectives to optimize information towards cognitive alignment among decentralized followers. Extensive comparisons with state-of-the-art approaches and ablation versions demonstrate the superiority of our framework. Physical experiments are also conducted to validate its practicality.
Yunjie Jia, Yong Song 0005, Jiyu Cheng, Heteng Zhang, Wei Zhang 0021, Rui Song 0002, Simon X. Yang, Sam Kwong
IEEE Trans. Ind. Informatics6
2025 Goal-Conditioned Reinforcement Learning With Adaptive Intrinsic Curiosity and Universal Value Network Fitting for Robotic Manipulation
abstract
Hindsight experience replay (HER) has greatly increased the possibility of using deep reinforcement learning (DRL) for robotic manipulation with sparse rewards. However, there are still concerns about low learning efficiency and poor performance due to its insufficient exploration ability and bias against the initial goal introduced by HER. In this article, to solve this problem, a multigoal robotic manipulation DRL method based on adaptive intrinsic curiosity and universal value network fitting (AIC-UVNF) is proposed to further improve the exploration ability and learning performance. Specifically, this method utilizes an improved curiosity mechanism to construct a joint intrinsic reward and adaptively adjust the proportion, which can enhance exploration ability and avoid excessive pursuit of novel states. In addition, a universal value network fitting approach is proposed to incorporate the initial goal into the value function fitting process, which employs the value of the initial goal to eliminate the bias of HER in the algorithm update. Combined with the off-policy soft actor-critic method, AIC-UVNF is verified on multigoal robotic manipulation tasks. The results show that the proposed method achieves better convergence efficiency and learning performance.
Xianfeng Yuan, Qingyang Xu, Bao Pang, Yong Song 0005, Rui Song 0002, Yibin Li 0001
IEEE Trans. Ind. Informatics6
2024 History-Aware Planning for Risk-free Autonomous Navigation on Unknown Uneven Terrain
abstract
It is challenging for the mobile robot to achieve autonomous and mapless navigation in the unknown environment with uneven terrain. In this study, we present a layered and systematic pipeline. At the local level, we maintain a tree structure that is dynamically extended with the navigation. This structure unifies the planning with the terrain identification. Besides, it contributes to explicitly identifying the hazardous areas on uneven terrain. In particular, certain nodes of the tree are consistently kept to form a sparse graph at the global level, which records the history of the exploration. A series of subgoals that can be obtained in the tree and the graph are utilized for leading the navigation. To determine a subgoal, we develop an evaluation method whose input elements can be efficiently obtained on the layered structure. We conduct both simulation and real-world experiments to evaluate the developed method and its key modules. The experimental results demonstrate the effectiveness and efficiency of our method. The robot can travel through the unknown uneven region safely and reach the target rapidly without a preconstructed map.
Yinchuan Wang, Nianfei Du, Yongsen Qin, Rui Song 0002, Chaoqun Wang 0009
ICRA5
2024 OBHMR: Robust Partial-to-full Generalized Point Set Registration with Overlap-guided Bidirectional Hybrid Mixture Model
abstract
In this paper, we introduce a novel overlap-based bidirectional point set registration approach, i.e., Overlap-guided Bidirectional Hybrid Mixture Registration (OBHMR), which incorporates geometric information (i.e., normal vectors) in both the correspondence and transformation stages and formulates the optimization objective of registration in a bidirectional manner. More importantly, to address the issue of partial-to-full registration, OBHMR utilises the predicted point-wise overlap score using networks to formulate the overlap-guided Hybrid Mixture Model consisting of the Gaussian Mixture Model (GMM) and Fisher Mixture Model (FMM). OBHMR contains four components: (1) the overlap-guided correspondence network that estimates the correspondence probabilities and calculates the point-wise overlap score; (2) the learning posterior module that estimates the overlap-guided HMM parameters; (3) the transformation module that computes the rigid transformation by formulating the optimisation objective in a bidirectional registration way, given correspondences and overlap-guided HMM parameters. Experiments using 1457 human femur and 1301 human hip models demonstrate significant improvements in partial-to-full registration performance (p < 0.01) under different overlapping ratios, compared to state-of-the-art registration approaches. Furthermore, individual contributions of three modules (i.e., additional normal vectors, overlap score estimation module and the bidirectional mechanism) in OBHMR have been validated in ablation studies. The results demonstrate OBHMR’s capability of tackling the challenging partial-to-full registration problems in computer-assisted orthopedic surgery. The codes are available at https://github.com/Dxinz/DeepOBHMR.
Xinzhe Du, Zhengyan Zhang, Rui Song 0002, Yibin Li 0001, Max Q.-H. Meng, Zhe Min
IROS4
2024 DeepBHMR: Learning Bidirectional Hybrid Mixture Models for Generalized Rigid Point Set Registration
abstract
In this paper, we introduce a novel normal-assisted learning-based rigid registration approach, i.e., Deep Bi-directional Hybrid Mixture Registration (DeepBHMR). Our approach utilises helpful normal vectors explicitly in both correspondence and transformation stages and formulates the optimization objective of registration in a bi-directional way that considers noise in both point sets. DeepBHMR consists of three modules: (1) the correspondence network that estimates the correspondence probability relating points within one generalized point set (i.e., positional and normal vectors) with components of Hybrid Mixture Models (HMMs) representing the other generalized point set; (2) the posterior module that computes HMMs parameters; (3) the transformation module that computes the rotation matrix and the translation vector given the estimated generalized-point to hybrid-distribution correspondences and HMMs parameters. DeepBHMR has been validated on 291 human femur and 260 hip models, and extensive experimental results demonstrate that DeepBHMR outperforms the state-of-the-art registration methods (p-value < 0.01). In the circumstance of femur bones, the mean rotation and translation error values are around 1° (i.e., 1.01°) and less than 1 mm (i.e., 0.36mm), respectively. Furthermore, even under the large transformation (i.e., in the range of [0,180]° and [0, 100] mm), the mean RMSE values being 3.05 mm is still satisfactory. Additionally, the results demonstrate the DeepBHMR’s favorable generalizability from femur shapes to hip shapes. We have carefully validated the significant benefits of incorporating normal vectors and the bidirectional mechanism. DeepBHMR can successfully handle the challenging scenario of large transformation and partial registration. The codes are available at https://github.com/zzyrobot/DeepBHMR.git.
Zhe Min, Zhengyan Zhang, Rui Song 0002, Yibin Li 0001, Max Q.-H. Meng
IROS4
2024 Bidirectional Partial-to-Full Non-Rigid Point Set Registration with Non-Overlapping Filtering
abstract
In this paper, we introduce Bidirectional Non-Overlapping Filtering Network (Bi-NOFNet), which registers the partial intraoperative point set with full preoperative point set for computer-assisted interventions (CAI). Our contributions are three-folds. First, Bi-NOFNet adopts customised feature extractor to extract distinctive features from both point sets, with which the per-point overlap mask is predicted and the overlapping region is segmented for the preoperative point set. Furthermore, we propose two methods to filter out the non-overlapping regions, at feature-level (i.e., Bi-NOFNet(Feature)) and point-level (i.e., Bi-NOFNet (Point)). For these two methods, we develop supervised registration strategy where the ground-truth overlap mask and displacement vectors are employed, and weakly-supervised registration strategies where only the ground-truth overlap mask is available. Additionally, to fully utilise the information in both space, we propose a bidirectional registration mechanism, which predicts the displacement vectors associated with the intraoperative point set (i.e., the forward way) and those warpping the preoperative point set (i.e., the backward way). Experiments have been conducted on the proposed DeformMedShapeNet dataset that contains 615 different liver shapes. Extensive results demonstrate that Bi-NOFNet performs well for partial-to-full registration tasks under various scenarios of noise, overlap ratios and deformation levels, outperforming existing non-rigid registration approaches.
Rui Song 0002, Yibin Li 0001, Max Q.-H. Meng, Zhe Min
IROS3
2024 Distributed LSTM-GCN-Based Spatial-Temporal Indoor Temperature Prediction in Multizone Buildings
abstract
Indoor temperature prediction of multiple zones in near future horizons is vital in developing an optimal regulation strategy of heating, ventilation, and air conditioning systems in large-scale complex buildings. This is, however, challenging due to the spatial–temporal correlation and multivariable coupling characteristics. This article proposes a novel deep learning framework incorporating the distributed long short-term memory and graph convolution network namely DL-GCN for indoor temperature prediction in large public buildings, aiming to learn the spatial–temporal correlation and multivariable coupling features. First, the indoor temperature and humidity data from different zones are handled by GCN networks to extract the temperature spatial features. Then, in the distributed LSTM module, other data, such as light and ac power consumption, are fused with the outputs of the GCN module, respectively, in a distributed way to learn the coupling interactions and temporal characteristics between these variables. Comparison study and ablation experiments are conducted using real datasets from a large-scale building to verify its effectiveness and superior performance in multizone indoor temperature prediction.
Xinli Wang, Xiaohong Yin, Kang Li 0002, Lei Wang 0184, Rui Song 0002
IEEE Trans. Ind. Informatics7
2023 Design and Development of a Rapidly Deployable Low-Cost Tensegrity In-Pipe Robot
abstract
Existing in-pipe robots have insufficient adaptability when dealing with accidents in unfamiliar pipe environments. Developing a pipe robot that can be designed and manufactured quickly is one solution. The tensegrity structure is a self-stressing spatial structure formed by the interaction of rigid members and flexible cables, which has the advantages of simple structure, good flexibility, deformability, and impact resistance. Inspired by this structure, we design a novel worm-like tensegrity robot for different pipe environments, which can be manufactured rapidly at low cost. Firstly, a robotic module based on the tensegrity structure is designed inspired by the motion patterns of worm-like organisms. Then, the design process of the module is presented based on the mathematical analysis of the deformation. Finally, a prototype of the tensegrity robot is developed using simple and low-cost parts in less than an hour. To test the motion performance, load performance, and inspection capability of the tensegrity robot, we designed a series of experiments on horizontal pipes, vertical pipes, elbows, and steel pipes. Experimental results show that the worm-like tensegrity robot is simple in structure, easy to manufacture, low in cost, and good in performance.
Yixiang Liu, Xiaolin Dai, Kai Guo 0004, Jiang Wu 0018, Rui Song 0002, Jie Zhao 0003, Yibin Li 0001
IROS5
2023 An LSTM and ANN Fusion Dynamic Model of a Proton Exchange Membrane Fuel Cell
abstract
A proton exchange membrane fuel cell (PEMFC) has great application prospects due to its low emission and high efficiency. An accurate model to predict the dynamic output voltage is essential for the optimal control of the PEMFC for the applications on vehicles and power stations. In this article, a novel deep learning framework with the long short-term memory (LSTM) and artificial neural network (ANN) fusion is proposed to develop the PEMFC dynamic model by extracting both the historical and current information. The LSTM extracts the temporal information from the past PEMFC states with its order determined with autocorrelation and partial autocorrelation functions, while the influence of the system current inputs is learnt by the ANN. Then, the outputs of LSTM and ANN are concatenated with the multiple information fused to predict the PEMFC dynamic output voltage. After validated by the operating data from a laboratory-scale PEMFC system, the LSTM and ANN fusion model is compared with the existing models, such as support vector regression, ANN, and LSTM methods. The comparison results show that the proposed LSTM and ANN fusion model can provide the best prediction performance with the lowest mean square error of 1.303. The proposed LSTM and ANN fusion model can be helpful to develop the optimal control strategy of the PEMFC.
Wei Li 0187, Xinli Wang, Lei Wang 0184, Lei Jia 0003, Rui Song 0002, Zhichao Fu
IEEE Trans. Ind. Informatics5
2022 Low-drift LiDAR-only Odometry and Mapping for UGVs in Environments with Non-level Roads
abstract
This study focuses on localization and mapping for UGVs when they are deployed in environments with non-level roads. In these scenarios, the vehicles need to travel through flat but not necessarily level grounds, i.e., ascent or descent, which may cause drifts of the robot pose and distortion of the map. We develop a low-drift LiDAR odometry and mapping approach for the UGV with LiDAR as the only exteroceptive sensor. A factor-graph based pose optimization method is developed with a specifically designed factor named slope factor. This factor includes the slope information that is estimated from a real-time LiDAR data stream. The slope information is also used to enhance the loop-closure detection procedure. Moreover, an incremental pitch estimation mechanism is designed to achieve further pose estimation refinement. We demonstrate the effectiveness of the developed framework in real-world environments. The odometry drift is lower and the map is more precise than experiments with the state-of-the-arts. Notably, on the Kitti dataset, our method also exhibits convincing performance, demonstrating its strength in more general application scenarios.
Yinchuan Wang, Chaoqun Wang 0009, Rui Song 0002, Yibin Li 0001
IROS4
2022 An In-pipe Crawling Robot based on Tensegrity Structures
abstract
This paper presents a novel concept to develop robots capable of crawling in tubular environments, inspired by the movement of earthworms and the biological musculoskeletal systems in nature. A tensegrity structures-based robotic module with shape changeability actuated by only one linear actuator is proposed. The mechanical structure of the robotic module is determined on the basis of force density method. By serially cascading three uniform modules, the in-pipe crawling robot is designed and manufactured. The robot has the abilities to crawl in both horizontal and vertical pipes with different inner diameters, and to pass through elbow pipes adaptively under the control of a simple actuation sequence. The effectiveness of the robot is demonstrated by experimental results on the prototype. Compared with existing robots, this proposed approach enables compact yet robust structures, along with enhanced compliance, mobility, and adaptability.
Yixiang Liu, Qing Bi, Xiaolin Dai, Rui Song 0002, Xizhe Zang, Yibin Li 0001
IROS4
2022 Adaptive neural control for mobile manipulator systems based on adaptive state observer
Yukun Zheng, Yixiang Liu, Rui Song 0002, Xin Ma 0001, Yibin Li 0001
Neurocomputing3
2021 Autonomous cognition development with lifelong learning: A self-organizing and reflecting cognitive network
Xin Ma 0001, Rui Song 0002, Xuewen Rong, Yibin Li 0001
Neurocomputing3
2021 Predicting academic performance of students in Chinese-foreign cooperation in running schools with graph convolutional network
Haitao Pu, Mingqu Fan, Hong-Bin Zhang, Bi-Zhen You, Jinjiao Lin, Chunfang Liu, Yanze Zhao, Rui Song 0002
Neural Comput. Appl.8
2021 Moment-based multi-lane detection and tracking
Rui Song 0002, Hui Chen 0009, Reinhard Klette, Yanyan Xu 0002
Signal Process. Image Commun.2
2020 A self-organizing developmental cognitive architecture with interactive reinforcement learning
Xin Ma 0001, Rui Song 0002, Xuewen Rong, Xincheng Tian, Yibin Li 0001
Neurocomputing3
2019 Robot skill acquisition in assembly process using deep reinforcement learning
Fengming Li, Sisi Zhang, Rui Song 0002
Neurocomputing5
2016 Nonlinear non-negative matrix factorization using deep learning
abstract
In this paper, we describe the deep learning method to reduce the dimension of the data samples under the framework Non-negative Matrix Factorization (NMF). That is to say, we try to find the good representation of the data samples for the task of NMF. To this end, a nonlinear NMF optimization model is constructed and the optimization algorithm is developed. The experimental results on some benchmark dataset show the nonlinear dimension reduction helps the NMF to improve the clustering performance.
Hui Zhang 0092, Huaping Liu 0001, Rui Song 0002, Fuchun Sun 0001
IJCNN3
2016 Nonlinear dictionary learning based deep neural networks
abstract
In this paper, we demonstrate nonlinear features extracted by deep neural network have better results in the task of dictionary learning. A nonlinear dictionary learning model is constructed and the optimization algorithm is developed. In the learning algorithm, we use the deep neural network to convey raw samples to feature space and learn a nonlinear dictionary. The extensive experimental results of classification on some benchmark dataset show the proposed nonlinear dictionary significantly improves the classification performance.
Hui Zhang 0092, Huaping Liu 0001, Rui Song 0002, Fuchun Sun 0001
IJCNN3
2015 Fingertip-based interactive projector-camera system
Jun Cheng 0002, Rui Song 0002, Xinyu Wu 0001
Signal Process.3
2014 Video anomaly detection based on a hierarchical activity discovery within spatio-temporal contexts
Dan Xu 0006, Rui Song 0002, Xinyu Wu 0001, Nannan Li 0001, Wei Feng 0009, Huihuan Qian
Neurocomputing2