EDBT 2026 Demo / reviewers in the wild / expert
Dihua Zhai
dblp:133/3604 · also Di-Hua Zhai
· DBLP profile ↗
47ranked-venue papers
10as first author
37since 2021 · last 2026
0000-0001-8653-8626ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 4 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GraphGrasp: Lightweight and Efficient Graph-Guided 6-DoF Robotic Grasp Pose Estimation Networkabstract6-DoF object grasping is a crucial skill for embodied intelligent robots. Previous methods often rely on large-scale networks for feature extraction, followed by grasp pose prediction, which increases the network's parameter count and overlooks the geometric and graph features of the point cloud. To address these challenges, we propose GraphGrasp, a graph-guided 6-DoF grasping pose prediction method. It performs graph analysis from the perspectives of scene, object, and grasping graphs. First, we introduce a graph feature embedding method based on local-global features to model the scene graph effectively. Then, we use a graph transformer strategy to represent spatial relationships between objects in the object graph. Finally, we propose a multi-metric, multi-level grasp pose evaluation algorithm to predict and explore graspable points, enabling effective construction of grasp graphs and accurate grasp pose evaluation. We test GraphGrasp on the GraspNet-1Billion dataset, and the results show that, compared to previous methods, it achieves nearly the same performance with about 1/5 of the parameters of state-of-the-art methods, significantly improving grasp pose prediction speed. Additionally, in real-world robot grasping scenarios, GraphGrasp outperforms previous methods in practical grasp pose prediction tasks. Sheng Yu 0009, Dihua Zhai, Yuanqing Xia |
AAAI | 2 |
| 2026 | MMSegRWKV: Enhancing Multimodal MRI Segmentation for Internet-of-Medical-Things-Enabled Healthcare With RWKV-Inspired Architectures
Yitong Cao, Yuanqing Xia, Ke Tian, Dihua Zhai |
IEEE Internet Things J. | 4 |
| 2026 | Real-Time High-Precision Control of Robot Manipulators: An Adaptive Data-Driven Linear MPC FrameworkabstractHigh-precision control of constrained robot manipulators under uncertainties remains a fundamental challenge. While conventional nonlinear model predictive control (NMPC) often struggles with real-time requirements due to its heavy computational burden, linear MPC (LMPC) typically relies on terminal invariant sets that are often computationally intractable for complex nonlinear systems. To address these limitations, this paper proposes a computationally efficient data-driven linear model predictive control (DLMPC) framework that achieves control accuracy comparable to NMPC while ensuring real-time performance. A variable-length sliding-window dynamic mode decomposition with control (DMDc) method is developed to identify a time-varying local affine model from recent input–output data, enabling accurate linearization under uncertainties. Based on this model, a novel time-varying terminal constraint is designed to substitute the conventional terminal set, thereby obviating the need for uncertainty upper bounds that are difficult to obtain in practice. The recursive feasibility and stability of the proposed framework are established using Lyapunov stability theory. Finally, experimental results on a Franka Emika Panda robot demonstrate the effectiveness and superior performance of the proposed method. Qianchen Guo, Zhihang Sun, Wentao Ning, Dihua Zhai, Yuanqing Xia |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2026 | RGB-Based Category-Level Object Pose Estimation With Multi Pose Maps for Robotic Grasp DetectionabstractRGB-D-based category-level object pose estimation has achieved very good pose estimation results in robotic grasp tasks. However, these methods rely on accurate depth information, and in industrial scenes where depth information is unknown or subject to significant interference, these methods are unable to apply. Therefore, this paper investigates the problem of category-level object pose estimation based solely on RGB images. First, to accurately predict the Normalized Object Coordinate Space (NOCS) map of objects in the scene, we build an encoding-decoding structure to achieve accurate NOCS map prediction. Then, to ensure an accurate transformation from NOCS maps to Intra-class Variation-Free Consensus (IVFC) maps, we propose a new deformable convolution that reduces the network’s computational load while improving the prediction accuracy of the IVFC maps. Finally, to fully utilize the information contained in multiple pose maps (NOCS maps and IVFC maps), we propose a Multi-Pose Map-based Pose (MPMPose) computation method to accurately predict the object pose. We test our method on the CAMERA25, REAL275, and Wild6D datasets, and the experimental results show that our proposed MPMPose can effectively complete the pose estimation task of unknown objects in scenes based solely on RGB images. Finally, we apply MPMPose to the robot grasping task in a real-world scenario. The experimental results show that MPMPose can effectively assist robots in completing the pose estimation task of objects in real scenes, enabling stable object grasping by robots. Sheng Yu 0009, Dihua Zhai, Yunqiao Zeng, Yuanqing Xia |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2026 | FreePose: Zero-Shot 6D Object Pose Estimation Using Pretrained Foundation ModelsabstractAn accurate 6D object pose estimation is essential for robotic manipulation and augmented reality applications. Existing methods typically require extensive training for new objects, limiting their effectiveness in dynamic environments where new objects are frequently introduced. In this paper, we propose FreePose, an efficient free-trained zero-shot 6D pose estimation method leveraging pre-trained visual and geometric foundation models. Our approach includes an offline onboarding stage, in which multiple viewpoint templates of a reference object are rendered, then visual and geometric features are extracted using visual and geometric pretrained models, respectively. These visual features are then back-projected onto corresponding 3D points, enabling a precise alignment between appearance and geometry, and subsequently fused with geometric features to form a robust unified representation. During inference stage, target object instances are segmented from RGB-D image using SAM2 coupled with an object-matching algorithm. Visual features of each target instance is similarly extracted, back-projected, and fused with geometric features. Robust 3D-3D correspondences are then established using nearest-neighbor search. Finally, pose estimation is obtained using the TEASER registration algorithm. Extensive evaluations conducted on the BOP5 core datasets show that our approach achieves results comparable to state-of-the-art methods. To highlight the effectiveness and potential of FreePose in real-world scenarios, FreePose is deployed on a real UR3 robot to perform grasping experiments reaching a success grasp rate of 65.0%. Abdulrahman Alsumeri, Dihua Zhai, Yuanqing Xia |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | KeyPose: Category-Level 6D Object Pose Estimation with Self-Adaptive KeypointsabstractCategory-level object pose estimation is an important task in computer vision. Some prior methods based on assumptions often struggle with drastic changes in object appearance. To address this challenge, we propose a new method for object pose estimation based on object-adaptive keypoints. In this paper, we first introduce a transformer-based keypoint prediction method for adaptive forecasting of point cloud keypoints. This method calculates the similarity between keypoint features and point cloud features, allowing keypoints to represent object geometry more effectively. Furthermore, to enhance the geometric feature construction of keypoints, we propose a graph-based keypoint feature aggregation method, which considers both the structural relationships between keypoints and the point cloud, strengthening the network's understanding of geometric structures. At this stage, keypoints remain at the geometric spatial level of the object and have not been predicted in NOCS. To improve the accuracy of keypoint prediction in NOCS, we design a NOCS voxelization method that divides NOCS into multiple voxels and accurately predicts NOCS keypoints within these voxels. Experimental results on multiple benchmark datasets demonstrate that our proposed KeyPose method outperforms all existing methods, achieving over 20% improvement in pose accuracy on some critical datasets. Sheng Yu 0009, Dihua Zhai, Yuanqing Xia |
AAAI | 2 |
| 2025 | RCGNet: RGB-based Category-Level 6D Object Pose Estimation with Geometric GuidanceabstractWhile most current RGB-D-based category-level object pose estimation methods achieve strong performance, they face significant challenges in scenes lacking depth information. In this paper, we propose a novel category-level object pose estimation approach that relies solely on RGB images. This method enables accurate pose estimation in real-world scenarios without the need for depth data. Specifically, we design a transformer-based neural network for category-level object pose estimation, where the transformer is employed to predict and fuse the geometric features of the target object. To ensure that these predicted geometric features faithfully capture the object’s geometry, we introduce a geometric feature-guided algorithm, which enhances the network’s ability to effectively represent the object’s geometric information. Finally, we utilize the RANSAC-PnP algorithm to compute the object’s pose, addressing the challenges associated with variable object scales in pose estimation. Experimental results on benchmark datasets demonstrate that our approach is not only highly efficient but also achieves superior accuracy compared to previous RGB-based methods. These promising results offer a new perspective for advancing category-level object pose estimation using RGB images. Sheng Yu 0009, Dihua Zhai, Yuanqing Xia |
IROS | 2 |
| 2025 | A Hybrid Framework Based on Bio-Signal and Built-in Force Sensor for Human-Robot Active Co-CarryingabstractHuman-robot collaboration represents a promising avenue for applications in future factory scenarios. Existing work predominantly concentrates on passive assistance based on real-time sensing sensor like force sensors, while only a few studies have ventured into the exploration and realization of active assistance provided by robots with the assistance of predictive sensors. This paper proposes an innovative hybrid framework that combines human bio-signal information with built-in robot force sensors to implement human-robot active collaboration. First, a Hill-type muscle-skeleton model is adopted and calibrated through partial swarm optimization (PSO). With this model, surface electromyography(sEMG) is used to estimate human limb stiffness. Then, an Radial Basis Function neural network (RBFNN) compensator is developed to account for the uncertainty in human-object-robot dynamics. Subsequently, we propose an adaptive variable impedance controller, incorporating a global bias into the neural network architecture. This innovative modification serves to augment the system’s robustness, streamline the network configuration by curtailing the number of hidden neurons, and consequently, facilitate more consistent and efficient human-robot interaction behavior. Finally, we substantiate the effectiveness of the proposed methodology through a two-link robotic simulation experiment and a real-world co-carrying task employing with the Baxter robot and human partner. These rigorous evaluations unveil a significant alleviation of task-related human workload attributed to our proposed framework.Note to Practitioners—This framework aims to address the existing research gap in human-robot collaboration, particularly involving bio-signal utilization, to facilitate perceptive active assistance within a typical industrial assembly scenario. In such representative tasks, many studies primarily employ real-time sensors such as force, position. These sensors, while essential, are limited by their detection principles and require collaborative operation to ascertain stiffness and realize passive assistance. Conversely, bio-signals intrinsically contain stiffness information and exhibit prospective characteristics that can be leveraged for stiffness prediction. In this typical task, we design a human-robot co-transport system with two crucial characteristics: first, the robot is capable of detecting human stiffness tendencies and comprehending human intent, leading to self-adjusting robotic behavior that provides enhanced protection for the transported object. Secondly, the newly proposed controller can manage sudden disturbances and execute self-repairs, thus increasing the task success rate and ensuring worker safety. Leyun Hu, Dihua Zhai, Dongdong Yu, Yuanqing Xia |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Robust Whole-Body Safety-Critical Control for Sampled-Data Robotic Manipulators via Control Barrier FunctionsabstractIn this paper, a novel robust safety-critical control method is proposed to ensure whole-body safety for the robotic manipulator, which is implemented as a sampled-data system with measurement errors. The manipulator and obstacles are approximated as several spherical enclosures, and a whole-body safety constraint with relative degree two is formulated based on the distance function. Robust control barrier function (CBF) constraints are first designed to handle the predefined joint velocity constraints. Building upon this, a robust high-order CBF constraint is derived to enforce the whole-body safety constraint. Each stage of the derivation incorporates the sample-and-hold error and measurement error. These robust CBF constraints are then unified with a nominal controller to form an optimization problem, ensuring that the velocity constraints and the safety constraint are satisfied. The effectiveness of the proposed algorithm is demonstrated through simulations and experiments on a 7-degree-of-freedom (DOF) Franka Emika Panda robot. Yuhan Xiong, Dihua Zhai, Yuanqing Xia |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Independent Observers-Based Fault Diagnosis for Multiple Sensor Faults of Full-Vehicle Active Suspension Systems With Inaccurate ModelsabstractIn this paper, an independent observers-based fault diagnosis method is proposed for multiple sensor faults of the full-vehicle active suspension system with an inaccurate model. To address the problem of model uncertainties brought by the linearized full-vehicle suspension model, unmodelled dynamics, parametric uncertainties, and external disturbances are first combined into an integrated uncertain term. Disturbance observers are designed to online track the integrated uncertain terms, whose estimates will be employed to decrease the influence of model uncertainties on fault diagnosis. An independent fault diagnosis observer is designed for each sensor separately, where only the measurement of the matched sensor is taken as the observer input. In this way, each fault diagnosis observer works independently and interactions between measurements of multiple faulty sensors can be decoupled to locate the sensor faults. Anomalies of each sensor are monitored by the fault diagnosis observer in a one-to-one relationship such that fault detection and isolation can be realized at the same time. In particular, the sensitivity and robustness of the fault diagnosis method are improved with the estimates of the integrated uncertain terms injected into the fault diagnosis observers to emulate model uncertainties. Effectiveness of the proposed scheme has been validated via simulation results considering multiple sensor faults occurring at different time or simultaneously. Note to Practitioners—Good reliability of vehicle subsystems is demanded to provide stable rideability and high maneuverability for vehicles, especially for electric vehicles. To guarantee the reliability of the essential components in practice motivates us to focus on the fault diagnosis problems for sensors of vehicle suspension systems. In this paper, considering the possible scenario of multiple sensor faults, an independent observers-based fault diagnosis scheme is designed for full-vehicle suspension systems. By using the proposed method, we can detect the sensor failure and isolate the faulty sensors simultaneously. An independent fault diagnosis observer is built for each sensor with its single sensor measurement taken as the input of the observer. In this way, interactions of the incorrect measurements from the multiple faulty sensors can be removed to avoid interference from these measurement couplings in locating the faulty sensors. Anomalies of each sensor are monitored by its matched fault diagnosis observer in a one-to-one relationship. Different from the existing methods, model uncertainties brought by the linearized full-vehicle suspension model are observed online and their estimates are employed to build the fault diagnosis observers to narrow the gap between the reality and model, and promote the accuracy and effectiveness of the model-based fault diagnosis method. However, incipient sensor faults cannot be effectively diagnosed with our method, which is one drawback of this proposed fault diagnosis scheme. For future work, fault accommodation strategy will be investigated to compensate for the sensor failure and retain satisfactory system performance. Real vehicle experiments will be conducted to test and improve the algorithm, and thus promote the designed method into practical use. The proposed fault diagnosis method for multiple sensor faults can be extended to other mechanical control systems with similar structure of system equations, in addition to vehicle suspension systems. Yuanqing Xia, Dihua Zhai |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Fast and Accurate Category-Level Object Pose Estimation Without Shape Priors for Robotic Grasp DetectionabstractCategory-level object pose estimation is crucial for enabling robot grasping. Currently, many methods rely on 3D shape priors for pose estimation, but obtaining priors for specific categories often requires a significant amount of time to generate. Although some methods do not rely on priors, However, these methods struggle to achieve a balance between speed and accuracy. Achieving fast and accurate category-level object pose estimation remains a challenging issue. In this paper, we propose an algorithm called FAPose, which aims to simultaneously achieve speed and accuracy in category-level object pose estimation without any shape priors. Firstly, we design an RGB-point cloud feature aggregation method based on transformers to fuse RGB and point cloud features. Secondly, we develop a dual-constraint object pose estimation method to effectively leverage both feature space and geometric space information. This approach constructs pose constraints in both feature space and geometric space to achieve optimal pose prediction. Finally, we validate FAPose through benchmark datasets and real robot grasping experiments. The experimental results demonstrate that our proposed method surpasses most existing state-of-the-art pose estimation methods, achieving superior pose estimation performance. The code for FAPose will be available after the paper is accepted for publication. Sheng Yu 0009, Jian Yin 0032, Dihua Zhai, Yuanqing Xia |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | ZSPose: Instance-Level Zero-Shot Object Pose Estimation With Segment Anything ModelabstractEstimating the poses of new objects is a challenging problem. Although many methods have been developed for instance-level object pose estimation, they often struggle when faced with new/unfamiliar objects. In this paper, we propose a zero-shot pose estimation method for new objects called ZSPose. We leverage SAM’s zero-shot feature to segment objects in cluttered environments and acquire masks for each object. To facilitate the matching of object masks with object models and categories, we propose a novel object matching strategy that aligns masks with the corresponding object models. Subsequently, based on the derived object masks, we produce object point clouds. Utilizing the RGB images of the objects alongside the point clouds, we present a feature-weight-based method for object pose estimation, achieving accurate pose estimation by predicting the matching weights between the model features and the point cloud features. We conduct performance testing on various instance-level object pose estimation datasets, and experimental results show that our proposed method significantly enhances the accuracy of object pose estimation. It demonstrates excellent generalization, making it applicable to pose estimation for a wide range of new objects. Finally, to validate the practical applicability of ZSPose, we apply it to real-world object pose estimation tasks and robotic grasping tasks. The experimental findings indicate that ZSPose effectively estimates the poses of new objects, assisting robots in performing practical grasping tasks, thus holding considerable practical value. Sheng Yu 0009, Dihua Zhai, Jian Yin 0032, Yuanqing Xia |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | TCRNet: Transparent Object Depth Completion With Cascade RefinementsabstractTransparent objects are commonly found in real life and industrial production. Unlike opaque objects, transparent objects are not easily identifiable in RGB images and often require depth information to determine their position in the image. However, due to the influence of other environmental factors such as reflection and refraction, the depth information of transparent objects is often inaccurate. This leads to difficulties for robots in grasping transparent objects, as incorrect depth information can result in the robot being unable to predict or predict incorrectly the grasping pose. Therefore, it is necessary to complete the depth information for transparent objects. Previous methods for depth completion of transparent objects often struggle to balance accuracy and real-time performance simultaneously. To achieve this goal, in this paper, we propose a transparent object depth completion network called TCRNet based on a cascade refinement structure, which balances accuracy and real-time performance simultaneously. First, the network incorporates a cascade refinement structure in the decoding stage to refine features multiple times, improving the accuracy of depth information. Additionally, an attention module is designed to adjust the extracted features, enabling the network to focus on depth information features in transparent object regions. Finally, a transformer-based error module is implemented in the network’s final output stage to predict and adjust the error between the depth image and the ground truth. TCRNet is trained and tested on three datasets: ClearGrasp, Omniverse Object, and TransCG. It outperforms previous methods in terms of performance. Furthermore, TCRNet is applied to existing grasp detection methods to conduct grasping experiments on transparent objects using a real Baxter robot.Note to Practitioners—With the development of RGB-D camera technology, RGB-D cameras are now widely used in various scenarios such as industrial production, autonomous driving, and robot grasping. However, in certain situations where the camera faces transparent or highly reflective objects, the depth information captured by the camera is often not accurate enough, which can lead to subsequent accidents. Therefore, it is necessary to repair and complete the depth images to achieve accurate understanding of the scene’s depth information. In recent years, with the advancement of deep learning, deep learning-based depth image processing and restoration techniques have been widely applied. In this paper, we propose a high-accuracy network for repairing depth images of transparent objects, which can accurately restore and estimate the depth information of transparent objects in various scenarios. Moreover, experimental results demonstrate that our proposed method can generalize well to other unknown scenes, achieving excellent results. Dihua Zhai, Sheng Yu 0009, Yuyin Guan, Yuanqing Xia |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | Prescribed-Time Safety Control for Unknown Systems and Its Application to Robotic ManipulatorabstractIn this paper, the concept of prescribed-time safety is introduced in the case that initial states are not in a safe set, which requires that the trajectories of systems visit a safe set within the prescribed time and then remain in the safe set. In contrast to safety-critical applications that require trajectories to always be in the safe set, prescribed-time safety control presents a greater challenge due to the introduction of a convergence time constraint. This paper presents a method to ensure prescribed-time safety for robotic systems. First, the prescribed-time control barrier function (PTCBF) is proposed for systems with relative degree of one, and it is extended to prescribed-time high order control barrier function (PTHOCBF) for systems with arbitrary relative degrees. Then, quadratic programming (QP) subject to PTCBF condition is constructed to solve control input. However, the proposed method cannot guarantee safety for systems with uncertain models. To address this limitation, a prescribed-time sliding mode disturbance observer (PTSMDO) is proposed to estimate the uncertainty. The estimated error is close to 0 before the time that states stay in the safe set. Based on the observed value of the uncertainty, the control input is solved by QP. Finally, the effectiveness of the proposed method is verified by a simulation and a physical experiment on Franka Emika robot. Note to Practitioners—This paper is motivated by the challenge encountered in robotic systems, which necessitate the control of the system’s state to reach a predefined safe set and subsequently carry out tasks within this safe set in various practical applications, particularly in the realm of robot motion planning. In comparison to existing methods such as finite-time CBF and fixed-time CBF, the proposed prescribed-time CBF approach in this paper offers the advantage of allowing users to preselect the arrival time. Additionally, given that robot systems are typically subject to uncertainties stemming from parameter errors and external disturbances, this paper also introduces a method to the issue of prescribed-time safety for robot systems featuring uncertain models. The uncertain term within the model can be accurately estimated prior to the moment when the state arrives at the safe set, employing the proposed disturbance observer. Subsequently, the prescribed-time CBF is integrated with the estimation results to construct QP for control input. The simulation and experiment conducted on the Franka Emika robot verify that the proposed method is feasible. Sihua Zhang, Dihua Zhai, Yuhan Xiong, Yuanqing Xia |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Focus-TransUnet3D: High-Precision Model for 3D Segmentation of Medical Point TargetsabstractDeep learning has been extensively applied in medical image segmentation, providing significant support for disease diagnosis. However, traditional encoder-decoder networks struggle with segmenting scale-sensitive point target lesions. To address this challenge, this paper proposes an innovative incremental fusion architecture that can integrate different models and achieve significant performance improvements through complementary fusion. Based on this architecture, we developed Focus-TransUnet3D by combining the Trans-FusionNet3D model and the 3D Unet model. This model adopts a global-to-local segmentation strategy, effectively addressing the challenges of medical point target segmentation, thereby expanding the application of deep learning in the field of medical image processing. Furthermore, we design a deep fusion strategy suitable for the transformer model to adapt to multi-scale feature learning. The integration of the transformer model with convolutional neural networks brings improvements in local and global feature extraction capabilities, enhancing the applicability of our model. We evaluate our model on three clinical datasets with different target scales: the Intracranial Artery dataset, the Intracranial Aneurysm dataset, and the LiTS17 dataset. The results indicate that in the external test for intracranial aneurysm auxiliary diagnosis, the model trained with only 47 annotated samples achieved the state-of-the-art performance, attaining a Dice coefficient of 84.14% and a sensitivity of 100%. This effectively addresses the challenges of annotation scarcity and tiny targets. Our code will be released athttps://github.com/caijilia/FTUnet3D. Dihua Zhai, Hao Li 0075, Ke Tian, Yi Yang 0009, Zhenyao Chang, Shuo Wang 0001, Yuanqing Xia |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | A novel distributed scheduling algorithm for maximizing total task allocations of multi-UAV systems
Shaokun Yan, Yuanqing Xia, Dihua Zhai |
J. Supercomput. | 3 |
| 2025 | Cooperative Online Learning for Multiagent System Control via Gaussian Processes With Event-Triggered MechanismabstractIn the realm of the cooperative control of multiagent systems (MASs) with unknown dynamics, Gaussian process (GP) regression is widely used to infer the uncertainties due to its modeling flexibility of nonlinear functions and the existence of a theoretical prediction error bound. Online learning, which involves incorporating newly acquired training data into GP models, promises to improve control performance by enhancing predictions during the operation. Therefore, this article investigates the online cooperative learning algorithm for MAS control. Moreover, an event-triggered data selection mechanism, inspired by the analysis of a centralized event-trigger (CET), is introduced to reduce the model update frequency and enhance the data efficiency. With the proposed learning-based control, the practical convergence of the MAS is validated with guaranteed tracking performance via the Lyapunov theory. Furthermore, the exclusion of the Zeno behavior for individual agents is shown. Finally, the effectiveness of the proposed event-triggered online learning method is demonstrated in simulations. Xiaobing Dai, Zewen Yang, Sihua Zhang, Dihua Zhai, Yuanqing Xia, Sandra Hirche |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Category-Level 6-D Object Pose Estimation With Shape Deformation for Robotic Grasp DetectionabstractCategory-level 6-D object pose estimation plays a crucial role in achieving reliable robotic grasp detection. However, the disparity between synthetic and real datasets hinders the direct transfer of models trained on synthetic data to real-world scenarios, leading to ineffective results. Additionally, creating large-scale real datasets is a time-consuming and labor-intensive task. To overcome these challenges, we propose CatDeform, a novel category-level object pose estimation network trained on synthetic data but capable of delivering good performance on real datasets. In our approach, we introduce a transformer-based fusion module that enables the network to leverage multiple sources of information and enhance prediction accuracy through feature fusion. To ensure proper deformation of the prior point cloud to align with scene objects, we propose a transformer-based attention module that deforms the prior point cloud from both geometric and feature perspectives. Building upon CatDeform, we design a two-branch network for supervised learning, bridging the gap between synthetic and real datasets and achieving high-precision pose estimation in real-world scenes using predominantly synthetic data supplemented with a small amount of real data. To minimize reliance on large-scale real datasets, we train the network in a self-supervised manner by estimating object poses in real scenes based on the synthetic dataset without manual annotation. We conduct training and testing on CAMERA25 and REAL275 datasets, and our experimental results demonstrate that the proposed method outperforms state-of-the-art (SOTA) techniques in both self-supervised and supervised training paradigms. Finally, we apply CatDeform to object pose estimation and robotic grasp experiments in real-world scenarios, showcasing a higher grasp success rate. Sheng Yu 0009, Dihua Zhai, Yuyin Guan, Yuanqing Xia |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | 6-D Object Pose Estimation Based on Point Pair Matching for Robotic Grasp DetectionabstractThe 6-D pose estimation is a critical work essential to achieve reliable robotic grasping. Currently, the prevalent method is reliant on keypoint correspondence. However, this approach hinges on the determination of object keypoint locations, alongside their detection and localization in real scenes. It also employs the random sample consensus (RANSAC)-based perspective-n-point (PnP) algorithm to solve the pose. Yet, it is nondifferentiable and incapable of backpropagation with loss during the training phase. Alternatively, the direct regression method, while speedy and differentiable, falls short in terms of pose estimation performance, and thus needs enhancement. In view of these gaps, we investigate PPM6D, a new method for 6-D object pose estimation based on regression and point pair matching. Our methodology begins with a proposed cross-fusion module, designed to achieve the fusion and complementation of RGB features and point cloud features. Subsequently, an attention module adjusts the features of the object's 3-D model. Finally, we design a point pair matching module for effective matching of points and characteristics, resulting in an integral matching and fusion. PPM6D is extensively trained and tested utilizing benchmark datasets like LINEMOD, occlusion LINEMOD (LINEMOD-occ), YCB-Video, and T-LESS dataset. Experimental results prove that PPM6D can outperform many keypoint-based pose estimation methods, given its relatively rapid speed, thereby offering novel regression-based pose estimation ideas. When applied to real-world scenarios of object pose estimation tasks and grasp tasks of an actual Baxter robot, PPM6D demonstrates superior performance as compared to most alternatives. Sheng Yu 0009, Dihua Zhai, Yufeng Zhan, Wencai Wang, Yuyin Guan, Yuanqing Xia |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | CatFormer: Category-Level 6D Object Pose Estimation with TransformerabstractAlthough there has been significant progress in category-level object pose estimation in recent years, there is still considerable room for improvement. In this paper, we propose a novel transformer-based category-level 6D pose estimation method called CatFormer to enhance the accuracy pose estimation. CatFormer comprises three main parts: a coarse deformation part, a fine deformation part, and a recurrent refinement part. In the coarse and fine deformation sections, we introduce a transformer-based deformation module that performs point cloud deformation and completion in the feature space. Additionally, after each deformation, we incorporate a transformer-based graph module to adjust fused features and establish geometric and topological relationships between points based on these features. Furthermore, we present an end-to-end recurrent refinement module that enables the prior point cloud to deform multiple times according to real scene features. We evaluate CatFormer's performance by training and testing it on CAMERA25 and REAL275 datasets. Experimental results demonstrate that CatFormer surpasses state-of-the-art methods. Moreover, we extend the usage of CatFormer to instance-level object pose estimation on the LINEMOD dataset, as well as object pose estimation in real-world scenarios. The experimental results validate the effectiveness and generalization capabilities of CatFormer. Our code and the supplemental materials are avaliable at https://github.com/BIT-robot-group/CatFormer. Sheng Yu 0009, Dihua Zhai, Yuanqing Xia |
AAAI | 2 |
| 2024 | MITDBA: Mitigating Dynamic Backdoor Attacks in Federated Learning for IoT ApplicationsabstractFederated learning (FL) is widely used in the Internet of Things (IoT) systems. However, FL is susceptible to backdoor attacks due to its inherently distributed and privacy-preserving nature. Existing studies assume that backdoor triggers on different malicious clients are universal, and most defense algorithms are designed to counter backdoor attacks based on this assumption. Recently, dynamic backdoor attacks have been proposed to undermine robust algorithms in centralized machine learning. We introduce dynamic backdoor attacks into the FL system and develop three types of dynamic backdoors named Aggregation, Single, and Continuous to target the FL system. To defend against such attacks, we propose a novel robust algorithm called MITDBA, which utilizes gramian information to capture high-order representations, then employs spectral signatures to detect and remove malicious clients, and finally utilizes clipping operations to filter the selected local models during the aggregation process. We conduct attack and defense experiments on MNIST, CIFAR-10, and GTSRB data sets. The experimental results demonstrate that our designed attack strategies can successfully insert dynamic backdoors into the global model, bypassing the existing state-of-the-art defenses, but these attacks can be effectively mitigated by MITDBA. Dihua Zhai, Dongyu Han, Yuyin Guan, Yuanqing Xia |
IEEE Internet Things J. | 2 |
| 2024 | RoPE: Defending against backdoor attacks in federated learning systems
Dihua Zhai, Yuanqing Xia |
Knowl. Based Syst. | 2 |
| 2024 | FANet: Fast and Accurate Robotic Grasp Detection Based on KeypointsabstractIn practice, the real-time and accuracy of robotic grasp detection are two very important metrics. In the past, researchers had to sacrifice the real-time nature of the detection network in order to obtain higher detection accuracy. How to make the real-time and accuracy of the network co-exist is a problem worth studying. In order to solve this problem, this paper proposes a network, FANet, based on grasp keypoints, which improves the accuracy of grasp detection while ensuring the real-time performance. The key of this paper is how to quickly and accurately detect grasped keypoints. To this end, this paper proposes a local refinement module that optimizes and de-duplicates each feature of the multi-scale feature map, enabling the network to make full use of the multi-scale features. We also propose a global feature refinement module that allows the network to make better use of global features. We also propose a grasp keypoint optimization module that predicts the offset between the actual keypoints and the predicted keypoints, enabling the network to predict the keypoints more accurately. Moreover, we develop two FANets specifically for grasp detection on CPU and GPU, both of which can accomplish real-time grasp detection in real-world scenes. We complete the training and testing of FANet on the Cornell dataset and the Jacquard dataset, achieving SOTA results on the Jacquard dataset. We also test FANet on a dataset of unknown objects, all with good results. Finally, we use the FANet in grasping experiments with an actual Baxter robot and achieve an average grasping success rate of 96%.Note to Practitioners—Real-time and accuracy are two very important metrics in robotic grasping detection. To achieve high accuracy, more time is often consumed for feature extraction. Similarly, in order to improve the real time performance, we need to reduce the time consumed in the feature extraction process, which may result in a drop in detection accuracy. How to coordinate the relationship between them, so as to have both, is a problem worth investigating. Current methods tend to focus on obtaining higher accuracy and they are willing to spend more time to achieve higher accuracy. But in some practical scenarios, such as on factory assembly lines, objects move fast, and the network needs to be able to detect the grasping position quickly, the real-time performance is more important, which makes some methods difficult to use. In addition, most of the current methods tend to focus on GPU-based robotic grasp detection methods, and in real-world scenarios we may not have such a powerful processing GPU available. In contrast, the CPU is an indispensable unit of the computer that we can use to process images without a high-performance GPU. However, compared to GPUs, the CPUs’ image processing capability is poor, making it difficult to achieve real-time processing. Faced with this situation, the problem of how to achieve real-time and high accuracy in a CPU-only robotic grasp detection network is worth studying, but most of the existing methods ignore this problem. To address these problems, we propose a Fast and Accurate robotic grasp detection Network (FANet), which not only enables the network to combine real-time and accuracy, but also enables real-time detection on CPU or GPU. Dihua Zhai, Sheng Yu 0009, Yuanqing Xia |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2024 | Lightweight Multiscale Spatiotemporal Locally Connected Graph Convolutional Networks for Single Human Motion ForecastingabstractHuman motion forecasting is an important and challenging task in many computer vision application domains. Recent work concentrates on utilizing the timing processing ability of recurrent neural networks (RNNs) to achieve smooth and reliable results in short-term prediction. However, as evidenced by previous works, RNNs suffer from error accumulation, leading to unreliable results. In this paper, we propose a simple feed-forward deep neural network for motion prediction, which takes into account temporal smoothness between frames and spatial dependencies between human body joints. We design Lightweight Multiscale Spatiotemporal Locally Connected Graph Convolutional Networks (MST-LCGCN) for Single Human Motion Forecasting to implicitly establish the spatiotemporal dependence in the process of human movement, where different scales fuse dynamically during training. The entire model is action-agnostic and follows a framework of encoder-decoder. The encoder consists of temporal GCNs (TGCNs) to capture motion features between frames and locally connected spatial GCNs (SGCNs) to extract spatial structure among joints. The decoder uses temporal convolution networks (TCNs) to maintain its extensibility for long-term prediction. Considerable experiments show that our approach outperforms previous methods on the Human3.6M and CMU Mocap datasets while only requiring much fewer parameters.Note to Practitioners—Accuracy and real-time performance are the two most significant evaluation factors for the challenge of human motion forecasting. Existing methods tend to use models with a huge amount of parameters, sacrificing operation speed to obtain a small increase in accuracy. However, in practical scenarios, the slowdown in speed makes predictions meaningless. Therefore, we propose a lightweight MST-LCGCN network to learn human action patterns over time. To obtain higher accuracy, we extract features from the spatial and temporal dimensions to contain more information; to obtain faster operation speed, we design our network while reducing unnecessary depth as much as possible. We demonstrate the advantages of our model in terms of efficiency and accuracy through extensive quantitative and qualitative experiments on two datasets. Our network will be helpful for robots to avoid obstacles in advance and compensate for network delays, and we will apply them to real life in the future. Dihua Zhai, Zigeng Yan, Yuanqing Xia |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2024 | ERDUnet: An Efficient Residual Double-Coding Unet for Medical Image SegmentationabstractMedical image segmentation is widely used in clinical diagnosis, and methods based on convolutional neural networks have been able to achieve high accuracy. However, it is still difficult to extract global context features, and the parameters are too large to be clinically applied. In this regard, we propose a novel network structure to improve the traditional encoder-decoder network model, which saves parameters while maintaining segmentation accuracy. We improve the feature extraction efficiency by constructing an encoder module that can simultaneously extract local features and global continuity information. A novel attention module is designed to optimize segmentation boundary regions while improving training efficiency. The feature transfer structure of the decoding part is also improved, which fully integrates the features of different levels to restore the spatial resolution more finely. We evaluate our model on seven different medical segmentation datasets, the 2018 Data Science Bowl Challenge (DSBC2018), the 2018 Lesion Boundary Segmentation Challenge (ISIC2018), the Gland Segmentation in Colon Histology Images Challenge (GlaS), Kvasir-SEG, CVC-ClinicDB, Kvasir-Instrument and Polypgen. Extensive experimental results show that our model can achieve good segmentation performance while maintaining a small number of parameters and computational load, which can further facilitate the generalization of the theoretical approach to clinical practice. Our code will be released athttps://github.com/caijilia/ERDUnet. Hao Li 0075, Dihua Zhai, Yuanqing Xia |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | An Efficient Robotic Pushing and Grasping Method in Cluttered SceneabstractPushing and grasping (PG) are crucial skills for intelligent robots. These skills enable robots to perform complex grasping tasks in various scenarios. These PG methods can be categorized into single-stage and multistage approaches. Single-stage methods are faster but less accurate, while multistage methods offer high accuracy at the expense of time efficiency. To address this issue, a novel end-to-end PG method called efficient PG network (EPGNet) is proposed in this article. EPGNet achieves both high accuracy and efficiency simultaneously. To optimize performance with fewer parameters, EfficientNet-B0 is used as the backbone of EPGNet. Additionally, a novel cross-fusion module is introduced to enhance network performance in robotic PG tasks. This module fuses and utilizes local and global features, aiding the network in handling objects of varying sizes in different scenes. EPGNet consists of two branches dedicated to predicting PG actions, respectively. Both branches are trained simultaneously within a Q-learning framework. Training data is collected through trial and error, involving the robot performing PG actions. To bridge the gap between simulation and reality, a unique PG dataset is proposed. Additionally, a YOLACT network is trained on the PG dataset to facilitate object detection and segmentation. A comprehensive set of experiments is conducted in simulated environments and real-world scenarios. The results demonstrate that EPGNet outperforms single-stage methods and offers competitive performance compared to multistage methods, all while utilizing fewer parameters. A video is available at https://youtu.be/HNKJjQH0MPc. Sheng Yu 0009, Dihua Zhai, Yuanqing Xia, Yuyin Guan |
IEEE Trans. Cybern. | 2 |
| 2024 | PerVK: A Robust Personalized Federated Framework to Defend Against Backdoor Attacks for IoT ApplicationsabstractRobustness and attacks have become prominent concerns in federated learning (FL)-based Internet of Things (IoT). Our focus primarily lies on robustness, as existing robust algorithms are limited by the data distribution and attacker quantity. Personalized FL has emerged as a paradigm to address data heterogeneity, providing personalized local models for participating clients. In this work, we aim to produce personalized models for clients and defend against backdoor attacks on IoT applications by harnessing personalized FL. We proposePerVK, a personalized FL framework that utilizes virtual learning, personalized learning, and knowledge distillation.PerVKeffectively reduces data heterogeneity and overcomes the limitations imposed by the number of malicious clients and data distributions. Empirical experiments are conducted on CIFAR-10 and GTSRB datasets, considering various attack scenarios, as well as compared the performance ofPerVKwith state-of-the-art baselines. The experimental results demonstrate thatPerVKsuccessfully defends against backdoor attacks and outperforms existing baselines. Dihua Zhai, Yuanqing Xia |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | CatTrack: Single-Stage Category-Level 6D Object Pose Tracking via Convolution and Vision TransformerabstractIn the current research, many researchers have focused on instance-level pose tracking, which requires a 3D model of the object in advance, making it challenging to apply in practice. To address this limitation, some researchers have proposed the category-level object pose tracking method. Achieving accurate and speedy monocular category-level pose tracking is an essential research goal. In this article, we propose CatTrack, a new single-stage keypoints-based monocular category-level multi-object pose tracking network. A significant issue in object pose tracking tasks is utilizing the information from the previous frame to guide pose estimation for the next frame. However, as the object poses and camera information in each frame are different, we need to remove irrelevant information and emphasize useful features. To this end, we propose a transformer-based temporal information capture module to leverage the position information of keypoints from the previous frame. Furthermore, we propose a new keypoint matching module to enable the grouping and matching of object keypoints in complex scenes. We have successfully applied CatTrack to the Objectron dataset and achieved superior results in comparison to existing methods. Furthermore, we have also evaluated the generalization of CatTrack and successfully applied it to track the 6D pose of unseen real-world objects. Sheng Yu 0009, Dihua Zhai, Yuanqing Xia |
IEEE Trans. Multim. | 2 |
| 2024 | A Novel Robotic Pushing and Grasping Method Based on Vision Transformer and ConvolutionabstractRobotic grasping techniques have been widely studied in recent years. However, it is always a challenging problem for robots to grasp in cluttered scenes. In this issue, objects are placed close to each other, and there is no space around for the robot to place the gripper, making it difficult to find a suitable grasping position. To solve this problem, this article proposes to use the combination of pushing and grasping (PG) actions to help grasp pose detection and robot grasping. We propose a pushing-grasping combined grasping network (GN), PG method based on transformer and convolution (PGTC). For the pushing action, we propose a vision transformer (ViT)-based object position prediction network pushing transformer network (PTNet), which can well capture the global and temporal features and can better predict the position of objects after pushing. To perform the grasping detection, we propose a cross dense fusion network (CDFNet), which can make full use of the RGB image and depth image, and fuse and refine them several times. Compared with previous networks, CDFNet is able to detect the optimal grasping position more accurately. Finally, we use the network for both simulation and actual UR3 robot grasping experiments and achieve SOTA performance. Video and dataset are available at https://youtu.be/Q58YE-Cc250. Sheng Yu 0009, Dihua Zhai, Yuanqing Xia |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Egret: Reinforcement Mechanism for Sequential Computation Offloading in Edge ComputingabstractAs an emerging computing paradigm, edge computing offers computational resources closer to the data sources, helping to improve the service quality of many real-time applications. A crucial problem is designing a rational pricing mechanism to maximize the revenue of the edge computing service provider (ECSP). However, prior works have considerable limitations: clients are static and are required to disclose their preferences, which is impractical. To address this issue, we propose a novel sequential computation offloading mechanism, where the ECSP posts prices of computational resources with different configurations to clients in turn. Clients independently choose which computational resources to rent and how to offload based on their prices. Then Egret, a deep reinforcement learning-based approach that achieves maximum revenue, is proposed. Egret determines the optimal price and visiting orders online without infringing on clients’ privacy. Experimental results show that the revenue of ECSP in Egret is only 1.29% lower than Oracle and 23.43% better than the state-of-the-art when the client arrives dynamically. Haosong Peng, Yufeng Zhan, Dihua Zhai, Xiaopu Zhang, Yuanqing Xia |
IEEE Trans. Serv. Comput. | 3 |
| 2024 | Robotic Grasp Detection With 6-D Pose Estimation Based on Graph Convolution and RefinementabstractSix-dimensional (6-D) object pose estimation plays a critical role in robotic grasp, which performs extensive usage in manufacturing. The current state-of-the-art pose estimation techniques primarily depend on matching keypoints. Typically, these methods establish a correspondence between 2-D keypoints in an image and the corresponding ones in a 3-D object model. And then they use the PnP-RANSAC algorithm to determine the 6-D pose of the object. However, this approach is not end-to-end trainable and may encounter difficulties when applied to scenarios necessitating differentiable poses. When employing a direct end-to-end regression method, the outcomes are often inferior. To tackle the mentioned problems, we present GR6D, which is a keypoint-and graph-convolution-based neural network for differentiable pose estimation based on RGB-D data. First, we propose a multiscale fusion method that utilizes convolution and graph convolution to exploit information contained in RGB and depth images. Additionally, we propose a transformer-based pose refinement module to further adjust features from RGB images and point clouds. We evaluate GR6D on three datasets: 1) LINEMOD; 2) occlusion LINEMOD; and 3) YCB-Video dataset, and it outperforms most state-of-the-art methods. Finally, we apply GR6D to pose estimation and the robotic grasping task in the real world, manifesting superior performance. Sheng Yu 0009, Dihua Zhai, Yuanqing Xia, Chengyu Zhang 0010 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2023 | SCFL: Mitigating backdoor attacks in federated learning based on SVD and clustering
Dihua Zhai, Yuanqing Xia |
Comput. Secur. | 2 |
| 2023 | An adaptive robust defending algorithm against backdoor attacks in federated learning
Dihua Zhai, Yongping He, Yuanqing Xia |
Future Gener. Comput. Syst. | 2 |
| 2023 | SKGNet: Robotic Grasp Detection With Selective Kernel ConvolutionabstractReal-time and accuracy are important evaluation metrics of robotic grasp detection algorithms. To further improve the accuracy on the premise of ensuring real-time performance, in this paper, a new Selective Kernel convolution Grasp detection Network (SKGNet) is proposed. Compared with previous methods, the attention mechanism and multi-scale fusion features are integrated into the SKGNet, which makes the network not only pay full attention to the grasp area but also flexibly adjust the grasp area according to the scale of the object, thus effectively distinguishing the object from the background. The SKGNet is trained and tested on the Cornell dataset and the Jacquard dataset, with the accuracy of 99.1% and 95.9% respectively, which is superior to SOTA methods. Moreover, SKGNet’s detection speed has reached 28fps. To demonstrate the performance of SKGNet, comparison studies and ablation experiments are performed in this paper. Finally, the grasp experiments of Baxter robot are also performed to verify the generalization of SKGNet in the actual scene, which achieves an average grasping success rate of 96.5%. Video is available athttps://youtu.be/j07sb_ChzWQ. Note to Practitioners—Autonomous grasping is an very important skill for the robotic systems in the real world. However, due to the low grasp detection accuracy, robotic grasp is still a challenging problem. Although some methods have been developed to improve the grasp detection accuracy, the time efficiency is poor. Grasp detection with good accuracy and efficiency is worthy of further study. In this view, a novel deep learning-based grasp detection network SKGNet is proposed in this paper. It takes RGB-D images as input, trains and tests on public datasets, and finally outputs a series of grasp rectangles. Compared with the existing works, it not only achieves state-of-the-art detection accuracy, but also has high efficiency. To demonstrate the generalization performance and effectiveness, the SKGNet is also tested in the real world, and applied to perform the actual grasp task of Baxter robot. The results show that the SKGNet has good robustness and can detect the unknown objects of different sizes and shapes in the real world well. Sheng Yu 0009, Dihua Zhai, Yuanqing Xia |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2023 | A Motion Planning Method for Robots Based on DMPs and Modified Obstacle-Avoiding AlgorithmabstractThis paper addresses the motion planning of the manipulator in task space. To improve the overall trajectory performance, a special motion planning method based on DMPs (Dynamic Movement Primitives) and the modified obstacle-avoiding algorithm is proposed. The proposed method solves the problems of trajectory jitter and inability to avoid obstacles in some scenarios, which are faced by the scheme of steering angle. Besides, it improves the retention of teaching intentions, reduces the loss of free space, and helps the system adapt to the dynamic environment. At the theoretical level, the convergence of the target state has been proved using Lyapunov stability theory-based analysis. The availability of the proposed method is validated and analyzed by performing a series of numerical simulations and Baxter robot experiments. The results indicate that the proposed method can provide reliable solutions for motion planning. Note to Practitioners—From the perspective of demonstration learning, this paper aims to elevate the motion planning of the manipulator in task space, especially in improving the obstacle avoidance performance. The existing DMPs-based motion planning algorithm uses the scheme of steering angle to achieve obstacle avoidance, and the obstacle is regarded as a mesh of points on the boundary. The obstacle-avoiding performance is limited. How to combine the obstacle-avoiding algorithm with the DMPs-based motion planning algorithm more effectively, so as to simultaneously achieve retaining teaching intentions as much as possible, still faces challenges. This paper designs a modified obstacle-avoiding algorithm, which solves the problems of trajectory jitter and inability to avoid obstacles in some circumstances, and improves the retention of teaching intentions. The modified obstacle-avoiding algorithm is also applied to the dynamic environment. The proposed method satisfies Lyapunov stability theory-based analysis, and the experiments on Baxter robot verify the feasibility. Dihua Zhai, Zhiqiang Xia, Haocun Wu, Yuanqing Xia |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2023 | Discrete-Time Control Barrier Function: High-Order Case and Adaptive CaseabstractThis article proposes the novel concepts of the high-order discrete-time control barrier function (CBF) and adaptive discrete-time CBF. The high-order discrete-time CBF is used to guarantee forward invariance of a safe set for discrete-time systems of high relative degree. An optimization problem is then established unifying high-order discrete-time CBFs with discrete-time control Lyapunov functions to yield a safe controller. To improve the feasibility of such optimization problems, the adaptive discrete-time CBF is designed, which can relax constraints on system control input through time-varying penalty functions. The effectiveness of the proposed methods in dealing with high relative degree constraints and improving feasibility is verified on the discrete-time system of a three-link manipulator. Yuhan Xiong, Dihua Zhai, Mahdi Tavakoli, Yuanqing Xia |
IEEE Trans. Cybern. | 2 |
| 2022 | An MVMD-CCA Recognition Algorithm in SSVEP-Based BCI and Its Application in Robot ControlabstractThis article proposes a novel recognition algorithm for the steady-state visual evoked potentials (SSVEP)-based brain-computer interface (BCI) system. By combining the advantages of multivariate variational mode decomposition (MVMD) and canonical correlation analysis (CCA), an MVMD-CCA algorithm is investigated to improve the detection ability of SSVEP electroencephalogram (EEG) signals. In comparison with the classical filter bank canonical correlation analysis (FBCCA), the nonlinear and non-stationary EEG signals are decomposed into a fixed number of sub-bands by MVMD, which can enhance the effect of SSVEP-related sub-bands. The experimental results show that MVMD-CCA can effectively reduce the influence of noise and EEG artifacts and improve the performance of SSVEP-based BCI. The offline experiments show that the average accuracies of MVMD-CCA in the training dataset and testing dataset are improved by 3.08% and 1.67%, respectively. In the SSVEP-based online robotic manipulator grasping experiment, the recognition accuracies of the four subjects are 92.5%, 93.33%, 90.83%, and 91.67%, respectively. Dihua Zhai, Yuhan Xiong, Leyun Hu, Yuanqing Xia |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | StainCNNs: An efficient stain feature learning method
Gaoyi Lei, Yuanqing Xia, Dihua Zhai, Wei Zhang 0169, Duanduan Chen, Defeng Wang |
Neurocomputing | 3 |
| 2020 | Velocity-Observer-Based Distributed Finite-Time Attitude Tracking Control for Multiple Uncertain Rigid SpacecraftabstractThis article addresses the distributed finite-time attitude tracking control problem for a group of uncertain rigid spacecraft in the presence of unavailable angular velocity under the directed topology condition. First, a finite-time adaptive neural network observer is proposed for each follower to estimate its own unavailable angular velocity. Unlike existing velocity-observer-based design methods, the proposed one does not need the exact knowledge of the system model, and works well for the systems with both vanishing and nonvanishing uncertainties. Further, another finite-time observer is provided to obtain the precise angular velocity information of the dynamic leader in a distributed manner. Based on these two observers and adding a power integrator technique, a continuous distributed finite-time control scheme with only attitude measurements is finally established. A rigorous theoretical proof shows that the entire finite-time stability of the combined observer-controller closed-loop system is ensured. Simulation results illustrate the benefits and effectiveness of the developed control scheme. Bing Cui, Yuanqing Xia, Kun Liu 0002, Yujuan Wang 0001, Dihua Zhai |
IEEE Trans. Ind. Informatics | 5 |
| 2018 | A Novel Switching-Based Control Framework for Improved Task Performance in Teleoperation System With Asymmetric Time-Varying DelaysabstractThis paper addresses the adaptive control for task-space teleoperation systems with constrained predefined synchronization error, where a novel switched control framework is investigated. Based on multiple Lyapunov-Krasovskii functionals method, the stability of the resulting closed-loop system is established in the sense of state-independent input-to-output stability. Compared with previous work, the developed method can simultaneously handle the unknown kinematics/dynamics, asymmetric varying time delays, and prescribed performance control in a unified framework. It is shown that the developed controller can guarantee the prescribed transient-state and steady-state synchronization performances between the master and slave robots, which is demonstrated by the simulation study. Dihua Zhai, Yuanqing Xia |
IEEE Trans. Cybern. | 1 |
| 2018 | Multilateral Telecoordinated Control of Multiple Robots With Uncertain KinematicsabstractThis paper addresses the telecoordinated control of multiple robots in the simultaneous presence of asymmetric time-varying delays, nonpassive external forces, and uncertain kinematics/dynamics. To achieve the control objective, a neuroadaptive controller with utilizing prescribed performance control and switching control technique is developed, where the basic idea is to employ the concept of motion synchronization in each pair of master-slave robots and among all slave robots. By using the multiple Lyapunov-Krasovskii functionals method, the state-independent input-to-output practical stability of the closed-loop system is established. Compared with the previous approaches, the new design is straightforward and easier to implement and is applicable to a wider area. Simulation results on three pairs of three degrees-of-freedom robots confirm the theoretical findings. Dihua Zhai, Yuanqing Xia |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Adaptive Control of Semi-Autonomous Teleoperation System With Asymmetric Time-Varying Delays and Input UncertaintiesabstractThis paper addresses the adaptive task-space bilateral teleoperation for heterogeneous master and slave robots to guarantee stability and tracking performance, where a novel semi-autonomous teleoperation framework is developed to ensure the safety and enhance the efficiency of the robot in remote site. The basic idea is to stabilize the tracking error in task space while enhancing the efficiency of complex teleoperation by using redundant slave robot with subtask control. To unify the study of the asymmetric time-varying delays, passive/nonpassive exogenous forces, dynamic parameter uncertainties and dead-zone input in the same framework, a novel switching technique-based adaptive control scheme is investigated, where a special switched error filter is developed. By replacing the derivatives of position errors with their filtered outputs in the coordinate torque design, and employing the multiple Lyapunov-Krasovskii functionals method, the complete closed-loop master (slave) system is proven to be state-independent input-to-output stable. It is shown that both the position tracking errors in task space and the adaptive parameter estimation errors remain bounded for any bounded exogenous forces. Moreover, by using the redundancy of the slave robot, the proposed teleoperation framework can autonomously achieve additional subtasks in the remote environment. Finally, the obtained results are demonstrated by the simulation. Dihua Zhai, Yuanqing Xia |
IEEE Trans. Cybern. | 1 |
| 2017 | Finite-Time Control of Teleoperation Systems With Input Saturation and Varying Time DelaysabstractThis paper develops a finite-time control approach for nonlinear teleoperation systems, which is capable of unifying the study of model uncertainties, actuator saturation, and asymmetric time-varying delays in the same framework. First, a novel anti-windup compensator is designed to analyze the effect of actuator saturation. To achieve the finite-time tracking, a nonsmooth generalized switched filter is also investigated. By introducing the anti-windup compensator and the generalized switched filter into the adaptive fuzzy control torque design, a novel finite-time controller is developed. By using the multiple Lypaunov-Krasovskii functionals method, the resulting closed-loop system is state-independent input-to-output practical stable (SIIOpS) and based on this, it is proved to be finite-time SIIOpS. It is shown that the asymptotic convergence of the adaptive estimation errors and the finite-time convergence of the position tracking errors are obtained, whether the robots contact with the human operator/environment or not. Finally, the effectiveness is demonstrated by the simulation results. Dihua Zhai, Yuanqing Xia |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2016 | Further results on cloud control systems
Yuanqing Xia, Yongming Qin, Dihua Zhai, Senchun Chai |
Sci. China Inf. Sci. | 3 |
| 2016 | On Input-to-State Stability of Switched Stochastic Nonlinear Systems Under Extended Asynchronous SwitchingabstractAn extended asynchronous switching model is investigated for a class of switched stochastic nonlinear retarded systems in the presence of both detection delay and false alarm, where the extended asynchronous switching is described by two independent and exponentially distributed stochastic processes, and further simplified as Markovian. Based on the Razumikhin-type theorem incorporated with average dwell-time approach, the sufficient criteria for global asymptotic stability in probability and stochastic input-to-state stability are given, whose importance and effectiveness are finally verified by numerical examples. Yu Kang 0001, Dihua Zhai, Guo-Ping Liu 0003, Yun-Bo Zhao |
IEEE Trans. Cybern. | 2 |
| 2016 | Neural Network-Based Control of Networked Trilateral Teleoperation With Geometrically Unknown ConstraintsabstractMost studies on bilateral teleoperation assume known system kinematics and only consider dynamical uncertainties. However, many practical applications involve tasks with both kinematics and dynamics uncertainties. In this paper, trilateral teleoperation systems with dual-master-single-slave framework are investigated, where a single robotic manipulator constrained by an unknown geometrical environment is controlled by dual masters. The network delay in the teleoperation system is modeled as Markov chain-based stochastic delay, then asymmetric stochastic time-varying delays, kinematics and dynamics uncertainties are all considered in the force-motion control design. First, a unified dynamical model is introduced by incorporating unknown environmental constraints. Then, by exact identification of constraint Jacobian matrix, adaptive neural network approximation method is employed, and the motion/force synchronization with time delays are achieved without persistency of excitation condition. The neural networks and parameter adaptive mechanism are combined to deal with the system uncertainties and unknown kinematics. It is shown that the system is stable with the strict linear matrix inequality-based controllers. Finally, the extensive simulation experiment studies are provided to demonstrate the performance of the proposed approach. Zhijun Li 0001, Yuanqing Xia, Dehong Wang, Dihua Zhai, Chun-Yi Su, Xingang Zhao |
IEEE Trans. Cybern. | 4 |
| 2016 | Adaptive Fuzzy Control of Multilateral Asymmetric Teleoperation for Coordinated Multiple Mobile ManipulatorsabstractThis paper addresses adaptive fuzzy control for multimaster-multislave teleoperation for multiple mobile manipulators carrying a common object in a cooperative manner that subjected to asymmetric time-varying delays and model parameter uncertainties. In the proposed control framework, a novel switched error filtering is designed. By introducing the filtering output in the control torque design, the complete closed-loop master/slave systems are modeled as a special class of switched system that are composed of two subsystems, i.e., the local master (slave) dynamics with well-defined auxiliary variable and the switched error filter subsystem, which is fairly different from the existing subsystem decomposition method. Utilizing the Lyapunov-Krasovskii method, the complete closed-loop master (slave) system is proved to be state-independent input-to-output stable. The proposed scheme has overcome some application limitations existing in the literature. It is shown that the position tracking errors and the parameter estimation errors can remain bounded under the proposed control laws, which are validated by simulation studies. Dihua Zhai, Yuanqing Xia |
IEEE Trans. Fuzzy Syst. | 1 |