Zhipeng Wang 0006

dblp:56/5818-6 · DBLP profile ↗
← Back
25ranked-venue papers
2as first author
25since 2021 · last 2026
0000-0003-4632-9170ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 11 since 2021Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A survey on robotic manipulation of deformable objects: Recent advances, open challenges and new frontiers
Feida Gu, Zhipeng Wang 0006, Zhongpan Zhu, Yanmin Zhou, Bin He 0003
Neurocomputing2
2026 Offline-Trained GAN-Augmented Highly Adaptive Control With Multi-DoF Fusion for Pneumatic Soft Surgical Robots
abstract
Pneumatic soft robots are well-suited for minimally invasive surgery owing to their compliance and safe interaction with tissues. However, achieving highly adaptive control is difficult owing to modeling inaccuracies, inter-chamber coupling, and disturbances from surgical instruments. Non-learning adaptive methods depend on simplified models and perform poorly in unstructured settings. Conversely, learning-based methods often impose high computational costs in multi-degree-of-freedom (multi-DoF) pneumatic systems. A previous study proposed a generative adversarial network (GAN)-based proportional–integral–derivative (G-PID) controller that combined PID stability with learning-based adaptability by aligning system behavior with a reference model. However, its performance in highly coupled multi-DoF pneumatic soft robots was unverified, and its online adversarial training was computationally intensive. We addressed these limitations by developing an offline-trained G-PID controller, shifting adversarial training offline to reduce computational overhead, achieving 23-fold faster convergence, and enabling real-time, model-free control with balanced adaptability and efficiency. We evaluated three multi-DoF data fusion strategies, showing effective coordination of DoF coupling while maintaining individual control fidelity. Validation on a multi-DoF soft robotic mechatronic system for single-port transvesical prostatectomy revealed tip errors below 0.16 mm across surgical instruments. Proposed controller enhances scalability and adaptability and may generalize to other mechatronic systems with nonlinear, coupled dynamics.
Yuxi Lu 0002, Zhongchao Zhou, Dongliang Zheng, Yanmin Zhou, Zhipeng Wang 0006, Wenwei Yu, Bin He 0003
IEEE Trans Autom. Sci. Eng.5
2026 EDIL: An End-to-End Decoupled Imitation Learning Method for Stable Long-Horizon Bimanual Manipulation
abstract
As a critical capability for automating complex industrial tasks like precision assembly, bimanual fine-grained manipulation has become a key area of research in robotics, where end-to-end imitation learning has emerged as a prominent paradigm. However, in long-horizon cooperative tasks, prevalent end-to-end methods suffer from an inadequate representation of critical task features, particularly those essential for fine-grained coordination between the arms and grippers. This deficiency often destabilizes the manipulation policy, leading to issues like spurious gripper activations and culminating in task failure. To address this challenge, we propose an end-to-end decoupled imitation learning (EDIL) method, which decouples the bimanual manipulation task into a multimodal arm policy for global trajectories and a temporally coordinated attentive gripper policy for fine end-effector actions. The arm policy leverages a Transformer-based encoder–decoder architecture to learn multimodal trajectories from expert demonstrations. The gripper policy leverages cross-attention to facilitate implicit, dynamic feature sharing between the arms, integrating sequential state history and visual data to ensure cooperative stability. We evaluated EDIL on three challenging long-horizon manipulation tasks. Experimental results demonstrate that our method significantly outperforms state-of-the-art approaches, particularly on more complex subtasks, showcasing its robustness and effectiveness.
Zhipeng Wang 0006, Chaoyun Yang, Rong Jiang 0003, Jinyu Zou, Gengdong Zhou, Yanmin Zhou, Bin He 0003
IEEE Trans. Ind. Informatics1
2025 Learning Efficient Robotic Garment Manipulation with Standardization
abstract
Garment manipulation is a significant challenge for robots due to the complex dynamics and potential self-occlusion of garments. Most existing methods of efficient garment unfolding overlook the crucial role of standardization of flattened garments, which could significantly simplify downstream tasks like folding, ironing, and packing. This paper presents APS-Net, a novel approach to garment manipulation that combines unfolding and standardization in a unified framework. APS-Net employs a dual-arm, multi-primitive policy with dynamic fling to quickly unfold crumpled garments and pick-and-place(p&p) for precise alignment. The purpose of garment standardization during unfolding involves not only maximizing surface coverage but also aligning the garment’s shape and orientation to predefined requirements. To guide effective robot learning, we introduce a novel factorized reward function for standardization, which incorporates garment coverage (Cov), keypoint distance (KD), and intersection-over-union (IoU) metrics. Additionally, we introduce a spatial action mask and an Action Optimized Module to improve unfolding efficiency by selecting actions and operation points effectively. In simulation, APS-Net outperforms state-of-the-art methods for long sleeves, achieving 3.9% better coverage, 5.2% higher IoU, and a 0.14 decrease in KD (7.09% relative reduction). Real-world folding tasks further demonstrate that standardization simplifies the folding process. Project page: https://hellohaia.github.io/APS/
Changshi Zhou, Feng Luan, Jiarui Hu 0005, Shaoqiang Meng, Zhipeng Wang 0006, Yanchao Dong, Yanmin Zhou, Bin He 0003
ICML5
2025 Rotation Invariant Spatial Networks for Single-View Point Cloud Classification
abstract
Point cloud classification is critical for three-dimensional scene understanding. However, in real-world scenarios, depth cameras often capture partial, single-view point clouds of objects with different poses, making their accurate classification a challenge. In this paper, we propose a novel point cloud classification network that captures the detailed spatial structure of objects by constructing tetrahedra, which is different from point-wise operations. Specifically, we propose a RISpaNet block to extract rotation-invariant features. A rotation-invariant property generation module is designed in RISpaNet for constructing rotation-invariant tetrahedron properties (RITPs). Meanwhile, a multi-scale pooling module and a hybrid encoder are used to process RITPs to generate integrated rotation-invariant features. Further, for single-view point clouds, a complete point cloud auxiliary branch and a part-whole correlation module are jointly employed to obtain complete point cloud features from partial point clouds. Experimental results show that this network performs better than other state-of-the-art methods, evaluated on four public datasets. We achieved an overall accuracy of 94.7% (+2.0%) on ModelNet40, 93.4% (+5.9%) on MVP, 94.7% (+6.3%) on PCN and 94.8% (+1.7%) on ScanObjectNN. Our project website is https://luxurylf.github.io/RISpaNet_project/.
Feng Luan, Jiarui Hu 0005, Changshi Zhou, Zhipeng Wang 0006, Jiguang Yue, Yanmin Zhou, Bin He 0003
IJCAI4
2025 Uni-Zipper: A Multi-modal Perception Framework of Deformable Objects with Unpaired Data
abstract
Multi-modal perception plays a crucial role in preventing deformation and damage during the robotic manipulation of deformable objects. However, integrating new heterogeneous modalities into existing robotic perception frameworks remains a significant challenge, primarily due to the need for massive amounts of paired data. In this paper, we propose Uni-Zipper, a scalable multi-modal fusion framework designed to expand new modalities with the help of semantic enhancement without relying on paired data. Uni-Zipper consists of a tokenizer that projects various modalities into a shared embedding space, a summary word embedding layer with a feature dictionary, a modality alignment space, and dynamic reconfigurable task heads. To facilitate efficient integration and extension of new modalities, the Zipper alignment mechanism is employed, effectively bridging the modality gap between different input types. Our experimental results demonstrate that Uni-Zipper successfully fuses four modalities and enhances performance in downstream tasks. Despite a 12% decrease in parameter count, Uni-Zipper maintains comparable performance.
Yanmin Zhou, Wei Wang 0515, Yiyang Jin, Zhipeng Wang 0006, Rong Jiang 0003, Xin Li 0082, Hongrui Sang, Bin He 0003
IROS5
2025 Sensing Differently: Unifying Vision, Language, Posture and Tactile in Robotic Perception
abstract
Multi-modal fusion perception enhances robotic performance in complex tasks by providing more comprehensive information than single modality. While tactile and proprioceptive sensing are effective for direct contact tasks like grasping, current research mainly focuses on vision-language fusion, neglecting other embodied modalities. The primary challenges of this limitation are the difficulty in generating natural language labels for embodied information like tactile and proprioception and aligning them with vision and language. To address this, we introduce VLaPT, a novel multi-modal grasping dataset that aligns vision and language (VL) with posture and tactile (PT), enabling robots to sense differently from environment to self. VLaPT includes 75 objects, 1,533 grasps, and over 78K synchronized vision-language-posturetactile pairs. The dataset incorporates structured, rich-text descriptions generated using modality-level language annotation templates, ensuring effective cross-modality alignment. Leveraging this dataset, we trained a lightweight multi-modal alignment framework, CLIP-ME, which enhances the performance of several downstream tasks with only a 5% increase in parameters. The VLaPT is publicly available in https://huggingface.co/datasets/xsdfasfgsa/VLaPT.
Yanmin Zhou, Yiyang Jin, Rong Jiang 0003, Xin Li 0093, Hongrui Sang, Zhipeng Wang 0006, Bin He 0003
IROS7
2025 Robot learning in the era of foundation models: a survey
Xuan Xiao 0002, Zhipeng Wang 0006, Yanmin Zhou, Bin He 0003
Neurocomputing3
2025 Learning Graph Dynamics With Interaction Effects Propagation for Deformable Linear Objects Shape Control
abstract
Robotic manipulation of deformable linear objects (DLOs) has broad application prospects, e.g., manufacturing and medical surgery. To achieve such tasks, a critical challenge is the precise control of the DLOs’ shapes, which requires an accurate dynamics model for deformation prediction. However, due to the infinite dimensionality of the DLOs and the complexity of their deformation mechanism, dynamics models are hard to theoretically calculate. In this paper, for representing the DLO, we use multiple particles being uniformly distributed along the DLO. For learning the dynamics model, we adopt Graph Neural Network (GNN) to learn local interaction effects between neighboring particles, and use the attention mechanism to aggregate the effects of these interactions for the purpose of effect propagation along the DLO (called GA-Net). For manipulation, the Model Predictive Control (MPC) coupled with the learned dynamics model is used to calculate the optimal robot movements, which can also generalize to unseen DLOs. Simulation and real-world experiments demonstrate that GA-Net shows better accuracy than existing methods, and the proposed control framework is effective for different DLOs. Specifically, for model prediction (150 steps), the prediction performance of GA-Net is 14.14% better than the strong baseline (IN-BiLSTM). Videos are available athttps://parkergu.github.io/work_dlo/. Note to Practitioners—This paper was motivated by the problem of shape control of DLOs (e.g., ropes, cables) but it also applies to other deformable objects. Robotic manipulation of DLOs has broad application prospects across various industries, including medical surgeries and manufacturing. Existing approaches to manipulate DLOs, such as reinforcement learning, suffer from sample inefficiency and challenges in generalization. To alleviate these issues, we propose a model-based framework. We adopt GNN and attention mechanism to learn DLOs’ dynamics. Then we use MPC coupled with the learned dynamics model for manipulation of DLOs. The framework is sample-efficient for manipulation, and can generalize to unseen DLOs. Previous works on GNN-based dynamics model do not consider instantaneous propagation of interaction effects, which leads to a false prediction. To alleviate this issue, we adopt GNN to learn interaction effects between neighboring particles, and use the attention mechanism to propagate local interaction effects along the DLO. Simulation and real-world experiments demonstrate that our dynamics model shows better accuracy than existing methods, and also demonstrate the effectiveness of the proposed control framework.
Feida Gu, Hongrui Sang, Yanmin Zhou, Rong Jiang 0003, Zhipeng Wang 0006, Bin He 0003
IEEE Trans Autom. Sci. Eng.6
2025 A Novel Human-in-the-Loop Multimodal Intention Fusion Method for Human-Robot Interaction
abstract
Understanding human intention plays a crucial role in the research of human-robot interaction (HRI) for service robots. Although multimodal sensing methods have shown promising results in certain scenarios, the fusion strategy cannot flexibly adapt to user preference and dynamic environment. To address this limitation, this paper proposed a human-in-the-loop multimodal intention fusion (HIL-MIF) algorithm that introduces weight factors to assess the importance of each modality. By dynamically adjusting these weight factors based on user feedback, we achieved high accuracy in intention understanding, enhancing the personalization and reliability of the interaction. In addition, a multimodal service robot system was proposed that supports three interaction modalities including gaze, voice and gestures, which can be flexibly configured into different combinations. Validation experiment was performed on target object grasping tasks and collected user subjective evaluations. The experimental results demonstrate that the HIL-MIF proposed in this paper is superior to other commonly used multimodal fusion methods in terms of both accuracy and reliability. This paper paves the way for the future tri-co intelligent robot deployment in the service industry.
Zhipeng Wang 0006, Yanmin Zhou, Bin He 0003
IEEE Trans Autom. Sci. Eng.5
2025 Toward Cognitive Digital Twin System of Human-Robot Collaboration Manipulation
abstract
Multielement decision-making is crucial for the robust deployment of human-robot collaboration (HRC) systems in flexible manufacturing environments with personalized tasks and dynamic scenes. Large Language Models (LLMs) have recently demonstrated remarkable reasoning capabilities in various robotic tasks, potentially offering this capability. However, the application of LLMs to actual HRC systems requires the timely and comprehensive capturing of real-scene information. In this study, we suggest incorporating real scene data into LLMs using digital twin (DT) technology and present a cognitive digital twin prototype system of HRC manipulation, known as HRC-CogiDT. Specifically, we initially construct a scene semantic graph encoding the geometric information of entities, spatial relations between entities, actions of humans and robots, and collaborative activities. Subsequently, we devise a prompt that merges scene semantics with prior knowledge of activities, linking the real scene with LLMs. To evaluate performance, we compile an HRC scene understanding dataset and set up a laboratory-level experimental platform. Empirical results indicate that HRC-CogiDT can swiftly perceive scene changes and make high-level decisions based on varying task requirements, such as task planning, anomaly detection, and schedule reasoning. This study provides promising insights for the future applications of LLMs in robotics.Note to Practitioners—Recently, LLMs have demonstrated significant success in various robotic tasks, suggesting their potential as a powerful tool for robotic decision-making. Motivated by this, to improve the production efficiency of HRC in flexible manufacturing, we innovatively combine LLMs with DT technology, and propose a cognitive DT system for HRC, aiming to integrate LLMs into the decision-making loop of HRC system. Experiments conducted in a laboratory-scale platform indicate that the proposed system can handle different decision-making needs in different HRC activities. This system can provide professional guidance to operators in a comprehensible form and serve as a medium for monitoring the safety and standardization of the manipulation process. Future work will explore the use of virtual space provided by the proposed system to optimize the decision outputs of LLMs to make the proposed system more broadly applicable.
Xin Li 0093, Bin He 0003, Zhipeng Wang 0006, Yanmin Zhou, Gang Li 0020, Xiang Li 0010
IEEE Trans Autom. Sci. Eng.3
2025 Long-Sequence Task Planning for Flexible Assembly With Spatio-Temporal Scene Graph
abstract
As the demand for personalized assembly increases, effective assembly task planning continues to present significant challenges, particularly in scenarios with diverse assembly goals and complex dependencies among components. To address these challenges, we propose the Spatio-Temporal Scene Graph-Enhanced Assembly Planning Model (STG-AP), which utilizes only the initial and goal visual observations of the assembly task to infer the assembly sequence from a global perspective, leveraging Graph Neural Network (GNN) and Transformer architectures to capture spatiotemporal dependencies. To evaluate the model’s performance, we conducted assessments on the IKEA ASM Dataset and also developed a lightweight Block-Assembly (Block ASM) Dataset designed to offer a rich variety of scenarios for training and evaluation. Through extensive experiments, we demonstrate that STG-AP consistently outperforms state-of-the-art methods across various sequence lengths. Specifically, our proposed Spatio-Temporal Scene Graph (STG) effectively enhances the performance of long-sequence task planning by skillfully learning the latent representations of the scene. Our method provides a more robust and effective solution for long-sequence flexible assembly planning in robotics.
Zhipeng Wang 0006, Qiao Pan, Yanmin Zhou, Rong Jiang 0003, Bin He 0003, Xin Li 0093
IEEE Trans Autom. Sci. Eng.1
2025 Robotic Motion Optimization for Tactile Object Recognition Learned From Human Behaviors
abstract
Touch is one of the most important human senses. With the development of artificial intelligence, an increasing number of scholars are investigating how robots can be endowed with the sense of touch. Tactile information acquisition relies heavily on action patterns, as touch is generated by direct contact. This paper proposes a robot active tactile perception framework inspired by human behaviors for object recognition. Human recognition experiments were designed to explore the action patterns used to recognize unknown objects. The analysis of these experiments revealed three primary action patterns: pressing, sliding, and rubbing. Robotic motions were conducted to collect tactile data from the same objects using these actions. Given the limited number of available datasets, which restricts the application of active perception learning methods inspired by human behavior patterns, this work aims to address this challenge by developing a novel framework. A tactile convolutional neural network was established for object recognition using the collected data. Additionally, a human-behavior-pattern-inspired tactile feedback control unit was integrated to optimize robotic motions and improve recognition accuracy. The proposed method achieved a recognition accuracy of 97.5%, surpassing the human experiment. This work provides valuable insights into the active tactile perception of robots, offering a reference for future research in this field.
Yanmin Zhou, Ping Lu 0013, Zhipeng Wang 0006, Bin He 0003
IEEE Trans Autom. Sci. Eng.4
2025 T-TD3: A Reinforcement Learning Framework for Stable Grasping of Deformable Objects Using Tactile Prior
abstract
Human tactile perception enables rapid assessment of deformable objects and the application of appropriate force to prevent slip or excessive deformation. However, this task remains challenging for robots. To address this issue, we propose the T-TD3 algorithm, which utilizes a multi-scale fusion neural network (MSF-Net) for the fused perception of multi-scale features, including the tactile prior information obtained through preprocessing. Our approach decomposes the robot task of grasping deformable objects into three subtasks: slip detection, stable grasping evaluation, and minimum grasping force tracking. We develop a simulation environment called CR5GraspStable-Env using PyBullet and TACTO for the network training. Our work reports a success rate of 94.81% in the robot task of grasping deformable objects in real, demonstrating an excellent sim-to-real capability. Moreover, the proposed approach has the potential to be extended to other stable grasping tasks that utilize tactile perception. Note to Practitioners—Traditional grasping strategies in stable grasping tasks typically apply significant grasping force to prevent slip. However, excessive grasping force for deformable objects may lead to excessive deformation and damage. While humans can flexibly control grasping force through tactile perception, it poses a significant challenge for robots. To overcome this challenge, we propose a novel method for the robot to learn an improved stable grasping strategy by incorporating tactile priors. Our method establishes a unified tactile prior representation for the visual-based tactile sensors mounted on the robot grippers, which enables them to sense the contact state. Additionally, it utilizes sensor distortion correction based on spatial symmetry to ensure the applicability of the tactile prior representation to any kind of visual-based tactile sensor. Furthermore, this method integrates multi-scale tactile priors and robot states, and utilizes reinforcement learning to autonomously make real-time decisions, aiming to minimize the grasping force while maintaining stable grasps on deformable objects. The primary research objective of this article is to address the challenge of achieving stable grasps on variable objects, while also being applicable to grasping rigid objects. We validated the practicality of this method by successfully achieving a 94.81% success rate in stably grasping deformable objects using the robotic arm CR5 in both simulated and real-world environments. Furthermore, this method can be readily applied to other robot systems. In the future, we plan to extend the application of our proposed method to more dexterous manipulators and perform more complex manipulation tasks. Additionally, we aim to introduce new fusion algorithms and decision-making strategies to further enhance the applicability of this method.
Yanmin Zhou, Yiyang Jin, Ping Lu 0013, Zhipeng Wang 0006, Bin He 0003
IEEE Trans Autom. Sci. Eng.5
2025 SSFold: Learning to Fold Arbitrary Crumpled Cloth Using Graph Dynamics From Human Demonstration
abstract
Robotic cloth manipulation poses significant challenges due to the fabric’s complex dynamics and the high dimensionality of configuration spaces. Previous approaches have focused on isolated smoothing or folding tasks and relied heavily on simulations, often struggling to bridge the sim-to-real gap. This gap arises as simulated cloth dynamics fail to capture real-world properties such as elasticity, friction, and occlusions, causing accuracy loss and limited generalization. To tackle these challenges, we propose a two-stream architecture with sequential and spatial pathways, unifying smoothing and folding tasks into a single adaptable policy model. The sequential stream determines pick-and-place positions, while the spatial stream, using a connectivity dynamics model, constructs a visibility graph from partial point cloud data, enabling the model to infer the cloth’s full configuration despite occlusions. To address the sim-to-real gap, we integrate real-world human demonstration data via a hand-tracking detection algorithm, enhancing real-world performance across diverse cloth configurations. Our method, validated on a UR5 robot across six distinct cloth folding tasks, consistently achieves desired folded states from arbitrary crumpled initial configurations, with success rates of 100.0%, 100.0%, 83.3%, 66.7%, 83.3%, and 66.7%. It outperforms state-of-the-art cloth manipulation techniques and generalizes to unseen fabrics with diverse colors, shapes, and stiffness. Project page: https://zcswdt.github.io/SSFold/.
Changshi Zhou, Haichuan Xu, Jiarui Hu 0005, Feng Luan, Zhipeng Wang 0006, Yanchao Dong, Yanmin Zhou, Bin He 0003
IEEE Trans Autom. Sci. Eng.5
2025 TrustNet: Deep Ensemble Learning for EEG-Based Trust Recognition in Human-Robot Cooperation
abstract
Recognizing human trust states is crucial for effective human–robot cooperation, as it enables robots to evaluate their decisions and align their actions with human expectations and preferences. However, the complex and dynamic nature of trust in these interactions has led to a lack of effective real-time, generalized trust recognition methods. Electroencephalography (EEG) offers a promising solution by exploring the relationship between trust and specific brain activity patterns. Nonetheless, individual variability in EEG data presents challenges in achieving consistently high recognition performance. In this article, we propose a systematic approach to recognize trust in human–robot cooperation based on EEG data. To address individual differences and achieve robust performance, we introduce TrustNet, a novel stacking ensemble learning model that combines the strengths of several heterogeneous deep learning architectures. Validated on our constructed dataset, EEGTrust, our method achieves an average accuracy of 91.31% in slicewise experiments and 66.11% in trialwise experiments, significantly outperforming baseline methods. Furthermore, we investigate the significance of EEG channels and frequency bands for trust recognition. Results highlight the importance of gamma and beta frequency bands, along with electrodes positioned over frontal and parietal scalp regions, with stronger effects observed at right-hemisphere electrode locations. Feature reduction analysis identified an optimal EEG configuration comprising gamma, beta, and alpha frequency bands from electrodes over nine key scalp regions, specifically excluding middle and left occipital electrode groups.
Caiyue Xu, Yanmin Zhou, Zhipeng Wang 0006, Bin He 0003
IEEE Trans. Comput. Soc. Syst.4
2025 A Novel Bidirectional Controlled Prosthetic Hand With Rigid-Flexible Coupled Structure and Skin Stretch Feedback
abstract
A bionic prosthetic hand is critical for individuals with upper limb disabilities to gain the ability to live independently. However, practical applications are often limited by low adaptability in grasping, complex control, and feedback difficulties. To address these challenges, this work designs a prosthetic hand with a rigid-flexible coupling structure, demonstrating promising adaptability in grasping. It can successfully grasp objects with common shapes, materials, and stiffness. To realize intuitive feedback, the prosthetic hand utilizes a skin-stretching actuator array to simulate the tactile feedback experienced when a normal human hand grasps a heavy object with the palm oriented downward. The fuzzy control encodes the actuator array to provide the user with information such as grasping location, stability, and force level, thus allowing the user to control the motion of the prosthetic hand more intuitively. The array of actuators enables the user to perceive the grasping state of all fingers and lowers the user's threshold of understanding based on fuzzy control, facilitating quicker and easier understanding of the grasping situation. Additionally, for bidirectional control, this prosthetic hand enables neural control using electromyographic (EMG) signals obtained from the proximal surface of the forearm. The EMG sensor's acquisition sites are optimized, and real-time bidirectional demonstration is achieved. This study paves the way for high-performance bionic prosthetics and proposes a novel tactile feedback system with easier understood cascaded spatiotemporal type-2 fuzzy tactile control mapping.
Sicheng Xuan, Funan Zeng, Chuyuan Bian, Zhipeng Wang 0006, Yanmin Zhou, Bin He 0003
IEEE Trans. Fuzzy Syst.4
2025 Transition-Aware Point Cloud Completion by a Progressive Refinement Generative Adversarial Network
abstract
Three-dimensional reconstruction can help robots and vehicles understand their surroundings for subsequent navigation and manipulation tasks. However, in the case of target occlusion, it is difficult for visual sensors to acquire complete information about objects. In this work, we propose a progressive refinement generative adversarial network (PR-GAN) to recover object shapes guided by transition-awareness. This method directly predicts the missing point cloud from the partial point cloud. Our PR-GAN contains a progressive generation module (PGM) and a discriminator. A self-attention-based encoder is proposed in PGM to capture contextual information between local and global features. To guide encoders in generating accurate point clouds, we further propose a progressive fusion module (PFM) that extracts transition information between point clouds of different scales. Moreover, a part-whole correlation module (PWCM) is designed to extract the transition-awareness between the partial and the whole point clouds to further preserve the details. With the above modules, we enhance the spatial logic perception capability of the network so that PR-GAN can fully extract point cloud features and predict the high-fidelity point cloud. Experimental results show that PR-GAN performs better compared to other methods, evaluated on three public datasets. The code is available at https://github.com/luxurylf/PR-GAN.
Feng Luan, Jiarui Hu 0005, Zhipeng Wang 0006, Jiguang Yue, Yanmin Zhou, Bin He 0003
IEEE Trans. Multim.3
2024 X-Tacformer : Spatio-tempral Attention Model for Tactile Recognition
abstract
Recently, tactile sensing has attracted great interests in robotics, especially for exploring unstructured objects. Sensor arrays play an important role in the exploration, which generates rich spatio-temporal information. In this work, we propose an efficient tactile recognition model, X-Tacformer. This model pays attention to both spatial and temporal features of tactile sequences from sensor arrays, which is verified by four public datasets, Ev-Objects, Ev-Containers, Augment8000 and BioTac-Dos. Comparative studies show that our model has resulted in a significant improvement of the recognition accuracy by 0.0223, 0.1416, 0.2735 and 0.1592 in these datasets. In order to verify its performances on dataset with rich spatio-temporal features, a self-designed dataset, ALU-Textures, was constructed with 10 fabrics from everyday textiles, aiming to extend the data collection action modes of current datasets by simulating human rubbing movements with the thumb and index fingers of an Allegro hand. Our model also demonstrates efficient salient feature learning capabilities on ALU-Textures, which is further augmented by tactile data augmentation methods.
Jiarui Hu 0005, Yanmin Zhou, Zhipeng Wang 0006, Xin Li 0093, Yongkang Jiang, Bin He 0003
ICRA3
2024 Trust Recognition in Human-Robot Cooperation Using EEG
abstract
Collaboration between humans and robots is becoming increasingly crucial in our daily life. In order to accomplish efficient cooperation, trust recognition is vital, empowering robots to predict human behaviors and make trust-aware decisions. Consequently, there is an urgent need for a generalized approach to recognize human-robot trust. This study addresses this need by introducing an EEG-based method for trust recognition during human-robot cooperation. A human-robot cooperation game scenario is used to stimulate various human trust levels when working with robots. To enhance recognition performance, the study proposes an EEG Vision Transformer model coupled with a 3-D spatial representation to capture the spatial information of EEG, taking into account the topological relationship among electrodes. To validate this approach, a public EEG-based human trust dataset called EEGTrust is constructed. Experimental results indicate the effectiveness of the proposed approach, achieving an accuracy of 74.99% in slice-wise cross-validation and 62.00% in trial-wise cross-validation. This outperforms baseline models in both recognition accuracy and generalization. Furthermore, an ablation study demonstrates a significant improvement in trust recognition performance of the spatial representation. The source code and EEGTrust dataset are available at https://github.com/CaiyueXu/EEGTrust.
Caiyue Xu, Yanmin Zhou, Zhipeng Wang 0006, Ping Lu 0013, Bin He 0003
ICRA4
2024 A digital twin system for Task-Replanning and Human-Robot control of robot manipulation
Xin Li 0093, Bin He 0003, Zhipeng Wang 0006, Yanmin Zhou, Gang Li 0020, Zhongpan Zhu
Adv. Eng. Informatics3
2024 Geometric-aware RGB-D representation learning for hand-object reconstruction
Yanmin Zhou, Zhipeng Wang 0006, Hongrui Sang, Rong Jiang 0003, Bin He 0003
Expert Syst. Appl.3
2022 A Data Agent Inspired by Interpersonal Interaction Behaviors for Wireless Sensor Networks
abstract
Event monitoring is the main purpose of monitoring applications based on wireless sensor networks (WSNs). To realize the autonomous representation of tunnel disasters in WSNs, this work proposes a data agent (DA) inspired by interpersonal interaction behaviors, which bridges the gap between data and humans. The DA divides event monitoring into four states: 1) active perception; 2) understanding; 3) thinking; and 4) representation, to realize a kind of human-like event monitoring. First, a DA behavior model based on a finite-state machine is proposed, which can adaptively switch among these four states. Second, four methods, including active perception based on a dual neuron perception structure, understanding for forming the information granularity, thinking for forming the disaster knowledge graph, and representation using grayscale and knowledge graphs are proposed to realize the four states of the DA. The experimental results show that the DA inspired by interpersonal interaction can make WSNs not only achieve efficient and accurate disaster self-monitoring but also represent the disaster situation intuitively and quickly through grayscale and knowledge graphs.
Gang Li 0020, Bin He 0003, Zhipeng Wang 0006, Yanmin Zhou
IEEE Internet Things J.3
2022 Blockchain-Enhanced Spatiotemporal Data Aggregation for UAV-Assisted Wireless Sensor Networks
abstract
Wireless sensor networks (WSNs) are widely used in the field of monitoring. For data collection of sensor nodes in large-scale monitoring scenarios, unmanned aerial vehicle (UAV)-assisted WSNs have emerged. For the security and validity of data collection, a blockchain-enhanced data collection framework for UAV-assisted WSNs is presented in this article. To reduce data redundancy in WSNs, a sparsity-optimized and compressed sensing-based spatiotemporal data aggregation model is built. A UAV identity authentication mechanism based on a Merkle tree is also designed to ensure the security of data transmission. By combining blockchain building and data aggregation, a disaster semantic blockchain (DSB) based on a data reconstruction-directed consensus mechanism is presented. Disaster semantics are extracted by analyzing the semantic association relationship of disaster, background, event, and sensor data. The experimental results show that the blockchain-enhanced spatiotemporal data aggregation effectively increases the network life cycle and data reconstruction accuracy. The disaster situation can be described accurately through the DSB.
Gang Li 0020, Bin He 0003, Zhipeng Wang 0006, Jie Chen 0003
IEEE Trans. Ind. Informatics3
2021 A data-efficient goal-directed deep reinforcement learning method for robot visuomotor skill
Rong Jiang 0003, Zhipeng Wang 0006, Bin He 0003, Yanmin Zhou, Gang Li 0020, Zhongpan Zhu
Neurocomputing2