Bin He 0003

dblp:78/4523-3 · DBLP profile ↗
← Back
72ranked-venue papers
3as first author
68since 2021 · last 2026
0000-0003-3193-6269ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 27 since 2021Applied, interdisciplinary, general and emerging computing · 27 · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 11 since 2021Computer networks · 7 · 1 first-author · 5 since 2021Systems, architecture and hardware · 6 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 A survey on robotic manipulation of deformable objects: Recent advances, open challenges and new frontiers
Feida Gu, Zhipeng Wang 0006, Zhongpan Zhu, Yanmin Zhou, Bin He 0003
Neurocomputing7
2026 Developing the robotic space-force boundary of physical interaction perception in an infant way
Yanmin Zhou, Chengjin Wang, Feng Luan, Xin Li 0093, Yongkang Jiang, Bin He 0003
Neurocomputing7
2026 On the Stability of Spatially Distributed Cavity Laser and Boundary of Resonant Beam SLIPT
abstract
Spatially distributed cavity (SDC) lasers are a promising technology for simultaneous light information and power transfer (SLIPT), offering benefits such as increased mobility and intrinsic safety, which are advantageous for various Internet of Things (IoT) devices. However, achieving beam transmission over meter-level long working distances presents significant challenges from cavity stability constraints, manufacturing/ assembly tolerances, and diffraction losses. This paper conducts a theoretical investigation of the fundamental restrictions limiting long-range resonant beam generation. We investigate cavity stability and beam characteristics, and propose a binary-search-based Monte Carlo simulation algorithm as well as a linear approximation algorithm to quantify the maximum acceptable tolerances for stable operation. Numerical results indicate that the stable region contracts sharply as distance increases. For fixed-component systems, an acceptable tolerance of 0.01 mm restricts the achievable transmission distance to less than 2 m. To address this limitation, we also prove the feasibility of long-range beam formation using precision adjustable elements, paving the way for advanced engineering applications. Experimental results verified this assumption, demonstrating that by tuning the stable region during assembly, the transmission distance could be extended to 2.8 m. This work provides essential theoretical insights and practical design guidelines for realizing stable, long-range SDC systems.
Mingliang Xiong, Zeqian Guo, Qingwen Liu 0001, Gang Wang 0014, Gang Li 0020, Bin He 0003
IEEE Internet Things J.7
2026 Self-Aligning Resonant Beam for Simultaneous Wireless Power Transfer and Duplex Communication
abstract
Sustainable energy supply and high-speed communications are two significant needs for the upcoming 6G applications. This paper introduces a self-aligning resonant beam system for simultaneous light information and power transfer (SLIPT), employing a novel coupled spatially distributed resonator (CSDR). The system utilizes a resonant beam for efficient power delivery and a second-harmonic beam for concurrent data transmission, inherently minimizing echo interference and enabling bidirectional communication. Through comprehensive analyses, we investigate the CSDR’s stable region, beam evolution, and power characteristics in relation to working distance and device parameters. Numerical simulations validate the CSDR-SLIPT system’s feasibility by identifying a stable beam waist location for achieving accurate mode-match coupling between two spatially distributed resonant cavities and demonstrating its operational range and efficient power delivery across varying distances. The research reveals the system’s benefits in terms of both safety and energy transmission efficiency. We also demonstrate the trade-off among the reflectivities of the cavity mirrors in the CSDR. Besides, an experiment was conducted to verified the feasibility of self-aligning beam generation and safety under the designed structure. These findings offer valuable design insights for resonant beam systems, advancing SLIPT with significant potential for remote device connectivity.
Mingliang Xiong, Qingwen Liu 0001, Hao Deng 0002, Gang Wang 0014, Jianchen Zhu, Gang Li 0020, Bin He 0003
IEEE J. Sel. Areas Commun.7
2026 Beyond textual rationales: Anatomy-grounded chain-of-thought for traceable radiology reasoning
Jun Yang 0056, Mengyuan Xu, Mingliang Xiong, Wen Fang 0001, Mingqing Liu 0002, Hao Deng 0002, Bin He 0003, Gang Li 0020, Qingwen Liu 0001
Knowl. Based Syst.9
2026 Heatmap Pooling Network for Action Recognition From RGB Videos
abstract
Human action recognition (HAR) in videos has garnered widespread attention due to the rich information in RGB videos. Nevertheless, existing methods for extracting deep features from RGB videos face challenges such as information redundancy, susceptibility to noise and high storage costs. To address these issues and fully harness the useful information in videos, we propose a novel heatmap pooling network (HP-Net) for action recognition from videos, which extracts information-rich, robust and concise pooled features of the human body in videos through a feedback pooling module. The extracted pooled features demonstrate obvious performance advantages over the previously obtained pose data and heatmap features from videos. In addition, we design a spatial-motion co-learning module and a text refinement modulation module to integrate the extracted pooled features with other multimodal data, enabling more robust action recognition. Extensive experiments on several benchmarks namely NTU RGB+D 60, NTU RGB+D 120, Toyota-Smarthome and uncrewed aerial vehicles (UAV)-Human consistently verify the effectiveness of our HP-Net, which outperforms the existing human action recognition methods.
Mengyuan Liu 0001, Yongkang Jiang, Bin He 0003
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Virtual-Real Integration in Unmanned Systems: Emerging Technologies, Applications, and Future Trends
abstract
The integration of virtual and real environments is driving transformative advancements in unmanned systems, offering new avenues for intelligent collaboration and adaptive deployment. This survey presents a comprehensive framework for virtual-real integration centered on unmanned systems, aiming to unify fragmented research efforts and guide future exploration. By adopting both global and local perspectives, the framework facilitates a cohesive understanding of system components, their interactions, and the structural relationships that underpin seamless coordination. This is further complemented by an application case study and a synthesis of the current limitations. To operationalize this framework, the paper systematically examines enabling technologies, existing constraints, and the collaborative architecture that supports dynamic interaction across physical and virtual domains. It further outlines the evolving trends and construction requirements of practical application scenarios. Finally, key challenges and emerging research opportunities are discussed to inform future work and encourage deeper exploration of this rapidly developing field.
Zhongpan Zhu, Shumaila Javaid, Bin Cheng 0008, Wei Li 0211, Bin He 0003
IEEE Trans Autom. Sci. Eng.7
2026 Offline-Trained GAN-Augmented Highly Adaptive Control With Multi-DoF Fusion for Pneumatic Soft Surgical Robots
abstract
Pneumatic soft robots are well-suited for minimally invasive surgery owing to their compliance and safe interaction with tissues. However, achieving highly adaptive control is difficult owing to modeling inaccuracies, inter-chamber coupling, and disturbances from surgical instruments. Non-learning adaptive methods depend on simplified models and perform poorly in unstructured settings. Conversely, learning-based methods often impose high computational costs in multi-degree-of-freedom (multi-DoF) pneumatic systems. A previous study proposed a generative adversarial network (GAN)-based proportional–integral–derivative (G-PID) controller that combined PID stability with learning-based adaptability by aligning system behavior with a reference model. However, its performance in highly coupled multi-DoF pneumatic soft robots was unverified, and its online adversarial training was computationally intensive. We addressed these limitations by developing an offline-trained G-PID controller, shifting adversarial training offline to reduce computational overhead, achieving 23-fold faster convergence, and enabling real-time, model-free control with balanced adaptability and efficiency. We evaluated three multi-DoF data fusion strategies, showing effective coordination of DoF coupling while maintaining individual control fidelity. Validation on a multi-DoF soft robotic mechatronic system for single-port transvesical prostatectomy revealed tip errors below 0.16 mm across surgical instruments. Proposed controller enhances scalability and adaptability and may generalize to other mechatronic systems with nonlinear, coupled dynamics.
Yuxi Lu 0002, Zhongchao Zhou, Dongliang Zheng, Yanmin Zhou, Zhipeng Wang 0006, Wenwei Yu, Bin He 0003
IEEE Trans Autom. Sci. Eng.8
2026 EALLMs: Environment-Aligned LLMs for Enhanced Exploration and Communication in Multi-Agent Reinforcement Learning
abstract
Leveraging large language models (LLMs) for collaborative sequential decision-making is a significant challenge, despite strong semantic understanding and extensive prior knowledge. Conversely, multi-agent reinforcement learning (MARL) can learn environment-aligned policies through interaction, but often suffers from inefficient exploration and heavy reliance on centralized global state information. To achieve complementary advantages, we propose the environment-aligned LLMs (EALLMs). In our framework, an LLM serves as a shared policy for all agents and is updated through online MARL to achieve alignment with the environment. Simultaneously, another LLM, fine-tuned with offline datasets, acts as an information integrator to generate global state for communication purposes. Additionally, we design robust, task-specific prompts tailored to multi-agent systems. Extensive experiments demonstrate that EALLMs outperform classical MARL and LLM-based baselines in both exploration efficiency and overall performance on the SMAC and SMACv2 benchmarks. Ablation studies further confirm EALLMs’ ability to achieve competitive results without relying on explicit global state, while preserving the original capabilities of the LLM during alignment.
Zhuohui Zhang, Bin Cheng 0008, Bin He 0003
IEEE Trans Autom. Sci. Eng.3
2026 Robot Few-Shot Manipulation Skills Learning Based on Meta Imitation Learning and Mixture of Experts Model
Jiahe Zhao, Xiu Su, Bin He 0003
IEEE Trans Autom. Sci. Eng.4
2026 LLM-Guided Adaptive Compensator: Bringing Adaptivity to Robotic Feedback Control With Large Language Model
abstract
Recent advances in code generation and reasoning have enabled the growing use of large language models (LLMs) in robotics. The design of traditional robotic feedback controllers often requires extensive expert knowledge and iterative computational effort, motivating the exploration of LLM-assisted controller development. However, existing efforts are largely restricted to overly simplified systems, provide limited comparison with human-designed controllers, and lack validation on real-world robotic platforms. To address these gaps, this work investigates the use of LLMs in robotic feedback control by focusing on adaptive control and proposing an LLM-guided adaptive compensator. Instead of generating a complete controller from scratch, the LLM is guided to design a compensator that modulates the response of an unknown system to match that of a predefined reference system, thereby achieving adaptivity. The proposed approach is evaluated using five adaptive controllers on a two-degree-of-freedom (DoF) soft robot and a one-DoF shoulder joint of a humanoid robot. Both simulation and real-world experiments demonstrate that the LLM-guided adaptive compensator achieves performance comparable to four conventional adaptive controllers. Lyapunov-based analysis further establishes the stability and generalization capability of the proposed compensator, while analysis of the reasoning process suggests a more structured design approach. This study introduces a novel framework for integrating LLMs into robotic feedback control, with promising practical applicability.
Zhongchao Zhou, Yuxi Lu 0002, Yaonan Zhu, Yifei Zhao 0003, Qian Niu, Bin He 0003, Liang He 0007, Wenwei Yu, Yusuke Iwasawa
IEEE Trans Autom. Sci. Eng.6
2026 Decoding Human Touch Noninvasively: Tactile Inference From EMG and Kinematics Using AET-TacNet
abstract
Deep understanding of human hand dexterity is crucial for making robotic hands more generalizable. While human hand manipulation skills, embedded in hand kinematics and tactile sensing, are typically recorded using instrumented gloves, these gloves can hinder natural hand movement and tactile feedback, potentially limiting the quality of recorded human manipulation data and adversely affecting the human manipulation understanding and the human–robot skill transfer process. We thus propose a novel approach for tactile inference by simultaneously capturing kinematic, electromyography, and tactile information during human manipulation without invasive data gloves. Autoencoder-transformer tactile network, a deep learning framework that leverages modality-specific autoencoders and a Transformer-based model, is introduced to extract compact latent representations from multiple modalities and accurately predict tactile information. We evaluated our approach using a dataset of human manipulation activities, where participants performed various tasks including frontal reaching for objects, pouring, screwing, and feeding, while their kinematics, electromyography, and tactile information were recorded. The proposed approach achieves a normalized root-mean-square error in tactile reconstruction of 0.032, a mean absolute error of 0.015, and a symmetric mean absolute percentage error of 13.4%, significantly outperforming standard baseline methods. These results demonstrated that our noninvasive approach could effectively infer tactile information while preserving natural hand movement and tactile feedback, leading to improved data quality that enhances both the understanding of human motor control and imitation learning for more nuanced and dexterous robotic control.
Huiming Pan, Kezhe Zhu, Dongxuan Li, Yueyuan Chen, Bin He 0003, Peter B. Shull
IEEE Trans. Ind. Informatics6
2026 EDIL: An End-to-End Decoupled Imitation Learning Method for Stable Long-Horizon Bimanual Manipulation
abstract
As a critical capability for automating complex industrial tasks like precision assembly, bimanual fine-grained manipulation has become a key area of research in robotics, where end-to-end imitation learning has emerged as a prominent paradigm. However, in long-horizon cooperative tasks, prevalent end-to-end methods suffer from an inadequate representation of critical task features, particularly those essential for fine-grained coordination between the arms and grippers. This deficiency often destabilizes the manipulation policy, leading to issues like spurious gripper activations and culminating in task failure. To address this challenge, we propose an end-to-end decoupled imitation learning (EDIL) method, which decouples the bimanual manipulation task into a multimodal arm policy for global trajectories and a temporally coordinated attentive gripper policy for fine end-effector actions. The arm policy leverages a Transformer-based encoder–decoder architecture to learn multimodal trajectories from expert demonstrations. The gripper policy leverages cross-attention to facilitate implicit, dynamic feature sharing between the arms, integrating sequential state history and visual data to ensure cooperative stability. We evaluated EDIL on three challenging long-horizon manipulation tasks. Experimental results demonstrate that our method significantly outperforms state-of-the-art approaches, particularly on more complex subtasks, showcasing its robustness and effectiveness.
Zhipeng Wang 0006, Chaoyun Yang, Rong Jiang 0003, Jinyu Zou, Gengdong Zhou, Yanmin Zhou, Bin He 0003
IEEE Trans. Ind. Informatics7
2026 PillarSLAM: A Pillar-Based Structural Semantic SLAM With Novel Relocalization for Autonomous Driving in Underground Parking Lot
abstract
Simultaneous localization and mapping (SLAM) is a fundamental technology of autonomous vehicle, yet its deployment in large-scale, low-texture indoor environments remains highly challenging. The paper introduces PillarSLAM, a multimodal, pillar-based semantic SLAM framework designed to address the limitations of conventional feature-driven approaches. First, a dynamic pillar modeling strategy is proposed to continuously refine pillar dimensions without singularity. Second, a geometry-driven data association method is developed, utilizing distance and pose constraints to uniquely identify individual pillar faces. Third, an object-centric multi-view loop closure is introduced, leveraging pillar positions instead of viewpoint-dependent keypoints to achieve robust long-range corrections. Finally, a topology-based relocalization method is designed, where subgraph of observed pillars are matched to the global map via graph isomorphism, enabling reliable place recognition independent of visual appearance. Experiments in both simulated and real-world environments demonstrate that PillarSLAM significantly outperforms state-of-the-art SLAM systems in mapping accuracy and relocalization robustness under challenging conditions.
Yanchao Dong, Jinfei Ye, Sixiong Xu, Bin He 0003
IEEE Trans. Intell. Transp. Syst.5
2026 Evolutionary Hyper-Transformation for Multi-AAV Path Planning to Visit Moving Targets
abstract
This article addresses a novel path planning problem for multiple fixed-wing autonomous aerial vehicles (AAVs) to visit a set of moving targets, originating from AAV cooperative missions such as emergency communication and target surveillance. This problem can be formulated as a multiple Dubins traveling salesman problem with moving targets (mDTSPMT). The key challenge lies in the strong cross-level coupling between target assignment, encounter sequences, and motion-constrained paths for multiple AAVs in the presence of moving targets. To solve mDTSPMT efficiently, we develop an efficient transformation method by sampling the access location and heading of each AAV to visit moving targets, constructing the mDTSPMT roadmap, and transferring it into an asymmetric multiple traveling salesman problem (AMTSP). This transformation allows the use of mature AMTSP solvers while preserving the essential motion and timing constraints of the original problem. However, the performance of the transformation method heavily depends on the quality of the samples. To improve the quality of samples, a hyper-transformation (HT) framework is proposed, which adaptively optimizes AAV sampling, guiding the search toward more promising configurations and enhancing both the solution quality and computational efficiency of the transformation method. Experiments with extensive instances show that the proposed method outperforms four competitive algorithms in generating coordinated and time-efficient Dubins paths for multiple AAVs encountering multiple targets.
Bin He 0003, Bin Xin 0002, Jie Chen 0003
IEEE Trans. Syst. Man Cybern. Syst.4
2025 Imagine: Image-Guided 3D Part Assembly with Structure Knowledge Graph
abstract
3D part assembly is a promising task in 3D computer vision and robotics, focusing on assembling 3D parts together by predicting their 6-DoF poses. Like most 3D shape understanding tasks, existing methods primarily address this task by memorizing the poses of parts during the training process, leading to inaccuracies in complex assemblies and poor generalization to novel categories. In order to essentially improve the performance, structure knowledge of the target assembly is indispensable before assembling, which abstracts the potential part composition and their structural relationships. An image of the target assembly can serve as a common source for constructing this structure knowledge. Nevertheless, the image is far from enough, as its knowledge can be incomplete and ambiguous due to part occlusion and varying views. To tackle these issues, we propose Imagine, a novel Image-guided 3D part assembly framework with structure knowledge graph. As a novel assembly prior, the structure knowledge graph originates from the image and is refined as understanding the 3D parts. It encodes robust part-aware structural and semantic information of the assembly, guides the 3D parts from a coarse super-structure to a fine assembly, and co-evolves progressively throughout the assembly process. Extensive experiments demonstrate the state-of-the-art performance of our framework, along with strong generalization to novel images and categories.
Mingyu You, Bin He 0003
AAAI4
2025 Align-A-Video: Deterministic Reward Tuning of Image Diffusion Models for Consistent Video Editing
abstract
Due to control limitations in the denoising process and the lack of training, zero-shot video editing methods often struggle to meet user instructions, resulting in generated videos that are visually unappealing and fail to fully satisfy expectations. To address this problem, we propose Align-A-Video, a video editing pipeline that incorporates human feedback through reward fine-tuning. Our approach consists of two key steps: 1) Deterministic Reward Fine-tuning. To reduce optimization costs for expected noise distributions, we propose a deterministic reward tuning strategy. This method improves tuning stability by increasing sample determinism, allowing the tuning process to be completed in minutes; 2) Feature Propagation Across Frames. We optimize a selected anchor frame and propagate its features to the remaining frames, improving both visual quality and semantic fidelity. This approach avoids temporal consistency degradation from reward optimization. Extensive qualitative and quantitative experiments confirm the effectiveness of using reward fine-tuning in Align-A-Video, significantly improving the overall quality of generated videos.
Yingkang Zhong, Jiangchuan Mu, Mingliang Xiong, Wen Fang 0001, Mingqing Liu 0002, Hao Deng 0001, Bin He 0003, Gang Li 0020, Qingwen Liu 0001
CVPR9
2025 Completing 3D Partial Assemblies with View-Consistent 2D-3D Correspondence
Mingyu You, Bin He 0003
ICCV4
2025 Learning Efficient Robotic Garment Manipulation with Standardization
abstract
Garment manipulation is a significant challenge for robots due to the complex dynamics and potential self-occlusion of garments. Most existing methods of efficient garment unfolding overlook the crucial role of standardization of flattened garments, which could significantly simplify downstream tasks like folding, ironing, and packing. This paper presents APS-Net, a novel approach to garment manipulation that combines unfolding and standardization in a unified framework. APS-Net employs a dual-arm, multi-primitive policy with dynamic fling to quickly unfold crumpled garments and pick-and-place(p&p) for precise alignment. The purpose of garment standardization during unfolding involves not only maximizing surface coverage but also aligning the garment’s shape and orientation to predefined requirements. To guide effective robot learning, we introduce a novel factorized reward function for standardization, which incorporates garment coverage (Cov), keypoint distance (KD), and intersection-over-union (IoU) metrics. Additionally, we introduce a spatial action mask and an Action Optimized Module to improve unfolding efficiency by selecting actions and operation points effectively. In simulation, APS-Net outperforms state-of-the-art methods for long sleeves, achieving 3.9% better coverage, 5.2% higher IoU, and a 0.14 decrease in KD (7.09% relative reduction). Real-world folding tasks further demonstrate that standardization simplifies the folding process. Project page: https://hellohaia.github.io/APS/
Changshi Zhou, Feng Luan, Jiarui Hu 0005, Shaoqiang Meng, Zhipeng Wang 0006, Yanchao Dong, Yanmin Zhou, Bin He 0003
ICML8
2025 Rotation Invariant Spatial Networks for Single-View Point Cloud Classification
abstract
Point cloud classification is critical for three-dimensional scene understanding. However, in real-world scenarios, depth cameras often capture partial, single-view point clouds of objects with different poses, making their accurate classification a challenge. In this paper, we propose a novel point cloud classification network that captures the detailed spatial structure of objects by constructing tetrahedra, which is different from point-wise operations. Specifically, we propose a RISpaNet block to extract rotation-invariant features. A rotation-invariant property generation module is designed in RISpaNet for constructing rotation-invariant tetrahedron properties (RITPs). Meanwhile, a multi-scale pooling module and a hybrid encoder are used to process RITPs to generate integrated rotation-invariant features. Further, for single-view point clouds, a complete point cloud auxiliary branch and a part-whole correlation module are jointly employed to obtain complete point cloud features from partial point clouds. Experimental results show that this network performs better than other state-of-the-art methods, evaluated on four public datasets. We achieved an overall accuracy of 94.7% (+2.0%) on ModelNet40, 93.4% (+5.9%) on MVP, 94.7% (+6.3%) on PCN and 94.8% (+1.7%) on ScanObjectNN. Our project website is https://luxurylf.github.io/RISpaNet_project/.
Feng Luan, Jiarui Hu 0005, Changshi Zhou, Zhipeng Wang 0006, Jiguang Yue, Yanmin Zhou, Bin He 0003
IJCAI7
2025 STC-Tracker: Spatiotemporal-Consistent Multi-Robot Collaboration Framework for Long-Term Dynamic Object Tracking
abstract
Multi-robot cooperative tracking, as a vital sub-field of multi-robot collaboration, exhibits significant potential in areas such as military reconnaissance and emergency rescue. Conventional dynamic object tracking methods often face issues of incomplete target detection and even loss in complex scenes, owing to variations in viewpoint or occlusion. To address these problems, this paper proposes STC-Tracker, a multi-robot collaborative tracking system aimed at extending the lifecycle of dynamic objects. On the one hand, the system restores the original appearance of objects by retracing historical point clouds from keyframes while monitoring their motion trajectories in real time. On the other hand, by estimating the motion model of each target, our system is capable of maintaining the lifecycle of specific objects, even in cases of brief disappearance. Experiments are conducted on public and self-collected datasets. The results demonstrate that our algorithm outperforms SOTAs in both single-robot and multi-robot configurations while exhibiting low computational resource consumption. In addition, our algorithm supports LiDARs of different scanning patterns, including spinning LiDARs and solid-state LiDARs, and is capable of real-time dynamic object tracking and global map construction.
Yanchao Dong, Bin He 0003
IROS4
2025 Uni-Zipper: A Multi-modal Perception Framework of Deformable Objects with Unpaired Data
abstract
Multi-modal perception plays a crucial role in preventing deformation and damage during the robotic manipulation of deformable objects. However, integrating new heterogeneous modalities into existing robotic perception frameworks remains a significant challenge, primarily due to the need for massive amounts of paired data. In this paper, we propose Uni-Zipper, a scalable multi-modal fusion framework designed to expand new modalities with the help of semantic enhancement without relying on paired data. Uni-Zipper consists of a tokenizer that projects various modalities into a shared embedding space, a summary word embedding layer with a feature dictionary, a modality alignment space, and dynamic reconfigurable task heads. To facilitate efficient integration and extension of new modalities, the Zipper alignment mechanism is employed, effectively bridging the modality gap between different input types. Our experimental results demonstrate that Uni-Zipper successfully fuses four modalities and enhances performance in downstream tasks. Despite a 12% decrease in parameter count, Uni-Zipper maintains comparable performance.
Yanmin Zhou, Wei Wang 0515, Yiyang Jin, Zhipeng Wang 0006, Rong Jiang 0003, Xin Li 0082, Hongrui Sang, Bin He 0003
IROS9
2025 Sensing Differently: Unifying Vision, Language, Posture and Tactile in Robotic Perception
abstract
Multi-modal fusion perception enhances robotic performance in complex tasks by providing more comprehensive information than single modality. While tactile and proprioceptive sensing are effective for direct contact tasks like grasping, current research mainly focuses on vision-language fusion, neglecting other embodied modalities. The primary challenges of this limitation are the difficulty in generating natural language labels for embodied information like tactile and proprioception and aligning them with vision and language. To address this, we introduce VLaPT, a novel multi-modal grasping dataset that aligns vision and language (VL) with posture and tactile (PT), enabling robots to sense differently from environment to self. VLaPT includes 75 objects, 1,533 grasps, and over 78K synchronized vision-language-posturetactile pairs. The dataset incorporates structured, rich-text descriptions generated using modality-level language annotation templates, ensuring effective cross-modality alignment. Leveraging this dataset, we trained a lightweight multi-modal alignment framework, CLIP-ME, which enhances the performance of several downstream tasks with only a 5% increase in parameters. The VLaPT is publicly available in https://huggingface.co/datasets/xsdfasfgsa/VLaPT.
Yanmin Zhou, Yiyang Jin, Rong Jiang 0003, Xin Li 0093, Hongrui Sang, Zhipeng Wang 0006, Bin He 0003
IROS8
2025 Robot learning in the era of foundation models: a survey
Xuan Xiao 0002, Zhipeng Wang 0006, Yanmin Zhou, Bin He 0003
Neurocomputing7
2025 Development and Application of Coverage Control Algorithms: A Concise Review
abstract
Coverage control is a foundational domain within multi-agent systems, which has recently undergone substantial advancements. Model-based and optimization-based control methods have achieved new breakthroughs and applications, while data-driven and learning-based algorithms in coverage control have also yielded significant results. This review examines the evolution and application of coverage control algorithms in multi-agent systems. Moreover, this review then focuses on how contemporary algorithms build on traditional control strategies and integrate cutting-edge data-driven and adaptive learning techniques. Finally, it explores potential future directions, emphasizing the importance of interdisciplinary approaches in overcoming existing challenges and seizing new opportunities in coverage control deployment. This review aims to be a valuable resource for researchers and practitioners, guiding continued exploration in this field.
Bin Cheng 0008, Mingyuan He, Zhongpan Zhu, Bin He 0003, Jie Chen 0003
IEEE Trans Autom. Sci. Eng.4
2025 Learning Graph Dynamics With Interaction Effects Propagation for Deformable Linear Objects Shape Control
abstract
Robotic manipulation of deformable linear objects (DLOs) has broad application prospects, e.g., manufacturing and medical surgery. To achieve such tasks, a critical challenge is the precise control of the DLOs’ shapes, which requires an accurate dynamics model for deformation prediction. However, due to the infinite dimensionality of the DLOs and the complexity of their deformation mechanism, dynamics models are hard to theoretically calculate. In this paper, for representing the DLO, we use multiple particles being uniformly distributed along the DLO. For learning the dynamics model, we adopt Graph Neural Network (GNN) to learn local interaction effects between neighboring particles, and use the attention mechanism to aggregate the effects of these interactions for the purpose of effect propagation along the DLO (called GA-Net). For manipulation, the Model Predictive Control (MPC) coupled with the learned dynamics model is used to calculate the optimal robot movements, which can also generalize to unseen DLOs. Simulation and real-world experiments demonstrate that GA-Net shows better accuracy than existing methods, and the proposed control framework is effective for different DLOs. Specifically, for model prediction (150 steps), the prediction performance of GA-Net is 14.14% better than the strong baseline (IN-BiLSTM). Videos are available athttps://parkergu.github.io/work_dlo/. Note to Practitioners—This paper was motivated by the problem of shape control of DLOs (e.g., ropes, cables) but it also applies to other deformable objects. Robotic manipulation of DLOs has broad application prospects across various industries, including medical surgeries and manufacturing. Existing approaches to manipulate DLOs, such as reinforcement learning, suffer from sample inefficiency and challenges in generalization. To alleviate these issues, we propose a model-based framework. We adopt GNN and attention mechanism to learn DLOs’ dynamics. Then we use MPC coupled with the learned dynamics model for manipulation of DLOs. The framework is sample-efficient for manipulation, and can generalize to unseen DLOs. Previous works on GNN-based dynamics model do not consider instantaneous propagation of interaction effects, which leads to a false prediction. To alleviate this issue, we adopt GNN to learn interaction effects between neighboring particles, and use the attention mechanism to propagate local interaction effects along the DLO. Simulation and real-world experiments demonstrate that our dynamics model shows better accuracy than existing methods, and also demonstrate the effectiveness of the proposed control framework.
Feida Gu, Hongrui Sang, Yanmin Zhou, Rong Jiang 0003, Zhipeng Wang 0006, Bin He 0003
IEEE Trans Autom. Sci. Eng.7
2025 A Novel Human-in-the-Loop Multimodal Intention Fusion Method for Human-Robot Interaction
abstract
Understanding human intention plays a crucial role in the research of human-robot interaction (HRI) for service robots. Although multimodal sensing methods have shown promising results in certain scenarios, the fusion strategy cannot flexibly adapt to user preference and dynamic environment. To address this limitation, this paper proposed a human-in-the-loop multimodal intention fusion (HIL-MIF) algorithm that introduces weight factors to assess the importance of each modality. By dynamically adjusting these weight factors based on user feedback, we achieved high accuracy in intention understanding, enhancing the personalization and reliability of the interaction. In addition, a multimodal service robot system was proposed that supports three interaction modalities including gaze, voice and gestures, which can be flexibly configured into different combinations. Validation experiment was performed on target object grasping tasks and collected user subjective evaluations. The experimental results demonstrate that the HIL-MIF proposed in this paper is superior to other commonly used multimodal fusion methods in terms of both accuracy and reliability. This paper paves the way for the future tri-co intelligent robot deployment in the service industry.
Zhipeng Wang 0006, Yanmin Zhou, Bin He 0003
IEEE Trans Autom. Sci. Eng.7
2025 Toward Cognitive Digital Twin System of Human-Robot Collaboration Manipulation
abstract
Multielement decision-making is crucial for the robust deployment of human-robot collaboration (HRC) systems in flexible manufacturing environments with personalized tasks and dynamic scenes. Large Language Models (LLMs) have recently demonstrated remarkable reasoning capabilities in various robotic tasks, potentially offering this capability. However, the application of LLMs to actual HRC systems requires the timely and comprehensive capturing of real-scene information. In this study, we suggest incorporating real scene data into LLMs using digital twin (DT) technology and present a cognitive digital twin prototype system of HRC manipulation, known as HRC-CogiDT. Specifically, we initially construct a scene semantic graph encoding the geometric information of entities, spatial relations between entities, actions of humans and robots, and collaborative activities. Subsequently, we devise a prompt that merges scene semantics with prior knowledge of activities, linking the real scene with LLMs. To evaluate performance, we compile an HRC scene understanding dataset and set up a laboratory-level experimental platform. Empirical results indicate that HRC-CogiDT can swiftly perceive scene changes and make high-level decisions based on varying task requirements, such as task planning, anomaly detection, and schedule reasoning. This study provides promising insights for the future applications of LLMs in robotics.Note to Practitioners—Recently, LLMs have demonstrated significant success in various robotic tasks, suggesting their potential as a powerful tool for robotic decision-making. Motivated by this, to improve the production efficiency of HRC in flexible manufacturing, we innovatively combine LLMs with DT technology, and propose a cognitive DT system for HRC, aiming to integrate LLMs into the decision-making loop of HRC system. Experiments conducted in a laboratory-scale platform indicate that the proposed system can handle different decision-making needs in different HRC activities. This system can provide professional guidance to operators in a comprehensible form and serve as a medium for monitoring the safety and standardization of the manipulation process. Future work will explore the use of virtual space provided by the proposed system to optimize the decision outputs of LLMs to make the proposed system more broadly applicable.
Xin Li 0093, Bin He 0003, Zhipeng Wang 0006, Yanmin Zhou, Gang Li 0020, Xiang Li 0010
IEEE Trans Autom. Sci. Eng.2
2025 Contrast, Imitate, Adapt: Learning Robotic Skills From Raw Human Videos
abstract
Learning robotic skills from raw human videos remains a non-trivial challenge. Previous works tackled this problem by leveraging behavior cloning or learning reward functions from videos. Despite their remarkable performances, they may introduce several issues, such as the necessity for robot actions, requirements for consistent viewpoints and similar layouts between human and robot videos, as well as low sample efficiency. To this end, our key insight is to learn task priors by contrasting videos and to learn action priors through imitating trajectories from videos, and to utilize the task priors to guide trajectories to adapt to novel scenarios. We propose a three-stage skill learning framework denoted as Contrast-Imitate-Adapt (CIA). An interaction-aware alignment transformer is proposed to learn task priors by temporally aligning video pairs. Then a trajectory generation model is used to learn action priors. To adapt to novel scenarios different from human videos, the Inversion-Interaction method is designed to initialize coarse trajectories and refine them by limited interaction. In addition, CIA introduces an optimization method based on semantic directions of trajectories for interaction security and sample efficiency. The alignment distances computed by IAAformer are used as the rewards. We evaluate CIA in six real-world everyday tasks, and empirically demonstrate that CIA significantly outperforms previous state-of-the-art works in terms of task success rate and generalization to diverse novel scenarios layouts and object instances.Note to Practitioners—This work aims to study robot skill learning from raw human videos. Compared with teleoperation or kinesthetic teaching in the laboratory, such learning method can flexibly utilize large-scale human videos available on the Internet, thereby improving the robot’s ability to generalize to various complex scenarios. Previous works on learning from videos usually have some issues, including requirements for robot actions, consistent viewpoints, similar layouts and low sample efficiency. To alleviate these issues, we propose a three-stage skill learning framework CIA. Temporal alignment is utilized to learn task priors through our proposed transformer-based model and self-supervised loss functions. A trajectory generation model is trained to learn the action priors. To further adapt to diverse scenarios, we propose a two-stage policy improvement method by initialization and interaction. An optimization method is introduced to ensure safe interaction and sample efficiency, where the optimization objective is guided by the learned task priors. The experimental results show that our CIA outperforms other state-of-the-art methods in task success rate and generalization to novel scenarios.
Zhifeng Qian, Mingyu You, Hongjun Zhou, Xuanhui Xu, Jinzhe Xue, Bin He 0003
IEEE Trans Autom. Sci. Eng.7
2025 Long-Sequence Task Planning for Flexible Assembly With Spatio-Temporal Scene Graph
abstract
As the demand for personalized assembly increases, effective assembly task planning continues to present significant challenges, particularly in scenarios with diverse assembly goals and complex dependencies among components. To address these challenges, we propose the Spatio-Temporal Scene Graph-Enhanced Assembly Planning Model (STG-AP), which utilizes only the initial and goal visual observations of the assembly task to infer the assembly sequence from a global perspective, leveraging Graph Neural Network (GNN) and Transformer architectures to capture spatiotemporal dependencies. To evaluate the model’s performance, we conducted assessments on the IKEA ASM Dataset and also developed a lightweight Block-Assembly (Block ASM) Dataset designed to offer a rich variety of scenarios for training and evaluation. Through extensive experiments, we demonstrate that STG-AP consistently outperforms state-of-the-art methods across various sequence lengths. Specifically, our proposed Spatio-Temporal Scene Graph (STG) effectively enhances the performance of long-sequence task planning by skillfully learning the latent representations of the scene. Our method provides a more robust and effective solution for long-sequence flexible assembly planning in robotics.
Zhipeng Wang 0006, Qiao Pan, Yanmin Zhou, Rong Jiang 0003, Bin He 0003, Xin Li 0093
IEEE Trans Autom. Sci. Eng.6
2025 Robotic Motion Optimization for Tactile Object Recognition Learned From Human Behaviors
abstract
Touch is one of the most important human senses. With the development of artificial intelligence, an increasing number of scholars are investigating how robots can be endowed with the sense of touch. Tactile information acquisition relies heavily on action patterns, as touch is generated by direct contact. This paper proposes a robot active tactile perception framework inspired by human behaviors for object recognition. Human recognition experiments were designed to explore the action patterns used to recognize unknown objects. The analysis of these experiments revealed three primary action patterns: pressing, sliding, and rubbing. Robotic motions were conducted to collect tactile data from the same objects using these actions. Given the limited number of available datasets, which restricts the application of active perception learning methods inspired by human behavior patterns, this work aims to address this challenge by developing a novel framework. A tactile convolutional neural network was established for object recognition using the collected data. Additionally, a human-behavior-pattern-inspired tactile feedback control unit was integrated to optimize robotic motions and improve recognition accuracy. The proposed method achieved a recognition accuracy of 97.5%, surpassing the human experiment. This work provides valuable insights into the active tactile perception of robots, offering a reference for future research in this field.
Yanmin Zhou, Ping Lu 0013, Zhipeng Wang 0006, Bin He 0003
IEEE Trans Autom. Sci. Eng.5
2025 Movement Primitive Categorization Balancing the Learnability and Adaptability
abstract
Given the rapid advancement of robotic technologies, robots will eventually enter our daily lives, performing complex long-horizon tasks. Complex long-horizon tasks, such as furniture assembly, typically contain dozens of subtasks with various scenes. Although complex, assembly is based on several reusable movements. Movement Primitive (MP) is a promising framework for learning reusable movements from demonstrations and adapting the learned movements to the test scenes. The critical step in employing MP methods is categorizing the unlabeled demonstrations into different MPs. However, current MP methods focus on individual MP learning using manually selected demonstrations, neglecting categorization. Manual categorization of demonstrations is easy to fall into suboptimal. If the demonstrations within the same category are too similar, the learned MP cannot be adapted to task scenes with various obstacles. Conversely, a significant distance between demonstrations leads to the MP’s failure in learning. To this end, we propose the following principle for MP categorization: balance the Learnability and Adaptability. Following this principle, we introduce an optimal transportation (OT)-based theoretical framework and a practical solution utilizing an auto-encoder network. We obtain the lower threshold of learnability by OT. Then we increase the adaptability of MP until it reaches the lower threshold of learnability. For complex long-horizon task learning, we propose a balanced MPs-based learning framework that contains four modules, termed BaMPs. BaMPs achieved success rates of 100% and 80%, respectively, in the 12-step and 20-step tasks.
Xuanhui Xu, Mingyu You, Hongjun Zhou, Zhifeng Qian, Jinzhe Xue, Weisheng Xu, Bin He 0003
IEEE Trans Autom. Sci. Eng.7
2025 Autonomous Trajectory Tracking of Bioinspired Soft Microrollers Through a Novel Direction-Speed Decoupled Control Strategy
abstract
Microrollers hold great potential for biomedical applications for its mobility and environmental adaptivity. Autonomous trajectory tracking is crucial for microrollers to conduct medical tasks precisely and safely in complex unstructured physiological environments but remains challenging. In this article, we propose a novel direction-speed decoupled control strategy for autonomous trajectory tracking of microrollers. A modified kinematic model based on extended high-order states is proposed to capture the dynamic characteristics of microrollers, enhancing adaptability to unstructured environments. Decoupled direction and speed controllers adjusting the lateral and longitudinal motions of microrollers are designed based on separate error states to offer improved autonomy and performance in trajectory tracking. A real-time online parameter identification approach is formulated that skips tedious preliminary trials and could generally be used for other variations of microrollers and environments. The trajectory tracking performance of the proposed control strategy is evaluated experimentally on bioinspired magnetic microrollers (BMMs) in biological-like environments. The BMM successfully navigates through cell clusters and reaches the target destination, showing the potential of the proposed control strategy in real physiological environments.Note to Practitioners—This article is motivated by the growing interest in enhancing the autonomy of microrollers in biomedical applications. Controlling the microrollers along an intended trajectory is crucial for applications such as targeted drug delivery and minimally invasive surgery. Traditional control strategies focus solely on spatial path following, disregarding the motion speed of the microrollers. Moreover, kinematic model generally neglects the high-order states for simplification, which cannot facilitate microllers to track rapid changing trajectories in unstructured environments. The identification in real-time and amidst unknown disturbances is also crucial for microrollers to operate efficiently in complex physiological environments. However, existing control strategies rely heavily on preliminary experiments to obtain kinematic parameters for each form of microrollers and environments, which is impractical for real-world applications. To address these challenges, we propose a novel control strategy encompassing an extended state kinematic model, decoupled direction and speed controllers, and an online approach to updating kinematic parameters. Precise trajectory tracking containing direction and speed of BMM is achieved in a biological-like environment consisting of real cell clusters, showing improved performances compared to conventional control methods, and can be adapted to other forms of microrollers without preliminary trials.
Wei Zhang 0383, Bin He 0003, Yin Zhen, Yanmin Zhou
IEEE Trans Autom. Sci. Eng.2
2025 A Biological Structure-Inspired Infrastructure Health Assessment Method
abstract
Structural health monitoring with wireless sensor networks(WSNs) plays an increasingly critical role in modern municipal. While there still remains a gap in getting a comprehensive understanding of complex structural data. Skin diseases can be well diagnosed with modern medical technology, from which similar methods can be learned to improve the intuitiveness and accuracy of structural health monitoring. This work proposes a multi-layered skin-like architecture based method(MSHA), which each layer has its own functions and on the whole presents the disaster situation. This biological structure-inspired architecture has three layers: 1) data substrate layer; 2) connectivity structure layer; 3) pathological manifestation layer, to simulate the three-layer structure of skin. First, a temporal feature extraction method is proposed, which can provide temporal correlations of each node. Second, a dimensional independent spacial feature extraction method is proposed, these first two methods form the connectivity structure layer, which is mainly composed of a spatio-temporal correlation model. Third, a structural health evaluation method for the pathological manifestation layer is proposed to fuse heterogeneous data and calculate multi-granularity structural risk with the features extracted from the connectivity structure layer. The experimental results show that MSHA can achieve not only accurate prediction and intuitive disaster situation, but also low energy consumption and longer lifetime. Note to Practitioners—This paper was motivated by the problem of effectively monitoring the structural health of infrastructure with wireless sensor networks. Existing methods for infrastructure structural health monitoring have space for improvement in maximizing the utilization of correlations between sensor nodes, temporal sequences, and between different sensor types, and fail to produce intuitive results to characterize structural health. In this paper, inspired by the multilayer structure of skin, we propose a new method for analyzing and presenting the structural changes of infrastructure under wireless sensor network data, which fully analyzes the relationship between sensor data in time, space, and type, and is able to predict the structural data in the future period and visually express the current safety situation with size and color information. In this paper, we mathematically describe the model construction method and feature transfer of the constructed intelligent monitoring method. We validate the proposed method using actual tunnel data. The experimental results show that the method is feasible and can accurately indicate the current safety situation in each area while achieving high accuracy in predicting future data, as well as good performance in terms of energy consumption and life cycle. In the future, we will explore optimal strategies for the placement of sensor nodes and improve the convenience and adaptability of the models across multiple scenarios.
Gang Li 0020, Bin He 0003, Bin Cheng 0008
IEEE Trans Autom. Sci. Eng.3
2025 T-TD3: A Reinforcement Learning Framework for Stable Grasping of Deformable Objects Using Tactile Prior
abstract
Human tactile perception enables rapid assessment of deformable objects and the application of appropriate force to prevent slip or excessive deformation. However, this task remains challenging for robots. To address this issue, we propose the T-TD3 algorithm, which utilizes a multi-scale fusion neural network (MSF-Net) for the fused perception of multi-scale features, including the tactile prior information obtained through preprocessing. Our approach decomposes the robot task of grasping deformable objects into three subtasks: slip detection, stable grasping evaluation, and minimum grasping force tracking. We develop a simulation environment called CR5GraspStable-Env using PyBullet and TACTO for the network training. Our work reports a success rate of 94.81% in the robot task of grasping deformable objects in real, demonstrating an excellent sim-to-real capability. Moreover, the proposed approach has the potential to be extended to other stable grasping tasks that utilize tactile perception. Note to Practitioners—Traditional grasping strategies in stable grasping tasks typically apply significant grasping force to prevent slip. However, excessive grasping force for deformable objects may lead to excessive deformation and damage. While humans can flexibly control grasping force through tactile perception, it poses a significant challenge for robots. To overcome this challenge, we propose a novel method for the robot to learn an improved stable grasping strategy by incorporating tactile priors. Our method establishes a unified tactile prior representation for the visual-based tactile sensors mounted on the robot grippers, which enables them to sense the contact state. Additionally, it utilizes sensor distortion correction based on spatial symmetry to ensure the applicability of the tactile prior representation to any kind of visual-based tactile sensor. Furthermore, this method integrates multi-scale tactile priors and robot states, and utilizes reinforcement learning to autonomously make real-time decisions, aiming to minimize the grasping force while maintaining stable grasps on deformable objects. The primary research objective of this article is to address the challenge of achieving stable grasps on variable objects, while also being applicable to grasping rigid objects. We validated the practicality of this method by successfully achieving a 94.81% success rate in stably grasping deformable objects using the robotic arm CR5 in both simulated and real-world environments. Furthermore, this method can be readily applied to other robot systems. In the future, we plan to extend the application of our proposed method to more dexterous manipulators and perform more complex manipulation tasks. Additionally, we aim to introduce new fusion algorithms and decision-making strategies to further enhance the applicability of this method.
Yanmin Zhou, Yiyang Jin, Ping Lu 0013, Zhipeng Wang 0006, Bin He 0003
IEEE Trans Autom. Sci. Eng.6
2025 SSFold: Learning to Fold Arbitrary Crumpled Cloth Using Graph Dynamics From Human Demonstration
abstract
Robotic cloth manipulation poses significant challenges due to the fabric’s complex dynamics and the high dimensionality of configuration spaces. Previous approaches have focused on isolated smoothing or folding tasks and relied heavily on simulations, often struggling to bridge the sim-to-real gap. This gap arises as simulated cloth dynamics fail to capture real-world properties such as elasticity, friction, and occlusions, causing accuracy loss and limited generalization. To tackle these challenges, we propose a two-stream architecture with sequential and spatial pathways, unifying smoothing and folding tasks into a single adaptable policy model. The sequential stream determines pick-and-place positions, while the spatial stream, using a connectivity dynamics model, constructs a visibility graph from partial point cloud data, enabling the model to infer the cloth’s full configuration despite occlusions. To address the sim-to-real gap, we integrate real-world human demonstration data via a hand-tracking detection algorithm, enhancing real-world performance across diverse cloth configurations. Our method, validated on a UR5 robot across six distinct cloth folding tasks, consistently achieves desired folded states from arbitrary crumpled initial configurations, with success rates of 100.0%, 100.0%, 83.3%, 66.7%, 83.3%, and 66.7%. It outperforms state-of-the-art cloth manipulation techniques and generalizes to unseen fabrics with diverse colors, shapes, and stiffness. Project page: https://zcswdt.github.io/SSFold/.
Changshi Zhou, Haichuan Xu, Jiarui Hu 0005, Feng Luan, Zhipeng Wang 0006, Yanchao Dong, Yanmin Zhou, Bin He 0003
IEEE Trans Autom. Sci. Eng.8
2025 TrustNet: Deep Ensemble Learning for EEG-Based Trust Recognition in Human-Robot Cooperation
abstract
Recognizing human trust states is crucial for effective human–robot cooperation, as it enables robots to evaluate their decisions and align their actions with human expectations and preferences. However, the complex and dynamic nature of trust in these interactions has led to a lack of effective real-time, generalized trust recognition methods. Electroencephalography (EEG) offers a promising solution by exploring the relationship between trust and specific brain activity patterns. Nonetheless, individual variability in EEG data presents challenges in achieving consistently high recognition performance. In this article, we propose a systematic approach to recognize trust in human–robot cooperation based on EEG data. To address individual differences and achieve robust performance, we introduce TrustNet, a novel stacking ensemble learning model that combines the strengths of several heterogeneous deep learning architectures. Validated on our constructed dataset, EEGTrust, our method achieves an average accuracy of 91.31% in slicewise experiments and 66.11% in trialwise experiments, significantly outperforming baseline methods. Furthermore, we investigate the significance of EEG channels and frequency bands for trust recognition. Results highlight the importance of gamma and beta frequency bands, along with electrodes positioned over frontal and parietal scalp regions, with stronger effects observed at right-hemisphere electrode locations. Feature reduction analysis identified an optimal EEG configuration comprising gamma, beta, and alpha frequency bands from electrodes over nine key scalp regions, specifically excluding middle and left occipital electrode groups.
Caiyue Xu, Yanmin Zhou, Zhipeng Wang 0006, Bin He 0003
IEEE Trans. Comput. Soc. Syst.5
2025 A Novel Bidirectional Controlled Prosthetic Hand With Rigid-Flexible Coupled Structure and Skin Stretch Feedback
abstract
A bionic prosthetic hand is critical for individuals with upper limb disabilities to gain the ability to live independently. However, practical applications are often limited by low adaptability in grasping, complex control, and feedback difficulties. To address these challenges, this work designs a prosthetic hand with a rigid-flexible coupling structure, demonstrating promising adaptability in grasping. It can successfully grasp objects with common shapes, materials, and stiffness. To realize intuitive feedback, the prosthetic hand utilizes a skin-stretching actuator array to simulate the tactile feedback experienced when a normal human hand grasps a heavy object with the palm oriented downward. The fuzzy control encodes the actuator array to provide the user with information such as grasping location, stability, and force level, thus allowing the user to control the motion of the prosthetic hand more intuitively. The array of actuators enables the user to perceive the grasping state of all fingers and lowers the user's threshold of understanding based on fuzzy control, facilitating quicker and easier understanding of the grasping situation. Additionally, for bidirectional control, this prosthetic hand enables neural control using electromyographic (EMG) signals obtained from the proximal surface of the forearm. The EMG sensor's acquisition sites are optimized, and real-time bidirectional demonstration is achieved. This study paves the way for high-performance bionic prosthetics and proposes a novel tactile feedback system with easier understood cascaded spatiotemporal type-2 fuzzy tactile control mapping.
Sicheng Xuan, Funan Zeng, Chuyuan Bian, Zhipeng Wang 0006, Yanmin Zhou, Bin He 0003
IEEE Trans. Fuzzy Syst.7
2025 Transition-Aware Point Cloud Completion by a Progressive Refinement Generative Adversarial Network
abstract
Three-dimensional reconstruction can help robots and vehicles understand their surroundings for subsequent navigation and manipulation tasks. However, in the case of target occlusion, it is difficult for visual sensors to acquire complete information about objects. In this work, we propose a progressive refinement generative adversarial network (PR-GAN) to recover object shapes guided by transition-awareness. This method directly predicts the missing point cloud from the partial point cloud. Our PR-GAN contains a progressive generation module (PGM) and a discriminator. A self-attention-based encoder is proposed in PGM to capture contextual information between local and global features. To guide encoders in generating accurate point clouds, we further propose a progressive fusion module (PFM) that extracts transition information between point clouds of different scales. Moreover, a part-whole correlation module (PWCM) is designed to extract the transition-awareness between the partial and the whole point clouds to further preserve the details. With the above modules, we enhance the spatial logic perception capability of the network so that PR-GAN can fully extract point cloud features and predict the high-fidelity point cloud. Experimental results show that PR-GAN performs better compared to other methods, evaluated on three public datasets. The code is available at https://github.com/luxurylf/PR-GAN.
Feng Luan, Jiarui Hu 0005, Zhipeng Wang 0006, Jiguang Yue, Yanmin Zhou, Bin He 0003
IEEE Trans. Multim.6
2025 Deep Learning for Low-Light Vision: A Comprehensive Survey
abstract
Visual recognition in low-light environments is a challenging problem since degraded images are the stacking of multiple degradations (noise, low light and blur, etc.). It has received extensive attention from academia and industry in the era of deep learning. Existing surveys focus on low-light image enhancement (LLIE) methods and normal-light visual recognition methods, while few comprehensive surveys of low-light-related vision tasks. This article provides a comprehensive survey of the latest advancements in low-light vision, including methods, datasets, and evaluation metrics, in two aspects: visual quality-driven and recognition quality-driven. On the visual quality-driven aspect, we survey a large number of very recent LLIE methods. On the recognition quality-driven aspect, we survey low-light object detection techniques in the deep learning era using more intuitive categorization method. Furthermore, a quantitative benchmarking of different methods is conducted on several widely adopted low-light vision-related datasets. Finally, we discuss the challenges that exist in low-light vision and future directions worth exploring. We provide a public website that will continue to track developments in this promising field.
Qian Zhao 0014, Gang Li 0020, Bin He 0003, Runjie Shen
IEEE Trans. Neural Networks Learn. Syst.3
2025 PhysFiT: Physical-aware 3D Shape Understanding for Finishing Incomplete Assembly
abstract
Understanding the part composition and structure of 3D shapes is crucial for a wide range of 3D applications, including 3D part assembly and 3D assembly completion. Compared to 3D part assembly, 3D assembly completion is more complicated, which involves repairing broken or incomplete furniture that miss several parts with a toolkit. Given an incomplete assembly, 3D assembly completion seeks to identify its missing parts from multiple candidates, determine their poses, and produce complete assembly that is well-connected, structurally stable, and aesthetically pleasing. This task necessitates not only specialized knowledge of part composition but, more importantly, an awareness of physical constraints, i.e., connectivity, stability, and symmetry. Neglecting these constraints often results in assemblies that, although visually plausible, are impractical. To address this challenge, we propose PhysFiT, a physical-aware 3D shape understanding framework. This framework is built upon attention-based part relation modeling and incorporates connection modeling, simulation-free stability optimization and symmetric transformation consistency. We evaluate its efficacy on 3D part assembly and 3D assembly completion, a novel assembly task presented in this work. Extensive experiments demonstrate the effectiveness of PhysFiT in constructing geometrically sound and physically compliant assemblies.
Mingyu You, Hongjun Zhou, Bin He 0003
ACM Trans. Graph.4
2025 Resonant Beam Enabled Multi-Target Localization
abstract
In the era of the Internet of everything (IoE) and the metaverse, there is a growing demand for high-accuracy indoor positioning for applications such as autonomous robots, virtual reality, and smartphones. This paper proposed a resonant beam phase-based passive localization (RBPPL) system optimized for high-precision indoor positioning in multi-access scenarios. By leveraging the self-alignment characteristic and integrating the analysis of resonant beam phase, angle of arrival (AoA) matching and binocular disparity method for 3D point coordinate acquisition, the RBPPL system achieves binocular passive multi-access 3D positioning with an error within 4 cm at a distance of 8 m. We present a novel multi-access AoA estimation method that overcomes the challenges of spot overlap in traditional CMOS-based angle analysis. We propose a telescope system to correct the phase and focus the propagation direction of optical resonant beam systems. Simulations demonstrate the system’s robustness and high accuracy. The proposed RBPPL system, optimized for multi-access scenarios, offers a promising solution for high-accuracy indoor positioning, supporting various IoE and metaverse applications. Future work will focus on real-world deployment and its potential in complex multi-access scenarios.
Guangkun Zhang, Mengyuan Xu, Wen Fang 0001, Mingliang Xiong, Mingqing Liu 0002, Gang Li 0020, Bin He 0003, Qingwen Liu 0001
IEEE Trans. Wirel. Commun.9
2024 X-Tacformer : Spatio-tempral Attention Model for Tactile Recognition
abstract
Recently, tactile sensing has attracted great interests in robotics, especially for exploring unstructured objects. Sensor arrays play an important role in the exploration, which generates rich spatio-temporal information. In this work, we propose an efficient tactile recognition model, X-Tacformer. This model pays attention to both spatial and temporal features of tactile sequences from sensor arrays, which is verified by four public datasets, Ev-Objects, Ev-Containers, Augment8000 and BioTac-Dos. Comparative studies show that our model has resulted in a significant improvement of the recognition accuracy by 0.0223, 0.1416, 0.2735 and 0.1592 in these datasets. In order to verify its performances on dataset with rich spatio-temporal features, a self-designed dataset, ALU-Textures, was constructed with 10 fabrics from everyday textiles, aiming to extend the data collection action modes of current datasets by simulating human rubbing movements with the thumb and index fingers of an Allegro hand. Our model also demonstrates efficient salient feature learning capabilities on ALU-Textures, which is further augmented by tactile data augmentation methods.
Jiarui Hu 0005, Yanmin Zhou, Zhipeng Wang 0006, Xin Li 0093, Yongkang Jiang, Bin He 0003
ICRA6
2024 Trust Recognition in Human-Robot Cooperation Using EEG
abstract
Collaboration between humans and robots is becoming increasingly crucial in our daily life. In order to accomplish efficient cooperation, trust recognition is vital, empowering robots to predict human behaviors and make trust-aware decisions. Consequently, there is an urgent need for a generalized approach to recognize human-robot trust. This study addresses this need by introducing an EEG-based method for trust recognition during human-robot cooperation. A human-robot cooperation game scenario is used to stimulate various human trust levels when working with robots. To enhance recognition performance, the study proposes an EEG Vision Transformer model coupled with a 3-D spatial representation to capture the spatial information of EEG, taking into account the topological relationship among electrodes. To validate this approach, a public EEG-based human trust dataset called EEGTrust is constructed. Experimental results indicate the effectiveness of the proposed approach, achieving an accuracy of 74.99% in slice-wise cross-validation and 62.00% in trial-wise cross-validation. This outperforms baseline models in both recognition accuracy and generalization. Furthermore, an ablation study demonstrates a significant improvement in trust recognition performance of the spatial representation. The source code and EEGTrust dataset are available at https://github.com/CaiyueXu/EEGTrust.
Caiyue Xu, Yanmin Zhou, Zhipeng Wang 0006, Ping Lu 0013, Bin He 0003
ICRA6
2024 A digital twin system for Task-Replanning and Human-Robot control of robot manipulation
Xin Li 0093, Bin He 0003, Zhipeng Wang 0006, Yanmin Zhou, Gang Li 0020, Zhongpan Zhu
Adv. Eng. Informatics2
2024 Designing novel adaptive dynamic event-triggered protocols for uncertain multi-agent systems
Bin Cheng 0008, Bin He 0003
Sci. China Inf. Sci.2
2024 Geometric-aware RGB-D representation learning for hand-object reconstruction
Yanmin Zhou, Zhipeng Wang 0006, Hongrui Sang, Rong Jiang 0003, Bin He 0003
Expert Syst. Appl.6
2024 Graph-based geometric structure line parsing
Feng Li 0046, Gang Li 0020, Bin He 0003, Ping Lu 0013, Bin Cheng 0008
Neurocomputing3
2024 Laser Ranger-Based Baseline Measurement for Collaborative Localization
abstract
To address challenges in outdoor multi-robot collaborative localization (MRCL) due to low GPS accuracy, we propose a system using three UGVs, each equipped with a shared camera and a laser rangefinder. Our trilateral localization algorithm combines least-squares matrix and gradient descent methods, resulting in an 80.4% improvement in accuracy compared to traditional methods. The system mitigates the limitations of GPS accuracy by utilizing accurate baseline measurements and optimizing the localization process. These advancements have potential applications in transportation, production, and logistics, enhancing MRCL performance in outdoor environments.
Mingqing Liu 0002, Yihan Zhu, Qingwen Liu 0001, Qunhui Yang, Gang Li 0020, Bin He 0003
IEEE Internet Things J.8
2024 A robust visual SLAM system for low-texture and semi-static environments
Bin He 0003, Sixiong Xu, Yanchao Dong, Senbo Wang, Jiguang Yue, Lingling Ji
Multim. Tools Appl.1
2024 Mean policy-based proximal policy optimization for maneuvering decision in multi-UAV air combat
Bin Xin 0002, Bin He 0003
Neural Comput. Appl.3
2024 Transductive Learning With Prior Knowledge for Generalized Zero-Shot Action Recognition
abstract
It is challenging to achieve generalized zero-shot action recognition. Different from the conventional zero-shot tasks which assume that the instances of the source classes are absent in the test set, the generalized zero-shot task studies the case that the test set contains both the source and the target classes. Due to the gap between visual feature and semantic embedding as well as the inherent bias of the learned classifier towards the source classes, the existing generalized zero-shot action recognition approaches are still far less effective than traditional zero-shot action recognition approaches. Facing these challenges, a novel transductive learning with prior knowledge (TLPK) model is proposed for generalized zero-shot action recognition. First, TLPK learns the prior knowledge which assists in bridging the gap between visual features and semantic embeddings, and preliminarily reduces the bias caused by the visual-semantic gap. Then, a transductive learning method that employs unlabeled target data is designed to overcome the bias problem in an effective manner. To achieve this, a target semantic-available approach and a target semantic-free approach are devised to utilize the target semantics in two different ways, where the target semantic-free approach exploits prior knowledge to produce well-performed semantic embeddings. By exploring the usage of the aforementioned prior-knowledge learning and transductive learning strategies, TLPK significantly bridges the visual-semantic gap and alleviates the bias between the source and the target classes. The experiments on the benchmark datasets of HMDB51 and UCF101 demonstrate the effectiveness of the proposed model compared to the state-of-the-art methods. The source code of this work can be found inhttps://mic.tongji.edu.cn
Taiyi Su, Hanli Wang, Qiuping Qi, Bin He 0003
IEEE Trans. Circuits Syst. Video Technol.5
2024 Adaptive Approximation Tracking Control of a Continuum Robot With Uncertainty Disturbances
abstract
Continuum robot has certain compliance and intrinsic safety, which makes it an excellent substitute for the traditional rigid robot in the various tasks, such as human-robot interaction or medical surgery. However, due to the complex nonlinearities which are induced by the compliance, parameter uncertainties, and unknown disturbances, good performance control scheme for the continuum robot is always a challenging task for the practical applications. Thus, to overcome this challenge of the uncertain dynamics and unknown external disturbances, this article develops a novel adaptive control scheme for a continuum robot using the function approximation technique (FAT). Specifically, for the proposed continuum robot, an adaptive FAT control (AFATC) strategy with no update laws is proposed to handle the uncertain parameters of the robot dynamics and external disturbances. The control law is expressed as a finite linear combination of the orthogonal basis functions by the FAT. The proposed AFATC scheme uses a fixed control structure, and the weight matrices are not updated in time. Then, the stability of the proposed controller is proved based on the Lyapunov function. Afterwards, the simulation results indicate the proposed AFATC scheme has good control performance compared with the regressor-free adaptive control (RFAC) method. Finally, the effectiveness of the proposed AFATC scheme is demonstrated in a group of real-time experiments.
Shoulin Xu, Bin He 0003
IEEE Trans. Cybern.2
2024 Multi-Modal Structure-Embedding Graph Transformer for Visual Commonsense Reasoning
abstract
Visual commonsense reasoning (VCR) is a challenging reasoning task that aims to not only answer the question based on a given image but also provide a rationale justifying for the choice. Graph-based networks are appropriate to represent and extract the correlation between image and language for reasoning, where how to construct and learn graphs based on such multi-modal Euclidean data is a fundamental problem. Most existing graph-based methods view visual regions and linguistic words as identical graph nodes, ignoring inherent characteristics of multi-modal data. In addition, these approaches typically only have one graph-learning layer, and the performance declines as the model goes deeper. To address these issues, a novel method named Multi-modal Structure-embedding Graph Transformer (MSGT) is proposed. Specifically, an answer-vision graph and an answer-question graph are constructed to represent and model intra-modal and inter-modal correlations in VCR simultaneously, where additional multi-modal structure representations are initialized and embedded according to visual region distances and linguistic word orders for more reasonable graph representation. Then, a structure-injecting graph transformer is designed to inject embedded structure priors into the semantic correlation matrix for the evolution of node features and structure representations, which can stack more layers to make model deeper and extract more powerful features with instructive priors. To adaptively fuse graph features, a scored pooling mechanism is further developed to select valuable clues for reasoning from learnt node features. Experiments demonstrate the superiority of the proposed MSGT framework compared with state-of-the-art methods on the VCR benchmark dataset.
Jian Zhu 0006, Hanli Wang, Bin He 0003
IEEE Trans. Multim.3
2024 Monotonic Quantile Network for Worst-Case Offline Reinforcement Learning
abstract
A key challenge in offline reinforcement learning (RL) is how to ensure the learned offline policy is safe, especially in safety-critical domains. In this article, we focus on learning a distributional value function in offline RL and optimizing a worst-case criterion of returns. However, optimizing a distributional value function in offline RL can be hard, since the crossing quantile issue is serious, and the distribution shift problem needs to be addressed. To this end, we propose monotonic quantile network (MQN) with conservative quantile regression (CQR) for risk-averse policy learning. First, we propose an MQN to learn the distribution over returns with non-crossing guarantees of the quantiles. Then, we perform CQR by penalizing the quantile estimation for out-of-distribution (OOD) actions to address the distribution shift in offline RL. Finally, we learn a worst-case policy by optimizing the conditional value-at-risk (CVaR) of the distributional value function. Furthermore, we provide theoretical analysis of the fixed-point convergence in our method. We conduct experiments in both risk-neutral and risk-sensitive offline settings, and the results show that our method obtains safe and conservative behaviors in robotic locomotion tasks.
Chenjia Bai, Ting Xiao 0002, Zhoufan Zhu, Lingxiao Wang 0003, Animesh Garg, Bin He 0003, Peng Liu 0008, Zhaoran Wang 0001
IEEE Trans. Neural Networks Learn. Syst.7
2023 3D Assembly Completion
abstract
Automatic assembly is a promising research topic in 3D computer vision and robotics. Existing works focus on generating assembly (e.g., IKEA furniture) from scratch with a set of parts, namely 3D part assembly. In practice, there are higher demands for the robot to take over and finish an incomplete assembly (e.g., a half-assembled IKEA furniture) with an off-the-shelf toolkit, especially in human-robot and multi-agent collaborations. Compared to 3D part assembly, it is more complicated in nature and remains unexplored yet. The robot must understand the incomplete structure, infer what parts are missing, single out the correct parts from the toolkit and finally, assemble them with appropriate poses to finish the incomplete assembly. Geometrically similar parts in the toolkit can interfere, and this problem will be exacerbated with more missing parts. To tackle this issue, we propose a novel task called 3D assembly completion. Given an incomplete assembly, it aims to find its missing parts from a toolkit and predict the 6-DoF poses to make the assembly complete. To this end, we propose FiT, a framework for Finishing the incomplete 3D assembly with Transformer. We employ the encoder to model the incomplete assembly into memories. Candidate parts interact with memories in a memory-query paradigm for final candidate classification and pose prediction. Bipartite part matching and symmetric transformation consistency are embedded to refine the completion. For reasonable evaluation and further reference, we design two standard toolkits of different difficulty, containing different compositions of candidate parts. We conduct extensive comparisons with several baseline methods and ablation studies, demonstrating the effectiveness of the proposed method.
Rufeng Zhang, Mingyu You, Hongjun Zhou, Bin He 0003
AAAI5
2023 HBOD: A Novel Dataset with Synchronized Hand, Body, and Object Manipulation Data for Human-Robot Interaction
abstract
Estimating hand and body posture is crucial for enabling human-robot collaboration, preventing occupational diseases, and training humanoid robots. Although advances in wearable motion sensors, such as Inertial Measurement Units (IMUs), have resulted in public datasets in industrial and occupational settings, these datasets rarely include movements with subjects holding and manipulating objects. However, it is crucial to have data on how humans move and interact with different objects so that we can better understand human motion intention and movement strategies in specific scenarios. We thus propose the HBOD dataset (hand-body-object dataset), which encompasses synchronized human pose data from an IMU sensor network, hand posture data from a smart data glove, and object position and attitude information obtained from a motion capture system, while subjects move and interact with a screwdriver, hammer, spanner, electric drill, and a rectangular workpiece. This paper provides an overview of the hardware setup, experimental protocol, data format, and data visualization results. This dataset provides crucial object information absent from existing datasets, thus offering highly valuable manipulation data for occupational diseases research, human-robot interaction, and robot skill acquisition.
Peiqi Kang, Kezhe Zhu, Bin He 0003, Peter B. Shull
BSN4
2023 Zero-Shot Object Goal Visual Navigation
abstract
Object goal visual navigation is a challenging task that aims to guide a robot to find the target object based on its visual observation, and the target is limited to the classes pre-defined in the training stage. However, in real households, there may exist numerous target classes that the robot needs to deal with, and it is hard for all of these classes to be contained in the training stage. To address this challenge, we study the zero-shot object goal visual navigation task, which aims at guiding robots to find targets belonging to novel classes without any training samples. To this end, we also propose a novel zero-shot object navigation framework called semantic similarity network (SSNet). Our framework use the detection results and the cosine similarity between semantic word embeddings as input. Such type of input data has a weak correlation with classes and thus our framework has the ability to generalize the policy to novel classes. Extensive experiments on the AI2-THOR platform show that our model outperforms the baseline models in the zero-shot object navigation task, which proves the generalization ability of our model. Our code is available at: https://github.com/pioneer-innovation/Zero-Shot-Object-Navigation.
Qianfan Zhao, Lu Zhang 0054, Bin He 0003, Hong Qiao, Zhiyong Liu 0001
ICRA3
2023 Knowledge-Enriched Attention Network With Group-Wise Semantic for Visual Storytelling
abstract
As a technically challenging topic, visual storytelling aims at generating an imaginary and coherent story with narrative multi-sentences from a group of relevant images. Existing methods often generate direct and rigid descriptions of apparent image-based contents, because they are not capable of exploring implicit information beyond images. Hence, these schemes could not capture consistent dependencies from holistic representation, impairing the generation of reasonable and fluent stories. To address these problems, a novel knowledge-enriched attention network with group-wise semantic model is proposed. Three main novel components are designed and supported by substantial experiments to reveal practical advantages. First, a knowledge-enriched attention network is designed to extract implicit concepts from external knowledge system, and these concepts are followed by a cascade cross-modal attention mechanism to characterize imaginative and concrete representations. Second, a group-wise semantic module with second-order pooling is developed to explore the globally consistent guidance. Third, a unified one-stage story generation model with encoder-decoder structure is proposed to simultaneously train and infer the knowledge-enriched attention network, group-wise semantic module and multi-modal story generation decoder in an end-to-end fashion. Substantial experiments on the visual storytelling datasets with both objective and subjective evaluation metrics demonstrate the superior performance of the proposed scheme as compared with other state-of-the-art methods. The source code of this work can be found in https://mic.tongji.edu.cn.
Tengpeng Li, Hanli Wang, Bin He 0003, Chang Wen Chen
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 A Survey of Crowdsensing and Privacy Protection in Digital City
abstract
The key pillar of developing digital city is the ubiquitous sensing of people and the environment. Crowdsensing requires a large number of users to participate in the collection of sensing data, and these data may carry sensitive information, such as identity and location related to the users or sensing object. If this information is eavesdropped, intercepted, and leaked, this may seriously harm the interests of individuals, organizations, and even countries. Therefore, from a privacy perspective, users may be reluctant to open data. While relying on mobile devices used by a large number of ordinary users as the basic sensing unit, it is necessary to include a variety of communication methods to realize the distribution of sensing tasks and to collect the sensing data. Then, to complete the complex crowdsensing tasks, it is important to ensure privacy security in the context of crowdsensing because it is a key problem. In this article, we comb through the development status of crowdsensing in the digital city, emphatically analyze the privacy protection in crowdsensing under the background of digital city, and qualitatively evaluate the existing privacy protection technologies for crowdsensing. Finally, this article presents research challenges and future directions that should be addressed to improve the performance of privacy protection technologies for crowdsensing systems.
Bin He 0003, Gang Li 0020, Bin Cheng 0008
IEEE Trans. Comput. Soc. Syst.2
2023 Robust Adaptive Fuzzy Fault Tolerant Control of Robot Manipulators With Unknown Parameters
abstract
This article resolves the safety tracking problem of the robot manipulator with the actuator faults, process faults, uncertain dynamics, and external disturbances. A novel robust adaptive fuzzy fault tolerant control (FTC) framework is presented. In this approach, a constraint mechanism of the zero overshoot error is proposed to keep the filter variable within an adjustable range, and the compact set can be available. An adaptive fuzzy logic system is, then, employed to approximate the time-varying, nonlinear, and unknown robot dynamics. Furthermore, a robust adaptive term is designed to compensate the actuator faults and approximation errors, and ensure the convergence and stability of the whole robot control system. The uniform ultimate boundedness of the closed-loop dynamics for the robot manipulator is proved by the Lyapunov function. Afterward, the simulation results demonstrate that the proposed robust adaptive fuzzy FTC (RAFFTC) approach of the robot manipulator has good control performance. Finally, the effectiveness of the proposed RAFFTC approach is verified by experiments.
Shoulin Xu, Bin He 0003
IEEE Trans. Fuzzy Syst.2
2023 Communication and Control in Collaborative UAVs: Recent Advances and Future Trends
abstract
The recent progress in unmanned aerial vehicles (UAV) technology has significantly advanced UAV-based applications for military, civil, and commercial domains. Nevertheless, the challenges of establishing high-speed communication links, flexible control strategies, and developing efficient collaborative decision-making algorithms for a swarm of UAVs limit their autonomy, robustness, and reliability. Thus, a growing focus has been witnessed on collaborative communication to allow a swarm of UAVs to coordinate and communicate autonomously for the cooperative completion of tasks in a short time with improved efficiency and reliability. This work presents a comprehensive review of collaborative communication in a multi-UAV system. We thoroughly discuss the characteristics of intelligent UAVs and their communication and control requirements for autonomous collaboration and coordination. Moreover, we review various UAV collaboration tasks, summarize the applications of UAV swarm networks for dense urban environments and present the use case scenarios to highlight the current developments of UAV-based applications in various domains. Finally, we identify several exciting future research direction that needs attention for advancing the research in collaborative UAVs.
Shumaila Javaid, Nasir Saeed, Zakria Qadir, Hamza Fahim, Bin He 0003, Houbing Song, Muhammad Bilal 0003
IEEE Trans. Intell. Transp. Syst.5
2023 Dual Stream Meta Learning for Road Surface Classification and Riding Event Detection on Shared Bikes
abstract
Road surface condition monitoring and bike riding event detection are crucial in densely populated cities for travel efficiency and rider safety. However, most current approaches are either costly, unreliable in different scenarios, or not adaptable in new environments. This article proposes a novel automated approach leveraging widely used shared bikes to intelligently detect road surface conditions and riding events suitable for interactive Internet of Things (IoT) cities. We propose a novel dual stream meta learning approach to solve the reliability problem when bike types for the training and testing are different with a limited set of new samples and the self-adaptive problem when classifying new classes without retraining the model, both via dual stream meta learning. Results demonstrate the feasibility of the proposed IoT-based solution with 98.9% accuracy for road surface conditions and 99.6% accuracy for riding events via the proposed dual stream deep learning method in the conventional scenario. With few samples per class, the proposed method is more reliable than other commonly used approaches in the different-bike scenario (e.g., proposed 92.4% versus random forest 74.6%). In cases of predicting new classes, the algorithm is 95.6% accurate using only one sample per class without explicit training (compared to 78.0% for$K $-nearest neighbor). This article proposes a robust IoT framework for smart cities involving road surface conditions and rider events which could be critical for many applications, including city mapping, shared bike rental maintenance and rider performance, and city maintenance services.
Zachary A. Strout, Bin He 0003, Daiyan Peng, Peter B. Shull, Benny P. L. Lo
IEEE Trans. Syst. Man Cybern. Syst.3
2022 A Data Agent Inspired by Interpersonal Interaction Behaviors for Wireless Sensor Networks
abstract
Event monitoring is the main purpose of monitoring applications based on wireless sensor networks (WSNs). To realize the autonomous representation of tunnel disasters in WSNs, this work proposes a data agent (DA) inspired by interpersonal interaction behaviors, which bridges the gap between data and humans. The DA divides event monitoring into four states: 1) active perception; 2) understanding; 3) thinking; and 4) representation, to realize a kind of human-like event monitoring. First, a DA behavior model based on a finite-state machine is proposed, which can adaptively switch among these four states. Second, four methods, including active perception based on a dual neuron perception structure, understanding for forming the information granularity, thinking for forming the disaster knowledge graph, and representation using grayscale and knowledge graphs are proposed to realize the four states of the DA. The experimental results show that the DA inspired by interpersonal interaction can make WSNs not only achieve efficient and accurate disaster self-monitoring but also represent the disaster situation intuitively and quickly through grayscale and knowledge graphs.
Gang Li 0020, Bin He 0003, Zhipeng Wang 0006, Yanmin Zhou
IEEE Internet Things J.2
2022 Blockchain-Enhanced Spatiotemporal Data Aggregation for UAV-Assisted Wireless Sensor Networks
abstract
Wireless sensor networks (WSNs) are widely used in the field of monitoring. For data collection of sensor nodes in large-scale monitoring scenarios, unmanned aerial vehicle (UAV)-assisted WSNs have emerged. For the security and validity of data collection, a blockchain-enhanced data collection framework for UAV-assisted WSNs is presented in this article. To reduce data redundancy in WSNs, a sparsity-optimized and compressed sensing-based spatiotemporal data aggregation model is built. A UAV identity authentication mechanism based on a Merkle tree is also designed to ensure the security of data transmission. By combining blockchain building and data aggregation, a disaster semantic blockchain (DSB) based on a data reconstruction-directed consensus mechanism is presented. Disaster semantics are extracted by analyzing the semantic association relationship of disaster, background, event, and sensor data. The experimental results show that the blockchain-enhanced spatiotemporal data aggregation effectively increases the network life cycle and data reconstruction accuracy. The disaster situation can be described accurately through the DSB.
Gang Li 0020, Bin He 0003, Zhipeng Wang 0006, Jie Chen 0003
IEEE Trans. Ind. Informatics2
2021 A data-efficient goal-directed deep reinforcement learning method for robot visuomotor skill
Rong Jiang 0003, Zhipeng Wang 0006, Bin He 0003, Yanmin Zhou, Gang Li 0020, Zhongpan Zhu
Neurocomputing3
2021 MABAN: Multi-Agent Boundary-Aware Network for Natural Language Moment Retrieval
abstract
The amount of videos over the Internet and electronic surveillant cameras is growing dramatically, meanwhile paired sentence descriptions are significant clues to select attentional contents from videos. The task of natural language moment retrieval (NLMR) has drawn great interests from both academia and industry, which aims to associate specific video moments with the text descriptions figuring complex scenarios and multiple activities. In general, NLMR requires temporal context to be properly comprehended, and the existing studies suffer from two problems: (1) limited moment selection and (2) insufficient comprehension of structural context. To address these issues, a multi-agent boundary-aware network (MABAN) is proposed in this work. To guarantee flexible and goal-oriented moment selection, MABAN utilizes multi-agent reinforcement learning to decompose NLMR into localizing the two temporal boundary points for each moment. Specially, MABAN employs a two-phase cross-modal interaction to exploit the rich contextual semantic information. Moreover, temporal distance regression is considered to deduce the temporal boundaries, with which the agents can enhance the comprehension of structural context. Extensive experiments are carried out on two challenging benchmark datasets of ActivityNet Captions and Charades-STA, which demonstrate the effectiveness of the proposed approach as compared to state-of-the-art methods. The project page can be found in https://mic.tongji.edu.cn/e5/23/c9778a189731/page.htm.
Hanli Wang, Bin He 0003
IEEE Trans. Image Process.3
2021 A Novel Texture-Less Object Oriented Visual SLAM System
abstract
Traditional mapping modules in Visual Simultaneous Localization and Mapping (i.e. Visual SLAM) systems can only estimate 3D information of isolated sparse or semi-dense feature points. But there are lots of object instances in the environments which geometric information can be utilized to enhance the quality of mapping and localization. Hence, it is required for the Visual SLAM system to utilize high-dimensional features like object instances or structural lines in mapping and localization. To meet the gap between the above requirements and the traditional implementation of Visual SLAM systems, we present in this paper a novel Visual SLAM method that can effectively utilize texture-less object instances for mapping and localization. The proposed Visual SLAM method includes newly designed feature extraction, matching, localization and mapping modules, which jointly use object features and point features to estimate camera 6-DOF poses and do richer map construction. A group of organized raster points is used to represent objects during feature matching and pose estimation process in the proposed Visual SLAM pipeline. Owing to the object feature fusion in the co-visibility graph it could conduct scale aware bundle adjustments to reduce accumulated error. The advantages of proposed Visual SLAM method are demonstrated through experiments conducted both on synthetic datasets and real-world datasets.
Yanchao Dong, Senbo Wang, Jiguang Yue, Ce Chen, Shibo He, Haotian Wang 0003, Bin He 0003
IEEE Trans. Intell. Transp. Syst.7
2018 A traffic congestion aware vehicle-to-vehicle communication framework based on Voronoi diagram and information granularity
Gang Li 0020, Bin He 0003, Aimin Du
Peer-to-Peer Netw. Appl.2
2017 Intelligent Self-Adaptation Data Behavior Control Inspired by Speech Acts
abstract
Wireless sensor networks (WSNs) are a promising technology for collecting information by utilizing various types of small-sized sensors. Implementing efficient data acquisition and transmission is important for real-time monitoring and data analysis, owing to massive and heterogeneous data with low precision. In this work, the data behavior model inspired by speech acts is built, including the data correlation model and the data comment model. Intelligent self-adaptation data behavior control is then proposed, of which the main idea is to make sensor nodes (SNs) process data intelligently. In the proposed data behavior control, the temporal compression behavior is achieved through variable-cycle transmission and the multivariate spatial compression behavior is motivated by compressed sensing (CS) theory. The multivariate spatial-temporal compression composed of the above two compression behaviors can adjust itself through data comment behaviors including the large-cycle error self-correction behavior based on the node credibility and the large-cycle event self-adaptation behavior based on the information granularity. The data behavior control presented has been validated by the experiments in the Shanghai metro tunnel. The experiment results show that these intelligent data behaviors inspired by speech acts can make WSNs more effective and intelligent with no change in the existing network structure.
Bin He 0003, Gang Li 0020
ACM Trans. Sens. Networks1
2015 Salient object detection based on meanshift filtering and fusion of colour information
abstract
Colour and its spatial distribution are the main information currently used to detect salient objects in an image, but this cannot always guarantee satisfying performance. To deal with this problem, a salient object detection algorithm has been presented based on meanshift filtering and fusion of colour information. Superpixel segmentation is used to analyse the images by sets of pixels instead of single pixel, which improves the robustness of the algorithm to noises, as well as the efficiency. Meanshift filtering is used to detect the modes of every superpixel in spatial domain and range domain, respectively, which is the basis of the subsequent calculation. Each target therefore offers almost the same saliency and the spatial distribution of which will be easier to analyse. The fusion of colour contrast and colour concentration as well as centre prior is used as criterion to evaluate the saliency of every single superpixel. According to the tests of the algorithm on the open popular dataset, it has been proved that the algorithm presented in this work shows better results in both the aspect of effectiveness and efficiency, compared with its traditional equivalents.
Gang Li 0020, Bin He 0003, Xiaojiao Tao
IET Image Process.4
2014 Spatial-temporal compression and recovery in a wireless sensor network in an underground tunnel environment
Bin He 0003, Haifeng Tang
Knowl. Inf. Syst.1