Zhaojie Ju

dblp:57/213 · DBLP profile ↗
← Back
86ranked-venue papers
10as first author
43since 2021 · last 2026
0000-0002-9524-7609ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 49 · 9 first-author · 21 since 2021Human-computer interaction and ubiquitous computing · 25 · 15 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 1 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 HALT: Hierarchical Attention Learning for Visual Tracking
abstract
Siamese tracking algorithms have gained widespread recognition due to their exceptional efficiency and scalability. However, they exhibit suboptimal performance when localizing arbitrary targets under various disturbances, particularly in complex environments involving challenges such as illumination variation, deformation, and background clutter. Therefore, this paper proposes a Hierarchical Attention Learning network (HAL) to enhance tracking performance. Drawing inspiration from the hybrid attention mechanism, HAL designs a Feature Mining Attention module (FMA), a Global Feature Attention module (GFA), and a hierarchical network structure. Concretely, FMA employs parallel branches to fully extract channel and spatial features, enabling preliminary enhancement of target features and establishing global associations. Since lower-level and higher-level features emphasize positional and semantic information, respectively, GFA comprehensively integrates multi-level features to obtain more accurate tracking predictions. In particular, to improve the model's representation capacity, a hierarchical network structure is developed to deepen the HAL network and strengthen the feature dependencies between the template and the search region. Finally, based on HAL, we propose a Hierarchical Attention Learning Tracker (HALT) for visual tracking in complex environments, which is capable of learning rich hierarchical features. Extensive experiments demonstrate that, compared to state-of-the-art trackers, our HALT achieves outstanding tracking performance across multiple benchmarks while maintaining a real-time speed of 48.5 fps.
Fengwei Gu, Ao Li 0002, Chen Chen 0086, Chengtao Cai, Renjie Qiao, Guangyao Zhai, Zhaojie Ju
IEEE Trans Autom. Sci. Eng.9
2026 Multistep Intent Estimation Guided Adaptive Passive Control for Safety-Aware Physical Human-Robot Collaboration
abstract
physical human-robot collaboration (pHRC) requires strict safety and efficiency guarantees, imposing heightened demands on accurate human intent estimation and adaptive control in a stable manner. To address these challenges, we propose a novel two-loop adaptive passive control framework guided by multistep human intent estimation to reduce human-robot disagreement and improve robot assistance level, facilitating safety-aware efficient pHRC. In the framework, outer loop's intent estimation guides the inner loop's adaptive passive controller, ensuring real-time robot behavior adjustment based on multistep intention. Specifically, the outer loop incorporates a transformer-based human intent estimator (THIE) that integrates the Transformer with a conditional variational autoencoder (CVAE) for multistep predictions, accurately estimating motion and force to guide the robot. The inner loop incorporates a goal-oriented reinforcement learning (GoRL)-based adaptive impedance control, which constructs multistep rewards based on prediction and probability from THIE to adjust impedance parameters, thereby balancing disagreement and assistance, and promoting locally optimal robot behaviors. Furthermore, an energy tank-based passive model predictive control (ET-PMPC) is employed to limit robot stored energy, avoiding the impact of variable impedance on safety. Experiments validate that our framework outperforms state-of-the-art (SOTA) methods, significantly improving intent estimation accuracy, robot assistance level, and safety, highlighting its potential to advance pHRC.
Zhengtao Zhang, Yuchuang Tong, Zhaojie Ju
IEEE Trans. Cybern.4
2025 IDAGC: Adaptive Generalized Human-Robot Collaboration via Human Intent Estimation and Multimodal Policy Learning
abstract
In Human-Robot Collaboration (HRC), which encompasses physical interaction and remote cooperation, accurate estimation of human intentions and seamless switching of collaboration modes to adjust robot behavior remain paramount challenges. To address these issues, we propose an Intent-Driven Adaptive Generalized Collaboration (IDAGC) framework that leverages multimodal data and human intent estimation to facilitate adaptive policy learning across multi-tasks in diverse scenarios, thereby facilitating autonomous inference of collaboration modes and dynamic adjustment of robotic actions. This framework overcomes the limitations of existing HRC methods, which are typically restricted to a single collaboration mode and lack the capacity to identify and transition between diverse states. Central to our framework is a predictive model that captures the interdependencies among vision, language, force, and robot state data to accurately recognize human intentions with a Conditional Variational Autoencoder (CVAE) and automatically switch collaboration modes. By employing dedicated encoders for each modality and integrating extracted features through a Transformer decoder, the framework efficiently learns multi-task policies, while force data optimizes compliance control and intent estimation accuracy during physical interactions. Experiments highlights our framework’s practical potential to advance the comprehensive development of HRC.
Yuchuang Tong, Guanchen Liu, Zhaojie Ju, Zhengtao Zhang
IROS4
2025 Fine-Grained Action Recognition Using Cross-Modal Attention Network for Human-Robot Sign Language Interaction
abstract
With the growing demand for barrier-free communication among the deaf and mute, human-robot sign language interaction has gradually gained attention as an auxiliary tool. Action recognition serves as a crucial information source for robots to understand human behavior, enabling robots to recognize signs and achieve natural interaction with deaf people through it. However, existing Human-Robot Interaction (HRI) technologies based on action recognition mainly focus on coarse-grained human movements, failing to capture and respond to nuanced actions in real-world scenarios. Additionally, multi-modal action recognition often employs early or late fusion methods to integrate various modalities, lacking the exploration of relationships between modalities, resulting in the loss of some correlated information. To enable robots to adapt to diverse scenarios for a more nuanced understanding of human behaviors, we propose a fine-grained action recognition framework using Cross-Modal Attention network (CMA) based on RGB and skeleton. Firstly, holistic features including face, hand, and body are extracted by a pose estimator, effectively representing intricate human actions. Subsequently, to fully leverage the extracted fine-grained features, skeleton is represented as heatmap volumes. Finally, a Cross-Attention Interaction (CAI) module is designed to explore the intrinsic connections between RGB and skeleton, facilitating mutual learning of their respective advantageous features in the deep layers of feature extraction, thereby achieving information interaction. Simultaneously, HRI experiments are conducted on the large-scale fine-grained action dataset, WLASL2000. In this HRI system, the robotic arm responds by performing sign language aligned with the human actions identified by CMA, showcasing the practicality and effectiveness of our proposed model in real-world scenarios.
Qing Gao 0002, Xianfeng Cheng, Xuerui Li, Zhaojie Ju
SMC5
2025 SignRobot: Sign Language Recognition for Robot Interaction Based on Dual-Stream Multi-Fusion with Frame Enhancement Network
abstract
Deaf and hard-of-hearing individuals often rely on signs for communication, but limited translation resources restrict their daily needs. Robots equipped with sign language recognition and interaction capabilities can assist in bridging this gap. To enhance sign language translation resources, we design a robotic system called SignRobot, which accurately recognizes and responds to signs. To improve recognition performance, we develop a sign language recognition network based on dual-stream multi-fusion and frame enhancement (MFE-Net), using RGB and heatmaps as inputs. Specifically, an frame enhancement module with parallel spatial and motion guidance is introduced to emphasize key spatial regions and movement changes. By employing a multi-fusion strategy, we achieve feature-level interaction and adaptive late fusion between modalities, improving accuracy and robustness. Experimental results show that MFE-Net surpasses state-of-the-art methods on the PHOENIX14 and PHOENIX14-T datasets. Additionally, our SignRobot successfully demonstrates sign recognition and robotic responses in sign language, representing a promising advancement in robot-assisted communication for deaf people.
Qing Gao 0002, Yuanchuan Lai, Yang Zhang 0028, Zhaojie Ju
SMC5
2025 Multi-Scale Video Demoiréing Based on Selective Time-Domain Fusion
abstract
Video demoiréing aims to remove Moiré patterns from video footage captured by cameras when recording electronic displays. Considering the limitation that existing methods often fail to effectively exploit temporal information across video frames, thereby hindering their ability to recover details and maintain temporal consistency, this letter proposed a novel Multi-scale Video Demoiréing network (MVDS) based on Selective Temporal Fusion. By employing deformable convolutions for frame alignment and a selective temporal fusion mechanism, MVDS efficiently exploits temporal information to improve demoiréing performance. An embedded multi-scale U-net with attention fusion further refines image details at different scales. Extensive experiments on the VDMoiré dataset demonstrate that MVDS outperforms state-of-the-art methods in terms of both qualitative and quantitative metrics, effectively removing Moiré patterns while preserving image details and temporal consistency.
Yinfeng Fang, Xiaohao Pan, Yuxi Wang 0002, Dalin Zhou, Zhaojie Ju
IEEE Signal Process. Lett.5
2025 ATTSF-Net: Attention-Based Similarity Fusion Network for Audio-Visual Emotion Recognition
abstract
Emotional factors play a pivotal role in fields such as autonomous driving and intelligent emotional robotics. The accurate extraction of emotional factors is instrumental in reducing error rates within these domains. With the continuous deepening of exploration in the emotional domain, rich multimodal data has progressively supplanted unimodal data. Nevertheless, current multimodal approaches still grapple with the following challenges: 1.Partial loss of information both within individual modalities and across different modalities. 2.Incorrect extraction of modality-invariant features. To facilitate multimodal interaction and address the aforementioned issues, this paper proposes an Attention-based Similarity Fusion Network (ATTSF-Net) for audio-visual emotion recognition. The network is based on multimodal data and comprises the proposed Cross-Multimodal Block (CMB), Similarity Adjustment Block (SAB), and Audio-Visual Auxiliary Modules (AVAM). CMB employs a cross-modal attention mechanism and model-level fusion to facilitate interactions between modalities. SAB is designed to learn modality-invariant features. AVAM utilizes additional audio-visual auxiliary networks to provide supplementary emotional information, enabling the full extraction of intra-modal information. A similarity loss function based on Kullback-Leibler (KL) divergence is designed to ensure the consistency of the learned audio-visual emotional information. The proposed model achieves an accuracy of 88.67% on the RAVDESS dataset (8.67% higher than human) and an unweighted accuracy (UA) of 81.93% and a weighted accuracy (WA) of 79.77% on the IEMOCAP dataset (4.01% and 5.86% higher than transfer learning, respectively). A series of visualization experiments are conducted to demonstrate the effectiveness of the proposed model.
Zhijia Zhang, Zhaojie Ju
IEEE Trans. Affect. Comput.3
2025 A Novel Hybrid 2.5D Map Representation Method Enabling 3D Reconstruction of Semantic Objects in Expansive Indoor Environments
abstract
This paper presents a novel method for creating a hybrid 2.5D semi-semantic map, merging a 2D geometric map with a sparse 3D object map, specifically designed for expansive indoor environments. The primary motivation behind this method is to tackle the issue of temporal and viewpoint discontinuity in RGB-D observations of individual objects within such environments. Notably, current RGB-D SLAM research mainly focuses on small-scale scenarios, often overlooking this specific issue. To address this challenge, our approach proposes to represent objects with a set of object-specific keyframes, and optimizes the spatial relationships between these keyframes in a deferred offline mode. Another objective is to alleviate the accumulation of trajectory drift, which can adversely impact object association/reidentification in expansive environments. To achieve this, we leverage a 2D pose graph SLAM module that indirectly provides initial poses of the RGB-D sensor, thereby facilitating the construction of the sparse 3D object map. Addressing the challenge arising from the dimensional disparity between the two sub-maps (2D vs 3D), we implement a joint optimization strategy to refine the hybrid map, ensuring accuracy compatibility between them. The effectiveness of our proposed method is validated through experiments conducted in real-world environments, and a time efficiency analysis demonstrates the potential of our algorithm to operate in real-time. Note to Practitioners—The motivation of this paper is to create 3D object maps for expansive indoor environments. The challenge arises from the fact that most existing RGB-D SLAM algorithms are designed for small-scale indoor scenarios, lacking suitability for expansive environments with empty areas. In real expansive environments, the short-range capacity of RGB-D sensors increases the risk of failures in continuously observing objects in empty areas, potentially causing the robot to lose its way. This results in two consequences: 1) discontinuity in the viewpoints of individual objects, challenging the online reconstruction of 3D objects; 2) accumulations of trajectory drift during the robot’s loss complicate the reidentification of an object upon its reappearance. Our approach addresses these issues by concurrently generating a 2D geometric map, effectively mitigating trajectory drift. Additionally, we employ a set of object-specific keyframes to represent object reconstruction, efficiently handling viewpoint discontinuity in object observations. Experimental results demonstrate the applicability of our algorithm to real-world environments beyond room-scale scenarios, closely aligning with the challenges encountered in practical applications. The analysis also indicates the potential for real-time operation, rendering our SLAM algorithm advantageous for practical deployment in robotic systems. An additional benefit is that, apart from the 3D object map, our method creates a 2D geometric map, supporting most basic navigation tasks for industrial robots. In future research, we will further improve the time-consuming process and enhance the efficiency of the algorithm.
Shuai Guo 0004, Yazhou Hu, Mingliang Xu 0001, Jinguo Liu, Zhaojie Ju
IEEE Trans Autom. Sci. Eng.5
2025 Recognizing Video Activities in the Wild via View-to-Scene Joint Learning
abstract
Recognizing video actions in the wild is challenging for visual control systems. In-the-wild videos show actions not seen in training data, recorded from various angles and scenes with the same labels. Most existing methods address this challenge by developing complex frameworks to extract spatiotemporal features. To achieve view robustness and scene generalization cost-effectively, we explore view consistency and scene joint understanding. Based on this, we propose a neural network (called Wild-VAR) to learn view and scene information jointly without any 3D pose ground truth labels, a new approach to recognizing video actions in the wild. Unlike most existing methods, first, we propose a Cubing module to self-learn body consistency between views instead of comprehensive image features, boosting the generalization performance of across-view settings. Specifically, we map 3D representations to multiple 2D features and then adopt a self-adaptive scheme to constrain 2D features from different perspectives. Moreover, we propose temporal neural networks (called T-Scene) to develop a recognizing framework, enabling Wild-VAR to flexibly learn scenes across time, including key interactors and context, in video sequences. Extensive experiments show that Wild-VAR consistently outperforms state-of-the-art methods on four benchmarks. Notably, with only half the computation costs, Wild-VAR improves accuracy by 2.2% and 1.3% on the Kinetics-400 and the Something-Somthing V2 datasets, respectively. Note to Practitioners—In human-robot interaction tasks, video action recognition technology is a prerequisite for visual control. In real applications, humans move freely in 3D space, which results in significant changes in the view of video capture and constantly changing scenes. Deep Neural Networks are limited by the perspectives and scenarios contained in the training data, resulting in most existing methods are only effective for identifying actions from 2–4 fixed views, and the background is single. Therefore, existing models are often difficult to generalize to unconstrained application environments. Human view and video scene understanding are often treated separately. Inspired by the human visual system, this paper proposes a view-to-scene video processing method in a cost-efficient way. In real-world applications, this lightweight method can be integrated into robots to help identify human behavior in complex environments. Fewer parameters indicate that the method can be easily migrated to different types of behaviors, and the reduced computational costs represent the ability to achieve real-time performance under limited hardware conditions.
Xuna Wang, Xu Cheng 0003, Zhaojie Ju, Yingke Xu
IEEE Trans Autom. Sci. Eng.5
2025 Holistic-Based Cross-Attention Modal Fusion Network for Video Sign Language Recognition
abstract
As a bridge between the deaf people and the outside, sign language primarily involves hand movements, complemented by intricate facial and body expressions. To enhance the engagement of sign language users in today's prevalent online social activities, it is necessary to incorporate video sign language recognition (SLR) technology into social multimedia. However, current multimodal video SLR methods predominantly rely on limited data, leading to poor robustness and susceptibility to overfitting. Additionally, most works employed simple concatenation, failing to explore effective interaction among different modalities. To address these issues, we propose a cross-attention modal fusion network (CAMFuse) based on red-green-blue (RGB) and skeleton to achieve more robust video SLR. First, CAMFuse introduces a comprehensive bimodal framework that considers coarse-grained features from body and fine-grained features from hands and face. Second, departing from previous skeleton-based video SLR methods represented through graphs, CAMFuse adopts heatmap volumes to reduce storage and promote subsequent interaction. Last, a space-time cross-attention fusion (ST-CAF) module is applied to the deeper feature extraction stages of RGB and skeleton, aiming to mine complementary relationships between them and learn the excellent information from each other, reducing the independence among modalities. Experimental results demonstrate the effectiveness of our proposed CAMFuse, outperforming the state-of-the-art methods on the popular isolated video sign language datasets WLASL-2000 (53.12%) and AUTSL (96.18%) dataset.
Qing Gao 0002, Haixing Mai, Zhaojie Ju
IEEE Trans. Comput. Soc. Syst.4
2025 A Hybrid sEMG-FMG Sensor Fusion Approach Under Muscle Fatigue Using Cascade Fuzzy Forest
abstract
Muscle-driven human–machine systems (HMSs) have advanced significantly in recent years, yet practical applications are often challenged by issues like muscle fatigue. This study introduces an innovative system that captures both surface electromyographic (sEMG) signals and muscle force myography (FMG) signals from the human arm simultaneously. The sEMG signals provide information about the electrical activity of muscle contractions, while the FMG signals monitor morphological changes in the muscles, all in a noninvasive manner. We developed a portable, wearable hybrid sEMG-FMG acquisition system, consisting of a signal acquisition module and an armband to gather both types of signals from the same skin area. Using this system, we examined the sensitivity of sEMG and FMG signals to muscle fatigue and assessed whether combining these two signals could mitigate the effects of muscle fatigue on gesture recognition accuracy. The sEMG and FMG signals were recorded during various hand gestures under both fatigued and nonfatigued conditions. Additionally, a cascade fuzzy forest (CFF) algorithm was developed to enhance hand motion recognition accuracy using the combined sEMG-FMG signals. Experimental results show that FMG signals are resilient to muscle fatigue, and the integration of sEMG with FMG signals significantly reduces the adverse effects of muscle fatigue. The CFF algorithm further improves recognition accuracy, demonstrating the effectiveness of the proposed approach.
Yinfeng Fang, Lingfeng Wu, Xixia Yu, Yong Peng 0001, Zhaojie Ju
IEEE Trans. Fuzzy Syst.6
2025 Spatiotemporal View-Reset Deep Learning With Attentional GRUs for Skeleton-Based Human Action Recognition
abstract
Skeleton-based human action recognition has attracted significant attention. However, the skeleton spatial invariance and temporal context modeling are recent challenges for most existing methods. This work proposes a view-reset network, which integrates two important branches: spatial view-reset module (SVRM) and temporal attention module (TAM). In the SVRM, the skeleton model from different perspectives is reset in a unified coordinate system, eliminating the influence of viewpoint changes. In the TAM, the perception of temporal features is jointly enhanced by weighting each frame’s importance in the context. Furthermore, the pretrained residual network (ResNet) is used for prediction. The sample size is increased through data augmentation to improve the robustness of the model. The SVRM, TAM, and ResNet form an end-to-end learning network. The ablation study proved that the model could record the key skeletons and frames in the sequence and then reset the human body to a new position, making it easy for learning. The proposed model is evaluated on four challenging benchmarks based on the performance of the cross-view evaluation metrics. Experiments prove that the proposed model has superior performance and surpasses many state-of-the-art algorithms, with an increase of 1.91% over the top ten on the NTU RGB+D 60 dataset.
Xuna Wang, Hongwei Gao 0002, Zide Liu, Zhaojie Ju
IEEE Trans. Hum. Mach. Syst.6
2025 Video Object Detection Considering Dynamic Neighborhood Feature Multiplexing
abstract
Video object detection is essential for human-interaction applications, including bimanual manipulation sensing (BMS). The effects of video detection in practical applications still need to be improved, as they are restricted by long-range spatiotemporal dependency analysis. How do humans sense bimanual manipulation in videos, especially for deteriorated clips? We argue that humans analyze the current clips based on earlier memory, namely, long-term spatial and temporal dependencies (LTSTD). However, most existing methods have yet to report significant results, as the limited exploration of these dependencies limits them. Developing an easy-to-integrate module is generally preferred for future applications rather than designing a complex end-to-end framework. Therefore, we propose a dynamic neighborhood feature multiplexing mechanism for online video object detection in this article, which is better at learning LTSTD in flexible and robust ways, boosting existing detection results, called DNFM. Specifically, we develop dynamic memory enhancement neural networks for better long-term feature aggregation with negligible additional computation costs. We multiplex each frame feature to aggregate key enhanced representations under the guidance of dynamic memory recall. The DNFM contributes to various famous detectors in BMS and other challenging detection tasks, and particular attention has been devoted to “low-quality” frame detection. Experimental results show that, while achieving state-of-the-art detection performance, DNFM clearly illustrates the easy-to-integrate operation for boosting the video object detection results.
Xuna Wang, Dalin Zhou, Yingke Xu, Zhaojie Ju
IEEE Trans. Syst. Man Cybern. Syst.8
2025 A Stability-Guaranteed Variable Admittance Control Architecture for Complex Physical Interaction Tasks With Multiple Scenarios
abstract
Interaction stability is an important concern in variable admittance control, and it is a prerequisite for achieving the desired compliant interaction with humans or uncertain environments. Different variable admittance controllers with essential stability analyses are required for multiple interaction scenarios in a complex task, where stability analysis is the major difficulty. In this article, we propose a unified stability-guaranteed variable admittance control architecture to decrease the control system complexity, in which we represent the admittance control model as a port-Hamiltonian (pH) system and design an energy tank with the optimization theory. The advantages of the proposed architecture are reflected in several aspects: 1) compatibility with common strategies; 2) perturbation resistance; and 3) passivity as a natural property. As the dissipated energy from the pH system is injected into the energy tank while the energy in the tank is used to support the interaction and perturbation resistance, the best passive approximation of the desired behaviors is generated, which can avoid the control signal failure (control signal failure denotes that the unexpected input$u=0$occurs due to a lack of energy in the tank). Three different groups of experiments are conducted to verify the feasibility, perturbation resistance, and practicality of the proposed architecture.
Hao Zhou 0042, Xin Zhang 0081, Jinguo Liu, Zhaojie Ju
IEEE Trans. Syst. Man Cybern. Syst.4
2024 Surface defect detection methods for industrial products with imbalanced samples: A review of progress in the 2020s
Dongxu Bai, Gongfa Li, Du Jiang, Juntong Yun, Bo Tao 0002, Guozhang Jiang, Ying Sun 0004, Zhaojie Ju
Eng. Appl. Artif. Intell.8
2024 Sharing-Net: Lightweight feedforward network for skeleton-based action recognition based on information sharing mechanism
Qing Gao 0002, Zhaojie Ju, Yulan Guo
Pattern Recognit.3
2024 EANTrack: An Efficient Attention Network for Visual Tracking
abstract
Recently, Siamese trackers have gained widespread attention in visual tracking due to their exceptional performance. However, many trackers still suffer from limitations in challenging scenarios, such as fast motion and scale variation, which hinder the full exploitation of target features. Consequently, the accuracy and efficiency of the trackers are limited. Therefore, this paper proposes an efficient attention network, called EAN, to improve tracking performance. The EAN comprises three primary components, namely a Transformer-s subnetwork, a Transformer-t subnetwork, and a Feature-Fused Attention Module (FFAM). The designed Transformer-s and Transformer-t subnetworks adopt complementary structures and functions to fully integrate and emphasize the relevant feature information, including channel and spatial features. The FFAM is responsible for fusing the multi-level features from both subnetworks, which establishes the global dependencies between the templates and search regions and enhances the discriminative power of the model. To further improve the tracking accuracy, a novel Feature-Aware Attention Module (FAAM) is introduced into the tracking prediction head to enhance the feature representation capability of the model. Finally, we propose an efficient EANTrack tracker based on EAN for robust tracking in complex scenarios, which exhibits significant advantages in challenging attributes. Experimental results on multiple benchmarks indicate that our approach achieves remarkable tracking performance with a real-time running speed of 55.6fps.Note to Practitioners—Siamese trackers have garnered considerable attention in the field of visual tracking due to their impressive performance. However, these trackers often face limitations in challenging scenarios, which impede the complete exploitation of target features. As a result, the accuracy and efficiency of many trackers are compromised. To address these issues, we propose an efficient tracker called EANTrack to enable robust tracking in complex scenarios. Our EANTrack exhibits significant advantages in handling challenging attributes. Please refer to our complete paper for detailed information on the EANTrack tracker and experimental results. Practitioners in the field can benefit from our research by leveraging our findings and methodologies in their work. We encourage further exploration and experimentation to enhance the performance and applicability of visual tracking systems.
Fengwei Gu, Chengtao Cai, Qidan Zhu, Zhaojie Ju
IEEE Trans Autom. Sci. Eng.5
2024 Adaptive Tracking Control of Robotic Manipulators With Unknown Kinematics and Uncertain Dynamics
abstract
This paper addresses a long-standing yet well-documented open problem on trajectory tracking control of manipulators, which is simultaneously affected by unknown kinematics and uncertain dynamics. A theoretical framework for implementing exponential tracking control is established by unifying two novel controllers, i.e., the Jacobian matrix adaption (JMA) controller and the observer-based estimation law, into an integrated control system. The proposed JMA controller converts internal, implicit and immeasurable model information to external, explicit and measurable input-output information for adaptive learning of unknown kinematics. The proposed observer-based estimation law can guarantee the global exponential stability of the estimation error for accurately estimating uncertain dynamics. In addition, the inherent measurement noise, hard nonlinearity, and limited sampling period lead to chattering phenomenon and accumulated errors occur during convergence process. Hence, an improved simple model-free adaptive sliding mode control (ASMC) scheme is proposed to compensate these limitations, which has fast adaptability and powerful tracking and chattering suppression capabilities. It is theoretically proved that the integrated control system is globally exponentially stable. Simulation, experiments and comparison verify the convergence performance of the proposed integrated control system.Note to Practitioners—Despite the enormous advantages provided by advanced robotic manipulators, developing effective tracking control schemes remains a challenging problem Unfortunately, however, the assumption of precise kinematic and dynamic parameters during tracking control is practically unrealistic due to the absence of precise parameters and measurements of the interaction between robotic manipulators and external environment. This uncertainty will reduce the control performance of manipulators, such as accuracy and repeatability. It is of practical significance to further solve the tracking control of manipulators while considering both the uncertain kinematics and the uncertain dynamics. Therefore, this paper addresses the tracking control problem of manipulators, which is simultaneously affected by unknown kinematics and uncertain dynamics. By unifying the two novel controllers into an integrated control system, a theoretical framework for implementing exponential tracking control is established. In addition, an improved scheme is proposed to compensate chattering phenomenon and accumulated errors during the convergence process. The superior convergence performance of the integrated control system are verified by the simulation and experiments.
Yuchuang Tong, Jinguo Liu, Hao Zhou 0042, Zhaojie Ju, Xin Zhang 0081
IEEE Trans Autom. Sci. Eng.4
2024 RTSformer: A Robust Toroidal Transformer With Spatiotemporal Features for Visual Tracking
abstract
In complex environments, trackers are extremely susceptible to some interference factors, such as fast motions, occlusion, and scale changes, which result in poor tracking performance. The reason is that trackers cannot sufficiently utilize the target feature information in these cases. Therefore, it has become a particularly critical issue in the field of visual tracking to utilize the target feature information efficiently. In this article, a composite transformer involving spatiotemporal features is proposed to achieve robust visual tracking. Our method develops a novel toroidal transformer to fully integrate features while designing a template refresh mechanism to provide temporal features efficiently. Combined with the hybrid attention mechanism, the composite of temporal and spatial feature information is more conducive to mining feature associations between the template and search region than a single feature. To further correlate the global information, the proposed method adopts a closed-loop structure of the toroidal transformer formed by the cross-feature fusion head to integrate features. Moreover, the designed score head is used as a basis for judging whether the template is refreshed. Ultimately, the proposed tracker can achieve the tracking task only through a simple network framework, which especially simplifies the existing tracking architectures. Experiments show that the proposed tracker outperforms extensive state-of-the-art methods on seven benchmarks at a real-time speed of 56.5 fps.
Fengwei Gu, Chengtao Cai, Qidan Zhu, Zhaojie Ju
IEEE Trans. Hum. Mach. Syst.5
2024 Propagation Structure Fusion for Rumor Detection Based on Node-Level Contrastive Learning
abstract
With the rise of social media, the rapid spread of rumors online has resulted in numerous negative effects on society and the economy. The methods for rumor detection have attracted great interest from both academia and industry. Given the widespread effectiveness of contrastive learning, many graph contrastive learning models for rumor detection have been proposed by using the event propagation structure as graph data. However, the existing contrastive models usually treat the propagation structure of other events similar to the anchor events as negative samples. While this design choice allows for discriminative learning, on the other hand, it also inevitably pushes apart semantically similar samples and, thus, degrades model performance. In this article, we propose a novel propagation fusion model called propagation structure fusion model based on node-level contrastive learning (PFNC) for rumor detection based on node-level contrastive learning. PFNC first obtains three augmented propagation structures by masking the text of each node in the propagation structure randomly and perturbing some edges in the propagation structure based on the importance of edges. Then, PFNC applies the node-level contrastive learning method between every two augmented propagation structures to prevent the samples with similar propagation structure from far away. Finally, a convolutional neural network (CNN)-based model is proposed to capture the relevant information that is consistent and supplementary among three augmented propagation structures by regarding the propagation structure of the event as a color picture, three augmented propagation structures as color channels, and each node as a pixel. The experimental results on real datasets show that the PFNC significantly outperforms the state-of-the-art models for rumor detection.
Jiachen Ma 0003, Yong Liu 0029, Chunqiang Hu, Zhaojie Ju
IEEE Trans. Neural Networks Learn. Syst.5
2024 Versatile Graph Neural Networks Toward Intuitive Human Activity Understanding
abstract
Benefiting from the advanced human visual system, humans naturally classify activities and predict motions in a short time. However, most existing computer vision studies consider those two tasks separately, resulting in an insufficient understanding of human actions. Moreover, the effects of view variations remain challenging for most existing skeleton-based methods, and the existing graph operators cannot fully explore multiscale relationship. In this article, a versatile graph-based model (Vers-GNN) is proposed to deal with those two tasks simultaneously. First, a skeleton representation self-regulated scheme is proposed. It is among the first trials that successfully integrate the idea of view adaptation into a graph-based human activity analysis system. Next, several novel graph operators are proposed to model the positional relationships and learn the abstract dynamics between different human joints and parts. Finally, a practical multitask learning framework and a multiobjective self-supervised learning scheme are proposed to promote both the tasks. The comparative experimental results show that Vers-GNN outperforms the recent state-of-the-art methods for both the tasks, with the to date highest recognition accuracies on the datasets of NTU RGB + D (CV: 97.2%), UWA3D (88.7%), and CMU (1000 ms: 1.13).
Yingke Xu, Zhaojie Ju
IEEE Trans. Neural Networks Learn. Syst.4
2023 Editorial
abstract
We are pleased to announce the publication of the Special Issue on Hybrid Control of Autonomous Mobile Robots: Architectures, Algorithms and Applications. The control problem of Autonomous Mobile Robots (AMR) in a dynamic environment is a fundament problem that has been receiving much attention from researchers from the world. The main issue here is how to obtain accurate, flexible, and reliable navigation? To perform a navigation task efficiently and effectively, the robot must have perception, decision-making and action capacities for interacting with the environment. The type and complexity of control architecture are usually related to the complexity of the environment and the task at hand. Navigation methods are classified into two main categories, namely, global planning methods (deliberative navigation) and local planning methods (reactive navigation). The main advantage of local planning methods is that they do not require a priori knowledge on the environment model and sometimes without the explicit model of the robot. In the recent years, several local planning methods have been developed. Most of them are based on artificial potential field, fuzzy logic and artificial neural networks. These methods are generally applicable to unknown environments and can be easily adapted to dynamically changing environments. However, such methods frequently suffer from the problem of local. In addition, the actual trajectory is not optimal in terms of distance and/or travel-time due to lack of global vision on the environment. In global planning methods, a navigation task can be achieved in two phases, namely, trajectory planning and tracking phases. Trajectory planning of a robot revolves around fulfilling some performance criteria (distance, travel-time, and energy consumption) and satisfying a certain number of constraints (geometric, kinematic, and/or dynamic). This ensures a safe and fast navigation solution taking into consideration kinematic and dynamic capacities of the robot, and the constraints related to the environment. However, these methods do not adapt to the dynamic of the environment (unexpected obstacles) or completely unknown environment. As regards to the trajectory planning, several approaches whereby the trajectory is generally made up of line segments connected via tangential circular arcs have been proposed. Most of these works deal with minimum-time trajectory-planning problems, under linear/angular velocity bounds of the platform. Some performance techniques have been developed to reach the goal as quickly as possible by smoothing transitions, thus achieving continuous-curvature trajectories. Concerning the problem of trajectory tracking, it revolves around following a reference trajectory by minimizing the position, orientation and sometimes speeds errors while maintaining the robot's stability. Many control methods have been proposed; some of them are the classic PID control, Lyapunov-based nonlinear control, sliding mode control, and fuzzy logic control. According to the available information on the navigation environment, methods of the first or second group are selected more often. This leads for three classes of control architectures, namely, reactive, deliberative and hybrid ones. Reactive control architectures are based on the “Sense & Act” principle that combines trajectory planning and its execution at the same level. Generally speaking, they are composed of a set of specific behavioral modules (task-specific behaviors). This allows the robot to make real-time decisions based on local perception and reactive interactions required in unknown and dynamically changing environments. The reference of most proposed solutions is the Subsumption Architecture which can be divided into two main classes based on competitive or cooperative mechanisms between behaviors modules. Deliberative control architecture is based on “Sense, Plan & Act” principle used in fully known environments. In fact, the robot model must be known and continually updated to plan the robot's actions. In this approach, one or more trajectories are first planned. Next, according to the actual state of the perceived information, the robot executes trajectory tracking strategies. Deliberative systems are considered as classical control architectures since they were the first to be tested. Given the drawbacks of the two types of methods, the combination of both types gives hybrid control architecture which enables navigation in partially known environments. This choice allows fast and reactive solution while avoiding unexpected obstacles and reducing the traveling time with introduction of partial knowledge of the environment. In fact, some interesting works adopting this approach have been reported in the literature. The last decade witnessed increasingly rapid progress in AI-powered hybrid control of AMR, mainly backed up by advances in the areas of artificial intelligence and deep learning. In particular, AI-based hybrid control architectures, convolutional, and recurrent neural networks, as well as the deep reinforcement learning paradigm have been proposed. These methodologies form a base for scene perception, path planning, behavior arbitration and motion control algorithms. Furthermore, modular perception-planning-action pipeline, where each module is built using deep learning methods which directly map sensory information to steering commands has been investigated. In this special issue, we included original contributions pertaining to architectures, algorithms and applications of hybrid control of AMR. We have performed a professional and strict review process in order to guarantee the quality of the special issue. We would also like to cordially thank all the reviewers who have participated in the review process of the articles submitted to this special issue, and the publishing team.
Meng Joo Er, Zhaojie Ju, Alexander Ferrein
Comput. Intell.2
2023 Repformer: a robust shared-encoder dual-pipeline transformer for visual tracking
Fengwei Gu, Chengtao Cai, Qidan Zhu, Zhaojie Ju
Neural Comput. Appl.5
2023 Distributed Resilient Tracking of Multiagent Systems Under Actuator and Sensor Faults
abstract
The distributed resilient tracking problem for multiagent systems (MASs) is investigated in the presence of actuator/sensor faults over directed topology. Both actuator fault and sensor fault are taken into account. Meanwhile, using the local information, the fault compensators are introduced. Then, based on the fuzzy-logic systems (FLSs) and modification technique of adaptive law, a novel distributed adaptive resilient control protocol is developed, which can compensate the effect of faults on the actuator and sensor. It turns out that all signals of MASs are bounded, while the tracking errors enter an adjustable bounded region around the origin. Toward the end, two simulations are provided to validate the effectiveness of the theoretical results.
Yanming Wu 0002, Jinguo Liu, Zhanshan Wang 0001, Zhaojie Ju
IEEE Trans. Cybern.4
2023 Oropharynx Visual Detection by Using a Multi-Attention Single-Shot Multibox Detector for Human-Robot Collaborative Oropharynx Sampling
abstract
The pandemic of COVID-19 has increased the demand for the oropharynx sampling robots. For an automatic oropharynx sampling, detection and localization of the oropharynx objects are essential. First, in response to the small-object and real-time needs of visual oropharynx detection, a lightweight multi-attention single-shot multibox detector (MASSD) method is designed. This method can effectively improve the detection accuracy of oropharynx sampling regions, especially small regions, while ensuring sufficient speed by introducing spatial attention, channel attention, and feature fusion mechanisms into the single-shot multibox detector. Second, the proposed MASSD is applied to an oropharyngeal swab (OP-swab) robot system to detect oropharynx sampling regions and conduct autonomous sampling. In the experiment, training and validation based on a custom oropharynx dataset verify the effectiveness and efficiency of the proposed MASSD. The detection accuracy can reach 81.3% of mean average [email protected]:0.95 at 104 frames per second and the application experiment on the OP-swab robot system performs oropharynx sampling with 100% success accuracy in human–robot collaboration strategy.
Qing Gao 0002, Yongquan Chen, Zhaojie Ju
IEEE Trans. Hum. Mach. Syst.3
2023 An Efficient RGB-D Hand Gesture Detection Framework for Dexterous Robot Hand-Arm Teleoperation System
abstract
Aiming at the problems of accurate and fast hand gesture detection and teleoperation mapping in the hand-based visual teleoperation of dexterous robots, an efficient hand gesture detection framework based on deep learning is proposed in this article. It can achieve an accurate and fast hand gesture detection and teleoperation of dexterous robots based on an anchor-free network architecture by using an RGB-D camera. First, an RGB-D early-fusion method based on the HSV space is proposed, effectively reducing background interference and enhancing hand information. Second, a hand gesture classification network (HandClasNet) is proposed to realize hand detection and localization by detecting the center and corner points of hands, and a HandClasNet is proposed to realize gesture recognition by using a parallel EfficientNet structure. Then, a dexterous robot hand-arm teleoperation system based on the hand gesture detection framework is designed to realize the hand-based teleoperation of a dexterous robot. Our method achieves high accuracy with fast speed on public and custom hand datasets and outperforms some state-of-the-art methods. In addition, the application of the proposed method in the hand-based teleoperation system can control the grasping of various objects by a dexterous hand-arm system in real time and accurately, which verifies the efficiency of our method.
Qing Gao 0002, Zhaojie Ju, Yongquan Chen, Chuliang Chi
IEEE Trans. Hum. Mach. Syst.2
2023 Parallel Dual-Hand Detection by Using Hand and Body Features for Robot Teleoperation
abstract
Visual hand-based robot teleoperation provides a powerful guarantee for robots to complete complex tasks. However, detection and distinction of dual hands on images are difficult because of the small differences between left and right hands. To solve this problem, a parallel dual-hand detection and distinction method that combines the features of hands with the relationship features between the dual hands and body pose is proposed to achieve robust and accurate dual-hand detection. This parallel dual-hand detection method includes a hand detection module, a body pose estimation module, and a fusion module. In the hand detection module, a hand detector that realizes fast and accurate hand detection by detecting the center and corner points of hands is designed. In the body pose estimation module, a body pose estimator with dual-hand positions is proposed. The fusion module is designed to fuse hand detection and dual-hand estimation results to achieve distinction between left and right hands. Finally, the parallel dual-hand detection method is applied to a bimanual robot teleoperation system by using a designed dual-hand teleoperation framework. The proposed parallel dual-hand detection method can achieve 98.54% mAP of hand detection with 18 frames per second on a custom dual-hand detection dataset, and the bimanual robot teleoperation method can achieve 95.4% average accuracy for teleoperation tasks. Experimental results show the high accuracy and speed of our proposed parallel dual-hand detection method and its practicability in bimanual robot teleoperation.
Qing Gao 0002, Zhaojie Ju, Yongquan Chen, Shiwu Lai
IEEE Trans. Hum. Mach. Syst.2
2023 Mouth Cavity Visual Analysis Based on Deep Learning for Oropharyngeal Swab Robot Sampling
abstract
The visual analysis of the mouth cavity plays a significant role in the pathogen specimen sampling and disease diagnosis of the mouth cavity. Aiming at performance defects of general detectors based on deep learning in detecting mouth cavity components, this article proposes a mouth cavity analysis network (MCNet), which is an instance segmentation method with spatial features, and a mouth cavity dataset (MCData), which is the first available dataset for mouth cavity detecting and segmentation. First, given the lack of a mouth cavity image dataset, the MCData for detecting and segmenting key parts in the mouth cavity was developed for model training and testing. Second, the MCNet was designed based on the mask region-based convolutional neural network. To improve the performance of feature extraction, a parallel multiattention module was designed. Besides, to solve low detection accuracy of small-sized objects, a multiscale region proposal network structure was designed. Then, the mouth cavity spatial structure features were introduced, and the detection confidence could be refined to increase the detection accuracy. The MCNet achieved 81.5% detection accuracy and 78.1% segmentation accuracy (intersection over union = 0.50:0.95) on the MCData. Comparative experiments with the MCData showed that the proposed MCNet outperformed state-of-the-art approaches with the task of mouth cavity instance segmentation. In addition, the MCNet has been used in an oropharyngeal swab robot for COVID-19 oropharyngeal sampling.
Qing Gao 0002, Zhaojie Ju, Yongquan Chen, Tianwei Zhang 0002, Yuquan Leng
IEEE Trans. Hum. Mach. Syst.2
2023 Marrying Global-Local Spatial Context for Image Patches in Computer-Aided Assessment
abstract
Computer-aided assessment using whole slide images (WSIs) is one of the critical steps in clinical procedures. How do doctors recognize cancer in a WSI? A quick answer is that they consider the spatial structure of a WSI rather than only considering single patches. We argue that two clues are essential for computer-aided deep learning: 1) global spatial context and 2) local semantic information. This is because local, semi-local, and global tissue observing are the principal assessment means of pathologists, perfectly corresponding with both clues. However, most existing methods only consider local spatial information learning within each patch rather than developing an effective local-to-global reaction, leading to an incapable of capturing robust and enriched representation. Toward a new area for computer-aided assessment, we propose novel neural networks to learn the global–local spatial context in WSIs, called GLSCL. The GLSCL is among the first trials that understand both clues for WSI understanding. Furthermore, the proposed novel operators enable the GLSCL to learn spatial semantic representation sufficiently. We evaluate the GLSCL using renal cell carcinoma (RCC) samples with synthetic ambiguity collected from the public benchmark and clinical procedures. Enhanced by global and local spatial information, the GLSCL achieves state-of-the-art performance, including classification accuracy, survival prediction index, and cancer tissue attention rate.
Maode Lai, Zhaojie Ju, Yingke Xu
IEEE Trans. Syst. Man Cybern. Syst.5
2022 Improvement of Unconstrained Appearance-Based Gaze Tracking with LSTM
abstract
Gaze tracking is not only an important research direction in computer vision but also an important non-verbal clue in human life. What is important is that the direction of gaze can be used as a reference for judging a person’s intentions. In order to improve the accuracy of predicting gaze direction, a model of 3D gaze tracking based on bidirectional Long Short-Term Memory (LSTM) is proposed in this paper. The backbone network of the model is ResNet and its variants. The output of the model is the angular error of gaze direction. To improve the accuracy of the model prediction, the attention mechanism is adopted in this work. The ablation experiments are conducted on the selected Gaze360, which is a dataset with sufficiently large and diverse data. The angular error of the proposed model decreases from 13.5° to 12.6°.
Guoxu Li, Lihong Dai, Qing Gao 0002, Hongwei Gao 0002, Zhaojie Ju
SMC5
2022 View-Robust Neural Networks for Unseen Human Action Recognition in Videos
abstract
Data-driven deep learning achieved excellent performance for human action recognition. However, unseen action recognition remains a challenge for most existing neural networks. Because the action categories, collection perspectives, and scenarios considered during data collection are limited. Compared with class-unseen action recognition, view-unseen action recognition in videos is under-explored. This paper proposes view-robust neural networks (VR-Net) to recognize unseen actions in videos. The VR-Net consists of a 3D pose estimation module, skeleton adaptive transformation neural networks, and classification modules. We first extract 3D skeleton models from the video sequence based on existing pose estimation methods. Next, we propose a skeleton representation transformation scheme and achieve it based on Convolutional Neural Networks (VR-CNN) and Graph Neural Networks (VR-GCN), resulting in the optimal skeleton representations. Futhermore, we explore an associate optimization scheme and a fused output method. We evaluate the proposed neural networks on three challenging benchmarks, i.e., NTU RGB-D dataset (NTU), Kinetics-400 dataset, and Human3.6M dataset (H3.6M). The experimental results show that view robust neural networks achieve the top performance compared to state-of-the-art RGB-based and skeleton-based works, such as 93.6% on the NTU (CV) and 94.6% on the Kinetics-400 dataset (Top-5). The proposed neural networks significantly improve the recognition performance for unseen action recognition, such as 86.8% on the H3.6M (View 2).
Zhaojie Ju, Yingke Xu
SMC3
2022 Modelling EMG driven wrist movements using a bio-inspired neural network
Yinfeng Fang, Jiani Yang, Dalin Zhou, Zhaojie Ju
Neurocomputing4
2022 A Physics-Guided Coordinated Distributed MPC Method for Shape Control of an Antenna Reflector
abstract
Active shape control for an antenna reflector is a significant procedure used to compensate for the impacts of a complicated space environment. In this article, a physics-guided distributed model predictive control (DMPC) framework for reflector shape control with input saturation is proposed. First, guided by the actual physical characteristics, an overall structural system is decomposed into multilevel subsystems with the help of a so-called substructuring technique. For each subsystem, a prediction model with information interaction is discretized by an explicit Newmark- β method. Then, to improve the system-wide control performance, a coordinator among all the subsystems is designed in an iterative fashion. The input saturation constraints are addressed by transforming the original problem into a linear complementarity problem (LCP). Finally, by solving the LCP, the input trajectory can be obtained. The performance of the proposed DMPC algorithm is validated through an experiment on the shape control of an antenna reflector structure.
Fei Li 0041, Haijun Peng, Xiangshuai Song, Jinguo Liu, Shujun Tan, Zhaojie Ju
IEEE Trans. Cybern.6
2022 Active Disturbance Rejection Control of Euler-Lagrange Systems Exploiting Internal Damping
abstract
Active disturbance rejection control (ADRC) is an efficient control technique to accommodate both internal uncertainties and external disturbances. In the typical ADRC framework, however, the design philosophy is to "force" the system dynamics into a double-integral form by an extended state observer (ESO) and then the controller is designed. Especially, the systems' physical structure has been neglected in such a design paradigm. In this article, a new ADRC framework is proposed by incorporating at a fundamental level the physical structure of the Euler-Lagrange (EL) systems. In particular, the differential feedback gain can be selected considerably small or even 0, due to the effective exploitation of the system's internal damping. The design principle stems from an analysis of the energy balance of EL systems, yielding a physically interpretable design. Moreover, the exploitation of the system's internal damping is thoroughly discussed, which is of practical significance for applications of the proposed design. Besides, a sliding-mode ESO is designed to improve the estimation performance over traditional linear ESO. Finally, the proposed control framework is illustrated through tracking control of an omnidirectional mobile robot. Extensive experimental tests are conducted to verify the proposed design as well as the discussions.
Chao Ren 0003, Yutong Ding, Liang Hu 0002, Jinguo Liu, Zhaojie Ju, Shugen Ma
IEEE Trans. Cybern.5
2022 Deep Temporal Model-Based Identity-Aware Hand Detection for Space Human-Robot Interaction
abstract
Hand detection is a crucial technology for space human-robot interaction (SHRI), and the awareness of hand identities is particularly critical. However, most advanced works have three limitations: 1) the low detection accuracy of small-size objects; 2) insufficient temporal feature modeling between frames in videos; and 3) the inability of real-time detection. In the article, a temporal detector (called TA-RSSD) is proposed based on the SSD and spatiotemporal long short-term memory (ST-LSTM) for real-time detection in SHRI applications. Next, based on the online tubelet analysis, a real-time identity-awareness module is designed for multiple hand object identification. Several notable properties are described as follows: 1) the hybrid structure of the Resnet-101 and the SSD improves the detection accuracy of small objects; 2) three-level feature pyramidal structure retains rich semantic information without losing detailed information; 3) a group of the redesigned temporal attentional LSTM (TA-LSTM) is utilized for three-level feature map modeling, which effectively achieves background suppression and scale suppression; 4) low-level attention maps are used to eliminate in-class similarity between hand objects, which improves the accuracy of identity awareness; and 5) a novel association training scheme enhances the temporal coherence between frames. The proposed model is evaluated on the SHRI-VID dataset (collected according to the task requirements), the AU-AIR dataset, and the ImageNet-VID benchmark. Extensive ablation studies and comparisons on detection and identity-awareness capacities show the superiority of the proposed model. Finally, a set of actual testing is conducted on a space robot, and the results show that the proposed model achieves a real-time speed and high accuracy.
Hongwei Gao 0002, Dalin Zhou, Jinguo Liu, Qing Gao 0002, Zhaojie Ju
IEEE Trans. Cybern.6
2022 Guest Editorial Special Issue on Cyborg Intelligence: Human Enhancement With Fuzzy Sets
abstract
The papers in this special section focus on cyborg intelligence. Well-known scientists and experts have expressed concern that robots may take over the world. More generally, there is a concern that robots could take over human jobs and leave billions of people suffering long-term unemployment. Yet, such concerns ignored the potential of intelligence techniques to enhance the natural capabilities of human beings with in-the-body technologies and so become cyborgs with superior capabilities to robots. Cyborg intelligence is dedicated to improving the natural capabilities of human beings by integrating artificial intelligence (AI) with biological intelligence and in-the-body technologies through tight integrations of machines and biological beings.
Zhijun Li 0001, Jian Huang 0001, Hang Su 0001, Zhaojie Ju
IEEE Trans. Fuzzy Syst.4
2022 Binocular Feature Fusion and Spatial Attention Mechanism Based Gaze Tracking
abstract
Gaze tracking is widely used in driver safety driving, visual impairment detection, virtual reality, human robot interaction, and reading process tracking. However, varying illumination, various head poses, different distances between human and cameras, occlusion of hair or glasses, and low-quality images pose huge challenges to accurate gaze tracking. In this article, based on binocular feature fusion and convolution neural network, a novel method of gaze tracking is proposed, in which local binocular spatial attention mechanism (LBSAM) and global binocular spatial attention mechanism (GBSAM) are integrated into the network model to improve the accuracy. Furthermore, the proposed method is validated on the GazeCapture database. In addition, four groups of comparative experiments have been conducted: between binocular feature fusion model and binocular data fusion model; among the local binocular spatial attention model, the local binocular channel attention model, and the model without local binocular attention mechanism; between the model with GBSAM and that without GBSAM; and between the proposed method and other state-of-the-art approaches. The experimental results verify the advantages of binocular feature fusion, LBSAM and GBSAM, and the effectiveness of the proposed method.
Lihong Dai, Jinguo Liu, Zhaojie Ju
IEEE Trans. Hum. Mach. Syst.3
2022 Deep Object Detector With Attentional Spatiotemporal LSTM for Space Human-Robot Interaction
abstract
Global temporal information and local semantic information are essential cues for high-performance online object detection in videos. However, despite their promising detection accuracy in most cases, most state-of-the-art approaches have following two limitations: invalid background/scale suppression and inadequate temporal information mining between frames. Many jobs currently focus on temporal information learning based on a single frame. In this article, we propose an attentional global–local information learning network; this is one of the first attempts to fully use both types of information between frames. Attention maps are creatively utilized to transfer temporal contexts between frames. This also effectively alleviates the adverse effects of scale changes. Furthermore, empowered by a detailed framework, a proposed detector effectively uses multilevel feature extraction. Given these contributions, the proposed detector achieves state-of-the-art performance on challenging benchmarks. Finally, practical experiments are conducted on a space human–robot interaction platform.
Hongwei Gao 0002, Yongquan Chen, Dalin Zhou, Jinguo Liu, Zhaojie Ju
IEEE Trans. Hum. Mach. Syst.6
2021 A Novel Curved Gaussian Mixture Model and Its Application in Motion Skill Encoding
abstract
The purpose of this paper is to present a novel curved Gaussian Mixture Model (CGMM) and to study the application of it in motion skill encoding. Primarily, Gaussian mixture model (GMM) has been widely applied on many occasions when a probability density function is needed to approximate a complex probability distribution. However, GMM cannot efficiently approach highly non-linear distributions. Thus, the proposed novel CGMM, as a weighted mixture of curved Gaussian models (CGM), is structured with non-linear transfers, which reshapes the flat GMM into a geo-metrically curved one. As a consequence, CGMM has more freedoms and flexibilities than the flat GMM so a CGMM requires fewer number of components in fitting highly non-linear motion trajectories. Moreover, we derive a dedicated iterative parameter estimation algorithm for the CGMM based on maximum likelihood estimation (MLE) theory. To evaluate the performance of the CGMM and its parameter estimation algorithm, a series of quantitative experiments are carried out. We first test the model performance in the data fitting task with the generated synthetic data. Then a motion skill encoding test is carried out on a human motion trajectory dataset built by a Virtual Reality (VR) based motion tracking system. The empirical results support that CGMM outperforms state-of-the-arts in the model performance test. Meanwhile, CGMM has a significant improvement in encoding high dimensional non-linear trajectory data compared to the GMM in motion skill encoding test with its dedicated parameter estimation algorithm.
Disi Chen, Gongfa Li, Dalin Zhou, Zhaojie Ju
IROS4
2021 Hand gesture recognition using multimodal data fusion and multiscale parallel convolutional neural network for human-robot interaction
abstract
Abstract Hand gesture recognition plays an important role in human–robot interaction. The accuracy and reliability of hand gesture recognition are the keys to gesture‐based human–robot interaction tasks. To solve this problem, a method based on multimodal data fusion and multiscale parallel convolutional neural network (CNN) is proposed in this paper to improve the accuracy and reliability of hand gesture recognition. First of all, data fusion is conducted on the sEMG signal, the RGB image, and the depth image of hand gestures. Then, the fused images are generated to two different scale images by downsampling, which are respectively input into two subnetworks of the parallel CNN to obtain two hand gesture recognition results. After that, hand gesture recognition results of the parallel CNN are combined to obtain the final hand gesture recognition result. Finally, experiments are carried out on a self‐made database containing 10 common hand gestures, which verify the effectiveness and superiority of the proposed method for hand gesture recognition. In addition, the proposed method is applied to a seven‐degree‐of‐freedom bionic manipulator to achieve robotic manipulation with hand gestures.
Qing Gao 0002, Jinguo Liu, Zhaojie Ju
Expert Syst. J. Knowl. Eng.3
2021 Attribute-Driven Granular Model for EMG-Based Pinch and Fingertip Force Grand Recognition
abstract
Fine multifunctional prosthetic hand manipulation requires precise control on the pinch-type and the corresponding force, and it is a challenge to decode both aspects from myoelectric signals. This paper proposes an attribute-driven granular model (AGrM) under a machine-learning scheme to solve this problem. The model utilizes the additionally captured attribute as the latent variable for a supervised granulation procedure. It was fulfilled for EMG-based pinch-type classification and the fingertip force grand prediction. In the experiments, 16 channels of surface electromyographic signals (i.e., main attribute) and continuous fingertip force (i.e., subattribute) were simultaneously collected while subjects performing eight types of hand pinches. The use of AGrM improved the pinch-type recognition accuracy to around 97.2% by 1.8% when constructing eight granules for each grasping type and received more than 90% force grand prediction accuracy at any granular level greater than six. Further, sensitivity analysis verified its robustness with respect to different channel combination and interferences. In comparison with other clustering-based granulation methods, AGrM achieved comparable pinch recognition accuracy but was of lowest computational cost and highest force grand prediction accuracy.
Yinfeng Fang, Dalin Zhou, Kairu Li, Zhaojie Ju, Honghai Liu 0001
IEEE Trans. Cybern.4
2021 Physical Human-Robot Collaboration: Robotic Systems, Learning Methods, Collaborative Strategies, Sensors, and Actuators
abstract
This article presents a state-of-the-art survey on the robotic systems, sensors, actuators, and collaborative strategies for physical human-robot collaboration (pHRC). This article starts with an overview of some robotic systems with cutting-edge technologies (sensors and actuators) suitable for pHRC operations and the intelligent assist devices employed in pHRC. Sensors being among the essential components to establish communication between a human and a robotic system are surveyed. The sensor supplies the signal needed to drive the robotic actuators. The survey reveals that the design of new generation collaborative robots and other intelligent robotic systems has paved the way for sophisticated learning techniques and control algorithms to be deployed in pHRC. Furthermore, it revealed the relevant components needed to be considered for effective pHRC to be accomplished. Finally, a discussion of the major advances is made, some research directions, and future challenges are presented.
Uchenna Emeoha Ogenyi, Jinguo Liu, Chenguang Yang 0001, Zhaojie Ju, Honghai Liu 0001
IEEE Trans. Cybern.4
2021 Robot Motor Skill Transfer With Alternate Learning in Two Spaces
abstract
Recent research achievements in learning from demonstration (LfD) illustrate that the reinforcement learning is effective for the robots to improve their movement skills. The current challenge mainly remains in how to generate new robot motions automatically to perform new tasks, which have a similar preassigned performance indicator but are different from the demonstration tasks. To deal with the abovementioned issue, this article proposes a framework to represent the policy and conduct imitation learning and optimization for robot intelligent trajectory planning, based on the improved locally weighted regression (iLWR) and policy improvement with path integral by dual perturbation (PI2-DP). Besides, the reward-guided weight searching and basis function’s adaptive evolving are performed alternately in two spaces, i.e., the basis function space and the weight space, to deal with the abovementioned problem. The alternate learning process constructs a sequence of two-tuples that join the demonstration task and new one together for motor skill transfer, so that the robot gradually acquires motor skill, from the task similar to demonstration to dissimilar tasks with different performance metrics. Classical via-points trajectory planning experiments are performed with the SCARA manipulator, a 10-degree of freedom (DOF) planar, and the UR robot. These results show that the proposed method is not only feasible but also effective.
Xiang Teng, Ce Cao, Zhaojie Ju, Ping Lou
IEEE Trans. Neural Networks Learn. Syst.4
2020 Combining Reinforcement Learning and Rule-based Method to Manipulate Objects in Clutter
abstract
Picking up the clustered objects is always a challenging task in robot research field. And reinforcement learning enables robot to adapt to different tasks through plenty of attempts. To reduce the complexity of strategy learning, we propose a framework for robots to pick up the objects in clutter on table based on deep reinforcement learning and rule-based method. To manipulate the objects on table, we mainly divide the robot actions into two categories: one is pushing that uses the reinforcement learning method, while the other one is grasping that is inferred by image morphological processing. The pushing action can separate the stacking objects, create a robust grasp point for the following grasp. The grasp detect algorithm determines if there is a suitable grasp point. Judging on the result of pushing, the grasp detect algorithm will return a reward for pushing learning. Taking images as input, our framework can keep a high grasp rate with low computational complexity, which makes it achieve clutter clearing quickly.
Zhaojie Ju, Chenguang Yang 0001
IJCNN2
2020 Robust real-time hand detection and localization for space human-robot interaction based on deep learning
Qing Gao 0002, Jinguo Liu, Zhaojie Ju
Neurocomputing3
2020 Neural networks and learning systems for human machine interfacing
abstract
This is a PDF file of an article that has undergone enhancements after acceptance, such as the addition of a cover page and metadata, and formatting for readability, but it is not yet the definitive version of record.This version will undergo additional copyediting, typesetting and review before it is published in its final form, but we are providing this version to give early visibility of the article.Please note that, during the production process, errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.
Zhaojie Ju, Jinguo Liu, Yongan Huang, Naoyuki Kubota, John Q. Gan
Neurocomputing1
2020 Attention mechanism-based CNN for facial expression recognition
Jing Li 0027, Kan Jin, Dalin Zhou, Naoyuki Kubota, Zhaojie Ju
Neurocomputing5
2020 Decomposition algorithm for depth image of human health posture based on brain health
Bowen Luo, Ying Sun 0004, Gongfa Li, Disi Chen, Zhaojie Ju
Neural Comput. Appl.5
2020 Multi-stage adaptive regression for online activity recognition
Bangli Liu, Haibin Cai, Zhaojie Ju, Honghai Liu 0001
Pattern Recognit.3
2019 Spatial Map Learning with Self-Organizing Adaptive Recurrent Incremental Network
abstract
Biological information inspires the advancement of a navigational mechanism for autonomous robots to help people explore and map real-world environments. However, the robot's ability to constantly acquire environmental information in real-world, dynamic environments has remained a challenge for many years. In this paper, we propose a self-organizing adaptive recurrent incremental network that models human episodic memory to learn spatiotemporal representations from novel sensory data. The proposed method termed as SOARIN consists of two main learning process that is active learning and episodic memory playback. For active learning (robot exploration), SOARIN quickly learns and adapts incoming novel sensory data as episodic neurons via competitive Hebbian Learning. Episodic neurons are connecting with each other and gradually forms a spatial map that can be used for robot localization. Episodic memory playback is triggered whenever the robot is in an inactive mode (charging or hibernating). During playback, SOARIN gradually integrates knowledge and experience into more consolidate spatial map structures that can overcome the catastrophic forgetting. The proposed method is analyzed and evaluated in term of map learning and localization through a series of real robot experiments in real-world indoor environments.
Wei Hong Chin, Naoyuki Kubota, Chu Kiong Loo, Zhaojie Ju, Honghai Liu 0001
IJCNN4
2019 A novel feature extraction method for machine learning based on surface electromyography from healthy brain
Gongfa Li, Jiahan Li, Zhaojie Ju, Ying Sun 0004, Jianyi Kong
Neural Comput. Appl.3
2019 RGB-D sensing based human action and interaction analysis: A survey
Bangli Liu, Haibin Cai, Zhaojie Ju, Honghai Liu 0001
Pattern Recognit.3
2019 Electrotactile Feedback in a Virtual Hand Rehabilitation Platform: Evaluation and Implementation
abstract
Tactile feedback plays an important role in hand manipulation, especially in the grasping process which is one of the major functions of the hand. However, few commercially available prosthetic hands or hand motor function rehabilitation systems are equipped with tactile feedback. The absence of suitable tactile feedback modules leads to an inferior rehabilitation performance with a large burden on user training and compromised usability. Thus, it is challenging and essential to integrate a proper tactile feedback module with the existing hand rehabilitation systems to achieve a better control performance and accelerate the rehabilitation process. This paper focuses on the implementation and evaluation of the electrotactile feedback (EF) enhanced rehabilitation system. A virtual hand rehabilitation platform is proposed comprising an surface electromyography (sEMG) acquisition module, an electrotactile stimulation module, a virtual environment with sEMG-driven humanlike hand and numerical feedbacks of grasping force and fingertip deformation, where a closed-loop control is formed. Three different feedback conditions including visual feedback (VF), EF, and no feedback (NF) are compared based on the proposed platform. Experiments were conducted on 10 able-bodied subjects, and multiple quantitative metrics for the rehabilitation performance evaluation including training burden estimation and success rate (SR) of tasks were adopted. Results indicate that the integration of EF is helpful to both reduce the rehabilitation duration and improve the virtual grasping SR in comparison with the NF condition while possessing a better practicality over VF. Note to Practitioners-This paper is motivated by the problem of hand grasp control for rehabilitation purposes, but it also applies to other hand motor function rehabilitation process. Existing hand motor function rehabilitation approaches generally lack a proper feedback and rely on the heavy burden during user training and the users experience. This paper suggests incorporating EF to improve the efficiency and efficacy of the rehabilitation process. The electrical stimulation is driven by the myoelectric-sensing-based force estimation and encoded in a manual scheme to fit each individual involved. In our work, a virtual hand rehabilitation platform is implemented to verify the feasibility of EF in reducing the burden of user training and improving the rehabilitation performance, which allows the expanding of the current system into a broader spectrum of motor function rehabilitation applications. Experiments on able-bodied subjects suggest that the EF in the proposed virtual hand rehabilitation platform is feasible, but it has not been tested on the limb-impaired subjects and confined to a manual encoding of electrical stimulation. In the future research, the design of a general EF enhanced hand rehabilitation platform with a standardized stimulation parameter optimization will be addressed and further validated on the subjects with limb impairments and amputation.
Kairu Li, Peter Boyd, Yu Zhou 0013, Zhaojie Ju, Honghai Liu 0001
IEEE Trans Autom. Sci. Eng.4
2018 Accurate Eye Center Localization via Hierarchical Adaptive Convolution
Haibin Cai, Bangli Liu, Zhaojie Ju, Serge Thill, Tony Belpaeme, Bram Vanderborght, Honghai Liu 0001
BMVC3
2018 Online Action Recognition based on Skeleton Motion Distribution
Bangli Liu, Zhaojie Ju, Naoyuki Kubota, Honghai Liu 0001
BMVC2
2018 Gesture Recognition Based on Depth Information and Convolutional Neural Network
abstract
Vision-based gesture recognition accords with natural communication habits of human and can carry out long-distance and non-contact interactions. So it has become a hot direction in human-computer interaction research whose recognition effect largely depends on the performance of image preprocessing and recognition algorithms. In this paper, a gesture recognition method using color image and depth image combined is designed. For the influence of the angle on the same gesture, the skeleton algorithm is optimized based on the layer-by-layer stripping concept. The fast refinement algorithm improves the process of repeated scanning, extracts the key node information in the skeleton map of the hand, and establishes the spatial axis of the hand to determine the gesture direction. The gesture recognition experiment was performed based on convolutional neural network. The results showed the recognition accuracy rate was 96.01%, and the robustness and accuracy of the proposed recognition method were verified.
Du Jiang, Gongfa Li, Guozhang Jiang, Disi Chen, Zhaojie Ju
SMC5
2018 Knowledge Representation and Knowledge Base System Modeling of Lean Evaluation Model
abstract
Aiming at the phenomenon of low lean level of Chinese enterprise and the over lean level of foreign enterprise, it is a significant to build an evaluation tool of the enterprise lean degree to guide the sustainablility of the enterprises' lean improvement. Combined with the current research results of lean, a design scheme of knowledge base system based on lean evaluation model with 5 layers structure is put forward. With the existing model representation method, a method of model knowledge representation are combined with the object-oriented and framework, and which makes the model knowledgeable. An UML technology is used to modelling the cases of the system, and the dynamic process and static class of the system are studied. Finally, a simulation example is given to verify the effectiveness of the system.
Guozhang Jiang, Xiaowu Chen 0002, Gongfa Li, Zhaojie Ju
SMC5
2018 A structured multi-feature representation for recognizing human action and interaction
Bangli Liu, Zhaojie Ju, Honghai Liu 0001
Neurocomputing2
2017 A force-driven granular model for EMG based grasp recognition
abstract
It is a challenge to precisely predict hand grasps based on EMG signals given practical scenarios, due to its inherent nature. This paper proposes a solution to tackle the challenge with a force-driven granular model (FDGM). The problem of n-class hand grasp classification has been represented as force-based granular modelling, in which a number of granules are constructed for each class relying on the synchronically captured grasping force. A rule based mechanism is formed for granule generation of each class, and a cross-testing algorithm is proposed to optimise the number of granules. The experiment based on 8-case grasp recognition reveals that the proposed method performs better in terms of motion recognition accuracy of multiple EMG channel combination, and is more insensitive to signal interferences. In comparison with other rules of information granulation, it is confirmed that the force-driven rule is of the most efficiency with comparable classification accuracy. The research outcomes pave the way for real-time prediction of grasps and corresponding force in human-centred environments.
Yinfeng Fang, Dalin Zhou, Kairu Li, Zhaojie Ju, Honghai Liu 0001
SMC4
2017 Activity recognition for asd children based on joints estimation
abstract
Human motion recognition is a trending topic and could be applied in many areas, the motion estimation of ASD children is more challenging because of the high uncertainty of their activities, we thus introduced a novel method which is designed for estimating the upper joints and recognising their special motions, we verified the proposed method on our recorded ASD children dataset and adult dataset, the experimental results show the proposed method is effective on the dataset.
Dongxu Gao, Zhaojie Ju, Yingfeng Fang, Jiangtao Cao, Chenguang Yang 0001, Honghai Liu 0001
SMC2
2017 The design of multi-task simulation manipulator based on motor imagery EEG
abstract
In this paper, a mind controlled multi-task manipulator based on motor imagery electroencephalogram (EEG) is proposed. Describe the system function first: In the case of only two types of control signal, the implementation of multi-task Manipulator relies on a toggle-confirmation mode of operation: the task is switched when imagining the left-hand movement, and the task is confirmed when the right-hand movement is imagined. In the BCI system, common spatial pattern (CSP) is used for feature extraction, mutual information for feature selection, and linear discriminant analysis (LDA) for pattern classification. The EEG signal is processed and classified into two categories, imagery of left-hand and right-hand movement. In this way, we can achieve the multi-task control of the manipulator under the premise of ensuring the accuracy of EEG recognition.
Yuhang Ye 0002, Chenguang Yang 0001, Zhaojie Ju, Zhijun Li 0001
SMC4
2017 Real-time visual tracking based on improved perceptual hashing
Mengjuan Fei, Zhaojie Ju, Xiantong Zhen, Jing Li 0027
Multim. Tools Appl.2
2017 Robot manipulator self-identification for surrounding obstacle detection
abstract
Obstacle detection plays an important role for robot collision avoidance and motion planning. This paper focuses on the study of the collision prediction of a dual-arm robot based on a 3D point cloud. Firstly, a self-identification method is presented based on the over-segmentation approach and the forward kinematic model of the robot. Secondly, a simplified 3D model of the robot is generated using the segmented point cloud. Finally, a collision prediction algorithm is proposed to estimate the collision parameters in real-time. Experimental studies using the Kinect Ⓡ sensor and the Baxter Ⓡ robot have been performed to demonstrate the performance of the proposed algorithms.
Xinyu Wang 0018, Chenguang Yang 0001, Zhaojie Ju, Hongbin Ma, Mengyin Fu
Multim. Tools Appl.3
2016 Data fusion-based real-time hand gesture recognition with Kinect V2
abstract
Hand gesture recognition is an important topic in human-computer interaction. However, most of the current methods are complicated and time-consuming, which limits the use of hand gesture recognition in real-time circumstances. In this paper, we propose a data fusion-based hand gesture recognition model by fusing depth information and skeleton data. Because of the accurate segmentation and tracking with Kinect V2, the model can achieve real-time performance, which is 18.7% faster than some of the state-of-the-art methods. Based on the experimental results, the proposed model is accurate and robust to rotation, flip, scale changes, lighting changes, cluttered background, and distortions. This ensures its use in different real-world human-computer interaction tasks.
Yuhai Lan, Jing Li 0027, Zhaojie Ju
HSI3
2016 Real-time 3D point cloud segmentation using Growing Neural Gas with Utility
abstract
This paper proposes a real-time feature extraction and segmentation method for a 3D point cloud. First of all, we apply Growing Neural Gas with Utility (GNG-U) to the point cloud for learning a topological structure. However, the standard GNG-U cannot learn the topological structure of 3D space environment and color information simultaneously. To this end, we then modify the GNG-U algorithm by using a weight vector. we propose a surface feature extraction and segmentation method by efficiently utilizing the topological structure. Our segmentation method is based on a region growing method whose similarity value uses the inner value of two normal vectors connected by the topological structure. We show experimental results of the proposed method and discuss the effectiveness of the proposed method.
Yuichiro Toda, Zhaojie Ju, Hui Yu 0001, Naoyuki Takesue, Kazuyoshi Wada, Naoyuki Kubota
HSI2
2016 Multi-view transition HMMs based view-invariant human action recognition method
Xiaofei Ji, Zhaojie Ju, Changhui Wang
Multim. Tools Appl.2
2016 A novel approach to extract hand gesture feature in depth images
Zhaojie Ju, Dongxu Gao, Jiangtao Cao, Honghai Liu 0001
Multim. Tools Appl.1
2015 Real time object tracking via a mixture model
abstract
Object tracking has been applied in many fields such as intelligent surveillance and computer vision. Although much progress has been made, there are still many puzzles which pose a huge challenge to object tracking. Currently, the problems are mainly caused by appearance model as well as real-time performance. A novel method was been proposed in this paper to handle both of these problems. Locally dense contexts feature and image information (i.e. the relationship between the object and its surrounding regions) are combined in a Bayes framework. Then the tracking problem can be seen as a prediction question which need to compute the posterior probability. Both scale variations and temple updating are considered in the proposed algorithm to assure the effectiveness. To make the algorithm runs in a real time system, a Fourier Transform (FT) is used when solving the Bayes equation. Therefore, the MMOT (Mixture model for object tracking) runs in real-time and performs better than state-of-the-art algorithms on some challenging image sequences in terms of accuracy, quickness and robustness.
Dongxu Gao, Zhaojie Ju, Jiangtao Cao, Honghai Liu 0001
RO-MAN2
2015 A New Wearable Ultrasound Muscle Activity Sensing System for Dexterous Prosthetic Control
abstract
In this paper we introduce a novel Wearable Ultrasound Radial Muscle Activity Detection System (WURMADS) and a demonstration of its ability to recognize forearm muscle activities of amputee subjects to control a dexterous prosthetic hand. The system consists of control electronics to capture and record the ultrasound echo signals in real-time and two wearable bands embedded with eight ultrasound transducers. Based on the principles of Sonomyography, we recognized the intended isotonic finger gestures (trans-radial amputees) by analysing the echo patterns captured by eight single element transducers. Conventional non-invasive my electric detection strategy, Surface Electro Myography (sEMG), cannot reliably identify the deeper muscle activations in the forearm due to crosstalk and signal attenuation. It also suffers from signal degradation due to muscle fatigue and non-linearity. For this study we custom designed 1D single element waterproof transducers to be assembled in a radial structure. A real-time capturing electronic system was also designed, which is portable and wearable. The system operates at 5MHz, the frequency that demonstrated optimum performance for forearm applications. Ten trans-radial amputee subjects were employed in this study to capture data by executing five supervised gestures. The results obtained from offline analysis show that good gesture recognition rates can be achieved, while maintaining a decent correlation coefficient margin to distinguish gestures.
Nalinda Hettiarachchi, Zhaojie Ju, Honghai Liu 0001
SMC2
2015 Automatic Reconstruction of Dense 3D Face Point Cloud with a Single Depth Image
abstract
Human face analysis is the basis for many other computer vision tasks, such as camera surveillance, entrance authorization and age estimation. With 3D face models, the vision task based on facial analysis can usually achieve a higher accuracy than the 2D cases since it provides more information with the additional dimension. However, most existing 3D face reconstruction methods suffer from complicated processing and high computation. This paper presents a novel method that simplifies the 3D face reconstruction process with only one shot of Kinect data. The output of the system is a high density of 3D face point cloud with smoother surface. This provides rich details of the human face for other computer vision tasks. Experiments with real world data show promising results using the proposed method.
Shu Zhang 0002, Hui Yu 0001, Junyu Dong, Ting Wang 0018, Zhaojie Ju, Honghai Liu 0001
SMC5
2015 Time series modeling of surface EMG based hand manipulation identification via expectation maximization algorithm
Yang Lu 0003, Zhaojie Ju, Yurong Liu, Yuxuan Shen, Honghai Liu 0001
Neurocomputing2
2014 Finger pinch force estimation through muscle activations using a surface EMG sleeve on the forearm
abstract
For prosthetic hand manipulation, the surface Electromyography(sEMG) has been widely applied. Researchers usually focus on the recognition of hand grasps or gestures, but ignore the hand force, which is equally important for robotic hand control. Therefore, this paper concentrates on the methods of finger forces estimation based on multichannel sEMG signal. A custom-made sEMG sleeve system omitting the stage of muscle positioning is utilised to capture the sEMG signal on the forearm. A mathematic model for muscle activation extraction is established to describe the relationship between finger pinch forces and sEMG signal, where the genetic algorithm is employed to optimise the coefficients. The results of experiments in this paper shows three main contributions: 1) There is a systematical relationship between muscle activations and the pinch finger forces. 2) To estimate the finger force, muscle precise positioning for electrodes placement is not inevitable. 3) In a multi-channel EMG system, selecting specific combinations of several channels can improve the estimation accuracy for specific gestures.
Yinfeng Fang, Zhaojie Ju, Honghai Liu 0001
FUZZ-IEEE2
2014 A modified EM algorithm for hand gesture segmentation in RGB-D data
abstract
This paper proposes a novel method with a modified Expectation-Maximisation (EM) Algorithm to segment hand gestures in the RGB-D data captured by Kinect. With the depth map and RGB image aligned by the genetic algorithm to estimate the key points from both depth and RGB images, a novel approach is proposed to refine the edge of the tracked hand gesture, which is used to segment the RGB image of the hand gestures, by applying a modified EM algorithm based on Bayesian networks. The experimental results demonstrated the modified EM algorithm effectively adjusts the RGB edges of the segmented hand gestures. The proposed methods have potential to improve the performance of hand gesture recognition in Human-Computer Interaction (HCI).
Zhaojie Ju, Yuehui Wang, Wei Zeng 0001, Haibin Cai, Honghai Liu 0001
FUZZ-IEEE1
2014 Grounding spatial relations in natural language by fuzzy representation for human-robot interaction
abstract
This paper addresses the issue of grounding spatial relations in natural language for human-robot interaction and robot control. The problem is approached by identifying two set of spatial relations, the image space-based and object-centered, and expressing them as fuzzy sets to capture the ambiguity inherent to the linguistic expressions for the relations. The sizes and shades of the scene objects have also been modeled as fuzzy sets for conditioning the spatial relations. To verify the validity of our approach and test its feasibility in a natural language-based interface, we have considered the typical scenarios of using the spatial relations in simple declarative and imperative sentences and designed simple grammars for parsing such sentences. Our experiment has shown that fuzzy spatial relation analysis provides a useful way for modeling the ambiguity or imprecision of the natural language in describing spatial relations and that it is possible to use the spatial relation models to support robot control and human-robot interaction in a natural language-based interface.
Jiacheng Tan, Zhaojie Ju, Honghai Liu 0001
FUZZ-IEEE2
2014 An algorithm for real-time object tracking in complex environment
abstract
The current sparse representation tracking algorithm is not suitable for the objects that illumination changes, scale changes, the object color is similar with the surrounding region, and occlusion etc, what's more, it is hard to realize real-time tracking for solving an l1 norm related minimization problems. An optimal algorithm is introduced by exploiting an accelerated proximal gradient approach which contains some improvements of particle filter function, sparse representation alterative weights and coefficient. These improvements not only reduce the influences of appearance change but also make the tracker runs in real time. Both qualitative and quantitative evaluations demonstrate that the proposed tracking algorithm has favorably better performance than several state-of-the-art trackers using challenging benchmark image sequences, and significantly reduces the computing cost.
Dongxu Gao, Jiangtao Cao, Zhaojie Ju
IJCNN3
2014 Image factorization and feature fusion for enhancing robot vision in human face recognition
abstract
Illumination variation has been a challenging problem for face recognition in robot vision. To reduce the effect caused by illumination variation, a lot of studies have been explored. The Total Variation (TV) method is particular used to factorize images into a low frequency component and a high frequency one. However, the low frequency component still contains significant intrinsic features resulting in failure in face recognition in some cases. In this paper, we propose to further extract illumination invariant features from face images under uncontrolled varying lighting conditions. The Nonsampled Contourlet Transform (NSCT) method is employed to enhance the extraction of intrinsic feature. The combined factorization model is very effective in the experiment on the Yale database.
Hui Yu 0001, Zhaojie Ju, Honghai Liu 0001
IJCNN2
2014 Dynamical Characteristics of Surface EMG Signals of Hand Grasps via Recurrence Plot
abstract
Recognizing human hand grasp movements through surface electromyogram (sEMG) is a challenging task. In this paper, we investigated nonlinear measures based on recurrence plot, as a tool to evaluate the hidden dynamical characteristics of sEMG during four different hand movements. A series of experimental tests in this study show that the dynamical characteristics of sEMG data with recurrence quantification analysis (RQA) can distinguish different hand grasp movements. Meanwhile, adaptive neuro-fuzzy inference system (ANFIS) is applied to evaluate the performance of the aforementioned measures to identify the grasp movements. The experimental results show that the recognition rate (99.1%) based on the combination of linear and nonlinear measures is much higher than those with only linear measures (93.4%) or nonlinear measures (88.1%). These results suggest that the RQA measures might be a potential tool to reveal the sEMG hidden characteristics of hand grasp movements and an effective supplement for the traditional linear grasp recognition methods.
Gaoxiang Ouyang, Zhaojie Ju, Honghai Liu 0001
IEEE J. Biomed. Health Informatics3
2012 A generalised framework for analysing human hand motions based on multisensor information
abstract
In this paper, an integrated framework with multiple sensory information for analysing human hand motions is proposed, and it consists of components of system integration, signal preprocessing, correlation study of sensory information and human motion recognition based on manipulation intention. Three types of sensors are employed in the framework to simultaneously capture the finger angle trajectory, the hand contact force and the forearm electromyography (EMG) signal. The signal preprocessing module is to facilitate the rapid acquisition of human hand tasks by automatically synchronising and segmenting the manipulation primitives. Correlations of the sensory information are studied by using Empirical Copula and demonstrate there exist significant relationships between muscle signals and finger trajectories and between muscle signals and contact forces. In addition, motion recognition based on the EMG intention is investigated by using both Gaussian Mixture Models (GMMs) and Support Vector Machine (SVM) and discussion of the comparative results is presented.
Zhaojie Ju, Honghai Liu 0001
FUZZ-IEEE1
2012 Surface EMG signals determinism analysis based on recurrence plot for hand grasps
abstract
This paper proposes determinism measure (DET) based on recurrence plot, which is capable of showing the recurrence property of a deterministic dynamical system, to evaluate the dynamical characteristics of the surface electromyogram (sEMG) during three different hand movements. In addition, the linear discriminant analysis (LDA) is applied to evaluate the performance of the above measures to identify these three hand grasp movements. The experimental result shows that the recognition rate, 96.7%, based on the combination of the linear and non-linear measures is much higher than that with only linear measures, and DET might be a potential tool to reveal the sEMG hidden characteristics of hand grasp movements.
Gaoxiang Ouyang, Zhaojie Ju, Honghai Liu 0001
IJCNN2
2012 Fuzzy Gaussian Mixture Models
Zhaojie Ju, Honghai Liu 0001
Pattern Recognit.1
2011 Hand motion recognition via fuzzy active curve axis Gaussian mixture models: A comparative study
abstract
Unconstrained human hand motions consisting grasp motion and in-hand manipulation lead to a fundamental challenge that many algorithms have to face in both theoretical and practical development, mainly due to the complexity and dexterity of the human hand. In this paper, fuzzy active curve axis Gaussian Mixture Model (FAcaGMM) is proposed by introducing a weighting exponent on the fuzzy membership into active curve axis Gaussian Mixture Models (AcaGMM) to improve its convergence efficiency, and then FAcaGMM is used to recognize human hand motions. In addition, a comparative study of recognition methods including FAcaGMM, Time Clustering (TC), Empirical Copula (EC), GMM and HMM is presented to recognize human hand motions including both grasps and in-hand manipulations from different subjects with varying training samples.
Zhaojie Ju, Honghai Liu 0001
FUZZ-IEEE1
2011 A Unified Fuzzy Framework for Human-Hand Motion Recognition
abstract
Unconstrained human-hand motions that consist grasp motions and in-hand manipulations lead to a fundamental challenge that many algorithms have to face in both theoretical and practical development, mainly due to the complexity and dexterity of the human hand. There is no effective solution reported to recognize in-hand manipulations, although recognition algorithms have been proposed to recognize grasp motions in constrained scenarios. This paper proposes a novel unified fuzzy framework of a set of recognition algorithms: time clustering, fuzzy active axis Gaussian mixture mode, and fuzzy empirical copula, from numerical clustering to data dependence structure in the context of optimally real-time human-hand motion recognition. Time clustering is a fuzzy time-modeling approach that is based on fuzzy clustering and Takagi--Sugeno modeling with a numerical value as output. The fuzzy active axis Gaussian mixture model effectively extract abstract Gaussian pattern to represent components of hand gestures with a fast convergence. A fuzzy empirical copula utilizes the dependence structure among the finger joint angles to recognize the motion type. The proposed algorithms have been evaluated on a wide range of scenarios of human-hand recognition: 1) datasets that include 13 grasps and ten in-hand manipulations; 2) single subject and multiple subjects; and 3) varying training samples. The experimental results have demonstrated that the proposed framework outperforms the hidden Markov model (HMM) and Gaussian mixture model in terms of both effectiveness and efficiency criteria.
Zhaojie Ju, Honghai Liu 0001
IEEE Trans. Fuzzy Syst.1
2010 Applying fuzzy EM algorithm with a fast convergence to GMMs
abstract
Inspired from the mechanism of Fuzzy C-means (FCMs) which introduces a degree of fuzziness on the dissimilarity function based on distances, a fuzzy Expectation Maximization (EM) algorithm for Gaussian Mixture Models (GMMs) is proposed in this paper. In the fuzzy EM algorithm, the dissimilarity function is defined as the multiplicative inverse of probability density function. Different from FCMs, the defined dissimilarity function is based on the exponential function of the distance. The fuzzy EM algorithm is compared with normal EM algorithm in terms of fitting degree and convergence speed. The experimental results in modeling random data and various characters demonstrate the ability of the proposed algorithm in reducing the computational cost of GMMs.
Zhaojie Ju, Honghai Liu 0001
FUZZ-IEEE1
2010 Human hand motion recognition using Empirical Copula
abstract
Programming by Demonstration (PbD) enables robotic hands to learn human manipulation skills through storing motion primitives and recognizing motion types. In this paper, Empirical Copula is introduced to recognize dynamic human hand motions for the first time using the proposed motion template and matching algorithm. The huge computational cost of Empirical Copula is alleviated by the proposed re-sampling processing. The experiments with human hand motions including grasps and in-hand manipulations demonstrate Empirical Copula outperforms the Time Clustering (TC) method, Gaussian Mixture Models (GMMs) and Hidden Markov Models (HMMs) in terms of recognition rate. In addition, Empirical Copula is also proved to be able to recognize different motions from different subjects.
Zhaojie Ju, Honghai Liu 0001
IROS1
2009 A switching fuzzy control method for the magnetic active suspension system
abstract
A switching fuzzy control system is proposed for dealing with the non-linear dynamics of electromagnetic suspension system. With two fuzzy sub-controllers and a switch engine, the switching fuzzy control system is flexible to cover changeable initial conditions with less computational cost. For satisfying the coupling constraint on positions of four floaters, a global self-supervisor with feedback structure is designed to control all four fuzzy subsystems. Simulations on the magnetic suspension with three different initial position settings demonstrated the efficiency of proposed method.
Jiangtao Cao, Zhaojie Ju, Xiaofei Ji, Honghai Liu 0001
FUZZ-IEEE2
2009 Fast estimating data dependence structure via fuzzy empirical copula
abstract
As a non-parametric algorithm, empirical copula is an effective way to estimate the dependence structure of high-dimension arbitrarily distributed data. However, it suffers from the problem of huge computation time because of its high computational complexity. In this paper, fuzzy empirical copula is proposed to solve this problem by combining the fuzzy clustering by local approximation of memberships (FLAME) with empirical copula. In the proposed algorithm, FLAME is extended from two-dimension data to high-dimension data and FLAME+is implemented to identify the highest density objects which represent the original dataset, and then empirical copula is used to estimate its independence structure according to the new dataset. Case studies have been carried out to demonstrate the effectiveness of the fuzzy empirical copula.
Zhaojie Ju, Honghai Liu 0001, Youlun Xiong
FUZZ-IEEE1