Honghai Liu 0001

dblp:10/4601 · DBLP profile ↗
← Back
206ranked-venue papers
14as first author
72since 2021 · last 2026
0000-0002-2880-4698ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 111 · 9 first-author · 32 since 2021Applied, interdisciplinary, general and emerging computing · 55 · 5 first-author · 21 since 2021Human-computer interaction and ubiquitous computing · 44 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 1 first-author · 16 since 2021Systems, architecture and hardware · 10 · 5 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 HKAFER: Achieve Visual Parameter-Efficient Fine-Tuning via Heterogeneous Kronecker Adaptation for Facial Expression Recognition
abstract
Facial Expression Recognition (FER) seeks to classify affective states from facial images, which remains a challenging problem due to variations in real-world conditions. FER task becomes particularly complex when handling unconstrained environments characterized by partial occlusions, different head poses, and so on. To address the above problems, current approaches rely on extensive learnable parameters and complex model architectures, which inevitably lead to overfitting and cause the FER model to focus on non-discriminative facial regions. In this work, we propose an HKAFER model that can adaptively enhance visual expression representations through efficiently fine-tuning the image encoder in large Visual Foundation Models (VFMs) and Vision-Language Models (VLMs). Specifically, we establish Heterogeneous Kronecker Adaptation (HeKA), which consists of multi-scale adapters based on Kronecker product in a parallel manner, offering significantly diverse subspaces to learn the incremental matrices. Besides, we also propose Dual-Branch Interactive Router (DBIR) to dynamically assign the weights of adapters, which promotes collaboration and information flow among them. In this way, our HKAFER can effectively capture robust spatial features and the regional associations. Experimental results demonstrate that our proposed model not only outperforms state-of-the-art methods on several FER benchmarks but also uses significantly fewer trainable parameters.
Yu Gao 0010, Haoyu Ji 0001, Zhiyong Wang 0009, Wenze Huang, Xueting Liu 0009, Weihong Ren, Honghai Liu 0001
AAAI9
2026 Efficient speech command recognition leveraging spiking neural networks and progressive time-scaled curriculum distillation
Jiaqi Wang 0003, Liutao Yu, Liwei Huang, Chenlin Zhou, Han Zhang 0035, Zhenxi Song, Honghai Liu 0001, Min Zhang 0005, Zhengyu Ma, Zhiguo Zhang 0001
Neural Networks7
2026 Unsupervised Cross-Domain 3D Human Pose Estimation via Pseudo-Label-Guided Global Transforms
abstract
Existing 3D human pose estimation methods often suffer in performance, when applied to cross-scenario inference, due to domain shifts in characteristics such as camera viewpoint, position, posture, and body size. Among these factors, camera viewpoints and locations have been shown to contribute significantly to the domain gap by influencing the global positions of human poses. To address this, we propose a novel framework that explicitly conducts global transformations between pose positions in the camera coordinate systems of source and target domains. We start with a Pseudo-Label Generation Module that is applied to the 2D poses of the target dataset to generate pseudo-3D poses. Then, a Global Transformation Module leverages a human-centered coordinate system as a novel bridging mechanism to seamlessly align the positional orientations of poses across disparate domains, ensuring consistent spatial referencing. To further enhance generalization, a Pose Augmentor is incorporated to address variations in human posture and body size. This process is iterative, allowing refined pseudo-labels to progressively improve guidance for domain adaptation. Our method is evaluated on various cross-dataset benchmarks, including Human3.6M, MPI-INF-3DHP, and 3DPW. The proposed method outperforms state-of-the-art approaches and even outperforms the target-trained model.
Zhiyong Wang 0009, Xinyu Fan 0009, Amirhossein Dadashzadeh, Honghai Liu 0001, Majid Mirmehdi
IEEE Trans. Circuits Syst. Video Technol.5
2026 Topology-Motion Decoupling Framework With Textual Regularization for Skeleton-Based Temporal Action Segmentation
abstract
Skeleton-based temporal action segmentation aims to capture key information in long skeleton motion sequences to temporally segment and identify actions at a fine-grained level. Existing approaches have achieved promising results by improving the modeling of topological spatial relationships and long-term temporal dependencies. However, current methods often overlook the distinct nature of motion and topological information, applying a monolithic modeling paradigm to both. This approach fails to fully exploit their differential contributions to precise boundary localization and effective class discrimination. To address these limitations, we propose a novel Topology-Motion Decoupling Framework (TMD). Our framework incorporates three key designs. First, an auxiliary Differential Motion Perception Branch explicitly models the temporal gradients of skeletal trajectory to decouple boundary-sensitive motion features. Second, we introduce two effective fusion modules that integrate the complementary features from both branches for mutual enhancement. Finally, a Boundary-Aware Textual Regularization scheme leverages a dual set of semantic prompts for boundary/non-boundary to differentially guide the feature learning process. By design, TMD explicitly mitigates semantic and temporal confusion between actions, thereby enhancing inter-class discriminability and boundary awareness. Extensive experiments on five challenging public datasets demonstrate that our TMD achieves state-of-the-art performance.
Haoyu Ji 0001, Bowen Chen 0004, Wenze Huang, Weihong Ren, Zhiyong Wang 0009, Honghai Liu 0001
IEEE Trans. Circuits Syst. Video Technol.7
2026 Context Modeling With Multimodal Prompts for Emotion Recognition in Conversation
abstract
Emotion Recognition in Conversation (ERC) plays an important role in driving the development of human-machine interaction. After the extensive exploration of the text modality, visual and audio information has attracted considerable attention. Most of the existing approaches adopt either the attention mechanism or graph neural networks to conduct multi-modal fusion by directly utilizing pre-extracted single modal features, but they ignore the inherent priors and characteristics (highly related to emotions) contained in different modalities. For example, for visual modality, facial expression can well reveal a person's emotional state, and the intonation, speech rate and volume of audio modality also reflect emotional fluctuation. Thus, in this work, we aim to promote context fusion among multi-modal features by exploring inherent modal priors. Firstly, we create textual descriptions for different emotion categories belonging to different modalities, denoted as multimodal prompts which can be generated either from common sense or using large language models (e.g., ChatGPT). Then, we propose an adaptive gating fusion module which dynamically learns the weights between unimodal features and prompts, and allows the multimodal prompts to participate in encoding and enriching modal information. Finally, to facilitate multimodal fusion, we design a multimodal progressive encoder to learn inter-modal interactions among conversational utterances. Experimental results show that our model outperforms state-of-the-art models in ERC on two popular benchmark datasets.
Weihong Ren, Yu Gao 0010, Jianzhuang Liu, Honghai Liu 0001
IEEE Trans. Multim.5
2025 Neural Parameter Search for Slimmer Fine-Tuned Models and Better Transfer
abstract
Guodong Du, Zitao Fang, Jing Li, Junlin Li, Runhua Jiang, Shuyang Yu, Yifei Guo, Yangneng Chen, Sim Kuan Goh, Ho-Kin Tang, Daojing He, Honghai Liu, Min Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Guodong Du 0002, Zitao Fang, Jing Li 0034, Runhua Jiang, Shuyang Yu, Yifei Guo, Yangneng Chen, Sim Kuan Goh, Ho-Kin Tang, Daojing He, Honghai Liu 0001, Min Zhang 0005
ACL (1)12
2025 Speed Up Your Code: Progressive Code Acceleration Through Bidirectional Tree Editing
abstract
Longhui Zhang, Jiahao Wang, Meishan Zhang, GaoXiong Cao, Ensheng Shi, Mayuchi Mayuchi, Jun Yu, Honghai Liu, Jing Li, Min Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Longhui Zhang, Meishan Zhang, GaoXiong Cao, Ensheng Shi, Mayuchi Mayuchi, Jun Yu 0002, Honghai Liu 0001, Jing Li 0034, Min Zhang 0005
ACL (1)8
2025 InstructHOI: Context-Aware Instruction for Multi-Modal Reasoning in Human-Object Interaction Detection
abstract
Recently, Large Foundation Models (LFMs), e.g., CLIP and GPT, have significantly advanced the Human-Object Interaction (HOI) detection, due to their superior generalization and transferability. Prior HOI detectors typically employ single- or multi-modal prompts to generate discriminative representations for HOIs from pretrained LFMs. However, such prompt-based approaches focus on transferring HOI-specific knowledge, but unexplore the potential reasoning capabilities of LFMs, which can provide informative context for ambiguous and open-world interaction recognition. In this paper, we propose InstructHOI, a novel method that leverages context-aware instructions to guide multi-modal reasoning for HOI detection. Specifically, to bridge knowledge gap and enhance reasoning abilities, we first perform HOI-domain fine-tuning on a pretrained multi-modal LFM, using a generated dataset with 140K interaction-reasoning image-text pairs. Then, we develop a Context-aware Instruction Generator (CIG) to guide interaction reasoning. Unlike traditional language-only instructions, CIG first mines visual interactive context at the human-object level, which is then fused with linguistic instructions, forming multi-modal reasoning guidance. Furthermore, an Interest Token Selector (ITS) is adopted to adaptively filter image tokens based on context-aware instructions, thereby aligning reasoning process with interaction regions. Extensive experiments on two public benchmarks demonstrate that our proposed method outperforms the state-of-the-art ones, under both supervised and zero-shot settings.
Jinguo Luo, Weihong Ren, Quanlong Zheng, Zhenlong Yuan, Zhiyong Wang 0009, Haonan Lu, Honghai Liu 0001
NeurIPS8
2025 JADFER: Exploring Spatial-Contextual Interaction With Joint Attention Dropping for Facial Expression Recognition
abstract
Facial Expression Recognition (FER) aims to categorize emotional expressions depicted on a human face, and is a challenging task under unconstrained conditions, such as face occlusions and pose variations. Recent methods usually adopt self attention or cross attention to explore global or local relationships among different level features. However, these methods are inclined to focus on the redundant facial regions, causing model overfitting. To address this problem, we propose a new FER model named JADFER, which drops the joint attention in the weight matrix to adaptively enhance facial expression representations. Specifically, our JADFER model consists of three components: Spatial Branch (SB), Contextual Branch (CB), and Spatial-Contextual Interaction (SCI). First, SB runs$N$paths in parallel, where a Variety loss is designed to guide the paths of SB to focus on different discriminative regions. Meanwhile, CB abstracts the contextual facial representations using self attention with Joint Attention Dropping (JAD). Then, the SCI adopts the spatial features from SB to query the contextual representations from CB through cross attention with JAD, which regulates the attention weights by dropping the similar activations to further enhance the facial embeddings. Experimental results demonstrate that the proposed model outperforms the state-of-the-art methods on several FER benchmarks.
Yu Gao 0010, Weihong Ren, Weibo Jiang, Honghai Liu 0001
IEEE Trans. Affect. Comput.7
2025 Text Prompt Region Decomposition for Effective Facial Expression Recognition
abstract
Facial expressions are conveyed through semantically distinct visual cues distributed across different facial regions, making region-aware feature modeling essential for accurate Facial Expression Recognition (FER). However, existing methods typically rely on implicit attention mechanisms or manually defined region cues without grounding in semantically aligned supervision, which often leads to suboptimal representation of regional expression cues, increased risk of overfitting, and limited interpretability. To address these limitations, we propose a Text Prompt Region Decomposition (TPRD) network that explicitly disentangles expression-relevant features across key facial regions via text prompt guidance. Specifically, TPRD comprises a visual-language pretrained encoder (e.g., CLIP), a Region Decomposition Module (RDM), and a Regional Integration Module (RIM). The visual encoder extracts global visual features from full-face images, while the text encoder embeds region-specific prompts (e.g., “mouth”, “eye”) into semantic vectors within a shared visual-language embedding space. The RDM employs a multi-branch architecture to project global visual features onto the semantic directions of region-specific text embeddings, enabling explicit extraction of local visual features. Subsequently, the RIM models the interaction between regional and global features, adaptively generating distinct regional contributions that modulate the global representation for different facial expression samples. Experimental results demonstrate that our proposed TPRD achieves leading performance in both within- and cross-dataset evaluations, as well as in scenarios involving occlusion and large head pose variations.
Zhiyong Wang 0009, Honghai Liu 0001
IEEE Trans. Affect. Comput.5
2025 Facial Expression Monitoring via Fine-Grained Vision-Language Alignment
abstract
In the fields of health care and clinical monitoring, vision-based Facial Expression Recognition (FER) has achieved significantly progress, but it still faces the challenge of poor generalization ability under unconstrained conditions of occlusions and pose variation. Recently, Vision-Language Model (VLM) has greatly advanced the FER task. However, the existing VLM-based FER methods typically leverage a hard-crafted prompt (e.g., “a photo of [class]”) and only focus on the holistic semantic alignment, which may suffer from modal heterogeneity. In this work, we propose a fine-grained vision-language model via Prompt Masking for FER (PMFER). Specifically, for each expression, we first create fine-grained prompts using facial action units to guide the image encoder to learn discriminative representations. Further, to finely align text prompts and visual action units, we randomly drop a phrase description in the prompts and then predict the dropped phrase by conducting modal cross attention, implicitly promoting fine-grained vision-language alignment. In addition, we also design a modal-adversarial strategy to holistically eliminate the modal difference between visual and textual embeddings in a common latent space. Experimental results demonstrate that our PMFER model outperforms the state-of-the-art methods on several FER benchmarks, especially under the conditions of occlusions and pose variations. Note to Practitioners—Facial expression recognition is very important in health care and clinical monitoring, which provides an useful tool to assess the psychological and physiological conditions of patients. Although FER has made significant progress with the development of deep learning technologies, it still faces problems in the complex environments (e.g., occlusions and pose variations). To address the above issues, we propose a novel FER method in this work based on the recent vision-language model. It takes RGB image and text prompts as input and finally predicts the expression classification. Different from the existing methods, the proposed PMFER can enable fine-grained modal alignment for facial key units. Compared with the state-of-the-art methods on the public datasets, it can achieve better results, especially under the conditions of occlusions and pose variations. Also, we evaluate the proposed method on a real-world pain dataset, and the results demonstrate that PMFER has a good generalization and can be applied to health care.
Weihong Ren, Yu Gao 0010, Xi'ai Chen, Zhi Han, Zhiyong Wang 0009, Jiaole Wang, Honghai Liu 0001
IEEE Trans Autom. Sci. Eng.7
2025 Multiscale Skeleton-Based Temporal Action Segmentation Using Hierarchical Temporal Modeling and Prediction Ensemble
abstract
Skeleton-based temporal action segmentation (TAS) decomposes untrimmed skeleton sequence into meaningful segments. The variance in temporal scale challenges the skeleton modeling network to seek a balance between over-segmentation and under-segmentation. Current methods often rely on parallel multiscale feature extractors and additional refinement modules to mitigate the multiscale issue, which brings significant computations and complexity. To address these issues, this article proposes multiscale skeleton-based TAS (MSTAS), consisting of temporal probability pyramid (TPP) and smoothed multiscale ensemble (SME). TPP represents each action as a collection of multiscale probability distributions using a U-shape hierarchical temporal pyramid. Subsequently, SME takes the average of distributions instead of deploying additional refinement stages to achieve action segmentation. Considering the over-confident issue that exists in each scale, SME incorporates a novel label smoothing phase to improve the probability distributions by dynamically calibrating the confidence of each scale. Experimental results on four public datasets show that the MSTAS achieves state-of-the-art performance with less computation overheads, such as +1.1% accuracy and +2.8% [email protected] on the challenging LARa dataset with 70% fewer parameters and 80% fewer GFLOPS. Benefiting from confidence calibration, the MSTAS efficiently utilizes more temporal scales while keeping better calibration for ambiguous action instances. Additionally, the U-shape pyramid demonstrates a strong compatibility with classical refinement module, enabling the efficient extraction of multiscale motion representations.
Bowen Chen 0004, Haoyu Ji 0001, Weihong Ren, Qiyi Tong, Zhiyong Wang 0009, Honghai Liu 0001
IEEE Trans. Cybern.7
2025 Interaction-Aware Transformer Network for Human-Object Interaction Detection
abstract
human-object interaction (HOI) detection tackles the problem of joint localization and classification of HOIs. Recent HOI detection methods are mainly based on transformer networks, where the explicit priors at the object level (e.g., scene layout, object appearance, or category) are usually fed into the transformer to improve the object query ability. Though these methods have achieved remarkable results, they did not pay enough attention to the implicit action-level information, which is the fundamental element of HOI. In this work, we propose an interaction-aware transformer network (IATN) to obtain the interaction-aware query, by jointly utilizing implicit action-level priors and explicit object-level priors. Specifically, we design an action-aware module (AAM) to aggregate implicit action priors from the scene level and instance level, respectively. Then, we design an action-oriented graph (AOG), where human feature and object feature are graph nodes and action semantics represent graph edges, to aggregate priors jointly from action level and object level. Afterwards, the interaction-aware query is acquired and finally adopted to obtain the HOI predictions. Besides, we leverage knowledge distillation to enhance the action-level priors by transferring the final HOI predictions to the intermediate features. Extensive experiments on HICO-DET and V-COCO datasets verify the effectiveness of our proposed interaction-aware model.
Weibo Jiang, Weihong Ren, Jiandong Tian, Hanwei Ma, Bowen Chen 0004, Honghai Liu 0001
IEEE Trans. Cybern.6
2025 Self-Supervised Learning for Intuitive Control of Prosthetic Hand Movements via Sonomyography
abstract
As a primary effector of humans, the hand plays a crucial role in many aspects of daily life. Recognizing multidegree-of-freedom hand movements from muscle activity helps infer human motion intentions. Solving this problem has direct applications in prosthetic and exoskeleton control. Here, we propose a self-supervised learning algorithm inspired by muscle synergies to achieve simultaneous estimation of wrist rotation (supination/pronation) and hand grasp (open/close) from sonomyography-the muscle deformation detected by a wearable ultrasound array. Unlike conventional methods collecting both muscle activity and hand kinematics for supervised model calibration, this algorithm only uses unlabeled forearm ultrasound signals for self-supervised wrist and hand movement estimation, where movement labels are auto-generated. The performance of the proposed algorithm was experimentally evaluated with ten participants including an amputee. Offline analysis demonstrated that the proposed algorithm can accurately estimate simultaneous wrist rotation and hand grasp movements and were 0.98 and 0.94 for the able-bodied, and 0.98 and 0.90 for the amputee, respectively). Notably, the performance of the self-supervised learning was superior to the supervised learning for the amputee. Online experiments demonstrated that intended wrist and hand movements can be deciphered in real time, enabling accurate control of a virtual hand. This study will open up a new avenue for the sonomyographic human-machine interaction.
Xingchen Yang, Zongtian Yin, Yixuan Sheng, Dario Farina, Honghai Liu 0001
IEEE Trans. Cybern.5
2025 Dual-Modal Gesture Recognition Using Adaptive Weight Hierarchical Soft Voting Mechanism
abstract
Muscle force and morphology information offer complementary perspectives for gesture recognition and its applications. Surface Electromyography (sEMG) provides force and electrophysiological information associated with muscles, while A-mode ultrasound (AUS) reveals muscle morphological information. By leveraging these two modalities, more comprehensive muscle motor unit information relevant to gesture recognition can be obtained. In this article, we introduce the adaptive weight classification (AWC) module and its enhanced version with hierarchical classifiers, adaptive weight hierarchical soft voting (AWHSV), to integrate AUS and sEMG into a fused modality. This approach dynamically adjusts the weights of individual and fused features, compensating for lost details during fusion, leading to a richer information representation and significantly improving algorithm robustness in gesture recognition. The experimental results demonstrate that the proposed method achieves recognition rates that are 0.66%, 2.36%, and 1.30% higher than those of its counterparts using sEMG, AUS, and sEMG-AUS, respectively. Moreover, the method outperforms state-of-the-art approaches, confirming its effectiveness in gesture recognition across both single and multiple modalities. This work demonstrates the advantages of the proposed AWHSV method, providing broader application scenarios for gesture recognition.
Yue Zhang 0072, Sheng Wei 0006, Zheng Wang 0048, Honghai Liu 0001
IEEE Trans. Cybern.4
2025 Onboard Operational Safety Filter for a Quadrotor in an Environment With Dynamic Obstacles
abstract
Quadrotors have been applied to a wide range of industrial applications in recent years. It is vital to ensure the safety of a quadrotor operated by a human pilot in a practical environment that usually involves dynamic obstacles. This article develops an onboard Operational Safety Filter (OSF) for a quadrotor in an environment with dynamic obstacles. The developed OSF integrates motion prediction of dynamic obstacles and an improved backup controller into the existing backup controller-based OSF framework. The motion prediction of dynamic obstacles aims to address the dynamic obstacles. The improved backup controller can provide enhanced operational freedom for a quadrotor with respect to the existing backup controllers. The developed onboard OSF is applied to quadrotors in simulation and in the real world to evaluate its effectiveness in addressing dynamic obstacles, improving operational freedom, and running on a real quadrotor.
Weifeng Zeng, Hao Xiong 0004, Hantao Jiang, Bernd R. Noack, Wenjie Lu 0004, Honghai Liu 0001
IEEE Trans. Ind. Informatics7
2025 Text-Derived Relational Graph-Enhanced Network for Skeleton-Based Action Segmentation
abstract
Skeleton-based Temporal Action Segmentation (STAS) aims to segment and recognize various actions from long, untrimmed sequences of human skeletal movements. Current STAS methods typically employ spatio-temporal modeling to establish dependencies among joints as well as frames, and utilize one-hot encoding with cross-entropy loss for frame-wise classification supervision. However, these methods overlook the intrinsic correlations among joints and actions within skeletal features, leading to a limited understanding of human movements. To address this, we propose a Text-Derived Relational Graph-Enhanced Network (TRG-Net) that leverages prior graphs generated by Large Language Models (LLM) to enhance both modeling and supervision. For modeling, the Dynamic Spatio-Temporal Fusion Modeling (DSFM) method incorporates Text-Derived Joint Graphs (TJG) with channel- and frame-level dynamic adaptation to effectively model spatial relations, while integrating spatio-temporal core features during temporal modeling. For supervision, the Absolute-Relative Inter-Class Supervision (ARIS) method employs contrastive learning between action features and text embeddings to regularize the absolute class distributions, and utilizes Text-Derived Action Graphs (TAG) to capture the relative inter-class relationships among action features. Additionally, we propose a Spatial-Aware Enhancement Processing (SAEP) method, which incorporates random joint occlusion and axial rotation to enhance spatial generalization. Performance evaluations on four public datasets demonstrate that TRG-Net achieves state-of-the-art results.
Haoyu Ji 0001, Bowen Chen 0004, Weihong Ren, Wenze Huang, Zhiyong Wang 0009, Honghai Liu 0001
IEEE Trans. Image Process.7
2025 Synergistic Prompting Learning for Human-Object Interaction Detection
abstract
Human-Object Interaction (HOI) detection, as a foundational task in human-centric understanding, aims to detect interactive triplets in real-world scenarios. To better distinguish diverse HOIs within an open-world context, current HOI detectors utilize pre-trained Visual-Language Models (VLMs) to extract prior knowledge through textual prompts (i.e., descriptive texts for each HOI instance). However, relying on predetermined descriptive texts, such approaches only acquire a fixed set of textual knowledge for HOI prediction, consequently resulting in inferior performance and limited generalization. To remedy this, we propose a novel VLM-based method, which jointly performs prompting learning from both visual and textual perspectives and synergizes visual-textual prompting for HOI detection. Initially, we design a hierarchical adaptation architecture to perform progressive prompting: visual prompting is facilitated through gradual token migration from VLM's image encoder, while textual prompting is initialized with progressively leveled interaction descriptions. In addition, to synergize the visual-textual prompting learning, a text-supervising and image-tuning loop is introduced, in which the text-supervising stage guides visual prompting learning through contrastive learning and the image-tuning stage refines textual prompting by modal matching. Finally, we employ an interaction-aware knowledge merging mechanism to effectively transfer visual-textual knowledge encapsulated within synergistic prompting for HOI detection. Extensive experiments on two benchmarks demonstrate that our proposed method outperforms the state-of-the-art ones, under both supervised and zero-shot settings.
Jinguo Luo, Weihong Ren, Zhiyong Wang 0009, Xi'ai Chen, Huijie Fan, Zhi Han, Honghai Liu 0001
IEEE Trans. Image Process.7
2025 Iris Geometric Transformation Guided Deep Appearance-Based Gaze Estimation
abstract
The geometric alterations in the iris's appearance are intricately linked to the gaze direction. However, current deep appearance-based gaze estimation methods mainly rely on latent feature sharing to leverage iris features for improving deep representation learning, often neglecting the explicit modeling of their geometric relationships. To address this issue, this paper revisits the physiological structure of the eyeball and introduces a set of geometric assumptions, such as "the normal vector of the iris center approximates the gaze direction". Building on these assumptions, we propose an Iris Geometric Transformation Guided Gaze estimation (IGTG-Gaze) module, which establishes an explicit geometric parameter sharing mechanism to link gaze direction and sparse iris landmark coordinates directly. Extensive experimental results demonstrate that IGTG-Gaze seamlessly integrates into various deep neural networks, flexibly extends from sparse iris landmarks to dense eye mesh, and consistently achieves leading performance in both within- and cross-dataset evaluations, all while maintaining end-to-end optimization. These advantages highlight IGTG-Gaze as a practical and effective approach for enhancing deep gaze representation from appearance.
Zhiyong Wang 0009, Weihong Ren, Honghai Liu 0001
IEEE Trans. Image Process.5
2025 Early Screening of Autism in Toddlers via Express-Needs-With-Pointing Protocol
abstract
The incidence of autism spectrum disorders (ASD), a neurodevelopmental condition associated with challenges in social communication, has witnessed a remarkable surge in recent years, with adverse effects on individuals, families, and society at large. Early screening for autism ensures timely access to interventions, yet screening lacks systematic and methodical approaches for objectively quantifying social behaviors. In response to this, we propose a protocol for early assistive screening, termed the Express-Needs-with-Pointing (ENP), which employs a multi-sensor platform to quantify the one of the social skills of toddler. A vision-based pointing behavior detection method is proposed, combining gaze estimation and pointing estimation, where the pointing estimation integrates forearm orientation and finger direction. We conduct an experiment involving twenty toddlers aged between 16 and 32 months, 4 of whom are typically developing (TD) children, 6 diagnosed with ASD, 8 diagnosed with global developmental delay (GDD), and 5 diagnosed with language disorders (LD). The results demonstrate that the automated assessment methods for pointing behavior achieved an impressive accuracy rate of 93.9%. These findings provide compelling evidence that the ENP is one of the highly effective protocols and holds significant implications for assisting in early autism screening.
Zhiyong Wang 0009, Haibo Qin, Bingrui Zhou, Huiping Li 0004, Xiu Xu, Honghai Liu 0001
IEEE J. Biomed. Health Informatics9
2025 Exploring Eye-Tracking Based Biomarkers to Assess Cognitive Abilities in Autistic Children: A Feasibility Study
abstract
Cognitive assessment can reveal a person's cognitive processing and behavioral patterns, making it an indispensable component of autism intervention and prognosis. Existing machine-assisted cognitive assessment methods primarily focus on children's performance outcomes, overlooking distinctive behavioral models, particularly characteristics of eye movement behavior, which have been demonstrated as the most direct indicators of cognitive abilities. In this study, we explore eye-tracking biomarkers for assisting cognitive assessment through a series of meticulously designed multi-level human-computer interaction protocols, encompassing three cognitive abilities: pairing and categorization, emotion recognition, and social interaction. A platform embedded with an eye-tracking module has been developed to reliably collect and analyze eye movement data, even in the presence of unrestricted large head movements in children. Experimental results indicate that there are significant group differences between autism and typically developing children in the eye-tracking features of total fixation duration, response latency, time to first fixation, mean fixation duration, and visit count in the absence of significant intergroup differences in the Wechsler Preschool and Primary Scale of Intelligence (WPPSI) and Wechsler Intelligence Scale for Children (WISC) assessment results. In addition, certain eye-tracking features in each group are correlated with WPPSI/WISC scale scores, enabling clinical cognitive assessments within each group based on these eye movement features. This study suggests that using eye-tracking features as biomarkers to assist detailed cognitive assessments holds significant potential for the intervention and prognosis of autism.
Chunchun Hu, Zhiyong Wang 0009, Bingrui Zhou, Qinyi Ye, Ruihan Lin, Xiu Xu, Honghai Liu 0001
IEEE J. Biomed. Health Informatics10
2025 Cortico-Ocular Coupling Analysis for Developmental and Behavioral Disorders: A Review
abstract
Developmental and behavioral disorders (DBD) have a significant impact on children's neurological activity and behavioral performance. Early diagnosis and treatment are known to be beneficial for improving DBD outcomes, yet existing unimodal neurophysiological assessment tools for DBD yield significant heterogeneity in results, highlighting the urgent need for exploring novel assessment tools. Cortico-ocular coupling (COC) refers to the information interaction between the cerebral cortex and eyes, and COC analysis is a technique for quantitatively measuring the correlation of neural oscillations and eye movements as biomarkers for assessment and mechanism disclosure. This review focuses on COC analysis for DBD from four perspectives: neural substrates, research paradigms, analysis methods, and applications. First, this review provides a comprehensive overview of the neural substrates and evocation paradigms related to COC analysis, aiming at helping target brain region selection, experimental result analysis, and paradigm design. The neural substrates and evocation paradigms are categorized according to functional domains, including social functioning, attention, cognition, early visual processing, and motor function. Then, this review summarizes the EEG and eye-tracking features, the analysis methods, and the validation datasets involved in COC analysis, aiming at helping implement COC analysis. Next, this review presents the applications of COC analysis in DBD, proving the validity and advance of COC analysis. In the end, the limitations, challenges, and future directions of COC analysis are discussed.
Zhiyong Wang 0009, Chunchun Hu, Peilian Chi, Xiu Xu, Honghai Liu 0001
IEEE J. Biomed. Health Informatics6
2025 Online 4D Ultrasound-Guided Robotic Tracking Enables 3D Ultrasound Localization Microscopy With Large Tissue Displacements
abstract
Super-Resolution Ultrasound (SRUS) imaging through localising and tracking microbubbles, also known as Ultrasound localization Microscopy (ULM), has demonstrated reconstruction of microvascular structure and flow with sub-diffraction resolution, and its potential in a range of clinical applications. However, imaging organs with large tissue movements, such as those caused by respiration, presents substantial challenges. Existing methods often require breath holding to maintain accumulation accuracy, which limits data acquisition time and ULM image saturation. To improve image quality in the presence of large tissue movements, this study introduces an approach integrating high-frame-rate volumetric ultrasound with online precise robotic probe control. Tested on a microvasculature phantom with slow but large translation motions, up to 5 mm/s in speed and 20 mm in distance- twice the aperture size of the matrix array used, our method achieved real-time tracking of the moving phantom and imaging volume rate at 85 Hz, keeping majority of the target volume in the imaging field of view. ULM images of the moving cross channels in the phantom were successfully reconstructed in post-processing, demonstrating the feasibility of super-resolution imaging under large tissue motions. This represents a significant step towards ULM imaging of organs with large motion.
Jipeng Yan 0001, Qingyuan Tan, Shusei Kawara, Bingxue Wang, Matthieu Toulemonde, Honghai Liu 0001, Ying Tan 0001, Meng-Xing Tang
IEEE Trans. Medical Imaging7
2025 Lifelong-MonoDepth: Lifelong Learning for Multidomain Monocular Metric Depth Estimation
abstract
With the rapid advancements in autonomous driving and robot navigation, there is a growing demand for lifelong learning (LL) models capable of estimating metric (absolute) depth. LL approaches potentially offer significant cost savings in terms of model training, data storage, and collection. However, the quality of RGB images and depth maps is sensor-dependent, and depth maps in the real world exhibit domain-specific characteristics, leading to variations in depth ranges. These challenges limit existing methods to LL scenarios with small domain gaps and relative depth map estimation. To facilitate lifelong metric depth learning, we identify three crucial technical challenges that require attention: 1) developing a model capable of addressing the depth scale variation through scale-aware depth learning; 2) devising an effective learning strategy to handle significant domain gaps; and 3) creating an automated solution for domain-aware depth inference in practical applications. Based on the aforementioned considerations, in this article, we present 1) a lightweight multihead framework that effectively tackles the depth scale imbalance; 2) an uncertainty-aware LL solution that adeptly handles significant domain gaps; and 3) an online domain-specific predictor selection method for real-time inference. Through extensive numerical studies, we show that the proposed method can achieve good efficiency, stability, and plasticity, leading the benchmarks by 8%-15%. The code is available at https://github.com/FreeformRobotics/Lifelong-MonoDepth.
Junjie Hu 0003, Chenyou Fan, Liguang Zhou, Qing Gao 0002, Honghai Liu 0001, Tin Lun Lam
IEEE Trans. Neural Networks Learn. Syst.5
2025 Snippet-Aware Transformer With Multiple Action Elements for Skeleton-Based Action Segmentation
abstract
The skeleton-based temporal action segmentation (STAS) aims to densely segment and classify human actions within lengthy untrimmed skeletal motion sequences. Current methods primarily rely on graph convolutional networks (GCNs) for intraframe spatial modeling and temporal convolutional networks (TCNs) for interframe temporal modeling to discern motion patterns. However, these approaches often overlook the distinctive nature of essential action elements across various actions, including engaged core body parts and key subactions. This oversight limits the ability to distinguish different actions within a given sequence. To address these limitations, the snippet-aware Transformer with multiple action element (ME-ST) is proposed to enhance the discrimination and segmentation among actions, which leverages intrasnippet attention along joints and sequences to identify core joints and key subactions at different scales. Specifically, in terms of the spatial domain, the intrasnippet cross-joint attention (CJA) module divides the sequence into distinct snippets and computes attention to establish intricate joint semantic relationships, emphasizing the identification of core motion joints. In terms of the temporal domain, in the encoder, the intrasnippet cross-frame attention (CFA) module segments the sequence in a blockwise expansion manner and establishes interframe relationships to highlight the most discriminative frames. In the decoder, clip-level representations at various temporal scales are initially generated through an hourglass-like sampling process, followed by the intrasnippet cross-scale attention (CSA) module to integrate the key clip information across different time scales. The performance evaluation on five public datasets demonstrates that ME-ST achieves state-of-the-art (SOTA) performance.
Haoyu Ji 0001, Bowen Chen 0004, Wenze Huang, Weihong Ren, Zhiyong Wang 0009, Honghai Liu 0001
IEEE Trans. Neural Networks Learn. Syst.6
2025 On the Passive Virtual Viscous Element Injection Method for Elastic Joint Robots
abstract
Increasing the viscosity of elastic joints can significantly improve the performance of elastic joint robots during physical human–robot interactions. However, current approaches for injecting viscous elements require an additional damper to be added in parallel with the elastic elements. In this paper, we propose a new concept called virtual viscous element injection (VVI), which enables a robot to exhibit viscoelasticity without altering its mechanical structure. VVI relies only on motor-side dynamics reshaping and state feedback. Interestingly, the VVI method allows high-resolution joint torque measurements in elastic joint robots, unlike in physical viscoelastic joint robots, which measure joint torque using higher-order derivatives of the positions. Furthermore, the VVI method is proved to preserve the passivity of robot dynamics, which provides numerous possibilities for the applications of combined passivity-based controllers. Specifically, we first emphasize the impedance control method using VVI. The results demonstrate that the VVI-DF method, which combines the direct feedback (DF) method with VVI, addresses the issue of excessive acceleration feedback in the controller. This provides looser constraints for achieving a high-gain torque loop in impedance control. Moreover, this paper also provides examples of the application of VVI combined with passivity-based position and torque controllers. Experiments and simulations demonstrate the effectiveness of the proposed methods. The proposed method can be extended to various robots, such as exoskeletons, and collaborative robots.
Tengyu Hou, Ye Ding 0001, Bo Zhang 0049, Honghai Liu 0001
IEEE Trans. Robotics5
2024 Exploring Self- and Cross-Triplet Correlations for Human-Object Interaction Detection
abstract
Human-Object Interaction (HOI) detection plays a vital role in scene understanding, which aims to predict the HOI triplet in the form of . Existing methods mainly extract multi-modal features (e.g., appearance, object semantics, human pose) and then fuse them together to directly predict HOI triplets. However, most of these methods focus on seeking for self-triplet aggregation, but ignore the potential cross-triplet dependencies, resulting in ambiguity of action prediction. In this work, we propose to explore Self- and Cross-Triplet Correlations (SCTC) for HOI detection. Specifically, we regard each triplet proposal as a graph where Human, Object represent nodes and Action indicates edge, to aggregate self-triplet correlation. Also, we try to explore cross-triplet dependencies by jointly considering instance-level, semantic-level, and layout-level relations. Besides, we leverage the CLIP model to assist our SCTC obtain interaction-aware feature by knowledge distillation, which provides useful action clues for HOI detection. Extensive experiments on HICO-DET and V-COCO datasets verify the effectiveness of our proposed SCTC.
Weibo Jiang, Weihong Ren, Jiandong Tian, Liangqiong Qu, Zhiyong Wang 0009, Honghai Liu 0001
AAAI6
2024 CRED: A Corneal Reflection and Environment Dataset
abstract
Existing studies have proved that corneal reflection images can not only visualize the human surroundings, but also accurately reflect the attention information of the eyes to the environment, which promotes the research and application of visual tracking and human posture localization in the field of human-computer interaction. The cornea is a small and transparent reflective surface with weak reflective ability, and its reflected images always have dull colors and low resolution. Some researches try to obtain clearer and brighter corneal reflection images when a person is facing a screen or outdoors, but the reflected images are highly susceptible to the interference of iris color and texture. However, strong corneal reflections are highly susceptible to obscuring the iris and pupil regions, affecting the accuracy of gaze tracking. In addition, a large number of reflected images interfered by the iris must rely on iris features for image enhancement. These two limitations make it difficult to directly apply eye images taken outdoors. We try to propose a corneal reflection and human eye surroundings dataset, CRED, which contains not only segmented images of human eye images and ocular structures (e.g., iris, pupil, and eyelid margins) with significant corneal reflections, but also corneal reflections and ground truth of the human eye surroundings scene separated from the ocular images. We believe that with the help of the CRED dataset, a large number of deep learning-based end-to-end works can be performed for iris and pupil position estimation and localization in the presence of strong corneal reflection interference. Similarly, the clarity and usability of corneal reflection images will be significantly improved.
Mengqi Du, Yue Zhang 0072, Jianhua Zhang 0002, Honghai Liu 0001
CSCWD4
2024 Discovering Syntactic Interaction Clues for Human-Object Interaction Detection
abstract
Recently, Vision-Language Model (VLM) has greatly ad-vanced the Human-Object Interaction (HOI) detection. The existing VLM-based HOI detectors typically adopt a hand-crafted template (e.g., a photo of a person [action] a/an [object]) to acquire text knowledge through the VLM text encoder. However, such approaches, only encoding the action-specific text prompts in vocabulary level, may suffer from learning ambiguity without exploring the fine-grained clues from the perspective of interaction context. In this paper, we propose a novel method to discover Syntactic Interaction Clues for HOI detection (SICHOI) by using VLM. Specifically, we first investigate what are the essen-tial elements for an interaction context, and then establish a syntactic interaction bank from three levels: spatial relationship, action-oriented posture and situational condition. Further, to align visual features with the syntactic interaction bank, we adopt a multi-view extractor to jointly aggre-gate visual features from instance, interaction, and image levels accordingly. In addition, we also introduce a dual cross-attention decoder to perform context propagation be-tween text knowledge and visual features, thereby enhancing the HOI detection. Experimental results demonstrate that our proposed method achieves state-of-the-art performance on HICO-DET and V-COCO.
Jinguo Luo, Weihong Ren, Weibo Jiang, Xi'ai Chen, Qiang Wang 0015, Zhi Han, Honghai Liu 0001
CVPR7
2024 Language-Assisted Skeleton Action Understanding for Skeleton-Based Temporal Action Segmentation
Haoyu Ji 0001, Bowen Chen 0004, Xinglong Xu, Weihong Ren, Zhiyong Wang 0009, Honghai Liu 0001
ECCV (54)6
2024 BNMTrans: A Brain Network Sequence-Driven Manifold-Based Transformer for Cognitive Impairment Detection Using EEG
abstract
Identifying mild cognitive impairment (MCI) is vital for Alzheimer’s disease prevention. As neurodegenerative diseases progress, synchronous activity in electroencephalography (EEG) - indicating functional connectivity - changes due to neural system deterioration. Thus, developing geometric learning to decode the functional brain structure is essential. Techniques such as graph neural networks and Riemannian manifolds show potential in analyzing non-Euclidean data. However, existing approaches neglect to combine synchronous activity with temporal dependence and still remain insufficient for MCI detection. This paper proposes the Brain Network sequence-driven Manifold-based Transformer (BNMTrans) to identify MCI patterns from EEG data. BNMTrans leverages its strengths by extracting features from sequential brain networks through the self-attention mechanism, guided by the geometric correlations within the Riemannian manifold. By integrating long-term temporal dynamics and structural relationships within manifold space based on functional connectivity, this approach outperforms others in EEG feature comparisons and state-of-the-art evaluations based on clinical data from 89 subjects (46 MCI, 43 healthy controls) at a local hospital. Our work has significance for both MCI clinical management and technical progression in the EEG field.
Ruihan Qin, Zhenxi Song, Huixia Ren, Zian Pei, Xue Shi, Yi Guo 0007, Honghai Liu 0001, Min Zhang 0005, Zhiguo Zhang 0001
ICASSP8
2024 Fusing Multi-Level Features from Audio and Contextual Sentence Embedding from Text for Interview-Based Depression Detection
abstract
Automatic depression detection based on audio and text representations from participants’ interviews has attracted widespread attention. However, most of previous researches only used one type of feature of one single modality for depression detection, so that the rich information of audio and text from interviews has not been fully utilized. Moreover, an effective multi-modal fusion approach to leverage the independence among audio and text representations is still lacking. To address these problems, we propose a multi-modal fusion depression detection model based on the interaction of multilevel audio features and text sentence embedding. Specifically, we first extract Low-Level Descriptors (LLDs), mel-spectrogram features, and wav2vec features from the audio. Then we design a Multi-level Audio Features Interaction Module (MAFIM) to fuse these three levels of features for a comprehensive audio representation. For interview text, we use pre-trained BERT to extract sentence-level embedding. Further, to effectively fuse audio and text representations, we design a Channel Attention-based Multi-modal Fusion Module (CAMFM) by taking into account the independence and correlation between two different modalities. Our proposed model shows better performance on two datasets, DAIC-WOZ and EATD-Corpus, than existing methods, so it has a high potential to be applied for interview-based depression detection in practice.
Junqi Xue, Ruihan Qin, Xinxu Zhou, Honghai Liu 0001, Min Zhang 0005, Zhiguo Zhang 0001
ICASSP4
2024 EmoTVR: A Hybrid Model to Estimate Continuous-Time and Continuous-Level Emotion from Electroencephalography
abstract
Emotion recognition from electroencephalography (EEG) has attracted widespread interest, but few studies have considered estimating the highly dynamic trajectories of emotion in a relatively long period, such as video watching. To address this problem, we first recruit participants to assign continuous-time and continuous-level emotion labels to videos from the SEED corpus. Then, we propose a hybrid model, namely Emotion Time-Varying Regression (EmoTVR), to estimate continuous-time and continuous-level emotion using EEG spatial-temporal representations. EmoTVR combines atrous convolutional networks for spatial feature extraction and a temporal self-attentive regressor using attention-based long short-term memory for temporal feature extraction and continuous estimation. Moreover, EmoTVR adopts the Domain Adversarial Neural Network to address the problem of individual difference. Experimental results through both within-subject and cross-subject cross-validations demonstrate the superiority of EmoTVR in recognizing dynamic and continuous emotion over traditional methods. The proposed EmoTVR method caters for the needs of dynamic and continuous emotion recognition in naturalistic conditions, so it is highly potential for practical applications of emotion recognition.
Xinxu Zhou, Weishan Ye, Junqi Xue, Honghai Liu 0001, Min Zhang 0005, Zhiguo Zhang 0001
ICASSP5
2024 MLPER: Multi-Level Prompts for Adaptively Enhancing Vision-Language Emotion Recognition
abstract
In the field of robotics, vision-based Emotion Recognition (ER) has achieved significant progress, but it still faces the challenge of poor generalization ability under unconstrained conditions (e.g., occlusions and pose variations). In this work, we propose MLPER model, which introduces Vision-Language Model for Emotion Recognition to learn discriminative representations adaptively. Specifically, different from typically leveraging a hand-crafted prompt (e.g., "a photo of a [class] person"), we first establish Multi-Level Prompts from three aspects: facial expression, human posture and situational condition using large language models, like ChatGPT. Correspondingly, we extract the visual tokens from three levels: the face, body, and context. Further, to achieve fine-grained alignment at each level, we adopt textual tokens from the positive and the hard negative to query visual tokens, predicting whether a pair of image and text is matched. Experimental results demonstrate that our MLPER model outperforms the state-of-the-art methods on several ER benchmarks, especially under the conditions of occlusions and pose variations.
Yu Gao 0010, Weihong Ren, Xinglong Xu, Zhiyong Wang 0009, Honghai Liu 0001
IROS6
2024 GroupTrack: Multi-Object Tracking by Using Group Motion Patterns
abstract
The main challenge of Multi-Object Tracking (MOT) lies in maintaining a distinctive identity for each target in dense crowds or occluded scenarios. Although the existing methods have achieved significantly progress by using robust object detectors or complex association strategies, they cannot effectively solve long-term tracking due to individually motion or appearance modeling for each single target. In this paper, we propose a novel 2D MOT tracker GroupTrack, to learn reliable motion state for each target using group motion patterns. Specifically, for each tracklet, we first choose its neighboring ones to form a group of motion patterns, which can provide informative clues for the motion estimation of the current tracklet. Then, we apply the group motion patterns to perform tracklet prediction and data association. By integrating prior from neighboring motion patterns into the data association process, GroupTrack provides a new paradigm for target motion modeling in extremely crowded and occluded scenarios. Through extensive experiments on the public MOT17 and MOT20 datasets, we demonstrate the effectiveness of our approach in challenging scenarios and show state-of-the-art performance at various MOT metrics.
Xinglong Xu, Weihong Ren, Gan Sun, Haoyu Ji 0001, Yu Gao 0010, Honghai Liu 0001
IROS6
2024 Automatic Recognition of Social Engagement for Children with Autism Spectrum Disorder
abstract
Estimating children's engagement levels improves their understanding of their social behaviors, since they can reflect their devotion to social interaction with others. This paper proposes an automatic method to recognize children's engagement levels in a triadic social interaction context. First, an overall metric function containing behavior, cognition, and affective dimensions is proposed to estimate children's multidimensional engagement levels. Then, the automatic feature extraction method based on gaze estimation, facial expression recognition, pose estimation, and object recognition models is illustrated to extract features to compute the engagement levels. Videos of 24 children, including 13 children with autism spectrum disorder (ASD), in triadic social interaction were collected for the engagement recognition experiment and cross-group analysis. The experimental results validate the effectiveness of the proposed automatic feature extraction method compared to human observations. Cross-group analyses revealed significant differences in affective engagement between children with ASD and typical developmental (TD) children.
Zhiyong Wang 0009, Xiu Xu, Honghai Liu 0001
SMC7
2024 Dual Transducers Co-Focusing Method for Ultrasound Stimulation in Calf Peripheral Nervous System: A Feasibility Study
abstract
As a neuromodulation technology, ultrasound stimulation exhibits great potential due to its non-invasive and targeted nature. Low-Intensity Focused Ultrasound (LIFU) allows for precise targeting of the peripheral nervous system, enabling the elicitation of diverse sensations in the human body, including tactile, cold, heat, and pain. Current research efforts predominantly focus on ultrasound stimulation of the upper limbs, specifically the fingers. However, there is a paucity of investigation pertaining to the stimulation of the peripheral nervous system in the lower limb of the human body. Furthermore, the utilization of a singularly focused transducer for stimulation introduces anisotropy, thereby rendering it challenging to enhance the axial resolution. This study aims to explore an innovative approach using dual transducers that synergistically focus ultrasound to achieve highly precise stimulation and reduce the disparity between axial and lateral resolution. The primary objective is to investigate the feasibility of utilizing this approach for deep muscle stimulation in the human calf muscles, thus delving into the realm of lower limb neuromodulation.
Honghai Liu 0001
SMC4
2024 Training Object Detectors from Scratch: An Empirical Study in the Era of Vision Transformer
abstract
Abstract Modeling in computer vision has long been dominated by convolutional neural networks (CNNs). Recently, in light of the excellent performance of self-attention mechanism in the language field, transformers tailored for visual data have drawn significant attention and triumphed over CNNs in various vision tasks. These vision transformers heavily rely on large-scale pre-training to achieve competitive accuracy, which not only hinders the freedom of architectural design in downstream tasks like object detection, but also causes learning bias and domain mismatch in the fine-tuning stages. To this end, we aim to get rid of the “pre-train and fine-tune” paradigm of vision transformer and train transformer based object detector from scratch. Some earlier works in the CNNs era have successfully trained CNNs based detectors without pre-training, unfortunately, their findings do not generalize well when the backbone is switched from CNNs to a vision transformer. Instead of proposing a specific vision transformer based detector, in this work, our goal is to reveal the insights of training vision transformer based detectors from scratch. In particular, we expect those insights to help other researchers and practitioners, and inspire more interesting research in other fields, such as remote sensing, visual-linguistic pre-training, etc. One of the key findings is that both architectural changes and more epochs play critical roles in training vision transformer based detectors from scratch. Experiments on the MS COCO dataset demonstrate that vision transformer based detectors trained from scratch can also achieve similar performance to their counterparts with ImageNet pre-training.
Weixiang Hong 0001, Wang Ren, Jiangwei Lao, Lele Xie, Liheng Zhong, Jian Wang 0108, Jingdong Chen, Honghai Liu 0001
Int. J. Comput. Vis.8
2024 Multimodal Emotion Recognition for Children with Autism Spectrum Disorder in Social Interaction
abstract
Autism Spectrum Disorders (ASD) remain a healthcare challenge and gain considerable attention due to the increasing prevalence rates and insupportable burden on families and society. It is noted that the recognition of children’s emotional states plays an important role in the evaluation and intervention process of ASD. In this paper, we aim to address the problem of automatic recognition of the emotional states of ASD children in social interactive scenarios. Since the child can be unconstrained in realistic scenarios, the face occlusion under pose variations and uncertain backgrounds become challenges of this task. To tackle this problem, we employ both facial expressions as well as body poses as cues to recognize the emotional states while most traditional methods only leverage the former. Firstly for the facial information, spatial features are extracted through convolutional neural networks followed by a temporal transformer to extract temporal information. Then for the body pose information, graph convolutional networks combined with the self-attention part are used to represent spatial features and temporal convolutional layers for temporal counterparts. Finally, different multimodal fusion ways are explored to generate final recognition results. We evaluate this method on a challenging database collected by us in real-world child-clinician interactive scenarios and the proposed method achieved significantly better results than baselines using only facial information. Thus it is suggested that there is a potential to assist in clinical practice by providing the recognized emotion as feedback.
Zhiyong Wang 0009, Bingrui Zhou, Jingxin Deng, Xiu Xu, Honghai Liu 0001
Int. J. Hum. Comput. Interact.10
2024 Unsupervised Time-Aware Sampling Network With Deep Reinforcement Learning for EEG-Based Emotion Recognition
abstract
Recognizing human emotions from complex, multivariate, and non-stationary electroencephalography (EEG) time series is essential in affective brain-computer interface. However, because continuous labeling of ever-changing emotional states is not feasible in practice, existing methods can only assign a fixed label to all EEG timepoints in a continuous emotion-evoking trial, which overlooks the highly dynamic emotional states and highly non-stationary EEG signals. To solve the problems of high reliance on fixed labels and ignorance of time-changing information, in this paper we propose a time-aware sampling network (TAS-Net) using deep reinforcement learning (DRL) for unsupervised emotion recognition, which is able to detect key emotion fragments and disregard irrelevant and misleading parts. Specifically, we formulate the process of mining key emotion fragments from EEG time series as a Markov decision process and train a time-aware agent through DRL without label information. First, the time-aware agent takes deep features from a feature extractor as input and generates sample-wise importance scores reflecting the emotion-related information each sample contains. Then, based on the obtained sample-wise importance scores, our method preserves top-Xcontinuous EEG fragments with relevant emotion and discards the rest. Finally, we treat these continuous fragments as key emotion fragments and feed them into a hypergraph decoding model for unsupervised clustering. Extensive experiments are conducted on three public datasets (SEED, DEAP, and MAHNOB-HCI) for emotion recognition using leave-one-subject-out cross-validation, and the results demonstrate the superiority of the proposed method against previous unsupervised emotion recognition methods. The proposed TAS-Net has great potential in achieving a more practical and accurate affective brain-computer interface in a dynamic and label-free circumstance. The source code is made available athttps://github.com/infinite-tao/TAS-Net.
Yue Pan 0010, Min Zhang 0005, Linling Li, Li Zhang 0041, Honghai Liu 0001, Zhiguo Zhang 0001
IEEE Trans. Affect. Comput.9
2024 Learning Self- and Cross-Triplet Context Clues for Human-Object Interaction Detection
abstract
Human-Object Interaction (HOI) detection aims to infer interactions between humans and objects, and it is very important for scene analysis and understanding. The existing methods usually focus on exploring instance-level (e.g., object appearance) or interaction-level (e.g., action semantic) features to conduct interaction prediction. However, most of these methods only consider the self-triplet feature aggregation, which may lead to learning ambiguity without exploring the cross-triplet context exchange. In this paper, from both visual and textual perspectives, we propose a novel method to jointly explore self-and cross-triplet interaction context clues for HOI detection. First, we employ a graph neural network to perform self-triplet aggregation, where human and object features represent graph nodes and visual interaction feature and textual prior knowledge are acted as two different edges. Furthermore, we also attempt to explore cross-triplet context exchange by incorporating symbiotic and layout relationships among different HOI triplets. Extensive experiments on two benchmarks demonstrate that our proposed method outperforms the state-of-the-art ones and achieves the impressive performance of 40.32 mAP on HICO-DET and 69.1 mAP on V-COCO datasets, respectively.
Weihong Ren, Jinguo Luo, Weibo Jiang, Liangqiong Qu, Zhi Han, Jiandong Tian, Honghai Liu 0001
IEEE Trans. Circuits Syst. Video Technol.7
2024 Gaze Estimation by Attention-Induced Hierarchical Variational Auto-Encoder
abstract
Appearance-based gaze estimation has been widely studied recently with promising performance. The majority of appearance-based gaze estimation methods are developed under the deterministic frameworks. However, the deterministic gaze estimation methods suffer from large performance drop upon challenging eye images in low-resolution, darkness, partial occlusions, etc. To alleviate this problem, in this article, we alternatively reformulate the appearance-based gaze estimation problem under a generative framework. Specifically, we propose a variational inference model, that is, variational gaze estimation network (VGE-Net), to generate multiple gaze maps as complimentary candidates simultaneously supervised by the ground-truth gaze map. To achieve robust estimation, we adaptively fuse the gaze directions predicted on these candidate gaze maps by a regression network through a simple attention mechanism. Experiments on three benchmarks, that is, MPIIGaze, EYEDIAP, and Columbia, demonstrate that our VGE-Net outperforms state-of-the-art gaze estimation methods, especially on challenging cases. Comprehensive ablation studies also validate the effectiveness of our contributions. The code will be publicly released.
Guanhe Huang, Jingyue Shi, Jun Xu 0019, Jing Li 0027, Shengyong Chen, Yingjun Du, Xiantong Zhen, Honghai Liu 0001
IEEE Trans. Cybern.8
2024 Dual Regression-Enhanced Gaze Target Detection in the Wild
abstract
Gaze is a vital feature in analyzing natural human behavior and social interaction. Existing gaze target detection studies learn gaze from gaze orientations and scene cues via a neural network to model gaze in unconstrained scenes. Though achieve decent accuracy, these studies either employ complex model architectures or leverage additional depth information, which limits the model application. This article proposes a simple and effective gaze target detection model that employs dual regression to improve detection accuracy while maintaining low model complexity. Specifically, in the training phase, the model parameters are optimized under the supervision of coordinate labels and corresponding Gaussian-smoothed heatmap labels. In the inference phase, the model outputs the gaze target in the form of coordinates as prediction rather than heatmaps. Extensive experimental results on within-dataset and cross-dataset evaluations on public datasets and clinical data of autism screening demonstrate that our model has high accuracy and inference speed with solid generalization capabilities.
Zhiyong Wang 0009, Weihong Ren, Xiu Xu, Honghai Liu 0001
IEEE Trans. Cybern.9
2024 FM-3DFR: Facial Manipulation-Based 3-D Face Reconstruction
abstract
3-D Morphable model (3DMM) has widely benefited 3-D face-involved challenges given its parametric facial geometry and appearance representation. However, previous 3-D face reconstruction methods suffer from limited power in facial expression representation due to the unbalanced training data distribution and insufficient ground-truth 3-D shapes. In this article, we propose a novel framework to learn personalized shapes so that the reconstructed model well fits the corresponding face images. Specifically, we augment the dataset following several principles to balance the facial shape and expression distribution. A mesh editing method is presented as the expression synthesizer to generate more face images with various expressions. Besides, we improve the pose estimation accuracy by transferring the projection parameter into the Euler angles. Finally, a weighted sampling method is proposed to improve the robustness of the training process, where we define the offset between the base face model and the ground-truth face model as the sampling probability of each vertex. The experiments on several challenging benchmarks have demonstrated that our method achieves state-of-the-art performance.
Shuwen Zhao, Dinghuang Zhang, Zhiyong Wang 0009, Honghai Liu 0001
IEEE Trans. Cybern.6
2024 Adaptive Monte Carlo Localization in Unstructured Environment via the Dimension Chain of Semantic Corners
abstract
This article investigates the relocalization of robots in indoor environments. To achieve this, indoor corners were classified into eight distinct categories, and a two-dimensional grid map was constructed using various sensors. Deep learning technology was employed to extract semantic information from corners, and the Bayesian method was used to build a semantic corner map incrementally. Additionally, the class attributes and positional relationships for each corner were explored, thereby facilitating the establishment of a dimension chain for semantic corners. Furthermore, a fast and efficient method was introduced for retrieving this dimension chain. During the relocalization process, the dimension chain of semantic corners was utilized for initial positioning, followed by the application of improved adaptive Monte Carlo localization (AMCL) algorithm for precise localization. Through comparative analysis using AMCL and several state-of-the-art methods, superior performance in both localization success rate and real-time implementation was demonstrated. Finally, extensive relocalization experiments were conducted to validate the effectiveness of the proposed method.
Yunfei Li 0001, Yufei Guo 0003, Honghai Liu 0001
IEEE Trans. Ind. Informatics6
2024 Computational Interpersonal Communication Model for Screening Autistic Toddlers: A Case Study of Response-to-Name
abstract
Interpersonal communication facilitates symptom measures of autistic sociability to enhance clinical decision-making in identifying children with autism spectrum disorder (ASD). Traditional methods are carried out by clinical practitioners with assessment scales, which are subjective to quantify. Recent studies employ engineering technologies to analyze children's behaviors with quantitative indicators, but these methods only generate specific rule-driven indicators that are not adaptable to diverse interaction scenarios. To tackle this issue, we propose a Computational Interpersonal Communication Model (CICM) based on psychological theory to represent dyadic interpersonal communication as a stochastic process, providing a scenario-independent theoretical framework for evaluating autistic sociability. We apply CICM to the response-to-name (RTN) with 48 subjects, including 30 toddlers with ASD and 18 typically developing (TD), and design a joint state transition matrix as quantitative indicators. Paired with machine learning, our proposed CICM-driven indicators achieve consistencies of 98.44% and 83.33% with RTN expert ratings and ASD diagnosis, respectively. Beyond outstanding screening results, we also reveal the interpretability between CICM-driven indicators and expert ratings based on statistical analysis.
Bingrui Zhou, Zhiyong Wang 0009, Bowen Chen 0004, Chunchun Hu, Huiping Li 0004, Xiu Xu, Honghai Liu 0001
IEEE J. Biomed. Health Informatics10
2024 Ultrasound as a Neurorobotic Interface: A Review
abstract
Neurorobotic devices, such as prostheses, exoskeletons, and muscle stimulators, can partly restore motor functions in individuals with disabilities, such as stroke, spinal cord injury (SCI), and amputations and musculoskeletal impairments. These devices require information transfer from and to the nervous system by neurorobotic interfaces. However, current interfacing systems have limitations of low-spatial and temporal resolution, and lack robustness, with sensitivity to, e.g., fatigue and sensor displacement. Muscle scanning and imaging by ultrasound technology has emerged as a neurorobotic interface alternative to more conventional electrophysiological recordings. While muscle ultrasound detects movement of muscle fibers, and therefore does not directly detect neural information, the muscle fibers are activated by neurons in the spinal cord and therefore their motions mirror the neural code sent from the spinal cord to muscles. In this view, muscle imaging by ultrasound provides information on the neural activation underlying movement intent and execution. Here, we critically review the literature on ultrasound applied as a neurorobotic interface, focusing on technological progresses and current achievements, machine learning algorithms, and applications in both upper-and lower-limb robotics. This critical review reveals that ultrasound in the human-machine interface field has evolved from bulky hardware to miniaturized systems, from multichannel imaging to sparse channel sensing, from simple muscle morphological analysis to input signal for musculoskeletal models and machine learning, from unimodal sensing to multimodal fusion, and from conventional statistical learning to deep learning. For future advances, we recommend exploring high-precision ultrasound imaging technology, improving the wearability and ergonomics of systems and transducers, and developing user-friendly real-time human-machine interaction models.
Xingchen Yang, Claudio Castellini, Dario Farina, Honghai Liu 0001
IEEE Trans. Syst. Man Cybern. Syst.4
2023 Fatigue Detection Based on Multiple Visual Features in Virtual Driving System
abstract
Fatigue driving poses a significant hazard, leading to numerous traffic accidents annually. However, the fatigue detection algorithm based on visual features still has the problem of low accuracy, and it is difficult to test extreme fatigue driving conditions. This study proposes a fatigue detection method that exploits multiple visual features, including facial landmarks, mouth aspect ratio (MAR), eye aspect ratio (EAR), PERCLOS and head pose. Furthermore, we introduce a novel approach by developing a virtual driving system dedicated to fatigue detection. This system offers a rich driving environment and enables the exploration of fatigue-related boundary conditions. To validate our method and assess the potential of the virtual driving system in detecting fatigue driving, we conduct an experiment involving five healthy adults using the aforementioned system. In our final results, the detection accuracy of mild fatigue reached 0.94, and the detection accuracy of severe fatigue reached 0.97. The results confirm both the feasibility of our approach and the promising prospects of the virtual driving system in fatigue detection.
Qiyi Tong, Zhiyong Wang 0009, Ruihan Lin, Honghai Liu 0001
IECON8
2023 Improving Stability of Gaze Target Detection in Videos
abstract
Obtaining accurate and stable results in gaze target detection is vital for the subsequent analysis of gaze meaning. However, existing image-based methods, which focus solely on enhancing accuracy, demonstrate poor stability when directly applied on videos. Especially when the video frame rate is low, even though the actual gaze target positions do not differ significantly between adjacent frames, the detected positions vary considerably. This inconsistency, stemming from the lack of temporal information, makes dynamic detection challenging and can lead to jarring outcomes. To reduce the jitter in gaze target detection in videos, we introduce an approach that integrates spatial and temporal modules to combine spatial with temporal information. Additionally, we propose a Jitter loss function to capture significant jitter and impose a strong penalty during training, which empowers our model with increased stability for dynamic detection. Based on a self-collected dataset, experiments demonstrate that our approach exhibits superior stability without compromising accuracy.
Zhiyong Wang 0009, Xiu Xu, Honghai Liu 0001
IECON6
2023 WSCFER: Improving Facial Expression Representations by Weak Supervised Contrastive Learning
abstract
The major challenge of Facial Expression Recog-nition (FER) is to learn class discriminative representations, and the existing works mainly address it by designing various classification networks from class level. However, learning representations at class level is limited due to the inconspicuous class discrimination among different facial expressions. Thus, in this paper, we propose a Weak Supervised Contrastive learning FER (WSCFER) method to improve facial expression representations by simultaneously learning instance-level representations which are highly complementary to the general class-level representations. Specifically, our proposed WSCFER consists of three components: a major task for FER classification, an auxiliary task for Weak Supervised Contrastive (WSC) learning which pulls augmented samples of the same image together while pushing apart instance samples from different classes, and a Partial Consistency Loss (PCL) for optimizing the two embedding spaces from both the class level and the instance level. We compare WSC with some state-of-the-art contrastive methods and find that it can efficiently learn instance-level representations but avoid overemphasizing irrelevant parts, which is crucial for FER. WSCFER achieves superior performance on several in-the-wild databases, and it also shows the promising potential for learning representations under noisy annotations.
Bowen Chen 0004, Xiu Xu, Weihong Ren, Honghai Liu 0001
IROS6
2023 Construction and Analysis of Deep Functional Corticomuscular Coupling Effect
abstract
Based on electroencephalogram (EEG) and electromyogram (EMG) signals, the function coupling between the cerebral cortex and muscles has been widely studied to evaluate the motor function and reveal various motor control and pathological mechanisms in healthy individuals or patients with movement disorders. However, the effect of the different signal sources on the functional corticomuscular coupling remains unclear. In this study, four different signal source combinations were constructed by EEG and high-density surface EMG (HD-sEMG) signals as well as their reconstructed source signals to analyze the corticomuscular coupling during isometric index finger contraction tasks at different levels of maximum voluntary contraction. A nonparametric coupling model was used to study the effect of deep source signals on changing the coherence magnitude indicators of corticomuscular coupling related to hand movements. The results showed that the reconstruction of EEG and HD-sEMG signals significantly improved the coherence peak and coherence strength under low-level force. However, as the force level increased, only the reconstructed brain source signals demonstrated a more substantial effect. In addition, the coherence peak was positively correlated with the finger force level. This study demonstrated the importance and positive impact of the reconstructed EEG and EMG signals for estimating corticomuscular coherence.
Jinbiao Liu, Manli Luo, Linqing Feng, Honghai Liu 0001, Yina Wei
SMC7
2023 Deep Depth Completion From Extremely Sparse Data: A Survey
abstract
Depth completion aims at predicting dense pixel-wise depth from an extremely sparse map captured from a depth sensor, e.g., LiDARs. It plays an essential role in various applications such as autonomous driving, 3D reconstruction, augmented reality, and robot navigation. Recent successes on the task have been demonstrated and dominated by deep learning based solutions. In this article, for the first time, we provide a comprehensive literature review that helps readers better grasp the research trends and clearly understand the current advances. We investigate the related studies from the design aspects of network architectures, loss functions, benchmark datasets, and learning strategies with a proposal of a novel taxonomy that categorizes existing methods. Besides, we present a quantitative comparison of model performance on three widely used benchmarks, including indoor and outdoor datasets. Finally, we discuss the challenges of prior works and provide readers with some insights for future research directions.
Junjie Hu 0003, Chenyu Bao, Mete Ozay, Chenyou Fan, Qing Gao 0002, Honghai Liu 0001, Tin Lun Lam
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Multi-Scale Attention Learning Network for Facial Expression Recognition
abstract
Facial Expression Recognition (FER) aims to identify emotional expressions in human faces, and it is a fundamental task in computer vision. Recently, some methods apply Vision Transformer (ViT) to FER and have achieved promising results. However, FER still suffers from two key issues: inter-class similarity and intra-class discrepancy. To address the issues, in this letter, we propose a Multi-Scale Attention Learning Network (MALN) based on ViT, which can learn facial expression embeddings in a multi-scale manner. Specifically, we adopt a multi-branch ViT architecture to adaptively explore multi-scale correlations without self-attention. Furthermore, we also design a Scale Distinction Loss (SDL) to dynamically regulate facial embeddings from multiple branches, which can guide ViT to capture discriminative facial regions. Experimental results on three public datasets (inluding RAF-DB, AffectNet and FERPlus) demonstrate the effectiveness of our proposed MALN for FER
Weihong Ren, Yu Gao 0010, Weibo Jiang, Honghai Liu 0001
IEEE Signal Process. Lett.5
2023 Deep EEG Superresolution via Correlating Brain Structural and Functional Connectivities
abstract
Electroencephalogram (EEG) excels in portraying rapid neural dynamics at the level of milliseconds, but its spatial resolution has often been lagging behind the increasing demands in neuroscience research or subject to limitations imposed by emerging neuroengineering scenarios, especially those centering on consumer EEG devices. Current superresolution (SR) methods generally do not suffice in the reconstruction of high-resolution (HR) EEG as it remains a grand challenge to properly handle the connection relationship amongst EEG electrodes (channels) and the intensive individuality of subjects. This study proposes a deep EEG SR framework correlating brain structural and functional connectivities (Deep-EEGSR), which consists of a compact convolutional network and an auxiliary fully connected network for filter generation (FGN). Deep-EEGSR applies graph convolution adapting to the structural connectivity amongst EEG channels when coding SR EEG. Sample-specific dynamic convolution is designed with filter parameters adjusted by FGN conforming to functional connectivity of intensive subject individuality. Overall, Deep-EEGSR operates on low-resolution (LR) EEG and reconstructs the corresponding HR acquisitions through an end-to-end SR course. The experimental results on three EEG datasets (autism spectrum disorder, emotion, and motor imagery) indicate that: 1) Deep-EEGSR significantly outperforms the state-of-the-art counterparts with normalized mean squared error (NMSE) decreased by 1%-6% and the improvement of signal-to-noise ratio (SNR) up to 1.2 dB and 2) the SR EEG manifests superiority to the LR alternative in ASD discrimination and spatial localization of typical ASD EEG characteristics, and this superiority even increases with the scale of SR.
Yunbo Tang, Dan Chen 0001, Honghai Liu 0001, Xiaoli Li 0002
IEEE Trans. Cybern.3
2023 A Multimodal Multilevel Converged Attention Network for Hand Gesture Recognition With Hybrid sEMG and A-Mode Ultrasound Sensing
abstract
Gesture recognition based on surface electromyography (sEMG) has been widely used in the field of human-machine interaction (HMI). However, sEMG has limitations, such as low signal-to-noise ratio and insensitivity to fine finger movements, so we consider adding A-mode ultrasound (AUS) to enhance the recognition impact. To explore the influence of multisource sensing data on gesture recognition and better integrate the features of different modules. We proposed a multimodal multilevel converged attention network (MMCANet) model for multisource signals composed of sEMG and AUS. The proposed model extracts the hidden features of the AUS signal with a convolutional neural network (CNN). Meanwhile, a CNN-LSTM (long-short memory network) hybrid structure extracts some spatial-temporal features from the sEMG signal. Then, two types of CNN features from AUS and sEMG are spliced and transmitted to a transformer encoder to fuse the information and interact with sEMG features to produce hybrid features. Finally, the classification results are output employing fully connected layers. Attention mechanisms are used to adjust the weights of feature channels. We compared MMCANet's feature extraction and classification performance with that of manually extracted sEMG-AUS features using four traditional machine-learning (ML) algorithms. The recognition accuracy increased by at least 5.15%. In addition, we tried deep learning (DL) methods with CNN on single modals. The experimental results showed that the proposed model improved 14.31% and 3.80% over the CNN method with single sEMG and AUS, respectively. Compared with some state-of-the-art fusion techniques, our method also achieved better results.
Sheng Wei 0006, Yue Zhang 0072, Honghai Liu 0001
IEEE Trans. Cybern.3
2022 An End-to-end Posture Perception Method for Soft Bending Actuators Based on Kirigami-inspired Piezoresistive Sensors
abstract
Posture sensing of soft actuators is critical for performing closed-loop control of soft robots. This paper presents a novel end-to-end posture perception method for soft actuators by developing long short-term memory (LSTM) neural networks. A novel flexible bending sensor developed from off-the-shelf conductive silicon material was proposed and used for posture sensing. In the proposed method, the hysteresis of the soft robot and non-linear sensing signals from the flexible bending sensors have also been considered. With one-step calibration from the sensor output, the posture of the soft actuator could be captured by the LSTM network. The method was validated on a finger-size one DOF pneumatic fiber-reinforced bending actuator. Four kirigami-inspired flexible piezoresistive transducers were placed on the top surface of the actuator. Results show that the transducers could sense the posture of the actuator with acceptable accuracy. We believe our work could benefit soft robot dynamic posture perception and closed-loop control.
Jing Shu, Junming Wang 0002, Yujie Su, Honghai Liu 0001, Zheng Li 0012, Raymond Kai-Yu Tong
BSN4
2022 A Swift Gaze Estimate Method Based On The Corneal Image System
abstract
With the development of intelligent manufacturing, the demand of the incoming Human-Machine Interaction such as the augment reality rapidly increasing. However, the existing interaction modes in the augment reality, rely heavily on the hands or head movement. The inflexible modes is inefficient in the busy work flow. In this paper, we propose a gaze estimation work based on the Corneal Image System which can improve the efficiency of the interaction. Several prior works have proved, single Corneal Image contains the subject’s gaze information. However, the quality of Corneal Image is always impacted by the color and texture of the iris or the light from the surrounding, is hard to be applied directly. In order to improve the quality of the Corneal Image, people usually import additional devices into their work, such as infrared camera or eye tracker. These extra devices cause their gaze estimation works to become cumbersome and hard to be re-implemented commonly. Our gaze estimation work requires no additional device, can be seamlessly integrated into the AR domain with the help of the AprilTag mark. An AprilTag mark contained in an eye image, is distinct enough to be recognized, meanwhile, owns the hybrid pose relationship information between the eye, camera, and the focused AprilTag mark. The gaze can be inferred through the rigid body coordinate transformation naturally from this relationship. Many experiments have demonstrated that our approach is much easier to be re-implemented than the previous Corneal Image System based gaze computing works, at the same time, have the near performance to the state of the art.
Mengqi Du, Kaiqi Chen 0001, Jianhua Zhang 0002, Honghai Liu 0001
CSCWD4
2022 AutoENP: An Auto Rating Pipeline for Expressing Needs via Pointing Protocol
abstract
Early screening for ASD (Autism Spectrum Disorder) is crucial and also challenging due to the limited medical resource. Expressing Needs with Pointing (ENP) is a low-cost yet effective protocol for early screening. However, the current methods need to manually trim video for analyzing ENP protocol, which is labour-intensive. Also, they detect discriminative signs with separately high-level clues (e.g., pose, object detection), but ignore the temporal action relationships between child and clinician, which usually leads to invalid detection. In contrast to previous approaches, we propose an Auto Rating Pipeline for Expressing Needs via Pointing Protocol, named AutoENP. Specifically, we introduce action segmentation into early screening, to capture temporal interaction relationships without manually intervention. To detect fine-grained hand motions, we fuse global, local and fine-grained features to fully understand the screening scene. Besides, we integrate focal loss and center loss to improve the detection accuracy for rare actions. To evaluate the proposed pipeline, we collected 22 ENP videos containing 7 actions with above 40,000 frames. Experimental results demonstrate that our model achieves 82.1% and 84.7% action accuracy for child and clinician, respectively. Moreover, 18 in 22 children’s ENP levels are reported correctly against the clinician’s diagnoses.
Bowen Chen 0004, Weihong Ren, Honghai Liu 0001, Huiping Li 0004, Xiu Xu, Bingrui Zhou
ICPR3
2022 Eye-hand Coordination based Properties for Social Capability Representation: a Case Study
abstract
This paper explores social capability properties via eye-hand coordination with a case study. In order to investigate social capability properties, an ICF based protocol is designed and experiments are conducted in a customized platform for investigating social capability for children with autism spectrum disorder. A set of eye-hand coordination based properties are extracted to represent individual social capability in terms of eye-hand coordination. A set of metric are also defined to measure the coordination behaviour such as engagement. It is evident that the eye-hand coordination based properties have huge potential for analyzing social capability, pave the way for machine-assisted intervention for ASD children.
Carrie M. Toptan, Dinghuang Zhang, Honghai Liu 0001
SMC3
2022 Appearance-Based Gaze Estimation for ASD Diagnosis
abstract
Biomarkers, such as magnetic resonance imaging (MRI) and electroencephalogram have been used to help diagnose autism spectrum disorder (ASD). However, the diagnosis needs the assist of specialized medical equipment in the hospital or laboratory. To diagnose ASD in a more effective and convenient way, in this article, we propose an appearance-based gaze estimation algorithm-AttentionGazeNet, to accurately estimate the subject's 3-D gaze from a raw video. The experimental results show its competitive performance on the MPIIGaze dataset and the improvement of 14.7% for static head pose and 46.7% for moving head pose on the EYEDIAP dataset compared with the state-of-the-art gaze estimation algorithms. After projecting the obtained gaze vector onto the screen coordinate, we apply accumulated histogram to taking into account both spatial and temporal information of estimated gaze-point and head-pose sequences. Finally, classification is conducted on our self-collected autistic children video dataset (ACVD), which contains 405 videos from 135 different ASD children, 135 typically developing (TD) children in a primary school, and 135 TD children in a kindergarten. The classification results on ACVD shows the effectiveness and efficiency of our proposed method, with the accuracy 94.8%, the sensitivity 91.1% and the specificity 96.7% for ASD.
Jing Li 0027, Zejin Chen, Yihao Zhong, Hak-Keung Lam, Junxia Han, Gaoxiang Ouyang, Xiaoli Li 0002, Honghai Liu 0001
IEEE Trans. Cybern.8
2022 Early Screening of Autism in Toddlers via Response-To-Instructions Protocol
abstract
Early screening of autism spectrum disorder (ASD) is crucial since early intervention evidently confirms significant improvement of functional social behavior in toddlers. This article attempts to bootstrap the response-to-instructions (RTIs) protocol with vision-based solutions in order to assist professional clinicians with an automatic autism diagnosis. The correlation between detected objects and toddler's emotional features, such as gaze, is constructed to analyze their autistic symptoms. Twenty toddlers between 16-32 months of age, 15 of whom diagnosed with ASD, participated in this study. The RTI method is validated against human codings, and group differences between ASD and typically developing (TD) toddlers are analyzed. The results suggest that the agreement between clinical diagnosis and the RTI method achieves 95% for all 20 subjects, which indicates vision-based solutions are highly feasible for automatic autistic diagnosis.
Zhiyong Wang 0009, Bin Ji 0004, Jingxin Deng, Xiu Xu, Honghai Liu 0001
IEEE Trans. Cybern.10
2022 Unsupervised Domain Adaptation for Gesture Identification Against Electrode Shift
abstract
Surface electromyogram (sEMG)-based hand gesture recognition, which interprets commands given by humans through sEMG signals, performs well in many studies. However, its recognition accuracy drops dramatically due to electrode shift since the distributions of motion classes are changed. Although calibrating the system with newly collected samples after electrode shift maintains the accuracy, collecting labeled samples is inconvenient and time-consuming since the procedure is rigid. However, the calibration may not work properly without label, especially when the change is significant. This study proposes a user friendly and convenient calibration method for hand gesture recognition by an unsupervised domain adaptation method, which only obtains the unlabeled samples of preselected benchmark classes from users in calibration. The change of benchmark classes is captured by unlabeled samples by a clustering method. The other classes are estimated based on the benchmark classes by regression models. As a result, the information of all classes is used to calibrate the system. Linear discriminant analysis is used to demonstrate our model. A dataset with ten subjects is collected to verify the performance empirically. Experimental results confirm that our method utilizes the unlabeled benchmark class samples in calibration and achieves 75.55% average accuracy. Our method is more robust to electrode shift and improves around 8.5% accuracy consistently on all subjects compared with the methods without calibration or label information in calibration. Although the accuracy of our method is slightly less than the ones using label calibration samples, our calibration data collection is more convenient and less complicated.
Patrick P. K. Chan, Qiu Xia Li, Yinfeng Fang, Linyi Xu, Kairu Li, Honghai Liu 0001, Daniel S. Yeung
IEEE Trans. Hum. Mach. Syst.6
2022 Toward Children's Empathy Ability Analysis: Joint Facial Expression Recognition and Intensity Estimation Using Label Distribution Learning
abstract
Empathy ability is one of the most important social communication skills in early childhood development. To analyze the children's empathy ability, facial expression analysis (FEA) is an effective way due to its ability to understand children's emotional states. Previous works mainly focus on recognizing the facial expression categories yet fail to estimate expression intensity, the latter of which is more important for fine-grained emotion analysis. To this end, this article first proposes to analyze children's empathy ability with both the categories and the intensities of facial expressions. A novel FEA method based on intensity label distribution learning is presented, which aims to recognize expression categories and estimate their intensity levels in an end-to-end framework. First, the intensity label distribution is generated for each frame in the expression sequence using a linear interpolation estimation and a Gaussian function to address the lack of reasonable annotations for expression intensity. Then, the extended intensity label distribution is presented to automatically encode the expression intensity in a multidimensional expression space, which aims to integrate the expression recognition and intensity estimation into a unified framework as well as boost the expression recognition performance by suppressing the variations in appearance caused by intensity and by emphasizing those variations among weak expressions. Finally, a Siamese-like convolutional neural network is presented to learn the expression model from a pair of frames that includes an expressive frame and its corresponding neutral frame using the extended intensity label distribution as the supervised information, thus effectively eliminating the expression-unrelated information's influence on FEA. Numerous experiments validate that the proposed method is promising in analysis of the differences in empathy ability between typically developing children and children with autism spectrum disorder.
Jingying Chen 0001, Ruyi Xu, Kun Zhang 0031, Zongkai Yang, Honghai Liu 0001
IEEE Trans. Ind. Informatics6
2022 A Wearable Ultrasound Interface for Prosthetic Hand Control
abstract
Ultrasound can non-invasively detect muscle deformations and has great potential applications in prosthetic hand control. Traditional ultrasound equipment was usually too bulky to be applied in wearable scenarios. This research presented a compact ultrasound device that could be integrated into a prosthetic hand socket. The miniaturized ultrasound system included four A-mode ultrasound transducers for sensing musculature deformations, a signal excitation/acquisition module, and a prosthetic hand control module. The size of the ultrasound system was 65*75*25 mm, weighing only 85 g. For the first time, we integrated the ultrasound system into a prosthetic hand socket to evaluate its performance in practical prosthetic hand control. We designed an experiment requiring twenty subjects to perform six commonly used gestures. The performance of decoding ultrasound signals was analyzed offline using four classification algorithms and then was assessed in online control. The average values of online classification accuracy with and without wearing the physical prosthetic were 91.5 [Formula: see text] and 96.5 [Formula: see text], respectively. We found that wearing the prosthetic hand influenced the ultrasound gestures classification accuracy, but remarkable online classification performance could still be maintained. These experimental results demonstrated the efficacy of the designed integrated ultrasound system for practical use, paving the way for an effective HMI system that could be widely used in prosthetic hand control.
Zongtian Yin, Hanwei Chen, Xingchen Yang, Yifan Liu 0006, Ning Zhang 0031, Jianjun Meng, Honghai Liu 0001
IEEE J. Biomed. Health Informatics7
2022 Fatigue-Sensitivity Comparison of sEMG and A-Mode Ultrasound based Hand Gesture Recognition
abstract
Though physiological signal based human-machine interfaces (HMIs) have recently developed rapidly, their practical use is restricted by many real-world environmental factors, one of which is muscle fatigue. This paper explores the sensitivities between surface electromyography (sEMG) and A-mode ultrasound (AUS) sensing modalities subject to muscle fatigue in the context of hand gesture recognition tasks. Two metrics, mean classification accuracy ( mCA) and decline rate ( DR), are proposed to evaluate the accuracy and muscle fatigue sensitivity between sEMG and AUS based HMIs. Muscle fatigue inducing experiment was designed and eight subjects were recruited to participate in the experiment. The gesture recognition accuracies of sEMG and AUS under non-fatigue state and fatigue state are compared through Mahalanobis distance based classifier linear discriminant analysis (LDA). In addition, Mahalanobis distance based metrics, repeatability index ( RI) and separability index ( SI), are introduced to evaluate the changes in the feature distribution during muscle fatigue and reveal the cause of the fatigue sensitivity difference between sEMG and AUS signals. The experimental results demonstrate that the fatigue robustness of AUS signal is better than that of sEMG signal. Specifically, with the employment of the LDA classifier trained under non-fatigue state, the testing accuracy of the sEMG signal on the non-fatigue state is 94.96%, while reduce to 68.26% on the fatigue state. The testing accuracy of the AUS signal on the corresponding states is 99.68% and 91.24% respectively. AUS signal attains higher mCA and lower DR, indicating that it has advantages over sEMG signal in terms of both accuracy and muscle fatigue sensitivity. In addition, the RI and RI/SI analysis reveal that before and after muscle fatigue, the consistency of AUS feature distribution is better than that of sEMG. These research outcomes validate that AUS is more tolerant to feature migration caused by muscle fatigue than sEMG.
Yu Zhou 0013, Yicheng Yang, Jipeng Yan 0001, Honghai Liu 0001
IEEE J. Biomed. Health Informatics5
2022 Guest Editorial Introduction to the Special Issue on Intelligent Transportation Systems in Epidemic Areas
abstract
The COVID-19 pandemic has posed significant challenges to transportation systems in various aspects, such as transferring patients and medical resources, enforcing physical distancing in public transportation, and controlling virus transmission through transportation networks. To address these challenges, a variety of artificial intelligence technologies, such as autonomous driving, big data analytics, intelligent vehicle routing and scheduling, and intelligent traffic control, have been employed in the design of intelligent transportation systems. This Special Issue provides a forum for researchers and practitioners to present the most recent advances in presenting and applying intelligent technologies to promote transportation systems in large-scale epidemics.
Yujun Zheng 0001, Honghai Liu 0001, Houxiang Zhang, Shengyong Chen
IEEE Trans. Intell. Transp. Syst.2
2022 Closed-Loop Construction and Analysis of Cortico-Muscular-Cortical Functional Network After Stroke
abstract
Brain networks allow a topological understanding into the pathophysiology of stroke-induced motor deficits, and have been an influential tool for investigating brain functions. Unfortunately, currently applied methods generally lack in the recognition of the dynamic changes in the cortical networks related to muscle activity, which is crucial to clarify the alterations of the cooperative working patterns in the motor control system after stroke. In this study, we integrate corticomuscular and intermuscular interactions to cortico-cortical network and propose a novel closed-loop construction of cortico-muscular-cortical functional network, named closed-loop network (CLN). Directional characteristic in terms of differentiating causal interactions is endowed on basis of the CLN framework, further expanding the definition of functional connectivity (FC) and effective connectivity (EC) dedicated to CLN. Next, CLN is applied to stroke patients to reveal the underlying after-effects mechanism of low frequency repetitive transcranial magnetic stimulation (rTMS) induced alterations of cortical physiologic functions during movement. Results show that the short-term modulation of rTMS is reflected in the enhancement of information interaction within the interhemispheric primary motor regions and inhibition of the coupling between motor cortex and effector muscles. CLN provides a new perspective for the study of motor-related cortical networks with muscle activities involvement instead of being restricted to brain network analysis of behaviors.
Jinbiao Liu, Gansheng Tan, Jixian Wang, Yina Wei, Yixuan Sheng, Honghai Liu 0001
IEEE Trans. Medical Imaging8
2021 Attribute-Driven Granular Model for EMG-Based Pinch and Fingertip Force Grand Recognition
abstract
Fine multifunctional prosthetic hand manipulation requires precise control on the pinch-type and the corresponding force, and it is a challenge to decode both aspects from myoelectric signals. This paper proposes an attribute-driven granular model (AGrM) under a machine-learning scheme to solve this problem. The model utilizes the additionally captured attribute as the latent variable for a supervised granulation procedure. It was fulfilled for EMG-based pinch-type classification and the fingertip force grand prediction. In the experiments, 16 channels of surface electromyographic signals (i.e., main attribute) and continuous fingertip force (i.e., subattribute) were simultaneously collected while subjects performing eight types of hand pinches. The use of AGrM improved the pinch-type recognition accuracy to around 97.2% by 1.8% when constructing eight granules for each grasping type and received more than 90% force grand prediction accuracy at any granular level greater than six. Further, sensitivity analysis verified its robustness with respect to different channel combination and interferences. In comparison with other clustering-based granulation methods, AGrM achieved comparable pinch recognition accuracy but was of lowest computational cost and highest force grand prediction accuracy.
Yinfeng Fang, Dalin Zhou, Kairu Li, Zhaojie Ju, Honghai Liu 0001
IEEE Trans. Cybern.5
2021 Physical Human-Robot Collaboration: Robotic Systems, Learning Methods, Collaborative Strategies, Sensors, and Actuators
abstract
This article presents a state-of-the-art survey on the robotic systems, sensors, actuators, and collaborative strategies for physical human-robot collaboration (pHRC). This article starts with an overview of some robotic systems with cutting-edge technologies (sensors and actuators) suitable for pHRC operations and the intelligent assist devices employed in pHRC. Sensors being among the essential components to establish communication between a human and a robotic system are surveyed. The sensor supplies the signal needed to drive the robotic actuators. The survey reveals that the design of new generation collaborative robots and other intelligent robotic systems has paved the way for sophisticated learning techniques and control algorithms to be deployed in pHRC. Furthermore, it revealed the relevant components needed to be considered for effective pHRC to be accomplished. Finally, a discussion of the major advances is made, some research directions, and future challenges are presented.
Uchenna Emeoha Ogenyi, Jinguo Liu, Chenguang Yang 0001, Zhaojie Ju, Honghai Liu 0001
IEEE Trans. Cybern.5
2021 Screening Early Children With Autism Spectrum Disorder via Response-to-Name Protocol
abstract
Incidence of children with autism spectrum disorder (ASD) has increased with an average rate of 1% worldwide. Clinical ASD screening, especially for children screening is a laborious and skilled task; however, there is no objective and effective method automating ASD children screening. Analyzing children ASD characteristics in predefined motion behavior protocols is attempted to provide automatic solutions to children ASD screening. A novel protocol, response to name (RTN), is proposed in this article for ASD clinical validation and diagnosis. The RTN method is jointly designed with clinical partners, and novel gaze estimation is developed for validating ASD characteristic behavior. Seventeen subjects including ten adults and seven children (five ASD subjects and two healthy subjects) have participated the experiment. The experiment results show that the proposed RTN system achieves an average classification score of 92.7% fully demonstrating that the principle of motion protocol based ASD screening has the potential to have early ASD screening automated.
Zhiyong Wang 0009, Keshi He, Xiu Xu, Honghai Liu 0001
IEEE Trans. Ind. Informatics6
2021 Multiscale Transfer Spectral Entropy for Quantifying Corticomuscular Interaction
abstract
Corticomuscular coupling reflects nonlinear interactions and multi-layer neural information transmission between the motor cortex and effector muscle in the sensorimotor system. Transfer spectral entropy (TSE) method has been used to describe corticomuscular coupling within single scale. As an extension of TSE, multiscale transfer spectral entropy (MSTSE) is proposed in this paper to depict multi-layer of neural information transfer between two coupling signals. The reliability and effectiveness of MSTSE were verified on data generated by nonlinear numerical models and those of a force tracking task. Compared with TSE, MSTSE is more robust to the embedding dimension and performs optimally in the detection of the coupling properties. Further analysis of the physiological signals showed that the MSTSE provided more detailed band characteristics than the single scale TSE measurement. MSTSE indicates significant coupling scattered in alpha, beta and low gamma bands during the force tracking task. Besides, the coupling strength in the descending direction of the beta band was significantly higher than that in the ascending direction. This study constructs multi-scale coupling information to provide a new perspective for exploring corticomuscular interaction.
Jinbiao Liu, Gansheng Tan, Yixuan Sheng, Honghai Liu 0001
IEEE J. Biomed. Health Informatics4
2021 A Wearable Ultrasound System for Sensing Muscular Morphological Deformations
abstract
Noninvasive monitoring of muscle contraction, which provides information about muscle morphological deformations, has a great potential in medical applications, such as stroke rehabilitation and prosthesis control. This paper presents a wearable multichannel A-mode ultrasound system for the multiperspective muscle contraction detection against its existing bulky ultrasound sensing counterpart. The system consists of a waveform generator and a waveform amplifier for ultrasound excitation, as well as a signal processing module for echo receiving, amplifying, and transmitting. In addition, a miniaturized transducer was optimized for muscle contraction monitoring using 1-3 piezoelectric composite. The system's superiorities on excitation pulse, detection depth, and axial resolution were validated by the evaluation experiments. And in vivo muscle deformation detection and virtual prosthesis control experiments (target achievement control test) demonstrated its ability in rehabilitation applications, with a task completion rate of 100% and path efficiency of 93.30%. These results confirmed the reliability of the proposed ultrasound system and paved the way for its applications in rehabilitation treatment and prosthesis research.
Xingchen Yang, Zhenfeng Chen, Nalinda Hettiarachchi, Jipeng Yan 0001, Honghai Liu 0001
IEEE Trans. Syst. Man Cybern. Syst.5
2020 Object-oriented Map Exploration and Construction Based on Auxiliary Task Aided DRL
abstract
Environment exploration by autonomous robots through deep reinforcement learning (DRL) based methods has attracted more and more attention. However, existing methods usually focus on robot navigation to single or multiple fixed goals, while ignoring the perception and construction of external environments. In this paper, we propose a novel environment exploration task based on DRL, which requires a robot fast and completely perceives all objects of interest, and reconstructs their poses in a global environment map, as much as the robot can do. To this end, we design an auxiliary task aided DRL model, which is integrated with the auxiliary object detection and 6-DoF pose estimation components. The outcome of auxiliary tasks can improve the learning speed and robustness of DRL, as well as the accuracy of object pose estimation. Comprehensive experimental results on the indoor simulation platform AI2-THOR have shown the effectiveness and robustness of our method.
Junzhe Xu 0001, Jianhua Zhang 0002, Shengyong Chen, Honghai Liu 0001
ICPR4
2020 Sampling-Based Path Planning in Heterogeneous Dimensionality-Reduced Spaces
abstract
Many sampling strategies often consider the goal and obstacle population to bias/restrict the search area, and they however become less effective when the robot has many degrees of freedom. This paper explores the nonhomogeneous restriction imposed by the obstacles and presents an improved SBP approach enhanced by heterogeneous dimensionality reduction of the full configuration space. Based on the projection residual, a new Dirichlet process (DP) mixture model is proposed to capture a number of Dimensionality-Reduced Spaces (DRSs), which offer the planning spaces with fewer dimensions than its single-DRS counterpart. Then, the sampling and planning procedures are unified with a proposed transversality condition, connecting sampled nodes across DRSs. At last, a quadratic programming is formulated and quickly solved to map the found path in DRSs to an output path in the full configuration space. Numerical simulations on path planning problems of a high-dimensional Intervention Autonomous Underwater Vehicle (I-AUV) have been conducted, showing the feasibility and efficiency of the proposed method.
Wenjie Lu 0004, Huan Yu 0004, Hao Xiong 0004, Honghai Liu 0001
IECON4
2020 A-mode Ultrasound Driven Sensor Fusion for Hand Gesture Recognition
abstract
Traditionally, Surface electromyography (sEMG) has been the predominant method of sensing muscle activity in order to control myoelectric prosthesis. While many prosthesis control schemes used simple direct control, an ever increasing focus has moved to pattern recognition based approaches which promise a greater degree of natural control such as to further improve an amputees quality of life. Although pattern recognition based approaches have shown great promise, they have innate limitations due to changes that may occur during long term use which prevent clinical acceptance. Due to these limitations, researchers have increasingly investigated alternative modalities to provide more robust control schemes. A particular modality that has seen increasing interest is ultrasound based sensing due to its capability to better understand deep tissue activity.Within this research, A-mode ultrasound based sensing is proposed not as a replacement for sEMG based sensing but instead to augment and drive sEMG based sensing during activities that may otherwise prove challenging to traditional sEMg based control schemes.
Peter Boyd, Honghai Liu 0001
IJCNN2
2020 CalibRCNN: Calibrating Camera and LiDAR by Recurrent Convolutional Neural Network and Geometric Constraints
abstract
In this paper, we present Calibration Recurrent Convolutional Neural Network (CalibRCNN) to infer a 6 degrees of freedom (DOF) rigid body transformation between 3D LiDAR and 2D camera. Different from the existing methods, our 3D-2D CalibRCNN not only uses the LSTM network to extract the temporal features between 3D point clouds and RGB images of consecutive frames, but also uses the geometric loss and photometric loss obtained by the interframe constraint to refine the calibration accuracy of the predicted transformation parameters. The CalibRCNN aims at inferring the correspondence between projected depth image and RGB image to learn the underlying geometry of 2D-3D calibration. Thus, the proposed calibration model achieves a good generalization ability to adapt to unknown initial calibration error ranges, and other 3D LiDAR and 2D camera pairs with different intrinsic parameters from the training dataset. Extensive experiments have demonstrated that our CalibRCNN can achieve state-of-the-art accuracy by comparison with other CNN based methods.
Jieying Shi, Ziheng Zhu, Jianhua Zhang 0002, Ruyu Liu, Zhenhua Wang 0003, Shengyong Chen, Honghai Liu 0001
IROS7
2020 Delay estimation for cortical-muscular interaction via the rate of voxels change
abstract
It is evident that corticomuscular coherence (CMC), representing the functional coupling between motor cortex and muscle tissues, plays a crucial role in neurophysiologic studies and applications. It is hypothesized that there is an unknown time delay comprising at least neural conduction time in the process of corticomuscular interaction. In this study, we developed a novel delay estimation method, defined as the rate of voxels change (RVC) for the estimation of time delay in two coupled physiological signals. The RVC is the dynamic variation of the local CMCs observed in different time offsets. Both simulation and physiological data confirm the capability of RVC in estimating cortical-muscular delay. The underlying mechanisms of individual discrepancy of the latency is also investigated via exploring the correlation between delays, brain activity and motor performance. Correlation analyses indicate an intrinsic link between the connectivity strength of the brain network and the length of time delay in cortical-muscular interactions.
Jinbiao Liu, Gansheng Tan, Yixuan Sheng, Jiaole Wang, Wenjie Lu 0004, Honghai Liu 0001
SMC6
2020 Free-Head Pose Estimation under Low-Resolution Scenarios
abstract
Head pose offers vital cues to infer one's social attention in wide applications. Most existing head pose estimation algorithms have demonstrated competitive results taking high resolution images of frontal view as input. However, these approaches still work poorly if they are fed with low-resolution images. In a more common realistic scene such as computer vision assisted autism screening, images of unconstrained patients that are taken from distant cameras often have low-resolution and non-frontal faces. To this end, we present a multi-view scheme based CNN for free-head pose classification. Residual Networks are taken as the backbone model to generate effective feature representations of the low-resolution images. A novel multi-view feature fusion layer is proposed to address facial appearance variation over multi perspectives owing to free movement. Also, a multi loss function by combining binned classification and regression losses of different pose angles is employed to obtain a more precise pose estimation. The proposed method is evaluated on two scenarios: (1) a publicly available dataset and (2) a practical application to explore the social attention of a group with social deficits: autistic children. Experimental results suggest that our method significantly outperforms state-of-the-art multi-view pose classification methods and achieves comparable pose estimation results. Moreover, the proposed method can be extended to applications of quantitative analysis of social deficits.
Zhiyong Wang 0009, Haibo Qin, Bin Ji 0004, Honghai Liu 0001
SMC6
2020 An Auxiliary Screening System for Autism Spectrum Disorder Based on Emotion and Attention Analysis
abstract
The screening and diagnosis of Autism Spectrum Disorder(ASD) suffer from great challenges due to insufficient professional clinicians and complex procedures. It is urgent to introduce an effective auxiliary system in the diagnosis and treatment process to assist in the completion of pathological information collection tasks, consequently simplifying the screening method and improving the accuracy of screening. We propose a computer vision-based early screening system for ASD to characterize the facial expressions and eye gaze attention which are considered to remarkable indicators for early screening of autism. The system provides the subjects with three different virtual interaction modes: video, picture, and virtual interactive game. During the interaction between the subject and the computer, the system extracts and analyzes the quantitative information of the subject's performance. Then, through computer vision-based emotion analysis and attention analysis methods, the subject's emotions and attention features in the three interaction modes are automatically calculated to assist in the early screening of autism. Finally, the accuracy and feasibility of the system are verified through experiments on both the publicly available dataset and the data collected from 10 ASD children.
Bin Ji 0004, Zhiyong Wang 0009, Honghai Liu 0001
SMC5
2020 Feature Fusion of sEMG and Ultrasound Signals in Hand Gesture Recognition
abstract
Multi-modal sensory fusion is believed to obtain higher accuracy in gesture recognition. Its difficulty lies in mining discriminative features and fusing features from different modalities. Surface electromyography(sEMG) and ultrasound signals are typical signal modalities in gesture recognition. It is expected that the fusion of them can take advantage of the complementarity of electrophysiological information and muscle morphology information. This paper proposed two kinds of feature fusion method. The one is concatenating the manual designed sEMG and ultrasound features, and the other is a convolutional neural network (CNN) based feature exaction and fusion method for sEMG and ultrasound signals. Eight able-bodied subjects were involved to participate in the experiments. In the experiments, four channels of sEMG and A-mode ultrasound signals corresponding to 20 gestures were collected synchronously to evaluate the proposed method. The experimental results demonstrated that the fusion sEMG-ultrasound feature always outperformed the separate sEMG or ultrasound feature regardless of the feature extraction method, and as for fusion sEMG-ultrasound feature, the CNN based method achieve a high accuracy (97.38±1.49%) in 20 gestures, which surpassed the method of concatenating the manual designed features and applying machine learning algorithm (LDA, KNN, SVM).
Yu Zhou 0013, Yicheng Yang, Jiaole Wang, Honghai Liu 0001
SMC5
2020 Two-dimensional discrete feature based spatial attention CapsNet For sEMG signal recognition
Guoqi Chen, Wanliang Wang, Zheng Wang 0048, Honghai Liu 0001, Zelin Zang, Weikun Li
Appl. Intell.4
2020 Upper-limb functional assessment after stroke using mirror contraction: A pilot study
Yu Zhou 0013, Hongze Jiang, Honghai Liu 0001
Artif. Intell. Medicine6
2020 Classifying ASD children with LSTM based on raw videos
Jing Li 0027, Yihao Zhong, Junxia Han, Gaoxiang Ouyang, Xiaoli Li 0002, Honghai Liu 0001
Neurocomputing6
2020 HDS-SP: A novel descriptor for skeleton-based human action recognition
Zhiyong Wang 0009, Honghai Liu 0001
Neurocomputing3
2020 Improved itracker combined with bidirectional long short-term memory for 3D gaze estimation using appearance cues
Xiaolong Zhou 0001, Jianing Lin, Zhuo Zhang 0012, Zhanpeng Shao, Shenyong Chen, Honghai Liu 0001
Neurocomputing6
2020 Surface EMG data aggregation processing for intelligent prosthetic action recognition
Gongfa Li, Guozhang Jiang, Disi Chen, Honghai Liu 0001
Neural Comput. Appl.5
2020 Research on gesture recognition of smart data fusion features in the IoT
Ying Sun 0004, Gongfa Li, Guozhang Jiang, Disi Chen, Honghai Liu 0001
Neural Comput. Appl.6
2020 Multi-stage adaptive regression for online activity recognition
Bangli Liu, Haibin Cai, Zhaojie Ju, Honghai Liu 0001
Pattern Recognit.4
2020 Guest Editorial: Special Section on Latest Advances on Industrial Intelligent Video Systems and Analytics
abstract
This Special Section collects the latest developments in video system design, data compression, target detection, object localization, behavior analysis, motion detection, and real-time implementation of industrial video systems and intelligent analytics to bring the latest ideas and solutions of the research community on practical video systems to our audience. The Special Section focuses on several topics that are recently concerned in the community, including, multicamera network, real-time hardware implementation, networked data analytics, bandwidth limited compression, motion pattern analysis, action understanding, three-dimensional (3-D) reconstruction, contextual recognition, object detection and tracking, intelligent robot vision, security surveillance, intelligent transportation, and other industrial applications. The Special Section presents 14 articles on intelligent video systems and analytics. These articles are briefly summarized here.
Shengyong Chen, Honghai Liu 0001, Naoyuki Kubota
IEEE Trans. Ind. Informatics2
2019 Spatial Map Learning with Self-Organizing Adaptive Recurrent Incremental Network
abstract
Biological information inspires the advancement of a navigational mechanism for autonomous robots to help people explore and map real-world environments. However, the robot's ability to constantly acquire environmental information in real-world, dynamic environments has remained a challenge for many years. In this paper, we propose a self-organizing adaptive recurrent incremental network that models human episodic memory to learn spatiotemporal representations from novel sensory data. The proposed method termed as SOARIN consists of two main learning process that is active learning and episodic memory playback. For active learning (robot exploration), SOARIN quickly learns and adapts incoming novel sensory data as episodic neurons via competitive Hebbian Learning. Episodic neurons are connecting with each other and gradually forms a spatial map that can be used for robot localization. Episodic memory playback is triggered whenever the robot is in an inactive mode (charging or hibernating). During playback, SOARIN gradually integrates knowledge and experience into more consolidate spatial map structures that can overcome the catastrophic forgetting. The proposed method is analyzed and evaluated in term of map learning and localization through a series of real robot experiments in real-world indoor environments.
Wei Hong Chin, Naoyuki Kubota, Chu Kiong Loo, Zhaojie Ju, Honghai Liu 0001
IJCNN5
2019 Dictionary Learning and Confidence Map Estimation-Based Tracker for Robot-Assisted Therapy System
Xiaolong Zhou 0001, Sixian Chan 0001, Shengyong Chen, Honghai Liu 0001
PRCV (1)5
2019 Electrotactile Stimulation Waveform Modulation Based on A Customized Portable Stimulator: A Pilot Study
abstract
Artificial tactile sensation (ATS) is of great importance in diverse fields, especially for amputees to explore and interact with the world. Although prosthesis are dexterous nowadays, lack of tactile sensation is expected to result in a high rejection rate. Electrical stimulation offers the potential of creating ATS to restore tactile sensation. This paper proposed a portable wireless electrotactile stimulator, which can output common square wave (CSW), sine wave (SW) and time-varying pulse width square wave (TPSW), with a superior resolution. A preliminary experiment showed that the amplitude was easy to be distinguished compared to the frequency for all waveforms. The pulse width changing frequency of TPSW was more discernible than that of CSW and SW. Besides, the TPSW felt the most comfortable while the CSW felt the worst. Furthermore, the SW felt better compared with CSW, especially under the low-frequency and high-amplitude condition. Therefore, the method have a potential to alleviate the discomfort of electrotactile stimulation and achieve a smooth sensation.
Yicheng Yang, Yu Zhou 0013, Keshi He, Honghai Liu 0001
SMC5
2019 Jointly network: a network based on CNN and RBM for gesture recognition
Ying Sun 0004, Gongfa Li, Guozhang Jiang, Honghai Liu 0001
Neural Comput. Appl.5
2019 RGB-D sensing based human action and interaction analysis: A survey
Bangli Liu, Haibin Cai, Zhaojie Ju, Honghai Liu 0001
Pattern Recognit.4
2019 Electrotactile Feedback in a Virtual Hand Rehabilitation Platform: Evaluation and Implementation
abstract
Tactile feedback plays an important role in hand manipulation, especially in the grasping process which is one of the major functions of the hand. However, few commercially available prosthetic hands or hand motor function rehabilitation systems are equipped with tactile feedback. The absence of suitable tactile feedback modules leads to an inferior rehabilitation performance with a large burden on user training and compromised usability. Thus, it is challenging and essential to integrate a proper tactile feedback module with the existing hand rehabilitation systems to achieve a better control performance and accelerate the rehabilitation process. This paper focuses on the implementation and evaluation of the electrotactile feedback (EF) enhanced rehabilitation system. A virtual hand rehabilitation platform is proposed comprising an surface electromyography (sEMG) acquisition module, an electrotactile stimulation module, a virtual environment with sEMG-driven humanlike hand and numerical feedbacks of grasping force and fingertip deformation, where a closed-loop control is formed. Three different feedback conditions including visual feedback (VF), EF, and no feedback (NF) are compared based on the proposed platform. Experiments were conducted on 10 able-bodied subjects, and multiple quantitative metrics for the rehabilitation performance evaluation including training burden estimation and success rate (SR) of tasks were adopted. Results indicate that the integration of EF is helpful to both reduce the rehabilitation duration and improve the virtual grasping SR in comparison with the NF condition while possessing a better practicality over VF. Note to Practitioners-This paper is motivated by the problem of hand grasp control for rehabilitation purposes, but it also applies to other hand motor function rehabilitation process. Existing hand motor function rehabilitation approaches generally lack a proper feedback and rely on the heavy burden during user training and the users experience. This paper suggests incorporating EF to improve the efficiency and efficacy of the rehabilitation process. The electrical stimulation is driven by the myoelectric-sensing-based force estimation and encoded in a manual scheme to fit each individual involved. In our work, a virtual hand rehabilitation platform is implemented to verify the feasibility of EF in reducing the burden of user training and improving the rehabilitation performance, which allows the expanding of the current system into a broader spectrum of motor function rehabilitation applications. Experiments on able-bodied subjects suggest that the EF in the proposed virtual hand rehabilitation platform is feasible, but it has not been tested on the limb-impaired subjects and confined to a manual encoding of electrical stimulation. In the future research, the design of a general EF enhanced hand rehabilitation platform with a standardized stimulation parameter optimization will be addressed and further validated on the subjects with limb impairments and amputation.
Kairu Li, Peter Boyd, Yu Zhou 0013, Zhaojie Ju, Honghai Liu 0001
IEEE Trans Autom. Sci. Eng.5
2019 Towards Zero Re-Training for Long-Term Hand Gesture Recognition via Ultrasound Sensing
abstract
While myoelectric pattern recognition is a prevailing way for gesture recognition, the inherent nonstationarity of electromyography signals hinders its long-term application. This study aims to prove a hypothesis that morphological information of muscle contraction detected by ultrasound image is potentially suitable for long-term use. A set of ultrasound-based algorithms are proposed to realize robust hand gesture recognition over multiple days, with user training only at the first day. A markerless calibration algorithm is first presented to position the ultrasound probe during donning and doffing; an algorithm combining speeded-up robust features and bag-of-features model being immune to ultrasound probe shift and rotation is then introduced; a self-enhancing classification method is next adopted to update classification model automatically by incorporating useful knowledge from testing data; finally the performance of long-term hand gesture recognition with zero re-training is validated by a six-day experiment of six healthy subjects, whose outcomes strongly support the hypothesis with about 94% of gesture recognition accuracy for each testing day. This study confirms the feasibility of adoption of ultrasound sensing for long-term musculature related applications.
Xingchen Yang, Dalin Zhou, Yu Zhou 0013, Youjia Huang, Honghai Liu 0001
IEEE J. Biomed. Health Informatics5
2018 Accurate Eye Center Localization via Hierarchical Adaptive Convolution
Haibin Cai, Bangli Liu, Zhaojie Ju, Serge Thill, Tony Belpaeme, Bram Vanderborght, Honghai Liu 0001
BMVC7
2018 Online Action Recognition based on Skeleton Motion Distribution
Bangli Liu, Zhaojie Ju, Naoyuki Kubota, Honghai Liu 0001
BMVC4
2018 Feature Selection Mechanism in CNNs for Facial Expression Recognition
Shuwen Zhao, Haibin Cai, Honghai Liu 0001, Jianhua Zhang 0002, Shengyong Chen
BMVC3
2018 Robust Eye Center Localization Based on an Improved SVR Method
Zhiyong Wang 0009, Haibin Cai, Honghai Liu 0001
ICONIP (7)3
2018 Automatic Estimation of Biceps Brachi Muscle Thickness in B-Mode Ultrasound Images
abstract
Muscle thickness is an important parameter used for quantifying musculoskeletal function. At present, the muscle thickness measurement relies on the manual method, which is subjective, time-consuming, and prone to error. In this paper, a novel automatic calculation method was proposed to achieve the continuous and quantitative measurement for the muscle thickness of biceps brachi in ultrasound images. The proposed method includes three steps: the detection and tracking of the seed point in the superficial aponeurosis, the detection and tracking of the feature line in the deep aponeurosis, and the muscle thickness estimation. The performance of the proposed method was firstly compared to the manual measured results, which demonstrates that they have a good agreement during the experiment and can be used for efficiently estimating the muscle thickness of biceps brachi in the musculoskeletal ultrasound images. Then, the validated method was employed to investigate the relationship between the muscle thickness and the joint angle. The results show that, the elbow flexion and extension angle has an approximately linear relationship with muscle thickness change.
Honghai Liu 0001, Xingchen Yang, Xueli Sun, Linwei Ye, Keshi He
SMC1
2018 ECG-Enhanced Multi-sensor Solution for Wearable Sports Devices
abstract
Electrocardiogram (ECG) plays a crucial role in the prevention of cardiovascular diseases in humans and various types of ECG monitoring equipment are continuously being developed. Most monitoring methods place the electrodes near the heart or on both arms, based on standard ECG leads. Although accurate monitoring effect has been achieved by these conventional approaches, the wearable performance still need to be improved. This paper proposed a novel ECG-enhanced multi-sensor solution for wearable sports devices. A wireless and wearable ECG detection system based on signal acquisition from left upper-arm was designed to verify this solution method. The system has been evaluated with solid experiments proving that the system has outperformed existing similar system. Moreover, the inertial measurement unit (IMU) and electromyography (EMG) data were detected and fused by the system to determine the validity of the ECG signal. It indicated that the system can achieve a good performance of heart rate detection under different body states.
Yu Zhou 0013, Yinfeng Fang, Honghai Liu 0001
SMC4
2018 Analysis of Stable sEMG Features for Bilateral Upper Limb Motion
abstract
It is evident that the objective physiology-based diagnostic methods are desired to improve the evaluation efficiency for stroke assessment, however the assessment relies mainly on the diagnosis scale and the experience of the therapists in clinic. This paper aims to evaluate the features based on sEMG from bilateral upper limb, which can be used to show the bilateral upper limb differences and to map the stroke scale. Preliminary experiments were conducted on eleven healthy subjects and some classical and innovative features were evaluated. With all the work of this paper, we anticipate that Root Mean Square (RMS), Crest Factor (CF) and Impulse Factor (I), are able to be used to evaluate stroke scale.
Yangbo Yu, Hongze Jiang, Yu Zhou 0013, Honghai Liu 0001
SMC4
2018 A structured multi-feature representation for recognizing human action and interaction
Bangli Liu, Zhaojie Ju, Honghai Liu 0001
Neurocomputing3
2018 Bayesian tensor factorization for multi-way analysis of multi-dimensional EEG
Yunbo Tang, Dan Chen 0001, Lizhe Wang 0001, Albert Y. Zomaya, Jingying Chen 0001, Honghai Liu 0001
Neurocomputing6
2018 Gesture Recognition Based on Kinect and sEMG Signal Fusion
Ying Sun 0004, Cuiqiao Li, Gongfa Li, Guozhang Jiang, Du Jiang, Honghai Liu 0001, Zhigao Zheng 0001, Wanneng Shu
Mob. Networks Appl.6
2018 Ultrasound-Based Sensing Models for Finger Motion Classification
abstract
Motions of the fingers are complex since hand grasping and manipulation are conducted by spatial and temporal coordination of forearm muscles and tendons. The dominant methods based on surface electromyography (sEMG) could not offer satisfactory solutions for finger motion classification due to its inherent nature of measuring the electrical activity of motor units at the skin's surface. In order to recognize morphological changes of forearm muscles for accurate hand motion prediction, ultrasound imaging is employed to investigate the feasibility of detecting mechanical deformation of deep muscle compartments in potential clinical applications. In this study, finger motion classification has been represented as subproblems: recognizing the discrete finger motions and predicting the continuous finger angles. Predefined 14 finger motions are presented in both sEMG signals and ultrasound images and captured simultaneously. Linear discriminant analysis classifier shows the ultrasound has better average accuracy (95.88%) than the sEMG (90.14%). On the other hand, the study of predicting the metacarpophalangeal (MCP) joint angle of each finger in nonperiod movements also confirms that classification method based on ultrasound achieves better results (average correlation 0.89 $\pm$ 0.07 and NRMSE 0.15 $\pm$ 0.05) than sEMG (0.81 $\pm$ 0.09 and 0.19 $\pm$ 0.05). The research outcomes evidently demonstrate that the ultrasound can be a feasible solution for muscle-driven machine interface, such as accurate finger motion control of prostheses and wearable robotic devices.
Youjia Huang, Xingchen Yang, Dalin Zhou, Keshi He, Honghai Liu 0001
IEEE J. Biomed. Health Informatics6
2017 Human-human interaction recognition based on spatial and motion trend feature
abstract
Human-human interaction recognition has attracted increasing attention in recent years due to its wide applications in computer vision fields. Currently there are few publicly available RGBD-based human-human interaction datasets collected. This paper introduces a new dataset for human-human interaction recognition. Furthermore, a novel feature descriptor based on spatial relationship and semantic motion trend similarity between body parts is proposed for human-human interaction recognition. The motion trend of each skeleton joint is firstly quantified into the specific semantic word and then a Kernel is built for measuring the similarity of either intra or inter body parts by histogram interaction. Finally, the proposed feature descriptor is evaluated on the SBU interaction dataset and the collected dataset. Experimental results demonstrate the outperformance of our method over the state-of-the-art methods.
Bangli Liu, Haibin Cai, Xiaofei Ji, Honghai Liu 0001
ICIP4
2017 Cascade support vector regression-based facial expression-aware face frontalization
abstract
The main aim of face frontalization is to synthesize the frontal facial appearances from non-frontal facial images. How to estimate the frontal face-shape is a crucial but very challenging problem in the frontalization task. Most existing methods use a single shape template to fit in with frontal facial appearances, which will result in a loss of expression-related information. In this work, we present a novel facial expression-aware face frontalization method which directly learns the pair-wise relations between non-frontal face-shape and its frontal counterpart. The support vector regression is explored to train the pair-wise regression model. Considered the pair-wise relationship is non-linear, an appropriate cascade manner is applied to iteratively adjust and optimize the model. With the estimated frontal shape, facial appearances are synthesized through a texture-fitting process formulated by solving a simple optimization problem. The proposed method has been evaluated on a in-the-wild facial expression database. The experimental results shows an outstanding performance of both visual effects of expression recovery and facial expression recognition.
Yiming Wang 0001, Hui Yu 0001, Junyu Dong, Muwei Jian, Honghai Liu 0001
ICIP5
2017 Two-eye model-based gaze estimation from a Kinect sensor
abstract
In this paper, we present an effective and accurate gaze estimation method based on two-eye model of a subject with the tolerance of free head movement from a Kinect sensor. To accurately and efficiently determine the point of gaze, i) we employ two-eye model to improve the estimation accuracy; ii) we propose an improved convolution-based means of gradients method to localize the iris center in 3D space; iii) we present a new personal calibration method that only needs one calibration point. The method approximates the visual axis as a line from the iris center to the gaze point to determine the eyeball centers and the Kappa angles. The final point of gaze can be calculated by using the calibrated personal eye parameters. We experimentally evaluate the proposed gaze estimation method on eleven subjects. Experimental results demonstrate that our gaze estimation method has an average estimation accuracy around 1.99°, which outperforms many leading methods in the state-of-the-art.
Xiaolong Zhou 0001, Haibin Cai, Youfu Li 0001, Honghai Liu 0001
ICRA4
2017 Embedded vision based automotive interior intrusion detection system
abstract
Motor vehicle theft has caused massive economic loss over the world. This paper proposes an embedded vision system to detect automotive interior intrusion. The system uses a fusion of an acceleration module and a vision module to meet the requirement of low power consumption for most motor vehicles. Furthermore, an effective intrusion detection algorithm is developed for the on-board vision module. The vision system is able to detect the intrusion even in the dark night due to the employment of infrared lights. Experimental evaluation is conducted under a variety of illumination conditions, such as day time, night time and even shining light. An intrusion detection accuracy of 91.7% is achieved, which shows that the developed embedded vision system is reliable for motor vehicle intrusion detection.
Haibin Cai, Hwang Joonkoo, Yinfeng Fang, Honghai Liu 0001
SMC6
2017 Exploring the relation between EMG sampling frequency and hand motion recognition accuracy
abstract
Myoelectric control with surface EMG signal has achieved great success in clinics, but only limited to the control of 2-Degrees-of-freedom prosthesis. With the appearance of multiple-channel and high-density EMG system and the advances of pattern recognition technology, it becomes possible to control a multi-degree smart prosthesis using EMG signals. However, it requires high performance EMG systems with high sampling frequency, which impedes the popularity of EMG-based applications. This study aims to explore a way to reduce the cost of EMG system by investigating the effect of sampling rate on gesture recognition accuracy. Two groups of experiments on inner-group and cross-group were designed to evaluate the classification accuracy at different EMG sampling frequency. In comparison with the sampling frequency at 1kHz, a lower sampling frequency at 400 Hz could achieve comparable accuracy, reduced by only 0.43% (KNN) and 0.83% (SVM) with the overall accuracy at 99.40% and 98.67%, respectively. It implies that appropriate reduction of the sampling frequency can be a good choice to balance the cost and performance of a multiple channel EMG system for feature-based hand gesture classification.
HongFeng Chen, Yue Zhang 0072, Zhuo Zhang 0012, Yinfeng Fang, Honghai Liu 0001, Chun-Yan Yao
SMC5
2017 A force-driven granular model for EMG based grasp recognition
abstract
It is a challenge to precisely predict hand grasps based on EMG signals given practical scenarios, due to its inherent nature. This paper proposes a solution to tackle the challenge with a force-driven granular model (FDGM). The problem of n-class hand grasp classification has been represented as force-based granular modelling, in which a number of granules are constructed for each class relying on the synchronically captured grasping force. A rule based mechanism is formed for granule generation of each class, and a cross-testing algorithm is proposed to optimise the number of granules. The experiment based on 8-case grasp recognition reveals that the proposed method performs better in terms of motion recognition accuracy of multiple EMG channel combination, and is more insensitive to signal interferences. In comparison with other rules of information granulation, it is confirmed that the force-driven rule is of the most efficiency with comparable classification accuracy. The research outcomes pave the way for real-time prediction of grasps and corresponding force in human-centred environments.
Yinfeng Fang, Dalin Zhou, Kairu Li, Zhaojie Ju, Honghai Liu 0001
SMC5
2017 Activity recognition for asd children based on joints estimation
abstract
Human motion recognition is a trending topic and could be applied in many areas, the motion estimation of ASD children is more challenging because of the high uncertainty of their activities, we thus introduced a novel method which is designed for estimating the upper joints and recognising their special motions, we verified the proposed method on our recorded ASD children dataset and adult dataset, the experimental results show the proposed method is effective on the dataset.
Dongxu Gao, Zhaojie Ju, Yingfeng Fang, Jiangtao Cao, Chenguang Yang 0001, Honghai Liu 0001
SMC6
2017 Toward an Enhanced Human-Machine Interface for Upper-Limb Prosthesis Control With Combined EMG and NIRS Signals
abstract
Advanced myoelectric prosthetic hands are currently limited due to the lack of sufficient signal sources on amputation residual muscles and inadequate real-time control performance. This paper presents a novel human-machine interface for prosthetic manipulation that combines the advantages of surface electromyography (EMG) and near-infrared spectroscopy (NIRS) to overcome the limitations of myoelectric control. Experiments including 13 able-bodied and three amputee subjects were carried out to evaluate both offline classification accuracy (CA) and online performance of the forearm motion recognition system based on three types of sensors (EMG-only, NIRS-only, and hybrid EMG-NIRS). The experimental results showed that both the offline CA and realtime performance for controlling a virtual prosthetic hand were significantly (p <; 0.05) improved by combining EMG and NIRS. These findings suggest that fusion of EMG and NIRS is feasible to improve the control of upper-limb prostheses, without increasing the number of sensor nodes or complexity of signal processing. The outcomes of this study have great potential to promote the development of dexterous prosthetic hands for transradial amputees.
Weichao Guo, Xinjun Sheng, Honghai Liu 0001
IEEE Trans. Hum. Mach. Syst.3
2016 Facial Expression-Aware Face Frontalization
Yiming Wang 0001, Hui Yu 0001, Junyu Dong, Brett Stevens, Honghai Liu 0001
ACCV (3)5
2016 Human-machine interface based on multi-channel single-element ultrasound transducers: A preliminary study
abstract
Ultrasound (US) imaging is a promising sensing technique in the field of human-machine interface, and many positive results have been reported in literature on hand gesture recognition or finger angle prediction based on US imaging. However, in most of these studies, linear array ultrasound probes were used to generate US images, which made the US device expensive and bulky. In this paper, a method of extracting forearm muscle information via multiple single-element US transducers is proposed. By using this kind of transducers, a low-cost and small-size human-machine interface can be expected. Preliminary results show that an average recognition accuracy of 96% can be achieved for six motions, including five finger flexions and rest state.
Keshi He, Xueli Sun, Honghai Liu 0001
HealthCom4
2016 Performances of surface EMG and Ultrasound signals in recognizing finger motion
abstract
This paper compared the performances of the surface electromyography (sEMG) with Ultrasound(US) signals in finger motion recognition. Compared to traditional sEMG-based human machine interface (HMI), the ultrasound can provide some information about the morphological changes of muscles and has higher resolution. It is possible for the US-based HMI to recognize some more dexterous hand gesture especially for some finger motions. In the experiment, the subjects were instructed to perform 14 different finger motions. The sEMG signals and the ultrasound images were collected simultaneously. The 2-fold cross-validation with LDA classifier was used to analyze the accuracy. The mean accuracy of US-based HMI was nearly 96.37% while the sEMG was 92.41%. One-way ANOVA analysis showed that the US-based HMI had a significantly higher classification accuracy than sEMG-based HMI. What's more, the difference between individuals and finger motions of ultrasound was significantly smaller than that of sEMG. The US-based HMI performed better and more stable than sEMG, which may imply that ultrasound has a great potential to be an alternative method to sEMG regarding precise control, or it can corporate with sEMG to achieve a better HMI.
Youjia Huang, Honghai Liu 0001
HSI2
2016 Comparison of online adaptive learning algorithms for myoelectric hand control
abstract
Pattern recognition (PR) based myoelectric hand control has become a research focus in the field of rehabilitative engineer and intelligent control. However, the state of the art method is hardly adopted for clinical use because of signal interfered by shift, fatigue and user-unfriendly of retraining. The aim of this study is to evaluate the performance of different kinds of online algorithms in classifying the myoelectric hand motions, and reveal the key factors to classification accuracy of online learning algorithms. Two groups of experiments on intra-session and inter-session were designed to evaluate the classification and recognition performance of overall methods. The comparison results show that the second-order online learning algorithms outperformed the first-order algorithms in classification and recognition. Soft confidence-weighted learning performs best with 99% classification rate in same session and over 85% recognition rate in different session. This paper uncovers the online learning with large margin and confidence weight can always acquire a good property. In addition, online learning algorithms retrain the classification model by incorporating the testing data to the previous model by measuring the changes between the predicted label and true label which can improve the performance in long-term use.
Yue Zhang 0072, Zheng Wang 0048, Zhuo Zhang 0012, Yinfeng Fang, Honghai Liu 0001
HSI5
2016 Combining 3D joints Moving Trend and Geometry property for human action recognition
abstract
Depth image based human action recognition has attracted many attentions due to the popularity of the depth sensors. However, accurate recognition still remains a challenge because of various object appearances, poses and video sequences. In this paper, a novel skeleton joints descriptor based on 3D Moving Trend and Geometry (3DMTG) property is proposed for human action recognition. Specifically, a histogram of 3D moving directions between consecutive frames for each joint is constructed to represent the 3D moving trend feature in spatial domain. The geometry information of joints in each frame is modelled by the relative motion with the initial status. The proposed feature descriptor is evaluated on two popular datasets. The experimental results demonstrate the superior performance of our method over the state-of-the-art methods, especially the higher recognition rates for complex actions.
Bangli Liu, Hui Yu 0001, Xiaolong Zhou 0001, Honghai Liu 0001
SMC5
2016 A wave detection method for air-coupled ultrasound system on human abdominal region
abstract
This paper describes an air-coupled ultrasound system by using DIO-2000. The system is aimed to use for inner muscle evaluation in rehabilitation process. The system evaluates inner muscle with low constrain than conventional method which measured by using MR image, X-ray CT and contacted ultrasound system. Our air-coupled ultrasound system measures transmitted ultrasound wave with very low power through human abdominal region by employing a pulsar-receiver with high sensitive preamplifier, and wave detection method based on fuzzy inference finds transmitted wave from noisy wave. The fuzzy inference is derived from characteristics of transmitted wave. In the experiment, we evaluate the accuracies of wave detection method for human body.
Takahiro Takeda, Takuaya Mabuchi, Naoyuki Takesue, Naoyuki Kubota, Honghai Liu 0001
SMC5
2016 Efficient vanishing point detection method in unstructured road environments based on dark channel prior
abstract
Vanishing point detection is a key technique in the fields such as road detection, camera calibration and visual navigation. This study presents a new vanishing point detection method, which delivers efficiency by using a dark channel prior‐based segmentation method and an adaptive straight lines search mechanism in the road region. First, the dark channel prior information is used to segment the image into a series of regions. Then the straight lines are extracted from the region contours, and the straight lines in the road region are estimated by a vertical envelope and a perspective quadrilateral constraint. The vertical envelope roughly divides the whole image into sky region, vertical region and road region. The perspective quadrilateral constraint, as the authors defined herein, eliminates the vertical lines interference inside the road region to extract the approximate straight lines in the road region. Finally, the vanishing point is estimated by the meanshift clustering method, which are computed based on the proposed grouping strategies and the intersection principles. Experiments have been conducted with a large number of road images under different environmental conditions, and the results demonstrate that the authors’ proposed algorithm can estimate vanishing point accurately and efficiently in unstructured road scenes.
Weili Ding, Yong Li 0028, Honghai Liu 0001
IET Comput. Vis.3
2016 Hand posture recognition based on heterogeneous features fusion of multiple kernels learning
Jiangtao Cao, Siquan Yu, Honghai Liu 0001, Ping Li 0012
Multim. Tools Appl.3
2016 A fusion method for robust face tracking
Xiaodong Jiang, Hui Yu 0001, Yang Lu 0003, Honghai Liu 0001
Multim. Tools Appl.4
2016 A novel approach to extract hand gesture feature in depth images
Zhaojie Ju, Dongxu Gao, Jiangtao Cao, Honghai Liu 0001
Multim. Tools Appl.4
2016 Guest Editorial: Advanced Understanding and Modelling of Human Motion in Multidimensional Spaces
Hui Yu 0001, Junyu Dong, Tuan D. Pham, Honghai Liu 0001
Multim. Tools Appl.4
2016 Accurately estimating rigid transformations in registration using a boosting-inspired mechanism
Yonghuai Liu, Honghai Liu 0001, Ralph R. Martin, Luigi De Dominicis, Ran Song 0001, Yitian Zhao
Pattern Recognit.2
2015 Dynamic facial expression recognition using local patch and LBP-TOP
abstract
Local binary pattern on three orthogonal planes (LBP-TOP) is one of the most popular method for dynamic texture analysis and has been successfully applied to facial expression analysis. Yet an effective LBP-TOP operator highly relies on preprocessing. And, like many appearance-based approaches, this approach reserves more identity-related cues rather than expression. In this work, we propose a fully automatic approach for facial expression recognition based on points registration, localized patch extraction and LBP-TOP feature representation. The efficiency of this method is evaluated on CK+ database. Results show that the proposed method has achieved a better performance compared with existing methods.
Yiming Wang 0001, Hui Yu 0001, Brett Stevens, Honghai Liu 0001
HSI4
2015 Combining Kinect and PnP for camera pose estimation
abstract
This paper presents a novel method to conduct camera pose estimation though combining Kinect and Perspective-n-points algorithms. Most existing camera pose estimation methods suffer from the errors caused by inevitable outliers between 2D-3D correspondences. To this end, we propose to use a random down sampling process to deal with outliers in this paper. The proposed method is divided into two main steps, which are 2D-3D correspondences generation and pose estimation. The method has been tested in a real project, and the experiment has shown encouraging results compared to the ground truth.
Shu Zhang 0002, Hui Yu 0001, Junyu Dong, Ting Wang 0018, Lin Qi 0004, Honghai Liu 0001
HSI6
2015 Real time object tracking via a mixture model
abstract
Object tracking has been applied in many fields such as intelligent surveillance and computer vision. Although much progress has been made, there are still many puzzles which pose a huge challenge to object tracking. Currently, the problems are mainly caused by appearance model as well as real-time performance. A novel method was been proposed in this paper to handle both of these problems. Locally dense contexts feature and image information (i.e. the relationship between the object and its surrounding regions) are combined in a Bayes framework. Then the tracking problem can be seen as a prediction question which need to compute the posterior probability. Both scale variations and temple updating are considered in the proposed algorithm to assure the effectiveness. To make the algorithm runs in a real time system, a Fourier Transform (FT) is used when solving the Bayes equation. Therefore, the MMOT (Mixture model for object tracking) runs in real-time and performs better than state-of-the-art algorithms on some challenging image sequences in terms of accuracy, quickness and robustness.
Dongxu Gao, Zhaojie Ju, Jiangtao Cao, Honghai Liu 0001
RO-MAN4
2015 A New Wearable Ultrasound Muscle Activity Sensing System for Dexterous Prosthetic Control
abstract
In this paper we introduce a novel Wearable Ultrasound Radial Muscle Activity Detection System (WURMADS) and a demonstration of its ability to recognize forearm muscle activities of amputee subjects to control a dexterous prosthetic hand. The system consists of control electronics to capture and record the ultrasound echo signals in real-time and two wearable bands embedded with eight ultrasound transducers. Based on the principles of Sonomyography, we recognized the intended isotonic finger gestures (trans-radial amputees) by analysing the echo patterns captured by eight single element transducers. Conventional non-invasive my electric detection strategy, Surface Electro Myography (sEMG), cannot reliably identify the deeper muscle activations in the forearm due to crosstalk and signal attenuation. It also suffers from signal degradation due to muscle fatigue and non-linearity. For this study we custom designed 1D single element waterproof transducers to be assembled in a radial structure. A real-time capturing electronic system was also designed, which is portable and wearable. The system operates at 5MHz, the frequency that demonstrated optimum performance for forearm applications. Ten trans-radial amputee subjects were employed in this study to capture data by executing five supervised gestures. The results obtained from offline analysis show that good gesture recognition rates can be achieved, while maintaining a decent correlation coefficient margin to distinguish gestures.
Nalinda Hettiarachchi, Zhaojie Ju, Honghai Liu 0001
SMC3
2015 Automatic Reconstruction of Dense 3D Face Point Cloud with a Single Depth Image
abstract
Human face analysis is the basis for many other computer vision tasks, such as camera surveillance, entrance authorization and age estimation. With 3D face models, the vision task based on facial analysis can usually achieve a higher accuracy than the 2D cases since it provides more information with the additional dimension. However, most existing 3D face reconstruction methods suffer from complicated processing and high computation. This paper presents a novel method that simplifies the 3D face reconstruction process with only one shot of Kinect data. The output of the system is a high density of 3D face point cloud with smoother surface. This provides rich details of the human face for other computer vision tasks. Experiments with real world data show promising results using the proposed method.
Shu Zhang 0002, Hui Yu 0001, Junyu Dong, Ting Wang 0018, Zhaojie Ju, Honghai Liu 0001
SMC6
2015 Visual tracking based on improved foreground detection and perceptual hashing
Mengjuan Fei, Jing Li 0027, Honghai Liu 0001
Neurocomputing3
2015 Time series modeling of surface EMG based hand manipulation identification via expectation maximization algorithm
Yang Lu 0003, Zhaojie Ju, Yurong Liu, Yuxuan Shen, Honghai Liu 0001
Neurocomputing5
2015 Pattern recognition technologies for multimedia information processing
Sang-Soo Yeo, Ken Chen 0002, Honghai Liu 0001
Multim. Tools Appl.3
2015 Flow field texture representation-based motion segmentation for crowd counting
Haiming He, Shukai Cao, Honghai Liu 0001
Mach. Vis. Appl.4
2014 Finger pinch force estimation through muscle activations using a surface EMG sleeve on the forearm
abstract
For prosthetic hand manipulation, the surface Electromyography(sEMG) has been widely applied. Researchers usually focus on the recognition of hand grasps or gestures, but ignore the hand force, which is equally important for robotic hand control. Therefore, this paper concentrates on the methods of finger forces estimation based on multichannel sEMG signal. A custom-made sEMG sleeve system omitting the stage of muscle positioning is utilised to capture the sEMG signal on the forearm. A mathematic model for muscle activation extraction is established to describe the relationship between finger pinch forces and sEMG signal, where the genetic algorithm is employed to optimise the coefficients. The results of experiments in this paper shows three main contributions: 1) There is a systematical relationship between muscle activations and the pinch finger forces. 2) To estimate the finger force, muscle precise positioning for electrodes placement is not inevitable. 3) In a multi-channel EMG system, selecting specific combinations of several channels can improve the estimation accuracy for specific gestures.
Yinfeng Fang, Zhaojie Ju, Honghai Liu 0001
FUZZ-IEEE4
2014 A modified EM algorithm for hand gesture segmentation in RGB-D data
abstract
This paper proposes a novel method with a modified Expectation-Maximisation (EM) Algorithm to segment hand gestures in the RGB-D data captured by Kinect. With the depth map and RGB image aligned by the genetic algorithm to estimate the key points from both depth and RGB images, a novel approach is proposed to refine the edge of the tracked hand gesture, which is used to segment the RGB image of the hand gestures, by applying a modified EM algorithm based on Bayesian networks. The experimental results demonstrated the modified EM algorithm effectively adjusts the RGB edges of the segmented hand gestures. The proposed methods have potential to improve the performance of hand gesture recognition in Human-Computer Interaction (HCI).
Zhaojie Ju, Yuehui Wang, Wei Zeng 0001, Haibin Cai, Honghai Liu 0001
FUZZ-IEEE5
2014 Grounding spatial relations in natural language by fuzzy representation for human-robot interaction
abstract
This paper addresses the issue of grounding spatial relations in natural language for human-robot interaction and robot control. The problem is approached by identifying two set of spatial relations, the image space-based and object-centered, and expressing them as fuzzy sets to capture the ambiguity inherent to the linguistic expressions for the relations. The sizes and shades of the scene objects have also been modeled as fuzzy sets for conditioning the spatial relations. To verify the validity of our approach and test its feasibility in a natural language-based interface, we have considered the typical scenarios of using the spatial relations in simple declarative and imperative sentences and designed simple grammars for parsing such sentences. Our experiment has shown that fuzzy spatial relation analysis provides a useful way for modeling the ambiguity or imprecision of the natural language in describing spatial relations and that it is possible to use the spatial relation models to support robot control and human-robot interaction in a natural language-based interface.
Jiacheng Tan, Zhaojie Ju, Honghai Liu 0001
FUZZ-IEEE3
2014 Image factorization and feature fusion for enhancing robot vision in human face recognition
abstract
Illumination variation has been a challenging problem for face recognition in robot vision. To reduce the effect caused by illumination variation, a lot of studies have been explored. The Total Variation (TV) method is particular used to factorize images into a low frequency component and a high frequency one. However, the low frequency component still contains significant intrinsic features resulting in failure in face recognition in some cases. In this paper, we propose to further extract illumination invariant features from face images under uncontrolled varying lighting conditions. The Nonsampled Contourlet Transform (NSCT) method is employed to enhance the extraction of intrinsic feature. The combined factorization model is very effective in the experiment on the Yale database.
Hui Yu 0001, Zhaojie Ju, Honghai Liu 0001
IJCNN3
2014 Linear regression for head pose analysis
abstract
Extensive research has been conducted to estimate and analyze head poses for various applications. Most existing methods tend to detect facial features and locate landmarks on a face for pose estimation. However, the sensitivity to occlusion of some face parts with key features and uncontrolled illumination of face images make the facial feature detection vulnerable. In this paper, we propose a framework for pose estimation without the need of face features or landmarks detection. Specifically, we formulate the pose estimation as a linear regression applied to the pose space. This method is based on the assumption that pose space cannot be linearly approximated in the pose subspace. The experimental results strongly support this assumption. In cases where the database does not obtain various poses in the intraclass, we propose to generate those poses through a 3D reconstruction and projection method. The experiment conducted on the CMU MultiPIE and IMM Face database has shown the effectiveness of the proposed method.
Hui Yu 0001, Honghai Liu 0001
IJCNN2
2014 Robust sEMG electrodes configuration for pattern recognition based prosthesis control
abstract
Electromyographic (EMG) signal is the electrical manifestation of a muscle contraction. Surface EMG signal can be obtained by electrodes on the skin to control prosthetic hand. However, surface EMG is sensitive to environmental interference, which leads to a low motion recognition rate of prosthesis control when encountering unexpected interferences, like electrodes shift. Electrodes shift occurs particularly in the day-to-day use of wearing electrodes. As a result, a long-term training procedure is necessary. To solve this problem, this paper proposes a new sEMG electrodes configuration to reduce the interference caused by electrodes shift. Experiments are designed to verify the improvements through evaluating the classification accuracy of discriminating eleven hand motions by pattern recognition approach. The comparison results show that the proposed electrodes configuration increases the pattern recognition rate by 4% and 8% when applied kNN and LDA classifier, respectively. This paper suggests that optimising electrodes configuration is able to improve the EMG pattern discrimination and the proposed electrodes configuration has reference value.
Yinfeng Fang, Honghai Liu 0001
SMC2
2014 A wireless wearable sEMG and NIRS acquisition system for an enhanced human-computer interface
abstract
Surface electromyography (sEMG) is extensively explored in human-computer interface (HCI); complementary to the electrophysiological activity of the muscles, the hemodynamic information that measured from near infrared spectroscopy (NIRS) is less investigated. Properly combining the sEMG and NIRS would provide a novel approach for HCI applications. This paper presents a multi-channel wireless wearable sEMG and NIRS acquisition system aiming for enhanced human-computer interaction, by providing more information about the muscle activity for subject's motor intention decoding. Extensive tests were carried out to evaluate the system performance. It showed that this novel system proved to be able to capture sEMG signals similar to those of the commercialized sEMG acquisition devices, and had a comparable NIRS sensor performance. Furthermore, simultaneously recording of sEMG and NIRS signals, the system had shown the ability to provide more information about the muscle activities for a better HCI performance. The classification accuracy of 13 hand gesture motions was significantly (P<;0.001) improved by using combined sEMG and NIRS features comparing to sEMG or NIRS features individually, suggesting that the proposed sEMG and NIRS system could be potentially available for an enhanced HCI.
Weichao Guo, Peng-Fei Yao, Xinjun Sheng, Honghai Liu 0001
SMC4
2014 Regression-Based Facial Expression Optimization
abstract
This paper presents an approach for reproducing optimal 3-D facial expressions based on blendshape regression. It aims to improve fidelity of facial expressions but maintain the efficiency of the blendshape method, which is necessary for applications such as human-machine interaction and avatars. The method intends to optimize the given facial expression using action units (AUs) based on the facial action coding system recorded from human faces. To help capture facial movements for the target face, an intermediate model space is generated, where both the target and source AUs have the same mesh topology and vertex number. The optimization is conducted interactively in the intermediate model space through adjusting the regulating parameter. The optimized facial expression model is transferred back to the target facial model to produce the final facial expression. We demonstrate that given a sketched facial expression with rough vertex positions indicating the intended facial expression, the proposed method approaches the sketched facial expression through automatically selecting blendshapes with corresponding weights. The sketched expression model is finally approximated through AUs representing true muscle movements, which improves the fidelity of facial expressions.
Hui Yu 0001, Honghai Liu 0001
IEEE Trans. Hum. Mach. Syst.2
2014 Dynamical Characteristics of Surface EMG Signals of Hand Grasps via Recurrence Plot
abstract
Recognizing human hand grasp movements through surface electromyogram (sEMG) is a challenging task. In this paper, we investigated nonlinear measures based on recurrence plot, as a tool to evaluate the hidden dynamical characteristics of sEMG during four different hand movements. A series of experimental tests in this study show that the dynamical characteristics of sEMG data with recurrence quantification analysis (RQA) can distinguish different hand grasp movements. Meanwhile, adaptive neuro-fuzzy inference system (ANFIS) is applied to evaluate the performance of the aforementioned measures to identify the grasp movements. The experimental results show that the recognition rate (99.1%) based on the combination of linear and nonlinear measures is much higher than those with only linear measures (93.4%) or nonlinear measures (88.1%). These results suggest that the RQA measures might be a potential tool to reveal the sEMG hidden characteristics of hand grasp movements and an effective supplement for the traditional linear grasp recognition methods.
Gaoxiang Ouyang, Zhaojie Ju, Honghai Liu 0001
IEEE J. Biomed. Health Informatics4
2013 Explore New Eye Tracking and Gaze Locating Methods
abstract
Eye tracking has been used extensively in research, often for plotting the gaze location of research participants. Many of the existing methods on tracking gaze locations, however, require special equipment. This equipment can be very expensive and/or cumbersome to use, restricting the use of it to specialist labs. By producing a method of eye tracking which requires minimal equipment and set up time, eye tracking might be more widely used as a research tool, especially in exploratory areas where the outcomes of the research are not known. This paper introduces two novel eye tracking methods which use only a web cam feed as input and require no specialist equipment. The camera used in this research is a standard web cam built into a normal laptop, any current commercial web cam could be used. Experiments using the proposed method showed more accurate tracking results than some existing eye trackers.
Alexander Kadyrov, Hui Yu 0001, Joe Eyles, Honghai Liu 0001
SMC4
2013 Ship Detection and Segmentation Using Image Correlation
abstract
There have been intensive research interests in ship detection and segmentation due to high demands on a wide range of civil applications in the last two decades. However, existing approaches, which are mainly based on statistical properties of images, fail to detect smaller ships and boats. Specifically, known techniques are not robust enough in view of inevitable small geometric and photometric changes in images consisting of ships. In this paper a novel approach for ship detection is proposed based on correlation of maritime images. The idea comes from the observation that a fine pattern of the sea surface changes considerably from time to time whereas the ship appearance basically keeps unchanged. We want to examine whether the images have a common unaltered part, a ship in this case. To this end, we developed a method - Focused Correlation (FC) to achieve robustness to geometric distortions of the image content. Various experiments have been conducted to evaluate the effectiveness of the propose.
Alexander Kadyrov, Hui Yu 0001, Honghai Liu 0001
SMC3
2013 New Perception of Fluffy Surfaces
abstract
Scene material recognition is a valuable perceptual ability for robots. However, current methods fail to provide robust solution for indoor surrounding material recognition. As human beings are able to immediately understand properties of surrounding materials, especially when it concerns smooth fluffy materials, we believe robots should be equipped with similar abilities. In this paper we explore a new idea for fluffy surface perception for robots based on above hypothesis. The aim is to enable robots to distinguish smooth and fluffy surfaces without touching them. This is achieved through calculating image match ability map from video cameras. Through measuring the similarity of images captured from different viewpoints, robots are able to immediately recognize whether it is a fluffy material. The method has been validated by primary experiments. Our results show that robots can have a sense of material properties for fluffy materials without touching them or any prior knowledge.
Alexander Kadyrov, Hui Yu 0001, Honghai Liu 0001
SMC3
2013 Stability Analysis of Polynomial-Fuzzy-Model-Based Control Systems Using Switching Polynomial Lyapunov Function
abstract
This paper investigates the stability problem of polynomial-fuzzy-model-based control system, which is formed by a polynomial fuzzy model and a polynomial fuzzy controller connected in a closed loop. A switching polynomial Lyapunov function consisting of a number of local polynomial Lyapunov functions is proposed to investigate the system stability. It demonstrates a nice property in favor of the stability analysis that each local polynomial Lyapunov function transits continuously to each other. As different local polynomial Lyapunov functions are employed to investigate the system stability according to the operating domain, relaxed stability conditions compared with the stability analysis result with a common Lyapunov function can be developed. In order to allow a greater design flexibility for the polynomial fuzzy controller, the proposed polynomial-fuzzy-model-based control scheme does not require that both the polynomial fuzzy model and polynomial fuzzy controller share the same premise membership functions. Stability conditions in terms of sum of squares are obtained to guarantee system stability and facilitate control synthesis. Simulation examples are given to verify the stability analysis results and demonstrate the effectiveness of the proposed polynomial fuzzy control scheme.
Hak-Keung Lam, Mohammad Narimani, Hongyi Li 0001, Honghai Liu 0001
IEEE Trans. Fuzzy Syst.4
2013 Intelligent Video Systems and Analytics: A Survey
abstract
Recent technology and market trends have demanded the significant need for feasible solutions to video/camera systems and analytics. This paper provides a comprehensive account on theory and application of intelligent video systems and analytics. It highlights the video system architectures, tasks, and related analytic methods. It clearly demonstrates that the importance of the role that intelligent video systems and analytics play can be found in a variety of domains such as transportation and surveillance. Research directions are outlined with a focus on what is essential to achieve the goals of intelligent video systems and analytics.
Honghai Liu 0001, Shengyong Chen, Naoyuki Kubota
IEEE Trans. Ind. Informatics1
2012 A generalised framework for analysing human hand motions based on multisensor information
abstract
In this paper, an integrated framework with multiple sensory information for analysing human hand motions is proposed, and it consists of components of system integration, signal preprocessing, correlation study of sensory information and human motion recognition based on manipulation intention. Three types of sensors are employed in the framework to simultaneously capture the finger angle trajectory, the hand contact force and the forearm electromyography (EMG) signal. The signal preprocessing module is to facilitate the rapid acquisition of human hand tasks by automatically synchronising and segmenting the manipulation primitives. Correlations of the sensory information are studied by using Empirical Copula and demonstrate there exist significant relationships between muscle signals and finger trajectories and between muscle signals and contact forces. In addition, motion recognition based on the EMG intention is investigated by using both Gaussian Mixture Models (GMMs) and Support Vector Machine (SVM) and discussion of the comparative results is presented.
Zhaojie Ju, Honghai Liu 0001
FUZZ-IEEE2
2012 Surface EMG signals determinism analysis based on recurrence plot for hand grasps
abstract
This paper proposes determinism measure (DET) based on recurrence plot, which is capable of showing the recurrence property of a deterministic dynamical system, to evaluate the dynamical characteristics of the surface electromyogram (sEMG) during three different hand movements. In addition, the linear discriminant analysis (LDA) is applied to evaluate the performance of the above measures to identify these three hand grasp movements. The experimental result shows that the recognition rate, 96.7%, based on the combination of the linear and non-linear measures is much higher than that with only linear measures, and DET might be a potential tool to reveal the sEMG hidden characteristics of hand grasp movements.
Gaoxiang Ouyang, Zhaojie Ju, Honghai Liu 0001
IJCNN3
2012 Motion Detection based on Simulated Depth Measurement
abstract
Depth information is a very important cue to understand human motion. In this paper, we establish that, even with no real depth camera, the concept of obtaining depth information is applicable for human motion detection. We propose a new motion detection method based on the concept of a real world video surveillance system enhanced with depth cameras. It is developed for detecting and analysing human motion. First, it imitates depth measuring process of a depth camera. Specially chosen in the image during the initialization process, view points play the role of cameras, whereas the depth is measured as a distance from these points to the human figure in the image. Initially, the body is partitioned into four segments to obtain the information about which part of the body is moving. Then, in course of the working cycle of the method, the received depth values are constantly subtracted from the previously obtained values, and the intensity of the body motion is calculated using root mean square. The method has been tested on actions taken from a standard motion dataset (IXMAS). It proved to be stable and reliable.
Chern Hong Lim, Alexander Kadyrov, Chee Seng Chan, Honghai Liu 0001
KES4
2012 Fuzzy Gaussian Mixture Models
Zhaojie Ju, Honghai Liu 0001
Pattern Recognit.2
2012 Reliable Fuzzy Control for Active Suspension Systems With Actuator Delay and Fault
abstract
This paper is focused on reliable fuzzy$H_{\infty }$controller design for active suspension systems with actuator delay and fault. The Takagi–Sugeno (T–S) fuzzy model approach is adapted in this study with the consideration of the sprung and the unsprung mass variation, the actuator delay and fault, and other suspension performances. By the utilization of the parallel-distributed compensation scheme, a reliable fuzzy$H_{\infty }$performance analysis criterion is derived for the proposed T–S fuzzy model. Then, a reliable fuzzy$H_{\infty }$controller is designed such that the resulting T–S fuzzy system is reliable in the sense that it is asymptotically stable and has the prescribed$H_{\infty }$performance under given constraints. The existence condition of the reliable fuzzy$H_{\infty }$controller is obtained in terms of linear matrix inequalities (LMIs) Finally, a quarter-vehicle suspension model is used to demonstrate the effectiveness and potential of the proposed design techniques.
Hongyi Li 0001, Honghai Liu 0001, Huijun Gao, Peng Shi 0001
IEEE Trans. Fuzzy Syst.2
2012 Guest Editorial Special Section on Intelligent Video Systems and Analytics
abstract
The 11 papers in this special section focus on intelligent video systems and analytics.
Honghai Liu 0001, Shengyong Chen, Naoyuki Kubota
IEEE Trans. Ind. Informatics1
2012 Classification of Upper Limb Motion Trajectories Using Shape Features
abstract
To understand and interpret human motion is a very active research area nowadays because of its importance in sports sciences, health care, and video surveillance. However, classification of human motion patterns is still a challenging topic because of the variations in kinetics and kinematics of human movements. In this paper, we present a novel algorithm for automatic classification of motion trajectories of human upper limbs. The proposed scheme starts from transforming 3-D positions and rotations of the shoulder/elbow/wrist joints into 2-D trajectories. Discriminative features of these 2-D trajectories are, then, extracted using a probabilistic shape-context method. Afterward, these features are classified using a k-means clustering algorithm. Experimental results demonstrate the superiority of the proposed method over the state-of-the-art techniques.
Huiyu Zhou 0001, Huosheng Hu, Honghai Liu 0001, Jinshan Tang
IEEE Trans. Syst. Man Cybern. Part C3
2012 Neural-Network-Based Decentralized Adaptive Output-Feedback Control for Large-Scale Stochastic Nonlinear Systems
abstract
This paper focuses on the problem of neural-network-based decentralized adaptive output-feedback control for a class of nonlinear strict-feedback large-scale stochastic systems. The dynamic surface control technique is used to avoid the explosion of computational complexity in the backstepping design process. A novel direct adaptive neural network approximation method is proposed to approximate the unknown and desired control input signals instead of the unknown nonlinear functions. It is shown that the designed controller can guarantee all the signals in the closed-loop system to be semiglobally uniformly ultimately bounded in a mean square. Simulation results are provided to demonstrate the effectiveness of the developed control design approach.
Qi Zhou 0002, Peng Shi 0001, Honghai Liu 0001, Shengyuan Xu 0001
IEEE Trans. Syst. Man Cybern. Part B3
2011 Hand motion recognition via fuzzy active curve axis Gaussian mixture models: A comparative study
abstract
Unconstrained human hand motions consisting grasp motion and in-hand manipulation lead to a fundamental challenge that many algorithms have to face in both theoretical and practical development, mainly due to the complexity and dexterity of the human hand. In this paper, fuzzy active curve axis Gaussian Mixture Model (FAcaGMM) is proposed by introducing a weighting exponent on the fuzzy membership into active curve axis Gaussian Mixture Models (AcaGMM) to improve its convergence efficiency, and then FAcaGMM is used to recognize human hand motions. In addition, a comparative study of recognition methods including FAcaGMM, Time Clustering (TC), Empirical Copula (EC), GMM and HMM is presented to recognize human hand motions including both grasps and in-hand manipulations from different subjects with varying training samples.
Zhaojie Ju, Honghai Liu 0001
FUZZ-IEEE2
2011 Actuator delayed active vehicle suspension control: A T-S fuzzy approach
abstract
This paper focuses on fuzzy H∞controller design for uncertain active suspension systems with actuator delay based on Takagi-Sugeno (T-S) model approach. This dynamic system is presented by taking into account the sprung and unsprung mass variations, the actuator delay, and the suspension performance. The fuzzy H∞controller is designed such that the resulting T-S fuzzy system is asymptotically stable and guarantees H∞performance, and simultaneous satisfying the constraint performance. The existence condition of fuzzy H∞control is obtained in terms of linear matrix inequalities (LMIs) and can be solved by using the standard software. A quarter-car suspension model is provided to validate the effectiveness of the proposed design procedures.
Hongyi Li 0001, Honghai Liu 0001, Huijun Gao
FUZZ-IEEE2
2011 Recent Advances in Fuzzy Qualitative Reasoning
abstract
A reliable human skin detection method that is adaptable to different human skin colors and illumination conditions is essential for better human skin segmentation. Even though different human skin-color detection solutions have been successfully applied, they are prone to false skin detection and are not able to cope with the variety of human skin colors across different ethnic. Moreover, existing methods require high computational cost. In this paper, we propose a novel human skin detection approach that combines a smoothed 2-D histogram and Gaussian model, for automatic human skin detection in color image(s). In our approach, an eye detector is used to refine the skin model for a specific person. The proposed approach reduces computational costs as no training is required, and it improves the accuracy of skin detection despite wide variation in ethnicity and illumination. To the best of our knowledge, this is the first method to employ fusion strategy for this purpose. Qualitative and quantitative results on three standard public datasets and a comparison with state-of-the-art methods have shown the effectiveness and robustness of the proposed approach.
Chee Seng Chan, George Macleod Coghill, Honghai Liu 0001
Int. J. Uncertain. Fuzziness Knowl. Based Syst.3
2011 Fuzzy Qualitative Reasoning about Dynamic Systems Containing Trigonometric Relationships
abstract
In this paper we present a system which incorporates the features of Fuzzy Qualitative Trigonometry (FQT) with those of the Fuzzy Qualitative Reasoning system, Morven. FQT is designed for modeling and reasoning about robot kinematic systems whereas Morven was designed for simulation and envisionment of straightforward dynamic systems. It has been the case that thus far QR systems have not been designed to reason about the behaviour of dynamic system containing fuzzy trigonometric relations. The resulting tool described in this paper goes some way to addressing this deficit.
George Macleod Coghill, Honghai Liu 0001, Allan M. Bruce, Carol Wisley
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2
2011 A Unified Fuzzy Framework for Human-Hand Motion Recognition
abstract
Unconstrained human-hand motions that consist grasp motions and in-hand manipulations lead to a fundamental challenge that many algorithms have to face in both theoretical and practical development, mainly due to the complexity and dexterity of the human hand. There is no effective solution reported to recognize in-hand manipulations, although recognition algorithms have been proposed to recognize grasp motions in constrained scenarios. This paper proposes a novel unified fuzzy framework of a set of recognition algorithms: time clustering, fuzzy active axis Gaussian mixture mode, and fuzzy empirical copula, from numerical clustering to data dependence structure in the context of optimally real-time human-hand motion recognition. Time clustering is a fuzzy time-modeling approach that is based on fuzzy clustering and Takagi--Sugeno modeling with a numerical value as output. The fuzzy active axis Gaussian mixture model effectively extract abstract Gaussian pattern to represent components of hand gestures with a fast convergence. A fuzzy empirical copula utilizes the dependence structure among the finger joint angles to recognize the motion type. The proposed algorithms have been evaluated on a wide range of scenarios of human-hand recognition: 1) datasets that include 13 grasps and ten in-hand manipulations; 2) single subject and multiple subjects; and 3) varying training samples. The experimental results have demonstrated that the proposed framework outperforms the hidden Markov model (HMM) and Gaussian mixture model in terms of both effectiveness and efficiency criteria.
Zhaojie Ju, Honghai Liu 0001
IEEE Trans. Fuzzy Syst.2
2011 Exploring Human Hand Capabilities Into Embedded Multifingered Object Manipulation
abstract
This paper provides a comprehensive computational account of hand-centered research, which is principles, methodologies and practical issues behind human hands, robot hands, rehabilitation hands, prosthetic hands and their applications. In order to help readers understand hand-centered research, this paper presents recent scientific findings and technologies including human hand analysis and synthesis, hand motion capture, recognition algorithms and applications, it serves the purpose of how to transfer human hand manipulation skills to related hand-centered applications in a computational context. The concluding discussion assesses the progress thus far and outlines some research challenges and future directions, and solution to which is essential to achieve the goals of human hand manipulation skill transfer. It is expected that the survey will also provide profound insights into an in-depth understanding of real-time hand-centered algorithms, human perception-action, and potential hand-centered healthcare solutions.
Honghai Liu 0001
IEEE Trans. Ind. Informatics1
2011 Computational Analysis of Sparse Datasets for Fault Diagnosis in Large Tribological Mechanisms
abstract
This paper presents the most up-to-date methods for the task of designing a system to accurately classify abnormal events, or faults, in a complex tribological mechanism, using elemental analysis of lubrication oil as an indicator of engine condition. The discussion combines perspectives from numerous fault diagnosis applications, both online and offline, to focus upon the task of offline event detection and diagnosis of datasets from elemental analysis, and although this does not suffer from complexity issues as in real-time processing, it introduces a number of other problems such as sparsity and selecting an accurate knowledge representation as well as reasoning under uncertainty and ignorance. The role of confounding variables is significant in sparse datasets, and as such this paper demonstrates an alternative perspective on both eliminating to an extent the effect of confounding variables and inferring unseen variables from measured variables. There has been little review work on this subject, and as a result this paper helps to join disparate research from a number of different domains to achieve some unification of alternative perspectives. This paper concludes by providing a case study to identify the methods that can be utilized in combination.
Ian Morgan, Honghai Liu 0001
IEEE Trans. Syst. Man Cybern. Part C2
2010 A two-stage pattern matching method for speaker recognition of partner robots
abstract
By using human speech information, different kinds of speaker and speech recognition systems have been developed for partner robots to efficiently cooperate with people in the daily life. For improving the recognition accuracy and robustness, a two-stage pattern matching algorithm for speaker recognition system of partner robots is proposed. In the first matching stage, by using fuzzy c-means and declustering in vector quantization(VQ) method, the recognition performance with limited training data is improved. For avoiding the phenomenon of similar cepstral features by different speakers, with three additional speech features, the second stage is designed to rematch the similar recognition results of the first stage. In order to evaluate the proposed structure, some experiments have implemented on a public database ELSDSR and an speech owners database for partner robots. The results verified the proposed method obtained more accurate recognition results with strong robustness.
Jiangtao Cao, Naoyuki Kubota, Honghai Liu 0001
FUZZ-IEEE3
2010 Fuzzy qualitative complex actions recognition
abstract
Understanding actions is a complex issue in many aspects. However, most of the literature on action recognition deals with only simple actions. In this paper, we proposed the fuzzy qualitative robot kinematics framework to complex actions over time, e.g. walk then run and over the body, walk while wave hand etc. The human limbs is modelled as articulated rigid bodies and its motion is represented by a series of such models in terms of time. With this, we eventually converted the human motion analysis into a conventional robotic problem which has been well studied. Experimental results has shown that the action model built in this manifold offers few advantages. e.g. handles the tradeoffs in the off-the-shelf tracking algorithm and avoid using generative model where the size of the training data typically goes as the square of the number of states.
Chee Seng Chan, Honghai Liu 0001, Weng-Kin Lai
FUZZ-IEEE2
2010 Human actions recognition using Fuzzy PCA and discriminative hidden model
abstract
As a temporal classification problem, visual-based human actions recognition is an important component for some potential applications. In this paper, we combine Fuzzy Principle Component Analysis(Fuzzy PCA) and hidden Conditional Random Fields(HCRFs) to achieve a viewpoint insensitive human action recognition. Fuzzy PCA is used to reduce the dimension of the silhouette image features to obtain the compact representation of action space. HCRFs is applied to model the human actions from different actors and different viewpoints. This method can relax the independence assumption of the generative model. Experiment results on a public dataset demonstrate the effectiveness and robustness of our method.
Xiaofei Ji, Honghai Liu 0001
FUZZ-IEEE2
2010 Applying fuzzy EM algorithm with a fast convergence to GMMs
abstract
Inspired from the mechanism of Fuzzy C-means (FCMs) which introduces a degree of fuzziness on the dissimilarity function based on distances, a fuzzy Expectation Maximization (EM) algorithm for Gaussian Mixture Models (GMMs) is proposed in this paper. In the fuzzy EM algorithm, the dissimilarity function is defined as the multiplicative inverse of probability density function. Different from FCMs, the defined dissimilarity function is based on the exponential function of the distance. The fuzzy EM algorithm is compared with normal EM algorithm in terms of fitting degree and convergence speed. The experimental results in modeling random data and various characters demonstrate the ability of the proposed algorithm in reducing the computational cost of GMMs.
Zhaojie Ju, Honghai Liu 0001
FUZZ-IEEE2
2010 Extending evolutionary Fuzzy Quantile Inference to classify partially occluded human motions
abstract
This work presents a framework that combines the concept of Fuzzy Quantile Inference (FQI) with Genetic Programming (GP) in order to accurately classify real natural 3d human Motion Capture data. FQI is a generalization of Fuzzy Gaussian Inference. It builds Fuzzy Membership Functions that map to hidden Probability Distributions underlying human motions, providing a suitable modelling paradigm for such noisy data. Genetic Programming (GP) is used to make a time dependent and context aware filter that improves the qualitative output of the classifier. Results show that FQI outperforms a GMM-based classifier when recognizing six different boxing stances simultaneously, and that the addition of the GP based filter improves the accuracy of the FQI classifier significantly. A mechanism allowing the FQI extended framework to deal with occluded data reasonably well is also integrated.
Mehdi Khoury, Honghai Liu 0001
FUZZ-IEEE2
2010 Human hand motion recognition using Empirical Copula
abstract
Programming by Demonstration (PbD) enables robotic hands to learn human manipulation skills through storing motion primitives and recognizing motion types. In this paper, Empirical Copula is introduced to recognize dynamic human hand motions for the first time using the proposed motion template and matching algorithm. The huge computational cost of Empirical Copula is alleviated by the proposed re-sampling processing. The experiments with human hand motions including grasps and in-hand manipulations demonstrate Empirical Copula outperforms the Time Clustering (TC) method, Gaussian Mixture Models (GMMs) and Hidden Markov Models (HMMs) in terms of recognition rate. In addition, Empirical Copula is also proved to be able to recognize different motions from different subjects.
Zhaojie Ju, Honghai Liu 0001
IROS2
2010 Viewpoint Insensitive Actions Recognition Using Hidden Conditional Random Fields
Xiaofei Ji, Honghai Liu 0001
KES (1)2
2010 An Interval Fuzzy Controller for Vehicle Active Suspension Systems
abstract
A novel interval type-2 fuzzy controller architecture is proposed to resolve nonlinear control problems of vehicle active suspension systems. It integrates the Takagi-Sugeno (T-S) fuzzy model, interval type-2 fuzzy reasoning, the Wu-Mendel uncertainty bound method, and selected optimization algorithms together to construct the switching routes between generated linear model control surfaces. The stability analysis of the proposed approach is presented. The proposed method is implemented into a numerical example and a case study on a nonlinear half-vehicle active suspension system. The simulation results demonstrate the effectiveness and efficiency of the proposed approach.
Jiangtao Cao, Ping Li 0012, Honghai Liu 0001
IEEE Trans. Intell. Transp. Syst.3
2010 Advances in View-Invariant Human Motion Analysis: A Review
abstract
As viewpoint issue is becoming a bottleneck for human motion analysis and its application, in recent years, researchers have been devoted to view-invariant human motion analysis and have achieved inspiring progress. The challenge here is to find a methodology that can recognize human motion patterns to reach increasingly sophisticated levels of human behavior description. This paper provides a comprehensive survey of this significant research with the emphasis on view-invariant representation, and recognition of poses and actions. In order to help readers understand the integrated process of visual analysis of human motion, this paper presents recent development in three major issues involved in a general human motion analysis system, namely, human detection, view-invariant pose representation and estimation, and behavior understanding. Public available standard datasets are recommended. The concluding discussion assesses the progress so far, and outlines some research challenges and future directions, and solution to what is essential to achieve the goals of human motion analysis.
Xiaofei Ji, Honghai Liu 0001
IEEE Trans. Syst. Man Cybern. Part C2
2009 A switching fuzzy control method for the magnetic active suspension system
abstract
A switching fuzzy control system is proposed for dealing with the non-linear dynamics of electromagnetic suspension system. With two fuzzy sub-controllers and a switch engine, the switching fuzzy control system is flexible to cover changeable initial conditions with less computational cost. For satisfying the coupling constraint on positions of four floaters, a global self-supervisor with feedback structure is designed to control all four fuzzy subsystems. Simulations on the magnetic suspension with three different initial position settings demonstrated the efficiency of proposed method.
Jiangtao Cao, Zhaojie Ju, Xiaofei Ji, Honghai Liu 0001
FUZZ-IEEE4
2009 An extended fuzzy logic system for uncertainty modelling
abstract
An extended fuzzy logic system (EFLS) based on interval fuzzy membership functions is proposed for covering more uncertainty in practical applications. With the degree of uncertainty in fuzzy membership functions, interval fuzzy membership functions are self-generated to include uncertainties which occur from understanding linguistic knowledge and fuzzy rules in fuzzy methods. A novel adaptive strategy is designed to self-tune the interval fuzzy membership functions and to deduce the crisp outputs with feedback structure. An inverse kinematics modelling study based on a two-joint robotic arm has demonstrated that proposed EFLS outperforms conventional fuzzy methods.
Jiangtao Cao, Ping Li 0012, Honghai Liu 0001
FUZZ-IEEE3
2009 GMM-QNT hybrid framework for vision-based human motion analysis
abstract
The understanding of human behaviour in video is a challenging task in that the same behaviour might have several different meanings depending upon the scene and task context in which it is performed. While human seem to perform scene interpretations without effort, this is a formidable and yet unsolved task for artificial vision systems. One of the main reasons is that there exists a gap between low-level vision at signal level and high-level representation of activities at symbolic level. In this paper, we present an intelligent connection framework using Gaussian mixture model-based clustering (GMM) to bridge the low-level vision data and the qualitative normalised templates (QNT) - a symbolic representation for human motion based on fuzzy qualitative robot kinematics, which could link the former with domain-dependent scenarios. The proposed method has been applied to the recognition of eight types of human motions and an empirical comparison with fuzzy hidden Markov-based human motion recognition system.
Chee Seng Chan, Honghai Liu 0001
FUZZ-IEEE2
2009 Fast estimating data dependence structure via fuzzy empirical copula
abstract
As a non-parametric algorithm, empirical copula is an effective way to estimate the dependence structure of high-dimension arbitrarily distributed data. However, it suffers from the problem of huge computation time because of its high computational complexity. In this paper, fuzzy empirical copula is proposed to solve this problem by combining the fuzzy clustering by local approximation of memberships (FLAME) with empirical copula. In the proposed algorithm, FLAME is extended from two-dimension data to high-dimension data and FLAME+is implemented to identify the highest density objects which represent the original dataset, and then empirical copula is used to estimate its independence structure according to the new dataset. Case studies have been carried out to demonstrate the effectiveness of the fuzzy empirical copula.
Zhaojie Ju, Honghai Liu 0001, Youlun Xiong
FUZZ-IEEE2
2009 Boxing motions classification through combining fuzzy gaussian inference with a context-aware rule-based system
abstract
This paper continues to explore the potential of newly introduced Fuzzy Gaussian Inference (FGI). It aims at constructing fuzzy membership functions by modelling hidden probability distributions underlying human motions. A fuzzy rule-based system has been employed to assist boxing motion classification from natural human Motion Capture data. In this experiment, FGI alone is able to recognise seven different boxing stances simultaneously with an accuracy superior to a GMM-based classifier. Results indicate that adding a Fuzzy Inference Engine on top of FGI improves the accuracy of the classifier in a consistent way.
Mehdi Khoury, Honghai Liu 0001
FUZZ-IEEE2
2009 Fuzzy qualitative trigonometry
Honghai Liu 0001, George Macleod Coghill, Dave P. Barnes
Int. J. Approx. Reason.1
2009 Fuzzy Qualitative Human Motion Analysis
abstract
This paper proposes a fuzzy qualitative approach to vision-based human motion analysis with an emphasis on human motion recognition. It achieves feasible computational cost for human motion recognition by combining fuzzy qualitative robot kinematics with human motion tracking and recognition algorithms. First, a data-quantization process is proposed to relax the computational complexity suffered from visual tracking algorithms. Second, a novel human motion representation, i.e., qualitative normalized template, is developed in terms of the fuzzy qualitative robot kinematics framework to effectively represent human motion. The human skeleton is modeled as a complex kinematic chain, and its motion is represented by a series of such models in terms of time. Finally, experiment results are provided to demonstrate the effectiveness of the proposed method. An empirical comparison with conventional hidden Markov model (HMM) and fuzzy HMM (FHMM) shows that the proposed approach consistently outperforms both HMMs in human motion recognition.
Chee Seng Chan, Honghai Liu 0001
IEEE Trans. Fuzzy Syst.2
2008 Adaptive fuzzy logic controller for vehicle active suspensions with interval type-2 fuzzy membership functions
abstract
Elicited from the least means squares optimal algorithm (LMS), an adaptive fuzzy logic controller (AFC) based on interval type-2 fuzzy sets is proposed for vehicle non-linear active suspension systems. The interval membership functions (IMF2s) are utilized in the AFC design to deal with not only non-linearity and uncertainty caused from irregular road inputs and immeasurable disturbance, but also the potential uncertainty of expertpsilas knowledge and experience. The adaptive strategy is designed to self-tune the active force between the lower bounds and upper bounds of interval fuzzy outputs. A case study based on a quarter active suspension model has demonstrated that the proposed type-2 fuzzy controller significantly outperforms conventional fuzzy controllers of an active suspension and a passive suspension.
Jiangtao Cao, Honghai Liu 0001, Ping Li 0012, David J. Brown 0002
FUZZ-IEEE2
2008 A fuzzy qualitative approach to human motion recognition
abstract
The understanding of human motions captured in image sequences pose two main difficulties which are often regarded as computationally ill-defined: 1) modelling the uncertainty in the training data, and 2) constructing a generic activity representation that can describe simple actions as well as complicated tasks that are performed by different humans. In this paper, these problems are addressed from a direction which utilises the concept of fuzzy qualitative robot kinematics [9]. First of all, the training data representing a typical activity is acquired by tracking the human anatomical landmarks in an image sequences. Then, the uncertainty arise when the limitations of the tracking algorithm are handled by transforming the continuous training data into a set of discrete symbolic representations - qualitative states in a quantisation process. Finally, in order to construct a template that is regarded as a combination ordered sequence of all body segments movements, robot kinematics, a well-defined solution to describe the resulting motion of rigid bodies that form the robot, has been employed. We defined these activity templates as qualitative normalised templates, a manifold trajectory of unique state transition patterns in the quantity space. Experimental results and a comparison with the hidden Markov models have demonstrated that the proposed method is very encouraging and shown a better successful recognition rate on the two available motion databases.
Chee Seng Chan, Honghai Liu 0001, David J. Brown 0002, Naoyuki Kubota
FUZZ-IEEE2
2008 A fuzzy qualitative framework for robot intelligent connection
abstract
This paper proposes a novel framework in fuzzy qualitative terms for an attempt of attacking the robot intelligent connection problem. Robot intelligent connection is in the context of closing the gap between symbolic or qualitative functions and numerical sensing and control tasks through a generalized robot kinematics. First, fuzzy qualitative robot kinematics is revisited which provides theoretical preliminaries for the proposed robot motion representation. Secondly, a motion representation based on Gaussian mixture models is presented where Gaussian functions are combined to model a multimodal density of fuzzy qualitative kinematics parameters of a robotic end effector using clustering. Finally, simulation results in a PUMA 600 robot demonstrated that the proposed method effectively provides a two-way connection for robot representations used for numerical and symbolic robot tasks.
Honghai Liu 0001
FUZZ-IEEE1
2008 Predictive unsupervised organisation in marine engine fault detection
abstract
This paper utilises topological learners, the self organising map in combination with the k means algorithm to organise potential engine faults and the respective location of faults, focussing on a 12 cylinder 2 stroke marine diesel engine. This method is applied to reduce the numerosity of the data presented to a user by selecting representative samples from a number of clusters to enable efficient diagnosis. The novelty of the approach centres around the sparsity of the dataset compared to the majority of fault diagnosis techniques, and the potential for improved safety and efficiency within the marine industry compared to existing diagnosis systems. The accuracy of the SOM and k means, as well as the neural gas algorithm is compared to the standard accuracy of the k means algorithm to validate the algorithmpsilas performance and application to this domain, where it can be seen that topological learners have much potential to be applied to the field of fault diagnosis.
Ian Morgan, Honghai Liu 0001, George Turnbull, David J. Brown 0002
IJCNN2
2008 Analysing Flight Data Using Clustering Methods
Christopher Jesse, Honghai Liu 0001, Edward Smart, David J. Brown 0002
KES (1)2
2008 Visual-Based View-Invariant Human Motion Analysis: A Review
Xiaofei Ji, Honghai Liu 0001, David J. Brown 0002
KES (1)2
2008 A Fuzzy Qualitative Framework for Connecting Robot Qualitative and Quantitative Representations
abstract
This paper proposes a novel framework for describing articulated robot kinematics motion with the goal of providing a unified representation by combining symbolic or qualitative functions and numerical sensing and control tasks in the context of intelligent robotics. First, fuzzy qualitative robot kinematics that provides theoretical preliminaries for the proposed robot motion representation is revisited. Second, a fuzzy qualitative framework based on clustering techniques is presented to connect numerical and symbolic robot representations. Built on thek-\bbAGOPoperator (an extension of the ordered weighted aggregation operators), k-means and Gaussian functions are adapted to model a multimodal density of fuzzy qualitative kinematics parameters of a robot in both Cartesian and joint spaces; on the other hand, a mixture regressor and interpolation method are employed to convert Gaussian symbols into numerical values. Finally, simulation results in a PUMA 560 robot demonstrated that the proposed method effectively provides a two-way connection for robot representations used for both numerical and symbolic robotic tasks.
Honghai Liu 0001
IEEE Trans. Fuzzy Syst.1
2008 Fuzzy Qualitative Robot Kinematics
abstract
We propose a fuzzy qualitative (FQ) version of robot kinematics with the goal of bridging the gap between symbolic or qualitative functions and numerical sensing and control tasks for intelligent robotics. First, we revisit FQ trigonometry, and then derive its derivative extension. Next, we replace the trigonometry role in robot kinematics using FQ trigonometry and the proposed derivative extension, which leads to a FQ version of robot kinematics. FQ transformation, position, and velocity of a serial kinematics robot are derived and discussed. Finally, we propose an aggregation operator to extract robot behaviors with the highlight of the impact of the proposed methods to intelligent robotics. The proposed methods have been integrated into XTRIG MATLAB toolbox and a case study on a PUMA robot has been implemented to demonstrate their effectiveness.
Honghai Liu 0001, David J. Brown 0002, George Macleod Coghill
IEEE Trans. Fuzzy Syst.1
2008 State of the Art in Vehicle Active Suspension Adaptive Control Systems Based on Intelligent Methodologies
abstract
This paper reviews computational-intelligence-involved approaches in active vehicle suspension control systems with a focus on the problems raised in practical implementations by their nonlinear and uncertain properties. After a brief introduction on active suspension models, the paper explores the state of the art in fuzzy inference systems, neural networks, genetic algorithms, and their combination for suspension control issues. Discussions and comments are provided based on the reviewed simulation and experimental results. The paper is concluded with remarks and future directions.
Jiangtao Cao, Honghai Liu 0001, Ping Li 0012, David J. Brown 0002
IEEE Trans. Intell. Transp. Syst.2
2008 Navigation Technologies for Autonomous Underwater Vehicles
abstract
With recent advances in battery capacity and the development of hydrogen fuel cells, autonomous underwater vehicles (AUVs) are being used to undertake longer missions that were previously performed by manned or tethered vehicles. As a result, more advanced navigation systems are needed to maintain an accurate position over a larger operational area. The accuracy of the navigation system is critical to the quality of the data collected during survey missions and the recovery of the AUV. Many different methods for navigation in different underwater environments have been proposed in the literature. In this correspondence paper, the state of the art in navigation technologies for AUVs is investigated for theoretical and operational systems. Their suitability for use in different environments is compared and current limitations of these methods are identified. In addition, new approaches to address these current problems and areas for future research are suggested. Finally, it is concluded that only geophysically referenced methods will enable AUVs to navigate accurately over large areas and that advances in underwater feature recognition are required before these methods can be implemented in operational AUVs.
L. Stutters, Honghai Liu 0001, C. Tiltman, David J. Brown 0002
IEEE Trans. Syst. Man Cybern. Part C2
2007 Accurate range image registration: Eliminating or modelling outliers
abstract
Automatic and accurate range image registration is often a prerequisite step for range image analysis and interpretation. Due to occlusion, appearance and disappearance of points in different images, outliers inevitably occur. In this case, various techniques to eliminate and model outliers have been proposed for accurate range image registration. The objective of this paper is to experimentally investigate which of the outlier elimination and modelling is more effective for the evaluation of possible correspondences established, so that a deep insight into how advanced range image registration algorithms will be developed can be obtained. The experimental results based on both synthetic data and real images show that the outlier modelling often outperforms the outlier elimination in the sense of producing more accurate and robust range image registration results.
Yonghuai Liu, Honghai Liu 0001, Longzhuang Li, Baogang Wei
ETFA2
2007 FHLS: Fuzzy Hierarchical Location Service for Mobile Ad Hoc Networks
abstract
Location services are used in mobile ad hoc and hybrid networks either to locate the geographic position of a given node in the network or for locating a data item. One of the main usages of position location services is in location based routing algorithms. In particular, geographic routing protocols can route messages more efficiently to their destinations based on the destination node's geographic position, which is provided by a location service. In this paper, we propose an adaptive location service on the basis of fuzzy logic called FHLS (fuzzy hierarchical location service) for mobile ad hoc networks. The FHLS uses the adaptive location update scheme using the fuzzy logic on the basis of the mobility and the call preference of mobile nodes. The performance of the FHLS is to be evaluated by using a simulation, and compared with that of existing HLS scheme.
Ihn-Han Bae, Honghai Liu 0001
FUZZ-IEEE2
2007 An Effective Human Motion Classification Approach using Knowledge Representation in Qualitative Normalised Templates
abstract
Classification of human motion in video data is essential in numerous applications. However, problems arise as the human exhibits complex and dynamic motion that is nonlinear and time varying. In this paper, we propose a knowledge-based human motion classification framework that employs fuzzy qualitative reasoning to address these problems. Our approach utilises the rich contextual information (e.g. structural and transitional characteristic of human motion) captured in video sequence to effectively study and recognise human motion. With the aid of domain knowledge, a set of fuzzy rules are defined in the knowledge base. This work is in contrast with previous attempts that depend solely on the trajectories of the body parts. Experimental results on two classes of motion (e.g. walking and running) that result in similar motions; and a comparison with the conventional method has demonstrated and validated the effectiveness of the proposed method in improving the perception of human motion.
Chee Seng Chan, Honghai Liu 0001, David J. Brown 0002
FUZZ-IEEE2
2007 An Approach to Robot Motion Behaviour Representation
abstract
We propose a set of algorithms for robot motion behaviour representation based on our previous research, fuzzy qualitative trigonometry and fuzzy qualitative robot kinematics. It aims at developing a fuzzy qualitative framework for the connection of qualitative and quantitative representations. Firstly, we revisit fuzzy qualitative robot kinematics and derive fuzzy qualitative robot kinematics in terms of D-H robotic structure, fuzzy qualitative trajectories of a PUMA robot have been provided to prove the proposed method. Then robot behaviour extraction has been studied using three algorithms which are an adopted ordered weighted averaging operators, k-means classifier and EM clustering. Finally, simulation results of the PUMA robot using the three algorithms have been demonstrated and discussed with concluding remarks.
Honghai Liu 0001, David J. Brown 0002
FUZZ-IEEE1
2007 Stabilization Control for a Class of Switched Fuzzy Discrete-Time Systems
abstract
Stability issues for switched systems whose subsystems are all Sugeno fuzzy discrete-time systems are studied and new stabilization control results derived. Innovated representation models for switched fuzzy systems are proposed. The single Lyapunov function method has been adopted to study the stability of this class of switched fuzzy systems. Sufficient conditions for quadratic asymptotic stability are presented and stabilizing switching laws of the state-dependent form employed in PDC scheme are designed. Illustrative examples demonstrate the effectiveness and the feasible performance of the proposed synthesis through the respective simulation results.
Honghai Liu 0001, Georgi M. Dimirovski, Jun Zhao 0002
FUZZ-IEEE2
2007 Fault-tolerant Control of Uncertain Time-delay Discrete-time Systems Using T-S Model
abstract
The investigated control problem of nonlinear time-delay discrete-time systems is addressed using Takagi-Sugeno model. Parametric uncertainty and time-delay terms are employed in building the Takagi-Sugeno model for the controlled plant for the purpose of close representation of the original plant system. The integrity and robustness are guaranteed for the closed-loop fuzzy system in sense of Lyapunov stability method via fuzzy state observers. Sufficient conditions for the fuzzy system to possess integrity against actuator failures and sensor failures in the closed-loop are derived in terms of linear matrix inequalities under assumption of known bounds on uncertainties. The results for trailer-truck example are used to illustrate the effectiveness of the method.
Yuanwei Jing, Rong-Zhen Chen, Honghai Liu 0001, Georgi M. Dimirovski
FUZZ-IEEE4
2007 Qualitative kinematics of planar robots: Intelligent connection
Honghai Liu 0001, George Macleod Coghill, David J. Brown 0002
Int. J. Approx. Reason.1
2006 An Extension to Fuzzy Qualitative Trigonometry and Its Application to Robot Kinematics
abstract
We propose an extension of fuzzy qualitative trigonometric derivatives to fuzzy qualitative trigonometry (FQT). First, we revisit fuzzy qualitative trigonometry, and then derive its derivative extension. Next, we replace the trigonometry role in robot kinematics using FQT and the proposed extension, which leads to a general version of robotic kinematics. Fuzzy qualitative transformation, position and velocity of a robot are derived and discussed. Finally, we highlight the impact of the proposed methods to robotics, especially the issue of closing the gap between low-level sensing & control tasks and symbolic cognitive functions, which is one of key open problems for AI robotics [1]. Simulation results are provided to support the theoretical improvement.
Honghai Liu 0001, David J. Brown 0002
FUZZ-IEEE1
2006 Combining Multi-Frame Images for Enhancement Using Self-Delaying Dynamic Networks
abstract
This paper presents the use of a newly created net-work structure known as a Self-Delaying Dynamic Network (SDN). The SDNs were created to process data which varies with time. They feature an inbuilt timing structure which allows them to pass different items of data at different rates. This allows the network to store data from one time period until it can be used with related data in a later time period. These SDNs are non-recurrent temporal neural networks which store input data and they feature dynamic logic based connections between layers. An application is shown to create a high resolution image from a set of time stepped input frames. Several low resolution images and one high resolution image of a number of scenes were presented to the SDN during training by a Genetic Algorithm (GA). The trained SDN was then used to enhance a number of unseen noisy image sets. The images formed by the SDN are superior in several ways to the images produced using bi-cubic interpolation.
Lewis Eric Hibell, Honghai Liu 0001, David J. Brown 0002
IJCNN2
2006 Human Arm-Motion Classification Using Qualitative Normalised Templates
Chee Seng Chan, Honghai Liu 0001, David J. Brown 0002
KES (1)2
2006 Spiking Neural Network Based Classification of Task-Evoked EEG Signals
Piyush Goel, Honghai Liu 0001, David J. Brown 0002, Avijit Datta
KES (1)2
2006 Temperature Field Estimation for the Pistons of Diesel Engine 4112
Zuoqin Qian, Honghai Liu 0001, Guangde Zhang, David J. Brown 0002
KES (1)2
2005 Fuzzy Qualitative Trigonometry
abstract
This paper proposes fuzzy qualitative representation of trigonometry (FQT) in order to bridge the gap between qualitative and quantitative representation of physical systems using Trigonometry. Fuzzy qualitative coordinates are defined by replacing a unit circle with a fuzzy qualitative circle; the Cartesian translation and orientation are replaced by their fuzzy membership functions. Trigonometric functions, rules and the extensions to triangles in Euclidean space are converted into their counterparts in fuzzy qualitative coordinates using fuzzy logic and qualitative reasoning techniques. We developed a MATLAB toolbox XTrig in terms of 4-tuple fuzzy numbers to demonstrate the characteristics of the FQT. This approach addresses a representation transformation interface to connect qualitative and quantitative descriptions of trigonometry-related systems (e.g., robotic systems).
Honghai Liu 0001, George Macleod Coghill
SMC1
2005 A model-based approach to robot fault diagnosis
Honghai Liu 0001, George Macleod Coghill
Knowl. Based Syst.1
2004 Qualitative Modelling of Planar Robots
Honghai Liu 0001, George Macleod Coghill
ECAI1