Zhiyong Wang 0009

dblp:62/234-9 · DBLP profile ↗
← Back
32ranked-venue papers
3as first author
28since 2021 · last 2026
0000-0001-5546-6666ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Human-computer interaction and ubiquitous computing · 4 · 2 since 2021Systems, architecture and hardware · 3 · 3 since 2021
YearPublicationVenuePosition
2026 HKAFER: Achieve Visual Parameter-Efficient Fine-Tuning via Heterogeneous Kronecker Adaptation for Facial Expression Recognition
abstract
Facial Expression Recognition (FER) seeks to classify affective states from facial images, which remains a challenging problem due to variations in real-world conditions. FER task becomes particularly complex when handling unconstrained environments characterized by partial occlusions, different head poses, and so on. To address the above problems, current approaches rely on extensive learnable parameters and complex model architectures, which inevitably lead to overfitting and cause the FER model to focus on non-discriminative facial regions. In this work, we propose an HKAFER model that can adaptively enhance visual expression representations through efficiently fine-tuning the image encoder in large Visual Foundation Models (VFMs) and Vision-Language Models (VLMs). Specifically, we establish Heterogeneous Kronecker Adaptation (HeKA), which consists of multi-scale adapters based on Kronecker product in a parallel manner, offering significantly diverse subspaces to learn the incremental matrices. Besides, we also propose Dual-Branch Interactive Router (DBIR) to dynamically assign the weights of adapters, which promotes collaboration and information flow among them. In this way, our HKAFER can effectively capture robust spatial features and the regional associations. Experimental results demonstrate that our proposed model not only outperforms state-of-the-art methods on several FER benchmarks but also uses significantly fewer trainable parameters.
Yu Gao 0010, Haoyu Ji 0001, Zhiyong Wang 0009, Wenze Huang, Xueting Liu 0009, Weihong Ren, Honghai Liu 0001
AAAI3
2026 Unsupervised Cross-Domain 3D Human Pose Estimation via Pseudo-Label-Guided Global Transforms
abstract
Existing 3D human pose estimation methods often suffer in performance, when applied to cross-scenario inference, due to domain shifts in characteristics such as camera viewpoint, position, posture, and body size. Among these factors, camera viewpoints and locations have been shown to contribute significantly to the domain gap by influencing the global positions of human poses. To address this, we propose a novel framework that explicitly conducts global transformations between pose positions in the camera coordinate systems of source and target domains. We start with a Pseudo-Label Generation Module that is applied to the 2D poses of the target dataset to generate pseudo-3D poses. Then, a Global Transformation Module leverages a human-centered coordinate system as a novel bridging mechanism to seamlessly align the positional orientations of poses across disparate domains, ensuring consistent spatial referencing. To further enhance generalization, a Pose Augmentor is incorporated to address variations in human posture and body size. This process is iterative, allowing refined pseudo-labels to progressively improve guidance for domain adaptation. Our method is evaluated on various cross-dataset benchmarks, including Human3.6M, MPI-INF-3DHP, and 3DPW. The proposed method outperforms state-of-the-art approaches and even outperforms the target-trained model.
Zhiyong Wang 0009, Xinyu Fan 0009, Amirhossein Dadashzadeh, Honghai Liu 0001, Majid Mirmehdi
IEEE Trans. Circuits Syst. Video Technol.2
2026 Topology-Motion Decoupling Framework With Textual Regularization for Skeleton-Based Temporal Action Segmentation
abstract
Skeleton-based temporal action segmentation aims to capture key information in long skeleton motion sequences to temporally segment and identify actions at a fine-grained level. Existing approaches have achieved promising results by improving the modeling of topological spatial relationships and long-term temporal dependencies. However, current methods often overlook the distinct nature of motion and topological information, applying a monolithic modeling paradigm to both. This approach fails to fully exploit their differential contributions to precise boundary localization and effective class discrimination. To address these limitations, we propose a novel Topology-Motion Decoupling Framework (TMD). Our framework incorporates three key designs. First, an auxiliary Differential Motion Perception Branch explicitly models the temporal gradients of skeletal trajectory to decouple boundary-sensitive motion features. Second, we introduce two effective fusion modules that integrate the complementary features from both branches for mutual enhancement. Finally, a Boundary-Aware Textual Regularization scheme leverages a dual set of semantic prompts for boundary/non-boundary to differentially guide the feature learning process. By design, TMD explicitly mitigates semantic and temporal confusion between actions, thereby enhancing inter-class discriminability and boundary awareness. Extensive experiments on five challenging public datasets demonstrate that our TMD achieves state-of-the-art performance.
Haoyu Ji 0001, Bowen Chen 0004, Wenze Huang, Weihong Ren, Zhiyong Wang 0009, Honghai Liu 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 InstructHOI: Context-Aware Instruction for Multi-Modal Reasoning in Human-Object Interaction Detection
abstract
Recently, Large Foundation Models (LFMs), e.g., CLIP and GPT, have significantly advanced the Human-Object Interaction (HOI) detection, due to their superior generalization and transferability. Prior HOI detectors typically employ single- or multi-modal prompts to generate discriminative representations for HOIs from pretrained LFMs. However, such prompt-based approaches focus on transferring HOI-specific knowledge, but unexplore the potential reasoning capabilities of LFMs, which can provide informative context for ambiguous and open-world interaction recognition. In this paper, we propose InstructHOI, a novel method that leverages context-aware instructions to guide multi-modal reasoning for HOI detection. Specifically, to bridge knowledge gap and enhance reasoning abilities, we first perform HOI-domain fine-tuning on a pretrained multi-modal LFM, using a generated dataset with 140K interaction-reasoning image-text pairs. Then, we develop a Context-aware Instruction Generator (CIG) to guide interaction reasoning. Unlike traditional language-only instructions, CIG first mines visual interactive context at the human-object level, which is then fused with linguistic instructions, forming multi-modal reasoning guidance. Furthermore, an Interest Token Selector (ITS) is adopted to adaptively filter image tokens based on context-aware instructions, thereby aligning reasoning process with interaction regions. Extensive experiments on two public benchmarks demonstrate that our proposed method outperforms the state-of-the-art ones, under both supervised and zero-shot settings.
Jinguo Luo, Weihong Ren, Quanlong Zheng, Zhenlong Yuan, Zhiyong Wang 0009, Haonan Lu, Honghai Liu 0001
NeurIPS6
2025 Text Prompt Region Decomposition for Effective Facial Expression Recognition
abstract
Facial expressions are conveyed through semantically distinct visual cues distributed across different facial regions, making region-aware feature modeling essential for accurate Facial Expression Recognition (FER). However, existing methods typically rely on implicit attention mechanisms or manually defined region cues without grounding in semantically aligned supervision, which often leads to suboptimal representation of regional expression cues, increased risk of overfitting, and limited interpretability. To address these limitations, we propose a Text Prompt Region Decomposition (TPRD) network that explicitly disentangles expression-relevant features across key facial regions via text prompt guidance. Specifically, TPRD comprises a visual-language pretrained encoder (e.g., CLIP), a Region Decomposition Module (RDM), and a Regional Integration Module (RIM). The visual encoder extracts global visual features from full-face images, while the text encoder embeds region-specific prompts (e.g., “mouth”, “eye”) into semantic vectors within a shared visual-language embedding space. The RDM employs a multi-branch architecture to project global visual features onto the semantic directions of region-specific text embeddings, enabling explicit extraction of local visual features. Subsequently, the RIM models the interaction between regional and global features, adaptively generating distinct regional contributions that modulate the global representation for different facial expression samples. Experimental results demonstrate that our proposed TPRD achieves leading performance in both within- and cross-dataset evaluations, as well as in scenarios involving occlusion and large head pose variations.
Zhiyong Wang 0009, Honghai Liu 0001
IEEE Trans. Affect. Comput.4
2025 Facial Expression Monitoring via Fine-Grained Vision-Language Alignment
abstract
In the fields of health care and clinical monitoring, vision-based Facial Expression Recognition (FER) has achieved significantly progress, but it still faces the challenge of poor generalization ability under unconstrained conditions of occlusions and pose variation. Recently, Vision-Language Model (VLM) has greatly advanced the FER task. However, the existing VLM-based FER methods typically leverage a hard-crafted prompt (e.g., “a photo of [class]”) and only focus on the holistic semantic alignment, which may suffer from modal heterogeneity. In this work, we propose a fine-grained vision-language model via Prompt Masking for FER (PMFER). Specifically, for each expression, we first create fine-grained prompts using facial action units to guide the image encoder to learn discriminative representations. Further, to finely align text prompts and visual action units, we randomly drop a phrase description in the prompts and then predict the dropped phrase by conducting modal cross attention, implicitly promoting fine-grained vision-language alignment. In addition, we also design a modal-adversarial strategy to holistically eliminate the modal difference between visual and textual embeddings in a common latent space. Experimental results demonstrate that our PMFER model outperforms the state-of-the-art methods on several FER benchmarks, especially under the conditions of occlusions and pose variations. Note to Practitioners—Facial expression recognition is very important in health care and clinical monitoring, which provides an useful tool to assess the psychological and physiological conditions of patients. Although FER has made significant progress with the development of deep learning technologies, it still faces problems in the complex environments (e.g., occlusions and pose variations). To address the above issues, we propose a novel FER method in this work based on the recent vision-language model. It takes RGB image and text prompts as input and finally predicts the expression classification. Different from the existing methods, the proposed PMFER can enable fine-grained modal alignment for facial key units. Compared with the state-of-the-art methods on the public datasets, it can achieve better results, especially under the conditions of occlusions and pose variations. Also, we evaluate the proposed method on a real-world pain dataset, and the results demonstrate that PMFER has a good generalization and can be applied to health care.
Weihong Ren, Yu Gao 0010, Xi'ai Chen, Zhi Han, Zhiyong Wang 0009, Jiaole Wang, Honghai Liu 0001
IEEE Trans Autom. Sci. Eng.5
2025 Multiscale Skeleton-Based Temporal Action Segmentation Using Hierarchical Temporal Modeling and Prediction Ensemble
abstract
Skeleton-based temporal action segmentation (TAS) decomposes untrimmed skeleton sequence into meaningful segments. The variance in temporal scale challenges the skeleton modeling network to seek a balance between over-segmentation and under-segmentation. Current methods often rely on parallel multiscale feature extractors and additional refinement modules to mitigate the multiscale issue, which brings significant computations and complexity. To address these issues, this article proposes multiscale skeleton-based TAS (MSTAS), consisting of temporal probability pyramid (TPP) and smoothed multiscale ensemble (SME). TPP represents each action as a collection of multiscale probability distributions using a U-shape hierarchical temporal pyramid. Subsequently, SME takes the average of distributions instead of deploying additional refinement stages to achieve action segmentation. Considering the over-confident issue that exists in each scale, SME incorporates a novel label smoothing phase to improve the probability distributions by dynamically calibrating the confidence of each scale. Experimental results on four public datasets show that the MSTAS achieves state-of-the-art performance with less computation overheads, such as +1.1% accuracy and +2.8% [email protected] on the challenging LARa dataset with 70% fewer parameters and 80% fewer GFLOPS. Benefiting from confidence calibration, the MSTAS efficiently utilizes more temporal scales while keeping better calibration for ambiguous action instances. Additionally, the U-shape pyramid demonstrates a strong compatibility with classical refinement module, enabling the efficient extraction of multiscale motion representations.
Bowen Chen 0004, Haoyu Ji 0001, Weihong Ren, Qiyi Tong, Zhiyong Wang 0009, Honghai Liu 0001
IEEE Trans. Cybern.6
2025 Robust Compact Human Pose Learning Against Open-World Visual Perturbations
abstract
Skeleton-based action recognition has achieved remarkable progress. However, in open-world scenarios, limited human visual labels, drifting skeletal structures, and novel action categories introduce complex visual disturbances that severely limit the robustness of pose representations. Herein, we propose UnicornPose, a Universal Compact Human Pose Representation, that learns robust skeleton correlations and recognizes action across various open-world scenarios. The core advantages include: 1) Continuously modeling human skeletal structures along the action timeline to construct a rich feature volume of human poses, ensuring sufficient information for universal representation. 2) Utilizing a multiview decoupling method to compress visual information further, obtaining robust pose representations that facilitate easier generalization across different open-world scenarios. 3) Coherence training and regularization constraint methods should be employed to enhance the generalization capability of noise-containing pose representations. These contributions enable UnicornPose to effectively counter noise interference and surpass the existing top results by 3-4%.
Xuna Wang, Yuping Guo, Weiming Fan, Zhiyong Wang 0009
IEEE Trans. Ind. Informatics5
2025 Text-Derived Relational Graph-Enhanced Network for Skeleton-Based Action Segmentation
abstract
Skeleton-based Temporal Action Segmentation (STAS) aims to segment and recognize various actions from long, untrimmed sequences of human skeletal movements. Current STAS methods typically employ spatio-temporal modeling to establish dependencies among joints as well as frames, and utilize one-hot encoding with cross-entropy loss for frame-wise classification supervision. However, these methods overlook the intrinsic correlations among joints and actions within skeletal features, leading to a limited understanding of human movements. To address this, we propose a Text-Derived Relational Graph-Enhanced Network (TRG-Net) that leverages prior graphs generated by Large Language Models (LLM) to enhance both modeling and supervision. For modeling, the Dynamic Spatio-Temporal Fusion Modeling (DSFM) method incorporates Text-Derived Joint Graphs (TJG) with channel- and frame-level dynamic adaptation to effectively model spatial relations, while integrating spatio-temporal core features during temporal modeling. For supervision, the Absolute-Relative Inter-Class Supervision (ARIS) method employs contrastive learning between action features and text embeddings to regularize the absolute class distributions, and utilizes Text-Derived Action Graphs (TAG) to capture the relative inter-class relationships among action features. Additionally, we propose a Spatial-Aware Enhancement Processing (SAEP) method, which incorporates random joint occlusion and axial rotation to enhance spatial generalization. Performance evaluations on four public datasets demonstrate that TRG-Net achieves state-of-the-art results.
Haoyu Ji 0001, Bowen Chen 0004, Weihong Ren, Wenze Huang, Zhiyong Wang 0009, Honghai Liu 0001
IEEE Trans. Image Process.6
2025 Synergistic Prompting Learning for Human-Object Interaction Detection
abstract
Human-Object Interaction (HOI) detection, as a foundational task in human-centric understanding, aims to detect interactive triplets in real-world scenarios. To better distinguish diverse HOIs within an open-world context, current HOI detectors utilize pre-trained Visual-Language Models (VLMs) to extract prior knowledge through textual prompts (i.e., descriptive texts for each HOI instance). However, relying on predetermined descriptive texts, such approaches only acquire a fixed set of textual knowledge for HOI prediction, consequently resulting in inferior performance and limited generalization. To remedy this, we propose a novel VLM-based method, which jointly performs prompting learning from both visual and textual perspectives and synergizes visual-textual prompting for HOI detection. Initially, we design a hierarchical adaptation architecture to perform progressive prompting: visual prompting is facilitated through gradual token migration from VLM's image encoder, while textual prompting is initialized with progressively leveled interaction descriptions. In addition, to synergize the visual-textual prompting learning, a text-supervising and image-tuning loop is introduced, in which the text-supervising stage guides visual prompting learning through contrastive learning and the image-tuning stage refines textual prompting by modal matching. Finally, we employ an interaction-aware knowledge merging mechanism to effectively transfer visual-textual knowledge encapsulated within synergistic prompting for HOI detection. Extensive experiments on two benchmarks demonstrate that our proposed method outperforms the state-of-the-art ones, under both supervised and zero-shot settings.
Jinguo Luo, Weihong Ren, Zhiyong Wang 0009, Xi'ai Chen, Huijie Fan, Zhi Han, Honghai Liu 0001
IEEE Trans. Image Process.3
2025 Iris Geometric Transformation Guided Deep Appearance-Based Gaze Estimation
abstract
The geometric alterations in the iris's appearance are intricately linked to the gaze direction. However, current deep appearance-based gaze estimation methods mainly rely on latent feature sharing to leverage iris features for improving deep representation learning, often neglecting the explicit modeling of their geometric relationships. To address this issue, this paper revisits the physiological structure of the eyeball and introduces a set of geometric assumptions, such as "the normal vector of the iris center approximates the gaze direction". Building on these assumptions, we propose an Iris Geometric Transformation Guided Gaze estimation (IGTG-Gaze) module, which establishes an explicit geometric parameter sharing mechanism to link gaze direction and sparse iris landmark coordinates directly. Extensive experimental results demonstrate that IGTG-Gaze seamlessly integrates into various deep neural networks, flexibly extends from sparse iris landmarks to dense eye mesh, and consistently achieves leading performance in both within- and cross-dataset evaluations, all while maintaining end-to-end optimization. These advantages highlight IGTG-Gaze as a practical and effective approach for enhancing deep gaze representation from appearance.
Zhiyong Wang 0009, Weihong Ren, Honghai Liu 0001
IEEE Trans. Image Process.2
2025 Early Screening of Autism in Toddlers via Express-Needs-With-Pointing Protocol
abstract
The incidence of autism spectrum disorders (ASD), a neurodevelopmental condition associated with challenges in social communication, has witnessed a remarkable surge in recent years, with adverse effects on individuals, families, and society at large. Early screening for autism ensures timely access to interventions, yet screening lacks systematic and methodical approaches for objectively quantifying social behaviors. In response to this, we propose a protocol for early assistive screening, termed the Express-Needs-with-Pointing (ENP), which employs a multi-sensor platform to quantify the one of the social skills of toddler. A vision-based pointing behavior detection method is proposed, combining gaze estimation and pointing estimation, where the pointing estimation integrates forearm orientation and finger direction. We conduct an experiment involving twenty toddlers aged between 16 and 32 months, 4 of whom are typically developing (TD) children, 6 diagnosed with ASD, 8 diagnosed with global developmental delay (GDD), and 5 diagnosed with language disorders (LD). The results demonstrate that the automated assessment methods for pointing behavior achieved an impressive accuracy rate of 93.9%. These findings provide compelling evidence that the ENP is one of the highly effective protocols and holds significant implications for assisting in early autism screening.
Zhiyong Wang 0009, Haibo Qin, Bingrui Zhou, Huiping Li 0004, Xiu Xu, Honghai Liu 0001
IEEE J. Biomed. Health Informatics1
2025 Exploring Eye-Tracking Based Biomarkers to Assess Cognitive Abilities in Autistic Children: A Feasibility Study
abstract
Cognitive assessment can reveal a person's cognitive processing and behavioral patterns, making it an indispensable component of autism intervention and prognosis. Existing machine-assisted cognitive assessment methods primarily focus on children's performance outcomes, overlooking distinctive behavioral models, particularly characteristics of eye movement behavior, which have been demonstrated as the most direct indicators of cognitive abilities. In this study, we explore eye-tracking biomarkers for assisting cognitive assessment through a series of meticulously designed multi-level human-computer interaction protocols, encompassing three cognitive abilities: pairing and categorization, emotion recognition, and social interaction. A platform embedded with an eye-tracking module has been developed to reliably collect and analyze eye movement data, even in the presence of unrestricted large head movements in children. Experimental results indicate that there are significant group differences between autism and typically developing children in the eye-tracking features of total fixation duration, response latency, time to first fixation, mean fixation duration, and visit count in the absence of significant intergroup differences in the Wechsler Preschool and Primary Scale of Intelligence (WPPSI) and Wechsler Intelligence Scale for Children (WISC) assessment results. In addition, certain eye-tracking features in each group are correlated with WPPSI/WISC scale scores, enabling clinical cognitive assessments within each group based on these eye movement features. This study suggests that using eye-tracking features as biomarkers to assist detailed cognitive assessments holds significant potential for the intervention and prognosis of autism.
Chunchun Hu, Zhiyong Wang 0009, Bingrui Zhou, Qinyi Ye, Ruihan Lin, Xiu Xu, Honghai Liu 0001
IEEE J. Biomed. Health Informatics3
2025 Cortico-Ocular Coupling Analysis for Developmental and Behavioral Disorders: A Review
abstract
Developmental and behavioral disorders (DBD) have a significant impact on children's neurological activity and behavioral performance. Early diagnosis and treatment are known to be beneficial for improving DBD outcomes, yet existing unimodal neurophysiological assessment tools for DBD yield significant heterogeneity in results, highlighting the urgent need for exploring novel assessment tools. Cortico-ocular coupling (COC) refers to the information interaction between the cerebral cortex and eyes, and COC analysis is a technique for quantitatively measuring the correlation of neural oscillations and eye movements as biomarkers for assessment and mechanism disclosure. This review focuses on COC analysis for DBD from four perspectives: neural substrates, research paradigms, analysis methods, and applications. First, this review provides a comprehensive overview of the neural substrates and evocation paradigms related to COC analysis, aiming at helping target brain region selection, experimental result analysis, and paradigm design. The neural substrates and evocation paradigms are categorized according to functional domains, including social functioning, attention, cognition, early visual processing, and motor function. Then, this review summarizes the EEG and eye-tracking features, the analysis methods, and the validation datasets involved in COC analysis, aiming at helping implement COC analysis. Next, this review presents the applications of COC analysis in DBD, proving the validity and advance of COC analysis. In the end, the limitations, challenges, and future directions of COC analysis are discussed.
Zhiyong Wang 0009, Chunchun Hu, Peilian Chi, Xiu Xu, Honghai Liu 0001
IEEE J. Biomed. Health Informatics2
2025 Snippet-Aware Transformer With Multiple Action Elements for Skeleton-Based Action Segmentation
abstract
The skeleton-based temporal action segmentation (STAS) aims to densely segment and classify human actions within lengthy untrimmed skeletal motion sequences. Current methods primarily rely on graph convolutional networks (GCNs) for intraframe spatial modeling and temporal convolutional networks (TCNs) for interframe temporal modeling to discern motion patterns. However, these approaches often overlook the distinctive nature of essential action elements across various actions, including engaged core body parts and key subactions. This oversight limits the ability to distinguish different actions within a given sequence. To address these limitations, the snippet-aware Transformer with multiple action element (ME-ST) is proposed to enhance the discrimination and segmentation among actions, which leverages intrasnippet attention along joints and sequences to identify core joints and key subactions at different scales. Specifically, in terms of the spatial domain, the intrasnippet cross-joint attention (CJA) module divides the sequence into distinct snippets and computes attention to establish intricate joint semantic relationships, emphasizing the identification of core motion joints. In terms of the temporal domain, in the encoder, the intrasnippet cross-frame attention (CFA) module segments the sequence in a blockwise expansion manner and establishes interframe relationships to highlight the most discriminative frames. In the decoder, clip-level representations at various temporal scales are initially generated through an hourglass-like sampling process, followed by the intrasnippet cross-scale attention (CSA) module to integrate the key clip information across different time scales. The performance evaluation on five public datasets demonstrates that ME-ST achieves state-of-the-art (SOTA) performance.
Haoyu Ji 0001, Bowen Chen 0004, Wenze Huang, Weihong Ren, Zhiyong Wang 0009, Honghai Liu 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 Exploring Self- and Cross-Triplet Correlations for Human-Object Interaction Detection
abstract
Human-Object Interaction (HOI) detection plays a vital role in scene understanding, which aims to predict the HOI triplet in the form of . Existing methods mainly extract multi-modal features (e.g., appearance, object semantics, human pose) and then fuse them together to directly predict HOI triplets. However, most of these methods focus on seeking for self-triplet aggregation, but ignore the potential cross-triplet dependencies, resulting in ambiguity of action prediction. In this work, we propose to explore Self- and Cross-Triplet Correlations (SCTC) for HOI detection. Specifically, we regard each triplet proposal as a graph where Human, Object represent nodes and Action indicates edge, to aggregate self-triplet correlation. Also, we try to explore cross-triplet dependencies by jointly considering instance-level, semantic-level, and layout-level relations. Besides, we leverage the CLIP model to assist our SCTC obtain interaction-aware feature by knowledge distillation, which provides useful action clues for HOI detection. Extensive experiments on HICO-DET and V-COCO datasets verify the effectiveness of our proposed SCTC.
Weibo Jiang, Weihong Ren, Jiandong Tian, Liangqiong Qu, Zhiyong Wang 0009, Honghai Liu 0001
AAAI5
2024 Language-Assisted Skeleton Action Understanding for Skeleton-Based Temporal Action Segmentation
Haoyu Ji 0001, Bowen Chen 0004, Xinglong Xu, Weihong Ren, Zhiyong Wang 0009, Honghai Liu 0001
ECCV (54)5
2024 MLPER: Multi-Level Prompts for Adaptively Enhancing Vision-Language Emotion Recognition
abstract
In the field of robotics, vision-based Emotion Recognition (ER) has achieved significant progress, but it still faces the challenge of poor generalization ability under unconstrained conditions (e.g., occlusions and pose variations). In this work, we propose MLPER model, which introduces Vision-Language Model for Emotion Recognition to learn discriminative representations adaptively. Specifically, different from typically leveraging a hand-crafted prompt (e.g., "a photo of a [class] person"), we first establish Multi-Level Prompts from three aspects: facial expression, human posture and situational condition using large language models, like ChatGPT. Correspondingly, we extract the visual tokens from three levels: the face, body, and context. Further, to achieve fine-grained alignment at each level, we adopt textual tokens from the positive and the hard negative to query visual tokens, predicting whether a pair of image and text is matched. Experimental results demonstrate that our MLPER model outperforms the state-of-the-art methods on several ER benchmarks, especially under the conditions of occlusions and pose variations.
Yu Gao 0010, Weihong Ren, Xinglong Xu, Zhiyong Wang 0009, Honghai Liu 0001
IROS5
2024 Automatic Recognition of Social Engagement for Children with Autism Spectrum Disorder
abstract
Estimating children's engagement levels improves their understanding of their social behaviors, since they can reflect their devotion to social interaction with others. This paper proposes an automatic method to recognize children's engagement levels in a triadic social interaction context. First, an overall metric function containing behavior, cognition, and affective dimensions is proposed to estimate children's multidimensional engagement levels. Then, the automatic feature extraction method based on gaze estimation, facial expression recognition, pose estimation, and object recognition models is illustrated to extract features to compute the engagement levels. Videos of 24 children, including 13 children with autism spectrum disorder (ASD), in triadic social interaction were collected for the engagement recognition experiment and cross-group analysis. The experimental results validate the effectiveness of the proposed automatic feature extraction method compared to human observations. Cross-group analyses revealed significant differences in affective engagement between children with ASD and typical developmental (TD) children.
Zhiyong Wang 0009, Xiu Xu, Honghai Liu 0001
SMC3
2024 Multimodal Emotion Recognition for Children with Autism Spectrum Disorder in Social Interaction
abstract
Autism Spectrum Disorders (ASD) remain a healthcare challenge and gain considerable attention due to the increasing prevalence rates and insupportable burden on families and society. It is noted that the recognition of children’s emotional states plays an important role in the evaluation and intervention process of ASD. In this paper, we aim to address the problem of automatic recognition of the emotional states of ASD children in social interactive scenarios. Since the child can be unconstrained in realistic scenarios, the face occlusion under pose variations and uncertain backgrounds become challenges of this task. To tackle this problem, we employ both facial expressions as well as body poses as cues to recognize the emotional states while most traditional methods only leverage the former. Firstly for the facial information, spatial features are extracted through convolutional neural networks followed by a temporal transformer to extract temporal information. Then for the body pose information, graph convolutional networks combined with the self-attention part are used to represent spatial features and temporal convolutional layers for temporal counterparts. Finally, different multimodal fusion ways are explored to generate final recognition results. We evaluate this method on a challenging database collected by us in real-world child-clinician interactive scenarios and the proposed method achieved significantly better results than baselines using only facial information. Thus it is suggested that there is a potential to assist in clinical practice by providing the recognized emotion as feedback.
Zhiyong Wang 0009, Bingrui Zhou, Jingxin Deng, Xiu Xu, Honghai Liu 0001
Int. J. Hum. Comput. Interact.2
2024 Dual Regression-Enhanced Gaze Target Detection in the Wild
abstract
Gaze is a vital feature in analyzing natural human behavior and social interaction. Existing gaze target detection studies learn gaze from gaze orientations and scene cues via a neural network to model gaze in unconstrained scenes. Though achieve decent accuracy, these studies either employ complex model architectures or leverage additional depth information, which limits the model application. This article proposes a simple and effective gaze target detection model that employs dual regression to improve detection accuracy while maintaining low model complexity. Specifically, in the training phase, the model parameters are optimized under the supervision of coordinate labels and corresponding Gaussian-smoothed heatmap labels. In the inference phase, the model outputs the gaze target in the form of coordinates as prediction rather than heatmaps. Extensive experimental results on within-dataset and cross-dataset evaluations on public datasets and clinical data of autism screening demonstrate that our model has high accuracy and inference speed with solid generalization capabilities.
Zhiyong Wang 0009, Weihong Ren, Xiu Xu, Honghai Liu 0001
IEEE Trans. Cybern.3
2024 FM-3DFR: Facial Manipulation-Based 3-D Face Reconstruction
abstract
3-D Morphable model (3DMM) has widely benefited 3-D face-involved challenges given its parametric facial geometry and appearance representation. However, previous 3-D face reconstruction methods suffer from limited power in facial expression representation due to the unbalanced training data distribution and insufficient ground-truth 3-D shapes. In this article, we propose a novel framework to learn personalized shapes so that the reconstructed model well fits the corresponding face images. Specifically, we augment the dataset following several principles to balance the facial shape and expression distribution. A mesh editing method is presented as the expression synthesizer to generate more face images with various expressions. Besides, we improve the pose estimation accuracy by transferring the projection parameter into the Euler angles. Finally, a weighted sampling method is proposed to improve the robustness of the training process, where we define the offset between the base face model and the ground-truth face model as the sampling probability of each vertex. The experiments on several challenging benchmarks have demonstrated that our method achieves state-of-the-art performance.
Shuwen Zhao, Dinghuang Zhang, Zhiyong Wang 0009, Honghai Liu 0001
IEEE Trans. Cybern.5
2024 Trajectory Planning for Jumping and Soft Landing With a New Wheeled Bipedal Robot
abstract
The balance control of the wheeled bipedal robots has been relatively mature, but the jump control of the wheeled bipedal robots is not perfect at present. Aiming at problems such as significant landing impact and less than expected jump height of the wheeled bipedal robots, a method of jump trajectory planning with specific soft landing ability is proposed. The dynamic model of the robot jumping and the basis for distinguishing the land and fly phases are established. The planning of the wheel and foot trajectories and the main body trajectories during the robot jumping is completed, and the tracking of the robot jumping trajectories is completed using virtual model control and virtual force compensation. Through simulations and prototype experiments, the robot can perform various jumping tasks with high precision jump heights while still being able to land without much impact. The stability and effectiveness of the proposed method are demonstrated.
Zongxing Lu, Ligang Yao, Zhiyong Wang 0009
IEEE Trans. Ind. Informatics5
2024 Computational Interpersonal Communication Model for Screening Autistic Toddlers: A Case Study of Response-to-Name
abstract
Interpersonal communication facilitates symptom measures of autistic sociability to enhance clinical decision-making in identifying children with autism spectrum disorder (ASD). Traditional methods are carried out by clinical practitioners with assessment scales, which are subjective to quantify. Recent studies employ engineering technologies to analyze children's behaviors with quantitative indicators, but these methods only generate specific rule-driven indicators that are not adaptable to diverse interaction scenarios. To tackle this issue, we propose a Computational Interpersonal Communication Model (CICM) based on psychological theory to represent dyadic interpersonal communication as a stochastic process, providing a scenario-independent theoretical framework for evaluating autistic sociability. We apply CICM to the response-to-name (RTN) with 48 subjects, including 30 toddlers with ASD and 18 typically developing (TD), and design a joint state transition matrix as quantitative indicators. Paired with machine learning, our proposed CICM-driven indicators achieve consistencies of 98.44% and 83.33% with RTN expert ratings and ASD diagnosis, respectively. Beyond outstanding screening results, we also reveal the interpretability between CICM-driven indicators and expert ratings based on statistical analysis.
Bingrui Zhou, Zhiyong Wang 0009, Bowen Chen 0004, Chunchun Hu, Huiping Li 0004, Xiu Xu, Honghai Liu 0001
IEEE J. Biomed. Health Informatics3
2023 Fatigue Detection Based on Multiple Visual Features in Virtual Driving System
abstract
Fatigue driving poses a significant hazard, leading to numerous traffic accidents annually. However, the fatigue detection algorithm based on visual features still has the problem of low accuracy, and it is difficult to test extreme fatigue driving conditions. This study proposes a fatigue detection method that exploits multiple visual features, including facial landmarks, mouth aspect ratio (MAR), eye aspect ratio (EAR), PERCLOS and head pose. Furthermore, we introduce a novel approach by developing a virtual driving system dedicated to fatigue detection. This system offers a rich driving environment and enables the exploration of fatigue-related boundary conditions. To validate our method and assess the potential of the virtual driving system in detecting fatigue driving, we conduct an experiment involving five healthy adults using the aforementioned system. In our final results, the detection accuracy of mild fatigue reached 0.94, and the detection accuracy of severe fatigue reached 0.97. The results confirm both the feasibility of our approach and the promising prospects of the virtual driving system in fatigue detection.
Qiyi Tong, Zhiyong Wang 0009, Ruihan Lin, Honghai Liu 0001
IECON6
2023 Improving Stability of Gaze Target Detection in Videos
abstract
Obtaining accurate and stable results in gaze target detection is vital for the subsequent analysis of gaze meaning. However, existing image-based methods, which focus solely on enhancing accuracy, demonstrate poor stability when directly applied on videos. Especially when the video frame rate is low, even though the actual gaze target positions do not differ significantly between adjacent frames, the detected positions vary considerably. This inconsistency, stemming from the lack of temporal information, makes dynamic detection challenging and can lead to jarring outcomes. To reduce the jitter in gaze target detection in videos, we introduce an approach that integrates spatial and temporal modules to combine spatial with temporal information. Additionally, we propose a Jitter loss function to capture significant jitter and impose a strong penalty during training, which empowers our model with increased stability for dynamic detection. Based on a self-collected dataset, experiments demonstrate that our approach exhibits superior stability without compromising accuracy.
Zhiyong Wang 0009, Xiu Xu, Honghai Liu 0001
IECON3
2022 Early Screening of Autism in Toddlers via Response-To-Instructions Protocol
abstract
Early screening of autism spectrum disorder (ASD) is crucial since early intervention evidently confirms significant improvement of functional social behavior in toddlers. This article attempts to bootstrap the response-to-instructions (RTIs) protocol with vision-based solutions in order to assist professional clinicians with an automatic autism diagnosis. The correlation between detected objects and toddler's emotional features, such as gaze, is constructed to analyze their autistic symptoms. Twenty toddlers between 16-32 months of age, 15 of whom diagnosed with ASD, participated in this study. The RTI method is validated against human codings, and group differences between ASD and typically developing (TD) toddlers are analyzed. The results suggest that the agreement between clinical diagnosis and the RTI method achieves 95% for all 20 subjects, which indicates vision-based solutions are highly feasible for automatic autistic diagnosis.
Zhiyong Wang 0009, Bin Ji 0004, Jingxin Deng, Xiu Xu, Honghai Liu 0001
IEEE Trans. Cybern.2
2021 Screening Early Children With Autism Spectrum Disorder via Response-to-Name Protocol
abstract
Incidence of children with autism spectrum disorder (ASD) has increased with an average rate of 1% worldwide. Clinical ASD screening, especially for children screening is a laborious and skilled task; however, there is no objective and effective method automating ASD children screening. Analyzing children ASD characteristics in predefined motion behavior protocols is attempted to provide automatic solutions to children ASD screening. A novel protocol, response to name (RTN), is proposed in this article for ASD clinical validation and diagnosis. The RTN method is jointly designed with clinical partners, and novel gaze estimation is developed for validating ASD characteristic behavior. Seventeen subjects including ten adults and seven children (five ASD subjects and two healthy subjects) have participated the experiment. The experiment results show that the proposed RTN system achieves an average classification score of 92.7% fully demonstrating that the principle of motion protocol based ASD screening has the potential to have early ASD screening automated.
Zhiyong Wang 0009, Keshi He, Xiu Xu, Honghai Liu 0001
IEEE Trans. Ind. Informatics1
2020 Free-Head Pose Estimation under Low-Resolution Scenarios
abstract
Head pose offers vital cues to infer one's social attention in wide applications. Most existing head pose estimation algorithms have demonstrated competitive results taking high resolution images of frontal view as input. However, these approaches still work poorly if they are fed with low-resolution images. In a more common realistic scene such as computer vision assisted autism screening, images of unconstrained patients that are taken from distant cameras often have low-resolution and non-frontal faces. To this end, we present a multi-view scheme based CNN for free-head pose classification. Residual Networks are taken as the backbone model to generate effective feature representations of the low-resolution images. A novel multi-view feature fusion layer is proposed to address facial appearance variation over multi perspectives owing to free movement. Also, a multi loss function by combining binned classification and regression losses of different pose angles is employed to obtain a more precise pose estimation. The proposed method is evaluated on two scenarios: (1) a publicly available dataset and (2) a practical application to explore the social attention of a group with social deficits: autistic children. Experimental results suggest that our method significantly outperforms state-of-the-art multi-view pose classification methods and achieves comparable pose estimation results. Moreover, the proposed method can be extended to applications of quantitative analysis of social deficits.
Zhiyong Wang 0009, Haibo Qin, Bin Ji 0004, Honghai Liu 0001
SMC2
2020 An Auxiliary Screening System for Autism Spectrum Disorder Based on Emotion and Attention Analysis
abstract
The screening and diagnosis of Autism Spectrum Disorder(ASD) suffer from great challenges due to insufficient professional clinicians and complex procedures. It is urgent to introduce an effective auxiliary system in the diagnosis and treatment process to assist in the completion of pathological information collection tasks, consequently simplifying the screening method and improving the accuracy of screening. We propose a computer vision-based early screening system for ASD to characterize the facial expressions and eye gaze attention which are considered to remarkable indicators for early screening of autism. The system provides the subjects with three different virtual interaction modes: video, picture, and virtual interactive game. During the interaction between the subject and the computer, the system extracts and analyzes the quantitative information of the subject's performance. Then, through computer vision-based emotion analysis and attention analysis methods, the subject's emotions and attention features in the three interaction modes are automatically calculated to assist in the early screening of autism. Finally, the accuracy and feasibility of the system are verified through experiments on both the publicly available dataset and the data collected from 10 ASD children.
Bin Ji 0004, Zhiyong Wang 0009, Honghai Liu 0001
SMC3
2020 HDS-SP: A novel descriptor for skeleton-based human action recognition
Zhiyong Wang 0009, Honghai Liu 0001
Neurocomputing2
2018 Robust Eye Center Localization Based on an Improved SVR Method
Zhiyong Wang 0009, Haibin Cai, Honghai Liu 0001
ICONIP (7)1