Xiangmin Xu 0001

dblp:28/9939-1 · DBLP profile ↗
← Back
155ranked-venue papers
0as first author
106since 2021 · last 2026
0000-0003-4573-5820ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 96 · 59 since 2021Artificial intelligence and machine learning · 67 · 55 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 13 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 DAVID: Dual-stage Adaptive Vision-text Integrated Decoupling for Multimodal KV Cache Eviction
abstract
With the rapid development of multimodal large language models (MLLMs), deploying them on low-resource devices remains challenging. Beyond the model size, long multimodal inputs cause substantial memory overhead in the KV cache, making efficient cache management critical. In this paper, we propose DAVID, a KV cache eviction strategy that adapts to the degree of modality fusion across layers. By analyzing the feature distributions of vision and text tokens, we observe low fusion in early layers and high fusion in deeper layers. Based on this observation, DAVID adopts a decoupled eviction strategy in shallow layers and a super-modal eviction strategy in deeper layers. To support this dynamic switching, we design a lightweight metric that quantifies cross-modal fusion and uses a threshold to determine which layers require decoupling. Experimental results show that DAVID achieves state-of-the-art performance on multiple benchmarks and offers a new perspective on KV cache eviction for MLLMs.
Yifeng Gu, Jianxiu Jin, Kailing Guo, Xiangmin Xu 0001
AAAI4
2026 One-shot federated learning via prototype-based client grouping for model ensemble
Xingcheng Jiang, Xiang Tian 0003, Xiangmin Xu 0001
Expert Syst. Appl.4
2026 A dual uncertainty-aware fusion framework for face expression recognition in the wild
abstract
Facial Expression Recognition(FER) is a key task in the broader landscape of affective computing and human-computer interaction, enabling machines to interpret human emotions. To better learn discriminative features under complex facial variations, recent FER research has increasingly adopted multi-branch fusion architectures that aim to capture complementary features from diverse perspectives. However, existing multi-branch fusion strategies, including static weighting, simple concatenation, or uncertainty-aware modeling, lack the capacity to comprehensively capture and reconcile the reliability variations across both individual instances and structural branches. To overcome these limitations, we propose a novel multi-branch fusion strategy, named Dual Uncertainty-Aware Fusion Framework(DUAFF), which improves the discriminability of integrated features by simultaneously modeling instance-wise uncertainty and inter-branch correlations. Specifically, the proposed method comprises two complementary modules: Instance-Discrepant Uncertainty-Aware Fusion Module (ID-UAFM) and Branch-Discrepant Uncertainty-Aware Fusion Module (BD-UAFM). ID-UAFM is introduced to perform channel-wise entropy analysis between semantically distinct samples to estimate instance-level uncertainty, enabling selective channel-wise fusion that emphasizes reliable representations while suppressing uncertain responses. BD-UAFM is further proposed to capture structural uncertainty by evaluating the relative reliability of features across multiple branches and adaptively weighting their contributions based on inter-branch discrepancies. Experimental results demonstrate that the proposed DUAFF consistently outperforms POSTER across three benchmark datasets, achieving accuracy improvements of 0.23 % on RAF-DB, 0.69 % on FER2013, and 0.29 % on AffectNet (7-class), thereby confirming its effectiveness in enhancing the reliability and discriminability of facial representations.
Wenfeng Jiang, Lin Wang 0004, Fang Liu 0030, Chunmei Qing, Xiaofen Xing, Xiangmin Xu 0001, Weiquan Fan, Zhanpeng Jin
Expert Syst. Appl.7
2026 EMS-Grasping: Providing Finger with Force Feedback for Grasping Virtual Objects Using Electrical Muscle Stimulation
abstract
There are some limitations in the application of electrical muscle stimulation (EMS) in virtual reality (VR) that provides precise force feedback to the fingers. To address this issue, we propose EMS-Grasping, an interaction technique that combines biomechanical simulation model with EMS. This work makes two key contributions: (1) an EMS-based musculoskeletal model of the hand is constructed to establish the relationship between the electrical stimulation intensity and the angle of the finger joint; (2) a simulation model-based method for EMS control of finger fixation at a target angle is proposed and applied to grasping virtual objects in VR. In the first experiment, we verified the reliability of the simulation model. In the second experiment, we compared three interaction models to evaluate the performance of EMS-Grasping. The results show that EMS-Grasping not only effectively reduces finger penetration with virtual objects, but also enhances the user experience in terms of realism and comfort. This implies that EMS-Grasping has good potential in expressing grasping virtual objects.
Zikang Dong, Jialong Liu, Hongbo Yao, Fenghan Zhou, Haoqiang Hua, Xiangmin Xu 0001
Int. J. Hum. Comput. Interact.7
2026 Random graph construction-based hierarchical attention multi-task graph convolution network for sEEG SOZ identification
Huachao Yan, Kailing Guo, Shiwei Song, Xiaofen Xing, Xiangmin Xu 0001
Neurocomputing5
2026 MMRepAgent: Explainable stock earnings forecasting via multimodal report agent framework
Xiangyu Li 0010, Yawen Zeng, Xiaofen Xing, Jin Xu 0014, Xiangmin Xu 0001
Knowl. Based Syst.5
2026 Knowledge-distillation based personalized federated learning with distribution constraints
Chang Mu, Kailing Guo, Xiang Tian 0003, Xiangmin Xu 0001
Neural Networks5
2026 Soft local reactivation for communication efficient federated learning
Chang Mu, Kailing Guo, Xiang Tian 0003, Xiangmin Xu 0001
Pattern Recognit.5
2026 SiaTalker: Siamese Emotion Injection for Coarse-to-Fine Speech-Driven 3D Facial Animation
abstract
During communication, people inadvertently express emotions through facial animations, with emotion constantly fluctuating throughout the conversation. Existing speech-driven facial animation works predominantly rely on sentence-level coarse-grained emotion labels to impart emotional expression information, thereby overlooking temporal emotion variations. How to make 3D talking heads articulate fluently while adeptly expressing mood fluctuations is the problem addressed by our work. We propose a framework, named SiaTalker, which exquisitely refines emotions of facial animations from coarse-grained levels to dual-perspective interactive fine-grained levels, resulting in highly coordinated and fluid emotional expressions. Specifically, we propose the coarse-to-fine dual granularity fusion decoding framework to connect global and local emotional aspects, allowing nuanced shifts within a consistent emotional tone. The pseudo-siamese cross-perspective submodule is embedded to interact with siamese fine-grained emotions across perspectives, further capturing subtle emotional fluctuations and amplifying the authenticity of facial movements. Furthermore, to comprehensively evaluate our work, the CREMA-D dataset is reconstructed into a 3D emotional talking face dataset with the largest number of subjects, serving as a valuable experimental supplement. Extensive experiments and user studies demonstrate that our approach outperforms state-of-the-art methods and exhibits superior emotional agility and expressiveness in facial movements.
Zhaojie Chu, Yilin Lan, Jianxiu Jin, Xiangmin Xu 0001, Xiaofen Xing
IEEE Trans. Affect. Comput.4
2026 EETalk: Expression Enhancement in Speech-Driven 3D Facial Animation
abstract
Speech-driven 3D facial animation aims to generate natural and expressive facial movements from speech. Although significant progress has been made, existing methods still face challenges in generating realistic upper facial expressions. Specifically, existing methods that jointly optimize the holistic face tend to overlook fine-grained spatial movements due to motion differences across facial regions. Pre-trained speech feature extractors, which emphasize long-term dependencies, provide limited fine-grained temporal cues. In this work, we propose a novel framework, EETalk, to enhance the realism of facial expressions, which captures fine-grained spatial information and f ine-grained temporal dynamics from speech. To alleviate loss of fine-grained spatial information, we propose a novel disassemble and-reassemble modeling strategy. This strategy constructs two independent motion representation spaces for the upper and lower faces, allowing for the capture of weak upper-face movements while preserving motion diversity. Then, we propose a Cross-Region Coordination Module to ensure the synchronization and coordination of the movements from the independent upper and lower faces. To effectively capture facial micro-expressions, we incorporate fine-grained time-varying features to compensate for the short-timescale details underrepresented in the long term semantic features extracted by pre-trained self-supervised models, thereby improving the generation of fast, subtle facial expressions. Experimental results demonstrate that our approach significantly improves motion accuracy, expression consistency, and perceptual quality compared to existing methods.
Zhaojie Chu, Kailing Guo, Xiaofen Xing, Bolun Cai, Lin Wang 0004, Xiangmin Xu 0001
IEEE Trans. Multim.6
2026 PEGCL: Pseudo-Entropy Guided Complementary Learning for Robust Facial Expression Recognition Under Label Noise
abstract
Facial Expression Recognition (FER) has recently plays a crucial role in advancing human-computer interaction systems, aiming to understand users' inner states and underlying intentions. However, FER in real-world scenarios remains challenging due to significant label noise, caused by ambiguous facial expressions in low-quality images and annotation bias. To tackle this issue, this paper proposes a novel framework, Pseudo-Entropy Guided Complementary Learning (PEGCL), designed to robustly handle noisy labels by leveraging complementary information, which trains networks using all complementary labels defined as “facial expression images that do not belong to complementary emotion labels.” This approach effectively utilizes non-target emotion labels to mitigate the impact of label noise, rather than relying solely on annotated emotion labels. Specifically, the proposed PEGCL framework consists of three components: logit normalization to stabilize predicted probabilities and prevent gradient explosions, transformed complementary learning to redistribute the optimization focus across complementary categories by leveraging pseudo-entropy guided, and random complementary label dropping to dynamically exclude subsets of complementary labels, enhancing generalization and preventing overfitting. These components collectively ensure robust and efficient optimization under noisy label conditions. Importantly, the proposed PEGCL does not require explicit noise estimation or complex label correction mechanisms, making it a simple and effective solution for real-world FER tasks. Extensive experiments on benchmark FER datasets demonstrate that PEGCL consistently outperforms existing methods, achieving the state-of-the-art robustness against label noise while maintaining high classification accuracy.
Lin Wang 0004, Dan Liao, Fang Liu 0030, Xiangmin Xu 0001, Kailing Guo, Zhanpeng Jin
IEEE Trans. Multim.4
2025 PsyDT: Using LLMs to Construct the Digital Twin of Psychological Counselor with Personalized Counseling Style for Psychological Counseling
abstract
Currently, large language models (LLMs) have made significant progress in the field of psychological counseling.However, existing mental health LLMs overlook a critical issue where they do not consider the fact that different psychological counselors exhibit different personal styles, including linguistic styles and therapeutic types, etc.As a result, these LLMs fail to satisfy the individual needs of clients who seek different counseling styles.To help bridge this gap, we propose PsyDT, a novel framework using LLMs to construct the Digital Twin of Psychological counselor with personalized counseling style.Compared to the timeconsuming and costly approach of collecting a large number of real-world counseling cases to create a specific counselor's digital twin, our framework offers a faster and more costeffective solution.To construct PsyDT, we utilize dynamic one-shot learning by using GPT-4 to capture counselor's unique counseling style, mainly focusing on linguistic style and therapy technique.Subsequently, using existing singleturn long-text dialogues with client personality, GPT-4 is guided to synthesize multi-turn dialogues of specific counselor.Finally, we finetune the LLMs on the synthesized dataset, Psy-DTCorpus, to achieve the digital twin of psychological counselor with personalized counseling style.Experimental results indicate that our proposed PsyDT framework can synthesize multi-turn dialogues that closely resemble realworld counseling cases and demonstrate better performance compared to other baselines, thereby show that our framework can effectively construct the digital twin of psychological counselor with a specific counseling style. 1
Haojie Xie, Yirong Chen, Xiaofen Xing, Jingkai Lin, Xiangmin Xu 0001
ACL (1)5
2025 Intermediate-Selective Feature Enhancement for Speech Emotion Recognition
abstract
Speech Emotion Recognition (SER) plays a vital role in enabling intelligent systems to perceive and respond to human emotions. Recent advances in large pretrained models have led to substantial improvements in SER performance. However, most existing approaches focus solely on a single layer representation or perform layer-wise fusion, which may not be optimal for downstream emotion recognition tasks. In this paper, we propose ISFE (Intermediate-Selective Feature Enhancement), a novel framework designed to enhance the emotional expressiveness of pretrained representations. Rather than fusing features across different layers, ISFE targets and enhances specific emotion-related signals from intermediate layer that are often suppressed in deeper layers. Extensive experiments on the IEMOCAP and MELD datasets demonstrate that ISFE achieves better performance than previous state-of-the-art methods, validating the effectiveness of enhancing intermediate-layer features for SER.
Yangbiao Li, Xiaofen Xing, Jialong Mai, Jingyuan Xing, Xiangmin Xu 0001
ASRU5
2025 Enhancing fNIRS Signal Classification with Test-Time Training by Improved Spatiotemporal Feature Extraction
abstract
Functional Near-Infrared Spectroscopy (fNIRS) is a convenient brain imaging technology that is adaptable to various complex environments. It can detect the brain's responses to external stimuli across different contexts. However, existing fNIRS processing methods fail to mine the brain's sequential association responses in long-term signals over time. To better capture the hidden spatiotemporal correlations and long-range dependencies inherent in fNIRS signals, we propose fNIRS-TTT, a novel classification architecture leveraging Test-Time Training (TTT). Our framework introduces two core innovations: the Cross-Attention Embedding module (CAE) and the TTT Block. The CAE module combines the wide-kernel Patch Conv (capturing cross-channel spatial patterns) and the Channel Conv (processing temporal features within individual channels) together. Their outputs are fused via Cross-Attention to generate spatiotemporal tokens enriched with multi-scale spatial information for fNIRS. The TTT components dynamically adapt to the sequential nature of fNIRS data, effectively capturing long-range temporal context and mitigating overfitting compared to former models. KFold Cross-Validation (KFold-CV) and Leave-One-Subject Cross-Validation (LOSO-CV) are conducted on three open datasets, which demonstrate the superior performance of the proposed fNIRS-TTT compared to state-of-the-art models.
Wanxiang Luo, Chunmei Qing, Junpeng Tan, Yihang Zou, Xiangmin Xu 0001
BIBM5
2025 Modeling Thousands of Human Annotators for Generalizable Text-to-Image Person Re-identification
abstract
Text-to-image person re-identification (ReID) aims to retrieve the images of an interested person based on textual descriptions. One main challenge for this task is the high cost in manually annotating large-scale databases, which affects the generalization ability of ReID models. Recent works handle this problem by leveraging Multi-modal Large Language Models (MLLMs) to describe pedestrian images automatically. However, the captions produced by MLLMs lack diversity in description styles. To address this issue, we propose a Human Annotator Modeling (HAM) approach to enable MLLMs to mimic the description styles of thousands of human annotators. Specifically, we first extract style features from human textual descriptions and perform clustering on them. This allows us to group textual descriptions with similar styles into the same cluster. Then, we employ a prompt to represent each of these clusters and apply prompt learning to mimic the description styles of different human annotators. Furthermore, we define a style feature space and perform uniform sampling in this space to obtain more diverse clustering prototypes, which further enriches the diversity of the MLLM-generated captions. Finally, we adopt HAM to automatically annotate a massive-scale database for text-to-image ReID. Extensive experiments on this database demonstrate that it significantly improves the generalization ability of ReID models. Code is available at https://github.com/sssaury/HAM.
Jiayu Jiang, Changxing Ding, Wentao Tan, Xiangmin Xu 0001
CVPR6
2025 SAKI-RAG: Mitigating Context Fragmentation in Long-Document RAG via Sentence-level Attention Knowledge Integration
abstract
Traditional Retrieval-Augmented Generation (RAG) frameworks often segment documents into larger chunks to preserve contextual coherence, inadvertently introducing redundant noise.Recent advanced RAG frameworks have shifted toward finer-grained chunking to improve precision.However, in long-document scenarios, such chunking methods lead to fragmented contexts, isolated chunk semantics, and broken inter-chunk relationships, making cross-paragraph retrieval particularly challenging.To address this challenge, maintaining granular chunks while recovering their intrinsic semantic connections, we propose SAKI-RAG (Sentence-level Attention Knowledge Integration Retrieval-Augmented Generation).Our framework introduces two core components: (1) the SentenceAttnLinker, which constructs a semantically enriched knowledge repository by modeling inter-sentence attention relationships, and (2) the Dual-Axis Retriever, which is designed to expand and filter the candidate chunks from the dual dimensions of semantic similarity and contextual relevance.Experimental results across four datasets-Dragonball, SQUAD, NFCORPUS, and SCI-DOCS demonstrate that SAKI-RAG achieves better recall and precision compared to other RAG frameworks in long-document retrieval scenarios, while also exhibiting higher information efficiency.
Wenyu Tao, Xiaofen Xing, Zeliang Li, Xiangmin Xu 0001
EMNLP4
2025 DecoupledSynth: Enhancing Zero-Shot Text-to-Speech Via Factors Decoupling
abstract
Studies of speech representation enhance zero-shot Text-to-Speech by mapping text to intermediate representations before generating speech. However, using representations often struggles to balance linguistic, para-linguistic, and non-linguistic information in speech during the synthesis phase. Additionally, it usually takes substantial resources for representation extraction training. To address these limitations, we propose DecoupledSynth. It combines different self-supervised models to extract comprehensive, decoupled representations. This structure enables more thorough and nuanced synthesis by leveraging reference speech with decoupled processing stages. Experiments on the VCTK and LibriTTS datasets support the potential of this new framework, showing that it can produce more consistent and realistic speech. Speech demos are available at https://test1634.github.io/DecoupledSynth/.
Jingyuan Xing, Shuaiqi Chen, Xiangmin Xu 0001, Xiaofen Xing
ICASSP3
2025 Bilateral Collaboration with Large Vision-Language Models for Open Vocabulary Human-Object Interaction Detection
abstract
Open vocabulary Human-Object Interaction (HOI) detection is a challenging task that detects all triplets of interest in an image, even those that are not pre-defined in the training set. Existing approaches typically rely on output features generated by large Vision-Language Models (VLMs) to enhance the generalization ability of interaction representations. However, the visual features produced by VLMs are holistic and coarse-grained, which contradicts the nature of detection tasks. To address this issue, we propose a novel Bilateral Collaboration framework for open vocabulary HOI detection (BC-HOI). This framework includes an Attention Bias Guidance (ABG) component, which guides the VLM to produce fine-grained instance-level interaction features according to the attention bias provided by the HOI detector. It also includes a Large Language Model (LLM)-based Supervision Guidance (LSG) component, which provides fine-grained token-level supervision for the HOI detector by the LLM component of the VLM. LSG enhances the ability of ABG to generate high-quality attention bias. We conduct extensive experiments on two popular benchmarks: HICO-DET and V-COCO, consistently achieving superior performance in the open vocabulary and closed settings. The code will be released in Github.
Yupeng Hu 0005, Changxing Ding, Shaoli Huang, Xiangmin Xu 0001
ICCV5
2025 Drawing Developmental Trajectory From Cortical Surface Reconstruction
Ruowen Qu, Zhongliang Liu, Zhuoyan Dai, Dongzi Shi, Sijin Yu, Tong Xiong, Shiping Liu, Xiangmin Xu 0001, Xiaofen Xing, Xin Zhang 0013
ICCV9
2025 DP-Net: A 3D Dilated Projection Framework For Precise Fetal Brain Tissue Segmentation
abstract
Precise segmentation of fetal brain tissues in MRI is essential for studying brain development and for the early diagnosis and treatment of neurological disorders. However, the complex and variable anatomy of the fetal brain, significant morphological changes at different gestational ages, and the low-quality MRI and inherent noise of fetal acquisition pose significant challenges. To address these, we propose a novel 3D Dilated Projection U-net segmentation framework, DP-Net, which incorporates large kernel convolutions, atrous convolution for receptive field expansion, and dual skip connections mechanism to enhance global semantic consistency. Specifically, we introduce a Dilated Projection Block (DPB) that leverages atrous convolution to capture global context across multiple anatomical regions without additional parameters. Furthermore, we propose a Dual Skip Connection (DSC) mechanism to maintain encoder-decoder global consistency by fusing low-level and projected high-level features, mitigating blind spots introduced by atrous convolution. Extensive experiments show that our method significantly outperforms state-of-the-art methods, demonstrating its robustness and effectiveness in addressing the challenges of fetal brain tissue segmentation.
Junpeng Tan, Mingjin Chen, Chunmei Qing, Xin Zhang 0013, Xiangmin Xu 0001
ICIP5
2025 MMLoRA: Multitask Memory Parameter-Efficient Fine-Tuning for Multimodal SER
Yuanbo Fang, Xiaofen Xing, Xueru Li, Xiangmin Xu 0001
INTERSPEECH5
2025 Long-Context Speech Synthesis with Context-Aware Memory
Xiaofen Xing, Jingyuan Xing, Hangrui Hu, Xiangmin Xu 0001
INTERSPEECH6
2025 SA-RAS: Speaker-Aware Style Retrieval Augmented Generation for Expressive Zero-Shot Text-to-Speech Synthesis
Xueru Li, Jingyuan Xing, Xiaofen Xing, Xiangmin Xu 0001
INTERSPEECH5
2025 AA-SLLM: An Acoustically Augmented Speech Large Language Model for Speech Emotion Recognition
Jialong Mai, Xiaofen Xing, Yuanbo Fang, Xiangmin Xu 0001
INTERSPEECH5
2025 Chain-of-Thought Distillation with Fine-Grained Acoustic Cues for Speech Emotion Recognition
Jialong Mai, Xiaofen Xing, Yangbiao Li, Xiangmin Xu 0001
INTERSPEECH4
2025 EATS-Speech: Emotion-Adaptive Transformation and Priority Synthesis for Zero-Shot Text-to-Speech
Jingyuan Xing, Shuaiqi Chen, Xiaofen Xing, Xiangmin Xu 0001
INTERSPEECH5
2025 Heterogeneous Masked Attention-Guided Path Convolution for Functional Brain Network Analysis
Jiakun Xu, Xin Zhang 0013, Tong Xiong, Shengxian Chen, Xiaofen Xing, Jindou Hao, Xiangmin Xu 0001
MICCAI (12)7
2025 Optimal Feature Embedding for Document Large Visual Language Model
abstract
Document Large Vision Language Models excel in document-centric tasks and have become a key focus of research. Existing frameworks embed features from a lightweight, document-specific encoder into the first layer of a general-purpose Vision Language Model (VLM). However, this introduces a feature mismatch problem. VLMs typically consist of many stacked layers, with the feature hierarchy becoming increasingly abstract at higher layers. Specifically, the first-layer feature in a VLM is token-level, whereas the feature from the encoder is task-level, resulting in a mismatch. Consequently, it is crucial to identify an optimal layer within the VLM for embedding the encoder's features. Inspired by physics, we reformulate the search for the optimal embedding as a problem of finding the shortest time curve. Leveraging the properties of the shortest time curve, we theoretically derive a task-agnostic proxy score that requires only partial training and propose our searching framework, Brac4VLM. Our theoretical derivation shows that Brac4VLM reduces search time by 97.8% compared to brute-force methods. Experimental results further demonstrate that Brac4VLM identifies embedding points that closely align with the true optima. Moreover, the DocVLM with the optimal embedding position identified achieves state-of-the-art performance across various document-centric tasks. Codes: https://github.com/MaxKinny/Brac4VLM.
Fan Yang 0082, Ling Deng, Zhiyong Gan, Qisheng He, Yuanbo Fang, Xiangmin Xu 0001, Shuangping Huang, Tianshui Chen
ACM Multimedia6
2025 SemGesture: Synthesizing Semantically Enhanced and Coherent Gestures
abstract
Co-speech gestures are generally categorized into rhythmic and semantic gestures: rhythmic gestures align with speech rhythm and intonation, while semantic gestures convey specific meanings or emotions, enriching verbal communication. Most previous studies have focused on synthesizing rhythmic gestures, while recent methods have explored integrating large language models (LLMs) to retrieve semantic gestures and merge them with rhythmic ones. However, existing approaches primarily rely on textual context for retrieval, which may not fully capture the emotional and tonal nuances of speech, sometimes leading to semantic gestures that do not align with the speaker's intended expression. Additionally, common gesture fusion techniques often merge rhythmic and semantic gestures directly, causing discontinuity due to differences in their movement styles. To address these challenges, we propose SemGesture, a system designed to generate smooth and semantically accurate gestures. Our approach incorporates Context-aware Retrieval powered by a Large Audio Language Model, enabling precise retrieval of gestures that align with both the semantic and emotional aspects of speech. Additionally, the Gesture Fusion Module dynamically adjusts semantic gestures to harmonize with rhythmic gestures, ensuring seamless and coherent motion transitions. Extensive experiments demonstrate that SemGesture significantly outperforms existing methods in generating contextually accurate and visually natural gestures.
Pengsheng Liu, Zhaojie Chu, Xiaofen Xing, Xiangmin Xu 0001
ACM Multimedia4
2025 Multimodal speech emotion recognition via dynamic multilevel contrastive loss under local enhancement network
Weiquan Fan, Xiangmin Xu 0001, Fang Liu 0030, Xiaofen Xing
Expert Syst. Appl.2
2025 Low-complexity speaker embedding module with feature segmentation, transformation and reconstruction for few-shot speaker identification
Yanxiong Li, Qisheng Huang, Xiaofen Xing, Xiangmin Xu 0001
Expert Syst. Appl.4
2025 Facial expression recognition based on multi-task self-distillation with coarse and fine grained labels
Kailing Guo, Xiangmin Xu 0001
Expert Syst. Appl.4
2025 Tensor self-representation network for subspace clustering via alternating direction method of multipliers
Kailing Guo, Xiangmin Xu 0001
Knowl. Based Syst.3
2025 rPPG-TFCL: Time-frequency consistency learning for robust remote physiological measurement
Kailing Guo, Fang Liu 0030, Xiaofen Xing, Lin Wang 0004, Xiangmin Xu 0001, Zhanpeng Jin
Knowl. Based Syst.6
2025 Towards zero-shot human-object interaction detection via vision-language integration
Weiying Xue, Qi Liu 0005, Yuxiao Wang 0003, Zhenao Wei, Xiaofen Xing, Xiangmin Xu 0001
Neural Networks6
2025 Coordination Attention based Transformers with bidirectional contrastive loss for multimodal speech emotion recognition
Weiquan Fan, Xiangmin Xu 0001, Guohua Zhou, Xiaofang Deng, Xiaofen Xing
Speech Commun.2
2025 Road Surface State Change Detection Based on Binocular Vision for Autonomous Driving System
abstract
Road surface condition monitoring is crucial for enhancing transportation safety and efficiency, with applications in autonomous driving and urban infrastructure management. Existing methods often rely on single-camera setups or manual inspections, which are either insufficient for real-time monitoring or labor-intensive. This system focuses on two critical factors: road slope and surface damage, both significantly impacting driving safety and experience, highlighting the need for timely detection. To ensure accuracy and robustness, the system employs a binocular camera for detailed road environment insights and integrates urban sensing techniques. Its hardware deployment processes stereo vision data on embedded platforms, ensuring compatibility with urban IoT networks. This approach surpasses single-camera systems in detecting road surface variations. The research motivation stems from the pressing need to enhance road safety and driving conditions in urban areas. By analyzing binocular camera data and urban sensing technologies, the system offers real-time road condition analysis for effective decision-making. Regarding results, the system showed robust performance in detecting both road slope and surface damage. Slope detection achieved high accuracy with minimal error, and road damage detection reached an overall accuracy of 84%. The system remained stable across diverse conditions, including adverse weather and varying lighting.
Liangtian Zhao, Xiangmin Xu 0001, Shanshan Pei, Xiyuan Hu, Qiwei Xie
ACM Trans. Auton. Adapt. Syst.2
2025 Individual-Aware Attention Modulation for Unseen Speaker Emotion Recognition
abstract
In practical human-computer interaction (HCI) applications, robust speech emotion recognition (SER) for unseen speakers is crucial. Prior research has primarily focused on extracting common representations to enhance the generalization of cross-individual SER. However, most methods ignore the positive effects of individual characteristics. Actually, each speaker can be regarded as an independent individual domain. Personalized SER can be improved if the emotional expressions of individual speech characteristics are effectively utilized. To address the challenges in recognizing emotions for unseen speakers, this paper proposes a novel individual-aware attention modulation (IAM) model. Specifically, the IAM uses meta-learning techniques to extract modulation parameters for obtaining individual-related emotion expressions from individual characteristics. The base model is then modulated to facilitate the transfer of the common emotion representation space to an individual-specific emotion representation space. This transformation is achieved by applying attention modulation within the transformer-based model developed in this paper. In addition, we employ a meta-learning-based method to optimize model parameters, enhancing the adaptability of the model to unseen speakers, and a control factor is introduced to regulate the degree of individual modulation, thus enhancing the robustness of the modulation process. Experimental results demonstrate that the proposed model achieves significantly improved cross-individual SER performance.
Yuanbo Fang, Xiaofen Xing, Zhaojie Chu, Yifeng Du, Xiangmin Xu 0001
IEEE Trans. Affect. Comput.5
2025 CSE-GResNet: A Simple and Highly Efficient Network for Facial Expression Recognition
abstract
Facial expression recognition (FER) has recently attracted extensive attention in computer vision. However, existing methods mostly focus on the explicit performance and overlook their computational resources. Hence, achieving competitive performance while maintaining the model efficiency is still a huge challenge. To tackle these issues, we propose a highly lightweight yet effective Channel Shift-Enhancement Gabor-ResNet (CSE-GResNet) to capture the crucial visual properties in facial images. Concretely, we incorporate the Gabor Convolution (GConv) into ResNet to produce the robust GResNet as our backbone with limited memory cost. Furthermore, we propose extremely efficient Channel-Shift Module and Channel-Enhancement Module to insert in the GResNet in cascade. They are adopted to obtain and aggregate the facial informative representation from adjacent channels for extracting the subtle facial expression representation. We conduct extensive experiments on three wild datasets: RAF-DB, FER2013 and SFEW. The results show that the proposed CSE-GResNet achieves superior performance against the state-of-the-art methods with less computational and memory cost.
Shaoping Jiang, Xiaofen Xing, Fang Liu 0030, Xiangmin Xu 0001, Lin Wang 0004, Kailing Guo
IEEE Trans. Affect. Comput.4
2025 Alleviating One-to-Many Mapping in Talking Head Synthesis With Dynamic Adaptation Context and Style Adapter
abstract
Speech-driven talking head synthesis technology has made remarkable progress, but it still faces the challenge of one-to-many pathological mapping. The challenge results in inaccurate lip movements, ambiguity in facial expressions, and a lack of coherence during transitions between facial motions. The phenomenon is primarily caused by: (1) for one speaker, the same phoneme corresponds to a wide range of mouth shapes and facial expressions due to contextual variations, and (2) for the same spoken content, different speakers exhibit diverse facial motions as a result of unique speaking styles. In this work, we propose a novel framework, called AllTalk, to alleviate one-to-many pathological mapping, which enables a more vivid and natural talking head. Specifically, considering the asymmetry and dynamic nature of mouth shapes’ dependence on phoneme context, we propose a Dynamic Adaptive Context encoder to capture the context around the phoneme and its dynamics, thereby reducing the ambiguity in mapping speech to facial movements. Moreover, to alleviate the uncertainty caused by differences in speaking style, we propose a Style Adapter that expands a generic discrete motion space for the target speaker. The Style Adapter not only effectively represents general facial motions but also captures the personalized nuances of facial movements. To further enhance the fidelity of output, we introduce a Dynamic Gaussian Renderer based on 3D Gaussian Splatting, capable of producing stable and realistic rendering videos. Extensive qualitative and quantitative experiments demonstrate that AllTalk surpasses existing state-of-the-art methods, providing an effective solution to the challenge of one-to-many mapping. Project page: https://zjchu.github.io/projects/AllTalk.
Zhaojie Chu, Kailing Guo, Xiaofen Xing, Bolun Cai, Xiangmin Xu 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 A Review of AIoT-Based Human Activity Recognition: From Application to Technique
abstract
This scoping review paper redefines the Artificial Intelligence-based Internet of Things (AIoT) driven Human Activity Recognition (HAR) field by systematically extrapolating from various application domains to deduce potential techniques and algorithms. We distill a general model with adaptive learning and optimization mechanisms by conducting a detailed analysis of human activity types and utilizing contact or non-contact devices. It presents various system integration mathematical paradigms driven by multimodal data fusion, covering predictions of complex behaviors and redefining valuable methods, devices, and systems for HAR. Additionally, this paper establishes benchmarks for behavior recognition across different application requirements, from simple localized actions to group activities. It summarizes open research directions, including data diversity and volume, computational limitations, interoperability, real-time recognition, data security, and privacy concerns. Finally, we aim to serve as a comprehensive and foundational resource for researchers delving into the complex and burgeoning realm of AIoT-enhanced HAR, providing insights and guidance for future innovations and developments.
Wen Qi 0005, Xiangmin Xu 0001, Kun Qian 0003, Björn W. Schuller, Giancarlo Fortino, Andrea Aliverti
IEEE J. Biomed. Health Informatics2
2025 DCPTalk: Speech-Driven 3D Face Animation With Personalized Facial Dynamic Coupling Properties
abstract
Speech-driven 3D facial animation has emerged as a hot topic. During this process, movements in different facial regions are interdependent, influenced by the intricate interactions among facial muscles, and manifest personalized differences. The existing methods typically simplify the facial animation generation task to an infinitely thin surface skin deformation without an underlying structure, thereby ignoring the intricate and personalized dynamics of facial muscle activity. These methods tend to produce static or weak upper-face animations with an average facial movement style. In this work, we propose a novel framework, called DCPTalk, to mimic the intricate dynamics of facial muscle activity and portray personalized facial animations. Based on facial dynamic coupling properties, we propose Mouth2Face to simulate the facial muscle control system, yielding realistic and coordinated facial animations evoked by mouth movements. Mouth movements are easily synthesized from speech signals due to their direct correlation with phonetic articulation and vocal tract dynamics. To further enhance the detail of facial movements, we employ surface skin deformation to refine the facial animation derived from Mouth2Face. Furthermore, personal factors, including inherent physical traits and acquired speaking styles, directly determine the uniqueness and realism of facial animations. Inherent physical traits are embedded into Mouth2Face for constructing personalized facial muscle control system, while acquired speaking styles are employed to modulate external driving signals. Extensive qualitative and quantitative experiments as well as a user study indicate that DCPTalk outperforms the existing state-of-the-art methods.
Zhaojie Chu, Kailing Guo, Xiaofen Xing, Pengsheng Liu, Bolun Cai, Xiangmin Xu 0001
IEEE Trans. Multim.6
2025 Artistic Style Transfer via Fine-Grained Text Guidance and Contrastive Semantics Similarity
abstract
Due to the development of text-image multimodal methods, text is used to guide the style transfer of images, which has attracted growing attention. Notably, The existing text-guided image style methods are limited to expressing specific artistic style through simple text. It can only accept coarse-grained text input such as “Van Gogh” and “White Cloud”, and cannot understand fine-grained text input such as “The Night Café by Vincent van Gogh”. To this end, this paper proposes a novel artistic style transfer network based on the fine-grained text guidance and the contrastive semantics similarity, named as TCStyler. It can accept images or texts as style guidance, which is more suitable for fine-grained content understanding stylization. In this network, to address the issue of text-image cross-modal discrepancy, the residual attention feature mapper (RAFM) is introduced to constrain the differences between different modalities in feature space. Then, the global cascading style-sharing module (GCSM) is proposed for performing content-style feature fusion and image-text modality fusion by adopting a global feature-sharing strategy. Furthermore, the contrastive semantics similarity loss is designed to address the problem of multimodal universality. Quantitative and visualization experiments demonstrate that our TCStyler can handle fine-grained artistic text inputs and maintain consistency in the style transfer results guided by different modalities.
Chunmei Qing, Junpeng Tan, Jianxiu Jin, Xiangmin Xu 0001
IEEE Trans. Multim.5
2025 Compact Model Training by Low-Rank Projection With Energy Transfer
abstract
Low-rankness plays an important role in traditional machine learning but is not so popular in deep learning. Most previous low-rank network compression methods compress networks by approximating pretrained models and retraining. However, the optimal solution in the Euclidean space may be quite different from the one with low-rank constraint. A well-pretrained model is not a good initialization for the model with low-rank constraints. Thus, the performance of a low-rank compressed network degrades significantly. Compared with other network compression methods such as pruning, low-rank methods attract less attention in recent years. In this article, we devise a new training method, low-rank projection with energy transfer (LRPET), that trains low-rank compressed networks from scratch and achieves competitive performance. We propose to alternately perform stochastic gradient descent training and projection of each weight matrix onto the corresponding low-rank manifold. Compared to retraining on the compact model, this enables full utilization of model capacity since solution space is relaxed back to Euclidean space after projection. The matrix energy (the sum of squares of singular values) reduction caused by projection is compensated by energy transfer. We uniformly transfer the energy of the pruned singular values to the remaining ones. We theoretically show that energy transfer eases the trend of gradient vanishing caused by projection. In modern networks, a batch normalization (BN) layer can be merged into the previous convolution layer for inference, thereby influencing the optimal low-rank approximation (LRA) of the previous layer. We propose BN rectification to cut off its effect on the optimal LRA, which further improves the performance. Comprehensive experiments on CIFAR-10 and ImageNet have justified that our method is superior to other low-rank compression methods and also outperforms recent state-of-the-art pruning methods. For object detection and semantic segmentation, our method still achieves good compression results. In addition, we combine LRPET with quantization and hashing methods and achieve even better compression than the original single method. We further apply it in Transformer-based models to demonstrate its transferability. Our code is available at https://github.com/BZQLin/LRPET.
Kailing Guo, Zhenquan Lin, Canyang Chen, Xiaofen Xing, Fang Liu 0030, Xiangmin Xu 0001
IEEE Trans. Neural Networks Learn. Syst.6
2025 Multi-view Panoramic Image Style Transfer with Multi-scale Attention and Global Sharing
abstract
Style transfer for panoramic images is a challenging task, due to the problems associated with its unique structure, including edge discontinuities, pole distortion, fuzzy details, and memory limitation. In this article, we propose a novel Multi-view Transformation network for Panorama Style Transfer (MuTPST). First, this architecture has a multi-view panoramic transformation mechanism, which includes a multi-view cubic projection and a multi-view equirectangular re-projection of panoramic images. This can address pole distortion and edge discontinuity by skillfully applying multiple types of projections and transformations. To capture different levels of context and structure in the stylization stage, we carefully design a multi-scale attention content encoder, which can coordinate the distribution of visual attention across space and channels. Besides, by the sharing of global style features in thumbnails and patches, MuTPST can process ultra-high-resolution panoramic images (e.g., 10,000 \(\times\) 5,000 pixels) with limited GPU memory. Extensive experiments illustrate that the proposed method outperforms the state-of-the-art with a discernible improvement in panoramic image style transfer. More results and interactive features can be found on https://weiyang001.github.io/MuTPST/ .
Chunmei Qing, Junpeng Tan, Xiangmin Xu 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Disentangled Pre-Training for Human-Object Interaction Detection
abstract
Detecting human-object interaction (HOI) has long been limited by the amount of supervised data available. Recent approaches address this issue by pre-training according to pseudo-labels, which align object regions with HOI triplets parsed from image captions. However, pseudo-labeling is tricky and noisy, making HOI pre-training a complex process. Therefore, we propose an efficient disentangled pre-training method for HOI detection (DP-HOI) to address this problem. First, DP-HOI utilizes object detection and action recognition datasets to pre-train the detection and interaction decoder layers, respectively. Then, we arrange these decoder layers so that the pre-training architecture is consistent with the downstream HOI detection task. This facilitates efficient knowledge transfer. Specifically, the detection decoder identifies reliable human instances in each action recognition dataset image, generates one corresponding query, and feeds it into the interaction decoder for verb classification. Next, we combine the human instance verb predictions in the same image and impose image-level supervision. The DP-HOI structure can be easily adapted to the HOI detection task, enabling effective model parameter initialization. Therefore, it significantly enhances the performance of existing HOI detection models on a broad range of rare categories. The code and pre-trained weight are available at https://github.com/xingaoliIDP-HOI.
Zhuolong Li, Xingao Li, Changxing Ding, Xiangmin Xu 0001
CVPR4
2024 Texture-Preserving Diffusion Models for High-Fidelity Virtual Try-On
abstract
Image-based virtual try-on is an increasingly important task for online shopping. It aims to synthesize images of a specific person wearing a specified garment. Diffusion model-based approaches have recently become popular, as they are excellent at image synthesis tasks. However, these approaches usually employ additional image encoders and rely on the cross-attention mechanism for texture transfer from the garment to the person image, which affects the try-on's efficiency and fidelity. To address these issues, we propose an Texture-Preserving Diffusion (TPD) model for virtual try-on, which enhances the fidelity of the results and introduces no additional image encoders. Accordingly, we make contributions from two aspects. First, we propose to concatenate the masked person and reference garment images along the spatial dimension and utilize the resulting image as the input for the diffusion model's denoising UNet. This enables the original self-attention layers contained in the diffusion model to achieve efficient and accurate texture transfer. Second, we propose a novel diffusion-based method that predicts a precise inpainting mask based on the person and reference garment images, further enhancing the reliability of the try-on results. In addition, we integrate mask prediction and image synthesis into a single compact model. The experimental results show that our approach can be applied to various try-on tasks, e.g., garment-to-person and person-to-person try-ons, and significantly outperforms state-of-the-art methods on popular VITON, VITON-HD databases. Code is available at https://github.com/Ga14way/TPD.
Changxing Ding, Zhibin Hong, Xiangmin Xu 0001
CVPR6
2024 Clinical Scores Prediction and Medication Adjustment for Course of Parkinson's Disease
abstract
Parkinson's Disease (PD) is the second most prevalent neurodegenerative disorder worldwide, characterized by progressive motor and non-motor symptoms. Unfortunately, there are no definitive PD modifying therapies, so accurate course prediction in advance and appropriate medical adjustment are essential to slow down degenerative process from onset. This work addresses a novel challenge in PD course prediction specifically at month 60 (m60) of both motor and non-motor indicator utilizing Magnetic Resonance Imaging (MRI) and demographic data of previous years. A medication adjustment network based on Reinforcement Learning (RL) is utilized as an agent to simulate medication from professionals in a prediction environment. The proposed approach achieve a more accurate prediction on motor and non-motor simultaneously, demonstrating significant promise for longer-term PD course prediction compared to existing works. Furthermore, adjustable medication branch shows consistent with our advance result and provide possible guidance on medication for healthcare practitioners.
Xiaofen Xing, Xiangmin Xu 0001
ICASSP4
2024 Consistent Panoramic Video Style Transfer via Temporal-Spatial Cross Perception
Chunmei Qing, Junpeng Tan, Xiangmin Xu 0001
ICIC (6)4
2024 DropFormer: A Dynamic Noise-Dropping Transformer for Speech Emotion Recognition
Jialong Mai, Xiaofen Xing, Xiangmin Xu 0001
INTERSPEECH4
2024 Fetal MRI Reconstruction by Global Diffusion and Consistent Implicit Representation
Junpeng Tan, Xin Zhang 0013, Chunmei Qing, Chaoxiang Yang, He Zhang 0023, Gang Li 0001, Xiangmin Xu 0001
MICCAI (7)7
2024 Cortical Surface Reconstruction from 2D MRI with Segmentation-Constrained Super-Resolution and Representation Learning
Ruowen Qu, Dongzi Shi, Tong Xiong, Xiangmin Xu 0001, Xiaofen Xing, Xin Zhang 0013
MICCAI (2)5
2024 RetrievalMMT: Retrieval-Constrained Multi-Modal Prompt Learning for Multi-Modal Machine Translation
abstract
As an extension of machine translation, the primary objective of multi-modal machine translation is to optimize the utilization of visual information. Technically, image information is integrated into multi-modal fusion and alignment as an auxiliary modality through concepts or latent semantics, which are typically based on the Transformer framework. However, current approaches often ignore one modality to design numerous handcrafted features (e.g. visual concept extraction) and require training of all parameters in their framework. Therefore, it is worthwhile to explore multi-modal concepts or features to enhance performance and an efficient approach to incorporate visual information with minimal cost. Meanwhile, with the development of multi-modal large language models (MLLMs), they are faced with the visual hallucination issue of compromising performance, despite their powerful capabilities. Inspired by pioneering techniques in the multi-modal field, such as prompt learning and MLLMs, this paper innovatively explores the possibility of applying multi-modal prompt learning to this multi-modal machine translation task.
Yan Wang 0140, Yawen Zeng, Xiaofen Xing, Jin Xu 0014, Xiangmin Xu 0001
ICMR6
2024 Local Reactivation for Communication Efficient Federated Learning Based on Sparse Gradient Deviation
Chang Mu, Xiang Tian 0003, Kailing Guo, Xiangmin Xu 0001
PRCV (4)5
2024 As-Speech: Adaptive Style For Speech Synthesis
abstract
In recent years, there has been significant progress in Text-to-Speech (TTS) synthesis technology, enabling the high-quality synthesis of voices in common scenarios. In unseen situations, adaptive TTS requires a strong generalization capability for speaker style characteristics. However, the existing adaptive methods can only extract and integrate coarse-grained timbre or mixed rhythm attributes separately. In this paper, we propose AS-Speech, an adaptive style methodology that integrates the speaker timbre characteristics and rhythmic attributes into a unified framework for text-to-speech synthesis. Specifically, AS-Speech can accurately simulate style characteristics through fine-grained text-based timbre features and global rhythm information, and achieve high-fidelity speech synthesis through the diffusion model. Experiments show that our proposed model produces voices with higher similarity in terms of timbre and rhythm compared to a series of adaptive TTS models while maintaining the naturalness of synthetic speech. Samples are available at https://leezp99.github.io/as-speech-demo/
Xiaofen Xing, Shuaiqi Chen, Guoqiao Yu, Guanglu Wan, Xiangmin Xu 0001
SLT7
2024 Domain Adaption and Unified Knowledge Base Motivate Better Retrieval Models in Dialog Systems With RAG
abstract
Retrieval augmented generation (RAG) has emerged as a paradigm to address problems like hallucination in dialog systems based on large language model (LLM). Retrieval model is a key component in RAG framework for recalling relevant information. This paper describes our solution for FutureDial-RAG Challenge Track 1. We identify two primary challenges in this track: domain specificity and heterogeneity of knowledge base. To address the two challenges, we first adopt continual pre-training of a pre-trained retrieval model on both labeled and unlabeled data for domain adaption. Subsequently, we modify and expand the knowledge base, ensuring that each piece of knowledge is uniformly structured in a question-answer (QA) format. Finally, we construct negative samples based on the labeled data and the unified knowledge base, and fine-tune the retrieval model using contrastive learning. Our solution achieves a score of 2.023 on the dev set, which significantly outperforms the baseline.
Huadong Lin, Yirong Chen, Wenyu Tao, Mingyu Chen 0016, Xiangmin Xu 0001, Xiaofen Xing
SLT5
2024 Contrastive topic-enhanced network for video captioning
Yawen Zeng, Dongliang Liao, Gongfu Li, Jin Xu 0014, Hong Man, Xiangmin Xu 0001
Expert Syst. Appl.8
2024 EBMGC-GNF: Efficient Balanced Multi-View Graph Clustering via Good Neighbor Fusion
abstract
Exploiting consistent structure from multiple graphs is vital for multi-view graph clustering. To achieve this goal, we propose an Efficient Balanced Multi-view Graph Clustering via Good Neighbor Fusion (EBMGC-GNF) model which comprehensively extracts credible consistent neighbor information from multiple views by designing a Cross-view Good Neighbors Voting module. Moreover, a novel balanced regularization term based on p-power function is introduced to adjust the balance property of clusters, which helps the model adapt to data with different distributions. To solve the optimization problem of EBMGC-GNF, we transform EBMGC-GNF into an efficient form with graph coarsening method and optimize it based on accelareted coordinate descent algorithm. In experiments, extensive results demonstrate that, in the majority of scenarios, our proposals outperform state-of-the-art methods in terms of both effectiveness and efficiency.
Danyang Wu, Jitao Lu, Jin Xu 0014, Xiangmin Xu 0001, Feiping Nie 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 MASANet: Multi-Aspect Semantic Auxiliary Network for Visual Sentiment Analysis
abstract
Recently, multi-modal affective computing has demonstrated that introducing multi-modal information can enhance performance. However, multi-modal research faces significant challenges due to its high requirements regarding data acquisition, modal integrity, and feature alignment. The widespread use of multi-modal pre-training methods offers the possibility of aiding visual sentiment analysis by introducing cross-domain knowledge. This paper proposes a Multi-Aspect Semantic Auxiliary Network (MASANet) for visual sentiment analysis. Specifically, MASANet achieves modality expansion through cross-modal generation, making it possible to introduce cross-domain semantic assistance. Then, a cross-modal gating module and an adaptive modal fusion module are proposed for aspect-level and cross-modal interaction, respectively. In addition, a designed semantic polarity constraint loss is presented to improve sentiment multi-classification performance. Evaluations of eight widely-used affective image datasets demonstrate that our proposed method outperforms the state-of-the-art methods. Further ablation experiments and visualization results also confirm the effectiveness of the proposed method and its modules.
Jinglun Cen, Chunmei Qing, Haochun Ou, Xiangmin Xu 0001, Junpeng Tan
IEEE Trans. Affect. Comput.4
2024 Vesper: A Compact and Effective Pretrained Model for Speech Emotion Recognition
abstract
This paper presents a paradigm that adapts general large-scale pretrained models (PTMs) to speech emotion recognition task. Although PTMs shed new light on artificial general intelligence, they are constructed with general tasks in mind, and thus, their efficacy for specific tasks can be further improved. Additionally, employing PTMs in practical applications can be challenging due to their considerable size. Above limitations spawn another research direction, namely, optimizing large-scale PTMs for specific tasks to generate task-specific PTMs that are both compact and effective. In this paper, we focus on the speech emotion recognition task and propose an improVedemotion-specificpretrained encodercalled Vesper. Vesper is pretrained on a speech dataset based on WavLM and takes into account emotional characteristics. To enhance sensitivity to emotional information, Vesper employs an emotion-guided masking strategy to identify the regions that need masking. Subsequently, Vesper employs hierarchical and cross-layer self-supervision to improve its ability to capture acoustic and semantic representations, both of which are crucial for emotion recognition. Experimental results on the IEMOCAP, MELD, and CREMA-D datasets demonstrate that Vesper with 4 layers outperforms WavLM Base with 12 layers, and the performance of Vesper with 12 layers surpasses that of WavLM Large with 24 layers.
Xiaofen Xing, Peihao Chen, Xiangmin Xu 0001
IEEE Trans. Affect. Comput.4
2024 Bridge Graph Attention Based Graph Convolution Network With Multi-Scale Transformer for EEG Emotion Recognition
abstract
In multichannel electroencephalograph (EEG) emotion recognition, most graph-based studies employ shallow graph model for spatial characteristics learning due to node over-smoothing caused by an increase in network depth. To address over-smoothing, we propose the bridge graph attention-based graph convolution network (BGAGCN). It bridges previous graph convolution layers to attention coefficients of the final layer by adaptively combining each graph convolution output based on the graph attention network, thereby enhancing feature distinctiveness. Considering that graph-based networks primarily focus on local EEG channel relationships, we introduce a transformer for global dependency. Inspired by the neuroscience finding that neural activities of different timescales reflect distinct spatial connectivities, we modify the transformer to a multi-scale transformer (MT) by applying multi-head attention to multichannel EEG signals after 1D convolutions at different scales. MT learns spatial features more elaborately to enhance feature representation ability. By combining BGAGCN and MT, our model BGAGCN-MT achieves state-of-the-art accuracy under subject-dependent and subject-independent protocols across three benchmark EEG emotion datasets (SEED, SEED-IV and DREAMER). Notably, our model effectively addresses over-smoothing in graph neural networks and provides an efficient solution to learning spatial relationships of EEG features at different scales. Our code is available athttps://github.com/LogzZ.
Huachao Yan, Kailing Guo, Xiaofen Xing, Xiangmin Xu 0001
IEEE Trans. Affect. Comput.4
2024 CorrTalk: Correlation Between Hierarchical Speech and Facial Activity Variances for 3D Animation
abstract
Speech-driven 3D facial animation is a challenging cross-modal task that has attracted growing research interest. During speaking activities, the mouth displays strong motions, while the other facial regions typically demonstrate comparatively weak activity levels. Existing approaches often simplify the process by directly mapping single-level speech features to the entire facial animation, which overlook the differences in facial activity intensity leading to overly smoothed facial movements. In this study, we propose a novel framework, CorrTalk, which effectively establishes the temporal correlation between hierarchical speech features and facial activities of different intensities across distinct regions. A novel facial activity intensity prior is defined to distinguish between strong and weak facial activity, obtained by statistically analyzing facial animations. Based on the facial activity intensity prior, we propose a dual-branch decoding framework to synchronously synthesize strong and weak facial activity, which guarantees wider intensity facial animation synthesis. Furthermore, a weighted hierarchical feature encoder is proposed to establish temporal correlation between hierarchical speech features and facial activity at different intensities, which ensures lip-sync and plausible facial expressions. Extensive qualitatively and quantitatively experiments as well as a user study indicate that our CorrTalk outperforms existing state-of-the-art methods. The source code and supplementary video are publicly available at: https://zjchu.github.io/projects/CorrTalk/.
Zhaojie Chu, Kailing Guo, Xiaofen Xing, Yilin Lan, Bolun Cai, Xiangmin Xu 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 Label-Guided Dynamic Spatial-Temporal Fusion for Video-Based Facial Expression Recognition
abstract
Video-based facial expression recognition (FER) in the wild is a common yet challenging task. Extracting spatial and temporal features simultaneously is a common approach but may not always yield optimal results due to the distinct nature of spatial and temporal information. Extracting spatial and temporal features cascadingly has been proposed as an alternative approach However, the results of video-based FER sometimes fall short compared to image-based FER, indicating underutilization of spatial information of each frame and suboptimal modeling of frame relations in spatial-temporal fusion strategies. Although frame label is highly related to video label, it is overlooked in previous video-based FER methods. This paper proposes label-guided dynamic spatial-temporal fusion (LG-DSTF) that adopts frame labels to enhance the discriminative ability of spatial features and guide temporal fusion. By assigning each frame a video label, two auxiliary classification loss functions are constructed to steer discriminative spatial feature learning at different levels. The cross entropy between a uniform distribution and label distribution of spatial features is utilized to measure the classification confidence of each frame. The confidence values serve as dynamic weights to emphasize crucial frames during temporal fusion of spatial features. Our LG-DSTF achieves state-of-the-art results on FER benchmarks.
Xiang Tian 0003, Kailing Guo, Xiangmin Xu 0001
IEEE Trans. Multim.5
2024 Fourier Domain Robust Denoising Decomposition and Adaptive Patch MRI Reconstruction
abstract
The sparsity of the Fourier transform domain has been applied to magnetic resonance imaging (MRI) reconstruction in k -space. Although unsupervised adaptive patch optimization methods have shown promise compared to data-driven-based supervised methods, the following challenges exist in MRI reconstruction: 1) in previous k -space MRI reconstruction tasks, MRI with noise interference in the acquisition process is rarely considered. 2) Differences in transform domains should be resolved to achieve the high-quality reconstruction of low undersampled MRI data. 3) Robust patch dictionary learning problems are usually nonconvex and NP-hard, and alternate minimization methods are often computationally expensive. In this article, we propose a method for Fourier domain robust denoising decomposition and adaptive patch MRI reconstruction (DDAPR). DDAPR is a two-step optimization method for MRI reconstruction in the presence of noise and low undersampled data. It includes the low-rank and sparse denoising reconstruction model (LSDRM) and the robust dictionary learning reconstruction model (RDLRM). In the first step, we propose LSDRM for different domains. For the optimization solution, the proximal gradient method is used to optimize LSDRM by singular value decomposition and soft threshold algorithms. In the second step, we propose RDLRM, which is an effective adaptive patch method by introducing a low-rank and sparse penalty adaptive patch dictionary and using a sparse rank-one matrix to approximate the undersampled data. Then, the block coordinate descent (BCD) method is used to optimize the variables. The BCD optimization process involves valid closed-form solutions. Extensive numerical experiments show that the proposed method has a better performance than previous methods in image reconstruction based on compressed sensing or deep learning.
Junpeng Tan, Xin Zhang 0013, Chunmei Qing, Xiangmin Xu 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 Superpoint Transformer for 3D Scene Instance Segmentation
abstract
Most existing methods realize 3D instance segmentation by extending those models used for 3D object detection or 3D semantic segmentation. However, these non-straightforward methods suffer from two drawbacks: 1) Imprecise bounding boxes or unsatisfactory semantic predictions limit the performance of the overall 3D instance segmentation framework. 2) Existing method requires a time-consuming intermediate step of aggregation. To address these issues, this paper proposes a novel end-to-end 3D instance segmentation method based on Superpoint Transformer, named as SPFormer. It groups potential features from point clouds into superpoints, and directly predicts instances through query vectors without relying on the results of object detection or semantic segmentation. The key step in this framework is a novel query decoder with transformers that can capture the instance information through the superpoint cross-attention mechanism and generate the superpoint masks of the instances. Through bipartite matching based on superpoint masks, SPFormer can implement the network training without the intermediate aggregation step, which accelerates the network. Extensive experiments on ScanNetv2 and S3DIS benchmarks verify that our method is concise yet efficient. Notably, SPFormer exceeds compared state-of-the-art methods by 4.3% on ScanNetv2 hidden test set in terms of mAP and keeps fast inference speed (247ms per frame) simultaneously. Code is available at https://github.com/sunjiahao1999/SPFormer.
Chunmei Qing, Junpeng Tan, Xiangmin Xu 0001
AAAI4
2023 DST: Deformable Speech Transformer for Emotion Recognition
abstract
Enabled by multi-head self-attention, Transformer has exhibited remarkable results in speech emotion recognition (SER). Compared to the original full attention mechanism, window-based attention is more effective in learning fine-grained features while greatly reducing model redundancy. However, emotional cues are present in a multi-granularity manner such that the pre-defined fixed window can severely degrade the model flexibility. In addition, it is difficult to obtain the optimal window settings manually. In this paper, we propose a Deformable Speech Transformer, named DST, for SER task. DST determines the usage of window sizes conditioned on in-put speech via a light-weight decision network. Meanwhile, data-dependent offsets derived from acoustic features are utilized to adjust the positions of the attention windows, allowing DST to adaptively discover and attend to the valuable in-formation embedded in the speech. Extensive experiments on IEMOCAP and MELD demonstrate the superiority of DST.
Xiaofen Xing, Xiangmin Xu 0001, Jianxin Pang
ICASSP3
2023 DWFormer: Dynamic Window Transformer for Speech Emotion Recognition
abstract
Speech emotion recognition is crucial to human-computer interaction. The temporal regions that represent different emotions scatter in different parts of the speech locally. Moreover, the temporal scales of important information may vary over a large range within and across speech segments. Although transformer-based models have made progress in this field, the existing models could not precisely locate important regions at different temporal scales. To address the issue, we propose Dynamic Window transFormer (DWFormer), a new architecture that leverages temporal importance by dynamically splitting samples into windows. Self-attention mechanism is applied within windows for capturing temporal important information locally in a fine-grained way. Cross-window information interaction is also taken into account for global communication. DWFormer is evaluated on both the IEMO-CAP and the MELD datasets. Experimental results show that the proposed model achieves better performance than the previous state-of-the-art methods.
Shuaiqi Chen, Xiaofen Xing, Xiangmin Xu 0001
ICASSP5
2023 MGAT: Multi-Granularity Attention Based Transformers for Multi-Modal Emotion Recognition
abstract
Multi-modal emotion recognition is crucial for human-computer interaction. Many existing algorithms attempt to achieve multi-modal interactions through a cross-attention mechanism. Due to the problems of noise introduction and heavy computation in the original attention mechanism, window attention has become a new trend. However, emotions are presented asynchronously between different modalities, which makes it difficult to interact with emotional information between windows. Furthermore, multi-modal data are temporally misaligned, so single fixed window size is hard to describe cross-modal information. In this paper, we put these two issues into a unified framework and propose the multi-granularity attention based Transformers (MGAT). It addresses the emotional asynchrony and modality misalignment issues through a multi-granularity attention mechanism. Experimental results confirm the effectiveness of our method and the state-of-the-art performance is achieved on IEMOCAP.
Weiquan Fan, Xiaofen Xing, Bolun Cai, Xiangmin Xu 0001
ICASSP4
2023 Speaker-Aware Hierarchical Transformer For Personality Recognition In Multiparty Dialogues
abstract
Personality recognition is one of the core technologies in human-machine interaction, which has received increasing attention. Previous works mainly focus on essays or monologues, while personality traits reveal more in the interactions with others. Due to the lack of appropriate datasets, a few approaches aim to recognize personality traits in conversations, and most of them ignore interdependence between speakers and connection between conversations. In this paper, we create a multiparty dialogue-based personality dataset derived from CPED containing 1,195 data samples. We center on one speaker and extract related dialogues to compose each data sample annotated with speaker’s Big-Five traits, which is conducive to fully describe a center speaker using diverse cues of personality in different dialogues. Along the same lines, we propose a Speaker-aware Hierarchical Transformer named SH-Transformer to address above concerns, in which Personalized Embeddings (PE) adopt special tokens to distinguish center speakers in complete conversations and hierarchical Transformer capture diverse cues in utterances and conversations. Experimental results show that our method outperforms the non-interactive baseline by 1.38%, which confirms the necessity of considering both interactive information and diverse cues among dialogues. Our code will be released at github.com/Chloehxxx/SH-Transformer.
Wenjing Han, Yirong Chen, Xiaofen Xing, Guohua Zhou, Xiangmin Xu 0001
ICASSP5
2023 Text-Guided Generative Adversarial Network for Image Emotion Transfer
Siqi Zhu, Chunmei Qing, Xiangmin Xu 0001
ICIC (2)3
2023 Multi-Scale Transformer Network for Saliency Prediction on 360-Degree Images
abstract
The latest methods for saliency prediction on 360° images show that better results can be obtained using equirectangular (ERP) images as input. Due to the limitation of the receptive field, existing convolution-based networks cannot capture long-range information in complex 360° images. Although the transformer has the innate ability to capture long-range correlations with self-attention, large dataset requirement limit its application in saliency prediction of 360° images. In this paper, we present a novel Multi-scale Transformer framework for Saliency prediction on 360° images (MTSal360). The Multi-scale Transformer Module (MTM) is designed in the network to aggregate the contextual long-range information, which includes a Convolutional Positional Encoder (CPE) to enable the model could train and test on cubic and ERP format separately to address the insufficient data. Experiments on two public datasets illustrate that MTSal360 achieves better results over the state-of-the-art methods.
Chunmei Qing, Junpeng Tan, Xiangmin Xu 0001
ICIP4
2023 Cross Range Quantization for Network Compression
abstract
Quantization is effective in reducing model memory and accelerating inference, and is an important way to deploy deep neural networks on mobile smart devices. However, current popular learnable quantization functions often take simple truncation operations for values beyond the quantization range. We find that the truncation operation cause information loss and and restricts the update of values out of quantization range. To address this problem, we propose a universal cross range quantization (CRQ) method to reduce the information loss caused by the conventional truncation operation. CRQ splits the values exceeding the quantization range into two parts for separate quantization, and thus retain the information efficiently for performance improvement. In addition, we define a new metric named performance improvement efficiency (PIE) to measure the relationship between increased computation and performance improvement. Experiments on public benchmark image classification datasets show that CRQ achieves a significant accuracy gain with only a small increase in computation compared to the original learnable quantization method, and also outperforms many sophisticated designed state-of-the-art quantization methods in terms of accuracy and PIE.
Yicai Yang, Xiaofen Xing, Kailing Guo, Xiangmin Xu 0001, Fang Liu 0030
IJCNN5
2023 Exploring Downstream Transfer of Self-Supervised Features for Speech Emotion Recognition
Yuanbo Fang, Xiaofen Xing, Xiangmin Xu 0001
INTERSPEECH3
2023 Multi-Scale Temporal Transformer For Speech Emotion Recognition
Xiaofen Xing, Yuanbo Fang, Hengsheng Fan, Xiangmin Xu 0001
INTERSPEECH6
2023 Path-Based Heterogeneous Brain Transformer Network for Resting-State Functional Connectivity Analysis
Ruiyan Fang, Yu Li 0043, Xin Zhang 0013, Shengxian Chen, Xiangmin Xu 0001, Jieling Wu, Weili Lin, Li Wang 0026, Zhengwang Wu, Gang Li 0001
MICCAI (8)6
2023 Predicting Diverse Functional Connectivity from Structural Connectivity Based on Multi-contexts Discriminator GAN
Xin Zhang 0013, Lu Zhang 0050, Xiangmin Xu 0001, Dajiang Zhu
MICCAI (8)4
2023 Lightweight Multi-level Information Fusion Network for Facial Expression Recognition
Xiang Tian 0003, Xiangmin Xu 0001
MMM (2)4
2023 Multi-depth Fusion Transformer and Batch Piecewise Loss for Visual Sentiment Analysis
Haochun Ou, Chunmei Qing, Jinglun Cen, Xiangmin Xu 0001
PRCV (10)4
2023 Ink painting style transfer using asymmetric cycle-consistent GAN
Weining Wang 0003, Huan Ye, Fenghua Ye, Xiangmin Xu 0001
Eng. Appl. Artif. Intell.5
2023 Emotional generative adversarial network for image emotion transfer
Siqi Zhu, Chunmei Qing, Canqiang Chen, Xiangmin Xu 0001
Expert Syst. Appl.4
2023 Enhanced discriminative global-local feature learning with priority for facial expression recognition
Xiang Tian 0003, Kailing Guo, Xiangmin Xu 0001
Inf. Sci.5
2023 SpeechFormer++: A Hierarchical Efficient Framework for Paralinguistic Speech Processing
abstract
Paralinguistic speech processing is important in addressing many issues, such as sentiment and neurocognitive disorder analyses. Recently, Transformer has achieved remarkable success in the natural language processing field and has demonstrated its adaptation to speech. However, previous works on Transformer in the speech field have not incorporated the properties of speech, leaving the full potential of Transformer unexplored. In this paper, we consider the characteristics of speech and propose a general structure-based framework, called SpeechFormer++, for paralinguistic speech processing. More concretely, following the component relationship in the speech signal, we design a unit encoder to model the intra- and inter-unit information (i.e., frames, phones, and words) efficiently. According to the hierarchical relationship, we utilize merging blocks to generate features at different granularities, which is consistent with the structural pattern in the speech signal. Moreover, a word encoder is introduced to integrate word-grained features into each unit encoder, which effectively balances fine-grained and coarse-grained information. SpeechFormer++ is evaluated on the speech emotion recognition (IEMOCAP & MELD), depression classification (DAIC-WOZ) and Alzheimer's disease detection (Pitt) tasks. The results show that SpeechFormer++ outperforms the standard Transformer while greatly reducing the computational cost. Furthermore, it delivers superior results compared to the state-of-the-art approaches.
Xiaofen Xing, Xiangmin Xu 0001, Jianxin Pang
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Context-Based Adaptive Multimodal Fusion Network for Continuous Frame-Level Sentiment Prediction
abstract
Recently, video sentiment computing has become the focus of research because of its benefits in many applications such as digital marketing, education, healthcare, and so on. The difficulty of video sentiment prediction mainly lies in the regression accuracy of long-term sequences and how to integrate different modalities. In particular, different modalities may express different emotions. In order to maintain the continuity of long time-series sentiments and mitigate the multimodal conflicts, this paper proposes a novel Context-Based Adaptive Multimodal Fusion Network (CAMFNet) for consecutive frame-level sentiment prediction. A Context-based Transformer (CBT) module was specifically designed to embed clip features into continuous frame features, leveraging its capability to enhance the consistency of prediction results. Moreover, to resolve the multi-modal conflict between modalities, this paper proposed an Adaptive multimodal fusion (AMF) method based on the self-attention mechanism. It can dynamically determines the degree of shared semantics across modalities, enabling the model to flexibly adapt its fusion strategy. Through adaptive fusion of multimodal features, the AMF method effectively resolves potential conflicts arising from diverse modalities, ultimately enhancing the overall performance of the model. The proposed CAMFNet for consecutive frame-level sentiment prediction can ensure the continuity of long time-series sentiments. Extensive experiments illustrate the superiority of the proposed method especially in multimodal conflicts videos.
Maochun Huang, Chunmei Qing, Junpeng Tan, Xiangmin Xu 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Reflective Learning With Label Noise
abstract
Learning with noisy labels is one of the most challenging tasks in semi-supervised learning, and it poses significant problems in various practical applications. In the network learning process, the noisy labels concealed in the training dataset are easy to remember, resulting in poor generalization performance. To overcome this problem, inspired by the correction ability of humans – “think and learn from the past,” an end-to-end dynamic correction framework against label noise called Reflective Learning (RL) is proposed. This solution incorporates valuable knowledge from the past network training process to assist in correcting noisy labels. Specifically, during network training, a dynamic iterative function is implemented to adaptively correct noisy labels by employing the network’s predictive distribution information of all training epochs. This dynamic iterative function takes the form of a Standard Normal Distribution function to effectively match the changes of noisy label correction information contained in the network’s predictive probabilities. The proposed method is general and applicable to any backbone network and different types of noise without auxiliary information. Experiments are conducted on datasets with synthetic and real-world label noise datasets, including CIFAR-10, CIFAR-100, Tiny-ImageNet, and Clothing1M. They demonstrate that the proposed method is superior to the state-of-the-art results.
Lin Wang 0004, Xiangmin Xu 0001, Kailing Guo, Bolun Cai, Fang Liu 0030
IEEE Trans. Circuits Syst. Video Technol.2
2023 Context Sensing Attention Network for Video-based Person Re-identification
abstract
Video-based person re-identification (ReID) is challenging due to the presence of various interferences in video frames. Recent approaches handle this problem using temporal aggregation strategies. In this work, we propose a novel Context Sensing Attention Network (CSA-Net), which improves both the frame feature extraction and temporal aggregation steps. First, we introduce the Context Sensing Channel Attention (CSCA) module, which emphasizes responses from informative channels for each frame. These informative channels are identified with reference not only to each individual frame, but also to the content of the entire sequence. Therefore, CSCA explores both the individuality of each frame and the global context of the sequence. Second, we propose the Contrastive Feature Aggregation (CFA) module, which predicts frame weights for temporal aggregation. Here, the weight for each frame is determined in a contrastive manner: i.e., not only by the quality of each individual frame, but also by the average quality of the other frames in a sequence. Therefore, it effectively promotes the contribution of relatively good frames. Extensive experimental results on four datasets show that CSA-Net consistently achieves state-of-the-art performance.
Kan Wang 0004, Changxing Ding, Jianxin Pang, Xiangmin Xu 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2022 A Dataset for Falling Risk Assessment of the Elderly using Wearable Plantar Pressure
abstract
Falling is characterized by high incidence and great harm among the elderly. Timely assessing falling risk in daily life is helpful for reducing the occurrence of severe health outcomes. Establishing dataset for falling risk assessment based on wearable devices in the elderly is important work. However, current existing datasets might not reflect the natural gait of the subject due to the discomfort in wearing. Relevant data processing methods based on these datasets have limited practicability and might not be applied to real scenes in daily life. To make daily falling risk assessment possible, we proposed a novel approach to set up a continuous and wearable plantar pressure dataset of 48 older adults along with falling risk labels. The dataset was collected by plantar pressure monitoring shoes which were suitable for daily living spaces. Moreover, the Conv-LSTM algorithm was applied on the dataset, and the average classification result was up to 95.57%, reflecting the effectiveness of this dataset. The dataset is helpful for the studies of falling risk assessment and health monitoring among the elderly.
Guohua Hu, Jianxiu Jin, Shibin Wu, Junan Xie, Jianlin Ou, Zhuoming Chen, Xiangmin Xu 0001
BIBM9
2022 Key-Sparse Transformer for Multimodal Speech Emotion Recognition
abstract
Speech emotion recognition is a challenging research topic that plays a critical role in human-computer interaction. Multimodal inputs further improve the performance as more emotional information is used. However, existing studies learn all the information in the sample while only a small portion of it is about emotion. The redundant information will become noises and limit the system performance. In this paper, a key-sparse Transformer is proposed for efficient emotion recognition by focusing more on emotion related information. The proposed method is evaluated on the IEMOCAP and LSSED. Experimental results show that the proposed method achieves better performance than the state-of-the-art approaches.
Xiaofeng Xing, Xiangmin Xu 0001, Jianxin Pang
ICASSP3
2022 CS-GResNet: A Simple and Highly Efficient Network for Facial Expression Recognition
abstract
Facial expression recognition (FER) has recently attracted attention in computer vision. However, existing methods mostly focus on the explicit performance and overlook their computational resources and memory consumption. Hence, achieving promising performance while maintaining the efficiency of models is still a huge challenge. In this work, we propose a highly efficient Channel-Shift Gabor-ResNet (CS-GResNet) to capture the crucial visual properties in facial images. Concretely, we incorporate the Gabor Convolution (GConv) into ResNet to produce the significant GResNet as our backbone with limited memory cost. Furthermore, we adopt an extremely simple yet effective Channel-Shift Module inserted into the GResNet to obtain the facial informative representation via facilitating information exchanged among neighboring channels. We conduct extensive experiments on three wild datasets: RAF-DB, FER2013 and SFEW. The results show that our proposed CS-GResNet achieves superior performance against the state-of-the-art methods with less computational and memory cost. Codes are available at https://github.com/jsesr/CS-GResNet-PyTorch.
Shaoping Jiang, Xiangmin Xu 0001, Fang Liu 0030, Xiaofen Xing, Lin Wang 0004
ICASSP2
2022 DunhuangGAN: A Generative Adversarial Network for Dunhuang Mural Art Style Transfer
abstract
Style transfer has been successfully applied to various visual art creation tasks. However, due to the inherent style of Dunhuang mural art, existing methods can not produce high-quality mural art stylized images. We propose a novel model of DunhuangGAN for Dunhuang mural art style transfer. DunhuangGAN is based on the improved contrastive learning framework and optimized under the proposed multiple loss. Firstly, we propose a content-biased contrastive loss to alleviate the negative impacts caused by the inter-domain style differences. Secondly, we propose the line loss and color loss to simulate the line drawing modeling and heavy color of Dunhuang mural art. In addition, we introduce semantic loss to improve the visual effect of certain content element areas in generated images that are rare in the Dunhuang murals. Extensive experiments based on the collected dataset show that our method outperforms existing methods in the Dunhuang mural art style transfer task.
Weining Wang 0003, Huan Ye, Fenghua Ye, Xiangmin Xu 0001
ICME5
2022 SpeechFormer: A Hierarchical Efficient Framework Incorporating the Characteristics of Speech
abstract
Transformer has obtained promising results on cognitive speech signal processing field, which is of interest in various applications ranging from emotion to neurocognitive disorder analysis.However, most works treat speech signal as a whole, leading to the neglect of the pronunciation structure that is unique to speech and reflects the cognitive process.Meanwhile, Transformer has heavy computational burden due to its full attention operation.In this paper, a hierarchical efficient framework, called SpeechFormer, which considers the structural characteristics of speech, is proposed and can be served as a generalpurpose backbone for cognitive speech signal processing.The proposed SpeechFormer consists of frame, phoneme, word and utterance stages in succession, each performing a neighboring attention according to the structural pattern of speech with high computational efficiency.SpeechFormer is evaluated on speech emotion recognition (IEMOCAP & MELD) and neurocognitive disorder detection (Pitt & DAIC-WOZ) tasks, and the results show that SpeechFormer outperforms the standard Transformer-based framework while greatly reducing the computational cost.Furthermore, our SpeechFormer achieves comparable results to the state-of-the-art approaches.
Xiaofen Xing, Xiangmin Xu 0001, Jianxin Pang
INTERSPEECH3
2022 Simple-action-guided dictionary learning for complex action recognition
Fang Liu 0030, Xiangmin Xu 0001, Xiaofen Xing, Kailing Guo, Lin Wang 0004
Neurocomputing2
2022 GCB-Net: Graph Convolutional Broad Network and Its Application in Emotion Recognition
abstract
In recent years, emotion recognition has become a research focus in the area of artificial intelligence. Due to its irregular structure, EEG data can be analyzed by applying graphical based algorithms or models much more efficiently. In this work, a Graph Convolutional Broad Network (GCB-net) was designed for exploring the deeper-level information of graph-structured data. It used the graph convolutional layer to extract features of graph-structured input and stacks multiple regular convolutional layers to extract relatively abstract features. The final concatenation utilized the broad concept, which preserves the outputs of all hierarchical layers, allowing the model to search features in broad spaces. To improve the performance of the proposed GCB-net, the broad learning system (BLS) was applied to enhance its features. For comparison, two individual experiments were conducted to examine the efficiency of the proposed GCB-net based on the SJTU emotion EEG dataset (SEED) and DREAMER dataset respectively. In SEED, compared with other state-of-art methods, the GCB-net could better promote the accuracy (reaching 94.24 percent) on the DE feature of the all-frequency band. In DREAMER dataset, GCB-net performed better than other models with the same setting. Furthermore, the GCB-net reached high accuracies of 86.99, 89.32 and 89.20 percent on dimensions of Valence, Arousal and Dominance respectively. The experimental results showed the robust classifying ability of the GCB-net and BLS in EEG emotion recognition.
Tong Zhang 0015, Xuehan Wang, Xiangmin Xu 0001, C. L. Philip Chen
IEEE Trans. Affect. Comput.3
2022 ISNet: Individual Standardization Network for Speech Emotion Recognition
abstract
Speech emotion recognition plays an essential role in human-computer interaction. However, cross-individual representation learning and individual-agnostic systems are challenging due to the distribution deviation caused by individual differences. The existing related approaches mostly use the auxiliary task of speaker recognition to eliminate individual differences. Unfortunately, although these methods can reduce interindividual voiceprint differences, it is difficult to dissociate interindividual expression differences since each individual has its unique expression habits. In this paper, we propose an individual standardization network (ISNet) for speech emotion recognition to alleviate the problem of interindividual emotion confusion caused by individual differences. Specifically, we model individual benchmarks as representations of nonemotional neutral speech, and ISNet realizes individual standardization using the automatically generated benchmark, which improves the robustness of individual-agnostic emotion representations. In response to individual differences, we also propose more comprehensive and meaningful individual-level evaluation metrics. In addition, we continue our previous work to construct a challenging large-scale speech emotion dataset (LSSED). We propose a more reasonable division method of the training set and testing set to prevent individual information leakage. Experimental results on datasets of both large and small scales have proven the effectiveness of ISNet, and the new state-of-the-art performance is achieved under the same experimental conditions on IEMOCAP and LSSED.
Weiquan Fan, Xiangmin Xu 0001, Bolun Cai, Xiaofen Xing
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Intra- and Inter-Reasoning Graph Convolutional Network for Saliency Prediction on 360° Images
abstract
Cubic projection can be utilized to divide 360° images into multiple rectilinear images, with little distortion. However, the existing saliency prediction models fail to integrate semantic information of these images. In this paper, we address this by proposing an intra- and inter-reasoning graph convolutional network for saliency prediction on 360° images (SalReGCN360). The whole framework contains six sub-networks, each of which contains two branches. In the training phase, after utilizing Multiple Cubic Projection (MCP), six rectilinear images are simultaneously put into corresponding sub-networks. In one of the branches, the global features of a single rectilinear image are extracted by the intra-graph inference module to finely predict local saliency of 360° images. In the other branch, the contextual features are extracted by the inter-graph inference module to effectively integrate semantic information of six rectilinear images. Finally, the feature maps are generated by the two branches fusion, and six corresponding rectilinear saliency maps are predicted. Extensive experiments on two popular saliency datasets illustrate the superiority of the proposed model, especially the improvement in KLD metric.
Dongwen Chen, Chunmei Qing, Mengtao Ye, Xiangmin Xu 0001, Patrick Dickinson
IEEE Trans. Circuits Syst. Video Technol.5
2022 Path Signature Neural Network of Cortical Features for Prediction of Infant Cognitive Scores
abstract
Studies have shown that there is a tight connection between cognition skills and brain morphology during infancy. Nonetheless, it is still a great challenge to predict individual cognitive scores using their brain morphological features, considering issues like the excessive feature dimension, small sample size and missing data. Due to the limited data, a compact but expressive feature set is desirable as it can reduce the dimension and avoid the potential overfitting issue. Therefore, we pioneer the path signature method to further explore the essential hidden dynamic patterns of longitudinal cortical features. To form a hierarchical and more informative temporal representation, in this work, a novel cortical feature based path signature neural network (CF-PSNet) is proposed with stacked differentiable temporal path signature layers for prediction of individual cognitive scores. By introducing the existence embedding in path generation, we can improve the robustness against the missing data. Benefiting from the global temporal receptive field of CF-PSNet, characteristics consisted in the existing data can be fully leveraged. Further, as there is no need for the whole brain to work for a certain cognitive ability, a top K selection module is used to select the most influential brain regions, decreasing the model size and the risk of overfitting. Extensive experiments are conducted on an in-house longitudinal infant dataset within 9 time points. By comparing with several recent algorithms, we illustrate the state-of-the-art performance of our CF-PSNet (i.e., root mean square error of 0.027 with the time latency of 518 milliseconds for each sample).
Xin Zhang 0013, Hao Ni 0001, Chenyang Li 0007, Xiangmin Xu 0001, Zhengwang Wu, Li Wang 0026, Weili Lin, Gang Li 0001
IEEE Trans. Medical Imaging5
2022 Brain Connectivity Based Graph Convolutional Networks and Its Application to Infant Age Prediction
abstract
Infancy is a critical period for the human brain development, and brain age is one of the indices for the brain development status associated with neuroimaging data. The difference between the predicted age based on neuroimaging and the chronological age can provide an important early indicator of deviation from the normal developmental trajectory. In this study, we utilize the Graph Convolutional Network (GCN) to predict the infant brain age based on resting-state fMRI data. The brain connectivity obtained from rs-fMRI can be represented as a graph with brain regions as nodes and functional connections as edges. However, since the brain connectivity is a fully connected graph with features on edges, current GCN cannot be directly used for it is a node-based method for sparse graphs. Hence, we propose an edge-based Graph Path Convolution (GPC) method, which aggregates the information from different paths and can be naturally applied on dense graphs. We refer the whole model as Brain Connectivity Graph Convolutional Networks (BC-GCN). Further, two upgraded network structures are proposed by including the residual and attention modules, referred as BC-GCN-Res and BC-GCN-SE to emphasize the information of the original data and enhance influential channels. Moreover, we design a two-stage coarse-to-fine framework, which determines the age group first and then predicts the age using group-specific BC-GCN-SE models. To avoid accumulated errors from the first stage, a cross-group training strategy is adopted for the second stage regression models. We conduct experiments on infant fMRI scans from 6 to 811 days of age. The coarse-to-fine framework shows significant improvements when being applied to several models (reducing error over 10 days). Comparing with state-of-the-art methods, our proposed model BC-GCN-SE with coarse-to-fine framework reduces the mean absolute error of the prediction from >70 days to 49.9 days. The code is now available at https://github.com/SCUT-Xinlab/BC-GCN.
Yu Li 0043, Xin Zhang 0013, Jingxin Nie, Ruiyan Fang, Xiangmin Xu 0001, Zhengwang Wu, Dan Hu 0004, Li Wang 0026, Han Zhang 0002, Weili Lin, Gang Li 0001
IEEE Trans. Medical Imaging6
2022 Cross Parallax Attention Network for Stereo Image Super-Resolution
abstract
Stereo super-resolution (SR) aims to enhance the spatial resolution of one camera view using additional information from the other. Previous deep-learning-based stereo SR methods indeed improved the SR performance effectively by employing additional information, but they are unable to super-resolve stereo images where there are large disparities, or different types of epipolar lines. Moreover, in these methods, one model can only super-solve images of a particular view, and for one specific scale factor. This paper proposes a cross parallax attention stereo super-resolution network (CPASSRnet) which can perform stereo SR of multiple scale factors for both views, with a single model. To overcome the difficulties of large disparity and different types of epipolar lines, a cross parallax attention module (CPAM) is presented, which captures the global correspondence of additional information for each view, relative to the other. CPAM allows the two views to exchange additional information with each other according to the generated attention maps. Quantitative and qualitative results compared with the state of the arts illustrate the superiority of CPASSRnet. Ablation experiments demonstrate that the proposed components are effective and noise tests verify the robustness of CPASSRnet.
Canqiang Chen, Chunmei Qing, Xiangmin Xu 0001, Patrick Dickinson
IEEE Trans. Multim.3
2021 Cross-subject And Cross-device Wearable EEG Emotion Recognition Using Frontal EEG Under Virtual Reality Scenes
abstract
In recent years, with the rise of brain computer interface, automatic emotion recognition based on Electroencephalography (EEG) has attracted more and more attention. However, the emotional stimuli used in the existing studies are limited to music, pictures and videos, which cannot induce emotion well. Virtual reality (VR) can provide highly immersive 3D scenes, which can accurately and effectively induce emotion. Therefore, our work induced the subjects’ emotions through VR scenes, and collected the frontal EEG data based on wearable technology, innovatively proposed a VR-induced wearable frontal EEG emotion recognition dataset, which contains two sub datasets for cross-subject and cross-device research. Based on the dataset, we proposed a multi-spatial domain adaptation network (MSDAN) to eliminate the differences caused by individuals and EEG acquisition devices, and improve the generalization performance of the model in complex situations. MSDAN aimed to align the feature distributions of the source and target domains in multiple spaces and obtain the common features related to emotion. In the two sub datasets, our method achieved adequate results, which can obtain 72.08% and 75.14% accuracy in across-subject experiment,67.71% and 61.42% accuracy in across-device experiment, showing the significance and potential of the wearable EEG monitoring application based on VR in real life situation.
Feng Kuang, Haoqiang Hua, Shibin Wu, Xiangmin Xu 0001, Yunhe Liu 0003, Man Jiang
BIBM6
2021 Two-stream Gabor-AGraph Convolutional Networks for Facial Expression Recognition
abstract
Facial expression recognition (FER) has recently attracted much attention in computer vision. However, existing methods mostly focus on the texture information of faces and overlook their inherent topological features. Hence more informative and significant contents are ignored for expression recognition. In this work, we propose a Two-stream Gabor-AGraph Convolutional Network (2s-GAGCN) to exploit the facial texture and topological features simultaneously. The Gabor Stream and the Attention-Graph (AGraph) Stream are respectively introduced to capture the salient visual properties and discriminative landmark features of faces. In particular, we adopt a flexible node attention mechanism in AGraph Stream through utilizing global and local information to enhance the potential relationships among landmarks. Furthermore, a novel landmark feature descriptor is proposed to alleviate the redundant topological features, which shows promising improvement for the recognition accuracy. We conduct extensive experiments on two wild datasets: RAF-DB and SFEW. The results show that the proposed 2s-GAGCN achieves superior performance against the state-of-the-art methods.
Shaoping Jiang, Xiangmin Xu 0001, Xiaofen Xing, Lin Wang 0004, Fang Liu 0030
FG2
2021 Two-stream Global-Guided Attention Network for Facial Expression Recognition
abstract
Facial expression recognition (FER) in the wild is an important yet challenging problem because of uncontrolled conditions, such as occlusions, pose, and illumination. Most existing methods utilize the global and local information, but ignore the potential correlation between the global and local faces. In this paper, we propose a two-stream global-guided attention network (TGGAN) for FER in the wild. To further exploit the complementary relationship between local and global, we design a global-guided attention module (GGAM). Especially, a global guidance mechanism is proposed in GGAM to utilize the global information to guide the capture of key local features. Furthermore, inspired by the transformer, a self-attention mechanism is introduced in GGAM to emphasize salient face regions and fully integrate features extracted from global and local streams. We validate the proposed TGGAN on two wild datasets (FERPlus, RAF -DB) and further conduct experiments on their test subsets of occlusion and multi-poses. Extensive experiments show that the proposed TGGAN achieves superior performance against the state-of-the-art methods.
Yaoli Wen, Xiangmin Xu 0001, Fang Liu 0030, Xiaofen Xing, Lin Wang 0004
FG2
2021 LSSED: A Large-Scale Dataset and Benchmark for Speech Emotion Recognition
abstract
Speech emotion recognition is a vital contributor to the next generation of human-computer interaction (HCI). However, current existing small-scale databases have limited the development of related research. In this paper, we present LSSED, a challenging large-scale english speech emotion dataset, which has data collected from 820 subjects to simulate real- world distribution. In addition, we release some pre-trained models based on LSSED, which can not only promote the development of speech emotion recognition, but can also be transferred to related downstream tasks such as mental health analysis where data is extremely difficult to collect. Finally, our experiments show the necessity of large-scale datasets and the effectiveness of pre-trained models. The dateset will be released on https://github.com/tobefans/LSSED.
Weiquan Fan, Xiangmin Xu 0001, Xiaofen Xing, Dong-Yan Huang
ICASSP2
2021 Pruning the Unimportant or Redundant Filters? Synergy Makes Better
abstract
Filter pruning is a hot topic in convolutional neural network compression due to its friendliness to hardware implementation. Most pruning methods prune filters according to their importance, i.e., removing the filters that have little effect on the final performance of the network. While from another perspective, some recent works propose to prune upon the redundancy of filters. Filters pruned in this way usually have non-negligible effects on the final performance, whereas those effects could be compensated by the remaining filters. Since importance and redundancy pruning respectively captures the local and global information of a convolutional layer, which are mutually complementary to some extent, we propose a new pruning criterion that synergies both of those previous pruning criteria to make full use of the filter information. Comprehensive experiments on benchmark image classification datasets show the effectiveness of our proposed pruning criterion.
Yucheng Cai, Zhuowen Yin, Kailing Guo, Xiangmin Xu 0001
IJCNN4
2021 Weight Evolution: Improving Deep Neural Networks Training through Evolving Inferior Weight Values
abstract
To obtain good performance, convolutional neural networks are usually over-parameterized. This phenomenon has stimulated two interesting topics: pruning the unimportant weights for compression and reactivating the unimportant weights to make full use of network capability. However, current weight reactivation methods usually reactivate the entire filters, which may not be precise enough. Looking back in history, the prosperity of filter pruning is mainly due to its friendliness to hardware implementation, but pruning at a finer structure level, i.e., weight elements, usually leads to better network performance. We study the problem of weight element reactivation in this paper. Motivated by evolution, we select the unimportant filters and update their unimportant elements by combining them with the important elements of important filters, just like gene crossover to produce better offspring, and the proposed method is called weight evolution (WE). WE is mainly composed of four strategies. We propose a global selection strategy and a local selection strategy and combine them to locate the unimportant filters. A forward matching strategy is proposed to find the matched important filters and a crossover strategy is proposed to utilize the important elements of the important filters for updating unimportant filters. WE is plug-in to existing network architectures. Comprehensive experiments show that WE outperforms the other reactivation methods and plug-in training methods with typical convolutional neural networks, especially lightweight networks. Our code is available at https://github.com/BZQLin/Weight-evolution.
Zhenquan Lin, Kailing Guo, Xiaofen Xing, Xiangmin Xu 0001
ACM Multimedia4
2021 Spatiotemporal and frequential cascaded attention networks for speech emotion recognition
Shuzhen Li, Xiaofen Xing, Weiquan Fan, Bolun Cai, Perry Fordson, Xiangmin Xu 0001
Neurocomputing6
2021 Attention-aware concentrated network for saliency prediction
Pengqian Li, Xiaofen Xing, Xiangmin Xu 0001, Bolun Cai
Neurocomputing3
2021 Hierarchical Lifelong Learning by Sharing Representations and Integrating Hypothesis
abstract
In lifelong machine learning (LML) systems, consecutive new tasks from changing circumstances are learned and added to the system. However, sufficiently labeled data are indispensable for extracting intertask relationships before transferring knowledge in classical supervised LML systems. Inadequate labels may deteriorate the performance due to the poor initial approximation. In order to extend the typical LML system, we propose a novel hierarchical lifelong learning algorithm (HLLA) consisting of two following layers: 1) the knowledge layer consisted of shared representations and integrated knowledge basis at the bottom and 2) parameterized hypothesis functions with features at the top. Unlabeled data is leveraged in HLLA for pretraining of the shared representations. We also have considered a selective inherited updating method to deal with intertask distribution shifting. Experiments show that our HLLA method outperforms many other recent LML algorithms, especially when dealing with higher dimensional, lower correlation, and fewer labeled data problems.
Tong Zhang 0015, Guoxi Su, Chunmei Qing, Xiangmin Xu 0001, Bolun Cai, Xiaofen Xing
IEEE Trans. Syst. Man Cybern. Syst.4
2020 Attention-Based Fine-Grained Classification of Bone Marrow Cells
Weining Wang 0003, Peirong Guo, Lemin Li, Hongxia Shi, Xiangmin Xu 0001
ACCV (5)7
2020 Learning Spatio-Temporal Convolutional Network for Real-Time Object Tracking
abstract
Siamese series of tracking networks have shown great potentials in achieving balanced accuracy and beyond real-time speed. However, most of existing siamese trackers only consider appearance features of first frame, and hardly benefit from interframe information. The lack of latest temporal transformation degrades the tracking performance during challenges such as deformation and partial occlusion. In this paper we focus on making using of the rich information in latest consecutive frames to improve the feature representation of initial template frame. Specifically, the latest frames after 3d convolution are used to generate an attention map, which is then point-wise multiplied by the features of first frame to obtain the updated template. With the attention map, the template can adaptively cope with the deformation and occlusion of the target. Since the first frame is always used as the basis of the template, there is no cumulative error when using the latest frames for attention. Due to the shared 2d convolution of all frames, the feature map results can be reused so that the added module has almost no time-consuming effects. This module is easily embedded into different siamese trackers. Through verification, the module has significantly improved tracking performance in different backbone situations.
Hanzao Chen, Xiaofen Xing, Xiangmin Xu 0001
ICASSP3
2020 Adaptive Domain-Aware Representation Learning for Speech Emotion Recognition
Weiquan Fan, Xiangmin Xu 0001, Xiaofen Xing, Dong-Yan Huang
INTERSPEECH2
2020 Joint Image Quality Assessment and Brain Extraction of Fetal MRI Using Deep Learning
Lufan Liao, Xin Zhang 0013, Fenqiang Zhao, Tao Zhong 0002, Yuchen Pei, Xiangmin Xu 0001, Li Wang 0026, He Zhang 0023, Dinggang Shen, Gang Li 0001
MICCAI (6)6
2020 Infant Cognitive Scores Prediction with Multi-stream Attention-Based Temporal Path Signature Features
Xin Zhang 0013, Hao Ni 0001, Chenyang Li 0007, Xiangmin Xu 0001, Zhengwang Wu, Li Wang 0026, Weili Lin, Dinggang Shen, Gang Li 0001
MICCAI (7)5
2020 SalBiNet360: Saliency Prediction on 360° Images with Local-Global Bifurcated Deep Network
abstract
With the development of the virtual reality applications, predicting human visual attention on 360° images is valuable to content creators and encoding algorithms, and becomes essential to understand user behaviour. In this paper, we propose a local-global bifurcated deep network for saliency prediction on 360° images, which is named as SalBiNet360. In the global deep sub-network, multiple multi-scale contextual modules and a multilevel decoder are utilized to integrate the features from the middle and deep layers of the network. In the local deep sub-network, only one multi-scale contextual module and a single-level decoder are utilized to reduce the redundancy of local saliency maps. Finally, fused saliency maps are generated by linear combination of the global and local saliency maps. Experiments on two publicly available datasets illustrate that the proposed SalBiNet360 outperforms the tested state-of-the-art methods.
Dongwen Chen, Chunmei Qing, Xiangmin Xu 0001, Huansheng Zhu
VR3
2020 Exploring privileged information from simple actions for complex action recognition
Fang Liu 0030, Xiangmin Xu 0001, Tong Zhang 0015, Kailing Guo, Lin Wang 0004
Neurocomputing2
2019 Meta-Learning Perspective for Personalized Image Aesthetics Assessment
abstract
Image aesthetic is a highly subjective task. Thus, generic aesthetics models may lead to inconsistent user agreements even on the same image. Personalized aesthetics models can be employed to remedy the inconsistency issue. In real situation, users shared very small number of annotated images, which makes this problem more challenging. To solve problems above, unlike previous works that focused on user interactive or extracting simple yet effective image features, we address this by meta-learning. Meta-learning is a framework designed for quick adaption of an existing model to a new task with limited labeled data samples. In this way, we can leverage a small amount of annotated data from user and generate an effective personalized aesthetics model quickly. In addition, we proposed a novel meta-learning strategy and a novel meta regularization for our task. Experimental results demonstrate that our approach can effectively learn personalized aesthetics preferences and outperform existing methods on quantitative comparisons with a strong generalization ability.
Weining Wang 0003, Junjie Su, Lemin Li, Xiangmin Xu 0001, Jiebo Luo 0001
ICIP4
2019 Image Aesthetic Assessment Based on Perception Consistency
Weining Wang 0003, Lemin Li, Xiangmin Xu 0001
PRCV (2)4
2019 Spatial-spectral classification of hyperspectral images: a deep learning framework with Markov Random fields based modelling
abstract
For the spatial‐spectral classification of hyperspectral images (HSIs), a deep learning framework is proposed in this study, which consists of convolutional neural networks (CNNs) and Markov random fields (MRFs). Firstly, a CNN model to learn the deep spectral feature from the HSI is built and the class posterior probability distribution is estimated. The CNN with a dropout layer can relieve the overfitting in classification. The CNN is utilised as a pixel‐classifier, so it only works in the spectral domain. Then, the spatial information will be encoded by MRF‐based multilevel logistic prior for regularising the classification. To derive the correlation of both spectral and spatial features for improving algorithm performance, the marginal probability distribution in HSI is learned using MRF‐based loopy belief propagation. In comparison with several state‐of‐the‐art approaches for data classification on three publicly available HSI datasets, experimental results have demonstrated the superior performance of the proposed methodology.
Chunmei Qing, Jiawei Ruan, Xiangmin Xu 0001, Jinchang Ren, Jaime Zabalza
IET Image Process.3
2019 Accelerating Flexible Manifold Embedding for Scalable Semi-Supervised Learning
abstract
In this paper, we address the problem of large-scale graph-based semi-supervised learning for multi-class classification. Most existing scalable graph-based semi-supervised learning methods are based on the hard linear constraint or cannot cope with the unseen samples, which limits their applications and learning performance. To this end, we build upon our previous work flexible manifold embedding (FME) [1] and propose two novel linear-complexity algorithms called fast flexible manifold embedding (f-FME) and reduced flexible manifold embedding (r-FME). Both of the proposed methods accelerate FME and inherit its advantages. Specifically, our methods address the hard linear constraint problem by combining a regression residue term and a manifold smoothness term jointly, which naturally provides the prediction model for handling unseen samples. To reduce computational costs, we exploit the underlying relationship between a small number of anchor points and all data points to construct the graph adjacency matrix, which leads to simplified closed-form solutions. The resultant f-FME and r-FME algorithms not only scale linearly in both time and space with respect to the number of training samples but also can effectively utilize information from both labeled and unlabeled data. Experimental results show the effectiveness and scalability of the proposed methods.
Suo Qiu, Feiping Nie 0001, Xiangmin Xu 0001, Chunmei Qing, Dong Xu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2018 EEG Emotion Recognition Using Dynamical Graph Convolutional Neural Networks and Broad Learning System
Xuehan Wang, Tong Zhang 0015, Xiangmin Xu 0001, Long Chen 0001, Xiao-Fen Xing, C. L. Philip Chen
BIBM3
2018 Recurrent Neural Networks for Automatic Replay Spoofing Attack Detection
abstract
In order to enhance the security of automatic speaker verification (ASV) systems, automatic spoofing attack detection, which discriminates the fake audio recordings from genuine human speech, has gain much attention recently. Among various ways of spoofing attacks, replay attacks are one of the most effective and economical methods. In this paper, we explore using recurrent neural networks for automatic replay spoofing attack detection. More specifically, we focus on recurrent neural networks with more sophisticated recurrent units that involve a gating mechanism, such as a long short term memory (LSTM) unit and a recently proposed gated recurrent unit (GRU). Our experimental results on the ASVspoof 2017 showed that neural networks significantly outperform Gaussian mixture models (GMM). In addition, we achieved the best equal error rate of 9.81 % on the ASVspoof2017 and 1.077% on the BTAS 2016 by using GRU models, which outperform the best feed-forward neural networks by 19% and 46%, relatively and respectively.
Zhuxin Chen, Xiangmin Xu 0001, Dongpeng Chen
ICASSP4
2018 Perception Preserving Decolorization
abstract
Decolorization is a basic tool to transform a color image into a grayscale image, which is used in digital printing, stylized black-and-white photography, and in many single-channel image processing applications. While recent researches focus on retaining as much as possible meaningful visual features and color contrast. In this paper, we explore how to use deep neural networks for decolorization, and propose an optimization approach aiming at perception preserving. The system uses deep representations to extract content information based on human visual perception, and automatically selects suitable grayscale for decolorization. The evaluation experiments show the effectiveness of the proposed method.
Bolun Cai, Xiangmin Xu 0001, Xiaofen Xing
ICIP2
2018 Learning Adaptive Selection Network for Real-Time Visual Tracking
abstract
Offline-trained trackers based on convolutional neural networks (CNNs) have shown great potential in achieving balanced accuracy and real-time speed. However, offline-trained trackers are prone to drift to background clutters. In this paper, we present an adaptive selection network tracker (ASNT) to address the tracking drift problem. Inspired by feature selection technique used in other vision problems, we introduce a learnable selection unit for Siamese network based trackers. The selection unit enables the tracker to select relevant feature map automatically for the target. Channel dropout is applied in the selection unit to improve generalization performance for convolutional layers. To further improve the discrimination between background clutters and the target, an adaptive method is used to initialize the tracker for each video sequence. Experiments on OTB-2013 and VOT2014 datasets demonstrate that our ASNT tracker has a comparable performance against state-of-the-art methods, yet can run at a speed of over 100 fps.
Jiangfeng Xiong, Xiangmin Xu 0001, Bolun Cai, Xiaofen Xing, Kailing Guo
ICME2
2018 FReLU: Flexible Rectified Linear Units for Improving Convolutional Neural Networks
abstract
Rectified linear unit (ReLU) is a widely used activation function for deep convolutional neural networks. However, because of the zero-hard rectification, ReLU networks lose the benefits from negative values. In this paper, we propose a novel activation function called flexible rectified linear unit (FReLU) to further explore the effects of negative values. By redesigning the rectified point of ReLU as a learnable parameter, FReLU expands the states of the activation output. When a network is successfully trained, FReLU tends to converge to a negative value, which improves the expressiveness and thus the performance. Furthermore, FReLU is designed to be simple and effective without exponential functions to maintain low-cost computation. For being able to easily used in various network architectures, FReLU does not rely on strict assumptions by self-adaption. We evaluate FReLU on three standard image classification datasets, including CIFAR-10, CIFAR-100, and ImageNet. Experimental results show that FReLU achieves fast convergence and competitive performance on both plain and residual networks.
Suo Qiu, Xiangmin Xu 0001, Bolun Cai
ICPR2
2018 Visual Sentiment Analysis with Noisy Labels by Reweighting Loss
abstract
Visual sentiment analysis of online user generated content is important for many social media analysis tasks. However, label noise is common in sentiment analysis datasets, which deteriorate classification performance. To address this issue, we propose a novel visual sentiment analysis method based on loss reweighting to improve model robustness for label noise. First, a CNN is pre-trained with softmax loss on noisy labels datasets. Second, noise matrix is estimated by resorting and repositioning predicted probability, which is predicted by the pre-trained CNN. Third, converting noise estimation to the loss weight, the degeneration of sentiment classifiers performance caused by noisy labels can be compensated by re-training neural network with this reweighing loss. We conduct experiments on public sentiment datasets including Sentibank and Twitter datasets, and demonstrate that the proposed method outperforms state-of-the-art results.
Kailing Guo, Xiangmin Xu 0001, Lin Wang 0004, Bolun Cai
SMC2
2018 Improved Quantification of 18O Labeled LC-MS Based on I-Ching Divination Evolutionary Algorithm
abstract
An innovative quantification method for 18O labeled LC-MS data is proposed based on I-Ching divination evolutionary algorithm(IDEA). Considering label efficiency for calculating the least squares regression function, traditional methods based on genetic algorithm(GA) or other optimized algorithms will bring high level of computation complexity. The proposed method applies very flexible I-Ching operators(ICOs)— intrication operator, turnover operator, and mutual operator. The objective is the function of determining coefficients, which include the 18O/16O ratio r, the label efficiency f, and the abundance a of 16O. Comparing with GA, the proposed algorithm can significantly improve the accuracy and precision of peptide ratio measurements and better performs in the evolution procedure over mathematically calculating the function. Simultaneously we run the experiment with mix peptide raw data of predefined ratio. The result shows that our proposal algorithm is superior to the conventional GA in exploring optimum solution for better quantification accuracy.
Tianjun Li, C. L. Philip Chen, Long Chen 0001, Tong Zhang 0015, Bianna Chen, Xiangmin Xu 0001
SMC6
2018 Design of Highly Nonlinear Substitution Boxes Based on I-Ching Operators
abstract
This paper is to design substitution boxes (S-Boxes) using innovative I-Ching operators (ICOs) that have evolved from ancient Chinese I-Ching philosophy. These three operators-intrication, turnover, and mutual- inherited from I-Ching are specifically designed to generate S-Boxes in cryptography. In order to analyze these three operators, identity, compositionality, and periodicity measures are developed. All three operators are only applied to change the output positions of Boolean functions. Therefore, the bijection property of S-Box is satisfied automatically. It means that our approach can avoid singular values, which is very important to generate S-Boxes. Based on the periodicity property of the ICOs, a new network is constructed, thus to be applied in the algorithm for designing S-Boxes. To examine the efficiency of our proposed approach, some commonly used criteria are adopted, such as nonlinearity, strict avalanche criterion, differential approximation probability, and linear approximation probability. The comparison results show that S-Boxes designed by applying ICOs have a higher security and better performance compared with other schemes. Furthermore, the proposed approach can also be used to other practice problems in a similar way.
Tong Zhang 0015, C. L. Philip Chen, Long Chen 0001, Xiangmin Xu 0001, Bin Hu 0001
IEEE Trans. Cybern.4
2018 GoDec+: Fast and Robust Low-Rank Matrix Decomposition Based on Maximum Correntropy
abstract
GoDec is an efficient low-rank matrix decomposition algorithm. However, optimal performance depends on sparse errors and Gaussian noise. This paper aims to address the problem that a matrix is composed of a low-rank component and unknown corruptions. We introduce a robust local similarity measure called correntropy to describe the corruptions and, in doing so, obtain a more robust and faster low-rank decomposition algorithm: GoDec+. Based on half-quadratic optimization and greedy bilateral paradigm, we deliver a solution to the maximum correntropy criterion (MCC)-based low-rank decomposition problem. Experimental results show that GoDec+ is efficient and robust to different corruptions including Gaussian noise, Laplacian noise, salt & pepper noise, and occlusion on both synthetic and real vision data. We further apply GoDec+ to more general applications including classification and subspace clustering. For classification, we construct an ensemble subspace from the low-rank GoDec+ matrix and introduce an MCC-based classifier. For subspace clustering, we utilize GoDec+ values low-rank matrix for MCC-based self-expression and combine it with spectral clustering. Face recognition, motion segmentation, and face clustering experiments show that the proposed methods are effective and robust. In particular, we achieve the state-of-the-art performance on the Hopkins 155 data set and the first 10 subjects of extended Yale B for subspace clustering.
Kailing Guo, Liu Liu 0014, Xiangmin Xu 0001, Dong Xu 0001, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.3
2017 Emotion classification using deep neural networks and emotional patches
abstract
Emotion is closely related to healthy and abnormal mood is the alarm of our body. This paper is concentrated on the objective and accurate emotion classification using EEG signal. We propose emotional patches and combine it with the deep belief network(DBN) to achieve high-precision emotion classification. DBN is able to fit the distribution of the EEG signal and mapping the extracted feature to the higher-level characteristics space where we can easily perform high-precision classification. Compared with the other method, our method uses the emotional patches which have considered the temporal information of emotion and reduce the influence of noise. In addition, our model doesn't need to be trained twice to complete higher classification accuracy. We divide the EEG signal and choose the vital β frequency band where we perform feature extraction. Based on the SJTU Emotion EEG Dataset(SEED), we perform the emotion classification experiment and compare our method with the commonly used classifiers such as SVM, LR and CCA ect. The experimental result demonstrates that our method achieves the highest classification accuracy and outperform the state-of-theart emotion classification approaches based on EEG.
Jungming Huang, Xiangmin Xu 0001, Tong Zhang 0015
BIBM2
2017 Estimation of valence of emotion using two frontal EEG channels
abstract
Emotion recognition using EEG signals has become a hot research topic in the last few years. This paper aims at providing a novel method for emotion recognition using less channels of frontal EEG signals. By employing the asymmetry theory of frontal brain, a new method fusing spatial and frequency features was presented, which only adopted two channels of frontal EEG signals at Fp1 and Fp2. In order to estimate the efficiency of the method, a GBDT classifier was evaluated and selected, and the method was implemented on the DEAP database. The maximum and mean classification accuracy were achieved as 76.34% and 75.18% respectively, which exhibited the best result comparing with other related studies. This method is extremely suitable for wearable EEG monitoring applications in human daily life.
Shiyi Wu, Xiangmin Xu 0001, Bin Hu 0001
BIBM2
2017 Improving Training of Deep Neural Networks via Singular Value Bounding
abstract
Deep learning methods achieve great success recently on many computer vision problems. In spite of these practical successes, optimization of deep networks remains an active topic in deep learning research. In this work, we focus on investigation of the network solution properties that can potentially lead to good performance. Our research is inspired by theoretical and empirical results that use orthogonal matrices to initialize networks, but we are interested in investigating how orthogonal weight matrices perform when network training converges. To this end, we propose to constrain the solutions of weight matrices in the orthogonal feasible set during the whole process of network training, and achieve this by a simple yet effective method called Singular Value Bounding (SVB). In SVB, all singular values of each weight matrix are simply bounded in a narrow band around the value of 1. Based on the same motivation, we also propose Bounded Batch Normalization (BBN), which improves Batch Normalization by removing its potential risk of ill-conditioned layer transform. We present both theoretical and empirical results to justify our proposed methods. Experiments on benchmark image classification datasets show the efficacy of our proposed SVB and BBN. In particular, we achieve the state-of-the-art results of 3.06% error rate on CIFAR10 and 16.90% on CIFAR100, using off-the-shelf network architectures (Wide ResNets). Our preliminary results on ImageNet also show the promise in large-scale learning. We release the implementation code of our methods at www.aperture-lab.net/research/svb.
Kui Jia, Dacheng Tao, Shenghua Gao, Xiangmin Xu 0001
CVPR4
2017 Aesthetic Quality Assessment of Photos with Faces
Weining Wang 0003, Jiexiong Huang, Xiangmin Xu 0001, Quanzeng You, Jiebo Luo 0001
ICIG (3)3
2017 Automatic Classification of Focal Liver Lesion in Ultrasound Images Based on Sparse Representation
Weining Wang 0003, Yizi Jiang, Tingting Shi, Longzhong Liu, Qinghua Huang, Xiangmin Xu 0001
ICIG (2)6
2017 Edge/structure preserving smoothing via relativity-of-Gaussian
abstract
This paper presents a novel edge/structure-preserving image smoothing via relativity-of-Gaussian. As a simple local regularization, it performs the local analysis of scale features and globally optimizes its results into a piecewise smooth. The central idea to ensure proper texture smoothing is based on cross-scale relative that captures the weak textures from the most prominent edges/structures. Our method outperforms the previous methods in removing the detail information while preserving main image content.
Bolun Cai, Xiaofen Xing, Xiangmin Xu 0001
ICIP3
2017 Multi-scale convolutional neural networks for crowd counting
abstract
Crowd counting on static images is a challenging problem due to scale variations. Recently deep neural networks have been shown to be effective in this task. However, existing neural-networks-based methods often use the multi-column or multi-network model to extract the scale-relevant features, which is more complicated for optimization and computation wasting. To this end, we propose a novel multi-scale convolutional neural network (MSCNN) for single image crowd counting. Based on the multi-scale blobs, the network is able to generate scale-relevant features for higher crowd counting performances in a single-column architecture, which is both accuracy and cost effective for practical applications. Complemental results show that our method outperforms the state-of-the-art methods on both accuracy and robustness with far less number of parameters.
Lingke Zeng, Xiangmin Xu 0001, Bolun Cai, Suo Qiu, Tong Zhang 0015
ICIP2
2017 ResNet and Model Fusion for Automatic Spoofing Detection
Zhuxin Chen, Xiangmin Xu 0001
INTERSPEECH4
2017 Robust object tracking based on sparse representation and incremental weighted PCA
Xiaofen Xing, Fuhao Qiu, Xiangmin Xu 0001, Chunmei Qing, Yinrong Wu
Multim. Tools Appl.3
2017 Multimodal Estimation of Distribution Algorithms
abstract
Taking the advantage of estimation of distribution algorithms (EDAs) in preserving high diversity, this paper proposes a multimodal EDA. Integrated with clustering strategies for crowding and speciation, two versions of this algorithm are developed, which operate at the niche level. Then these two algorithms are equipped with three distinctive techniques: 1) a dynamic cluster sizing strategy; 2) an alternative utilization of Gaussian and Cauchy distributions to generate offspring; and 3) an adaptive local search. The dynamic cluster sizing affords a potential balance between exploration and exploitation and reduces the sensitivity to the cluster size in the niching methods. Taking advantages of Gaussian and Cauchy distributions, we generate the offspring at the niche level through alternatively using these two distributions. Such utilization can also potentially offer a balance between exploration and exploitation. Further, solution accuracy is enhanced through a new local search scheme probabilistically conducted around seeds of niches with probabilities determined self-adaptively according to fitness values of these seeds. Extensive experiments conducted on 20 benchmark multimodal problems confirm that both algorithms can achieve competitive performance compared with several state-of-the-art multimodal algorithms, which is supported by nonparametric tests. Especially, the proposed algorithms are very promising for complex problems with many local optima.
Qiang Yang 0008, Weineng Chen, Yun Li 0002, C. L. Philip Chen, Xiangmin Xu 0001, Jun Zhang 0003
IEEE Trans. Cybern.5
2016 Efficient action recognition from compressed depth maps
abstract
We propose an efficient action recognition scheme based solely on compressed depth maps. Each depth map is coded by a recently proposed scalable encoder that employs multi-scale breakpoints and an adaptive discrete wavelet transform (DWT). DWT coefficients describe smooth variations in depth while breakpoints communicate sharp boundaries. Both of these attributes are extracted from the bit-stream and utilized to construct features which are subject to a classification scheme for human action recognition. By extracting features from the compressed bit-stream computational complexity is significantly reduced thereby making the proposed scheme suitable for real-time applications. A L2-regularized collaborative representation classifier is employed for classification. The proposed scheme is computationally more efficient when compared with conventional approaches. Experimental results on the MSR 3D action dataset validate the effectiveness and efficiency of our proposed scheme.
Jie Miao, Xiaoyi Jia, Reji Mathew, Xiangmin Xu 0001, David S. Taubman, Chunmei Qing
ICIP4
2016 An electronic stethoscope for heart diseases based on micro-electro-mechanical-system microphone
abstract
In this paper, an electronic stethoscope for heart diseases based on micro-electro-mechanical-system microphone is developed. It consists of a stethoscope head, an electronic circuit, and an APP based on Android system. The circuit amplifies heart sounds picked up by the stethoscope head, and includes a band-pass filter to remove background noise. The supporting APP can record, replay, visualize, and upload the heart sounds. Experiment results show that the electronic stethoscope can collect high-quality heart sounds reliably.
Demiao Ou, Liping OuYang, Zhijun Tan, Hongqiang Mo, Xiang Tian 0003, Xiangmin Xu 0001
INDIN6
2016 Improved Music Genre Classification with Convolutional Neural Networks
Wenkang Lei, Xiangmin Xu 0001, Xiaofeng Xing
INTERSPEECH3
2016 Image and video dehazing using view-based cluster segmentation
abstract
To avoid distortion in sky regions and make the sky and white objects clear, in this paper we propose a new image and video dehazing method utilizing the view-based cluster segmentation. Firstly, GMM(Gaussian Mixture Model)is utilized to cluster the depth map based on the distant view to estimate the sky region and then the transmission estimation is modified to reduce distortion. Secondly, we present to use GMM based on Color Attenuation Prior to divide a single hazy image into K classifications, so that the atmospheric light estimation is refined to improve global contrast. Finally, online GMM cluster is applied to video dehazing. Extensive experimental results demonstrate that the proposed algorithm can have superior haze removing and color balancing capabilities.
Chunmei Qing, Xiangmin Xu 0001, Bolun Cai
VCIP3
2016 Wireless and sensorless 3D ultrasound imaging
Haitao Gao, Qinghua Huang, Xiangmin Xu 0001, Xuelong Li 0001
Neurocomputing3
2016 Synthesized computational aesthetic evaluation of photos
Weining Wang 0003, Dong Cai, Li Wang 0026, Qinghua Huang, Xiangmin Xu 0001, Xuelong Li 0001
Neurocomputing5
2016 A multi-scene deep learning model for image aesthetic evaluation
Weining Wang 0003, Mingquan Zhao, Li Wang 0026, Jiexiong Huang, Chengjia Cai, Xiangmin Xu 0001
Signal Process. Image Commun.6
2016 DehazeNet: An End-to-End System for Single Image Haze Removal
abstract
Single image haze removal is a challenging ill-posed problem. Existing methods use various constraints/priors to get plausible dehazing solutions. The key to achieve haze removal is to estimate a medium transmission map for an input hazy image. In this paper, we propose a trainable end-to-end system called DehazeNet, for medium transmission estimation. DehazeNet takes a hazy image as input, and outputs its medium transmission map that is subsequently used to recover a haze-free image via atmospheric scattering model. DehazeNet adopts convolutional neural network-based deep architecture, whose layers are specially designed to embody the established assumptions/priors in image dehazing. Specifically, the layers of Maxout units are used for feature extraction, which can generate almost all haze-relevant features. We also propose a novel nonlinear activation function in DehazeNet, called bilateral rectified linear unit, which is able to improve the quality of recovered haze-free image. We establish connections between the components of the proposed DehazeNet and those used in existing methods. Experiments on benchmark images show that DehazeNet achieves superior performance over existing methods, yet keeps efficient and easy to use.
Bolun Cai, Xiangmin Xu 0001, Kui Jia, Chunmei Qing, Dacheng Tao
IEEE Trans. Image Process.2
2016 BIT: Biologically Inspired Tracker
abstract
Visual tracking is challenging due to image variations caused by various factors, such as object deformation, scale change, illumination change, and occlusion. Given the superior tracking performance of human visual system (HVS), an ideal design of biologically inspired model is expected to improve computer visual tracking. This is, however, a difficult task due to the incomplete understanding of neurons' working mechanism in the HVS. This paper aims to address this challenge based on the analysis of visual cognitive mechanism of the ventral stream in the visual cortex, which simulates shallow neurons (S1 units and C1 units) to extract low-level biologically inspired features for the target appearance and imitates an advanced learning mechanism (S2 units and C2 units) to combine generative and discriminative models for target location. In addition, fast Gabor approximation and fast Fourier transform are adopted for real-time learning and detection in this framework. Extensive experiments on large-scale benchmark data sets show that the proposed biologically inspired tracker performs favorably against the state-of-the-art methods in terms of efficiency, accuracy, and robustness. The acceleration technique in particular ensures that biologically inspired tracker maintains a speed of approximately 45 frames/s.
Bolun Cai, Xiangmin Xu 0001, Xiaofen Xing, Kui Jia, Jie Miao, Dacheng Tao
IEEE Trans. Image Process.2
2016 Simple to Complex Transfer Learning for Action Recognition
abstract
Recognizing complex human actions is very challenging, since training a robust learning model requires a large amount of labeled data, which is difficult to acquire. Considering that each complex action is composed of a sequence of simple actions which can be easily obtained from existing data sets, this paper presents a simple to complex action transfer learning model (SCA-TLM) for complex human action recognition. SCA-TLM improves the performance of complex action recognition by leveraging the abundant labeled simple actions. In particular, it optimizes the weight parameters, enabling the complex actions to be learned to be reconstructed by simple actions. The optimal reconstruct coefficients are acquired by minimizing the objective function, and the target weight parameters are then represented as a combination of source weight parameters. The main advantage of the proposed SCA-TLM compared with existing approaches is that we exploit simple actions to recognize complex actions instead of only using complex actions as training samples. To validate the proposed SCA-TLM, we conduct extensive experiments on two well-known complex action data sets: 1) Olympic Sports data set and 2) UCF50 data set. The results show the effectiveness of the proposed SCA-TLM for complex action recognition.
Fang Liu 0030, Xiangmin Xu 0001, Shuoyang Qiu, Chunmei Qing, Dacheng Tao
IEEE Trans. Image Process.2
2015 BIT: Bio-inspired tracker
abstract
Visual tracking is a challenging problem due to various factors such as deformation, rotation and illumination. As is well known, given the superior tracking performance of human vision, bio-inspired model is expected to improve the computer visual tracking. However, the design of bio-inspired tracking framework is challenging, due to the incomplete comprehension and hyper-scale of senior neurons, which will influence the effectiveness and real-time performance of the tracker. According to the ventral stream in visual cortex, a novel bio-inspired tracker (BIT) is proposed, which simulates shallow neurons (S1 and C1) to extract low-level bio-inspired feature for target appearance and imitates senior learning mechanism (S2 and C2) to combine generative and discriminative model for position estimation. In addition, Fast Fourier Transform (FFT) is adopted for real-time learning and detection in this framework. On the recent benchmark[1], extensive experimental results show BIT performs favorably against state-of-the-art methods in terms of accuracy and robustness.
Bolun Cai, Xiangmin Xu 0001, Xiaofen Xing, Chunmei Qing
ICIP2
2015 Residue boundary histograms for action recognition in the compressed domain
abstract
Traditional action recognition approaches are too slow for real-time or large-scale applications. This problem has been tackled by replacing optical flow with motion vectors from the compressed domain. Yet further usage of compressed domain information for action recognition is possible. Discrete cosine transform (DCT) coefficients, which correspond to residue data, represent information which the block based motion vectors fail to capture. We propose a set of residue boundary histograms (RBH) features for action recognition, separating each DCT block into four parts to obtain four small residue maps and then encoding each residue map by histogram-based descriptors to obtain local features. Experimental results on three challenging datasets show that proposed RBH features improve upon motion vector based features significantly. While more than 100× faster, the results are highly competitive compared with traditional action recognition approaches.
Jie Miao, Xiangmin Xu 0001, Reji Mathew
ICIP2
2015 An automatic energy-based region growing method for ultrasound image segmentation
abstract
Segmentation for lesion region in ultrasound images is crucial for computer-aided diagnosis system. But it has always been a difficult task due to the defects inherent in the ultrasound images. In this paper, we propose an automatic energy-based region growing (AERG) method to automatically segment the lesion region in ultrasound images of liver. At first, the seed point of lesion region is automatically selected by sparse reconstruction algorithm. Then the region growing process is controlled by a novel energy function including both internal and external energy, so as to make the edge of the region converge to the contour of the lesion accurately and keep a small internal difference at the same time. Experiment results show that our method could improve the segmentation accuracy in comparison with other four often used segmentation methods.
Weining Wang 0003, Jiachang Li, Yizi Jiang, Yi Xing, Xiangmin Xu 0001
ICIP5
2015 An efficient image aesthetic analysis system using Hadoop
Weining Wang 0003, Weijian Zhao, Chengjia Cai, Jiexiong Huang, Xiangmin Xu 0001
Signal Process. Image Commun.5
2015 Temporal Variance Analysis for Action Recognition
abstract
Slow feature analysis (SFA) extracts slowly varying signals from input data and has been used to model complex cells in the primary visual cortex (V1). It transmits information to both ventral and dorsal pathways to process appearance and motion information, respectively. However, SFA only uses slowly varying features for local feature extraction, because they represent appearance information more effectively than motion information. To better utilize temporal information, we propose temporal variance analysis (TVA) as a generalization of SFA. TVA learns a linear transformation matrix that projects multidimensional temporal data to temporal components with temporal variance. Inspired by the function of V1, we learn receptive fields by TVA and apply convolution and pooling to extract local features. Embedded in the improved dense trajectory framework, TVA for action recognition is proposed to: 1) extract appearance and motion features from gray using slow and fast filters, respectively; 2) extract additional motion features using slow filters from horizontal and vertical optical flows; and 3) separately encode extracted local features with different temporal variances and concatenate all the encoded features as final features. We evaluate the proposed TVA features on several challenging data sets and show that both slow and fast features are useful in the low-level feature extraction. Experimental results show that the proposed TVA features outperform the conventional histogram-based features, and excellent results can be achieved by combining all TVA features.
Jie Miao, Xiangmin Xu 0001, Shuoyang Qiu, Chunmei Qing, Dacheng Tao
IEEE Trans. Image Process.2
2014 LDA based compact and discriminative dictionary learning for sparse coding
abstract
The dictionary response usually affects the recognition results directly as it represents the original data and usually serves as the input of the classifier. However, the over-complete dictionary usually results in high dimensional response and redundancy. The application of the linear discriminant analysis (LDA)-based mapping method transforms the original dictionary response to be more discriminative for compact dictionary learning, resulting in high intra-class similarity and high inter-class dissimilarity in the response domain for better classification. By analyzing the recognition rate, the compactness and the purity, the proposed method can learn a small size of compact and discriminative dictionary with global optimization, and it can get a comparable or even better performance than the over-complete dictionary with much less computation cost. Experimental results demonstrate that the proposed approach also outperforms several recently proposed compact dictionary learning methods on human action recognition and object classification.
Jiayong Chen, Xiangmin Xu 0001, Chunmei Qing, Jianxiu Jin
ICIP2
2014 High speed deep networks based on Discrete Cosine Transformation
abstract
The traditional deep networks take raw pixels of data as input, and automatically learn features using unsupervised learning algorithms. In this configuration, in order to learn good features, the networks usually have multi-layer and many hidden units which lead to extremely high training time costs. As a widely used image compression algorithm, Discrete Cosine Transformation (DCT) is utilized to reduce image information redundancy because only a limited number of the DCT coefficients can preserve the most important image information. In this paper, it is proposed that a novel framework by combining DCT and deep networks for high speed object recognition system. The use of a small subset of DCT coefficients of data to feed into a 2-layer sparse auto-encoders instead of raw pixels. Because of the excellent decorrelation and energy compaction properties of DCT, this approach is proved experimentally not only efficient, but also it is a computationally attractive approach for processing high-resolution images in a deep architecture.
Xiaoyi Zou, Xiangmin Xu 0001, Chunmei Qing, Xiaofen Xing
ICIP2
2014 Visual saliency detection based on region descriptors and prior knowledge
Weining Wang 0003, Dong Cai, Xiangmin Xu 0001, Alan Wee-Chung Liew
Signal Process. Image Commun.3
2013 A novel incremental weighted PCA algorithm for visual tracking
abstract
This paper addresses the drifting problem in online visual tracking. The tracking result is usually described by a bounding box, which inevitably contains background in the box and causes drifting. This paper tries to treat the background part and the truly target part discriminatively to reduce the effect of background. A novel incremental weighted PCA (IWPCA) algorithm is proposed. The most important contribution of this paper is an approximation method which limits the great and increasing computational cost, caused by the weighted form, to a constant. Therefore it is suitable for online tracking. Combined with particle filter, the proposed algorithm achieves superior results in several challenging video sequences in terms of stableness and accuracy, and greatly alleviates the drifting problem.
Kailing Guo, Xiangmin Xu 0001, Fuhao Qiu, Jiayong Chen
ICIP2