VLDB 2026 Research / reviewers in the wild / expert
Xiaofen Xing
dblp:41/9939
· DBLP profile ↗
74ranked-venue papers
1as first author
65since 2021 · last 2026
0000-0002-0016-9055ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 48 · 1 first-author · 39 since 2021Artificial intelligence and machine learning · 38 · 37 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A dual uncertainty-aware fusion framework for face expression recognition in the wildabstractFacial Expression Recognition(FER) is a key task in the broader landscape of affective computing and human-computer interaction, enabling machines to interpret human emotions. To better learn discriminative features under complex facial variations, recent FER research has increasingly adopted multi-branch fusion architectures that aim to capture complementary features from diverse perspectives. However, existing multi-branch fusion strategies, including static weighting, simple concatenation, or uncertainty-aware modeling, lack the capacity to comprehensively capture and reconcile the reliability variations across both individual instances and structural branches. To overcome these limitations, we propose a novel multi-branch fusion strategy, named Dual Uncertainty-Aware Fusion Framework(DUAFF), which improves the discriminability of integrated features by simultaneously modeling instance-wise uncertainty and inter-branch correlations. Specifically, the proposed method comprises two complementary modules: Instance-Discrepant Uncertainty-Aware Fusion Module (ID-UAFM) and Branch-Discrepant Uncertainty-Aware Fusion Module (BD-UAFM). ID-UAFM is introduced to perform channel-wise entropy analysis between semantically distinct samples to estimate instance-level uncertainty, enabling selective channel-wise fusion that emphasizes reliable representations while suppressing uncertain responses. BD-UAFM is further proposed to capture structural uncertainty by evaluating the relative reliability of features across multiple branches and adaptively weighting their contributions based on inter-branch discrepancies. Experimental results demonstrate that the proposed DUAFF consistently outperforms POSTER across three benchmark datasets, achieving accuracy improvements of 0.23 % on RAF-DB, 0.69 % on FER2013, and 0.29 % on AffectNet (7-class), thereby confirming its effectiveness in enhancing the reliability and discriminability of facial representations. Wenfeng Jiang, Lin Wang 0004, Fang Liu 0030, Chunmei Qing, Xiaofen Xing, Xiangmin Xu 0001, Weiquan Fan, Zhanpeng Jin |
Expert Syst. Appl. | 6 |
| 2026 | Random graph construction-based hierarchical attention multi-task graph convolution network for sEEG SOZ identification
Huachao Yan, Kailing Guo, Shiwei Song, Xiaofen Xing, Xiangmin Xu 0001 |
Neurocomputing | 4 |
| 2026 | MMRepAgent: Explainable stock earnings forecasting via multimodal report agent framework
Xiangyu Li 0010, Yawen Zeng, Xiaofen Xing, Jin Xu 0014, Xiangmin Xu 0001 |
Knowl. Based Syst. | 3 |
| 2026 | SiaTalker: Siamese Emotion Injection for Coarse-to-Fine Speech-Driven 3D Facial AnimationabstractDuring communication, people inadvertently express emotions through facial animations, with emotion constantly fluctuating throughout the conversation. Existing speech-driven facial animation works predominantly rely on sentence-level coarse-grained emotion labels to impart emotional expression information, thereby overlooking temporal emotion variations. How to make 3D talking heads articulate fluently while adeptly expressing mood fluctuations is the problem addressed by our work. We propose a framework, named SiaTalker, which exquisitely refines emotions of facial animations from coarse-grained levels to dual-perspective interactive fine-grained levels, resulting in highly coordinated and fluid emotional expressions. Specifically, we propose the coarse-to-fine dual granularity fusion decoding framework to connect global and local emotional aspects, allowing nuanced shifts within a consistent emotional tone. The pseudo-siamese cross-perspective submodule is embedded to interact with siamese fine-grained emotions across perspectives, further capturing subtle emotional fluctuations and amplifying the authenticity of facial movements. Furthermore, to comprehensively evaluate our work, the CREMA-D dataset is reconstructed into a 3D emotional talking face dataset with the largest number of subjects, serving as a valuable experimental supplement. Extensive experiments and user studies demonstrate that our approach outperforms state-of-the-art methods and exhibits superior emotional agility and expressiveness in facial movements. Zhaojie Chu, Yilin Lan, Jianxiu Jin, Xiangmin Xu 0001, Xiaofen Xing |
IEEE Trans. Affect. Comput. | 5 |
| 2026 | EETalk: Expression Enhancement in Speech-Driven 3D Facial AnimationabstractSpeech-driven 3D facial animation aims to generate natural and expressive facial movements from speech. Although significant progress has been made, existing methods still face challenges in generating realistic upper facial expressions. Specifically, existing methods that jointly optimize the holistic face tend to overlook fine-grained spatial movements due to motion differences across facial regions. Pre-trained speech feature extractors, which emphasize long-term dependencies, provide limited fine-grained temporal cues. In this work, we propose a novel framework, EETalk, to enhance the realism of facial expressions, which captures fine-grained spatial information and f ine-grained temporal dynamics from speech. To alleviate loss of fine-grained spatial information, we propose a novel disassemble and-reassemble modeling strategy. This strategy constructs two independent motion representation spaces for the upper and lower faces, allowing for the capture of weak upper-face movements while preserving motion diversity. Then, we propose a Cross-Region Coordination Module to ensure the synchronization and coordination of the movements from the independent upper and lower faces. To effectively capture facial micro-expressions, we incorporate fine-grained time-varying features to compensate for the short-timescale details underrepresented in the long term semantic features extracted by pre-trained self-supervised models, thereby improving the generation of fast, subtle facial expressions. Experimental results demonstrate that our approach significantly improves motion accuracy, expression consistency, and perceptual quality compared to existing methods. Zhaojie Chu, Kailing Guo, Xiaofen Xing, Bolun Cai, Lin Wang 0004, Xiangmin Xu 0001 |
IEEE Trans. Multim. | 3 |
| 2026 | Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head AnimationabstractSinging-driven 3D head animation is a compelling yet underexplored task with broad applications in virtual avatars, entertainment, and education. Existing speech-driven approaches, which typically map audio directly to motion through implicit phoneme-to-viseme correspondences, often yield over-smoothed, emotionally flat, and semantically inconsistent results. These limitations render them inadequate for the unique demands of singing-driven animation. To address this challenge, we propose Think2Sing, a unified diffusion-based framework that integrates pretrained large language models to generate semantically consistent and temporally coherent 3D head animations conditioned on both lyrics and acoustics. Central to our framework is the introduction of motion subtitles, a structured, time-aligned representation generated via a Singing Chain-of-Thought process with acoustic-guided retrieval. These subtitles provide region-specific expressive cues that serve as interpretable priors for animation synthesis. We further formulate head animation as motion intensity prediction over key facial regions, enabling fine-grained control and more faithful expressive modeling. To support this paradigm, we construct the first multimodal singing dataset with synchronized 3D motion, acoustic descriptors, and aligned motion subtitles, enabling semantically grounded and expressive motion learning. Extensive experiments demonstrate that Think2Sing significantly outperforms state-of-the-art methods in realism, expressiveness, and emotional fidelity. Furthermore, our framework supports flexible subtitle-conditioned editing, enabling precise and user-controllable animation synthesis. Zikai Huang, Xuemiao Xu, Xiaofen Xing, Harry Qin, Shengfeng He |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | PsyDT: Using LLMs to Construct the Digital Twin of Psychological Counselor with Personalized Counseling Style for Psychological CounselingabstractCurrently, large language models (LLMs) have made significant progress in the field of psychological counseling.However, existing mental health LLMs overlook a critical issue where they do not consider the fact that different psychological counselors exhibit different personal styles, including linguistic styles and therapeutic types, etc.As a result, these LLMs fail to satisfy the individual needs of clients who seek different counseling styles.To help bridge this gap, we propose PsyDT, a novel framework using LLMs to construct the Digital Twin of Psychological counselor with personalized counseling style.Compared to the timeconsuming and costly approach of collecting a large number of real-world counseling cases to create a specific counselor's digital twin, our framework offers a faster and more costeffective solution.To construct PsyDT, we utilize dynamic one-shot learning by using GPT-4 to capture counselor's unique counseling style, mainly focusing on linguistic style and therapy technique.Subsequently, using existing singleturn long-text dialogues with client personality, GPT-4 is guided to synthesize multi-turn dialogues of specific counselor.Finally, we finetune the LLMs on the synthesized dataset, Psy-DTCorpus, to achieve the digital twin of psychological counselor with personalized counseling style.Experimental results indicate that our proposed PsyDT framework can synthesize multi-turn dialogues that closely resemble realworld counseling cases and demonstrate better performance compared to other baselines, thereby show that our framework can effectively construct the digital twin of psychological counselor with a specific counseling style. 1 Haojie Xie, Yirong Chen, Xiaofen Xing, Jingkai Lin, Xiangmin Xu 0001 |
ACL (1) | 3 |
| 2025 | Intermediate-Selective Feature Enhancement for Speech Emotion RecognitionabstractSpeech Emotion Recognition (SER) plays a vital role in enabling intelligent systems to perceive and respond to human emotions. Recent advances in large pretrained models have led to substantial improvements in SER performance. However, most existing approaches focus solely on a single layer representation or perform layer-wise fusion, which may not be optimal for downstream emotion recognition tasks. In this paper, we propose ISFE (Intermediate-Selective Feature Enhancement), a novel framework designed to enhance the emotional expressiveness of pretrained representations. Rather than fusing features across different layers, ISFE targets and enhances specific emotion-related signals from intermediate layer that are often suppressed in deeper layers. Extensive experiments on the IEMOCAP and MELD datasets demonstrate that ISFE achieves better performance than previous state-of-the-art methods, validating the effectiveness of enhancing intermediate-layer features for SER. Yangbiao Li, Xiaofen Xing, Jialong Mai, Jingyuan Xing, Xiangmin Xu 0001 |
ASRU | 2 |
| 2025 | SAKI-RAG: Mitigating Context Fragmentation in Long-Document RAG via Sentence-level Attention Knowledge IntegrationabstractTraditional Retrieval-Augmented Generation (RAG) frameworks often segment documents into larger chunks to preserve contextual coherence, inadvertently introducing redundant noise.Recent advanced RAG frameworks have shifted toward finer-grained chunking to improve precision.However, in long-document scenarios, such chunking methods lead to fragmented contexts, isolated chunk semantics, and broken inter-chunk relationships, making cross-paragraph retrieval particularly challenging.To address this challenge, maintaining granular chunks while recovering their intrinsic semantic connections, we propose SAKI-RAG (Sentence-level Attention Knowledge Integration Retrieval-Augmented Generation).Our framework introduces two core components: (1) the SentenceAttnLinker, which constructs a semantically enriched knowledge repository by modeling inter-sentence attention relationships, and (2) the Dual-Axis Retriever, which is designed to expand and filter the candidate chunks from the dual dimensions of semantic similarity and contextual relevance.Experimental results across four datasets-Dragonball, SQUAD, NFCORPUS, and SCI-DOCS demonstrate that SAKI-RAG achieves better recall and precision compared to other RAG frameworks in long-document retrieval scenarios, while also exhibiting higher information efficiency. Wenyu Tao, Xiaofen Xing, Zeliang Li, Xiangmin Xu 0001 |
EMNLP | 2 |
| 2025 | MEIJU - The 1st Multimodal Emotion and Intent Joint Understanding ChallengeabstractMultimodal Emotion and Intent Joint Understanding (MEIJU) aims to decode the semantic information expressed in the multimodal dialogues while inferring the emotions and intents, providing users with a more humanized human-machine interaction experience. However, challenges such as difficulties in data acquisition and imbalance annotations have made it difficult for current methods to meet the demands of practical applications. Therefore, we have organized two tracks focusing on the key themes of semi-supervised learning and class imbalance. Additionally, we have prepared data in two different languages (English and Mandarin) for each track, treating each language as a sub-track, to encourage participants to explore solutions in more diverse linguistic environments. Our code can be found at https://github.com/AI-S2-Lab/MEIJU2025-baseline. Rui Liu 0008, Xiaofen Xing, Zheng Lian 0004, Haizhou Li 0001, Björn W. Schuller, Haolin Zuo |
ICASSP | 2 |
| 2025 | SoCov: Semi-Orthogonal Parametric Pooling of Covariance Matrix for Speaker RecognitionabstractIn conventional deep speaker embedding frameworks, the pooling layer aggregates all frame-level features over time and computes their mean and standard deviation statistics as inputs to subsequent segment-level layers. Such statistics pooling strategy produces fixed-length representations from variable-length speech segments. However, this method treats different frame-level features equally and discards covariance information. In this paper, we propose the Semi-orthogonal parameter pooling of Covariance matrix (SoCov) method. The SoCov pooling computes the covariance matrix from the self-attentive frame-level features and compresses it into a vector using the semi-orthogonal parametric vectorization, which is then concatenated with the weighted standard deviation vector to form inputs to the segment-level layers. Deep embedding based on SoCov is called "sc-vector". The proposed sc-vector is compared to several different baselines on the SRE21 development and evaluation sets. The sc-vector system significantly outperforms the conventional x-vector system, with a relative reduction in EER of 15.5% on SRE21Eval. When using self-attentive deep feature, SoCov helps to reduce EER on SRE21Eval by about 30.9% relatively to the conventional "mean + standard deviation" statistics. Rongjin Li, Dongpeng Chen, Jintao Kang, Xiaofen Xing |
ICASSP | 5 |
| 2025 | DecoupledSynth: Enhancing Zero-Shot Text-to-Speech Via Factors DecouplingabstractStudies of speech representation enhance zero-shot Text-to-Speech by mapping text to intermediate representations before generating speech. However, using representations often struggles to balance linguistic, para-linguistic, and non-linguistic information in speech during the synthesis phase. Additionally, it usually takes substantial resources for representation extraction training. To address these limitations, we propose DecoupledSynth. It combines different self-supervised models to extract comprehensive, decoupled representations. This structure enables more thorough and nuanced synthesis by leveraging reference speech with decoupled processing stages. Experiments on the VCTK and LibriTTS datasets support the potential of this new framework, showing that it can produce more consistent and realistic speech. Speech demos are available at https://test1634.github.io/DecoupledSynth/. Jingyuan Xing, Shuaiqi Chen, Xiangmin Xu 0001, Xiaofen Xing |
ICASSP | 4 |
| 2025 | Drawing Developmental Trajectory From Cortical Surface Reconstruction
Ruowen Qu, Zhongliang Liu, Zhuoyan Dai, Dongzi Shi, Sijin Yu, Tong Xiong, Shiping Liu, Xiangmin Xu 0001, Xiaofen Xing, Xin Zhang 0013 |
ICCV | 10 |
| 2025 | Enhancing User-Oriented Proactivity in Open-Domain Dialogues with Critic GuidanceabstractOpen-domain dialogue systems aim to generate natural and engaging conversations, providing significant practical value in real applications such as social robotics and personal assistants. The advent of large language models (LLMs) has greatly advanced this field by improving context understanding and conversational fluency. However, existing LLM-based dialogue systems often fall short in proactively understanding the user's chatting preferences and guiding conversations toward user-centered topics. This lack of user-oriented proactivity can lead users to feel unappreciated, reducing their satisfaction and willingness to continue the conversation in human-computer interactions. To address this issue, we propose a User-oriented Proactive Chatbot (UPC) to enhance the user-oriented proactivity. Specifically, we first construct a critic to evaluate this proactivity inspired by the LLM-as-a-judge strategy. Given the scarcity of high-quality training data, we then employ the critic to guide dialogues between the chatbot and user agents, generating a corpus with enhanced user-oriented proactivity. To ensure the diversity of the user backgrounds, we introduce the ISCO-800, a diverse user background dataset for constructing user agents. Moreover, considering the communication difficulty varies among users, we propose an iterative curriculum learning method that trains the chatbot from easy-to-communicate users to more challenging ones, thereby gradually enhancing its performance. Experiments demonstrate that our proposed training method is applicable to different LLMs, improving user-oriented proactivity and attractiveness in open-domain dialogues. Code and appendix are available at github.com/wang678/LLM-UPC. Yufeng Wang 0004, Jinwu Hu, Ziteng Huang, Kunyang Lin, Zitian Zhang, Peihao Chen, Yu Hu 0004, Qianyue Wang, Zhu Liang Yu, Bin Sun 0001, Xiaofen Xing, Mingkui Tan |
IJCAI | 11 |
| 2025 | MMLoRA: Multitask Memory Parameter-Efficient Fine-Tuning for Multimodal SER
Yuanbo Fang, Xiaofen Xing, Xueru Li, Xiangmin Xu 0001 |
INTERSPEECH | 2 |
| 2025 | Long-Context Speech Synthesis with Context-Aware Memory
Xiaofen Xing, Jingyuan Xing, Hangrui Hu, Xiangmin Xu 0001 |
INTERSPEECH | 2 |
| 2025 | SA-RAS: Speaker-Aware Style Retrieval Augmented Generation for Expressive Zero-Shot Text-to-Speech Synthesis
Xueru Li, Jingyuan Xing, Xiaofen Xing, Xiangmin Xu 0001 |
INTERSPEECH | 3 |
| 2025 | AA-SLLM: An Acoustically Augmented Speech Large Language Model for Speech Emotion Recognition
Jialong Mai, Xiaofen Xing, Yuanbo Fang, Xiangmin Xu 0001 |
INTERSPEECH | 2 |
| 2025 | Chain-of-Thought Distillation with Fine-Grained Acoustic Cues for Speech Emotion Recognition
Jialong Mai, Xiaofen Xing, Yangbiao Li, Xiangmin Xu 0001 |
INTERSPEECH | 2 |
| 2025 | EATS-Speech: Emotion-Adaptive Transformation and Priority Synthesis for Zero-Shot Text-to-Speech
Jingyuan Xing, Shuaiqi Chen, Xiaofen Xing, Xiangmin Xu 0001 |
INTERSPEECH | 4 |
| 2025 | Heterogeneous Masked Attention-Guided Path Convolution for Functional Brain Network Analysis
Jiakun Xu, Xin Zhang 0013, Tong Xiong, Shengxian Chen, Xiaofen Xing, Jindou Hao, Xiangmin Xu 0001 |
MICCAI (12) | 5 |
| 2025 | SemGesture: Synthesizing Semantically Enhanced and Coherent GesturesabstractCo-speech gestures are generally categorized into rhythmic and semantic gestures: rhythmic gestures align with speech rhythm and intonation, while semantic gestures convey specific meanings or emotions, enriching verbal communication. Most previous studies have focused on synthesizing rhythmic gestures, while recent methods have explored integrating large language models (LLMs) to retrieve semantic gestures and merge them with rhythmic ones. However, existing approaches primarily rely on textual context for retrieval, which may not fully capture the emotional and tonal nuances of speech, sometimes leading to semantic gestures that do not align with the speaker's intended expression. Additionally, common gesture fusion techniques often merge rhythmic and semantic gestures directly, causing discontinuity due to differences in their movement styles. To address these challenges, we propose SemGesture, a system designed to generate smooth and semantically accurate gestures. Our approach incorporates Context-aware Retrieval powered by a Large Audio Language Model, enabling precise retrieval of gestures that align with both the semantic and emotional aspects of speech. Additionally, the Gesture Fusion Module dynamically adjusts semantic gestures to harmonize with rhythmic gestures, ensuring seamless and coherent motion transitions. Extensive experiments demonstrate that SemGesture significantly outperforms existing methods in generating contextually accurate and visually natural gestures. Pengsheng Liu, Zhaojie Chu, Xiaofen Xing, Xiangmin Xu 0001 |
ACM Multimedia | 3 |
| 2025 | Multimodal speech emotion recognition via dynamic multilevel contrastive loss under local enhancement network
Weiquan Fan, Xiangmin Xu 0001, Fang Liu 0030, Xiaofen Xing |
Expert Syst. Appl. | 4 |
| 2025 | Low-complexity speaker embedding module with feature segmentation, transformation and reconstruction for few-shot speaker identification
Yanxiong Li, Qisheng Huang, Xiaofen Xing, Xiangmin Xu 0001 |
Expert Syst. Appl. | 3 |
| 2025 | rPPG-TFCL: Time-frequency consistency learning for robust remote physiological measurement
Kailing Guo, Fang Liu 0030, Xiaofen Xing, Lin Wang 0004, Xiangmin Xu 0001, Zhanpeng Jin |
Knowl. Based Syst. | 4 |
| 2025 | Towards zero-shot human-object interaction detection via vision-language integration
Weiying Xue, Qi Liu 0005, Yuxiao Wang 0003, Zhenao Wei, Xiaofen Xing, Xiangmin Xu 0001 |
Neural Networks | 5 |
| 2025 | Coordination Attention based Transformers with bidirectional contrastive loss for multimodal speech emotion recognition
Weiquan Fan, Xiangmin Xu 0001, Guohua Zhou, Xiaofang Deng, Xiaofen Xing |
Speech Commun. | 5 |
| 2025 | Individual-Aware Attention Modulation for Unseen Speaker Emotion RecognitionabstractIn practical human-computer interaction (HCI) applications, robust speech emotion recognition (SER) for unseen speakers is crucial. Prior research has primarily focused on extracting common representations to enhance the generalization of cross-individual SER. However, most methods ignore the positive effects of individual characteristics. Actually, each speaker can be regarded as an independent individual domain. Personalized SER can be improved if the emotional expressions of individual speech characteristics are effectively utilized. To address the challenges in recognizing emotions for unseen speakers, this paper proposes a novel individual-aware attention modulation (IAM) model. Specifically, the IAM uses meta-learning techniques to extract modulation parameters for obtaining individual-related emotion expressions from individual characteristics. The base model is then modulated to facilitate the transfer of the common emotion representation space to an individual-specific emotion representation space. This transformation is achieved by applying attention modulation within the transformer-based model developed in this paper. In addition, we employ a meta-learning-based method to optimize model parameters, enhancing the adaptability of the model to unseen speakers, and a control factor is introduced to regulate the degree of individual modulation, thus enhancing the robustness of the modulation process. Experimental results demonstrate that the proposed model achieves significantly improved cross-individual SER performance. Yuanbo Fang, Xiaofen Xing, Zhaojie Chu, Yifeng Du, Xiangmin Xu 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | CSE-GResNet: A Simple and Highly Efficient Network for Facial Expression RecognitionabstractFacial expression recognition (FER) has recently attracted extensive attention in computer vision. However, existing methods mostly focus on the explicit performance and overlook their computational resources. Hence, achieving competitive performance while maintaining the model efficiency is still a huge challenge. To tackle these issues, we propose a highly lightweight yet effective Channel Shift-Enhancement Gabor-ResNet (CSE-GResNet) to capture the crucial visual properties in facial images. Concretely, we incorporate the Gabor Convolution (GConv) into ResNet to produce the robust GResNet as our backbone with limited memory cost. Furthermore, we propose extremely efficient Channel-Shift Module and Channel-Enhancement Module to insert in the GResNet in cascade. They are adopted to obtain and aggregate the facial informative representation from adjacent channels for extracting the subtle facial expression representation. We conduct extensive experiments on three wild datasets: RAF-DB, FER2013 and SFEW. The results show that the proposed CSE-GResNet achieves superior performance against the state-of-the-art methods with less computational and memory cost. Shaoping Jiang, Xiaofen Xing, Fang Liu 0030, Xiangmin Xu 0001, Lin Wang 0004, Kailing Guo |
IEEE Trans. Affect. Comput. | 2 |
| 2025 | Alleviating One-to-Many Mapping in Talking Head Synthesis With Dynamic Adaptation Context and Style AdapterabstractSpeech-driven talking head synthesis technology has made remarkable progress, but it still faces the challenge of one-to-many pathological mapping. The challenge results in inaccurate lip movements, ambiguity in facial expressions, and a lack of coherence during transitions between facial motions. The phenomenon is primarily caused by: (1) for one speaker, the same phoneme corresponds to a wide range of mouth shapes and facial expressions due to contextual variations, and (2) for the same spoken content, different speakers exhibit diverse facial motions as a result of unique speaking styles. In this work, we propose a novel framework, called AllTalk, to alleviate one-to-many pathological mapping, which enables a more vivid and natural talking head. Specifically, considering the asymmetry and dynamic nature of mouth shapes’ dependence on phoneme context, we propose a Dynamic Adaptive Context encoder to capture the context around the phoneme and its dynamics, thereby reducing the ambiguity in mapping speech to facial movements. Moreover, to alleviate the uncertainty caused by differences in speaking style, we propose a Style Adapter that expands a generic discrete motion space for the target speaker. The Style Adapter not only effectively represents general facial motions but also captures the personalized nuances of facial movements. To further enhance the fidelity of output, we introduce a Dynamic Gaussian Renderer based on 3D Gaussian Splatting, capable of producing stable and realistic rendering videos. Extensive qualitative and quantitative experiments demonstrate that AllTalk surpasses existing state-of-the-art methods, providing an effective solution to the challenge of one-to-many mapping. Project page: https://zjchu.github.io/projects/AllTalk. Zhaojie Chu, Kailing Guo, Xiaofen Xing, Bolun Cai, Xiangmin Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | CLIP-Vision Guided Few-Shot Metal Surface Defect RecognitionabstractMetal surface defect recognition (MSDR) based on deep learning encounters the challenge of few-shot expert-labeled data. In this study, we proposed a CLIP-vision guided self supervised learning (CVGSSL) framework for representation learning of unlabeled data, completing MSDR using few-shot labeled data. This framework initially generates rich and diverse representation information through multiple CLIP-Vs to ensure effective SSL pretraining, followed by the design of an MLP-adapter to distill knowledge and adapt these representations to recognition tasks. In addition, we constructed a self-constrained loss to address the inherent problem of intraclass and interclass distance ambiguity that causes the representation to fall into an equivocal decision margin. Following label-free pretraining of CVGSSL, the downstream model adapts to one-shot to four-shot defect recognition tasks through fine-tuning. Experimental results demonstrate that CVGSSL outperforms state-of-the-art SSL methods across three public metal surface defect datasets, with the efficacy of the approach validated through extensive ablation experiments. Tianlei Wang, Zeliang Li, Ying Xu 0005, Yikui Zhai, Xiaofen Xing, Kailing Guo, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
IEEE Trans. Ind. Informatics | 5 |
| 2025 | DCPTalk: Speech-Driven 3D Face Animation With Personalized Facial Dynamic Coupling PropertiesabstractSpeech-driven 3D facial animation has emerged as a hot topic. During this process, movements in different facial regions are interdependent, influenced by the intricate interactions among facial muscles, and manifest personalized differences. The existing methods typically simplify the facial animation generation task to an infinitely thin surface skin deformation without an underlying structure, thereby ignoring the intricate and personalized dynamics of facial muscle activity. These methods tend to produce static or weak upper-face animations with an average facial movement style. In this work, we propose a novel framework, called DCPTalk, to mimic the intricate dynamics of facial muscle activity and portray personalized facial animations. Based on facial dynamic coupling properties, we propose Mouth2Face to simulate the facial muscle control system, yielding realistic and coordinated facial animations evoked by mouth movements. Mouth movements are easily synthesized from speech signals due to their direct correlation with phonetic articulation and vocal tract dynamics. To further enhance the detail of facial movements, we employ surface skin deformation to refine the facial animation derived from Mouth2Face. Furthermore, personal factors, including inherent physical traits and acquired speaking styles, directly determine the uniqueness and realism of facial animations. Inherent physical traits are embedded into Mouth2Face for constructing personalized facial muscle control system, while acquired speaking styles are employed to modulate external driving signals. Extensive qualitative and quantitative experiments as well as a user study indicate that DCPTalk outperforms the existing state-of-the-art methods. Zhaojie Chu, Kailing Guo, Xiaofen Xing, Pengsheng Liu, Bolun Cai, Xiangmin Xu 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Compact Model Training by Low-Rank Projection With Energy TransferabstractLow-rankness plays an important role in traditional machine learning but is not so popular in deep learning. Most previous low-rank network compression methods compress networks by approximating pretrained models and retraining. However, the optimal solution in the Euclidean space may be quite different from the one with low-rank constraint. A well-pretrained model is not a good initialization for the model with low-rank constraints. Thus, the performance of a low-rank compressed network degrades significantly. Compared with other network compression methods such as pruning, low-rank methods attract less attention in recent years. In this article, we devise a new training method, low-rank projection with energy transfer (LRPET), that trains low-rank compressed networks from scratch and achieves competitive performance. We propose to alternately perform stochastic gradient descent training and projection of each weight matrix onto the corresponding low-rank manifold. Compared to retraining on the compact model, this enables full utilization of model capacity since solution space is relaxed back to Euclidean space after projection. The matrix energy (the sum of squares of singular values) reduction caused by projection is compensated by energy transfer. We uniformly transfer the energy of the pruned singular values to the remaining ones. We theoretically show that energy transfer eases the trend of gradient vanishing caused by projection. In modern networks, a batch normalization (BN) layer can be merged into the previous convolution layer for inference, thereby influencing the optimal low-rank approximation (LRA) of the previous layer. We propose BN rectification to cut off its effect on the optimal LRA, which further improves the performance. Comprehensive experiments on CIFAR-10 and ImageNet have justified that our method is superior to other low-rank compression methods and also outperforms recent state-of-the-art pruning methods. For object detection and semantic segmentation, our method still achieves good compression results. In addition, we combine LRPET with quantization and hashing methods and achieve even better compression than the original single method. We further apply it in Transformer-based models to demonstrate its transferability. Our code is available at https://github.com/BZQLin/LRPET. Kailing Guo, Zhenquan Lin, Canyang Chen, Xiaofen Xing, Fang Liu 0030, Xiangmin Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Clinical Scores Prediction and Medication Adjustment for Course of Parkinson's DiseaseabstractParkinson's Disease (PD) is the second most prevalent neurodegenerative disorder worldwide, characterized by progressive motor and non-motor symptoms. Unfortunately, there are no definitive PD modifying therapies, so accurate course prediction in advance and appropriate medical adjustment are essential to slow down degenerative process from onset. This work addresses a novel challenge in PD course prediction specifically at month 60 (m60) of both motor and non-motor indicator utilizing Magnetic Resonance Imaging (MRI) and demographic data of previous years. A medication adjustment network based on Reinforcement Learning (RL) is utilized as an agent to simulate medication from professionals in a prediction environment. The proposed approach achieve a more accurate prediction on motor and non-motor simultaneously, demonstrating significant promise for longer-term PD course prediction compared to existing works. Furthermore, adjustable medication branch shows consistent with our advance result and provide possible guidance on medication for healthcare practitioners. Xiaofen Xing, Xiangmin Xu 0001 |
ICASSP | 3 |
| 2024 | Modality-Guided Collaborative Filtering for Recommendation
Kai Zhang 0078, Linping Gao, Jinda Lu, Yutong Yuan, Xiaofen Xing |
ICIC (13) | 5 |
| 2024 | Reinforce Tokens for the Next Recommendation Generation
Kai Zhang 0078, Linping Gao, Zhaobao Yang, Xiaofen Xing |
ICIC (13) | 5 |
| 2024 | DropFormer: A Dynamic Noise-Dropping Transformer for Speech Emotion Recognition
Jialong Mai, Xiaofen Xing, Xiangmin Xu 0001 |
INTERSPEECH | 2 |
| 2024 | Cortical Surface Reconstruction from 2D MRI with Segmentation-Constrained Super-Resolution and Representation Learning
Ruowen Qu, Dongzi Shi, Tong Xiong, Xiangmin Xu 0001, Xiaofen Xing, Xin Zhang 0013 |
MICCAI (2) | 6 |
| 2024 | RetrievalMMT: Retrieval-Constrained Multi-Modal Prompt Learning for Multi-Modal Machine TranslationabstractAs an extension of machine translation, the primary objective of multi-modal machine translation is to optimize the utilization of visual information. Technically, image information is integrated into multi-modal fusion and alignment as an auxiliary modality through concepts or latent semantics, which are typically based on the Transformer framework. However, current approaches often ignore one modality to design numerous handcrafted features (e.g. visual concept extraction) and require training of all parameters in their framework. Therefore, it is worthwhile to explore multi-modal concepts or features to enhance performance and an efficient approach to incorporate visual information with minimal cost. Meanwhile, with the development of multi-modal large language models (MLLMs), they are faced with the visual hallucination issue of compromising performance, despite their powerful capabilities. Inspired by pioneering techniques in the multi-modal field, such as prompt learning and MLLMs, this paper innovatively explores the possibility of applying multi-modal prompt learning to this multi-modal machine translation task. Yan Wang 0140, Yawen Zeng, Xiaofen Xing, Jin Xu 0014, Xiangmin Xu 0001 |
ICMR | 4 |
| 2024 | As-Speech: Adaptive Style For Speech SynthesisabstractIn recent years, there has been significant progress in Text-to-Speech (TTS) synthesis technology, enabling the high-quality synthesis of voices in common scenarios. In unseen situations, adaptive TTS requires a strong generalization capability for speaker style characteristics. However, the existing adaptive methods can only extract and integrate coarse-grained timbre or mixed rhythm attributes separately. In this paper, we propose AS-Speech, an adaptive style methodology that integrates the speaker timbre characteristics and rhythmic attributes into a unified framework for text-to-speech synthesis. Specifically, AS-Speech can accurately simulate style characteristics through fine-grained text-based timbre features and global rhythm information, and achieve high-fidelity speech synthesis through the diffusion model. Experiments show that our proposed model produces voices with higher similarity in terms of timbre and rhythm compared to a series of adaptive TTS models while maintaining the naturalness of synthetic speech. Samples are available at https://leezp99.github.io/as-speech-demo/ Xiaofen Xing, Shuaiqi Chen, Guoqiao Yu, Guanglu Wan, Xiangmin Xu 0001 |
SLT | 2 |
| 2024 | Domain Adaption and Unified Knowledge Base Motivate Better Retrieval Models in Dialog Systems With RAGabstractRetrieval augmented generation (RAG) has emerged as a paradigm to address problems like hallucination in dialog systems based on large language model (LLM). Retrieval model is a key component in RAG framework for recalling relevant information. This paper describes our solution for FutureDial-RAG Challenge Track 1. We identify two primary challenges in this track: domain specificity and heterogeneity of knowledge base. To address the two challenges, we first adopt continual pre-training of a pre-trained retrieval model on both labeled and unlabeled data for domain adaption. Subsequently, we modify and expand the knowledge base, ensuring that each piece of knowledge is uniformly structured in a question-answer (QA) format. Finally, we construct negative samples based on the labeled data and the unified knowledge base, and fine-tune the retrieval model using contrastive learning. Our solution achieves a score of 2.023 on the dev set, which significantly outperforms the baseline. Huadong Lin, Yirong Chen, Wenyu Tao, Mingyu Chen 0016, Xiangmin Xu 0001, Xiaofen Xing |
SLT | 6 |
| 2024 | Style-conditioned music generation with Transformer-GANsabstractRecently, various algorithms have been developed for generating appealing music. However, the style control in the generation process has been somewhat overlooked. Music style refers to the representative and unique appearance presented by a musical work, and it is one of the most salient qualities of music. In this paper, we propose an innovative music generation algorithm capable of creating a complete musical composition from scratch based on a specified target style. A style-conditioned linear Transformer and a style-conditioned patch discriminator are introduced in the model. The style-conditioned linear Transformer models musical instrument digital interface (MIDI) event sequences and emphasizes the role of style information. Simultaneously, the style-conditioned patch discriminator applies an adversarial learning mechanism with two innovative loss functions to enhance the modeling of music sequences. Moreover, we establish a discriminative metric for the first time, enabling the evaluation of the generated music’s consistency concerning music styles. Both objective and subjective evaluations of our experimental results indicate that our method’s performance with regard to music production is better than the performances encountered in the case of music production with the use of state-of-the-art methods in available public datasets. Weining Wang 0003, Xiaofen Xing |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2024 | Vesper: A Compact and Effective Pretrained Model for Speech Emotion RecognitionabstractThis paper presents a paradigm that adapts general large-scale pretrained models (PTMs) to speech emotion recognition task. Although PTMs shed new light on artificial general intelligence, they are constructed with general tasks in mind, and thus, their efficacy for specific tasks can be further improved. Additionally, employing PTMs in practical applications can be challenging due to their considerable size. Above limitations spawn another research direction, namely, optimizing large-scale PTMs for specific tasks to generate task-specific PTMs that are both compact and effective. In this paper, we focus on the speech emotion recognition task and propose an improVedemotion-specificpretrained encodercalled Vesper. Vesper is pretrained on a speech dataset based on WavLM and takes into account emotional characteristics. To enhance sensitivity to emotional information, Vesper employs an emotion-guided masking strategy to identify the regions that need masking. Subsequently, Vesper employs hierarchical and cross-layer self-supervision to improve its ability to capture acoustic and semantic representations, both of which are crucial for emotion recognition. Experimental results on the IEMOCAP, MELD, and CREMA-D datasets demonstrate that Vesper with 4 layers outperforms WavLM Base with 12 layers, and the performance of Vesper with 12 layers surpasses that of WavLM Large with 24 layers. Xiaofen Xing, Peihao Chen, Xiangmin Xu 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | Bridge Graph Attention Based Graph Convolution Network With Multi-Scale Transformer for EEG Emotion RecognitionabstractIn multichannel electroencephalograph (EEG) emotion recognition, most graph-based studies employ shallow graph model for spatial characteristics learning due to node over-smoothing caused by an increase in network depth. To address over-smoothing, we propose the bridge graph attention-based graph convolution network (BGAGCN). It bridges previous graph convolution layers to attention coefficients of the final layer by adaptively combining each graph convolution output based on the graph attention network, thereby enhancing feature distinctiveness. Considering that graph-based networks primarily focus on local EEG channel relationships, we introduce a transformer for global dependency. Inspired by the neuroscience finding that neural activities of different timescales reflect distinct spatial connectivities, we modify the transformer to a multi-scale transformer (MT) by applying multi-head attention to multichannel EEG signals after 1D convolutions at different scales. MT learns spatial features more elaborately to enhance feature representation ability. By combining BGAGCN and MT, our model BGAGCN-MT achieves state-of-the-art accuracy under subject-dependent and subject-independent protocols across three benchmark EEG emotion datasets (SEED, SEED-IV and DREAMER). Notably, our model effectively addresses over-smoothing in graph neural networks and provides an efficient solution to learning spatial relationships of EEG features at different scales. Our code is available athttps://github.com/LogzZ. Huachao Yan, Kailing Guo, Xiaofen Xing, Xiangmin Xu 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2024 | CorrTalk: Correlation Between Hierarchical Speech and Facial Activity Variances for 3D AnimationabstractSpeech-driven 3D facial animation is a challenging cross-modal task that has attracted growing research interest. During speaking activities, the mouth displays strong motions, while the other facial regions typically demonstrate comparatively weak activity levels. Existing approaches often simplify the process by directly mapping single-level speech features to the entire facial animation, which overlook the differences in facial activity intensity leading to overly smoothed facial movements. In this study, we propose a novel framework, CorrTalk, which effectively establishes the temporal correlation between hierarchical speech features and facial activities of different intensities across distinct regions. A novel facial activity intensity prior is defined to distinguish between strong and weak facial activity, obtained by statistically analyzing facial animations. Based on the facial activity intensity prior, we propose a dual-branch decoding framework to synchronously synthesize strong and weak facial activity, which guarantees wider intensity facial animation synthesis. Furthermore, a weighted hierarchical feature encoder is proposed to establish temporal correlation between hierarchical speech features and facial activity at different intensities, which ensures lip-sync and plausible facial expressions. Extensive qualitatively and quantitatively experiments as well as a user study indicate that our CorrTalk outperforms existing state-of-the-art methods. The source code and supplementary video are publicly available at: https://zjchu.github.io/projects/CorrTalk/. Zhaojie Chu, Kailing Guo, Xiaofen Xing, Yilin Lan, Bolun Cai, Xiangmin Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | DST: Deformable Speech Transformer for Emotion RecognitionabstractEnabled by multi-head self-attention, Transformer has exhibited remarkable results in speech emotion recognition (SER). Compared to the original full attention mechanism, window-based attention is more effective in learning fine-grained features while greatly reducing model redundancy. However, emotional cues are present in a multi-granularity manner such that the pre-defined fixed window can severely degrade the model flexibility. In addition, it is difficult to obtain the optimal window settings manually. In this paper, we propose a Deformable Speech Transformer, named DST, for SER task. DST determines the usage of window sizes conditioned on in-put speech via a light-weight decision network. Meanwhile, data-dependent offsets derived from acoustic features are utilized to adjust the positions of the attention windows, allowing DST to adaptively discover and attend to the valuable in-formation embedded in the speech. Extensive experiments on IEMOCAP and MELD demonstrate the superiority of DST. Xiaofen Xing, Xiangmin Xu 0001, Jianxin Pang |
ICASSP | 2 |
| 2023 | DWFormer: Dynamic Window Transformer for Speech Emotion RecognitionabstractSpeech emotion recognition is crucial to human-computer interaction. The temporal regions that represent different emotions scatter in different parts of the speech locally. Moreover, the temporal scales of important information may vary over a large range within and across speech segments. Although transformer-based models have made progress in this field, the existing models could not precisely locate important regions at different temporal scales. To address the issue, we propose Dynamic Window transFormer (DWFormer), a new architecture that leverages temporal importance by dynamically splitting samples into windows. Self-attention mechanism is applied within windows for capturing temporal important information locally in a fine-grained way. Cross-window information interaction is also taken into account for global communication. DWFormer is evaluated on both the IEMO-CAP and the MELD datasets. Experimental results show that the proposed model achieves better performance than the previous state-of-the-art methods. Shuaiqi Chen, Xiaofen Xing, Xiangmin Xu 0001 |
ICASSP | 2 |
| 2023 | MGAT: Multi-Granularity Attention Based Transformers for Multi-Modal Emotion RecognitionabstractMulti-modal emotion recognition is crucial for human-computer interaction. Many existing algorithms attempt to achieve multi-modal interactions through a cross-attention mechanism. Due to the problems of noise introduction and heavy computation in the original attention mechanism, window attention has become a new trend. However, emotions are presented asynchronously between different modalities, which makes it difficult to interact with emotional information between windows. Furthermore, multi-modal data are temporally misaligned, so single fixed window size is hard to describe cross-modal information. In this paper, we put these two issues into a unified framework and propose the multi-granularity attention based Transformers (MGAT). It addresses the emotional asynchrony and modality misalignment issues through a multi-granularity attention mechanism. Experimental results confirm the effectiveness of our method and the state-of-the-art performance is achieved on IEMOCAP. Weiquan Fan, Xiaofen Xing, Bolun Cai, Xiangmin Xu 0001 |
ICASSP | 2 |
| 2023 | Speaker-Aware Hierarchical Transformer For Personality Recognition In Multiparty DialoguesabstractPersonality recognition is one of the core technologies in human-machine interaction, which has received increasing attention. Previous works mainly focus on essays or monologues, while personality traits reveal more in the interactions with others. Due to the lack of appropriate datasets, a few approaches aim to recognize personality traits in conversations, and most of them ignore interdependence between speakers and connection between conversations. In this paper, we create a multiparty dialogue-based personality dataset derived from CPED containing 1,195 data samples. We center on one speaker and extract related dialogues to compose each data sample annotated with speaker’s Big-Five traits, which is conducive to fully describe a center speaker using diverse cues of personality in different dialogues. Along the same lines, we propose a Speaker-aware Hierarchical Transformer named SH-Transformer to address above concerns, in which Personalized Embeddings (PE) adopt special tokens to distinguish center speakers in complete conversations and hierarchical Transformer capture diverse cues in utterances and conversations. Experimental results show that our method outperforms the non-interactive baseline by 1.38%, which confirms the necessity of considering both interactive information and diverse cues among dialogues. Our code will be released at github.com/Chloehxxx/SH-Transformer. Wenjing Han, Yirong Chen, Xiaofen Xing, Guohua Zhou, Xiangmin Xu 0001 |
ICASSP | 3 |
| 2023 | Cross Range Quantization for Network CompressionabstractQuantization is effective in reducing model memory and accelerating inference, and is an important way to deploy deep neural networks on mobile smart devices. However, current popular learnable quantization functions often take simple truncation operations for values beyond the quantization range. We find that the truncation operation cause information loss and and restricts the update of values out of quantization range. To address this problem, we propose a universal cross range quantization (CRQ) method to reduce the information loss caused by the conventional truncation operation. CRQ splits the values exceeding the quantization range into two parts for separate quantization, and thus retain the information efficiently for performance improvement. In addition, we define a new metric named performance improvement efficiency (PIE) to measure the relationship between increased computation and performance improvement. Experiments on public benchmark image classification datasets show that CRQ achieves a significant accuracy gain with only a small increase in computation compared to the original learnable quantization method, and also outperforms many sophisticated designed state-of-the-art quantization methods in terms of accuracy and PIE. Yicai Yang, Xiaofen Xing, Kailing Guo, Xiangmin Xu 0001, Fang Liu 0030 |
IJCNN | 2 |
| 2023 | Exploring Downstream Transfer of Self-Supervised Features for Speech Emotion Recognition
Yuanbo Fang, Xiaofen Xing, Xiangmin Xu 0001 |
INTERSPEECH | 2 |
| 2023 | Multi-Scale Temporal Transformer For Speech Emotion Recognition
Xiaofen Xing, Yuanbo Fang, Hengsheng Fan, Xiangmin Xu 0001 |
INTERSPEECH | 2 |
| 2023 | SpeechFormer++: A Hierarchical Efficient Framework for Paralinguistic Speech ProcessingabstractParalinguistic speech processing is important in addressing many issues, such as sentiment and neurocognitive disorder analyses. Recently, Transformer has achieved remarkable success in the natural language processing field and has demonstrated its adaptation to speech. However, previous works on Transformer in the speech field have not incorporated the properties of speech, leaving the full potential of Transformer unexplored. In this paper, we consider the characteristics of speech and propose a general structure-based framework, called SpeechFormer++, for paralinguistic speech processing. More concretely, following the component relationship in the speech signal, we design a unit encoder to model the intra- and inter-unit information (i.e., frames, phones, and words) efficiently. According to the hierarchical relationship, we utilize merging blocks to generate features at different granularities, which is consistent with the structural pattern in the speech signal. Moreover, a word encoder is introduced to integrate word-grained features into each unit encoder, which effectively balances fine-grained and coarse-grained information. SpeechFormer++ is evaluated on the speech emotion recognition (IEMOCAP & MELD), depression classification (DAIC-WOZ) and Alzheimer's disease detection (Pitt) tasks. The results show that SpeechFormer++ outperforms the standard Transformer while greatly reducing the computational cost. Furthermore, it delivers superior results compared to the state-of-the-art approaches. Xiaofen Xing, Xiangmin Xu 0001, Jianxin Pang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | CS-GResNet: A Simple and Highly Efficient Network for Facial Expression RecognitionabstractFacial expression recognition (FER) has recently attracted attention in computer vision. However, existing methods mostly focus on the explicit performance and overlook their computational resources and memory consumption. Hence, achieving promising performance while maintaining the efficiency of models is still a huge challenge. In this work, we propose a highly efficient Channel-Shift Gabor-ResNet (CS-GResNet) to capture the crucial visual properties in facial images. Concretely, we incorporate the Gabor Convolution (GConv) into ResNet to produce the significant GResNet as our backbone with limited memory cost. Furthermore, we adopt an extremely simple yet effective Channel-Shift Module inserted into the GResNet to obtain the facial informative representation via facilitating information exchanged among neighboring channels. We conduct extensive experiments on three wild datasets: RAF-DB, FER2013 and SFEW. The results show that our proposed CS-GResNet achieves superior performance against the state-of-the-art methods with less computational and memory cost. Codes are available at https://github.com/jsesr/CS-GResNet-PyTorch. Shaoping Jiang, Xiangmin Xu 0001, Fang Liu 0030, Xiaofen Xing, Lin Wang 0004 |
ICASSP | 4 |
| 2022 | SpeechFormer: A Hierarchical Efficient Framework Incorporating the Characteristics of SpeechabstractTransformer has obtained promising results on cognitive speech signal processing field, which is of interest in various applications ranging from emotion to neurocognitive disorder analysis.However, most works treat speech signal as a whole, leading to the neglect of the pronunciation structure that is unique to speech and reflects the cognitive process.Meanwhile, Transformer has heavy computational burden due to its full attention operation.In this paper, a hierarchical efficient framework, called SpeechFormer, which considers the structural characteristics of speech, is proposed and can be served as a generalpurpose backbone for cognitive speech signal processing.The proposed SpeechFormer consists of frame, phoneme, word and utterance stages in succession, each performing a neighboring attention according to the structural pattern of speech with high computational efficiency.SpeechFormer is evaluated on speech emotion recognition (IEMOCAP & MELD) and neurocognitive disorder detection (Pitt & DAIC-WOZ) tasks, and the results show that SpeechFormer outperforms the standard Transformer-based framework while greatly reducing the computational cost.Furthermore, our SpeechFormer achieves comparable results to the state-of-the-art approaches. Xiaofen Xing, Xiangmin Xu 0001, Jianxin Pang |
INTERSPEECH | 2 |
| 2022 | Simple-action-guided dictionary learning for complex action recognition
Fang Liu 0030, Xiangmin Xu 0001, Xiaofen Xing, Kailing Guo, Lin Wang 0004 |
Neurocomputing | 3 |
| 2022 | ISNet: Individual Standardization Network for Speech Emotion RecognitionabstractSpeech emotion recognition plays an essential role in human-computer interaction. However, cross-individual representation learning and individual-agnostic systems are challenging due to the distribution deviation caused by individual differences. The existing related approaches mostly use the auxiliary task of speaker recognition to eliminate individual differences. Unfortunately, although these methods can reduce interindividual voiceprint differences, it is difficult to dissociate interindividual expression differences since each individual has its unique expression habits. In this paper, we propose an individual standardization network (ISNet) for speech emotion recognition to alleviate the problem of interindividual emotion confusion caused by individual differences. Specifically, we model individual benchmarks as representations of nonemotional neutral speech, and ISNet realizes individual standardization using the automatically generated benchmark, which improves the robustness of individual-agnostic emotion representations. In response to individual differences, we also propose more comprehensive and meaningful individual-level evaluation metrics. In addition, we continue our previous work to construct a challenging large-scale speech emotion dataset (LSSED). We propose a more reasonable division method of the training set and testing set to prevent individual information leakage. Experimental results on datasets of both large and small scales have proven the effectiveness of ISNet, and the new state-of-the-art performance is achieved under the same experimental conditions on IEMOCAP and LSSED. Weiquan Fan, Xiangmin Xu 0001, Bolun Cai, Xiaofen Xing |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Two-stream Gabor-AGraph Convolutional Networks for Facial Expression RecognitionabstractFacial expression recognition (FER) has recently attracted much attention in computer vision. However, existing methods mostly focus on the texture information of faces and overlook their inherent topological features. Hence more informative and significant contents are ignored for expression recognition. In this work, we propose a Two-stream Gabor-AGraph Convolutional Network (2s-GAGCN) to exploit the facial texture and topological features simultaneously. The Gabor Stream and the Attention-Graph (AGraph) Stream are respectively introduced to capture the salient visual properties and discriminative landmark features of faces. In particular, we adopt a flexible node attention mechanism in AGraph Stream through utilizing global and local information to enhance the potential relationships among landmarks. Furthermore, a novel landmark feature descriptor is proposed to alleviate the redundant topological features, which shows promising improvement for the recognition accuracy. We conduct extensive experiments on two wild datasets: RAF-DB and SFEW. The results show that the proposed 2s-GAGCN achieves superior performance against the state-of-the-art methods. Shaoping Jiang, Xiangmin Xu 0001, Xiaofen Xing, Lin Wang 0004, Fang Liu 0030 |
FG | 3 |
| 2021 | Two-stream Global-Guided Attention Network for Facial Expression RecognitionabstractFacial expression recognition (FER) in the wild is an important yet challenging problem because of uncontrolled conditions, such as occlusions, pose, and illumination. Most existing methods utilize the global and local information, but ignore the potential correlation between the global and local faces. In this paper, we propose a two-stream global-guided attention network (TGGAN) for FER in the wild. To further exploit the complementary relationship between local and global, we design a global-guided attention module (GGAM). Especially, a global guidance mechanism is proposed in GGAM to utilize the global information to guide the capture of key local features. Furthermore, inspired by the transformer, a self-attention mechanism is introduced in GGAM to emphasize salient face regions and fully integrate features extracted from global and local streams. We validate the proposed TGGAN on two wild datasets (FERPlus, RAF -DB) and further conduct experiments on their test subsets of occlusion and multi-poses. Extensive experiments show that the proposed TGGAN achieves superior performance against the state-of-the-art methods. Yaoli Wen, Xiangmin Xu 0001, Fang Liu 0030, Xiaofen Xing, Lin Wang 0004 |
FG | 4 |
| 2021 | LSSED: A Large-Scale Dataset and Benchmark for Speech Emotion RecognitionabstractSpeech emotion recognition is a vital contributor to the next generation of human-computer interaction (HCI). However, current existing small-scale databases have limited the development of related research. In this paper, we present LSSED, a challenging large-scale english speech emotion dataset, which has data collected from 820 subjects to simulate real- world distribution. In addition, we release some pre-trained models based on LSSED, which can not only promote the development of speech emotion recognition, but can also be transferred to related downstream tasks such as mental health analysis where data is extremely difficult to collect. Finally, our experiments show the necessity of large-scale datasets and the effectiveness of pre-trained models. The dateset will be released on https://github.com/tobefans/LSSED. Weiquan Fan, Xiangmin Xu 0001, Xiaofen Xing, Dong-Yan Huang |
ICASSP | 3 |
| 2021 | Cross-Lingual Voice Conversion with Disentangled Universal Linguistic Representations
Zhenchuan Yang, Xiaofen Xing |
Interspeech | 4 |
| 2021 | Weight Evolution: Improving Deep Neural Networks Training through Evolving Inferior Weight ValuesabstractTo obtain good performance, convolutional neural networks are usually over-parameterized. This phenomenon has stimulated two interesting topics: pruning the unimportant weights for compression and reactivating the unimportant weights to make full use of network capability. However, current weight reactivation methods usually reactivate the entire filters, which may not be precise enough. Looking back in history, the prosperity of filter pruning is mainly due to its friendliness to hardware implementation, but pruning at a finer structure level, i.e., weight elements, usually leads to better network performance. We study the problem of weight element reactivation in this paper. Motivated by evolution, we select the unimportant filters and update their unimportant elements by combining them with the important elements of important filters, just like gene crossover to produce better offspring, and the proposed method is called weight evolution (WE). WE is mainly composed of four strategies. We propose a global selection strategy and a local selection strategy and combine them to locate the unimportant filters. A forward matching strategy is proposed to find the matched important filters and a crossover strategy is proposed to utilize the important elements of the important filters for updating unimportant filters. WE is plug-in to existing network architectures. Comprehensive experiments show that WE outperforms the other reactivation methods and plug-in training methods with typical convolutional neural networks, especially lightweight networks. Our code is available at https://github.com/BZQLin/Weight-evolution. Zhenquan Lin, Kailing Guo, Xiaofen Xing, Xiangmin Xu 0001 |
ACM Multimedia | 3 |
| 2021 | Spatiotemporal and frequential cascaded attention networks for speech emotion recognition
Shuzhen Li, Xiaofen Xing, Weiquan Fan, Bolun Cai, Perry Fordson, Xiangmin Xu 0001 |
Neurocomputing | 2 |
| 2021 | Attention-aware concentrated network for saliency prediction
Pengqian Li, Xiaofen Xing, Xiangmin Xu 0001, Bolun Cai |
Neurocomputing | 2 |
| 2021 | Hierarchical Lifelong Learning by Sharing Representations and Integrating HypothesisabstractIn lifelong machine learning (LML) systems, consecutive new tasks from changing circumstances are learned and added to the system. However, sufficiently labeled data are indispensable for extracting intertask relationships before transferring knowledge in classical supervised LML systems. Inadequate labels may deteriorate the performance due to the poor initial approximation. In order to extend the typical LML system, we propose a novel hierarchical lifelong learning algorithm (HLLA) consisting of two following layers: 1) the knowledge layer consisted of shared representations and integrated knowledge basis at the bottom and 2) parameterized hypothesis functions with features at the top. Unlabeled data is leveraged in HLLA for pretraining of the shared representations. We also have considered a selective inherited updating method to deal with intertask distribution shifting. Experiments show that our HLLA method outperforms many other recent LML algorithms, especially when dealing with higher dimensional, lower correlation, and fewer labeled data problems. Tong Zhang 0015, Guoxi Su, Chunmei Qing, Xiangmin Xu 0001, Bolun Cai, Xiaofen Xing |
IEEE Trans. Syst. Man Cybern. Syst. | 6 |
| 2020 | Learning Spatio-Temporal Convolutional Network for Real-Time Object TrackingabstractSiamese series of tracking networks have shown great potentials in achieving balanced accuracy and beyond real-time speed. However, most of existing siamese trackers only consider appearance features of first frame, and hardly benefit from interframe information. The lack of latest temporal transformation degrades the tracking performance during challenges such as deformation and partial occlusion. In this paper we focus on making using of the rich information in latest consecutive frames to improve the feature representation of initial template frame. Specifically, the latest frames after 3d convolution are used to generate an attention map, which is then point-wise multiplied by the features of first frame to obtain the updated template. With the attention map, the template can adaptively cope with the deformation and occlusion of the target. Since the first frame is always used as the basis of the template, there is no cumulative error when using the latest frames for attention. Due to the shared 2d convolution of all frames, the feature map results can be reused so that the added module has almost no time-consuming effects. This module is easily embedded into different siamese trackers. Through verification, the module has significantly improved tracking performance in different backbone situations. Hanzao Chen, Xiaofen Xing, Xiangmin Xu 0001 |
ICASSP | 2 |
| 2020 | Adaptive Domain-Aware Representation Learning for Speech Emotion Recognition
Weiquan Fan, Xiangmin Xu 0001, Xiaofen Xing, Dong-Yan Huang |
INTERSPEECH | 3 |
| 2018 | Perception Preserving DecolorizationabstractDecolorization is a basic tool to transform a color image into a grayscale image, which is used in digital printing, stylized black-and-white photography, and in many single-channel image processing applications. While recent researches focus on retaining as much as possible meaningful visual features and color contrast. In this paper, we explore how to use deep neural networks for decolorization, and propose an optimization approach aiming at perception preserving. The system uses deep representations to extract content information based on human visual perception, and automatically selects suitable grayscale for decolorization. The evaluation experiments show the effectiveness of the proposed method. Bolun Cai, Xiangmin Xu 0001, Xiaofen Xing |
ICIP | 3 |
| 2018 | Learning Adaptive Selection Network for Real-Time Visual TrackingabstractOffline-trained trackers based on convolutional neural networks (CNNs) have shown great potential in achieving balanced accuracy and real-time speed. However, offline-trained trackers are prone to drift to background clutters. In this paper, we present an adaptive selection network tracker (ASNT) to address the tracking drift problem. Inspired by feature selection technique used in other vision problems, we introduce a learnable selection unit for Siamese network based trackers. The selection unit enables the tracker to select relevant feature map automatically for the target. Channel dropout is applied in the selection unit to improve generalization performance for convolutional layers. To further improve the discrimination between background clutters and the target, an adaptive method is used to initialize the tracker for each video sequence. Experiments on OTB-2013 and VOT2014 datasets demonstrate that our ASNT tracker has a comparable performance against state-of-the-art methods, yet can run at a speed of over 100 fps. Jiangfeng Xiong, Xiangmin Xu 0001, Bolun Cai, Xiaofen Xing, Kailing Guo |
ICME | 4 |
| 2017 | Edge/structure preserving smoothing via relativity-of-GaussianabstractThis paper presents a novel edge/structure-preserving image smoothing via relativity-of-Gaussian. As a simple local regularization, it performs the local analysis of scale features and globally optimizes its results into a piecewise smooth. The central idea to ensure proper texture smoothing is based on cross-scale relative that captures the weak textures from the most prominent edges/structures. Our method outperforms the previous methods in removing the detail information while preserving main image content. Bolun Cai, Xiaofen Xing, Xiangmin Xu 0001 |
ICIP | 2 |
| 2017 | Robust object tracking based on sparse representation and incremental weighted PCA
Xiaofen Xing, Fuhao Qiu, Xiangmin Xu 0001, Chunmei Qing, Yinrong Wu |
Multim. Tools Appl. | 1 |
| 2016 | BIT: Biologically Inspired TrackerabstractVisual tracking is challenging due to image variations caused by various factors, such as object deformation, scale change, illumination change, and occlusion. Given the superior tracking performance of human visual system (HVS), an ideal design of biologically inspired model is expected to improve computer visual tracking. This is, however, a difficult task due to the incomplete understanding of neurons' working mechanism in the HVS. This paper aims to address this challenge based on the analysis of visual cognitive mechanism of the ventral stream in the visual cortex, which simulates shallow neurons (S1 units and C1 units) to extract low-level biologically inspired features for the target appearance and imitates an advanced learning mechanism (S2 units and C2 units) to combine generative and discriminative models for target location. In addition, fast Gabor approximation and fast Fourier transform are adopted for real-time learning and detection in this framework. Extensive experiments on large-scale benchmark data sets show that the proposed biologically inspired tracker performs favorably against the state-of-the-art methods in terms of efficiency, accuracy, and robustness. The acceleration technique in particular ensures that biologically inspired tracker maintains a speed of approximately 45 frames/s. Bolun Cai, Xiangmin Xu 0001, Xiaofen Xing, Kui Jia, Jie Miao, Dacheng Tao |
IEEE Trans. Image Process. | 3 |
| 2015 | BIT: Bio-inspired trackerabstractVisual tracking is a challenging problem due to various factors such as deformation, rotation and illumination. As is well known, given the superior tracking performance of human vision, bio-inspired model is expected to improve the computer visual tracking. However, the design of bio-inspired tracking framework is challenging, due to the incomplete comprehension and hyper-scale of senior neurons, which will influence the effectiveness and real-time performance of the tracker. According to the ventral stream in visual cortex, a novel bio-inspired tracker (BIT) is proposed, which simulates shallow neurons (S1 and C1) to extract low-level bio-inspired feature for target appearance and imitates senior learning mechanism (S2 and C2) to combine generative and discriminative model for position estimation. In addition, Fast Fourier Transform (FFT) is adopted for real-time learning and detection in this framework. On the recent benchmark[1], extensive experimental results show BIT performs favorably against state-of-the-art methods in terms of accuracy and robustness. Bolun Cai, Xiangmin Xu 0001, Xiaofen Xing, Chunmei Qing |
ICIP | 3 |
| 2014 | High speed deep networks based on Discrete Cosine TransformationabstractThe traditional deep networks take raw pixels of data as input, and automatically learn features using unsupervised learning algorithms. In this configuration, in order to learn good features, the networks usually have multi-layer and many hidden units which lead to extremely high training time costs. As a widely used image compression algorithm, Discrete Cosine Transformation (DCT) is utilized to reduce image information redundancy because only a limited number of the DCT coefficients can preserve the most important image information. In this paper, it is proposed that a novel framework by combining DCT and deep networks for high speed object recognition system. The use of a small subset of DCT coefficients of data to feed into a 2-layer sparse auto-encoders instead of raw pixels. Because of the excellent decorrelation and energy compaction properties of DCT, this approach is proved experimentally not only efficient, but also it is a computationally attractive approach for processing high-resolution images in a deep architecture. Xiaoyi Zou, Xiangmin Xu 0001, Chunmei Qing, Xiaofen Xing |
ICIP | 4 |