Huan Zhao 0003

dblp:15/3548-3 · DBLP profile ↗
← Back
51ranked-venue papers
17as first author
44since 2021 · last 2026
0000-0001-6286-5868ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 9 first-author · 21 since 2021Artificial intelligence and machine learning · 20 · 7 first-author · 18 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1Theory of computation · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 PLUM-Net: Prototype-Induced Label Structuring for Disentangled Multimodal Representation Network
abstract
Existing multimodal representation learning approaches often rely on simple feature concatenation or unified transformations, which fail to effectively disentangle and leverage common and private information across different modalities in a progressive manner. Moreover, they typically lack adaptive modeling tailored to specific task requirements. To address these limitations, we propose a Prototype-Induced Label Structuring for Disentangled Multimodal Representation Network (PLUM-Net). It first employs a multilevel semantic alignment module to synchronize global and local semantics across audio, visual and textual streams. On this aligned foundation, a prototype-based single-modal label generation module derives modality-specific hard and soft-labels that subtly steer the network toward a cleaner split between shared and private cues. Guided by these labels, the task-conditioned feature bifurcator module channels information through the most beneficial common or private pathway for the given task, after which a private refinement module polishes and fuses each modality’s idiosyncratic signals. Extensive experiments show that PLUM-Net delivers strong performance on datasets such as CMU-MOSI, CMU-MOSEI and UR-FUNNY, achieving an ACC-2 of 90.3% on CMU-MOSI, representing a 2%–4% improvement over previous SOTA models.
Huan Zhao 0003, Xupeng Zha, Guanghui Ye, Zixing Zhang 0001
AAAI2
2026 Making Visual Dialogue More Engaging: A New Task, Method, and Metric
abstract
Large language model (LLM)-based visual dialogue (VD) systems have made response generation for image-grounded conversations more correct and coherent. However, user engagement - the extent to which a user is interested, emotionally involved, and willing to continue the conversation - remains a challenge. To fully explore engaging VD, we propose: (i) a new task named Audio-enhanced VD (AVD), which introduces additional audio dialogue contexts that can more vividly convey the speaker's emotions as input, with the aim of generating correct but more engaging dialogue responses. Specifically, we employ a text-to-speech model as the modality translator to generate the paired acoustic utterances from the inputting textual utterances; (ii) an accompanying approach named Visually-grounded and Interleaved Text-Audio Dialogue Modeling (VITA-DM), which utilizes both image-grounded information and interleaved text-audio utterances for visual dialogue modeling, differentiating from previous multi-modal LLM (MLLM)-based methods that normally model text and audio modalities separately. We also present three pre-training tasks to better learn multi-modal interactions across language, vision, and audio; (iii) a novel metric named Multi-Modal Engagement (MME), which fills the gap of engagement estimation in VD and can provide a fine-grained assessment along emotional, attentional, and reply engagement dimensions (EE, AE, RE). We experiment on two popular datasets and provide extensive evaluations (automatic, engagement-specific, and human), supporting the validity of our approach. Furthermore, based on empirical results that reveal that emotions contribute the most to engagement, we justify our emphasis on the emotional aspect throughout the definition, solution, and evaluation of our task.
Guanghui Ye, Huan Zhao 0003, Yingxue Gao, Zhixue Zhao, Xupeng Zha, Zhihua Jiang
AAAI2
2026 Automatic and Reliable Evaluation for Academic Caption-to-Figure Generation with LMMs
abstract
Guanghui Ye, Huan Zhao, Qin Zhu, Fengnan Li, Jiaqi Li, Yixian Shen, Zhonghao Ren, Zhihua Jiang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Guanghui Ye, Huan Zhao 0003, Fengnan Li, Jiaqi Li 0008, Yixian Shen, Zhonghao Ren, Zhihua Jiang
ACL (1)2
2026 DyGIN: Modality-Unified Dynamic Graph Inception Network for Context-Aware Emotion Recognition
abstract
Graph representation learning has attracted considerable attention for its ability to generate node representations by aggregating information from neighboring nodes. However, existing graph-based methods for sequential data merely focus on a static graph constructed on an entire unimodal sequence, largely ignoring their dynamic evolution. To overcome this limitation, we introduce a modality-unifiedDynamicGraphInceptionNetwork (DyGIN) that models dynamic evolutionary graphs for different modalities. DyGIN constructs dynamic graphs from temporal subsequences using a sliding window, and incorporates a Temporal Graph Evolution Gated Recurrent Unit (TGE-GRU) to update graph weights at each segment. Additionally, we propose a context-aware similarity matrix, which updates node representations based on the degree of neighboring nodes, replacing the traditional averaging method. Our model is optimized with a combination of graph structure loss, classification loss, and a learnable pooling function. We validate DyGIN on emotion recognition tasks across three public datasets—RML, IEMOCAP, and DEAP—where it outperforms state-of-the-art models by 1.6% on RML, 1.81% and 2.33% on IEMOCAP, and 2.78% and 6.81% on DEAP, based on accuracy and weighted F1 score. Our code is available at https://github.com/G22-web/DyGIN.
Yingxue Gao, Huan Zhao 0003, Zixing Zhang 0001
IEEE Internet Things J.2
2026 Generating Multi-Modal Knowledge Clues as an Image: Toward Improving Image-Sequence Reasoning With Assisted Visual Input
abstract
Recent multi-modal large language models (MLLMs) have exhibited powerful abilities in addressing complex vision-language tasks such as image-sequence reasoning (ISR). However, significant challenges remain, e.g., it is still difficult for the MLLMs to fully capture and represent cross-image visual knowledge such as scene relations, attributes, and entity links between multiple images, which hinders them from better solving ISR. To alleviate these issues, we introduce a novel concept Visualized Knowledge Clue (VizKC) - synthetic images that encode key visual and external knowledge from a sequence of input images and are then used alongside the original input images within a multi-image MLLM to enhance reasoning performance. Accordingly, we propose an accompanying approach named VizKC-ISR, composed of two modules - VizKC generation and VizKC utilization. Specifically, in the generation module, VizKC-ISR follows aSee-Find-Fusepipeline: (i) “See - Scene Perception”, to construct an initial VizKC that incorporates scene relations of key visual entities detected from an original image; (ii) “Find - Knowledge Generation”, to generate enriched image captions with real-world knowledge and fine-grained entity details and then extract structured knowledge tuples from generated captions; (iii) “Fuse - Image Editing”, to introduce relevant knowledge tuples into the VizKC via iterative image editing. In the utilization module, we employ a multi-image MLLM (e.g., mPLUG-Owl3) to solve the VizKC-assisted ISR tasks by reasoning with generated knowledge clues. We evaluate VizKC-ISR on nine ISR benchmarks categorized into three multi-image scenarios. The results show that our VizKC-ISR performs best in all tasks, e.g., obtaining the highest average accuracy of 63.1% and surpassing the mPLUG-Owl3 baseline by 6.4 absolute points, due to the bridge between visually-grounded reasoning and multi-modal knowledge challenges.
Guanghui Ye, Huan Zhao 0003, Yixian Shen, Jiaqi Li 0008, Fengnan Li, Zhihua Jiang, Keqin Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Dual-View Learning for Conversational Emotion Recognition Through Context and Emotion-Shift Modeling
abstract
Conversational Emotion Recognition (CER) has recently been explored through conversational context modeling to learn the emotion distribution, i.e., the likelihood over emotion categories associated with each utterance. While these methods have shown promising results in emotion classification, they often focus on the interactions between utterances (utterance-view) and overlook shifts in the speaker's emotions (emotion-view). This emphasis on homogeneous view modeling limits their overall effectiveness. To address this limitation, we propose DVL-CER, a novel Dual-View Learning approach for CER. DVL-CER integrates both the utterance-view and emotion-view using two projection heads, enabling cross-view projection of emotion distributions. Our approach offers several key advantages: (1) We introduce an emotion-view that captures shifts in a speaker's emotions from initial to subsequent states within a conversation. This view enriches the conversation modeling and supports seamless integration with various CER baseline models. (2) Our dual-view projection learning strategy flexibly balances consistency and independence between the two heterogeneous views, promoting view-specific adaptation learning and incorporating the emotion verification capability within CER. We validate DVL-CER through extensive experiments on two widely-used datasets, IEMOCAP and EmoryNLP. The results demonstrate that DVL-CER achieves state-of-the-art performance, delivering robust and high-quality emotion distributions compared with existing CER methods and other dual-view learning strategies.
Xupeng Zha, Huan Zhao 0003, Guanghui Ye, Zixing Zhang 0001
AAAI2
2025 Knowledge Image Matters: Improving Knowledge-Based Visual Reasoning with Multi-Image Large Language Models
abstract
We revisit knowledge-based visual reasoning (KB-VR) in light of modern advances in multimodal large language models (MLLMs), and make the following contributions: (i) We propose Visual Knowledge Card (VKC) -a novel image that incorporates not only internal visual knowledge (e.g., scene-aware information) detected from the raw image, but also external world knowledge (e.g., attribute or object knowledge) produced by a knowledge generator; (ii) We present VKC-enhanced Multi-Image Reasoning (VKC-MIR) -a fourstage pipeline which harnesses a state-of-theart scene perception engine to construct an initial VKC (Stage-1), a powerful LLM to generate relevant domain knowledge (Stage-2), an excellent image editing toolkit to introduce generated knowledge into an iteratively-edited VKC (Stage-3), and finally, an emerging multiimage MLLM to solve the VKC-enhanced task (Stage-4).By performing experiments on three popular KB-VR benchmarks, our approach achieves new state-of-the-art results compared to previous top-performing models.Our code is available at: https://github. com/yyy1103/VKC.
Guanghui Ye, Huan Zhao 0003, Zhixue Zhao, Xupeng Zha, Zhihua Jiang
ACL (1)2
2025 XDGesture: An xLSTM-based Diffusion Model for Co-speech Gesture Generation
abstract
In multimodal human-computer interaction, generating co-speech gestures is crucial for enhancing interaction naturalness and user experience. However, achieving synchronized and natural gesture sequences remains a significant challenge due to the complexity of modeling temporal dependencies across different modalities. Existing methods often rely on simple concatenation techniques, which are limited in effectively handling multimodal information. To address this issue, we propose XDGesture, a diffusion-based framework that integrates a Cross-Modal Fusion module and xLSTM. The Cross-Modal Fusion module efficiently merges information from different modalities, providing the model with rich contextual conditions. Meanwhile, xLSTM, with its enhanced memory structure and exponential gating mechanism, processes the fused multimodal data, capturing long-range dependencies between speech and gestures. This enables the generation of high-quality gesture sequences that are naturally synchronized with speech. Experimental results demonstrate that XDGesture remarkably outperforms existing baselines on multiple datasets, particularly in terms of gesture quality, naturalness, and synchronization with speech.
Zixing Zhang 0001, Huan Zhao 0003, Björn W. Schuller
ICASSP5
2025 Enhanced Multimodal Emotion Recognition in Conversations via Contextual Filtering and Multi-Frequency Graph Propagation
abstract
Multimodal Emotion Recognition in Conversations (ERC) plays a crucial role in understanding human language and behavior in real-world scenarios. However, existing research tends to simply concatenate multimodal representations, failing to capture the complex relationships between modalities. Recent advances have shown that Graph Neural Networks (GNNs) are effective in capturing complex data relationships, offering a promising solution for multimodal ERC. Despite this, current GNN-based methods still face challenges, including weak interactions between modalities, neglecting the information entropy of utterances, and erasure of high-frequency signals that capture key variations and discrepancies between closely related nodes. To address these limitations, we propose a GNNs-based multi-frequency propagation method enhanced by contextual filtering for multimodal ERC. Our approach introduces a context filtering module that combines a similarity matrix and an information entropy matrix, enabling GNNs to effectively capture the inherent relationships among utterances and provide sufficient multimodal and contextual modeling. Additionally, our method explores multivariate relationships by recognizing the varying importance of emotional discrepancies and commonalities through multi-frequency signals. Experimental results on two benchmark datasets, IEMOCAP and MELD, demonstrate that our method outperforms the latest (non-)graph-based works. Our method is available at https://github.com/G22-web/ConFilMER.
Huan Zhao 0003, Yingxue Gao, Haijiao Chen, Guanghui Ye, Zixing Zhang 0001
ICASSP1
2025 Parameter-Efficient Federal-Tuning Enhances Privacy Preserving for Speech Emotion Recognition
abstract
The Pre-trained Speech Models (PSMs) generate universal speech representations using self-supervised or weakly-supervised learning from large-scale datasets. It achieves promising performance when fine-tuned for specific tasks such as Speech Emotion Recognition (SER). However, fine-tuning on various datasets requires storing the entire model’s weight parameters, complicating real-world deployment. Additionally, centralized fine-tuning relies on user data, posing significant privacy risks. To address these challenges, we propose employing Federated Learning (FL) for fine-tuning PSMs with a Parameter-Efficient Fine-Tuning (PEFT) method. By embedding trainable layers in the feed-forward layers of the pre-trained model, we keep the backbone model frozen and only update the trainable layer parameters during federated training, significantly reducing parameter transmission. Specifically, we evaluated the performance of downstream model fine-tuning, adapter tuning, embedding prompt tuning, and LoRA within a federated fine-tuning framework for PSMs, demonstrating the framework’s feasibility and effectiveness. Furthermore, attribute inference attack tests showed that gender inference results on three datasets were at chance levels.
Haijiao Chen, Huan Zhao 0003, Yingxue Gao, Zixing Zhang 0001
ICASSP2
2025 DSSM: Dual State Space Model For Human Motions Generation
abstract
Text-driven human motion generation has attracted considerable critical attention in recent years. The task requires generating movements that are diverse, natural, and comfortable in accordance with the text description. However, while generating the human motion, there is a significant gap in the amount of feature information contained in word-based text description modality and joint-based human motion modality. This results in an extreme imbalance of feature information in the latent space across the modalities, which seriously affects the effect of feature fusion. To alleviate the imbalance between modalities, we propose the Dual State Space Model (DSSM), which reconstruction the fused feature from coarse-to-fine. The DSSM contains two unit structures: the Masked State Space Model (MSSM) and the Hierarchical State Space Block (HSSB). At the same time, in order to make better use of timing information and reduce the computational complexity of the model, the DSSM is also the first method to introduce the state space model (SSM) into the text-driven motion sequence generation. We evaluated the DSSM on the HumanML3D and KITML benchmark datasets, and the experimental results show that our approach achieves state-of-the-art performance. The source code is available on GitHub at https://github.com/yimingliu123/DSSM.
Huan Zhao 0003, Yaqian Liu, Haijiao Chen, Guanghui Ye, Zixing Zhang 0001
ICASSP2
2025 PTalker: Personalized Speech-Driven 3D Talking Head Animation via Style Disentanglement and Modality Alignment
abstract
Speech-driven 3D talking head generation aims to produce lifelike facial animations precisely synchronized with speech. While considerable progress has been made in achieving high lip-synchronization accuracy, existing methods largely overlook the intricate nuances of individual speaking styles, which limits personalization and realism. In this work, we present a novel framework for personalized 3D talking head animation, namely ''PTalker''. This framework preserves speaking style through style disentanglement from audio and facial motion sequences and enhances lip-synchronization accuracy through a three-level alignment mechanism between audio and mesh modalities. Specifically, to effectively disentangle style and content, we design disentanglement constraints that encode driven audio and motion sequences into distinct style and content spaces to enhance speaking style representation. To improve lip-synchronization accuracy, we adopt a modality alignment mechanism incorporating three aspects: spatial alignment using Graph Attention Networks to capture vertex connectivity in the 3D mesh structure, temporal alignment using cross-attention to capture and synchronize temporal dependencies, and feature alignment by top-k bidirectional contrastive losses and KL divergence constraints to ensure consistency between speech and mesh modalities. Extensive qualitative and quantitative experiments on public datasets demonstrate that PTalker effectively generates realistic, stylized 3D talking heads that accurately match identity-specific speaking styles, outperforming state-of-the-art methods. The source code and supplementary videos are available at: PTalker.
Yang Xu 0013, Huan Zhao 0003, Hao Zhang 0139, Zixing Zhang 0001
ACM Multimedia3
2025 AudioFab: Building A General and Intelligent Audio Factory through Tool Learning
abstract
Currently, artificial intelligence is profoundly transforming the audio domain; however, numerous advanced algorithms and tools remain fragmented, lacking a unified and efficient framework to unlock their full potential. Existing audio agent frameworks often suffer from complex environment configurations and inefficient tool collaboration. To address these limitations, we introduce AudioFab, an open-source agent framework aimed at establishing an open and intelligent audio-processing ecosystem. Compared to existing solutions, AudioFab's modular design resolves dependency conflicts, simplifying tool integration and extension. It also optimizes tool learning through intelligent selection and few-shot learning, improving efficiency and accuracy in complex audio tasks. Furthermore, AudioFab provides a user-friendly natural language interface tailored for non-expert users. As a foundational framework, AudioFab's core contribution lies in offering a stable and extensible platform for future research and development in audio and multimodal AI. The code is available at https://github.com/SmileHnu/AudioFab.
Jing Han 0010, Qianshuai Xue, Huan Zhao 0003, Zixing Zhang 0001
ACM Multimedia5
2025 Federal parameter-efficient fine-tuning for speech emotion recognition
Haijiao Chen, Huan Zhao 0003, Zixing Zhang 0001, Keqin Li 0001
Expert Syst. Appl.2
2025 UniDE: A multi-level and low-resource framework for automatic dialogue evaluation via LLM-based data augmentation and multitask learning
Guanghui Ye, Huan Zhao 0003, Zixing Zhang 0001, Zhihua Jiang
Inf. Process. Manag.2
2025 CCDE: A Compact and Competitive Dialogue Evaluation Framework via Knowledge Distillation of Large Language Models
abstract
Automatic evaluation metrics not only play a vital role in developing dialogue and interactive systems but also have a great impact on social activities in our daily life. However, previous specialized metrics for evaluating dialogues exhibit a relatively low correlation with human judgments. In addition, today’s state-of-the-art (SOTA) evaluators that leverage large language models (LLMs) are challenging to deploy in real-world applications due to their sheer size. To this end, we propose a novel evaluation framework, compact and competitive dialogue evaluation (CCDE), which leverages knowledge distillation of LLMs to generate training data and sequentially learn a multitask evaluator regarding diversified quality dimensions. Specifically, we first employ ChatGPT asteacherto generate a high-quality and rich-annotation corpus, CCDE-data. Then, we implement astudentevaluator CCDE (1.3B) via using InstructGPT as the backbone model that is trained and fine-tuned on CCDE-data. We conduct extensive experiments on three public benchmarks: fine-grained evaluation of dialog (FED), PersonaChat, and TopicalChat. The results demonstrate that our model CCDE can outperform the current SOTA model G-Eval which calls GPT-4 ($\boldsymbol{\geq}$175B) by 4.3 on the FED dataset, 3.5 on the PersonaChat dataset, and 0.3 on the TopicalChat dataset, in terms of the Spearman correlation metric (%). We release the data and code at:https://anonymous.4open.science/r/ccde-3827.
Guanghui Ye, Huan Zhao 0003, Haijiao Chen, Zhixue Zhao, Zhihua Jiang, Keqin Li 0001
IEEE Trans. Comput. Soc. Syst.2
2025 Robust Hashing With Bilinear Drift for Image-Text Retrieval
abstract
Supervised hashing models for image-text retrieval are fundamental and versatile in social media analysis and cross-lingual web search. Among them, supervised bilinear drift hashing is one of the most popular approaches. However, it still faces several challenges. For instance, how to leverage the power of bilinear drift hashing to distinguish similar and dissimilar data samples effectively; how to strengthen the semantic relationship between similar data and supervision. To solve these problems, we propose Robust Hashing with Bilinear Drift (RHBD) to improve the accuracy and robustness of the supervised model. The key idea of this work is to generate effective hash codes between image-text feature representations by combining robust data distributions and multiple supervision information. The benefits of bilinear drift with robust hashing, which enhance the discrimination of hash binary, are manifested mainly in two ways: (1) RHBD employs a semantic autoencoder with a linear drift to get a discriminative common feature representation between image and text modalities; (2) RHBD explores iteration quantization with a linear drift to well generate similarity-preserving hash codes. Moreover, we introduce multiple supervision learning to promote the consistency between data information and supervision knowledge for semantic complementarity. Results on three public datasets show that RHBD is effective in image-text retrieval, consistently outperforming other state-of-the-art models with comparable training efficiency to competitive baselines.
Huan Zhao 0003, Song Wang 0016, Zixing Zhang 0001, Keqin Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2024 Adaptive Speech Emotion Representation Learning Based On Dynamic Graph
abstract
Graph representation learning has become a hot research topic due to its powerful nonlinear fitting capability in extracting representative node embeddings. However, for sequential data such as speech signals, most traditional methods merely focus on the static graph created within a sequence, and largely overlook the intrinsic evolving patterns of these data. This may reduce the efficiency of graph representation learning for sequential data. For this reason, we propose an adaptive graph representation learning method based on dynamically evolved graphs, which are consecutively constructed on a series of subsequences segmented by a sliding window. In doing this, it is better to capture local and global context information within a long sequence. Moreover, we introduce a weighted approach to update the node representation rather than the conventional average one, where the weights are calculated by a novel matrix computation based on the degree of neighboring nodes. Finally, we construct a learnable graph convolutional layer that combines the graph structure loss and classification loss to optimize the graph structure. To verify the effectiveness of the proposed method, we conducted experiments for speech emotion recognition on the IEMOCAP and RAVDESS datasets. Experimental results show that the proposed method outperforms the latest (non-)graph-based models.
Yingxue Gao, Huan Zhao 0003, Zixing Zhang 0001
ICASSP2
2024 Customising General Large Language Models for Specialised Emotion Recognition Tasks
abstract
The advent of large language models (LLMs) has gained tremendous attention over the past year. Previous studies have shown the astonishing performance of LLMs not only in other tasks but also in emotion recognition in terms of accuracy, universality, explanation, robustness, few/zero-shot learning, and others. Leveraging the capability of LLMs inevitably becomes an essential solution for emotion recognition. To this end, we further comprehensively investigate how LLMs perform in linguistic emotion recognition if we concentrate on this specific task. Specifically, we exemplify a publicly available and widely used LLM – Chat General Language Model, and customise it for our target by using two different model adaptation techniques, i.e., deep prompt tuning and low-rank adaptation. The experimental results obtained on six widely used datasets present that the adapted LLM can easily outperform other state-of-the-art but specialised deep models. This indicates the strong transferability and feasibility of LLMs in the field of emotion recognition.
Liyizhe Peng, Zixing Zhang 0001, Jing Han 0010, Huan Zhao 0003, Björn W. Schuller
ICASSP5
2024 Esihgnn: Event-State Interactions Infused Heterogeneous Graph Neural Network for Conversational Emotion Recognition
abstract
Conversational Emotion Recognition (CER) aims to predict the emotion expressed by an utterance (referred to as an "event") during a conversation. Existing graph-based methods mainly focus on event interactions to comprehend the conversational context, while overlooking the direct influence of the speaker’s emotional state on the events. In addition, real-time modeling of the conversation is crucial for real-world applications but is rarely considered. Toward this end, we propose a novel graph-based approach, namely Event-State Interactions infused Heterogeneous Graph Neural Network (ESIHGNN), which incorporates the speaker’s emotional state and constructs a heterogeneous event-state interaction graph to model the conversation. Specifically, a heterogeneous directed acyclic graph neural network is employed to dynamically update and enhance the representations of events and emotional states at each turn, thereby improving conversational coherence and consistency. Furthermore, to further improve the performance of CER, we enrich the graph’s edges with external knowledge. Experimental results on four publicly available CER datasets show the superiority of our approach and the effectiveness of the introduced heterogeneous event-state interaction graph.
Xupeng Zha, Huan Zhao 0003, Zixing Zhang 0001
ICASSP2
2024 Bilevel Relational Graph Representation Learning-based Multimodal Emotion Recognition in Conversation
abstract
Emotion recognition in conversation (ERC) is a key research topic in natural language processing, helping computers understand human emotions. Although substantial strides have been made in deep learning methods, the exploration of graphs in multimodal ERC is still in its infancy. The bottleneck of graph-based methods lies in the neighborhood aggregation strategy, a mechanism through which node attributes are transmitted and gathered. Existing strategies suffer from the issue of redundant irrelevant information, causing interference with the discriminative information of nodes. Moreover, traditional single-layer graph convolutional networks face challenges in efficiently extracting long-range contextual information. To address these issues, we propose a multimodal conversational emotion recognition approach based on a bilevel relational graph (BiGraph). Specifically, we construct two graphs: a global affinity graph, clustered by assessing node similarity between the target node and its neighborhood nodes to preserve discriminative information. Another is a local context dependency graph based on information from different speakers. The edges of these graphs are mainly determined by speaker context and temporal relations. Our method yields compelling results in extensive experiments conducted on the IEMOCAP and MOSEI datasets, which demonstrate the effectiveness and superiority of the proposed model. Our code is available at https://github.com/LiMei0329/BiGraph.
Huan Zhao 0003, Yi Ju, Yingxue Gao
ICME1
2024 LSTDial: Enhancing Dialogue Generation via Long- and Short-Term Measurement Feedback
abstract
Guanghui Ye, Huan Zhao, Zixing Zhang, Xupeng Zha, Zhihua Jiang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Guanghui Ye, Huan Zhao 0003, Zixing Zhang 0001, Xupeng Zha, Zhihua Jiang
NAACL-HLT2
2024 Knowledge enhancement for speech emotion recognition via multi-level acoustic feature
abstract
Speech emotion recognition (SER) has become an increasingly attractive machine learning task for domain applications.It aims to improve the discriminative capacity of speech emotion utilising a certain type of features (e.g.MFCC, Spectrograms, Wav2vec2) or multi-type combination features.However, the potential of acousticrelated deep features is frequently overlooked in existing approaches that rely solely on a single type of feature or employ a basic combination of multiple feature types.To address this challenge, a multi-level acoustic feature cross-fusion approach is proposed, aiming to compensate for missing information between various features.It helps to enhance the SER performance by integrating different types of knowledge through the cross-fusion mechanism.Moreover, multitask learning is utilised to share useful information through gender recognition, which can also obtain multiple common representations in a fine-grained space.Experimental results show that the fusion approach can capture the inner connections between multilevel acoustic features to refine the knowledge.The SOTA results were obtained under the same experimental conditions.
Huan Zhao 0003, Nianxin Huang, Haijiao Chen
Connect. Sci.1
2024 Individual mapping and asymmetric dual supervision for discrete cross-modal hashing
Song Wang 0016, Huan Zhao 0003, Zixing Zhang 0001, Keqin Li 0001
Expert Syst. Appl.2
2024 Discriminative Feature Learning-Based Federated Lightweight Distillation Against Multiple Attacks
abstract
Thanks to the advantages of cloud and edge computing, federated learning (FL)–based speech emotion recognition (SER) tasks can be well-scaled to cloud-edge-terminal ecosystems. It aims to characterize emotions while protecting data privacy. However, catastrophic forgetting caused by data heterogeneity, potential system attacks, and possible privacy leakage and communication overhead from parameter sharing have constrained its breakthrough. Some schemes that attempt to tackle the FL bottleneck do not consider these issues comprehensively. We propose a federated distillation-based multiple defense approach (FedMud), which simultaneously considers how to balance system performance, privacy security, and communication overhead. First, It employs a server-side lightweight generator to learn global view knowledge and guides client-side updates through distillation, further mitigating catastrophic forgetting and improving system performance. In addition, we design a multi-path integrated defense paradigm to counter potential system attacks, with a data perturbation technique based on gradient modification, a dynamically weighted selection method, and a privacy-enhanced strategy by capturing discriminative features. Moreover, to minimize parameter leakage, the parameter-decoupled hierarchical sharing mechanism is utilized, which also significantly reduces the communication overhead. The experimental results show that our approach is effective, with gender predictions down to chance levels while maintaining SER performance enhancements.
Haijiao Chen, Huan Zhao 0003, Zixing Zhang 0001, Keqin Li 0001
IEEE Internet Things J.2
2024 Gradient-Level Differential Privacy Against Attribute Inference Attack for Speech Emotion Recognition
abstract
The Federated Learning (FL) paradigm for distributed privacy preservation is valued for its ability to collaboratively train Speech Emotion Recognition (SER) models while keeping data localized. However, recent studies reveal privacy leakage in the model sharing process. Existing differential privacy schemes face increasing inference attack risks as clients expose more model updates. To address these challenges, we propose aGradient-levelHierarchicalDifferentialPrivacy (GHDP) strategy to mitigate attribute inference attacks. GHDP employs normalization to distinguish gradient importance, clipping significant gradients and filtering out sensitive information that may lead to privacy leaks. Additionally, increased random perturbations are applied to early model layers during backpropagation, achieving hierarchical differential privacy through layered noise addition. This theoretically grounded approach offers enhanced protection for critical information. Our experiments show that GHDP maintains stable SER performance while providing robust privacy protection, unaffected by the number of model updates.
Haijiao Chen, Huan Zhao 0003, Zixing Zhang 0001
IEEE Signal Process. Lett.2
2024 Refashioning Emotion Recognition Modeling: The Advent of Generalized Large Models
abstract
After its inception, emotion recognition or affective computing has increasingly become an active research topic due to its broad applications. The corresponding computational models have gradually migrated from statistically shallow models to neural-network-based deep models, which can significantly boost the performance of emotion recognition and consistently achieve the best results on different benchmarks, and thus has been considered the first option for emotion recognition. However, the debut of large language models (LLMs), such as ChatGPT and GPT4, has remarkably astonished the world due to their emerged capabilities of zero/few-shot learning, in-context learning (ICL), chain-of-thought, and others that are never shown in previous deep models. In the present article, we comprehensively investigate how the LLMs perform in emotion recognition in terms of diverse aspects, including ICL, few-shot prompting, accuracy, generalization, and explanation. Moreover, we offer some insights and pose other potential challenges, hoping to ignite broader discussions about enhancing emotion recognition in the new era of advanced and more generalized models.
Zixing Zhang 0001, Liyizhe Peng, Jing Han 0010, Huan Zhao 0003, Björn W. Schuller
IEEE Trans. Comput. Soc. Syst.5
2023 Privacy-Enhanced Federated Learning Against Attribute Inference Attack for Speech Emotion Recognition
abstract
Federal learning-based (FL) Speech Emotion Recognition (SER) framework aims to protect data privacy when characterizing emotions. However, previous studies have shown that the framework is vulnerable, because curious servers can indirectly infer user private information. To address this challenge, we propose a novel privacy- enhanced SER approach against attribute inference attack. It helps filter sensitive information and attends to highlight emotion features before uploading the shared model updates under the FL. Firstly, a bi-directional recurrent neural network captures the latent representations in sequences to discard partial redundant features. Then, a feature attention mechanism is applied to focus on the salient regions in the latent representations, further hiding emotion-irrelevant attributes. The experimental results show that the introduced model is effective. The attack capability of a gender prediction model is reduced to a chance level while retaining SER performance.
Huan Zhao 0003, Haijiao Chen, Yufeng Xiao, Zixing Zhang 0001
ICASSP1
2023 Speaker-aware Cross-modal Fusion Architecture for Conversational Emotion Recognition
Huan Zhao 0003, Zixing Zhang 0001
INTERSPEECH1
2023 Frequency Domain Feature Learning with Wavelet Transform for Image Translation
Huan Zhao 0003, Yujiang Wang 0005, Song Wang 0016, Lixuan Li, Xupeng Zha, Zixing Zhang 0001
PRICAI (3)1
2022 Cross-Modal Image-Text Matching via Coupled Projection Learning Hashing
abstract
Hashing, which aims to explore the correlations between modalities for multimedia data, shows numerous ad-vantages in cross-modal image-text matching tasks. However, most studies simply consider inter-modal correlations and ignore the semantic information contained within individual modalities, which produces suboptimal correlation representations between modalities. Based on this, we propose a Coupled Projection Learning Hashing method (CPLH). It directly maps the image-text data pairs into the different Hamming spaces to preserve as much of the original feature information as possible. Next, the CPLH joins the heterogeneous data over the Hamming spaces to maximize the correlation between two distinct modalities. In this way, the discriminative hash functions applied to the image-text matching process have semantic information from all modalities, which improves the matching accuracy. Besides, to further reduce the computational complexity of the proposed method, we propose a variant called CPLH-ts. It separates the learning process of the hash functions from the objective function by exploiting a two-step hashing strategy. Adequate experiments on three public datasets illustrate the superiority of our method and its variant compared to several state-of-the-art methods in matching accuracy and training efficiency.
Huan Zhao 0003, Haoqian Wang, Xupeng Zha, Song Wang 0016
DSAA1
2022 Coarse-to-Fine Response Generation for Document Grounded Conversations
Huan Zhao 0003, Xupeng Zha
PRICAI (1)1
2022 Deliberation Selector for Knowledge-Grounded Conversation Generation
Huan Zhao 0003, Song Wang 0016, Zixing Zhang 0001, Xupeng Zha
PRICAI (3)1
2022 ADMS: An online attack detection and mitigation system for LDoS attacks via SDN
Dan Tang 0003, Xiyin Wang, Yudong Yan, Dongshuo Zhang, Huan Zhao 0003
Comput. Commun.5
2022 Affective feature knowledge interaction for empathetic conversation generation
abstract
A popular chatbot can generate natural and human-like responses, and the crucial technology is the ability to understand and appreciate the emotions and demands expressed from the perspective of the user. However, some empathetic dialogue generation models only specialise in commonsense and neglect emotion, which can only get a one-sided understanding of the user's situation and makes the model unable to express emotion better. In this paper, we propose a novel affective feature knowledge interactive model named AFKI, to enhance response generation performance, which enriches conversation history to obtain emotional interactive context by leveraging fine-grained emotional features and commonsense knowledge. Furthermore, we utilise an emotional interactive context encoder to learn higher-level affective interaction information and distill the emotional state feature to guide the empathetic response generation. The emotional features are to well capture the subtle differences of the user's emotional expression, and the commonsense knowledge improves the representation of affective information on generated responses. Extensive experiments on the empathetic conversation task demonstrate that our model generates multiple responses with higher emotion accuracy and stronger empathetic ability compared with baseline model approaches for empathetic response generation.
Ensi Chen, Huan Zhao 0003, Xupeng Zha, Haoqian Wang, Song Wang 0016
Connect. Sci.2
2022 Cost-sensitive regression learning on small dataset through intra-cluster product favoured feature selection
abstract
Massive regression and forecasting tasks are generally cost-sensitive regression learning problems with asymmetric costs between over-prediction and under-prediction. However, existing classic methods, such as clustering and feature selection, are subject to difficulties in dealing with small datasets. As one of the key challenges, it is difficult to statistically validate the importance of features using traditional algorithms (e.g. the Boruta algorithm) owing to insufficient available data. By leveraging the feature information intra-cluster (item group with similar attributes), we propose an intra-cluster product favoured (ICPF) feature selection algorithm to select the information based on the traditional filtering method (specifically the Boruta algorithm in our study). The experimental results show that the ICPF algorithm significantly reduces the number of dimensions of the selected feature set and improves the performance of cost-sensitive regression learning. The misprediction cost decreased by 33.5% (linear-linear cost function) and 32.4% (quadratic-quadratic cost function) after adopting the ICPF algorithm. In addition, the advantage of the ICPF algorithm is robust to other regression models, such as random forest and XGboost.
Fangfang Xu, Huan Zhao 0003
Connect. Sci.2
2022 Cross-modal image-text search via Efficient Discrete Class Alignment Hashing
Song Wang 0016, Huan Zhao 0003, Yunbo Wang, Jing Huang 0012, Keqin Li 0001
Inf. Process. Manag.2
2022 Cross-domain image translation with a novel style-guided diversity loss design
Huan Zhao 0003, Jing Huang 0012, Keqin Li 0001
Knowl. Based Syst.2
2022 Discrete Joint Semantic Alignment Hashing for Cross-Modal Image-Text Search
abstract
Supervised cross-modal image-text hashing has aroused extensive concentrations in comprehending the correspondence between vision and language for data search tasks. Existing methods learn the compact hash codes by leveraging a given image-text data pairs or supervised information to explore such correspondence. However, they still confront obvious drawbacks. First, there is no engagement between multiple semantic information that yields the suboptimal search performance. Second, most of them adopt continuous relaxation strategy by discarding the discrete constraints, which results in large binary quantization errors. To deal with these problems, we propose a novel supervised hashing method, termed Discrete Joint Semantic Alignment Hashing (DJSAH). Specifically, it builds a connection between semantics (a.k.a. class labels and pairwise similarities) by the joint semantic alignment learning. And thus the high-level discriminative semantics can be preserved into the hash codes. Besides, a well-designed discrete optimization algorithm with linear computation and memory cost is developed to reduce the information loss of the hash codes with no need for relaxation. Extensive experiments and analyses on three benchmark datasets validate the superiority of the proposed DJSAH against several state-of-the-art hashing methods.
Song Wang 0016, Huan Zhao 0003, Keqin Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 Fast Discrete Matrix Factorization Hashing for Large-Scale Cross-Modal Retrieval
Huan Zhao 0003, Xiaolin She, Song Wang 0016, Kaili Ma 0002
MMM (1)1
2021 An Asymmetric Two-Sided Penalty Term for CT-GAN
Huan Zhao 0003, Yu Wang 0118
MMM (1)1
2021 Upgraded Attention-Based Local Feature Learning Block for Speech Emotion Recognition
Huan Zhao 0003, Yingxue Gao, Yufeng Xiao
PAKDD (2)1
2021 Work in Progress: Network Attack Detection Towards Smart Factory
abstract
With the continuous development of network communication and Internet of Things technology, the smart factory of new energy vehicles is increasingly dependent on network communication technology. Due to its increasing openness, which leads to increasing security risks, the attackers' system vulnerability discovery ability and attack techniques are also improving, making the security threats of smart factories escalating. To improve the autonomous sensing and defence capability of smart production lines for security vulnerabilities in the collaborative manufacturing environment, we put forward an adaptive LDoS attack detection scheme based on RF-GMM algorithm in an SDN environment. The method distinguishes normal and abnormal states of networks in smart factories by establishing a multi-feature selection model and profiling network anomalies to achieve the detection for external intrusions.
Dan Tang 0003, Dongshuo Zhang, Huan Zhao 0003, Dashun Liu, Yudong Yan
RTAS3
2021 Learning a maximized shared latent factor for cross-modal hashing
Song Wang 0016, Huan Zhao 0003, Kei Nai
Knowl. Based Syst.2
2019 Compact Convolutional Recurrent Neural Networks via Binarization for Speech Emotion Recognition
abstract
Despite the great advances, most of the recently developed automatic speech recognition systems focus on working in a server-client manner, and thus often require a high computational cost, such as the storage size and memory accesses. This, however, does not satisfy the increasing demand for a succinct model that can run smoothly in embedded devices like smartphones. To this end, in this paper we propose a neural network compression method, in the way of quantizing the weights of the neural networks from the original full-precised values into binary values that then can be stored and processed with only one bit per value. In doing this, the traditional neural network-based large-size speech emotion recognition models can be greatly compressed into smaller ones, which demand lower computational cost. To evaluate the feasibility of the proposed approach, we take a state-of-the-art speech emotion recognition model, i. e., convolutional recurrent neural networks, as an example, and conduct experiments on two widely used emotional databases. We find that the proposed binary neural networks are able to yield a remarkable model compression rate but at limited expense of model performance.
Huan Zhao 0003, Yufeng Xiao, Jing Han 0010, Zixing Zhang 0001
ICASSP1
2019 Indoor Security Localization Algorithm Based on Location Discrimination Ability of AP
Juan Luo, Huan Zhao 0003
NSS4
2018 Centroid-based KNN query in mobile service of LBS
abstract
With the increasing use of GPS-embedded mobile phone in location-based services (LBSs), people's location privacy is facing considerable challenge. Therefore, we propose a centroid- based method of K-nearest neighbor (KNN) query for LBSs. Unlike previous dummy-based approaches, our method sends a request to the LBS server that does not contain the genuine user location, and only requires the server to return the results of the client needs thus to reduce the downstream communication cost. We also propose an effective KNN(k-nearest neighbor) search algorithm to reduce the processing time of the server side . This algorithm can reduce bandwidth usages and efficiently support KNN queries without revealing the private information of the query issuer. Security analysis and empirical evaluation results verify the effectiveness and efficiency of our scheme.
Huan Zhao 0003
NOMS1
2017 A Sentiment Classification Model Using Group Characteristics of Writing Style Features
abstract
Sentiment analysis is becoming increasingly important mainly because of the growth of web comments. Sentiment polarity classification is a popular process in this field. Writing style features, such as lexical and word-based features, are often used in the authorship identification and gender classification of online messages. However, writing style features were only used in feature selection for sentiment classification. This research presents an exploratory study of the group characteristics of writing style features on the Internet Movie Database (IMDb) movie sentiment data set. Furthermore, this study utilizes the specific group characteristics of writing style in improving the performance of sentiment classification. We determine the optimum clustering number of user reviews based on writing style features distribution. According to the classification model trained on a training subset with specific writing style clustering tags, we determine that the model trained on the data set of a specific writing style group has an optimal effect on the classification accuracy, which is better than the model trained on the entire data set in a particular positive or negative polarity. Through the polarity characteristics of specific writing style groups, we propose a general model in improving the performance of the existing classification approach. Results of the experiments on sentiment classification using the IMDb data set demonstrate that the proposed model improves the performance in terms of classification accuracy.
Huan Zhao 0003, Xixiang Zhang, Keqin Li 0001
Int. J. Pattern Recognit. Artif. Intell.1
2017 Automatic syllable segmentation algorithm of Chinese speech based on MF-DFA
Shaofang He, Huan Zhao 0003
Speech Commun.2
2016 A novel activation function for multilayer feed-forward neural networks
Aboubakar Nasser Samatin Njikam, Huan Zhao 0003
Appl. Intell.2
2015 Automatic Chinese Personality Recognition Based on Prosodic Features
Huan Zhao 0003, Zeying Yang, Zuo Chen, Xixiang Zhang
MMM (1)1