Xupeng Zha

dblp:275/0532 · DBLP profile ↗
← Back
11ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0001-6928-9998ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PLUM-Net: Prototype-Induced Label Structuring for Disentangled Multimodal Representation Network
abstract
Existing multimodal representation learning approaches often rely on simple feature concatenation or unified transformations, which fail to effectively disentangle and leverage common and private information across different modalities in a progressive manner. Moreover, they typically lack adaptive modeling tailored to specific task requirements. To address these limitations, we propose a Prototype-Induced Label Structuring for Disentangled Multimodal Representation Network (PLUM-Net). It first employs a multilevel semantic alignment module to synchronize global and local semantics across audio, visual and textual streams. On this aligned foundation, a prototype-based single-modal label generation module derives modality-specific hard and soft-labels that subtly steer the network toward a cleaner split between shared and private cues. Guided by these labels, the task-conditioned feature bifurcator module channels information through the most beneficial common or private pathway for the given task, after which a private refinement module polishes and fuses each modality’s idiosyncratic signals. Extensive experiments show that PLUM-Net delivers strong performance on datasets such as CMU-MOSI, CMU-MOSEI and UR-FUNNY, achieving an ACC-2 of 90.3% on CMU-MOSI, representing a 2%–4% improvement over previous SOTA models.
Huan Zhao 0003, Xupeng Zha, Guanghui Ye, Zixing Zhang 0001
AAAI4
2026 Making Visual Dialogue More Engaging: A New Task, Method, and Metric
abstract
Large language model (LLM)-based visual dialogue (VD) systems have made response generation for image-grounded conversations more correct and coherent. However, user engagement - the extent to which a user is interested, emotionally involved, and willing to continue the conversation - remains a challenge. To fully explore engaging VD, we propose: (i) a new task named Audio-enhanced VD (AVD), which introduces additional audio dialogue contexts that can more vividly convey the speaker's emotions as input, with the aim of generating correct but more engaging dialogue responses. Specifically, we employ a text-to-speech model as the modality translator to generate the paired acoustic utterances from the inputting textual utterances; (ii) an accompanying approach named Visually-grounded and Interleaved Text-Audio Dialogue Modeling (VITA-DM), which utilizes both image-grounded information and interleaved text-audio utterances for visual dialogue modeling, differentiating from previous multi-modal LLM (MLLM)-based methods that normally model text and audio modalities separately. We also present three pre-training tasks to better learn multi-modal interactions across language, vision, and audio; (iii) a novel metric named Multi-Modal Engagement (MME), which fills the gap of engagement estimation in VD and can provide a fine-grained assessment along emotional, attentional, and reply engagement dimensions (EE, AE, RE). We experiment on two popular datasets and provide extensive evaluations (automatic, engagement-specific, and human), supporting the validity of our approach. Furthermore, based on empirical results that reveal that emotions contribute the most to engagement, we justify our emphasis on the emotional aspect throughout the definition, solution, and evaluation of our task.
Guanghui Ye, Huan Zhao 0003, Yingxue Gao, Zhixue Zhao, Xupeng Zha, Zhihua Jiang
AAAI6
2025 Dual-View Learning for Conversational Emotion Recognition Through Context and Emotion-Shift Modeling
abstract
Conversational Emotion Recognition (CER) has recently been explored through conversational context modeling to learn the emotion distribution, i.e., the likelihood over emotion categories associated with each utterance. While these methods have shown promising results in emotion classification, they often focus on the interactions between utterances (utterance-view) and overlook shifts in the speaker's emotions (emotion-view). This emphasis on homogeneous view modeling limits their overall effectiveness. To address this limitation, we propose DVL-CER, a novel Dual-View Learning approach for CER. DVL-CER integrates both the utterance-view and emotion-view using two projection heads, enabling cross-view projection of emotion distributions. Our approach offers several key advantages: (1) We introduce an emotion-view that captures shifts in a speaker's emotions from initial to subsequent states within a conversation. This view enriches the conversation modeling and supports seamless integration with various CER baseline models. (2) Our dual-view projection learning strategy flexibly balances consistency and independence between the two heterogeneous views, promoting view-specific adaptation learning and incorporating the emotion verification capability within CER. We validate DVL-CER through extensive experiments on two widely-used datasets, IEMOCAP and EmoryNLP. The results demonstrate that DVL-CER achieves state-of-the-art performance, delivering robust and high-quality emotion distributions compared with existing CER methods and other dual-view learning strategies.
Xupeng Zha, Huan Zhao 0003, Guanghui Ye, Zixing Zhang 0001
AAAI1
2025 Knowledge Image Matters: Improving Knowledge-Based Visual Reasoning with Multi-Image Large Language Models
abstract
We revisit knowledge-based visual reasoning (KB-VR) in light of modern advances in multimodal large language models (MLLMs), and make the following contributions: (i) We propose Visual Knowledge Card (VKC) -a novel image that incorporates not only internal visual knowledge (e.g., scene-aware information) detected from the raw image, but also external world knowledge (e.g., attribute or object knowledge) produced by a knowledge generator; (ii) We present VKC-enhanced Multi-Image Reasoning (VKC-MIR) -a fourstage pipeline which harnesses a state-of-theart scene perception engine to construct an initial VKC (Stage-1), a powerful LLM to generate relevant domain knowledge (Stage-2), an excellent image editing toolkit to introduce generated knowledge into an iteratively-edited VKC (Stage-3), and finally, an emerging multiimage MLLM to solve the VKC-enhanced task (Stage-4).By performing experiments on three popular KB-VR benchmarks, our approach achieves new state-of-the-art results compared to previous top-performing models.Our code is available at: https://github. com/yyy1103/VKC.
Guanghui Ye, Huan Zhao 0003, Zhixue Zhao, Xupeng Zha, Zhihua Jiang
ACL (1)4
2024 Esihgnn: Event-State Interactions Infused Heterogeneous Graph Neural Network for Conversational Emotion Recognition
abstract
Conversational Emotion Recognition (CER) aims to predict the emotion expressed by an utterance (referred to as an "event") during a conversation. Existing graph-based methods mainly focus on event interactions to comprehend the conversational context, while overlooking the direct influence of the speaker’s emotional state on the events. In addition, real-time modeling of the conversation is crucial for real-world applications but is rarely considered. Toward this end, we propose a novel graph-based approach, namely Event-State Interactions infused Heterogeneous Graph Neural Network (ESIHGNN), which incorporates the speaker’s emotional state and constructs a heterogeneous event-state interaction graph to model the conversation. Specifically, a heterogeneous directed acyclic graph neural network is employed to dynamically update and enhance the representations of events and emotional states at each turn, thereby improving conversational coherence and consistency. Furthermore, to further improve the performance of CER, we enrich the graph’s edges with external knowledge. Experimental results on four publicly available CER datasets show the superiority of our approach and the effectiveness of the introduced heterogeneous event-state interaction graph.
Xupeng Zha, Huan Zhao 0003, Zixing Zhang 0001
ICASSP1
2024 LSTDial: Enhancing Dialogue Generation via Long- and Short-Term Measurement Feedback
abstract
Guanghui Ye, Huan Zhao, Zixing Zhang, Xupeng Zha, Zhihua Jiang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Guanghui Ye, Huan Zhao 0003, Zixing Zhang 0001, Xupeng Zha, Zhihua Jiang
NAACL-HLT4
2023 Frequency Domain Feature Learning with Wavelet Transform for Image Translation
Huan Zhao 0003, Yujiang Wang 0005, Song Wang 0016, Lixuan Li, Xupeng Zha, Zixing Zhang 0001
PRICAI (3)6
2022 Cross-Modal Image-Text Matching via Coupled Projection Learning Hashing
abstract
Hashing, which aims to explore the correlations between modalities for multimedia data, shows numerous ad-vantages in cross-modal image-text matching tasks. However, most studies simply consider inter-modal correlations and ignore the semantic information contained within individual modalities, which produces suboptimal correlation representations between modalities. Based on this, we propose a Coupled Projection Learning Hashing method (CPLH). It directly maps the image-text data pairs into the different Hamming spaces to preserve as much of the original feature information as possible. Next, the CPLH joins the heterogeneous data over the Hamming spaces to maximize the correlation between two distinct modalities. In this way, the discriminative hash functions applied to the image-text matching process have semantic information from all modalities, which improves the matching accuracy. Besides, to further reduce the computational complexity of the proposed method, we propose a variant called CPLH-ts. It separates the learning process of the hash functions from the objective function by exploiting a two-step hashing strategy. Adequate experiments on three public datasets illustrate the superiority of our method and its variant compared to several state-of-the-art methods in matching accuracy and training efficiency.
Huan Zhao 0003, Haoqian Wang, Xupeng Zha, Song Wang 0016
DSAA3
2022 Coarse-to-Fine Response Generation for Document Grounded Conversations
Huan Zhao 0003, Xupeng Zha
PRICAI (1)4
2022 Deliberation Selector for Knowledge-Grounded Conversation Generation
Huan Zhao 0003, Song Wang 0016, Zixing Zhang 0001, Xupeng Zha
PRICAI (3)6
2022 Affective feature knowledge interaction for empathetic conversation generation
abstract
A popular chatbot can generate natural and human-like responses, and the crucial technology is the ability to understand and appreciate the emotions and demands expressed from the perspective of the user. However, some empathetic dialogue generation models only specialise in commonsense and neglect emotion, which can only get a one-sided understanding of the user's situation and makes the model unable to express emotion better. In this paper, we propose a novel affective feature knowledge interactive model named AFKI, to enhance response generation performance, which enriches conversation history to obtain emotional interactive context by leveraging fine-grained emotional features and commonsense knowledge. Furthermore, we utilise an emotional interactive context encoder to learn higher-level affective interaction information and distill the emotional state feature to guide the empathetic response generation. The emotional features are to well capture the subtle differences of the user's emotional expression, and the commonsense knowledge improves the representation of affective information on generated responses. Extensive experiments on the empathetic conversation task demonstrate that our model generates multiple responses with higher emotion accuracy and stronger empathetic ability compared with baseline model approaches for empathetic response generation.
Ensi Chen, Huan Zhao 0003, Xupeng Zha, Haoqian Wang, Song Wang 0016
Connect. Sci.4