Haochen Xue

dblp:349/0893 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 FedCD: Towards Consolidated Distillation for Heterogeneous Federated Learning
abstract
Knowledge Distillation (KD) serves as an effective approach to addressing heterogeneity issues in Federated Learning (FL), leveraging additional datasets to align local and global models better. There are two primary distillation paradigms: feature-based distillation, which utilizes intermediate-layer features of the network, and logit-based distillation, which employs the final layer's logit outputs. However, existing studies often select distillation methods based on intuitive and empirical evidence when facing different heterogeneous settings, neglecting the intrinsic relationship between distillation paradigms and heterogeneity. This oversight may result in suboptimal federated knowledge distillation performance under heterogeneous conditions. In this paper, we propose the Consolidated Distillation for Heterogeneous Federated Learning - FedCD that balances knowledge representations from both feature-based and logit-based distillation to enhance performance. Specifically, to address the misalignment between knowledge conveyed by features and logits, we aggregate features from different layers via cross-layer attention to preserve semantic knowledge, followed by distribution modeling using Gaussian Mixture Models. This process strengthens knowledge distillation by constraining the transformation of different network layers' features under a consolidated distribution, thereby mitigating impacts from both data and model heterogeneity. Extensive experiments demonstrate that FedCD outperforms state-of-the-art methods by over 10.72% and validate the effectiveness of our approach.
Yichen Li 0006, Huifa Li, Xinlin Zhuang, Haochen Xue, Haozhao Wang, Muhammad Imran Razzak
AAAI6
2026 CNText2Sign and CNSign: Unified Chinese Sign Language Datasets for Bidirectional Accessibility
abstract
Sign language is the primary communication mode for 72 million hearing-impaired individuals worldwide, necessitating effective bidirectional Sign Language Production and Sign Language Translation systems. However, functional bidirectional systems require a unified linguistic environment, hindered by the lack of suitable unified datasets, particularly those providing the necessary pose information for accurate Sign Language Production (SLP) evaluation. Concurrently, current SLP evaluation methods like back-translation ignore pose accuracy, and high-quality coordinated generation remains challenging. To create this crucial environment and overcome these challenges, we introduce CNText2Sign and CNSign, which together constitute the first unified dataset aimed at supporting bidirectional accessibility systems for Chinese sign language; CNText2Sign provides 15,000 natural language-to-sign mappings and standardized skeletal keypoints for 8,643 vocabulary items supporting pose assessment. Building upon this foundation, we propose the AuraLLM model, which leverages a decoupled architecture with CNText2Sign's pose data for novel direct gesture accuracy assessment. The model employs retrieval augmentation and Cascading Vocabulary Resolution to handle semantic mapping and out-of-vocabulary words, and achieves all-scenario production with controllable coordination of gestures and facial expressions via pose-conditioned video synthesis. Concurrently, our Sign Language Translation model SignMST-C employs targeted self-supervised pretraining for dynamic feature capture, achieving new SOTA results on PHOENIX2014-T with BLEU-4 scores up to 32.08. AuraLLM establishes a strong performance baseline on CNText2Sign with a BLEU-4 score of 50.41 under direct evaluation.
Yulong Li 0002, Zhixiang Lu, Haochen Xue, Jianghao Wu 0001, Mian Zhou, Kang Dang, Yifang Wang 0006, Muhammad Imran Razzak, Jionglong Su
KDD (1)6
2026 Rhythm of Opinion: Interpretable Hawkes-Graph Networks for Hierarchical Opinion Propagation
Yulong Li 0002, Zhixiang Lu, Peixin Guo, Simin Lai, Haochen Xue, Xiwei Liu, Yichen Li 0006, Zhaodong Wu, Mian Zhou, Muhammad Imran Razzak, Qingxia Li, Jionglong Su
WWW6
2025 MMRC: A Large-Scale Benchmark for Understanding Multimodal Large Language Model in Real-World Conversation
abstract
Recent multimodal large language models (MLLMs) have demonstrated significant potential in open-ended conversation, generating more accurate and personalized responses. However, their abilities to memorize, recall, and reason in sustained interactions within real-world scenarios remain underexplored. This paper introduces MMRC, a Multi-Modal Real-world Conversation benchmark for evaluating six core open-ended abilities of MLLMs: information extraction, multi-turn reasoning, information update, image management, memory recall, and answer refusal. With data collected from real-world scenarios, MMRC comprises 5,120 conversations and 28,720 corresponding manually labeled questions, posing a significant challenge to existing MLLMs. Evaluations on 20 MLLMs in MMRC indicate an accuracy drop during open-ended interactions. We identify four common failure patterns: long-term memory degradation, inadequacies in updating factual knowledge, accumulated assumption of error propagation, and reluctance to “say no.” To mitigate these issues, we propose a simple yet effective NOTE-TAKING strategy, which can record key information from the conversation and remind the model during its responses, enhancing conversational capabilities. Experiments across six MLLMs demonstrate significant performance improvements.
Haochen Xue, Yexin Liu, Qidong Huang, Yulong Li 0002, Zhongxing Xu, Chong Zhang 0006, Yutong Xie 0001, Muhammad Imran Razzak, ZongYuan Ge, Jionglong Su, Junjun He, Yu Qiao 0001
ACL (1)1
2025 Decoding the Flow: CauseMotion for Emotional Causality Analysis in Long-form Conversations
abstract
Long-sequence causal reasoning seeks to uncover causal relationships within extended time series data but is hindered by complex dependencies and the challenges of validating causal links. To address the limitations of large-scale language models (e.g., GPT-4) in capturing intricate emotional causality within extended dialogues, we propose CauseMotion, an innovative framework combining emotional causal dynamic mapping with multimodal feature fusion. CauseMotion implements dynamic mapping through a sliding window mechanism and fusion strategies, while integrating audio features—vocal emotion, intensity, and speech rate—to enrich semantic representations. This design enables efficient retrieval of contextually relevant information and precise inference of emotional causal chains spanning multiple conversational turns. We constructed the first benchmark dataset for long-sequence emotional causal reasoning, featuring dialogues with over 70 turns. Experimental results show that CauseMotion significantly enhances emotional understanding and causal inference capabilities in large language models. A GLM-4 integrated with CauseMotion achieves an 8.7% improvement in causal accuracy over the original model and surpasses GPT-4o by 1.2%. On the DiaASQ dataset, CauseMotion-GLM-4 achieves state-of-the-art results in accuracy, F1 score, and causal reasoning accuracy.
Yulong Li 0002, Zichen Yu, Zhixiang Lu, Haochen Xue, Zhaodong Wu, Kang Dang, Muhammad Imran Razzak, Jionglong Su
AVSS7
2025 Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
abstract
Recent advancements in multimodal large language models (MLLMs) have significantly improved performance in visual question answering. However, they often suffer from hallucinations. In this work, hallucinations are categorized into two main types: initial hallucinations and snowball hallucinations. We argue that adequate contextual information can be extracted directly from the token interaction process. Inspired by causal inference in the decoding strategy, we propose to leverage causal masks to establish information propagation between multimodal tokens. The hypothesis is that insufficient interaction between those tokens may lead the model to rely on outlier tokens, overlooking dense and rich contextual cues. Therefore, we propose to intervene in the propagation process by tackling outlier tokens to enhance in-context inference. With this goal, we present FarSight, a versatile plug-and-play decoding strategy to reduce attention interference from outlier tokens merely by optimizing the causal mask. The heart of our method is effective token propagation. We design an attention register structure within the upper triangular matrix of the causal mask, dynamically allocating attention to capture attention diverted to outlier tokens. Moreover, a positional awareness encoding method with a diminishing masking rate is proposed, allowing the model to attend to further preceding tokens, especially for video sequence tasks. With extensive experiments, FarSight demonstrates significant hallucination-mitigating performance across different MLLMs on both image and video benchmarks, proving its effectiveness.
Zhongxing Xu, Zile Huang, Haochen Xue, Ziyang Chen 0003, Zelin Peng, Sijin Zhou, Wenxue Li 0003, Yulong Li 0002, Wenxuan Song, Shiyan Su, Wei Feng 0015, Jionglong Su, Mingquan Lin, Yifan Peng 0002, Xuelian Cheng, Muhammad Imran Razzak, ZongYuan Ge
CVPR6
2025 Advancing Low-Resource Machine Translation: A Unified Data Selection and Scoring Optimization Framework
Zhixiang Lu, Peichen Ji, Yulong Li 0002, Ding Sun, Chenyu Xue 0002, Haochen Xue, Mian Zhou, Angelos Stefanidis, Jionglong Su, Zhengyong Jiang
ICIC (24)6
2025 MSWAL: 3D Multi-class Segmentation of Whole Abdominal Lesions Dataset
Zhaodong Wu, Qiaochu Zhao, Yulong Li 0002, Haochen Xue, Zhengyong Jiang, Angelos Stefanidis, Muhammad Imran Razzak, ZongYuan Ge, Junjun He, Yu Qiao 0001, Kang Dang, Jionglong Su
MICCAI (2)5
2025 Genesis: A Large-Scale Benchmark for Multimodal Large Language Model in Emotional Causality Analysis
Yulong Li 0002, Zhixiang Lu, Jianghao Wu 0001, Haochen Xue, Mian Zhou, Jionglong Su, Muhammad Imran Razzak
ACM Multimedia8
2025 Two-stage effective attentional generative adversarial network
Mingyu Jin, Qinkai Yu, Chong Zhang 0006, Haochen Xue
Appl. Intell.4
2024 Bridging the Projection Gap: Overcoming Projection Bias Through Parameterized Distance Learning
Chong Zhang 0006, Mingyu Jin, Qinkai Yu, Haochen Xue, Shreyank N. Gowda, Xiao-Bo Jin
ACCV (8)4
2024 Goal-Guided Generative Prompt Injection Attack on Large Language Models
abstract
Current large language models (LLMs) provide a strong foundation for large-scale user-oriented natural language tasks. Numerous users can easily inject adversarial text or instructions through the user interface, thus causing LLM model security challenges. Although there is much research on prompt injection attacks, most black-box attacks use heuristic strategies. It is unclear how these heuristic strategies relate to the success rate of attacks and thus effectively improve model robustness. To solve this problem, we redefine the goal of the attack: to maximize the KL divergence between the conditional probabilities of the clean text and the adversarial text. Furthermore, we prove that maximizing the KL divergence is equivalent to maximizing the Mahalanobis distance between the embedded representation$x$and$x^{\prime}$of the clean text and the adversarial text when the conditional probability is a Gaussian distribution and gives a quantitative relationship on$x$and$x^{\prime}$. Then we designed a simple and effective goal-guided generative prompt injection strategy (G2PIA) to find an injection text that satisfies specific constraints to achieve the optimal attack effect approximately. Notably, our attack method is a query-free black-box attack method with a low computational cost. Experimental results on seven LLM models and four datasets show the effectiveness of our attack method.
Chong Zhang 0006, Mingyu Jin, Qinkai Yu, Haochen Xue, Xiao-Bo Jin
ICDM5
2024 Multi-task Prompt Words Learning for Social Media Content Generation
abstract
The rapid development of the Internet has profoundly changed human life. Humans are increasingly expressing themselves and interacting with others on social media platforms. However, although artificial intelligence technology has been widely used in many aspects of life, its application in social media content creation is still blank. To solve this problem, we propose a new prompt word generation framework based on multi-modal information fusion, which combines multiple tasks including topic classification, sentiment analysis, scene recognition and keyword extraction to generate more comprehensive prompt words. Subsequently, we use a template containing a set of prompt words to guide ChatGPT to generate high-quality tweets. Furthermore, in the absence of effective and objective evaluation criteria in the field of content generation, we use the ChatGPT tool to evaluate the results generated by the algorithm, making large-scale evaluation of content generation algorithms possible. Evaluation results on extensive content generation demonstrate that our cue word generation framework generates higher quality content compared to manual methods and other cueing techniques, while topic classification, sentiment analysis, and scene recognition significantly enhance content clarity and its consistency with the image.
Haochen Xue, Chong Zhang 0006, Chenzhi Liu, Fangyu Wu 0001, Xiao-Bo Jin
IJCNN1
2023 Image Blending Algorithm with Automatic Mask Generation
Haochen Xue, Mingyu Jin, Chong Zhang 0006, Qian Weng, Xiao-Bo Jin
ICONIP (8)1