EDBT 2026 Demo / reviewers in the wild / expert
Xianwei Zhuang
dblp:339/2318
· DBLP profile ↗
22ranked-venue papers
10as first author
22since 2021 · last 2026
0009-0004-4392-6126ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 8 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination MitigationabstractLarge vision-language models (LVLMs) have demonstrated impressive capabilities across diverse multimodal tasks, yet they remain highly susceptible to visual hallucinations (VH), often producing confident but inaccurate descriptions of visual content. Building on the insight that not all tokens and attention heads contribute equally to VH mitigation, we introduce VisFlow, a lightweight and training-free framework that alleviates hallucinations by directly modulating attention patterns during inference. To address two primary challenges of VH, namely insufficient visual attention and the dominance of language priors, we identify three problematic attention behaviors in LVLMs: (1) disproportionate allocation of attention to uninformative or trailing visual tokens, (2) over-dependence on the previously generated token, and (3) excessive fixation on system prompts that hinders multimodal integration. To overcome these issues, VisFlow introduces a dual-level Attention Intervention, consisting of Token-level Attention Intervention (TAI), which reinforces attention to salient visual regions, and Head-level Attention Intervention (HAI), which suppresses undue focus on system prompts and adjacent text tokens. Together, these interventions strengthen visual alignment while reducing linguistic bias. Extensive experiments across diverse models and benchmarks demonstrate that VisFlow effectively mitigates hallucinations with minimal computational overhead. Lexiang Tang, Xianwei Zhuang, Bang Yang, Jinghan Ru, Yuexian Zou |
AAAI | 2 |
| 2025 | ATRI: Mitigating Multilingual Audio Text Retrieval Inconsistencies by Reducing Data Distribution ErrorsabstractMultilingual audio-text retrieval (ML-ATR) is a challenging task that aims to retrieve audio clips or multilingual texts from databases. However, existing ML-ATR schemes suffer from inconsistencies for instance similarity matching across languages. To address the inconsistency issue in multilingual audio-text retrieval, we first identify two intuitive factors that contribute to inconsistency: misalignment between audio and multilingual text embeddings, and error propagation in model optimization. By systematically analyzing these factors, we derive theoretical weight error upper bounds for quantifying their effects and find that the main source of inconsistency is the data distribution error during training. This finding motivates our solution to reduce data distribution errors.We propose a consistent ML-ATR scheme using 1-to-k contrastive learning and audio-English co-anchor contrastive learning, aiming to mitigate the negative impact of data distribution error on recall and consistency in ML-ATR. Experimental results on the translated AudioCaps and Clotho datasets show that our scheme achieves state-of-the-art performance on recall and consistency metrics for eight mainstream languages, including English. Our code will be available at https://github.com/ATRI-ACL/ATRI-ACL. Yuguo Yin, Yuxin Xie 0004, Dongchao Yang, Jinghan Ru, Xianwei Zhuang, Yuexian Zou |
ACL (1) | 6 |
| 2025 | VASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token SparsificationabstractLarge Vision-Language Models (LVLMs) may produce outputs that are unfaithful to reality, also known as visual hallucinations (VH), which significantly impedes their real-world usage. To alleviate VH, various decoding strategies have been proposed to enhance visual information. However, many of these methods may require secondary decoding and rollback, which significantly reduces inference speed. In this work, we propose an efficient plug-and-play decoding algorithm via Visual-Aware Sparsification (VASparse) from the perspective of token sparsity for mitigating VH. VAS-parse is inspired by empirical observations: (1) the sparse activation of attention in LVLMs, and (2) visual-agnostic tokens sparsification exacerbates VH. Based on these insights, we propose a novel token sparsification strategy that balances efficiency and trustworthiness. Specifically, VASparse implements a visual-aware token selection strategy during decoding to reduce redundant tokens while preserving visual context effectively. Additionally, we innovatively introduce a sparse-based visual contrastive decoding method to recalibrate the distribution of hallucinated outputs without the time overhead associated with secondary decoding. Subsequently, VASparse recalibrates attention scores to penalize attention sinking of LVLMs towards text tokens. Extensive experiments across four popular benchmarks confirm the effectiveness of VASparse in mitigating VH across different LVLM families without requiring additional training or post-processing. Impressively, VASparse achieves state-of-the-art performance for mitigating VH while maintaining competitive decoding speed. Code is available at https://github.com/mengchuang123/VASparse-github. Xianwei Zhuang, Zhihong Zhu 0001, Yuxin Xie 0004, Yuexian Zou |
CVPR | 1 |
| 2025 | HCoTT: Hierarchical Chain-of-Thought DistillationabstractChains of Thought (CoT) have shown potential in augmenting the reasoning capabilities of language models, yet their effectiveness is predominantly observed in large language models (LLMs). Recently, several attempts have been made to inject CoT into small language models (SLMs) using distillation and achieved promising results. However, current methods (1) ignore the rationality and hierarchical logic of reasoning when constructing CoT; (2) fail to inject hierarchical reasoning priors into SLMs. In this paper, we design a Hierarchical CoT distillation framework termed HCoTT, whose core component is a hierarchical recursive sampling module and a hierarchical learning module. Specifically, hierarchical recursive sampling utilizes a hierarchical logic process to generate more diverse explanations and a Hierarchical Chain of Thought (HCoT). Furthermore, hierarchical learning encompasses hierarchical supervision and representation learning, which is designed to augment learning and representation of implicit explanatory priors in HCoT for SLMs. Experimental results show that HCoTT can effectively improve the performance of SLMs on Faculty-Reasoning and Multiple-Choice QA tasks. More impressively, our method is model-independent and can consistently improve performance with existing language model fusions of different scales. Zhichang Wang, Xianwei Zhuang, Zhihong Zhu 0001, Yuexian Zou |
ICASSP | 2 |
| 2025 | UniCoTT: A Unified Framework for Structural Chain-of-Thought DistillationabstractChains of thought (CoTs) have achieved success in enhancing the reasoning capabilities of large language models (LLMs), while their effectiveness is predominantly observed in LLMs. Existing solutions methods adopt distillation to inject chain-of-thought capabilities into small models (SLMs). However, they: (1) can not guarantee the rationality of the generated explanation due to hallucinations; (2) ignore diverse structures of CoT during knowledge transfer. In this paper, we propose a unified CoT distillation framework termed UniCoTT for considering diverse structural CoTs (\emph{i.e.}, chain, tree, and graph). UniCoTT contains two core strategies: iterative construction for structured CoTs and the structural constraint strategy. Specifically, UniCoTT prompts LLMs to iteratively produce accurate explanations with answers and unifies structured explanations as UniCoT which is seen as a bridge for knowledge transfer. Furthermore, UniCoTT utilizes the proposed unified supervised learning and structural consistency learning strategies to transfer knowledge of structured CoT to SLMs. Experimental results show that UniCoTT can significantly improve the performance of SLMs on multiple datasets across different NLP tasks. Our code is available at https://github.com/mengchuang123/UniCoTT. Xianwei Zhuang, Zhihong Zhu 0001, Zhichang Wang, Xuxin Cheng, Yuexian Zou |
ICLR | 1 |
| 2025 | FoleyMaster: High-Quality Video-to-Audio Synthesis via MLLM-Augmented Prompt Tuning and Joint Semantic-Temporal Adaptation
Yuehan Jin, Xianwei Zhuang, Yongkang Yin, Yuexian Zou |
INTERSPEECH | 4 |
| 2025 | SpeechSEC: A Unified Multi-Task Framework for Speech Synthesis, Editing, and Continuation
Dongchao Yang, Xianwei Zhuang, Yuxin Xie 0004, Yuehan Jin, Yuexian Zou |
INTERSPEECH | 3 |
| 2025 | SemiGMMPoint: Semi-supervised point cloud segmentation based on Gaussian mixture models
Xianwei Zhuang, Hualiang Wang, Xiaoxuan He, Siming Fu, Haoji Hu |
Pattern Recognit. | 1 |
| 2024 | Towards Multi-Intent Spoken Language Understanding via Hierarchical Attention and Optimal TransportabstractMulti-Intent spoken language understanding (SLU) can handle complicated utterances expressing multiple intents, which has attracted increasing attention from researchers. Although existing models have achieved promising performance, most of them still suffer from two leading problems: (1) each intent has its specific scope and the semantic information outside the scope might potentially hinder accurate predictions, i.e. scope barrier; (2) only the guidance from intent to slot is modeled but the guidance from slot to intent is often neglected, i.e. unidirectional guidance. In this paper, we propose a novel Multi-Intent SLU framework termed HAOT, which utilizes hierarchical attention to divide the scopes of each intent and applies optimal transport to achieve the mutual guidance between slot and intent. Experiments demonstrate that our model achieves state-of-the-art performance on two public Multi-Intent SLU datasets, obtaining the 3.4 improvement on MixATIS dataset compared to the previous best models in overall accuracy. Xuxin Cheng, Zhihong Zhu 0001, Hongxiang Li 0004, Yaowei Li 0001, Xianwei Zhuang, Yuexian Zou |
AAAI | 5 |
| 2024 | Towards Explainable Joint Models via Information Theory for Multiple Intent Detection and Slot FillingabstractRecent joint models for multi-intent detection and slot filling have obtained promising results through modeling the unidirectional or bidirectional guidance between intent and slot. However, existing works design joint models heuristically and lack some theoretical exploration, including (1) theoretical measurement of the joint-interaction quality; (2) explainability of design and optimization methods of joint models, which may limit the performance and efficiency of designs. In this paper, we mathematically define the cross-task information gain (CIG) to measure the quality of joint processes from an information-theoretic perspective and discover an implicit optimization of CIG in previous models. Based on this, we propose a novel multi-stage iterative framework with theoretical effectiveness, explainability, and convergence, which can explicitly optimize information for cross-task interactions. Further, we devise an information-based joint model (InfoJoint) that conforms to this theoretical framework to gradually reduce the cross-task propagation of erroneous semantics through CIG iterative maximization. Extensive experiment results on two public datasets show that InfoJoint outperforms the state-of-the-art models by a large margin. Xianwei Zhuang, Xuxin Cheng, Yuexian Zou |
AAAI | 1 |
| 2024 | PCAD: Towards ASR-Robust Spoken Language Understanding via Prototype Calibration and Asymmetric DecouplingabstractXianwei Zhuang, Xuxin Cheng, Liming Liang, Yuxin Xie, Zhichang Wang, Zhiqi Huang, Yuexian Zou. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Xianwei Zhuang, Xuxin Cheng, Yuxin Xie 0004, Zhichang Wang, Zhiqi Huang 0001, Yuexian Zou |
ACL (1) | 1 |
| 2024 | Uncertainty-Aware Sign Language Video Retrieval with Probability Distribution Modeling
Yuanjiang Luo, Xuxin Cheng, Xianwei Zhuang, Keren Fu |
ECCV (44) | 5 |
| 2024 | KDProR: A Knowledge-Decoupling Probabilistic Framework for Video-Text Retrieval
Xianwei Zhuang, Hongxiang Li 0004, Xuxin Cheng, Zhihong Zhu 0001, Yuxin Xie 0004, Yuexian Zou |
ECCV (34) | 1 |
| 2024 | Relevance Is a Guiding Light: Relevance-aware Adaptive Learning for End-to-end Task-oriented Dialogue SystemabstractRetrieving accurate domain knowledge and providing helpful information are crucial in developing an effective end-to-end task-oriented dialogue system (E2ETOD).Existing approaches to this field follow a retrieve-then-generate paradigm and train their systems on one specific domain.However, existing approaches still suffer from the Distractive Attributes Problem (DAP): struggling to deal with false but similar knowledge (a.k.a hard negative entities), which is even more intractable when countless pieces of knowledge from different domains are blended in a real-world scenario.To alleviate DAP, we propose the Relevanceaware Adaptive Learning (ReAL), a novel twostage training framework that eliminates hard negatives step-by-step and aligns retrieval with generation.In the first stage, we introduce a top-k adaptive contrastive loss and utilize the divergence-driven feedback from the frozen generator to pre-train the retriever.In the second stage, we propose using the metric score distribution as an anchor to align retrieval with generation.Thorough experiments on three benchmark datasets demonstrate ReAL's superiority over existing methods, with extensive analysis validating its strong capabilities of overcoming in-and cross-domain distractions. Zhanpeng Chen, Zhihong Zhu 0001, Wanshi Xu, Xianwei Zhuang, Yuexian Zou |
EMNLP | 4 |
| 2024 | Dual-oriented Disentangled Network with Counterfactual Intervention for Multimodal Intent DetectionabstractMultimodal intent detection leverages diverse modalities for a comprehensive understanding of user intentions in real-world scenarios, playing a critical role in modern taskoriented dialogue systems.While existing methods have made progress in modal alignment and fusion, they overlook two vital limitations: (I) Close entanglement of multimodal semantics with modal structures; (II) Insufficient learning of the causal effects of semantic and modality-specific information on final predictions in end-to-end training.To address these limitations, we introduce the Dualoriented Disentangled Network with Counterfactual Intervention (DuoDN).DuoDN consists of a Dual-oriented Disentangled Encoder that decouples semantics-and modality-oriented representations, and a Counterfactual Intervention Module that uses causal inference to understand causal effects by injecting confounders.Experiments on three benchmark datasets demonstrate DuoDN's superiority over existing methods, with extensive analysis validating its advantages. Zhanpeng Chen, Zhihong Zhu 0001, Xianwei Zhuang, Zhiqi Huang 0001, Yuexian Zou |
EMNLP | 3 |
| 2024 | What are the Generator Preferences for End-to-end Task-Oriented Dialog System?abstractFully end-to-end task-oriented dialogue (EToD) systems have shown excellent performance, which requires the ability to retrieve entities accurately for generation.Existing methods improve the accuracy of entity retrieval and construct data flows between retrieval results and response generator, achieving promising results.However, most of them suffer from the following issues: (i) The entity is retrieved by directly interacting with the context at a coarse-grained level, so the similarity score may be disturbed by irrelevant attributes; (ii) The generator pays equal attention to retrieved entities and the context and does not learn the generation preferences for the current turn.In this paper, we propose a framework called Regulating Preferences of Generator (RPG) based on retrieval results, which includes a generator preference extractor, an entity retriever, and a generator with the gate-controlled preference regulator.The generator preference extractor not only improves the entity retriever by filtering the interference of irrelevant attributes but also provides more focused guidance to the generator by performing inter-turn attribute prediction.Experiments and analyses on three standard benchmarks show that our framework outperforms existing methods and improves the quality of the dialogue. Wanshi Xu, Xianwei Zhuang, Zhanpeng Chen, Zhihong Zhu 0001, Xuxin Cheng, Yuexian Zou |
EMNLP | 2 |
| 2024 | Game on Tree: Visual Hallucination Mitigation via Coarse-to-Fine View Tree and Game TheoryabstractLarge Vision-Language Models (LVLMs) may produce outputs that are unfaithful to reality, also known as visual hallucinations (VH), which hinders their application in multimodal understanding and decision-making.In this work, we introduce a novel plug-and-play trainfree decoding algorithm named Game and Tree based Hallucination Mitigation (GTHM), designed for mitigating VH.GTHM is inspired by empirical observations that the fuzziness of multi-granularity view perception exacerbates VH.Based on this, GTHM leverages visual information to construct a coarse-to-fine visual view tree (CFTree) that organizes visual objects, attributes, and relationships in a hierarchical manner.Additionally, we innovatively model the optimal visual-token matching process on the CFTree as the cooperative game.Specifically, we define the Tree-based Shapley Value (TSV) for each visual view on the CFTree to assess its significant contribution to the overall visual understanding, thereby determining the optimal visual granularity.Subsequently, we utilize the TSV as guidance to implement adaptive weight contrastive decoding to achieve vision-aware decoding.Extensive experiments on four popular benchmarks confirm the effectiveness of our GTHM in alleviating VH across different LVLM families without additional training or post-processing.Our code is published at https://github.com/ mengchuang123/GTHM. Xianwei Zhuang, Zhihong Zhu 0001, Zhanpeng Chen, Yuxin Xie 0004, Yuexian Zou |
EMNLP | 1 |
| 2024 | TFCD: Towards Multi-modal Sarcasm Detection via Training-Free Counterfactual Debiasing
Zhihong Zhu 0001, Xianwei Zhuang, Yunyan Zhang, Derong Xu, Guimin Hu, Xian Wu 0001, Yefeng Zheng 0001 |
IJCAI | 2 |
| 2024 | GPA: Global and Prototype Alignment for Audio-Text Retrieval
Yuxin Xie 0004, Zhihong Zhu 0001, Xianwei Zhuang, Zhichang Wang, Yuexian Zou |
INTERSPEECH | 3 |
| 2024 | Towards Multimodal-augmented Pre-trained Language Models via Self-balanced Expectation-Maximization IterationabstractPre-trained language models (PLMs) that rely solely on textual corpus may present limitations in multimodal semantics comprehension. Existing studies attempt to alleviate this issue by incorporating additional modal information through image retrieval or generation. However, these methods: (1) inevitably encounter modality gaps and noise; (2) treat all modalities indiscriminately; and (3) ignore visual or acoustic semantics of key entities. To tackle these challenges, we propose a novel principled iterative framework for multimodal-augmented PLMs termed MASE, which achieves efficient and balanced injection of multimodal semantics under the proposed Expectation Maximization (EM) based iterative algorithm. Initially, MASE utilizes multimodal proxies instead of explicit data to enhance PLMs, which avoids noise and modality gaps. In E-step, MASE adopts a novel information-driven self-balanced strategy to estimate allocation weights. Furthermore, MASE employs heterogeneous graph attention to capture entity-level fine-grained semantics on the proposed multimodal-semantic scene graph. In M-step, MASE injects global multimodal knowledge into PLMs through a cross-modal contrastive loss. Experimental results show that MASE consistently outperforms competitive baselines on multiple tasks across various architectures. More impressively, MASE is compatible with existing efficient parameter fine-tuning methods, such as prompt learning. Xianwei Zhuang, Xuxin Cheng, Zhihong Zhu 0001, Zhanpeng Chen, Hongxiang Li 0004, Yuexian Zou |
ACM Multimedia | 1 |
| 2024 | MaCSC: Towards Multimodal-augmented Pre-trained Language Models via Conceptual Prototypes and Self-balancing CalibrationabstractXianwei Zhuang, Zhichang Wang, Xuxin Cheng, Yuxin Xie, Liming Liang, Yuexian Zou. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Xianwei Zhuang, Zhichang Wang, Xuxin Cheng, Yuxin Xie 0004, Yuexian Zou |
NAACL-HLT | 1 |
| 2022 | Residual Swin Transformer Unet with Consistency Regularization for Automatic Breast Ultrasound Tumor SegmentationabstractAutomatic Breast Ultrasound (ABUS) image segmentation is of great significance for breast cancer diagnosis and treatment. However, similar to most medical datasets, ABUS image datasets are often small-scale and seriously imbalanced, which makes ABUS image segmentation become a challenge. To solve this problem, we propose the Residual Swin Transformer Unet with Consistency Regularization (RSTUnet-CR) which can make full use of non-lesion and unlabeled images for high-precision tumor segmentation on ABUS images. We design a consistency-regularization decoder to reconstruct the input image, which can learn well from non-lesion and unlabeled data. The reconstruction task makes the model more suitable for the imbalanced medical image datasets. In addition, observing that the ABUS images have global semantic correlation, we establish long-distance dependence of images by the residual Swin Transformer block to improve segmentation performance. We evaluate our method on the ABUS dataset collected from 256 subjects and demonstrate the superiority of the proposed method over other state-of-the-art methods in this imbalanced dataset. Xianwei Zhuang, Xiner Zhu, Haoji Hu, Jincao Yao, Na Feng, Dong Xu 0006 |
ICIP | 1 |