Xin Zou 0001

dblp:18/6081-1 · DBLP profile ↗
← Back
25ranked-venue papers
7as first author
25since 2021 · last 2026
0000-0003-4771-3836ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 4 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 PLA-MGRA: Multi-Granularity and Relation-Aware Learning for Efficient and Generalizable Protein-Ligand Binding Affinity Prediction
abstract
Protein-Ligand Affinity (PLA) prediction quantifies the interaction strength to guide rational drug design. Existing approaches typically analyze interaction at a single granularity and overlook tightly coupled relationships between protein and ligand in both structure and functionality, consequently yielding suboptimal representations, leading to significant performance drops in real-world scenarios. To address this problem, we propose PLA-MGRA, a minimalist and effective PLA prediction framework. Specifically, PLA-MGRA captures both fine-grained atomic details and coarse grained functional semantics within the 3D structure of protein–ligand complexes, through multi-granularity learning. To further parse the coupled protein–ligand relationships, we design relation-aware learning to enhance the binding nature of representations. Extensive experiments demonstrate that our method achieves state-of-the-art performance on multiple protein–ligand affinity prediction benchmarks, while also offering generalizability and interpretability.
Shunfan Li, Jiangkai Long, Xin Zou 0001, Chang Tang, Yuanyuan Liu 0004, Xiao He 0010
AAAI3
2026 Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
abstract
Current Video Large Language Models (VideoLLMs) suffer from quadratic computational complexity and key-value cache scaling, due to their reliance on processing excessive redundant visual tokens. To address this problem, we propose SharpV, a minimalist and efficient method for adaptive pruning of visual tokens and KV cache. Different from most uniform compression approaches, SharpV dynamically adjusts pruning ratios based on spatial-temporal information. Remarkably, this adaptive mechanism occasionally achieves performance gains over dense models, offering a novel paradigm for adaptive pruning. During the KV cache pruning stage, based on observations of visual information degradation, SharpV prunes degraded visual features via a self-calibration manner, guided by similarity to original visual features. In this way, SharpV achieves hierarchical cache pruning from the perspective of information bottleneck, offering a new insight into VideoLLMs' information flow. Experiments on multiple public benchmarks demonstrate the superiority of SharpV. Moreover, to the best of our knowledge, SharpV is notably the first two-stage pruning framework that operates without requiring access to exposed attention scores, ensuring full compatibility with hardware acceleration techniques like Flash Attention.
Jialong Qin, Xin Zou 0001, Xuming Hu
AAAI2
2026 Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods
abstract
Chenfei Liao, Wensong Wang, Zichen Wen, Xu Zheng, Yiyu Wang, Haocong He, Yuanhuiyi Lyu, Lutao Jiang, Xin Zou, Yuqian Fu, Bin Ren, Linfeng Zhang, Xuming Hu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Chenfei Liao, Wensong Wang, Zichen Wen, Xu Zheng 0002, Haocong He, Yuanhuiyi Lyu, Lutao Jiang, Xin Zou 0001, Yuqian Fu, Bin Ren 0005, Linfeng Zhang 0001, Xuming Hu
ACL (1)9
2026 Vision-Language Introspection: Mitigating Overconfident Hallucinations in MLLMs via Interpretable Bi-Causal Steering
abstract
Shuliang Liu, Songbo Yang, Dong Fang, Sihang Jia, Yuqi Tang, Lingfeng Su, Ruoshui Peng, Yibo Yan, Xin Zou, Xuming Hu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Songbo Yang, Sihang Jia, Lingfeng Su, Ruoshui Peng, Xin Zou 0001, Xuming Hu
ACL (1)9
2026 Curriculum trustworthy multi-modal learning
abstract
Trustworthy multi-modal learning reliably integrates multiple data sources. However, current methods often face a challenge, which is the inherent non-convex nature of deep neural networks. It leads to their susceptibility to local minima, ultimately resulting in a reduced capacity for generalization. To address this issue, we first create a theoretical framework, which extends the application of curriculum learning in multi-modal scenarios. Secondly, we propose a novel curriculum termed the Dynamic SRM Curriculum (DSRMC). It consists of two modules: a scoring function and a training schedule. The scoring function sorts samples from simple to complex. The training scheduler aims to manage the quantity of samples supplied at each round during training. DSRMC facilitates positioning the learned model in a flatter region of the loss landscape, thereby enhancing its overall generalization ability. Building on DSRMC, we eventually propose an innovative method termed as Curriculum Trustworthy Multi-modal Learning (CTML). It applies DSRMC in multi-modal learning application scenarios. Extensive experiments conducted on three open datasets show that the proposed CTML outperforms state-of-the-art methods, with a maximum improvement of 6.7% in macro F1 score. Our code and dataset are publicly available on https://github.com/HackerHyper/DSRMC.git .
Xin Zou 0001, Jun Sun 0014, Lingfang Zeng, Linqing Feng, Lei Liu 0029, Chang Tang
Expert Syst. Appl.2
2026 Learning Disentangled Representations for Generalized Multi-View Clustering
abstract
Multi-View Clustering (MVC) has gained significant attention for its ability to leverage complementary information across diverse views. However, existing deep MVC methods often struggle with view-distribution entanglement during cross-view fusion, which hampers the quality of the shared latent space and leads to suboptimal clustering performance. To address this issue, we propose the Generalized Multi-view Auto-Encoder (GMAE), a framework designed to preserve cross-view complementarity through disentangled representation learning. Specifically, GMAE employs dual-path autoencoders to decouple source features into view-specific and view-common embeddings, facilitating the discovery of clearer clustering structures. We further construct cross-view adversarial discriminators to guide view-specific encoders in capturing more discriminative features. By strategically modulating mutual information, GMAE effectively aligns distributions and prevents representation collapse, ensuring the generation of robust, non-trivial embeddings. Comprehensive experiments on 13 benchmark datasets demonstrate that GMAE consistently outperforms state-of-the-art methods in both complete and incomplete MVC tasks.
Xin Zou 0001, Ruimeng Liu, Chang Tang, Zhenglai Li, Xinwang Liu 0002, Kunlun He, Wanqing Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 Generative Diffusion Contrastive Network for Multi-View Clustering
abstract
In recent years, Multi-View Clustering (MVC) has been significantly advanced under the influence of deep learning. By integrating heterogeneous data from multiple views, MVC enhances clustering analysis, making multi-view fusion critical to clustering performance. However, multi-view fusion remains challenged by low-quality data, primarily stemming from two reasons: 1) Certain views are contaminated by noisy data. 2) Some views suffer from missing data. This paper proposes a novel Stochastic Generative Diffusion Fusion (SGDF) method to address this problem. SGDF leverages a multiple generative mechanism for the multi-view feature of each sample. It exhibits robustness against low-quality data. Building on SGDF, we further present the Generative Diffusion Contrastive Network (GDCN). Extensive experiments show that GDCN achieves the state-of-the-art results in deep MVC tasks. The source code is publicly available athttps://github.com/HackerHyper/GDCN.
Xin Zou 0001, Lei Liu 0029, Chang Tang, Li-Rong Dai 0001
IEEE Signal Process. Lett.2
2025 Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading Scenarios
abstract
Yunkai Dang, Mengxi Gao, Yibo Yan, Xin Zou, Yanggan Gu, Jungang Li, Jingyu Wang, Peijie Jiang, Aiwei Liu, Jia Liu, Xuming Hu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yunkai Dang, Mengxi Gao, Xin Zou 0001, Yanggan Gu, Jungang Li, Peijie Jiang, Aiwei Liu, Xuming Hu
EMNLP4
2025 Pierce the Mists, Greet the Sky: Decipher Knowledge Overshadowing via Knowledge Circuit Analysis
abstract
Large Language Models (LLMs), despite their remarkable capabilities, are hampered by hallucinations.A particularly challenging variant, knowledge overshadowing, occurs when one piece of activated knowledge inadvertently masks another relevant piece, leading to erroneous outputs even with high-quality training data.Current understanding of overshadowing is largely confined to inference-time observations, lacking deep insights into its origins and internal mechanisms during model training.Therefore, we introduce PHANTOMCIRCUIT, a novel framework designed to comprehensively analyze and detect knowledge overshadowing.By innovatively employing knowledge circuit analysis, PHANTOMCIRCUIT dissects the function of key components in the circuit and how the attention pattern dynamics contribute to the overshadowing phenomenon and its evolution throughout the training process.Extensive experiments demonstrate PHANTOMCIRCUIT 's effectiveness in identifying such instances, offering novel insights into this elusive hallucination and providing the research community with a new methodological lens for its potential mitigation.Our code can be found in https://github.com/halfmorepiece/PhantomCircuit.
Haoming Huang, Jiahao Huo, Xin Zou 0001, Xinfeng Li, Kun Wang 0056, Xuming Hu
EMNLP4
2025 Trusted Mamba Contrastive Network for Multi-View Clustering
abstract
Multi-view clustering can partition data samples into their categories by learning a consensus representation in an unsupervised way and has received more and more attention in recent years. However, there is an untrusted fusion problem. The reasons for this problem are as follows: 1) The current methods ignore the presence of noise or redundant information in the view; 2) The similarity of contrastive learning comes from the same sample rather than the same cluster in deep multi-view clustering. It causes multi-view fusion in the wrong direction. This paper proposes a novel multi-view clustering network to address this problem, termed as Trusted Mamba Contrastive Network (TMCN). Specifically, we present a new Trusted Mamba Fusion Network (TMFN), which achieves a trusted fusion of multi-view data through a selective mechanism. Moreover, we align the fused representation and the view-specific representation using the Average-similarity Contrastive Learning (AsCL) module. AsCL increases the similarity of view presentation from the same cluster, not merely from the same sample. Extensive experiments show that the proposed method achieves state-of-the-art results in deep multi-view clustering tasks. The source code is available at https://github.com/HackerHyper/TMCN.
Xin Zou 0001, Lei Liu 0029, Zhangmin Huang, Chang Tang, Li-Rong Dai 0001
ICASSP2
2025 Dynamic SRM Curriculum for Trustworthy Multi-modal Classification
abstract
Trustworthy multi-modal learning integrates multiple sources of data reliably. However, the current methods still focus on performance improvement by developing deep multi-modal networks. These approaches frequently encounter challenges due to the inherent non-convex nature of deep neural networks and their vulnerability to local minima, ultimately leading to a diminished ability for generalization. To address this problem, we present a novel curriculum termed the Dynamic SRM Curriculum (DSRMC). Within DSRMC, the deep trustworthy multi-modal networks undergo training with data provided sequentially, progressing from simple to complex samples. This training strategy mimics the human learning process, commencing with fundamental concepts and gradually advancing to tackle more complex and abstract ideas. Building upon DSRMC, we propose an innovative Curriculum Trustworthy Multi-modal Learning (CTML) method. CTML makes it easier to place the learned model in a flatter area, which improves its overall ability for generalization. Comprehensive experiments on three public datasets demonstrate that the proposed CTML performs better than state-of-the-art methods, achieving a maximum improvement of 6.7% on macroF1.
Cui Yu, Xin Zou 0001, Zhangmin Huang, Chenshu Hu, Jun Sun 0014, Bo Lyu, Lei Liu 0029, Chang Tang, Li-Rong Dai 0001
ICASSP3
2025 Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality
abstract
Multimodal Large Language Models (MLLMs) have emerged as a central focus in both industry and academia, but often suffer from biases introduced by visual and language priors, which can lead to multimodal hallucination. These biases arise from the visual encoder and the Large Language Model (LLM) backbone, affecting the attention mechanism responsible for aligning multimodal inputs. Existing decoding-based mitigation methods focus on statistical correlations and overlook the causal relationships between attention mechanisms and model output, limiting their effectiveness in addressing these biases. To tackle this issue, we propose a causal inference framework termed CausalMM that applies structural causal modeling to MLLMs, treating modality priors as a confounder between attention mechanisms and output. Specifically, by employing backdoor adjustment and counterfactual reasoning at both the visual and language attention levels, our method mitigates the negative effects of modality priors and enhances the alignment of MLLM's inputs and outputs, with a maximum score improvement of 65.3% on 6 VLind-Bench indicators and 164 points on MME Benchmark compared to conventional methods. Extensive experiments validate the effectiveness of our approach while being a plug-and-play solution. Our code is available at: https://github.com/The-Martyr/CausalMM.
Guanyu Zhou, Xin Zou 0001, Kun Wang 0056, Aiwei Liu, Xuming Hu
ICLR3
2025 Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models
abstract
Despite their impressive capabilities, Multimodal Large Language Models (MLLMs) are prone to hallucinations, i.e., the generated content that is nonsensical or unfaithful to input sources. Unlike in LLMs, hallucinations in MLLMs often stem from the sensitivity of text decoder to visual tokens, leading to a phenomenon akin to "amnesia" about visual information. To address this issue, we propose MemVR, a novel decoding paradigm inspired by common cognition: when the memory of an image seen the moment before is forgotten, people will look at it again for factual answers. Following this principle, we treat visual tokens as supplementary evidence, re-injecting them into the MLLM through Feed Forward Network (FFN) as “key-value memory” at the middle trigger layer. This look-twice mechanism occurs when the model exhibits high uncertainty during inference, effectively enhancing factual alignment. Comprehensive experimental evaluations demonstrate that MemVR significantly mitigates hallucination across various MLLMs and excels in general benchmarks without incurring additional time overhead.
Xin Zou 0001, Yuanhuiyi Lyu, Kening Zheng, Sirui Huang, Junkai Chen, Peijie Jiang, Chang Tang, Xuming Hu
ICML1
2025 RealRAG: Retrieval-augmented Realistic Image Generation via Self-reflective Contrastive Learning
abstract
Recent text-to-image generative models, e.g., Stable Diffusion V3 and Flux, have achieved notable progress. However, these models are strongly restricted to their limited knowledge, a.k.a., their own fixed parameters, that are trained with closed datasets. This leads to significant hallucinations or distortions when facing fine-grained and unseen novel real-world objects, e.g., the appearance of the Tesla Cybertruck. To this end, we present the first real-object-based retrieval-augmented generation framework (RealRAG), which augments fine-grained and unseen novel object generation by learning and retrieving real-world images to overcome the knowledge gaps of generative models. Specifically, to integrate missing memory for unseen novel object generation, we train a reflective retriever by self-reflective contrastive learning, which injects the generator’s knowledge into the sef-reflective negatives, ensuring that the retrieved augmented images compensate for the model’s missing knowledge. Furthermore, the real-object-based framework integrates fine-grained visual knowledge for the generative models, tackling the distortion problem and improving the realism for fine-grained object generation. Our Real-RAG is superior in its modular application to all types of state-of-the-art text-to-image generative models and also delivers remarkable performance boosts with all of them, such as a gain of 16.18% FID score with the auto-regressive model on the Stanford Car benchmark.
Yuanhuiyi Lyu, Xu Zheng 0002, Lutao Jiang, Xin Zou 0001, Huiyu Zhou 0005, Linfeng Zhang 0001, Xuming Hu
ICML5
2025 From Guesswork to Guarantee: Towards Faithful Multimedia Web Forecasting with TimeSieve
abstract
The domain of time series forecasting has gained significant attention due to its critical applications in multimedia-rich web traffic (including video streaming workloads and dynamic content delivery) and cross-platform advertisement click predictions, which are essential for web operations planning. While models like TimeSieve have demonstrated strong capabilities in predicting web visitation metrics, they suffer from critical unfaithfulness issues, including sensitivity to random seeds, input noise, layer noise, and parametric perturbations. To address these limitations, we propose Faithful TimeSieve (FTS), an enhanced framework designed to improve prediction reliability and robustness. Our approach systematically detects and mitigates unfaithfulness in TimeSieve, significantly enhancing its stability and consistency. Experimental results demonstrate that FTS substantially improves the model's faithfulness, setting a new standard for temporal forecasting methods. This advancement not only increases TimeSieve's reliability but also contributes to more robust temporal modeling, particularly crucial for web traffic forecasting where prediction accuracy directly impacts operational decisions. Our work thus represents a significant step toward more dependable time series predictions in web-related applications.
Songning Lai, Ninghui Feng, Jiechao Gao, Hao Wang 0220, Haochen Sui, Xin Zou 0001, Wenshuo Chen, Lijie Hu, Hang Zhao 0010, Xuming Hu, Yutao Yue
ACM Multimedia6
2025 SparseMVC: Probing Cross-view Sparsity Variations for Multi-view Clustering
abstract
Existing multi-view clustering methods employ various strategies to address data-level sparsity and view-level dynamic fusion. However, we identify a critical yet overlooked issue: varying sparsity across views. Cross-view sparsity variations lead to encoding discrepancies, heightening sample-level semantic heterogeneity and making view-level dynamic weighting inappropriate. To tackle these challenges, we propose Adaptive Sparse Autoencoders for Multi-View Clustering (SparseMVC), a framework with three key modules. Initially, the sparse autoencoder probes the sparsity of each view and adaptively adjusts encoding formats via an entropy-matching loss term, mitigating cross-view inconsistencies. Subsequently, the correlation-informed sample reweighting module employs attention mechanisms to assign weights by capturing correlations between early-fused global and view-specific features, reducing encoding discrepancies and balancing contributions. Furthermore, the cross-view distribution alignment module aligns feature distributions during the late fusion stage, accommodating datasets with an arbitrary number of views. Extensive experiments demonstrate that SparseMVC achieves state-of-the-art clustering performance. Our framework advances the field by extending sparsity handling from the data-level to view-level and mitigating the adverse effects of encoding discrepancies through sample-level dynamic weighting. The source code is publicly available at https://github.com/cleste-pome/SparseMVC.
Ruimeng Liu, Xin Zou 0001, Chang Tang, Xingchen Hu 0001, Kun Sun 0002, Xinwang Liu 0002
NeurIPS2
2025 Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
abstract
Despite their powerful capabilities, multimodal large language models (MLLMs) suffer from considerable computational overhead due to their reliance on massive visual tokens. Recent studies have explored token pruning to alleviate this problem, which typically uses text-vision cross-attention or [CLS] attention to assess and discard redundant visual tokens. In this work, we identify a critical limitation of such attention-first pruning approaches, i.e., they tend to preserve semantically similar tokens, resulting in pronounced performance drops under high pruning rates. To this end, we propose HoloV, a simple yet effective, plug-and-play visual token pruning framework for efficient inference. Distinct from previous attention-first schemes, HoloV rethinks token retention from a holistic perspective. By adaptively distributing the pruning budget across different spatial crops, HoloV ensures that the retained tokens capture the global visual context rather than isolated salient features. This strategy minimizes representational collapse and maintains task-relevant information even under aggressive pruning. Experimental results demonstrate that our HoloV achieves superior performance across various tasks, MLLM architectures, and pruning ratios compared to SOTA methods. For instance, LLaVA1.5 equipped with HoloV preserves 95.8% of the original performance after pruning 88.9% of visual tokens, achieving superior efficiency-accuracy trade-offs.
Xin Zou 0001, Yuanhuiyi Lyu, Xu Zheng 0002, Linfeng Zhang 0001, Xuming Hu
NeurIPS1
2025 Cancer-drug response prediction via feature aggregation and association graph learning
Kaiyi Xu, Minhui Wang, Xin Zou 0001, Chengfu Ji, Chang Tang
Eng. Appl. Artif. Intell.3
2025 Dual alignment feature embedding network for multi-omics data clustering
Yuang Xiao, Xin Zou 0001, Chang Tang
Knowl. Based Syst.4
2025 HSTrans: Homogeneous substructures transformer for predicting frequencies of drug-side effects
Kaiyi Xu, Minhui Wang, Xin Zou 0001, Ao Wei, Jiajia Chen 0010, Chang Tang
Neural Networks3
2024 DAI-Net: Dual Adaptive Interaction Network for Coordinated Medication Recommendation
abstract
Medication recommendation is a productive task for AI-driven healthcare systems, which can assist clinicians in prescribing judicious and effective treatments. However, existing medication recommendation methods omit two key pieces of information: Coarse-grained interaction information between distinct types of symptoms in a patient's medical history and corresponding medication representations can serve as attention for predicting the current medication combinations of the patient. Fine-grained interaction information between medication substructure representations and different types of symptoms can facilitate the construction of molecular-level disentangled medication representations. To address this dilemma, we propose a novelDualAdaptiveInteractionNetwork (DAI-Net), which encodes comprehensive interaction knowledge between patients' multifaceted health records and medication molecules to improve the performance of medication recommendation and heighten interpretability of the model. Specifically, we design a symptom-aware medication matching module to extract coordinated associations between patient symptoms and medication molecules, coarse-grained interaction learning. The medication embeddings are utilized to transform patient-medication matching properties into a symptom-substructure matching matrix for fine-grained interaction. The patient's Longitudinal representation is employed as a query to decode both symptom-medication and symptom-substructure matching information for coordinated medication representation. DAI-Net is an end-to-end recommendation model. Extensive experiments on the real-world EHR datasets, i.e., the public benchmark MIMIC-III, MIMIC-IV, and eICU, demonstrate that the proposed DAI-Net achieves competitive performance compared to other state-of-the-art ones, with an average improvement of 1.8%, 2.1% in Jaccard on MIMIC-III and -IV dataset.
Xin Zou 0001, Xiao He 0010, Wei Zhang 0049, Jiajia Chen 0010, Chang Tang
IEEE J. Biomed. Health Informatics1
2023 Hierarchical Attention Learning for Multimodal Classification
abstract
Multimodal learning aims to integrate complementary information from different modalities for more reliable decisions. However, existing multimodal classification methods simply integrate the learned local features, which ignore the underlying structure of each modality and the higher-order correlation across modalities. In this paper, we propose a novel Hierarchical Attention Learning Network (HALNet) for multimodal classification. Specifically, HALNet has three merits: 1) A hierarchical feature fusion module is proposed to learn multilevel features, aggregating multi-level features for a global feature representation with the attention mechanism and progressive fusion tactics. 2) A cross-modal higher-order fusion module is introduced to capture the prospective cross-modal correlations at label space. 3) A dual prediction pattern is designed to generate credible decisions. Extensive experiments on three real-world multimodal datasets demonstrate that HALNet achieves competitive performance compared to the state-of-the-art.
Xin Zou 0001, Chang Tang, Wei Zhang 0049, Kun Sun 0002, Liangxiao Jiang
ICME1
2023 Multispectral Object Detection via Cross-Modal Conflict-Aware Learning
abstract
Multispectral object detection has gained significant attention due to its potential in all-weather applications, particularly those involving visible (RGB) and infrared (IR) images. Despite substantial advancements in this domain, current methodologies primarily rely on rudimentary accumulation operations to combine complementary information from disparate modalities, overlooking the semantic conflicts that arise from the intrinsic heterogeneity among modalities. To address this issue, we propose a novel learning network, the Cross-modal Conflict-Aware Learning Network (CALNet), that takes into account semantic conflicts and complementary information within multi-modal input. Our network comprises two pivotal modules: the Cross-Modal Conflict Rectification Module (CCR) and the Selected Cross-modal Fusion (SCF) Module. The CCR module mitigates modal heterogeneity by examining contextual information of analogous pixels, thus alleviating multi-modal information with semantic conflicts. Subsequently, semantically coherent information is supplied to the SCF module, which fuses multi-modal features by assessing intra-modal importance to select semantically rich features and mining inter-modal complementary information. To assess the effectiveness of our proposed method, we develop a two-stream one-stage detector based on CALNet for multispectral object detection. Comprehensive experimental outcomes demonstrate that our approach considerably outperforms existing methods in resolving the cross-modal semantic conflict issue and achieving state-of-the-art accuracy in detection results.
Xiao He 0010, Chang Tang, Xin Zou 0001, Wei Zhang 0049
ACM Multimedia3
2023 DPNET: Dynamic Poly-attention Network for Trustworthy Multi-modal Classification
abstract
With advances in sensing technology, multi-modal data collected from different sources are increasingly available. Multi-modal classification aims to integrate complementary information from multi-modal data to improve model classification performance. However, existing multi-modal classification methods are basically weak in integrating global structural information and providing trustworthy multi-modal fusion, especially in safety-sensitive practical applications (e.g., medical diagnosis). In this paper, we propose a novel Dynamic Poly-attention Network (DPNET) for trustworthy multi-modal classification. Specifically, DPNET has four merits: (i) To capture the intrinsic modality-specific structural information, we design a structure-aware feature aggregation module to learn the corresponding structure-preserved global compact feature representation. (ii) A transparent fusion strategy based on the modality confidence estimation strategy is induced to track information variation within different modalities for dynamical fusion. (iii) To facilitate more effective and efficient multi-modal fusion, we introduce a cross-modal low-rank fusion module to reduce the complexity of tensor-based fusion and activate the implication of different rank-wise features via a rank attention mechanism. (iv) A label confidence estimation module is devised to drive the network to generate more credible confidence. An intra-class attention loss is introduced to supervise the network training. Extensive experiments on four real-world multi-modal biomedical datasets demonstrate that the proposed method achieves competitive performance compared to other state-of-the-art ones.
Xin Zou 0001, Chang Tang, Zhenglai Li, Xiao He 0010, Shan An, Xinwang Liu 0002
ACM Multimedia1
2023 Inclusivity induced adaptive graph learning for multi-view clustering
Xin Zou 0001, Chang Tang, Kun Sun 0002, Wei Zhang 0049, Deqiong Ding
Knowl. Based Syst.1