VLDB 2026 Research / reviewers in the wild / expert
Chengliang Liu 0003
dblp:30/5444-3
· DBLP profile ↗
52ranked-venue papers
10as first author
52since 2021 · last 2026
0000-0001-5983-8981ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 8 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 4 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IAD-R1: Reinforcing Consistent Reasoning in Industrial Anomaly DetectionabstractIndustrial anomaly detection is a critical component of modern manufacturing, yet the scarcity of defective samples restricts traditional detection methods to scenario-specific applications. Although Vision-Language Models (VLMs) demonstrate significant advantages in generalization capabilities, their performance in industrial anomaly detection remains limited. To address this challenge, we propose IAD-R1, a universal post-training framework applicable to VLMs of different architectures and parameter scales, which substantially enhances their anomaly detection capabilities. IAD-R1 employs a two-stage training strategy: the Perception Activation Supervised Fine-Tuning (PA-SFT) stage utilizes a meticulously constructed high-quality Chain-of-Thought dataset (Expert-AD) for training, enhancing anomaly perception capabilities and establishing reasoning-to-answer correlations; the Structured Control Group Relative Policy Optimization (SC-GRPO) stage employs carefully designed reward functions to achieve a capability leap from "Anomaly Perception" to "Anomaly Interpretation". Experimental results demonstrate that IAD-R1 achieves significant improvements across 7 VLMs, the largest improvement was on the DAGM dataset, with average accuracy 43.3% higher than the 0.5B baseline. Notably, the 0.5B parameter model trained with IAD-R1 surpasses commercial models including GPT-4.1 and Claude-Sonnet-4 in zero-shot settings, demonstrating the effectiveness and superiority of IAD-R1. Yunkang Cao, Chengliang Liu 0003, Yuan Xiong, Xinghui Dong, Chao Huang 0008 |
AAAI | 3 |
| 2026 | Detecting Fake News in Short Videos Through Multi-View AggregationabstractThe increasing prominence of short video platforms has positioned them as a primary channel for public awareness of current events, while also facilitating the widespread dissemination of fake news, thus highlighting the critical need for automated detection technologies. In contrast to fake news confined to text and images, short video news encompasses multiple modalities and extensive information, presenting heightened challenges. Most existing research emphasizes the analysis of news content or user comments alone, while overlooking the crucial role of publishers, leading to poor model performance when handling fake news lacking obvious false signals. Therefore, we propose a Publisher Profiling Module to identify new false signals. To enable a more comprehensive detection of misinformation, we design a Multi-View Aggregation (MVA) model, simultaneously evaluating news from three distinct perspectives: sentiment analysis, content understanding, and publisher profiling. Late fusion is applied at the decision level to leverage the complementary strengths of these perspectives, addressing the limitations of single-view methods. Our experiments conducted on the FakeSV and FVC datasets demonstrate the superior performance of the proposed method. Yuan Xiong, Chengliang Liu 0003, Jie Wen 0001, Chao Huang 0008 |
AAAI | 3 |
| 2026 | Frequency-Aligned Cross-Modal Learning with Top-K Wavelet Fusion and Dynamic Expert Routing for Enhanced Retinal Disease DiagnosisabstractMultimodal fusion of color fundus photography (CFP) and optical coherence tomography (OCT) B-scan images has demonstrated superior diagnostic potential for retinal diseases compared to single-modality approaches. However, existing fusion paradigms - whether through naive concatenation or attention mechanisms - treat cross-modal interactions indiscriminately, lacking adaptive modulation of modality-specific contributions under varying clinical scenarios. We propose an adaptive fusion framework that dynamically routes and refines multimodal signals for enhancing disease recognition. The framework comprises two key components: 1) Dynamic Cross-Modal Expert Routing (CMER), which selectively activates convolutional neural network (CNN) experts from one modality based on contextual guidance from the other, ensuring only the most relevant feature extractors contribute to fusion; and 2) Top-K Expert-Guided Wavelet Fusion (TEWF), which performs discrete wavelet transform (DWT) to decompose selected features into low- and high-frequency subbands. Cross-modal attention is then applied specifically to high-frequency components, where lesion-specific microstructures reside, enabling frequency-aware fusion. Finally, inverse DWT (IDWT) reconstructs the fused representation, weighted by CMER-derived importance scores to amplify informative modality cues while suppressing redundancy. Experimental validation on two multimodal retinal datasets demonstrates that our method achieves state-of-the-art performance, outperforming existing fusion strategies by significant margins in disease classification accuracy and robustness. Haoran Li 0024, Haoyu Cao 0002, Yongting Hu, Qihao Xu, Chengliang Liu 0003, Xiaoling Luo 0001, Zhihao Wu 0002, Yong Xu 0001, Wei Wang 0169 |
AAAI | 6 |
| 2026 | Vision-Language Models Guided Graph Concept Reasoning for Interpretable Diabetic Retinopathy DiagnosisabstractDeep neural networks (DNNs) have significantly advanced diabetic retinopathy (DR) diagnosis, yet their black-box nature limits clinical acceptance due to a lack of interpretability. Concept bottleneck model (CBM) offers a promising solution by enabling concept-level reasoning and test-time intervention, with recent DR studies modeling lesions as concepts and grades as outcomes. However, current methods often ignore relationships between lesion concepts across different DR grades and struggle when fine-grained lesion concepts are unavailable, limiting their interpretability and real-world applicability. To bridge these gaps, we propose VLM-GCR, a vision-language model guided graph concept reasoning framework for interpretable DR diagnosis. VLM-GCR emulates the diagnostic process of ophthalmologists by constructing a grading-aware lesion concept graph that explicitly models the interactions among lesions and their relationships to disease grades. In concept-free clinical scenarios, our method introduces a vision-language guided dynamic concept pseudo-labeling mechanism to mitigate the challenges of existing concept-based models in fine-grained lesion recognition. Additionally, we introduce a multi-level intervention method that supports error correction, enabling transparent and robust human-AI collaboration. Experiments on two public DR benchmarks show that VLM-GCR achieves strong performance in both lesion and grading tasks, while delivering clear and clinically meaningful reasoning steps. Qihao Xu, Xiaoling Luo 0001, Chengliang Liu 0003, Yongting Hu, Xinheng Lyu, Yong Xu 0001 |
AAAI | 4 |
| 2026 | Generative data-engine foundation model for universal few-shot 2D vascular image segmentationabstractThe segmentation of 2D vascular structures via deep learning holds significant clinical value but is hindered by the scarcity of annotated data, severely limiting its widespread application. Developing a universal few-shot vascular segmentation model is highly desirable, yet remains challenging due to the need for extensive training and the inherent complexities of vascular imaging. In this work, we propose UniVG (Generative Data-engine Foundation Model for Universal Few-shot 2D Vascular Image Segmentation), a novel approach that learns the compositionality of vascular images and constructing a generative foundation model for robust vascular segmentation. UniVG enables the synthesis and learning of diverse and realistic vascular images through two key innovations: 1) Compositional learning for flexible and diverse vascular synthesis: It decomposes and recombines vascular structures with varying morphological features and diverse foreground-background configurations to generate richly diverse synthetic image-label pairs. 2) Few-shot generative adaptation for transferable segmentation: It fine-tunes pre-trained models with minimal annotated data to bridge the gap between synthetic and real vascular domains, synthesizing authentic and diverse vessel images for downstream few-shot vascular segmentation learning. To support our approach, we develop UniVG-58K, a large dataset comprising 58,689 vascular images across five imaging modalities, facilitating robust large-scale generative pre-training. Extensive experiments on 11 vessel segmentation tasks cross 5 modalties (only with 5 labeled images on each task) demonstrate that UniVG achieves performance comparable to fully supervised models, significantly reducing data collection and annotation costs. All code and datasets will be made publicly available at https://github.com/XinAloha/UniVG. Rongjun Ge, Yuxing Liu, Chengliang Liu 0003, Pinzheng Zhang, Jiong Zhang 0004, Jian Yang 0009, Jean-Louis Dillenseger, Yuting He 0001, Yang Chen 0008 |
Medical Image Anal. | 4 |
| 2026 | Learning Compact Semantic Information and Reliable Pseudo-Labels for Incomplete Multi-View Multi-Label ClassificationabstractMulti-view data encompasses various data types, including multi-feature, multi-sequence, and multi-modal data. Multi-view multi-label classification aims to leverage the rich semantic information contained in multiple views to achieve enhanced multi-label classification performance. In practical applications, the absence of views and labels poses a significant challenge to multi-view multi-label classification tasks. Premised on the assumption that shared semantic information across multiple views is sufficient to support the downstream task, we propose CTRL, a novel incomplete multi-view multi-label classification framework to address the multi-view learning challenge on the data with partially missing views and missing labels in this paper. The core mechanism of CTRL lies in learning a high-purity, low-redundancy condensed representation that adequately captures the essential information of the original data. Specifically, we design a new objective loss to enhance the semantic information of shared cross-view within the joint representation learning process while simultaneously suppressing intra-view redundant information that is irrelevant to the downstream task. This enables CTRL to extract task-relevant representations even when views are incomplete. Furthermore, we employ the Beta Evidential Neural Network to model the label distribution. This network is then integrated with Dempster-Shafer theory, enabling our model to perform label-level classification uncertainty estimation. This also allows us to use the estimated uncertainty and belief mass to create high-reliability pseudo-labels, resulting in further gains in model performance. Experimental results on multiple benchmark datasets demonstrate the superior performance of our proposed model in terms of accuracy, robustness, and reliability. Chengliang Liu 0003, Jie Wen 0001, Li Shen 0008, Bob Zhang 0001, Yong Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Causal Interventional Prompt Tuning for Few-Shot Out-of-Distribution GeneralizationabstractFine-tuning pre-trained vision-language models (VLMs) has shown substantial benefits in a wide range of downstream tasks, often achieving impressive performance with minimal labeled data. Parameter-efficient fine-tuning techniques, in particular, have demonstrated their effectiveness in enhancing downstream task performance. However, these methods frequently struggle to generalize to out-of-distribution (OOD) data due to their reliance on non-causal representations, which can introduce biases and spurious correlations that negatively impact decision-making. Such spurious factors hinder the model's generalization ability beyond the training distribution. To address these challenges, in this paper, we propose a novel causal intervention-based prompt tuning method to adapt VLMs to few-shot OOD generalization. Specifically, we leverage the front-door adjustment technique from causal inference to mitigate the effects of spurious correlations and enhance the model's focus on causal relationships. Built upon VLMs, our approach begins by decoupling causal and non-causal representations in the vision-language alignment process. The causal representation that captures only essential semantically relevant information can serve as a mediator variable between the input image and output label, mitigating the biases from the latent confounder. To further enrich this causal representation, we propose a novel text-based diversity augmentation technique that uses textual features to provide additional semantic context. This augmentation technique can enhance the diversity of the causal representation, making it more robust and generalizable to various OOD scenarios. Experimental results across multiple OOD datasets demonstrate that our method significantly outperforms existing approaches, achieving state-of-the-art generalization performance. Jie Wen 0001, Chao Huang 0008, Chengliang Liu 0003, Yong Xu 0001, Xiaochun Cao |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Partial Multiview Incomplete Multilabel Learning via Uncertainty-Driven Reliable Dynamic FusionabstractCurrently, an increasing number of researchers are focusing on partial multiview incomplete multilabel learning. However, many methods generally integrate features from multiple views via an average weighting strategy, which overlooks the potential mismatch between the contribution of each view and their assigned fusion weights and thus generates unreliable fused features. To address this issue, we propose a novel uncertainty-driven reliable dynamic fusion framework for partial multiview incomplete multilabel learning. Unlike existing methods, the proposed uncertainty-driven reliable sample-level dynamic fusion module operates on the principle that samples exhibiting greater uncertainty possess fewer reliable features. This module evaluates the uncertainty of each sample and, in turn, estimates the reliability of features with the uncertainty of sample judgement, thereby obtaining reliable weights to guide the information fusion of multiple views. Furthermore, many existing approaches for handling incomplete multilabel scenarios typically concentrate on the information from annotated labels, neglecting the potential information of unknown tags. To bridge this gap, we incorporate an innovative pseudolabelling strategy that effectively identifies trustworthy pseudolabels that correspond to those unannotated uncertain labels, thereby adding additional supervisory information to assist model training. Moreover, we also devise a feature masking strategy to further augment the encoder's representation learning capabilities. The experimental results across five datasets demonstrate that our method outperforms current state-of-the-art methods. Jie Wen 0001, Xiaohuan Lu, Chengliang Liu 0003, Xiaozhao Fang, Yong Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Disentangling Consistent and Specific Information for Double Incomplete Multi-View Multi-Label ClassificationabstractAs a prominent research topic, multi-view multi-label classification (MvMlC) aims to assign multiple labels to samples by integrating information from various perspectives. However, in real-world scenarios, MvMlC frequently faces the learning challenge of data with missing views and labels, typically resulting from sensor malfunctions, or the costly and time-consuming process of manual annotation. In addition, learning robust representations that are both consistent across views and specific to individual views remains a challenge. To address these issues, we propose a novel double incomplete multi-view multi-label classification framework based on Disentangling Consistent and Specific Information (DCSI). Specifically, we employ a dual-channel encoder with identical architecture but distinct objectives to extract cross-view consistent information and view-specific unique information from all views, respectively. Meanwhile, a view discriminator is constructed to decouple these two types of information, facilitating the extraction of pure consistent and specific information. Moreover, we meticulously design fusion strategies tailored to each representation type. Regarding consistent representations, we propose a dynamic-confidence-aware fusion mechanism that assesses the reliability of each view's representations in relation to the classification task, enabling the model to prioritize information from trustworthy representations. For specific representations, in light of their complementary rather than redundant property, we suggest treating such representations from each view equally to ensure fairness. Through experimental validation on five datasets, the results demonstrate that our method outperforms existing state-of-the-art methods. Jie Wen 0001, Lian Zhao, Xiaohuan Lu, Chengliang Liu 0003, Li Shen 0008, Chao Huang 0008, Yong Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | AdaSAM-AD: Boosting SAM2 for fine-grained pixel-level anomaly detection via spatial-channel calibration and deformable cascades
Xianbing Zhao, Xinyang Yang, Wai Keung Wong, Chengliang Liu 0003 |
Pattern Recognit. | 6 |
| 2026 | Noise-Induced Cross-Modal Information Interaction and Dual-Prompt Learning for Medical Image SegmentationabstractAccurate medical image segmentation plays a vital role in clinical diagnostics by facilitating the precise delineation of anatomical structures and pathological regions. However, the performance of existing segmentation methods is often constrained by the scarcity of high-quality annotated datasets, as manual labeling is both labor-intensive and reliant on domain-specific expertise. To address this limitation without requiring additional annotations, we propose a novel multimodal segmentation framework that leverages medical text annotations as an auxiliary modality to complement visual information. In particular, our approach introduces a learnable encoding strategy for joint distribution modeling of image and text, which enables discriminative fusion and effectively suppresses cross-modal redundancy. Moreover, we innovatively design a frequency-domain prompt encoder based on the discrete wavelet transform (DWT) to capture multi-frequency features, thereby significantly enhancing the model's ability to delineate fine-grained boundaries. Overall, our framework integrates cross-attention for effective cross-modal interaction, employs joint distribution modeling to enable discriminative and redundancy-reduced multimodal fusion, and incorporates auxiliary supervision to strengthen the learning of task-relevant features. Extensive experiments on nine public datasets across three clinical tasks-including cell, lung infection, and polyp segmentation-demonstrate that our method achieves competitive segmentation performance while maintaining favorable computational efficiency. Comprehensive ablation studies and feature distribution visualizations further validate the effectiveness and robustness of our proposed components. The code will be made publicly available at https://github.com/chenpeng052/MDFP. Chao Huang 0008, Jie Wen 0001, Wei Wang 0335, Li Shen 0008, Wenqi Ren, Xiaochun Cao, Chengliang Liu 0003 |
IEEE Trans. Image Process. | 8 |
| 2026 | Multi-View Hilbert Curve-Based Hierarchical Information Aggregation for Incomplete Multimodal Alzheimer's Disease DiagnosisabstractTimely identification of Alzheimer's disease (AD) benefits from combining neuroimaging, fluid biomarkers, and cognitive assessments, yet in practice one or more modalities are often unavailable due to various factors such as cost, patient compliance, and procedural risks. Furthermore, conventional convolutional neural network (CNN) architectures and even Transformer-based models struggle to efficiently capture both local and global dependencies, especially when dealing with high-dimensional and highly heterogeneous medical data. In this study, we introduce a novel hierarchical information aggregation and dynamic fusion (HI-AD) framework for incomplete multimodal AD diagnosis. Our method couples a multi-view Hilbert curve-guided Mamba block with hierarchical spatial feature extraction to retain spatial continuity, model long-range dependencies, and integrate local context in neuroimaging data. To balance semantic alignment and modality-specific information, we propose a unified mutual information-driven learning objective with an active confidence evaluation mechanism, thereby preventing modality collapse and promoting robust representation learning. Extensive experiments on real-world datasets validate that our HI-AD framework consistently outperforms existing state-of-the-art methods across a diverse range of modality-missing scenarios, establishing an effective and generalizable solution for early-stage AD screening in heterogeneous clinical data environments. Chengliang Liu 0003, Yuanxi Que, Wai Keung Wong, Xiaoling Luo 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2025 | Multi-view Evidential Learning-based Medical Image SegmentationabstractMedical image segmentation provides useful information about the shape and size of organs, which is beneficial for improving diagnosis, analysis, and treatment. Despite traditional deep learning-based models can extract domain-specific knowledge, they face a generalization bottleneck due to the limited embedded knowledge scope. Vision foundation models have been demonstrated to be effective in extracting generalizable knowledge, but they cannot extract domain-specific knowledge without fine-tuning. In this work, we propose a novel multi-view evidential learning-based framework, which can extract both domain-specific and generalizable knowledge from multi-view features by combining the advantages of traditional and vision foundation models. Specifically, a novel multi-view state space model (MV-SSM) is designed to extract task-related knowledge while removing redundant information within multi-view features. The proposed MV-SSM utilizes Mamba, a state space model, to model cross-view contextual dependencies between domain-specific and generalizable features. Additionally, evidential learning is adopted to quantify the segmentation uncertainty of the model for boundary. In special, variational Dirichlet is introduced to characterize the distribution of the result probabilities, parameterized with collected evidence to quantify uncertainty. As a result, the model can reduce the segmentation uncertainties of boundaries by optimizing the parameters of the Dirichlet distribution. Experimental results on three datasets show that our method obtains superior segmentation performance. Chao Huang 0008, Yushu Shi, Wai Keung Wong, Chengliang Liu 0003, Wei Wang 0169, Zhihua Wang 0002, Jie Wen 0001 |
AAAI | 4 |
| 2025 | Deep Hierarchies and Invariant Disease-Indicative Feature Learning for Computer Aided Diagnosis of Multiple Fundus DiseasesabstractWith the advancement of computer vision, numerous models have been proposed for screening of fundus diseases. However, the recognition of multiple fundus diseases is often hampered by the simultaneous presence of multiple disease types and the confluence of lesion types in fundus images. This paper addresses these challenges by conceptualizing them as multi-level feature fusion and self-supervised disease-indicative feature learning problems. We decode fundus images at various levels of granularity to delineate scenarios wherein multiple diseases and lesions co-occur. To effectively integrate these features, we introduce a hierarchical vision transformer (HVT) that adeptly captures both inter-level and intra-level dependencies. A novel forward-attention module is proposed to enhance the integration of lower-level semantic information into higher semantic layers, thereby enriching the representation of complex features. Additionally, we introduce a novel self-supervised mask-consistent feature learner (MCFL). Unlike traditional mask-autoencoders that reconstruct original images using encoder-decoder structures, MCFL utilizes a teacher-student framework to reconstruct mask-consistent feature maps. In this setup, exponential moving averaging is employed to derive classification-guided features, serving as labels for reconstruction rather than merely reconstructing the original images. This innovative approach facilitates the extraction of disease-indicative features. Extensive experiments demonstrate that our method significantly outperforms existing state-of-the-art models. Wei Wang 0169, Xiaoling Luo 0001, Zhihao Wu 0002, Chengliang Liu 0003, Jie Wen 0001, Yong Xu 0001 |
AAAI | 5 |
| 2025 | Like an Ophthalmologist: Dynamic Selection Driven Multi-View Learning for Diabetic Retinopathy GradingabstractDiabetic retinopathy (DR), with its large patient population, has become a formidable threat to human visual health. In the clinical diagnosis of DR, multi-view fundus images are considered to be more suitable for DR diagnosis because of the wide coverage of the field of view. Therefore, different from most of the previous single-view DR grading methods, we design a dynamic selection-driven multi-view DR grading method to fit clinical scenarios better. Since lesion information plays a key role in DR diagnosis, previous methods usually boost the model performance by enhancing the lesion feature. However, during the actual diagnosis, ophthalmologists not only focus on the crucial parts, but also exclude irrelevant features to ensure the accuracy of judgment. To this end, we introduce the idea of dynamic selection and design a series of selection mechanisms from fine granularity to coarse granularity. In this work, we first introduce an Ophthalmic Image Reader (OIR) agent to provide the model with pixel-level prompts of suspected lesion areas. Moreover, a Multi-View Token Selection Module (MVTSM) is designed to prune redundant feature tokens and realize dynamic selection of key information. In the final decision stage, we dynamically fuse multi-view features through the novel Multi-View Mixture of Experts Module (MVMoEM), to enhance key views and reduce the impact of conflicting views. Extensive experiments on a large multi-view fundus image dataset with 34,452 images demonstrate that our method performs favorably against state-of-the-art models. Xiaoling Luo 0001, Qihao Xu, Huisi Wu, Chengliang Liu 0003, Zhihui Lai 0001, LinLin Shen |
AAAI | 4 |
| 2025 | Wavelet-based Global-Local Interaction Network with Cross-Attention for Multi-View Diabetic Retinopathy DetectionabstractMulti-view diabetic retinopathy (DR) detection has recently emerged as a promising method to address the issue of incomplete lesions faced by single-view DR. However, it is still challenging due to the variable sizes and scattered locations of lesions. Furthermore, existing multi-view DR methods typically merge multiple views without considering the correlations and redundancies of lesion information across them. Therefore, we propose a novel method to overcome the challenges of difficult lesion information learning and inadequate multi-view fusion. Specifically, we introduce a two-branch network to obtain both local lesion features and their global dependencies. The high-frequency component of the wavelet transform is used to exploit lesion edge information, which is then enhanced by global semantic to facilitate difficult lesion learning. Additionally, we present a cross-view fusion module to improve multi-view fusion and reduce redundancy. Experimental results on large public datasets demonstrate the effectiveness of our method. The code is open sourced on https://github.com/HuYongting/WGLIN. Yongting Hu, Chengliang Liu 0003, Xiaoling Luo 0001, Xiaoyan Dou, Qihao Xu, Yong Xu 0001 |
ICME | 3 |
| 2025 | Learning Compact Semantic Information for Incomplete Multi-View Missing Multi-Label ClassificationabstractMulti-view data involves various data forms, such as multi-feature, multi-sequence and multimodal data, providing rich semantic information for downstream tasks. The inherent challenge of incomplete multi-view missing multi-label learning lies in how to effectively utilize limited supervision and insufficient data to learn discriminative representation. Starting from the sufficiency of multi-view shared information for downstream tasks, we argue that the existing contrastive learning paradigms on missing multi-view data show limited consistency representation learning ability, leading to the bottleneck in extracting multi-view shared information. In response, we propose to minimize task-independent redundant information by pursuing the maximization of cross-view mutual information. Additionally, to alleviate the hindrance caused by missing labels, we develop a dual-branch soft pseudo-label cross-imputation strategy to improve classification performance. Extensive experiments on multiple benchmarks validate our advantages and demonstrate strong compatibility with both missing and complete data. Jie Wen 0001, Zhanyan Tang, Yuting He 0001, Mu Li 0005, Chengliang Liu 0003 |
ICML | 7 |
| 2025 | Mutual Learning for SAM Adaptation: A Dual Collaborative Network Framework for Source-Free Domain TransferabstractSegment Anything Model (SAM) has demonstrated remarkable zero-shot segmentation capabilities across various visual tasks. However, its performance degrades significantly when deployed in new target domains with substantial distribution shifts. While existing self-training methods based on fixed teacher-student architectures have shown improvements, they struggle to ensure that the teacher network consistently outperforms the student under severe domain shifts. To address this limitation, we propose a novel Collaborative Mutual Learning Framework for source-free SAM adaptation, leveraging dual-networks in a dynamic and cooperative manner. Unlike fixed teacher-student paradigms, our method dynamically assigns the teacher and student roles by evaluating the reliability of each collaborative network in each training iteration. Our framework incorporates a dynamic mutual learning mechanism with three key components: a direct alignment loss for knowledge transfer, a reverse distillation loss to encourage diversity, and a triplet relationship loss to refine feature representations. These components enhance the adaptation capabilities of the collaborative networks, enabling them to generalize effectively to target domains while preserving their pre-trained knowledge. Extensive experiments on diverse target domains demonstrate that our proposed framework achieves state-of-the-art adaptation performance. Wai Keung Wong, Chengliang Liu 0003, Xiaoling Luo 0001, Yong Xu 0001 |
ICML | 3 |
| 2025 | Enhancing Multimodal Protein Function Prediction Through Dual-Branch Dynamic Selection with Reconstructive Pre-TrainingabstractMultimodal protein features play a crucial role in protein function prediction. However, these features encompass a wide range of information, ranging from structural data and sequence features to protein attributes and interaction networks, making it challenging to decipher their complex interconnections. In this work, we propose a multimodal protein function prediction method (DSRPGO) by utilizing dynamic selection and reconstructive pre-training mechanisms. To acquire complex protein information, we introduce reconstructive pre-training to mine more fine-grained information with low semantic levels. Moreover, we put forward the Bidirectional Interaction Module (BInM) to facilitate interactive learning among multimodal features. Additionally, to address the difficulty of hierarchical multi-label classification in this task, a Dynamic Selection Module (DSM) is designed to select the feature representation that is most conducive to current protein function prediction. Our proposed DSRPGO model improves significantly in BPO, MFO, and CCO on human datasets, thereby outperforming other benchmark models. Xiaoling Luo 0001, Chengliang Liu 0003, Xiaopeng Jin, Jie Wen 0001 |
IJCAI | 3 |
| 2025 | Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-ThoughtabstractRecent advancements in reasoning capability of Multimodal Large Language Models (MLLMs) demonstrate its effectiveness in tackling complex visual tasks. However, existing MLLM-based Video Anomaly Detection (VAD) methods remain limited to shallow anomaly descriptions without deep reasoning. In this paper, we propose a new task named Video Anomaly Reasoning (VAR), which aims to enable deep analysis and understanding of anomalies in the video by requiring MLLMs to think explicitly before answering. To this end, we propose Vad-R1, an end-to-end MLLM-based framework for VAR. Specifically, we design a Perception-to-Cognition Chain-of-Thought (P2C-CoT) that simulates the human process of recognizing anomalies, guiding the MLLMs to reason about anomalies step-by-step. Based on the structured P2C-CoT, we construct Vad-Reasoning, a dedicated dataset for VAR. Furthermore, we propose an improved reinforcement learning algorithm AVA-GRPO, which explicitly incentivizes the anomaly reasoning capability of MLLMs through a self-verification mechanism with limited annotations. Experimental results demonstrate that Vad-R1 achieves superior performance, outperforming both open-source and proprietary models on VAD and VAR tasks. Chao Huang 0008, Benfeng Wang, Wei Wang 0169, Jie Wen 0001, Chengliang Liu 0003, Li Shen 0008, Xiaochun Cao |
NeurIPS | 5 |
| 2025 | Hierarchical Information Aggregation for Incomplete Multimodal Alzheimer's Disease DiagnosisabstractAlzheimer's Disease (AD) poses a significant health threat to the aging population, underscoring the critical need for early diagnosis to delay disease progression and improve patient quality of life. Recent advances in heterogeneous multimodal artificial intelligence (AI) have facilitated comprehensive joint diagnosis, yet practical clinical scenarios frequently encounter incomplete modalities due to factors like high acquisition costs or radiation risks. Moreover, traditional convolution-based architecture face inherent limitations in capturing long-range dependencies and handling heterogeneous medical data efficiently. To address these challenges, in our proposed heterogeneous multimodal diagnostic framework (HAD), we develop a multi-view Hilbert curve-based Mamba block and a hierarchical spatial feature extraction module to simultaneously capture local spatial features and global dependencies, effectively alleviating spatial discontinuities introduced by voxel serialization. Furthermore, to balance semantic consistency and modal specificity, we build a unified mutual information learning objective in the heterogeneous multimodal embedding space, which maintains effective learning of modality-specific information to avoid modality collapse caused by model preference. Extensive experiments demonstrate that our HAD significantly outperforms state-of-the-art methods in various modality-missing scenarios, providing an efficient and reliable solution for early-stage AD diagnosis. Chengliang Liu 0003, Yuanxi Que, Qihao Xu, Jie Wen 0001, Xiaoling Luo 0001 |
NeurIPS | 1 |
| 2025 | Saliency-Guided Selection Driven Multi-scale Network for Breast Tumor Detection
Xinfei Gu, Chengliang Liu 0003, Xiaoling Luo 0001, Qihao Xu, Zhihui Lai 0001, Heng Kong |
PRCV (13) | 2 |
| 2025 | Reliable Representation Learning for Incomplete Multi-View Missing Multi-Label ClassificationabstractAs a cross-topic of multi-view learning and multi-label classification, multi-view multi-label classification has gradually gained traction in recent years. The application of multi-view contrastive learning has further facilitated this process; however, the existing multi-view contrastive learning methods crudely separate the so-called negative pair, which largely results in the separation of samples belonging to the same category or similar ones. Besides, plenty of multi-view multi-label learning methods ignore the possible absence of views and labels. To address these issues, in this paper, we propose an incomplete multi-view missing multi-label classification network named RANK. In this network, a label-driven multi-view contrastive learning strategy is proposed to leverage supervised information to preserve the intra-view structure and perform the cross-view consistency alignment. Furthermore, we break through the view-level weights inherent in existing methods and propose a quality-aware subnetwork to dynamically assign quality scores to each view of each sample. The label correlation information is fully utilized in the final multi-label cross-entropy classification loss, effectively improving the discriminative power. Last but not least, our model is not only able to handle complete multi-view multi-label data, but also works on datasets with missing instances and labels. Extensive experiments confirm that our RANK outperforms existing state-of-the-art methods. Chengliang Liu 0003, Jie Wen 0001, Yong Xu 0001, Bob Zhang 0001, Liqiang Nie, Min Zhang 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Multi-view diabetic retinopathy grading via cross-view spatial alignment and adaptive vessel reinforcing
Xiaoyan Dou, Xiaoling Luo 0001, Zhihao Wu 0002, Chengliang Liu 0003, Tianyi Luo, Jie Wen 0001, Bingo Wing-Kuen Ling, Yong Xu 0001, Wei Wang 0169 |
Pattern Recognit. | 5 |
| 2025 | A Lesion-Fusion Neural Network for Multi-View Diabetic Retinopathy GradingabstractAs the most common complication of diabetes, diabetic retinopathy (DR) is one of the main causes of irreversible blindness. Automatic DR grading plays a crucial role in early diagnosis and intervention, reducing the risk of vision loss in people with diabetes. In these years, various deep-learning approaches for DR grading have been proposed. Most previous DR grading models are trained using the dataset of single-field fundus images, but the entire retina cannot be fully visualized in a single field of view. There are also problems of scattered location and great differences in the appearance of lesions in fundus images. To address the limitations caused by incomplete fundus features, and the difficulty in obtaining lesion information. This work introduces a novel multi-view DR grading framework, which solves the problem of incomplete fundus features by jointly learning fundus images from multiple fields of view. Furthermore, the proposed model combines multi-view inputs such as fundus images and lesion snapshots. It utilizes heterogeneous convolution blocks (HCB) and scalable self-attention classes (SSAC), which enhance the ability of the model to obtain lesion information. The experimental results show that our proposed method performs better than the benchmark methods on the large-scale dataset. Xiaoling Luo 0001, Qihao Xu, Zhihua Wang 0002, Chao Huang 0008, Chengliang Liu 0003, Xiaopeng Jin, Jianguo Zhang 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Foregroundness-Aware Task Disentanglement and Self-Paced Curriculum Learning for Domain Adaptive Object DetectionabstractUnsupervised domain adaptive object detection (UDA-OD) is a challenging problem since it needs to locate and recognize objects while maintaining the generalization ability across domains. Most existing UDA-OD methods directly integrate the adaptive modules into the detectors. This integration procedure can significantly sacrifice the detection performances, though it enhances the generalization ability. To solve this problem, we propose an effective framework, named foregroundness-aware task disentanglement and self-paced curriculum adaptation (FA-TDCA), to disentangle the UDA-OD task into four independent subtasks of source detector pretraining, classification adaptation, location adaptation, and target detector training. The disentanglement can transfer the knowledge effectively while maintaining the detection performance of our model. In addition, we propose a new metric, i.e., foregroundness, and use it to evaluate the confidence of the location result. We use both foregroundness and classification confidence to assess the label quality of the proposals. For effective knowledge transfer across domains, we utilize a self-paced curriculum learning paradigm to train adaptors and gradually improve the quality of the pseudolabels associated with the target samples. Experiment results indicate that our method achieves state-of-the-art results on four cross-domain object detection tasks. Linhui Xiao, Chengliang Liu 0003, Zhihao Wu 0002, Yong Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Spatial Continuity and Nonequal Importance in Salient Object Detection With Image-Category SupervisionabstractDue to the inefficiency of pixel-level annotations, weakly supervised salient object detection with image-category labels (WSSOD) has been receiving increasing attention. Previous works usually endeavor to generate high-quality pseudolabels to train the detectors in a fully supervised manner. However, we find that the detection performance is often limited by two types of noise contained in pseudolabels: 1) holes inside the object or at the edge and outliers in the background and 2) missing object portions and redundant surrounding regions. To mitigate the adverse effects caused by them, we propose local pixel correction (LPC) and key pixel attention (KPA), respectively, based on two key properties of desirable pseudolabels: 1) spatial continuity, meaning an object region consists of a cluster of adjacent points; and 2) nonequal importance, meaning pixels have different importance for training. Specifically, LPC fills holes and filters out outliers based on summary statistics of the neighborhood as well as its size. KPA directs the focus of training toward ambiguous pixels in multiple pseudolabels to discover more accurate saliency cues. To evaluate the effectiveness of our method, we design a simple yet strong baseline we call weakly supervised saliency detector with Transformer (WSSDT) and unify the proposed modules into WSSDT. Extensive experiments on five datasets demonstrate that our method significantly improves the baseline and outperforms all existing congeneric methods. Moreover, we establish the first benchmark to evaluate WSSOD robustness. The results show that our method can improve detection robustness as well. The code and robustness benchmark are available at https://github.com/Horatio9702/SCNI. Zhihao Wu 0002, Chengliang Liu 0003, Jie Wen 0001, Yong Xu 0001, Jian Yang 0003, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Attention-Induced Embedding Imputation for Incomplete Multi-View Partial Multi-Label ClassificationabstractAs a combination of emerging multi-view learning methods and traditional multi-label classification tasks, multi-view multi-label classification has shown broad application prospects. The diverse semantic information contained in heterogeneous data effectively enables the further development of multi-label classification. However, the widespread incompleteness problem on multi-view features and labels greatly hinders the practical application of multi-view multi-label classification. Therefore, in this paper, we propose an attention-induced missing instances imputation technique to enhance the generalization ability of the model. Different from existing incomplete multi-view completion methods, we attempt to approximate the latent features of missing instances in embedding space according to cross-view joint attention, instead of recovering missing views in kernel space or original feature space. Accordingly, multi-view completed features are dynamically weighted by the confidence derived from joint attention in the late fusion phase. In addition, we propose a multi-view multi-label classification framework based on label-semantic feature learning, utilizing the statistical weak label correlation matrix and graph attention network to guide the learning process of label-specific features. Finally, our model is compatible with missing multi-view and partial multi-label data simultaneously and extensive experiments on five datasets confirm the advancement and effectiveness of our embedding imputation method and multi-view multi-label classification model. Chengliang Liu 0003, Jinlong Jia, Jie Wen 0001, Xiaoling Luo 0001, Chao Huang 0008, Yong Xu 0001 |
AAAI | 1 |
| 2024 | A Two-Stage Information Extraction Network for Incomplete Multi-View Multi-Label ClassificationabstractRecently, multi-view multi-label classification (MvMLC) has received a significant amount of research interest and many methods have been proposed based on the assumptions of view completion and label completion. However, in real-world scenarios, multi-view multi-label data tends to be incomplete due to various uncertainties involved in data collection and manual annotation. As a result, the conventional MvMLC methods fail. In this paper, we propose a new two-stage MvMLC network to solve this incomplete MvMLC issue with partial missing views and missing labels. Different from the existing works, our method attempts to leverage the diverse information from the partially missing data based on the information theory. Specifically, our method aims to minimize task-irrelevant information while maximizing task-relevant information through the principles of information bottleneck theory and mutual information extraction. The first stage of our network involves training view-specific classifiers to concentrate the task-relevant information. Subsequently, in the second stage, the hidden states of these classifiers serve as input for an alignment model, an autoencoder-based mutual information extraction framework, and a weighted fusion classifier to make the final prediction. Extensive experiments performed on five datasets validate that our method outperforms other state-of-the-art methods. Code is available at https://github.com/KevinTan10/TSIEN. Ce Zhao, Chengliang Liu 0003, Jie Wen 0001, Zhanyan Tang |
AAAI | 3 |
| 2024 | HACDR-Net: Heterogeneous-Aware Convolutional Network for Diabetic Retinopathy Multi-Lesion SegmentationabstractDiabetic Retinopathy (DR), the leading cause of blindness in diabetic patients, is diagnosed by the condition of retinal multiple lesions. As a difficult task in medical image segmentation, DR multi-lesion segmentation faces the main concerns as follows. On the one hand, retinal lesions vary in location, shape, and size. On the other hand, because some lesions occupy only a very small part of the entire fundus image, the high proportion of background leads to difficulties in lesion segmentation. To solve the above problems, we propose a heterogeneous-aware convolutional network (HACDR-Net) that composes heterogeneous cross-convolution, heterogeneous modulated deformable convolution, and optional near-far-aware convolution. Our network introduces an adaptive aggregation module to summarize the heterogeneous feature maps and get diverse lesion areas in the heterogeneous receptive field along the channels and space. In addition, to solve the problem of the highly imbalanced proportion of focal areas, we design a new medical image segmentation loss function, Noise Adjusted Loss (NALoss). NALoss balances the predictive feature distribution of background and lesion by jointing Gaussian noise and hard example mining, thus enhancing awareness of lesions. We conduct the experiments on the public datasets IDRiD and DDR, and the experimental results show that the proposed method achieves better performance than other state-of-the-art methods. The code is open-sourced on github.com/xqh180110910537/HACDR-Net. Qihao Xu, Xiaoling Luo 0001, Chao Huang 0008, Chengliang Liu 0003, Jie Wen 0001, Yong Xu 0001 |
AAAI | 4 |
| 2024 | Deep Variational Incomplete Multi-View Clustering: Exploring Shared Clustering StructuresabstractIncomplete multi-view clustering (IMVC) aims to reveal shared clustering structures within multi-view data, where only partial views of the samples are available. Existing IMVC methods primarily suffer from two issues: 1) Imputation-based methods inevitably introduce inaccurate imputations, which in turn degrade clustering performance; 2) Imputation-free methods are susceptible to unbalanced information among views and fail to fully exploit shared information. To address these issues, we propose a novel method based on variational autoencoders. Specifically, we adopt multiple view-specific encoders to extract information from each view and utilize the Product-of-Experts approach to efficiently aggregate information to obtain the common representation. To enhance the shared information in the common representation, we introduce a coherence objective to mitigate the influence of information imbalance. By incorporating the Mixture-of-Gaussians prior information into the latent representation, our proposed method is able to learn the common representation with clustering-friendly structures. Extensive experiments on four datasets show that our method achieves competitive clustering performance compared with state-of-the-art methods. Gehui Xu, Jie Wen 0001, Chengliang Liu 0003, Lunke Fei, Wei Wang 0169 |
AAAI | 3 |
| 2024 | Partial Multi-View Multi-Label Classification via Semantic Invariance Learning and Prototype ModelingabstractThe difficulty of partial multi-view multi-label learning lies in coupling the consensus of multi-view data with the task relevance of multi-label classification, under the condition where partial views and labels are unavailable. In this paper, we seek to compress cross-view representation to maximize the proportion of shared information to better predict semantic tags. To achieve this, we establish a model consistent with the information bottleneck theory for learning cross-view shared representation, minimizing non-shared information while maintaining feature validity to help increase the purity of task-relevant information. Furthermore, we model multi-label prototype instances in the latent space and learn label correlations in a data-driven manner. Our method outperforms existing state-of-the-art methods on multiple public datasets while exhibiting good compatibility with both partial and complete data. Finally, we experimentally reveal the importance of condensing shared information under the premise of information balancing, in the process of multi-view information encoding and compression. Chengliang Liu 0003, Gehui Xu, Jie Wen 0001, Chao Huang 0008, Yong Xu 0001 |
ICML | 1 |
| 2024 | Language-Driven Cross-Modal Classifier for Zero-Shot Multi-Label Image RecognitionabstractLarge-scale pre-trained vision-language models (e.g., CLIP) have shown powerful zero-shot transfer capabilities in image recognition tasks. Recent approaches typically employ supervised fine-tuning methods to adapt CLIP for zero-shot multi-label image recognition tasks. However, obtaining sufficient multi-label annotated image data for training is challenging and not scalable. In this paper, we propose a new language-driven framework for zero-shot multi-label recognition that eliminates the need for annotated images during training. Leveraging the aligned CLIP multi-modal embedding space, our method utilizes language data generated by LLMs to train a cross-modal classifier, which is subsequently transferred to the visual modality. During inference, directly applying the classifier to visual inputs may limit performance due to the modality gap. To address this issue, we introduce a cross-modal mapping method that maps image embeddings to the language modality while retaining crucial visual information. Comprehensive experiments demonstrate that our method outperforms other zero-shot multi-label recognition methods and achieves competitive results compared to few-shot methods. Jie Wen 0001, Chengliang Liu 0003, Xiaozhao Fang, Yong Xu 0001, Zheng Zhang 0006 |
ICML | 3 |
| 2024 | Long Short-Term Dynamic Prototype Alignment Learning for Video Anomaly Detection
Chao Huang 0008, Jie Wen 0001, Chengliang Liu 0003 |
IJCAI | 3 |
| 2024 | Deep dual incomplete multi-view multi-label classification via label semantic-guided contrastive learning
Jinrong Cui, Yazi Xie, Chengliang Liu 0003, Qiong Huang 0001, Mu Li 0005, Jie Wen 0001 |
Neural Networks | 3 |
| 2024 | Weakly Supervised Video Anomaly Detection via Self-Guided Temporal Discriminative TransformerabstractWeakly supervised video anomaly detection is generally formulated as a multiple instance learning (MIL) problem, where an anomaly detector learns to generate frame-level anomaly scores under the supervision of MIL-based video-level classification. However, most previous works suffer from two drawbacks: 1) they lack ability to model temporal relationships between video segments and 2) they cannot extract sufficient discriminative features to separate normal and anomalous snippets. In this article, we develop a weakly supervised temporal discriminative (WSTD) paradigm, that aims to leverage both temporal relation and feature discrimination to mitigate the above drawbacks. To this end, we propose a transformer-styled temporal feature aggregator (TTFA) and a self-guided discriminative feature encoder (SDFE). Specifically, TTFA captures multiple types of temporal relationships between video snippets from different feature subspaces, while SDFE enhances the discriminative powers of features by clustering normal snippets and maximizing the separability between anomalous snippets and normal centers in embedding space. Experimental results on three public benchmarks indicate that WSTD outperforms state-of-the-art unsupervised and weakly supervised methods, which verifies the superiority of the proposed method. Chao Huang 0008, Chengliang Liu 0003, Jie Wen 0001, Lian Wu, Yong Xu 0001, Qiuping Jiang, Yaowei Wang 0001 |
IEEE Trans. Cybern. | 2 |
| 2024 | Projective Incomplete Multi-View ClusteringabstractDue to the rapid development of multimedia technology and sensor technology, multi-view clustering (MVC) has become a research hotspot in machine learning, data mining, and other fields and has been developed significantly in the past decades. Compared with single-view clustering, MVC improves clustering performance by exploiting complementary and consistent information among different views. Such methods are all based on the assumption of complete views, which means that all the views of all the samples exist. It limits the application of MVC, because there are always missing views in practical situations. In recent years, many methods have been proposed to solve the incomplete MVC (IMVC) problem and a kind of popular method is based on matrix factorization (MF). However, such methods generally cannot deal with new samples and do not take into account the imbalance of information between different views. To address these two issues, we propose a new IMVC method, in which a novel and simple graph regularized projective consensus representation learning model is formulated for incomplete multi-view data clustering task. Compared with the existing methods, our method not only can obtain a set of projections to handle new samples but also can explore information of multiple views in a balanced way by learning the consensus representation in a unified low-dimensional subspace. In addition, a graph constraint is imposed on the consensus representation to mine the structural information inside the data. Experimental results on four datasets show that our method successfully accomplishes the IMVC task and obtain the best clustering performance most of the time. Our implementation is available at https://github.com/Dshijie/PIMVC. Jie Wen 0001, Chengliang Liu 0003, Ke Yan 0003, Gehui Xu, Yong Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Information Recovery-Driven Deep Incomplete Multiview Clustering NetworkabstractIncomplete multiview clustering (IMC) is a hot and emerging topic. It is well known that unavoidable data incompleteness greatly weakens the effective information of multiview data. To date, existing IMC methods usually bypass unavailable views according to prior missing information, which is considered a second-best scheme based on evasion. Other methods that attempt to recover missing information are mostly applicable to specific two-view datasets. To handle these problems, in this article, we propose an information-recovery-driven-deep IMC network, termed as RecFormer. Concretely, a two-stage autoencoder network with self-attention structure is built to synchronously extract high-level semantic representations of multiple views and recover the missing data. Besides, we develop a recurrent graph reconstruction mechanism that cleverly leverages the restored views to promote representation learning and further data reconstruction. Visualization of recovery results are given and sufficient experimental results confirm that our RecFormer has obvious advantages over other top methods. Chengliang Liu 0003, Jie Wen 0001, Zhihao Wu 0002, Xiaoling Luo 0001, Chao Huang 0008, Yong Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Deep Double Incomplete Multi-View Multi-Label Learning With Incomplete Labels and Missing ViewsabstractView missing and label missing are two challenging problems in the applications of multi-view multi-label classification scenery. In the past years, many efforts have been made to address the incomplete multi-view learning or incomplete multi-label learning problem. However, few works can simultaneously handle the challenging case with both the incomplete issues. In this article, we propose a new incomplete multi-view multi-label learning network to address this challenging issue. The proposed method is composed of four major parts: view-specific deep feature extraction network, weighted representation fusion module, classification module, and view-specific deep decoder network. By, respectively, integrating the view missing information and label missing information into the weighted fusion module and classification module, the proposed method can effectively reduce the negative influence caused by two such incomplete issues and sufficiently explore the available data and label information to obtain the most discriminative feature extractor and classifier. Furthermore, our method can be trained in both supervised and semi-supervised manners, which has important implications for flexible deployment. Experimental results on five benchmarks in supervised and semi-supervised cases demonstrate that the proposed method can greatly enhance the classification performance on the difficult incomplete multi-view multi-label classification tasks with missing labels and missing views. Jie Wen 0001, Chengliang Liu 0003, Lunke Fei, Ke Yan 0003, Yong Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | DICNet: Deep Instance-Level Contrastive Network for Double Incomplete Multi-View Multi-Label ClassificationabstractIn recent years, multi-view multi-label learning has aroused extensive research enthusiasm. However, multi-view multi-label data in the real world is commonly incomplete due to the uncertain factors of data collection and manual annotation, which means that not only multi-view features are often missing, and label completeness is also difficult to be satisfied. To deal with the double incomplete multi-view multi-label classification problem, we propose a deep instance-level contrastive network, namely DICNet. Different from conventional methods, our DICNet focuses on leveraging deep neural network to exploit the high-level semantic representations of samples rather than shallow-level features. First, we utilize the stacked autoencoders to build an end-to-end multi-view feature extraction framework to learn the view-specific representations of samples. Furthermore, in order to improve the consensus representation ability, we introduce an incomplete instance-level contrastive learning scheme to guide the encoders to better extract the consensus information of multiple views and use a multi-view weighted fusion module to enhance the discrimination of semantic features. Overall, our DICNet is adept in capturing consistent discriminative representations of multi-view multi-label data and avoiding the negative effects of missing views and missing labels. Extensive experiments performed on five datasets validate that our method outperforms other state-of-the-art methods. Chengliang Liu 0003, Jie Wen 0001, Xiaoling Luo 0001, Chao Huang 0008, Zhihao Wu 0002, Yong Xu 0001 |
AAAI | 1 |
| 2023 | Incomplete Multi-View Multi-Label Learning via Label-Guided Masked View- and Category-Aware TransformersabstractAs we all know, multi-view data is more expressive than single-view data and multi-label annotation enjoys richer supervision information than single-label, which makes multi-view multi-label learning widely applicable for various pattern recognition tasks. In this complex representation learning problem, three main challenges can be characterized as follows: i) How to learn consistent representations of samples across all views? ii) How to exploit and utilize category correlations of multi-label to guide inference? iii) How to avoid the negative impact resulting from the incompleteness of views or labels? To cope with these problems, we propose a general multi-view multi-label learning framework named label-guided masked view- and category-aware transformers in this paper. First, we design two transformer-style based modules for cross-view features aggregation and multi-label classification, respectively. The former aggregates information from different views in the process of extracting view-specific features, and the latter learns subcategory embedding to improve classification performance. Second, considering the imbalance of expressive power among views, an adaptively weighted view fusion module is proposed to obtain view-consistent embedding features. Third, we impose a label manifold constraint in sample-level representation learning to maximize the utilization of supervised information. Last but not least, all the modules are designed under the premise of incomplete views and labels, which makes our method adaptable to arbitrary multi-view and multi-label data. Extensive experiments on five datasets confirm that our method has clear advantages over other state-of-the-art methods. Chengliang Liu 0003, Jie Wen 0001, Xiaoling Luo 0001, Yong Xu 0001 |
AAAI | 1 |
| 2023 | MVCINN: Multi-View Diabetic Retinopathy Detection Using a Deep Cross-Interaction Neural NetworkabstractDiabetic retinopathy (DR) is the main cause of irreversible blindness for working-age adults. The previous models for DR detection have difficulties in clinical application. The main reason is that most of the previous methods only use single-view data, and the single field of view (FOV) only accounts for about 13% of the FOV of the retina, resulting in the loss of most lesion features. To alleviate this problem, we propose a multi-view model for DR detection, which takes full advantage of multi-view images covering almost all of the retinal field. To be specific, we design a Cross-Interaction Self-Attention based Module (CISAM) that interfuses local features extracted from convolutional blocks with long-range global features learned from transformer blocks. Furthermore, considering the pathological association in different views, we use the feature jigsaw to assemble and learn the features of multiple views. Extensive experiments on the latest public multi-view MFIDDR dataset with 34,452 images demonstrate the superiority of our method, which performs favorably against state-of-the-art models. To the best of our knowledge, this work is the first study on the public large-scale multi-view fundus images dataset for DR detection. Xiaoling Luo 0001, Chengliang Liu 0003, Wai Keung Wong, Jie Wen 0001, Xiaopeng Jin, Yong Xu 0001 |
AAAI | 2 |
| 2023 | Highly Confident Local Structure Based Consensus Graph Learning for Incomplete Multi-view ClusteringabstractGraph-based multi-view clustering has attracted extensive attention because of the powerful clustering-structure representation ability and noise robustness. Considering the reality of a large amount of incomplete data, in this paper, we propose a simple but effective method for incomplete multi-view clustering based on consensus graph learning, termed as HCLS_CGL. Unlike existing methods that utilize graph constructed from raw data to aid in the learning of consistent representation, our method directly learns a consensus graph across views for clustering. Specifically, we design a novel confidence graph and embed it to form a confidence structure driven consensus graph learning model. Our confidence graph is based on an intuitive similar-nearest-neighbor hypothesis, which does not require any additional information and can help the model to obtain a high-quality consensus graph for better clustering. Numerous experiments are performed to confirm the effectiveness of our method. Jie Wen 0001, Chengliang Liu 0003, Gehui Xu, Zhihao Wu 0002, Chao Huang 0008, Lunke Fei, Yong Xu 0001 |
CVPR | 2 |
| 2023 | Localized and Balanced Efficient Incomplete Multi-view ClusteringabstractIn recent years, many incomplete multi-view clustering methods have been proposed to address the challenging unsupervised clustering issue on the multi-view data with missing views. However, most of the existing works are inapplicable to large-scale clustering task and their clustering results are unstable since these methods have high computational complexities and their results are produced by kmeans rather than their designed learning models. In this paper, we propose a new one-step incomplete multi-view clustering model, called Localized and Balanced Incomplete Multi-view Clustering (LBIMVC), to address these issues. Specifically, LBIMVC develops a new graph regularized incomplete multi-matrix-factorization model to obtain the unique clustering result by learning a consensus probability representation, where each element of the consensus representation can directly reflect the probability of the corresponding sample to the class. In addition, the proposed graph regularized model integrates geometric preserving and consensus representation learning into one term without introducing any extra constraint terms and parameters to explore the structure of data. Moreover, to avoid that samples are over divided into a few clusters, a balanced constraint is introduced to the model. Experimental results on four databases demonstrate that our method not only obtains competitive clustering performance, but also performs faster than some state-of-the-art methods. Jie Wen 0001, Gehui Xu, Chengliang Liu 0003, Lunke Fei, Chao Huang 0008, Wei Wang 0169, Yong Xu 0001 |
ACM Multimedia | 3 |
| 2023 | Masked Two-channel Decoupling Framework for Incomplete Multi-view Weak Multi-label LearningabstractMulti-view learning has become a popular research topic in recent years, but research on the cross-application of classic multi-label classification and multi-view learning is still in its early stages. In this paper, we focus on the complex yet highly realistic task of incomplete multi-view weak multi-label learning and propose a masked two-channel decoupling framework based on deep neural networks to solve this problem. The core innovation of our method lies in decoupling the single-channel view-level representation, which is common in deep multi-view learning methods, into a shared representation and a view-proprietary representation. We also design a cross-channel contrastive loss to enhance the semantic property of the two channels. Additionally, we exploit supervised information to design a label-guided graph regularization loss, helping the extracted embedding features preserve the geometric structure among samples. Inspired by the success of masking mechanisms in image and text analysis, we develop a random fragment masking strategy for vector features to improve the learning ability of encoders. Finally, it is important to emphasize that our model is fully adaptable to arbitrary view and label absences while also performing well on the ideal full data. We have conducted sufficient and convincing experiments to confirm the effectiveness and advancement of our model. Chengliang Liu 0003, Jie Wen 0001, Chao Huang 0008, Zhihao Wu 0002, Xiaoling Luo 0001, Yong Xu 0001 |
NeurIPS | 1 |
| 2023 | Balance guided incomplete multi-view spectral clustering
Lilei Sun, Jie Wen 0001, Chengliang Liu 0003, Lunke Fei, Lusi Li |
Neural Networks | 3 |
| 2023 | Selecting High-Quality Proposals for Weakly Supervised Object Detection With Bottom-Up Aggregated Attention and Phase-Aware LossabstractWeakly supervised object detection (WSOD) has received widespread attention since it requires only image-category annotations for detector training. Many advanced approaches solve this problem by a two-phase learning framework, that is, instance mining that classifies generated proposals via multiple instance learning, and instance refinement that iteratively refines bounding boxes using the supervision produced by the preceding stage. In this paper, we observe that the detection performance is usually limited by imprecise supervision, including part domination and untight boxes. To mitigate their adverse effects, we focus on selecting high-quality proposals as the supervision for WSOD. To be specific, for the issue of part domination, we propose bottom-up aggregated attention which incorporates low-level features from shallow layers to improve location representation of top-level features. In this manner, the proposals corresponding to entire objects can get high scores. Its advantage is that it can be flexibly plugged into the WSOD framework since there is no need to attach learnable parameters or learning branches. As regards the problem of untight boxes, we propose a phase-aware loss, which is the first work to measure supervision quality by the loss in the instance mining phase, to highlight correct boxes and suppress untight ones. In this work, we unify the proposed two modules into the framework of online instance classifier refinement. Extensive experiments on the PASCAL VOC and the MS COCO demonstrate that our method can significantly improve the performance of WSOD and achieve the state-of-the-art results. The code is available at https://github.com/Horatio9702/BUAA_PALoss. Zhihao Wu 0002, Chengliang Liu 0003, Jie Wen 0001, Yong Xu 0001, Jian Yang 0003, Xuelong Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Localized Sparse Incomplete Multi-View ClusteringabstractIncomplete multi-view clustering, which aims to solve the clustering problem on the incomplete multi-view data with partial view missing, has received more and more attention in recent years. Although numerous methods have been developed, most of the methods either cannot flexibly handle the incomplete multi-view data with arbitrary missing views or do not consider the negative factor of information imbalance among views. Moreover, some methods do not fully explore the local structure of all incomplete views. To tackle these problems, this paper proposes a simple but effective method, named localized sparse incomplete multi-view clustering (LSIMVC). Different from the existing methods, LSIMVC intends to learn a sparse and structured consensus latent representation from the incomplete multi-view data by optimizing a sparse regularized and novel graph embedded multi-view matrix factorization model. Specifically, in such a novel model based on the matrix factorization, a norm based sparse constraint is introduced to obtain the sparse low-dimensional individual representations and the sparse consensus representation. Moreover, a novel local graph embedding term is introduced to learn the structured consensus representation. Different from the existing works, our local graph embedding term aggregates the graph embedding task and consensus representation learning task into a concise term. Furthermore, to reduce the imbalance factor of incomplete multi-view learning, an adaptive weighted learning scheme is introduced to LSIMVC. Comprehensive experimental results performed on six incomplete multi-view databases verify that the performance of our LSIMVC is superior to the state-of-the-art IMC approaches. Chengliang Liu 0003, Zhihao Wu 0002, Jie Wen 0001, Yong Xu 0001, Chao Huang 0008 |
IEEE Trans. Multim. | 1 |
| 2022 | Deep Object Detection with Example Attribute Based Prediction ModulationabstractDeep object detectors suffer from the gradient contribution imbalance during training. In this paper, we point out that such imbalance can be ascribed to the imbalance in example attributes, e.g., difficulty and shape variation degree. We further propose example attribute based prediction modulation (EAPM) to address it. In EAPM, first, the attribute of an example is defined by the prediction and the corresponding ground truth. Then, a modulating factor w.r.t the example attribute is introduced to modulate the prediction error. Finally, the new prediction and the ground-truth are input into the loss function. Essentially, we adjust the gradients of examples with specific attributes to reweight their contribution on the global gradients. We apply EAPM with focal loss and balanced L1 loss to simultaneously solve the imbalance in classification and localization. The experimental results on MS COCO demonstrate that EAPM can bring substantial improvement for deep object detectors. Zhihao Wu 0002, Chengliang Liu 0003, Chao Huang 0008, Jie Wen 0001, Yong Xu 0001 |
ICASSP | 2 |
| 2022 | Hierarchical Graph Embedded Pose Regularity Learning via Spatio-Temporal Transformer for Abnormal Behavior DetectionabstractAbnormal behavior detection in surveillance video is a fundamental task in modern public security. Different from typical pixel-based solutions, pose-based approaches leverage low-dimensional and strongly-structured skeleton feature, which enables the anomaly detector to be immune to complex background noise and obtain higher efficiency. However, existing pose-based methods only utilize the pose of each individual independently while ignore the important interactions between individuals. In this paper, we present a hierarchical graph embedded pose regularity learning framework via spatio-temporal transformer, which leverages the strength of graph representation in encoding strongly-structured skeleton feature. Specifically, skeleton feature is encoded as the hierarchical graph representation, which jointly models the interactions among multiple individuals and the correlations among body joints within the same individual. Furthermore, a novel task-specific spatial-temporal graph transformer is designed to encode the hierarchical spatio-temporal graph embeddings of human skeletons and learn the regular patterns within normal training videos. Experimental results indicate that our method obtains superior performance over state-of-the-art methods on several challenging datasets. Chao Huang 0008, Zheng Zhang 0006, Chengliang Liu 0003, Jie Wen 0001, Yong Xu 0001, Yaowei Wang 0001 |
ACM Multimedia | 4 |
| 2022 | Pixel-Level Anomaly Detection via Uncertainty-aware Prototypical TransformerabstractPixel-level visual anomaly detection, which aims to recognize the abnormal areas from images, plays an important role in industrial fault detection and medical diagnosis. However, it is a challenging task due to the following reasons: i) the large variation of anomalies; and ii) the ambiguous boundary between anomalies and their normal surroundings. In this work, we present an uncertainty-aware prototypical transformer (UPformer), which takes into account both the diversity and uncertainty of anomaly to achieve accurate pixel-level visual anomaly detection. To this end, we first design a memory-guided prototype learning transformer encoder to learn and memorize the prototypical representations of anomalies for enabling the model to capture the diversity of anomalies. Additionally, an anomaly detection uncertainty quantizer is designed to learn the distributions of anomaly detection for measuring the anomaly detection uncertainty. Furthermore, an uncertainty-aware transformer decoder is proposed to leverage the detection uncertainties to guide the model to focus on the uncertain areas and generate the final detection results. As a result, our method achieves more accurate anomaly detection by combining the benefits of prototype learning and uncertainty estimation. Experimental results on five datasets indicate that our method achieves state-of-the-art anomaly detection performance. Chao Huang 0008, Chengliang Liu 0003, Zheng Zhang 0006, Zhihao Wu 0002, Jie Wen 0001, Qiuping Jiang, Yong Xu 0001 |
ACM Multimedia | 2 |
| 2022 | Weakly Supervised Video Anomaly Detection via Transformer-Enabled Temporal Relation LearningabstractWeakly supervised video anomaly detection is a challenging problem due to the lack of refined frame-level labels in training videos. Most prior works typically address it with the multiple instance learning paradigm, which divides a video into multiple snippets and trains a snippet classifier to distinguish anomalies from normal snippets via video-level classification loss. However, these solutions are limited in the insufficient representations. In this paper, we propose a novel weakly supervised temporal relation learning framework for anomaly detection, which efficiently explores the temporal relation between snippets and enhances the discriminative powers of features using only video-level labelled videos. To this end, we design a transformer-enabled feature encoder to convert the input task-agnostic features into discriminative task-specific features by mining the semantic similarity and position relation between snippets. As a result, our model can make a more accurate anomaly detection for current video snippet based on the learned discriminative features. Experimental results indicate that the proposed method is superior to existing state-of-the-art approaches, which demonstrates the effectiveness of our model. Dasheng Zhang, Chao Huang 0008, Chengliang Liu 0003, Yong Xu 0001 |
IEEE Signal Process. Lett. | 3 |