Yishu Liu 0001

dblp:91/3156-1 · DBLP profile ↗
← Back
32ranked-venue papers
7as first author
32since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 28 · 5 first-author · 28 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Cross Modal Fine-grained Alignment via Granularity-aware and Region-uncertain Modeling
abstract
Fine-grained image-text alignment is a pivotal challenge in multimodal learning, underpinning key applications such as visual question answering, image captioning, and vision-language navigation. Unlike global alignment, fine-grained alignment requires precise correspondence between localized visual regions and textual tokens, often hindered by noisy attention mechanisms and oversimplified modeling of cross-modal relationships. In this work, we identify two fundamental limitations of existing approaches: the lack of robust intra-modal mechanisms to assess the significance of visual and textual tokens, leading to poor generalization in complex scenes; and the absence of fine-grained uncertainty modeling, which fails to capture the one-to-many and many-to-one nature of region-word correspondences. To address these issues, we propose a unified approach that incorporates significance-aware and granularity-aware modeling and region-level uncertainty modeling. Our method leverages modality-specific biases to identify salient features without relying on brittle cross-modal attention, and represents region features as a mixture of Gaussian distributions to capture fine-grained uncertainty. Extensive experiments on Flickr30K and MS-COCO demonstrate that our approach achieves state-of-the-art performance across various backbone architectures, significantly enhancing the robustness and interpretability of fine-grained image-text alignment.
Haoming Zhou, Yishu Liu 0001, Bingzhi Chen, Yuncheng Jiang 0004
AAAI3
2026 DS2VP: Dynamically-Selected Spatially Visual Prompting
abstract
The significant effectiveness of prompt tuning for computer vision tasks has been extensively demonstrated in numerous studies. As a widely feasible solution, the spatial modeling paradigm aims to overcome the limitations of sequence modeling paradigm in capturing spatial relationships within images by learning a prompt token map and aligning it spatially with the image token map. However, such spatial modeling paradigms of visual prompt tuning still face two potential challenges: 1) Most existing methods fail to design individual prompts for different images, and the learned prompts have the same static effect on all images. 2) The strategy of existing methods overlooks the selection of key spatial information and indiscriminately prompts all information within the image. In this work, we propose a novel Dynamically-Selected and Spatial Visual Prompting, termed as DS2VP, which aims to effectively utilize the key spatial information of the input image and enable dynamic visual prompt selection. Specifically, our DS2VP approach is meticulously designed to leverage the key index generator to filter key regions of the image for determining the spatial target of prompts, thus enabling dynamic selection of prompts for different images. By adding prompt tokens at selected key locations, an image prompt fusion module is deployed by adapting the learnable prompt tokens into the input image tokens, further achieving a fine-grained spatial alignment. Moreover, we propose a multi-level prompt interaction module that facilitates interactions between visual prompts at different levels to enhance feature representations across various semantic levels. Extensive experiments on two challenging benchmarks for image classification have demonstrated the superiority of DS2VP over other state-of-the-art methods for visual prompt tuning.
Yishu Liu 0001, Bingzhi Chen, Guangming Lu 0002
IEEE Trans. Circuits Syst. Video Technol.1
2026 InfoARD: Enhancing Adversarial Robustness Distillation With Attack-Strength Adaptation and Mutual-Information Maximization
abstract
Adversarial distillation (AD) aims to mitigate deep neural networks' inherent vulnerability to adversarial attacks, thereby providing robust protection for compact models through teacher-student interactions. Despite advancements, existing AD studies still suffer from insufficient robustness due to the limitations of fixed attack strength and attention region shifts. To address these challenges, we propose a strength-adaptive Info-maximizing Adversarial Robustness Distillation paradigm, namely "InfoARD", which strategically incorporates the Attack-Strength Adaptation (ASA) and Mutual-Information Maximization (MIM) to enhance adversarial robustness against adversarial attacks and perturbations. Unlike previous adversarial training (AT) methods that utilize fixed attack strength, the ASA mechanism is designed to capture smoother and generalized classification boundaries by dynamically tailoring the attack strength based on the characteristics of individual instances. Benefiting from mutual information constraints, our MIM strategy ensures the student model effectively learns from various levels of feature representations and attention patterns, thereby deepening the student model's understanding of the teacher model's decision-making processes. Furthermore, a comprehensive multi-granularity distillation is conducted to capture knowledge across multiple dimensions, enabling a more effective transfer of knowledge from the teacher model to the student model. Note that our InfoARD can be seamlessly integrated into existing AD frameworks, further boosting the adversarial robustness of deep learning models. Extensive experiments on various challenging datasets consistently demonstrate the effectiveness and robustness of our InfoARD, surpassing previous state-of-the-art methods.
Ruihan Liu, Jieyi Cai, Yishu Liu 0001, Sudong Cai, Bingzhi Chen, Yulan Guo, Mohammed Bennamoun
IEEE Trans. Image Process.3
2025 Towards Robust Visual Question Answering via Prompt-Driven Geometric Harmonization
abstract
Visual Question Answering (VQA) has garnered significant attention as a crucial link between vision and language, aimed at generating accurate responses to visual queries. However, current VQA models still struggle with the challenges of minority class collapse and spurious semantic correlations posed by language bias and imbalanced distributions. To address these challenges, this paper proposes a novel Prompt-Driven Geometric Harmonization (PDGH) paradigm, which integrates both geometric structure and information entropy principles to enhance the ability of VQA models to generalize effectively across diverse scenarios. Specifically, our PDGH approach is meticulously designed to generate image-generated prompts that are guided by specific question cues, facilitating a more accurate and context-aware understanding of the visual content. Moreover, we project the prompt-visual-question and visual-question joint representations into a unified hypersphere space, applying feature weight self-orthogonality and prompt-information entropy correction constraints to optimize the margin, further alleviating minority class collapse and correcting language bias. To maintain the geometric integrity of the representation space, we introduce multi-space geometric contrast constraints to minimize the impact of spurious priors introduced during training. Finally, a semantic matrix is constructed for the coordinated joint representation to ensure that the learned instances are semantically consistent and improve reasoning ability. Extensive experiments on various general and medical VQA datasets demonstrate the consistent superiority of our PDGH approach over existing state-of-the-art baselines.
Yishu Liu 0001, Congcong Wen, Guangming Lu 0002, Bingzhi Chen
AAAI1
2025 OralXrays-9: Towards Hospital-Scale Panoramic X-ray Anomaly Detection via Personalized Multi-Object Query-Aware Mining
abstract
In clinical practice, panoramic dental radiography is a widely employed imaging technique that can provide a detailed and comprehensive view of dental structures and surrounding tissues for identifying various oral anomalies. However, due to the complexity of oral anomalies and the scarcity of available data, existing research still suffers from substantial challenges in automated oral anomaly detection. To this end, this paper presents a new hospital-scale panoramic X-ray benchmark, namely "OralXrays-91", which consists of 12,688 panoramic X-ray images with 84,113 meticulously annotated instances across nine common oral anomalies. Correspondingly, we propose a personalized Multi-Object Query-Aware Mining (MOQAM) paradigm, which jointly incorporates the Distribution-IoU Region Proposal Network (DI-RPN) and Class-Balanced Spherical Contrastive Regularization (CB-SCR) mechanisms to address the challenges posed by multi-scale variations and class-imbalanced distributions. To the best of our knowledge, this is the first attempt to develop AI-driven diagnostic systems specifically designed for multi-object oral anomaly detection, utilizing publicly available data resources. Extensive experiments on the newly-published OralXrays-9 dataset and real-world nature scenarios consistently demonstrate the superiority of our MOQAM in revolutionizing oral healthcare practices.
Bingzhi Chen, Sisi Fu, Xiaocheng Fang, Jieyi Cai, Minhua Lu, Yishu Liu 0001
CVPR7
2025 Revisiting DETR for Small Object Detection via Noise-Resilient Query Optimization
abstract
Despite advancements in Transformer-based detectors for small object detection (SOD), recent studies show that these detectors still face challenges due to inherent noise sensitivity in feature pyramid networks (FPN) and diminished query quality in existing label assignment strategies. In this paper, we propose a novel Noise-Resilient Query Optimization (NRQO) paradigm, which innovatively incorporates the Noise-Tolerance Feature Pyramid Network (NT-FPN) and the Pairwise-Similarity Region Proposal Network (PS-RPN). Specifically, NTFPN mitigates noise during feature fusion in FPN by preserving spatial and semantic information integrity. Unlike existing label assignment strategies, PS-RPN generates a sufficient number of high-quality positive queries by enhancing anchor-ground truth matching through position and shape similarities, without the need for additional hyperparameters. Extensive experiments on multiple benchmarks consistently demonstrate the superiority of NRQO over state-of-the-art baselines.
Xiaocheng Fang, Jieyi Cai, Wenxiu Cai, Yishu Liu 0001, Bingzhi Chen
ICME5
2025 Task-Aware Knowledge Prompt and Distillation for Cross-Domain Few-Shot Learning
abstract
Cross-Domain Few-Shot Learning (CD-FSL) aims to recognize unseen classes from target domains using only limited labeled samples. However, mainstream CD-FSL methods face two key challenges: (1) domain gap, arising from distributional differences between the source and target domains, and (2) overfitting, which occurs due to the small number of labeled samples in target domains, causing the model to overfit to these few samples. To address these challenges, we propose a novel Task-Aware Knowledge Prompt and Distillation (TKPD) method for CD-FSL, which integrates the Vision-Text Domain Prompt (VTDP) and Attribute-Task Knowledge Distillation (ATKD) modules. VTDP mitigates the domain gap by generating domain prompts to acquire domain-relevant knowledge, while ATKD strengthens the model by incorporating both homologous and heterogeneous knowledge, extending the knowledge base beyond the limited labeled samples and mitigating overfitting. Extensive experiments on 13 benchmark datasets validate the effectiveness of the proposed TKPD method.
Jun Liang 0002, Yunyu Zou, Yalong Cheng, Yishu Liu 0001, Bingzhi Chen
ICME6
2025 Enhancing Few-Shot Class-Incremental Learning via Cross-Modal Bias Alignment
abstract
Few-Shot Class-Incremental Learning (FSCIL) aims to learn new classes from limited samples while retaining knowledge of previously learned classes. Prompt tuning on large-scale pre-trained models has achieved certain success in FSCIL. However, current prompt-based approaches still suffer from the challenges of modality bias and catastrophic forgetting. To address these challenges, we propose a prompt tuning-based Cross-Modal Bias Alignment (CMBA) framework, which integrates Prompting Bias Alignment (PBA) and Dynamic Prototype Tuning (DPT). PBA mitigates modality bias by aligning the prompting biases between the visual and textual modalities. Meanwhile, DPT relieves catastrophic forgetting by introducing pseudo-sample features of old classes in the incremental sessions, which are generated by dynamically tuning prototypes that adapt to the evolving feature space. Extensive experiments on multiple FSCIL datasets demonstrate the consistent superiority of our CMBA approach over the existing state-of-the-art baselines.
Desen Wang, Yishu Liu 0001, Bingzhi Chen
ICME4
2025 Cause-Effect Driven Optimization for Robust Medical Visual Question Answering with Language Biases
abstract
Existing Medical Visual Question Answering (Med-VQA) models often suffer from language biases, where spurious correlations between question types and answer categories are inadvertently established. To address these issues, we propose a novel Cause-Effect Driven Optimization framework called CEDO, that incorporates three well-established mechanisms, i.e., Modality-driven Heterogeneous Optimization (MHO), Gradient-guided Modality Synergy (GMS), and Distribution-adapted Loss Rescaling (DLR), for comprehensively mitigating language biases from both causal and effectual perspectives. Specifically, MHO employs adaptive learning rates for specific modalities to achieve heterogeneous optimization, thus enhancing robust reasoning capabilities. Additionally, GMS leverages the Pareto optimization method to foster synergistic interactions between modalities and enforce gradient orthogonality to eliminate bias updates, thereby mitigating language biases from the effect side, i.e., shortcut bias. Furthermore, DLR is designed to assign adaptive weights to individual losses to ensure balanced learning across all answer categories, effectively alleviating language biases from the cause side, i.e., imbalance biases within datasets. Extensive experiments on multiple traditional and bias-sensitive benchmarks consistently demonstrate the robustness of CEDO over state-of-the-art competitors.
Huanjia Zhu, Yishu Liu 0001, Xiaozhao Fang, Guangming Lu 0002, Bingzhi Chen
IJCAI2
2025 Med-BiasX: Robust Medical Visual Question Answering with Language Biases
Huanjia Zhu, Yishu Liu 0001, Chengju Zhou, Guangming Lu 0002, Bingzhi Chen
MICCAI (14)2
2025 CauRDG: Enhancing Domain Generalization with Causal-Driven Semantic Consistency Reasoning
abstract
Domain generalization (DG) plays a pivotal role in enabling models to maintain robust performance across heterogeneous environments. However, existing DG methods are fundamentally constrained by two intertwined limitations: (1) causal misalignment, which stems from undifferentiated feature encoding that entangles causal mechanisms with environmental biases; (2)semantic conflict arises when conventional adaptation methods find it challenging to balance the preservation of class discriminability with the mitigation of domain-specific distribution discrepancies. To address these challenges of DG, we propose a novel Causal-Driven Semantic Consistency Reasoning (CauRDG) method, which synergistically integrates Prototype-Guided Causal Disentanglement (PGCD) and Dual-Space Semantic Disambiguation (DSSD). Specifically, PGCD constructs a causal framework that identifies stable relationships and decouples invariant mechanisms from domain-specific variations, preserving causal consistency while adapting to contextual differences. DSSD harnesses a dual-space paradigm, enhancing local categorical clarity and maintaining global conceptual unity, thus balancing domain-specific precision with cross-domain coherence. The robustness provided by CauRDG ensures robust extraction and interpretation of essential features by preserving invariant causal structures, thereby harmonizing discriminative semantics with domain-varying contexts. Extensive experiments on multiple benchmark datasets consistently demonstrate the effectiveness and superiority of our CauRDG over state-of-the-art baselines.
Zongxin Liu 0003, Yishu Liu 0001, Guangming Lu 0002, Xiaoling Luo 0001, Bingzhi Chen
ACM Multimedia2
2025 PET-GPRA: Rethinking PET with Gradient-Aware Prompting and Router-Free Adapters for Few-shot Class-Incremental Learning
abstract
Few-Shot Class-Incremental Learning (FSCIL) aims to continuously learn novel concepts from limited training samples without forgetting previously encountered classes. Recent advancements have leveraged Parameter-Efficient Tuning (PET) strategies on pre-trained models to enhance FSCIL performance. However, current PET-based FSCIL approaches still suffer from the challenges posed by catastrophic collapse of general prompt and limited adaptability of specific prompt . To this end, we redefine the function of the PET paradigm with both gradient-aware prompting (GAP) and router-free adapters (RFA) to boost the performance of FSCIL, termed as "PET-GPRA". To dynamically balance the retention of previously learned general knowledge and the acquisition of novel class information across sessions, the GAP paradigm adaptively adjusts the updated gradient of the general prompt by leveraging the angular relationship between the general knowledge gradient and the novel knowledge gradient. Meanwhile, the RFA mechanism utilizes the semantic similarity between class attributes to replace the routing network, guiding the integration of adapter information, in which adapters serve as specific prompts to enhance the adaptability. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority and effectiveness of our proposed PET-GPRA framework over state-of-the-art baselines.
Yishu Liu 0001, Desen Wang, Xiaoling Luo 0001, Bingzhi Chen, Guangming Lu 0002
ACM Multimedia1
2025 SG-FSL: Cross-Domain Few-Shot Learning with Style-Decoupled Augmentation and Gradient-Conflict Adjustment
abstract
Cross-Domain Few-Shot Learning (CD-FSL) aims to transfer knowledge acquired from a source domain with abundant data to the target domain with limited labeled samples. Recent advancements have enhanced model generalization through Perturbation Augmentation (PA), facilitating more effective knowledge transfer. However, PA-based CD-FSL methods still suffer from two critical challenges, i.e., (1) limited diversity of augmented samples, making it difficult to cover the true distribution of unseen domains, and (2) conflicting gradients during model optimization, where augmented and original samples drive the model's optimization in opposing directions. To address these issues, we propose a novel PA-based framework with Style-Decoupled Augmentation (SDA) and Gradient-Conflict Adjustment (GCA) for Cross-Domain Few-Shot Learning, which is termed ''SG-FSL''. Specifically, SDA decouples the source domain style into style weights and basis styles, generating diverse unseen styles by perturbing the style weights to reweight the basis styles. Meanwhile, GCA leverages the angular relationships between the domain-specific gradient directions of augmented and original features, adaptively adjusting the gradient directions of original features to ensure that the model acquires diverse domain knowledge without interference, guiding it toward conflict-free optimization. Comprehensive experiments on multiple benchmark datasets consistently demonstrate the effectiveness and superiority of our method over state-of-the-art baselines.
Yunyu Zou, Yishu Liu 0001, Jun Liang 0002, Bingzhi Chen
ACM Multimedia2
2025 Language-Bias-Resilient Visual Question Answering via Adaptive Multi-Margin Collaborative Debiasing
abstract
Language bias in Visual Question Answering (VQA) arises when models exploit spurious statistical correlations between question templates and answers, particularly in out-of-distribution scenarios, thereby neglecting essential visual cues and compromising genuine multimodal reasoning. Despite numerous efforts to enhance the robustness of VQA models, a principled understanding of how such bias originates and influences model behavior remains underdeveloped. In this paper, we address this gap through a comprehensive empirical and theoretical analysis, revealing that modality-specific gradient imbalances, which originate from the inherent heterogeneity of multimodal data, lead to skewed feature fusion and biased classifier weights. To alleviate these issues, we propose a novel Multi-Margin Collaborative Debiasing (MMCD) framework that adaptively integrates frequency-, confidence-, and difficulty-aware angular margins with a dynamic difficulty-aware contrastive learning mechanism, to dynamically reshape decision boundaries. Extensive experiments across multiple challenging VQA benchmarks confirm the consistent superiority of our proposed MMCD over state-of-the-art baselines in combating language bias.
Huanjia Zhu, Shuyuan Zheng, Yishu Liu 0001, Sudong Cai, Bingzhi Chen
NeurIPS3
2025 LBF-VQA: Towards Language Bias-Free Visual Question Answering With Multi-Space Collaborative Debiasing
Yishu Liu 0001, Huanjia Zhu, Bingzhi Chen, Xiaozhao Fang, Guangming Lu 0002, Shengli Xie 0001
IEEE Trans. Knowl. Data Eng.1
2025 Toward Robust Semi-Supervised Distribution Alignment Against Label Distribution Shift With Noisy Annotations
abstract
Deep learning-based AI models typically require a large amount of high-quality annotated data to achieve optimal performance. However, thelabel distribution shiftcaused by noisy annotations can lead to perturbations in the classification boundary, reducing the robustness and generalization capabilities of deep learning models. To mitigate this issue, we transform the problem of learning from noisy labels into a semi-supervised learning problem, and propose a novel Semi-Supervised Distribution Alignment (SSDA) framework that strategically integrates noise-robust distribution alignment within a unified semi-supervised learning paradigm for combating noisy labels. By leveraging the similarity distribution between historical predictions, the proposed SSDA approach benefits from a flexible multi-historical regression modeling strategy, which aims to identify high-confidence samples/pairs and recalibrate the label shift through pseudo-labels. Furthermore, our approach employs a comprehensive multi-granularity distribution adaptation strategy, incorporating both instance-wise and class-aware distribution alignment to quantitatively minimize semantic discrepancies across different mixed feature domains. In this way, our SSDA approach ultimately achieves more resilient and generalizable performance against label noise, even in the presence of substantial noise. Extensive experiments conducted on multiple simulated and real-world noisy benchmark datasets consistently demonstrate the superiority and effectiveness of our SSDA method compared to existing state-of-the-art baselines.
Bingzhi Chen, Zhanhao Ye, Yishu Liu 0001, Xiaozhao Fang, Guangming Lu 0002, Shengli Xie 0001, Xuelong Li 0001
IEEE Trans. Multim.3
2024 CariesXrays: Enhancing Caries Detection in Hospital-Scale Panoramic Dental X-rays via Feature Pyramid Contrastive Learning
abstract
Dental caries has been widely recognized as one of the most prevalent chronic diseases in the field of public health. Despite advancements in automated diagnosis across various medical domains, it remains a substantial challenge for dental caries detection due to its inherent variability and intricacies. To bridge this gap, we release a hospital-scale panoramic dental X-ray benchmark, namely “CariesXrays”, to facilitate the advancements in high-precision computer-aided diagnosis for dental caries. It comprises 6,000 panoramic dental X-ray images, with a total of 13,783 instances of dental caries, all meticulously annotated by dental professionals. In this paper, we propose a novel Feature Pyramid Contrastive Learning (FPCL) framework, that jointly incorporates feature pyramid learning and contrastive learning within a unified diagnostic paradigm for automated dental caries detection. Specifically, a robust dual-directional feature pyramid network (D2D-FPN) is designed to adaptively capture rich and informative contextual information from multi-level feature maps, thus enhancing the generalization ability of caries detection across different scales. Furthermore, our model is augmented with an effective proposals-prototype contrastive regularization learning (P2P-CRL) mechanism, which can flexibly bridge the semantic gaps among diverse dental caries with varying appearances, resulting in high-quality dental caries proposals. Extensive experiments on our newly-established CariesXrays benchmark demonstrate the potential of FPCL to make a significant social impact on caries diagnosis.
Bingzhi Chen, Sisi Fu, Yishu Liu 0001, Jiahui Pan 0003, Guangming Lu 0002, Zheng Zhang 0006
AAAI3
2024 Decoupled Self-Adaptive Distribution Regularization for Few-Shot Image Classification
abstract
The feature dispersion, arising from the inherent constraints of data scarcity, has emerged as a prominent challenge in the domain of few-shot learning. In this paper, we propose a novel Self-adaptive Distribution Regularization (SADR) approach, which can adaptively bridge the semantic gaps across distribution patterns and boundaries for learning from limited labeled data. Specifically, the technical core of our SADR approach is to decouple the feature embeddings into two discrete spaces: the intra-class and inter-class distributions, leading to robust and discriminative feature representations in a self-adaptive manner. To achieve meticulous similarity measurements while mitigating redundant feature information, an innovative regularized Brownian Distance Covariance (R-BDC) metric is strategically designed to simultaneously explore both the joint and marginal distributions present among diverse input samples. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority of our SADR approach over state-of-the-art baselines.
Bingzhi Chen, Haoming Zhou, Yishu Liu 0001, Guangming Lu 0002, Zheng Zhang 0006
ICASSP3
2024 Medical Vision-Language Representation Learning with Cross-Modal Multi-Teacher Contrastive Distillation
abstract
Medical vision-language representation learning has garnered considerable attention owing to its applicability to extracting generic representations from the image and text modality. However, it still remains challenging to acquire a more comprehensive understanding of intra- and inter-modal semantic knowledge. In this paper, we propose a Cross-Modal Multi-Teacher Contrastive Distillation (CMCD) architecture, which aims to comprehensively learn medical vision-language representation in a unified multi-teacher framework. Specifically, a cross-modal knowledge distillation (CKD) module is designed to refine reconstructed semantics under an additional supervision signal generated by momentum teachers from the other modality, achieving more robust semantic interaction across modalities. To better alleviate the heterogeneity and semantic gaps, the multi-level contrastive learning (MCL) module is conceived to align features of both intra- and inter-modal via contrastive learning from multi-level perspectives. Extensive experiments on two medical downstream tasks, i.e., Med-VQA and Med-ITC, demonstrate that our CMCD consistently outperforms the state-of-the-art methods.
Bingzhi Chen, Yishu Liu 0001, Jiahui Pan 0003, Meirong Ding
ICASSP3
2024 Rethinking Adversarial Robustness Distillation VIA Strength-Dependent Adaptive Regularization
abstract
Despite the progress achieved by existing adversarial distillation (AD) approaches, most mainstream models suffer from inadequate adversarial robustness, due to the challenges of fixed attack strength and unreliable teacher guidance. In this paper, we propose a novel Strength-Dependent Adaptive Regularization (SDAR) paradigm to reinforce the function of adversarial distillation with strength-adaptive adversarial attack (SAA) and multi-dimensional knowledge distillation (MKD). Different from the traditional adversarial training (AT) methods, the proposed SAA scheme dynamically assigns an adaptive and efficient attack strength for each instance, which aims to facilitate smoother classification boundaries. By incorporating dynamic strength coefficients, a comprehensive MKD strategy is designed to fully explore the valuable context information and narrow distribution discrepancies across teacher-student domains. Particularly, our SDAR paradigm can seamlessly integrate with the current AD frameworks, further enhancing the adversarial robustness of deep learning models. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority of SDAR over state-of-the-art baselines.
Bingzhi Chen, Shuobin Lin, Yishu Liu 0001, Zheng Zhang 0006, Guangming Lu 0002, Lewei He
ICME3
2024 Enhancing Few-Shot Classification without Forgetting Through Multi-level Contrastive Constraints
abstract
Most recent few-shot learning approaches are based on meta-learning with episodic training. However, prior studies encounter two crucial problems: (1) the presence of inductive bias, and (2) the occurrence of catastrophic forgetting. In this paper, we propose a novel Multi-Level Contrastive Constraints (MLCC) framework, that jointly integrates within-episode learning and across-episode learning into a unified interactive learning paradigm to solve these issues. Specifically, we employ a space-aware interaction modeling scheme to explore the correct inductive paradigms for each class between within-episode similarity/dis-similarity distributions. Additionally, with the aim of better utilizing former prior knowledge, a cross-stage distribution adaption strategy is designed to align the across-episode distributions from different time stages, thus reducing the semantic gap between existing and past prediction distribution. Extensive experiments on multiple few-shot datasets demonstrate the consistent superiority of MLCC approach over the existing state-of-the-art baselines.
Bingzhi Chen, Haoming Zhou, Yishu Liu 0001, Jiahui Pan 0003, Guangming Lu 0002
ICME3
2024 Robust Visual Question Answering With Contrastive-Adversarial Consistency Constraints
abstract
Visual cues and question semantics contribute to final answer predictions from distinct perspectives. However, inherent language bias confounds the relationship between visual and question cues, leading to a misguided preference for question semantics. Different from the existing studies that focus on inter-class discrimination, this paper proposes a robust visual question answering framework with contrastive-adversarial consistency constraints (CACC) at both inter- and intra-instance levels. From a fine-grained instance-level perspective, our approach initially introduces an effective inter-instance contrastive constraint to perform adaptive bias rectification. To enhance intra-instance invariance and reduce information redundancy, we refine the concept of semantic structure relationships by constructing intra-instance adversarial constraints using the Hilbert-Schmidt Independence Criterion (HSIC) independence criterion. Benefitting from both inter- and intra-instance perspectives, our method can effectively alleviate these language biases, enhancing the overall robustness of the representation. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority of our CACC over state-of-the-art baselines.
Meirong Ding, Yishu Liu 0001, Guangming Lu 0002, Bingzhi Chen
ICME3
2024 Enhancing Cross-Modal Retrieval via Visual-Textual Prompt Hashing
Bingzhi Chen, Zhongqi Wu, Yishu Liu 0001, Guangming Lu 0002, Zheng Zhang 0006
IJCAI3
2024 Medical Cross-Modal Prompt Hashing with Robust Noisy Correspondence Learning
Yishu Liu 0001, Zhongqi Wu, Bingzhi Chen, Zheng Zhang 0006, Guangming Lu 0002
MICCAI (3)1
2024 Stay Focused is All You Need for Adversarial Robustness
Bingzhi Chen, Ruihan Liu, Yishu Liu 0001, Xiaozhao Fang, Jiahui Pan 0003, Guangming Lu 0002, Zheng Zhang 0006
ACM Multimedia3
2024 Prototype-Guided Dual-Transformer Reasoning for Video Individual Counting
abstract
Video Individual Counting (VIC), which focuses on accurately tallying the total number of individuals in a video without duplication, is crucial for urban public space management and densely-populated areas planning. Existing methods suffer from limitations in terms of expensive manual annotation, and the efficiency of location or detection algorithms. In this work, we contribute a novel Prototype-guided Dual-Transformer Reasoning framework, termed PDTR, which takes both similarity and difference of adjacent frames into account to achieve accurate counting in an end-to-end regression manner. Specifically, we first design a multi-receptive field feature fusion module to acquire initial comprehensive representations. Subsequently, the dynamic prototype generation module memorizes consistent representations of similar information to generate prototypes. Additionally, to further dig out the shared and private features from different frames, a prototype cross-guided decoder and a privacy-decoupling module are designed. Extensive experiments conducted on two existing VIC datasets, consistently demonstrate the superiority of PDTR over state-of-the-art baselines.
Yishu Liu 0001, Huafeng Li 0001, Jinxing Li 0003, Guangming Lu 0002
ACM Multimedia2
2024 Combating Visual Question Answering Hallucinations via Robust Multi-Space Co-Debias Learning
Yishu Liu 0001, Huanjia Zhu, Yuncheng Jiang 0004, Zheng Zhang 0006, Bingzhi Chen
ACM Multimedia2
2024 Deep Fuzzy Multiteacher Distillation Network for Medical Visual Question Answering
abstract
Medical visual question answering (medical VQA) is a critical cross-modal interaction task that garnered considerable attention in the medical domain. Several existing methods commonly leverage the vision-and-language pretraining paradigms to mitigate the limitation of small-scale data. Nevertheless, most of them still suffer from two challenges that remain for further research: 1) limited research focuses on distilling representation from a complete modality to guide the representation learning of masked data in other modalities. 2) Multimodal fusion based on self-attention mechanisms cannot effectively handle the inherent uncertainty and vagueness of information interaction across modalities. To mitigate these issues, in this article, we propose a novel deep fuzzy multiteacher distillation (DFMD) network for medical VQA, which can take advantage of fuzzy logic to model the uncertainties from vison-language representations across modalities in a multiteacher framework. Specifically, a multiteacher knowledge distillation module is conceived to assist in reconstructing the missing semantics under the supervision signal generated by teachers from the other complete modality, achieving more robust semantic interaction across modalities. Incorporating insights from the fuzzy logic theory, we propose a noise-robust encoder called FuzBERT that enables our DFMD model to reduce the imprecision and ambiguity in feature representation during the multimodal interaction process. To the best of our knowledge, our work isthe first attemptto combine the fuzzy logic theory with the transformer-based encoder to effectively learn multimodal representation for medical VQA. Experimental results on the VQA-RAD and SLAKE datasets consistently demonstrate the superiority of our proposed DFMD method over state-of-the-art baselines.
Yishu Liu 0001, Bingzhi Chen, Shuihua Wang, Guangming Lu 0002, Zheng Zhang 0006
IEEE Trans. Fuzzy Syst.1
2024 Contrastive Multi-Bit Collaborative Learning for Deep Cross-Modal Hashing
abstract
Deep cross-modal hashing, as a promising fast similarity search technique, has attracted broad interest and obtained great success owing to its outstanding representation capability and computational efficiency. Since the inconsistent feature representations and distributions of different modalities (i.e., image and text), prior studies primarily focus on preserving pairwise similarity with global embedding, but fail to further utilize detailed local representations to effectively align such heterogeneous data to jointly bridge the heterogeneous and semantic gaps across modalities. Meanwhile, typical learning networks can learn onlyonefixed-length hash code rather than multi-length ones, leading to extremely limited flexibility and scalability. To tackle these issues, this paper proposes a novelContrastive Multi-bit Collaborative Learning(CMCL) network, which hierarchically aligns both global and local features among different modalities and simultaneously generates multi-length hash codes (i.e., 16-, 32-, 64-bits) in one unified transformer-based framework. Specifically, we design a novel cross-modal contrastive alignment module to simultaneously bridge the heterogeneous and semantic gaps across modalities via global and local contrastive learning. Moreover, we propose a multi-bit collaborative optimization module to synchronously produce multi-length hash codes under the explicit guidance of one auxiliary online hash learner with a longer length (i.e., 128-bit). As such, our CMCL framework can jointly alleviate the heterogeneity among modalities from a hierarchical perspective and collaboratively explore the correlations between multi-bit hash codes, thereby yielding multi-length discriminative hash codes in a one-stop learning manner. Comprehensive experiments demonstrate the consistent superiority of our CMCL in multi-bit hash code learning over the state-of-the-art cross-modal hashing baselines.
Qingpeng Wu, Zheng Zhang 0006, Yishu Liu 0001, Liqiang Nie
IEEE Trans. Knowl. Data Eng.3
2023 Combating Medical Label Noise via Robust Semi-supervised Contrastive Learning
Bingzhi Chen, Zhanhao Ye, Yishu Liu 0001, Zheng Zhang 0006, Jiahui Pan 0003, Guangming Lu 0002
MICCAI (1)3
2023 Multi-Granularity Interactive Transformer Hashing for Cross-modal Retrieval
abstract
With the powerful representation ability and privileged efficiency, deep cross-modal hashing (DCMH) has become an emerging fast similarity search technique. Prior studies primarily focus on exploring pairwise similarities across modalities, but fail to comprehensively capture the multi-grained semantic correlations during intra- and inter-modal negotiation. To tackle this issue, this paper proposes a novel Multi-granularity Interactive Transformer Hashing (MITH) network, which hierarchically considers both coarse- and fine-grained similarity measurements across different modalities in one unified transformer-based framework. To the best of our knowledge, this is the first attempt for multi-granularity transformer-based cross-modal hashing. Specifically, a well-designed distilled intra-modal interaction module is deployed to excavate modality-specific concept knowledge with global-local knowledge distillation under the guidance of implicit conceptual category-level representations. Moreover, we construct a contrastive inter-modal alignment module to mine modality-independent semantic concept correspondences with instance- and token-wise contrastive learning, respectively. Such a collaborative learning paradigm can jointly alleviate the heterogeneity and semantic gaps among different modalities from a multi-granularity perspective, yielding discriminative modality-invariant hash codes. Extensive experiments on multiple representative cross-modal datasets demonstrate the consistent superiority of MITH over the existing state-of-the-art baselines. The codes are available at https://github.com/DarrenZZhang/MITH.
Yishu Liu 0001, Qingpeng Wu, Zheng Zhang 0006, Guangming Lu 0002
ACM Multimedia1
2021 Deep Active Context Estimation for Automated COVID-19 Diagnosis
abstract
Many studies on automated COVID-19 diagnosis have advanced rapidly with the increasing availability of large-scale CT annotated datasets. Inevitably, there are still a large number of unlabeled CT slices in the existing data sources since it requires considerable consuming labor efforts. Notably, cinical experience indicates that the neighboring CT slices may present similar symptoms and signs. Inspired by such wisdom, we propose DACE, a novel CNN-based deep active context estimation framework, which leverages the unlabeled neighbors to progressively learn more robust feature representations and generate a well-performed classifier for COVID-19 diagnosis. Specifically, the backbone of the proposed DACE framework is constructed by a well-designed Long-Short Hierarchical Attention Network (LSHAN), which effectively incorporates two complementary attention mechanisms, i.e., short-range channel interactions (SCI) module and long-range spatial dependencies (LSD) module, to learn the most discriminative features from CT slices. To make full use of such available data, we design an efficient context estimation criterion to carefully assign the additional labels to these neighbors. Benefiting from two complementary types of informative annotations from -nearest neighbors, i.e., the majority of high-confidence samples with pseudo labels and the minority of low-confidence samples with hand-annotated labels, the proposed LSHAN can be fine-tuned and optimized in an incremental learning manner. Extensive experiments on the Clean-CC-CCII dataset demonstrate the superior performance of our method compared with the state-of-the-art baselines.
Bingzhi Chen, Yishu Liu 0001, Zheng Zhang 0006, Yingjian Li 0001, Zhao Zhang 0001, Guangming Lu 0002, Hongbing Yu
ACM Trans. Multim. Comput. Commun. Appl.2