Bingzhi Chen

dblp:34/7694 · DBLP profile ↗
← Back
62ranked-venue papers
19as first author
57since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 48 · 14 first-author · 46 since 2021Artificial intelligence and machine learning · 14 · 5 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 BayesVQA: Energy-Guided Bayesian Debiasing for Language-Bias-Robust Visual Question Answering
abstract
Numerous studies have demonstrated that Visual Question Answering (VQA) models are vulnerable to language priors and dataset biases, often leading to spurious correlations between questions and answers. As a result, these models excessively rely on linguistic cues, neglecting essential visual information and causing representational distortions. To address this challenge, we propose a novel Bayesian debiasing framework termed BayesVQA, which integrates three carefully designed mechanisms: Energy-guided Prior Variance (EPV), Energy-guided Posterior Sampling (EPS), and Energy-guided Likelihood Reweighting (ELR). Specifically, we explicitly decompose each sample's latent representation into a biased feature and a stochastic corrective perturbation δ. Using a Bayesian formulation, we model the posterior distribution of the perturbation δ conditioned on the predictive uncertainty, quantified via calibrated energy scores. To mitigate language bias, the posterior is optimized through energy-driven variational inference with an uncertainty-adaptive prior and sampling strategy. Moreover, the ELR mechanism incorporates an energy-based weighting of the reconstruction objective and enforces an energy-coherence constraint to emphasize challenging, high-uncertainty instances and align model confidence before and after debiasing. Extensive experiments conducted across multiple standard VQA benchmarks consistently validate the superior performance of our BayesVQA method over state-of-the-art competitors under distributional shifts and challenging bias conditions.
Huanjia Zhu, Xiangwen Deng, Qinghao Zhong, Bingzhi Chen
AAAI5
2026 Cross Modal Fine-grained Alignment via Granularity-aware and Region-uncertain Modeling
abstract
Fine-grained image-text alignment is a pivotal challenge in multimodal learning, underpinning key applications such as visual question answering, image captioning, and vision-language navigation. Unlike global alignment, fine-grained alignment requires precise correspondence between localized visual regions and textual tokens, often hindered by noisy attention mechanisms and oversimplified modeling of cross-modal relationships. In this work, we identify two fundamental limitations of existing approaches: the lack of robust intra-modal mechanisms to assess the significance of visual and textual tokens, leading to poor generalization in complex scenes; and the absence of fine-grained uncertainty modeling, which fails to capture the one-to-many and many-to-one nature of region-word correspondences. To address these issues, we propose a unified approach that incorporates significance-aware and granularity-aware modeling and region-level uncertainty modeling. Our method leverages modality-specific biases to identify salient features without relying on brittle cross-modal attention, and represents region features as a mixture of Gaussian distributions to capture fine-grained uncertainty. Extensive experiments on Flickr30K and MS-COCO demonstrate that our approach achieves state-of-the-art performance across various backbone architectures, significantly enhancing the robustness and interpretability of fine-grained image-text alignment.
Haoming Zhou, Yishu Liu 0001, Bingzhi Chen, Yuncheng Jiang 0004
AAAI4
2026 S2PD: Physically grounded augmentations and stable parameter updates for cross-domain few-shot learning
Shuai Kang, Yunyu Zou, Ziteng Hong, Bingzhi Chen, Guangming Lu 0002
Neurocomputing5
2026 DS2VP: Dynamically-Selected Spatially Visual Prompting
abstract
The significant effectiveness of prompt tuning for computer vision tasks has been extensively demonstrated in numerous studies. As a widely feasible solution, the spatial modeling paradigm aims to overcome the limitations of sequence modeling paradigm in capturing spatial relationships within images by learning a prompt token map and aligning it spatially with the image token map. However, such spatial modeling paradigms of visual prompt tuning still face two potential challenges: 1) Most existing methods fail to design individual prompts for different images, and the learned prompts have the same static effect on all images. 2) The strategy of existing methods overlooks the selection of key spatial information and indiscriminately prompts all information within the image. In this work, we propose a novel Dynamically-Selected and Spatial Visual Prompting, termed as DS2VP, which aims to effectively utilize the key spatial information of the input image and enable dynamic visual prompt selection. Specifically, our DS2VP approach is meticulously designed to leverage the key index generator to filter key regions of the image for determining the spatial target of prompts, thus enabling dynamic selection of prompts for different images. By adding prompt tokens at selected key locations, an image prompt fusion module is deployed by adapting the learnable prompt tokens into the input image tokens, further achieving a fine-grained spatial alignment. Moreover, we propose a multi-level prompt interaction module that facilitates interactions between visual prompts at different levels to enhance feature representations across various semantic levels. Extensive experiments on two challenging benchmarks for image classification have demonstrated the superiority of DS2VP over other state-of-the-art methods for visual prompt tuning.
Yishu Liu 0001, Bingzhi Chen, Guangming Lu 0002
IEEE Trans. Circuits Syst. Video Technol.3
2026 InfoARD: Enhancing Adversarial Robustness Distillation With Attack-Strength Adaptation and Mutual-Information Maximization
abstract
Adversarial distillation (AD) aims to mitigate deep neural networks' inherent vulnerability to adversarial attacks, thereby providing robust protection for compact models through teacher-student interactions. Despite advancements, existing AD studies still suffer from insufficient robustness due to the limitations of fixed attack strength and attention region shifts. To address these challenges, we propose a strength-adaptive Info-maximizing Adversarial Robustness Distillation paradigm, namely "InfoARD", which strategically incorporates the Attack-Strength Adaptation (ASA) and Mutual-Information Maximization (MIM) to enhance adversarial robustness against adversarial attacks and perturbations. Unlike previous adversarial training (AT) methods that utilize fixed attack strength, the ASA mechanism is designed to capture smoother and generalized classification boundaries by dynamically tailoring the attack strength based on the characteristics of individual instances. Benefiting from mutual information constraints, our MIM strategy ensures the student model effectively learns from various levels of feature representations and attention patterns, thereby deepening the student model's understanding of the teacher model's decision-making processes. Furthermore, a comprehensive multi-granularity distillation is conducted to capture knowledge across multiple dimensions, enabling a more effective transfer of knowledge from the teacher model to the student model. Note that our InfoARD can be seamlessly integrated into existing AD frameworks, further boosting the adversarial robustness of deep learning models. Extensive experiments on various challenging datasets consistently demonstrate the effectiveness and robustness of our InfoARD, surpassing previous state-of-the-art methods.
Ruihan Liu, Jieyi Cai, Yishu Liu 0001, Sudong Cai, Bingzhi Chen, Yulan Guo, Mohammed Bennamoun
IEEE Trans. Image Process.5
2026 Toward Bidirectional Adaptability for Few-Shot Class-Incremental Learning With Forward-Backward Knowledge Transfer
abstract
The development of Deep Neural Networks (DNNs) has enabled AI-driven models to excel in recognizing a limited set of classes within static environments. As AI systems progress, few-shot class-incremental learning (FSCIL) aims to expand their understanding of novel classes from minimal samples while retaining knowledge of previously encountered ones. However, most existing FSCIL models face significant challenges, includinginadequate adaptabilityandcatastrophic forgetting, which hinder their ability to maintain robust forward and backward learning capabilities. To address these issues, this paper proposes a novel Forward-Backward Knowledge Transfer (FBKT) paradigm, which strategically integrates forward distribution adaptation (FDA) and backward semantic alignment (BSA) mechanisms to achieve bidirectional adaptability in knowledge transfer. The FDA mechanism enhances forward adaptability by expanding and reserving the embedding space for new classes using semantic-irrelevant masked images as virtual negative classes, thereby mitigating data overfitting. It also employs self-supervised representation learning to utilize semantic-relevant local embeddings as additional positive samples, fostering class separation and generalization. Meanwhile, the BSA mechanism ensures the semantic consistency of previously learned classes across sessions during class-incremental learning, promoting smoother backward adaptability and reducing model degradation. Extensive experiments conducted on multiple benchmark datasets consistently highlight the superior performance and effectiveness of our FBKT compared to state-of-the-art methods.
Bingzhi Chen, Sudong Cai, Xiaozhao Fang, Mohammed Bennamoun, Shengli Xie 0001
IEEE Trans. Multim.1
2025 Towards Robust Visual Question Answering via Prompt-Driven Geometric Harmonization
abstract
Visual Question Answering (VQA) has garnered significant attention as a crucial link between vision and language, aimed at generating accurate responses to visual queries. However, current VQA models still struggle with the challenges of minority class collapse and spurious semantic correlations posed by language bias and imbalanced distributions. To address these challenges, this paper proposes a novel Prompt-Driven Geometric Harmonization (PDGH) paradigm, which integrates both geometric structure and information entropy principles to enhance the ability of VQA models to generalize effectively across diverse scenarios. Specifically, our PDGH approach is meticulously designed to generate image-generated prompts that are guided by specific question cues, facilitating a more accurate and context-aware understanding of the visual content. Moreover, we project the prompt-visual-question and visual-question joint representations into a unified hypersphere space, applying feature weight self-orthogonality and prompt-information entropy correction constraints to optimize the margin, further alleviating minority class collapse and correcting language bias. To maintain the geometric integrity of the representation space, we introduce multi-space geometric contrast constraints to minimize the impact of spurious priors introduced during training. Finally, a semantic matrix is constructed for the coordinated joint representation to ensure that the learned instances are semantically consistent and improve reasoning ability. Extensive experiments on various general and medical VQA datasets demonstrate the consistent superiority of our PDGH approach over existing state-of-the-art baselines.
Yishu Liu 0001, Congcong Wen, Guangming Lu 0002, Bingzhi Chen
AAAI6
2025 OralXrays-9: Towards Hospital-Scale Panoramic X-ray Anomaly Detection via Personalized Multi-Object Query-Aware Mining
abstract
In clinical practice, panoramic dental radiography is a widely employed imaging technique that can provide a detailed and comprehensive view of dental structures and surrounding tissues for identifying various oral anomalies. However, due to the complexity of oral anomalies and the scarcity of available data, existing research still suffers from substantial challenges in automated oral anomaly detection. To this end, this paper presents a new hospital-scale panoramic X-ray benchmark, namely "OralXrays-91", which consists of 12,688 panoramic X-ray images with 84,113 meticulously annotated instances across nine common oral anomalies. Correspondingly, we propose a personalized Multi-Object Query-Aware Mining (MOQAM) paradigm, which jointly incorporates the Distribution-IoU Region Proposal Network (DI-RPN) and Class-Balanced Spherical Contrastive Regularization (CB-SCR) mechanisms to address the challenges posed by multi-scale variations and class-imbalanced distributions. To the best of our knowledge, this is the first attempt to develop AI-driven diagnostic systems specifically designed for multi-object oral anomaly detection, utilizing publicly available data resources. Extensive experiments on the newly-published OralXrays-9 dataset and real-world nature scenarios consistently demonstrate the superiority of our MOQAM in revolutionizing oral healthcare practices.
Bingzhi Chen, Sisi Fu, Xiaocheng Fang, Jieyi Cai, Minhua Lu, Yishu Liu 0001
CVPR1
2025 Learning Hierarchical Attribute Prompt for Vision-Language Models
abstract
Prompt learning is a common strategy for adapting Visual Language Models (VLMs) to downstream tasks by fine-tuning prompts for task-specific performance. However, existing methods face two key challenges: overfitting to base classes, which limits generalization to novel classes, and the dependence on manually generated or LLM-based descriptions, which are time-consuming and error-prone. To address these issues, we propose the Learning Hierarchical Attribute Prompt (LHAP) method, which introduces fine-grained semantic alignment through hierarchical prompts. By autonomously extracting visual attributes from images, LHAP generates local-level prompts (LLP) to capture fine-grained semantics and global-level prompts (GLP) to model overall semantics. The combination of LLP and GLP not only improves generalization but also mitigates errors and inefficiencies from manual or LLM-based descriptions. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority and robustness of LHAP over state-of-the-art methods.
Jun Liang 0002, Yunyu Zou, Yalong Cheng, Bingzhi Chen
ICASSP6
2025 A2GP-SF: Enhancing Few-shot Class Incremental Learning via Attribute Generative Prompting and Adaptive Sharpness Flattening
abstract
Few-shot Class Incremental Learning (FSCIL) aims to incrementally learn new classes with limited examples while retaining knowledge of previously learned classes. Recent advancements in prompt tuning for large pre-trained models have shown promise in FSCIL. However, current FSCIL methods still suffer from challenges like insufficient plasticity and limited generalization. To tackle these challenges, we propose a novel prompt tuning-based framework named A2GP-SF, which integrates attribute generative prompting (AGP) and adaptive sharpness flattening (ASF). The proposed AGP paradigm dynamically generates attribute-aware prompts for each instance, facilitating better semantics learning and enhancing plasticity. Additionally, the ASF mechanism aims to mitigate overfitting by applying adaptive perturbations to flatten sharpness, with these perturbations adjusted based on gradient norm changes, thereby enhancing the model’s robustness and generalization. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority of our proposed A2GP-SF framework.
Desen Wang, Sisi Fu, Congcong Wen, Bingzhi Chen
ICASSP6
2025 Towards Differential Optimization: Rehearsal-Free Class-Incremental Learning with Slow Learners and Fast Adapters
abstract
Class-incremental learning (CIL) enables models to learn new tasks without forgetting previously acquired knowledge. However, existing CIL approaches often struggle with inadequate adaptation to task-specific feature spaces and catastrophic forgetting of previously-acquired knowledge, compromising the models’ plasticity and stability. To address these challenges, this paper proposes a novel differential optimization paradigm called DO-CIL, which incorporates task-agnostic slow learner (TSL) with task-specific fast adapter (TFA) for rehearsal-free CIL. Specifically, TSL aims to effectively capture shared knowledge with low learning rates for robust generalization, while TFA allows pre-trained models to adapt to new task-specific feature spaces. Benefitting from the classifier retraining strategy, a learnable semantic shift network is also proposed to align prototypes with the evolving model representation, facilitating the retraining of task-specific classifiers based on these updated prototypes. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority and effectiveness of our DO-CIL approach compared to state-of-the-art baselines.
Yinghong Chen, Huanjia Zhu, Jieyi Cai, Jun Liang 0002, Bingzhi Chen
ICASSP6
2025 Enhancing Incomplete Multimodal Learning via Modal Complementary Recovering
abstract
Multimodal learning presents significant challenges arising from the unpredictable absence of modalities during both training and testing phases. Existing recovery methods struggle to leverage the available data, which can introduce additional noise during the recovery process and degrade performance. To mitigate these issues, we introduce a novel Modal Complementary Recovering (MCR) paradigm that strategically integrates both Complementary Graph-based Recovery (CGR) and Topological Low-Rank Adaptation (ToRA) mechanisms to enhance the effectiveness and reliability of incomplete multimodal learning. For effective exploitation of the complementarity among different modalities, the main objective of CGR is to employ bidirectional mapping flows trained on a small subset of complete data to learn complementary graphs across all modalities. By constructing entity-relationship diagram specific to dataset, ToRA is designed to enhance the fine-tuning process by incorporating topological prompt. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority of our MCR paradigm in comparison to state-of-the-art baselines.
Meirong Ding, Chuang Zou, Wenxiu Cai, Bingzhi Chen
ICASSP6
2025 Advancing Few-Shot Class-Incremental Learning with Virtual Prototype Guidance Prompting
abstract
Few-Shot Class-Incremental Learning (FSCIL) aims to incrementally learn new class knowledge from limited samples while preserving previously knowledge from encountered classes. However, existing FSCIL methods encounter two primary challenges: (1) inadequate adaptation, where overfitting to new classes compromises the model’s adaptability, and (2) catastrophic forgetting, where previously learned knowledge is not well preserved. In this paper, we propose the Virtual Prototype Guidance Prompting (VPGP) paradigm, integrating the Multi-Grained Prompt (MGP) and Virtual-Prototype Guidance (VPG) strategies. Specifically, MGP enhances adaptation and prevents overfitting by introducing domain-general and fine-grained prompts, expanding the embedding space to capture core feature representations of novel classes. Meanwhile, VPG mitigates catastrophic forgetting by employing a dynamic fusion strategy to retrieve robust old class knowledge and generate virtual prototype, guiding the model to maintain learned knowledge across different sessions. Extensive experiments on multiple benchmark datasets demonstrate the superiority of our proposed VPGP framework.
Huanjia Zhu, Xiaocheng Fang, Jun Liang 0002, Bingzhi Chen
ICASSP5
2025 Revisiting DETR for Small Object Detection via Noise-Resilient Query Optimization
abstract
Despite advancements in Transformer-based detectors for small object detection (SOD), recent studies show that these detectors still face challenges due to inherent noise sensitivity in feature pyramid networks (FPN) and diminished query quality in existing label assignment strategies. In this paper, we propose a novel Noise-Resilient Query Optimization (NRQO) paradigm, which innovatively incorporates the Noise-Tolerance Feature Pyramid Network (NT-FPN) and the Pairwise-Similarity Region Proposal Network (PS-RPN). Specifically, NTFPN mitigates noise during feature fusion in FPN by preserving spatial and semantic information integrity. Unlike existing label assignment strategies, PS-RPN generates a sufficient number of high-quality positive queries by enhancing anchor-ground truth matching through position and shape similarities, without the need for additional hyperparameters. Extensive experiments on multiple benchmarks consistently demonstrate the superiority of NRQO over state-of-the-art baselines.
Xiaocheng Fang, Jieyi Cai, Wenxiu Cai, Yishu Liu 0001, Bingzhi Chen
ICME6
2025 Task-Aware Knowledge Prompt and Distillation for Cross-Domain Few-Shot Learning
abstract
Cross-Domain Few-Shot Learning (CD-FSL) aims to recognize unseen classes from target domains using only limited labeled samples. However, mainstream CD-FSL methods face two key challenges: (1) domain gap, arising from distributional differences between the source and target domains, and (2) overfitting, which occurs due to the small number of labeled samples in target domains, causing the model to overfit to these few samples. To address these challenges, we propose a novel Task-Aware Knowledge Prompt and Distillation (TKPD) method for CD-FSL, which integrates the Vision-Text Domain Prompt (VTDP) and Attribute-Task Knowledge Distillation (ATKD) modules. VTDP mitigates the domain gap by generating domain prompts to acquire domain-relevant knowledge, while ATKD strengthens the model by incorporating both homologous and heterogeneous knowledge, extending the knowledge base beyond the limited labeled samples and mitigating overfitting. Extensive experiments on 13 benchmark datasets validate the effectiveness of the proposed TKPD method.
Jun Liang 0002, Yunyu Zou, Yalong Cheng, Yishu Liu 0001, Bingzhi Chen
ICME7
2025 Enhancing Few-Shot Class-Incremental Learning via Cross-Modal Bias Alignment
abstract
Few-Shot Class-Incremental Learning (FSCIL) aims to learn new classes from limited samples while retaining knowledge of previously learned classes. Prompt tuning on large-scale pre-trained models has achieved certain success in FSCIL. However, current prompt-based approaches still suffer from the challenges of modality bias and catastrophic forgetting. To address these challenges, we propose a prompt tuning-based Cross-Modal Bias Alignment (CMBA) framework, which integrates Prompting Bias Alignment (PBA) and Dynamic Prototype Tuning (DPT). PBA mitigates modality bias by aligning the prompting biases between the visual and textual modalities. Meanwhile, DPT relieves catastrophic forgetting by introducing pseudo-sample features of old classes in the incremental sessions, which are generated by dynamically tuning prototypes that adapt to the evolving feature space. Extensive experiments on multiple FSCIL datasets demonstrate the consistent superiority of our CMBA approach over the existing state-of-the-art baselines.
Desen Wang, Yishu Liu 0001, Bingzhi Chen
ICME5
2025 Cause-Effect Driven Optimization for Robust Medical Visual Question Answering with Language Biases
abstract
Existing Medical Visual Question Answering (Med-VQA) models often suffer from language biases, where spurious correlations between question types and answer categories are inadvertently established. To address these issues, we propose a novel Cause-Effect Driven Optimization framework called CEDO, that incorporates three well-established mechanisms, i.e., Modality-driven Heterogeneous Optimization (MHO), Gradient-guided Modality Synergy (GMS), and Distribution-adapted Loss Rescaling (DLR), for comprehensively mitigating language biases from both causal and effectual perspectives. Specifically, MHO employs adaptive learning rates for specific modalities to achieve heterogeneous optimization, thus enhancing robust reasoning capabilities. Additionally, GMS leverages the Pareto optimization method to foster synergistic interactions between modalities and enforce gradient orthogonality to eliminate bias updates, thereby mitigating language biases from the effect side, i.e., shortcut bias. Furthermore, DLR is designed to assign adaptive weights to individual losses to ensure balanced learning across all answer categories, effectively alleviating language biases from the cause side, i.e., imbalance biases within datasets. Extensive experiments on multiple traditional and bias-sensitive benchmarks consistently demonstrate the robustness of CEDO over state-of-the-art competitors.
Huanjia Zhu, Yishu Liu 0001, Xiaozhao Fang, Guangming Lu 0002, Bingzhi Chen
IJCAI5
2025 PerioDet: Large-Scale Panoramic Radiograph Benchmark for Clinical-Oriented Apical Periodontitis Detection
Xiaocheng Fang, Jieyi Cai, Chengju Zhou, Minhua Lu, Bingzhi Chen
MICCAI (16)6
2025 Med-BiasX: Robust Medical Visual Question Answering with Language Biases
Huanjia Zhu, Yishu Liu 0001, Chengju Zhou, Guangming Lu 0002, Bingzhi Chen
MICCAI (14)5
2025 Label Prediction Inherited Hashing for Cross-Modal Retrieval: Applying Supervised Hashing to Unsupervised Tasks
abstract
Supervised cross-modal hashing has achieved remarkable progress in retrieving related items across different modalities. However, in practical applications, a significant portion of data remains unlabeled, such as online data on websites, which must be included for effective retrieval. To address this challenge, while maintaining the high accuracy and efficiency of supervised methods, few works have attempted to adapt existing supervised techniques to handle unsupervised tasks through a general modular approach. To this end, we introduce a novel cross-modal hashing method, termed Label Prediction Inherited Hashing (LPIH). Initially, LPIH leverages labeled data to learn high-quality general label functions using supervised methods. Subsequently, it inherits the existing hash codes from existing supervised methods to further refine the pseudo-label information. Finally, LPIH integrates the refined pseudo-label information with the existing hash functions to learn new hash functions specifically tailored for unsupervised tasks. Extensive experimental results on three public datasets demonstrate the superior performance of LPIH compared to state-of-the-art (SOTA) cross-modal hashing methods. Specifically, LPIH achieves an average precision improvement of 5% over SOTA methods, highlighting its effectiveness in bridging the gap between supervised and unsupervised learning in the context of cross-modal retrieval.
Kaihang Jiang, Wai Keung Wong, Jianyang Qin, Xiaozhao Fang, Jie Wen 0001, Bingzhi Chen, Hongbo Gao 0001
ACM Multimedia6
2025 CauRDG: Enhancing Domain Generalization with Causal-Driven Semantic Consistency Reasoning
abstract
Domain generalization (DG) plays a pivotal role in enabling models to maintain robust performance across heterogeneous environments. However, existing DG methods are fundamentally constrained by two intertwined limitations: (1) causal misalignment, which stems from undifferentiated feature encoding that entangles causal mechanisms with environmental biases; (2)semantic conflict arises when conventional adaptation methods find it challenging to balance the preservation of class discriminability with the mitigation of domain-specific distribution discrepancies. To address these challenges of DG, we propose a novel Causal-Driven Semantic Consistency Reasoning (CauRDG) method, which synergistically integrates Prototype-Guided Causal Disentanglement (PGCD) and Dual-Space Semantic Disambiguation (DSSD). Specifically, PGCD constructs a causal framework that identifies stable relationships and decouples invariant mechanisms from domain-specific variations, preserving causal consistency while adapting to contextual differences. DSSD harnesses a dual-space paradigm, enhancing local categorical clarity and maintaining global conceptual unity, thus balancing domain-specific precision with cross-domain coherence. The robustness provided by CauRDG ensures robust extraction and interpretation of essential features by preserving invariant causal structures, thereby harmonizing discriminative semantics with domain-varying contexts. Extensive experiments on multiple benchmark datasets consistently demonstrate the effectiveness and superiority of our CauRDG over state-of-the-art baselines.
Zongxin Liu 0003, Yishu Liu 0001, Guangming Lu 0002, Xiaoling Luo 0001, Bingzhi Chen
ACM Multimedia5
2025 PET-GPRA: Rethinking PET with Gradient-Aware Prompting and Router-Free Adapters for Few-shot Class-Incremental Learning
abstract
Few-Shot Class-Incremental Learning (FSCIL) aims to continuously learn novel concepts from limited training samples without forgetting previously encountered classes. Recent advancements have leveraged Parameter-Efficient Tuning (PET) strategies on pre-trained models to enhance FSCIL performance. However, current PET-based FSCIL approaches still suffer from the challenges posed by catastrophic collapse of general prompt and limited adaptability of specific prompt . To this end, we redefine the function of the PET paradigm with both gradient-aware prompting (GAP) and router-free adapters (RFA) to boost the performance of FSCIL, termed as "PET-GPRA". To dynamically balance the retention of previously learned general knowledge and the acquisition of novel class information across sessions, the GAP paradigm adaptively adjusts the updated gradient of the general prompt by leveraging the angular relationship between the general knowledge gradient and the novel knowledge gradient. Meanwhile, the RFA mechanism utilizes the semantic similarity between class attributes to replace the routing network, guiding the integration of adapter information, in which adapters serve as specific prompts to enhance the adaptability. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority and effectiveness of our proposed PET-GPRA framework over state-of-the-art baselines.
Yishu Liu 0001, Desen Wang, Xiaoling Luo 0001, Bingzhi Chen, Guangming Lu 0002
ACM Multimedia5
2025 SG-FSL: Cross-Domain Few-Shot Learning with Style-Decoupled Augmentation and Gradient-Conflict Adjustment
abstract
Cross-Domain Few-Shot Learning (CD-FSL) aims to transfer knowledge acquired from a source domain with abundant data to the target domain with limited labeled samples. Recent advancements have enhanced model generalization through Perturbation Augmentation (PA), facilitating more effective knowledge transfer. However, PA-based CD-FSL methods still suffer from two critical challenges, i.e., (1) limited diversity of augmented samples, making it difficult to cover the true distribution of unseen domains, and (2) conflicting gradients during model optimization, where augmented and original samples drive the model's optimization in opposing directions. To address these issues, we propose a novel PA-based framework with Style-Decoupled Augmentation (SDA) and Gradient-Conflict Adjustment (GCA) for Cross-Domain Few-Shot Learning, which is termed ''SG-FSL''. Specifically, SDA decouples the source domain style into style weights and basis styles, generating diverse unseen styles by perturbing the style weights to reweight the basis styles. Meanwhile, GCA leverages the angular relationships between the domain-specific gradient directions of augmented and original features, adaptively adjusting the gradient directions of original features to ensure that the model acquires diverse domain knowledge without interference, guiding it toward conflict-free optimization. Comprehensive experiments on multiple benchmark datasets consistently demonstrate the effectiveness and superiority of our method over state-of-the-art baselines.
Yunyu Zou, Yishu Liu 0001, Jun Liang 0002, Bingzhi Chen
ACM Multimedia4
2025 Language-Bias-Resilient Visual Question Answering via Adaptive Multi-Margin Collaborative Debiasing
abstract
Language bias in Visual Question Answering (VQA) arises when models exploit spurious statistical correlations between question templates and answers, particularly in out-of-distribution scenarios, thereby neglecting essential visual cues and compromising genuine multimodal reasoning. Despite numerous efforts to enhance the robustness of VQA models, a principled understanding of how such bias originates and influences model behavior remains underdeveloped. In this paper, we address this gap through a comprehensive empirical and theoretical analysis, revealing that modality-specific gradient imbalances, which originate from the inherent heterogeneity of multimodal data, lead to skewed feature fusion and biased classifier weights. To alleviate these issues, we propose a novel Multi-Margin Collaborative Debiasing (MMCD) framework that adaptively integrates frequency-, confidence-, and difficulty-aware angular margins with a dynamic difficulty-aware contrastive learning mechanism, to dynamically reshape decision boundaries. Extensive experiments across multiple challenging VQA benchmarks confirm the consistent superiority of our proposed MMCD over state-of-the-art baselines in combating language bias.
Huanjia Zhu, Shuyuan Zheng, Yishu Liu 0001, Sudong Cai, Bingzhi Chen
NeurIPS5
2025 LBF-VQA: Towards Language Bias-Free Visual Question Answering With Multi-Space Collaborative Debiasing
Yishu Liu 0001, Huanjia Zhu, Bingzhi Chen, Xiaozhao Fang, Guangming Lu 0002, Shengli Xie 0001
IEEE Trans. Knowl. Data Eng.3
2025 Toward Robust Semi-Supervised Distribution Alignment Against Label Distribution Shift With Noisy Annotations
abstract
Deep learning-based AI models typically require a large amount of high-quality annotated data to achieve optimal performance. However, thelabel distribution shiftcaused by noisy annotations can lead to perturbations in the classification boundary, reducing the robustness and generalization capabilities of deep learning models. To mitigate this issue, we transform the problem of learning from noisy labels into a semi-supervised learning problem, and propose a novel Semi-Supervised Distribution Alignment (SSDA) framework that strategically integrates noise-robust distribution alignment within a unified semi-supervised learning paradigm for combating noisy labels. By leveraging the similarity distribution between historical predictions, the proposed SSDA approach benefits from a flexible multi-historical regression modeling strategy, which aims to identify high-confidence samples/pairs and recalibrate the label shift through pseudo-labels. Furthermore, our approach employs a comprehensive multi-granularity distribution adaptation strategy, incorporating both instance-wise and class-aware distribution alignment to quantitatively minimize semantic discrepancies across different mixed feature domains. In this way, our SSDA approach ultimately achieves more resilient and generalizable performance against label noise, even in the presence of substantial noise. Extensive experiments conducted on multiple simulated and real-world noisy benchmark datasets consistently demonstrate the superiority and effectiveness of our SSDA method compared to existing state-of-the-art baselines.
Bingzhi Chen, Zhanhao Ye, Yishu Liu 0001, Xiaozhao Fang, Guangming Lu 0002, Shengli Xie 0001, Xuelong Li 0001
IEEE Trans. Multim.1
2024 CariesXrays: Enhancing Caries Detection in Hospital-Scale Panoramic Dental X-rays via Feature Pyramid Contrastive Learning
abstract
Dental caries has been widely recognized as one of the most prevalent chronic diseases in the field of public health. Despite advancements in automated diagnosis across various medical domains, it remains a substantial challenge for dental caries detection due to its inherent variability and intricacies. To bridge this gap, we release a hospital-scale panoramic dental X-ray benchmark, namely “CariesXrays”, to facilitate the advancements in high-precision computer-aided diagnosis for dental caries. It comprises 6,000 panoramic dental X-ray images, with a total of 13,783 instances of dental caries, all meticulously annotated by dental professionals. In this paper, we propose a novel Feature Pyramid Contrastive Learning (FPCL) framework, that jointly incorporates feature pyramid learning and contrastive learning within a unified diagnostic paradigm for automated dental caries detection. Specifically, a robust dual-directional feature pyramid network (D2D-FPN) is designed to adaptively capture rich and informative contextual information from multi-level feature maps, thus enhancing the generalization ability of caries detection across different scales. Furthermore, our model is augmented with an effective proposals-prototype contrastive regularization learning (P2P-CRL) mechanism, which can flexibly bridge the semantic gaps among diverse dental caries with varying appearances, resulting in high-quality dental caries proposals. Extensive experiments on our newly-established CariesXrays benchmark demonstrate the potential of FPCL to make a significant social impact on caries diagnosis.
Bingzhi Chen, Sisi Fu, Yishu Liu 0001, Jiahui Pan 0003, Guangming Lu 0002, Zheng Zhang 0006
AAAI1
2024 Enhancing DETRs for Small Object Detection via Multi-Scale Refinement and Query-Aided Mining
Sisi Fu, Xiaocheng Fang, Jieyi Cai, Huosheng Wen, Bingzhi Chen
ACML7
2024 Progressive Stepwise Diffusion Model with Dual Decoders for Semi-Supervised Medical Image Segmentation
abstract
Semi-supervised medical image segmentation tasks aim to harness the potential of vast amounts of unlabeled data using a limited amount of annotated data. Denoising Diffusion Probabilistic Models, which have achieved significant success in image generation, are gradually being explored for their potential in semantic image segmentation. However, their application in semi-supervised medical image segmentation is still in its early stages. Initially, due to the high randomness of diffusion models, the pseudo-labels generated during the early training phase may mislead the processing of unlabeled data. Additionally, the use of fixed-time steps for random sampling during training limits the ability of the model to learn effective denoising functions at an early stage. To address these issues, we propose an innovative framework named Progressive Stepwise Diffusion Network with Dual Decoders (PSDD) for semi-supervised medical image segmentation. This framework incorporates an additional normal decoder into the denoising diffusion encoder-decoder structure to provide more accurate labels and employs a Progressive Incremental Step strategy to gradually train the model for longer generation processes. Evaluated on two 2D colon polyp segmentation datasets and a 3D Left Atrium dataset, the experimental results demonstrate significant performance improvements over current advanced methods, thereby validating the effectiveness and potential of this framework in handling complex semi-supervised learning scenarios.
Xiaolin Huang, Jingchun Lin, Bingzhi Chen, Guangming Lu 0002
BIBM5
2024 Decoupled Self-Adaptive Distribution Regularization for Few-Shot Image Classification
abstract
The feature dispersion, arising from the inherent constraints of data scarcity, has emerged as a prominent challenge in the domain of few-shot learning. In this paper, we propose a novel Self-adaptive Distribution Regularization (SADR) approach, which can adaptively bridge the semantic gaps across distribution patterns and boundaries for learning from limited labeled data. Specifically, the technical core of our SADR approach is to decouple the feature embeddings into two discrete spaces: the intra-class and inter-class distributions, leading to robust and discriminative feature representations in a self-adaptive manner. To achieve meticulous similarity measurements while mitigating redundant feature information, an innovative regularized Brownian Distance Covariance (R-BDC) metric is strategically designed to simultaneously explore both the joint and marginal distributions present among diverse input samples. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority of our SADR approach over state-of-the-art baselines.
Bingzhi Chen, Haoming Zhou, Yishu Liu 0001, Guangming Lu 0002, Zheng Zhang 0006
ICASSP1
2024 Medical Vision-Language Representation Learning with Cross-Modal Multi-Teacher Contrastive Distillation
abstract
Medical vision-language representation learning has garnered considerable attention owing to its applicability to extracting generic representations from the image and text modality. However, it still remains challenging to acquire a more comprehensive understanding of intra- and inter-modal semantic knowledge. In this paper, we propose a Cross-Modal Multi-Teacher Contrastive Distillation (CMCD) architecture, which aims to comprehensively learn medical vision-language representation in a unified multi-teacher framework. Specifically, a cross-modal knowledge distillation (CKD) module is designed to refine reconstructed semantics under an additional supervision signal generated by momentum teachers from the other modality, achieving more robust semantic interaction across modalities. To better alleviate the heterogeneity and semantic gaps, the multi-level contrastive learning (MCL) module is conceived to align features of both intra- and inter-modal via contrastive learning from multi-level perspectives. Extensive experiments on two medical downstream tasks, i.e., Med-VQA and Med-ITC, demonstrate that our CMCD consistently outperforms the state-of-the-art methods.
Bingzhi Chen, Yishu Liu 0001, Jiahui Pan 0003, Meirong Ding
ICASSP1
2024 Rethinking Adversarial Robustness Distillation VIA Strength-Dependent Adaptive Regularization
abstract
Despite the progress achieved by existing adversarial distillation (AD) approaches, most mainstream models suffer from inadequate adversarial robustness, due to the challenges of fixed attack strength and unreliable teacher guidance. In this paper, we propose a novel Strength-Dependent Adaptive Regularization (SDAR) paradigm to reinforce the function of adversarial distillation with strength-adaptive adversarial attack (SAA) and multi-dimensional knowledge distillation (MKD). Different from the traditional adversarial training (AT) methods, the proposed SAA scheme dynamically assigns an adaptive and efficient attack strength for each instance, which aims to facilitate smoother classification boundaries. By incorporating dynamic strength coefficients, a comprehensive MKD strategy is designed to fully explore the valuable context information and narrow distribution discrepancies across teacher-student domains. Particularly, our SDAR paradigm can seamlessly integrate with the current AD frameworks, further enhancing the adversarial robustness of deep learning models. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority of SDAR over state-of-the-art baselines.
Bingzhi Chen, Shuobin Lin, Yishu Liu 0001, Zheng Zhang 0006, Guangming Lu 0002, Lewei He
ICME1
2024 Enhancing Few-Shot Classification without Forgetting Through Multi-level Contrastive Constraints
abstract
Most recent few-shot learning approaches are based on meta-learning with episodic training. However, prior studies encounter two crucial problems: (1) the presence of inductive bias, and (2) the occurrence of catastrophic forgetting. In this paper, we propose a novel Multi-Level Contrastive Constraints (MLCC) framework, that jointly integrates within-episode learning and across-episode learning into a unified interactive learning paradigm to solve these issues. Specifically, we employ a space-aware interaction modeling scheme to explore the correct inductive paradigms for each class between within-episode similarity/dis-similarity distributions. Additionally, with the aim of better utilizing former prior knowledge, a cross-stage distribution adaption strategy is designed to align the across-episode distributions from different time stages, thus reducing the semantic gap between existing and past prediction distribution. Extensive experiments on multiple few-shot datasets demonstrate the consistent superiority of MLCC approach over the existing state-of-the-art baselines.
Bingzhi Chen, Haoming Zhou, Yishu Liu 0001, Jiahui Pan 0003, Guangming Lu 0002
ICME1
2024 Ambiguity Consistency and Uncertainty Minimization for Semi-Supervised Medical Image Segmentation
abstract
Co-training and pseudo-supervision are two common strategies in semi-supervised medical image segmentation. However, co-training may lead to a ’resonance’ problem, and the effectiveness of generating pseudo-labels by setting thresholds may greatly depend on manual efforts. To address these issues, we propose an innovative framework for Ambiguity Consistency and Uncertainty Minimization (ACUM) in semi-supervised medical image segmentation. Specifically, ACUM comprises two main components: (1) Ambiguity Consistency Constraint (ACC), which encourages model differentiation and applies dynamic pixel-level consistency constraints through ambiguous areas between sub-networks; (2) Pixel Uncertainty Minimization (PUM), which generates high-confidence pseudo-labels by selecting labels with relatively low uncertainty based on the uncertainty maps of sub-networks. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority of our proposed ACUM approach over state-of-the-art techniques.
Xiaolin Huang, Yujiang Yao, Bingzhi Chen
ICME6
2024 Deep Unfolding 3D Non-Local Transformer Network for Hyperspectral Snapshot Compressive Imaging
abstract
Hyperspectral compressive imaging has shown remarkable advancements through the adoption of deep unfolding frameworks, which integrate the proximal mapping prior into the data fidelity term to formulate the reconstruction problem. However, existing technologies still face challenges in effectively capturing spatial-spectral features during the iterative deep prior learning stage, leading to unsatisfactory performance degradation. To address this issue, we propose a deep unfolding 3D non-local transformer (3DNLT) network for hyperspectral compressive imaging. A learnable half-quadratic splitting (HQS) algorithm is utilized to iteratively update the linear projection. Furthermore, a 3D non-local attention ushaped transformer is presented as the deep proximal mapping prior module to obtain the spatial-spectral long-range dependency features, leading to enhance the network’s ability to capture fine-grained hyperspectral and spatial details. Experimental results on both synthetic and real hyperspectral image reconstruction have demonstrated the superior performance of the 3DNLT network compared to state-of-the-art methods.
Yongyong Chen, Bingzhi Chen, Yicong Zhou
ICME4
2024 Robust Visual Question Answering With Contrastive-Adversarial Consistency Constraints
abstract
Visual cues and question semantics contribute to final answer predictions from distinct perspectives. However, inherent language bias confounds the relationship between visual and question cues, leading to a misguided preference for question semantics. Different from the existing studies that focus on inter-class discrimination, this paper proposes a robust visual question answering framework with contrastive-adversarial consistency constraints (CACC) at both inter- and intra-instance levels. From a fine-grained instance-level perspective, our approach initially introduces an effective inter-instance contrastive constraint to perform adaptive bias rectification. To enhance intra-instance invariance and reduce information redundancy, we refine the concept of semantic structure relationships by constructing intra-instance adversarial constraints using the Hilbert-Schmidt Independence Criterion (HSIC) independence criterion. Benefitting from both inter- and intra-instance perspectives, our method can effectively alleviate these language biases, enhancing the overall robustness of the representation. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority of our CACC over state-of-the-art baselines.
Meirong Ding, Yishu Liu 0001, Guangming Lu 0002, Bingzhi Chen
ICME6
2024 Enhancing Cross-Modal Retrieval via Visual-Textual Prompt Hashing
Bingzhi Chen, Zhongqi Wu, Yishu Liu 0001, Guangming Lu 0002, Zheng Zhang 0006
IJCAI1
2024 Medical Cross-Modal Prompt Hashing with Robust Noisy Correspondence Learning
Yishu Liu 0001, Zhongqi Wu, Bingzhi Chen, Zheng Zhang 0006, Guangming Lu 0002
MICCAI (3)3
2024 Stay Focused is All You Need for Adversarial Robustness
Bingzhi Chen, Ruihan Liu, Yishu Liu 0001, Xiaozhao Fang, Jiahui Pan 0003, Guangming Lu 0002, Zheng Zhang 0006
ACM Multimedia1
2024 Partial Multi-label Learning Based On Near-Far Neighborhood Label Enhancement And Nonlinear Guidance
Na Han, Xiaozhao Fang, Bingzhi Chen, Jie Wen 0001
ACM Multimedia5
2024 Combating Visual Question Answering Hallucinations via Robust Multi-Space Co-Debias Learning
Yishu Liu 0001, Huanjia Zhu, Yuncheng Jiang 0004, Zheng Zhang 0006, Bingzhi Chen
ACM Multimedia7
2024 A feature-enhanced hybrid attention network for traffic sign recognition in real scenes
abstract
Abstract Currently, traffic sign recognition techniques have been brought into the assistive driving of automobiles. However, small traffic sign recognition in real scenes is still a challenging task due to the class imbalance issue and the size limit of the traffic signs. To address the above issues, a feature‐enhanced hybrid attention network is proposed based on YOLOv5s for a small, fast, and accurate traffic sign detector. First, a series of online data augmentation strategies are designed in the preprocessing module for the model training. Second, the hybrid channel and spatial attention module CSAM are integrated into the backbone for a better feature extraction ability. Third, the channel attention module CAM is used in the detection head for a more efficient feature fusion ability. To validate the approach, extensive experiments are conducted based on the Tsinghua‐Tencent 100K dataset. It is found that the novel method achieves state‐of‐the‐art performance with only negligible increases in the model parameter and computational overhead. Specifically, the , parameters, and FLOPs are 85.8%, 7.13 M, and 16.1 G, respectively.
Lewei He, Fucai Lan, Chuanzhe Zhou, Yaoguang Ye, Wencong Zhang, Bingzhi Chen, Jiahui Pan 0003
IET Image Process.6
2024 Context-aware graph embedding with gate and attention for session-based recommendation
Junlong Chi, Peilin Hong, Guangming Lu 0002, David Zhang 0001, Bingzhi Chen
Neurocomputing6
2024 Deep Fuzzy Multiteacher Distillation Network for Medical Visual Question Answering
abstract
Medical visual question answering (medical VQA) is a critical cross-modal interaction task that garnered considerable attention in the medical domain. Several existing methods commonly leverage the vision-and-language pretraining paradigms to mitigate the limitation of small-scale data. Nevertheless, most of them still suffer from two challenges that remain for further research: 1) limited research focuses on distilling representation from a complete modality to guide the representation learning of masked data in other modalities. 2) Multimodal fusion based on self-attention mechanisms cannot effectively handle the inherent uncertainty and vagueness of information interaction across modalities. To mitigate these issues, in this article, we propose a novel deep fuzzy multiteacher distillation (DFMD) network for medical VQA, which can take advantage of fuzzy logic to model the uncertainties from vison-language representations across modalities in a multiteacher framework. Specifically, a multiteacher knowledge distillation module is conceived to assist in reconstructing the missing semantics under the supervision signal generated by teachers from the other complete modality, achieving more robust semantic interaction across modalities. Incorporating insights from the fuzzy logic theory, we propose a noise-robust encoder called FuzBERT that enables our DFMD model to reduce the imprecision and ambiguity in feature representation during the multimodal interaction process. To the best of our knowledge, our work isthe first attemptto combine the fuzzy logic theory with the transformer-based encoder to effectively learn multimodal representation for medical VQA. Experimental results on the VQA-RAD and SLAKE datasets consistently demonstrate the superiority of our proposed DFMD method over state-of-the-art baselines.
Yishu Liu 0001, Bingzhi Chen, Shuihua Wang, Guangming Lu 0002, Zheng Zhang 0006
IEEE Trans. Fuzzy Syst.2
2024 Vague-Segment Technique: Automatic Computation of Tumor Stroma Ratio for Breast Cancer on Whole Slides
abstract
The calculation of Tumor Stroma Ratio (TSR) is a challenging medical issue that could improve predictions of neoadjuvant chemotherapy benefits and patient prognoses. Although several studies on breast cancer and deep learning methods have achieved promising results, the drawbacks that pixel-level semantic segmentation processes could not extract core tumor regions containing both tumor pixels and stroma pixels make it difficult to accurately calculate TSR. In this paper, we propose a Vague-Segment Technique (VST) consisting of a designed SwinV2UNet module and a modified Suzuki algorithm. Specifically, the SwinV2UNet identifies tumor pixels and generate pixel-level classification results, based on which the modified Suzuki algorithm extracts the contour of core tumor regions in terms of cosine angle. Through this way, VST obtains vaguely segmentation results of core tumor regions containing both tumor pixels and stroma pixels, where the TSR could be calculated by the formula of Intersection over Union (IOU). For the training and evaluation, we utilize the well-known The Cancer Genome Atlas (TCGA) database to create an annotated dataset, while 150 images with TSR annotations from real cases are also collected. The experimental results illustrate that the proposed VST could generate better tumor identification results compared with state-of-the-art methods, where the extracted core tumor regions lead to more consistencies of calculated TSR with senior experts compared to junior pathologists. The experimental results demonstrate the superiority of our proposed pipeline, which has promise for future clinical application.
Xinsen Lian, Kunping Yang, Bingzhi Chen, Xiuhong Cai, Xinling Lu, Jinlin Chen, Ming Tian, Pengtao Lin
IEEE J. Biomed. Health Informatics3
2024 Deep Learning-Based Segment Trend Removal Approach for Prognostics and Health Management Signals of Rail Vehicles
abstract
To ensure the safety and reliability of rail vehicles, prognostics and health management (PHM) is crucial. However, PHM signals can often be disrupted by trend components, which fundamentally undermine the accuracy of damage identification and prediction. Removing trend components efficiently is a challenge due to the intermittent and nonlinear character of railway vehicle PHM signals. This study proposes a segmented trend removal approach (STRA) for PHM signals based on deep learning. To evaluate the effectiveness of trend removal, a drift ratio and amplitude ratio assessment metric is defined. An identification model based on convolutional neural networks is also proposed to detect start-stop features and determine overall trend points for signal segmentation. A segmented spline baseline method is presented to eliminate running segment trends, while a least squares approach (LSA) is employed to address stationary segment trends. To increase processing and training efficiency with various sampling frequencies, a frequency conversion technique is implemented. The STRA outperforms conventional methods such as LSA, spline baseline approach, and empirical mode decomposition in nonlinear trend reduction. Additionally, this method can be extended to eliminate the zero drift of intermittent acceleration signals in PHM and is similarly suitable for use in other PHM applications, including automobiles, vessels, and bridges.
Shaoze Zhou, Pengfei Zhao 0005, Bingzhi Chen
IEEE Trans. Reliab.4
2023 Combating Medical Label Noise via Robust Semi-supervised Contrastive Learning
Bingzhi Chen, Zhanhao Ye, Yishu Liu 0001, Zheng Zhang 0006, Jiahui Pan 0003, Guangming Lu 0002
MICCAI (1)1
2023 Deep Margin-Sensitive Representation Learning for Cross-Domain Facial Expression Recognition
abstract
Cross-domain Facial Expression Recognition (FER) aims to safely transfer the learned knowledge from labeled source data to unlabeled target data, which is challenging due to the subtle difference between various expressions and the large discrepancy between domains. Existing methods mainly focus on reducing the domain shift for transferable features but fail to learn discriminative representations for recognizing facial expression, which may result in negative transfer under cross-domain settings. To this end, we propose a novel Deep Margin-Sensitive Representation Learning (DMSRL) framework, which can extract multi-level discriminative features during sematic-aware domain adaptation. Specifically, we design a semantic metric learning module based on the category prior of source data and generated pseudo labels of target data, which can facilitate discriminative intra-domain representation learning and transferable inter-domain knowledge discovery by enlarging the category margin. Moreover, we develop a mutual information minimization module by simultaneously distilling the domain-invariant components and eliminating the domain-sensitive ones, which benefits discriminative transferable feature learning by generating accurate pseudo target labels. Furthermore, instead of only utilizing the global features, we formulate a multi-level feature extracting module to concurrently get the local ones, which contain detailed information to distinguish the small changes among different expressions. These modules are jointly utilized in our DMSRL in an end-to-end manner to ensure the positive transfer of source knowledge. Extensive experimental results on seven databases demonstrate that our DMSRL can achieve superior performance against state-of-the-art baselines.
Yingjian Li 0001, Zheng Zhang 0006, Bingzhi Chen, Guangming Lu 0002, David Zhang 0001
IEEE Trans. Multim.3
2022 Multi-Label Chest X-Ray Image Classification via Semantic Similarity Graph Embedding
abstract
Automated multi-label chest X-ray (CXR) image classification has recently made significant progress in clinical diagnosis based on the advanced deep learning techniques. However, most existing methods mainly focus on analyzing locality visual cues from a single image but fail to leverage the underlying explicit correlations among different images for precise disease diagnosis. By contrast, an experienced radiologist expertizes in transferring knowledge from previous tasks to diagnose the present radiograph. To enable the machine like a radiologist, this paper proposes a novel Semantic Similarity Graph Embedding (SSGE) framework, which explicitly explores the semantic similarities among images to optimize the visual feature embedding for improving the performance of multi-label CXR images classification. Specifically, the proposed SSGE framework contains three main components: the image feature embedding (IFE) module, similarity graph construction (SGC) module, and semantic similarity learning (SSL) module. To realize interactive teaching and learning between visual and semantic information, the proposed SSGE framework is built on the “Teacher-Student” (semantic-visual) learning mechanism. With the guidance and supervision of the cross-image similarity graph generated by the SGC module, the SSL module leverages Graph Convolutional Network (GCN) to adaptively recalibrate the multi-image feature representations extracted from the IFE module, which guarantees their semantic consistency. Furthermore, we propose a novel re-weighting strategy to learn a more optimal semantic-similarity graph for the information propagation of the GCN layers. Extensive experiments on two benchmark datasets demonstrate the effectiveness of the proposed method in comparison with some state-of-the-art baselines.
Bingzhi Chen, Zheng Zhang 0006, Yingjian Li 0001, Guangming Lu 0002, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 Self-Supervised Exclusive-Inclusive Interactive Learning for Multi-Label Facial Expression Recognition in the Wild
abstract
Facial Expression Recognition (FER) is a long-standing but challenging research problem in computer vision. Existing approaches mainly focus on single-label emotional prediction, which cannot handle the complex multi-label FER task because of the coupling behavior of multiple emotions on a single facial image. To this end, in this paper, we propose a novel Self-supervised Exclusive-Inclusive Interactive Learning (SEIIL) method to facilitate discriminative multi-label FER in the wild, which can effectively handle the coupled multiple sentiments with limited unconstrained training data. Specifically, we construct an emotion disentangling module to capture the inclusive and exclusive characteristics of facial expressions, which can decouple the compound numerous emotions on an image. Moreover, an adaptively-weighted ensemble technique is conceived to aggregate category-level latent exclusive embeddings, and then a conditional adversarial interactive learning module is designed to fully leverage the complementary between the inclusive and formulated latent representations. Furthermore, to tackle the insufficient data for training, we introduce a self-supervised learning strategy to augment the amount and diversity of facial images, which can endow the model with advanced generalization ability. Under this strategy, the proposed two modules can be concurrently utilized in our SEIIL to jointly handle the coupled emotions and alleviate the overfitting problem. Extensive experimental results on six databases illustrate the superb performance of our method against state-of-the-art baselines.
Yingjian Li 0001, Yingnan Gao, Bingzhi Chen, Zheng Zhang 0006, Guangming Lu 0002, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 Learning Informative and Discriminative Features for Facial Expression Recognition in the Wild
abstract
The informativeness and discriminativeness of features collaboratively ensure high-accuracy Facial Expression Recognition (FER) in the wild. Most of existing methods use the single-path deep convolutional neural network with softmax loss for basic FER, while they cannot deal with the challenging situations of the compound FER in the wild, because they fail to learn informative and discriminative features in a targeted manner. To this end, we present an Informative and Discriminative Feature Learning (IDFL) framework that consists of two key components: the Multi-Path Attention Convolutional Neural Network (MPACNN) and Balanced Separate loss (BS loss), for both basic and compound high-accuracy FER in the wild. Specifically, MPACNN leverages different paths to learn diverse features. These features are then adaptively fused into informative ones via an attention module, such that the model can adequately capture detailed information for both basic and compound FER. The BS loss maximizes the inter-class distance of features and minimizes the intra-class one. In this way, the features are discriminative enough for high-accuracy FER in the wild. Particularly, the BS loss is invoked as the objective function of MPACNN, so the model can learn informative and discriminative features at the same time, yielding better performance. Seven databases are utilized to evaluate the proposed method, and the results demonstrate that our method achieves state-of-the-art performance on both basic and compound expressions with good generalization ability. Moreover, our model contains fewer parameters and can be trained faster than other related models.
Yingjian Li 0001, Yao Lu 0008, Bingzhi Chen, Zheng Zhang 0006, Jinxing Li 0003, Guangming Lu 0002, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 Semantic-Interactive Graph Convolutional Network for Multilabel Image Recognition
abstract
Multilabel image recognition, a critically practical task in computer vision, aims to predict multiple objects present in each image. The existing studies mainly focus on conceptual visual cues but fail to reconcile the visual information with their semantic guidance. Intuitively, humans can not only associate extra topological concepts but also imagine other approximate scenes based on a semantic description. Inspired by such semantic-interactive capability, two different types of semantic priors, i.e., the concept correlations of the same scene and semantic similarities among different scenes, should be further explored for the recognition decisions. To efficiently interact with these semantic relationships, in this article, we propose a novel semantic-interactive graph convolutional network (SI-GCN), which can leverage the topological information learned from knowledge graphs to boost the performance of multilabel recognition. Specifically, the proposed SI-GCN framework consists of two different GCN-based branches in parallel, i.e., concept correlations learning (CCL) branch and semantic similarity learning (SSL) branch. Inputting the semantic-embedding vectors of all the concepts, the CCL branch maps the label co-occurrence graph into a set of interdependent concept classifiers. Recalibrating the image feature embedding with the standardized supervision of the semantic similarity graph, the SSL branch learns the semantically consistent in-batch visual representations. Finally, a well-established interactive learning scheme is formulated to concurrently optimize the obtained concept classifiers and the visual representation learning in an end-to-end manner. Extensive experiments on the MS-COCO and Pascal VOC 2007 & 2012 benchmarks demonstrate the superiorities of the proposed SI-GCN method compared to the state-of-the-art baselines.
Bingzhi Chen, Zheng Zhang 0006, Yao Lu 0008, Fanglin Chen 0001, Guangming Lu 0002, David Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2021 Hierarchical Network Based on the Fusion of Static and Dynamic Features for Speech Emotion Recognition
abstract
Many studies on automatic speech emotion recognition (SER) have been devoted to extracting meaningful emotional features for generating emotion-relevant representations. However, they generally ignore the complementary learning of static and dynamic features, leading to limited performances. In this paper, we propose a novel hierarchical network called HNSD that can efficiently integrate the static and dynamic features for SER. Specifically, the proposed HNSD framework consists of three different modules. To capture the discriminative features, an effective encoding module is firstly designed to simultaneously encode both static and dynamic features. By taking the obtained features as inputs, the Gated Multi-features Unit (GMU) is conducted to explicitly determine the emotional intermediate representations for frame-level features fusion, instead of directly fusing these acoustic features. In this way, the learned static and dynamic features can jointly and comprehensively generate the unified feature representations. Benefiting from a well-designed attention mechanism, the last classification module is applied to predict the emotional states at the utterance level. Extensive experiments on the IEMOCAP benchmark dataset demonstrate the superiority of our method in comparison with state-of-the-art baselines.
Mi-Xiao Hou, Bingzhi Chen, Zheng Zhang 0006, Guangming Lu 0002
ICASSP3
2021 JDMAN: Joint Discriminative and Mutual Adaptation Networks for Cross-Domain Facial Expression Recognition
abstract
Cross-domain Facial Expression Recognition (FER) is challenging due to the difficulty of concurrently handling the domain shift and semantic gap during domain adaptation. Existing methods mainly focus on reducing the domain discrepancy for transferable features but fail to decrease the semantic one, which may result in negative transfer. To this end, we propose Joint Discriminative and Mutual Adaptation Networks (JDMAN), which collaboratively bridge the domain shift and semantic gap by domain- and category-level co-adaptation based on mutual information and discriminative metric learning techniques. Specifically, we design a mutual information minimization module for domain-level adaptation, which narrows the domain shift by simultaneously distilling the domain-invariant components and eliminating the untransferable ones lying in different domains. Moreover, we propose a semantic metric learning module for category-level adaptation, which can close the semantic discrepancy during discriminative intra-domain representation learning and transferable inter-domain knowledge discovery. These two modules are jointly leveraged in our JDMAN to safely transfer the source knowledge to target data in an end-to-end manner. Extensive experimental results on six databases show that our method achieves state-of-the-art performance. The code of our JDMAN is available at https://github.com/YingjianLi/JDMAN.
Yingjian Li 0001, Yingnan Gao, Bingzhi Chen, Zheng Zhang 0006, Lei Zhu 0002, Guangming Lu 0002
ACM Multimedia3
2021 An Embarrassingly Simple Approach to Discrete Supervised Hashing
abstract
Prior hashing works typically learn a projection function from high-dimensional visual feature space to low-dimensional latent space. However, such a projection function remains several crucial bottlenecks: 1) information loss and coding redundancy are inevitable; 2) the available information of semantic labels is not well-explored; 3) the learned latent embedding lacks explicit semantic meaning. To overcome these limitations, we propose a novel supervised Discrete Auto-Encoder Hashing (DAEH) framework, in which a linear auto-encoder can effectively project the semantic labels of images into a latent representation space. Instead of using the visual feature projection, the proposed DAEH framework skillfully explores the semantic information of supervised labels to refine the latent feature embedding and further optimizes hashing function. Meanwhile, we reformulate the objective and relax the discrete constraints for the binary optimization problem. Extensive experiments on Caltech-256, CIFAR-10, and MNIST datasets demonstrate that our method can outperform the state-of-the-art hashing baselines.
Shuguang Zhao, Bingzhi Chen, Zheng Zhang 0006, Guangming Lu 0002
MMAsia2
2021 Multimodal Emotion Recognition With Temporal and Semantic Consistency
abstract
Automated multimodal emotion recognition has become an emerging but challenging research topic in the fields of affective learning and sentiment analysis. The existing works mainly focus on developing multimodal fusion strategies to incorporate different emotion-related features. However, they fail to explore the inherent contextual consistency to reconcile the emotional information across modalities. In this paper, we propose a novel Time and Semantic Interaction Network (TSIN), which concurrently incorporates the advantages of temporal and semantic consistency into the multimodal emotion recognition task. Specifically, a well-designed Speech and Text Embedding (STE) module is devoted to formulating the initial embedding spaces by respectively building the modality-specific representations of speech and text. Instead of separately learning or directly fusing the acoustic and textual features, we propose a well-defined Time and Semantic Interaction (TSI) module to conduct the emotional parsing and sentiment refining by performing the fine-grained temporal alignment and cross-modal semantic interaction. Benefitting from temporal and semantic consistency constraints, both speech-text embeddings can be interactively optimized and fine-tuned in the learning process. In this way, the learnt acoustics and textual features can jointly and efficiently predict the final emotional state. Extensive experiments on the IEMOCAP dataset demonstrate the superiorities of our TSIN framework in comparison with state-of-the-art baselines.
Bingzhi Chen, Mi-Xiao Hou, Zheng Zhang 0006, Guangming Lu 0002, David Zhang 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Deep Active Context Estimation for Automated COVID-19 Diagnosis
abstract
Many studies on automated COVID-19 diagnosis have advanced rapidly with the increasing availability of large-scale CT annotated datasets. Inevitably, there are still a large number of unlabeled CT slices in the existing data sources since it requires considerable consuming labor efforts. Notably, cinical experience indicates that the neighboring CT slices may present similar symptoms and signs. Inspired by such wisdom, we propose DACE, a novel CNN-based deep active context estimation framework, which leverages the unlabeled neighbors to progressively learn more robust feature representations and generate a well-performed classifier for COVID-19 diagnosis. Specifically, the backbone of the proposed DACE framework is constructed by a well-designed Long-Short Hierarchical Attention Network (LSHAN), which effectively incorporates two complementary attention mechanisms, i.e., short-range channel interactions (SCI) module and long-range spatial dependencies (LSD) module, to learn the most discriminative features from CT slices. To make full use of such available data, we design an efficient context estimation criterion to carefully assign the additional labels to these neighbors. Benefiting from two complementary types of informative annotations from -nearest neighbors, i.e., the majority of high-confidence samples with pseudo labels and the minority of low-confidence samples with hand-annotated labels, the proposed LSHAN can be fine-tuned and optimized in an incremental learning manner. Extensive experiments on the Clean-CC-CCII dataset demonstrate the superior performance of our method compared with the state-of-the-art baselines.
Bingzhi Chen, Yishu Liu 0001, Zheng Zhang 0006, Yingjian Li 0001, Zhao Zhang 0001, Guangming Lu 0002, Hongbing Yu
ACM Trans. Multim. Comput. Commun. Appl.1
2020 Two-stream collaborative network for multi-label chest X-ray Image classification with lung segmentation
Bingzhi Chen, Zheng Zhang 0006, Jianyong Lin, Yi Chen 0023, Guangming Lu 0002
Pattern Recognit. Lett.1
2020 Lesion Location Attention Guided Network for Multi-Label Thoracic Disease Classification in Chest X-Rays
abstract
Traditional clinical experiences have shown the benefit of lesion location attention for improving clinical diagnosis tasks. Inspired by this point of interest, in this paper we propose a novel lesion location attention guided network named LLAGnet to focus on the discriminative features from lesion locations for multi-label thoracic disease classification in chest X-rays (CXRs). By revealing the equivalence of the region-level attention (RLA) and channel-level attention (CLA), we find that the RLA is available as priors for object localization while the CLA implicitly provides high weights to the attractive channels, which both enable lesion location attention excitation. To integrate the advantages from both mechanisms, the proposed LLAGnet is structured with two corresponding attention modules, i.e., the RLA and CLA modules. Specifically, the RLA module consists of the global and local branches. And the weakly supervised attention mechanism embedded in the global branch can obtain visual regions of lesion locations by back-propagating gradients. Then the optimal attention region is amplified and applied to the local branch to provide more fine-grained features for the image classification. Finally, the CLA module adaptively enhances the weights of channel-wise features from the lesion locations by modeling interdependencies among channels. Extensive experiments on the ChestX-ray14 dataset clearly substantiate the effectiveness of LLAGnet as compared with the state-of-the-art baselines.
Bingzhi Chen, Jinxing Li 0003, Guangming Lu 0002, David Zhang 0001
IEEE J. Biomed. Health Informatics1
2020 Label Co-Occurrence Learning With Graph Convolutional Networks for Multi-Label Chest X-Ray Image Classification
abstract
Existing multi-label medical image learning tasks generally contain rich relationship information among pathologies such as label co-occurrence and interdependency, which is of great importance for assisting in clinical diagnosis and can be represented as the graph-structured data. However, most state-of-the-art works only focus on regression from the input to the binary labels, failing to make full use of such valuable graph-structured information due to the complexity of graph data. In this paper, we propose a novel label co-occurrence learning framework based on Graph Convolution Networks (GCNs) to explicitly explore the dependencies between pathologies for the multi-label chest X-ray (CXR) image classification task, which we term the "CheXGCN". Specifically, the proposed CheXGCN consists of two modules, i.e., the image feature embedding (IFE) module and label co-occurrence learning (LCL) module. Thanks to the LCL model, the relationship between pathologies is generalized into a set of classifier scores by introducing the word embedding of pathologies and multi-layer graph information propagation. During end-to-end training, it can be flexibly integrated into the IFE module and then adaptively recalibrate multi-label outputs with these scores. Extensive experiments on the ChestX-Ray14 and CheXpert datasets have demonstrated the effectiveness of CheXGCN as compared with the state-of-the-art baselines.
Bingzhi Chen, Jinxing Li 0003, Guangming Lu 0002, Hongbing Yu, David Zhang 0001
IEEE J. Biomed. Health Informatics1
2019 Mask-Most Net: Mask Approximation Based Multi-oriented Scene Text Detection Network
abstract
In this paper, a novel multi-task cascade framework, which jointly takes the detection and the segmentation into account, is presented for the scene text detection. To address the issue of multi-oriented scene text detection, we propose an instance-level mask approximation method through the auxiliary regression task on center and corner points. Specifically, the text instance in the image is first coarsely detected, followed by a contextual module which can capture more accurate instances. To cope with the scale variation existing in these detected instances, a combination of high-level semantic and low-level features is further exploited, achieving more robust and better performance. A series of experiments conducted on different benchmark datasets demonstrate the effectiveness of the proposed method.
Xiaobao Guo, Jinxing Li 0003, Bingzhi Chen, Guangming Lu 0002
ICME3
2019 Multi-label Chest X-Ray Image Classification via Label Co-occurrence Learning
Bingzhi Chen, Yao Lu 0008, Guangming Lu 0002
PRCV (2)1