EDBT 2026 Demo / reviewers in the wild / expert
Bingzhi Chen
dblp:34/7694
· DBLP profile ↗
62ranked-venue papers
19as first author
57since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 48 · 14 first-author · 46 since 2021Artificial intelligence and machine learning · 14 · 5 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BayesVQA: Energy-Guided Bayesian Debiasing for Language-Bias-Robust Visual Question AnsweringabstractNumerous studies have demonstrated that Visual Question Answering (VQA) models are vulnerable to language priors and dataset biases, often leading to spurious correlations between questions and answers. As a result, these models excessively rely on linguistic cues, neglecting essential visual information and causing representational distortions. To address this challenge, we propose a novel Bayesian debiasing framework termed BayesVQA, which integrates three carefully designed mechanisms: Energy-guided Prior Variance (EPV), Energy-guided Posterior Sampling (EPS), and Energy-guided Likelihood Reweighting (ELR). Specifically, we explicitly decompose each sample's latent representation into a biased feature and a stochastic corrective perturbation δ. Using a Bayesian formulation, we model the posterior distribution of the perturbation δ conditioned on the predictive uncertainty, quantified via calibrated energy scores. To mitigate language bias, the posterior is optimized through energy-driven variational inference with an uncertainty-adaptive prior and sampling strategy. Moreover, the ELR mechanism incorporates an energy-based weighting of the reconstruction objective and enforces an energy-coherence constraint to emphasize challenging, high-uncertainty instances and align model confidence before and after debiasing. Extensive experiments conducted across multiple standard VQA benchmarks consistently validate the superior performance of our BayesVQA method over state-of-the-art competitors under distributional shifts and challenging bias conditions. Huanjia Zhu, Xiangwen Deng, Qinghao Zhong, Bingzhi Chen |
AAAI | 5 |
| 2026 | Cross Modal Fine-grained Alignment via Granularity-aware and Region-uncertain ModelingabstractFine-grained image-text alignment is a pivotal challenge in multimodal learning, underpinning key applications such as visual question answering, image captioning, and vision-language navigation. Unlike global alignment, fine-grained alignment requires precise correspondence between localized visual regions and textual tokens, often hindered by noisy attention mechanisms and oversimplified modeling of cross-modal relationships. In this work, we identify two fundamental limitations of existing approaches: the lack of robust intra-modal mechanisms to assess the significance of visual and textual tokens, leading to poor generalization in complex scenes; and the absence of fine-grained uncertainty modeling, which fails to capture the one-to-many and many-to-one nature of region-word correspondences. To address these issues, we propose a unified approach that incorporates significance-aware and granularity-aware modeling and region-level uncertainty modeling. Our method leverages modality-specific biases to identify salient features without relying on brittle cross-modal attention, and represents region features as a mixture of Gaussian distributions to capture fine-grained uncertainty. Extensive experiments on Flickr30K and MS-COCO demonstrate that our approach achieves state-of-the-art performance across various backbone architectures, significantly enhancing the robustness and interpretability of fine-grained image-text alignment. Haoming Zhou, Yishu Liu 0001, Bingzhi Chen, Yuncheng Jiang 0004 |
AAAI | 4 |
| 2026 | S2PD: Physically grounded augmentations and stable parameter updates for cross-domain few-shot learning
Shuai Kang, Yunyu Zou, Ziteng Hong, Bingzhi Chen, Guangming Lu 0002 |
Neurocomputing | 5 |
| 2026 | DS2VP: Dynamically-Selected Spatially Visual PromptingabstractThe significant effectiveness of prompt tuning for computer vision tasks has been extensively demonstrated in numerous studies. As a widely feasible solution, the spatial modeling paradigm aims to overcome the limitations of sequence modeling paradigm in capturing spatial relationships within images by learning a prompt token map and aligning it spatially with the image token map. However, such spatial modeling paradigms of visual prompt tuning still face two potential challenges: 1) Most existing methods fail to design individual prompts for different images, and the learned prompts have the same static effect on all images. 2) The strategy of existing methods overlooks the selection of key spatial information and indiscriminately prompts all information within the image. In this work, we propose a novel Dynamically-Selected and Spatial Visual Prompting, termed as DS2VP, which aims to effectively utilize the key spatial information of the input image and enable dynamic visual prompt selection. Specifically, our DS2VP approach is meticulously designed to leverage the key index generator to filter key regions of the image for determining the spatial target of prompts, thus enabling dynamic selection of prompts for different images. By adding prompt tokens at selected key locations, an image prompt fusion module is deployed by adapting the learnable prompt tokens into the input image tokens, further achieving a fine-grained spatial alignment. Moreover, we propose a multi-level prompt interaction module that facilitates interactions between visual prompts at different levels to enhance feature representations across various semantic levels. Extensive experiments on two challenging benchmarks for image classification have demonstrated the superiority of DS2VP over other state-of-the-art methods for visual prompt tuning. Yishu Liu 0001, Bingzhi Chen, Guangming Lu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | InfoARD: Enhancing Adversarial Robustness Distillation With Attack-Strength Adaptation and Mutual-Information MaximizationabstractAdversarial distillation (AD) aims to mitigate deep neural networks' inherent vulnerability to adversarial attacks, thereby providing robust protection for compact models through teacher-student interactions. Despite advancements, existing AD studies still suffer from insufficient robustness due to the limitations of fixed attack strength and attention region shifts. To address these challenges, we propose a strength-adaptive Info-maximizing Adversarial Robustness Distillation paradigm, namely "InfoARD", which strategically incorporates the Attack-Strength Adaptation (ASA) and Mutual-Information Maximization (MIM) to enhance adversarial robustness against adversarial attacks and perturbations. Unlike previous adversarial training (AT) methods that utilize fixed attack strength, the ASA mechanism is designed to capture smoother and generalized classification boundaries by dynamically tailoring the attack strength based on the characteristics of individual instances. Benefiting from mutual information constraints, our MIM strategy ensures the student model effectively learns from various levels of feature representations and attention patterns, thereby deepening the student model's understanding of the teacher model's decision-making processes. Furthermore, a comprehensive multi-granularity distillation is conducted to capture knowledge across multiple dimensions, enabling a more effective transfer of knowledge from the teacher model to the student model. Note that our InfoARD can be seamlessly integrated into existing AD frameworks, further boosting the adversarial robustness of deep learning models. Extensive experiments on various challenging datasets consistently demonstrate the effectiveness and robustness of our InfoARD, surpassing previous state-of-the-art methods. Ruihan Liu, Jieyi Cai, Yishu Liu 0001, Sudong Cai, Bingzhi Chen, Yulan Guo, Mohammed Bennamoun |
IEEE Trans. Image Process. | 5 |
| 2026 | Toward Bidirectional Adaptability for Few-Shot Class-Incremental Learning With Forward-Backward Knowledge TransferabstractThe development of Deep Neural Networks (DNNs) has enabled AI-driven models to excel in recognizing a limited set of classes within static environments. As AI systems progress, few-shot class-incremental learning (FSCIL) aims to expand their understanding of novel classes from minimal samples while retaining knowledge of previously encountered ones. However, most existing FSCIL models face significant challenges, includinginadequate adaptabilityandcatastrophic forgetting, which hinder their ability to maintain robust forward and backward learning capabilities. To address these issues, this paper proposes a novel Forward-Backward Knowledge Transfer (FBKT) paradigm, which strategically integrates forward distribution adaptation (FDA) and backward semantic alignment (BSA) mechanisms to achieve bidirectional adaptability in knowledge transfer. The FDA mechanism enhances forward adaptability by expanding and reserving the embedding space for new classes using semantic-irrelevant masked images as virtual negative classes, thereby mitigating data overfitting. It also employs self-supervised representation learning to utilize semantic-relevant local embeddings as additional positive samples, fostering class separation and generalization. Meanwhile, the BSA mechanism ensures the semantic consistency of previously learned classes across sessions during class-incremental learning, promoting smoother backward adaptability and reducing model degradation. Extensive experiments conducted on multiple benchmark datasets consistently highlight the superior performance and effectiveness of our FBKT compared to state-of-the-art methods. Bingzhi Chen, Sudong Cai, Xiaozhao Fang, Mohammed Bennamoun, Shengli Xie 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | Towards Robust Visual Question Answering via Prompt-Driven Geometric HarmonizationabstractVisual Question Answering (VQA) has garnered significant attention as a crucial link between vision and language, aimed at generating accurate responses to visual queries. However, current VQA models still struggle with the challenges of minority class collapse and spurious semantic correlations posed by language bias and imbalanced distributions. To address these challenges, this paper proposes a novel Prompt-Driven Geometric Harmonization (PDGH) paradigm, which integrates both geometric structure and information entropy principles to enhance the ability of VQA models to generalize effectively across diverse scenarios. Specifically, our PDGH approach is meticulously designed to generate image-generated prompts that are guided by specific question cues, facilitating a more accurate and context-aware understanding of the visual content. Moreover, we project the prompt-visual-question and visual-question joint representations into a unified hypersphere space, applying feature weight self-orthogonality and prompt-information entropy correction constraints to optimize the margin, further alleviating minority class collapse and correcting language bias. To maintain the geometric integrity of the representation space, we introduce multi-space geometric contrast constraints to minimize the impact of spurious priors introduced during training. Finally, a semantic matrix is constructed for the coordinated joint representation to ensure that the learned instances are semantically consistent and improve reasoning ability. Extensive experiments on various general and medical VQA datasets demonstrate the consistent superiority of our PDGH approach over existing state-of-the-art baselines. Yishu Liu 0001, Congcong Wen, Guangming Lu 0002, Bingzhi Chen |
AAAI | 6 |
| 2025 | OralXrays-9: Towards Hospital-Scale Panoramic X-ray Anomaly Detection via Personalized Multi-Object Query-Aware MiningabstractIn clinical practice, panoramic dental radiography is a widely employed imaging technique that can provide a detailed and comprehensive view of dental structures and surrounding tissues for identifying various oral anomalies. However, due to the complexity of oral anomalies and the scarcity of available data, existing research still suffers from substantial challenges in automated oral anomaly detection. To this end, this paper presents a new hospital-scale panoramic X-ray benchmark, namely "OralXrays-91", which consists of 12,688 panoramic X-ray images with 84,113 meticulously annotated instances across nine common oral anomalies. Correspondingly, we propose a personalized Multi-Object Query-Aware Mining (MOQAM) paradigm, which jointly incorporates the Distribution-IoU Region Proposal Network (DI-RPN) and Class-Balanced Spherical Contrastive Regularization (CB-SCR) mechanisms to address the challenges posed by multi-scale variations and class-imbalanced distributions. To the best of our knowledge, this is the first attempt to develop AI-driven diagnostic systems specifically designed for multi-object oral anomaly detection, utilizing publicly available data resources. Extensive experiments on the newly-published OralXrays-9 dataset and real-world nature scenarios consistently demonstrate the superiority of our MOQAM in revolutionizing oral healthcare practices. Bingzhi Chen, Sisi Fu, Xiaocheng Fang, Jieyi Cai, Minhua Lu, Yishu Liu 0001 |
CVPR | 1 |
| 2025 | Learning Hierarchical Attribute Prompt for Vision-Language ModelsabstractPrompt learning is a common strategy for adapting Visual Language Models (VLMs) to downstream tasks by fine-tuning prompts for task-specific performance. However, existing methods face two key challenges: overfitting to base classes, which limits generalization to novel classes, and the dependence on manually generated or LLM-based descriptions, which are time-consuming and error-prone. To address these issues, we propose the Learning Hierarchical Attribute Prompt (LHAP) method, which introduces fine-grained semantic alignment through hierarchical prompts. By autonomously extracting visual attributes from images, LHAP generates local-level prompts (LLP) to capture fine-grained semantics and global-level prompts (GLP) to model overall semantics. The combination of LLP and GLP not only improves generalization but also mitigates errors and inefficiencies from manual or LLM-based descriptions. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority and robustness of LHAP over state-of-the-art methods. Jun Liang 0002, Yunyu Zou, Yalong Cheng, Bingzhi Chen |
ICASSP | 6 |
| 2025 | A2GP-SF: Enhancing Few-shot Class Incremental Learning via Attribute Generative Prompting and Adaptive Sharpness FlatteningabstractFew-shot Class Incremental Learning (FSCIL) aims to incrementally learn new classes with limited examples while retaining knowledge of previously learned classes. Recent advancements in prompt tuning for large pre-trained models have shown promise in FSCIL. However, current FSCIL methods still suffer from challenges like insufficient plasticity and limited generalization. To tackle these challenges, we propose a novel prompt tuning-based framework named A2GP-SF, which integrates attribute generative prompting (AGP) and adaptive sharpness flattening (ASF). The proposed AGP paradigm dynamically generates attribute-aware prompts for each instance, facilitating better semantics learning and enhancing plasticity. Additionally, the ASF mechanism aims to mitigate overfitting by applying adaptive perturbations to flatten sharpness, with these perturbations adjusted based on gradient norm changes, thereby enhancing the model’s robustness and generalization. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority of our proposed A2GP-SF framework. Desen Wang, Sisi Fu, Congcong Wen, Bingzhi Chen |
ICASSP | 6 |
| 2025 | Towards Differential Optimization: Rehearsal-Free Class-Incremental Learning with Slow Learners and Fast AdaptersabstractClass-incremental learning (CIL) enables models to learn new tasks without forgetting previously acquired knowledge. However, existing CIL approaches often struggle with inadequate adaptation to task-specific feature spaces and catastrophic forgetting of previously-acquired knowledge, compromising the models’ plasticity and stability. To address these challenges, this paper proposes a novel differential optimization paradigm called DO-CIL, which incorporates task-agnostic slow learner (TSL) with task-specific fast adapter (TFA) for rehearsal-free CIL. Specifically, TSL aims to effectively capture shared knowledge with low learning rates for robust generalization, while TFA allows pre-trained models to adapt to new task-specific feature spaces. Benefitting from the classifier retraining strategy, a learnable semantic shift network is also proposed to align prototypes with the evolving model representation, facilitating the retraining of task-specific classifiers based on these updated prototypes. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority and effectiveness of our DO-CIL approach compared to state-of-the-art baselines. Yinghong Chen, Huanjia Zhu, Jieyi Cai, Jun Liang 0002, Bingzhi Chen |
ICASSP | 6 |
| 2025 | Enhancing Incomplete Multimodal Learning via Modal Complementary RecoveringabstractMultimodal learning presents significant challenges arising from the unpredictable absence of modalities during both training and testing phases. Existing recovery methods struggle to leverage the available data, which can introduce additional noise during the recovery process and degrade performance. To mitigate these issues, we introduce a novel Modal Complementary Recovering (MCR) paradigm that strategically integrates both Complementary Graph-based Recovery (CGR) and Topological Low-Rank Adaptation (ToRA) mechanisms to enhance the effectiveness and reliability of incomplete multimodal learning. For effective exploitation of the complementarity among different modalities, the main objective of CGR is to employ bidirectional mapping flows trained on a small subset of complete data to learn complementary graphs across all modalities. By constructing entity-relationship diagram specific to dataset, ToRA is designed to enhance the fine-tuning process by incorporating topological prompt. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority of our MCR paradigm in comparison to state-of-the-art baselines. Meirong Ding, Chuang Zou, Wenxiu Cai, Bingzhi Chen |
ICASSP | 6 |
| 2025 | Advancing Few-Shot Class-Incremental Learning with Virtual Prototype Guidance PromptingabstractFew-Shot Class-Incremental Learning (FSCIL) aims to incrementally learn new class knowledge from limited samples while preserving previously knowledge from encountered classes. However, existing FSCIL methods encounter two primary challenges: (1) inadequate adaptation, where overfitting to new classes compromises the model’s adaptability, and (2) catastrophic forgetting, where previously learned knowledge is not well preserved. In this paper, we propose the Virtual Prototype Guidance Prompting (VPGP) paradigm, integrating the Multi-Grained Prompt (MGP) and Virtual-Prototype Guidance (VPG) strategies. Specifically, MGP enhances adaptation and prevents overfitting by introducing domain-general and fine-grained prompts, expanding the embedding space to capture core feature representations of novel classes. Meanwhile, VPG mitigates catastrophic forgetting by employing a dynamic fusion strategy to retrieve robust old class knowledge and generate virtual prototype, guiding the model to maintain learned knowledge across different sessions. Extensive experiments on multiple benchmark datasets demonstrate the superiority of our proposed VPGP framework. Huanjia Zhu, Xiaocheng Fang, Jun Liang 0002, Bingzhi Chen |
ICASSP | 5 |
| 2025 | Revisiting DETR for Small Object Detection via Noise-Resilient Query OptimizationabstractDespite advancements in Transformer-based detectors for small object detection (SOD), recent studies show that these detectors still face challenges due to inherent noise sensitivity in feature pyramid networks (FPN) and diminished query quality in existing label assignment strategies. In this paper, we propose a novel Noise-Resilient Query Optimization (NRQO) paradigm, which innovatively incorporates the Noise-Tolerance Feature Pyramid Network (NT-FPN) and the Pairwise-Similarity Region Proposal Network (PS-RPN). Specifically, NTFPN mitigates noise during feature fusion in FPN by preserving spatial and semantic information integrity. Unlike existing label assignment strategies, PS-RPN generates a sufficient number of high-quality positive queries by enhancing anchor-ground truth matching through position and shape similarities, without the need for additional hyperparameters. Extensive experiments on multiple benchmarks consistently demonstrate the superiority of NRQO over state-of-the-art baselines. Xiaocheng Fang, Jieyi Cai, Wenxiu Cai, Yishu Liu 0001, Bingzhi Chen |
ICME | 6 |
| 2025 | Task-Aware Knowledge Prompt and Distillation for Cross-Domain Few-Shot LearningabstractCross-Domain Few-Shot Learning (CD-FSL) aims to recognize unseen classes from target domains using only limited labeled samples. However, mainstream CD-FSL methods face two key challenges: (1) domain gap, arising from distributional differences between the source and target domains, and (2) overfitting, which occurs due to the small number of labeled samples in target domains, causing the model to overfit to these few samples. To address these challenges, we propose a novel Task-Aware Knowledge Prompt and Distillation (TKPD) method for CD-FSL, which integrates the Vision-Text Domain Prompt (VTDP) and Attribute-Task Knowledge Distillation (ATKD) modules. VTDP mitigates the domain gap by generating domain prompts to acquire domain-relevant knowledge, while ATKD strengthens the model by incorporating both homologous and heterogeneous knowledge, extending the knowledge base beyond the limited labeled samples and mitigating overfitting. Extensive experiments on 13 benchmark datasets validate the effectiveness of the proposed TKPD method. Jun Liang 0002, Yunyu Zou, Yalong Cheng, Yishu Liu 0001, Bingzhi Chen |
ICME | 7 |
| 2025 | Enhancing Few-Shot Class-Incremental Learning via Cross-Modal Bias AlignmentabstractFew-Shot Class-Incremental Learning (FSCIL) aims to learn new classes from limited samples while retaining knowledge of previously learned classes. Prompt tuning on large-scale pre-trained models has achieved certain success in FSCIL. However, current prompt-based approaches still suffer from the challenges of modality bias and catastrophic forgetting. To address these challenges, we propose a prompt tuning-based Cross-Modal Bias Alignment (CMBA) framework, which integrates Prompting Bias Alignment (PBA) and Dynamic Prototype Tuning (DPT). PBA mitigates modality bias by aligning the prompting biases between the visual and textual modalities. Meanwhile, DPT relieves catastrophic forgetting by introducing pseudo-sample features of old classes in the incremental sessions, which are generated by dynamically tuning prototypes that adapt to the evolving feature space. Extensive experiments on multiple FSCIL datasets demonstrate the consistent superiority of our CMBA approach over the existing state-of-the-art baselines. Desen Wang, Yishu Liu 0001, Bingzhi Chen |
ICME | 5 |
| 2025 | Cause-Effect Driven Optimization for Robust Medical Visual Question Answering with Language BiasesabstractExisting Medical Visual Question Answering (Med-VQA) models often suffer from language biases, where spurious correlations between question types and answer categories are inadvertently established. To address these issues, we propose a novel Cause-Effect Driven Optimization framework called CEDO, that incorporates three well-established mechanisms, i.e., Modality-driven Heterogeneous Optimization (MHO), Gradient-guided Modality Synergy (GMS), and Distribution-adapted Loss Rescaling (DLR), for comprehensively mitigating language biases from both causal and effectual perspectives. Specifically, MHO employs adaptive learning rates for specific modalities to achieve heterogeneous optimization, thus enhancing robust reasoning capabilities. Additionally, GMS leverages the Pareto optimization method to foster synergistic interactions between modalities and enforce gradient orthogonality to eliminate bias updates, thereby mitigating language biases from the effect side, i.e., shortcut bias. Furthermore, DLR is designed to assign adaptive weights to individual losses to ensure balanced learning across all answer categories, effectively alleviating language biases from the cause side, i.e., imbalance biases within datasets. Extensive experiments on multiple traditional and bias-sensitive benchmarks consistently demonstrate the robustness of CEDO over state-of-the-art competitors. Huanjia Zhu, Yishu Liu 0001, Xiaozhao Fang, Guangming Lu 0002, Bingzhi Chen |
IJCAI | 5 |
| 2025 | PerioDet: Large-Scale Panoramic Radiograph Benchmark for Clinical-Oriented Apical Periodontitis Detection
Xiaocheng Fang, Jieyi Cai, Chengju Zhou, Minhua Lu, Bingzhi Chen |
MICCAI (16) | 6 |
| 2025 | Med-BiasX: Robust Medical Visual Question Answering with Language Biases
Huanjia Zhu, Yishu Liu 0001, Chengju Zhou, Guangming Lu 0002, Bingzhi Chen |
MICCAI (14) | 5 |
| 2025 | Label Prediction Inherited Hashing for Cross-Modal Retrieval: Applying Supervised Hashing to Unsupervised TasksabstractSupervised cross-modal hashing has achieved remarkable progress in retrieving related items across different modalities. However, in practical applications, a significant portion of data remains unlabeled, such as online data on websites, which must be included for effective retrieval. To address this challenge, while maintaining the high accuracy and efficiency of supervised methods, few works have attempted to adapt existing supervised techniques to handle unsupervised tasks through a general modular approach. To this end, we introduce a novel cross-modal hashing method, termed Label Prediction Inherited Hashing (LPIH). Initially, LPIH leverages labeled data to learn high-quality general label functions using supervised methods. Subsequently, it inherits the existing hash codes from existing supervised methods to further refine the pseudo-label information. Finally, LPIH integrates the refined pseudo-label information with the existing hash functions to learn new hash functions specifically tailored for unsupervised tasks. Extensive experimental results on three public datasets demonstrate the superior performance of LPIH compared to state-of-the-art (SOTA) cross-modal hashing methods. Specifically, LPIH achieves an average precision improvement of 5% over SOTA methods, highlighting its effectiveness in bridging the gap between supervised and unsupervised learning in the context of cross-modal retrieval. Kaihang Jiang, Wai Keung Wong, Jianyang Qin, Xiaozhao Fang, Jie Wen 0001, Bingzhi Chen, Hongbo Gao 0001 |
ACM Multimedia | 6 |
| 2025 | CauRDG: Enhancing Domain Generalization with Causal-Driven Semantic Consistency ReasoningabstractDomain generalization (DG) plays a pivotal role in enabling models to maintain robust performance across heterogeneous environments. However, existing DG methods are fundamentally constrained by two intertwined limitations: (1) causal misalignment, which stems from undifferentiated feature encoding that entangles causal mechanisms with environmental biases; (2)semantic conflict arises when conventional adaptation methods find it challenging to balance the preservation of class discriminability with the mitigation of domain-specific distribution discrepancies. To address these challenges of DG, we propose a novel Causal-Driven Semantic Consistency Reasoning (CauRDG) method, which synergistically integrates Prototype-Guided Causal Disentanglement (PGCD) and Dual-Space Semantic Disambiguation (DSSD). Specifically, PGCD constructs a causal framework that identifies stable relationships and decouples invariant mechanisms from domain-specific variations, preserving causal consistency while adapting to contextual differences. DSSD harnesses a dual-space paradigm, enhancing local categorical clarity and maintaining global conceptual unity, thus balancing domain-specific precision with cross-domain coherence. The robustness provided by CauRDG ensures robust extraction and interpretation of essential features by preserving invariant causal structures, thereby harmonizing discriminative semantics with domain-varying contexts. Extensive experiments on multiple benchmark datasets consistently demonstrate the effectiveness and superiority of our CauRDG over state-of-the-art baselines. Zongxin Liu 0003, Yishu Liu 0001, Guangming Lu 0002, Xiaoling Luo 0001, Bingzhi Chen |
ACM Multimedia | 5 |
| 2025 | PET-GPRA: Rethinking PET with Gradient-Aware Prompting and Router-Free Adapters for Few-shot Class-Incremental LearningabstractFew-Shot Class-Incremental Learning (FSCIL) aims to continuously learn novel concepts from limited training samples without forgetting previously encountered classes. Recent advancements have leveraged Parameter-Efficient Tuning (PET) strategies on pre-trained models to enhance FSCIL performance. However, current PET-based FSCIL approaches still suffer from the challenges posed by catastrophic collapse of general prompt and limited adaptability of specific prompt . To this end, we redefine the function of the PET paradigm with both gradient-aware prompting (GAP) and router-free adapters (RFA) to boost the performance of FSCIL, termed as "PET-GPRA". To dynamically balance the retention of previously learned general knowledge and the acquisition of novel class information across sessions, the GAP paradigm adaptively adjusts the updated gradient of the general prompt by leveraging the angular relationship between the general knowledge gradient and the novel knowledge gradient. Meanwhile, the RFA mechanism utilizes the semantic similarity between class attributes to replace the routing network, guiding the integration of adapter information, in which adapters serve as specific prompts to enhance the adaptability. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority and effectiveness of our proposed PET-GPRA framework over state-of-the-art baselines. Yishu Liu 0001, Desen Wang, Xiaoling Luo 0001, Bingzhi Chen, Guangming Lu 0002 |
ACM Multimedia | 5 |
| 2025 | SG-FSL: Cross-Domain Few-Shot Learning with Style-Decoupled Augmentation and Gradient-Conflict AdjustmentabstractCross-Domain Few-Shot Learning (CD-FSL) aims to transfer knowledge acquired from a source domain with abundant data to the target domain with limited labeled samples. Recent advancements have enhanced model generalization through Perturbation Augmentation (PA), facilitating more effective knowledge transfer. However, PA-based CD-FSL methods still suffer from two critical challenges, i.e., (1) limited diversity of augmented samples, making it difficult to cover the true distribution of unseen domains, and (2) conflicting gradients during model optimization, where augmented and original samples drive the model's optimization in opposing directions. To address these issues, we propose a novel PA-based framework with Style-Decoupled Augmentation (SDA) and Gradient-Conflict Adjustment (GCA) for Cross-Domain Few-Shot Learning, which is termed ''SG-FSL''. Specifically, SDA decouples the source domain style into style weights and basis styles, generating diverse unseen styles by perturbing the style weights to reweight the basis styles. Meanwhile, GCA leverages the angular relationships between the domain-specific gradient directions of augmented and original features, adaptively adjusting the gradient directions of original features to ensure that the model acquires diverse domain knowledge without interference, guiding it toward conflict-free optimization. Comprehensive experiments on multiple benchmark datasets consistently demonstrate the effectiveness and superiority of our method over state-of-the-art baselines. Yunyu Zou, Yishu Liu 0001, Jun Liang 0002, Bingzhi Chen |
ACM Multimedia | 4 |
| 2025 | Language-Bias-Resilient Visual Question Answering via Adaptive Multi-Margin Collaborative DebiasingabstractLanguage bias in Visual Question Answering (VQA) arises when models exploit spurious statistical correlations between question templates and answers, particularly in out-of-distribution scenarios, thereby neglecting essential visual cues and compromising genuine multimodal reasoning. Despite numerous efforts to enhance the robustness of VQA models, a principled understanding of how such bias originates and influences model behavior remains underdeveloped. In this paper, we address this gap through a comprehensive empirical and theoretical analysis, revealing that modality-specific gradient imbalances, which originate from the inherent heterogeneity of multimodal data, lead to skewed feature fusion and biased classifier weights. To alleviate these issues, we propose a novel Multi-Margin Collaborative Debiasing (MMCD) framework that adaptively integrates frequency-, confidence-, and difficulty-aware angular margins with a dynamic difficulty-aware contrastive learning mechanism, to dynamically reshape decision boundaries. Extensive experiments across multiple challenging VQA benchmarks confirm the consistent superiority of our proposed MMCD over state-of-the-art baselines in combating language bias. Huanjia Zhu, Shuyuan Zheng, Yishu Liu 0001, Sudong Cai, Bingzhi Chen |
NeurIPS | 5 |
| 2025 | LBF-VQA: Towards Language Bias-Free Visual Question Answering With Multi-Space Collaborative Debiasing
Yishu Liu 0001, Huanjia Zhu, Bingzhi Chen, Xiaozhao Fang, Guangming Lu 0002, Shengli Xie 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Toward Robust Semi-Supervised Distribution Alignment Against Label Distribution Shift With Noisy AnnotationsabstractDeep learning-based AI models typically require a large amount of high-quality annotated data to achieve optimal performance. However, thelabel distribution shiftcaused by noisy annotations can lead to perturbations in the classification boundary, reducing the robustness and generalization capabilities of deep learning models. To mitigate this issue, we transform the problem of learning from noisy labels into a semi-supervised learning problem, and propose a novel Semi-Supervised Distribution Alignment (SSDA) framework that strategically integrates noise-robust distribution alignment within a unified semi-supervised learning paradigm for combating noisy labels. By leveraging the similarity distribution between historical predictions, the proposed SSDA approach benefits from a flexible multi-historical regression modeling strategy, which aims to identify high-confidence samples/pairs and recalibrate the label shift through pseudo-labels. Furthermore, our approach employs a comprehensive multi-granularity distribution adaptation strategy, incorporating both instance-wise and class-aware distribution alignment to quantitatively minimize semantic discrepancies across different mixed feature domains. In this way, our SSDA approach ultimately achieves more resilient and generalizable performance against label noise, even in the presence of substantial noise. Extensive experiments conducted on multiple simulated and real-world noisy benchmark datasets consistently demonstrate the superiority and effectiveness of our SSDA method compared to existing state-of-the-art baselines. Bingzhi Chen, Zhanhao Ye, Yishu Liu 0001, Xiaozhao Fang, Guangming Lu 0002, Shengli Xie 0001, Xuelong Li 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | CariesXrays: Enhancing Caries Detection in Hospital-Scale Panoramic Dental X-rays via Feature Pyramid Contrastive LearningabstractDental caries has been widely recognized as one of the most prevalent chronic diseases in the field of public health. Despite advancements in automated diagnosis across various medical domains, it remains a substantial challenge for dental caries detection due to its inherent variability and intricacies. To bridge this gap, we release a hospital-scale panoramic dental X-ray benchmark, namely “CariesXrays”, to facilitate the advancements in high-precision computer-aided diagnosis for dental caries. It comprises 6,000 panoramic dental X-ray images, with a total of 13,783 instances of dental caries, all meticulously annotated by dental professionals. In this paper, we propose a novel Feature Pyramid Contrastive Learning (FPCL) framework, that jointly incorporates feature pyramid learning and contrastive learning within a unified diagnostic paradigm for automated dental caries detection. Specifically, a robust dual-directional feature pyramid network (D2D-FPN) is designed to adaptively capture rich and informative contextual information from multi-level feature maps, thus enhancing the generalization ability of caries detection across different scales. Furthermore, our model is augmented with an effective proposals-prototype contrastive regularization learning (P2P-CRL) mechanism, which can flexibly bridge the semantic gaps among diverse dental caries with varying appearances, resulting in high-quality dental caries proposals. Extensive experiments on our newly-established CariesXrays benchmark demonstrate the potential of FPCL to make a significant social impact on caries diagnosis. Bingzhi Chen, Sisi Fu, Yishu Liu 0001, Jiahui Pan 0003, Guangming Lu 0002, Zheng Zhang 0006 |
AAAI | 1 |
| 2024 | Enhancing DETRs for Small Object Detection via Multi-Scale Refinement and Query-Aided Mining
Sisi Fu, Xiaocheng Fang, Jieyi Cai, Huosheng Wen, Bingzhi Chen |
ACML | 7 |
| 2024 | Progressive Stepwise Diffusion Model with Dual Decoders for Semi-Supervised Medical Image SegmentationabstractSemi-supervised medical image segmentation tasks aim to harness the potential of vast amounts of unlabeled data using a limited amount of annotated data. Denoising Diffusion Probabilistic Models, which have achieved significant success in image generation, are gradually being explored for their potential in semantic image segmentation. However, their application in semi-supervised medical image segmentation is still in its early stages. Initially, due to the high randomness of diffusion models, the pseudo-labels generated during the early training phase may mislead the processing of unlabeled data. Additionally, the use of fixed-time steps for random sampling during training limits the ability of the model to learn effective denoising functions at an early stage. To address these issues, we propose an innovative framework named Progressive Stepwise Diffusion Network with Dual Decoders (PSDD) for semi-supervised medical image segmentation. This framework incorporates an additional normal decoder into the denoising diffusion encoder-decoder structure to provide more accurate labels and employs a Progressive Incremental Step strategy to gradually train the model for longer generation processes. Evaluated on two 2D colon polyp segmentation datasets and a 3D Left Atrium dataset, the experimental results demonstrate significant performance improvements over current advanced methods, thereby validating the effectiveness and potential of this framework in handling complex semi-supervised learning scenarios. Xiaolin Huang, Jingchun Lin, Bingzhi Chen, Guangming Lu 0002 |
BIBM | 5 |
| 2024 | Decoupled Self-Adaptive Distribution Regularization for Few-Shot Image ClassificationabstractThe feature dispersion, arising from the inherent constraints of data scarcity, has emerged as a prominent challenge in the domain of few-shot learning. In this paper, we propose a novel Self-adaptive Distribution Regularization (SADR) approach, which can adaptively bridge the semantic gaps across distribution patterns and boundaries for learning from limited labeled data. Specifically, the technical core of our SADR approach is to decouple the feature embeddings into two discrete spaces: the intra-class and inter-class distributions, leading to robust and discriminative feature representations in a self-adaptive manner. To achieve meticulous similarity measurements while mitigating redundant feature information, an innovative regularized Brownian Distance Covariance (R-BDC) metric is strategically designed to simultaneously explore both the joint and marginal distributions present among diverse input samples. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority of our SADR approach over state-of-the-art baselines. Bingzhi Chen, Haoming Zhou, Yishu Liu 0001, Guangming Lu 0002, Zheng Zhang 0006 |
ICASSP | 1 |
| 2024 | Medical Vision-Language Representation Learning with Cross-Modal Multi-Teacher Contrastive DistillationabstractMedical vision-language representation learning has garnered considerable attention owing to its applicability to extracting generic representations from the image and text modality. However, it still remains challenging to acquire a more comprehensive understanding of intra- and inter-modal semantic knowledge. In this paper, we propose a Cross-Modal Multi-Teacher Contrastive Distillation (CMCD) architecture, which aims to comprehensively learn medical vision-language representation in a unified multi-teacher framework. Specifically, a cross-modal knowledge distillation (CKD) module is designed to refine reconstructed semantics under an additional supervision signal generated by momentum teachers from the other modality, achieving more robust semantic interaction across modalities. To better alleviate the heterogeneity and semantic gaps, the multi-level contrastive learning (MCL) module is conceived to align features of both intra- and inter-modal via contrastive learning from multi-level perspectives. Extensive experiments on two medical downstream tasks, i.e., Med-VQA and Med-ITC, demonstrate that our CMCD consistently outperforms the state-of-the-art methods. Bingzhi Chen, Yishu Liu 0001, Jiahui Pan 0003, Meirong Ding |
ICASSP | 1 |
| 2024 | Rethinking Adversarial Robustness Distillation VIA Strength-Dependent Adaptive RegularizationabstractDespite the progress achieved by existing adversarial distillation (AD) approaches, most mainstream models suffer from inadequate adversarial robustness, due to the challenges of fixed attack strength and unreliable teacher guidance. In this paper, we propose a novel Strength-Dependent Adaptive Regularization (SDAR) paradigm to reinforce the function of adversarial distillation with strength-adaptive adversarial attack (SAA) and multi-dimensional knowledge distillation (MKD). Different from the traditional adversarial training (AT) methods, the proposed SAA scheme dynamically assigns an adaptive and efficient attack strength for each instance, which aims to facilitate smoother classification boundaries. By incorporating dynamic strength coefficients, a comprehensive MKD strategy is designed to fully explore the valuable context information and narrow distribution discrepancies across teacher-student domains. Particularly, our SDAR paradigm can seamlessly integrate with the current AD frameworks, further enhancing the adversarial robustness of deep learning models. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority of SDAR over state-of-the-art baselines. Bingzhi Chen, Shuobin Lin, Yishu Liu 0001, Zheng Zhang 0006, Guangming Lu 0002, Lewei He |
ICME | 1 |
| 2024 | Enhancing Few-Shot Classification without Forgetting Through Multi-level Contrastive ConstraintsabstractMost recent few-shot learning approaches are based on meta-learning with episodic training. However, prior studies encounter two crucial problems: (1) the presence of inductive bias, and (2) the occurrence of catastrophic forgetting. In this paper, we propose a novel Multi-Level Contrastive Constraints (MLCC) framework, that jointly integrates within-episode learning and across-episode learning into a unified interactive learning paradigm to solve these issues. Specifically, we employ a space-aware interaction modeling scheme to explore the correct inductive paradigms for each class between within-episode similarity/dis-similarity distributions. Additionally, with the aim of better utilizing former prior knowledge, a cross-stage distribution adaption strategy is designed to align the across-episode distributions from different time stages, thus reducing the semantic gap between existing and past prediction distribution. Extensive experiments on multiple few-shot datasets demonstrate the consistent superiority of MLCC approach over the existing state-of-the-art baselines. Bingzhi Chen, Haoming Zhou, Yishu Liu 0001, Jiahui Pan 0003, Guangming Lu 0002 |
ICME | 1 |
| 2024 | Ambiguity Consistency and Uncertainty Minimization for Semi-Supervised Medical Image SegmentationabstractCo-training and pseudo-supervision are two common strategies in semi-supervised medical image segmentation. However, co-training may lead to a ’resonance’ problem, and the effectiveness of generating pseudo-labels by setting thresholds may greatly depend on manual efforts. To address these issues, we propose an innovative framework for Ambiguity Consistency and Uncertainty Minimization (ACUM) in semi-supervised medical image segmentation. Specifically, ACUM comprises two main components: (1) Ambiguity Consistency Constraint (ACC), which encourages model differentiation and applies dynamic pixel-level consistency constraints through ambiguous areas between sub-networks; (2) Pixel Uncertainty Minimization (PUM), which generates high-confidence pseudo-labels by selecting labels with relatively low uncertainty based on the uncertainty maps of sub-networks. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority of our proposed ACUM approach over state-of-the-art techniques. Xiaolin Huang, Yujiang Yao, Bingzhi Chen |
ICME | 6 |
| 2024 | Deep Unfolding 3D Non-Local Transformer Network for Hyperspectral Snapshot Compressive ImagingabstractHyperspectral compressive imaging has shown remarkable advancements through the adoption of deep unfolding frameworks, which integrate the proximal mapping prior into the data fidelity term to formulate the reconstruction problem. However, existing technologies still face challenges in effectively capturing spatial-spectral features during the iterative deep prior learning stage, leading to unsatisfactory performance degradation. To address this issue, we propose a deep unfolding 3D non-local transformer (3DNLT) network for hyperspectral compressive imaging. A learnable half-quadratic splitting (HQS) algorithm is utilized to iteratively update the linear projection. Furthermore, a 3D non-local attention ushaped transformer is presented as the deep proximal mapping prior module to obtain the spatial-spectral long-range dependency features, leading to enhance the network’s ability to capture fine-grained hyperspectral and spatial details. Experimental results on both synthetic and real hyperspectral image reconstruction have demonstrated the superior performance of the 3DNLT network compared to state-of-the-art methods. Yongyong Chen, Bingzhi Chen, Yicong Zhou |
ICME | 4 |
| 2024 | Robust Visual Question Answering With Contrastive-Adversarial Consistency ConstraintsabstractVisual cues and question semantics contribute to final answer predictions from distinct perspectives. However, inherent language bias confounds the relationship between visual and question cues, leading to a misguided preference for question semantics. Different from the existing studies that focus on inter-class discrimination, this paper proposes a robust visual question answering framework with contrastive-adversarial consistency constraints (CACC) at both inter- and intra-instance levels. From a fine-grained instance-level perspective, our approach initially introduces an effective inter-instance contrastive constraint to perform adaptive bias rectification. To enhance intra-instance invariance and reduce information redundancy, we refine the concept of semantic structure relationships by constructing intra-instance adversarial constraints using the Hilbert-Schmidt Independence Criterion (HSIC) independence criterion. Benefitting from both inter- and intra-instance perspectives, our method can effectively alleviate these language biases, enhancing the overall robustness of the representation. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority of our CACC over state-of-the-art baselines. Meirong Ding, Yishu Liu 0001, Guangming Lu 0002, Bingzhi Chen |
ICME | 6 |
| 2024 | Enhancing Cross-Modal Retrieval via Visual-Textual Prompt Hashing
Bingzhi Chen, Zhongqi Wu, Yishu Liu 0001, Guangming Lu 0002, Zheng Zhang 0006 |
IJCAI | 1 |
| 2024 | Medical Cross-Modal Prompt Hashing with Robust Noisy Correspondence Learning
Yishu Liu 0001, Zhongqi Wu, Bingzhi Chen, Zheng Zhang 0006, Guangming Lu 0002 |
MICCAI (3) | 3 |
| 2024 | Stay Focused is All You Need for Adversarial Robustness
Bingzhi Chen, Ruihan Liu, Yishu Liu 0001, Xiaozhao Fang, Jiahui Pan 0003, Guangming Lu 0002, Zheng Zhang 0006 |
ACM Multimedia | 1 |
| 2024 | Partial Multi-label Learning Based On Near-Far Neighborhood Label Enhancement And Nonlinear Guidance
Na Han, Xiaozhao Fang, Bingzhi Chen, Jie Wen 0001 |
ACM Multimedia | 5 |
| 2024 | Combating Visual Question Answering Hallucinations via Robust Multi-Space Co-Debias Learning
Yishu Liu 0001, Huanjia Zhu, Yuncheng Jiang 0004, Zheng Zhang 0006, Bingzhi Chen |
ACM Multimedia | 7 |
| 2024 | A feature-enhanced hybrid attention network for traffic sign recognition in real scenesabstractAbstract Currently, traffic sign recognition techniques have been brought into the assistive driving of automobiles. However, small traffic sign recognition in real scenes is still a challenging task due to the class imbalance issue and the size limit of the traffic signs. To address the above issues, a feature‐enhanced hybrid attention network is proposed based on YOLOv5s for a small, fast, and accurate traffic sign detector. First, a series of online data augmentation strategies are designed in the preprocessing module for the model training. Second, the hybrid channel and spatial attention module CSAM are integrated into the backbone for a better feature extraction ability. Third, the channel attention module CAM is used in the detection head for a more efficient feature fusion ability. To validate the approach, extensive experiments are conducted based on the Tsinghua‐Tencent 100K dataset. It is found that the novel method achieves state‐of‐the‐art performance with only negligible increases in the model parameter and computational overhead. Specifically, the , parameters, and FLOPs are 85.8%, 7.13 M, and 16.1 G, respectively. Lewei He, Fucai Lan, Chuanzhe Zhou, Yaoguang Ye, Wencong Zhang, Bingzhi Chen, Jiahui Pan 0003 |
IET Image Process. | 6 |
| 2024 | Context-aware graph embedding with gate and attention for session-based recommendation
Junlong Chi, Peilin Hong, Guangming Lu 0002, David Zhang 0001, Bingzhi Chen |
Neurocomputing | 6 |
| 2024 | Deep Fuzzy Multiteacher Distillation Network for Medical Visual Question AnsweringabstractMedical visual question answering (medical VQA) is a critical cross-modal interaction task that garnered considerable attention in the medical domain. Several existing methods commonly leverage the vision-and-language pretraining paradigms to mitigate the limitation of small-scale data. Nevertheless, most of them still suffer from two challenges that remain for further research: 1) limited research focuses on distilling representation from a complete modality to guide the representation learning of masked data in other modalities. 2) Multimodal fusion based on self-attention mechanisms cannot effectively handle the inherent uncertainty and vagueness of information interaction across modalities. To mitigate these issues, in this article, we propose a novel deep fuzzy multiteacher distillation (DFMD) network for medical VQA, which can take advantage of fuzzy logic to model the uncertainties from vison-language representations across modalities in a multiteacher framework. Specifically, a multiteacher knowledge distillation module is conceived to assist in reconstructing the missing semantics under the supervision signal generated by teachers from the other complete modality, achieving more robust semantic interaction across modalities. Incorporating insights from the fuzzy logic theory, we propose a noise-robust encoder called FuzBERT that enables our DFMD model to reduce the imprecision and ambiguity in feature representation during the multimodal interaction process. To the best of our knowledge, our work isthe first attemptto combine the fuzzy logic theory with the transformer-based encoder to effectively learn multimodal representation for medical VQA. Experimental results on the VQA-RAD and SLAKE datasets consistently demonstrate the superiority of our proposed DFMD method over state-of-the-art baselines. Yishu Liu 0001, Bingzhi Chen, Shuihua Wang, Guangming Lu 0002, Zheng Zhang 0006 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2024 | Vague-Segment Technique: Automatic Computation of Tumor Stroma Ratio for Breast Cancer on Whole SlidesabstractThe calculation of Tumor Stroma Ratio (TSR) is a challenging medical issue that could improve predictions of neoadjuvant chemotherapy benefits and patient prognoses. Although several studies on breast cancer and deep learning methods have achieved promising results, the drawbacks that pixel-level semantic segmentation processes could not extract core tumor regions containing both tumor pixels and stroma pixels make it difficult to accurately calculate TSR. In this paper, we propose a Vague-Segment Technique (VST) consisting of a designed SwinV2UNet module and a modified Suzuki algorithm. Specifically, the SwinV2UNet identifies tumor pixels and generate pixel-level classification results, based on which the modified Suzuki algorithm extracts the contour of core tumor regions in terms of cosine angle. Through this way, VST obtains vaguely segmentation results of core tumor regions containing both tumor pixels and stroma pixels, where the TSR could be calculated by the formula of Intersection over Union (IOU). For the training and evaluation, we utilize the well-known The Cancer Genome Atlas (TCGA) database to create an annotated dataset, while 150 images with TSR annotations from real cases are also collected. The experimental results illustrate that the proposed VST could generate better tumor identification results compared with state-of-the-art methods, where the extracted core tumor regions lead to more consistencies of calculated TSR with senior experts compared to junior pathologists. The experimental results demonstrate the superiority of our proposed pipeline, which has promise for future clinical application. Xinsen Lian, Kunping Yang, Bingzhi Chen, Xiuhong Cai, Xinling Lu, Jinlin Chen, Ming Tian, Pengtao Lin |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Deep Learning-Based Segment Trend Removal Approach for Prognostics and Health Management Signals of Rail VehiclesabstractTo ensure the safety and reliability of rail vehicles, prognostics and health management (PHM) is crucial. However, PHM signals can often be disrupted by trend components, which fundamentally undermine the accuracy of damage identification and prediction. Removing trend components efficiently is a challenge due to the intermittent and nonlinear character of railway vehicle PHM signals. This study proposes a segmented trend removal approach (STRA) for PHM signals based on deep learning. To evaluate the effectiveness of trend removal, a drift ratio and amplitude ratio assessment metric is defined. An identification model based on convolutional neural networks is also proposed to detect start-stop features and determine overall trend points for signal segmentation. A segmented spline baseline method is presented to eliminate running segment trends, while a least squares approach (LSA) is employed to address stationary segment trends. To increase processing and training efficiency with various sampling frequencies, a frequency conversion technique is implemented. The STRA outperforms conventional methods such as LSA, spline baseline approach, and empirical mode decomposition in nonlinear trend reduction. Additionally, this method can be extended to eliminate the zero drift of intermittent acceleration signals in PHM and is similarly suitable for use in other PHM applications, including automobiles, vessels, and bridges. Shaoze Zhou, Pengfei Zhao 0005, Bingzhi Chen |
IEEE Trans. Reliab. | 4 |
| 2023 | Combating Medical Label Noise via Robust Semi-supervised Contrastive Learning
Bingzhi Chen, Zhanhao Ye, Yishu Liu 0001, Zheng Zhang 0006, Jiahui Pan 0003, Guangming Lu 0002 |
MICCAI (1) | 1 |
| 2023 | Deep Margin-Sensitive Representation Learning for Cross-Domain Facial Expression RecognitionabstractCross-domain Facial Expression Recognition (FER) aims to safely transfer the learned knowledge from labeled source data to unlabeled target data, which is challenging due to the subtle difference between various expressions and the large discrepancy between domains. Existing methods mainly focus on reducing the domain shift for transferable features but fail to learn discriminative representations for recognizing facial expression, which may result in negative transfer under cross-domain settings. To this end, we propose a novel Deep Margin-Sensitive Representation Learning (DMSRL) framework, which can extract multi-level discriminative features during sematic-aware domain adaptation. Specifically, we design a semantic metric learning module based on the category prior of source data and generated pseudo labels of target data, which can facilitate discriminative intra-domain representation learning and transferable inter-domain knowledge discovery by enlarging the category margin. Moreover, we develop a mutual information minimization module by simultaneously distilling the domain-invariant components and eliminating the domain-sensitive ones, which benefits discriminative transferable feature learning by generating accurate pseudo target labels. Furthermore, instead of only utilizing the global features, we formulate a multi-level feature extracting module to concurrently get the local ones, which contain detailed information to distinguish the small changes among different expressions. These modules are jointly utilized in our DMSRL in an end-to-end manner to ensure the positive transfer of source knowledge. Extensive experimental results on seven databases demonstrate that our DMSRL can achieve superior performance against state-of-the-art baselines. Yingjian Li 0001, Zheng Zhang 0006, Bingzhi Chen, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | Multi-Label Chest X-Ray Image Classification via Semantic Similarity Graph EmbeddingabstractAutomated multi-label chest X-ray (CXR) image classification has recently made significant progress in clinical diagnosis based on the advanced deep learning techniques. However, most existing methods mainly focus on analyzing locality visual cues from a single image but fail to leverage the underlying explicit correlations among different images for precise disease diagnosis. By contrast, an experienced radiologist expertizes in transferring knowledge from previous tasks to diagnose the present radiograph. To enable the machine like a radiologist, this paper proposes a novel Semantic Similarity Graph Embedding (SSGE) framework, which explicitly explores the semantic similarities among images to optimize the visual feature embedding for improving the performance of multi-label CXR images classification. Specifically, the proposed SSGE framework contains three main components: the image feature embedding (IFE) module, similarity graph construction (SGC) module, and semantic similarity learning (SSL) module. To realize interactive teaching and learning between visual and semantic information, the proposed SSGE framework is built on the “Teacher-Student” (semantic-visual) learning mechanism. With the guidance and supervision of the cross-image similarity graph generated by the SGC module, the SSL module leverages Graph Convolutional Network (GCN) to adaptively recalibrate the multi-image feature representations extracted from the IFE module, which guarantees their semantic consistency. Furthermore, we propose a novel re-weighting strategy to learn a more optimal semantic-similarity graph for the information propagation of the GCN layers. Extensive experiments on two benchmark datasets demonstrate the effectiveness of the proposed method in comparison with some state-of-the-art baselines. Bingzhi Chen, Zheng Zhang 0006, Yingjian Li 0001, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Self-Supervised Exclusive-Inclusive Interactive Learning for Multi-Label Facial Expression Recognition in the WildabstractFacial Expression Recognition (FER) is a long-standing but challenging research problem in computer vision. Existing approaches mainly focus on single-label emotional prediction, which cannot handle the complex multi-label FER task because of the coupling behavior of multiple emotions on a single facial image. To this end, in this paper, we propose a novel Self-supervised Exclusive-Inclusive Interactive Learning (SEIIL) method to facilitate discriminative multi-label FER in the wild, which can effectively handle the coupled multiple sentiments with limited unconstrained training data. Specifically, we construct an emotion disentangling module to capture the inclusive and exclusive characteristics of facial expressions, which can decouple the compound numerous emotions on an image. Moreover, an adaptively-weighted ensemble technique is conceived to aggregate category-level latent exclusive embeddings, and then a conditional adversarial interactive learning module is designed to fully leverage the complementary between the inclusive and formulated latent representations. Furthermore, to tackle the insufficient data for training, we introduce a self-supervised learning strategy to augment the amount and diversity of facial images, which can endow the model with advanced generalization ability. Under this strategy, the proposed two modules can be concurrently utilized in our SEIIL to jointly handle the coupled emotions and alleviate the overfitting problem. Extensive experimental results on six databases illustrate the superb performance of our method against state-of-the-art baselines. Yingjian Li 0001, Yingnan Gao, Bingzhi Chen, Zheng Zhang 0006, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Learning Informative and Discriminative Features for Facial Expression Recognition in the WildabstractThe informativeness and discriminativeness of features collaboratively ensure high-accuracy Facial Expression Recognition (FER) in the wild. Most of existing methods use the single-path deep convolutional neural network with softmax loss for basic FER, while they cannot deal with the challenging situations of the compound FER in the wild, because they fail to learn informative and discriminative features in a targeted manner. To this end, we present an Informative and Discriminative Feature Learning (IDFL) framework that consists of two key components: the Multi-Path Attention Convolutional Neural Network (MPACNN) and Balanced Separate loss (BS loss), for both basic and compound high-accuracy FER in the wild. Specifically, MPACNN leverages different paths to learn diverse features. These features are then adaptively fused into informative ones via an attention module, such that the model can adequately capture detailed information for both basic and compound FER. The BS loss maximizes the inter-class distance of features and minimizes the intra-class one. In this way, the features are discriminative enough for high-accuracy FER in the wild. Particularly, the BS loss is invoked as the objective function of MPACNN, so the model can learn informative and discriminative features at the same time, yielding better performance. Seven databases are utilized to evaluate the proposed method, and the results demonstrate that our method achieves state-of-the-art performance on both basic and compound expressions with good generalization ability. Moreover, our model contains fewer parameters and can be trained faster than other related models. Yingjian Li 0001, Yao Lu 0008, Bingzhi Chen, Zheng Zhang 0006, Jinxing Li 0003, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Semantic-Interactive Graph Convolutional Network for Multilabel Image RecognitionabstractMultilabel image recognition, a critically practical task in computer vision, aims to predict multiple objects present in each image. The existing studies mainly focus on conceptual visual cues but fail to reconcile the visual information with their semantic guidance. Intuitively, humans can not only associate extra topological concepts but also imagine other approximate scenes based on a semantic description. Inspired by such semantic-interactive capability, two different types of semantic priors, i.e., the concept correlations of the same scene and semantic similarities among different scenes, should be further explored for the recognition decisions. To efficiently interact with these semantic relationships, in this article, we propose a novel semantic-interactive graph convolutional network (SI-GCN), which can leverage the topological information learned from knowledge graphs to boost the performance of multilabel recognition. Specifically, the proposed SI-GCN framework consists of two different GCN-based branches in parallel, i.e., concept correlations learning (CCL) branch and semantic similarity learning (SSL) branch. Inputting the semantic-embedding vectors of all the concepts, the CCL branch maps the label co-occurrence graph into a set of interdependent concept classifiers. Recalibrating the image feature embedding with the standardized supervision of the semantic similarity graph, the SSL branch learns the semantically consistent in-batch visual representations. Finally, a well-established interactive learning scheme is formulated to concurrently optimize the obtained concept classifiers and the visual representation learning in an end-to-end manner. Extensive experiments on the MS-COCO and Pascal VOC 2007 & 2012 benchmarks demonstrate the superiorities of the proposed SI-GCN method compared to the state-of-the-art baselines. Bingzhi Chen, Zheng Zhang 0006, Yao Lu 0008, Fanglin Chen 0001, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2021 | Hierarchical Network Based on the Fusion of Static and Dynamic Features for Speech Emotion RecognitionabstractMany studies on automatic speech emotion recognition (SER) have been devoted to extracting meaningful emotional features for generating emotion-relevant representations. However, they generally ignore the complementary learning of static and dynamic features, leading to limited performances. In this paper, we propose a novel hierarchical network called HNSD that can efficiently integrate the static and dynamic features for SER. Specifically, the proposed HNSD framework consists of three different modules. To capture the discriminative features, an effective encoding module is firstly designed to simultaneously encode both static and dynamic features. By taking the obtained features as inputs, the Gated Multi-features Unit (GMU) is conducted to explicitly determine the emotional intermediate representations for frame-level features fusion, instead of directly fusing these acoustic features. In this way, the learned static and dynamic features can jointly and comprehensively generate the unified feature representations. Benefiting from a well-designed attention mechanism, the last classification module is applied to predict the emotional states at the utterance level. Extensive experiments on the IEMOCAP benchmark dataset demonstrate the superiority of our method in comparison with state-of-the-art baselines. Mi-Xiao Hou, Bingzhi Chen, Zheng Zhang 0006, Guangming Lu 0002 |
ICASSP | 3 |
| 2021 | JDMAN: Joint Discriminative and Mutual Adaptation Networks for Cross-Domain Facial Expression RecognitionabstractCross-domain Facial Expression Recognition (FER) is challenging due to the difficulty of concurrently handling the domain shift and semantic gap during domain adaptation. Existing methods mainly focus on reducing the domain discrepancy for transferable features but fail to decrease the semantic one, which may result in negative transfer. To this end, we propose Joint Discriminative and Mutual Adaptation Networks (JDMAN), which collaboratively bridge the domain shift and semantic gap by domain- and category-level co-adaptation based on mutual information and discriminative metric learning techniques. Specifically, we design a mutual information minimization module for domain-level adaptation, which narrows the domain shift by simultaneously distilling the domain-invariant components and eliminating the untransferable ones lying in different domains. Moreover, we propose a semantic metric learning module for category-level adaptation, which can close the semantic discrepancy during discriminative intra-domain representation learning and transferable inter-domain knowledge discovery. These two modules are jointly leveraged in our JDMAN to safely transfer the source knowledge to target data in an end-to-end manner. Extensive experimental results on six databases show that our method achieves state-of-the-art performance. The code of our JDMAN is available at https://github.com/YingjianLi/JDMAN. Yingjian Li 0001, Yingnan Gao, Bingzhi Chen, Zheng Zhang 0006, Lei Zhu 0002, Guangming Lu 0002 |
ACM Multimedia | 3 |
| 2021 | An Embarrassingly Simple Approach to Discrete Supervised HashingabstractPrior hashing works typically learn a projection function from high-dimensional visual feature space to low-dimensional latent space. However, such a projection function remains several crucial bottlenecks: 1) information loss and coding redundancy are inevitable; 2) the available information of semantic labels is not well-explored; 3) the learned latent embedding lacks explicit semantic meaning. To overcome these limitations, we propose a novel supervised Discrete Auto-Encoder Hashing (DAEH) framework, in which a linear auto-encoder can effectively project the semantic labels of images into a latent representation space. Instead of using the visual feature projection, the proposed DAEH framework skillfully explores the semantic information of supervised labels to refine the latent feature embedding and further optimizes hashing function. Meanwhile, we reformulate the objective and relax the discrete constraints for the binary optimization problem. Extensive experiments on Caltech-256, CIFAR-10, and MNIST datasets demonstrate that our method can outperform the state-of-the-art hashing baselines. Shuguang Zhao, Bingzhi Chen, Zheng Zhang 0006, Guangming Lu 0002 |
MMAsia | 2 |
| 2021 | Multimodal Emotion Recognition With Temporal and Semantic ConsistencyabstractAutomated multimodal emotion recognition has become an emerging but challenging research topic in the fields of affective learning and sentiment analysis. The existing works mainly focus on developing multimodal fusion strategies to incorporate different emotion-related features. However, they fail to explore the inherent contextual consistency to reconcile the emotional information across modalities. In this paper, we propose a novel Time and Semantic Interaction Network (TSIN), which concurrently incorporates the advantages of temporal and semantic consistency into the multimodal emotion recognition task. Specifically, a well-designed Speech and Text Embedding (STE) module is devoted to formulating the initial embedding spaces by respectively building the modality-specific representations of speech and text. Instead of separately learning or directly fusing the acoustic and textual features, we propose a well-defined Time and Semantic Interaction (TSI) module to conduct the emotional parsing and sentiment refining by performing the fine-grained temporal alignment and cross-modal semantic interaction. Benefitting from temporal and semantic consistency constraints, both speech-text embeddings can be interactively optimized and fine-tuned in the learning process. In this way, the learnt acoustics and textual features can jointly and efficiently predict the final emotional state. Extensive experiments on the IEMOCAP dataset demonstrate the superiorities of our TSIN framework in comparison with state-of-the-art baselines. Bingzhi Chen, Mi-Xiao Hou, Zheng Zhang 0006, Guangming Lu 0002, David Zhang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | Deep Active Context Estimation for Automated COVID-19 DiagnosisabstractMany studies on automated COVID-19 diagnosis have advanced rapidly with the increasing availability of large-scale CT annotated datasets. Inevitably, there are still a large number of unlabeled CT slices in the existing data sources since it requires considerable consuming labor efforts. Notably, cinical experience indicates that the neighboring CT slices may present similar symptoms and signs. Inspired by such wisdom, we propose DACE, a novel CNN-based deep active context estimation framework, which leverages the unlabeled neighbors to progressively learn more robust feature representations and generate a well-performed classifier for COVID-19 diagnosis. Specifically, the backbone of the proposed DACE framework is constructed by a well-designed Long-Short Hierarchical Attention Network (LSHAN), which effectively incorporates two complementary attention mechanisms, i.e., short-range channel interactions (SCI) module and long-range spatial dependencies (LSD) module, to learn the most discriminative features from CT slices. To make full use of such available data, we design an efficient context estimation criterion to carefully assign the additional labels to these neighbors. Benefiting from two complementary types of informative annotations from -nearest neighbors, i.e., the majority of high-confidence samples with pseudo labels and the minority of low-confidence samples with hand-annotated labels, the proposed LSHAN can be fine-tuned and optimized in an incremental learning manner. Extensive experiments on the Clean-CC-CCII dataset demonstrate the superior performance of our method compared with the state-of-the-art baselines. Bingzhi Chen, Yishu Liu 0001, Zheng Zhang 0006, Yingjian Li 0001, Zhao Zhang 0001, Guangming Lu 0002, Hongbing Yu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2020 | Two-stream collaborative network for multi-label chest X-ray Image classification with lung segmentation
Bingzhi Chen, Zheng Zhang 0006, Jianyong Lin, Yi Chen 0023, Guangming Lu 0002 |
Pattern Recognit. Lett. | 1 |
| 2020 | Lesion Location Attention Guided Network for Multi-Label Thoracic Disease Classification in Chest X-RaysabstractTraditional clinical experiences have shown the benefit of lesion location attention for improving clinical diagnosis tasks. Inspired by this point of interest, in this paper we propose a novel lesion location attention guided network named LLAGnet to focus on the discriminative features from lesion locations for multi-label thoracic disease classification in chest X-rays (CXRs). By revealing the equivalence of the region-level attention (RLA) and channel-level attention (CLA), we find that the RLA is available as priors for object localization while the CLA implicitly provides high weights to the attractive channels, which both enable lesion location attention excitation. To integrate the advantages from both mechanisms, the proposed LLAGnet is structured with two corresponding attention modules, i.e., the RLA and CLA modules. Specifically, the RLA module consists of the global and local branches. And the weakly supervised attention mechanism embedded in the global branch can obtain visual regions of lesion locations by back-propagating gradients. Then the optimal attention region is amplified and applied to the local branch to provide more fine-grained features for the image classification. Finally, the CLA module adaptively enhances the weights of channel-wise features from the lesion locations by modeling interdependencies among channels. Extensive experiments on the ChestX-ray14 dataset clearly substantiate the effectiveness of LLAGnet as compared with the state-of-the-art baselines. Bingzhi Chen, Jinxing Li 0003, Guangming Lu 0002, David Zhang 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2020 | Label Co-Occurrence Learning With Graph Convolutional Networks for Multi-Label Chest X-Ray Image ClassificationabstractExisting multi-label medical image learning tasks generally contain rich relationship information among pathologies such as label co-occurrence and interdependency, which is of great importance for assisting in clinical diagnosis and can be represented as the graph-structured data. However, most state-of-the-art works only focus on regression from the input to the binary labels, failing to make full use of such valuable graph-structured information due to the complexity of graph data. In this paper, we propose a novel label co-occurrence learning framework based on Graph Convolution Networks (GCNs) to explicitly explore the dependencies between pathologies for the multi-label chest X-ray (CXR) image classification task, which we term the "CheXGCN". Specifically, the proposed CheXGCN consists of two modules, i.e., the image feature embedding (IFE) module and label co-occurrence learning (LCL) module. Thanks to the LCL model, the relationship between pathologies is generalized into a set of classifier scores by introducing the word embedding of pathologies and multi-layer graph information propagation. During end-to-end training, it can be flexibly integrated into the IFE module and then adaptively recalibrate multi-label outputs with these scores. Extensive experiments on the ChestX-Ray14 and CheXpert datasets have demonstrated the effectiveness of CheXGCN as compared with the state-of-the-art baselines. Bingzhi Chen, Jinxing Li 0003, Guangming Lu 0002, Hongbing Yu, David Zhang 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2019 | Mask-Most Net: Mask Approximation Based Multi-oriented Scene Text Detection NetworkabstractIn this paper, a novel multi-task cascade framework, which jointly takes the detection and the segmentation into account, is presented for the scene text detection. To address the issue of multi-oriented scene text detection, we propose an instance-level mask approximation method through the auxiliary regression task on center and corner points. Specifically, the text instance in the image is first coarsely detected, followed by a contextual module which can capture more accurate instances. To cope with the scale variation existing in these detected instances, a combination of high-level semantic and low-level features is further exploited, achieving more robust and better performance. A series of experiments conducted on different benchmark datasets demonstrate the effectiveness of the proposed method. Xiaobao Guo, Jinxing Li 0003, Bingzhi Chen, Guangming Lu 0002 |
ICME | 3 |
| 2019 | Multi-label Chest X-Ray Image Classification via Label Co-occurrence Learning
Bingzhi Chen, Yao Lu 0008, Guangming Lu 0002 |
PRCV (2) | 1 |