Longzhen Yang

dblp:321/3428 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0002-5791-145XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A General Framework for Efficient Medical Image Analysis via Shared Attention Vision Transformer
abstract
Vision Transformers (ViTs) demonstrate significant promise in medical image analysis but face two critical challenges: 1) their limited ability to capture local features in data-scarce scenarios, leading to data inefficiency, and 2) their high computational and storage demands of the full fine-tuning process in transfer learning, resulting in parameter inefficiency. To achieve efficient and accurate medical image analysis, we propose Shared Attention Vision Transformer (SAViT) that comprises three innovative modules: i) Shared Prior Attention (SPA) that enhances data efficiency by innovatively employing a visual prompt to sequentially share consistent attention weights across local image regions, thereby enabling the learning of translational invariance to capture locality; ii) MixPool that preserves global modeling ability by aggregating local features after SPA through a multi-pooling mechanism, thus effectively facilitating long-range dependency across local image regions; and iii) Low-rank Multi-head Self-Attention (Lr-MSA) that improves parameter efficiency by using low-rank weights of multi-head self-attention, hence reducing computational complexity while maintaining accuracy in medical image analysis. SAViT demonstrates strong generalization across multiple medical imaging modalities, including retinopathy, dermoscopy, and radiography. Extensive experiments are conducted. The results indicate its high data efficiency and outstanding performance in comparison with more than 20 medical-specific and ViT-based models when all of them are trained from scratch. It excels in parameter-efficient tuning by surpassing 17 models across 6 datasets in transfer learning, with only ${0}.{17}$ M/ ${0}.{23}$ M trainable parameters on ViT-B/SwinViT-B backbones requiring ${86}.{60}$ M/ ${88}.{00}$ M parameters. Source code can be found at: https://github.com/LYH-hh/SAViT.
Ying Wen 0003, Longzhen Yang, Lianghua He, MengChu Zhou
IEEE Trans. Medical Imaging3
2025 AFiRe: Anatomy-Driven Self-Supervised Learning for Fine-Grained Representation in Radiographic Images
abstract
Current self-supervised methods, such as contrastive learning, predominantly focus on global discrimination, neglecting the critical fine-grained anatomical details required for accurate radiographic analysis. To address this challenge, we propose the Anatomy-driven self-supervised framework for enhancing Fine-grained Representation in radiographic image analysis (AFiRe). The core idea of AFiRe is to align the anatomical consistency with the unique token-processing characteristics of Vision Transformer. Specifically, AFiRe synergistically performs two self-supervised schemes: (i) Token-wise anatomy-guided contrastive learning, which aligns image tokens based on structural and categorical consistency to enhance fine-grained spatial-anatomical discrimination; (ii) Pixel-level anomaly-removal restoration, which particularly focuses on local anomalies, thereby refining the learned discrimination with detailed geometrical information. Additionally, we propose the Synthetic Lesion Mask to enhance anatomical diversity while preserving intra-consistency, which is typically corrupted by traditional data augmentations, such as Cropping and Affine transformations. Experimental results show that AFiRe: (i) provides robust anatomical discrimination, achieving more cohesive feature clusters compared to state-of-the-art contrastive learning methods; (ii) demonstrates superior generalization, surpassing 7 radiography-specific self-supervised methods in multi-label classification tasks with limited labeling; and (iii) integrates fine-grained information, enabling precise anomaly detection using only image-level annotations.
Lianghua He, Ying Wen 0003, Longzhen Yang, Hongzhou Chen
AAAI4
2025 CoSMIC: Continual Self-Supervised Learning for Multi-Domain Medical Imaging Via Conditional Mutual Information Maximization
Ying Wen 0003, Longzhen Yang, Lianghua He, Heng Tao Shen
ICCV3
2025 RadLAS: A Foundation Model for Interpretable Radiography Image Analysis with Lesion-Aware Self-Supervised Pre-training
abstract
Medical Foundation Models (MFMs) are revolutionizing radiography image analysis with scalable and generalized diagnostic capabilities. However, their effectiveness in real-world clinical practice is limited due to insufficient interpretability. To address this limitation, we propose RadLAS, a novel MFM for interpretable Radiographic image analysis by introducing Lesion-Aware Self-supervised pre-training. Unlike conventional MFMs that rely on post-hoc explanations, RadLAS innovates by directly emulating human diagnostic reasoning to first grounding lesion evidence and then making decisions accordingly. Specifically, RadLAS introduces two self-supervised tasks: (I) Lesion-grounded Reconstruction, which learns structured anatomical representations by restoring lesion-aware image patches into their healthy counterparts, thereby facilitating pixel-level grounding of lesion evidence via input-normal contrast. (II) Lesion-discrimination Contrastive Learning, which enhances lesion-aware pattern in representations by explicitly decoupling grounded lesion evidence as clinical cues and aligning them with global semantics, thereby enabling direct lesion-oriented diagnosis while preserving global context. RadLAS demonstrates excellent performance across diverse downstream radiographic datasets, offering verifiable explanations by deriving specific diagnoses (Task II) based on grounded lesion evidence (Task I), while preserving generalized representations essential for high diagnostic accuracy. Extensive experiments demonstrate that RadLAS (i) achieves superior interpretability with highly correlated lesion prediction and localization, surpassing 11 interpretable medical models; (ii) delivers scalable representation learning, outperforming 14 SOTA supervised and self-supervised MFMs.
Ying Wen 0003, Longzhen Yang, Lianghua He, Heng Tao Shen
ACM Multimedia3
2025 Self-Supervised Anatomical Consistency Learning for Vision-Grounded Medical Report Generation
abstract
Vision-grounded medical report generation aims to produce clinically accurate descriptions of medical images, anchored in explicit visual evidence to improve interpretability and facilitate integration into clinical workflows. However, existing methods often rely on separately trained detection modules that require extensive expert annotations, introducing high labeling costs and limiting generalizability due to pathology distribution bias across datasets. To address these challenges, we propose Self-Supervised Anatomical Consistency Learning (SS-ACL)-a novel and annotation-free framework that aligns generated reports with corresponding anatomical regions using simple textual prompts. SS-ACL constructs a hierarchical anatomical graph inspired by the invariant top-down inclusion structure of human anatomy, organizing entities by spatial location. It recursively reconstructs fine-grained anatomical regions to enforce intra-sample spatial alignment, inherently guiding attention maps toward visually relevant areas prompted by text. To further enhance inter-sample semantic alignment for abnormality recognition, SS-ACL introduces a region-level contrastive learning based on anatomical consistency. These aligned embeddings serve as priors for report generation, enabling attention maps to provide interpretable visual evidence. Extensive experiments demonstrate that SS-ACL, without relying on expert annotations, (i) generates accurate and visually grounded reports-outperforming state-of-the-art methods by 10% in lexical accuracy and 25% in clinical efficacy, and (ii) achieves competitive performance on various downstream visual tasks, surpassing current leading visual foundation models by 8% in zero-shot visual grounding. Our code is available at https://github.com/kaelsunkiller/ssacl.
Longzhen Yang, Zhangkai Ni, Ying Wen 0003, Lianghua He, Heng Tao Shen
ACM Multimedia1
2025 FeaInfNet: Diagnosis of Medical Images With Feature-Driven Inference and Visual Explanations
abstract
Interpretable deep-learning models have received widespread attention in the field of image recognition. However, owing to the coexistence of medical-image categories and the challenge of identifying subtle decision-making regions, many proposed interpretable deep-learning models suffer from insufficient accuracy and interpretability in diagnosing images of medical diseases. Therefore, this study proposed a feature-driven inference network (FeaInfNet) that incorporates a feature-based network reasoning structure. Specifically, local feature masks (LFM) were developed to extract feature vectors, thereby providing global information for these vectors and enhancing the expressive ability of FeaInfNet. Second, FeaInfNet compares the similarity of the feature vector corresponding to each subregion image patch with the disease and normal prototype templates that may appear in the region. It then combines the comparison of each subregion when making the final diagnosis. This strategy simulates the diagnosis process of doctors, making the model interpretable during the reasoning process, while avoiding misleading results caused by the participation of normal areas during reasoning. Finally, we proposed adaptive dynamic masks (Adaptive-DM) to interpret feature vectors and prototypes into human-understandable image patches to provide an accurate visual interpretation. Extensive experiments on multiple publicly available medical datasets, including RSNA, iChallenge-PM, COVID-19, ChinaCXRSet, MontgomerySet, and CBIS-DDSM, demonstrated that our method achieves state-of-the-art classification accuracy and interpretability compared with baseline methods in the diagnosis of medical images. Additional ablation studies were performed to verify the effectiveness of each component.
Yitao Peng, Lianghua He, Die Hu 0002, Longzhen Yang, Shaohua Shang
IEEE J. Biomed. Health Informatics5
2025 Variational Transformer: A Framework Beyond the Tradeoff Between Accuracy and Diversity for Image Captioning
abstract
Accuracy and diversity represent two critical quantifiable performance metrics in the generation of natural and semantically accurate captions. While efforts are made to enhance one of them, the other suffers due to the inherent conflicting and complex relationship between them. In this study, we demonstrate that the suboptimal accuracy levels derived from human annotations are unsuitable for machine-generated captions. To boost diversity while maintaining high accuracy, we propose an innovative variational transformer (VaT) framework. By integrating "invisible information prior (IIP)" and "auto-selectable Gaussian mixture model (AGMM)," we enable its encoder to learn precise linguistic information and object relationships in various scenes, thus ensuring high accuracy. By incorporating the "range-median reward (RMR)" baseline into it, we preserve a wider range of candidates with higher rewards during the reinforcement-learning-based training process, thereby guaranteeing outstanding diversity. Experimental results indicate that our method achieves simultaneous improvements in accuracy and diversity by up to 1.1% and 4.8%, respectively, over the state-of-the-art. Furthermore, our approach demonstrates its performance that is the closest to human annotations in semantic retrieval, with its score of 50.3 versus the human score of 50.6. Thus, the method can be readily put into industrial use.
Longzhen Yang, Lianghua He, Die Hu 0002, Yitao Peng, Hongzhou Chen, MengChu Zhou
IEEE Trans. Neural Networks Learn. Syst.1
2024 Hierarchical Dynamic Masks for Visual Explanation of Neural Networks
abstract
Despite the remarkable accomplishments of deep neural networks in computer vision tasks, the inherent opacity of their operations remains a pressing concern. Attribution methods generating visual explanatory maps representing the importance of image pixels for model classification are popular for explaining neural network decisions. However, the small and diverse decision regions in fine-grained or medical images limit the precision and comprehensiveness of the existing attribution methods when explaining decisions made for such a data type. This paper introduces a novel attribution method called hierarchical dynamic masks (HDM) to overcome these concerns to generate saliency maps with high recognition reliability and localization capability. Specifically, we suggest dynamic masks (DM), which enable multiple small-sized benchmark mask vectors to learn the image's critical information roughly through an optimization method. The benchmark mask vectors guide the learning of the large-sized combination mask vectors so that their overlay mask accurately learns detailed pixel importance information. Additionally, we construct the HDM by hierarchically concatenating DM modules. These DM modules search and combine the regions of interest in the remaining neural network classification decisions within the masked image in a learning-based way. Since HDM forces DM to perform importance analysis in different areas, it makes the fused saliency map more comprehensive. The experiments reveal that the proposed method outperforms existing approaches significantly regarding recognition credibility and positioning ability when qualitatively and quantitatively tested on CUB-200-2011 and iChallenge-PM datasets.
Yitao Peng, Lianghua He, Die Hu 0002, Longzhen Yang, Shaohua Shang
IEEE Trans. Multim.5
2024 Decoupling Deep Learning for Enhanced Image Recognition Interpretability
abstract
The quest for enhancing the interpretability of neural networks has become a prominent focus in recent research endeavors. Prototype-based neural networks have emerged as a promising avenue for imbuing models with interpretability by gauging the similarity between image components and category prototypes to inform decision-making. However, these networks face challenges as they share similarity activations during both the inference and explanation processes, creating a tradeoff between accuracy and interpretability. To address this issue and ensure that a network achieves high accuracy and robust interpretability in the classification process, this article introduces a groundbreaking prototype-based neural network termed the “Decoupling Prototypical Network” (DProtoNet). This novel architecture comprises encoder, inference, and interpretation modules. In the encoder module, we introduce decoupling feature masks to facilitate the generation of feature vectors and prototypes, enhancing the generalization capabilities of the model. The inference module leverages these feature vectors and prototypes to make predictions based on similarity comparisons, thereby preserving an interpretable inference structure. Meanwhile, the interpretation module advances the field by presenting a novel approach: a “multiple dynamic masks decoder” that replaces conventional upsampling similarity activations. This decoder operates by perturbing images with mask vectors of varying sizes and learning saliency maps through consistent activation. This methodology offers a precise and innovative means of interpreting prototype-based networks. DProtoNet effectively separates the inference and explanation components within prototype-based networks. By eliminating the constraints imposed by shared similarity activations during the inference and explanation phases, our approach concurrently elevates accuracy and interpretability. Experimental evaluations on diverse public natural datasets, including CUB-200-2011, Stanford Cars, and medical datasets like RSNA and iChallenge-PM, corroborate the substantial enhancements achieved by our method compared to previous state-of-the-art approaches. Furthermore, ablation studies are conducted to provide additional evidence of the effectiveness of our proposed components.
Yitao Peng, Lianghua He, Die Hu 0002, Longzhen Yang, Shaohua Shang
ACM Trans. Multim. Comput. Commun. Appl.5