Yang Zhao 0019

dblp:50/2082-19 · DBLP profile ↗
← Back
26ranked-venue papers
5as first author
23since 2021 · last 2026
0000-0001-5252-658XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 4 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DOEI: Dual optimization of embedding information for attention-enhanced class activation maps
Zeyu Zhang 0006, Huazhang Wang, Shimin Wen, Daji Ergu, Ying Cai 0002, Yang Zhao 0019
Neurocomputing10
2026 Advancing federated domain generalization in ophthalmology: Vision enhancement and consistency assurance for multicenter fundus image segmentation
Yang Zhao 0019, Xianxun Zhu, Jun Wang 0121, Yan Liu 0052
Pattern Recognit.3
2025 ProjectedEx: Enhancing Generation in Explainable AI for Prostate Cancer
abstract
Prostate cancer, a growing global health concern, necessitates precise diagnostic tools, with Magnetic Resonance Imaging (MRI) offering high-resolution soft tissue imaging that significantly enhances diagnostic accuracy. Recent advancements in explainable AI and representation learning have significantly improved prostate cancer diagnosis by enabling automated and precise lesion classification. However, existing explainable AI methods, particularly those based on frameworks like generative adversarial networks (GANs), are predominantly developed for natural image generation, and their application to medical imaging often leads to suboptimal performance due to the unique characteristics and complexity of medical image. To address these challenges, our paper introduces three key contributions. First, we propose ProjectedEx, a generative framework that provides interpretable, multi-attribute explanations, effectively linking medical image features to classifier decisions. Second, we enhance the encoder module by incorporating feature pyramids, which enables multiscale feedback to refine the latent space and improves the quality of generated explanations. Additionally, we conduct comprehensive experiments on both the generator and classifier, demonstrating the clinical relevance and effectiveness of ProjectedEx in enhancing interpretability and supporting the adoption of AI in medical settings. Code will be released at https://github.com/Richardqiyi/ProjectedEx.
Xuyin Qi, Zeyu Zhang 0006, Aaron Berliano Handoko, Huazhan Zheng, Mingxi Chen, Ta Duc Huy, Vu Minh Hieu Phan, Linqi Cheng, Zhibin Liao, Yang Zhao 0019, Minh-Son To
CBMS13
2025 PedDet: Adaptive Spectral Optimization for Multimodal Pedestrian Detection
abstract
Pedestrian detection in intelligent transportation systems has made significant progress but faces two critical challenges: (1) insufficient fusion of complementary information between visible and infrared spectra, particularly in complex scenarios, and (2) sensitivity to illumination changes, such as low-light or overexposed conditions, leading to degraded performance. To address these issues, we propose PedDet, an adaptive spectral optimization complementarity framework which specifically enhanced and optimized for multispectral pedestrian detection. PedDet introduces the Multi-scale Spectral Feature Perception Module (MSFPM) to adaptively fuse visible and infrared features, enhancing robustness and flexibility in feature extraction. Additionally, the Illumination Robustness Feature Decoupling Module (IRFDM) improves detection stability under varying lighting by decoupling pedestrian and background features. We further design a contrastive alignment to enhance intermodal feature discrimination. Experiments on LLVIP and MSDS datasets demonstrate that PedDet achieves state-of-the-art performance, improving the mAP by 6.6 % with superior detection accuracy even in low-light conditions, marking a significant step forward for road safety.
Zeyu Zhang 0006, Wenxin Zhang 0005, Zirui Song, Xiuying Chen, Yang Zhao 0019
ECAI9
2025 MSDet: Receptive Field Enhanced Multiscale Detection for Tiny Pulmonary Nodule
abstract
Pulmonary nodules are critical for early lung cancer diagnosis, but traditional CT imaging methods suffer from low detection rates and poor localization. Small nodule detection is challenging due to subtle differences in density and issues like occlusion. Existing methods such as FPN, with its fixed feature fusion and limited receptive field, struggle to effectively overcome these issues. To address these challenges, our paper proposed three key contributions: Firstly, we proposed MSDet, a multiscale attention and receptive field network for detecting tiny pulmonary nodules. Secondly, we proposed the extended receptive domain (ERD) strategy to capture richer contextual information and reduce false positives caused by nodule occlusion. We also proposed the position channel attention mechanism (PCAM) to optimize feature learning and reduce multiscale detection errors, and designed the tiny object detection block (TODB) to enhance the detection of tiny nodules. Experiments on the LUNA16 dataset show an 8.8% improvement in mAP over YOLOv8, achieving state-of-the-art performance. The code is available at https://github.com/CaiGuoHui123/MSDet.
Guohui Cai, Ruicheng Zhang, Hongyang He, Zeyu Zhang 0006, Daji Ergu, Yuanzhouhan Cao, Jinman Zhao, Binbin Hu, Zhibin Liao, Yang Zhao 0019, Ying Cai 0002
ICME10
2025 SS-MPP: Semi-Supervised Shape-Aware Medical Image Segmentation Based on Multi-Scale Pixel-Wise Prototype
abstract
Semi-supervised methods, which efficiently leverage a small amount of labeled data, have been widely used in the field of medical imaging. This paper proposes a multi-level prototype-based morphological perception model that leverages the consistency of pixel intensities in medical images to enhance overall and edge segmentation. Deep prototypes capture global morphology, while shallow prototypes focus on edge details. The prototype loss encourages the encoder to separate foreground and background features. This multi-scale pixel-level prototype consistency is utilized to supervise the model’s predictions on unlabeled data, enabling better utilization of the unlabeled data. Additionally, we observed that this consistency constraint may cause the encoder to lose detailed features. To address this, we designed a sparse multi-scale attention module that accelerates the recovery of detailed information by fusing global and local information from other layers. We conducted comparative validation against other semi-supervised methods on well-known public medical segmentation datasets.
Kanqi Wang, Haoyun Wang, Yang Zhao 0019
ICME4
2025 MedConv: Convolutions Beat Transformers on Long-Tailed Bone Density Prediction
abstract
Bone density prediction via CT scans to estimate T-scores is crucial, providing a more precise assessment of bone health compared to traditional methods like X-ray bone density tests, which lack spatial resolution and the ability to detect localized changes. However, CT-based prediction faces two major challenges: the high computational complexity of transformer-based architectures, which limits their deployment in portable and clinical settings, and the imbalanced, long-tailed distribution of real-world hospital data that skews predictions. To address these issues, we introduce MedConv, a convolutional model for bone density prediction that outperforms transformer models with lower computational demands. We also adapt Bal-CE loss and post-hoc logit adjustment to improve class balance. Extensive experiments on our AustinSpine dataset shows that our approach achieves up to 21% improvement in accuracy and 20% in ROC AUC over previous state-of-the-art methods. Code will be available at https://github.com/Richardqiyi/MedConv.
Xuyin Qi, C. Zeyu Zhang, Huazhan Zheng, Mingxi Chen, Numan Kutaiba, Ruth Lim, Cherie Chiang, Zi En Tham, Xuan Ren, Wenxin Zhang 0005, Wenbing Lv, Guangzhen Yao, Renda Han, Kangsheng Wang, Hongtao Mao, Yu Li 0047, Zhibin Liao, Yang Zhao 0019, Minh-Son To
IJCNN21
2025 Gate-ViT: Gated Vision Transformer for Fine-Grained Visual Classification
Kanqi Wang, Peiyu Wang, Qin Zhang 0011, Yang Zhao 0019, Xiaohan Yu 0001
PAKDD (3)5
2025 Medical artificial intelligence for early detection of lung cancer: A survey
Guohui Cai, Ying Cai 0002, Zeyu Zhang 0006, Yuanzhouhan Cao, Daji Ergu, Zhibin Liao, Yang Zhao 0019
Eng. Appl. Artif. Intell.8
2025 CIT: Rethinking class-incremental semantic segmentation with a Class Independent Transformation
abstract
Class-incremental semantic segmentation (CSS) requires that a model learn to segment new classes without forgetting how to segment previous ones: this is typically achieved by distilling the current knowledge and incorporating the latest data. However, bypassing iterative distillation by directly transferring outputs of initial classes to the current learning task is not supported in existing class-specific CSS methods. Via Softmax, they enforce dependency between classes and adjust the output distribution at each learning step, resulting in a large probability distribution gap between initial and current tasks. We introduce a simple, yet effective Class Independent Transformation (CIT) that converts the outputs of existing semantic segmentation models into class-independent forms with negligible cost or performance loss. By utilizing class-independent predictions facilitated by CIT, we establish an accumulative distillation framework, ensuring equitable incorporation of all class information. We conduct extensive experiments on various segmentation architectures, including DeepLabV3, Mask2Former, and SegViTv2. Results from these experiments show minimal task forgetting across different datasets, with less than 5% for ADE20K in the most challenging 11 task configurations and less than 1% across all configurations for the PASCAL VOC 2012 dataset. • Softmax interdependency causes incremental forgetting in continual learning. • We introduce a class-independent transformation (CIT) to reduce forgetting. • CIT reformulates segmentation as class-agnostic, enhancing CSS training pipelines. • Our method significantly reduces forgetting on ADE20K compared to CSS baselines. • CIT achieves near-zero forgetting ( ≤ 1%) in Pascal-VOC 2012 settings.
Jinchao Ge, Bowen Zhang 0009, Akide Liu, Vu Minh Hieu Phan, Qi Chen 0014, Yangyang Shu, Yang Zhao 0019
Pattern Recognit.7
2025 Self-Supervised Lie Algebra Representation Learning via Optimal Canonical Metric
abstract
Learning discriminative representation with limited training samples is emerging as an important yet challenging visual categorization task. While prior work has shown that incorporating self-supervised learning can improve performance, we found that the direct use of canonical metric in a Lie group is theoretically incorrect. In this article, we prove that a valid optimization measurement should be a canonical metric on Lie algebra. Based on the theoretical finding, this article introduces a novel self-supervised Lie algebra network (SLA-Net) representation learning framework. Via minimizing canonical metric distance between target and predicted Lie algebra representation within a computationally convenient vector space, SLA-Net avoids computing nontrivial geodesic (locally length-minimizing curve) metric on a manifold (curved space). By simultaneously optimizing a single set of parameters shared by self-supervised learning and supervised classification, the proposed SLA-Net gains improved generalization capability. Comprehensive evaluation results on eight public datasets show the effectiveness of SLA-Net for visual categorization with limited samples.
Xiaohan Yu 0001, Zicheng Pan, Yang Zhao 0019, Yongsheng Gao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 MedDet: Generative Adversarial Distillation for Efficient Cervical Disc Herniation Detection
abstract
Cervical disc herniation (CDH) is a prevalent musculoskeletal disorder that significantly impacts health and requires labor-intensive analysis from experts. Despite advancements in automated detection of medical imaging, two significant challenges hinder the real-world application of these methods. First, the computational complexity and resource demands present a significant gap for real-time application. Second, noise in MRI reduces the effectiveness of existing methods by distorting feature extraction. To address these challenges, we propose three key contributions: Firstly, we introduced MedDet, which leverages the multi-teacher single-student knowledge distillation for model compression and efficiency, meanwhile integrating generative adversarial training to enhance performance. Additionally, we customize the second-order nmODE to improve the model’s resistance to noise in MRI. Lastly, we conducted comprehensive experiments on the CDH-1848 dataset, achieving up to a 5% improvement in mAP compared to previous methods. Our approach also delivers over 5 times faster inference speed, with approximately 67.8% reduction in parameters and 36.9% reduction in FLOPs compared to the teacher model. These advancements significantly enhance the performance and efficiency of automated CDH detection, demonstrating promising potential for future application in clinical practice.
Zeyu Zhang 0006, Nengmin Yi, Shengbo Tan, Ying Cai 0002, Yi Yang 0001, Lei Xu 0001, Qingtai Li, Daji Ergu, Yang Zhao 0019
BIBM10
2024 Motion Avatar: Generate Human and Animal Avatars with Arbitrary Motion
Zeyu Zhang 0006, Biao Wu 0006, Shiya Huang, Wenbo Zhang 0009, Ling Chen 0006, Yang Zhao 0019
BMVC10
2024 Occluded Person Retrieval with Hierarchical Feature Optimization
abstract
Occluded person retrieval aims to match images from occluded pedestrians. It pushes forward progress of person retrieval towards applications in real-world scenarios, thus attracting increasing attention in recent years. A key challenge is to learn discriminative representation within limited informative regions due to obstacle or pedestrian occlusion. To that end, we propose a hierarchical feature optimization model (HFO) that jointly optimizes image-level, object-level and part-level features for improved occluded person retrieval. A hierarchical discriminative feature grouping (HDFG) module is developed to generate hierarchical object/part masks for comprehensive feature extraction. Via learning a set of part prototypes, HDFG localizes hierarchical informative object/parts by grouping intermediate feature vectors based on their similarity to these prototypes. The proposed HFO is trained in an end-to-end manner using only identity labels, making it a practical solution for occluded person retrieval. We verify the effectiveness of the proposed method on three challenging occluded datasets and two holistic datasets, i.e., Occluded-DukeMTMC, Occluded-REID, P-DukeMTMC-reID, Market1501, and DukeMTMC-reID. Extensive experiments and ablation studies demonstrate superior or comparable performance of the proposed method over the state-of-the-art methods. The code is available at https://github.com/Patrickzad/HFO.
Yang Zhao 0019, Pengcheng Zhang 0003, Xiaohan Yu 0001, Zhibin Liao, Johan Verjans, Xiao Bai 0001
FG1
2023 Mix-ViT: Mixing attentive vision transformer for ultra-fine-grained visual categorization
Xiaohan Yu 0001, Jun Wang 0121, Yang Zhao 0019, Yongsheng Gao 0001
Pattern Recognit.3
2023 Gait-Assisted Video Person Retrieval
abstract
Video person retrieval aims at matching video clips of the same person across non-overlapping camera views, where video sequences contain more comprehensive information, e.g., temporal cues. How to extract useful temporal cues is the key to the success of a video person retrieval system. Gait, as a unique biometric modality indicating the way people walk, contains informative temporal information. To date, it is not clear how to fully utilize gait to boost the performance of video person retrieval. In this paper, to validate whether gait could help retrieve person in videos, we build a two-stream architecture, named appearance-gait network (AGNet), to jointly learn the appearance features and gait features from RGB video clips and silhouette video clips. We further explore how to fully utilize gait features to enhance the video feature representation. Specifically, we propose an appearance-gait attention module (AGA) to fuse a discriminative feature representation for the person retrieval task. Furthermore, to eliminate the requirement of silhouette video clips during inference, we propose a simple yet effective appearance-gait distillation module (AGD) which transfers the gait knowledge to appearance stream. As such, we are able to perform the enhanced video person retrieval without silhouette video clips, which makes the inference more flexible and practical. To the best of our knowledge, our work is the first to successfully introduce such appearance-gait knowledge distillation design for video person retrieval. We verify the effectiveness of the proposed methods on two large-scale challenging benchmarks of MARS and DukeMTMC-VideoReID. Extensive experiments demonstrate superior or comparable performance compared to the state-of-the-art methods while being much simpler. Source code is publicly available athttps://github.com/yangyangkiki/Gait-Assisted-Video-Reid.
Yang Zhao 0019, Xiaohan Yu 0001, Chunlei Liu 0001, Yongsheng Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 SPARE: Self-supervised part erasing for ultra-fine-grained visual categorization
Xiaohan Yu 0001, Yang Zhao 0019, Yongsheng Gao 0001
Pattern Recognit.2
2022 Learning discriminative region representation for person retrieval
Yang Zhao 0019, Xiaohan Yu 0001, Yongsheng Gao 0001, Chunhua Shen
Pattern Recognit.1
2022 RB-Net: Training Highly Accurate and Efficient Binary Neural Networks With Reshaped Point-Wise Convolution and Balanced Activation
abstract
In this paper, we find that the conventional convolution operation becomes the bottleneck for extremely efficient binary neural networks (BNNs). To address this issue, we open up a new direction by introducing a reshaped point-wise convolution (RPC) to replace the conventional one to build BNNs. Specifically, we conduct a point-wise convolution after rearranging the spatial information into depth, with which at least$2.25\times $computation reduction can be achieved. Such an efficient RPC allows us to explore more powerful representational capacity of BNNs under a given computation complexity budget. Moreover, we propose to use a balanced activation (BA) to adjust the distribution of the scaled activations after binarization, which enables significant performance improvement of BNNs. After integrating RPC and BA, the proposed network, dubbed as RB-Net, strikes a good trade-off between accuracy and efficiency, achieving superior performance with lower computational cost against the state-of-the-art BNN methods. Specifically, our RB-Net achieves 66.8% Top-1 accuracy with ResNet-18 backbone on ImageNet, exceeding the state-of-the-art Real-to-Binary Net (65.4%) by 1.4% while achieving more than$3\times $reduction (52M vs. 165M) in computational complexity.
Chunlei Liu 0001, Wenrui Ding, Peng Chen 0037, Bohan Zhuang, Yufeng Wang 0004, Yang Zhao 0019, Baochang Zhang 0001, Yuqi Han
IEEE Trans. Circuits Syst. Video Technol.6
2021 Benchmark Platform for Ultra-Fine-Grained Visual Categorization Beyond Human Performance
abstract
Deep learning methods have achieved remarkable success in fine-grained visual categorization. Such successful categorization at sub-ordinate level, e.g., different animal or plant species, however relies heavily on the visual differences that human can observe and the ground-truths are labelled on the basis of such human visual observation. In contrast, few research has been done for visual categorization at the ultra-fine-grained level, i.e., a granularity where even human experts can hardly identify the visual differences or are not yet able to give affirmative labels by inferring observed pattern differences. This paper reports our efforts towards mitigating this research gap. We introduce the ultra-fine-grained (UFG) image dataset, a large collection of 47,114 images from 3,526 categories. All the images in the proposed UFG image dataset are grouped into categories with different confirmed cultivar names. In addition, we perform an extensive evaluation of state-of-the-art fine-grained classification methods on the proposed UFG image dataset as comparative baselines. The proposed UFG image dataset and evaluation protocols is intended to serve as a benchmark platform that can advance research of visual classification from approaching human performance to beyond human ability, via facilitating benchmark data of artificial intelligence (AI) not to be limited by the labels of human intelligence (HI). The dataset is available online at https://githuh.com/XiaohanYu-GU/Ultra-FGVC.
Xiaohan Yu 0001, Yang Zhao 0019, Yongsheng Gao 0001, Shengwu Xiong 0001
ICCV2
2021 Deep High-Resolution Representation Learning for Visual Recognition
abstract
High-resolution representations are essential for position-sensitive vision problems, such as human pose estimation, semantic segmentation, and object detection. Existing state-of-the-art frameworks first encode the input image as a low-resolution representation through a subnetwork that is formed by connecting high-to-low resolution convolutions in series (e.g., ResNet, VGGNet), and then recover the high-resolution representation from the encoded low-resolution representation. Instead, our proposed network, named as High-Resolution Network (HRNet), maintains high-resolution representations through the whole process. There are two key characteristics: (i) Connect the high-to-low resolution convolution streams in parallel and (ii) repeatedly exchange the information across resolutions. The benefit is that the resulting representation is semantically richer and spatially more precise. We show the superiority of the proposed HRNet in a wide range of applications, including human pose estimation, semantic segmentation, and object detection, suggesting that the HRNet is a stronger backbone for computer vision problems. All the codes are available at https://github.com/HRNet.
Jingdong Wang 0001, Ke Sun 0009, Tianheng Cheng, Borui Jiang, Chaorui Deng, Yang Zhao 0019, Dong Liu 0002, Yadong Mu, Mingkui Tan, Xinggang Wang, Wenyu Liu 0001, Bin Xiao 0004
IEEE Trans. Pattern Anal. Mach. Intell.6
2021 MaskCOV: A random mask covariance network for ultra-fine-grained visual categorization
Xiaohan Yu 0001, Yang Zhao 0019, Yongsheng Gao 0001, Shengwu Xiong 0001
Pattern Recognit.2
2021 Learning deep part-aware embedding for person retrieval
Yang Zhao 0019, Chunhua Shen, Xiaohan Yu 0001, Hao Chen 0041, Yongsheng Gao 0001, Shengwu Xiong 0001
Pattern Recognit.1
2020 Patchy Image Structure Classification Using Multi-Orientation Region Transform
abstract
Exterior contour and interior structure are both vital features for classifying objects. However, most of the existing methods consider exterior contour feature and internal structure feature separately, and thus fail to function when classifying patchy image structures that have similar contours and flexible structures. To address above limitations, this paper proposes a novel Multi-Orientation Region Transform (MORT), which can effectively characterize both contour and structure features simultaneously, for patchy image structure classification. MORT is performed over multiple orientation regions at multiple scales to effectively integrate patchy features, and thus enables a better description of the shape in a coarse-to-fine manner. Moreover, the proposed MORT can be extended to combine with the deep convolutional neural network techniques, for further enhancement of classification accuracy. Very encouraging experimental results on the challenging ultra-fine-grained cultivar recognition task, insect wing recognition task, and large variation butterfly recognition task are obtained, which demonstrate the effectiveness and superiority of the proposed MORT over the state-of-the-art methods in classifying patchy image structures. Our code and three patchy image structure datasets are available at: https://github.com/XiaohanYu-GU/MReT2019.
Xiaohan Yu 0001, Yang Zhao 0019, Yongsheng Gao 0001, Shengwu Xiong 0001
AAAI2
2020 MobileFAN: Transferring deep hidden representation for face alignment
Yang Zhao 0019, Yifan Liu 0001, Chunhua Shen, Yongsheng Gao 0001, Shengwu Xiong 0001
Pattern Recognit.1
2016 Research on campus traffic congestion detection using BP neural network and Markov model
Xiaohan Yu 0001, Shengwu Xiong 0001, W. Eric Wong, Yang Zhao 0019
J. Inf. Secur. Appl.5