Yikui Zhai

dblp:59/7722 · DBLP profile ↗
← Back
36ranked-venue papers
8as first author
29since 2021 · last 2026
0000-0003-0154-9743ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 20 · 5 first-author · 19 since 2021Artificial intelligence and machine learning · 8 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 GAformer: Low-Light Image Enhancement Based on Gradient-Aware Kernel and Frequency-Modulated Transformer
abstract
Low-light image enhancement aims to improve contrast and detail representation under insufficient illumination. However, existing methods primarily rely on deepening or widening convolutional layers, while neglecting image prior information, often resulting in detail loss and visual distortion. Moreover, convolutional neural networks (CNNs) struggle to capture long-range dependencies. To address these limitations, we propose a Gradient-Aware Transformer (GAformer), which integrates gradient-aware convolutions with Transformer-based global modelling. By leveraging gradient priors for enhanced local structural representation and exploiting the global interaction capability of Transformers, GAformer achieves more comprehensive and stable enhancement. Specifically, Gradient-Aware Kernels (GAK) are introduced to optimise edge feature extraction, followed by an Illumination Map-Guided Attention (IGA) mechanism that selectively enhances low-illumination regions. Furthermore, a Frequency-Modulated Calibration (FMC) module facilitates interaction between low- and high-frequency components for progressive guided recovery. Experimental results on multiple benchmark datasets demonstrate that GAformer outperforms state-of-the-art methods in both quantitative evaluation and visual quality.
Yifan Shuai, Ke Wang 0068, Weiming Feng 0005, Shuai Pang, Dehua Zhou, Yikui Zhai
ICMR6
2026 A real-time vehicle detection method in unmanned aerial vehicle images with selective contextual features
Wanxia Huang, Chaojun Dong, Xiankun Liu, Ye Li 0002, Yikui Zhai, Kaitong Ou, Hao Quan 0002
Eng. Appl. Artif. Intell.5
2026 Diffusion-augmented direct classification: A few-shot learning framework for Synthetic Aperture Radar image automatic target recognition
Zilu Ying, Wenyu Ke, Yikui Zhai, Xinglin Liu, Pasquale Coscia, Angelo Genovese
Eng. Appl. Artif. Intell.3
2026 DLGCNet: Multimodal remote sensing semantic segmentation via dual diagonal low-rank adaptation and graph convolutional feature fusion
Jun-Ying Zeng, Xudong Jia 0001, Bin Deng 0003, Yikui Zhai, Chuanbo Qin, Pasquale Coscia, Angelo Genovese
Knowl. Based Syst.5
2026 DUR-Net+: Semi-Supervised Abdominal CT Pheochromocytoma Segmentation via Dynamic Uncertainty Rectified and Prior Knowledge From SAM-Med3D
abstract
Pheochromocytoma is a rare urological adrenal tumor disease. Automated segmentation of pheochromocytomas from computed tomography (CT) is essential for diagnosis and treatment. However, this task is a challenging one due to issues such as blurred boundaries, irregular shapes, variations in location and size, and the lack of annotated images for training. To address these issues, we propose a semi-supervised framework for pheochromocytoma segmentation that primarily consists of a dynamic uncertainty rectification mechanism and a supervised strategy based on SAM-Med3D prior knowledge. First, we design a semi-supervised segmentation model comprising a shared encoder and multiple independent decoders that dynamically select pseudo labels from the different decoder outputs. To mitigate the risk of unreliable predictions caused by sparse annotations during training, we introduce uncertainty estimation to prioritize reliable outputs. Additionally, an Attentional Convolution Block (ACB) is designed in the encoding stage to fully utilize both global and local features, improving tumor recognition in segmentation. Furthermore, SAM-Med3D prior knowledge is incorporated into the framework as supplementary supervisory information, aiding the model in learning from limited labeled data. To eliminate the labor-intensive requirement for manual prompts in SAM-Med3D, we leverage pseudo labels to generate high-quality mask prompts, thus transforming the clinical workflow. Experiments on two pheochromocytoma datasets from different centers demonstrate that our proposed method achieves competitive performance.
Chuanbo Qin, Zhuyuan Chen, Dong Wang 0083, Jun-Ying Zeng, Xudong Jia 0001, Maoqing Hu, Yikui Zhai, Pasquale Coscia, Angelo Genovese
IEEE J. Biomed. Health Informatics10
2026 Bidirectional Interactive Multi-Scale Aggregation Network for Vehicle Detection in Urban Traffic
abstract
Existing UAV vehicle-detection datasets, typically captured under static and uniform illumination, fail to adequately represent the variable lighting conditions, dense traffic, and frequent occlusions observed in real-world transportation hubs. To bridge this gap, a new dataset, UAV-HubSurveillance, is introduced to capture complex vehicle interactions across urban transportation nodes under diverse environmental scenarios. Although UAV-HubSurveillance provides rich and multidimensional interaction data, it still suffers from severe occlusions and adverse weather conditions that hinder detection and identification accuracy. To address these limitations, a novel vehicle detection framework, termed bidirectional interactive multi-scale aggregation-yolo (BIMSA-YOLO), is proposed, which integrates bidirectional feature interaction with adaptive multi-scale aggregation to enhance detection robustness. First, the bidirectional shallow fusion module (BSFM) facilitates cross-resolution information exchange through a lightweight gating strategy, preserving fine-grained details of small objects. Second, the interactive deep fusion module (IDFM) reinforces contextual coherence via attention-guided cross-level semantic fusion. Third, the multi-scale adaptive aggregation module (MSAAM) dynamically aligns and integrates multi-scale features to improve robustness against scale variation. Extensive experiments conducted on the UAV-HubSurveillance dataset demonstrate that BIMSA-YOLO significantly enhances detection performance under dynamic, occluded, and adverse-weather conditions. Specifically, the proposed model achieves an mAP0.5of 63.2%, surpassing the baseline by 5.3 percentage points. Furthermore, BIMSA-YOLO also exhibits strong generalization capabilities on VisDrone and CARPK datasets. Our code and dataset are available athttps://github.com/yikuizhai/BIMSA-YOLO
Chaojun Dong, Wenkang Qiu, Ye Li 0002, Yikui Zhai, Xiankun Liu, Chaoyun Mai, Hufei Zhu, Pasquale Coscia, Angelo Genovese, C. L. Philip Chen
IEEE Trans. Intell. Transp. Syst.4
2026 Useg-PanoDepth:Unified $360^{\circ }$ Depth Estimation for Indoor and Outdoor Scenes With Semantic Assistance
abstract
In complex$360^{\circ }$scenes, depth estimation is challenging for small objects and the depth of object boundaries, which cannot be effectively solved with existing works.$360^{\circ }$depth estimation is unable to produce uniform depth estimate findings in both indoor and outdoor settings due to the datasets. In this paper, the Useg-PanoDepth and PanoDepth dataset is proposed to improve the above problems effectively. The Diagonal-aware Attention Module (DAM) effectively estimates small objects in complex scenes. Enhanced Boundary Module (EBM), for enhancing boundary information,can also effectively solve the problem of depth unification of indoor and outdoor scenes. Extensive experiments on our constructed PanoDepth dataset, Useg-PanoDepth achieves SOTA results. The Relative accuracy (deltahttps://github.com/xjh6/Useg-PanoDepth.
Qingling Chang, Jingheng Xu, Yan Cui 0011, Yikui Zhai, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Multim.4
2025 Layer-Wise Unlearning for Model Adaption in Non-Stationary Environments
abstract
Although the fine-tuning-based deep transfer learning method performs well in adapting a pre-trained model to downstream tasks, it struggles in non-stationary environments where the data distribution changes dynamically over time. In such scenarios, the trained model rapidly becomes obsolete. Directly fine-tuning it for new environments will inevitably lead to negative transfer, degrading the model's generalization performance. Inspired by neuroscience, some existing studies have proposed a so-called “unlearning-and-relearning” training paradigm to alleviate this issue. However, existing methods face limitations; for example, they may destroy general features in shallow layers and lack adaptability when performing deep-layer unlearning. This paper proposes a Layer-wise Unlearning (LwU) method. Our approach first identifies the layers requiring unlearning by estimating the transferability of each layer in the trained model. Then, within the selected layers, it preserves high-sensitivity parameters and re-initializes low-sensitivity ones based on a parameter sensitivity analysis to achieve unlearning. Extensive experiments on synthetic, real-world, and image datasets demonstrate that LwU effectively overcomes negative transfer and substantially improves the model's generalization performance in new environments. The code of the proposed method and data are available at https://github.com/mlmmwym/LwU.
Yanbing Zhou, Yimin Wen, Zhanhua Liu, Hang Yu 0006, Yikui Zhai
ICDM6
2025 F2MCANet: joint frequency fusion and multi-channels attention neural network for surface defect recognition
Tianlei Wang, Zeliang Li, Yikui Zhai, Jiajie Tian
Soft Comput.5
2025 PBSD-Net: Prismatic Battery Surface Defect Detection via Sliding Slice Amplification and Shunted Dynamic Snake Convolution
abstract
Automatically detecting surface defects in prismatic battery is crucial for ensuring quality meets established standards. Traditional methods face challenges in accurately identifying these defects due to their minute and varied shapes and high density of distribution. To address these issues, we propose an innovative network for prismatic battery surface defect (PBSD-Net), which employs shunted dynamic snake convolution and focal modulation to detect surface defects in prismatic battery. This network is integrated into the 2D-AOI system. Firstly, we introduce sliding slice amplification (SSA) as a training strategy to enhance the network’s ability to recognize densely clustered tiny defects. Secondly, we develop a novel method using the shunted dynamic snake convolution (SDSC) module and focal modulation (FM) to improve the extraction of deformation features, thereby addressing complex and sporadically scattered surface defects. By integrating the SDSC module and FM mechanism, the receptive field of the defect feature extraction network is expanded, enabling the acquisition of comprehensive defect edge features. Additionally, we introduce the quality focal loss (QFL) function to effectively tackle the issue of imbalanced sample types. Experimental results on the PBSD-RGB dataset demonstrate that our method achieves a mAP@50 of 85.8%, representing an improvement of approximately 7.7% over the baseline network. We have applied the PBSD-Net to an automatic defect detection system in a well-known battery production company. This enhancement significantly boosts the accuracy of surface defect detection in prismatic battery. The relevant code is at the https://github.com/yikuizhai/PBSD-Net.
Ying Xu 0005, Bo Li 0165, Yikui Zhai, Feng Ke, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans Autom. Sci. Eng.3
2025 GMTNet: Dense Object Detection via Global Dynamically Matching Transformer Network
abstract
In recent years, object detection models have been extensively applied across various industries, leveraging learned samples to recognize and locate objects. However, industrial environments present unique challenges, including complex backgrounds, dense object distributions, object stacking, and occlusion. To address these challenges, we propose the Global Dynamic Matching Transformer Network (GMTNet). GMTNet partitions images into blocks and employs a sliding window approach to capture information from each block and their interrelationships, mitigating background interference while acquiring global information for dense object recognition. By reweighting key-value pairs in multi-scale feature maps, GMTNet enhances global information relevance and effectively handles occlusion and overlap between objects. Furthermore, we introduce a dynamic sample matching method to tackle the issue of excessive candidate boxes in dense detection tasks. This method adaptively adjusts the number of matched positive samples according to the specific detection task, enabling the model to reduce the learning of irrelevant features and simplify post-processing. Experimental results demonstrate that GMTNet excels in dense detection tasks and outperforms current mainstream algorithms. The code will be available athttp://github.com/yikuizhai/GMTNet.
Chaojun Dong, Chengxuan Wang, Yikui Zhai, Ye Li 0002, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Circuits Syst. Video Technol.3
2025 AEGL-Net: Adaptive Multiscale Global-Local Feature Fusion Network for Remote Sensing Change Detection
abstract
With the rapid advancements in deep learning technology, the field of remote sensing change detection (RSCD) has witnessed significant improvements and innovations. In this context, bitemporal image processing, using features directly extracted by the backbone for subsequent fusion operations, may be obstructed by external environmental factors, potentially limiting the effective capture of complex feature variations. Moreover, overlooking local features during the fusion of bitemporal features can significantly affect the final detection results. As a result, achieving accurate change detection (CD) still encounters various challenges. To tackle these issues, this paper proposes a CD network (AEGL-Net) with Adaptive Multiscale Enhancement (AME) and Global-Local Feature Fusion (GLFF) modules. First, AME enhances features at each stage of backbone extraction through an adaptive strategy, balancing the enhancement of semantic information and texture details. Then, GLFF is used to fuse the bitemporal image features, which enhances the modeling of global dependencies while also fusing shared and context-aware weights to enhance the local features. Finally, the merged features are fed into the decoder to generate precise change maps. Experiments conducted with four open RSCD datasets (LEVIR-CD, S2Looking, SYSU-CD, and UAV-CD) demonstrate that our proposed AEGL-Net outperforms ten state-of-the-art models in the RSCD field. Our code is available at https://github.com/yikuizhai/AEGL-Net.
Zilu Ying, Yikui Zhai, Hufei Zhu, Hongsheng Zhang 0001, Pasquale Coscia, Angelo Genovese, Fabio Scotti, Vincenzo Piuri, C. L. Philip Chen
IEEE Trans. Geosci. Remote. Sens.3
2025 Spatial Reconstruction and Joint Training in Transformer Network for Cross-Domain Remote Sensing Images Semantic Segmentation
abstract
Recently, Unsupervised Domain Adaptation (UDA) methods have attracted considerable attention in Remote Sensing Images (RSI) semantic segmentation. However, cross-domain RSI exhibit diverse scales, imbalanced distributions within domains, and significant inter-domain variations. In response to these challenges, we combine Spatial reconstruction and Joint training with the Transformer Network (SJT-Net). This framework introduces a spatial reconstruction method to address the issue of inconsistent ground sampling distances in cross domain RSI, which is rarely considered in existing approaches. Transferring domain knowledge at a similar spatial scale improves the spatial representation ability of UDA models. Unlike traditional adversarial training using ResNet for feature extraction, the SJT-Net employs Segformer, which enhances the model’s ability to capture in-class features across domains and improves global dependency modeling. Transmitting these refined features to the discriminator allows for more precise feature-level domain alignment. To enhance feature decoding, an interactive global-local decoder is constructed to efficiently capture both global relationships and local details of landform objects. Our framework leverages adversarial training to generate highly confident model weights and pseudo-labels for self-training in the target domain. Through iterative updates, the model’s generalization capability is gradually improved, eventually achieving optimal segmentation performance. Experimental results demonstrate that SJT-Net outperforms current UDA approaches and accomplishes state-of-the-art (SOTA) segmentation accuracy. The repository can be accessed at https://github.com/AnsonD0820/SJT-Net.
Jun-Ying Zeng, Senyao Deng, Yikui Zhai, Xudong Jia 0001, Chuanbo Qin, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Geosci. Remote. Sens.3
2025 Multimodal Feature Fusion Network With Text Difference Enhancement for Remote Sensing Change Detection
abstract
Although deep learning has advanced remote sensing change detection (RSCD), most methods rely solely on image modality, limiting feature representation, change pattern modeling, and generalization—especially under illumination and noise disturbances. To address this, we propose MMChange, a multimodal RSCD method that combines image and text modalities to enhance accuracy and robustness. An Image Feature Refinement (IFR) module is introduced to highlight key regions and suppress environmental noise. To overcome the semantic limitations of image features, we employ a vision-language model (VLM) to generate semantic descriptions of bi-temporal images. A Textual Difference Enhancement (TDE) module then captures fine-grained semantic shifts, guiding the model toward meaningful changes. To bridge the heterogeneity between modalities, we design an Image-Text Feature Fusion (ITFF) module that enables deep cross-modal integration. Extensive experiments on LEVIR-CD, WHU-CD, and SYSU-CD demonstrate that MMChange consistently surpasses state-of-the-art methods across multiple metrics, validating its effectiveness for multimodal RSCD. Code is available at: https://github.com/yikuizhai/MMChange.
Yikui Zhai, Zilu Ying, Tingfeng Xian, Wenlve Zhou, Zhiheng Zhou 0001, Xudong Jia 0001, Hongsheng Zhang 0001, C. L. Philip Chen
IEEE Trans. Geosci. Remote. Sens.2
2025 CLIP-Vision Guided Few-Shot Metal Surface Defect Recognition
abstract
Metal surface defect recognition (MSDR) based on deep learning encounters the challenge of few-shot expert-labeled data. In this study, we proposed a CLIP-vision guided self supervised learning (CVGSSL) framework for representation learning of unlabeled data, completing MSDR using few-shot labeled data. This framework initially generates rich and diverse representation information through multiple CLIP-Vs to ensure effective SSL pretraining, followed by the design of an MLP-adapter to distill knowledge and adapt these representations to recognition tasks. In addition, we constructed a self-constrained loss to address the inherent problem of intraclass and interclass distance ambiguity that causes the representation to fall into an equivocal decision margin. Following label-free pretraining of CVGSSL, the downstream model adapts to one-shot to four-shot defect recognition tasks through fine-tuning. Experimental results demonstrate that CVGSSL outperforms state-of-the-art SSL methods across three public metal surface defect datasets, with the efficacy of the approach validated through extensive ablation experiments.
Tianlei Wang, Zeliang Li, Ying Xu 0005, Yikui Zhai, Xiaofen Xing, Kailing Guo, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Ind. Informatics4
2024 MSFA-Net : Multiple Spatial-Channel Feature Aggregation Network for Change Detection and a UAV-CD Dataset
abstract
Change detection in remote sensing images is pivotal for monitoring and comprehending dynamic environmental phenomena. Nonetheless, conventional change detection models have grappled with correlating information between channel and spatial dimensions due to inherent feature extraction limitations. Hence, this paper proposes an innovative change detection framework and a high resolution UAV change detection dataset named UAV-CD dataset. The introduced network embraces a Siamese network, amalgamating a feature extraction backbone network along with spatial and channel reconstruction convolution (ScConv) and Ghost modules. A primary contribution of this paper is the incorporation of ScConv into the change detection network, facilitating the reconstruction of information in both spatial and channel dimensions. Additionally, the Ghost module is employed to fortify information across distinct channel dimensions within the feature maps. Compared with current state-of-the-art methods, it is indicated that the proposed approach achieves superior performance on the LEVIR-CD, SYSU-CD, and our proposed dataset UAV-CD.
Yikui Zhai, Haolin Lv, Tingfeng Xian, Zilu Ying, Hao Quan 0002, Xudong Jia 0001
IGARSS2
2024 A Scale-Temporal Interaction Network For Remote Sensing Image Change Detection And A UAV-CD Dataset
abstract
Remote sensing (RS) image change detection (CD) is a challenging visual task due to its rich and complex image information. Nowadays, CD has yielded fruitful results. However, insufficient feature interaction hinders further improvement of CD performance. In this paper, we introduce a scale-temporal interaction network (STI-Net). It extracts multi-scale bitemporal features using a depth-separable convolution-based Siamese encoder, followed by both Cross-Scale Feature Interaction (CSFI) and Cross-Temporal Feature Interaction (CTFI). Finally, we employ a straightforward decoder to generate the change map. Additionally, to enrich the CD data, we introduced a new dataset based on UAV optical image, named UAV-CD. This dataset comprises 2660 pairs of images sized at 768×768 pixels, focusing primarily on building and land changes. Experiments demonstrate that our method outperforms existing state-of-the-art methods on two public CD datasets as well as UAV-CD, showcasing excellent performance.
Tingfeng Xian, Zilu Ying, Haolin Lv, Yikui Zhai, Hao Quan 0002, Xudong Jia 0001
IGARSS5
2024 Mutual Information Compensation for High-Fidelity Image Generation With Limited Data
abstract
Limited data availability poses a perennial challenge in the field of generative image generation. However, current up-sampling methods suffer from inherent deficiencies, resulting in generators' failure to faithfully restore images with identical resolution and distribution. This is primarily evident in the diminishing mutual information as image resolution increases. Considering the above issue, we propose MICGAN, a data-efficient Generative Adversarial Network (GAN) that compensates for mutual information in every layer output feature of the generator through dense frequency skip connections. MICGAN is grounded in information theory, offering a robust theoretical foundation for mitigating mutual information decay. Additionally, we demonstrate that the augmented mutual information primarily stems from increased high-frequency components in images. Furthermore, we introduce three sub-architectures–MICGAN-A, MICGAN-C, and MICGAN-S–—each employing distinct feature fusion methods. Extensive experiments across thirteen diverse datasets validate the advancements and effectiveness of MICGAN compared to state-of-the-art methods.
Yikui Zhai, Zhihao Long, Wenfeng Pan, C. L. Philip Chen
IEEE Signal Process. Lett.1
2024 LDCL: Low-Confidence Discriminant Contrastive Learning for Small-Sample SAR ATR
abstract
Synthetic Aperture Radar (SAR) target image acquisition presents challenges and incurs high annotation costs. The emergence of self-supervised contrastive learning shows promise for SAR automatic target recognition (ATR) with limited data. However, SAR images suffer from poor discriminability and high sample similarity, hindering instance discrimination in contrastive learning. To address this, we propose Low-confidence Discriminant Contrastive Learning (LDCL), which integrates group-instance contrast and batch mixed training for SAR ATR. LDCL consists of two branches: classical instance discrimination and group-instance discrimination. We refine the SAR-group instance discrimination loss function by incorporating distance calculations to guide feature vectors towards nearest clusters, enhancing discrimination within the feature space. Additionally, we introduce a batch image mixing training strategy to reduce confidence in SAR instance discrimination while preserving intra-class consistency. Experimental results on small sample MSTAR and FUSAR-Ship datasets demonstrate that LDCL outperforms traditional transfer learning and self-supervised learning methods, achieving significantly higher recognition rates in SAR ATR tasks.
Jinrui Liao, Yikui Zhai, Qingsong Wang 0003, Bing Sun 0002, Vincenzo Piuri
IEEE Trans. Geosci. Remote. Sens.2
2024 DGMA2-Net: A Difference-Guided Multiscale Aggregation Attention Network for Remote Sensing Change Detection
abstract
Remote sensing change detection (RSCD) focuses on identifying regions that have undergone changes between two remote sensing images captured at different times. Recently, convolutional neural networks (CNNs) have shown promising results in the challenging task of RSCD. However, these methods do not efficiently fuse bitemporal features and extract useful information that is beneficial to subsequent RSCD tasks. In addition, they did not consider multilevel feature interactions in feature aggregation and ignore relationships between difference features and bitemporal features, which thus affects the RSCD results. To address the above problems, a difference-guided multiscale aggregation attention network, DGMA2-Net, is developed. Bitemporal features at different levels are extracted through a Siamese convolutional network and a multiscale difference fusion module (MDFM) is then created to fuse bitemporal features and extract, in a multiscale manner, difference features containing rich contextual information. After the MDFM treatment, two difference aggregation modules (DAMs) are used to aggregate difference features at different levels for multilevel feature interactions. The features through DAMs are sent to the difference-enhanced attention modules (DEAMs) to strengthen the connections between bitemporal features and difference features and further refine change features. Finally, refined change features are superimposed from deep to shallow and a change map is produced. In validating the effectiveness of DGMA2-Net, a series of experiments are conducted on three public RSCD benchmark datasets (LEVIR-CD, BCDD, and SYSU-CD). The experimental results demonstrate that DGMA2-Net surpasses the current eight state-of-the-art methods in RSCD. Our code is released at https://github.com/yikuizhai/DGMA2-Net.
Zilu Ying, Zijun Tan, Yikui Zhai, Xudong Jia 0001, Wenba Li, Jun-Ying Zeng, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Geosci. Remote. Sens.3
2024 DS-HyFA-Net: A Deeply Supervised Hybrid Feature Aggregation Network With Multiencoders for Change Detection in High-Resolution Imagery
abstract
With the advancement of deep learning (DL) technologies, remarkable progress has been achieved in change detection (CD). Existing DL-based methods primarily focus on the discrepancy in bitemporal images, while overlooking the commonality in bitemporal images. However, one of the reasons hindering the improvement of CD performance is the inadequate utilization of image information. To address the above issue, we propose a Deeply Supervised Hybrid Feature Aggregation Network (DS-HyFA-Net). This network predicts changes by integrating the distinctness and the commonality in bitemporal images. Specifically, the DS-HyFA-Net primarily consists of a set of encoders and a Hybrid Feature Aggregation (HyFA) module. It uses a Siamese encoder (or Encoder I) and a specialized encoder (or Encoder II) to extract distinct and common features (CFs) in bitemporal images, respectively. The HyFA module efficiently aggregates distinct and common features (or hybrid features) and generates a change map using a predictor. In addition, a common feature learning strategy (CFLS) is introduced, based on deeply supervised (DS) techniques, to guide Encoder II in learning CFs. Experimental results on three well-recognized datasets demonstrate the effectiveness of the innovative DS-HyFA-Net, achieving F1-Scores of 93.33% on WHU-CD, 90.98% on LEVIR-CD, and 81.14% on SYSU-CD. Our code is available athttps://github.com/yikuizhai/DS-HyFA-Net.
Zilu Ying, Tingfeng Xian, Yikui Zhai, Xudong Jia 0001, Hongsheng Zhang 0001, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Geosci. Remote. Sens.3
2024 CAS-Net: Comparison-Based Attention Siamese Network for Change Detection With an Open High-Resolution UAV Image Dataset
abstract
Change detection (CD) is a process of extracting changes on the Earth’s surface from bitemporal images. Current CD methods that use high-resolution remote sensing images require extensive computational resources and are vulnerable to the presence of irrelevant noises in the images. In addressing these challenges, a comparison-based attention Siamese network (CAS-Net) is proposed. The network utilizes contrastive attention modules (CAMs) for feature fusion and employs a classifier to determine similarities and differences of bitemporal image patches. It simplifies pixel-level CDs by comparing image patches. As such, the influences of image background noises on change predictions are reduced. Along with the CAS-Net, an unmanned aerial vehicle (UAV) similarity detection (UAV-SD) dataset is built using high-resolution remote sensing images. This dataset, serving as a benchmark for CD, comprises 10000 pairs of UAV images with a size of$256 \times 256$. Experiments of the CAS-Net on the UAV-SD dataset demonstrate that the CAS-Net is superior to other baseline CD networks. The CAS-Net detection accuracy is 93.1% on the UAV-SD dataset. The code and the dataset can be found athttps://github.com/WenbaLi/CAS-Net.
Yikui Zhai, Wenba Li, Tingfeng Xian, Xudong Jia 0001, Hongsheng Zhang 0001, Zijun Tan, Jun-Ying Zeng, C. L. Philip Chen
IEEE Trans. Geosci. Remote. Sens.1
2024 Efficient Adjacent Feature Harmonizer Network With UAV-CD+ Dataset for Remote Sensing Change Detection
abstract
Remote sensing change detection (RSCD) aims to identify changes within bi-temporal registered images. However, existing deep learning (DL)-based RSCD networks often suffer from large numbers of parameters, high computational complexity, and low inference speed, making it challenging to achieve efficient inference in real-world deployments. In addition, current models lack robust feature-fitting capabilities, necessitating the development of an efficient and powerful RSCD model to address this issue. Therefore, we propose a novel RSCD network named efficient adjacent feature harmonizer network (EAFH-Net) with fast computational speed and lightweight design. It is based on MobileNetV2, considering that change maps of different sizes contain temporal information of bitemporal features and spatial information at various scales, we introduce a multiscale feature neighbor fusion module (MFNFM) to address the lack of interaction between sophisticated-level and elementary-level features, and spatial and channel feature harmonizer module (SCFHM) to harmonize the spatiotemporal information of the change maps. Moreover, data-driven DL algorithms face another challenge due to insufficient granularity and the need for more practical datasets. Therefore, we present unmanned aerial vehicle (UAV)-CD+, a dataset comprising 2002 pairs of bi-temporal UAV low-altitude images, each sized at$1024\times 1024$. We performed experiments on three publicly accessible datasets in conjunction with UAV-CD+, comparing the results with other state-of-the-art (SOTA) methods. EAFH-Net attains the utmost precision, obtaining 91.74% on LEVIR-CD, 84.28% on SYSU-CD, 95.07% on WHU-CD, 79.12% on CLCD, and 70.12% on UAV-CD+. We have our model code available at the following link:https://github.com/yikuizhai/UCSFH-Net.
Yikui Zhai, Hongsheng Zhang 0001, Tingfeng Xian, Ying Xu 0005, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti, C. L. Philip Chen
IEEE Trans. Geosci. Remote. Sens.1
2024 Learning to Count Arbitrary Industrial Manufacturing Workpieces
abstract
Man-made workpiece counting is a routine job for manufactory workers; however, this is an error-prone task. In this article, we are interested in detecting and counting arbitrary workpieces in industrial manufacturing. Therefore, we construct a comprehensive and large-scale open-world public benchmark dataset for workpiece counting, called workpiece counting dataset, which includes 121 475 instances of workpieces from 351 different categories. We also propose a novel method for workpiece detection and counting, named two-stage workpiece counting network. The first stage of the network is to develop a class-agnostic detector to localize each workpiece instance, followed by the second stage to employ an unsupervised deep clustering strategy with the backbone network pretrained in a workpiece convolutional autoencoder for decision boundary prediction, achieving workpiece clustering under unknownKvalues. Finally, our experiments show that the proposed method outperforms current mainstream methods, greatly enhancing the efficiency of factory operations.
Yikui Zhai, Feng Ke, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Ind. Informatics2
2024 Large-Scale High-Altitude UAV-Based Vehicle Detection via Pyramid Dual Pooling Attention Path Aggregation Network
abstract
UAVs can collect vehicle data in high-altitude scenes, playing a significant role in intelligent urban management due to their wide of view. Nevertheless, the current datasets for UAV-based vehicle detection are acquired at altitude below 150 meters. This contrasts with the data perspective obtained from high-altitude scenes, potentially leading to incongruities in data distribution. Consequently, it is challenging to apply these datasets effectively in high-altitude scenes, and there is an ongoing obstacle. To resolve this challenge, we developed a comprehensive vehicle dataset named LH-UAV-Vehicle, specifically collected at flight altitudes ranging from 250 to 400 meters. Collecting data at higher flight altitudes offers a broader perspective, but it concurrently introduces complexity and diversity in the background, which consequently impacts vehicle localization and recognition accuracy. In response, we proposed the pyramid dual pooling attention path aggregation network (PDPA-PAN), an innovative framework that improves detection performance in high-altitude scenes by combining spatial and semantic information. Object attention integration in both spatial and channel dimensions is aimed by the pyramid dual pooling attention module (PDPAM), which is achieved through the parallel integration of two distinct attention mechanisms. Furthermore, we have individually developed the pyramid pooling attention module (PPAM) and the dual pooling attention module (DPAM). The PPAM emphasizes channel attention, while the DPAM prioritizes spatial attention. This design aims to enhance vehicle information and suppress background interference more effectively. Extensive experiments conducted on the LH-UAV-Vehicle conclusively demonstrate the efficacy of the proposed vehicle detection method. Our code and dataset can be found at https://github.com/yikuizhai/PDPA-PAN.
Zilu Ying, Yikui Zhai, Hao Quan 0002, Wenba Li, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Intell. Transp. Syst.3
2023 UAV-BCD: A UAV Building Change Detection Dataset
abstract
Remote sensing change detection (RSCD) holds significant prominence as a research topic within the realm of computer vision. However, previous RSCD datasets have been constructed based on satellite remote sensing images. Traditional satellite remote sensing images have problems such as insufficient resolution, difficult data acquisition, and complex processing processes, and there is a certain gap between data distribution and actual needs. UAVs not only have the advantages of flexibility and high-speed, but also can capture high-resolution images, which are especially suitable for high-precision RSCD in small areas. Therefore, this paper proposed a new UAV RSCD dataset — UAV Building Change Detection Dataset (UAV-BCD). The proposed dataset contains 2024 pairs of finely registered high-resolution images collected by UAVs and their corresponding pixel-level labels, which can provide a new benchmark for RSCD. We evaluate the effectiveness of UAV-BCD with the five state-of-art deep neural networks in RSCD.
Zilu Ying, Zijun Tan, Wenba Li, Zhangzhao Liang, Yikui Zhai
IGARSS6
2023 SAS-NET: Similarity Attention Siamese Network for Building Change Detection in UAV Images
abstract
Change detection refers to extract change information using deep learning or traditional image processing methods to quantitatively analyze and characterize landmark changes on bi-temporal images. Currently, change detection is mainly a pixel-level task, and obtaining accurate change detection segmentation predictions requires a more elaborate and complex model architecture design. To simplify the change detection task, we proposed a novel similarity detection model, Similarity Attention Siamese Network (SAS-NET). It analyzed and predicted if the bi-temporal image patches were similar, and simplified pixel-level change detection tasks to patch-level similarity classification prediction tasks. In this work, a UAV Similarity Detection Dataset (UAV-SD) was also proposed to explore the advantages of patch-level prediction tasks over pixel-level change detection tasks. The proposed method achieved 90.5% accuracy on UAV-SD, which proves that it is more effective than other advanced change detection methods.
Yikui Zhai, Wenba Li, Zijun Tan, Zilu Ying
IGARSS1
2022 Joint Transformer and Multi-scale CNN for DCE-MRI Breast Cancer Segmentation
abstract
Abstract Automatic segmentation of breast cancer lesions in dynamic contrast-enhanced magnetic resonance imaging is challenged by low accuracy of delineation of the infiltration area, variable structure and shapes, large intensity heterogeneity changes, and low boundary contrast. This study constructed a two-stage breast cancer image segmentation framework and proposes a novel breast cancer lesion segmentation model (TR-IMUnet). The benchmark U-Net network model enables a rough delineation of the breast area in the acquired images and eliminates the influence of unrelated tissues (chest muscle, fat, and heart) on breast tumor segmentation. Based on the extracted results of the region of interest, the rectified linear unit (ReLU) function of the encoding–decoding structure in the model was replaced by an improved ReLU function to reserve and adjust the data dynamically according to input information. The segmentation accuracy of breast cancer lesions was improved by embedding a multi-scale fusion block and a transformer module in the coding path of the model, thereby obtaining multi-scale and global attention information. The experimental results showed that the breast tumor segmentation indexes Dice coefficient (Dice), Intersection over Union (IoU), Sensitivity (SEN), and Positive Predictive Value (PPV) increased by 4.27, 5.21, 3.37, and 3.68%, respectively, relative to the U-Net reference model. The proposed model improves the segmentation results of breast cancer lesions and reduces small area mis-segmentation and calcification segmentation.
Chuanbo Qin, Jun-Ying Zeng, Lianfang Tian, Yikui Zhai, Xiaozhi Zhang
Soft Comput.5
2022 Weakly Contrastive Learning via Batch Instance Discrimination and Feature Clustering for Small Sample SAR ATR
abstract
In recent years, impressive performance of deep learning technology has been recognized in synthetic aperture radar (SAR) automatic target recognition (ATR). Since a large amount of annotated data are required in this technique, it poses a trenchant challenge to the issue of obtaining a high recognition rate through less labeled data. To overcome this problem, inspired by the contrastive learning, we proposed a novel framework named batch instance discrimination and feature clustering (BIDFC). In this framework, different from that of the objective of general contrastive learning methods, embedding distance between samples should be moderate because of the high similarity between samples in the SAR images. Consequently, our flexible framework is equipped with adjustable distance between embedding, which we term as weakly contrastive learning. Technically, instance labels are assigned to the unlabeled data in per batch, and random augmentation and training are performedfewtimes on these augmented data. Meanwhile, a novel dynamic-weighted variance loss (DWV loss) function is also posed to cluster the embedding of enhanced versions for each sample. The experimental results on the moving and stationary target acquisition and recognition (MSTAR) database indicate a 91.25% classification accuracy of our method fine-tuned on only 3.13% training data. Even though a linear evaluation is performed on the same training data, the accuracy can still reach 90.13%. We also verified the effectiveness of BIDFC in OpenSarShip database, indicating that our method can be generalized to other data sets. Our code is available at:https://github.com/Wenlve-Zhou/BIDFC-master.
Yikui Zhai, Wenlve Zhou, Bing Sun 0002, Jingwen Li 0003, Qirui Ke, Zilu Ying, Junying Gan, Chaoyun Mai, Ruggero Donida Labati, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Geosci. Remote. Sens.1
2020 Weakly supervised facial expression recognition via transferred DAL-CNN and active incremental learning
Ying Xu 0005, Yikui Zhai, Junying Gan, Jun-Ying Zeng, He Cao, Fabio Scotti, Vincenzo Piuri, Ruggero Donida Labati
Soft Comput.3
2019 Region-division-based joint sparse representation classification for hyperspectral images
abstract
In this article, a region‐division‐based joint sparse representation classification (RDJSRC) method is proposed to solve the heterogeneous region problem in the joint sparse representation classification (JSRC) method used in hyperspectral image (HSI) classification. The RDJSRC method incorporates regional information, obtained by the hidden Markov random field (HMRF), into the JSRC to reduce the interference of heterogeneous pixels in the neighbourhood of the test pixel and finally improve the classification performance. The framework of this method is as follows. The first several principal components (PCs) are initially selected to be the new HSI by transforming the original HSI with the PC analysis algorithm. Then, the regional information containing the spatial structure of the HSI is obtained by applying the HMRF algorithm to the first PC. Through incorporating this regional information into the JSRC procedure, the initial label of the test pixel can be jointly determined by the new HSI pixels within the homogeneity in the search window. Ultimately, the final label of the test pixel is determined by a voting strategy based on multiple classification results. Compared with several classification methods, experimental results, indicate that this method achieves improvement from 2 to 3% in HSI classification.
Yikui Zhai, Lei Liu 0032
IET Image Process.3
2017 SAR Automatic Target Recognition Based on Deep Convolutional Neural Network
Ying Xu 0005, Kaipin Liu, Zilu Ying, Lijuan Shang, Yikui Zhai, Vincenzo Piuri, Fabio Scotti
ICIG (3)6
2017 Deep Convolutional Neural Network for Facial Expression Recognition
Yikui Zhai, Jun-Ying Zeng, Vincenzo Piuri, Fabio Scotti, Zilu Ying, Ying Xu 0005, Junying Gan
ICIG (1)1
2014 Deep self-taught learning for facial beauty prediction
Junying Gan, Lichen Li 0002, Yikui Zhai, Yinhua Liu
Neurocomputing3
2013 Disguised face recognition via local phase quantization plus geometry coverage
abstract
Disguised face recognition (FR) is considered as one of the difficult and important problems in FR field. Rather than disguised modeling, a disguised face recognition algorithm based on local phase quantization (LPQ) feature and geometry coverage is presented in this paper. LPQ method is applied to extract the phase statistics feature which is robust to the disguised mode, and hyper sausage neuron based on biomimetic pattern recognition (BPR) theory is adopted to construct high-dimensional geometry coverage of different classes, which makes full use of continuous characteristics of different class face features while avoids the interruption of the disguised mode. Experiments on AR face database and disguised face database established by police face combination software show that, compared with the state-of-the-art method, the proposed recognition algorithm can achieve high recognition results under disguised conditions.
Yikui Zhai, Junying Gan, Jun-Ying Zeng, Ying Xu 0005
ICASSP1
2012 A Novel Artificial Fish Swarm Algorithm Based on Multi-objective Optimization
Yikui Zhai, Ying Xu 0005, Junying Gan
ICIC (2)1