EDBT 2026 Demo / reviewers in the wild / expert
Zilu Ying
dblp:41/7370
· DBLP profile ↗
15ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0002-3074-5586ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Diffusion-augmented direct classification: A few-shot learning framework for Synthetic Aperture Radar image automatic target recognition
Zilu Ying, Wenyu Ke, Yikui Zhai, Xinglin Liu, Pasquale Coscia, Angelo Genovese |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | AEGL-Net: Adaptive Multiscale Global-Local Feature Fusion Network for Remote Sensing Change DetectionabstractWith the rapid advancements in deep learning technology, the field of remote sensing change detection (RSCD) has witnessed significant improvements and innovations. In this context, bitemporal image processing, using features directly extracted by the backbone for subsequent fusion operations, may be obstructed by external environmental factors, potentially limiting the effective capture of complex feature variations. Moreover, overlooking local features during the fusion of bitemporal features can significantly affect the final detection results. As a result, achieving accurate change detection (CD) still encounters various challenges. To tackle these issues, this paper proposes a CD network (AEGL-Net) with Adaptive Multiscale Enhancement (AME) and Global-Local Feature Fusion (GLFF) modules. First, AME enhances features at each stage of backbone extraction through an adaptive strategy, balancing the enhancement of semantic information and texture details. Then, GLFF is used to fuse the bitemporal image features, which enhances the modeling of global dependencies while also fusing shared and context-aware weights to enhance the local features. Finally, the merged features are fed into the decoder to generate precise change maps. Experiments conducted with four open RSCD datasets (LEVIR-CD, S2Looking, SYSU-CD, and UAV-CD) demonstrate that our proposed AEGL-Net outperforms ten state-of-the-art models in the RSCD field. Our code is available at https://github.com/yikuizhai/AEGL-Net. Zilu Ying, Yikui Zhai, Hufei Zhu, Hongsheng Zhang 0001, Pasquale Coscia, Angelo Genovese, Fabio Scotti, Vincenzo Piuri, C. L. Philip Chen |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Multimodal Feature Fusion Network With Text Difference Enhancement for Remote Sensing Change DetectionabstractAlthough deep learning has advanced remote sensing change detection (RSCD), most methods rely solely on image modality, limiting feature representation, change pattern modeling, and generalization—especially under illumination and noise disturbances. To address this, we propose MMChange, a multimodal RSCD method that combines image and text modalities to enhance accuracy and robustness. An Image Feature Refinement (IFR) module is introduced to highlight key regions and suppress environmental noise. To overcome the semantic limitations of image features, we employ a vision-language model (VLM) to generate semantic descriptions of bi-temporal images. A Textual Difference Enhancement (TDE) module then captures fine-grained semantic shifts, guiding the model toward meaningful changes. To bridge the heterogeneity between modalities, we design an Image-Text Feature Fusion (ITFF) module that enables deep cross-modal integration. Extensive experiments on LEVIR-CD, WHU-CD, and SYSU-CD demonstrate that MMChange consistently surpasses state-of-the-art methods across multiple metrics, validating its effectiveness for multimodal RSCD. Code is available at: https://github.com/yikuizhai/MMChange. Yikui Zhai, Zilu Ying, Tingfeng Xian, Wenlve Zhou, Zhiheng Zhou 0001, Xudong Jia 0001, Hongsheng Zhang 0001, C. L. Philip Chen |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | MSFA-Net : Multiple Spatial-Channel Feature Aggregation Network for Change Detection and a UAV-CD DatasetabstractChange detection in remote sensing images is pivotal for monitoring and comprehending dynamic environmental phenomena. Nonetheless, conventional change detection models have grappled with correlating information between channel and spatial dimensions due to inherent feature extraction limitations. Hence, this paper proposes an innovative change detection framework and a high resolution UAV change detection dataset named UAV-CD dataset. The introduced network embraces a Siamese network, amalgamating a feature extraction backbone network along with spatial and channel reconstruction convolution (ScConv) and Ghost modules. A primary contribution of this paper is the incorporation of ScConv into the change detection network, facilitating the reconstruction of information in both spatial and channel dimensions. Additionally, the Ghost module is employed to fortify information across distinct channel dimensions within the feature maps. Compared with current state-of-the-art methods, it is indicated that the proposed approach achieves superior performance on the LEVIR-CD, SYSU-CD, and our proposed dataset UAV-CD. Yikui Zhai, Haolin Lv, Tingfeng Xian, Zilu Ying, Hao Quan 0002, Xudong Jia 0001 |
IGARSS | 5 |
| 2024 | A Scale-Temporal Interaction Network For Remote Sensing Image Change Detection And A UAV-CD DatasetabstractRemote sensing (RS) image change detection (CD) is a challenging visual task due to its rich and complex image information. Nowadays, CD has yielded fruitful results. However, insufficient feature interaction hinders further improvement of CD performance. In this paper, we introduce a scale-temporal interaction network (STI-Net). It extracts multi-scale bitemporal features using a depth-separable convolution-based Siamese encoder, followed by both Cross-Scale Feature Interaction (CSFI) and Cross-Temporal Feature Interaction (CTFI). Finally, we employ a straightforward decoder to generate the change map. Additionally, to enrich the CD data, we introduced a new dataset based on UAV optical image, named UAV-CD. This dataset comprises 2660 pairs of images sized at 768×768 pixels, focusing primarily on building and land changes. Experiments demonstrate that our method outperforms existing state-of-the-art methods on two public CD datasets as well as UAV-CD, showcasing excellent performance. Tingfeng Xian, Zilu Ying, Haolin Lv, Yikui Zhai, Hao Quan 0002, Xudong Jia 0001 |
IGARSS | 2 |
| 2024 | Performance Evaluation of Anomaly Detection with a New Battery Surface Anomaly Dataset
Zilu Ying, Haolin Lv, Yingwen Chen 0002, Kanghong Tan |
PRCV (11) | 2 |
| 2024 | DGMA2-Net: A Difference-Guided Multiscale Aggregation Attention Network for Remote Sensing Change DetectionabstractRemote sensing change detection (RSCD) focuses on identifying regions that have undergone changes between two remote sensing images captured at different times. Recently, convolutional neural networks (CNNs) have shown promising results in the challenging task of RSCD. However, these methods do not efficiently fuse bitemporal features and extract useful information that is beneficial to subsequent RSCD tasks. In addition, they did not consider multilevel feature interactions in feature aggregation and ignore relationships between difference features and bitemporal features, which thus affects the RSCD results. To address the above problems, a difference-guided multiscale aggregation attention network, DGMA2-Net, is developed. Bitemporal features at different levels are extracted through a Siamese convolutional network and a multiscale difference fusion module (MDFM) is then created to fuse bitemporal features and extract, in a multiscale manner, difference features containing rich contextual information. After the MDFM treatment, two difference aggregation modules (DAMs) are used to aggregate difference features at different levels for multilevel feature interactions. The features through DAMs are sent to the difference-enhanced attention modules (DEAMs) to strengthen the connections between bitemporal features and difference features and further refine change features. Finally, refined change features are superimposed from deep to shallow and a change map is produced. In validating the effectiveness of DGMA2-Net, a series of experiments are conducted on three public RSCD benchmark datasets (LEVIR-CD, BCDD, and SYSU-CD). The experimental results demonstrate that DGMA2-Net surpasses the current eight state-of-the-art methods in RSCD. Our code is released at https://github.com/yikuizhai/DGMA2-Net. Zilu Ying, Zijun Tan, Yikui Zhai, Xudong Jia 0001, Wenba Li, Jun-Ying Zeng, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | DS-HyFA-Net: A Deeply Supervised Hybrid Feature Aggregation Network With Multiencoders for Change Detection in High-Resolution ImageryabstractWith the advancement of deep learning (DL) technologies, remarkable progress has been achieved in change detection (CD). Existing DL-based methods primarily focus on the discrepancy in bitemporal images, while overlooking the commonality in bitemporal images. However, one of the reasons hindering the improvement of CD performance is the inadequate utilization of image information. To address the above issue, we propose a Deeply Supervised Hybrid Feature Aggregation Network (DS-HyFA-Net). This network predicts changes by integrating the distinctness and the commonality in bitemporal images. Specifically, the DS-HyFA-Net primarily consists of a set of encoders and a Hybrid Feature Aggregation (HyFA) module. It uses a Siamese encoder (or Encoder I) and a specialized encoder (or Encoder II) to extract distinct and common features (CFs) in bitemporal images, respectively. The HyFA module efficiently aggregates distinct and common features (or hybrid features) and generates a change map using a predictor. In addition, a common feature learning strategy (CFLS) is introduced, based on deeply supervised (DS) techniques, to guide Encoder II in learning CFs. Experimental results on three well-recognized datasets demonstrate the effectiveness of the innovative DS-HyFA-Net, achieving F1-Scores of 93.33% on WHU-CD, 90.98% on LEVIR-CD, and 81.14% on SYSU-CD. Our code is available athttps://github.com/yikuizhai/DS-HyFA-Net. Zilu Ying, Tingfeng Xian, Yikui Zhai, Xudong Jia 0001, Hongsheng Zhang 0001, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Large-Scale High-Altitude UAV-Based Vehicle Detection via Pyramid Dual Pooling Attention Path Aggregation NetworkabstractUAVs can collect vehicle data in high-altitude scenes, playing a significant role in intelligent urban management due to their wide of view. Nevertheless, the current datasets for UAV-based vehicle detection are acquired at altitude below 150 meters. This contrasts with the data perspective obtained from high-altitude scenes, potentially leading to incongruities in data distribution. Consequently, it is challenging to apply these datasets effectively in high-altitude scenes, and there is an ongoing obstacle. To resolve this challenge, we developed a comprehensive vehicle dataset named LH-UAV-Vehicle, specifically collected at flight altitudes ranging from 250 to 400 meters. Collecting data at higher flight altitudes offers a broader perspective, but it concurrently introduces complexity and diversity in the background, which consequently impacts vehicle localization and recognition accuracy. In response, we proposed the pyramid dual pooling attention path aggregation network (PDPA-PAN), an innovative framework that improves detection performance in high-altitude scenes by combining spatial and semantic information. Object attention integration in both spatial and channel dimensions is aimed by the pyramid dual pooling attention module (PDPAM), which is achieved through the parallel integration of two distinct attention mechanisms. Furthermore, we have individually developed the pyramid pooling attention module (PPAM) and the dual pooling attention module (DPAM). The PPAM emphasizes channel attention, while the DPAM prioritizes spatial attention. This design aims to enhance vehicle information and suppress background interference more effectively. Extensive experiments conducted on the LH-UAV-Vehicle conclusively demonstrate the efficacy of the proposed vehicle detection method. Our code and dataset can be found at https://github.com/yikuizhai/PDPA-PAN. Zilu Ying, Yikui Zhai, Hao Quan 0002, Wenba Li, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | UAV-BCD: A UAV Building Change Detection DatasetabstractRemote sensing change detection (RSCD) holds significant prominence as a research topic within the realm of computer vision. However, previous RSCD datasets have been constructed based on satellite remote sensing images. Traditional satellite remote sensing images have problems such as insufficient resolution, difficult data acquisition, and complex processing processes, and there is a certain gap between data distribution and actual needs. UAVs not only have the advantages of flexibility and high-speed, but also can capture high-resolution images, which are especially suitable for high-precision RSCD in small areas. Therefore, this paper proposed a new UAV RSCD dataset — UAV Building Change Detection Dataset (UAV-BCD). The proposed dataset contains 2024 pairs of finely registered high-resolution images collected by UAVs and their corresponding pixel-level labels, which can provide a new benchmark for RSCD. We evaluate the effectiveness of UAV-BCD with the five state-of-art deep neural networks in RSCD. Zilu Ying, Zijun Tan, Wenba Li, Zhangzhao Liang, Yikui Zhai |
IGARSS | 1 |
| 2023 | SAS-NET: Similarity Attention Siamese Network for Building Change Detection in UAV ImagesabstractChange detection refers to extract change information using deep learning or traditional image processing methods to quantitatively analyze and characterize landmark changes on bi-temporal images. Currently, change detection is mainly a pixel-level task, and obtaining accurate change detection segmentation predictions requires a more elaborate and complex model architecture design. To simplify the change detection task, we proposed a novel similarity detection model, Similarity Attention Siamese Network (SAS-NET). It analyzed and predicted if the bi-temporal image patches were similar, and simplified pixel-level change detection tasks to patch-level similarity classification prediction tasks. In this work, a UAV Similarity Detection Dataset (UAV-SD) was also proposed to explore the advantages of patch-level prediction tasks over pixel-level change detection tasks. The proposed method achieved 90.5% accuracy on UAV-SD, which proves that it is more effective than other advanced change detection methods. Yikui Zhai, Wenba Li, Zijun Tan, Zilu Ying |
IGARSS | 6 |
| 2022 | Weakly Contrastive Learning via Batch Instance Discrimination and Feature Clustering for Small Sample SAR ATRabstractIn recent years, impressive performance of deep learning technology has been recognized in synthetic aperture radar (SAR) automatic target recognition (ATR). Since a large amount of annotated data are required in this technique, it poses a trenchant challenge to the issue of obtaining a high recognition rate through less labeled data. To overcome this problem, inspired by the contrastive learning, we proposed a novel framework named batch instance discrimination and feature clustering (BIDFC). In this framework, different from that of the objective of general contrastive learning methods, embedding distance between samples should be moderate because of the high similarity between samples in the SAR images. Consequently, our flexible framework is equipped with adjustable distance between embedding, which we term as weakly contrastive learning. Technically, instance labels are assigned to the unlabeled data in per batch, and random augmentation and training are performedfewtimes on these augmented data. Meanwhile, a novel dynamic-weighted variance loss (DWV loss) function is also posed to cluster the embedding of enhanced versions for each sample. The experimental results on the moving and stationary target acquisition and recognition (MSTAR) database indicate a 91.25% classification accuracy of our method fine-tuned on only 3.13% training data. Even though a linear evaluation is performed on the same training data, the accuracy can still reach 90.13%. We also verified the effectiveness of BIDFC in OpenSarShip database, indicating that our method can be generalized to other data sets. Our code is available at:https://github.com/Wenlve-Zhou/BIDFC-master. Yikui Zhai, Wenlve Zhou, Bing Sun 0002, Jingwen Li 0003, Qirui Ke, Zilu Ying, Junying Gan, Chaoyun Mai, Ruggero Donida Labati, Vincenzo Piuri, Fabio Scotti |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2017 | SAR Automatic Target Recognition Based on Deep Convolutional Neural Network
Ying Xu 0005, Kaipin Liu, Zilu Ying, Lijuan Shang, Yikui Zhai, Vincenzo Piuri, Fabio Scotti |
ICIG (3) | 3 |
| 2017 | Deep Convolutional Neural Network for Facial Expression Recognition
Yikui Zhai, Jun-Ying Zeng, Vincenzo Piuri, Fabio Scotti, Zilu Ying, Ying Xu 0005, Junying Gan |
ICIG (1) | 6 |
| 2009 | Facial Expression Recognition with Local Binary Pattern and Laplacian Eigenmaps
Zilu Ying, Linbo Cai, Junying Gan, Sibin He |
ICIC (1) | 1 |