Ying Zhu 0002

dblp:01/2082-2 · DBLP profile ↗
← Back
16ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0002-5708-3252ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Unsupervised deep hashing based on multi-scale aggregation and optimal transport matching for image retrieval
Lei Ma 0004, Hao Pei, Lei Wang 0068, Ying Zhu 0002, Yu Shi 0004, Hanyu Hong, Xinyu Dai, Fanman Meng, Qingbo Wu 0001
Neurocomputing4
2025 Multimodal Remote Sensing Sparse Registration With a Global-Local Descriptor
abstract
Multimodal image registration is a key procedure in remote sensing applications (such as remote sensing image stitching), which faces significant challenges including radiometric discrepancies and local geometric deformations caused by the differences of both sensor and imaging parameters. Traditional methods remove coarse error using global features, making it difficult to identify misregistrations at early stage, thus limiting registration accuracy improvement. When existing convolutional registration neural networks extract deep features, shallow local feature information is usually lost because the network gradually focuses on high-level abstract features, causing local details to be simplified or lost in the global feature construction. Solving this problem will greatly increase the complexity of the model, and the network needs to reorganize and train the data according to specific tasks, which is time-consuming. To address these issues, this letter develops a hybrid registration model with a global-local descriptor. Specifically, we first obtain improved RIFT keypoints via combining rotated and scale invariant corner points produced by the integral scale detection Min-moment with extracted edge points generated by the FAST detection Max-moment. Then, a global-local descriptor is constructed by combining the improved RIFT descriptor with the LoFTR coarse-grained feature descriptor. Finally, a 0–1 distance allocation matrix is formulated to improve the registration success rate (SR). The experimental results show that the proposed method has a powerful capability in improving both generalization and accuracy and outperforms mainstream methods, even the average number of correctly registered correspondences is about two times and 1.7 times higher than LoFTR and RIFT, respectively.
Yaozong Zhang, Yuanyin Lei, Ying Zhu 0002, Lei Wang 0068, Hanyu Hong, Zhenghua Huang
IEEE Geosci. Remote. Sens. Lett.3
2025 Progressive Learning-Based Jitter Distortion Correction for Remote Sensing Images of Time Delay and Integration Camera
abstract
The widespread use of time delay and integration charge-coupled device (TDI CCD) technology in high-resolution spaceborne optical cameras has made high-frequency jitter effects a common issue, resulting in different levels of distortion in images. Current methods mostly concentrate on correction of obviously high levels of geometric distortion. Focusing on low levels of geometric distortion, which are more difficult to accurately detect, this paper proposes a progressive learning-based correction method for high-frequency jitter distortion in remote sensing images from spaceborne TDI CCD cameras, utilizing a Generative Adversarial Network (GAN). First, a distorted dataset with diverse jitter levels for progressive training is generated through jitter simulation model by adjusting the parameters. Then, a GAN model is employed for the correction task. The generator consists of the Distortion Net for geometric distortion correction and the Detail Enhancement Net for image detail restoration. Finally, a progressive learning strategy is used to gradually enhance the ability of network to correct minor geometric distortion. The proposed method is validated using simulated images and real-world satellite images. Experimental results demonstrate that the proposed method outperforms existing restoration methods both in simulated datasets and practical scenarios.
Ying Zhu 0002, Mi Wang, Jun Pan 0001, Hanyu Hong, Lei Ma 0004, Lei Wang 0068
IEEE Trans. Geosci. Remote. Sens.1
2024 Generative Adversarial Network-Based Jitter Distortion Correction for High Resolution Spaceborne Images
abstract
This paper presents a Generative Adversarial Network (GAN)-based jitter distortion correction method for spaceborne images of Time Delay Integration (TDI) Charge-Coupled Device (CCD) camera. This method leverages the advantages of GANs and combines content loss, adversarial loss, and perceptual loss to effectively repair distorted images while preserving image details, which does not rely on jitter information captured by high-frequency attitude sensors, nor depends on the analysis of overlapping areas between different bands in multispectral images. The experimental results show that the proposed method achieves automated correction of geometric distortions and has shown promising restoration results on real distorted images captured by Yaogan-26 satellite and GaoFen satellite, which achieves better results than other blind restoration methods.
Ying Zhu 0002, Lei Wang 0068, Lei Ma 0004, Jinmeng Wu
IGARSS2
2024 Rigorous Parallax Observation Model-Based Remote Sensing Panchromatic and Multispectral Images Jitter Distortion Correction for Time Delay Integration Cameras
abstract
Time delay integration charge-coupled device (TDI CCD) is sensitive to the platform’s stability during push-broom imaging. Due to variations in total integration time, panchromatic and multispectral images suffer varying degrees of geometric distortion caused by satellite jitter with high frequency, which leads to different inner distortion in different band images and different band-to-band mismatching errors between different band combinations. To address this problem, this paper proposes a rigorous parallax observation model considering multi-stage integration time and presents a jitter distortion correction method for remote sensing panchromatic and multispectral images captured by TDI cameras based on it. First, the law of the amplitude attenuation and phase offset of platform jitter deviation on the image under different TDI stages is determined through simulation verification. Then, the rigorous parallax observation model is proposed to establish an accurate relationship between the relative jitter error of two multispectral images with multi-stage integration and the absolute single-stage integration jitter error by introducing the amplitude attenuation factor and phase offset. Finally, the jitter distortion curves of images with different integration stages and integration time can be reconstructed based on the estimated absolute jitter error and the imaging parameters. Subsequently, the jitter distortion can be further corrected by image resampling. The proposed method was verified through both simulation and real data experiments using GaoFen-9 satellite images. Experimental results show that the proposed method can effectively correct high-frequency jitter distortion in panchromatic and multispectral images, which cannot be corrected by traditional single-stage integration jitter detection model.
Ying Zhu 0002, Mi Wang, Jun Pan 0001, Guo Ye, Hanyu Hong, Lei Wang 0068
IEEE Trans. Geosci. Remote. Sens.1
2023 Unsupervised Encoder-Decoder Model for Anomaly Prediction Task
Jinmeng Wu, Pengcheng Shu, Hanyu Hong, Xingxun Li, Lei Ma 0004, Yaozong Zhang, Ying Zhu 0002, Lei Wang 0068
MMM (2)7
2023 Joint ordinal regression and multiclass classification for diabetic retinopathy grading with transformers and CNNs fusion network
Lei Ma 0004, Qihang Xu, Hanyu Hong, Yu Shi 0004, Ying Zhu 0002, Lei Wang 0068
Appl. Intell.5
2023 Complementary Parts Contrastive Learning for Fine-Grained Weakly Supervised Object Co-Localization
abstract
The aim of weakly supervised object co-localization is to locate different objects of the same superclass in a dataset. Recent methods achieve impressive co-localization performance by multiple instance learning and self-supervised learning. However, these methods ignore the common part information shared by fine-grained objects and the influence of the complementary parts on the co-localization of the fine-grained objects. To solve these issues, we propose a complementary parts contrastive learning method for fine-grained weakly supervised object co-localization. The proposed method follows such an assumption that fine-grained object parts with the same/different semantic meaning should have similar/dissimilar feature representations in the feature space. The proposed method tackles two critical issues in this task:$i)$how to spread the model’s attention and suppress the complex background noise, and$ii)$how to leverage the cross-category common parts information to mitigate the context co-occurrence problem. To address$i)$, we attempt to integrate local and context cues via three types of attention including self-supervised attention, channel, and spatial attention to spread the model’s attention toward automatically identifying and localizing most discriminative parts of objects in the fine-grained images. To solve$ii)$, we propose a cross-category object complementarity part contrastive learning module to identify the extracted part regions with different semantic information by pulling the same part features closer and pushing different part features away, which can mitigate the confounding bias caused by the co-occurrence surroundings within specific classes. Extensive qualitative and quantitative evaluations demonstrate the effectiveness of the proposed method on four fine-grained co-localization datasets: CUB-200–2011, Stanford Cars, FGVC-Aircraft, and Stanford Dogs. Code and models are available athttps://github.com/Zhao-fan/CPCL.
Lei Ma 0004, Hanyu Hong, Lei Wang 0068, Ying Zhu 0002
IEEE Trans. Circuits Syst. Video Technol.5
2022 Semi-Supervised Semantic Segmentation of SAR Images Based on Cross Pseudo-Supervision
abstract
Due to the unique imaging mechanism and wide application of synthetic aperture radar (SAR), SAR image interpretation has been researched by more and more scholars. The supervised SAR image semantic segmentation methods that based on deep learning require a large number of accurate pixel-level labels, which are very hard to obtain. The lack of labeled samples limits the practical application of deep learning methods in SAR image semantic segmentation. To reduce the requirement of labeled data, we decided to introduce the cross pseudo-supervision network (CPS-Net) into SAR image semi-supervised semantic segmentation and promote the development of semi-supervised learning in SAR image interpretation. The semi-supervised segmentation based on CPS-Net has the following advantages: (1) CPS-Net encourages high similarity between two networks with the same input data, which helps improve the performance. (2) CPS-Net can make better use of the pseudo-supervision of unlabeled data to guide the network training. Experimental results show that CPS-Net achieves excellent semi-supervised semantic segmentation results on Sentinel-1 dual-polarization data with less labeled data. Compared with well-known semantic segmentation methods U-Net and DeeplabV3+, the performance of SAR image segmentation is significantly improved.
Hanyu Hong, Ying Zhu 0002, Yaozong Zhang, Pengtian Wang, Lei Wang 0068
IGARSS3
2022 Quantitative Evaluation of Multi-Sensor Image Registraction Feature Descriptor
abstract
Multi-sensor image registration is a basic and important issue in the field of remote sensing applications. At present, many algorithms have not directly evaluated and analyzed the feature descriptor design of the algorithm. Taking the feature descriptors of RIFT, SIFT, SAR-SIFT and HAPCG as the analysis objects, this paper designs experiments to analyze their stability under gray distortion and local geometric distortion, gives a quantitative evaluation, and reveals the contribution of the feature descriptor of each multi-sensor image registration algorithm in the process of multi-sensor image registration.
Yaozong Zhang, Zhenghua Huang, Lei Wang 0068, Ying Zhu 0002, Hanyu Hong
IGARSS5
2022 MINet: Multilevel Inheritance Network-Based Aerial Scene Classification
abstract
Scene classification of aerial images is the basis of automatic recognition of complex scenes, and it is also a challenging computer vision task. In recent years, with the rapid development of deep learning, the semantic feature extraction method based on a convolutional neural network (CNN) has made great progress. Moreover, a recent study indicates that combining the semantic information of deep-layer features with the detailed texture information of shallow-layer features in CNN can further improve the performance of classification. In this letter, an end-to-end multilevel feature-based network named multilevel inheritance network (MINet) is proposed for aerial scene classification. First, the feature extraction module based on the feature pyramid network (FPN) is used to get multilevel feature maps. In the process of merging shallow features, high-level semantics of deep-layer are inherited. Then, an attention mechanism is added after the multilevel features to reduce the interference of redundant information and noise. Finally, we use a feature fusion module to automatically learn the weight of each feature layer and make a comprehensive decision. The effectiveness of the proposed method is verified in AID, WHU-RS19 and NWPU-RESISC45 datasets. Results show that the proposed method achieves competitive classification accuracy.
Jiarui Hu 0001, Qidi Shu, Jun Pan 0001, Jianguang Tu, Ying Zhu 0002, Mi Wang
IEEE Geosci. Remote. Sens. Lett.5
2022 PolSAR-SSN: An End-to-End Superpixel Sampling Network for PolSAR Image Classification
abstract
Polarimetric synthetic aperture radar (PolSAR) image classification is one of the fundamental research areas in remote sensing. Superpixels can provide boundary constraint information and are widely used in PolSAR image interpretation. However, traditional machine learning superpixel algorithms have many limitations for PolSAR image interpretation. Pseudo-color images are usually used as the superpixel algorithm inputs, and the loss of polarimetric information will decrease the performance. In addition, the superpixel algorithms are difficult to incorporate into state-of-the-art deep learning models and cannot be trained in an end-to-end manner. In this letter, a trainable end-to-end deep superpixel network is proposed for PolSAR image classification. The inputs of the proposed method can be any low/middle-level polarimetric features of a PolSAR image and the rich polarimetric feature representation can be learned. The produced superpixels of the proposed method are more concentrated near the land cover boundaries and can significantly improve the performance of PolSAR image classification. Experimental results show that the overall accuracies of the proposed method are approximately 2.57% and 1.44% higher than traditional superpixel algorithms on two PolSAR datasets and surpass some well-known deep learning methods.
Lei Wang 0068, Hanyu Hong, Yaozong Zhang, Jinmeng Wu, Lei Ma 0004, Ying Zhu 0002
IEEE Geosci. Remote. Sens. Lett.6
2021 Jitter Detection and Image Restoration Based on Continue Dynamic Shooting Model for High-Resolution TDI CCD Satellite Images
abstract
Although time delay integration charge-coupled devices (TDI CCDs) have been widely used in high-resolution spaceborne optical cameras, they are sensitive to satellite jitter: the images obtained by them are affected by both distortion and blur. Therefore, according to the multistage integral imaging characteristics of TDI CCDs, this article not only proposes a continue dynamic shooting model (CDSM) to reflect the real push-broom mode of the satellite but also presents a method containing jitter detection and image restoration based on it. In the presented method, the CDSM subdivides the TDI CCD integration intervals. The subdivision number of CDSM is determined by the proposed integral transformation function (ITF). Then, it feeds back into the ITF and also contributes to the point spread function (PSF) estimation. Among the abovementioned, ITF defines the relationship between the parallax images and the jitter curve, and aims to improve the jitter detection performance. Finally, an adaptive image restoration based on context is conducted, which combines time, space, and spectrum information. Besides the simulated images, multispectral images of GaoFen-1 02 satellite were also adopted to validate the performance of the presented method. Experimental results indicate that the accuracy of the jitter detection is increased, and the geometric and radiometric qualities of restored images are also improved.
Jun Pan 0001, Guo Ye, Ying Zhu 0002, Fen Hu, Mi Wang
IEEE Trans. Geosci. Remote. Sens.3
2020 Atmospheric Refraction Calibration of Geometric Positioning for Optical Remote Sensing Satellite
abstract
Owing to the effects of atmospheric refraction, the path of light propagation is bent, making the three-point collinear principle inapplicable and influencing the geometric accuracy of high-resolution optical satellite geometric positioning. This letter presents a novel geometric positioning method with atmospheric refraction calibration for optical remote sensing satellites. The atmospheric ellipsoid model is established using the measured atmospheric parameters and accurately describes the shape and characteristics of the real atmosphere. With iterative processing of geometric positioning and atmospheric refraction calibration, the path of light propagation in atmospheric ellipsoids is calibrated, and the real coordinates of the ground object are positioned accurately. With the advantages of simplicity and independence of sensors, the rational function model with atmospheric refraction calibration is proposed to achieve geometric positioning with higher geometric accuracy. Experimental results demonstrate that the proposed model can calibrate atmospheric refraction error and improve the geometric accuracy of the optical imagery with a large view angle. Furthermore, compared with the refraction index using the measurement data, it is proven that atmospheric refraction calibration should be implemented according to the measured atmospheric parameters during imaging instead of using the empirical model.
Ying Zhu 0002, Mi Wang, Shuying Jin, Qilong Rao
IEEE Geosci. Remote. Sens. Lett.2
2020 Object Detection in High Resolution Remote Sensing Imagery Based on Convolutional Neural Networks With Suitable Object Scale Features
abstract
Object detection in high spatial resolution remote sensing images (HSRIs) is an important part of image information automatic extraction, analysis, and understanding. The region of interest (ROI) scale of object detection and the object feature representation are two vital factors in HSRI object detection. With respect to these two issues, this article presents a novel HSRI object detection method based on convolutional neural networks (CNNs) with suitable object scale features. First, the suitable ROI scale of object detection is obtained by compiling statistics for the scale range of objects in HSRIs. Then, a CNN framework for object detection in HSRIs is designed using a suitable ROI scale of object detection. The object features obtained using a CNN have good universality and robustness. Finally, a CNN framework with a suitable ROI scale of object detection is trained and tested. Using the WHU-RSONE data set, the proposed method is compared with the faster region-based CNN (Faster-RCNN) framework. The experimental results show that the proposed method outperforms the Faster-RCNN framework and provides good object detection results in HSRIs.
Mi Wang, Ying Zhu 0002
IEEE Trans. Geosci. Remote. Sens.4
2013 An automatic accuracy evaluation approach of band registration for multi-spectral imagery
abstract
Considering band misalignment caused by attitude jittering or other factors, band registration becomes the most critical pre-processing step for multispectral imagery as the registration result will directly influence the following applications. So band registration accuracy evaluation is necessary before registered imagery going through the next processing step. This paper proposes an automatic approach to evaluate the band registration accuracy for multi-spectral imagery. The proposed method is based on the theory of image matching, which includes three main steps: 1) feature points detection, and then 2) corresponding points matching, and 3) accuracy evaluation. Experiments are designed for the validation of the proposed approach with RGB image of Toronto in Canada captured by the Microsoft Vexcel's UltraCam-D (UCD) camera, in which quantitative analyses are applied to assess the accuracy and reliability of the method. And evaluation result for multi-spectral images of satellite, i.e., ZiYuan-3, by this method were presented. The result shows that the proposed method can automatic evaluate band registration accuracy for multi-spectral imagery accurately, efficiently and objectively.
Ying Zhu 0002, Mi Wang, Jun Pan 0001
IGARSS1