VLDB 2026 Research / reviewers in the wild / expert
Yaozong Zhang
dblp:210/6714
· DBLP profile ↗
19ranked-venue papers
3as first author
16since 2021 · last 2025
0000-0002-3727-7957ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Lightweight and Real-Time Asymmetric Multi-output Thermal Radiation Effects Correction in Infrared Images
Dongming Xie, Xinyuan Deng, Yaozong Zhang |
ICIG (2) | 6 |
| 2025 | FUT: Frequency-aware U-shaped transformer for image denoising
Yaozong Zhang, Zhenghua Huang |
J. Vis. Commun. Image Represent. | 1 |
| 2025 | RUST: Residual U-shaped transformer to approximate Taylor expansion for image denoising
Zhenghua Huang, Yu Shi 0004, Yaozong Zhang |
Knowl. Based Syst. | 6 |
| 2025 | DMSDA-YOLO: Dynamic Multiscale Dilated Attention for Remote Sensing Object DetectionabstractIt is an extremely challenging task to detect multiscale targets (especially small objects) in remote sensing (RS) images with complex backgrounds. This letter develops a novel RS object detection model, namely dynamic multiscale dilated attention based on YOLOv5 (DMSDA-YOLO), of which the key improvements include: one is that, in the backbone, a multiscale dilated attention fusion module (MDAFM) is proposed to capture multiscale feature information and a coordinate anchor attention (CAA) mechanism is incorporated to increase the focus on target regions while suppressing background interference. The other is that a spatial attention pyramid neck network is proposed to improve its feature fusion capability while a dynamic attention-aware feature extraction module (DAFEM) is introduced to enhance the network’s adaptability to multiscale targets in the neck. Objective and subjective results of experiments on the DIOR, HRRSD, and NWPU VHR-10 datasets demonstrate that our DMSDA-YOLO outperforms existing state-of-the-art object detection approaches in detecting multiscale targets under complex backgrounds, and its competitive computational complexity is beneficial for its extensive application. Zhenghua Huang, Zijian Xu 0010, Yaozong Zhang, Yu Shi 0004 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | Multimodal Remote Sensing Sparse Registration With a Global-Local DescriptorabstractMultimodal image registration is a key procedure in remote sensing applications (such as remote sensing image stitching), which faces significant challenges including radiometric discrepancies and local geometric deformations caused by the differences of both sensor and imaging parameters. Traditional methods remove coarse error using global features, making it difficult to identify misregistrations at early stage, thus limiting registration accuracy improvement. When existing convolutional registration neural networks extract deep features, shallow local feature information is usually lost because the network gradually focuses on high-level abstract features, causing local details to be simplified or lost in the global feature construction. Solving this problem will greatly increase the complexity of the model, and the network needs to reorganize and train the data according to specific tasks, which is time-consuming. To address these issues, this letter develops a hybrid registration model with a global-local descriptor. Specifically, we first obtain improved RIFT keypoints via combining rotated and scale invariant corner points produced by the integral scale detection Min-moment with extracted edge points generated by the FAST detection Max-moment. Then, a global-local descriptor is constructed by combining the improved RIFT descriptor with the LoFTR coarse-grained feature descriptor. Finally, a 0–1 distance allocation matrix is formulated to improve the registration success rate (SR). The experimental results show that the proposed method has a powerful capability in improving both generalization and accuracy and outperforms mainstream methods, even the average number of correctly registered correspondences is about two times and 1.7 times higher than LoFTR and RIFT, respectively. Yaozong Zhang, Yuanyin Lei, Ying Zhu 0002, Lei Wang 0068, Hanyu Hong, Zhenghua Huang |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2025 | RCST: Residual Context-Sharing Transformer Cascade to Approximate Taylor Expansion for Remote Sensing Image DenoisingabstractTaylor expansion is a polynomial for approximating a function constructed by the coefficients of its derivatives at a certain point, where it is a challenging research to utilize the powerful learning ability of deep learning (DL) to characterize the polynomial parts for pursuing its approximate solution. In this article, we develop a cascading residual context-sharing Transformer (RCST) to approximate Taylor expansion for remote sensing (RS) image denoising. Our RCST method includes the following key procedures. First, a mapping function about a latent clean RS image patch is built by employing the low-rank characteristic of its neighborhood RS image blocks, and is expanded into a polynomial with Taylor expansion for its approximate solution. Second, the intrinsic recursive relationship of the neighborhood derivatives is analyzed and is mathematically formulated, which provides a theoretical interpretability for the construction of our RCST model. Third, a lightweight residual network (LRNet) is developed to estimate the base layer, while the RCST is shared to calculate the derivative parts. Finally, to transfer as many rich multiscale details from noisy RS images to estimated results as possible, we adopt a down-/upsampling architecture. Specifically, a spatial-Fourier upsampling (SFUS) operator is reported to preserve both local and global information. Quantitatively and qualitatively experimental results validate that our RCST denoising method can achieve competitive performance and is even superior to other SOTA denoising approaches. Zhenghua Huang, Yu Shi 0004, Yaozong Zhang |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | Corrections to "Semi-Supervised Learning for Infrared Thermal Radiation Correction in the Real World"
Yu Shi 0004, Xinyuan Deng, Lei Wang 0068, Yaozong Zhang, Zhenghua Huang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | RSTC: Residual Swin Transformer Cascade to approximate Taylor expansion for image denoising
Biyun Xu, Yaozong Zhang, Zhenghua Huang |
Comput. Vis. Image Underst. | 5 |
| 2024 | Semi-Supervised Learning for Infrared Thermal Radiation Correction in the Real WorldabstractInfrared images are susceptible to thermal radiation. Infrared thermal radiation correction methods based on physical prior may fail while correcting real-world images, because assumed priors do not always hold in the real world, resulting in the presence of thermal radiation residuals. Supervised learning-based methods have the potential to achieve favorable outcomes in the correction of synthetic images. However, due to the unavailability of labeled datasets, their efficacy is limited when applied to real-world images. To address this problem, in this article, to the best of our knowledge, we propose the first semi-supervised learning network for infrared radiation correction in the real world, named SIRCNet. The network is trained using a semi-supervised strategy, which includes a supervised training stage and a self-supervised training stage. In the supervised training stage, we constructed a multilevel wavelet decomposition and reconstruction correction (MWDRC) module for latent image correction and an efficient generalized feature extraction (EGFE) module for bias field estimation. Furthermore, EGFE is composed of one partial channel interactive (PCI) attention block and three effective residual blocks (ERBs). Surface fitting can approximate the thermal radiation bias field of the thermal radiation degradation images. The fit bias field can provide critical prior knowledge that enhances EGFE’s estimation of the thermal radiation bias field. Hence, in the self-supervised training stage, when fine-tuning MWDRC and EGFE using a generator, surface fitting is employed to constrain EGFE. Comparative experiments demonstrate that SIRCNet outperforms existing correction methods on both real and synthetic datasets, achieving the best metrics as well as visualization. Yu Shi 0004, Xinyuan Deng, Lei Wang 0068, Yaozong Zhang, Zhenghua Huang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Unsupervised Encoder-Decoder Model for Anomaly Prediction Task
Jinmeng Wu, Pengcheng Shu, Hanyu Hong, Xingxun Li, Lei Ma 0004, Yaozong Zhang, Ying Zhu 0002, Lei Wang 0068 |
MMM (2) | 6 |
| 2023 | Scribble-attention hierarchical network for weakly supervised salient object detection in optical remote sensing images
Lei Ma 0004, Hanyu Hong, Yaozong Zhang, Lei Wang 0068, Jinmeng Wu |
Appl. Intell. | 4 |
| 2023 | DGDNet: Deep Gradient Descent Network for Remotely Sensed Image DenoisingabstractGradient descent strategy, viewed as an important model optimization method, has been widely used for various tasks (such as model-based image denoising) of computer vision. In the gradient descent denoising model, the learning rate (LR) and residual component are two important parts to be adaptively estimated for its stable point. This letter proposes a deep gradient descent network (DGDNet), including two key points: one is that the LR is designed with eigenvalues of Hessian matrix of remotely sensed images (RSIs) and their local weighted factor (LWF), which can recognize structures from RSIs degraded by additive white Gaussian noise (AWGN). The other is that the residual part is calculated by an U-shaped network (USNet) to speed up the DGDNet convergent to a fixed point. Finally, the two components are plugged into the gradient descent scheme and contribute to an enjoyable result with a few iterations. Quantitatively and qualitatively experimental results demonstrate that the proposed DGDNet can obtain a stable solution efficiently, and produce competitive denoising performance which is even better than that yielded by the state-of-the-art noise reduction methods. Zhenghua Huang, Zifan Zhu, Zhicheng Wang 0004, Yu Shi 0004, Yaozong Zhang |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Semi-Supervised Semantic Segmentation of SAR Images Based on Cross Pseudo-SupervisionabstractDue to the unique imaging mechanism and wide application of synthetic aperture radar (SAR), SAR image interpretation has been researched by more and more scholars. The supervised SAR image semantic segmentation methods that based on deep learning require a large number of accurate pixel-level labels, which are very hard to obtain. The lack of labeled samples limits the practical application of deep learning methods in SAR image semantic segmentation. To reduce the requirement of labeled data, we decided to introduce the cross pseudo-supervision network (CPS-Net) into SAR image semi-supervised semantic segmentation and promote the development of semi-supervised learning in SAR image interpretation. The semi-supervised segmentation based on CPS-Net has the following advantages: (1) CPS-Net encourages high similarity between two networks with the same input data, which helps improve the performance. (2) CPS-Net can make better use of the pseudo-supervision of unlabeled data to guide the network training. Experimental results show that CPS-Net achieves excellent semi-supervised semantic segmentation results on Sentinel-1 dual-polarization data with less labeled data. Compared with well-known semantic segmentation methods U-Net and DeeplabV3+, the performance of SAR image segmentation is significantly improved. Hanyu Hong, Ying Zhu 0002, Yaozong Zhang, Pengtian Wang, Lei Wang 0068 |
IGARSS | 4 |
| 2022 | Quantitative Evaluation of Multi-Sensor Image Registraction Feature DescriptorabstractMulti-sensor image registration is a basic and important issue in the field of remote sensing applications. At present, many algorithms have not directly evaluated and analyzed the feature descriptor design of the algorithm. Taking the feature descriptors of RIFT, SIFT, SAR-SIFT and HAPCG as the analysis objects, this paper designs experiments to analyze their stability under gray distortion and local geometric distortion, gives a quantitative evaluation, and reveals the contribution of the feature descriptor of each multi-sensor image registration algorithm in the process of multi-sensor image registration. Yaozong Zhang, Zhenghua Huang, Lei Wang 0068, Ying Zhu 0002, Hanyu Hong |
IGARSS | 1 |
| 2022 | DLRP: Learning Deep Low-Rank Prior for Remotely Sensed Image DenoisingabstractRemotely sensed images degraded by additive white Gaussian noise (AWGN) are not beneficial for the analysis of their contents. Such a phenomenon is usually modeled as an inverse problem which can be solved by model-based optimization methods or discriminative learning approaches. The former pursue their pleasing performance at the cost of a highly computational burden while the latter are impressive for their fast testing speed but are limited by their application range. To join their merits, this letter proposes a nonlocal self-similar (NSS) block-based deep image denoising scheme, namely deep low-rank prior (DLRP), which includes the following key points: First, the low-rank property of the neighboring NSS patches ordered lexicographically is utilized to model a global objective function (GOF). Second, with the aid of an alternative iteration strategy, the GOF can be easily decomposed into two independent subproblems. One is a quadratic optimization problem, and has a closed-form solution. While the other is a low-rank minimization denoising problem and is learned by deep convolutional neural network (DCNN). Then, the deep denoiser, acted as a modular part, is plugged into the model-based optimization method with adaptive noise level estimation to solve the inverse problem. In the experiments, we first discuss parameter setting and the convergence. Then, quantitative/qualitative comparisons of experimental results validate that the DLRP is a flexible and powerful denoising method to achieve competitive performance which even outperforms those produced by state-of-the-arts. Zhenghua Huang, Zhicheng Wang 0004, Zifan Zhu, Yaozong Zhang, Yu Shi 0004, Tianxu Zhang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | PolSAR-SSN: An End-to-End Superpixel Sampling Network for PolSAR Image ClassificationabstractPolarimetric synthetic aperture radar (PolSAR) image classification is one of the fundamental research areas in remote sensing. Superpixels can provide boundary constraint information and are widely used in PolSAR image interpretation. However, traditional machine learning superpixel algorithms have many limitations for PolSAR image interpretation. Pseudo-color images are usually used as the superpixel algorithm inputs, and the loss of polarimetric information will decrease the performance. In addition, the superpixel algorithms are difficult to incorporate into state-of-the-art deep learning models and cannot be trained in an end-to-end manner. In this letter, a trainable end-to-end deep superpixel network is proposed for PolSAR image classification. The inputs of the proposed method can be any low/middle-level polarimetric features of a PolSAR image and the rich polarimetric feature representation can be learned. The produced superpixels of the proposed method are more concentrated near the land cover boundaries and can significantly improve the performance of PolSAR image classification. Experimental results show that the overall accuracies of the proposed method are approximately 2.57% and 1.44% higher than traditional superpixel algorithms on two PolSAR datasets and surpass some well-known deep learning methods. Lei Wang 0068, Hanyu Hong, Yaozong Zhang, Jinmeng Wu, Lei Ma 0004, Ying Zhu 0002 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2020 | Joint Analysis and Weighted Synthesis Sparsity Priors for Simultaneous Denoising and Destriping Optical Remote Sensing ImagesabstractStripe and random noise are two different degradation phenomena that commonly coexist in optical remote sensing images, and they are often modeled as inverse problems. In model-based inverse problems, analysis and synthesis sparse representations (SSRs) are used as regularization terms to obtain approximate solutions due to their respective merits, i.e., the nonzero coefficients in SSR are usually used to describe an image, while the indexes of zeros in analysis sparse representation (ASR) are used to characterize the stripe. Inspired by these merits, we propose a unified variational framework, called a joint analysis and weighted synthesis (JAWS) sparsity model, to simultaneously separate the clean image and the stripe from a single optical remote sensing image. To solve the JAWS sparsity model efficiently, an alternating minimization optimization strategy is first employed to separate it into two subproblems that are used for different tasks. One called as weighted SSR (WSSR) is the main for optical remote sensing image denoising, which can be effectively solved by employing the weighted singular value thresholding operator, while the other called as ASR is the main approach for optical remote sensing image destriping, which is optimized by adopting the split Bregman iteration. By minimizing the two subproblems alternatively, the proposed JAWS sparsity model is efficiently solved. Finally, both quantitative and qualitative results of experiments on synthetic and real-world optical remote sensing images validate that the proposed approach is effective and even better than the state of the arts. Zhenghua Huang, Yaozong Zhang, Qian Li 0019, Tianxu Zhang, Nong Sang, Hanyu Hong |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Principal component dictionary-based patch grouping for image denoising
Shoukui Yao, Yi Chang 0002, Xiaojuan Qin, Yaozong Zhang, Tianxu Zhang |
J. Vis. Commun. Image Represent. | 4 |
| 2018 | Progressive Dual-Domain Filter for Enhancing and Denoising Optical Remote-Sensing ImagesabstractEnhancement and denoising have always been a pair of conflicting problems in image processing of computer vision. Inspired by an earlier dual-domain filter (DDF), this letter proposes a progressive DDF to simultaneously enhance and denoise low-quality optical remote-sensing images. The main procedure of the proposed enhancement filter has two parts. First, a bilateral filter is exploited as a guide filter to obtain high-contrast images, which are enhanced by a histogram modification method. Then, low-contrast useful structures are restored by a short-time Fourier transform and are enhanced using an adaptive correction parameter. Both the quantitative and qualitative results of experiments on synthetic and real-world low-quality remote-sensing images demonstrate that the proposed method performs well on contrast enhancement, structure preservation, and noise reduction. Moreover, its satisfactory computation time resulting from its simple implementation makes it suitable for extensive application. Zhenghua Huang, Yaozong Zhang, Qian Li 0019, Tianxu Zhang, Nong Sang, Hanyu Hong |
IEEE Geosci. Remote. Sens. Lett. | 2 |