EDBT 2026 Demo / reviewers in the wild / expert
Yu Shi 0004
dblp:55/4736-4
· DBLP profile ↗
21ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0002-8511-2110ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unsupervised deep hashing based on multi-scale aggregation and optimal transport matching for image retrieval
Lei Ma 0004, Hao Pei, Lei Wang 0068, Ying Zhu 0002, Yu Shi 0004, Hanyu Hong, Xinyu Dai, Fanman Meng, Qingbo Wu 0001 |
Neurocomputing | 5 |
| 2025 | RUST: Residual U-shaped transformer to approximate Taylor expansion for image denoising
Zhenghua Huang, Yu Shi 0004, Yaozong Zhang |
Knowl. Based Syst. | 5 |
| 2025 | DMSDA-YOLO: Dynamic Multiscale Dilated Attention for Remote Sensing Object DetectionabstractIt is an extremely challenging task to detect multiscale targets (especially small objects) in remote sensing (RS) images with complex backgrounds. This letter develops a novel RS object detection model, namely dynamic multiscale dilated attention based on YOLOv5 (DMSDA-YOLO), of which the key improvements include: one is that, in the backbone, a multiscale dilated attention fusion module (MDAFM) is proposed to capture multiscale feature information and a coordinate anchor attention (CAA) mechanism is incorporated to increase the focus on target regions while suppressing background interference. The other is that a spatial attention pyramid neck network is proposed to improve its feature fusion capability while a dynamic attention-aware feature extraction module (DAFEM) is introduced to enhance the network’s adaptability to multiscale targets in the neck. Objective and subjective results of experiments on the DIOR, HRRSD, and NWPU VHR-10 datasets demonstrate that our DMSDA-YOLO outperforms existing state-of-the-art object detection approaches in detecting multiscale targets under complex backgrounds, and its competitive computational complexity is beneficial for its extensive application. Zhenghua Huang, Zijian Xu 0010, Yaozong Zhang, Yu Shi 0004 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2025 | Optimal Transport Quantization Based on Cross-X Semantic Hypergraph Learning for Fine-Grained Image RetrievalabstractLarge-scale fine-grained image retrieval aims to learn compact discriminative feature representations based on mining the subtle distinctions between visually similar objects. However, existing fine-grained image retrieval methods focus on enhancing the attention to the discriminative regions within single images, which barely exploit the high-order relational information between the global features and local region features across different images. Thus, the over-fitting problem of complex personalized differences cannot be effectively solved. In addition, existing unconstrained vector quantization methods tend to assign unquantized feature vectors to a few major codewords, which are unable to effectively distinguish the quantized features and reduce the redundant information. To address these issues, we propose a novel optimal transport quantization method based on cross-X semantic hypergraph learning for large-scale fine-grained image retrieval. Specifically, we first introduce a cross-layer multi-scale aggregation module to extract the global features and local region features. Subsequently, we build a semantic hypergraph to model the high-order correlations between the global features and local region features extracted from different layers, different scales and different images, which can alleviate the over-fitting problem of complex personalized differences by suppressing sample-level and background noise. Moreover, we introduce an error regularization term into the progressive asymmetric quantization loss to reduce the quantization errors and preserve the semantic similarity. Finally, we attempt to introduce the code balance and uncorrelated constraints into the multi-codebook quantization framework to improve the utilization efficiency of codewords and reduce the redundant information, which can be approximated by solving the optimal transport problem. Experimental results on several fine-grained image datasets demonstrate that the proposed method outperforms the state-of-the-art fine-grained image retrieval methods. Lei Ma 0004, Yu Shi 0004, Fanman Meng, Qingbo Wu 0001, Hanyu Hong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | RCST: Residual Context-Sharing Transformer Cascade to Approximate Taylor Expansion for Remote Sensing Image DenoisingabstractTaylor expansion is a polynomial for approximating a function constructed by the coefficients of its derivatives at a certain point, where it is a challenging research to utilize the powerful learning ability of deep learning (DL) to characterize the polynomial parts for pursuing its approximate solution. In this article, we develop a cascading residual context-sharing Transformer (RCST) to approximate Taylor expansion for remote sensing (RS) image denoising. Our RCST method includes the following key procedures. First, a mapping function about a latent clean RS image patch is built by employing the low-rank characteristic of its neighborhood RS image blocks, and is expanded into a polynomial with Taylor expansion for its approximate solution. Second, the intrinsic recursive relationship of the neighborhood derivatives is analyzed and is mathematically formulated, which provides a theoretical interpretability for the construction of our RCST model. Third, a lightweight residual network (LRNet) is developed to estimate the base layer, while the RCST is shared to calculate the derivative parts. Finally, to transfer as many rich multiscale details from noisy RS images to estimated results as possible, we adopt a down-/upsampling architecture. Specifically, a spatial-Fourier upsampling (SFUS) operator is reported to preserve both local and global information. Quantitatively and qualitatively experimental results validate that our RCST denoising method can achieve competitive performance and is even superior to other SOTA denoising approaches. Zhenghua Huang, Yu Shi 0004, Yaozong Zhang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Corrections to "Semi-Supervised Learning for Infrared Thermal Radiation Correction in the Real World"
Yu Shi 0004, Xinyuan Deng, Lei Wang 0068, Yaozong Zhang, Zhenghua Huang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Semi-Supervised Learning for Infrared Thermal Radiation Correction in the Real WorldabstractInfrared images are susceptible to thermal radiation. Infrared thermal radiation correction methods based on physical prior may fail while correcting real-world images, because assumed priors do not always hold in the real world, resulting in the presence of thermal radiation residuals. Supervised learning-based methods have the potential to achieve favorable outcomes in the correction of synthetic images. However, due to the unavailability of labeled datasets, their efficacy is limited when applied to real-world images. To address this problem, in this article, to the best of our knowledge, we propose the first semi-supervised learning network for infrared radiation correction in the real world, named SIRCNet. The network is trained using a semi-supervised strategy, which includes a supervised training stage and a self-supervised training stage. In the supervised training stage, we constructed a multilevel wavelet decomposition and reconstruction correction (MWDRC) module for latent image correction and an efficient generalized feature extraction (EGFE) module for bias field estimation. Furthermore, EGFE is composed of one partial channel interactive (PCI) attention block and three effective residual blocks (ERBs). Surface fitting can approximate the thermal radiation bias field of the thermal radiation degradation images. The fit bias field can provide critical prior knowledge that enhances EGFE’s estimation of the thermal radiation bias field. Hence, in the self-supervised training stage, when fine-tuning MWDRC and EGFE using a generator, surface fitting is employed to constrain EGFE. Comparative experiments demonstrate that SIRCNet outperforms existing correction methods on both real and synthetic datasets, achieving the best metrics as well as visualization. Yu Shi 0004, Xinyuan Deng, Lei Wang 0068, Yaozong Zhang, Zhenghua Huang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | WDTSNet: Wavelet Decomposition Two-Stage Network for Infrared Thermal Radiation Effect CorrectionabstractRecently, infrared thermal radiation effect correction methods are dominated by removing bias field in spatial domain. Since they do not consider the low-frequency characteristics of thermal radiation bias field and the high-frequency information of image content, these methods often fail in the enhancement of contrast and details. To address this problem, we propose a novel wavelet decomposition two-stage network for infrared thermal radiation effect correction, named WDTSNet. Through wavelet decomposition, we construct a low-frequency thermal radiation effect coarse correction subnetwork (LFCCSN) and a high-frequency detail enhancement fine correction subnetwork (HFFCSN), respectively. Firstly, we take the small size low-frequency component of the degraded image after discrete wavelet transformation (DWT) as the input of the first stage LFCCSN and propose an intra-block multiscale residual dense module (IMRDM) to complete the coarse correction and contrast enhancement through different scales of receptive fields and intra-block channel information interaction. Secondly, we perform inverse discrete wavelet transformation (IDWT) to obtain the input of the second stage HFFCSN, and build a high-frequency gated residual module (HGRM) in HFFCSN to remove residual thermal radiation bias field and acquire the enhanced high-frequency information. In addition, we further design dual-branch cross-scale attention fusion module (DCAFM) between encoders and decoders to effectively aggregate the cross-scale information flow. Extensive experiments on simulated and real infrared images demonstrate that the proposed WDTSNet performs well on enhancing contrast and details than existing methods. The code will be publicly available upon acceptance. Yu Shi 0004, Yixin Zhou, Lei Ma 0004, Lei Wang 0068, Hanyu Hong |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Joint ordinal regression and multiclass classification for diabetic retinopathy grading with transformers and CNNs fusion network
Lei Ma 0004, Qihang Xu, Hanyu Hong, Yu Shi 0004, Ying Zhu 0002, Lei Wang 0068 |
Appl. Intell. | 4 |
| 2023 | DDABNet: a dense Do-conv residual network with multisupervision and mixed attention for image deblurring
Yu Shi 0004, Zhigao Huang, Jisong Chen, Lei Ma 0004, Lei Wang 0068, Hanyu Hong |
Appl. Intell. | 1 |
| 2023 | DGDNet: Deep Gradient Descent Network for Remotely Sensed Image DenoisingabstractGradient descent strategy, viewed as an important model optimization method, has been widely used for various tasks (such as model-based image denoising) of computer vision. In the gradient descent denoising model, the learning rate (LR) and residual component are two important parts to be adaptively estimated for its stable point. This letter proposes a deep gradient descent network (DGDNet), including two key points: one is that the LR is designed with eigenvalues of Hessian matrix of remotely sensed images (RSIs) and their local weighted factor (LWF), which can recognize structures from RSIs degraded by additive white Gaussian noise (AWGN). The other is that the residual part is calculated by an U-shaped network (USNet) to speed up the DGDNet convergent to a fixed point. Finally, the two components are plugged into the gradient descent scheme and contribute to an enjoyable result with a few iterations. Quantitatively and qualitatively experimental results demonstrate that the proposed DGDNet can obtain a stable solution efficiently, and produce competitive denoising performance which is even better than that yielded by the state-of-the-art noise reduction methods. Zhenghua Huang, Zifan Zhu, Zhicheng Wang 0004, Yu Shi 0004, Yaozong Zhang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | Dynamic scene deblurring with continuous cross-layer attention transmission
Junxiong Fei, Jianguo Liu 0004, Yu Shi 0004, Hanyu Hong |
Pattern Recognit. | 5 |
| 2023 | Question-aware dynamic scene graph of local semantic representation learning for visual question answering
Jinmeng Wu, Fulin Ge, Hanyu Hong, Yu Shi 0004, Yanbin Hao, Lei Ma 0004 |
Pattern Recognit. Lett. | 4 |
| 2022 | DLRP: Learning Deep Low-Rank Prior for Remotely Sensed Image DenoisingabstractRemotely sensed images degraded by additive white Gaussian noise (AWGN) are not beneficial for the analysis of their contents. Such a phenomenon is usually modeled as an inverse problem which can be solved by model-based optimization methods or discriminative learning approaches. The former pursue their pleasing performance at the cost of a highly computational burden while the latter are impressive for their fast testing speed but are limited by their application range. To join their merits, this letter proposes a nonlocal self-similar (NSS) block-based deep image denoising scheme, namely deep low-rank prior (DLRP), which includes the following key points: First, the low-rank property of the neighboring NSS patches ordered lexicographically is utilized to model a global objective function (GOF). Second, with the aid of an alternative iteration strategy, the GOF can be easily decomposed into two independent subproblems. One is a quadratic optimization problem, and has a closed-form solution. While the other is a low-rank minimization denoising problem and is learned by deep convolutional neural network (DCNN). Then, the deep denoiser, acted as a modular part, is plugged into the model-based optimization method with adaptive noise level estimation to solve the inverse problem. In the experiments, we first discuss parameter setting and the convergence. Then, quantitative/qualitative comparisons of experimental results validate that the DLRP is a flexible and powerful denoising method to achieve competitive performance which even outperforms those produced by state-of-the-arts. Zhenghua Huang, Zhicheng Wang 0004, Zifan Zhu, Yaozong Zhang, Yu Shi 0004, Tianxu Zhang |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Removing Atmospheric Turbulence Effects Via Geometric Distortion and Blur RepresentationabstractRemoving the geometric distortion and space-time-varying blur caused by atmospheric turbulence from a given image sequence remains a challenge. Since geometric distortion and blur are two different kinds of distortions and interact with each other in the process of image restoration, it is difficult to extract the features that are useful to the restoration process when the images experience multiple distortions. In this article, we propose a new scheme based on geometric distortion and blur representation. The blur invariants and maximum gradient are used to represent the geometric distortion and sharpness of an image frame, respectively. The proposed scheme consists of three parts. First, two fast frame selection algorithms based on independent evaluations of the sharpness and geometric distortion are proposed to subsample a sharp subsequence and obtain a reference image. Next, to suppress the geometric distortion, a moment-blur-invariant-based method is presented to estimate the deformation vector between two degraded frames, and the selected sharp frames are registered to the reference image. Finally, a blind deconvolution method is applied to deblur the fused image, generating a final restoration result. Various experimental results show that the proposed method can effectively alleviate distortion and blur, as well as significantly improve the visual quality of real atmospheric turbulence-degraded images. Chao Pan 0004, Yu Shi 0004, Jianguo Liu 0004, Hanyu Hong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Multi-scale Image Partitioning and Saliency Detection for Single Image Blind Deblurring
Jiaqian Yan, Yu Shi 0004, Zhigao Huang, Ruzhou Li |
PRCV (4) | 2 |
| 2021 | Learning discrete class-specific prototypes for deep semantic hashing
Lei Ma 0004, Yu Shi 0004, Likun Huang, Zhenghua Huang, Jinmeng Wu |
Neurocomputing | 3 |
| 2020 | Correlation Filtering-Based Hashing for Fine-Grained Image RetrievalabstractThe low storage and strong representation capabilities of hash codes for image retrievalhas made hashing technologies very popular. Several existing deep hashing methods focuson the task of general image retrieval, while neglecting the task of fine-grained image retrieval. Recently, some fine-grained hashing methods have been proposed to capture the subtle differences, which mainly utilize the single-modality visual features to solve the discriminative region localization while ignoring the semantic information. In this letter, we propose a correlation filtering hashing (CFH) method to learn discrete binary codes, which can adequately take advantage of the cross-modal correlation between the semantic information and the visual features for discriminative region localization. Specifically, we utilize a feature pyramid network to learn multi-level visual features. Subsequently, the label vector is embedded into the visual space, which can be used as a correlation filter on the feature maps to capture the latent location of objects. Finally, weperform global average pooling over the output maps and concatenate the features of different levels to produce the hash codes of query images. Extensive experiments on two fine-grained datasets show that the proposed CFH outperforms the state-of-the-art hashing methods. Lei Ma 0004, Yu Shi 0004, Jinmeng Wu |
IEEE Signal Process. Lett. | 3 |
| 2018 | Nonuniformity Correction Method of Thermal Radiation Effects in Infrared Images
Hanyu Hong, Yu Shi 0004, Tianxu Zhang |
PRCV (1) | 2 |
| 2017 | Iteratively reweighted blind deconvolution for passive millimeter-wave images
Houzhang Fang, Yu Shi 0004, Donghui Pan |
Signal Process. | 2 |
| 2012 | An Improved Method of Identification Based on Thermal Palm Vein Image
Ran Wang 0005, Guoyou Wang, Jianguo Liu 0004, Yu Shi 0004 |
ICONIP (2) | 5 |