EDBT 2026 Demo / reviewers in the wild / expert
Haitao Yin
dblp:09/6938
· DBLP profile ↗
23ranked-venue papers
16as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 8 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Frequency-enhanced wavelet transformer based decoder for medical image segmentation
Haitao Yin, Yongchang Xu |
Pattern Recognit. | 1 |
| 2026 | FA-DETR: Frequency-Aware DETR Based Object Detection for Uncrewed Aerial Vehicle ImagesabstractDetection Transformer (DETR) has achieved remarkable performance across diverse object detection tasks. However, the long-range similarity computation inherent to its self-attention mechanism exhibits some limitations in detecting weak and small objects, particularly for structurally similar and densely distributed unmanned aerial vehicle (UAV) objects. To alleviate this limitation, this letter proposes a Frequency-Aware DETR (FA-DETR) that integrates frequency domain learning. FA-DETR incorporates three key modules into RT-DETR, including Adaptive Frequency Adjustment (AdaFA) module, Spatial-Frequency Feature Fusion (S3F) module, and Cross-Level Feature Compensation (CLFC) module. AdaFA splits input features into several frequency sub-bands and refines them through dynamic reconfiguration, which can enhance feature informativeness of weak and small objects. Through learnable parameters, S3F selectively combines spatial-frequency features across different levels and mitigates inherent feature biases. To further preserve semantic-structural consistency, CLFC employs a self-compensation approach that enriches high-level representations with structural details from low-level features using multiscale cross-attention. FA-DETR is built by plugging the AdaFA, S3F and CLFC modules into RT-DETR. Extensive experiments on benchmark datasets demonstrate that FA-DETR obtains state-of-the-art detection performance. Specifically, it provides 2.8% mAP improvements over RT-DETR on the VisDrone-2019 dataset. Haitao Yin |
IEEE Signal Process. Lett. | 2 |
| 2025 | SPViT-FER: A Sparse Pruning Based Vision Transformer for Facial Expression Recognition
Haitao Yin |
ICIG (2) | 2 |
| 2025 | Transformer-Mamba U-Net With Dual Cross-Attention for Medical Image Segmentation
Yongchang Xu, Haitao Yin |
PRCV (13) | 2 |
| 2025 | Progressive Dynamic Queries Reformation-Based DETR for Remote Sensing Object DetectionabstractObject queries-based detection transformer (DETR) makes remarkable achievements in object detection. However, most object queries design approaches are initialized with only one input and shared among all samples, which may result in the propagation of probing errors and lacking understanding of remote sensing objects with diversified structures and complex backgrounds. To address these issues, this letter proposes a progressive dynamic queries reformation (PDQR) for DETR-based remote sensing object detection, which consists of multihierarchical dynamic object queries and progressive reformation. A group of unique object queries are dynamically weighted, which are then fed into the current stage of decoder to reform the updated object queries of previous stage. This progressive reformation can suppress error propagation from earlier stages and reduce the influences of backgrounds. Moreover, the dynamic object queries can enhance the awareness ability of fine-grained features. PDQR can be flexibly plugged into various DETRs. The experimental results on different benchmark datasets demonstrate the superiority of PDQR over several state-of-the-art DETRs. Specifically, the PDQR-based DINO achieves 95.9%, 80.2%, and 97.3% mAPs on NWPU VHR-10, DIOR, and RSOD datasets, respectively. Haitao Yin, Zhuyun Zhu |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2025 | A Text-Guided Query Adaptive Vision Transformer for PansharpeningabstractPansharpening aims to fuse a low resolution multispectral (LRMS) image and a high resolution panchromatic (PAN) image, and then synthesize a high resolution MS (HRMS) image. Benefiting from its potent capacity for global information modeling, vision Transformer (ViT) has achieved remarkable performance in pansharpening. However, most ViT-based methods heavily rely on cross-attention with static modality queries to extract and integrate complementary information from PAN and LRMS images, neglecting the adaptivity and specificity of different modalities. Moreover, vision-only representation framework often struggles to balance spatial enhancement with spectral fidelity. To tackle these issues, this letter proposes a Text guided Query Adaptive ViT (Text-QAViT) for pansharpening. Specifically, a Query Adaptive Cross-Attention (QACA) block is designed to enhance cross-modal interaction, which comprises cross-modal query projection and gating selection mechanism. Compared to static swapped queries, the queries in the QACA block are dynamic and exhibit better adaptability. Moreover, inspired by the powerful semantic understanding of language-vision model, we design a Text-Conditional Batch Normalization (TCBN) block which utilizes the text embeddings as a guidance to modulate image features. By leveraging high-level language semantics, the TCBN block flexibly balances spatial and spectral features, thereby effectively enhancing visual quality. Extensive experiments on both reduced-resolution and full-resolution datasets acquired by different satellites demonstrate the superiority of Text-QAViT over state-of-the-art methods. Haitao Yin, Yamin Zhu |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2025 | SED-DETR: A Scale-Enhanced Deformable Detection Transformer for Remote Sensing ImagesabstractDetection Transformer (DETR) has emerged as a highly promising approach in object detection and has attracted significant interest. However, most DETR-like methods cannot simultaneously leverage the shape and scale priors for attention calculation, resulting in limited performance in detecting remote sensing objects with diverse shapes and scales. To address this issue, this article proposes a Scale-Enhanced Deformable DETR (SED-DETR) for remote sensing object detection (RSOD). The core component of SED-DETR is Scale-Enhanced Deformable Attention (SEDA), which is designed based on the principles of deformable shape and dynamic scale. Specifically, the SEDA module utilizes multi-scale attention heads. First, conventional multiple attention heads are consolidated into several scale-heads through an adaptive scale aggregation approach, which dynamically adjusts the distributions of different scales to enhance the scale-aware modeling ability. For each scale-head, dilated sampling is applied at a specific dilation rate to capture multi-scale receptive fields. The sampled positions are further refined by learnable offsets predicted from query features, enabling a deformable dilated mechanism for fine-grained feature extraction of multi-scale instances. Finally, we adopt the mixed query selection and the denoising training defined in DINO to implement SED-DETR. Experimental results on the xView, DIOR, NWPU VHR-10 and COCO datasets demonstrate that SED-DETR outperforms state-of-the-art DETR-like methods. Specifically, SED-DETR achieves 5.6%, 10.9%, and 8.6% mAP gains over the baseline Deformable DETR on the xView, DIOR, and NWPU VHR-10 datasets, respectively. The source code is available at https://github.com/zzy599/SEDDETR. Haitao Yin, Zhuyun Zhu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Deep side group sparse coding network for image denoisingabstractAbstract Recently, deep learning has made significant progress in image denoising. However, most of existing deep learning based methods are purely data‐driven, without considering the knowledge of image denoising. Moreover, the parameters of deep denoising network are not explainable. According to these issues, this paper proposes a deep side group sparse coding network for image denoising, named a side group sparse coding (SGSC)‐Net. First, SGSC model for image denoising by exploiting prior information regarding the group sparse coefficients consistency is developed. Specifically, the side information is constructed as the weighted combination of intermediate estimations, and updated iteratively. Then, the optimisation solution of SGSC model is turned into a deep neural network using deep unfolding, that is, SGSC‐Net. The computational path of SGSC‐Net fully follows the iterations of optimisation solution, and consequently the network parameters are interpretable. Furthermore, the design of SGSC‐Net employs the insight of SGSC denoising model. The experimental results on well‐known datasets quantitatively and qualitatively demonstrate that SGSC‐Net is competitive to existing deep unfolding‐based and typical deep neural network‐based methods. Haitao Yin, Tianyou Wang |
IET Image Process. | 1 |
| 2022 | RAiA-Net: A Multi-Stage Network With Refined Attention in Attention Module for Single Image DerainingabstractImage deraining is an important task for the outdoor computer vision system in the rain days. Deep learning with attention mechanism has shown promising performance for rain removal. However, most of existing attention mechanisms select the locations of attention module empirically, and the receptive field of network is small and fixed. Motivated by the Attention in Attention (AiA) and multi-stage strategy, we propose a multi-stage deep neural network for image deraining, which is equipped with the refined AiA (RAiA) module. Specifically, RAiA extends the original AiA by exploiting the joint dependencies of channel and spatial information to generate the dynamic allocated weights. Moreover, the proposed network is carried out by a multi-stage U-Net fashion, which can extract features progressively, enlarge the receptive field of network, and improve the ability of image structures representation. We demonstrate the superiority of proposed method on the public synthetic datasets and real rainy images through comparing to seven state-of-the-art deep learning based methods. Haitao Yin |
IEEE Signal Process. Lett. | 1 |
| 2022 | CSformer: Cross-Scale Features Fusion Based Transformer for Image DenoisingabstractWindow self-attention based Transformer receives the advanced results in image denoising. However, the current methods still have some limitations in capturing the global dependencies and local responses. To tackle these problems, this paper proposes a novel Transformer based image denoising method, called as CSformer, which is equipped with two key blocks, including the cross-scale features fusion (CS2F) block and mixed global-local Swin (M-Swin) Transformer block. The CSformer has a specific multi-scale framework, in which the multi-scale features, extracted by M-Swin Transformer, are fused using CS2F block. Such cross-scale fusion not only enriches the features, but also yields the multi-scale self-attention. In addition, the M-Swin Transformer block consists of the Swin Transformer block and the separable convolution based convolutional local-extraction (CLE) block, which can boost the ability of Transformer in local representation. We demonstrate the superiority of CSformer on some well-known datasets at different noise levels with comparisons to several state-of-the-art methods. Haitao Yin |
IEEE Signal Process. Lett. | 1 |
| 2022 | Laplacian Pyramid Generative Adversarial Network for Infrared and Visible Image FusionabstractGenerative adversarial network (GAN) has recently demonstrated a powerful tool for infrared and visible image fusion. However, existing methods extract the features incompletely, miss some textures, and lack the stability of training. To cope with these issues, this article proposes a novel image fusion Laplacian pyramid GAN (IF-LapGAN). Firstly, a generator is constructed which consists of shallow features extraction module, Laplacian pyramid module, and reconstruction module. Specifically, the Laplacian pyramid module is a pyramid-style encoder-decoder architecture, which progressively extracts the multi-scale features. Moreover, the attention module is equipped in the decoder to effectively decode the salient features. Then, two discriminators are adopted to discriminate the fused image and two different modalities respectively. To improve the stability of adversarial learning, we propose to develop another side supervised loss based on the side pre-trained fusion network. Extensive experiments show that IF-LapGAN achieves 3.27%, 27.28%, 6.32%, 1.39%, 3.14%, 1.15% and 1.07% improvement gains in terms of$Q_{NMI}$,$Q_{M}$,$Q_{Yang}$,$Q^{AB/F}$, MI, VIF, and FMI, respectively, compared with the second best values. Haitao Yin, Jinghu Xiao |
IEEE Signal Process. Lett. | 1 |
| 2022 | PSCSC-Net: A Deep Coupled Convolutional Sparse Coding Network for PansharpeningabstractGiven a low-resolution multispectral (MS) image and a high-resolution panchromatic image, the task of pansharpening is to generate a high-resolution MS image. Deep learning (DL)-based methods receive extensive attention recently. Different from the existing DL-based methods, this article proposes a novel deep neural network for pansharpening inspired by the learned iterative soft thresholding algorithm. First, a coupled convolutional sparse coding-based pansharpening (PSCSC) model and related traditional optimization algorithm are proposed. Then, following the procedures of traditional algorithm for solving PSCSC, an interpretable end-to-end deep pansharpening network is developed using a deep unfolding strategy. The designed deep architecture can also be understood in the view of details injection (DI)-based scheme. This work offers a solution that integrates the DL-, DI-, and variational optimization-based schemes into a framework. The experimental results on the reduced- and full-scale datasets demonstrate that the proposed deep pansharpening network outperforms popular traditional methods and some current DL-based methods. Haitao Yin |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Panchromatic Side Sparsity Model-Based Deep Unfolding Network for PansharpeningabstractDeep learning (DL) recently receives state-of-the-art results in pansharpening. However, most of the existing pansharpening neural networks are purely data-driven, without taking account of the characteristics of the pansharpening task. To address this issue, we propose a novel deep unfolding network for pansharpening that combines the insight of the variational optimization (VO) model and the capability of deep neural network (DNN). We first develop a panchromatic side sparsity (PASS) prior-based VO model for pansharpend image reconstruction, which is formulated as the$\ell _{1}-\ell _{1}$minimization. In particular, the PASS prior is defined using the transform sparsity, which can alleviate the influences of the irregular outliers between the multispectral (MS) and panchromatic (PAN) images. The iterations of half-quadratic splitting algorithm for solving the$\ell _{1}-\ell _{1}$minimization are then deeply unfolded into a DNN, referred as PASS-Net. To capture the nonlinear relationship between MS and PAN images, the linear transforms used in PASS prior are extended into the subnetworks in PASS-Net. Moreover, a pair of learnable downsampling and upsampling modules are designed to realize the downsampling and upsampling operations, which can improve the flexibility. The experimental results on different satellite datasets confirm that PASS-Net is superior to some representational traditional methods and state-of-the-art DL-based methods. Haitao Yin |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Hyperspectral Image Denoising Based on Graph-Structured Low Rank and Non-local Constraint
Haitao Yin |
PRCV (1) | 2 |
| 2019 | PAN-Guided Cross-Resolution Projection for Local Adaptive Sparse Representation- Based PansharpeningabstractSparse representation (SR)-based methods solve pansharpening as an image superresolution problem and receive great popularity. Conventional approaches assume that the high- and low-resolution images have the same sparse coefficients. However, the identity mapping is not universal and also limits the performance. To overcome this limitation, this paper proposes a PAN-guided cross-resolution projection-based pan-sharpening (PGCP-PS) which incorporates the SR image superresolution and details injection pansharpening scheme into a framework. The basic idea of PGCP-PS is to inject a possible offset into the SR superresolution reconstructed part. In addition, the same sparse coefficients assumption across different resolutions is relaxed as the same sparse support with a local adaptive cross-resolution projection. By exploiting the similarity between panchromatic (PAN) and multispectral (MS) images, the cross-resolution projection and offset for sharpening the MS image are estimated from a simulated PAN image superresolution scenario. The high- and low-resolution dictionaries used in the stage of SR image superresolution are learned from PAN image and its degraded version. A series of experimental results on the reduced-scale and full-scale data sets demonstrates that the PGCP-PS outperforms some advanced methods and existing SR-based methods. Haitao Yin |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | Learning category distance metric for data clustering
Baoguo Chen, Haitao Yin |
Neurocomputing | 2 |
| 2017 | A Joint Sparse and Low-Rank Decomposition for Pansharpening of Multispectral ImagesabstractPansharpening aims to fuse a high-resolution panchromatic (PAN) image and a low-resolution multispectral (MS) image. Several synthesis techniques have been reported to solve the problem of pansharpening. Details injection (DI) consists of the cascaded processes of details extraction and injection. The former is crucial for performance. By exploiting the relationship among multiple data acquired on the same scene through different sensors, this paper first develops a joint sparse and low-rank (JSLR) decomposition with an assumption that multiple data have a common low-rank component. Then, a novel DI-type pansharpening method is proposed based on JSLR decomposition, named as JSLR-based pansharpening (JSLRP). In JSLRP, the injected spatial details are calculated as a linear combination of JSLR decomposed components. To ensure the low-rank condition, the JSLR is implemented on the PAN and MS images in the nonlocal similar patches form by adopting the nonlocal self-similarity. Finally, the superiority of JSLRP is demonstrated by comparing with several well-known methods on the reduced-scale data and full-scale data. Haitao Yin |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | Sparse representation with learned multiscale dictionary for image fusion
Haitao Yin |
Neurocomputing | 1 |
| 2015 | Sparse representation based pansharpening with details injection model
Haitao Yin |
Signal Process. | 1 |
| 2015 | Pansharpening With Multiscale Normalized Nonlocal Means Filter: A Two-Step ApproachabstractPansharpening aims to synthesize a high-spatial-resolution multispectral (MS) image by fusing a panchromatic (PAN) image and a low-resolution MS image. The multiresolution analysis (MRA)-based methods are a popular group of pansharpening methods. However, in the MRA-based methods, spatial distortions may occur in the pansharpened product due to the misalignment of PAN and MS data. To address the spatial distortion issue in MRA-based methods, this paper proposes a two-step approach, which consists of the coarse step and the refined step. The coarse step produces a preliminary result using the traditional details injection model. Then, the preliminary product is refined with a second details injection operation in the refined step. Moreover, in our proposed two-step approach, a novel multiscale decomposition based on a normalized nonlocal means (NNLM) filter is developed to extract the spatial detail. Compared with the original nonlocal means filter, the designed NNLM makes the similarity measure more robust and accurate by exploiting the normalized intensity value and the mean value jointly. The experimental results on various satellite data demonstrate the superiority of the proposed pansharpening scheme by comparing with ten well-known methods. Haitao Yin, Shutao Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2013 | Remote Sensing Image Fusion via Sparse Representations Over Learned DictionariesabstractRemote sensing image fusion can integrate the spatial detail of panchromatic (PAN) image and the spectral information of a low-resolution multispectral (MS) image to produce a fused MS image with high spatial resolution. In this paper, a remote sensing image fusion method is proposed with sparse representations over learned dictionaries. The dictionaries for PAN image and low-resolution MS image are learned from the source images adaptively. Furthermore, a novel strategy is designed to construct the dictionary for unknown high-resolution MS images without training set, which can make our proposed method more practical. The sparse coefficients of the PAN image and low-resolution MS image are sought by the orthogonal matching pursuit algorithm. Then, the fused high-resolution MS image is calculated by combining the obtained sparse coefficients and the dictionary for the high-resolution MS image. By comparing with six well-known methods in terms of several universal quality evaluation indexes with or without references, the simulated and real experimental results on QuickBird and IKONOS images demonstrate the superiority of our method. Shutao Li 0001, Haitao Yin, Leyuan Fang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2012 | Multitemporal Image Change Detection Using a Detail-Enhancing Approach With Nonsubsampled Contourlet TransformabstractIn this letter, we propose an unsupervised approach for change detection in multitemporal satellite images based on a novel detail-enhancing algorithm. The multitemporal source images are first used to generate the difference image, which is decomposed into low-pass approximation and high-pass directional subbands by the nonsubsampled contourlet transform. The coefficients from the directional subbands are fused at intrascale and interscale to extract the meaningful details of the difference image. After that, the extracted details are injected into one base image selected from the approximation subbands, which results in a detail-enhanced difference image. For each pixel in the enhanced difference image, a dimension-reduced feature vector is created using the principal component analysis (PCA). The final change detection map is achieved by clustering the feature vectors using a PCA-guidedk-means algorithm into “changed” and “unchanged” classes. Experimental results demonstrate the superior performance of the proposed approach compared with several well-known change detection techniques. Shutao Li 0001, Leyuan Fang, Haitao Yin |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2011 | Single image super resolution via texture constrained sparse representationabstractImage super resolution is a challenging highly ill-posed inverse problem. In this paper, we proposed a texture constrained sparse representation for single image super resolution. Firstly, the low resolution observed image is segmented into different texture regions. Through preprepared texture databases, the low resolution regions are classified into different texture categories using the designed texture classifier. Then, the high resolution segments are reconstructed by sparse representation with relevant texture dictionaries. Integrating all segments, the high resolution result is obtained. The proposed method is compared with sparse representation method and some existing methods. The experimental results show that our method achieves better results in visual inspection and quantitative analysis. Haitao Yin, Shutao Li 0001, Jianwen Hu |
ICIP | 1 |