EDBT 2026 Demo / reviewers in the wild / expert
Zonghao Han
dblp:329/9768
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2025
0009-0009-9613-6458ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 7 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MSFD: Multiscale Feature Decomposition for Cross-Modality Visible-to-Infrared Drone Image TranslationabstractIn the global landscape of the Internet of Things (IoT), drone IoT technology has gained widespread application. This technology can monitor and analyze land use and land cover more quickly and more accurately. Currently, the images collected by drone IoT technology are mostly visible images, which are highly susceptible to external environmental factors, while the acquisition of infrared images is relatively more challenging. Visible-to-infrared drone image translation seeks to convert visible drone images into their corresponding infrared counterparts. Although existing GAN-based image-to-image translation methods have demonstrated impressive results in the domain of natural images, they still face challenges in generating highly realistic infrared drone images. Therefore, a novel Multi-Scale Feature Decomposition (MSFD) method is introduced for visible-to-infrared drone image translation. The proposed approach accomplishes the translation through spectral feature disentanglement and cross-modal recombination. In our model, spectral feature disentanglement is based on the separation of modality-specific spectral information and modality-invariant shared structural content from the image representation. Subsequently, the spectral features and underlying content from different modalities can be recombined by generators to facilitate cross-modality image translation. To enhance the quality of generated images, our method integrates a multi-scale spectral feature encoder to address significant spectral discrepancies between targets and backgrounds in drone images by extracting and fusing spectral features at different scales. Additionally, the strategy of multi-scale generators and discriminators further enhances the generation quality of infrared drone images. The experimental results highlight the superior performance of our model in visible-to-infrared drone image translation. Zhiquan Liu 0001, Zonghao Han, Mingyang Ma 0004, Jian Zhao 0002 |
IEEE Internet Things J. | 3 |
| 2025 | CAMCFormer: Cross-Attention and Multicorrelation Aided Transformer for Few-Shot Object Detection in Optical Remote Sensing ImagesabstractFew-shot object detection (FSOD) enables the detection of novel-class objects in remote sensing images (RSIs) with limited labeled samples. Although convolutional neural networks (CNNs) are commonly used for this task, they suffer from two inherent constraints. First, their limited local receptive field fails to capture global context within a single image and the relational dependencies between query and support images. Second, an additional feature alignment mechanism is typically required to bridge the gap between query and support images. To address these challenges, this work introduces a novel cross-attention and multicorrelation aided transformer (CAMCFormer) FSOD framework tailored for global feature representation and multicorrelation modeling in complex and large-scale RSIs. Specifically, a long-distance cross-attention module (LDCAM) is devised to capture dependencies between distant elements across query and support images at each feature extraction layer. This module facilitates the exchange of contextual information between images, resulting in more comprehensive feature representations and eliminating the need for separate feature alignment and fusion modules. Multicorrelation aided heads (MAHs) are constructed to enhance detection performance further to model various relational aspects, i.e., channel-correlation detection head (CCDH), spatial-correlation detection head (SCDH), and cross-attention detection head (CADH). These aided heads contribute to more robust and accurate classification and localization. Comprehensive experiments have been conducted, demonstrating the superiority of the proposed framework compared to several state-of-the-art detectors, highlighting its potential as an effective solution for FSOD in remote sensing scenarios. Lefan Wang, Shaohui Mei, Yi Wang 0068, Jiawei Lian, Zonghao Han, Yan Feng 0005 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | DR-AVIT: Toward Diverse and Realistic Aerial Visible-to-Infrared Image TranslationabstractImage-to-image (I2I) translation methods based on Generative Adversarial Networks (GANs) have shown general solutions for aerial visible-to-infrared image translation (AVIT) task. Though existing approaches have made impressive results, they still struggle to produce diverse or high-realism translated aerial infrared images (AIIs). In this paper, a novel model is proposed to achieve both diverse and realistic AVIT, named DR-AVIT. Specifically, we introduce disentangled representation learning to disentangle the image representation of aerial visible images (AVIs) and AIIs into a domain-invariant semantic structure space and two domain-specific imaging style spaces. By leveraging this disentanglement, our model can perform the translation process conditioned on semantic structure information derived from the input AVI and randomly sampled imaging style features from the AII domain to obtain diverse outputs. Furthermore, a new constraint is present to encourage GANs to learn efficient mappings between AVI and AII domains by integrating geometry-consistency constraint and a dual learning framework, named dual geometry-consistency constraint. Coping with these two designs, our method exhibits superiority in both realism and diversity of the translation results over several state-of-the-art I2I translation methods on AVIID dataset and two new benchmark datasets for AVIT, which are obtained by extracting data from publicly available datasets. Code of DR-AVIT and proposed benchmark datasets are available at https://github.com/silver-hzh/DR-AVIT. Zonghao Han, Yuru Su, Shaohui Mei |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | A Novel Center-Boundary Metric Loss to Learn Discriminative Features for Hyperspectral Image ClassificationabstractLearning discriminative features is of crucial for hyperspectral image (HSI) classification. Though metric learning has been applied to learn effective features in HSI classification tasks, existing metric loss functions only consider distance among features of sample pairs but ignore the feature centers and boundaries in the embedding feature space, which limits the discrimination of learned features. In this paper, a novel metric loss function named center-boundary metric loss (CBML) is proposed to learn more discriminative features so as to improve HSI classification performance. Unlike the existing metric loss functions, CBML not only considers the distance between sample pairs to enhance intra-class similarity and inter-class separability but also pays more attention to the feature centers and boundaries in the embedding feature space that could greatly determine and affect the category of features. Specifically, CBML forces the distance of a sample to its corresponding feature center to be explicitly smaller than that to samples from other classes by a predefined threshold. As a result, the boundaries of different classes will separate an actual distance, which improves the discrimination of learned features. Moreover, in order to improve the training efficiency, a cross mini-batch sampling strategy is further proposed to break through the limitation within the mini-batch by using features between several contiguous mini-batches to sample pairs without increasing the size of the mini-batch. Accordingly, the sampling range of sample pairs is greatly expanded, and the training data is more fully exploited. Experimental results over four benchmark datasets with a typical network for HSI classification demonstrate our proposed method outperforms several state-of-the-arts. Shaohui Mei, Zonghao Han, Mingyang Ma 0004, Fulin Xu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Few-Shot Object Detection With Multilevel Information Interaction for Optical Remote Sensing ImagesabstractMetalearning has been widely applied to solve the few-shot object detection (FSOD) problem in natural scenes, which performs similarity measurement and information aggregation of the support set and the query set. However, regarding remote sensing images (RSIs), many difficulties caused by their disparities need to be further addressed, such as inconsistencies in imaging scale, direction, and background between support and query images. These result in feature misalignment and attention bias, interfering with model performance. In this article, a multilevel information interaction (MLII) strategy is proposed for FSOD to alleviate feature misalignment and attention bias. Information interactions are conducted within multiple scales of features and highlight similar regions of query and support features. A semantic enhancement module (SEM) is proposed to assist MLII in extracting key information and achieving more discriminative feature representation. Moreover, a feature cross-aggregation module (FCM) with separate classification losses is designed to train the detector to identify objects that coexist in query and support images. Extensive experiments demonstrate that the proposed method outperforms several state-of-the-art few-shot object detectors over commonly used benchmark datasets, i.e., DIOR and NWPU-10. Lefan Wang, Shaohui Mei, Yi Wang 0068, Jiawei Lian, Zonghao Han |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Adaptive Composite Feature Generation for Object Detection in Remote Sensing ImagesabstractObject detection in remote sensing images identifies and extracts the acquired Earth surface information, providing data support and research basis for multiple fields. Remote sensing image object detection based on knowledge distillation (KD) can transfer the knowledge of a large teacher model to a smaller student model, achieving the effect of low parameter volume and high accuracy. Mainstream methods directly imitate teacher features to improve student performance, ignoring the generation of high-ranking features through teacher features instructing student feature maps in this knowledge transfer process. In this article, an adaptive composite feature generation (ACFG) strategy is proposed to achieve end-to-end trainable KD for object detection in remote sensing images, in which the robustness of feature points under composite masks is improved through adaptive feature mapping. In particular, a composite mask generator (CMG) module is proposed to select student instance-related features and point background features. Furthermore, a global and local projection layer (GLPL) module is proposed to connect the local information and global information of the feature map under the mask generator to adaptively realize the global recovery mapping of the feature map with partial feature points. Finally, balanced decoupling loss (BDL) is improved to handle foreground and background loss separately, so that the two decoupled features can better enable the student model to learn instance-related information. Note that the proposed ACFG is capable of conducting KD for both single-stage and two-stage object detectors. Experimental results using both anchor-based and anchor-free detectors on the DIOR dataset and DOTA dataset demonstrate that the proposed ACFG clearly achieved better performance than several state-of-the-art (SOTA) algorithms for KD. Ziye Zhang 0007, Shaohui Mei, Mingyang Ma 0004, Zonghao Han |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Hierarchical Feature Fusion of Transformer With Patch Dilating for Remote Sensing Scene ClassificationabstractRecently, the Transformer-based technique has emerged as a promising solution for modeling contextual information in Remote Sensing (RS) scenes and has found widespread applications in RS scene classification. However, how to make full use of intermediate features learned in Transformers is of crucial importance in the RS scene classification tasks. Therefore, this paper proposes a Hierarchical Feature Fusion of Transformer with Patch Dilating (HFFT-PD), which aims to capture rich contextual information from hierarchical features to enhance the performance of RS scene classification. Specifically, the HFFT-PD model consists of a Hierarchical Transformer Merging (HTM) block and a Lightweight Adaptive Channel Compression (LACC) module, in which the HTM is specially designed for the Transformer architecture to bridge the semantic gaps between features from different hierarchical blocks, and the LACC accounts for the significance of distinct channels in the ultimate classification features. In addition, a brand-new Patch Dilating strategy is uniquely designed for the Transformer paradigm, functioning as a reassembly operator predicated on patch features. Contrasting with conventional upsampling techniques, Patch Dilating facilitates upsampling without requiring supplementary information, while concurrently preserving the semantic content of local spatial structure. Extensive and rigorous experiments conducted on the UCM, AID, and NWPU-45 datasets, with training ratios of 80%, 50%, and 20% respectively, demonstrate that our proposed HFFT-PD outperforms the baseline at least by 0.59%, 0.44%, and 0.99% respectively, showcasing the significant superiority of our HFFT-PD over contemporary state-of-the-art methodologies. Mingyang Ma 0004, Yong Li 0036, Shaohui Mei, Zonghao Han, Jian Zhao 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Diversity Measurement-Based Meta-Learning for Few-Shot Object Detection of Remote Sensing ImagesabstractMost object detection methods based on deep learning require large amounts of labeled data and can detect only the categories in the training set. Such issues significantly limit applications in remote sensing scenarios where it usually needs to recognize novel, unseen objects given very few training examples. To address these limitations, a novel meta-learning-based object detection method using Faster R-CNN framework is proposed for optical remote sensing image. Specifically, a diversity measurement module is proposed to measure diversity information between support images and query images on base classes so as to acquire more meta-knowledge. Experiments on DIOR dataset demonstrate our method has achieved superior performance than state-of-the-art meta-learning detection models in the field of remote sensing. Lefan Wang, Zonghao Han, Yan Feng 0005, Jiang Wei, Shaohui Mei |
IGARSS | 3 |