VLDB 2026 Research / reviewers in the wild / expert
Kun Gao 0001
dblp:46/2802-1
· DBLP profile ↗
24ranked-venue papers
0as first author
18since 2021 · last 2026
0000-0001-6666-8036ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Topology-aware dynamic high-order graph learning for hyperspectral image classificationabstract• Topology-aware dynamic high-order graph learning proves robustness & high performance. • Dynamic high-order graph module enables boundary & spatial topology perception. • Spectral-spatial image convolution module reduces redundant band & mine local feature. • Spectral-spatial feature fusion module captures rich spectral-spatial features. Hyperspectral image (HSI) classification is challenging due to intra-class spectral variability caused by mixed pixels at land cover boundaries in complex spatial distributions. Additionally, the classification model easily suffers from overfitting due to high-dimensional features and limited available samples. Recently, GCNs have emerged to handle these challenges, as they enable convolution on non-Euclidean data with arbitrary structures and modeling topological relationships between nodes. However, existing GCNs either use predefined adjacency matrices that cannot be updated throughout the training process or perform feature aggregation only between the central node and its first-order neighbors, which hinders the building of spatial topology. This paper introduces a novel topology-aware dynamic high-order graph learning (TDHGL) to address the aforementioned issues. Specifically, we first introduce a spectral-spatial image convolution module (SICM) to reduce spectral redundancy and extract local spectral-spatial features. Then, we design a dynamic high-order graph module (DHGM) to extract the spatial topology of the HSI. It allows for dynamic updating of adjacency matrices and facilitates high-order message interactions between the central node and its higher-order neighbors, mitigating classification inaccuracies induced by intra-class spectral variability while enhancing the TDHGL’s generalization capacity. Finally, the spectral-spatial feature fusion module (SFFM) aggregates the local spectral-spatial features with spatial topology to capture rich spectral-spatial features. Our TDHGL achieves overall accuracies of 94.37 %, 97.17 %, 95.87 %, and 98.21 % on the Indian Pines, Salinas, University of Pavia, and WHU-Hi-LongKou datasets, respectively, and outperforms other state-of-the-art methods. Hong Wang 0025, Kun Gao 0001, Xiaodian Zhang, Zhijia Yang, Wei Li 0032 |
Expert Syst. Appl. | 2 |
| 2025 | Efficient Grounding DINO: Efficient Cross-Modality Fusion and Efficient Label Assignment for Visual Grounding in Remote SensingabstractVisual grounding for remote sensing (RSVG) aims to detect objects in remote sensing scenes based on textual descriptions. While existing methods perform well on RSVG datasets, they are limited to single-object predictions, making them unsuitable for multi-object candidate category datasets. Open-set methods can be applied to both RSVG and candidate datasets, but their use in remote sensing remains rare. To bridge this gap, we introduce the open-set approach to RSVG and propose Efficient Grounding DINO, using Grounding DINO as a baseline. Open-set methods rely on two key modules: cross-modality fusion and label assignment. Existing cross-modality fusion methods simultaneously update text and multi-scale visual features, which hampers the model’s ability to generalize under different texts and increases learning complexity. Existing methods predict a single object, allowing direct use as a positive example for loss calculation, while open-set methods for multi-objects require one-to-one matching to assign positive and negative samples. However, background interference in the RSVG datasets causes frequent misassignments, slowing model convergence. We address these issues with two innovations: the multi-scale image-to-text fusion module (MSITFM), which updates text features using self-attention to maintain independence from visual features and employs scale-specific cross-attention for multi-scale visual feature fusion to reduce learning complexity, achieving a 3% parameter and 21.6% GFLOPs reduction. Text confidence matching (TCM) incorporates IoU-based confidence into label assignment to reduce mismatches and enhance model performance. Experiments on DIOR-RSVG, RSVG-HR, and DOTA datasets validate the effectiveness of our approach. Zibo Hu, Kun Gao 0001, Xiaodian Zhang, Zhijia Yang, Mingfeng Cai, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | DSFuse: A Dual-Diffusion Structure for Feature Fidelity Infrared and Visible Image FusionabstractImage fusion aims to combine the complementary features of different modalities to produce an informative fused image. Due to the different imaging mechanisms, information conflicts may arise from infrared and visible source images. Existing infrared and visible fusion methods are devoted to preserving the features of source images as much as possible. However, handling conflicting information is often overlooked. Thus, we leverage the powerful generative priors of diffusion and propose a dual-diffusion structure, termed DSFuse, to handle conflicting information and achieve feature fidelity during image fusion processing. Diffusion modules are introduced to guide the fusion network to understand the meaningful information of the source image easily. First, the fusion network is used to retain features in the fused image as much as possible. Then, diffusion modules are used to reconstruct source images from noise based on the output of the fusion network. Finally, feedback from the diffusion modules forces the fusion network to aggregate modality information to ensure fidelity; the high quality of the fusion result is also profitable for a better reconstruction of diffusion modules, forming a positive feedback loop. In addition, we release a new dataset for infrared/visible fusion to support the fusion network training and evaluation, named the multiscene infrared and visible (MSIV) images dataset. Extensive experiments demonstrate that DSFuse outperforms other state-of-the-art (SOTA) fusion methods. Zhijia Yang, Kun Gao 0001, Yanzheng Zhang, Xiaodian Zhang, Zibo Hu, Wei Li 0032 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Unsupervised Blind Hyperspectral Super-Resolution for Unregistered ImagesabstractHyperspectral images super-resolution (HSI-SR) aims to fuse low-resolution HSI (LR-HSIs) and high-resolution multispectral images (HR-MSIs) for high-resolution HSIs (HR-HSIs). Most existing methods require registered image pairs and prior knowledge of spectral response functions (SRFs), which requires effort to realize in practical applications. To overcome this limitation, this paper proposes an unsupervised blind HSI-SR method (UBHSI-SR) for unregistered HSIs and MSIs. UBHSI-SR consists of two unmixing branches, each having its own encoder while sharing the decoder. First, the HSI unmixing branch learns to predict abundance maps and learns precise endmember spectra. Then, the learnable SRF transfers LR-HSIs to the registered LR-MSIs. The abundance similarity constraint between LR-HSIs and LR-MSIs guides the learning of the MSI encoder. With the abundance maps of HR-MSI, the shared decoder predicts the HR-HSIs as final results. Experiments on three remote sensing datasets validate the superior performance of UBHSI-SR to existing fusion methods. Baiyang Hu, Xiaodian Zhang, Kun Gao 0001 |
IGARSS | 3 |
| 2024 | Computer Vision Target Detection-Aided High-Frequency Satellite-Ground CommunicationsabstractSatellite-to-ground communication systems typically operate in environments with high interference levels, complex topologies, and stringent platform constraints. Therefore, intelligent, anti-interference, and low-power systems are required to achieve the desired transmission performance. This paper proposes a system for optimizing high-frequency satellite-to-ground communications using computer vision (CV) technology, like millimeter-wave (mmWave) satellite communication systems. The system uniquely combines CV-based target localization with adaptive beamforming and power control to optimize communication links with ground targets such as base stations, ships, and aircraft. This approach significantly outperforms traditional radio frequency-based methods in accuracy and efficiency, particularly in dynamic mmWave scenarios. Simulation results confirm the superiority of our system in terms of sum rate and energy efficiency, demonstrating its potential to revolutionize high-frequency satellite communications by providing reliable, high-quality service to terrestrial targets. Finally, simulation results are presented to demonstrate the efficiency of the proposed schemes. Zizheng Hua, Ying Ke, Shuai Wang 0013, Gaofeng Pan, Kun Gao 0001 |
IEEE Internet Things J. | 6 |
| 2024 | Optical Imaging Degradation Simulation and Transformer-Based Image Restoration for Remote SensingabstractDue to atmospheric turbulence, optical system limitations, satellite platform jitter, and other reasons, remote sensing images inevitably undergo different degrees of degradation. Employing the deep learning method to improve the on-orbit image quality faces many challenges such as lack of data, limited computing resources, network architecture design, and so on. Among these factors, establishing a physics-guided dataset during the image restoration stage and avoiding unforeseen effects such as ringing pose a significant challenge for remote sensing image restoration. This letter proposes an optical imaging degradation simulation model and Transformer-based algorithm to improve remote sensing image quality. First, we model the degradation result from phase to image of optical remote sensing imaging using Zernike Polynomials, thus, a large-scale paired dataset is constructed. Then, a multi-level feature fusion transformer is introduced to mitigate the defect during restoration. The proposed algorithm incorporates a multi-level feature fusion module to fuse feature information from multi-scales effectively. Additionally, a multi-level space and frequency loss function is introduced to enhance the learning of high-frequency information to ensure that the edge suppresses noise amplification and ringing effects during recovery. Finally, experimental results on synthetic data show that our method improved by 25.4% and 22.3% with the blurred images on the PSNR index and SSIM index. Visual results on the GaoFen-1/2A PMS images have enhanced clarity and suppressed artifacts such as ringing which demonstrate the effectiveness and capability of our proposed method. Hua Wei 0007, Kun Gao 0001, Qiuyan Tang, Xiongxin Tang, Fanjiang Xu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | A Novel Intercalibration Method for Fengyun(FY)-3 VIRR Using MERSI Onboard the Same Satellite Based on Pseudo-Invariant PixelsabstractThis study presents a novel approach to the radiometric inter-calibration between two sensors onboard the same satellite based on pseudo-invariant pixels (PIPs) using iteratively re-weighted multivariate alteration detection (IR-MAD) method. The IR-MAD algorithm can statistically select pseudo-invariant pixels from the multispectral image pair to assess the radiometric differences between them. Analysis of multiple image pairs from different acquisition times can provide long-term inter-calibration results of the two sensors. The procedure is applied to Fengyun(FY)-3A&3B Visible Infrared Radiometer (VIRR), with the Medium Resolution Spectral Imager (MERSI) onboard the same platform as the reference. Consistency of the spatial distribution of the PIPs selected by IR-MAD with pseudo-invariant calibration sites (PICS) given by other scientists demonstrates the effectiveness of our method. The long-term time series trending of top-of-atmosphere VIRR reflectance over LIBYA1 and LIBYA4 after inter-calibration correction shows that the inter-calibrated VIRR has good agreement with MERSI, with a mean bias of less than 1% and an uncertainty of less than 2% for most channels. The approach requires no prior knowledge of the inter-calibration targets and extends PICS to the pixel-level targets, which results in more diverse samples, broader dynamic ranges and lower uncertainty, yielding consistent and reliable long-term inter-calibration results. Xiuqing Hu, Kun Gao 0001, Guorong Li, Na Xu 0001, Peng Zhang 0024 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Transformer and CNN Hybrid Network for Super-Resolution Semantic Segmentation of Remote Sensing ImageryabstractSuper-resolution semantic segmentation (SRSS) based on Convolutional neural network (CNN) cannot establish long-range dependencies due to limited receptive field, which limits the SRSS to obtain accurate high-resolution (HR) segmentation results from the low-resolution (LR) input images. In this paper, we design a Transformer and CNN hybrid SRSS network that consists of two branches: Transformer and CNN hybrid SRSS branch and super-resolution guided branch. In the Transformer and CNN hybrid SRSS branch, Transformer extracts global context information from the feature map of the CNN, while skip connection is used to retain the local context information extracted from the CNN and combines both features to further improve the segmentation performance. In addition, the super-resolution guided branch is designed to supplement rich structure information and guide the semantic segmentation (SS). We test the proposed method on the ISPRS Vaihingen benchmark data set, and our network is superior to other state-of-the-art methods. Kun Gao 0001, Hong Wang 0025, Xiaodian Zhang, Shuzhong Li |
IGARSS | 2 |
| 2023 | Pixel- And Patch-Wise Context-Aware Learning with CNN and GCN Collaboration for Hyperspectral Image ClassificationabstractGraph convolutional network (GCN) gains increasing attention in the hyperspectral image (HSI) classification by the ability to flexibly capture arbitrarily irregular objects. However, due to expensive computation, the graph construction is usually based on superpixel-wise nodes, which ignore the subtle pixel-wise features. In contrast, the convolution neural network (CNN) can mine pixel-wise spectral-spatial features but is limited to capturing local features in small square windows. In this paper, we design a new CNN and GCN collaborative network to simultaneously introduce pixel- and patch-wise contextual information. Concretely, we use the depthwise separable convolution to perform pixel-wise local feature extraction. To further mine the long-range contextual information between land covers, we concatenate a GCN. Finally, we further fuse the complementary features and decode them to obtain the classification map. Extensive experiments reveal that our method achieves competitive performance. Hong Wang 0025, Kun Gao 0001, Xiaodian Zhang, Zibo Hu, Zhijia Yang, Yuxuan Mao |
IGARSS | 2 |
| 2023 | Computer Vision-Aided mmWave UAV Communication SystemsabstractUnmanned aerial vehicle (UAV) communication systems usually operate in harsh scenarios, which require accurate information about the topology and wireless channel to achieve the desired transmission performance. Therefore, when millimeter-wave (mmWave) communication with its intrinsic Line-of-Sight (LoS) condition is adopted, accurate target localization is essential to determine the spatial relationship between the UAV and the grounded receivers (Rxs). In this article, a computer-vision (CV)-aided jointly optimization scheme of flight trajectory and power allocation is designed for mmWave UAV communication systems by utilizing the visual information captured via cameras equipped at the UAV. Compared with traditional schemes, the implementation cost and overhead can be greatly saved as no radio frequency transmissions are required in the proposed localization scheme. In addition, the transmit power at the UAV is jointly optimized with its flight trajectory in two different cases. Finally, simulation results are presented to demonstrate the efficiency of the proposed schemes. Zizheng Hua, Yang Lu 0008, Gaofeng Pan, Kun Gao 0001, Daniel B. da Costa 0001 |
IEEE Internet Things J. | 4 |
| 2023 | Successive Clustering-Based Outlier Resistant Band Selection Method for Hyperspectral Images With Spatial Information Difference MetricsabstractIn hyperspectral classification applications, band selection (BS) is an effective preprocessing method that reduces image redundancy without changing the original data. The property whereby different objects can be spatially separated is used for image classification, but BS methods based on quantitation of this property have not gotten enough attention. A cluster-based BS method that uses the dilation distances (DDs) with respect to the metric of spatial distances has been proposed, but the DD is strongly affected by outliers and calculating DD is time-consuming. Moreover, there is a mismatch between DD and the method of clustering and selecting representative band. In this letter, we propose a BS method based on pixel sorting-feature-based DD (SFDD) to accurately determine spatial information differences (SIDs) metric and design a method of successive clustering as well as a method of representative BS to match the features of this metric. We optimize the method to calculate the SFDD to reduce the time needed for it. In contrast to most BS methods, the bands selected by our method have a large SID among them such that objects at different positions are clearly differentiated in the spectral dimension after dimension reduction. The results of experiments showed that the proposed approach provides results that are competitive with those of several state-of-the-art methods. Kun Gao 0001, Xiaodian Zhang, Yunpeng Feng |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | EMO2-DETR: Efficient-Matching Oriented Object Detection With TransformersabstractObject detection in remote sensing is a challenging task due to the arbitrary orientations of objects and the vast variation in the number of objects within a single image. For instance, one image may contain hundreds of small vehicles, while another may only have a single football field. Recently, DEtection TRansformer (DETR) and its variants have achieved great success in object detection by setting a fixed number of object queries and using bipartite graph matching for one-to-one label assignment. However, we have observed that bipartite graph matching can result in relative redundancy of object queries when the number of objects changes dramatically in an image. This relative redundancy can cause two problems: slower convergence during training and redundant bounding boxes during inference. To analyze the aforementioned problems, we proposed a metric, Redundancy of Object Query (ROQ), to quantitatively analyze the redundancy. Through experiments, we discovered that the reason for the two issues is the difficulty in distinguishing between high-quality negative samples and positive samples. In this paper, we proposed Efficient-Matching Oriented Object Detection with Transformers(EMO2-DETR) consisting of three dedicated components to address the aforementioned issues. Specifically, Reassign Bipartite Graph Matching (RBGM) is proposed to extract high-quality negative samples from the negative samples. And Ignored Sample Predicted Head (ISPH) is proposed to predict high-quality negative samples. Then, Reassigned Hungarian loss is used to better involve high-quality negative samples in the update of model parameters. Extensive experiments on DOTAv1 and DOTAv1.5 datasets demonstrated that our proposed method achieves competitive results. Zibo Hu, Kun Gao 0001, Xiaodian Zhang, Hong Wang 0025, Zhijia Yang, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Hyperspectral Time-Series Target Detection Based on Spectral Perception and Spatial-Temporal Tensor DecompositionabstractThe detection of camouflaged targets in the complex background is a hot topic of current research. Existing hyperspectral target detection algorithms do not take advantage of spatial information and rarely use temporal information. It is difficult to obtain the required targets, and the detection performance in hyperspectral sequences with complex background will be low. Therefore, a hyperspectral time-series target detection method based on spectral perception and spatial-temporal tensor decomposition (SPSTT) is proposed. Firstly, a sparse target perception strategy based on spectral matching is proposed. To initially acquire the sparse targets, the matching results are adjusted by using the correlation mean of the prior spectrum, the pixel to be measured and the four-neighborhood pixel spectra. The separation of target and background is enhanced by making full use of local spatial structure information through local topology graph representation of the pixel to be measured. Secondly, in order to obtain a more accurate rank and make full use of temporal continuity and spatial correlation, a spatial-temporal tensor model based on the Gamma norm andL2,1norm is constructed. Furthermore, an excellent alternating direction method of multipliers is proposed to solve this model. Finally, spectral matching is fused with spatial-temporal tensor decomposition in order to reduce false alarms and retain more right targets. A 176-band hyperspectral image sequence (BIT-HSIS-I) dataset is collected for the hyperspectral target detection task. It is found by testing on the collected dataset that the proposed SPSTT has superior performance over the state-of-the-art algorithms. Xiaobin Zhao, Kun Gao 0001, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Constrained-Target Band Selection Based on Band Combination for Hyperspectral Target Detection Using CEMabstractSelecting an appropriate band subset is a vital problem for hyperspectral target detection. The constrained-target band selection method, derived from constrained energy minimization (CEM), can select different band subsets according to different targets. The selected band will be used to detect the target by CEM. However, the current methods did not consider the adaptation of the selected band with CEM. We propose a constrained target band selection method based on band combination, which treats the selected band as a whole to match the input requirements of CEM to detect the specific target. Firstly, Criteria for evaluating band combination is proposed based on the constrained-target band selection method. Secondly, the features of these band combinations are proposed. Thirdly, the key band set is proposed to reduce the range of search band combinations from the whole band to a subset of the whole band, based on these features and the sparse constrained band selection (SCBS) method. Finally, the desired band combination is searched in the key band set based on these features. Experiments show that the proposed method offers promising results. Kun Gao 0001, Xiaodian Zhang |
IGARSS | 2 |
| 2022 | Probability Differential-Based Class Label Noise Purification for Object Detection in Aerial ImagesabstractModern object detection for aerial images requires numerous annotated data. However, the data annotation process inevitably introduces noise due to the bird’s eye view perspective of aerial images and the professional requirements of annotations. While recent noise-robust object detection methods achieved great success, the noise side effect during the early training stage was still a problem. As demonstrated in this letter, noise during the early training stage will cumulatively affect the final performance. Based on the abovementioned observations, we propose a training strategy called correction maximization training to purify the noisy annotations and then train models. In particular, we design a novel noise filter called the probability differential (PD) to identify and revise wrong labels. After purification, we train the detector with the revised dataset. Compared with the existing works, the proposed method could be adapted in most modern object detectors (e.g., Faster RCNN and RetinaNet) and requires little hyperparameter tuning across different datasets and models. Extensive experiments on DOTA show that the proposed method achieves the state-of-the-art results with both symmetric and asymmetric noise. Zibo Hu, Kun Gao 0001, Xiaodian Zhang, Hong Wang 0025, Jiawei Han 0008 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Gradient Enhanced Dual Regression Network: Perception-Preserving Super-Resolution for Multi-Sensor Remote Sensing ImageryabstractMost existing learning-based single image super-resolution (SISR) methods mainly focus on improving reconstruction accuracy, but they always generate overly smoothed results that fail to match the visual perception. Although perceptual quality can be greatly improved via introducing adversarial loss, image fidelity may decrease to some extent. Moreover, most methods are trained and evaluated on simulated datasets and their performance would drop significantly on real remote sensing imagery. To solve the above problems, we propose a new SISR algorithm named gradient enhanced dual regression network (GEDRN). Based on the dual regression framework, we use share-source residual structure and non-local operation to learn abundant low-frequency information and long-distance spatial correlations. Besides, we not only introduce additional gradient information to avoid blurry results but also apply gradient loss and perceptual loss to further improve the perceptual quality. Our GEDRN is trained and tested on real-world multi-sensor satellite images. Experimental results demonstrate the superiority of the proposed method in achieving much better perceptual quality and ensuring high fidelity. Zhenzhou Zhang, Kun Gao 0001, Lei Min, Shijing Ji, Chong Ni, Dayu Chen |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | Band Selection of Hyperspectral Images Using Attention-Based AutoencodersabstractBand selection is an effective method to reduce redundancy in a hyperspectral image (HSI) without compromising the original contents. Popular band selection methods usually use strong assumptions, such as linear or nonlinear assumptions with simple predefined Kernel functions, to model the correlations between bands. However, this kind of strong assumption may not valid in the real environment due to the complex interactions between bands. In this letter, we treat hyperspectral band selection as a spectral reconstruction task. By assuming that an HSI can be sparsely reconstructed from a few informative bands, we propose an attention-based autoencoder to model the underlying nonlinear interdependencies between bands. The proposed model consists of two parts: an attention module and an autoencoder. The attention module is used to produce the attention mask which selects the most informative bands for every pixel. The autoencoder uses these informative bands to reconstruct the raw HSI. The final band selection is conducted via clustering column vectors of the attention mask and exploring the most representative band for each cluster. Different from most of the existing band selection methods, the proposed method directly learns global nonlinear correlations between bands without strong assumptions. The proposed model is easy to implement and all the parameters can be jointly optimized using the stochastic gradient descend algorithm. Experiments on three open public data sets show that the proposed method offers the promising results. Zeyang Dou, Kun Gao 0001, Xiaodian Zhang, Hong Wang 0025 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | Improving Performance and Adaptivity of Anchor-Based Detector Using Differentiable Anchoring With Efficient Target GenerationabstractMost anchor-based object detection methods have adopted predefined anchor boxes as regression references. However, the proper setting of anchor boxes may vary significantly across different datasets, improperly designed anchors severely limit the performances and adaptabilities of detectors. Recently, some works have tackled this problem by learning anchor shapes from datasets. However, all of these works explicitly or implicitly rely on predefined anchors, limiting universalities of detectors. In this paper, we propose a simple learning anchoring scheme with an effective target generation method to cast off predefined anchor dependencies. The proposed anchoring scheme, named as differentiable anchoring, simplifies learning anchor shape process by adding only one branch in parallel with the existing classification and bounding box regression branches. The proposed target generation method, including the$L_{p}$norm ball approximation and the optimization difficulty-based pyramid level assignment approach, generates positive samples for the new branch. Compared with existing learning anchoring-based approaches, the proposed method doesn’t require any predefined anchors, while tremendously improving performances and adaptiveness of detectors. The proposed method can be seamlessly integrated to Faster RCNN, RetinaNet, and SSD, improving the detection mAP by 2.8%, 2.1% and 2.3% respectively on MS COCO 2017 test-dev set. Moreover, the differentiable anchoring-based detectors can be directly applied to specific scenarios without any modification of the hyperparameters or using a specialized optimization. Specifically, the differentiable anchoring-based RetinaNet achieves very competitive performances on tiny face detection and text detection tasks, which are not well handled by the conventional and guided anchoring based RetinaNets for the MS COCO dataset. Zeyang Dou, Kun Gao 0001, Xiaodian Zhang, Hong Wang 0025 |
IEEE Trans. Image Process. | 2 |
| 2020 | Blind Hyperspectral Unmixing using Dual Branch Deep Autoencoder with Orthogonal Sparse PriorabstractBlind hyperspectral unmixing has become an important task for hyperspectral applications. In this paper, we propose a dual branch autoencoder with a novel sparse prior to simultaneously extract endmembers and abundances from the raw HSI. The dual branch structure extends the linear mixing model by only modeling linear mixtures of the endmembers and treating the bilinear interactions as error. In this way, the proposed model doesn't require the assumptions of explicit forms of bilinear interactions. The proposed sparse prior, named as orthogonal sparse prior, is based on the key observation that the abundance vector of one pixel is very sparse, there are often no more than two non-zero elements. Different from the conventional norm-based sparse prior which assumes the abundance maps are independent, the orthogonal sparse prior explores the orthogonality between the abundance maps. Extensive experiments on two real datasets show that the proposed method significantly and consistently outperforms the compared state-of-the-art methods, with up to 50% improvements. Zeyang Dou, Kun Gao 0001, Xiaodian Zhang, Hong Wang 0025 |
ICASSP | 2 |
| 2020 | Deep Learning-Based Hyperspectral Target Detection without Extra Labeled DataabstractTarget detection from hyperspectral images is an important problem. Recently, several deep learning-based target detection algorithms have been proposed. However, most of them require extra well-labeled data to train detectors. In this paper, we propose a deep learning-based target detection algorithm that doesn't require any extra labeled data. The proposed detector is based on the siamese network and the low-rank-sparse autoencoder. The autoencoder separates the test spectrum into a low-rank component and a sparse component, based on the assumption that the normal spectrum space has a low-rank structure while outliers sparsely spread in the image. The low-rank output of the autoencoder and the target spectrum are then separately fed into the Siamese network to get two high level features, and the final cosine similarity score is computed based on two features. To properly train the proposed detector, we develop a data creation method that creates numerous simulative training data. Extensive experiments show that the proposed method achieves state-of-the-art results. Zeyang Dou, Kun Gao 0001, Xiaodian Zhang, Hong Wang 0025 |
IGARSS | 2 |
| 2020 | Noise Resistant Focal Loss for Object Detection
Zibo Hu, Kun Gao 0001, Xiaodian Zhang, Zeyang Dou |
PRCV (2) | 2 |
| 2020 | On-Orbit MTF Estimation for GF-4 Satellite Using Spatial Multisampling on a New TargetabstractGF-(Gaofen- means high resolution in Chinese) satellite launched in 2015 is the first geosynchronous orbit remote sensing satellite in China. To evaluate the on-orbit modulation transfer function (MTF) of the space-borne panchromatic camera in GF-4, a modified pulse target is proposed. This new type target is laid in the uniform low-reflection background region whose center region is a high-reflection square with the size of 3~4 ground sampled distance (GSD), and then two low-reflection rectangle stripes are extended along the center square toward both ends with the length over 4 GSD. In consideration of the staring imaging mechanism of GF-4, the target image sequences are accessed with relative random distance taken by the space-borne camera in a short period to reduce accidental error of one sampling and enhance the accuracy of estimation. Based on the multi-sampling calibrating images, maximum a posteriori estimation model and gradient descent solution with grid searching method are used to fit the pulse response function (PRF) when taking the target center positions as references. MTF is then calculated from PRF via Fourier transformation. Numerical simulation results reveal our method can keep the accuracy stable despite of different kinds of noise. Actual calibration result using this method shows that on-orbit MTF of GF-4 camera at Nyquist frequency is 0.1473. Kun Gao 0001, Zeyang Dou, Hong Wang 0025, Xingke Fu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2020 | Hyperspectral Unmixing Using Orthogonal Sparse Prior-Based Autoencoder With Hyper-Laplacian Loss and Data-Driven Outlier DetectionabstractHyperspectral unmixing, which estimates end-members and their corresponding abundance fractions simultaneously, is an important task for hyperspectral applications. In this article, we propose a new autoencoder-based hyperspectral unmixing model with three novel components. First, we propose a new sparse prior to abundance maps. The proposed prior, called orthogonal sparse prior (OSP), is based on the observations that different abundance maps are close to orthogonal because, generally, no more than two end-members are mixed within one pixel. As opposed to the conventional norm-based sparse prior that assumes the abundance maps are independent, the proposed OSP explores the orthogonality between the abundance maps. Second, we propose the hyper-Laplacian loss to model the reconstruction error. The key observation is that the reconstruction error distribution usually has a heavy-tailed shape, which is better modeled by the hyper-Laplacian distribution rather than the commonly used Gaussian distribution. Third, to ease the side effect of outliers for end-member initializations, we develop a data-driven approach to detect outliers from the raw hyperspectral images. Extensive experiments on both synthetic and real-world data sets show that the proposed method significantly and consistently outperforms the compared state-of-the-art methods, with up to more than 50% improvements. Zeyang Dou, Kun Gao 0001, Xiaodian Zhang, Hong Wang 0025 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Spatio-temporal super-resolution for multi-videos based on belief propagation
Tinghua Zhang, Kun Gao 0001, Guoqiang Ni, Guihua Fan |
Signal Process. Image Commun. | 2 |