VLDB 2026 Research / reviewers in the wild / expert
Xikun Hu
dblp:192/6484
· DBLP profile ↗
12ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0002-1077-2560ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust Fine-Grained Oriented Ship Detection for Remote Sensing Imagery via Controllable Generative PretrainingabstractFine-grained ship recognition in remote sensing imagery is essential for maritime applications. However, its development is hindered by two challenges: 1) the limited granularity of existing ship detection datasets, and 2) the disturbance of complex maritime conditions as well as the arbitrary ship orientations and distributions. To address the first issue, we annotated a large-scale fine-grained ship instance detection dataset (LAFI), comprising 48,717 ship instances worldwide with 49 categories. To tackle the challenges of marine disturbance and diverse ship status, we proposed a controllable generative knowledge-driven ship detection framework (COSD). It employs a controllable diffusion model guided by ship-marine textual prompt to generate millions of synthetic images that not only preserve ship structures but also cover diverse sea and weather conditions for robust pretraining. The pretraining stage then utilizes masked reconstruction to learn component-level cues under occlusion, clutter, fog, and illumination changes. Furthermore, a heterogeneous feature alignment decoder is designed to align multi-modal metrics of orientation and distribution features in the latent space, allowing for accurate representation of diverse ship status. Extensive experiments on two benchmark datasets showed that our method respectively increased 0.011 and 0.030 mean average precision (mAP@50) over SOTA methods, particularly in scenarios involving small, densely packed and arbitrary oriented ships. Da He, Xikun Hu, Ping Zhong 0001, Qian Shi 0001, Xiaoping Liu 0001, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | Spectral-Feature-Guided Controllable Diffusion for SAR-to-Optical Satellite Imagery Generation in Wildfire Mapping
Yushan Zou, Xikun Hu, Ping Zhong 0001 |
PRCV (15) | 2 |
| 2025 | Discriminative Latent-Space Learning for Fine-Grained Object Detection in Remote Sensing ImagesabstractAbstract—Fine-grained object detection (FOD) is essential in many remote sensing image interpretation tasks. Existing FOD methods have achieved remarkable progress in modeling discriminative features for FOD in remote sensing images. However, they receive unsatisfactory recognition accuracy due to the curse of dimensionality (CoD) problem. In this paper, we propose an orthogonal constraint-based discriminative latent-space learning (DLL) method to address the CoD problem. We first optimize a sparse optimization paradigm with convex relaxation to extract shared features between fine-grained objects into a low-dimensional latent-space. Then, we solve the orthogonal space of the latent-space for extracting irrelevant features related to common features from input features, i.e., discriminative features. The sparse optimization paradigm reprojects the underlying trends hidden in high-dimensional feature space into a low-dimensional latent space and hence addresses the CoD problem. We use neural parameters with latent-space and orthogonal constraints to approximately solve the proposed DLL, which can be efficiently optimized under convex programming tools. We theoretically prove the effectiveness of the proposed DLL. By adding our DLL to the existing deep learning-based object detection method, extensive experiments conducted on two datasets demonstrate that our method achieves superior performance compared with other state-of-the-art methods. Xikun Hu, Wenlin Liu, Xiangsheng Wang, Chen Chen 0152, Ya Jiang, Ping Zhong 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | MSP-O: A Two-Stage Mask Pretraining Method With Object Focus for Remote Sensing Object DetectionabstractPretraining is showing fundamental importance in computer vision, to initialize the backbones of deep learning models with large natural image datasets, e.g., ImageNet. However, there is seldom a special pretraining technique focused on remote sensing object detection due to the distinct characteristics of remote sensing images, such as specific imaging angles, complex backgrounds, and various target categories. This article aims to bridge the gap between natural image data and remote sensing image data and to facilitate pretraining for the specific object detection task. This article introduces masked remote sensing pretraining with object (MSP-O), an innovative two-stage masked pretraining method. Traditional single-stage pretraining frameworks aim to extract intrinsic features from large-scale image datasets to learn object discrimination and localization ability. MSP-O builds upon this by adding a second-stage pretraining phase to learn specific features related to remote sensing objects of interest themselves, facilitating the downstream object detection task. Specifically, MSP-O incorporates a pixel-level restoration task focused on the objects of interest, guiding the model to prioritize the objects to be detected. Compared with traditional pretraining methods, MSP-O is able to learn feature representations that are more relevant to the detection task, especially when the data scale is limited. Extensive experiments on the DOTA and DIOR datasets demonstrate the effectiveness of MSP-O, showing substantial improvements in detection performance compared with conventional methods. Wenlin Liu, Xikun Hu, Ping Zhong 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Active Object Detection for UAV Remote Sensing via Behavior Cloning and Enhanced Q-Network With Shallow FeaturesabstractObject detection in Unmanned Aerial Vehicle (UAV) remote sensing imagery faces two critical challenges: the inability to effectively handle multi-scale targets and the difficulty in addressing partial occlusions, which significantly impact detection accuracy in real-world applications. To overcome these limitations, we propose an active object detection (AOD) method that dynamically adjusts viewing angles and scales during the detection process. Our key innovation lies in a novel architecture that uniquely combines shallow Feature Pyramid Network features with detector output bounding boxes, combining with a self-supervised learning framework. The technical originality of our approach is further enhanced by our two-stage learning methodology, which initially employs behavior cloning to establish robust foundational performance, followed by Q-learning fine-tuning to optimize detection strategies. To facilitate comprehensive evaluation, we introduce CARLA-AOD, a new benchmark dataset that encompasses 18 diverse scenarios across hemispheric space above targets, specifically designed for UAV remote sensing AOD applications. Extensive experimental validation demonstrates the effectiveness of our approach, achieving substantial improvements over baseline detectors across multiple datasets: 7.9% on Small Airport, 14.0% on Virtual Park, and 10.2% on our CARLA-AOD dataset. The two-stage learning process proves particularly effective, with Q-learning fine-tuning providing an additional performance boost of up to 7.3% beyond the behavior cloning baseline. Moreover, our method achieves a 60.2% reduction in inference time compared to the state-of-the-art DCCL algorithm, making it particularly suitable for time-critical applications such as emergency response and real-time surveillance. The dataset is available at IEEE Dataport. IEEE Dataport. Zhuocheng Zou, Xikun Hu, Ping Zhong 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Weakly Misalignment-Free Adaptive Feature Alignment for UAVs-Based Multimodal Object DetectionabstractVisible-infrared (RGB-IR) image fusion has shown great potentials in object detection based on unmanned aerial ve-hicles (UAVs). However, the weakly misalignment problem between multimodal image pairs limits its performance in object detection. Most existing methods often ignore the modality gap and emphasize a strict alignment, resulting in an upper bound of alignment quality and an increase of implementation costs. To address these challenges, we propose a novel method named Offset-guided Adaptive Feature Alignment (OAFA), which could adaptively adjust the relative positions between multimodal features. Considering the impact of modality gap on the cross-modality spa-tial matching, a Cross-modality Spatial Offset Modeling (CSOM) module is designed to establish a common sub-space to estimate the precise feature-level offsets. Then, an Offset-guided Deformable Alignment and Fusion (ODAF) module is utilized to implicitly capture optimal fusion po-sitions for detection task rather than conducting a strict alignment. Comprehensive experiments demonstrate that our method not only achieves state-of-the-art performance in the UAVs-based object detection task but also shows strong robustness to the weakly misalignment problem. Chen Chen 0152, Jiahao Qi, Kangcheng Bin, Ruigang Fu, Xikun Hu, Ping Zhong 0001 |
CVPR | 6 |
| 2024 | Near Real-Time Burned Area Progression Mapping With Multispectral Data Using Ensemble LearningabstractMonitoring the wildfire progression is essential to quantify the fire-disturbance areas for emergency responses. To combine the advantages of pixelwise machine learning (ML) method and region-based deep learning (DL) segmentation model, this study proposes a two-phase hybrid framework for near real-time burned area progression mapping: the first one intends to depict burned area delimitation using a contextual algorithm HRNet to exclude the unburned areas outside the perimeter and minimize omission errors, which partially remain unburned patches within the delimitation as commission errors. The second phase refines the burned area spatially using ensemble fusion based on an updating support vector machine (SVM) model under the voting scheme as new imagery arrives to reduce the commission errors consecutively. The validation results showed that the accuracy of perimeter prediction using the HRNet can reach 96.77% in Kappa. The iterative optimization can improve the average Kappa value from 62.55% to 70.75% for burned area pixel classification using pixelwise SVM alone. The proposed ensemble learning framework can further refine the burned area progression results, reaching an average Kappa up to 85.19%, at four acquisition dates with Sentinel-2 and Landsat-8 available during the Sand fire event that occurred in California. Xikun Hu, Puzhao Zhang, Ka-Veng Yuen, Ping Zhong 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2024 | DSRNet: Diagonal Subsampling Reconstruction Network for Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection seeks to locate pixels in a scene exhibiting substantial spectral discrepancies from the surrounding background pixels, holding essential applications in both civilian and military domains. However, in real-world scenarios, the inherent properties of hyperspectral imaging, the irregular forms of anomalous targets, and the lack of prior information pose significant challenges for anomaly detection methodologies. To address these issues, we first explore a novel reconstruction-based modeling approach for hyperspectral anomaly detection, offering a rational motivation for the modeling approach and a detailed exposition of its effective implementation. Furthermore, we propose a diagonal subsampling reconstruction network (DSRNet) for anomaly detection of hyperspectral data. Specifically, DSRNet consists of a paired training data generation algorithm using subsampling and a self-supervision training process enforced with a reconstruction consistency constraint. The training input is derived by a randomly diagonal averaging subsampler, where training pairs are derived from the same original hyperspectral data. Extensive experiments on four public datasets demonstrate the superiority of our DSRNet compared over several state-of-the-art baselines, with an average AUC score increase of 0.0156. Xikun Hu, Ping Zhong 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Empowering Physical Attacks With Jacobian Matrix Regularization Against ViT-Based Detectors in UAV Remote Sensing ImagesabstractVision transformers (ViTs) have achieved great success in unmanned aerial vehicle (UAV) target detection tasks. However, little attention has been paid to the adversarial attack against ViT-based detectors, and the generated adversarial examples cannot take physical realizability and attack transferability into account at the same time. To overcome the limitation, we focus on transferable attacks toward ViT-based detectors in optical UAV-based remote sensing images and generate adversarial examples in the physical world. Concretely, we design unique perturbation patches deployed within and beyond the target object rather than requiring the patches to be aligned with image tokens. To narrow the gap between limited digital samples and complex physical scenarios, we conduct data augmentation on training images at global and local levels. In addition, we propose a novel transferable attack method named Jacobian matrix regularization (JMR), which consists of feature variance regularization (FVR) and attention weight regularization (AWR). Specifically, FVR calculates feature variances of different channels within specific layers and then sets the features as zeros for channels with top variances. AWR is achieved by masking the largest self-attention weights. We conduct extensive transferable experiments with typical detectors in both digital and physical UAV-based remote sensing scenarios. The results indicate that our method could achieve competitive transferability compared with state-of-the-art methods. Yu Zhang 0221, Zhiqiang Gong, Wenlin Liu, Jiahao Qi, Xikun Hu, Ping Zhong 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | Remote Respiratory and Cardiac Motion Patterns Separation With 4D Imaging RadarsabstractRadar-based noncontact physiological signals monitoring is meaningful for daily health monitoring, post-disaster rescue, and public security. This paper focuses on the theoretical and experimental study of noncontact respiratory and cardiac motion signals separation by remote sensing using a four-dimensional (4D) imaging radar. To adaptively separate respiratory and cardiac motion patterns, we propose a variational mode separation (VMS) algorithm. VMS is established on optimizing a variational problem to separate different modes. It minimizes the energy overlap of the heartbeat and respiration signals as well as their harmonics with an equality constraint. Both simulation and real scene data results show that the proposed VMS algorithm is suitable for separating the weaker cardiac motion pattern from the strong respiratory motion pattern, restraining the influence of respiration harmonics on the heartbeat component. Furthermore, we have implemented continuous remote monitoring of respiratory rate (RR) and heart rate (HR) by employing the proposed method in a real scene. The results validate the consistency with the reference respiration belt and electrocardiogram (ECG). The root mean square errors (RMSEs) of RR and HR for the remote measurement are 0.13 breaths per minute (brpm) and 1.7 beats per minute (bpm), respectively. Zhi Li 0081, Tian Jin 0001, Xikun Hu, Yongkun Song, Zhenqun Sang |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | Gan-based SAR to Optical Image Translation in Fire-Disturbed RegionsabstractClimate change by anthropogenic warming leads to increases in dry fuels and promotes forest fires. Multispectral images' quality is easily affected by poor atmospheric conditions. SAR satellite sensors can penetrate through clouds and image day and night. However, the burned area mapping methods widely used for optical data are not feasible to be applied for SAR data owing to the differences in imaging mechanisms. Recent advances in deep image translation can fill this gap by using Generative Adversarial Networks (GAN). In this research, we apply a GAN-based model for SAR to optical image translation over fire-disturbed regions. Specifically, Sentinel-1 SAR images are translated into Sentinel-2 images using the ResNet-based Pix2Pix model, which is trained on 281 large fire events and tested on the other 23 events in Canada. The generated images preserve the spectral characteristics well and show high similarity to the real images with Structure Similarity Index Measure (SSIM) over 0.59. Xikun Hu, Puzhao Zhang, Yifang Ban |
IGARSS | 1 |
| 2022 | Wildfire-S1S2-Canada: A Large-Scale Sentinel-1/2 Wildfire Burned Area Mapping Dataset Based on the 2017-2019 Wildfires in CanadaabstractWildfires vary across space and time, precisely and timely mapping on the wildfire affected areas is critical for wildfire management, population and property protection, and environmental impact assessment. In this study, we established a large-scale annotated wildfire burned area dataset based on freely available Sentinel-1 SAR and Sentinel-2 multispectral instrument (MSI) data and Canada Wildfire Burned Area Database. This dataset includes bi-temporal Sentinel-1 and Sentinel-2 images, which allows users to exploit remotely sensed data acquired in both optical and microwave domains. On the proposed dataset, we achieved the highest IoU score of 0.86 on the Sentinel-2 data with Siamese U-Net, and the highest IoU score of 0.80 on the Sentinel-1 data using U-Net with early fusion. The combined use of Sentinel-1 and Sentinel-2 failed to bring significant improvement compared to Sentinel-2 based results, but this dataset may have the potential to boost Sentinel-1 based results with Sentinel-2 data for near real-time wildfire progression mapping. Puzhao Zhang, Xikun Hu, Yifang Ban |
IGARSS | 2 |