EDBT 2026 Demo / reviewers in the wild / expert
Shitian He
dblp:287/8996
· DBLP profile ↗
13ranked-venue papers
4as first author
13since 2021 · last 2026
0000-0001-9696-8865ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A mutual information-based framework for generalized image fusion via common-unique decoupling
Liyuan Pan, Huanxin Zou, Jun Li 0020, Hao Chen 0046, Xinyi Ying, Shitian He, Yingqian Wang 0002 |
Knowl. Based Syst. | 7 |
| 2026 | D2-DETR:DETR With Dual-Domain frequency-spatial modeling for unmanned aerial vehicle imagery object detection
Xuanming Liu, Huanxin Zou, Jun Li 0020, Liyuan Pan, Shitian He, Jiangshan Li, Wanyu Chen |
Knowl. Based Syst. | 5 |
| 2025 | Exploring cross-branch information for semi-supervised remote sensing object detection
Shitian He, Huanxin Zou, Yingqian Wang 0002, Hao Chen 0046, Ning Jing |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | MambaRSIS: Context-aware multi-scale feature aggregation with selective state space model for remote sensing instance segmentationabstractRemote sensing instance segmentation aims to detect and assign pixel-level labels to each instance in remote sensing images, which holds critical engineering significance for both civil and military applications. While existing domain-specific methods have made progress, they still struggle with three persistent challenges: ineffective context modeling in cluttered backgrounds, information loss during multi-scale feature fusion, and blurred boundaries for densely clustered small objects. To address these limitations, we propose a novel remote sensing instance segmentation framework with three artificial intelligence (AI) methodological innovations, which comprises: a Context Perception Module (CPM) for context modeling, a Context Guided Multi-Scale Feature Aggregation (CGFA) method for multi-scale feature fusion, and a Multi-Path Region Proposal Extractor (MPRPE) with boundary-refined segmentation. The CPM leverages the selective state space model (Mamba) to capture long-range contextual information, effectively addressing the issue of cluttered backgrounds in remote sensing images. The CGFA replaces standard feature pyramid network architecture which is limited by direct summation or concatenation, preserving fine-grained spatial details with context guidance. The MPRPE and boundary-aware segmentation head mitigate the challenges of missed detection of small objects and blurred edge predictions, which arise from the clustered distribution of small objects and semantic ambiguity. Extensive experiments on the challenging iSAID and NWPU VHR-10 datasets validate the proposed method’s consistent improvements across metrics while demonstrating its practical engineering impact on remote sensing interpretation systems. Liyuan Pan, Huanxin Zou, Hao Chen 0046, Shitian He, Xuanming Liu, Jiangshan Li, Wanyu Chen |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Multimodal image generation and fusion through content-style hybrid disentanglementabstract• Research highlight 1: We propose a novel cross-task hybrid training methodology for multimodal images, offering a simple yet unified solution that simultaneously addresses both image generation and fusion tasks. • Research highlight 2: Building upon mutual-supervised multimodal image pairs, we innovatively integrate single-modality self-supervision to develop a hybrid-supervised decoupling framework with a dedicated loss function, achieving robust separation of content-style representations. • Research highlight 3: Extensive experiments spanning on four modalities and seven popular datasets demonstrate our method’s consistent superiority and impressive cross-task capability. Ablation studies further reveal that our framework learns generalized representations transferable across different image processing tasks. Multimodal image fusion and cross-modal translation are fundamental yet challenging tasks in computer vision, with their performance directly impacting downstream applications. Existing approaches typically treat these tasks independently, developing specialized models that fail to exploit the intrinsic relationships between different modalities. This limitation not only restricts model generalizability but also hinders further performance improvements. In this paper, we propose a joint optimization framework for image generation and fusion. Specifically, we generalize multimodal image tasks as the fusion and transformation of cross-modal features, and design a hybrid task training strategy. At the data level, we introduce a self-supervised and mutual-supervised hybrid mechanism for content-style feature decoupling, which achieves superior feature separation through stepwise training on intra-modal and cross-modal data. At the model level, we construct a triple-branch decoupling head along with fusion and transformation modules to ensure synchronous and efficient execution of dual tasks. Our method not only breaks through the single task limitation of the model, but also innovatively introduces mixed supervision into multimodal processing. We conduct comprehensive experiments covering four modalities fusion tasks on seven popular datasets. Extensive experimental results demonstrate that our method achieves superior performance on two tasks as compared of the respective state-of-the-art methods, and show impressive cross-task generalization capability. Huanxin Zou, Jun Li 0020, Hao Chen 0046, Xinyi Ying, Shitian He, Yingqian Wang 0002, Liyuan Pan |
Knowl. Based Syst. | 6 |
| 2024 | YOLOX-Drone: An Improved Object Detection Method for UAV ImagesabstractUnmanned aerial vehicles (UAV) are widely used for their small size and flexibility. However, the large number of small objects and the significant difference in object size in UAV images bring great challenges to the detection task. Therefore, we propose an object detection method for UAV images with four improvements on the strong baseline model YOLOX-S, which is robust to detect small objects and multi-scale objects. Firstly, we introduce a high-resolution feature map to retain rich detailed information about small objects. Secondly, we propose new up-sampling and down-sampling modules to reduce the feature information loss during the sampling process. Thirdly, we present the triple-scale feature fusion module (TSFFM) to fuse more abundant multi-scale features in the neck’s bottom-up feature fusion process. Finally, the parrell dilated convolution attention module (PD-CAM) is proposed to learn the multi-receptive field features. Experiment results on the VisDrone-VID2019 dataset validate the effectiveness and superiority of the proposed method. Huanxin Zou, Shitian He, Shuo Liu 0015, Liyuan Pan |
IGARSS | 3 |
| 2024 | Learning Remote Sensing Object Detection With Single Point SupervisionabstractPointly Supervised Object Detection (PSOD) has attracted considerable interests due to its lower labeling cost as compared to box-level supervised object detection. However, the complex scenes, densely packed and dynamic-scale objects in Remote Sensing (RS) images hinder the development of PSOD methods in RS field. In this paper, we make the first attempt to achieve RS object detection with single point supervision, and propose a PSOD method tailored for RS images. Specifically, we design a point label upgrader (PLUG) to generate pseudo box labels from single point labels, and then use the pseudo boxes to supervise the optimization of existing detectors. Moreover, to handle the challenge of the densely packed objects in RS images, we propose a sparse feature guided semantic prediction module which can generate high-quality semantic maps by fully exploiting informative cues from sparse objects. Extensive ablation studies on the DOTA dataset have validated the effectiveness of our method. Our method can achieve significantly better performance as compared to state-of-the-art image-level and point-level supervised detection methods, and reduce the performance gap between PSOD and box-level supervised object detection. Code is available at https://github.com/heshitian/PLUG. Shitian He, Huanxin Zou, Yingqian Wang 0002, Boyang Li 0007, Ning Jing |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Dense Contrastive Learning Based Object Detection for Remote Sensing ImagesabstractSupervised learning based object detectors suffer from the high cost and difficulty of labeling datasets. Self-supervised learning methods require no manual annotations. However, the misalignment between the pretext task designed for image classification and the downstream task affects the detection performance. Therefore, this paper proposes a self-supervised dense contrastive learning method to improve performance of object detection in remote sensing images. Specifically, first, Swin Transformer substitutes popular CNN to extract features of augmented multiple views. Second, global and local features are extracted using parallel global and dense projector heads, respectively. Third, a predictor head is added to increase the nonlinear transformations in the network. Extensive experiments on the NWPU VHR-10 dataset show that the proposed method outperforms two representative strong baseline methods, including MoCoV2 and DenseCL. Shuo Liu 0015, Huanxin Zou, Shitian He, Li Sun 0009 |
IGARSS | 5 |
| 2022 | Semantic Segmentation of High-Resolution Remote Sensing Images Based on Sparse Self-AttentionabstractSemantic segmentation of high-resolution optical remote sensing images is an important but challenging task. To solve the problem that many semantic segmentation networks fail to efficiently utilize global and local context information to improve the segmentation performance, this paper proposes a semantic segmentation network based on sparse self-attention (SDANet) to model the global context dependencies. Specifically, the feature maps are first divided into four regions in spatial and channel dimensions, respectively, and the divided feature maps are rearranged to form new regions. Second, the position and channel self-attention operations are performed on the rearranged regions. Third, the feature maps are restored to the original combination and the position together with channel self-attention operations are performed again to obtain the output feature maps. Finally, semantic segmentation is completed based on the output feature maps. Extensive experiments conducted on the ISPRS Vaihingen dataset demonstrate that the proposed method is superior to self-attention-based DANet, CCNet, and other general semantic segmentation networks, such as FCN, Deeplabv3+, HRNet, etc. Li Sun 0009, Huanxin Zou, Shitian He, Shuo Liu 0015 |
IGARSS | 6 |
| 2022 | Generative Adversarial Network for SAR-to-Optical Image Translation with Feature Cross-Fusion InferenceabstractThe translation of synthetic aperture radar (SAR) to optical images provides a new solution for the interpretation of SAR images. Most of the existing translation networks are based on generative adversarial networks and use 9-residual blocks or U-Net structures in the feature inference phase. Both structures cause a large amount of information lost during the conversion of SAR image features to optical features, making the outline of the translated image blurred or semantic information lost. Aiming at this problem, this paper proposes a cross-fusion inference network structure, which preserves both high-resolution features and low-resolution features in the whole process of feature inference. Our proposed method broadens the network horizontally while deepening it vertically and improving the image translation performance. The experiments conducted on the public dataset sen1-2 show that the proposed method is superior to other networks. Huanxin Zou, Li Sun 0009, Shitian He, Shuo Liu 0015 |
IGARSS | 6 |
| 2022 | Enhancing Mid-Low-Resolution Ship Detection With High-Resolution Feature DistillationabstractTo enhance mid–low-resolution ship detection, existing methods generally use image super-resolution (SR) as a preprocessing step and feed the super-resolved images to the detectors. However, these methods only use high-resolution (HR) images as ground-truth labels to supervise the training of their SR module but overlook the rich HR information in the detection stage. Inspired by the recent advances in knowledge distillation, in this letter, we design a feature distillation framework to fully exploit the information in ground-truth HR images to handle mid–low-resolution ship detection. Our framework consists of a student network and a teacher network. The student network first super-resolves input images using an SR module and then feeds the super-resolved images to the detection module. The teacher network whose architecture is the same as the student detection module directly takes HR images as input to generate HR feature representation and then distills these HR features to the student network through a distillation loss. Using our feature distillation framework, HR images are not only used as ground-truth labels to train the SR module but also provide “ground-truth” features to train the detection module, which enhances the detection performance of the student network. We apply our framework to several popular detectors, includingFCOS,Faster-RCNN,Mask-RCNN, andCascase-RCNN, and conduct extensive ablation studies to validate its effectiveness and generality. Experimental results on the HRSC2016, DOTA, and NWPU VHR-10 datasets demonstrate that, when applying our framework toFaster-RCNN, our method can outperform several state-of-the-art detection methods in terms of mAP50 and mAP75. Shitian He, Huanxin Zou, Yingqian Wang 0002, Runlin Li |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2021 | Rotated Hybrid Task Cascade Network for Remote Sensing Aircraft Target RecognitionabstractAutomatic aircraft target recognition, including direction detection and fine-grained classification, is an important but challenging problem. Multi-directional densely arranged targets and the tiny differences between classes cause difficulties in recognition and direction prediction. To overcome the aforementioned problems, a rotated hybrid task cascade (RHTC) network is proposed. Specifically, RHTC cascades the segmentation branch and the bounding-box (bbox) branch to fuse the semantic feature in a coarse- to- fine manner. In addition, a new oriented bounding box regressor (OBBR) is proposed to predict the direction of target, and a new directionalloss function is added to further optimize the regressor. Moreover, we design fine masks in preprocessing to achieve improved recognition performance. The experimental results evaluated on the datasets collected from Google Earth show that RHTC can achieve the state-of-the-art performance on self-defined direction precision (DP) and mean average precision (mAP). Huanxin Zou, Runlin Li, Shitian He, Li Sun 0009 |
IGARSS | 5 |
| 2021 | Shipsrdet: An End-to-End Remote Sensing Ship Detector Using Super-Resolved Feature RepresentationabstractHigh-resolution remote sensing images can provide abundant appearance information for ship detection. Although several existing methods use image super-resolution (SR) approaches to improve the detection performance, they consider image SR and ship detection as two separate processes and overlook the internal coherence between these two correlated tasks. In this paper, we explore the potential benefits introduced by image SR to ship detection, and propose an end-to-end network named ShipSRDet. In our method, we not only feed the super-resolved images to the detector but also integrate the intermediate features of the SR network with those of the detection network. In this way, the informative feature representation extracted by the SR network can be fully used for ship detection. Experimental results on the HRSC dataset validate the effectiveness of our method. Our ShipSRDet can recover the missing details from the input image and achieves promising ship detection performance. Shitian He, Huanxin Zou, Yingqian Wang 0002, Runlin Li |
IGARSS | 1 |