EDBT 2026 Demo / reviewers in the wild / expert
Zhicheng Zhao 0002
dblp:55/7547-2
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0002-2761-7399ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MambaEVT: Event Stream-Based Visual Object Tracking Using State Space ModelabstractEvent camera-based visual tracking has drawn more and more attention in recent years due to the unique imaging principle and advantages of low energy consumption, high dynamic range, and dense temporal resolution. Current event-based tracking algorithms are gradually hitting their performance bottlenecks, due to the utilization of vision Transformer and the static template for target object localization. In this paper, we propose a novel Mamba-based visual tracking framework that adopts the state space model with linear complexity as a backbone network. The search regions and target template are fed into the vision Mamba network for simultaneous feature extraction and interaction. The output tokens of search regions will be fed into the tracking head for target localization. More importantly, we consider introducing a dynamic template update strategy into the tracking framework using the Memory Mamba network. By considering the diversity of samples in the target template library and making appropriate adjustments to the template memory module, a more effective dynamic template can be integrated. The effective combination of dynamic and static templates allows our Mamba-based tracking algorithm to achieve a good balance between accuracy and computational cost on multiple large-scale datasets, including EventVOT, VisEvent, and FE240hz. The source code and checkpoint have been released on https://github.com/Event-AHU/MambaEVT. Xiao Wang 0014, Shiao Wang, Xixi Wang 0005, Zhicheng Zhao 0002, Lin Zhu 0012, Bo Jiang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | UAV Video Vehicle Detection: Benchmark and BaselineabstractWith the increasing application of unmanned aerial vehicles (UAVs) in intelligent transportation systems, vehicle object detection in UAV videos has received increasing attention. Precise categorization and detection for vehicles in UAVs is important in many practical applications. However, existing object detection methods, tailored for natural images, often fall short of accurately identifying vehicle objects. Additionally, high-altitude UAV imaging mainly employs horizontal bounding box annotation, frequently leading to significant obstruction and overlapping. Hence, we propose a new task called UAV video vehicle detection (VVD) to achieve precise detection and categorization of vehicles in high-altitude UAV imaging environments. To facilitate the research and development of UAV VVD, we construct the first large-scale well-annotated benchmark UAV VVD dataset, which includes 70 UAV videos captured at a 500-m altitude, with 361489 vehicle instances annotated by the oriented bounding boxes and vehicle categories. Moreover, we introduce a novel category refinement network (CRNet) approach that extracts and refines vehicle object features from the bounding box of the detection results to classify vehicle categories. This approach effectively eliminates the interference of the background and other vehicle objects in candidate boxes. Notably, the vehicle object features are projected into subspace, enabling the category refinement module (CRM) to focus more on the distinctive characteristics of the vehicle object itself through normalization operations. We conduct extensive experiments on the proposed VVD dataset. Experimental results demonstrate the superiority and effectiveness of the proposed CRNet method. The relevant code and dataset are available athttps://github.com/mmic-lcl. Yun Xiao 0003, Jinfa Wang, Zhicheng Zhao 0002, Bo Jiang 0002, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Reflectance-Guided Progressive Feature Alignment Network for All-Day UAV Object DetectionabstractObject detection using visible-infrared images has become increasingly crucial for all-day applications of unmanned aerial vehicle (UAV). However, existing multi-modal detection methods face significant challenges in low-light conditions, where degraded visible image quality exacerbates weak alignment issues and compromises feature fusion effectiveness. Although recent approaches have attempted to address these issues through cross-attention mechanisms or feature alignment strategies, they often suffer from unstable performance and limited generalization capability in challenging nighttime scenarios. To address these limitations, we propose a novel Reflectance-Guided Progressive Feature Alignment Network (RGFNet) for robust UAV object detection. Our proposed method leverages the illumination-invariant characteristic of reflectance features decomposed from visible images via Retinex theory to guide cross-modal alignment and fusion. Specifically, we design a Reflectance-Guided Collaborative Alignment Module (RCAM) that utilizes reflectance guidance to perform bidirectional feature alignment between visible and infrared modalities, effectively reducing position misalignment under varying lighting conditions. Furthermore, we introduce a Light-Aware Selective Fusion Module (LSFM) that maps multi-modal features into a shared hidden state space through selective state space mechanism, enabling efficient feature interaction while maintaining linear computational complexity. Extensive experiments on two challenging UAV detection benchmarks, DroneVehicle and DVTOD, demonstrate the superiority of our method. RGFNet achieves state-of-the-art performance with 81.4% mAP on DroneVehicle and 88.5% mAP on DVTOD. The code is available at https://github.com/uavdet/RGFNet. Zhicheng Zhao 0002, Wei Zhang 0393, Yun Xiao 0003, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Dense Tiny Object Detection: A Scene Context Guided Approach and a Unified BenchmarkabstractWith the continuous advancement of remote sensing observation technology, wide-area observation and high-resolution imaging make remote sensing images contain a large number of dense tiny objects. The detection of dense tiny objects is a very challenging task since these objects are with very low resolution and might stick together. Existing work lacks further exploration of the contextual scene information and inherent characteristics of dense tiny objects, which are crucial for performance improvement of dense tiny object detection. In this work, we propose a novel Scene Contextualized Detection Network (SCDNet) by decoupling scene contextual information through a dedicated scene classification sub-network, thereby enabling an enhanced exploration of the relationship between tiny objects and their surrounding environments. In particular, we design a lightweight scene context guided fusion module in SCDNet to incorporate scene context information around dense tiny objects more effectively. Moreover, we further develop the scene context guided foreground enhancement module to suppress the background information while enhancing the foreground information based on the scene information. In addition, this research field still lacks a large-scale benchmark dataset with dense tiny objects, which is crucial for the training and comprehensive evaluation of detection methods. To this end, we construct a large-scale dataset for dense tiny object detection. It contains 11,600 images with 1,019,800 instances, the average absolute size of objects is smaller than 13 pixels, and each image contains 88 objects on average. Extensive experiments are conducted on the proposed dataset, and the results demonstrate the superiority and effectiveness of SCDNet compared to existing methods. The dataset and evaluation code are available at https://github.com/mmic-lcl. Zhicheng Zhao 0002, Chenglong Li 0002, Yun Xiao 0003, Jin Tang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Modality Conversion Meets Superresolution: A Collaborative Framework for High- Resolution Thermal UAV Image GenerationabstractDue to the limitations and costs of thermal sensors, unmanned aerial vehicle (UAV) platforms often equip with high-resolution (HR) visible imaging and low-resolution (LR) thermal imaging cameras for all-day monitoring capability. Existing works generate the high-resolution thermal UAV images by either super-resolution (SR) from high-resolution visible and low-resolution thermal images or modality conversion (MC) from high-resolution visible images. However, the modality gap between visible and thermal sources might degrade the generation quality. We observe that the MC task is beneficial in addressing the cross-modal gap in the SR task, while the SR task can provide the condition of thermal information to boost the MC task. Moreover, these two tasks have the same output and can thus be carried out simultaneously without any additional annotation. Based on this observation, we propose a collaborative enhancement network (CENet), which performs thermal UAV image SR and visible image MC in a joint manner, for high-resolution thermal UAV image generation. In particular, we design a mutual guidance module to interact the features from SR and MC tasks in an alternating bidirectional manner. Considering that low-level vision tasks are position-sensitive, to further enhance the feature alignment between the two tasks, we design a bidirectional alignment fusion module to maintain feature consistency of the MC and SR branches. The proposed collaborative framework not only achieves joint and unified training of the two tasks, but also generates two types of complementary high-resolution images. Extensive experiments on public datasets demonstrate that the proposed CENet outperforms current state-of-the-art super-resolution (SR) methods in generating high-resolution thermal UAV images, as quantified by peak signal-to-noise ratio (PSNR) and structural similarity index measure (SSIM). Zhicheng Zhao 0002, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Category-Oriented Localization Distillation for SAR Object Detection and a Unified BenchmarkabstractDespite much research progress in synthetic aperture radar (SAR) object detection, the performance of SAR object detection has encountered a bottleneck limited by the imaging mechanism of SAR. In this work, we investigate how to perform robust SAR object detection by distilling the category knowledge from optical images in the training stage. To this end, we propose a novel knowledge distillation method called Category-oriented Localization Distillation (CoLD), which employs the optical object detection network as the teacher to guide the SAR object detection network. To introduce the category prior knowledge of the teacher network in the localization knowledge transferring, a category-oriented partition module is designed in CoLD to decouple candidate bounding boxes into target and non-target ones according to the category information in optical images. Through box decoupling, the accuracy and efficiency of SAR object detection can be significantly improved. Moreover, an IoU-based weighting module is introduced in CoLD to guide the student network focusing more on high-quality candidate boxes by adaptively changing the weight of each candidate bounding box based on the corresponding IoU score in the teacher network. In addition, a unified benchmark dataset is created for the evaluation of optical information guided SAR object detection, which consists of 14,665 optical and SAR image pairs in the training set and 3,666 SAR images in the testing set. Extensive experiments on the dataset demonstrate the effectiveness of our CoLD against state-of-the-art methods. The dataset is available at: https://github.com/mmic-lcl/Datasets-and-benchmark-code. Rui Ruan, Zhicheng Zhao 0002, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Thermal UAV Image Super-Resolution Guided by Multiple Visible CuesabstractUnmanned aerial vehicle (UAV) thermal-imaging has received much attention, but the insufficient image resolution caused by thermal imaging systems is still a crucial problem that limits the understanding of thermal UAV images. However, high-resolution visible images are relatively easy to access, and it is thus valuable for exploring useful information from visible image to assist thermal UAV image super-resolution (SR). In this article, we propose a novel multiconditioned guidance network (MGNet) to effectively mine the information of visible images for thermal UAV image SR. High-resolution visible UAV images usually contain salient appearance, semantic, and edge information, which plays a critical role in boosting the performance of thermal UAV image SR. Therefore, we design an effective multicue guidance module (MGM) to leverage the appearance, edge, and semantic cues from visible images to guide thermal UAV image SR. In addition, we build the first benchmark dataset for the task of thermal UAV image SR guided by visible images. It is collected by a multimodal UAV platform and composes of 1025 pairs of manually aligned visible and thermal images. Extensive experiments on the built dataset show that our MGNet can effectively leverage useful information from visible images to improve the performance of thermal UAV image SR and perform well against several state-of-the-art methods. The dataset is available at:https://github.com/mmic-lcl/Datasets-and-benchmark-code. Zhicheng Zhao 0002, Chenglong Li 0002, Yun Xiao 0003, Jin Tang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Remote Sensing Image Scene Classification Based on an Enhanced Attention ModuleabstractClassifying different satellite remote sensing scenes is a very important subtask in the field of remote sensing image interpretation. With the recent development of convolutional neural networks (CNNs), remote sensing scene classification methods have continued to improve. However, the use of recognition methods based on CNNs is challenging because the background of remote sensing image scenes is complex and many small objects often appear in these scenes. In this letter, to improve the feature extraction and generalization abilities of deep neural networks so that they can learn more discriminative features, an enhanced attention module (EAM) was designed. Our proposed method achieved very competitive performance—94.29% accuracy on NWPU-RESISC45 and state-of-the-art performance on different remote sensing scene recognition data sets. The experimental results show that the proposed method can learn more discriminative features than state-of-the-art methods, and it can effectively improve the accuracy of scene classification for remote sensing images. Our code is available athttps://github.com/williamzhao95/Pay-More-Attention. Zhicheng Zhao 0002, Ze Luo |
IEEE Geosci. Remote. Sens. Lett. | 1 |