EDBT 2026 Demo / reviewers in the wild / expert
Zicong Zhu
dblp:279/0333
· DBLP profile ↗
12ranked-venue papers
6as first author
12since 2021 · last 2026
0000-0003-3897-919XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Two-Stage Method With Lightweight Network and Active Contour Model for Remote Sensing Image SegmentationabstractActive contour models (ACMs) have shown effectiveness on remote sensing (RS) image segmentation tasks. However, this type of method still faces two important problems when segmenting RS images. First, not using high-level semantic information makes it difficult for ACM to distinguish targets and backgrounds with similar textures. Second, the manual contour initialization required by ACMs is inconvenient and inefficient. To address the above problems, we propose a two-stage segmentation method that consists of an improved U-Net using a mixed pooling attention and an ACM (UMPA-ACM). In the first stage, a lightweight network based on U-Net structure is developed to extract semantic features while reducing computational cost. A mixed pooling attention (MPA) module is designed to enhance the ability of our proposed network to extract high-level semantic information from RS images. In the second stage, an adaptive feature enhancement (AFE) module computes grayscale information from original images and feature maps produced by the lightweight network and then fuses them to improve the intensity of target edges; a morphological-threshold process (MTP) module automatically generates appropriate initial contours for targets from the semantic feature maps instead of manual contour initialization. Then, a new ACM, proposed based on pre-fitting foreground and background in local regions, uses the statistical characteristics of local intensities and the bias field correction to suppress the interference of non-target regions, thereby improving segmentation accuracy. Experimental results show that the mean Dice Similarity Coefficient (mDSC) and the mean Intersection over Union (mIoU) of our method are higher than those of the suboptimal method by 1.21% and 1.60%, respectively, on average for segmenting images from six RS datasets, which verify the advantage of our proposed method. Bin Dong 0005, Zicong Zhu, Qianqian Bu, Mengya Wu, Jingen Ni |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | An active contour model with adaptive weighted mean filtering and anisotropic diffusion filtering
Bin Dong 0005, Qianqian Bu, Zicong Zhu, Jingen Ni |
Signal Process. | 3 |
| 2025 | Active contour model with improved second-order differential driven term
Bin Dong 0005, Zicong Zhu, Qianqian Bu, Jingen Ni |
Signal Process. | 2 |
| 2025 | Align and Complete Samples in Remote Sensing Fine-Grained Rigid Object DetectionabstractCurrently, the remote sensing fine-grained rigid object detectors mainly face two challenges: fuzzy localization and inaccurate classification (including misclassifying and multi-classifying). Firstly, mainstream detectors based on the “dense prediction” paradigm suffer from a misalignment between anchor points (APs) and ground truths (GTs). Commonly, they adopt multi-stage regression as a solution, which brings an ambiguous sample definition problem and redundant computational costs. Secondly, two factors seriously affect the classification. They are the insufficient sample learning caused by the long-tail distribution of categories, and the difficulty in extracting discriminative features caused by slight inter-class variance. To address the issues above, we propose an efficient aligning and completing detector (ACDet) based on a single-stage structure. Firstly, the Adaptive Anchor Alignment Mechanism decouples the centripetal sampling bias in the categorical features and leverages it to learn the APs aligned with GTs. It is plug-and-play, requiring no additional supervision annotations. Secondly, a novel Online Tail-sample Supplementation algorithm is proposed. It dynamically maintains class balance during training and can be easily added as post-processing behind existing sample assignment strategies. Thirdly, an Adaptive Group Perceptron is designed to effectively enrich the diversity of features and enhance the model’s ability to extract discriminative features. Experiments on four public datasets demonstrate that the proposed ACDet achieves the state-of-the-art (SOTA) level, even surpassing competitive multi-stage detectors, with fewer computing resources and a faster inference speed. The code will be public after the paper is published. Zicong Zhu, Jian Kang 0005, Wenhui Diao, Bing Wang 0015, Jingen Ni |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Active Contour Model Driven by Non-Local Feature Fitting Energy Function With Scalable NormalizationabstractIt is challenging for active contour models (ACMs) to segment weak-edge and noisy images efficiently and accurately. To solve this problem, a novel ACM is proposed in this work. The proposed ACM achieves high-precision segmentation for weak-edge and noisy images using a non-local feature fitting energy function and a scalable normalization method. The non-local feature fitting energy function is constructed based on the distances calculated by Jeffreys divergence between non-local weighted fitting images and the image processed by the non-local means (NLM) algorithm. The non-local weighted fitting images include the fitting foreground and background with image edge features. The images processed by the NLM algorithm is used to reduce the influence of noise. The data-driven term, obtained by minimizing the non-local feature fitting energy function, is computed before the level set iteration, which improves the computation speed. In addition, a scalable normalization method is proposed to normalize the data-driven term. The ability to distinguish the targets from the background for different types of images is enhanced by adjusting a scaling factor, improving the robustness and accuracy of the proposed model. Experimental results demonstrate the advantages of the proposed model. Qianqian Bu, Bin Dong 0005, Zicong Zhu, Jingen Ni |
IEEE Trans. Image Process. | 3 |
| 2024 | A Parameter-Efficient Differentiable Active Contour Network for Precisely Building Instance SegmentationabstractMethods combining deep neural networks (DNN) and active contour models (ACM) have shown effective performance in building segmentation tasks. However, due to the independent structural design, it suffers from a large amount of parameters and computational cost. To address the above problem and further explore the performance effect of the coupling DNN and ACM methods, we present a Parameter-efficient Differentiable Active Contour (PDAC) network for precisely segmenting buildings in remote sensing scenarios. Specifically, an adapter is applied to fine-tune the encoder, instead of initializing a new one. Based on it, a semantic segmentation model can achieve further localization improvement by introducing a minor amount of parameters. Experiments on two BE datasets show that PDAC achieves better performance with nearly half of the computational resources than the baseline. Zicong Zhu, Bin Dong 0005, Qianqian Bu, Jingen Ni |
IGARSS | 1 |
| 2024 | An active contour model based on shadow image and reflection edge for image segmentation
Bin Dong 0005, Guirong Weng, Qianqian Bu, Zicong Zhu, Jingen Ni |
Expert Syst. Appl. | 4 |
| 2024 | SIRS: Multitask Joint Learning for Remote Sensing Foreground-Entity Image-Text RetrievalabstractThe essence of improving the effect of cross-modal image-text retrieval (CIR) lies in the finer-grained modeling of homogeneous features between modalities. However, in remote sensing (RS) scenarios, existing methods usually apply the image-sentence granular feature alignment paradigm, bringing significant difficulties to the fine-grained representation of homogeneous features between modalities. Besides, more complex background noise and extreme scale ranges of foreground targets are hard to distinguish, causing the feature mottle problem. To address the above issues, we propose a novel Semantic-guided Image-text Retrieval framework with Segmentation (SIRS). It is a multi-task joint learning framework for plug-and-play and end-to-end training RS CIR models efficiently, including Semantic-guided Spatial Attention (SSA) and Adaptive Multi-scale Weighting (AMW) modules. First, SSA introduces a background reconstruction branch based on noise perception and a semantic segmentation branch based on pixel-level prediction. It explores a joint learning strategy that concisely filters background noise and refines foreground features considerably. Secondly, AMW performs multi-scale weighting on various layers of feature map output by the encoder, effectively improving the learning efficiency of foreground targets at different scales. It is worth mentioning that SIRS outputs combination results with image and segmentation mask, which is not available in other methods. Based on the RSITMD dataset, we complete the semantic segmentation annotation RSITMD-SS to verify the performance of the proposed method. Sufficient and complete experiments verify the effectiveness of the proposed method. With SIRS, the mainstream SVP and CLIP-based methods improve about 7 mR and derive segmentation prediction with acceptable computational cost optionally. The code and associated dataset will be available on https://github.com/StarBurstStream0/SIRS. Zicong Zhu, Jian Kang 0005, Wenhui Diao, Yingchao Feng, Junxi Li, Jingen Ni |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | RingMo: A Remote Sensing Foundation Model With Masked Image ModelingabstractDeep learning approaches have contributed to the rapid development of remote sensing (RS) image interpretation. The most widely used training paradigm is to use ImageNet pretrained models to process RS data for specified tasks. However, there are issues such as domain gap between natural and RS scenes and the poor generalization capacity of RS models. It makes sense to develop a foundation model with general RS feature representation. Since a large amount of unlabeled data is available, the self-supervised method has more development significance than the fully supervised method in RS. However, most of the current self-supervised methods use contrastive learning, whose performance is sensitive to data augmentation, additional information, and selection of positive and negative pairs. In this article, we leverage the benefits of generative self-supervised learning (SSL) for RS images and propose an RS foundationmodel framework called RingMo, which consists of two parts. First, a large-scale dataset is constructed by collecting two million RS images from satellite and aerial platforms, covering multiple scenes and objects around the world. Second, we propose an RS foundation model training method designed for dense and small objects in complicated RS scenes. We show that the foundation model trained on our dataset with RingMo method achieves state-of-the-art (SOTA) on eight datasets across four downstream tasks, demonstrating the effectiveness of the proposed framework. Through in-depth exploration, we believe it is time for RS researchers to embrace generative SSL and leverage its general representation capabilities to speed up the development of RS applications. Xian Sun 0001, Peijin Wang, Wanxuan Lu, Zicong Zhu, Qibin He 0001, Junxi Li, Xuee Rong, Zhujun Yang, Qinglin He, Ruiping Wang 0001, Jiwen Lu, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | CODet: Component Object Detector Extracting Structural Features Based on Target CharacteristicsabstractDeep learning technology has promoted the object detection task in the remote sensing (RS) field to move toward better performance and more demanding requirements. Except for rigid body objects, component objects (COs) with more complex characteristics remain a detection challenge. Its “partial rules and overall disorder” characteristic limits the model learning ability to the structural features. And the internal noise and relatively sparse arrangement are not conducive to optimizing the model by the existing sample assignment strategies. We propose CODet to detect COs in RS scenes. It consists of a cross-hierarchy feature fusion module (CFM) and a noise-sparse sample assignment (NSA) strategy. CFM learns the potential representation and relative position relationship of components by fusing different level features. NSA redefines the optimization process of sample assignment. It aims to alleviate the problems of classification–localization misalignment (CLM) and the positive–negative sample imbalance (PNI) caused by the object’s internal noise and sparse arrangement. The method is verified on the proposed COD dataset of six categories of COs, reaching an average mAP/mAP50of 54.3/86.0. To be closer to the task requirements of the practical RS scene, we also propose a RS large-scale images inference framework. It includes a dataset (APRoI, labeled with COs and rigid body objects), a large-scale image inference strategy, and a set of evaluation metrics. With CODet as the core, the framework can effectively reduce the inference time by three to four times on images with an average of more than 100 million pixels. Zicong Zhu, Xian Sun 0001, Wenhui Diao, Kaiqiang Chen, Qibin He 0001, Guangluan Xu, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | AOPDet: Automatic Organized Points Detector for Precisely Localizing Objects in Aerial ImageryabstractWith the development of deep convolutional neural networks, detecting rotating objects in remote-sensing images is of great significance in various fields. Existing rotating object detectors most suffer the problem of ambiguous supervision caused by inappropriate rotating object representations. This problem may result in fuzzy object localization and further lead to misclassification. In this article, we propose an Automatic Organized Points Detector (AOPDet), which derives precise localization results by applying a novel rotating object representation called nonsequential corners representation. To achieve the proposed representation, an Automatic Organization Mechanism (AOM) technique is designed to guide the model to organize points to object corners automatically. An Automatic-Organized-Points-specific (AOP-specific) head structure is also designed and equipped in the model to better focus on the rotating object detection task. On public aerial datasets, experiments show that the AOPDet achieves 17.0 mAP higher than the compared baseline model, reaching the state-of-the-art (SOTA) level. Detailed ablation experiments and error analysis strongly reveal the effectiveness of the proposed model. Zicong Zhu, Xian Sun 0001, Wenhui Diao, Kaiqiang Chen, Guangluan Xu, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Invariant Structure Representation for Remote Sensing Object Detection Based on Graph ModelingabstractDue to the characteristics of vertical orthophoto imaging, the apparent structural features of the object in the remote sensing image are relatively stable, such as the cross-shaped structure of the aircraft, the rectangular structure of the vehicle, etc. Compared with the traditional visual features, using these features is conducive to improving the accuracy of object detection. However, there are few studies on such characteristics. In this paper, we systematically study the invariant structural features of remote sensing objects and propose a Graph Focusing Aggregation Network (GFA-Net) to represent the structural features of remote sensing objects. Among them, in view of the problem that traditional convolutional neural networks (CNNs) are sensitive to the changes in rotation, scale, and other factors, which makes it difficult to extract structural features, we propose the Graph Focusing Process (GFP) based on the idea of graph convolution. Analysis and experiments show that graph structure has significant advantages over Euclidean feature space under CNN in expressing such structural features. In order to realize the end-to-end efficient training of the above model, we design Graph Aggregation Network (GAN) to update the weight of nodes. We verify the effectiveness of our method on the proposed multi-task datasets ACSD and large-scale fine-grained remote sensing dataset FAIR1M. Experiments conducted on the object detection data sets of DOTA and HRSC2016 prove that the proposed method is superior to the current state-of-the-art method. Zicong Zhu, Xian Sun 0001, Wenhui Diao, Kaiqiang Chen, Guangluan Xu, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |