EDBT 2026 Demo / reviewers in the wild / expert
Zitong Wu
dblp:283/6804
· DBLP profile ↗
16ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0002-0449-9465ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 2 first-author · 14 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Direct Data-Driven Event-Triggered Consensus Under Disturbances and Digraph
Guanglei Zhao, Zitong Wu, Changchun Hua, Weili Ding |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Interpretable Fine-Grained Aircraft Classification Network for Remote Sensing Image With Image Pair Interaction and Neural TreeabstractThis paper proposes a novel interpretable framework for fine-grained aircraft classification in high-stakes remote sensing applications. Our approach addresses three key challenges: small inter-class variance, large intra-class variance, and the need for model interpretability. Specifically, our framework is built on the Swin Transformer (SwinT) backbone and includes three main modules. First, we present the Dynamic Attention Fusion Module (DAFM), which adaptively fuses multi-stage attention maps from the SwinT backbone. By leveraging a dispersion-based weighting mechanism, DAFM balances the contributions of coarse and fine-grained features, capturing both global structures and localized details. Second, we propose the Adaptive Image Pair Interaction Module (AIPI), which dynamically adjusts feature interaction strategies based on intra-class and inter-class similarity, effectively enhancing informative regions and improving robustness. To further optimize discriminative power, we incorporate an AIPI loss function that enforces intra-class consistency and inter-class separability. Finally, we develop a Binary Neural Tree Module (BNTM) to hierarchically select and propagate informative image patches, enhancing both feature refinement and interpretability through explicit path-based decision-making. Extensive experiments on benchmark datasets demonstrate that our framework significantly improves classification accuracy and interpretability, making it well-suited for applications requiring transparent and reliable decision-making. The Code can be found at https://github.com/StarmanGzx/BNTM. Zhengxi Guo, Biao Hou, Xianpeng Guo, Chen Yang 0027, Zitong Wu, Bo Ren 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Masked Angle-Aware Autoencoder for Remote Sensing Images
Zhihao Li 0005, Biao Hou, Siteng Ma, Zitong Wu, Xianpeng Guo, Bo Ren 0001, Licheng Jiao |
ECCV (8) | 4 |
| 2024 | Weakly supervised object localization via knowledge distillation based on foreground-background contrast
Siteng Ma, Biao Hou, Zhihao Li 0005, Zitong Wu, Xianpeng Guo, Chen Yang 0027, Licheng Jiao |
Neurocomputing | 4 |
| 2024 | Unreliable Pixel Contrast Based on von Mises-Fisher Distribution for Semi-Supervised SAR SegmentationabstractThe unique visual properties and the huge size of synthetic aperture radar (SAR) images pose challenges in labeling data. The scarcity of labeled data limits the training of SAR image segmentation networks. To address this issue, semi-supervised methods are used to train the network. However, traditional semi-supervised algorithms like Mean Teacher often generate erroneous pseudo-labels that carry over to subsequent training epochs, leading to overfitting and affecting network performance. This overfitting stems from an overreliance on unreliable pixels in Mean Teacher. In this study, an enhanced approach is proposed called unreliable pixel contrast (UPCo), where a von Mises-Fisher distribution is applied to constrain unreliable pixels in feature space. We augment the segmentation network with a feature output header for pixel-level contrastive learning in UPCo. Moreover, to minimize the computational effort during the training phase, hard sample selection and negative object non-uniform selection strategies are designed to facilitate contrastive learning. The proposed UPCo was evaluated on two large-scene SAR images and demonstrated its superiority over other comparative algorithms, achieving more optimal performance in semi-supervised segmentation of SAR images. Zitong Wu, Biao Hou, Xianpeng Guo, Bo Ren 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2024 | MSRIP-Net: Addressing Interpretability and Accuracy Challenges in Aircraft Fine-Grained Recognition of Remote Sensing ImagesabstractThe task of fine-grained aircraft recognition is crucial in the field of remote sensing. Despite some progress achieved by traditional deep learning methods in addressing this challenge, they are often perceived as a “black box,” lacking transparent explanations for model decisions. Current interpretable methods based on attention mechanisms, although providing some interpretability, do not align with human thought logic. Therefore, we propose a multiscale rotation-invariant prototype network (MSRIP-Net). Our approach simulates the intuitive reasoning process of humans in identifying objects by segmenting them into multiple components. Importantly, MSRIP-Net has the capability to automatically recognize rigid components on aircraft targets without relying on additional part annotations, using only image-level class labels. In addition, our approach effectively addresses challenges presented by noise, deformations, and multiscale variations in remote sensing targets and has been comprehensively evaluated on datasets FAIR1M1.0 and Rareplane. Our results demonstrate that MSRIP-Net achieves higher accuracy compared with existing fine-grained recognition methods. Furthermore, we provide insights into the model’s decision-making process to illustrate the interpretability of our approach. Zhengxi Guo, Biao Hou, Xianpeng Guo, Zitong Wu, Chen Yang 0027, Bo Ren 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | MGC: MLP-Guided CNN Pretraining Using a Small-Scale Dataset for Remote Sensing ImagesabstractTo overcome the inherent domain gap between natural images and remote sensing images (RSIs), it is highly desirable to develop pretraining methods specifically for RSIs. Considering the lack of widely recognized large-scale benchmarks like ImageNet in the RSI community and limited computational resources, this article proposes multilayer perceptron (MLP)-guided convolutional neural network (CNN) (MGC), a method that employs an MLP to guide the pretraining of a CNN from small-scale datasets for RSIs. MGC has two encoders, each consisting of a CNN branch and an MLP branch. We first contrast pairwise samples from the same type of branches or different types of branches across the encoders and employ a positive-pair guidance strategy to explore consistency. Due to the inherent locality issue of shallow layers in a CNN, the CNN branches often do not attend to correct foreground regions such as objects, regions of interest, and land coverage. Therefore, we further propose an attention guidance strategy to guide the CNN branches to focus on foreground regions and learn discriminative representations effectively. The proposed MGC method is validated by pretraining a CNN model using the MGC and applying it to different downstream tasks including scene classification, rotated object detection, semantic segmentation, and change detection on ten datasets. Results have confirmed the effectiveness of the proposed MGC. Our code will be released at:https://github.com/benesakitam/MGC. Zhihao Li 0005, Biao Hou, Wanqing Li 0001, Zitong Wu, Bo Ren 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Rebalancing Gaussian Location Loss for High-Precision Detection on Remote Sensing ImagesabstractAerial image objects are usually orientated arbitrarily, with a large scale range, and densely distributed. Traditional horizontal bounding box (HBB) detectors tend to filter out densely distributed objects leading to missed detections, such as ship (SH) and vehicle. Therefore, oriented object detection has become a mainstream solution in recent years. The 2-D Gaussian distribution representation of the oriented bounding boxes (OBBs) solves the problem of angular discontinuity and boundary discontinuity and thus gets more attention. However, as the aspect ratio of the object gradually decreases, its predicted angular performance continues to decrease. We find that the angular gradient of an object decreases sharply as the aspect ratio decreases, resulting in a large gradient gap between a small aspect ratio object (SARO) and a large aspect ratio object (LARO). It makes the detector prefer to ignore SARO during training, which weakens the high precision performance of SARO. We call this phenomenon shape imbalance. To solve the problem, we proposed a simple gradient rebalancing strategy named shape balance. Since the shape imbalance is only related to the aspect ratio of the object, we designed a modulation function with an inverse aspect ratio to calculate the balance coefficient. The principle of the function is that the larger the aspect ratio, the smaller the balance coefficient; the smaller the aspect ratio, the larger the balance coefficient. We aim to get the balance coefficients for objects with different aspect ratios. Location loss multiplied by a balance coefficient can directly adjust the gradient gap between objects with different aspect ratios to achieve a rebalancing effect. Extensive experiments conducted on DOTA-v1.0 dataset and DIOR-R dataset verify the effectiveness of our proposed method. Our method improves the detection performance of Gaussian location loss by an average of 2.08%/1.01%(AP75/mAP) metrics on the DOTA-v1.0 dataset and 1.17%/0.82%(AP75/mAP) improvements for DIOR-R dataset. Biao Hou, Zitong Wu, Xianpeng Guo, Bo Ren 0001, Zhongle Ren, Chen Yang 0027, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Complete Rotated Localization Loss Based on Super-Gaussian Distribution for Remote Sensing ImagesabstractLocalization regression in oriented object detection tasks has long faced boundary discontinuity and angular discontinuity problems induced by periodic angles. These problems were successfully resolved by using a 2d Gaussian distribution to modelling the oriented bounding box (OBB). However, the angular information of square-like objects will be lost when they are converted to 2d Gaussian distribution, forming a systematic problem. Its fundamental reason is that when the aspect ratio of the object tends to 1, the equiprobability curve of 2d Gaussian distribution degenerates from an ellipse to a circle, thus losing the orientation information of the rotated object. This results in the bounding boxes of such square-like objects not being learned effectively. To resolve this problem, we used the Lamé curve (or superellipse) to modify the existing 2d Gaussian function and designed a super-Gaussian distribution. This distribution can maintain anisotropy at arbitrary aspect ratios, thus preserving the angular information of the oriented object. We used the Kullback-Leibler (KL) divergence to measure the distance between two super-Gaussian distributions and convert it into a localization loss (SGKLD) by a function. SGKLD is an improved version of KLD loss. By modifying the form of the probability distribution, we elegantly fix the angle missing problem of the traditional Gaussian distribution. We validated the effectiveness of the proposed algorithm on several datasets and obtained the performance of SOTA. Our algorithm achieves a mean average precision (mAP) of 80.07, 76.59, 62.27, and 90.55/98.13 on the DOTA-v1.0, DOTA-v1.5, DOTA-v2.0, and HRSC2016 datasets, respectively. Biao Hou, Zitong Wu, Zhengxi Guo, Bo Ren 0001, Xianpeng Guo, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Gaussian Synthesis for High-Precision Location in Oriented Object DetectionabstractIn aerial image scenes, the objects have properties of arbitrary orientation, large-scale range, and dense distribution. Thus, the object detector uses oriented bounding box (OBB) to locate objects, which is more complex and challenging than horizontal bounding box (HBB) detector. Mainstream OBB detectors mostly use one-to-many label assignment strategy to predict multiple bounding boxes for the same object, and filter out repeat predictions by non-maximum suppression (NMS). NMS ranks with confidence and drops the detection box with IoU higher than the threshold, which is easy to get the local optimum result. The clustered synthesis method gets more accurate results than the original NMS, but applying it to the OBB detector leads to border shift, which arises from the angular discontinuity problem. Therefore, we use Gaussian OBB (G-OBB) to deal with the angular discontinuity and thus eliminate the offset generated by direct synthesis. G-OBB is not an easy to understand and describe representation. For this reason, we analyze the properties of G-OBB, and design a decoding method to convert a G-OBB to a rotated rectangular box, further discussing its conditions. Based on the decoding method, we propose a Gaussian synthesis algorithm (GauS), which transforms the OBB into Gaussian space, followed by synthesis, and finally transforms the synthesis result back into a new OBB. We have derived the synthesis and decoding methods, and further verified their effectiveness. The extensive experiments on several existing models show that GauS takes very little computation and improves detector’s high-precision performance. Extensive experiments verify the effectiveness, stability, and universality of the proposed algorithm. In addition, The RTMDet using GauS achieves a performance of 81.61 AP50and gains a 0.39% improvement in mAP, which achieves the SOTA performance. Our implementation is available at: https://github.com/lzh420202/GauS. Biao Hou, Zitong Wu, Bo Ren 0001, Zhongle Ren, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Unsupervised Prototype-Wise Contrastive Learning for Domain Adaptive Semantic Segmentation in Remote Sensing ImageabstractLabeling data in the field of remote sensing is time-consuming and labor-intensive, making domain adaptation between different domains an urgently needed solution. To address the domain gap between diverse datasets in the remote sensing domain, numerous methods tailored for domain adaptation in high-resolution remote sensing imagery have emerged. Some of the existing methods focus on reducing the domain gap at either the feature level or the pixel level, often overlooking their underlying connection. To tackle this issue, we introduce a prototype-wise contrastive feature alignment paradigm (PCFA) aimed at bridging the representations between the feature and pixel levels. By dynamically updating, we acquire prototype information encompassed by different mini-batches and employ an optimal transport mechanism to reasonably apply the prototype feature distribution in guiding the learning of target domain features. We conduct extensive domain adaptation semantic segmentation (DASS) experiments on the ISPRS Vaihingen and Potsdam datasets, achieving an improvement about 4%~5% in mIoU (mean Intersection over Union) compared to previous methods using the DeepLabV2 framework. Siteng Ma, Biao Hou, Xianpeng Guo, Zitong Wu, Zhihao Li 0005, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Automatic Aug-Aware Contrastive Proposal Encoding for Few-Shot Object Detection of Remote Sensing ImagesabstractIn the annotation of remote sensing images (RSIs), the effectiveness of common object detection methods trained on only a few samples decreases instantly, which has prompted increasing research on the few-shot problem in remote sensing. RSIs often exhibit suboptimal performance in few-shot scenarios due to the intricate nature of scene information interference and the high degree of cosine similarity, both of which present significant challenges to their effectiveness. In this paper, a two-stage detection framework based on fine-tuning is selected to deal with the common problems in the few-shot task of remote sensing domain. Considering the excessive scale variation of instances in remote sensing datasets, we introduce an automatically learned aug-aware search module to provide an intelligent data augmentation solution for Faster R-CNN using different optimal augmentation policies searched by the network to fit the current dataset. We introduce a contrastive RoI branch to better classify novel class proposal features that are easily confused by the base class. We named our work AACE and conducted extensive experiments on two common object detection datasets in remote sensing, NWPU VHR 10 and DIOR, on which AACE achieved about 2.30% and 2.61% improvement, respectively, in the number of shots listed in the paper, compared to other algorithms. Siteng Ma, Biao Hou, Zitong Wu, Zhihao Li 0005, Xianpeng Guo, Bo Ren 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Incremental Land Cover Classification via Label Strategy and Adaptive WeightsabstractDuring incremental learning tasks, catastrophic forgetting occurs when old models are updated with new information. To address this issue, we propose a novel method called label strategy and adaptive weights (LSAW) that improves the incremental learning process. The label strategy introduces the old classes and solves the problem of how to reasonably use the wrong samples predicted by the old model. In the cross-entropy (CE) loss, we apply a threshold to filter the pseudolabels predicted by the old model. Subsequently, we merge the pixel samples with high probability with the current label. The probability here refers to the probability that the pixel belongs to the true class. This process enables the introduction of information from old classes that are not directly accessible in the current stage. Moreover, this information is relatively reliable, and the model exhibits confidence in its accuracy. For the remaining pixels, we retain all classes’ information through label smoothing. In the distillation function, the old class and background pixel samples are selected for distillation according to the prediction map of the old classes. The weights of the classes are adaptively updated and adjusted using specific label information from each batch and the different stages of incremental learning. As demonstrated by the results of our experiment, on three remote sensing image datasets: China Computer Federation (CCF), Potsdam, and Vaihingen, our method achieves the best results. Bo Ren 0001, Zhao Wang 0011, Biao Hou, Bo Liu 0009, Zitong Wu, Jocelyn Chanussot, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | A Neural Network Based on Consistency Learning and Adversarial Learning for Semisupervised Synthetic Aperture Radar Ship DetectionabstractShip detection in synthetic aperture radar (SAR) images has important application value. Sea clutter, complex scenes, a large size change in ships, and the arbitrary directionality of ships make ship detection challenging. With the development of deep learning, many deep learning algorithms have been applied to SAR images. These algorithms need a lot of labeled data for training. It is time-consuming to label SAR data, and the unlabeled data are easy to obtain. It is necessary to use the unlabeled data effectively to improve the performance of the algorithm. In this study, a semisupervised SAR ship detection network, named the semisupervised consistency learning adversarial network (SCLANet), is presented. SCLANet is a two-stage detection network. The local features around the ship can be extracted by the SCLANet, and the features generated from unlabeled data become closer to those generated from labeled data by using adversarial learning. There are two consistency learning modules in SCLANet: noise robustness consistency learning and output encoding consistency learning. Noise robustness consistency learning can increase the robustness of the SCLANet. Maintaining consistency between the noisy results and the original results can train the unlabeled data. In output encoding consistency learning, outputs are mapped to a picture that is fed into an encoder to obtain the intermediate representation embedding. Another embedding is a layer in the main network of the SCLANet. Reducing the error between two embeddings can train the SCLANet with unlabeled data. Two types of consistency learning can be used as pretext tasks for semisupervised learning. Experiments were conducted on two SAR ship datasets. Compared with other algorithms, the SCLANet achieved the highest detection accuracy, indicating that it is more advantageous to use in ship detection. Biao Hou, Zitong Wu, Bo Ren 0001, Xianpeng Guo, Shuang Wang 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Multicrop Fusion Strategy Based on Prototype Assignment for Remote Sensing Image Scene ClassificationabstractThe gap between self-supervised visual representation learning and supervised learning is gradually closing. Self-supervised learning does not rely on a large amount of labeled data and reduces the loss of human labeled information. Compared with natural images, remote sensing images require rich samples and human annotation by experts. Moreover, many algorithms have poor interpretability and unconvincing results. Therefore, this paper proposes a self-supervised method based on prototype assignment by designing a pretext task so that the network maps features to prototypes in the process of learning, swaps the code corresponding to the obtained features, combines them with another data-enhancing feature, and then optimizes the network. The prototype is introduced to explain the clustering idea embodied in the whole process. Considering the existence of the scene information-rich characteristic of remote sensing images, we introduce multiple views with different resolutions to capture more detailed information on the images. Finally, if the data enhancement method is not powerful enough, the network can easily fall into an overfitting state, which prevents the network from learning subtle differences and detailed information. To address this shortcoming, we propose a fusion strategy to flatten the decision boundary of the framework so that the model can also learn the soft similarity between sample pairs. We name the whole framework MFPC. In extensive experiments conducted on three common remote sensing image datasets (i.e., UCMerced, AID, and NWPU45), MFPC achieves a maximum improvement of 4.3% over some existing self-supervised algorithms, indicating that it can achieve good results. Siteng Ma, Biao Hou, Xianpeng Guo, Zhihao Li 0005, Zitong Wu, Shuang Wang 0001, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | Multiscale CNN With Autoencoder Regularization Joint Contextual Attention Network for SAR Image ClassificationabstractSynthetic aperture radar (SAR) image classification is a fundamental research direction in image interpretation. With the development of various intelligent technologies, deep learning techniques are gradually being applied to SAR image classification. In this study, a new SAR classification algorithm known as the multiscale convolutional neural network with an autoencoder regularization joint contextual attention network (MCAR-CAN) is proposed. The MCAR-CAN has two branches: the autoencoder regularization branch and the context attention branch. First, autoencoder regularization is used for the reconstruction of the input to regularize the classification in the autoencoder regularization branch. Multiscale input and an asymmetric structure of the autoencoder branch cause the network more to be focused on classification than on reconstruction. Second, the attention mechanism is used to produce an attention map in which each attention weight corresponds to a context correlation in attention branch. The robust features are obtained by the attention mechanism. Finally, the features obtained by the two branches are spliced for classification. In addition, a new training strategy and a postprocessing method are designed to further improve the classification accuracy. Experiments performed on the data from three SAR images demonstrated the effectiveness and robustness of the proposed algorithm. Zitong Wu, Biao Hou, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 1 |