VLDB 2026 Research / reviewers in the wild / expert
Zhitong Xiong
dblp:202/2877
· DBLP profile ↗
42ranked-venue papers
10as first author
35since 2021 · last 2025
0000-0002-3953-585XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 26 · 6 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards a Unified Copernicus Foundation Model for Earth VisionabstractAdvances in Earth observation (EO) foundation models have unlocked the potential of big satellite data to learn generic representations from space, benefiting a wide range of downstream applications crucial to our planet. However, most existing efforts remain limited to fixed spectral sensors, focus solely on the Earth's surface, and overlook valuable metadata beyond imagery. In this work, we take a step towards next-generation EO foundation models with three key components: 1) Copernicus-Pretrain, a massive-scale pretraining dataset that integrates 18.7M aligned images from all major Copernicus Sentinel missions, spanning from the Earth's surface to its atmosphere; 2) Copernicus-FM, a unified foundation model capable of processing any spectral or non-spectral sensor modality using extended dynamic hypernetworks and flexible metadata encoding; and 3) Copernicus-Bench, a systematic evaluation benchmark with 15 hierarchical downstream tasks ranging from preprocessing to specialized applications for each Sentinel mission. Our dataset, model, and benchmark greatly improve the scalability, versatility, and multimodal adaptability of EO foundation models, while also creating new opportunities to connect EO, weather, and climate research. Codes, datasets and models are available at https://github.com/zhu-xlab/Copernicus-FM. Yi Wang 0072, Zhitong Xiong, Chenying Liu 0001, Adam J. Stewart, Thomas Dujardin, Nikolaos-Ioannis Bountos, Angelos Zavras, Franziska Gerken, Ioannis Papoutsis, Laura Leal-Taixé, Xiao Xiang Zhu 0001 |
ICCV | 2 |
| 2025 | REOBench: Benchmarking Robustness of Earth Observation Foundation ModelsabstractEarth observation foundation models have shown strong generalization across multiple Earth observation tasks, but their robustness under real-world perturbations remains underexplored. To bridge this gap, we introduce REOBench, the first comprehensive benchmark for evaluating the robustness of Earth observation foundation models across six tasks and twelve types of image corruptions, including both appearance-based and geometric perturbations. To ensure realistic and fine-grained evaluation, our benchmark focuses on high-resolution optical remote sensing images, which are widely used in critical applications such as urban planning and disaster response. We conduct a systematic evaluation of a broad range of models trained using masked image modeling, contrastive learning, and vision-language pre-training paradigms. Our results reveal that (1) existing Earth observation foundation models experience significant performance degradation when exposed to input corruptions. (2) The severity of degradation varies across tasks, model architectures, backbone sizes, and types of corruption, with performance drop varying from less than 1% to over 25%. (3) Vision-language models show enhanced robustness, particularly in multimodal tasks. REOBench underscores the vulnerability of current Earth observation foundation models to real-world corruptions and provides actionable insights for developing more robust and reliable models. Xiang Li 0001, Siwei Liu 0001, Zhitong Xiong, Chunbo Luo, Lu Liu 0001, Mykola Pechenizkiy, Xiao Xiang Zhu 0001, Tianjin Huang |
NeurIPS | 5 |
| 2025 | Semi-Supervised Building Footprint Extraction Using Debiased Pseudo-LabelsabstractAccurate extraction of building footprints from satellite imagery is of high value. Currently, deep learning methods are predominant in this field due to their powerful representation capabilities. However, they generally require extensive pixel-wise annotations, which constrains their practical application. Semi-supervised learning (SSL) significantly mitigates this requirement by leveraging large volumes of unlabeled data for model self-training (ST), thus enhancing the viability of building footprint extraction. Despite its advantages, SSL faces a critical challenge: the imbalanced distribution between the majority background class and the minority building class, which often results in model bias toward the background during training. To address this issue, this article introduces a novel method called DeBiased matching (DBMatch) for semi-supervised building footprint extraction. DBMatch comprises three main components: 1) a basic supervised learning module (SUP) that uses labeled data for initial model training; 2) a classical weak-to-strong ST module that generates pseudo-labels from unlabeled data for further model ST; and 3) a novel logit debiasing (LDB) module that calculates a global logit bias between building and background, allowing for dynamic pseudo-label calibration. To verify the effectiveness of the proposed DBMatch, extensive experiments are performed on three public building footprint extraction datasets covering six global cities in SSL setting. The experimental results demonstrate that our method significantly outperforms some advanced SSL methods in semi-supervised building footprint extraction. Our codes will be publicly provided athttps://github.com/zhu-xlab/SSL_Buildings. Wei Huang 0068, Ziqi Gu, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Height-Assisted Semi-Supervised Building Footprint Extraction From Optical Remote Sensing ImagesabstractAutomatic building footprint extraction from optical remote sensing (RS) images is popular and crucial for various downstream applications. Current building footprint extraction methods are mainly based on deep learning, which requires large amounts of manually labeled data for model training, limiting their practical deployment. Semi-supervised semantic segmentation (SSS), which leverages limited labeled data for supervised learning and abundant unlabeled data for unsupervised self-training, offers a promising solution to reduce this reliance. Nonetheless, directly applying existing SSS methods to building footprint extraction with limited labels fails to fully exploit the geometric structural features of buildings—key characteristics that distinguish them from background. To tackle this challenge, we propose a semi-supervised learning framework, HeightMatch, which integrates real or synthetic height information with RS images to extract more comprehensive and discriminative feature representations of buildings, particularly in limited-label scenarios. During training, these height maps effectively enhance the model’s ability to capture geometric structures, leading to more accurate pseudo-labels for unlabeled data and thereby enabling more effective self-training. At inference, building predictions rely solely on RS images, ensuring the practicality of the proposed method. Extensive experimental results on five widely-used building footprint extraction datasets demonstrate the effectiveness and superiority of our method in comparison with multiple state-of-the-art SSS methods. Our code is available at https://github.com/zhu-xlab/HeightMatch. Wei Huang 0068, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | RainScaler: A Physics-Inspired Network for Precipitation Correction and DownscalingabstractSpatial downscaling of precipitation, in which finegrained regional precipitation patterns are recovered from coarse-resolution images, plays a crucial role in various weather and meteorological analyses. However, the intricate noise information presented in the observation data intertwines with the fine-scale characteristics, which poses challenges for subsequent feature extraction. Regional precipitation suffers from complex spatial patterns. Moreover, the real observatory data contains information inconsistent with the established physical principle, due either to inaccurate or incomplete physical models or limited data quality, thus making the implementation of physicallyinformed deep learning more difficult. For example, strong physical constraints may lead to over-regularization, in which the model becomes too rigid and fails to capture certain complexities in the data. In this work, we propose RainScaler, a physicsinspired deep neural network, to tackle these issues. First, to remove the noise and preserve the vital precipitation patterns effectively, the proposed RainScaler exploits an Inconsistencyaware Denoising Net to explicitly model the spatial variability of noise in the input. In addition, a graph module is designed to learn the geographical-dependent fine-grained patterns in high dimensional feature space at a moderate computation cost. Finally, multi-scale physical constraints are skillfully embedded to incorporate additional insights into the data-driven framework. We test our approach on a public dataset consisting of over 60,000 real low-resolution and high-resolution precipitation map pairs collected by different sensors. Our method produces realisticlooking precipitation maps with better discernment capability and corrects the structural error of precipitation distribution, especially for extreme events. Moreover, we evaluate the potential risks of incorporating physical constraints in real-world data applications. Our method unveils opportunities for multi-source data fusion and provides possible solutions to improve the physical feasibility of data-driven models. The codes are available in https://github.com/zhu-xlab/RainScaler.git. Shan Zhao 0007, Zhitong Xiong, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Mono3DVG: 3D Visual Grounding in Monocular ImagesabstractWe introduce a novel task of 3D visual grounding in monocular RGB images using language descriptions with both appearance and geometry information. Specifically, we build a large-scale dataset, Mono3DRefer, which contains 3D object targets with their corresponding geometric text descriptions, generated by ChatGPT and refined manually. To foster this task, we propose Mono3DVG-TR, an end-to-end transformer-based network, which takes advantage of both the appearance and geometry information in text embeddings for multi-modal learning and 3D object localization. Depth predictor is designed to explicitly learn geometry features. The dual text-guided adapter is proposed to refine multiscale visual and geometry features of the referred object. Based on depth-text-visual stacking attention, the decoder fuses object-level geometric cues and visual appearance into a learnable query. Comprehensive benchmarks and some insightful analyses are provided for Mono3DVG. Extensive comparisons and ablation studies show that our method significantly outperforms all baselines. The dataset and code will be released. Yang Zhan 0007, Yuan Yuan 0001, Zhitong Xiong |
AAAI | 3 |
| 2024 | Representation Enhancement-Stabilization: Reducing Bias-Variance of Domain Generalization
Wei Huang 0068, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001 |
ECCV (36) | 3 |
| 2024 | Decoupling Common and Unique Representations for Multimodal Self-supervised Learning
Yi Wang 0072, Conrad M. Albrecht, Nassim Ait Ali Braham, Chenying Liu 0001, Zhitong Xiong, Xiao Xiang Zhu 0001 |
ECCV (29) | 5 |
| 2024 | Disentangling Semi-Supervised Semantic Segmentation of Remote Sensing ImagesabstractIn Earth observation, semantic understanding of Remote Sensing (RS) images holds significant importance, yet it is hindered in practice by the need for extensive manual pixel-level labeling. Semi-supervised semantic segmentation (SSS) of RS images would be a promising solution, which fully utilizes unlabeled data for model self-training under the guidance of limited labeled data. The mainstream SSS methods use pseudo-labels of the unlabeled data for model training, however, their performance is bottlenecked because of confirmation bias, i.e., stubborn incorrect pseudo-labels. To counter this, our study introduces a novel disentanglement learning (DL) method tailored for RS-SSS. It separates the predictions of the labeled and unlabeled data by two individual prediction heads during current training, and then integrates them during follow-up training. The experimental results verify its effectiveness on two widely-used RS semantic segmentation datasets in semi-supervised setting. Wei Huang 0068, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001 |
IGARSS | 3 |
| 2024 | One for All: Toward Unified Foundation Models for Earth VisionabstractFoundation models characterized by extensive parameters and trained on large-scale datasets have demonstrated remarkable efficacy across various downstream tasks for remote sensing data. Current remote sensing foundation models typically specialize in a single modality or a specific spatial resolution range, limiting their versatility for downstream datasets. While there have been attempts to develop multi-modal remote sensing foundation models, they typically employ separate vision encoders for each modality or spatial resolution, necessitating a switch in backbones contingent upon the input data. To address this issue, we introduce a simple yet effective method, termed OFA-Net (One-For-All Network): employing a single, shared Transformer backbone for multiple data modalities with different spatial resolutions. Using the masked image modeling mechanism, we pre-train a single Transformer backbone on a curated multi-modal dataset with this simple design. Then the backbone model can be used in different downstream tasks, thus forging a path towards a unified foundation backbone model in Earth vision. The proposed method is evaluated on 12 distinct downstream tasks and demonstrates promising performance. Zhitong Xiong, Yi Wang 0072, Fahong Zhang 0001, Xiao Xiang Zhu 0001 |
IGARSS | 1 |
| 2024 | Efficient Subseasonal Weather Forecast Using Teleconnection-Informed TransformersabstractSubseasonal forecasting, which is pivotal for agriculture, water resource management, and early warning of disasters, faces challenges due to the chaotic nature of the atmosphere. Recent advances in machine learning (ML) have revolutionized weather forecasting by achieving competitive predictive skills to numerical models. However, training such foundation models requires thousands of GPU days, which causes substantial carbon emissions and limits their broader applicability. Moreover, ML models tend to fool the pixel-wise error scores by producing smoothed results which lack physical consistency and meteorological meaning. To deal with the aforementioned problems, we propose a teleconnection-informed transformer. Our architecture leverages the pretrained Pangu model to achieve good initial weights and integrates a teleconnection-informed temporal module to improve predictability in an extended temporal range. Remarkably, by adjusting 1.1% of the Pangu model’s parameters, our method enhances predictability on four surface and five upper-level atmospheric variables at a two-week lead time. Furthermore, the teleconnection-filtered features improve the spatial granularity of outputs significantly, indicating their potential physical consistency. Our research underscores the importance of atmospheric and oceanic teleconnections in driving future weather conditions. Besides, it presents a resource-efficient pathway for researchers to leverage existing foundation models on versatile downstream tasks. Shan Zhao 0007, Zhitong Xiong, Xiao Xiang Zhu 0001 |
IGARSS | 2 |
| 2024 | Detail-Preserving and Diverse Image Translation for Adverse Visual Object DetectionabstractThe effectiveness of object detection is significantly hampered in challenging nighttime or rainy scenarios. This is due to the severe domain shifts between daytime and adverse-visual images. Previous methods have demonstrated that using image-to-image translation methods for data augmentation can effectively address domain shifts, but they may still fail in preserving image objects when faced with extreme adverse images like rainy nights. In addition, achieving diversity in the generated results remains challenging. To this end, we propose a Progressive Adverse Image Translation (PAIT) framework that tackles domain shifts by generating diverse and detail-preserving images. The main contributions of this paper are as follows. 1) We propose a novel PAIT framework, which incorporates an iterative mapping module and a slicing layer. This framework enables the progressive generation of increasingly challenging images in a fine-to-coarse manner. 2) To preserve the details of the images, we innovatively introduce an iterative mapping module to generate smooth style transform curves. 3) To enhance the diversity of synthesized images, a simple but efficient end-to-end optimization method is proposed. 4) We found a strong correlation between the style diversity of augmented images and the performance of the detection model through a quantitative analysis, highlighting the crucial role of style diversity in enhancing the model’s generalizability. Our framework achieves state-of-the-art performance on multiple challenging visual datasets, surpassing the current state-of-the-art methods by 27%(+8.0AP). Moreover, our approach and modules can be easily extended to different detectors and other domain adaptation methods, making it a versatile solution for object detection in adverse visual environments. Our code will be available athttps://github.com/ssunguotu/Diverse-Aug. Guolong Sun, Zhitong Xiong, Yuan Yuan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Self-Supervised Pretraining With Monocular Height Estimation for Semantic SegmentationabstractMonocular height estimation (MHE) is key for generating 3-D city models, essential for swift disaster response. Moving beyond the traditional focus on performance enhancement, our study breaks new ground by probing the interpretability of MHE networks. We have pioneeringly discovered that neurons within MHE models demonstrate selectivity for both height and semantic classes. This insight sheds light on the complex inner workings of MHE models and inspires innovative strategies for leveraging elevation data more effectively. Informed by this insight, we propose a pioneering framework that employs MHE as a self-supervised pretraining method for remote sensing (RS) imagery. This approach significantly enhances the performance of semantic segmentation tasks. Furthermore, we develop a disentangled latent transformer (DLT) module that leverages explainable deep representations from pretrained MHE networks for unsupervised semantic segmentation. Our method demonstrates the significant potential of MHE tasks in developing foundation models for sophisticated pixel-level semantic analyses. Additionally, we present a new dataset designed to benchmark the performance of both semantic segmentation and height estimation tasks. The dataset and code will be publicly available athttps://github.com/zhu-xlab/DLT-MHE.pytorch. Zhitong Xiong, Sining Chen, Yilei Shi, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Few-Shot Object Detection in Remote Sensing: Lifting the Curse of Incompletely Annotated Novel ObjectsabstractObject detection is an essential and fundamental task in computer vision and satellite image processing. Existing deep learning methods have achieved impressive performance thanks to the availability of large-scale annotated datasets. Yet, in real-world applications the availability of labels is limited. In this context, few-shot object detection (FSOD) has emerged as a promising direction, which aims at enabling the model to detect novel objects with only few of them annotated. However, many existing FSOD algorithms overlook a critical issue: when an input image contains multiple novel objects and only a subset of them are annotated, the unlabeled objects will be considered as background during training. This can cause confusions and severely impact the model’s ability to recall novel objects. To address this issue, we propose a self-training-based FSOD (ST-FSOD) approach, which incorporates the self-training mechanism into the few-shot fine-tuning process. ST-FSOD aims to enable the discovery of novel objects that are not annotated, and take them into account during training. On the one hand, we devise a two-branch region proposal networks (RPN) to separate the proposal extraction of base and novel objects, On another hand, we incorporate the student-teacher mechanism into RPN and the region of interest (RoI) head to include those highly confident yet unlabeled targets as pseudo labels. Experimental results demonstrate that our proposed method outperforms the state-of-the- art in various FSOD settings by a large margin. The codes will be publicly available at https://github.com/zhu-xlab/ST-FSOD. Fahong Zhang 0001, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Polyhedron-Based Graph Neural Network for Compact Building Model ReconstructionabstractThree-dimensional (3D) building models play a crucial role in shaping digital twin cities and enabling a wide range of urban applications. However, one challenge remains in obtaining a compact representation of buildings from remote sensing. This paper introduces a novel deep learning approach to reconstructing polygonal building models from LiDAR point clouds. Our method leverages a graph neural network to assemble the polyhedra generated through space partitioning, thereby formulating building surface reconstruction as a graph node classification problem. To facilitate network training, we construct a synthetic dataset by simulating aerial LiDAR point clouds on building surface meshes. Experimental results demonstrate the effectiveness of our method, achieving a polyhedral classification accuracy of 96.4%. Moreover, our approach offers high efficiency and interpretability through end-to-end optimization. Zhaiyu Chen, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001 |
IGARSS | 3 |
| 2023 | Adaptive Bins for Monocular Height Estimation from Single Remote Sensing ImagesabstractMonocular height estimation is of great importance in generating 3D city models from single remote sensing images, while it is a challenging task due to the ill-posed nature of the problem. To address the issue, we propose to adopt adaptive bins (AdaBins) for the network design, which enhances the representation capability of the network with the classification-regression paradigm and the incorporation of local features and global context via a vision transformer encoder. Besides, to weaken the biases of the trained networks caused by the long-tailed nature of the dataset, a head-tail cut is conducted for different treatments of head and tail pixels. Experiments show that improvements are expected with the proposed network on the proposed GBH dataset. Sining Chen, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001 |
IGARSS | 3 |
| 2023 | RSSOD-Bench: a Large-Scale Benchmark Dataset for Salient Object Detection in Optical Remote Sensing ImageryabstractWe present the RSSOD-Bench dataset for salient object detection (SOD) in optical remote sensing imagery. While SOD has achieved success in natural scene images with deep learning, research in SOD for remote sensing imagery (RSSOD) is still in its early stages. Existing RSSOD datasets have limitations in terms of scale, and scene categories, which make them misaligned with real-world applications. To address these shortcomings, we construct the RSSOD-Bench dataset, which contains images from four different cities in the USA1. The dataset provides annotations for various salient object categories, such as buildings, lakes, rivers, highways, bridges, aircraft, ships, athletic fields, and more. The salient objects in RSSOD-Bench exhibit large-scale variations, cluttered backgrounds, and different seasons. Unlike existing datasets, RSSOD-Bench offers uniform distribution across scene categories. We benchmark 23 different state-of-the-art approaches from both the computer vision and remote sensing communities. Experimental results demonstrate that more research efforts are required for the RSSOD task. Zhitong Xiong, Yanfeng Liu, Qi Wang 0009, Xiao Xiang Zhu 0001 |
IGARSS | 1 |
| 2023 | Exploring Geometric Deep Learning for Precipitation NowcastingabstractPrecipitation nowcasting (up to a few hours) remains a challenge due to the highly complex local interactions that need to be captured accurately. Convolutional Neural Networks rely on convolutional kernels convolving with grid data and the extracted features are trapped by limited receptive field, typically expressed in excessively smooth output compared to ground truth. Thus they lack the capacity to model complex spatial relationships among the grids. Geometric deep learning aims to generalize neural network models to non-Euclidean domains. Such models are more flexible in defining nodes and edges and can effectively capture dynamic spatial relationship among geographical grids. Motivated by this, we explore a geometric deep learning-based temporal Graph Convolutional Network (GCN) for precipitation nowcasting. The adjacency matrix that simulates the interactions among grid cells is learned automatically by minimizing the L1 loss between prediction and ground truth pixel value during the training procedure. Then, the spatial relationship is refined by GCN layers while the temporal information is extracted by 1D convolution with various kernel lengths. The neighboring information is fed as auxiliary input layers to improve the final result. We test the model on sequences of radar reflectivity maps over the Trento/Italy area. The results show that GCNs improves the effectiveness of modeling the local details of the cloud profile as well as the prediction accuracy by achieving decreased error measures. Shan Zhao 0007, Sudipan Saha, Zhitong Xiong, Niklas Boers, Xiao Xiang Zhu 0001 |
IGARSS | 3 |
| 2023 | HTC-DC Net: Monocular Height Estimation From Single Remote Sensing Imagesabstract3D geo-information is of great significance for understanding the living environment; however, 3D perception from remote sensing data, especially on a large scale, is restricted, mainly due to the high costs of 3D sensors such as LiDAR. To tackle this problem, we propose a method for monocular height estimation from optical imagery, which is currently one of the richest sources of remote sensing data. As an ill-posed problem, monocular height estimation requires well-designed networks for enhanced representations to improve performance. Moreover, the distribution of height values is long-tailed with the low-height pixels, e.g., the background, as the head, and thus trained networks are usually biased and tend to underestimate building heights. To solve the problems, instead of formalizing the problem as a regression task, we propose HTC-DC Net following the classification-regression paradigm, with the head-tail cut (HTC) and the distribution-based constraints (DCs) as the main contributions. HTC-DC Net is composed of the backbone network as the feature extractor, the HTC-AdaBins module, and the hybrid regression process. The HTC-AdaBins module serves as the classification phase to determine bins adaptive to each input image. It is equipped with a vision transformer encoder to incorporate local context with holistic information and involves an HTC to address the long-tailed problem in monocular height estimation for balancing the performances of foreground and background pixels. The hybrid regression process does the regression via the smoothing of bins from the classification phase, which is trained via DCs. The proposed network is tested on datasets of different resolutions, namely, DFC19 (1.3 m) and GBH (3 m). Experimental results show the superiority of the proposed network over existing methods by large margins. Extensive ablation studies demonstrate the effectiveness of each design component. Codes and trained models are published at https://github.com/zhu-xlab/HTC-DC-Net. Sining Chen, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | AdaptMatch: Adaptive Matching for Semisupervised Binary Segmentation of Remote Sensing ImagesabstractThere are various binary semantic segmentation tasks in remote sensing (RS) that aim to extract the foreground areas of interest, such as buildings and roads, from the background in satellite images. In particular, semi-supervised learning, which can use limited labeled data to guide a large amount of unlabeled data for model training, can significantly promote the fast applications of these tasks in practice. However, due to the predominance of the background in RS images, the foreground only accounts for a small proportion of the pixels. It poses a challenge: models are biased toward the majority class of the background, leading to poor performance on the minority class of the foreground. To address this issue, this paper proposes a novel and effective semi-supervised learning framework, Adaptive Matching (AdaptMatch), for RS binary segmentation. AdaptMatch calculates individual and adaptive thresholds of the foreground and background based on their convergence difficulty in an online manner at the training stage; the adaptive thresholds are then used to select the high-confidence pseudo-labeled data of the two classes for model self-training in turn. Extensive experiments are conducted on two widely-studied RS binary segmentation tasks, building footprint extraction and road extraction, to demonstrate the effectiveness and generalizability of the proposed method. The results show that the proposed AdaptMatch achieves superior performance compared with some state-of-the-art semi-supervised methods in RS binary segmentation tasks. The codes will be publicly available at https://github.com/zhu-xlab/AdaptMatch. Wei Huang 0068, Yilei Shi, Zhitong Xiong, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Distilling Knowledge From Super-Resolution for Efficient Remote Sensing Salient Object DetectionabstractCurrent state-of-the-art remote sensing salient object detectors always require high-resolution spatial context to ensure excellent performance, which incurs enormous computation costs and hinders real-time efficiency. In this work, we propose a universal super-resolution assisted learning (SRAL) framework to boost performance and accelerate the inference efficiency of existing approaches. To this end, we propose to reduce the spatial resolution of the input remote sensing images (RSIs), which is model-agnostic, and can be applied to existing algorithms without extra computation cost. Specifically, a transposed saliency detection decoder (TSDD) is designed to upsample interim features progressively. On top of it, an auxiliary super-resolution decoder (ASRD) is proposed to build a multitask learning (MTL) framework to investigate an efficient complementary paradigm of saliency detection and super-resolution. Furthermore, a novel task-fusion guidance module (TFGM) is proposed to effectively distill domain knowledge from the super-resolution auxiliary task to the salient object detection task in optical RSIs. The presented ASRD and TFGM can be omitted in the inference phase without any extra computational budget. Extensive experiments on three datasets show that the presented SRAL with 224×224 input is superior to more than 20 algorithms. Moreover, it can be successfully generalized to existing typical networks with significant accuracy improvements in a parameter-free manner. Codes and models are available at https://github.com/lyf0801/SRAL. Yanfeng Liu, Zhitong Xiong, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Transcending Pixels: Boosting Saliency Detection via Scene Understanding From Aerial ImageryabstractExisting remote sensing image salient object detection (RSI-SOD) methods widely perform object-level semantic understanding with pixel-level supervision, but ignore the image-level scene information. As a fundamental attribute of RSIs, the scene has a complex intrinsic correlation with salient objects, which may bring hints to improve saliency detection performance. However, existing RSI-SOD datasets lack both pixel- and image-level labels, and it is non-trivial to effectively transfer the scene domain knowledge for more accurate saliency localization. To address these challenges, we first annotate the image-level scene labels of three RSI-SOD datasets inspired by remote sensing scene classification. On top of it, we present a novel scene-guided dual-stream network (SDNet), which can perform cross-task knowledge distillation from the scene classification to facilitate accurate saliency detection. Specifically, a scene knowledge transfer module (SKTM) and a conditional dynamic guidance module (CDGM) are designed for extracting saliency key area as spatial attention from the scene subnet and guiding the saliency subnet to generate scene-enhanced saliency features, respectively. Finally, an object contour awareness module (OCAM) is introduced to enable the model to focus more on irregular spatial details of salient objects from the complicated background. Extensive experiments reveal that our SDNet outperforms over 20 state-of-the-art algorithms on three datasets. Moreover, we prove that the proposed framework is model-agnostic, and its extension to six baselines can bring significant performance benefits. Code will be available at https://github.com/lyf0801/SDNet. Yanfeng Liu, Zhitong Xiong, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | THE Benchmark: Transferable Representation Learning for Monocular Height EstimationabstractGenerating 3D city models rapidly is crucial for many applications. Monocular height estimation is one of the most efficient and timely ways to obtain large-scale geometric information. However, existing works focus primarily on training and testing models using unbiased datasets, which does not align well with real-world applications. Therefore, we propose a new benchmark dataset to study the transferability of height estimation models in a cross-dataset setting. To this end, we first design and construct a large-scale benchmark dataset for cross-dataset transfer learning on the height estimation task. This benchmark dataset includes a newly proposed large-scale synthetic dataset, a newly collected real-world dataset, and four existing datasets from different cities. Next, a new experimental protocol,few-shot cross-dataset transfer, is designed. Furthermore, in this paper, we propose a scale-deformable convolution module to enhance the window-based Transformer for handling the scale-variation problem in the height estimation task. Experimental results have demonstrated the effectiveness of the proposed methods in traditional and cross-dataset transfer settings. The datasets and codes are publicly available at https://mediatum.ub.tum.de/1662763 and https://thebenchmarkh.github.io/. Zhitong Xiong, Wei Huang 0068, Jingtao Hu, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Parameter-Efficient Transfer Learning for Remote Sensing Image-Text RetrievalabstractVision-and-language pre-training (VLP) models have experienced a surge in popularity recently. By fine-tuning them on specific datasets, significant performance improvements have been observed in various tasks. However, full fine-tuning of VLP models not only consumes a significant amount of computational resources but also has a significant environmental impact. Moreover, as remote sensing (RS) data is constantly being updated, full fine-tuning may not be practical for real-world applications. To address this issue, in this work, we investigate the parameter-efficient transfer learning (PETL) method to effectively and efficiently transfer visual-language knowledge from the natural domain to the RS domain on the image-text retrieval task. To this end, we make the following contributions. 1) We construct a novel and sophisticated PETL framework for the RS image-text retrieval (RSITR) task, which includes the pretrained CLIP model, a multimodal remote sensing adapter, and a hybrid multi-modal contrastive (HMMC) learning objective; 2) To deal with the problem of high intra-modal similarity in RS data, we design a simple yet effective HMMC loss; 3) We provide comprehensive empirical studies for PETL-based RS image-text retrieval. Our results demonstrate that the proposed method is promising and of great potential for practical applications. 4) We benchmark extensive state-of-the-art PETL methods on the RSITR task. Our proposed model only contains 0.16M training parameters, which can achieve a parameter reduction of 98.9% compared to full fine-tuning, resulting in substantial savings in training costs. Our retrieval performance exceeds traditional methods by 7-13% and achieves comparable or better performance than full fine-tuning. This work can provide new ideas and useful insights for RS vision-language tasks. Yuan Yuan 0001, Yang Zhan 0007, Zhitong Xiong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | RSVG: Exploring Data and Models for Visual Grounding on Remote Sensing DataabstractIn this article, we introduce the task of visual grounding for remote sensing data (RSVG). RSVG aims to localize the referred objects in remote sensing (RS) images with the guidance of natural language. To retrieve rich information from RS imagery using natural language, many research tasks, such as RS image visual question answering, RS image captioning, and RS image–text retrieval, have been investigated a lot. However, the object-level visual grounding on RS images is still underexplored. Thus, in this work, we propose to construct the dataset and explore deep learning models for the RSVG task. Specifically, our contributions can be summarized as follows. First, we build the new large-scale benchmark of RSVG based on detection in optical remote sensing (DIOR) dataset, termed DIOR-RSVG, to fully advance the research of RSVG. This new dataset includes image/expression/box triplets for training and evaluating visual grounding models. Second, we benchmark extensive state-of-the-art (SOTA) natural image visual grounding methods on the constructed DIOR-RSVG dataset, and some insightful analyses are provided based on the results. Third, a novel transformer-based multigranularity visual language fusion (MGVLF) module is proposed. Remotely sensed images are usually with large-scale variations and cluttered backgrounds. To deal with the scale-variation problem, the MGVLF module takes advantage of multiscale visual features and multigranularity textual embeddings to learn more discriminative representations. To cope with the cluttered background problem, MGVLF adaptively filters irrelevant noise and enhances salient features. In this way, our proposed model can incorporate more effective multilevel and multimodal features to boost performance. This work can provide useful insights for developing better RSVG models. Yang Zhan 0007, Zhitong Xiong, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Pseudo Features-Guided Self-Training for Domain Adaptive Semantic Segmentation of Satellite ImagesabstractSemantic segmentation is a fundamental and crucial task that is of great importance to real-world satellite image-based applications. Yet a widely acknowledged issue that occurs when applying the semantic segmentation models to unseen scenery is that the model will perform much poorer than when it was applied to scenery similar to the training data. This phenomenon is usually termed as the domain shift problem. To tackle it, this article presents a self-training-based unsupervised domain adaptation (UDA) method. Different from the previous self-training approaches which focus on rectifying and improving the quality of the pseudo labels, we instead seek to exploit feature-level relation among neighboring pixels to structure and regularize the prediction of the adapted model. Based on the assumption that spatial topological relation is maintained despite the impact of the domain shift, we propose a novel self-training mechanism to perform DA by exploiting local relation in the feature space spanned by the teacher model, from which the pseudo labels are generated. Quantitative experiments on four different public benchmarks demonstrate that the proposed method can outperform the other UDA methods. Besides, analytical experiments also intuitively verify the proposed assumption. Codes will be publicly available athttps://github.com/zhu-xlab/PFST. Fahong Zhang 0001, Yilei Shi, Zhitong Xiong, Wei Huang 0068, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Doubly Deformable Aggregation of Covariance Matrices for Few-Shot Segmentation
Zhitong Xiong, Haopeng Li 0001, Xiao Xiang Zhu 0001 |
ECCV (20) | 1 |
| 2022 | Towards Global Forest Biomass Estimators from Tree Height DataabstractIn order to estimate tree biomass, allometric equations take tree parameters such as tree height, wood density, circumference of trunk, and crown diameter as input parameters. Given that most of these quantities are challenging to be extracted from remote sensing data, we evaluate the option to approximate biomass by tree height only. We study our approach by evaluating linear regression, random forest, and Gaussian process regressor models when applied to the 2016 Jucker dataset. Results indicate that linear models fail to properly capture the relationship between biomass and tree height, but the Gaussian process regressor outperms the other two candidate models. Qian Song, Conrad M. Albrecht, Zhitong Xiong, Xiao Xiang Zhu 0001 |
IGARSS | 3 |
| 2022 | Knowledge Transfer for Label-Efficient Monocular Height EstimationabstractEstimating height from monocular remote sensing images is one of the most efficient ways for building large-scale 3D city models. However, existing deep learning based methods usu-ally require a large amount of training data, which could be cost-consuming or even not possible to obtain. Towards a label-efficient deep learning model, we propose a new task and dataset for weak-shot monocular height estimation. In this task, only the relative height labels between pairs of a small portion of points are given, which is cheaper and more friendly for humans to annotate. In addition, to enhance the model performance under the sparse and weak-shot super-vision, we propose a Transformer-based network for trans-ferring the learned knowledge from a large-scale synthetic dataset to real-world data. Experimental results have shown the effectiveness of the proposed method on a public dataset under the sparse and weak supervision. Zhitong Xiong, Xiao Xiang Zhu 0001 |
IGARSS | 1 |
| 2022 | Gradient Matters: Designing Binarized Neural Networks via Enhanced Information-FlowabstractBinarized neural networks (BNNs) have drawn significant attention in recent years, owing to great potential in reducing computation and storage consumption. While it is attractive, traditional BNNs usually suffer from slow convergence speed and dramatical accuracy-degradation on large-scale classification datasets. To minimize the gap between BNNs and deep neural networks (DNNs), we propose a new framework of designing BNNs, dubbed Hyper-BinaryNet, from the aspect of enhanced information-flow. Our contributions are threefold: 1) Considering the capacity-limitation in the backward pass, we propose an 1-bit convolution module named HyperConv. By exploiting the capacity of auxiliary neural networks, BNNs gain better performance on large-scale image classification task. 2) Considering the slow convergence speed in BNNs, we rethink the gradient accumulation mechanism and propose a hyper accumulation technique. By accumulating gradients in multiple variables rather than one as before, the gradient paths for each weight increase, which escapes BNNs from the gradient bottleneck problem during training. 3) Considering the ill-posed optimization problem, a novel gradient estimation warmup strategy, dubbed STE-Warmup, is developed. This strategy prevents BNNs from the unstable optimization process by progressively transferring neural networks from 32-bit to 1-bit. We conduct evaluations with variant architectures on three public datasets: CIFAR-10/100 and ImageNet. Compared with state-of-the-art BNNs, Hyper-BinaryNet shows faster convergence speed and outperforms existing BNNs by a large margin. Qi Wang 0009, Nianhui Guo, Zhitong Xiong, Zeping Yin, Xuelong Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Hybrid Feature Aligned Network for Salient Object Detection in Optical Remote Sensing ImageryabstractRecently, salient object detection in optical remote sensing images (RSI-SOD) has attracted great attention. Benefiting from the success of deep learning and the inspiration of natural SOD task, RSI-SOD has achieved fast progress over the past two years. However, existing methods usually suffer from the intrinsic problems of optical RSIs, 1) cluttered background; 2) scale variation of salient objects; 3) complicated edges and irregular topology. To remedy these problems, we propose a hybrid feature aligned network (HFANet) jointly modeling boundary learning to detect salient objects effectively. Specifically, we design a hybrid encoder by unifying two components to capture global context for mitigating the disturbance of complex background. Then, to detect multiscale salient objects effectively, we propose a Gated Fold-ASPP (GF-ASPP) to extract abundant context in the deep semantic features. Furthermore, an adjacent feature aligned module (AFAM) is presented for integrating adjacent features with unparameterized alignment strategy. Finally, we propose a novel interactive guidance loss (IGLoss) to combine saliency and edge detection, which can adaptively perform mutual supervision of the two sub-tasks to facilitate detection of salient objects with blurred edges and irregular topology. Adequate experimental results on three optical RSI-SOD datasets reveal that the presented approach exceeds 18 state-of-the-art ones. All codes and detection results are available athttps://github.com/lyf0801/HFANet. Qi Wang 0009, Yanfeng Liu, Zhitong Xiong, Yuan Yuan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Change Detection Meets Visual Question AnsweringabstractThe Earth’s surface is continually changing, and identifying changes plays an important role in urban planning and sustainability. Although change detection techniques have been successfully developed for many years, these techniques are still limited to experts and facilitators in related fields. In order to provide every user with flexible access to change information and help them better understand land-cover changes, we introduce a novel task: change detection-based visual question answering (CDVQA) on multi-temporal aerial images. In particular, multi-temporal images can be queried to obtain high level change-based information according to content changes between two input images. We first build a CDVQA dataset including multi-temporal image-question-answer triplets using an automatic question-answer generation method. Then, a baseline CDVQA framework is devised in this work, and it contains four parts: multi-temporal feature encoding, multi-temporal fusion, multi-modal fusion, and answer prediction. In addition, we also introduce a change enhancing module to multi-temporal feature encoding, aiming at incorporating more change-related information. Finally, effects of different backbones and multi-temporal fusion strategies are studied on the performance of CDVQA task. The experimental results provide useful insights for developing better CDVQA models, which are important for future research on this task. The dataset will be available at https://github.com/YZHJessica/CDVQA. Zhenghang Yuan, Lichao Mou, Zhitong Xiong, Xiao Xiang Zhu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | CM-Net: Concentric Mask Based Arbitrary-Shaped Text DetectionabstractRecently fast arbitrary-shaped text detection has become an attractive research topic. However, most existing methods are non-real-time, which may fall short in intelligent systems. Although a few real-time text methods are proposed, the detection accuracy is far behind non-real-time methods. To improve the detection accuracy and speed simultaneously, we propose a novel fast and accurate text detection framework, namely CM-Net, which is constructed based on a new text representation method and a multi-perspective feature (MPF) module. The former can fit arbitrary-shaped text contours by concentric mask (CM) in an efficient and robust way. The latter encourages the network to learn more CM-related discriminative features from multiple perspectives and brings no extra computational cost. Benefiting the advantages of CM and MPF, the proposed CM-Net only needs to predict one CM of the text instance to rebuild the text contour and achieves the best balance between detection accuracy and speed compared with previous works. Moreover, to ensure that multi-perspective features are effectively learned, the multi-factor constraints loss is proposed. Extensive experiments demonstrate the proposed CM is efficient and robust to fit arbitrary-shaped text instances, and also validate the effectiveness of MPF and constraints loss for discriminative text features recognition. Furthermore, experimental results show that the proposed CM-Net is superior to existing state-of-the-art (SOTA) real-time text detection methods in both detection speed and accuracy on MSRA-TD500, CTW1500, Total-Text, and ICDAR2015 datasets. Chuang Yang 0003, Mulin Chen, Zhitong Xiong, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Image Process. | 3 |
| 2022 | Looking Closer at the Scene: Multiscale Representation Learning for Remote Sensing Image Scene ClassificationabstractRemote sensing image scene classification has attracted great attention because of its wide applications. Although convolutional neural network (CNN)-based methods for scene classification have achieved excellent results, the large-scale variation of the features and objects in remote sensing images limits the further improvement of the classification performance. To address this issue, we present multiscale representation for scene classification, which is realized by a global-local two-stream architecture. This architecture has two branches of the global stream and local stream, which can individually extract the global features and local features from the whole image and the most important area. In order to locate the most important area in the whole image using only image-level labels, a weakly supervised key area detection strategy of structured key area localization (SKAL) is specially designed to connect the above two streams. To verify the effectiveness of the proposed SKAL-based two-stream architecture, we conduct comparative experiments based on three widely used CNN models, including AlexNet, GoogleNet, and ResNet18, on four public remote sensing image scene classification data sets, and achieve the state-of-the-art results on all the four data sets. Our codes are provided in https://github.com/hw2hwei/SKAL. Qi Wang 0009, Wei Huang 0068, Zhitong Xiong, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | ASK: Adaptively Selecting Key Local Features for RGB-D Scene RecognitionabstractIndoor scene images usually contain scattered objects and various scene layouts, which make RGB-D scene classification a challenging task. Existing methods still have limitations for classifying scene images with great spatial variability. Thus, how to extract local patch-level features effectively using only image label is still an open problem for RGB-D scene recognition. In this article, we propose an efficient framework for RGB-D scene recognition, which adaptively selects important local features to capture the great spatial variability of scene images. Specifically, we design a differentiable local feature selection (DLFS) module, which can extract the appropriate number of key local scene-related features. Discriminative local theme-level and object-level representations can be selected with DLFS module from the spatially-correlated multi-modal RGB-D features. We take advantage of the correlation between RGB and depth modalities to provide more cues for selecting local features. To ensure that discriminative local features are selected, the variational mutual information maximization loss is proposed. Additionally, the DLFS module can be easily extended to select local features of different scales. By concatenating the local-orderless and global-structured multi-modal features, the proposed framework can achieve state-of-the-art performance on public RGB-D scene recognition datasets. Zhitong Xiong, Yuan Yuan 0001, Qi Wang 0009 |
IEEE Trans. Image Process. | 1 |
| 2020 | Variational Context-Deformable ConvNets for Indoor Scene ParsingabstractContext information is critical for image semantic segmentation. Especially in indoor scenes, the large variation of object scales makes spatial-context an important factor for improving the segmentation performance. Thus, in this paper, we propose a novel variational context-deformable (VCD) module to learn adaptive receptive-field in a structured fashion. Different from standard ConvNets, which share fixed-size spatial context for all pixels, the VCD module learns a deformable spatial-context with the guidance of depth information: depth information provides clues for identifying real local neighborhoods. Specifically, adaptive Gaussian kernels are learned with the guidance of multimodal information. By multiplying the learned Gaussian kernel with standard convolution filters, the VCD module can aggregate flexible spatial context for each pixel during convolution. The main contributions of this work are as follows: 1) a novel VCD module is proposed, which exploits learnable Gaussian kernels to enable feature learning with structured adaptive-context; 2) variational Bayesian probabilistic modeling is introduced for the training of VCD module, which can make it continuous and more stable; 3) a perspective-aware guidance module is designed to take advantage of multi-modal information for RGB-D segmentation. We evaluate the proposed approach on three widely-used datasets, and the performance improvement has shown the effectiveness of the proposed method. Zhitong Xiong, Yuan Yuan 0001, Nianhui Guo, Qi Wang 0009 |
CVPR | 1 |
| 2020 | KALM: Key Area Localization Mechanism for Abnormality Detection in Musculoskeletal RadiographsabstractRecently abnormality detection in musculoskeletal radio-graphs has attracted many attentions. For abnormality detection, it is crucial to locate the most important area in the musculoskeletal radiographs. To achieve this goal, we propose a key area localization mechanism (KALM) for abnormality detection for the first time in this paper. The proposed KALM explicitly defines the process of selecting the most important area from the whole image with using only image-level label. Based on KALM, we further present a joint global and local feature representation strategy for abnormality detection which takes as input both the entire image and the selected local area. The experimental results based on several classical convolutional neural network (CNN) architectures of MURA, the largest abnormality detection dataset of musculoskeletal radiographs, demonstrate the effectiveness of our KALM. Wei Huang 0068, Zhitong Xiong, Qi Wang 0009, Xuelong Li 0001 |
ICASSP | 2 |
| 2020 | MSN: Modality separation networks for RGB-D scene recognition
Zhitong Xiong, Yuan Yuan 0001, Qi Wang 0009 |
Neurocomputing | 1 |
| 2019 | ACM: Adaptive Cross-Modal Graph Convolutional Neural Networks for RGB-D Scene RecognitionabstractRGB image classification has achieved significant performance improvement with the resurge of deep convolutional neural networks. However, mono-modal deep models for RGB image still have several limitations when applied to RGB-D scene recognition. 1) Images for scene classification usually contain more than one typical object with flexible spatial distribution, so the object-level local features should also be considered in addition to global scene representation. 2) Multi-modal features in RGB-D scene classification are still under-utilized. Simply combining these modal-specific features suffers from the semantic gaps between different modalities. 3) Most existing methods neglect the complex relationships among multiple modality features. Considering these limitations, this paper proposes an adaptive crossmodal (ACM) feature learning framework based on graph convolutional neural networks for RGB-D scene recognition. In order to make better use of the modal-specific cues, this approach mines the intra-modality relationships among the selected local features from one modality. To leverage the multi-modal knowledge more effectively, the proposed approach models the inter-modality relationships between two modalities through the cross-modal graph (CMG). We evaluate the proposed method on two public RGB-D scene classification datasets: SUN-RGBD and NYUD V2, and the proposed method achieves state-of-the-art performance. Yuan Yuan 0001, Zhitong Xiong, Qi Wang 0009 |
AAAI | 2 |
| 2019 | VSSA-NET: Vertical Spatial Sequence Attention Network for Traffic Sign DetectionabstractAlthough traffic sign detection has been studied for years and great progress has been made with the rise of deep learning technique, there are still many problems remaining to be addressed. For complicated real-world traffic scenes, there are two main challenges. First, traffic signs are usually small-sized objects, which makes them more difficult to detect than large ones; second, it is hard to distinguish false targets which resemble real traffic signs in complex street scenes without context information. To handle these problems, we propose a novel end-to-end deep learning method for traffic sign detection in complex environments. Our contributions are as follows: 1) we propose a multi-resolution feature fusion network architecture which exploits densely connected deconvolution layers with skip connections, and can learn more effective features for a small-size object and 2) we frame the traffic sign detection as a spatial sequence classification and regression task, and propose a vertical spatial sequence attention module to gain more context information for better detection performance. To comprehensively evaluate the proposed method, we experiment on several traffic sign datasets as well as the general object detection dataset, and the results have shown the effectiveness of our proposed method. Yuan Yuan 0001, Zhitong Xiong, Qi Wang 0009 |
IEEE Trans. Image Process. | 2 |
| 2018 | AI-NET: Attention Inception Neural Networks for Hyperspectral Image ClassificationabstractRecently, deep learning methods have dominated many fields thanks to its powerful discriminative feature learning ability. While for hyperspectral images (HSI) analysis, these deep neural networks methods suffer from overfitting as the number of labeled training samples are limited. Thus more efficient neural network architecture should be designed to improve the performance of HSI classification task. In this paper, a novel attention inception module is introduced to extract features dynamically from multi-resolution convolutional filters. The AI-NET constructed by stacking the proposed attention inception module can adaptively learn the network architecture by dynamically routing between the attention inception modules. By exploiting different spatial size convolutional filters and dynamic CNN architecture, more representative feature can be learned with limited training samples. Extensive experimental results have shown that the proposed method can adaptively adjust the network architecture and obtain better classification performance. Zhitong Xiong, Yuan Yuan 0001, Qi Wang 0009 |
IGARSS | 1 |
| 2017 | An Incremental Framework for Video-Based Traffic Sign Detection, Tracking, and RecognitionabstractVideo-based traffic sign detection, tracking, and recognition is one of the important components for the intelligent transport systems. Extensive research has shown that pretty good performance can be obtained on public data sets by various state-of-the-art approaches, especially the deep learning methods. However, deep learning methods require extensive computing resources. In addition, these approaches mostly concentrate on single image detection and recognition task, which is not applicable in real-world applications. Different from previous research, we introduce a unified incremental computational framework for traffic sign detection, tracking, and recognition task using the mono-camera mounted on a moving vehicle under non-stationary environments. The main contributions of this paper are threefold: (1) to enhance detection performance by utilizing the contextual information, this paper innovatively utilizes the spatial distribution prior of the traffic signs; (2) to improve the tracking performance and localization accuracy under non-stationary environments, a new efficient incremental framework containing off-line detector, online detector, and motion model predictor together is designed for traffic sign detection and tracking simultaneously; and (3) to get a more stable classification output, a scale-based intra-frame fusion method is proposed. We evaluate our method on two public data sets and the performance has shown that the proposed system can obtain results comparable with the deep learning method with less computing resource in a near-real-time manner. Yuan Yuan 0001, Zhitong Xiong, Qi Wang 0009 |
IEEE Trans. Intell. Transp. Syst. | 2 |