EDBT 2026 Demo / reviewers in the wild / expert
Zhuo Zheng
dblp:217/4050
· DBLP profile ↗
35ranked-venue papers
12as first author
26since 2021 · last 2026
0000-0003-1811-6725ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 16 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SALR: Sparsity-Aware Low-Rank Representation for Efficient Fine-Tuning of Large Language ModelsabstractAdapting large pre-trained language models to downstream tasks often entails fine-tuning millions of parameters or deploying costly dense weight updates, which hinders their use in resource-constrained environments. Low-rank Adaptation (LoRA) reduces trainable parameters by factorizing weight updates, yet the underlying dense weights still impose high storage and computation costs. Magnitude-based pruning can yield sparse models but typically degrades LoRA’s performance when applied naively. In this paper, we introduce SALR (Sparsity-Aware Low-Rank Representation), a novel fine-tuning paradigm that unifies low-rank adaptation with sparse pruning under a rigorous mean-squared-error framework. We prove that statically pruning only the frozen base weights minimizes the pruning error bound, and we recover the discarded residual information via a truncated-SVD low-rank adapter, which provably reduces per-entry MSE by a factor of (1 - r/min(d, k)). To maximize hardware efficiency, we fuse multiple low-rank adapters into a single concatenated GEMM, and we adopt a bitmap-based encoding with a two-stage pipelined decoding + GEMM design to achieve true model compression and speedup. Empirically, SALR attains 50% sparsity on various LLMs while matching the performance of LoRA on GSM8K and MMLU, reduces model size by 2x, and delivers up to a 1.7x inference speedup. Longteng Zhang, Sen Wu 0001, Zhengyu Qing, Zhuo Zheng, Danning Ke, Qihong Lin, Qiang Wang 0022, Shaohuai Shi, Xiaowen Chu 0001 |
AAAI | 5 |
| 2025 | TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation DataabstractLarge vision and language assistants have enabled new capabilities for interpreting natural images. These approaches have recently been adapted to earth observation data, but they are only able to handle single image inputs, limiting their use for many real-world tasks. In this work, we develop a new vision and language assistant called TEOChat that can engage in conversations about temporal sequences of earth observation data. To train TEOChat, we curate an instruction-following dataset composed of many single image and temporal tasks including building change and damage assessment, semantic change detection, and temporal scene classification. We show that TEOChat can perform a wide variety of spatial and temporal reasoning tasks, substantially outperforming previous vision and language assistants, and even achieving comparable or better performance than several specialist models trained to perform specific tasks. Furthermore, TEOChat achieves impressive zero-shot performance on a change detection and change question answering dataset, outperforms GPT-4o and Gemini 1.5 Pro on multiple temporal tasks, and exhibits stronger single image capabilities than a comparable single image instruction-following model on scene classification, visual question answering, and captioning. We publicly release our data, models, and code at https://github.com/ermongroup/TEOChat . Jeremy Irvin, Emily Ruoyu Liu, Joyce Chuyi Chen, Ines Dormoy, Samar Khanna, Zhuo Zheng, Stefano Ermon |
ICLR | 7 |
| 2025 | Open-CD: A Comprehensive Toolbox for Change DetectionabstractWe present Open-CD, a change detection toolbox that contains a rich set of change detection methods as well as related components and modules. The toolbox started from a series of open source general vision task tools, including OpenMMLab Toolkits, PyTorch Image Models (Timm), etc. It gradually evolves into a unified platform that covers many popular change detection methods and contemporary modules. It not only includes training and inference codes, but also provides some useful scripts for data analysis. We believe this toolbox is by far the most comprehensive change detection toolbox. In this report, we introduce the features, supported methods and applications of Open-CD. In addition, we also conduct a benchmarking study on different methods and components. We wish that the toolbox and benchmark could serve the growing research community by providing a flexible toolkit to re-implement existing methods and develop their own new change detectors. Code and models are available at https://github.com/likyoo/open-cd. Kaiyu Li 0001, Chengxi Han, Yupeng Deng 0001, Keyan Chen 0001, Zhuo Zheng, Hao Chen 0045, Ziyuan Liu 0006, Yuantao Gu, Zhengxia Zou, Zhenwei Shi 0001, Sheng Fang 0001, Deyu Meng, Zhi Wang 0002, Xiangyong Cao |
ACM Multimedia | 6 |
| 2025 | DisasterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and ResponseabstractLarge vision-language models (VLMs) have made great achievements in Earth vision. However, complex disaster scenes with diverse disaster types, geographic regions, and satellite sensors have posed new challenges for VLM applications. To fill this gap, we curate the first remote sensing vision-language dataset (DisasterM3) for global-scale disaster assessment and response. DisasterM3 includes 26,988 bi-temporal satellite images and 123k instruction pairs across 5 continents, with three characteristics: **1) Multi-hazard**: DisasterM3 involves 36 historical disaster events with significant impacts, which are categorized into 10 common natural and man-made disasters. **2) Multi-sensor**: Extreme weather during disasters often hinders optical sensor imaging, making it necessary to combine Synthetic Aperture Radar (SAR) imagery for post-disaster scenes. **3) Multi-task**: Based on real-world scenarios, DisasterM3 includes 9 disaster-related visual perception and reasoning tasks, harnessing the full potential of VLM's reasoning ability with progressing from disaster-bearing body recognition to structural damage assessment and object relational reasoning, culminating in the generation of long-form disaster reports. We extensively evaluated 14 generic and remote sensing VLMs on our benchmark, revealing that state-of-the-art models struggle with the disaster tasks, largely due to the lack of a disaster-specific corpus, cross-sensor gap, and damage object counting insensitivity. Focusing on these issues, we fine-tune four VLMs using our dataset and achieve stable improvements (up to 10.4\%$\uparrow$QA, 2.1$\uparrow$Report, 40.8\%$\uparrow$Referring Seg.) with robust cross-sensor and cross-disaster generalization capabilities. Project: https://github.com/Junjue-Wang/DisasterM3. Weihao Xuan, Heli Qi, Kunyi Liu, Hongruixuan Chen, Jian Song 0010, Junshi Xia, Zhuo Zheng, Naoto Yokoya |
NeurIPS | 10 |
| 2025 | DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City UnderstandingabstractMultimodal large language models (MLLMs) have demonstrated remarkable capabilities in visual understanding, but their application to long-term Earth observation analysis remains limited, primarily focusing on single-temporal or bi-temporal imagery. To address this gap, we introduce DVL-Suite, a comprehensive framework for analyzing long-term urban dynamics through remote sensing imagery. Our suite comprises 14,871 high-resolution (1.0m) multi-temporal images spanning 42 major cities in the U.S. from 2005 to 2023, organized into two components: DVL-Bench and DVL-Instruct. The DVL-Bench includes six urban understanding tasks, from fundamental change detection (pixel-level) to quantitative analyses (regional-level) and comprehensive urban narratives (scene-level), capturing diverse urban dynamics including expansion/transformation patterns, disaster assessment, and environmental challenges. We evaluate 18 state-of-the-art MLLMs and reveal their limitations in long-term temporal understanding and quantitative analysis. These challenges motivate the creation of DVL-Instruct, a specialized instruction-tuning dataset designed to enhance models' capabilities in multi-temporal Earth observation. Building upon this dataset, we develop DVLChat, a baseline model capable of both image-level question-answering and pixel-level segmentation, facilitating a comprehensive understanding of city dynamics through language interactions. Project: https://github.com/weihao1115/dynamicvl. Weihao Xuan, Heli Qi, Zihang Chen 0001, Zhuo Zheng, Yanfei Zhong, Junshi Xia, Naoto Yokoya |
NeurIPS | 5 |
| 2025 | Changen2: Multi-Temporal Remote Sensing Generative Change Foundation ModelabstractOur understanding of the temporal dynamics of the Earth's surface has been significantly advanced by deep vision models, which often require a massive amount of labeled multi-temporal images for training. However, collecting, preprocessing, and annotating multi-temporal remote sensing images at scale is non-trivial since it is expensive and knowledge-intensive. In this paper, we present scalable multi-temporal change data generators based on generative models, which are cheap and automatic, alleviating these data problems. Our main idea is to simulate a stochastic change process over time. We describe the stochastic change process as a probabilistic graphical model, namely the generative probabilistic change model (GPCM), which factorizes the complex simulation problem into two more tractable sub-problems, i.e., condition-level change event simulation and image-level semantic change synthesis. To solve these two problems, we present Changen2, a GPCM implemented with a resolution-scalable diffusion transformer which can generate time series of remote sensing images and corresponding semantic and change labels from labeled and even unlabeled single-temporal images. Changen2 is a "generative change foundation model" that can be trained at scale via self-supervision, and is capable of producing change supervisory signals from unlabeled single-temporal images. Unlike existing "foundation models", our generative change foundation model synthesizes change data to train task-specific foundation models for change detection. The resulting model possesses inherent zero-shot change detection capabilities and excellent transferability. Comprehensive experiments suggest Changen2 has superior spatiotemporal scalability in data generation, e.g., Changen2 model trained on 256 pixel single-temporal images can yield time series of any length and resolutions of 1,024 pixels. Changen2 pre-trained models exhibit superior zero-shot performance (narrowing the performance gap to 3% on LEVIR-CD and approximately 10% on both S2Looking and SECOND, compared to fully supervised counterpart) and transferability across multiple types of change tasks, including ordinary and off-nadir building change, land-use/land-cover change, and disaster assessment. Zhuo Zheng, Stefano Ermon, Liangpei Zhang 0001, Yanfei Zhong |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Learning Temporal Consistency for High Spatial Resolution Remote Sensing Imagery Semantic Change DetectionabstractSemantic change detection (SCD) is a crucial task in remote sensing imagery interpretation, which identifies where the changes are and what categories the objects before and after the changes belong to. Post-classification comparison (PCC), as the simplest method, suffers from severe false alarms. While multi-task methods face another issue, in that the changed areas in the binary change map have same object categories in the semantic maps, i.e., semantic change inconsistency. In this study, we leveraged the commutative property of temporal semantic labels in unchanged areas to generate temporal consistency embedding and propose temporal-semantic feature blending (TSFB), which is a feature interaction operation that can dynamically adjust the distance between bi-temporal features by controlling a blending weight. In order to realize adaptive temporal consistency learning, we propose the TEmporal-Semantic feature CalibratiOn (TESCO) module, which can estimate the optimal value of the blending weight for TSFB automatically and make the segmentation network learn temporal consistency from coarse to fine through recurrence and the end-to-end training process. The TESCO module is a plug-and-play module that can be combined with any SCD method. In this study, we added the TESCO module to both PCC and multi-task based SCD methods and conducted comprehensive experiments on three large-scale and different application scenario SCD datasets with different semantic change numbers and time series. The experimental results show that the TESCO module can effectively learn temporal consistency, resulting in a significant improvement in the performance of PCC methods and the semantic consistency of the multi-task based methods. The implementation of TESCO module will be available at https://github.com/Daisy-7/TESCO. Shiqi Tian, Ailong Ma, Zhuo Zheng, Xicheng Tan, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Advancing Weakly-Supervised Change Detection in Satellite Images via Adversarial Class PromptingabstractWeakly-Supervised Change Detection (WSCD) aims to distinguish specific object changes (e.g., objects appearing or disappearing) from background variations (e.g., environmental changes due to light, weather, or seasonal shifts) in paired satellite images, relying only on paired image (i.e., image-level) classification labels. This technique significantly reduces the need for dense annotations required in fully-supervised change detection. However, as image-level supervision only indicates whether objects have changed in a scene, WSCD methods often misclassify background variations as object changes, especially in complex remote-sensing scenarios. In this work, we propose an Adversarial Class Prompting (AdvCP) method to address this co-occurring noise problem, including two phases: a) Adversarial Prompt Mining: After each training iteration, we introduce adversarial prompting perturbations, using incorrect one-hot image-level labels to activate erroneous feature mappings. This process reveals co-occurring adversarial samples under weak supervision, namely background variation features that are likely to be misclassified as object changes. b) Adversarial Sample Rectification: We integrate these adversarially prompt-activated pixel samples into training by constructing an online global prototype. This prototype is built from an exponentially weighted moving average of the current batch and all historical training data. Serving as an unbiased anchor, the global prototype guides the rectification of adversarial pixel samples. Our AdvCP can be seamlessly integrated into current WSCD methods without adding additional inference cost. Experiments on ConvNet, Transformer, and Segment Anything Model (SAM)-based baselines demonstrate significant performance enhancements, achieving up to 7.37%, 7.46%, and 6.56% IoU improvements on the WHU-CD, LEVIR-CD, and DSIFN-CD datasets. Furthermore, we demonstrate the generalizability of AdvCP to other multi-class weakly-supervised dense prediction scenarios. Code is available at https://github.com/zhenghuizhao/AdvCP. Zhenghui Zhao, Chen Wu 0003, Di Wang 0023, Hongruixuan Chen, Cuiqun Chen, Zhuo Zheng, Bo Du 0001, Liangpei Zhang 0001 |
IEEE Trans. Image Process. | 6 |
| 2024 | EarthVQA: Towards Queryable Earth via Relational Reasoning-Based Remote Sensing Visual Question AnsweringabstractEarth vision research typically focuses on extracting geospatial object locations and categories but neglects the exploration of relations between objects and comprehensive reasoning. Based on city planning needs, we develop a multi-modal multi-task VQA dataset (EarthVQA) to advance relational reasoning-based judging, counting, and comprehensive analysis. The EarthVQA dataset contains 6000 images, corresponding semantic masks, and 208,593 QA pairs with urban and rural governance requirements embedded. As objects are the basis for complex relational reasoning, we propose a Semantic OBject Awareness framework (SOBA) to advance VQA in an object-centric way. To preserve refined spatial locations and semantics, SOBA leverages a segmentation network for object semantics generation. The object-guided attention aggregates object interior features via pseudo masks, and bidirectional cross-attention further models object external relations hierarchically. To optimize object counting, we propose a numerical difference loss that dynamically adds difference penalties, unifying the classification and regression tasks. Experimental results show that SOBA outperforms both advanced general and remote sensing methods. We believe this dataset and framework provide a strong benchmark for Earth vision's complex analysis. The project page is at https://Junjue-Wang.github.io/homepage/EarthVQA. Zhuo Zheng, Zihang Chen 0001, Ailong Ma, Yanfei Zhong |
AAAI | 2 |
| 2024 | LUWA Dataset: Learning Lithic Use-Wear Analysis on Microscopic ImagesabstractLithic Use-Wear Analysis (LUWA) using microscopic images is an underexplored vision-for-science research area. It seeks to distinguish the worked material, which is critical for understanding archaeological artifacts, material interactions, tool functionalities, and dental records. However, this challenging task goes beyond the well-studied image classification problem for common objects. It is affected by many confounders owing to the complex wear mechanism and microscopic imaging, which makes it difficult even for human experts to identify the worked material successfully. In this paper, we investigate the following three questions on this unique vision task for the first time:(i) How well can state-of-the-art pre-trained models (like DINOv2) generalize to the rarely seen domain? (ii) How can few-shot learning be exploited for scarce microscopic images? (iii) How do the ambiguous magnification and sensing modality influence the classification accuracy? To study these, we collaborated with archaeologists and built the first open-source and the largest LUWA dataset containing 23,130 microscopic images with different magnifications and sensing modalities. Extensive experiments show that existing pretrained models notably outperform human experts but still leave a large gap for improvements. Most importantly, the LUWA dataset provides an underexplored opportunity for vision and learning communities and complements existing image classification problems on common objects. Irving Fang, Akshat Kaushik, Alice Rodriguez, Hanwen Zhao, Juexiao Zhang, Zhuo Zheng, Radu Iovita, Chen Feng 0002 |
CVPR | 8 |
| 2024 | MapChange: Enhancing Semantic Change Detection with Temporal-Invariant Historical Maps Based on Deep Triplet NetworkabstractSemantic Change Detection (SCD) is recognized as both a crucial and challenging task in the field of image analysis. Traditional methods for SCD have predominantly relied on the comparison of image pairs. However, this approach is significantly hindered by substantial imaging differences, which arise due to variations in shooting times, atmospheric conditions, and angles. Such discrepancies lead to two primary issues: the under-detection of minor yet significant changes, and the generation of false alarms due to temporal variances. These factors often result in unchanged objects appearing markedly different in multi-temporal images. In response to these challenges, the MapChange framework has been developed. This framework introduces a novel paradigm that synergizes temporal-invariant historical map data with contemporary high-resolution images. By employing this combination, the temporal variance inherent in conventional image pair comparisons is effectively mitigated. The efficacy of the MapChange framework has been empirically validated through comprehensive testing on two public datasets. These tests have demonstrated the framework's marked superiority over existing state-of-the-art SCD methods. Yinhe Liu, Sunan Shi, Zhuo Zheng, Jue Wang 0011, Shiqi Tian, Yanfei Zhong |
IGARSS | 3 |
| 2024 | Segment Any ChangeabstractVisual foundation models have achieved remarkable results in zero-shot image classification and segmentation, but zero-shot change detection remains an open problem.
In this paper, we propose the segment any change models (AnyChange), a new type of change detection model that supports zero-shot prediction and generalization on unseen change types and data distributions.
AnyChange is built on the segment anything model (SAM) via our training-free adaptation method, bitemporal latent matching.
By revealing and exploiting intra-image and inter-image semantic similarities in SAM's latent space, bitemporal latent matching endows SAM with zero-shot change detection capabilities in a training-free way.
We also propose a point query mechanism to enable AnyChange's zero-shot object-centric change detection capability.
We perform extensive experiments to confirm the effectiveness of AnyChange for zero-shot change detection.
AnyChange sets a new record on the SECOND benchmark for unsupervised change detection, exceeding the previous SOTA by up to 4.4\% F$_1$ score, and achieving comparable accuracy with negligible manual annotations (1 pixel per image) for supervised change detection. Code is available at https://github.com/Z-Zheng/pytorch-change-models. Zhuo Zheng, Yanfei Zhong, Liangpei Zhang 0001, Stefano Ermon |
NeurIPS | 1 |
| 2024 | Single-Temporal Supervised Learning for Universal Remote Sensing Change Detection
Zhuo Zheng, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
Int. J. Comput. Vis. | 1 |
| 2024 | Explicable Fine-Grained Aircraft Recognition Via Deep Part Parsing Prior Framework for High-Resolution Remote Sensing ImageryabstractAircraft recognition is crucial in both civil and military fields, and high-spatial resolution remote sensing has emerged as a practical approach. However, existing data-driven methods fail to locate discriminative regions for effective feature extraction due to limited training data, leading to poor recognition performance. To address this issue, we propose a knowledge-driven deep learning method called the explicable aircraft recognition framework based on a part parsing prior (APPEAR). APPEAR explicitly models the aircraft's rigid structure as a pixel-level part parsing prior, dividing it into five parts: 1) the nose; 2) left wing; 3) right wing; 4) fuselage; and 5) tail. This fine-grained prior provides reliable part locations to delineate aircraft architecture and imposes spatial constraints among the parts, effectively reducing the search space for model optimization and identifying subtle interclass differences. A knowledge-driven aircraft part attention (KAPA) module uses this prior to achieving a geometric-invariant representation for identifying discriminative features. Part features are generated by part indexing in a specific order and sequentially embedded into a compact space to obtain a fixed-length representation for each part, invariant to aircraft orientation and scale. The part attention module then takes the embedded part features, adaptively reweights their importance to identify discriminative parts, and aggregates them for recognition. The proposed APPEAR framework is evaluated on two aircraft recognition datasets and achieves superior performance. Moreover, experiments with few-shot learning methods demonstrate the robustness of our framework in different tasks. Ablation analysis illustrates that the fuselage and wings of the aircraft are the most effective parts for recognition. Yanfei Zhong, Ailong Ma, Zhuo Zheng, Liangpei Zhang 0001 |
IEEE Trans. Cybern. | 4 |
| 2024 | Adaptive Self-Supporting Prototype Learning for Remote Sensing Few-Shot Semantic SegmentationabstractThe semantic segmentation of remote sensing images with few shots has important theoretical and application value. Most of the existing few-shot semantic segmentation frameworks are based on prototype learning methods, in which a single support prototype is designed to guide the query set for prediction. However, the visual differences between the support set and the query set make it difficult for a single support prototype, generated from the support set, to comprehensively encapsulate the semantic information of all the query images. This article introduces an adaptive self-supporting prototype learning network designed for few-shot segmentation (FSS), in order to tackle the challenges mentioned earlier. We propose adaptive hyperprototype representation (HPR), which consists of hyperprototype clustering (HPC) and guided prototype matching (GPM), to generate and assign multiple representative prototypes to compensate for the limitations of a single prototype in representing the semantic information of the query images. Specifically, HPC is a parameter-free and adaptive approach, which can extract more representative prototypes by aggregating similar feature vectors utilizing superpixel feature clustering. Meanwhile, GPM can select matched prototypes to provide more accurate guidance, allowing for uniformly aligned representation of multiple prototypes and complex image semantic information. We also introduce self-supporting matching (SSM) prototype learning, which can accurately guide the query set segmentation by acquiring query set prototypes. SSM generates initial pseudo labels for the query set based on the support set prototypes, and further guides the query set using the pseudo labels, along with the query prototypes generated by its own features, thus effectively avoiding visual differences between the support set and query set. The proposed adaptive self-supporting prototype learning network substantially improves the prototype quality and achieves a superior performance on object-level remote sensing datasets. Weihao Shen, Ailong Ma, Zhuo Zheng, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Scalable Multi-Temporal Remote Sensing Change Data Generation via Simulating Stochastic Change ProcessabstractUnderstanding the temporal dynamics of Earth’s surface is a mission of multi-temporal remote sensing image analysis, significantly promoted by deep vision models with its fuel—labeled multi-temporal images. However, collecting, preprocessing, and annotating multi-temporal remote sensing images at scale is non-trivial since it is expensive and knowledge-intensive. In this paper, we present a scalable multi-temporal remote sensing change data generator via generative modeling, which is cheap and automatic, alleviating these problems. Our main idea is to simulate a stochastic change process over time. We consider the stochastic change process as a probabilistic semantic state transition, namely generative probabilistic change model (GPCM), which decouples the complex simulation problem into two more trackable sub-problems, i.e., change event simulation and semantic change synthesis. To solve these two problems, we present the change generator (Changen), a GAN-based GPCM, enabling controllable object change data generation, including customizable object property, and change event. The extensive experiments suggest that our Changen has superior generation capability, and the change detectors with Changen pre-training exhibit excellent transferability to real-world change datasets. Zhuo Zheng, Shiqi Tian, Ailong Ma, Liangpei Zhang 0001, Yanfei Zhong |
ICCV | 1 |
| 2023 | FarSeg++: Foreground-Aware Relation Network for Geospatial Object Segmentation in High Spatial Resolution Remote Sensing ImageryabstractGeospatial object segmentation, a fundamental Earth vision task, always suffers from scale variation, the larger intra-class variance of background, and foreground-background imbalance in high spatial resolution (HSR) remote sensing imagery. Generic semantic segmentation methods mainly focus on the scale variation in natural scenarios. However, the other two problems are insufficiently considered in large area Earth observation scenarios. In this paper, we propose a foreground-aware relation network (FarSeg++) from the perspectives of relation-based, optimization-based, and objectness-based foreground modeling, alleviating the above two problems. From the perspective of the relations, the foreground-scene relation module improves the discrimination of the foreground features via the foreground-correlated contexts associated with the object-scene relation. From the perspective of optimization, foreground-aware optimization is proposed to focus on foreground examples and hard examples of the background during training to achieve a balanced optimization. Besides, from the perspective of objectness, a foreground-aware decoder is proposed to improve the objectness representation, alleviating the objectness prediction problem that is the main bottleneck revealed by an empirical upper bound analysis. We also introduce a new large-scale high-resolution urban vehicle segmentation dataset to verify the effectiveness of the proposed method and push the development of objectness prediction further forward. The experimental results suggest that FarSeg++ is superior to the state-of-the-art generic semantic segmentation methods and can achieve a better trade-off between speed and accuracy. Zhuo Zheng, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | GRE and Beyond: A Global Road Extraction DatasetabstractAccurate and timely road mapping that describes the road network geometry and topology is the key element of intelligent transport systems and smart city management. However, current global road maps like OpenStreetMap (OSM) are typically outdated and spatially incomplete with uneven accuracies. Although the development of remote sensing satellite technology and the advance of computer vision technology have made it possible to quickly extract road networks from massive very-high-resolution (VHR) remote sensing imagery, existing road extraction methods are limited by the problem: lacking of an accurate and diverse training dataset for global-scale road extraction, and manually labelling millions of road samples for training a global model is labor intensive. To address this problem, we utilized VHR satellite imagery and open-source crowdsourcing geospatial big data to build a robust global-scale road training dataset, termed GlobalRoadNet, for global road extraction (GRE) and beyond. The proposed GlobalRoadNet contains 47210 samples from 121 capital cities of six continents in Europe, Africa, Asia, South America, Oceania, and North America. Experimental results show that GlobalRoadNet can significantly improve model performance, not only can be applied for road extraction, but also has the potential to update OSM road data. Yanfei Zhong, Zhuo Zheng |
IGARSS | 3 |
| 2022 | A Spectral-Spatial-Dependent Global Learning Framework for Insufficient and Imbalanced Hyperspectral Image ClassificationabstractDeep learning techniques have been widely applied to hyperspectral image (HSI) classification and have achieved great success. However, the deep neural network model has a large parameter space and requires a large number of labeled data. Deep learning methods for HSI classification usually follow a patchwise learning framework. Recently, a fast patch-free global learning (FPGA) architecture was proposed for HSI classification according to global spatial context information. However, FPGA has difficulty in extracting the most discriminative features when the sample data are imbalanced. In this article, a spectral-spatial-dependent global learning (SSDGL) framework based on the global convolutional long short-term memory (GCL) and global joint attention mechanism (GJAM) is proposed for insufficient and imbalanced HSI classification. In SSDGL, the hierarchically balanced (H-B) sampling strategy and the weighted softmax loss are proposed to address the imbalanced sample problem. To effectively distinguish similar spectral characteristics of land cover types, the GCL module is introduced to extract the long short-term dependency of spectral features. To learn the most discriminative feature representations, the GJAM module is proposed to extract attention areas. The experimental results obtained with three public HSI datasets show that the SSDGL has powerful performance in insufficient and imbalanced sample problems and is superior to other state-of-the-art methods. Qiqi Zhu, Weihuan Deng, Zhuo Zheng, Yanfei Zhong, Qingfeng Guan 0001, Weihua Lin, Liangpei Zhang 0001, DeRen Li |
IEEE Trans. Cybern. | 3 |
| 2022 | Cascaded Multi-Task Road Extraction Network for Road Surface, Centerline, and Edge ExtractionabstractRoad extraction from very high-resolution (VHR) remote sensing imagery remains a huge challenge, due to the shadows and occlusions of trees and buildings. Such complex backgrounds result in deep networks often producing fragmented roads with poor connectivity. Road extraction has three typical tasks: road surface segmentation (SS), centerline extraction (CE), and edge detection (ED), which are conducted in a wide range of real applications. Also, the three tasks have a symbiotic relationship, i.e., the road SS determines the location of the centerline and edges, and the CE and ED can allow the generation of more continuous road surfaces. However, most of the previous works have completed these three tasks separately, without exploiting the symbiotic relationship between them to boost the road connectivity. In this article, in order to improve road connectivity, a cascaded multitask (CasMT) road extraction framework for simultaneously extracting the road surface, centerline, and edges is proposed. In the proposed framework, topology-aware learning is applied to capture the long-distance topological relationships, and hard example mining (HEM) loss is employed to focus more on hard samples, to further enhance the road completeness. Extensive experiments were conducted on the DeepGlobe road dataset and a large-scale road dataset (called the LSCC dataset) from the three Chinese cities of Beijing, Shanghai, and Wuhan. The experimental results obtained on the public DeepGlobe dataset demonstrate that the proposed CasMT framework can significantly outperform the current state-of-the-art method. Moreover, the generalization capability of the model was verified on the LSCC dataset, where the proposed CasMT framework achieved the best performance in the average path length similarity (APLS) road topology metric, which further confirms the superiority of the proposed framework. Yanfei Zhong, Zhuo Zheng, Ailong Ma, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | FactSeg: Foreground Activation-Driven Small Object Semantic Segmentation in Large-Scale Remote Sensing ImageryabstractThe small object semantic segmentation task is aimed at automatically extracting key objects from high-resolution remote sensing (HRS) imagery. Compared with the large-scale coverage areas for remote sensing imagery, the key objects, such as cars and ships, in HRS imagery often contain only a few pixels. In this article, to tackle this problem, the foreground activation (FA)-driven small object semantic segmentation (FactSeg) framework is proposed from perspectives of structure and optimization. In the structure design, FA object representation is proposed to enhance the awareness of the weak features in small objects. The FA object representation framework is made up of a dual-branch decoder and collaborative probability (CP) loss. In the dual-branch decoder, the FA branch is designed to activate the small object features (activation) and suppress the large-scale background, and the semantic refinement (SR) branch is designed to further distinguish small objects (refinement). The CP loss is proposed to effectively combine the activation and refinement outputs of the decoder under the CP hypothesis. During the collaboration, the weak features of the small objects are enhanced with the activation output, and the refined output can be viewed as the refinement of the binary outputs. In the optimization stage, small object mining (SOM)-based network optimization is applied to automatically select effective samples and refine the direction of the optimization while addressing the imbalanced sample problem between the small objects and the large-scale background. The experimental results obtained with two benchmark HRS imagery segmentation datasets demonstrate that the proposed framework outperforms the state-of-the-art semantic segmentation methods and achieves a good tradeoff between accuracy and efficiency. Code will be available at:http://rsidea.whu.edu.cn/FactSeg.htm Ailong Ma, Yanfei Zhong, Zhuo Zheng |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | A Supervised Progressive Growing Generative Adversarial Network for Remote Sensing Image Scene ClassificationabstractRemote sensing image scene classification is a challenging task. With the development of deep learning, methods based on convolutional neural networks (CNNs) have made great achievements in remote sensing image scene classification. Since the training of a CNN requires a large number of labeled samples, a generative adversarial network (GAN) for sample generation represents a new opportunity to solve the problem of the limited samples. However, most of the existing GAN-based sample generation methods can only generate unlabeled samples, instead of samples labeled with the corresponding scene category. In this article, to solve the problem, a supervised progressive growing generative adversarial network (SPG-GAN) is proposed for remote sensing image scene classification. The proposed method can generate labeled samples for the remote sensing image scene classification, significantly improving the classification accuracy in the case of limited samples. The SPG-GAN method has two main improvements. First, a conditional generative framework for labeled samples is proposed, in which the label information is added in the channel dimension as the input. By considering the constraints of the label information in the loss function, the network can be trained in the direction of a specific category. As a result, the network can generate remote sensing image scene classification samples with label categories. Second, a progressive growing sample generation method is introduced. In order to ensure that the generated samples have more spatial details, they are generated by progressively adding modules to the generator and discriminator, thereby ensuring that the generated sample is of better quality. After testing on two benchmark datasets and carrying out a large-scale experiment in the central area of the city of Wuhan in China, it was found that the proposed method can obtain a superior scene classification accuracy in the case of limited samples. Ailong Ma, Zhuo Zheng, Yanfei Zhong, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Change is Everywhere: Single-Temporal Supervised Object Change Detection in Remote Sensing ImageryabstractFor high spatial resolution (HSR) remote sensing images, bitemporal supervised learning always dominates change detection using many pairwise labeled bitemporal images. However, it is very expensive and time-consuming to pairwise label large-scale bitemporal HSR remote sensing images. In this paper, we propose single-temporal supervised learning (STAR) for change detection from a new perspective of exploiting object changes in unpaired images as supervisory signals. STAR enables us to train a high-accuracy change detector only using unpaired labeled images and generalize to real-world bitemporal images. To evaluate the effectiveness of STAR, we design a simple yet effective change detector called ChangeStar, which can reuse any deep semantic segmentation architecture by the ChangeMixin module. The comprehensive experimental results show that ChangeStar outperforms the baseline with a large margin under single-temporal super-vision and achieves superior performance under bitemporal supervision. Code is available at https://github.com/Z-Zheng/ChangeStar. Zhuo Zheng, Ailong Ma, Liangpei Zhang 0001, Yanfei Zhong |
ICCV | 1 |
| 2021 | Sensor-Specific Adversarial Network for Transferable Land-Cover ClassificationabstractAs the multi-source high-spatial-resolution (HSR) images are being daily acquired from different sensors, it brings the challenge of transferring the recognition model from labeled images to new unlabelled images obtained from other sensors. Existing deep transfer learning methods encode the land-cover features in the same architecture, which ignores the sensor divergence. In this paper, we tackle this problem by proposing a sensor-specific adversarial network for HSR land-cover classification. Specifically, the sensor-specific normalization (SN) is designed for decoupling the sensor divergence in different normalization weights. Moreover, the transferable adversarial optimization is proposed for effectively optimizing the source-related, target-related, and discriminator weights. Considering the sensor-specific characteristics, our proposed method improves the transferability of deep learning models between airborne and spaceborne sensors. The mutual transferability experiments on a self-constructed cross-sensor land-cover dataset demonstrate that the proposed method outperforms the state-of-the-art deep transfer learning methods. Yanfei Zhong, Zhuo Zheng, Ailong Ma |
IGARSS | 3 |
| 2021 | Weakly Supervised Semantic Change Detection via Label Refinement FrameworkabstractSemantic change detection is a meaningful but challenging task in the remote sensing community. The currently dominant approaches are mainly based on deep learning. However, the lack of high-resolution annotations is the main bottleneck for semantic change detection at scale when using these state-of-the-art deep learning models. In this paper, the label refinement framework is proposed for weakly-supervised semantic change detection, which allows the deep network to learn from low-resolution labels and produce high-resolution semantic change maps, thus alleviating the data-hungry problem. This framework contains four parts: coarse label training’ pseudo-label refinement, multitask change detection and post-process. The experimental results on 2021 IEEE GRSS Data Fusion Contest Track MSD dataset confirmed the effectiveness of the proposed method. Additionally, our method wins 4th place in the 2021 IEEE GRSS Data Fusion Contest Track MSD (DFC21-MSD). Zhuo Zheng, Yinhe Liu, Shiqi Tian, Ailong Ma, Yanfei Zhong |
IGARSS | 1 |
| 2021 | RSNet: The Search for Remote Sensing Deep Neural Networks in Recognition TasksabstractDeep learning algorithms, especially convolutional neural networks (CNNs), have recently emerged as a dominant paradigm for high spatial resolution remote sensing (HRS) image recognition. A large amount of CNNs have already been successfully applied to various HRS recognition tasks, such as land-cover classification and scene classification. However, they are often modifications of the existing CNNs derived from natural image processing, in which the network architecture is inherited without consideration of the complexity and specificity of HRS images. In this article, the remote sensing deep neural network (RSNet) framework is proposed using an automatically search strategy to find the appropriate network architecture for HRS image recognition tasks. In RSNet, the hierarchical search space is first designed to include module- and transition-level spaces. The module-level space defines the basic structure block, where a series of lightweight operations as candidates, including depthwise separable convolutions, is proposed to ensure the efficiency. The transition-level space controls the spatial resolution transformations of the features. In the hierarchical search space, a gradient-based search strategy is used to find the appropriate architecture. In RSNet, the task-driven architecture training process can acquire the optimal model parameters of the switchable recognition module for HRS image recognition tasks. The experimental results obtained using four benchmark data sets for land-cover classification and scene classification tasks demonstrate that the searched RSNet can achieve a satisfactory accuracy with a high computational efficiency and, hence, provides an effective option for the processing of HRS imagery. Yanfei Zhong, Zhuo Zheng, Ailong Ma, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Foreground-Aware Relation Network for Geospatial Object Segmentation in High Spatial Resolution Remote Sensing ImageryabstractGeospatial object segmentation, as a particular semantic segmentation task, always faces with larger-scale variation, larger intra-class variance of background, and foreground-background imbalance in the high spatial resolution (HSR) remote sensing imagery. However, general semantic segmentation methods mainly focus on scale variation in the natural scene, with inadequate consideration of the other two problems that usually happen in the large area earth observation scene. In this paper, we argue that the problems lie on the lack of foreground modeling and propose a foreground-aware relation network (FarSeg) from the perspectives of relation-based and optimization-based foreground modeling, to alleviate the above two problems. From perspective of relation, FarSeg enhances the discrimination of foreground features via foreground-correlated contexts associated by learning foreground-scene relation. Meanwhile, from perspective of optimization, a foreground-aware optimization is proposed to focus on foreground examples and hard examples of background during training for a balanced optimization. The experimental results obtained using a large scale dataset suggest that the proposed method is superior to the state-of-the-art general semantic segmentation methods and achieves a better trade-off between speed and accuracy. Zhuo Zheng, Yanfei Zhong, Ailong Ma |
CVPR | 1 |
| 2020 | A Novel Global-Aware Deep Network for Road Detection of Very High Resolution Remote Sensing ImageryabstractRoad detection from very high-resolution (VHR) remote sensing imagery has great importance in a broad array of applications. However, the most advanced deep learning-based methods often produce fragmented road segments, due to the complex backgrounds of images, such as the occlusions and shadows caused by the trees and buildings, or the surrounding objects with similar textures. In this paper, the characteristics of existing models were analyzed and an effective road recognition method was explored, we found that capturing long-range dependencies helps improve road recognition. Therefore, a novel global-aware deep network (GAN) for road detection is proposed, in which the spatial-aware module (SAM) was applied to capture spatial context dependencies and the channel-aware module (CAM) was applied to capture the interchannel dependencies. Through establishing the relationships between spatial contexts and between channels, the GAN could effectively alleviate the road recognition problems, and the advantages of the proposed approach were validated on the public DeepGlobe road dataset. The experimental result demonstrates the superiority of our method. Yanfei Zhong, Zhuo Zheng |
IGARSS | 3 |
| 2020 | FPGA: Fast Patch-Free Global Learning Framework for Fully End-to-End Hyperspectral Image ClassificationabstractDeep learning techniques have provided significant improvements in hyperspectral image (HSI) classification. The current deep learning-based HSI classifiers follow a patch-based learning framework by dividing the image into overlapping patches. As such, these methods are local learning methods, which have a high computational cost. In this article, a fast patch-free global learning (FPGA) framework is proposed for HSI classification. The proposed framework consists of three main parts: 1) a designed sampling strategy; 2) an encoder-decoder-based fully convolutional network (FCN); and 3) lateral connections between the encoder and decoder. In FPGA, an encoder-decoder-based FCN is utilized to consider the global spatial information by processing the whole image, which results in fast inference. However, it is difficult to directly utilize the encoder-decoder-based FCN for HSI classification as it always fails to converge due to the insufficiently diverse gradients caused by the limited training samples. To solve the divergence problem and maintain the FCNs abilities of fast inference and global spatial information mining, a global stochastic stratified (GS2) sampling strategy is first proposed by transforming all the training samples into a stochastic sequence of stratified samples. This strategy can obtain diverse gradients to guarantee the convergence of the FCN in the FPGA framework. For a better design of FCN architecture, FreeNet, which is a fully end-to-end network for HSI classification, is proposed to maximize the exploitation of the global spatial information and boost the performance via a spectral attention-based encoder and a lightweight decoder. A lateral connection module is also designed to connect the encoder and decoder, fusing the spatial details in the encoder and the semantic features in the decoder. The experimental results obtained using three public benchmark data sets suggest that the FPGA framework is superior to the patch-based framework in both speed and accuracy for HSI classification. Zhuo Zheng, Yanfei Zhong, Ailong Ma, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | COLOR: Cycling, Offline Learning, and Online Representation Framework for Airport and Airplane Detection Using GF-2 Satellite ImagesabstractMonitoring airports using remote sensing imagery require us to first detect the airports and then perform airplane detection. Detecting airports and airplanes with large-scale remote sensing imagery are significant and challenging tasks in the field of remote sensing. Although many detection algorithms have been developed for detecting airports and airplanes in remote sensing imagery, the efficiency of the processing does not meet the needs of real applications in large-scale remote sensing imagery. In recent years, deep learning techniques, such as deep convolutional neural networks (DCNNs), have achieved great progress in image recognition. However, training a DCNN needs a large number of training examples to accurately fit the data distribution. Annotating training examples in large-scale remote sensing imagery is time-consuming, which makes the pipeline inefficient. In this article, to overcome the above two weaknesses, we propose a novel cycling data-driven framework for efficient and robust airport localization and airplane detection. The proposed method consists of three modules: cycling by example refinement (C), offline learning (OL), and online representation (OR), namely cycling, offline learning, and online representation (COLOR). The OR module is a coarse-to-fine cascaded convolutional neural network, which is used to detect airports and airplanes. The example refinement (ER) module implements the cycling and makes use of the unlabeled remote sensing images and the corresponding predictions obtained by the OR module, to generate training examples. The OL module aims to use the training examples from the ER module to update the OR module, to further improve the performance. The whole workflow involves COLOR. The COLOR framework was used to detect airplanes and airports in 512 large-scale Gaofen-2 (GF-2) remote sensing images with 29$200\times27$ 620 pixels. The results showed that the proposed method obtained a mean average precision (mAP) of 88.32% for the airplane detection. In addition due to the proposed coarse-to-fine cascaded OR module the proposed method is much faster than the traditional approaches in real-world applications. Yanfei Zhong, Zhuo Zheng, Ailong Ma, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | S3NET: Towards Real-Time Hyperspectral Imagery ClassificationabstractFast hyperspectral image classification methods are required to real time processing on UAV and airborne platform. But current state of the art methods are under patch based framework that is slow due to duplicate computation, while fully convolutional neural network (FCN) is difficult to run inference in one shot due to patch based training strategy. In this paper, we firstly identify overparameterization and inconsistent class ratio per mini-batch as the central causes impeding training of FCN based methods so we propose a novel framework based on a new lightweight network and a stratified sample based training strategy to overcome above problems. To refine redundant spectrum information for fast computation, the spectral attention module is proposed as a soft band selection. The proposed training strategy enables the stability of training by keeping class ratio consistent in mini-batches. On the WHU-HongHu UAV dataset, our method achieves great performance by obtaining a OA of 98.94 and running at 144 fps. Additionally, our method can obtain a flexible trade-off between speed and accuracy. Zhuo Zheng, Yanfei Zhong |
IGARSS | 1 |
| 2019 | Pop-Net: Encoder-Dual Decoder for Semantic Segmentation and Single-View Height EstimationabstractThe single-view semantic 3D challenge in 2019 Data Fusion Contest is to predict both semantic labels and normalized digital surface model (nDSM) for urban scenes from single-view satellite images. We propose a novel pyramid on pyramid network (Pop-Net) based on Encoder-Dual Decoder framework to end-to-end multi-task learning. The encoder is a deformable ResNet-101 backbone network. Two feature pyramid networks, as decoders, are responsible for semantic segmentation and height estimation, respectively. Semantic information is crucial to estimate height. Therefore, regression pyramid on the semantic pyramid is introduced to leverage semantic features to help height estimation. To deal with outliers in heights, we leverage anchor-based regression and smooth L1 loss for optimization to obtain more robust height estimation. Without bells and whistles, our single model entry achieves 77.78% mIoU and 53.40% mIoU-3 on test set, ranking 2nd in the Single-view Semantic 3D Challenge of the 2019 IEEE GRSS Data Fusion Contest. The code is available at https://github.com/Z-Zheng/PopNet. Zhuo Zheng, Yanfei Zhong |
IGARSS | 1 |
| 2019 | Multi-Scale and Multi-Task Deep Learning Framework for Automatic Road ExtractionabstractRoad detection and centerline extraction from very high-resolution (VHR) remote sensing imagery are of great significance in various practical applications. Road detection and centerline extraction operations depend on each other, to a certain extent. The road detection constrains the appearance of the centerline, and the centerline enhances the linear features of the road detection. However, most of the previous works have addressed these two tasks separately and have not considered the symbiotic relationship between them, making it difficult to obtain smooth and complete roads. In this paper, a novel multi-scale and multi-task deep learning framework for automatic road extraction (MSMT-RE) is proposed to build the relationship between them and simultaneously complete the road detection and centerline extraction tasks. U-Net is selected as the basic network for multi-task learning due to its strong ability to preserve spatial details. Multi-scale feature integration is also applied in the framework to increase the robustness of the feature extraction. Meanwhile, an adaptive loss function is introduced to solve the problems of roads taking up a small percentage of the training samples, and the fact that the positive samples of the two tasks are unbalanced. Finally, experiments were conducted on two public road data sets and two large images from Google Earth, and the proposed framework was compared with other state-of-the-art deep learning-based road extraction methods, both quantitatively and qualitatively. The proposed approach outperformed all the compared methods, confirming its advantages in automatic road extraction. Yanfei Zhong, Zhuo Zheng, Ji Zhao 0006, Ailong Ma, Jie Yang 0040 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2018 | Multi-Channel Pose-Aware Convolution Neural Networks for Multi-View Facial Expression RecognitionabstractAlthough tremendous strides have been made in facial expression recognition(FER), recognizing facial expressions in non-frontal views remains an open challenge due to the limited access to large scale training data with various poses. To make full use of the limited data, we propose a novel multi-channel pose-aware convolution neural network (MPCNN) that consists of three parts: the multi-channel feature extraction, jointly multi-scale feature fusion, and the pose-aware recognition. The feature extraction part has 3 sub-CNNs and it learns convolutional features from different features. The joint fusion part fuses multi-scale features to enhance high-level feature representation in a hierarchical way. The fused features are fed to the pose-aware recognition part that includes pose-specific recognition branches and a pose estimation sub-network. According to the estimated pose, MPCNN finally classifies the facial expression through a conditional weighted combination of the pose-specific recognition branches. MPCNN is end-to-end trainable by minimizing the joint loss of pose and expression recognition. We evaluated the proposed method on two public multi-view FER datasets (BU-3DFE and KDEF) and a FER dataset in the wild (SFEW). The experimental results demonstrate that MPCNN outperforms the state-of-the-art FER methods with both within-dataset and cross-dataset settings. Yuanyuan Liu 0004, Jiabei Zeng, Shiguang Shan, Zhuo Zheng |
FG | 4 |
| 2018 | Color: Cycling Offline Learning and Online Representing for Remote Sensing DataflowabstractIn recent years, many model driven frameworks have achieved great success in remote sensing. For large scale remote sensing dataflow, however, its ability of processing is far from enough. In this paper, we propose a novel cycling data driven framework based on deep learning, which consists of three components: offline learning, online representing and sample refining. That online representing makes predictions on unlabeled remote sensing images and uses them to retrain the model by offline learning while sample refining as a router refines coarse predictions and forwards refined predictions to offline learning. Because the process is like circling between offline learning and online representing, we refer to our framework as circling offline learning and online representing (COLOR). We show the excellent performance and potentiality on airplane object detection by applying a fairly straightforward implement of COLOR using our GF2 airplane object detection dataset. Code and dataset will be made available. Zhuo Zheng, Yanfei Zhong |
IGARSS | 1 |