VLDB 2026 Research / reviewers in the wild / expert
Zhaocong Wu
dblp:38/1798
· DBLP profile ↗
11ranked-venue papers
3as first author
5since 2021 · last 2023
0000-0003-2435-5538ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Crisscross-Global Vision Transformers Model for Very High Resolution Aerial Image Semantic SegmentationabstractSemantic segmentation is a key means for understanding very-high resolution (VHR) aerial imagery. With the explosive development of deep learning, deep learning methods are being applied to the segmentation of VHR images, with convolutional neural networks (CNNs) as the basic framework. However, owing to the highly complex details present in VHR images and the high spatial dependence of geographical objects, CNN-based methods are inadequate. This is because the inherent locality of CNNs limits the size of the receptive field, thus limiting the ability to obtain long-range context information. To solve this problem, in this paper, we propose a transformer-based novel deep learning model called crisscross-global vision transformers (CGVT). CGVT exploits the transformer’s inherent ability to obtain long-range context information to solve the restricted receptive field problem. Specifically, we redesign the self-attention mechanism in the transformer and call it crisscross-global attention. It consists of two parts: crisscross transformer encoder block (CC-TEB) and global squeeze transformer encoder block (GS-TEB). CC-TEB overcomes the limitation of the traditional self-attention design (specifically, difficulty applying it to VHR aerial image segmentation) and further increases the local feature representation ability of the model. GS-TEB increases the global feature representation ability of the model. The results of experiments conducted on the popular ISPRS Vaihingen, IEEE GRSS Data Fusion Contest Zeebrugge, and LoveDA Semantic Segmentation Challenge datasets verify the effectiveness and superiority of our proposed method. Specifically, it achieved state-of-the-art performance on both Zeebrugge and LoveDA datasets, and is currently ranked second in Vaihingen dataset. Guohui Deng, Zhaocong Wu, Miaozhong Xu, Zhiye Wang, Zhongyuan Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Mapping Forest Canopy Height at Large Scales Using ICESat-2 and Landsat: An Ecological Zoning Random Forest ApproachabstractForest canopy height (FCH) is a crucial indicator in the calculation of forest biomass and carbon sinks. There are various methods to measure FCH, such as space-borne light detection and ranging (LiDAR), but their data are spatially discrete and do not provide continuous FCH maps. Therefore, an FCH estimation method that associates sparse LiDAR data with spatially continuous variables is required. The traditional approach of constructing a single model overlooks the spatial variability in forest growth, which will limit the FCH accuracy. Considering the distinct nature of forest in different ecological zones, the following hold. First, we proposed an ecological zoning random forest (EZRF) model in 33 ecological zones in China. Compared with the total zone RF (TZRF) model, the EZRF model showed a greater potential, which was 21.5%–36.5% more accurate than the TZRF model. Second, we analyzed a total of 62 variables related to forest growth, including Landsat variables and ancillary variables (forest canopy cover, bioclimatic, topographic, and hillshade factors). An insight into variable selection in FCH modeling was provided by analyzing the prediction accuracy of FCH under different categorical variables and analyzing the importance of variables in different ecological zones. Third, finally, we produced a 30-m continuous FCH map by the EZRF model. Compared with the airborne LiDAR data, the FCH prediction results produced a root mean square error (RMSE) of 2.50–5.35 m, which were 41%–72% more precisely than the Global FCH, 2019. The results demonstrate the effectiveness of our proposed method and contribute to the study of forest carbon sinks. Zhaocong Wu, Fanglin Shi |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Selection of Optimal Bands for Hyperspectral Local Feature DescriptorabstractThe use of hyperspectral images (HSIs) and 3-D data has proven to be an efficient combination for numerous applications. A common method of obtaining corresponding data is 3-D reconstruction from HSI, for which the local feature descriptor is vital to the final accuracy. However, redundant bands may hamper the performance of the descriptor, which is similar to the so-called Hughes phenomenon. Band selection (BS) is effective in overcoming such problems. Existing BS methods fail to select the band in a differentiable manner and, thus, cannot be jointly optimized with downstream tasks. In this letter, we propose a novel end-to-end HSI local feature descriptor network (with joint optimal BS) called HyperDesc. It implements a true band selection (TBS) module by turning BS into a CONCRETE random variable-based differentiable sampling operation. The discrete distribution of each selected band can be learned from input HSI with a nonlocal spectral–spatial attention network. Finally, the selected band is sent to a descriptor network to extract the local feature descriptor. The whole network is trained in an end-to-end manner. Experiments are conducted on multiview close-range HSI. The results show that spectral information provided by selected bands can boost the performance of the descriptor. Zhaocong Wu |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | CCANet: Class-Constraint Coarse-to-Fine Attentional Deep Network for Subdecimeter Aerial Image Semantic SegmentationabstractSemantic segmentation is important for the understanding of subdecimeter aerial images. In recent years, deep convolutional neural networks (DCNNs) have been used widely for semantic segmentation in the field of remote sensing. However, because of the highly complex subdecimeter resolution of aerial images, inseparability often occurs among some geographic entities of interest in the spectral domain. In addition, the semantic segmentation methods based on DCNNs mostly obtain context information using extra information within the added receptive field. However, the context information obtained this way is not explicit. We propose a novel class-constraint coarse-to-fine attentional (CCA) deep network, which enables the formation of class information constraints to obtain explicit long-range context information. Further, the performance of subdecimeter aerial image semantic segmentation can be improved, particularly for fine-structured geographic entities. Based on coarse-to-fine technology, we obtained a coarse segmentation result and constructed an image class feature library. We propose the use of the attention mechanism to obtain strong class-constrained features. Consequently, pixels of different geographic entities can adaptively match the corresponding categories in the class feature library. Additionally, we employed a novel loss function, CCA-loss to realize end-to-end training. The experimental results obtained using two popular open benchmarks, International Society for Photogrammetry and Remote Sensing (ISPRS) 2-D semantic labeling Vaihingen data set and Institute of Electrical and Electronics Engineers (IEEE) Geoscience and Remote Sensing Society (GRSS) Data Fusion Contest Zeebrugge data set, validated the effectiveness and superiority of our proposed model. The proposed method achieved state-of-the-art performance on the IEEE GRSS Data Fusion Contest Zeebrugge data set. Guohui Deng, Zhaocong Wu, Miaozhong Xu, Yanfei Zhong |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A Lightweight Deep Learning-Based Cloud Detection Method for Sentinel-2A Imagery Fusing Multiscale Spectral and Spatial FeaturesabstractClouds are a very important factor in the availability of optical remote sensing images. Recently, deep learning (DL)-based cloud detection methods have surpassed classical methods based on rules and physical models of clouds. However, most of these deep models are very large, which limits their applicability and explainability, while other models do not make use of the full spectral information in multispectral images, such as Sentinel-2. In this article, we propose a lightweight network for cloud detection, fusing multiscale spectral and spatial features (CD-FM3SFs) and tailored for processing all spectral bands in Sentinel-2A images. The proposed method consists of an encoder and a decoder. In the encoder, three input branches are designed to handle spectral bands at their native resolution and extract multiscale spectral features. Three novel components are designed: a mixed depthwise separable convolution (MDSC) and a shared and dilated residual block (SDRB) to extract multiscale spatial features, and a concatenation and sum (CS) operation to fuse multiscale spectral and spatial features with little calculation and no additional parameters. The decoder of CD-FM3SF outputs three cloud masks at the same resolution as input bands to enhance the supervision information of small, middle, and large clouds. To validate the performance of the proposed method, we manually labeled 36 Sentinel-2A scenes evenly distributed over mainland China. The experiment results demonstrate that CD-FM3SF outperforms traditional cloud detection methods and state-of-the-art DL-based methods in both accuracy and speed. Jun Li 0087, Zhaocong Wu, Zhongwen Hu, Canliang Jian, Shaojie Luo, Lichao Mou, Xiao Xiang Zhu 0001, Matthieu Molinier |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Self-Attentive Generative Adversarial Network for Cloud Detection in High Resolution Remote Sensing ImagesabstractCloud detection is an important step in the processing of remote sensing images. Most methods based on convolutional neural networks (CNNs) for cloud detection require pixel-level labels, which are time-consuming and expensive to annotate. To overcome this challenge, this letter proposes a novel semisupervised algorithm for cloud detection by training a self-attentive generative adversarial network (SAGAN) to extract the feature difference between cloud images and cloud-free images. Our main idea is to introduce visual attention into the process of generating “real” cloud-free images. The training of SAGAN is based on three guiding principles: expansion of attention maps of cloud regions which will be replaced with translated cloud-free images, reduction of attention maps to coincide with cloud boundaries, and optimization of a self-attentive network to handle the extreme cases. The inputs for SAGAN training are the images and image-level labels, which are easier, cheaper, and more time-saving than the existing methods based on CNN. To test the performance of SAGAN, experiments are conducted on the Sentinel-2A Level 1C image data. The results show that the proposed method achieves very promising results with only the image-level labels of training samples. Zhaocong Wu, Jun Li 0087, Zhongwen Hu, Matthieu Molinier |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2017 | Unsupervised Simplification of Image Hierarchies via Evolution Analysis in Scale-Sets FrameworkabstractRegion-based hierarchical image representation is crucial in many computer vision applications. However, in practice, an image hierarchy is usually dense, and contains many less informative branches. It is expected that a hierarchy should be accurate and simplified, which is not only desirable for different applications, but also saves considerable computational load for the further analysis. To achieve this target, this paper proposes a novel approach for unsupervised simplification of region-based image hierarchies, which employs the global and local evolution analyses of a hierarchy. First, we introduce a global evolution analysis in the scale-sets framework, which provides clues for eliminating less informative branches. Moreover, a hybrid unsupervised simplification method is designed, utilizing the information from global and local evolution functions. A number of experiments on various images have shown that the proposed approach is effective and efficient in removing less informative nodes (averagely about 90% of the whole nodes), while preserving salient image details and retaining the accuracy. Zhongwen Hu, Qingquan Li 0001, Qian Zhang 0046, Qin Zou 0001, Zhaocong Wu |
IEEE Trans. Image Process. | 5 |
| 2013 | A Spatially-Constrained Color-Texture Model for Hierarchical VHR Image SegmentationabstractThis letter presents a novel spatially-constrained color–texture model for hierarchical segmentation of very high resolution images. The segmentation starts with an initial partition, where the image is partitioned into many homogeneous regions. Then, the regions are regarded as node sets of a region adjacency graph, in which the distances of each pair of adjacent regions are calculated combining color and textural features with spatial constraint. Finally, a stepwise optimized region merging process is applied to obtain hierarchical segmentation results. Experiments and comparisons by using different satellite images are carried out to demonstrate the encouraging performance as well as the high efficiency of the proposed method. Zhongwen Hu, Zhaocong Wu, Qian Zhang 0046 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2012 | Gross primary production estimation by combining MODIS products and Ameriflux data through Artificial Neural Network for croplandsabstractVegetation productivity is the basis of all the biosphere activities on the land surface that relate to global biogeochemical cycles of carbon and nitrogen. The accurate quantification of gross primary production (GPP) in crops is important for regional and global studies of carbon budgets. Many flux observation nets have been established to help us monitoring the carbon cycling. However, estimation of GPP of terrestrial ecosystems for regions, continents, or the globe can improve our understanding of the feedbacks between the terrestrial biosphere and the atmosphere in the context of global change and facilitate climate policymaking. Remote sensing is a potentially powerful technology with which to extrapolate eddy covariance-based GPP to continental scales. In this paper, we combined MODIS products and Ameriflux networks data to simulate and predict GPP at four different cropland sites, using the Artificial Neural Networks (ANN). The results were quite approving compared to MODIS GPP product and tower-based measurements, which indicated it could be an applicable approach for GPP estimation. Zhaocong Wu, Wanshou Jiang |
IGARSS | 2 |
| 2012 | A Scale-Synthesis Method for High Spatial Resolution Remote Sensing Image SegmentationabstractMultiscale segmentation is always needed to extract semantic meaningful objects for object-based remote sensing image analysis. Choosing the appropriate segmentation scales for distinct ground objects and intelligently combining them together are two crucial issues to get the appropriate segmentation result for target applications. With respect to these two issues, this paper proposes a simple scale-synthesis method which is highly flexible to be adjusted to meet the segmentation requirements of varying image-analysis tasks. The main idea of this method is to first divide the whole image area into multiple regions; each region consisted of ground objects that have similar optimal segmentation scale. Then, synthesize the suboptimal segmentations of each region to get the final segmentation result. The result is the combination of suboptimal scales of objects and is therefore more coherent to ground objects. To validate this method, the land-cover-category map is used to guide the scale synthesis of multiscale image segmentations for the Quickbird-image land-use classification. First, the image is coarsely divided into multiple regions; each region belongs to a certain land-cover category. Then, multiscale-segmentation results are generated by the Mumford-Shah function based region-merging method. For each land-cover category, the optimal segmentation scale is selected by the supervised segmentation-accuracy-assessment method. Finally, the optimal scales of segmentation results are synthesized under the guide of land-cover category. It is proved that the proposed scale-synthesis method can generate a more accurate segmentation result that benefits the latter classification. The land-use-classification accuracy reaches to 77.8%. Lina Yi, Guifeng Zhang, Zhaocong Wu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2010 | An Edge Embedded Marker-Based Watershed Algorithm for High Spatial Resolution Remote Sensing Image SegmentationabstractThis correspondence proposes an edge embedded marker-based watershed algorithm for high spatial resolution remote sensing image segmentation. Two improvement techniques are proposed for the two key steps of maker extraction and pixel labeling, respectively, to make it more effective and efficient for high spatial resolution image segmentation. Moreover, the edge information, detected by the edge detector embedded with confidence, is used to direct the two key steps for detecting objects with weak boundary and improving the positional accuracy of the objects boundary. Experiments on different images show that the proposed method has a good generality in producing good segmentation results. It performs well both in retaining the weak boundary and reducing the undesired over-segmentation. DeRen Li, Guifeng Zhang, Zhaocong Wu, Lina Yi |
IEEE Trans. Image Process. | 3 |