EDBT 2026 Demo / reviewers in the wild / expert
Linlin Zhang 0007
dblp:68/1772-7
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0001-5073-1694ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Instance-Level Multitask Learning for 3-D Building Extraction From Monocular Off-Nadir Satellite Sensor ImageryabstractExtracting 3D building information from monocular satellite sensor imagery remains a formidable challenge in the field of remote sensing. Multitask frameworks based on deep learning, which particularly for simultaneously predicting 2D building outlines and their respective heights using ortho-rectified satellite imagery, have shown promise in addressing this challenge. Height estimation is notably complex due to the absence of explicit height indicators, limited interaction between semantic-height features, and inadequate representation of building relationships. Moreover, the availability of data sources is a limiting factor for broader application. To overcome these issues, this study introduces an innovative instance-level multitask learning model (named BDH-Net) that leverages off-nadir perspectives and roof-to-footprint offset vectors to enhance modeling. This model comprises four key components: a pixel-wise feature extraction image encoder-decoder, a query transformer decoder, a multitask decoder, and a height decoder that employs intra-instance and inter-instance attention for precise building height estimation. Additionally, we pioneer the use of Google Earth imagery to construct an off-nadir satellite dataset with roof-to-footprint offset vectors specifically designed for building instance segmentation and height prediction, known as the BDH dataset. Comprehensive experiments demonstrate that the proposed BDH-Net significantly improves the accuracy of monocular 3D building data extraction by integrating roof-to-footprint offset vectors and leveraging context specific to each building instance. With the extensive coverage and regular updates of Google Earth imagery, BDH-Net holds substantial potential for wide-ranging and long-term applications. The source code of the proposed BDH-Net and the BDH dataset are publicly available at https://github.com/wishx98/BDHNet. Wenxu Shi, Qingyan Meng, Linlin Zhang 0007, Maofan Zhao, Guinan Guo, Peter M. Atkinson |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | HR-UVFormer: A Top-Down and Multimodal Hierarchical Extraction Approach for Urban VillagesabstractUrban Villages (UVs) renovation has been incorporated into the Sustainable Development Goals (SDGs) as a result of the inequality issue among residents garnering substantial social attention. However, existing deep-learning techniques for UVs extraction have been limited to a single spatial scale (e.g., patch-level or pixel-level extraction), leading to inadequate precision and integrity in their extraction outcomes. To overcome this limitation, our study introduces HR-UVFormer, a top-down and multimodal hierarchical extraction approach that extracts UVs from a coarse scale (patch) to a fine granularity (pixel), aiming to enhance the internal completeness and boundary accuracy of the extraction results. The multimodal approach can effectively fuse multimodal features (e.g., building footprints (BF)) with remote sensing images (RSI) to enhance UVs extraction. The Shenzhen results indicate that the coarse-scale extraction accuracy achieves an overall accuracy (OA) of 98.79%, and the fine-grained extraction accuracy achieves a mean Intersection over Union (mIoU) of 93.60%. Furthermore, ablation experiments demonstrate a notable 7.14% improvement in mIoU with the hierarchical extraction strategy compared to the traditional pixel-based extraction strategy, and the fusion of BF and RSI yields further improvements of 2.78% and 0.65% in OA and mIoU, respectively. This finding confirms the synergistic effect between RSI and BF in UVs extraction, which has been further analyzed in this study. Additionally, the proposed model outperforms other deep learning models and exhibits the potential to support more modal features (e.g., POI). Finally, the experimental dataset and code can be publicly accessed at https://github.com/q1310546582/HR-UVFormer-code. Qingyan Meng, Fei Zhao 0002, Linlin Zhang 0007, Xinli Hu, Tamás Jancsó |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Beyond Pixel-Level Annotation: Exploring Self-Supervised Learning for Change Detection With Image-Level SupervisionabstractChange detection (CD) in high-resolution remote sensing has received large attention due to its wide range of applications. Many methods have been proposed in the literature and achieved excellent performance. However, they are often fully supervised, thus requiring abundant pixel-level labeled samples, which is time-consuming and labor-intensive. Especially compared to the common single-temporal interpretation, labeling bi-temporal images is often more complicated. Therefore, this study combines weakly supervised learning (WSL) to reduce label acquisition costs. But changed regions are small, fragmented, and similar to the background, which increase the gap between weakly supervised and fully supervised tasks. To address these difficulties, we explore self-supervised methods to construct a WSL framework based on image-level labels for general CD, termed WSLCD in this paper. First, we design a double-branch siamese network to derive embeddings and initial class attention maps (CAMs), which inputs the original image pair and the spatially transformed image pair. Second, mutual learning and equivariant regularization (MLER) is enforced on CAMs from different views, which implements consistency constraints in confusion regions and makes CAMs learn from each other based on saliency regions. Furthermore, prototype-based contrastive learning (PCL) is designed such that unreliable pixels can learn from prototypes computed from reliable pixels. PCL includes intra-view contrast and cross-view contrast depending on whether the prototypes and class embeddings are from the same view. With the above strategies, we narrow the gap between image-level weakly supervised CD and fully supervised CD. Experiments are conducted on three CD datasets, including CLCD, DSIFN and GCD. Our method achieves state-of-the-art performance on pseudo label generation and CD. The code is available at https://github.com/mfzhao1998/WSLCD. Maofan Zhao, Xinli Hu, Linlin Zhang 0007, Qingyan Meng, Yuxing Chen 0002, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | PanDiff: A Novel Pansharpening Method Based on Denoising Diffusion Probabilistic ModelabstractPansharpening is a crucial image processing technique for numerous remote sensing downstream tasks, aiming to recover high spatial resolution multispectral (HRMS) images by fusing high spatial resolution panchromatic (PAN) images and low spatial resolution multispectral (LRMS) images. Most current mainstream pansharpening fusion frameworks directly learn the mapping relationships from PAN and LRMS images to HRMS images by extracting key features. However, we propose a novel pansharpening method based on the denoising diffusion probabilistic model (DDPM) called PanDiff, which learns the data distribution of the difference maps (DM) between HRMS and interpolated MS (IMS) images from a new perspective. Specifically, PanDiff decomposes the complex fusion process of PAN and LRMS images into a multi-step Markov process, and the U-Net is employed to reconstruct each step of the process from random Gaussian noise. Notably, the PAN and LRMS images serve as the injected conditions to guide the U-Net in PanDiff, rather than being the fusion objects as in other pansharpening methods. Furthermore, we propose a modal intercalibration module (MIM) to enhance the guidance effect of the PAN and LRMS images. The experiments are conducted on a freely available benchmark dataset, including GaoFen-2, QuickBird, and WorldView-3 images. The experimental results from the fusion and generalization tests effectively demonstrate the outstanding fusion performance and high robustness of PanDiff. Fig. 1 depicts the results of the proposed method performed on various scenes. Additionally, the ablation experiments confirm the rationale behind PanDiff’s construction. Qingyan Meng, Wenxu Shi, Linlin Zhang 0007 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Local and Long-Range Collaborative Learning for Remote Sensing Scene ClassificationabstractWith the development of high-resolution satellites, more and more attention has been paid to remote sensing (RS) scene classification. Convolutional neural networks (CNNs), which replace the traditional handcrafted features with a learning-based feature extraction mechanism, are widely used in scene classification. But CNNs are less effective in deriving long-range contextual relations, which limits the further improvement. Visual transformer (VT), an emerging image processing method, provides a new perspective for RS scene classification by directly acquiring long-range features. Although there have been limited works combining CNN and VT through simple concatenation, the collaborations between them are insufficient. To address these issues, we propose a local and long-range collaborative framework (L2RCF). First, we design a dual-stream structure to extract the local and long-range features. Second, a cross-feature calibration (CFC) module is designed for them to improve representation of the fusion features. Then, combining deep supervision (DS) and deep mutual learning (DML), a novel joint loss is proposed to enhance the dual-stream feature extractor and further improve the fused features. Finally, a two-stage semi-supervised training strategy is designed to improve performance with unlabeled samples. To demonstrate the effectiveness of L2RCF, we conducted experiments on three widely used RS scene classification data sets: RSSCN7, AID, and NWPU. The results show that L2RCF performs significantly better compared with some state-of-the-art scene classification methods. Maofan Zhao, Qingyan Meng, Linlin Zhang 0007, Xinli Hu, Lorenzo Bruzzone |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Multilayer Feature Fusion Network With Spatial Attention and Gated Mechanism for Remote Sensing Scene ClassificationabstractRemote sensing (RS) scene classification has attracted extensive attention due to its large number of applications. Recently, convolutional neural networks (CNNs) methods have shown impressive ability of feature learning in RS scene classification. However, the performance is still limited by large-scale variance and complex background. To address these problems, we present a multilayer feature fusion network with spatial attention and gated mechanism (MLF2Net_SAGM) for RS scene classification. At first, the backbone is employed to extract multilayer convolutional features. Then, a residual spatial attention module (RSAM) is proposed to enhance discriminative regions of the multilayer feature maps, and key areas can be harvested. Finally, the multilayer spatial calibration features are fused to form the final feature map, and a gated fusion module (GFM) is designed to eliminate feature redundancy and mutual exclusion (FRME). To verify the effectiveness of the proposed method, we conduct comparative experiments based on three widely used RS image scene classification benchmarks. The results show that the direct fusion of multilayer features via element-wise addition leads to FRME, whereas our method fuses multilayer features more effectively and improves the performance of scene classification. Qingyan Meng, Maofan Zhao, Linlin Zhang 0007, Wenxu Shi, Lorenzo Bruzzone |
IEEE Geosci. Remote. Sens. Lett. | 3 |