EDBT 2026 Demo / reviewers in the wild / expert
Jie Ma 0004
dblp:62/5110-4
· DBLP profile ↗
18ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0001-7957-4274ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | KD-RSCC: A Karras Diffusion Framework for Efficient Remote Sensing Change CaptioningabstractRemote Sensing Image Change Captioning (RSICC) is a challenging task that involves describing surface changes between bi-temporal or multi-temporal satellite images using natural language. This task requires both fine-grained visual understanding and expressive language generation. Transformer-based and LSTM-based models have shown promising results in this domain. However, they may encounter difficulties in generating flexible and diverse captions, particularly when training data is limited or imbalanced. While diffusion models provide richer textual outputs, they are often constrained by long inference times. To address these issues, we propose a novel diffusion-based framework, KD-RSCC, for efficient and expressive remote sensing change captioning. This framework utilizes the Karras sampling method to significantly reduce the number of steps required during inference, while preserving the quality and diversity of the generated captions. Additionally, we introduce a Large Language Model (LLM)-based evaluation strategy G-EvalRSCC, to conduct a more comprehensive assessment of the semantic accuracy, fluency, and linguistic diversity of the generated descriptions. Experimental results demonstrate that KD-RSCC achieves an optimal balance between generation quality and inference speed, enhancing the flexibility and readability of its outputs. The code and supplementary materials are available at https://github.com/Fay-Y/KD_RSCC. Jie Ma 0004, Liqiang Qiao |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | Latent Diffusion, Implicit Amplification: Efficient Continuous-Scale Super-Resolution for Remote Sensing ImagesabstractRecent advancements in diffusion models have significantly improved performance in super-resolution (SR) tasks. However, previous research often overlooks the fundamental differences between SR and general image generation. General image generation involves creating images from scratch, while SR focuses specifically on enhancing existing low-resolution (LR) images by adding typically missing high-frequency details. This oversight not only increases the training difficulty but also limits their inference efficiency. Furthermore, previous diffusion-based SR methods are typically trained and inferred at fixed integer scale factors, lacking flexibility to meet the needs of upsampling with non-integer scale factors. To address these issues, this paper proposes latent diffusion for continuous-scale SR (LDCSR), an efficient and flexible SR model designed for remote sensing imagery. LDCSR employs a two-stage latent diffusion paradigm. During the first stage, an autoencoder is trained to capture the differential priors between high-resolution (HR) and LR images. The encoder intentionally ignores the existing LR content to alleviate the encoding burden, while the decoder introduces an SR branch equipped with a continuous scale upsampling module to accomplish the reconstruction under the guidance of the differential prior. In the second stage, a conditional diffusion model is learned within the latent space to predict the true differential prior encoding. Experimental results demonstrate that LDCSR achieves superior objective metrics and visual quality compared to the state-of-the-art SR methods. Additionally, it reduces the inference time of diffusion-based SR methods to a level comparable to that of non-diffusion methods. The code is available https://github.com/MoooJianG/LDCSR. Jiangwei Mo, Jie Ma 0004 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Diffusion-RSCC: Diffusion Probabilistic Model for Change Captioning in Remote Sensing ImagesabstractRemote sensing image change captioning (RSICC) aims at generating human-like language to describe the semantic changes between bitemporal remote sensing image (RSI) pairs. It provides valuable insight into environmental dynamics and land management. Unlike conventional change captioning (CC) tasks, RSICC involves not only retrieving relevant information across different modalities and generating fluent captions but also mitigating the impact of pixel-level differences on the terrain change localization. Pixel-level discrepancies over a long time span decrease caption accuracy. To address these problems, we propose a probabilistic diffusion-based model that leverages its remarkable generative capability to produce flexible captions. In the training phase, we construct a condition denoiser to efficiently map the real caption distribution to a standard Gaussian distribution. This denoiser incorporates cross-mode fusion (CMF) and stacking self-attention (SSA) modules to enhance cross-modal alignment and reduce pixel interference, thereby improving caption accuracy. In the training phase, the condition denoiser provides a new strategy for mean value estimation and helps to generate captions step by step. Extensive experiments on the LEVIR-CC dataset and DUBAI-CC dataset demonstrate the effectiveness of our Diffusion-RSCC and each of its individual components. The quantitative results showcase superior performance over existing methods across both traditional and newly introduced metrics. The code is available at:https://github.com/Fay-Y/Diffusion-RSCC. Jie Ma 0004 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | A Saliency-Aware Method for Arbitrary Style TransferabstractArbitrary style transfer has drawn widespread attention from academic communities for its extensive application in reality. However, existing methods fail to make a proper trade-off between flexibility and capability. All of these methods directly synthesize style textures on the whole content image indiscriminately, leading to local distortions when there are obvious visual distinctions in different regions. In this paper, we propose a novel Saliency-Aware Instance Normalization (SAIN) module considering the visual distinctions between salient area and non-salient area for arbitrary style transfer. To extract region-specific stylized features, SAIN separately performs the local feature alignment in each region. Besides, for the full integration of global and local features, we design a global branch in our model, where content features are adjusted with the globally calculated statistics of style features. Comprehensive evaluations demonstrate our method’s superiority over six state-of-the-art arbitrary style transfer methods. Source code is available at https://github.com/yitongli123/SAIN. Jie Ma 0004 |
ICIP | 2 |
| 2023 | Weakly-Supervised ROI Extraction Method Based on Contrastive Learning for Remote Sensing ImagesabstractROI extraction is an active but challenging task in remote sensing because of the complicated landform, the complex boundaries and the requirement of annotations. Weakly supervised learning (WSL) aims at learning a mapping from input image to pixel-wise prediction under image-wise labels, which can dramatically decrease the labor cost. However, due to the imprecision of labels, the accuracy and time consumption of WSL methods are relatively unsatisfactory. In this paper, we propose a two-step ROI extraction based on contractive learning. Firstly, we present to integrate multiscale Grad-CAM to obtain pseudo pixelwise annotations with well boundaries. Then, to reduce the compact of misjudgments in pseudo annotations, we construct a contrastive learning strategy to encourage the features inside ROI as close as possible and separate background features from foreground features. Comprehensive experiments demonstrate the superiority of our proposal. Code is available at https://github.com/HE-Lingfeng/ROI-Extraction Mengze Xu, Jie Ma 0004 |
IGARSS | 3 |
| 2023 | Scribble-Supervised Target Extraction Method Based on Inner Structure-Constraint for Remote Sensing ImagesabstractWeakly supervised learning based on scribble annotations in target extraction of remote sensing images has drawn much interest due to scribbles’ flexibility in denoting winding objects and low cost of manually labeling. However, scribbles are too sparse to identify object structure and detailed information, bringing great challenges in target localization and boundary description. To alleviate these problems, in this paper, we construct two inner structure- constraints, a deformation consistency loss and a trainable active contour loss, together with a scribble-constraint to supervise the optimization of the encoder-decoder network without introducing any auxiliary module or extra operation based on prior cues. Comprehensive experiments demonstrate our method’s superiority over five state-of-the- art algorithms in this field. Source code is available at https://github.com/yitongli123/ISC-TE. Chang Liu 0071, Jie Ma 0004 |
IGARSS | 3 |
| 2023 | Dual-Diffusion: Dual Conditional Denoising Diffusion Probabilistic Models for Blind Super-Resolution Reconstruction in RSIsabstractPrevious super-resolution reconstruction (SR) works are always designed on the assumption that the degradation operation is fixed, such as bicubic downsampling. However, as for remote sensing images, some unexpected factors can cause the blurred visual performance, like weather factors, orbit altitude, etc. Blind SR methods are proposed to deal with various degradations. There are two main challenges of blind SR in RSIs: 1) the accurate estimation of degradation kernels; 2) the realistic image generation in the ill-posed problem. To rise to the challenge, we propose a novel blind SR framework based on dual conditional denoising diffusion probabilistic models (DDSR). In our work, we introduce conditional denoising diffusion probabilistic models (DDPM) from two aspects: kernel estimation progress and reconstruction progress, named as the dual-diffusion. As for kernel estimation progress, conditioned on low-resolution (LR) images, a new DDPM-based kernel predictor is constructed by studying the invertible mapping between the kernel distribution and the latent distribution. As for reconstruction progress, regarding the predicted degradation kernels and LR images as conditional information, we construct a DDPM-based reconstructor to learning the mapping from the LR images to HR images. Comprehensive experiments show the priority of our proposal compared with SOTA blind SR methods. Source Code and supplementary materials are available at https://github.com/Lincoln20030413/DDSR. Mengze Xu, Jie Ma 0004, Yuanyuan Zhu 0003 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Remote Sensing Image Super-Resolution via Saliency-Guided Feedback GANsabstractIn remote sensing images (RSIs), the visual characteristics of different regions are versatile, which poses a considerable challenge to single image super-resolution (SISR). Most existing SISR methods for RSIs ignore the diverse reconstruction needs of different regions and thus face a serious contradiction between high perception quality and less spatial distortion. The mean square error (MSE) optimization-based methods produce results of unsatisfactory visual quality, while generative adversarial networks (GANs) can produce photo-realistic but severely distorted results caused by pseudotextures. In addition, increasingly deeper networks, although providing powerful feature representations, also face problems of overfitting and occupying too much storage space. In this article, we propose a new saliency-guided feedback GAN (SG-FBGAN) to address these problems. The proposed SG-FBGAN applies different reconstruction principles for areas with varying levels of saliency and uses feedback (FB) connections to improve the expressivity of the network while reducing parameters. First, we propose a saliency-guided FB generator with our carefully designed paired-feedback block (PFBB). The PFBB uses two branches, a salient and a nonsalient branch, to handle the FB information and generate powerful high-level representations for salient and nonsalient areas, respectively. Then, we measure the visual perception quality of salient areas, nonsalient areas, and the global image with a saliency-guided multidiscriminator, which can dramatically eliminate pseudotextures. Finally, we introduce a curriculum learning strategy to enable the proposed SG-FBGAN to handle complex degradation models. Comprehensive evaluations and ablation studies validate the effectiveness of our proposal. Libao Zhang, Jie Ma 0004 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Scribble-Supervised ROI Extraction Using Residual Dense Dilated Network for Remote Sensing ImagesabstractIn this paper, we focus on ROI extraction with only scribble annotations for remote sensing images (RSIs). The main challenges exist in predicting precise region boundary and suppressing complex background interference in RSIs. To address these issue, we propose a scribble-supervised residual dense dilated network (RD-DN) trained with a novel label update strategy to produce integral ROIs with accurate boundary. Specifically, the hybrid dilated convolution block is introduced as the basic module of the RD-DN, aiming to help provide much denser results. The RD-DN is first trained with only scribble labels to generate initial results. Then, as training phase goes, the label update strategy updates the labels iteratively by combining the initial results and the scribble annotations with a morphological dilation procedure. Comprehensive experiments on the GeoEye-1 dataset demonstrate the superiority of our proposal compared with state-of-the-art ROIs extraction methods. Jie Ma 0004 |
IGARSS | 1 |
| 2021 | Salient Object Detection Based on Progressively Supervised Learning for Remote Sensing ImagesabstractSalient object detection (SOD) is a crucial task in the field of remote sensing image (RSI) processing. Weakly supervised SOD methods, which generate saliency maps by classification convolutional neural networks (CNNs), considerably reduce labor costs. However, due to the complexity of remote sensing scenes, concerns remain about weakly supervised SOD for RSIs: 1) since the pooling operations are applied in the classification CNNs, the boundary maintenance of weakly supervised methods is unsatisfactory and 2) several sophisticated postprocessing procedures are used in previous weakly supervised methods, which are inevitably time-consuming. To solve these problems, we combine the benefits of weakly and fully supervised learning and propose a new SOD method named progressively supervised learning (PSL) for RSIs. The proposed method realizes end-to-end SOD with a lightweight model under imagewise annotations. First, to reduce the demands on large-scale pixelwise annotations, we propose a pseudo-label generation method based on a classification network and gradient-weighted class activation mapping (Grad-CAM) to compute pseudo saliency maps (PSMs) for training samples and auxiliary images in a weakly supervised manner. Then, to improve the computational efficiency, we construct a feedback saliency analysis network (FSAN), where the generated PSMs are regarded as pixelwise labels. Finally, inspired by curriculum learning, we design a new denoising loss function to further reduce the effect brought by missing judgment in PSMs and enhance the detection accuracy. Comprehensive evaluations with two remote sensing data sets and a comparison with 11 methods validate the superiority of the proposed PSL model. Libao Zhang, Jie Ma 0004 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | SC-PNN: Saliency Cascade Convolutional Neural Network for PansharpeningabstractIn many remote sensing tasks, different types of regions or targets differ in requirements for spectral and spatial quality. The discrepancy reveals that a uniform pansharpening strategy applying to the entire image may not fulfill the varying demands of different regions appropriately. From this aspect, we resort to saliency analysis to distinguish regions with different spatial and spectral requirements and then propose a new saliency cascade convolutional neural network for pansharpening (SC-PNN). SC-PNN is composed of two parts: a dilated deformable convolutional network (DDCN) for saliency analysis and a saliency cascade residual dense network (SC-RDN) for pansharpening. DDCN is a fully convolutional network based on hybrid dilated convolution and deformable convolution, aiming to separate salient regions, such as residential areas from nonsalient areas, including mountains and vegetation areas, with well-defined boundaries and integrity. In the fusion process, SC-RDN is specially designed with the help of saliency analysis. We first construct a deep regression network to estimate a primarily sharpened image and subsequently leverage the saliency map produced by DDCN to develop a saliency enhancement module. In this module, the quality of salient and nonsalient areas is further improved by two independent deep residual dense networks. Thus, a precise fused image can be predicted. Experiments on SPOT5, GeoEye-1, and WorldView-3 data sets reveal that, compared to state-of-the-art pansharpening methods, our proposal has a superior ability to improve the spatial quality and preserve spectral information. The effectiveness of the saliency enhancement module is also validated in the experiment. Libao Zhang, Jue Zhang 0001, Jie Ma 0004, Xiuping Jia |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | SD-FB-GAN: Saliency-Driven Feedback Gan for Remote Sensing Image Super-Resolution ReconstructionabstractThe visual characteristics of different regions in remote sensing images are significantly versatile, which poses a huge challenge to single image super-resolution. Although generative adversarial network (GAN) has shown great potential in generating photo-realistic results, it provides unsatisfactory performance in objective metrics owning to pseudo textures brought by adversarial learning. In this paper, we propose a new saliency-driven feedback GAN to cope with these problems. We design a saliency-driven feedback generator based on paired-feedback blocks (PFBBs) and recurrent structure to provide strong reconstruction ability. In the PFBB, the saliency map serves as an indicator to reflect the texture complexity, so different reconstruction principles can be applied to restore areas with varying levels of saliency. Besides, we propose to measure the visual quality of salient areas, non-salient areas, and the whole image with multi-discriminators, which can dramatically eliminate pseudo textures. Comprehensive evaluations and ablation studies validate the superiority of our proposal. Jie Ma 0004, Jue Zhang 0001, Libao Zhang |
ICIP | 1 |
| 2020 | SD-GAN: Saliency-Discriminated GAN for Remote Sensing Image SuperresolutionabstractRecently, convolutional neural networks have shown superior performance in single-image superresolution. Although existing mean-square-error-based methods achieve high peak signal-to-noise ratio (PSNR), they tend to generate oversmooth results. Generative adversarial network (GAN)-based methods can provide high-resolution (HR) images with higher perceptual quality, but produce pseudotextures in images, which generally leads to lower PSNR. Besides, different regions in remote sensing images (RSIs) reflect discrepant surface topography and visual characteristics. This means a uniform reconstruction strategy may not be suitable for all targets in RSIs. To solve these problems, we propose a novel saliency-discriminated GAN for RSI superresolution. First, hierarchical weakly supervised saliency analysis is introduced to compute a saliency map, which is subsequently employed to distinguish the diverse demands of regions in the following generator and discriminator part. Different from previous GANs, the proposed residual dense saliency generator takes saliency maps as a supplementary condition in the generator. Simultaneously, combining the characteristic of RSIs, we design a new paired discriminator to enhance the perceptual quality, which measures the distance between generated images and HR images in salient areas and nonsalient areas, respectively. Comprehensive evaluations validate the superiority of the proposed model. Jie Ma 0004, Libao Zhang, Jue Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2020 | Hierarchical Weakly Supervised Learning for Residential Area Semantic Segmentation in Remote Sensing ImagesabstractResidential-area segmentation is one of the most fundamental tasks in the field of remote sensing. Recently, fully supervised convolutional neural network (CNN)-based methods have shown superiority in the field of semantic segmentation. However, a serious problem for those CNN-based methods is that pixel-level annotations are expensive and laborious. In this study, a novel hierarchical weakly supervised learning (HWSL) method is proposed to realize pixel-level semantic segmentation in remote sensing images. First, a weakly supervised hierarchical saliency analysis is proposed to capture a sequence of class-specific hierarchical saliency maps by computing the gradient maps with respect to the middle layers of the CNN. Then, superpixels and low-rank matrix recovery are introduced to highlight the common salient areas and fuse class-specific saliency maps with adaptive weights. Finally, a subtraction operation between class-specific saliency maps is conducted to generate hierarchical residual saliency maps and fulfill residential-area segmentation. Comprehensive evaluations with two remote sensing data sets and comparison with seven methods validate the superiority of the proposed HWSL model. Libao Zhang, Jie Ma 0004, Xinran Lv |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2020 | ROI Extraction Based on Multiview Learning and Attention Mechanism for Unbalanced Remote Sensing Data SetabstractWith an increasing number of remote sensing images (RSIs), the automatic region of interest (ROI) extraction based on convolutional neural networks (CNNs) has attracted much interest in recent years. Although fully supervised CNN-based methods have shown superiority in the field of object extraction, pixelwise annotations are expensive and time-consuming. Moreover, due to the unstable distribution of ROIs in complex landscapes, the ratios of foreground and background areas are quite different in RSIs. Training CNNs with such unbalanced data sets lead to over-fitting and low accuracy. In this article, we propose a framework that combines multiview learning and attention mechanism (MLAM) to solve the above mentioned problems. First, we develop a CNN-based weakly supervised method with a weight-balanced loss function to solve the problems caused by an unbalanced data set. It also helps to generate imagewise saliency maps by computing the gradient maps with respect to the input images. Then, we design a multiview strategy to dramatically reduce the missing inspection. Finally, we design a feedback attention mechanism based on the stage neighbor binary pattern to further modify the extraction result. In summary, the proposed framework achieves pixelwise ROI extraction under imagewise annotations through data-driven ML and a knowledge-driven visual attention mechanism. We evaluate the performance of the MLAM framework on two challenging data sets with complex backgrounds. The experimental results indicate that the proposed framework can achieve better performance than other eight ROI extraction models for unbalanced remote sensing data sets. Jie Ma 0004, Libao Zhang, Yang Sun 0007 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Remote-Sensing Image Superresolution Based on Visual Saliency Analysis and Unequal Reconstruction NetworksabstractRemote-sensing images (RSIs) generally have strong spatial characteristics for surface features. Various ground objects, such as residential areas, roads, forests, and rivers, differ substantially. According to this visual attention characteristic, regions with complicated texture features require more realistic details to reflect a better description of the topography, while regions such as farmlands should be smooth and have less noise. However, most existing single-image superresolution (SISR) methods fail to fully utilize these properties and therefore apply a uniform reconstruction strategy to the whole image. In this article, we propose a novel saliency-driven unequal single-image reconstruction network in which the demands of various regions in the superresolution (SR) process are distinguished by saliency maps. First, we design a new gradient-based saliency analysis method to produce more accurate saliency maps with imagewise annotations. The method utilizes the superiority of a multireception field to extract both high-level features and low-level features. Second, we propose a novel saliency-driven gate conditional generative adversarial network, where the saliency map is regarded as a medium during the training procedure of the whole network. The saliency map is regarded as a pixelwise condition in a generator to enhance the training capability of the network. Additionally, we design a new loss function that combines normalized content loss, saliency-driven perceptual loss, and gate-control adversarial loss to further refine details of texture-complex areas for RSIs. We evaluate the performance of our algorithm and compare it with many other state-of-the-art SR methods using a remote-sensing data set. The experimental results show that our approach achieves the optimal outcome in salient areas. Our method attains the best effect on global quality and visual performance. Libao Zhang, Jie Ma 0004, Jue Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | A New Pansharpening Method Using Objectness Based Saliency Analysis and Saliency Guided Deep Residual NetworkabstractPansharpening is a fundamental and crucial task in the remote sensing community. For remote sensing images, there is a significant difference in demands for spatial and spectral resolution in different regions. From this perspective, we propose a new pansharpening method using objectness based saliency analysis and saliency guided deep residual network to boost the fusion accuracy. We first develop an objectness based saliency analysis by incorporating texture feature and objectness measurements to estimate saliency values in images and thereby help discriminate different demands for spatial improvement and spectral preservation. Inspired by the impressive performance of deep learning, we subsequently construct a saliency guided deep residual network to implement pansharpening. In addition, in order to produce images with subtler details, we design a new loss function, the normalized mean square error, particularly for the pansharpening task. Experiments support the superiority of our proposal over six competing methods. Libao Zhang, Jue Zhang 0001, Xinran Lyu, Jie Ma 0004 |
ICIP | 4 |
| 2019 | Target heat-map network: An end-to-end deep network for target detection in remote sensing images
Huai Chen, Libao Zhang, Jie Ma 0004, Jue Zhang 0001 |
Neurocomputing | 3 |