Chengli Peng

dblp:272/3784 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Multi-source attention autoencoder network for hyperspectral unmixing with LiDAR data
Jiwei Hu, Yangrui Bai, Qiwen Jin, Chengli Peng
Neurocomputing5
2025 Task Knowledge Injection: Training-Free Adaptation of Multimodal Large Language Models for Remote Sensing Image Understanding
abstract
Parameter fine-tuning is the mainstream approach for adapting Multimodal Large Langauge Models (MLLMs) to downstream remote sensing tasks. However, such method risks degrading pre-trained knowledge and also incur significant costs. This paper argues that downstream adaptation of MLLMs essentially involves effective injection of task-specific knowledge, which does not necessarily require parameter updates. Based on this perspective, we propose a training-free knowledge injection method. By constructing a multi-task knowledge base (MTKB), the model can dynamically retrieve task-related knowledge to serve as context during inference, thereby enhancing its understanding. Specifically, we design a three-part framework. (1) Task knowledge construction: Diverse texts are unified into a key and value structure for image-text matching, forming the MTKB. (2) Two-stage retrieval: A coarse-to-fine process is employed to match query images with task knowledge in the MTKB, utilizing both unimodal and cross-modal similarity. (3) Knowledge injection: Matched knowledge is integrated into the MLLM via extended embeddings, without altering parameters. Experimental results across multiple datasets demonstrate that our method significantly enhances the image understanding capabilities of the model. Our approach achieves about 5% improvement in accuracy and related metrics across several datasets, with performance on the RSVQ-LR dataset comparable to specialized models.
Haifeng Li 0007, Qiujun Li, Wang Guo, Hongyuan Yuan, Run Shao, Chengli Peng
IEEE Geosci. Remote. Sens. Lett.8
2025 HASNet: A Foreground Association-Driven Siamese Network With Hard Sample Optimization for Remote Sensing Image Change Detection
abstract
Remote sensing change detection (RS-CD) relies on the model’s ability to learn features of marked change objects, known as foreground targets. Beyond foreground targets, the background targets are more valuable samples for change detection, such as unlabeled ones, semantically ambiguous ones, pseudo-changes, and non-interesting changes, referred to as hard case samples (HCSs) in this article. There are two additional challenges to learning HCSs: 1) the loss function focusing on the foreground targets with rich labels and ignoring the HCSs in the background, called the imbalance problem and 2) it is difficult for a model to learn the change information of HCSs directly, which is called HCSs missingness. This article proposed a foreground association-driven Siamese network with hard sample optimization (HASNet). To deal with the imbalance problem, we propose an equilibrium optimization loss (EO-loss) function to regulate the optimization focus of the foreground and background, determine the HCSs through the distribution of the loss values, and introduce dynamic weights in the loss term to gradually shift the optimization focus of the loss from the foreground to the background hard cases as the training progresses. To address the HCSs missingness, we propose the scene-foreground association module by using potential remote sensing spatial scene information to model the association between the target of interest in the foreground and the related context to obtain scene embedding to reinforce the feature of hard cases. Experiments on four public datasets with 11 baselines show that HASNet outperforms current state-of-the-art CD methods, particularly in detecting HCSs. The source code is available athttps://github.com/GeoX-Lab/HASNet.
Chao Tao 0001, Dongsheng Kuang, Zhenyang Huang, Chengli Peng, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.4
2024 A more reliable local-global-guided network for correspondence pruning
Chengli Peng, Zizhuo Li, Qiwen Jin
Pattern Recognit. Lett.1
2023 MSINet: Mining scale information from digital surface models for semantic segmentation of aerial images
Chengli Peng, Haifeng Li 0007, Chao Tao 0001, Yansheng Li 0001, Jiayi Ma 0001
Pattern Recognit.1
2023 GraSS: Contrastive Learning With Gradient-Guided Sampling Strategy for Remote Sensing Image Semantic Segmentation
abstract
Self-supervised contrastive learning (SSCL) has achieved significant milestones in remote sensing image (RSI) understanding. Its essence lies in designing an unsupervised instance discrimination pretext task to extract image features from a large number of unlabeled images that are beneficial for downstream tasks. However, existing instance discrimination based SSCL suffers from two limitations when applied to the RSI semantic segmentation task: 1) Positive sample confounding issue, SSCL treats different augmentations of the same RSI as positive samples, but the richness, complexity, and imbalance of RSI ground objects lead to the model actually pulling a variety of different ground objects closer while pulling positive samples closer, which confuse the feature of different ground objects. 2) Feature adaptation bias, SSCL treats RSI patches containing various ground objects as individual instances for discrimination and obtains instance-level features, which are not fully adapted to pixel-level or object-level semantic segmentation tasks. To address the above limitations, we consider constructing samples containing single ground objects to alleviate positive sample confounding issue, and make the model obtain object-level features from the contrastive between single ground objects. Meanwhile, we observed that the discrimination information can be mapped to specific regions in RSI through the gradient of unsupervised contrastive loss, these specific regions tend to contain single ground objects. Based on this, we propose contrastive learning with Gradient guided Sampling Strategy (GraSS) for RSI semantic segmentation. GraSS consists of two stages: 1) the instance discrimination warm-up stage to provide initial discrimination information to the contrastive loss gradients, 2) the gradient guided sampling contrastive training stage to adaptively construct samples containing more singular ground objects using the discrimination information. Experimental results on three open datasets demonstrate that GraSS effectively enhances the performance of SSCL in high-resolution RSI semantic segmentation. Compared to eight baseline methods from six different types of SSCL, GraSS achieves an average improvement of 1.57% and a maximum improvement of 3.58% in terms of mean intersection over the union. Additionally, we discovered that the unsupervised contrastive loss gradients contain rich feature information, which inspires us to utilize gradient information more extensively during model training to attain additional model capacity. The source code is available at https://github.com/GeoX-Lab/GraSS.
Chao Tao 0001, Yunsheng Zhang 0001, Chengli Peng, Haifeng Li 0007
IEEE Trans. Geosci. Remote. Sens.5
2022 Robust Feature Matching via Local Consensus
abstract
Feature matching is the foundation and key task of remote sensing image registration, which is to establish a reliable point corresponding relationship between the feature points of two images. In this article, a simple and effective local consensus method for rigid and nonrigid feature matching is proposed and applied to solve the problem of high outliers ratio caused by nonrigid transformation, nonlinear radiation difference, and speckle noise in the remote sensing image registration task. We first establish the putative feature correspondences according to the similarity between local descriptors and then use local consensus constraints (including neighborhood consensus and motion vector consensus) to remove outliers. The specific steps are given as follows. First, we use the neighborhood consensus constraint of feature points to carry out preliminary filtering to remove outliers with obvious errors and retain a large number of inliers, so as to obtain a clean reliable set. Then, the reliable set space is grided into several nonoverlapping cells, and the estimated motion vector is calculated for each cell. By taking the comprehensive deviation between the ordinary motion vectors and estimated motion vectors, we transform the matching problem into a mathematical optimization model and derive a closed-form solution with linear time and linear space complexities. In this way, our method can also significantly increase the speed of operation without sacrificing accuracy. A large number of feature matching experiments on remote sensing prove that our method is superior to existing methods and also has good results in the general scene.
Jun Chen 0019, Meng Yang 0031, Chengli Peng, Linbo Luo 0002, Wenping Gong
IEEE Trans. Geosci. Remote. Sens.3
2022 Cross Fusion Net: A Fast Semantic Segmentation Network for Small-Scale Semantic Information Capturing in Aerial Scenes
abstract
Capturing accurate multiscale semantic information from the images is of great importance for high-quality semantic segmentation. Over the past years, a large number of methods attempt to improve the multiscale information capturing ability of the networks via various means. However, these methods always suffer unsatisfactory efficiency (e.g., speed or accuracy) on the images that include a large number of small-scale objects, for example, aerial images. In this article, we propose a new network named cross fusion net (CF-Net) for fast and effective extraction of the multiscale semantic information, especially for small-scale semantic information. In particular, the proposed CF-Net can capture more accurate small-scale semantic information from two aspects. On the one hand, we develop a channel attention refinement block to select the informative features. On the other hand, we propose a cross fusion block to enlarge the receptive field of the low-level feature maps. As a result, the network can encode more accurate semantic information from the small-scale objects, and the segmentation accuracy of the small-scale objects is improved accordingly. We have compared the proposed CF-Net with several state-of-the-art semantic segmentation methods on two popular aerial image segmentation data sets. Experimental results reveal that the average$F_{1}$score gain brought by our CF-Net is about 0.43% and the$F_{1}$score gain of the small-scale objects (e.g., cars) is about 2.61%. In addition, our CF-Net has the fastest inference speed, which proves its superiority in the aerial scenes. Our code will be released at:https://github.com/pcl111/CF-Net.
Chengli Peng, Kaining Zhang, Yong Ma 0001, Jiayi Ma 0001
IEEE Trans. Geosci. Remote. Sens.1
2022 DBDnet: A Deep Boosting Strategy for Image Denoising
abstract
In this paper, we propose a new deep network architecture named deep boosting denoising net (DBDnet) for image denoising. It is a residual learning network that can generate a noise map from a noisy observation. In detail, it first generates a coarse noise map via a simple structure, and then updates the noise map gradually via a boosting function. The motivation of our DBDnet stems from the observation that the noise map recovered by any algorithm cannot ideally equal the ground-truth noise map, which typically contains noise. We call this noise NoN,i.e., noise of noise map. Based on this observation, we formulate the denoising as a process of reducing NoN, and the role of DBDnet is to eliminate the NoN from the coarse noise map. In particular, we analyze the process of reducing NoN theoretically, and propose an NoN eliminating module to simulate it accordingly. We evaluate the proposed DBDnet on images polluted by different levels of additive white Gaussian noise and real noise. Experiment results demonstrate that our DBDnet can attain better denoising performance compared with state-of-the-art methods on several kinds of image denoising tasks. In particular, for the Gaussian denoising and real image denoising tasks, the average improvements of the PSNR values brought by our DBDnet are about 0.25 dB and 1.01 dB, respectively. In addition, we find and verify that the deep boosting insight can be easily introduced into the state-of-the-art image denoising network, and promotes its denoising performance. Our code is publicly available athttps://github.com/jiayi-ma/DBDNet.
Jiayi Ma 0001, Chengli Peng, Xin Tian 0006, Junjun Jiang
IEEE Trans. Multim.2
2021 Bilateral attention decoder: A lightweight decoder for real-time semantic segmentation
Chengli Peng, Tian Tian 0006, Chen Chen 0001, Xiaojie Guo 0001, Jiayi Ma 0001
Neural Networks1
2020 Semantic segmentation using stride spatial pyramid pooling and dual attention decoder
Chengli Peng, Jiayi Ma 0001
Pattern Recognit.1