Chenguang Dai

dblp:119/1951 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
YearPublicationVenuePosition
2026 P3D: Plug-and-play prompt-driven framework for RGB-thermal semantic segmentation
abstract
• A plug-and-play prompt-driven framework for RGB-thermal image semantic segmentation. • LoRA-based fine-tuning strategy for SAM series model integration. • A model-agnostic encoder to generate statistical distributed prompts for training. The semantic segmentation of RGB-thermal images is critical for applications with low-light conditions. Existing works primarily focus on feature fusion strategies and model design to enhance performance. While Visual Foundation Models (VFMs) have been introduced in previous studies to improve generalization and segmentation accuracy, they suffer from poor compatibility with other models thus requiring full model retraining. Additionally, the domain gap and modality gap between VFM pre-training datasets and RGB-thermal semantic segmentation datasets pose significant challenges to VFM adaptation for downstream tasks. To address these issues, in this paper a plug-and-play prompt driven framework P 3 D is proposed. Unlike existing VFM-based methods that require complete retraining for each specific architecture, P 3 D is designed with a model-agnostic training strategy that enables one-time training and seamless integration with various existing methods without requiring retraining. First, a dual-branch LoRA (Low-Rank Adaptation) fine-tuned (DBLF) image encoder for the RGB and thermal image branches is proposed to narrow the domain gap and modality gap when incorporating SAM series models into our task. Second, a unified prompt generation and representation (UPGR) encoder is proposed. It generates diverse prompts using semantic labels during the training stage, ensuring the generated prompts are model-agnostic and compatible with existing methods. Finally, a cross-modality spatial-channel attention (CM-SCA) decoder is developed to fuse the embeddings from two-modality images and prompts for the final prediction. Extensive experiments are conducted on three popular benchmarks. Results demonstrate that P 3 D not only improves the performance of existing models but also outperforms current state-of-the-art (SOTA) methods leveraging < 1% trainable parameters. More importantly, by simply plugging P 3 D into existing methods, we consistently achieve significant performance improvements without retraining these base models, demonstrating the practical value of our plug-and-play design.
Yongqi Sun, Chenguang Dai, Hanyun Wang, Longguang Wang, Wenke Li, Anzhu Yu
Pattern Recognit.2
2025 Ms-DANet: Multiscale Difference-Aware Network for 3-D Point Cloud Change Detection
abstract
With the rapid advancements of 3-D acquisition technology, 3-D change detection has gained lots of attentions recently. Existing deep learning-based point cloud change detection methods usually adopt a common encoder-decoder structure to learn pointwise features. However, these feature learning backbones are not specifically designed for change detection task, and ignore the local structure discrepancies during feature learning. To address these issues, this article proposes a multiscale difference-aware network (Ms-DANet) for 3-D point cloud change detection. First, we propose a difference-guided multiscale feature learning (DG-MsFL) module to enhance the feature differences between bi-temporal point clouds at multiple scales during feature encoding, and use these differences to guide the network focusing more on the local structures with large discrepancies. Next, we introduce a multiscale difference feature fusion (Ms-DFF) module to fuse the multiscale feature differences to learn more discriminative features during feature decoding. Finally, we treat the point cloud change detection task as a semantic classification problem, and propose a multiscale loss (Ms-Loss) function to promote the network training. We conduct experiments on the real-world street-level point cloud change detection dataset SLPCCD and the simulated airborne urban point cloud change detection dataset URB3DCD. The experimental results show that Ms-DANet obtains a significant improvement on both the real-world and simulated point cloud change detection datasets, demonstrating its effectiveness and robustness across various sensors and data modalities.
Jinhao Lu, Chenguang Dai, Zhenchao Zhang 0001, Xuanguang Liu, Ruqin Zhou, Song Ji, Haiyan Guan, Hanyun Wang
IEEE Trans. Geosci. Remote. Sens.2
2025 Intelligent Detection of Sports Fields Based on Very High-Resolution Stereo Satellite Imagery
abstract
Sports fields are primary locations for people to engage in sports activities and cultural exchanges, possessing significant social and economic value and being an important object for remote-sensing mapping. To extract sports field data efficiently and accurately, the first multicategory, fine-grained sports field dataset, SF_VHR_SSI, was constructed based on various measured very-high-resolution stereo satellite data. It includes comprehensive metadata information and diverse sample data, capable of providing data support for various types of object detection tasks. To ascertain the scientific validity and rationality of the dataset, extensive benchmarking experiments were conducted utilizing 15 established algorithms encompassing 54 distinct models. Based on these experiments, the complexity and challenges inherent in the SF_VHR_SSI dataset were analyzed, thereby providing insights that inform the design of subsequent algorithms. To improve the accuracy of sports field detection further, partial convolution was introduced, a squeeze-and-excitation mechanism was incorporated, and a multi-detection box fusion strategy was employed. The FasterCSPELAN4 feature extraction module, SEAttention attention module, and Merge-NMS postprocessing module were designed. The novel SF_GELAN network was constructed based on the generalized efficient layer aggregation network (GELAN) architecture, which is both efficient and accurate. Furthermore, by utilizing model scaling techniques, three versions of the network—large, medium, and small—were developed to cater to diverse application scenarios. A comparison of the experimental results demonstrates that all three versions of the SF_GELAN network achieved state-of-the-art performance on the SF_VHR_SSI dataset, providing robust algorithmic support for high-precision sports field detection globally.
Dashuai Shang, Chenguang Dai
IEEE Trans. Geosci. Remote. Sens.3
2025 Supervised Contrastive Learning for Indoor Point Cloud Oversegmentation
abstract
Point cloud oversegmentation method can obtain a series of superpoints by grouping points that are semantically and geometrically consistent. The generated superpoints can be treated as the basic processing units in various downstream tasks to improve task performance and processing efficiency. However, due to the high semantic and geometric complexity of point cloud scenes, obtaining high-quality superpoints is still challenging. Aiming to generate high-quality indoor superpoints, we propose an end-to-end supervised contrastive learning framework SCL-OverSeg for indoor point cloud oversegmentation. Firstly, to solve the challenge of balancing the importance of geometric similarity and spatial proximity constraint between points and superpoints in indoor scenes, we integrate the geometric similarity and spatial proximity constraint into the supervision signal by generating the superpoint ground truth. To solve the challenge of superpoints crossing objects, we propose to utilize instance labels rather than semantic labels to generate the ideal superpoint ground truth as the object-level supervision signal. Secondly, to construct the distinguishable embedding space facilitating to the assignments of points to superpoints, we propose point-superpoint contrastive learning to compel the network to project each point to be closer to the reasonable superpoint in embedding space. Besides, with the instance labels, to improve the superpoint performance on object boundaries, we propose the object boundary contrastive learning to enhance the feature distinguishability between tough points across the object boundaries. Extensive experiments demonstrate that SCL-OverSeg can effectively improve indoor oversegmentation performance, especially on object boundaries. The relevant codes will be available onhttps://github.com/sssssyf/SCL-OverSeg.
Yifan Sun 0008, Chenguang Dai, Wenke Li, Song Ji, Anzhu Yu, Yiping Chen 0002, Hanyun Wang
IEEE Trans. Multim.2
2024 TBSCD-Net: A Siamese Multitask Network Integrating Transformers and Boundary Regularization for Semantic Change Detection From VHR Satellite Images
abstract
Semantic change detection (SCD) from very high-resolution images involves two key challenges: (1) the global features of bitemporal images tend to be extracted insufficiently, leading to imprecise land cover semantic classification results, and (2) the detected changed objects exhibit ambiguous boundaries, resulting in low geometric accuracy. To address these two issues, we propose an SCD method called TBSCD-Net based on a multi-task learning framework to simultaneously identify different types of semantic changes and regularize change boundaries. Firstly, we construct a hybrid encoder combining transformer and convolutional neural network (TCEncoder) to enhance the extraction of global context information. A bitemporal semantic linkage module (Bi-SLM) is embedded into the TCEncoder to enhance the semantic correlations between bitemporal images. Secondly, we introduce a boundary-region joint extractor based on Laplacian operators (LOBRE) to regularize the changed objects. We evaluated the effectiveness of the proposed method using the SECOND dataset and a Fuzhou GF-2 SCD dataset (FZ-SCD) and compared it with seven existing methods. The proposed method performed better than the other evaluated methods as it achieved 24.42% Sek and 20.18% GTC on the SECOND dataset and 23.10% Sek and 23.15% GTC on the FZ-SCD dataset. The results of ablation studies on the FZ-SCD dataset also verified the effectiveness of the developed modules for SCD.
Xuanguang Liu, Chenguang Dai, Zhenchao Zhang 0001, Mengmeng Li 0002, Hanyun Wang, Hongliang Ji, Yujie Li 0010
IEEE Geosci. Remote. Sens. Lett.2
2024 Deep Semantic Graph Matching for Large-Scale Outdoor Point Cloud Registration
abstract
Current point cloud registration methods are mainly based on local geometric information and usually ignore the semantic information contained in the scenes. In this paper, we treat the point cloud registration problem as a semantic instance matching and registration task, and propose a deep semantic graph matching method (DeepSGM) for large-scale outdoor point cloud registration. Firstly, the semantic categorical labels of 3D points are obtained using a semantic segmentation network. The adjacent points with the same category labels are then clustered together using the Euclidean clustering algorithm to obtain the semantic instances, which are represented by three kinds of attributes including spatial location information, semantic categorical information, and global geometric shape information. Secondly, the semantic adjacency graph is constructed based on the spatial adjacency relations of semantic instances. To fully explore the topological structures between semantic instances in the same scene and across different scenes, the spatial distribution features and the semantic categorical features are learned with graph convolutional networks, and the global geometric shape features are learned with a PointNet-like network. These three kinds of features are further enhanced with the self-attention and cross-attention mechanisms. Thirdly, the semantic instance matching is formulated as an optimal transport problem, and solved through an optimal matching layer. Finally, the geometric transformation matrix between two point clouds is first estimated by the SVD algorithm and then refined by the ICP algorithm. Experimental results conducted on the KITTI Odometry dataset demonstrate that the proposed method improves the registration performance and outperforms various state-of-the-art methods.
Shaocong Liu, Tao Wang 0075, Yan Zhang 0159, Ruqin Zhou, Li Li 0100, Chenguang Dai, Longguang Wang, Hanyun Wang
IEEE Trans. Geosci. Remote. Sens.6
2024 Combined Adjustment for Very High-Resolution Satellite Stereo Images and ICESat-2 Laser Altimetry Data
abstract
As we all know, the combined block adjustment of very high-resolution satellite stereo images (VHRSIs) and spaceborne laser altimetry data (Ice, Cloud, and land Elevation Satellite-2 (ICESat-2) data) can provide support for 1:10000 topographic mapping without the use of ground control points (GCPs). However, there are still many challenges in the combined adjustment processing, such as the selection of elevation control points and the registration of two types of data. To address these issues in combined adjustment, a preliminary method for selecting elevation control points is proposed based on various laser attribute information, building utilizing an analysis of ATL08 product characteristics. Furthermore, virtual planar control points and elevation checkpoints are introduced, and a combined block adjustment strategy is designed to refine the application of laser elevation control points and resolve registration issues between the stereo image data and laser altimetry data during the adjustment process. The proposed strategy was implemented and evaluated across three representative experimental areas. In the experiments, the elevation root-mean-square errors (RMSEs) of processed stereo images in all three areas were found to be below 1 m, while the planimetric accuracy remained consistent with the results obtained from the free network adjustment. The findings suggest that the developed strategy effectively addresses the challenges associated with elevation control point selection and data registration in the combined adjustment process, significantly enhancing elevation accuracy while preserving the planar accuracy of stereo images. Therefore, the proposed combined block adjustment strategy supports the endeavor of global 1:10000 mapping without the need for GCPs and promotes the application of VHRSI and ICESat-2 data in land surface surveying and mapping.
Dashuai Shang, Chenguang Dai
IEEE Trans. Geosci. Remote. Sens.3
2023 Scene Overlap Prediction for LiDAR-Based Place Recognition
abstract
Recently, LiDAR (Light Detection and Ranging)-based place recognition, has been widely concerned because of its robustness to light conditions, seasonal changes, and viewpoint variations. Unlike most of existing methods which represent the whole point cloud scenes with global descriptors, we treat the LiDAR-based place recognition problem as a scene overlap prediction task and propose an end-to-end overlap prediction network, which consists of a feature learning backbone, a feature enhancement module, and an overlap prediction module. Based on the prediction result for each point, the overlapping ratios between two point clouds are computed and used to predict whether these two point clouds are at the same place. In addition, to promote the computational efficiency and reduce the model complexity, a lightweight feature learning backbone is also adopted. The experiments conducted on the KITTI Odometry dataset demonstrate that the proposed method achieves superior performance compared with state-of-the-art methods. The lightweight method also obtains 2x inference speed with little performance degradation compared with the vanilla method.
Yingjian Zhang, Chenguang Dai, Ruqin Zhou, Zhenchao Zhang 0001, Hongliang Ji, Huixin Fan, Hanyun Wang
IEEE Geosci. Remote. Sens. Lett.2
2022 The First Challenge on Moving Object Detection and Tracking in Satellite Videos: Methods and Results
abstract
In this paper, we briefly summarize the first challenge on moving object detection and tracking in satellite videos (SatVideoDT). This challenge has three tracks related to satellite video analysis, including moving object detection (Track 1), single object tracking (Track 2), and multiple-object tracking (Track 3). 123, 89, and 70 participants successfully registered, while 37, 42, and 29 teams submitted their final results on the test datasets for Tracks 1-3, respectively. The top-performing methods and their results in each track are described with details. This challenge establishes a new benchmark for satellite video analysis.
Yulan Guo, Qingyong Hu, Feng Zhang 0046, Ye Zhang 0037, Hanyun Wang, Chenguang Dai, Weilong Guo, Xiyu Qi, Kelong Tu, Shudan Zhu, Lai Chen, Bin Lin 0013, Chaocan Xue, Jinlei Zheng, Limei Qin, Ying Li 0017, Manqi Zhao, Lu Ruan 0003, Mingpeng Cui, Guanchen Ding, Guangwei Jiang, Zhenzhong Chen 0001, Kaiyang Cao, Lingyu Kong, Shaodong Chen, Zhicheng Zhao 0001, Qin Shen, Lei Liu 0049, Chenglong Li 0002, Yun Xiao 0003
ICPR8
2022 MaskNet++: Inlier/outlier identification for two point clouds
Ruqin Zhou, Hanyun Wang, Xixing Li, Yulan Guo, Chenguang Dai, Wanshou Jiang
Comput. Graph.5
2022 Contour deformation network for instance segmentation
Kefeng Lv, Hanyun Wang, Huaigang Jiang, Chenguang Dai
Pattern Recognit. Lett.7
2022 Extraction Strategy for ICESat-2 Elevation Control Points Based on ATL08 Product
abstract
ICESat-2 can obtain high-precision three-dimensional measurement information of targets and has a unique advantage in determining global elevation control points. However, due to the influence of the atmospheric environment, target characteristics, hardware equipment, and other factors, its elevation accuracy is not highly reliable, and not all data points can be used as control points. To obtain high-precision elevation control points from ICESat-2 data products, an extraction strategy that combines the accuracy and location requirements for control points was developed based on the ATL08 product. Multiple attribute parameters were incorporated into the extraction procedure, including terrain factor information, segment elevation information, cloud confidence flag, surface coverage data, topographic photon quantity, and photon height difference information. The extraction approach was then performed and analyzed using experimental data from the Hanzhong area in Shanxi province and the Songshan area in Henan province. In the experiments, the root-mean-square error (RMSE) of the extracted data points in two areas were both about 0.5 m and could reach 0.3 m after eliminating gross error points caused by surface changes and misclassification. The results suggest that the developed strategy takes into account the point location requirements of control points, can overcome the influence of complex terrain and ground objects and significantly improve the overall accuracy of obtained control points. Therefore, the proposed extraction strategy can support the establishment of the global elevation control point database and promote the application of ICESat-2 data in land surface surveying and mapping.
Dashuai Shang, Chenguang Dai, Qifang Ma, Ziquan Wang
IEEE Trans. Geosci. Remote. Sens.3
2020 NGM: Neural Gaussian Mirror for Controlled Feature Selection in Neural Networks
abstract
Deep neural networks (DNNs) have become increasingly popular and achieved outstanding performance in predictive tasks. However, the DNN framework itself cannot inform the user which features are more or less relevant for making the prediction, which limits its applicability in many scientific fields. We introduce neural Gaussian mirrors (NGMs), in which mirrored features are created, via a structured perturbation based on a kernel-based conditional dependence measure, to help evaluate feature importance. We design two modifications of the DNN architecture for incorporating mirrored features and providing mirror statistics to measure feature importance. As shown in simulated and real data examples, the proposed method controls the feature selection error rate at a predefined level and maintains a high selection power even with the presence of highly correlated features.
Yu Gui, Chenguang Dai, Jun S. Liu
ICMLA3