EDBT 2026 Demo / reviewers in the wild / expert
Yi Wan 0001
dblp:63/506-1
· DBLP profile ↗
22ranked-venue papers
0as first author
19since 2021 · last 2026
0000-0001-6777-6047ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SkySplat: Generalizable 3D Gaussian Splatting from Multi-Temporal Sparse Satellite ImagesabstractThree-dimensional scene reconstruction from sparse-view satellite images is a long-standing and challenging task. While 3D Gaussian Splatting (3DGS) and its variants have recently attracted attention for its high efficiency, existing methods remain unsuitable for satellite images due to incompatibility with rational polynomial coefficient (RPC) models and limited generalization capability. Recent advances in generalizable 3DGS approaches show potential, but they perform poorly on multi-temporal sparse satellite images due to limited geometric constraints, transient objects, and radiometric inconsistencies. To address these limitations, we propose SkySplat, a novel self-supervised framework that integrates the RPC model into the generalizable 3DGS pipeline, enabling more effective use of sparse geometric cues for improved reconstruction. SkySplat relies only on RGB images and radiometric-robust relative height supervision, thereby eliminating the need for ground-truth height maps. Key components include a Cross-Self Consistency Module (CSCM), which mitigates transient object interference via consistency-based masking, and a multi-view consistency aggregation strategy that refines reconstruction results. Compared to per-scene optimization methods, SkySplat achieves an 86 times speedup over EOGS with higher accuracy. It also outperforms generalizable 3DGS baselines, reducing MAE from 13.18 m to 1.80 m on the DFC19 dataset significantly, and demonstrates strong cross-dataset generalization on the MVS3D benchmark. Xuejun Huang, Xinyi Liu 0002, Yi Wan 0001, Bin Zhang 0046, Mingtao Xiong, Yingying Pei, Yongjun Zhang 0002 |
AAAI | 3 |
| 2026 | FreNTS: Neural Texture Synthesis in Frequency DomainabstractAlthough existing texture synthesis methods perform well in generating large images with irregularly repeated textures to avoid visually unrealistic repetitions, they still face significant challenges in synthesizing regular textures with densely interconnected structures. In this paper, we propose a novel neural texture synthesis method, FreNTS, which uses frequency domain information to enhance the texture synthesis process, synthesizing textures with continuous, complete, and visually realistic overall structures. The core idea is to perform the Discrete Cosine Transform on image patches to obtain the corresponding frequency domain rate information features, and then use the designed adaptive guided correspondence (AGC) loss to calculate the correlation difference between the source image and the target image in the frequency domain and spatial domains, thereby constraining the optimization of the target image to achieve high-quality texture synthesis. In addition, to better evaluate the effect of texture synthesis, we introduce Tile LPIPS as the metric for quantitative evaluation. Experimental results show that the proposed FreNTS can effectively accelerate the process of neural texture synthesis and use high-frequency information to capture better structural details to synthesize realistic textures. Dongdong Yue, Xinyi Liu 0002, Yongjun Zhang 0002, Zeshuang Zheng, Yi Wan 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | CasP: Improving Semi-Dense Feature Matching Pipeline Leveraging Cascaded Correspondence Priors for GuidanceabstractSemi-dense feature matching methods have shown strong performance in challenging scenarios. However, the existing pipeline relies on a global search across the entire feature map to establish coarse matches, limiting further improvements in accuracy and efficiency. Motivated by this limitation, we propose a novel pipeline, CasP, which leverages cascaded correspondence priors for guidance. Specifically, the matching stage is decomposed into two progressive phases, bridged by a region-based selective cross-attention mechanism designed to enhance feature discriminability. In the second phase, one-to-one matches are determined by restricting the search range to the one-to-many prior areas identified in the first phase. Additionally, this pipeline benefits from incorporating high-level features, which helps reduce the computational costs of low-level feature extraction. The acceleration gains of CasP increase with higher resolution, and our lite model achieves a speedup of $\sim2.2\times$ at a resolution of 1152 compared to the most efficient method, ELoFTR. Furthermore, extensive experiments demonstrate its superiority in geometric estimation, particularly with impressive cross-domain generalization. These advantages highlight its potential for latency-sensitive and high-robustness applications, such as SLAM and UAV systems. Code is available at https://github.com/pq-chen/CasP. Peiqi Chen, Lei Yu 0005, Yi Wan 0001, Yingying Pei, Xinyi Liu 0002, Yongxiang Yao, Lixiang Ru, Liheng Zhong, Jingdong Chen, Ming Yang 0007, Yongjun Zhang 0002 |
ICCV | 3 |
| 2025 | StereoINR: Cross-View Geometry Consistent Stereo Super Resolution with Implicit Neural RepresentationabstractStereo image super-resolution (SSR) aims to enhance high-resolution details by leveraging information from stereo image pairs. However, existing stereo super-resolution (SSR) upsampling methods (e.g., pixel shuffle) often overlook cross-view geometric consistency and are limited to fixed-scale upsampling. The key issue is that previous upsampling methods use convolutions to independently process deep features of different views, lacking cross-view and non-local information perception, making it difficult to select beneficial information from multi-view scenes adaptively. In this work, we propose Stereo Implicit Neural Representation (StereoINR), which innovatively models stereo image pairs as continuous implicit representations. This continuous representation breaks through the scale limitations, providing a unified solution for arbitrary-scale stereo super-resolution reconstruction of left-right views. Furthermore, by incorporating spatial warping and cross-attention mechanisms, StereoINR enables effective cross-view information fusion and achieves significant improvements in pixel-level geometric consistency. Extensive experiments on multiple datasets demonstrate that StereoINR outperforms out-of-training-distribution scale upsampling and matches state-of-the-art SSR methods within training-distribution scales. Xinyi Liu 0002, Yi Wan 0001, Panwang Xia, Yongjun Zhang 0002 |
ACM Multimedia | 3 |
| 2025 | Confidence-Aware Superpixel Self-Supervised Learning for Cross-Modal Point Cloud RecognitionabstractThe high cost of 3D annotation severely limits the availability of high-quality labeled data, posing a critical bottleneck to point cloud recognition tasks. In contrast, acquiring and annotating natural images is far more cost-effective, making 2D image-based methods more accessible. With the rise of large-scale pre-trained 2D feature extraction models, which have demonstrated remarkable embedding capabilities across diverse domains, an opportunity emerges to leverage them for 3D point cloud tasks. By transferring knowledge from well-established 2D models, reliance on 3D annotation can be significantly reduced while improving feature extraction efficiency. However, adapting these 2D models, originally trained on natural images, to remote sensing data introduces several challenges. In particularly, the generalization ability of 2D models is limited, and widely used point-pixel constraint methods suffer from high computational complexity and low error tolerance. To address these issues, this letter proposes the confidence-aware superpixel self-supervised learning method (CS-SSP). By incorporating confidence estimation, CS-SSP enhances the reliability of cross-domain feature embeddings, ensuring accurate feature transfer from natural images to remote sensing data. Additionally, a superpixel-based constraint is introduced to reduce computational complexity. Moreover, CS-SSP employs a coarse-to-fine training strategy, progressively refining constraints to mitigate training difficulties. The experiments demonstrate that CS-SSP outperforms baseline methods across all metrics. Notably, when fine-tuning with extremely limited labeled samples, CS-SSP achieves the state-of-the-art accuracy across all key metrics, highlighting its strong advantage in low-data scenarios. Yameng Wang, Bin Zhang 0046, Yi Wan 0001, Zhaoxi Yue, Yongjun Zhang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | MVSR3D: An End-to-End Framework for Semantic 3-D Reconstruction Using Multiview Satellite ImageryabstractSemantic 3D reconstruction from multi-view images is essential for applications such as 3D city modeling and robot navigation. However, existing methods treat semantic segmentation and height estimation as separate tasks, leading to suboptimal reconstruction results. To bridge this gap, we introduce MVSR3D, the first end-to-end framework for semantic 3D reconstruction using multi-view satellite images. MVSR3D employs a dual-stream architecture, consisting of the segmentation branch (MVSAM) based on Segment Anything Model (SAM) and the height estimation branch based on multi-view stereo. To enhance multi-view feature fusion, we propose the Epipolar Cross Attention (ECA) module in the MVSAM branch, which integrates image embeddings primarily along epipolar line to exploit complementary multi-view information. Unlike conventional multi-task learning approaches, we design dedicated interaction modules—the SAM Feature-Guided (SAM-FG) module and the Elevation-Guided Sparse Prompts Generator (EGSPG)—to facilitate multi-task interaction and feature fusion. Extensive evaluations on the DFC19 and SpaceNet4 datasets demonstrate that MVSR3D significantly outperforms the state-of-the-art multi-view multi-task learning method, improving the mIoU3 metric at a 2.5-meter threshold by 37.09%–45.11%. Xuejun Huang, Xinyi Liu 0002, Yi Wan 0001, Bin Zhang 0046, Yameng Wang, Yongjun Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | GS2Poly: Textured Polygonal Building Reconstruction Guided by Gaussian Opacity FieldsabstractCompact low-poly building models with concise structures and texture fidelity are essential infrastructure for digital twin cities. Traditional point cloud-based reconstruction methods often rely on surface normals, and the presence of missing data and noise poses significant challenges for accurate reconstruction. In this paper, we propose GS2Poly, a textured polygonal mesh reconstruction method for buildings based on the 3D Gaussian Splatting (3DGS) framework. Firstly, 3DGS of the building scene is reconstructed under planar structure constraints. A density-weighted Gaussian sampling method is utilized to sample high-quality surface point clouds and extract planar primitives from 3DGS reconstruction results. Next, GS2Poly applies an adaptive spatial partitioning strategy to generate a set of candidate convex polyhedra. Finally, guided by the Gaussian opacity field, a Markov random field is constructed to extract the polygonal mesh surface, followed by high-fidelity texture mapping using an optimal rendering strategy. Experimental results across diverse building scenarios demonstrate that GS2Poly exhibits higher geometric fidelity than spatial partitioning-based or 3DGS-based mesh simplification methods. Additionally, the proposed texture mapping strategy effectively avoids typical texture artifacts such as occlusion, seams and distortions. Xinyi Liu 0002, Weiwei Fan, Yongjun Zhang 0002, Zexu Zhang, Yi Wan 0001, Dongdong Yue, Jiachen Zhong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Automatic On-Orbit Geometric Calibration for High-Resolution Optical Satellites by Distance Transformation ModelabstractOn-orbit geometric calibration is crucial for generating high-quality satellite products. Existing methods are often costly or complex, requiring manual control points or specialized orbit relations. This paper introduces a novel approach using accurate 3D contours from city-level aerial LiDAR point clouds. Our method, leveraging open-source LiDAR data, is both free and highly efficient. We employ a distance-transformation-based registration model and an iterative robust solver for patch-level image-point-cloud registration and global optimization. During the registration process, an a-contrario judgment model helps detect and filter out false results. Experiments with Santa Clara and Las Vegas data, using quality-level-1 LiDAR point clouds from the 3D Elevation Program of the US Geological Survey, showed that sub-meter resolution satellite images acquired by WorldView-2, GeoEye-2, and the newly launched Wuhan-1 satellite can achieve subpixel accuracy via the proposed automatic on-orbit calibration approach without any manual operation. Applied to future satellite clusters, the proposed approach has great potential for fast on-orbit production of remote sensing products and significantly reducing the maintenance costs of optical satellites. Yi Wan 0001, Yongjun Zhang 0002, Shuangming Zhao, Yansong Duan, Mingtao Xiong, Zhonghua Hu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Multimodal Remote Sensing Image Robust Matching Based on Second-Order Tensor Orientation Feature TransformationabstractNonrigid deformation (NRD) and image noise in multimodal remote sensing images (MRSI) lead to abrupt changes in feature directions, resulting in sensitivity to rotational variation, sparse correct matches, and high false match rates. In order to address these challenges, this article proposes a second-order tensor orientation feature transformation (SOFT) method to improve the rotational invariance of MRSI matching and increase the number of correct matches (NCMs). The SOFT method has two main contributions: 1) a novel second-order tensor orientation descriptor is constructed by generating a tensor orientation feature map using a designed second-order tensor function, which is then combined with a gradient location and orientation histogram (GLOH)-like descriptor framework to achieve robust rotational invariance in multimodal image matching and 2) an error-removal global-local iterative optimization (EGIO) is introduced, employing a skewness of mixed pixel intensity (SMPI) function to automatically select matching seed points, followed by an iterative partition optimization strategy for refining corresponding points. Experiments on 744 groups of typical MRSIs demonstrate that the SOFT method significantly outperforms nine state-of-the-art methods, achieving an average 97% improvement in the NCMs, an average 25.51% improvement in the rate of correct matches (RCMs), and an average reduction in RMSE of 2.69 pixels. The proposed SOFT method, thus, offers robust MRSI matching with strong rotational invariance and precise identification of corresponding points, proving its effectiveness for complex remote sensing scenarios. Access to experiment-related data and codes will be provided athttps://skyearth.org/research/. Yongjun Zhang 0002, Peihao Wu, Yongxiang Yao, Yi Wan 0001, Wenfei Zhang, Yansheng Li 0001, Xiaohu Yan |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | EcoMatcher: Efficient Clustering Oriented Matcher for Detector-Free Image Matching
Peiqi Chen, Lei Yu 0005, Yi Wan 0001, Yongjun Zhang 0002, Jian Wang 0108, Liheng Zhong, Jingdong Chen, Ming Yang 0007 |
ECCV (68) | 3 |
| 2024 | Enhancing Cross-View Geo-Localization With Domain Alignment and Scene ConsistencyabstractCross-View Geo-Localization task is aimed at establishing correspondences between images captured from different perspectives within the same geographical region. The major challenge lies in the significant appearance variations of the same scene in different views. Current methods predominantly rely on learning a representation of the coarse-grained information from images and then evaluating the similarity, while the fine-grained features are usually not well-treated. In this paper, a novel method, named DAC (Domain Alignment and scene Consistency) is proposed, which leverages contrastive learning to acquire the global information of images and simultaneously employs a domain space alignment module to align the fine-grained features. The comprehensive utilization of multi-grained vision information guarantees better feature representations. Additionally, a cross-batch scene consistency strategy is proposed in the network to establish the global supervision of the positive samples based on scene correspondence, which improves the distinctiveness of the image representations. Advanced performance is shown by our method in drone-view target localization and drone navigation applications, outperforming state-of-the-art methods in comprehensive tests on the popular public datasets University-1652 and SUES-200. Our method also outperforms existing methods in cross-region localization, showing an average improvement of 5.6% in the R@1. Our codes and models are available athttps://github.com/SummerpanKing/DAC. Panwang Xia, Yi Wan 0001, Yongjun Zhang 0002, Jiwei Deng |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | CAMP: A Cross-View Geo-Localization Method Using Contrastive Attributes Mining and Position-Aware PartitioningabstractCross-view geo-localization (CVGL) task aims to utilize geographic data, such as maps or high-resolution satellite images, as reference to estimate the positions of a ground- or near-ground- captured query image. This task is particularly challenging due to the significant changes in visual appearance resulting from the extreme viewpoint variations. To address this challenge, a range of innovative methods have been proposed. However, intra-scene geometric information and inter-scene discriminative representation are not fully explored. In this article, we propose a novel CVGL method using contrastive attributes mining and position-aware partitioning (CAMP), which incorporates a position-aware partition branch (PPB) and a contrastive attributes mining (CAM) strategy. PPB learns fine-grained local features of different parts and captures their spatial information, providing a comprehensive understanding of scenes from both textual and spatial perspectives. CAM establishes supervision of the negative samples based on the images from the same platform, empowering the model to better discern differences between distinct scenes without extra memory cost. The proposed CAMP surpasses existing methods, achieving state-of-the-art results on the satellite-drone CVGL datasets University-1652 and SUES-200. Additionally, our method also outperforms existing methods in cross-dataset generalization, achieving an 8.85% increase in R@1 when trained on the University-1652 dataset and tested on the SUES-200 dataset at a height of 150 m. Our code and model are available athttps://github.com/Mabel0403/CAMP. Yi Wan 0001, Yongjun Zhang 0002, Guangshuai Wang, Zhenyang Zhao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Semantic Information-Aided Geometric Correction of High Resolution Satellite ImagesabstractWith the rapid development of deep learning technology, the automatic classification and semantic segmentation of remote sensing imagery become more and more accurate. Meanwhile, with the improvement of the generalization, this technology is more and more widely used in the industry [1] . This paper proposes a new framework of sematic-information-aided geometric correction of high-resolution satellite images (HRSIs) to achieve higher accuracy and automation. Yi Wan 0001, Yongjun Zhang 0002 |
IGARSS | 2 |
| 2023 | M2-CDNet: A Multi-Scale and Multi-Level Network for Remote Sensing Image Change DetectionabstractChange detection plays a crucial role in environmental monitoring and earth observation tasks, leveraging the abundant data acquired by remote sensing platforms. While learning-based methods have shown promise in strictly registered datasets, their practical applicability in real-world scenarios remains challenging. This paper addresses the limitations of existing methods by proposing M2-CDNet, a novel approach that integrates the U-Net architecture with the multi-scale fusion (MSF) strategy, deformable convolutions, and multi-scale outputs. Experiments on the public and self-collected datasets demonstrate that M2-CDNet achieves superior accuracy-efficiency trade-offs compared to state-of-the-art methods. Moreover, M2-CDNet shows better robustness against image projection bias and registration errors. Yi Wan 0001, Yongjun Zhang 0002 |
IGARSS | 3 |
| 2023 | CloudViT: A Lightweight Vision Transformer Network for Remote Sensing Cloud DetectionabstractClouds inevitably exist in satellite images, which limit the processing and application of satellite images to a certain extent. Therefore, cloud detection is a preprocessing task in satellite image extraction and analysis processing. However, the existing methods are difficult to mine robust features, and the number of parameters and computation are large, which is not conducive to the deployment of the model. In this letter, cloud vision transformer (CloudViT), a lightweight vision transformer network for cloud detection from satellite imagery, is proposed. In detail, to utilize dark channel priors in multispectral imagery to guide the network to learn features, a multiscale dark channel extractor is used to first predict dark channels, and then, the dark channel features and image features are input to the attention mechanism-based dark channel-guided context aggregation module to enhance image features, which in turn makes cloud detection results more accurate. At the same time, to enhance the transfer ability of the network between different satellite sensors, a plug-and-play channel adaptive module is proposed to deal with the inconsistency of the number of different satellite sensor bands. The experimental results on the Landsat7 dataset show that our network CloudViT outperforms the state-of-the-art methods while keeping the number of parameters and computation small. At the same time, the experimental results on transfer to three other datasets show that using the channel adaptation module can greatly improve the transfer ability of the model. Bin Zhang 0046, Yongjun Zhang 0002, Yansheng Li 0001, Yi Wan 0001, Yongxiang Yao |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | ELSR: Efficient Line Segment Reconstruction with Planes and Points GuidanceabstractThree-dimensional (3D) line segments are helpful for scene reconstruction. Most of the existing 3D-line-segment reconstruction algorithms deal with two views or dozens of small-size images; while in practice there are usually hundreds or thousands of large-size images. In this paper, we propose an efficient line segment reconstruction method called ELSR11Available at https://skyearth.org/publication/project/ELSR. ELSR exploits scene planes that are commonly seen in city scenes and sparse 3D points that can be acquired easily from the structure-from-motion (SfM) approach. For two views, ELSR efficiently finds the local scene plane to guide the line matching and exploits sparse 3D points to accelerate and constrain the matching. To reconstruct a 3D line segment with multiple views, ELSR utilizes an efficient abstraction approach that selects representative 3D lines based on their spatial consistence. Our experiments demonstrated that ELSR had a higher accuracy and efficiency than the existing methods. Moreover, our results showed that ELSR could reconstruct 3D lines efficiently for large and complex scenes that contain thousands of large-size images. Yi Wan 0001, Yongjun Zhang 0002, Xinyi Liu 0002, Bin Zhang 0046, Xiqi Wang |
CVPR | 2 |
| 2022 | SCTRANS: A TRANSFORMER NETWORK BASED ON THE SPATIAL AND CHANNEL ATTENTION FOR CLOUD DETECTIONabstractCloud detection is an important preprocessing step for remote sensing image processing and analysis. The current deep-learning-based cloud detection methods are mostly based on Convolutional Neural Network (CNN) which pay more attention to local information. To make more use of the global information, in this article, we propose a transformer-based cloud detection method (SCTrans) based on the spatial and channel attention mechanism. The experiment results show that when using only three-band images on the Landsat7 dataset, the mIoU of the validation set reaches 85.92% and the mIoU of the test set reaches 87.86%. The experimental results show that the proposed network has a higher mIoU and F1 score than Fmask and other networks. Wenke Jiao, Yongjun Zhang 0002, Bin Zhang 0046, Yi Wan 0001 |
IGARSS | 4 |
| 2022 | A Cascaded Cross-Modal Network for Semantic Segmentation from High-Resolution Aerial Imagery and RAW Lidar DataabstractAs various sensors appear, extracting information from multimodal data becomes a prominent topic. Current multimodal approaches for image and LiDAR normally discard the point-to-point topology relationship of the latter to keep the dimension matched. To tackle this task, we propose a cascaded cross-modal network (CCMN) to extract the joint-features from high-resolution aerial imagery and LiDAR point directly, instead of their abridged derivatives. Firstly, point-wise features are extract from raw LiDAR data by a forepart 3D extractor. Subsequently, the LiDAR-derived features are executed spatial reference conversion to project and align to the imagery coordinate space. Finally, the cross-modal compounds containing the obtained feature maps and the corresponding images are placed into a U-shape structure to generate segmentation results. The experiment results indicate that our strategy surpasses the popular multimodal method by 6% on mIoU. Yameng Wang, Bin Zhang 0046, Yi Wan 0001, Yongjun Zhang 0002 |
IGARSS | 3 |
| 2022 | Multi-Modal Remote Sensing Image Matching Considering Co-Occurrence FilterabstractTraditional image feature matching methods cannot obtain satisfactory results for multi-modal remote sensing images (MRSIs) in most cases because different imaging mechanisms bring significant nonlinear radiation distortion differences (NRD) and complicated geometric distortion. The key to MRSI matching is trying to weakening or eliminating the NRD and extract more edge features. This paper introduces a new robust MRSI matching method based on co-occurrence filter (CoF) space matching (CoFSM). Our algorithm has three steps: (1) a new co-occurrence scale space based on CoF is constructed, and the feature points in the new scale space are extracted by the optimized image gradient; (2) the gradient location and orientation histogram algorithm is used to construct a 152-dimensional log-polar descriptor, which makes the multi-modal image description more robust; and (3) a position-optimized Euclidean distance function is established, which is used to calculate the displacement error of the feature points in the horizontal and vertical directions to optimize the matching distance function. The optimization results then are rematched, and the outliers are eliminated using a fast sample consensus algorithm. We performed comparison experiments on our CoFSM method with the scale-invariant feature transform (SIFT), upright-SIFT, PSO-SIFT, and radiation-variation insensitive feature transform (RIFT) methods using a multi-modal image dataset. The algorithms of each method were comprehensively evaluated both qualitatively and quantitatively. Our experimental results show that our proposed CoFSM method can obtain satisfactory results both in the number of corresponding points and the accuracy of its root mean square error. The average number of obtained matches is namely 489.52 of CoFSM, and 412.52 of RIFT. As mentioned earlier, the matching effect of the proposed method was significantly greater than the three state-of-art methods. Our proposed CoFSM method achieved good effectiveness and robustness. Executable programs of CoFSM and MRSI datasets are published: https://skyearth.org/publication/project/CoFSM/. Yongxiang Yao, Yongjun Zhang 0002, Yi Wan 0001, Xinyi Liu 0002, Xiaohu Yan, Jiayuan Li 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | DEM Extraction from Airborne Lidar Point Cloud in Thick-Forested Areas via Convolutional Neural NetworkabstractDigital Elevation Model (DEM), representing the height of the earth terrain, is one of the crucial geographic information products. One of the main data source of DEM is the airborne LiDAR point cloud with its non-ground-reflections filtered out. Point cloud filtering in thick-forested areas is difficult without enough ground control points when using conventional methods. In this paper, a supervised method is proposed to handle the problem of automatic DEM extraction with little ground control points. The design of the method is inspired by the successful application of the convolutional neural networks (CNN) in the image super resolution (SR) process. First, with the given LiDAR point cloud, the digital surface model (DSM) is resampled with regular grid. Then, by learning the spatial autocorrelation between the DSM and its corresponding DEM, a robust CNN model is established. Finally, the DEM in thick-forested areas can be generated from the DSM with the trained model. Experimental results at two different mountain sites in China validate the effectiveness of the proposed method of high-precision DEM generation. Yongjun Zhang 0002, Sizhe Xiang, Yi Wan 0001, Yimin Luo |
IGARSS | 3 |
| 2020 | Band-Independent Encoder-Decoder Network for Pan-Sharpening of Remote Sensing ImagesabstractPan-sharpening is a fundamental task for remote sensing image processing. It aims at creating a high-resolution multispectral (HRMS) image from a multispectral (MS) image and a panchromatic (PAN) image. In this article, a new band-independent encoder-decoder network is proposed for pan-sharpening. The network takes a single band of the MS (BMS) image, the PAN image, and the low-resolution PAN (LRPAN) image as inputs. The output of the network is the corresponding band of high-resolution MS (HRBMS) image. In this way, the network can process MS images with any number of bands. The overall structure of the network consists of two encoder-decoder modules at low-resolution and high-resolution, respectively. An auxiliary LRPAN image is used to speed up the training and improve the performance. The partly shared network and hierarchical structure for low-resolution and high-resolution enable a better fusion of features extracted from different scales. With a fast fine-tuning strategy, the trained model can be applied to images from different sensors. Experiments performed on different data sets demonstrate that the proposed method outperforms several state-of-the-art pan-sharpening methods in both visual appearance and objective indexes, and the single-band evaluation results further verify the superiority of the proposed method. Chi Liu 0004, Yongjun Zhang 0002, Shugen Wang, Yangjun Ou, Yi Wan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2016 | DEM-Assisted RFM Block Adjustment of Pushbroom Nadir Viewing HRS ImageryabstractNadir viewing satellite image is an effective data source to generate orthomosaics. Because of the georeferencing error of satellite images, block adjustment is the first step of orthomosaic generation over a large area. However, the geometric relationship of the neighboring orbits of the nadir viewing images is not rigid enough. This paper proposes a new rational function model (RFM) block adjustment approach that constrains the tie point elevation to enhance the relative geometric rigidity. By interpolating the elevations of tie points in a digital elevation model (DEM) and estimating the a priori errors of the interpolated elevations, better overall relative accuracy is obtained, and the local optimal solution problem is avoided. By constraining the adjusted model parameters according to the a priori error of RFMs, block adjustment without ground control point (GCP) is performed. By optimal initializing the object-space positions of tie points with multi-backprojection method, the needed iteration times of block adjustment are reduced. The proposed approach is investigated with 46 Ziyuan-3 sensor-corrected images, a 1:50 000 scale DEM, and 586 GCPs. Compared with Teo's approach that constrains the horizontal coordinates and elevations of tie points, the approach in this paper converges much faster when the GCPs are sparse, and meanwhile, the absolute and relative accuracy of the two approaches are almost the same. The result of block adjustment with only four GCPs shows that no accuracy degeneration occurred in the test area and the root-mean-square error of independent check point reaches about 1.5 ground resolutions. Different DEMs and number of tie points are used to investigate whether the block adjustment result is influenced by these factors. The results show that better DEM accuracy and denser tie points do improve the accuracy when the images have large side-sway angles. The proposed approach is also tested with 5118 IKONOS-2 images that cover the southern Europe without GCP. The result shows that the relative mosaicking accuracy is much better than that of Grodecki's approach. Yongjun Zhang 0002, Yi Wan 0001, Xinhui Huang |
IEEE Trans. Geosci. Remote. Sens. | 2 |