VLDB 2026 Research / reviewers in the wild / expert
Yongjun Zhang 0002
dblp:43/5828-2
· DBLP profile ↗
108ranked-venue papers
18as first author
62since 2021 · last 2026
0000-0001-9845-4251ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 72 · 14 first-author · 35 since 2021Artificial intelligence and machine learning · 26 · 4 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 14 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SkySplat: Generalizable 3D Gaussian Splatting from Multi-Temporal Sparse Satellite ImagesabstractThree-dimensional scene reconstruction from sparse-view satellite images is a long-standing and challenging task. While 3D Gaussian Splatting (3DGS) and its variants have recently attracted attention for its high efficiency, existing methods remain unsuitable for satellite images due to incompatibility with rational polynomial coefficient (RPC) models and limited generalization capability. Recent advances in generalizable 3DGS approaches show potential, but they perform poorly on multi-temporal sparse satellite images due to limited geometric constraints, transient objects, and radiometric inconsistencies. To address these limitations, we propose SkySplat, a novel self-supervised framework that integrates the RPC model into the generalizable 3DGS pipeline, enabling more effective use of sparse geometric cues for improved reconstruction. SkySplat relies only on RGB images and radiometric-robust relative height supervision, thereby eliminating the need for ground-truth height maps. Key components include a Cross-Self Consistency Module (CSCM), which mitigates transient object interference via consistency-based masking, and a multi-view consistency aggregation strategy that refines reconstruction results. Compared to per-scene optimization methods, SkySplat achieves an 86 times speedup over EOGS with higher accuracy. It also outperforms generalizable 3DGS baselines, reducing MAE from 13.18 m to 1.80 m on the DFC19 dataset significantly, and demonstrates strong cross-dataset generalization on the MVS3D benchmark. Xuejun Huang, Xinyi Liu 0002, Yi Wan 0001, Bin Zhang 0046, Mingtao Xiong, Yingying Pei, Yongjun Zhang 0002 |
AAAI | 8 |
| 2026 | Full-Scope Vectorization of Geographical Elements from Large-Size Remote Sensing ImageryabstractLarge-size very-high-resolution (VHR) remote sensing imagery has emerged as a critical data source for high-precision vector mapping of multi-scale geographical elements such as building, water, road and etc. When dealing with the large-size image, due to the limited memory of GPU, the deep learning-based vector mapping methods often employ the sliding block strategy. This inevitably leads to the degenerated performance because of the stitching difficulty of the sliding blocks' vector mapping results. Therefore, it is necessary to conduct full-scope vector mapping via mining the consistent cue in large-size remote sensing imagery. To this end, this paper presents a novel global context-aware local point optimization method. To leverage the global context, this paper proposes a novel pyramid fusion network (PFNet) to conduct semantic segmentation of the large-size image in an end-to-end manner. Under the constraint of the global semantic segmentation result, a new inflection-point perception network (IPNet) is proposed to generate a set of stable points to depict the boundary of each element. Extensive experiments on building, water and road datasets, where each image has over 100 million pixels, show that our method obviously outperforms the existing methods. Yansheng Li 0001, Wanchun Li, Bo Dang 0002, Yu Wang 0222, Wei Chen 0089, Bingnan Yang, Yongjun Zhang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2026 | Travel Like a Train: Detection and Reconstruction of Rail Tracks From Aerial Images
Yongjun Zhang 0002, Guangshuai Wang, Bin Zhang 0046, Jiwei Deng, Ziqian Huang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2026 | InfiniGS: Toward Efficient Ultra-High Resolution 3D Reconstruction via Gaussian Splattingabstract3D Gaussian Splatting (3DGS) has attracted considerable attention due to its remarkable rendering efficiency and superior visual quality. However, typical implementations struggle with ultra-high-resolution imagery due to excessive GPU memory demands, a critical limitation particularly in domains such as aerial mapping and remote sensing, where fine-grained details are crucial for accurate 3D reconstruction. In this paper, we introduce InfiniGS, an enhanced Gaussian Splatting framework that scales up training image resolution while maintaining fidelity and efficiency, enabling reconstruction from ultra-high-resolution images without encountering out-of-memory (OOM) issues. Our framework consists of two modules; firstly, we propose a Multi-Resolution Training strategy that guides the optimization from recovering global structures toward fine-grained details, while mitigating erosion artifacts during resolution transitions. Secondly, we introduce an Image Block Partitioning and Asynchronous Optimization scheme, which significantly reduces GPU memory overhead. Most importantly, we reveal the scaling rule for the densification threshold of Gaussian primitives and the hyperparameters of the Adam optimizer during resolution transitions and image partitioning. This ensures consistently high rendering quality regardless of the multi-resolution factor or partitioning strategy employed. Extensive experiments demonstrate that InfiniGS achieves superior Novel View Synthesis (NVS) for high-resolution scenes, while requiring less memory and shorter training time. To the best of our knowledge, InfiniGS is the first method capable of reconstructing scenes with resolutions up to 10 K and beyond, at full resolution and high fidelity, without exceeding the memory capacity of mainstream powerful GPUs (e.g., NVIDIA RTX 3090 with 24 GB VRAM). Xinyi Liu 0002, Yongyu Tian, Yuan Kou, Yongjun Zhang 0002 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2026 | FreNTS: Neural Texture Synthesis in Frequency DomainabstractAlthough existing texture synthesis methods perform well in generating large images with irregularly repeated textures to avoid visually unrealistic repetitions, they still face significant challenges in synthesizing regular textures with densely interconnected structures. In this paper, we propose a novel neural texture synthesis method, FreNTS, which uses frequency domain information to enhance the texture synthesis process, synthesizing textures with continuous, complete, and visually realistic overall structures. The core idea is to perform the Discrete Cosine Transform on image patches to obtain the corresponding frequency domain rate information features, and then use the designed adaptive guided correspondence (AGC) loss to calculate the correlation difference between the source image and the target image in the frequency domain and spatial domains, thereby constraining the optimization of the target image to achieve high-quality texture synthesis. In addition, to better evaluate the effect of texture synthesis, we introduce Tile LPIPS as the metric for quantitative evaluation. Experimental results show that the proposed FreNTS can effectively accelerate the process of neural texture synthesis and use high-frequency information to capture better structural details to synthesize realistic textures. Dongdong Yue, Xinyi Liu 0002, Yongjun Zhang 0002, Zeshuang Zheng, Yi Wan 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | HeMoRa: Unsupervised Heuristic Consensus Sampling for Robust Point Cloud RegistrationabstractHeuristic information for consensus set sampling is essential for correspondence-based point cloud registration, but existing approaches typically rely on supervised learning or expert-driven parameter tuning. In this work, we propose HeMoRa, a new unsupervised framework that trains a Heuristic information Generator (HeGen) to estimate sampling probabilities for correspondences using a Multi-order Reward Aggregator (MoRa) loss. The core of MoRa is to train HeGen through extensive trials and feedback, enabling unsupervised learning. While this process can be implemented using policy optimization, directly applying the policy gradient to optimize HeGen presents challenges such as sensitivity to noise and low reward efficiency. To address these issues, we propose a Maximal Reward Propagation (MRP) mechanism that enhances the training process by prioritizing noise-free signals and improving reward utilization. Experimental results show that equipped with HeMoRa, the consensus set sampler achieves improvements in both robustness and accuracy. For example, on the 3DMatch dataset with FCGF feature, the registration recall of our unsupervised methods (Ours+SM and Ours+SC2) even outperforms the state-of-the-art supervised method VBreg. Our code is available at HeMoRa. Shaocheng Yan, Kaiyan Zhao, Zhenjun Zhao, Yongjun Zhang 0002, Jiayuan Li 0001 |
CVPR | 6 |
| 2025 | CasP: Improving Semi-Dense Feature Matching Pipeline Leveraging Cascaded Correspondence Priors for GuidanceabstractSemi-dense feature matching methods have shown strong performance in challenging scenarios. However, the existing pipeline relies on a global search across the entire feature map to establish coarse matches, limiting further improvements in accuracy and efficiency. Motivated by this limitation, we propose a novel pipeline, CasP, which leverages cascaded correspondence priors for guidance. Specifically, the matching stage is decomposed into two progressive phases, bridged by a region-based selective cross-attention mechanism designed to enhance feature discriminability. In the second phase, one-to-one matches are determined by restricting the search range to the one-to-many prior areas identified in the first phase. Additionally, this pipeline benefits from incorporating high-level features, which helps reduce the computational costs of low-level feature extraction. The acceleration gains of CasP increase with higher resolution, and our lite model achieves a speedup of $\sim2.2\times$ at a resolution of 1152 compared to the most efficient method, ELoFTR. Furthermore, extensive experiments demonstrate its superiority in geometric estimation, particularly with impressive cross-domain generalization. These advantages highlight its potential for latency-sensitive and high-robustness applications, such as SLAM and UAV systems. Code is available at https://github.com/pq-chen/CasP. Peiqi Chen, Lei Yu 0005, Yi Wan 0001, Yingying Pei, Xinyi Liu 0002, Yongxiang Yao, Lixiang Ru, Liheng Zhong, Jingdong Chen, Ming Yang 0007, Yongjun Zhang 0002 |
ICCV | 12 |
| 2025 | StereoINR: Cross-View Geometry Consistent Stereo Super Resolution with Implicit Neural RepresentationabstractStereo image super-resolution (SSR) aims to enhance high-resolution details by leveraging information from stereo image pairs. However, existing stereo super-resolution (SSR) upsampling methods (e.g., pixel shuffle) often overlook cross-view geometric consistency and are limited to fixed-scale upsampling. The key issue is that previous upsampling methods use convolutions to independently process deep features of different views, lacking cross-view and non-local information perception, making it difficult to select beneficial information from multi-view scenes adaptively. In this work, we propose Stereo Implicit Neural Representation (StereoINR), which innovatively models stereo image pairs as continuous implicit representations. This continuous representation breaks through the scale limitations, providing a unified solution for arbitrary-scale stereo super-resolution reconstruction of left-right views. Furthermore, by incorporating spatial warping and cross-attention mechanisms, StereoINR enables effective cross-view information fusion and achieves significant improvements in pixel-level geometric consistency. Extensive experiments on multiple datasets demonstrate that StereoINR outperforms out-of-training-distribution scale upsampling and matches state-of-the-art SSR methods within training-distribution scales. Xinyi Liu 0002, Yi Wan 0001, Panwang Xia, Yongjun Zhang 0002 |
ACM Multimedia | 6 |
| 2025 | Confidence-Aware Superpixel Self-Supervised Learning for Cross-Modal Point Cloud RecognitionabstractThe high cost of 3D annotation severely limits the availability of high-quality labeled data, posing a critical bottleneck to point cloud recognition tasks. In contrast, acquiring and annotating natural images is far more cost-effective, making 2D image-based methods more accessible. With the rise of large-scale pre-trained 2D feature extraction models, which have demonstrated remarkable embedding capabilities across diverse domains, an opportunity emerges to leverage them for 3D point cloud tasks. By transferring knowledge from well-established 2D models, reliance on 3D annotation can be significantly reduced while improving feature extraction efficiency. However, adapting these 2D models, originally trained on natural images, to remote sensing data introduces several challenges. In particularly, the generalization ability of 2D models is limited, and widely used point-pixel constraint methods suffer from high computational complexity and low error tolerance. To address these issues, this letter proposes the confidence-aware superpixel self-supervised learning method (CS-SSP). By incorporating confidence estimation, CS-SSP enhances the reliability of cross-domain feature embeddings, ensuring accurate feature transfer from natural images to remote sensing data. Additionally, a superpixel-based constraint is introduced to reduce computational complexity. Moreover, CS-SSP employs a coarse-to-fine training strategy, progressively refining constraints to mitigate training difficulties. The experiments demonstrate that CS-SSP outperforms baseline methods across all metrics. Notably, when fine-tuning with extremely limited labeled samples, CS-SSP achieves the state-of-the-art accuracy across all key metrics, highlighting its strong advantage in low-data scenarios. Yameng Wang, Bin Zhang 0046, Yi Wan 0001, Zhaoxi Yue, Yongjun Zhang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2025 | STAR: A First-Ever Dataset and a Large-Scale Benchmark for Scene Graph Generation in Large-Size Satellite ImageryabstractScene graph generation (SGG) in satellite imagery (SAI) benefits promoting understanding of geospatial scenarios from perception to cognition. In SAI, objects exhibit great variations in scales and aspect ratios, and there exist rich relationships between objects (even between spatially disjoint objects), which makes it attractive to holistically conduct SGG in large-size very-high-resolution (VHR) SAI. However, there lack such SGG datasets. Due to the complexity of large-size SAI, mining triplets subject, relationship, object heavily relies on long-range contextual reasoning. Consequently, SGG models designed for small-size natural imagery are not directly applicable to large-size SAI. This paper constructs a large-scale dataset for SGG in large-size VHR SAI with image sizes ranging from 512 × 768 to 27,860 × 31,096 pixels, named STAR (Scene graph generaTion in lArge-size satellite imageRy), encompassing over 210K objects and over 400K triplets. To realize SGG in large-size SAI, we propose a context-aware cascade cognition (CAC) framework to understand SAI regarding object detection (OBD), pair pruning and relationship prediction for SGG. We also release a SAI-oriented SGG toolkit with about 30 OBD and 10 SGG methods which need further adaptation by our devised modules on our challenging STAR dataset. The dataset and toolkit are available at: https://linlin-dev.github.io/project/STAR. Yansheng Li 0001, Tingzhu Wang, Xue Yang 0005, Qi Wang 0009, Youming Deng, Xian Sun 0001, Haifeng Li 0007, Bo Dang 0002, Yongjun Zhang 0002, Yi Yu 0010, Junchi Yan |
IEEE Trans. Pattern Anal. Mach. Intell. | 12 |
| 2025 | MVSR3D: An End-to-End Framework for Semantic 3-D Reconstruction Using Multiview Satellite ImageryabstractSemantic 3D reconstruction from multi-view images is essential for applications such as 3D city modeling and robot navigation. However, existing methods treat semantic segmentation and height estimation as separate tasks, leading to suboptimal reconstruction results. To bridge this gap, we introduce MVSR3D, the first end-to-end framework for semantic 3D reconstruction using multi-view satellite images. MVSR3D employs a dual-stream architecture, consisting of the segmentation branch (MVSAM) based on Segment Anything Model (SAM) and the height estimation branch based on multi-view stereo. To enhance multi-view feature fusion, we propose the Epipolar Cross Attention (ECA) module in the MVSAM branch, which integrates image embeddings primarily along epipolar line to exploit complementary multi-view information. Unlike conventional multi-task learning approaches, we design dedicated interaction modules—the SAM Feature-Guided (SAM-FG) module and the Elevation-Guided Sparse Prompts Generator (EGSPG)—to facilitate multi-task interaction and feature fusion. Extensive evaluations on the DFC19 and SpaceNet4 datasets demonstrate that MVSR3D significantly outperforms the state-of-the-art multi-view multi-task learning method, improving the mIoU3 metric at a 2.5-meter threshold by 37.09%–45.11%. Xuejun Huang, Xinyi Liu 0002, Yi Wan 0001, Bin Zhang 0046, Yameng Wang, Yongjun Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | Unify Points and Lines for Efficient and Accurate Plane SegmentationabstractArtificial scenes often contain abundant planar and linear structures, which are essential for various remote sensing tasks. However, most existing plane segmentation methods primarily rely on point features, neglecting the structural and guiding roles of line features. To address this, we propose a unified RANSAC-based framework that integrates point and line features for efficient and robust plane segmentation. It introduces new sampling patterns for points and lines, guided by an adaptive probability model that dynamically adjusts the sampling strategy based on the quality of the generated candidate planes. Additionally, we design a boundary optimization strategy leveraging line constraints to improve segmentation accuracy near structural edges. In this way, the method is particularly suitable for urban scenarios where planar structures are often accompanied by structural lines, while still remaining applicable to raw point clouds in natural scenes without such features. To evaluate our method, we conducted extensive experiments on both synthetic and real-world datasets. The results show that it achieves an averageF1planeof 93.55% andF1pointof 96.2%, outperforming the best-performing baseline by about 1.4%, 1.2%, and 2.2% inF1plane,F1point, andmIoU, respectively, while also reducing runtime by approximately 18.8%. The core implementation of our method is publicly available at https://github.com/dreamer198/PSUPL. Ziqian Huang, Yongjun Zhang 0002, Xianzhang Zhu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | GS2Poly: Textured Polygonal Building Reconstruction Guided by Gaussian Opacity FieldsabstractCompact low-poly building models with concise structures and texture fidelity are essential infrastructure for digital twin cities. Traditional point cloud-based reconstruction methods often rely on surface normals, and the presence of missing data and noise poses significant challenges for accurate reconstruction. In this paper, we propose GS2Poly, a textured polygonal mesh reconstruction method for buildings based on the 3D Gaussian Splatting (3DGS) framework. Firstly, 3DGS of the building scene is reconstructed under planar structure constraints. A density-weighted Gaussian sampling method is utilized to sample high-quality surface point clouds and extract planar primitives from 3DGS reconstruction results. Next, GS2Poly applies an adaptive spatial partitioning strategy to generate a set of candidate convex polyhedra. Finally, guided by the Gaussian opacity field, a Markov random field is constructed to extract the polygonal mesh surface, followed by high-fidelity texture mapping using an optimal rendering strategy. Experimental results across diverse building scenarios demonstrate that GS2Poly exhibits higher geometric fidelity than spatial partitioning-based or 3DGS-based mesh simplification methods. Additionally, the proposed texture mapping strategy effectively avoids typical texture artifacts such as occlusion, seams and distortions. Xinyi Liu 0002, Weiwei Fan, Yongjun Zhang 0002, Zexu Zhang, Yi Wan 0001, Dongdong Yue, Jiachen Zhong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | AllSpark: A Multimodal Spatiotemporal General Intelligence Model With Ten Modalities via Language as a Reference FrameworkabstractRGB, multispectral, point and other spatio-temporal modal data fundamentally represent different observational approaches for the same geographic object. Therefore, leveraging multimodal data is an inherent requirement for comprehending geographic objects. However, due to the high heterogeneity in structure and semantics among various spatio-temporal modalities, the joint interpretation of multimodal spatio-temporal data has long been an extremely challenging problem. The primary challenge resides in striking a trade-off between the cohesion and autonomy of diverse modalities. This trade-off becomes progressively nonlinear as the number of modalities expands. Inspired by the human cognitive system and linguistic philosophy, where perceptual signals from the five senses converge into language, we introduce the Language as Reference Framework (LaRF), a fundamental principle for constructing a multimodal unified model. Building upon this, we propose AllSpark, a multimodal spatio-temporal general artificial intelligence model. Our model integrates ten different modalities into a unified framework, including one-dimensional (language, code, table), two-dimensional (RGB, SAR, multispectral, hyperspectral, graph, trajectory), and three-dimensional (point cloud) modalities. To achieve modal cohesion, AllSpark introduces a modal bridge and multimodal large language model (LLM) to map diverse modal features into the language feature space. To maintain modality autonomy, AllSpark uses modality-specific encoders to extract the tokens of various spatio-temporal modalities. Finally, observing a gap between the model’s interpretability and downstream tasks, we designed modality-specific prompts and task heads, enhancing the model’s generalization capability across specific tasks. Experiments indicate that the incorporation of language enables AllSpark to excel in few-shot classification tasks for RGB and point cloud modalities without additional training, surpassing baseline performance by up to 41.82%. Additionally, AllSpark, despite lacking expert knowledge in most spatio-temporal modalities and utilizing a unified structure, demonstrates strong adaptability across ten modalities. LaRF and AllSpark contribute to the shift in the research paradigm in spatio-temporal intelligence, transitioning from a modality-specific and task-specific paradigm to a general paradigm. The source code is available at https://github.com/GeoX-Lab/AllSpark. Run Shao, Qiujun Li, Linrui Xu, Qing Zhu 0012, Yongjun Zhang 0002, Yansheng Li 0001, Yu Liu 0003, Shizhong Yang, Haifeng Li 0007 |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2025 | Automatic On-Orbit Geometric Calibration for High-Resolution Optical Satellites by Distance Transformation ModelabstractOn-orbit geometric calibration is crucial for generating high-quality satellite products. Existing methods are often costly or complex, requiring manual control points or specialized orbit relations. This paper introduces a novel approach using accurate 3D contours from city-level aerial LiDAR point clouds. Our method, leveraging open-source LiDAR data, is both free and highly efficient. We employ a distance-transformation-based registration model and an iterative robust solver for patch-level image-point-cloud registration and global optimization. During the registration process, an a-contrario judgment model helps detect and filter out false results. Experiments with Santa Clara and Las Vegas data, using quality-level-1 LiDAR point clouds from the 3D Elevation Program of the US Geological Survey, showed that sub-meter resolution satellite images acquired by WorldView-2, GeoEye-2, and the newly launched Wuhan-1 satellite can achieve subpixel accuracy via the proposed automatic on-orbit calibration approach without any manual operation. Applied to future satellite clusters, the proposed approach has great potential for fast on-orbit production of remote sensing products and significantly reducing the maintenance costs of optical satellites. Yi Wan 0001, Yongjun Zhang 0002, Shuangming Zhao, Yansong Duan, Mingtao Xiong, Zhonghua Hu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Multimodal Remote Sensing Image Robust Matching Based on Second-Order Tensor Orientation Feature TransformationabstractNonrigid deformation (NRD) and image noise in multimodal remote sensing images (MRSI) lead to abrupt changes in feature directions, resulting in sensitivity to rotational variation, sparse correct matches, and high false match rates. In order to address these challenges, this article proposes a second-order tensor orientation feature transformation (SOFT) method to improve the rotational invariance of MRSI matching and increase the number of correct matches (NCMs). The SOFT method has two main contributions: 1) a novel second-order tensor orientation descriptor is constructed by generating a tensor orientation feature map using a designed second-order tensor function, which is then combined with a gradient location and orientation histogram (GLOH)-like descriptor framework to achieve robust rotational invariance in multimodal image matching and 2) an error-removal global-local iterative optimization (EGIO) is introduced, employing a skewness of mixed pixel intensity (SMPI) function to automatically select matching seed points, followed by an iterative partition optimization strategy for refining corresponding points. Experiments on 744 groups of typical MRSIs demonstrate that the SOFT method significantly outperforms nine state-of-the-art methods, achieving an average 97% improvement in the NCMs, an average 25.51% improvement in the rate of correct matches (RCMs), and an average reduction in RMSE of 2.69 pixels. The proposed SOFT method, thus, offers robust MRSI matching with strong rotational invariance and precise identification of corresponding points, proving its effectiveness for complex remote sensing scenarios. Access to experiment-related data and codes will be provided athttps://skyearth.org/research/. Yongjun Zhang 0002, Peihao Wu, Yongxiang Yao, Yi Wan 0001, Wenfei Zhang, Yansheng Li 0001, Xiaohu Yan |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | GLH-Water: A Large-Scale Dataset for Global Surface Water Detection in Large-Size Very-High-Resolution Satellite ImageryabstractGlobal surface water detection in very-high-resolution (VHR) satellite imagery can directly serve major applications such as refined flood mapping and water resource assessment. Although achievements have been made in detecting surface water in small-size satellite images corresponding to local geographic scales, datasets and methods suitable for mapping and analyzing global surface water have yet to be explored. To encourage the development of this task and facilitate the implementation of relevant applications, we propose the GLH-water dataset that consists of 250 satellite images and 40.96 billion pixels labeled surface water annotations that are distributed globally and contain water bodies exhibiting a wide variety of types (e.g. , rivers, lakes, and ponds in forests, irrigated fields, bare areas, and urban areas). Each image is of the size 12,800 × 12,800 pixels at 0.3 meter spatial resolution. To build a benchmark for GLH-water, we perform extensive experiments employing representative surface water detection models, popular semantic segmentation models, and ultra-high resolution segmentation models. Furthermore, we also design a strong baseline with the novel pyramid consistency loss (PCL) to initially explore this challenge, increasing IoU by 2.4% over the next best baseline. Finally, we implement the cross-dataset generalization and pilot area application experiments, and the superior performance illustrates the strong generalization and practical application value of GLH-water dataset. Project page: https://jack-bo1220.github.io/project/GLH-water.html Yansheng Li 0001, Bo Dang 0002, Wanchun Li, Yongjun Zhang 0002 |
AAAI | 4 |
| 2024 | SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation ImageryabstractPrior studies on Remote Sensing Foundation Model (RSFM) reveal immense potential towards a generic model for Earth Observation. Nevertheless, these works primar-ily focus on a single modality without temporal and geo-context modeling, hampering their capabilities for diverse tasks. In this study, we present SkySense, a generic billion-scale model, pretrained on a curated multimodal Remote Sensing Imagery (RSI) dataset with 21.5 million temporal sequences. SkySense incorporates a factorized multimodal spatiotemporal encoder taking temporal sequences of opti-cal and Synthetic Aperture Radar (SAR) data as input. This encoder is pretrained by our proposed Multi-Granularity Contrastive Learning to learn representations across different modal and spatial granularities. To further enhance the RSI representations by the geo-context clue, we introduce Geo-Context Prototype Learning to learn region-aware prototypes upon RSI's multimodal spatiotemporal features. To our best knowledge, SkySense is the largest Multi-Modal RSFM to date, whose modules can be flexibly combined or used individually to accommodate various tasks. It demonstrates remarkable generalization capabilities on a thor-ough evaluation encompassing 16 datasets over 7 tasks, from single- to multimodal, static to temporal, and classification to localization. SkySense surpasses 18 recent RSFMs in all test scenarios. Specifically, it outperforms the latest models such as GFM, SatLas and Scale-MAE by a large margin, i.e., 2.76%, 3.67% and 3.61% on average respectively. We will release the pretrained weights to facilitate future research and Earth Observation applications. Xin Guo 0010, Jiangwei Lao, Bo Dang 0002, Lei Yu 0005, Lixiang Ru, Liheng Zhong, Dingxiang Hu, Huimei He, Jian Wang 0108, Jingdong Chen, Ming Yang 0007, Yongjun Zhang 0002, Yansheng Li 0001 |
CVPR | 15 |
| 2024 | EcoMatcher: Efficient Clustering Oriented Matcher for Detector-Free Image Matching
Peiqi Chen, Lei Yu 0005, Yi Wan 0001, Yongjun Zhang 0002, Jian Wang 0108, Liheng Zhong, Jingdong Chen, Ming Yang 0007 |
ECCV (68) | 4 |
| 2024 | MMKDGAT: Multi-modal Knowledge graph-aware Deep Graph Attention Network for remote sensing image recommendation
Fei Wang 0094, Xianzhang Zhu, Yongjun Zhang 0002, Yansheng Li 0001 |
Expert Syst. Appl. | 4 |
| 2024 | Learning to Holistically Detect Bridges From Large-Size VHR Remote Sensing ImageryabstractBridge detection in remote sensing images (RSIs) plays a crucial role in various applications, but it poses unique challenges compared to the detection of other objects. In RSIs, bridges exhibit considerable variations in terms of their spatial scales and aspect ratios. Therefore, to ensure the visibility and integrity of bridges, it is essential to perform holistic bridge detection in large-size very-high-resolution (VHR) RSIs. However, the lack of datasets with large-size VHR RSIs limits the deep learning algorithms' performance on bridge detection. Due to the limitation of GPU memory in tackling large-size images, deep learning-based object detection methods commonly adopt the cropping strategy, which inevitably results in label fragmentation and discontinuous prediction. To ameliorate the scarcity of datasets, this paper proposes a large-scale dataset named GLH-Bridge comprising 6,000 VHR RSIs sampled from diverse geographic locations across the globe. These images encompass a wide range of sizes, varying from 2,048 × 2,048 to 16,384 × 16,384 pixels, and collectively feature 59,737 bridges. These bridges span diverse backgrounds, and each of them has been manually annotated, using both an oriented bounding box (OBB) and a horizontal bounding box (HBB). Furthermore, we present an efficient network for holistic bridge detection (HBD-Net) in large-size RSIs. The HBD-Net presents a separate detector-based feature fusion (SDFF) architecture and is optimized via a shape-sensitive sample re-weighting (SSRW) strategy. The SDFF architecture performs inter-layer feature fusion (IFF) to incorporate multi-scale context in the dynamic image pyramid (DIP) of the large-size image, and the SSRW strategy is employed to ensure an equitable balance in the regression weight of bridges with various aspect ratios. Based on the proposed GLH-Bridge dataset, we establish a bridge detection benchmark including the OBB and HBB tasks, and validate the effectiveness of the proposed HBD-Net. Additionally, cross-dataset generalization experiments on two publicly available datasets illustrate the strong generalization capability of the GLH-Bridge dataset. Yansheng Li 0001, Yongjun Zhang 0002, Yihua Tan, Jin-Gang Yu, Song Bai 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Enhancing Cross-View Geo-Localization With Domain Alignment and Scene ConsistencyabstractCross-View Geo-Localization task is aimed at establishing correspondences between images captured from different perspectives within the same geographical region. The major challenge lies in the significant appearance variations of the same scene in different views. Current methods predominantly rely on learning a representation of the coarse-grained information from images and then evaluating the similarity, while the fine-grained features are usually not well-treated. In this paper, a novel method, named DAC (Domain Alignment and scene Consistency) is proposed, which leverages contrastive learning to acquire the global information of images and simultaneously employs a domain space alignment module to align the fine-grained features. The comprehensive utilization of multi-grained vision information guarantees better feature representations. Additionally, a cross-batch scene consistency strategy is proposed in the network to establish the global supervision of the positive samples based on scene correspondence, which improves the distinctiveness of the image representations. Advanced performance is shown by our method in drone-view target localization and drone navigation applications, outperforming state-of-the-art methods in comprehensive tests on the popular public datasets University-1652 and SUES-200. Our method also outperforms existing methods in cross-region localization, showing an average improvement of 5.6% in the R@1. Our codes and models are available athttps://github.com/SummerpanKing/DAC. Panwang Xia, Yi Wan 0001, Yongjun Zhang 0002, Jiwei Deng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Scene Graph-Aware Hierarchical Fusion Network for Remote Sensing Image Retrieval With Text FeedbackabstractIn the realm of image retrieval with text feedback, existing studies have predominantly concentrated on the intrinsic attribute of target objects, neglecting extrinsic information essential for remote sensing (RS) images, such as spatial relationships. This research addresses this gap by incorporating RS image scene graphs as side information, given their capacity to encapsulate internal object attributes, external structural features between objects, and the relationships among images. To fully leverage the features from the reference RS image, scene graph, and modifier sentence, we propose a Scene graph-aware Hierarchical Fusion Network (SHF), which optimally integrates the multimodal features in a two-stage fusion process. Initially, image and scene graph features are fused hierarchically, followed by transforming content information with a proposed Multi-modal Global Content block (MGC), ultimately transforming style information. To validate the superiority of SHF, we constructed three datasets with images from several popular RS datasets, named Airplane (3461 image+text-image pairs), Tennis (1924 image+text-image pairs), and WHIRT (3344 image+text-image pairs). Extensive experiments conducted on these datasets show that SHF significantly outperforms state-of-the-art methods. Fei Wang 0094, Xianzhang Zhu, Xiaojian Liu 0004, Yongjun Zhang 0002, Yansheng Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | CAMP: A Cross-View Geo-Localization Method Using Contrastive Attributes Mining and Position-Aware PartitioningabstractCross-view geo-localization (CVGL) task aims to utilize geographic data, such as maps or high-resolution satellite images, as reference to estimate the positions of a ground- or near-ground- captured query image. This task is particularly challenging due to the significant changes in visual appearance resulting from the extreme viewpoint variations. To address this challenge, a range of innovative methods have been proposed. However, intra-scene geometric information and inter-scene discriminative representation are not fully explored. In this article, we propose a novel CVGL method using contrastive attributes mining and position-aware partitioning (CAMP), which incorporates a position-aware partition branch (PPB) and a contrastive attributes mining (CAM) strategy. PPB learns fine-grained local features of different parts and captures their spatial information, providing a comprehensive understanding of scenes from both textual and spatial perspectives. CAM establishes supervision of the negative samples based on the images from the same platform, empowering the model to better discern differences between distinct scenes without extra memory cost. The proposed CAMP surpasses existing methods, achieving state-of-the-art results on the satellite-drone CVGL datasets University-1652 and SUES-200. Additionally, our method also outperforms existing methods in cross-dataset generalization, achieving an 8.85% increase in R@1 when trained on the University-1652 dataset and tested on the SUES-200 dataset at a height of 150 m. Our code and model are available athttps://github.com/Mabel0403/CAMP. Yi Wan 0001, Yongjun Zhang 0002, Guangshuai Wang, Zhenyang Zhao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Scene Adaptive Building Individual Segmentation Based on Large-Scale Airborne LiDAR Point CloudsabstractBuilding individual segmentation plays a crucial role in building querying, management, analysis, and attribute addition. Previous research on this topic has primarily concentrated on small-scale scenes and single-type buildings. However, when dealing with complex scenes that contain diverse buildings, existing methods for building individual segmentation often encounter challenges, such as excessive undersegmentation and oversegmentation. To tackle this issue, we propose a scene adaptive building individual segmentation (SABIS) based on large-scale airborne LiDAR point clouds. The method first segments the roof object and then extract elevation feature and area feature of the roof object. Based on these features, the building point cloud is classified into two categories: urban scene buildings and rural residential scene buildings. Finally, for urban scene buildings, the building individual segmentation method based on the cylinder model consistency is used. For rural residential scene buildings, the building individual segmentation method based on bidirectional saliency features is employed. In this article, the proposed SABIS algorithm is quantitatively evaluated by using three large scene datasets at home and abroad and four benchmark methods. All kinds of accuracy are significantly better than the most advanced algorithms. Wangshan Yang, Yongjun Zhang 0002, Xinyi Liu 0002, Boyong Gao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Progressive Learning With Cross-Window Consistency for Semi-Supervised Semantic SegmentationabstractSemi-supervised semantic segmentation focuses on the exploration of a small amount of labeled data and a large amount of unlabeled data, which is more in line with the demands of real-world image understanding applications. However, it is still hindered by the inability to fully and effectively leverage unlabeled images. In this paper, we reveal that cross-window consistency (CWC) is helpful in comprehensively extracting auxiliary supervision from unlabeled data. Additionally, we propose a novel CWC-driven progressive learning framework to optimize the deep network by mining weak-to-strong constraints from massive unlabeled data. More specifically, this paper presents a biased cross-window consistency (BCC) loss with an importance factor, which helps the deep network explicitly constrain confidence maps from overlapping regions in different windows to maintain semantic consistency with larger contexts. In addition, we propose a dynamic pseudo-label memory bank (DPM) to provide high-consistency and high-reliability pseudo-labels to further optimize the network. Extensive experiments on three representative datasets of urban views, medical scenarios, and satellite scenes with consistent performance gain demonstrate the superiority of our framework. Our code is released at https://jack-bo1220.github.io/project/CWC.html. Bo Dang 0002, Yansheng Li 0001, Yongjun Zhang 0002, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | Augmented Maximum Correntropy Criterion for Robust Geometric PerceptionabstractMaximum correntropy criterion (MCC) is a robust and powerful technique to handle heavy-tailed nonGaussian noise, which has many applications in the fields of vision, signal processing, machine learning, etc. In this article, we introduce several contributions to the MCC and propose an augmented MCC (AMCC), which raises the robustness of classic MCC variants for robust fitting to an unprecedented level. Our first contribution is to present an accurate bandwidth estimation algorithm based on the probability density function (PDF) matching, which solves the instability problem of the Silverman's rule. Our second contribution is to introduce the idea of graduated nonconvexity (GNC) and a worst-rejection strategy into MCC, which compensates for the sensitivity of MCC to high outlier ratios. Our third contribution is to provide a definition of local distribution measure to evaluate the quality of inliers, which makes the MCC no longer limited to random outliers but is generally suitable for both random and clustered outliers. Our fourth contribution is to show the generalizability of the proposed AMCC by providing eight application examples in geometry perception and performing comprehensive evaluations on five of them. Our experiments demonstrate that 1) AMCC is empirically robust to 80%$-$90% of random outliers across applications, which is much better than Cauchy M-estimation, MCC, and GNC-GM; 2) AMCC achieves excellent performance in clustered outliers, whose success rate is 60%$-$70% percentage points higher than the second-ranked method at 80% of outliers; 3) AMCC can run in real-time, which is 10$-$100 times faster than RANSAC-type methods in low-dimensional estimation problems with high outlier ratios. This gap will increase exponentially with the model dimension. Jiayuan Li 0001, Qingwu Hu, Xinyi Liu 0002, Yongjun Zhang 0002 |
IEEE Trans. Robotics | 4 |
| 2023 | Semantic Information-Aided Geometric Correction of High Resolution Satellite ImagesabstractWith the rapid development of deep learning technology, the automatic classification and semantic segmentation of remote sensing imagery become more and more accurate. Meanwhile, with the improvement of the generalization, this technology is more and more widely used in the industry [1] . This paper proposes a new framework of sematic-information-aided geometric correction of high-resolution satellite images (HRSIs) to achieve higher accuracy and automation. Yi Wan 0001, Yongjun Zhang 0002 |
IGARSS | 3 |
| 2023 | A New Multi-Level Attention Feature Fusion Method for Hyperspectral and Lidar Data Joint ClassificationabstractJoint classification of multisource data for better Earth observation becomes an interesting but challenging problem. However, existing methods usually fail to be optimal due to the limitations in the heterogeneous feature representation and complementary information fusion. In this paper, we propose a new multi-level attention-based feature fusion method for the joint classification of HSI and LiDAR data. First, a two-stream deep network is built to extract the spectral-spatial feature of HSI and the elevation feature of LiDAR, respectively. To fully use the complementary and correlated information of HSI and LiDAR data, we adopt attention-based feature extraction and fusion module to deliver a high-discrimination feature representation both for cross-source and single-source data. Then, the extracted features are fed into fully connected layers to generate class probabilities. Finally, a decision-level fusion strategy is adopted to further improve the classification results. Extensive experiments on the Houston dataset demonstrate the effectiveness of the proposed method over some state-of-the-art approaches. Zhi Gao 0005, Leyuan Fang, Yongjun Zhang 0002 |
IGARSS | 4 |
| 2023 | M2-CDNet: A Multi-Scale and Multi-Level Network for Remote Sensing Image Change DetectionabstractChange detection plays a crucial role in environmental monitoring and earth observation tasks, leveraging the abundant data acquired by remote sensing platforms. While learning-based methods have shown promise in strictly registered datasets, their practical applicability in real-world scenarios remains challenging. This paper addresses the limitations of existing methods by proposing M2-CDNet, a novel approach that integrates the U-Net architecture with the multi-scale fusion (MSF) strategy, deformable convolutions, and multi-scale outputs. Experiments on the public and self-collected datasets demonstrate that M2-CDNet achieves superior accuracy-efficiency trade-offs compared to state-of-the-art methods. Moreover, M2-CDNet shows better robustness against image projection bias and registration errors. Yi Wan 0001, Yongjun Zhang 0002 |
IGARSS | 4 |
| 2023 | MFVNet: a deep adaptive fusion network with multiple field-of-views for remote sensing image semantic segmentation
Yansheng Li 0001, Wei Chen 0089, Xin Huang 0002, Zhi Gao 0005, Tao He 0002, Yongjun Zhang 0002 |
Sci. China Inf. Sci. | 7 |
| 2023 | To see further: Knowledge graph-aware deep graph convolutional network for recommender systems
Fei Wang 0094, Yongjun Zhang 0002, Yansheng Li 0001, Chenming Zhu |
Inf. Sci. | 3 |
| 2023 | CloudViT: A Lightweight Vision Transformer Network for Remote Sensing Cloud DetectionabstractClouds inevitably exist in satellite images, which limit the processing and application of satellite images to a certain extent. Therefore, cloud detection is a preprocessing task in satellite image extraction and analysis processing. However, the existing methods are difficult to mine robust features, and the number of parameters and computation are large, which is not conducive to the deployment of the model. In this letter, cloud vision transformer (CloudViT), a lightweight vision transformer network for cloud detection from satellite imagery, is proposed. In detail, to utilize dark channel priors in multispectral imagery to guide the network to learn features, a multiscale dark channel extractor is used to first predict dark channels, and then, the dark channel features and image features are input to the attention mechanism-based dark channel-guided context aggregation module to enhance image features, which in turn makes cloud detection results more accurate. At the same time, to enhance the transfer ability of the network between different satellite sensors, a plug-and-play channel adaptive module is proposed to deal with the inconsistency of the number of different satellite sensor bands. The experimental results on the Landsat7 dataset show that our network CloudViT outperforms the state-of-the-art methods while keeping the number of parameters and computation small. At the same time, the experimental results on transfer to three other datasets show that using the channel adaptation module can greatly improve the transfer ability of the model. Bin Zhang 0046, Yongjun Zhang 0002, Yansheng Li 0001, Yi Wan 0001, Yongxiang Yao |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | Edge-Preserving Stereo Matching Using LiDAR Points and Image Line FeaturesabstractThis letter proposes a LiDAR and image line-guided stereo matching method (L2GSM), which combines sparse but high-accuracy LiDAR points and sharp object edges of images to generate accurate and fine-structure point clouds. After extracting depth discontinuity lines on the image by using LiDAR depth information, we propose a trilateral update of cost volume and depth discontinuity lines-aware semi-global matching (SGM) strategies to integrate LiDAR data and depth discontinuity lines into the dense matching algorithm. The experimental results for the indoor and aerial datasets show that our method significantly improves the results of the original SGM and outperforms two state-of-the-art LiDAR constraints’ SGM methods, especially in recovering the 3-D structure of low-textured and depth discontinuity regions. In addition, the 3-D point clouds generated by our proposed method outperform the LiDAR data and dense matching point clouds generated by Metashape and SURE aerial in terms of completeness and edge accuracy. Siyuan Zou, Xinyi Liu 0002, Xu Huang 0005, Yongjun Zhang 0002, Senyuan Wang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | QGORE: Quadratic-Time Guaranteed Outlier Removal for Point Cloud RegistrationabstractWith the development of 3D matching technology, correspondence-based point cloud registration gains more attention. Unfortunately, 3D keypoint techniques inevitably produce a large number of outliers, i.e., outlier rate is often larger than 95%. Guaranteed outlier removal (GORE) Bustos and Chin has shown very good robustness to extreme outliers. However, the high computational cost (exponential in the worst case) largely limits its usages in practice. In this paper, we propose the first$O(N^{2})$time GORE method, called quadratic-time GORE (QGORE), which preserves the globally optimal solution while largely increases the efficiency. QGORE leverages a simple but effective voting idea via geometric consistency for upper bound estimation, which achieves almost the same tightness as the one in GORE. We also present a one-point RANSAC by exploring “rotation correspondence” for lower bound estimation, which largely reduces the number of iterations of traditional 3-point RANSAC. Further, we propose al$_{p}$p-like adaptive estimator for optimization. Extensive experiments show that QGORE achieves the same robustness and optimality as GORE while being 1$\sim$2 orders faster. The source code will be made publicly available. Jiayuan Li 0001, Qingwu Hu, Yongjun Zhang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Joint Learning of Semantic Segmentation and Height Estimation for Remote Sensing Image Leveraging Contrastive LearningabstractSemantic segmentation and height estimation are two critical tasks in remote sensing scene understanding that are highly correlated with each other. To address both tasks simultaneously, it is natural to consider designing a unified deep learning model that aims to improve performance by jointly learning complementary information among the associated tasks. In this paper, we learn the two tasks jointly under a deep multi-task learning framework and propose two novel objective functions, called cross-task contrastive loss and cross-pixel contrastive loss, respectively, to enhance multi-task learning performance through contrastive learning. Specifically, the cross-task contrastive loss is designed to maximize the mutual information of different task features and enforce the model to learn the consistency between semantic segmentation and height estimation. In addition, our method goes beyond previous approaches that only apply contrastive learning at the instance level. Instead, we design a pixel-wise contrastive loss function that pulls together pixel embeddings belonging to the same semantic class, while pushing apart pixel embeddings from different semantic classes. Furthermore, we find that this semantic-guided contrastive loss simultaneously improves the performance of the height estimation task. Our proposed approach is simple and effective and does not introduce any additional overhead to the model during the testing phase. We extensively evaluate our method on the Vaihingen and Potsdam datasets, and the experimental results demonstrate that our approach significantly outperforms the state-of-the-art methods in both height estimation and semantic segmentation. Zhi Gao 0005, Yongjun Zhang 0002, Ruifang Zhai |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Semantic-Aware Attack and Defense on Deep Hashing Networks for Remote-Sensing Image RetrievalabstractDeep hashing networks have been successful in retrieving interesting images from massive remote sensing images. There is no doubt that security and reliability are critical in remote sensing image retrieval. Recent studies about natural image retrieval have shown the vulnerability of deep hashing networks to adversarial examples, but there do not exist any researches about the attack and defense on deep hashing networks in remote sensing image retrieval. Due to the large intra-class difference and high inter-class similarity of remote sensing images, the attack and defense methods on deep hashing networks for natural images cannot be directly applied to the remote sensing images. Different from the widely adopted instance-aware hash codes which often present the suboptimum performance of the attack and defense on deep hashing networks, this paper recommends the usage of semantic-aware hash codes, which take into account multiple samples in the given semantic categories, in both attack and defense. To pursue the strongest attack on remote sensing image retrieval, a novel semantic-aware attack with weights via multiple random initialization (RWC) is proposed. To alleviate the retrieval degradation caused by adversarial attacks, a new adversarial training defense method on deep hashing networks with the adversarial semantic-aware consistency constraint (ACN) is proposed. Extensive experiments on three typical open remote sensing image datasets (i.e., UCM, AID, NWPU-RESISC45) show the proposed attack and defense methods on various deep hashing networks achieve better performance compared with the state-of-the-art methods. The source code will be made publicly available along with this paper. Yansheng Li 0001, Mengze Hao, Hu Zhu, Yongjun Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | SPGAN-DA: Semantic-Preserved Generative Adversarial Network for Domain Adaptive Remote Sensing Image Semantic SegmentationabstractUnsupervised domain adaptation for remote sensing semantic segmentation seeks to adapt a model trained on the labeled source domain to the unlabeled target domain. One of the most promising ways is to translate images from the source domain to the target domain to align the spectral information or imaging mode by the generative adversarial network (GAN). However, source-to-target translation often brings bias in the translated images causing limited performance, as semantic information is not well considered in the translation procedure. To overcome this limitation, we present an innovative semantic-preserved generative adversarial network (SPGAN), designed to mitigate the image translation bias and then leverage the translated images as well as unlabeled target images by class distribution alignment (CDA) module to train a domain adaptive semantic segmentation model. The above two stages are coupled together to form a unified framework called SPGAN-DA. Specifically, we first conduct semantic invariant translation from source to target domain, which is achieved by introducing representation-invariant and semantic-preserved constraints to the GAN model. To further narrow the landscape layout gap between the translated and target images, CDA semantic segmentation is proposed. CDA semantic segmentation consists of two aspects. At the model input level, object discrepancy is eliminated by introducing the ClassMix operation. At the model output level, boundary enhancement is proposed to refine the performance of object boundaries. Extensive experiments on three typical remote sensing cross-domain semantic segmentation benchmarks demonstrate the effectiveness and generality of our proposed method, which competes favorably against existing state-of-the-art methods. Yansheng Li 0001, Te Shi 0001, Yongjun Zhang 0002, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | A Radiometric Block Adjustment Method for Unmanned Aerial Vehicle Images Considering the Image VignettingabstractUnmanned aerial vehicles (UAVs) equipped with different sensors can provide data with high spatiotemporal resolution and have broad application prospects. During the flight of the UAV, changes in illumination, exposure time, etc., will cause different degrees of radiometric differences between images, resulting in a calibration relationship established on a single image that cannot be applied to other images; in addition, the vignetting effect also significantly changes the brightness distribution inside an image, thus posing challenges for radiometric calibration of UAV images. In this paper, based on block adjustment (BA), we proposed a radiometric block adjustment model under the consideration of vignetting and the light-dark differences between images. The proposed method requires only a small number of calibration blankets, thus reducing the complexity of the experiment. The results from two study areas showed that the proposed method could compensate for vignetting to a certain extent and the radiometric consistency of the two datasets was improved from 12.9%~21.8% to 4.7%~12.7%. Validated using ground samples, the mean RMSE and MRPE of all five bands were 0.054, 21.8%, and 0.037, 20.4% in the two study areas, respectively. The total uncertainty was less than 8.1%. When there were obvious light-dark differences between images, such as in the visible light bands, our method could significantly improve the accuracy of the radiometric calibration. Wanshan Peng, Shenghui Fang, Yongjun Zhang 0002, Jadunandan Dash, Jiacai Mo |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Hashing-Based Deep Metric Learning for the Classification of Hyperspectral and LiDAR DataabstractMultisource remote sensing data provide abundant and complementary information for land cover classification. Existing classification methods mainly focus on designing a multi-stream deep network to extract separate features of each single-source data, then adopting a fusing strategy to combine these extracted features for final classification. However, this kind of method neglects the sample correlation of single-source and cross-source data, which may deliver an unsatisfactory classification result when dealing with high intraclass-variability and low interclass-variability samples. To this end, a novel hashing-based deep metric learning (HDML) method is proposed for hyperspectral images (HSIs) and light detection and ranging (LiDAR) data classification in this paper. First, a two-stream deep network is built to extract the spectral-spatial features of HSI and the elevation features of LiDAR, respectively. To fully use the complementary and correlated information of HSI and LiDAR data, we adopt attention-based feature fusion (AFF) modules to deliver a high-discrimination fused feature both for cross-source and single-source feature fusion. Then, the extracted features are fed into fully connected layers to generate class probabilities, respectively. Different from most existing methods that only utilize semantic information of samples, we elaborately designed a loss function to simultaneously consider the label-based semantic loss and hashing-based metric loss. Finally, a decision-level fusion strategy is adopted to further improve the classification results. Extensive experiments on three public HSI and LiDAR data sets demonstrate the effectiveness of the proposed method over some state-of-the-art approaches. Zhi Gao 0005, Leyuan Fang, Yongjun Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | ELSR: Efficient Line Segment Reconstruction with Planes and Points GuidanceabstractThree-dimensional (3D) line segments are helpful for scene reconstruction. Most of the existing 3D-line-segment reconstruction algorithms deal with two views or dozens of small-size images; while in practice there are usually hundreds or thousands of large-size images. In this paper, we propose an efficient line segment reconstruction method called ELSR11Available at https://skyearth.org/publication/project/ELSR. ELSR exploits scene planes that are commonly seen in city scenes and sparse 3D points that can be acquired easily from the structure-from-motion (SfM) approach. For two views, ELSR efficiently finds the local scene plane to guide the line matching and exploits sparse 3D points to accelerate and constrain the matching. To reconstruct a 3D line segment with multiple views, ELSR utilizes an efficient abstraction approach that selects representative 3D lines based on their spatial consistence. Our experiments demonstrated that ELSR had a higher accuracy and efficiency than the existing methods. Moreover, our results showed that ELSR could reconstruct 3D lines efficiently for large and complex scenes that contain thousands of large-size images. Yi Wan 0001, Yongjun Zhang 0002, Xinyi Liu 0002, Bin Zhang 0046, Xiqi Wang |
CVPR | 3 |
| 2022 | Hierarchical Memory Learning for Fine-Grained Scene Graph Generation
Youming Deng, Yansheng Li 0001, Yongjun Zhang 0002, Xiang Xiang 0001, Jian Wang 0108, Jingdong Chen, Jiayi Ma 0001 |
ECCV (27) | 3 |
| 2022 | SCTRANS: A TRANSFORMER NETWORK BASED ON THE SPATIAL AND CHANNEL ATTENTION FOR CLOUD DETECTIONabstractCloud detection is an important preprocessing step for remote sensing image processing and analysis. The current deep-learning-based cloud detection methods are mostly based on Convolutional Neural Network (CNN) which pay more attention to local information. To make more use of the global information, in this article, we propose a transformer-based cloud detection method (SCTrans) based on the spatial and channel attention mechanism. The experiment results show that when using only three-band images on the Landsat7 dataset, the mIoU of the validation set reaches 85.92% and the mIoU of the test set reaches 87.86%. The experimental results show that the proposed network has a higher mIoU and F1 score than Fmask and other networks. Wenke Jiao, Yongjun Zhang 0002, Bin Zhang 0046, Yi Wan 0001 |
IGARSS | 2 |
| 2022 | Discriminative Feature Extraction and Fusion for Classification of Hyperspectral and Lidar DataabstractMultisource remote sensing data provide the abundant and complementary information for land cover classification. In this paper, we propose a deep hashing-based feature extraction and fusion framework for joint classification of hyper-spectral and LiDAR data. Firstly, HSIs and LiDAR data are fed into a two-stream network to extract deep features after data preprocessing. Then, we adopt hashing technique to constrain single-source and cross-source similarities, i.e., samples with same classes should have small feature distance and samples with different classes should have large feature distance. Furthermore, a feature-level fusion strategy is exploited to fuse the two kind of multisource information. Finally, we design an object function to consider the similarity information between sample pairs and semantic information of each sample, which can deliver the discriminative features for classification. The experiments on Houston data demonstrate the effectiveness of the proposed method over some competitive approaches. Zhi Gao 0005, Yongjun Zhang 0002 |
IGARSS | 3 |
| 2022 | A Cascaded Cross-Modal Network for Semantic Segmentation from High-Resolution Aerial Imagery and RAW Lidar DataabstractAs various sensors appear, extracting information from multimodal data becomes a prominent topic. Current multimodal approaches for image and LiDAR normally discard the point-to-point topology relationship of the latter to keep the dimension matched. To tackle this task, we propose a cascaded cross-modal network (CCMN) to extract the joint-features from high-resolution aerial imagery and LiDAR point directly, instead of their abridged derivatives. Firstly, point-wise features are extract from raw LiDAR data by a forepart 3D extractor. Subsequently, the LiDAR-derived features are executed spatial reference conversion to project and align to the imagery coordinate space. Finally, the cross-modal compounds containing the obtained feature maps and the corresponding images are placed into a U-shape structure to generate segmentation results. The experiment results indicate that our strategy surpasses the popular multimodal method by 6% on mIoU. Yameng Wang, Bin Zhang 0046, Yi Wan 0001, Yongjun Zhang 0002 |
IGARSS | 4 |
| 2022 | RLPath: a knowledge graph link prediction method using reinforcement learning based attentive relation path searching and representation learning
Ling Chen 0001, Jun Cui 0003, Xing Tang 0006, Yuntao Qian, Yansheng Li 0001, Yongjun Zhang 0002 |
Appl. Intell. | 6 |
| 2022 | KLGCN: Knowledge graph-aware Light Graph Convolutional Network for recommender systems
Fei Wang 0094, Yansheng Li 0001, Yongjun Zhang 0002 |
Expert Syst. Appl. | 3 |
| 2022 | Combining deep learning and ontology reasoning for remote sensing image semantic segmentation
Yansheng Li 0001, Song Ouyang, Yongjun Zhang 0002 |
Knowl. Based Syst. | 3 |
| 2022 | Few-Shot Scene Classification of Optical Remote Sensing Images Leveraging Calibrated Pretext TasksabstractSmall data holds big AI potential. As one of the promising small data AI approaches, few-shot learning has the goal to learn a model efficiently that can recognize novel classes with extremely limited training samples. Therefore it is critical to accumulate useful prior knowledge obtained from large-scale base class dataset. To realize few-shot scene classification of optical remote sensing images, we start from a baseline model that trains all base classes using a standard cross-entropy loss leveraging two auxiliary objectives to capture intrinsical characteristics across the semantic classes. Specifically, rotation prediction learns to recognize the 2D rotation of an input to guide the learning of class-transferable knowledge, and contrastive learning aims to pull together the positive pairs while pushing apart the negative pairs to promote intra-class consistency and inter-class inconsistency. We jointly optimize such two pretext tasks and semantic class prediction task in an end-to-end manner. To further overcome the overfitting issue, we introduce a regularization technique, adversarial model perturbation, to calibrate the pretext tasks so as to enhance the generalization ability. Extensive experiments on public remote sensing benchmarks including NWPU-RESISC45, AID, and WHU-RS-19 demonstrate that our method works effectively and achieves best performance that significantly outperforms many state-of-the-art approaches. Zhi Gao 0005, Yongjun Zhang 0002, Can Li 0016, Tiancan Mei |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | LNIFT: Locally Normalized Image for Rotation Invariant Multimodal Feature MatchingabstractSevere nonlinear radiation distortion (NRD) is the bottleneck problem of multimodal image matching. Although many efforts have been made in the past few years, such as the radiation-variation insensitive feature transform (RIFT) and the histogram of orientated phase congruency (HOPC), almost all these methods are based on frequency-domain information that suffers from high computational overhead and memory footprint. In this article, we propose a simple but very effective multimodal feature matching algorithm in the spatial domain, called locally normalized image feature transform (LNIFT). We first propose a local normalization filter to convert original images into normalized images for feature detection and description, which largely reduces the NRD between multimodal images. We demonstrate that normalized matching pairs have a much larger correlation coefficient than the original ones. We then detect oriented FAST and rotated brief (ORB) keypoints on the normalized images and use an adaptive nonmaximal suppression (ANMS) strategy to improve the distribution of keypoints. We also describe keypoints on the normalized images based on a histogram of oriented gradient (HOG), such as a descriptor. Our LNIFT achieves rotation invariance the same as ORB without any additional computational overhead. Thus, LNIFT can be performed in near real-time on images with 1024$\times 1024$pixels (only costs 0.32 s with 2500 keypoints). Four multimodal image datasets with a total of 4000 matching pairs are used for comprehensive evaluations, including synthetic aperture radar (SAR)–optical, infrared–optical, and depth–optical datasets. Experimental results show that LNIFT is far superior to RIFT in terms of efficiency (0.49 s versus 47.8 s on a$1024 \times 1024$image), success rate (99.9% versus 79.85%), and number of correct matches (309 versus 119). The source code and datasets will be publicly available athttps://ljy-rs.github.io/web. Jiayuan Li 0001, Wangyi Xu, Yongjun Zhang 0002, Qingwu Hu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | An Automatic Radiometric Cross-Calibration Method for Wide-Angle Medium-Resolution Multispectral Satellite Sensor Using Landsat DataabstractRadiometric calibration of the medium-resolution satellite data is critical for monitoring and quantifying changes in the Earth’s environment and resources. Many medium-resolution satellite sensors have irregular revisits and, sometimes, have a large difference in illumination viewing geometry compared with a reference sensor, posing a great challenge for routine cross-calibration practices. To overcome these issues, this study proposed a cross-calibration method to calibrate medium-resolution multispectral data. The Chinese Gaofen-4 (GF-4) panchromatic and multispectral sensor (PMS) data with large viewing angles were used as the test data, and Landsat-8 operational land imager (OLI) data were used as the reference data. A bidirectional reflectance distribution function (BRDF) correction method was proposed to eliminate the effects of differences in illumination viewing geometry between GF-4 and Landsat-8. The validation using concurrent image shows that the mean relative error (MRE) of cross calibration is less than 6.65%. Validation using ground measurements shows that our calibration results have an improvement of around 14.8% compared with the official released calibration coefficients. The time series cross calibration reveals that, without the requirements of simultaneous nadir observations (SNOs), our calibration activities can be carried out more often in practice. Gradual and continuous radiometric sensor degradation is identified with the monthly updated calibration coefficients, demonstrating the reliability and importance of the timely cross calibration. Besides, the cross-calibration approach does not rely on any specific calibration site, and the difference in illumination viewing geometry can be well considered. Thus, it can be easily adapted and applied to other optical satellite data. Tao He 0002, Shunlin Liang, Yongjun Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Asymmetric Hash Code Learning for Remote Sensing Image RetrievalabstractRemote sensing image retrieval (RSIR), aiming at searching for a set of similar items to a given query image, is a very important task in remote sensing applications. Deep hashing learning as the current mainstream method has achieved satisfactory retrieval performance. On one hand, various deep neural networks are used to extract semantic features of remote sensing images. On the other hand, the hashing techniques are subsequently adopted to map the high-dimensional deep features to the low-dimensional binary codes. This kind of method attempts to learn one hash function for both the query and database samples in a symmetric way. However, with the number of database samples increasing, it is typically time-consuming to generate the hash codes of large-scale database images. In this article, we propose a novel deep hashing method, named asymmetric hash code learning (AHCL), for RSIR. The proposed AHCL generates the hash codes of query and database images in an asymmetric way. In more detail, the hash codes of query images are obtained by binarizing the output of the network, while the hash codes of database images are directly learned by solving the designed objective function. In addition, we combine the semantic information of each image and the similarity information of pairs of images as supervised information to train a deep hashing network, which improves the representation ability of deep features and hash codes. The experimental results on three public datasets demonstrate that the proposed method outperforms symmetric methods in terms of retrieval accuracy and efficiency. The source code is available athttps://github.com/weiweisong415/Demo_AHCL_for_TGRS2022. Zhi Gao 0005, Renwei Dian, Pedram Ghamisi, Yongjun Zhang 0002, Jón Atli Benediktsson |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Multi-Modal Remote Sensing Image Matching Considering Co-Occurrence FilterabstractTraditional image feature matching methods cannot obtain satisfactory results for multi-modal remote sensing images (MRSIs) in most cases because different imaging mechanisms bring significant nonlinear radiation distortion differences (NRD) and complicated geometric distortion. The key to MRSI matching is trying to weakening or eliminating the NRD and extract more edge features. This paper introduces a new robust MRSI matching method based on co-occurrence filter (CoF) space matching (CoFSM). Our algorithm has three steps: (1) a new co-occurrence scale space based on CoF is constructed, and the feature points in the new scale space are extracted by the optimized image gradient; (2) the gradient location and orientation histogram algorithm is used to construct a 152-dimensional log-polar descriptor, which makes the multi-modal image description more robust; and (3) a position-optimized Euclidean distance function is established, which is used to calculate the displacement error of the feature points in the horizontal and vertical directions to optimize the matching distance function. The optimization results then are rematched, and the outliers are eliminated using a fast sample consensus algorithm. We performed comparison experiments on our CoFSM method with the scale-invariant feature transform (SIFT), upright-SIFT, PSO-SIFT, and radiation-variation insensitive feature transform (RIFT) methods using a multi-modal image dataset. The algorithms of each method were comprehensively evaluated both qualitatively and quantitatively. Our experimental results show that our proposed CoFSM method can obtain satisfactory results both in the number of corresponding points and the accuracy of its root mean square error. The average number of obtained matches is namely 489.52 of CoFSM, and 412.52 of RIFT. As mentioned earlier, the matching effect of the proposed method was significantly greater than the three state-of-art methods. Our proposed CoFSM method achieved good effectiveness and robustness. Executable programs of CoFSM and MRSI datasets are published: https://skyearth.org/publication/project/CoFSM/. Yongxiang Yao, Yongjun Zhang 0002, Yi Wan 0001, Xinyi Liu 0002, Xiaohu Yan, Jiayuan Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | DACHA: A Dual Graph Convolution Based Temporal Knowledge Graph Representation Learning Method Using Historical RelationabstractTemporal knowledge graph (TKG) representation learning embeds relations and entities into a continuous low-dimensional vector space by incorporating temporal information. Latest studies mainly aim at learning entity representations by modeling entity interactions from the neighbor structure of the graph. However, the interactions of relations from the neighbor structure of the graph are neglected, which are also of significance for learning informative representations. In addition, there still lacks an effective historical relation encoder to model the multi-range temporal dependencies. In this article, we propose a d ual gr a ph c onvolution network based TKG representation learning method using h istorical rel a tions (DACHA). Specifically, we first construct the primal graph according to historical relations, as well as the edge graph by regarding historical relations as nodes. Then, we employ the dual graph convolution network to capture the interactions of both entities and historical relations from the neighbor structure of the graph. In addition, the temporal self-attentive historical relation encoder is proposed to explicitly model both local and global temporal dependencies. Extensive experiments on two event based TKG datasets demonstrate that DACHA achieves the state-of-the-art results. Ling Chen 0001, Xing Tang 0006, Yuntao Qian, Yansheng Li 0001, Yongjun Zhang 0002 |
ACM Trans. Knowl. Discov. Data | 6 |
| 2021 | Representation Learning of Remote Sensing Knowledge Graph for Zero-Shot Remote Sensing Image Scene ClassificationabstractAlthough deep learning has revolutionized remote sensing image scene classification, current deep learning-based approaches highly depend on the massive supervision of the predetermined scene categories and have disappointingly poor performance on new categories which go beyond the predetermined scene categories. In reality, the classification task often has to be extended along with the emergence of new applications that inevitably involve new categories of remote sensing image scenes, so how to make the deep learning model own the inference ability to recognize the remote sensing image scenes from unseen categories becomes incredibly important. By fully exploiting the remote sensing domain characteristic, this paper proposes a novel remote sensing knowledge graph-guided deep alignment network to address zero-shot remote sensing image scene classification. To improve the semantic representation ability of remote sensing-oriented scene categories, this paper, for the first time, tries to generate the semantic representations of remote sensing scene categories by representation learning of remote sensing knowledge graph (SR-RSKG). In addition, this paper proposes a novel deep alignment network with a series of constraints (DAN) to conduct robust cross-modal alignment between visual features and semantic representations. Extensive experiments on one merged remote sensing image scene dataset, which is the integration of multiple publicly open remote sensing image scene datasets, show that the presented SR-RSKG obviously outperforms the existing semantic representation methods (e.g., the natural language processing models and manually annotated attribute vectors), and our proposed DAN shows better performance compared with the state-of- the-art methods under different kinds of semantic representations. Yansheng Li 0001, Yongjun Zhang 0002, Ruixian Chen, Jingdong Chen |
IGARSS | 3 |
| 2021 | Rotation Consistency-Preserved Generative Adversarial Networks for Cross-Domain Aerial Image Semantic SegmentationabstractDue to its wide applications, aerial image semantic segmentation attracts increasing research interest in recent years. As well known, deep semantic segmentation network (DSSN) has been widely used to deal with aerial image segmentation and achieves spectacular success. However, when applying the DSSN trained with the labeled aerial images (i.e., the source domain) to predict the aerial images acquired with different acquisition conditions (i.e., the target domain), the performance often dramatically degrades. To alleviate the negative influence of cross-domain data shift, this paper proposes a domain adaptation approach to deal with cross-domain aerial image semantic segmentation. More precisely, this paper proposes a novel rotation consistency-preserved generative adversarial network (RCP-GAN) to carry out domain adaptation for mapping aerial images in the source domain to the target domain. Furthermore, the mapped aerial imageries with labels are used to train DSSN, which is further used to classify aerial imagery in the target domain. To verify the validity of the presented approach, we give two cross-domain experimental settings including: (I) variation of geographic location; (II) variation of both geographic location and imaging mode. Extensive experiments under two typical cross-domain settings show that our proposed method can effectively address the domain shift problem and outperform the state-of-the-art methods with a large margin. Te Shi 0001, Yansheng Li 0001, Yongjun Zhang 0002 |
IGARSS | 3 |
| 2021 | Robust line segment matching via reweighted random walks on the homography graph
Yongjun Zhang 0002, Chang Li 0004 |
Pattern Recognit. | 2 |
| 2021 | AG3line: Active grouping and geometry-gradient combined validation for fast line segment extraction
Yongjun Zhang 0002, Yansheng Li 0001 |
Pattern Recognit. | 1 |
| 2021 | Error-Tolerant Deep Learning for Remote Sensing Image Scene ClassificationabstractDue to its various application potentials, the remote sensing image scene classification (RSSC) has attracted a broad range of interests. While the deep convolutional neural network (CNN) has recently achieved tremendous success in RSSC, its superior performances highly depend on a large number of accurately labeled samples which require lots of time and manpower to generate for a large-scale remote sensing image scene dataset. In contrast, it is not only relatively easy to collect coarse and noisy labels but also inevitable to introduce label noise when collecting large-scale annotated data in the remote sensing scenario. Therefore, it is of great practical importance to robustly learn a superior CNN-based classification model from the remote sensing image scene dataset containing non-negligible or even significant error labels. To this end, this article proposes a new RSSC-oriented error-tolerant deep learning (RSSC-ETDL) approach to mitigate the adverse effect of incorrect labels of the remote sensing image scene dataset. In our proposed RSSC-ETDL method, learning multiview CNNs and correcting error labels are alternatively conducted in an iterative manner. It is noted that to make the alternative scheme work effectively, we propose a novel adaptive multifeature collaborative representation classifier (AMF-CRC) that benefits from adaptively combining multiple features of CNNs to correct the labels of uncertain samples. To quantitatively evaluate the performance of error-tolerant methods in the remote sensing domain, we construct remote sensing image scene datasets with: 1) simulated noisy labels by corrupting the open datasets with varying error rates and 2) real noisy labels by deploying the greedy annotation strategies that are practically used to accelerate the process of annotating remote sensing image scene datasets. Extensive experiments on these datasets demonstrate that our proposed RSSC-ETDL approach outperforms the state-of-the-art approaches. Yansheng Li 0001, Yongjun Zhang 0002, Zhihui Zhu |
IEEE Trans. Cybern. | 2 |
| 2021 | Simultaneous Cloud Detection and Removal From Bitemporal Remote Sensing Images Using Cascade Convolutional Neural NetworksabstractClouds and cloud shadows heavily affect the quality of the remote sensing images and their application potential. Algorithms have been developed for detecting, removing, and reconstructing the shaded regions with the information from the neighboring pixels or multisource data. In this article, we propose an integrated cloud detection and removal framework using cascade convolutional neural networks, which provides accurate cloud and shadow masks and repaired images. First, a novel fully convolutional network (FCN), embedded with multiscale aggregation and the channel-attention mechanism, is developed for detecting clouds and shadows from a cloudy image. Second, another FCN, with the masks of the detected cloud and shadow, the cloudy image, and a temporal image as the input, is used for the cloud removal and missing-information reconstruction. The reconstruction is realized through a self-training strategy that is designed to learn the mapping between the clean-pixel pairs of the bitemporal images, which bypasses the high demand of manual labels. Experiments showed that our proposed framework can simultaneously detect and remove the clouds and shadows from the images and the detection accuracy surpassed several recent cloud-detection methods; the effects of image restoring outperform the mainstream methods in every indicator by a large margin. The data set used for cloud detection and removal is made open. Shunping Ji, Peiyu Dai, Yongjun Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Learning Deep Cross-Modal Embedding Networks for Zero-Shot Remote Sensing Image Scene ClassificationabstractDue to its wide applications, remote sensing (RS) image scene classification has attracted increasing research interest. When each category has a sufficient number of labeled samples, RS image scene classification can be well addressed by deep learning. However, in the RS big data era, it is extremely difficult or even impossible to annotate RS scene samples for all the categories in one time as the RS scene classification often needs to be extended along with the emergence of new applications that inevitably involve a new class of RS images. Hence, the RS big data era fairly requires a zero-shot RS scene classification (ZSRSSC) paradigm in which the classification model learned from training RS scene categories obeys the inference ability to recognize the RS image scenes from unseen categories, in common with the humans’ evolutionary perception ability. Unfortunately, zero-shot classification is largely unexploited in the RS field. This article proposes a novel ZSRSSC method based on locality-preservation deep cross-modal embedding networks (LPDCMENs). The proposed LPDCMENs, which can fully assimilate the pairwise intramodal and intermodal supervision in an end-to-end manner, aim to alleviate the problem of class structure inconsistency between two hybrid spaces (i.e., the visual image space and the semantic space). To pursue a stable and generalization ability, which is highly desired for ZSRSSC, a set of explainable constraints is specially designed to optimize LPDCMENs. To fully verify the effectiveness of the proposed LPDCMENs, we collect a new large-scale RS scene data set, including the instance-level visual images and class-level semantic representations (RSSDIVCS), where the general and domain knowledge is exploited to construct the class-level semantic representations. Extensive experiments show that the proposed ZSRSSC method based on LPDCMENs can obviously outperform the state-of-the-art methods, and the domain knowledge further improves the performance of ZSRSSC compared with the general knowledge. The collected RSSDIVCS will be made publicly available along with this article. Yansheng Li 0001, Zhihui Zhu, Jin-Gang Yu, Yongjun Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | SemiCDNet: A Semisupervised Convolutional Neural Network for Change Detection in High Resolution Remote-Sensing ImagesabstractChange detection (CD) is one of the main applications of remote sensing. With the increasing popularity of deep learning, most recent developments of CD methods have introduced the use of deep learning techniques to increase the accuracy and automation level over traditional methods. However, when using supervised CD methods, a large amount of labeled data is needed to train deep convolutional networks with millions of parameters. These labeled data are difficult to acquire for CD tasks. To address this limitation, a novel semisupervised convolutional network for CD (SemiCDNet) is proposed based on a generative adversarial network (GAN). First, both the labeled data and unlabeled data are input into the segmentation network to produce initial predictions and entropy maps. Then, to exploit the potential of unlabeled data, two discriminators are adopted to enforce the feature distribution consistency of segmentation maps and entropy maps between the labeled and unlabeled data. During the competitive training, the generator is continuously regularized by utilizing the unlabeled information, thus improving its generalization capability. The effectiveness and reliability of our proposed method are verified on two high-resolution remote sensing data sets. Extensive experimental results demonstrate the superiority of the proposed method against other state-of-the-art approaches. Daifeng Peng, Lorenzo Bruzzone, Yongjun Zhang 0002, Haiyan Guan, Haiyong Ding, Xu Huang 0005 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Small Object Detection Leveraging on Simultaneous Super-resolutionabstractDespite the impressive advancement achieved in object detection, the detection performance of small object is still far from satisfactory due to the lack of sufficient detailed appearance to distinguish it from similar objects. Inspired by the positive effects of super-resolution for object detection, we propose a framework that can be incorporated with detector networks to improve the performance of small object detection, in which the low-resolution image is super-resolved via generative adversarial network (GAN) in an unsupervised manner. In our method, the super-resolution network and the detection network are trained jointly. In particular, the detection loss is back-propagated into the super-resolution network during training to facilitate detection. Compared with available simultaneous super-resolution and detection methods which heavily rely on low-/high-resolution image pairs, our work breaks through such restriction via applying the CycleGAN strategy, achieving increased generality and applicability, while remaining an elegant structure. Extensive experiments on datasets from both computer vision and remote sensing communities demonstrate that our method obtains competitive performance on a wide range of complex scenarios. Zhi Gao 0005, Xiaodong Liu 0008, Yongjun Zhang 0002, Tiancan Mei |
ICPR | 4 |
| 2020 | DEM Extraction from Airborne Lidar Point Cloud in Thick-Forested Areas via Convolutional Neural NetworkabstractDigital Elevation Model (DEM), representing the height of the earth terrain, is one of the crucial geographic information products. One of the main data source of DEM is the airborne LiDAR point cloud with its non-ground-reflections filtered out. Point cloud filtering in thick-forested areas is difficult without enough ground control points when using conventional methods. In this paper, a supervised method is proposed to handle the problem of automatic DEM extraction with little ground control points. The design of the method is inspired by the successful application of the convolutional neural networks (CNN) in the image super resolution (SR) process. First, with the given LiDAR point cloud, the digital surface model (DSM) is resampled with regular grid. Then, by learning the spatial autocorrelation between the DSM and its corresponding DEM, a robust CNN model is established. Finally, the DEM in thick-forested areas can be generated from the DSM with the trained model. Experimental results at two different mountain sites in China validate the effectiveness of the proposed method of high-precision DEM generation. Yongjun Zhang 0002, Sizhe Xiang, Yi Wan 0001, Yimin Luo |
IGARSS | 1 |
| 2020 | Deep Networks Under Block-Level Supervision for Pixel-Level Cloud Detection in Multi-Spectral Satellite ImageryabstractCloud cover hinders the usability of optical remote sensing imagery. Existing cloud detection methods either require hand-crafted features or utilize deep networks. Generally, deep networks perform better than hand-crafted features. However, deep networks for cloud detection need massive and expensive pixel-level annotation labels. To alleviate that, this paper proposes a weakly supervised deep learning-based cloud detection method using only block-level labels, with a new global convolutional pooling operation and a local pooling pruning strategy to improve the performance. For evaluating, we collect a training dataset containing over 160,000 image blocks with block-level labels and a testing dataset including ten large image scenes with pixel-level labels. Even under extremely weak supervision, our method performed well with the average overall accuracy reached 97.2 %. Experiments demonstrate that our proposed method obviously outperforms the state-of-the-art methods. Wei Chen 0089, Yansheng Li 0001, Yongjun Zhang 0002, Xiaolong Hao |
IGARSS | 3 |
| 2020 | A CNN-GCN Framework for Multi-Label Aerial Image Scene ClassificationabstractAs one of the fundamental tasks in aerial image understanding, multi-label aerial image scene classification attracts increasing research interest. In general, the semantic category of a scene is reflected by the object information and the topological relations among objects. Most of existing deep learning-based aerial image scene classification methods (e.g., convolutional neural network (CNN)) classify the image scene by perceiving object information, while how to learn spatial relationships from image scene is still a challenging problem. In literature, graph convolutional network (GCN) has been successfully used for learning spatial characteristics of topological data, but it is rarely adopted in aerial image scene classification. To simultaneously mine both the object visual information and spatial relationships among multiple objects, this paper proposes a novel framework combining CNN and GCN to address multi-label aerial image scene classification. Extensive experimental results on two public datasets show that our proposed method can achieve better performance than the state-of-the-art methods. Yansheng Li 0001, Ruixian Chen, Yongjun Zhang 0002 |
IGARSS | 3 |
| 2020 | Unsupervised Style Transfer via Dualgan for Cross-Domain Aerial Image ClassificationabstractDue to its wide applications, aerial image classification, which is also called semantic segmentation of aerial imagery, attracts increasing research interest in recent years. Until now, deep semantic segmentation network (DSSN) has been widely adopted to address aerial image classification and achieves tremendous success. However, the superior performance of DSSN highly depends on massive targeted data with labels. When DSSN is trained on data from the source domain but tested on data from the target domain, the performance of DSSN is often very limited due to the data shift between source and target domains. To alleviate the disadvantage influence of data shift, this paper proposes a domain adaptation approach via unsupervised style transfer to cope with cross-domain aerial image classification. More specifically, this paper innovatively recommends DualGAN to conduct unsupervised style transfer for mapping aerial images in the source domain to the target domain. The mapped aerial imagery with labels is adopted to train DSSN, which is further used to classify aerial imagery in the target domain. To verify the validity of the presented approach, we give two cross-domain experimental settings including: (I) variation of geographic location; (II) variation of both geographic location and imaging mode. Extensive experiments under two typical cross-domain settings show that our proposed method can obviously outperform the state-of-the-art methods. Yansheng Li 0001, Te Shi 0001, Wei Chen 0089, Yongjun Zhang 0002, Zhibin Wang 0004, Hao Li 0030 |
IGARSS | 4 |
| 2020 | A Target Tracking and Positioning Framework for Video Satellites Based on SLAMabstractWith the booming development in aerospace technology, the video satellite which observes the live phenomena on the ground by video shooting has gradually emerged as a new Earth observation method. And remote sensing comes into a "dynamic" era with the demand for new processing techniques, especially the near-real-time tracking and geo-positioning algorithm for ground moving targets. However, many researchers merely extract pixel-level trajectories in post-processed video products, resulting in fairly limited applications. We regard the video satellite as a robot flying in space and adopt the SLAM framework for the positioning of ground moving targets. The designed framework is based on the representative ORB-SLAM and we make improvements mainly in feature extraction, satellite pose estimation, moving target tracking and positioning. We coordinate a moving fishing boat with GPS-RTK (Real-time Kinematic) devices and a video satellite observing it simultaneously for verification and evaluation of our method. Experiments demonstrate that our framework provides reasonable geolocation of the moving target in satellite videos. Finally, some open problems and potential research directions are discussed. Zhi Gao 0005, Yongjun Zhang 0002, Ben M. Chen |
IROS | 3 |
| 2020 | Multimodal image registration using histogram of oriented gradient distance and data-driven grey wolf optimizer
Xiaohu Yan, Yongjun Zhang 0002, Dejun Zhang, Neng Hou |
Neurocomputing | 2 |
| 2020 | Registration of Multimodal Remote Sensing Images Using Transfer OptimizationabstractMultimodal image registration is critical yet challenging for remote sensing image processing. Due to the large nonlinear intensity differences between the multimodal images, conventional search algorithms tend to get trapped into local optima when optimizing the transformation parameters by maximizing mutual information (MI). To address this problem, inspired by transfer learning, we propose a novel search algorithm named transfer optimization (TO), which can be applied to any optimizer. In TO, an optimizer transfers its better individuals to the other optimizer in each iteration. Thus, TO can share information between two optimizers and take advantage of their search mechanisms, which is helpful to avoid the local optima. Then, the registration of the multimodal remote sensing images using TO is presented. We compare the proposed algorithm with several state-of-the-art algorithms on real and simulated image pairs. Experimental results demonstrate the superiority of our algorithm in terms of registration accuracy. Xiaohu Yan, Yongjun Zhang 0002, Dejun Zhang, Neng Hou, Bin Zhang 0046 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2020 | Band-Independent Encoder-Decoder Network for Pan-Sharpening of Remote Sensing ImagesabstractPan-sharpening is a fundamental task for remote sensing image processing. It aims at creating a high-resolution multispectral (HRMS) image from a multispectral (MS) image and a panchromatic (PAN) image. In this article, a new band-independent encoder-decoder network is proposed for pan-sharpening. The network takes a single band of the MS (BMS) image, the PAN image, and the low-resolution PAN (LRPAN) image as inputs. The output of the network is the corresponding band of high-resolution MS (HRBMS) image. In this way, the network can process MS images with any number of bands. The overall structure of the network consists of two encoder-decoder modules at low-resolution and high-resolution, respectively. An auxiliary LRPAN image is used to speed up the training and improve the performance. The partly shared network and hierarchical structure for low-resolution and high-resolution enable a better fusion of features extracted from different scales. With a fast fine-tuning strategy, the trained model can be applied to images from different sensors. Experiments performed on different data sets demonstrate that the proposed method outperforms several state-of-the-art pan-sharpening methods in both visual appearance and objective indexes, and the single-band evaluation results further verify the superiority of the proposed method. Chi Liu 0004, Yongjun Zhang 0002, Shugen Wang, Yangjun Ou, Yi Wan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | A Rigorous on-orbit Geometric Calibration Method for High-Resolution Optical Sensor of Chinese Mapping SatelliteabstractInternationally, the technology of optical earth observation satellite is developing in the direction of miniaturization, high-resolution and high maneuverability. High-precision on-orbit geometric calibration of the imaging sensors is the most important guarantee of interior geometric quality and direct positioning accuracy of the satellite images. The high-resolution camera of Chinese Mapping Satellite-1 is a multi-chip TDICCD optical sensor in triangular mechanical staggered stitching structure. To improve the accuracy and robustness of its geometric calibration while reducing the correlation between calibration parameters, this paper proposed a new multi-chip pointing angle model for internal calibration and an optimized imaging model for external calibration. The overlapping connection relation between multi-chip TDICCDs of an image and multi-scene images are fully utilized for constructing the calibration error equations with constraints. Then, the internal and external calibration are iteratively solved through relaxation block adjustment method. Experimental results demonstrate that the proposed rigorous on-orbit geometric calibration method can effectively and stably eliminate various system errors of the imaging system, thus improving the interior and exterior geometric quality of image production. Kun Hu 0017, Yongjun Zhang 0002, Xu Huang 0005 |
IGARSS | 2 |
| 2019 | Learning Deep Networks under Noisy Labels for Remote Sensing Image Scene ClassificationabstractThe deep convolutional neural network (DCNN) has been successfully applied in remote sensing (RS) image scene classification, and the superior performances of DCNNs highly depend on not only the large number of samples, but also the high accuracy of labels. In practice, collecting definitely accurate labels for a large-scale RS image scene dataset needs large amounts of manual intervention, but collecting roughly noisy labels for a dataset would be significantly simplified in the RS scenario. Therefore, how to robustly learn the superior DCNN-based classification model from one RS image scene dataset containing some error labels is a practical problem of great importance. To this end, this paper proposes a simple but effective RS-oriented error-tolerant deep learning (RS-ETDL) approach to mitigate the adverse effect of incorrect labels of the corrupted RS image scene dataset. In our proposed RS-ETDL method, learning multi-view DCNN models and correcting error labels are jointly conducted in an iteratively alternative manner. Extensive experiments on the noisy RS image scene dataset demonstrate that our proposed method outperforms the state-of-the-art approaches with a large margin. Yansheng Li 0001, Yongjun Zhang 0002, Zhihui Zhu |
IGARSS | 2 |
| 2019 | A Mixture Likelihood Model of the Anisotropic Gaussian and Uniform Distributions for Accurate Oblique Image Point MatchingabstractIn this letter, we propose a mixture likelihood model for accurate oblique image point matching. The basic prior assumption is that the noises are anisotropic with zero mean and different covariances in x- and y -directions for inliers, while the outliers have uniform distribution, which is more suitable for tilted scenes or viewpoint changes. Furthermore, the oblique image point matching problem is formulated as an improved maximum a posteriori (IMAP) estimation of a Bayesian model. In this model, based on the vector field interpolation framework, we combined the mixture likelihood model and our previous adaptive image mismatch removal method, where a two-order term of the regularization coefficient is introduced into the regularized risk function, and a parameter self-adaptive Gaussian kernel function is imposed to construct the regularization term. Subsequently, the expectation-maximization algorithm is utilized to solve the IMAP estimation, in which all the latent variances are able to obtain excellent estimation. Experimental results on real data sets verified that our method was superior to some similar methods in terms of precision and also had better self-adaptability characteristic than some hypothesis-and-verify methods. More experiments on viewpoint changes demonstrated our method's effectiveness without loss of precision-recall tradeoffs, besides significant efficiency improvement. Xunwei Xie, Yongjun Zhang 0002, Daifeng Peng |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2019 | Automatic and Unsupervised Water Body Extraction Based on Spectral-Spatial Features Using GF-1 Satellite ImageryabstractWater body extraction from remote sensing imagery is an essential and nontrivial issue due to the complexity of the spectral characteristics of various kinds of water bodies and the redundant background information. An automatic multifeature water body extraction (MFWE) method integrating spectral and spatial features is proposed in this letter for water body extraction from GF-1 multispectral imagery in an unsupervised way. This letter first discusses a spatial feature index, called the pixel region index (PRI), to describe the smoothness in a local area surrounding a pixel. PRI is advantageous for assisting the normalized difference water index (NDWI) in detecting major water bodies, especially in urban areas. On the other hand, part of the water pixels near the borders may not be included in major water bodies, k-means clustering is subsequently conducted to cluster all the water pixels into the same group as a guide map. Finally, the major water bodies and the guide map are merged to obtain the final water mask. Our experimental results demonstrate that accurate water masks were achieved for all seven GF-1 imagery scenes examined. Three images with a complex background and water conditions were used to quantitatively compare the proposed method to NDWI thresholding and support vector machine classification, which verified the higher accuracy and effectiveness of the proposed method. Yongjun Zhang 0002, Xinyi Liu 0002, Xu Huang 0005 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2019 | Pan-Sharpening Using an Efficient Bidirectional Pyramid NetworkabstractPan-sharpening is an important preprocessing step for remote sensing image processing tasks; it fuses a low-resolution multispectral image and a high-resolution (HR) panchromatic (PAN) image to reconstruct a HR multispectral (MS) image. This paper introduces a new end-to-end bidirectional pyramid network for pan-sharpening. The overall structure of the proposed network is a bidirectional pyramid, which permits the network to process MS and PAN images in two separate branches level by level. At each level of the network, spatial details extracted from the PAN image are injected into the upsampled MS image to reconstruct the pan-sharpened image from coarse resolution to fine resolution. Subpixel convolutional layers and the enhanced residual blocks are used to make the network efficient. Comparison of the results obtained with our proposed method and the results using other widely used state-of-the-art approaches confirms that our proposed method outperforms the others in visual appearance and objective indexes. Yongjun Zhang 0002, Chi Liu 0004, Yangjun Ou |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | A Coarse-to-Fine Framework for Cloud Removal in Remote Sensing Image SequenceabstractClouds and accompanying shadows, which exist in optical remote sensing images with high possibility, can degrade or even completely occlude certain ground-cover information in images, limiting their applicabilities for Earth observation, change detection, or land-cover classification. In this paper, we aim to deal with cloud contamination problems with the objective of generating cloud-removed remote sensing images. Inspired by low-rank representation together with sparsity constraints, we propose a coarse-to-fine framework for cloud removal in the remote sensing image sequence. Leveraging on group-sparsity constraint, we first decompose the observed cloud image sequence of the same area into the low-rank component, group-sparse outliers, and sparse noise, corresponding to cloud-free land-covers, clouds (and accompanying shadows), and noise respectively. Subsequently, a discriminative robust principal component analysis (RPCA) algorithm is utilized to assign aggressive penalizing weights to the initially detected cloud pixels to facilitate cloud removal and scene restoration. Moreover, we incorporate geometrical transformation into a low-rank model to address the misalignment of the image sequence. Significantly superior to conventional cloud-removal methods, neither cloud-free reference image(s) nor additional operations of cloud and shadow detection are required in our method. Extensive experiments on both simulated data and real data demonstrate that our method works effectively, outperforming many state-of-the-art approaches. Yongjun Zhang 0002, Fei Wen 0004, Zhi Gao 0005 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2018 | Salient Object Detection Via Double Sparse Representations Under Visual Attention GuidanceabstractThis paper introduces a novel method for salient object detection from the perspective of sparse representation under visual attention guidance. After pretreatment and regional analysis with eye fixation detection and multi scale segmentation, regions that are used to make up the foreground and background dictionaries are respectively selected by sorting the visual attraction level of all image regions. For saliency measurement, the reconstruction errors instead of common local and global contrasts are used as the saliency indicator, which is expected to improve the object integrity. In addition, the multi scale workflow is conductive to enhance the robustness for objects of different sizes. The proposed method was compared to six state-of-the-art saliency detection methods using three benchmark datasets, and it was confirmed to have more favorable performance in the detection of multiple objects as well as maintaining the integrity of the object area. Yongjun Zhang 0002, Xunwei Xie, Yansheng Li 0001 |
IGARSS | 2 |
| 2018 | A New Registration Algorithm for Multimodal Remote Sensing ImagesabstractAutomatic registration of remote sensing images is a challenging problem in the applications of remote sensing. The multimodal remote sensing images have significant nonlinear radiometric differences, which lead to the failure of area-based and feature-based registration methods. In this paper, to overcome significant nonlinear radiometric differences and large scale differences of multimodal remote sensing images, we propose a new registration algorithm, which can meet the need of initial registration of multimodal remote sensing images that conform to similarity transformation model. Our synthetic and real-data experimental results demonstrate the effectiveness and good performance of the proposed method in terms of visualization and registration accuracy. Xunwei Xie, Yongjun Zhang 0002 |
IGARSS | 2 |
| 2018 | A Novel Fine Registration Technique for Very High Resolution Remote Sensing ImagesabstractThis paper presents a novel registration noise (RN) estimation technique for fine registration of very high resolution (VHR) images. This is accomplished by using a two-step strategy to estimate and mitigate residual local misalignments in standardly registered VHR images. The first step takes advantages of the superpixel segmentation and frequency filtering to generate sparse superpixels as the basic objects for RN estimation. Then local rectification is employed for fine registration of the input image under the aid of RN information. More factors are taken into consideration in order to enhance the RN estimation performance. The proposed approach is designed in a fine registration strategy, which can effectively improve the pre-registration result. The experimental results obtained with real datasets confirm the effectiveness of the proposed method. Xianzhang Zhu, Yongjun Zhang 0002 |
IGARSS | 2 |
| 2018 | Two-Pass Robust Component Analysis for Cloud Removal in Satellite Image SequenceabstractDue to the inevitable existence of clouds and their shadows in optical remote sensing images, certain ground-cover information is degraded or even appears to be missing, which limits analysis and utilization. Thus, cloud removal is of great importance to facilitate downstream applications. Motivated by the sparse representation techniques which have obtained a stunning performance in a variety of applications, including target detection, anomaly detection, and so on; we propose a two-pass robust principal component analysis (RPCA) framework for cloud removal in the satellite image sequence. First, a plain RPCA is applied for initial cloud region detection, followed by a straightforward morphological operation to ensure that the cloud region is completely detected. Subsequently, a discriminative RPCA algorithm is proposed to assign aggressive penalizing weights to the detected cloud pixels to facilitate cloud removal and scene restoration. Significantly superior to currently available methods, neither a cloud-free reference image nor a specific algorithm of cloud detection is required in our method. Experiments on both simulated and real images yield visually plausible and numerically verified results, demonstrating the effectiveness of our method. Fei Wen 0004, Yongjun Zhang 0002, Zhi Gao 0005 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2018 | Object-Based Change Detection for VHR Images Based on Multiscale Uncertainty AnalysisabstractScale is of great significance in image analysis and interpretation. In order to utilize scale information, multiscale fusion is usually employed to combine change detection (CD) results from different scales. However, CD results from different scales are usually treated independently, which ignores the scale contextual information. To overcome this drawback, this letter introduces a novel object-based change detection (OBCD) technique for unsupervised CD in very high-resolution (VHR) images by incorporating multiscale uncertainty analysis. First, two temporal images are stacked and segmented using a series of optimal segmentation scales ranging from coarse to fine. Second, an initial CD result is obtained by fusing the pixel-based CD result and OBCD result based on Dempter-Shafer (DS) evidence theory. Third, multiscale uncertainty analysis is implemented from coarse scale to fine scale by support vector machine classification. Finally, a CD map is generated by combining all the available information in all the scales. The experimental results employing SPOT5 and GF-1 images demonstrate the effectiveness and superiority of the proposed approach. Yongjun Zhang 0002, Daifeng Peng, Xu Huang 0005 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2018 | Fine Registration for VHR Images Based on Superpixel Registration-Noise EstimationabstractLocal nonlinear geometric distortion is problematic in the registration of very high-resolution (VHR) images. In the standard registration approach, the precision of control points generated from salient feature matching cannot be guaranteed. This letter introduces a novel superpixel registration-noise (RN) estimation method based on a two-step fine registration technique that can be estimate and mitigate the local residual misalignments in VHR images. The first step employs superpixel sparse representation and multiple displacement analysis to estimate RN information of the preregistered image. The second step optimizes the control points obtained in preregistration by combining the RN information and gross error information, and finally fine registers the input image by employing local rectification. The experiments using two data sets generated from Chinese GF2, GF1, and ZY3 satellites are discussed in this letter, and the promising results verify the effectiveness of the proposed new method. Xianzhang Zhu, Yongjun Zhang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2018 | Robust infrared small target detection using local steering kernel reconstruction
Yansheng Li 0001, Yongjun Zhang 0002 |
Pattern Recognit. | 2 |
| 2018 | Learning Source-Invariant Deep Hashing Convolutional Neural Networks for Cross-Source Remote Sensing Image RetrievalabstractDue to the urgent demand for remote sensing big data analysis, large-scale remote sensing image retrieval (LSRSIR) attracts increasing attention from researchers. Generally, LSRSIR can be divided into two categories as follows: uni-source LSRSIR (US-LSRSIR) and cross-source LSRSIR (CS-LSRSIR). More specifically, US-LSRSIR means the inquiry remote sensing image and images in the searching data set come from the same remote sensing data source, whereas CS-LSRSIR is designed to retrieve remote sensing images with a similar content to the inquiry remote sensing image that are from a different remote sensing data source. In the literature, US-LSRSIR has been widely exploited, but CS-LSRSIR is rarely discussed. In practical situations, remote sensing images from different kinds of remote sensing data sources are continually increasing, so there is a great motivation to exploit CS-LSRSIR. Therefore, this paper focuses on CS-LSRSIR. To cope with CS-LSRSIR, this paper proposes source-invariant deep hashing convolutional neural networks (SIDHCNNs), which can be optimized in an end-to-end manner using a series of well-designed optimization constraints. To quantitatively evaluate the proposed SIDHCNNs, we construct a dual-source remote sensing image data set that contains eight typical land-cover categories and 10 000 dual samples in each category. Extensive experiments show that the proposed SIDHCNNs can yield substantial improvements over several baselines involving the most recent techniques. Yansheng Li 0001, Yongjun Zhang 0002, Xin Huang 0002, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Large-Scale Remote Sensing Image Retrieval by Deep Hashing Neural NetworksabstractAs one of the most challenging tasks of remote sensing big data mining, large-scale remote sensing image retrieval has attracted increasing attention from researchers. Existing large-scale remote sensing image retrieval approaches are generally implemented by using hashing learning methods, which take handcrafted features as inputs and map the high-dimensional feature vector to the low-dimensional binary feature vector to reduce feature-searching complexity levels. As a means of applying the merits of deep learning, this paper proposes a novel large-scale remote sensing image retrieval approach based on deep hashing neural networks (DHNNs). More specifically, DHNNs are composed of deep feature learning neural networks and hashing learning neural networks and can be optimized in an end-to-end manner. Rather than requiring to dedicate expertise and effort to the design of feature descriptors, we can automatically learn good feature extraction operations and feature hashing mapping under the supervision of labeled samples. To broaden the application field, DHNNs are evaluated under two representative remote sensing cases: scarce and sufficient labeled samples. To make up for a lack of labeled samples, DHNNs can be trained via transfer learning for the former case. For the latter case, DHNNs can be trained via supervised learning from scratch with the aid of a vast number of labeled samples. Extensive experiments on one public remote sensing image data set with a limited number of labeled samples and on another public data set with plenty of labeled samples show that the proposed remote sensing image retrieval approach based on DHNNs can remarkably outperform state-of-the-art methods under both of the examined conditions. Yansheng Li 0001, Yongjun Zhang 0002, Xin Huang 0002, Hu Zhu, Jiayi Ma 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | 3D building roof reconstruction from airborne LiDAR point clouds: a framework based on a spatial databaseabstractThree-dimensional (3D) building models are essential for 3D Geographic Information Systems and play an important role in various urban management applications. Although several light detection and ranging (LiDAR) data-based reconstruction approaches have made significant advances toward the fully automatic generation of 3D building models, the process is still tedious and time-consuming, especially for massive point clouds. This paper introduces a new framework that utilizes a spatial database to achieve high performance via parallel computation for fully automatic 3D building roof reconstruction from airborne LiDAR data. The framework integrates data-driven and model-driven methods to produce building roof models of the primary structure with detailed features. The framework is composed of five major components: (1) a density-based clustering algorithm to segment individual buildings, (2) an improved boundary-tracing algorithm, (3) a hybrid method for segmenting planar patches that selects seed points in parameter space and grows the regions in spatial space, (4) a boundary regularization approach that considers outliers and (5) a method for reconstructing the topological and geometrical information of building roofs using the intersections of planar patches. The entire process is based on a spatial database, which has the following advantages: (a) managing and querying data efficiently, especially for millions of LiDAR points, (b) utilizing the spatial analysis functions provided by the system, reducing tedious and time-consuming computation, and (c) using parallel computing while reconstructing 3D building roof models, improving performance. Rujun Cao, Yongjun Zhang 0002, Xinyi Liu 0002, Zongze Zhao |
Int. J. Geogr. Inf. Sci. | 2 |
| 2017 | Automatic Reference Image Selection for Color Balancing in Remote Sensing Imagery MosaicabstractSelection of a reference image is an important step in color balancing. However, the past research and currently available methods do not focus on it, leading to the lack of an effective way to select the reference image for color balancing in remote sensing imagery mosaic. This letter proposes a novel automatic reference image selection method that aims to select the reference images by assessing multifactors according to the land surface types of the target images. The proposed method addresses the limitations caused by the use of a single assessment factor as well as the selection of a single image as the reference in traditional methods. In addition, the proposed method has a wider range of applications than those requiring no reference image. The visual experimental results indicate that the proposed method can select the suitable reference images, which benefits the color balancing result, and outperforms the other comparative methods. Moreover, the absolute mean value of skewness metric of the proposed method is 0.0831, which is lower than the values of the other comparison methods. It indicates that the result of the proposed method had the best performance in the color information. The quantitative analyses with the metric of absolute difference of mean value indicate that the proposed method has a good ability in maintaining the spectral information, and the spectral changing rates had been reduced at least 10.66% by the proposed method when compared with the other methods. Lei Yu 0005, Yongjun Zhang 0002, Yihui Lu |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2017 | A Two-Step Semiglobal Filtering Approach to Extract DTM From Middle Resolution DSMabstractMany filtering algorithms have been developed to extract the digital terrain model (DTM) from dense urban light detection and ranging data or the high-resolution digital surface model (DSM), assuming a smooth variation of topographic relief. However, this assumption breaks for a middle-resolution DSM because of the diminished distinction between steep terrains and nonground points. This letter introduces a two-step semiglobal filtering (TSGF) workflow to separate those two components. The first SGF step uses the digital elevation model of the Shuttle Radar Topography Mission to obtain a flat-terrain mask for the input DSM; then, a segmentation-constrained SGF is used to remove the nonground points within the flat-terrain mask while maintaining the shape of the terrain. Experiments are conducted using DSMs generated from Chinese ZY3 satellite imageries, verified the effectiveness of the proposed method. Compared with the conventional progressive morphological filter method, the usage of flat-terrain mask reduced the average root-mean-square error of DTM from 9.76 to 4.03 m, which is further reduced to 2.42 m by the proposed TSGF method. Yanfeng Zhang 0004, Yongjun Zhang 0002, Zhang Yunjun, Zongze Zhao |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2017 | A Simple and Efficient Method for Radial Distortion Estimation by Relative OrientationabstractIn order to solve the accuracy problem caused by lens distortions of nonmetric digital cameras mounted on an unmanned aerial vehicle, the estimation for initial values of lens distortion must be studied. Based on the fact that radial lens distortions are the most significant of lens distortions, a simple and efficient method for radial lens distortion estimation is proposed in this paper. Starting from the coplanar equation, the geometric characteristics of the relative orientation equations are explored. This paper further proves that the radial lens distortion can be linearly estimated in a continuous relative orientation model. The proposed procedure only requires a sufficient number of point correspondences between two or more images obtained by the same camera; thus it is suitable for a natural scene where the lack of straight lines and calibration objects precludes most previous techniques. Both computer simulation and real data have been used to test the proposed method; the experimental results show that the proposed method is easy to use and flexible. Yansong Duan, Yongjun Zhang 0002, Zuxun Zhang, Xinyi Liu 0002, Kun Hu 0017 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2017 | A Mixed Radiometric Normalization Method for Mosaicking of High-Resolution Satellite ImageryabstractA new mixed radiometric normalization (MRN) method is introduced in this paper which aims to eliminate the radiometric difference in image mosaicking. The radiometric normalization methods can be classified as the absolute and relative approaches in traditional solutions. Though the absolute methods could get the precise surface reflectance values of the images, rigorous conditions required for them are usually difficult to obtain, which makes the absolute methods impractical in many cases. The relative methods, which are simple and practicable, are more widely applied. However, the standard for designating the reference image needed for these methods is not unified. Moreover, the color error propagation and the two-body problems are common obstacles for the relative methods. The proposed MRN approach combines absolute and relative radiometric normalization methods, by which the advantages of both can be fully used and the limitations can be effectively avoided. First, suitable image after absolute radiometric calibration is selected as the reference image. Then, the invariant feature probability between the pixels of the target image and that of the reference image is obtained. Afterward, an adaptive local approach is adopted to obtain a suitable linear regression model for each block. Finally, a bilinear interpolation method is employed to obtain the radiometric calibration parameters for each pixel. Moreover, the CIELAB color space is adopted to evaluate the results quantitatively. Experimental results of ZY-3, GF-1, and GF-2 data indicate that the proposed method can eliminate the radiometric differences between images from the same or even different sensors. Yongjun Zhang 0002, Lei Yu 0005 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2016 | A novel spatio-temporal saliency approach for robust dim moving target detection from airborne infrared image sequences
Yansheng Li 0001, Yongjun Zhang 0002, Jin-Gang Yu, Yihua Tan, Jinwen Tian, Jiayi Ma 0001 |
Inf. Sci. | 2 |
| 2016 | DEM-Aided Bundle Adjustment With Multisource Satellite Imagery: ZY-3 and GF-1 in Large AreasabstractIn this letter, a new digital elevation model (DEM)-aided bundle block adjustment (BBA) method is proposed which utilizes a rational-polynomial-coefficient affine transformation model and a preconditioned conjugate gradient (PCG) algorithm with multisource satellite imagery (ZY-3 and GF-1) for producing and updating ortho maps of large areas. To deal with the weak geometry of the large blocks, a reference DEM is used in this method as an additional constraint in the BBA. The PCG algorithm is applied to solve the large normal matrix produced by the massive data of the large areas. Our proposed method was tested on three blocks of real data collected by GF-1 panchromatic and multispectral sensors and ZY-3 three-line-camera sensors. The preliminary results show that the proposed method can achieve an accuracy of better than 0.5 pixels in planimetry and is suitable for wide application in ortho-map production. It also has great potential for the ortho-map production of superlarge areas such as the country of China as one block. Maoteng Zheng, Yongjun Zhang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2016 | DEM-Assisted RFM Block Adjustment of Pushbroom Nadir Viewing HRS ImageryabstractNadir viewing satellite image is an effective data source to generate orthomosaics. Because of the georeferencing error of satellite images, block adjustment is the first step of orthomosaic generation over a large area. However, the geometric relationship of the neighboring orbits of the nadir viewing images is not rigid enough. This paper proposes a new rational function model (RFM) block adjustment approach that constrains the tie point elevation to enhance the relative geometric rigidity. By interpolating the elevations of tie points in a digital elevation model (DEM) and estimating the a priori errors of the interpolated elevations, better overall relative accuracy is obtained, and the local optimal solution problem is avoided. By constraining the adjusted model parameters according to the a priori error of RFMs, block adjustment without ground control point (GCP) is performed. By optimal initializing the object-space positions of tie points with multi-backprojection method, the needed iteration times of block adjustment are reduced. The proposed approach is investigated with 46 Ziyuan-3 sensor-corrected images, a 1:50 000 scale DEM, and 586 GCPs. Compared with Teo's approach that constrains the horizontal coordinates and elevations of tie points, the approach in this paper converges much faster when the GCPs are sparse, and meanwhile, the absolute and relative accuracy of the two approaches are almost the same. The result of block adjustment with only four GCPs shows that no accuracy degeneration occurred in the test area and the root-mean-square error of independent check point reaches about 1.5 ground resolutions. Different DEMs and number of tie points are used to investigate whether the block adjustment result is influenced by these factors. The results show that better DEM accuracy and denser tie points do improve the accuracy when the images have large side-sway angles. The proposed approach is also tested with 5118 IKONOS-2 images that cover the southern Europe without GCP. The result shows that the relative mosaicking accuracy is much better than that of Grodecki's approach. Yongjun Zhang 0002, Yi Wan 0001, Xinhui Huang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | An Improved Method for Transforming GPS/INS Attitude to National Map Projection FrameabstractGlobal Positioning System/Inertial Navigation System (GPS/INS) integrated navigation systems play a very important role in modern photogrammetry and laser scanning by virtue of their capability of direct measurement of high-precision position and attitude data in the WGS 84 datum. In practice, as georeferencing is often conducted in national coordinates, there is a need to transform GPS/INS data to the required national map projection frame first. This letter presents an improved coordinate-transformation-based method for the GPS/INS attitude transformation by taking the datum scale distortion and the length distortion into account. Experimental results show that the transformation errors of our improved method are on the order of magnitude of 1 × 10-5°, which can be safely ignored in aerial photogrammetric processing, whereas the maximum error of the previous coordinate-transformation-based method can be up to several 0.001°. Yongjun Zhang 0002, Qingquan Li 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2015 | Multistrip Bundle Block Adjustment of ZY-3 Satellite Imagery by Rigorous Sensor Model Without Ground Control PointabstractExtensive applications of Zi-Yuan 3 (ZY-3) satellite imagery of China have commenced since its on-orbit test was finished. Most of the data are processed scene by scene with a few ground control points (GCPs) for each scene; this conventional method is mature and widely used all over the world. However, very little work has focused on its application to super large area blocks. This letter aims to study mapping applications without GCPs for a super large area, which is defined as a range of interprovincial or even nationwide areas. The automatic matching and bundle block adjustment (BBA) software developed by our research team are applied to deal with two blocks of ZY-3 three-line camera imagery which covers most of the provinces in eastern China. Our comparison analysis of different data processing methods and the geolocation accuracies of the overlapped areas between adjacent strips are presented in this letter, as well as the possibility of nationwide BBA. The preliminary test results show that multistrip BBA without GCPs can achieve accuracies of about 13-15 m in both planimetry and height, which means that nationwide BBA is considered practical and feasible. Yongjun Zhang 0002, Maoteng Zheng, Xiaodong Xiong, Jinxin Xiong |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2015 | LiDAR Strip Adjustment Using Multifeatures Matched With Aerial ImagesabstractAirborne light detecting and ranging (LiDAR) systems have been widely used for the fast acquisition of dense topographic data. Regrettably, coordinate errors always exist in LiDAR-acquired points. The errors are attributable to several sources, such as laser ranging errors, sensor mounting errors, and position and orientation system (POS) systematic errors, among others. LiDAR strip adjustment (LSA) is the solution to eliminating the errors, but most state-of-the-art LSA methods neglect the influence from POS systematic errors by assuming that the POS is precise enough. Unfortunately, many of the LiDAR systems used in China are equipped with a low-precision POS due to cost considerations. Subsequently, POS systematic errors should be also considered in the LSA. This paper presents an aerotriangulation-aided LSA (AT-aided LSA) method whose major task is eliminating position and angular errors of the laser scanner caused by boresight angular errors and POS systematic errors. The aerial images, which cover the same area with LiDAR strips, are aerotriangulated and serve as the reference data for LSA. Two types of conjugate features are adopted as control elements (i.e., the conjugate points matched between the LiDAR intensity images and the aerial images and the conjugate corner features matched between LiDAR point clouds and aerial images). Experiments using the AT-aided LSA method are conducted using a real data set, and a comparison with the three-dimensional similarity transformation (TDST) LSA method is also performed. Experimental results support the feasibility of the proposed AT-aided LSA method and its superiority over the TDST LSA method. Yongjun Zhang 0002, Xiaodong Xiong, Maoteng Zheng, Xu Huang 0005 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | Self-Calibration Adjustment of CBERS-02B Long-Strip ImageryabstractDue to hardware limitations, such as the poor accuracy of its onboard Global Positioning System receiver and star tracks, the direct georeferencing accuracy of the China and Brazil Earth Resource Satellite 02B (CBERS-02B) by its onboard position and attitude measurements is less than 1000 m at times. Thus, the image data cannot be directly used in surveying applications. This paper presents a self-calibration bundle adjustment strategy to improve the georeferencing accuracy of the onboard high-resolution camera (HRC). An adequate number of automatically matched ground control points (GCPs) are used to perform the bundle adjustment. Both the systematic error compensation model and the orientation image model along with the interior self-calibration parameters are used in the bundle adjustment to eliminate the systematic errors. A self-calibration strategy is used to compensate for the time delay and integrated charge-coupled device translation and rotation errors by introducing a total of ten interior orientation parameters. The preliminary results show that the accuracy of self-calibration bundle adjustment is two pixels better than that of bundle adjustment without self-calibration, and the planimetric accuracy of the check points is about 10 m. The unusual variations of the exterior orientation parameters in some cases are eliminated after enlarging the orientation image intervals and increasing the weights of the onboard position and attitude observations. Maoteng Zheng, Yongjun Zhang 0002, Xiaodong Xiong |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2014 | Camera Pose Determination and 3-D Measurement From Monocular Oblique Images With Horizontal Right Angle ConstraintsabstractThis letter introduces a novel method for camera pose determination from monocular urban oblique images. Horizontal right angles that widely exist in urban scenes are used as geometric constraints in the camera pose determination, and the proposed 3-D measurement method using a monocular image is presented and then used to check the accuracy of the recovered image's exterior orientation parameters. Compared to the available vertical-line-based camera pose determination method, our new method is more accurate. Xiaodong Xiong, Yongjun Zhang 0002, Maoteng Zheng |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2014 | Quantitative Analysis on Geometric Size of LiDAR FootprintabstractA light detection and ranging (LiDAR) footprint is the spot area illuminated by a single laser beam, which varies with the beam direction and the regional terrain encountered. The geometric size of the LiDAR footprint is one of the most critical parameters of LiDAR point cloud data. It plays a very important role in the high-precision geometric and radiometric calibration of LiDAR systems. This letter utilizes space analytic geometry to derive LiDAR footprint equations and strictly considers laser beam attitude and terrain slope. Compared to the conventional plane geometry solution, the proposed approach is not only more rigorous in theory but also more powerful in practical applications. Yongjun Zhang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2014 | Road Centerline Extraction in Complex Urban Scenes From LiDAR Data Based on Multiple FeaturesabstractAutomatic extraction of roads from images of complex urban areas is a very difficult task due to the occlusions and shadows of contextual objects, and complicated road structures. As light detection and ranging (LiDAR) data explicitly contain direct 3-D information of the urban scene and are less affected by occlusions and shadows, they are a good data source for road detection. This paper proposes to use multiple features to detect road centerlines from the remaining ground points after filtering. The main idea of our method is to effectively detect smooth geometric primitives of potential road centerlines and to separate the connected nonroad features (parking lots and bare grounds) from the roads. The method consists of three major steps, i.e., spatial clustering based on multiple features using an adaptive mean shift to detect the center points of roads, stick tensor voting to enhance the salient linear features, and a weighted Hough transform to extract the arc primitives of the road centerlines. In short, we denote our method as Mean shift, Tensor voting, Hough transform (MTH). We evaluated the method using the Vaihingen and Toronto data sets from the International Society for Photogrammetry and Remote Sensing Test Project on Urban Classification and 3-D Building Reconstruction. The completeness of the extracted road network on the Vaihingen data and the Toronto data are 81.7% and 72.3%, respectively, and the correctness are 88.4% and 89.2%, respectively, yielding the best performance compared with template matching and phase-coded disk methods. Xiangyun Hu, Jie Shan, Jianqing Zhang, Yongjun Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2014 | On-Orbit Geometric Calibration of ZY-3 Three-Line Array Imagery With Multistrip Data SetsabstractZY-3, which was launched on January 9, 2012, is the first stereo mapping satellite in China. The initial accuracy of direct georeferencing with the onboard three-line camera (TLC) imagery is low. Sensor geometric calibration with bundle block adjustment is used to improve the georeferencing accuracy. A new on-orbit sensor calibration method that can correct the misalignment angles between the spacecraft and the TLC and the misalignment of charge-coupled device is described. All of the calibration processes are performed using a multistrip data set. The control points are automatically matched from existing digital ortho map and digital elevation model. To fully evaluate the accuracy of different calibration methods, the calibrated parameters are used as input data to conduct georeferencing and bundle adjustment with a total of 19 strips of ZY-3 TLC data. A systematic error compensation model is introduced as the sensor model in bundle adjustment to compensate for the position and attitude errors. Numerous experiments demonstrate that the new calibration model can largely improve the external accuracy of direct georeferencing from the kilometer level to better than 20 m in both plane and height. A further bundle block adjustment with medium-accuracy ground control points (GCPs), using these calibrated parameters, can achieve external accuracy of about 4 m in plane and 3 m in height. Higher accuracy of about 1.3 m in plane and 1.7 m in height can be achieved by bundle adjustment using high-accuracy GCPs. Yongjun Zhang 0002, Maoteng Zheng, Jinxin Xiong, Yihui Lu, Xiaodong Xiong |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2013 | Fast Filtering of LiDAR Point Cloud in Urban Areas Based on Scan Line Segmentation and GPU AccelerationabstractThe fast filtering of massive point cloud data from light detection and ranging (LiDAR) systems is important for many applications, such as the automatic extraction of digital elevation models in urban areas. We propose a simple scan-line-based algorithm that detects local lowest points first and treats them as the seeds to grow into ground segments by using slope and elevation. The scan line segmentation algorithm can be naturally accelerated by parallel computing due to the independent processing of each line. Furthermore, modern graphics processing units (GPUs) can be used to speed up the parallel process significantly. Using a strip of a LiDAR point cloud, with up to 48 million points, we test the algorithm in terms of both error rate and time performance. The tests show that the method can produce satisfactory results in less than 0.6 s of processing time using the GPU acceleration. Xiangyun Hu, Yongjun Zhang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2013 | Approximate Correction of Length Distortion for Direct Georeferencing in Map Projection FrameabstractMany geometric distortions, such as earth curvature distortion and length distortion, exist in the map projection frame. Therefore, in aerial photogrammetry, if direct georeferencing is performed in the map projection frame, the crucial work becomes compensating for the effect of the various geometric distortions. This letter mainly focuses on length distortion and proposes two new correction approaches: the changing image coordinates method and the changing object coordinates method. The experimental results show that the changing object coordinates method is less influenced by terrain fluctuation, and its correction accuracy is therefore commonly higher than the changing image coordinates method as well as two existing approaches (i.e., the changing flight height method and the changing focal length method). Yongjun Zhang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2012 | A New Approach on Optimization of the Rational Function Model of High-Resolution Satellite ImageryabstractOverparameterization is one of the major problems that the rational function model (RFM) faces. A new approach of RFM parameter optimization is proposed in this paper. The proposed RFM parameter optimization method can resolve the ill-posed problem by removing all of the unnecessary parameters based on scatter matrix and elimination transformation strategies. The performances of conventional ridge estimation and the proposed method are evaluated with control and check grids generated from Satellites d'observation de la Terre (SPOT-5) high-resolution satellite data. Experimental results show that the precision of the proposed method, with about 35 essential parameters, is 10% to 20% higher than that of the conventional model with all 78 parameters. Moreover, the ill-posed problem is effectively alleviated by the proposed method, and thus, the stability of the estimated parameters is significantly improved. Yongjun Zhang 0002, Yihui Lu, Xu Huang 0005 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2006 | Automatic measurement of industrial sheetmetal parts with CAD data and non-metric image sequence
Yongjun Zhang 0002, Zuxun Zhang, Jianqing Zhang |
Comput. Vis. Image Underst. | 1 |
| 2005 | Automatic Inspection of Industrial Sheetmetal Parts with Single Non-metric CCD Camera
Yongjun Zhang 0002 |
ADMA | 1 |
| 2004 | Deformation visual inspection of industrial parts with image sequence
Yongjun Zhang 0002, Zuxun Zhang, Jianqing Zhang |
Mach. Vis. Appl. | 1 |