Yanpeng Cao

dblp:91/7629 · DBLP profile ↗
← Back
36ranked-venue papers
12as first author
17since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Systems, architecture and hardware · 3 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2025 A hybrid perceptron with cross-domain transferability towards active steady-state non-line-of-sight imaging
Xi Tong, Jiangxin Yang, Yanpeng Cao
Signal Process.4
2025 SC3EF: A Joint Self-Correlation and Cross-Correspondence Estimation Framework for Visible and Thermal Image Registration
abstract
Multispectral imaging plays a critical role in a range of intelligent transportation applications, including advanced driver assistance systems (ADAS), traffic monitoring, and night vision. However, accurate visible and thermal (RGB-T) image registration poses a significant challenge due to the considerable modality differences. In this paper, we present a novel joint Self-Correlation and Cross-Correspondence Estimation Framework (SC3EF), leveraging both local representative features and global contextual cues to effectively generate RGB-T correspondences. For this purpose, we design a convolution-transformer-based pipeline to extract local representative features and encode global correlations of intra-modality for inter-modality correspondence estimation between unaligned visible and thermal images. After merging the local and global correspondence estimation results, we further employ a hierarchical optical flow estimation decoder to progressively refine the estimated dense correspondence maps. Extensive experiments demonstrate the effectiveness of our proposed method, outperforming the current state-of-the-art (SOTA) methods on representative RGB-T datasets. Furthermore, it also shows competitive generalization capabilities across challenging scenarios, including large parallax, severe occlusions, adverse weather, and other cross-modal datasets (e.g., RGB-N and RGB-D).
Xi Tong, Jiangxin Yang, Xin Li 0003, Yanpeng Cao
IEEE Trans. Intell. Transp. Syst.5
2023 Light field angular super-resolution based on structure and scene information
Jiangxin Yang, Lingyu Wang 0005, Lifei Ren, Yanpeng Cao, Yanlong Cao
Appl. Intell.4
2023 A sub-region Unet for weak defects segmentation with global information and mask-aware loss
Jiangxin Yang, Yanlong Cao, Guizhong Fu, Yanpeng Cao
Eng. Appl. Artif. Intell.6
2023 Infrared and visible image fusion based on a two-stage class conditioned auto-encoder network
Yanpeng Cao, Xi Tong, Jiangxin Yang, Yanlong Cao
Neurocomputing1
2023 A deep thermal-guided approach for effective low-light visible image enhancement
Yanpeng Cao, Xi Tong, Fan Wang 0022, Jiangxin Yang, Yanlong Cao, Sabin Tiberius Strat, Christel-Loïc Tisse
Neurocomputing1
2023 View position prior-supervised light field angular super-resolution network with asymmetric feature extraction and spatial-angular interaction
Yanlong Cao, Lingyu Wang 0005, Lifei Ren, Jiangxin Yang, Yanpeng Cao
Neurocomputing5
2023 Light field angular super-resolution based on intrinsic and geometric information
Lingyu Wang 0005, Lifei Ren, Xiaoyao Wei, Jiangxin Yang, Yanlong Cao, Yanpeng Cao
Knowl. Based Syst.6
2023 Single image super-resolution based on progressive fusion of orientation-aware features
Zewei He, Yanpeng Cao, Jiangxin Yang, Yanlong Cao, Xin Li 0003, Siliang Tang, Yueting Zhuang, Zheming Lu 0001
Pattern Recognit.3
2023 Multi-Modal Image Fusion via Deep Laplacian Pyramid Hybrid Network
abstract
Fusion of images acquired using different sensors generates a single output with enhanced information for high-level visual perception applications. The transformer architecture has demonstrated its powerful ability to obtain important global contextual dependencies for multi-modal image fusion tasks. However, transformer-based image fusion methods face many critical issues, such as incurring huge computational burdens, limited ability to learn local features, and the difficulty of handling images of arbitrary sizes. To address the above limits, we proposed a novel Laplacian Pyramid Hybrid (LapH) network to combine the advantages of CNN and transformer architectures for multi-modal image fusion tasks. With the divide-and-conquer philosophy, we first build a light-weight CNN-based branch, performing effective extraction and fusion of texture/edge features via central difference convolutions, to process the high-resolution components with abundant details encoded in the lower pyramid levels of the Laplacian pyramid. Then, we design a transformer-based branch to process the low-resolution base components, learning long-range dependencies of global-contextual features without incurring extensive computational loads. Here, we design a multi-scale recurrent modulation mechanism to integrate the edge/texture features from the CNN branch as guidance to progressively refine the feature extraction and fusion on low-frequency components. Finally, we propose a new multi-scale spatial consistency loss term based on the neighbor contrast in source images, generating fused images with more natural and realistic appearances. Extensive experiments on two different multi-modal image fusion tasks verify the superiority of our method. The source codes are made publicly available athttps://github.com/rgttadv/LapH.
Guizhong Fu, Jiangxin Yang, Yanlong Cao, Yanpeng Cao
IEEE Trans. Circuits Syst. Video Technol.5
2023 PPI Edge Infused Spatial-Spectral Adaptive Residual Network for Multispectral Filter Array Image Demosaicing
abstract
Multispectral filter array (MSFA) sensors provide a cost-effective and one-shot acquisition solution to obtain well-aligned multi-band images, which are helpful for various optical and remote sensing applications. However, the sparse spatial sampling rate and strong spectral cross-correlation make MSFA image demosaicing a challenging problem. Therefore, it is essential to develop effective MSFA demosaicing solutions to reconstruct full-resolution and high-fidelity multispectral images from the raw mosaic image. In this paper, we present a Pseudo-panchromatic Image (PPI) Edge infused Spatial-Spectral Adaptive Residual Network (PPIE-SSARN) for multispectral filter array image demosaicing. The proposed two-branch model deploys a residual sub-branch to adaptively compensate for the spatial and spectral differences of reconstructed multispectral images and a PPI edge infusion sub-branch to enrich the edge-related information. Moreover, we design an effective mosaic initial feature extraction module with a spatial- and spectral-adaptive weight-sharing strategy whose kernel weights can change adaptively with spatial locations and spectral bands to avoid artifacts and aliasing problems. Experimental results demonstrate the superiority of our proposed method, outperforming the state-of-the-art MSFA demosaicing approaches and achieving satisfying demosaicing results in terms of spatial accuracy and spectral fidelity. Our models and code will be publicly available.
Jiesi Zheng, Yafei Dong, Jiangxin Yang, Yanlong Cao, Yanpeng Cao
IEEE Trans. Geosci. Remote. Sens.7
2023 Multi-Range View Aggregation Network With Vision Transformer Feature Fusion for 3D Object Retrieval
abstract
View-based methods have achieved state-of-the-art performance in 3D object retrieval. However, view-based methods still encounter two major challenges. The first is how to leverage the inter-view correlation to enhance view-level visual features. The second is how to effectively fuse view-level features into a discriminative global descriptor. Towards these two challenges, we propose a multi-range view aggregation network (MRVANet) with a vision transformer based feature fusion scheme for 3D object retrieval. Unlike the existing methods which only consider aggregating neighboring or adjacent views which could bring in redundant information, we propose a multi-range view aggregation module to enhance individual view representations through view aggregation beyond only neighboring views but also incorporate the views at different ranges. Furthermore, to generate the global descriptor from view-level features, we propose to employ the multi-head self-attention mechanism introduced by vision transformer to fuse the view-level features. Extensive experiments conducted on three public datasets including ModelNet40, ShapeNet Core55 and MCB-A demonstrate the superiority of the proposed network over the state-of-the-art methods in 3D object retrieval.
Dongyun Lin, Shitala Prasad, Aiyuan Guo, Yanpeng Cao
IEEE Trans. Multim.6
2022 Spatio-Temporal 3-D Residual Networks for Simultaneous Detection and Depth Estimation of CFRP Subsurface Defects in Lock-In Thermography
abstract
Nondestructive thermography is a high-speed, low-cost, and safe solution for subsurface defects detection of carbon fiber reinforced polymer (CFRP) materials, providing essential quality control in aerospace, automobile, and sports industries. In this article, we build a reflective lock-in thermography system and construct a dataset that contains real-captured thermal image sequences of CFRP samples with various simulated internal defects under different excitation frequencies. Then, we present a novel 3-D convolutional neural network (CNN) model incorporating a combination of spatial and temporal convolutional filters and batch-size independent group normalization (GN) as a unified framework to process thermal image sequences captured by lock-in thermography for simultaneous subsurface defect detection and depth estimation. Finally, we define a multitask loss function to perform end-to-end training of both defect detection and depth estimation tasks based on the real-captured infrared sequences. Comparative experiments are carried out on CFRP specimens with artificial defects of various sizes/shapes and at different depths. Qualitative and quantitative results illustrate that our 3-D CNN model is capable of predicting accurate locations and depths of subsurface defects and performs favorably against the hand-crafted and CNN-based methods in lock-in thermography for individual defect detection and depth estimation tasks. The captured dataset and the source codes will be made publicly available.
Yafei Dong, Chenjie Xia, Jiangxin Yang, Yanlong Cao, Yanpeng Cao, Xin Li 0003
IEEE Trans. Ind. Informatics5
2022 Uncertainty-Aware Unsupervised Domain Adaptation in Object Detection
abstract
Unsupervised domain adaptive object detection aims to adapt detectors from a labelled source domain to an unlabelled target domain. Most existing works take a two-stage strategy that first generates region proposals and then detects objects of interest, where adversarial learning is widely adopted to mitigate the inter-domain discrepancy in both stages. However, adversarial learning may impair the alignment of well-aligned samples as it merely aligns the global distributions across domains. To address this issue, we design an uncertainty-aware domain adaptation network (UaDAN) that introduces conditional adversarial learning to align well-aligned and poorly-aligned samples separately in different manners. Specifically, we design an uncertainty metric that assesses the alignment of each sample and adjusts the strength of adversarial learning for well-aligned and poorly-aligned samples adaptively. In addition, we exploit the uncertainty metric to achieve curriculum learning that first performs easier image-level alignment and then more difficult instance-level alignment progressively. Extensive experiments over four challenging domain adaptive object detection datasets show that UaDAN achieves superior performance as compared with state-of-the-art methods.
Dayan Guan, Jiaxing Huang 0001, Aoran Xiao, Shijian Lu, Yanpeng Cao
IEEE Trans. Multim.5
2021 Real-Time Super-Resolution System of 4K-Video Based on Deep Learning
abstract
Video super-resolution (VSR) technology excels in reconstructing low-quality video, avoiding unpleasant blur effect caused by interpolation-based algorithms. However, vast computation complexity and memory occupation hampers the edge of deplorability and the runtime inference in real-life applications, especially for large-scale VSR task. This paper explores the possibility of real-time VSR system and designs an efficient and generic VSR network, termed EGVSR. The proposed EGVSR is based on spatio-temporal adversarial learning for temporal coherence. In order to pursue faster VSR processing ability up to 4K resolution, this paper tries to choose lightweight network structure and efficient upsampling method to reduce the computation required by EGVSR network under the guarantee of high visual quality. Besides, we implement the batch normalization computation fusion, convolutional acceleration algorithm and other neural network acceleration techniques on the actual hardware platform to optimize the inference process of EGVSR network. Finally, our EGVSR achieves the real-time processing capacity of [email protected]. Compared with TecoGAN, the most advanced VSR network at present, we achieve 85.04% reduction of computation density and 7.92× performance speedups. In terms of visual quality, the proposed EGVSR tops the list of most metrics (such as LPIPS, tOF, tLP, etc.) on the public test dataset Vid4 and surpasses other state-of-the-art methods in overall performance score.
Yanpeng Cao, Changjun Song, Yongming Tang, He Li 0008
ASAP1
2021 Few-Shot Defect Segmentation Leveraging Abundant Defect-Free Training Samples Through Normal Background Regularization And Crop-And-Paste Operation
abstract
In industrial quality assessment, it is challenging to conduct automated and accurate defect segmentation under the condition that abundant defect-free images but very limited anomalous images are available. This paper tackles the challenging few-shot defect segmentation task under such condition. We propose two regularization techniques via incorporating abundant defect-free images into the training of an encoder-decoder segmentation network. We first propose a Normal Background Regularization (NBR) loss which is jointly minimized with the segmentation loss, enhancing the encoder network to produce discriminative representations for normal regions. Secondly, we crop/paste defective regions to the randomly selected normal images for data augmentation and propose a weighted binary cross-entropy loss to enhance the training by emphasizing more realistic crop-and-pasted augmented images based on feature-level similarity comparison. Extensive experiments on MVTec AD and MTSD datasets demonstrate the superiority of the proposed method over the competing methods under few-shot settings.
Dongyun Lin, Yanpeng Cao
ICME2
2021 ESKN: Enhanced selective kernel network for single image super-resolution
Zewei He, Guizhong Fu, Yanpeng Cao, Yanlong Cao, Jiangxin Yang, Xin Li 0003
Signal Process.3
2020 Explore Efficient LUT-based Architecture for Quantized Convolutional Neural Networks on FPGA
abstract
The vast computations of the convolutional neural network have limited the speed of the forward inference running in hardware. In recent years, network quantization technique has made it possible to quantize network into low bit-wide and retain the original performance simultaneously, while the complexity of the quantized network is still considerable. FPGA is a highly parallelized platform, which contains a mass of configurable logic resources. We study on the feasibility of implementing convolution calculation based on pure LUTs, introduce the shift multipliers and addition trees, and propose an efficient architecture for QNN on FPGA. With the optimization of Winograd algorithm for QNN, we demonstrate that our scheme significantly reduces the number of multipliers and saves the usage of LUT resources by $2.25 \times $ at least without using DSP resources. As a result, our LUT-based architecture for QNN shortens the latency up to $19.3 \times $ and represents more effective performance compared to other methods.
Yanpeng Cao, Yongming Tang
FCCM1
2020 Realization of Quantized Neural Network for Super-resolution on PYNQ
abstract
Vision tasks usually require vast amount of computation and memory resources, which create barriers to edge computing applications. Quantized neural network can provide memory saving, scalability and energy efficiency, while the accuracies of results may decrease. In this paper, we adjust the data-width of feature maps, weights and temporary variables in SRCNN to achieve a trade-off between precision and accuracy. Also, we design a dedicated convolutional acceleration for data stream under the heterogeneous CPU-FPGA platform: PYNQ, including the changed data streaming order and im2col for convolution. Results show when data-width was set to 12-bit, quantization had almost no effect on visual perception of superresolved images. The acceleration of quantized convolution on FPGA can achieve a speed up ratio of 120x at 250MHz, compared with ARM CPU.
Feng Yu 0006, Yanpeng Cao, Yongming Tang
FCCM2
2020 Photo Stream Question Answer
abstract
Understanding and reasoning over partially observed visual clues are often regarded as a challenging real-world problem even for human beings. In this paper, we present a new visual question answering (VQA) task -- Photo Stream QA, which aims to answer the open-ended questions about a narrative photo stream. Photo Stream QA is more challenging and interesting than the existing VQA tasks, since the temporal and visual variance among photos in the stream is huge and hard to observe. Therefore, instead of learning simple vision-text mappings, the AI algorithms must fill these variance gaps with more recollection, reasoning, even the knowledge from our daily experiences. To tackle the problems in Photo Stream QA, we propose an end-to-end baseline (E-TAA) with a novel Experienced Unit (E-unit) and Three-stage Alternating Attention (TAA). E-unit yields a better visual representation which captures the temporal semantic relation among visual clues in the photo stream, while TAA creates three levels of attention that gradually refines visual features by using the textual representation from the question as the guidance. Experimental results on our developed dataset demonstrate that, as the first attempt at the Photo Stream QA task, E-TAA provides promising results outperforming all the other baseline methods.
Wenqiao Zhang, Siliang Tang, Yanpeng Cao, Jun Xiao 0001, Shiliang Pu, Fei Wu 0001, Yueting Zhuang
ACM Multimedia3
2020 WaterNet: An adaptive matching pipeline for segmenting water with volatile appearance
abstract
We develop a novel network to segment water with significant appearance variation in videos. Unlike existing state-of-the-art video segmentation approaches that use a pre-trained feature recognition network and several previous frames to guide segmentation, we accommodate the object’s appearance variation by considering features observed from the current frame. When dealing with segmentation of objects such as water, whose appearance is non-uniform and changing dynamically, our pipeline can produce more reliable and accurate segmentation results than existing algorithms.
Yongqing Liang 0001, Navid H. Jafari, Qin Chen 0003, Yanpeng Cao, Xin Li 0003
Comput. Vis. Media5
2020 MRFN: Multi-Receptive-Field Network for Fast and Accurate Single Image Super-Resolution
abstract
Recently, convolutional neural network (CNN) based models have shown great potential in the task of single image superresolution (SISR). However, many state-of-the-art SISR solutions are reproducing some tricks proven effective in other vision tasks, such as pursuing a deeper model. In this paper, we propose a new solution (named as Multi-Receptive-Field Network - MRFN), which outperforms existing SISR solutions in three different aspects. First, from receptive field: a novel multi-receptive-field (MRF) module is proposed to extract and fuse features in different receptive fields from local to global. Integrating these hierarchical features can generate better mappings on recovering high-fidelity details at different scales. Second, from network architectures: both dense skip connections and deep supervision are utilized to combine features from the current MRF module and preceding ones for training more representative features. Moreover, a deconvolution layer is embedded at the end of the network to avoid artificial priors induced by numerical data pre-processing (e.g., bicubic stretching), and speed up the restoration process. Finally, from error modeling: different from L1 and L2 loss functions, we proposed a novel two-parameter training loss called Weighted Huber loss function which can adaptively adjust the value of back-propagated derivative according to the residual value, thus fit the reconstruction error more effectively. Extensive qualitative and quantitative evaluation results on benchmark datasets demonstrate that our proposed MRFN can achieve more accurate recovering results than most state-of-the-art methods with significantly less complexity.
Zewei He, Yanpeng Cao, Baobei Xu, Jiangxin Yang, Yanlong Cao, Siliang Tang, Yueting Zhuang
IEEE Trans. Multim.2
2020 Frame Augmented Alternating Attention Network for Video Question Answering
abstract
Vision and language understanding is one of the most fundamental and challenging problems in Multimedia Intelligence. Simultaneously understanding video actions with a related natural language question, and further produces accurate answer is even more challenging since it requires joint modeling information across modality. In the past few years, some studies begin to attack this problem by utilizing attention enhanced deep neural networks. However, simple attention mechanisms such as unidirectional attention fail to yield a better mapping between different modalities. Moreover, none of these Video QA models explore high-level semantics in augmented video-frame level. In this paper, we augmented each frame representation with its context information by a novel feature extractor that combines the advantages of Resnet and a variant of C3D. In addition, we proposed a novel alternating attention network which can alternately attend frame regions, video frames and words in the question in multi-turns. This yields better joint representations of video and question, further help the deep model to discover the deeper relationship between two modalities. Our method outperforms the state-of-the-art Video QA models on two existing video question answering datasets. Further ablation studies proved that our feature extractor and the alternating attention mechanism can improve the performance jointly.
Wenqiao Zhang, Siliang Tang, Yanpeng Cao, Shiliang Pu, Fei Wu 0001, Yueting Zhuang
IEEE Trans. Multim.3
2019 Deep Learning for Semantic Segmentation of UAV Videos
abstract
As one of the key problems in both remote sensing and computer vision, video semantic segmentation has been attracting increasing amounts of attention. Using video segmentation technique for Unmanned Aerial Vehicle (UAV) data processing is also a popular application. Previous methods extended single image segmentation approaches to multiple frames. The temporal dependencies are ignored in these methods. This paper proposes a novel segmentation method to solve this problem. Combining the fully convolutional networks (FCN) and the Convolution Long Short Term Memory (Conv-LSTM) together, we segment the sequence of the video frames instead of segmenting each individual frame separately. FCN serves as the frame-based segmentation method. Conv-LSTM makes use of the temporal information between consecutive frames. Experimental results show the superiority of this method especially in some classes compared to the single image segmentation model using video dataset from UAV.
Ye Lyu, Yanpeng Cao, Michael Ying Yang
IGARSS3
2019 Fast and accurate single image super-resolution via an energy-aware improved deep residual network
Yanpeng Cao, Zewei He, Zhangyu Ye, Xin Li 0003, Yanlong Cao, Jiangxin Yang
Signal Process.1
2019 Accurate salient object detection via dense recurrent connections and residual-based hierarchical feature integration
Yanpeng Cao, Guizhong Fu, Jiangxin Yang, Yanlong Cao, Michael Ying Yang
Signal Process. Image Commun.1
2019 Cascaded Deep Networks With Multiple Receptive Fields for Infrared Image Super-Resolution
abstract
Infrared images have a wide range of military and civilian applications, including night vision, surveillance, and robotics. However, high-resolution infrared detectors are difficult to fabricate and their manufacturing cost is expensive. In this paper, we present a cascaded architecture of deep neural networks with multiple receptive fields to increase the spatial resolution of infrared images by a large scale factor (x8). Instead of reconstructing a high-resolution image from its low-resolution version using a single complex deep network, the key idea of our approach is to set up a mid-point (scale x2) between scale x1 and x8 such that lost information can be divided into two components. Lost information within each component contains similar patterns thus can be more accurately recovered even using a simpler deep network. In our proposed cascaded architecture, two consecutive deep networks with different receptive fields are jointly trained through a multi-scale loss function. The first network with a large receptive field is applied to recover largescale structure information, while the second one uses a relatively smaller receptive field to reconstruct small-scale image details. Our proposed method is systematically evaluated using realistic infrared images. Compared with state-of-the-art super-resolution methods, our proposed cascaded approach achieves improved reconstruction accuracy using significantly fewer parameters.
Zewei He, Siliang Tang, Jiangxin Yang, Yanlong Cao, Michael Ying Yang, Yanpeng Cao
IEEE Trans. Circuits Syst. Video Technol.6
2018 A multi-scale non-uniformity correction method based on wavelet decomposition and guided filtering for uncooled long wave infrared camera
Yanlong Cao, Zewei He, Jiangxin Yang, Xiaoping Ye, Yanpeng Cao
Signal Process. Image Commun.5
2016 Bi-layer dictionary learning for remote sensing image classification
abstract
With the widely application of high-resolution remote sensing images, its classification has attracted a lot of attention. Usually, some different categories share common patterns, which make these categories look similar. This makes the classification of such categories a challenging task. In this paper, we propose a novel dictionary learning based bilayer classification algorithm to solve this problem. Using SIFT descriptor, instead of directly classifying an image, we separate the classification in two steps. In the first step, the similar categories are clustered to be as a new category for the first classification layer. In this step, the inter-class variation are maximized. The second layer is designed to classify the similar categories clustered in the same group. Experimental results show the superiority of our method compared to the state-of-the-art methods using UCMerced LandUse dataset.
Michael Ying Yang, Saif Dawood Salman Al-Shaikhli, Yanpeng Cao, Bodo Rosenhahn
IGARSS4
2016 Effective Strip Noise Removal for Low-Textured Infrared Images Based on 1-D Guided Filtering
abstract
Infrared images typically contain obvious strip noise. It is a challenging task to eliminate such noise without blurring fine image details in low-textured infrared images. In this paper, we introduce an effective single-image-based algorithm to accurately remove strip-type noise present in infrared images without causing blurring effects. First, a 1-D row guided filter is applied to perform edge-preserving image smoothing in the horizontal direction. The extracted high-frequency image part contains both strip noise and a significant amount of image details. Through a thermal calibration experiment, we discover that a local linear relationship exists between infrared data and strip noise of pixels within a column. Based on the derived strip noise behavioral model, strip noise components are accurately decomposed from the extracted high-frequency signals by applying a 1-D column guided filter. Finally, the estimated noise terms are subtracted from the raw infrared images to remove strips without blurring image details. The performance of the proposed technique is thoroughly investigated and is compared with the state-of-the-art 1-D and 2-D denoising algorithms using captured infrared images.
Yanpeng Cao, Michael Ying Yang, Christel-Loïc Tisse
IEEE Trans. Circuits Syst. Video Technol.1
2015 Learning human photo shooting patterns from large-scale community photo collections
Yanpeng Cao, Kay O'Halloran
Multim. Tools Appl.1
2015 Descriptor evaluation and feature regression for multimodal image analysis
Xuanzi Yong, Michael Ying Yang, Yanpeng Cao, Bodo Rosenhahn
Mach. Vis. Appl.3
2012 Improved feature extraction and matching in urban environments based on 3D viewpoint normalization
Yanpeng Cao, John McDonald 0001
Comput. Vis. Image Underst.1
2011 Robust alignment of wide baseline terrestrial laser scans via 3D viewpoint normalization
abstract
The complexity of natural scenes and the amount of information acquired by terrestrial laser scanners turn the registration among scans into a complex problem. This problem becomes even more challenging when two individual scans captured at significantly changed viewpoints (wide baseline). Since laser-scanning instruments nowadays are often equipped with an additional image sensor, it stands to reason making use of the image content to improve the registration process of 3D scanning data. In this paper, we present a novel improvement to the existing feature techniques to enable automatic alignment between two widely separated 3D scans. The key idea consists of extracting dominant planar structures from 3D point clouds and then utilizing the recovered 3D geometry to improve the performance of 2D image feature extraction and matching. The resulting features are very discriminative and robust to perspective distortions and viewpoint changes due to exploiting the underlying 3D structure. Using this novel viewpoint invariant feature, the corresponding 3D points are automatically linked in terms of wide baseline image matching. Initial experiments with real data demonstrate the potential of the proposed method for the challenging wide baseline 3D scanning data alignment tasks.
Yanpeng Cao, Michael Ying Yang, John McDonald 0001
WACV1
2009 Robust feature correspondences from a large set of unsorted wide baseline images
abstract
Given a set of unordered images taken in a wide area, an effective solution is proposed for establishing robust feature correspondences among them. Two major improvements are made in our work as follows: firstly, a robust technique is proposed for the self-organization of a large number of images without spatial orderings; secondly, a novel wide-baseline matching approach is developed to obtain good correspondences over images taken from substantially different viewpoints. The output consists of many sets of reliable pair-wise feature correspondences which are essential in various computer vision applications. Realistic experiments were carried out to evaluate the performances of the proposed method by using a large amount of images captured from our university's campus.
Yanpeng Cao, John McDonald 0001
ICIP1
2009 Viewpoint invariant features from single images using 3D geometry
abstract
In this paper we present a novel approach for generating viewpoint invariant features from single images and demonstrate their application for robust matching over widely separated views. The key idea consists of retrieving building structure from single images and then utlising the recovered 3D geometry to improve the performances of feature extraction and matching. Urban environments usually contain many structured regularities, so that the images of those environments contain straight parallel lines and vanishing points, which can be efficiently exploited for 3D reconstruction. We present an effective scheme to recover 3D planar surfaces using the extracted line segments and their associated vanishing points. The viewpoint invariant features are then computed on the normalized front-parallel views of the obtained 3D planes. The advantages of the proposed approach include: (1) the new feature is very robust against perspective distortions and viewpoint changes due to its consideration of 3D geometry; (2) the features are completely computed from single images and do not need information from additional devices (e.g. stereo cameras, or active ranging devices). Experiments are carried out to demonstrate the proposed scheme ability to effectively handle very difficult wide baseline matching tasks in the presence of repetitive building structures and significant viewpoint changes.
Yanpeng Cao, John McDonald 0001
WACV1