Jian Yao 0002

dblp:40/4105-2 · DBLP profile ↗
← Back
46ranked-venue papers
0as first author
16since 2021 · last 2026
0000-0002-9134-5084ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 7 since 2021Artificial intelligence and machine learning · 16 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 5 since 2021Systems, architecture and hardware · 6 · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-view multilayer perceptron attention for 3D object recognition
Yeduozi Yao, Jiaqin Jiang, Li Li 0047, Jian Yao 0002
J. Vis. Commun. Image Represent.5
2026 SMDGS: Scale-Aligned Monocular Depth-Guided 3D Gaussian Splatting for Rendering and Surface Reconstruction
abstract
3D Gaussian Splatting (3DGS) has been explored for surface reconstruction, however unstructured and discontinuous Gaussian point clouds lead to uneven surface reconstruction accuracy as well as frequent loss of Novel View Synthesis (NVS) quality. To address this problem, we propose a scale-aligned monocular depth-guided 3DGS, a promising novel framework that combines geometric prior regularization and consistency supervision to achieve high-quality rendering and surface reconstruction. Specifically, monocular depth, estimated by some recently proposed monocular depth estimation models, contain implicitly abundant valuable geometric cues, but scale ambiguity limits its application. Therefore we first propose a $K$K-Nearest Neighbor (KNN)-based depth alignment framework that utilizes the full-domain gradient at monocular depth map to align to the sparse point cloud obtained during the Structure from Motion (SfM), which is employed for regularization to enhance geometric representation. Then a pseudo-mesh-based multi-view consistency module is introduced to fine-tune and guide the model to recover the accurate surface. Finally, a pixel-level isotropic gradient aware method guides the appropriate growth of the Gaussians to further improve the surface and rendering quality. Experiments on dozens of indoor, outdoor, and object-centered/non-object-centered datasets demonstrate that our method achieves accurate surface reconstruction with excellent NVS performance.
Xiaosong Wei, Pengwei Zhou, Annan Zhou, Li Li 0047, Jian Yao 0002
IEEE Trans. Vis. Comput. Graph.7
2026 SGGS: Semantic-Guided 3D Gaussian Splatting With Adaptive Rendering
abstract
3D Gaussian Splatting (3DGS) has shown great promise in a variety of applications due to its exceptional real-time rendering quality and explicit representation, leading to numerous improvements across various fields. However, existing methods lack consideration of main objects and important structural information in their overall optimization strategies. This results in blurring of main objects in adaptive rendering and the loss of high-frequency details on targets that are insufficiently captured. In this work, we introduce a semantic-guided 3DGS method with adaptive rendering, which optimizes important structures through the guidance of boundary Gaussians, while leveraging semantic features to enhance the rendering of main objects in adaptive rendering. Experiments show that the proposed semantic-guided method can enhance important structures and high-frequency information in corner regions without significantly increasing the total number of Gaussians. This method also improves the separability between objects. At the same time, a semantic-guided Level-of-Detail (LoD) rendering approach enables the rapid display of main targets and the rendering of a complete scene. The semantic-guided methodology we have presented exhibits compatibility with a range of existing techniques.
Annan Zhou, Li Li 0047, Jian Yao 0002
IEEE Trans. Vis. Comput. Graph.6
2025 JPG-SLAM: Joint Point-Gaussian Splatting Representation for Dense Dynamic SLAM
abstract
This paper presents a simultaneous localization and mapping (SLAM) system to provide accurate pose estimation and dynamic scene reconstruction. Our approach proposes a Joint Point-Gaussian Splatting representation, which fully integrates the robustness of isotropic feature points in pose estimation and the flexibility of anisotropic 3D Gaussians in scene representation. This system does not need to suppress the anisotropic representation of Gaussian elements, which enables the mapping module to achieve finer scene representation with lower memory consumption. Additionally, in order to enhance the adaptability of the system in dynamic environments, we introduced a dynamic region recognition module and utilized 3D Gaussian Splatting and 4D Gaussian Splatting representations to represent static and dynamic regions respectively. Furthermore, we developed a local map management strategy for Gaussian Splatting mapping, effectively reducing the memory and computational resource usage in the mapping process. Experiments on public datasets demonstrate that our system achieves state-of-the-art tracking and mapping accuracy compared to existing baselines.
Kunrui Huang, Wennan Yang, Pengwei Zhou, Li Li 0047, Jian Yao 0002
ICRA5
2025 A Coarse-to-Fine Boundary Relabeling Approach for Roof Plane Segmentation
abstract
Building roof plane segmentation is important for three-dimensional (3D) building model reconstruction from airborne light detection and ranging (LiDAR) point data. During the roof plane segmentation, challenges such as pseudo planes, over- and under-segmentation often arise, particularly evident in the boundary regions. To improve the accuracy of plane segmentation, various energy optimization-based methods have been proposed to refine the roof planes. However, the existing methods optimize the energy function at the point level, which may lead to getting stuck in local optima. To address these problems, we propose a coarse-to-fine boundary relabeling approach for roof plane segmentation. Starting from an initial plane segmentation result, the proposed method iteratively refines the planes by adjusting the boundaries from the voxel level to the point level. In addition, we also design a new energy function that considers accurateness, smoothness and compactness to guide the optimization. The experimental results constructed on two datasets demonstrate that the proposed method outperforms the existing roof plane segmentation methods, achieving high accuracy and smooth boundary extraction. The source code of the proposed approach will be publicly available at https://github.com/Li-Li-Whu/Coarse2FineRoofPlane.
Guozheng Xu, Siyuan You, Li Li 0047, Jian Yao 0002
IEEE Geosci. Remote. Sens. Lett.5
2025 A Parallelizable Global Color Consistency Optimization Algorithm for Multiple Images
abstract
The global optimization-based color correction approach aims to minimize the color differences of multiple images by optimizing the correction model for each image. The color differences in multisource and multitemporal remote sensing images are difficult to express using a simple correction model with few parameters. When employing a more flexible correction model, the number of correction parameters and optimization equations grows rapidly with the increase in the number and resolution of input images. In addition, the correction parameters of all images are coupled together and need to be solved simultaneously. An excessive number of parameters results in solving slowly or potential failure. To solve this problem, we propose a parallelizable color correction approach that decouples the correlation of correction parameters in the optimization equations and optimizes each image separately. First, we introduce auxiliary variables that replace values related to other images in the cost function. Second, we construct optimization equations for each image and parallelly solve the correction parameters. Finally, we correct the input images through a weighted correction model to better eliminate correction artifacts. Our approach iteratively optimizes auxiliary variables and correction parameters until the correction results converge. The experimental results on several challenging datasets show that our approach significantly improves execution efficiency and obtains the global optimal solution using the flexible correction model.
Hongche Yin, Pengwei Zhou, Guozheng Xu, Gaoming He, Li Li 0047, Jian Yao 0002
IEEE Geosci. Remote. Sens. Lett.6
2025 A Transformer-Based Roof Plane Segmentation Approach for Airborne LiDAR Point Clouds
abstract
In the fields of photogrammetry and computer vision, three-dimensional (3D) urban building model reconstruction from airborne Light Detection and Ranging (LiDAR) point clouds has attracted significant attention in recent years. Accurately and automatically extracting local geometric structures, such as planar patches, from 3D point cloud data directly determines the quality of subsequent 3D model reconstruction. Considering that the roof is a crucial component of a real building, roof plane segmentation is a critical procedure in building 3D reconstruction. In this paper, a novel dual-branch transformer-based network is designed to accurately segment roof planes from airborne LiDAR point clouds. We first use PointNet++ followed with a transformer encoder to extract point-wise feature embeddings. Then, in the first branch, a transformer decoder module is applied to directly learn the instance centers of planar patches by giving a set of learned queries. Because the transformer can effectively model the relations of the queries and the global context information, the instance center positions of all planes included in the input point clouds can be accurately predicted. In this way, the number and center positions of roof planes are known before performing roof plane segmentation. In the second branch, we predict the offsets for each point using its point-wise feature to shift it towards the corresponding instance center. After that, the plane parameters for each plane instance can be estimated using the shifted points around the predicted centers, and the rest of points are assigned to its nearest plane to generate the final roof planes. The experimental results illustrate that our approach can successfully address the plane segmentation challenge for diverse building roof structures while achieving performance superior to the current state-of-the-art techniques. We will make the source code of our approach publicly available at https://github.com/Li-Li-Whu/PlaneTransformer.
Siyuan You, Guozheng Xu, Pengwei Zhou, Yubing Wei, Jian Yao 0002, Li Li 0047
IEEE Trans. Geosci. Remote. Sens.5
2024 iMVS: Integrating multi-view information on multiple scales for 3D object recognition
Jiaqin Jiang, JingMin Tu, Li Li 0047, Jian Yao 0002
J. Vis. Commun. Image Represent.6
2024 A multi-view references image super-resolution framework for generating the large-FOV and high-resolution image
Jiaqin Jiang, Li Li 0047, Lunhao Duan, Jian Yao 0002
J. Vis. Commun. Image Represent.5
2024 Workpiece classification based on transfer component analysis
Liyong Qiao, Shuang Zhang 0009, Chungang Liu, Huilong Jin, Jian Yao 0002, Lingru Cao, Yujia Ji
Wirel. Networks6
2023 Multi-view convolutional vision transformer for 3D object recognition
Li Li 0047, Junqin Lin, Jian Yao 0002, JingMin Tu
J. Vis. Commun. Image Represent.5
2022 JPV-Net: Joint Point-Voxel Representations for Accurate 3D Object Detection
abstract
Voxel and point representations are widely applied in recent 3D object detection tasks from LiDAR point clouds. Voxel representations contribute to efficiently and rapidly locating objects, whereas point representations are capable of describing intra-object spatial relationship for detection refinement. In this work, we aim to exploit the strengths of both two representations, and present a novel two-stage detector, named Joint Point-Voxel Network (JPV-Net). Specifically, our framework is equipped with a Dual Encoders-Fusion Decoder, which consists of the dual encoders to extract voxel features of sketchy 3D scenes and point features rich in geometric context, respectively, and the Feature Propagation Fusion (FP-Fusion) decoder to attentively fuse them from coarse to fine. By making use of the advantages of these features, the refinement network can effectively eliminate false detection and provide better accuracy. Besides, to further develop the perception characteristics of voxel CNN and point backbone, we design two novel intersection-over-union (IoU) estimation modules for proposal generation and refinement, both of which can alleviate the misalignment between the localization and the classification confidence. Extensive experiments on the KITTI dataset and ONCE dataset demonstrate that our proposed JPV-Net outperforms other state-of-the-art methods with remarkable margins.
Nan Song, Tianyuan Jiang, Jian Yao 0002
AAAI3
2022 DeepMLE: A Robust Deep Maximum Likelihood Estimator for Two-view Structure from Motion
abstract
Two-view structure from motion (SfM) is the cornerstone of 3D reconstruction and visual SLAM (vSLAM). Many existing end-to-end learning-based methods usually formulate it as a brute regression problem. However, the inadequate utilization of traditional geometry model makes the model not robust in unseen environments. To improve the generalization capability and robustness of end-to-end two-view SfM network, we formulate the two-view SfM problem as a maximum likelihood estimation (MLE) and solve it with the proposed framework, denoted as DeepMLE. First, we propose to take the deep multi-scale correlation maps to depict the visual similarities of 2D image matches decided by ego-motion. In addition, in order to increase the robustness of our framework, we formulate the likelihood function of the correlations of 2D image matches as a Gaussian and Uniform mixture distribution which takes the uncertainty caused by illumination changes, image noise and moving objects into account. Meanwhile, an uncertainty prediction module is presented to predict the pixel-wise distribution parameters. Finally, we iteratively refine the depth and relative camera pose using the gradient-like information to maximize the likelihood function of the correlations. Extensive experimental results on several datasets prove that our method significantly outperforms the state-of-the-art end-to-end two-view SfM approaches in accuracy and generalization capability.
Yuxi Xiao, Li Li 0047, Jian Yao 0002
IROS4
2021 VIC-Net: Voxelization Information Compensation Network for Point Cloud 3D Object Detection
abstract
Voxel-based methods have been widely used in point cloud 3D object detection. These methods usually transform points into voxels while suffering from information loss during point cloud voxelization. To address this problem, we propose a novel one-stage Voxelization Information Compensation Network (VIC-Net), which has the ability of loss-free feature extraction. The whole framework consists of a point branch for geometry detail extraction and a voxel branch for efficient proposals generation. Firstly, PointNet++ is adopted to efficiently encode geometry structure features from the raw point clouds. Then based on the encoded point features, two Point2Voxel (P2V) feature fusion modules are proposed to fuse point features with a voxel backbone, including Local P2V and Multi-Scale P2V. The P2V modules respectively integrate local detail features and multi-scale semantic contexts into a sparse voxel backbone. Thirdly, an auxiliary reconstruction loss is employed on the point branch to explicitly guide the point backbone to be aware of real geometry structures. In addition, we extend VIC-Net to a two-stage approach, namely VIC-RCNN, which further utilizes the fine geometry features to refine object locations. Experiments on the KITTI dataset demonstrate that our proposed VIC-Net outperforms other onestage methods and our two-stage method VIC-RCNN achieves new state-of-the-art performance.
Tianyuan Jiang, Nan Song, Ruihao Yin, Ye Gong, Jian Yao 0002
ICRA6
2021 Grid Model-Based Global Color Correction for Multiple Image Mosaicking
abstract
Color consistency optimization for multiple images is a challenging problem in image mosaicking. To facilitate the global color optimization, existing approaches mainly use less flexible models, e.g., linear or gamma function, to eliminate the color differences between multiple images. However, these models often struggle to eliminate the color differences that existed in the local areas and preserve the image gradient information. To solve this problem, we creatively propose a novel color-correction model, which comprised a series of local grid linear models. This model is simple, but it is flexible enough to approximate a variety of complicated local color variations. To obtain the optimal model parameters for each image globally, a specific cost function that considers both color consistency and gradient preservation is designed and solved. The aim of our approach is to generate a composite image with visually consistent color. The original color information may be destroyed. Thus, this approach is unsuitable for the quantitative remote sensing applications. The experimental results on several challenging data sets show that the proposed approach outperforms state-of-the-art approaches in both visual quality and quantitative metrics.
Li Li 0047, Yunmeng Li, Menghan Xia, Yinxuan Li, Jian Yao 0002, Bin Wang 0100
IEEE Geosci. Remote. Sens. Lett.5
2021 Extraction of Street Pole-Like Objects Based on Plane Filtering From Mobile LiDAR Data
abstract
Pole-like objects provide important street infra- structure for road inventory and road mapping. In this article, we proposed a novel pole-like object extraction algorithm based on plane filtering from mobile Light Detection and Ranging (LiDAR) data. The proposed approach is composed of two parts. In the first part, a novel octree-based split scheme was proposed to fit initial planes from off-ground points. The results of the plane fitting contribute to the extraction of pole-like objects. In the second part, we proposed a novel method of pole-like object extraction by plane filtering based on local geometric feature restriction and isolation detection. The proposed approach is a new solution for detecting pole-like objects from mobile LiDAR data. The innovation in this article is that we assumed that each of the pole-like objects can be represented by a plane. Thus, the essence of extracting pole-like objects will be converted to plane selecting problem. The proposed method has been tested on three data sets captured from different scenes. The average completeness, correctness, and quality of our approach can reach up to 87.66%, 88.81%, and 79.03%, which is superior to state-of-the-art approaches. The experimental results indicate that our approach can extract pole-like objects robustly and efficiently.
JingMin Tu, Jian Yao 0002, Li Li 0047, Binbin Xiang
IEEE Trans. Geosci. Remote. Sens.2
2020 Single Image Super-Resolution Using Depth Map as Constraint
abstract
Single image super-resolution based on the deep neural network has achieved great performance recently, but generating photo-realistic images remains a challenging problem. To tackle this issue, we propose a method that uses depth maps as a constraint to get better visual quality. Specifically, we propose a self-adaptive feature transform (AFT) layer, which can perform affine transformation on the feature map based on the depth map to constrain the plausible solution space of the SR image. Furthermore, we propose a hierarchical residual multi-scale fusion block to improve the representational ability of the network. Experimental results on benchmark datasets demonstrate that our method is superior to other perceptually-oriented SISR methods in terms of visual quality and also achieves state-of-the-art performance on quantitative metrics.
Jiaqin Jiang, Jian Yao 0002
ICIP3
2020 Deformable Spatial Propagation Networks For Depth Completion
abstract
Depth completion has attracted extensive attention recently due to the development of autonomous driving, which aims to recover dense depth map from sparse depth measurements. Convolutional spatial propagation network (CSPN) is one of the state-of-the-art methods in this task, which adopt a linear propagation model to refine coarse depth maps with local context. However, the propagation of each pixel occurs in a fixed receptive field. This may not be the optimal for refinement since different pixel needs different local context. To tackle this issue, in this paper, we propose a deformable spatial propagation network (DSPN) to adaptively generates different receptive field and affinity matrix for each pixel. It allows the network obtain information with much fewer but more relevant pixels for propagation. Experimental results on KITTI depth completion benchmark demonstrate that our proposed method achieves the state-of-the-art performance.
Hongche Yin, Jian Yao 0002
ICIP3
2020 Joint learning of image detail and transmission map for single image dehazing
Shengdong Zhang, Fazhi He, Wenqi Ren, Jian Yao 0002
Vis. Comput.4
2019 GHRNet: Guided Hierarchical Refinement Network for Stereo Matching
abstract
Recently, deep convolutional neural networks (CNNs) have been well developed in the task of stereo matching. However, most existing CNNs-based methods cannot accurately estimate details especially for some tiny objects and shape boundaries. To solve this problem, in this paper, we propose a guided hierarchical refinement network (GHRNet) with a specially designed guidance block for details refinement. Specifically, an initial low-resolution disparity map is first estimated. It is then fed into a hierarchical architecture to recover disparity details from coarse to fine. The foremost contribution of our proposed method is an embedded guidance block, which fully makes use of clues from input images to refine predicated disparity map. We evaluate our network on several popular datasets. The experimental results demonstrate that the proposed network can significantly improve disparity details, and the overall result also achieves the state-of-the-art performance.
Kai Chen 0024, Jian Yao 0002
ICIP3
2019 Data-Adaptive Packing Method for Compression of Dynamic Point Cloud Sequences
abstract
This paper proposes a data-adaptive packing method for compression of dynamic point cloud sequences. Due to the increasing popularity in emerging immersive applications, the interest in representing the virtual world with dynamic point cloud sequences has never been higher. However, it remains a challenging work to compress these sequences with high efficiency. The proposed method divides the input sequences into some groups adaptively, in which patches in all frames are packed in a consistent way. Thus, the occupancy maps as well as the corresponding depth and texture images, are highly consistent both in temporal and spatial, which benefits the subsequent 2D video encoding greatly. Experimental results show that the proposed method achieves significant improvement in coding efficiency over the reference method of MPEG Point Cloud Compression (PCC).
Jian Yao 0002, JingMin Tu
ICME2
2019 DeepCrack: A deep hierarchical feature learning architecture for crack segmentation
Jian Yao 0002, Xiaohu Lu, Renping Xie, Li Li 0047
Neurocomputing2
2019 RoadNet: Learning to Comprehensively Analyze Road Networks in Complex Urban Scenes From High-Resolution Remotely Sensed Images
abstract
It is a classical task to automatically extract road networks from very high-resolution (VHR) images in remote sensing. This paper presents a novel method for extracting road networks from VHR remotely sensed images in complex urban scenes. Inspired by image segmentation, edge detection, and object skeleton extraction, we develop a multitask convolutional neural network (CNN), called RoadNet, to simultaneously predict road surfaces, edges, and centerlines, which is the first work in such field. The RoadNet solves seven important issues in this vision problem: 1) automatically learning multiscale and multilevel features [gained by the deeply supervised nets (DSN) providing integrated direct supervision] to cope with the roads in various scenes and scales; 2) holistically training the mentioned tasks in a cascaded end-to-end CNN model; 3) correlating the predictions of road surfaces, edges, and centerlines in a network model to improve the multitask prediction; 4) designing elaborate architecture and loss function, by which the well-trained model produces approximately single-pixel width road edges/centerlines without nonmaximum suppression postprocessing; 5) cropping and bilinear blending to deal with the large VHR images with finite-computing resources; 6) introducing rough and simple user interaction to obtain desired predictions in the challenging regions; and 7) establishing a benchmark data set which consists of a series of VHR remote sensing images with pixelwise annotation. Different from the previous works, we pay more attention to the challenging situations, in which there are lots of shadows and occlusions along the road regions. Experimental results on two benchmark data sets show the superiority of our proposed approaches.
Jian Yao 0002, Xiaohu Lu, Menghan Xia, Xingbo Wang 0001
IEEE Trans. Geosci. Remote. Sens.2
2019 A Novel Octree-Based 3-D Fully Convolutional Neural Network for Point Cloud Classification in Road Environment
abstract
The automatic classification of 3-D point clouds is publicly known as a challenging task in a complex road environment. Specifically, each point is automatically classified into a unique category label, and then, the labels are used as clues for semantic analysis and scene recognition. Instead of heuristically extracting handcrafted features in traditional methods to classify all points, we put forward an end-to-end octree-based fully convolutional network (FCN) to classify 3-D point clouds in an urban road environment. There are four contributions in this paper. The first is that the integration and comprehensive uses of OctNet and FCN greatly decrease the computing time and memory demands compared with a dense 3-D convolutional neural network (CNN). The second is that the octree-based network is strengthened by means of modifying the cross-entropy loss function to solve the problems of an unbalanced category distribution. The third is that an Inception-ResNet block is united with our network, which enables our 3-D CNN to effectively learn how to classify scenes containing objects at multiple scales and improve classification accuracy. The last is that an open source data set (HuangshiRoad data set) with ten different classes is introduced for 3-D point cloud classification. Three representative data sets [Semantic3D, WHU_MLS (blocks I and II), and HuangshiRoad] with different covered areas and numbers of points and classes are selected to evaluate our proposed method. The experimental results show that the overall classification accuracy is appreciable, with 89.4% for Semantic3D, 82.9% for WHU_MLS block I, 91.4% for WHU_MLS block II, and 94% for HuangshiRoad. Our deep learning approach can efficiently classify 3-D dense point clouds in an urban road environment measured by a mobile laser scanning (MLS) system or static LiDAR.
Binbin Xiang, JingMin Tu, Jian Yao 0002, Li Li 0047
IEEE Trans. Geosci. Remote. Sens.3
2018 Regularizing RNNs for Caption Generation by Reconstructing the Past With the Present
abstract
Recently, caption generation with an encoder-decoder framework has been extensively studied and applied in different domains, such as image captioning, code captioning, and so on. In this paper, we propose a novel architecture, namely Auto-Reconstructor Network (ARNet), which, coupling with the conventional encoder-decoder framework, works in an end-to-end fashion to generate captions. ARNet aims at reconstructing the previous hidden state with the present one, besides behaving as the input-dependent transition operator. Therefore, ARNet encourages the current hidden state to embed more information from the previous one, which can help regularize the transition dynamics of recurrent neural networks (RNNs). Extensive experimental results show that our proposed ARNet boosts the performance over the existing encoder-decoder models on both image captioning and source code captioning tasks. Additionally, ARNet remarkably reduces the discrepancy between training and inference processes for caption generation. Furthermore, the performance on permuted sequential MNIST demonstrates that ARNet can effectively regularize RNN, especially on modeling long-term dependencies. Our code is available at: https://github.com/chenxinpeng/ARNet.
Lin Ma 0002, Jian Yao 0002, Wei Liu 0005
CVPR4
2018 Towards Less Generic Responses in Neural Conversation Models: A Statistical Re-weighting Method
abstract
Sequence-to-sequence neural generation models have achieved promising performance on short text conversation tasks.However, they tend to generate generic/dull responses, leading to unsatisfying dialogue experience.We observe that in conversation tasks, each query could have multiple responses, which forms a 1-to-n or m-to-n relationship in the view of the total corpus.The objective function used in standard sequence-to-sequence models will be dominated by loss terms with generic patterns.Inspired by this observation, we introduce a statistical re-weighting method that assigns different weights for the multiple responses of the same query, and trains the standard neural generation model with the weights.Experimental results on a large Chinese dialogue corpus show that our method improves the acceptance rate of generated responses compared with several baseline models and significantly reduces the number of generated generic responses.
Wei Bi, Xiaojiang Liu, Jian Yao 0002, Shuming Shi 0001
EMNLP5
2018 Multiple Combined Constraints for Image Stitching
abstract
Several approaches to image stitching use different constraints to estimate the motion model between image pairs. These constraints can be roughly divided into two categories: geometric constraints and photometric constraints. In this paper, geometric and photometric constraints are combined to improve the alignment quality, which is based on the observation that these two kinds of constraints are complementary. On the one hand, geometric constraints (e.g., point and line correspondences) are usually spatially biased and are insufficient in some extreme scenes, while photometric constraints are always evenly and densely distributed. On the other hand, photometric constraints are sensitive to displacements and are not suitable for images with large parallaxes, while geometric constraints are usually imposed by feature matching and are more robust to handle parallaxes. The proposed method therefore combines them together in an efficient mesh-based image warping framework. It achieves better alignment quality than methods only with geometric constraints, and can handle larger parallax than photometric-constraint-based method. Experimental results on various images illustrate that the proposed method outperforms representative state-of-the-art image stitching methods reported in the literature.
Kai Chen 0024, JingMin Tu, Binbin Xiang, Li Li 0047, Jian Yao 0002
ICIP5
2018 Feed-Net: Fully End-to-End Dehazing
abstract
This paper proposes an image dehazing model built with a fully convolutional neural network (CNN), called Fully End-to-End Dehazing Network (FEED-Net). In contrast to estimate the transmission map and the atmospheric light separately as most previous deep learning methods, FEED-Net recovers the hazy-free image directly from a hazy image via a light-weight CNN. In addition, we introduce contextual information into dehazing via dilated convolution and use dense skip connection for feature fusion, which makes end-to-end dehazing possible. Experimental results show our method outperforms the state-of-the-art algorithms on both synthetic dataset and real-world images.
Shengdong Zhang, Wenqi Ren, Jian Yao 0002
ICME3
2018 Video Stitching with Extended-MeshFlow
abstract
In this paper, we present a method that stitches multiple videos captured with a fixed camera rig. We propose an Extended-MeshFlow motion model for video stitching. Firstly, uniform features are detected and matched at the overlapping region, from which the Exended-MeshFlow model is estimated. The model then warps the adjacent views to the common central view to eliminate the spatial misalignment. The motions located on the feature positions are interpolated to the mesh vertexes by Multilevel B-Spline Approximation(MBA). Collecting the motions on the vertexes form the vertex profiles, which are smoothed for temporal consistency. During the smoothing, only previous frames are required, thus the proposed method can stitch videos in an online mode. Experimental results on various of videos demonstrate that the proposed method can produce comparable stitching results in aspects of spatial alignment and temporal coherence.
Kai Chen 0024, Jian Yao 0002, Binbin Xiang, JingMin Tu
ICPR2
2018 A Monocular SLAM System Leveraging Structural Regularity in Manhattan World
abstract
The structural features in Manhattan world encode useful geometric information of parallelism, orthogonality and/or coplanarity in the scene. By fully exploiting these structural features, we propose a novel monocular SLAM system which provides accurate estimation of camera poses and 3D map. The foremost contribution of the proposed system is a structural feature-based optimization module which contains three novel optimization strategies. First, a rotation optimization strategy using the parallelism and orthogonality of 3D lines is presented. We propose a global binding method to compute an accurate estimation of the absolute rotation of the camera. Then we propose an approach for calculating the relative rotation to further refine the absolute rotation. Second, a translation optimization strategy leveraging coplanarity is proposed. Coplanar features are effectively identified, and we leverage them by a unified model handling both points and lines to calculate the relative translation, and then the optimal absolute translation. Third, a 3D line optimization strategy utilizing parallelism, orthogonality and coplanarity simultaneously is proposed to obtain an accurate 3D map consisting of structural line segments with low computational complexity. Experiments in man-made environments have demonstrated that the proposed system outperforms existing state-of-the-art monocular SLAM systems in terms of accuracy and robustness.
Haoang Li, Jian Yao 0002, Jean-Charles Bazin, Xiaohu Lu, Yazhou Xing, Kang Liu 0003
ICRA2
2018 Robust Camera Pose Estimation via Consensus on Ray Bundle and Vector Field
abstract
Estimating the camera pose requires point correspondences. However, in practice, correspondences are inevitably corrupted by outliers, which affects the pose estimation. We propose a general and accurate outlier removal strategy for robust camera pose estimation. The proposed strategy can detect outliers by leveraging the fact that only inliers comply with two effective consensuses, i.e., 3D ray bundle consensus and 2D vector field consensus. Our strategy has a nested structure. First, the outer module utilizes the 3D ray bundle consensus. We define the likelihood based on the probabilistic mixture model and maximize it by the expectation-maximization (EM) algorithm. The inlier probability of each correspondence and the camera pose are determined alternately. Second, the inner module exploits the 2D vector field consensus to refine the probabilities obtained by the outer module. The refinement based on the Bayesian rule facilitates the convergence of the outer module and improves the accuracy of the entire framework. Our strategy can be integrated into various existing camera pose estimation methods which are originally vulnerable to outliers. Experiments on both synthesized data and real images have shown that our approach outperforms state-of-the-art outlier rejection methods in terms of accuracy and robustness.
Haoang Li, Ji Zhao 0001, Jean-Charles Bazin, Jian Yao 0002
IROS6
2018 Edge-Enhanced Optimal Seamline Detection for Orthoimage Mosaicking
abstract
In this letter, we propose a novel algorithm to detect the optimal seamline by using the edge-enhanced energy function for orthoimage mosaicking. First, we generate the energy cost map by using traditional intensity and gradient difference. Then, to ensure that the detected optimal seamlines avoid crossing the obvious objects, the energy map is enhanced in the regions of valid edge segments around the obvious objects. The edge segments are separately detected from input images, and the invalid edge segments are removed by using texture complexity index, image difference, and consistency constraint. Finally, we detect the optimal seamline from this edge-enhanced energy cost map via graph cuts. Experimental results on several groups of challenging orthoimages show that the proposed algorithm is capable of creating high-quality seamlines for orthoimage mosaicking and outperforms state-of-the-art algorithms and the commercial software based on the visual comparison and statistical evaluation.
Li Li 0047, Jian Yao 0002, Renping Xie
IEEE Geosci. Remote. Sens. Lett.2
2018 Subtracted Histogram: Utilizing Mutual Relation Between Features for Thresholding
abstract
In this paper, we propose a thresholding method utilizing the mutual relation between the two used features, which should be both robust and of low correlation. The mutual relation is exploited through a histogram subtraction transformation, which tries to reduce only the background part of each of the histograms for the features so that we can easily differentiate between the background part and the target part. A speedup scheme that requires prior knowledge to exclude the possible local minimum(s) is also presented. To sufficiently validate the effectiveness of our histogram subtraction transformation, the two developed thresholding methods are applied to shadow detection and vegetation detection and compared with some state-of-the-art methods in both fields. The comparative results on multiple data sets indicate that with the thresholds automatically provided by the proposed thresholding methods, even the simple binarization methods can obtain good detection results.
Xiaojian Chen, Jian Yao 0002
IEEE Trans. Geosci. Remote. Sens.4
2017 Combining points and lines for camera pose estimation and optimization in monocular visual odometry
abstract
In this paper, we propose a unified model for camera pose estimation and a novel strategy for pose optimization by combining points and lines in monocular visual odometry. Our proposed unified model treats point and line features equivalently, which is applicable for all the minimal cases requiring the minimum number 3 of point or/and line features and can be easily extended for various circumstances with more additional observations. The core idea is to directly retrieve all stationary points of a cost function which is minimized by the first-order optimality condition without initialization or iteration. The estimated pose is reliable due to robust geometric constraints and the reliable algebraic solver. To refine the camera pose, we propose a novel optimization strategy to minimize the unconstrained Sampson error by taking specific uncertainty for each feature into account to penalize noise more reasonably. Moreover, it is simpler than the conventional bundle adjustment by avoiding the high-dimensional parameter searching. Experimental results on simulated data and real images have sufficiently demonstrated the superiority of our proposed camera pose estimation and optimization method by comparing with state-of-the-art monocular algorithms.
Haoang Li, Jian Yao 0002, Xiaohu Lu
IROS2
2017 Single Image Dehazing via Image Generating
Shengdong Zhang, Jian Yao 0002, Edel B. García Reyes
PSIVT2
2017 2-Line Exhaustive Searching for Real-Time Vanishing Point Estimation in Manhattan World
abstract
This paper presents a very simple and efficient algorithm to estimate 1, 2 or 3 orthogonal vanishing point(s) on a calibrated image in Manhattan world. Unlike the traditional methods which apply 1, 3, 4, or 6 line(s) to generate vanishing point hypotheses, we propose to use 2 lines to get the first vanishing point v1, then uniformly take sample of the second vanishing point v2 on the great circle of v1 on the equivalent sphere, and finally calculate the third vanishing point v3 by the cross-product of v1 and v2. There are three advantages of the proposed method over traditional multi-line method. First, the 2-line model is much more robust and reliable than the multi-line method, which can be applied in the scene with 1, 2 or 3 orthogonal vanishing point(s). Second, the probability of the 2-line model being formed of inner line segments can be calculated given the outlier ratio, which means that the number of iterations can be determined, and thus the estimation of vanishing points can be performed in a very simple exhaustive way instead of the traditional RANSAC method. Third, the real-time performance is achieved by building a polar grid for the line intersection points, which functions as a lookup table for the validation of vanishing point hypotheses. Our algorithm has been validated successfully in the YUD dataset and sets of challenging real images.
Xiaohu Lu, Jian Yao 0002, Haoang Li
WACV2
2017 Optimal seamline detection in dynamic scenes via graph cuts for image mosaicking
Li Li 0047, Jian Yao 0002, Haoang Li, Menghan Xia, Wei Zhang 0021
Mach. Vis. Appl.2
2017 Globally consistent alignment for planar mosaicking via topology analysis
Menghan Xia, Jian Yao 0002, Renping Xie, Li Li 0047, Wei Zhang 0021
Pattern Recognit.2
2016 Edge chain detection by applying Helmholtz principle on gradient magnitude map
abstract
In this paper, we present an efficient edge chain detection algorithm by applying the Helmholtz principle on the gradient magnitude map of an image. An edge chain validation method is proposed which uses the “relative number of false alarms” (RNFA) instead of the traditional “number of false alarms” (NFA). The edge chains are detected first and then validated according to their RNFA values. In this way, edge chains that are weak in gradients but meaningful in vision can be detected. To evaluate the proposed edge chain detector in quantity, an edge chain detection benchmark which consists of 25 labeled images in different scenes was built. The proposed edge chain detector was tested in this benchmark, and the experimental results sufficiently demonstrate that the proposed edge chain detector outperforms the state-of-the-art methods.
Xiaohu Lu, Jian Yao 0002, Li Li 0047, Wei Zhang 0021
ICPR2
2016 Line segment matching: A benchmark
abstract
As the vital procedure for exploiting line segments extracted from images for solving computer vision problems, Line Segment Matching (LSM) has received growing attentions from researcher in recent years, and a considerable number of methods have been proposed. However, no one has attempted to solve two major problems in this area. The first is how to evaluate different methods in an unbiased way. All proposed methods were evaluated using images and line segment detectors selected by the authors themselves, making the conclusions based on the somewhat biased experiments less convincing. The second problem is that there is no reliably automatic way to access the correctness of obtained line segment matches, which can often be up to hundreds in quantity. Checking them one by one by visual inspection is the only reliable, but very tedious and error-prone way. In this paper, we target to solve the two problems. We introduce a benchmark which provides the ground truth matches among 30 pairs of line segment sets extracted from 15 representative image pairs using two state-of-the-art line segment detectors. With the benchmark, we evaluated some of the existing LSM methods.
Kai Li 0015, Jian Yao 0002, Mengsheng Lu, Yuan Heng, Yinxuan Li
WACV2
2016 Joint point and line segment matching on wide-baseline stereo images
abstract
This paper presents an method that matches points and line segments jointly on wide-baseline stereo images. In both two images to be matched, line segments are extracted and those spatially adjacent ones are intersected to generate V-junctions. To match V-junctions from the two images, we extract for each of them an affine and scale invariant local region and describe it with SIFT. The putative V-junction matches obtained from evaluating their description vectors are refined subsequently by the epipolar line constraint and topological distribution constraint among neighbor V-junctions. Since once a pair of V-junctions are matched, the two pairs of line segments forming them are matched accordingly. A part of line segments from the two images are therefore matched along with V-junction matches. To get more line segment matches, we further match those left unmatched line segments by the local homographies estimated from their adjacent V-junction matches. Experiments verify the robustness of the proposed method and its superiority to both some famous point and line segment matching methods on wide-baseline stereo images. In addition, we also show the proposed method can make it easier for 3D line segment reconstruction.
Kai Li 0015, Jian Yao 0002, Menghan Xia, Li Li 0047
WACV2
2016 Hierarchical line matching based on Line-Junction-Line structure descriptor and local homography estimation
Kai Li 0015, Jian Yao 0002, Xiaohu Lu, Li Li 0047
Neurocomputing2
2016 Text-Attentional Convolutional Neural Network for Scene Text Detection
abstract
Recent deep learning models have demonstrated strong capabilities for classifying text and non-text components in natural images. They extract a high-level feature globally computed from a whole image component (patch), where the cluttered background information may dominate true text features in the deep representation. This leads to less discriminative power and poorer robustness. In this paper, we present a new system for scene text detection by proposing a novel text-attentional convolutional neural network (Text-CNN) that particularly focuses on extracting text-related regions and features from the image components. We develop a new learning mechanism to train the Text-CNN with multi-level and rich supervised information, including text region mask, character label, and binary text/non-text information. The rich supervision information enables the Text-CNN with a strong capability for discriminating ambiguous texts, and also increases its robustness against complicated background components. The training process is formulated as a multi-task learning problem, where low-level supervised information greatly facilitates the main task of text/non-text classification. In addition, a powerful low-level detector called contrast-enhancement maximally stable extremal regions (MSERs) is developed, which extends the widely used MSERs by enhancing intensity contrast between text patterns and background. This allows it to detect highly challenging text patterns, resulting in a higher recall. Our approach achieved promising results on the ICDAR 2013 data set, with an F-measure of 0.82, substantially improving the state-of-the-art results.
Tong He 0001, Yu Qiao 0001, Jian Yao 0002
IEEE Trans. Image Process.4
2015 Line-based Multi-Label Energy Optimization for fisheye image rectification and calibration
abstract
Fisheye image rectification and estimation of intrinsic parameters for real scenes have been addressed in the literature by using line information on the distorted images. In this paper, we propose an easily implemented fisheye image rectification algorithm with line constrains in the undistorted perspective image plane. A novel Multi-Label Energy Optimization (MLEO) method is adopted to merge short circular arcs sharing the same or the approximately same circular parameters and select long circular arcs for camera rectification. Further we propose an efficient method to estimate intrinsic parameters of the fisheye camera by automatically selecting three properly arranged long circular arcs from previously obtained circular arcs in the calibration procedure. Experimental results on a number of real images and simulated data show that the proposed method can achieve good results and outperforms the existing approaches and the commercial software in most cases.
Mi Zhang 0004, Jian Yao 0002, Menghan Xia, Kai Li 0015
CVPR2
2015 CannyLines: A parameter-free line segment detector
abstract
In this paper, we present a robust line segment detection algorithm to efficiently detect the line segments from an input image. Firstly a parameter-free Canny operator, named as CannyPF, is proposed to robustly extract the edge map from an input image by adaptively setting the low and high thresholds for the traditional Canny operator. Secondly, both efficient edge linking and splitting techniques are proposed to collect collinear point clusters directly from the edge map, which are used to fit the initial line segments based on the least-square fitting method. Thirdly, longer and more complete line segments are produced via efficient extending and merging. Finally, all the detected line segments are validated due to the Helmholtz principle [1, 2] in which both the gradient orientation and magnitude information are considered. Experimental results on a set of representative images illustrate that our proposed line segment detector, named as CannyLines, can extract more meaningful line segments than two popularly used line segment detectors, LSD [3] and ED-Lines [4], especially on the man-made scenes.
Xiaohu Lu, Jian Yao 0002, Kai Li 0015, Li Li 0047
ICIP2
2014 Scale selection based on Moran's I for segmentation of high resolution remotely sensed images
abstract
Image segmentation is a prerequisite for object-based image analysis (OBIA). However, selecting an optimal segmentation scale is often time consuming and needs trial-and-error. This paper presents an unsupervised scale selection method based on the rate of change of a spatial autocorrelation indicator - the global Moran's I for segmentation of high resolution remotely sensed images. It was compared with other two scale selection methods and its effectiveness is validated through both visual analysis and by referencing to multiple manual segmentations. Experimental results on our own data and statistical data from an external reference showed that the optimal scale could be easily selected through the proposed method.
Weihong Cui, Jian Yao 0002
IGARSS4