VLDB 2026 Research / reviewers in the wild / expert
Li Li 0047
dblp:53/2189-47
· DBLP profile ↗
24ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0003-3138-6033ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-view multilayer perceptron attention for 3D object recognition
Yeduozi Yao, Jiaqin Jiang, Li Li 0047, Jian Yao 0002 |
J. Vis. Commun. Image Represent. | 4 |
| 2026 | SMDGS: Scale-Aligned Monocular Depth-Guided 3D Gaussian Splatting for Rendering and Surface Reconstructionabstract3D Gaussian Splatting (3DGS) has been explored for surface reconstruction, however unstructured and discontinuous Gaussian point clouds lead to uneven surface reconstruction accuracy as well as frequent loss of Novel View Synthesis (NVS) quality. To address this problem, we propose a scale-aligned monocular depth-guided 3DGS, a promising novel framework that combines geometric prior regularization and consistency supervision to achieve high-quality rendering and surface reconstruction. Specifically, monocular depth, estimated by some recently proposed monocular depth estimation models, contain implicitly abundant valuable geometric cues, but scale ambiguity limits its application. Therefore we first propose a $K$K-Nearest Neighbor (KNN)-based depth alignment framework that utilizes the full-domain gradient at monocular depth map to align to the sparse point cloud obtained during the Structure from Motion (SfM), which is employed for regularization to enhance geometric representation. Then a pseudo-mesh-based multi-view consistency module is introduced to fine-tune and guide the model to recover the accurate surface. Finally, a pixel-level isotropic gradient aware method guides the appropriate growth of the Gaussians to further improve the surface and rendering quality. Experiments on dozens of indoor, outdoor, and object-centered/non-object-centered datasets demonstrate that our method achieves accurate surface reconstruction with excellent NVS performance. Xiaosong Wei, Pengwei Zhou, Annan Zhou, Li Li 0047, Jian Yao 0002 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2026 | SGGS: Semantic-Guided 3D Gaussian Splatting With Adaptive Renderingabstract3D Gaussian Splatting (3DGS) has shown great promise in a variety of applications due to its exceptional real-time rendering quality and explicit representation, leading to numerous improvements across various fields. However, existing methods lack consideration of main objects and important structural information in their overall optimization strategies. This results in blurring of main objects in adaptive rendering and the loss of high-frequency details on targets that are insufficiently captured. In this work, we introduce a semantic-guided 3DGS method with adaptive rendering, which optimizes important structures through the guidance of boundary Gaussians, while leveraging semantic features to enhance the rendering of main objects in adaptive rendering. Experiments show that the proposed semantic-guided method can enhance important structures and high-frequency information in corner regions without significantly increasing the total number of Gaussians. This method also improves the separability between objects. At the same time, a semantic-guided Level-of-Detail (LoD) rendering approach enables the rapid display of main targets and the rendering of a complete scene. The semantic-guided methodology we have presented exhibits compatibility with a range of existing techniques. Annan Zhou, Li Li 0047, Jian Yao 0002 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | JPG-SLAM: Joint Point-Gaussian Splatting Representation for Dense Dynamic SLAMabstractThis paper presents a simultaneous localization and mapping (SLAM) system to provide accurate pose estimation and dynamic scene reconstruction. Our approach proposes a Joint Point-Gaussian Splatting representation, which fully integrates the robustness of isotropic feature points in pose estimation and the flexibility of anisotropic 3D Gaussians in scene representation. This system does not need to suppress the anisotropic representation of Gaussian elements, which enables the mapping module to achieve finer scene representation with lower memory consumption. Additionally, in order to enhance the adaptability of the system in dynamic environments, we introduced a dynamic region recognition module and utilized 3D Gaussian Splatting and 4D Gaussian Splatting representations to represent static and dynamic regions respectively. Furthermore, we developed a local map management strategy for Gaussian Splatting mapping, effectively reducing the memory and computational resource usage in the mapping process. Experiments on public datasets demonstrate that our system achieves state-of-the-art tracking and mapping accuracy compared to existing baselines. Kunrui Huang, Wennan Yang, Pengwei Zhou, Li Li 0047, Jian Yao 0002 |
ICRA | 4 |
| 2025 | A Coarse-to-Fine Boundary Relabeling Approach for Roof Plane SegmentationabstractBuilding roof plane segmentation is important for three-dimensional (3D) building model reconstruction from airborne light detection and ranging (LiDAR) point data. During the roof plane segmentation, challenges such as pseudo planes, over- and under-segmentation often arise, particularly evident in the boundary regions. To improve the accuracy of plane segmentation, various energy optimization-based methods have been proposed to refine the roof planes. However, the existing methods optimize the energy function at the point level, which may lead to getting stuck in local optima. To address these problems, we propose a coarse-to-fine boundary relabeling approach for roof plane segmentation. Starting from an initial plane segmentation result, the proposed method iteratively refines the planes by adjusting the boundaries from the voxel level to the point level. In addition, we also design a new energy function that considers accurateness, smoothness and compactness to guide the optimization. The experimental results constructed on two datasets demonstrate that the proposed method outperforms the existing roof plane segmentation methods, achieving high accuracy and smooth boundary extraction. The source code of the proposed approach will be publicly available at https://github.com/Li-Li-Whu/Coarse2FineRoofPlane. Guozheng Xu, Siyuan You, Li Li 0047, Jian Yao 0002 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | A Parallelizable Global Color Consistency Optimization Algorithm for Multiple ImagesabstractThe global optimization-based color correction approach aims to minimize the color differences of multiple images by optimizing the correction model for each image. The color differences in multisource and multitemporal remote sensing images are difficult to express using a simple correction model with few parameters. When employing a more flexible correction model, the number of correction parameters and optimization equations grows rapidly with the increase in the number and resolution of input images. In addition, the correction parameters of all images are coupled together and need to be solved simultaneously. An excessive number of parameters results in solving slowly or potential failure. To solve this problem, we propose a parallelizable color correction approach that decouples the correlation of correction parameters in the optimization equations and optimizes each image separately. First, we introduce auxiliary variables that replace values related to other images in the cost function. Second, we construct optimization equations for each image and parallelly solve the correction parameters. Finally, we correct the input images through a weighted correction model to better eliminate correction artifacts. Our approach iteratively optimizes auxiliary variables and correction parameters until the correction results converge. The experimental results on several challenging datasets show that our approach significantly improves execution efficiency and obtains the global optimal solution using the flexible correction model. Hongche Yin, Pengwei Zhou, Guozheng Xu, Gaoming He, Li Li 0047, Jian Yao 0002 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2025 | A Transformer-Based Roof Plane Segmentation Approach for Airborne LiDAR Point CloudsabstractIn the fields of photogrammetry and computer vision, three-dimensional (3D) urban building model reconstruction from airborne Light Detection and Ranging (LiDAR) point clouds has attracted significant attention in recent years. Accurately and automatically extracting local geometric structures, such as planar patches, from 3D point cloud data directly determines the quality of subsequent 3D model reconstruction. Considering that the roof is a crucial component of a real building, roof plane segmentation is a critical procedure in building 3D reconstruction. In this paper, a novel dual-branch transformer-based network is designed to accurately segment roof planes from airborne LiDAR point clouds. We first use PointNet++ followed with a transformer encoder to extract point-wise feature embeddings. Then, in the first branch, a transformer decoder module is applied to directly learn the instance centers of planar patches by giving a set of learned queries. Because the transformer can effectively model the relations of the queries and the global context information, the instance center positions of all planes included in the input point clouds can be accurately predicted. In this way, the number and center positions of roof planes are known before performing roof plane segmentation. In the second branch, we predict the offsets for each point using its point-wise feature to shift it towards the corresponding instance center. After that, the plane parameters for each plane instance can be estimated using the shifted points around the predicted centers, and the rest of points are assigned to its nearest plane to generate the final roof planes. The experimental results illustrate that our approach can successfully address the plane segmentation challenge for diverse building roof structures while achieving performance superior to the current state-of-the-art techniques. We will make the source code of our approach publicly available at https://github.com/Li-Li-Whu/PlaneTransformer. Siyuan You, Guozheng Xu, Pengwei Zhou, Yubing Wei, Jian Yao 0002, Li Li 0047 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | iMVS: Integrating multi-view information on multiple scales for 3D object recognition
Jiaqin Jiang, JingMin Tu, Li Li 0047, Jian Yao 0002 |
J. Vis. Commun. Image Represent. | 5 |
| 2024 | A multi-view references image super-resolution framework for generating the large-FOV and high-resolution image
Jiaqin Jiang, Li Li 0047, Lunhao Duan, Jian Yao 0002 |
J. Vis. Commun. Image Represent. | 2 |
| 2023 | Multi-view convolutional vision transformer for 3D object recognition
Li Li 0047, Junqin Lin, Jian Yao 0002, JingMin Tu |
J. Vis. Commun. Image Represent. | 3 |
| 2022 | DeepMLE: A Robust Deep Maximum Likelihood Estimator for Two-view Structure from MotionabstractTwo-view structure from motion (SfM) is the cornerstone of 3D reconstruction and visual SLAM (vSLAM). Many existing end-to-end learning-based methods usually formulate it as a brute regression problem. However, the inadequate utilization of traditional geometry model makes the model not robust in unseen environments. To improve the generalization capability and robustness of end-to-end two-view SfM network, we formulate the two-view SfM problem as a maximum likelihood estimation (MLE) and solve it with the proposed framework, denoted as DeepMLE. First, we propose to take the deep multi-scale correlation maps to depict the visual similarities of 2D image matches decided by ego-motion. In addition, in order to increase the robustness of our framework, we formulate the likelihood function of the correlations of 2D image matches as a Gaussian and Uniform mixture distribution which takes the uncertainty caused by illumination changes, image noise and moving objects into account. Meanwhile, an uncertainty prediction module is presented to predict the pixel-wise distribution parameters. Finally, we iteratively refine the depth and relative camera pose using the gradient-like information to maximize the likelihood function of the correlations. Extensive experimental results on several datasets prove that our method significantly outperforms the state-of-the-art end-to-end two-view SfM approaches in accuracy and generalization capability. Yuxi Xiao, Li Li 0047, Jian Yao 0002 |
IROS | 2 |
| 2021 | Grid Model-Based Global Color Correction for Multiple Image MosaickingabstractColor consistency optimization for multiple images is a challenging problem in image mosaicking. To facilitate the global color optimization, existing approaches mainly use less flexible models, e.g., linear or gamma function, to eliminate the color differences between multiple images. However, these models often struggle to eliminate the color differences that existed in the local areas and preserve the image gradient information. To solve this problem, we creatively propose a novel color-correction model, which comprised a series of local grid linear models. This model is simple, but it is flexible enough to approximate a variety of complicated local color variations. To obtain the optimal model parameters for each image globally, a specific cost function that considers both color consistency and gradient preservation is designed and solved. The aim of our approach is to generate a composite image with visually consistent color. The original color information may be destroyed. Thus, this approach is unsuitable for the quantitative remote sensing applications. The experimental results on several challenging data sets show that the proposed approach outperforms state-of-the-art approaches in both visual quality and quantitative metrics. Li Li 0047, Yunmeng Li, Menghan Xia, Yinxuan Li, Jian Yao 0002, Bin Wang 0100 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2021 | Extraction of Street Pole-Like Objects Based on Plane Filtering From Mobile LiDAR DataabstractPole-like objects provide important street infra- structure for road inventory and road mapping. In this article, we proposed a novel pole-like object extraction algorithm based on plane filtering from mobile Light Detection and Ranging (LiDAR) data. The proposed approach is composed of two parts. In the first part, a novel octree-based split scheme was proposed to fit initial planes from off-ground points. The results of the plane fitting contribute to the extraction of pole-like objects. In the second part, we proposed a novel method of pole-like object extraction by plane filtering based on local geometric feature restriction and isolation detection. The proposed approach is a new solution for detecting pole-like objects from mobile LiDAR data. The innovation in this article is that we assumed that each of the pole-like objects can be represented by a plane. Thus, the essence of extracting pole-like objects will be converted to plane selecting problem. The proposed method has been tested on three data sets captured from different scenes. The average completeness, correctness, and quality of our approach can reach up to 87.66%, 88.81%, and 79.03%, which is superior to state-of-the-art approaches. The experimental results indicate that our approach can extract pole-like objects robustly and efficiently. JingMin Tu, Jian Yao 0002, Li Li 0047, Binbin Xiang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | DeepCrack: A deep hierarchical feature learning architecture for crack segmentation
Jian Yao 0002, Xiaohu Lu, Renping Xie, Li Li 0047 |
Neurocomputing | 5 |
| 2019 | A Novel Octree-Based 3-D Fully Convolutional Neural Network for Point Cloud Classification in Road EnvironmentabstractThe automatic classification of 3-D point clouds is publicly known as a challenging task in a complex road environment. Specifically, each point is automatically classified into a unique category label, and then, the labels are used as clues for semantic analysis and scene recognition. Instead of heuristically extracting handcrafted features in traditional methods to classify all points, we put forward an end-to-end octree-based fully convolutional network (FCN) to classify 3-D point clouds in an urban road environment. There are four contributions in this paper. The first is that the integration and comprehensive uses of OctNet and FCN greatly decrease the computing time and memory demands compared with a dense 3-D convolutional neural network (CNN). The second is that the octree-based network is strengthened by means of modifying the cross-entropy loss function to solve the problems of an unbalanced category distribution. The third is that an Inception-ResNet block is united with our network, which enables our 3-D CNN to effectively learn how to classify scenes containing objects at multiple scales and improve classification accuracy. The last is that an open source data set (HuangshiRoad data set) with ten different classes is introduced for 3-D point cloud classification. Three representative data sets [Semantic3D, WHU_MLS (blocks I and II), and HuangshiRoad] with different covered areas and numbers of points and classes are selected to evaluate our proposed method. The experimental results show that the overall classification accuracy is appreciable, with 89.4% for Semantic3D, 82.9% for WHU_MLS block I, 91.4% for WHU_MLS block II, and 94% for HuangshiRoad. Our deep learning approach can efficiently classify 3-D dense point clouds in an urban road environment measured by a mobile laser scanning (MLS) system or static LiDAR. Binbin Xiang, JingMin Tu, Jian Yao 0002, Li Li 0047 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2018 | Multiple Combined Constraints for Image StitchingabstractSeveral approaches to image stitching use different constraints to estimate the motion model between image pairs. These constraints can be roughly divided into two categories: geometric constraints and photometric constraints. In this paper, geometric and photometric constraints are combined to improve the alignment quality, which is based on the observation that these two kinds of constraints are complementary. On the one hand, geometric constraints (e.g., point and line correspondences) are usually spatially biased and are insufficient in some extreme scenes, while photometric constraints are always evenly and densely distributed. On the other hand, photometric constraints are sensitive to displacements and are not suitable for images with large parallaxes, while geometric constraints are usually imposed by feature matching and are more robust to handle parallaxes. The proposed method therefore combines them together in an efficient mesh-based image warping framework. It achieves better alignment quality than methods only with geometric constraints, and can handle larger parallax than photometric-constraint-based method. Experimental results on various images illustrate that the proposed method outperforms representative state-of-the-art image stitching methods reported in the literature. Kai Chen 0024, JingMin Tu, Binbin Xiang, Li Li 0047, Jian Yao 0002 |
ICIP | 4 |
| 2018 | Edge-Enhanced Optimal Seamline Detection for Orthoimage MosaickingabstractIn this letter, we propose a novel algorithm to detect the optimal seamline by using the edge-enhanced energy function for orthoimage mosaicking. First, we generate the energy cost map by using traditional intensity and gradient difference. Then, to ensure that the detected optimal seamlines avoid crossing the obvious objects, the energy map is enhanced in the regions of valid edge segments around the obvious objects. The edge segments are separately detected from input images, and the invalid edge segments are removed by using texture complexity index, image difference, and consistency constraint. Finally, we detect the optimal seamline from this edge-enhanced energy cost map via graph cuts. Experimental results on several groups of challenging orthoimages show that the proposed algorithm is capable of creating high-quality seamlines for orthoimage mosaicking and outperforms state-of-the-art algorithms and the commercial software based on the visual comparison and statistical evaluation. Li Li 0047, Jian Yao 0002, Renping Xie |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2017 | Optimal seamline detection in dynamic scenes via graph cuts for image mosaicking
Li Li 0047, Jian Yao 0002, Haoang Li, Menghan Xia, Wei Zhang 0021 |
Mach. Vis. Appl. | 1 |
| 2017 | Globally consistent alignment for planar mosaicking via topology analysis
Menghan Xia, Jian Yao 0002, Renping Xie, Li Li 0047, Wei Zhang 0021 |
Pattern Recognit. | 4 |
| 2016 | Edge chain detection by applying Helmholtz principle on gradient magnitude mapabstractIn this paper, we present an efficient edge chain detection algorithm by applying the Helmholtz principle on the gradient magnitude map of an image. An edge chain validation method is proposed which uses the “relative number of false alarms” (RNFA) instead of the traditional “number of false alarms” (NFA). The edge chains are detected first and then validated according to their RNFA values. In this way, edge chains that are weak in gradients but meaningful in vision can be detected. To evaluate the proposed edge chain detector in quantity, an edge chain detection benchmark which consists of 25 labeled images in different scenes was built. The proposed edge chain detector was tested in this benchmark, and the experimental results sufficiently demonstrate that the proposed edge chain detector outperforms the state-of-the-art methods. Xiaohu Lu, Jian Yao 0002, Li Li 0047, Wei Zhang 0021 |
ICPR | 3 |
| 2016 | Joint point and line segment matching on wide-baseline stereo imagesabstractThis paper presents an method that matches points and line segments jointly on wide-baseline stereo images. In both two images to be matched, line segments are extracted and those spatially adjacent ones are intersected to generate V-junctions. To match V-junctions from the two images, we extract for each of them an affine and scale invariant local region and describe it with SIFT. The putative V-junction matches obtained from evaluating their description vectors are refined subsequently by the epipolar line constraint and topological distribution constraint among neighbor V-junctions. Since once a pair of V-junctions are matched, the two pairs of line segments forming them are matched accordingly. A part of line segments from the two images are therefore matched along with V-junction matches. To get more line segment matches, we further match those left unmatched line segments by the local homographies estimated from their adjacent V-junction matches. Experiments verify the robustness of the proposed method and its superiority to both some famous point and line segment matching methods on wide-baseline stereo images. In addition, we also show the proposed method can make it easier for 3D line segment reconstruction. Kai Li 0015, Jian Yao 0002, Menghan Xia, Li Li 0047 |
WACV | 4 |
| 2016 | Hierarchical line matching based on Line-Junction-Line structure descriptor and local homography estimation
Kai Li 0015, Jian Yao 0002, Xiaohu Lu, Li Li 0047 |
Neurocomputing | 4 |
| 2015 | CannyLines: A parameter-free line segment detectorabstractIn this paper, we present a robust line segment detection algorithm to efficiently detect the line segments from an input image. Firstly a parameter-free Canny operator, named as CannyPF, is proposed to robustly extract the edge map from an input image by adaptively setting the low and high thresholds for the traditional Canny operator. Secondly, both efficient edge linking and splitting techniques are proposed to collect collinear point clusters directly from the edge map, which are used to fit the initial line segments based on the least-square fitting method. Thirdly, longer and more complete line segments are produced via efficient extending and merging. Finally, all the detected line segments are validated due to the Helmholtz principle [1, 2] in which both the gradient orientation and magnitude information are considered. Experimental results on a set of representative images illustrate that our proposed line segment detector, named as CannyLines, can extract more meaningful line segments than two popularly used line segment detectors, LSD [3] and ED-Lines [4], especially on the man-made scenes. Xiaohu Lu, Jian Yao 0002, Kai Li 0015, Li Li 0047 |
ICIP | 4 |
| 2015 | Globally consistent alignment for mosaicking aerial imagesabstractIn this paper, we present a robust method to efficiently create a globally consistent and seamless mosaic from aerial images. Firstly, a globally consistent registration strategy is proposed to align the aerial images in a common coordinate system, which combines the affine model with the homographic model effectively. To suppress the accumulation of perspective distortions induced by a sequential set of aerial images taken from a wide-range region, we proposed to initially align each image by an affine model and then perform a homographic refinement in groups to increase the global consistency. Secondly, to efficiently conceal the parallax between aligned images in overlap regions with large depth differences where it is impossible to recover a highly accurate consistent image registration, a novel optimized seamline detection algorithm in the graph cuts energy minimization framework is proposed to find optimal seamlines within overlap regions for image mosaicking through rounding visually obvious foreground objects. Finally, experimental results on several representative image sets illustrate the superiority of our proposed approaches. Menghan Xia, Man Yao, Li Li 0047, Xiaohu Lu |
ICIP | 3 |