Li He 0002

dblp:02/1558-2 · DBLP profile ↗
← Back
22ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0003-0261-4068ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 7 since 2021Systems, architecture and hardware · 9 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Accurate and Robust UWB Localization With Incomplete Measurements Based on Multi-Modal Diffusion Model
Ming Sun 0025, Bo Yang 0019, Li He 0002, Hong Zhang 0013
IEEE Trans Autom. Sci. Eng.4
2026 Two-Step Nyström Sampling for Large-Scale Kernel Approximation
abstract
Nystrom approximation is one of the most popular approximation methods to accelerate kernel analysis on largescale data sets. Nystrom employs one single landmark set to ¨ obtain eigenvectors (low-rank decomposition) and projects the entire data set to the eigenvectors (embedding). Most existing methods focus on accelerating landmark selection. For extremely large-scale data sets, however, the embedding time cost, rather than that of low-rank decomposition, is critical. In addition, both accuracy and embedding time cost are dominated by the landmark set size. As a result, using more landmarks is the only way to improve accuracy at the cost of extremely high embedding costs. In this paper, we propose a method for the first time to decouple embedding cost from that of low-rank decomposition. We first obtain the eigenvectors from a large landmark set for a low error, and then optimize a small landmark set that minimizes the landmark-set-embedding error to ensure a low embedding cost. In return, our accuracy is close to that of the large landmark set but the small one dominates the embedding time cost. Our method can deal with popular kernels and be plugged into most existing methods. Experimental results demonstrate the superiority of the proposed method.
Li He 0002, Hong Zhang 0013
IEEE Trans. Big Data1
2024 PISR: Polarimetric Neural Implicit Surface Reconstruction for Textureless and Specular Objects
Guangcheng Chen, Yicheng He, Li He 0002, Hong Zhang 0013
ECCV (8)3
2024 SWCF-Net: Similarity-weighted Convolution and Local-global Fusion for Efficient Large-scale Point Cloud Semantic Segmentation
abstract
Large-scale point cloud consists of a multitude of individual objects, thereby encompassing rich structural and underlying semantic contextual information, resulting in a challenging problem in efficiently segmenting a point cloud. Most existing researches mainly focus on capturing intricate local features without giving due consideration to global ones, thus failing to leverage semantic context. In this paper, we propose a Similarity-Weighted Convolution and local-global Fusion Network, named SWCF-Net, which takes into account both local and global features. We propose a Similarity-Weighted Convolution (SWConv) to effectively extract local features, where similarity weights are incorporated into the convolution operation to enhance the generalization capabilities. Then, we employ a downsampling operation on the K and V channels within the attention module, thereby reducing the quadratic complexity to linear, enabling Transformer to deal with large-scale point cloud. At last, orthogonal components are extracted in the global features and then aggregated with local features, thereby eliminating redundant information between local and global features and consequently promoting efficiency. We evaluate SWCF-Net on large-scale outdoor datasets SemanticKITTI and Toronto3D. Our experimental results demonstrate the effectiveness of the proposed network. Our method achieves a competitive result with less computational cost, and is able to handle large-scale point clouds efficiently. The code is available at https://github.com/Sylva-Lin/SWCF-Net.
Zhenchao Lin, Li He 0002, Hongqiang Yang, Xiaoqun Sun, Guojin Zhang, Weinan Chen, Yisheng Guan, Hong Zhang 0013
IROS2
2024 Robust Data Association Against Detection Deficiency for Semantic SLAM
abstract
Robust and accurate object association is essential for precise 3D object landmark inference in semantic Simultaneous Localization and Mapping (SLAM), and yet remains challenging due to the detection deficiency caused by high miss detection rate, false alarm, occlusion and limited field-of-view, etc. The 2D location of an object is a crucial complementary cue to the appearance feature, especially in the case of associating objects across frames under large viewpoint changes. However, motion model or trajectory pattern based methods struggle to infer object motion reliably with a moving camera. In this paper, by exploiting the local projective warping consistency, a local homography based 2D motion inference method is proposed to sequentially estimate the object location along with uncertainty. By integrating the deep appearance feature and semantic information, an object association method, named HOA, which is robust to detection deficiency is proposed. Experimental evaluations suggest that the proposed motion prediction method is capable of maintaining a low cumulative error over a long duration, which enhances the object association performance in both accuracy and robustness. Note to Practitioners—This work aims to consistently associate 2D detection boxes corresponding to the same 3D object across images. In tasks of landmark-based navigation, collision avoidance, grasping and manipulation, objects in the task space are commonly simplified into 3D enveloping surfaces (e.g. cuboid or ellipsoid) by using 2D object detection boxes from multiple image views, and accurate data association is a prerequisite for precise enveloping surface reconstruction. This problem remains challenging considering the imperfect object detections, the appearance similarity of objects and the unpredictable trajectory of the moving camera. This work proposes a long-term reliable 2D location prediction algorithm that is capable of handling the complex motion of the target. Along with the appearance feature extracted by a retrain-free deep learning based model, this work proposes an object association method that can simultaneously deal with multiple objects with unknown object categories under the moving camera scenario.
Xubin Lin, Jiahao Ruan, Yirui Yang, Li He 0002, Yisheng Guan, Hong Zhang 0013
IEEE Trans Autom. Sci. Eng.4
2023 Combining Scene Coordinate Regression and Absolute Pose Regression for Visual Relocalization
abstract
Visual relocalization is a fundamental problem in computer vision and robotics. Recently, regression-based methods become popular and they can be categorized into two classes: absolute pose regression and scene coordinate regression. In this work, we present a combined regression network that jointly learns scene coordinate regression and absolute pose regression for single-image visual relocalization. The proposed network composes of a feature encoder and two regression branches with uncertainty modeling. In particular, we design a deep feature conditioning module, aiming at propagating the coarse pose information in absolute pose regression to inform the predictions in scene coordinate regression. The proposed network is trained in an end-to-end fashion to learn both regression tasks. Moreover, we propose an uncertainty-driven RANSAC algorithm that incorporates the predicted scene coordinates and their uncertainties to solve the camera pose during inference. To the best of our knowledge, this work is the first to combine scene coordinate regression and pose regression in a hierarchical framework for visual relocalization. Experiments on indoor and outdoor benchmarks demonstrate the effectiveness and the superiority of the proposed method over the state-of-the-art methods.
Jiahao Ruan, Li He 0002, Yisheng Guan, Hong Zhang 0013
ICRA2
2023 Robot Person Following Under Partial Occlusion
abstract
Robot person following (RPF) is a capability that supports many useful human-robot-interaction (HRI) applications. However, existing solutions to person following often as-sume full observation of the tracked person. As a consequence, they cannot track the person reliably under partial occlusion where the assumption of full observation is not satisfied. In this paper, we focus on the problem of robot person following under partial occlusion caused by a limited field of view of a monocular camera. Based on the key insight that it is possible to locate the target person when one or more of hislher joints are visible, we propose a method in which each visible joint contributes a location estimate of the followed person. Experiments on a public person-following dataset show that, even under partial occlusion, the proposed method can still locate the person more reliably than the existing SOTA methods. As well, the application of our method is demonstrated in real experiments on a mobile robot.
Hanjing Ye, Jieting Zhao, Yaling Pan, Weinan Chen, Li He 0002, Hong Zhang 0013
ICRA5
2023 Doubly Stochastic Distance Clustering
abstract
In doubly stochastic (DS) clustering, it is common to initialize the DS matrix with a similarity matrix and use the eigen-decomposition of the DS-scaled similarity matrix to obtain the optimal cluster indicators. The selection of a proper initial similarity measure, however, is a difficult problem and the eigen-decomposition is time-consuming, with time complexity of$O(n^{3})$where$n$is the data size. In this paper, we propose to replace the DS similarity matrix with the DS Euclidean distance matrix for clustering. We show that the optimal cluster indicators minimize the$k$-medoids error of data with DS Euclidean distance. We propose a fast method to obtain data with DS distance for clustering. Compared with DS similarity clustering, DS distance clustering is kernel-free and of low time complexity, typically$O(nd^{2}+d^{3})$where$d$is the input dimension. Experimental results on real-world datasets and the image segmentation task verify the superiority of our DS distance clustering over several competing methods.
Li He 0002, Hong Zhang 0013
IEEE Trans. Circuits Syst. Video Technol.1
2022 Perspective Phase Angle Model for Polarimetric 3D Reconstruction
Guangcheng Chen, Li He 0002, Yisheng Guan, Hong Zhang 0013
ECCV (2)2
2022 NDD: A 3D Point Cloud Descriptor Based on Normal Distribution for Loop Closure Detection
abstract
Loop closure detection is a key technology for long-term robot navigation in complex environments. In this paper, we present a global descriptor, named Normal Distribution Descriptor (NDD), for 3D point cloud loop closure detection. The descriptor encodes both the probability density score and entropy of a point cloud as the descriptor. We also propose a fast rotation alignment process and use correlation coefficient as the similarity between descriptors. Experimental results show that our approach outperforms the state-of-the-art point cloud descriptors in both accuracy and efficency. The source code is available and can be integrated into existing LiDAR odometry and mapping (LOAM) systems.
Li He 0002, Hong Zhang 0013, Xubin Lin, Yisheng Guan
IROS2
2022 Curvature-Variation-Inspired Sampling for Point Cloud Classification and Segmentation
abstract
Point cloud is a discrete and unordered expression of 3D data. A lot of methods have been proposed to solve the problem in 3D object classification and scene recognition. To handle the huge amount of unordered point cloud, down-sampling before processing is needed. The shortage of existing sampling methods is the lack of geometry information consideration, which is essential for point cloud classification and segmentation tasks. Our method is mainly motivated by the observation that points with a high curvature variation can depict the outlines of objects. Thus, we propose a curvature variation based sampling method for point cloud classification and segmentation tasks. We aim to sample points with high curvature variations, which are considered to be more suitable for classification and segmentation tasks than the traditional sampling method. We combine the proposed sampling algorithm with the existing sampling method for multiple information fusion, and a higher accuracy and mean IoU can be achieved. The experimental results verify the advantage of considering curvature variation in classification and segmentation tasks.
Weinan Chen, Xubin Lin, Li He 0002, Yisheng Guan
IEEE Signal Process. Lett.4
2021 Robust Improvement in 3D Object Landmark Inference for Semantic Mapping
abstract
Recent works on semantic Simultaneous Localization and Mapping (SLAM) utilizing object landmarks have shown superiority in terms of robustness and accuracy in tracking and localization. 3D object landmarks represented by a cubic or quadric surface are inferred from 2D object bounding boxes which are typically captured from multiple views by an object detector. Nevertheless, bounding box noises and small camera baseline may lead to an inaccurate 3D object landmark inference. Inspired by the dual quadric enveloping property, in this work, we introduce the horizontal support assumption to constrain rotation w.r.t. roll and pitch for a quadric representation. As the result, we reduce the number of quadric parameters and narrow down the solution space, and ultimately produce a relatively accurate inference. Extensive experimental evaluations under both simulated and real scenarios are conducted in this paper. Quantitative results demonstrate that our approach outperforms the state-of-the-art.
Xubin Lin, Yirui Yang, Li He 0002, Weinan Chen, Yisheng Guan, Hong Zhang 0013
ICRA3
2020 Keypoint Description by Descriptor Fusion Using Autoencoders
abstract
Keypoint matching is an important operation in computer vision and its applications such as visual simultaneous localization and mapping (SLAM) in robotics. This matching operation heavily depends on the descriptors of the keypoints, and it must be performed reliably when images undergo conditional changes such as those in illumination and viewpoint. In this paper, a descriptor fusion model (DFM) is proposed to create a robust keypoint descriptor by fusing CNN-based descriptors using autoencoders. Our DFM architecture can be adapted to either trained or pre-trained CNN models. Based on the performance of existing CNN descriptors, we choose HardNet and DenseNet169 as representatives of trained and pre-trained descriptors. Our proposed DFM is evaluated on the latest benchmark datasets in computer vision with challenging conditional changes. The experimental results show that DFM is able to achieve state-of-the-art performance, with the mean mAP that is 6.45% and 6.53% higher than HardNet and DenseNet169, respectively.
Zhuang Dai, Xinghong Huang, Weinan Chen, Chuangbing Chen, Li He 0002, Shuhuan Wen, Hong Zhang 0013
ICRA5
2019 A Comparison of CNN-Based and Hand-Crafted Keypoint Descriptors
abstract
Keypoint matching is an important operation in computer vision and its applications such as visual simultaneous localization and mapping (SLAM) in robotics. This matching operation heavily depends on the descriptors of the keypoints, and it must be performed reliably when images undergo condition changes such as those in illumination and viewpoint. Previous research in keypoint description has pursued three classes of descriptors: hand-crafted, those from trained convolutional neural networks (CNN), and those from pre-trained CNNs. This paper provides a comparative study of the three classes of keypoint descriptors, in terms of their ability to handle conditional changes. The study is conducted on the latest benchmark datasets in computer vision with challenging conditional changes. Our study finds that (a) in general CNN-based descriptors outperform hand-crafted descriptors, (b) the trained CNN descriptors perform better than pre-trained CNN descriptors with respect to viewpoint changes, and (c) pre-trained CNN descriptors perform better than trained CNN descriptors with respect to illumination changes. These findings can serve as a basis for selecting appropriate keypoint descriptors for various applications.
Zhuang Dai, Xinghong Huang, Weinan Chen, Li He 0002, Hong Zhang 0013
ICRA4
2019 Improving Keypoint Matching Using a Landmark-Based Image Representation
abstract
Motivated by the need to improve the performance of visual loop closure verification via multi-view geometry (MVG) under significant illumination and viewpoint changes, we propose a keypoint matching method that uses landmarks as an intermediate image representation in order to leverage the power of deep learning. In environments with various changes, the traditional verification method via MVG may encounter difficulty because of their inability to generate a sufficient number of correctly matched keypoints. Our method exploits the excellent invariance properties of convolutional neural network (ConvNet) features, which have shown outstanding performance for matching landmarks between images. By generating and matching landmarks first in the images and then matching the keypoints within the matched landmark pairs, we can significantly improve the quality of matched keypoints in terms of precision and recall measures. The proposed method is validated on challenging datasets that involve significant illumination and viewpoint changes, to establish its superior performance to the standard keypoint matching method.
Xinghong Huang, Zhuang Dai, Weinan Chen, Li He 0002, Hong Zhang 0013
ICRA4
2019 Fast Large-Scale Spectral Clustering via Explicit Feature Mapping
abstract
We propose an efficient spectral clustering method for large-scale data. The main idea in our method consists of employing random Fourier features to explicitly represent data in kernel space. The complexity of spectral clustering thus is shown lower than existing Nyström approximations on largescale data. With m training points from a total of n data points, Nyström method requires O(nmd + m3+ nm2) operations, where d is the input dimension. In contrast, our proposed method requires O(nDd + D3+ n'D2), where n' is the number of data points needed until convergence and D is the kernel mapped dimension. In large-scale datasets where n ≪ n hold true, our explicitly mapping method can significantly speed up eigenvector approximation and benefit prediction speed in spectral clustering. For instance, on MNIST (60000 data points), the proposed method is similar in clustering accuracy to Nyström methods while its speed is twice as fast as Nyström.
Li He 0002, Nilanjan Ray, Yisheng Guan, Hong Zhang 0013
IEEE Trans. Cybern.1
2019 An Efficient Edge Artificial Intelligence MultiPedestrian Tracking Method With Rank Constraint
abstract
Characterized by the ability to handle varying number of objects, tracking by detection framework becomes increasingly popular in multiobject tracking (MOT) problem. However, the tracking performance heavily depends on the object detector. Considering that data association optimization and association affinity model are two key parts in MOT, an online multipedestrian tracking method is proposed to formulate a more effective association affinity model. It includes a two-step data association taking advantage of rank-based dynamic motion affinity model. The rank-based dynamic motion affinity model is used to estimate the object state and refine the trajectory for each of target to achieve the noiseless trajectory. Both strategies are beneficial to eliminate ambiguous detection responses during association. To fairly verify the proposed method, three public datasets are adopted. Both qualitative and quantitative experiment results demonstrate the superiorities of the proposed tracking algorithm in comparison with its counterparts.
Honghong Yang, Jinming Wen, Xiaojun Wu 0002, Li He 0002, Shahid Mumtaz
IEEE Trans. Ind. Informatics4
2018 Online multiple objects tracking with detection reliability prior constraint
Honghong Yang, Li He 0002
Multim. Tools Appl.2
2018 Kernel K-Means Sampling for Nyström Approximation
abstract
A fundamental problem in Nyström-based kernel matrix approximation is the sampling method by which training set is built. In this paper, we suggest to use kernel -means sampling, which is shown in our works to minimize the upper bound of a matrix approximation error. We first propose a unified kernel matrix approximation framework, which is able to describe most existing Nyström approximations under many popular kernels, including Gaussian kernel and polynomial kernel. We then show that, the matrix approximation error upper bound, in terms of the Frobenius norm, is equal to the -means error of data points in kernel space plus a constant. Thus, the -means centers of data in kernel space, or the kernel -means centers, are the optimal representative points with respect to the Frobenius norm error upper bound. Experimental results, with both Gaussian kernel and polynomial kernel, on real-world data sets and image segmentation tasks show the superiority of the proposed method over the state-of-the-art methods.
Li He 0002, Hong Zhang 0013
IEEE Trans. Image Process.1
2017 Error bound of Nyström-approximated NCut eigenvectors and its application to training size selection
Li He 0002, Nilanjan Ray, Hong Zhang 0013
Neurocomputing1
2016 M2DP: A novel 3D point cloud descriptor and its application in loop closure detection
abstract
In this paper, we present a novel global descriptor M2DP for 3D point clouds, and apply it to the problem of loop closure detection. In M2DP, we project a 3D point cloud to multiple 2D planes and generate a density signature for points for each of the planes. We then use the left and right singular vectors of these signatures as the descriptor of the 3D point cloud. Our experimental results show that the proposed algorithm outperforms state-of-the-art global 3D descriptors in both accuracy and efficiency.
Li He 0002, Xiaolong Wang 0005, Hong Zhang 0013
IROS1
2016 Iterative ensemble normalized cuts
Li He 0002, Hong Zhang 0013
Pattern Recognit.1