Fulin Tang

dblp:231/6787 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-8474-2671ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 8 since 2021Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SLNMapping: Super Lightweight Neural Mapping in Large-Scale Scenes
Chenhui Shi 0001, Fulin Tang, Hao Wei 0008, Yihong Wu 0002
Int. J. Comput. Vis.2
2025 3D-SLNR: A Super Lightweight Neural Representation for Large-scale 3D Mapping
abstract
We propose 3D-SLNR, a new and ultra-lightweight neural representation with outstanding performance for large-scale 3D mapping. The representation defines a global signed distance function (SDF) in near-surface space based on a set of band-limited local SDFs anchored at support points sampled from point clouds. These SDFs are parameterized only by a tiny multi-layer perceptron (MLP) with no latent features, and the state of each SDF is modulated by three learnable geometric properties: position, rotation, and scaling, which make the representation adapt to complex geometries. Then, we develop a novel parallel algorithm tailored for this unordered representation to efficiently detect local SDFs where each sampled point is located, allowing for real-time updates of local SDF states during training. Additionally, a prune-and-expand strategy is introduced to enhance adaptability further. The synergy of our low-parameter model and its adaptive capabilities results in an extremely compact representation with excellent expressiveness. Extensive experiments demonstrate that our method achieves state-of-the-art reconstruction performance with less than 1/5 of the memory footprint compared with previous advanced methods.
Chenhui Shi 0001, Fulin Tang, Ning An 0002, Yihong Wu 0002
CVPR2
2025 A Novel Framework for Learning Bézier Decomposition From 3D Point Clouds
abstract
This paper proposes a fully differentiable and end-to-end framework for learning Bézier decomposition on 3D point clouds. The framework aims to partition input point clouds into multiple Bézier primitive patches through a learned Bézier decomposition process. Unlike previous approaches that handle different primitive types separately, thus being limited to specific shape categories, our method seeks to achieve a generalized primitive segmentation on point clouds. Drawing inspiration from Bézier decomposition on NURBS models, we adapt it to guide point cloud segmentation without relying on pre-defined primitive types. To achieve this, we introduce a joint optimization framework that simultaneously learns Bézier primitive segmentation and geometric fitting in a cascaded architecture. Additionally, we propose a soft voting regularizer to enhance primitive segmentation and an auto-weight embedding module to effectively cluster point features, making the network more robust and applicable to various scenarios. Furthermore, we incorporate a reconstruction module capable of processing multiple CAD models with different primitives simultaneously. Extensive experiments were conducted on both synthetic ABC datasets and real-scan datasets to validate and compare our approach against several baseline methods. The results demonstrate that our method outperforms previous work in terms of segmentation accuracy, while also exhibiting significantly faster inference speed.
Rao Fu 0004, Qian Li 0075, Cheng Wen 0001, Ning An 0002, Fulin Tang
IEEE Trans. Circuits Syst. Video Technol.5
2025 CornerVINS: Accurate Localization and Layout Mapping for Structural Environments Leveraging Hierarchical Geometric Representations
abstract
A compact and consistent map of surroundings is critical for intelligent robots to understand their situations and realize robust navigation. Most existing techniques rely on infinite planes, which are sensitive to pose drift and may lead to confusing maps. Towards high-level perception in indoor environments, we propose CornerVINS, an innovative RGB-D inertial localization and layout mapping method leveraging hierarchical geometric features, i.e., points, planes, and box corners. Specifically, points are enhanced by fusing depth information, and planes are modeled as bounded patches using convex hulls to increase their discriminability. More importantly, box corners, lying at the intersection of three orthogonal planes, are parameterized with a 6-dimensional vector and integrated into the extended Kalman filter for the first time. We introduce a hierarchical mechanism to effectively extract and associate planes and corners, which are considered as layout components of scenes and serve as long-term landmarks to correct camera poses. Extensive experiments prove that the proposed box corners bring significant improvements, enabling accurate localization and consistent layout mapping at low computational cost. Overall, the proposed CornerVINS outperforms state-of-the-art systems in both accuracy and efficiency. The core implementations are available athttps://github.com/zydddd/CornerVINS.
Fulin Tang, Yihong Wu 0002
IEEE Trans. Robotics2
2024 A Region-Growing Supervised Geometry-Weighted Transformer for Normal Estimation
abstract
This paper addresses the challenge of accurately estimating unoriented normals for 3D point clouds, particularly in the presence of noise, varying densities, and sharp features. To overcome these challenges, we introduce a novel network architecture that integrates a geometry-weighted transformer module with a dual backbone network. This combined architecture enables accurate prediction of normals from patch embeddings that encode edge and spatial features. Additionally, we propose a region-growing loss to supervise the network, which can aid in normal estimation for sharp features by segmenting input patches based on normal discontinuities, thus avoiding the over-smoothness for normals in the sharp feature areas. Extensive experiments on PCPNet and SceneNN datasets highlight the effectiveness of our approach, leading to state-of-the-art performance against baseline methods.
Rao Fu 0004, Qian Li 0075, Cheng Wen 0001, Ning An 0002, Fulin Tang
ICME5
2024 Depth assisted novel view synthesis using few images
abstract
In this paper, we introduce a novel approach to improve the performance of Neural Radiance Fields (NeRF) from limited input views. NeRF has exhibited impressive capabilities in producing photo-realistic renderings when trained on dense input views, but its performance degrades as the number of training views decreases. Our key insight is that the original NeRF lacks geometric regularization and appearance information due to limited inputs, resulting in an over-fitting issue. To address this challenge, we present a novel method: first, a global sampling method with geometric regularization is employed by utilizing warped images as additional pseudo-views, which optimizes the multi-view consistency during the training. Second, we introduce a local patch sampling technique with perceptual regularization to ensure pixel correspondence in appearance. Furthermore, we incorporate depth information for explicit geometry regularization . We evaluate our method on the DTU dataset and LLFF dataset from a different number of inputs. Extensive evaluations demonstrate that our approach outperforms existing benchmarks across various metrics, achieving state-of-the-art results.
Qian Li 0075, Rao Fu 0004, Fulin Tang
Image Vis. Comput.3
2024 Novel camera self-calibration method with clustering prior and nonlinear optimization from an image sequence
Xiaohui Jiang, Haijiang Zhu, Ning An 0002, Binjian Xie, Hao Wei 0008, Fulin Tang, Yihong Wu 0002
Multim. Tools Appl.6
2024 Covariant Peak Constraint for Accurate Keypoint Detection and Keypoint-Specific Descriptor Learning
abstract
Local feature extraction consists of keypoint detection and local descriptor extraction. Firstly, in keypoint detector learning, existing covariance constraint loss functions cannot constrain the probability distribution shapes in local probability maps that surround keypoints. And existing auxiliary peak loss functions, which are used to alleviate the problem, impair the performance of local feature methods. To solve this problem, we propose a novel Covariant Peak constraint Loss (CP Loss) which is defined as the expectations of local probability maps' position errors. Minimizing our CP Loss can make local probability maps accurately peak at reliable keypoints. Secondly, in descriptor learning, the Neural Reprojection Error (NRE) aims at constraining dense descriptor maps of images. But we argue that only those descriptors of keypoints need to be constrained. Thus, we propose a novel Conditional Neural Reprojection Error (CNRE) that is only conditioned on keypoints. Compared with the NRE, our CNRE can achieve much higher efficiency and produce more keypoint-specific descriptors with better matching performance. We use our CP Loss and CNRE to train a local feature network named as CPCN-Feat. Experimental results show that our CPCN-Feat achieves state-of-the-art performance on four challenging benchmarks.
Yujie Fu, Fulin Tang, Yihong Wu 0002
IEEE Trans. Multim.3
2023 PLPL-VIO: A Novel Probabilistic Line Measurement Model for Point-Line-Based Visual-Inertial Odometry
abstract
Point and line features are complementary in Visual-Inertial Odometry (VIO) or Visual-Inertial Simultaneous Localization And Mapping (VI-SLAM) systems. The advantage of combining these two types of features relies on their proper weighting in the cost function, usually set by their uncertainty. Compared with point features, setting line segment endpoints' uncertainty with isotropic distribution is unreasonable. But the uncertainty of line feature observation, especially for the endpoints' uncertainty along the line, is difficult to set due to occlusion and fragmentation problems. In this article, we use infinite lines as the line feature observations and prove that the uncertainty of these observations is only related to the vertical uncertainty of the endpoints, thus avoiding setting the parallel uncertainty of the endpoints. Besides, we introduce a novel consistent measurement model for line features. Furthermore, for long-time constraints, we add 3D line segments into the state vector and derive how to update them properly. Finally, we construct a point-line-based VIO system that takes into account the uncertainty of line feature observations and the consistency of line feature measurements. The proposed VIO system is validated on two public datasets. The results show that the proposed method obtains the best accuracy compared with the state-of-the-art point-based VIO systems (OpenVINS, VINS-Mono), a point-line-based VIO system (PL-VINS), and a structural line-based system (StructVIO).
Zewen Xu, Hao Wei 0008, Fulin Tang, Yihong Wu 0002, Gang Ma 0007, Shuzhe Wu, Xin Jin 0004
IROS3
2023 Efficient 6-DoF camera pose tracking with circular edges
Fulin Tang, Shaohuan Wu, Zhengda Qian, Yihong Wu 0002
Comput. Vis. Image Underst.1
2022 Automatic Detection and Fitting of Ellipse Markers Using EllipseNet
abstract
Ellipses are important elements in projective geometry. Accurate extraction of ellipse information is the first step in many computer vision applications, such as ellipse-based camera calibration and camera pose estimation. At present, most ellipse detection algorithms rely on the edge features extracted by Canny, which leads to a number of wrong detection results since non-ellipse edges features are also involved. To address this problem, this paper proposes a novel ellipse marker detection neural network, called EllipseNet. Notably, we propose a new loss function to enhance the rotation perception ability of EllipseNet. Furthermore, a novel ellipse marker data enhancement method is proposed for saving the time cost of labelling ellipse parameters. Experiments show that EllipseNet can improve the detection precision of ellipse regions by more than 3% improvement compared with other SOTA general object detectors.
Zhengda Qian, Fulin Tang, Bingxi Liu 0001, Yujie Fu, Shaohuan Wu, Xiaohong Jia 0001, Yihong Wu 0002
ICPR2
2022 VidSfM: Robust and Accurate Structure-From-Motion for Monocular Videos
abstract
With the popularization of smartphones, larger collection of videos with high quality is available, which makes the scale of scene reconstruction increase dramatically. However, high-resolution video produces more match outliers, and high frame rate video brings more redundant images. To solve these problems, a tailor-made framework is proposed to realize an accurate and robust structure-from-motion based on monocular videos. The key ideas include two points: one is to use the spatial and temporal continuity of video sequences to improve the accuracy and robustness of reconstruction; the other is to use the redundancy of video sequences to improve the efficiency and scalability of system. Our technical contributions include an adaptive way to identify accurate loop matching pairs, a cluster-based camera registration algorithm, a local rotation averaging scheme to verify the pose estimate and a local images extension strategy to reboot the incremental reconstruction. In addition, our system can integrate data from different video sequences, allowing multiple videos to be simultaneously reconstructed. Extensive experiments on both indoor and outdoor monocular videos demonstrate that our method outperforms the state-of-the-art approaches in robustness, accuracy and scalability.
Hainan Cui, Diantao Tu, Fulin Tang, Pengfei Xu 0013, Hongmin Liu 0001, Shuhan Shen
IEEE Trans. Image Process.3
2021 A Flexible and Efficient Loop Closure Detection Based on Motion Knowledge
abstract
Loop closure detection (LCD) is an essential module for simultaneous localization and mapping (SLAM), which can correct accumulated errors after long-term explorations. The widely used bag-of-words (BoW) model can not satisfy well the requirements of both low time consumption and high accuracy for a mobile platform. In this paper, we propose a novel LCD algorithm based on motion knowledge. We give a flexible and efficient detection strategy and also give flexible and efficient combinations of a global binary feature extracted by convolutional neural network (CNN) and a hand-crafted local binary feature. We take a continuous motion model, grid-based motion statistics (GMS) and motion states as motion knowledge. Furthermore, we fuse the proposed LCD with a visual-inertial odometry (VIO) system to correct localization errors by a pose graph optimization. Comparative experiments with state-of-the-art LCD algorithms on typical datasets have been carried out, and the results demonstrate that our proposed method achieves quite high recall rates and quite high speed at 100% precision. Moreover, experimental results from VIO further validate the effectiveness of the proposed method.
Bingxi Liu 0001, Fulin Tang, Yujie Fu, Yanqun Yang, Yihong Wu 0002
ICRA2
2021 Highly Efficient Line Segment Tracking with an IMU-KLT Prediction and a Convex Geometric Distance Minimization
abstract
Line segment features become popular in SLAM community. Usually, line-based SLAM systems utilize local appearance descriptors for line segment tracking. However, traditional descriptor-based line segment tracking algorithms suffer from the problem that accuracy and speed cannot be possessed simultaneously, which affects the performance of line-based SLAM systems negatively. We propose a novel line segment tracking method with an IMU-KLT line segment prediction and a convex geometric distance minimization to boost line segment tracking performance in both accuracy and speed. Particularly, the proposed convex geometric distance minimization uses a ℓ1-norm model to minimize geometric constraints between predicted line segments and extracted line segments efficiently. Furthermore, the line segment tracking is embedded into a VIO system and we adapt it to obtain more reliable point tracking. Experimental results on public datasets show that the proposed line segment tracking method achieves much higher accuracy and much less time cost than state-of-the-art level, where not only the number of correct matches increases but also the inlier ratios are increased by at least 35.1% along with a 3 times faster speed. Besides, the VIO system combining the proposed line segment tracking is improved in terms of accuracy.
Hao Wei 0008, Fulin Tang, Chaofan Zhang, Yihong Wu 0002
ICRA2
2021 High Power-Efficient and Performance-Density FPGA Accelerator for CNN-Based Object Detection
Chaofan Zhang, Fulin Tang, Yihong Wu 0002, Xuezhi Yang
PRCV (1)4
2020 3D Mapping and 6D Pose Computation for Real Time Augmented Reality on Cylindrical Objects
abstract
Visual Augmented Reality (AR) typically overlays virtual computer graphics or other virtual contents on the real world videos, attracting much interest from both academic and industrial communities. Although AR techniques on planes are well studied, cylindrical objects are seldom used for augmented reality. In this paper, we propose a new method for 3D reconstruction and 6D pose computation for augmented reality on a cylindrical object. The 6D pose is the relative pose between the camera and the cylindrical object, which is very convenient to make augmented reality. First, we capture some images with a cylindrical object and then reconstruct its 3D model with textures offline by using projective invariance and image contours. Second, according to the 3D model, we track the 6D relative pose between the camera and the cylindrical object online, where we propose a linear P3P RANSAC to remove outliers. Finally, the virtual images are exactly aligned with the cylindrical object in the real world. Experimental results show that the proposed method outperforms the state of the arts in terms of 3D mapping and 6D pose computation on cylindrical objects.
Fulin Tang, Yihong Wu 0002, Xiaohui Hou, Haibin Ling
IEEE Trans. Circuits Syst. Video Technol.1
2019 FMD Stereo SLAM: Fusing MVG and Direct Formulation Towards Accurate and Fast Stereo SLAM
abstract
We propose a novel stereo visual SLAM framework considering both accuracy and speed at the same time. The framework makes full use of the advantages of key-feature-based multiple view geometry (MVG) and direct-based formulation. At the front-end, our system performs direct formulation and constant motion model to predict a robust initial pose, reprojects local map to find 3D-2D correspondence and finally refines pose by the reprojection error minimization. This frontend process makes our system faster. At the back-end, MVG is used to estimate 3D structure. When a new keyframe is inserted, new mappoints are generated by triangulating. In order to improve the accuracy of the proposed system, bad mappoints are removed and a global map is kept by bundle adjustment. Especially, the stereo constraint is performed to optimize the map. This back-end process makes our system more accurate. Experimental evaluation on EuRoC dataset shows that the proposed algorithm can run at more than 100 frames per second on a consumer computer while achieving highly competitive accuracy.
Fulin Tang, Heping Li, Yihong Wu 0002
ICRA1
2019 Efficient conic fitting with an analytical Polar-N-Direction geometric distance
Yihong Wu 0002, Haoren Wang, Fulin Tang, Zhiheng Wang 0001
Pattern Recognit.3
2019 Real-time human segmentation by BowtieNet and a SLAM-based human AR system
abstract
Generally, it is difficult to obtain accurate pose and depth for a non-rigid moving object from a single RGB camera to create augmented reality (AR). In this study, we build an augmented reality system from a single RGB camera for a non-rigid moving human by accurately computing pose and depth, for which two key tasks are segmentation and monocular Simultaneous Localization and Mapping (SLAM). Most existing monocular SLAM systems are designed for static scenes, while in this AR system, the human body is always moving and non-rigid. In order to make the SLAM system suitable for a moving human, we first segment the rigid part of the human in each frame. A segmented moving body part can be regarded as a static object, and the relative motions between each moving body part and the camera can be considered the motion of the camera. Typical SLAM systems designed for static scenes can then be applied. In the segmentation step of this AR system, we first employ the proposed BowtieNet, which adds the atrous spatial pyramid pooling (ASPP) of DeepLab between the encoder and decoder of SegNet to segment the human in the original frame, and then we use color information to extract the face from the segmented human area. Based on the human segmentation results and a monocular SLAM, this system can change the video background and add a virtual object to humans. The experiments on the human image segmentation datasets show that BowtieNet obtains state-of-the-art human image segmentation performance and enough speed for real-time segmentation. The experiments on videos show that the proposed AR system can robustly add a virtual object to humans and can accurately change the video background.
Xiaomei Zhao, Fulin Tang, Yihong Wu 0002
Virtual Real. Intell. Hardw.2