Guowei Wan

dblp:50/8109 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
4since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 3 since 2021Systems, architecture and hardware · 4 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
3D vision · 43% Robot navigation and mapping · 36% Autonomous driving · 13%
Computer graphics and multimedia
1 paper
Geometric modeling and processing · 54% Image and video processing · 46%

Topics — the 19 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
point cloud registration
1.432023
A Unified BEV Model for Joint Learning of 3D Local Features and Overlap Estimation · ICRA 2023
DeepVCP: An End-to-End Deep Neural Network for Point Cloud Registration · ICCV 2019
L3-Net: Towards Learning Based LiDAR Localization for Autonomous Driving · CVPR 2019
Robotics › Robot navigation and mapping
localization
1.132020
LiDAR Inertial Odometry Aided Robust LiDAR Localization System in Changing City Scenes · ICRA 2020
L3-Net: Towards Learning Based LiDAR Localization for Autonomous Driving · CVPR 2019
Robust and Precise Vehicle Localization Based on Multi-Sensor Fusion in Diverse City Scenes · ICRA 2018
Robotics › Robot navigation and mapping › localization › range-based localization
LiDAR localization
0.822020
LiDAR Inertial Odometry Aided Robust LiDAR Localization System in Changing City Scenes · ICRA 2020
L3-Net: Towards Learning Based LiDAR Localization for Autonomous Driving · CVPR 2019
Robotics › Autonomous driving › HD map construction
lane graph extraction
0.812024
RoadPainter: Points Are Ideal Navigators for Topology TransformER · ECCV (12) 2024
Computer vision › 3D vision › point cloud processing
point cloud navigation
0.812024
RoadPainter: Points Are Ideal Navigators for Topology TransformER · ECCV (12) 2024
Machine learning › Graph learning › geometric learning › topological deep learning
topological transformer
0.812024
RoadPainter: Points Are Ideal Navigators for Topology TransformER · ECCV (12) 2024
Computer vision › 3D vision › 3d shape representation › 3d shape representation learning
3d descriptor learning
0.712023
A Unified BEV Model for Joint Learning of 3D Local Features and Overlap Estimation · ICRA 2023
Computer vision › 3D vision › local feature descriptor
keypoint description
0.712023
A Unified BEV Model for Joint Learning of 3D Local Features and Overlap Estimation · ICRA 2023
Robotics › Robot navigation and mapping › localization › odometry
LiDAR-inertial odometry
0.412020
LiDAR Inertial Odometry Aided Robust LiDAR Localization System in Changing City Scenes · ICRA 2020
Robotics › Autonomous driving
perception
0.412020
LiDAR Inertial Odometry Aided Robust LiDAR Localization System in Changing City Scenes · ICRA 2020
Computer vision › 3D vision
visual localization
0.412020
DA4AD: End-to-End Deep Attention-Based Visual Localization for Autonomous Driving · ECCV (28) 2020
Robotics › Robot navigation and mapping › localization
multi-sensor localization
0.312018
Robust and Precise Vehicle Localization Based on Multi-Sensor Fusion in Diverse City Scenes · ICRA 2018
Robotics › Robot navigation and mapping › SLAM
loop closure detection
0.212023
A Unified BEV Model for Joint Learning of 3D Local Features and Overlap Estimation · ICRA 2023
Robotics › Robot navigation and mapping
SLAM
0.212023
A Unified BEV Model for Joint Learning of 3D Local Features and Overlap Estimation · ICRA 2023
Robotics › Robot navigation and mapping › robot mapping › map management
map update
0.112020
LiDAR Inertial Odometry Aided Robust LiDAR Localization System in Changing City Scenes · ICRA 2020
Image and video processing › image restoration › inverse problem
missing data interpolation
0.112010
Non-local scan consolidation for 3D urban scenes · ACM Trans. Graph. 2010
Image and video processing › image restoration › image denoising › patch-based denoising
non-local means
0.112010
Non-local scan consolidation for 3D urban scenes · ACM Trans. Graph. 2010
Geometric modeling and processing › point cloud processing
point cloud consolidation
0.112010
Non-local scan consolidation for 3D urban scenes · ACM Trans. Graph. 2010
Geometric modeling and processing
surface reconstruction
0.112010
Non-local scan consolidation for 3D urban scenes · ACM Trans. Graph. 2010

Methods — techniques the papers use, named apart from their topics

keypoint detection · 1.0transformer · 0.8point-based navigation · 0.8sparse UNet · 0.7cross-attention · 0.7bird's-eye-view representation · 0.7occupancy grid · 0.4end-to-end learning · 0.4deep attention · 0.4MAP estimation · 0.4surface clustering · 0.1non-local filtering · 0.1in-plane and off-plane denoising · 0.1
YearPublicationVenuePosition
2024 RoadPainter: Points Are Ideal Navigators for Topology TransformER
Zhongxing Ma, Yongkun Wen, Weixin Lu, Guowei Wan
ECCV (12)5
2024 EgoVM: Achieving Precise Ego-Localization using Lightweight Vectorized Maps
abstract
Accurate and reliable ego-localization is critical for autonomous driving. In this paper, we present EgoVM, an end-to-end localization network that achieves comparable localization accuracy to prior state-of-the-art methods, but uses lightweight vectorized maps instead of heavy point-based maps. To begin with, we extract BEV features from online multi-view images and LiDAR point cloud. Then, we employ a set of learnable semantic embeddings to encode the semantic types of map elements and supervise them with semantic segmentation, to make their feature representation consistent with BEV features. After that, we feed map queries, composed of learnable semantic embeddings and coordinates of map elements, into a transformer decoder to perform cross-modality matching with BEV features. Finally, we adopt a robust histogram-based pose solver to estimate the optimal pose by searching exhaustively over candidate poses. We comprehensively validate the effectiveness of our method using both the nuScenes dataset and a newly collected dataset. The experimental results show that our method achieves centimeter-level localization accuracy, and outperforms existing methods using vectorized maps by a large margin. Furthermore, our model has been extensively tested in a large fleet of autonomous vehicles under various challenging urban scenes.
Yuzhe He, Xiaofei Rui, Chengying Cai, Guowei Wan
IROS5
2023 A Unified BEV Model for Joint Learning of 3D Local Features and Overlap Estimation
abstract
Pairwise point cloud registration is a critical task for many applications, which heavily depends on finding correct correspondences from the two point clouds. However, the low overlap between input point clouds causes the registration to fail easily, leading to mistaken overlapping and mismatched correspondences, especially in scenes where non-overlapping regions contain similar structures. In this paper, we present a unified bird's-eye view (BEV) model for jointly learning of 3D local features and overlap estimation to fulfill pairwise registration and loop closure. Feature description is performed by a sparse UNet-like network based on BEV representation, and 3D keypoints are extracted by a detection head for 2D locations, and a regression head for heights. For overlap detection, a cross-attention module is applied for interacting contextual information of input point clouds, followed by a classification head to estimate the overlapping region. We evaluate our unified model extensively on the KITTI dataset and Apollo-SouthBay dataset. The experiments demonstrate that our method significantly outperforms existing methods on overlap estimation, especially in scenes with small overlaps. It also achieves top registration performance on both datasets in terms of translation and rotation errors.
Lin Li 0091, Wendong Ding, Yongkun Wen, Yufei Liang, Yong Liu 0007, Guowei Wan
ICRA6
2022 ACDet: Attentive Cross-view Fusion for LiDAR-based 3D Object Detection
abstract
Recent works on 3D object detection take the range image as input, which have achieved comparable performance with bird's eye view (BEV) based methods. Compared to BEV, range view provides dense and compact observations which allows for more popular feature encoders. To leverage complementary information of range view and BEV, we present ACDet - a novel single-stage multi-view fusion method. Rather than fusing point-level features from range view and BEV at early stage, the key contribution is that we introduce an attentive cross-view fusion module based on transformer to fuse higher level features, and further adopt a supervised foreground mask learned from BEV features to enhance the fused features. Notably, a geometric-attention kernel is proposed to enhance features extracted from range image. Finally, we design an anchor-free detection head with optimized label assignment strategy, and its performance exceeds the existing anchor-based and anchor-free 3D detection heads by a large margin. We evaluate our ACDet model extensively on the KITTI dataset and Waymo Open Dataset (WOD). ACDet outperforms most of singlestage models on KITTI dataset in terms of multi-class 3D and BEV mean average precision. ACDet also outperforms both range-view and multi-view fusion methods on WOD.
Jiaolong Xu, Guowei Wan
3DV4
2020 DA4AD: End-to-End Deep Attention-Based Visual Localization for Autonomous Driving
Guowei Wan, Shenhua Hou, Xiaofei Rui, Shiyu Song
ECCV (28)2
2020 LiDAR Inertial Odometry Aided Robust LiDAR Localization System in Changing City Scenes
abstract
Environmental fluctuations pose crucial challenges to a localization system in autonomous driving. We present a robust LiDAR localization system that maintains its kinematic estimation in changing urban scenarios by using a dead reckoning solution implemented through a LiDAR inertial odometry. Our localization framework jointly uses information from complementary modalities such as global matching and LiDAR inertial odometry to achieve accurate and smooth localization estimation. To improve the performance of the LiDAR odometry, we incorporate inertial and LiDAR intensity cues into an occupancy grid based LiDAR odometry to enhance frame-to-frame motion and matching estimation. Multi-resolution occupancy grid is implemented yielding a coarse-to-fine approach to balance the odometry's precision and computational requirement. To fuse both the odometry and global matching results, we formulate a MAP estimation problem in a pose graph fusion framework that can be efficiently solved. An effective environmental change detection method is proposed that allows us to know exactly when and what portion of the map requires an update. We comprehensively validate the effectiveness of the proposed approaches using both the Apollo-SouthBay dataset and our internal dataset. The results confirm that our efforts lead to a more robust and accurate localization system, especially in dynamically changing urban scenarios.
Wendong Ding, Shenhua Hou, Guowei Wan, Shiyu Song
ICRA4
2019 L3-Net: Towards Learning Based LiDAR Localization for Autonomous Driving
abstract
We present L3-Net - a novel learning-based LiDAR localization system that achieves centimeter-level localization accuracy, comparable to prior state-of-the-art systems with hand-crafted pipelines. Rather than relying on these hand-crafted modules, we innovatively implement the use of various deep neural network structures to establish a learning-based approach. L3-Net learns local descriptors specifically optimized for matching in different real-world driving scenarios. 3D convolutions over a cost volume built in the solution space significantly boosts the localization accuracy. RNNs are demonstrated to be effective in modeling the vehicle's dynamics, yielding better temporal smoothness and accuracy. We comprehensively validate the effectiveness of our approach using freshly collected datasets. Multiple trials of repetitive data collection over the same road and areas make our dataset ideal for testing localization systems. The SunnyvaleBigLoop sequences, with a year's time interval between the collected mapping and testing data, made it quite challenging, but the low localization error of our method in these datasets demonstrates its maturity for real industrial implementation.
Weixin Lu, Guowei Wan, Shenhua Hou, Shiyu Song
CVPR3
2019 DeepVCP: An End-to-End Deep Neural Network for Point Cloud Registration
abstract
We present DeepVCP - a novel end-to-end learning-based 3D point cloud registration framework that achieves comparable registration accuracy to prior state-of-the-art geometric methods. Different from other keypoint based methods where a RANSAC procedure is usually needed, we implement the use of various deep neural network structures to establish an end-to-end trainable network. Our keypoint detector is trained through this end-to-end structure and enables the system to avoid the interference of dynamic objects, leverages the help of sufficiently salient features on stationary objects, and as a result, achieves high robustness. Rather than searching the corresponding points among existing points, the key contribution is that we innovatively generate them based on learned matching probabilities among a group of candidates, which can boost the registration accuracy. We comprehensively validate the effectiveness of our approach using both the KITTI dataset and the Apollo-SouthBay dataset. Results demonstrate that our method achieves comparable registration accuracy and runtime efficiency to the state-of-the-art geometry-based methods, but with higher robustness to inaccurate initial poses. Detailed ablation and visualization analysis are included to further illustrate the behavior and insights of our network. The low registration error and high robustness of our method make it attractive to the substantial applications relying on the point cloud registration task.
Weixin Lu, Guowei Wan, Xiangyu Fu, Pengfei Yuan, Shiyu Song
ICCV2
2018 Robust and Precise Vehicle Localization Based on Multi-Sensor Fusion in Diverse City Scenes
abstract
We present a robust and precise localization system that achieves centimeter-level localization accuracy in disparate city scenes. Our system adaptively uses information from complementary sensors such as GNSS, LiDAR, and IMU to achieve high localization accuracy and resilience in challenging scenes, such as urban downtown, highways, and tunnels. Rather than relying only on LiDAR intensity or 3D geometry, we make innovative use of LiDAR intensity and altitude cues to significantly improve localization system accuracy and robustness. Our GNSS RTK module utilizes the help of the multi-sensor fusion framework and achieves a better ambiguity resolution success rate. An error-state Kalman filter is applied to fuse the localization measurements from different sources with novel uncertainty estimation. We validate, in detail, the effectiveness of our approaches, achieving 5-10cm RMS accuracy and outperforming previous state-of-the-art systems. Importantly, our system, while deployed in a large autonomous driving fleet, made our vehicles fully autonomous in crowded city streets despite road construction that occurred from time to time. A dataset including more than 60 km real traffic driving in various urban roads is used to comprehensively test our system.
Guowei Wan, Renlan Cai, Shiyu Song
ICRA1
2016 Mobility Fitting using 4D RANSAC
abstract
Abstract Capturing the dynamics of articulated models is becoming increasingly important. Dynamics, better than geometry, encode the functional information of articulated objects such as humans, robots and mechanics. Acquired dynamic data is noisy, sparse, and temporarily incoherent. The latter property is especially prominent for analysis of dynamics. Thus, processing scanned dynamic data is typically an ill‐posed problem. We present an algorithm that robustly computes the joints representing the dynamics of a scanned articulated object. Our key idea is to by‐pass the reconstruction of the underlying surface geometry and directly solve for motion joints. To cope with the often‐times extremely incoherent scans, we propose a space‐time fitting‐and‐voting approach in the spirit of RANSAC. We assume a restricted set of articulated motions defined by a set of joints which we fit to the 4D dynamic data and measure their fitting quality. Thus, we repeatedly select random subsets and fit with joints, searching for an optimal candidate set of mobility parameters. Without having to reconstruct surfaces as intermediate means, our approach gains the advantage of being robust and efficient. Results demonstrate the ability to reconstruct dynamics of various articulated objects consisting of a wide range of complex and compound motions.
Hao Li 0015, Guowei Wan, Honghua Li, Andrei Sharf, Kai Xu 0004, Baoquan Chen
Comput. Graph. Forum2
2012 Grammar-based 3D facade segmentation and reconstruction
Guowei Wan, Andrei Sharf
Comput. Graph.1
2012 Sorting unorganized photo sets for urban reconstruction
Guowei Wan, Noah Snavely, Daniel Cohen-Or, Baoquan Chen, Sikun Li
Graph. Model.1
2010 Non-local scan consolidation for 3D urban scenes
abstract
Recent advances in scanning technologies, in particular devices that extract depth through active sensing, allow fast scanning of urban scenes. Such rapid acquisition incurs imperfections: large regions remain missing, significant variation in sampling density is common, and the data is often corrupted with noise and outliers. However, buildings often exhibit large scale repetitions and self-similarities. Detecting, extracting, and utilizing such large scale repetitions provide powerful means to consolidate the imperfect data. Our key observation is that the same geometry, when scanned multiple times over reoccurrences of instances, allow application of a simple yet effective non-local filtering. The multiplicity of the geometry is fused together and projected to abase-geometrydefined by clustering corresponding surfaces. Denoising is applied by separating the process into off-plane and in-plane phases. We show that the consolidation of the reoccurrences provides robust denoising and allow reliable completion of missing parts. We present evaluation results of the algorithm on several LiDAR scans of buildings of varying complexity and styles.
Andrei Sharf, Guowei Wan, Yangyan Li, Niloy J. Mitra, Daniel Cohen-Or, Baoquan Chen
ACM Trans. Graph.3
2009 An incremental extremely random forest classifier for online learning and tracking
abstract
Decision trees have been widely used for online learning classification. Many approaches usually need large data stream to finish decision trees induction, as show notable limitations (even fail) with small data stream. In fact, there exist many real instances with small data stream. In the paper, we propose a novel incremental extremely random forest algorithm, dealing with online learning classification with small streaming data. In our method, arriving examples are stored at the leaf nodes and used to determine when to split the leaf nodes combined with Gini index, so the trees can be expanded efficiently with a few examples. Our algorithm has been applied to solve both online learning and video object tracking problems, and the results on UCI datasets and challenging video sequences demonstrate its effectiveness and robustness.
Aiping Wang, Guowei Wan, Zhi-Quan Cheng, Sikun Li
ICIP2