Enxu Li

dblp:285/4934 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2023
0009-0000-4113-8864ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
3D vision · 36% Segmentation and scene understanding · 29% Autonomous driving · 20%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › point cloud segmentation › LiDAR segmentation
LiDAR panoptic segmentation
1.732023
CPSeg: Cluster-free Panoptic Segmentation of 3D LiDAR Point Clouds · ICRA 2023
SMAC-Seg: LiDAR Panoptic Segmentation via Sparse Multi-directional Attention Clustering · ICRA 2022
GP-S3Net: Graph-based Panoptic Sparse Semantic Segmentation Network · ICCV 2021
Computer vision › Segmentation and scene understanding
instance segmentation
1.222023
CPSeg: Cluster-free Panoptic Segmentation of 3D LiDAR Point Clouds · ICRA 2023
SMAC-Seg: LiDAR Panoptic Segmentation via Sparse Multi-directional Attention Clustering · ICRA 2022
Robotics › Autonomous driving › perception
LiDAR perception
1.222023
CPSeg: Cluster-free Panoptic Segmentation of 3D LiDAR Point Clouds · ICRA 2023
SMAC-Seg: LiDAR Panoptic Segmentation via Sparse Multi-directional Attention Clustering · ICRA 2022
Computer vision › 3D vision › point cloud segmentation
point cloud semantic segmentation
1.222023
MemorySeg: Online LiDAR Semantic Segmentation with a Latent Memory · ICCV 2023
(AF)2-S3Net: Attentive Feature Fusion With Adaptive Feature Selection for Sparse Semantic Segmentation Network · CVPR 2021
Robotics › Autonomous driving
perception
0.722023
(AF)2-S3Net: Attentive Feature Fusion With Adaptive Feature Selection for Sparse Semantic Segmentation Network · CVPR 2021
MemorySeg: Online LiDAR Semantic Segmentation with a Latent Memory · ICCV 2023
Computer vision › 3D vision
3d scene understanding
0.712023
MemorySeg: Online LiDAR Semantic Segmentation with a Latent Memory · ICCV 2023
Machine learning › Deep learning architectures and training › memory-augmented neural networks
memory network
0.712023
MemorySeg: Online LiDAR Semantic Segmentation with a Latent Memory · ICCV 2023
Computer vision › Video understanding and tracking
temporal modeling
0.712023
MemorySeg: Online LiDAR Semantic Segmentation with a Latent Memory · ICCV 2023
Computer vision › Segmentation and scene understanding
attention-based clustering
0.612022
SMAC-Seg: LiDAR Panoptic Segmentation via Sparse Multi-directional Attention Clustering · ICRA 2022
Computer vision › Segmentation and scene understanding
panoptic segmentation
0.512021
GP-S3Net: Graph-based Panoptic Sparse Semantic Segmentation Network · ICCV 2021
Computer vision › Segmentation and scene understanding
semantic segmentation
0.512021
(AF)2-S3Net: Attentive Feature Fusion With Adaptive Feature Selection for Sparse Semantic Segmentation Network · CVPR 2021
Machine learning › Graph learning › graph neural network
graph convolutional network
0.112021
GP-S3Net: Graph-based Panoptic Sparse Semantic Segmentation Network · ICCV 2021

Methods — techniques the papers use, named apart from their topics

temporal regularization · 0.7sparse 3d representation · 0.7pairwise embedding comparison · 0.7latent memory · 0.7dual-decoder network · 0.7depth completion · 0.7sparse multi-directional attention · 0.6centroid-aware repel loss · 0.6attentive feature fusion · 0.5adaptive feature selection · 0.5
YearPublicationVenuePosition
2023 MemorySeg: Online LiDAR Semantic Segmentation with a Latent Memory
abstract
Semantic segmentation of LiDAR point clouds has been widely studied in recent years, with most existing methods focusing on tackling this task using a single scan of the environment, However, leveraging the temporal stream of observations can provide very rich contextual information on regions of the scene with poor visibility (e.g., occlusions) or sparse observations (e.g., at long range), and can help reduce redundant computation frame after frame. In this paper, we tackle the challenge of exploiting the information from the past frames to improve the predictions of the current frame in an online fashion. To address this challenge, we propose a novel framework for semantic segmentation of a temporal sequence of LiDAR point clouds that utilizes a memory network to store, update and retrieve past information. Our framework also includes a novel regularizer that penalizes prediction variations in the neighborhood of the point cloud. Prior works have attempted to incorporate memory in range view representations for semantic segmentation, but these methods fail to handle occlusions and the range view representation of the scene changes drastically as agents nearby move. Our proposed framework overcomes these limitations by building a sparse 3D latent representation of the surroundings. We evaluate our method on SemanticKITTI, nuScenes, and PandaSet. Our experiments demonstrate the effectiveness of the proposed framework compared to the state-of-the-art. For more information, visit the project website: https://waabi.ai/research/memoryseg.
Enxu Li, Sergio Casas 0002, Raquel Urtasun
ICCV1
2023 CPSeg: Cluster-free Panoptic Segmentation of 3D LiDAR Point Clouds
abstract
A fast and accurate panoptic segmentation system for LiDAR point clouds is crucial for autonomous driving vehicles to understand the surrounding objects and scenes. Existing approaches usually rely on proposals or clustering to segment foreground instances. As a result, they struggle to achieve real-time performance. In this paper, we propose a novel real-time end-to-end panoptic segmentation network for LiDAR point clouds, called CPSeg. In particular, CPSeg comprises a shared encoder, a dual-decoder, and a cluster-free instance segmentation head, which is able to dynamically pillarize foreground points according to the learned embedding. Then, it acquires instance labels by finding connected pillars with a pairwise embedding comparison. Thus, the conventional proposal-based or clustering-based instance segmentation is transformed into a binary segmentation problem on the pairwise embedding comparison matrix. To help the network regress instance embedding, a fast and deterministic depth completion algorithm is proposed to calculate the surface normal of each point cloud in real-time. The proposed method is benchmarked on two large-scale autonomous driving datasets: SemanticKITTI and nuScenes. Notably, extensive experimental results show that CPSeg achieves state-of-the-art results among real-time approaches on both datasets.
Enxu Li, Ryan Razani, Yixuan Xu 0004
ICRA1
2022 SMAC-Seg: LiDAR Panoptic Segmentation via Sparse Multi-directional Attention Clustering
abstract
Panoptic segmentation aims to address semantic and instance segmentation simultaneously in a unified framework. However, an efficient solution of panoptic segmentation in applications like autonomous driving is still an open research problem. In this work, we propose a novel LiDAR-based panoptic system, called SMAC-Seg. We present a learnable sparse multi-directional attention clustering to segment multi-scale foreground instances. SMAC-Seg is a real-time clustering-based approach, which removes the complex proposal network to segment instances. Most existing clustering-based methods use the difference of the predicted and ground truth center offset as the only loss to supervise the instance centroid regression. However, this loss function only considers the centroid of the current object, but its relative position with respect to the neighbouring objects is not considered when learning to cluster. Thus, we propose to use a novel centroid-aware repel loss as an additional term to effectively supervise the network in order to differentiate each object cluster with its neighbours. Our experimental results show that SMAC-Seg achieves state-of-the-art performance among all real-time deployable networks on both large-scale public SemanticKITTI and nuScenes panoptic segmentation datasets.
Enxu Li, Ryan Razani, Yixuan Xu 0004
ICRA1
2021 (AF)2-S3Net: Attentive Feature Fusion With Adaptive Feature Selection for Sparse Semantic Segmentation Network
abstract
Autonomous robotic systems and self driving cars rely on accurate perception of their surroundings as the safety of the passengers and pedestrians is the top priority. Semantic segmentation is one of the essential components of road scene perception that provides semantic information of the surrounding environment. Recently, several methods have been introduced for 3D LiDAR semantic segmentation. While they can lead to improved performance, they are either afflicted by high computational complexity, therefore are inefficient, or they lack fine details of smaller object instances. To alleviate these problems, we propose (AF)2-S3Net, an end-to-end encoder-decoder CNN network for 3D LiDAR semantic segmentation. We present a novel multibranch attentive feature fusion module in the encoder and a unique adaptive feature selection module with feature map re-weighting in the decoder. Our (AF)2-S3Net fuses the voxel-based learning and point-based learning methods into a unified framework to effectively process the potentially large 3D scene. Our experimental results show that the proposed method outperforms the state-of-the-art approaches on the large-scale nuScenes-lidarseg and SemanticKITTI benchmark, ranking 1ston both competitive public leaderboard competitions upon publication.
Ryan Razani, Ehsan Taghavi, Enxu Li
CVPR4
2021 GP-S3Net: Graph-based Panoptic Sparse Semantic Segmentation Network
abstract
Panoptic segmentation as an integrated task of both static environmental understanding and dynamic object identification, has recently begun to receive broad research interest. In this paper, we propose a new computationally efficient LiDAR based panoptic segmentation framework, called GP-S3Net. GP-S3Net is a proposal-free approach in which no object proposals are needed to identify the objects in contrast to conventional two-stage panoptic systems, where a detection network is incorporated for capturing instance information. Our new design consists of a novel instance-level network to process the semantic results by constructing a graph convolutional network to identify objects (foreground), which later on are fused with the back-ground classes. Through the fine-grained clusters of the foreground objects from the semantic segmentation back-bone, over-segmentation priors are generated and subsequently processed by 3D sparse convolution to embed each cluster. Each cluster is treated as a node in the graph and its corresponding embedding is used as its node feature. Then a GCNN predicts whether edges exist between each cluster pair. We utilize the instance label to generate ground truth edge labels for each constructed graph in order to supervise the learning. Extensive experiments demonstrate that GP-S3Net outperforms the current state-of-the-art approaches, by a significant margin across available datasets such as, nuScenes and SemanticPOSS, ranking 1ston the competitive public SemanticKITTI leaderboard upon publication.
Ryan Razani, Enxu Li, Ehsan Taghavi
ICCV3