Guofeng Mei

dblp:197/7875 · DBLP profile ↗
← Back
24ranked-venue papers
11as first author
22since 2021 · last 2026
0000-0002-0494-5031ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 6 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 9 first-author · 16 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion
abstract
Denoising Diffusion Probabilistic Models (DDPMs) have shown success in robust 3D object detection tasks. Existing methods often rely on the score matching from 3D boxes or pre-trained diffusion priors. However, they typically require multi-step iterations in inference, which limits efficiency. To address this, we propose a Robust single-stage fully Sparse 3D object Detection Network with a Detachable Latent Framework (DLF) of DDPMs, named RSDNet. Specifically, RSDNet learns the denoising process in latent feature spaces through lightweight denoising networks like multi-level denoising autoencoders (DAEs). This enables RSDNet to effectively understand scene distributions under multi-level perturbations, achieving robust and reliable detection. Meanwhile, we reformulate the noising and denoising mechanisms of DDPMs, enabling DLF to construct multi-type and multi-level noise samples and targets, enhancing RSDNet robustness to multiple perturbations. Furthermore, a semantic-geometric conditional guidance is introduced to perceive the object boundaries and shapes, alleviating the center feature missing problem in sparse representations, enabling RSDNet to perform in a fully sparse detection pipeline. Moreover, the detachable denoising network design of DLF enables RSDNet to perform single-step detection in inference, further enhancing detection efficiency. Extensive experiments on public benchmarks show that RSDNet can outperform existing methods, achieving state-of-the-art detection.
Wentao Qu, Guofeng Mei, Jing Wang 0201, Yujiao Wu, Xiaoshui Huang, Liang Xiao 0001
AAAI2
2026 Masked Clustering Prediction for Unsupervised Point Cloud Pre-training
abstract
Vision transformers (ViTs) have recently been widely applied to 3D point cloud understanding, with masked autoencoding as the predominant pre-training paradigm. However, the challenge of learning dense and informative semantic features from point clouds via standard ViTs remains underexplored. We propose MaskClu, a novel unsupervised pre-training method for ViTs on 3D point clouds that integrates masked point modeling with clustering-based learning. MaskClu is designed to reconstruct both cluster assignments and cluster centers from masked point clouds, thus encouraging the model to capture dense semantic information. Additionally, we introduce a global contrastive learning mechanism that enhances instance-level feature learning by contrasting different masked views of the same point cloud. By jointly optimizing these complementary objectives, i.e., dense semantic reconstruction, and instance-level contrastive learning. MaskClu enables ViTs to learn richer and more semantically meaningful representations from 3D point clouds. We validate the effectiveness of MaskClu via multiple 3D tasks, including part segmentation, semantic segmentation, object detection, and classification, setting new competitive results.
Bin Ren 0005, Xiaoshui Huang, Mengyuan Liu 0001, Hong Liu 0008, Fabio Poiesi, Nicu Sebe, Guofeng Mei
AAAI7
2025 Vocabulary-Free 3D Instance Segmentation with Vision-Language Assistant
abstract
Most recent 3D instance segmentation methods are open vocabulary, offering a greater flexibility than closedvocabulary methods. Yet, they are limited to reasoning within a specific set of concepts, i.e., the vocabulary, prompted by the user at test time. In essence, these models cannot reason in an open-ended fashion, i.e., answering “List the objects in the scene.” We introduce the first method to address 3D instance segmentation in a setting that is void of any vocabulary prior, namely a vocabularyfree setting. We leverage a large vision-language assistant and an open-vocabulary$2 D$instance segmenter to discover and ground semantic categories on the posed images. To form 3D instance masks, we first partition the input point cloud into dense superpoints, which are then merged into 3D instance masks. We propose a novel superpoint merging strategy via spectral clustering, accounting for both mask coherence and semantic coherence that are estimated from the 2D object instance masks. We evaluate our method using ScanNet200 and Replica, outperforming existing methods in both vocabulary-free and open-vocabulary settings.
Guofeng Mei, Luigi Riz, Yiming Wang 0002, Fabio Poiesi
3DV1
2025 Fully-Geometric Cross-Attention for Point Cloud Registration
abstract
Point cloud registration approaches often fail when the overlap between point clouds is low due to noisy point correspondences. This work introduces a novel cross-attention mechanism tailored for Transformer-based architectures that tackles this problem, by fusing information from coordinates and features at the super-point level between point clouds. This formulation has remained unexplored primarily because it must guarantee rotation and translation invariance since point clouds reside in different and independent reference frames. We integrate the Gromov-Wasserstein distance into the cross-attention formulation to jointly compute distances between points across different point clouds and account for their geometric structure. By doing so, points from two distinct point clouds can attend to each other under arbitrary rigid transformations. At the point level, we also devise a self-attention mechanism that aggregates the local geometric structure information into point features for fine matching. Our formulation boosts the number of inlier correspondences, thereby yielding more precise registration results compared to state-of-the-art approaches. We have conducted an extensive evaluation on 3DMatch, 3DLoMatch, KITTI, and 3DCSR datasets. Project page: https://github.com/twowwj/FLAT.
Weijie Wang 0002, Guofeng Mei, Jian Zhang 0002, Nicu Sebe, Bruno Lepri, Fabio Poiesi
3DV2
2025 PerLA: Perceptive 3D Language Assistant
abstract
Enabling Large Language Models (LLMs) to understand the 3D physical world is an emerging yet challenging research direction. Current strategies for processing point clouds typically downsample the scene or divide it into smaller parts for separate analysis. However, both approaches risk losing key local details or global contextual information. In this paper, we introduce PerLA, a 3D language assistant designed to be more perceptive to both details and context, making visual representations more informative for the LLM. PerLA captures high-resolution (local) details in parallel from different point cloud areas and integrates them with (global) context obtained from a lower-resolution whole point cloud. We present a novel algorithm that preserves point cloud locality through the Hilbert curve and effectively aggregates local-to-global information via cross-attention and a graph neural network. Lastly, we introduce a novel loss for local representation consensus to promote training stability. PerLA outperforms state-of-the-art 3D language assistants, with gains of up to +1.34 CiDEr on ScanQA for question answering, and +4.22 on ScanRefer and +3.88 on Nr3D for dense captioning. Project page: https://gfmei.github.io/PerLA
Guofeng Mei, Luigi Riz, Yujiao Wu, Fabio Poiesi, Yiming Wang 0002
CVPR1
2025 PointGAC: Geometric-Aware Codebook for Masked Point Cloud Modeling
abstract
Most masked point cloud modeling (MPM) methods follow a regression paradigm to reconstruct the coordinate or feature of masked regions. However, they tend to over-constrain the model to learn the details of the masked region, resulting in failure to capture generalized features. To address this limitation, we propose \textbf{\textit{PointGAC}}, a novel clustering-based MPM method that aims to align the feature distribution of masked regions. Specially, it features an online codebook-guided teacher-student framework. Firstly, it presents a geometry-aware partitioning strategy to extract initial patches. Then, the teacher model updates a codebook via online k-means based on features extracted from the complete patches. This procedure facilitates codebook vectors to become cluster centers. Afterward, we assigns the unmasked features to their corresponding cluster centers, and the student model aligns the assignment for the reconstructed masked features. This strategy focuses on identifying the cluster centers to which the masked features belong, enabling the model to learn more generalized feature representations. Benefiting from a proposed codebook maintenance mechanism, codebook vectors are actively updated, which further increases the efficiency of semantic feature learning. Experiments validate the effectiveness of the proposed method on various downstream tasks. Code is available at https://github.com/LAB123-tech/PointGAC
Abiao Li, Chenlei Lv, Yuming Fang 0001, Yifan Zuo 0001, Jian Zhang 0002, Guofeng Mei
ICCV6
2025 Free-form language-based robotic reasoning and grasping
abstract
Performing robotic grasping from a cluttered bin based on human instructions is a challenging task, as it requires understanding both the nuances of free-form language and the spatial relationships between objects. Vision-Language Models (VLMs) trained on web-scale data, such as GPT-4o, have demonstrated remarkable reasoning capabilities across both text and images. But can they truly be used for this task in a zero-shot setting? And what are their limitations? In this paper, we explore these research questions via the free-form language-based robotic grasping task and propose a novel method, FreeGrasp, leveraging the pre-trained VLMs’ world knowledge to reason about human instructions and object spatial arrangements. Our method detects all objects as keypoints and uses these keypoints to annotate marks on images, aiming to facilitate GPT-4o’s zero-shot spatial reasoning. This allows our method to determine whether a requested object is directly graspable or if other objects must be grasped and removed first. Since no existing dataset is specifically designed for this task, we introduce a synthetic dataset FreeGraspData by extending the MetaGraspNetV2 dataset with human-annotated instructions and ground-truth grasping sequences. We conduct extensive analyses with both FreeGraspData and real-world validation with a gripper-equipped robotic arm, demonstrating state-of-the-art performance in grasp reasoning and execution. Project website: https://tev-fbk.github.io/FreeGrasp/.
Runyu Jiao, Alice Fasoli, Francesco Giuliari, Matteo Bortolon, Sergio Povoli, Guofeng Mei, Yiming Wang 0002, Fabio Poiesi
IROS6
2025 The Consistent Transmission Laws Under the Consistent Intervention Policy
abstract
COVID-19 has been spreading globally from 2019 to 2022, prompting many countries to establish consistent, timely, and strict intervention policies during this period. However, the existing analysis of central epidemiological parameters (CPPs) falls short in exploring the consistency of transmission laws implied by these policies. To resolve this limitation, we propose a consistent transmission law inference framework. This framework first provides a coarse-to-fine detection method to extract outbreaks and devises an autonomous inference sliding window (ASW) algorithm to estimate the infected-case-dependent CPPs of these outbreaks. Then, we reveal consistent transmission laws through a differential analysis of the estimated CPPs. Focusing on the transmission of COVID-19 in China, a typical country with a consistent intervention policy, our model reveals remarkably consistent outbreak laws. In short, the policy controls the effective growth rate almost to zero (to a 1E-03 scale) within three days since the outbreak started. Specifically: 1) the policy was very effective at the beginning of the outbreak, leading to the exposure rate hitting the bottom on one day between the second and eleventh days and then keeping it at a plateau; 2) under large-scale nucleic acid testing and contact tracking, the incubation rate decreased and reached a plateau mainly on the third day; and 3) the recovery rate fluctuated extremely little on a 1E-03 scale, showing no significant breakthrough in COVID-19 treatment that helped the policy work. Our method can be readily adapted to other countries and SEIR epidemic.
Juan Liu 0006, Guofeng Mei, Yuanqing Xia
IEEE Trans. Comput. Soc. Syst.2
2024 Bringing Masked Autoencoders Explicit Contrastive Properties for Point Cloud Self-supervised Learning
Bin Ren 0005, Guofeng Mei, Danda Pani Paudel, Weijie Wang 0002, Yawei Li 0001, Mengyuan Liu 0001, Rita Cucchiara, Luc Van Gool, Nicu Sebe
ACCV (7)2
2024 Geometrically-Driven Aggregation for Zero-Shot 3D Point Cloud Understanding
abstract
Zero-shot 3D point cloud understanding can be achieved via 2D Vision-Language Models (VLMs). Existing strategies directly map VLM representations from 2D pixels of rendered or captured views to 3D points, overlooking the inherent and expressible point cloud geometric structure. Geometrically similar or close regions can be exploited for bolstering point cloud understanding as they are likely to share semantic information. To this end, we introduce the first training-free aggregation technique that leverages the point cloud's 3D geometric structure to improve the quality of the transferred VLM representations. Our approach operates iteratively, performing local-to-global aggregation based on geometric and semantic point-level reasoning. We benchmark our approach on three downstream tasks, including classification, part segmentation, and semantic segmentation, with a variety of datasets representing both synthetic/real-world, and indoor/outdoor scenarios. Our approach achieves new state-of-the-art results in all benchmarks. Code and dataset are available at https://luigiriz.github.io/geoze-website/
Guofeng Mei, Luigi Riz, Yiming Wang 0002, Fabio Poiesi
CVPR1
2024 Point Cloud Pre-Training with Diffusion Models
abstract
Pre-training a model and then fine-tuning it on down-stream tasks has demonstrated significant success in the 2D image and NLP domains. However, due to the unordered and non-uniform density characteristics of point clouds, it is non-trivial to explore the prior knowledge of point clouds and pre-train a point cloud backbone. In this paper, we propose a novel pre-training method called Point cloud Diffusion pre-training (PointDif). We consider the point cloud pre-training task as a conditional point-to-point gen-eration problem and introduce a conditional point genera-tor. This generator aggregates the features extracted by the backbone and employs them as the condition to guide the point-to-point recovery from the noisy point cloud, thereby assisting the backbone in capturing both local and global geometric priors as well as the global point density distri-bution of the object. We also present a recurrent uniform sampling optimization strategy, which enables the model to uniformly recover from various noise levels and learn from balanced supervision. Our PointDif achieves substan-tial improvement across various real-world datasets for di-verse downstream tasks such as classification, segmentation and detection. Specifically, PointDif attains 70.0% mIoU on S3DIS Area 5 for the segmentation task and achieves an average improvement of 2.4% on ScanObjectNN for the classification task compared to TAP. Furthermore, our pre-training framework can be flexibly applied to diverse point cloud backbones and bring considerable gains. Code is available at https://github.com/zhengxiaozx/PointDif.
Xiaoshui Huang, Guofeng Mei, Yuenan Hou, Zhaoyang Lyu, Bo Dai 0002, Wanli Ouyang, Yongshun Gong
CVPR3
2024 GSTran: Joint Geometric and Semantic Coherence for Point Cloud Segmentation
Abiao Li, Chenlei Lv, Guofeng Mei, Yifan Zuo 0001, Jian Zhang 0002, Yuming Fang 0001
ICPR (18)3
2024 Unsupervised Point Cloud Representation Learning by Clustering and Neural Rendering
abstract
Abstract Data augmentation has contributed to the rapid advancement of unsupervised learning on 3D point clouds. However, we argue that data augmentation is not ideal, as it requires a careful application-dependent selection of the types of augmentations to be performed, thus potentially biasing the information learned by the network during self-training. Moreover, several unsupervised methods only focus on uni-modal information, thus potentially introducing challenges in the case of sparse and textureless point clouds. To address these issues, we propose an augmentation-free unsupervised approach for point clouds, named CluRender, to learn transferable point-level features by leveraging uni-modal information for soft clustering and cross-modal information for neural rendering. Soft clustering enables self-training through a pseudo-label prediction task, where the affiliation of points to their clusters is used as a proxy under the constraint that these pseudo-labels divide the point cloud into approximate equal partitions. This allows us to formulate a clustering loss to minimize the standard cross-entropy between pseudo and predicted labels. Neural rendering generates photorealistic renderings from various viewpoints to transfer photometric cues from 2D images to the features. The consistency between rendered and real images is then measured to form a fitting loss, combined with the cross-entropy loss to self-train networks. Experiments on downstream applications, including 3D object detection, semantic segmentation, classification, part segmentation, and few-shot learning, demonstrate the effectiveness of our framework in outperforming state-of-the-art techniques.
Guofeng Mei, Cristiano Saltori, Elisa Ricci 0001, Nicu Sebe, Qiang Wu 0001, Jian Zhang 0002, Fabio Poiesi
Int. J. Comput. Vis.1
2023 Unsupervised Deep Probabilistic Approach for Partial Point Cloud Registration
abstract
Deep point cloud registration methods face challenges to partial overlaps and rely on labeled data. To address these issues, we propose UDPReg, an unsupervised deep probabilistic registration framework for point clouds with partial overlaps. Specifically, we first adopt a network to learn posterior probability distributions of Gaussian mixture models (GMMs) from point clouds. To handle partial point cloud registration, we apply the Sinkhorn algorithm to predict the distribution-level correspondences under the constraint of the mixing weights of GMMs. To enable unsupervised learning, we design three distribution consistency-based losses: self-consistency, cross-consistency, and local contrastive. The self-consistency loss is formulated by encouraging GMMs in Euclidean and feature spaces to share identical posterior distributions. The cross-consistency loss derives from the fact that the points of two partially overlapping point clouds belonging to the same clusters share the cluster centroids. The cross-consistency loss allows the network to flexibly learn a transformation-invariant posterior distribution of two aligned point clouds. The local contrastive loss facilitates the network to extract discriminative local features. Our UDPReg achieves competitive performance on the 3DMatch/3DLoMatch and ModelNet/ModelLoNet benchmarks.
Guofeng Mei, Hao Tang 0005, Xiaoshui Huang, Weijie Wang 0002, Juan Liu 0006, Jian Zhang 0002, Luc Van Gool, Qiang Wu 0001
CVPR1
2023 Graph Matching Optimization Network for Point Cloud Registration
abstract
Point Cloud Registration is a fundamental and challenging problem in 3D computer vision. Recent works often utilize geometric structure features in downsampled points (patches) to seek correspondences, then propagate these sparse patch correspondences to the dense level in the corresponding patches' neighborhood. However, they neglect the explicit global scale rigid constraint at the dense level point matching. We claim that the explicit isometry-preserving constraint in the dense level on a global scale is also important for improving feature representation in the training stage. To this end, we propose a Graph Matching Optimization based Network (GMONet for short), which utilizes the graph-matching optimizer to explicitly exert the isometry preserving constraints in the point feature training to improve the point feature representation. Specifically, we exploit a partial graph-matching optimizer to enhance the super point (i.e., down-sampled key points) features and a full graph-matching optimizer to improve the dense level point features in the overlap region. Meanwhile, we leverage the inexact proximal point method and the mini-batch sampling technique to accelerate these two graph-matching optimizers. Given high discriminative point features in the evaluation stage, we utilize the RANSAC approach to estimate the transformation between the scanned pairs. The proposed method has been evaluated on the 3DMatch/3DLoMatch and the KITTI datasets. The experimental results show that our method performs competitively compared to state-of-the-art baselines.
Qianliang Wu, Yaqi Shen, Haobo Jiang, Guofeng Mei, Yaqing Ding 0001, Lei Luo 0001, Jin Xie 0001, Jian Yang 0003
IROS4
2023 Overlap-guided Gaussian Mixture Models for Point Cloud Registration
abstract
Probabilistic 3D point cloud registration methods have shown competitive performance in overcoming noise, outliers, and density variations. However, registering point cloud pairs in the case of partial overlap is still a challenge. This paper proposes a novel overlap-guided probabilistic registration approach that computes the optimal transformation from matched Gaussian Mixture Model (GMM) parameters. We reformulate the registration problem as the problem of aligning two Gaussian mixtures such that a statistical discrepancy measure between the two corresponding mixtures is minimized. We introduce a Transformer-based detection module to detect overlapping regions, and represent the input point clouds using GMMs by guiding their alignment through overlap scores computed by this detection module. Experiments show that our method achieves superior registration accuracy and efficiency than state-of-the-art methods when handling point clouds with partial overlap and different densities on synthetic and real-world datasets. https://github.com/gfmei/ogmm
Guofeng Mei, Fabio Poiesi, Cristiano Saltori, Jian Zhang 0002, Elisa Ricci 0001, Nicu Sebe
WACV1
2023 Cross-source point cloud registration: Challenges, progress and prospects
Xiaoshui Huang, Guofeng Mei, Jian Zhang 0002
Neurocomputing2
2022 Data Augmentation-free Unsupervised Learning for 3D Point Cloud Understanding
Guofeng Mei, Cristiano Saltori, Fabio Poiesi, Jian Zhang 0002, Elisa Ricci 0001, Nicu Sebe, Qiang Wu 0001
BMVC1
2022 Unsupervised Point Cloud Pre-Training Via Contrasting and Clustering
abstract
The annotation for large-scale point clouds is still time-consuming and unavailable for many complex real-world tasks. Point cloud pre-training is a promising direction to auto-extract features without labeled data. Therefore, this paper proposes a general unsupervised approach, named ConClu for point cloud pre-training by jointly performing contrasting and clustering. Specifically, the contrasting is formulated by maximizing the similarity feature vectors produced by encoders fed with two augmentations of the same point cloud. The clustering simultaneously clusters the data while enforcing consistency between cluster assignments produced different augmentations. Experimental evaluations on downstream applications outperform state-of-the-art techniques, which demonstrates the effectiveness of our framework.
Guofeng Mei, Xiaoshui Huang, Juan Liu 0006, Jian Zhang 0002, Qiang Wu 0001
ICIP1
2022 Partial Point Cloud Registration Via Soft Segmentation
abstract
Most existing correspondence-free registration methods suffer from performance degradation in partial overlapped point clouds. To solve the partial overlapped point cloud registration, this paper proposes, SegReg, a soft Segmentation-based correspondence-free Registration approach. Specifically, we first softly segment both source and target point clouds into a discrete number of geometric partitions, respectively. Then registration is achieved through iteratively using the IC-LK algorithm to minimize the distance between the feature descriptors of the corresponded partitions. Extensive experiments on synthetic synthetic dataset ModelNet40 and real dataset 7Scene show that the proposed method achieves state-of-the-art performance.
Guofeng Mei, Xiaoshui Huang, Jian Zhang 0002, Qiang Wu 0001
ICIP1
2022 Overlap-Guided Coarse-to-Fine Correspondence Prediction for Point Cloud Registration
abstract
Establishing reliable correspondences between a pair of point clouds is essential for registration with partial overlaps. However, existing correspondence estimation works usually struggle to distinguish the points in overlap and non-overlap regions. This paper thus proposes an Overlap-guided Coarse-to-Fine Network, named OCFNet, which first establishes correspondences at a coarse level and then refines them at a point level. Specifically, at the coarse level, our model first aggregates two point clouds into smaller sets of super-points with associated features and overlap scores, followed by establishing coarse-level correspondences between the two sets of super-points under the guidance of overlap scores. On the fine stage, a decoder recovers the raw points while jointly learning the associated features and overlap scores. Coarse-level proposals are then expanded to patches, and point-level correspondences are sequentially refined from the corresponding patches. We conducted comprehensive experiments on 3DMatch, 3DLoMatch, and KITTI benchmarks to show the effectiveness of the proposed method. [code]
Guofeng Mei, Xiaoshui Huang, Jian Zhang 0002, Qiang Wu 0001
ICME1
2022 Robust real-world point cloud registration by inlier detection
Xiaoshui Huang, Yangfu Wang, Sheng Li 0020, Guofeng Mei, Zongyi Xu, Yucheng Wang 0003, Jian Zhang 0002, Mohammed Bennamoun
Comput. Vis. Image Underst.4
2020 Feature-Metric Registration: A Fast Semi-Supervised Approach for Robust Point Cloud Registration Without Correspondences
abstract
We present a fast feature-metric point cloud registration framework, which enforces the optimisation of registration by minimising a feature-metric projection error without correspondences. The advantage of the feature-metric projection error is robust to noise, outliers and density difference in contrast to the geometric projection error. Besides, minimising the feature-metric projection error does not need to search the correspondences so that the optimisation speed is fast. The principle behind the proposed method is that the feature difference is smallest if point clouds are aligned very well. We train the proposed method in a semi-supervised or unsupervised approach, which requires limited or no registration label data. Experiments demonstrate our method obtains higher accuracy and robustness than the state-of-the-art methods. Besides, experimental results show that the proposed method can handle significant noise and density difference, and solve both same-source and cross-source point cloud registration.
Xiaoshui Huang, Guofeng Mei, Jian Zhang 0002
CVPR2
2018 Compressive-Sensing-Based Structure Identification for Multilayer Networks
abstract
The coexistence of multiple types of interactions within social, technological, and biological networks has motivated the study of the multilayer nature of real-world networks. Meanwhile, identifying network structures from dynamical observations is an essential issue pervading over the current research on complex networks. This paper addresses the problem of structure identification for multilayer networks, which is an important topic but involves a challenging inverse problem. To clearly reveal the formalism, the simplest two-layer network model is considered and a new approach to identifying the structure of one layer is proposed. Specifically, if the interested layer is sparsely connected and the node behaviors of the other layer are observable at a few time points, then a theoretical framework is established based on compressive sensing and regularization. Some numerical examples illustrate the effectiveness of the identification scheme, its requirement of a relatively small number of observations, as well as its robustness against small noise. It is noteworthy that the framework can be straightforwardly extended to multilayer networks, thus applicable to a variety of real-world complex systems.
Guofeng Mei, Xiaoqun Wu, Yingfei Wang, Mi Hu, Jun-An Lu, Guanrong Chen
IEEE Trans. Cybern.1