Xianzhi Li 0001

dblp:126/1233-1 · DBLP profile ↗
← Back
37ranked-venue papers
4as first author
30since 2021 · last 2026
0000-0001-6835-5607ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 4 first-author · 23 since 2021Artificial intelligence and machine learning · 19 · 14 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Rethinking Point Cloud Representation Learning for Freeing Transformer to Perceive Local
abstract
Transformers are widely utilized in the point cloud domain. However, existing methods tend to overburden Transformer with the dual task of local geometric perception and global feature extraction, limiting its ability to capture highlevel semantic knowledge. To address this issue, we present Representation Decoder (R-Decoder), a novel representation extraction module compatible with various point cloud Transformer methods, enabling the Transformer to focus on its excellent local perception. The R-Decoder iteratively extracts multiple global features from tokens generated by Transformer, refining them to construct an overall representation of point cloud. To ensure full adaptation of the R-Decoder to the knowledge of pre-trained Transformers, we design a cross-modal representation alignment task that leverages multimodal knowledge to specifically pre-train the R-Decoder. As a post-processing module, the R-Decoder seamlessly integrates with Transformers, while decoupling local perception and global representation. This design allows the Transformer to focus on the semantic encoding role for point tokens. Extensive experiments show that our RDecoder significantly boosts the capabilities of 3D representation learning in various point cloud Transformer methods. Notably, it achieves impressive classification accuracies of 95.1% on the ScanObjectNN dataset and 95.3% on the ModelNet40 dataset. Moreover, our method obtains new SOTA on all benchmarks of few-shot and zero-shot classification, while enhancing the multimodal task capabilities of pre-trained Transformers. Code and weights are available athttps://github.com/TangYuan96/RDecoder.
Yunlong Yu 0002, Xianzhi Li 0001, Rui Wang 0077, Jinfeng Xu 0002, Qiao Yu 0002, Yixue Hao, Long Hu, Min Chen 0003
IEEE Trans. Multim.3
2026 DTSNet: Dynamic Transformer Slimming for Efficient Vision Recognition
abstract
Transformer-based models have recently adopted increasingly complex structure (e.g., deeper or wider stacked network) to promote the representation learning capabilities of vision recognition. However, progressively deeper or wider stacked network cause the expensive computation cost, which hinders their effective deployment in resource-constrained edge clouds or end devices. In this paper, we propose DTSNet, a dynamic transformer slimming model, which scales vision transformers (ViTs) down across layers from both of the model depth and input width. This is the first time to explore the joint reduction of input tokens and model parameters for ViTs under maintaining performance. Specifically, DTSNet adopts a diversity-enhanced weight sharing module to reduce network parameters, where the weight knowledge of multiple adjacent blocks is effectively integrated into one block. Furthermore, DTSNet designs a unified and massively scalable token pruning mechanism that dynamically discarding less important tokens with a model-driven manner, by introducing a series of discriminant parameters, which is a simple change to the common architecture of vision transformers. Extensive experiments are conducted to verify that DTSNet is able to yield high efficacy in compressing parameter space and accelerating model inference. DTSNet-T/-S/-B on ImageNet achieves 3.0M/11.1M/42.9M parameters and 0.8/2.9/13.7 GFLOPs, where number of parameters are reduced by 48%$\sim$51% and inference speed are improved by 1.3$\times \sim 1.5\times$. Experiments results on semantic segmentation and object detection dataset further demonstrate the potential of DTSNet on complex dense prediction tasks. Code will be available upon publication.
Wenjing Xiao, Xianzhi Li 0001, Long Hu, Yixue Hao, Min Chen 0003
IEEE Trans. Multim.2
2025 More Text, Less Point: Towards 3D Data-Efficient Point-Language Understanding
abstract
Enabling Large Language Models (LLMs) to comprehend the 3D physical world remains a significant challenge. Due to the lack of large-scale 3D-text pair datasets, the success of LLMs has yet to be replicated in 3D understanding. In this paper, we rethink this issue and propose a new task: 3D Data-Efficient Point-Language Understanding. The goal is to enable LLMs to achieve robust 3D object understanding with minimal 3D point cloud and text data pairs. To address this task, we introduce GreenPLM, which leverages more text data to compensate for the lack of 3D data. First, inspired by using CLIP to align images and text, we utilize a pre-trained point cloud-text encoder to map the 3D point cloud space to the text space. This mapping leaves us to seamlessly connect the text space with LLMs. Once the point-text-LLM connection is established, we further enhance text-LLM alignment by expanding the intermediate text space, thereby reducing the reliance on 3D point cloud data. Specifically, we generate 6M free-text descriptions of 3D objects, and design a three-stage training strategy to help LLMs better explore the intrinsic connections between different modalities. To achieve efficient modality alignment, we design a zero-parameter cross-attention module for token pooling. Extensive experimental results show that GreenPLM requires only 12% of the 3D training data used by existing state-of-the-art models to achieve superior 3D understanding. Remarkably, GreenPLM also achieves competitive performance using text-only data.
Xu Han 0016, Xianzhi Li 0001, Qiao Yu 0002, Jinfeng Xu 0002, Yixue Hao, Long Hu, Min Chen 0003
AAAI3
2025 SASep: Saliency-Aware Structured Separation of Geometry and Feature for Open Set Learning on Point Clouds
abstract
Recent advancements in deep learning have greatly enhanced 3D object recognition, but most models are limited to closed-set scenarios, unable to handle unknown samples in real-world applications. Open-set recognition (OSR) addresses this limitation by enabling models to both classify known classes and identify novel classes. However, current OSR methods rely on global features to differentiate known and unknown classes, treating the entire object uniformly and overlooking the varying semantic importance of its different parts. To address this gap, we propose Salience-Aware Structured Separation (SASep), which includes (i) a tunable semantic decomposition (TSD) module to semantically decompose objects into important and unimportant parts, (ii) a geometric synthesis strategy (GSS) to generate pseudo-unknown objects by combining these unimportant parts, and (iii) a synth-aided margin separation (SMS) module to enhance feature-level separation by expanding the feature distributions between classes. Together, these components improve both geometric and feature representations, enhancing the model’s ability to effectively distinguish known and unknown classes. Experimental results show that SASep achieves superior performance in 3D OSR, outperforming existing state-of-the-art methods. The codes are available at https://github.com/JinfengX/SASep.
Jinfeng Xu 0002, Xianzhi Li 0001, Xu Han 0016, Qiao Yu 0002, Yixue Hao, Long Hu, Min Chen 0003
CVPR2
2025 MoST: Efficient Monarch Sparse Tuning for 3D Representation Learning
abstract
We introduce Monarch Sparse Tuning (MoST), the first reparameterization-based parameter-efficient fine-tuning (PEFT) method tailored for 3D representation learning. Unlike existing adapter-based and prompt-tuning 3D PEFT methods, MoST introduces no additional inference overhead and is compatible with many 3D representation learning backbones. At its core, we present a new family of structured matrices for 3D point clouds, Point Monarch, which can capture local geometric features of irregular points while offering high expressiveness. MoST reparameterizes the dense update weight matrices as our sparse Point Monarch matrices, significantly reducing parameters while retaining strong performance. Experiments on various backbones show that MoST is simple, effective, and highly generalizable. It captures local features in point clouds, achieving state-of-the-art results on multiple benchmarks, e.g., 97.5% acc. on ScanOb-jectNN(PB_50_RS) and 96.2% on ModelNet40 classification, while it can also combine with other matrix decompositions (e.g., Low-rank, Kronecker) to further reduce parameters.
Xu Han 0016, Jinfeng Xu 0002, Xianzhi Li 0001
CVPR4
2025 Fancy123: One Image to High-Quality 3D Mesh Generation via Plug-and-Play Deformation
abstract
Generating 3D meshes from a single image is an important but ill-posed task. Existing methods mainly adopt 2D multiview diffusion models to generate intermediate multiview images, and use the Large Reconstruction Model (LRM) to create the final meshes. However, the multiview images exhibit local inconsistencies, and the meshes often lack fidelity to the input image or look blurry. We propose Fancy123, featuring two enhancement modules and an unprojection operation to address the above three issues, respectively. The appearance enhancement module deforms the 2D multiview images to realign misaligned pixels for better multiview consistency. The fidelity enhancement module deforms the 3D mesh to match the input image. The unprojection of the input image and deformed multiview images onto LRM’s generated mesh ensures high clarity, discarding LRM’s predicted blurry-looking mesh colors. Extensive qualitative and quantitative experiments verify Fancy123’s SoTA performance with significant improvement. Also, the two enhancement modules are plug-and-play and work at inference time, allowing seamless integration into various existing single-image-to-3D methods. Project page: https://github.com/YuQiao0303/Fancy123.
Qiao Yu 0002, Xianzhi Li 0001, Xu Han 0016, Long Hu, Yixue Hao, Min Chen 0003
CVPR2
2025 Geometrically-plausible and Semantically-consistent Generation of Indoor Panoramas
abstract
We present PanoGPS, a new approach for generating indoor panoramas, capable of following a room layout sketch and description text to generate panoramas with geometrically-plausible room layouts and semantically-consistent contents. Overall, our idea is to inform the generative model collectively with the two inputs, to inspire it to implicitly generate and refine a unified latent code for panorama generation. Specifically, we first propose using the semantic layout map to encode the room geometry to condition the generative model with plausible geometric information. Second, we establish a pipeline to integrate features from the two inputs and generate a single unified latent code that describes both the room layout and room contents of the target panorama. Third, we progressively refine the latent code to produce a coarse panorama to enforce seamless left-right border connections, while iteratively upsizing the coarse panorama with details. Our method outperforms existing approaches by generating high-quality panoramas with geometrically plausible structures and semantically meaningful content. Code is available at: https://github.com/wumengyangok/PanoGPS
Zhiliang Zeng, Mengyang Wu, Xianzhi Li 0001, Wenzhao Gao, Shaohui Jiao, Chi-Wing Fu
ICME3
2025 PointDreamer: Zero-Shot 3D Textured Mesh Reconstruction From Colored Point Cloud
abstract
Faithfully reconstructing textured meshes is crucial for many applications. Compared to text or image modalities, leveraging 3D colored point clouds as input (colored-PC-to-mesh) offers inherent advantages in comprehensively and precisely replicating the target object's 360$^{\circ }$∘ characteristics. While most existing colored-PC-to-mesh methods suffer from blurry textures or require hard-to-acquire 3D training data, we propose PointDreamer, a novel framework that harnesses 2D diffusion prior for superior texture quality. Crucially, unlike prior 2D-diffusion-for-3D works driven by text or image inputs, PointDreamer successfully adapts 2D diffusion models to 3D point cloud data by a novel project-inpaint-unproject pipeline. Specifically, it first projects the point cloud into sparse 2D images and then performs diffusion-based inpainting. After that, diverging from most existing 3D reconstruction or generation approaches that predict texture in 3D/UV space thus often yielding blurry texture, PointDreamer achieves high-quality texture by directly unprojecting the inpainted 2D images to the 3D mesh. Furthermore, we identify for the first time a typical kind of unprojection artifact appearing in occlusion borders, which is common in other multiview-image-to-3D pipelines but less-explored. To address this, we propose a novel solution named the Non-Border-First (NBF) unprojection strategy. Extensive qualitative and quantitative experiments on various synthetic and real-scanned datasets demonstrate that PointDreamer, though zero-shot, exhibits SoTA performance (30% improvement on LPIPS score from 0.118 to 0.068), and is robust to noisy, sparse, or even incomplete input data.
Qiao Yu 0002, Xianzhi Li 0001, Xu Han 0016, Jinfeng Xu 0002, Long Hu, Min Chen 0003
IEEE Trans. Vis. Comput. Graph.2
2025 JIMR: Joint Semantic and Geometry Learning for Point Scene Instance Mesh Reconstruction
abstract
Point scene instance mesh reconstruction is a challenging task since it requires both scene-level instance segmentation and instance-level mesh reconstruction from partial observations simultaneously. Previous works either adopt a detection backbone or a segmentation one, and then directly employ a mesh reconstruction network to produce complete meshes from incomplete instance point clouds. To further boost the mesh reconstruction quality with both local details and global smoothness, in this work, we propose JIMR, a joint framework with two cascaded stages for semantic and geometry understanding. In the first stage, we propose to perform both instance segmentation and object detection simultaneously. By making both tasks promote each other, this design facilitates subsequent mesh reconstruction by providing more precisely-segmented instance points and better alignment benefiting from predicted complete bounding boxes. In the second stage, we propose a complete-then-reconstruct procedure, where the completion module explicitly disentangles completion from reconstruction, and enables the usage of pre-trained weights of existing powerful completion and reconstruction networks. Moreover, we propose a comprehensive confidence score to filter proposals considering the quality of instance segmentation, bounding box detection, semantic classification, and mesh reconstruction at the same time. Experiments show that our proposed JIMR outperforms state-of-the-art methods regarding instance reconstruction qualitatively and quantitatively.
Qiao Yu 0002, Xianzhi Li 0001, Jinfeng Xu 0002, Long Hu, Yixue Hao, Min Chen 0003
IEEE Trans. Vis. Comput. Graph.2
2024 PDF: A Probability-Driven Framework for Open World 3D Point Cloud Semantic Segmentation
abstract
Existing point cloud semantic segmentation networks cannot identify unknown classes and update their knowledge, due to a closed-set and static perspective of the real world, which would induce the intelligent agent to make bad decisions. To address this problem, we propose a Probability-Driven Framework (PDF)11Code available at: https://github.com/JinfengX/PointCloudPDF. for open world semantic segmentation that includes (i) a lightweight U-decoder branch to identify unknown classes by estimating the uncertainties, (ii) a flexible pseudo-labeling scheme to supply geometry features along with probability distribution features of unknown classes by generating pseudo labels, and (iii) an incremental knowledge distillation strategy to incorporate novel classes into the existing knowledge base gradually. Our framework enables the model to behave like human beings, which could recognize unknown objects and incrementally learn them with the corresponding knowledge. Experimental results on the S3DIS and ScanNetv2 datasets demonstrate that the proposed PDF outperforms other methods by a large margin in both important tasks of open world semantic segmentation.
Jinfeng Xu 0002, Xianzhi Li 0001, Yixue Hao, Long Hu, Min Chen 0003
CVPR3
2024 Geo-Encoder: A Chunk-Argument Bi-Encoder Framework for Chinese Geographic Re-Ranking
abstract
Yong Cao, Ruixue Ding, Boli Chen, Xianzhi Li, Min Chen, Daniel Hershcovich, Pengjun Xie, Fei Huang. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yong Cao 0001, Ruixue Ding, Boli Chen, Xianzhi Li 0001, Min Chen 0003, Daniel Hershcovich, Pengjun Xie, Fei Huang 0002
EACL (1)4
2024 Mamba3D: Enhancing Local Features for 3D Point Cloud Analysis via State Space Model
abstract
Existing Transformer-based models for point cloud analysis suffer from quadratic complexity, leading to compromised point cloud resolution and information loss. In contrast, the newly proposed Mamba model, based on state space models (SSM), outperforms Transformer in multiple areas with only linear complexity. However, the straightforward adoption of Mamba does not achieve satisfactory performance on point cloud tasks. In this work, we present Mamba3D, a state space model tailored for point cloud learning to enhance local feature extraction, achieving superior performance, high efficiency, and scalability potential. Specifically, we propose a simple yet effective Local Norm Pooling (LNP) block to extract local geometric features. Additionally, to obtain better global features, we introduce a bidirectional SSM (bi-SSM) with both a token forward SSM and a novel backward SSM that operates on the feature channel. Extensive experimental results show that Mamba3D surpasses Transformer-based counterparts and concurrent works in multiple tasks, with or without pre-training. Notably, Mamba3D achieves multiple SoTA, including an overall accuracy of 92.6% (train from scratch) on the ScanObjectNN and 95.1% (with single-modal pre-training) on the ModelNet40 classification task, with only linear complexity. Our code and weights are available at https://github.com/xhanxu/Mamba3D.
Xu Han 0016, Zhaoxuan Wang, Xianzhi Li 0001
ACM Multimedia4
2024 MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors
abstract
Large 2D vision-language models (2D-LLMs) have gained significant attention by bridging Large Language Models (LLMs) with images using a simple projector. Inspired by their success, large 3D point cloud-language models (3D-LLMs) also integrate point clouds into LLMs. However, directly aligning point clouds with LLM requires expensive training costs, typically in hundreds of GPU-hours on A100, which hinders the development of 3D-LLMs. In this paper, we introduce MiniGPT-3D, an efficient and powerful 3D-LLM that achieves multiple SOTA results while training for only 27 hours on one RTX 3090. Specifically, we propose to align 3D point clouds with LLMs using 2D priors from 2D-LLMs, which can leverage the similarity between 2D and 3D visual information. We introduce a novel four-stage training strategy for modality alignment in a cascaded way, and a mixture of query experts module to adaptively aggregate features with high efficiency. Moreover, we utilize parameter-efficient fine-tuning methods LoRA and Norm fine-tuning, resulting in only 47.8M learnable parameters, which is up to 260x fewer than existing methods. Extensive experiments show that MiniGPT-3D achieves SOTA on 3D object classification and captioning tasks, with significantly cheaper training costs. Notably, MiniGPT-3D gains an 8.12 increase on GPT-4 evaluation score for the challenging object captioning task compared to ShapeLLM-13B, while the latter costs 160 total GPU-hours on 8 A800. We are the first to explore the efficient 3D-LLM, offering new insights to the community. Code and weights are available at https://github.com/TangYuan96/MiniGPT-3D.
Xu Han 0016, Xianzhi Li 0001, Qiao Yu 0002, Yixue Hao, Long Hu, Min Chen 0003
ACM Multimedia3
2024 Efficient Crowd Counting via Dual Knowledge Distillation
abstract
Most researchers focus on designing accurate crowd counting models with heavy parameters and computations but ignore the resource burden during the model deployment. A real-world scenario demands an efficient counting model with low-latency and high-performance. Knowledge distillation provides an elegant way to transfer knowledge from a complicated teacher model to a compact student model while maintaining accuracy. However, the student model receives the wrong guidance with the supervision of the teacher model due to the inaccurate information understood by the teacher in some cases. In this paper, we propose a dual-knowledge distillation (DKD) framework, which aims to reduce the side effects of the teacher model and transfer hierarchical knowledge to obtain a more efficient counting model. First, the student model is initialized with global information transferred by the teacher model via adaptive perspectives. Then, the self-knowledge distillation forces the student model to learn the knowledge by itself, based on intermediate feature maps and target map. Specifically, the optimal transport distance is utilized to measure the difference of feature maps between the teacher and the student to perform the distribution alignment of the counting area. Extensive experiments are conducted on four challenging datasets, demonstrating the superiority of DKD. When there are only approximately 6% of the parameters and computations from the original models, the student model achieves a faster and more accurate counting performance as the teacher model even surpasses it.
Rui Wang 0077, Yixue Hao, Long Hu, Xianzhi Li 0001, Min Chen 0003, Yiming Miao, Iztok Humar
IEEE Trans. Image Process.4
2024 Point-LGMask: Local and Global Contexts Embedding for Point Cloud Pre-Training With Multi-Ratio Masking
abstract
Self-supervised learning has achieved great success in both natural language processing and 2D vision, where masked modeling is a quite popular pre-training scheme. However, extending masking to 3D point cloud understanding that combines local and global features poses a new challenge. In our work, we present Point-LGMask, a novel method to embed both local and global contexts with multi-ratio masking, which is quite effective for self-supervised feature learning of point clouds but is unfortunately ignored by existing pre-training works. Specifically, to avoid fitting to a fixed masking ratio, we first propose multi-ratio masking, which prompts the encoder to fully explore representative features thanks to tasks of different difficulties. Next, to encourage the embedding of both local and global features, we formulate a compound loss, which consists of (i) a global representation contrastive loss to encourage the cluster assignments of the masked point clouds to be consistent to that of the completed input, and (ii) a local point cloud prediction loss to encourage accurate prediction of masked points. Equipped with our Point-LGMask, we show that our learned representations transfer well to various downstream tasks, including few-shot classification, shape classification, object part segmentation, as well as real-world scene-based 3D object detection and 3D semantic segmentation. Particularly, our model largely advances existing pre-training methods on the difficult few-shot classification task using the real-captured ScanObjectNN dataset by surpassing over 4% to the second-best method. Also, our Point-LGMask achieves 0.4%$AP_{25}$and 0.8%$AP_{50}$gains on 3D object detection task over the second-best method. 0.4% mAcc and 0.5% mIoU. Codes have been released athttps://github.com/TangYuan96/Point-LGMask.
Xianzhi Li 0001, Jinfeng Xu 0002, Qiao Yu 0002, Long Hu, Yixue Hao, Min Chen 0003
IEEE Trans. Multim.2
2023 CasFusionNet: A Cascaded Network for Point Cloud Semantic Scene Completion by Dense Feature Fusion
abstract
Semantic scene completion (SSC) aims to complete a partial 3D scene and predict its semantics simultaneously. Most existing works adopt the voxel representations, thus suffering from the growth of memory and computation cost as the voxel resolution increases. Though a few works attempt to solve SSC from the perspective of 3D point clouds, they have not fully exploited the correlation and complementarity between the two tasks of scene completion and semantic segmentation. In our work, we present CasFusionNet, a novel cascaded network for point cloud semantic scene completion by dense feature fusion. Specifically, we design (i) a global completion module (GCM) to produce an upsampled and completed but coarse point set, (ii) a semantic segmentation module (SSM) to predict the per-point semantic labels of the completed points generated by GCM, and (iii) a local refinement module (LRM) to further refine the coarse completed points and the associated labels from a local perspective. We organize the above three modules via dense feature fusion in each level, and cascade a total of four levels, where we also employ feature fusion between each level for sufficient information usage. Both quantitative and qualitative results on our compiled two point-based datasets validate the effectiveness and superiority of our CasFusionNet compared to state-of-the-art methods in terms of both scene completion and semantic segmentation. The codes and datasets are available at: https://github.com/JinfengX/CasFusionNet.
Jinfeng Xu 0002, Xianzhi Li 0001, Qiao Yu 0002, Yixue Hao, Long Hu, Min Chen 0003
AAAI2
2023 On Improving Boundary Quality of Instance Segmentation in Cluttered and Chaotic Scenarios
abstract
Instance segmentation is a long-standing task for supporting robotic bin picking. However, objects of diverse classes can be closely packed with occlusions in cluttered and chaotic scenes, hence, even recent methods could have difficulty in locating clear and precise boundaries to distinguish nearby objects. In this work, we aim to improve the boundary quality of the instance masks for robust and precise instance segmentation in these challenging scenarios. Technical-wise, we first formulate an IoU-based Boundary-aware Mask head (IBM head) for predicting the instance-level mask, boundary, and their corresponding IoU scores. With this core module, we then follow the coarse-to-fine strategy and design our pipeline with two stages: an 1IoUNet to learn localization-based objectness cue and a hierarchical mask refiner to produce sharper and cleaner boundaries. We deploy the IBM head throughout the framework. Extensive experimental results on three grasping benchmarks manifest that our method attains the best instance segmentation performance, compared with the state-of-the-art approaches. Practically, we conduct real-world picking tests to show that with the objectness and boundary IoU scores as guidance, we are able to filter invalid (occluded) instances and select high-fidelity (exposed) instances for grasping.
Biqi Yang, Xianzhi Li 0001, Yun-Hui Liu 0001, Chi-Wing Fu, Pheng-Ann Heng
ICRA3
2023 Joint-MAE: 2D-3D Joint Masked Autoencoders for 3D Point Cloud Pre-training
abstract
Masked Autoencoders (MAE) have shown promising performance in self-supervised learning for both 2D and 3D computer vision. However, existing MAE-style methods can only learn from the data of a single modality, i.e., either images or point clouds, which neglect the implicit semantic and geometric correlation between 2D and 3D. In this paper, we explore how the 2D modality can benefit 3D masked autoencoding, and propose Joint-MAE, a 2D-3D joint MAE framework for self-supervised 3D point cloud pre-training. Joint-MAE randomly masks an input 3D point cloud and its projected 2D images, and then reconstructs the masked information of the two modalities. For better cross-modal interaction, we construct our JointMAE by two hierarchical 2D-3D embedding modules, a joint encoder, and a joint decoder with modal-shared and model-specific decoders. On top of this, we further introduce two cross-modal strategies to boost the 3D representation learning, which are local-aligned attention mechanisms for 2D-3D semantic cues, and a cross-reconstruction loss for 2D-3D geometric constraints. By our pre-training paradigm, Joint-MAE achieves superior performance on multiple downstream tasks, e.g., 92.4% accuracy for linear SVM on ModelNet40 and 86.07% accuracy on the hardest split of ScanObjectNN.
Renrui Zhang, Longtian Qiu, Xianzhi Li 0001, Pheng-Ann Heng
IJCAI4
2023 V2Depth: Monocular Depth Estimation via Feature-Level Virtual-View Simulation and Refinement
abstract
Due to the lack of spatial cues giving merely a single image, many monocular depth estimation methods have been developed to leverage stereo or multi-view images to learn the spatial information of a scene in a self-supervised manner. However, these methods have limited performance gain since they are not able to exploit sufficient 3D geometry cues during inference, where only monocular images are available. In this work, we present V2Depth, a novel coarse-to-fine framework with Virtual View feature simulation for supervised monocular Depth estimation. Specifically, we first design a virtual-view feature simulator by leveraging the technique of novel view synthesis and contrastive learning to generate virtual view feature maps. In this way, we explicitly provide representative spatial geometry for subsequent depth estimation in both the training and inference stages. Then we introduce a 3DVA-Refiner to iteratively optimize the predicted depth map. During the optimization process, 3D-aware virtual attention is developed to capture the global spatial-context correlations to maintain the feature consistency of different views and estimation integrity of the 3D scene such as objects with occlusion relationships. Decisive improvements over state-of-the-art approaches on three benchmark datasets across all metrics demonstrate the superiority of our method.
Zizhang Wu, Zhuozheng Li, Zhi-Gang Fan, Yunzhe Wu, Jian Pu, Xianzhi Li 0001
ACM Multimedia6
2023 Prototypical Variational Autoencoder for 3D Few-shot Object Detection
abstract
Few-Shot 3D Point Cloud Object Detection (FS3D) is a challenging task, aiming to detect 3D objects of novel classes using only limited annotated samples for training. Considering that the detection performance highly relies on the quality of the latent features, we design a VAE-based prototype learning scheme, named prototypical VAE (P-VAE), to learn a probabilistic latent space for enhancing the diversity and distinctiveness of the sampled features. The network encodes a multi-center GMM-like posterior, in which each distribution centers at a prototype. For regularization, P-VAE incorporates a reconstruction task to preserve geometric information. To adopt P-VAE for the detection framework, we formulate Geometric-informative Prototypical VAE (GP-VAE) to handle varying geometric components and Class-specific Prototypical VAE (CP-VAE) to handle varying object categories. In the first stage, we harness GP-VAE to aid feature extraction from the input scene. In the second stage, we cluster the geometric-informative features into per-instance features and use CP-VAE to refine each instance feature with category-level guidance. Experimental results show the top performance of our approach over the state of the arts on two FS3D benchmarks. Quantitative ablations and qualitative prototype analysis further demonstrate that our probabilistic modeling can significantly boost prototype learning for FS3D.
Weiliang Tang, Biqi Yang, Xianzhi Li 0001, Yun-Hui Liu 0001, Pheng-Ann Heng, Chi-Wing Fu
NeurIPS3
2023 Point Set Self-Embedding
abstract
This work presents an innovative method for point set self-embedding, that encodes the structural information of a dense point set into its sparser version in a visual but imperceptible form. The self-embedded point set can function as the ordinary downsampled one and be visualized efficiently on mobile devices. Particularly, we can leverage the self-embedded information to fully restore the original point set for detailed analysis on remote servers. This task is challenging, since both the self-embedded point set and the restored point set should resemble the original one. To achieve a learnable self-embedding scheme, we design a novel framework with two jointly-trained networks: one to encode the input point set into its self-embedded sparse point set and the other to leverage the embedded information for inverting the original point set back. Further, we develop a pair of up-shuffle and down-shuffle units in the two networks, and formulate loss terms to encourage the shape similarity and point distribution in the results. Extensive qualitative and quantitative results demonstrate the effectiveness of our method on both synthetic and real-scanned datasets. The source code and trained models will be publicly available at https://github.com/liruihui/Self-Embedding.
Ruihui Li, Xianzhi Li 0001, Tien-Tsin Wong, Chi-Wing Fu
IEEE Trans. Vis. Comput. Graph.2
2022 SEPL-Net: A Semantics-Enhanced Pseudo Labeling Network for Semi-Supervised Image Analysis
abstract
As the mainstream solution for semi-supervised learning (SSL), pseudo-labeling-based approaches have achieved re-markable success. However, an obvious drawback of existing methods is that the valuable semantic relationships among categories are often ignored, thus leading to suboptimal encoded embeddings. To address this, we present a novel Semantics-Enhanced Pseudo Labeling Network, called SEPL-Net, for image analysis in a semi-supervised manner. SEPL-Net explores the prior knowledge of visual similarity between different classes to improve the quality of pseudo label decision making. Particularly, we encode semantic labels combined with the one-hot label to jointly train our network by exploiting their disagreement. To alleviate the difficulty of labeling unlabeled images due to the introduction of semantic labels, we further design different classifiers with differentiated strong augmentation modes to enable cooperative pseudo labeling. Extensive experimental results show that our SEPL-Net outperforms existing SSL methods with the averaged 1.84% accuracy improvement on image classification task. Code is available at https://github.com/sweetvicky/SEPLNet.git.
Wenjing Xiao, Kai Hwang 0001, Min Chen 0003, Xianzhi Li 0001
ICME4
2022 Towards Robust Part-aware Instance Segmentation for Industrial Bin Picking
abstract
Industrial bin picking is a challenging task that requires accurate and robust segmentation of individual object instances. Particularly, industrial objects can have irregular shapes, that is, thin and concave, whereas in bin-picking scenarios, objects are often closely packed with strong occlusion. To address these challenges, we formulate a novel part-aware instance segmentation pipeline. The key idea is to decompose industrial objects into correlated approximate convex parts and enhance the object-level segmentation with part-level segmentation. We design a part-aware network to predict part masks and part-to-part offsets, followed by a part aggregation module to assemble the recognized parts into instances. To guide the network learning, we also propose an automatic label decoupling scheme to generate ground-truth part-level labels from instance-level labels. Finally, we contribute the first instance segmentation dataset, which contains a variety of industrial objects that are thin and have non-trivial shapes. Extensive experimental results on various industrial objects demonstrate that our method can achieve the best segmentation results compared with the state-of-the-art approaches.
Yidan Feng, Biqi Yang, Xianzhi Li 0001, Chi-Wing Fu, Kai Chen 0028, Qi Dou 0001, Mingqiang Wei, Yun-Hui Liu 0001, Pheng-Ann Heng
ICRA3
2022 SESR: Self-Ensembling Sim-to-Real Instance Segmentation for Auto-Store Bin Picking
abstract
Instance segmentation is an important task for supporting robotic grasping in auto-store scenarios. Accurate segmentation usually relies on the quantity and quality of available annotated training data. However, it requires tremendous cost to obtain these labels. In this work, without requiring any human annotations on real data, our proposed self-ensembling sim-to-real network, namely SESR, is able to generate precise instance masks for a wide variety of supermarket goods. We design our SESR with a teacher model and a student model trained with a self-ensembling strategy. We adopt different levels of consistency to bridge the sim-to-real gap and boost the model generalization ability. Also, we compile an auto-store bin-picking dataset covering various goods. Extensive experiments on both unseen scenarios and unseen objects validate the effectiveness and superiority of our method over others, and the robot arm demonstrations further show that our segmentation results can support real-time auto-store bin picking.
Biqi Yang, Kai Chen 0028, Yidan Feng, Xianzhi Li 0001, Qi Dou 0001, Chi-Wing Fu, Yun-Hui Liu 0001, Pheng-Ann Heng
IROS6
2022 RePCD-Net: Feature-Aware Recurrent Point Cloud Denoising Network
Honghua Chen, Zeyong Wei, Xianzhi Li 0001, Yabin Xu, Mingqiang Wei, Jun Wang 0039
Int. J. Comput. Vis.3
2022 TIF: Trajectory and Information Flow Coupling Mechanism for Behavior Analysis in Autonomous Driving
abstract
The significant achievements have been made in crowd detection and tracking due to the advancement of artificial intelligence in the autonomous driving. However, the image-based methods have strict requirements for the collection conditions of video, and the development of the new generation of flexible fabrics has become potential sensors to perceive context. In this paper, an intelligent fabric space enabled by multi-sensing sensors is established to track the motion objects. We propose a behavior analysis pipeline including the modules of data preparation, trajectory coupling, motion scenario segmentation, and motion pattern measurement to capture the crowd information from micro-level and macro-level over the intelligent fabric space. After making preprocess for the multi-sensing data, a coupling mechanism is formulated to fuse the video-based trajectory and fabric-based trajectory. And an automatic motion scenario segmentation model divides the surrounding scenario into main-crowd, sub-crowd, and background according to the motion behavior. Further, we define measurement metrics to analyze the motion pattern for the different crowds. Extensive experiments prove that our proposed methods effectively fuse multiple trajectories and realize the crowd segmentation and the motion description. This will greatly help autonomous vehicles and control system perceive the surrounding pedestrians and the environment to make precise driving decisions.
Rui Wang 0077, Jinfeng Xu 0002, Jia Liu 0009, Di Wu 0001, Yixue Hao, Xianzhi Li 0001, Min Chen 0003
IEEE Trans. Intell. Transp. Syst.6
2022 A Rotation-Invariant Framework for Deep Point Cloud Analysis
abstract
Recently, many deep neural networks were designed to process 3D point clouds, but a common drawback is that rotation invariance is not ensured, leading to poor generalization to arbitrary orientations. In this article, we introduce a new low-level purely rotation-invariant representation to replace common 3D Cartesian coordinates as the network inputs. Also, we present a network architecture to embed these representations into features, encoding local relations between points and their neighbors, and the global shape structure. To alleviate inevitable global information loss caused by the rotation-invariant representations, we further introduce a region relation convolution to encode local and non-local information. We evaluate our method on multiple point cloud analysis tasks, including (i) shape classification, (ii) part segmentation, and (iii) shape retrieval. Extensive experimental results show that our method achieves consistent, and also the best performance, on inputs at arbitrary orientations, compared with all the state-of-the-art methods.
Xianzhi Li 0001, Ruihui Li, Guangyong Chen, Chi-Wing Fu, Daniel Cohen-Or, Pheng-Ann Heng
IEEE Trans. Vis. Comput. Graph.1
2021 Point Cloud Upsampling via Disentangled Refinement
abstract
Point clouds produced by 3D scanning are often sparse, non-uniform, and noisy. Recent upsampling approaches aim to generate a dense point set, while achieving both distribution uniformity and proximity-to-surface, and possibly amending small holes, all in a single network. After revisiting the task, we propose to disentangle the task based on its multi-objective nature and formulate two cascaded sub-networks, a dense generator and a spatial refiner. The dense generator infers a coarse but dense out-put that roughly describes the underlying surface, while the spatial refiner further fine-tunes the coarse output by adjusting the location of each point. Specifically, we design a pair of local and global refinement units in the spatial refiner to evolve a coarse feature map. Also, in the spatial refiner, we regress a per-point offset vector to further adjust the coarse outputs in fine scale. Extensive qualitative and quantitative results on both synthetic and real-scanned datasets demonstrate the superiority of our method over the state-of-the-arts. The code is publicly available at https://github.com/liruihui/Dis-PU.
Ruihui Li, Xianzhi Li 0001, Pheng-Ann Heng, Chi-Wing Fu
CVPR2
2021 SP-GAN: sphere-guided 3D shape generation and manipulation
abstract
We present SP-GAN, a new unsupervised sphere-guided generative model for direct synthesis of 3D shapes in the form of point clouds. Compared with existing models, SP-GAN is able to synthesize diverse and high-quality shapes with fine details and promote controllability for part-aware shape generation and manipulation, yet trainable without any parts annotations. In SP-GAN, we incorporate a global prior (uniform points on a sphere) to spatially guide the generative process and attach a local prior (a random latent code) to each sphere point to provide local details. The key insight in our design is to disentangle the complex 3D shape generation task into a global shape modeling and a local structure adjustment, to ease the learning process and enhance the shape generation quality. Also, our model forms an implicit dense correspondence between the sphere points and points in every generated shape, enabling various forms of structure-aware shape manipulations such as part editing, part-wise shape interpolation, and multi-shape part composition, etc., beyond the existing generative models. Experimental results, which include both visual and quantitative evaluations, demonstrate that our model is able to synthesize diverse point clouds with fine details and less noise, as compared with the state-of-the-art models.
Ruihui Li, Xianzhi Li 0001, Ka-Hei Hui, Chi-Wing Fu
ACM Trans. Graph.2
2021 DNF-Net: A Deep Normal Filtering Network for Mesh Denoising
abstract
This article presents a deep normal filtering network, called DNF-Net, for mesh denoising. To better capture local geometry, our network processes the mesh in terms of local patches extracted from the mesh. Overall, DNF-Net is an end-to-end network that takes patches of facet normals as inputs and directly outputs the corresponding denoised facet normals of the patches. In this way, we can reconstruct the geometry from the denoised normals with feature preservation. Besides the overall network architecture, our contributions include a novel multi-scale feature embedding unit, a residual learning strategy to remove noise, and a deeply-supervised joint loss function. Compared with the recent data-driven works on mesh denoising, DNF-Net does not require manual input to extract features and better utilizes the training data to enhance its denoising performance. Finally, we present comprehensive experiments to evaluate our method and demonstrate its superiority over the state of the art on both synthetic and real-scanned meshes.
Xianzhi Li 0001, Ruihui Li, Lei Zhu 0003, Chi-Wing Fu, Pheng-Ann Heng
IEEE Trans. Vis. Comput. Graph.1
2020 PointAugment: An Auto-Augmentation Framework for Point Cloud Classification
abstract
We present PointAugment, a new auto-augmentation framework that automatically optimizes and augments point cloud samples to enrich the data diversity when we train a classification network. Different from existing auto-augmentation methods for 2D images, PointAugment is sample-aware and takes an adversarial learning strategy to jointly optimize an augmentor network and a classifier network, such that the augmentor can learn to produce augmented samples that best fit the classifier. Moreover, we formulate a learnable point augmentation function with a shape-wise transformation and a point-wise displacement, and carefully design loss functions to adopt the augmented samples based on the learning progress of the classifier. Extensive experiments also confirm PointAugment's effectiveness and robustness to improve the performance of various networks on shape classification and retrival.
Ruihui Li, Xianzhi Li 0001, Pheng-Ann Heng, Chi-Wing Fu
CVPR2
2020 Unsupervised Detection of Distinctive Regions on 3D Shapes
abstract
This article presents a novel approach to learn and detect distinctive regions on 3D shapes. Unlike previous works, which require labeled data, our method is unsupervised. We conduct the analysis on point sets sampled from 3D shapes, then formulate and train a deep neural network for an unsupervised shape clustering task to learn local and global features for distinguishing shapes with respect to a given shape set. To drive the network to learn in an unsupervised manner, we design a clustering-based nonparametric softmax classifier with an iterative re-clustering of shapes, and an adapted contrastive loss for enhancing the feature embedding quality and stabilizing the learning process. By then, we encourage the network to learn the point distinctiveness on the input shapes. We extensively evaluate various aspects of our approach and present its applications for distinctiveness-guided shape retrieval, sampling, and view selection in 3D scenes.
Xianzhi Li 0001, Lequan Yu, Chi-Wing Fu, Daniel Cohen-Or, Pheng-Ann Heng
ACM Trans. Graph.1
2019 PU-GAN: A Point Cloud Upsampling Adversarial Network
abstract
Point clouds acquired from range scans are often sparse, noisy, and non-uniform. This paper presents a new point cloud upsampling network called PU-GAN1, which is formulated based on a generative adversarial network (GAN), to learn a rich variety of point distributions from the latent space and upsample points over patches on object surfaces. To realize a working GAN network, we construct an up-down-up expansion unit in the generator for upsampling point features with error feedback and self-correction, and formulate a self-attention unit to enhance the feature integration. Further, we design a compound loss with adversarial, uniform and reconstruction terms, to encourage the discriminator to learn more latent patterns and enhance the output point distribution uniformity. Qualitative and quantitative evaluations demonstrate the quality of our results over the state-of-the-arts in terms of distribution uniformity, proximity-to-surface, and 3D reconstruction quality.
Ruihui Li, Xianzhi Li 0001, Chi-Wing Fu, Daniel Cohen-Or, Pheng-Ann Heng
ICCV2
2019 Deep Floor Plan Recognition Using a Multi-Task Network With Room-Boundary-Guided Attention
abstract
This paper presents a new approach to recognize elements in floor plan layouts. Besides walls and rooms, we aim to recognize diverse floor plan elements, such as doors, windows and different types of rooms, in the floor layouts. To this end, we model a hierarchy of floor plan elements and design a deep multi-task neural network with two tasks: one to learn to predict room-boundary elements, and the other to predict rooms with types. More importantly, we formulate the room-boundary-guided attention mechanism in our spatial contextual module to carefully take room-boundary features into account to enhance the room-type predictions. Furthermore, we design a cross-and-within-task weighted loss to balance the multi-label tasks and prepare two new datasets for floor plan recognition. Experimental results demonstrate the superiority and effectiveness of our network over the state-of-the-art methods.
Zhiliang Zeng, Xianzhi Li 0001, Ying Kin Yu, Chi-Wing Fu
ICCV2
2018 PU-Net: Point Cloud Upsampling Network
abstract
Learning and analyzing 3D point clouds with deep networks is challenging due to the sparseness and irregularity of the data. In this paper, we present a data-driven point cloud upsampling technique. The key idea is to learn multi-level features per point and expand the point set via a multi-branch convolution unit implicitly in feature space. The expanded feature is then split to a multitude of features, which are then reconstructed to an upsampled point set. Our network is applied at a patch-level, with a joint loss function that encourages the upsampled points to remain on the underlying surface with a uniform distribution. We conduct various experiments using synthesis and scan data to evaluate our method and demonstrate its superiority over some baseline methods and an optimization-based method. Results show that our upsampled points have better uniformity and are located closer to the underlying surfaces.
Lequan Yu, Xianzhi Li 0001, Chi-Wing Fu, Daniel Cohen-Or, Pheng-Ann Heng
CVPR2
2018 EC-Net: An Edge-Aware Point Set Consolidation Network
Lequan Yu, Xianzhi Li 0001, Chi-Wing Fu, Daniel Cohen-Or, Pheng-Ann Heng
ECCV (7)2
2018 Non-Local Low-Rank Normal Filtering for Mesh Denoising
abstract
Abstract This paper presents a non‐local low‐rank normal filtering method for mesh denoising. By exploring the geometric similarity between local surface patches on 3D meshes in the form of normal fields, we devise a low‐rank recovery model that filters normal vectors by means of patch groups. In summary, our method has the following key contributions. First, we present the guided normal patch covariance descriptor to analyze the similarity between patches. Second, we pack normal vectors on similar patches into the normal‐field patch‐group (NPG) matrix for rank analysis. Third, we formulate mesh denoising as a low‐rank matrix recovery problem based on the prior that the rank of the NPG matrix is high for raw meshes with noise, but can be significantly reduced for denoised meshes, whose normal vectors across similar patches should be more strongly correlated. Furthermore, we devise an objective function based on an improved truncated γ norm, and derive an op tim ization procedure using the alternative direction method of multipliers and iteratively re‐weighted least squares techniques. We conducted several experiments to evaluate our method using various 3D models, and compared our results against several state‐of‐the‐art methods. Experimental results show that our method consistently outperforms other methods and better preserves the fine details.
Xianzhi Li 0001, Lei Zhu 0003, Chi-Wing Fu, Pheng-Ann Heng
Comput. Graph. Forum1