Jinfeng Xu 0002

dblp:45/2083-2 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
12since 2021 · last 2026
0009-0005-3357-7724ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Rethinking Point Cloud Representation Learning for Freeing Transformer to Perceive Local
abstract
Transformers are widely utilized in the point cloud domain. However, existing methods tend to overburden Transformer with the dual task of local geometric perception and global feature extraction, limiting its ability to capture highlevel semantic knowledge. To address this issue, we present Representation Decoder (R-Decoder), a novel representation extraction module compatible with various point cloud Transformer methods, enabling the Transformer to focus on its excellent local perception. The R-Decoder iteratively extracts multiple global features from tokens generated by Transformer, refining them to construct an overall representation of point cloud. To ensure full adaptation of the R-Decoder to the knowledge of pre-trained Transformers, we design a cross-modal representation alignment task that leverages multimodal knowledge to specifically pre-train the R-Decoder. As a post-processing module, the R-Decoder seamlessly integrates with Transformers, while decoupling local perception and global representation. This design allows the Transformer to focus on the semantic encoding role for point tokens. Extensive experiments show that our RDecoder significantly boosts the capabilities of 3D representation learning in various point cloud Transformer methods. Notably, it achieves impressive classification accuracies of 95.1% on the ScanObjectNN dataset and 95.3% on the ModelNet40 dataset. Moreover, our method obtains new SOTA on all benchmarks of few-shot and zero-shot classification, while enhancing the multimodal task capabilities of pre-trained Transformers. Code and weights are available athttps://github.com/TangYuan96/RDecoder.
Yunlong Yu 0002, Xianzhi Li 0001, Rui Wang 0077, Jinfeng Xu 0002, Qiao Yu 0002, Yixue Hao, Long Hu, Min Chen 0003
IEEE Trans. Multim.5
2025 More Text, Less Point: Towards 3D Data-Efficient Point-Language Understanding
abstract
Enabling Large Language Models (LLMs) to comprehend the 3D physical world remains a significant challenge. Due to the lack of large-scale 3D-text pair datasets, the success of LLMs has yet to be replicated in 3D understanding. In this paper, we rethink this issue and propose a new task: 3D Data-Efficient Point-Language Understanding. The goal is to enable LLMs to achieve robust 3D object understanding with minimal 3D point cloud and text data pairs. To address this task, we introduce GreenPLM, which leverages more text data to compensate for the lack of 3D data. First, inspired by using CLIP to align images and text, we utilize a pre-trained point cloud-text encoder to map the 3D point cloud space to the text space. This mapping leaves us to seamlessly connect the text space with LLMs. Once the point-text-LLM connection is established, we further enhance text-LLM alignment by expanding the intermediate text space, thereby reducing the reliance on 3D point cloud data. Specifically, we generate 6M free-text descriptions of 3D objects, and design a three-stage training strategy to help LLMs better explore the intrinsic connections between different modalities. To achieve efficient modality alignment, we design a zero-parameter cross-attention module for token pooling. Extensive experimental results show that GreenPLM requires only 12% of the 3D training data used by existing state-of-the-art models to achieve superior 3D understanding. Remarkably, GreenPLM also achieves competitive performance using text-only data.
Xu Han 0016, Xianzhi Li 0001, Qiao Yu 0002, Jinfeng Xu 0002, Yixue Hao, Long Hu, Min Chen 0003
AAAI5
2025 SASep: Saliency-Aware Structured Separation of Geometry and Feature for Open Set Learning on Point Clouds
abstract
Recent advancements in deep learning have greatly enhanced 3D object recognition, but most models are limited to closed-set scenarios, unable to handle unknown samples in real-world applications. Open-set recognition (OSR) addresses this limitation by enabling models to both classify known classes and identify novel classes. However, current OSR methods rely on global features to differentiate known and unknown classes, treating the entire object uniformly and overlooking the varying semantic importance of its different parts. To address this gap, we propose Salience-Aware Structured Separation (SASep), which includes (i) a tunable semantic decomposition (TSD) module to semantically decompose objects into important and unimportant parts, (ii) a geometric synthesis strategy (GSS) to generate pseudo-unknown objects by combining these unimportant parts, and (iii) a synth-aided margin separation (SMS) module to enhance feature-level separation by expanding the feature distributions between classes. Together, these components improve both geometric and feature representations, enhancing the model’s ability to effectively distinguish known and unknown classes. Experimental results show that SASep achieves superior performance in 3D OSR, outperforming existing state-of-the-art methods. The codes are available at https://github.com/JinfengX/SASep.
Jinfeng Xu 0002, Xianzhi Li 0001, Xu Han 0016, Qiao Yu 0002, Yixue Hao, Long Hu, Min Chen 0003
CVPR1
2025 MoST: Efficient Monarch Sparse Tuning for 3D Representation Learning
abstract
We introduce Monarch Sparse Tuning (MoST), the first reparameterization-based parameter-efficient fine-tuning (PEFT) method tailored for 3D representation learning. Unlike existing adapter-based and prompt-tuning 3D PEFT methods, MoST introduces no additional inference overhead and is compatible with many 3D representation learning backbones. At its core, we present a new family of structured matrices for 3D point clouds, Point Monarch, which can capture local geometric features of irregular points while offering high expressiveness. MoST reparameterizes the dense update weight matrices as our sparse Point Monarch matrices, significantly reducing parameters while retaining strong performance. Experiments on various backbones show that MoST is simple, effective, and highly generalizable. It captures local features in point clouds, achieving state-of-the-art results on multiple benchmarks, e.g., 97.5% acc. on ScanOb-jectNN(PB_50_RS) and 96.2% on ModelNet40 classification, while it can also combine with other matrix decompositions (e.g., Low-rank, Kronecker) to further reduce parameters.
Xu Han 0016, Jinfeng Xu 0002, Xianzhi Li 0001
CVPR3
2025 PointDreamer: Zero-Shot 3D Textured Mesh Reconstruction From Colored Point Cloud
abstract
Faithfully reconstructing textured meshes is crucial for many applications. Compared to text or image modalities, leveraging 3D colored point clouds as input (colored-PC-to-mesh) offers inherent advantages in comprehensively and precisely replicating the target object's 360$^{\circ }$∘ characteristics. While most existing colored-PC-to-mesh methods suffer from blurry textures or require hard-to-acquire 3D training data, we propose PointDreamer, a novel framework that harnesses 2D diffusion prior for superior texture quality. Crucially, unlike prior 2D-diffusion-for-3D works driven by text or image inputs, PointDreamer successfully adapts 2D diffusion models to 3D point cloud data by a novel project-inpaint-unproject pipeline. Specifically, it first projects the point cloud into sparse 2D images and then performs diffusion-based inpainting. After that, diverging from most existing 3D reconstruction or generation approaches that predict texture in 3D/UV space thus often yielding blurry texture, PointDreamer achieves high-quality texture by directly unprojecting the inpainted 2D images to the 3D mesh. Furthermore, we identify for the first time a typical kind of unprojection artifact appearing in occlusion borders, which is common in other multiview-image-to-3D pipelines but less-explored. To address this, we propose a novel solution named the Non-Border-First (NBF) unprojection strategy. Extensive qualitative and quantitative experiments on various synthetic and real-scanned datasets demonstrate that PointDreamer, though zero-shot, exhibits SoTA performance (30% improvement on LPIPS score from 0.118 to 0.068), and is robust to noisy, sparse, or even incomplete input data.
Qiao Yu 0002, Xianzhi Li 0001, Xu Han 0016, Jinfeng Xu 0002, Long Hu, Min Chen 0003
IEEE Trans. Vis. Comput. Graph.5
2025 JIMR: Joint Semantic and Geometry Learning for Point Scene Instance Mesh Reconstruction
abstract
Point scene instance mesh reconstruction is a challenging task since it requires both scene-level instance segmentation and instance-level mesh reconstruction from partial observations simultaneously. Previous works either adopt a detection backbone or a segmentation one, and then directly employ a mesh reconstruction network to produce complete meshes from incomplete instance point clouds. To further boost the mesh reconstruction quality with both local details and global smoothness, in this work, we propose JIMR, a joint framework with two cascaded stages for semantic and geometry understanding. In the first stage, we propose to perform both instance segmentation and object detection simultaneously. By making both tasks promote each other, this design facilitates subsequent mesh reconstruction by providing more precisely-segmented instance points and better alignment benefiting from predicted complete bounding boxes. In the second stage, we propose a complete-then-reconstruct procedure, where the completion module explicitly disentangles completion from reconstruction, and enables the usage of pre-trained weights of existing powerful completion and reconstruction networks. Moreover, we propose a comprehensive confidence score to filter proposals considering the quality of instance segmentation, bounding box detection, semantic classification, and mesh reconstruction at the same time. Experiments show that our proposed JIMR outperforms state-of-the-art methods regarding instance reconstruction qualitatively and quantitatively.
Qiao Yu 0002, Xianzhi Li 0001, Jinfeng Xu 0002, Long Hu, Yixue Hao, Min Chen 0003
IEEE Trans. Vis. Comput. Graph.4
2024 PDF: A Probability-Driven Framework for Open World 3D Point Cloud Semantic Segmentation
abstract
Existing point cloud semantic segmentation networks cannot identify unknown classes and update their knowledge, due to a closed-set and static perspective of the real world, which would induce the intelligent agent to make bad decisions. To address this problem, we propose a Probability-Driven Framework (PDF)11Code available at: https://github.com/JinfengX/PointCloudPDF. for open world semantic segmentation that includes (i) a lightweight U-decoder branch to identify unknown classes by estimating the uncertainties, (ii) a flexible pseudo-labeling scheme to supply geometry features along with probability distribution features of unknown classes by generating pseudo labels, and (iii) an incremental knowledge distillation strategy to incorporate novel classes into the existing knowledge base gradually. Our framework enables the model to behave like human beings, which could recognize unknown objects and incrementally learn them with the corresponding knowledge. Experimental results on the S3DIS and ScanNetv2 datasets demonstrate that the proposed PDF outperforms other methods by a large margin in both important tasks of open world semantic segmentation.
Jinfeng Xu 0002, Xianzhi Li 0001, Yixue Hao, Long Hu, Min Chen 0003
CVPR1
2024 Point-LGMask: Local and Global Contexts Embedding for Point Cloud Pre-Training With Multi-Ratio Masking
abstract
Self-supervised learning has achieved great success in both natural language processing and 2D vision, where masked modeling is a quite popular pre-training scheme. However, extending masking to 3D point cloud understanding that combines local and global features poses a new challenge. In our work, we present Point-LGMask, a novel method to embed both local and global contexts with multi-ratio masking, which is quite effective for self-supervised feature learning of point clouds but is unfortunately ignored by existing pre-training works. Specifically, to avoid fitting to a fixed masking ratio, we first propose multi-ratio masking, which prompts the encoder to fully explore representative features thanks to tasks of different difficulties. Next, to encourage the embedding of both local and global features, we formulate a compound loss, which consists of (i) a global representation contrastive loss to encourage the cluster assignments of the masked point clouds to be consistent to that of the completed input, and (ii) a local point cloud prediction loss to encourage accurate prediction of masked points. Equipped with our Point-LGMask, we show that our learned representations transfer well to various downstream tasks, including few-shot classification, shape classification, object part segmentation, as well as real-world scene-based 3D object detection and 3D semantic segmentation. Particularly, our model largely advances existing pre-training methods on the difficult few-shot classification task using the real-captured ScanObjectNN dataset by surpassing over 4% to the second-best method. Also, our Point-LGMask achieves 0.4%$AP_{25}$and 0.8%$AP_{50}$gains on 3D object detection task over the second-best method. 0.4% mAcc and 0.5% mIoU. Codes have been released athttps://github.com/TangYuan96/Point-LGMask.
Xianzhi Li 0001, Jinfeng Xu 0002, Qiao Yu 0002, Long Hu, Yixue Hao, Min Chen 0003
IEEE Trans. Multim.3
2023 CasFusionNet: A Cascaded Network for Point Cloud Semantic Scene Completion by Dense Feature Fusion
abstract
Semantic scene completion (SSC) aims to complete a partial 3D scene and predict its semantics simultaneously. Most existing works adopt the voxel representations, thus suffering from the growth of memory and computation cost as the voxel resolution increases. Though a few works attempt to solve SSC from the perspective of 3D point clouds, they have not fully exploited the correlation and complementarity between the two tasks of scene completion and semantic segmentation. In our work, we present CasFusionNet, a novel cascaded network for point cloud semantic scene completion by dense feature fusion. Specifically, we design (i) a global completion module (GCM) to produce an upsampled and completed but coarse point set, (ii) a semantic segmentation module (SSM) to predict the per-point semantic labels of the completed points generated by GCM, and (iii) a local refinement module (LRM) to further refine the coarse completed points and the associated labels from a local perspective. We organize the above three modules via dense feature fusion in each level, and cascade a total of four levels, where we also employ feature fusion between each level for sufficient information usage. Both quantitative and qualitative results on our compiled two point-based datasets validate the effectiveness and superiority of our CasFusionNet compared to state-of-the-art methods in terms of both scene completion and semantic segmentation. The codes and datasets are available at: https://github.com/JinfengX/CasFusionNet.
Jinfeng Xu 0002, Xianzhi Li 0001, Qiao Yu 0002, Yixue Hao, Long Hu, Min Chen 0003
AAAI1
2023 Crowd Intelligent Grouping Collaboration Evacuation via Multi-agent Reinforcement Learning
abstract
The crowd evacuation strategy seeks to arrange crowd evacuation in an orderly manner to protect people’s lives and reduce property damage in case of sudden emergencies in crowded and complex places. In recent years, there have been several works to apply deep learning to crowd evacuation to make evacuation strategies more intelligent. However, existing researches rarely consider the integration of scene perception and crowd evacuation, which leads to evacuation methods that are detached from the scene and also ignore the crowd collaboration in the evacuation process. To this end, we propose Intelligent Crowd Evacuation Architecture based on Visual features using Multi-Agent Reinforcement Learning (ICEA-VMARL). Subsequently, we present modeling analysis on the crowd grouping and group evacuation modules of the architecture. First, we propose the Population Grouping algorithm based on Continuous Spatiotemporal individual Similarity (PGCSS), which combines crowd features to group crowds. Then, we propose a Group Collaborative Evacuation algorithm based on Multi-Agent Reinforcement Learning (GCE-MARL), which considers group collaboration while evacuating to achieve global optimal evacuation. Finally, we build an experimental crowd simulation system, and the results demonstrate that the crowd grouping algorithm and group evacuation algorithm proposed have better performance compared with other methods.
Rui Wang 0077, Jinfeng Xu 0002, Long Hu, Yixue Hao
CSCWD4
2022 Drone enabled Smart Air-Agent for 6G Network
abstract
The future ubiquitous network, which is mainly characterized by full coverage communication, air-ground integration, multidimensional fusion, network reconfiguration and sensing-communication-computing integration, has become the development trend of 6G technology. The realization of ubiquitous coverage and perceptive fusion of IoT-UAV-Edge is an urgent problem to be solved for complex fusion services. Therefore, this paper proposes a drone-enabled smart air agent in 6G edge fusion system. Firstly, the energy efficient dynamic routing strategy based on joint air-ground control optimization is designed to improve the fusion sensing performance and prolong the service time of drone swarm. Then, the system integration of user-IoT-UAV-Edge is realized to achieve the functionalities of perception, transmission, computing and analysis. Finally, an airborne data fusion mechanism based on multi-source sensing is designed to solve the associated cognitive optimization problem for multi-modal information. The experimental results invalidate the effectiveness and practicability of our system on autonomous path planning, effective computing offloading and accurate airborne fusion.
Yiming Miao, Jinfeng Xu 0002, Min Chen 0003, Kai Hwang 0001
ICC2
2022 TIF: Trajectory and Information Flow Coupling Mechanism for Behavior Analysis in Autonomous Driving
abstract
The significant achievements have been made in crowd detection and tracking due to the advancement of artificial intelligence in the autonomous driving. However, the image-based methods have strict requirements for the collection conditions of video, and the development of the new generation of flexible fabrics has become potential sensors to perceive context. In this paper, an intelligent fabric space enabled by multi-sensing sensors is established to track the motion objects. We propose a behavior analysis pipeline including the modules of data preparation, trajectory coupling, motion scenario segmentation, and motion pattern measurement to capture the crowd information from micro-level and macro-level over the intelligent fabric space. After making preprocess for the multi-sensing data, a coupling mechanism is formulated to fuse the video-based trajectory and fabric-based trajectory. And an automatic motion scenario segmentation model divides the surrounding scenario into main-crowd, sub-crowd, and background according to the motion behavior. Further, we define measurement metrics to analyze the motion pattern for the different crowds. Extensive experiments prove that our proposed methods effectively fuse multiple trajectories and realize the crowd segmentation and the motion description. This will greatly help autonomous vehicles and control system perceive the surrounding pedestrians and the environment to make precise driving decisions.
Rui Wang 0077, Jinfeng Xu 0002, Jia Liu 0009, Di Wu 0001, Yixue Hao, Xianzhi Li 0001, Min Chen 0003
IEEE Trans. Intell. Transp. Syst.2