Lupeng Liu

dblp:238/7086 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0003-2517-7593ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Hi-RWKV: Hierarchical RWKV Modeling for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI) classification demands models that can jointly capture long-range spatial relations and high-dimensional spectral structures while remaining scalable to large scenes and robust under limited supervision. Existing CNN-, Transformer-, and state-space-based approaches either suffer from restricted receptive fields, quadratic attention complexity, or directional biases that hinder dense pixel-wise prediction. To address these limitations, we propose Hi-RWKV, a hierarchical recurrent weighted key-value framework tailored for hyperspectral analysis. Hi-RWKV introduces three key innovations: 1) a spatial structure-guided bidirectional propagation mechanism that integrates global spatial context while preserving boundary fidelity via edge-aware gating; 2) a spectral identity-driven channel mixing module that incorporates learnable band embeddings and whitening transforms to enhance cross-band discriminability; and 3) a multi-stage hierarchical encoder that progressively refines spectral-spatial representations with strictly linear complexity. Together, these designs enable efficient, direction-free spectral-spatial reasoning essential for large-scale HSI interpretation. Extensive experiments on four benchmarks demonstrate that Hi-RWKV consistently achieves state-of-the-art accuracy under diverse training regimes. Ablation studies confirm that each proposed module offers complementary gains in boundary preservation, spectral discrimination, and data efficiency. By unifying scalable recurrence with hyperspectral-specific structural modeling, Hi-RWKV establishes a strong and efficient paradigm for high-resolution remote sensing. The logs and source data of this article are available at https://github.com/HSI-Lab/Hi-RWKV.
Yunbiao Wang, Dongbo Yu, Hengyu Niu, Daifeng Xiao, Lupeng Liu, Jun Xiao 0005
IEEE Trans. Image Process.6
2025 Activating Sparse Part Concepts for 3D Class Incremental Learning
abstract
This work tackles the challenge of 3D Class-Incremental Learning (CIL), where a model must learn to classify new 3D objects while retaining knowledge of previously learned classes. Existing methods often struggle with catastrophic forgetting, misclassifying old objects due to overreliance on shortcut local features. Our approach addresses this issue by learning a set of part concepts for part-aware features. Particularly, we only activate a small subset of part concepts for the feature representation of each part-aware feature. This facilitates better generalization across categories and mitigates catastrophic forgetting. We further improve the task-wise classification through a part relation-aware Transformer design. At last, we devise learnable affinities to fuse task-wise classification heads and avoid confusion among different tasks. We evaluate our method on three 3D CIL benchmarks, achieving state-of-the-art performance. Code is available at https://github.com/zhenyatian/ILPC.
Zhenya Tian, Jun Xiao 0005, Lupeng Liu, Haiyong Jiang
CVPR3
2025 FC-MonoDETR: A Monocular 3D Object Detection Network Based on Foreground Constraint
abstract
Estimating 3D information about objects from a single image is a challenging problem in computer vision due to the lack of multi-view information for depth estimation. The Transformer-based methods propagate the target's 3D center depth within its 2D bounding box to construct object-level depth labels. By treating the 2D box area as a unified entity, these methods can perform sufficient feature sampling and parsing within the aforementioned area. This rough but robust detection strategy effectively avoids the dependence of 3D detection on accurate depth estimation of the target center and a few key points nearby. However, due to the lack of effective foreground constraints, these methods struggle to ensure that sampling points are located inside the target, while exterior points lack valid depth values for supervision, which affects both detection accuracy and stability. To address this issue, we propose a Foreground-Constrained Monocular 3D Object Detector (FC-MonoDETR). First, we leverage 2D annotations to generate target segmentation masks using the Segment Anything Model (SAM), directly establishing depth supervision under foreground constraints. Second, we design an attention-based feature fusion module that utilizes contour information to refine visual features and emphasizes the role of foreground regions in depth estimation, guiding the network to focus more effectively on the foreground during holistic 3D information parsing. Finally, we model the relative depth relationships between targets and optimize the estimation of the target's center depth through a specially designed target center depth loss function. Considering the stability issues of Transformer-based methods, we recommend using a more comprehensive evaluation strategy. The sufficient migration experiments have verified the effectiveness of our constructed foreground-constrained depth supervision and feature fusion module in optimizing Transformer-based methods.
Daifeng Xiao, Dongbo Yu, Yunbiao Wang, Jun Xiao 0005, Ying Wang 0030, Lupeng Liu
ICMR6
2025 GraphSplat: Sparse-View Generalizable 3D Gaussian Splatting is Worth Graph of Nodes
abstract
Generalizable 3D Gaussian Splatting (G-3DGS) has recently emerged as a promising solution for efficient 3D scene representation and novel view synthesis. However, sparse-view scenarios pose a critical challenge for accurate depth estimation. In such cases, viewpoint overlaps are minimal, and many regions are visible from only a single view. As a result, reliable multi-view matching is unavailable in these areas, leading to significant reconstruction quality degradation. To tackle this bottleneck, we propose GraphSplat, a feed-forward framework for novel view synthesis that dynamically incorporates both cross-view and monocular cues through a graph-based feature aggregation strategy. Central to our approach is a Multi-view Aggregate Graph Attention (MAGA) mechanism, which adaptively reweights intra-view and inter-view node connections to compensate for unreliable multi-view correspondences with robust single-view depth priors. In addition, we design a Hierarchical Depth Fusion Estimator (HDFE) module to integrate monocular and multi-view depth cues, effectively reducing ghosting artifacts and improving geometric consistency. Extensive evaluations on RealEstate10K and ACID benchmarks show that GraphSplat achieves competitive performance against prior SOTA methods, with improvements in appearance fidelity and cross-dataset generalization particularly under challenging sparse-view conditions.
Zeyang Bai, Yunbiao Wang, Dongbo Yu, Jun Xiao 0005, Lupeng Liu
ACM Multimedia5
2025 Robust Vegetation Filtering for Rock-Mass Scene Point Clouds via Bidirectional Mamba and Adaptive Triplane Feature Representation
abstract
The irregular and chaotically interwoven distributions of vegetation and rock mass in natural environments pose significant challenges for vegetation filtering of rock point clouds, as existing methods struggle with inefficient contextual modeling and insufficient feature discriminability. To address these limitations, we propose a robust vegetation filtering method for rock mass scene point clouds featuring two key innovations: (1) Bidirectional Point Cloud Mamba module that adopts an alternating allocation strategy to divide the sampled superpoints into two complementary groups, and constructs optimized sequences from edge to center for each group via the Traveling Salesman Problem (TSP) algorithm to enable spatially continuous context capture while eliminating dependency on regular geometric priors; (2) Structure-based adaptive Tri-Plane feature representation includes Principal Component Analysis (PCA)-based support plane selection, projective 2D feature mapping, and learnable multi-scale aggregation. By exploiting the distribution differences of rock mass and vegetation structure information in 2D projections while emulating human visual perception mechanisms that rely on optimal viewing perspectives and different scales, this module enables more discriminative feature extraction in complex natural rock-mass scenarios. Extensive comparative and ablation experiments on real-world datasets demonstrate that our method achieves superior accuracy, robustness, and generalization, and excels in preserving intricate details compared to existing methods.
Shuaichen Guo, Lupeng Liu, Daifeng Xiao, Wenniu Zhang, Ying Wang 0030, Jun Xiao 0005, Dongbo Yu
IEEE Trans. Geosci. Remote. Sens.2
2025 Effective Spatial-Spectral Feature Representation for Hyperspectral Image Classification
abstract
Hyperspectral image (HSI), with their rich spectral information and spatial details, have demonstrated significant potential for classification tasks in fields such as remote sensing, agriculture, and environmental monitoring. However, existing methods still exhibit limitations in feature representation, primarily manifested in insufficient contextual modeling and the inability to effectively address spectral redundancy and significant variations in spatial scales. To address these challenges, this paper proposes an enhanced MambaHSI-based framework for HSI classification, focusing on improving the representation capability of spatial-spectral features. The proposed method consists of three key innovations: (1) A hierarchical DualGroupMamba module that progressively models intra-group and inter-group spectral dependencies to enhance fine-grained spectral discrimination and global contextual awareness; (2) a lightweight Hyperspectral Channel Attention (HCA) that dynamically adjusts the importance of feature channels based on the spatial–spectral information of different bands, effectively suppressing redundant information and highlighting discriminative features. (3) a Hybrid Feature Enhancer (HFE) module that effectively represents and fuses multi-scale spatial features by extracting local texture details and perceiving the overall spatial distribution of scenes, thereby enhancing the model’s adaptability to complex spatial structures. Through a systematic evaluation on four benchmark hyperspectral datasets, the proposed method achieved an average overall classification accuracy of 95.64%, outperforming the current best method by 1.89%. The experimental results validate the superior performance of the proposed approach in enhancing the representation of spatial–spectral features. The latest logs are now available at https://github.com/Tomyaya/EFR.
Dongbo Yu, Yunbiao Wang, Ying Wang 0030, Jun Xiao 0005, Lupeng Liu
IEEE Trans. Geosci. Remote. Sens.6
2025 MambaHSI+: Multidirectional State Propagation for Efficient Hyperspectral Image Classification
abstract
Hyperspectral image classification faces significant challenges due to high-dimensional spectral redundancy and complex spatial-spectral dependencies. While existing MambaHSI models leverage state-space modeling to enhance representation learning, their unidirectional formulation fails to fully capture bidirectional spatial interactions and cross-band contextual dependencies. Moreover, the uniform projection mechanism struggles to effectively distinguish spectral variations across different wavelengths. To address these limitations, we propose MambaHSI+, a novel framework that integrates bidirectional state-space modeling with spectral trajectory learning. The proposed architecture introduces three key innovations: (1) A bidirectional context modeling module, enhanced by reverse-order modeling, enables multi-directional spatial information propagation through recursive state transitions, facilitating comprehensive local-global feature aggregation while maintaining linear computational efficiency; (2) a spectral trajectory learning paradigm that formulates spectral evolution as continuous state-space process with bidirectional propagation, effectively encoding cross-band relationships; and (3) a Mamba-enhanced channel attention mechanism that adaptively emphasizes discriminative spectral features via selective state-space transformations. Extensive experiments on four benchmark datasets demonstrate state-of-the-art performance, achieving an average accuracy improvement of 2.51% over MambaHSI. By integrating state-space systems with advanced Mamba mechanisms, MambaHSI+ establishes a new paradigm for spectral-spatial representation learning, significantly advancing classification accuracy while ensuring computational efficiency. The latest logs are now available at https://github.com/RockAilab/MambaHSI_Plus.
Yunbiao Wang, Lupeng Liu, Jun Xiao 0005, Dongbo Yu, Wenniu Zhang
IEEE Trans. Geosci. Remote. Sens.2
2025 BGPSeg: Boundary-Guided Primitive Instance Segmentation of Point Clouds
abstract
Point cloud primitive instance segmentation is critical for understanding the geometric shapes of man-made objects. Existing learning-based methods mainly focus on learning high-dimensional feature representations of points and further perform clustering or region growing to obtain corresponding primitive instances. However, these features generally cannot accurately represent the discriminability between instances, especially near the boundaries or in regions with small differences in geometric properties. This limitation often leads to over- or under-segmentation of geometric primitives. On the other hand, the boundaries of different primitives are the direct features that distinguish them and thus utilizing boundary information to guide feature learning and clustering is crucial for this task. In this paper, we propose a novel framework BGPSeg for point cloud primitive instance segmentation that utilizes boundary-guided feature extraction and clustering. Specifically, we first introduce a boundary-guided feature extractor with the additional input of a boundary probability map, which utilizes boundary-guided sampling and a boundary transformer to enhance feature discrimination among points crossing geometric boundaries. Furthermore, we propose a boundary-guided primitive clustering module, which combines boundary clues and geometric feature discrimination for clustering to further improve the segmentation performance. Finally, we demonstrate the effectiveness of our BGPSeg with a series of comparison and ablation experiments while achieving the state-of-the-art primitive instance segmentation. Our code is available at https://github.com/fz-20/BGPSeg.
Chuanqing Zhuang, Zhengda Lu, Yiqun Wang 0001, Lupeng Liu, Jun Xiao 0005
IEEE Trans. Image Process.5
2022 A Novel Rock-Mass Point Cloud Registration Method Based on Feature Line Extraction and Feature Point Matching
abstract
Registration will directly affect the quality of overall rock-mass point cloud, which is the basis of 3-D reconstruction for rock mass. Advanced methods establish correspondence by extracting various features that remain unchanged. Although these methods have made great progress, they analyze the local characteristics of each sample point, which leads to be inefficient. In this article, we select registration interesting points from feature lines that were extracted based on supervoxel and innovatively introduce the “clustering, primary matching, and coarse registration” strategy, which effectively reduces the complexity of calculating the corresponding relationship during point cloud registration. Finally, the iterative closest point (ICP) algorithm is used to optimize the result of coarse registration. By selecting registration interesting points from the extracted feature lines, the proposed method inherits the robustness of feature lines to noise, initial position, and so on. The experimental results prove that the coarse registration and refined registration results of the proposed method both have high accuracy and efficiency.
Lupeng Liu, Jun Xiao 0005, Yunbiao Wang, Zhengda Lu, Ying Wang 0030
IEEE Trans. Geosci. Remote. Sens.1
2021 Accurate Rock-Mass Extraction From Terrestrial Laser Point Clouds via Multiscale and Multiview Convolutional Feature Representation
abstract
Existing 3-D object extraction methods on terrestrial laser point clouds are further developed through filtering and labeling. However, such predefined features are heuristically designed to process generic object point clouds. Thus, existing abilities are insufficient to handle specific rock-mass point clouds. Given the complexity and diversity of terrestrial environments, the effective removal of vegetation points from rock-mass point clouds is particularly challenging. To address such problems, this study presents a novel approach for 3-D rock-mass point clouds labeling by using convolutional feature learning based on distribution priors with multiple scales and views. First, to extract discriminative features of each point for classification, we propose novel multiview supporting planes to analyze the spatial distribution and structure of its neighboring points for each category. Second, we define the multiscale spatial distribution matrix on a grid representation (e.g., the number of points projected into each cell). Last, the statistical information of points is nonlinearly combined and hierarchically compressed to generate a compact and effective convolutional feature representation for classification. The effectiveness of the proposed method is evaluated via experiments on rock-mass point clouds from different scenes. Compared with existing extraction approaches, experimental results indicate the superiority of the proposed method in terms of the precision and recall.
Yunbiao Wang, Shibiao Xu, Jun Xiao 0005, Ying Wang 0030, Lupeng Liu
IEEE Trans. Geosci. Remote. Sens.6