Huilong Pi

dblp:400/7443 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
13since 2021 · last 2026
0000-0002-0884-4051ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LiNeXt: Revisiting LiDAR Completion with Efficient Non-Diffusion Architectures
abstract
3D LiDAR scene completion from point clouds is a fundamental component of perception systems in autonomous vehicles. Previous methods have predominantly employed diffusion models for high‑fidelity reconstruction. However, their multi-step iterative sampling incurs significant computational overhead, limiting its real-time applicability. To address this, we propose LiNeXt: a lightweight, non‐diffusion network optimized for rapid and accurate point cloud completion. Specifically, LiNeXt first applies the Noise‑to‑Coarse (N2C) Module to denoise the input noisy point cloud in a single pass, thereby obviating the multi‑step iterative sampling of diffusion‑based methods. The Refine Module then takes the coarse point cloud and its intermediate features from the N2C Module to perform more precise refinement, further enhancing structural completeness. Furthermore, we observe that LiDAR point clouds exhibit a distance-dependent spatial distribution, being densely sampled at proximal ranges and sparsely sampled at distal ranges. Accordingly, we propose the Distance‑aware Selected Repeat strategy to generate a more uniformly distributed noisy point cloud. On the SemanticKITTI dataset, LiNeXt achieves a 199.8 times speedup in inference, reduces Chamfer Distance by 50.7 percent, and uses only 6.1 percent of the parameters compared with LiDiff. These results demonstrate the superior efficiency and effectiveness of LiNeXt for real-time scene completion.
Wenzhe He, Ruihui Li, Huilong Pi, Jiapeng Zhang 0001, Zhuo Tang, Kenli Li 0001
AAAI5
2026 Leveraging pre-trained large language models with refined prompting for online task and motion planning
Huihui Guo, Huilong Pi, Yunchuan Qin, Zhuo Tang, Kenli Li 0001
Neurocomputing2
2026 SONet: An Efficient, Lightweight, and Accurate RGB-T Crowd Counting Network for Smart City Edge Computing
Zuodong Niu, Guoqing Xiao 0001, Huilong Pi, Shenghong Yang, Zhuo Tang
IEEE Internet Things J.3
2026 Application of LLM-powered Multimodal Driver Emotion Recognition in IoV System
abstract
In the Internet of Vehicles (IoV) systems, recognizing driver emotions is crucial to alleviate dangerous driving behaviors caused by emotional instability. Current research predominantly utilizes multimodal data generated by various types of sensors in IoV systems as input to analyze driver emotion changes using multimodal models. However, existing methods are not enough to fully exploit the advantages of large language models (LLM) in information extraction and multimodal feature fusion, which limits the inference capability of emotion recognition models. Therefore, this article proposes an LLM-auxiliary supervision module, which assists in the training phase through LLM to enhance the performance of multimodal emotion recognition models. Specifically, we designed a label text feature extraction (LTFE) module that employs LLM for text data augmentation and extraction, converting label text into semantically informative feature representations. Additionally, we proposed the label-auxiliary supervision (LAS) strategy, which effectively integrates the LLM label text features learned from the LTFE module with the multimodal emotion recognition model during the training phase to enhance the model’s inference ability. Notably, the LTFE and LAS modules are used only during the training phase, ensuring that the backbone model requires minimal computational resources during inference, making it compatible with the computational constraints of intelligent vehicular devices. Extensive experiments conducted on the PPB-Emo, RAVDESS, and IEMOCAP datasets demonstrate that the proposed method outperforms existing approaches in driver emotion recognition tasks.
Yiming Wu 0004, Ronghui Cao, Zhuo Tang, Wangdong Yang, Huilong Pi
ACM Trans. Internet Things6
2026 PI-Net: Point-to-Image Knowledge Distillation for Camera-Based 3D Semantic Scene Completion
abstract
Camera-based Semantic Scene Completion (SSC) aims to infer the geometric structure and semantic information in the entire 3D scene from limited 2D images. However, due to the lack of geometric information in the image, existing methods tend to generate fuzzy completion and incorrect semantic boundaries. In this paper, we propose cross-modal knowledge distillation to address this issue, namely PI-Net, which guides the camera-based model to learn accurate 3D geometry to compensate for spatial surroundings information during training. Specifically, we propose a point cloud occupancy prediction model as the teacher, leveraging its output for strong depth supervision signals and spatial voxel information to enhance the student model. To facilitate effective distillation, we design depth guidance distillation to improve geometric predictions, and spatial guidance distillation to assist the student model in better capturing the structural information of the surrounding environment. Finally, prediction domain distillation is incorporated to facilitate holistic learning from point cloud to image. Experimental results demonstrate that PI-Net outperforms state-of-the-art camera-based methods on challenging benchmarks—SemanticKITTI and SSCBench-KITTI-360.
Yujie Xue, Huilong Pi, Zhuo Tang, Kenli Li 0001, Ruihui Li
IEEE Trans. Multim.2
2025 VLScene: Vision-Language Guidance Distillation for Camera-Based 3D Semantic Scene Completion
abstract
Camera-based 3D semantic scene completion (SSC) provides dense geometric and semantic perception for autonomous driving. However, images provide limited information making the model susceptible to geometric ambiguity caused by occlusion and perspective distortion. Existing methods often lack explicit semantic modeling between objects, limiting their perception of 3D semantic context. To address these challenges, we propose a novel method VLScene: Vision-Language Guidance Distillation for Camera-based 3D Semantic Scene Completion. The key insight is to use the vision-language model to introduce high-level semantic priors to provide the object spatial context required for 3D scene understanding. Specifically, we design a vision-language guidance distillation process to enhance image features, which can effectively capture semantic knowledge from the surrounding environment and improve spatial context reasoning. In addition, we introduce a geometric-semantic sparse awareness mechanism to propagate geometric structures in the neighborhood and enhance semantic information through contextual sparse interactions. Experimental results demonstrate that VLScene achieves rank-1st performance on challenging benchmarks—SemanticKITTI and SSCBench-KITTI-360, yielding remarkably mIoU scores of 17.52 and 19.10, respectively.
Meng Wang 0040, Huilong Pi, Ruihui Li, Yunchuan Qin, Zhuo Tang, Kenli Li 0001
AAAI2
2025 TextHair3D: Text-driven 3D Hair Editing with Generative Priors
abstract
Text-driven hair editing on 3D heads is a challenging problem in computer vision and graphics. In this paper, we propose TextHair3D, a NeRF-based text-driven 3D hair editing method that uses 3D perception to generate priors, edit hair attributes from user-provided text, and preserve facial features. TextHair3D uses the Contrastive Language-Image Pre-training (CLIP) model to encode textual conditions. To address the complexity and roughness of local editing, we design a combined conditional mapping module to map image and text conditions into latent space for learning generative priors. This enables high-quality, photo-realistic hair editing and 3D head reproduction. Extensive experiments show Tex-tHair3D’s superiority in visual realism and attribute accuracy.
Huilong Pi, Yunchuan Qin, Ruihui Li, Kenli Li 0001
ICASSP2
2025 Single-View Reconstruction via Decoupled 3D Gaussian Splatting
abstract
Creating high-quality 3D object representations from a single-view image is challenging. Existing methods tend to infer the geometry and texture information simultaneously within a shared network. However, decoding geometry and texture from a unified network often leads to their entanglement, causing geometric structure collapse or floating artifacts. After revisiting this task, we propose a single-view reconstruction framework based on 3D Gaussian Splatting. The key idea is to decouple Gaussian position attribute generation from texture feature generation. Technically, our framework combines a Geometry Generator, a Texture Generator, and a Gaussian Attributes Decoder. Two parallel branches, Geometry Generator and Texture Generator, aim for point cloud prediction and texture optimization, respectively. Then the Gaussian Attributes Decoder integrates the generated position and texture attributes into a coherent Gaussian point cloud, facilitating efficient novel view synthesis. Extensive qualitative and quantitative evaluations of public datasets demonstrate that our method consistently outperforms existing methods in terms of reconstruction quality and inferring efficiency.
Shiming Zhu, Huilong Pi, Yunchuan Qin, Zhuo Tang, Ruihui Li
ICASSP3
2025 SDFormer: Vision-Based 3D Semantic Scene Completion via SAM-Assisted Dual-Channel Voxel Transformer
Yujie Xue, Huilong Pi, Jiapeng Zhang 0001, Yunchuan Qin, Zhuo Tang, Kenli Li 0001, Ruihui Li
ICCV2
2025 Self-Supervised Point Cloud Completion based on Multi-View Augmentations of Single Partial Point Cloud
abstract
Point cloud completion aims to reconstruct complete shapes from partial observations. Although current methods have achieved remarkable performance, they still have some limitations: Supervised methods heavily rely on ground truth, which limits their generalization to real-world datasets due to the synthetic-to-real domain gap. Unsupervised methods require complete point clouds to compose unpaired training data, and weakly-supervised methods need multi-view observations of the object. Existing self-supervised methods frequently produce unsatisfactory predictions due to the limited capabilities of their self-supervised signals. To overcome these challenges, we propose a novel self-supervised point cloud completion method. We design a set of novel self-supervised signals based on multi-view augmentations of the single partial point cloud. Additionally, to enhance the model’s learning ability, we first incorporate Mamba into self-supervised point cloud completion task, encouraging the model to generate point clouds with better quality. Experiments on synthetic and real-world datasets demonstrate that our method achieves state-of-the-art results.
Jingjing Lu, Huilong Pi, Yunchuan Qin, Zhuo Tang, Ruihui Li
ICME2
2025 CSubBT: A modular execution framework with self-adjusting capability for mobile manipulation system
Huihui Guo, Huizhang Luo, Huilong Pi, Mingxing Duan, Kenli Li 0001, Chubo Liu
Neurocomputing3
2025 Low-Light Domain Enhancement and Multidomain Progressive Fusion for RGB-T Day-Night Crowd Counting
abstract
In IoT-driven intelligent perception systems, multi-modal crowd counting using visible-thermal (RGB-T) sensor arrays deployed on edge nodes plays a vital role in urban management. Existing methods primarily focus on inter-modal feature fusion but suffer from noise amplification when directly integrating thermal features with degraded visible images under low-light conditions, leading to inaccurate density estimation and poor nighttime performance. To address this, we propose the first domain-adaptive network for RGB-T crowd counting. First, we introduce a dark-light domain adaptation module based on Retinex decomposition. Unlike traditional approaches that extract only bright and thermal domain features, our module employs a dual decomposition-recomposition process to learn illumination-invariant features from RGB images, enhancing discriminability in dark regions while preserving bright-domain semantics. Second, the module is supervised by feature alignment, reflectance consistency, and decomposition invariance losses, strengthening its robustness under extreme illumination degradation. Finally, a multi-domain progressive decision fusion module is designed to leverage shared and domain-specific features through public feature extraction, dual-domain fusion, and weighted decision processes, improving generalization across varying lighting conditions. Trained solely on well-lit data, our method generalizes effectively to dark scenarios. Experiments on DroneRGBT and RGBTCC datasets demonstrate superior performance over state-of-the-art fusion methods, particularly excelling in low-light crowd counting.
Zuodong Niu, Huilong Pi, Guoqing Xiao 0001, Shenghong Yang, Zhuo Tang, Dazheng Liu
IEEE Internet Things J.2
2025 A Completing Missing Pedestrian Trajectories Method Driven by Prior-Posterior Knowledge and Interactive Information
abstract
Pedestrian historical trajectory completion significantly bolsters the predictive accuracy of models. However, traditional statistical models such as Hidden Markov Models (HMM), which focus solely on individual pedestrian trajectories, often fall short in terms of generalization. Conversely, data-driven deep learning approaches demand extensive and meticulous data annotation as well as large datasets. Additionally, leveraging sequential historical data and uncovering the correlation between neighboring pedestrians during absences presents a significant challenge. To address these issues, we introduce a novel trajectory completion method that harnesses prior-posterior knowledge and interactive information, termed CMPT. Our approach commences with the design of a Neighbor Pedestrian Selection module (NPS), adept at identifying neighboring pedestrians through a composite scoring system that evaluates feature similarity and proximity. Subsequently, we employ a Top-Graph Attention Network (T-GAT) to extract multiple correlation sets between preceding and succeeding moments within the scenario. These correlations are then fed into the Markov-Inverse Recovery module (MR), which utilizes prior and posterior insights to flesh out the neighbor influence at the unobserved intervals. Culminating in the Trajectory Reconstruction module (TR), we integrate the completed neighbor influence data with the historical trajectory of the missing pedestrian to finalize the missing trajectory reconstruction. Empirical evidence from our experiments indicates that the Final Distance Error (FDE) of the trajectories completed by CMPT is a commendable 0.30. The source code for CMPT is available from https://github.com/ZYueliang/CMPT-Net.
Mingxing Duan, Xinyue Zheng, Huilong Pi, Yan Ding 0004, Zhuo Tang
IEEE Trans. Intell. Transp. Syst.3