EDBT 2026 Demo / reviewers in the wild / expert
Ruihui Li
dblp:204/0720
· DBLP profile ↗
45ranked-venue papers
6as first author
42since 2021 · last 2026
0000-0002-4266-6420ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 6 first-author · 29 since 2021Artificial intelligence and machine learning · 21 · 3 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LiNeXt: Revisiting LiDAR Completion with Efficient Non-Diffusion Architecturesabstract3D LiDAR scene completion from point clouds is a fundamental component of perception systems in autonomous vehicles. Previous methods have predominantly employed diffusion models for high‑fidelity reconstruction. However, their multi-step iterative sampling incurs significant computational overhead, limiting its real-time applicability. To address this, we propose LiNeXt: a lightweight, non‐diffusion network optimized for rapid and accurate point cloud completion. Specifically, LiNeXt first applies the Noise‑to‑Coarse (N2C) Module to denoise the input noisy point cloud in a single pass, thereby obviating the multi‑step iterative sampling of diffusion‑based methods. The Refine Module then takes the coarse point cloud and its intermediate features from the N2C Module to perform more precise refinement, further enhancing structural completeness. Furthermore, we observe that LiDAR point clouds exhibit a distance-dependent spatial distribution, being densely sampled at proximal ranges and sparsely sampled at distal ranges. Accordingly, we propose the Distance‑aware Selected Repeat strategy to generate a more uniformly distributed noisy point cloud. On the SemanticKITTI dataset, LiNeXt achieves a 199.8 times speedup in inference, reduces Chamfer Distance by 50.7 percent, and uses only 6.1 percent of the parameters compared with LiDiff. These results demonstrate the superior efficiency and effectiveness of LiNeXt for real-time scene completion. Wenzhe He, Ruihui Li, Huilong Pi, Jiapeng Zhang 0001, Zhuo Tang, Kenli Li 0001 |
AAAI | 4 |
| 2026 | PointSLAM++: Robust Dense Neural Gaussian Point Cloud-based SLAMabstractReal-time 3D reconstruction is crucial for robotics and augmented reality, yet current simultaneous localization and mapping(SLAM) approaches often struggle to maintain structural consistency and robust pose estimation in the presence of depth noise. This work introduces PointSLAM++, a novel RGB-D SLAM system that leverages a hierarchically constrained neural Gaussian representation to preserve structural relationships while generating Gaussian primitives for scene mapping. It also employs progressive pose optimization to mitigate depth sensor noise, significantly enhancing localization accuracy. Furthermore, it utilizes a dynamic neural representation graph that adjusts the distribution of Gaussian nodes based on local geometric complexity, enabling the map to adapt to intricate scene details in real time. This combination yields high-precision 3D mapping and photorealistic scene rendering. Experimental results show PointSLAM++ outperforms existing 3DGS-based SLAM methods in reconstruction accuracy and rendering quality, demonstrating its advantages for large-scale AR and robotics. Boyao Han, Ruihui Li |
AAAI | 5 |
| 2026 | LaCo: Layer-wise Compensation for Pruned Large Language ModelsabstractPruning is essential for the efficient deployment of Large Language Models (LLMs); however, it causes severe performance degradation due to the structural distortion induced by sparsity.Existing recovery strategies, such as LoRA, predominantly employ global finetuning, often overlooking the mechanistic root of this degradation: the layer-wise accumulation and amplification of local errors.To address this limitation, we propose LaCo (Layerwise Compensation), a framework that reorients the recovery paradigm from global adaptation to hierarchical representation alignment.By sequentially optimizing each layer to reconstruct the model's hidden states, LaCo effectively intercepts the error propagation chain at its source.Extensive experiments demonstrate that LaCo surpasses parameter-efficient baselines in both perplexity reduction and zeroshot reasoning.Notably, it reduces recoverytime memory usage to approximately 1/7 of the baseline and requires only 2,048 unlabeled samples to match a LoRA model trained on 50k examples-achieving a ∼ 25× improvement in data efficiency. Yingen Liu, Fan Wu 0016, Xuyan Pan, Ruihui Li, Zhuo Tang, Kenli Li 0001 |
ACL (1) | 4 |
| 2026 | SCPid: Point cloud semantic scene completion via spatial-chunked perception and image prior distillation
Ying Liu 0027, Ruihui Li |
Expert Syst. Appl. | 4 |
| 2026 | FacialTalk: Audio-driven high-fidelity facial portrait generation using 3D facial prior
Daowu Yang, Qiyun Yang, Ruihui Li |
Pattern Recognit. | 4 |
| 2026 | ISDNet: High-Fidelity Single-View Reconstruction of Indoor Scenes via Instance Separation and DeformationabstractIn this work, we aim to reconstruct the 3D shape of an indoor scene from a single view, which includes multiple objects and the background. This task is challenging for existing methods since those instances of indoor scenes regularly occlude each other and contain diverse topologies. To address this, we propose a novel framework, ISDNet, to adaptively separate mixed instances and perform topology-aware reconstruction. Specifically, Specifically, ISDNet consists of two cascaded subnetworks: an instance separation module (ISM) and an instance deformation module (IDM). The ISM learns to separate occluded objects through stepwise sampling, inferring clean features for each instance. On the basis of these features, IDM generates an instance-topology-aware template and deforms it with learned offsets to reconstruct detailed geometry. Quantitative and qualitative experiments on the SUNRGB-D and 3D-FRONT datasets demonstrate that ISDNet outperforms the state-of-theart methods in terms of local details and overall shapes. Xiaolin He, Ying Liu 0027, Yiming Han, Junxian Chen, Ruihui Li |
IEEE Trans. Multim. | 5 |
| 2026 | PI-Net: Point-to-Image Knowledge Distillation for Camera-Based 3D Semantic Scene CompletionabstractCamera-based Semantic Scene Completion (SSC) aims to infer the geometric structure and semantic information in the entire 3D scene from limited 2D images. However, due to the lack of geometric information in the image, existing methods tend to generate fuzzy completion and incorrect semantic boundaries. In this paper, we propose cross-modal knowledge distillation to address this issue, namely PI-Net, which guides the camera-based model to learn accurate 3D geometry to compensate for spatial surroundings information during training. Specifically, we propose a point cloud occupancy prediction model as the teacher, leveraging its output for strong depth supervision signals and spatial voxel information to enhance the student model. To facilitate effective distillation, we design depth guidance distillation to improve geometric predictions, and spatial guidance distillation to assist the student model in better capturing the structural information of the surrounding environment. Finally, prediction domain distillation is incorporated to facilitate holistic learning from point cloud to image. Experimental results demonstrate that PI-Net outperforms state-of-the-art camera-based methods on challenging benchmarks—SemanticKITTI and SSCBench-KITTI-360. Yujie Xue, Huilong Pi, Zhuo Tang, Kenli Li 0001, Ruihui Li |
IEEE Trans. Multim. | 5 |
| 2026 | HSG-Net: Point Cloud Completion via Heuristic Structure GrowingabstractExisting point cloud completion methods rely on extracting latent codes from a partial point cloud to reconstruct a complete structure. However, the complexity of the partial point clouds, making the completion results of such methods less satisfactory, especially in long-distance (away from partial point cloud) areas. To tackle this challenge, we propose a point cloud completion network via heuristic structure growing (HSG-Net), which progressively completes the close-distance structure through an iterative heuristic structure growth strategy. Particularly, a novel data preprocessing (DP) method is proposed to obtain ground truth (GT) with specific structural integrity, guiding the network to learn close-distance structural information. In addition, the proposed consistency constraint displacement module (CCDM) is employed to fulfill structure growth, and a feature memory module (FMM) further enhances the quality of the grown structure. Furthermore, a proposed local information generator is used to further refine the structure-grown point cloud, fetching the final result. Extensive quantitative and qualitative results demonstrate that our HSG-Net outperforms the state-of-the-art methods. Junxian Chen, Ying Liu 0027, Ruihui Li |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | VLScene: Vision-Language Guidance Distillation for Camera-Based 3D Semantic Scene CompletionabstractCamera-based 3D semantic scene completion (SSC) provides dense geometric and semantic perception for autonomous driving. However, images provide limited information making the model susceptible to geometric ambiguity caused by occlusion and perspective distortion. Existing methods often lack explicit semantic modeling between objects, limiting their perception of 3D semantic context. To address these challenges, we propose a novel method VLScene: Vision-Language Guidance Distillation for Camera-based 3D Semantic Scene Completion. The key insight is to use the vision-language model to introduce high-level semantic priors to provide the object spatial context required for 3D scene understanding. Specifically, we design a vision-language guidance distillation process to enhance image features, which can effectively capture semantic knowledge from the surrounding environment and improve spatial context reasoning. In addition, we introduce a geometric-semantic sparse awareness mechanism to propagate geometric structures in the neighborhood and enhance semantic information through contextual sparse interactions. Experimental results demonstrate that VLScene achieves rank-1st performance on challenging benchmarks—SemanticKITTI and SSCBench-KITTI-360, yielding remarkably mIoU scores of 17.52 and 19.10, respectively. Meng Wang 0040, Huilong Pi, Ruihui Li, Yunchuan Qin, Zhuo Tang, Kenli Li 0001 |
AAAI | 3 |
| 2025 | TextHair3D: Text-driven 3D Hair Editing with Generative PriorsabstractText-driven hair editing on 3D heads is a challenging problem in computer vision and graphics. In this paper, we propose TextHair3D, a NeRF-based text-driven 3D hair editing method that uses 3D perception to generate priors, edit hair attributes from user-provided text, and preserve facial features. TextHair3D uses the Contrastive Language-Image Pre-training (CLIP) model to encode textual conditions. To address the complexity and roughness of local editing, we design a combined conditional mapping module to map image and text conditions into latent space for learning generative priors. This enables high-quality, photo-realistic hair editing and 3D head reproduction. Extensive experiments show Tex-tHair3D’s superiority in visual realism and attribute accuracy. Huilong Pi, Yunchuan Qin, Ruihui Li, Kenli Li 0001 |
ICASSP | 4 |
| 2025 | Single-View Reconstruction via Decoupled 3D Gaussian SplattingabstractCreating high-quality 3D object representations from a single-view image is challenging. Existing methods tend to infer the geometry and texture information simultaneously within a shared network. However, decoding geometry and texture from a unified network often leads to their entanglement, causing geometric structure collapse or floating artifacts. After revisiting this task, we propose a single-view reconstruction framework based on 3D Gaussian Splatting. The key idea is to decouple Gaussian position attribute generation from texture feature generation. Technically, our framework combines a Geometry Generator, a Texture Generator, and a Gaussian Attributes Decoder. Two parallel branches, Geometry Generator and Texture Generator, aim for point cloud prediction and texture optimization, respectively. Then the Gaussian Attributes Decoder integrates the generated position and texture attributes into a coherent Gaussian point cloud, facilitating efficient novel view synthesis. Extensive qualitative and quantitative evaluations of public datasets demonstrate that our method consistently outperforms existing methods in terms of reconstruction quality and inferring efficiency. Shiming Zhu, Huilong Pi, Yunchuan Qin, Zhuo Tang, Ruihui Li |
ICASSP | 6 |
| 2025 | TD-GS: Few-shot Object View Synthesis via Task-Disentangled 3D Gaussian Splattingabstract3D Gaussian Splatting (3D-GS) has exhibited impressive progress in novel view synthesis. When given the sparse views, its performance degrades severely, causing many problems like novel views collapse and excessive floaters. Many recent methods take into account fitting input views, inferring missing scene information and optimizing the final scene representation, all through a single stage. After revisiting the task, we propose a novel framework, Task-Disentangled 3D Gaussian Splatting, abbreviated to TD-GS. It splits the sparse views synthesis task into two subtasks: (i) Dense Generation. (ii) Enhanced Synthesis. In the subtask of Dense Generation, we estimate dense views from sparse input. Then in the subtask of Enhanced Synthesis, both the dense views and the sparse input participate in the training of Gaussians to obtain the final Gaussian representation of the scene. In the process of completing the first subtask, we carefully design Gaussian Cloud Denoising to directly edit 3D Gaussians. Also, we introduce two regularization methods to guide the geometric optimization towards an optimal solution. The purpose is to estimate more reliable outputs. Many experiments have validated that our TD-GS outperforms other state-of-the-art methods. Ying Liu 0027, Xiaohao Zhang, Zhuo Tang, Ruihui Li |
ICASSP | 6 |
| 2025 | High-Fidelity Single-View Reconstruction of Indoor Scenes using 3D Shape Prior Template and Pixel-Aligned DeformationabstractThis paper presents a novel pipeline for estimating room layouts and reconstructing the 3D shapes of indoor objects. This task remains challenging due to occlusions of indoor scenes, which lead to incomplete shape and poor geometric quality manifested as non-smooth meshes. Our key insight is that occlusions of indoor objects inherently lead to insufficient information in images. Pixel-level features alone are inadequate to recover the complete structure of objects; Therefore, additional prior information is required to supplement the missing occluded details. To address this, we propose a two-stage training strategy. First, a VQ-VAE encodes 3D shapes into a latent space, with a decoder leveraging priors to predict occluded regions. Second, pixel-aligned features are used for deformation, ensuring consistency between the reconstructed shape and the image. Quantitative and qualitative evaluations on the 3D-FRONT and SUNRGB-D datasets demonstrate that our approach surpasses state-of-the-art methods in reconstructing more complete geometric topologies and smoother meshes. Xiaohao Zhang, Xiaolin He, Zhuo Tang, Ruihui Li |
ICASSP | 6 |
| 2025 | SDFormer: Vision-Based 3D Semantic Scene Completion via SAM-Assisted Dual-Channel Voxel Transformer
Yujie Xue, Huilong Pi, Jiapeng Zhang 0001, Yunchuan Qin, Zhuo Tang, Kenli Li 0001, Ruihui Li |
ICCV | 7 |
| 2025 | MROSS: Multi-Round Region-based Optimization for Scene SketchingabstractScene sketching is to convert a scene into a simplified, abstract representation that captures the essential elements and composition of the original scene. It requires a semantic understanding of the scene and consideration of different regions within the scene. Since scenes often contain diverse visual information across various regions, such as foreground objects, background elements, and spatial divisions, dealing with these different regions poses unique difficulties. In this paper, we define a sketch as some sets of Bézier curves because of their smooth and versatile characteristics. We optimize different regions of input scene in multiple rounds. In each optimization round, the strokes sampled from the next region can seamlessly be integrated into the sketch generated in the previous optimization round. We propose an additional stroke initialization method to ensure the integrity of the scene and the convergence of optimization. A novel CLIP-based Semantic Loss and a VGG-based Feature Loss are utilized to guide our multi-round optimization. Extensive experimental results on the quality and quantity of the generated sketches confirm the effectiveness of our method. Yiqi Liang, Ying Liu 0027, Dandan Long, Ruihui Li |
ICME | 4 |
| 2025 | Self-Supervised Point Cloud Completion based on Multi-View Augmentations of Single Partial Point CloudabstractPoint cloud completion aims to reconstruct complete shapes from partial observations. Although current methods have achieved remarkable performance, they still have some limitations: Supervised methods heavily rely on ground truth, which limits their generalization to real-world datasets due to the synthetic-to-real domain gap. Unsupervised methods require complete point clouds to compose unpaired training data, and weakly-supervised methods need multi-view observations of the object. Existing self-supervised methods frequently produce unsatisfactory predictions due to the limited capabilities of their self-supervised signals. To overcome these challenges, we propose a novel self-supervised point cloud completion method. We design a set of novel self-supervised signals based on multi-view augmentations of the single partial point cloud. Additionally, to enhance the model’s learning ability, we first incorporate Mamba into self-supervised point cloud completion task, encouraging the model to generate point clouds with better quality. Experiments on synthetic and real-world datasets demonstrate that our method achieves state-of-the-art results. Jingjing Lu, Huilong Pi, Yunchuan Qin, Zhuo Tang, Ruihui Li |
ICME | 5 |
| 2025 | SA-MVSNet: Spatial-aware Multi-view Stereo Network with Attention Cost VolumeabstractDeep learning-based multi-view stereo (MVS) methods enable dense point cloud reconstruction in texture-rich areas. However, existing methods incur significant computational costs to capture pixel dependencies for complete reconstruction in low-texture regions. Additionally, discrete depth layers in occluded environments hinder the cost volume’s ability to model object information effectively. To address these issues, we propose a spatial-aware multi-view stereo network with attention cost volume, termed SA-MVSNet. The network introduces the pixel-driven spatial interaction (PDSI) module, which integrates the hierarchical spatial location enhancement mechanism (HSLE) and the spatial context aggregation mechanism (SCA). Leveraging an efficient parallel architecture, the PDSI module captures pixel-level spatial dependencies with the HSLE and strengthens global contextual information through the SCA. This design improves the network’s ability to represent features in low-texture regions while maintaining high inference efficiency. Furthermore, SA-MVSNet incorporates an attention weight generation branch that refines the cost volume by aggregating multi-scale depth cues, effectively mitigating the impact of occlusion. Experiments on the DTU dataset and the Tanks and Temples dataset show that our method outperforms other learning-based methods, achieving superior performance and strong generalization ability. Haoran Kong, Fanzi Zeng, Longbao Dai, Jingyang Hu, Jiang-hao Cai, Jianxia Chen, Ruihui Li, Hongbo Jiang 0001 |
IROS | 7 |
| 2025 | RWKV-PCSSC: Exploring RWKV Model for Point Cloud Semantic Scene CompletionabstractSemantic Scene Completion (SSC) aims to generate a complete semantic scene from an incomplete input. Existing approaches often employ dense network architectures with a high parameter count, leading to increased model complexity and resource demands. To address these limitations, we propose RWKV-PCSSC, a lightweight point cloud semantic scene completion network inspired by the Receptance Weighted Key Value (RWKV) mechanism. Specifically, we introduce a RWKV Seed Generator (RWKV-SG) module that can aggregate features from a partial point cloud to produce a coarse point cloud with coarse features. Subsequently, the point-wise feature of the point cloud is progressively restored through multiple stages of the RWKV Point Deconvolution (RWKV-PD) modules. By leveraging a compact and efficient design, our method achieves a lightweight model representation. Experimental results demonstrate that RWKV-PCSSC reduces the parameter count by 4.18× and improves memory efficiency by 1.37× compared to state-of-the-art methods PointSSC[51]. Furthermore, our network achieves state-of-the-art performance on established indoor (SSC-PC, NYUCAD-PC) and outdoor (PointSSC) scene dataset, as well as on our proposed datasets (NYUCAD-PC-V2, 3D-FRONT-PC). Wenzhe He, Wentang Chen, Ying Liu 0027, Ruihui Li |
ACM Multimedia | 6 |
| 2025 | Learning Temporal 3D Semantic Scene Completion via Optical Flow Guidanceabstract3D Semantic Scene Completion (SSC) provides comprehensive scene geometry and semantics for autonomous driving perception, which is crucial for enabling accurate and reliable decision-making. However, existing SSC methods are limited to capturing sparse information from the current frame or naively stacking multi-frame temporal features, thereby failing to acquire effective scene context. These approaches ignore critical motion dynamics and struggle to achieve temporal consistency. To address the above challenges, we propose a novel temporal SSC method FlowScene: Learning Temporal 3D Semantic Scene Completion via Optical Flow Guidance. By leveraging optical flow, FlowScene can integrate motion, different viewpoints, occlusions, and other contextual cues, thereby significantly improving the accuracy of 3D scene completion. Specifically, our framework introduces two key components: (1) a Flow-Guided Temporal Aggregation module that aligns and aggregates temporal features using optical flow, capturing motion-aware context and deformable structures; and (2) an Occlusion-Guided Voxel Refinement module that injects occlusion masks and temporally aggregated features into 3D voxel space, adaptively refining voxel representations for explicit geometric modeling.
Experimental results demonstrate that FlowScene achieves state-of-the-art performance, with mIoU of 17.70 and 20.81 on the SemanticKITTI and SSCBench-KITTI-360 benchmarks. Meng Wang 0040, Fan Wu 0016, Ruihui Li, Yunchuan Qin, Zhuo Tang, Li Ken Li |
NeurIPS | 3 |
| 2025 | MoNeRF: Deformable Neural Rendering for Talking Heads via Latent Motion NavigationabstractAbstract Novel view synthesis for talking heads presents significant challenges due to the complex and diverse motion transformations involved. Conventional methods often resort to reliance on structure priors, like facial templates, to warp observed images into a canonical space conducive to rendering. However, the incorporation of such priors introduces a trade‐off‐while aiding in synthesis, they concurrently amplify model complexity, limiting generalizability to other deformable scenes. Departing from this paradigm, we introduce a pioneering solution: the motion‐conditioned neural radiance field, MoNeRF, designed to model talking heads through latent motion navigation. At the core of MoNeRF lies a novel approach utilizing a compact set of latent codes to represent orthogonal motion directions. This innovative strategy empowers MoNeRF to efficiently capture and depict intricate scene motion by linearly combining these latent codes. In an extended capability, MoNeRF facilitates motion control through latent code adjustments, supports view transfer based on reference videos, and seamlessly extends its applicability to model human bodies without necessitating structural modifications. Rigorous quantitative and qualitative experiments unequivocally demonstrate MoNeRF's superior performance compared to state‐of‐the‐art methods in talking head synthesis. We will release the source code upon publication. Yan Ding 0004, Ruihui Li, Zhuo Tang, Kenli Li 0001 |
Comput. Graph. Forum | 3 |
| 2025 | Vision-based 3D semantic scene completion via capture dynamic representationsabstractThe vision-based semantic scene completion task aims to predict dense geometric and semantic 3D scene representations from 2D images. However, the presence of dynamic objects in the scene seriously affects the accuracy of the model inferring 3D structures from 2D images. Existing methods simply stack multiple frames of image input to increase dense scene semantic information, but ignore the fact that dynamic objects and non-texture areas violate multi-view consistency and matching reliability. To address these issues, we propose a novel method, CDScene: Vision-based 3D Semantic Scene Completion via Capturing Dynamic Representations. First, we leverage a large multi-modal model to extract 2D explicit semantics and align them into 3D space. Second, we exploit the characteristics of monocular and stereo depth to decouple scene information into dynamic and static features. The dynamic features contain structural relationships around dynamic objects, and the static features contain dense contextual spatial information. Finally, we design a dynamic-static adaptive fusion module to effectively extract and aggregate complementary features, achieving robust and accurate semantic scene completion in autonomous driving scenarios. Extensive experimental results on the SemanticKITTI, SSCBench-KITTI360, and SemanticKITTI-C datasets demonstrate the superiority and robustness of CDScene over existing state-of-the-art methods. Meng Wang 0040, Fan Wu 0016, Yunchuan Qin, Ruihui Li, Zhuo Tang, Kenli Li 0001 |
Knowl. Based Syst. | 4 |
| 2025 | Decoupling upper and lower face transformers for binary interactive video generation
Daowu Yang, Ying Liu 0027, Qiyun Yang, Ruihui Li |
Neural Networks | 4 |
| 2025 | MixSSC: Forward-Backward Mixture for Vision-Based 3D Semantic Scene CompletionabstractVision-based semantic scene completion task aims to predict dense geometric and semantic 3D scene representations from 2D images. However, 3D modeling from a single view is an ill-posed problem, limited by the field of view and occlusion problems caused by image input. Moreover, existing methods tend to produce erroneous scene hallucinations and overly smooth boundary segmentation due to a lack of information. To address this problem, we propose MixSSC, which mixes the sparsity of forward projection with the denseness of depth-prior backward projection. The aim is to use sparse features to fill information-poor regions and dense features to enhance visible regions. Specifically, we develop the forward-backward mixture module, which enables the generation of scene mixture voxel representation by leveraging the benefits of both forward and backward projection. Subsequently, we design the semantic-spatial fusion module, which utilizes a coarse-to-fine approach to process mixture voxel features at the semantic-spatial level. Extensive experimental results on the SemanticKITTI, SSCBench-KITTI-360 and nuScenes datasets demonstrate the superiority of MixSSC. Our code is available on https://github.com/willemeng/MixSSC. Meng Wang 0040, Yan Ding 0004, Yunchuan Qin, Ruihui Li, Zhuo Tang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Bi-SSC: Geometric-Semantic Bidirectional Fusion for Camera-Based 3D Semantic Scene CompletionabstractCamera-based Semantic Scene Completion (SSC) is to infer the full geometry of objects and scenes from only 2D images. The task is particularly challenging for those in-visible areas, due to the inherent occlusions and lighting ambiguity. Existing works ignore the information missing or ambiguous in those shaded and occluded areas, resulting in distorted geometric prediction. To address this issue, we propose a novel method, Bi-SSC, bidirectional geomet-ric semantic fusion for camera-based 3D semantic scene completion. The key insight is to use the neighboring structure of objects in the image and the spatial differences from different perspectives to compensate for the lack of information in occluded areas. Specifically, we introduce a spatial sensory fusion module with multiple association attention to improve semantic correlation in geometric distributions. This module works within single view and across stereo views to achieve global spatial consistency. Experimental results demonstrate that Bi-SSC outperforms state-of-the-art camera-based methods on SemanticKITTI, particularly excelling in those invisible and shaded areas. Yujie Xue, Ruihui Li, Fan Wu 0016, Zhuo Tang, Kenli Li 0001, Mingxing Duan |
CVPR | 2 |
| 2024 | MC-SORT: A Motion Correction-Based Framework for Long-Term Multiple Object TrackingabstractLong-term occlusion is one of the most formidable challenges in Multi-Object Tracking (MOT). The motion models of existing SORT-based trackers are unreliable in estimating the motion states of long-term occluded targets. This is mainly because as the occlusion period increases, the increases speed of estimation errors in the motion model increases faster. In practical applications, we believe that the estimation error of the tracker during long-term occlusion is mainly concentrated in the estimation error of the motion model on the velocity of the occluded target. In this work, we have demonstrated that in the long-term occlusion period, appropriately correcting the estimated values of the motion model on the target motion velocity and fully utilizing the temporal and attribute information of the target’s historical trajectory as calculation indicators of correlation are beneficial for improving the robustness of the tracker in long-term occlusion. We refer to our proposed motion correction-based framework as MC-SORT, which mainly consists of a Momentum Compensation Module (MCM) and a Backtracking Re-association (BRA) module. The former can correct the estimated value of the target’s motion state during long-term occlusion, the latter uses the temporal and attribute information of the target’s historical trajectory during long-term occlusion as correlation indicators to measure the degree of correlation between the target and trajectory. Our proposed MC-SORT has the characteristics of simplicity, online, real-time, and plug-and-play, particularly improving the robustness of the tracker in long-term occlusion. The extensive experimental results on the MOT17 and MOT20 datasets demonstrate the robustness and superiority of our framework. Yunchuan Qin, Ruihui Li, Guanghua Tan, Zhuo Tang, Kenli Li 0001 |
ECAI | 3 |
| 2024 | DeformingNet: Deforming Multiple Uniform 3D Priors for 3D Point Cloud CompletionabstractWe propose DeformingNet, an effective 3D point cloud completion network. Unlike existing methods that complete partial point cloud by directly learning the morphing function from 2D grids to 3D shapes, which limits the model’s inference capability due to the intrinsic gaps between feature spaces with different dimensions, we design a deforming-based point generator that emulates the deforming from multiple uniform 3D priors (i.e., pre-defined 3D point clouds in cube shape) into 3D shapes. In addition, we design a MAE-based encoder, which introduces the MAE encoder from point-MAE pre-trained on ShapeNet dataset to learn the correlation of local regions and fuse local and global features to enrich the latent representation. The 3D shapes generated by DeformingNet have both accurate local details and faithful global structure with less noise. Experiments demonstrate that our DeformingNet outperforms state-of-the-art methods in terms of quantitative metrics and visual quality. Jingjing Lu, Yunchuan Qin, Fan Wu 0016, Kenli Li 0001, Ruihui Li |
ICME | 6 |
| 2024 | Talking Portrait with Discrete Motion Priors in Neural Radiation FieldabstractSpeech-driven facial video is a one-to-many mapping problem where each input audio can have multiple plausible facial outputs, leading to overly smooth facial movements results. To overcome this problem, we introduce discrete motion priors to reduce the uncertainty in facial movements and enhance realism. Simultaneously through adversarial training, we achieve domain adaptation, mapping facial movements to a low-dimensional grid space to enhance adaptability to external audio. Following that we establish an explicit connection between facial movements region and spatial regions using a spatial position attention to improve the precision of dynamic portrait modeling. Finally we expedite the generation of speaking face videos using a grid-based dynamic neural radiation field to render the head and torso separately. Experimental results demonstrate that our approach outperforms previous methods, yielding more extensive and realistic speech-driven facial video. Daowu Yang, Ying Liu 0027, Qiyun Yang, Ruihui Li |
ICME | 4 |
| 2024 | 3D head-talk: speech synthesis 3D head movement face animation
Daowu Yang, Ruihui Li, Yuyi Peng, Xibei Huang |
Soft Comput. | 2 |
| 2024 | AWDepth: Monocular Depth Estimation for Adverse Weather via Masked EncodingabstractMonocular depth estimation has made considerable advances under clear weather conditions. However, how to learn accurate scene depth under rain and fog conditions and alleviate the negative influence of occlusion, light, visibility, etc., is an open problem. To address this problem, in this article, we split the adverse weather depth estimation network into two subbranches: the depth prediction branch and the masked encoding branch. The depth prediction branch is used for depth estimation. The masked encoding branch, inspired by masked image modeling, uses random masks to simulate occlusion or low visibility often seen in rain and fog, forcing this branch to learn to infer the prediction of masked regions from the context. In order to make the masked encoding better enhance the depth prediction, we designed the mask feature fusion module, which can fuse the depth and spatial context features of the two branches to produce a fine-level depth map. The experimental results on the Foggy Cityscapes and RainCityscapes datasets demonstrate that our method achieves state-of-the-art performance, significantly outperforming previous methods across all evaluation metrics. Meng Wang 0040, Yunchuan Qin, Ruihui Li, Zhuo Tang, Kenli Li 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | Neural Wavelet-domain Diffusion for 3D Shape Generation, Inversion, and ManipulationabstractThis paper presents a new approach for 3D shape generation, inversion, and manipulation, through a direct generative modeling on a continuous implicit representation in wavelet domain. Specifically, we propose acompact wavelet representationwith a pair of coarse and detail coefficient volumes to implicitly represent 3D shapes via truncated signed distance functions and multi-scale biorthogonal wavelets. Then, we design a pair of neural networks: a diffusion-basedgeneratorto produce diverse shapes in the form of the coarse coefficient volumes and adetail predictorto produce compatible detail coefficient volumes for introducing fine structures and details. Further, we may jointly train anencoder networkto learn a latent space for inverting shapes, allowing us to enable a rich variety of whole-shape and region-aware shape manipulations. Both quantitative and qualitative experimental results manifest the compelling shape generation, inversion, and manipulation capabilities of our approach over the state-of-the-art methods. Jingyu Hu 0001, Ka-Hei Hui, Zhengzhe Liu, Ruihui Li, Chi-Wing Fu |
ACM Trans. Graph. | 4 |
| 2023 | ISS: Image as Stepping Stone for Text-Guided 3D Shape Generation
Zhengzhe Liu, Peng Dai 0003, Ruihui Li, Xiaojuan Qi 0001, Chi-Wing Fu |
ICLR | 3 |
| 2023 | SD-Net: Spatially-Disentangled Point Cloud Completion NetworkabstractPoint clouds obtained from 3D scanning are typically incomplete, noisy, and sparse. Previous completion methods aim to generate complete point clouds, while taking into account the densification of point clouds, filling small holes, and proximity-to-surface, all through a single network. After revisiting the task, we propose SDNet, which disentangles the task based on the spatial characteristics of point clouds and formulates two sub-networks, a Dense Refiner and a Missing Generator. Given a partial input, the Dense Refiner produces a dense and clean point cloud, as a more reliable partial surface, which assists the Missing Generator to better infer the remaining point cloud structure. To promote the alignment and interaction across these two modules, we propose a Cross Fusion Unit with designed Non-Symmetrical Cross Transformers to capture geometric relationships between partial and missing regions, contributing to a complete, dense and well-aligned output. Extensive quantitative and qualitative results demonstrate that our method outperforms the state-of-the-art methods. Junxian Chen, Ying Liu 0027, Yiqi Liang, Dandan Long, Xiaolin He, Ruihui Li |
ACM Multimedia | 6 |
| 2023 | DreamStone: Image as a Stepping Stone for Text-Guided 3D Shape GenerationabstractThis paper presents a new text-guided 3D shape generation approach DreamStone that uses images as a stepping stone to bridge the gap between the text and shape modalities for generating 3D shapes without requiring paired text and 3D data. The core of our approach is a two-stage feature-space alignment strategy that leverages a pre-trained single-view reconstruction (SVR) model to map CLIP features to shapes: to begin with, map the CLIP image feature to the detail-rich 3D shape space of the SVR model, then map the CLIP text feature to the 3D shape space through encouraging the CLIP-consistency between the rendered images and the input text. Besides, to extend beyond the generative capability of the SVR model, we design the text-guided 3D shape stylization module that can enhance the output shapes with novel structures and textures. Further, we exploit pre-trained text-to-image diffusion models to enhance the generative diversity, fidelity, and stylization capability. Our approach is generic, flexible, and scalable. It can be easily integrated with various SVR models to expand the generative space and improve the generative fidelity. Extensive experimental results demonstrate that our approach outperforms the state-of-the-art methods in terms of generative quality and consistency with the input text. Zhengzhe Liu, Peng Dai 0003, Ruihui Li, Xiaojuan Qi 0001, Chi-Wing Fu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Point Set Self-EmbeddingabstractThis work presents an innovative method for point set self-embedding, that encodes the structural information of a dense point set into its sparser version in a visual but imperceptible form. The self-embedded point set can function as the ordinary downsampled one and be visualized efficiently on mobile devices. Particularly, we can leverage the self-embedded information to fully restore the original point set for detailed analysis on remote servers. This task is challenging, since both the self-embedded point set and the restored point set should resemble the original one. To achieve a learnable self-embedding scheme, we design a novel framework with two jointly-trained networks: one to encode the input point set into its self-embedded sparse point set and the other to leverage the embedded information for inverting the original point set back. Further, we develop a pair of up-shuffle and down-shuffle units in the two networks, and formulate loss terms to encourage the shape similarity and point distribution in the results. Extensive qualitative and quantitative results demonstrate the effectiveness of our method on both synthetic and real-scanned datasets. The source code and trained models will be publicly available at https://github.com/liruihui/Self-Embedding. Ruihui Li, Xianzhi Li 0001, Tien-Tsin Wong, Chi-Wing Fu |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2022 | Neural Template: Topology-aware Reconstruction and Disentangled Generation of 3D MeshesabstractThis paper introduces a novel framework called DTNet for 3D mesh reconstruction and generation via Disentangled Topology. Beyond previous works, we learn a topology-aware neural template specific to each input then deform the template to reconstruct a detailed mesh while preserving the learned topology. One key insight is to decouple the complex mesh reconstruction into two sub-tasks: topology formulation and shape deformation. Thanks to the decoupling, DT-Net implicitly learns a disentangled representation for the topology and shape in the latent space. Hence, it can enable novel disentangled controls for supporting various shape generation applications, e.g., remix the topologies of 3D objects, that are not achievable by previous reconstruction works. Extensive experimental results demonstrate that our method11Code available at https://github.com/edward1997104/Neural-Template. is able to produce high-quality meshes, particularly with diverse topologies, as compared with the state-of-the-art methods. Ka-Hei Hui, Ruihui Li, Jingyu Hu 0001, Chi-Wing Fu |
CVPR | 2 |
| 2022 | PC2-PU: Patch Correlation and Point Correlation for Effective Point Cloud UpsamplingabstractPoint cloud upsampling is to densify a sparse point set acquired from 3D sensors, providing a denser representation for the underlying surface. Existing methods divide the input points into small patches and upsample each patch separately, however, ignoring the global spatial consistency between patches. In this paper, we present a novel method PC$^2$-PU, which explores patch-to-patch and point-to-point correlations for more effective and robust point cloud upsampling. Specifically, our network has two appealing designs: (i) We take adjacent patches as supplementary inputs to compensate the loss structure information within a single patch and introduce a Patch Correlation Module to capture the difference and similarity between patches. (ii) After augmenting each patch's geometry, we further introduce a Point Correlation Module to reveal the relationship of points inside each patch to maintain the local spatial consistency. Extensive experiments on both synthetic and real scanned datasets demonstrate that our method surpasses previous upsampling methods, particularly with the noisy inputs. The code and data are at: https://github.com/chenlongwhu/PC2-PU.git. Chen Long, Ruihui Li, Hao Wang 0057, Zhen Dong 0005, Bisheng Yang |
ACM Multimedia | 3 |
| 2022 | Neural Wavelet-domain Diffusion for 3D Shape GenerationabstractThis paper presents a new approach for 3D shape generation, enabling direct generative modeling on a continuous implicit representation in wavelet domain. Specifically, we propose a compact wavelet representation with a pair of coarse and detail coefficient volumes to implicitly represent 3D shapes via truncated signed distance functions and multi-scale biorthogonal wavelets, and formulate a pair of neural networks: a generator based on the diffusion model to produce diverse shapes in the form of coarse coefficient volumes; and a detail predictor to further produce compatible detail coefficient volumes for enriching the generated shapes with fine structures and details. Both quantitative and qualitative experimental results manifest the superiority of our approach in generating diverse and high-quality shapes with complex topology and structures, clean surfaces, and fine details, exceeding the 3D generation capabilities of the state-of-the-art models. Ka-Hei Hui, Ruihui Li, Jingyu Hu 0001, Chi-Wing Fu |
SIGGRAPH Asia | 2 |
| 2022 | Inferring RNA-binding protein target preferences using adversarial domain adaptationabstractPrecise identification of target sites of RNA-binding proteins (RBP) is important to understand their biochemical and cellular functions. A large amount of experimental data is generated by in vivo and in vitro approaches. The binding preferences determined from these platforms share similar patterns but there are discernable differences between these datasets. Computational methods trained on one dataset do not always work well on another dataset. To address this problem which resembles the classic "domain shift" in deep learning, we adopted the adversarial domain adaptation (ADDA) technique and developed a framework (RBP-ADDA) that can extract RBP binding preferences from an integration of in vivo and vitro datasets. Compared with conventional methods, ADDA has the advantage of working with two input datasets, as it trains the initial neural network for each dataset individually, projects the two datasets onto a feature space, and uses an adversarial framework to derive an optimal network that achieves an optimal discriminative predictive power. In the first step, for each RBP, we include only the in vitro data to pre-train a source network and a task predictor. Next, for the same RBP, we initiate the target network by using the source network and use adversarial domain adaptation to update the target network using both in vitro and in vivo data. These two steps help leverage the in vitro data to improve the prediction on in vivo data, which is typically challenging with a lower signal-to-noise ratio. Finally, to further take the advantage of the fused source and target data, we fine-tune the task predictor using both data. We showed that RBP-ADDA achieved better performance in modeling in vivo RBP binding data than other existing methods as judged by Pearson correlations. It also improved predictive performance on in vitro datasets. We further applied augmentation operations on RBPs with less in vivo data to expand the input data and showed that it can improve prediction performances. Lastly, we explored the predictive interpretability of RBP-ADDA, where we quantified the contribution of the input features by Integrated Gradients and identified nucleotide positions that are important for RBP recognition. Ying Liu 0027, Ruihui Li, Jiawei Luo 0001, Zhaolei Zhang |
PLoS Comput. Biol. | 2 |
| 2022 | A Rotation-Invariant Framework for Deep Point Cloud AnalysisabstractRecently, many deep neural networks were designed to process 3D point clouds, but a common drawback is that rotation invariance is not ensured, leading to poor generalization to arbitrary orientations. In this article, we introduce a new low-level purely rotation-invariant representation to replace common 3D Cartesian coordinates as the network inputs. Also, we present a network architecture to embed these representations into features, encoding local relations between points and their neighbors, and the global shape structure. To alleviate inevitable global information loss caused by the rotation-invariant representations, we further introduce a region relation convolution to encode local and non-local information. We evaluate our method on multiple point cloud analysis tasks, including (i) shape classification, (ii) part segmentation, and (iii) shape retrieval. Extensive experimental results show that our method achieves consistent, and also the best performance, on inputs at arbitrary orientations, compared with all the state-of-the-art methods. Xianzhi Li 0001, Ruihui Li, Guangyong Chen, Chi-Wing Fu, Daniel Cohen-Or, Pheng-Ann Heng |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | Point Cloud Upsampling via Disentangled RefinementabstractPoint clouds produced by 3D scanning are often sparse, non-uniform, and noisy. Recent upsampling approaches aim to generate a dense point set, while achieving both distribution uniformity and proximity-to-surface, and possibly amending small holes, all in a single network. After revisiting the task, we propose to disentangle the task based on its multi-objective nature and formulate two cascaded sub-networks, a dense generator and a spatial refiner. The dense generator infers a coarse but dense out-put that roughly describes the underlying surface, while the spatial refiner further fine-tunes the coarse output by adjusting the location of each point. Specifically, we design a pair of local and global refinement units in the spatial refiner to evolve a coarse feature map. Also, in the spatial refiner, we regress a per-point offset vector to further adjust the coarse outputs in fine scale. Extensive qualitative and quantitative results on both synthetic and real-scanned datasets demonstrate the superiority of our method over the state-of-the-arts. The code is publicly available at https://github.com/liruihui/Dis-PU. Ruihui Li, Xianzhi Li 0001, Pheng-Ann Heng, Chi-Wing Fu |
CVPR | 1 |
| 2021 | SP-GAN: sphere-guided 3D shape generation and manipulationabstractWe present SP-GAN, a new unsupervised sphere-guided generative model for direct synthesis of 3D shapes in the form of point clouds. Compared with existing models, SP-GAN is able to synthesize diverse and high-quality shapes with fine details and promote controllability for part-aware shape generation and manipulation, yet trainable without any parts annotations. In SP-GAN, we incorporate a global prior (uniform points on a sphere) to spatially guide the generative process and attach a local prior (a random latent code) to each sphere point to provide local details. The key insight in our design is to disentangle the complex 3D shape generation task into a global shape modeling and a local structure adjustment, to ease the learning process and enhance the shape generation quality. Also, our model forms an implicit dense correspondence between the sphere points and points in every generated shape, enabling various forms of structure-aware shape manipulations such as part editing, part-wise shape interpolation, and multi-shape part composition, etc., beyond the existing generative models. Experimental results, which include both visual and quantitative evaluations, demonstrate that our model is able to synthesize diverse point clouds with fine details and less noise, as compared with the state-of-the-art models. Ruihui Li, Xianzhi Li 0001, Ka-Hei Hui, Chi-Wing Fu |
ACM Trans. Graph. | 1 |
| 2021 | DNF-Net: A Deep Normal Filtering Network for Mesh DenoisingabstractThis article presents a deep normal filtering network, called DNF-Net, for mesh denoising. To better capture local geometry, our network processes the mesh in terms of local patches extracted from the mesh. Overall, DNF-Net is an end-to-end network that takes patches of facet normals as inputs and directly outputs the corresponding denoised facet normals of the patches. In this way, we can reconstruct the geometry from the denoised normals with feature preservation. Besides the overall network architecture, our contributions include a novel multi-scale feature embedding unit, a residual learning strategy to remove noise, and a deeply-supervised joint loss function. Compared with the recent data-driven works on mesh denoising, DNF-Net does not require manual input to extract features and better utilizes the training data to enhance its denoising performance. Finally, we present comprehensive experiments to evaluate our method and demonstrate its superiority over the state of the art on both synthetic and real-scanned meshes. Xianzhi Li 0001, Ruihui Li, Lei Zhu 0003, Chi-Wing Fu, Pheng-Ann Heng |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2020 | PointAugment: An Auto-Augmentation Framework for Point Cloud ClassificationabstractWe present PointAugment, a new auto-augmentation framework that automatically optimizes and augments point cloud samples to enrich the data diversity when we train a classification network. Different from existing auto-augmentation methods for 2D images, PointAugment is sample-aware and takes an adversarial learning strategy to jointly optimize an augmentor network and a classifier network, such that the augmentor can learn to produce augmented samples that best fit the classifier. Moreover, we formulate a learnable point augmentation function with a shape-wise transformation and a point-wise displacement, and carefully design loss functions to adopt the augmented samples based on the learning progress of the classifier. Extensive experiments also confirm PointAugment's effectiveness and robustness to improve the performance of various networks on shape classification and retrival. Ruihui Li, Xianzhi Li 0001, Pheng-Ann Heng, Chi-Wing Fu |
CVPR | 1 |
| 2019 | PU-GAN: A Point Cloud Upsampling Adversarial NetworkabstractPoint clouds acquired from range scans are often sparse, noisy, and non-uniform. This paper presents a new point cloud upsampling network called PU-GAN1, which is formulated based on a generative adversarial network (GAN), to learn a rich variety of point distributions from the latent space and upsample points over patches on object surfaces. To realize a working GAN network, we construct an up-down-up expansion unit in the generator for upsampling point features with error feedback and self-correction, and formulate a self-attention unit to enhance the feature integration. Further, we design a compound loss with adversarial, uniform and reconstruction terms, to encourage the discriminator to learn more latent patterns and enhance the output point distribution uniformity. Qualitative and quantitative evaluations demonstrate the quality of our results over the state-of-the-arts in terms of distribution uniformity, proximity-to-surface, and 3D reconstruction quality. Ruihui Li, Xianzhi Li 0001, Chi-Wing Fu, Daniel Cohen-Or, Pheng-Ann Heng |
ICCV | 1 |
| 2017 | Aggregating complementary boundary contrast with smoothing for salient region detection
Ruihui Li, Jianrui Cai, Hanling Zhang, Taihong Wang |
Vis. Comput. | 1 |