Ruijie Lu

dblp:125/9394 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 MOVIS: Enhancing Multi-Object Novel View Synthesis for Indoor Scenes
abstract
Repurposing pre-trained diffusion models has been proven to be effective for NVS. However, these methods are mostly limited to a single object; directly applying such methods to compositional multi-object scenarios yields inferior results, especially incorrect object placement and inconsistent shape and appearance under novel views. How to enhance and systematically evaluate the cross-view consistency of such models remains under-explored. To address this issue, we propose MOVIS to enhance the structural awareness of the view-conditioned diffusion model for multi-object NVS in terms of model inputs, auxiliary tasks, and training strategy. First, we inject structure-aware features, including depth and object mask, into the denoising U-Net to enhance the model’s comprehension of object instances and their spatial relationships. Second, we introduce an auxiliary task requiring the model to simultaneously predict novel view object masks, further improving the model’s capability in differentiating and placing objects. Finally, we conduct an in-depth analysis of the diffusion sampling process and carefully devise a structure-guided timestep sampling scheduler during training, which balances the learning of global object placement and fine-grained detail recovery. To systematically evaluate the plausibility of synthesized images, we propose to assess cross-view consistency and novel view object placement alongside existing image-level NVS metrics. Extensive experiments on challenging synthetic and realistic datasets demonstrate that our method exhibits strong generalization capabilities and produces consistent novel view synthesis, highlighting its potential to guide future 3D-aware multi-object NVS tasks. Our project page is available at https://jason-aplp.github.io/MOVIS/.
Ruijie Lu, Yixin Chen 0003, Junfeng Ni, Baoxiong Jia, Yu Liu 0110, Diwen Wan, Siyuan Huang 0001
CVPR1
2025 Decompositional Neural Scene Reconstruction with Generative Diffusion Prior
abstract
Decompositional reconstruction of 3D scenes, with complete shapes and detailed texture of all objects within, is intriguing for downstream applications but remains challenging, particularly with sparse views as input. Recent approaches incorporate semantic or geometric regularization to address this issue, but they suffer significant degradation in underconstrained areas and fail to recover occluded regions. We argue that the key to solving this problem lies in supplementing missing information for these areas. To this end, we propose DP-Recon, which employs diffusion priors in the form of Score Distillation Sampling (SDS) to optimize the neural representation of each individual object under novel views. This provides additional information for the underconstrained areas, but directly incorporating diffusion prior raises potential conflicts between the reconstruction and generative guidance. Therefore, we further introduce a visibility-guided approach to dynamically adjust the per-pixel SDS loss weights. Together these components enhance both geometry and appearance recovery while remaining faithful to input images. Extensive experiments across Replica and ScanNet++ demonstrate that our method significantly outperforms state-of-the-art methods. Notably, it achieves better object reconstruction under 10 views than the baselines under 100 views. Our method enables seamless text-based editing for geometry and appearance through SDS optimization and produces decomposed object meshes with detailed UV maps that support photo-realistic Visual effects (VFX) editing. The project page is available at https://dp-recon.github.io/.
Junfeng Ni, Yu Liu 0110, Ruijie Lu, Zirui Zhou, Song-Chun Zhu, Yixin Chen 0003, Siyuan Huang 0001
CVPR3
2025 TACO: Taming Diffusion for In-the-Wild Video Amodal Completion
Ruijie Lu, Yixin Chen 0003, Yu Liu 0110, Jiaxiang Tang, Junfeng Ni, Diwen Wan, Siyuan Huang 0001
ICCV1
2025 Building Interactable Replicas of Complex Articulated Objects via Gaussian Splatting
abstract
Building interactable replicas of articulated objects is a key challenge in computer vision. Existing methods often fail to effectively integrate information across different object states, limiting the accuracy of part-mesh reconstruction and part dynamics modeling, particularly for complex multi-part articulated objects. We introduce ArtGS, a novel approach that leverages 3D Gaussians as a flexible and efficient representation to address these issues. Our method incorporates canonical Gaussians with coarse-to-fine initialization and updates for aligning articulated part information across different object states, and employs a skinning-inspired part dynamics modeling module to improve both part-mesh reconstruction and articulation learning. Extensive experiments on both synthetic and real-world datasets, including a new benchmark for complex multi-part objects, demonstrate that ArtGS achieves state-of-the-art performance in joint parameter estimation and part mesh reconstruction. Our approach significantly improves reconstruction quality and efficiency, especially for multi-part articulated objects. Additionally, we provide comprehensive analyses of our design choices, validating the effectiveness of each component to highlight potential areas for future improvement.
Yu Liu 0110, Baoxiong Jia, Ruijie Lu, Junfeng Ni, Song-Chun Zhu, Siyuan Huang 0001
ICLR3
2025 Fast SP-GS: Reconstructing Dynamic Scenes in Minutes
abstract
Despite recent advances in Gaussian Splatting techniques-such as Superpoint Gaussian Splatting (SP-GS), which enables real-time, high-fidelity rendering-3D reconstruction of dynamic scenes remains a significant challenge in computer vision. However, SP-GS requires nearly an hour for dynamic scene optimization, severely limiting its practical applications in AR and VR. To address this limitation, we propose Fast SP-GS, an efficient approach that reduces training time to mere minutes. Building upon acceleration methods for static scenes (e.g., Mini-Splatting, Taming 3DGS, FlashGS), our novel 2D-GS-based framework enhances speed and quality via three key innovations: First, an aggressive 2D-GS densification strategy reduces required training iterations, while a Gaussian simplification strategy minimizes redundant parameters. Second, a novel 2D-GS optical flow loss provides explicit motion supervision, accelerating convergence. Third, an optimized CUDA implementation maximizes rendering efficiency. Extensive experiments on synthetic and real-world datasets confirm that Fast SP-GS reconstructs dynamic scenes in minutes, surpassing SP-GS in both rendering quality and computational efficiency. The source code is available at https://github.com/dnvtmf/Fast-SP-GS.
Diwen Wan, Jiaxiang Tang, Ruijie Lu, Yuxiang Wang 0015
ISMAR3
2025 Efficient Part-level 3D Object Generation via Dual Volume Packing
abstract
Recent progress in 3D object generation has greatly improved both the quality and efficiency. However, most existing methods generate a single mesh with all parts fused together, which limits the ability to edit or manipulate individual parts. A key challenge is that different objects may have a varying number of parts. To address this, we propose a new end-to-end framework for part-level 3D object generation. Given a single input image, our method generates high-quality 3D objects with an arbitrary number of complete and semantically meaningful parts. We introduce a dual volume packing strategy that organizes all parts into two complementary volumes, allowing for the creation of complete and interleaved parts that assemble into the final object. Experiments show that our model achieves better quality, diversity, and generalization than previous image-based part-level generation methods. Our project page is at \url{https://research.nvidia.com/labs/dir/partpacker/}.
Jiaxiang Tang, Ruijie Lu, Max Li, Zekun Hao, Xuan Li 0015, Fangyin Wei, Shuran Song, Ming-Yu Liu 0001, Tsung-Yi Lin
NeurIPS2
2025 Generating Objects with Part-Articulation from a Single Image
abstract
Generating articulated objects, such as laptops and microwaves, is a crucial yet challenging task with extensive applications in Embodied AI and AR/VR. Current image-to-3D methods primarily focus on surface geometry and texture, neglecting part decomposition and articulation modeling. Meanwhile, neural reconstruction approaches (e.g., NeRF or Gaussian Splatting) rely on dense multi-view or interaction data, limiting their scalability. In this paper, we introduce DreamArt, a novel framework for generating high-fidelity, interactable articulated assets from single-view images. DreamArt employs a three-stage pipeline: firstly, it reconstructs part‑segmented and complete 3D object meshes through a combination of image-to-3D generation, mask-prompted 3D segmentation, and part amodal completion. Second, we fine-tune a video diffusion model to capture part-level articulation priors, leveraging movable part masks as prompt and amodal images to mitigate ambiguities caused by occlusion. Finally, DreamArt optimizes the articulation motion, represented by a dual quaternion, and conducts global texture refinement and repainting to ensure coherent, high-quality textures across all parts. Experimental results demonstrate that DreamArt effectively generates high-quality articulated objects, possessing accurate part shape, high appearance fidelity, and plausible articulation, thereby providing a scalable solution for articulated asset generation.
Ruijie Lu, Yu Liu 0110, Jiaxiang Tang, Junfeng Ni, Yuxiang Wang 0015, Diwen Wan, Yixin Chen 0003, Siyuan Huang 0001
SIGGRAPH Asia1
2024 RoomTex: Texturing Compositional Indoor Scenes via Iterative Inpainting
Qi Wang 0105, Ruijie Lu, Xudong Xu, Jingbo Wang 0003, Michael Yu Wang, Bo Dai 0002, Dan Xu 0002
ECCV (68)2
2024 Superpoint Gaussian Splatting for Real-Time High-Fidelity Dynamic Scene Reconstruction
abstract
Rendering novel view images in dynamic scenes is a crucial yet challenging task. Current methods mainly utilize NeRF-based methods to represent the static scene and an additional time-variant MLP to model scene deformations, resulting in relatively low rendering quality as well as slow inference speed. To tackle these challenges, we propose a novel framework named Superpoint Gaussian Splatting (SP-GS). Specifically, our framework first employs explicit 3D Gaussians to reconstruct the scene and then clusters Gaussians with similar properties (e.g., rotation, translation, and location) into superpoints. Empowered by these superpoints, our method manages to extend 3D Gaussian splatting to dynamic scenes with only a slight increase in computational expense. Apart from achieving state-of-the-art visual quality and real-time rendering under high resolutions, the superpoint representation provides a stronger manipulation capability. Extensive experiments demonstrate the practicality and effectiveness of our approach on both synthetic and real-world datasets. Please see our project page at https://dnvtmf.github.io/SP_GS.github.io.
Diwen Wan, Ruijie Lu
ICML2
2024 Template-free Articulated Gaussian Splatting for Real-time Reposable Dynamic View Synthesis
abstract
While novel view synthesis for dynamic scenes has made significant progress, capturing skeleton models of objects and re-posing them remains a challenging task. To tackle this problem, in this paper, we propose a novel approach to automatically discover the associated skeleton model for dynamic objects from videos without the need for object-specific templates. Our approach utilizes 3D Gaussian Splatting and superpoints to reconstruct dynamic objects. Treating superpoints as rigid parts, we can discover the underlying skeleton model through intuitive cues and optimize it using the kinematic model. Besides, an adaptive control strategy is applied to avoid the emergence of redundant superpoints. Extensive experiments demonstrate the effectiveness and efficiency of our method in obtaining re-posable 3D objects. Not only can our approach achieve excellent visual fidelity, but it also allows for the real-time rendering of high-resolution images.
Diwen Wan, Yuxiang Wang 0015, Ruijie Lu
NeurIPS3
2005 Tree-ring precipitation records since 1860 at Changling Mountain, China
Shangyu Gao, Ruijie Lu, Hong Xia, Mingrui Qiang, Dengshan Zhang
IGARSS2
2005 Study on land cover change detection method based on NDVI time series batasets: change detection indexes design
abstract
The normalized difference vegetation index (NDVI) time-series database, derived from NOAA/AVHRR, SPOT/VEGETATION, TERRA or AQUA/MODIS, is increasingly being recognized as a valuable data source for extracting land cover and its change information at global, continental and large regional scale. However, existing approaches, such as principal component analysis (PCA) and change vector analysis (CVA) present considerable difficulties in taking full advantage the NDVI dataset for land cover change detection. Based on the assumptions that different land cover types have different NDVI temporal profiles and that the NDVI profile curve can be regarded as a spectrum in which an NDVI value for a certain date corresponds to on band value of this spectrum, we analyzed the existing change detection indexes and develop a new land cover change detection method based on Lance distance and a cross correlogram spectral matching (CCSM) technique. The new method was validated in the simulation experiments and a case study area of Beijing. From the results, we have demonstrated that the new method takes the shape and value features of NDVI profile curve into consideration. The relatively better performance of the new method can be attributed to two advantages: (1) the new method can discriminate long-term land cover changes form other changes by excluding false changes caused by vegetation phenology changes, climate events, atmospheric variability and sensor noise; (2) it is similarly sensitive to all kinds of land cover changes no matter where the changes have occurred. The better results compared with the CVA method suggest that the new method is effective and has potential for land cover change detection using an NDVI time-series dataset. Furthermore, it is worth noting that the method can not only be applied to NDVI datasets but also to other index datasets reflection surface conditions sampled at different time interval. It can also be applied to datasets for different satellites without the need to normalized sensor differences.
Jin Chen 0001, Ruijie Lu, Tianxiang Yue
IGARSS3
2005 An effective approach to remove cloud-fog cover and enhance remotely sensed imagery
abstract
A new approach is proposed to remove cloud-fog cover and to enhance satellite imagery based on Laplacian enhancement and histogram transformation. The new approach is also compared with other two methods: homomorphism filter and histogram match. Information entropy and spectral correlation coefficient are used to evaluate the results. A multi- spectral SPOT image, which is contaminated by thin cloud-fog noises, is selected for the research. The results from experimental tests show that the new approach is better than the two others, not only removing cloud-fog cover and enhangcing the spatial information of images but also keeping the true spectral feature of the images.
Jin Chen 0001, Ruijie Lu
IGARSS4