Diwen Wan

dblp:227/6394 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
7since 2021 · last 2025
0000-0002-3640-0511ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 MOVIS: Enhancing Multi-Object Novel View Synthesis for Indoor Scenes
abstract
Repurposing pre-trained diffusion models has been proven to be effective for NVS. However, these methods are mostly limited to a single object; directly applying such methods to compositional multi-object scenarios yields inferior results, especially incorrect object placement and inconsistent shape and appearance under novel views. How to enhance and systematically evaluate the cross-view consistency of such models remains under-explored. To address this issue, we propose MOVIS to enhance the structural awareness of the view-conditioned diffusion model for multi-object NVS in terms of model inputs, auxiliary tasks, and training strategy. First, we inject structure-aware features, including depth and object mask, into the denoising U-Net to enhance the model’s comprehension of object instances and their spatial relationships. Second, we introduce an auxiliary task requiring the model to simultaneously predict novel view object masks, further improving the model’s capability in differentiating and placing objects. Finally, we conduct an in-depth analysis of the diffusion sampling process and carefully devise a structure-guided timestep sampling scheduler during training, which balances the learning of global object placement and fine-grained detail recovery. To systematically evaluate the plausibility of synthesized images, we propose to assess cross-view consistency and novel view object placement alongside existing image-level NVS metrics. Extensive experiments on challenging synthetic and realistic datasets demonstrate that our method exhibits strong generalization capabilities and produces consistent novel view synthesis, highlighting its potential to guide future 3D-aware multi-object NVS tasks. Our project page is available at https://jason-aplp.github.io/MOVIS/.
Ruijie Lu, Yixin Chen 0003, Junfeng Ni, Baoxiong Jia, Yu Liu 0110, Diwen Wan, Siyuan Huang 0001
CVPR6
2025 TACO: Taming Diffusion for In-the-Wild Video Amodal Completion
Ruijie Lu, Yixin Chen 0003, Yu Liu 0110, Jiaxiang Tang, Junfeng Ni, Diwen Wan, Siyuan Huang 0001
ICCV6
2025 Fast SP-GS: Reconstructing Dynamic Scenes in Minutes
abstract
Despite recent advances in Gaussian Splatting techniques-such as Superpoint Gaussian Splatting (SP-GS), which enables real-time, high-fidelity rendering-3D reconstruction of dynamic scenes remains a significant challenge in computer vision. However, SP-GS requires nearly an hour for dynamic scene optimization, severely limiting its practical applications in AR and VR. To address this limitation, we propose Fast SP-GS, an efficient approach that reduces training time to mere minutes. Building upon acceleration methods for static scenes (e.g., Mini-Splatting, Taming 3DGS, FlashGS), our novel 2D-GS-based framework enhances speed and quality via three key innovations: First, an aggressive 2D-GS densification strategy reduces required training iterations, while a Gaussian simplification strategy minimizes redundant parameters. Second, a novel 2D-GS optical flow loss provides explicit motion supervision, accelerating convergence. Third, an optimized CUDA implementation maximizes rendering efficiency. Extensive experiments on synthetic and real-world datasets confirm that Fast SP-GS reconstructs dynamic scenes in minutes, surpassing SP-GS in both rendering quality and computational efficiency. The source code is available at https://github.com/dnvtmf/Fast-SP-GS.
Diwen Wan, Jiaxiang Tang, Ruijie Lu, Yuxiang Wang 0015
ISMAR1
2025 Generating Objects with Part-Articulation from a Single Image
abstract
Generating articulated objects, such as laptops and microwaves, is a crucial yet challenging task with extensive applications in Embodied AI and AR/VR. Current image-to-3D methods primarily focus on surface geometry and texture, neglecting part decomposition and articulation modeling. Meanwhile, neural reconstruction approaches (e.g., NeRF or Gaussian Splatting) rely on dense multi-view or interaction data, limiting their scalability. In this paper, we introduce DreamArt, a novel framework for generating high-fidelity, interactable articulated assets from single-view images. DreamArt employs a three-stage pipeline: firstly, it reconstructs part‑segmented and complete 3D object meshes through a combination of image-to-3D generation, mask-prompted 3D segmentation, and part amodal completion. Second, we fine-tune a video diffusion model to capture part-level articulation priors, leveraging movable part masks as prompt and amodal images to mitigate ambiguities caused by occlusion. Finally, DreamArt optimizes the articulation motion, represented by a dual quaternion, and conducts global texture refinement and repainting to ensure coherent, high-quality textures across all parts. Experimental results demonstrate that DreamArt effectively generates high-quality articulated objects, possessing accurate part shape, high appearance fidelity, and plausible articulation, thereby providing a scalable solution for articulated asset generation.
Ruijie Lu, Yu Liu 0110, Jiaxiang Tang, Junfeng Ni, Yuxiang Wang 0015, Diwen Wan, Yixin Chen 0003, Siyuan Huang 0001
SIGGRAPH Asia6
2024 Open-set Hierarchical Semantic Segmentation for 3D Scene
abstract
The Segment-Anything Model (SAM) shows exceptional zero-shot capabilities for 2D images. Developing a similar model for 3D, however, is challenging due to limited datasets. In this paper, we introduce a zero-shot algorithm to segment a 3D scene into elements at various levels of detail, and further organize the results in a hierarchical tree structure. We propose a tree quality metric to evaluate the algorithm’s performance. Notably, our algorithm eliminates the need for 3D annotations. It uses robust 2D models to generate a 2D segmentation tree for each rendered image. Then, using graph neural networks, it aggregates these 2D trees to form a unified 3D segmentation tree. Extensive experiments on the PartNet dataset and complex 3D scenes validate the algorithm’s effectiveness. We release the source code at https://github.com/dnvtmf/OTS.
Diwen Wan, Jiaxiang Tang, Jingbo Wang 0003, Xiaokang Chen, Lingyun Gan
ICME1
2024 Superpoint Gaussian Splatting for Real-Time High-Fidelity Dynamic Scene Reconstruction
abstract
Rendering novel view images in dynamic scenes is a crucial yet challenging task. Current methods mainly utilize NeRF-based methods to represent the static scene and an additional time-variant MLP to model scene deformations, resulting in relatively low rendering quality as well as slow inference speed. To tackle these challenges, we propose a novel framework named Superpoint Gaussian Splatting (SP-GS). Specifically, our framework first employs explicit 3D Gaussians to reconstruct the scene and then clusters Gaussians with similar properties (e.g., rotation, translation, and location) into superpoints. Empowered by these superpoints, our method manages to extend 3D Gaussian splatting to dynamic scenes with only a slight increase in computational expense. Apart from achieving state-of-the-art visual quality and real-time rendering under high resolutions, the superpoint representation provides a stronger manipulation capability. Extensive experiments demonstrate the practicality and effectiveness of our approach on both synthetic and real-world datasets. Please see our project page at https://dnvtmf.github.io/SP_GS.github.io.
Diwen Wan, Ruijie Lu
ICML1
2024 Template-free Articulated Gaussian Splatting for Real-time Reposable Dynamic View Synthesis
abstract
While novel view synthesis for dynamic scenes has made significant progress, capturing skeleton models of objects and re-posing them remains a challenging task. To tackle this problem, in this paper, we propose a novel approach to automatically discover the associated skeleton model for dynamic objects from videos without the need for object-specific templates. Our approach utilizes 3D Gaussian Splatting and superpoints to reconstruct dynamic objects. Treating superpoints as rigid parts, we can discover the underlying skeleton model through intuitive cues and optimize it using the kinematic model. Besides, an adaptive control strategy is applied to avoid the emergence of redundant superpoints. Extensive experiments demonstrate the effectiveness and efficiency of our method in obtaining re-posable 3D objects. Not only can our approach achieve excellent visual fidelity, but it also allows for the real-time rendering of high-resolution images.
Diwen Wan, Yuxiang Wang 0015, Ruijie Lu
NeurIPS1
2020 Controllable Orthogonalization in Training DNNs
abstract
Orthogonality is widely used for training deep neural networks (DNNs) due to its ability to maintain all singular values of the Jacobian close to 1 and reduce redundancy in representation. This paper proposes a computationally efficient and numerically stable orthogonalization method using Newton's iteration (ONI), to learn a layer-wise orthogonal weight matrix in DNNs. ONI works by iteratively stretching the singular values of a weight matrix towards 1. This property enables it to control the orthogonality of a weight matrix by its number of iterations. We show that our method improves the performance of image classification networks by effectively controlling the orthogonality to provide an optimal tradeoff between optimization benefits and representational capacity reduction. We also show that ONI stabilizes the training of generative adversarial networks (GANs) by maintaining the Lipschitz continuity of a network, similar to spectral normalization (SN), and further outperforms SN by providing controllable orthogonality.
Lei Huang 0015, Li Liu 0004, Fan Zhu 0001, Diwen Wan, Zehuan Yuan, Bo Li 0026, Ling Shao 0001
CVPR4
2020 Deep quantization generative networks
Diwen Wan, Fumin Shen, Li Liu 0004, Fan Zhu 0001, Lei Huang 0015, Mengyang Yu, Heng Tao Shen, Ling Shao 0001
Pattern Recognit.1
2018 TBN: Convolutional Neural Network with Ternary Inputs and Binary Weights
Diwen Wan, Fumin Shen, Li Liu 0004, Fan Zhu 0001, Jie Qin 0004, Ling Shao 0001, Heng Tao Shen
ECCV (2)1