Zhiwen Yan

dblp:309/7477 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021
YearPublicationVenuePosition
2026 1000FPS+ Novel View Synthesis from End-To-End Opaque Triangle Optimization
abstract
Recent implicit and primitive-based radiance field methods, such as NeRF and 3DGS, have demonstrated impressive capabilities in novel view synthesis from multi-view images. However, their custom representations are often incompatible with conventional graphics pipelines, limiting their application in areas like editing, relighting, physics simulation, and particle effects. Additionally, their volume rendering approach requires alpha-blending multiple colors per pixel, which slows down rendering. Traditional differentiable rendering methods, while offering higher compatibility and speed, rely on known mesh topology, making them unsuitable for complex scene-level reconstruction. To address these limitations, we introduce a novel end-toend optimization process of disjoint opaque triangles, natively compatible with standard graphics engines for a wide range of applications. To enable gradient-based optimization over the highly non-differentiable rasterization process, we employ a$2 D$SDF approximation and a two-layer occlusion approximation. We also incorporate density controls to ensure detailed and complete scene reconstructions. Our paper tackles the challenging end-to-end optimization of scene-level novel view synthesis with opaque representation only. Our approach achieves over 1000 FPS rendering on a single desktop GPU, providing high compatibility and similar novel view synthesis quality to existing methods.
Zhiwen Yan, Weng Fei Low, Tianxin Huang, Gim Hee Lee
3DV1
2025 Particle Rendering: Implicitly Aggregating Incident and Outgoing Light Fields for Novel View Synthesis
abstract
This paper presents Particle Rendering (PR), a new implicit rendering approach that extends Neural Radiance Fields (NeRF) by incorporating incident light along with traditional outgoing light modeling. In our framework, a 3D scene consists of a mass of particles, each offering a deeper understanding of light interactions by reflecting and emitting light in all directions. Our methodology involves a three-phase training pipeline: 1) Estimating the outgoing light field through a NeRF model; 2) Distilling the incident light field. A simple metric is introduced to assess the quality of the ray for better supervision; 3) Implicit rendering. We propose an implicit method to aggregate incident and outgoing fields that leverages Multilayer Perceptrons (MLP) to directly infer final pixel values, thus avoiding the limitation of traditional physically-based rendering techniques. The effectiveness of PR is demonstrated through state-of-the-art results in various challenging indoor and outdoor scenes, emphasizing its capability to handle complex lighting and reflective materials.
Tao Hu 0011, Zhiwen Yan, Xiaogang Xu 0002, Gim Hee Lee
3DV2
2025 OD-NeRF: Efficient Training of On-the-Fly Dynamic Neural Radiance Fields
abstract
Dynamic neural radiance fields (dynamic NeRFs) have achieved remarkable successes in synthesizing novel views for 3D dynamic scenes. Traditional approaches typically necessitate full video sequences for the training phase prior to the synthesis of new views, akin to replaying a recording of a dynamic 3D event. In contrast, on-the-fly training allows for the immediate processing and rendering of dynamic scenes without the need for pre-training on full sequences, offering a more flexible and time-efficient solution for dynamic scene rendering tasks. In this paper, we propose a highly efficient on-the-fly training algorithm for dynamic NeRFs, named OD-NeRF. To accelerate the training process, our method minimizes the training required for the model at each frame by using: 1) a NeRF model conditioned on multi-view projected colors, which exhibits superior generalization across multiple frames with minimal training, and 2) a transition and update algorithm that leverages the occupancy grid from the last frame to sample efficiently at the current frame. Our algorithm can achieve an interactive training speed of 10FPS on synthetic dynamic scenes on-the-fly, and a 3×-9× training speed-up compared to the state-of-the-art on-the-fly NeRF on real-world dynamic scenes.
Zhiwen Yan, Chen Li 0038, Gim Hee Lee
3DV1
2025 ComPC: Completing a 3D Point Cloud with 2D Diffusion Priors
abstract
3D point clouds directly collected from objects through sensors are often incomplete due to self-occlusion. Conventional methods for completing these partial point clouds rely on manually organized training sets and are usually limited to object categories seen during training. In this work, we propose a test-time framework for completing partial point clouds across unseen categories without any requirement for training. Leveraging point rendering via Gaussian Splatting, we develop techniques of Partial Gaussian Initialization, Zero-shot Fractal Completion, and Point Cloud Extraction that utilize priors from pre-trained 2D diffusion models to infer missing regions and extract uniform completed point clouds. Experimental results on both synthetic and real-world scanned point clouds demonstrate that our approach outperforms existing methods in completing a variety of objects. Our project page is at \url{https://tianxinhuang.github.io/projects/ComPC/}.
Tianxin Huang, Zhiwen Yan, Gim Hee Lee
ICLR2
2025 GenXD: Generating Any 3D and 4D Scenes
abstract
Recent developments in 2D visual generation have been remarkably successful. However, 3D and 4D generation remain challenging in real-world applications due to the lack of large-scale 4D data and effective model design. In this paper, we propose to jointly investigate general 3D and 4D generation by leveraging camera and object movements commonly observed in daily life. Due to the lack of real-world 4D data in the community, we first propose a data curation pipeline to obtain camera poses and object motion strength from videos. Based on this pipeline, we introduce a large-scale real-world 4D scene dataset: CamVid-30K. By leveraging all the 3D and 4D data, we develop our framework, GenXD, which allows us to produce any 3D or 4D scene. We propose multiview-temporal modules, which disentangle camera and object movements, to seamlessly learn from both 3D and 4D data. Additionally, GenXD employs masked latent conditions to support a variety of conditioning views. GenXD can generate videos that follow the camera trajectory as well as consistent 3D views that can be lifted into 3D representations. We perform extensive evaluations across various real-world and synthetic datasets, demonstrating GenXD's effectiveness and versatility compared to previous methods in 3D and 4D generation.
Chung-Ching Lin, Zhiwen Yan, Zhengyuan Yang, Gim Hee Lee
ICLR4
2024 Multi-Scale 3D Gaussian Splatting for Anti-Aliased Rendering
abstract
3D Gaussians have recently emerged as a highly efficient representation for 3D reconstruction and rendering. Despite its high rendering quality and speed at high resolutions, they both deteriorate drastically when rendered at lower resolutions or from far away camera position. During low resolution or far away rendering, the pixel size of the image can fall below the Nyquist frequency compared to the screen size of each splatted 3D Gaussian and leads to aliasing effect. The rendering is also drastically slowed down by the sequential alpha blending of more splatted Gaussians per pixel. To address these issues, we propose a multi-scale 3D Gaussian splatting algorithm, which maintains Gaussians at different scales to represent the same scene. Higher-resolution images are rendered with more small Gaussians, and lower-resolution images are rendered with fewer larger Gaussians. With similar training time, our algorithm can achieve 13%-66% PSNR and 160%-2400% rendering speed improvement at 4 × -128 × scale rendering on Mip-NeRF360 dataset compared to the single scale 3D Gaussian splatting. More results and code are released on our project page.
Zhiwen Yan, Weng Fei Low, Gim Hee Lee
CVPR1
2024 Fashion Image Retrieval with Occlusion
Jimin Sohn, Haeji Jung, Zhiwen Yan, Vibha Masti, Bhiksha Raj
ICPR (21)3
2023 NeRF-DS: Neural Radiance Fields for Dynamic Specular Objects
abstract
Dynamic Neural Radiance Field (NeRF) is a powerful algorithm capable of rendering photo-realistic novel view images from a monocular RGB video of a dynamic scene. Although it warps moving points across frames from the observation spaces to a common canonical space for rendering, dynamic NeRF does not model the change of the reflected color during the warping. As a result, this approach often fails drastically on challenging specular objects in motion. We address this limitation by reformulating the neural radiance field function to be conditioned on surface position and orientation in the observation space. This allows the specular surface at different poses to keep the different reflected colors when mapped to the common canonical space. Additionally, we add the mask of moving objects to guide the deformation field. As the specular surface changes color during motion, the mask mitigates the problem of failure to find temporal correspondences with only RGB supervision. We evaluate our model based on the novel view synthesis quality with a self-collected dataset of different moving specular objects in realistic environments. The experimental results demonstrate that our method significantly improves the reconstruction quality of moving specular objects from monocular RGB videos compared to the existing NeRF models. Our code and data are available at the project website11https://github.com/JokerYan/NeRF-DS.
Zhiwen Yan, Chen Li 0038, Gim Hee Lee
CVPR1
2023 GNeSF: Generalizable Neural Semantic Fields
abstract
3D scene segmentation based on neural implicit representation has emerged recently with the advantage of training only on 2D supervision. However, existing approaches still requires expensive per-scene optimization that prohibits generalization to novel scenes during inference. To circumvent this problem, we introduce a \textit{generalizable} 3D segmentation framework based on implicit representation. Specifically, our framework takes in multi-view image features and semantic maps as the inputs instead of only spatial information to avoid overfitting to scene-specific geometric and semantic information. We propose a novel soft voting mechanism to aggregate the 2D semantic information from different views for each 3D point. In addition to the image features, view difference information is also encoded in our framework to predict the voting scores. Intuitively, this allows the semantic information from nearby views to contribute more compared to distant ones. Furthermore, a visibility module is also designed to detect and filter out detrimental information from occluded views. Due to the generalizability of our proposed method, we can synthesize semantic maps or conduct 3D semantic segmentation for novel scenes with solely 2D semantic supervision. Experimental results show that our approach achieves comparable performance with scene-specific approaches. More importantly, our approach can even outperform existing strong supervision-based approaches with only 2D annotations.
Chen Li 0038, Mengqi Guo, Zhiwen Yan, Gim Hee Lee
NeurIPS4
2022 Energy and Spectrum Efficient Radio Frequency Fingerprint Intelligent Blind Identification
abstract
Radio frequency fingerprint identification (RFFI) technology identifies the emitter by extracting one or more unintentional features of the signal from the emitter. To solve the problem that the traditional deep learning network is not highly adaptable for the contour features extracted from the signal, this paper proposes a novel RFFI method based on a deformable convolutional network. This network makes the convolution operation more biased towards the useful information content in the feature map with higher energy, and ignores part of the background noise information. The proposed blind identification method requires less information and no training sequences and pilots, Thus, it achieves energy and spectrum efficient radio communications. Simulation verifies that the proposed method can achieve better recognition performance and is beneficial for green radios.
Mingqian Liu, Zhiwen Yan, Junlin Zhang
VTC Spring2