EDBT 2026 Demo / reviewers in the wild / expert
Rundi Wu
dblp:241/5506
· DBLP profile ↗
18ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0003-0133-8196ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 10 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Pixel Cube: Diffusion-based Portrait Video Relighting Through Realistic Lighting ReproductionabstractWe present a diffusion-based method for relighting dynamic portrait videos with photorealism and temporal consistency. Our method is fueled by a hybrid training dataset that consists of real-captured and rendered dynamic portrait videos with diverse subject appearances, facial motions, head poses, and known lighting conditions. Specifically, we construct an LED-based lighting system for realistic lighting emulation and high-speed video relighting data acquisition. By leveraging the image priors embedded in pre-trained video diffusion models, and using per-frame high dynamic range (HDR) environment map as lighting control, we train a high-performance generative model for realistic and identity-preserving dynamic portrait video relighting. In addition to the environment map control, our model uses a synthesized background image to enable control on the camera's exposure level and color tone. Our model can produce temporally consistent relit portrait video that looks realistic and harmonious under a provided new environment and faithfully preserve the subject's expression and fine facial features, including skin tone, wrinkles, and facial hair. Our model generalizes well to unseen data, in terms of the subject appearance, motion, and lighting condition. We perform extensive experiments on relighting in-the-wild videos with various environment maps and demonstrate practical applications on portrait photography. Results show that our method achieves state-of-the-art performance in photorealism, lighting harmony, and temporal consistency. Our project page: https://yufanzhang82.github.io/PixelCube/. Yufan Zhang 0001, Yu Ji 0001, Ayo Ajiboye, Rundi Wu, Yu Guo 0007, Changxi Zheng, Jinwei Ye |
ACM Trans. Graph. | 4 |
| 2025 | SimVS: Simulating World Inconsistencies for Robust View SynthesisabstractNovel-view synthesis techniques achieve impressive results for static scenes but struggle when faced with the inconsistencies inherent to casual capture settings: varying illumination, scene motion, and other unintended effects that are difficult to model explicitly. We present an approach for leveraging generative video models to simulate the inconsistencies in the world that can occur during capture. We use this process, along with existing multi-view datasets, to create synthetic data for training a multi-view harmonization network that is able to reconcile inconsistent observations into a consistent 3D scene. We demonstrate that our world-simulation strategy significantly outperforms traditional augmentation methods in handling real-world scene variations, thereby enabling highly accurate static 3D reconstructions in the presence of a variety of challenging inconsistencies. Project page: https://alextrevithick.com/simvs Alex Trevithick, Roni Paiss, Philipp Henzler, Dor Verbin, Rundi Wu, Hadi Alzayer, Ruiqi Gao, Ben Poole, Jonathan T. Barron, Aleksander Holynski, Ravi Ramamoorthi, Pratul P. Srinivasan |
CVPR | 5 |
| 2025 | CAT4D: Create Anything in 4D with Multi-View Video Diffusion ModelsabstractWe present CAT4D, a method for creating 4D (dynamic 3D) scenes from monocular video. CAT4D leverages a multi-view video diffusion model trained on a diverse combination of datasets to enable novel view synthesis at any specified camera poses and timestamps. Combined with a novel sampling approach, this model can transform a single monocular video into a multi-view video, enabling robust 4D reconstruction via optimization of a deformable 3D Gaussian representation. We demonstrate competitive performance on novel view synthesis and dynamic scene reconstruction benchmarks, and highlight the creative capabilities for 4D scene generation from real or generated videos. See our project page for results and interactive demos: cat-4d.github.io. Rundi Wu, Ruiqi Gao, Ben Poole, Alex Trevithick, Changxi Zheng, Jonathan T. Barron, Aleksander Holynski |
CVPR | 1 |
| 2025 | VLMaterial: Procedural Material Generation with Large Vision-Language ModelsabstractProcedural materials, represented as functional node graphs, are ubiquitous in computer graphics for photorealistic material appearance design. They allow users to perform intuitive and precise editing to achieve desired visual appearances. However, creating a procedural material given an input image requires professional knowledge and significant effort. In this work, we leverage the ability to convert procedural materials into standard Python programs and fine-tune a large pre-trained vision-language model (VLM) to generate such programs from input images. To enable effective fine-tuning, we also contribute an open-source procedural material dataset and propose to perform program-level augmentation by prompting another pre-trained large language model (LLM). Through extensive evaluation, we show that our method outperforms previous methods on both synthetic and real-world examples. Beichen Li 0005, Rundi Wu, Armando Solar-Lezama, Changxi Zheng, Liang Shi 0003, Bernd Bickel, Wojciech Matusik |
ICLR | 2 |
| 2025 | A 1.92mW 1.6-to-3.2GHz Third-Harmonic Blocker-Tolerant Receiver Front-End Using Dual Current-Reuse N-Path LNA and Band-Switchable BalunabstractThis paper presents an ultra-low power third-harmonic blocker-tolerant receiver (RX) using dual current-reuse N-path low noise amplifier (LNA) and band-switchable balun. The proposed RX front-end uses I&Q third-harmonic mixing based on non-overlapping 12-phase local oscillator (LO), which reduces the operating frequency and tunable range of RF source by 6 times to save the power of LO generator. A dual current-reuse N-path LNA is proposed to reduce the LNA power consumption by half without sacrificing out-of-band suppression. Moreover, the trade-off between bandwidth matching and high-Q selectivity is broken under extremely low power consumption by combining band-switchable balun and N-path LNA. The RX is designed in a 40 nm CMOS process with a core area of 0.69 mm2. The simulated results indicate that the RX achieves wideband tunable high-Q selectivity from 1.6 to 3.2 GHz with the double sideband (DSB) NF from 3.8 to 4.4 dB and +12.2 dBm out-of-band (OB) IIP3 at 100 MHz offset. Over the frequency range of interest, total RX power consumption is 1.52 to 1.92 mW, with the LO driver requiring only 0.19 mW/GHz. Rundi Wu, Yetong Wang, Zongle Ma, Keping Wang |
ISCAS | 1 |
| 2024 | ReconFusion: 3D Reconstruction with Diffusion Priorsabstract3D reconstruction methods such as Neural Radiance Fields (NeRFs) excel at rendering photorealistic novel views of complex scenes. However, recovering a high-quality NeRF typically requires tens to hundreds of input images, resulting in a time-consuming capture process. We present ReconFusion to reconstruct real-world scenes using only a few photos. Our approach leverages a diffusion prior for novel view synthesis, trained on synthetic and multiview datasets, which regularizes a NeRF-based 3D reconstruction pipeline at novel camera poses beyond those captured by the set of input images. Our method synthesizes realistic geometry and texture in underconstrained regions while preserving the appearance of observed regions. We perform an extensive evaluation across various real-world datasets, including forward-facing and 360-degree scenes, demonstrating significant performance improvements over previous few-view NeRF reconstruction approaches. Please see our project page at reconfusion.github. io. Rundi Wu, Ben Mildenhall, Philipp Henzler, Keunhong Park, Ruiqi Gao, Daniel Watson, Pratul P. Srinivasan, Dor Verbin, Jonathan T. Barron, Ben Poole, Aleksander Holynski |
CVPR | 1 |
| 2024 | Generative Camera Dolly: Extreme Monocular Dynamic Novel View Synthesis
Basile Van Hoorick, Rundi Wu, Ege Ozguroglu, Kyle Sargent, Ruoshi Liu, Pavel Tokmakov, Achal Dave, Changxi Zheng, Carl Vondrick |
ECCV (24) | 2 |
| 2024 | PhysDreamer: Physics-Based Interaction with 3D Objects via Video Generation
Hong-Xing Yu, Rundi Wu, Brandon Yushan Feng, Changxi Zheng, Noah Snavely, Jiajun Wu 0001, William T. Freeman |
ECCV (2) | 3 |
| 2024 | Sin3DM: Learning a Diffusion Model from a Single 3D Textured ShapeabstractSynthesizing novel 3D models that resemble the input example as long been pursued by graphics artists and machine learning researchers. In this paper, we present Sin3DM, a diffusion model that learns the internal patch distribution from a single 3D textured shape
and generates high-quality variations with fine geometry and texture details. Training a diffusion model directly in 3D would induce large memory and computational cost. Therefore, we first compress the input into a lower-dimensional latent space and then train a diffusion model on it. Specifically, we encode the input 3D textured shape into triplane feature maps that represent the signed distance and texture fields of the input. The denoising network of our diffusion model has a limited receptive field to avoid overfitting, and uses triplane-aware 2D convolution blocks to improve the result quality. Aside from randomly generating new samples, our model also facilitates applications such as retargeting, outpainting and local editing. Through extensive qualitative and quantitative evaluation, we show that our method outperforms prior methods in generation quality of 3D shapes. Rundi Wu, Ruoshi Liu, Carl Vondrick, Changxi Zheng |
ICLR | 1 |
| 2023 | Zero-1-to-3: Zero-shot One Image to 3D ObjectabstractWe introduce Zero-1-to-3, a framework for changing the camera viewpoint of an object given just a single RGB image. To perform novel view synthesis in this under-constrained setting, we capitalize on the geometric priors that large-scale diffusion models learn about natural images. Our conditional diffusion model uses a synthetic dataset to learn controls of the relative camera viewpoint, which allow new images to be generated of the same object under a specified camera transformation. Even though it is trained on a synthetic dataset, our model retains a strong zero-shot generalization ability to out-of-distribution datasets as well as in-the-wild images, including impressionist paintings. Our viewpoint-conditioned diffusion approach can further be used for the task of 3D reconstruction from a single image. Qualitative and quantitative experiments show that our method significantly outperforms state-of-the-art single-view 3D reconstruction and novel view synthesis models by leveraging Internet-scale pre-training. Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tokmakov, Sergey Zakharov, Carl Vondrick |
ICCV | 2 |
| 2023 | Implicit Neural Spatial Representations for Time-dependent PDEsabstractImplicit Neural Spatial Representation (INSR) has emerged as an effective representation of spatially-dependent vector fields. This work explores solving time-dependent PDEs with INSR. Classical PDE solvers introduce both temporal and spatial discretizations. Common spatial discretizations include meshes and meshless point clouds, where each degree-of-freedom corresponds to a location in space. While these explicit spatial correspondences are intuitive to model and understand, these representations are not necessarily optimal for accuracy, memory usage, or adaptivity. Keeping the classical temporal discretization unchanged (e.g., explicit/implicit Euler), we explore INSR as an alternative spatial discretization, where spatial information is implicitly stored in the neural network weights. The network weights then evolve over time via time integration. Our approach does not require any training data generated by existing solvers because our approach is the solver itself. We validate our approach on various PDEs with examples involving large elastic deformations, turbulent fluids, and multi-scale phenomena. While slower to compute than traditional representations, our approach exhibits higher accuracy and lower memory consumption. Whereas classical solvers can dynamically adapt their spatial representation only by resorting to complex remeshing algorithms, our INSR approach is intrinsically adaptive. By tapping into the rich literature of classic time integrators, e.g., operator-splitting schemes, our method enables challenging simulations in contact mechanics and turbulent flows where previous neural-physics approaches struggle. Videos and codes are available on the project page: http://www.cs.columbia.edu/cg/INSR-PDE/ Rundi Wu, Eitan Grinspun, Changxi Zheng, Peter Yichen Chen |
ICML | 2 |
| 2022 | Dynamic Sliding Window for Realtime Denoising NetworksabstractRealtime speech denoising has been long studied. Almost all existing methods process the incoming data stream using a sliding window of fixed-size. Yet, we show that the use of fixed-size sliding window may lead to an accumulating lag, especially in presence of other background computing processes that may occupy CPU resources. In response, we propose a new sliding window strategy and a lightweight neural network to leverage it. Our experiments show that the proposed approach achieves denoising quality on a par with the stateof-the-art realtime denoising models. More importantly, our approach is faster, maintaining a stable realtime performance even when the available computing power fluctuates. Jinxu Xiang, Yuyang Zhu, Rundi Wu, Ruilin Xu 0001, Yuko Ishiwaka, Changxi Zheng |
ICASSP | 3 |
| 2022 | Learning to Generate 3D Shapes from a Single ExampleabstractExisting generative models for 3D shapes are typically trained on a large 3D dataset, often of a specific object category. In this paper, we investigate the deep generative model that learns from only a single reference 3D shape. Specifically, we present a multi-scale GAN-based model designed to capture the input shape's geometric features across a range of spatial scales. To avoid large memory and computational cost induced by operating on the 3D volume, we build our generator atop the tri-plane hybrid representation, which requires only 2D convolutions. We train our generative model on a voxel pyramid of the reference shape, without the need of any external supervision or manual annotation. Once trained, our model can generate diverse and high-quality 3D shapes possibly of different sizes and aspect ratios. The resulting shapes present variations across different scales, and at the same time retain the global structure of the reference shape. Through extensive evaluation, both qualitative and quantitative, we demonstrate that our model can generate 3D shapes of various types. 1 Rundi Wu, Changxi Zheng |
ACM Trans. Graph. | 1 |
| 2021 | DeepCAD: A Deep Generative Network for Computer-Aided Design ModelsabstractDeep generative models of 3D shapes have received a great deal of research interest. Yet, almost all of them generate discrete shape representations, such as voxels, point clouds, and polygon meshes. We present the first 3D generative model for a drastically different shape representation— describing a shape as a sequence of computer-aided design (CAD) operations. Unlike meshes and point clouds, CAD models encode the user creation process of 3D shapes, widely used in numerous industrial and engineering design tasks. However, the sequential and irregular structure of CAD operations poses significant challenges for existing 3D generative models. Drawing an analogy between CAD operations and natural language, we propose a CAD generative network based on the Transformer. We demonstrate the performance of our model for both shape autoencoding and random shape generation. To train our network, we create a new CAD dataset consisting of 178,238 models and their CAD construction sequences. We have made this dataset publicly available to promote future research on this topic. Rundi Wu, Chang Xiao 0003, Changxi Zheng |
ICCV | 1 |
| 2020 | PQ-NET: A Generative Part Seq2Seq Network for 3D ShapesabstractWe introduce PQ-NET, a deep neural network which represents and generates 3D shapes via sequential part assembly. The input to our network is a 3D shape segmented into parts, where each part is first encoded into a feature representation using a part autoencoder. The core component of PQ-NET is a sequence-to-sequence or Seq2Seq autoencoder which encodes a sequence of part features into a latent vector of fixed size, and the decoder reconstructs the 3D shape, one part at a time, resulting in a sequential assembly. The latent space formed by the Seq2Seq encoder encodes both part structure and fine part geometry. The decoder can be adapted to perform several generative tasks including shape autoencoding, interpolation, novel shape generation, and single-view 3D reconstruction, where the generated shapes are all composed of meaningful parts. Rundi Wu, Yixin Zhuang, Kai Xu 0004, Hao (Richard) Zhang, Baoquan Chen |
CVPR | 1 |
| 2020 | Multimodal Shape Completion via Conditional Generative Adversarial Networks
Rundi Wu, Xuelin Chen, Yixin Zhuang, Baoquan Chen |
ECCV (4) | 1 |
| 2020 | Listening to Sounds of Silence for Speech DenoisingabstractWe introduce a deep learning model for speech denoising, a long-standing challenge in audio analysis arising in numerous applications. Our approach is based on a key observation about human speech: there is often a short pause between each sentence or word. In a recorded speech signal, those pauses introduce a series of time periods during which only noise is present. We leverage these incidental silent intervals to learn a model for automatic speech denoising given only mono-channel audio. Detected silent intervals over time expose not just pure noise but its time-varying features, allowing the model to learn noise dynamics and suppress it from the speech signal. Experiments on multiple datasets confirm the pivotal role of silent interval detection for speech denoising, and our method outperforms several state-of-the-art denoising methods, including those that accept only audio input (like ours) and those that denoise based on audiovisual input (and hence require more information). We also show that our method enjoys excellent generalization properties, such as denoising spoken languages not seen during training. Ruilin Xu 0001, Rundi Wu, Yuko Ishiwaka, Carl Vondrick, Changxi Zheng |
NeurIPS | 2 |
| 2019 | Learning character-agnostic motion for motion retargeting in 2DabstractAnalyzing human motion is a challenging task with a wide variety of applications in computer vision and in graphics. One such application, of particular importance in computer animation, is the retargeting of motion from one performer to another. While humans move in three dimensions, the vast majority of human motions are captured using video, requiring 2D-to-3D pose and camera recovery, before existing retargeting approaches may be applied. In this paper, we present a new method for retargeting video-captured motion between different human performers, without the need to explicitly reconstruct 3D poses and/or camera parameters. In order to achieve our goal, we learn to extract, directly from a video, a high-level latent motion representation, which is invariant to the skeleton geometry and the camera view. Our key idea is to train a deep neural network to decompose temporal sequences of 2D poses into three components: motion, skeleton, and camera view-angle. Having extracted such a representation, we are able to re-combine motion with novel skeletons and camera views, and decode a retargeted temporal sequence, which we compare to a ground truth from a synthetic dataset. We demonstrate that our framework can be used to robustly extract human motion from videos, bypassing 3D reconstruction, and outperforming existing retargeting methods, when applied to videos in-the-wild. It also enables additional applications, such as performance cloning, video-driven cartoons, and motion retrieval. Kfir Aberman, Rundi Wu, Dani Lischinski, Baoquan Chen, Daniel Cohen-Or |
ACM Trans. Graph. | 2 |