EDBT 2026 Demo / reviewers in the wild / expert
Kai Zhang 0045
dblp:55/957-45
· DBLP profile ↗
28ranked-venue papers
5as first author
27since 2021 · last 2025
0000-0002-1727-1689ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 5 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 17 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LazyDiT: Lazy Learning for the Acceleration of Diffusion TransformersabstractDiffusion Transformers have emerged as the preeminent models for a wide array of generative tasks, demonstrating superior performance and efficacy across various applications. The promising results come at the cost of slow inference, as each denoising step requires running the whole transformer model with a large amount of parameters. In this paper, we show that performing the full computation of the model at each diffusion step is unnecessary, as some computations can be skipped by lazily reusing the results of previous steps. Furthermore, we show that the lower bound of similarity between outputs at consecutive steps is notably high, and this similarity can be linearly approximated using the inputs. To verify our demonstrations, we propose the **LazyDiT**, a lazy learning framework that efficiently leverages cached results from earlier steps to skip redundant computations. Specifically, we incorporate lazy learning layers into the model, effectively trained to maximize laziness, enabling dynamic skipping of redundant computations. Experimental results show that LazyDiT outperforms the DDIM sampler across multiple diffusion transformer models at various resolutions. Furthermore, we implement our method on mobile devices, achieving better performance than DDIM with similar latency. Xuan Shen, Zhao Song 0002, Yufa Zhou 0001, Bo Chen 0029, Yanyu Li, Yifan Gong 0004, Kai Zhang 0045, Hao Tan 0002, Jason Kuen, Henghui Ding, Zhihao Shu, Wei Niu 0002, Pu Zhao 0001, Yanzhi Wang 0001, Jiuxiang Gu |
AAAI | 7 |
| 2025 | Generating 3D-Consistent Videos from Unposed Internet PhotosabstractWe address the problem of generating videos from unposed internet photos. A handful of input images serve as keyframes, and our model interpolates between them to simulate a path moving between the cameras. Given random images, a model’s ability to capture underlying geometry, recognize scene identity, and relate frames in terms of camera position and orientation reflects a fundamental understanding of 3D structure and scene layout. However, existing video models such as Luma Dream Machine fail at this task. We design a self-supervised method that takes advantage of the consistency of videos and variability of multiview internet photos to train a scalable, 3D-aware video model without any 3D annotations such as camera parameters. We validate that our method outperforms all baselines in terms of geometric and appearance consistency. We also show our model benefits applications that enable camera control, such as 3D Gaussian Splatting. Our results suggest that we can scale up scene-level 3D learning using only 2D data such as videos and multiview internet photos. Gene Chou, Kai Zhang 0045, Sai Bi, Hao Tan 0002, Zexiang Xu, Fujun Luan, Bharath Hariharan, Noah Snavely |
CVPR | 2 |
| 2025 | Turbo3D: Ultra-fast Text-to-3D GenerationabstractWe present Turbo3D, an ultra-fast text-to-3D system capable of generating high-quality Gaussian splatting assets in under one second. Turbo3D employs a rapid 4-step, 4-view diffusion generator and an efficient feed-forward Gaussian reconstructor, both operating in latent space, as shown in Fig. 1. The 4-step, 4-view generator is a student model distilled through a novel Dual-Teacher approach, which encourages the student to learn view consistency from a multiview teacher and photo-realism from a single-view teacher. By shifting the Gaussian reconstructor’s inputs from pixel space to latent space, we eliminate the extra image decoding time and halve the transformer sequence length for maximum efficiency. Our method demonstrates superior 3D generation results compared to previous baselines, while operating in a fraction of their runtime. Hanzhe Hu, Tianwei Yin, Fujun Luan, Hao Tan 0002, Zexiang Xu, Sai Bi, Shubham Tulsiani, Kai Zhang 0045 |
CVPR | 9 |
| 2025 | MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized DataabstractWe propose scaling up 3D scene reconstruction by training with synthesized data. At the core of our work is MegaSynth, a procedurally generated 3D dataset comprising 700K scenes—over 50 times larger than the prior real dataset DL3DV—dramatically scaling the training data. To enable scalable data generation, our key idea is eliminating semantic information, removing the need to model complex semantic priors such as object affordances and scene composition. Instead, we model scenes with basic spatial structures and geometry primitives, offering scalability. Besides, we control data complexity to facilitate training while loosely aligning it with real-world data distribution to benefit real-world generalization. We explore training LRMs with both MegaSynth and available real data. Experiment results show that joint training or pre-training with MegaSynth improves reconstruction quality by 1.2 to 1.8 dB PSNR across diverse image domains. Moreover, models trained solely on MegaSynth perform comparably to those trained on real data, underscoring the low-level nature of 3D reconstruction. Additionally, we provide an in-depth analysis of MegaSynth’s properties for enhancing model capability, training stability, and generalization, as well as application to other tasks. Hanwen Jiang, Zexiang Xu, Desai Xie, Haian Jin, Fujun Luan, Zhixin Shu, Kai Zhang 0045, Sai Bi, Xin Sun 0014, Jiuxiang Gu, Qixing Huang, Georgios Pavlakos, Hao Tan 0002 |
CVPR | 8 |
| 2025 | Buffer Anytime: Zero-Shot Video Depth and Normal from Image PriorsabstractWe present Buffer Anytime, a framework for estimation of depth and normal maps (which we call geometric buffers) from video that eliminates the need for paired video–depth and video–normal training data. Instead of relying on large-scale annotated video datasets, we demonstrate high-quality video buffer estimation by leveraging single-image priors with temporal consistency constraints. Our zero-shot training strategy combines state-of-the-art image estimation models based on optical flow smoothness through a hybrid loss function, implemented via a lightweight temporal attention architecture. Applied to leading image models like Depth Anything V2 and Marigold-E2E-FT, our approach significantly improves temporal consistency while maintaining accuracy. Experiments show that our method not only outperforms image-based approaches but also achieves results comparable to state-of-the-art video models trained on large-scale paired video datasets, despite using no such paired video data. Zhengfei Kuang, Kai Zhang 0045, Hao Tan 0002, Sai Bi, Zexiang Xu, Milos Hasan, Gordon Wetzstein, Fujun Luan |
CVPR | 3 |
| 2025 | RandAR: Decoder-only Autoregressive Visual Generation in Random OrdersabstractWe introduce RandAR, a decoder-only visual autoregressive (AR) model capable of generating images in arbitrary token orders. Unlike previous decoder-only AR models that rely on a predefined generation order, RandAR removes this inductive bias, unlocking new capabilities in decoder-only generation. Our essential design enables random order by inserting a "position instruction token" before each image token to be predicted, representing the spatial location of the next image token. Trained on randomly permuted token sequences - a more challenging task than fixed-order generation, RandAR achieves comparable performance to its conventional raster-order counterpart. More importantly, decoder-only transformers trained from random orders acquire new capabilities. For the efficiency bottleneck of AR models, RandAR adopts parallel decoding with KV-Cache at inference time, enjoying 2.5 × acceleration without sacrificing generation quality. Additionally, RandAR supports inpainting, outpainting and resolution extrapolation in a zero-shot manner. We hope RandAR inspires new directions for decoder-only visual generation models and broadens their applications across diverse scenarios. Our project page is at https://rand-ar.github.io/. Ziqi Pang, Fujun Luan, Yunze Man, Hao Tan 0002, Kai Zhang 0045, William T. Freeman, Yu-Xiong Wang |
CVPR | 6 |
| 2025 | Baking Gaussian Splatting Into Diffusion Denoiser for Fast and Scalable Single-Stage Image-to-3D Generation and Reconstruction
Yuanhao Cai, He Zhang 0004, Kai Zhang 0045, Yixun Liang, Mengwei Ren, Fujun Luan, Qing Liu 0017, Soo Ye Kim, Jianming Zhang 0001, Yuqian Zhou, Yulun Zhang 0001, Xiaokang Yang 0001, Zhe Lin 0001, Alan L. Yuille |
ICCV | 3 |
| 2025 | Long-LRM: Long-Sequence Large Reconstruction Model for Wide-Coverage Gaussian SplatsabstractWe propose Long-LRM, a feed-forward 3D Gaussian reconstruction model for instant, high-resolution, 360° wide-coverage, scene-level reconstruction. Specifically, it takes in 32 input images at a resolution of 960x540 and produces the Gaussian reconstruction in just 1 second on a single A100 GPU. To handle the long sequence of 250K tokens brought by the large input size, Long-LRM features a mixture of the recent Mamba2 blocks and the classical transformer blocks, enhanced by a light-weight token merging module and Gaussian pruning steps that balance between quality and efficiency. We evaluate Long-LRM on the large-scale DL3DV benchmark and Tanks&Temples, demonstrating reconstruction quality comparable to the optimization-based methods while achieving an 800x speedup w.r.t. the optimization-based approaches and an input size at least 60x larger than the previous feed-forward approaches. We conduct extensive ablation studies on our model design choices for both rendering quality and computation efficiency. We also explore Long-LRM's compatibility with other Gaussian variants such as 2D GS, which enhances Long-LRM's ability in geometry reconstruction. Project page: https://arthurhero.github.io/projects/llrm Hao Tan 0002, Kai Zhang 0045, Sai Bi, Fujun Luan, Yicong Hong, Fuxin Li, Zexiang Xu |
ICCV | 3 |
| 2025 | Rayzer: a Self-Supervised Large View Synthesis Model
Hanwen Jiang, Hao Tan 0002, Peng Wang 0099, Hai Jin 0001, Sai Bi, Kai Zhang 0045, Fujun Luan, Kalyan Sunkavalli, Qixing Huang, Georgios Pavlakos |
ICCV | 7 |
| 2025 | LVSM: A Large View Synthesis Model with Minimal 3D Inductive BiasabstractWe propose the Large View Synthesis Model (LVSM), a novel transformer-based approach for scalable and generalizable novel view synthesis from sparse-view inputs. We introduce two architectures: (1) an encoder-decoder LVSM, which encodes input image tokens into a fixed number of 1D latent tokens, functioning as a fully learned scene representation, and decodes novel-view images from them; and (2) a decoder-only LVSM, which directly maps input images to novel-view outputs, completely eliminating intermediate scene representations. Both models bypass the 3D inductive biases used in previous methods---from 3D representations (e.g., NeRF, 3DGS) to network designs (e.g., epipolar projections, plane sweeps)---addressing novel view synthesis with a fully data-driven approach. While the encoder-decoder model offers faster inference due to its independent latent representation, the decoder-only LVSM achieves superior quality, scalability, and zero-shot generalization, outperforming previous state-of-the-art methods by 1.5 to 3.5 dB PSNR. Comprehensive evaluations across multiple datasets demonstrate that both LVSM variants achieve state-of-the-art novel view synthesis quality, delivering superior performance even with reduced computational resources (1-2 GPUs). Please see our anonymous website for more details: https://haian-jin.github.io/projects/LVSM/ Haian Jin, Hanwen Jiang, Hao Tan 0002, Kai Zhang 0045, Sai Bi, Fujun Luan, Noah Snavely, Zexiang Xu |
ICLR | 4 |
| 2025 | RelitLRM: Generative Relightable Radiance for Large Reconstruction ModelsabstractWe propose RelitLRM, a Large Reconstruction Model (LRM) for generating high-quality Gaussian splatting representations of 3D objects under novel illuminations from sparse (4-8) posed images captured under unknown static lighting. Unlike prior inverse rendering methods requiring dense captures and slow optimization, often causing artifacts like incorrect highlights or shadow baking, RelitLRM adopts a feed-forward transformer-based model with a novel combination of a geometry reconstructor and a relightable appearance generator based on diffusion. The model is trained end-to-end on synthetic multi-view renderings of objects under varying known illuminations. This architecture design enables to effectively decompose geometry and appearance, resolve the ambiguity between material and lighting, and capture the multi-modal distribution of shadows and specularity in the relit appearance. We show our sparse-view feed-forward RelitLRM offers competitive relighting results to state-of-the-art dense-view optimization-based baselines while being significantly faster. Our project page is available at: https://relit-lrm.github.io/. Zhengfei Kuang, Haian Jin, Zexiang Xu, Sai Bi, Hao Tan 0002, He Zhang 0004, Milos Hasan, William T. Freeman, Kai Zhang 0045, Fujun Luan |
ICLR | 11 |
| 2025 | Gaussian Mixture Flow Matching ModelsabstractDiffusion models approximate the denoising distribution as a Gaussian and predict its mean, whereas flow matching models reparameterize the Gaussian mean as flow velocity. However, they underperform in few-step sampling due to discretization error and tend to produce over-saturated colors under classifier-free guidance (CFG). To address these limitations, we propose a novel Gaussian mixture flow matching (GMFlow) model: instead of predicting the mean, GMFlow predicts dynamic Gaussian mixture (GM) parameters to capture a multi-modal flow velocity distribution, which can be learned with a KL divergence loss. We demonstrate that GMFlow generalizes previous diffusion and flow matching models where a single Gaussian is learned with an $L_2$ denoising loss. For inference, we derive GM-SDE/ODE solvers that leverage analytic denoising distributions and velocity fields for precise few-step sampling. Furthermore, we introduce a novel probabilistic guidance scheme that mitigates the over-saturation issues of CFG and improves image generation quality. Extensive experiments demonstrate that GMFlow consistently outperforms flow matching baselines in generation quality, achieving a Precision of 0.942 with only 6 sampling steps on ImageNet 256$\times$256. Hansheng Chen 0001, Kai Zhang 0045, Hao Tan 0002, Zexiang Xu, Fujun Luan, Leonidas J. Guibas, Gordon Wetzstein, Sai Bi |
ICML | 2 |
| 2025 | 4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any TimeabstractCan we scale 4D pretraining to learn general space-time representations that reconstruct an object from a few views at some times to any view at any time? We provide an affirmative answer with 4D-LRM, the first large-scale 4D reconstruction model that takes input from unconstrained views and timestamps and renders arbitrary novel view-time combinations. Unlike prior 4D approaches, e.g., optimization-based, geometry-based, or generative, that struggle with efficiency, generalization, or faithfulness, 4D-LRM learns a unified space-time representation and directly predicts per-pixel 4D Gaussian primitives from posed image tokens across time, enabling fast, high-quality rendering at, in principle, infinite frame rate. Our results demonstrate that scaling spatiotemporal pretraining enables accurate and efficient 4D reconstruction. We show that 4D-LRM generalizes to novel objects, interpolates across time, and handles diverse camera setups. It reconstructs 24-frame sequences in one forward pass with less than 1.5 seconds on a single A100 GPU. Ziqiao Ma 0001, Xuweiyi Chen, Shoubin Yu, Sai Bi, Kai Zhang 0045, Sihan Xu, Zexiang Xu, Kalyan Sunkavalli, Mohit Bansal, Joyce Y. Chai, Hao Tan 0002 |
NeurIPS | 5 |
| 2024 | Neural Directional Encoding for Efficient and Accurate View-Dependent Appearance ModelingabstractNovel-view synthesis of specular objects like shiny metals or glossy paints remains a significant challenge. Not only the glossy appearance but also global illumination effects, including reflections of other objects in the environment, are critical components to faithfully reproduce a scene. In this paper, we present Neural Directional Encoding (NDE), a view-dependent appearance encoding of neural radiance fields (NeRF) for rendering specular objects. NDE transfers the concept of feature-grid-based spatial encoding to the angular domain, significantly improving the ability to model high-frequency angular signals. In contrast to previous methods that use encoding functions with only angular input, we additionally cone-trace spatial features to obtain a spatially varying directional encoding, which addresses the challenging interreflection effects Extensive experiments on both synthetic and real datasets show that a NeRF model with NDE (1) outperforms the state of the art on view synthesis of specular objects, and (2) works with small networks to allow fast (real-time) inference. The source code is available at: https://github.com/lwwu2/nde Liwen Wu, Sai Bi, Zexiang Xu, Fujun Luan, Kai Zhang 0045, Iliyan Georgiev, Kalyan Sunkavalli, Ravi Ramamoorthi |
CVPR | 5 |
| 2024 | GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D GenerationabstractDespite recent advances in text-to-3D generative methods, there is a notable absence of reliable evaluation metrics. Existing metrics usually focus on a single criterion each, such as how well the asset aligned with the input text. These metrics lack the flexibility to generalize to different evaluation criteria and might not align well with human preferences. Conducting user preference studies is an alternative that offers both adaptability and human-aligned results. User studies, however, can be very ex-pensive to scale. This paper presents an automatic, ver-satile, and human-aligned evaluation metric for text-to-3D generative models. To this end, we first develop a prompt generator using GPT-4V to generate evaluating prompts, which serve as input to compare text-to-3D models. We further design a method instructing GPT-4V to compare two 3D assets according to user-defined crite-ria. Finally, we use these pairwise comparison results to assign these models Elo ratings. Experimental results suggest our metric strongly aligns with human preference across different evaluation criteria. Our code is available at https://github.com/3DTopia/GPTEval3D. Guandao Yang, Zhibing Li, Kai Zhang 0045, Ziwei Liu 0002, Leonidas J. Guibas, Dahua Lin, Gordon Wetzstein |
CVPR | 4 |
| 2024 | DATENeRF: Depth-Aware Text-Based Editing of NeRFs
Sara Rojas 0001, Julien Philip, Kai Zhang 0045, Sai Bi, Fujun Luan, Bernard Ghanem, Kalyan Sunkavalli |
ECCV (11) | 3 |
| 2024 | MegaScenes: Scene-Level View Synthesis at Scale
Joseph Tung, Gene Chou, Ruojin Cai, Guandao Yang, Kai Zhang 0045, Gordon Wetzstein, Bharath Hariharan, Noah Snavely |
ECCV (29) | 5 |
| 2024 | GS-LRM: Large Reconstruction Model for 3D Gaussian Splatting
Kai Zhang 0045, Sai Bi, Hao Tan 0002, Yuanbo Xiangli, Nanxuan Zhao, Kalyan Sunkavalli, Zexiang Xu |
ECCV (22) | 1 |
| 2024 | PF-LRM: Pose-Free Large Reconstruction Model for Joint Pose and Shape PredictionabstractWe propose a Pose-Free Large Reconstruction Model (PF-LRM) for reconstructing a 3D object from a few unposed images even with little visual overlap, while simultaneously estimating the relative camera poses in ~1.3 seconds on a single A100 GPU. PF-LRM is a highly scalable method utilizing self-attention blocks to exchange information between 3D object tokens and 2D image tokens; we predict a coarse point cloud for each view, and then use a differentiable Perspective-n-Point (PnP) solver to obtain camera poses. When trained on a huge amount of multi-view posed data of ~1M objects, PF-LRM shows strong cross-dataset generalization ability, and outperforms baseline methods by a large margin in terms of pose prediction accuracy and 3D reconstruction quality on various unseen evaluation datasets. We also demonstrate our model's applicability in downstream text/image-to-3D task with fast feed-forward inference. Our project website is at: https://totoro97.github.io/pf-lrm. Peng Wang 0099, Hao Tan 0002, Sai Bi, Yinghao Xu 0001, Fujun Luan, Kalyan Sunkavalli, Wenping Wang 0001, Zexiang Xu, Kai Zhang 0045 |
ICLR | 9 |
| 2024 | LRM: Large Reconstruction Model for Single Image to 3DabstractWe propose the first Large Reconstruction Model (LRM) that predicts the 3D model of an object from a single input image within just 5 seconds. In contrast to many previous methods that are trained on small-scale datasets such as ShapeNet in a category-specific fashion, LRM adopts a highly scalable transformer-based architecture with 500 million learnable parameters to directly predict a neural radiance field (NeRF) from the input image. We train our model in an end-to-end manner on massive multi-view data containing around 1 million objects, including both synthetic renderings from Objaverse and real captures from MVImgNet. This combination of a high-capacity model and large-scale training data empowers our model to be highly generalizable and produce high-quality 3D reconstructions from various testing inputs, including real-world in-the-wild captures and images created by generative models. Video demos and interactable 3D meshes can be found on our LRM project webpage: https://yiconghong.me/LRM. Yicong Hong, Kai Zhang 0045, Jiuxiang Gu, Sai Bi, Yang Zhou 0009, Difan Liu, Feng Liu 0015, Kalyan Sunkavalli, Trung Bui, Hao Tan 0002 |
ICLR | 2 |
| 2024 | Instant3D: Fast Text-to-3D with Sparse-view Generation and Large Reconstruction ModelabstractText-to-3D with diffusion models has achieved remarkable progress in recent years. However, existing methods either rely on score distillation-based optimization which suffer from slow inference, low diversity and Janus problems, or are feed-forward methods that generate low-quality results due to the scarcity of 3D training data. In this paper, we propose Instant3D, a novel method that generates high-quality and diverse 3D assets from text prompts in a feed-forward manner. We adopt a two-stage paradigm, which first generates a sparse set of four structured and consistent views from text in one shot with a fine-tuned 2D text-to-image diffusion model, and then directly regresses the NeRF from the generated images with a novel transformer-based sparse-view reconstructor. Through extensive experiments, we demonstrate that our method can generate diverse 3D assets of high visual quality within 20 seconds, which is two orders of magnitude faster than previous optimization-based methods that can take 1 to 10 hours. Our project webpage is: https://jiahao.ai/instant3d/. Hao Tan 0002, Kai Zhang 0045, Zexiang Xu, Fujun Luan, Yinghao Xu 0001, Yicong Hong, Kalyan Sunkavalli, Gregory Shakhnarovich, Sai Bi |
ICLR | 3 |
| 2024 | DMV3D: Denoising Multi-view Diffusion Using 3D Large Reconstruction ModelabstractWe propose DMV3D, a novel 3D generation approach that uses a transformer-based 3D large reconstruction model to denoise multi-view diffusion. Our reconstruction model incorporates a triplane NeRF representation and, functioning as a denoiser, can denoise noisy multi-view images via 3D NeRF reconstruction and rendering, achieving single-stage 3D generation in the 2D diffusion denoising process. We train DMV3D on large-scale multi-view image datasets of extremely diverse objects using only image reconstruction losses, without accessing 3D assets. We demonstrate state-of-the-art results for the single-image reconstruction problem where probabilistic modeling of unseen object parts is required for generating diverse reconstructions with sharp textures. We also show high-quality text-to-3D generation results outperforming previous 3D diffusion models. Our project website is at: https://dmv3d.github.io/. Yinghao Xu 0001, Hao Tan 0002, Fujun Luan, Sai Bi, Peng Wang 0099, Zifan Shi, Kalyan Sunkavalli, Gordon Wetzstein, Zexiang Xu, Kai Zhang 0045 |
ICLR | 11 |
| 2024 | Neural Gaffer: Relighting Any Object via DiffusionabstractSingle-image relighting is a challenging task that involves reasoning about the complex interplay between geometry, materials, and lighting. Many prior methods either support only specific categories of images, such as portraits, or require special capture conditions, like using a flashlight. Alternatively, some methods explicitly decompose a scene into intrinsic components, such as normals and BRDFs, which can be inaccurate or under-expressive. In this work, we propose a novel end-to-end 2D relighting diffusion model, called Neural Gaffer, that takes a single image of any object and can synthesize an accurate, high-quality relit image under any novel environmental lighting condition, simply by conditioning an image generator on a target environment map, without an explicit scene decomposition. Our method builds on a pre-trained diffusion model, and fine-tunes it on a synthetic relighting dataset, revealing and harnessing the inherent understanding of lighting present in the diffusion model. We evaluate our model on both synthetic and in-the-wild Internet imagery and demonstrate its advantages in terms of generalization and accuracy. Moreover, by combining with other generative methods, our model enables many downstream 2D tasks, such as text-based relighting and object insertion. Our model can also operate as a strong relighting prior for 3D tasks, such as relighting a radiance field. Haian Jin, Fujun Luan, Yuanbo Xiangli, Sai Bi, Kai Zhang 0045, Zexiang Xu, Jin Sun 0009, Noah Snavely |
NeurIPS | 6 |
| 2024 | LRM-Zero: Training Large Reconstruction Models with Synthesized DataabstractWe present LRM-Zero, a Large Reconstruction Model (LRM) trained entirely on synthesized 3D data, achieving high-quality sparse-view 3D reconstruction. The core of LRM-Zero is our procedural 3D dataset, Zeroverse, which is automatically synthesized from simple primitive shapes with random texturing and augmentations (e.g., height fields, boolean differences, and wireframes). Unlike previous 3D datasets (e.g., Objaverse) which are often captured or crafted by humans to approximate real 3D data, Zeroverse completely ignores realistic global semantics but is rich in complex geometric and texture details that are locally similar to or even more intricate than real objects. We demonstrate that our LRM-Zero, trained with our fully synthesized Zeroverse, can achieve high visual quality in the reconstruction of real-world objects, competitive with models trained on Objaverse. We also analyze several critical design choices of Zeroverse that contribute to LRM-Zero's capability and training stability. Our work demonstrates that 3D reconstruction, one of the core tasks in 3D vision, can potentially be addressed without the semantics of real-world objects. The Zeroverse's procedural synthesis code and interactive visualization are available at: https://desaixie.github.io/lrm-zero/. Desai Xie, Sai Bi, Zhixin Shu, Kai Zhang 0045, Zexiang Xu, Yi Zhou 0023, Sören Pirk, Arie E. Kaufman, Xin Sun 0014, Hao Tan 0002 |
NeurIPS | 4 |
| 2022 | IRON: Inverse Rendering by Optimizing Neural SDFs and Materials from Photometric ImagesabstractWe propose a neural inverse rendering pipeline called IRON that operates on photometric images and outputs high-quality 3D content in the format of triangle meshes and material textures readily deployable in existing graphics pipelines. Our method adopts neural representations for geometry as signed distance fields (SDFs) and materials during optimization to enjoy their flexibility and compactness, and features a hybrid optimization scheme for neural SDFs: first, optimize using a volumetric radiance field approach to recover correct topology, then optimize further using edgeaware physics-based surface rendering for geometry refinement and disentanglement of materials and lighting. In the second stage, we also draw inspiration from mesh-based differentiable rendering, and design a novel edge sampling algorithm for neural SDFs to further improve performance. We show that our IRON achieves significantly better inverse rendering quality compared to prior works. Kai Zhang 0045, Fujun Luan, Zhengqi Li, Noah Snavely |
CVPR | 1 |
| 2022 | ARF: Artistic Radiance Fields
Kai Zhang 0045, Nicholas I. Kolkin, Sai Bi, Fujun Luan, Zexiang Xu, Eli Shechtman, Noah Snavely |
ECCV (31) | 1 |
| 2021 | PhySG: Inverse Rendering With Spherical Gaussians for Physics-Based Material Editing and RelightingabstractWe present PhySG, an end-to-end inverse rendering pipeline that includes a fully differentiable renderer, and can reconstruct geometry, materials, and illumination from scratch from a set of images. Our framework represents specular BRDFs and environmental illumination using mixtures of spherical Gaussians, and represents geometry as a signed distance function parameterized as a Multi-Layer Perceptron. The use of spherical Gaussians allows us to efficiently solve for approximate light transport, and our method works on scenes with challenging non-Lambertian reflectance captured under natural, static illumination. We demonstrate, with both synthetic and real data, that our reconstructions not only enable rendering of novel viewpoints, but also physics-based appearance editing of materials and illumination. Kai Zhang 0045, Fujun Luan, Qianqian Wang 0002, Kavita Bala, Noah Snavely |
CVPR | 1 |
| 2020 | Depth Sensing Beyond LiDAR RangeabstractDepth sensing is a critical component of autonomous driving technologies, but today's LiDAR- or stereo camera- based solutions have limited range. We seek to increase the maximum range of self-driving vehicles' depth perception modules for the sake of better safety. To that end, we propose a novel three-camera system that utilizes small field of view cameras. Our system, along with our novel algorithm for computing metric depth, does not require full pre-calibration and can output dense depth maps with practically acceptable accuracy for scenes and objects at long distances not well covered by most commercial LiDARs. Kai Zhang 0045, Jiaxin Xie, Noah Snavely, Qifeng Chen 0001 |
CVPR | 1 |