VLDB 2026 Research / reviewers in the wild / expert
Yiyi Liao
dblp:139/0761
· DBLP profile ↗
56ranked-venue papers
6as first author
45since 2021 · last 2026
0000-0001-6662-3022ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 5 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 3 first-author · 31 since 2021Systems, architecture and hardware · 8 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FreeFix: Boosting 3D Gaussian Splatting via Fine-Tuning-Free Diffusion ModelsabstractNeural Radiance Fields and 3D Gaussian Splatting have advanced novel view synthesis, yet still rely on dense inputs and often degrade at extrapolated views. Recent approaches leverage generative models, such as diffusion models, to provide additional supervision, but face a tradeoff between generalization and fidelity: fine-tuning diffusion models for artifact removal improves fidelity but risks overfitting, while fine-tuning-free methods preserve generalization but often yield lower fidelity. We introduce FreeFix, a fine-tuning-free approach that pushes the boundary of this trade-off by enhancing extrapolated rendering with pretrained image diffusion models. We present an interleaved 2D-3D refinement strategy, showing that image diffusion models can be leveraged for consistent refinement without relying on costly video diffusion models. Furthermore, we take a closer look at the guidance signal for 2D refinement and propose a per-pixel confidence mask to identify uncertain regions for targeted improvement. Experiments across multiple datasets show that FreeFix improves multiframe consistency and achieves performance comparable to or surpassing fine-tuning-based methods, while retaining strong generalization ability. Our project page is at https://xdimlab.github.io/freefix. Zisen Shao, Sheng Miao, Dongfeng Bai, Yiyi Liao |
3DV | 7 |
| 2026 | MPEG Explorations Toward 3D Gaussian Splat Coding and Standardizationabstract3D Gaussian splats (3DGS) have rapidly gained traction as a 3D scene representation technique that enables efficient real-time rendering and highfidelity novel view synthesis. This paper reports on ongoing MPEG GSC efforts conducted jointly by the MPEG Video Coding group (WG 4) and the Coding of 3D Graphics and Haptics group (WG 7) to define a practical and interoperable compression framework for 3DGS. MPEG GSC is planning GSC standardization with short-term and long-term timelines to address market requirements. The short-term objective is to standardize coding tools that build on proven MPEG ecosystems while introducing only the minimal set of extensions, syntax, and processing required for INRIA-3DGS format (referred to in MPEG as I-3DGS). In the long-term, MPEG is also investigating broader alternatives for 3DGS representation and compression, including approaches that integrate training during compression, collectively referred to as Alternative-3DGS (A-3DGS). More specifically, this paper focuses on the I-3DGS and explores both geometrybased and video-based coding frameworks within MPEG GSC. Gun Bang, Yiyi Liao, Alexandre Zaghetto, Marius Preda, Lu Yu 0003 |
DCC | 2 |
| 2026 | Dual-Hilbert Scan for Efficient Video-Based Gaussian Splatting Compression
Yiyi Liao, Lu Yu 0003 |
ISCAS | 2 |
| 2026 | Towards Depth Foundation Models: Recent Trends in Vision-Based Depth EstimationabstractDepth estimation is a fundamental task in 3D computer vision, crucial for applications such as 3D reconstruction, free-viewpoint rendering, robotics, autonomous driving, and AR/VR technologies. Traditional methods relying on hardware sensors like LiDAR are often limited by their high costs, low resolution, and sensitivity to the environment, limiting their applicability to real-world scenarios. Recent advances in vision-based methods offer a promising alternative, yet they face challenges in generalization and stability due to either the low capacity of model architectures or reliance on domain-specific and small-scale datasets. The emergence of scaling laws and foundation models in other domains has inspired the development of “depth foundation models”: deep neural networks trained on large datasets with strong zeroshot generalization capabilities. This paper surveys the evolution of deep learning architectures and paradigms for depth estimation across monocular, stereo, multiview, and monocular video settings. We explore the potential of these models to address existing challenges and we also provide a comprehensive overview of large-scale datasets that can facilitate their development. By identifying key architectures and training strategies, we aim to highlight the path towards robust depth foundation models, offering insights for future research and applications. Zhen Xu 0008, Sida Peng, Haotong Lin, Jiahao Shao, Peishan Yang, Qinglin Yang, Sheng Miao, Yifan Wang 0026, Ruizhen Hu, Yiyi Liao, Xiaowei Zhou 0001, Hujun Bao |
Comput. Vis. Media | 14 |
| 2026 | HUGSIM: A Real-Time, Photo-Realistic and Closed-Loop Simulator for Autonomous DrivingabstractIn the past few decades, autonomous driving algorithms have made significant progress in perception, planning, and control. However, evaluating individual components does not fully reflect the performance of entire systems, highlighting the need for more holistic assessment methods. This motivates the development of HUGSIM, a closed-loop, photo-realistic, and real-time simulator for evaluating autonomous driving algorithms. We achieve this by lifting captured 2D RGB images into the 3D space via 3D Gaussian Splatting, improving the rendering quality for closed-loop scenarios, and building the closed-loop environment. In terms of rendering, we tackle challenges of novel view synthesis in closed-loop scenarios, including viewpoint extrapolation and 360-degree vehicle rendering. Beyond novel view synthesis, HUGSIM further enables the full closed simulation loop, dynamically updating the ego and actor states and observations based on control commands. Moreover, HUGSIM offers a comprehensive benchmark across more than 70 sequences from KITTI-360, Waymo, nuScenes, and PandaSet, along with over 400 varying scenarios, providing a fair and realistic evaluation platform for existing autonomous driving algorithms. HUGSIM not only serves as an intuitive evaluation benchmark but also unlocks the potential for fine-tuning autonomous driving algorithms in a photorealistic closed-loop setting. Longzhong Lin, Yichong Lu, Dongfeng Bai, Yue Wang 0020, Andreas Geiger 0001, Yiyi Liao |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2026 | GSCodec Studio: A Modular Framework for Gaussian Splat Compressionabstract3D Gaussian Splatting and its extension to 4D dynamic scenes enable photorealistic, real-time rendering from real-world captures, positioning Gaussian Splats (GS) as a promising format for next-generation immersive media. However, their high storage requirements pose significant challenges for practical use in sharing, transmission, and storage. Despite various studies exploring GS compression from different perspectives, these efforts remain scattered across separate repositories, complicating benchmarking and the integration of best practices. To address this gap, we present GSCodec Studio, a unified and modular framework for GS reconstruction, compression, and rendering. The framework incorporates a diverse set of 3D/4D GS reconstruction methods and GS compression techniques as modular components, facilitating flexible combinations and comprehensive comparisons. By integrating best practices from community research and our own explorations, GSCodec Studio supports the development of compact representation and compression solutions for static and dynamic Gaussian Splats. Specifically, we present Static and Dynamic GSCodec: Static GSCodec achieves competitive 3D Gaussian Splat rate-distortion performance with low decoding complexity, while Dynamic GSCodec delivers advanced 4D Gaussian Splat compression performance. The code for our framework is publicly available at https://github.com/JasonLSC/GSCodec_Studio, to advance the research on Gaussian Splats compression. Sicheng Li 0003, Chengzhen Wu, Hao Li 0069, Yiyi Liao, Lu Yu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Low-Rank Approximation for Efficient Compression of Gaussian Splatting Spherical Harmonicsabstract3D Gaussian Splatting (3DGS) enables real-time, high-fidelity rendering but suffers from large model sizes, mainly due to the Spherical Harmonics (SH) coefficients used for view-dependent appearance modeling, whose parameter count scales quadratically(O(L2))with the degreeL. This paper presents the first systematic study that reveals and leverages the intrinsic low-rank structure of SH coefficients in 3DGS. Unlike prior approaches that truncate spectral energy, the proposed low-rank paradigm compactly preserves spectral information, achieving high visual quality with substantially reduced storage. Two complementary approaches are introduced. (1) SHAC-PCA (Principal Component Analysis) is a plug-and-play post-hoc compressor that retains principal spectral variance for high-fidelity compression. (2) SHAC-LST (Learned Subset Transformation) is a training-integrated approach that decomposes Alternating Current (AC) of SH coefficients into low-dimensional subset coefficients and a shared transformation matrix, offering superior compression and even quality improvements through regularization. Both methods effectively reduce SH coefficients storage complexity toO(L). Extensive experiments demonstrate that our approaches significantly reduce memory usage while maintaining or even improving rendering quality. The proposed techniques are highly versatile: SHAC-PCA can be applied to any pre-trained 3DGS model, while SHAC-LST supports end-to-end training or fine-tuning. Both methods are compatible with existing 3DGS compression pipelines, providing a practical and general solution for efficient compression of SH coefficients in 3DGS. Sicheng Li 0003, Yiyi Liao, Lu Yu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | GIFStream: 4D Gaussian-based Immersive Video with Feature StreamabstractImmersive video offers a 6-Dof-Free viewing experience, potentially playing a key role in future video technology. Recently, 4D Gaussian Splatting has gained attention as an effective approach for immersive video due to its high rendering efficiency and quality, though maintaining quality with manageable storage remains challenging. To address this, we introduce GIFStream, a novel 4D Gaussian representation using a canonical space and a deformation field enhanced with time-dependent feature streams. These feature streams enable complex motion modeling and allow efficient compression by leveraging their motion-awareness and temporal correspondence. Additionally, we incorporate both temporal and spatial compression networks for end-to-end compression. Experimental results show that GIFStream delivers high-quality immersive video at 30 Mbps, with real-time rendering and fast decoding on an RTX 4090. Hao Li 0069, Sicheng Li 0003, Abudouaihati Batuer, Lu Yu 0003, Yiyi Liao |
CVPR | 6 |
| 2025 | UrbanCAD: Towards Highly Controllable and Photorealistic 3D Vehicles for Urban Scene SimulationabstractPhotorealistic 3D vehicle models with high controllability are essential for autonomous driving simulation and data augmentation. While handcrafted CAD models provide flexible controllability, free CAD libraries often lack the high-quality materials necessary for photorealistic rendering. Conversely, reconstructed 3D models offer high-fidelity rendering but lack controllability. In this work, we introduce UrbanCAD, a framework that generates highly controllable and photorealistic 3D vehicle digital twins from a single urban image, leveraging a large collection of free 3D CAD models and handcrafted materials. To achieve this, we propose a novel pipeline that follows a retrieval-optimization manner, adapting to observational data while preserving fine-grained expert-designed priors for both geometry and material. This enables vehicles’ realistic 360° rendering, background insertion, material transfer, relighting, and component manipulation. Furthermore, given multi-view background perspective and fisheye images, we approximate environment lighting using fisheye images and reconstruct the background with 3DGS, enabling the photorealistic insertion of optimized CAD models into rendered novel view backgrounds. Experimental results demonstrate that UrbanCAD outperforms baselines in terms of photorealism. Additionally, we show that various perception models maintain their accuracy when evaluated on UrbanCAD with in-distribution configurations but degrade when applied to realistic out-of-distribution data generated by our method. This suggests that UrbanCAD is a significant advancement in creating photorealistic, safety-critical driving scenarios for downstream applications. Yichong Lu, Yichi Cai, Shangzhan Zhang, Haoji Hu, Andreas Geiger 0001, Yiyi Liao |
CVPR | 8 |
| 2025 | EVolSplat: Efficient Volume-based Gaussian Splatting for Urban View SynthesisabstractNovel view synthesis of urban scenes is essential for autonomous driving-related applications. Existing NeRF and 3DGS-based methods show promising results in achieving photorealistic renderings but require slow, per-scene optimization. We introduce EVolSplat, an efficient 3D Gaussian Splatting model for urban scenes that works in a feed-forward manner. Unlike existing feed-forward, pixelaligned 3DGS methods, which often suffer from issues like multi-view inconsistencies and duplicated content, our approach predicts 3D Gaussians across multiple frames within a unified volume using a 3D convolutional network. This is achieved by initializing 3D Gaussians with noisy depth predictions, and then refining their geometric properties in 3D space and predicting color based on 2D textures. Our model also handles distant views and the sky with a flexible hemisphere background model. This enables us to perform fast, feed-forward reconstruction while achieving real-time rendering. Experimental evaluations on the KITTI-360 and Waymo datasets show that our method achieves state-of-the-art quality compared to existing feedforward 3DGS- and NeRF-based methods. Sheng Miao, Jiaxin Huang 0012, Dongfeng Bai, Xu Yan 0005, Yue Wang 0020, Andreas Geiger 0001, Yiyi Liao |
CVPR | 9 |
| 2025 | Learning Temporally Consistent Video Depth from Video Diffusion PriorsabstractThis work addresses the challenge of streamed video depth estimation, which expects not only per-frame accuracy but, more importantly, cross-frame consistency. We argue that sharing contextual information between frames or clips is pivotal in fostering temporal consistency. Therefore, we reformulate depth prediction into a conditional generation problem to provide contextual information within a clip and across clips. Specifically, we propose a consistent context-aware training and inference strategy for arbitrarily long videos to provide cross-clip context. We sample independent noise levels for each frame within a clip during training while using a sliding window strategy and initializing overlapping frames with previously predicted frames without adding noise. Moreover, we design an effective training strategy to provide context within a clip. Extensive experimental results validate our design choices and demonstrate the superiority of our approach, dubbed ChronoDepth. Project page: xdimlab.github.io/ChronoDepth. Jiahao Shao, Youmin Zhang 0008, Yujun Shen, Vitor Campagnolo Guizilini, Yue Wang 0041, Matteo Poggi, Yiyi Liao |
CVPR | 9 |
| 2025 | Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene GenerationabstractIn this work, we introduce Prometheus, a 3D-aware latent diffusion model for text-to-3D generation at both object and scene levels in seconds. We formulate 3D scene generation as multi-view, feed-forward, pixel-aligned 3D Gaussian generation within the latent diffusion paradigm. To ensure generalizability, we build our model upon pretrained text-to-image generation model with only minimal adjustments, and further train it using a large number of images from both single-view and multi-view datasets. Furthermore, we introduce an RGB-D latent space into 3D Gaussian generation to disentangle appearance and geometry information, enabling efficient feed-forward generation of 3D Gaussians with better fidelity and geometry. Extensive experimental results demonstrate the effectiveness of our method in both feed-forward 3D Gaussian reconstruction and text-to-3D generation. Project page: Prometheus. Jiahao Shao, Yujun Shen, Andreas Geiger 0001, Yiyi Liao |
CVPR | 6 |
| 2025 | Vivid4D: Improving 4D Reconstruction from Monocular Video by Video Inpainting
Jiaxin Huang 0012, Sheng Miao, Bangbang Yang, Yuewen Ma, Yiyi Liao |
ICCV | 5 |
| 2025 | Orientation Matters: Making 3D Generative Models Orientation-AlignedabstractHumans intuitively perceive object shape and orientation from a single image, guided by strong priors about canonical poses. However, existing 3D generative models often produce misaligned results due to inconsistent training data, limiting their usability in downstream tasks. To address this gap, we introduce the task of orientation-aligned 3D object generation: producing 3D objects from single images with consistent orientations across categories. To facilitate this, we construct Objaverse-OA, a dataset of 14,832 orientation-aligned 3D models spanning 1,008 categories. Leveraging Objaverse-OA, we fine-tune two representative 3D generative models based on multi-view diffusion and 3D variational autoencoder frameworks to produce aligned objects that generalize well to unseen objects across various categories. Experimental results demonstrate the superiority of our method over post-hoc alignment approaches. Furthermore, we showcase downstream applications enabled by our aligned object generation, including zero-shot object orientation estimation via analysis-by-synthesis and efficient arrow-based object rotation manipulation. Yichong Lu, Yuzhuo Tian, Zijin Jiang, Hao Ouyang, Haoji Hu, Yujun Shen, Yiyi Liao |
NeurIPS | 10 |
| 2025 | Eliminating Geometric Representation Redundancy for 3D Gaussian Splat Codingabstract3D Gaussian Splatting (3DGS) enables photorealistic, real-time rendering, yet its native representation imposes significant storage and transmission overhead, hindering widespread deployment. Most existing approaches focus on reducing data volume or improving the efficiency of lossy coding. However, they overlook the high data entropy caused by inherent representation ambiguity, where multiple geometric attribute values can define the same geometry. To address this, we introduce a lightweight, plug-and-play preprocessing method that lowers raw data entropy by canonicalizing scale and quaternion attributes. Specifically, our method first resolves geometric representation ambiguity via a deterministic regularization rule that enforces a unique representation; reduces dimensionality by converting 4D quaternions to minimal 3D Rodrigues parameters; and addresses numerical redundancy by clamping perceptually insignificant scale values. Our method is a generic, plug-and-play preprocessing module, fully orthogonal to existing 3DGS compression schemes, and effectively boosts their coding efficiency. When combined with a baseline compression pipeline, it yields an average BD-Rate reduction of 15.81% compared to the same pipeline without our preprocessing. Shanchuan Liu, Sicheng Li 0003, Yiyi Liao, Lu Yu 0003 |
VCIP | 4 |
| 2025 | Neural mesh refinementabstractSubdivision is a widely used technique for mesh refinement. Classic methods rely on fixed manually defined weighting rules and struggle to generate a finer mesh with appropriate details, while advanced neural subdivision methods achieve data-driven nonlinear subdivision but lack robustness, suffering from limited subdivision levels and artifacts on novel shapes. To address these issues, this paper introduces a neural mesh refinement (NMR) method that uses the geometric structural priors learned from fine meshes to adaptively refine coarse meshes through subdivision, demonstrating robust generalization. Our key insight is that it is necessary to disentangle the network from non-structural information such as scale, rotation, and translation, enabling the network to focus on learning and applying the structural priors of local patches for adaptive refinement. For this purpose, we introduce an intrinsic structure descriptor and a locally adaptive neural filter. The intrinsic structure descriptor excludes the non-structural information to align local patches, thereby stabilizing the input feature space and enabling the network to robustly extract structural priors. The proposed neural filter, using a graph attention mechanism, extracts local structural features and adapts learned priors to local patches. Additionally, we observe that Charbonnier loss can alleviate over-smoothing compared to L2 loss. By combining these design choices, our method gains robust geometric learning and locally adaptive capabilities, enhancing generalization to various situations such as unseen shapes and arbitrary refinement levels. We evaluate our method on a diverse set of complex three-dimensional (3D) shapes, and experimental results show that it outperforms existing subdivision methods in terms of geometry quality. See https://zhuzhiwei99.github.io/NeuralMeshRefinement for the project page. Lu Yu 0003, Yiyi Liao |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2025 | PanopticNeRF-360: Panoramic 3D-to-2D Label Transfer in Urban ScenesabstractTraining perception systems for self-driving cars requires substantial 2D annotations that are labor-intensive to manual label. While existing datasets provide rich annotations on pre-recorded sequences, they fall short in labeling rarely encountered viewpoints, potentially hampering the generalization ability for perception models. In this paper, we present PanopticNeRF-360, a novel approach that combines coarse 3D annotations with noisy 2D semantic cues to generate high-quality panoptic labels and images from any viewpoint. Our key insight lies in exploiting the complementarity of 3D and 2D priors to mutually enhance geometry and semantics. Specifically, we propose to leverage coarse 3D bounding primitives and noisy 2D semantic and instance predictions to guide geometry optimization, by encouraging predicted labels to match panoptic pseudo ground truth. Simultaneously, the improved geometry assists in filtering 3D&2D annotation noise by fusing semantics in 3D space via a learned semantic field. To further enhance appearance, we combine MLP and hash grids to yield hybrid scene features, striking a balance between high-frequency appearance and contiguous semantics. Our experiments demonstrate PanopticNeRF-360's state-of-the-art performance over label transfer methods on the challenging urban scenes of the KITTI-360 dataset. Moreover, PanopticNeRF-360 enables omnidirectional rendering of high-fidelity, multi-view and spatiotemporally consistent appearance, semantic and instance labels. Shangzhan Zhang, Tianrun Chen, Yichong Lu, Xiaowei Zhou 0001, Andreas Geiger 0001, Yiyi Liao |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2024 | NeRFCodec: Neural Feature Compression Meets Neural Radiance Fields for Memory-Efficient Scene Representation
Sicheng Li 0003, Hao Li 0069, Yiyi Liao, Lu Yu 0003 |
CVPR | 3 |
| 2024 | HUGS: Holistic Urban 3D Scene Understanding via Gaussian SplattingabstractHolistic understanding of urban scenes based on RGB images is a challenging yet important problem. It encompasses understanding both the geometry and appearance to enable novel view synthesis, parsing semantic labels, and tracking moving objects. Despite considerable progress, existing approaches often focus on specific aspects of this task and require additional inputs such as LiDAR scans or manually annotated 3D bounding boxes. In this paper, we introduce a novel pipeline that utilizes 3D Gaussian Splatting for holistic urban scene understanding. Our main idea involves the joint optimization of geometry, appearance, semantics, and motion using a combination of static and dynamic 3D Gaussians, where moving object poses are regularized via physical constraints. Our approach offers the ability to render new viewpoints in real-time, yielding 2D and 3D semantic information with high accuracy, and reconstruct dynamic scenes, even in scenarios where 3D bounding box detection are highly noisy. Experimental results on KITTI, KITTI-360, and Virtual KITTI 2 demonstrate the effectiveness of our approach. Our project page is at https://xdimlab.github.io/hugs_website. Jiahao Shao, Dongfeng Bai, Weichao Qiu, Yue Wang 0020, Andreas Geiger 0001, Yiyi Liao |
CVPR | 9 |
| 2024 | Learning 3D-Aware GANs from Unposed Images with Template Feature Field
Xinya Chen, Hanlei Guo, Yanrui Bin, Shangzhan Zhang, Yujun Shen, Yiyi Liao |
ECCV (16) | 8 |
| 2024 | REFRAME: Reflective Surface Real-Time Rendering for Mobile Devices
Chaojie Ji, Yiyi Liao |
ECCV (45) | 3 |
| 2024 | Efficient Depth-Guided Urban View Synthesis
Sheng Miao, Jiaxin Huang 0012, Dongfeng Bai, Weichao Qiu, Andreas Geiger 0001, Yiyi Liao |
ECCV (30) | 7 |
| 2024 | NGEL-SLAM: Neural Implicit Representation-based Global Consistent Low-Latency SLAM SystemabstractNeural implicit representations have emerged as a promising solution for providing dense geometry in Simultaneous Localization and Mapping (SLAM). However, existing methods in this direction fall short in terms of global consistency and low latency. This paper presents NGEL-SLAM to tackle the above challenges. To ensure global consistency, our system leverages a traditional feature-based tracking module that incorporates loop closure. Additionally, we maintain a global consistent map by representing the scene using multiple neural implicit fields, enabling quick adjustment to the loop closure. Moreover, our system allows for fast convergence through the use of octree-based implicit representations. The combination of rapid response to loop closure and fast convergence makes our system a truly low-latency system that achieves global consistency. Our system enables rendering high-fidelity RGB-D images, along with extracting dense and complete surfaces. Experiments on both synthetic and real-world datasets suggest that our system achieves state-of-the-art tracking and mapping accuracy while maintaining low latency. Yunxuan Mao, Zhuqing Zhang, Yue Wang 0020, Rong Xiong, Yiyi Liao |
ICRA | 7 |
| 2024 | ν-DBA: Neural Implicit Dense Bundle Adjustment Enables Image-Only Driving Scene ReconstructionabstractThe joint optimization of the sensor trajectory and 3D map is a crucial characteristic of bundle adjustment (BA), essential for autonomous driving. This paper presents ν-DBA, a novel framework implementing geometric dense bundle adjustment (DBA) using 3D neural implicit surfaces for map parametrization, which optimizes both the map surface and trajectory poses using geometric error guided by dense optical flow prediction. Additionally, we fine-tune the optical flow model with per-scene self-supervision to further improve the quality of the dense mapping. Our experimental results on multiple driving scene datasets demonstrate that our method achieves superior trajectory optimization and dense reconstruction accuracy. We also investigate the influences of photometric error and different neural geometric priors on the performance of surface reconstruction and novel view synthesis. Our method stands as a significant step towards leveraging neural implicit representations in dense bundle adjustment for more accurate trajectories and detailed environmental mapping. Yunxuan Mao, Bingqi Shen, Rong Xiong, Yiyi Liao, Yue Wang 0020 |
IROS | 6 |
| 2024 | PanopticRecon: Leverage Open-vocabulary Instance Segmentation for Zero-shot Panoptic ReconstructionabstractPanoptic reconstruction is a challenging task in 3D scene understanding. However, most existing methods heavily rely on pre-trained semantic segmentation models and known 3D object bounding boxes for 3D panoptic segmentation, which is not available for in-the-wild scenes. In this paper, we propose a novel zero-shot panoptic reconstruction method from RGB-D images of scenes. For zero-shot segmentation, we leverage open-vocabulary instance segmentation, but it has to face partial labeling and instance association challenges. We tackle both challenges by propagating partial labels with the aid of dense generalized features and building a 3D instance graph for associating 2D instance IDs. Specifically, we exploit partial labels to learn a classifier for generalized semantic features to provide complete labels for scenes with dense distilled features. Moreover, we formulate instance association as a 3D instance graph segmentation problem, allowing us to fully utilize the scene geometry prior and all 2D instance masks to infer global unique pseudo 3D instance ID. Our method outperforms state-of-the-art methods on the indoor dataset ScanNet V2 and the outdoor dataset KITTI-360, demonstrating the effectiveness of our graph segmentation method and reconstruction network. Yili Liu, Chenrui Han, Sitong Mao, Shunbo Zhou, Rong Xiong, Yiyi Liao, Yue Wang 0020 |
IROS | 7 |
| 2024 | PET-NeRV: Bridging Generalized Video Codec and Content-Specific Neural RepresentationabstractGeneralized models dominate neural video compression methods to compress arbitrary video with the same model. However, obtaining a universal neural codec with high compression efficiency on all videos is challenging. Existing work attempts to adapt the decoder-side model per video content through full parameter tuning, but this requires a large Group of Pictures (GOP) to compensate for the cost of transmitting updated parameters. To tackle this challenge, we propose to tune generalized video codecs per video content in a parameter-efficient manner. The per-content tuned parameters are further compressed with entropy coding using adaptive distribution estimations. This allows for enhancing the compression efficiency while maintaining a normal GOP size for random access capabilities. To validate the generality and validity of our approach, we apply it to two representative methods: CNN-based, DCVC-HEM, and Transformer-based, VCT. Our results demonstrate that introducing content-specific representation leads to a notable improvement in compression efficiency compared to the original methods. Hao Li 0069, Lu Yu 0003, Yiyi Liao |
VCIP | 3 |
| 2024 | Recent Trends in 3D Reconstruction of General Non-Rigid ScenesabstractAbstract Reconstructing models of the real world, including 3D geometry, appearance, and motion of real scenes, is essential for computer graphics and computer vision. It enables the synthesizing of photorealistic novel views, useful for the movie industry and AR/VR applications. It also facilitates the content creation necessary in computer games and AR/VR by avoiding laborious manual design processes. Further, such models are fundamental for intelligent computing systems that need to interpret real‐world scenes and actions to act and interact safely with the human world. Notably, the world surrounding us is dynamic, and reconstructing models of dynamic, non‐rigidly moving scenes is a severely underconstrained and challenging problem. This state‐of‐the‐art report (STAR) offers the reader a comprehensive summary of state‐of‐the‐art techniques with monocular and multi‐view inputs such as data from RGB and RGB‐D sensors, among others, conveying an understanding of different approaches, their potential applications, and promising further research directions. The report covers 3D reconstruction of general non‐rigid scenes and further addresses the techniques for scene decomposition, editing and controlling, and generalizable and generative modeling. More specifically, we first review the common and fundamental concepts necessary to understand and navigate the field and then discuss the state‐of‐the‐art techniques by reviewing recent approaches that use traditional and machine‐learning‐based neural representations, including a discussion on the newly enabled applications. The STAR is concluded with a discussion of the remaining limitations and open challenges. Raza Yunus, Jan Eric Lenssen, Michael Niemeyer, Yiyi Liao, Christian Rupprecht 0001, Christian Theobalt, Gerard Pons-Moll, Jia-Bin Huang 0001, Vladislav Golyanik, Eddy Ilg |
Comput. Graph. Forum | 4 |
| 2024 | Reality3DSketch: Rapid 3D Modeling of Objects From Single Freehand SketchesabstractThe emerging trend of AR/VR places great demands on 3D content. However, most existing software requires expertise and is difficult for novice users to use. In this paper, we aim to create sketch-based modeling tools for user-friendly 3D modeling. We introduce Reality3DSketch with a novel application of an immersive 3D modeling experience, in which a user can capture the surrounding scene using a monocular RGB camera and can draw a single sketch of an object in the real-time reconstructed 3D scene. A 3D object is generated and placed in the desired location, enabled by our novel neural network with the input of a single sketch. Our neural network can predict the pose of a drawing and can turn a single sketch into a 3D model with view and structural awareness, which addresses the challenge of sparse sketch input and view ambiguity. We conducted extensive experiments synthetic and real-world datasets and achieved state-of-the-art (SOTA) results in both sketch view estimation and 3D modeling performance. According to our user study, our method of performing 3D modeling in a scene is$>$5x faster than conventional methods. Users are also more satisfied with the generated 3D model than the results of existing methods. Tianrun Chen, Chaotao Ding, Lanyun Zhu, Ying Zang, Yiyi Liao, Zejian Li, Lingyun Sun |
IEEE Trans. Multim. | 5 |
| 2023 | SteerNeRF: Accelerating NeRF Rendering via Smooth Viewpoint TrajectoryabstractNeural Radiance Fields (NeRF) have demonstrated superior novel view synthesis performance but are slow at rendering. To speed up the volume rendering process, many acceleration methods have been proposed at the cost of large memory consumption. To push the frontier of the efficiency-memory trade-off, we explore a new perspective to accelerate NeRF rendering, leveraging a key fact that the view-point change is usually smooth and continuous in interactive viewpoint control. This allows us to leverage the information of preceding viewpoints to reduce the number of rendered pixels as well as the number of sampled points along the ray of the remaining pixels. In our pipeline, a low-resolution feature map is rendered first by volume rendering, then a lightweight 2D neural renderer is applied to generate the output image at target resolution leveraging the features of preceding and current frames. We show that the proposed method can achieve competitive rendering quality while reducing the rendering time with little memory overhead, enabling 30FPS at 1080P image resolution with a low memory footprint. Sicheng Li 0003, Hao Li 0069, Yue Wang 0020, Yiyi Liao, Lu Yu 0003 |
CVPR | 4 |
| 2023 | Learning 3D-Aware Image Synthesis with Unknown Pose DistributionabstractExisting methods for 3D-aware image synthesis largely depend on the 3D pose distribution pre-estimated on the training set. An inaccurate estimation may mislead the model into learning faulty geometry. This work proposes PoF3D that frees generative radiance fields from the requirements of 3D pose priors. We first equip the generator with an efficient pose learner, which is able to infer a pose from a latent code, to approximate the underlying true pose distribution automatically. We then assign the discriminator a task to learn pose distribution under the supervision of the generator and to differentiate real and synthesized images with the predicted pose as the condition. The pose-free generator and the pose-aware discriminator are jointly trained in an adversarial manner. Extensive results on a couple of datasets confirm that the performance of our approach, regarding both image quality and geometry quality, is on par with state of the art. To our best knowledge, PoF3D demonstrates the feasibility of learning high-quality 3D-aware image synthesis without using 3D pose priors for the first time. Project page can be found here. Zifan Shi, Yujun Shen, Yinghao Xu 0001, Sida Peng, Yiyi Liao, Qifeng Chen 0001, Dit-Yan Yeung |
CVPR | 5 |
| 2023 | Painting 3D Nature in 2D: View Synthesis of Natural Scenes from a Single Semantic MaskabstractWe introduce a novel approach that takes a single semantic mask as input to synthesize multi-view consistent color images of natural scenes, trained with a collection of single images from the Internet. Prior works on 3D-aware image synthesis either require multi-view supervision or learning category-level prior for specific classes of objects, which are inapplicable to natural scenes. Our key idea to solve this challenge is to use a semantic field as the intermediate representation, which is easier to reconstruct from an input semantic mask and then translated to a radiance field with the assistance of off-the-shelf semantic image synthesis models. Experiments show that our method outperforms baseline methods and produces photorealistic and multi-view consistent videos of a variety of natural scenes. The project website is https://zju3dv.github.io/paintingnature/. Shangzhan Zhang, Sida Peng, Tianrun Chen, Linzhan Mou, Haotong Lin, Kaicheng Yu, Yiyi Liao, Xiaowei Zhou 0001 |
CVPR | 7 |
| 2023 | VeRi3D: Generative Vertex-based Radiance Fields for 3D Controllable Human Image SynthesisabstractUnsupervised learning of 3D-aware generative adversarial networks has lately made much progress. Some recent work demonstrates promising results of learning human generative models using neural articulated radiance fields, yet their generalization ability and controllability lag behind parametric human models, i.e., they do not perform well when generalizing to novel pose/shape and are not part controllable. To solve these problems, we propose VeRi3D, a generative human vertex-based radiance field parameterized by vertices of the parametric human template, SMPL. We map each 3D point to the local coordinate system defined on its neighboring vertices, and use the corresponding vertex feature and local coordinates for mapping it to color and density values. We demonstrate that our simple approach allows for generating photorealistic human images with free control over camera pose, human pose, shape, as well as enabling part-level editing. Xinya Chen, Jiaxin Huang 0012, Yanrui Bin, Lu Yu 0003, Yiyi Liao |
ICCV | 5 |
| 2023 | RICO: Regularizing the Unobservable for Indoor Compositional ReconstructionabstractRecently, neural implicit surfaces have become popular for multi-view reconstruction. To facilitate practical applications like scene editing and manipulation, some works extend the framework with semantic masks input for the object-compositional reconstruction rather than the holistic perspective. Though achieving plausible disentanglement, the performance drops significantly when processing the indoor scenes where objects are usually partially observed. We propose RICO to address this by regularizing the unobservable regions for indoor compositional reconstruction. Our key idea is to first regularize the smoothness of the occluded background, which then in turn guides the foreground object reconstruction in unobservable regions based on the object-background relationship. Particularly, we regularize the geometry smoothness of occluded background patches. With the improved background surface, the signed distance function and the reversedly rendered depth of objects can be optimized to bound them within the background range. Extensive experiments show our method outperforms other methods on synthetic and real-world indoor scenes and prove the effectiveness of proposed regularizations. The code is available at https://github.com/kyleleey/RICO Zizhang Li, Xiaoyang Lyu, Yuanyuan Ding, Mengmeng Wang 0005, Yiyi Liao, Yong Liu 0007 |
ICCV | 5 |
| 2023 | UrbanGIRAFFE: Representing Urban Scenes as Compositional Generative Neural Feature FieldsabstractGenerating photorealistic images with controllable camera pose and scene contents is essential for many applications including AR/VR and simulation. Despite the fact that rapid progress has been made in 3D-aware generative models, most existing methods focus on object-centric images and are not applicable to generating urban scenes for free camera viewpoint control and scene editing. To address this challenging task, we propose UrbanGIRAFFE, which uses a coarse 3D panoptic prior, including the layout distribution of uncountable stuff and countable objects, to guide a 3D-aware generative model. Our model is compositional and controllable as it breaks down the scene into stuff, objects, and sky. Using stuff prior in the form of semantic voxel grids, we build a conditioned stuff generator that effectively incorporates the coarse semantic and geometry information. The object layout prior further allows us to learn an object generator from cluttered scenes. With proper loss functions, our approach facilitates photorealistic 3D-aware image synthesis with diverse controllability, including large camera movement, stuff editing, and object manipulation. We validate the effectiveness of our model on both synthetic and real-world datasets, including the challenging KITTI-360 dataset. Hanlei Guo, Rong Xiong, Yue Wang 0020, Yiyi Liao |
ICCV | 6 |
| 2023 | DPCN++: Differentiable Phase Correlation Network for Versatile Pose RegistrationabstractPose registration is critical in vision and robotics. This article focuses on the challenging task of initialization-free pose registration up to 7DoF for homogeneous and heterogeneous measurements. While recent learning-based methods show promise using differentiable solvers, they either rely on heuristically defined correspondences or require initialization. Phase correlation seeks solutions in the spectral domain and is correspondence-free and initialization-free. Following this, we propose a differentiable solver and combine it with simple feature extraction networks, namely DPCN++. It can perform registration for homo/hetero inputs and generalizes well on unseen objects. Specifically, the feature extraction networks first learn dense feature grids from a pair of homogeneous/heterogeneous measurements. These feature grids are then transformed into a translation and scale invariant spectrum representation based on Fourier transform and spherical radial aggregation, decoupling translation and scale from rotation. Next, the rotation, scale, and translation are independently and efficiently estimated in the spectrum step-by-step. The entire pipeline is differentiable and trained end-to-end. We evaluate DCPN++ on a wide range of tasks taking different input modalities, including 2D bird's-eye view images, 3D object and scene measurements, and medical images. Experimental results demonstrate that DCPN++ outperforms both classical and learning-based baselines, especially on partially observed and heterogeneous measurements. Zexi Chen, Yiyi Liao, Haozhe Du, Xuecheng Xu, Haojian Lu, Rong Xiong, Yue Wang 0020 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | KITTI-360: A Novel Dataset and Benchmarks for Urban Scene Understanding in 2D and 3DabstractFor the last few decades, several major subfields of artificial intelligence including computer vision, graphics, and robotics have progressed largely independently from each other. Recently, however, the community has realized that progress towards robust intelligent systems such as self-driving cars requires a concerted effort across the different fields. This motivated us to develop KITTI-360, successor of the popular KITTI dataset. KITTI-360 is a suburban driving dataset which comprises richer input modalities, comprehensive semantic instance annotations and accurate localization to facilitate research at the intersection of vision, graphics and robotics. For efficient annotation, we created a tool to label 3D scenes with bounding primitives and developed a model that transfers this information into the 2D image domain, resulting in over 150k images and 1B 3D points with coherent semantic instance annotations across 2D and 3D. Moreover, we established benchmarks and baselines for several tasks relevant to mobile perception, encompassing problems from computer vision, graphics, and robotics on the same dataset, e.g., semantic scene understanding, novel view synthesis and semantic SLAM. KITTI-360 will enable progress at the intersection of these research areas and thus contribute towards solving one of today's grand challenges: the development of fully autonomous self-driving systems. Yiyi Liao, Andreas Geiger 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | RING++: Roto-Translation Invariant Gram for Global Localization on a Sparse Scan MapabstractGlobal localization plays a critical role in many robot applications. LiDAR-based global localization draws the community's focus with its robustness against illumination and seasonal changes. To further improve the localization under large viewpoint differences, we propose RING++ that has roto-translation-invariant representation for place recognition and global convergence for both rotation and translation estimation. With the theoretical guarantee, RING++ is able to address the large viewpoint difference using a lightweight map with sparse scans. In addition, we derive sufficient conditions of feature extractors for the representation preserving the roto-translation invariance, making RING++ a framework applicable to generic multichannel features. To the best of our knowledge, this is the first learning-free framework to address all the subtasks of global localization in the sparse scan map. Validations on real-world datasets show that our approach demonstrates better performance than state-of-the-art learning-free methods and competitive performance with learning-based methods. Finally, we integrate RING++ into a multirobot/session simultaneous localization and mapping system, performing its effectiveness in collaborative applications. Xuecheng Xu, Jun Wu 0003, Haojian Lu, Qiuguo Zhu, Yiyi Liao, Rong Xiong, Yue Wang 0020 |
IEEE Trans. Robotics | 6 |
| 2022 | Panoptic NeRF: 3D-to-2D Label Transfer for Panoptic Urban Scene SegmentationabstractLarge-scale training data with high-quality annotations is critical for training semantic and instance segmentation models. Unfortunately, pixel-wise annotation is labor-intensive and costly, raising the demand for more efficient labeling strategies. In this work, we present a novel 3D-to-2D label transfer method, Panoptic NeRF1, which aims for obtaining per-pixel 2D semantic and instance labels from easy-to-obtain coarse 3D bounding primitives. Our method utilizes NeRF as a differentiable tool to unify coarse 3D annotations and 2D semantic cues transferred from existing datasets. We demonstrate that this combination allows for improved geometry guided by semantic information, enabling rendering of accurate semantic maps across multiple views. Furthermore, this fusion process resolves label ambiguity of the coarse 3D annotations and filters noise in the 2D predictions. By inferring in 3D space and rendering to 2D labels, our 2D semantic and instance labels are multiview consistent by design. Experimental results show that Panoptic NeRF outperforms existing label transfer methods in terms of accuracy and multi-view consistency on challenging urban scenes of the KITTI-360 dataset. Shangzhan Zhang, Tianrun Chen, Yichong Lu, Lanyun Zhu, Xiaowei Zhou 0001, Andreas Geiger 0001, Yiyi Liao |
3DV | 8 |
| 2022 | A Visual Navigation Perspective for Category-Level Object Pose Estimation
Fangxun Zhong, Rong Xiong, Yun-Hui Liu 0001, Yue Wang 0020, Yiyi Liao |
ECCV (6) | 6 |
| 2022 | VoxGRAF: Fast 3D-Aware Image Synthesis with Sparse Voxel GridsabstractState-of-the-art 3D-aware generative models rely on coordinate-based MLPs to parameterize 3D radiance fields. While demonstrating impressive results, querying an MLP for every sample along each ray leads to slow rendering.Therefore, existing approaches often render low-resolution feature maps and process them with an upsampling network to obtain the final image. Albeit efficient, neural rendering often entangles viewpoint and content such that changing the camera pose results in unwanted changes of geometry or appearance.Motivated by recent results in voxel-based novel view synthesis, we investigate the utility of sparse voxel grid representations for fast and 3D-consistent generative modeling in this paper.Our results demonstrate that monolithic MLPs can indeed be replaced by 3D convolutions when combining sparse voxel grids with progressive growing, free space pruning and appropriate regularization.To obtain a compact representation of the scene and allow for scaling to higher voxel resolutions, our model disentangles the foreground object (modeled in 3D) from the background (modeled in 2D).In contrast to existing approaches, our method requires only a single forward pass to generate a full 3D scene. It hence allows for efficient rendering from arbitrary viewpoints while yielding 3D consistent results with high visual fidelity. Code and models are available at https://github.com/autonomousvision/voxgraf. Katja Schwarz, Axel Sauer, Michael Niemeyer, Yiyi Liao, Andreas Geiger 0001 |
NeurIPS | 4 |
| 2021 | SMD-Nets: Stereo Mixture Density NetworksabstractDespite stereo matching accuracy has greatly improved by deep learning in the last few years, recovering sharp boundaries and high-resolution outputs efficiently remains challenging. In this paper, we propose Stereo Mixture Density Networks (SMD-Nets), a simple yet effective learning framework compatible with a wide class of 2D and 3D architectures which ameliorates both issues. Specifically, we exploit bimodal mixture densities as output representation and show that this allows for sharp and precise disparity estimates near discontinuities while explicitly modeling the aleatoric uncertainty inherent in the observations. Moreover, we formulate disparity estimation as a continuous problem in the image domain, allowing our model to query disparities at arbitrary spatial precision. We carry out comprehensive experiments on a new high-resolution and highly realistic synthetic stereo dataset, consisting of stereo pairs at 8Mpx resolution, as well as on real-world stereo datasets. Our experiments demonstrate increased depth accuracy near object boundaries and prediction of ultra high-resolution disparity maps on standard GPUs. We demonstrate the flexibility of our technique by improving the performance of a variety of stereo backbones. Fabio Tosi, Yiyi Liao, Carolin Schmitt, Andreas Geiger 0001 |
CVPR | 2 |
| 2021 | KiloNeRF: Speeding up Neural Radiance Fields with Thousands of Tiny MLPsabstractNeRF synthesizes novel views of a scene with unprecedented quality by fitting a neural radiance field to RGB images. However, NeRF requires querying a deep Multi-Layer Perceptron (MLP) millions of times, leading to slow rendering times, even on modern GPUs. In this paper, we demonstrate that real-time rendering is possible by utilizing thousands of tiny MLPs instead of one single large MLP. In our setting, each individual MLP only needs to represent parts of the scene, thus smaller and faster-to-evaluate MLPs can be used. By combining this divide-and-conquer strategy with further optimizations, rendering is accelerated by three orders of magnitude compared to the original NeRF model without incurring high storage costs. Further, using teacher-student distillation for training, we show that this speed-up can be achieved without sacrificing visual quality. Christian Reiser, Songyou Peng, Yiyi Liao, Andreas Geiger 0001 |
ICCV | 3 |
| 2021 | Shape As Points: A Differentiable Poisson SolverabstractIn recent years, neural implicit representations gained popularity in 3D reconstruction due to their expressiveness and flexibility. However, the implicit nature of neural implicit representations results in slow inference times and requires careful initialization. In this paper, we revisit the classic yet ubiquitous point cloud representation and introduce a differentiable point-to-mesh layer using a differentiable formulation of Poisson Surface Reconstruction (PSR) which allows for a GPU-accelerated fast solution of the indicator function given an oriented point cloud. The differentiable PSR layer allows us to efficiently and differentiably bridge the explicit 3D point representation with the 3D mesh via the implicit indicator field, enabling end-to-end optimization of surface reconstruction metrics such as Chamfer distance. This duality between points and meshes hence allows us to represent shapes as oriented point clouds, which are explicit, lightweight and expressive. Compared to neural implicit representations, our Shape-As-Points (SAP) model is more interpretable, lightweight, and accelerates inference time by one order of magnitude. Compared to other explicit representations such as points, patches, and meshes, SAP produces topology-agnostic, watertight manifold surfaces. We demonstrate the effectiveness of SAP on the task of surface reconstruction from unoriented point clouds and learning-based reconstruction. Songyou Peng, Chiyu Max Jiang, Yiyi Liao, Michael Niemeyer, Marc Pollefeys, Andreas Geiger 0001 |
NeurIPS | 3 |
| 2021 | On the Frequency Bias of Generative ModelsabstractThe key objective of Generative Adversarial Networks (GANs) is to generate new data with the same statistics as the provided training data. However, multiple recent works show that state-of-the-art architectures yet struggle to achieve this goal. In particular, they report an elevated amount of high frequencies in the spectral statistics which makes it straightforward to distinguish real and generated images. Explanations for this phenomenon are controversial: While most works attribute the artifacts to the generator, other works point to the discriminator. We take a sober look at those explanations and provide insights on what makes proposed measures against high-frequency artifacts effective. To achieve this, we first independently assess the architectures of both the generator and discriminator and investigate if they exhibit a frequency bias that makes learning the distribution of high-frequency content particularly problematic. Based on these experiments, we make the following four observations: 1) Different upsampling operations bias the generator towards different spectral properties. 2) Checkerboard artifacts introduced by upsampling cannot explain the spectral discrepancies alone as the generator is able to compensate for these artifacts. 3) The discriminator does not struggle with detecting high frequencies per se but rather struggles with frequencies of low magnitude. 4) The downsampling operations in the discriminator can impair the quality of the training signal it provides.In light of these findings, we analyze proposed measures against high-frequency artifacts in state-of-the-art GAN training but find that none of the existing approaches can fully resolve spectral artifacts yet. Our results suggest that there is great potential in improving the discriminator and that this could be key to match the distribution of the training data more closely. Katja Schwarz, Yiyi Liao, Andreas Geiger 0001 |
NeurIPS | 2 |
| 2021 | Learning Steering Kernels for Guided Depth CompletionabstractThis paper addresses the guided depth completion task in which the goal is to predict a dense depth map given a guidance RGB image and sparse depth measurements. Recent advances on this problem nurture hopes that one day we can acquire accurate and dense depth at a very low cost. A major challenge of guided depth completion is to effectively make use of extremely sparse measurements, e.g., measurements covering less than 1% of the image pixels. In this paper, we propose a fully differentiable model that avoids convolving on sparse tensors by jointly learning depth interpolation and refinement. More specifically, we propose a differentiable kernel regression layer that interpolates the sparse depth measurements via learned kernels. We further refine the interpolated depth map using a residual depth refinement layer which leads to improved performance compared to learning absolute depth prediction using a vanilla network. We provide experimental evidence that our differentiable kernel regression layer not only enables end-to-end training from very sparse measurements using standard convolutional network architectures, but also leads to better depth interpolation results compared to existing heuristically motivated methods. We demonstrate that our method outperforms many state-of-the-art guided depth completion techniques on both NYUv2 and KITTI. We further show the generalization ability of our method with respect to the density and spatial statistics of the sparse depth measurements. Lina Liu 0010, Yiyi Liao, Yue Wang 0020, Andreas Geiger 0001, Yong Liu 0007 |
IEEE Trans. Image Process. | 2 |
| 2020 | Towards Unsupervised Learning of Generative Models for 3D Controllable Image SynthesisabstractIn recent years, Generative Adversarial Networks have achieved impressive results in photorealistic image synthesis. This progress nurtures hopes that one day the classical rendering pipeline can be replaced by efficient models that are learned directly from images. However, current image synthesis models operate in the 2D domain where disentangling 3D properties such as camera viewpoint or object pose is challenging. Furthermore, they lack an interpretable and controllable representation. Our key hypothesis is that the image generation process should be modeled in 3D space as the physical world surrounding us is intrinsically three-dimensional. We define the new task of 3D controllable image synthesis and propose an approach for solving it by reasoning both in 3D space and in the 2D image domain. We demonstrate that our model is able to disentangle latent 3D factors of simple multi-object scenes in an unsupervised fashion from raw images. Compared to pure 2D baselines, it allows for synthesizing scenes that are consistent wrt. changes in viewpoint or object pose. We further evaluate various 3D representations in terms of their usefulness for this challenging task. Yiyi Liao, Katja Schwarz, Lars M. Mescheder, Andreas Geiger 0001 |
CVPR | 1 |
| 2020 | GRAF: Generative Radiance Fields for 3D-Aware Image SynthesisabstractWhile 2D generative adversarial networks have enabled high-resolution image synthesis, they largely lack an understanding of the 3D world and the image formation process. Thus, they do not provide precise control over camera viewpoint or object pose. To address this problem, several recent approaches leverage intermediate voxel-based representations in combination with differentiable rendering. However, existing methods either produce low image resolution or fall short in disentangling camera and scene properties, e.g., the object identity may vary with the viewpoint. In this paper, we propose a generative model for radiance fields which have recently proven successful for novel view synthesis of a single scene. In contrast to voxel-based representations, radiance fields are not confined to a coarse discretization of the 3D space, yet allow for disentangling camera and scene properties while degrading gracefully in the presence of reconstruction ambiguity. By introducing a multi-scale patch-based discriminator, we demonstrate synthesis of high-resolution images while training our model from unposed 2D images alone. We systematically analyze our approach on several challenging synthetic and real-world datasets. Our experiments reveal that radiance fields are a powerful representation for generative image synthesis, leading to 3D consistent models that render with high fidelity. Katja Schwarz, Yiyi Liao, Michael Niemeyer, Andreas Geiger 0001 |
NeurIPS | 2 |
| 2019 | Connecting the Dots: Learning Representations for Active Monocular Depth EstimationabstractWe propose a technique for depth estimation with a monocular structured-light camera, i.e., a calibrated stereo set-up with one camera and one laser projector. Instead of formulating the depth estimation via a correspondence search problem, we show that a simple convolutional architecture is sufficient for high-quality disparity estimates in this setting. As accurate ground-truth is hard to obtain, we train our model in a self-supervised fashion with a combination of photometric and geometric losses. Further, we demonstrate that the projected pattern of the structured light sensor can be reliably separated from the ambient information. This can then be used to improve depth boundaries in a weakly supervised fashion by modeling the joint statistics of image and depth edges. The model trained in this fashion compares favorably to the state-of-the-art on challenging synthetic and real-world datasets. In addition, we contribute a novel simulator, which allows to benchmark active depth prediction algorithms in controlled conditions. Gernot Riegler, Yiyi Liao, Simon Donné, Vladlen Koltun, Andreas Geiger 0001 |
CVPR | 2 |
| 2018 | Deep Marching Cubes: Learning Explicit Surface RepresentationsabstractExisting learning based solutions to 3D surface prediction cannot be trained end-to-end as they operate on intermediate representations (e.g., TSDF) from which 3D surface meshes must be extracted in a post-processing step (e.g., via the marching cubes algorithm). In this paper, we investigate the problem of end-to-end 3D surface prediction. We first demonstrate that the marching cubes algorithm is not differentiable and propose an alternative differentiable formulation which we insert as a final layer into a 3D convolutional neural network. We further propose a set of loss functions which allow for training our model with sparse point supervision. Our experiments demonstrate that the model allows for predicting sub-voxel accurate 3D shapes of arbitrary topology. Additionally, it learns to complete shapes and to separate an object's inside from its outside even in the presence of sparse and incomplete ground truth. We investigate the benefits of our approach on the task of inferring shapes from 3D point clouds. Our model is flexible and can be combined with a variety of shape encoder and shape inference techniques. Yiyi Liao, Simon Donné, Andreas Geiger 0001 |
CVPR | 1 |
| 2017 | Parse geometry from a line: Monocular depth estimation with partial laser observationabstractMany standard robotic platforms are equipped with at least a fixed 2D laser range finder and a monocular camera. Although those platforms do not have sensors for 3D depth sensing capability, knowledge of depth is an essential part in many robotics activities. Therefore, recently, there is an increasing interest in depth estimation using monocular images. As this task is inherently ambiguous, the data-driven estimated depth might be unreliable in robotics applications. In this paper, we have attempted to improve the precision of monocular depth estimation by introducing 2D planar observation from the remaining laser range finder without extra cost. Specifically, we construct a dense reference map from the sparse laser range data, redefining the depth estimation task as estimating the distance between the real and the reference depth. To solve the problem, we construct a novel residual of residual neural network, and tightly combine the classification and regression losses for continuous depth estimation. Experimental results suggest that our method achieves considerable promotion compared to the state-of-the-art methods on both NYUD2 and KITTI, validating the effectiveness of our method on leveraging the additional sensory information. We further demonstrate the potential usage of our method in obstacle avoidance where our methodology provides comprehensive depth information compared to the solution using monocular camera or 2D laser range finder alone. Yiyi Liao, Lichao Huang, Yue Wang 0020, Sarath Kodagoda, Yinan Yu, Yong Liu 0007 |
ICRA | 1 |
| 2017 | Graph Regularized Auto-Encoders for Image RepresentationabstractImage representation has been intensively explored in the domain of computer vision for its significant influence on the relative tasks such as image clustering and classification. It is valuable to learn a low-dimensional representation of an image which preserves its inherent information from the original image space. At the perspective of manifold learning, this is implemented with the local invariant idea to capture the intrinsic low-dimensional manifold embedded in the high-dimensional input space. Inspired by the recent successes of deep architectures, we propose a local invariant deep nonlinear mapping algorithm, called graph regularized auto-encoder (GAE). With the graph regularization, the proposed method preserves the local connectivity from the original image space to the representation space, while the stacked auto-encoders provide explicit encoding model for fast inference and powerful expressive capacity for complex modeling. Theoretical analysis shows that the graph regularizer penalizes the weighted Frobenius norm of the Jacobian matrix of the encoder mapping, where the weight matrix captures the local property in the input space. Furthermore, the underlying effects on the hidden representation space are revealed, providing insightful explanation to the advantage of the proposed method. Finally, the experimental results on both clustering and classification tasks demonstrate the effectiveness of our GAE as well as the correctness of the proposed theoretical analysis, and it also suggests that GAE is a superior solution to the current deep representation learning techniques comparing with variant auto-encoders and existing local invariant methods. Yiyi Liao, Yue Wang 0020, Yong Liu 0007 |
IEEE Trans. Image Process. | 1 |
| 2017 | Scalable Learning Framework for Traversable Region Detection Fusing With Appearance and Geometrical InformationabstractIn this paper, we present an online learning framework for traversable region detection fusing both appearance and geometry information. Our framework proposes an appearance classifier supervised by the sparse geometric clues to capture the variation in online data, yielding dense detection result in real time. It provides superior detection performance using appearance information with weak geometric prior and can be further improved with more geometry from external sensors. The learning process is divided into three steps: First, we construct features from the super-pixel level, which reduces the computational cost compared with the pixel level processing. Then we classify the multi-scale super-pixels to vote the label of each pixel. Second, we use weighted extreme learning machine as our classifier to deal with the imbalanced data distribution since the weak geometric prior only initializes the labels in a small region. Finally, we employ the online learning process so that our framework can be adaptive to the changing scenes. Experimental results on three different styles of image sequences, i.e., shadow road, rain sequence, and variational sequence, demonstrate the adaptability, stability, and parameter insensitivity of our weak geometry motivated method. We further demonstrate the performance of learning framework on additional five challenging data sets captured by Kinect V2 and stereo camera, validating the method's effectiveness and efficiency. Yue Wang 0020, Yong Liu 0007, Yiyi Liao, Rong Xiong |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2016 | Understand scene categories by objects: A semantic regularized scene classifier using Convolutional Neural NetworksabstractScene classification is a fundamental perception task for environmental understanding in today's robotics. In this paper, we have attempted to exploit the use of popular machine learning technique of deep learning to enhance scene understanding, particularly in robotics applications. As scene images have larger diversity than the iconic object images, it is more challenging for deep learning methods to automatically learn features from scene images with less samples. Inspired by human scene understanding based on object knowledge, we address the problem of scene classification by encouraging deep neural networks to incorporate object-level information. This is implemented with a regularization of semantic segmentation. With only 5 thousand training images, as opposed to 2.5 million images, we show the proposed deep architecture achieves superior scene classification results to the state-of-the-art on a publicly available SUN RGB-D dataset. In addition, performance of semantic segmentation, the regularizer, also reaches a new record with refinement derived from predicted scene labels. Finally, we apply our model trained on SUN RGB-D dataset to a set of images captured in our university using a mobile robot, demonstrating the generalization ability of the proposed algorithm. Yiyi Liao, Sarath Kodagoda, Yue Wang 0020, Lei Shi 0013, Yong Liu 0007 |
ICRA | 1 |
| 2016 | General subspace constrained non-negative matrix factorization for data representation
Yong Liu 0007, Yiyi Liao, Weicong Liu |
Neurocomputing | 2 |
| 2015 | LOIND: An illumination and scale invariant RGB-D descriptorabstractWe introduce a novel RGB-D descriptor called local ordinal intensity and normal descriptor (LOIND) with the integration of texture information in RGB image and geometric information in depth image. We implement the descriptor with a 3-D histogram supported by orders of intensities and angles between normal vectors, in addition with the spatial sub-divisions. The former ordering information which is invariant under the transformation of illumination, scale and rotation provides the robustness of our descriptor, while the latter spatial distribution provides higher information capacity so that the discriminative performance is promoted. Comparable experiments with the state-of-art descriptors, e.g. SIFT, SURF, CSHOT and BRAND, show the effectiveness of our LOIND to the complex illumination changes and scale transformation. We also provide a new method to estimate the dominant orientation with only the geometric information, which can ensure the rotation invariance under extremely poor illumination. Guanghua Feng, Yong Liu 0007, Yiyi Liao |
ICRA | 3 |
| 2015 | Traversable region detection with a learning frameworkabstractIn this paper, we present a novel learning framework for traversable region detection. Firstly, we construct features from the super-pixel level which can reduce the computational cost compared to pixel level. Multi-scale super-pixels are extracted to give consideration to both outline and detail information. Then we classify the multiple-scale super-pixels and merge the labels in pixel level. Meanwhile, we use weighted ELM as our classifier which can deal with the imbalanced class distribution since we only assume that a small region in front of robot is traversable at the beginning of learning. Finally, we employ the online learning process so that our framework can be adaptive to varied scenes. Experimental results on three different style of image sequences, i.e. shadow road, rain sequence and variational sequence, demonstrate the adaptability, stability and parameter insensitivity of our method to the varied scenes and complex illumination. Yong Liu 0007, Yiyi Liao, Yue Wang 0020 |
ICRA | 3 |