Zehao Yu 0002

dblp:203/9233-2 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
9since 2021 · last 2024
0000-0002-6559-9830ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
10 papers
3D vision · 55% Robot manipulation · 13% Deep learning architectures and training · 10%
Computer graphics and multimedia
4 papers
Geometric modeling and processing · 57% Rendering · 43%

Topics — the 29 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Rendering
novel view synthesis
1.022024
Mip-Splatting: Alias-Free 3D Gaussian Splatting · CVPR 2024
Gaussian Opacity Fields: Efficient Adaptive Surface Reconstruction in Unbounded Scenes · ACM Trans. Graph. 2024
Computer vision › 3D vision
depth estimation
0.922020
P2Net: Patch-Match and Plane-Regularization for Unsupervised Indoor Depth Estimation · ECCV (24) 2020
Fast-MVSNet: Sparse-to-Dense Multi-View Stereo With Learned Propagation and Gauss-Newton Refinement · CVPR 2020
Robotics › Robot manipulation › grasping › grasp detection
6-dof grasp detection
0.812024
Efficient End-to-End Detection of 6-DoF Grasps for Robotic Bin Picking · ICRA 2024
Computer vision › 3D vision › 3d reconstruction › surface reconstruction
gaussian splatting surface reconstruction
0.812024
Gaussian Opacity Fields: Efficient Adaptive Surface Reconstruction in Unbounded Scenes · ACM Trans. Graph. 2024
Robotics › Robot manipulation
grasping
0.812024
Efficient End-to-End Detection of 6-DoF Grasps for Robotic Bin Picking · ICRA 2024
Computer vision › 3D vision › 3d reconstruction
surface reconstruction
0.812024
Gaussian Opacity Fields: Efficient Adaptive Surface Reconstruction in Unbounded Scenes · ACM Trans. Graph. 2024
Rendering › gaussian splatting
3d gaussian splatting
0.812024
Mip-Splatting: Alias-Free 3D Gaussian Splatting · CVPR 2024
Rendering
antialiasing
0.812024
Mip-Splatting: Alias-Free 3D Gaussian Splatting · CVPR 2024
Geometric modeling and processing › 3d reconstruction
indoor scene reconstruction
0.812024
DebSDF: Delving Into the Details and Bias of Neural Indoor Scene Reconstruction · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Geometric modeling and processing › surface reconstruction › implicit surface reconstruction
neural implicit surface reconstruction
0.812024
DebSDF: Delving Into the Details and Bias of Neural Indoor Scene Reconstruction · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Geometric modeling and processing › shape representation › implicit representation
signed distance function
0.812024
DebSDF: Delving Into the Details and Bias of Neural Indoor Scene Reconstruction · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Robotics › Autonomous driving
end-to-end driving
0.712023
TransFuser: Imitation With Transformer-Based Sensor Fusion for Autonomous Driving · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Machine learning › Reinforcement learning
imitation learning
0.712023
TransFuser: Imitation With Transformer-Based Sensor Fusion for Autonomous Driving · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Computer vision › 3D vision › depth estimation
monocular depth estimation
0.712023
PlaneDepth: Self-Supervised Depth Estimation via Orthogonal Planes · CVPR 2023
Computer vision › 3D vision › depth estimation
self-supervised depth estimation
0.712023
PlaneDepth: Self-Supervised Depth Estimation via Orthogonal Planes · CVPR 2023
Robotics › Robot navigation and mapping
sensor fusion
0.712023
TransFuser: Imitation With Transformer-Based Sensor Fusion for Autonomous Driving · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Computer vision › 3D vision › 3d shape representation
implicit surface representation
0.612022
MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface Reconstruction · NeurIPS 2022
Machine learning › Deep learning architectures and training › feedforward neural network
multilayer perceptron
0.612022
AS-MLP: An Axial Shifted MLP Architecture for Vision · ICLR 2022
Computer vision › 3D vision › 3d reconstruction › surface reconstruction › neural surface reconstruction
neural implicit surface reconstruction
0.612022
MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface Reconstruction · NeurIPS 2022
Machine learning › Deep learning architectures and training › feedforward neural network › multilayer perceptron
vision MLP
0.612022
AS-MLP: An Axial Shifted MLP Architecture for Vision · ICLR 2022
Computer vision › 3D vision › 3d reconstruction
multi-view stereo
0.412020
Fast-MVSNet: Sparse-to-Dense Multi-View Stereo With Learned Propagation and Gauss-Newton Refinement · CVPR 2020
Computer vision › 3D vision
3d reconstruction
0.412019
Single-Image Piece-Wise Planar 3D Reconstruction via Associative Embedding · CVPR 2019
Computer vision › Segmentation and scene understanding
instance segmentation
0.412019
Single-Image Piece-Wise Planar 3D Reconstruction via Associative Embedding · CVPR 2019
Computer vision › 3D vision › 3d reconstruction › surface reconstruction
piecewise planar reconstruction
0.412019
Single-Image Piece-Wise Planar 3D Reconstruction via Associative Embedding · CVPR 2019
Computer vision › 3D vision › depth image analysis
depth map processing
0.212024
Efficient End-to-End Detection of 6-DoF Grasps for Robotic Bin Picking · ICRA 2024
Machine learning › Trustworthy machine learning
uncertainty modeling
0.212024
DebSDF: Delving Into the Details and Bias of Neural Indoor Scene Reconstruction · IEEE Trans. Pattern Anal. Mach. Intell. 2024
Geometric modeling and processing
surface reconstruction
0.212024
3D Neural Edge Reconstruction · CVPR 2024
Computer vision › 3D vision › 3d reconstruction
multi-view reconstruction
0.212022
MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface Reconstruction · NeurIPS 2022
Computer vision › 3D vision › geometric estimation › geometric model fitting
plane fitting
0.112019
Single-Image Piece-Wise Planar 3D Reconstruction via Associative Embedding · CVPR 2019

Methods — techniques the papers use, named apart from their topics

volume rendering · 3.0smoothness regularization · 1.5marching tetrahedra · 1.5importance-guided ray sampling · 1.53d gaussian splatting · 1.5convolutional neural network · 0.8unsigned distance function · 0.8probabilistic grasp modeling · 0.8power-spherical distribution · 0.8multi-view edge maps · 0.8mip filter · 0.83d smoothing filter · 0.8laplacian mixture model · 0.7bilateral occlusion mask · 0.7
YearPublicationVenuePosition
2024 3D Neural Edge Reconstruction
abstract
Real-world objects and environments are predominantly composed of edge features, including straight lines and curves. Such edges are crucial elements for various applications, such as CAD modeling, surface meshing, lane mapping, etc. However, existing traditional methods only prioritize lines over curves for simplicity in geometric modeling. To this end, we introduce EMAP, a new method for learning 3D edge representations with a focus on both lines and curves. Our method implicitly encodes 3D edge distance and direction in Unsigned Distance Functions (UDF) from multi-view edge maps. On top of this neural representation, we propose an edge extraction algorithm that robustly abstracts parametric 3D edges from the inferred edge points and their directions. Comprehensive evaluations demonstrate that our method achieves better 3D edge reconstruction on multiple challenging datasets. We further show that our learned UDF field enhances neural surface reconstruction by capturing more details.
Songyou Peng, Zehao Yu 0002, Shaohui Liu, Rémi Pautrat, Xiaochuan Yin, Marc Pollefeys
CVPR3
2024 Mip-Splatting: Alias-Free 3D Gaussian Splatting
abstract
Recently, 3D Gaussian Splatting has demonstrated impressive novel view synthesis results, reaching high fidelity and efficiency. However, strong artifacts can be observed when changing the sampling rate, e.g., by changing focal length or camera distance. We find that the source for this phenomenon can be attributed to the lack of 3D frequency constraints and the usage of a 2D dilation filter. To address this problem, we introduce a 3D smoothing filter to constrains the size of the 3D Gaussian primitives based on the maximal sampling frequency induced by the input views. It eliminates high-frequency artifacts when zooming in. Moreover, replacing 2D dilation with a 2D Mip filter, which simulates a 2D box filter, effectively mitigates aliasing and dilation issues. Our evaluation, including scenarios such a training on single-scale images and testing on multiple scales, validates the effectiveness of our approach.
Zehao Yu 0002, Anpei Chen, Binbin Huang 0004, Torsten Sattler, Andreas Geiger 0001
CVPR1
2024 Efficient End-to-End Detection of 6-DoF Grasps for Robotic Bin Picking
abstract
Bin picking is an important building block for many robotic systems, in logistics, production or in household use-cases. In recent years, machine learning methods for the prediction of 6-DoF grasps on diverse and unknown objects have shown promising progress. However, existing approaches only consider a single ground truth grasp orientation at a grasp location during training and therefore can only predict limited grasp orientations which leads to a reduced number of feasible grasps in bin picking with restricted reachability. In this paper, we propose a novel approach for learning dense and diverse 6-DoF grasps for parallel-jaw grippers in robotic bin picking. We introduce a parameterized grasp distribution model based on Power-Spherical distributions that enables a training based on all possible ground truth samples. Thereby, we also consider the grasp uncertainty enhancing the model’s robustness to noisy inputs. As a result, given a single top-down view depth image, our model can generate diverse grasps with multiple collision-free grasp orientations. Experimental evaluations in simulation and on a real robotic bin picking setup demonstrate the model’s ability to generalize across various object categories achieving an object clearing rate of around 90% in simulation and real-world experiments. We also outperform state of the art approaches. Moreover, the proposed approach exhibits its usability in real robot experiments without any refinement steps, even when only trained on a synthetic dataset, due to the probabilistic grasp distribution modeling.
Alexander Qualmann, Zehao Yu 0002, Miroslav Gabriel, Philipp Schillinger, Markus Spies, Ngo Anh Vien, Andreas Geiger 0001
ICRA3
2024 DebSDF: Delving Into the Details and Bias of Neural Indoor Scene Reconstruction
abstract
In recent years, the neural implicit surface has emerged as a powerful representation for multi-view surface reconstruction due to its simplicity and State-of-the-Art performance. However, reconstructing smooth and detailed surfaces in indoor scenes from multi-view images presents unique challenges. Indoor scenes typically contain large texture-less regions, making the photometric loss unreliable for optimizing the implicit surface. Previous work utilizes monocular geometry priors to improve the reconstruction in indoor scenes. However, monocular priors often contain substantial errors in thin structure regions due to domain gaps and the inherent inconsistencies when derived independently from different views. This paper presents DebSDF to address these challenges, focusing on the utilization of uncertainty in monocular priors and the bias in SDF-based volume rendering. We propose an uncertainty modeling technique that associates larger uncertainties with larger errors in the monocular priors. High-uncertainty priors are then excluded from optimization to prevent bias. This uncertainty measure also informs an importance-guided ray sampling and adaptive smoothness regularization, enhancing the learning of fine structures. We further introduce a bias-aware signed distance function to density transformation that takes into account the curvature and the angle between the view direction and the SDF normals to reconstruct fine details better. Our approach has been validated through extensive experiments on several challenging datasets, demonstrating improved qualitative and quantitative results in reconstructing thin structures in indoor scenes, thereby outperforming previous work.
Zehao Yu 0002, Shenghua Gao
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Gaussian Opacity Fields: Efficient Adaptive Surface Reconstruction in Unbounded Scenes
abstract
Recently, 3D Gaussian Splatting (3DGS) has demonstrated impressive novel view synthesis results, while allowing the rendering of high-resolution images in real-time. However, leveraging 3D Gaussians for surface reconstruction poses significant challenges due to the explicit and disconnected nature of 3D Gaussians. In this work, we present Gaussian Opacity Fields (GOF), a novel approach for efficient, high-quality, and adaptive surface reconstruction in unbounded scenes. Our GOF is derived from ray-tracing-based volume rendering of 3D Gaussians, enabling direct geometry extraction from 3D Gaussians by identifying its levelset, without resorting to Poisson reconstruction or TSDF fusion as in previous work. We approximate the surface normal of Gaussians as the normal of the ray-Gaussian intersection plane, enabling the application of regularization that significantly enhances geometry. Furthermore, we develop an efficient geometry extraction method utilizing Marching Tetrahedra, where the tetrahedral grids are induced from 3D Gaussians and thus adapt to the scene's complexity. Our evaluations reveal that GOF surpasses existing 3DGS-based methods in surface reconstruction and novel view synthesis. Further, it compares favorably to or even outperforms, neural implicit methods in both quality and speed.
Zehao Yu 0002, Torsten Sattler, Andreas Geiger 0001
ACM Trans. Graph.1
2023 PlaneDepth: Self-Supervised Depth Estimation via Orthogonal Planes
abstract
Multiple near frontal-parallel planes based depth representation demonstrated impressive results in self-supervised monocular depth estimation (MDE). Whereas, such a representation would cause the discontinuity of the ground as it is perpendicular to the frontal-parallel planes, which is detrimental to the identification of drivable space in autonomous driving. In this paper, we propose the PlaneDepth, a novel orthogonal planes based presentation, including vertical planes and ground planes. PlaneDepth estimates the depth distribution using a Laplacian Mixture Model based on orthogonal planes for an input image. These planes are used to synthesize a reference view to provide the self-supervision signal. Further, we find that the widely used resizing and cropping data augmentation breaks the orthogonality assumptions, leading to inferior plane predictions. We address this problem by explicitly constructing the resizing cropping transformation to rectify the predefined planes and predicted camera pose. Moreover, we propose an augmented self-distillation loss supervised with a bilateral occlusion mask to boost the robustness of orthogonal planes representation for occlusions. Thanks to our orthogonal planes representation, we can extract the ground plane in an unsupervised manner, which is important for autonomous driving. Extensive experiments on the KITTI dataset demonstrate the effectiveness and efficiency of our method. The code is available at https://github.com/svip-lab/PlaneDepth.
Ruoyu Wang 0014, Zehao Yu 0002, Shenghua Gao
CVPR2
2023 TransFuser: Imitation With Transformer-Based Sensor Fusion for Autonomous Driving
abstract
How should we integrate representations from complementary sensors for autonomous driving? Geometry-based fusion has shown promise for perception (e.g., object detection, motion forecasting). However, in the context of end-to-end driving, we find that imitation learning based on existing sensor fusion methods underperforms in complex driving scenarios with a high density of dynamic agents. Therefore, we propose TransFuser, a mechanism to integrate image and LiDAR representations using self-attention. Our approach uses transformer modules at multiple resolutions to fuse perspective view and bird's eye view feature maps. We experimentally validate its efficacy on a challenging new benchmark with long routes and dense traffic, as well as the official leaderboard of the CARLA urban driving simulator. At the time of submission, TransFuser outperforms all prior work on the CARLA leaderboard in terms of driving score by a large margin. Compared to geometry-based fusion, TransFuser reduces the average collisions per kilometer by 48%.
Kashyap Chitta, Aditya Prakash 0001, Bernhard Jaeger, Zehao Yu 0002, Katrin Renz, Andreas Geiger 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 AS-MLP: An Axial Shifted MLP Architecture for Vision
Dongze Lian, Zehao Yu 0002, Shenghua Gao
ICLR2
2022 MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface Reconstruction
abstract
In recent years, neural implicit surface reconstruction methods have become popular for multi-view 3D reconstruction. In contrast to traditional multi-view stereo methods, these approaches tend to produce smoother and more complete reconstructions due to the inductive smoothness bias of neural networks. State-of-the-art neural implicit methods allow for high-quality reconstructions of simple scenes from many input views. Yet, their performance drops significantly for larger and more complex scenes and scenes captured from sparse viewpoints. This is caused primarily by the inherent ambiguity in the RGB reconstruction loss that does not provide enough constraints, in particular in less-observed and textureless areas. Motivated by recent advances in the area of monocular geometry prediction, we systematically explore the utility these cues provide for improving neural implicit surface reconstruction. We demonstrate that depth and normal cues, predicted by general-purpose monocular estimators, significantly improve reconstruction quality and optimization time. Further, we analyse and investigate multiple design choices for representing neural implicit surfaces, ranging from monolithic MLP models over single-grid to multi-resolution grid representations. We observe that geometric monocular priors improve performance both for small-scale single-object as well as large-scale multi-object scenes, independent of the choice of representation.
Zehao Yu 0002, Songyou Peng, Michael Niemeyer, Torsten Sattler, Andreas Geiger 0001
NeurIPS1
2020 Fast-MVSNet: Sparse-to-Dense Multi-View Stereo With Learned Propagation and Gauss-Newton Refinement
abstract
Almost all previous deep learning-based multi-view stereo (MVS) approaches focus on improving reconstruction quality. Besides quality, efficiency is also a desirable feature for MVS in real scenarios. Towards this end, this paper presents a Fast-MVSNet, a novel sparse-to-dense coarse-to-fine framework, for fast and accurate depth estimation in MVS. Specifically, in our Fast-MVSNet, we first construct a sparse cost volume for learning a sparse and high-resolution depth map. Then we leverage a small-scale convolutional neural network to encode the depth dependencies for pixels within a local region to densify the sparse high-resolution depth map. At last, a simple but efficient Gauss-Newton layer is proposed to further optimize the depth map. On one hand, the high-resolution depth map, the data-adaptive propagation method and the Gauss-Newton layer jointly guarantee the effectiveness of our method. On the other hand, all modules in our Fast-MVSNet are lightweight and thus guarantee the efficiency of our approach. Besides, our approach is also memory-friendly because of the sparse depth representation. Extensive experimental results show that our method is 5 times and 14 times faster than Point-MVSNet and R-MVSNet, respectively, while achieving comparable or even better results on the challenging Tanks and Temples dataset as well as the DTU dataset. Code is available at https://github.com/svip-lab/FastMVSNet.
Zehao Yu 0002, Shenghua Gao
CVPR1
2020 P2Net: Patch-Match and Plane-Regularization for Unsupervised Indoor Depth Estimation
Zehao Yu 0002, Shenghua Gao
ECCV (24)1
2019 Single-Image Piece-Wise Planar 3D Reconstruction via Associative Embedding
abstract
Single-image piece-wise planar 3D reconstruction aims to simultaneously segment plane instances and recover 3D plane parameters from an image. Most recent approaches leverage convolutional neural networks (CNNs) and achieve promising results. However, these methods are limited to detecting a fixed number of planes with certain learned order. To tackle this problem, we propose a novel two-stage method based on associative embedding, inspired by its recent success in instance segmentation. In the first stage, we train a CNN to map each pixel to an embedding space where pixels from the same plane instance have similar embeddings. Then, the plane instances are obtained by grouping the embedding vectors in planar regions via an efficient mean shift clustering algorithm. In the second stage, we estimate the parameter for each plane instance by considering both pixel-level and instance-level consistencies. With the proposed method, we are able to detect an arbitrary number of planes. Extensive experiments on public datasets validate the effectiveness and efficiency of our method. Furthermore, our method runs at 30 fps at the testing time, thus could facilitate many real-time applications such as visual SLAM and human-robot interaction. Code is available at https://github.com/svip-lab/PlanarReconstruction.
Zehao Yu 0002, Jia Zheng 0002, Dongze Lian, Zihan Zhou 0001, Shenghua Gao
CVPR1
2018 Believe It or Not, We Know What You Are Looking At!
Dongze Lian, Zehao Yu 0002, Shenghua Gao
ACCV (3)2