EDBT 2026 Demo / reviewers in the wild / expert
Zehao Yu 0002
dblp:203/9233-2
· DBLP profile ↗
13ranked-venue papers
6as first author
9since 2021 · last 2024
0000-0002-6559-9830ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
10 papers |
3D vision · 55% Robot manipulation · 13% Deep learning architectures and training · 10% | |
| Computer graphics and multimedia
4 papers |
Geometric modeling and processing · 57% Rendering · 43% |
Topics — the 29 heaviest of 31, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Rendering
novel view synthesis |
1.0 | 2 | 2024 | Mip-Splatting: Alias-Free 3D Gaussian Splatting · CVPR 2024 Gaussian Opacity Fields: Efficient Adaptive Surface Reconstruction in Unbounded Scenes · ACM Trans. Graph. 2024 |
Computer vision › 3D vision
depth estimation |
0.9 | 2 | 2020 | P2Net: Patch-Match and Plane-Regularization for Unsupervised Indoor Depth Estimation · ECCV (24) 2020 Fast-MVSNet: Sparse-to-Dense Multi-View Stereo With Learned Propagation and Gauss-Newton Refinement · CVPR 2020 |
Robotics › Robot manipulation › grasping › grasp detection
6-dof grasp detection |
0.8 | 1 | 2024 | Efficient End-to-End Detection of 6-DoF Grasps for Robotic Bin Picking · ICRA 2024 |
Computer vision › 3D vision › 3d reconstruction › surface reconstruction
gaussian splatting surface reconstruction |
0.8 | 1 | 2024 | Gaussian Opacity Fields: Efficient Adaptive Surface Reconstruction in Unbounded Scenes · ACM Trans. Graph. 2024 |
Robotics › Robot manipulation
grasping |
0.8 | 1 | 2024 | Efficient End-to-End Detection of 6-DoF Grasps for Robotic Bin Picking · ICRA 2024 |
Computer vision › 3D vision › 3d reconstruction
surface reconstruction |
0.8 | 1 | 2024 | Gaussian Opacity Fields: Efficient Adaptive Surface Reconstruction in Unbounded Scenes · ACM Trans. Graph. 2024 |
Rendering › gaussian splatting
3d gaussian splatting |
0.8 | 1 | 2024 | Mip-Splatting: Alias-Free 3D Gaussian Splatting · CVPR 2024 |
Rendering
antialiasing |
0.8 | 1 | 2024 | Mip-Splatting: Alias-Free 3D Gaussian Splatting · CVPR 2024 |
Geometric modeling and processing › 3d reconstruction
indoor scene reconstruction |
0.8 | 1 | 2024 | DebSDF: Delving Into the Details and Bias of Neural Indoor Scene Reconstruction · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Geometric modeling and processing › surface reconstruction › implicit surface reconstruction
neural implicit surface reconstruction |
0.8 | 1 | 2024 | DebSDF: Delving Into the Details and Bias of Neural Indoor Scene Reconstruction · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Geometric modeling and processing › shape representation › implicit representation
signed distance function |
0.8 | 1 | 2024 | DebSDF: Delving Into the Details and Bias of Neural Indoor Scene Reconstruction · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Robotics › Autonomous driving
end-to-end driving |
0.7 | 1 | 2023 | TransFuser: Imitation With Transformer-Based Sensor Fusion for Autonomous Driving · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Machine learning › Reinforcement learning
imitation learning |
0.7 | 1 | 2023 | TransFuser: Imitation With Transformer-Based Sensor Fusion for Autonomous Driving · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Computer vision › 3D vision › depth estimation
monocular depth estimation |
0.7 | 1 | 2023 | PlaneDepth: Self-Supervised Depth Estimation via Orthogonal Planes · CVPR 2023 |
Computer vision › 3D vision › depth estimation
self-supervised depth estimation |
0.7 | 1 | 2023 | PlaneDepth: Self-Supervised Depth Estimation via Orthogonal Planes · CVPR 2023 |
Robotics › Robot navigation and mapping
sensor fusion |
0.7 | 1 | 2023 | TransFuser: Imitation With Transformer-Based Sensor Fusion for Autonomous Driving · IEEE Trans. Pattern Anal. Mach. Intell. 2023 |
Computer vision › 3D vision › 3d shape representation
implicit surface representation |
0.6 | 1 | 2022 | MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface Reconstruction · NeurIPS 2022 |
Machine learning › Deep learning architectures and training › feedforward neural network
multilayer perceptron |
0.6 | 1 | 2022 | AS-MLP: An Axial Shifted MLP Architecture for Vision · ICLR 2022 |
Computer vision › 3D vision › 3d reconstruction › surface reconstruction › neural surface reconstruction
neural implicit surface reconstruction |
0.6 | 1 | 2022 | MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface Reconstruction · NeurIPS 2022 |
Machine learning › Deep learning architectures and training › feedforward neural network › multilayer perceptron
vision MLP |
0.6 | 1 | 2022 | AS-MLP: An Axial Shifted MLP Architecture for Vision · ICLR 2022 |
Computer vision › 3D vision › 3d reconstruction
multi-view stereo |
0.4 | 1 | 2020 | Fast-MVSNet: Sparse-to-Dense Multi-View Stereo With Learned Propagation and Gauss-Newton Refinement · CVPR 2020 |
Computer vision › 3D vision
3d reconstruction |
0.4 | 1 | 2019 | Single-Image Piece-Wise Planar 3D Reconstruction via Associative Embedding · CVPR 2019 |
Computer vision › Segmentation and scene understanding
instance segmentation |
0.4 | 1 | 2019 | Single-Image Piece-Wise Planar 3D Reconstruction via Associative Embedding · CVPR 2019 |
Computer vision › 3D vision › 3d reconstruction › surface reconstruction
piecewise planar reconstruction |
0.4 | 1 | 2019 | Single-Image Piece-Wise Planar 3D Reconstruction via Associative Embedding · CVPR 2019 |
Computer vision › 3D vision › depth image analysis
depth map processing |
0.2 | 1 | 2024 | Efficient End-to-End Detection of 6-DoF Grasps for Robotic Bin Picking · ICRA 2024 |
Machine learning › Trustworthy machine learning
uncertainty modeling |
0.2 | 1 | 2024 | DebSDF: Delving Into the Details and Bias of Neural Indoor Scene Reconstruction · IEEE Trans. Pattern Anal. Mach. Intell. 2024 |
Geometric modeling and processing
surface reconstruction |
0.2 | 1 | 2024 | 3D Neural Edge Reconstruction · CVPR 2024 |
Computer vision › 3D vision › 3d reconstruction
multi-view reconstruction |
0.2 | 1 | 2022 | MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface Reconstruction · NeurIPS 2022 |
Computer vision › 3D vision › geometric estimation › geometric model fitting
plane fitting |
0.1 | 1 | 2019 | Single-Image Piece-Wise Planar 3D Reconstruction via Associative Embedding · CVPR 2019 |
Methods — techniques the papers use, named apart from their topics
volume rendering · 3.0smoothness regularization · 1.5marching tetrahedra · 1.5importance-guided ray sampling · 1.53d gaussian splatting · 1.5convolutional neural network · 0.8unsigned distance function · 0.8probabilistic grasp modeling · 0.8power-spherical distribution · 0.8multi-view edge maps · 0.8mip filter · 0.83d smoothing filter · 0.8laplacian mixture model · 0.7bilateral occlusion mask · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | 3D Neural Edge ReconstructionabstractReal-world objects and environments are predominantly composed of edge features, including straight lines and curves. Such edges are crucial elements for various applications, such as CAD modeling, surface meshing, lane mapping, etc. However, existing traditional methods only prioritize lines over curves for simplicity in geometric modeling. To this end, we introduce EMAP, a new method for learning 3D edge representations with a focus on both lines and curves. Our method implicitly encodes 3D edge distance and direction in Unsigned Distance Functions (UDF) from multi-view edge maps. On top of this neural representation, we propose an edge extraction algorithm that robustly abstracts parametric 3D edges from the inferred edge points and their directions. Comprehensive evaluations demonstrate that our method achieves better 3D edge reconstruction on multiple challenging datasets. We further show that our learned UDF field enhances neural surface reconstruction by capturing more details. Songyou Peng, Zehao Yu 0002, Shaohui Liu, Rémi Pautrat, Xiaochuan Yin, Marc Pollefeys |
CVPR | 3 |
| 2024 | Mip-Splatting: Alias-Free 3D Gaussian SplattingabstractRecently, 3D Gaussian Splatting has demonstrated impressive novel view synthesis results, reaching high fidelity and efficiency. However, strong artifacts can be observed when changing the sampling rate, e.g., by changing focal length or camera distance. We find that the source for this phenomenon can be attributed to the lack of 3D frequency constraints and the usage of a 2D dilation filter. To address this problem, we introduce a 3D smoothing filter to constrains the size of the 3D Gaussian primitives based on the maximal sampling frequency induced by the input views. It eliminates high-frequency artifacts when zooming in. Moreover, replacing 2D dilation with a 2D Mip filter, which simulates a 2D box filter, effectively mitigates aliasing and dilation issues. Our evaluation, including scenarios such a training on single-scale images and testing on multiple scales, validates the effectiveness of our approach. Zehao Yu 0002, Anpei Chen, Binbin Huang 0004, Torsten Sattler, Andreas Geiger 0001 |
CVPR | 1 |
| 2024 | Efficient End-to-End Detection of 6-DoF Grasps for Robotic Bin PickingabstractBin picking is an important building block for many robotic systems, in logistics, production or in household use-cases. In recent years, machine learning methods for the prediction of 6-DoF grasps on diverse and unknown objects have shown promising progress. However, existing approaches only consider a single ground truth grasp orientation at a grasp location during training and therefore can only predict limited grasp orientations which leads to a reduced number of feasible grasps in bin picking with restricted reachability. In this paper, we propose a novel approach for learning dense and diverse 6-DoF grasps for parallel-jaw grippers in robotic bin picking. We introduce a parameterized grasp distribution model based on Power-Spherical distributions that enables a training based on all possible ground truth samples. Thereby, we also consider the grasp uncertainty enhancing the model’s robustness to noisy inputs. As a result, given a single top-down view depth image, our model can generate diverse grasps with multiple collision-free grasp orientations. Experimental evaluations in simulation and on a real robotic bin picking setup demonstrate the model’s ability to generalize across various object categories achieving an object clearing rate of around 90% in simulation and real-world experiments. We also outperform state of the art approaches. Moreover, the proposed approach exhibits its usability in real robot experiments without any refinement steps, even when only trained on a synthetic dataset, due to the probabilistic grasp distribution modeling. Alexander Qualmann, Zehao Yu 0002, Miroslav Gabriel, Philipp Schillinger, Markus Spies, Ngo Anh Vien, Andreas Geiger 0001 |
ICRA | 3 |
| 2024 | DebSDF: Delving Into the Details and Bias of Neural Indoor Scene ReconstructionabstractIn recent years, the neural implicit surface has emerged as a powerful representation for multi-view surface reconstruction due to its simplicity and State-of-the-Art performance. However, reconstructing smooth and detailed surfaces in indoor scenes from multi-view images presents unique challenges. Indoor scenes typically contain large texture-less regions, making the photometric loss unreliable for optimizing the implicit surface. Previous work utilizes monocular geometry priors to improve the reconstruction in indoor scenes. However, monocular priors often contain substantial errors in thin structure regions due to domain gaps and the inherent inconsistencies when derived independently from different views. This paper presents DebSDF to address these challenges, focusing on the utilization of uncertainty in monocular priors and the bias in SDF-based volume rendering. We propose an uncertainty modeling technique that associates larger uncertainties with larger errors in the monocular priors. High-uncertainty priors are then excluded from optimization to prevent bias. This uncertainty measure also informs an importance-guided ray sampling and adaptive smoothness regularization, enhancing the learning of fine structures. We further introduce a bias-aware signed distance function to density transformation that takes into account the curvature and the angle between the view direction and the SDF normals to reconstruct fine details better. Our approach has been validated through extensive experiments on several challenging datasets, demonstrating improved qualitative and quantitative results in reconstructing thin structures in indoor scenes, thereby outperforming previous work. Zehao Yu 0002, Shenghua Gao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Gaussian Opacity Fields: Efficient Adaptive Surface Reconstruction in Unbounded ScenesabstractRecently, 3D Gaussian Splatting (3DGS) has demonstrated impressive novel view synthesis results, while allowing the rendering of high-resolution images in real-time. However, leveraging 3D Gaussians for surface reconstruction poses significant challenges due to the explicit and disconnected nature of 3D Gaussians. In this work, we present Gaussian Opacity Fields (GOF), a novel approach for efficient, high-quality, and adaptive surface reconstruction in unbounded scenes. Our GOF is derived from ray-tracing-based volume rendering of 3D Gaussians, enabling direct geometry extraction from 3D Gaussians by identifying its levelset, without resorting to Poisson reconstruction or TSDF fusion as in previous work. We approximate the surface normal of Gaussians as the normal of the ray-Gaussian intersection plane, enabling the application of regularization that significantly enhances geometry. Furthermore, we develop an efficient geometry extraction method utilizing Marching Tetrahedra, where the tetrahedral grids are induced from 3D Gaussians and thus adapt to the scene's complexity. Our evaluations reveal that GOF surpasses existing 3DGS-based methods in surface reconstruction and novel view synthesis. Further, it compares favorably to or even outperforms, neural implicit methods in both quality and speed. Zehao Yu 0002, Torsten Sattler, Andreas Geiger 0001 |
ACM Trans. Graph. | 1 |
| 2023 | PlaneDepth: Self-Supervised Depth Estimation via Orthogonal PlanesabstractMultiple near frontal-parallel planes based depth representation demonstrated impressive results in self-supervised monocular depth estimation (MDE). Whereas, such a representation would cause the discontinuity of the ground as it is perpendicular to the frontal-parallel planes, which is detrimental to the identification of drivable space in autonomous driving. In this paper, we propose the PlaneDepth, a novel orthogonal planes based presentation, including vertical planes and ground planes. PlaneDepth estimates the depth distribution using a Laplacian Mixture Model based on orthogonal planes for an input image. These planes are used to synthesize a reference view to provide the self-supervision signal. Further, we find that the widely used resizing and cropping data augmentation breaks the orthogonality assumptions, leading to inferior plane predictions. We address this problem by explicitly constructing the resizing cropping transformation to rectify the predefined planes and predicted camera pose. Moreover, we propose an augmented self-distillation loss supervised with a bilateral occlusion mask to boost the robustness of orthogonal planes representation for occlusions. Thanks to our orthogonal planes representation, we can extract the ground plane in an unsupervised manner, which is important for autonomous driving. Extensive experiments on the KITTI dataset demonstrate the effectiveness and efficiency of our method. The code is available at https://github.com/svip-lab/PlaneDepth. Ruoyu Wang 0014, Zehao Yu 0002, Shenghua Gao |
CVPR | 2 |
| 2023 | TransFuser: Imitation With Transformer-Based Sensor Fusion for Autonomous DrivingabstractHow should we integrate representations from complementary sensors for autonomous driving? Geometry-based fusion has shown promise for perception (e.g., object detection, motion forecasting). However, in the context of end-to-end driving, we find that imitation learning based on existing sensor fusion methods underperforms in complex driving scenarios with a high density of dynamic agents. Therefore, we propose TransFuser, a mechanism to integrate image and LiDAR representations using self-attention. Our approach uses transformer modules at multiple resolutions to fuse perspective view and bird's eye view feature maps. We experimentally validate its efficacy on a challenging new benchmark with long routes and dense traffic, as well as the official leaderboard of the CARLA urban driving simulator. At the time of submission, TransFuser outperforms all prior work on the CARLA leaderboard in terms of driving score by a large margin. Compared to geometry-based fusion, TransFuser reduces the average collisions per kilometer by 48%. Kashyap Chitta, Aditya Prakash 0001, Bernhard Jaeger, Zehao Yu 0002, Katrin Renz, Andreas Geiger 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | AS-MLP: An Axial Shifted MLP Architecture for Vision
Dongze Lian, Zehao Yu 0002, Shenghua Gao |
ICLR | 2 |
| 2022 | MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface ReconstructionabstractIn recent years, neural implicit surface reconstruction methods have become popular for multi-view 3D reconstruction. In contrast to traditional multi-view stereo methods, these approaches tend to produce smoother and more complete reconstructions due to the inductive smoothness bias of neural networks. State-of-the-art neural implicit methods allow for high-quality reconstructions of simple scenes from many input views. Yet, their performance drops significantly for larger and more complex scenes and scenes captured from sparse viewpoints. This is caused primarily by the inherent ambiguity in the RGB reconstruction loss that does not provide enough constraints, in particular in less-observed and textureless areas. Motivated by recent advances in the area of monocular geometry prediction, we systematically explore the utility these cues provide for improving neural implicit surface reconstruction. We demonstrate that depth and normal cues, predicted by general-purpose monocular estimators, significantly improve reconstruction quality and optimization time. Further, we analyse and investigate multiple design choices for representing neural implicit surfaces, ranging from monolithic MLP models over single-grid to multi-resolution grid representations. We observe that geometric monocular priors improve performance both for small-scale single-object as well as large-scale multi-object scenes, independent of the choice of representation. Zehao Yu 0002, Songyou Peng, Michael Niemeyer, Torsten Sattler, Andreas Geiger 0001 |
NeurIPS | 1 |
| 2020 | Fast-MVSNet: Sparse-to-Dense Multi-View Stereo With Learned Propagation and Gauss-Newton RefinementabstractAlmost all previous deep learning-based multi-view stereo (MVS) approaches focus on improving reconstruction quality. Besides quality, efficiency is also a desirable feature for MVS in real scenarios. Towards this end, this paper presents a Fast-MVSNet, a novel sparse-to-dense coarse-to-fine framework, for fast and accurate depth estimation in MVS. Specifically, in our Fast-MVSNet, we first construct a sparse cost volume for learning a sparse and high-resolution depth map. Then we leverage a small-scale convolutional neural network to encode the depth dependencies for pixels within a local region to densify the sparse high-resolution depth map. At last, a simple but efficient Gauss-Newton layer is proposed to further optimize the depth map. On one hand, the high-resolution depth map, the data-adaptive propagation method and the Gauss-Newton layer jointly guarantee the effectiveness of our method. On the other hand, all modules in our Fast-MVSNet are lightweight and thus guarantee the efficiency of our approach. Besides, our approach is also memory-friendly because of the sparse depth representation. Extensive experimental results show that our method is 5 times and 14 times faster than Point-MVSNet and R-MVSNet, respectively, while achieving comparable or even better results on the challenging Tanks and Temples dataset as well as the DTU dataset. Code is available at https://github.com/svip-lab/FastMVSNet. Zehao Yu 0002, Shenghua Gao |
CVPR | 1 |
| 2020 | P2Net: Patch-Match and Plane-Regularization for Unsupervised Indoor Depth Estimation
Zehao Yu 0002, Shenghua Gao |
ECCV (24) | 1 |
| 2019 | Single-Image Piece-Wise Planar 3D Reconstruction via Associative EmbeddingabstractSingle-image piece-wise planar 3D reconstruction aims to simultaneously segment plane instances and recover 3D plane parameters from an image. Most recent approaches leverage convolutional neural networks (CNNs) and achieve promising results. However, these methods are limited to detecting a fixed number of planes with certain learned order. To tackle this problem, we propose a novel two-stage method based on associative embedding, inspired by its recent success in instance segmentation. In the first stage, we train a CNN to map each pixel to an embedding space where pixels from the same plane instance have similar embeddings. Then, the plane instances are obtained by grouping the embedding vectors in planar regions via an efficient mean shift clustering algorithm. In the second stage, we estimate the parameter for each plane instance by considering both pixel-level and instance-level consistencies. With the proposed method, we are able to detect an arbitrary number of planes. Extensive experiments on public datasets validate the effectiveness and efficiency of our method. Furthermore, our method runs at 30 fps at the testing time, thus could facilitate many real-time applications such as visual SLAM and human-robot interaction. Code is available at https://github.com/svip-lab/PlanarReconstruction. Zehao Yu 0002, Jia Zheng 0002, Dongze Lian, Zihan Zhou 0001, Shenghua Gao |
CVPR | 1 |
| 2018 | Believe It or Not, We Know What You Are Looking At!
Dongze Lian, Zehao Yu 0002, Shenghua Gao |
ACCV (3) | 2 |