EDBT 2026 Demo / reviewers in the wild / expert
Diantao Tu
dblp:304/4585
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2025
0009-0007-6115-3771ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
3D vision · 87% Autonomous driving · 10% Segmentation and scene understanding · 2% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
structure from motion |
1.7 | 3 | 2025 | MGSfM: Multi-Camera Geometry Driven Global Structure-from-Motion · ICCV 2025 VidSfM: Robust and Accurate Structure-From-Motion for Monocular Videos · IEEE Trans. Image Process. 2022 PanoPose: Self-supervised Relative Pose Estimation for Panoramic Images · CVPR 2024 |
Computer vision › 3D vision › structure from motion
global structure from motion |
1.1 | 2 | 2025 | MGSfM: Multi-Camera Geometry Driven Global Structure-from-Motion · ICCV 2025 PanoPose: Self-supervised Relative Pose Estimation for Panoramic Images · CVPR 2024 |
Computer vision › 3D vision
camera pose estimation |
0.9 | 1 | 2025 | MGSfM: Multi-Camera Geometry Driven Global Structure-from-Motion · ICCV 2025 |
Computer vision › 3D vision
pose estimation |
0.8 | 1 | 2024 | PanoPose: Self-supervised Relative Pose Estimation for Panoramic Images · CVPR 2024 |
Computer vision › 3D vision › camera pose estimation
relative pose estimation |
0.8 | 1 | 2024 | PanoPose: Self-supervised Relative Pose Estimation for Panoramic Images · CVPR 2024 |
Computer vision › 3D vision › structure from motion
incremental structure from motion |
0.6 | 1 | 2022 | VidSfM: Robust and Accurate Structure-From-Motion for Monocular Videos · IEEE Trans. Image Process. 2022 |
Computer vision › 3D vision › 3d reconstruction
dense 3d reconstruction |
0.5 | 1 | 2021 | Semantically Guided Multi-View Stereo for Dense 3D Road Mapping · ICRA 2021 |
Computer vision › 3D vision › 3d reconstruction
multi-view stereo |
0.5 | 1 | 2021 | Semantically Guided Multi-View Stereo for Dense 3D Road Mapping · ICRA 2021 |
Robotics › Autonomous driving
perception |
0.5 | 1 | 2021 | Semantically Guided Multi-View Stereo for Dense 3D Road Mapping · ICRA 2021 |
Robotics › Autonomous driving › perception › camera-based perception
multi-camera perception |
0.3 | 1 | 2025 | MGSfM: Multi-Camera Geometry Driven Global Structure-from-Motion · ICCV 2025 |
Computer vision › 3D vision
depth estimation |
0.2 | 1 | 2024 | PanoPose: Self-supervised Relative Pose Estimation for Panoramic Images · CVPR 2024 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.1 | 1 | 2021 | Semantically Guided Multi-View Stereo for Dense 3D Road Mapping · ICRA 2021 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
soft constraints |
0.1 | 1 | 2021 | Semantically Guided Multi-View Stereo for Dense 3D Road Mapping · ICRA 2021 |
Methods — techniques the papers use, named apart from their topics
rotation averaging · 1.4translation averaging · 0.9convex optimization · 0.9self-supervised learning · 0.8posenet · 0.8panoramic camera · 0.8depth-net · 0.8loop closure · 0.6semantic segmentation · 0.5patchmatch · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MGSfM: Multi-Camera Geometry Driven Global Structure-from-MotionabstractMulti-camera systems are increasingly vital in the environmental perception of autonomous vehicles and robotics. Their physical configuration offers inherent fixed relative pose constraints that benefit Structure-from-Motion (SfM). However, traditional global SfM systems struggle with robustness due to their optimization framework. We propose a novel global motion averaging framework for multi-camera systems, featuring two core components: a decoupled rotation averaging module and a hybrid translation averaging module. Our rotation averaging employs a hierarchical strategy by first estimating relative rotations within rigid camera units and then computing global rigid unit rotations. To enhance the robustness of translation averaging, we incorporate both camera-to-camera and camera-to-point constraints to initialize camera positions and 3D points with a convex distance-based objective function and refine them with an unbiased non-bilinear angle-based objective function. Experiments on large-scale datasets show that our system matches or exceeds incremental SfM accuracy while significantly improving efficiency. Our framework outperforms existing global SfM methods, establishing itself as a robust solution for real-world multi-camera SfM applications. The code is available at https://github.com/3dv-casia/MGSfM/. Peilin Tao, Hainan Cui, Diantao Tu, Shuhan Shen |
ICCV | 3 |
| 2024 | PanoPose: Self-supervised Relative Pose Estimation for Panoramic ImagesabstractScaled relative pose estimation, i.e., estimating relative rotation and scaled relative translation between two images, has always been a major challenge in global Structure-from-Motion (SfM). This difficulty arises because the two-view relative translation computed by traditional geometric vision methods, e.g. the five-point algorithm, is scaleless. Many researchers have proposed diverse translation averaging methods to solve this problem. Instead of solving the problem in the motion averaging phase, we focus on estimating scaled relative pose with the help of panoramic cameras and deep neural networks. In this paper, a novel network, namely PanoPose, is proposed to estimate the relative motion in a fully self-supervised manner and a global SfM pipeline is built for panorama images. The proposed PanoPose comprises a depth-net and a pose-net, with self-supervision achieved by reconstructing the reference image from its neighboring images based on the estimated depth and relative pose. To maintain precise pose estimation under large viewing angle differences, we randomly rotate the panoramic images and pre-train the posenet with images before and after the rotation. To enhance scale accuracy, a fusion block is introduced to incorporate depth information into pose estimation. Extensive experiments on panoramic SfM datasets demonstrate the effectiveness of PanoPose compared with state-of-the-arts. Diantao Tu, Hainan Cui, Xianwei Zheng, Shuhan Shen |
CVPR | 1 |
| 2022 | Multi-Camera-LiDAR Auto-Calibration by Joint Structure-from-MotionabstractMultiple sensors, especially cameras and LiDARs, are widely used in autonomous vehicles. In order to fuse data from different sensors accurately, precise calibrations are required, including camera intrinsic parameters, and relative poses between multiple cameras and LiDARs. However, most existing camera-LiDAR calibration methods need to place manually designed calibration objects in multiple locations and multiple times, which are time-consuming and labor-intensive, and are not suitable for frequent use. To address that, in this paper we proposed a novel calibration pipeline that can automatically calibrate multiple cameras and multiple LiDARs in a Structure-from-Motion (SfM) process. In our pipeline, we first perform a global SfM on all images with the help of rough LiDAR data to get the initial poses of all sensors. Then, feature points on lines and planes are extracted from both SfM point cloud and LiDARs. With these features, a global Bundle Adjustment is performed to minimize the point reprojection errors, point-to-line errors, and point-to-plane errors together. During this minimization process, camera intrinsic parameters, camera and LiDAR poses, and SfM point cloud are refined jointly. The proposed method uses the characteristics of natural scenes, does not require manually designed calibration objects, and incorporates all calibration parameters into a unified optimization framework. Experiments on autonomous vehicles with different sensor configurations demonstrate the effectiveness and robustness of the proposed method. Diantao Tu, Baoyu Wang, Hainan Cui, Yuqian Liu, Shuhan Shen |
IROS | 1 |
| 2022 | VidSfM: Robust and Accurate Structure-From-Motion for Monocular VideosabstractWith the popularization of smartphones, larger collection of videos with high quality is available, which makes the scale of scene reconstruction increase dramatically. However, high-resolution video produces more match outliers, and high frame rate video brings more redundant images. To solve these problems, a tailor-made framework is proposed to realize an accurate and robust structure-from-motion based on monocular videos. The key ideas include two points: one is to use the spatial and temporal continuity of video sequences to improve the accuracy and robustness of reconstruction; the other is to use the redundancy of video sequences to improve the efficiency and scalability of system. Our technical contributions include an adaptive way to identify accurate loop matching pairs, a cluster-based camera registration algorithm, a local rotation averaging scheme to verify the pose estimate and a local images extension strategy to reboot the incremental reconstruction. In addition, our system can integrate data from different video sequences, allowing multiple videos to be simultaneously reconstructed. Extensive experiments on both indoor and outdoor monocular videos demonstrate that our method outperforms the state-of-the-art approaches in robustness, accuracy and scalability. Hainan Cui, Diantao Tu, Fulin Tang, Pengfei Xu 0013, Hongmin Liu 0001, Shuhan Shen |
IEEE Trans. Image Process. | 2 |
| 2021 | Semantically Guided Multi-View Stereo for Dense 3D Road MappingabstractCompared to widely used LiDAR-based mapping in autonomous driving field, image-based mapping method has the advantages of low cost, high resolution, and no need for complex calibration. However, the image-based 3D mapping depends heavily on the texture richness and always leaves holes and outliers in low-textured areas, such as the road surface. To this end, this paper proposed a novel semantically guided Multi-View Stereo method for dense 3D road mapping, which integrates semantic information into PatchMatch-based MVS pipeline and uses image semantic segmentation as soft constraints in neighbor views selection, depth-map initialization, depth propagation, and depth-map completion. Experimental results on public and our own datasets show that, with the help of semantics, the proposed method achieves superior completeness with comparable accuracy for 3D road mapping compared to state-of-the-art MVS methods. Mingzhe Lv, Diantao Tu, Xincheng Tang, Yuqian Liu, Shuhan Shen |
ICRA | 2 |