Botao Ye

dblp:227/4610 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
7since 2021 · last 2025
0009-0006-0656-0174ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
3D vision · 70% Generative modeling · 7% Reinforcement learning · 7%
Computer graphics and multimedia
1 paper
Rendering · 61% Computer animation and physical simulation · 30% Virtual and augmented reality · 9%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d reconstruction
1.522025
No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images · ICLR 2025
Self-Supervised Super-Plane for Neural 3D Reconstruction · CVPR 2023
Computer vision › 3D vision › neural rendering
3d gaussian splatting
0.912025
No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images · ICLR 2025
Computer vision › 3D vision
camera pose estimation
0.912025
No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images · ICLR 2025
Robotics › Autonomous driving › trajectory prediction
ego motion prediction
0.912025
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control · CVPR 2025
Computer vision › 3D vision
novel view synthesis
0.912025
No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images · ICLR 2025
Computer vision › 3D vision › 3d reconstruction › uncalibrated reconstruction
pose-free reconstruction
0.912025
No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images · ICLR 2025
Computer vision › 3D vision › camera pose estimation
sparse-view pose estimation
0.912025
No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images · ICLR 2025
Machine learning › Generative modeling
video generation
0.912025
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control · CVPR 2025
Machine learning › Reinforcement learning › model-based reinforcement learning
world model
0.912025
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control · CVPR 2025
Rendering
gaussian splatting
0.912025
HoliGS: Holistic Gaussian Splatting for Embodied View Synthesis · NeurIPS 2025
Rendering
novel view synthesis
0.912025
HoliGS: Holistic Gaussian Splatting for Embodied View Synthesis · NeurIPS 2025
Computer vision › 3D vision › 3d reconstruction › surface reconstruction › neural surface reconstruction
neural implicit surface reconstruction
0.712023
Self-Supervised Super-Plane for Neural 3D Reconstruction · CVPR 2023
Computer vision › 3D vision › 3d reconstruction › surface reconstruction
planar surface reconstruction
0.712023
Self-Supervised Super-Plane for Neural 3D Reconstruction · CVPR 2023
Machine learning › Deep learning architectures and training
data augmentation
0.612022
Exploring Geometric Consistency for Monocular 3D Object Detection · CVPR 2022
Computer vision › 3D vision › multi-view geometry
geometric consistency
0.612022
Exploring Geometric Consistency for Monocular 3D Object Detection · CVPR 2022
Computer vision › 3D vision › 3d object detection › image-based 3d object detection
monocular 3d object detection
0.612022
Exploring Geometric Consistency for Monocular 3D Object Detection · CVPR 2022
Computer vision › Video understanding and tracking
object tracking
0.612022
Joint Feature Learning and Relation Modeling for Tracking: A One-Stream Framework · ECCV (22) 2022
Computer vision › 3D vision › geometric deep learning › 3d representation learning
self-supervised 3d representation learning
0.212023
Self-Supervised Super-Plane for Neural 3D Reconstruction · CVPR 2023

Methods — techniques the papers use, named apart from their topics

skeleton-driven articulation · 0.9rigid transformation · 0.9photometric loss · 0.9multimodal learning · 0.9invertible neural flow · 0.9feed-forward reconstruction · 0.9autoregressive generation · 0.93d gaussian splatting · 0.9self-supervised super-plane constraint · 0.7iterative training · 0.7joint feature learning · 0.6geometry-aware augmentation · 0.6camera perturbation · 0.6
YearPublicationVenuePosition
2025 Synthesizing Consistent Novel Views Via 3D Epipolar Attention Without Re-Training
abstract
Large diffusion models demonstrate remarkable zeroshot capabilities in novel view synthesis from a single image. However, these models often face challenges in maintaining consistency across novel and reference views. A crucial factor leading to this issue is the limited utilization of contextual information from reference views. Specifically, when there is an overlap in the viewing frustum between two views, it is essential to ensure that the corresponding regions maintain consistency in both geometry and appearance. This observation leads to a simple yet effective approach, where we propose to use epipolar geometry to locate and retrieve overlapping information from the input view. This information is then incorporated into the generation of target views, eliminating the need for training or fine-tuning, as the process requires no learnable parameters. Furthermore, to enhance the overall consistency of generated views, we extend the utilization of epipolar attention to a multi-view setting, allowing retrieval of overlapping information from the input view and other target views. Qualitative and quantitative experimental results demonstrate the effectiveness of our method in significantly improving the consistency of synthesized views without the need for any fine-tuning. Moreover, This enhancement also boosts the performance of downstream applications such as 3D reconstruction. The code is available at https://github.com/botaoye/ConsisSyn.
Botao Ye, Sifei Liu, Marc Pollefeys, Ming-Hsuan Yang 0001
3DV1
2025 GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control
abstract
We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, our model has precise control over object dynamics, ego-agent motion and human poses. GEM generates paired RGB and depth outputs for richer spatial understanding. We introduce autoregressive noise schedules to enable stable long-horizon generations. Our dataset is comprised of 4000+ hours of multimodal data across domains like autonomous driving, egocentric human activities, and drone flights. Pseudo-labels are used to get depth maps, ego-trajectories, and human poses. We use a comprehensive evaluation framework, including a new Control of Object Manipulation (COM) metric, to assess controllability. Experiments show GEM excels at generating diverse, controllable scenarios and temporal consistency over long generations. Code, models, and datasets are fully open-sourced1.
Mariam Hassan, Sebastian Stapf, Ahmad Rahimi, Pedro M. B. Rezende, Yasaman Haghighi, David Brüggemann, Isinsu Katircioglu, Xiaoran Chen, Marco Cannici, Elie Aljalbout, Botao Ye, Xi Wang 0021, Aram Davtyan, Mathieu Salzmann, Davide Scaramuzza 0001, Marc Pollefeys, Paolo Favaro, Alexandre Alahi
CVPR13
2025 No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images
abstract
We introduce NoPoSplat, a feed-forward model capable of reconstructing 3D scenes parameterized by 3D Gaussians from unposed sparse multi-view images. Our model, trained exclusively with photometric loss, achieves real-time 3D Gaussian reconstruction during inference. To eliminate the need for accurate pose input during reconstruction, we anchor one input view's local camera coordinates as the canonical space and train the network to predict Gaussian primitives for all views within this space. This approach obviates the need to transform Gaussian primitives from local coordinates into a global coordinate system, thus avoiding errors associated with per-frame Gaussians and pose estimation. To resolve scale ambiguity, we design and compare various intrinsic embedding methods, ultimately opting to convert camera intrinsics into a token embedding and concatenate it with image tokens as input to the model, enabling accurate scene scale prediction. We utilize the reconstructed 3D Gaussians for novel view synthesis and pose estimation tasks and propose a two-stage coarse-to-fine pipeline for accurate pose estimation. Experimental results demonstrate that our pose-free approach can achieve superior novel view synthesis quality compared to pose-required methods, particularly in scenarios with limited input image overlap. For pose estimation, our method, trained without ground truth depth or explicit matching loss, significantly outperforms the state-of-the-art methods with substantial improvements. This work makes significant advances in pose-free generalizable 3D reconstruction and demonstrates its applicability to real-world scenarios. Code and trained models are available at https://noposplat.github.io/.
Botao Ye, Sifei Liu, Haofei Xu, Marc Pollefeys, Ming-Hsuan Yang 0001, Songyou Peng
ICLR1
2025 HoliGS: Holistic Gaussian Splatting for Embodied View Synthesis
abstract
We propose HoliGS, a novel deformable Gaussian splatting framework that addresses embodied view synthesis from long monocular RGB videos. Unlike prior 4D Gaussian splatting and dynamic NeRF pipelines, which struggle with training overhead in minute-long captures, our method leverages invertible Gaussian Splatting deformation networks to reconstruct large-scale, dynamic environments accurately. Specifically, we decompose each scene into a static background plus time-varying objects, each represented by learned Gaussian primitives undergoing global rigid transformations, skeleton-driven articulation, and subtle non-rigid deformations via an invertible neural flow. This hierarchical warping strategy enables robust free-viewpoint novel-view rendering from various embodied camera trajectories by attaching Gaussians to a complete canonical foreground shape (e.g., egocentric or third-person follow), which may involve substantial viewpoint changes and interactions between multiple actors. Our experiments demonstrate that HoliGS achieves superior reconstruction quality on challenging datasets while significantly reducing both training and rendering time compared to state-of-the-art monocular deformable NeRFs. These results highlight a practical and scalable solution for EVS in real-world scenarios. The source code will be released.
Botao Ye, Xiaojun Shan, Weijie Lyu, Lu Qi 0001, Kelvin C. K. Chan, Yinxiao Li, Ming-Hsuan Yang 0001
NeurIPS3
2023 Self-Supervised Super-Plane for Neural 3D Reconstruction
abstract
Neural implicit surface representation methods show impressive reconstruction results but struggle to handle texture-less planar regions that widely exist in indoor scenes. Existing approaches addressing this leverage image prior that requires assistive networks trained with large-scale annotated datasets. In this work, we introduce a self-supervised super-plane constraint by exploring the free geometry cues from the predicted surface, which can further regularize the reconstruction of plane regions without any other ground truth annotations. Specifically, we introduce an iterative training scheme, where (i) grouping of pixels to formulate a super-plane (analogous to super-pixels), and (ii) optimizing of the scene reconstruction network via a super-plane constraint, are progressively conducted. We demonstrate that the model trained with superplanes surprisingly outperforms the one using conventional annotated planes, as individual super-plane statistically occupies a larger area and leads to more stable training. Extensive experiments show that our self-supervised super-plane constraint significantly improves 3D reconstruction quality even better than using ground truth plane segmentation. Additionally, the plane reconstruction results from our model can be used for auto-labeling for other vision tasks. The code and models are available at https://github.com/botaoye/S3PRecon.
Botao Ye, Sifei Liu, Ming-Hsuan Yang 0001
CVPR1
2022 Exploring Geometric Consistency for Monocular 3D Object Detection
abstract
This paper investigates the geometric consistency for monocular 3D object detection, which suffers from the ill-posed depth estimation. We first conduct a thorough analysis to reveal how existing methods fail to consistently localize objects when different geometric shifts occur. In particular, we design a series of geometric manipulations to diagnose existing detectors and then illustrate their vulnerability to consistently associate the depth with object apparent sizes and positions. To alleviate this issue, we propose four geometry-aware data augmentation approaches to enhance the geometric consistency of the detectors. We first modify some commonly used data augmentation methods for 2D images so that they can maintain geometric consistency in 3D spaces. We demonstrate such modifications are important. In addition, we propose a 3D-specific image perturbation method that employs the camera movement. During the augmentation process, the camera system with the corresponding image is manipulated, while the geometric visual cues for depth recovery are preserved. We show that by using the geometric consistency constraints, the proposed augmentation techniques lead to improvements on the KITTI and nuScenes monocular 3D detection benchmarks with state-of-the-art results. In addition, we demonstrate that the augmentation methods are well suited for semisupervised training and cross-dataset generalization.
Qing Lian, Botao Ye, Ruijia Xu, Weilong Yao, Tong Zhang 0001
CVPR2
2022 Joint Feature Learning and Relation Modeling for Tracking: A One-Stream Framework
Botao Ye, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001
ECCV (22)1
2018 An Overview of Event Based Directional Change for Algorithmic Trading
abstract
This paper outlines a framework of a practicable scheme to facilitate algorithm trading of securities. The proposed scheme is capable to intelligently identify, analyze, and implement the intrinsic directional changes in the price movement of the stock market. An overall qualitative assessment is provided together with survey of existing theoretical and empirical foundations towards the success of such an algorithm. Potential loopholes and roadmap for further improvement are suggested.
Botao Ye, Dejun Xie
SERA1