EDBT 2026 Demo / reviewers in the wild / expert
Beilei Cui
dblp:253/0485
· DBLP profile ↗
9ranked-venue papers
4as first author
8since 2021 · last 2025
0009-0009-7900-8032ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Advancing Dense Endoscopic Reconstruction with Gaussian Splatting-Driven Surface Normal-Aware Tracking and MappingabstractSimultaneous Localization and Mapping (SLAM) is essential for precise surgical interventions and robotic tasks in minimally invasive procedures. While recent advancements in 3D Gaussian Splatting (3DGS) have improved SLAM with high-quality novel view synthesis and fast rendering, these systems struggle with accurate depth and surface reconstruction due to multi-view inconsistencies. Simply incorporating SLAM and 3DGS leads to mismatches between the reconstructed frames. In this work, we present Endo-2DTAM, a real-time endoscopic SLAM system with 2D Gaussian Splatting (2DGS) to address these challenges. Endo-2DTAM incorporates a surface normal-aware pipeline, which consists of tracking, mapping, and bundle adjustment modules for geometrically accurate reconstruction. Our robust tracking module combines point-topoint and point-to-plane distance metrics, while the mapping module utilizes normal consistency and depth distortion to enhance surface reconstruction quality. We also introduce a pose-consistent strategy for efficient and geometrically coherent keyframe sampling. Extensive experiments on public endoscopic datasets demonstrate that Endo-2DTAM achieves an RMSE of$1.87 \pm 0.63 \mathbf{m m}$for depth reconstruction of surgical scenes while maintaining computationally efficient tracking, high-quality visual appearance, and real-time rendering. Our code will be released at github.com/lastbasket/Endo-2DTAM. Yiming Huang 0007, Beilei Cui, Long Bai 0008, Zhen Chen 0018, Jinlin Wu, Zhen Li 0026, Hongbin Liu 0001, Hongliang Ren 0001 |
ICRA | 2 |
| 2025 | Endo-4DGX: Robust Endoscopic Scene Reconstruction and Illumination Correction with Gaussian Splatting
Yiming Huang 0007, Long Bai 0008, Beilei Cui, Yanheng Li 0002, Tong Chen 0011, Jie Wang 0097, Jinlin Wu, Zhen Lei 0001, Hongbin Liu 0001, Hongliang Ren 0001 |
MICCAI (9) | 3 |
| 2025 | SurgTPGS: Semantic 3D Surgical Scene Understanding with Text Promptable Gaussian Splatting
Yiming Huang 0007, Long Bai 0008, Beilei Cui, Kun Yuan 0004, Guankun Wang, Mobarak I. Hoque, Nicolas Padoy, Nassir Navab, Hongliang Ren 0001 |
MICCAI (9) | 3 |
| 2025 | V²-SfMLearner: Learning Monocular Depth and Ego-Motion for Multimodal Wireless Capsule EndoscopyabstractDeep learning can predict depth maps and capsule ego-motion from capsule endoscopy videos, aiding in 3D scene reconstruction and lesion localization. However, the collisions of the capsule endoscopies within the gastrointestinal tract cause vibration perturbations in the training data. Existing solutions focus solely on vision-based processing, neglecting other auxiliary signals like vibrations that could reduce noise and improve performance. Therefore, we propose V2-SfMLearner, a multimodal approach integrating vibration signals into vision-based depth and capsule motion estimation for monocular capsule endoscopy. We construct a multimodal capsule endoscopy dataset containing vibration and visual signals, and our artificial intelligence solution develops an unsupervised method using vision-vibration signals, effectively eliminating vibration perturbations through multimodal learning. Specifically, we carefully design a vibration network branch and a Fourier fusion module, to detect and mitigate vibration noises. The fusion framework is compatible with popular vision-only algorithms. Extensive validation on the multimodal dataset demonstrates superior performance and robustness against vision-only algorithms. Without the need for large external equipment, our V2-SfMLearner has the potential for integration into clinical capsule robots, providing real-time and dependable digestive examination tools. The findings show promise for practical implementation in clinical settings, enhancing the diagnostic capabilities of doctors. Note to Practitioners—This paper is motivated by the problem of estimating the depth and ego-motion information for the wireless capsule endoscopy in the human gastrointestinal tract to realize accurate, efficient, robust, and real-time inspection. Our estimation method does not engage any external localization equipment. Instead, inspired by the existing research on integrating capsule endoscopy and inertial measurement units, we introduce vibration signals into vision-based depth and ego-motion estimation approaches, improving the accuracy and robustness of the estimation results based on multimodal learning methods. Research on capsule robots or computer vision can readily be combined with our framework for various clinical and industrial applications. Long Bai 0008, Beilei Cui, Yanheng Li 0002, Shilong Yao, Sishen Yuan, Yanan Wu 0003, Yang Zhang 0053, Max Q.-H. Meng, Zhen Li 0026, Weiping Ding 0001, Hongliang Ren 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2024 | EndoDAC: Efficient Adapting Foundation Model for Self-Supervised Depth Estimation from Any Endoscopic Camera
Beilei Cui, Mobarakol Islam, Long Bai 0008, An Wang 0007, Hongliang Ren 0001 |
MICCAI (6) | 1 |
| 2024 | Endo-4DGS: Endoscopic Monocular Scene Reconstruction with 4D Gaussian Splatting
Yiming Huang 0007, Beilei Cui, Long Bai 0008, Mengya Xu, Mobarakol Islam, Hongliang Ren 0001 |
MICCAI (6) | 2 |
| 2023 | Rectifying Noisy Labels with Sequential Prior: Multi-scale Temporal Feature Affinity Learning for Robust Video Segmentation
Beilei Cui, Minqing Zhang, Mengya Xu, An Wang 0007, Wu Yuan 0001, Hongliang Ren 0001 |
MICCAI (9) | 1 |
| 2021 | Self-Guided Deep Multi-View Subspace Clustering NetworkabstractTo cluster the data with complex structures, Deep Subspace Clustering Network (DSCN) extracts the subspace relations among non-linear latent features. However, the performance improvement has encountered bottlenecks due to the lack of supervision. Meantime, in multi-view settings, most of DSCN-based methods underestimate the significance of view-fusion, which always adopt simple tactics. To address these issues, we propose a self-supervised model for simultaneous subspace clustering, consensus construction and self-guided learning, named as Self-Guided Deep Multi-view Subspace Clustering Network (SG-DMSC). We utilize DSCN to learn a complex subspace representation for each single-view. Considering their different importance, we design the view-fusion layer to establish the agreement. We construct a novel loss term, the spectral supervisor, so that the consensus can be more clustering-friendly by the self-guidance of pseudo labels. Theoretical support is provided to reflect the validity of this self-guided strategy. An alternate iterative optimization algorithm is presented to handle SG-DMSC. Experiments on real-world datasets confirm its efficacy compared with others. Beilei Cui, Hong Yu 0005, Linlin Zong |
ICME | 1 |
| 2019 | Self-Weighted Multi-View Clustering with Deep Matrix FactorizationabstractDue to the efficiency of exploring multiple views of the real-word data, Multi-View Clustering (MVC) has attracted extensive attention from the scholars and researches based on it have made significant progress. However, multi-view data with numerous complementary information is vulnerable to various factors (such as noise). So it is an important and challenging task to discover the intrinsic characteristics hidden deeply in the data. In this paper, we present a novel MVC algorithm based on deep matrix factorization, named Self-Weighted Multi-view Clustering with Deep Matrix Factorization (SMDMF). By performing the deep decomposition structure, SMDMF can eliminate interference and reveal semantic information of the multi-view data. To properly integrate the complementary information among views, it assigns an automatic weight for each view without introducing supernumerary parameters. We also analyze the convergence of the algorithm and discuss the hierarchical parameters. The experimental results on four datasets show our algorithm is superior to other comparisons in all aspects. Beilei Cui, Hong Yu 0005, Siwen Li |
ACML | 1 |