EDBT 2026 Demo / reviewers in the wild / expert
Ruofeng Wei
dblp:286/6491
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-7071-4135ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Overlap-Aware Online-Adaptive Non-Rigid Registration of Intraoperative Tissue in Minimally Invasive SurgeryabstractNon-rigid registration of intraoperative tissue is essential for surgical navigation and scene reconstruction in minimally invasive surgery. However, accurate registration remains challenging due to significant tissue deformation and partial overlaps caused by laparoscope movement. We propose an Overlap-Aware Online-Adaptive Non-Rigid Registration Method (OANRM) to address these challenges. The framework introduces a Hierarchical Matching Network (HMNet) that simultaneously predicts overlapping regions and their correspondences through a novel similarity-based approach. Our method uniquely incorporates an online adaptation mechanism that continuously fine-tunes the network parameters using unsupervised losses, enabling robust performance across varying surgical scenarios without requiring additional training data. A Transform Displacement Deformation Prediction (TDDP) module further enhances the framework by handling non-overlapping regions through integrating Random Sample Consensus with distance-based interpolation. The method is validated on both artificial datasets with controlled deformations and clinical datasets from real surgical procedures. Experimental results demonstrate that OANRM achieves state-of-the-art performance, significantly outperforming existing methods in handling complex tissue deformations and varying overlap ratios. https://github.com/AIGCer0807/OANRM. Hangjie Mo, Weizhao Cheng, Ziming Shen, Ruofeng Wei, Xiaojian Li 0003, Shanlin Yang |
IEEE Trans. Medical Imaging | 4 |
| 2026 | Real-Time Monocular 2-D and 3-D Perception of Endoluminal Scenes for Controlling Flexible Robotic Endoscopic InstrumentsabstractEndoluminal surgery offers a minimally invasive option for early-stage gastrointestinal and urinary tract, but is limited by basic surgical tools and a steep learning curve. Robotic systems, particularly continuum robots, provide flexible instruments that enable precise, intuitive tissue resection in confined spaces, potentially improving outcomes. This paper presents an integrated visual perception platform for a continuum robotic system in endoluminal surgery. Our objective is to leverage monocular endoscopic image-based perception algorithms to accurately identify the position and orientation of flexible instruments and measure their distances from surrounding tissues. This thorough understanding of continuum robots and surgical scenes enhances the robustness of robotic procedures. We introduce 2D and 3D learning-based perception algorithms and develop a physically-realistic simulator that models the dynamics of flexible instruments. This simulator features a pipeline for generating realistic endoluminal scenes, enabling control of flexible robots in a realistic environment and substantial data collection. Using a continuum robot prototype, we conducted extensive evaluations, including module assessments and system-level evaluation of the perception platform. Results demonstrate that our perception algorithms significantly improve control of flexible instruments, reducing manipulation time by over 70% for trajectory-following tasks and enhancing the understanding of complex surgical scenarios, leading to robust endoluminal surgeries. Ruofeng Wei, Kai Chen 0024, Yui-Lun Ng, Yiyao Ma, Justin D. L. Ho, Hon-Sing Tong, Ka-Wai Kwok, Qi Dou 0001 |
IEEE Trans. Robotics | 1 |
| 2025 | Data-Efficient Learning Control of Continuum Robots in Constrained EnvironmentsabstractThis research investigates learning-based control of continuum robots in constrained environments without relying on analytical models. We propose a data-efficient stochastic control strategy incorporating online model updates to achieve precise manipulation even when arbitrary robot deformations occur due to environmental interactions. A localized Gaussian process regression approach accounting for state stochasticity is first presented to approximate the forward kinematics. The learned model enables uncertainty-aware stochastic predictions via the proposed scaled unscented transform (SUT)-based method for efficient exploration. Leveraging new data, online model updates are performed in a highly sample-efficient manner. Furthermore, a probabilistic model predictive control approach integrating the learned models and chance constraints based on Chebyshev’s inequality is developed for searching an optimal control sequence. Simulations and experiments are performed to demonstrate the effectiveness of the proposed approach for controlling continuum robots in constrained environments using limited observational data.Note to Practitioners—The motivation of this research is to solve the problem of controlling continuum robots in constraint environment. The flexibility of continuum robots significantly affects the manipulation accuracy, and the interaction between the continuum robot and environmental constraints can also lead to unpredictable behavior. Learning control methods that rely only on sensory data, provide a feasible solution to the aforementioned problem. However, current methods lack sample efficiency and the capability to handle unknown environmental constraints. This research proposes a learning control method which can control a flexible continuum robot in constrained environments with high data-efficiency and robustness even when the robot shape undergoes sudden deformations due to contact with obstacles. Hangjie Mo, Ruofeng Wei, Xiaowen Kong, Yun-Hui Liu 0001, Dong Sun 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Absolute Monocular Depth Estimation on Robotic Visual and Kinematics Data via Self-Supervised LearningabstractAccurate estimation of absolute depth from a monocular endoscope is a fundamental task for automatic navigation systems in robotic surgery. Previous works solely rely on uni-modal data (i.e., monocular images), which can only estimate depth values arbitrarily scaled with the real world. In this paper, we present a novel framework, SADER, which explores vision and robot kinematics to estimate the high-quality absolute depth for monocular surgical scenes. To jointly learn the multi-modal data, we introduce a self-distillation based two-stage training policy in the framework. In the first stage, a boosting depth module based on vision transformer is proposed to improve the relative depth estimation network that is trained in a self-supervised method. Then, we develop an algorithm to automatically compute the scale from robot kinematics. By coupling the scale and relative depth data, pseudo absolute depth labels for all images are yielded. In the second stage, we re-train the network with 3D loss supervised by pseudo labels. To make our method generalize to different endoscopes, the learning of endoscopic intrinsics is integrated into the network. In addition, we did cadaver experiments to collect new surgical depth estimation data about robotic laparoscopy for evaluation. Experimental results on public SCARED and cadaver data demonstrate that the SADER outperforms previous state-of-art even stereo-based methods with an accuracy error under 1.90 mm, proving the feasibility of our approach to recover the absolute depth with monocular inputs. Note to Practitioners—This paper aims to solve the problem of absolute monocular depth estimation in automatic surgical navigation by leveraging the multi-modal data from the robot-based endoscopic system. Accurate depth perception with real scales of the monocular scene is essential for the control of surgical robots in automatic navigation. However, current methods can only predict the relative depth of the surgical scene using monocular images. In this article, we propose a self-supervised learning-based method to achieve high-quality absolute depth estimation of monocular endoscopic images. It neither needs manual data annotation, nor other imaging modalities. The experiments extensively validate the feasibility and high performance of our framework for absolute depth estimation on monocular endoscopes. This absolute depth perception framework can be potentially encapsulated into the automatic navigation system in the near future. Ruofeng Wei, Bin Li 0082, Fangxun Zhong, Hangjie Mo, Qi Dou 0001, Yun-Hui Liu 0001, Dong Sun 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | UC-NeRF: Uncertainty-Aware Conditional Neural Radiance Fields From Endoscopic Sparse ViewsabstractVisualizing surgical scenes is crucial for revealing internal anatomical structures during minimally invasive procedures. Novel View Synthesis is a vital technique that offers geometry and appearance reconstruction, enhancing understanding, planning, and decision-making in surgical scenes. Despite the impressive achievements of Neural Radiance Field (NeRF), its direct application to surgical scenes produces unsatisfying results due to two challenges: endoscopic sparse views and significant photometric inconsistencies. In this paper, we propose uncertainty-aware conditional NeRF for novel view synthesis to tackle the severe shape-radiance ambiguity from sparse surgical views. The core of UC-NeRF is to incorporate the multi-view uncertainty estimation to condition the neural radiance field for modeling the severe photometric inconsistencies adaptively. Specifically, our UC-NeRF first builds a consistency learner in the form of multi-view stereo network, to establish the geometric correspondence from sparse views and generate uncertainty estimation and feature priors. In neural rendering, we design a base-adaptive NeRF network to exploit the uncertainty estimation for explicitly handling the photometric inconsistencies. Furthermore, an uncertainty-guided geometry distillation is employed to enhance geometry learning. Experiments on the SCARED and Hamlyn datasets demonstrate our superior performance in rendering appearance and geometry, consistently outperforming the current state-of-the-art approaches. Our code will be released at https://github.com/wrld/UC-NeRF. Jiangliu Wang, Ruofeng Wei, Qi Dou 0001, Yun-Hui Liu 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2024 | Shape-Guided Configuration-Aware Learning for Endoscopic-Image-Based Pose Estimation of Flexible Robotic Instruments
Yiyao Ma, Kai Chen 0028, Hon-Sing Tong, Ruofeng Wei, Yui-Lun Ng, Ka-Wai Kwok, Qi Dou 0001 |
ECCV (22) | 4 |
| 2024 | Enhanced Scale-Aware Depth Estimation for Monocular Endoscopic Scenes with Geometric Modeling
Ruofeng Wei, Bin Li 0082, Kai Chen 0028, Yiyao Ma, Yun-Hui Liu 0001, Qi Dou 0001 |
MICCAI (6) | 1 |
| 2023 | Autonomous Intelligent Navigation for Flexible Endoscopy Using Monocular Depth Guidance and 3-D Shape PlanningabstractRecent advancements toward perception and decision-making of flexible endoscopes have shown great potential in computer-aided surgical interventions. However, owing to modeling uncertainty and inter-patient anatomical variation in flexible endoscopy, the challenge remains for efficient and safe navigation in patient-specific scenarios. This paper presents a novel data-driven framework with self-contained visual-shape fusion for autonomous intelligent navigation of flexible endoscopes requiring no priori knowledge of system models and global environments. A learning-based adaptive visual servoing controller is proposed to online update the eye-in-hand vision-motor configuration and steer the endoscope, which is guided by monocular depth estimation via a vision transformer (ViT). To prevent unnecessary and excessive interactions with surrounding anatomy, an energy-motivated shape planning algorithm is introduced through entire endoscope 3-D proprioception from embedded fiber Bragg grating (FBG) sensors. Furthermore, a model predictive control (MPC) strategy is developed to minimize the elastic potential energy flow and simultaneously optimize the steering policy. Dedicated navigation experiments on a robotic-assisted flexible endoscope with an FBG fiber in several phantom environments demonstrate the effectiveness and adaptability of the proposed framework. Yiang Lu, Ruofeng Wei, Bin Li 0082, Wei Chen 0068, Jianshu Zhou, Qi Dou 0001, Dong Sun 0001, Yun-Hui Liu 0001 |
ICRA | 2 |
| 2022 | 3D Perception based Imitation Learning under Limited Demonstration for Laparoscope Control in Robotic SurgeryabstractAutomatic laparoscope motion control is fundamentally important for surgeons to efficiently perform operations. However, its traditional control methods based on tool tracking without considering information hidden in surgical scenes are not intelligent enough, while the latest supervised imitation learning (IL)-based methods require expensive sensor data and suffer from distribution mismatch issues caused by limited demonstrations. In this paper, we propose a novel Imitation Learning framework for Laparoscope Control (ILLC) with reinforcement learning (RL), which can efficiently learn the control policy from limited surgical video clips. Specially, we first extract surgical laparoscope trajectories from unlabeled videos as the demonstrations and reconstruct the corresponding surgical scenes. To fully learn from limited motion trajectory demonstrations, we propose Shape Preserving Trajectory Augmentation (SPTA) to augment these data, and build a simulation environment that supports parallel RGB-D rendering to reinforce the RL policy for interacting with the environment efficiently. With adversarial training for IL, we obtain the laparoscope control policy based on the generated rollouts and surgical demonstrations. Extensive experiments are conducted in unseen reconstructed surgical scenes, and our method outperforms the previous IL methods, which proves the feasibility of our unified learning-based framework for laparoscope control. Bin Li 0082, Ruofeng Wei, Bo Lu 0001, Chi Hang Yee, Chi-Fai Ng, Pheng-Ann Heng, Qi Dou 0001, Yun-Hui Liu 0001 |
ICRA | 2 |
| 2022 | Distilled Visual and Robot Kinematics Embeddings for Metric Depth Estimation in Monocular Scene ReconstructionabstractEstimating precise metric depth and scene reconstruction from monocular endoscopy is a fundamental task for surgical navigation in robotic surgery. However, traditional stereo matching adopts binocular images to perceive the depth information, which is difficult to transfer to the soft robotics-based surgical systems due to the use of monocular endoscopy. In this paper, we present a novel framework that combines robot kinematics and monocular endoscope images with deep unsupervised learning into a single network for metric depth estimation and then achieve 3D reconstruction of complex anatomy. Specifically, we first obtain the relative depth maps of surgical scenes by leveraging a brightness-aware monocular depth estimation method. Then, the corresponding endoscope poses are computed based on non-linear optimization of geo-metric and photometric reprojection residuals. Afterwards, we develop a Depth-driven Sliding Optimization (DDSO) algorithm to extract the scaling coefficient from kinematics and calculated poses offline. By coupling the metric scale and relative depth data, we form a robust ensemble that represents the metric and consistent depth. Next, we treat the ensemble as supervisory labels to train a metric depth estimation network for surgeries (i.e., MetricDepthS-Net) that distills the embeddings from the robot kinematics, endoscopic videos, and poses. With accurate metric depth estimation, we utilize a dense visual reconstruction method to recover the 3D structure of the whole surgical site. We have extensively evaluated the proposed framework on public SCARED and achieved comparable performance with stereo-based depth estimation methods. Our results demon-strate the feasibility of the proposed approach to recover the metric depth and 3D structure with monocular inputs. Ruofeng Wei, Bin Li 0082, Hangjie Mo, Fangxun Zhong, Yonghao Long 0001, Qi Dou 0001, Yun-Hui Liu 0001, Dong Sun 0001 |
IROS | 1 |