EDBT 2026 Demo / reviewers in the wild / expert
Yonghao Long 0001
dblp:155/8698-1
· DBLP profile ↗
13ranked-venue papers
2as first author
13since 2021 · last 2025
0000-0003-4474-7854ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning dissection trajectories from expert surgical videos via imitation learning with equivariant diffusion
Yonghao Long 0001, Yueyao Chen, Hon-Chi Yip, Markus Scheppach, Philip W. Y. Chiu, Yeung Yam, Helen M. Meng, Qi Dou 0001 |
Medical Image Anal. | 2 |
| 2024 | Multi-objective Cross-task Learning via Goal-conditioned GPT-based Decision Transformers for Surgical Robot Task AutomationabstractSurgical robot task automation has been a promising research topic for improving surgical efficiency and quality. Learning-based methods have been recognized as an interesting paradigm and been increasingly investigated. However, existing approaches encounter difficulties in long-horizon goal-conditioned tasks due to the intricate compositional structure, which requires decision-making for a sequence of sub-steps and understanding of inherent dynamics of goal-reaching tasks. In this paper, we propose a new learning-based framework by leveraging the strong reasoning capability of the GPT-based architecture to automate surgical robotic tasks. The key to our approach is developing a goal-conditioned decision transformer to achieve sequential representations with goal-aware future indicators in order to enhance temporal reasoning. Moreover, considering to exploit a general understanding of dynamics inherent in manipulations, thus making the model’s reasoning ability to be task-agnostic, we also design a cross-task pretraining paradigm that uses multiple training objectives associated with data from diverse tasks. We have conducted extensive experiments on 10 tasks using the surgical robot learning simulator SurRoL [1]. The results show that our new approach achieves promising performance and task versatility compared to existing methods. The learned trajectories can be deployed on the da Vinci Research Kit (dVRK) for validating its practicality in real surgical robot settings. Our project website is at: https://med-air.github.io/SurRoL. Jiawei Fu 0001, Yonghao Long 0001, Kai Chen 0028, Qi Dou 0001 |
ICRA | 2 |
| 2024 | Self-Supervised Cyclic Diffeomorphic Mapping for Soft Tissue Deformation Recovery in Robotic Surgery ScenesabstractThe ability to recover tissue deformation from surgical video is fundamental for many downstream applications in robotic surgery. Despite noticeable advancements, this task remains under-explored due to the complex dynamics of soft tissues manipulated by surgical instruments. Achieving dense and accurate tissue tracking is further complicated by ambiguous pixel correspondence in regions with homogeneous texture. In this paper, we introduce a novel self-supervised framework to recover tissue deformations from stereo surgical videos. Our approach integrates semantics, cross-frame motion flow, and long-range temporal dependencies to accurately represent tissue dynamics for deformation recovery. Moreover, we incorporate diffeomorphic mapping to regularize the warping field to be physically more realistic. To comprehensively evaluate our method, we collected stereo surgical video clips containing three types of tissue manipulation (i.e., pushing, dissection and retraction) from two surgical procedures (i.e., hemicolectomy and mesorectal excision). Our method demonstrates promising results in capturing tissue 3D deformation, and generalizes well across different actions and procedures. It also outperforms current state-of-the-art approaches based on non-rigid registration and optical flow estimation. To the best of our knowledge, this is the first work on self-supervised learning for dense tissue deformation modeling from stereo surgical videos. The paper's code is available at: https://github.com/ med-air/RecoverTissueDeform. Shizhan Gong, Yonghao Long 0001, Kai Chen 0024, Yuliang Xiao, Alexis Cheng, Zerui Wang, Qi Dou 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2023 | Value-Informed Skill Chaining for Policy Learning of Long-Horizon Tasks with Surgical RobotabstractReinforcement learning is still struggling with solving long-horizon surgical robot tasks which involve multiple steps over an extended duration of time due to the policy exploration challenge. Recent methods try to tackle this problem by skill chaining, in which the long-horizon task is decomposed into multiple subtasks for easing the exploration burden and subtask policies are temporally connected to complete the whole long-horizon task. However, smoothly connecting all subtask policies is difficult for surgical robot scenarios. Not all states are equally suitable for connecting two adjacent subtasks. An undesired terminate state of the previous subtask would make the current subtask policy unstable and result in a failed execution. In this work, we introduce value-informed skill chaining (ViSkill), a novel reinforcement learning framework for long-horizon surgical robot tasks. The core idea is to distinguish which terminal state is suitable for starting all the following subtask policies. To achieve this target, we introduce a state value function that estimates the expected success probability of the entire task given a state. Based on this value function, a chaining policy is learned to instruct subtask policies to terminate at the state with the highest value so that all subsequent policies are more likely to be connected for accomplishing the task. We demonstrate the effectiveness of our method on three complex surgical robot tasks from SurRoL, a comprehensive surgical simulation platform, achieving high task success rates and execution efficiency. Code is available at https: / /github. com/med-air/ViSkill. Kai Chen 0028, Jianan Li 0006, Yonghao Long 0001, Qi Dou 0001 |
IROS | 5 |
| 2023 | Visual-Kinematics Graph Learning for Procedure-Agnostic Instrument Tip Segmentation in Robotic SurgeriesabstractAccurate segmentation of surgical instrument tip is an important task for enabling downstream applications in robotic surgery, such as surgical skill assessment, tool-tissue interaction and deformation modeling, as well as surgical autonomy. However, this task is very challenging due to the small sizes of surgical instrument tips, and significant variance of surgical scenes across different procedures. Although much effort has been made on visual-based methods, existing segmentation models still suffer from low robustness thus not usable in practice. Fortunately, kinematics data from the robotic system can provide reliable prior for instrument location, which is consistent regardless of different surgery types. To make use of such multi-modal information, we propose a novel visual-kinematics graph learning framework to accurately segment the instrument tip given various surgical procedures. Specifically, a graph learning framework is proposed to encode relational features of instrument parts from both image and kinematics. Next, a cross-modal contrastive loss is designed to incorporate robust geometric prior from kinematics to image for tip segmentation. We have conducted experiments on a private paired visual-kinematics dataset including multiple procedures, i.e., prostatectomy, total mesorectal excision, fundoplication and distal gastrectomy on cadaver, and distal gastrectomy on porcine. The leave-one-procedure-out cross validation demon-strated that our proposed multi-modal segmentation method significantly outperformed current image-based state-of-the-art approaches, exceeding averagely 11.2% on Dice. Yonghao Long 0001, Kai Chen 0028, Cheuk Hei Leung, Zerui Wang, Qi Dou 0001 |
IROS | 2 |
| 2023 | Comparative validation of machine learning algorithms for surgical workflow and skill analysis with the HeiChole benchmarkabstractPURPOSE: Surgical workflow and skill analysis are key technologies for the next generation of cognitive surgical assistance systems. These systems could increase the safety of the operation through context-sensitive warnings and semi-autonomous robotic assistance or improve training of surgeons via data-driven feedback. In surgical workflow analysis up to 91% average precision has been reported for phase recognition on an open data single-center video dataset. In this work we investigated the generalizability of phase recognition algorithms in a multicenter setting including more difficult recognition tasks such as surgical action and surgical skill. METHODS: To achieve this goal, a dataset with 33 laparoscopic cholecystectomy videos from three surgical centers with a total operation time of 22 h was created. Labels included framewise annotation of seven surgical phases with 250 phase transitions, 5514 occurences of four surgical actions, 6980 occurences of 21 surgical instruments from seven instrument categories and 495 skill classifications in five skill dimensions. The dataset was used in the 2019 international Endoscopic Vision challenge, sub-challenge for surgical workflow and skill analysis. Here, 12 research teams trained and submitted their machine learning algorithms for recognition of phase, action, instrument and/or skill assessment. RESULTS: F1-scores were achieved for phase recognition between 23.9% and 67.7% (n = 9 teams), for instrument presence detection between 38.5% and 63.8% (n = 8 teams), but for action recognition only between 21.8% and 23.3% (n = 5 teams). The average absolute error for skill assessment was 0.78 (n = 1 team). CONCLUSION: Surgical workflow and skill analysis are promising technologies to support the surgical team, but there is still room for improvement, as shown by our comparison of machine learning algorithms. This novel HeiChole benchmark can be used for comparable evaluation and validation of future work. In future studies, it is of utmost importance to create more open, high-quality datasets in order to allow the development of artificial intelligence and cognitive robotics in surgery. Martin Wagner 0001, Beat P. Müller-Stich, Anna Kisilenko, Patrick Heger, Lars Mündermann, David M. Lubotsky, Tornike Davitashvili, Manuela Capek, Annika Reinke, Carissa Reid, Tong Yu 0009, Armine Vardazaryan, Chinedu Innocent Nwoye, Nicolas Padoy, Eungjoo Lee 0001, Constantin Disch, Hans Meine, Tong Xia, Fucang Jia, Satoshi Kondo, Wolfgang Reiter, Yueming Jin, Yonghao Long 0001, Meirui Jiang, Qi Dou 0001, Pheng-Ann Heng, Isabell Twick, Kadir Kirtaç, Enes Hosgor, Jon Lindström Bolmgren, Michael Stenzel, Björn von Siemens, Zhenxiao Ge, Haiming Sun, Di Xie, Mengqi Guo, Daochang Liu, Hannes Kenngott, Felix Nickel, Moritz von Frankenberg, Franziska Mathis-Ullrich, Annette Kopp-Schneider, Lena Maier-Hein, Stefanie Speidel, Sebastian Bodenstedt |
Medical Image Anal. | 26 |
| 2022 | Distilled Visual and Robot Kinematics Embeddings for Metric Depth Estimation in Monocular Scene ReconstructionabstractEstimating precise metric depth and scene reconstruction from monocular endoscopy is a fundamental task for surgical navigation in robotic surgery. However, traditional stereo matching adopts binocular images to perceive the depth information, which is difficult to transfer to the soft robotics-based surgical systems due to the use of monocular endoscopy. In this paper, we present a novel framework that combines robot kinematics and monocular endoscope images with deep unsupervised learning into a single network for metric depth estimation and then achieve 3D reconstruction of complex anatomy. Specifically, we first obtain the relative depth maps of surgical scenes by leveraging a brightness-aware monocular depth estimation method. Then, the corresponding endoscope poses are computed based on non-linear optimization of geo-metric and photometric reprojection residuals. Afterwards, we develop a Depth-driven Sliding Optimization (DDSO) algorithm to extract the scaling coefficient from kinematics and calculated poses offline. By coupling the metric scale and relative depth data, we form a robust ensemble that represents the metric and consistent depth. Next, we treat the ensemble as supervisory labels to train a metric depth estimation network for surgeries (i.e., MetricDepthS-Net) that distills the embeddings from the robot kinematics, endoscopic videos, and poses. With accurate metric depth estimation, we utilize a dense visual reconstruction method to recover the 3D structure of the whole surgical site. We have extensively evaluated the proposed framework on public SCARED and achieved comparable performance with stereo-based depth estimation methods. Our results demon-strate the feasibility of the proposed approach to recover the metric depth and 3D structure with monocular inputs. Ruofeng Wei, Bin Li 0082, Hangjie Mo, Fangxun Zhong, Yonghao Long 0001, Qi Dou 0001, Yun-Hui Liu 0001, Dong Sun 0001 |
IROS | 5 |
| 2022 | Neural Rendering for Stereo 3D Reconstruction of Deformable Tissues in Robotic Surgery
Yuehao Wang, Yonghao Long 0001, Siu Hin Fan, Qi Dou 0001 |
MICCAI (8) | 2 |
| 2022 | AutoLaparo: A New Dataset of Integrated Multi-tasks for Image-guided Surgical Automation in Laparoscopic Hysterectomy
Ziyi Wang 0006, Bo Lu 0001, Yonghao Long 0001, Fangxun Zhong, Tak Hong Cheung, Qi Dou 0001, Yun-Hui Liu 0001 |
MICCAI (8) | 3 |
| 2021 | Relational Graph Learning on Visual and Kinematics Embeddings for Accurate Gesture Recognition in Robotic SurgeryabstractAutomatic surgical gesture recognition is fundamentally important to enable intelligent cognitive assistance in robotic surgery. With recent advancement in robot-assisted minimally invasive surgery, rich information including surgical videos and robotic kinematics can be recorded, which provide complementary knowledge for understanding surgical gestures. However, existing methods either solely adopt uni-modal data or directly concatenate multi-modal representations, which can not sufficiently exploit the informative correlations inherent in visual and kinematics data to boost gesture recognition accuracies. In this regard, we propose a novel online approach of multi-modal relational graph network (i.e., MRG-Net) to dynamically integrate visual and kinematics information through interactive message propagation in the latent feature space. In specific, we first extract embeddings from video and kinematics sequences with temporal convolutional networks and LSTM units. Next, we identify multi-relations in these multi-modal embeddings and leverage them through a hierarchical relational graph learning module. The effectiveness of our method is demonstrated with state-of-the-art results on the public JIGSAWS dataset, outperforming current uni-modal and multi-modal methods on both suturing and knot typing tasks. Furthermore, we validated our method on in-house visual-kinematics datasets collected with da Vinci Research Kit (dVRK) platforms in two centers, with consistent promising performance achieved. Our code and data are released at: https://www.cse.cuhk.edu.hk/~yhlong/mrgnet.html. Yonghao Long 0001, Jie Ying Wu, Bo Lu 0001, Yueming Jin, Mathias Unberath, Yun-Hui Liu 0001, Pheng-Ann Heng, Qi Dou 0001 |
ICRA | 1 |
| 2021 | Trans-SVNet: Accurate Phase Recognition from Surgical Videos via Hybrid Embedding Aggregation Transformer
Yueming Jin, Yonghao Long 0001, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (4) | 3 |
| 2021 | E-DSSR: Efficient Dynamic Surgical Scene Reconstruction with Transformer-Based Stereoscopic Depth Perception
Yonghao Long 0001, Zhaoshuo Li, Chi Hang Yee, Chi-Fai Ng, Russell H. Taylor, Mathias Unberath, Qi Dou 0001 |
MICCAI (4) | 1 |
| 2021 | Temporal Memory Relation Network for Workflow Recognition From Surgical VideoabstractAutomatic surgical workflow recognition is a key component for developing context-aware computer-assisted systems in the operating theatre. Previous works either jointly modeled the spatial features with short fixed-range temporal information, or separately learned visual and long temporal cues. In this paper, we propose a novel end-to-end temporal memory relation network (TMRNet) for relating long-range and multi-scale temporal patterns to augment the present features. We establish a long-range memory bank to serve as a memory cell storing the rich supportive information. Through our designed temporal variation layer, the supportive cues are further enhanced by multi-scale temporal-only convolutions. To effectively incorporate the two types of cues without disturbing the joint learning of spatio-temporal features, we introduce a non-local bank operator to attentively relate the past to the present. In this regard, our TMRNet enables the current feature to view the long-range temporal dependency, as well as tolerate complex temporal extents. We have extensively validated our approach on two benchmark surgical video datasets, M2CAI challenge dataset and Cholec80 dataset. Experimental results demonstrate the outstanding performance of our method, consistently exceeding the state-of-the-art methods by a large margin (e.g., 67.0% v.s. 78.9% Jaccard on Cholec80 dataset). Yueming Jin, Yonghao Long 0001, Cheng Chen 0013, Qi Dou 0001, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 2 |