Jie Ying Wu

dblp:233/0069 · DBLP profile ↗
← Back
15ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0002-7306-8140ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Systems, architecture and hardware · 6 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Neural-Augmented Kelvinlet for Real-Time Soft Tissue Deformation Modeling
abstract
Accurate and efficient modeling of soft-tissue interactions is fundamental for advancing surgical simulation, surgical robotics, and model-based surgical automation. To achieve real-time latency, classical Finite Element Method (FEM) solvers are often replaced with neural approximations; however, naively training such models in a fully data-driven manner without incorporating physical priors frequently leads to poor generalization and physically implausible predictions. We present a novel physics-informed neural simulation framework that enables real-time prediction of soft-tissue deformations under complex single- and multi-grasper interactions. Our approach integrates Kelvinlet-based analytical priors with large-scale FEM data, capturing both linear and nonlinear tissue responses. This hybrid design improves predictive accuracy and physical plausibility across diverse neural architectures while maintaining the low-latency performance required for interactive applications. We validate our method on challenging surgical manipulation tasks involving standard laparoscopic grasping tools, demonstrating substantial improvements in deformation fidelity and temporal stability over existing baselines. These results establish Kelvinlet-augmented learning as a principled and computationally efficient paradigm for real-time, physics-aware soft-tissue simulation in surgical AI.
Ashkan Shahbazi, Kyvia Pereira, Jon S. Heiselman, Elaheh Akbari, Annie C. Benson, Sepehr Seifi, Garrison L. H. Johnston, Jie Ying Wu, Nabil Simaan, Michael I. Miga, Soheil Kolouri
AAAI9
2026 EndoPBR: Photorealistic Synthetic Data for Surgical 3D Vision via Physically-based Rendering
abstract
Synthetic data has played a pivotal role in developing large-scale 3D vision models due to its high-quality annotations and ease of curation. In domains where labeled data collection is difficult, such as endoscopy, synthetic data holds promise as a means to generate the large-scale annotated datasets required to train modern neural networks. In this work, we address a core question for data-scarce applications in 3D vision: how can we generate synthetic labeled data, and how useful would the data be for training downstream vision models? First, we introduce a novel data generation module that takes images with known geometry and camera poses as input and estimates the material and lighting conditions of the scene. To stabilize training, we leverage domain-specific properties like non-stationary lighting and anatomical material priors. We model the material properties as a bidirectional reflectance distribution function, parameterized by a neural network. Via the rendering equation, we can generate photorealistic images at arbitrary camera poses. We demonstrate that this method produces competitive novel view synthesis results compared to previous work while being more lightweight, flexible, and efficient. Second, we use our synthetic data to train models on various downstream 3D vision tasks and find that models trained solely on our synthetic data generally outperform those trained on real data across various metrics and tasks. Our experiments show that synthetic data is a promising avenue towards robust 3D vision in surgical scenes.1
John J. Han, Jie Ying Wu
WACV2
2025 From Monocular Vision to Autonomous Action: Guiding Tumor Resection via 3D Reconstruction
abstract
Surgical automation requires precise guidance and understanding of the scene. Current methods in the literature rely on bulky depth cameras to create maps of the anatomy; however, this does not translate well to space-limited clinical applications. Monocular cameras are small and allow minimally invasive surgeries in tight spaces, but additional processing is required to generate 3D scene understanding. We propose a 3D mapping pipeline that uses only RGB images to create segmented point clouds of the target anatomy. To ensure the most accurate reconstruction, we compare different structure from motion algorithms’ performance on mapping the central airway obstructions, and test the pipeline on a downstream task of tumor resection. In several metrics, including post-procedure percentage tissue charring, our pipeline performs comparably to RGB-D cameras and, in some cases, even surpasses their downstream task performance. These promising results demonstrate that automation guidance can be achieved in minimally invasive procedures with monocular cameras. This study is a step toward the complete autonomy of surgical robots.
Ayberk Acar, Mariana E. Smith, Lidia Al-Zogbi, Tanner Watts, Fangjie Li, Hao Li 0108, Nural Yilmaz, Paul Maria Scheikl, Jesse F. d'Almeida, Susheela Sharma, Lauren Branscombe, Tayfun Efe Ertop, Robert J. Webster III, Ipek Oguz, Alan Kuntz, Axel Krieger, Jie Ying Wu
IROS17
2025 NAVIUS: Navigated Augmented Reality Visualization for Ureteroscopic Surgery
Ayberk Acar, Jumanh Atoum, Peter S. Connor, Clifford Pierre, Carisa N. Lynch, Nicholas L. Kavoussi, Jie Ying Wu
MICCAI (11)7
2025 From Sight to Skill: A Surgeon-Centered Augmented Reality System for Ureteroscopy Training
Jumanh Atoum, Fangjie Li, Ayberk Acar, Nicholas L. Kavoussi, Jie Ying Wu
MICCAI (11)5
2025 Augmented Reality-Based Guidance with Deformable Registration in Head and Neck Tumor Resection
Qingyun Yang, Fangjie Li, Sindhura Sridhar, Whitney Jin, Jennifer Du, Jon S. Heiselman, Michael I. Miga, Michael Topf, Jie Ying Wu
MICCAI (9)11
2024 A Robotic Mediation Device for Skill Assessment and Training During Colonoscopy
abstract
Colonoscopy demands multi-finger coordinated motion to achieve safe navigation. As a result, training for colonoscopists is challenging and skill assessment currently relies on subjective scoring by expert proctors. There is a need to provide tools for skill assessment and aid with training new interventionalists. This paper presents a new concept of an in-hand robotic mediation device that can be used for both skill assessment and training. The robotic device can be used to infer the kinematic motion as well as the power input of a user - both of which are proposed to be used for skill assessment and subtask skill classification. Preliminary results collected expert and novice users performing colonoscopy navigation are used to demonstrate this device as a skill assessment tool. A machine-learning model (classification and regression trees) is used for subtask classification of skill and evaluating the most important classification features. A user study demonstrates the effectiveness of this in hand haptic training and assessment tool. We believe that, in the future, this device will enable accelerated skill assessment and training and possible semi-automation of difficult maneuvers.
Olivia K. Richards, Elan Z. Ahronovich, Neel Shihora, Ahmet Yildiz, Jumana Atoum, Jie Ying Wu, Keith Obstein, Nabil Simaan
IROS6
2024 A Hybrid Model and Learning-Based Force Estimation Framework for Surgical Robots
abstract
Haptic feedback to the surgeon during robotic surgery would enable safer and more immersive surgeries but estimating tissue interaction forces at the tips of robotically controlled surgical instruments has proven challenging. Few existing surgical robots can measure interaction forces directly and the additional sensor may limit the life of instruments. We present a hybrid model and learning-based framework for force estimation for the Patient Side Manipulators (PSM) of a da Vinci Research Kit (dVRK). The model-based component identifies the dynamic parameters of the robot and estimates free-space joint torque, while the learning-based component compensates for environmental factors, such as the additional torque caused by trocar interaction between the PSM instrument and the patient’s body wall. We evaluate our method in an abdominal phantom and achieve an error in force estimation of under 10% normalized root-mean-squared error. We show that by using a model-based method to perform dynamics identification, we reduce reliance on the training data covering the entire workspace. Although originally developed for the dVRK, the proposed method is a generalizable framework for other compliant surgical robots. The code is available at https://github.com/vu-maple-lab/dvrk_force_estimation.
Hao Yang 0011, Haoying Zhou, Gregory S. Fischer, Jie Ying Wu
IROS4
2024 MeshBrush: Painting the Anatomical Mesh with Neural Stylization for Endoscopy
John J. Han, Ayberk Acar, Nicholas L. Kavoussi, Jie Ying Wu
MICCAI (6)4
2022 CaRTS: Causality-Driven Robot Tool Segmentation from Vision and Kinematics Data
Hao Ding 0021, Jintan Zhang, Peter Kazanzides, Jie Ying Wu, Mathias Unberath
MICCAI (8)4
2021 Relational Graph Learning on Visual and Kinematics Embeddings for Accurate Gesture Recognition in Robotic Surgery
abstract
Automatic surgical gesture recognition is fundamentally important to enable intelligent cognitive assistance in robotic surgery. With recent advancement in robot-assisted minimally invasive surgery, rich information including surgical videos and robotic kinematics can be recorded, which provide complementary knowledge for understanding surgical gestures. However, existing methods either solely adopt uni-modal data or directly concatenate multi-modal representations, which can not sufficiently exploit the informative correlations inherent in visual and kinematics data to boost gesture recognition accuracies. In this regard, we propose a novel online approach of multi-modal relational graph network (i.e., MRG-Net) to dynamically integrate visual and kinematics information through interactive message propagation in the latent feature space. In specific, we first extract embeddings from video and kinematics sequences with temporal convolutional networks and LSTM units. Next, we identify multi-relations in these multi-modal embeddings and leverage them through a hierarchical relational graph learning module. The effectiveness of our method is demonstrated with state-of-the-art results on the public JIGSAWS dataset, outperforming current uni-modal and multi-modal methods on both suturing and knot typing tasks. Furthermore, we validated our method on in-house visual-kinematics datasets collected with da Vinci Research Kit (dVRK) platforms in two centers, with consistent promising performance achieved. Our code and data are released at: https://www.cse.cuhk.edu.hk/~yhlong/mrgnet.html.
Yonghao Long 0001, Jie Ying Wu, Bo Lu 0001, Yueming Jin, Mathias Unberath, Yun-Hui Liu 0001, Pheng-Ann Heng, Qi Dou 0001
ICRA2
2021 An Interpretable Approach to Automated Severity Scoring in Pelvic Trauma
Anna Zapaishchykova, David Dreizin, Zhaoshuo Li, Jie Ying Wu, Shahrooz Faghih Roohi, Mathias Unberath
MICCAI (3)4
2020 Neural Network based Inverse Dynamics Identification and External Force Estimation on the da Vinci Research Kit
abstract
Most current surgical robotic systems lack the ability to sense tool/tissue interaction forces, which motivates research in methods to estimate these forces from other available measurements, primarily joint torques. These methods require the internal joint torques, due to the robot inverse dynamics, to be subtracted from the measured joint torques. This paper presents the use of neural networks to estimate the inverse dynamics of the da Vinci surgical robot, which enables estimation of the external environment forces. Experiments with motions in free space demonstrate that the neural networks can estimate the internal joint torques within 10% normalized rootmean-square error (NRMSE), which outperforms model-based approaches in the literature. Comparison with an external force sensor shows that the method is able to estimate environment forces within about 10% NRMSE.
Nural Yilmaz, Jie Ying Wu, Peter Kazanzides, Ugur Tümerdem
ICRA2
2019 LumiPath - Towards Real-Time Physically-Based Rendering on Embedded Devices
Laura Fink, Sing Chun Lee, Jie Ying Wu, Xingtong Liu, Tianyu Song 0002, Yordanka Velikova, Marc Stamminger, Nassir Navab, Mathias Unberath
MICCAI (5)3
2018 FPGA-Based Velocity Estimation for Control of Robots with Low-Resolution Encoders
abstract
Robot control algorithms often rely on measurements of robot joint velocities, which can be estimated by measuring the time between encoder edges. When encoder edges occur infrequently, such as at low velocities and/or with low resolution encoders, this measurement delay may affect the stability of closed-loop control. This is evident in both the joint position control and Cartesian impedance control of the da Vinci Research Kit (dVRK), which contains several low-resolution encoders. We present a hardware-based method that gives more frequent velocity updates and is not affected by common encoder imperfections such as non-uniform duty cycles and quadrature phase error. The proposed method measures the time between consecutive edges of the same type but, unlike prior methods, is implemented for the rising and falling edges of both channels. Additionally, it estimates acceleration to enable software compensation of the measurement delay. The method is shown to improve Cartesian impedance control of the dVRK.
Jie Ying Wu, Zihan Chen 0004, Anton Deguet, Peter Kazanzides
IROS1