EDBT 2026 Demo / reviewers in the wild / expert
Sophia Bano
dblp:159/3825
· DBLP profile ↗
21ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0003-1329-4565ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SurgflowNet: Leveraging unannotated video for consistent endoscopic pituitary surgery workflow recognitionabstract-score and 13.4% in Edit Score over the SOTA, SurgflowNetdemonstrates a significant improvement in workflow recognition for endoscopic pituitary surgery. Anjana Wijekoon, Adrito Das, Zhehua Mao, Danyal Z. Khan, John G. Hanrahan, Danail Stoyanov, Hani J. Marcus, Sophia Bano |
Artif. Intell. Medicine | 8 |
| 2026 | TDN-PPO: An Automatic Control Framework of a Surgical Robot for Posterior Segment Ophthalmic SurgeryabstractIn ophthalmic surgery, particularly the procedures involving the posterior segment, clinicians face significant challenges in maintaining precise control of handheld instruments to peel fragile membranes without damaging the surrounding healthy fundus tissue, even for seasoned clinicians employing specialised ophthalmic surgical robots. The implementation of autonomous control in robot-assisted surgical systems holds promise for overcoming these obstacles and simplifying intricate surgical tasks. This paper introduces an autonomous control framework, integrating a Tip Detection Network (TDN) with a Proximal Policy Optimization (PPO) network, designed to autonomously navigate the Tip of the Surgical Instrument (ToSI) towards the intended lesion site in a real-world scenario. Results indicate that the accuracy of the TDN module in detecting the ToSI position in images of varying sizes can be reliably maintained within a 4.6-pixel range. The autonomous control deviation for the PPO module ranges between [0.6585μm, 7.995μm], with an average discrepancy of 5.118μm. The communication frequency across the modules is maintained at 35.7 Hz. In physical environments, the TDN-PPO framework adeptly navigates the ToSI to autonomously and precisely converge on the preset Target Lesion (TL) macular hole, maintaining a tip-to-target distance error within a margin of 2.094 pixels (38μm). Throughout the autonomous navigation phase, the maximal contact force between the ToSI and the TL is capped at 41.1 mN, aligning with the upper threshold for contact force between the ToSI and tissue prescribed in clinical surgical settings. Ning Wang 0042, Sophia Bano, Danail Stoyanov, Ziting Liang, Agostino Stilli |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | PitVis-2023 challenge: Workflow recognition in videos of endoscopic pituitary surgeryabstractThe field of computer vision applied to videos of minimally invasive surgery is ever-growing. Workflow recognition pertains to the automated recognition of various aspects of a surgery, including: which surgical steps are performed; and which surgical instruments are used. This information can later be used to assist clinicians when learning the surgery or during live surgery. The Pituitary Vision (PitVis) 2023 Challenge tasks the community to step and instrument recognition in videos of endoscopic pituitary surgery. This is a particularly challenging task when compared to other minimally invasive surgeries due to: the smaller working space, which limits and distorts vision; and higher frequency of instrument and step switching, which requires more precise model predictions. Participants were provided with 25-videos, with results presented at the MICCAI-2023 conference as part of the Endoscopic Vision 2023 Challenge in Vancouver, Canada, on 08-Oct-2023. There were 18-submissions from 9-teams across 6-countries, using a variety of deep learning models. The top performing model for step recognition utilised a transformer based architecture, uniquely using an autoregressive decoder with a positional encoding input. The top performing model for instrument recognition utilised a spatial encoder followed by a temporal encoder, which uniquely used a 2-layer temporal architecture. In both cases, these models outperformed purely spatial based models, illustrating the importance of sequential and temporal information. This PitVis-2023 therefore demonstrates state-of-the-art computer vision models in minimally invasive surgery are transferable to a new dataset. Benchmark results are provided in the paper, and the dataset is publicly available at: https://doi.org/10.5522/04/26531686. Adrito Das, Danyal Z. Khan, Dimitris Psychogyios, John G. Hanrahan, Francisco Vasconcelos 0001, You Pang, Zhen Chen 0018, Jinlin Wu, Xiaoyang Zou, Guoyan Zheng, Abdul Qayyum 0002, Moona Mazher, Muhammad Imran Razzak, Tianbin Li, Jin Ye 0002, Junjun He, Szymon Plotka, Joanna Kaleta, Amine Yamlahi, Antoine Jund, Patrick Godau, Satoshi Kondo, Satoshi Kasai, Kousuke Hirasawa, Dominik Rivoir, Stefanie Speidel, Alejandra Pérez, Santiago Rodríguez, Pablo Andrés Arbeláez, Danail Stoyanov, Hani J. Marcus, Sophia Bano |
Medical Image Anal. | 33 |
| 2025 | Vision-Based Automatic Control of a Surgical Robot for Posterior Segment Ophthalmic SurgeryabstractIn ophthalmic surgery, especially in posterior segment procedures, clinicians face significant challenges, like the inherent tremor of the surgeon’s arm, restricted visibility, and heavy reliance on the surgeon’s skills for precise control of hand-held tools during micro-surgical movements. Automatic control of robotic-assisted ophthalmic surgical systems has the potential to overcome these challenges, simplifying complex surgical procedures. This paper proposes a novel image-guided automatic control method for an Ophthalmic micro-Surgical Robot (OmSR), specifically designed for posterior segment eye surgery. The method relies on forceps shadow tracking. The paper introduces a tip detection network (Net-SR), which accurately calculates the coordinates of the Tips of Surgical Forceps (ToSF) and Tips of Shadow (ToS) to enable automatic navigation. Additionally, through the Non-Uniform Rational B-Spline (NURBS) curve interpolation and speed look-ahead algorithm, dense and time-continuous data points are obtained to improve control accuracy and smoothness. The accuracy of the Net-SR network and motion of the ToSF, and the effectiveness of the proposed automatic controller are experimentally evaluated. Results demonstrate a significant 98.21% improvement in the Net-SR network accuracy over the normal keypoint detection network. The use of the speed look-ahead algorithm leads to a notable 41.7% improvement in optimal speed, and the ToSF successfully reaches the target lesion with vision-based navigation and no overscale motion. Note to Practitioners—The practical problem that motivated this research is the need for safer and more efficient surgical procedures, focusing on minimizing the risk of fundus tissue damage associated with intraoperative surgical instruments. To overcome challenges related to handheld and tele-operated control, we explore automatic control as a promising solution. In this paper, the tip of the instrument can consistently and accurately reach the target lesion with high precision and no overscale motion, allowing for deskilling of complex and repetitive tasks. This capability holds potential for the clinical needle insertion operation and membrane peeling operation. The proposed control methods can also be extended to other surgical procedures. Ning Wang 0042, Sophia Bano, Danail Stoyanov, Agostino Stilli |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2024 | Gaussian Pancakes: Geometrically-Regularized 3D Gaussian Splatting for Realistic Endoscopic Reconstruction
Sierra Bonilla, Shuai Zhang 0029, Dimitris Psychogyios, Danail Stoyanov, Francisco Vasconcelos 0001, Sophia Bano |
MICCAI (6) | 6 |
| 2024 | PitVQA: Image-Grounded Text Embedding LLM for Visual Question Answering in Pituitary Surgery
Runlong He, Mengya Xu, Adrito Das, Danyal Z. Khan, Sophia Bano, Hani J. Marcus, Danail Stoyanov, Matthew J. Clarkson, Mobarakol Islam |
MICCAI (6) | 5 |
| 2024 | Placental vessel segmentation and registration in fetoscopy: Literature review and MICCAI FetReg2021 challenge findingsabstractFetoscopy laser photocoagulation is a widely adopted procedure for treating Twin-to-Twin Transfusion Syndrome (TTTS). The procedure involves photocoagulation pathological anastomoses to restore a physiological blood exchange among twins. The procedure is particularly challenging, from the surgeon's side, due to the limited field of view, poor manoeuvrability of the fetoscope, poor visibility due to amniotic fluid turbidity, and variability in illumination. These challenges may lead to increased surgery time and incomplete ablation of pathological anastomoses, resulting in persistent TTTS. Computer-assisted intervention (CAI) can provide TTTS surgeons with decision support and context awareness by identifying key structures in the scene and expanding the fetoscopic field of view through video mosaicking. Research in this domain has been hampered by the lack of high-quality data to design, develop and test CAI algorithms. Through the Fetoscopic Placental Vessel Segmentation and Registration (FetReg2021) challenge, which was organized as part of the MICCAI2021 Endoscopic Vision (EndoVis) challenge, we released the first large-scale multi-center TTTS dataset for the development of generalized and robust semantic segmentation and video mosaicking algorithms with a focus on creating drift-free mosaics from long duration fetoscopy videos. For this challenge, we released a dataset of 2060 images, pixel-annotated for vessels, tool, fetus and background classes, from 18 in-vivo TTTS fetoscopy procedures and 18 short video clips of an average length of 411 frames for developing placental scene segmentation and frame registration for mosaicking techniques. Seven teams participated in this challenge and their model performance was assessed on an unseen test dataset of 658 pixel-annotated images from 6 fetoscopic procedures and 6 short clips. For the segmentation task, overall baseline performed was the top performing (aggregated mIoU of 0.6763) and was the best on the vessel class (mIoU of 0.5817) while team RREB was the best on the tool (mIoU of 0.6335) and fetus (mIoU of 0.5178) classes. For the registration task, overall the baseline performed better than team SANO with an overall mean 5-frame SSIM of 0.9348. Qualitatively, it was observed that team SANO performed better in planar scenarios, while baseline was better in non-planner scenarios. The detailed analysis showed that no single team outperformed on all 6 test fetoscopic videos. The challenge provided an opportunity to create generalized solutions for fetoscopic scene understanding and mosaicking. In this paper, we present the findings of the FetReg2021 challenge, alongside reporting a detailed literature review for CAI in TTTS fetoscopy. Through this challenge, its analysis and the release of multi-center fetoscopic data, we provide a benchmark for future research in this field. Sophia Bano, Alessandro Casella, Francisco Vasconcelos 0001, Abdul Qayyum 0002, Abdessalam Benzinou, Moona Mazher, Fabrice Mériaudeau, Chiara Lena, Ilaria A. Cintorrino, Gaia Romana De Paolis, Jessica Biagioli, Daria Grechishnikova, Jing Jiao, Bizhe Bai, Yanyan Qiao, Binod Bhattarai, Rebati Raman Gaire, Ronast Subedi, Eduard Vazquez, Szymon Plotka, Aneta Lisowska, Arkadiusz Sitek, George Attilakos, Ruwan Wimalasundera, Anna L. David, Dario Paladini, Jan Deprest, Elena De Momi, Leonardo S. Mattos, Sara Moccia, Danail Stoyanov |
Medical Image Anal. | 1 |
| 2024 | SurgT challenge: Benchmark of soft-tissue trackers for robotic surgery
João Cartucho, Alistair Weld, Samyakh Tukra, Haozheng Xu, Hiroki Matsuzaki, Taiyo Ishikawa, Minjun Kwon, Yongeun Jang, Kwang-Ju Kim, Gwang Lee, Bizhe Bai, Lüder A. Kahrs, Lars Boecking, Simeon Allmendinger, Leopold Müller, Yueming Jin, Sophia Bano, Francisco Vasconcelos 0001, Wolfgang Reiter, Jonas Hajek, Estevão Lima, João L. Vilaça, Sandro F. Queiros, Stamatia Giannarou |
Medical Image Anal. | 18 |
| 2024 | SimCol3D - 3D reconstruction during colonoscopy challengeabstractColorectal cancer is one of the most common cancers in the world. While colonoscopy is an effective screening technique, navigating an endoscope through the colon to detect polyps is challenging. A 3D map of the observed surfaces could enhance the identification of unscreened colon tissue and serve as a training platform. However, reconstructing the colon from video footage remains difficult. Learning-based approaches hold promise as robust alternatives, but necessitate extensive datasets. Establishing a benchmark dataset, the 2022 EndoVis sub-challenge SimCol3D aimed to facilitate data-driven depth and pose prediction during colonoscopy. The challenge was hosted as part of MICCAI 2022 in Singapore. Six teams from around the world and representatives from academia and industry participated in the three sub-challenges: synthetic depth prediction, synthetic pose prediction, and real pose prediction. This paper describes the challenge, the submitted methods, and their results. We show that depth prediction from synthetic colonoscopy images is robustly solvable, while pose estimation remains an open research question. Anita Rau, Sophia Bano, Yueming Jin, Pablo Azagra, Javier Morlana, Rawen Kader, Edward Sanderson, Bogdan J. Matuszewski, Erez Posner, Netanel Frank, Varshini Elangovan, Sista Raviteja, Zhengwen Li, Jiquan Liu, Seenivasan Lalithkumar, Mobarakol Islam, Hongliang Ren 0001, Laurence B. Lovat, J. M. M. Montiel, Danail Stoyanov |
Medical Image Anal. | 2 |
| 2024 | MonoLoT: Self-Supervised Monocular Depth Estimation in Low-Texture Scenes for Automatic Robotic EndoscopyabstractThe self-supervised monocular depth estimation framework is well-suited for medical images that lack ground-truth depth, such as those from digestive endoscopes, facilitating navigation and 3D reconstruction in the gastrointestinal tract. However, this framework faces several limitations, including poor performance in low-texture environments, limited generalisation to real-world datasets, and unclear applicability in downstream tasks like visual servoing. To tackle these challenges, we propose MonoLoT, a self-supervised monocular depth estimation framework featuring two key innovations: point matching loss and batch image shuffle. Extensive ablation studies on two publicly available datasets, namely C3VD and SimCol, have shown that methods enabled by MonoLoT achieve substantial improvements, with accuracies of 0.944 on C3VD and 0.959 on SimCol, surpassing both depth-supervised and self-supervised baselines on C3VD. Qualitative evaluations on real-world endoscopic data underscore the generalisation capabilities of our methods, outperforming both depth-supervised and self-supervised baselines. To demonstrate the feasibility of using monocular depth estimation for visual servoing, we have successfully integrated our method into a proof-of-concept robotic platform, enabling real-time automatic intervention and control in digestive endoscopy. In summary, our method represents a significant advancement in monocular depth estimation for digestive endoscopy, overcoming key challenges and opening promising avenues for medical applications. Qi He 0008, Sophia Bano, Danail Stoyanov, Siyang Zuo |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Why is the Winner the Best?abstractInternational benchmarking competitions have become fundamental for the comparative performance assessment of image analysis methods. However, little attention has been given to investigating what can be learnt from these competitions. Do they really generate scientific progress? What are common and successful participation strategies? What makes a solution superior to a competing method? To address this gap in the literature, we performed a multicenter study with all 80 competitions that were conducted in the scope of IEEE ISBI 2021 and MICCAI 2021. Statistical analyses performed based on comprehensive descriptions of the submitted algorithms linked to their rank as well as the underlying participation strategies revealed common characteristics of winning solutions. These typically include the use of multi-task learning (63%) and/or multi-stage pipelines (61%), and a focus on augmentation (100%), image preprocessing (97%), data curation (79%), and post-processing (66%). The “typical” lead of a winning team is a computer scientist with a doctoral degree, five years of experience in biomedical image analysis, and four years of experience in deep learning. Two core general development strategies stood out for highly-ranked teams: the reflection of the metrics in the method design and the focus on analyzing and handling failure cases. According to the organizers, 43% of the winning algorithms exceeded the state of the art but only 11% completely solved the respective domain problem. The insights of our study could help researchers (1) improve algorithm development strategies when approaching new problems, and (2) focus on open research questions revealed by this work. Matthias Eisenmann, Annika Reinke, Vivienn Weru, Minu Tizabi, Fabian Isensee, Tim Adler, Sharib Ali, Vincent Andrearczyk, Marc Aubreville, Ujjwal Baid, Spyridon Bakas, Niranjan Balu, Sophia Bano, Jorge Bernal, Sebastian Bodenstedt, Alessandro Casella, Veronika Cheplygina, Marie Daum, Marleen de Bruijne, Adrien Depeursinge, Reuben Dorent, Jan Egger, David Gage Ellis, Sandy Engelhardt, Melanie Ganz-Benjaminsen, Noha M. Ghatwary, Gabriel Girard, Patrick Godau, Anubha Gupta, Lasse Hansen, Kanako Harada, Mattias P. Heinrich, Nicholas Heller, Alessa Hering, Arnaud Huaulmé, Pierre Jannin, A. Emre Kavur, Oldrich Kodym, Michal Kozubek 0001, Jianning Li 0002, Hongwei Li 0004, Jun Ma 0016, Carlos Martín-Isla, Bjoern Menze, J. Alison Noble, Valentin Oreiller, Nicolas Padoy, Sarthak Pati, Kelly Payette, Tim Rädsch, Jonathan Rafael-Patino, Vivek Singh Bawa, Stefanie Speidel, Carole H. Sudre, Kimberlin M. H. van Wijnen, Martin Wagner 0001, D. Wei, Amine Yamlahi, Moi Hoon Yap, C. Yuan, Maximilian Zenk, A. Zia, David Zimmerer, Dogu Baran Aydogan, Binod Bhattarai, Louise Bloch, Raphael Brüngel, J. Cho, C. Choi, Qi Dou 0001, Ivan Ezhov, Christoph M. Friedrich, C. Fuller, Rebati Raman Gaire, Adrian Galdran, Álvaro García-Faura, Maria Grammatikopoulou, S. Hong, Mostafa Jahanifar, I. Jang, Abdolrahim Kadkhodamohammadi, I. Kang, Florian Kofler, S. Kondo, Hugo J. Kuijf, M. Luu, Tomaz Martincic, Pedro Morais, Mohamed A. Naser, Bruno Oliveira 0002, David Owen 0001, S. Pang, Szymon Plotka, Élodie Puybareau, Nasir M. Rajpoot, K. Ryu, Numan Saeed, Adam J. Shephard, Dejan Stepec, Ronast Subedi, Guillaume Tochon, Helena R. Torres, Hélène Urien, João L. Vilaça, Kareem A. Wahid, Benedikt Wiestler, Marek Wodzinski, F. Xia, J. Xie, Z. Xiong, Sen Yang 0006, Klaus H. Maier-Hein, Paul F. Jaeger, Annette Kopp-Schneider, Lena Maier-Hein |
CVPR | 13 |
| 2023 | A Multi-task Network for Anatomy Identification in Endoscopic Pituitary Surgery
Adrito Das, Danyal Z. Khan, Simon C. Williams, John G. Hanrahan, Anouk Borg, Neil L. Dorward, Sophia Bano, Hani J. Marcus, Danail Stoyanov |
MICCAI (9) | 7 |
| 2022 | BiometryNet: Landmark-based Fetal Biometry Estimation from Standard Ultrasound Planes
Netanell Avisdris, Leo Joskowicz, Brian Dromey, Anna L. David, Donald Peebles, Danail Stoyanov, Dafna Ben-Bashat, Sophia Bano |
MICCAI (4) | 8 |
| 2022 | Retrieval of Surgical Phase Transitions Using Reinforcement Learning
Sophia Bano, Ann-Sophie Page, Jan Deprest, Danail Stoyanov, Francisco Vasconcelos 0001 |
MICCAI (8) | 2 |
| 2021 | AutoFB: Automating Fetal Biometry Estimation from Standard Ultrasound Planes
Sophia Bano, Brian Dromey, Francisco Vasconcelos 0001, Raffaele Napolitano, Anna L. David, Donald Peebles, Danail Stoyanov |
MICCAI (7) | 1 |
| 2021 | Deep learning for detection and segmentation of artefact and disease instances in gastrointestinal endoscopyabstractThe Endoscopy Computer Vision Challenge (EndoCV) is a crowd-sourcing initiative to address eminent problems in developing reliable computer aided detection and diagnosis endoscopy systems and suggest a pathway for clinical translation of technologies. Whilst endoscopy is a widely used diagnostic and treatment tool for hollow-organs, there are several core challenges often faced by endoscopists, mainly: 1) presence of multi-class artefacts that hinder their visual interpretation, and 2) difficulty in identifying subtle precancerous precursors and cancer abnormalities. Artefacts often affect the robustness of deep learning methods applied to the gastrointestinal tract organs as they can be confused with tissue of interest. EndoCV2020 challenges are designed to address research questions in these remits. In this paper, we present a summary of methods developed by the top 17 teams and provide an objective comparison of state-of-the-art methods and methods designed by the participants for two sub-challenges: i) artefact detection and segmentation (EAD2020), and ii) disease detection and segmentation (EDD2020). Multi-center, multi-organ, multi-class, and multi-modal clinical endoscopy datasets were compiled for both EAD2020 and EDD2020 sub-challenges. The out-of-sample generalization ability of detection algorithms was also evaluated. Whilst most teams focused on accuracy improvements, only a few methods hold credibility for clinical usability. The best performing teams provided solutions to tackle class imbalance, and variabilities in size, origin, modality and occurrences by exploring data augmentation, data fusion, and optimal class thresholding techniques. Sharib Ali, Mariia Dmitrieva, Noha M. Ghatwary, Sophia Bano, Gorkem Polat, Alptekin Temizel, Adrian Krenzer, Amar Hekalo, Bogdan J. Matuszewski, Mourad Gridach, Irina Voiculescu, Vishnusai Yoganand, Arnav Chavan, Aryan Raj, Nhan T. Nguyen, Dat Q. Tran, Lê Duy Huynh, Nicolas Boutry, Shahadate Rezvy, Haijian Chen, Yoon Ho Choi, Anand Subramanian 0004, Velmurugan Balasubramanian, Xiaohong W. Gao, Hongyu Hu, Yusheng Liao, Danail Stoyanov, Christian Daul, Stefano Realdon, Renato Cannizzaro, Dominique Lamarque, Terry Tran-Nguyen, Adam Bailey, Barbara Braden, James E. East, Jens Rittscher |
Medical Image Anal. | 4 |
| 2020 | Deep Placental Vessel Segmentation for Fetoscopic Mosaicking
Sophia Bano, Francisco Vasconcelos 0001, Luke M. Shepherd, Emmanuel B. Vander Poorten, Tom Vercauteren, Sébastien Ourselin, Anna L. David, Jan Deprest, Danail Stoyanov |
MICCAI (3) | 1 |
| 2019 | Deep Sequential Mosaicking of Fetoscopic Videos
Sophia Bano, Francisco Vasconcelos 0001, Marcel Tella-Amo, George Dwyer, Caspar Gruijthuijsen, Jan Deprest, Sébastien Ourselin, Emmanuel B. Vander Poorten, Tom Vercauteren, Danail Stoyanov |
MICCAI (1) | 1 |
| 2016 | ViComp: composition of user-generated videosabstractWe propose ViComp, an automatic audio-visual camera selection framework for composing uninterrupted recordings from multiple user-generated videos (UGVs) of the same event. We design an automatic audio-based cut-point selection method to segment the UGV. ViComp combines segments of UGVs using a rank-based camera selection strategy by considering audio-visual quality and camera selection history. We analyze the audio to maintain audio continuity. To filter video segments which contain visual degradations, we perform spatial and spatio-temporal quality assessment. We validate the proposed framework with subjective tests and compare it with state-of-the-art methods. Sophia Bano, Andrea Cavallaro |
Multim. Tools Appl. | 1 |
| 2015 | Gyro-based Camera-motion Detection in User-generated VideosabstractWe propose a gyro-based camera-motion detection method for videos captured with smartphones. First, the delay between the acquisition of video and gyroscope data is estimated using similarities induced by camera motion in the two sensor modalities. Pan, tilt and shake are then detected using the dominant motions and high frequencies in the gyroscope data. Morphological operations are applied to remove outliers and to identify segments with continuous camera-motion. We compare the proposed method with existing methods that use visual or inertial sensor data. Sophia Bano, Andrea Cavallaro, Xavier Parra Llanas |
ACM Multimedia | 1 |
| 2015 | Discovery and organization of multi-camera user-generated videos of the same event
Sophia Bano, Andrea Cavallaro |
Inf. Sci. | 1 |