EDBT 2026 Demo / reviewers in the wild / expert
Sharib Ali
dblp:14/11430
· DBLP profile ↗
27ranked-venue papers
6as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 3 first-author · 20 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Nested resolution mesh-graph CNN for automated extraction of liver surface anatomical landmarksabstractThe anatomical landmarks on the liver (mesh) surface, including the falciform ligament and liver ridge, are composed of triangular meshes of varying shapes, sizes, and positions, making them highly complex. Extracting and segmenting these landmarks is critical for augmented reality-based intraoperative navigation and monitoring. The key to this task lies in comprehensively understanding the overall geometric shape and local topological information of the liver mesh. However, due to the liver's variations in shape and appearance, coupled with limited data, deep learning methods often struggle with automatic liver landmark segmentation. To address this, we propose a two-stage automatic framework combining mesh-CNN and graph-CNN. In the first stage, dynamic graph convolution (DGCNN) is employed on low-resolution meshes to achieve rapid global understanding, generating initial landmark proposals at two levels, "dilation" and "erosion", and mapping them onto the original high-resolution surface. Subsequently, a refinement network based on mesh convolution fuses these landmark proposals from edge features along the local topology of the high-resolution mesh surface, producing refined segmentation results. Additionally, we incorporate an anatomy-aware Dice loss to address resolution imbalance and better handle sparse anatomical regions. Extensive experiments on two liver datasets, both in-distribution and out-of-distribution, demonstrate that our method accurately processes liver meshes of different resolutions, outperforming state-of-the-art methods. The reconstructed liver mesh dataset and the source code are available at https://github.com/xukun-zhang/MeshGraphCNN. Xukun Zhang, Jinghui Feng, Peng Liu 0074, Minghao Han, Yanlan Kang, Sharib Ali, Lihua Zhang 0002 |
Medical Image Anal. | 9 |
| 2026 | Self-Supervised Monocular Depth and Pose Estimation for Endoscopy With Latent PriorsabstractAccurate 3D reconstruction in endoscopy enables quantitative and holistic lesion characterization within the gastrointestinal (GI) tract. To achieve this, reliable depth and pose estimation is required. However, endoscopy systems are monocular, and existing methods relying on synthetic datasets or complex models often lack generalizability in challenging endoscopic conditions. We propose a robust self-supervised monocular depth and pose estimation framework that incorporates a StyleGAN-based generator and a Variational Autoencoder (VAE). The StyleGAN generator leverages extensive depth scenes from natural images to condition the depth network, enhancing realism and robustness of depth predictions through latent feature priors. For pose estimation, we reformulate it within a VAE framework, treating pose transitions as latent variables to regularize scale, stabilize z-axis prominence, and improve x-y sensitivity. To further enhance pose stability and generalizability, we introduce a prior transfer module that distills motion knowledge from natural scene SLAM systems. Specifically, pose priors from a pretrained SLAM model-supervised on large-scale natural scene datasets-are used to guide the latent distribution of pose through a KL-divergence reparameterization. This mechanism effectively transfers structural motion priors into the endoscopic domain, improving trajectory consistency under challenging conditions. This dual refinement pipeline enables accurate depth and pose predictions, effectively addressing the GI tract's complex textures and lighting. Extensive evaluations on SimCol, C3VD, and EndoSLAM datasets confirm our framework's superior performance over published self-supervised methods in endoscopic depth and pose estimation. All data descriptions and code are available at https://github.com/EricXuziang/Self-supervised-with-Latent-Priors.git. Ziang Xu 0001, James East, Sharib Ali, Jens Rittscher |
IEEE Trans. Medical Imaging | 6 |
| 2026 | NeeCo: Image Synthesis of Novel Instrument States Based on Dynamic and Deformable 3-D Gaussian ReconstructionabstractComputer vision-based technologies significantly enhance surgical automation by advancing tool tracking, detection, and localization. However, Current data-driven approaches are data-voracious, requiring large, high-quality labeled image datasets. Our Work introduces a novel dynamic Gaussian Splatting technique to address the data scarcity in surgical image datasets. We propose a dynamic Gaussian model to represent dynamic surgical scenes, enabling the rendering of surgical instruments from unseen viewpoints and deformations with real tissue backgrounds. We utilize a dynamic training adjustment strategy to address challenges posed by poorly calibrated camera poses from real-world scenarios. Additionally, automatically generate annotations for our synthetic data. For evaluation, we constructed a new dataset featuring seven scenes with 14,000 frames of tool and camera motion and tool jaw articulation, with a background of an ex-vivo porcine model. Using this dataset, we synthetically replicate the scene deformation from the ground truth data, allowing direct comparisons of synthetic image quality. Experimental results illustrate that our method generates photo-realistic labeled image datasets with the highest PSNR (29.87). We further evaluate the performance of medical-specific neural networks trained on real and synthetic images using an unseen real-world image dataset. Our results show that the performance of models trained on synthetic images generated by the proposed method outperforms those trained with state-of-the-art standard data augmentation by 10%, leading to an overall improvement in model performances by nearly 15%. Tianle Zeng, Junlei Hu, Gerardo Loza Galindo, Sharib Ali, Duygu Sarikaya, Pietro Valdastri, Dominic Jones |
IEEE Trans. Medical Imaging | 4 |
| 2025 | ESPNet: Edge-Aware Feature Shrinkage Pyramid for Polyp Segmentation
Raneem Toman, Venkataraman Subramanian, Sharib Ali |
MICCAI (11) | 3 |
| 2025 | An objective comparison of methods for augmented reality in laparoscopic liver resection by preoperative-to-intraoperative image fusion from the MICCAI2022 challengeabstractAugmented reality for laparoscopic liver resection is a visualisation mode that allows a surgeon to localise tumours and vessels embedded within the liver by projecting them on top of a laparoscopic image. Preoperative 3D models extracted from Computed Tomography (CT) or Magnetic Resonance (MR) imaging data are registered to the intraoperative laparoscopic images during this process. Regarding 3D-2D fusion, most algorithms use anatomical landmarks to guide registration, such as the liver's inferior ridge, the falciform ligament, and the occluding contours. These are usually marked by hand in both the laparoscopic image and the 3D model, which is time-consuming and prone to error. Therefore, there is a need to automate this process so that augmented reality can be used effectively in the operating room. We present the Preoperative-to-Intraoperative Laparoscopic Fusion challenge (P2ILF), held during the Medical Image Computing and Computer Assisted Intervention (MICCAI 2022) conference, which investigates the possibilities of detecting these landmarks automatically and using them in registration. The challenge was divided into two tasks: (1) A 2D and 3D landmark segmentation task and (2) a 3D-2D registration task. The teams were provided with training data consisting of 167 laparoscopic images and 9 preoperative 3D models from 9 patients, with the corresponding 2D and 3D landmark annotations. A total of 6 teams from 4 countries participated in the challenge, whose results were assessed for each task independently. All the teams proposed deep learning-based methods for the 2D and 3D landmark segmentation tasks and differentiable rendering-based methods for the registration task. The proposed methods were evaluated on 16 test images and 2 preoperative 3D models from 2 patients. In Task 1, the teams were able to segment most of the 2D landmarks, while the 3D landmarks showed to be more challenging to segment. In Task 2, only one team obtained acceptable qualitative and quantitative registration results. Based on the experimental outcomes, we propose three key hypotheses that determine current limitations and future directions for research in this domain. Sharib Ali, Yamid Espinel, Yueming Jin, Peng Liu 0074, Bianca Güttner, Xukun Zhang, Lihua Zhang 0002, Thomas Dowrick, Matthew J. Clarkson, Shiting Xiao, Yifan Wu 0021, Lei Zhu 0003, Dai Sun, Micha Pfeiffer, Shahid Farid, Lena Maier-Hein, Emmanuel Buc, Adrien Bartoli |
Medical Image Anal. | 1 |
| 2025 | Multi-task learning with cross-task consistency for improved depth estimation in colonoscopyabstractaccuracy over the most accurate baseline state-of-the-art Big-to-Small (BTS) approach. All experiments are conducted on a recently released C3VD dataset, and thus, we provide a first benchmark of state-of-the-art methods on this dataset. Pedro Esteban Chavarrias-Solano, Andrew J. Bulpitt, Venkataraman Subramanian, Sharib Ali |
Medical Image Anal. | 4 |
| 2025 | Validating polyp and instrument segmentation methods in colonoscopy through Medico 2020 and MedAI 2021 ChallengesabstractAutomatic analysis of colonoscopy images has been an active field of research motivated by the importance of early detection of precancerous polyps. However, detecting polyps during the live examination can be challenging due to various factors such as variation of skills and experience among the endoscopists, lack of attentiveness, and fatigue leading to a high polyp miss-rate. Therefore, there is a need for an automated system that can flag missed polyps during the examination and improve patient care. Deep learning has emerged as a promising solution to this challenge as it can assist endoscopists in detecting and classifying overlooked polyps and abnormalities in real time, improving the accuracy of diagnosis and enhancing treatment. In addition to the algorithm’s accuracy, transparency and interpretability are crucial to explaining the whys and hows of the algorithm’s prediction. Further, conclusions based on incorrect decisions may be fatal, especially in medicine. Despite these pitfalls, most algorithms are developed in private data, closed source, or proprietary software, and methods lack reproducibility. Therefore, to promote the development of efficient and transparent methods, we have organized the “Medico automatic polyp segmentation (Medico 2020)” and “MedAI: Transparency in Medical Image Segmentation (MedAI 2021)” competitions. The Medico 2020 challenge received submissions from 17 teams, while the MedAI 2021 challenge also gathered submissions from another 17 distinct teams in the following year. We present a comprehensive summary and analyze each contribution, highlight the strength of the best-performing methods, and discuss the possibility of clinical translations of such methods into the clinic. Our analysis revealed that the participants improved dice coefficient metrics from 0.8607 in 2020 to 0.8993 in 2021 despite adding diverse and challenging frames (containing irregular, smaller, sessile, or flat polyps), which are frequently missed during a routine clinical examination. For the instrument segmentation task, the best team obtained a mean Intersection over union metric of 0.9364. For the transparency task, a multi-disciplinary team, including expert gastroenterologists, accessed each submission and evaluated the team based on open-source practices, failure case analysis, ablation studies, usability and understandability of evaluations to gain a deeper understanding of the models’ credibility for clinical deployment. The best team obtained a final transparency score of 21 out of 25. Through the comprehensive analysis of the challenge, we not only highlight the advancements in polyp and surgical instrument segmentation but also encourage subjective evaluation for building more transparent and understandable AI-based colonoscopy systems. Moreover, we discuss the need for multi-center and out-of-distribution testing to address the current limitations of the methods to reduce the cancer burden and improve patient care. • We present a detailed analysis of the Medico 2020 and MedAI 2021 challenges that are aimed at advancing automated polyp and instrument segmentation in colonoscopy for early colorectal cancer diagnosis by using novel deep learning methods. • To the best of our knowledge, MedAI 2021 is the first challenge to evaluate the transparency in both GI endoscopy and colonoscopy. Through the challenge, we invited the participants to list package dependencies and architecture code (with instructions for building, compiling, and training) and share trained model weights in a standardized format. Additionally, we invited participants to include the code for model evaluation and provide repository licensing information to enable others to use the code and the trained model responsibly. Moreover, we asked the participants to explain model predictions using intermediate heatmaps, perform ablation studies, conduct a thorough failure analysis, and share their code for reproducing the results. Finally, we performed a subjective evaluation by including an expert gastroenterologist in the group and gave the final transparency score based on the usefulness and understandability of the results. Our initiative aims to promote transparency in AI research and foster the development of reliable, interpretable, and trustworthy algorithms for use in medical image segmentation. • We provide a comparative analysis of the 34 proposed methods in both challenges (3 subtasks), covering small details of each team in the form of Tables, qualitative and quantitative results (failure analysis), and an in-depth analysis of the findings. • We explore trust, safety, interpretability, transparency, and generalizability issues and provide future strategies to overcome the current limitations of developed algorithms. Debesh Jha, Vanshali Sharma, Debapriya Banik, Debayan Bhattacharya, Kaushiki Roy, Steven Alexander Hicks, Nikhil Kumar Tomar, Vajira Thambawita, Adrian Krenzer, Ge-Peng Ji, Sahadev Poudel, George Batchkala, Saruar Alam, Awadelrahman M. A. Ahmed, Quoc-Huy Trinh, Zeshan Khan, Tien-Phat Nguyen, Shruti Shrestha, Sabari Nathan, Jeonghwan Gwak, Ritika Kumari Jha, Zheyuan Zhang 0001, Alexander Schlaefer, Debotosh Bhattacharjee, Manas Kamal Bhuyan, Pradip K. Das, Deng-Ping Fan, Sravanthi Parasa, Sharib Ali, Michael Riegler 0001, Pål Halvorsen, Thomas de Lange, Ulas Bagci |
Medical Image Anal. | 29 |
| 2024 | Ultrasound Image Segmentation and its Evaluation using Various Encoder ArchitecturesabstractUltrasound imaging faces particular challenges with high inter-operator variability and manual inspection of abnormalities. Deep learning segmentation methods are progressing rapidly to address these clinical challenges with improved automatic segmentation performance combining convolutional neural network (CNN) and Transformer approaches. However, challenges still remain with poor performance in boundary areas due to high speckle noise and training models with limited training data. This paper demonstrates a comparison of EfficientNet B2, EfficientNet B7 and PVT-v2-B5 pre-trained encoder backbones in a U-Net architecture. A noticeable improvement for all three pre-trained backbones is shown particularly in smaller Breast Ultrasound datasets with limited training data (EfficientNet B2: 4.5%, EfficientNet B7: 3.9%, PVT-v2-B5: 2.4%). However, the improvement is marginal (less than 1%) in the larger Nerve Ultrasound dataset. In addition, we noticed that the performance across all backbones is better in segmenting regular regions of interest (e.g. benign breast lesions), over irregular shapes (e.g. malignant breast lesions). The code used for this study is available at: https://github.com/aimsgroup-Leeds/IEEECBMS2024_US_Seg Edward Ellis, Andrew J. Bulpitt, Sharib Ali |
CBMS | 3 |
| 2024 | Multi-modal detection transformer with data engineering technique to stratify patients with Ulcerative ColitisabstractData engineering has become a powerful tool for machine learning applications over the last few years. In computer vision for generative AI, the necessity of large amounts of data for training models has become a significant bottleneck. Data augmentation is a technique with limitations, even if it is useful. Training models with synthetic data are a solution due to the flexibility and scalability of the data creation that they can offer. When creating a synthetic dataset, one of the biggest challenges, however, is generating accurate and valuable data that can guarantee that the samples are a factual representation of the area of interest or an image; therefore, validating the dataset by subject matter experts becomes crucial. Examples of the multiple applications this method can use are image captioning, question-answering applications, generative AI, and overall multimodal problems. For image captioning, an image and its description are needed. Regions-of-Interest (ROI) in the image can be associated with text, forming a multimodal relationship associating an ROI inside an image and describing it. This work proposes a methodology to create automatic ROI-description multimodal Ulcerative Colitis (UC) dataset construction. To deal with the requirement of a large dataset for model training, we introduce stable diffusion for generating images that represent these patch-level characteristics widely used to classify a sample into an MES score. We utilise the clinically accepted phenotypes for informed decision-making. These include ulcers, bleeding, and erosions. We use this dataset to train a transformer-based detection pipeline (DETR) to find the characteristic inside the raw UC image to generate an ROI and associate it with a text template that describes the region. Finally, we compare our results against a baseline ROI dataset that medical experts have validated. Alexis Lopez, Flor Helena Valencia, Gilberto Ochoa-Ruiz, Sharib Ali |
CBMS | 4 |
| 2024 | Domain Generalization for Endoscopic Image Segmentation by Disentangling Style-Content Information and SuperPixel ConsistencyabstractFrequent monitoring is necessary to stratify individuals based on their likelihood of developing gastrointestinal (GI) cancer precursors. In the clinical practice, white-light imaging (WLI), and complimentary modalities such as narrow-band imaging (NBI) and fluorescence imaging are used to assess risk areas. However, conventional deep learning (DL) models have depleted performance due to domain gap when a model is trained on one modality and tested on a different one. In our earlier approach we used superpixel based method referred to as “SUPRA” to effectively learn domain-invariant information using color and space distances to generate groups of pixels. One of the main limitations of this early work is that the aggregation does not exploit structural information, making it sub-optimal for segmentation tasks, especially for polyps and heterogeneous color distributions. Therefore, in this work, we propose an approach for style-content disentanglement using instance normalization and instance selective whitening (ISW) for an improved domain generalization when combined with SUPRA. We evaluate our approach on two datasets: EndoUDA Barret’s Esophagus and EndoUDA polyps and compare its performance with previous three state-of-the-art (SOTA) methods. Our findings demonstrate a notable enhancement in performance compared to both baseline and state-of-the-art methods across the target domain data. Specifically, our approach exhibited improvements of 14%, 10%, 8%, and 18% over the baseline and three SOTA methods on the polyp dataset. Additionally, it surpassed the second best method (EndoUDA) on the BE dataset by nearly 2%. Mansoor Ali Teevno, Rafael Martinez Garcia Peña, Gilberto Ochoa-Ruiz, Sharib Ali |
CBMS | 4 |
| 2024 | SSL-CPCD: Self-Supervised Learning With Composite Pretext-Class Discrimination for Improved Generalisability in Endoscopic Image AnalysisabstractData-driven methods have shown tremendous progress in medical image analysis. In this context, deep learning-based supervised methods are widely popular. However, they require a large amount of training data and face issues in generalisability to unseen datasets that hinder clinical translation. Endoscopic imaging data is characterised by large inter- and intra-patient variability that makes these models more challenging to learn representative features for downstream tasks. Thus, despite the publicly available datasets and datasets that can be generated within hospitals, most supervised models still underperform. While self-supervised learning has addressed this problem to some extent in natural scene data, there is a considerable performance gap in the medical image domain. In this paper, we propose to explore patch-level instance-group discrimination and penalisation of inter-class variation using additive angular margin within the cosine similarity metrics. Our novel approach enables models to learn to cluster similar representations, thereby improving their ability to provide better separation between different classes. Our results demonstrate significant improvement on all metrics over the state-of-the-art (SOTA) methods on the test set from the same and diverse datasets. We evaluated our approach for classification, detection, and segmentation. SSL-CPCD attains notable Top 1 accuracy of 79.77% in ulcerative colitis classification, an 88.62% mean average precision (mAP) for detection, and an 82.32% dice similarity coefficient for polyp segmentation tasks. These represent improvements of over 4%, 2%, and 3%, respectively, compared to the baseline architectures. We demonstrate that our method generalises better than all SOTA methods to unseen datasets, reporting over 7% improvement. Ziang Xu 0001, Jens Rittscher, Sharib Ali |
IEEE Trans. Medical Imaging | 3 |
| 2023 | Why is the Winner the Best?abstractInternational benchmarking competitions have become fundamental for the comparative performance assessment of image analysis methods. However, little attention has been given to investigating what can be learnt from these competitions. Do they really generate scientific progress? What are common and successful participation strategies? What makes a solution superior to a competing method? To address this gap in the literature, we performed a multicenter study with all 80 competitions that were conducted in the scope of IEEE ISBI 2021 and MICCAI 2021. Statistical analyses performed based on comprehensive descriptions of the submitted algorithms linked to their rank as well as the underlying participation strategies revealed common characteristics of winning solutions. These typically include the use of multi-task learning (63%) and/or multi-stage pipelines (61%), and a focus on augmentation (100%), image preprocessing (97%), data curation (79%), and post-processing (66%). The “typical” lead of a winning team is a computer scientist with a doctoral degree, five years of experience in biomedical image analysis, and four years of experience in deep learning. Two core general development strategies stood out for highly-ranked teams: the reflection of the metrics in the method design and the focus on analyzing and handling failure cases. According to the organizers, 43% of the winning algorithms exceeded the state of the art but only 11% completely solved the respective domain problem. The insights of our study could help researchers (1) improve algorithm development strategies when approaching new problems, and (2) focus on open research questions revealed by this work. Matthias Eisenmann, Annika Reinke, Vivienn Weru, Minu Tizabi, Fabian Isensee, Tim Adler, Sharib Ali, Vincent Andrearczyk, Marc Aubreville, Ujjwal Baid, Spyridon Bakas, Niranjan Balu, Sophia Bano, Jorge Bernal, Sebastian Bodenstedt, Alessandro Casella, Veronika Cheplygina, Marie Daum, Marleen de Bruijne, Adrien Depeursinge, Reuben Dorent, Jan Egger, David Gage Ellis, Sandy Engelhardt, Melanie Ganz-Benjaminsen, Noha M. Ghatwary, Gabriel Girard, Patrick Godau, Anubha Gupta, Lasse Hansen, Kanako Harada, Mattias P. Heinrich, Nicholas Heller, Alessa Hering, Arnaud Huaulmé, Pierre Jannin, A. Emre Kavur, Oldrich Kodym, Michal Kozubek 0001, Jianning Li 0002, Hongwei Li 0004, Jun Ma 0016, Carlos Martín-Isla, Bjoern Menze, J. Alison Noble, Valentin Oreiller, Nicolas Padoy, Sarthak Pati, Kelly Payette, Tim Rädsch, Jonathan Rafael-Patino, Vivek Singh Bawa, Stefanie Speidel, Carole H. Sudre, Kimberlin M. H. van Wijnen, Martin Wagner 0001, D. Wei, Amine Yamlahi, Moi Hoon Yap, C. Yuan, Maximilian Zenk, A. Zia, David Zimmerer, Dogu Baran Aydogan, Binod Bhattarai, Louise Bloch, Raphael Brüngel, J. Cho, C. Choi, Qi Dou 0001, Ivan Ezhov, Christoph M. Friedrich, C. Fuller, Rebati Raman Gaire, Adrian Galdran, Álvaro García-Faura, Maria Grammatikopoulou, S. Hong, Mostafa Jahanifar, I. Jang, Abdolrahim Kadkhodamohammadi, I. Kang, Florian Kofler, S. Kondo, Hugo J. Kuijf, M. Luu, Tomaz Martincic, Pedro Morais, Mohamed A. Naser, Bruno Oliveira 0002, David Owen 0001, S. Pang, Szymon Plotka, Élodie Puybareau, Nasir M. Rajpoot, K. Ryu, Numan Saeed, Adam J. Shephard, Dejan Stepec, Ronast Subedi, Guillaume Tochon, Helena R. Torres, Hélène Urien, João L. Vilaça, Kareem A. Wahid, Benedikt Wiestler, Marek Wodzinski, F. Xia, J. Xie, Z. Xiong, Sen Yang 0006, Klaus H. Maier-Hein, Paul F. Jaeger, Annette Kopp-Schneider, Lena Maier-Hein |
CVPR | 7 |
| 2023 | Anatomical-Aware Point-Voxel Network for Couinaud Segmentation in Liver CT
Xukun Zhang, Yang Liu 0007, Sharib Ali, Minghao Han, Tao Liu 0050, Peng Zhai, Zhiming Cui 0001, Peixuan Zhang, Lihua Zhang 0002 |
MICCAI (3) | 3 |
| 2023 | FANet: A Feedback Attention Network for Improved Biomedical Image SegmentationabstractThe increase of available large clinical and experimental datasets has contributed to a substantial amount of important contributions in the area of biomedical image analysis. Image segmentation, which is crucial for any quantitative analysis, has especially attracted attention. Recent hardware advancement has led to the success of deep learning approaches. However, although deep learning models are being trained on large datasets, existing methods do not use the information from different learning epochs effectively. In this work, we leverage the information of each training epoch to prune the prediction maps of the subsequent epochs. We propose a novel architecture called feedback attention network (FANet) that unifies the previous epoch mask with the feature map of the current training epoch. The previous epoch mask is then used to provide hard attention to the learned feature maps at different convolutional layers. The network also allows rectifying the predictions in an iterative fashion during the test time. We show that our proposed feedback attention model provides a substantial improvement on most segmentation metrics tested on seven publicly available biomedical imaging datasets demonstrating the effectiveness of FANet. The source code is available at https://github.com/nikhilroxtomar/FANet. Nikhil Kumar Tomar, Debesh Jha, Michael Riegler 0001, Håvard D. Johansen, Dag Johansen, Jens Rittscher, Pål Halvorsen, Sharib Ali |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2022 | GMSRF-Net: An Improved generalizability with Global Multi-Scale Residual Fusion Network for Polyp SegmentationabstractColonoscopy is a gold standard procedure but is highly operator-dependent. Efforts have been made to automate the detection and segmentation of polyps, a precancerous precursor, to effectively minimize missed rate. Widely used computer-aided polyp segmentation systems actuated by encoder-decoder have achieved high performance in terms of accuracy. However, polyp segmentation datasets collected from varied centers can follow different imaging protocols leading to difference in data distribution. As a result, most methods suffer from performance drop when trained and tested on different distributions and therefore, require re-training for each specific dataset. We address this generalizability issue by proposing a global multi-scale residual fusion network (GMSRF-Net). Our proposed network maintains high-resolution representations by performing multi-scale fusion operations across all resolution scales through dense connections while preserving low-level information. To further leverage scale information, we design cross multi-scale attention (CMSA) module that uses multi-scale features to identify, keep, and propagate informative features. Additionally, we introduce multi-scale feature selection (MSFS) modules to perform channel-wise attention that gates irrelevant features gathered through global multi-scale fusion within the GMSRF-Net. The repeated fusion operations gated by CMSA and MSFS demonstrate improved generalizability of our network.Experiments conducted on two different polyp segmentation datasets show that our proposed GMSRF-Net outperforms the previous top-performing state-of-the-art method by 8.34% and 10.31% on unseen CVC-ClinicDB and on unseen Kvasir-SEG, in terms of dice coefficient. Additionally, when tested on unseen CVC-ColonDB, we surpass the state-of-the-art method by 9.38% and 4.04% in terms of dice coefficient, when source dataset is Kvasir-SEG and CVC-ClinicDB, respectively. Sukalpa Chanda, Debesh Jha, Umapada Pal 0001, Sharib Ali |
ICPR | 5 |
| 2022 | TGANet: Text-Guided Attention for Improved Polyp Segmentation
Nikhil Kumar Tomar, Debesh Jha, Ulas Bagci, Sharib Ali |
MICCAI (3) | 4 |
| 2022 | Real-time instance segmentation of surgical instruments using attention and multi-scale feature fusionabstractPrecise instrument segmentation aids surgeons to navigate the body more easily and increases patient safety. While accurate tracking of surgical instruments in real-time plays a crucial role in minimally invasive computer-assisted surgeries, it is a challenging task to achieve, mainly due to: (1) a complex surgical environment, and (2) model design trade-off in terms of both optimal accuracy and speed. Deep learning gives us the opportunity to learn complex environment from large surgery scene environments and placements of these instruments in real world scenarios. The Robust Medical Instrument Segmentation 2019 challenge (ROBUST-MIS) provides more than 10,000 frames with surgical tools in different clinical settings. In this paper, we propose a light-weight single stage instance segmentation model complemented with a convolutional block attention module for achieving both faster and accurate inference. We further improve accuracy through data augmentation and optimal anchor localization strategies. To our knowledge, this is the first work that explicitly focuses on both real-time performance and improved accuracy. Our approach out-performed top team performances in the most recent edition of ROBUST-MIS challenge with over 44% improvement on area-based multi-instance dice metric MI_DSC and 39% on distance-based multi-instance normalized surface dice MI_NSD. We also demonstrate real-time performance (>60 frames-per-second) with different but competitive variants of our final approach. Juan Carlos Angeles Ceron, Gilberto Ochoa-Ruiz, Leonardo Chang 0001, Sharib Ali |
Medical Image Anal. | 4 |
| 2022 | MSRF-Net: A Multi-Scale Residual Fusion Network for Biomedical Image SegmentationabstractMethods based on convolutional neural networks have improved the performance of biomedical image segmentation. However, most of these methods cannot efficiently segment objects of variable sizes and train on small and biased datasets, which are common for biomedical use cases. While methods exist that incorporate multi-scale fusion approaches to address the challenges arising with variable sizes, they usually use complex models that are more suitable for general semantic segmentation problems. In this paper, we propose a novel architecture called Multi-Scale Residual Fusion Network (MSRF-Net), which is specially designed for medical image segmentation. The proposed MSRF-Net is able to exchange multi-scale features of varying receptive fields using a Dual-Scale Dense Fusion (DSDF) block. Our DSDF block can exchange information rigorously across two different resolution scales, and our MSRF sub-network uses multiple DSDF blocks in sequence to perform multi-scale fusion. This allows the preservation of resolution, improved information flow and propagation of both high- and low-level features to obtain accurate segmentation maps. The proposed MSRF-Net allows to capture object variabilities and provides improved results on different biomedical datasets. Extensive experiments on MSRF-Net demonstrate that the proposed method outperforms the cutting-edge medical image segmentation methods on four publicly available datasets. We achieve the Dice Coefficient (DSC) of 0.9217, 0.9420, and 0.9224, 0.8824 on Kvasir-SEG, CVC-ClinicDB, 2018 Data Science Bowl dataset, and ISIC-2018 skin lesion segmentation challenge dataset respectively. We further conducted generalizability tests and achieved DSC of 0.7921 and 0.7575 on CVC-ClinicDB and Kvasir-SEG, respectively. Debesh Jha, Sukalpa Chanda, Umapada Pal 0001, Håvard D. Johansen, Dag Johansen, Michael Riegler 0001, Sharib Ali, Pål Halvorsen |
IEEE J. Biomed. Health Informatics | 8 |
| 2021 | NanoNet: Real-Time Polyp Segmentation in Video Capsule Endoscopy and ColonoscopyabstractDeep learning in gastrointestinal endoscopy can assist to improve clinical performance and be helpful to assess lesions more accurately. To this extent, semantic segmentation methods that can perform automated real-time delineation of a region-of-interest, e.g., boundary identification of cancer or pre-cancerous lesions, can benefit both diagnosis and interventions. However, accurate and real-time segmentation of endoscopic images is extremely challenging due to its high operator dependence and high-definition image quality. To utilize automated methods in clinical settings, it is crucial to design lightweight models with low latency such that they can be integrated with low-end endoscope hardware devices. In this work, we propose NanoNet, a novel architecture for the segmentation of video capsule endoscopy and colonoscopy images. Our proposed architecture allows real-time performance and has higher segmentation accuracy compared to other more complex ones. We use video capsule endoscopy and standard colonoscopy datasets with polyps, and a dataset consisting of endoscopy biopsies and surgical instruments, to evaluate the effectiveness of our approach. Our experiments demonstrate the increased performance of our architecture in terms of a trade-off between model complexity, speed, model parameters, and metric performances. Moreover, the resulting models' size is relatively tiny, with only nearly 36,000 parameters compared to traditional deep learning approaches having millions of parameters. Debesh Jha, Nikhil Kumar Tomar, Sharib Ali, Michael Riegler 0001, Håvard D. Johansen, Dag Johansen, Thomas de Lange, Pål Halvorsen |
CBMS | 3 |
| 2021 | EndoUDA: A Modality Independent Segmentation Approach for Endoscopy Imaging
Numan Celik, Sharib Ali, Soumya Gupta 0001, Barbara Braden, Jens Rittscher |
MICCAI (3) | 2 |
| 2021 | Kvasir-Instrument: Diagnostic and Therapeutic Tool Segmentation Dataset in Gastrointestinal Endoscopy
Debesh Jha, Sharib Ali, Krister Emanuelsen, Steven Alexander Hicks, Vajira Thambawita, Enrique Garcia-Ceja, Michael Riegler 0001, Thomas de Lange, Peter Thelin Schmidt, Håvard D. Johansen, Dag Johansen, Pål Halvorsen |
MMM (2) | 2 |
| 2021 | Deep learning for detection and segmentation of artefact and disease instances in gastrointestinal endoscopyabstractThe Endoscopy Computer Vision Challenge (EndoCV) is a crowd-sourcing initiative to address eminent problems in developing reliable computer aided detection and diagnosis endoscopy systems and suggest a pathway for clinical translation of technologies. Whilst endoscopy is a widely used diagnostic and treatment tool for hollow-organs, there are several core challenges often faced by endoscopists, mainly: 1) presence of multi-class artefacts that hinder their visual interpretation, and 2) difficulty in identifying subtle precancerous precursors and cancer abnormalities. Artefacts often affect the robustness of deep learning methods applied to the gastrointestinal tract organs as they can be confused with tissue of interest. EndoCV2020 challenges are designed to address research questions in these remits. In this paper, we present a summary of methods developed by the top 17 teams and provide an objective comparison of state-of-the-art methods and methods designed by the participants for two sub-challenges: i) artefact detection and segmentation (EAD2020), and ii) disease detection and segmentation (EDD2020). Multi-center, multi-organ, multi-class, and multi-modal clinical endoscopy datasets were compiled for both EAD2020 and EDD2020 sub-challenges. The out-of-sample generalization ability of detection algorithms was also evaluated. Whilst most teams focused on accuracy improvements, only a few methods hold credibility for clinical usability. The best performing teams provided solutions to tackle class imbalance, and variabilities in size, origin, modality and occurrences by exploring data augmentation, data fusion, and optimal class thresholding techniques. Sharib Ali, Mariia Dmitrieva, Noha M. Ghatwary, Sophia Bano, Gorkem Polat, Alptekin Temizel, Adrian Krenzer, Amar Hekalo, Bogdan J. Matuszewski, Mourad Gridach, Irina Voiculescu, Vishnusai Yoganand, Arnav Chavan, Aryan Raj, Nhan T. Nguyen, Dat Q. Tran, Lê Duy Huynh, Nicolas Boutry, Shahadate Rezvy, Haijian Chen, Yoon Ho Choi, Anand Subramanian 0004, Velmurugan Balasubramanian, Xiaohong W. Gao, Hongyu Hu, Yusheng Liao, Danail Stoyanov, Christian Daul, Stefano Realdon, Renato Cannizzaro, Dominique Lamarque, Terry Tran-Nguyen, Adam Bailey, Barbara Braden, James E. East, Jens Rittscher |
Medical Image Anal. | 1 |
| 2021 | A deep learning framework for quality assessment and restoration in video endoscopyabstractEndoscopy is a routine imaging technique used for both diagnosis and minimally invasive surgical treatment. Artifacts such as motion blur, bubbles, specular reflections, floating objects and pixel saturation impede the visual interpretation and the automated analysis of endoscopy videos. Given the widespread use of endoscopy in different clinical applications, robust and reliable identification of such artifacts and the automated restoration of corrupted video frames is a fundamental medical imaging problem. Existing state-of-the-art methods only deal with the detection and restoration of selected artifacts. However, typically endoscopy videos contain numerous artifacts which motivates to establish a comprehensive solution. In this paper, a fully automatic framework is proposed that can: 1) detect and classify six different artifacts, 2) segment artifact instances that have indefinable shapes, 3) provide a quality score for each frame, and 4) restore partially corrupted frames. To detect and classify different artifacts, the proposed framework exploits fast, multi-scale and single stage convolution neural network detector. In addition, we use an encoder-decoder model for pixel-wise segmentation of irregular shaped artifacts. A quality score is introduced to assess video frame quality and to predict image restoration success. Generative adversarial networks with carefully chosen regularization and training strategies for discriminator-generator networks are finally used to restore corrupted frames. The detector yields the highest mean average precision (mAP) of 45.7 and 34.7, respectively for 25% and 50% IoU thresholds, and the lowest computational time of 88 ms allowing for near real-time processing. The restoration models for blind deblurring, saturation correction and inpainting demonstrate significant improvements over previous methods. On a set of 10 test videos, an average of 68.7% of video frames successfully passed the quality score (≥0.9) after applying the proposed restoration framework thereby retaining 25% more frames compared to the raw videos. The importance of artifacts detection and their restoration on improved robustness of image analysis methods is also demonstrated in this work. Sharib Ali, Felix Zhou 0001, Adam Bailey, Barbara Braden, James E. East, Jens Rittscher |
Medical Image Anal. | 1 |
| 2021 | A comprehensive analysis of classification methods in gastrointestinal endoscopy imagingabstractGastrointestinal (GI) endoscopy has been an active field of research motivated by the large number of highly lethal GI cancers. Early GI cancer precursors are often missed during the endoscopic surveillance. The high missed rate of such abnormalities during endoscopy is thus a critical bottleneck. Lack of attentiveness due to tiring procedures, and requirement of training are few contributing factors. An automatic GI disease classification system can help reduce such risks by flagging suspicious frames and lesions. GI endoscopy consists of several multi-organ surveillance, therefore, there is need to develop methods that can generalize to various endoscopic findings. In this realm, we present a comprehensive analysis of the Medico GI challenges: Medical Multimedia Task at MediaEval 2017, Medico Multimedia Task at MediaEval 2018, and BioMedia ACM MM Grand Challenge 2019. These challenges are initiative to set-up a benchmark for different computer vision methods applied to the multi-class endoscopic images and promote to build new approaches that could reliably be used in clinics. We report the performance of 21 participating teams over a period of three consecutive years and provide a detailed analysis of the methods used by the participants, highlighting the challenges and shortcomings of the current approaches and dissect their credibility for the use in clinical settings. Our analysis revealed that the participants achieved an improvement on maximum Mathew correlation coefficient (MCC) from 82.68% in 2017 to 93.98% in 2018 and 95.20% in 2019 challenges, and a significant increase in computational speed over consecutive years. Debesh Jha, Sharib Ali, Steven Alexander Hicks, Vajira Thambawita, Hanna Borgli, Pia H. Smedsrud, Thomas de Lange, Konstantin Pogorelov, Philipp Harzig, Minh-Triet Tran, Wenhua Meng, Trung-Hieu Hoang, Danielle Dias, Tobey H. Ko, Taruna Agrawal, Olga Ostroukhova, Zeshan Khan, Muhammad Atif Tahir, Yang Liu 0007, Mathias Kirkerød, Dag Johansen, Mathias Lux, Håvard D. Johansen, Michael Riegler 0001, Pål Halvorsen |
Medical Image Anal. | 2 |
| 2016 | Illumination invariant optical flow using neighborhood descriptors
Sharib Ali, Christian Daul, Ernest Galbrun, Walter Blondel |
Comput. Vis. Image Underst. | 1 |
| 2016 | Anisotropic motion estimation on edge preserving Riesz wavelets for robust video mosaicing
Sharib Ali, Christian Daul, Ernest Galbrun, François Guillemin, Walter Blondel |
Pattern Recognit. | 1 |
| 2013 | Fast mosaicing of cystoscopic images from dense correspondence: Combined SURF and TV-L1 optical flow methodabstractIn white light cystoscopy, bladder images are characterized by a strong texture and scene illumination variability which complicates image mosaicing. State-of-art methods exhibit high image registration accuracy at the expense of computational time. We propose an algorithm which selects either a feature based method or an optical flow method according to the image texture; for fast and accurate bladder wall mosaicing. Total variation (TV) optical flow method (deduced by duality) guarantees robust registration of poorly textured images. Realistic phantom images are registered with subpixel accuracy with a processing speed-up by a factor of 8 and 16 for two reference methods. Patient data results also illustrate the performance of the algorithm. Sharib Ali, Christian Daul, Thomas Weibel, Walter Blondel |
ICIP | 1 |