EDBT 2026 Demo / reviewers in the wild / expert
Ulas Bagci
dblp:83/6388
· DBLP profile ↗
81ranked-venue papers
9as first author
43since 2021 · last 2026
0000-0001-7379-6829ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 47 · 5 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 44 · 6 first-author · 15 since 2021Artificial intelligence and machine learning · 18 · 1 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 7 · 6 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | What They Saw, Not Just Where They Looked: Semantic Scanpath Similarity via VLMs and NLP metricsabstractScanpath similarity metrics are central to eye-movement research, yet existing methods predominantly evaluate spatial and temporal alignment while neglecting semantic equivalence between attended image regions. We present a semantic scanpath similarity framework that integrates vision-language models (VLMs) into eye-tracking analysis. Each fixation is encoded under controlled visual context (patch-based and marker-based strategies) and transformed into concise textual descriptions, which are aggregated into scanpath-level representations. Semantic similarity is then computed using embedding-based and lexical NLP metrics and compared against established spatial measures, including MultiMatch and DTW. Experiments on free-viewing eye-tracking data demonstrate that semantic similarity captures partially independent variance from geometric alignment, revealing cases of high content agreement despite spatial divergence. We further analyze the impact of contextual encoding on description fidelity and metric stability. Our findings suggest that multimodal foundation models enable interpretable, content-aware extensions of classical scanpath analysis, providing a complementary dimension for gaze research within the ETRA community. Mohamed Amine Kerkouri, Marouane Tliba, Bin Wang 0068, Aladine Chetouani, Ulas Bagci, Alessandro Bruno |
ETRA | 5 |
| 2026 | GazeVaLM: A Multi-Observer Eye-Tracking Benchmark for Evaluating Clinical Realism in AI-Generated X-RaysabstractWe introduce GazeVaLM, a public eye-tracking dataset for studying clinical perception during chest radiograph authenticity assessment. The dataset comprises 960 gaze recordings from 16 expert radiologists interpreting 30 real and 30 synthetic chest X-rays (generated by diffusion based generative AI) under two conditions: diagnostic assessment and real-fake classification (Visual Turing test). For each image–observer pair, we provide raw gaze samples, fixation maps, scanpaths, saliency density maps, structured diagnostic labels, and authenticity judgments. We extend the protocol to 6 state-of-the-art multimodal LLMs, releasing their predicted diagnoses, authenticity labels, and confidence scores under matched conditions — enabling direct human–AI comparison at both decision and uncertainty levels. We further provide analyses of gaze agreement, inter-observer consistency, and benchmarking of radiologists versus LLMs in diagnostic accuracy and authenticity detection. GazeVaLM supports research in gaze modeling, clinical decision-making, human–AI comparison, generative image realism assessment, and uncertainty quantification. By jointly releasing visual attention data, clinical labels, and model predictions, we aim to facilitate reproducible research on how experts and AI systems perceive, interpret, and evaluate medical images. The dataset is available at https://huggingface.co/datasets/davidcwong/GazeVaLM. David C. Wong 0005, Zeynep Isik, Bin Wang 0068, Marouane Tliba, Gorkem Durak, Elif Keles, Halil Ertugrul Aktas, Aladine Chetouani, Cagdas Topel, Nicolo Gennaro, Camila Lopes Vendrami, Tugce Agirlar Trabzonlu, Amir Ali Rahsepar, Laetitia Perronne, Matthew Antalek, Onural Ozturk, Gokcan Okur, Andrew C. Gordon, Ayis Pyrros, Frank H. Miller, Amir Borhani, Hatice Savas, Eric M. Hart, Elizabeth A. Krupinski, Ulas Bagci |
ETRA | 25 |
| 2026 | UP2D: Uncertainty-aware progressive pseudo-label denoising for source-free domain adaptive medical image segmentation
Thanh-Huy Nguyen, Quang-Khai Bui-Tran, Manh D. Ho, Ba Thinh Lam, Vi Vu, Hoang-Thien Nguyen, Phat K. Huynh, Ulas Bagci |
Neurocomputing | 8 |
| 2026 | AdverIN: Monotonic adversarial intensity attack for domain generalization in medical image segmentation
Zheyuan Zhang 0001, Bin Wang 0068, Lanhong Yao, Elif Keles, Debesh Jha, Matthew Antalek, Gorkem Durak, Alpay Medetalibeyoglu, Concetto Spampinato, Baris Turkbey, Boqing Gong, Ulas Bagci |
Medical Image Anal. | 12 |
| 2026 | VHU-Net: Variational hadamard U-Net for body MRI bias field correction
Xin Zhu 0005, A. Enis Çetin, Gorkem Durak, Batuhan Gündogdu, Ziliang Hong, Hongyi Pan, Halil Ertugrul Aktas, Elif Keles, Hatice Savas, Aytekin Oto, Hiten D. Patel, Adam B. Murphy, Ashley Ross, Frank H. Miller, Baris Turkbey, Ulas Bagci |
Medical Image Anal. | 16 |
| 2026 | SAM-guided prompt learning for Multiple Sclerosis lesion segmentationabstractAccurate segmentation of Multiple Sclerosis (MS) lesions remains a critical challenge in medical image analysis due to their small size, irregular shape, and sparse distribution. Despite recent progress in vision foundation models — such as SAM and its medical variant MedSAM — these models have not yet been explored in the context of MS lesion segmentation. Moreover, their reliance on manually crafted prompts and high inference-time computational cost limits their applicability in clinical workflows, especially in resource-constrained environments. In this work, we introduce a novel training-time framework for effective and efficient MS lesion segmentation. Our method leverages SAM solely during training to guide a prompt learner that automatically discovers task-specific embeddings. At inference, SAM is replaced by a lightweight convolutional aggregator that maps the learned embeddings directly into segmentation masks—enabling fully automated, low-cost deployment. We show that our approach significantly outperforms existing specialized methods on the public MSLesSeg dataset, establishing new performance benchmarks in a domain where foundation models had not previously been applied. To assess generalizability, we also evaluate our method on pancreas and prostate segmentation tasks, where it achieves competitive accuracy while requiring an order of magnitude fewer parameters and computational resources compared to SAM-based pipelines. By eliminating the need for foundation models at inference time, our framework enables efficient segmentation without sacrificing accuracy. This design bridges the gap between large-scale pretraining and real-world clinical deployment, offering a scalable and practical solution for MS lesion segmentation and beyond. Code is available at https://github.com/perceivelab/MS-SAM-LESS . Federica Proietto Salanitri, Giovanni Bellitto, Salvatore Calcagno 0002, Ulas Bagci, Concetto Spampinato, Manuela Pennisi |
Pattern Recognit. Lett. | 4 |
| 2026 | REN: Anatomically-Informed Mixture-of-Experts for Interstitial Lung Disease DiagnosisabstractMixture-of-Experts (MoE) architectures achieve scalable learning by routing inputs to specialized subnetworks through conditional computation. However, conventional MoE designs assume homogeneous expert capability and domain-agnostic routing-assumptions that are fundamentally misaligned with medical imaging, where anatomical structure and regional disease heterogeneity govern pathological patterns. We introduce Regional Expert Networks (REN), the first anatomically-informed MoE framework for medical image classification. REN encodes anatomical priors by training seven specialized experts, each dedicated to a distinct lung lobe or bilateral lung combination, enabling precise modeling of region-specific pathological variation. Multi-modal gating mechanisms dynamically integrate radiomics biomarkers with deep learning (DL) features extracted by convolutional (CNN), Transformer (ViT), and state-space (Mamba) architectures to weight expert contributions at inference. Applied to interstitial lung disease (ILD) classification on a 597-patient, 1,898-scan longitudinal cohort, REN achieves consistently superior performance: the radiomics-guided ensemble attains an average AUC of $0.8646~\pm ~0.0467$ , a +12.5% improvement over the SwinUNETR single-model baseline (AUC 0.7685, ${p}={0}.{031}$ ). Lower-lobe experts reach AUCs of 0.88-0.90, outperforming DL baselines (CNN: 0.76-0.79) and mirroring known patterns of basal ILD progression. Evaluated under rigorous patient-level cross-validation, REN demonstrates strong generalizability and clinical interpretability, establishing a scalable, anatomically-guided framework potentially extensible to other structured medical imaging tasks. Code is available on our GitHub https://github.com/NUBagciLab/MoE-REN. Alec Peltekian, Halil Ertugrul Aktas, Gorkem Durak, Kevin M. Grudzinski, Bradford C. Bemiss, Carrie Richardson, Jane E. Dematte, G. R. Scott Budinger, Anthony J. Esposito, Alexander V. Misharin, Alok N. Choudhary, Ankit Agrawal 0001, Ulas Bagci |
IEEE Trans. Medical Imaging | 13 |
| 2025 | Shifts in Doctors' Eye Movements Between Real and AI-Generated Medical Images
David C. Wong 0005, Bin Wang 0068, Gorkem Durak, Marouane Tliba, Mohamed Amine Kerkouri, Aladine Chetouani, A. Enis Çetin, Cagdas Topel, Nicolo Gennaro, Camila Lopes Vendrami, Tugce Agirlar Trabzonlu, Amir Ali Rahsepar, Laetitia Perronne, Matthew Antalek, Onural Ozturk, Gokcan Okur, Andrew C. Gordon, Ayis Pyrros, Frank H. Miller, Amir Borhani, Hatice Savas, Eric M. Hart, Elizabeth A. Krupinski, Ulas Bagci |
ETRA | 24 |
| 2025 | MDNet: Multi-Decoder Network for Abdominal CT Organs SegmentationabstractAccurate segmentation of organs from abdominal CT scans is essential for clinical applications such as diagnosis, treatment planning, and patient monitoring. To handle challenges of heterogeneity in organ shapes, sizes, and complex anatomical relationships, we propose a Multi decoder network (MDNet), an encoder-decoder network that uses the pre-trained MiT-B2 as the encoder and multiple different decoder networks. Each decoder network is connected to a different part of the encoder via a multi-scale feature enhancement dilated block. With each decoder, we increase the depth of the network iteratively and refine segmentation masks, enriching feature maps by integrating previous decoders’ feature maps. To refine the feature map further, we also utilize the predicted masks from the previous decoder to the current decoder to provide spatial attention across foreground and background regions. MDNet effectively refines the segmentation mask with a high dice similarity coefficient (DSC) of 0.9013 and 0.9169 on the Liver Tumor segmentation (LiTS) and MSD Spleen datasets. Additionally, it reduces Hausdorff distance (HD) to 3.79 for the LiTS dataset and 2.26 for the spleen segmentation dataset, underscoring the precision of MDNet in capturing the complex contours. Moreover, MDNet is more interpretable and robust compared to the other baseline models. The code for our architecture is available at https://github.com/DebeshJha/MDNet. Debesh Jha, Nikhil Kumar Tomar, Koushik Biswas, Gorkem Durak, Matthew Antalek, Zheyuan Zhang 0001, Bin Wang 0068, Md Mostafijur Rahman, Hongyi Pan, Alpay Medetalibeyoglu, Vandan Gorade, Yury Velichko, Daniela P. Ladner, Amir Borhani, Ulas Bagci |
ICASSP | 16 |
| 2025 | Frequency-Based Federated Domain Generalization for Polyp SegmentationabstractFederated Learning (FL) offers a powerful strategy for training machine learning models across decentralized datasets while maintaining data privacy, yet domain shifts among clients can degrade performance, particularly in medical imaging tasks like polyp segmentation. This paper introduces a novel Frequency-Based Domain Generalization (FDG) framework, utilizing soft-thresholding and hard-thresholding in the Fourier domain to address these challenges. By applying soft-thresholding and hard-thresholding to Fourier coefficients, our method generates new images with reduced background noise and enhances the model’s ability to generalize across diverse medical imaging domains. Extensive experiments demonstrate substantial improvements in segmentation accuracy and domain robustness over baseline methods. This innovation integrates frequency domain techniques into FL, presenting a resilient approach to overcoming domain variability in decentralized medical image analysis. Hongyi Pan, Debesh Jha, Koushik Biswas, Ulas Bagci |
ICASSP | 4 |
| 2025 | Transformer-Enhanced Iterative Feedback Mechanism For Polyp SegmentationabstractColorectal cancer (CRC) is the third most common cause of cancer diagnosed in the United States. Notably, CRC is the leading cause of cancer in younger men less than 50 years old. Colonoscopy is considered the gold standard for the early diagnosis of CRC. Skills vary significantly among endoscopists, and a high miss rate is reported. Automated polyp segmentation can reduce the missed rates, and timely treatment is possible in the early stage. To address this challenge, we introduce Feedback Attention Network-v2 (FANetv2), an advanced encoder-decoder network designed to accurately segment polyps from colonoscopy images. Leveraging an initial input mask generated by Otsu thresholding, FANetv2 iteratively refines its binary segmentation masks through a novel feedback attention mechanism informed by the mask predictions of previous epochs. Additionally, it employs a text-guided approach that integrates essential information about the number (one or many) and size (small, medium, large) of polyps to further enhance its feature representation capabilities. This dual-task approach facilitates accurate polyp segmentation and aids in the auxiliary classification of polyp attributes, significantly boosting the model’s performance. Our comprehensive evaluations on the publicly available BKAI-IGH and CVC-ClinicDB datasets demonstrate the superior performance of FANetv2, evidenced by high dice similarity coefficients (DSC) of 0.9186 and 0.9481, along with low Hausdorff distances of 2.83 and 3.19, respectively. The source code for FANetv2 will be made available at https://github.com/nikhilroxtomar/FANetv2. Nikhil Kumar Tomar, Debesh Jha, Koushik Biswas, Ulas Bagci |
ICASSP | 4 |
| 2025 | ViCTr: Vital Consistency Transfer for Pathology Aware Image SynthesisabstractSynthesizing medical images remains challenging due to limited annotated pathological data, modality domain gaps, and the complexity of representing diffuse pathologies such as liver cirrhosis. Existing methods often struggle to maintain anatomical fidelity while accurately modeling pathological features, frequently relying on priors derived from natural images or inefficient multi-step sampling. In this work, we introduce ViCTr (Vital Consistency Transfer), a novel two-stage framework that combines a rectified flow trajectory with a Tweedie-corrected diffusion process to achieve high-fidelity, pathology-aware image synthesis. First, we pretrain ViCTr on the ATLAS-8k dataset using Elastic Weight Consolidation (EWC) to preserve critical anatomical structures. We then fine-tune the model adversarially with Low-Rank Adaptation (LoRA) modules for precise control over pathology severity. By reformulating Tweedie's formula within a linear trajectory framework, ViCTr supports one-step sampling, reducing inference from 50 steps to just 4, without sacrificing anatomical realism. We evaluate ViCTr on BTCV (CT), AMOS (MRI), and CirrMRI600+ (cirrhosis) datasets. Results demonstrate state-of-the-art performance, achieving a Medical Frechet Inception Distance (MFID) of 17.01 for cirrhosis synthesis 28% lower than existing approaches and improving nnUNet segmentation by +3.8% mDSC when used for data augmentation. Radiologist reviews indicate that ViCTr-generated liver cirrhosis MRIs are clinically indistinguishable from real scans. To our knowledge, ViCTr is the first method to provide fine-grained, pathology-aware MRI synthesis with graded severity control, closing a critical gap in AI-driven medical imaging research. Onkar Susladkar, Gayatri Deshmukh, Yalcin Tur, Gorkem Durak, Ulas Bagci |
ICCV | 5 |
| 2025 | VideoAds for Fast-Paced Video Understanding
Zheyuan Zhang 0001, Wanying Dou, Linkai Peng, Hongyi Pan, Ulas Bagci, Boqing Gong |
ICCV | 5 |
| 2025 | Order-aware Interactive SegmentationabstractInteractive segmentation aims to accurately segment target objects with minimal user interactions. However, current methods often fail to accurately separate target objects from the background, due to a limited understanding of order, the relative depth between objects in a scene. To address this issue, we propose OIS: order-aware interactive segmentation, where we explicitly encode the relative depth between objects into order maps. We introduce a novel order-aware attention, where the order maps seamlessly guide the user interactions (in the form of clicks) to attend to the image features. We further present an object-aware attention module to incorporate a strong object-level understanding to better differentiate objects with similar order. Our approach allows both dense and sparse integration of user clicks, enhancing both accuracy and efficiency as compared to prior works. Experimental results demonstrate that OIS achieves state-of-the-art performance, improving mIoU after one click by 7.61 on the HQSeg44K dataset and 1.32 on the DAVIS dataset as compared to the previous state-of-the-art SegNext, while also doubling inference speed compared to current leading methods. Bin Wang 0068, Anwesa Choudhuri, Meng Zheng 0002, Zhongpai Gao, Benjamin Planche, Andong Deng, Qin Liu 0008, Terrence Chen, Ulas Bagci, Ziyan Wu 0001 |
ICLR | 9 |
| 2025 | Optimizing Neural Network Effectiveness via Non-monotonicity RefinementabstractActivation functions play a crucial role in artificial neural networks by introducing non-linearities that enable networks to learn complex patterns in data. An appropriate choice of an activation function plays a crucial role in the training dynamics of a neural network, which can boost network performance significantly. Rectified Linear Unit (ReLU) and its variants, like leaky ReLU and parametric ReLU, have emerged as the most popular activations due to their ability to enable faster training and generalization in deep neural networks despite having some significant issues like vanishing gradient problems. In this paper, we have proposed smooth functions, which we call the AMSU family, which are smooth approximations of the maximum function. We derive three activations from the AMSU family, namely AMSU-1, AMSU-2, & AMSU-3, and show their effectiveness in different deep learning problems. By simply replacing the ReLU function, Top-1 accuracy improves by 5.88%, 5.96%, and 5.32% on the CIFAR100 dataset on the ShuffleNet V2 model. Also, replacing ReLU with AMSU-1, AMSU-2, and AMSU-3, Top-1 accuracy improves by 8.50%, 8.29%, and 7.70% on the CIFAR100 dataset on the ShuffleNet V2 model with FGSM attack. Also, Replacing ReLU with AMSU-1, AMSU-2, and AMSU-3 on ImageNet-1K data, we got 3%-5% improvement on ShuffleNet and MobileNet models. The source code is publicly available at https://github.com/koushik313/AMSU. Koushik Biswas, Amit Reza, Meghana Karri, Debesh Jha, Hongyi Pan, Nikhil Kumar Tomar, Aliza Subedi, Smriti Regmi, Ulas Bagci |
WACV | 9 |
| 2025 | Frequency-Domain Refinement of Vision Transformers for Robust Medical Image Segmentation Under DegradationabstractMedical image segmentation is crucial for precise diagnosis, treatment planning, and disease monitoring in clinical settings. While convolutional neural networks (CNNs) have achieved remarkable success, they struggle with modeling long-range dependencies. Vision Transformers (ViTs) address this limitation by leveraging self-attention mechanisms to capture global contextual information. However, ViTs often fall short in local feature description, which is crucial for precise segmentation. To address this issue, we reformulate self-attention in the frequency domain to enhance both local and global feature representation. Our approach, the Enhanced Wave Vision Transformer (EW-ViT), incorporates wavelet decomposition within the self-attention block to adaptively refine feature representation in low and high-frequency components. We also introduce the Prompt-Guided High-Frequency Refiner (PGHFR) module to handle image degradation, which mainly affects high-frequency components. This module uses implicit prompts to encode degradation-specific information and adjust high-frequency representations accordingly. Additionally, we apply a contrastive learning strategy to maintain feature consistency and ensure robustness against noise, leading to state-of-the-art (SOTA) performance in medical image segmentation, especially under various conditions of degradation. Source code is available at GitHub. Sanaz Karimijafarbigloo, Sina Ghorbani Kolahi, Reza Azad, Ulas Bagci, Dorit Merhof |
WACV | 4 |
| 2025 | Uncertainty-Guided Cross Attention Ensemble Mean Teacher for Semi-Supervised Medical Image SegmentationabstractThis work proposes a novel framework, UncertaintyGuided Cross Attention Ensemble Mean Teacher (UGCEMT), for achieving state-of-the-art performance in semisupervised medical image segmentation. UG-CEMT leverages the strengths of co-training and knowledge distillation by combining a Cross-attention Ensemble Mean Teacher framework (CEMT) inspired by Vision Transformers (ViT) with uncertainty-guided consistency regularization and Sharpness-Aware Minimization emphasizing uncertainty. UG-CEMT improves semi-supervised performance while maintaining a consistent network architecture and task setting by fostering high disparity between sub-networks. Experiments demonstrate significant advantages over existing methods like Mean Teacher and Crosspseudo Supervision in terms of disparity, domain generalization, and medical image segmentation performance. UG-CEMT achieves state-of-the-art results on multi-center prostate MRI and cardiac MRI datasets, where object segmentation is particularly challenging. Our results show that using only 10% labeled data, UG-CEMT approaches the performance of fully supervised methods, demonstrating its effectiveness in exploiting unlabeled data for robust medical image segmentation. The code is publicly available at https://github.com/Meghnak13/UG-CEMT Meghana Karri, Amit Soni Arya, Koushik Biswas, Nicolo Gennaro, Vedat Cicek, Gorkem Durak, Yuri S. Velichko, Ulas Bagci |
WACV | 8 |
| 2025 | Position of artificial intelligence in healthcare and future perspective
Vedat Cicek, Ulas Bagci |
Artif. Intell. Medicine | 2 |
| 2025 | Validating polyp and instrument segmentation methods in colonoscopy through Medico 2020 and MedAI 2021 ChallengesabstractAutomatic analysis of colonoscopy images has been an active field of research motivated by the importance of early detection of precancerous polyps. However, detecting polyps during the live examination can be challenging due to various factors such as variation of skills and experience among the endoscopists, lack of attentiveness, and fatigue leading to a high polyp miss-rate. Therefore, there is a need for an automated system that can flag missed polyps during the examination and improve patient care. Deep learning has emerged as a promising solution to this challenge as it can assist endoscopists in detecting and classifying overlooked polyps and abnormalities in real time, improving the accuracy of diagnosis and enhancing treatment. In addition to the algorithm’s accuracy, transparency and interpretability are crucial to explaining the whys and hows of the algorithm’s prediction. Further, conclusions based on incorrect decisions may be fatal, especially in medicine. Despite these pitfalls, most algorithms are developed in private data, closed source, or proprietary software, and methods lack reproducibility. Therefore, to promote the development of efficient and transparent methods, we have organized the “Medico automatic polyp segmentation (Medico 2020)” and “MedAI: Transparency in Medical Image Segmentation (MedAI 2021)” competitions. The Medico 2020 challenge received submissions from 17 teams, while the MedAI 2021 challenge also gathered submissions from another 17 distinct teams in the following year. We present a comprehensive summary and analyze each contribution, highlight the strength of the best-performing methods, and discuss the possibility of clinical translations of such methods into the clinic. Our analysis revealed that the participants improved dice coefficient metrics from 0.8607 in 2020 to 0.8993 in 2021 despite adding diverse and challenging frames (containing irregular, smaller, sessile, or flat polyps), which are frequently missed during a routine clinical examination. For the instrument segmentation task, the best team obtained a mean Intersection over union metric of 0.9364. For the transparency task, a multi-disciplinary team, including expert gastroenterologists, accessed each submission and evaluated the team based on open-source practices, failure case analysis, ablation studies, usability and understandability of evaluations to gain a deeper understanding of the models’ credibility for clinical deployment. The best team obtained a final transparency score of 21 out of 25. Through the comprehensive analysis of the challenge, we not only highlight the advancements in polyp and surgical instrument segmentation but also encourage subjective evaluation for building more transparent and understandable AI-based colonoscopy systems. Moreover, we discuss the need for multi-center and out-of-distribution testing to address the current limitations of the methods to reduce the cancer burden and improve patient care. • We present a detailed analysis of the Medico 2020 and MedAI 2021 challenges that are aimed at advancing automated polyp and instrument segmentation in colonoscopy for early colorectal cancer diagnosis by using novel deep learning methods. • To the best of our knowledge, MedAI 2021 is the first challenge to evaluate the transparency in both GI endoscopy and colonoscopy. Through the challenge, we invited the participants to list package dependencies and architecture code (with instructions for building, compiling, and training) and share trained model weights in a standardized format. Additionally, we invited participants to include the code for model evaluation and provide repository licensing information to enable others to use the code and the trained model responsibly. Moreover, we asked the participants to explain model predictions using intermediate heatmaps, perform ablation studies, conduct a thorough failure analysis, and share their code for reproducing the results. Finally, we performed a subjective evaluation by including an expert gastroenterologist in the group and gave the final transparency score based on the usefulness and understandability of the results. Our initiative aims to promote transparency in AI research and foster the development of reliable, interpretable, and trustworthy algorithms for use in medical image segmentation. • We provide a comparative analysis of the 34 proposed methods in both challenges (3 subtasks), covering small details of each team in the form of Tables, qualitative and quantitative results (failure analysis), and an in-depth analysis of the findings. • We explore trust, safety, interpretability, transparency, and generalizability issues and provide future strategies to overcome the current limitations of developed algorithms. Debesh Jha, Vanshali Sharma, Debapriya Banik, Debayan Bhattacharya, Kaushiki Roy, Steven Alexander Hicks, Nikhil Kumar Tomar, Vajira Thambawita, Adrian Krenzer, Ge-Peng Ji, Sahadev Poudel, George Batchkala, Saruar Alam, Awadelrahman M. A. Ahmed, Quoc-Huy Trinh, Zeshan Khan, Tien-Phat Nguyen, Shruti Shrestha, Sabari Nathan, Jeonghwan Gwak, Ritika Kumari Jha, Zheyuan Zhang 0001, Alexander Schlaefer, Debotosh Bhattacharjee, Manas Kamal Bhuyan, Pradip K. Das, Deng-Ping Fan, Sravanthi Parasa, Sharib Ali, Michael Riegler 0001, Pål Halvorsen, Thomas de Lange, Ulas Bagci |
Medical Image Anal. | 33 |
| 2025 | Large-scale multi-center CT and MRI segmentation of pancreas with deep learningabstractAutomated volumetric segmentation of the pancreas on cross-sectional imaging is needed for diagnosis and follow-up of pancreatic diseases. While CT-based pancreatic segmentation is more established, MRI-based segmentation methods are understudied, largely due to a lack of publicly available datasets, benchmarking research efforts, and domain-specific deep learning methods. In this retrospective study, we collected a large dataset (767 scans from 499 participants) of T1-weighted (T1 W) and T2-weighted (T2 W) abdominal MRI series from five centers between March 2004 and November 2022. We also collected CT scans of 1,350 patients from publicly available sources for benchmarking purposes. We introduced a new pancreas segmentation method, called PanSegNet , combining the strengths of nnUNet and a Transformer network with a new linear attention module enabling volumetric computation. We tested PanSegNet ’s accuracy in cross-modality (a total of 2,117 scans) and cross-center settings with Dice and Hausdorff distance (HD95) evaluation metrics. We used Cohen’s kappa statistics for intra and inter-rater agreement evaluation and paired t-tests for volume and Dice comparisons, respectively. For segmentation accuracy, we achieved Dice coefficients of 88.3% (±7.2%, at case level) with CT, 85.0% (±7.9%) with T1 W MRI, and 86.3% (±6.4%) with T2 W MRI. There was a high correlation for pancreas volume prediction with R 2 of 0.91, 0.84, and 0.85 for CT, T1 W, and T2 W, respectively. We found moderate inter-observer (0.624 and 0.638 for T1 W and T2 W MRI, respectively) and high intra-observer agreement scores. All MRI data is made available at https://osf.io/kysnj/ . Our source code is available at https://github.com/NUBagciLab/PaNSegNet . • We develop a first-ever cross-platform compatible (T1 W, T2 W, and CT) pancreas segmentation tool, named PanSegNet . • PaNSegNet has innovative “linear self-attention” blocks to reduce computational cost significantly while operating on 3D. • We shared our both source code and multi-center multi-contrast MRI datasets with ground truths. • PaNSegNet underwent rigorous validation, including cross-domain and multi-center comparisons between CT and MRI scans. Zheyuan Zhang 0001, Elif Keles, Gorkem Durak, Yavuz Taktak, Onkar Susladkar, Vandan Gorade, Debesh Jha, Asli C. Ormeci, Alpay Medetalibeyoglu, Lanhong Yao, Bin Wang 0068, Ilkin Isler, Linkai Peng, Hongyi Pan, Camila Lopes Vendrami, Amir Bourhani, Yury Velichko, Boqing Gong, Concetto Spampinato, Ayis Pyrros, Pallavi Tiwari, Derk C. F. Klatte, Megan Engels, Sanne Hoogenboom, Candice W. Bolan, Emil Agarunov, Nassier Harfouch, Chenchan Huang, Marco J. Bruno, Ivo Schoots, Rajesh Keswani, Frank H. Miller, Tamas Gonda, Cemal Yazici, Temel Tirkes, Baris Turkbey, Michael B. Wallace, Ulas Bagci |
Medical Image Anal. | 38 |
| 2025 | DiffBoost: Enhancing Medical Image Segmentation via Text-Guided Diffusion ModelabstractLarge-scale, big-variant, high-quality data are crucial for developing robust and successful deep-learning models for medical applications since they potentially enable better generalization performance and avoid overfitting. However, the scarcity of high-quality labeled data always presents significant challenges. This paper proposes a novel approach to address this challenge by developing controllable diffusion models for medical image synthesis, called DiffBoost. We leverage recent diffusion probabilistic models to generate realistic and diverse synthetic medical image data that preserve the essential characteristics of the original medical images by incorporating edge information of objects to guide the synthesis process. In our approach, we ensure that the synthesized samples adhere to medically relevant constraints and preserve the underlying structure of imaging data. Due to the random sampling process by the diffusion model, we can generate an arbitrary number of synthetic images with diverse appearances. To validate the effectiveness of our proposed method, we conduct an extensive set of medical image segmentation experiments on multiple datasets, including Ultrasound breast (+13.87%), CT spleen (+0.38%), and MRI prostate (+7.78%), achieving significant improvements over the baseline segmentation methods. The promising results demonstrate the effectiveness of our DiffBoost for medical image segmentation tasks and show the feasibility of introducing a first-ever text-guided diffusion model for general medical image segmentation tasks. With carefully designed ablation experiments, we investigate the influence of various data augmentations, hyper-parameter settings, patch size for generating random merging mask settings, and combined influence with different network architectures. Source code with checkpoints are available at https://github.com/NUBagciLab/DiffBoost. Zheyuan Zhang 0001, Lanhong Yao, Bin Wang 0068, Debesh Jha, Gorkem Durak, Elif Keles, Alpay Medetalibeyoglu, Ulas Bagci |
IEEE Trans. Medical Imaging | 8 |
| 2024 | PGDS: Pose-Guidance Deep Supervision for Mitigating Clothes-Changing in Person Re-IdentificationabstractPerson Re-Identification (Re-ID) task seeks to enhance the tracking of multiple individuals by surveillance cameras. It supports multimodal tasks, including text-based person retrieval and human matching. One of the most significant challenges faced in Re-ID is clothes-changing, where the same person may appear in different outfits. While previous methods have made notable progress in maintaining clothing data consistency and handling clothing change data, they still rely excessively on clothing information, which can limit performance due to the dynamic nature of human appearances. To mitigate this challenge, we propose the Pose-Guidance Deep Supervision (PGDS), an effective framework for learning pose guidance within the Re-ID task. It consists of three modules: a human encoder, a pose encoder, and a Pose-to-Human Projection module(PHP). Our framework guides the human encoder, i.e., the main re-identification model, with pose information from the pose encoder through multiple layers via the knowledge transfer mechanism from the PHP module, helping the human encoder learn body parts information without increasing computation resources in the inference stage. Through extensive experiments, our method surpasses the performance of current state-of-the-art methods, demonstrating its robustness and effectiveness for real-world applications. Our code is available at https://github.com/huyquoctrinh/PGDS. Quoc-Huy Trinh, Nhat-Tan Bui, Dinh-Hieu Hoang, Phuoc-Thao Vo Thi, Hai-Dang Nguyen, Debesh Jha, Ulas Bagci, T. Hoang Ngan Le, Minh-Triet Tran |
AVSS | 7 |
| 2024 | SAM-EG: Segment Anything Model with Egde Guidance framework for efficient Polyp Segmentation
Quoc-Huy Trinh, Hai-Dang Nguyen, Bao-Tram Nguyen Ngoc, Debesh Jha, Ulas Bagci, Minh-Triet Tran |
BMVC | 5 |
| 2024 | Prediction of MRI-Induced Power Absorption in Patients with DBS LeadsabstractThe interaction between deep brain stimulation (DBS) systems and magnetic resonance imaging (MRI) can induce tissue heating in patients. While electromagnetic (EM) simulations can be used to estimate the specific absorption rate (SAR) values in the presence of an implanted DBS system, they are computationally expensive. To address this drawback, we predict local SAR values in the tips of DBS leads with machine learning based efficient algorithms, specifically XgBoost and deep learning. We significantly outperformed the previous state of the art, and adapted new machine learning models based on Residual Networks family as well as XgBoost models. We observed that already extracted limited features are better suited for ensemble learning via XgBoost than deep networks due the small-data regime. Although we conclude that boosting gradient algorithm is more suitable for this non-linear regression problem due to structured nature of the data and small data regime, we found that width plays a more critical role than depth in network design and it has a strong potential for future research. Our experimental results, using a dataset of 260 instances that are patient-derived and artificial, reached an outstanding RMSE of 17.8 W/kg with XgBoost, 78 W/kg with deep networks, given that the previous study on this problem reached a state-of-the-art root mean square error value (RMSE) of 168 W/kg. Yalcin Tur, Jasmine Vu, Selam Waktola, Alpay Medetalibeyoglu, Laleh Golestanirad, Ulas Bagci |
CBMS | 6 |
| 2024 | Domain Generalization with fourier Transform and soft thresholdingabstractDomain generalization aims to train models on multiple source domains so that they can generalize well to unseen target domains. Among many domain generalization methods, Fourier-transformbased domain generalization methods have gained popularity primarily because they exploit the power of Fourier transformation to capture essential patterns and regularities in the data, making the model more robust to domain shifts. The mainstream Fouriertransform-based domain generalization swaps the Fourier amplitude spectrum while preserving the phase spectrum between the source and the target images. However, it neglects background interference in the amplitude spectrum. To overcome this limitation, we introduce a soft-thresholding function in the Fourier domain. We apply this newly designed algorithm to retinal fundus image segmentation, which is important for diagnosing ocular diseases but the neural network’s performance can degrade across different sources due to domain shifts. The proposed technique basically enhances fundus image augmentation by eliminating small values in the Fourier domain and providing better generalization. The innovative nature of the soft thresholding fused with Fourier-transform-based domain generalization improves neural network models’ performance by reducing the target images’ background interference significantly. Experiments on public data validate our approach’s effectiveness over conventional and state-of-the-art methods with superior segmentation metrics. Hongyi Pan, Bin Wang 0068, Zheyuan Zhang 0001, Xin Zhu 0005, Debesh Jha, A. Enis Çetin, Concetto Spampinato, Ulas Bagci |
ICASSP | 8 |
| 2024 | ProFONet: Prototypical Feature Space Optimized Network for Few Shot Classification
Vandan Gorade, Debesh Jha, Koushik Biswas, Pethuru Raj Chelliah, Ulas Bagci |
ICPR (7) | 6 |
| 2024 | Harmonized Spatial and Spectral Learning for Generalized Medical Image Segmentation
Vandan Gorade, Sparsh Mittal, Debesh Jha, Rekha Singhal, Ulas Bagci |
ICPR (13) | 5 |
| 2024 | Evidential Federated Learning for Skin Lesion Image Classification
Rutger Hendrix, Federica Proietto Salanitri, Concetto Spampinato, Simone Palazzo, Ulas Bagci |
ICPR (29) | 5 |
| 2024 | Adaptive Smooth Activation Function for Improved Organ Segmentation and Disease Diagnosis
Koushik Biswas, Debesh Jha, Nikhil Kumar Tomar, Meghana Karri, Amit Reza, Gorkem Durak, Alpay Medetalibeyoglu, Matthew Antalek, Yury Velichko, Daniela P. Ladner, Amir Borhani, Ulas Bagci |
MICCAI (9) | 12 |
| 2024 | Beyond Self-Attention: Deformable Large Kernel Attention for Medical Image SegmentationabstractMedical image segmentation has seen significant improvements with transformer models, which excel in grasping far-reaching contexts and global contextual information. However, the increasing computational demands of these models, proportional to the squared token count, limit their depth and resolution capabilities. Most current methods process D volumetric image data slice-by-slice (called pseudo 3D), missing crucial inter-slice information and thus reducing the model’s overall performance. To address these challenges, we introduce the concept of Deformable Large Kernel Attention (D-LKA Attention), a streamlined attention mechanism employing large convolution kernels to fully appreciate volumetric context. This mechanism operates within a receptive field akin to self-attention while sidestepping the computational overhead. Additionally, our proposed attention mechanism benefits from deformable convolutions to flexibly warp the sampling grid, enabling the model to adapt appropriately to diverse data patterns. We designed both 2D and 3D adaptations of the D-LKA Attention, with the latter excelling in cross-depth data understanding. Together, these components shape our novel hierarchical Vision Transformer architecture, the D-LKA Net. Evaluations of our model against leading methods on popular medical segmentation datasets (Synapse, NIH Pancreas, and Skin lesion) demonstrate its superior performance. Our code is publicly available at GitHub. Reza Azad, Leon Niggemeier, Michael Huttemann, Amirhossein Kazerouni, Ehsan Khodapanah Aghdam, Yury Velichko, Ulas Bagci, Dorit Merhof |
WACV | 7 |
| 2024 | SynergyNet: Bridging the Gap between Discrete and Continuous Representations for Precise Medical Image SegmentationabstractIn recent years, continuous latent space (CLS) and discrete latent space (DLS) deep learning models have been proposed for medical image analysis for improved performance. However, these models encounter distinct challenges. CLS models capture intricate details but often lack interpretability in terms of structural representation and robustness due to their emphasis on low-level features. Conversely, DLS models offer interpretability, robustness, and the ability to capture coarse-grained information thanks to their structured latent space. However, DLS models have limited efficacy in capturing fine-grained details. To address the limitations of both DLS and CLS models, we propose SynergyNet, a novel bottleneck architecture designed to enhance existing encoder-decoder segmentation frameworks. SynergyNet seamlessly integrates discrete and continuous representations to harness complementary information and successfully preserves both fine and coarse-grained details in the learned representations. Our extensive experiment on multi-organ segmentation and cardiac datasets demonstrates that SynergyNet outperforms other state of the art methods including TransUNet: dice scores improving by 2.16%, and Hausdorff scores improving by 11.13%, respectively. When evaluating skin lesion and brain tumor segmentation datasets, we observe a remarkable improvement of 1.71% in Intersection-over-Union scores for skin lesion segmentation and of 8.58% for brain tumor segmentation. Our innovative approach paves the way for enhancing the overall performance and capabilities of deep learning models in the critical domain of medical image analysis. Vandan Gorade, Sparsh Mittal, Debesh Jha, Ulas Bagci |
WACV | 4 |
| 2024 | INCODE: Implicit Neural Conditioning with Prior Knowledge EmbeddingsabstractImplicit Neural Representations (INRs) have revolutionized signal representation by leveraging neural networks to provide continuous and smooth representations of complex data. However, existing INRs face limitations in capturing fine-grained details, handling noise, and adapting to diverse signal types. To address these challenges, we introduce INCODE, a novel approach that enhances the control of the sinusoidal-based activation function in INRs using deep prior knowledge. INCODE comprises a harmonizer network and a composer network, where the harmonizer network dynamically adjusts key parameters of the activation function. Through a task-specific pre-trained model, INCODE adapts the task-specific parameters to optimize the representation process. Our approach not only excels in representation, but also extends its prowess to tackle complex tasks such as audio, image, and 3D shape reconstructions, as well as intricate challenges such as neural radiance fields (NeRFs), and inverse problems, including denoising, super-resolution, inpainting, and CT reconstruction. Through comprehensive experiments, INCODE demonstrates its superiority in terms of robustness, accuracy, quality, and convergence rate, broadening the scope of signal representation. Please visit the project’s website for details on the proposed method and access to the code. Amirhossein Kazerouni, Reza Azad, Alireza Hosseini, Dorit Merhof, Ulas Bagci |
WACV | 5 |
| 2024 | GazeGNN: A Gaze-Guided Graph Neural Network for Chest X-ray ClassificationabstractEye tracking research is important in computer vision because it can help us understand how humans interact with the visual world. Specifically for high-risk applications, such as in medical imaging, eye tracking can help us to comprehend how radiologists and other medical professionals search, analyze, and interpret images for diagnostic and clinical purposes. Hence, the application of eye tracking techniques in disease classification has become increasingly popular in recent years. Contemporary works usually transform gaze information collected by eye tracking devices into visual attention maps (VAMs) to supervise the learning process. However, this is a time-consuming preprocessing step, which stops us from applying eye tracking to radiologists’ daily work. To solve this problem, we propose a novel gaze-guided graph neural network (GNN), GazeGNN, to leverage raw eye-gaze data without being converted into VAMs. In GazeGNN, to directly integrate eye gaze into image classification, we create a unified representation graph that models both images and gaze pattern information. With this benefit, we develop a real-time, real-world, end-to-end disease classification algorithm for the first time in the literature. This achievement demonstrates the practicality and feasibility of integrating real-time eye tracking techniques into the daily work of radiologists. To our best knowledge, GazeGNN is the first work that adopts GNN to integrate image and eye-gaze data. Our experiments on the public chest X-ray dataset show that our proposed method exhibits the best classification performance compared to existing methods. The code is available at https://github.com/ukaukaaaa/GazeGNN. Bin Wang 0068, Hongyi Pan, Armstrong Aboah, Zheyuan Zhang 0001, Elif Keles, Drew A. Torigian, Baris Turkbey, Elizabeth A. Krupinski, Jayaram K. Udupa, Ulas Bagci |
WACV | 10 |
| 2024 | Domain Generalization with Correlated Style UncertaintyabstractDomain generalization (DG) approaches intend to extract domain invariant features that can lead to a more robust deep learning model. In this regard, style augmentation is a strong DG method taking advantage of instance-specific feature statistics containing informative style characteristics to synthetic novel domains. While it is one of the state-of-the-art methods, prior works on style augmentation have either disregarded the interdependence amongst distinct feature channels or have solely constrained style augmentation to linear interpolation. To address these research gaps, in this work, we introduce a novel augmentation approach, named Correlated Style Uncertainty (CSU), surpassing the limitations of linear interpolation in style statistic space and simultaneously preserving vital correlation information. Our method's efficacy is established through extensive experimentation on diverse cross-domain computer vision and medical imaging classification tasks: PACS, Office-Home, and Camelyon17 datasets, and the Duke-Market1501 instance retrieval task. The results showcase a remarkable improvement margin over existing state-of-the-art techniques. The source code is available https://github.com/freshman97/CSU. Zheyuan Zhang 0001, Bin Wang 0068, Debesh Jha, Ugur Demir, Ulas Bagci |
WACV | 5 |
| 2024 | Federated Learning for Medical Applications: A Taxonomy, Current Trends, Challenges, and Future Research DirectionsabstractWith the advent of the Internet of Things (IoT), artificial intelligence (AI), machine learning (ML), and deep learning (DL) algorithms, the landscape of data-driven medical applications has emerged as a promising avenue for designing robust and scalable diagnostic and prognostic models from medical data. This has gained a lot of attention from both academia and industry, leading to significant improvements in healthcare quality. However, the adoption of AI-driven medical applications still faces tough challenges, including meeting security, privacy, and Quality-of-Service (QoS) standards. Recent developments in federated learning (FL) have made it possible to train complex machine-learned models in a distributed manner and have become an active research domain, particularly processing the medical data at the edge of the network in a decentralized way to preserve privacy and address security concerns. To this end, in this article, we explore the present and future of FL technology in medical applications where data sharing is a significant challenge. We delve into the current research trends and their outcomes, unraveling the complexities of designing reliable and scalable FL models. This article outlines the fundamental statistical issues in FL, tackles device-related problems, addresses security challenges, and navigates the complexity of privacy concerns, all while highlighting its transformative potential in the medical field. Our study primarily focuses on medical applications of FL, particularly in the context of global cancer diagnosis. We highlight the potential of FL to enable computer-aided diagnosis tools that address this challenge with greater effectiveness than traditional data-driven methods. Recent literature has shown that FL models are robust and generalize well to new data, which is essential for medical applications. We hope that this comprehensive review will serve as a checkpoint for the field, summarizing the current state of the art and identifying open problems and future research directions. Ashish Rauniyar, Desta Haileselassie Hagos, Debesh Jha, Jan Erik Håkegård, Ulas Bagci, Danda B. Rawat, Vladimir Vlassov |
IEEE Internet Things J. | 5 |
| 2024 | COVID-19 Detection From Respiratory Sounds With Hierarchical Spectrogram TransformersabstractMonitoring of prevalent airborne diseases such as COVID-19 characteristically involves respiratory assessments. While auscultation is a mainstream method for preliminary screening of disease symptoms, its utility is hampered by the need for dedicated hospital visits. Remote monitoring based on recordings of respiratory sounds on portable devices is a promising alternative, which can assist in early assessment of COVID-19 that primarily affects the lower respiratory tract. In this study, we introduce a novel deep learning approach to distinguish patients with COVID-19 from healthy controls given audio recordings of cough or breathing sounds. The proposed approach leverages a novel hierarchical spectrogram transformer (HST) on spectrogram representations of respiratory sounds. HST embodies self-attention mechanisms over local windows in spectrograms, and window size is progressively grown over model stages to capture local to global context. HST is compared against state-of-the-art conventional and deep-learning baselines. Demonstrations on crowd-sourced multi-national datasets indicate that HST outperforms competing methods, achieving over 90% area under the receiver operating characteristic curve (AUC) in detecting COVID-19 cases. Idil Aytekin, Onat Dalmaz, Kaan Gönç, Haydar Ankishan, Emine Ulku Saritas, Ulas Bagci, Haydar Celik, Tolga Çukur |
IEEE J. Biomed. Health Informatics | 6 |
| 2023 | Laplacian-Former: Overcoming the Limitations of Vision Transformers in Local Texture Detection
Reza Azad, Amirhossein Kazerouni, Babak Azad, Ehsan Khodapanah Aghdam, Yury Velichko, Ulas Bagci, Dorit Merhof |
MICCAI (3) | 6 |
| 2023 | A Privacy-Preserving Walk in the Latent Space of Generative Models for Medical Applications
Matteo Pennisi, Federica Proietto Salanitri, Giovanni Bellitto, Simone Palazzo, Ulas Bagci, Concetto Spampinato |
MICCAI (3) | 5 |
| 2023 | DilatedSegNet: A Deep Dilated Segmentation Network for Polyp Segmentation
Nikhil Kumar Tomar, Debesh Jha, Ulas Bagci |
MMM (1) | 3 |
| 2022 | Video Capsule Endoscopy Classification using Focal Modulation Guided Convolutional Neural NetworkabstractVideo capsule endoscopy is a hot topic in computer vision and medicine. Deep learning can have a positive impact on the future of video capsule endoscopy technology. It can improve the anomaly detection rate, reduce physicians' time for screening, and aid in real-world clinical analysis. Computer-Aided diagnosis (CADx) classification system for video capsule endoscopy has shown a great promise for further improvement. For example, detection of cancerous polyp and bleeding can lead to swift medical response and improve the survival rate of the patients. To this end, an automated CADx system must have high throughput and decent accuracy. In this study, we propose FocalConvNet, a focal modulation network integrated with lightweight convolutional layers for the classification of small bowel anatomical landmarks and luminal findings. FocalCon-vNet leverages focal modulation to attain global context and allows global-local spatial interactions throughout the forward pass. Moreover, the convolutional block with its intrinsic induetive/learning bias and capacity to extract hierarchical features allows our FocalConvNet to achieve favourable results with high throughput. We compare our FocalConvNet with other state-of-the-art (SOTA) on Kvasir-Capsule, a large-scale VCE dataset with 44,228 frames with 13 classes of different anomalies. We achieved the weighted F1-score, recall and Matthews correlation coefficient (MCC) of 0.6734, 0.6373 and 0.2974, respectively, outperforming SOTA methodologies. Further, we obtained the highest throughput of 148.02 images/second rate to establish the potential of FocalConvNet in a real-time clinical environ-ment. The code of the proposed FocalConvNet is available at https://github.com/NoviceMAn-prog/FocalConvNet. Nikhil Kumar Tomar, Ulas Bagci, Debesh Jha |
CBMS | 3 |
| 2022 | Automatic Polyp Segmentation with Multiple Kernel Dilated Convolution NetworkabstractThe detection and removal of precancerous polyps through colonoscopy is the primary technique for the prevention of colorectal cancer worldwide. However, the miss rate of colorectal polyp varies significantly among the endoscopists. It is well known that a computer-aided diagnosis (CAD) system can assist endoscopists in detecting colon polyps and minimize the variation among endoscopists. In this study, we introduce a novel deep learning architecture, named MKDCNet, for automatic polyp segmentation robust to significant changes in polyp data distribution. MKDCNet is simply an encoder-decoder neural network that uses the pre-trained ResNet50 as the encoder and novel multiple kernel dilated convolution (MKDC) block that expands the field of view to learn more robust and heterogeneous representation. Extensive experiments on four publicly available polyp datasets and cell nuclei dataset show that the proposed MKDCNet outperforms the state-of-the-art methods when trained and tested on the same dataset as well when tested on unseen polyp datasets from different distributions. With rich results, we demonstrated the robustness of the proposed architecture. From an efficiency perspective, our algorithm can process at ($\approx 45$) frames per second on RTX 3090 GPU. MKDCNet can be a strong benchmark for building real-time systems for clinical colonoscopies. The code of the proposed MKDCNet is available at https://github.com/nikhilroxtomar/MKDCNet. Nikhil Kumar Tomar, Ulas Bagci, Debesh Jha |
CBMS | 3 |
| 2022 | TGANet: Text-Guided Attention for Improved Polyp Segmentation
Nikhil Kumar Tomar, Debesh Jha, Ulas Bagci, Sharib Ali |
MICCAI (3) | 3 |
| 2021 | Capsules for biomedical image segmentation
Rodney LaLonde, Ziyue Xu 0001, Ismail Irmakci, Sanjay Jain 0002, Ulas Bagci |
Medical Image Anal. | 5 |
| 2020 | Deep Recurrent-Convolutional Model for Automated Segmentation of Craniomaxillofacial CT ScansabstractIn this paper we define a deep learning architecture for automated segmentation of anatomical structures in Craniomaxillofacial (CMF) CT scans that leverages the recent success of encoder-decoder models for semantic segmentation of natural images. In particular, we propose a fully convolutional deep network that combines the advantages of recent fully convolutional models, such as Tiramisu, with squeeze-and-excitation blocks for feature recalibration, integrated with convolutional LSTMs to model spatio-temporal correlations between consecutive slices. The proposed segmentation network shows superior performance and generalization capabilities (to different structures and imaging modalities) than state of the art methods on automated segmentation of CMF structures (e.g., mandibles and airways) in several standard benchmarks (e.g., MICCAI datasets) and on new datasets proposed herein, effectively facing shape variability. Francesca Murabito, Simone Palazzo, Federica Proietto Salanitri, Francesco Rundo, Ulas Bagci, Daniela Giordano, Rosalia Leonardi, Concetto Spampinato |
ICPR | 5 |
| 2020 | Deep Multi-stage Model for Automated Landmarking of Craniomaxillofacial CT ScansabstractIn this paper we define a deep multi-stage architecture for automated landmarking of craniomaxillofacial (CMF) CT images. Our model is composed of three subnetworks that first localize, on reduced-resolution images, areas where landmarks may be found and then refine the search, at full-resolution scale, through a hierarchical structure aiming at increasing the granularity of the investigated region. The multi-stage pipeline is designed to deal with full resolution data and does not require any additional pre-processing step to reduce search space, as opposed to existing methods that can be only adopted for searching landmarks located in well-defined anatomical structures (e.g., mandibles). The automated landmarking system is tested on identifying landmarks located in several CMF regions, achieving an average error of 0.8 mm, significantly lower than expert readings. The proposed model also outperforms baselines and is on par with existing models that employ additional upstream segmentation, on state-of-the-art benchmarks. Simone Palazzo, Giovanni Bellitto, Luca Prezzavento, Francesco Rundo, Ulas Bagci, Daniela Giordano, Rosalia Leonardi, Concetto Spampinato |
ICPR | 5 |
| 2020 | Variational Capsule EncoderabstractWe propose a novel capsule network based variational encoder architecture, called Bayesian capsules (B-Caps), to modulate the mean and standard deviation of the sampling distribution in the latent space. We hypothesized that this approach can learn a better representation of features in the latent space than traditional approaches. Our hypothesis was tested by using the learned latent variables for image reconstruction task, where for MNIST and Fashion-MNIST datasets, different classes were separated successfully in the latent space using our proposed model. Our experimental results have shown improved reconstruction and classification performances for both datasets adding credence to our hypothesis. We also showed that by increasing the latent space dimension, the proposed B-Caps was able to learn a better representation when compared to the traditional variational auto-encoders (VAE). Hence our results indicate the strength of capsule networks in representation learning which has never been examined under the VAE settings before. Harish RaviPrakash, Syed Muhammad Anwar, Ulas Bagci |
ICPR | 3 |
| 2020 | Encoding Visual Attributes in Capsules for Explainable Medical Diagnoses
Rodney LaLonde, Drew A. Torigian, Ulas Bagci |
MICCAI (1) | 3 |
| 2020 | Instance-Level Microtubule TrackingabstractWe propose a new method of instance-level microtubule (MT) tracking in time-lapse image series using recurrent attention. Our novel deep learning algorithm segments individual MTs at each frame. Segmentation results from successive frames are used to assign correspondences among MTs. This ultimately generates a distinct path trajectory for each MT through the frames. Based on these trajectories, we estimate MT velocities. To validate our proposed technique, we conduct experiments using real and simulated data. We use statistics derived from real time-lapse series of MT gliding assays to simulate realistic MT time-lapse image series in our simulated data. This data set is employed as pre-training and hyperparameter optimization for our network before training on the real data. Our experimental results show that the proposed supervised learning algorithm improves the precision for MT instance velocity estimation drastically to 71.3% from the baseline result (29.3%). We also demonstrate how the inclusion of temporal information into our deep network can reduce the false negative rates from 67.8% (baseline) down to 28.7% (proposed). Our findings in this work are expected to help biologists characterize the spatial arrangement of MTs, specifically the effects of MT-MT interactions. Samira Masoudi, Afsaneh Razi, Cameron H. G. Wright, Jesse C. Gatlin, Ulas Bagci |
IEEE Trans. Medical Imaging | 5 |
| 2019 | PAN: Projective Adversarial Network for Medical Image Segmentation
Naji Khosravan, Aliasghar Mortazi, Michael B. Wallace, Ulas Bagci |
MICCAI (6) | 4 |
| 2019 | INN: Inflated Neural Networks for IPMN Diagnosis
Rodney LaLonde, Irene Tanner, Katerina Nikiforaki, Georgios Z. Papadakis, Pujan Kandel, Candice W. Bolan, Michael B. Wallace, Ulas Bagci |
MICCAI (5) | 8 |
| 2019 | A collaborative computer aided diagnosis (C-CAD) system with eye-tracking, sparse attentional model, and deep learning
Naji Khosravan, Haydar Celik, Baris Turkbey, Elizabeth C. Jones, Bradford J. Wood, Ulas Bagci |
Medical Image Anal. | 6 |
| 2019 | Evaluation of algorithms for Multi-Modality Whole Heart Segmentation: An open-access grand challengeabstractKnowledge of whole heart anatomy is a prerequisite for many clinical applications. Whole heart segmentation (WHS), which delineates substructures of the heart, can be very valuable for modeling and analysis of the anatomy and functions of the heart. However, automating this segmentation can be challenging due to the large variation of the heart shape, and different image qualities of the clinical data. To achieve this goal, an initial set of training data is generally needed for constructing priors or for training. Furthermore, it is difficult to perform comparisons between different methods, largely due to differences in the datasets and evaluation metrics used. This manuscript presents the methodologies and evaluation results for the WHS algorithms selected from the submissions to the Multi-Modality Whole Heart Segmentation (MM-WHS) challenge, in conjunction with MICCAI 2017. The challenge provided 120 three-dimensional cardiac images covering the whole heart, including 60 CT and 60 MRI volumes, all acquired in clinical environments with manual delineation. Ten algorithms for CT data and eleven algorithms for MRI data, submitted from twelve groups, have been evaluated. The results showed that the performance of CT WHS was generally better than that of MRI WHS. The segmentation of the substructures for different categories of patients could present different levels of challenge due to the difference in imaging and variations of heart shapes. The deep learning (DL)-based methods demonstrated great potential, though several of them reported poor results in the blinded evaluation. Their performance could vary greatly across different network structures and training strategies. The conventional algorithms, mainly based on multi-atlas segmentation, demonstrated good performance, though the accuracy and computational efficiency could be limited. The challenge, including provision of the annotated training data and the blinded evaluation for submitted algorithms on the test data, continues as an ongoing benchmarking resource via its homepage (www.sdspeople.fudan.edu.cn/zhuangxiahai/0/mmwhs/). Xiahai Zhuang, Lei Li 0020, Christian Payer, Darko Stern, Martin Urschler, Mattias P. Heinrich, Julien Oster, Chunliang Wang, Örjan Smedby, Cheng Bian, Xin Yang 0009, Pheng-Ann Heng, Aliasghar Mortazi, Ulas Bagci, Guanyu Yang 0001, Chenchen Sun, Gaetan Galisot, Jean-Yves Ramel, Guang Yang 0006 |
Medical Image Anal. | 14 |
| 2019 | RETOUCH: The Retinal OCT Fluid Detection and Segmentation Benchmark and ChallengeabstractRetinal swelling due to the accumulation of fluid is associated with the most vision-threatening retinal diseases. Optical coherence tomography (OCT) is the current standard of care in assessing the presence and quantity of retinal fluid and image-guided treatment management. Deep learning methods have made their impact across medical imaging, and many retinal OCT analysis methods have been proposed. However, it is currently not clear how successful they are in interpreting the retinal fluid on OCT, which is due to the lack of standardized benchmarks. To address this, we organized a challenge RETOUCH in conjunction with MICCAI 2017, with eight teams participating. The challenge consisted of two tasks: fluid detection and fluid segmentation. It featured for the first time: all three retinal fluid types, with annotated images provided by two clinical centers, which were acquired with the three most common OCT device vendors from patients with two different retinal diseases. The analysis revealed that in the detection task, the performance on the automated fluid detection was within the inter-grader variability. However, in the segmentation task, fusing the automated methods produced segmentations that were superior to all individual methods, indicating the need for further improvements in the segmentation performance. Hrvoje Bogunovic, Freerk G. Venhuizen, Sophie Riedl 0001, Stefanos Apostolopoulos, Alireza Bab-Hadiashar, Ulas Bagci, Mirza Faisal Beg, Loza Bekalo, Qiang Chen 0004, Carlos Ciller, Karthik Gopinath, Amirali Khodadadian Gostar, Kiwan Jeon, Zexuan Ji, Sung Ho Kang, Dara Koozekanani, Donghuan Lu, Dustin Morley, Keshab K. Parhi, Hyoung Suk Park, Abdolreza Rashno, Marinko Sarunic, Saad Shaikh, Jayanthi Sivaswamy, Ruwan B. Tennakoon, Shivin Yadav, Sandro De Zanet, Sebastian M. Waldstein, Bianca S. Gerendas, Caroline C. W. Klaver, Clara I. Sánchez, Ursula Schmidt-Erfurth |
IEEE Trans. Medical Imaging | 6 |
| 2019 | Lung and Pancreatic Tumor Characterization in the Deep Learning Era: Novel Supervised and Unsupervised Learning ApproachesabstractRisk stratification (characterization) of tumors from radiology images can be more accurate and faster with computer-aided diagnosis (CAD) tools. Tumor characterization through such tools can also enable non-invasive cancer staging, prognosis, and foster personalized treatment planning as a part of precision medicine. In this papet, we propose both supervised and unsupervised machine learning strategies to improve tumor characterization. Our first approach is based on supervised learning for which we demonstrate significant gains with deep learning algorithms, particularly by utilizing a 3D convolutional neural network and transfer learning. Motivated by the radiologists' interpretations of the scans, we then show how to incorporate task-dependent feature representations into a CAD system via a graph-regularized sparse multi-task learning framework. In the second approach, we explore an unsupervised learning algorithm to address the limited availability of labeled training data, a common problem in medical imaging applications. Inspired by learning from label proportion approaches in computer vision, we propose to use proportion-support vector machine for characterizing tumors. We also seek the answer to the fundamental question about the goodness of "deep features" for unsupervised tumor classification. We evaluate our proposed supervised and unsupervised learning algorithms on two different tumor diagnosis challenges: lung and pancreas with 1018 CT and 171 MRI scans, respectively, and obtain the state-of-the-art sensitivity and specificity results in both problems. Sarfaraz Hussein, Pujan Kandel, Candice W. Bolan, Michael B. Wallace, Ulas Bagci |
IEEE Trans. Medical Imaging | 5 |
| 2019 | Deep Geodesic Learning for Segmentation and Anatomical LandmarkingabstractIn this paper, we propose a novel deep learning framework for anatomy segmentation and automatic landmarking. Specifically, we focus on the challenging problem of mandible segmentation from cone-beam computed tomography (CBCT) scans and identification of 9 anatomical landmarks of the mandible on the geodesic space. The overall approach employs three inter-related steps. In the first step, we propose a deep neural network architecture with carefully designed regularization, and network hyper-parameters to perform image segmentation without the need for data augmentation and complex post-processing refinement. In the second step, we formulate the landmark localization problem directly on the geodesic space for sparsely-spaced anatomical landmarks. In the third step, we utilize a long short-term memory network to identify the closely-spaced landmarks, which is rather difficult to obtain using other standard networks. The proposed fully automated method showed superior efficacy compared to the state-of-the-art mandible segmentation and landmarking approaches in craniofacial anomalies and diseased states. We used a very challenging CBCT data set of 50 patients with a high-degree of craniomaxillofacial variability that is realistic in clinical practice. The qualitative visual inspection was conducted for distinct CBCT scans from 250 patients with high anatomical variability. We have also shown the state-of-the-art performance in an independent data set from the MICCAI Head-Neck Challenge (2015). Neslisah Torosdagli, Denise K. Liberton, Payal Verma, Murat Sincan, Janice Lee, Ulas Bagci |
IEEE Trans. Medical Imaging | 6 |
| 2018 | S4ND: Single-Shot Single-Scale Lung Nodule Detection
Naji Khosravan, Ulas Bagci |
MICCAI (2) | 2 |
| 2018 | Joint solution for PET image segmentation, denoising, and partial volume correction
Ziyue Xu 0001, Mingchen Gao, Georgios Z. Papadakis, Brian Luna, Sanjay Jain 0002, Daniel J. Mollura, Ulas Bagci |
Medical Image Anal. | 7 |
| 2017 | CardiacNET: Segmentation of Left Atrium and Proximal Pulmonary Veins from MRI Using Multi-view CNN
Aliasghar Mortazi, Rashed Karim, Kawal S. Rhode, Jeremy Burt, Ulas Bagci |
MICCAI (2) | 5 |
| 2017 | Automatic response assessment in regions of language cortex in epilepsy patients using ECoG-based functional mapping and machine learningabstractAccurate localization of brain regions responsible for language and cognitive functions in Epilepsy patients should be carefully determined prior to surgery. Electrocorticography (ECoG)-based Real Time Functional Mapping (RTFM) has been shown to be a safer alternative to the electrical cortical stimulation mapping (ESM), which is currently the clinical/gold standard. Conventional methods for analyzing RTFM signals are based on statistical comparison of signal power at certain frequency bands. Compared to gold standard (ESM), they have limited accuracies when assessing channel responses. In this study, we address the accuracy limitation of the current RTFM signal estimation methods by analyzing the full frequency spectrum of the signal and replacing signal power estimation methods with machine learning algorithms, specifically random forest (RF), as a proof of concept. We train RF with power spectral density of the time-series RTFM signal in supervised learning framework where ground truth labels are obtained from the ESM. Results obtained from RTFM of six adult patients in a strictly controlled experimental setup reveal the state of the art detection accuracy of ≈78% for the language comprehension task, an improvement of 23% over the conventional RTFM estimation method. To the best of our knowledge, this is the first study exploring the use of machine learning approaches for determining RTFM signal characteristics, and using the whole-frequency band for better region localization. Our results demonstrate the feasibility of machine learning based RTFM signal analysis method over the full spectrum to be a clinical routine in the near future. Harish RaviPrakash, Milena Korostenskaja, Ki Lee, James Baumgartner 0001, Eduardo Castillo, Ulas Bagci |
SMC | 6 |
| 2017 | CorteXpert: A model-based method for automatic renal cortex segmentation
Dehui Xiang, Ulas Bagci, Weifang Zhu, Jianhua Yao 0001, Milan Sonka, Xinjian Chen 0001 |
Medical Image Anal. | 2 |
| 2017 | Single-Channel Sparse Non-Negative Blind Source Separation Method for Automatic 3-D Delineation of Lung Tumor in PET ImagesabstractIn this paper, we propose a novel method for single-channel blind separation of nonoverlapped sources and, to the best of our knowledge, apply it for the first time to automatic segmentation of lung tumors in positron emission tomography (PET) images. Our approach first converts a 3-D PET image into a pseudo-multichannel image. Afterward, regularization free sparseness constrained non-negative matrix factorization is used to separate tumor from other tissues. By using complexity based criterion, we select tumor component as the one with minimal complexity. We have compared the proposed method with threshold based on 40% and 50% maximum standardized uptake value (SUV), graph cuts (GC), random walks (RW), and affinity propagation (AP) algorithms on 18 nonsmall cell lung cancer datasets with respect to ground truth (GT) provided by two radiologists. Dice similarity coefficient averaged with respect to two GTs is: 0.78 ± 0.12 by the proposed algorithm, 0.78 ± 0.1 by GC, 0.77 ± 0.13 by AP, 0.77 ± 0.07 by RW, and 0.75 ± 0.13 by 50% maximum SUV threshold. Since the proposed method achieved performance comparable with interactive methods, considering the unique challenges of lung tumor segmentation from PET images, our findings support possibility of using our fully automated method in routine clinics. The source codes will be available at www.mipav.net/English/research/research.html. Ivica Kopriva, Wei Ju 0002, Bin Zhang 0049, Dehui Xiang, Kai Yu 0009, Ximing Wang, Ulas Bagci, Xinjian Chen 0001 |
IEEE J. Biomed. Health Informatics | 8 |
| 2017 | Automatic Segmentation and Quantification of White and Brown Adipose Tissues from PET/CT ScansabstractIn this paper, we investigate the automatic detection of white and brown adipose tissues using Positron Emission Tomography/Computed Tomography (PET/CT) scans, and develop methods for the quantification of these tissues at the whole-body and body-region levels. We propose a patient-specific automatic adiposity analysis system with two modules. In the first module, we detect white adipose tissue (WAT) and its two sub-types from CT scans: Visceral Adipose Tissue (VAT) and Subcutaneous Adipose Tissue (SAT). This process relies conventionally on manual or semi-automated segmentation, leading to inefficient solutions. Our novel framework addresses this challenge by proposing an unsupervised learning method to separate VAT from SAT in the abdominal region for the clinical quantification of central obesity. This step is followed by a context driven label fusion algorithm through sparse 3D Conditional Random Fields (CRF) for volumetric adiposity analysis. In the second module, we automatically detect, segment, and quantify brown adipose tissue (BAT) using PET scans because unlike WAT, BAT is metabolically active. After identifying BAT regions using PET, we perform a co-segmentation procedure utilizing asymmetric complementary information from PET and CT. Finally, we present a new probabilistic distance metric for differentiating BAT from non-BAT regions. Both modules are integrated via an automatic body-region detection unit based on one-shot learning. Experimental evaluations conducted on 151 PET/CT scans achieve state-of-the-art performances in both central obesity as well as brown adiposity quantification. Sarfaraz Hussein, Aileen Green, Arjun Watane, David A. Reiter, Xinjian Chen 0001, Georgios Z. Papadakis, Bradford J. Wood, Aaron Cypess, Medhat M. Osman, Ulas Bagci |
IEEE Trans. Medical Imaging | 10 |
| 2016 | Characterization of Lung Nodule Malignancy Using Hybrid Shape and Appearance Features
Mario Buty, Ziyue Xu 0001, Mingchen Gao, Ulas Bagci, Aaron Wu, Daniel J. Mollura |
MICCAI (1) | 4 |
| 2015 | A hybrid method for airway segmentation and automated measurement of bronchial wall thickness on CT
Ziyue Xu 0001, Ulas Bagci, Brent Foster, Awais Mansoor, Jayaram K. Udupa, Daniel J. Mollura |
Medical Image Anal. | 2 |
| 2015 | Correction to "A Generic Approach to Pathological Lung Segmentation"abstractIn the above paper (ibid., vol. 33, no. 12, pp. 2293-2310, Dec. 2014), Awais Mansoor was incorrectly indicated as the corresponding author. Ulas Bagci should have been indicated as the corresponding author. Awais Mansoor, Ulas Bagci, Ziyue Xu 0001, Brent Foster, Kenneth N. Olivier, Jason M. Elinoff, Anthony F. Suffredini, Jayaram K. Udupa, Daniel J. Mollura |
IEEE Trans. Medical Imaging | 2 |
| 2014 | Optimally Stabilized PET Image Denoising Using Trilateral Filtering
Awais Mansoor, Ulas Bagci, Daniel J. Mollura |
MICCAI (1) | 2 |
| 2014 | Segmentation Based Denoising of PET Images: An Iterative Approach via Regional Means and Affinity Propagation
Ziyue Xu 0001, Ulas Bagci, Jurgen Seidel, David Thomasson, Jeffrey M. Solomon, Daniel J. Mollura |
MICCAI (1) | 2 |
| 2014 | A Generic Approach to Pathological Lung SegmentationabstractIn this study, we propose a novel pathological lung segmentation method that takes into account neighbor prior constraints and a novel pathology recognition system. Our proposed framework has two stages; during stage one, we adapted the fuzzy connectedness (FC) image segmentation algorithm to perform initial lung parenchyma extraction. In parallel, we estimate the lung volume using rib-cage information without explicitly delineating lungs. This rudimentary, but intelligent lung volume estimation system allows comparison of volume differences between rib cage and FC based lung volume measurements. Significant volume difference indicates the presence of pathology, which invokes the second stage of the proposed framework for the refinement of segmented lung. In stage two, texture-based features are utilized to detect abnormal imaging patterns (consolidations, ground glass, interstitial thickening, tree-inbud, honeycombing, nodules, and micro-nodules) that might have been missed during the first stage of the algorithm. This refinement stage is further completed by a novel neighboring anatomy-guided segmentation approach to include abnormalities with weak textures, and pleura regions. We evaluated the accuracy and efficiency of the proposed method on more than 400 CT scans with the presence of a wide spectrum of abnormalities. To our best of knowledge, this is the first study to evaluate all abnormal imaging patterns in a single segmentation framework. The quantitative results show that our pathological lung segmentation method improves on current standards because of its high sensitivity and specificity and may have considerable potential to enhance the performance of routine clinical tasks. Awais Mansoor, Ulas Bagci, Ziyue Xu 0001, Brent Foster, Kenneth N. Olivier, Jason M. Elinoff, Anthony F. Suffredini, Jayaram K. Udupa, Daniel J. Mollura |
IEEE Trans. Medical Imaging | 2 |
| 2013 | Denoising PET Images Using Singular Value Thresholding and Stein's Unbiased Risk Estimate
Ulas Bagci, Daniel J. Mollura |
MICCAI (3) | 1 |
| 2013 | Spatially Constrained Random Walk Approach for Accurate Estimation of Airway Wall Surfaces
Ziyue Xu 0001, Ulas Bagci, Brent Foster, Awais Mansoor, Daniel J. Mollura |
MICCAI (2) | 2 |
| 2013 | Joint segmentation of anatomical and functional images: Applications in quantification of lesions from PET, PET-CT, MRI-PET, and MRI-PET-CT images
Ulas Bagci, Jayaram K. Udupa, Neil Mendhiratta, Brent Foster, Ziyue Xu 0001, Jianhua Yao 0001, Xinjian Chen 0001, Daniel J. Mollura |
Medical Image Anal. | 1 |
| 2012 | Co-segmentation of Functional and Anatomical Images
Ulas Bagci, Jayaram K. Udupa, Jianhua Yao 0001, Daniel J. Mollura |
MICCAI (3) | 1 |
| 2012 | Medical Image Segmentation by Combining Graph Cuts and Oriented Active Appearance ModelsabstractIn this paper, we propose a novel method based on a strategic combination of the active appearance model (AAM), live wire (LW), and graph cuts (GCs) for abdominal 3-D organ segmentation. The proposed method consists of three main parts: model building, object recognition, and delineation. In the model building part, we construct the AAM and train the LW cost function and GC parameters. In the recognition part, a novel algorithm is proposed for improving the conventional AAM matching method, which effectively combines the AAM and LW methods, resulting in the oriented AAM (OAAM). A multiobject strategy is utilized to help in object initialization. We employ a pseudo-3-D initialization strategy and segment the organs slice by slice via a multiobject OAAM method. For the object delineation part, a 3-D shape-constrained GC method is proposed. The object shape generated from the initialization step is integrated into the GC cost computation, and an iterative GC-OAAM method is used for object delineation. The proposed method was tested in segmenting the liver, kidneys, and spleen on a clinical CT data set and also on the MICCAI 2007 Grand Challenge liver data set. The results show the following: 1) The overall segmentation accuracy of true positive volume fraction TPVF > 94.3% and false positive volume fraction can be achieved; 2) the initialization performance can be improved by combining the AAM and LW; 3) the multiobject strategy greatly facilitates initialization; 4) compared with the traditional 3-D AAM method, the pseudo-3-D OAAM method achieves comparable performance while running 12 times faster; and 5) the performance of the proposed method is comparable to state-of-the-art liver segmentation algorithm. The executable version of the 3-D shape-constrained GC method with a user interface can be downloaded from http://xinjianchen.wordpress.com/research/. Xinjian Chen 0001, Jayaram K. Udupa, Ulas Bagci, Ying Zhuge, Jianhua Yao 0001 |
IEEE Trans. Image Process. | 3 |
| 2012 | Hierarchical Scale-Based Multiobject Recognition of 3-D Anatomical StructuresabstractSegmentation of anatomical structures from medical images is a challenging problem, which depends on the accurate recognition (localization) of anatomical structures prior to delineation. This study generalizes anatomy segmentation problem via attacking two major challenges: 1) automatically locating anatomical structures without doing search or optimization, and 2) automatically delineating the anatomical structures based on the located model assembly. For 1), we propose intensity weighted ball-scale object extraction concept to build a hierarchical transfer function from image space to object (shape) space such that anatomical structures in 3-D medical images can be recognized without the need to perform search or optimization. For 2), we integrate the graph-cut (GC) segmentation algorithm with prior shape model. This integrated segmentation framework is evaluated on clinical 3-D images consisting of a set of 20 abdominal CT scans. In addition, we use a set of 11 foot MR images to test the generalizability of our method to the different imaging modalities as well as robustness and accuracy of the proposed methodology. Since MR image intensities do not possess a tissue specific numeric meaning, we also explore the effects of intensity nonstandardness on anatomical object recognition. Experimental results indicate that: 1) effective recognition can make the delineation more accurate; 2) incorporating a large number of anatomical structures via a model assembly in the shape model improves the recognition and delineation accuracy dramatically; 3) ball-scale yields useful information about the relationship between the objects and the image; 4) intensity variation among scenes in an ensemble degrades object recognition performance. Ulas Bagci, Xinjian Chen 0001, Jayaram K. Udupa |
IEEE Trans. Medical Imaging | 1 |
| 2011 | A New Prior Shape Model for Level Set Segmentation
Poay Hoon Lim, Ulas Bagci, Li Bai 0001 |
CIARP | 2 |
| 2011 | Learning Shape and Texture Characteristics of CT Tree-in-Bud Opacities for CAD Systems
Ulas Bagci, Jianhua Yao 0001, Jesus J. Caban, Anthony F. Suffredini, Tara N. Palmore, Daniel J. Mollura |
MICCAI (3) | 1 |
| 2010 | 3D automatic anatomy segmentation based on graph cut-oriented active appearance modelsabstractIn this paper, we propose a novel 3D automatic anatomy segmentation method based on the synergistic combination of active appearance models (AAM), live wire (LW) and graph cut (GC). The proposed method consists of three main parts: model building, initialization and segmentation. For the model building part, an AAM model is constructed and the LW cost function is trained. For the initialization part, an improved iterative model refinement algorithm is proposed for the AAM optimization, which synergistically combines the AAM and LW method (OAAM). And a multi-object strategy is applied to help the object initialization. A pseudo 3D initialization strategy is employed to segment the organs slice by slice via multi-object OAAM method. The model constraints are applied to the initialization result. For the segmentation part, the object shape information generated from the initialization step is integrated into the GC cost computation. And an iterative GCOAAM method is proposed for object delineation. This method is a general method and can be applied to any organ segmentation. The proposed method was tested on the clinical liver and kidney CT data sets. The results showed the following: (a) an overall segmentation accuracy of true positive fraction>93.5%, and false positive fraction<0.2% can be achieved. (b) The initialization performance is improved by combining the AAM and LW. (c) The multi-object strategy greatly helps the initialization due to inter-object constraints. Xinjian Chen 0001, Jianhua Yao 0001, Ying Zhuge, Ulas Bagci |
ICIP | 4 |
| 2010 | The role of intensity standardization in medical image registration
Ulas Bagci, Jayaram K. Udupa, Li Bai 0001 |
Pattern Recognit. Lett. | 1 |
| 2010 | Automatic Best Reference Slice Selection for Smooth Volume Reconstruction of a Mouse Brain From Histological ImagesabstractIn this paper, we present a novel and effective method for registering histological slices of a mouse brain to reconstruct a 3-D volume. First, intensity variations in images are corrected through an intensity standardization process so that intensity values remain constant across slices. Second, the image space is transformed to a feature space where continuous variables are taken as high fidelity image features for accurate registration. Third, in order to improve the quality of the reconstructed volume, an automatic best reference slice selection algorithm is developed based on iterative assessment of image entropy and mean square error of the registration process. Fourth, a novel metric for evaluating the quality of the reconstructed volume is developed. Finally, the effect of optimal reference slice selection on the quality of registration and subsequent reconstruction is demonstrated. Ulas Bagci, Li Bai 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2008 | Parallel AdaBoost algorithm for Gabor wavelet selection in face recognitionabstractIn this paper, the problem of automatic Gabor wavelet selection for face recognition is tackled by introducing an automatic algorithm based on Parallel AdaBoosting method. Incorporating mutual information into the algorithm leads to the selection procedure not only based on classification accuracy but also on efficiency. Effective image features are selected by using properly chosen Gabor wavelets optimised with Parallel AdaBoost method and mutual information to get high recognition rates with low computational cost. Experiments are conducted using the well-known FERET face database. In proposed framework, memory and computation costs are reduced significantly and high classification accuracy is obtained. Ulas Bagci, Li Bai 0001 |
ICIP | 1 |
| 2007 | Automatic Classification of Musical Genres Using Inter-Genre SimilarityabstractMusical genre classification is an essential tool for music information retrieval systems and it has potential to become a highly demanded application in various media platforms. Two important problems of the automatic musical genre classification are feature extraction and classifier design. In this letter, we propose two novel classifiers using inter-genre similarity (IGS) modeling and investigate the use of dynamic timbral texture features in order to improve automatic musical genre classification performance. Inter-genre similarity is modeled over hard-to-classify samples of the musical genre feature space. In the classification, samples within inter-genre similarity class are eliminated to reduce inter-genre confusion and to improve genre classification performance. Experimental results show that the proposed classifiers provide better classification rates than the existing methods. Ulas Bagci, Engin Erzin |
IEEE Signal Process. Lett. | 1 |