Mehmet Saygin Seyfioglu

dblp:153/9328 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Vision and language · 59% Information extraction and text analysis · 24% Multi-agent systems · 13%
Interdisciplinary, comprehensive, and emerging computing
4 papers
Medical and health informatics · 100%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Medical and health informatics
computational pathology
2.332025
PathFinder: A Multi-Modal Multi-Agent System for Medical Diagnostic Decision-Making Applied to Histopathology · ICCV 2025
Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos · CVPR 2024
Quilt-1M: One Million Image-Text Pairs for Histopathology · NeurIPS 2023
Computer vision › Vision and language › vision-language model › domain-specific vision-language model
medical vision-language model
1.522025
MedicalNarratives: Connecting Medical Vision and Language with Localized Narratives · NeurIPS 2025
Quilt-1M: One Million Image-Text Pairs for Histopathology · NeurIPS 2023
Computer vision › Vision and language
vision-language pretraining
1.522025
MedicalNarratives: Connecting Medical Vision and Language with Localized Narratives · NeurIPS 2025
Quilt-1M: One Million Image-Text Pairs for Histopathology · NeurIPS 2023
Knowledge, reasoning and agents › Multi-agent systems
multi-agent decision making
0.912025
PathFinder: A Multi-Modal Multi-Agent System for Medical Diagnostic Decision-Making Applied to Histopathology · ICCV 2025
Medical and health informatics › medical imaging
medical image analysis
0.912025
MedicalNarratives: Connecting Medical Vision and Language with Localized Narratives · NeurIPS 2025
Computer vision › Vision and language › vision-language model › multimodal large language model
visual instruction tuning
0.812024
Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos · CVPR 2024
Medical and health informatics › computational pathology › histopathology image analysis
whole slide image analysis
0.812024
Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos · CVPR 2024
Machine learning › Trustworthy machine learning › interpretability
natural language explanation
0.312025
PathFinder: A Multi-Modal Multi-Agent System for Medical Diagnostic Decision-Making Applied to Histopathology · ICCV 2025

Methods — techniques the papers use, named apart from their topics

CLIP · 3.1large language model · 2.8transformer models · 1.7multiple instance learning · 1.7multi-agent framework · 1.7localized narratives · 1.7visual instruction tuning · 1.5automatic speech recognition · 1.3
YearPublicationVenuePosition
2025 PathFinder: A Multi-Modal Multi-Agent System for Medical Diagnostic Decision-Making Applied to Histopathology
abstract
Diagnosing diseases through histopathology whole slide images (WSIs) is fundamental in modern pathology but is challenged by the gigapixel scale and complexity of WSIs. Trained histopathologists overcome this challenge by navigating the WSI, looking for relevant patches, taking notes, and compiling them to produce a final holistic diagnostic. Traditional AI approaches, such as multiple instance learning and transformer-based models, fail short of such a holistic, iterative, multi-scale diagnostic procedure, limiting their adoption in the real-world. We introduce PathFinder, a multi-modal, multi-agent framework that emulates the decision-making process of expert pathologists. PathFinder integrates four AI agents, the Triage Agent, Navigation Agent, Description Agent, and Diagnosis Agent, that collaboratively navigate WSIs, gather evidence, and provide comprehensive diagnoses with natural language explanations. The Triage Agent classifies the WSI as benign or risky; if risky, the Navigation and Description Agents iteratively focus on significant regions, generating importance maps and descriptive insights of sampled patches. Finally, the Diagnosis Agent synthesizes the findings to determine the patient's diagnostic classification. Our Experiments show that PathFinder outperforms state-of-the-art methods in skin melanoma diagnosis by 8% while offering inherent explainability through natural language descriptions of diagnostically relevant patches. Qualitative analysis by pathologists shows that the Description Agent's outputs are of high quality and comparable to GPT-4o. PathFinder is also the first AI-based system to surpass the average performance of pathologists in this challenging melanoma classification task by 9%, setting a new record for efficient, accurate, and interpretable AI-assisted diagnostics in pathology. Data, code and models available at https://pathfinder-dx.github.io/
Fatemeh Ghezloo, Mehmet Saygin Seyfioglu, Rustin Soraki, Wisdom Oluchi Ikezogwo, Beibin Li, Tejoram Vivekanandan, Joann G. Elmore, Ranjay Krishna, Linda G. Shapiro
ICCV2
2025 MedicalNarratives: Connecting Medical Vision and Language with Localized Narratives
abstract
Multi-modal models are data hungry. While datasets with natural images are abundant, medical image datasets can not afford the same luxury. To enable representation learning for medical images at scale, we turn to YouTube, a platform with a large reservoir of open-source medical pedagogical videos. We curate MedicalNarratives, a dataset 4.7M medical image-text pairs, with 1M samples containing dense annotations in the form of traces spatial traces (and bounding boxes), and 118K videos centered on the trace event (with aligned text), enabling spatiotemporal grounding beyond single frames. Similar to think-aloud studies where instructors speak while hovering their mouse cursor movements over relevant image regions, 1M images in MedicalNarratives contains localized mouse traces in image pixels, creating a spatial association between the text and pixels. To evaluate the utility of MedicalNarratives, we train GenMedClip with a CLIP-like objective using our dataset spanning 12 medical domains. GenMedClip outperforms previous state-of-the-art models on all 12 domains on a newly constructed medical imaging benchmark. Data, demo, code, and models will be made available.
Wisdom Oluchi Ikezogwo, Kevin M. Zhang, Mehmet Saygin Seyfioglu
NeurIPS3
2024 Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos
abstract
Diagnosis in histopathology requires a global whole slide images (WSIs) analysis, requiring pathologists to compound evidence from different WSI patches. The gigapixel scale of WSIs poses a challenge for histopathology multimodal models. Training multi-model models for histopathology requires instruction tuning datasets, which currently contain information for individual image patches, without a spatial grounding of the concepts within each patch and without a wider view of the WSI. To bridge this gap, we introduce QUILT-INSTRUCT, a large-scale dataset of107, 131 histopathology-specific instruction question/answer pairs, grounded within diagnostically relevant image patches that make up the WSI. Our dataset is collected by leveraging educational histopathology videos from YouTube, which provides spatial localization of narrations by automatically extracting the narrators' cursor positions. QUILT-INSTRUCT supports contextual reasoning by extracting diagnosis and supporting facts from the entire WSI. Using QUILT-INSTRUCT, we train QUILT-LLAVA, which can reason beyond the given single image patch, enabling diagnostic reasoning across patches. To evaluate QUILT-LLAVA, we propose a compre-hensive evaluation dataset created from 985 images and 1283 human-generated question-answers. We also thor-oughly evaluate QUILT-LLAVA using public histopathology datasets, where QUILT-LLAVA significantly outperforms SOTA by over 10% on relative GPT-4 score and 4% and 9% on open and closed set VQA11Our code, data, and model is publicly accessible at quilt-llava.github.io..
Mehmet Saygin Seyfioglu, Wisdom Oluchi Ikezogwo, Fatemeh Ghezloo, Ranjay Krishna, Linda G. Shapiro
CVPR1
2023 Quilt-1M: One Million Image-Text Pairs for Histopathology
abstract
Recent accelerations in multi-modal applications have been made possible with the plethora of image and text data available online. However, the scarcity of analogous data in the medical field, specifically in histopathology, has slowed comparable progress. To enable similar representation learning for histopathology, we turn to YouTube, an untapped resource of videos, offering $1,087$ hours of valuable educational histopathology videos from expert clinicians.From YouTube, we curate QUILT: a large-scale vision-language dataset consisting of $802, 144$ image and text pairs.QUILT was automatically curated using a mixture of models, including large language models, handcrafted algorithms, human knowledge databases, and automatic speech recognition.In comparison, the most comprehensive datasets curated for histopathology amass only around $200$K samples.We combine QUILT with datasets from other sources, including Twitter, research papers, and the internet in general, to create an even larger dataset: QUILT-1M, with $1$M paired image-text samples, marking it as the largest vision-language histopathology dataset to date. We demonstrate the value of QUILT-1M by fine-tuning a pre-trained CLIP model. Our model outperforms state-of-the-art models on both zero-shot and linear probing tasks for classifying new histopathology images across $13$ diverse patch-level datasets of $8$ different sub-pathologies and cross-modal retrieval tasks.
Wisdom Oluchi Ikezogwo, Mehmet Saygin Seyfioglu, Fatemeh Ghezloo, Dylan Stefan Chan Geva, Fatwir Sheikh Mohammed, Pavan Kumar Anand, Ranjay Krishna, Linda G. Shapiro
NeurIPS2
2022 Brain-Aware Replacements for Supervised Contrastive Learning in Detection of Alzheimer's Disease
Mehmet Saygin Seyfioglu, Zixuan Liu 0001, Pranav Kamath, Sadjyot Gangolli, Sheng Wang 0012, Thomas J. Grabowski, Linda G. Shapiro
MICCAI (1)1
2020 Leveraging Unlabeled Data for Glioma Molecular Subtype and Survival Prediction
abstract
In this paper, we address two long-standing radio-genomic challenges in glioma subtype and survival prediction: (1) how to leverage large amounts of unlabeled magnetic resonance (MR) imaging data and (2) how to unite MR data and genomic data. We propose a novel application of multi-task learning (MTL) that leverages unlabeled MR data by jointly learning an auxiliary tumor segmentation task with glioma subtype prediction and that can learn from patients with and without genomic data. We analyze multi-parametric MR data from 542 patients in the combined training, validation, and testing sets of the 2018 Multimodal Brain Tumor Segmentation Challenge and somatic copy number alteration (SCNA) data from 1090 patients in The Cancer Genome Atlas' (TCGA) lower-grade glioma and glioblastoma projects. Our MTL model significantly outperforms comparable classification models trained only on labeled MR data for both IDH1/2 mutation and 1p/19q co-deletion subtype prediction tasks. We also show that embeddings produced by our MTL models improve survival predictions beyond MR or SCNA on their own. Our code is available at https://github.com/nknuecht/glioma_mtl.
Nicholas Nuechterlein, Beibin Li, Mehmet Saygin Seyfioglu, Sachin Mehta, Patrick J. Cimino, Linda G. Shapiro
ICPR3
2017 A Hierarchical Approach for Sentiment Analysis and Categorization of Turkish Written Customer Relationship Management Data
abstract
Today, large scale companies are receiving tens of thousands of feedback from their customers every day, which makes it impossible for them to evaluate the feedbacks manually.As sentiments expressed by the customers are vitally important for companies, an accurate and swift analysis is needed.In this paper, a hierarchical approach is proposed for sentiment analysis and further categorization of Turkish written customer feedback to a private airline company.First, the word embeddings of customer feedbacks are computed by using Word2Vec then averaged in proportion with the inverse of their frequency in the document.For binary sentiment analysis, i.e determination of 'positive' and 'negative' sentiments, an extreme gradient boosting (xgboost) classifier is trained on averaged review vectors and an overall accuracy of 92.5% is obtained which is 16.8% higher than that of the baseline model.For further categorization of negative sentiments in one of twelve pre determined classes, an xgboost classifier is trained upon document embeddings of negatively classified comments, which were calculated using Doc2Vec.An overall accuracy of 71.16% is obtained for the task of categorization of 12 different classes using the Doc2Vec approach, thereby yielding a classification accuracy 19.1% higher than that of the baseline model.
Mehmet Saygin Seyfioglu, Mustafa Umut Demirezen
FedCSIS1
2017 Deep Neural Network Initialization Methods for Micro-Doppler Classification With Low Training Sample Support
abstract
Deep neural networks (DNNs) require large-scale labeled data sets to prevent overfitting while having good generalization. In radar applications, however, acquiring a measured data set of the order of thousands is challenging due to constraints on manpower, cost, and other resources. In this letter, the efficacy of two neural network initialization techniques-unsupervised pretraining and transfer learning-for dealing with training DNNs on small data sets is compared. Unsupervised pretraining is implemented through the design of a convolutional autoencoder (CAE), while transfer learning from two popular convolutional neural network architectures (VGGNet and GoogleNet) is used to augment measured RF data for training. A 12-class problem for discrimination of micro-Doppler signatures for indoor human activities is utilized to analyze activation maps, bottleneck features, class model, and classification accuracy with respect to training sample size. Results show that on meager data sets, transfer learning outperforms unsupervised pretraining and random initialization by 10% and 25%, respectively, but that when the sample size exceeds 650, unsupervised pretraining surpasses transfer learning and random initialization by 5% and 10%, respectively. Visualization of activation layers and learned models reveals how the CAE succeeds in representing the micro-Doppler signature.
Mehmet Saygin Seyfioglu, Sevgi Zubeyde Gurbuz
IEEE Geosci. Remote. Sens. Lett.1
2015 Automatic spectral signature extraction for hyperspectral target detection
abstract
There are two main approaches to hyperspectral target detection: anomaly detection techniques, which detect outliers substantially different from the background, and spectral signature techniques, which require as an input a user-defined target signature. Oftentimes, however, the target signature may not be known, or there may be unexpected targets in the image, which are unknown but still of interest. As a result, algorithms that can automatically extract potential target signatures without any a priori knowledge are of great interest. In this work, a fusion-based algorithm is developed that takes advantage of both spatial and spectral information to automatically extract the spectral signatures of potential targets of interest. The performance of several target detection algorithms is compared for both the proposed spatially-spectrally estimated (SSE) target signature and the initial target signature used by the Automatic Target Detection and Classification Algorithm (ATDCA). It is shown that the SSE signature leads to improved automatic spectral target recognition (ATSR) performance than the ATDCA algorithm for the test conducted on the AVIRIS Indian Pines dataset.
Mehmet Saygin Seyfioglu, Seyma Bayindir, Sevgi Zubeyde Gurbuz
IGARSS1
2014 Airborne radar clutter simulation using hyperspectral and LiDAR imagery
abstract
A key factor degrading the performance of airborne radar is clutter. Many current clutter mitigation algorithms rely on statistical clutter models; however, such models fail to incorporate the time varying, site dependent nature of unwanted scattering. Hyperspectral imagery combined with lidar-derived elevation data offers a unique and comprehensive assessment of the radar clutter environment. Elevation data can be used to compute the range of scatters to the radar, while hyperspectral data processing can yield information on terrain features, which directly impacts electromagnetic backscattering. A simulated airborne clutter map is provided, and statistics compared to typical clutter distributions.
Mehmet Saygin Seyfioglu, Sevgi Zubeyde Gurbuz
IGARSS1