Sezgin Er

dblp:342/4699 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0001-7266-9844ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Vision and language · 48% Segmentation and scene understanding · 32% Generative modeling · 16%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Medical and health informatics · 100%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 11 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Medical and health informatics
medical imaging
1.722025
Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical Imaging · NeurIPS 2025
RadGPT: Constructing 3D Image-Text Tumor Datasets · ICCV 2025
Computer vision › Segmentation and scene understanding
medical image segmentation
0.912025
RadGPT: Constructing 3D Image-Text Tumor Datasets · ICCV 2025
Computer vision › Vision and language › vision-language model › domain-specific vision-language model
medical vision-language
0.912025
RadGPT: Constructing 3D Image-Text Tumor Datasets · ICCV 2025
Computer vision › Vision and language › vision-language model › domain-specific vision-language model
medical vision-language model
0.912025
Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical Imaging · NeurIPS 2025
Computer vision › Vision and language › medical report generation
radiology report generation
0.912025
RadGPT: Constructing 3D Image-Text Tumor Datasets · ICCV 2025
Machine learning › Generative modeling › image generation › conditional image synthesis
text-conditioned image synthesis
0.912025
Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical Imaging · NeurIPS 2025
Computer vision › Segmentation and scene understanding › medical image segmentation
tumor segmentation
0.912025
RadGPT: Constructing 3D Image-Text Tumor Datasets · ICCV 2025
Medical and health informatics › medical imaging
3d medical imaging
0.912025
Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical Imaging · NeurIPS 2025
Medical and health informatics › medical report generation
CT report generation
0.912025
RadGPT: Constructing 3D Image-Text Tumor Datasets · ICCV 2025
Natural language and speech › Language models and text generation › text generation › data-to-text generation
report generation
0.312025
Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical Imaging · NeurIPS 2025
Medical and health informatics › medical imaging
medical image synthesis
0.212024
GenerateCT: Text-Conditional Generation of 3D Chest CT Volumes · ECCV (79) 2024

Methods — techniques the papers use, named apart from their topics

vision-language model · 1.7tokenization · 1.7causal convolutional encoder-decoder · 1.7anatomy-aware segmentation · 1.7contrastive pretraining · 0.9contrastive pre-training · 0.9
YearPublicationVenuePosition
2025 RadGPT: Constructing 3D Image-Text Tumor Datasets
abstract
With over 85 million CT scans performed annually in the United States, creating tumor-related reports is a challenging and time-consuming task for radiologists. To address this need, we present RadGPT, an Anatomy-Aware Vision-Language AI Agent for generating detailed reports from CT scans. RadGPT first segments tumors, including benign cysts and malignant tumors, and their surrounding anatomical structures, then transforms this information into both structured reports and narrative reports. These reports provide tumor size, shape, location, attenuation, volume, and interactions with surrounding blood vessels and organs. Extensive evaluation on unseen hospitals shows that RadGPT can produce accurate reports, with high sensitivity/specificity for small tumor (<2 cm) detection: 80/73% for liver tumors, 92/78% for kidney tumors, and 77/77% for pancreatic tumors. For large tumors, sensitivity ranges from 89% to 97%. The results significantly surpass the state-of-the-art in abdominal CT report generation. RadGPT generated reports for 17 public datasets. Through radiologist review and refinement, we have ensured the reports' accuracy, and created the first publicly available image-text 3D medical dataset, comprising over 1.8 million text tokens and 2.7 million images from 9,262 CT scans, including 2,947 tumor scans/reports of 8,562 tumor instances. Our reports can: (1) localize tumors in eight liver sub-segments and three pancreatic sub-segments annotated per-voxel; (2) determine pancreatic tumor stage (T1-T4) in 260 reports; and (3) present individual analyses of multiple tumors--rare in human-made reports. Importantly, 948 of the reports are for early-stage tumors.
Pedro R. A. S. Bassi, Mehmet Can Yavuz, Ibrahim Ethem Hamamci, Sezgin Er, Xiaoxi Chen, Bjoern Menze, Sergio Decherchi, Andrea Cavalli, Kang Wang 0016, Yang Yang 0009, Alan L. Yuille, Zongwei Zhou
ICCV4
2025 Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical Imaging
abstract
Recent progress in vision-language modeling for 3D medical imaging has been fueled by large-scale computed tomography (CT) corpora with paired free-text reports, stronger architectures, and powerful pretrained models. This has enabled applications such as automated report generation and text-conditioned 3D image synthesis. Yet, current approaches struggle with high-resolution, long-sequence volumes: contrastive pretraining often yields vision encoders that are misaligned with clinical language, and slice-wise tokenization blurs fine anatomy, reducing diagnostic performance on downstream tasks. We introduce BTB3D (Better Tokens for Better 3D), a causal convolutional encoder-decoder that unifies 2D and 3D training and inference while producing compact, frequency-aware volumetric tokens. A three-stage training curriculum enables (i) local reconstruction, (ii) overlapping-window tiling, and (iii) long-context decoder refinement, during which the model learns from short slice excerpts yet generalizes to scans exceeding $300$ slices without additional memory overhead. BTB3D sets a new state-of-the-art on two key tasks: it improves BLEU scores and increases clinical F1 by 40\% over CT2Rep, CT-CHAT, and Merlin for report generation; and it reduces FID by 75\% and halves FVD compared to GenerateCT and MedSyn for text-to-CT synthesis, producing anatomically consistent $512\times512\times241$ volumes. These results confirm that precise three-dimensional tokenization, rather than larger language backbones alone, is essential for scalable vision-language modeling in 3D medical imaging. The codebase is available at: https://github.com/ibrahimethemhamamci/BTB3D
Ibrahim Ethem Hamamci, Sezgin Er, Suprosanna Shit, Hadrien Reynaud, Dong Yang 0005, Marc Edgar, Daguang Xu, Bernhard Kainz, Bjoern Menze
NeurIPS2
2024 GenerateCT: Text-Conditional Generation of 3D Chest CT Volumes
Ibrahim Ethem Hamamci, Sezgin Er, Anjany Sekuboyina, Enis Simsar, Alperen Tezcan, Ayse Gulnihan Simsek, Sevval Nil Esirgun, Furkan Almas, Irem Dogan, Muhammed Furkan Dasdelen, Chinmay Prabhakar, Hadrien Reynaud, Sarthak Pati, Christian Bluethgen, Mehmet Kemal Özdemir, Bjoern Menze
ECCV (79)2
2024 CT2Rep: Automated Radiology Report Generation for 3D Medical Imaging
Ibrahim Ethem Hamamci, Sezgin Er, Bjoern Menze
MICCAI (12)2
2023 Diffusion-Based Hierarchical Multi-label Object Detection to Analyze Panoramic Dental X-Rays
Ibrahim Ethem Hamamci, Sezgin Er, Enis Simsar, Anjany Sekuboyina, Mustafa Gundogar, Bernd Stadlinger, Albert Mehl, Bjoern Menze
MICCAI (6)2