Cosmin Bercea

dblp:186/7940 · also Cosmin I. Bercea · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0003-2628-2766ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Knowledge to Sight: Reasoning over Visual Attributes via Knowledge Decomposition for Abnormality Grounding
abstract
In this work, we address the problem of grounding abnormalities in medical images, where the goal is to localize clinical findings based on textual descriptions. While generalist Vision-Language Models (VLMs) excel in natural grounding tasks, they often struggle in the medical domain due to rare, compositional, and domain-specific terms that are poorly aligned with visual patterns. Specialized medical VLMs address this challenge via large-scale domain pretraining, but at the cost of substantial annotation and computational resources. To overcome these limitations, we propose Knowledge to Sight (K2Sight), a framework that introduces structured semantic supervision by decomposing clinical concepts into interpretable visual attributes, such as shape, density, and anatomical location. These attributes are distilled from domain ontologies and encoded into concise instruction-style prompts, which guide region-text alignment during training. Unlike conventional report-level supervision, our approach explicitly bridges domain knowledge and spatial structure, enabling data-efficient training of compact models. We train compact models with 0.23B and 2B parameters using only 1.5% of the data required by state-of-the-art medical VLMs. Despite their small size and limited training data, these models achieve performance on par with or better than 7B+ medical VLMs, with up to 9.82% improvement in mAP50. Code and models: https://lijunrio.github.io/K2Sight/.
Che Liu 0002, Wenjia Bai, Rossella Arcucci, Cosmin Bercea, Julia A. Schnabel
WACV6
2026 TomoGraphView: 3D medical image classification with omnidirectional slice representations and graph neural networks
abstract
The sharp rise in medical tomography examinations has created a demand for automated systems that can reliably extract informative features for downstream tasks such as tumor characterization. Although 3D volumes contain richer information than individual slices, effective 3D classification remains difficult: volumetric data encode complex spatial dependencies, and the scarcity of large-scale 3D datasets has constrained progress toward 3D foundation models. As a result, many recent approaches rely on 2D vision foundation models trained on natural images, repurposing them as feature extractors for medical scans with surprisingly strong performance. Despite their practical success, current methods that apply 2D foundation models to 3D scans via slice-based decomposition remain fundamentally limited. Standard slicing along axial, sagittal, and coronal planes often fails to capture the true spatial extent of a structure when its orientation does not align with these canonical views. More critically, most approaches aggregate slice features independently, ignoring the underlying 3D geometry and losing spatial coherence across slices. To overcome these limitations, we propose TomoGraphView, a novel framework that integrates omnidirectional volume slicing with spherical graph-based feature aggregation. Instead of restricting the model to axial, sagittal, or coronal planes, our method samples both canonical and non-canonical cross-sections generated from uniformly distributed points on a sphere enclosing the volume. Triangulating these viewpoints yields a spherical graph that captures spatial relationships among views, and we use a graph neural network to aggregate their features accordingly. Experiments across six oncology 3D medical image classification datasets demonstrate that omnidirectional volume slicing improves the average performance in Area Under the Receiver Operating Characteristic Curve (AUROC) from 0.7701 to 0.8154 compared with traditional slicing approaches relying on canonical view planes. Moreover, we can further improve AUROC performance from 0.8198 to 0.8372 by leveraging our proposed graph neural network-based feature aggregation. Notably, TomoGraphView also surpasses large-scale pretrained 3D medical imaging models across all datasets and tasks, underscoring its effectiveness as a powerful framework for volumetric analysis and therefore represents a key step toward bridging the gap until fully native 3D foundation models become available in medical image analysis. We provide a user-friendly library for omnidirectional volume slicing at https://pypi.org/project/OmniSlicer.
Johannes Kiechle, Stefan M. Fischer, Daniel Lang 0003, Cosmin Bercea, Matthew Nyflot, Lina Felsner, Julia A. Schnabel, Jan Peeken
Medical Image Anal.4
2025 Influence of Classification Task and Distribution Shift Type on OOD Detection in Fetal Ultrasound
Chun Kit Wong, Anders Nymark Christensen, Cosmin Bercea, Julia A. Schnabel, Martin Grønnebæk Tolsgaard, Aasa Feragen
MICCAI (7)3
2025 NOVA: A Benchmark for Rare Anomaly Localization and Clinical Reasoning in Brain MRI
abstract
In many real-world applications, deployed models encounter inputs that differ from the data seen during training. Open-world recognition ensures that such systems remain robust as ever-emerging, previously _unknown_ categories appear and must be addressed without retraining.Foundation and vision-language models are pre-trained on large and diverse datasets with the expectation of broad generalization across domains, including medical imaging.However, benchmarking these models on test sets with only a few common outlier types silently collapses the evaluation back to a closed-set problem, masking failures on rare or truly novel conditions encountered in clinical use.We therefore present NOVA, a challenging, real-life _evaluation-only_ benchmark of $\sim$900 brain MRI scans that span 281 rare pathologies and heterogeneous acquisition protocols. Each case includes rich clinical narratives and double-blinded expert bounding-box annotations. Together, these enable joint assessment of anomaly localisation, visual captioning, and diagnostic reasoning. Because NOVA is never used for training, it serves as an _extreme_ stress-test of out-of-distribution generalisation: models must bridge a distribution gap both in sample appearance and in semantic space. Baseline results with leading vision-language models (GPT-4o, Gemini 2.0 Flash, and Qwen2.5-VL-72B) reveal substantial performance drops, with approximately a 65\% gap in localisation compared to natural-image benchmarks and 40\% and 20\% gaps in captioning and reasoning, respectively, compared to resident radiologists. Therefore, NOVA establishes a testbed for advancing models that can detect, localize, and reason about truly unknown anomalies.
Cosmin Bercea, Philipp Raffler, Evamaria O. Riedel, Lena Schmitzer, Angela Kurz, Felix Bitzer, Paula Roßmüller, Julian Canisius, Mirjam L. Beyrle, Che Liu 0002, Wenjia Bai, Bernhard Kainz, Julia A. Schnabel, Benedikt Wiestler
NeurIPS1
2024 Diffusion Models with Implicit Guidance for Medical Anomaly Detection
Cosmin Bercea, Benedikt Wiestler, Daniel Rueckert, Julia A. Schnabel
MICCAI (11)1
2024 Interpretable Representation Learning of Cardiac MRI via Attribute Regularization
Maxime Di Folco, Cosmin Bercea, Emily Chan, Julia A. Schnabel
MICCAI (10)2
2023 What Do AEs Learn? Challenging Common Assumptions in Unsupervised Anomaly Detection
Cosmin Bercea, Daniel Rueckert, Julia A. Schnabel
MICCAI (5)1
2023 Reversing the Abnormal: Pseudo-Healthy Generative Networks for Anomaly Detection
Cosmin Bercea, Benedikt Wiestler, Daniel Rueckert, Julia A. Schnabel
MICCAI (5)1
2016 Confidence-aware Levenberg-Marquardt optimization for joint motion estimation and super-resolution
abstract
Motion estimation across low-resolution frames and the reconstruction of high-resolution images are two coupled subproblems of multi-frame super-resolution. This paper introduces a new joint optimization approach for motion estimation and image reconstruction to address this interdependence. Our method is formulated via non-linear least squares optimization and combines two principles of robust super-resolution. First, to enhance the robustness of the joint estimation, we propose a confidence-aware energy minimization framework augmented with sparse regularization. Second, we develop a tailor-made Levenberg-Marquardt iteration scheme to jointly estimate motion parameters and the high-resolution image along with the corresponding model confidence parameters. Our experiments on simulated and real images confirm that the proposed approach outperforms decoupled motion estimation and image reconstruction as well as related state-of-the-art joint estimation algorithms.
Cosmin Bercea, Andreas K. Maier, Thomas Köhler 0004
ICIP1