Petru-Daniel Tudosiu

dblp:258/4838 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
10since 2021 · last 2024
0000-0001-6435-5079ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2024 MULAN: A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation
abstract
Text-to-image generation has achieved astonishing results, yet precise spatial controllability and prompt fidelity remain highly challenging. This limitation is typically addressed through cumbersome prompt engineering, scene layout conditioning, or image editing techniques which often require hand drawn masks. Nonetheless, pre-existing works struggle to take advantage of the natural instance-level compositionality of scenes due to the typically flat nature of rasterized RGB output images. Towards adressing this challenge, we introduce MuLAn: a novel dataset comprising over 44K MUlti-Layer ANnotations of RGB images as multi-layer, instance-wise RGBA decompositions, and over 100K instance images. To build MuLAn, we developed a training free pipeline which decomposes a monocular RGB image into a stack of RGBA layers comprising of background and isolated instances. We achieve this through the use of pre-trained general-purpose models, and by developing three modules: image decomposition for instance discovery and extraction, instance completion to reconstruct occluded areas, and image re-assembly. We use our pipeline to create MuLAn-COCO and MuLAn-LAION datasets, which contain a variety of image decompositions in terms of style, composition and complexity. With MuLAn, we provide the first photorealistic resource providing instance decompo-sition and occlusion information for high quality images, opening up new avenues for text-to-image generative AI re-search. With this, we aim to encourage the development of novel generation and editing technology, in particular layer-wise solutions. MuLAn data resources are available at https://MuLAn-dataset.github.io/.
Petru-Daniel Tudosiu, Yongxin Yang, Steven McDonagh 0001, Gerasimos Lampouras, Ignacio Iacobacci, Sarah Parisot
CVPR1
2024 Generating compositional scenes via Text-to-image RGBA Instance Generation
abstract
Text-to-image diffusion generative models can generate high quality images at the cost of tedious prompt engineering. Controllability can be improved by introducing layout conditioning, however existing methods lack layout editing ability and fine-grained control over object attributes. The concept of multi-layer generation holds great potential to address these limitations, however generating image instances concurrently to scene composition limits control over fine-grained object attributes, relative positioning in 3D space and scene manipulation abilities. In this work, we propose a novel multi-stage generation paradigm that is designed for fine-grained control, flexibility and interactivity. To ensure control over instance attributes, we devise a novel training paradigm to adapt a diffusion model to generate isolated scene components as RGBA images with transparency information. To build complex images, we employ these pre-generated instances and introduce a multi-layer composite generation process that smoothly assembles components in realistic scenes. Our experiments show that our RGBA diffusion model is capable of generating diverse and high quality instances with precise control over object attributes. Through multi-layer composition, we demonstrate that our approach allows to build and manipulate images from highly complex prompts with fine-grained control over object appearance and location, granting a higher degree of control than competing methods.
Alessandro Fontanella, Petru-Daniel Tudosiu, Yongxin Yang, Sarah Parisot
NeurIPS2
2024 Generating multi-pathological and multi-modal images and labels for brain MRI
abstract
The last few years have seen a boom in using generative models to augment real datasets, as synthetic data can effectively model real data distributions and provide privacy-preserving, shareable datasets that can be used to train deep learning models. However, most of these methods are 2D and provide synthetic datasets that come, at most, with categorical annotations. The generation of paired images and segmentation samples that can be used in downstream, supervised segmentation tasks remains fairly uncharted territory. This work proposes a two-stage generative model capable of producing 2D and 3D semantic label maps and corresponding multi-modal images. We use a latent diffusion model for label synthesis and a VAE-GAN for semantic image synthesis. Synthetic datasets provided by this model are shown to work in a wide variety of segmentation tasks, supporting small, real datasets or fully replacing them while maintaining good performance. We also demonstrate its ability to improve downstream performance on out-of-distribution data.
Virginia Fernandez, Walter H. L. Pinaya, Pedro Borges, Mark S. Graham, Petru-Daniel Tudosiu, Tom Vercauteren, Manuel Jorge Cardoso
Medical Image Anal.5
2023 Unsupervised 3D Out-of-Distribution Detection with Latent Diffusion Models
Mark S. Graham, Walter H. L. Pinaya, Paul Wright 0001, Petru-Daniel Tudosiu, Yee-Haur Mah, James T. Teo, Hans Rolf Jäger, David Werring, Parashkev Nachev, Sébastien Ourselin, Manuel Jorge Cardoso
MICCAI (1)4
2023 Geometry-Invariant Abnormality Detection
Ashay Patel, Petru-Daniel Tudosiu, Walter H. L. Pinaya, Olusola Adeleke, Gary J. Cook, Vicky Goh, Sébastien Ourselin, Manuel Jorge Cardoso
MICCAI (1)2
2023 InverseSR: 3D Brain MRI Super-Resolution Using a Latent Diffusion Model
Jueqi Wang, Jacob Levman, Walter H. L. Pinaya, Petru-Daniel Tudosiu, Manuel Jorge Cardoso, Razvan V. Marinescu
MICCAI (10)4
2023 Latent Transformer Models for out-of-distribution detection
abstract
Any clinically-deployed image-processing pipeline must be robust to the full range of inputs it may be presented with. One popular approach to this challenge is to develop predictive models that can provide a measure of their uncertainty. Another approach is to use generative modelling to quantify the likelihood of inputs. Inputs with a low enough likelihood are deemed to be out-of-distribution and are not presented to the downstream predictive model. In this work, we evaluate several approaches to segmentation with uncertainty for the task of segmenting bleeds in 3D CT of the head. We show that these models can fail catastrophically when operating in the far out-of-distribution domain, often providing predictions that are both highly confident and wrong. We propose to instead perform out-of-distribution detection using the Latent Transformer Model: a VQ-GAN is used to provide a highly compressed latent representation of the input volume, and a transformer is then used to estimate the likelihood of this compressed representation of the input. We demonstrate this approach can identify images that are both far- and near- out-of-distribution, as well as provide spatial maps that highlight the regions considered to be out-of-distribution. Furthermore, we find a strong relationship between an image's likelihood and the quality of a model's segmentation on it, demonstrating that this approach is viable for filtering out unsuitable images.
Mark S. Graham, Petru-Daniel Tudosiu, Paul Wright 0001, Walter H. L. Pinaya, Petteri Teikari, Ashay Patel, Jean-Marie U.-King-Im, Yee-Haur Mah, James T. Teo, Hans Rolf Jäger, David Werring, Geraint Rees 0001, Parashkev Nachev, Sébastien Ourselin, Manuel Jorge Cardoso
Medical Image Anal.2
2023 ICAM-Reg: Interpretable Classification and Regression With Feature Attribution for Mapping Neurological Phenotypes in Individual Scans
abstract
An important goal of medical imaging is to be able to precisely detect patterns of disease specific to individual scans; however, this is challenged in brain imaging by the degree of heterogeneity of shape and appearance. Traditional methods, based on image registration, historically fail to detect variable features of disease, as they utilise population-based analyses, suited primarily to studying group-average effects. In this paper we therefore take advantage of recent developments in generative deep learning to develop a method for simultaneous classification, or regression, and feature attribution (FA). Specifically, we explore the use of a VAE-GAN (variational autoencoder - general adversarial network) for translation called ICAM, to explicitly disentangle class relevant features, from background confounds, for improved interpretability and regression of neurological phenotypes. We validate our method on the tasks of Mini-Mental State Examination (MMSE) cognitive test score prediction for the Alzheimer's Disease Neuroimaging Initiative (ADNI) cohort, as well as brain age prediction, for both neurodevelopment and neurodegeneration, using the developing Human Connectome Project (dHCP) and UK Biobank datasets. We show that the generated FA maps can be used to explain outlier predictions and demonstrate that the inclusion of a regression module improves the disentanglement of the latent space. Our code is freely available on GitHub https://github.com/CherBass/ICAM.
Cher Bass, Mariana da Silva, Carole H. Sudre, Logan Z. J. Williams, Helena S. Sousa, Petru-Daniel Tudosiu, Fidel Alfaro-Almagro, Sean P. Fitzgibbon, Matthew F. Glasser, Stephen M. Smith 0001, Emma C. Robinson
IEEE Trans. Medical Imaging6
2022 Fast Unsupervised Brain Anomaly Detection and Segmentation with Diffusion Models
Walter H. L. Pinaya, Mark S. Graham, Robert J. Gray, Pedro F. Da Costa, Petru-Daniel Tudosiu, Paul Wright 0001, Yee-Haur Mah, Andrew D. MacKinnon, James T. Teo, Hans Rolf Jäger, David Werring, Geraint Rees 0001, Parashkev Nachev, Sébastien Ourselin, Manuel Jorge Cardoso
MICCAI (8)5
2022 Unsupervised brain imaging 3D anomaly detection and segmentation with transformers
abstract
Pathological brain appearances may be so heterogeneous as to be intelligible only as anomalies, defined by their deviation from normality rather than any specific set of pathological features. Amongst the hardest tasks in medical imaging, detecting such anomalies requires models of the normal brain that combine compactness with the expressivity of the complex, long-range interactions that characterise its structural organisation. These are requirements transformers have arguably greater potential to satisfy than other current candidate architectures, but their application has been inhibited by their demands on data and computational resources. Here we combine the latent representation of vector quantised variational autoencoders with an ensemble of autoregressive transformers to enable unsupervised anomaly detection and segmentation defined by deviation from healthy brain imaging data, achievable at low computational cost, within relative modest data regimes. We compare our method to current state-of-the-art approaches across a series of experiments with 2D and 3D data involving synthetic and real pathological lesions. On real lesions, we train our models on 15,000 radiologically normal participants from UK Biobank and evaluate performance on four different brain MR datasets with small vessel disease, demyelinating lesions, and tumours. We demonstrate superior anomaly detection performance both image-wise and pixel/voxel-wise, achievable without post-processing. These results draw attention to the potential of transformers in this most challenging of imaging tasks.
Walter H. L. Pinaya, Petru-Daniel Tudosiu, Robert J. Gray, Geraint Rees 0001, Parashkev Nachev, Sébastien Ourselin, Manuel Jorge Cardoso
Medical Image Anal.2
2020 ICAM: Interpretable Classification via Disentangled Representations and Feature Attribution Mapping
abstract
Feature attribution (FA), or the assignment of class-relevance to different locations in an image, is important for many classification problems but is particularly crucial within the neuroscience domain, where accurate mechanistic models of behaviours, or disease, require knowledge of all features discriminative of a trait. At the same time, predicting class relevance from brain images is challenging as phenotypes are typically heterogeneous, and changes occur against a background of significant natural variation. Here, we present a novel framework for creating class specific FA maps through image-to-image translation. We propose the use of a VAE-GAN to explicitly disentangle class relevance from background features for improved interpretability properties, which results in meaningful FA maps. We validate our method on 2D and 3D brain image datasets of dementia (ADNI dataset), ageing (UK Biobank), and (simulated) lesion detection. We show that FA maps generated by our method outperform baseline FA methods when validated against ground truth. More significantly, our approach is the first to use latent space sampling to support exploration of phenotype variation.
Cher Bass, Mariana da Silva, Carole H. Sudre, Petru-Daniel Tudosiu, Stephen M. Smith 0001, Emma C. Robinson
NeurIPS4