Sarah Parisot

dblp:20/10169 · DBLP profile ↗
← Back
31ranked-venue papers
11as first author
9since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 9 first-author · 6 since 2021Artificial intelligence and machine learning · 20 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 7 first-author
YearPublicationVenuePosition
2024 Improving Object Detection via Local-global Contrastive Learning
Danai Triantafyllidou, Sarah Parisot, Ales Leonardis, Steven McDonagh 0001
BMVC2
2024 MULAN: A Multi Layer Annotated Dataset for Controllable Text-to-Image Generation
abstract
Text-to-image generation has achieved astonishing results, yet precise spatial controllability and prompt fidelity remain highly challenging. This limitation is typically addressed through cumbersome prompt engineering, scene layout conditioning, or image editing techniques which often require hand drawn masks. Nonetheless, pre-existing works struggle to take advantage of the natural instance-level compositionality of scenes due to the typically flat nature of rasterized RGB output images. Towards adressing this challenge, we introduce MuLAn: a novel dataset comprising over 44K MUlti-Layer ANnotations of RGB images as multi-layer, instance-wise RGBA decompositions, and over 100K instance images. To build MuLAn, we developed a training free pipeline which decomposes a monocular RGB image into a stack of RGBA layers comprising of background and isolated instances. We achieve this through the use of pre-trained general-purpose models, and by developing three modules: image decomposition for instance discovery and extraction, instance completion to reconstruct occluded areas, and image re-assembly. We use our pipeline to create MuLAn-COCO and MuLAn-LAION datasets, which contain a variety of image decompositions in terms of style, composition and complexity. With MuLAn, we provide the first photorealistic resource providing instance decompo-sition and occlusion information for high quality images, opening up new avenues for text-to-image generative AI re-search. With this, we aim to encourage the development of novel generation and editing technology, in particular layer-wise solutions. MuLAn data resources are available at https://MuLAn-dataset.github.io/.
Petru-Daniel Tudosiu, Yongxin Yang, Steven McDonagh 0001, Gerasimos Lampouras, Ignacio Iacobacci, Sarah Parisot
CVPR8
2024 Generating compositional scenes via Text-to-image RGBA Instance Generation
abstract
Text-to-image diffusion generative models can generate high quality images at the cost of tedious prompt engineering. Controllability can be improved by introducing layout conditioning, however existing methods lack layout editing ability and fine-grained control over object attributes. The concept of multi-layer generation holds great potential to address these limitations, however generating image instances concurrently to scene composition limits control over fine-grained object attributes, relative positioning in 3D space and scene manipulation abilities. In this work, we propose a novel multi-stage generation paradigm that is designed for fine-grained control, flexibility and interactivity. To ensure control over instance attributes, we devise a novel training paradigm to adapt a diffusion model to generate isolated scene components as RGBA images with transparency information. To build complex images, we employ these pre-generated instances and introduce a multi-layer composite generation process that smoothly assembles components in realistic scenes. Our experiments show that our RGBA diffusion model is capable of generating diverse and high quality instances with precise control over object attributes. Through multi-layer composition, we demonstrate that our approach allows to build and manipulate images from highly complex prompts with fine-grained control over object appearance and location, granting a higher degree of control than competing methods.
Alessandro Fontanella, Petru-Daniel Tudosiu, Yongxin Yang, Sarah Parisot
NeurIPS5
2023 Learning to Name Classes for Vision and Language Models
abstract
Large scale vision and language models can achieve impressive zero-shot recognition performance by mapping class specific text queries to image content. Two distinct challenges that remain however, are high sensitivity to the choice of handcrafted class names that define queries, and the difficulty of adaptation to new, smaller datasets. Towards addressing these problems, we propose to leverage available data to learn, for each class, an optimal word embedding as a function of the visual content. By learning new word embeddings on an otherwise frozen model, we are able to retain zero-shot capabilities for new classes, easily adapt models to new datasets, and adjust potentially erroneous, non-descriptive or ambiguous class names. We show that our solution can easily be integrated in image classification and object detection pipelines, yields significant performance gains in multiple scenarios and provides insights into model biases and labelling errors.
Sarah Parisot, Yongxin Yang, Steven McDonagh 0001
CVPR1
2023 CLAD: A realistic Continual Learning benchmark for Autonomous Driving
Eli Verwimp, Sarah Parisot, Lanqing Hong, Steven McDonagh 0001, Eduardo Pérez-Pellitero, Matthias De Lange, Tinne Tuytelaars
Neural Networks3
2022 Content-Diverse Comparisons improve IQA
William Thong, José Costa Pereira, Sarah Parisot, Ales Leonardis, Steven McDonagh 0001
BMVC3
2022 Re-examining Distillation for Continual Object Detection
Eli Verwimp, Sarah Parisot, Lanqing Hong, Steven McDonagh 0001, Eduardo Pérez-Pellitero, Matthias De Lange, Tinne Tuytelaars
BMVC3
2022 Long-tail Recognition via Compositional Knowledge Transfer
abstract
In this work, we introduce a novel strategy for long-tail recognition that addresses the tail classes’ few-shot problem via training-free knowledge transfer. Our objective is to transfer knowledge acquired from information-rich common classes to semantically similar, and yet data-hungry, rare classes in order to obtain stronger tail class representations. We leverage the fact that class prototypes and learned cosine classifiers provide two different, complementary representations of class cluster centres in feature space, and use an attention mechanism to select and recompose learned classifier features from common classes to obtain higher quality rare class representations. Our knowledge transfer process is training free, reducing overfitting risks, and can afford continual extension of classifiers to new classes. Experiments show that our approach can achieve significant performance boosts on rare classes while maintaining robust common class performance, outperforming directly comparable state-of-the-art models
Sarah Parisot, Pedro M. Esperança, Steven McDonagh 0001, Tamas J. Madarasz, Yongxin Yang, Zhenguo Li
CVPR1
2022 A Continual Learning Survey: Defying Forgetting in Classification Tasks
abstract
Artificial neural networks thrive in solving the classification problem for a particular rigid task, acquiring knowledge through generalized learning behaviour from a distinct training phase. The resulting network resembles a static entity of knowledge, with endeavours to extend this knowledge without targeting the original task resulting in a catastrophic forgetting. Continual learning shifts this paradigm towards networks that can continually accumulate knowledge over different tasks without the need to retrain from scratch. We focus on task incremental classification, where tasks arrive sequentially and are delineated by clear boundaries. Our main contributions concern: (1) a taxonomy and extensive overview of the state-of-the-art; (2) a novel framework to continually determine the stability-plasticity trade-off of the continual learner; (3) a comprehensive experimental comparison of 11 state-of-the-art continual learning methods; and (4) baselines. We empirically scrutinize method strengths and weaknesses on three benchmarks, considering Tiny Imagenet and large-scale unbalanced iNaturalist and a sequence of recognition datasets. We study the influence of model capacity, weight decay and dropout regularization, and the order in which the tasks are presented, and qualitatively compare methods in terms of required memory, computation time, and storage.
Matthias De Lange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia 0012, Ales Leonardis, Gregory Slabaugh, Tinne Tuytelaars
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 EHSOD: CAM-Guided End-to-End Hybrid-Supervised Object Detection with Cascade Refinement
abstract
Object detectors trained on fully-annotated data currently yield state of the art performance but require expensive manual annotations. On the other hand, weakly-supervised detectors have much lower performance and cannot be used reliably in a realistic setting. In this paper, we study the hybrid-supervised object detection problem, aiming to train a high quality detector with only a limited amount of fully-annotated data and fully exploiting cheap data with image-level labels. State of the art methods typically propose an iterative approach, alternating between generating pseudo-labels and updating a detector. This paradigm requires careful manual hyper-parameter tuning for mining good pseudo labels at each round and is quite time-consuming. To address these issues, we present EHSOD, an end-to-end hybrid-supervised object detection system which can be trained in one shot on both fully and weakly-annotated data. Specifically, based on a two-stage detector, we proposed two modules to fully utilize the information from both kinds of labels: 1) CAM-RPN module aims at finding foreground proposals guided by a class activation heat-map; 2) hybrid-supervised cascade module further refines the bounding-box position and classification with the help of an auxiliary head compatible with image-level data. Extensive experiments demonstrate the effectiveness of the proposed method and it achieves comparable results on multiple object detection benchmarks with only 30% fully-annotated data, e.g. 37.5% mAP on COCO. We will release the code and the trained models.
Linpu Fang, Hang Xu 0004, Zhili Liu, Sarah Parisot, Zhenguo Li
AAAI4
2020 A Multi-Hypothesis Approach to Color Constancy
abstract
Contemporary approaches frame the color constancy problem as learning camera specific illuminant mappings. While high accuracy can be achieved on camera specific data, these models depend on camera spectral sensitivity and typically exhibit poor generalisation to new devices. Additionally, regression methods produce point estimates that do not explicitly account for potential ambiguities among plausible illuminant solutions, due to the ill-posed nature of the problem. We propose a Bayesian framework that naturally handles color constancy ambiguity via a multi-hypothesis strategy. Firstly, we select a set of candidate scene illuminants in a data-driven fashion and apply them to a target image to generate a set of corrected images. Secondly, we estimate, for each corrected image, the likelihood of the light source being achromatic using a camera-agnostic CNN. Finally, our method explicitly learns a final illumination estimate from the generated posterior probability distribution. Our likelihood estimator learns to answer a camera-agnostic question and thus enables effective multi-camera training by disentangling illuminant estimation from the supervised learning task. We extensively evaluate our proposed approach and additionally set a benchmark for novel sensor generalisation without re-training. Our method provides state-of-the-art accuracy on multiple public datasets (up to 11% median angular error improvement) while maintaining real-time execution.
Daniel Hernández Juárez, Sarah Parisot, Benjamin Busam, Ales Leonardis, Gregory Slabaugh, Steven McDonagh 0001
CVPR2
2020 Unsupervised Model Personalization While Preserving Privacy and Scalability: An Open Problem
abstract
This work investigates the task of unsupervised model personalization, adapted to continually evolving, unlabeled local user images. We consider the practical scenario where a high capacity server interacts with a myriad of resource-limited edge devices, imposing strong requirements on scalability and local data privacy. We aim to address this challenge within the continual learning paradigm and provide a novel Dual User-Adaptation framework (DUA) to explore the problem. This framework flexibly disentangles user-adaptation into model personalization on the server and local data regularization on the user device, with desirable properties regarding scalability and privacy constraints. First, on the server, we introduce incremental learning of task-specific expert models, subsequently aggregated using a concealed unsupervised user prior. Aggregation avoids retraining, whereas the user prior conceals sensitive raw user data, and grants unsupervised adaptation. Second, local user-adaptation incorporates a domain adaptation point of view, adapting regularizing batch normalization parameters to the user data. We explore various empirical user configurations with different priors in categories and a tenfold of transforms for MIT Indoor Scene recognition, and classify numbers in a combined MNIST and SVHN setup. Extensive experiments yield promising results for data-driven local adaptation and elicit user priors for server adaptation to depend on the model rather than user data. Hence, although user-adaptation remains a challenging open problem, the DUA framework formalizes a principled foundation for personalizing both on server and user device, while maintaining privacy and scalability.
Matthias De Lange, Xu Jia 0012, Sarah Parisot, Ales Leonardis, Gregory Slabaugh, Tinne Tuytelaars
CVPR3
2020 DeepLPF: Deep Local Parametric Filters for Image Enhancement
abstract
Digital artists often improve the aesthetic quality of digital photographs through manual retouching. Beyond global adjustments, professional image editing programs provide local adjustment tools operating on specific parts of an image. Options include parametric (graduated, radial filters) and unconstrained brush tools. These highly expressive tools enable a diverse set of local image enhancements. However, their use can be time consuming, and requires artistic capability. State-of-the-art automated image enhancement approaches typically focus on learning pixel-level or global enhancements. The former can be noisy and lack interpretability, while the latter can fail to capture fine-grained adjustments. In this paper, we introduce a novel approach to automatically enhance images using learned spatially local filters of three different types (Elliptical Filter, Graduated Filter, Polynomial Filter). We introduce a deep neural network, dubbed Deep Local Parametric Filters (DeepLPF), which regresses the parameters of these spatially localized filters that are then automatically applied to enhance the image. DeepLPF provides a natural form of model regularization and enables interpretable, intuitive adjustments that lead to visually pleasing results. We report on multiple benchmarks and show that DeepLPF produces state-of-the-art performance on two variants of the MIT-Adobe 5k dataset, often using a fraction of the parameters required for competing methods.
Sean Moran, Pierre Marza, Steven McDonagh 0001, Sarah Parisot, Gregory Slabaugh
CVPR4
2020 More Classifiers, Less Forgetting: A Generic Multi-classifier Paradigm for Incremental Learning
Yu Liu 0012, Sarah Parisot, Gregory Slabaugh, Xu Jia 0012, Ales Leonardis, Tinne Tuytelaars
ECCV (26)2
2020 Many-Shot from Low-Shot: Learning to Annotate Using Mixed Supervision for Object Detection
Carlo Biffi, Steven McDonagh 0001, Philip Torr 0001, Ales Leonardis, Sarah Parisot
ECCV (8)5
2020 Few-Shot Single-View 3-D Object Reconstruction with Compositional Priors
Mateusz Michalkiewicz, Sarah Parisot, Stavros Tsogkas, Mahsa Baktash, Anders P. Eriksson, Eugene Belilovsky
ECCV (25)2
2020 Low Light Video Enhancement Using Synthetic Data Produced with an Intermediate Domain Mapping
Danai Triantafyllidou, Sean Moran, Steven McDonagh 0001, Sarah Parisot, Gregory Slabaugh
ECCV (13)4
2020 Probabilistic 3D Surface Reconstruction from Sparse MRI Information
Katarína Tóthová, Sarah Parisot, Matthew C. H. Lee, Esther Puyol-Antón, Andrew P. King, Marc Pollefeys, Ender Konukoglu
MICCAI (1)2
2018 Learning Conditioned Graph Structures for Interpretable Visual Question Answering
abstract
Visual Question answering is a challenging problem requiring a combination of concepts from Computer Vision and Natural Language Processing. Most existing approaches use a two streams strategy, computing image and question features that are consequently merged using a variety of techniques. Nonetheless, very few rely on higher level image representations, which can capture semantic and spatial relationships. In this paper, we propose a novel graph-based approach for Visual Question Answering. Our method combines a graph learner module, which learns a question specific graph representation of the input image, with the recent concept of graph convolutions, aiming to learn image representations that capture question specific interactions. We test our approach on the VQA v2 dataset using a simple baseline architecture enhanced by the proposed graph learner module. We obtain promising results with 66.18% accuracy and demonstrate the interpretability of the proposed method. Code can be found at github.com/aimbrain/vqa-project.
Will Norcliffe-Brown, Efstathios Vafeias, Sarah Parisot
NeurIPS3
2018 Disease prediction using graph convolutional networks: Application to Autism Spectrum Disorder and Alzheimer's disease
Sarah Parisot, Sofia Ira Ktena, Enzo Ferrante, Matthew C. H. Lee, Ricardo Guerrero, Ben Glocker, Daniel Rueckert
Medical Image Anal.1
2017 Distance Metric Learning Using Graph Convolutional Networks: Application to Functional Brain Networks
Sofia Ira Ktena, Sarah Parisot, Enzo Ferrante, Martin Rajchl, Matthew C. H. Lee, Ben Glocker, Daniel Rueckert
MICCAI (1)2
2017 Spectral Graph Convolutions for Population-Based Disease Prediction
Sarah Parisot, Sofia Ira Ktena, Enzo Ferrante, Matthew C. H. Lee, Ricardo Guerrero, Ben Glocker, Daniel Rueckert
MICCAI (3)1
2016 Boundary Mapping Through Manifold Learning for Connectivity-Based Cortical Parcellation
Salim Arslan, Sarah Parisot, Daniel Rueckert
MICCAI (1)2
2016 GraMPa: Graph-Based Multi-modal Parcellation of the Cortex Using Fusion Moves
Sarah Parisot, Ben Glocker, Markus Schirmer, Daniel Rueckert
MICCAI (1)1
2016 (Hyper)-graphical models in biomedical image analysis
Nikos Paragios, Enzo Ferrante, Ben Glocker, Nikos Komodakis, Sarah Parisot, Evangelia I. Zacharaki
Medical Image Anal.5
2015 A Continuous Flow-Maximisation Approach to Connectivity-Driven Cortical Parcellation
Sarah Parisot, Martin Rajchl, Jonathan Passerat-Palmbach, Daniel Rueckert
MICCAI (3)1
2014 Concurrent tumor segmentation and registration with uncertainty-based sparse non-uniform graphs
Sarah Parisot, William M. Wells III, Stéphane Chemouny, Hugues Duffau, Nikos Paragios
Medical Image Anal.1
2013 Uncertainty-Driven Efficiently-Sampled Sparse Graphical Models for Concurrent Tumor Segmentation and Atlas Registration
abstract
Graph-based methods have become popular in recent years and have successfully addressed tasks like segmentation and deformable registration. Their main strength is optimality of the obtained solution while their main limitation is the lack of precision due to the grid-like representations and the discrete nature of the quantized search space. In this paper we introduce a novel approach for combined segmentation/registration of brain tumors that adapts graph and sampling resolution according to the image content. To this end we estimate the segmentation and registration marginals towards adaptive graph resolution and intelligent definition of the search space. This information is considered in a hierarchical framework where uncertainties are propagated in a natural manner. State of the art results in the joint segmentation/registration of brain images with low-grade gliomas demonstrate the potential of our approach.
Sarah Parisot, William M. Wells III, Stéphane Chemouny, Hugues Duffau, Nikos Paragios
ICCV1
2012 Graph-based detection, segmentation & characterization of brain tumors
abstract
In this paper we propose a novel approach for detection, segmentation and characterization of brain tumors. Our method exploits prior knowledge in the form of a sparse graph representing the expected spatial positions of tumor classes. Such information is coupled with image-based classification techniques along with spatial smoothness constraints towards producing a reliable detection map within a unified graphical model formulation. Towards optimal use of prior knowledge, a two layer interconnected graph is considered with one layer corresponding to the low-grade glioma type (characterization) and the second layer to voxel-based decisions of tumor presence. Efficient linear programming both in terms of performance as well as in terms of computational load is considered to recover the lowest potential of the objective function. The outcome of the method refers to both tumor segmentation as well as their characterization. Promising results on substantial data sets demonstrate the extreme potentials of our method.
Sarah Parisot, Hugues Duffau, Stéphane Chemouny, Nikos Paragios
CVPR1
2012 Joint Tumor Segmentation and Dense Deformable Registration of Brain MR Images
Sarah Parisot, Hugues Duffau, Stéphane Chemouny, Nikos Paragios
MICCAI (2)1
2011 Graph Based Spatial Position Mapping of Low-Grade Gliomas
Sarah Parisot, Hugues Duffau, Stéphane Chemouny, Nikos Paragios
MICCAI (2)1