EDBT 2026 Demo / reviewers in the wild / expert
Alasdair Newson
dblp:123/3502
· DBLP profile ↗
19ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0003-4247-0249ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Restyling Unsupervised Concept Based Interpretable Networks with Generative ModelsabstractDeveloping inherently interpretable models for prediction has gained prominence in recent years. A subclass of these models, wherein the interpretable network relies on learning high-level concepts, are valued because of closeness of concept representations to human communication. However, the visualization and understanding of the learnt unsupervised dictionary of concepts encounters major limitations, especially for large-scale images. We propose here a novel method that relies on mapping the concept features to the latent space of a pretrained generative model. The use of a generative model enables high quality visualization, and lays out an intuitive and interactive procedure for better interpretation of the learnt concepts by imputing concept activations and visualizing generated modifications. Furthermore, leveraging pretrained generative models has the additional advantage of making the training of the system more efficient. We quantitatively ascertain the efficacy of our method in terms of accuracy of the interpretable prediction network, fidelity of reconstruction, as well as faithfulness and consistency of learnt concepts. The experiments are conducted on multiple image recognition benchmarks for large-scale images. Project page available at https://jayneelparekh.github.io/VisCoIN_project_page/ Jayneel Parekh, Quentin Bouniot, Pavlo Mozharovskyi, Alasdair Newson, Florence d'Alché-Buc |
ICLR | 4 |
| 2025 | Learning to Steer: Input-dependent Steering for Multimodal LLMsabstractSteering has emerged as a practical approach to enable post-hoc guidance of LLMs towards enforcing a specific behavior.
However, it remains largely underexplored for multimodal LLMs (MLLMs); furthermore, existing steering techniques, such as \textit{mean} steering, rely on a single steering vector, applied independently of the input query. This paradigm faces limitations when the desired behavior is dependent on the example at hand. For example, a safe answer may consist in abstaining from answering when asked for an illegal activity, or may point to external resources or consultation with an expert when asked about medical advice. In this paper, we investigate a fine-grained steering that uses an input-specific linear shift. This shift is computed using contrastive input-specific prompting. However, the input-specific prompts required for this approach are not known at test time. Therefore, we propose to train a small auxiliary module to predict the input-specific steering vector. Our approach, dubbed as L2S (Learn-to-Steer), demonstrates that it reduces hallucinations and enforces safety in MLLMs, outperforming other static baselines. We will open-source our code. Jayneel Parekh, Pegah Khayatan, Mustafa Shukor, Arnaud Dapogny, Alasdair Newson, Matthieu Cord |
NeurIPS | 5 |
| 2025 | Infusion: Internal Diffusion for Inpainting of Dynamic Textures and Complex MotionabstractAbstract Video inpainting is the task of filling a region in a video in a visually convincing manner It is very challenging due to the high dimensionality of the data and the temporal consistency required for obtaining convincing results. Recently, diffusion models have shown impressive results in modeling complex data distributions, including images and videos. Such models remain nonetheless very expensive to train and to perform inference with, which strongly reduce their applicability to videos, and yields unreasonable computational loads. We show that in the case of video inpainting, thanks to the highly auto‐similar nature of videos, the training data of a diffusion model can be restricted to the input video and still produce very satisfying results. With this internal learning approach, where the training data is limited to a single video, our lightweight models perform very well with only half a million parameters, in contrast to the very large networks with billions of parameters typically found in the literature. We also introduce a new method for efficient training and inference of diffusion models in the context of internal learning, by splitting the diffusion process into different learning intervals corresponding to different noise levels of the diffusion process. We show qualitative and quantitative results, demonstrating that our method reaches or exceeds state of the art performance in the case of dynamic textures and complex dynamic backgrounds. Nicolas Cherel, Andrés Almansa, Yann Gousseau, Alasdair Newson |
Comput. Graph. Forum | 4 |
| 2025 | Neural Film Grain RenderingabstractAbstract Film grain refers to the specific texture of film‐acquired images, due to the physical nature of photographic film. Being a visual signature of such images, there is a strong interest in the film‐industry for the rendering of these textures for digital images. Some previous works are able to closely mimic the physics of films and produce high quality results, but are computationally expensive. We propose a method based on a lightweight neural network and a texture aware loss function, achieving realistic results with very low complexity, even for large grains and high resolutions. We evaluate our algorithm both quantitatively and qualitatively with respect to previous work. Gwilherm Lesné, Yann Gousseau, Saïd Ladjal, Alasdair Newson |
Comput. Graph. Forum | 4 |
| 2024 | A Concept-Based Explainability Framework for Large Multimodal ModelsabstractLarge multimodal models (LMMs) combine unimodal encoders and large language models (LLMs) to perform multimodal tasks. Despite recent advancements towards the interpretability of these models, understanding internal representations of LMMs remains largely a mystery. In this paper, we present a novel framework for the interpretation of LMMs. We propose a dictionary learning based approach, applied to the representation of tokens. The elements of the learned dictionary correspond to our proposed concepts. We show that these concepts are well semantically grounded in both vision and text. Thus we refer to these as ``multi-modal concepts''.
We qualitatively and quantitatively evaluate the results of the learnt concepts. We show that the extracted multimodal concepts are useful to interpret representations of test samples. Finally, we evaluate the disentanglement between different concepts and the quality of grounding concepts visually and textually. Our implementation is publicly available: https://github.com/mshukor/xl-vlms. Jayneel Parekh, Pegah Khayatan, Mustafa Shukor, Alasdair Newson, Matthieu Cord |
NeurIPS | 4 |
| 2024 | Patch-based stochastic attention for image editing
Nicolas Cherel, Andrés Almansa, Yann Gousseau, Alasdair Newson |
Comput. Vis. Image Underst. | 4 |
| 2022 | A Style-Based GAN Encoder for High Fidelity Reconstruction of Images and Videos
Alasdair Newson, Yann Gousseau, Pierre Hellier |
ECCV (15) | 2 |
| 2022 | A Patch-Based Algorithm for Diverse and High Fidelity Single Image GenerationabstractImage generation is the task of producing new samples from one or several example images. Until recently, this has been done using large image databases, in particular using Generative Adversarial Networks (GANs). However, Shaham et al. [1] recently proposed the SinGAN method, which achieves this generation using a single image example. At the same time, researchers are realizing that classical patch- based methods can replace certain neural networks, with no costly training. In this paper, we present a purely patch-based method, named Patches for Single image generation (PSin), which requires no training and generates samples in seconds. Our algorithm is based on the minimization of a global, patch- based energy functional, which ensures the visual fidelity of the result to the original image. We also ensure diversity of the results by carefully choosing the initialization of the algorithm. We propose two initialization variants. We compare our results to both the original SinGAN and another recent patch-based image generation approach, both qualitatively and quantitatively using multiple metrics. Nicolas Cherel, Andrés Almansa, Yann Gousseau, Alasdair Newson |
ICIP | 4 |
| 2021 | Multi-View Radar Semantic SegmentationabstractUnderstanding the scene around the ego-vehicle is key to assisted and autonomous driving. Nowadays, this is mostly conducted using cameras and laser scanners, despite their reduced performance in adverse weather conditions. Automotive radars are low-cost active sensors that measure properties of surrounding objects, including their relative speed, and have the key advantage of not being impacted by rain, snow or fog. However, they are seldom used for scene understanding due to the size and complexity of radar raw data and the lack of annotated datasets. Fortunately, recent open-sourced datasets have opened up research on classification, object detection and semantic segmentation with raw radar signals using end-to-end trainable models. In this work, we propose several novel architectures, and their associated losses, which analyse multiple "views" of the range-angle-Doppler radar tensor to segment it semantically. Experiments conducted on the recent CARRADA dataset demonstrate that our best model outperforms alternative models, derived either from the semantic segmentation of natural images or from radar scene understanding, while requiring significantly fewer parameters. Both our code and trained models are available at https://github.com/valeoai/MVRSS. Arthur Ouaknine, Alasdair Newson, Patrick Pérez, Florence Tupin, Julien Rebut |
ICCV | 2 |
| 2021 | A Latent Transformer for Disentangled Face Editing in Images and VideosabstractHigh quality facial image editing is a challenging problem in the movie post-production industry, requiring a high degree of control and identity preservation. Previous works that attempt to tackle this problem may suffer from the entanglement of facial attributes and the loss of the person’s identity. Furthermore, many algorithms are limited to a certain task. To tackle these limitations, we propose to edit facial attributes via the latent space of a StyleGAN generator, by training a dedicated latent transformation network and incorporating explicit disentanglement and identity preservation terms in the loss function. We further introduce a pipeline to generalize our face editing to videos. Our model achieves a disentangled, controllable, and identity-preserving facial attribute editing, even in the challenging case of real (i.e., non-synthetic) images and videos. We conduct extensive experiments on image and video datasets and show that our model outperforms other state-of-the-art methods in visual quality and quantitative evaluation. Source codes are available at https://github.com/InterDigitalInc/latent-transformer. Alasdair Newson, Yann Gousseau, Pierre Hellier |
ICCV | 2 |
| 2021 | Learning Non-Linear Disentangled Editing For StyleganabstractRecent work has demonstrated the great potential of image editing in the latent space of powerful deep generative models such as StyleGAN. However, the success of such methods relies on the assumption that a linear hyperplane may separate the latent space into two subspaces for a binary attribute. In this work, we show that this hypothesis is a significant limitation and propose to learn a non-linear, regularized and identity-preserving latent space transformation that leads to more accurate and disentangled manipulations of facial attributes. Alasdair Newson, Yann Gousseau, Pierre Hellier |
ICIP | 2 |
| 2020 | CARRADA Dataset: Camera and Automotive Radar with Range- Angle- Doppler AnnotationsabstractHigh quality perception is essential for autonomous driving (AD) systems. To reach the accuracy and robustness thatare required by such systems, several types of sensors must be combined. Currently, mostly cameras and laser scanners (lidar) are deployed to build a representation of the world around the vehicle. While radar sensors have been used fora long time in the automotive industry, they are still under-used for AD despite their appealing characteristics (notably, their ability to measure the relative speed of obstacles and to operate even in adverse weather conditions). To alarge extent, this situation is due to the relative lack of automotive datasets with real radar signals that are both raw and annotated. In this work, we introduce CARRADA, a dataset of synchronized camera and radar recordings with range-angle-Doppler annotations. We also present a semi-automatic annotation approach, which was used to annotate the dataset, and a radar semantic segmentation baseline, which we evaluate on several metrics. Both our code and dataset are available online. Arthur Ouaknine, Alasdair Newson, Julien Rebut, Florence Tupin, Patrick Pérez |
ICPR | 2 |
| 2020 | High Resolution Face Age EditingabstractFace age editing has become a crucial task in film post-production, and is also becoming popular for general purpose photography. Recently, adversarial training has produced some of the most visually impressive results for image manipulation, including the face aging/de-aging task. In spite of considerable progress, current methods often present visual artifacts and can only deal with low-resolution images. In order to achieve aging/de-aging with the high quality and robustness necessary for wider use, these problems need to be addressed. This is the goal of the present work. We present an encoder-decoder architecture for face age editing. The core idea of our network is to encode a face image to age-invariant features, and learn a modulation vector corresponding to a target age. We then combine these two elements to produce a realistic image of the person with the desired target age. Our architecture is greatly simplified with respect to other approaches, and allows for fine-grained age editing on high resolution images in a single unified model. Source codes are available at https://github.com/InterDigitalInc/HRFAE. Gilles Puy, Alasdair Newson, Yann Gousseau, Pierre Hellier |
ICPR | 3 |
| 2017 | A Stochastic Film Grain Model for Resolution-Independent RenderingabstractAbstract The realistic synthesis and rendering of film grain is a crucial goal for many amateur and professional photographers and film‐makers whose artistic works require the authentic feel of analogue photography. The objective of this work is to propose an algorithm that reproduces the visual aspect of film grain texture on any digital image. Previous approaches to this problem either propose unrealistic models or simply blend scanned images of film grain with the digital image, in which case the result is inevitably limited by the quality and resolution of the initial scan. In this work, we introduce a stochastic model to approximate the physical reality of film grain, and propose a resolution‐free rendering algorithm to simulate realistic film grain for any digital input image. By varying the parameters of this model, we can achieve a wide range of grain types. We demonstrate this by comparing our results with film grain examples from dedicated software, and show that our rendering results closely resemble these real film emulsions. In addition to realistic grain rendering, our resolution‐free algorithm allows for any desired zoom factor, even down to the scale of the microscopic grains themselves. Alasdair Newson, Julie Delon, Bruno Galerne |
Comput. Graph. Forum | 1 |
| 2015 | Low-Rank Spatio-Temporal Video SegmentationabstractRecently, a great deal of interest has been generated by the technique known as Robust Principle Component Analysis (RPCA) of Candes et al. [1], which addresses the problem of separating a matrix into a low-rank and a sparse component. This very general formulation can be used for tasks such as background estimation in videos and face recognition. In the case of background estimation, the low-rank matrix models the background, and the sparse matrix corresponds to the foreground. A considerable drawback of this approach is its poor robustness to local lighting conditions. If lighting conditions vary locally, one of two things may happen. Either the method incorporates the lighting variation into the foreground, which is clearly undesirable, or the rank of the background model is allowed to increase. Unfortunately, this second option means that the true foreground is likely to become included in the background, especially for objects which are static for a short while. Here, we propose to model the background as a piece-wise low-rank matrix. In this manner, it will be possible to extract several localised models which correspond to coherent lighting conditions. However, for this we need to segment the input video into such coherent regions. We refer to this problem as a low-rank spatio-temporal video segmentation. We present an algorithm to address this segmentation problem, based on region merging and spectral clustering techniques. We show that by carrying out a local RPCA in each region, the results of foreground/background separation are greatly improved, in comparison with both the standard RPCA and several other well-known background estimation techniques. Let X ∈ Rm×n represent an input video, in matrix form. Each frame contains m pixels, and there are a total of n frames in our video. The goal of RPCA is to decompose X as X≈ L+S, where L is the low-rank matrix and S is the sparse matrix. Unfortunately, the rank of a matrix is a non-convex function, so a surrogate function, the nuclear norm is used. Thus, the background/foreground separation problem may be formulated as follows: Alasdair Newson, Mariano Tepper, Guillermo Sapiro |
BMVC | 1 |
| 2015 | Multi-temporal foreground detection in videosabstractA common task in video processing is the binary separation of a video's content into either background or moving foreground. However, many situations require a foreground analysis with a finer temporal granularity, in particular for objects or people which remain immobile for a certain period of time. We propose an efficient method which detects foreground at different timescales, by exploiting the desirable theoretical and practical properties of Robust Principal Component Analysis. Our algorithm can be used in a variety of scenarios such as detecting people who have fallen in a video, or analysing the fluidity of road traffic, while avoiding costly computations needed for nearest neighbours searches or optical flow analysis. Finally, our algorithm has the useful ability to perform motion analysis without explicitly requiring computationally expensive motion estimation. Mariano Tepper, Alasdair Newson, Pablo Sprechmann, Guillermo Sapiro |
ICIP | 2 |
| 2014 | Video Inpainting of Complex ScenesabstractWe propose an automatic video inpainting algorithm which relies on the optimization of a global, patch-based functional. Our algorithm is able to deal with a variety of challenging situations which naturally arise in video inpainting, such as the correct reconstruction of dynamic textures, multiple moving objects, and moving background. Furthermore, we achieve this in an order of magnitude less execution time with respect to the state-of-the-art. We are also able to achieve good quality results on high-definition videos. Finally, we provide specific algorithmic details to make implementation of our algorithm as easy as possible. The resulting algorithm requires no segmentation or manual input other than the definition of the inpainting mask and can deal with a wider variety of situations than is handled by previous work. Alasdair Newson, Andrés Almansa, Matthieu Fradet, Yann Gousseau, Patrick Pérez |
SIAM J. Imaging Sci. | 1 |
| 2014 | Robust Automatic Line Scratch Detection in FilmsabstractLine scratch detection in old films is a particularly challenging problem due to the variable spatiotemporal characteristics of this defect. Some of the main problems include sensitivity to noise and texture, and false detections due to thin vertical structures belonging to the scene. We propose a robust and automatic algorithm for frame-by-frame line scratch detection in old films, as well as a temporal algorithm for the filtering of false detections. In the frame-by-frame algorithm, we relax some of the hypotheses used in previous algorithms in order to detect a wider variety of scratches. This step's robustness and lack of external parameters is ensured by the combined use of an a contrario methodology and local statistical estimation. In this manner, over-detection in textured or cluttered areas is greatly reduced. The temporal filtering algorithm eliminates false detections due to thin vertical structures by exploiting the coherence of their motion with that of the underlying scene. Experiments demonstrate the ability of the resulting detection procedure to deal with difficult situations, in particular in the presence of noise, texture, and slanted or partial scratches. Comparisons show significant advantages over previous work. Alasdair Newson, Andrés Almansa, Yann Gousseau, Patrick Pérez |
IEEE Trans. Image Process. | 1 |
| 2013 | Temporal filtering of line scratch detections in degraded filmsabstractThe film defect known as the line scratch is difficult to restore automatically due to the large number of false alarms present in scratch detection algorithms. In this paper, an algorithm for dealing with these false alarms is proposed. Validating true scratches, which is the approach generally proposed in the literature, is a difficult task since scratch characteristics are hard to determine, making tracking these defects problematic. Instead, we eliminate false alarms by analysing their compatibility with a global motion estimation. We compare our algorithm with two other scratch detection methods from the literature. Experiments show that our algorithm outperforms these two, and that the proposed temporal filtering greatly improves precision while maintaining high recall. Alasdair Newson, Andrés Almansa, Yann Gousseau, Patrick Pérez |
ICIP | 1 |