Axel Davy

dblp:163/2221 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0002-2143-1542ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021
YearPublicationVenuePosition
2025 S3VD Self-Supervised Spatial Video Downsampling Loss: A Method for Training Video FPN Denoising Networks
abstract
Fixed pattern noise (FPN) is a temporally constant noise present on videos due to the non-uniformities of the sensors that may exhibit spatial correlation, typically across columns and/or rows. Acquiring real clean/noisy data is particularly challenging in the case of FPN, leading supervised FPN denoising networks to train using generated data. Self-supervised approaches for denoising allow training directly on real noisy sequences, avoiding the biases introduced by synthetic data. However, the spatial and temporal correlation of FPN violates noise independence assumptions underlying most self-supervised approaches. In this paper, we propose for the first time, a method for training video column FPN denoising networks in a self-supervised way. Our approach consists of spatial downsampling on rows or columns to obtain quasi-two independent noisy observations from the same images to train a network on. The proposed method can be applied to any network architecture. We demonstrate the effectiveness of our method with extensive experiments on synthetic FPN and publicly available real infrared data.
Hortensia Barral, Pablo Arias 0001, Axel Davy
ICIP3
2024 Fixed Pattern Noise Removal For Multi-View Single-Sensor Infrared Camera
abstract
Fixed pattern noise (FPN) is a temporally coherent noise present on videos due to the non-uniformities in the response of the imaging sensor. It is a common problem for infrared videos which degrades the quality of the observation and hinders subsequent applications. In this work we introduce a generalization of the FPN removal problem where the input data consists of several different sequences with the same FPN. This is motivated by infrared cameras that capture multiple views with a single sensor via a periodic motion pattern of a mirror or the camera itself, such as those used in surveillance. This multi-view setting allows for a much more accurate estimation of the FPN in comparison with the standard FPN removal problem from a single view. We propose a novel energy minimization approach for multi-view FPN removal, and two optimization algorithms that can be applied both in an off-line and online manner. In addition, we show that the proposed energy can be adapted to the problem of FPN removal from a single view with a rolling window approach, obtaining a significant improvement over the state of the art. We demonstrate the performance of the proposed method with synthetic data and real data from surveillance infrared cameras.
Arnaud Barral, Pablo Arias 0001, Axel Davy
WACV3
2024 On the Importance of Large Objects in CNN Based Object Detection Algorithms
abstract
Object detection models, a prominent class of machine learning algorithms, aim to identify and precisely locate objects in images or videos. However, this task might yield uneven performances sometimes caused by the objects sizes and the quality of the images and labels used for training. In this paper, we highlight the importance of large objects in learning features that are critical for all sizes. Given these findings, we propose to introduce a weighting term into the training loss. This term is a function of the object area size. We show that giving more weight to large objects leads to improved detection scores across all object sizes and so an overall improvement in Object Detectors performances (+2 p.p. of mAP on small objects, +2 p.p. on medium and +4 p.p. on large on COCO val 2017 with InternImage-T). Additional experiments and ablation studies with different models and on a different dataset further confirm the robustness of our findings.
Ahmed Ben Saad, Gabriele Facciolo, Axel Davy
WACV3
2023 Improving Pixel-Level Contrastive Learning by Leveraging Exogenous Depth Information
abstract
Self-supervised representation learning based on Contrastive Learning (CL) has been the subject of much attention in recent years. This is due to the excellent results obtained on a variety of subsequent tasks (in particular classification), without requiring a large amount of labeled samples. However, most reference CL algorithms (such as SimCLR and MoCo, but also BYOL and Barlow Twins) are not adapted to pixel-level downstream tasks. One existing solution known as PixPro proposes a pixel-level approach that is based on filtering of pairs of positive/negative image crops of the same image using the distance between the crops in the whole image. We argue that this idea can be further enhanced by incorporating semantic information provided by exogenous data as an additional selection filter, which can be used (at training time) to improve the selection of the pixel-level positive/negative samples. In this paper we will focus on the depth information, which can be obtained by using a depth estimation network or measured from available data (stereovision, parallax motion, LiDAR, etc.). Scene depth can provide meaningful cues to distinguish pixels belonging to different objects based on their depth. We show that using this exogenous information in the contrastive loss leads to improved results and that the learned representations better follow the shapes of objects. In addition, we introduce a multi-scale loss that alleviates the issue of finding the training parameters adapted to different object sizes. We demonstrate the effectiveness of our ideas on the Breakout Segmentation on Borehole Images where we achieve an improvement of 1.9% over PixPro and nearly 5% over the supervised baseline. We further validate our technique on the indoor scene segmentation tasks with ScanNet and outdoor scenes with CityScapes (1.6% and 1.1% improvement over PixPro respectively).
Ahmed Ben Saad, Kristina Prokopetc, Josselin Kherroubi, Axel Davy, Adrien Courtois, Gabriele Facciolo
WACV4
2022 Self-Supervised Super-Resolution for Multi-Exposure Push-Frame Satellites
abstract
Modern Earth observation satellites capture multi-exposure bursts of push-frame images that can be super-resolved via computational means. In this work, we propose a super-resolution method for such multi-exposure sequences, a problem that has received very little attention in the literature. The proposed method can handle the signal-dependent noise in the inputs, process sequences of any length, and be robust to inaccuracies in the exposure times. Furthermore, it can be trained end-to-end with self-supervision, without requiring ground truth high resolution frames, which makes it especially suited to handle real data. Central to our method are three key contributions: i) a base-detail decomposition for handling errors in the exposure times, ii) a noise-level-aware feature encoding for improved fusion of frames with varying signal-to-noise ratio and iii) a permutation invariant fusion strategy by temporal pooling operators. We evaluate the proposed method on synthetic and real data and show that it outperforms by a significant margin existing single-exposure approaches that we adapted to the multi-exposure case.
Ngoc Long Nguyen, Jérémy Anger, Axel Davy, Pablo Arias 0001, Gabriele Facciolo
CVPR3
2022 Self-Supervised Push-Frame Super-Resolution With Detail-Preserving Control And Outlier Detection
abstract
Self-supervised training enables the application of deep-learning based methods for multi-image super-resolution of satellite imagery. In this work we propose two improvements on the self-supervised Deep-Shift-and-Add (DSA) method introduced by Nguyen et al. First, we demonstrate how the self-supervised loss of DSA can be extended to provide the image interpreter with a spatially varying parameter to control the trade-off between detail preservation and noise removal at test time. Second, we endow the DSA architecture with a mechanism that enables the network to be robust to outliers produced for example by dead pixels, reflections or registration errors.
Ngoc Long Nguyen, Jérémy Anger, Axel Davy, Pablo Arias 0001, Gabriele Facciolo
IGARSS3
2021 PROBA-V-REF: Repurposing the PROBA-V Challenge for Reference-Aware Super Resolution
abstract
The PROBA-V Super-Resolution challenge distributes real low-resolution image series and corresponding high-resolution targets to advance research on Multi-Image Super Resolution (MISR) for satellite images. However, in the PROBA-V dataset the low-resolution image corresponding to the high-resolution target is not identified. We argue that in doing so, the challenge ranks the proposed methods not only by their MISR performance, but mainly by the heuristics used to guess which image in the series is the most similar to the high-resolution target. We demonstrate this by improving the performance obtained by the two winners of the challenge only by using a different reference image, which we compute following a simple heuristic. Based on this, we propose PROBA-V-REF a variant of the PROBA-V dataset, in which the reference image in the low-resolution series is provided, and show that the ranking between the methods changes in this setting. This is relevant to many practical use cases of MISR where the goal is to super-resolve a specific image of the series, i.e. the reference is known. The proposed PROBA-V-REF should better reflect the performance of the different methods for this reference-aware MISR problem.
Ngoc Long Nguyen, Jérémy Anger, Axel Davy, Pablo Arias 0001, Gabriele Facciolo
IGARSS3
2021 Self-supervised training for blind multi-frame video denoising
abstract
We propose a self-supervised approach for training multi-frame video denoising networks. These networks predict each frame from a stack of frames around it. Our self-supervised approach benefits from the temporal consistency in the video by minimizing a loss that penalizes the difference between the predicted frame and a neighboring one, after aligning them using an optical flow. We use the proposed strategy to denoise a video contaminated with an unknown noise type, by fine-tuning a pre-trained denoising network on the noisy video. The proposed fine-tuning reaches and sometimes surpasses the performance of state-of-the-art networks trained with supervision. We demonstrate this by showing extensive results on video blind denoising of different synthetic and real noises. In addition, the proposed fine-tuning can be applied to any parameter that controls the denoising performance of the network. We show how this can be expoited to perform joint denoising and noise level estimation for heteroscedastic noise.
Valéry Dewil, Jérémy Anger, Axel Davy, Thibaud Ehret, Gabriele Facciolo, Pablo Arias 0001
WACV3
2021 Fast, Nonlocal and Neural: A Lightweight High Quality Solution to Image Denoising
abstract
With the widespread application of convolutional neural networks (CNNs), the traditional model based denoising algorithms are now outperformed. However, CNNs face two problems. First, they are computationally demanding, which makes their deployment especially difficult for mobile terminals. Second, experimental evidence shows that CNNs often over-smooth regular textures present in images, in contrast to traditional non-local models. In this letter, we propose a solution to both issues by combining a nonlocal algorithm with a lightweight residual CNN. s solution gives full latitude to the advantages of both models. We apply this framework to two GPU implementations of classic nonlocal algorithms (NLM and BM3D) and observe a substantial gain in both cases, performing better than the state-of-the-art with low computational requirements. Our solution is between 10 and 20 times faster than CNNs with equivalent performance and attains higher PSNR. In addition the final method shows a notable gain on images containing complex textures like the ones of the MIT Moiré dataset.
Yu Guo 0008, Axel Davy, Gabriele Facciolo, Jean-Michel Morel, Qiyu Jin
IEEE Signal Process. Lett.2
2019 Model-Blind Video Denoising via Frame-To-Frame Training
abstract
Modeling the processing chain that has produced a video is a difficult reverse engineering task, even when the camera is available. This makes model based video processing a still more complex task. In this paper we propose a fully blind video denoising method, with two versions off-line and on-line. This is achieved by fine-tuning a pre-trained AWGN denoising network to the video with a novel frame-to-frame training strategy. Our denoiser can be used without knowledge of the origin of the video or burst and the post-processing steps applied from the camera sensor. The on-line process only requires a couple of frames before achieving visually pleasing results for a wide range of perturbations. It nonetheless reaches state-of-the-art performance for standard Gaussian noise, and can be used off-line with still better performance.
Thibaud Ehret, Axel Davy, Jean-Michel Morel, Gabriele Facciolo, Pablo Arias 0001
CVPR2
2019 Joint Demosaicking and Denoising by Fine-Tuning of Bursts of Raw Images
abstract
Demosaicking and denoising are the first steps of any camera image processing pipeline and are key for obtaining high quality RGB images. A promising current research trend aims at solving these two problems jointly using convolutional neural networks. Due to the unavailability of ground truth data these networks cannot be currently trained using real RAW images. Instead, they resort to simulated data. In this paper we present a method to learn demosaicking directly from mosaicked images, without requiring ground truth RGB data. We apply this to learn joint demosaicking and denoising only from RAW images, thus enabling the use of real data. In addition we show that for this application fine-tuning a network to a specific burst improves the quality of restoration for both demosaicking and denoising.
Thibaud Ehret, Axel Davy, Pablo Arias 0001, Gabriele Facciolo
ICCV2
2019 Detection of Small Anomalies on Moving Background
abstract
We consider the problem of detecting small targets in videos where the textured background is also possibly moving. The proposed method is based on a two steps statistical framework. In a first step, the optical flow is computed using a pyramidal scheme incorporating statistical tests for a result with reliability guarantees. In the second step, the detection of targets is changed into a problem of anomaly detection in noise, and statistical testing ensures a control of the number of false detections.
Axel Davy, Agnès Desolneux, Jean-Michel Morel
ICIP1
2019 A Non-Local CNN for Video Denoising
abstract
Non-local patch-based methods were until recently state-of-the-art for image denoising but are now outperformed by convolutional neural networks (CNNs). Yet they are still the best ones for video denoising, as video redundancy is a key factor to attain high denoising performance. In this work we propose a novel video denoising CNN. Non-local self-similarity is incorporated into the network via a first non-trainable layer which finds for each patch in the input image its most similar patches in a 3D spatio-temporal search region centered at the target patch. The central values of these patches are then gathered in a feature vector which is assigned to each image pixel. This information is presented to a CNN which is trained to predict a clean image. The proposed architecture achieves state-of-the-art results. To the best of our knowledge, this is the first successful application of CNNs to video denoising.
Axel Davy, Thibaud Ehret, Jean-Michel Morel, Pablo Arias 0001, Gabriele Facciolo
ICIP1
2019 A Novel Activity Detector Applied to Sentinel-1 for Surveillance
abstract
Change detection is the challenging process of identifying meaningful changes on an image sequence. In this work, we propose a new change detector for SAR images. We show that by controlling the a priori number of false alarms, one can detect events and filter out very small detections without first applying an anti-speckle filter. A detection is made in the absolute log-ratio image when enough pixels in a neighborhood exceed a statistically pre-determined threshold. This flexible approach allows several neighborhood shapes and sizes to be tested while controlling the number of false alarms. A global bound on the desired expected number of false alarms automatically determines the detection thresholds for each tested configuration.
Axel Davy, Max Dunitz
IGARSS1
2018 Reducing Anomaly Detection in Images to Detection in Noise
abstract
Anomaly detectors address the difficult problem of detecting automatically exceptions in an arbitrary background image. Detection methods have been proposed by the thousands because each problem requires a different background model. By analyzing the existing approaches, we show that the problem can be reduced to detecting anomalies in residual images (extracted from the target image) in which noise and anomalies prevail. Hence, the general and impossible background modeling problem is replaced by simpler noise modeling, and allows the calculation of rigorous thresholds based on the a contrario detection theory. Our approach is therefore unsupervised and works on arbitrary images.
Axel Davy, Thibaud Ehret, Jean-Michel Morel, Mauricio Delbracio
ICIP1
2017 Brain tumor segmentation with Deep Neural Networks
Mohammad Havaei, Axel Davy, David Warde-Farley, Antoine Biard, Aaron C. Courville, Yoshua Bengio, Christopher Joseph Pal, Pierre-Marc Jodoin, Hugo Larochelle
Medical Image Anal.2