Olivier Le Meur

dblp:35/3436 · DBLP profile ↗
← Back
71ranked-venue papers
20as first author
13since 2021 · last 2025
0000-0001-9883-0296ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 65 · 19 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 3 since 2021Computer networks · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Style-FG: A Style-based Framework for Film Grain Analysis and Synthesis
abstract
Film grain which used to be a by-product of the chemical processing in the analog film stock is a desirable feature in the era of digital cameras. Besides participating to the artistic intent during content creation, film grain has also interesting properties in the video compression chain such as its ability to mask compression artifacts. In this article, we use a deep learning-based framework for film grain analysis, generation, and synthesis. Our framework Style-FG consists of three modules: a style encoder performing film grain style analysis, a mapping network responsible for film grain style generation, and a synthesis network that generates and blends a specific grain style to a given content in a content-adaptive manner. All modules are trained jointly, thanks to dedicated loss functions, on a new large and diverse dataset of pairs of grain-free and grainy images that we made publicly available to the community. 1 Quantitative and qualitative evaluations show that fidelity to the reference grain, diversity of grain styles as well as a perceptually pleasant grain synthesis are achieved, demonstrating that each module outperforms the state-of-the-art in the task it was designed for. To contribute further to the sustainability necessary effort of the digital information and communication field, a light-weight version of Style-FG is also proposed, which demonstrates similar quantitative and qualitative performances, while reducing the number of network parameters by a factor of 92%.
Zoubida Ameur, Claire-Hélène Demarty, Olivier Le Meur, Daniel Ménard
ACM Trans. Multim. Comput. Commun. Appl.3
2024 3R-INN: How to Be Climate Friendly While Consuming/Delivering Videos?
Zoubida Ameur, Claire-Hélène Demarty, Daniel Ménard, Olivier Le Meur
ECCV (75)4
2023 Display Power Modeling for Energy Consumption Control
abstract
As the most consuming devices in the video chain, it is necessary to master the power consumption of displays. However it necessitates to have a precise modeling of their power consumption. This paper proposes two approaches that outperform the state-of-the-art, to model the power consumption of emissive displays: first through the estimation of the display parameters from a theoretical RGBW model, then through two deep-based models agnostic of the display technology, that output either a global power value or a power map corresponding to an input image. All three modelings were made possible thanks to the release of DISPLAYPOWER2k, a large dataset of power consumption values for OLED displays.
Claire-Hélène Demarty, Laurent Blondé, Olivier Le Meur
ICIP3
2023 Deep-Learning-Based Energy Aware Images
abstract
In this paper, we present a method to compute energy-aware images, that aims to reduce the energy consumption of displays. This method relies on a lightweight unsupervised deep model which finds out the best trade-off between visual quality and energy reduction. From an input image and an energy reduction rate, a dimming map is inferred. We show that the proposed model performs as good as state-of-the-art methods, while being much more simple. In addition, the dimming map computation is constrained in order to ease its distribution throughout the video chain.
Olivier Le Meur, Claire-Hélène Demarty, Laurent Blondé
ICIP1
2023 Energy-Aware HDR Content End-to-End VVC Encoding
abstract
The proposed demonstration combines several commercially viable state-of-the-art technologies to achieve significant energy reduction while delivering HDR video without compromising QoE. The components of the demo are: (i) The demo is based our solution on VVC, the most efficient, highest quality video codec available today. (ii) We leverage Advanced HDR by Technicolor, a solution that uniquely can deliver HDR and SDR video using a single stream. (iii) We introduce a solution for dynamically adjusting peak brightness of displays via display adaptation controlled through the network using dynamic metadata encapsulated in MPEG SEI messages. This solution has been tested and shows a capability of reducing power by 10% or more without any subjective quality loss, and far more with minimal subjective impact. (iv) While we demonstrate a streaming (unicast) scenario in our live demo, our system is also capable of working in a dynamic unicast/multicast environment based on DVB transmission standards. (v) The entire demonstration is resulting from an industrial collaboration and is based on ready to market technologies and prototypes.
Olivier Le Meur, Franck Aumont, Juan-Carlos Vargas, Pierre-Loup Cabarat, Daniel Ménard, Oussama Hammami, Thomas Guionnet
MMSP1
2023 Style-based film grain analysis and synthesis
abstract
Film grain which used to be a by-product of the chemical processing in the analog film stock, is a desirable feature in the era of digital cameras. Besides participating to the artistic intent during content creation, film grain has also interesting properties in the video compression chain such as its ability to mask compression artifacts. In this paper, we use a deep learning-based framework for film grain analysis, generation and synthesis. Our framework consists of three modules: a style encoder performing film grain style analysis, a mapping network responsible for film grain style generation, and a synthesis network that generates and blends a specific grain style to a given content in a content-adaptive manner. All modules are trained jointly, thanks to dedicated loss functions, on a new large and diverse dataset of pairs of grain-free and grainy images that we made publicly available to the community1. Quantitative and qualitative evaluations show that fidelity to the reference grain, diversity of grain styles as well as a perceptually pleasant grain synthesis are achieved, demonstrating that each module outperforms the state-of-the-art in the task it was designed for.
Zoubida Ameur, Claire-Hélène Demarty, Olivier Le Meur, Daniel Ménard, Edouard François
MMSys3
2023 HDR-LFNet: Inverse tone mapping using fusion network
Mathieu Chambe, Ewa Kijak, Zoltán Miklós 0001, Olivier Le Meur, Rémi Cozot, Kadi Bouatouch
Comput. Graph.4
2023 Invertible Energy-Aware Images
abstract
Displaying video content requires significant amounts of energy. In face of the energy crisis and climate emergency, we propose a new approach to produce energy-aware images. The purpose of such images is to consume less energy when displayed onscreen while maximizing the quality of experience. A new invertible neural network called Invertible Energy-Aware Network (InvEAN) is proposed to produce invertible energy-aware images, allowing to reduce the energy consumption of display devices and offering the possibility to recover the original image if required. Experimental results show that the InvEAN network outperforms two existing methods over three datasets.
Olivier Le Meur, Claire-Hélène Demarty
IEEE Signal Process. Lett.1
2022 Deep learning for assessing the aesthetics of professional photographs
abstract
Abstract Aesthetic quality assessment for photographs is an important research topic since it can be used for a number of applications, such as image database management or image browsing. In 2012, the aesthetic visual analysis (AVA) dataset has been proposed. It has since then been used to train the majority of computational models of aesthetics assessment. We observe that AVA is mainly composed of competitive photographs which notion of aesthetics differs from other kinds of photographs, such as professional photographs. In this paper, we evaluate whether or not recent aesthetics assessment models generalize well and perform well over professional photographs. We noticed that the different models we tested behave differently on both categories, and therefore do not generalize well. Besides, we fine‐tuned one of the tested model using professional photographs and the results show that this fine‐tuning is effectively improving the coverage of the methods.
Mathieu Chambe, Rémi Cozot, Olivier Le Meur
Comput. Animat. Virtual Worlds3
2022 Trajectory Saliency Detection Using Consistency-Oriented Latent Codes From a Recurrent Auto-Encoder
abstract
In this paper, we are concerned with the detection of progressive dynamic saliency from video sequences. More precisely, we are interested in saliency related to motion and likely to appear progressively over time. It can be relevant to trigger alarms, to dedicate additional processing or to detect specific events. Trajectories represent the best way to support progressive dynamic saliency detection. Accordingly, we will talk about trajectory saliency. A trajectory will be qualified as salient if it deviates from normal trajectories that share a common motion pattern related to a given context. First, we need a compact while discriminative representation of trajectories. We adopt a (nearly) unsupervised learning-based approach. The latent code estimated by a recurrent auto-encoder provides the desired representation. In addition, we enforce consistency for normal (similar) trajectories through the auto-encoder loss function. The distance of the trajectory code to a prototype code accounting for normality is the means to detect salient trajectories. We validate our trajectory saliency detection method on synthetic and real trajectory datasets, and highlight the contributions of its different components. We compare our method favourably to existing methods on several saliency configurations constructed from the publicly available large dataset of pedestrian trajectories acquired in a railway station.
Léo Maczyta, Patrick Bouthemy, Olivier Le Meur
IEEE Trans. Circuits Syst. Video Technol.3
2021 Example-Based Colour Transfer for 3D Point Clouds
abstract
Abstract Example‐based colour transfer between images, which has raised a lot of interest in the past decades, consists of transferring the colour of an image to another one. Many methods based on colour distributions have been proposed, and more recently, the efficiency of neural networks has been demonstrated again for colour transfer problems. In this paper, we propose a new pipeline with methods adapted from the image domain to automatically transfer the colour from a target point cloud to an input point cloud. These colour transfer methods are based on colour distributions and account for the geometry of the point clouds to produce a coherent result. The proposed methods rely on simple statistical analysis, are effective, and succeed in transferring the colour style from one point cloud to another. The qualitative results of the colour transfers are evaluated and compared with existing methods.
Ific Goudé, Rémi Cozot, Olivier Le Meur, Kadi Bouatouch
Comput. Graph. Forum3
2021 Deep saliency models : The quest for the loss function
Alexandre Bruckert, Hamed Rezazadegan Tavakoli, Zhi Liu 0003, Marc Christie, Olivier Le Meur
Neurocomputing5
2021 Predicting atypical visual saliency for autism spectrum disorder via scale-adaptive inception module and discriminative region enhancement loss
Weijie Wei 0001, Zhi Liu 0003, Lijin Huang, Alexis Nebout, Olivier Le Meur, Tianhong Zhang, Jijun Wang 0003, Lihua Xu
Neurocomputing5
2020 Eye-Gaze Activity in Crowds: Impact of Virtual Reality and Density
abstract
When we are walking in crowds, we mainly use visual information to avoid collisions with other pedestrians. Thus, gaze activity should be considered to better understand interactions between people in a crowd. In this work, we use Virtual Reality (VR) to facilitate motion and gaze tracking, as well as to accurately control experimental conditions, in order to study the effect of crowd density on eye-gaze behavior. Our motivation is to better understand how interaction neighborhood (i.e., the subset of people actually influencing one’s locomotion trajectory) changes with density. To this end, we designed two experiments. The first one evaluates the biases introduced by the use of VR on the visual activity when walking among people, by comparing eye-gaze activity while walking in a real and virtual street. We then designed a second experiment where participants walked in a virtual street with different levels of pedestrian density. We demonstrate that gaze fixations are performed at the same frequency despite increases in pedestrian density, while the eyes scan a narrower portion of the street. These results suggest that in such situations walkers focus more on people in front and closer to them. These results provide valuable insights regarding eye-gaze activity during interactions between people in a crowd, and suggest new recommendations in designing more realistic crowd simulations.
Florian Berton, Ludovic Hoyet, Anne-Hélène Olivier, Julien Bruneau 0002, Olivier Le Meur, Julien Pettré
VR5
2020 Effective schizophrenia recognition using discriminative eye movement features and model-metric based features
Lijin Huang, Weijie Wei 0001, Zhi Liu 0003, Tianhong Zhang, Jijun Wang 0003, Lihua Xu, Olivier Le Meur
Pattern Recognit. Lett.8
2020 FANet: Features Adaptation Network for 360$^{\circ }$ Omnidirectional Salient Object Detection
abstract
Salient object detection (SOD) in 360° omnidirectional images has become an eye-catching problem because of the popularity of affordable 360° cameras. In this paper, we propose a Features Adaptation Network (FANet) to highlight salient objects in 360° omnidirectional images reliably. To utilize the feature extraction capability of convolutional neural networks and capture global object information, we input the equirectangular 360° images and corresponding cube-map 360° images to the feature extraction network (FENet) simultaneously to obtain multi-level equirectangular and cube-map features. Furthermore, we fuse these two kinds of features at each level of the FENetby a projection features adaptation (PFA) module, for selecting these two kinds of features adaptively. Finally, we combine the preliminary adaptation features at different levels by a multi-level features adaptation (MLFA) module, which weights these different-level features adaptively and produces the final saliency maps. Experiments show our FANet outperforms the state-of-the-art methods on the 360° omnidirectional SOD datasets.
Mengke Huang, Zhi Liu 0003, Gongyang Li, Xiaofei Zhou 0003, Olivier Le Meur
IEEE Signal Process. Lett.5
2019 How Well Current Saliency Prediction Models Perform on UAVs Videos?
Anne-Flore Perrin, Lu Zhang 0037, Olivier Le Meur
CAIP (1)3
2019 Deep Learning For Inter-Observer Congruency Prediction
abstract
According to the literature regarding visual saliency, observers may exhibit considerable variations in their gaze behaviors. These variations are influenced by aspects such as cultural background, age or prior experiences, but also by features in the observed images. The dispersion between the gaze of different observers looking at the same image is commonly referred as inter-observer congruency (IOC). Predicting this congruence can be of great interest when it comes to study the visual perception of an image. In this paper, we introduce a new method based on deep learning techniques to predict the IOC of an image. This is achieved by first extracting features from an image through a deep convolutional network. We then show that using such features to train a model with a shallow network regression technique significantly improves the precision of the prediction over existing approaches.
Alexandre Bruckert, Yat Hong Lam, Marc Christie, Olivier Le Meur
ICIP4
2019 Unsupervised Motion Saliency Map Estimation Based On Optical Flow Inpainting
abstract
The paper addresses the problem of motion saliency in videos, that is, identifying regions that undergo motion departing from its context. We propose a new unsupervised paradigm to compute motion saliency maps. The key ingredient is the flow inpainting stage. Candidate regions are determined from the optical flow boundaries. The residual flow in these regions is given by the difference between the optical flow and the flow inpainted from the surrounding areas. It provides the cue for motion saliency. The method is flexible and general by relying on motion information only. Experimental results on the DAVIS 2016 benchmark demonstrate that the method compares favourably with state-of-the-art video saliency methods.
Léo Maczyta, Patrick Bouthemy, Olivier Le Meur
ICIP3
2019 Quality Metric Aggregation for HDR/WCG Images
abstract
High Dynamic Range (HDR) and Wide Color Gamut (WCG) screens are able to display images with brighter and darker pixels with more vivid colors than ever. Automatically assessing the quality of these HDR/WCG images is of critical importance to evaluate the performances of image compression schemes. In recent years, full-reference metrics, such as HDR-VDP-2, PU-encoding metrics, have been designed for this purpose. However, none of these metrics consider chromatic artifacts. In this paper, we propose our own full-reference quality metric adapted to HDR and WCG content that is sensitive to chromatic distortions. The proposed metric is based on two existing HDR quality metrics and color image features. A support vector machine regression is used to combine the aforementioned features. Experimental results demonstrate the effectiveness of the proposed metric in the context of image compression.
Maxime Rousselot, Xavier Ducloux, Olivier Le Meur, Rémi Cozot
ICIP3
2019 Saliency-aware inter-image color transfer for image manipulation
Zhi Liu 0003, Qihan Jiao, Olivier Le Meur, Wanlei Zhao
Multim. Tools Appl.4
2019 CNN-based temporal detection of motion saliency in videos
Léo Maczyta, Patrick Bouthemy, Olivier Le Meur
Pattern Recognit. Lett.3
2019 Context in Photo Albums: Understanding and Modeling User Behavior in Clustering and Selection
abstract
Recent progress in digital photography and storage availability has drastically changed our approach to photo creation. While in the era of film cameras, careful forethought would usually precede the capture of a photo; nowadays, a large number of pictures can be taken with little effort. One of the consequences is the creation of numerous photos depicting the same moment in slightly different ways, which makes the process of organizing photos laborious for the photographer. Nevertheless, photo collection organization is important both for exploring photo albums and for simplifying the ultimate task of selecting the best photos. In this work, we conduct a user study to explore how users tend to organize or cluster similar photos in albums, to what extent different users agree in their clustering decisions, and to investigate how the clustering-defined photo context affects the subsequent photo-selection process. We also propose an automatic hierarchical clustering solution for modeling user clustering decisions. To demonstrate the usefulness of our approach, we apply it to the task of automatic photo evaluation within photo albums and propose a clustering-based context adaptation.
Dmitry Kuzovkin, Tania Pouli, Olivier Le Meur, Rémi Cozot, Jonathan Kervec, Kadi Bouatouch
ACM Trans. Appl. Percept.3
2018 Deepcomics: saliency estimation for comics
abstract
A key requirement for training deep learning saliency models is large training eye tracking datasets. Despite the fact that the accessibility of eye tracking technology has greatly increased, collecting eye tracking data on a large scale for very specific content types is cumbersome, such as comic images, which are different from natural images such as photographs because text and pictorial content is integrated. In this paper, we show that a deep network trained on visual categories where the gaze deployment is similar to comics outperforms existing models and models trained with visual categories for which the gaze deployment is dramatically different from comics. Further, we find that it is better to use a computationally generated dataset on visual category close to comics one than real eye tracking data of a visual category that has different gaze deployment. These findings hold implications for the transference of deep networks to different domains.
Kévin Bannier, Eakta Jain, Olivier Le Meur
ETRA3
2018 How Old Do You Look? Inferring Your Age from Your Gaze
abstract
The visual exploration of a scene, represented by a visual scanpath, depends on a number of factors. Among them, the age of the observer plays a significant role. For instance, young kids are making shorter saccades and longer fixations than adults. In the light of these observations, we propose a new method for inferring the age of the observer from its scanpath. The proposed method is based on a 1D CNN network which is trained by real eye tracking data collected on five age groups. In order to boost the performance, the training dataset is augmented by predicting a high number of scan-paths thanks to the use of an age-dependent computational saccadic model. The proposed method brings a new momentum in this field not only by significantly outperforming existing method but also by being robust to noise and data erasure.
Tianyi Zhang 0013, Olivier Le Meur
ICIP2
2018 Image Selection in Photo Albums
abstract
The selection of the best photos in personal albums is a task that is often faced by photographers. This task can become laborious when the photo collection is large and it contains multiple similar photos. Recent advances on image aesthetics and photo importance evaluation has led to the creation of different metrics for automatically assessing a given image. However, these metrics are intended for the independent assessment of an image, without considering the possible context implicitly present within photo albums. In this work, we perform a user study for assessing how users select photos when provided with a complete photo album---a task that better reflects how users may review their personal photos and collections. Using the data provided by our study, we evaluate how existing state-of-the-art photo assessment methods perform relative to user selection, focusing in particular on deep learning based approaches. Finally, we explore a recent framework for adapting independent image scores to collections and evaluate in which scenarios such an adaptation can prove beneficial.
Dmitry Kuzovkin, Tania Pouli, Rémi Cozot, Olivier Le Meur, Jonathan Kervec, Kadi Bouatouch
ICMR4
2018 RGBD co-saliency detection via multiple kernel boosting and fusion
Lishan Wu, Zhi Liu 0003, Hangke Song, Olivier Le Meur
Multim. Tools Appl.4
2018 Multi-purpose bi-local CAT-based guidance filter
Hristina Hristova, Olivier Le Meur, Rémi Cozot, Kadi Bouatouch
Signal Process. Image Commun.2
2018 Transformation of the Multivariate Generalized Gaussian Distribution for Image Editing
abstract
Multivariate generalized Gaussian distributions (MGGDs) have aroused a great interest in the image processing community thanks to their ability to describe accurately various image features, such as image gradient fields. However, so far their applicability has been limited by the lack of a transformation between two of these parametric distributions. In this paper, we propose a novel transformation between MGGDs, consisting of an optimal transportation of the second-order statistics and a stochastic-based shape parameter transformation. We employ the proposed transformation between MGGDs for a color transfer and a gradient transfer between images. We also propose a new simultaneous transfer of color and gradient, which we apply for image color correction.
Hristina Hristova, Olivier Le Meur, Rémi Cozot, Kadi Bouatouch
IEEE Trans. Vis. Comput. Graph.2
2017 Perceptual metric for color transfer methods
abstract
In this paper, we propose a perceptual model for evaluating results from color transfer methods. We conduct a user study, which provides a set of subjective scores for triplets of input, target and result images. Then, for each triplet, we compute a number of image features, which objectively characterize a color transfer. To describe the relationship between these features and the subjective scores, we build a regression model with random forests. An analysis and a cross-validation show that the predictions of our model are highly accurate.
Hristina Hristova, Olivier Le Meur, Rémi Cozot, Kadi Bouatouch
ICIP2
2017 Age-dependent saccadic models for predicting eye movements
abstract
How people look at visual information reveals fundamental information about themselves, their interests and their state of mind. While previous visual attention models output static 2-dimensional saliency maps, saccadic models predict not only what observers look at but also how they move their eyes to explore the scene. Here we demonstrate that saccadic models are a flexible framework that can be tailored to emulate the gaze patterns from childhood to adulthood. The proposed age-dependent saccadic model not only outputs human-like, i.e. age-specific visual scanpath, but also significantly outperforms other state-of-the-art saliency models.
Olivier Le Meur, Antoine Coutrot, Adrien Le Roch, Andrea Helo, Pia Rämä, Zhi Liu 0003
ICIP1
2017 Saliency-based navigation in omnidirectional image
abstract
Omnidirectional images describe the color information at a given position from all directions. Affordable 360° cameras have recently been developed leading to an explosion of the 360° data shared on social networks. However, an omnidirectional image does not contain interesting content everywhere. Some part of the images are indeed more likely to be looked at by some users than others. Knowing these regions of interest might be useful for 360° image compression, streaming, retargeting or even editing. In this paper, we aim at modelling the user navigation within a 360° image, and detecting which parts of an omnidirectional content might draw users' attention. In particular, the paper proposes to aggregate and analyze 2D saliency detectors in different map projections, and also proposes a smooth navigation through the image to maximize saliency.
Thomas Maugey, Olivier Le Meur, Zhi Liu 0003
MMSP2
2017 Depth-Guided Disocclusion Inpainting of Synthesized RGB-D Images
abstract
We propose to tackle the problem of RGB-D image disocclusion inpainting when synthesizing new views of a scene by changing its viewpoint. Indeed, such a process creates holes both in depth and color images. First, we propose a novel algorithm to perform depth-map disocclusion inpainting. Our intuitive approach works particularly well for recovering the lost structures of the objects and to inpaint the depth-map in a geometrically plausible manner. Then, we propose a depth-guided patch-based inpainting method to fill-in the color image. Depth information coming from the reconstructed depth-map is added to each key step of the classical patch-based algorithm from Criminisi et al. in an intuitive manner. Relevant comparisons to the state-of-the-art inpainting methods for the disocclusion inpainting of both depth and color images are provided and illustrate the effectiveness of our proposed algorithms.
Pierre Buyssens, Olivier Le Meur, Maxime Daisy, David Tschumperlé, Olivier Lézoray
IEEE Trans. Image Process.2
2017 Visual Attention Saccadic Models Learn to Emulate Gaze Patterns From Childhood to Adulthood
abstract
How people look at visual information reveals fundamental information about themselves, their interests and their state of mind. While previous visual attention models output static 2D saliency maps, saccadic models aim to predict not only where observers look at but also how they move their eyes to explore the scene. In this paper, we demonstrate that saccadic models are a flexible framework that can be tailored to emulate observer's viewing tendencies. More specifically, we use fixation data from 101 observers split into five age groups (adults, 8-10 y.o., 6-8 y.o., 4-6 y.o., and 2 y.o.) to train our saccadic model for different stages of the development of human visual system. We show that the joint distribution of saccade amplitude and orientation is a visual signature specific to each age group, and can be used to generate age-dependent scan paths. Our age-dependent saccadic model does not only output human-like, age-specific visual scan paths, but also significantly outperforms other state-of-the-art saliency models. We demonstrate that the computational modeling of visual attention, through the use of saccadic model, can be efficiently adapted to emulate the gaze behavior of a specific group of observers.
Olivier Le Meur, Antoine Coutrot, Zhi Liu 0003, Pia Rämä, Adrien Le Roch, Andrea Helo
IEEE Trans. Image Process.1
2017 Depth-Aware Salient Object Detection and Segmentation via Multiscale Discriminative Saliency Fusion and Bootstrap Learning
abstract
This paper proposes a novel depth-aware salient object detection and segmentation framework via multiscale discriminative saliency fusion (MDSF) and bootstrap learning for RGBD images (RGB color images with corresponding Depth maps) and stereoscopic images. By exploiting low-level feature contrasts, mid-level feature weighted factors and high-level location priors, various saliency measures on four classes of features are calculated based on multiscale region segmentation. A random forest regressor is learned to perform the discriminative saliency fusion (DSF) and generate the DSF saliency map at each scale, and DSF saliency maps across multiple scales are combined to produce the MDSF saliency map. Furthermore, we propose an effective bootstrap learning-based salient object segmentation method, which is bootstrapped with samples based on the MDSF saliency map and learns multiple kernel support vector machines. Experimental results on two large datasets show how various categories of features contribute to the saliency detection performance and demonstrate that the proposed framework achieves the better performance on both saliency detection and salient object segmentation.
Hangke Song, Zhi Liu 0003, Huan Du, Guangling Sun, Olivier Le Meur, Tongwei Ren
IEEE Trans. Image Process.5
2017 Creating Segments and Effects on Comics by Clustering Gaze Data
abstract
Traditional comics are increasingly being augmented with digital effects, such as recoloring, stereoscopy, and animation. An open question in this endeavor is identifying where in a comic panel the effects should be placed. We propose a fast, semi-automatic technique to identify effects-worthy segments in a comic panel by utilizing gaze locations as a proxy for the importance of a region. We take advantage of the fact that comic artists influence viewer gaze towards narrative important regions. By capturing gaze locations from multiple viewers, we can identify important regions and direct a computer vision segmentation algorithm to extract these segments. The challenge is that these gaze data are noisy and difficult to process. Our key contribution is to leverage a theoretical breakthrough in the computer networks community towards robust and meaningful clustering of gaze locations into semantic regions, without needing the user to specify the number of clusters. We present a method based on the concept of relative eigen quality that takes a scanned comic image and a set of gaze points and produces an image segmentation. We demonstrate a variety of effects such as defocus, recoloring, stereoscopy, and animations. We also investigate the use of artificially generated gaze locations from saliency models in place of actual gaze locations.
Ishwarya Thirunarayanan, Khimya Khetarpal, Sanjeev J. Koppal, Olivier Le Meur, John M. Shea, Eakta Jain
ACM Trans. Multim. Comput. Commun. Appl.4
2017 High-dynamic-range image recovery from flash and non-flash image pairs
Hristina Hristova, Olivier Le Meur, Rémi Cozot, Kadi Bouatouch
Vis. Comput.2
2015 Spatiotemporal saliency detection based on superpixel-level trajectory
Zhi Liu 0003, Xiang Zhang 0006, Olivier Le Meur, Liquan Shen
Signal Process. Image Commun.4
2015 Special issue on recent advances in saliency models, applications and evaluations
Zhi Liu 0003, Olivier Le Meur, Ali Borji, Hongliang Li 0001
Signal Process. Image Commun.2
2015 Video Inpainting With Short-Term Windows: Application to Object Removal and Error Concealment
abstract
In this paper, we propose a new video inpainting method which applies to both static or free-moving camera videos. The method can be used for object removal, error concealment, and background reconstruction applications. To limit the computational time, a frame is inpainted by considering a small number of neighboring pictures which are grouped into a group of pictures (GoP). More specifically, to inpaint a frame, the method starts by aligning all the frames of the GoP. This is achieved by a region-based homography computation method which allows us to strengthen the spatial consistency of aligned frames. Then, from the stack of aligned frames, an energy function based on both spatial and temporal coherency terms is globally minimized. This energy function is efficient enough to provide high quality results even when the number of pictures in the GoP is rather small, e.g. 20 neighboring frames. This drastically reduces the algorithm complexity and makes the approach well suited for near real-time video editing applications as well as for loss concealment applications. Experiments with several challenging video sequences show that the proposed method provides visually pleasing results for object removal, error concealment, and background reconstruction context.
Mounira Ebdelli, Olivier Le Meur, Christine Guillemot
IEEE Trans. Image Process.2
2014 Saliency Aggregation: Does Unity Make Strength?
Olivier Le Meur, Zhi Liu 0003
ACCV (4)1
2014 HEVC Intra coding of ultra HD video with reduced complexity
abstract
The HEVC (High Efficiency Video Coding) standard brings the necessary quality versus rate performance for efficient transmission of Ultra High Definition formats (UHD). However, one of the remaining barriers to its adoption for UHD content is its high encoding complexity. In this paper, we address the problem of HEVC encoding complexity reduction by proposing a strategy to infer UHD coding modes and quadtree structure from those optimized for the lower (HD) resolution version of the input video. A speed-up factor of 3 is achieved compared to directly encoding the UHD format at the expense of a limited quality loss.
Nicolas Dhollande, Olivier Le Meur, Christine Guillemot
ICIP2
2014 Single image super-resolution using sparse representations with structure constraints
abstract
This paper describes a new single-image super-resolution algorithm based on sparse representations with image structure constraints. A structure tensor based regularization is introduced in the sparse approximation in order to improve the sharpness of edges. The new formulation allows reducing the ringing artefacts which can be observed around edges reconstructed by existing methods. The proposed method, named Sharper Edges based Adaptive Sparse Domain Selection (SE-ASDS), achieves much better results than many state-of-the-art algorithms, showing significant improvements in terms of PSNR (average of 29.63, previously 29.19), SSIM (average of 0.8559, previously 0.8471) and visual quality perception.
Júlio César Ferreira, Olivier Le Meur, Christine Guillemot, Eduardo A. B. da Silva, Gilberto Arantes Carrijo
ICIP2
2014 Co-saliency detection based on region-level fusion and pixel-level refinement
abstract
This paper addresses the problem of co-saliency detection, which aims to identify the common salient objects in a set of images and is important for many applications such as object co-segmentation and co-recognition. First, the segmentation driven low-rank matrix recovery model is used for intra saliency detection in each individual image of the image set, to highlight the regions whose features are sparse in each image. Then, a region-level fusion method, which exploits inter-region dissimilarities on color histograms and global consistency of regions over the image set, adjusts the intra saliency maps to obtain the region-level co-saliency maps, which can highlight co-salient object regions and suppress irrelevant regions. Finally, a pixel-level refinement method, which integrates color-spatial similarity between pixel and region with image border connectivity based object prior, generates the pixel-level co-saliency maps with better quality. Extensive experiments on two benchmark datasets demonstrate that the proposed co-saliency model consistently outperforms the state-of-the-art co-saliency models in both subjective and objective evaluation.
Zhi Liu 0003, Wenbin Zou, Xiang Zhang 0006, Olivier Le Meur
ICME5
2014 Superpixel-Based Spatiotemporal Saliency Detection
abstract
This paper proposes a superpixel-based spatiotemporal saliency model for saliency detection in videos. Based on the superpixel representation of video frames, motion histograms and color histograms are extracted at the superpixel level as local features and frame level as global features. Then, superpixel-level temporal saliency is measured by integrating motion distinctiveness of superpixels with a scheme of temporal saliency prediction and adjustment, and superpixel-level spatial saliency is measured by evaluating global contrast and spatial sparsity of superpixels. Finally, a pixel-level saliency derivation method is used to generate pixel-level temporal and spatial saliency maps, and an adaptive fusion method is exploited to integrate them into the spatiotemporal saliency map. Experimental results on two public datasets demonstrate that the proposed model outperforms six state-of-the-art spatiotemporal saliency models in terms of both saliency detection and human fixation prediction.
Zhi Liu 0003, Xiang Zhang 0006, Shuhua Luo, Olivier Le Meur
IEEE Trans. Circuits Syst. Video Technol.4
2014 Saliency Tree: A Novel Saliency Detection Framework
abstract
This paper proposes a novel saliency detection framework termed as saliency tree. For effective saliency measurement, the original image is first simplified using adaptive color quantization and region segmentation to partition the image into a set of primitive regions. Then, three measures, i.e., global contrast, spatial sparsity, and object prior are integrated with regional similarities to generate the initial regional saliency for each primitive region. Next, a saliency-directed region merging approach with dynamic scale control scheme is proposed to generate the saliency tree, in which each leaf node represents a primitive region and each non-leaf node represents a non-primitive region generated during the region merging process. Finally, by exploiting a regional center-surround scheme based node selection criterion, a systematic saliency tree analysis including salient node selection, regional saliency adjustment and selection is performed to obtain final regional saliency measures and to derive the high-quality pixel-wise saliency map. Extensive experimental results on five datasets with pixel-wise ground truths demonstrate that the proposed saliency tree model consistently outperforms the state-of-the-art saliency models.
Zhi Liu 0003, Wenbin Zou, Olivier Le Meur
IEEE Trans. Image Process.3
2013 Analysis of patch-based similarity metrics: Application to denoising
abstract
This paper presents a performance analysis of measures used for assessing similarities between patches. Compared to subjective ground thruth, our results indicate that some metrics are more suitable than others in a context of patch matching. This conclusion is confirmed by an experiment on non-local means (NLM) denoising algorithm. The denoising quality depends on the chosen similarity metric. In the best case, the gain is of 1.3dB compared to a classical SSD-based denoising algorithm.
Mounira Ebdelli, Olivier Le Meur, Christine Guillemot
ICASSP2
2013 Image inpainting using LLE-LDNR and linear subspace mappings
abstract
The paper first describes an examplar-based image inpainting algorithm using a locally linear neighbor embedding technique with low-dimensional neighborhood representation (LLE-LDNR). The inpainting algorithm first searches the K nearest neighbors ( ) of the input patch to be filled-in and linearly combine them with LLE-LDNR to synthesize the missing pixels. Linear regression is then introduced for improving the K-NN search. The performance of the LLE-LDNR with the enhanced K-NN search method is assessed for two applications: loss concealment and object removal.
Christine Guillemot, Mehmet Türkan, Olivier Le Meur, Mounira Ebdelli
ICASSP3
2013 Memorability of natural scenes: The role of attention
abstract
The image memorability consists in the faculty of an image to be recalled after a period of time. Recently, the memorability of an image database was measured and some factors responsible for this memorability were highlighted. In this paper, we investigate the role of visual attention in image memorability around two axis. The first one is experimental and uses results of eye-tracking performed on a set of images of different memorability scores. The second investigation axis is predictive and we show that attention-related features can advantageously replace low-level features in image memorability prediction. From our work it appears that the role of visual attention is important and should be more taken into account along with other low-level features.
Matei Mancas, Olivier Le Meur
ICIP2
2013 3D view synthesis with inter-view consistency
abstract
In this paper, we propose a new pipeline to synthesize virtual views by extrapolation. It allows us to generate virtual views far away from each other, each presenting the exact same level of quality. This inter-view consistency is key to seamlessly navigate between viewpoints. Its computational cost is also lower than that of existing approaches. We compare the proposed approach with state-of-the-art methods and show the effectiveness of this new view synthesis pipeline.
David Wolinski, Olivier Le Meur, Josselin Gautier
ACM Multimedia2
2013 Object removal and loss concealment using neighbor embedding methods
Christine Guillemot, Mehmet Türkan, Olivier Le Meur, Mounira Ebdelli
Signal Process. Image Commun.3
2013 Hierarchical Super-Resolution-Based Inpainting
abstract
This paper introduces a novel framework for examplar-based inpainting. It consists in performing first the inpainting on a coarse version of the input image. A hierarchical super-resolution algorithm is then used to recover details on the missing areas. The advantage of this approach is that it is easier to inpaint low-resolution pictures than high-resolution ones. The gain is both in terms of computational complexity and visual quality. However, to be less sensitive to the parameter setting of the inpainting method, the low-resolution input picture is inpainted several times with different configurations. Results are efficiently combined with a loopy belief propagation and details are recovered by a single-image super-resolution algorithm. Experimental results in a context of image editing and texture synthesis demonstrate the effectiveness of the proposed method. Results are compared to five state-of-the-art inpainting methods.
Olivier Le Meur, Mounira Ebdelli, Christine Guillemot
IEEE Trans. Image Process.1
2012 Super-Resolution-Based Inpainting
Olivier Le Meur, Christine Guillemot
ECCV (6)1
2012 Examplar-based video inpainting with motion-compensated neighbor embedding
abstract
This paper describes a video inpainting algorithm based on motion-compensated neighbor embedding. The unknown pixels are estimated as a linear combination of the K closest patches using motion-compensated neighbor embedding. The algorithm is first assessed by assuming the motion information of the masked pixels to be known. This assumption is not realistic in video editing (object removal) applications. It however helps isolating the various problems for the sake of analysis. Different approaches are then assessed in the context where the motion information of missing pixels is unknown. Experiments on several videos show the benefits of the proposed approach which lead to natural looking videos with less annoying artefacts than when using a template matching technique.
Mounira Ebdelli, Christine Guillemot, Olivier Le Meur
ICIP3
2012 Efficient depth map compression based on lossless edge coding and diffusion
abstract
The multi-view plus depth video (MVD) format has recently been introduced for 3DTV and free-viewpoint video (FVV) scene rendering. Given one view (or several views) with its depth information, depth image-based rendering techniques have the ability to generate intermediate views. The MVD format however generates large volumes of data which need to be compressed for storage and transmission. This paper describes a new depth map encoding algorithm which aims at exploiting the intrinsic depth maps properties. Depth images indeed represent the scene surface and are characterized by areas of smoothly varying grey levels separated by sharp edges at the position of object boundaries. Preserving these characteristics is important to enable high quality view rendering at the receiver side. The proposed algorithm proceeds in three steps: the edges at object boundaries are first detected using a Sobel operator. The positions of the edges are encoded using the JBIG algorithm. The luminance values of the pixels along the edges are then encoded using an optimized path encoder. The decoder runs a fast diffusion-based inpainting algorithm which fills in the unknown pixels within the objects by starting from their boundaries. The performance of the algorithm is assessed against JPEG-2000 and HEVC, both in terms of PSNR of the depth maps versus rate as well as in terms of PSNR of the synthesized virtual views.
Josselin Gautier, Olivier Le Meur, Christine Guillemot
PCS2
2011 Robustness and repeatability of saliency models subjected to visual degradations
abstract
The present study investigates the sensitivity of computationnal models of visual attention when subjected to visual degradations. One hundred and twenty natural color pictures were degraded using 6 filtering operations. By using different settings, five state-of-the-art models are used to compute 11400 saliency maps. The comparison of these maps to human saliency maps indicates that the tested models are robust to most of the visual degradations they were subjected to. These findings have implications on saliency-based applications, such as quality assessment and coding. A last point concerns the high repeatability of saliency models that might be used in a context of image retrieval.
Olivier Le Meur
ICIP1
2011 Examplar-based inpainting based on local geometry
abstract
In this paper, we propose a novel inpainting algorithm combining the advantages of PDE-based schemes and examplar-based approaches. The proposed algorithm relies on the use of structure tensors to define the filling order priority and template matching. The structure tensors are computed in a hierarchic manner whereas the template matching is based on a K-nearest neighbor algorithm. The value K is adaptively set in function of the local texture information. Compared to two state of the art approaches, the proposed method provides more coherent results.
Olivier Le Meur, Josselin Gautier, Christine Guillemot
ICIP1
2011 Predicting saliency using two contextual priors: The dominant depth and the horizon line
abstract
A computational model of visual attention using visual inferences is proposed. The dominant depth and the horizon line position are inferred from low-level visual features. This prior knowledge helps to find salient areas on still color pictures. Regarding the dominant depth, the idea is to favor the lowest spatial frequencies on close-up scenes whereas the highest spatial frequencies are used to predict salient areas on panoramic view. Some studies showed that the horizon line is a natural attractor of our gaze. Horizon detection is then used to improve the saliency prediction. Results show that the proposed model outperforms existing approaches. However, the dominant depth does not bring any gain in the saliency prediction.
Olivier Le Meur
ICME1
2011 Prediction of the inter-observer visual congruency (IOVC) and application to image ranking
abstract
This paper proposes an automatic method for predicting the inter-observer visual congruency (IOVC). The IOVC reflects the congruence or the variability among different subjects looking at the same image. Predicting this congruence is of interest for image processing applications where the visual perception of a picture matters such as website design, advertisement, etc. This paper makes several new contributions. First, a computational model of the IOVC is proposed. This new model is a mixture of low-level visual features extracted from the input picture where model's parameters are learned by using a large eye-tracking database. Once the parameters have been learned, it can be used for any new picture. Second, regarding low-level visual feature extraction, we propose a new scheme to compute the depth of field of a picture. Finally, once the training and the feature extraction have been carried out, a score ranging from 0 (minimal congruency) to 1 (maximal congruency) is computed. A value of 1 indicates that observers would focus on the same locations and suggests that the picture presents strong locations of interest. A second database of eye movements is used to assess the performance of the proposed model. Results show that our IOVC criterion outperforms the Feature Congestion measure \cite{Rosenholtz2007}. To illustrate the interest of the proposed model, we have used it to automatically rank personalized photograph.
Olivier Le Meur, Thierry Baccino, Aline Roumy
ACM Multimedia1
2010 Spatio-temporal combination of saliency maps and eye-tracking assessment of different strategies
abstract
The modeling of the human visual attention into a computational attention model leads to the split of visual features into several independent channels. Then, a difficult problem arises to combine these maps, having different dynamic ranges or distribution. When several maps are considered, such process is mandatory in order to compute a single measure of interest for each location, regardless of which features contributed to the salience. Several strategies of cue combination are proposed in this paper for the spatial cues as well as the temporal saliency. Finally, some user tests on still image and video databases leads to highlight one operator.
Christel Chamaret, Jean-Claude Chevet, Olivier Le Meur
ICIP3
2010 Overt visual attention for free-viewing and quality assessment tasks: Impact of the regions of interest on a video quality metric
Olivier Le Meur, Alexandre Ninassi, Patrick Le Callet, Dominique Barba
Signal Process. Image Commun.1
2010 Do video coding impairments disturb the visual attention deployment?
Olivier Le Meur, Alexandre Ninassi, Patrick Le Callet, Dominique Barba
Signal Process. Image Commun.1
2010 Relevance of a Feed-Forward Model of Visual Attention for Goal-Oriented and Free-Viewing Tasks
abstract
A purely bottom-up model of visual attention is proposed and compared to five state-of-the-art models. The role of the low-level visual features is examined in two contexts. Two datasets are used: one containing data coming from an eye tracking experiment obtained in a free-viewing task and a second containing 5000 hand-label pictures (observers had to enclose the most visually interesting objects in a rectangle). The relevancy of the bottom-up models, i.e. the ability of a model to predict where the salient areas are located, is evaluated. Whatever the metrics and the datasets, the degree of similarity between predictions and ground truth is significantly above chance. The proposed model, resting on a small number of features, is shown to be a good predictor of the human visual fixations but also a good predictor of the objects chosen as interesting by observers. This study suggests that the low-level of visual features have a significant role in a free-viewing task but also in a high-level visual task, such as the choice of the object of interest in a complex visual scene. Another outcome concerns the viewing duration used in eye tracking experiments. Results suggest that this parameter is finally not as critical as one would expect.
Olivier Le Meur, Jean-Claude Chevet
IEEE Trans. Image Process.1
2009 What we see is most likely to be what matters: Visual attention and applications
abstract
The computational modeling of the visual attention is receiving increasing attention from the computer vision community. Several bottom-up models have been proposed. In spite of their complexities, these models are still a basic description of our visual system. Review of resulting approaches of these efforts are presented in the first part of this paper. Limitations of these approaches are introduced and several research trends are given. Among them, the most important one might be the use of prior knowledge, conjointly with the low-level visual features. Concomitantly with visual attention (VA) modeling progress, the image and video processing community is increasingly considering VA models in different fields or services. Current and future applications of VA models are discussed in the second part.
Olivier Le Meur, Patrick Le Callet
ICIP1
2008 Which semi-local visual masking model forwavelet based image quality metric?
abstract
Properties and models of the human visual system (HVS) are the fundaments for most of sufficient objective image or video quality metrics. Among HVS properties, visual masking is a sensitive issue. Many models exist in literature. Simplest models can only predict visibility threshold for very simple cue while for natural images one should consider more complex approaches such as semi-local masking. Our previous work has shown the positive impact of incorporating semi-local masking in image quality metric according to one subjective study. It is important to consolidate this work with different subjective experiments. In this paper, different visual masking models, including contrast masking and semi-local masking, are evaluated according to three subjective studies. These subjective experiments were conducted with different protocols, different types of display devices, different contents and different populations.
Alexandre Ninassi, Olivier Le Meur, Patrick Le Callet, Dominique Barba
ICIP2
2008 Attention-based video reframing: Validation using eye-tracking
abstract
Watching TV shows on cell phones is starting to become a reality. Nevertheless, there still exist some significant issues due to the small size of cell phone screens. The direct transfer of contents that are not specifically shot for the mobile device will provide indistinguishable objects. An automated way, delivering the best viewing experience is proposed in this paper. This solution significantly improves the visual comfort, by zooming in on the regions of interest. The relevance of this solution rests on its capability to preserve the visually important areas as well as the temporal stability. Eye-tracking experiments are one metric to assess the reframing quality. Involving 16 observers, they show that more than 90% of the visually important regions are kept in the reframed clip.
Christel Chamaret, Olivier Le Meur
ICPR2
2007 Does where you Gaze on an Image Affect your Perception of Quality? Applying Visual Attention to Image Quality Metric
abstract
The aim of an objective image quality assessment is to find an automatic algorithm that evaluates the quality of pictures or video as a human observer would do. To reach this goal, researchers try to simulate the Human Visual System (HVS). Visual attention is a main feature of the HVS, but few studies have been done on using it in image quality assessment. In this work, we investigate the use of the visual attention information in their final pooling step. The rationale of this choice is that an artefact is likely more annoying in a salient region than in other areas. To shed light on this point, a quality assessment campaign has been conducted during which eye movements have been recorded. The results show that applying the visual attention to image quality assessment is not trivial, even with the ground truth.
Alexandre Ninassi, Olivier Le Meur, Patrick Le Callet, Dominique Barba
ICIP (2)2
2006 Efficient Saliency-Based Repurposing Method
abstract
Images play a very relevant role in our daily life. People now can easily shoot and share pictures thanks to the exponential growth of the portable medias, such as digital cameras, mobile phone. As the display size of those devices is relatively small, browsing large pictures remains difficult. Content re-purposing is an elegant solution to deal with this problem. It consists in cropping the images in order to display only the most interesting parts of the picture. A new algorithm is proposed in this paper; the experiments described herein, leading to a qualitative and a quantitative assessment, show that the proposed solution outperforms the conventional method.
Olivier Le Meur, Xavier Castellani, Patrick Le Callet, Dominique Barba
ICIP1
2006 A Coherent Computational Approach to Model Bottom-Up Visual Attention
abstract
Visual attention is a mechanism which filters out redundant visual information and detects the most relevant parts of our visual field. Automatic determination of the most visually relevant areas would be useful in many applications such as image and video coding, watermarking, video browsing, and quality assessment. Many research groups are currently investigating computational modeling of the visual attention system. The first published computational models have been based on some basic and well-understood Human Visual System (HVS) properties. These models feature a single perceptual layer that simulates only one aspect of the visual system. More recent models integrate complex features of the HVS and simulate hierarchical perceptual representation of the visual input. The bottom-up mechanism is the most occurring feature found in modern models. This mechanism refers to involuntary attention (i.e., salient spatial visual features that effortlessly or involuntary attract our attention). This paper presents a coherent computational approach to the modeling of the bottom-up visual attention. This model is mainly based on the current understanding of the HVS behavior. Contrast sensitivity functions, perceptual decomposition, visual masking, and center-surround interactions are some of the features implemented in this model. The performances of this algorithm are assessed by using natural images and experimental measurements from an eye-tracking system. Two adequate well-known metrics (correlation coefficient and Kullbacl-Leibler divergence) are used to validate this model. A further metric is also defined. The results from this model are finally compared to those from a reference bottom-up model.
Olivier Le Meur, Patrick Le Callet, Dominique Barba, Dominique Thoreau
IEEE Trans. Pattern Anal. Mach. Intell.1
2005 A spatio-temporal model of the selective human visual attention
abstract
A new spatio-temporal model for simulating the bottom-up visual attention is proposed. It has been built from numerous important properties of the human visual system (HVS). This paper focuses both on the architecture of the model and on its performances. Given that the spatial model of the bottom-up visual attention has already been defined [O. Le Meur et al., 2004], the temporal dimension is more accurately described. A qualitative and quantitative comparison with human fixations collected from an eye tracking apparatus is undertaken. From the former, the quality of the prediction is deemed very good whereas the latter illustrates that the best predictor of the human fixation consists of the sum all visual features (achromatic, chromatic and motion).
Olivier Le Meur, Dominique Thoreau, Patrick Le Callet, Dominique Barba
ICIP (3)1
2004 Performance assessment of a visual attention system entirely based on a human vision modeling
abstract
It is now commonly assumed that the human visual attention, which is a selecting process of the most relevant locations in a scene according to a particular behavior, is driven by both top-down (task-dependent) and bottom-up (signal-dependent) control. A new model attempting to simulate the bottom-up process has been designed Le Meur, O et al., (2004). This model is purely based on visual system properties that provides noticeable advantages compared to the classical published approaches. This paper focuses on the performance assessment of this model by achieving a comparison with real fixation points stemming from eye-tracking apparatus both subjectively and objectively.
Olivier Le Meur, Patrick Le Callet, Dominique Barba, Dominique Thoreau
ICIP1