EDBT 2026 Demo / reviewers in the wild / expert
Olivier Le Meur
dblp:35/3436
· DBLP profile ↗
71ranked-venue papers
20as first author
13since 2021 · last 2025
0000-0001-9883-0296ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 65 · 19 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 3 since 2021Computer networks · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Style-FG: A Style-based Framework for Film Grain Analysis and SynthesisabstractFilm grain which used to be a by-product of the chemical processing in the analog film stock is a desirable feature in the era of digital cameras. Besides participating to the artistic intent during content creation, film grain has also interesting properties in the video compression chain such as its ability to mask compression artifacts. In this article, we use a deep learning-based framework for film grain analysis, generation, and synthesis. Our framework Style-FG consists of three modules: a style encoder performing film grain style analysis, a mapping network responsible for film grain style generation, and a synthesis network that generates and blends a specific grain style to a given content in a content-adaptive manner. All modules are trained jointly, thanks to dedicated loss functions, on a new large and diverse dataset of pairs of grain-free and grainy images that we made publicly available to the community. 1 Quantitative and qualitative evaluations show that fidelity to the reference grain, diversity of grain styles as well as a perceptually pleasant grain synthesis are achieved, demonstrating that each module outperforms the state-of-the-art in the task it was designed for. To contribute further to the sustainability necessary effort of the digital information and communication field, a light-weight version of Style-FG is also proposed, which demonstrates similar quantitative and qualitative performances, while reducing the number of network parameters by a factor of 92%. Zoubida Ameur, Claire-Hélène Demarty, Olivier Le Meur, Daniel Ménard |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | 3R-INN: How to Be Climate Friendly While Consuming/Delivering Videos?
Zoubida Ameur, Claire-Hélène Demarty, Daniel Ménard, Olivier Le Meur |
ECCV (75) | 4 |
| 2023 | Display Power Modeling for Energy Consumption ControlabstractAs the most consuming devices in the video chain, it is necessary to master the power consumption of displays. However it necessitates to have a precise modeling of their power consumption. This paper proposes two approaches that outperform the state-of-the-art, to model the power consumption of emissive displays: first through the estimation of the display parameters from a theoretical RGBW model, then through two deep-based models agnostic of the display technology, that output either a global power value or a power map corresponding to an input image. All three modelings were made possible thanks to the release of DISPLAYPOWER2k, a large dataset of power consumption values for OLED displays. Claire-Hélène Demarty, Laurent Blondé, Olivier Le Meur |
ICIP | 3 |
| 2023 | Deep-Learning-Based Energy Aware ImagesabstractIn this paper, we present a method to compute energy-aware images, that aims to reduce the energy consumption of displays. This method relies on a lightweight unsupervised deep model which finds out the best trade-off between visual quality and energy reduction. From an input image and an energy reduction rate, a dimming map is inferred. We show that the proposed model performs as good as state-of-the-art methods, while being much more simple. In addition, the dimming map computation is constrained in order to ease its distribution throughout the video chain. Olivier Le Meur, Claire-Hélène Demarty, Laurent Blondé |
ICIP | 1 |
| 2023 | Energy-Aware HDR Content End-to-End VVC EncodingabstractThe proposed demonstration combines several commercially viable state-of-the-art technologies to achieve significant energy reduction while delivering HDR video without compromising QoE. The components of the demo are: (i) The demo is based our solution on VVC, the most efficient, highest quality video codec available today. (ii) We leverage Advanced HDR by Technicolor, a solution that uniquely can deliver HDR and SDR video using a single stream. (iii) We introduce a solution for dynamically adjusting peak brightness of displays via display adaptation controlled through the network using dynamic metadata encapsulated in MPEG SEI messages. This solution has been tested and shows a capability of reducing power by 10% or more without any subjective quality loss, and far more with minimal subjective impact. (iv) While we demonstrate a streaming (unicast) scenario in our live demo, our system is also capable of working in a dynamic unicast/multicast environment based on DVB transmission standards. (v) The entire demonstration is resulting from an industrial collaboration and is based on ready to market technologies and prototypes. Olivier Le Meur, Franck Aumont, Juan-Carlos Vargas, Pierre-Loup Cabarat, Daniel Ménard, Oussama Hammami, Thomas Guionnet |
MMSP | 1 |
| 2023 | Style-based film grain analysis and synthesisabstractFilm grain which used to be a by-product of the chemical processing in the analog film stock, is a desirable feature in the era of digital cameras. Besides participating to the artistic intent during content creation, film grain has also interesting properties in the video compression chain such as its ability to mask compression artifacts. In this paper, we use a deep learning-based framework for film grain analysis, generation and synthesis. Our framework consists of three modules: a style encoder performing film grain style analysis, a mapping network responsible for film grain style generation, and a synthesis network that generates and blends a specific grain style to a given content in a content-adaptive manner. All modules are trained jointly, thanks to dedicated loss functions, on a new large and diverse dataset of pairs of grain-free and grainy images that we made publicly available to the community1. Quantitative and qualitative evaluations show that fidelity to the reference grain, diversity of grain styles as well as a perceptually pleasant grain synthesis are achieved, demonstrating that each module outperforms the state-of-the-art in the task it was designed for. Zoubida Ameur, Claire-Hélène Demarty, Olivier Le Meur, Daniel Ménard, Edouard François |
MMSys | 3 |
| 2023 | HDR-LFNet: Inverse tone mapping using fusion network
Mathieu Chambe, Ewa Kijak, Zoltán Miklós 0001, Olivier Le Meur, Rémi Cozot, Kadi Bouatouch |
Comput. Graph. | 4 |
| 2023 | Invertible Energy-Aware ImagesabstractDisplaying video content requires significant amounts of energy. In face of the energy crisis and climate emergency, we propose a new approach to produce energy-aware images. The purpose of such images is to consume less energy when displayed onscreen while maximizing the quality of experience. A new invertible neural network called Invertible Energy-Aware Network (InvEAN) is proposed to produce invertible energy-aware images, allowing to reduce the energy consumption of display devices and offering the possibility to recover the original image if required. Experimental results show that the InvEAN network outperforms two existing methods over three datasets. Olivier Le Meur, Claire-Hélène Demarty |
IEEE Signal Process. Lett. | 1 |
| 2022 | Deep learning for assessing the aesthetics of professional photographsabstractAbstract Aesthetic quality assessment for photographs is an important research topic since it can be used for a number of applications, such as image database management or image browsing. In 2012, the aesthetic visual analysis (AVA) dataset has been proposed. It has since then been used to train the majority of computational models of aesthetics assessment. We observe that AVA is mainly composed of competitive photographs which notion of aesthetics differs from other kinds of photographs, such as professional photographs. In this paper, we evaluate whether or not recent aesthetics assessment models generalize well and perform well over professional photographs. We noticed that the different models we tested behave differently on both categories, and therefore do not generalize well. Besides, we fine‐tuned one of the tested model using professional photographs and the results show that this fine‐tuning is effectively improving the coverage of the methods. Mathieu Chambe, Rémi Cozot, Olivier Le Meur |
Comput. Animat. Virtual Worlds | 3 |
| 2022 | Trajectory Saliency Detection Using Consistency-Oriented Latent Codes From a Recurrent Auto-EncoderabstractIn this paper, we are concerned with the detection of progressive dynamic saliency from video sequences. More precisely, we are interested in saliency related to motion and likely to appear progressively over time. It can be relevant to trigger alarms, to dedicate additional processing or to detect specific events. Trajectories represent the best way to support progressive dynamic saliency detection. Accordingly, we will talk about trajectory saliency. A trajectory will be qualified as salient if it deviates from normal trajectories that share a common motion pattern related to a given context. First, we need a compact while discriminative representation of trajectories. We adopt a (nearly) unsupervised learning-based approach. The latent code estimated by a recurrent auto-encoder provides the desired representation. In addition, we enforce consistency for normal (similar) trajectories through the auto-encoder loss function. The distance of the trajectory code to a prototype code accounting for normality is the means to detect salient trajectories. We validate our trajectory saliency detection method on synthetic and real trajectory datasets, and highlight the contributions of its different components. We compare our method favourably to existing methods on several saliency configurations constructed from the publicly available large dataset of pedestrian trajectories acquired in a railway station. Léo Maczyta, Patrick Bouthemy, Olivier Le Meur |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Example-Based Colour Transfer for 3D Point CloudsabstractAbstract Example‐based colour transfer between images, which has raised a lot of interest in the past decades, consists of transferring the colour of an image to another one. Many methods based on colour distributions have been proposed, and more recently, the efficiency of neural networks has been demonstrated again for colour transfer problems. In this paper, we propose a new pipeline with methods adapted from the image domain to automatically transfer the colour from a target point cloud to an input point cloud. These colour transfer methods are based on colour distributions and account for the geometry of the point clouds to produce a coherent result. The proposed methods rely on simple statistical analysis, are effective, and succeed in transferring the colour style from one point cloud to another. The qualitative results of the colour transfers are evaluated and compared with existing methods. Ific Goudé, Rémi Cozot, Olivier Le Meur, Kadi Bouatouch |
Comput. Graph. Forum | 3 |
| 2021 | Deep saliency models : The quest for the loss function
Alexandre Bruckert, Hamed Rezazadegan Tavakoli, Zhi Liu 0003, Marc Christie, Olivier Le Meur |
Neurocomputing | 5 |
| 2021 | Predicting atypical visual saliency for autism spectrum disorder via scale-adaptive inception module and discriminative region enhancement loss
Weijie Wei 0001, Zhi Liu 0003, Lijin Huang, Alexis Nebout, Olivier Le Meur, Tianhong Zhang, Jijun Wang 0003, Lihua Xu |
Neurocomputing | 5 |
| 2020 | Eye-Gaze Activity in Crowds: Impact of Virtual Reality and DensityabstractWhen we are walking in crowds, we mainly use visual information to avoid collisions with other pedestrians. Thus, gaze activity should be considered to better understand interactions between people in a crowd. In this work, we use Virtual Reality (VR) to facilitate motion and gaze tracking, as well as to accurately control experimental conditions, in order to study the effect of crowd density on eye-gaze behavior. Our motivation is to better understand how interaction neighborhood (i.e., the subset of people actually influencing one’s locomotion trajectory) changes with density. To this end, we designed two experiments. The first one evaluates the biases introduced by the use of VR on the visual activity when walking among people, by comparing eye-gaze activity while walking in a real and virtual street. We then designed a second experiment where participants walked in a virtual street with different levels of pedestrian density. We demonstrate that gaze fixations are performed at the same frequency despite increases in pedestrian density, while the eyes scan a narrower portion of the street. These results suggest that in such situations walkers focus more on people in front and closer to them. These results provide valuable insights regarding eye-gaze activity during interactions between people in a crowd, and suggest new recommendations in designing more realistic crowd simulations. Florian Berton, Ludovic Hoyet, Anne-Hélène Olivier, Julien Bruneau 0002, Olivier Le Meur, Julien Pettré |
VR | 5 |
| 2020 | Effective schizophrenia recognition using discriminative eye movement features and model-metric based features
Lijin Huang, Weijie Wei 0001, Zhi Liu 0003, Tianhong Zhang, Jijun Wang 0003, Lihua Xu, Olivier Le Meur |
Pattern Recognit. Lett. | 8 |
| 2020 | FANet: Features Adaptation Network for 360$^{\circ }$ Omnidirectional Salient Object DetectionabstractSalient object detection (SOD) in 360° omnidirectional images has become an eye-catching problem because of the popularity of affordable 360° cameras. In this paper, we propose a Features Adaptation Network (FANet) to highlight salient objects in 360° omnidirectional images reliably. To utilize the feature extraction capability of convolutional neural networks and capture global object information, we input the equirectangular 360° images and corresponding cube-map 360° images to the feature extraction network (FENet) simultaneously to obtain multi-level equirectangular and cube-map features. Furthermore, we fuse these two kinds of features at each level of the FENetby a projection features adaptation (PFA) module, for selecting these two kinds of features adaptively. Finally, we combine the preliminary adaptation features at different levels by a multi-level features adaptation (MLFA) module, which weights these different-level features adaptively and produces the final saliency maps. Experiments show our FANet outperforms the state-of-the-art methods on the 360° omnidirectional SOD datasets. Mengke Huang, Zhi Liu 0003, Gongyang Li, Xiaofei Zhou 0003, Olivier Le Meur |
IEEE Signal Process. Lett. | 5 |
| 2019 | How Well Current Saliency Prediction Models Perform on UAVs Videos?
Anne-Flore Perrin, Lu Zhang 0037, Olivier Le Meur |
CAIP (1) | 3 |
| 2019 | Deep Learning For Inter-Observer Congruency PredictionabstractAccording to the literature regarding visual saliency, observers may exhibit considerable variations in their gaze behaviors. These variations are influenced by aspects such as cultural background, age or prior experiences, but also by features in the observed images. The dispersion between the gaze of different observers looking at the same image is commonly referred as inter-observer congruency (IOC). Predicting this congruence can be of great interest when it comes to study the visual perception of an image. In this paper, we introduce a new method based on deep learning techniques to predict the IOC of an image. This is achieved by first extracting features from an image through a deep convolutional network. We then show that using such features to train a model with a shallow network regression technique significantly improves the precision of the prediction over existing approaches. Alexandre Bruckert, Yat Hong Lam, Marc Christie, Olivier Le Meur |
ICIP | 4 |
| 2019 | Unsupervised Motion Saliency Map Estimation Based On Optical Flow InpaintingabstractThe paper addresses the problem of motion saliency in videos, that is, identifying regions that undergo motion departing from its context. We propose a new unsupervised paradigm to compute motion saliency maps. The key ingredient is the flow inpainting stage. Candidate regions are determined from the optical flow boundaries. The residual flow in these regions is given by the difference between the optical flow and the flow inpainted from the surrounding areas. It provides the cue for motion saliency. The method is flexible and general by relying on motion information only. Experimental results on the DAVIS 2016 benchmark demonstrate that the method compares favourably with state-of-the-art video saliency methods. Léo Maczyta, Patrick Bouthemy, Olivier Le Meur |
ICIP | 3 |
| 2019 | Quality Metric Aggregation for HDR/WCG ImagesabstractHigh Dynamic Range (HDR) and Wide Color Gamut (WCG) screens are able to display images with brighter and darker pixels with more vivid colors than ever. Automatically assessing the quality of these HDR/WCG images is of critical importance to evaluate the performances of image compression schemes. In recent years, full-reference metrics, such as HDR-VDP-2, PU-encoding metrics, have been designed for this purpose. However, none of these metrics consider chromatic artifacts. In this paper, we propose our own full-reference quality metric adapted to HDR and WCG content that is sensitive to chromatic distortions. The proposed metric is based on two existing HDR quality metrics and color image features. A support vector machine regression is used to combine the aforementioned features. Experimental results demonstrate the effectiveness of the proposed metric in the context of image compression. Maxime Rousselot, Xavier Ducloux, Olivier Le Meur, Rémi Cozot |
ICIP | 3 |
| 2019 | Saliency-aware inter-image color transfer for image manipulation
Zhi Liu 0003, Qihan Jiao, Olivier Le Meur, Wanlei Zhao |
Multim. Tools Appl. | 4 |
| 2019 | CNN-based temporal detection of motion saliency in videos
Léo Maczyta, Patrick Bouthemy, Olivier Le Meur |
Pattern Recognit. Lett. | 3 |
| 2019 | Context in Photo Albums: Understanding and Modeling User Behavior in Clustering and SelectionabstractRecent progress in digital photography and storage availability has drastically changed our approach to photo creation. While in the era of film cameras, careful forethought would usually precede the capture of a photo; nowadays, a large number of pictures can be taken with little effort. One of the consequences is the creation of numerous photos depicting the same moment in slightly different ways, which makes the process of organizing photos laborious for the photographer. Nevertheless, photo collection organization is important both for exploring photo albums and for simplifying the ultimate task of selecting the best photos. In this work, we conduct a user study to explore how users tend to organize or cluster similar photos in albums, to what extent different users agree in their clustering decisions, and to investigate how the clustering-defined photo context affects the subsequent photo-selection process. We also propose an automatic hierarchical clustering solution for modeling user clustering decisions. To demonstrate the usefulness of our approach, we apply it to the task of automatic photo evaluation within photo albums and propose a clustering-based context adaptation. Dmitry Kuzovkin, Tania Pouli, Olivier Le Meur, Rémi Cozot, Jonathan Kervec, Kadi Bouatouch |
ACM Trans. Appl. Percept. | 3 |
| 2018 | Deepcomics: saliency estimation for comicsabstractA key requirement for training deep learning saliency models is large training eye tracking datasets. Despite the fact that the accessibility of eye tracking technology has greatly increased, collecting eye tracking data on a large scale for very specific content types is cumbersome, such as comic images, which are different from natural images such as photographs because text and pictorial content is integrated. In this paper, we show that a deep network trained on visual categories where the gaze deployment is similar to comics outperforms existing models and models trained with visual categories for which the gaze deployment is dramatically different from comics. Further, we find that it is better to use a computationally generated dataset on visual category close to comics one than real eye tracking data of a visual category that has different gaze deployment. These findings hold implications for the transference of deep networks to different domains. Kévin Bannier, Eakta Jain, Olivier Le Meur |
ETRA | 3 |
| 2018 | How Old Do You Look? Inferring Your Age from Your GazeabstractThe visual exploration of a scene, represented by a visual scanpath, depends on a number of factors. Among them, the age of the observer plays a significant role. For instance, young kids are making shorter saccades and longer fixations than adults. In the light of these observations, we propose a new method for inferring the age of the observer from its scanpath. The proposed method is based on a 1D CNN network which is trained by real eye tracking data collected on five age groups. In order to boost the performance, the training dataset is augmented by predicting a high number of scan-paths thanks to the use of an age-dependent computational saccadic model. The proposed method brings a new momentum in this field not only by significantly outperforming existing method but also by being robust to noise and data erasure. Tianyi Zhang 0013, Olivier Le Meur |
ICIP | 2 |
| 2018 | Image Selection in Photo AlbumsabstractThe selection of the best photos in personal albums is a task that is often faced by photographers. This task can become laborious when the photo collection is large and it contains multiple similar photos. Recent advances on image aesthetics and photo importance evaluation has led to the creation of different metrics for automatically assessing a given image. However, these metrics are intended for the independent assessment of an image, without considering the possible context implicitly present within photo albums. In this work, we perform a user study for assessing how users select photos when provided with a complete photo album---a task that better reflects how users may review their personal photos and collections. Using the data provided by our study, we evaluate how existing state-of-the-art photo assessment methods perform relative to user selection, focusing in particular on deep learning based approaches. Finally, we explore a recent framework for adapting independent image scores to collections and evaluate in which scenarios such an adaptation can prove beneficial. Dmitry Kuzovkin, Tania Pouli, Rémi Cozot, Olivier Le Meur, Jonathan Kervec, Kadi Bouatouch |
ICMR | 4 |
| 2018 | RGBD co-saliency detection via multiple kernel boosting and fusion
Lishan Wu, Zhi Liu 0003, Hangke Song, Olivier Le Meur |
Multim. Tools Appl. | 4 |
| 2018 | Multi-purpose bi-local CAT-based guidance filter
Hristina Hristova, Olivier Le Meur, Rémi Cozot, Kadi Bouatouch |
Signal Process. Image Commun. | 2 |
| 2018 | Transformation of the Multivariate Generalized Gaussian Distribution for Image EditingabstractMultivariate generalized Gaussian distributions (MGGDs) have aroused a great interest in the image processing community thanks to their ability to describe accurately various image features, such as image gradient fields. However, so far their applicability has been limited by the lack of a transformation between two of these parametric distributions. In this paper, we propose a novel transformation between MGGDs, consisting of an optimal transportation of the second-order statistics and a stochastic-based shape parameter transformation. We employ the proposed transformation between MGGDs for a color transfer and a gradient transfer between images. We also propose a new simultaneous transfer of color and gradient, which we apply for image color correction. Hristina Hristova, Olivier Le Meur, Rémi Cozot, Kadi Bouatouch |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2017 | Perceptual metric for color transfer methodsabstractIn this paper, we propose a perceptual model for evaluating results from color transfer methods. We conduct a user study, which provides a set of subjective scores for triplets of input, target and result images. Then, for each triplet, we compute a number of image features, which objectively characterize a color transfer. To describe the relationship between these features and the subjective scores, we build a regression model with random forests. An analysis and a cross-validation show that the predictions of our model are highly accurate. Hristina Hristova, Olivier Le Meur, Rémi Cozot, Kadi Bouatouch |
ICIP | 2 |
| 2017 | Age-dependent saccadic models for predicting eye movementsabstractHow people look at visual information reveals fundamental information about themselves, their interests and their state of mind. While previous visual attention models output static 2-dimensional saliency maps, saccadic models predict not only what observers look at but also how they move their eyes to explore the scene. Here we demonstrate that saccadic models are a flexible framework that can be tailored to emulate the gaze patterns from childhood to adulthood. The proposed age-dependent saccadic model not only outputs human-like, i.e. age-specific visual scanpath, but also significantly outperforms other state-of-the-art saliency models. Olivier Le Meur, Antoine Coutrot, Adrien Le Roch, Andrea Helo, Pia Rämä, Zhi Liu 0003 |
ICIP | 1 |
| 2017 | Saliency-based navigation in omnidirectional imageabstractOmnidirectional images describe the color information at a given position from all directions. Affordable 360° cameras have recently been developed leading to an explosion of the 360° data shared on social networks. However, an omnidirectional image does not contain interesting content everywhere. Some part of the images are indeed more likely to be looked at by some users than others. Knowing these regions of interest might be useful for 360° image compression, streaming, retargeting or even editing. In this paper, we aim at modelling the user navigation within a 360° image, and detecting which parts of an omnidirectional content might draw users' attention. In particular, the paper proposes to aggregate and analyze 2D saliency detectors in different map projections, and also proposes a smooth navigation through the image to maximize saliency. Thomas Maugey, Olivier Le Meur, Zhi Liu 0003 |
MMSP | 2 |
| 2017 | Depth-Guided Disocclusion Inpainting of Synthesized RGB-D ImagesabstractWe propose to tackle the problem of RGB-D image disocclusion inpainting when synthesizing new views of a scene by changing its viewpoint. Indeed, such a process creates holes both in depth and color images. First, we propose a novel algorithm to perform depth-map disocclusion inpainting. Our intuitive approach works particularly well for recovering the lost structures of the objects and to inpaint the depth-map in a geometrically plausible manner. Then, we propose a depth-guided patch-based inpainting method to fill-in the color image. Depth information coming from the reconstructed depth-map is added to each key step of the classical patch-based algorithm from Criminisi et al. in an intuitive manner. Relevant comparisons to the state-of-the-art inpainting methods for the disocclusion inpainting of both depth and color images are provided and illustrate the effectiveness of our proposed algorithms. Pierre Buyssens, Olivier Le Meur, Maxime Daisy, David Tschumperlé, Olivier Lézoray |
IEEE Trans. Image Process. | 2 |
| 2017 | Visual Attention Saccadic Models Learn to Emulate Gaze Patterns From Childhood to AdulthoodabstractHow people look at visual information reveals fundamental information about themselves, their interests and their state of mind. While previous visual attention models output static 2D saliency maps, saccadic models aim to predict not only where observers look at but also how they move their eyes to explore the scene. In this paper, we demonstrate that saccadic models are a flexible framework that can be tailored to emulate observer's viewing tendencies. More specifically, we use fixation data from 101 observers split into five age groups (adults, 8-10 y.o., 6-8 y.o., 4-6 y.o., and 2 y.o.) to train our saccadic model for different stages of the development of human visual system. We show that the joint distribution of saccade amplitude and orientation is a visual signature specific to each age group, and can be used to generate age-dependent scan paths. Our age-dependent saccadic model does not only output human-like, age-specific visual scan paths, but also significantly outperforms other state-of-the-art saliency models. We demonstrate that the computational modeling of visual attention, through the use of saccadic model, can be efficiently adapted to emulate the gaze behavior of a specific group of observers. Olivier Le Meur, Antoine Coutrot, Zhi Liu 0003, Pia Rämä, Adrien Le Roch, Andrea Helo |
IEEE Trans. Image Process. | 1 |
| 2017 | Depth-Aware Salient Object Detection and Segmentation via Multiscale Discriminative Saliency Fusion and Bootstrap LearningabstractThis paper proposes a novel depth-aware salient object detection and segmentation framework via multiscale discriminative saliency fusion (MDSF) and bootstrap learning for RGBD images (RGB color images with corresponding Depth maps) and stereoscopic images. By exploiting low-level feature contrasts, mid-level feature weighted factors and high-level location priors, various saliency measures on four classes of features are calculated based on multiscale region segmentation. A random forest regressor is learned to perform the discriminative saliency fusion (DSF) and generate the DSF saliency map at each scale, and DSF saliency maps across multiple scales are combined to produce the MDSF saliency map. Furthermore, we propose an effective bootstrap learning-based salient object segmentation method, which is bootstrapped with samples based on the MDSF saliency map and learns multiple kernel support vector machines. Experimental results on two large datasets show how various categories of features contribute to the saliency detection performance and demonstrate that the proposed framework achieves the better performance on both saliency detection and salient object segmentation. Hangke Song, Zhi Liu 0003, Huan Du, Guangling Sun, Olivier Le Meur, Tongwei Ren |
IEEE Trans. Image Process. | 5 |
| 2017 | Creating Segments and Effects on Comics by Clustering Gaze DataabstractTraditional comics are increasingly being augmented with digital effects, such as recoloring, stereoscopy, and animation. An open question in this endeavor is identifying where in a comic panel the effects should be placed. We propose a fast, semi-automatic technique to identify effects-worthy segments in a comic panel by utilizing gaze locations as a proxy for the importance of a region. We take advantage of the fact that comic artists influence viewer gaze towards narrative important regions. By capturing gaze locations from multiple viewers, we can identify important regions and direct a computer vision segmentation algorithm to extract these segments. The challenge is that these gaze data are noisy and difficult to process. Our key contribution is to leverage a theoretical breakthrough in the computer networks community towards robust and meaningful clustering of gaze locations into semantic regions, without needing the user to specify the number of clusters. We present a method based on the concept of relative eigen quality that takes a scanned comic image and a set of gaze points and produces an image segmentation. We demonstrate a variety of effects such as defocus, recoloring, stereoscopy, and animations. We also investigate the use of artificially generated gaze locations from saliency models in place of actual gaze locations. Ishwarya Thirunarayanan, Khimya Khetarpal, Sanjeev J. Koppal, Olivier Le Meur, John M. Shea, Eakta Jain |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2017 | High-dynamic-range image recovery from flash and non-flash image pairs
Hristina Hristova, Olivier Le Meur, Rémi Cozot, Kadi Bouatouch |
Vis. Comput. | 2 |
| 2015 | Spatiotemporal saliency detection based on superpixel-level trajectory
Zhi Liu 0003, Xiang Zhang 0006, Olivier Le Meur, Liquan Shen |
Signal Process. Image Commun. | 4 |
| 2015 | Special issue on recent advances in saliency models, applications and evaluations
Zhi Liu 0003, Olivier Le Meur, Ali Borji, Hongliang Li 0001 |
Signal Process. Image Commun. | 2 |
| 2015 | Video Inpainting With Short-Term Windows: Application to Object Removal and Error ConcealmentabstractIn this paper, we propose a new video inpainting method which applies to both static or free-moving camera videos. The method can be used for object removal, error concealment, and background reconstruction applications. To limit the computational time, a frame is inpainted by considering a small number of neighboring pictures which are grouped into a group of pictures (GoP). More specifically, to inpaint a frame, the method starts by aligning all the frames of the GoP. This is achieved by a region-based homography computation method which allows us to strengthen the spatial consistency of aligned frames. Then, from the stack of aligned frames, an energy function based on both spatial and temporal coherency terms is globally minimized. This energy function is efficient enough to provide high quality results even when the number of pictures in the GoP is rather small, e.g. 20 neighboring frames. This drastically reduces the algorithm complexity and makes the approach well suited for near real-time video editing applications as well as for loss concealment applications. Experiments with several challenging video sequences show that the proposed method provides visually pleasing results for object removal, error concealment, and background reconstruction context. Mounira Ebdelli, Olivier Le Meur, Christine Guillemot |
IEEE Trans. Image Process. | 2 |
| 2014 | Saliency Aggregation: Does Unity Make Strength?
Olivier Le Meur, Zhi Liu 0003 |
ACCV (4) | 1 |
| 2014 | HEVC Intra coding of ultra HD video with reduced complexityabstractThe HEVC (High Efficiency Video Coding) standard brings the necessary quality versus rate performance for efficient transmission of Ultra High Definition formats (UHD). However, one of the remaining barriers to its adoption for UHD content is its high encoding complexity. In this paper, we address the problem of HEVC encoding complexity reduction by proposing a strategy to infer UHD coding modes and quadtree structure from those optimized for the lower (HD) resolution version of the input video. A speed-up factor of 3 is achieved compared to directly encoding the UHD format at the expense of a limited quality loss. Nicolas Dhollande, Olivier Le Meur, Christine Guillemot |
ICIP | 2 |
| 2014 | Single image super-resolution using sparse representations with structure constraintsabstractThis paper describes a new single-image super-resolution algorithm based on sparse representations with image structure constraints. A structure tensor based regularization is introduced in the sparse approximation in order to improve the sharpness of edges. The new formulation allows reducing the ringing artefacts which can be observed around edges reconstructed by existing methods. The proposed method, named Sharper Edges based Adaptive Sparse Domain Selection (SE-ASDS), achieves much better results than many state-of-the-art algorithms, showing significant improvements in terms of PSNR (average of 29.63, previously 29.19), SSIM (average of 0.8559, previously 0.8471) and visual quality perception. Júlio César Ferreira, Olivier Le Meur, Christine Guillemot, Eduardo A. B. da Silva, Gilberto Arantes Carrijo |
ICIP | 2 |
| 2014 | Co-saliency detection based on region-level fusion and pixel-level refinementabstractThis paper addresses the problem of co-saliency detection, which aims to identify the common salient objects in a set of images and is important for many applications such as object co-segmentation and co-recognition. First, the segmentation driven low-rank matrix recovery model is used for intra saliency detection in each individual image of the image set, to highlight the regions whose features are sparse in each image. Then, a region-level fusion method, which exploits inter-region dissimilarities on color histograms and global consistency of regions over the image set, adjusts the intra saliency maps to obtain the region-level co-saliency maps, which can highlight co-salient object regions and suppress irrelevant regions. Finally, a pixel-level refinement method, which integrates color-spatial similarity between pixel and region with image border connectivity based object prior, generates the pixel-level co-saliency maps with better quality. Extensive experiments on two benchmark datasets demonstrate that the proposed co-saliency model consistently outperforms the state-of-the-art co-saliency models in both subjective and objective evaluation. Zhi Liu 0003, Wenbin Zou, Xiang Zhang 0006, Olivier Le Meur |
ICME | 5 |
| 2014 | Superpixel-Based Spatiotemporal Saliency DetectionabstractThis paper proposes a superpixel-based spatiotemporal saliency model for saliency detection in videos. Based on the superpixel representation of video frames, motion histograms and color histograms are extracted at the superpixel level as local features and frame level as global features. Then, superpixel-level temporal saliency is measured by integrating motion distinctiveness of superpixels with a scheme of temporal saliency prediction and adjustment, and superpixel-level spatial saliency is measured by evaluating global contrast and spatial sparsity of superpixels. Finally, a pixel-level saliency derivation method is used to generate pixel-level temporal and spatial saliency maps, and an adaptive fusion method is exploited to integrate them into the spatiotemporal saliency map. Experimental results on two public datasets demonstrate that the proposed model outperforms six state-of-the-art spatiotemporal saliency models in terms of both saliency detection and human fixation prediction. Zhi Liu 0003, Xiang Zhang 0006, Shuhua Luo, Olivier Le Meur |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2014 | Saliency Tree: A Novel Saliency Detection FrameworkabstractThis paper proposes a novel saliency detection framework termed as saliency tree. For effective saliency measurement, the original image is first simplified using adaptive color quantization and region segmentation to partition the image into a set of primitive regions. Then, three measures, i.e., global contrast, spatial sparsity, and object prior are integrated with regional similarities to generate the initial regional saliency for each primitive region. Next, a saliency-directed region merging approach with dynamic scale control scheme is proposed to generate the saliency tree, in which each leaf node represents a primitive region and each non-leaf node represents a non-primitive region generated during the region merging process. Finally, by exploiting a regional center-surround scheme based node selection criterion, a systematic saliency tree analysis including salient node selection, regional saliency adjustment and selection is performed to obtain final regional saliency measures and to derive the high-quality pixel-wise saliency map. Extensive experimental results on five datasets with pixel-wise ground truths demonstrate that the proposed saliency tree model consistently outperforms the state-of-the-art saliency models. Zhi Liu 0003, Wenbin Zou, Olivier Le Meur |
IEEE Trans. Image Process. | 3 |
| 2013 | Analysis of patch-based similarity metrics: Application to denoisingabstractThis paper presents a performance analysis of measures used for assessing similarities between patches. Compared to subjective ground thruth, our results indicate that some metrics are more suitable than others in a context of patch matching. This conclusion is confirmed by an experiment on non-local means (NLM) denoising algorithm. The denoising quality depends on the chosen similarity metric. In the best case, the gain is of 1.3dB compared to a classical SSD-based denoising algorithm. Mounira Ebdelli, Olivier Le Meur, Christine Guillemot |
ICASSP | 2 |
| 2013 | Image inpainting using LLE-LDNR and linear subspace mappingsabstractThe paper first describes an examplar-based image inpainting algorithm using a locally linear neighbor embedding technique with low-dimensional neighborhood representation (LLE-LDNR). The inpainting algorithm first searches the K nearest neighbors ( ) of the input patch to be filled-in and linearly combine them with LLE-LDNR to synthesize the missing pixels. Linear regression is then introduced for improving the K-NN search. The performance of the LLE-LDNR with the enhanced K-NN search method is assessed for two applications: loss concealment and object removal. Christine Guillemot, Mehmet Türkan, Olivier Le Meur, Mounira Ebdelli |
ICASSP | 3 |
| 2013 | Memorability of natural scenes: The role of attentionabstractThe image memorability consists in the faculty of an image to be recalled after a period of time. Recently, the memorability of an image database was measured and some factors responsible for this memorability were highlighted. In this paper, we investigate the role of visual attention in image memorability around two axis. The first one is experimental and uses results of eye-tracking performed on a set of images of different memorability scores. The second investigation axis is predictive and we show that attention-related features can advantageously replace low-level features in image memorability prediction. From our work it appears that the role of visual attention is important and should be more taken into account along with other low-level features. Matei Mancas, Olivier Le Meur |
ICIP | 2 |
| 2013 | 3D view synthesis with inter-view consistencyabstractIn this paper, we propose a new pipeline to synthesize virtual views by extrapolation. It allows us to generate virtual views far away from each other, each presenting the exact same level of quality. This inter-view consistency is key to seamlessly navigate between viewpoints. Its computational cost is also lower than that of existing approaches. We compare the proposed approach with state-of-the-art methods and show the effectiveness of this new view synthesis pipeline. David Wolinski, Olivier Le Meur, Josselin Gautier |
ACM Multimedia | 2 |
| 2013 | Object removal and loss concealment using neighbor embedding methods
Christine Guillemot, Mehmet Türkan, Olivier Le Meur, Mounira Ebdelli |
Signal Process. Image Commun. | 3 |
| 2013 | Hierarchical Super-Resolution-Based InpaintingabstractThis paper introduces a novel framework for examplar-based inpainting. It consists in performing first the inpainting on a coarse version of the input image. A hierarchical super-resolution algorithm is then used to recover details on the missing areas. The advantage of this approach is that it is easier to inpaint low-resolution pictures than high-resolution ones. The gain is both in terms of computational complexity and visual quality. However, to be less sensitive to the parameter setting of the inpainting method, the low-resolution input picture is inpainted several times with different configurations. Results are efficiently combined with a loopy belief propagation and details are recovered by a single-image super-resolution algorithm. Experimental results in a context of image editing and texture synthesis demonstrate the effectiveness of the proposed method. Results are compared to five state-of-the-art inpainting methods. Olivier Le Meur, Mounira Ebdelli, Christine Guillemot |
IEEE Trans. Image Process. | 1 |
| 2012 | Super-Resolution-Based Inpainting
Olivier Le Meur, Christine Guillemot |
ECCV (6) | 1 |
| 2012 | Examplar-based video inpainting with motion-compensated neighbor embeddingabstractThis paper describes a video inpainting algorithm based on motion-compensated neighbor embedding. The unknown pixels are estimated as a linear combination of the K closest patches using motion-compensated neighbor embedding. The algorithm is first assessed by assuming the motion information of the masked pixels to be known. This assumption is not realistic in video editing (object removal) applications. It however helps isolating the various problems for the sake of analysis. Different approaches are then assessed in the context where the motion information of missing pixels is unknown. Experiments on several videos show the benefits of the proposed approach which lead to natural looking videos with less annoying artefacts than when using a template matching technique. Mounira Ebdelli, Christine Guillemot, Olivier Le Meur |
ICIP | 3 |
| 2012 | Efficient depth map compression based on lossless edge coding and diffusionabstractThe multi-view plus depth video (MVD) format has recently been introduced for 3DTV and free-viewpoint video (FVV) scene rendering. Given one view (or several views) with its depth information, depth image-based rendering techniques have the ability to generate intermediate views. The MVD format however generates large volumes of data which need to be compressed for storage and transmission. This paper describes a new depth map encoding algorithm which aims at exploiting the intrinsic depth maps properties. Depth images indeed represent the scene surface and are characterized by areas of smoothly varying grey levels separated by sharp edges at the position of object boundaries. Preserving these characteristics is important to enable high quality view rendering at the receiver side. The proposed algorithm proceeds in three steps: the edges at object boundaries are first detected using a Sobel operator. The positions of the edges are encoded using the JBIG algorithm. The luminance values of the pixels along the edges are then encoded using an optimized path encoder. The decoder runs a fast diffusion-based inpainting algorithm which fills in the unknown pixels within the objects by starting from their boundaries. The performance of the algorithm is assessed against JPEG-2000 and HEVC, both in terms of PSNR of the depth maps versus rate as well as in terms of PSNR of the synthesized virtual views. Josselin Gautier, Olivier Le Meur, Christine Guillemot |
PCS | 2 |
| 2011 | Robustness and repeatability of saliency models subjected to visual degradationsabstractThe present study investigates the sensitivity of computationnal models of visual attention when subjected to visual degradations. One hundred and twenty natural color pictures were degraded using 6 filtering operations. By using different settings, five state-of-the-art models are used to compute 11400 saliency maps. The comparison of these maps to human saliency maps indicates that the tested models are robust to most of the visual degradations they were subjected to. These findings have implications on saliency-based applications, such as quality assessment and coding. A last point concerns the high repeatability of saliency models that might be used in a context of image retrieval. Olivier Le Meur |
ICIP | 1 |
| 2011 | Examplar-based inpainting based on local geometryabstractIn this paper, we propose a novel inpainting algorithm combining the advantages of PDE-based schemes and examplar-based approaches. The proposed algorithm relies on the use of structure tensors to define the filling order priority and template matching. The structure tensors are computed in a hierarchic manner whereas the template matching is based on a K-nearest neighbor algorithm. The value K is adaptively set in function of the local texture information. Compared to two state of the art approaches, the proposed method provides more coherent results. Olivier Le Meur, Josselin Gautier, Christine Guillemot |
ICIP | 1 |
| 2011 | Predicting saliency using two contextual priors: The dominant depth and the horizon lineabstractA computational model of visual attention using visual inferences is proposed. The dominant depth and the horizon line position are inferred from low-level visual features. This prior knowledge helps to find salient areas on still color pictures. Regarding the dominant depth, the idea is to favor the lowest spatial frequencies on close-up scenes whereas the highest spatial frequencies are used to predict salient areas on panoramic view. Some studies showed that the horizon line is a natural attractor of our gaze. Horizon detection is then used to improve the saliency prediction. Results show that the proposed model outperforms existing approaches. However, the dominant depth does not bring any gain in the saliency prediction. Olivier Le Meur |
ICME | 1 |
| 2011 | Prediction of the inter-observer visual congruency (IOVC) and application to image rankingabstractThis paper proposes an automatic method for predicting the inter-observer visual congruency (IOVC). The IOVC reflects the congruence or the variability among different subjects looking at the same image. Predicting this congruence is of interest for image processing applications where the visual perception of a picture matters such as website design, advertisement, etc. This paper makes several new contributions. First, a computational model of the IOVC is proposed. This new model is a mixture of low-level visual features extracted from the input picture where model's parameters are learned by using a large eye-tracking database. Once the parameters have been learned, it can be used for any new picture. Second, regarding low-level visual feature extraction, we propose a new scheme to compute the depth of field of a picture. Finally, once the training and the feature extraction have been carried out, a score ranging from 0 (minimal congruency) to 1 (maximal congruency) is computed. A value of 1 indicates that observers would focus on the same locations and suggests that the picture presents strong locations of interest. A second database of eye movements is used to assess the performance of the proposed model. Results show that our IOVC criterion outperforms the Feature Congestion measure \cite{Rosenholtz2007}. To illustrate the interest of the proposed model, we have used it to automatically rank personalized photograph. Olivier Le Meur, Thierry Baccino, Aline Roumy |
ACM Multimedia | 1 |
| 2010 | Spatio-temporal combination of saliency maps and eye-tracking assessment of different strategiesabstractThe modeling of the human visual attention into a computational attention model leads to the split of visual features into several independent channels. Then, a difficult problem arises to combine these maps, having different dynamic ranges or distribution. When several maps are considered, such process is mandatory in order to compute a single measure of interest for each location, regardless of which features contributed to the salience. Several strategies of cue combination are proposed in this paper for the spatial cues as well as the temporal saliency. Finally, some user tests on still image and video databases leads to highlight one operator. Christel Chamaret, Jean-Claude Chevet, Olivier Le Meur |
ICIP | 3 |
| 2010 | Overt visual attention for free-viewing and quality assessment tasks: Impact of the regions of interest on a video quality metric
Olivier Le Meur, Alexandre Ninassi, Patrick Le Callet, Dominique Barba |
Signal Process. Image Commun. | 1 |
| 2010 | Do video coding impairments disturb the visual attention deployment?
Olivier Le Meur, Alexandre Ninassi, Patrick Le Callet, Dominique Barba |
Signal Process. Image Commun. | 1 |
| 2010 | Relevance of a Feed-Forward Model of Visual Attention for Goal-Oriented and Free-Viewing TasksabstractA purely bottom-up model of visual attention is proposed and compared to five state-of-the-art models. The role of the low-level visual features is examined in two contexts. Two datasets are used: one containing data coming from an eye tracking experiment obtained in a free-viewing task and a second containing 5000 hand-label pictures (observers had to enclose the most visually interesting objects in a rectangle). The relevancy of the bottom-up models, i.e. the ability of a model to predict where the salient areas are located, is evaluated. Whatever the metrics and the datasets, the degree of similarity between predictions and ground truth is significantly above chance. The proposed model, resting on a small number of features, is shown to be a good predictor of the human visual fixations but also a good predictor of the objects chosen as interesting by observers. This study suggests that the low-level of visual features have a significant role in a free-viewing task but also in a high-level visual task, such as the choice of the object of interest in a complex visual scene. Another outcome concerns the viewing duration used in eye tracking experiments. Results suggest that this parameter is finally not as critical as one would expect. Olivier Le Meur, Jean-Claude Chevet |
IEEE Trans. Image Process. | 1 |
| 2009 | What we see is most likely to be what matters: Visual attention and applicationsabstractThe computational modeling of the visual attention is receiving increasing attention from the computer vision community. Several bottom-up models have been proposed. In spite of their complexities, these models are still a basic description of our visual system. Review of resulting approaches of these efforts are presented in the first part of this paper. Limitations of these approaches are introduced and several research trends are given. Among them, the most important one might be the use of prior knowledge, conjointly with the low-level visual features. Concomitantly with visual attention (VA) modeling progress, the image and video processing community is increasingly considering VA models in different fields or services. Current and future applications of VA models are discussed in the second part. Olivier Le Meur, Patrick Le Callet |
ICIP | 1 |
| 2008 | Which semi-local visual masking model forwavelet based image quality metric?abstractProperties and models of the human visual system (HVS) are the fundaments for most of sufficient objective image or video quality metrics. Among HVS properties, visual masking is a sensitive issue. Many models exist in literature. Simplest models can only predict visibility threshold for very simple cue while for natural images one should consider more complex approaches such as semi-local masking. Our previous work has shown the positive impact of incorporating semi-local masking in image quality metric according to one subjective study. It is important to consolidate this work with different subjective experiments. In this paper, different visual masking models, including contrast masking and semi-local masking, are evaluated according to three subjective studies. These subjective experiments were conducted with different protocols, different types of display devices, different contents and different populations. Alexandre Ninassi, Olivier Le Meur, Patrick Le Callet, Dominique Barba |
ICIP | 2 |
| 2008 | Attention-based video reframing: Validation using eye-trackingabstractWatching TV shows on cell phones is starting to become a reality. Nevertheless, there still exist some significant issues due to the small size of cell phone screens. The direct transfer of contents that are not specifically shot for the mobile device will provide indistinguishable objects. An automated way, delivering the best viewing experience is proposed in this paper. This solution significantly improves the visual comfort, by zooming in on the regions of interest. The relevance of this solution rests on its capability to preserve the visually important areas as well as the temporal stability. Eye-tracking experiments are one metric to assess the reframing quality. Involving 16 observers, they show that more than 90% of the visually important regions are kept in the reframed clip. Christel Chamaret, Olivier Le Meur |
ICPR | 2 |
| 2007 | Does where you Gaze on an Image Affect your Perception of Quality? Applying Visual Attention to Image Quality MetricabstractThe aim of an objective image quality assessment is to find an automatic algorithm that evaluates the quality of pictures or video as a human observer would do. To reach this goal, researchers try to simulate the Human Visual System (HVS). Visual attention is a main feature of the HVS, but few studies have been done on using it in image quality assessment. In this work, we investigate the use of the visual attention information in their final pooling step. The rationale of this choice is that an artefact is likely more annoying in a salient region than in other areas. To shed light on this point, a quality assessment campaign has been conducted during which eye movements have been recorded. The results show that applying the visual attention to image quality assessment is not trivial, even with the ground truth. Alexandre Ninassi, Olivier Le Meur, Patrick Le Callet, Dominique Barba |
ICIP (2) | 2 |
| 2006 | Efficient Saliency-Based Repurposing MethodabstractImages play a very relevant role in our daily life. People now can easily shoot and share pictures thanks to the exponential growth of the portable medias, such as digital cameras, mobile phone. As the display size of those devices is relatively small, browsing large pictures remains difficult. Content re-purposing is an elegant solution to deal with this problem. It consists in cropping the images in order to display only the most interesting parts of the picture. A new algorithm is proposed in this paper; the experiments described herein, leading to a qualitative and a quantitative assessment, show that the proposed solution outperforms the conventional method. Olivier Le Meur, Xavier Castellani, Patrick Le Callet, Dominique Barba |
ICIP | 1 |
| 2006 | A Coherent Computational Approach to Model Bottom-Up Visual AttentionabstractVisual attention is a mechanism which filters out redundant visual information and detects the most relevant parts of our visual field. Automatic determination of the most visually relevant areas would be useful in many applications such as image and video coding, watermarking, video browsing, and quality assessment. Many research groups are currently investigating computational modeling of the visual attention system. The first published computational models have been based on some basic and well-understood Human Visual System (HVS) properties. These models feature a single perceptual layer that simulates only one aspect of the visual system. More recent models integrate complex features of the HVS and simulate hierarchical perceptual representation of the visual input. The bottom-up mechanism is the most occurring feature found in modern models. This mechanism refers to involuntary attention (i.e., salient spatial visual features that effortlessly or involuntary attract our attention). This paper presents a coherent computational approach to the modeling of the bottom-up visual attention. This model is mainly based on the current understanding of the HVS behavior. Contrast sensitivity functions, perceptual decomposition, visual masking, and center-surround interactions are some of the features implemented in this model. The performances of this algorithm are assessed by using natural images and experimental measurements from an eye-tracking system. Two adequate well-known metrics (correlation coefficient and Kullbacl-Leibler divergence) are used to validate this model. A further metric is also defined. The results from this model are finally compared to those from a reference bottom-up model. Olivier Le Meur, Patrick Le Callet, Dominique Barba, Dominique Thoreau |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | A spatio-temporal model of the selective human visual attentionabstractA new spatio-temporal model for simulating the bottom-up visual attention is proposed. It has been built from numerous important properties of the human visual system (HVS). This paper focuses both on the architecture of the model and on its performances. Given that the spatial model of the bottom-up visual attention has already been defined [O. Le Meur et al., 2004], the temporal dimension is more accurately described. A qualitative and quantitative comparison with human fixations collected from an eye tracking apparatus is undertaken. From the former, the quality of the prediction is deemed very good whereas the latter illustrates that the best predictor of the human fixation consists of the sum all visual features (achromatic, chromatic and motion). Olivier Le Meur, Dominique Thoreau, Patrick Le Callet, Dominique Barba |
ICIP (3) | 1 |
| 2004 | Performance assessment of a visual attention system entirely based on a human vision modelingabstractIt is now commonly assumed that the human visual attention, which is a selecting process of the most relevant locations in a scene according to a particular behavior, is driven by both top-down (task-dependent) and bottom-up (signal-dependent) control. A new model attempting to simulate the bottom-up process has been designed Le Meur, O et al., (2004). This model is purely based on visual system properties that provides noticeable advantages compared to the classical published approaches. This paper focuses on the performance assessment of this model by achieving a comparison with real fixation points stemming from eye-tracking apparatus both subjectively and objectively. Olivier Le Meur, Patrick Le Callet, Dominique Barba, Dominique Thoreau |
ICIP | 1 |