Claire-Hélène Demarty

dblp:08/1526 · DBLP profile ↗
← Back
23ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0001-6549-584XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 A Review of Computational Memorability: A Benchmark Framework
abstract
Abstract One of the powers of visual media lies in its ability to create a lasting impression on the viewer’s memory. In this digital age, where media is abundant and attention spans are fleeting, the task of predicting which content will stick in the viewer’s mind has become a critical challenge in computer vision. Computational memorability seeks to address this by developing models that estimate how memorable a piece of media is likely to be. In this review we focus on the MediaEval Predicting Video Memorability benchmark, a recurring evaluation task that has run annually since 2018. This benchmark provides a unique and consistent framework for researchers to compare and refine their memorability prediction techniques using standardised datasets and metrics. Its reproducible framework has proven invaluable for tracking progress and fostering innovation in this rapidly evolving domain. We analyse the evolution of the benchmark across its 2018–2023 editions, discussing the challenges that still remain, such as the need for more interpretability in models and the difficulty of predicting subjective and context-dependent memorability. By analysing and synthesising the collective insights gained from this task, we endeavour to inspire new avenues of inquiry and drive progress towards a more comprehensive understanding of this topic.
Mihai Gabriel Constantin, Claire-Hélène Demarty, Camilo Fosco, Sebastian Halder 0001, Graham Healy, Bogdan Ionescu, Stefan Valentin Luncanu, Iván Martín-Fernández, Ana Matran-Fernandez, Rukiye Savran Kiziltepe, Alan F. Smeaton, Liviu-Daniel Stefan, Lorin Sweeney, Alba Garcia Seco de Herrera
Int. J. Comput. Vis.2
2025 ACC: Alternating Complementary Colors for Display Energy Reduction
abstract
Displays are by far the most consuming devices in the video chain. In this paper, a novel method to reduce their consumption is disclosed that leverages the human vision’s properties and replaces each image pixel color by two other complementary colors, whose average power is lower than the original one. When applied just before a display panel, such a replacement allows to reproduce identical visual stimuli and does not impact the quality of experience (QoE). Two versions of the method are proposed depending on how the complementary colors temporally or spatially alternate, that achieve an average power reduction gain of up to 3.8% (max. 11%).
Kilian Ravon, Claire-Hélène Demarty, Laurent Blondé
ICIP2
2025 Style-FG: A Style-based Framework for Film Grain Analysis and Synthesis
abstract
Film grain which used to be a by-product of the chemical processing in the analog film stock is a desirable feature in the era of digital cameras. Besides participating to the artistic intent during content creation, film grain has also interesting properties in the video compression chain such as its ability to mask compression artifacts. In this article, we use a deep learning-based framework for film grain analysis, generation, and synthesis. Our framework Style-FG consists of three modules: a style encoder performing film grain style analysis, a mapping network responsible for film grain style generation, and a synthesis network that generates and blends a specific grain style to a given content in a content-adaptive manner. All modules are trained jointly, thanks to dedicated loss functions, on a new large and diverse dataset of pairs of grain-free and grainy images that we made publicly available to the community. 1 Quantitative and qualitative evaluations show that fidelity to the reference grain, diversity of grain styles as well as a perceptually pleasant grain synthesis are achieved, demonstrating that each module outperforms the state-of-the-art in the task it was designed for. To contribute further to the sustainability necessary effort of the digital information and communication field, a light-weight version of Style-FG is also proposed, which demonstrates similar quantitative and qualitative performances, while reducing the number of network parameters by a factor of 92%.
Zoubida Ameur, Claire-Hélène Demarty, Olivier Le Meur, Daniel Ménard
ACM Trans. Multim. Comput. Commun. Appl.2
2024 3R-INN: How to Be Climate Friendly While Consuming/Delivering Videos?
Zoubida Ameur, Claire-Hélène Demarty, Daniel Ménard, Olivier Le Meur
ECCV (75)2
2023 Display Power Modeling for Energy Consumption Control
abstract
As the most consuming devices in the video chain, it is necessary to master the power consumption of displays. However it necessitates to have a precise modeling of their power consumption. This paper proposes two approaches that outperform the state-of-the-art, to model the power consumption of emissive displays: first through the estimation of the display parameters from a theoretical RGBW model, then through two deep-based models agnostic of the display technology, that output either a global power value or a power map corresponding to an input image. All three modelings were made possible thanks to the release of DISPLAYPOWER2k, a large dataset of power consumption values for OLED displays.
Claire-Hélène Demarty, Laurent Blondé, Olivier Le Meur
ICIP1
2023 Deep-Learning-Based Energy Aware Images
abstract
In this paper, we present a method to compute energy-aware images, that aims to reduce the energy consumption of displays. This method relies on a lightweight unsupervised deep model which finds out the best trade-off between visual quality and energy reduction. From an input image and an energy reduction rate, a dimming map is inferred. We show that the proposed model performs as good as state-of-the-art methods, while being much more simple. In addition, the dimming map computation is constrained in order to ease its distribution throughout the video chain.
Olivier Le Meur, Claire-Hélène Demarty, Laurent Blondé
ICIP2
2023 Style-based film grain analysis and synthesis
abstract
Film grain which used to be a by-product of the chemical processing in the analog film stock, is a desirable feature in the era of digital cameras. Besides participating to the artistic intent during content creation, film grain has also interesting properties in the video compression chain such as its ability to mask compression artifacts. In this paper, we use a deep learning-based framework for film grain analysis, generation and synthesis. Our framework consists of three modules: a style encoder performing film grain style analysis, a mapping network responsible for film grain style generation, and a synthesis network that generates and blends a specific grain style to a given content in a content-adaptive manner. All modules are trained jointly, thanks to dedicated loss functions, on a new large and diverse dataset of pairs of grain-free and grainy images that we made publicly available to the community1. Quantitative and qualitative evaluations show that fidelity to the reference grain, diversity of grain styles as well as a perceptually pleasant grain synthesis are achieved, demonstrating that each module outperforms the state-of-the-art in the task it was designed for.
Zoubida Ameur, Claire-Hélène Demarty, Olivier Le Meur, Daniel Ménard, Edouard François
MMSys2
2023 Invertible Energy-Aware Images
abstract
Displaying video content requires significant amounts of energy. In face of the energy crisis and climate emergency, we propose a new approach to produce energy-aware images. The purpose of such images is to consume less energy when displayed onscreen while maximizing the quality of experience. A new invertible neural network called Invertible Energy-Aware Network (InvEAN) is proposed to produce invertible energy-aware images, allowing to reduce the energy consumption of display devices and offering the possibility to recover the original image if required. Experimental results show that the InvEAN network outperforms two existing methods over three datasets.
Olivier Le Meur, Claire-Hélène Demarty
IEEE Signal Process. Lett.2
2023 Deep-Based Film Grain Removal and Synthesis
abstract
In this paper, deep learning-based techniques for film grain removal and synthesis that can be applied in video coding are proposed. Film grain is inherent in analog film content because of the physical process of capturing images and video on film. It can also be present in digital content where it is purposely added to reflect the era of analog film and to evoke certain emotions in the viewer or enhance the perceived quality. In the context of video coding, the random nature of film grain makes it both difficult to preserve and very expensive to compress. To better preserve it while compressing the content efficiently, film grain is removed and modeled before video encoding and then restored after video decoding. In this paper, a film grain removal model based on an encoder-decoder architecture and a film grain synthesis model based on a conditional generative adversarial network (cGAN) are proposed. Both models are trained on a large dataset of pairs of clean (grain-free) and grainy images. Quantitative and qualitative evaluations of the developed solutions were conducted and showed that the proposed film grain removal model is effective in filtering film grain at different intensity levels using two configurations: 1) a non-blind configuration where the film grain level of the grainy input is known and provided as input; and 2) a blind configuration where the film grain level is unknown. As for the film grain synthesis task, the experimental results show that the proposed model is able to reproduce realistic film grain with a controllable intensity level specified as input.
Zoubida Ameur, Wassim Hamidouche, Edouard François, Milos Radosavljevic, Daniel Ménard, Claire-Hélène Demarty
IEEE Trans. Image Process.6
2022 Affect in Multimedia: Benchmarking Violent Scenes Detection
abstract
In this article, we report on the creation of a publicly available, common evaluation framework for Violent Scenes Detection (VSD) in Hollywood and YouTube videos. We propose a robust data set, the VSD96, with more than 96 hours of video of various genres, annotations at different levels of detail (e.g., shot-level, segment-level), annotations of mid-level concepts (e.g., blood, fire), various pre-computed multi-modal descriptors, and over 230 system output results as baselines. This is the most comprehensive data set available to this date tailored to the VSD task and was extensively validated during the MediaEval benchmarking campaigns. Furthermore, we provide an in-depth analysis of the crucial components of VSD algorithms, by reviewing the capabilities and the evolution of existing systems (e.g., overall trends and outliers, the influence of the employed features and fusion techniques, the influence of deep learning approaches). Finally, we discuss the possibility of going beyond state-of-the-art performance via an ad-hoc late fusion approach. Experimentation is carried out on the VSD96 data. We provide the most important lessons learned and gained insights. The increasing number of publications using the VSD96 data underline the importance of the topic. The presented and published resources are a practitioner's guide and also a strong baseline to overcome, which will help researchers for the coming years in analyzing aspects of audio-visual affect and violence detection in movies and videos.
Mihai Gabriel Constantin, Liviu-Daniel Stefan, Bogdan Ionescu, Claire-Hélène Demarty, Mats Sjöberg, Markus Schedl, Guillaume Gravier
IEEE Trans. Affect. Comput.4
2021 Visual Interestingness Prediction: A Benchmark Framework and Literature Review
Mihai Gabriel Constantin, Liviu-Daniel Stefan, Bogdan Ionescu, Ngoc Q. K. Duong, Claire-Hélène Demarty, Mats Sjöberg
Int. J. Comput. Vis.5
2019 VideoMem: Constructing, Analyzing, Predicting Short-Term and Long-Term Video Memorability
abstract
Humans share a strong tendency to memorize/forget some of the visual information they encounter. This paper focuses on understanding the intrinsic memorability of visual content. To address this challenge, we introduce a large scale dataset (VideoMem) composed of 10,000 videos with memorability scores. In contrast to previous work on image memorability - where memorability was measured a few minutes after memorization - memory performance is measured twice: a few minutes and again 24-72 hours after memorization. Hence, the dataset comes with short-term and long-term memorability annotations. After an in-depth analysis of the dataset, we investigate various deep neural network-based models for the prediction of video memorability. Our best model using a ranking loss achieves a Spearman's rank correlation of 0.494 (respectively 0.256) for short-term (resp. long-term) memorability prediction, while our model with attention mechanism provides insights of what makes a content memorable. The VideoMem dataset with pre-extracted features is publicly available.
Romain Cohendet, Claire-Hélène Demarty, Ngoc Q. K. Duong, Martin Engilberge
ICCV2
2018 Interestingness Prediction & its Application to Immersive Content
abstract
Which parts or objects are interesting in a content? In this paper we first propose three computational models to automatically predict interestingness rankings of areas/objects inside a 2D picture. We based our modeling on previous experimental findings to ensure reliability of the prediction when compared to the human assessement of interestingness. Our two first models are based on low level features, extracted from image regions, which have been stated as useful in the human interest process. A baseline model is built by estimating a linear regression from a small dataset of 49 images. The second model estimates a rewarding term based on additional experimental observations. By adding image semantics, we then construct a last model, which more generally benefits from a better understanding of the content. It also integrates notions such that unusualness or human beings' presence that have proven to play key roles in the interestingness process. Finally, targeting VR applications, we extend our models to immersive content, both images and videos, and propose an innovative application to guide the viewer in his/her navigation based on intuitive visual or audio cues.
Gwenaëlle Marquant, Claire-Hélène Demarty, Christel Chamaret, Joel Sirot, Louis Chevallier
CBMI2
2018 Deep Learning for Predicting Image Memorability
abstract
Memorability of media content such as images and videos has recently become an important research subject in computer vision. This paper presents our computation model for predicting image memorability, which is based on a deep learning architecture designed for a classification task. We exploit the use of both convolutional neural network (CNN) - based visual features and semantic features related to image captioning for the task. We train and test our model on the large-scale benchmarking memorability dataset: LaMem. Experiment result shows that the proposed computational model obtains better prediction performance than the state of the art, and even outperforms human consistency. We further investigate the genericity of our model on other memorability datasets. Finally, by validating the model on interestingness datasets, we reconfirm the uncorrelation between memorability and interestingness of images.
Hammad Squalli-Houssaini, Ngoc Q. K. Duong, Gwenaëlle Marquant, Claire-Hélène Demarty
ICASSP4
2018 Annotating, Understanding, and Predicting Long-term Video Memorability
abstract
Memorability can be regarded as a useful metric of video importance to help make a choice between competing videos. Research on computational understanding of video memorability is however in its early stages. There is no available dataset for modelling purposes, and the few previous attempts provided protocols to collect video memorability data that would be difficult to generalize. Furthermore, the computational features needed to build a robust memorability predictor remain largely undiscovered. In this article, we propose a new protocol to collect long-term video memorability annotations. We measure the memory performances of 104 participants from weeks to years after memorization to build a dataset of 660 videos for video memorability prediction. This dataset is made available for the research community. We then analyze the collected data in order to better understand video memorability, in particular the effects of response time, duration of memory retention and repetition of visualization on video memorability. We finally investigate the use of various types of audio and visual features and build a computational model for video memorability prediction. We conclude that high level visual semantics help better predict the memorability of videos.
Romain Cohendet, Karthik Yadati, Ngoc Q. K. Duong, Claire-Hélène Demarty
ICMR4
2017 Deep learning for multimodal-based video interestingness prediction
abstract
Predicting interestingness of media content remains an important, but challenging research subject. The difficulty comes first from the fact that, besides being a high-level semantic concept, interestingness is highly subjective and its global definition has not been agreed yet. This paper presents the use of up-to-date deep learning techniques for solving the task. We perform experiments with both social-driven (i.e., Flickr videos) and content-driven (i.e., videos from the MediaEval 2016 interestingness task) datasets. To account for the temporal aspect and multimodality of videos, we tested various deep neural network (DNN) architectures, including a new combination of several recurrent neural networks (RNNs), to handle several temporal samples at the same time. We then investigated different strategies for dealing with unbalanced datasets. Multimodality, as the mid-level fusion of audio and visual information, brought benefit to the task. We also established that social interestingness differs from content interestingness.
Yuesong Shen, Claire-Hélène Demarty, Ngoc Q. K. Duong
ICME2
2015 VSD, a public dataset for the detection of violent scenes in movies: design, annotation, analysis and evaluation
Claire-Hélène Demarty, Cédric Penet, Mohammad Soleymani 0001, Guillaume Gravier
Multim. Tools Appl.1
2015 Variability modelling for audio events detection in movies
Cédric Penet, Claire-Hélène Demarty, Guillaume Gravier, Patrick Gros
Multim. Tools Appl.2
2014 Classification-oriented structure learning in Bayesian networks for multimodal event detection in videos
Guillaume Gravier, Claire-Hélène Demarty, Siwar Baghdadi, Patrick Gros
Multim. Tools Appl.2
2012 Multimodal information fusion and temporal integration for violence detection in movies
abstract
This paper presents a violent shots detection system that studies several methods for introducing temporal and multimodal information in the framework. It also investigates different kinds of Bayesian network structure learning algorithms for modelling these problems. The system is trained and tested using the MediaEval 2011 Affect Task corpus, which comprises of 15 Hollywood movies. It is experimentally shown that both multimodality and temporality add interesting information into the system. Moreover, the analysis of the links between the variables of the resulting graphs yields important observations about the quality of the structure learning algorithms. Overall, our best system achieved 50% false alarms and 3% missed detection, which is among the best submissions in the MediaEval campaign.
Cédric Penet, Claire-Hélène Demarty, Guillaume Gravier, Patrick Gros
ICASSP2
2012 Electrodermal activity applied to violent scenes impact measurement and user profiling
abstract
Identifying violent scenes in a movie may be of high interest as soon as the associated content has to be shown to a specific audience, as children for instance. However, defining what a violent scene is as well as extracting the violent excerpts from the only audiovisual cue are two hard tasks. In this article, we propose a pilot study to evaluate the interest of an approach based on the use of the electrodermal activity (EDA) to address these problems in an objective and human-centered manner. Assuming a consumer context, we especially focus on the use of a commercial sensor to capture the EDA. A comparison with a more professional device is initially performed to validate the accuracy of the commercial sensor. Two main aspects of the violent scene impact measurement with the EDA are then discussed. To tackle the violent scene detection problem, a methodology to correlate the EDA information with manual annotations of violent excerpts is first proposed. The user sensitivity to violence is then addressed and an approach to identify different profiles in the audience is qualitatively proposed. The advantages and limitations of such an approach are discussed and potential improvements are finally proposed.
Julien Fleureau, Cédric Penet, Philippe Guillotel, Claire-Hélène Demarty
SMC4
2008 Structure learning in a Bayesian network-based video indexing framework
abstract
Several stochastic models provide an effective framework to identify the temporal structure of audiovisual data. Most of them need as input a first video structure, i.e. connections between features and video events. Provided that this structure is given as input, the parameters are then estimated from training data. Bayesian networks offer an additional feature, namely structure learning, which allows the automatic construction of the model structure from training data. Structure learning obviously leads to an increased generality of the model building process. This paper investigates the trade-off between the increase of generality and the quality of the results in video analysis. We model video data using dynamic Bayesian networks (DBNs) where the static part of the network accounts for the correlations between low-level features extracted from the raw data and between these features and the events considered. It is precisely this part of the network whose structure is automatically constructed from training data. Experimental results on a commercial detection case study application show that, even though the model structure is determined in a non supervised manner, the resulting model is effective for the detection of commercial segments in video data.
Siwar Baghdadi, Guillaume Gravier, Claire-Hélène Demarty, Patrick Gros
ICME3
2006 A Video Fingerprint Based on Visual Digest and Local Fingerprints
abstract
A fingerprinting design extracts discriminating features, called fingerprints. The extracted features are unique and specific to each image/video. The visual hash is usually a global fingerprinting technique with crypto-system constraints. In this paper, we propose an innovative video content identification process which combines a visual hash function and a local fingerprinting. Thanks to a visual hash function, we observe the video content variation and we detect key frames. A local image fingerprint technique characterizes the detected key frames. The set of local fingerprints for the whole video summarizes the video or fragments of the video. The video fingerprinting algorithm identifies an unknown video or a fragment of video within a video fingerprint database. It compares the local fingerprints of the candidate video with all local fingerprints of a database even if strong distortions are applied to an original content.
Ayoub Massoudi, Frédéric Lefèbvre, Claire-Hélène Demarty, Lionel Oisel, Bertrand Chupeau
ICIP3