VLDB 2026 Research / reviewers in the wild / expert
Jenny Benois-Pineau
dblp:59/4268
· DBLP profile ↗
115ranked-venue papers
9as first author
20since 2021 · last 2026
0000-0003-0659-8894ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 98 · 8 first-author · 18 since 2021Artificial intelligence and machine learning · 34 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | There Is More to Attention: Statistical Filtering Enhances Explanations in Vision Transformers
Meghna Ayyar, Jenny Benois-Pineau, Akka Zemmari |
ICPR (5) | 2 |
| 2025 | First-Person Human Sensing for Upper Limb Neuroprosthesis Control: 6D Pose Estimation of Objects to GraspabstractHuman behaviour sensing and understanding with the analysis of human intentions is a mandatory part of many assistive and healthcare IT scenarios. In the present work we propose an integrative approach for understanding human intentions and its context for vision-assisted control of upper limb prostheses. In particular, we focus on the estimation of the 6D pose, i.e. the spatial coordinates and the rotation angles around the axes of the egocentric system of first-person-view camera mounted on the glasses. We propose a solution based on the DenseFusion backbone, which is applicable to a real-world prosthesis-assisting scenario due to its fairly good accuracy. Indeed, the error in the object position measured as the mean error of object model points with a fine-tuned model on a controllable synthetic dataset is of 0.8 cm, which is an acceptable range for Robotic API in our scenario of assisting amputees wearing upper limb prostheses. Ander Etxezarreta, Jenny Benois-Pineau, Renaud Péteri, Lucas Bardisbanian, Aymar de Rugy |
CBMI | 2 |
| 2024 | A Hybrid AI System for Fusion of Object and Context Information: Application to the Rail Line Defect DetectionabstractA hybrid artificial intelligence (Hybrid AI) which represents a convergence of a classical (symbolic) AI with recent machine learning approaches has become a very quickly developing research axis. The combination of rule-based reasoning and statistical learning is required whenever the domain knowledge has to be incorporated in the decision system. In this work we present a system on the basis of Deep Neural Networks (DNNs) as object detectors, such as You Only Look Once version 8 (YOLOv8), transformers and logical rules which link objects and their context in the problem of rail line defect detection. Fusion of information is performed at the intermediate level - in the feature space, mixing sets of elements of this space delimited due to the object and context element detectors. Combination of objects and context elements is performed accordingly to the domain-defined rules, and fusion is ensured by a vision transformer. Experiments have been conducted on the domainrecorded dataset of rail defects. The proposed hybrid system outperforms base-line objects detection up to 0.28 of accuracy increase. Alexey Zhukov, Jenny Benois-Pineau, Alain Rivero, Akka Zemmari, Mohamed Mosbah 0001, Danilo Crispiani |
CBMI | 2 |
| 2024 | ET: Explain to Train: Leveraging Explanations to Enhance the Training of A Multimodal TransformerabstractExplainable Artificial Intelligence (XAI) has become increasingly vital for improving the transparency and reliability of neural network decisions. Transformer architectures have emerged as the state-of-the-art for various tasks across single modalities such as video, language, or signals, as well as for multimodal approaches. Although XAI methods for transformers are available, their potential impact during model training remains underexplored. Thus, we propose Explanation-guided Training (ET), leveraging an XAI method to identify salient input regions and guide the model to focus solely on these salient regions during training. We develop ET in a typical multimodal analysis framework using a multimodal transformer that operates on videos and signals. ET enhances the input by masking the non-salient regions for videos and enhances the signals with weights based on explanation scores for the sensor modality. Comparative evaluation with baseline vanilla training and the state-of-the-art XAI-based IFI method [1] shows that ET consistently outperforms them. We benchmark our method on the publicly available UCF50 video dataset to demonstrate that ET is better than vanilla training and IFI. A risk detection corpus comprising egocentric videos and wearable sensor data is used for multimodal evaluation. Our code is available at https://gitub.u-bordeaux.fr/mayyar/explain-to-train Meghna Ayyar, Jenny Benois-Pineau, Akka Zemmari |
ICIP | 2 |
| 2024 | ESL: Explain to Improve Streaming Learning for Transformers
Meghna Ayyar, Jenny Benois-Pineau, Akka Zemmari |
ICPR (9) | 2 |
| 2024 | IFI: Interpreting for Improving: A Multimodal Transformer with an Interpretability Technique for Recognition of Risk Events
Rupayan Mallick, Jenny Benois-Pineau, Akka Zemmari |
MMM (4) | 2 |
| 2024 | A hybrid transformer with domain adaptation using interpretability techniques for the application to the detection of risk situations
Rupayan Mallick, Jenny Benois-Pineau, Akka Zemmari, Kamel Guerda, Boris Mansencal, Hélène Amieva, Laura Middleton |
Multim. Tools Appl. | 2 |
| 2024 | Asymmetric Multi-Task Learning for Interpretable Gaze-Driven Grasping Action ForecastingabstractThis work tackles the automatic prediction of grasping intention of humans observing their environment. Our target application is the assistance to people with motor disabilities and potential cognitive impairments, using assistive robotics. Our proposal leverages the analysis of human attention captured in the form of gaze fixations recorded by an eye-tracker on the first person video, as the anticipation of prehension actions is a well studied and well known phenomenon. We propose a multi-task system that simultaneously addresses the prediction of human attention in the near future, and the anticipation of grasping actions. Visual attention is modeled as a competitive process between a discrete set of states, each one associated to a well-known gaze movement pattern from visual psychology. We additionally consider an asymmetric multi-task problem, where attention modeling is an auxiliary task that helps to regularize the learning process of the main action prediction task, and propose a constrained multi-task loss that naturally deals with this asymmetry. Our model shows superior performance than other losses for dynamic multi-task learning, current dominant deep architectures for general action forecasting and particularly-tailored models for predicting grasping intention. In particular, it provides state-of-the-art performance in three datasets for egocentric action anticipation, with an average precision of 0.569 and 0.524 in GITW and Sharon datasets, respectively, and an accuracy of 89.2% and a success rate of 51.7% in Invisible dataset. Iván González-Díaz 0001, Miguel Molina-Moreno, Jenny Benois-Pineau, Aymar de Rugy |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Entropy-based Sampling for Streaming learning with Move-to-Data approach on VideoabstractThe current paradigm of training deep neural networks relies on large, annotated and representative datasets. They assume a static world where the target domain does not change. However, in the real-world, data changes over time and is often available on the fly. Naive retraining on new data causes catastrophic forgetting and the network is unable to generalize on old data. Streaming learning is a type of incremental learning where networks learn sequentially and as soon as a sample is available from the data stream. Instead of training on every new sample, we propose an uncertainty based selection criteria to improve our previously proposed fast streaming learning method Move-to-Data (MTD), called Entropy-based MTD (EMTD). Besides, streaming learning methods have so far mostly used Convolutional Neural Networks (CNNs) but in recent times Vision Transformers (ViTs) have shown much better performances for many vision tasks. Therefore, we use ViT based Video Transformer to analyse MTD, EMTD and their gradient descent based "retargeting" steps. We have compared the performances of EMTD with MTD (w/wo retargeting) and a popular streaming learning method ExStream for the transformer. EMTD is able to outperform baseline MTD, and EMTD with retargeting achieves close results as ExStream and is ∼ 1.2 times faster. Meghna Ayyar, Jenny Benois-Pineau, Akka Zemmari, Hélène Amieva, Laura Middleton |
CBMI | 2 |
| 2023 | On the stability, correctness and plausibility of visual explanation methods based on feature importanceabstractIn the field of Explainable AI, multiples evaluation metrics have been proposed in order to assess the quality of explanation methods w.r.t. a set of desired properties. In this work, we study the articulation between the stability, correctness and plausibility of explanations based on feature importance for image classifiers. We show that the existing metrics for evaluating these properties do not always agree, raising the issue of what constitutes a good evaluation metric for explanations. Finally, in the particular case of stability and correctness, we show the possible limitations of some evaluation metrics and propose new ones that take into account the local behaviour of the model under test. Romain Xu-Darme, Jenny Benois-Pineau, Romain Giot, Georges Quénot, Zakaria Chihani, Marie-Christine Rousset, Alexey Zhukov |
CBMI | 2 |
| 2022 | Streaming learning with Move-to-Data approach for image classificationabstractIn Deep Neural Network training, the availability of a large amount of representative training data is the sine qua non-condition for a good generalization capacity of the model. In many real-world applications, data is not available at a glance, but coming on the fly. If a pre-trained model is fine-tuned on the new data, then catastrophic forgetting happens mostly. Incremental learning mechanisms propose ways to overcome catastrophic forgetting. Streaming learning is a type of incremental learning where models learn from new data instances as soon as they become available in a single training pass. In this work, we conduct an experimental study, on a large dataset, of an incremental/streaming learning method Move-to-Data we previously proposed, and propose an updated approach by ”re-targeting” with gradient descent which is faster than the popular streaming learning method ExStream. The method achieves better performances and computational efficiency compared to ExStream. Move-to-Data with gradient is on average 3.5 times faster than ExStream and has a similar accuracy, with 0.5% improvement compared to ExStream. Abel Kahsay Gebreslassie, Jenny Benois-Pineau, Akka Zemmari |
CBMI | 2 |
| 2022 | I Saw: A Self-Attention Weighted Method for Explanation of Visual TransformersabstractRecently, visual transformers have shown promising results in tasks such as image classification, segmentation, object detection, etc. The explanation of their decision remains a challenge. This paper focuses on exploiting self-attention for an explanation. We propose a generalized interpretation of the transformers i.e model agnostic but class-specific explanations. The main principle is in the use and weighting self-attention maps of a visual transformer. To evaluate it, we use the popular hypothesis that an explanation is good if it correlates with human perception of a visual scene. Thus, the method has been evaluated against the Gaze Fixation Density Maps obtained in a psycho-visual experiment on a public database. It has been compared with other popular explainers such as Grad-Cam, LRP, Rollout, and Adaptive Relevance methods. The proposed method outperforms the best baseline by 2% in a standard Pearson Correlation Coefficient (PCC) metric. Rupayan Mallick, Jenny Benois-Pineau, Akka Zemmari |
ICIP | 2 |
| 2022 | Pooling Transformer for Detection of Risk Events in In-The-Wild Video Ego DataabstractThe paper proposes a video transformer architecture for detection of risk events on frail adults with ego video monitoring data. First we introduce an extended taxonomy for risk events, and then we propose a transformer based video recognition model for detection of these risk events. The proposed transformer architecture consists of separable attention for spatial and temporal data. We also introduce a pooling operation on the temporal video data by learning of their importance. The experiments have been conducted on visual data of in-the-wild recorded BIRDS dataset and on Kinetics-400 for benchmarking. The use of the pooling operation in transformers gives an increment of 3% on BIRDS dataset. Rupayan Mallick, Jenny Benois-Pineau, Akka Zemmari, Thinhinane Yebda, Marion Pech, Hélène Amieva, Laura Middleton |
ICPR | 2 |
| 2022 | BG-3DM2F: Bidirectional gated 3D multi-scale feature fusion for Alzheimer's disease diagnosis
Ibtissam Bakkouri, Karim Afdel, Jenny Benois-Pineau, Gwénaëlle Catheline |
Multim. Tools Appl. | 3 |
| 2022 | Visual vs internal attention mechanisms in deep neural networks for image classification and object detection
Abraham Montoya Obeso, Jenny Benois-Pineau, Mireya S. García-Vázquez, Alejandro Alvaro Ramírez-Acosta |
Pattern Recognit. | 2 |
| 2021 | A GRU Neural Network with attention mechanism for detection of risk situations on multimodal lifelog dataabstractMultimedia today is also in multimodality. Working with heterogeneous signals we use multimedia techniques of data fusion and mining. Classification from real world datasets are often challenging. The paper is devoted to the detection of personal risk situations of fragile people from multi-modal sensing real world lifelog data named BIRDS. Using a real-world data is challenging as the risk situations are rare and last just a few seconds compared to the global volume of the dataset. In this paper we propose a GRU architecture with global attention block to recognise semantic risk situations from a limited taxonomy. Attention is also focused on data organisation and pre-processing with imputation and normalisation. The proposed method is applied to a real-world collected multimodal dataset and to the OpenSource dataset UCI-HAR for the sake of comparison with the state-of-the-art. Rupayan Mallick, Thinhinane Yebda, Jenny Benois-Pineau, Akka Zemmari, Marion Pech, Hélène Amieva |
CBMI | 3 |
| 2021 | Explaining 3D CNNs for Alzheimer's Disease Classification on sMRI Images with Multiple ROIsabstractClassification of Alzheimer’s disease from 3D structural Magnetic Resonance Imaging (sMRI) with deep neural networks has shown promising results in recent years. The decision interpretation of these networks is essential to aid medical experts to understand and rely on the results provided by such models. In this paper, we propose an adaptation of a recently developed feature-based explanation method and apply it to a 3D CNN architecture for the binary classification of Alzheimer’s disease and Normal Control from the hippocampal ROIs of brain sMRIs. We also compare our method to the state-of-the-art LRP method. Meghna Ayyar, Jenny Benois-Pineau, Akka Zemmari, Gwénaëlle Catheline |
ICIP | 2 |
| 2021 | Analysis of Deep Neural Networks Correlations with Human Subjects on a Perception TaskabstractIn information visualization, it has become mandatory to assess visualization techniques efficiency either to write a survey, optimize a technique or even design a new one. To do so, the common way is to conduct user evaluations through which human subjects are asked to solve a task on different visualization techniques while their performances are measured to assess which technique is the most efficient. These evaluations can be complex to design and setup in order not to be biased and, in the end, their results can become contestable when the evaluation methods standards evolve. To overcome these flaws, new evaluation methods are emerging, mostly making use of modern and efficient computer vision techniques such as deep learning. These new methods rely on a strong assumption that has not been studied deeply enough yet: humans and deep learning models performances can be correlated. This paper explores the performances of both a state-of-the-art deep neural network and human subjects on an outlier detection task taken from a previous experiment of the literature. The objective is to study whether the machine and humans behaviors were different or if some correlations can be observed. Our study shows that their results are significantly correlated and a machine learning model efficiently learned to predict human performances using deep neural network metrics as input. Hence, this work presents a use case where using a deep neural network to assess human subjects performances is efficient. Loann Giovannangeli, Romain Giot, David Auber, Jenny Benois-Pineau, Romain Bourqui |
IV | 4 |
| 2021 | Multimodal Sensor Data Analysis for Detection of Risk Situations of Fragile People in @home Environments
Thinhinane Yebda, Jenny Benois-Pineau, Marion Pech, Hélène Amieva, Laura Middleton, Max Bergelt |
MMM (2) | 2 |
| 2021 | Special issue on content-based multimedia indexing in the era of artificial intelligence
Stevan Rudinac, Jenny Benois-Pineau, Stéphane Marchand-Maillet |
Multim. Tools Appl. | 2 |
| 2020 | 3D attention mechanism for fine-grained classification of table tennis strokes using a Twin Spatio-Temporal Convolutional Neural NetworksabstractThe paper addresses the problem of recognition of actions in video with low inter-class variability such as Table Tennis strokes. Two stream, “twin” convolutional neural networks are used with 3D convolutions both on RGB data and optical flow. Actions are recognized by classification of temporal windows. We introduce 3D attention modules and examine their impact on classification efficiency. In the context of the study of sportsmen performances, a corpus of the particular actions of table tennis strokes is considered. The use of attention blocks in the network speeds up the training step and improves the classification scores up to 5% with our twin model. We visualize the impact on the obtained features and notice correlation between attention and player movements and position. Score comparison of state-of-the-art action classification method and proposed approach with attentional blocks is performed on the corpus. Proposed model with attention blocks outperforms previous model without them and our baseline. Pierre-Etienne Martin, Jenny Benois-Pineau, Renaud Péteri, Julien Morlier |
ICPR | 2 |
| 2020 | Detection of Semantic Risk Situations in Lifelog Data for Improving Life of Frail PeopleabstractThe automatic recognition of risk situations for frail people is an urgent research topic for the interdisciplinary artificial intelligence and multimedia community. Risky situations can be recognized from lifelog data recorded with wearable devices. In this paper, we present a new approach for the detection of semantic risk situations for frail people in lifelog data. Concept matching between general lifelog and risk taxonomies was realized and tuned AlexNet was deployed for detection of two semantic risks situations such as risk of domestic accident and risk of fraud with promising results. Thinhinane Yebda, Jenny Benois-Pineau, Marion Pech, Hélène Amieva, Cathal Gurrin |
ICMR | 2 |
| 2020 | Instrument Recognition in Laparoscopy for Technical Skill Assessment
Sabrina Kletz, Klaus Schöffmann, Andreas Leibetseder, Jenny Benois-Pineau, Heinrich Husslein |
MMM (2) | 4 |
| 2020 | Fine grained sport action recognition with Twin spatio-temporal convolutional neural networks
Pierre-Etienne Martin, Jenny Benois-Pineau, Renaud Péteri, Julien Morlier |
Multim. Tools Appl. | 2 |
| 2019 | Recognition of Alzheimer's Disease on sMRI based on 3D Multi-Scale CNN Features and a Gated Recurrent Fusion UnitabstractAccurate diagnosis of Alzheimer's Disease (AD) is still a public health challenge, and has been studied for several years now to make it efficient and more automatic. In this paper, we propose a novel Computer-Aided Diagnosis (CAD) system based on 3D Multi-scale Feature (3DMF) blocks and Gated Recurrent Fusion Unit (GRFU). Hippocampal Volumes Of Interest (VOI) are used as input. The method is applied on sMRI imaging modality standard for patients screening. First, multiscale features are extracted via 3D Convolutional Neural Network (CNN). They are then taken as input to Gated Recurrent Units (GRU) for performance improvement. Extensive experiments are performed on the public Alzheimers Disease Neuroimaging Initiative (ADNI) dataset. The experimental results demonstrate that our 3DMF model combined with GRFU obtains the state-of-the-art performance compared with the existing conventional methods. Furthermore, proposed approach yields a significant enhancement in terms of avoiding high similarity between classes and overfitting issue. Hence, our proposed CAD has the potential to significantly improve the conventional recognition and classification strategies for use in clinical applications. Ibtissam Bakkouri, Karim Afdel, Jenny Benois-Pineau, Gwénaëlle Catheline |
CBMI | 3 |
| 2019 | Identifying Surgical Instruments in Laparoscopy Using Deep Learning Instance SegmentationabstractRecorded videos from surgeries have become an increasingly important information source for the field of medical endoscopy, since the recorded footage shows every single detail of the surgery. However, while video recording is straightforward these days, automatic content indexing - the basis for content-based search in a medical video archive - is still a great challenge due to the very special video content. In this work, we investigate segmentation and recognition of surgical instruments in videos recorded from laparoscopic gynecology. More precisely, we evaluate the achievable performance of segmenting surgical instruments from their background by using a region-based fully convolutional network for instance-aware (1) instrument segmentation as well as (2) instrument recognition. While the first part addresses only binary segmentation of instances (i.e., distinguishing between instrument or background) we also investigate multi-class instrument recognition (i.e., identifying the type of instrument). Our evaluation results show that even with a moderately low number of training examples, we are able to localize and segment instrument regions with a pretty high accuracy. However, the results also reveal that determining the particular instrument is still very challenging, due to the inherently high similarity of surgical instruments. Sabrina Kletz, Klaus Schöffmann, Jenny Benois-Pineau, Heinrich Husslein |
CBMI | 3 |
| 2019 | Dropping Activations in Convolutional Neural Networks with Visual Attention MapsabstractThe introduction of visual attention models in data selection and features selection in CNNs for the task of image classification is an intensive and interesting research topic. In CNNs, the strategy of dropping activations, after features extraction layers, shown an increase in the generalization gap in large-scale datasets and avoiding over-fitting. Dropout has been studied in the literature in a fully-randomized manner to take down activations during training. In this paper, we introduce a saliency-based dropping strategy to take down activations in our AlexNet-like architecture. Our experiments are conducted for the specific task of specific Mexican architectural recognition, in 67 categories. The results are promising: the proposed approach outperformed other models reducing training time and reaching a higher accuracy. Abraham Montoya Obeso, Jenny Benois-Pineau, Mireya S. García-Vázquez, Alejandro Alvaro Ramírez-Acosta |
CBMI | 2 |
| 2019 | Multi-sensing of fragile persons for risk situation detection: devices, methods, challengesabstractThe ageing of the world population has raised ever-increasing demands for measuring physical conditions and assisting the elderly and/or fragile population at their homes. Physiological, motor and even environmental measurements are excellent indicators of their health status. Furthermore, connected wearable technologies in the framework of the Internet of Things allow for risk situations prevention. Recent advances of Artificial Intelligence techniques make possible high-accuracy decision making on multi-sensory data. Nevertheless, to train models and to perform online real-time detection of risk events from heterogeneous taxonomy, robust wearable devices are required. In this paper, existing solutions are reviewed for human sensing for these purposes. Moreover, we present the implementation of a multi-sensor device for the recognition of risk situations with a focus on the data synchronisation. Thinhinane Yebda, Jenny Benois-Pineau, Hélène Amieva, Benjamin Frolicher |
CBMI | 2 |
| 2019 | FPGA-based SIFT implementation for wearable computingabstractThe article describes the first steps to achieve control over a robotic or prosthetic arm based on analysis of visual environment acquired in real-time by video cameras on glasses and on the prosthesis. One of the main goals of the research is to develop a wearable, portable, lightweight, and low power consumption device for visual scene analysis. This paper will discuss the critical steps of its implementation on an FPGA board. We implemented some time-consuming parts of the SFT algorithm needed for the analysis in C/C++ language on TUL PYNQ-Z2 FPGA board. This implementation allows for a low power consumption of the programmable logic part of the system. The obtained value is 0. 274W. Processing capacity is 96.45 images per second on a small wearable size device which allow for the real-time implementation of the whole analysis in the future. Attila Fejér, Zoltán Nagy 0001, Jenny Benois-Pineau, Péter Szolgay, Aymar de Rugy, Jean-Philippe Domenger |
DDECS | 3 |
| 2019 | Fine-Grained Action Detection and Classification in Table Tennis with Siamese Spatio-Temporal Convolutional Neural NetworkabstractHuman action recognition in videos is one of the key problems in visual data interpretation. Despite intensive research, the recognition of actions with low inter-class variability remains a challenge. To answer this problem, my thesis focus on fine-grained classification challenge using a Siamese Spatio-Temporal Convolutional Neural Network and apply it to a new dataset we have introduced TTStroke-21. Our model take as input data RGB images and Optical Flow and is able to reach an accuracy of 91.4% against 43.1% for our baseline on temporal segmented videos. Detection and classification in videos using a sliding temporal window leads to a score of 81.3% over the whole dataset. Pierre-Etienne Martin, Jenny Benois-Pineau, Renaud Péteri |
ICIP | 2 |
| 2019 | Optimal Choice of Motion Estimation Methods for Fine-Grained Action Classification with 3D Convolutional NetworksabstractDetecting and classifying human actions in videos is one of the current challenges in visual content analysis and mining. This paper presents a method for performing a finegrained classification of sport actions using a Siamese SpatioTemporal Convolutional Neural Network (SSTCNN) model. This model takes RGB images and Optical Flow field as input data. Our first contribution is the comparison of different Optical flow methods and a study of their influence on the classification score. We also present different normalization methods for the optical flow that drastically impact results, boosting performances from 44% to 74% of accuracy. Our second contribution is the detection and classification of actions in videos performed using a sliding temporal window. It leads to a satisfying score of 81.3% over the whole dataset TTStroke-21. Pierre-Etienne Martin, Jenny Benois-Pineau, Renaud Péteri, Julien Morlier |
ICIP | 2 |
| 2019 | ChaboNet : Design of a deep CNN for prediction of visual saliency in natural video
Souad Chaabouni, Jenny Benois-Pineau, Chokri Ben Amar |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | Saliency-based selection of visual content for deep convolutional neural networks - Application to architectural style classification
Abraham Montoya Obeso, Jenny Benois-Pineau, Mireya S. García-Vázquez, Alejandro Alvaro Ramírez-Acosta |
Multim. Tools Appl. | 2 |
| 2019 | Perceptually-guided deep neural networks for ego-action prediction: Object grasping
Iván González-Díaz 0001, Jenny Benois-Pineau, Jean-Philippe Domenger, Daniel Cattaert, Aymar de Rugy |
Pattern Recognit. | 2 |
| 2018 | Sport Action Recognition with Siamese Spatio-Temporal CNNs: Application to Table TennisabstractHuman action recognition in video is one of the key problems in visual data interpretation. Despite intensive research, the recognition of actions with low inter-class variability remains a challenge. This paper presents a new Siamese Spatio-Temporal Convolutional neural network (SSTC) for this purpose. When applied to table tennis, it is possible to detect and recognize 20 table tennis strokes. The model has been trained on a specific dataset, TTStroke-21, recorded in natural condition (markerless) at the Faculty of Sports of the University of Bordeaux. Our model takes as inputs a RGB image sequence and its computed Optical Flow. After 3 spatio-temporal convolutions, data are fused in a fully connected layer of a proposed siamese network architecture. Our method reaches an accuracy of 91.4% against 43.1% for our baseline. Pierre-Etienne Martin, Jenny Benois-Pineau, Renaud Péteri, Julien Morlier |
CBMI | 2 |
| 2018 | Introduction of Explicit Visual Saliency in Training of Deep CNNs: Application to Architectural Styles ClassificationabstractIntroduction of visual saliency or interestingness in the content selection for image classification tasks is an intensively researched topic. It has been namely fulfilled for feature selection in feature-based methods. Nowadays, in the winner classifiers of visual content such as Deep Convolutional Neural Networks, visual saliency maps have not been introduced explicitly. Pooling features in CNNs is known as a good strategy to reduce data dimensionality, computational complexity and summarize representative features for subsequent layers. In this paper we introduce visual saliency in network pooling layers to spatially filter relevant features for deeper layers. Our experiments are conducted in a specific task to identify Mexican architectural styles. The results are promising: proposed approach reduces model loss and training time keeping the same accuracy as the base-line CNN. Abraham Montoya Obeso, Jenny Benois-Pineau, Mireya S. García-Vázquez, Alejandro Alvaro Ramírez-Acosta |
CBMI | 2 |
| 2018 | Classification of Alzheimer Disease on Imaging Modalities with Deep CNNs Using Cross-Modal Transfer LearningabstractA recent imaging modality Diffusion Tensor Imaging completes information used from Structural MRI in studies of Alzheimer disease. A large number of recent studies has explored pathologic staging of Alzheimer disease using the Mean Diffusivity maps extracted from the Diffusion Tensor Imaging modality. The Deep Neural Networks are seducing tools for classification of subjects' imaging data in computer-aided diagnosis of Alzheimer's disease. The major problem here is the lack of a publicly available large amount of training data in both modalities. The lack number of training data yields over-fitting phenomena. We propose a method of a cross-modal transfer learning: from Structural MRI to Diffusion Tensor Imaging modality. Models pre-trained on a structural MRI dataset with domain-depended data augmentation are used as initialization of network parameters to train on Mean Diffusivity data. The method shows a reduction of the over-fitting phenomena, improves learning performance, and thus increases the accuracy of prediction. Classifiers are then fused by a majority vote resulting in augmented scores of classification between Normal Control, Alzheimer Patients and Mild Cognitive Impairment subjects on a subset of ADNI dataset. Karim Aderghal, Alexander V. Khvostikov, Andrey S. Krylov, Jenny Benois-Pineau, Karim Afdel, Gwénaëlle Catheline |
CBMS | 4 |
| 2018 | Early and Late Fusion of Temporal Information for Classification of Surgical Actions in Laparoscopic GynecologyabstractThe most essential step towards semiautomatic extraction of relevant surgery scenes is semantic understanding of surgical actions in surgery videos. Currently, Convolutional Neural Networks (CNNs) are a de-facto standard for automatic content classification in many domain, including medical imaging. We aim to include increase the predictive performance of surgical action recognition within gynecologic laparoscopy, a subfield of endoscopic surgery, by fusing temporal information to the input layer of CNNs (early fusion), as well as temporal aggregation of single-frame prediction results (late fusion). Our evaluation shows that the proposed early fusion approaches are able to outperform a single-frame baseline when using the GoogLeNet architecture. Moreover, early fusion of motion information benefits the classification performance regardless of late fusion strategy. Late fusion has a high impact on classification performance, and its increase is additive to the performance increase of early fusion. Eventually, we found that the CNN capacity influences these results drastically. We conclude that the proposed methods in combination with a sufficiently high CNN capacity allow for a substantial increase in predictive performance.q Stefan Petscharnig, Klaus Schöffmann, Jenny Benois-Pineau, Souad Chaabouni, Jörg Keckstein |
CBMS | 3 |
| 2018 | Increasing Training Stability for Deep CNNSabstractIn the present work, we investigate the possibility to expand existing Deep Learning solvers in order to improve the quality and stability of the training. We propose different new solvers, all of them based on the filtering of the neural network parameters, and experimentally prove that properly tuning their respective hyper-parameters leads to a clear improvement of the training and validation results, in quality and in stability. Pierre Gillot, Jenny Benois-Pineau, Akka Zemmari, Yurii E. Nesterov |
ICIP | 2 |
| 2018 | Perceptually-guided Understanding of Egocentric Video Content: Recognition of Objects to GraspabstractIncorporating user perception into visual content search and understanding tasks has become one of the major trends in multimedia retrieval. We tackle the problem of object recognition guided by user perception, as indicated by his gaze during visual exploration, in the application domain of assistance to upper-limb amputees. Although selecting the object to be grasped represents a task-driven visual search, human gaze recordings are noisy due to several physiological factors. Hence, since gaze does not always point to the object of interest, we use video-level weak annotations indicating the object to be grasped, and propose a video-level weak loss in classification with Deep CNNs. Our results show that the method achieves notably better performance than other approaches over a complex real-life dataset specifically recorded, with optimal performance for fixation times around 400-800ms, producing a minimal impact on subjects' behavior. Iván González-Díaz 0001, Jenny Benois-Pineau, Jean-Philippe Domenger, Aymar de Rugy |
ICMR | 2 |
| 2018 | Multi-modal activity recognition from egocentric vision, semantic enrichment and lifelogging applications for the care of dementia
Georgios Meditskos, Pierre-Marie Plans, Thanos G. Stavropoulos, Jenny Benois-Pineau, Vincent Buso, Ioannis Kompatsiaris |
J. Vis. Commun. Image Represent. | 4 |
| 2017 | Classification of sMRI for Alzheimer's disease Diagnosis with CNN: Single Siamese Networks with 2D+? Approach and Fusion on ADNIabstractThe methods of Content-Based visual information indexing and retrieval penetrate into Healthcare and become popular in Computer-Aided Diagnostics. The PhD research we have started 13 months ago is devoted to the multimodal classification of MRI brain scans for Alzheimer Disease diagnostics. We use the winner classifier, such as CNN. We first proposed an original 2D+ approach. It avoids heavy volumetric computations and uses domain knowledge on Alzheimer biomarkers. We study discriminative power of different brain projections. Three binary classification tasks are considered separating Alzheimer Disease (AD) patients from Mild Cognitive Impairment (MCI) and Normal Control subject (NC). Two fusion methods on FC layer and on the single-projection CNN output show better performances, up to 91% of accuracy is achieved. The results are competitive with the SOA which uses heavier algorithmic chain. Karim Aderghal, Jenny Benois-Pineau, Karim Afdel |
ICMR | 2 |
| 2017 | Classification of sMRI for AD Diagnosis with Convolutional Neuronal Networks: A Pilot 2-D+ \epsilon Study on ADNI
Karim Aderghal, Manuel Boissenin, Jenny Benois-Pineau, Gwénaëlle Catheline, Karim Afdel |
MMM (1) | 3 |
| 2017 | Saliency Driven Object recognition in egocentric videos with deep CNN: toward application in assistance to Neuroprostheses
Philippe Pérez de San Roman, Jenny Benois-Pineau, Jean-Philippe Domenger, Florent Paclet, Daniel Cattaert, Aymar de Rugy |
Comput. Vis. Image Underst. | 2 |
| 2017 | Recognition of Alzheimer's disease and Mild Cognitive Impairment with multimodal image-derived biomarkers and Multiple Kernel Learning
Olfa Ben Ahmed, Jenny Benois-Pineau, Michèle Allard, Gwénaëlle Catheline, Chokri Ben Amar |
Neurocomputing | 2 |
| 2017 | Prediction of visual attention with deep CNN on artificially degraded videos for studies of attention of patients with Dementia
Souad Chaabouni, Jenny Benois-Pineau, Francois Tison, Chokri Ben Amar, Akka Zemmari |
Multim. Tools Appl. | 2 |
| 2017 | Segmentation of left ventricle on dynamic MRI sequences for blood flow cancellation in Thermotherapy
Samah Bouzidi, Aurelie Emilien, Jenny Benois-Pineau, Bruno Quesson, Chokri Ben Amar, Pascal Desbarats |
Signal Process. Image Commun. | 3 |
| 2016 | Transfer learning with deep networks for saliency prediction in natural videoabstractThe main purpose of transfer learning is to resolve the problem of different data distribution, generally, when the training samples of source domain are different from the training samples of the target domain. Prediction of salient areas in natural video suffers from the lack of large video benchmarks with human gaze fixations. Different databases only provide dozens up to one or two hundred of videos. The only public large database is HOLLYWOOD with 1707 videos available with gaze recordings. The main idea of this paper is to transfer the knowledge learned with the deep network on a large dataset to train the network on a small dataset to predict salient areas. The results show an improvement on two small publicly available video datasets. Souad Chaabouni, Jenny Benois-Pineau, Chokri Ben Amar |
ICIP | 2 |
| 2016 | A scalable summary generation method based on cross-modal consensus clustering and OLAP cube modeling
Gabriel Sargent, Karina Perez-Daniel, Andrei Stoian, Jenny Benois-Pineau, Sofian Maabout, Henri Nicolas, Mariko Nakano-Miyatake, Jean Carrive |
Multim. Tools Appl. | 4 |
| 2016 | Semantic Event Fusion of Different Visual Modality Concepts for Activity RecognitionabstractCombining multimodal concept streams from heterogeneous sensors is a problem superficially explored for activity recognition. Most studies explore simple sensors in nearly perfect conditions, where temporal synchronization is guaranteed. Sophisticated fusion schemes adopt problem-specific graphical representations of events that are generally deeply linked with their training data and focused on a single sensor. This paper proposes a hybrid framework between knowledge-driven and probabilistic-driven methods for event representation and recognition. It separates semantic modeling from raw sensor data by using an intermediate semantic representation, namely concepts. It introduces an algorithm for sensor alignment that uses concept similarity as a surrogate for the inaccurate temporal information of real life scenarios. Finally, it proposes the combined use of an ontology language, to overcome the rigidity of previous approaches at model definition, and a probabilistic interpretation for ontological models, which equips the framework with a mechanism to handle noisy and ambiguous concept observations, an ability that most knowledge-driven methods lack. We evaluate our contributions in multimodal recordings of elderly people carrying out IADLs. Results demonstrated that the proposed framework outperforms baseline methods both in event recognition performance and in delimiting the temporal boundaries of event instances. Carlos Fernando Crispim, Vincent Buso, Konstantinos Avgerinakis, Georgios Meditskos, Alexia Briassouli, Jenny Benois-Pineau, Ioannis Kompatsiaris, François Brémond |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2016 | Perceptual modeling in the problem of active object recognition in visual scenes
Iván González-Díaz 0001, Vincent Buso, Jenny Benois-Pineau |
Pattern Recognit. | 3 |
| 2016 | Fast Action Localization in Large-Scale Video ArchivesabstractFinding content in large video archives has so far required textual annotation to enable search by keywords. Our aim is to support retrieval from such archives using queries based on the example video clips that contain meaningful human actions. We propose a solution for the scalable search of actions in large-scale archives by leveraging the complementarity between the description at the frame level and the aggregation in time of descriptors. To permit fast search, we introduce a two-level cascade. The inexpensive first level employs aggregation to filter out a large part of the video. At the second level, aided by feature selection, a more discriminative comparison by frame alignment ranks the remaining video sequences. We improve upon the state of the art on popular data sets, and we introduce and show the results on a novel video archive data set that is significantly larger than previous ones. Andrei Stoian, Marin Ferecatu, Jenny Benois-Pineau, Michel Crucianu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Features-based approach for Alzheimer's disease diagnosis using visual pattern of water diffusion in tensor diffusion imagingabstractIn this paper, we propose a feature-based classification framework for Alzheimer's disease (AD) recognition using Tensor Diffusion Imaging (DTI). The main contribution consists in considering the visual pattern of water molecules diffusion in the most involved region in AD (hippocampal area). We use the Circular Harmonic Functions (CHFs) and the Bag-of-Visual-Words approach to build an AD related-signature. The experiments were accomplished first with a subset of participants from the Alzheimer's Disease Neuroimaging Initiative (ADNI) dataset and then with the DTI scans of a French epidemiological study: ”Bordeaux-3City”. Experimental results demonstrate that our features-based method applied on the MD maps is able to capture the AD-related atrophy and then classify between AD subjects. Olfa Ben Ahmed, Jenny Benois-Pineau, Chokri Ben Amar, Michele Aliara, Gwénaëlle Catheline |
ICIP | 2 |
| 2015 | Object recognition with top-down visual attention modeling for behavioral studiesabstractBehavioural analysis in instrumental activities of daily living has become a powerful tool in clinical studies and rises the question of what objects are manipulated by patients. In this paper we present a top-down probabilistic visual attention model for manipulated object recognition in egocentric video content. Although arms often occlude objects and are usually seen as a burden for many vision systems, they become an asset in our approach, as we extract both global and local features describing their geometric layout and pose, as well as the objects being manipulated. We integrate this information in a probabilistic generative model, provide update equations that automatically compute the model parameters optimizing the likelihood of the data, and design a method to generate maps of visual attention that are later used in an object-recognition framework. This task-driven assessment reveals that the proposed method outperforms the state of the art in object recognition for egocentric video content. Vincent Buso, Iván González-Díaz 0001, Jenny Benois-Pineau |
ICIP | 3 |
| 2015 | Scalable action localization with kernel-space hashingabstractTo detect and locate complex human actions in video, one trains a detector for each target class and applies it to the video content. This approach can scale to large video databases if the application of the detector can be made sublinear in the size of the database. Sublin-ear retrieval methods have been successfully explored for query-by-example but few were devised for these more challenging queries by detector. We put forward here a novel approximate search method that relies on LSH to support query-by-detector. We evaluate our method on a recent large action localization dataset and show it has significantly better efficiency than linear search. Andrei Stoian, Marin Ferecatu, Jenny Benois-Pineau, Michel Crucianu |
ICIP | 3 |
| 2015 | Classification of Alzheimer's disease subjects from MRI using hippocampal visual features
Olfa Ben Ahmed, Jenny Benois-Pineau, Michèle Allard, Chokri Ben Amar, Gwénaëlle Catheline |
Multim. Tools Appl. | 2 |
| 2015 | Geometrical cues in visual saliency models for active object recognition in egocentric videos
Vincent Buso, Jenny Benois-Pineau, Jean-Philippe Domenger |
Multim. Tools Appl. | 2 |
| 2015 | Guest editorial: Content-Based Multimedia Indexing
Klaus Schöffmann, Jenny Benois-Pineau, Bernard Mérialdo, Tamás Szirányi |
Multim. Tools Appl. | 2 |
| 2015 | Goal-oriented top-down probabilistic visual attention model for recognition of manipulated objects in egocentric videos
Vincent Buso, Iván González-Díaz 0001, Jenny Benois-Pineau |
Signal Process. Image Commun. | 3 |
| 2014 | A Comparative Study of Irregular Pyramid Matching in Bag-of-Bags of Words Model for Image Retrieval
Jenny Benois-Pineau, Aurélie Bugeau |
ICISP | 2 |
| 2014 | Hierarchical Hidden Markov Model in detecting activities of daily living in wearable videos for studies of dementia
Svebor Karaman, Jenny Benois-Pineau, Vladislavs Dovgalecs, Rémi Mégret, Julien Pinquier, Régine André-Obrecht, Yann Gaëstel, Jean-François Dartigues |
Multim. Tools Appl. | 2 |
| 2014 | Preface to the special issue on Content-Based Multimedia Indexing
Stéphane Marchand-Maillet, Patrick Lambert, Bernard Mérialdo, Jenny Benois-Pineau |
Multim. Tools Appl. | 4 |
| 2013 | Adaptive rejection of outliers for robust motion compensation in cardiac MR-thermometryabstractNew Magnetic Resonance (MR) imaging applications include real time monitoring of temperature changes during cardiac radiofrequency ablations. MR-thermometry requires online robust motion compensation to cope with the complex motion of the heart resulting from respiratory activity and cardiac contraction (potentially in presence of arrhythmia), together with the presence of noise in MR images. We propose a method to adaptively and automatically tune parameters of motion compensation algorithms that use robustness function. The core of the method is the estimation of the probability density function (pdf) of the error for each pixel in a reference frame using the Rician noise pdf model in MRI. Then parameter map is derived from estimated pdf. The proposed method leads to better results than using a fixed control parameter of the robustness function, which should facilitate the use of such methods for clinical purpose. Aurelie Emilien, Jenny Benois-Pineau, Delphine Elbes, Bruno Quesson |
ICIP | 2 |
| 2013 | ACM MM MIIRH 2013: workshop on multimedia indexing and information retrieval for healthcareabstractHealthcare systems are depending on increasingly sophisticated and ubiquitous technology, while telehealth is rapidly gaining importance with the advent of low-cost and effective technological solutions in medicine. The increase in the worldwide elderly population and the burden this is inflicting upon the workforce, societies and economies are making remote care and independent living at home a necessity. MIIRH is the first workshop on multimedia analysis for remote care of and assisted living solutions which enable people that are incapacitated in some regard to continue living independently at home and remain active members of society. The topics addressed in MIIRH are extremely timely, as multitudes of cost-effective and high quality care solutions are already being developed and used, rendering the examination of new medical, healthcare paradigms an absolute necessity. Jenny Benois-Pineau, Alexia Briassouli, Alex Hauptmann 0001 |
ACM Multimedia | 1 |
| 2013 | Preface for the special issue of MTAP following CBMI 2011
José María Martínez Sanchez, Bernard Mérialdo, Jenny Benois-Pineau, Joemon M. Jose |
Multim. Tools Appl. | 3 |
| 2012 | Feature-based brain MRI retrieval for Alzheimer disease diagnosisabstractIn this paper we consider the application of the feature-based approach to medical image retrieval, particularly brain MRI scans for early Alzheimer's disease diagnosis. The key idea is to provide the doctor with the images which have similar visual properties and have full case record, giving the ability to make more informed decision in the prodromal phase of the disease. With regard to the state-of-the art SIFT features in a Bag-of-Visual-Words approach we propose to use the Laguerre Circular Harmonic Functions coefficients as feature vectors. An additional pre-classification step based on estimation of Alzheimer's disease early image abnormalities is proposed to improve overall precision. Maxim M. Mizotin, Jenny Benois-Pineau, Michèle Allard, Gwénaëlle Catheline |
ICIP | 2 |
| 2012 | Strategies for multiple feature fusion with Hierarchical HMM: Application to activity recognition from wearable audiovisual sensors
Julien Pinquier, Svebor Karaman, Laetitia Letoupin, Patrice Guyot, Rémi Mégret, Jenny Benois-Pineau, Yann Gaëstel, Jean-François Dartigues |
ICPR | 6 |
| 2012 | Multi-layer Local Graph Words for Object Recognition
Svebor Karaman, Jenny Benois-Pineau, Rémi Mégret, Aurélie Bugeau |
MMM | 2 |
| 2012 | Content Based Image Retrieval Using Bag-Of-Regions
Rémi Vieux, Jenny Benois-Pineau, Jean-Philippe Domenger |
MMM | 2 |
| 2012 | Preface for the special issue of MTAP following CBMI 2010
Georges Quénot, Jenny Benois-Pineau, Régine André-Obrecht |
Multim. Tools Appl. | 2 |
| 2012 | Segmentation-based multi-class semantic object detection
Rémi Vieux, Jenny Benois-Pineau, Jean-Philippe Domenger, Achille J.-P. Braquelaire |
Multim. Tools Appl. | 2 |
| 2012 | Robust Real-Time-Constrained Estimation of Respiratory Motion for Interventional MRI on Mobile OrgansabstractReal-time magnetic resonance imaging is a promising tool for image-guided interventions. For applications such as thermotherapy on moving organs, a precise image-based compensation of motion is required in real time to allow quantitative analysis, retrocontrol of the interventional device, or determination of the therapy endpoint. Reduced field-of-view imaging represents a promising way to improve spatial and/or temporal resolution. However, it introduces new challenges for target motion estimation, since structures near the target may appear transiently due to the respiratory motion and the limited spatial coverage. In this paper, a new image-based motion estimation method is proposed combining a global motion estimation with a novel optical flow approach extending the initial Horn and Schunck (H&S) method by an additional regularization term. This term integrates the displacement of physiological landmarks into the variational formulation of the optical flow problem. This allowed for a better control of the optical flow in presence of transient structures. The method was compared to the same registration pipeline employing the H&S approach on a synthetic dataset and in vivo image sequences. Compared to the H&S approach, a significant improvement (p<0.05) of the Dice's similarity criterion computed between the reference and the registered organ positions was achieved. Sébastien Roujol, Jenny Benois-Pineau, Baudouin Denis de Senneville, Mario Ries, Bruno Quesson, Chrit T. W. Moonen |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2011 | A metric for no-reference video quality assessment for HD TV delivery based on saliency mapsabstractThis paper contributes to objective video quality assessment of broadcasted HDTV content without reference. In this context we present a new No-Reference video quality metric taking into account the behavior of Human Visual System. This new metric called WMBER is based on macro-blocks error detection weighted by saliency maps computed at the decoder side. Moreover, both macro-block error detection and saliency maps processing require only partial decoding allowing real-time performance. A subjective experiment has been carried out to evaluate the performances of the proposed metric. The results are compared to the Full Reference metric MSE. The evaluation of the results shows that the proposed method provides a very good prediction of subjective measures. Hugo Boujut, Jenny Benois-Pineau, Toufik Ahmed, Ofer Hadar, Patrick Bonnet |
ICME | 2 |
| 2011 | Detection of moving foreground objects in videos with strong camera motion
Dániel Szolgay, Jenny Benois-Pineau, Rémi Mégret, Yann Gaëstel, Jean-François Dartigues |
Pattern Anal. Appl. | 2 |
| 2011 | Local object-based super-resolution mosaicing from low-resolution video
Petra Gomez-Krämer, Jenny Benois-Pineau, Jean-Philippe Domenger |
Signal Process. | 2 |
| 2010 | Real time constrained motion estimation for ECG-gated cardiac MRIabstractContinuous magnetic resonance (MR) imaging is a promising tool for image-guided cardiac interventions. However, cardiac imaging is complicated by both, cardiac and respiratory motion. While the former is commonly addressed by cardiac gating, the latter requires a real time motion compensation. Here, a new image based motion estimation is proposed, extending the initial Horn & Schunck optical flow approach by an additional regularization term. This term allows to integrate displacement of physiological landmarks, which are obtained in a preparation step using pattern matching. The proposed method was evaluated on the left ventricle (LV) of cardiac MR-images and showed a better estimation accuracy compared to the original Horn & Schunck approach. Sébastien Roujol, Jenny Benois-Pineau, Baudouin Denis de Senneville, Bruno Quesson, Mario Ries, Chrit T. W. Moonen |
ICIP | 2 |
| 2010 | Human Daily Activities Indexing in Videos from Wearable Cameras for Monitoring of Patients with Dementia DiseasesabstractOur research focuses on analysing human activities according to a known behaviorist scenario, in case of noisy and high dimensional collected data. The data come from the monitoring of patients with dementia diseases by wearable cameras. We define a structural model of video recordings based on a Hidden Markov Model. New spatio-temporal features, color features and localization features are proposed as observations. First results in recognition of activities are promising. Svebor Karaman, Jenny Benois-Pineau, Rémi Mégret, Vladislavs Dovgalecs, Jean-François Dartigues, Yann Gaëstel |
ICPR | 2 |
| 2010 | The IMMED project: wearable video monitoring of people with age dementiaabstractIn this paper, we describe a new application for multimedia indexing, using a system that monitors the instrumental activities of daily living to assess the cognitive decline caused by dementia. The system is composed of a wearable camera device designed to capture audio and video data of the instrumental activities of a patient, which is leveraged with multimedia indexing techniques in order to allow medical specialists to analyze several hour long observation shots efficiently. Rémi Mégret, Vladislavs Dovgalecs, Hazem Wannous, Svebor Karaman, Jenny Benois-Pineau, Elie Khoury 0001, Julien Pinquier, Philippe Joly, Régine André-Obrecht, Yann Gaëstel, Jean-François Dartigues |
ACM Multimedia | 5 |
| 2010 | Significance Delta Reasoning with p-Adic Neural Networks: Application to Shot Change Detection in VideoabstractA new possibility to extract significant changes is given by the so-called p-adic number system. p-Adic metric as opposed to a conventional Euclidean metric is based on hierarchical encoding of information. p-Adic and more generally ultrametric spaces have been already used for modeling the functioning of cognitive systems for their property of inducing hierarchy in decision-making process. In this paper, we benefit from this property in modeling delta reasoning. Indeed p-adics give a possibility to extract significant information, in particular significant delta-changes and furthermore, to preserve this hierarchical structure in the process of performing various operations on the data. Hence, we can create p-adic neural networks operating on hierarchical strings of information. In this paper, an algorithm of random learning of p-adic neural networks is applied to the problem of detection of changes in streams of video information such as shot boundaries. Jenny Benois-Pineau, Andrei Yu. Khrennikov |
Comput. J. | 1 |
| 2010 | Scalable object-based video retrieval in HD video databases
Claire Morand, Jenny Benois-Pineau, Jean-Philippe Domenger, Joaquin Zepeda, Ewa Kijak, Christine Guillemot |
Signal Process. Image Commun. | 2 |
| 2009 | A multi-resolution particle filter tracking in a multi-camera environmentabstractThis paper presents a novel tracking method with the multi-resolution technique and a Kolmogrov-Smirnov test for model update to track a non-rigid target in an uncalibrated multi-camera environment. It is based on particle filter method using color appearance model. Compared to the related work, our method improves the tracking performance by proposing: i) a multi-resolution technique to rapidly locate the estimate of the target state and refine it gradually, ii) the Kolmogrov-Smirnov test to evaluate the reliability of the estimate so as to take the decision on further updating/ reinitialization of the estimate, as well as iii) an interaction of cameras approach to reinitialize the estimate by information detected in other cameras in case of tracking failures. After being tested in a multi-camera environment for one person tracking, our system is shown to give a better tracking result in comparison with mono-camera tracking, especially when occlusions occur. Henri Nicolas, Jenny Benois-Pineau |
ICIP | 3 |
| 2008 | HD motion estimation in a wavelet pyramid in JPEG2000 contextabstractGlobal motion estimation (GME) is a step of primary importance in video analysis and Indexing. In this paper, we propose a method for estimating the global motion in HD sequences compressed with JPE2000, as it is the compression standard adopted by the digital cinema. To respect the scalability nature of the standard, our method works preferentially on the lower resolution sub-bands of the wavelet compressed frames. We combine a hierarchical block matching (BM) with a robust GME. This not only allows a regularization of the background motion by approximating the BM estimated motion vectors with the model found, but also indicates where motion has to be estimated with a better precision, typically on the foreground objects. Claire Morand, Jenny Benois-Pineau, Jean-Philippe Domenger |
ICIP | 2 |
| 2007 | PCA-Based Image Registration : Application to On-Line MR Temperature Monitoring of Moving TissuesabstractReal-time magnetic resonance (MR) thermometry provides continuous temperature mapping inside the human body and is therefore a promising tool to monitor and control interventional therapies based on thermal ablation. Temperature information must be mapped to a reference position of observed organs in order to allow thermal dose computation, as the history of temperature is required for each pixel. Motion compensated MR-thermometry for thermotherapy has to cope with radio-frequency (RF) artifacts and relaxation-time changes of the monitored tissue. While purely optical-flow-based realignment may lead to temperature map computation errors for the case of local or global intensity changes, principal component analysis based realignment results in accurately registered temperature maps. The motion estimation process described in this paper consists of two steps : a parameterized flow models is initially computed using a principal component analysis during a preparative learning step; during the intervention, motion is characterized with a small set of parameters using a least square solver. Gregory Maclair, Baudouin Denis de Senneville, Mario Ries, Bruno Quesson, Pascal Desbarats, Jenny Benois-Pineau, Chrit T. W. Moonen |
ICIP (3) | 6 |
| 2007 | PCA-Based Magnetic Field Modeling : Application for On-Line MR Temperature Monitoring
Gregory Maclair, Baudouin Denis de Senneville, Mario Ries, Bruno Quesson, Pascal Desbarats, Jenny Benois-Pineau, Chrit T. W. Moonen |
MICCAI (2) | 6 |
| 2007 | Retrieval of objects in video by similarity based on graph matching
Fanny Chevalier, Jean-Philippe Domenger, Jenny Benois-Pineau, Maylis Delest |
Pattern Recognit. Lett. | 3 |
| 2007 | Signal processing: Image communication, special issue on content-based multimedia indexing and retrieval
Ebroul Izquierdo, Jenny Benois-Pineau, Régine André-Obrecht |
Signal Process. Image Commun. | 2 |
| 2007 | The ARGOS campaign: Evaluation of video analysis and indexing tools
Philippe Joly, Jenny Benois-Pineau, Ewa Kijak, Georges Quénot |
Signal Process. Image Commun. | 2 |
| 2007 | Super-resolution mosaicing from MPEG compressed video
Petra Gomez-Krämer, Ofer Hadar, Jenny Benois-Pineau, Jean-Philippe Domenger |
Signal Process. Image Commun. | 3 |
| 2006 | Knowledge-Based Supervised Learning Methods in a Classical Problem of Video Object TrackingabstractIn this paper we present a new scheme for detection and tracking of specific objects in a knowledge-based framework. The scheme uses a supervised learning method: support vector machines. Both problems, detection and tracking, are solved by a common approach: objects are located in video sequences by a SVM classifier. They are next tracked along the time by a SVM tracker with complete 6 parameters affine model. The method is applied in a video surveillance application for detection and tracking of frontal view faces. Real time application constraints are met by reduction of support vector set. Lionel Carminati, Jenny Benois-Pineau, Christian Jennewein |
ICIP | 2 |
| 2006 | Use of Motion Information in Super-Resolution MosaicingabstractIn this paper, we present a super-resolution (SR) method based on iterative backprojections. Motion information is used for the synthesis of the restoration filter in the SR method. Both, the blur estimation and the choice of the degradation model, are based on estimated global motion. The method is applied to highly under-sampled images such as DC images of MPEG compressed video and medical magnetic resonance (MR) images. Results of the comparison of our blur model with some common blur models are encouraging. Petra Gomez-Krämer, Ofer Hadar, Jenny Benois-Pineau, Jean-Philippe Domenger |
ICIP | 3 |
| 2006 | Scene similarity measure for video content segmentation in the framework of a rough indexing paradigmabstractThis article presents a scene similarity measure for video content segmentation. In the context of the rough indexing paradigm, we extract only partial information from MPEG compressed streams to measure the similarity of video frames through time. The similarity measure of I-Frames is defined based on motion compensation of DC images and local contrast computation. The method allows a real-time segmentation of the video content. © 2006 Wiley Periodicals, Inc. Int J Int Syst 21: 765–783, 2006. Petra Gomez-Krämer, Jenny Benois-Pineau, Jean-Philippe Domenger |
Int. J. Intell. Syst. | 2 |
| 2006 | DAG-based visual interfaces for navigation in indexed video content
Maylis Delest, Anthony Don, Jenny Benois-Pineau |
Multim. Tools Appl. | 3 |
| 2005 | Gaussian mixture classification for moving object detection in video surveillance environmentabstractThe paper deals with detection of moving objects by modelling pixel grey level distribution along the time. The detection of moving objects is based on learning and update of background pixel distributions. The choice of appropriate mixture's component for a given pixel is performed by likelihood maximization. An original Markov regularization is proposed to smooth detection. The method performs in real time on CIF resolution video and low cost commercial hardware. Lionel Carminati, Jenny Benois-Pineau |
ICIP (3) | 2 |
| 2005 | Super-resolution mosaicing from MPEG compressed videoabstractIn this paper, we describe a method for the construction of super-resolution (SR) mosaics. The low-resolution (LR) input sequence of the SR algorithm consists of DC images of I-frames (DCI-frames) extracted from MPEG compressed streams. For the registration of the LR sequence, the motion information from P-frames is used. The novelty of this approach is the determination of the optical transfer function (OTF) of the blur for the image restoration in the SR algorithm. First results are promising. Petra Gomez-Krämer, Ofer Hadar, Jenny Benois-Pineau, Jean-Philippe Domenger |
ICIP (1) | 3 |
| 2005 | Comparison of shot boundary detectorsabstractA video cut detector (CD), a member of the shot boundary detector (SBD) group, is an essential element for spatio-temporal audiovisual (AV) segmentation and various video-processing technologies. Platform, processing and performance constraints forced the development of various dedicated CDs. Future platforms allow the usage of advanced CD algorithms with higher reliability. In order to enable an appropriate trade-off decision to be made between reliability and the required processing power, benchmarking of four CD algorithms has taken place on bases of a generic, culture-diverse multi-genre AV corpus. In terms of complexity/performance trade-off, a field-difference-based CD proved to be optimal. Jan Nesvadba, Fabian Ernst, Jernej Perhavc, Jenny Benois-Pineau, Laurent Primaux |
ICME | 4 |
| 2005 | Real-Time and Distributed AV Content Analysis System for Consumer Electronics NetworksabstractThe ever-increasing complexity of generic multimedia-content-analysis-based (MCA) solutions, their processing power demanding nature and the need to prototype and assess solutions in a fast and cost-saving manner motivated the development of the Cassandra framework. The combination of state-of-the-art network and grid-computing solutions and recently standardized interfaces facilitated the set-up of this framework, forming the basis for multiple cross-domain and cross-organizational collaborations. It enables distributed computing scenario simulations for e.g. distributed content analysis (DCA) across consumer electronics (CE) in-home networks, but also the rapid development and assessment of complex multi-MCA-algorithm-based applications and system solutions. Furthermore, the framework's modular nature-logical MCA units are wrapped into so-called service units (SU)-ease the split between system-architecture- and algorithmic-related work and additionally facilitate reusability, extensibility and upgrade ability of those SUs Jan Nesvadba, Pedro Fonseca 0002, Alexander Sinitsyn, Fons de Lange, Martijn Thijssen, Patrick van Kaam, Hong Liu 0008, Rien van Leeuwen, Johan J. Lukkien, Andrei Korostelev, Jan Ypma, Bart Kroon, Hasan Celik, Alan Hanjalic, Suphi Umut Naci, Jenny Benois-Pineau, Peter H. N. de With, Jungong Han |
ICME | 16 |
| 2004 | Grouping video shots into scenes based on 1D mosaic descriptorsabstractThis paper describes an original approach for structuring video documents into scenes by grouping video shots. The method is based on the construction of 1D mosaics. 1D mosaics are built based on X-ray projections of color video frames representing integration along vertical and horizontal axes. The mosaicing is realized by motion compensation in a 1D domain. Grouping of shots in a scene is done by local and global matching of mosaics, based on piecewise linear approximation and hierarchical clustering. The results obtained on feature documentaries are promising. Henri Nicolas, Anne Manoury, Jenny Benois-Pineau, William Dupuis, Dominique Barba |
ICIP | 3 |
| 2003 | Human detection and tracking for video surveillance applications in a low-density environment
Lionel Carminati, Jenny Benois-Pineau, Marc Gelgon |
VCIP | 2 |
| 2002 | Full scheme of MPEG4-like codec based on wavelet transformabstractWe propose a full scheme of video codec with enhanced features, which are very useful for multimedia applications. Our codec supports video object based compression: the encoder automatically detects and tracks moving video objects with minor human interactivity. Each object is represented in constrained Delaunay mesh structure and encoded in motion-compensation manner to reduce the needed bit-rate for transmission. By transforming the residual errors (used as correction-values after motion-estimation) into wavelet domain followed by quantising phase, we not only reduce the bandwidth to even lower limit but also guarantee the proper quality (up to the possibly maximum deployed level) for various end-users with a bitstream possessing a famous property - the scalability - regardless the speed of the user-access to the communication network. Simulation is also implemented to demonstrate these virtues. Son Minh Tran, Kalman Fazekas, Jenny Benois-Pineau, Andras Gschwindt |
ICASSP | 3 |
| 2002 | Mesh-based error-scalable video object codec for variable bandwidth multimedia communicationsabstractThe work introduces a complete chain of video object compression. The process is based on an automatic extraction of video objects from raw video. The recent MPEG-4 standard philosophy, including mesh models and wavelet-based compression are involved in the scheme. Constrained Delaunay meshes are used to represent articulated video objects in a flexible manner conveying shape and motion information. The wavelet transform is applied to residual errors for scalable and efficient compression. Results on MPEG-4 test sequences for very low bitrate video communications are encouraging. Son Minh Tran, Kalman Fazekas, Andras Gschwindt, Jenny Benois-Pineau |
ICIP (1) | 4 |
| 2002 | A New Method for Region-Based Depth Ordering in a Video Sequence: Application to Frame Interpolation
Jenny Benois-Pineau, Henri Nicolas |
J. Vis. Commun. Image Represent. | 1 |
| 2001 | Joint tracking of polygonal and triangulated meshes of objects in moving sequences with time varying contentabstractThis paper proposes a method for tracking of objects contained in video sequences. Each video object is represented both by a triangulated mesh and a polygonal mesh. The tracking of such models along a moving sequence is based on a full region-based polygonal tracking. The triangulated mesh is a union of Delaunay meshes on each polygonal region of the VOP. The mesh is articulated, that is, each polygonal region in a VOP is Delaunay-triangulated separately and all partial meshes are connected in the global triangulation. Tracking of triangulated meshes is based on detecting and indexing new objects in the video scene along the time in a polygonal tracking phase. Amal Mahboubi, Jenny Benois-Pineau, Dominique Barba |
ICIP (2) | 2 |
| 2000 | Joint tracking of region-based and mesh models of 2D VOPs in video sequences
Jenny Benois-Pineau, Pierre Verbert, Dominique Barba |
VCIP | 1 |
| 1998 | Extraction of the Relative Depth Information of Objects in Video SequencesabstractThe paper presents a method for an automatic extraction of relative depth of objects in monocular video sequences. Video sequences are represented as a 2D and 1/2 combination of plane objects corresponding to video object plans (MPEG4) with associated affine motion. The method is based on the the analysis of integral profile of the grey-level function of objects which is computed by a discrete version of Radon transform. F. X. Martinez, Jenny Benois-Pineau, Dominique Barba |
ICIP (1) | 2 |
| 1998 | Relative Depth Estimation of Video Objects for Image InterpolationabstractThis paper presents a new method to determine efficiently the relative depth of video objects and its use for region-based interpolation. The depth is extracted by an analyse of occlusion areas of video objects along the time and the image interpolation takes into account this occlusion phenomena to increase the interpolated image quality. This region-based interpolation using relative depth of objects is very important for object-based video compression and manipulation. Franck Morier, Henri Nicolas, Jenny Benois-Pineau, Dominique Barba, Henri Sanson |
ICIP (1) | 3 |
| 1998 | Hierarchical segmentation of video sequences for content manipulation and adaptive coding
Jenny Benois-Pineau, Franck Morier, Dominique Barba, Henri Sanson |
Signal Process. | 1 |
| 1997 | Video coding for wireless varying bit-rate communications based on area of interest and region representationabstractA coding method for videophone communications over wireless, varying bit-rate, channels, is proposed. This method is based on defining one or more areas of interest and breaking them into a number of arbitrarily shaped regions which are homogeneous in their motion. The interest areas in the current implementation are human faces. They are detected with the help of a 2D scene model which is a-priori constructed and is based on knowledge of the application context. A mixed intra/inter-frame coding method is used for the description of the shape and topology of the regions where each polygonal vertex is encoded in a costless manner, either with regards to its predecessor in the same frame, or to the corresponding vertex in the previous frame. This coding is based on rate/distortion optimization constrained by the instantaneously available bit-rate. The error coding is also performed in an adaptive prioritized manner and is restricted within the detected areas of interest. Jenny Benois-Pineau, Dominique Barba, Nikolaos Sarris, Michael G. Strintzis |
ICIP (3) | 1 |
| 1997 | Robust Segmentation of Moving Image SequencesabstractThis paper introduces a method for a hierarchical representation of moving image sequences. The basic level of this hierarchy is a spatial segmentation, rich in regions and whose contours are localised very precisely. The next levels are constructed by a motion-based progressive merging of regions. Different levels of this hierarchy are nested. This representation is sufficiently flexible to allow a "logical zoom" in the image content and to provide a scalable bit rate. So the two goals of content manipulation and of image compression can be reached. Franck Morier, Jenny Benois-Pineau, Dominique Barba, Henri Sanson |
ICIP (1) | 2 |
| 1997 | Detection of human faces in color image sequences with arbitrary motions for very low bit-rate videophone coding
M. Kapfer, Jenny Benois-Pineau |
Pattern Recognit. Lett. | 2 |
| 1996 | Coding of structure in the region-based coder as a problem of optimization on graphsabstractThis paper deals with the coding of structure in region-based coders for moving sequences of images. The sequence of images is represented as spatio-temporal map of regions homogeneous regarded to motion-based criterion and characterized by their shape model, topological relations and motion parameters. Topological and geometrical models of spatio-temporal segmentation are introduced. The coding of geometrical information is formulated as the problem of the construction of the Eulerian circuit in planar non-connected graph. A suboptimal algorithm is proposed. Some results for teleconference sequence are given. Jenny Benois-Pineau, Ali Khenchaf, Dominique Barba |
ICPR | 1 |
| 1996 | Spatio-temporal segmentation of image sequences for object-oriented low bit-rate image coding
Ling Wu 0001, Jenny Benois-Pineau, Philippe Delagnes, Dominique Barba |
Signal Process. Image Commun. | 2 |
| 1995 | Spatio-temporal segmentation of image sequences for object-oriented low bit-rate image codingabstractA method for coding of video sequences based on semantic decomposition into motion homogeneous regions is presented. The set of regions-the spatio-temporal segmentation-is intitialized for the first couple of frames of the sequence and then, tracked along the time axis, that allows to maintain its stability. Based on the spatio-temporal segmentation a predictive coding scheme is developed. Motion parameter vector and boundary description are encoded for each spatio temporal region. The bit-rate obtained for motion parameters is very low (<0.01 bits/pixel). The bit-rate for the boundary description strongly depends on the stability of the segmentation and varies around 0.1 bits/pixel. A costless transmission mode is choosen to encode the boundary component. The high visual quality of predicted frames with almost absolute absence of artifacts and a specific structure of prediction error allows the use of a selective coding of error signal. Ling Wu 0001, Jenny Benois-Pineau, Dominique Barba |
ICIP | 2 |
| 1995 | Active contours approach to object tracking in image sequences with complex background
Philippe Delagnes, Jenny Benois-Pineau, Dominique Barba |
Pattern Recognit. Lett. | 2 |
| 1994 | Motion and structure based image segmentation for object oriented time-varying sequences codingabstractThe paper presents an ascending approach to the construction of the spatio-temporal segmentation of time varying sequences. The particularity of the approach is to combine motion homogeneity of the segments with spatial fidelity of occlusion boundaries. The tracking process with the developed procedures of spatial and temporal adjustment allows a high quality reconstruction of the sequence frames by motion. It makes the approach especially interesting in the framework of object-oriented image sequence coding. Jenny Benois-Pineau, Ling Wu 0001, Philippe Delagnes, Dominique Barba |
ICPR (1) | 1 |
| 1992 | Image segmentation by region-contour cooperation for image codingabstractDescribes a method of image segmentation for image coding based on contour detection and region-growing procedures. Contour detection allows one to find the most evident frontiers of homogeneous regions. The region-growing procedure serves to close contours and to obtain more precise segmentation. The original idea of the method is the choice of growing centres which are placed on the skeleton of non-closed regions and used in split-and-merge procedure of region growing. The method allows one to obtain staircase closure of original contours which is easy to code.> Jenny Benois-Pineau, Dominique Barba |
ICPR (3) | 1 |