Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jean Martinet

dblp:03/1471 · DBLP profile ↗
← Back
32ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0001-8821-5556ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-authorHuman-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Visualization and visual analytics · 50% Multimedia analysis and retrieval · 50%

Topics — the 1 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visualization and visual analytics
eye tracking analysis
0.112008
Analyzing eye fixations and gaze orientations on films and pictures · ACM Multimedia 2008

Methods — techniques the papers use, named apart from their topics

projection aggregation · 0.1dispersion computation · 0.1clustering of eye positions · 0.1
YearPublicationVenuePosition
2025 Cross-Modal Neuromorphic Semantic Segmentation based on Knowledge Distillation
abstract
Event-based cameras, with their high temporal resolution and low energy consumption, offer significant advantages over conventional cameras for computer vision tasks. However, current semantic segmentation algorithms for event-based data face challenges in achieving optimal performance for two main reasons: 1) the mismatch between sparse event streams (even when encoded into pseudo-frames) and the dense data structure of traditional frame-based images for which most existing implementations were designed, and 2) the lack of texture information in events, which solely detect temporal variations in brightness. To address this challenge, we propose a novel Cross-Modal (CM) Knowledge Distillation (KD) approach. Our method transfers knowledge from a high-performing Artificial Neural Network (ANN) processing fused grey-scale images and events to a Spiking Neural Network (SNN) - a bio-inspired, energy-efficient computing paradigm - operating on event data alone. Experiments on the DDD17 and DSEC-semantic datasets demonstrate that our approach significantly improves semantic segmentation results while reducing SNN energy consumption. This work represents the first application of cross-modal knowledge distillation to neuromorphic semantic segmentation, paving the way for more efficient event-based vision systems.
Dalia Hareb, Jean Martinet, Benoît Miramond
IJCNN2
2024 EvSegSNN: Neuromorphic Semantic Segmentation for Event Data
abstract
Semantic segmentation is an important computer vision task, particularly for scene understanding and navigation of autonomous vehicles and UAVs. Several variations of deep neural network architectures have been designed to tackle this task. However, due to their huge computational costs and their high memory consumption, these models are not meant to be deployed on resource-constrained systems. To address this limitation, we introduce an end-to-end biologically inspired semantic segmentation approach by combining Spiking Neural Networks (SNNs, a low-power alternative to classical neural networks) with event cameras whose output data can directly feed these neural network inputs. We have designed EvSegSNN, a biologically plausible encoder-decoder U-shaped architecture relying on Parametric Leaky Integrate and Fire neurons in an objective to trade-off resource usage against performance. The experiments conducted on DDD17 demonstrate that EvSegSNN outperforms the closest state-of-the-art model in terms of MIoU while reducing the number of parameters by a factor of 1.6 and sparing a batch normalization stage.
Dalia Hareb, Jean Martinet
IJCNN2
2023 Simultaneous neuromorphic selection of multiple salient objects for event vision
abstract
The combined use of spiking neural networks and event cameras is gaining momentum in the field of embedded computer vision as they promise to reduce latency and computational resource requests. However, state-of-the-art embedded neuromorphic models show little interest in modifying input data to optimise model performance, memory usage, latency, and power consumption. This work addresses this optimisation trade-off by implementing a neuromorphic model of salient selection, which simultaneously outputs multiple segregated objects of interest detected in an event-based scene. This work extends previous ones and identifies regions of interest as those corresponding to a high spatiotemporal density of events. Without any training and with a limited number of neurons, the proposed model is able to simultaneously detect different objects with a delay of only 14ms at most, and filtered objects maintain 73% of the original data's classification performance. We are thus confident that the method proposed in this paper will allow for improving the subsequent neuromorphic processing of event data on embedded systems. To the best of our knowledge, it is the first neuromorphic model able to simultaneously select multiple objects of interest. Our code can be found here: github.com/amygruel/FoveationStakes_DVS/.
Amélie Gruel, Jean Martinet, Michele Magno
IJCNN2
2023 Performance comparison of DVS data spatial downscaling methods using Spiking Neural Networks
abstract
Dynamic Vision Sensors (DVS) are an unconventional type of camera that produces sparse and asynchronous event data, which has recently led to a strong increase in its use for computer vision tasks namely in robotics. Embedded systems face limitations in terms of energy resources, memory, computational power, and communication bandwidth. Hence, this application calls for a way to reduce the amount of data to be processed while keeping the relevant information for the task at hand. We thus believe that a formal definition of event data reduction methods will provide a step further towards sparse data processing.The contributions of this paper are twofold: we introduce two complementary neuromorphic methods based on Spiking Neural Networks for DVS data spatial reduction, which is to best of our knowledge the first proposal of neuromorphic event data reduction; then we study for each method the trade-off between the amount of information kept after reduction, the performance of gesture classification after reduction and their capacity to handle events in real time. We demonstrate here that the proposed SNN-based methods outperform existing methods in a classification task for most dividing factors and are significantly better at handling data in real time, and make therefore the optimal choice for fully-integrated energy-efficient event data reduction running dynamically on a neuromorphic platform. Our code is publicly available online at: https://github.com/amygruel/EvVisu.
Amélie Gruel, Jean Martinet, Bernabé Linares-Barranco, Teresa Serrano-Gotarredona
WACV2
2021 Bio-inspired visual attention for silicon retinas based on spiking neural networks applied to pattern classification
abstract
Visual attention can be defined as the behavioral and cognitive process of selectively focusing on a discrete aspect of sensory cues while disregarding other perceivable information. This biological mechanism, more specifically saliency detection, has long been used in multimedia indexing to drive the analysis only on relevant parts of images or videos for further processing.The recent advent of silicon retinas (or event cameras - sensors that measure pixel-wise changes in brightness and output asynchronous events accordingly) raises the question of how to adapt attention and saliency to the unconventional type of such sensors' output. Silicon retina aims to reproduce the biological retina behaviour. In that respect, they produce punctual events in time that can be construed as neural spikes and interpreted as such by a neural network.In particular, Spiking Neural Networks (SNNs) represent an asynchronous type of artificial neural network closer to biology than traditional artificial networks, mainly because they seek to mimic the dynamics of neural membrane and action potentials over time. SNNs receive and process information in the form of spike trains. Therefore, they make for a suitable candidate for the efficient processing and classification of incoming event patterns measured by silicon retinas. In this paper, we review the biological background behind the attentional mechanism, and introduce a case study of event videos classification with SNNs, using a biology-grounded low-level computational attention mechanism, with interesting preliminary results.
Amélie Gruel, Jean Martinet
CBMI2
2021 Comparing HMAX and BoVW Models for Large-Scale Image Classification
abstract
Image classification is one of the most important topics in computer vision. It became crucial for large image datasets. In the literature, several image classification approaches are proposed. In this context, Bag-of-Visual Words (BoVW) model has been widely used. The BoVW model relies on building visual vocabulary and images are represented as histograms of visual words. However, recently, attention has been shifted to the use of complex architectures which are characterized by multilevel processing. HMAX (Hierarchical Max-pooling model) model has attracted a great deal of attention in image classification, due to its architecture, which alternates layers of feature extraction with layers of pooling. This paper aims at comparing bags of visual words model to HMAX model for image classification using large datasets. To achieve this goal, we study the use of image features obtained by BoVW model with SIFT (Scale-Invariant Feature Transform) descriptors, and we compare them to HMAX features. Image classification is performed by using the support vector machine (SVM) classifiers. Both HMAX and BoVW models are tested on ImageNet and OpenImages datasets and results have shown that the classification performance obtained by HMAX model outperforms the classification using BoVW model.
Jalila Filali, Hajer Baazaoui Zghal, Jean Martinet
KES3
2021 OntoAnnClass: ontology-based image annotation driven by classification using HMAX features
Jalila Filali, Hajer Baazaoui Zghal, Jean Martinet
Multim. Tools Appl.3
2020 A critical survey of STDP in Spiking Neural Networks for Pattern Recognition
abstract
The bio-inspired concept of Spike-Timing-Dependent Plasticity (STDP) derived from neurobiology is increasingly used in Spiking Neural Networks (SNNs) nowadays. Mostly found in unsupervised learning, though recent work has shown its usefulness in supervised or reinforced paradigms too, STDP is a key element to understanding SNN architectures' learning process. This review introduces a categorisation of its several variants and discusses their specificities and applications, from a pattern recognition perspective. It gathers a variety of definitions used in machine learning for pattern recognition. It provides relevant information for research communities of various backgrounds looking for an overview of this field.
Alex Vigneron, Jean Martinet
IJCNN2
2020 Ontology-Based Image Classification and Annotation
abstract
With the rapid growth of image collections, image classification and annotation has been active areas of research with notable recent progress. Bag-of-Visual-Words (BoVW) model, which relies on building visual vocabulary, has been widely used in this area. Recently, attention has been shifted to the use of advanced architectures which are characterized by multi-level processing. Hierarchical Max-Pooling (HMAX) model has attracted a great deal of attention in image classification. To improve image classification and annotation, several approaches based on ontologies have been proposed. However, image classification and annotation remain a challenging problem due to many related issues like the problem of ambiguity between classes. This problem can affect the quality of both classification and annotation results. In this paper, we propose an ontology-based image classification and annotation approach. Our contributions consist of the following: (1) exploiting ontological relationships between classes during both image classification and annotation processes; (2) combining the outputs of hypernym–hyponym classifiers to lead to a better discrimination between classes; and (3) annotating images by combining hypernym and hyponym classification results in order to improve image annotation and to reduce the ambiguous and inconsistent annotations. The aim is to improve image classification and annotation by using ontologies. Several strategies have been experimented, and the obtained results have shown that our proposal improves image classification and annotation.
Jalila Filali, Hajer Baazaoui Zghal, Jean Martinet
Int. J. Pattern Recognit. Artif. Intell.3
2018 Dynamic Index Finger Gesture Video Dataset for Mobile Interaction
abstract
This paper introduces an original video dataset containing dynamic index finger gestures in a mobile context. The dataset consists of 746 video sequences of 6 index finger gestures: left, right, up, down, tap, circle. The video sequences are obtained from a mobile phone's rear camera equipped with a wide angle lens, thus the dataset features particular challenges due to background motion and the distorted field of view. We present a baseline method that uses optical flow features in the form of a histogram to represent the motion information in the video sequences.
Cagan Arslan, Ioan Marius Bilasco, Jean Martinet
CBMI3
2017 Visually Supporting Image Annotation Based on Visual Features and Ontologies
abstract
Automatic Image Annotation (AIA) is a challenging problem in the field of image retrieval, and several methods have been proposed. However, visually supporting this important tasks and reducing the semantic gap between low-level image features and high-level semantic concepts still remains a key issue. In this paper, we propose a visually supporting image annotation framework based on visual features and ontologies. Our framework relies on three main components: (i) extraction and classification of features component, (ii) ontology’s building component and (iii) image annotation component. Our goal consists on improving the visual image annotation by:(1) extracting invariant and complex visual features; (2) integrating feature classification results and semantic concepts to build ontology and (3) combining both visual and semantic similarities during the image annotation process.
Jalila Filali, Hajer Baazaoui Zghal, Jean Martinet
IV3
2016 Towards Visual Vocabulary and Ontology-based Image Retrieval System
abstract
International audience
Jalila Filali, Hajer Baazaoui Zghal, Jean Martinet
ICAART (2)3
2015 Space-time Histograms And Their Application To Person Re-identification In TV Shows
abstract
The annotation of video streams by automatic content analysis is a growing field of research. The possibility of recognising persons appearing in TV shows allows to automatically structure ever-growing video archives. We propose a new descriptor to re-identify persons featured in videos, that is to say, to spot all occurrences of persons throughout a video. Our approach is dynamic as it benefits from motion information contained in videos, whereas the static approaches are solely based on still images. We extract person-tracks from videos and match them using a new descriptor and its associated similarity measure: the space-time histogram. The originality of our approach is the integration of temporal data into the descriptor. Experiments show that it provides a better estimation of the similarity between person-tracks. Our contribution has been evaluated using a corpus of real life french TV shows broadcasted on BFMTV and LCP TV channels and on some annotated episodes from "Buffy: the Vampire Slayer". Experimental results show that our approach significantly improves the precision of the re-identification process thanks to the use of the temporal dimension.
Rémi Auguste, Jean Martinet, Pierre Tirilly
ICMR2
2015 Boosting gender recognition performance with a fuzzy inference system
Taner Danisman, Ioan Marius Bilasco, Jean Martinet
Expert Syst. Appl.3
2014 DLBP: A novel descriptor for depth image based face recognition
abstract
This paper presents a novel descriptor for face depth images, generalizing the well-known Local Binary Pattern (LBP), in order to enhance its discriminative power for smooth depth images. The proposed descriptor is based on detecting shape patterns from face surfaces and enables accurate and fast description of shape variation in depth images. It is in the same form as conventional LBP, so patterns can be readily combined to form joint histograms to represent depth faces. The descriptor is computationally very simple, rapid and it is totally training-free. When we associate our descriptor in a face recognition scheme based on nearest neighbor classifier, it shows its discriminative power in depth based face recognition comparing to the conventional LBP and other extensions proposed for 3D face recognition. Many experiments are conducted on different databases in order to evaluate our method.
Amel Aissaoui, Jean Martinet, Chaabane Djeraba
ICIP2
2014 Elementary block extraction for mobile image search
José Mennesson, Pierre Tirilly, Jean Martinet
ICIP3
2014 Multimodal understanding for person recognition in video broadcasts
abstract
International audience
Frédéric Béchet, Meriem Bendris, Delphine Charlet, Géraldine Damnati, Benoît Favre, Mickael Rouvier, Rémi Auguste, Benjamin Bigot, Richard Dufour, Corinne Fredouille, Georges Linarès, Jean Martinet, Grégory Senay, Pierre Tirilly
INTERSPEECH12
2014 Iterative Random Visual Word Selection
abstract
In content based image retrieval, one of the most important step is the construction of image signatures. To do so, a part of state-of-the-art approaches propose to build a visual vocabulary. In this paper, we propose a new methodology for visual vocabulary construction that obtains high retrieval results. Moreover, it is computationally inexpensive to build and needs no prior knowledge on features or dataset used.
Thierry Urruty, Syntyche Gbèhounou, Huu Ton Le, Jean Martinet, Christine Fernandez-Maloigne
ICMR4
2014 Rapid and accurate face depth estimation in passive stereo systems
Amel Aissaoui, Jean Martinet, Chaabane Djeraba
Multim. Tools Appl.2
2013 Intelligent pixels of interest selection with application to facial expression recognition using multilayer perceptron
Taner Danisman, Ioan Marius Bilasco, Jean Martinet, Chaabane Djeraba
Signal Process.3
2012 3D face reconstruction in a binocular passive stereoscopic system using face properties
abstract
In this paper, we introduce a novel approach for face stereo reconstruction in passive stereo vision system. Our approach is based on the generation of a facial disparity map, requiring neither expensive devices nor generic face models. It consists of incorporating face properties in the disparity estimation to enhance the 3D face reconstruction. An algorithm based on the Active Shape Model (ASM) is proposed to acquire 3D sparse estimation of the face with a high confidence. Using sparse estimation as guidance and considering the face symmetry and smoothness, the dense disparity is completed. Experimental results are presented to demonstrate the reconstruction accuracy of the proposed method.
Amel Aissaoui, Jean Martinet, Chaabane Djeraba
ICIP2
2012 Toward a higher-level visual representation for content-based image retrieval
Ismail Elsayad, Jean Martinet, Thierry Urruty, Chaabane Djeraba
Multim. Tools Appl.2
2011 A semantically significant visual representation for social image retrieval
abstract
Having effective methods to access the desired images is essential nowadays with the availability of a huge amount of digital images. We propose a higher-level visual representation that enhances the traditional part-based Bag of Visual Words (BOW) representation in two aspects. Firstly, we introduce a new multilayer semantic significance analysis (MSSA) model to select Semantically Significant Visual Words (SSVWs) from the classical visual words in order to overcome the noisiness of the feature quantization process. Secondly, we strengthen the discrimination power of SSVWs by constructing Semantically Significant Visual Phrases (SSVPs) from frequently co-occurring SSVWs in the same local context that are semantically coherent. Finally, the large-scale extensive experimental results show that the proposed higher-level visual representation outperforms the traditional part-based image representation in social image retrieval.
Ismail Elsayad, Jean Martinet, Thierry Urruty, Yassine Benabbas, Chaabane Djeraba
ICME2
2011 A Semantic Higher-Level Visual Representation for Object Recognition
Ismail Elsayad, Jean Martinet, Thierry Urruty, Chaabane Djeraba
MMM (1)2
2011 A relational vector space model using an advanced weighting scheme for image retrieval
Jean Martinet, Yves Chiaramella, Philippe Mulhem
Inf. Process. Manag.1
2010 Toward a higher-level visual representation for content-based image retrieval
abstract
Having effective methods to access the desired images is essential nowadays with the availability of huge amount of digital images. The proposed approach is based on an analogy between content-based image retrieval and text retrieval. The aim of the approach is to build a meaningful mid-level representation of images to be used later for matching between a query image and other images in the desired database. The approach is based firstly on constructing different visual words using local patch extraction and fusion of descriptors. Secondly, we introduce a new method using multilayer pLSA to eliminate the noisiest words generated by the vocabulary building process. Thirdly, a new spatial weighting scheme is introduced that consists in weighting visual words according to the probability of each visual word to belong to each of the n Gaussian. Finally, we construct visual phrases from groups of visual words that are involved in strong association rules. Experimental results show that our approach outperforms the results of traditional image retrieval techniques.
Ismail Elsayad, Jean Martinet, Thierry Urruty, Samir Amir, Chaabane Djeraba
MoMM2
2008 Analyzing eye fixations and gaze orientations on films and pictures
abstract
Eye movements are arguably the most natural and repetitive movement of a human being. The most mundane activity, such as watching television or reading a newspaper, involves this automatic activity which consists of shifting our gaze from one point to another. Identification of the components of eye movements (fixations and saccades) is an essential part in the analysis of visual behavior because these types of movements provide the basic elements used by further investigations of human vision. However, many of the algorithms that detect fixations present a number of problems. In this paper, we present the results of a new fixation identification technique that is based on clustering of eye positions, using projections and a projection aggregation applied to static pictures. We also present results of a new method that computes dispersion of eye fixations in videos considering a multi-user environment.
Anthony Martinet, Jean Martinet, Nacim Ihaddadene, Stanislas Lew, Chaabane Djeraba
ACM Multimedia2
2008 Media objects for user-centered similarity matching
Jean Martinet, Shin'ichi Satoh 0001, Yves Chiaramella, Philippe Mulhem
Multim. Tools Appl.1
2007 Using Visual-Textual Mutual Information and Entropy for Inter-modal Document Indexing
Jean Martinet, Shin'ichi Satoh 0001
ECIR1
2007 Image-Based Quizzes from News Video Archives
abstract
This paper proposes a method for generating image-based quizzes from news video achieves. Although there are many types of quizzes, in this work we focus on matching quizzes in which an image is to be matched to one of several choices that are statements. The key to making a successful quiz of this type is to extract choice statements that are similar but nonidentical to the true statement regarding the image. For image selection, we need to select images that are suitable for image-based quizzes. In this paper, we report some preliminary work and highlight some interesting issues in making quizzes. We also propose a method for automatically generating quizzes using feature-based clustering analyses and assess the outcomes of trials actually using our method.
Masanori Sano, Nobuyuki Yagi, Jean Martinet, Norio Katayama, Shin'ichi Satoh 0001
ICME3
2005 A model for weighting image objects in home photographs
abstract
The paper presents a contribution to image indexing consisting in a weighting model for visible objects -- or image objects -- in home photographs. To improve its effectiveness this weighting model has been designed according to human perception criteria about what is estimated as important in photographs. Four basic hypotheses related to human perception are presented, and their validity is estimated as compared to actual observations from a user study. Finally a formal definition of this weighting model is presented and its consistence with the user study is evaluated.
Jean Martinet, Yves Chiaramella, Philippe Mulhem
CIKM1
2003 A Weighting Scheme for Star-Graphs
Jean Martinet, Iadh Ounis, Yves Chiaramella, Philippe Mulhem
ECIR1