Stéphane Ayache

dblp:92/3983 · DBLP profile ↗
← Back
30ranked-venue papers
6as first author
13since 2021 · last 2025
0000-0003-2982-7127ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 1 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Automatic objective metric for the optimization of nonverbal behavior generative models
abstract
Evaluating the quality of generated nonverbal behavior remains a major challenge in the development of generative models.While human evaluations are reliable, they are costly and impractical for large-scale or iterative optimization.In this work, we propose an objective evaluation framework based on aggregated ranks across multiple fidelity and diversity metrics, computed from both raw features and learned latent representations.Compared to existing works, our framework emphasizes consistency across multiple metrics, aiming to provide a more holistic assessment.
Alice Delbosc, Nicolas Sabouret, Brian Ravenet, Stéphane Ayache, Magalie Ochs
IVA4
2024 Implicit Regularization in Deep Tucker Factorization: Low-Rankness via Structured Sparsity
abstract
We theoretically analyze the implicit regularization of deep learning for tensor completion. We show that deep Tucker factorization trained by gradient descent induces a structured sparse regularization. This leads to a characterization of the effect of the depth of the neural network on the implicit regularization and provides a potential explanation for the bias of gradient descent towards solutions with low multilinear rank. Numerical experiments confirm our theoretical findings and give insights into the behavior of gradient descent in deep tensor factorization.
Kais Hariz, Hachem Kadri, Stéphane Ayache, Maher Moakher, Thierry Artières
AISTATS3
2024 Mitigation of gender bias in automatic facial non-verbal behaviors generation
abstract
Research on non-verbal behavior generation for social interactive agents focuses mainly on the believability and synchronization of non-verbal cues with speech. However, existing models, predominantly based on deep learning architectures, often perpetuate biases inherent in the training data. This raises ethical concerns, depending on the intended application of these agents. This paper addresses these issues by first examining the influence of gender on facial non-verbal behaviors. We concentrate on gaze, head movements, and facial expressions. We introduce a classifier capable of discerning the gender of a speaker from their non-verbal cues. This classifier achieves high accuracy on both real behavior data, extracted using state-of-the-art tools, and synthetic data, generated from a model developed in previous work. Building upon this work, we present a new model, FairGenderGen, which integrates a gender discriminator and a gradient reversal layer into our previous behavior generation model. This new model generates facial non-verbal behaviors from speech features, mitigating gender sensitivity in the generated behaviors. Our experiments demonstrate that the classifier, developed in the initial phase, is no longer effective in distinguishing the gender of the speaker from the generated non-verbal behaviors.
Alice Delbosc, Magalie Ochs, Nicolas Sabouret, Brian Ravenet, Stéphane Ayache
ICMI5
2024 Measuring hallucination in disentangled representations
abstract
Disentanglement is a key challenge in representation learning as it may enable several downstream tasks including edition operation at a high semantic level or privacy-preserving applications. While much effort has been put into the design of disentanglement methods and on how to evaluate their disentanglement performance no real studies put in evidence nor proposed to measure the hallucination that may occur in such disentangled representation spaces. This study focuses on characterizing, measuring and investigating hallucination in representation space learnt by state-of-the-art disentanglement methods.
Hamed Benazha, Stéphane Ayache, Hachem Kadri, Thierry Artières
IJCNN2
2024 Opti-CAM: Optimizing saliency maps for interpretability
abstract
Methods based on class activation maps (CAM) provide a simple mechanism to interpret predictions of convolutional neural networks by using linear combinations of feature maps as saliency maps. By contrast, masking-based methods optimize a saliency map directly in the image space or learn it by training another network on additional data. In this work we introduce Opti-CAM, combining ideas from CAM-based and masking-based approaches. Our saliency map is a linear combination of feature maps, where weights are optimized per image such that the logit of the masked image for a given class is maximized. We also fix a fundamental flaw in two of the most common evaluation metrics of attribution methods. On several datasets, Opti-CAM largely outperforms other CAM-based approaches according to the most relevant classification metrics. We provide empirical evidence supporting that localization and classifier interpretability are not necessarily aligned.
Hanwei Zhang 0001, Felipe Torres, Ronan Sicre, Yannis Avrithis, Stéphane Ayache
Comput. Vis. Image Underst.5
2024 Distillation of weighted automata from recurrent neural networks using a spectral approach
Rémi Eyraud, Stéphane Ayache
Mach. Learn.2
2023 DP-Net: Learning Discriminative Parts for Image Recognition
abstract
This paper1presents Discriminative Part Network (DP-Net), a deep architecture with strong interpretation capabilities, which exploits a pretrained Convolutional Neural Network (CNN) combined with a part-based recognition module. This system learns and detects parts in the images that are discriminative among categories, without the need for fine-tuning the CNN, making it more scalable than other part-based models. While part-based approaches naturally offer interpretable representations, we propose explanations at image and category levels and introduce specific constraints on the part learning process to make them more discrimative.
Ronan Sicre, Hanwei Zhang 0001, Julien Dejasmin, Chiheb Daaloul, Stéphane Ayache, Thierry Artières
ICIP5
2022 Are Vision-Language Transformers Learning Multimodal Representations? A Probing Perspective
abstract
In recent years, joint text-image embeddings have significantly improved thanks to the development of transformer-based Vision-Language models. Despite these advances, we still need to better understand the representations produced by those models. In this paper, we compare pre-trained and fine-tuned representations at a vision, language and multimodal level. To that end, we use a set of probing tasks to evaluate the performance of state-of-the-art Vision-Language models and introduce new datasets specifically for multimodal probing. These datasets are carefully designed to address a range of multimodal capabilities while minimizing the potential for models to rely on bias. Although the results confirm the ability of Vision-Language models to understand color at a multimodal level, the models seem to prefer relying on bias in text data for object position and size. On semantically adversarial examples, we find that those models are able to pinpoint fine-grained multimodal differences. Finally, we also notice that fine-tuning a Vision-Language model on multimodal tasks does not necessarily improve its multimodal ability. We make all datasets and code available to replicate experiments.
Emmanuelle Salin, Badreddine Farah, Stéphane Ayache, Benoît Favre
AAAI3
2022 Do Vision-and-Language Transformers Learn Grounded Predicate-Noun Dependencies?
abstract
Recent advances in vision-and-language modeling have seen the development of Transformer architectures that achieve remarkable performance on multimodal reasoning tasks.Yet, the exact capabilities of these black-box models are still poorly understood.While much of previous work has focused on studying their ability to learn meaning at the word-level, their ability to track syntactic dependencies between words has received less attention.We take a first step in closing this gap by creating a new multimodal task targeted at evaluating understanding of predicate-noun dependencies in a controlled setup.We evaluate a range of state-of-the-art models and find that their performance on the task varies considerably, with some models performing relatively well and others at chance level.In an effort to explain this variability, our analyses indicate that the quality (and not only sheer quantity) of pretraining data is essential.Additionally, the best performing models leverage fine-grained multimodal pretraining objectives in addition to the standard image-text matching objectives.This study highlights that targeted and controlled evaluations are a crucial step for a precise and rigorous test of the multimodal knowledge of vision-and-language models.
Mitja Nikolaus, Emmanuelle Salin, Stéphane Ayache, Abdellah Fourtassi, Benoît Favre
EMNLP3
2022 Implicit Regularization with Polynomial Growth in Deep Tensor Factorization
abstract
We study the implicit regularization effects of deep learning in tensor factorization. While implicit regularization in deep matrix and ’shallow’ tensor factorization via linear and certain type of non-linear neural networks promotes low-rank solutions with at most quadratic growth, we show that its effect in deep tensor factorization grows polynomially with the depth of the network. This provides a remarkably faithful description of the observed experimental behaviour. Using numerical experiments, we demonstrate the benefits of this implicit regularization in yielding a more accurate estimation and better convergence properties.
Kais Hariz, Hachem Kadri, Stéphane Ayache, Maher Moakher, Thierry Artières
ICML3
2022 Modeling, Recognizing, and Explaining Apparent Personality From Videos
abstract
Explainability and interpretability are two critical aspects of decision support systems. Despite their importance, it is only recently that researchers are starting to explore these aspects. This paper provides an introduction to explainability and interpretability in the context of apparent personality recognition. To the best of our knowledge, this is the first effort in this direction. We describe a challenge we organized on explainability in first impressions analysis from video. We analyze in detail the newly introduced data set, evaluation protocol, proposed solutions and summarize the results of the challenge. We investigate the issue of bias in detail. Finally, derived from our study, we outline research opportunities that we foresee will be relevant in this area in the near future.
Hugo Jair Escalante, Heysem Kaya, Albert Ali Salah, Sergio Escalera, Yagmur Güçlütürk, Umut Güçlü, Xavier Baró, Isabelle Guyon, Júlio C. S. Jacques Júnior, Meysam Madadi, Stéphane Ayache, Evelyne Viegas, Furkan Gürpinar, Achmadnoer Sukma Wicaksana, Cynthia C. S. Liem, Marcel van Gerven, Rob van Lier
IEEE Trans. Affect. Comput.11
2021 PSM-nets: Compressing Neural Networks with Product of Sparse Matrices
abstract
Over-parameterization of neural networks is a well known issue that comes along with their great performance. Among the many approaches proposed to tackle this problem, low-rank tensor decompositions are largely investigated to compress deep neural networks. Such techniques rely on a low-rank assumption of the layer weight tensors that does not always hold in practice. Following this observation, this paper studies sparsity inducing techniques to build new sparse matrix product layers for high-rate neural networks compression. Specifically, we explore recent advances in sparse optimization to replace each layer's weight matrix, either convolutional or fully connected, by a product of sparse matrices. Our experiments validate that our approach provides a better compression-accuracy trade-off than most popular low-rank-based compression techniques.
Luc Giffon, Stéphane Ayache, Hachem Kadri, Thierry Artières, Ronan Sicre
IJCNN2
2021 Implicit Regularization in Deep Tensor Factorization
abstract
Attempts of studying implicit regularization associated to gradient descent (GD) have identified matrix completion as a suitable test-bed. Late findings suggest that this phenomenon cannot be phrased as a minimization-norm problem, implying that a paradigm shift is required and that dynamics has to be taken into account. In the present work we address the more general setup of tensor completion by leveraging two popularized tensor factorization, namely Tucker and TensorTrain (TT). We track relevant quantities such as tensor nuclear norm, effective rank, generalized singular values and we introduce deep Tucker and TT unconstrained factorization to deal with the completion task. Experiments on both synthetic and real data show that gradient descent promotes solution with low-rank, and validate the conjecture saying that the phenomenon has to be addressed from a dynamical perspective.
Paolo Milanesi, Hachem Kadri, Stéphane Ayache, Thierry Artières
IJCNN3
2020 Partial Trace Regression and Low-Rank Kraus Decomposition
abstract
The trace regression model, a direct extension of the well-studied linear regression model, allows one to map matrices to real-valued outputs. We here introduce an even more general model, namely the partial-trace regression model, a family of linear mappings from matrix-valued inputs to matrix-valued outputs; this model subsumes the trace regression model and thus the linear regression model. Borrowing tools from quantum information theory, where partial trace operators have been extensively studied, we propose a framework for learning partial trace regression models from data by taking advantage of the so-called low-rank Kraus representation of completely positive maps. We show the relevance of our framework with synthetic and real-world experiments conducted for both i) matrix-to-matrix regression and ii) positive semidefinite matrix completion, two tasks which can be formulated as partial trace regression problems.
Hachem Kadri, Stéphane Ayache, Riikka Huusari, Alain Rakotomamonjy, Liva Ralaivola
ICML2
2020 Transfer Learning by Weighting Convolution
abstract
Transferring pretrained deep architectures to datasets with few labels is still a challenge in many real-world situations. This paper presents a new framework to understand convolutional neural networks, by establishing connections between Kronecker factorization and convolutional layers. We then introduce Convolution Weighting Layers that learn a vector of weights for each channel, allowing efficient transfer learning in small training settings, as well as enabling pruning the transferred models. Experiments are conducted on two main settings with few labeled data: transfer learning for classification and transfer learning for retrieval. Two well known convolutional architectures are evaluated on five public datasets. We show that weighting convolutions is efficient to adapt pretrained models to new tasks and that pruned networks conserve good performance.
Stéphane Ayache, Ronan Sicre, Thierry Artières
IJCNN1
2020 Mapping individual differences in cortical architecture using multi-view representation learning
abstract
In neuroscience, understanding inter-individual differences has recently emerged as a major challenge, for which functional magnetic resonance imaging (fMRI) has proven invaluable. For this, neuroscientists rely on basic methods such as univariate linear correlations between single brain features and a score that quantifies either the severity of a disease or the subject's performance in a cognitive task. However, to this date, task-fMRI and resting-state fMRI have been exploited separately for this question, because of the lack of methods to effectively combine them. In this paper, we introduce a novel machine learning method which allows combining the activation- and connectivity-based information respectively measured through these two fMRI protocols to identify markers of individual differences in the functional organization of the brain. It combines a multi-view deep autoencoder which is designed to fuse the two fMRI modalities into a joint representation space within which a predictive model is trained to guess a scalar score that characterizes the patient. Our experimental results demonstrate the ability of the proposed method to outperform competitive approaches and to produce interpretable and biologically plausible results.
Akrem Sellami, François-Xavier Dupé, Bastien Cagna, Hachem Kadri, Stéphane Ayache, Thierry Artières, Sylvain Takerkart
IJCNN5
2020 Guest Editorial: Image and Video Inpainting and Denoising
abstract
The papers in this special issue comprise all aspects of computer vision and pattern recognition devoted to image and video inpainting, including related tasks like denoising, debluring, sampling, super-resolutkon enhancement, restoration, hallucination, etc. The special issue was associated to the 2018 Chalearn Looking at People Satellite ECCV Workshop1 and the 2018 ChaLearn Challenges on Image and Video Inpainting.
Sergio Escalera, Hugo Jair Escalante, Xavier Baró, Isabelle Guyon, Meysam Madadi, Jun Wan 0001, Stéphane Ayache, Yagmur Güçlütürk, Umut Güçlü
IEEE Trans. Pattern Anal. Mach. Intell.7
2019 Deep Networks with Adaptive Nyström Approximation
abstract
Recent work has focused on combining kernel methods and deep learning to exploit the best of the two approaches. Here, we introduce a new architecture of neural networks in which we replace the top dense layers of standard convolutional architectures with an approximation of a kernel function by relying on the Nyström approximation. Our approach is easy and highly flexible. It is compatible with any kernel function and it allows exploiting multiple kernels. We show that our architecture has the same performance than standard architecture on datasets like SVHN and CIFAR100. One benefit of the method lies in its limited number of learnable parameters which makes it particularly suited for small training set sizes, e.g. from 5 to 20 samples per class.
Luc Giffon, Stéphane Ayache, Thierry Artières, Hachem Kadri
IJCNN2
2017 Design of an explainable machine learning challenge for video interviews
abstract
This paper reviews and discusses research advances on “explainable machine learning” in computer vision. We focus on a particular area of the “Looking at People” (LAP) thematic domain: first impressions and personality analysis. Our aim is to make the computational intelligence and computer vision communities aware of the importance of developing explanatory mechanisms for computer-assisted decision making applications, such as automating recruitment. Judgments based on personality traits are being made routinely by human resource departments to evaluate the candidates' capacity of social insertion and their potential of career growth. However, inferring personality traits and, in general, the process by which we humans form a first impression of people, is highly subjective and may be biased. Previous studies have demonstrated that learning machines can learn to mimic human decisions. In this paper, we go one step further and formulate the problem of explaining the decisions of the models as a means of identifying what visual aspects are important, understanding how they relate to decisions suggested, and possibly gaining insight into undesirable negative biases. We design a new challenge on explainability of learning machines for first impressions analysis. We describe the setting, scenario, evaluation metrics and preliminary outcomes of the competition. To the best of our knowledge this is the first effort in terms of challenges for explainability in computer vision. In addition our challenge design comprises several other quantitative and qualitative elements of novelty, including a “coopetition” setting, which combines competition and collaboration.
Hugo Jair Escalante, Isabelle Guyon, Sergio Escalera, Júlio C. S. Jacques Júnior, Meysam Madadi, Xavier Baró, Stéphane Ayache, Evelyne Viegas, Yagmur Güçlütürk, Umut Güçlü, Marcel van Gerven, Rob van Lier
IJCNN7
2013 The Multi-Task Learning View of Multimodal Data
abstract
We study the problem of learning from multiple views using kernel methods in a supervised setting. We approach this problem from a multi-task learning point of view and illustrate how to capture the interesting multimodal structure of the data using multi-task kernels. Our analysis shows that the multi-task perspective offers the flexibility to design more efficient multiple-source learning algorithms, and hence the ability to exploit multiple descriptions of the data. In particular, we formulate the multimodal learning framework using vector-valued reproducing kernel Hilbert spaces, and we derive specific multi-task kernels that can operate over multiple modalities. Finally, we analyze the vector-valued regularized least squares algorithm in this context, and demonstrate its potential in a series of experiments with a real-world multimodal data set.
Hachem Kadri, Stéphane Ayache, Cécile Capponi, Sokol Koço, François-Xavier Dupé, Emilie Morvant
ACML2
2012 Active Cleaning for Video Corpus Annotation
Bahjat Safadi, Stéphane Ayache, Georges Quénot
MMM2
2012 Parsimonious unsupervised and semi-supervised domain adaptation with good similarity functions
Emilie Morvant, Amaury Habrard, Stéphane Ayache
Knowl. Inf. Syst.3
2011 Sparse Domain Adaptation in Projection Spaces Based on Good Similarity Functions
abstract
We address the problem of domain adaptation for binary classification which arises when the distributions generating the source learning data and target test data are somewhat different. We consider the challenging case where no target labeled data is available. From a theoretical standpoint, a classifier has better generalization guarantees when the two domain marginal distributions are close. We study a new direction based on a recent framework of Balcan et al. allowing to learn linear classifiers in an explicit projection space based on similarity functions that may be not symmetric and not positive semi-definite. We propose a general method for learning a good classifier on target data with generalization guarantees and we improve its efficiency thanks to an iterative procedure by reweighting the similarity function - compatible with Balcan et al. framework - to move closer the two distributions in a new projection space. Hyper parameters and reweighting quality are controlled by a reverse validation procedure. Our approach is based on a linear programming formulation and shows good adaptation performances with very sparse models. We evaluate it on a synthetic problem and on real image annotation task.
Emilie Morvant, Amaury Habrard, Stéphane Ayache
ICDM3
2010 Content-based search in multilingual audiovisual documents using the International Phonetic Alphabet
Georges Quénot, Tien Ping Tan, Viet Bac Le, Stéphane Ayache, Laurent Besacier, Philippe Mulhem
Multim. Tools Appl.4
2009 Efficient image concept indexing by harmonic & arithmetic profiles entropy
abstract
We propose new efficient visual features called Profile Entropy Features (PEF), giving information on the structure of the image content, and defined as the entropy of the distribution of a projection of the pixels. We analyse two simple projection operators (arithmetic or harmonic mean), and two orientations (horizontal and vertical). PEF are fast to compute (10 images per sec. on a PentiumIV) and of small dimension. Moreover, we show on High Level Feature task in TrecVid2008 that PEF performs in average better than the features of the state of the art (usual color features, edge direction, Gabor, and Local Binary Pattern). Moreover, we show on another international image retrieval campaign, the Visual Concept Detection of ImageCLEF2008, that the arithmetic and harmonic projections give complementary informations, yielding to the third best rank system in the official run of this campaign. Other properties of the PEF are discussed.
Hervé Glotin, Zhong-Qiu Zhao, Stéphane Ayache
ICIP3
2008 Video Corpus Annotation Using Active Learning
Stéphane Ayache, Georges Quénot
ECIR1
2007 Classifier Fusion for SVM-Based Multimedia Semantic Indexing
Stéphane Ayache, Georges Quénot, Jérôme Gensel
ECIR1
2007 Evaluation of active learning strategies for video indexing
Stéphane Ayache, Georges Quénot
Signal Process. Image Commun.1
2006 Context-Based Conceptual Image Indexing
abstract
Automatic semantic classification of image databases is very useful for users searching and browsing but it is at the same time a very challenging research problem. Local features based image classification is one of the promising way to bridge the semantic gap in detecting concepts. This paper proposes a framework for incorporating contextual information into the concept detection process. The proposed method combines local and global classifiers (SVMs) with stacking. We studied the impact of topologic and semantic contexts in concept detection performance and proposed solutions to handle the large amount of dimensions involved in classified data. We conducted experiments on TRECVID'04 data set with 48104 images and 5 concepts. We found that the use of context yields a significant improvement both for the topologic and semantic contexts
Stéphane Ayache, Georges Quénot, Shin'ichi Satoh 0001
ICASSP (2)1
2005 Video Shot Classification Using Lexical Context
Stéphane Ayache, Georges Quénot, Mbarek Charhad
ECIR1