Slim Essid

dblp:53/6904 · DBLP profile ↗
← Back
75ranked-venue papers
9as first author
27since 2021 · last 2025
0000-0002-0028-327XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 55 · 7 first-author · 17 since 2021Artificial intelligence and machine learning · 26 · 2 first-author · 16 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 iKnow-audio: Integrating Knowledge Graphs with Audio-Language Models
abstract
Contrastive Language-Audio Pretraining (CLAP) models learn by aligning audio and text in a shared embedding space, enabling powerful zero-shot recognition.However, their performance is highly sensitive to prompt formulation and language nuances, and they often inherit semantic ambiguities and spurious correlations from noisy pretraining data.While prior work has explored prompt engineering, adapters, and prefix tuning to address these limitations, the use of structured prior knowledge remains largely unexplored.We present iKnow-audio, a framework that integrates knowledge graphs with audio-language models to provide robust semantic grounding.iKnow-audio builds on the Audio-centric Knowledge Graph (AKG), which encodes ontological relations comprising semantic, causal, and taxonomic connections reflective of everyday sound scenes and events.By training knowlege graph embedding models on the AKG and refining CLAP predictions through this structured knowledge, iKnow-audio improves disambiguation of acoustically similar sounds and reduces reliance on prompt engineering.Comprehensive zero-shot evaluations across six benchmark datasets demonstrate consistent gains over baseline CLAP, supported by embedding-space analyses that highlight improved relational grounding.Resources are publicly available at
Michel Olvera, Changhong Wang 0002, Paraskevas Stamatiadis, Gaël Richard, Slim Essid
EMNLP5
2025 Perceptual Noise-Masking with Music through Deep Spectral Envelope Shaping
abstract
People often listen to music in noisy environments, seeking to isolate themselves from ambient sounds. Indeed, a music signal can mask some of the noise’s frequency components due to the effect of simultaneous masking. In this article, we propose a neural network based on a psychoacoustic masking model, designed to enhance the music’s ability to mask ambient noise by reshaping its spectral envelope with predicted filter frequency responses. The model is trained with a perceptual loss function that balances two constraints: effectively masking the noise while preserving the original music mix and the user’s chosen listening level. We evaluate our approach on simulated data replicating a user’s experience of listening to music with headphones in a noisy environment. The results, based on defined objective metrics, demonstrate that our system improves the state of the art.
Clémentine Berger, Roland Badeau, Slim Essid
ICASSP3
2025 O-EENC-SD: Efficient Online End-to-End Neural Clustering for Speaker Diarization
abstract
We introduce O-EENC-SD: an end-to-end online speaker diarization system based on EEND-EDA, featuring a novel RNN-based stitching mechanism for online prediction. In particular, we develop a novel centroid refinement decoder whose usefulness is assessed through a rigorous ablation study. Our system provides key advantages over existing methods: a hyperparameter-free solution compared to unsupervised clustering approaches, and a more efficient alternative to current online end-to-end methods, which are computationally costly. We demonstrate that O-EENC-SD is competitive with the state of the art in the two-speaker conversational telephone speech domain, as tested on the CallHome dataset. Our results show that O-EENC-SD provides a great trade-off between DER and complexity, even when working on independent chunks with no overlap, making the system extremely efficient.
Elio Gruttadauria, Mathieu Fontaine 0002, Jonathan Le Roux, Slim Essid
ICASSP4
2025 Multiple Choice Learning for Efficient Speech Separation with Many Speakers
abstract
Training speech separation models in the supervised setting raises a permutation problem: finding the best assignation between the model predictions and the ground truth separated signals. This inherently ambiguous task is customarily solved using Permutation Invariant Training (PIT). In this article, we instead consider using the Multiple Choice Learning (MCL) framework, which was originally introduced to tackle ambiguous tasks. We demonstrate experimentally on the popular WSJ0-mix and LibriMix benchmarks that MCL matches the performances of PIT, while being computationally advantageous. This opens the door to a promising research direction, as MCL can be naturally extended to handle a variable number of speakers, or to tackle speech separation in the unsupervised setting.
David Perera, François Derrida, Théo Mariotte, Gaël Richard, Slim Essid
ICASSP5
2025 Masked Latent Prediction and Classification for Self-Supervised Audio Representation Learning
abstract
Recently, self-supervised learning methods based on masked latent prediction have proven to encode input data into powerful representations. However, during training, the learned latent space can be further transformed to extract higher-level information that could be more suited for down-stream classification tasks. Therefore, we propose a new method: MAsked latenT Prediction And Classification (MATPAC), which is trained with two pretext tasks solved jointly. As in previous work, the first pretext task is a masked latent prediction task, ensuring a robust input representation in the latent space. The second one is unsupervised classification, which utilises the latent representations of the first pretext task to match probability distributions between a teacher and a student. We validate the MATPAC method by comparing it to other state-of-the-art proposals and conducting ablations studies. MATPAC reaches state-of-the-art self-supervised learning results on reference audio classification datasets such as OpenMIC, GTZAN, ESC-50 and US8K and outperforms comparable supervised methods’ results for musical auto-tagging on Magna-tag-a-tune.
Aurian Quelennec, Pierre Chouteau, Geoffroy Peeters, Slim Essid
ICASSP4
2025 Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
abstract
Personalized speech enhancement (PSE) has shown convincing results when it comes to extracting a known target voice among interfering ones. The corresponding systems usually incorporate a representation of the target voice within the enhancement system, which is extracted from an enrollment clip of the target voice with upstream models. Those models are generally heavy as the speaker embedding’s quality directly affects PSE performances. Yet, embeddings generated beforehand cannot account for the variations of the target voice during inference time. In this paper, we propose to perform on-the-fly refinement of the speaker embedding using a tiny speaker encoder. We first introduce a novel contrastive knowledge distillation methodology in order to train a 150k-parameter encoder from complex embeddings. We then use this encoder within the enhancement system during inference and show that the proposed method greatly improves PSE performances while maintaining a low computational load.
Thomas Serre, Mathieu Fontaine 0002, Éric Benhaim, Slim Essid
ICASSP4
2025 MTSE: Multi-Target Speaker Extraction for Conversation Scenarios
Thomas Serre, Mathieu Fontaine 0002, Éric Benhaim, Slim Essid
INTERSPEECH4
2025 Speech self-supervised representations benchmarking: A case for larger probing heads
Mohamed Salah Zaïem, Youcef Kemiche, Titouan Parcollet, Slim Essid, Mirco Ravanelli
Comput. Speech Lang.4
2024 Collaborating Foundation Models for Domain Generalized Semantic Segmentation
abstract
Domain Generalized Semantic Segmentation (DGSS) deals with training a model on a labeled source domain with the aim of generalizing to unseen domains during inference. Existing DGSS methods typically effectuate robust features by means of Domain Randomization (DR). Such an approach is often limited as it can only account for style diversification and not content. In this work, we take an orthogonal approach to DGSS and propose to use an assembly of CoLlaborative FOUndation models for Domain Generalized Semantic Segmentation (CLOUDS). In detail, CLOUDS is a framework that integrates Foundation Models of various kinds: (i) CLIP backbone for its robust feature representation, (ii) Diffusion Model to diversify the content, thereby covering various modes of the possible target distribution, and (iii) Segment Anything Model (SAM) for iteratively refining the predictions of the segmentation model. Extensive experiments show that our CLOUDS excels in adapting from synthetic to real DGSS benchmarks and under varying weather conditions, notably outperforming prior methods by 5.6% and 6.7% on averaged mIoU, respectively. Our code is available at https://github.com/yasserben/CLOUDS
Yasser Benigmim, Subhankar Roy, Slim Essid, Vicky Kalogeiton, Stéphane Lathuilière
CVPR3
2024 Adapting Pitch-Based Self Supervised Learning Models for Tempo Estimation
abstract
Tempo estimation is the task of estimating the periodicity of the dominant rhythm pulse of a music audio signal. It has therefore a close relationship with dominant pitch estimation. Recently, both tasks have been addressed in a Self-Supervised Learning (SSL) fashion so as to leverage unlabelled data for training. In this work, we study the applicability of two successful pitch-based SSL models, SPICE and PESTO, for the purpose of tempo estimation. Both successfully exploit Siamese networks with a pitch-shifting view generation between the two branches. To apply these models for tempo estimation, we represent the audio signal by the Constant-Q transform (CQT) of its onset-strength-function and adapt their view generation using time-stretching (instead of pitch shifting), which is efficiently implemented by shifting the CQT. In a large experiment, we show that simply adapting PESTO in this way yields superior results than the previous SSL approach to tempo estimation for most datasets used in the reference benchmark. Further, since PESTO is light-weight, requiring only a few training data, we study a new learning scheme where the downstream datasets are processed directly in a SSL fashion (without access to labels) showing that this is an interesting alternative further improving the performance for some datasets.
Antonin Gagneré, Slim Essid, Geoffroy Peeters
ICASSP2
2024 Online Speaker Diarization of Meetings Guided by Speech Separation
abstract
Overlapped speech is notoriously problematic for speaker diarization systems. Consequently, the use of speech separation has recently been proposed to improve their performance. Although promising, speech separation models struggle with realistic data because they are trained on simulated mixtures with a fixed number of speakers. In this work, we introduce a new speech separation-guided diarization scheme suitable for the online speaker diarization of long meeting recordings with a variable number of speakers, as present in the AMI corpus. We envisage ConvTasNet and DPRNN as alternatives for the separation networks, with two or three output sources. To obtain the speaker diarization result, voice activity detection is applied on each estimated source. The final model is fine-tuned end-to-end, after first adapting the separation to real data using AMI. The system operates on short segments, and inference is performed by stitching the local predictions using speaker embeddings and incremental clustering. The results show that our system improves the state-of-the-art on the AMI headset mix, using no oracle information and under full evaluation (no collar and including overlapped speech). Finally, we show the strength of our system particularly on overlapped speech sections.
Elio Gruttadauria, Mathieu Fontaine 0002, Slim Essid
ICASSP3
2024 On The Choice of the Optimal Temporal Support for Audio Classification with Pre-Trained Embeddings
abstract
Current state-of-the-art audio analysis systems rely on pre-trained embedding models, often used off-the-shelf as (frozen) feature extractors. Choosing the best one for a set of tasks is the subject of many recent publications. However, one aspect often overlooked in these works is the influence of the duration of audio input considered to extract an embedding, which we refer to as Temporal Support (TS). In this work, we study the influence of the TS for well-established or emerging pre-trained embeddings, chosen to represent different types of architectures and learning paradigms. We conduct this evaluation using both musical instrument and environmental sound datasets, namely OpenMIC, TAU Urban Acoustic Scenes 2020 Mobile, and ESC-50. We especially highlight that Audio Spectrogram Transformer-based systems (PaSST and BEATs) remain effective with smaller TS, which therefore allows for a drastic reduction in memory and computational cost. Moreover, we show that by choosing the optimal TS we reach competitive results across all tasks. In particular, we improve the state-of-the-art results on OpenMIC, using BEATs and PaSST without any fine-tuning.
Aurian Quelennec, Michel Olvera, Geoffroy Peeters, Slim Essid
ICASSP4
2024 Winner-takes-all learners are geometry-aware conditional density estimators
abstract
Winner-takes-all training is a simple learning paradigm, which handles ambiguous tasks by predicting a set of plausible hypotheses. Recently, a connection was established between Winner-takes-all training and centroidal Voronoi tessellations, showing that, once trained, hypotheses should quantize optimally the shape of the conditional distribution to predict. However, the best use of these hypotheses for uncertainty quantification is still an open question. In this work, we show how to leverage the appealing geometric properties of the Winner-takes-all learners for conditional density estimation, without modifying its original training scheme. We theoretically establish the advantages of our novel estimator both in terms of quantization and density estimation, and we demonstrate its competitiveness on synthetic and real-world datasets, including audio data.
Victor Letzelter, David Perera, Cédric Rommel, Mathieu Fontaine 0002, Slim Essid, Gaël Richard, Patrick Pérez
ICML5
2024 An eye for an ear: zero-shot audio description leveraging an image captioner with audio-visual token distribution matching
abstract
Multimodal large language models have fueled progress in image captioning. These models, fine-tuned on vast image datasets, exhibit a deep understanding of semantic concepts. In this work, we show that this ability can be re-purposed for audio captioning, where the joint image-language decoder can be leveraged to describe auditory content associated with image sequences within videos featuring audiovisual content. This can be achieved via multimodal alignment. Yet, this multimodal alignment task is non-trivial due to the inherent disparity between audible and visible elements in real-world videos. Moreover, multimodal representation learning often relies on contrastive learning, facing the challenge of the so-called modality gap which hinders smooth integration between modalities. In this work, we introduce a novel methodology for bridging the audiovisual modality gap by matching the distributions of tokens produced by an audio backbone and those of an image captioner. Our approach aligns the audio token distribution with that of the image tokens, enabling the model to perform zero-shot audio captioning in an unsupervised fashion. This alignment allows for the use of either audio or audiovisual input by combining or substituting the image encoder with the aligned audio encoder. Our method achieves significantly improved performances in zero-shot audio captioning, compared to existing approaches.
Hugo Malard, Michel Olvera, Stéphane Lathuilière, Slim Essid
NeurIPS4
2024 Annealed Multiple Choice Learning: Overcoming limitations of Winner-takes-all with annealing
abstract
We introduce Annealed Multiple Choice Learning (aMCL) which combines simulated annealing with MCL. MCL is a learning framework handling ambiguous tasks by predicting a small set of plausible hypotheses. These hypotheses are trained using the Winner-takes-all (WTA) scheme, which promotes the diversity of the predictions. However, this scheme may converge toward an arbitrarily suboptimal local minimum, due to the greedy nature of WTA. We overcome this limitation using annealing, which enhances the exploration of the hypothesis space during training. We leverage insights from statistical physics and information theory to provide a detailed description of the model training trajectory. Additionally, we validate our algorithm by extensive experiments on synthetic datasets, on the standard UCI benchmark, and on speech separation.
David Perera, Victor Letzelter, Théo Mariotte, Adrien Cortés, Mickaël Chen, Slim Essid, Gaël Richard
NeurIPS6
2024 Self-Supervised Learning of Multi-Level Audio Representations for Music Segmentation
abstract
The task of music structure analysis refers to automatically identifying the location and the nature of musical sections within a song. In the supervised scenario, structural annotations generally result from exhaustive data collection processes, which represents one of the main challenges of this task. Moreover, both the subjectivity of music structure and the hierarchical characteristics it exhibits make the obtained structural annotations not fully reliable, in the sense that they do not convey a “universal ground-truth” unlike other tasks in music information retrieval. On the other hand, the quickly growing quantity of available music data has enabled weakly supervised and self-supervised approaches to achieve impressive results on a wide range of music-related problems. In this work, a self-supervised learning method is proposed to learn robust multi-level music representations prior to structural segmentation using contrastive learning. To this end, sets of frames sampled at different levels of detail are used to train a deep neural network in a disentangled manner. The proposed method is evaluated on both flat and multi-level segmentation. We show that each distinct sub-region of the output embeddings can efficiently account for structural similarity at their own targeted level of detail, which ultimately improves performance of downstream flat and multi-level segmentation. Finally, complementary experiments are carried out to study how the obtained representations can be further adapted to specific datasets using a supervised fine-tuning objective in order to facilitate structure retrieval in domains where human annotations remain scarce.
Morgan Buisson, Brian McFee, Slim Essid, Hélène C. Crayencour
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Cosmopolite Sound Monitoring (CoSMo): A Study of Urban Sound Event Detection Systems Generalizing to Multiple Cities
abstract
Measuring noise in cities and automatically identifying the corresponding sound sources are a crucial challenge for policymakers. Indeed, such information helps addressing noise pollution and improving the well-being of urban dwellers. In recent years, researchers have provided annotated datasets recorded in two major cities to foster the development of urban sound event detection (SED) systems. This paper presents an in-depth study of the behaviour of state-of-the-art SED systems well suited to our problem, combining three far-field real recordings datasets which can be used jointly during training. In our evaluation, we highlight the performance gaps existing between simple and hard recording examples based on the salience of sound events and the polyphony of the recordings. We provide new proximity annotations for this analysis. We evaluate the ability of urban SED systems to generalize across cities with varying degrees of training supervision. We show that such generalization is hindered mostly by the difficulties current urban SED systems have to detect sound events with low salience along with sound events in highly polyphonic soundscapes.
Florian Angulo, Slim Essid, Geoffroy Peeters, Christophe Mietlicki
ICASSP2
2023 Speech Self-Supervised Representation Benchmarking: Are We Doing it Right?
abstract
International audience
Mohamed Salah Zaïem, Youcef Kemiche, Titouan Parcollet, Slim Essid, Mirco Ravanelli
INTERSPEECH4
2023 Automatic Data Augmentation for Domain Adapted Fine-Tuning of Self-Supervised Speech Representations
abstract
International audience
Mohamed Salah Zaïem, Titouan Parcollet, Slim Essid
INTERSPEECH3
2023 Resilient Multiple Choice Learning: A learned scoring scheme with application to audio scene analysis
abstract
We introduce Resilient Multiple Choice Learning (rMCL), an extension of the MCL approach for conditional distribution estimation in regression settings where multiple targets may be sampled for each training input. Multiple Choice Learning is a simple framework to tackle multimodal density estimation, using the Winner-Takes-All (WTA) loss for a set of hypotheses. In regression settings, the existing MCL variants focus on merging the hypotheses, thereby eventually sacrificing the diversity of the predictions. In contrast, our method relies on a novel learned scoring scheme underpinned by a mathematical framework based on Voronoi tessellations of the output space, from which we can derive a probabilistic interpretation. After empirically validating rMCL with experiments on synthetic data, we further assess its merits on the sound source localization problem, demonstrating its practical usefulness and the relevance of its interpretation.
Victor Letzelter, Mathieu Fontaine 0002, Mickaël Chen, Patrick Pérez, Slim Essid, Gaël Richard
NeurIPS5
2022 Automatic Data Augmentation Selection and Parametrization in Contrastive Self-Supervised Speech Representation Learning
abstract
International audience
Mohamed Salah Zaïem, Titouan Parcollet, Slim Essid
INTERSPEECH3
2022 Opinions in Interactions : New Annotations of the SEMAINE Database
abstract
In this paper, we present the process we used in order to collect new annotations of opinions over the multimodal corpus SEMAINE composed of dyadic interactions. The dataset had already been annotated continuously in two affective dimensions related to the emotions: Valence and Arousal. We annotated the part of SEMAINE called Solid SAL composed of 79 interactions between a user and an operator playing the role of a virtual agent designed to engage a person in a sustained, emotionally colored conversation. We aligned the audio at the word level using the available high-quality manual transcriptions. The annotated dataset contains 5627 speech turns for a total of 73,944 words, corresponding to 6 hours 20 minutes of dyadic interactions. Each interaction has been labeled by three annotators at the speech turn level following a three-step process. This method allows us to obtain a precise annotation regarding the opinion of a speaker. We obtain thus a dataset dense in opinions, with more than 48% of the annotated speech turns containing at least one opinion. We then propose a new baseline for the detection of opinions in interactions improving slightly a state of the art model with RoBERTa embeddings. The obtained results on the database are promising with a F1-score at 0.72.
Valentin Barrière, Slim Essid, Chloé Clavel
LREC2
2021 Neuro-Steered Music Source Separation With EEG-Based Auditory Attention Decoding And Contrastive-NMF
abstract
We propose a novel informed music source separation paradigm, which can be referred to as neuro-steered music source separation. More precisely, the source separation process is guided by the user’s selective auditory attention decoded from his/her EEG response to the stimulus. This high-level prior information is used to select the desired instrument to isolate and to adapt the generic source separation model to the observed signal. To this aim, we leverage the fact that the attended instrument’s neural encoding is substantially stronger than the one of the unattended sources left in the mixture. This "contrast" is extracted using an attention decoder and used to inform a source separation model based on non-negative matrix factorization named Contrastive-NMF. The results are promising and show that the EEG information can automatically select the desired source to enhance and improve the separation quality.
Giorgia Cantisani, Slim Essid, Gaël Richard
ICASSP2
2021 Distributed Speech Separation in Spatially Unconstrained Microphone Arrays
abstract
Speech separation with several speakers is a challenging task because of the non-stationarity of the speech and the strong signal similarity between interferent sources. Current state-of-the-art solutions can separate well the different sources using sophisticated deep neural networks which are very tedious to train. When several microphones are available, spatial information can be exploited to design much simpler algorithms to discriminate speakers. We propose a distributed algorithm that can process spatial information in a spatially unconstrained microphone array. The algorithm relies on a convolutional recurrent neural network that can exploit the signal diversity from the distributed nodes. In a typical case of a meeting room, this algorithm can capture an estimate of each source in a first step and propagate it over the microphone array in order to increase the separation performance in a second step. We show that this approach performs even better when the number of sources and nodes increases. We also study the influence of a mismatch in the number of sources between the training and testing conditions.
Nicolas Furnon, Romain Serizel, Irina Illina, Slim Essid
ICASSP4
2021 Conditional Independence for Pretext Task Selection in Self-Supervised Speech Representation Learning
abstract
Through solving pretext tasks, self-supervised learning (SSL) leverages unlabeled data to extract useful latent representations replacing traditional input features in the downstream task. A common pretext task consists in pretraining a SSL model on pseudo-labels derived from the original signal. This technique is particularly relevant for speech data where various meaningful signal processing features may serve as pseudo-labels. However, the process of selecting pseudo-labels, for speech or other types of data, remains mostly unexplored and currently relies on observing the results on the final downstream task. Nevertheless, this methodology is not sustainable at scale due to substantial computational (hence carbon) costs. Thus, this paper introduces a practical and theoretical framework to select relevant pseudo-labels with respect to a given downstream task. More precisely, we propose a functional estimator of the pseudo-label utility grounded in the conditional independence theory, which does not require any training. The experiments conducted on speaker recognition and automatic speech recognition validate our estimator, showing a significant correlation between the performance observed on the downstream task and the utility estimates obtained with our approach, facilitating the prospection of relevant pseudo-labels for self-supervised speech representation learning.
Mohamed Salah Zaïem, Titouan Parcollet, Slim Essid
Interspeech3
2021 Early Detection of User Engagement Breakdown in Spontaneous Human-Humanoid Interaction
abstract
This paper presents a supervised classification system for forecasting a potential user engagement breakdown in human-robot interaction. We define engagement breakdown as a failure to successfully complete a predefined interaction scenario, where the user leaves before the expected end. The goal is thus to detect as early as possible such a potential engagement breakdown during the interaction between a human and a humanoid robot. To this end, we exploit a dataset that we have collected in real-world conditions where a set of participants were left to spontaneously engage in an interaction with the robot. The dataset is labeled according to the presence/absence of engagement breakdown. This study investigates the use of a multimodal approach to this problem, where a set of non-verbal features is considered to characterize the users' behavior. The use of combined multimodal features is found to effectively improve the performance of the system. The optimal set of data streams useful for this task is the combination of the distance to the robot, gaze and head motion, as well as facial expressions and speech. We study the time extent over which a user's departure can be anticipated. We find that this ability to anticipate the departure depends on the window during which we observe the user behavior.
Atef Ben Youssef, Chloé Clavel, Slim Essid
IEEE Trans. Affect. Comput.3
2021 DNN-Based Mask Estimation for Distributed Speech Enhancement in Spatially Unconstrained Microphone Arrays
abstract
Deep neural network (DNN)-based speech enhancement algorithms in microphone arrays have now proven to be efficient solutions to speech understanding and speech recognition in noisy environments. However, in the context of ad-hoc microphone arrays, many challenges remain and raise the need for distributed processing. In this paper, we propose to extend a previously introduced distributed DNN-based time-frequency mask estimation scheme that can efficiently use spatial information in form of so-called compressed signals which are pre-filtered target estimations. We study the performance of this algorithm named Tango under realistic acoustic conditions and investigate practical aspects of its optimal application. We show that the nodes in the microphone array cooperate by taking profit of their spatial coverage in the room. We also propose to use the compressed signals not only to convey the target estimation but also the noise estimation in order to exploit the acoustic diversity recorded throughout the microphone array.
Nicolas Furnon, Romain Serizel, Slim Essid, Irina Illina
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 DNN-based Distributed Multichannel Mask Estimation for Speech Enhancement in Microphone Arrays
abstract
Multichannel processing is widely used for speech enhancement but several limitations appear when trying to deploy these solutions in the real world. Distributed sensor arrays that consider several devices with a few microphones is a viable solution which allows for exploiting the multiple devices equipped with microphones that we are using in our everyday life. In this context, we propose to extend the distributed adaptive node-specific signal estimation approach to a neural network framework. At each node, a local filtering is performed to send one signal to the other nodes where a mask is estimated by a neural network in order to compute a global multichannel Wiener filter. In an array of two nodes, we show that this additional signal can be leveraged to predict the masks and leads to better speech enhancement performance than when the mask estimation relies only on the local signals.
Nicolas Furnon, Romain Serizel, Irina Illina, Slim Essid
ICASSP4
2020 Weakly Supervised Representation Learning for Audio-Visual Scene Analysis
abstract
Audio-visual (AV) representation learning is an important task from the perspective of designing machines with the ability to understand complex events. To this end, we propose a novel multimodal framework that instantiates multiple instance learning. Specifically, we develop methods that identify events and localize corresponding AV cues in unconstrained videos. Importantly, this is done using weak labels where only video-level event labels are known without any information about their location in time. We show that the learnt representations are useful for performing several tasks such as event/object classification, audio event detection, audio source separation and visual object localization. An important feature of our method is its capacity to learn from unsynchronized audio-visual events. We also demonstrate our framework's ability to separate out the audio source of interest through a novel use of nonnegative matrix factorization. State-of-the-art classification results, with a F1-score of 65.0, are achieved on DCASE 2017 smart cars challenge data with promising generalization to diverse object types such as musical instruments. Visualizations of localized visual regions and audio segments substantiate our system's efficacy, especially when dealing with noisy situations where modality-specific cues appear asynchronously.
Sanjeel Parekh, Slim Essid, Alexey Ozerov, Ngoc Q. K. Duong, Patrick Pérez, Gaël Richard
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 From the Token to the Review: A Hierarchical Multimodal approach to Opinion Mining
abstract
Alexandre Garcia, Pierre Colombo, Florence d’Alché-Buc, Slim Essid, Chloé Clavel. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Alexandre Garcia 0001, Pierre Colombo, Florence d'Alché-Buc, Slim Essid, Chloé Clavel
EMNLP/IJCNLP (1)4
2019 A Music Structure Informed Downbeat Tracking System Using Skip-chain Conditional Random Fields and Deep Learning
abstract
In recent years the task of downbeat tracking has received increasing attention and the state of the art has been improved with the introduction of deep learning methods. Among proposed solutions, existing systems exploit short-term musical rules as part of their language modelling. In this work we show in an oracle scenario how including longer-term musical rules, in particular music structure, can enhance downbeat estimation. We introduce a skip-chain conditional random field language model for downbeat tracking designed to include section information in an unified and flexible framework. We combine this model with a state-of-the-art convolutional-recurrent network and we contrast the system’s performance to the commonly used Bar Pointer model. Our experiments on the popular Beatles dataset show that incorporating structure information in the language model leads to more consistent and more robust downbeat estimations.
Magdalena Fuentes, Brian McFee, Hélène C. Crayencour, Slim Essid, Juan Pablo Bello
ICASSP4
2018 Attitude Classification in Adjacency Pairs of a Human-Agent Interaction with Hidden Conditional Random Fields
abstract
In this paper, the main goal is to classify, in a human-agent interaction, the attitude of the user using hidden conditional random fields. This model allows us to capture the dynamics of the interaction in the pairs of speech turns (adjacency pairs) analyzed by our system. High level linguistic features are computed at word level. The features include syntactic features, a statistical word embedding model and subjectivity lexicons. The proposed system is evaluated on the SEMAINE corpus. We obtain a Fl-score of 0.80, labeling using the most probable sequence of hidden states.
Valentin Barrière, Chloé Clavel, Slim Essid
ICASSP3
2018 An Ensemble Learning Approach to Detect Epileptic Seizures from Long Intracranial EEG Recordings
abstract
This paper proposes a patient-specific supervised classification algorithm to detect seizures in long offline intracranial electroencephalographic (iEEG) recordings. The main idea of the proposed algorithm is to combine a set of probabilistic classifiers, trained on a dataset of 1 s epochs, into a weighted ensemble classifier which can be used to analyze longer 5 s data segments. The method is trained and evaluated on 24 patients, all suffering from focal medically intractable epilepsy, from the Epilepsiae database. The evaluation of the method, conducted using an average of 113 hours (min: 32 h, max: 229 h) of iEEG data per patient, shows that the proposed algorithm improves upon existing methods for seizure detection with iEEG.
Jean-Baptiste Schiratti, Jean-Eudes Le Douget, Michel Le Van Quyen, Slim Essid, Alexandre Gramfort
ICASSP4
2018 Structured Output Learning with Abstention: Application to Accurate Opinion Prediction
abstract
Motivated by Supervised Opinion Analysis, we propose a novel framework devoted to Structured Output Learning with Abstention (SOLA). The structure prediction model is able to abstain from predicting some labels in the structured output at a cost chosen by the user in a flexible way. For that purpose, we decompose the problem into the learning of a pair of predictors, one devoted to structured abstention and the other, to structured output prediction. To compare fully labeled training data with predictions potentially containing abstentions, we define a wide class of asymmetric abstention-aware losses. Learning is achieved by surrogate regression in an appropriate feature space while prediction with abstention is performed by solving a new pre-image problem. Thus, SOLA extends recent ideas about Structured Output Prediction via surrogate problems and calibration theory and enjoys statistical guarantees on the resulting excess risk. Instantiated on a hierarchical abstention-aware loss, SOLA is shown to be relevant for fine-grained opinion mining and gives state-of-the-art results on this task. Moreover, the abstention-aware representations can be used to competitively predict user-review ratings based on a sentence-level opinion predictor.
Alexandre Garcia 0001, Chloé Clavel, Slim Essid, Florence d'Alché-Buc
ICML3
2017 Overlapping sound event detection with supervised Nonnegative Matrix Factorization
abstract
In this paper we propose a supervised Nonnegative Matrix Factorization (NMF) model for overlapping sound event detection in real life audio. We start by highlighting the usefulness of non-euclidean NMF to learn representations for detecting and classifying acoustic events in a multi-label setting. Then, we propose to learn a classifier and the NMF decomposition in a joint optimization problem. This is done with a general β-divergence version of the nonnegative task-driven dictionary learning model. An experimental evaluation is performed on the development set of the DCASE 2016 task3 challenge. The proposed supervised NMF-based system improves performance over the baseline and the submitted systems.
Victor Bisot, Slim Essid, Gaël Richard
ICASSP2
2017 Motion informed audio source separation
abstract
In this paper we tackle the problem of single channel audio source separation driven by descriptors of the sounding object's motion. As opposed to previous approaches, motion is included as a soft-coupling constraint within the nonnegative matrix factorization framework. The proposed method is applied to a multimodal dataset of instruments in string quartet performance recordings where bow motion information is used for separation of string instruments. We show that the approach offers better source separation result than an audio-based baseline and the state-of-the-art multimodal-based approaches on these very challenging music mixtures.
Sanjeel Parekh, Slim Essid, Alexey Ozerov, Ngoc Q. K. Duong, Patrick Pérez, Gaël Richard
ICASSP2
2017 Supervised group nonnegative matrix factorisation with similarity constraints and applications to speaker identification
abstract
This paper presents supervised feature learning approaches for speaker identification that rely on nonnegative matrix factorisation. Recent studies have shown that group nonnegative matrix factorisation and task-driven supervised dictionary learning can help performing effective feature learning for audio classification problems. This paper proposes to integrate a recent method that relies on group nonnegative matrix factorisation into a task-driven supervised framework for speaker identification. The goal is to capture both the speaker variability and the session variability while exploiting the discriminative learning aspect of the task-driven approach. Results on a subset of the ESTER corpus prove that the proposed approach can be competitive with I-vectors.
Romain Serizel, Victor Bisot, Slim Essid, Gaël Richard
ICASSP3
2017 UE-HRI: a new dataset for the study of user engagement in spontaneous human-robot interactions
abstract
In this paper, we present a new dataset of spontaneous interactions between a robot and humans, of which 54 interactions (between 4 and 15-minute duration each) are freely available for download and use. Participants were recorded while holding spontaneous conversations with the robot Pepper. The conversations started automatically when the robot detected the presence of a participant and kept the recording if he/she accepted the agreement (i.e. to be recorded). Pepper was in a public space where the participants were free to start and end the interaction when they wished. The dataset provides rich streams of data that could be used by research and development groups in a variety of areas.
Atef Ben Youssef, Chloé Clavel, Slim Essid, Miriam Bilac, Marine Chamoux, Angelica Lim
ICMI3
2017 Opinion Dynamics Modeling for Movie Review Transcripts Classification with Hidden Conditional Random Fields
abstract
In this paper, the main goal is to detect a movie reviewer's opinion using hidden conditional random fields. This model allows us to capture the dynamics of the reviewer's opinion in the transcripts of long unsegmented audio reviews that are analyzed by our system. High level linguistic features are computed at the level of inter-pausal segments. The features include syntactic features, a statistical word embedding model and subjectivity lexicons. The proposed system is evaluated on the ICT-MMMO corpus. We obtain a F1-score of 82\%, which is better than logistic regression and recurrent neural network approaches. We also offer a discussion that sheds some light on the capacity of our system to adapt the word embedding model learned from general written texts data to spoken movie reviews and thus model the dynamics of the opinion.
Valentin Barrière, Chloé Clavel, Slim Essid
INTERSPEECH3
2017 Feature Learning With Matrix Factorization Applied to Acoustic Scene Classification
abstract
In this paper, we study the usefulness of various matrix factorization methods for learning features to be used for the specific acoustic scene classification (ASC) problem. A common way of addressing ASC has been to engineer features capable of capturing the specificities of acoustic environments. Instead, we show that better representations of the scenes can be automatically learned from time–frequency representations using matrix factorization techniques. We mainly focus on extensions including sparse, kernel-based, convolutive and a novel supervised dictionary learning variant of principal component analysis and nonnegative matrix factorization. An experimental evaluation is performed on two of the largest ASC datasets available in order to compare and discuss the usefulness of these methods for the task. We show that the unsupervised learning methods provide better representations of acoustic scenes than the best conventional hand-crafted features on both datasets. Furthermore, the introduction of a novel nonnegative supervised matrix factorization model and deep neural networks trained on spectrograms, allow us to reach further improvements.
Victor Bisot, Romain Serizel, Slim Essid, Gaël Richard
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 Acoustic scene classification with matrix factorization for unsupervised feature learning
abstract
In this paper we study the use of unsupervised feature learning for acoustic scene classification (ASC). The acoustic environment recordings are represented by time-frequency images from which we learn features in an unsupervised manner. After a set of preprocessing and pooling steps, the images are decomposed using matrix factorization methods. By decomposing the data on a learned dictionary, we use the projection coefficients as features for classification. An experimental evaluation is done on a large ASC dataset to study popular matrix factorization methods such as Principal Component Analysis (PCA) and Non-negative Matrix Factorization (NMF) as well as some of their extensions including sparse, kernel based and convolutive variants. The results show the compared variants lead to significant improvement compared to the state-of-the-art results in ASC.
Victor Bisot, Romain Serizel, Slim Essid, Gaël Richard
ICASSP3
2016 Group nonnegative matrix factorisation with speaker and session variability compensation for speaker identification
abstract
This paper presents a feature learning approach for speaker identification that is based on nonnegative matrix factorisation. Recent studies have shown that with such models, the dictionary atoms can represent well the speaker identity. The approaches proposed so far focused only on speaker variability and not on session variability. However, this later point is a crucial aspect in the success of the I-vector approach that is now the state-of-the-art in speaker identification. This paper proposes a method that relies on group nonnegative matrix factorisation and that is inspired by the I-vector training procedure. By doing so the proposed approach intends to capture both the speaker variability and the session variability. Results on a small corpus prove that the proposed approach can be competitive with I-vectors.
Romain Serizel, Slim Essid, Gaël Richard
ICASSP2
2016 Machine listening techniques as a complement to video image analysis in forensics
abstract
Video is now one of the major sources of information for forensics. However, video documents can be originating from various recording devices (CCTV, mobile devices, etc.) with inconsistent quality and can sometimes be recorded in challenging light or motion conditions. Therefore, the amount of information that can be extracted relying solely on video image can vary to a great extent. Most of the videos however generally include audio recording as well. Machine listening can then become a valuable complement to video image analysis in challenging scenarios. In this paper, the authors present a brief overview of some machine listening techniques and their application to the analysis of video documents for forensics. The applicability of these techniques to forensics problems is then discussed in the light of machine listening system performances.
Romain Serizel, Victor Bisot, Slim Essid, Gaël Richard
ICIP3
2015 A Conditional Random Field system for beat tracking
abstract
In the present work, we introduce a new probabilistic model for the task of estimating beat positions in a musical audio recording, instantiating the Conditional Random Field (CRF) framework. Our approach takes its strength from a sophisticated temporal modeling of the audio observations, accounting for local tempo variations which are readily represented in the CRF model proposed using well-chosen potentials. The system is experimentally evaluated by studying its performance on 3 datasets of 1394 music excerpts of various western music styles and comparatively to 4 reference systems in the light of 6 reference evaluation metrics. The results show that the proposed system tracks perceptively coherent pulses and is very effective in estimating the beat positions while further work is needed to find the correct salient tempo.
Thomas Fillon, Cyril Joder, Simon Durand, Slim Essid
ICASSP4
2014 Assessment of new spectral features for eeg-based emotion recognition
abstract
The choice of appropriate features for automatic emotion recognition based on electroencephalographic (EEG) signals remains to date an open research question. In this paper we explore a wide range of potentially useful features, including original ones, comparing them to previous proposals through a rigorous experimental evaluation, using a strict cross-validation protocol. In particular we assess the effectiveness of new spectral features-both in multi-channel and single-channel EEG setups-for the problem of discriminating positively and negatively excited emotions. The evaluation is conducted using the ENTERFACE'06 dataset allowing us to study the behaviour of the tested features across different subjects. Our results prove the usefulness of various new spectral features even in single-channel setups. We also observe that the optimal selection of features is highly subject-dependent. Finally combining different groups of features we find the valence recognition accuracy to be possibly as high as 78%.
Anne-Claire Conneau, Slim Essid
ICASSP2
2014 Gesture recognition using a NMF-based representation of motion-traces extracted from depth silhouettes
abstract
We present a novel approach that classifies full-body human gestures using original spatio-temporal features obtained by applying non-negative matrix factorisation (NMF) to an extended depth silhouette representation. This extended representation, the motion-trace representation, incorporates temporal dimensions as it is built by superimposition of consecutive depth silhouettes. From this representation, a dictionary of local motion features is learned using NMF. Thus the projection of these local motion feature components on the incoming motion-traces results in a compact spatio-temporal feature representation. Those new features are then exploited using hidden Markov models for gesture recognition. Our experiments on a gesture dataset show that our approach outperforms more traditional methods that use pose features or decomposition techniques such as principal component analysis.
Aymeric Masurelle, Slim Essid, Gaël Richard
ICASSP2
2014 Piecewise constant nonnegative matrix factorization
abstract
In this paper we propose a non-negative matrix factorization (NMF) model with piecewise-constant activation coefficients. This structure is enforced using a total variation penalty on the rows of the activation matrix. The resulting optimization problem is solved with a majorization-minimization procedure. The proposed algorithm is well suited to analyze data explained by underlying piecewise-constant sequences of states. Its properties are first illustrated using synthetic data. We then use it to solve a video structuring problem that involves both segmentation and clustering tasks. An improvement over a state-of-the-art temporally smoothed NMF algorithm of both clustering and segmentation quality measures is observed.
Nicolas Seichepine, Slim Essid, Cédric Févotte, Olivier Cappé
ICASSP2
2013 Non-negative matrix factorization for single-channel EEG artifact rejection
abstract
New applications of Electroencephalographic recording (EEG) pose new challenges in terms of artifact removal. In our work we target applications where the EEG is to be captured by a single electrode and a number of additional lightweight sensors are allowed. Thus, this paper introduces a new method for artifact removal for single-channel EEG recordings using nonnegative matrix factorisation (NMF) in a Gaussian source separation framework. We focus the study on ocular artifacts and show that by properly exploiting prior information on the latter, through the analysis of electrooculographic recordings, our artifact removal results on single-channel EEG are comparable to the results obtained with the classic multi-channel Independent Component Analysis technique.
Cécilia Damon, Antoine Liutkus, Alexandre Gramfort, Slim Essid
ICASSP4
2013 Probabilistic dance performance alignment by fusion of multimodal features
abstract
This paper presents a probabilistic framework for the multimodal alignment of dance movements. The approach is based on a Hidden Markov Model (HMM) and considers different feature functions, each corresponding to a particular modality, namely motion features, extracted from depth maps, and audio features, extracted from audio recordings of dancers' steps. We show that this approach allows performing accurate dancer alignment, while constituting a general framework for various multimodal alignment tasks.
Angélique Dremeau, Slim Essid
ICASSP2
2013 Soft nonnegative matrix co-factorizationwith application to multimodal speaker diarization
abstract
This paper presents a new method for bimodal nonnegative matrix factorization (NMF). This method is well-suited to situations where two streams of data are concurrently analyzed and are expected to be related by loosely common factors. It allows for a soft co-factorization, which takes into account the relationship that exists between the modalities being processed, but returns different factors for distinct modalities. There is no need that the data related with each modality live in the same feature space; there is also no need that they have the same dimensionality. The co-factorization is obtained via a majorization-minimization (MM) algorithm. The behavior of the method is illustrated on both synthetic and real-world data. In particular, we show that exploiting the correlation between audio and video modalities in edited talk-show videos improve speaker diarization results.
Nicolas Seichepine, Slim Essid, Cédric Févotte, Olivier Cappé
ICASSP2
2013 Learning Optimal Features for Polyphonic Audio-to-Score Alignment
abstract
This paper addresses the design of feature functions for the matching of a musical recording to the symbolic representation of the piece (the score). These feature functions are defined as dissimilarity measures between the audio observations and template vectors corresponding to the score. By expressing the template construction as a linear mapping from the symbolic to the audio representation, one can learn the feature functions by optimizing the linear transformation. In this paper, we explore two different learning strategies. The first one uses a best-fit criterion (minimum divergence), while the second one exploits a discriminative framework based on a Conditional Random Fields model (maximum likelihood criterion). We evaluate the influence of the feature functions in an audio-to-score alignment task, on a large database of popular and classical polyphonic music. The results show that with several types of models, using different temporal constraints, the learned mappings have the potential to outperform the classic heuristic mappings. Several representations of the audio observations, along with several distance functions are compared in this alignment task. Our experiments elect the symmetric Kullback-Leibler divergence. Moreover, both the spectrogram and a CQT-based representation turn out to provide very accurate alignments, detecting more than 97% of the onsets with a precision of 100 ms with our most complex system.
Cyril Joder, Slim Essid, Gaël Richard
IEEE Trans. Speech Audio Process.2
2013 Smooth Nonnegative Matrix Factorization for Unsupervised Audiovisual Document Structuring
abstract
This paper introduces a new paradigm for unsupervised audiovisual document structuring. In this paradigm, a novel Nonnegative Matrix Factorization (NMF) algorithm is applied on histograms of counts (relating to a bag of features representation of the content) to jointly discover latent structuring patterns and their activations in time. Our NMF variant employs the Kullback-Leibler divergence as a cost function and imposes a temporal smoothness constraint to the activations. It is solved by a majorization-minimization technique. The approach proposed is meant to be generic and is particularly well suited to applications where the structuring patterns may overlap in time. As such, it is evaluated on two person-oriented video structuring tasks (one using the visual modality and the second the audio). This is done using a challenging database of political debate videos. Our results outperform reference results obtained by a method using Hidden Markov Models. Further, we show the potential that our general approach has for audio speaker diarization.
Slim Essid, Cédric Févotte
IEEE Trans. Multim.1
2013 A Multimodal Approach to Speaker Diarization on TV Talk-Shows
abstract
In this article, we propose solutions to the problem of speaker diarization of TV talk-shows, a problem for which adapted multimodal approaches, relying on other streams of data than only audio, remain largely under exploited. Hence we propose an original system that leverages prior knowledge on the structure of this type of content, especially the visual information relating to the active speakers, for an improved diarization performance. The architecture of this system can be decomposed into two main stages. First a reliable training set is created, in an unsupervised fashion, for each participant of the TV program being processed. This data is assembled by the association of visual and audio descriptors carefully selected in a clustering cascade. Then, Support Vector Machines are used for the classification of the speech data (of a given TV program). The performance of this new architecture is assessed on two French talk-show collections: Le Grand Échiquier and On n'a pas tout dit. The results show that our new system outperforms state-of-the-art methods, thus evidencing the effectiveness of kernel-based methods, as well as visual cues, in multimodal approaches to speaker diarization of challenging contents such as TV talk-shows.
Félicien Vallet, Slim Essid, Jean Carrive
IEEE Trans. Multim.2
2012 A single-class SVM based algorithm for computing an identifiable NMF
abstract
The geometric interpretation of Nonnegative Matrix Factorisation (NMF) as the problem of determining a convex cone that “well describes” the data under analysis has been key for addressing a major shortcoming of the “mainstream” NMF algorithms, that is the non-identifiability of the factorisation. On the basis of such geometric motivations, this paper proposes a novel algorithm that makes use of single-class support vector machines to recover the targeted NMF components. Not only does this new approach alleviate the NMF illposedness issue, but also it allows for automatically estimating the number of relevant NMF components, as demonstrated through experiments described in the paper. Moreover, it is readily kernelised thus opening the way for non-linear factorisations of the data.
Slim Essid
ICASSP1
2012 An advanced virtual dance performance evaluator
abstract
The ever increasing availability of high speed Internet access has led to a leap in technologies that support real-time realistic interaction between humans in online virtual environments. In the context of this work, we wish to realise the vision of an online dance studio where a dance class is to be provided by an expert dance teacher and to be delivered to online students via the web. In this paper we study some of the technical issues that need to be addressed in this challenging scenario. In particular, we describe an automatic dance analysis tool that would be used to evaluate a student's performance and provide him/her with meaningful feedback to aid improvement.
Slim Essid, Dimitrios S. Alexiadis, Robin Tournemenne, Marc Gowing, Philip Kelly, David S. Monaghan, Petros Daras, Angélique Dremeau, Noel E. O'Connor
ICASSP1
2012 A regressive boosting approach to automatic audio tagging based on soft annotator fusion
abstract
Automatic tagging of music has mostly been treated as a classification problem. In this framework, the association of a tag to a song is characterized in a “hard” fashion: the tag is either relevant or not. Yet, the relevance of a tag to a song is not always evident. Indeed, during the ground-truth annotation process, several annotators may express doubts, or disagree with each other. In this paper, we propose to fuse annotators' decisions in a way to keep information about this uncertainty. This fusion provides us continuous scores, that are used for training a regressive boosting algorithm. Our experiments show that regression with this soft ground truth leads to a more accurate learning, and better predictions, compared to traditionally used binary classification.
Rémi Foucard, Slim Essid, Mathieu Lagrange, Gaël Richard
ICASSP2
2012 Decomposing the video editing structure of a talk-show using nonnegative matrix factorization
abstract
We introduce a novel video structuring scheme that exploits nonnegative matrix factorization (NMF) on count data (in a bag of features representation of the visual stream) to jointly discover latent structuring patterns and their activations in time. Our NMF variant employs the Kullback-Leibler divergence as a cost function and imposes a temporal smoothness constraint to the activations. It is solved by a majorization-minimization technique. Our method is shown to be successful for decomposing the high-level editing structure of talk-shows. It is evaluated using a challenging database of TV political-debate programs, and found to clearly outperform a reference HMM method.
Slim Essid, Cédric Févotte
ICIP1
2012 Analysis of dance movements using gaussian processes: extended abstract
abstract
This work addresses the Huawei/3DLife Grand Challenge, presenting a novel method for the analysis of dance movements. The approach focuses on the decomposition of the dance movements into elementary motions. Placing this problem into a probabilistic framework, we propose to exploit Gaussian processes to accurately model the different components of the decomposition. The preliminary results, presented in this paper, are very promising. In particular, two applications are considered, illustrating the relevance of the proposed approach, namely the correction of tracking errors and the smoothing of some movements of the teacher to help toward the dance learning.
Antoine Liutkus, Angélique Dremeau, Dimitrios S. Alexiadis, Slim Essid, Petros Daras
ACM Multimedia4
2011 Hidden Discrete Tempo Model: A tempo-aware timing model for audio-to-score alignment
abstract
In this paper, we present the Hidden Discrete Tempo Model, an effective Dynamic Bayesian Network for audio to score matching. Its main feature is an explicit modeling of tempo, which directly in fluences the timing model of the musical performance. Thanks to a discretization of the tempo set, it allows for an efficient decoding by the Viterbi algorithm, and facilitates the introduction of features which directly depend on the local tempo. We take advantage of this property by using the cyclic tempogram descriptor in addition to chroma vectors and onset detection features. Experiment run on both classical piano and pop music show the very high accuracy of this model for audio to score alignment, as well as the usefulness of die tempo feature used.
Cyril Joder, Slim Essid, Gaël Richard
ICASSP2
2011 An audio-driven virtual dance-teaching assistant
abstract
This work addresses the Huawei/3Dlife Grand challenge proposing a set of audio tools for a virtual dance-teaching assistant. These tools are meant to help the dance student develop a sense of rhythm to correctly synchronize his/her movements and steps to the musical timing of the choreographies to be executed. They consist of three main components, namely a music (beat) analysis module, a source separation and remastering module and a dance step segmentation module. These components enable to create augmented tutorial videos highlighting the rhythmic information using, for instance, a synthetic dance teacher voice, but also videos highlighting the steps executed by a student to help in the evaluation of his/her performance.
Slim Essid, Yves Grenier, Mounira Maazaoui, Gaël Richard, Robin Tournemenne
ACM Multimedia1
2011 Enhanced visualisation of dance performance from automatically synchronised multimodal recordings
abstract
The Huawei/3DLife Grand Challenge Dataset provides multimodal recordings of Salsa dancing, consisting of audiovisual streams along with depth maps and inertial measurements. In this paper, we propose a system for augmented reality-based evaluations of Salsa dancer performances. An essential step for such a system is the automatic temporal synchronisation of the multiple modalities captured from different sensors, for which we propose efficient solutions. Furthermore, we contribute modules for the automatic analysis of dance performances and present an original software application, specifically designed for the evaluation scenario considered, which enables an enhanced dance visualisation experience, through the augmentation of the original media with the results of our automatic analyses.
Marc Gowing, Philip Kelly, Noel E. O'Connor, Cyril Concolato, Slim Essid, Jean Le Feuvre, Robin Tournemenne, Ebroul Izquierdo, Vlado Kitanovski, Qianni Zhang
ACM Multimedia5
2011 A Conditional Random Field Framework for Robust and Scalable Audio-to-Score Matching
abstract
In this paper, we introduce the use of conditional random fields (CRFs) for the audio-to-score alignment task. This framework encompasses the statistical models which are used in the literature and allows for more flexible dependency structures. In particular, it allows observation functions to be computed from several analysis frames. Three different CRF models are proposed for our task, for different choices of tradeoff between accuracy and complexity. Three types of features are used, characterizing the local harmony, note attacks and tempo. We also propose a novel hierarchical approach, which takes advantage of the score structure for an approximate decoding of the statistical model. This strategy reduces the complexity, yielding a better overall efficiency than the classic beam search method used in HMM-based models. Experiments run on a large database of classical piano and popular music exhibit very accurate alignments. Indeed, with the best performing system, more than 95% of the note onsets are detected with a precision finer than 100 ms. We additionally show how the proposed framework can be modified in order to be robust to possible structural differences between the score and the musical performance.
Cyril Joder, Slim Essid, Gaël Richard
IEEE ACM Trans. Audio Speech Lang. Process.2
2010 A comparative study of tonal acoustic features for a symbolic level music-to-score alignment
abstract
In this paper we review the acoustic features used for music-to-score alignment and study their influence on the performance in a challenging alignment task, where the audio data is polyphonic and may contain percussion. Furthermore, as we aim at using “real world” scores, we follow an approach which does exploit the rhythm information (considered unreliable) and test its robustness to score errors. We use a unified framework to handle different state-of-the-art features, and propose a simple way to exploit either a model of the feature values, or an audio synthesis of a musical score, in an audio-to-score alignment system. We confirm that chroma vectors drawn from representations using a logarithmic frequency scale are the most efficient features, and lead to a good precision, even with a simple alignment strategy. Robustness tests also show that the relative performance of the features do not depend on possible musical score degradations.
Cyril Joder, Slim Essid, Gaël Richard
ICASSP2
2010 Robust visual features for the multimodal identification of unregistered speakers in TV talk-shows
abstract
In this paper we propose a novel multimodal method for identifying unregistered speakers in a TV talk-show using a semi-supervised learning approach based on Support Vector Machines. Our study highlights the fact that specific visual features prove to be very efficient for this particular type of video content which is edited from multi-camera recordings. These visual features, motivated by prior knowledge on the approach followed by the TV director in choosing the appropriate shots, are found to bring a significant improvement in identification accuracy when used together with classic audio Mel-frequency cepstral coefficients (+8% compared to various baseline systems, in particular a standard audio only system).
Félicien Vallet, Slim Essid, Jean Carrive, Gaël Richard
ICIP2
2010 A conditional random field viewpoint of symbolic audio-to-score matching
abstract
We present a new approach of symbolic audio-to-score alignment, with the use of Conditional Random Fields (CRFs). Unlike Hidden Markov Models, these graphical models allow the calculation of state conditional probabilities to be made on the basis of several audio frames. The CRF models that we propose exploit this property to take into account the rhythmic information of the musical score. Assuming that the tempo is locally constant, they confront the neighborhood of each frame with several tempo hypotheses.
Cyril Joder, Slim Essid, Gaël Richard
ACM Multimedia2
2009 Incorporating prior knowledge on the digital media creation process into audio classifiers
abstract
In the process of music content creation, a wide range of typical audio effects such as reverberation, equalization or dynamic compression are very commonly used. Despite the fact that such effects have a clear impact on the audio features, they are rarely taken into account when building an automatic audio classifier. In this paper, it is shown that the incorporation of prior knowledge of the digital media creation chain can clearly improve the robustness of the audio classifiers, which is demonstrated on a task of musical instrument recognition. The proposed system is based on a robust feature selection strategy, on a novel use of the virtual support vector machines technique and a specific equalization used to normalize the signals to be classified. The robustness of the proposed system is experimentally evidenced using a rather large and varied sound database.
Maxime Lardeur, Slim Essid, Gaël Richard, Martin Haller, Thomas Sikora
ICASSP2
2009 Temporal Integration for Audio Classification With Application to Musical Instrument Classification
abstract
Nowadays, it appears essential to design automatic indexing tools which provide meaningful and efficient means to describe the musical audio content. There is in fact a growing interest for music information retrieval (MIR) applications amongst which the most popular are related to music similarity retrieval, artist identification, musical genre or instrument recognition. Current MIR-related classification systems usually do not take into account the mid-term temporal properties of the signal (over several frames) and lie on the assumption that the observations of the features in different frames are statistically independent. The aim of this paper is to demonstrate the usefulness of the information carried by the evolution of these characteristics over time. To that purpose, we propose a number of methods for early and late temporal integration and provide an in-depth experimental study on their interest for the task of musical instrument recognition on solo musical phrases. In particular, the impact of the time horizon over which the temporal integration is performed will be assessed both for fixed and variable frame length analysis. Also, a number of proposed alignment kernels will be used for late temporal integration. For all experiments, the results are compared to a state of the art musical instrument recognition system.
Cyril Joder, Slim Essid, Gaël Richard
IEEE Trans. Speech Audio Process.2
2008 A collaborative approach to automatic rushes video summarization
abstract
Video summarization is a useful tool which allows a user to grasp rapidly the essence of a video. In the development of this research topic we propose a new method based on different individual content segmentation and selection tools in a collaborative system. The main innovation of this work is to merge results from different approaches, so as to benefit from their respective qualities. Our system is organized in two phases: first segmentation of the video, second identification of relevant and redundant segments. The final list of selected segments is used to concatenate the video segments and build the final summary. In order to assess the effectiveness of this organization, we evaluate our system with a method based on the TRECVID 2007 BBC rushes summarization evaluation pilot and compare our performance with existing systems.
Werner Bailer, Emilie Dumont, Slim Essid, Bernard Mérialdo
ICIP3
2007 Combined Supervised and Unsupervised Approaches for Automatic Segmentation of Radiophonic Audio Streams
abstract
Speech/music discrimination is one of the most studied topics in the domain of audio data segmentation. In this paper, we propose and evaluate a novel method that includes feature selection and a combined supervised and unsupervised strategy for audio streams segmentation. A number of alternatives solutions for each component are assessed and the optimized system is compared to the approaches proposed in the framework of the ESTER campaign.
Gaël Richard, Mathieu Ramona, Slim Essid
ICASSP (2)3
2007 On the Correlation of Automatic Audio and Visual Segmentations of Music Videos
abstract
The study of the associations between audio and video content has numerous important applications in the fields of information retrieval and multimedia content authoring. In this work, we focus on music videos which exhibit a broad range of structural and semantic relationships between the music and the video content. To identify such relationships, a two-level automatic structuring of the music and the video is achieved separately. Note onsets are detected from the music signal, along with section changes. The latter is achieved by a novel algorithm which makes use of feature selection and statistical novelty detection approaches based on kernel methods. The video stream is independently segmented to detect changes in motion activity, as well as shot boundaries. Based on this two-level segmentation of both streams, four audio–visual correlation measures are computed. The usefulness of these correlation measures is illustrated by a query by video experiment on a 100 music video database, which also exhibits interesting genre dependencies.
Olivier Gillet, Slim Essid, Gaël Richard
IEEE Trans. Circuits Syst. Video Technol.2
2006 Hierarchical Classification of Musical Instruments on Solo Recordings
abstract
We propose a study on the use of hierarchical taxonomies for musical instrument recognition on solo recordings. Both a natural taxonomy (inspired by instrument families) and a taxonomy inferred automatically by means of hierarchical clustering are examined. They are used to build a hierarchical classification scheme based on support vector machine classifiers and an efficient selection of features from a wide set of candidate descriptors. The classification results found with each taxonomy are compared and analysed. The automatic taxonomy is found to perform slightly better than the "natural" one. However, our analysis of the confusion matrices related to these taxonomies suggest that both are limited. In fact, it shows that it could be more advantageous to utilise taxonomies such that the instruments which are commonly confused are put in distinct decision nodes
Slim Essid, Gaël Richard, Bertrand David 0002
ICASSP (5)1
2006 Instrument recognition in polyphonic music based on automatic taxonomies
abstract
We propose a new approach to instrument recognition in the context of real music orchestrations ranging from solos to quartets. The strength of our approach is that it does not require prior musical source separation. Thanks to a hierarchical clustering algorithm exploiting robust probabilistic distances, we obtain a taxonomy of musical ensembles which is used to efficiently classify possible combinations of instruments played simultaneously. Moreover, a wide set of acoustic features is studied including some new proposals. In particular, signal to mask ratios are found to be useful features for audio classification. This study focuses on a single music genre (i.e., jazz) but combines a variety of instruments among which are percussion and singing voice. Using a varied database of sound excerpts from commercial recordings, we show that the segmentation of music with respect to the instruments played can be achieved with an average accuracy of 53%.
Slim Essid, Gaël Richard, Bertrand David 0002
IEEE Trans. Speech Audio Process.1
2006 Musical instrument recognition by pairwise classification strategies
abstract
Musical instrument recognition is an important aspect of music information retrieval. In this paper, statistical pattern recognition techniques are utilized to tackle the problem in the context of solo musical phrases. Ten instrument classes from different instrument families are considered. A large sound database is collected from excerpts of musical phrases acquired from commercial recordings translating different instrument instances, performers, and recording conditions. More than 150 signal processing features are studied including new descriptors. Two feature selection techniques, inertia ratio maximization with feature space projection and genetic algorithms are considered in a class pairwise manner whereby the most relevant features are fetched for each instrument pair. For the classification task, experimental results are provided using Gaussian mixture models (GMMs) and support vector machines (SVMs). It is shown that higher recognition rates can be reached with pairwise optimized subsets of features in association with SVM classification using a radial basis function kernel
Slim Essid, Gaël Richard, Bertrand David 0002
IEEE Trans. Speech Audio Process.1
2005 Instrument recognition in polyphonic music
abstract
We propose a method for the recognition of musical instruments in polyphonic music excerpted from commercial recordings. By exploiting some cues on the common structures of musical ensembles, we show that it is possible to recognize up to 4 instruments playing concurrently. The system associates a hierarchical classification tree with a class-pairwise feature selection technique and Gaussian mixture models to discriminate possible combinations of instruments. Successful identification is achieved over short-time windows, enabling the system to be employed for segmentation purposes.
Slim Essid, Gaël Richard, Bertrand David 0002
ICASSP (3)1
2002 Dynamic temporal segmentation in parametric non-stationary modeling for percussive musical signals
abstract
An audio signal parametric modeling scheme is proposed that permits higher performance for representing strong sound transients. The exponentially damped sinusoids (EDS) model is considered in association with a high resolution parameter estimation approach. Such a technique is well adapted to almost every audio signal but is unfortunately not efficient when dealing with signals presenting strong temporal variations, such as percussive music signals, and causes pre-echo artifacts and weak onset dynamic reproduction which are prejudicial to listening. A system, based on the EDS model, has been developed with a transient detector and dynamic time segmentation and modeling that allows to overcome such artifacts.
Rémy Boyer, Slim Essid, Nicolas Moreau
ICME (1)2