EDBT 2026 Demo / reviewers in the wild / expert
Jean-Luc Rouas
dblp:69/2845
· DBLP profile ↗
34ranked-venue papers
11as first author
14since 2021 · last 2026
0000-0003-1933-0504ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 9 first-author · 8 since 2021Artificial intelligence and machine learning · 18 · 6 first-author · 8 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Medispeech: A French Reading and Spontaneous Speech Corpus for Sleepiness EstimationabstractInternational audience Colleen Beaumard, Vincent Martin 0002, Charles Brazier, Julien Coelho, Jean-Luc Rouas, Pierre Philip |
LREC | 5 |
| 2026 | SOMVOICE: A First Dataset to Study the Effects of Sleep Deprivation on Voice Characteristics of Healthy French SpeakersabstractInternational audience Vincent Martin 0002, Jean-Luc Rouas, Colleen Beaumard, Pierre Philip |
LREC | 2 |
| 2025 | Systolic Arrays and Structured Pruning Co-design for Efficient Transformers in Edge SystemsabstractInternational audience Pedro Palacios, Rafael Medina 0001, Jean-Luc Rouas, Giovanni Ansaloni, David Atienza 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | Adaptive Compression of Supervised and Self-Supervised Models for Green Speech RecognitionabstractComputational power is crucial for the development and deployment of artificial intelligence capabilities, as the large size of deep learning models often requires significant resources. Compression methods aim to reduce model size making artificial intelligence more sustainable and accessible. Compression techniques are often applied uniformly across model layers, without considering their individual characteristics. In this paper, we introduce a customized approach that optimizes compression for each layer individually. Some layers undergo both pruning and/or quantization, while others are only quantized, with fuzzy logic guiding these decisions. The quantization precision is further adjusted based on the importance of each layer. Our experiments on both supervised and self-supervised models using the librispeech dataset show only a slight decrease in performance, with about 85% memory footprint reduction. Mouaad Oujabour, Leila Ben Letaifa, Jean-François Dollinger, Jean-Luc Rouas |
ICASSP | 4 |
| 2025 | Network of acoustic characteristics for the automatic detection of suicide risk from speech. Contribution to the 2025 SpeechWellness challenge by the Semawave teamabstractInternational audience Vincent Martin 0002, Charles Brazier, Maxime Amblard, Michel Musiol, Jean-Luc Rouas |
INTERSPEECH | 5 |
| 2025 | Structured pruning for efficient systolic array accelerated cascade Speech-to-Text Translation
Jean-Luc Rouas, Charles Brazier, Leila Ben Letaifa, Rafael Medina 0001, Pedro Palacios, David Atienza 0001, Giovanni Ansaloni |
INTERSPEECH | 1 |
| 2024 | Why Voice Biomarkers of Psychiatric Disorders Are Not Used in Clinical Practice? Deconstructing the Myth of the Need for Objective DiagnosisabstractGiven the high prevalence of mental disorders and the significant diagnostic delays and difficulties in patient follow-up, voice biomarkers hold the promise of improving access to care and therapeutic follow-up for people with psychiatric disorders. Yet, despite many years of successful research in the field, none of these voice biomarkers are implemented in clinical practice. Beyond the reductive explanation of the lack of explainability of the involved machine learning systems, we look for arguments in the epistemology and sociology of psychiatry. We show that the estimation of diagnoses, the major task in the literature, is of little interest to both clinicians and patients. After tackling the common misbeliefs about diagnosis in psychiatry in a didactic way, we propose a paradigm shift towards the estimation of clinical symptoms and signs, which not only address the limitations raised against diagnosis estimation but also enable the formulation of new machine learning tasks. We hope that this paradigm shift will empower the use of vocal biomarkers in clinical practice. It is however conditional on a change in database labeling practices, but also on a profound change in the speech processing community’s practices towards psychiatry. Vincent Martin 0002, Jean-Luc Rouas |
LREC/COLING | 2 |
| 2024 | FVLLMONTI: The 3D Neural Network Compute Cube $(N^{2}C^{2})$ Concept for Efficient Transformer Architectures Towards Speech-to-Speech TranslationabstractThis multi-partner-project contribution introduces the midway results of the Horizon 2020 FVLLMONTI project. In this project we develop a new and ultra-efficient class of ANN accelerators, the neural network compute cube$(N^{2}C^{2})$, which is specifically designed to execute complex machine learning tasks in a 3D technology, in order to provide the high computing power and ultra-high efficiency needed for future edgeAI applications. We showcase its effectiveness by targeting the challenging class of Transformer ANNs, tailored for Automatic Speech Recognition and Machine Translation, the two fundamental components of speech-to-speech translation. To gain the full benefit of the accelerator design, we develop disruptive vertical transistor technologies and execute design-technology-co-optimization (DTCO) loops from single device, to cell and compute cube level. Further, a hardware-software-co-optimization is executed, e.g. by compressing the executed speech recognition and translation models for energy efficient executing without substantial loss in precision. Ian O'Connor, Sara Mannaa, Alberto Bosio, Bastien Deveautour, Damien Deleruyelle, Tetiana Obukhova, Cédric Marchand 0002, Jens Trommer, Çigdem Çakirlar, Bruno Neckel Wesling, Thomas Mikolajick, Oskar Baumgartner, Mischa Thesberg, David Pirker, Christoph Lenz, Zlatan Stanojevic, Markus Karner, Guilhem Larrieu, Sylvain Pelloquin, Konstantinous Moustakas, Giovanni Ansaloni, Alireza Amirshahi, David Atienza 0001, Jean-Luc Rouas, Leila Ben Letaifa, Georgeta Bordeall, Charles Brazier, C. Mukherjee 0001, Marina Deng, Marc François, Houssem Rezgui, Reveil Lucas, Cristell Maneux |
DATE | 25 |
| 2024 | Estimating Symptoms and Clinical Signs Instead of Disorders: The Path Toward The Clinical Use of Voice and Speech Biomarkers In PsychiatryabstractDespite the continuous innovation in voice biomarkers domain for more than a decade and the apparent need for clinicians to have objective diagnostic tools, no device has yet been implemented in real clinical settings or widely adopted by clinicians. After giving a short overview of the literature, we argue that in addition to the factors usually mentioned in the literature (low performance, database sizes, transparency, etc.), an underestimated but crucial factor preventing the use of such systems is the therapeutic relationship. We also discuss the "objectivity" of such systems, and the place of diagnosis in clinical practice and its conceptual limitations. In order to shape useful and relevant voice biomarkers, we propose to estimate symptoms instead of diagnosis, and draw perspectives related to this paradigm, which will require databases annotated with patients’ symptoms rather than only their pathological status. Vincent Martin 0002, Jean-Luc Rouas |
ICASSP | 2 |
| 2024 | Automatic Detection Of Sleepiness-Related Syndromes and Symptoms Using Voice and Speech BiomarkersabstractThis article is about the automatic estimation of sleepiness in hypersomnia patients recorded during a reading task. Based on the Multiple Sleep Latency Corpus, our main contribution is to explore new formulations of sleepiness detection in speech by specifying and performing five sleepiness-related classification tasks. We automatically classify three symptoms, and two syndromes, i.e. combinations of symptoms that are closer to clinical reasoning. Another contribution of this paper is the use of a simple and interpretable pipeline integrating selecting voice biomarkers of sleepiness, i.e. features that are both sensible and specific to sleepiness. In particular, specificity is adressed integrating a decorrelation step in the pipeline, which allows to certify that the descriptors selected by the pipeline are indeed specific of sleepiness with respect to 7 cofactors (age, BMI, etc.). Vincent Martin 0002, Jean-Luc Rouas, Pierre Philip |
ICASSP | 2 |
| 2023 | "Prediction of Sleepiness Ratings from Voice by Man and Machine": A Perceptual Experiment Replication StudyabstractFollowing the release of the SLEEP corpus during the Interspeech 2019 paralinguistic continuous sleepiness estimation challenge, a paper presented at Interspeech 2020 by Huckvale et al. examined the reasons for the poor performance of the models proposed for this task. Careful analyses of the corpus led to the conclusion that its bias makes it hazardous to use for training machine learning systems, but a perceptual experiment on a subset of this corpus seemed to indicate that human hearing is however able to estimate sleepiness on this corpus.In this study, we present the results of the Endymion replication study, in which the same samples were rated by thirty French-speaking naive listeners. We then discuss the causes of the differences between the two studies and examine the effect of listener and sample characteristics on annotation performances. Vincent Martin 0002, Aymeric Ferron, Jean-Luc Rouas, Pierre Philip |
ICASSP | 3 |
| 2023 | Affective attributes of French caregivers' professional speechabstractInternational audience Jean-Luc Rouas, Yaru Wu, Takaaki Shochi |
INTERSPEECH | 1 |
| 2022 | Fine-grained analysis of the transformer model for efficient pruningabstractIn automatic speech recognition, deep learning models such as transformers are increasingly used for their high performance. However, they suffer from their large size, which makes it very difficult to use them in real contexts. Hence the idea of pruning them. Conventional pruning methods are not optimal and sometimes not efficient since they operate blindly without taking into account the nature of the layers or their number of parameters or their distribution. In this work, we propose to perform a fine-grained analysis of the transformer model layers in order to determine the most efficient pruning approach. We show that it is more appropriate to prune some layers than others and underline the importance of knowing the behavior of the layers to choose the pruning approach. Leila Ben Letaifa, Jean-Luc Rouas |
ICMLA | 2 |
| 2021 | Automatic Speech Recognition Systems Errors for Objective Sleepiness Detection Through VoiceabstractInternational audience Vincent Martin 0002, Jean-Luc Rouas, Florian Boyer, Pierre Philip |
Interspeech | 2 |
| 2020 | The Objective and Subjective Sleepiness Voice CorporaabstractFollowing patients with chronic sleep disorders involves multiple appointments between doctors and patients which often results in episodic follow-ups with unevenly spaced interviews. Speech technologies and virtual doctors can help improve this follow-up. However, there are still some challenges to overcome: sleepiness measurements are diverse and are not always correlated, and most past research focused on detecting nstantaneous sleepiness levels of healthy sleep-deprived subjects. This article presents a large database to assess the sleepiness level of highly phenotyped patients that complain from excessive daytime sleepiness. Based on the Multiple Sleep Latency Test, it differs from existing databases by multiple aspects. First, it is omposed of recordings from patients suffering from excessive daytime sleepiness instead of sleep deprived healthy subjects. Second, it incites the subjects to sleep contrary to existing stressing sleepiness deprivation experimental paradigms. Third, the sleepiness level of the patients is evaluated with different temporal granularities - long term sleepiness and short term sleepiness - and both objective and subjective sleepiness measures are collected. Finally, it relies on the recordings of 94 highly phenotyped patients, allowing to unravel the influences of different physical factors (age, sex, weight, ... ) on voice. Vincent Martin 0002, Jean-Luc Rouas, Jean-Arthur Micoulaud-Franchi, Pierre Philip |
LREC | 2 |
| 2018 | Cultural Differences in Pattern Matching: Multisensory Recognition of Socio-affective ProsodyabstractInternational audience Takaaki Shochi, Jean-Luc Rouas, Marine Guerry, Donna Erickson |
INTERSPEECH | 2 |
| 2018 | Regularized Optimal Transport and the Rot Mover's DistanceabstractThis paper presents a unified framework for smooth convex regularization of discrete optimal transport problems. In this context, the regularized optimal transport turns out to be equivalent to a matrix nearness problem with respect to Bregman divergences. Our framework thus naturally generalizes a previously proposed regularization based on the Boltzmann-Shannon entropy related to the Kullback-Leibler divergence, and solved with the Sinkhorn-Knopp algorithm. We call the regularized optimal transport distance the rot mover's distance in reference to the classical earth mover's distance. By exploiting alternate Bregman projections, we develop the alternate scaling algorithm and non-negative alternate scaling algorithm, to compute efficiently the regularized optimal plans depending on whether the domain of the regularizer lies within the non-negative orthant or not. We further enhance the separable case with a sparse extension to deal with high data dimensions. We also instantiate our framework and discuss the inherent specificities for well-known regularizers and statistical divergences in the machine learning and information geometry communities. Finally, we demonstrate the merits of our methods with experiments using synthetic data to illustrate the effect of different regularizers, penalties and dimensions, as well as real-world data for a pattern recognition application to audio scene classification. Arnaud Dessein, Nicolas Papadakis, Jean-Luc Rouas |
J. Mach. Learn. Res. | 3 |
| 2016 | Automatic Classification of Phonation Modes in Singing Voice: Towards Singing Style Characterisation and Application to Ethnomusicological RecordingsabstractInternational audience Jean-Luc Rouas, Leonidas Ioannidis |
INTERSPEECH | 1 |
| 2015 | Synthetic Evidential Study as Augmented Collective Thought Process - Preliminary Report
Toyoaki Nishida, Masakazu Abe, Takashi Ookaki, Divesh Lala, Sutasinee Thovutikul, Hengjie Song, Yasser Mohammad, Christian Nitschke, Yoshimasa Ohmoto, Atsushi Nakazawa, Takaaki Shochi, Jean-Luc Rouas, Aurélie Bugeau, Fabien Lotte, Zuheng Ming, Geoffrey Letournel, Marine Guerry, Dominique Fourer |
ACIIDS (1) | 12 |
| 2013 | Feedback-based gameplay metrics
Raphaël Marczak, Gareth Schott, Pierre Hanna, Jean-Luc Rouas |
FDG | 4 |
| 2011 | In Search of Cues Discriminating West-African Accents in FrenchabstractInternational audience Philippe Boula de Mareüil, Jean-Luc Rouas, Manuela Yapomo |
INTERSPEECH | 2 |
| 2010 | Comparison of Spectral Properties of Read, Prepared and Casual Speech in French
Jean-Luc Rouas, Mayumi Beppu, Martine Adda-Decker |
LREC | 1 |
| 2008 | Portuguese variety identification on broadcast newsabstractThis paper describes an accent identification system for Portuguese, that explores different type of properties: acoustic, phono tactic and prosodic. The system is designed to be used as a pre-processing module for the Portuguese automatic speech recognition system developed at INESC-ID. In terms of variety identification, the overall rate of correct identification is 69.0% if all 7 varieties are considered, and the best results are obtained for Brazilian Portuguese, also the variety that proved easiest to identify in perceptual experiments. When distinguishing between European, Brazilian and African Portuguese, the identification rate goes up to 94.7%. The fact that the prosodic system alone can achieve an identification rate of 77% is also worth investigating. Jean-Luc Rouas, Isabel Trancoso, Céu Viana, Mónica Abreu |
ICASSP | 1 |
| 2008 | Language and variety verification on broadcast news for Portuguese
Jean-Luc Rouas, Isabel Trancoso, Céu Viana, Mónica Abreu |
Speech Commun. | 1 |
| 2007 | Automatic Prosodic Variations Modeling for Language and Dialect DiscriminationabstractThis paper addresses the problem of modeling prosody for language identification. The aim is to create a system that can be used prior to any linguistic work to show if prosodic differences among languages or dialects can be automatically determined. In previous papers, we defined a prosodic unit, the pseudosyllable. Rhythmic modeling has proven the relevance of the pseudosyllable unit for automatic language identification. In this paper, we propose to model the prosodic variations, that is to say model sequences of prosodic units. This is achieved by the separation of phrase and accentual components of intonation. We propose an independent coding of those components on differentiated scales of duration. Short-term and long-term language-dependent sequences of labels are modeled by n-gram models. The performance of the system is demonstrated by experiments on read speech and evaluated by experiments on spontaneous speech. Finally, an experiment is described on the discrimination of Arabic dialects, for which there is a lack of linguistic studies, notably on prosodic comparisons. We show that our system is able to clearly identify the dialectal areas, leading to the hypothesis that those dialects have prosodic differences. Jean-Luc Rouas |
IEEE Trans. Speech Audio Process. | 1 |
| 2005 | Modeling long and short-term prosody for language identificationabstractInternational audience Jean-Luc Rouas |
INTERSPEECH | 1 |
| 2005 | Rhythmic unit extraction and modelling for automatic language identification
Jean-Luc Rouas, Jérôme Farinas, François Pellegrino, Régine André-Obrecht |
Speech Commun. | 1 |
| 2004 | Fusing language identification systems using performance confidence indexesabstractIn the field of automatic language identification, several, mostly-empirical, arithmetic fusion operations are currently done to make a consensus decision from a set of acoustics-based identification systems whose estimated performance is taken into account by means of weighting techniques. The paper presents how to apply the discriminant factor analysis method formally to compute and use weighting performance confidence indexes at the expert and class levels. Moreover, the observation level is also explored. These confidence indexes allow us not only to qualify the identification decision with additional insight by means of a certainty degree, but also to provide acoustics-based identification systems with powerful uncertainty-based inference techniques where the systems' a priori performance knowledge is a key heuristic-like element to improve language identification capabilities. Jorge Gutiérrez, Jean-Luc Rouas, Régine André-Obrecht |
ICASSP (1) | 2 |
| 2003 | A fusion study in speech/music classificationabstractWe present and merge two speech/music classification approaches that we have developed. The first one is a differentiated modeling approach based on a spectral analysis, which is implemented using GMM (Gaussian mixture model). The other one is based on three original features: entropy modulation, stationary segment duration and number of segments. They are merged with the classical 4 Hertz modulation energy. Our classification system is a fusion of the two approaches. It is divided in two classifications (speech/non-speech and music/non-music) and provides 94% of accuracy for speech detection and 90% for music detection, with one second of input signal. Beside the spectral information and GMM, classically used in speech/music discrimination, simple parameters bring complementary and efficient information. Julien Pinquier, Jean-Luc Rouas, Régine André-Obrecht |
ICASSP (2) | 2 |
| 2003 | Modeling prosody for language identification on read and spontaneous speechabstractInternational audience Jean-Luc Rouas, Jérôme Farinas, François Pellegrino, Régine André-Obrecht |
ICASSP (1) | 1 |
| 2003 | A fusion study in speech / music classificationabstractIn this paper, we present and merge two speech / music classification approaches of that we have developed. The first one is a differentiated modeling approach based on a spectral analysis, which is implemented with GMM. The other one is based on three original features: entropy modulation, stationary segment duration and number of segments. They are merged with the classical 4 Hertz modulation energy. Our classification system is a fusion of the two approaches. It is divided in two classifications (speech/non-speech and music/non-music) and provides 94 % of accuracy for speech detection and 90 % for music detection, with one second of input signal. Beside the spectral information and GMM, classically used in speech / music discrimination, simple parameters bring complementary and efficient information. Julien Pinquier, Jean-Luc Rouas, Régine André-Obrecht |
ICME | 2 |
| 2003 | Modeling prosody for language identification on read and spontaneous speechabstractThis paper deals with an approach to automatic language identification using only prosodic modeling. The actual approach for language identification focuses mainly on phonotactics because it gives the best results. We propose here to evaluate the relevance of prosodic information for language identification with read studio recording and spontaneous telephone speech. For read speech, experiments were performed on the five languages of the MULTEXT database [E. Campoine and J. Veronis, Nov. 1998]. On the MULTEXT corpus, our prosodic system achieved an identification rate of 79% on the five languages discrimination task. For spontaneous speech, experiments are made on the ten languages of the OGI multilingual telephone speech corpus [Y. K. Muthusamy et al., October 1992]. On the OGI MLTS corpus, the results are given for languages pair discrimination tasks, and are compared with results from [F. Cummins et al., 1999]. As a conclusion, if our prosodic system achieves good performance on read speech, it might not take into account the complexity of spontaneous speech prosody. Jean-Luc Rouas, Jérôme Farinas, François Pellegrino, Régine André-Obrecht |
ICME | 1 |
| 2002 | Merging segmental and rhythmic features for Automatic Language IdentificationabstractThis paper deals with an approach to Automatic Language Identification based on rhythmic modeling and vowel system modeling. Experiments are performed on read speech for 5 European languages. They show that rhythm and stress may be automatically extracted and are relevant in language identification: using cross-validation, 78% of correct identification is reached with 21 s. utterances. The Vowel System Modeling, tested in the same conditions (cross-validation), is efficient and results in a 70% of correct identification for the 21 s. utterances. Last. merging the two models slightly improves the results. Jérôme Farinas, François Pellegrino, Jean-Luc Rouas, Régine André-Obrecht |
ICASSP | 3 |
| 2002 | Robust speech / music classification in audio documentsabstractInternational audience Julien Pinquier, Jean-Luc Rouas, Régine André-Obrecht |
INTERSPEECH | 2 |