VLDB 2026 Research / reviewers in the wild / expert
Mohamed Salah Zaïem
dblp:254/0980 · also Salah Zaiem
· DBLP profile ↗
14ranked-venue papers
6as first author
12since 2021 · last 2025
0000-0002-0788-3219ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Speech self-supervised representations benchmarking: A case for larger probing heads
Mohamed Salah Zaïem, Youcef Kemiche, Titouan Parcollet, Slim Essid, Mirco Ravanelli |
Comput. Speech Lang. | 1 |
| 2024 | TARIC-SLU: A Tunisian Benchmark Dataset for Spoken Language UnderstandingabstractIn recent years, there has been a significant increase in interest in developing Spoken Language Understanding (SLU) systems. SLU involves extracting a list of semantic information from the speech signal. A major issue for SLU systems is the lack of sufficient amount of bi-modal (audio and textual semantic annotation) training data. Existing SLU resources are mainly available in high-resource languages such as English, Mandarin and French. However, one of the current challenges concerning low-resourced languages is data collection and annotation. In this work, we present a new freely available corpus, named TARIC-SLU, composed of railway transport conversations in Tunisian dialect that is continuously annotated in dialogue acts and slots. We describe the semantic model of the dataset, the data and experiments conducted to build ASR-based and SLU-based baseline models. To facilitate its use, a complete recipe, including data preparation, training and evaluation scripts, has been built and will be integrated to SpeechBrain, a popular open-source conversational AI toolkit based on PyTorch. Salima Mdhaffar, Fethi Bougares, Renato De Mori, Mohamed Salah Zaïem, Mirco Ravanelli, Yannick Estève |
LREC/COLING | 4 |
| 2024 | Leveraging Data Collection and Unsupervised Learning for Code-Switched Tunisian Arabic Automatic Speech RecognitionabstractCrafting an effective Automatic Speech Recognition (ASR) solution for dialects demands innovative approaches that not only address the data scarcity issue but also navigate the intricacies of linguistic diversity. In this paper, we address the aforementioned ASR challenge, focusing on the Tunisian dialect. First, textual and audio data is collected and in some cases annotated. Second, we explore self-supervision, semi-supervision and few-shot code-switching approaches to push the state-of-the-art on different Tunisian test sets; covering different acoustic, linguistic and prosodic conditions. Finally, and given the absence of conventional spelling, we produce a human evaluation of our transcripts to avoid the noise coming from spelling inadequacies in our testing references. Our models, allowing to transcribe audio samples in a linguistic mix involving Tunisian Arabic, English and French, and all the data used during training and testing are released for public use and further improvements. Ahmed Amine Ben Abdallah, Ata Kabboudi, Amir Kanoun, Mohamed Salah Zaïem |
ICASSP | 4 |
| 2024 | How Should We Extract Discrete Audio Tokens from Self-Supervised Models?abstractDiscrete audio tokens have recently gained attention for their potential to bridge the gap between audio and language processing.Ideal audio tokens must preserve content, paralinguistic elements, speaker identity, and many other audio details.Current audio tokenization methods fall into two categories: Semantic tokens, acquired through quantization of Self-Supervised Learning (SSL) models, and Neural compression-based tokens (codecs).Although previous studies have benchmarked codec models to identify optimal configurations, the ideal setup for quantizing pretrained SSL models remains unclear.This paper explores the optimal configuration of semantic tokens across discriminative and generative tasks.We propose a scalable solution to train a universal vocoder across multiple SSL layers.Furthermore, an attention mechanism is employed to identify task-specific influential layers, enhancing the adaptability and performance of semantic tokens in diverse audio applications. Pooneh Mousavi, Jarod Duret, Mohamed Salah Zaïem, Luca Della Libera, Artem Ploujnikov, Cem Subakan, Mirco Ravanelli |
INTERSPEECH | 3 |
| 2024 | Open-Source Conversational AI with SpeechBrain 1.0abstractSpeechBrain is an open-source Conversational AI toolkit based on PyTorch, focused particularly on speech processing tasks such as speech recognition, speech enhancement, speaker recognition, text-to-speech, and much more. It promotes transparency and replicability by releasing both the pre-trained models and the complete recipes of code and algorithms required for training them. This paper presents SpeechBrain 1.0, a significant milestone in the evolution of the toolkit, which now has over 200 recipes for speech, audio, and language processing tasks, and more than 100 models available on Hugging Face. SpeechBrain 1.0 introduces new technologies to support diverse learning modalities, Large Language Model (LLM) integration, and advanced decoding strategies, along with novel models, tasks, and modalities. It also includes a new benchmark repository, offering researchers a unified platform for evaluating models across diverse tasks. Mirco Ravanelli, Titouan Parcollet, Adel Moumen, Sylvain de Langen, Cem Subakan, Peter Plantinga, Yingzhi Wang 0002, Pooneh Mousavi, Luca Della Libera, Artem Ploujnikov, Francesco Paissan, Davide Borra, Mohamed Salah Zaïem, Zeyu Zhao 0004, Shucong Zhang, Georgios Karakasidis, Sung-Lin Yeh, Pierre Champion, Aku Rouhe, Rudolf Braun, Florian Mai, Juan Zuluaga-Gomez, Seyed Mahed Mousavi, Andreas Nautsch, Xuechen Liu 0001, Sangeet Sagar, Jarod Duret, Salima Mdhaffar, Gaëlle Laperrière, Mickael Rouvier, Renato De Mori, Yannick Estève |
J. Mach. Learn. Res. | 13 |
| 2024 | CL-MASR: A Continual Learning Benchmark for Multilingual ASRabstractModern multilingual automatic speech recognition (ASR) systems like Whisper have made it possible to transcribe audio in multiple languages with a single model. However, current state-of-the-art ASR models are typically evaluated on individual languages or in a multi-task setting, overlooking the challenge of continually learning new languages. There is insufficient research on how to add new languages without losing valuable information from previous data. Furthermore, existing continual learning benchmarks focus mostly on vision and language tasks, leaving continual learning for multilingual ASR largely unexplored. To bridge this gap, we propose CL-MASR, a benchmark designed for studying multilingual ASR in a continual learning setting. CL-MASR provides a diverse set of continual learning methods implemented on top of large-scale pretrained ASR models, along with common metrics to assess the effectiveness of learning new languages while addressing the issue of catastrophic forgetting. To the best of our knowledge, CL-MASR is the first continual learning benchmark for the multilingual ASR task. Luca Della Libera, Pooneh Mousavi, Mohamed Salah Zaïem, Cem Subakan, Mirco Ravanelli |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Speech Self-Supervised Representation Benchmarking: Are We Doing it Right?abstractInternational audience Mohamed Salah Zaïem, Youcef Kemiche, Titouan Parcollet, Slim Essid, Mirco Ravanelli |
INTERSPEECH | 1 |
| 2023 | Automatic Data Augmentation for Domain Adapted Fine-Tuning of Self-Supervised Speech RepresentationsabstractInternational audience Mohamed Salah Zaïem, Titouan Parcollet, Slim Essid |
INTERSPEECH | 1 |
| 2022 | End-to-End Speech Recognition from Federated Acoustic ModelsabstractTraining Automatic Speech Recognition (ASR) models under federated learning (FL) settings has attracted a lot of attention recently. However, the FL scenarios often presented in the literature are artificial and fail to capture the complexity of real FL systems. In this paper, we construct a challenging and realistic ASR federated experimental setup consisting of clients with heterogeneous data distributions using the French and Italian sets of the CommonVoice dataset, a large heterogeneous dataset containing thousands of different speakers, acoustic environments and noises. We present the first empirical study on an attention-based sequence-to-sequence End-to-End (E2E) ASR model with three aggregation weighting strategies – standard FedAvg, loss-based aggregation and a novel word error rate (WER)-based aggregation, compared in two realistic FL scenarios: cross-silo with 10 clients and cross-device with 2K and 4K clients. This 4K cross-device ASR experiment is the largest ever performed. Our first-of-its-kind analysis on E2E ASR from heterogeneous and realistic federated acoustic models provides the foundations for future research and development of realistic FL ASR applications. Yan Gao 0016, Titouan Parcollet, Mohamed Salah Zaïem, Javier Fernández-Marqués, Pedro Porto Buarque de Gusmão, Daniel J. Beutel, Nicholas D. Lane |
ICASSP | 3 |
| 2022 | Automatic Data Augmentation Selection and Parametrization in Contrastive Self-Supervised Speech Representation LearningabstractInternational audience Mohamed Salah Zaïem, Titouan Parcollet, Slim Essid |
INTERSPEECH | 1 |
| 2022 | DP-Parse: Finding Word Boundaries from Raw Speech with an Instance LexiconabstractAbstract Finding word boundaries in continuous speech is challenging as there is little or no equivalent of a ‘space’ delimiter between words. Popular Bayesian non-parametric models for text segmentation (Goldwater et al., 2006, 2009) use a Dirichlet process to jointly segment sentences and build a lexicon of word types. We introduce DP-Parse, which uses similar principles but only relies on an instance lexicon of word tokens, avoiding the clustering errors that arise with a lexicon of word types. On the Zero Resource Speech Benchmark 2017, our model sets a new speech segmentation state-of-the-art in 5 languages. The algorithm monotonically improves with better input representations, achieving yet higher scores when fed with weakly supervised inputs. Despite lacking a type lexicon, DP-Parse can be pipelined to a language model and learn semantic and syntactic representations as assessed by a new spoken word embedding benchmark. 1 Robin Algayres, Tristan Ricoul, Julien Karadayi, Hugo Laurençon, Mohamed Salah Zaïem, Abdel-rahman Mohamed, Benoît Sagot, Emmanuel Dupoux |
Trans. Assoc. Comput. Linguistics | 5 |
| 2021 | Conditional Independence for Pretext Task Selection in Self-Supervised Speech Representation LearningabstractThrough solving pretext tasks, self-supervised learning (SSL) leverages unlabeled data to extract useful latent representations replacing traditional input features in the downstream task. A common pretext task consists in pretraining a SSL model on pseudo-labels derived from the original signal. This technique is particularly relevant for speech data where various meaningful signal processing features may serve as pseudo-labels. However, the process of selecting pseudo-labels, for speech or other types of data, remains mostly unexplored and currently relies on observing the results on the final downstream task. Nevertheless, this methodology is not sustainable at scale due to substantial computational (hence carbon) costs. Thus, this paper introduces a practical and theoretical framework to select relevant pseudo-labels with respect to a given downstream task. More precisely, we propose a functional estimator of the pseudo-label utility grounded in the conditional independence theory, which does not require any training. The experiments conducted on speaker recognition and automatic speech recognition validate our estimator, showing a significant correlation between the performance observed on the downstream task and the utility estimates obtained with our approach, facilitating the prospection of relevant pseudo-labels for self-supervised speech representation learning. Mohamed Salah Zaïem, Titouan Parcollet, Slim Essid |
Interspeech | 1 |
| 2020 | Evaluating the Reliability of Acoustic Speech EmbeddingsabstractInternational audience Robin Algayres, Mohamed Salah Zaïem, Benoît Sagot, Emmanuel Dupoux |
INTERSPEECH | 2 |
| 2019 | Sequence to Sequence Learning for Query ExpansionabstractAs fas as we are aware, using Sequence to Sequence algorithms for query expansion has not been explored yet in Information Retrieval literature. We tried to fill this gap in the literature with a custom Query Expansion system trained and tested on open datasets. One specificity of our engine compared to classic ones is that it does not need the documents to expand the introduced query. We test our expansions on two different tasks : Information Retrieval and Answer preselection. Our method yielded a slight improvement in performance in both two tasks . Our main contributions are :• Starting from open datasets, we built a Query Expansion training set using sentence-embeddings-based Keyword Extraction.• We assess the ability of the Sequence to Sequence neural networks to capture expanding relations in the words embeddings’ space.We afterwards started a quantitative and qualitative analysis of the weights learned by our network. In the second part, I will discuss what is learned by a Recurrent Neural Network compared to what we know about human language learning. Mohamed Salah Zaïem, Fatiha Sadat |
AAAI | 1 |