Fethi Bougares

dblp:75/9232 · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Machine translation · 59% Language models and text generation · 22% Representation and self-supervised learning · 19%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › language modeling
continuous space language models
0.212015
Investigating Continuous Space Language Models for Machine Translation Quality Estimation · EMNLP 2015
Natural language and speech › Machine translation › machine translation evaluation
translation quality estimation
0.212015
Investigating Continuous Space Language Models for Machine Translation Quality Estimation · EMNLP 2015
Natural language and speech › Machine translation
neural machine translation
0.212014
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation · EMNLP 2014
Machine learning › Representation and self-supervised learning › text embedding
phrase representation learning
0.212014
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation · EMNLP 2014
Natural language and speech › Machine translation
statistical machine translation
0.212014
Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation · EMNLP 2014

Methods — techniques the papers use, named apart from their topics

neural network features · 0.2continuous space language model · 0.2recurrent neural network · 0.2encoder-decoder · 0.2
YearPublicationVenuePosition
2026 WhiteHouse: Translation of the Casablanca Corpus for Multi-dialectal Arabic Speech Translation
Fethi Bougares, Salima Mdhaffar, Yannick Estève
LREC1
2026 SLURP-TN : Resource for Tunisian Dialect Spoken Language Understanding
Haroun Elleuch, Salima Mdhaffar, Yannick Estève, Fethi Bougares
LREC4
2025 ADI-20: Arabic Dialect Identification dataset and models
abstract
Published in Interspeech 2025
Haroun Elleuch, Salima Mdhaffar, Yannick Estève, Fethi Bougares
INTERSPEECH4
2024 TunArTTS: Tunisian Arabic Text-To-Speech Corpus
abstract
Being labeled as a low-resource language, the Tunisian dialect has no existing prior TTS research. In this paper, we present a speech corpus for Tunisian Arabic Text-to-Speech (TunArTTS) to initiate the development of end-to-end TTS systems for the Tunisian dialect. Our Speech corpus is extracted from an online English and Tunisian Arabic dictionary. We were able to extract a mono-speaker speech corpus of +3 hours of a male speaker sampled at 44100 kHz. The corpus is processed and manually diacritized. Furthermore, we develop various TTS systems based on two approaches: training from scratch and transfer learning. Both Tacotron2 and FastSpeech2 were used and evaluated using subjective and objective metrics. The experimental results show that our best results are obtained with the transfer learning from a pre-trained model on the English LJSpeech dataset. This model obtained a mean opinion score (MOS) of 3.88. TunArTTS will be publicly available for research purposes along with the baseline TTS system demo. Keywords: Tunisian Dialect, Text-To-Speech, Low-resource, Transfer Learning, TunArTTS
Imen Laouirine, Rami Kammoun, Fethi Bougares
LREC/COLING3
2024 TARIC-SLU: A Tunisian Benchmark Dataset for Spoken Language Understanding
abstract
In recent years, there has been a significant increase in interest in developing Spoken Language Understanding (SLU) systems. SLU involves extracting a list of semantic information from the speech signal. A major issue for SLU systems is the lack of sufficient amount of bi-modal (audio and textual semantic annotation) training data. Existing SLU resources are mainly available in high-resource languages such as English, Mandarin and French. However, one of the current challenges concerning low-resourced languages is data collection and annotation. In this work, we present a new freely available corpus, named TARIC-SLU, composed of railway transport conversations in Tunisian dialect that is continuously annotated in dialogue acts and slots. We describe the semantic model of the dataset, the data and experiments conducted to build ASR-based and SLU-based baseline models. To facilitate its use, a complete recipe, including data preparation, training and evaluation scripts, has been built and will be integrated to SpeechBrain, a popular open-source conversational AI toolkit based on PyTorch.
Salima Mdhaffar, Fethi Bougares, Renato De Mori, Mohamed Salah Zaïem, Mirco Ravanelli, Yannick Estève
LREC/COLING2
2024 ALLIES: A Speech Corpus for Segmentation, Speaker Diarization, Speech Recognition and Speaker Change Detection
abstract
This paper presents ALLIES, a meta corpus which gathers and extends existing French corpora collected from radio and TV shows. The corpus contains 1048 audio files for about 500 hours of speech. Agglomeration of data is always a difficult issue, as the guidelines used to collect, annotate and transcribe speech are generally different from one corpus to another. ALLIES intends to homogenize and correct speaker labels among the different files by integrated human feedback within a speaker verification system. The main contribution of this article is the design of a protocol in order to evaluate properly speech segmentation (including music and overlap detection), speaker diarization, speech transcription and speaker change detection. As part of it, a test partition has been carefully manually 1) segmented and annotated according to speech, music, noise, speaker labels with specific guidelines for overlap speech, 2) orthographically transcribed. This article also provides as a second contribution baseline results for several speech processing tasks.
Marie Tahon, Anthony Larcher, Martin Lebourdais, Fethi Bougares, Anna Silnova, Pablo Gimeno
LREC/COLING4
2022 Speech Resources in the Tamasheq Language
abstract
In this paper we present two datasets for Tamasheq, a developing language mainly spoken in Mali and Niger. These two datasets were made available for the IWSLT 2022 low-resource speech translation track, and they consist of collections of radio recordings from daily broadcast news in Niger (Studio Kalangou) and Mali (Studio Tamani). We share (i) a massive amount of unlabeled audio data (671 hours) in five languages: French from Niger, Fulfulde, Hausa, Tamasheq and Zarma, and (ii) a smaller 17 hours parallel corpus of audio recordings in Tamasheq, with utterance-level translations in the French language. All this data is shared under the Creative Commons BY-NC-ND 3.0 license. We hope these resources will inspire the speech community to develop and benchmark models using the Tamasheq language.
Marcely Zanon Boito, Fethi Bougares, Florentin Barbier, Souhir Gahbiche-Braham, Loïc Barrault, Mickael Rouvier, Yannick Estève
LREC2
2020 Investigating Self-Supervised Pre-Training for End-to-End Speech Translation
abstract
International audience
Fethi Bougares, Natalia A. Tomashenko, Yannick Estève, Laurent Besacier
INTERSPEECH2
2020 Text and Speech-based Tunisian Arabic Sub-Dialects Identification
abstract
Dialect IDentification (DID) is a challenging task, and it becomes more complicated when it is about the identification of dialects that belong to the same country. Indeed, dialects of the same country are closely related and exhibit a significant overlapping at the phonetic and lexical levels. In this paper, we present our first results on a dialect classification task covering four sub-dialects spoken in Tunisia. We use the term ’sub-dialect’ to refer to the dialects belonging to the same country. We conducted our experiments aiming to discriminate between Tunisian sub-dialects belonging to four different cities: namely Tunis, Sfax, Sousse and Tataouine. A spoken corpus of 1673 utterances is collected, transcribed and freely distributed. We used this corpus to build several speech- and text-based DID systems. Our results confirm that, at this level of granularity, dialects are much better distinguishable using the speech modality. Indeed, we were able to reach an F-1 score of 93.75% using our best speech-based identification system while the F-1 score is limited to 54.16% using text-based DID on the same test set.
Najla Ben Abdallah, Saméh Kchaou, Fethi Bougares
LREC3
2020 Addressing data sparsity for neural machine translation between morphologically rich languages
Mercedes García-Martínez, Walid Aransa, Fethi Bougares, Loïc Barrault
Mach. Transl.3
2019 Extrinsic Plagiarism Detection for French Language with Word Embeddings
Maryam Elamine, Fethi Bougares, Seifeddine Mechti, Lamia Hadrich Belguith
ISDA2
2016 Conditional Random Fields for the Tunisian Dialect Grapheme-to-Phoneme Conversion
Abir Masmoudi 0001, Mariem Ellouze, Fethi Bougares, Yannick Estève, Lamia Hadrich Belguith
INTERSPEECH3
2015 Investigating Continuous Space Language Models for Machine Translation Quality Estimation
abstract
We present novel features designed with a deep neural network for Machine Translation (MT) Quality Estimation (QE).The features are learned with a Continuous Space Language Model to estimate the probabilities of the source and target segments.These new features, along with standard MT system-independent features, are benchmarked on a series of datasets with various quality labels, including postediting effort, human translation edit rate, post-editing time and METEOR.Results show significant improvements in prediction over the baseline, as well as over systems trained on state of the art feature sets for all datasets.More notably, the addition of the newly proposed features improves over the best QE systems in WMT12 and WMT14 by a significant margin.
Kashif Shah, Raymond W. M. Ng, Fethi Bougares, Lucia Specia
EMNLP3
2015 Continuous Adaptation to User Feedback for Statistical Machine Translation
abstract
Frédéric Blain, Fethi Bougares, Amir Hazem, Loïc Barrault, Holger Schwenk. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Frédéric Blain, Fethi Bougares, Amir Hazem, Loïc Barrault, Holger Schwenk
HLT-NAACL2
2014 Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
abstract
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, Yoshua Bengio. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2014.
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, Yoshua Bengio
EMNLP5
2012 Low latency combination of parallelized single-pass LVCSR systems
abstract
International audience
Fethi Bougares, Mickael Rouvier, Yannick Estève, Georges Linarès
INTERSPEECH1
2011 Bag of n-gram driven decoding for LVCSR system harnessing
abstract
This paper focuses on automatic speech recognition systems combination based on driven decoding paradigms. The driven decoding algorithm (DDA) involves the use of a 1-best hypothesis provided by an auxiliary system as another knowledge source in the search algorithm of a primary system. In previous studies, it was shown that DDA outperforms ROVER when the primary system is guided by a more accurate system. In this paper we propose a new method to manage auxiliary transcriptions which are presented as a bag-of-n-grams (BONG) without temporal matching. These modifications allow to make easier the combination of several hypotheses given by different auxiliary systems. Using BONG combination with hypotheses provided by two auxiliary systems, each of which obtained more than 23% of WER on the same data, our experiments show that a CMU Sphinx based ASR system can reduce its WER from 19.85% to 18.66% which is better than the results reached with DDA or classical ROVER combination.
Fethi Bougares, Yannick Estève, Paul Deléglise, Georges Linarès
ASRU1
2010 Unsupervised model adaptation on targeted speech segments for LVCSR system combination
abstract
In context of Large-Vocabulary Continuous Speech Recognition, systems can reach a high level of performance when dealing with prepared speech, while their performance drops on spontaneous speech. This decrease is due to the fact that these two kinds of speech are marked by strong acoustic and linguistic differences. Previous research works had been done to detect and repair some peculiarities of spontaneous speech, as disfluencies, and to create specific models to improve recognition accuracy: a large amount of data is needed to see improvements and is expensive to collect. In this paper, we present a solution to create specialized acoustic and language models, by automatically extracting a data subset from the initial training corpus containing spontaneous speech, and adapting initial acoustic and linguistic models on it. As we assume these models can be complementary, we propose to combine general and adapted ASR system outputs. Experimental results show statistically significant gain, for a negligible cost (no additional training data and no human intervention).
Richard Dufour, Fethi Bougares, Yannick Estève, Paul Deléglise
INTERSPEECH2