Thomas Pellegrini

dblp:32/7534 · DBLP profile ↗
← Back
44ranked-venue papers
21as first author
16since 2021 · last 2025
0000-0001-8984-1399ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 18 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 15 first-author · 13 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Hierarchical Label Propagation: A Model-Size-Dependent Performance Booster for AudioSet Tagging
abstract
AudioSet is one of the most used and largest datasets in audio tagging, containing about 2 million audio samples that are manually labeled with 527 event categories organized into an ontology. However, the annotations contain inconsistencies, particularly where categories that should be labeled as positive according to the ontology are frequently mislabeled as negative. To address this issue, we apply Hierarchical Label Propagation (HLP), which propagates labels up the ontology hierarchy, resulting in a mean increase in positive labels per audio clip from 1.98 to 2.39 and affecting 109 out of the 527 classes. Our results demonstrate that HLP provides performance benefits across various model architectures, including convolutional neural networks (PANN’s CNN6 and ConvNeXT) and transformers (PaSST), with smaller models showing more improvements. Finally, on FSD50K, another widely used dataset, models trained on AudioSet with HLP consistently outperformed those trained without HLP. Our source code will be made available on GitHub.
Ludovic Tuncay, Etienne Labbé, Thomas Pellegrini
ICASSP3
2024 Specializing Self-Supervised Speech Representations for Speaker Segmentation
abstract
International audience
Séverin Baroudi, Thomas Pellegrini, Hervé Bredin
INTERSPEECH2
2024 Self-Supervised Models for Phoneme Recognition: Applications in Children's Speech for Reading Learning
abstract
Child speech recognition is still an underdeveloped area of research due to the lack of data (especially on non-English languages) and the specific difficulties of this task. Having explored various architectures for child speech recognition in previous work, in this article we tackle recent self-supervised models. We first compare wav2vec 2.0, HuBERT and WavLM models adapted to phoneme recognition in French child speech, and continue our experiments with the best of them, WavLM base+. We then further adapt it by unfreezing its transformer blocks during fine-tuning on child speech, which greatly improves its performance and makes it significantly outperform our base model, a Transformer+CTC. Finally, we study in detail the behaviour of these two models under the real conditions of our application, and show that WavLM base+ is more robust to various reading tasks and noise levels. Index Terms: speech recognition, child speech, self-supervised learning
Lucas Block Medin, Thomas Pellegrini, Lucile Gelin
INTERSPEECH2
2024 CoNeTTE: An Efficient Audio Captioning System Leveraging Multiple Datasets With Task Embedding
abstract
Automated Audio Captioning (AAC) involves generating natural language descriptions of audio content, using encoder-decoder architectures. An audio encoder produces audio embeddings fed to a decoder, usually a Transformer decoder, for caption generation. In this work, we describe our model, which novelty, compared to existing models, lies in the use of a ConvNeXt architecture as audio encoder, adapted from the vision domain to audio classification. This model, called CNext-trans, achieved state-of-the-art scores on the AudioCaps (AC) dataset and performed competitively on Clotho (CL), while using four to forty times fewer parameters than existing models. We examine potential biases in the AC dataset due to its origin from AudioSet by investigating unbiased encoder's impact on performance. Using the well-known PANN's CNN14, for instance, as an unbiased encoder, we observed a 0.017 absolute reduction in SPIDEr score (where higher scores indicate better performance). To improve cross-dataset performance, we conducted experiments by combining multiple AAC datasets (AC, CL, MACS, WavCaps) for training. Although this strategy enhanced overall model performance across datasets, it still fell short compared to models trained specifically on a single target dataset, indicating the absence of a one-size-fits-all model. To mitigate performance gaps between datasets, we introduced a Task Embedding (TE) token, allowing the model to identify the source dataset for each input sample. We provide insights into the impact of these TEs on both the form (words) and content (sound event types) of the generated captions. The resulting model, named CoNeTTE, an unbiased CNext-trans model enriched with dataset-specific Task Embeddings, achieved SPIDEr scores of 0.467 and 0.310 on AC and CL, respectively.
Etienne Labbé, Thomas Pellegrini, Julien Pinquier
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Dilated convolution with learnable spacings
Ismail Khalfaoui Hassani, Thomas Pellegrini, Timothée Masquelier
ICLR2
2023 Adapting a ConvNeXt Model to Audio Classification on AudioSet
abstract
International audience
Thomas Pellegrini, Ismail Khalfaoui Hassani, Etienne Labbé, Timothée Masquelier
INTERSPEECH1
2023 Audio-video fusion strategies for active speaker detection in meetings
Lionel Pibre, Francisco Madrigal, Cyrille Equoy, Frédéric Lerasle, Thomas Pellegrini, Julien Pinquier, Isabelle Ferrané
Multim. Tools Appl.5
2022 PCEDNet: A Lightweight Neural Network for Fast and Interactive Edge Detection in 3D Point Clouds
abstract
In recent years, Convolutional Neural Networks (CNN) have proven to be efficient analysis tools for processing point clouds, e.g., for reconstruction, segmentation, and classification. In this article, we focus on the classification of edges in point clouds, where both edges and their surrounding are described. We propose a new parameterization adding to each point a set of differential information on its surrounding shape reconstructed at different scales. These parameters, stored in a Scale-Space Matrix (SSM) , provide a well-suited information from which an adequate neural network can learn the description of edges and use it to efficiently detect them in acquired point clouds. After successfully applying a multi-scale CNN on SSMs for the efficient classification of edges and their neighborhood, we propose a new lightweight neural network architecture outperforming the CNN in learning time, processing time, and classification capabilities. Our architecture is compact, requires small learning sets, is very fast to train, and classifies millions of points in seconds.
Chems-Eddine Himeur, Thibault Lejemble, Thomas Pellegrini, Mathias Paulin, Loïc Barthe, Nicolas Mellado
ACM Trans. Graph.3
2021 Automatic macro segmentation into interaction sequence: a silence-based approach for meeting structuring
abstract
Meetings are a common activity in professional contexts, and it remains difficult to analyze them because they are not always structured and people cut each other off (in a debate of ideas for example). A first step, to facilitate their analysis, is to segment the meeting into homogeneous zones at interaction level. To do so, we studied the typology of the non-speech segments (pauses and silences) in order to determine the different sequences during a meeting. Indeed, information such as the frequency and lengths of the non-speech segments will be different during a presentation or a debate. In this article, we propose an original approach to segment meetings using only the non-speech segments. We apply a Voice Activity Detection (VAD) to find the non-speech segments from which a set of parameters are extracted to study the typology of silence segments. We then use a sliding window on the whole meeting and we apply an unsupervised approach on each of these windows. We have validated our approaches using purity and coverage metrics on part of the AMI corpus (38 meetings of about 28 minutes each). This approach is non-invasive and relies only on acoustic information and does not analyze speech content since moments containing speech, and potentially sensitive information, are not processed.
Lionel Pibre, Sélim Mechrouh, Thomas Pellegrini, Julien Pinquier, Isabelle Ferrané
CBMI3
2021 Weakly supervised discourse segmentation for multiparty oral conversations
abstract
Discourse segmentation, the first step of discourse analysis, has been shown to improve results for text summarization, translation and other NLP tasks.While segmentation models for written text tend to perform well, they are not directly applicable to spontaneous, oral conversation, which has linguistic features foreign to written text.Segmentation is less studied for this type of language, where annotated data is scarce, and existing corpora more heterogeneous.We develop a weak supervision approach to adapt, using minimal annotation, a state of the art discourse segmenter trained on written text to French conversation transcripts.Supervision is given by a latent model bootstrapped by manually defined heuristic rules that use linguistic and acoustic information.The resulting model improves the original segmenter, especially in contexts where information on speaker turns is lacking or noisy, gaining up to 13% in F-score.Evaluation is performed on data like those used to define our heuristic rules, but also on transcripts from two other corpora.
Lila Gravellier, Julie Hunter 0001, Philippe Muller, Thomas Pellegrini, Isabelle Ferrané
EMNLP (1)4
2021 Comparison of Deep Co-Training and Mean-Teacher Approaches for Semi-Supervised Audio Tagging
abstract
Recently, a number of semi-supervised learning (SSL) methods, in the framework of deep learning (DL), were shown to achieve state-of-the-art results on image datasets, while using a (very) limited amount of labeled data. To our knowledge, these approaches adapted and applied to audio data are still sparse, in particular for audio tagging (AT). In this work, we adapted the Deep-Co-Training algorithm (DCT) to perform AT, and compared it to another SSL approach called Mean Teacher (MT), that has been used by the winning participants of the DCASE competitions these last two years. Experiments were performed on three standard audio datasets: Environmental Sound classification (ESC-10), UrbanSound8K, and Google Speech Commands. We show that both DCT and MT achieved performance approaching that of a fully supervised training setting, while using a fraction of the labeled data available, and the remaining data as unlabeled data. In some cases, DCT even reached the best accuracy, for instance, 72.6% using half of the labeled data, compared to 74.4% using all the labeled data. DCT also consistently outperformed MT in almost all configurations. For instance, the most significant relative gains brought by DCT reached 12.2% on ESC-10, compared to 7.6% with MT. Our code is available online1.
Léo Cances, Thomas Pellegrini
ICASSP2
2021 Fast Threshold Optimization for Multi-Label Audio Tagging Using Surrogate Gradient Learning
abstract
Multi-label audio tagging consists of assigning sets of tags to audio recordings. At inference time, thresholds are applied on the confidence scores outputted by a probabilistic classifier, in order to decide which classes are detected active. In this work, we consider having at disposal a trained classifier and we seek to automatically optimize the decision thresholds according to a performance metric of interest, in our case F-measure (micro-F1). We propose a new method, called SGL-Thresh for Surrogate Gradient Learning of Thresholds, that makes use of gradient descent. Since F1 is not differentiable, we propose to approximate the thresholding operation gradients with the gradients of a sigmoid function. We report experiments on three datasets, using state-of-the-art pre-trained deep neural networks. In all cases, SGLThresh outperformed three other approaches: a default threshold value (defThresh), an heuristic search algorithm and a method estimating F1 gradients numerically. It reached 54.9% F1 on AudioSet eval, compared to 50.7% with defThresh. SGLThresh is very fast and scalable to a large number of tags1.
Thomas Pellegrini, Timothée Masquelier
ICASSP1
2021 Simulating Reading Mistakes for Child Speech Transformer-Based Phone Recognition
abstract
International audience
Lucile Gelin, Thomas Pellegrini, Julien Pinquier, Morgane Daniel
Interspeech2
2021 Deep-Learning-Based Central African Primate Species Classification with MixUp and SpecAugment
Thomas Pellegrini
Interspeech1
2021 Low-Activity Supervised Convolutional Spiking Neural Networks Applied to Speech Commands Recognition
abstract
Deep Neural Networks (DNNs) are the current state-of-the-art models in many speech related tasks. There is a growing interest, though, for more biologically realistic, hardware friendly and energy efficient models, named Spiking Neural Networks (SNNs). Recently, it has been shown that SNNs can be trained efficiently, in a supervised manner, using backpropagation with a surrogate gradient trick. In this work, we report speech command (SC) recognition experiments using supervised SNNs. We explored the Leaky-Integrate-Fire (LIF) neuron model for this task, and show that a model comprised of stacked dilated convolution spiking layers can reach an error rate very close to standard DNNs on the Google SC v1 dataset: 5.5%, while keeping a very sparse spiking activity, below 5%, thank to a new regularization term. We also show that modeling the leakage of the neuron membrane potential is useful, since the LIF model outperformed its non-leaky model counterpart significantly.
Thomas Pellegrini, Romain Zimmer, Timothée Masquelier
SLT1
2021 End-to-end acoustic modelling for phone recognition of young readers
Lucile Gelin, Morgane Daniel, Julien Pinquier, Thomas Pellegrini
Speech Commun.4
2019 Cosine-similarity penalty to discriminate sound classes in weakly-supervised sound event detection
abstract
The design of new methods and models when only weakly-labeled data are available is of paramount importance in order to reduce the costs of manual annotation and the considerable human effort associated with it. In this work, we address Sound Event Detection in the case where a weakly annotated dataset is available for training. The weak annotations provide tags of audio events but do not provide temporal boundaries. The objective is twofold: 1) audio tagging, i.e. multi-label classification at recording level, 2) sound event detection, i.e. localization of the event boundaries within the recordings. This work focuses mainly on the second objective. We explore an approach inspired by Multiple Instance Learning, in which we train a convolutional recurrent neural network to give predictions at frame-level, using a custom loss function based on the weak labels and the statistics of the frame-based predictions. Since some sound classes cannot be distinguished with this approach, we improve the method by penalizing similarity between the predictions of the positive classes during training. On the test set used in the DCASE 2018 challenge, consisting of 288 recordings and 10 sound classes, the addition of a penalty resulted in a localization F-score of 34.75%, and brought 10% relative improvement compared to not using the penalty. Our best model achieved a 26.20% F-score on the DCASE-2018 official Eval subset close to the 10-system ensemble approach that ranked second in the challenge with a 29.9% F-score.
Thomas Pellegrini, Léo Cances
IJCNN1
2019 Char+CV-CTC: Combining Graphemes and Consonant/Vowel Units for CTC-Based ASR Using Multitask Learning
abstract
International audience
Abdelwahab Heba, Thomas Pellegrini, Jean-Pierre Lorré, Régine André-Obrecht
INTERSPEECH2
2019 The Airbus Air Traffic Control Speech Recognition 2018 Challenge: Towards ATC Automatic Transcription and Call Sign Detection
abstract
In this paper, we describe the outcomes of the challenge organized and run by Airbus and partners in 2018. The challenge consisted of two tasks applied to Air Traffic Control (ATC) speech in English: 1) automatic speech-to-text transcription, 2) call sign detection (CSD). The registered participants were provided with 40 hours of speech along with manual transcriptions. Twenty-two teams submitted predictions on a five hour evaluation set. ATC speech processing is challenging for several reasons: high speech rate, foreign-accented speech with a great diversity of accents, noisy communication channels. The best ranked team achieved a 7.62% Word Error Rate and a 82.41% CSD F1-score. Transcribing pilots' speech was found to be twice as harder as controllers' speech. Remaining issues towards solving ATC ASR are also discussed.
Thomas Pellegrini, Jérôme Farinas, Estelle Delpech, François Lancelot
INTERSPEECH1
2018 Group emotion recognition strategies for entertainment robots
abstract
In this paper, a system to determine the emotion of a group of people via facial expression analysis is proposed for the Waseda Entertainment Robots. General models and standard methods for emotion definition and recognition are briefly described, as well as strategies for computing the group global emotion, knowing the individual emotions of group members. This work is based on Ekman's extended “Big Six” emotional model, popular in Computer Science and Affective Computing. Emotion recognition via facial expression analysis is performed with a cloud-computing based solution, using Microsoft Azure Cognitive services. First, the performances of both the Face API to detect faces, and Emotion API, to compute emotion via face expression analysis, are tested. After that, a solution to compute the emotion of a group of people has been implemented and its performances compared to human perceptions. This work presents concepts and strategies which can be generalized for applications within the scope of assistive robotics and, more broadly, affective computing, wherever it will be necessary to determine the emotion of a group of people.
Sarah Cosentino, Estelle I. S. Randria, Jia-Yeu Lin, Thomas Pellegrini, Salvatore Sessa 0001, Atsuo Takanishi
IROS4
2016 Sinusoidal Modelling for Ecoacoustics
abstract
Biodiversity assessment is a central and urgent task, necessary to monitoring the changes to ecological systems and under- standing the factors which drive these changes. Technological advances are providing new approaches to monitoring, which are particularly useful in remote regions. Situated within the framework of the emerging field of ecoacoustics, there is grow- ing interest in the possibility of extracting ecological informa- tion from digital recordings of the acoustic environment. Rather than focusing on identification of individual species, an increas- ing number of automated indices attempt to summarise acoustic activity at the community level, in order to provide a proxy for biodiversity. Originally designed for speech processing, sinu- soidal modelling has previously been used as a bioacoustic tool, for example to detect particular bird species. In this paper, we demonstrate the use of sinusoidal modelling as a proxy for bird abundance. Using data from acoustic surveys made during the breeding season in UK woodland, the number of extracted sinusoidal tracks is shown to correlate with estimates of bird abundance made by expert ornithologists listening to the recordings. We also report ongoing work exploring a new approach to investigate the composition of calls in spectro-temporal space that constitutes a promising new method for Ecoaoustic biodiversity assessment.
Patrice Guyot, Alice C. Eldridge, Ying Chen Eyre-Walker, Alison Johnston, Thomas Pellegrini, Mika Peck
INTERSPEECH5
2016 Pronunciation Assessment of Japanese Learners of French with GOP Scores and Phonetic Information
abstract
In this paper, we report automatic pronunciation assessment experiments at phone-level on a read speech corpus in French, collected from 23 Japanese speakers learning French as a foreign language. We compare the standard approach based on Goodness Of Pronunciation (GOP) scores and phone-specific score thresholds to the use of logistic regressions (LR) models. French native speech corpus, in which artificial pronunciation errors were introduced, was used as training set. Two typical errors of Japanese speakers were considered: /ö/ and /v/ of ten mispronounced as [l] and [b], respectively. The LR classifier achieved a 64.4% accuracy similar to the 63.8% accuracy of the baseline threshold method, when using GOP scores and the expected phone identity as input features only. A significant performance gain of 20.8% relative was obtained by adding phonetic and phonological features as input to the LR model, leading to a 77.1% accuracy. This LR model also outperformed another baseline approach based on linear discriminant models trained on raw f-BANK coefficient features.
Vincent Laborde, Thomas Pellegrini, Lionel Fontan, Julie Mauclair, Halima Sahraoui, Jérôme Farinas
INTERSPEECH2
2016 CNN-Based Phone Segmentation Experiments in a Less-Represented Language
abstract
These last years, there has been a regain of interest in unsupervised sub-lexical and lexical unit discovery. Speech segmentation into phone-like units may be a first interesting step for such a task. In this article, we report speech segmentation experiments in Xitsonga, a less-represented language spoken in South Africa. We chose to use convolutional neural networks (CNN) with FBANK static coefficients as input. The models take binary decisions whether a boundary is present or not at each signal sliding frame. We compare the use of a model trained exclusively on Xitsonga data to the use of a bootstrap model trained on a larger corpus of another language, the BUCKEYE U.S. English corpus. Using a two-convolution-layer model, a 79% F-measure was obtained on BUCKEYE, with a 20 ms error tolerance. This performance is equal to the human inter-annotator agreement rate. We then used this bootstrap model to segment Xitsonga data and compared the results when adapting it with 1 to 20 minutes of Xitsonga data.
Céline Manenti, Thomas Pellegrini, Julien Pinquier
INTERSPEECH2
2016 Inferring Phonemic Classes from CNN Activation Maps Using Clustering Techniques
abstract
National audience
Thomas Pellegrini, Sandrine Mouysset
INTERSPEECH1
2015 Comparing SVM, softmax, and shallow neural networks for eating condition classification
abstract
International audience
Thomas Pellegrini
INTERSPEECH1
2014 The goodness of pronunciation algorithm applied to disordered speech
abstract
International audience
Thomas Pellegrini, Lionel Fontan, Julie Mauclair, Jérôme Farinas, Marina Robert
INTERSPEECH1
2014 Speaker age estimation for elderly speech recognition in European Portuguese
abstract
International audience
Thomas Pellegrini, Vahid Hedayati, Isabel Trancoso, Annika Hämäläinen, José Miguel Salles Dias
INTERSPEECH1
2014 Segmentation in singer turns with the Bayesian information criterion
abstract
As part of a project on indexing ethno-musicological audio recordings, segmentation in singer turns automatically appeared to be essential. In this article, we present the problem of segmentation in singer turns of musical recordings and our first experiments in this direction by exploring a method based on the Bayesian Information Criterion (BIC), which are used in numerous works in audio segmentation, to detect singer turns. The BIC penalty coefficient was shown to vary when determining its value to achieve the best performance for each recording. In order to avoid the decision about which single value is best for all the documents, we propose to combine several segmentations obtained with different values of this parameter. This method consists of taking a posteriori decisions on which segment boundaries are to be kept. A gain of 7.1% in terms of F-measure was obtained compared to a standard coefficient.
Marwa Thlithi, Thomas Pellegrini, Julien Pinquier, Régine André-Obrecht
INTERSPEECH2
2014 El-WOZ: a client-server wizard-of-oz interface
Thomas Pellegrini, Vahid Hedayati, Ângela Costa
LREC1
2013 A corpus-based study of elderly and young speakers of European Portuguese: acoustic correlates and their impact on speech recognition performance
abstract
This paper presents a study of European Portuguese elderly speech, in which the acoustic characteristics of two groups of elderly speakers (aged 60-75 and over 75) are compared with those of young adult speakers (aged 19-30). The correlation between age and a set of 14 acoustic features was investigated, and decision trees were used to establish the relative importance of the features. A greater use of pauses characterized speakers aged 60 and over. For female speakers, speech rate also appeared to correlate with age. For male speakers, jitter distinguished between speakers aged 60-75 and older. The correlation between the features and speech recognition performance was also investigated. Word error rate correlated mostly with the use of pauses, speech rate, and the ratio of long phone realizations. Finally, by comparing the phone sequences used by the recognizer on the most frequent words, we observed that the young adult speakers reduced schwas more than the elderly speakers. This result seems to confirm the common idea that young speakers reduce articulation more than older speakers. Further investigation is needed to confirm this result by determining whether this is due to ageing or to the generation gap.
Thomas Pellegrini, Annika Hämäläinen, Philippe Boula de Mareüil, Michael Tjalve, Isabel Trancoso, Sara Candeias, José Miguel Salles Dias, Daniela Braga
INTERSPEECH1
2013 ASR-based exercises for listening comprehension practice in European Portuguese
Thomas Pellegrini, Rui Correia, Isabel Trancoso, Jorge Baptista, Nuno J. Mamede, Maxine Eskénazi
Comput. Speech Lang.1
2012 Overview of Computer-assisted Language Learning for European Portuguese at L2f
Thomas Pellegrini, Wang Ling, Rui Correia, Isabel Trancoso, Jorge Baptista, Nuno J. Mamede
CSEDU (2)1
2012 Less errors with TTS? A dictation experiment with foreign language learners
Thomas Pellegrini, Ângela Costa, Isabel Trancoso
INTERSPEECH1
2011 Automatic Generation of Listening Comprehension Learning Material in European Portuguese
abstract
The goal of this work is the automatic selection of materials for a listening comprehension game. We would like to select automatically transcribed sentences from recent broadcast news corpora, in order to gather material for the games with little human effort. The recognized words are used as the ground solution of the exercises, thus sentences with misrecognitions need to be filtered out. Our experiments confirmed the feasibility of the filter chain that automatically selects sentences, although harder confidence thresholds may be needed. Together with the correct words, wrong candidates, namely distractors, are also needed to build the exercises. Two techniques of distractor generation are presented, either based on the confusion networks produced by the recognizer, or on phonetic distances. The experiments confirmed the complementarity of both approaches. Index Terms: CALL, Listening Comprehension, European Portuguese, ASR, distractors
Thomas Pellegrini, Rui Correia, Isabel Trancoso, Jorge Baptista, Nuno J. Mamede
INTERSPEECH1
2010 Context dependent modelling approaches for hybrid speech recognizers
abstract
Speech recognition based on connectionist approaches is one of the most successful alternatives to widespread Gaussian systems. One of the main claims against hybrid recognizers is the increased complexity for context-dependent phone modeling, which is a key aspect in medium to large size vocabulary tasks. In this paper, we investigate the use of context-dependent triphone models in a connectionist speech recognizer. Thus, most common triphone state clustering procedures for Gaussian models are compared and applied to our hybrid recognizer. The developed systems with clustered context-dependent triphones show above 20% relative word error rate reduction compared to a baseline hybrid system in two selected WSJ evaluation test sets. Additionally, the recent porting efforts of the proposed context modelling approaches to a LVCSR system for English Broadcast News transcription are reported. Index Terms: speech recognition, context modeling, connectionist system
Alberto Abad, Thomas Pellegrini, Isabel Trancoso, João Paulo da Silva Neto
INTERSPEECH2
2010 Improving ASR error detection with non-decoder based features
abstract
This study reports error detection experiments in large vocabulary automatic speech recognition (ASR) systems, by using statistical classifiers. We explored new features gathered from other knowledge sources than the decoder itself: a binary feature that compares outputs from two different ASR systems (word by word), a feature based on the number of hits of the hypothesized bigrams, obtained by queries entered into a very popular Web search engine, and finally a feature related to automatically infered topics at sentence and word levels. Experiments were conducted on a European Portuguese broadcast news corpus. The combination of baseline decoder-based features and two of these additional features led to significant improvements, from 13.87% to 12.16% classification error rate (CER) with a maximum entropy model, and from 14.01% to 12.39% CER with linear-chain conditional random fields, comparing to a baseline using only decoder-based features.
Thomas Pellegrini, Isabel Trancoso
INTERSPEECH1
2010 Multimedia learning materials
abstract
This paper describes the integration of multimedia documents in the Portuguese version of REAP, a tutoring system for vocabulary learning. The documents result from the pipeline processing of Broadcast News videos that automatically segments the audio files, transcribes them, adds punctuation and capitalization, and breaks them into stories classified by topics. The integration of these materials in REAP was done in a way that tries to decrease the impact of potential errors of the automatic chain in the learning process.
José Lopes 0001, Isabel Trancoso, Rui Correia, Thomas Pellegrini, Hugo Meinedo, Nuno J. Mamede, Maxine Eskénazi
SLT4
2009 Audio contributions to semantic video search
abstract
This paper summarizes the contributions to semantic video search that can be derived from the audio signal. Because of space restrictions, the emphasis will be on non-linguistic cues. The paper thus covers what is generally known as audio segmentation, as well as audio event detection. Using machine learning approaches, we have built detectors for over 50 semantic audio concepts.
Isabel Trancoso, Thomas Pellegrini, José Portelo, Hugo Meinedo, Miguel M. F. Bugalho, Alberto Abad, João Paulo da Silva Neto
ICME2
2009 Detecting audio events for semantic video search
abstract
This paper describes our work on audio event detection, one of our tasks in the European project VIDIVIDEO. Preliminary experiments with a small corpus of sound effects have shown the potential of this type of corpus for training purposes. This paper describes our experiments with SVM classifiers, and different features, using a 290-hour corpus of sound effects, which allowed us to build detectors for almost 50 semantic concepts. Although the performance of these detectors on the development set is quite good (achieving an average F-measure of 0.87), preliminary experiments on documentaries and films showed that the task is much harder in real-life videos, which so often include overlapping audio events. Index Terms: event detection, audio segmentation 1.
Miguel M. F. Bugalho, José Portelo, Isabel Trancoso, Thomas Pellegrini, Alberto Abad
INTERSPEECH4
2009 Automatic Word Decompounding for ASR in a Morphologically Rich Language: Application to Amharic
abstract
This paper investigates a data-driven word decompounding algorithm for use in automatic speech recognition. An existing algorithm, called ldquoMorfessor,rdquo has been enhanced in order to address the problem of increased phonetic confusability arising from word decompounding by incorporating phonetic properties and some constraints on recognition units derived from forced alignments experiments. Speech recognition experiments have been carried out on a broadcast news task for the Amharic language to validate the approach. The out of vocabulary (OOV) word rates were reduced by 35% to 50% and a small reduction in word error rate (WER) has been achieved. The algorithm is relatively language independent and requires minimal adaptation to be applied to other languages.
Thomas Pellegrini, Lori Lamel
IEEE Trans. Speech Audio Process.1
2008 Developments of "Lëtzebuergesch" Resources for Automatic Speech Processing and Linguistic Studies
Martine Adda-Decker, Thomas Pellegrini, Eric Bilinski, Gilles Adda
LREC2
2007 Using phonetic features in unsupervised word decompounding for ASR with application to a less-represented language
abstract
In this paper, a data-driven word decompounding algorithm is described and applied to a broadcast news corpus in Amharic. The baseline algorithm has been enhanced in order to address the problem of increased phonetic confusability arising from word decompounding by incorporating phonetic properties and some constraints on recognition units derived from prior forced alignment experiments. Speech recognition experiments have been carried out to validate the approach. Out of vocabulary (OOV) words rates can be reduced by 30% to 40% and an absolute Word Error Rate (WER) reduction of 0.4% has been achieved. The algorithm is relatively language independent and requires minimal adaptation to be applied to other languages. Index Terms: automatic speech recognition, unsupervised word decompounding, less-represented languages
Thomas Pellegrini, Lori Lamel
INTERSPEECH1
2006 Investigating automatic decomposition for ASR in less represented languages
abstract
This paper addresses the use of an automatic decomposition method to reduce lexical variety and thereby improve speech recognition of less well-represented languages. The Amharic language has been selected for these experiments since only a small quantity of resources are available compared to well-covered languages. Inspired by the Harris algorithm, the method automatically generates plausible affixes, that combined with decompounding can reduce the size of the lexicon and the OOV rate. Recognition experiments are carried out for four different configurations (full-word and decompounded) and using supervised training with a corpus containing only two hours of manually transcribed data.
Thomas Pellegrini, Lori Lamel
INTERSPEECH1
2006 Experimental detection of vowel pronunciation variants in Amharic
Thomas Pellegrini, Lori Lamel
LREC1