VLDB 2026 Research / reviewers in the wild / expert
Thomas Pellegrini
dblp:32/7534
· DBLP profile ↗
44ranked-venue papers
21as first author
16since 2021 · last 2025
0000-0001-8984-1399ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 18 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 15 first-author · 13 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hierarchical Label Propagation: A Model-Size-Dependent Performance Booster for AudioSet TaggingabstractAudioSet is one of the most used and largest datasets in audio tagging, containing about 2 million audio samples that are manually labeled with 527 event categories organized into an ontology. However, the annotations contain inconsistencies, particularly where categories that should be labeled as positive according to the ontology are frequently mislabeled as negative. To address this issue, we apply Hierarchical Label Propagation (HLP), which propagates labels up the ontology hierarchy, resulting in a mean increase in positive labels per audio clip from 1.98 to 2.39 and affecting 109 out of the 527 classes. Our results demonstrate that HLP provides performance benefits across various model architectures, including convolutional neural networks (PANN’s CNN6 and ConvNeXT) and transformers (PaSST), with smaller models showing more improvements. Finally, on FSD50K, another widely used dataset, models trained on AudioSet with HLP consistently outperformed those trained without HLP. Our source code will be made available on GitHub. Ludovic Tuncay, Etienne Labbé, Thomas Pellegrini |
ICASSP | 3 |
| 2024 | Specializing Self-Supervised Speech Representations for Speaker SegmentationabstractInternational audience Séverin Baroudi, Thomas Pellegrini, Hervé Bredin |
INTERSPEECH | 2 |
| 2024 | Self-Supervised Models for Phoneme Recognition: Applications in Children's Speech for Reading LearningabstractChild speech recognition is still an underdeveloped area of research due to the lack of data (especially on non-English languages) and the specific difficulties of this task. Having explored various architectures for child speech recognition in previous work, in this article we tackle recent self-supervised models. We first compare wav2vec 2.0, HuBERT and WavLM models adapted to phoneme recognition in French child speech, and continue our experiments with the best of them, WavLM base+. We then further adapt it by unfreezing its transformer blocks during fine-tuning on child speech, which greatly improves its performance and makes it significantly outperform our base model, a Transformer+CTC. Finally, we study in detail the behaviour of these two models under the real conditions of our application, and show that WavLM base+ is more robust to various reading tasks and noise levels. Index Terms: speech recognition, child speech, self-supervised learning Lucas Block Medin, Thomas Pellegrini, Lucile Gelin |
INTERSPEECH | 2 |
| 2024 | CoNeTTE: An Efficient Audio Captioning System Leveraging Multiple Datasets With Task EmbeddingabstractAutomated Audio Captioning (AAC) involves generating natural language descriptions of audio content, using encoder-decoder architectures. An audio encoder produces audio embeddings fed to a decoder, usually a Transformer decoder, for caption generation. In this work, we describe our model, which novelty, compared to existing models, lies in the use of a ConvNeXt architecture as audio encoder, adapted from the vision domain to audio classification. This model, called CNext-trans, achieved state-of-the-art scores on the AudioCaps (AC) dataset and performed competitively on Clotho (CL), while using four to forty times fewer parameters than existing models. We examine potential biases in the AC dataset due to its origin from AudioSet by investigating unbiased encoder's impact on performance. Using the well-known PANN's CNN14, for instance, as an unbiased encoder, we observed a 0.017 absolute reduction in SPIDEr score (where higher scores indicate better performance). To improve cross-dataset performance, we conducted experiments by combining multiple AAC datasets (AC, CL, MACS, WavCaps) for training. Although this strategy enhanced overall model performance across datasets, it still fell short compared to models trained specifically on a single target dataset, indicating the absence of a one-size-fits-all model. To mitigate performance gaps between datasets, we introduced a Task Embedding (TE) token, allowing the model to identify the source dataset for each input sample. We provide insights into the impact of these TEs on both the form (words) and content (sound event types) of the generated captions. The resulting model, named CoNeTTE, an unbiased CNext-trans model enriched with dataset-specific Task Embeddings, achieved SPIDEr scores of 0.467 and 0.310 on AC and CL, respectively. Etienne Labbé, Thomas Pellegrini, Julien Pinquier |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Dilated convolution with learnable spacings
Ismail Khalfaoui Hassani, Thomas Pellegrini, Timothée Masquelier |
ICLR | 2 |
| 2023 | Adapting a ConvNeXt Model to Audio Classification on AudioSetabstractInternational audience Thomas Pellegrini, Ismail Khalfaoui Hassani, Etienne Labbé, Timothée Masquelier |
INTERSPEECH | 1 |
| 2023 | Audio-video fusion strategies for active speaker detection in meetings
Lionel Pibre, Francisco Madrigal, Cyrille Equoy, Frédéric Lerasle, Thomas Pellegrini, Julien Pinquier, Isabelle Ferrané |
Multim. Tools Appl. | 5 |
| 2022 | PCEDNet: A Lightweight Neural Network for Fast and Interactive Edge Detection in 3D Point CloudsabstractIn recent years, Convolutional Neural Networks (CNN) have proven to be efficient analysis tools for processing point clouds, e.g., for reconstruction, segmentation, and classification. In this article, we focus on the classification of edges in point clouds, where both edges and their surrounding are described. We propose a new parameterization adding to each point a set of differential information on its surrounding shape reconstructed at different scales. These parameters, stored in a Scale-Space Matrix (SSM) , provide a well-suited information from which an adequate neural network can learn the description of edges and use it to efficiently detect them in acquired point clouds. After successfully applying a multi-scale CNN on SSMs for the efficient classification of edges and their neighborhood, we propose a new lightweight neural network architecture outperforming the CNN in learning time, processing time, and classification capabilities. Our architecture is compact, requires small learning sets, is very fast to train, and classifies millions of points in seconds. Chems-Eddine Himeur, Thibault Lejemble, Thomas Pellegrini, Mathias Paulin, Loïc Barthe, Nicolas Mellado |
ACM Trans. Graph. | 3 |
| 2021 | Automatic macro segmentation into interaction sequence: a silence-based approach for meeting structuringabstractMeetings are a common activity in professional contexts, and it remains difficult to analyze them because they are not always structured and people cut each other off (in a debate of ideas for example). A first step, to facilitate their analysis, is to segment the meeting into homogeneous zones at interaction level. To do so, we studied the typology of the non-speech segments (pauses and silences) in order to determine the different sequences during a meeting. Indeed, information such as the frequency and lengths of the non-speech segments will be different during a presentation or a debate. In this article, we propose an original approach to segment meetings using only the non-speech segments. We apply a Voice Activity Detection (VAD) to find the non-speech segments from which a set of parameters are extracted to study the typology of silence segments. We then use a sliding window on the whole meeting and we apply an unsupervised approach on each of these windows. We have validated our approaches using purity and coverage metrics on part of the AMI corpus (38 meetings of about 28 minutes each). This approach is non-invasive and relies only on acoustic information and does not analyze speech content since moments containing speech, and potentially sensitive information, are not processed. Lionel Pibre, Sélim Mechrouh, Thomas Pellegrini, Julien Pinquier, Isabelle Ferrané |
CBMI | 3 |
| 2021 | Weakly supervised discourse segmentation for multiparty oral conversationsabstractDiscourse segmentation, the first step of discourse analysis, has been shown to improve results for text summarization, translation and other NLP tasks.While segmentation models for written text tend to perform well, they are not directly applicable to spontaneous, oral conversation, which has linguistic features foreign to written text.Segmentation is less studied for this type of language, where annotated data is scarce, and existing corpora more heterogeneous.We develop a weak supervision approach to adapt, using minimal annotation, a state of the art discourse segmenter trained on written text to French conversation transcripts.Supervision is given by a latent model bootstrapped by manually defined heuristic rules that use linguistic and acoustic information.The resulting model improves the original segmenter, especially in contexts where information on speaker turns is lacking or noisy, gaining up to 13% in F-score.Evaluation is performed on data like those used to define our heuristic rules, but also on transcripts from two other corpora. Lila Gravellier, Julie Hunter 0001, Philippe Muller, Thomas Pellegrini, Isabelle Ferrané |
EMNLP (1) | 4 |
| 2021 | Comparison of Deep Co-Training and Mean-Teacher Approaches for Semi-Supervised Audio TaggingabstractRecently, a number of semi-supervised learning (SSL) methods, in the framework of deep learning (DL), were shown to achieve state-of-the-art results on image datasets, while using a (very) limited amount of labeled data. To our knowledge, these approaches adapted and applied to audio data are still sparse, in particular for audio tagging (AT). In this work, we adapted the Deep-Co-Training algorithm (DCT) to perform AT, and compared it to another SSL approach called Mean Teacher (MT), that has been used by the winning participants of the DCASE competitions these last two years. Experiments were performed on three standard audio datasets: Environmental Sound classification (ESC-10), UrbanSound8K, and Google Speech Commands. We show that both DCT and MT achieved performance approaching that of a fully supervised training setting, while using a fraction of the labeled data available, and the remaining data as unlabeled data. In some cases, DCT even reached the best accuracy, for instance, 72.6% using half of the labeled data, compared to 74.4% using all the labeled data. DCT also consistently outperformed MT in almost all configurations. For instance, the most significant relative gains brought by DCT reached 12.2% on ESC-10, compared to 7.6% with MT. Our code is available online1. Léo Cances, Thomas Pellegrini |
ICASSP | 2 |
| 2021 | Fast Threshold Optimization for Multi-Label Audio Tagging Using Surrogate Gradient LearningabstractMulti-label audio tagging consists of assigning sets of tags to audio recordings. At inference time, thresholds are applied on the confidence scores outputted by a probabilistic classifier, in order to decide which classes are detected active. In this work, we consider having at disposal a trained classifier and we seek to automatically optimize the decision thresholds according to a performance metric of interest, in our case F-measure (micro-F1). We propose a new method, called SGL-Thresh for Surrogate Gradient Learning of Thresholds, that makes use of gradient descent. Since F1 is not differentiable, we propose to approximate the thresholding operation gradients with the gradients of a sigmoid function. We report experiments on three datasets, using state-of-the-art pre-trained deep neural networks. In all cases, SGLThresh outperformed three other approaches: a default threshold value (defThresh), an heuristic search algorithm and a method estimating F1 gradients numerically. It reached 54.9% F1 on AudioSet eval, compared to 50.7% with defThresh. SGLThresh is very fast and scalable to a large number of tags1. Thomas Pellegrini, Timothée Masquelier |
ICASSP | 1 |
| 2021 | Simulating Reading Mistakes for Child Speech Transformer-Based Phone RecognitionabstractInternational audience Lucile Gelin, Thomas Pellegrini, Julien Pinquier, Morgane Daniel |
Interspeech | 2 |
| 2021 | Deep-Learning-Based Central African Primate Species Classification with MixUp and SpecAugment
Thomas Pellegrini |
Interspeech | 1 |
| 2021 | Low-Activity Supervised Convolutional Spiking Neural Networks Applied to Speech Commands RecognitionabstractDeep Neural Networks (DNNs) are the current state-of-the-art models in many speech related tasks. There is a growing interest, though, for more biologically realistic, hardware friendly and energy efficient models, named Spiking Neural Networks (SNNs). Recently, it has been shown that SNNs can be trained efficiently, in a supervised manner, using backpropagation with a surrogate gradient trick. In this work, we report speech command (SC) recognition experiments using supervised SNNs. We explored the Leaky-Integrate-Fire (LIF) neuron model for this task, and show that a model comprised of stacked dilated convolution spiking layers can reach an error rate very close to standard DNNs on the Google SC v1 dataset: 5.5%, while keeping a very sparse spiking activity, below 5%, thank to a new regularization term. We also show that modeling the leakage of the neuron membrane potential is useful, since the LIF model outperformed its non-leaky model counterpart significantly. Thomas Pellegrini, Romain Zimmer, Timothée Masquelier |
SLT | 1 |
| 2021 | End-to-end acoustic modelling for phone recognition of young readers
Lucile Gelin, Morgane Daniel, Julien Pinquier, Thomas Pellegrini |
Speech Commun. | 4 |
| 2019 | Cosine-similarity penalty to discriminate sound classes in weakly-supervised sound event detectionabstractThe design of new methods and models when only weakly-labeled data are available is of paramount importance in order to reduce the costs of manual annotation and the considerable human effort associated with it. In this work, we address Sound Event Detection in the case where a weakly annotated dataset is available for training. The weak annotations provide tags of audio events but do not provide temporal boundaries. The objective is twofold: 1) audio tagging, i.e. multi-label classification at recording level, 2) sound event detection, i.e. localization of the event boundaries within the recordings. This work focuses mainly on the second objective. We explore an approach inspired by Multiple Instance Learning, in which we train a convolutional recurrent neural network to give predictions at frame-level, using a custom loss function based on the weak labels and the statistics of the frame-based predictions. Since some sound classes cannot be distinguished with this approach, we improve the method by penalizing similarity between the predictions of the positive classes during training. On the test set used in the DCASE 2018 challenge, consisting of 288 recordings and 10 sound classes, the addition of a penalty resulted in a localization F-score of 34.75%, and brought 10% relative improvement compared to not using the penalty. Our best model achieved a 26.20% F-score on the DCASE-2018 official Eval subset close to the 10-system ensemble approach that ranked second in the challenge with a 29.9% F-score. Thomas Pellegrini, Léo Cances |
IJCNN | 1 |
| 2019 | Char+CV-CTC: Combining Graphemes and Consonant/Vowel Units for CTC-Based ASR Using Multitask LearningabstractInternational audience Abdelwahab Heba, Thomas Pellegrini, Jean-Pierre Lorré, Régine André-Obrecht |
INTERSPEECH | 2 |
| 2019 | The Airbus Air Traffic Control Speech Recognition 2018 Challenge: Towards ATC Automatic Transcription and Call Sign DetectionabstractIn this paper, we describe the outcomes of the challenge organized and run by Airbus and partners in 2018. The challenge consisted of two tasks applied to Air Traffic Control (ATC) speech in English: 1) automatic speech-to-text transcription, 2) call sign detection (CSD). The registered participants were provided with 40 hours of speech along with manual transcriptions. Twenty-two teams submitted predictions on a five hour evaluation set. ATC speech processing is challenging for several reasons: high speech rate, foreign-accented speech with a great diversity of accents, noisy communication channels. The best ranked team achieved a 7.62% Word Error Rate and a 82.41% CSD F1-score. Transcribing pilots' speech was found to be twice as harder as controllers' speech. Remaining issues towards solving ATC ASR are also discussed. Thomas Pellegrini, Jérôme Farinas, Estelle Delpech, François Lancelot |
INTERSPEECH | 1 |
| 2018 | Group emotion recognition strategies for entertainment robotsabstractIn this paper, a system to determine the emotion of a group of people via facial expression analysis is proposed for the Waseda Entertainment Robots. General models and standard methods for emotion definition and recognition are briefly described, as well as strategies for computing the group global emotion, knowing the individual emotions of group members. This work is based on Ekman's extended “Big Six” emotional model, popular in Computer Science and Affective Computing. Emotion recognition via facial expression analysis is performed with a cloud-computing based solution, using Microsoft Azure Cognitive services. First, the performances of both the Face API to detect faces, and Emotion API, to compute emotion via face expression analysis, are tested. After that, a solution to compute the emotion of a group of people has been implemented and its performances compared to human perceptions. This work presents concepts and strategies which can be generalized for applications within the scope of assistive robotics and, more broadly, affective computing, wherever it will be necessary to determine the emotion of a group of people. Sarah Cosentino, Estelle I. S. Randria, Jia-Yeu Lin, Thomas Pellegrini, Salvatore Sessa 0001, Atsuo Takanishi |
IROS | 4 |
| 2016 | Sinusoidal Modelling for EcoacousticsabstractBiodiversity assessment is a central and urgent task, necessary to monitoring the changes to ecological systems and under- standing the factors which drive these changes. Technological advances are providing new approaches to monitoring, which are particularly useful in remote regions. Situated within the framework of the emerging field of ecoacoustics, there is grow- ing interest in the possibility of extracting ecological informa- tion from digital recordings of the acoustic environment. Rather than focusing on identification of individual species, an increas- ing number of automated indices attempt to summarise acoustic activity at the community level, in order to provide a proxy for biodiversity. Originally designed for speech processing, sinu- soidal modelling has previously been used as a bioacoustic tool, for example to detect particular bird species. In this paper, we demonstrate the use of sinusoidal modelling as a proxy for bird abundance. Using data from acoustic surveys made during the breeding season in UK woodland, the number of extracted sinusoidal tracks is shown to correlate with estimates of bird abundance made by expert ornithologists listening to the recordings. We also report ongoing work exploring a new approach to investigate the composition of calls in spectro-temporal space that constitutes a promising new method for Ecoaoustic biodiversity assessment. Patrice Guyot, Alice C. Eldridge, Ying Chen Eyre-Walker, Alison Johnston, Thomas Pellegrini, Mika Peck |
INTERSPEECH | 5 |
| 2016 | Pronunciation Assessment of Japanese Learners of French with GOP Scores and Phonetic InformationabstractIn this paper, we report automatic pronunciation assessment experiments at phone-level on a read speech corpus in French, collected from 23 Japanese speakers learning French as a foreign language. We compare the standard approach based on Goodness Of Pronunciation (GOP) scores and phone-specific score thresholds to the use of logistic regressions (LR) models. French native speech corpus, in which artificial pronunciation errors were introduced, was used as training set. Two typical errors of Japanese speakers were considered: /ö/ and /v/ of ten mispronounced as [l] and [b], respectively. The LR classifier achieved a 64.4% accuracy similar to the 63.8% accuracy of the baseline threshold method, when using GOP scores and the expected phone identity as input features only. A significant performance gain of 20.8% relative was obtained by adding phonetic and phonological features as input to the LR model, leading to a 77.1% accuracy. This LR model also outperformed another baseline approach based on linear discriminant models trained on raw f-BANK coefficient features. Vincent Laborde, Thomas Pellegrini, Lionel Fontan, Julie Mauclair, Halima Sahraoui, Jérôme Farinas |
INTERSPEECH | 2 |
| 2016 | CNN-Based Phone Segmentation Experiments in a Less-Represented LanguageabstractThese last years, there has been a regain of interest in unsupervised sub-lexical and lexical unit discovery. Speech segmentation into phone-like units may be a first interesting step for such a task. In this article, we report speech segmentation experiments in Xitsonga, a less-represented language spoken in South Africa. We chose to use convolutional neural networks (CNN) with FBANK static coefficients as input. The models take binary decisions whether a boundary is present or not at each signal sliding frame. We compare the use of a model trained exclusively on Xitsonga data to the use of a bootstrap model trained on a larger corpus of another language, the BUCKEYE U.S. English corpus. Using a two-convolution-layer model, a 79% F-measure was obtained on BUCKEYE, with a 20 ms error tolerance. This performance is equal to the human inter-annotator agreement rate. We then used this bootstrap model to segment Xitsonga data and compared the results when adapting it with 1 to 20 minutes of Xitsonga data. Céline Manenti, Thomas Pellegrini, Julien Pinquier |
INTERSPEECH | 2 |
| 2016 | Inferring Phonemic Classes from CNN Activation Maps Using Clustering TechniquesabstractNational audience Thomas Pellegrini, Sandrine Mouysset |
INTERSPEECH | 1 |
| 2015 | Comparing SVM, softmax, and shallow neural networks for eating condition classificationabstractInternational audience Thomas Pellegrini |
INTERSPEECH | 1 |
| 2014 | The goodness of pronunciation algorithm applied to disordered speechabstractInternational audience Thomas Pellegrini, Lionel Fontan, Julie Mauclair, Jérôme Farinas, Marina Robert |
INTERSPEECH | 1 |
| 2014 | Speaker age estimation for elderly speech recognition in European PortugueseabstractInternational audience Thomas Pellegrini, Vahid Hedayati, Isabel Trancoso, Annika Hämäläinen, José Miguel Salles Dias |
INTERSPEECH | 1 |
| 2014 | Segmentation in singer turns with the Bayesian information criterionabstractAs part of a project on indexing ethno-musicological audio recordings, segmentation in singer turns automatically appeared to be essential. In this article, we present the problem of segmentation in singer turns of musical recordings and our first experiments in this direction by exploring a method based on the Bayesian Information Criterion (BIC), which are used in numerous works in audio segmentation, to detect singer turns. The BIC penalty coefficient was shown to vary when determining its value to achieve the best performance for each recording. In order to avoid the decision about which single value is best for all the documents, we propose to combine several segmentations obtained with different values of this parameter. This method consists of taking a posteriori decisions on which segment boundaries are to be kept. A gain of 7.1% in terms of F-measure was obtained compared to a standard coefficient. Marwa Thlithi, Thomas Pellegrini, Julien Pinquier, Régine André-Obrecht |
INTERSPEECH | 2 |
| 2014 | El-WOZ: a client-server wizard-of-oz interface
Thomas Pellegrini, Vahid Hedayati, Ângela Costa |
LREC | 1 |
| 2013 | A corpus-based study of elderly and young speakers of European Portuguese: acoustic correlates and their impact on speech recognition performanceabstractThis paper presents a study of European Portuguese elderly speech, in which the acoustic characteristics of two groups of elderly speakers (aged 60-75 and over 75) are compared with those of young adult speakers (aged 19-30). The correlation between age and a set of 14 acoustic features was investigated, and decision trees were used to establish the relative importance of the features. A greater use of pauses characterized speakers aged 60 and over. For female speakers, speech rate also appeared to correlate with age. For male speakers, jitter distinguished between speakers aged 60-75 and older. The correlation between the features and speech recognition performance was also investigated. Word error rate correlated mostly with the use of pauses, speech rate, and the ratio of long phone realizations. Finally, by comparing the phone sequences used by the recognizer on the most frequent words, we observed that the young adult speakers reduced schwas more than the elderly speakers. This result seems to confirm the common idea that young speakers reduce articulation more than older speakers. Further investigation is needed to confirm this result by determining whether this is due to ageing or to the generation gap. Thomas Pellegrini, Annika Hämäläinen, Philippe Boula de Mareüil, Michael Tjalve, Isabel Trancoso, Sara Candeias, José Miguel Salles Dias, Daniela Braga |
INTERSPEECH | 1 |
| 2013 | ASR-based exercises for listening comprehension practice in European Portuguese
Thomas Pellegrini, Rui Correia, Isabel Trancoso, Jorge Baptista, Nuno J. Mamede, Maxine Eskénazi |
Comput. Speech Lang. | 1 |
| 2012 | Overview of Computer-assisted Language Learning for European Portuguese at L2f
Thomas Pellegrini, Wang Ling, Rui Correia, Isabel Trancoso, Jorge Baptista, Nuno J. Mamede |
CSEDU (2) | 1 |
| 2012 | Less errors with TTS? A dictation experiment with foreign language learners
Thomas Pellegrini, Ângela Costa, Isabel Trancoso |
INTERSPEECH | 1 |
| 2011 | Automatic Generation of Listening Comprehension Learning Material in European PortugueseabstractThe goal of this work is the automatic selection of materials for a listening comprehension game. We would like to select automatically transcribed sentences from recent broadcast news corpora, in order to gather material for the games with little human effort. The recognized words are used as the ground solution of the exercises, thus sentences with misrecognitions need to be filtered out. Our experiments confirmed the feasibility of the filter chain that automatically selects sentences, although harder confidence thresholds may be needed. Together with the correct words, wrong candidates, namely distractors, are also needed to build the exercises. Two techniques of distractor generation are presented, either based on the confusion networks produced by the recognizer, or on phonetic distances. The experiments confirmed the complementarity of both approaches. Index Terms: CALL, Listening Comprehension, European Portuguese, ASR, distractors Thomas Pellegrini, Rui Correia, Isabel Trancoso, Jorge Baptista, Nuno J. Mamede |
INTERSPEECH | 1 |
| 2010 | Context dependent modelling approaches for hybrid speech recognizersabstractSpeech recognition based on connectionist approaches is one of the most successful alternatives to widespread Gaussian systems. One of the main claims against hybrid recognizers is the increased complexity for context-dependent phone modeling, which is a key aspect in medium to large size vocabulary tasks. In this paper, we investigate the use of context-dependent triphone models in a connectionist speech recognizer. Thus, most common triphone state clustering procedures for Gaussian models are compared and applied to our hybrid recognizer. The developed systems with clustered context-dependent triphones show above 20% relative word error rate reduction compared to a baseline hybrid system in two selected WSJ evaluation test sets. Additionally, the recent porting efforts of the proposed context modelling approaches to a LVCSR system for English Broadcast News transcription are reported. Index Terms: speech recognition, context modeling, connectionist system Alberto Abad, Thomas Pellegrini, Isabel Trancoso, João Paulo da Silva Neto |
INTERSPEECH | 2 |
| 2010 | Improving ASR error detection with non-decoder based featuresabstractThis study reports error detection experiments in large vocabulary automatic speech recognition (ASR) systems, by using statistical classifiers. We explored new features gathered from other knowledge sources than the decoder itself: a binary feature that compares outputs from two different ASR systems (word by word), a feature based on the number of hits of the hypothesized bigrams, obtained by queries entered into a very popular Web search engine, and finally a feature related to automatically infered topics at sentence and word levels. Experiments were conducted on a European Portuguese broadcast news corpus. The combination of baseline decoder-based features and two of these additional features led to significant improvements, from 13.87% to 12.16% classification error rate (CER) with a maximum entropy model, and from 14.01% to 12.39% CER with linear-chain conditional random fields, comparing to a baseline using only decoder-based features. Thomas Pellegrini, Isabel Trancoso |
INTERSPEECH | 1 |
| 2010 | Multimedia learning materialsabstractThis paper describes the integration of multimedia documents in the Portuguese version of REAP, a tutoring system for vocabulary learning. The documents result from the pipeline processing of Broadcast News videos that automatically segments the audio files, transcribes them, adds punctuation and capitalization, and breaks them into stories classified by topics. The integration of these materials in REAP was done in a way that tries to decrease the impact of potential errors of the automatic chain in the learning process. José Lopes 0001, Isabel Trancoso, Rui Correia, Thomas Pellegrini, Hugo Meinedo, Nuno J. Mamede, Maxine Eskénazi |
SLT | 4 |
| 2009 | Audio contributions to semantic video searchabstractThis paper summarizes the contributions to semantic video search that can be derived from the audio signal. Because of space restrictions, the emphasis will be on non-linguistic cues. The paper thus covers what is generally known as audio segmentation, as well as audio event detection. Using machine learning approaches, we have built detectors for over 50 semantic audio concepts. Isabel Trancoso, Thomas Pellegrini, José Portelo, Hugo Meinedo, Miguel M. F. Bugalho, Alberto Abad, João Paulo da Silva Neto |
ICME | 2 |
| 2009 | Detecting audio events for semantic video searchabstractThis paper describes our work on audio event detection, one of our tasks in the European project VIDIVIDEO. Preliminary experiments with a small corpus of sound effects have shown the potential of this type of corpus for training purposes. This paper describes our experiments with SVM classifiers, and different features, using a 290-hour corpus of sound effects, which allowed us to build detectors for almost 50 semantic concepts. Although the performance of these detectors on the development set is quite good (achieving an average F-measure of 0.87), preliminary experiments on documentaries and films showed that the task is much harder in real-life videos, which so often include overlapping audio events. Index Terms: event detection, audio segmentation 1. Miguel M. F. Bugalho, José Portelo, Isabel Trancoso, Thomas Pellegrini, Alberto Abad |
INTERSPEECH | 4 |
| 2009 | Automatic Word Decompounding for ASR in a Morphologically Rich Language: Application to AmharicabstractThis paper investigates a data-driven word decompounding algorithm for use in automatic speech recognition. An existing algorithm, called ldquoMorfessor,rdquo has been enhanced in order to address the problem of increased phonetic confusability arising from word decompounding by incorporating phonetic properties and some constraints on recognition units derived from forced alignments experiments. Speech recognition experiments have been carried out on a broadcast news task for the Amharic language to validate the approach. The out of vocabulary (OOV) word rates were reduced by 35% to 50% and a small reduction in word error rate (WER) has been achieved. The algorithm is relatively language independent and requires minimal adaptation to be applied to other languages. Thomas Pellegrini, Lori Lamel |
IEEE Trans. Speech Audio Process. | 1 |
| 2008 | Developments of "Lëtzebuergesch" Resources for Automatic Speech Processing and Linguistic Studies
Martine Adda-Decker, Thomas Pellegrini, Eric Bilinski, Gilles Adda |
LREC | 2 |
| 2007 | Using phonetic features in unsupervised word decompounding for ASR with application to a less-represented languageabstractIn this paper, a data-driven word decompounding algorithm is described and applied to a broadcast news corpus in Amharic. The baseline algorithm has been enhanced in order to address the problem of increased phonetic confusability arising from word decompounding by incorporating phonetic properties and some constraints on recognition units derived from prior forced alignment experiments. Speech recognition experiments have been carried out to validate the approach. Out of vocabulary (OOV) words rates can be reduced by 30% to 40% and an absolute Word Error Rate (WER) reduction of 0.4% has been achieved. The algorithm is relatively language independent and requires minimal adaptation to be applied to other languages. Index Terms: automatic speech recognition, unsupervised word decompounding, less-represented languages Thomas Pellegrini, Lori Lamel |
INTERSPEECH | 1 |
| 2006 | Investigating automatic decomposition for ASR in less represented languagesabstractThis paper addresses the use of an automatic decomposition method to reduce lexical variety and thereby improve speech recognition of less well-represented languages. The Amharic language has been selected for these experiments since only a small quantity of resources are available compared to well-covered languages. Inspired by the Harris algorithm, the method automatically generates plausible affixes, that combined with decompounding can reduce the size of the lexicon and the OOV rate. Recognition experiments are carried out for four different configurations (full-word and decompounded) and using supervised training with a corpus containing only two hours of manually transcribed data. Thomas Pellegrini, Lori Lamel |
INTERSPEECH | 1 |
| 2006 | Experimental detection of vowel pronunciation variants in Amharic
Thomas Pellegrini, Lori Lamel |
LREC | 1 |