Isabel Trancoso

dblp:72/6829 · DBLP profile ↗
← Back
175ranked-venue papers
27as first author
25since 2021 · last 2026
0000-0001-5874-6313ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 134 · 23 first-author · 18 since 2021Artificial intelligence and machine learning · 132 · 15 first-author · 21 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 FalAR: A Large-scale Speaker-Annotated European Portuguese Speech Corpus of Parliamentary Sessions
Francisco Teixeira, Carlos Carvalho 0003, Mariana Julião, Catarina Botelho, Rubén Solera-Ureña, Sérgio Paulo, Thomas Rolland, Ben Peters, Isabel Trancoso, Alberto Abad
LREC9
2026 Exploring features for membership inference in ASR model auditing
Francisco Teixeira, Karla Pizzi, Raphaël Olivier, Alberto Abad, Bhiksha Raj, Isabel Trancoso
Comput. Speech Lang.6
2025 CAMÕES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese
abstract
Existing resources for Automatic Speech Recognition in Portuguese are mostly focused on Brazilian Portuguese, leaving European Portuguese (EP) and other varieties underexplored. To bridge this gap, we introduce CAMÕES, the first open framework for EP and other Portuguese varieties. It consists of (1) a comprehensive evaluation benchmark, including 46 h of EP test data spanning multiple domains; and (2) a collection of state-of-the-art models. For the latter, we consider multiple foundation models, evaluating their zero-shot and fine-tuned performances, as well as E-Branchformer models trained from scratch. A curated set of $\mathbf{4 2 5 h}$ of EP was used for both fine-tuning and training. Our results show comparable performance for EP between fine-tuned foundation models and the E-Branchformer. Furthermore, the best-performing models achieve relative improvements above 35% WER, compared to the strongest zero-shot foundation model, establishing a new state-of-the-art for EP and other varieties.
Carlos Carvalho 0003, Francisco Teixeira, Catarina Botelho, Anna Pompili, Rubén Solera-Ureña, Sérgio Paulo, Mariana Julião, Thomas Rolland, John Mendonça, Diogo A. P. Nunes, Isabel Trancoso, Alberto Abad
ASRU11
2025 Acoustic and Linguistic Biomarkers for Cognitive Impairment Detection from Speech
Catarina Botelho, David Gimeno-Gómez, Francisco Teixeira, John Mendonça, Patrícia Pereira, Diogo A. P. Nunes, Thomas Rolland, Anna Pompili, Rubén Solera-Ureña, Maria Ponte, David Martins de Matos, Carlos D. Martínez-Hinarejos, Isabel Trancoso, Alberto Abad
INTERSPEECH13
2025 Children's Voice Privacy: First Steps and Emerging Challenges
abstract
International audience
Ajinkya Kulkarni, Francisco Teixeira, Enno Hermann, Thomas Rolland, Isabel Trancoso, Mathew Magimai-Doss
INTERSPEECH5
2025 Speech Reference Intervals: An Assessment of Feasibility in Depression Symptom Severity Prediction
abstract
Major Depressive Disorder (MDD) is a prevalent mental disorder. Combining speech features and machine learning has promise for predicting MDD, but interpretability is crucial for clinical applications. Reference intervals (RIs) represent a typical range for a speech feature in a population. RIs could increase interpretability and help clinicians identify deviations from norms. They could also replace conventional speech features in machine learning models. However, no work has yet assessed the feasibility of speech RIs in MDD. We generated and compared RIs from three reference datasets varying in size, elicitation prompt, and health information. We then calculated deviations from each RI set for people with MDD to compare performance on a depression symptom severity prediction task. Our RI-based models trained with demographic data performed similarly to each other and equivalent models using conventional features or demographics only, demonstrating the value of RI-derived features.
Lauren L. White, Ewan Carr, Judith Dineley, Catarina Botelho, Pauline Conde, Faith Matcham, Carolin Oetzmann, Amos Folarin, George Fairs, Agnes Norbury, Stefano Goria, Srinivasan Vairavan, Til Wykes, Richard J. B. Dobson, Vaibhav A. Narayan, Matthew Hotopf, Alberto Abad, Isabel Trancoso, Nicholas Cummins
INTERSPEECH18
2025 Towards Cyberbullying Detection: Building, Benchmarking and Longitudinal Analysis of Aggressiveness and Conflicts/Attacks Datasets From Twitter
abstract
Offense and hate speech are a source of online conflicts which have become common in social media and, as such, their study is a growing topic of research in machine learning and natural language processing. This article presents two Portuguese language offense-related datasets that deepen the study of the subject: an Aggressiveness dataset and a Conflicts/Attacks dataset. While the former is similar to other offense detection related datasets, the latter constitutes a novelty due to the use of the history of the interaction between users. Several studies were carried out to construct and analyze the data in the datasets. The first study included gathering expressions of verbal aggression witnessed by adolescents to guide data extraction for the datasets. The second study included extracting data from Twitter (in Portuguese) that matched the most frequent expressions/words/sentences that were identified in the previous study. The third study consisted in the development of the Aggressiveness dataset, the Conflicts/Attacks dataset, and classification models. In our fourth study, we proposed to examine whether online aggression and conflicts/attacks revealed any trend changes over time with a sample of 86 adolescents. With this study, we also proposed to investigate whether the amount of tweets sent over a period of 273 days was related to online aggression and conflicts/attacks. Finally, we analyzed the percentage of participants who participated in the aggressions and/or attacks/conflicts.
Paula Ferreira 0004, Nádia Salgado Pereira, Hugo Rosa, Sofia Oliveira, Luísa Coheur, Sofia Mateus Francisco, Sidclay Bezerra de Souza, Ricardo Ribeiro 0001, João Paulo Carvalho 0001, Paula Paulino, Isabel Trancoso, Ana Margarida Veiga Simão
IEEE Trans. Affect. Comput.11
2024 Macro-descriptors for Alzheimer's disease detection using large language models
Catarina Botelho, John Mendonça, Anna Pompili, Tanja Schultz, Alberto Abad, Isabel Trancoso
INTERSPEECH6
2024 Unveiling Biases while Embracing Sustainability: Assessing the Dual Challenges of Automatic Speech Recognition Systems
abstract
In this paper, we present a bias and sustainability focused investigation of Automatic Speech Recognition (ASR) systems, namely Whisper and Massively Multilingual Speech (MMS), which have achieved state-of-the-art (SOTA) performances. Despite their improved performance in controlled settings, there remains a critical gap in understanding their efficacy and equity in real-world scenarios. We analyze ASR biases w.r.t. gender, accent, and age group, as well as their effect on downstream tasks. In addition, we examine the environmental impact of ASR systems, scrutinizing the use of large acoustic models on carbon emission and energy consumption. We also provide insights into our empirical analyses, offering a valuable contribution to the claims surrounding bias and sustainability in ASR systems.
Ajinkya Kulkarni, Atharva Kulkarni, Miguel Couceiro, Isabel Trancoso
INTERSPEECH4
2024 Towards Responsible Speech Processing
Isabel Trancoso
INTERSPEECH1
2024 ECoh: Turn-level Coherence Evaluation for Multilingual Dialogues
abstract
Despite being heralded as the new standard for dialogue evaluation, the closed-source nature of OpenAI's GPT-4 model poses challenges for the research community.Motivated by the need for lightweight, open source, and multilingual automated dialogue evaluators, this paper introduces GENRESCOH (Generated Responses targeting Coherence).GENRESCOH is a novel LLM-generated dataset comprising over 130k negative and positive responses and accompanying explanations seeded from XDailyDialog and XPersona covering English, French, German, Italian, and Chinese.Leveraging GEN-RESCOH, we propose ECOH 1 (Evaluation of Coherence), a family of evaluators trained to assess response coherence across multiple languages.Experimental results demonstrate that ECOH achieves multilingual coherence detection capabilities superior to the teacher model (GPT-3.5-Turbo)on GENRESCOH, despite being based on a much smaller architecture.Furthermore, the explanations provided by ECOH closely align in terms of quality with those generated by the teacher model.
John Mendonça, Isabel Trancoso, Alon Lavie
SIGDIAL2
2023 Privacy-Preserving Automatic Speaker Diarization
abstract
Automatic Speaker Diarization (ASD) is an enabling technology with numerous applications, which deals with recordings of multiple speakers, raising special concerns in terms of privacy. In fact, in remote settings, where recordings are shared with a server, clients relinquish not only the privacy of their conversation, but also of all the information that can be inferred from their voices. However, to the best of our knowledge, the development of privacy-preserving ASD systems has been overlooked thus far. In this work, we tackle this problem using a combination of two cryptographic techniques, Secure Multiparty Computation (SMC) and Secure Modular Hashing, and apply them to the two main steps of a cascaded ASD system: speaker embedding extraction and agglomerative hierarchical clustering. Our system is able to achieve a reasonable trade-off between performance and efficiency, presenting real-time factors of 1.1 and 1.6, for two different SMC security settings.
Francisco Teixeira, Alberto Abad, Bhiksha Raj, Isabel Trancoso
ICASSP4
2023 Towards Reference Speech Characterization for Health Applications
Catarina Botelho, Alberto Abad, Tanja Schultz, Isabel Trancoso
INTERSPEECH4
2023 Laughter in task-based settings: whom we talk to affects how, when, and how often we laugh
abstract
Map task corpora are not typically used to study laughter, but they allow an interesting analysis of multiple factors such as familiarity between the participants, their gender, and eye contact. We conducted linear/generalized mixed-effects analysis to study if co-laughter, laughter rate, and the percentage of voiced frames in laughs are influenced by such factors. Our results show that, in conversations without eye contact, the gender of the participant was statistically relevant regarding laughter rate and the percentage of voiced frames, and the difference in gender was relevant regarding co-laughter. On the other hand, with eye contact, familiarity was statistically relevant with respect to co-laughter, laughter rate, and the percentage of voiced frames. Most of our results align and extend what has been previously found, except for voiced laughs between friends. This study emphasizes the highly variable character of laughter and its dependence on interlocutors' characteristics.
Catarina Branco, Isabel Trancoso, Paulo Infante, Khiet P. Truong
INTERSPEECH2
2023 Prediction of the Gender-based Violence Victim Condition using Speech: What do Machine Learning Models rely on?
Emma Reyner-Fuentes, Esther Rituerto-González, Isabel Trancoso, Carmen Peláez-Moreno
INTERSPEECH3
2023 Towards Multilingual Automatic Open-Domain Dialogue Evaluation
abstract
The main limiting factor in the development of robust multilingual open-domain dialogue evaluation metrics is the lack of multilingual data and the limited availability of open-sourced multilingual dialogue systems.In this work, we propose a workaround for this lack of data by leveraging a strong multilingual pretrained encoder-based Language Model and augmenting existing English dialogue data using Machine Translation.We empirically show that the naive approach of finetuning a pretrained multilingual encoder model with translated data is insufficient to outperform the strong baseline of finetuning a multilingual model with only source data.Instead, the best approach consists in the careful curation of translated data using MT Quality Estimation metrics, excluding low quality translations that hinder its performance.
John Mendonça, Alon Lavie, Isabel Trancoso
SIGDIAL3
2022 Exploring Dementia Detection from Speech: Cross Corpus Analysis
abstract
In this work, we present a qualitative and quantitative analysis of speech and language features derived from two different corpora with the aim to predict early signs of dementia. One corpus consists of the Interdisciplinary Longitudinal Study on Adult Development and Aging (ILSE) designed to investigate satisfying and healthy aging. It consists of more than 6500 hours of biographic interviews from 1000 participants recorded over the course of 20 years. The other corpus is a cross sectional data set created for the ADReSS challenge 2020. In an experimental study we describe a large variety of acoustic and linguistic features that are automatically extracted from speech and corresponding transcriptions. We compare different traditional classifiers, i.e. Gaussian Mixture Models, Linear Discriminant Analysis, and Support Vector Machines. Our final performance results surpass the ADReSS benchmarks.
Ayimnisagul Ablimit, Catarina Botelho, Alberto Abad, Tanja Schultz, Isabel Trancoso
ICASSP5
2022 Challenges of using longitudinal and cross-domain corpora on studies of pathological speech
Catarina Botelho, Tanja Schultz, Alberto Abad, Isabel Trancoso
INTERSPEECH4
2022 Towards End-to-End Private Automatic Speaker Recognition
abstract
The development of privacy-preserving automatic speaker verification systems has been the focus of a number of studies with the intent of allowing users to authenticate themselves without risking the privacy of their voice. However, current privacy-preserving methods assume that the template voice representations (or speaker embeddings) used for authentication are extracted locally by the user. This poses two important issues: first, knowledge of the speaker embedding extraction model may create security and robustness liabilities for the authentication system, as this knowledge might help attackers in crafting adversarial examples able to mislead the system; second, from the point of view of a service provider the speaker embedding extraction model is arguably one of the most valuable components in the system and, as such, disclosing it would be highly undesirable. In this work, we show how speaker embeddings can be extracted while keeping both the speaker's voice and the service provider's model private, using Secure Multiparty Computation. Further, we show that it is possible to obtain reasonable trade-offs between security and computational cost. This work is complementary to those showing how authentication may be performed privately, and thus can be considered as another step towards fully private automatic speaker recognition.
Francisco Teixeira, Alberto Abad, Bhiksha Raj, Isabel Trancoso
INTERSPEECH4
2022 Towards Speaker Verification for Crowdsourced Speech Collections
abstract
Crowdsourcing the collection of speech provides a scalable setting to access a customisable demographic according to each dataset’s needs. The correctness of speaker metadata is especially relevant for speaker-centred collections - ones that require the collection of a fixed amount of data per speaker. This paper identifies two different types of misalignment present in these collections: Multiple Accounts misalignment (different contributors map to the same speaker), and Multiple Speakers misalignment (multiple speakers map to the same contributor). Based on state-of-the-art approaches to Speaker Verification, this paper proposes an unsupervised method for measuring speaker metadata plausibility of a collection, i.e., evaluating the match (or lack thereof) between contributors and speakers. The solution presented is composed of an embedding extractor and a clustering module. Results indicate high precision in automatically classifying contributor alignment (>0.94).
John Mendonça, Rui Correia, Mariana Lourenço, João Freitas, Isabel Trancoso
LREC5
2022 QualityAdapt: an Automatic Dialogue Quality Estimation Framework
abstract
Despite considerable advances in open-domain neural dialogue systems, their evaluation remains a bottleneck.Several automated metrics have been proposed to evaluate these systems, however, they mostly focus on a single notion of quality, or, when they do combine several sub-metrics, they are computationally expensive.This paper attempts to solve the latter: QualityAdapt leverages the Adapter framework for the task of Dialogue Quality Estimation.Using well defined semi-supervised tasks, we train Adapters for different subqualities and score generated responses with Adapter-Fusion.This compositionality provides an easy to adapt metric to the task at hand that incorporates multiple subqualities.It also reduces computational costs as individual predictions of all subqualities are obtained in a single forward pass.This approach achieves comparable results to state-of-the-art metrics on several datasets, whilst keeping the previously mentioned advantages.
John Mendonça, Alon Lavie, Isabel Trancoso
SIGDIAL3
2021 The in-the-Wild Speech Medical Corpus
abstract
Automatic detection of speech affecting (SA) diseases has received significant attention, particularly in clinical scenarios. However, the same task in in-the-wild conditions is often neglected, in part, due to the lack of appropriate datasets.In this work, we present the in-the-Wild Speech Medical (WSM) Corpus, a collection of in-the-wild videos, featuring subjects potentially affected by a SA disease - specifically, depression or Parkinson’s disease. The WSM Corpus contains a total 928 videos, and over 131 hours of speech. Each video is accompanied by a crowdsourced annotation for perceived age/gender, and self-reported health status of the speaker. The WSM Corpus is balanced over all the labels.In this work we present a detailed description of the collection, and annotation processes of the WSM corpus. Furthermore, we present present several baseline systems for the detection of SA diseases using speech alone, thus motivating the use of this type of in-the-wild data in paralinguistic audiovisual tasks.
Maria Joana Correia, Francisco Teixeira, Catarina Botelho, Isabel Trancoso, Bhiksha Raj
ICASSP4
2021 FoolHD: Fooling Speaker Identification by Highly Imperceptible Adversarial Disturbances
abstract
Speaker identification models are vulnerable to carefully designed adversarial perturbations of their input signals that induce misclassification. In this work, we propose a white-box steganography-inspired adversarial attack that generates imperceptible adversarial perturbations against a speaker identification model. Our approach, FoolHD, uses a Gated Convolutional Autoencoder that operates in the DCT domain and is trained with a multi-objective loss function, to generate and conceal the adversarial perturbation within the original audio files. In addition to hindering speaker identification performance, this multi-objective loss accounts for human perception through a frame-wise cosine similarity between MFCC feature vectors extracted from the original and adversarial audio files. We validate the effectiveness of FoolHD with a 250-speaker identification x-vector network, trained using VoxCeleb, in terms of accuracy, success rate, and imperceptibility. Our results show that FoolHD generates highly imperceptible adversarial audio files (average PESQ scores above 4.30), while achieving a success rate of 99.6% and 99.2% in misleading the speaker identification model, for untargeted and targeted settings, respectively.
Ali Shahin Shamsabadi, Francisco Teixeira, Alberto Abad, Bhiksha Raj, Andrea Cavallaro, Isabel Trancoso
ICASSP6
2021 Visual Speech for Obstructive Sleep Apnea Detection
Catarina Botelho, Alberto Abad, Tanja Schultz, Isabel Trancoso
Interspeech4
2021 Transfer Learning-Based Cough Representations for Automatic Detection of COVID-19
abstract
In the last months, there has been an increasing interest in developing reliable, cost-effective, immediate and easy to use machine learning based tools that can help health care operators, institutions, companies, etc. to optimize their screening campaigns.In this line, several initiatives emerged aimed at the automatic detection of COVID-19 from speech, breathing and coughs, with inconclusive preliminary results.The ComParE 2021 COVID-19 Cough Sub-challenge provides researchers from all over the world a suitable test-bed for the evaluation and comparison of their work.In this paper, we present the INESC-ID contribution to the ComParE 2021 COVID-19 Cough Sub-challenge.We leverage transfer learning to develop a set of three expert classifiers based on deep cough representation extractors.A calibrated decision-level fusion system provides the final classification of coughs recordings as either COVID-19 positive or negative.Results show unweighted average recalls of 72.3% and 69.3% in the development and test sets, respectively.Overall, the experimental assessment shows the potential of this approach although much more research on extended respiratory sounds datasets is needed.
Rubén Solera-Ureña, Catarina Botelho, Francisco Teixeira, Thomas Rolland, Alberto Abad, Isabel Trancoso
Interspeech6
2020 Toward Silent Paralinguistics: Speech-to-EMG - Retrieving Articulatory Muscle Activity from Speech
abstract
Electromyographic (EMG) signals recorded during speech production encode information on articulatory muscle activity and also on the facial expression of emotion, thus representing a speech-related biosignal with strong potential for paralinguistic applications.In this work, we estimate the electrical activity of the muscles responsible for speech articulation directly from the speech signal.To this end, we first perform a neural conversion of speech features into electromyographic time domain features, and then attempt to retrieve the original EMG signal from the time domain features.We propose a feed forward neural network to address the first step of the problem (speech features to EMG features) and a neural network composed of a convolutional block and a bidirectional long short-term memory block to address the second problem (true EMG features to EMG signal).We observe that four out of the five originally proposed time domain features can be estimated reasonably well from the speech signal.Further, the five time domain features are able to predict the original speech-related EMG signal with a concordance correlation coefficient of 0.663.We further compare our results with the ones achieved on the inverse problem of generating acoustic speech features from EMG features.
Catarina Botelho, Lorenz Diener, Dennis Küster, Kevin Scheck, Shahin Amiriparian, Björn W. Schuller, Tanja Schultz, Alberto Abad, Isabel Trancoso
INTERSPEECH9
2020 Towards Silent Paralinguistics: Deriving Speaking Mode and Speaker ID from Electromyographic Signals
abstract
Silent Computational Paralinguistics (SCP) -the assessment of speaker states and traits from non-audibly spoken communication -has rarely been targeted in the rich body of either Computational Paralinguistics or Silent Speech Processing.Here, we provide first steps towards this challenging but potentially highly rewarding endeavour: Paralinguistics can enrich spoken language interfaces, while Silent Speech Processing enables confidential and unobtrusive spoken communication for everybody, including mute speakers.We approach SCP by using speech-related biosignals stemming from facial muscle activities captured by surface electromyography (EMG).To demonstrate the feasibility of SCP, we select one speaker trait (speaker identity) and one speaker state (speaking mode).We introduce two promising strategies for SCP: (1) deriving paralinguistic speaker information directly from EMG of silently produced speech versus (2) first converting EMG into an audible speech signal followed by conventional computational paralinguistic methods.We compare traditional feature extraction and decision making approaches to more recent deep representation and transfer learning by convolutional and recurrent neural networks, using openly available EMG data.We find that paralinguistics can be assessed not only from acoustic speech but also from silent speech captured by EMG.
Lorenz Diener, Shahin Amiriparian, Catarina Botelho, Kevin Scheck, Dennis Küster, Isabel Trancoso, Björn W. Schuller, Tanja Schultz
INTERSPEECH6
2020 Analyzing Breath Signals for the Interspeech 2020 ComParE Challenge
John Mendonça, Francisco Teixeira, Isabel Trancoso, Alberto Abad
INTERSPEECH3
2020 Automatic In-the-wild Dataset Annotation with Deep Generalized Multiple Instance Learning
abstract
The automation of the diagnosis and monitoring of speech affecting diseases in real life situations, such as Depression or Parkinson’s disease, depends on the existence of rich and large datasets that resemble real life conditions, such as those collected from in-the-wild multimedia repositories like YouTube. However, the cost of manually labeling these large datasets can be prohibitive. In this work, we propose to overcome this problem by automating the annotation process, without any requirements for human intervention. We formulate the annotation problem as a Multiple Instance Learning (MIL) problem, and propose a novel solution that is based on end-to-end differentiable neural networks. Our solution has the additional advantage of generalizing the MIL framework to more scenarios where the data is stil organized in bags but does not meet the MIL bag label conditions. We demonstrate the performance of the proposed method in labeling the in-the-Wild Speech Medical (WSM) Corpus, using simple textual cues extracted from videos and their metadata. Furthermore we show what is the contribution of each type of textual cues for the final model performance, as well as study the influence of the size of the bags of instances in determining the difficulty of the learning problem
Maria Joana Correia, Isabel Trancoso, Bhiksha Raj
LREC2
2019 In-the-Wild End-to-End Detection of Speech Affecting Diseases
abstract
Speech is a complex bio-signal that has the potential to provide a rich bio-marker for health. It enables the development of non-invasive routes to early diagnosis and monitoring of speech affecting diseases, such as the ones studied in this work: Depression, and Parkinson's Disease. However, the major limitation of current speech based diagnosis and monitoring tools is the lack of large and diverse datasets. Existing datasets are small, and collected under very controlled conditions. As such, there is an upper bound in the complexity of the models that can be trained using these datasets. There is also limited applicability in real life scenarios where the channel and noise conditions, among others, are impossible to control. In this work, we show that datasets collected from in-the-wild sources, such as collections of vlogs, can contribute to improve the performance of diagnosis tools both in controlled and in-the-wild conditions, even though the data are noisier. Moreover, we show that it is possible to successfully move away from hand-crafted features (i.e. features that are computed based on predefined algorithms, that based on human expertise) and adopt end-to-end modeling paradigms, such as CNN-LSTMs, that extract data driven features from the raw spectrograms of the speech signal, and capture temporal information from the speech signals.
Maria Joana Correia, Isabel Trancoso, Bhiksha Raj
ASRU2
2019 Speech as a Biomarker for Obstructive Sleep Apnea Detection
abstract
Obstructive sleep apnea (OSA) is a prevalent sleep disorder, responsible for a decrease of people's quality of life, and significant morbidity and mortality associated with hypertension and cardiovascular diseases. OSA is caused by anatomical and functional alterations in the upper airways, thus we hypothesize that the speech properties of OSA patients are altered, making it possible to detect OSA through voice analysis. To address this hypothesis, we collected speech recordings from 25 OSA subjects and 20 controls, designed a feature set, and compared different machine learning algorithms for binary classification. We achieved a True-Positive-Rate of 88% and a True-Negative-Rate of 80% with a majority vote ensemble of SVM, LDA and kNN classifiers. These results were validated with in-the-wild data acquired from Youtube. Moreover, the negative impact of sleep disorders on working memory was also shown by the results obtained in one of the recorded verbal tasks.
Catarina Botelho, Isabel Trancoso, Alberto Abad, Teresa Paiva
ICASSP2
2019 Privacy-preserving Paralinguistic Tasks
abstract
Speech is one of the primary means of communication for humans. It can be viewed as a carrier for information on several levels as it conveys not only the meaning and intention predetermined by a speaker, but also paralinguistic and extralinguistic information about the speaker's age, gender, personality, emotional state, health state and affect. This makes it a particularly sensitive biometric, that should be protected. In this work we intent to explore how Leveled Homomorphic Encryption can be combined with a Neural Network to create a privacy-preserving machine learning framework for speech-based health-related tasks. In particular, we will apply this framework to the detection and assessment of a Cold, Depression and Parkinson's Disease. Moreover, we will show how using a Quantized Neural Network, with discretized weights, allows us to apply a Leveled Homomorphic Encryption technique called batching that can be utilized to reduce the effective computational cost of this framework.
Francisco Teixeira, Alberto Abad, Isabel Trancoso
ICASSP3
2019 Unbabel Talk - Human Verified Translations for Voice Instant Messaging
Luís Bernardo, Mathieu Giquel, Sebastião Quintas, Paulo Dimas, Helena Moniz, Isabel Trancoso
INTERSPEECH6
2019 Recognition of Latin American Spanish Using Multi-Task Learning
Carlos Mendes, Alberto Abad, João Paulo da Silva Neto, Isabel Trancoso
INTERSPEECH4
2019 The GDPR & Speech Data: Reflections of Legal and Technology Communities, First Steps Towards a Common Understanding
abstract
International audience
Andreas Nautsch, Catherine Jasserand, Els Kindt, Massimiliano Todisco, Isabel Trancoso, Nicholas W. D. Evans
INTERSPEECH5
2018 Exploring Hashing and Cryptonet Based Approaches for Privacy-Preserving Speech Emotion Recognition
abstract
The outsourcing of machine learning classification and data mining tasks can be an effective solution for those parties that need machine learning services, but lack the appropriate resources, knowledge and/or tools to carry them out, in their own premises. This solution, however, raises major privacy concerns, in particular, when irrevocable biometric data such as speech is involved. In this work, we focus on the development of privacy-preserving schemes in a speech emotion recognition task, as a proof of concept that could be extended to other speech analytics tasks. Our aim is to prove that the implementation of privacy-preserving speech mining schemes in challenging tasks involving paralinguistic features is not only feasible, but also accurate. Using distance-preserving hashing techniques in a first approach, and homomorphic encryption in a second approach, we successfully protect sensitive data with little degradation costs regarding the accuracy of the predictive models.
José Miguel Salles Dias, Alberto Abad, Isabel Trancoso
ICASSP3
2018 Acoustic-prosodic Entrainment in Structural Metadata Events
abstract
This paper presents an acoustic-prosodic analysis of entrain- ment in a Portuguese map-task corpus. Our aim is to ana- lyze how turn-by-turn entrainment varies with distinct structural metadata events: types of sentence-like units (SU) in consecu- tive turns (e.g. interrogatives followed by declaratives, or both declaratives), and with the presence of discourse markers, affir- mative cue words, and disfluencies in the beginning of turns. Entrainment at turn-exchanges may be observed in terms of pitch, energy, duration, and voice quality. Regarding SU types, question-answer turns are the ones with stronger similarity, and declarative-interrogative pairs are the ones where less entrain- ment occurs, as expected. Moreover, in question-answer pairs, there is also stronger evidence of entrainment with Yes/No and Tag questions than with Wh- questions. In fact, these subtypes are coded in distinctive prosodic ways (moreover, the first sub- type has no associated lexical-syntactic cues in Portuguese, only prosodic). As for turn-initial structures, entrainment is stronger when the second turn begins with an affirmative cue word; less strong with ambiguous structures (such as ‘OK’), emphatic af- firmative answers, and negative answers; and scarce with dis- fluencies and discourse markers. The different degrees of local entrainment may be related with the informative structure of distinct structural metadata events.
Vera Cabarrão, Fernando Batista, Helena Moniz, Isabel Trancoso, Ana Isabel Mata
INTERSPEECH4
2018 Mining Multimodal Repositories for Speech Affecting Diseases
Maria Joana Correia, Bhiksha Raj, Isabel Trancoso, Francisco Teixeira
INTERSPEECH3
2018 Patient Privacy in Paralinguistic Tasks
Francisco Teixeira, Alberto Abad, Isabel Trancoso
INTERSPEECH3
2018 Querying Depression Vlogs
abstract
Speech based diagnosis-aid tools for depression typically depend on few and small datasets, that are expensive to collect. The limited availability of training data poses a limitation to the quality that these systems can achieve. An unexplored alternative for large scale source of data are vlogs collected from online multimedia repositories. Along with the automation of the mining process, it is necessary to automate the labeling process too.In this work, we propose a framework to automatically label a corpus of in-the-wild vlogs of possibly depressed subjects, and we estimate the quality of the predicted labels, without ever having access to a ground truth for the majority of the corpus. The framework uses a small subset to train a model and estimate the labels for the remainder of the corpus. Then, using the predicted labels, we train a noisy model and attempt to reconstruct the labels of the original labeled subset. We hypothesize that the quality of the estimated labels for the unlabelled subset of the corpus is correlated to the quality of the label reconstruction of the labeled subset.The results of the bi-modal experiment using in-the-wild data are compared to the ones obtained using controlled data.
Maria Joana Correia, Bhiksha Raj, Isabel Trancoso
SLT3
2017 Speech Technologies: Reaching Maturity?
Isabel Trancoso
ICPRAM1
2017 Segment Level Voice Conversion with Recurrent Neural Networks
Miguel Varela Ramos, Alan W. Black, Ramón Fernandez Astudillo, Isabel Trancoso, Nuno Fonseca
INTERSPEECH4
2017 A Semi-Supervised Learning Approach for Acoustic-Prosodic Personality Perception in Under-Resourced Domains
abstract
Automatic personality analysis has gained attention in the last years as a fundamental dimension in human-To-human and human-To-machine interaction. However, it still suffers from limited number and size of speech corpora for specific domains, such as the assessment of children's personality. This paper investigates a semi-supervised training approach to tackle this scenario. We devise an experimental setup with age and language mismatch and two training sets: A small labeled training set from the Interspeech 2012 Personality Sub-challenge, containing French adult speech labeled with personality OCEAN traits, and a large unlabeled training set of Portuguese children's speech. As test set, a corpus of Portuguese children's speech labeled with OCEAN traits is used. Based on this setting, we investigate a weak supervision approach that iteratively refines an initial model trained with the labeled data-set using the unlabeled data-set. We also investigate knowledge-based features, which leverage expert knowledge in acoustic-prosodic cues and thus need no extra data. Results show that, despite the large mismatch imposed by language and age differences, it is possible to attain improvements with these techniques, pointing both to the benefits of using a weak supervision and expert-based acoustic-prosodic features across age and language.
Rubén Solera-Ureña, Helena Moniz, Fernando Batista, Vera Cabarrão, Anna Pompili, Ramón Fernandez Astudillo, Joana Campos 0001, Ana Paiva 0001, Isabel Trancoso
INTERSPEECH9
2016 Exploiting Phone Log-Likelihood Ratio Features for the Detection of the Native Language of Non-Native English Speakers
Alberto Abad, Eugénio Ribeiro, Fábio N. Kepler, Ramón Fernandez Astudillo, Isabel Trancoso
INTERSPEECH5
2016 SPA: Web-based Platform for easy Access to Speech Processing Modules
Fernando Batista, Pedro Curto, Isabel Trancoso, Alberto Abad, Jaime Ferreira, Eugénio Ribeiro, Helena Moniz, David Martins de Matos, Ricardo Ribeiro 0001
LREC3
2016 Adaptation of SVM for MIL for inferring the polarity of movies and movie reviews
abstract
Polarity detection is a research topic of major interest, with many applications including detecting the polarity of product reviews. However, in some cases, the polarity of the product reviews might not be available while the polarity of the product itself might be, prohibiting the use of any form of fully supervised learning technique. This scenario, while different, is close to that of multiple instance learning (MIL). In this work we propose two new adaptations of support vector machines (SVM) for MIL, θ-MIL, to suit this new scenario, and infer the polarity of products and product reviews. We perform experiments on the proposed methods using the IMDb movie review corpus, and compare the performance of the proposed methods to the traditional SVM for MIL approach. Although we make weaker assumptions about the data, the proposed methods achieve a comparable performance to the SVM for MIL in accurately detecting the polarity of movies and movie reviews.
Maria Joana Correia, Isabel Trancoso, Bhiksha Raj
SLT2
2016 Mining Parallel Corpora from Sina Weibo and Twitter
abstract
Microblogs such as Twitter, Facebook, and Sina Weibo (China's equivalent of Twitter) are a remarkable linguistic resource. In contrast to content from edited genres such as newswire, microblogs contain discussions of virtually every topic by numerous individuals in different languages and dialects and in different styles. In this work, we show that some microblog users post “self-translated” messages targeting audiences who speak different languages, either by writing the same message in multiple languages or by retweeting translations of their original posts in a second language. We introduce a method for finding and extracting this naturally occurring parallel data. Identifying the parallel content requires solving an alignment problem, and we give an optimally efficient dynamic programming algorithm for this. Using our method, we extract nearly 3M Chinese–English parallel segments from Sina Weibo using a targeted crawl of Weibo users who post in multiple languages. Additionally, from a random sample of Twitter, we obtain substantial amounts of parallel data in multiple language pairs. Evaluation is performed by assessing the accuracy of our extraction approach relative to a manual annotation as well as in terms of utility as training data for a Chinese–English machine translation system. Relative to traditional parallel data resources, the automatically extracted parallel data yield substantial translation quality improvements in translating microblog text and modest improvements in translating edited news content.
Wang Ling, Luís Marujo, Chris Dyer, Alan W. Black, Isabel Trancoso
Comput. Linguistics5
2015 Learning Word Representations from Scarce and Noisy Data with Embedding Subspaces
abstract
Ramon F. Astudillo, Silvio Amir, Wang Ling, Mário Silva, Isabel Trancoso. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Ramón Fernandez Astudillo, Silvio Amir, Wang Ling, Mário J. Silva, Isabel Trancoso
ACL (1)5
2015 Finding Function in Form: Compositional Character Models for Open Vocabulary Word Representation
abstract
Wang Ling, Chris Dyer, Alan W Black, Isabel Trancoso, Ramón Fermandez, Silvio Amir, Luís Marujo, Tiago Luís. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015.
Wang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso, Ramon Fermandez, Silvio Amir, Luís Marujo, Tiago Luís
EMNLP4
2015 Not All Contexts Are Created Equal: Better Word Representations with Variable Attention
abstract
Wang Ling, Yulia Tsvetkov, Silvio Amir, Ramón Fermandez, Chris Dyer, Alan W Black, Isabel Trancoso, Chu-Cheng Lin. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing. 2015.
Wang Ling, Yulia Tsvetkov, Silvio Amir, Ramon Fermandez, Chris Dyer, Alan W. Black, Isabel Trancoso, Chu-Cheng Lin
EMNLP7
2015 Privacy-preserving Query-by-Example Speech Search
abstract
This paper investigates a new privacy-preserving paradigm for the task of Query-by-Example Speech Search using Secure Binary Embeddings, a hashing method that converts vector data to bit strings through a combination of random projections followed by banded quantization. The proposed method allows performing spoken query search in an encrypted domain, by analyzing ciphered information computed from the original recordings. Unlike other hashing techniques, the embeddings allow the computation of the distance between vectors that are close enough, but are not perfect matches. This paper shows how these hashes can be combined with Dynamic Time Warping based on posterior derived features to perform secure speech search. Experiments performed on a sub-set of the Speech-Dat Portuguese corpus showed that the proposed privacy-preserving system obtains similar results to its non-private counterpart.
José Portelo, Alberto Abad, Bhiksha Raj, Isabel Trancoso
ICASSP4
2015 Integration of DNN based speech enhancement and ASR
abstract
Speech enhancement employing Deep Neural Networks (DNNs) is gaining strength as a data-driven alternative to classical Minimum Mean Square Error (MMSE) enhancement approaches. In the past, Observation Uncertainty approaches to integrate MMSE speech enhancement with Automatic Speech Recognition (ASR) have yielded good results as a lightweight alternative for robust ASR. In this paper we thus explore the integration of DNN-based speech enhancement with ASR by employing Observation Uncertainty techniques. For this purpose, we explore various techniques and approximations that allow propagating the uncertainty of inference of the DNN into feature domain. This uncertainty can then be used to dynamically compensate the ASR model utilizing techniques like uncertainty decoding. We test the proposed techniques on the AURORA4 corpus and show that notable improvements can be attained over the already effective DNN enhancement.
Ramón Fernandez Astudillo, Maria Joana Correia, Isabel Trancoso
INTERSPEECH3
2015 Detecting repetitions in spoken dialogue systems using phonetic distances
abstract
This paper addresses the problem of automatic detection of re-peated turns in Spoken Dialogue Systems. Repetitions can be a symptom of problematic communication between users and systems. Such repetitions are often due to speech recognition errors, which in turn makes it hard to use speech recognition to detect repetitions. We present an approach to detect rep-etition using the phonetic distance to find the best alignment between turns in the same dialogue. The alignment score ob-tained is combined with different features to improve repeti-tion detection. To evaluate the method proposed we compare several alignment techniques from edit distance to DTW-based distance, previously used in Spoken-Term detection tasks. We also compare two different methods to compute the phonetic distance: the first one using the phoneme sequence, and the second one using the distance between the phone posterior vec-tors. Two different datasets were used in this evaluation: a bus-schedule information system (in English) and a call routing system (in Swedish). The results show that approaches using phoneme distances over-perform approaches using Levenshtein distances between ASR outputs for repetition detection. Index Terms: spoken dialogue systems, repetition detection, phonetic distance
José Lopes 0001, Giampiero Salvi, Gabriel Skantze, Alberto Abad, Joakim Gustafson, Fernando Batista, Raveesh Meena, Isabel Trancoso
INTERSPEECH8
2015 Combining multiple approaches to predict the degree of nativeness
abstract
Automatic speaker nativeness assessment has multiple applications, such as second language learning and IVR systems. In this paper we view this as a regression problem, since the available labels are on a continuous scale. Multiple approaches were applied, such as phonotactic models, i-vectors, and goodness of pronunciation, covering both segmental and suprasegmental features. Different phonotactic models were adopted, either trained with the challenge data, or using additional multilingual data from other domains. The obtained values were later combined in multiple ways and fed to a support vector machine regressor. Results on the test set surpass the provided baseline and are in line with the results obtained on the remaining sets. This suggests that our models generalize well to other datasets
Eugénio Ribeiro, Jaime Ferreira, Julia Olcoz, Alberto Abad, Helena Moniz, Fernando Batista, Isabel Trancoso
INTERSPEECH7
2015 Two/Too Simple Adaptations of Word2Vec for Syntax Problems
abstract
Wang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015.
Wang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso
HLT-NAACL4
2015 From rule-based to data-driven lexical entrainment models in spoken dialog systems
José Lopes 0001, Maxine Eskénazi, Isabel Trancoso
Comput. Speech Lang.3
2015 Introduction to the Special Section on Continuous Space and Related Methods in Natural Language Processing
abstract
The articles in this special section discuss some latest findings on research problems related to the application of continuous space and related models in Natural Language Processing (NLP).
Haizhou Li 0001, Marcello Federico, Xiaodong He 0001, Helen M. Meng, Isabel Trancoso
IEEE ACM Trans. Audio Speech Lang. Process.5
2015 Graph-Based Lexicon Regularization for PCFG With Latent Annotations
abstract
This paper aims at learning a better probabilistic context-free grammar with latent annotations (PCFG-LA) by using a graph propagation (GP) technique. We propose leveraging the GP to regularize the lexical model of the grammar. The proposed approach constructs k-nearest neighbor ( k-NN) similarity graphs over words with identical pre-terminal (part-of-speech) tags, for propagating the probabilities of latent annotations given the words. The graphs demonstrate the relationship between words in syntactic and semantic levels, estimated by using a neural word representation method based on Recursive autoencoder (RAE). We modify the conventional PCFG-LA parameter estimation algorithm, expectation maximization (EM), by incorporating a GP process subsequent to the M-step. The GP encourages the smoothness among the graph vertices, where different words under similar syntactic and semantic environments should have approximate posterior distributions of nonterminal subcategories. The proposed PCFG-LA learning approach was evaluated together with a hierarchical split-and-merge training strategy, on parsing tasks for English, Chinese and Portuguese. The empirical results reveal two crucial findings: 1) regularizing the lexicons with GP results in positive effects to parsing accuracy; and 2) learning with unlabeled data can also expand the PCFG-LA lexicons.
Xiaodong Zeng, Derek F. Wong, Lidia S. Chao, Isabel Trancoso
IEEE ACM Trans. Audio Speech Lang. Process.4
2014 Toward Better Chinese Word Segmentation for SMT via Bilingual Constraints
abstract
This study investigates on building a better Chinese word segmentation model for statistical machine translation.It aims at leveraging word boundary information, automatically learned by bilingual character-based alignments, to induce a preferable segmentation model.We propose dealing with the induced word boundaries as soft constraints to bias the continuous learning of a supervised CRFs model, trained by the treebank data (labeled), on the bilingual data (unlabeled).The induced word boundary information is encoded as a graph propagation constraint.The constrained model induction is accomplished by using posterior regularization algorithm.The experiments on a Chinese-to-English machine translation task reveal that the proposed model can bring positive segmentation effects to translation quality.
Xiaodong Zeng, Lidia S. Chao, Derek F. Wong, Isabel Trancoso
ACL (1)4
2014 Accounting for the residual uncertainty of multi-layer perceptron based features
abstract
Multi-Layer Perceptrons (MLPs) are often interpreted as modeling a posterior distribution over classes given input features using the mean field approximation. This approximation is fast but neglects the residual uncertainty of inference at each layer, making inference less robust. In this paper we introduce a new approximation of MLP inference that takes under consideration this residual uncertainty. The proposed algorithm propagates not only the mean, but also the variance of inference through the network. At the current stage, the proposed method can not be used with soft-max layers. Therefore, we illustrate the benefits of this algorithm in a tandem scheme. We use the residual uncertainty of inference of MLP-based features to compensate a GMM-HMM backend with uncertainty decoding. Experiments on the Aurora4 corpus show consistent improvement of performance against conventional MLPs for all scenarios, in particular for clean speech and multi-style training.
Ramón Fernandez Astudillo, Alberto Abad, Isabel Trancoso
ICASSP3
2014 Verbal description of LEGO blocks
abstract
Query specification for 3D object retrieval still relies on traditional interaction paradigms. The goal of our study was to identify the most natural methods to describe 3D objects, focusing on verbal and gestural expressions. Our case study uses LEGO R blocks. We started by collecting a corpus involving ten pairs of subjects, in which one participant requests blocks for building a model from another participant. This small corpus suggests that users prefer to describe 3D objects verbally, rarely resorting to gestures, and using them only as complement. The paper describes this corpus, addressing the challenges that such verbal descriptions create for a speech understanding system, namely the long complex verbal descriptions, involving dimensions, shapes, colors, metaphors, and diminutives. The latter connote small size, endearment or insignificance, and are only very common in informal language. In this corpus, they occurred in one out of seven requests. This experiment was the first step of the development of a prototype for searching LEGO R blocks combining speech and stereoscopic 3D. Although the verbal interaction in the first version is limited to relatively simple queries, its combination with immersive visualization allows the user to explore query results in a dataset with virtual blocks. Index Terms: multimodal corpus, 3D objects, voice search
Diogo Henriques, Isabel Trancoso, Daniel Mendes, Alfredo Ferreira
INTERSPEECH2
2014 Speaker age estimation for elderly speech recognition in European Portuguese
abstract
International audience
Thomas Pellegrini, Vahid Hedayati, Isabel Trancoso, Annika Hämäläinen, José Miguel Salles Dias
INTERSPEECH3
2014 OpenLogos Semantico-Syntactic Knowledge-Rich Bilingual Dictionaries
Anabela Barreiro, Fernando Batista, Ricardo Ribeiro 0001, Helena Moniz, Isabel Trancoso
LREC5
2014 Linguistic Evaluation of Support Verb Constructions by OpenLogos and Google Translate
Anabela Barreiro, Johanna Monti, Brigitte Orliac, Susanne Preuß, Kutz Arrieta, Wang Ling, Fernando Batista, Isabel Trancoso
LREC8
2014 Revising the annotation of a Broadcast News corpus: a linguistic approach
Vera Cabarrão, Helena Moniz, Fernando Batista, Ricardo Ribeiro 0001, Nuno J. Mamede, Hugo Meinedo, Isabel Trancoso, Ana Isabel Mata, David Martins de Matos
LREC7
2014 Exploiting magnitude and phase spectral information for converted speech detection
abstract
Speaker verification systems have been shown to be vulnerable in situations where voice conversion techniques are used to try to fool them, evidencing an important security breach in these applications. This work focuses on the development of a new converted speech detector able to robustly address this problem. The proposed detector uses four spectral features extracted from the magnitude and the phase spectrum of the speech signal. To evaluate the performance of the detector we use a subset of the core task of the NIST SRE2006 corpus as the natural data. The converted data was produced with two different voice conversion methods: Gaussian mixture model and unit selection, from other NIST SRE2006 conditions. The converted speech detector achieved a detection accuracy of 99.1% and 98.5% for natural and converted utterances, respectively.
Maria Joana Correia, Alberto Abad, Isabel Trancoso
SLT3
2014 Lexicon expansion for latent variable grammars
Xiaodong Zeng, Derek F. Wong, Lidia S. Chao, Isabel Trancoso, Liangye He, Qiuping Huang
Pattern Recognit. Lett.4
2014 Speaking style effects in the production of disfluencies
Helena Moniz, Fernando Batista, Ana Isabel Mata, Isabel Trancoso
Speech Commun.4
2013 Microblogs as Parallel Corpora
Wang Ling, Guang Xiang, Chris Dyer, Alan W. Black, Isabel Trancoso
ACL (1)5
2013 Graph-based Semi-Supervised Model for Joint Chinese Word Segmentation and Part-of-Speech Tagging
Xiaodong Zeng, Derek F. Wong, Lidia S. Chao, Isabel Trancoso
ACL (1)4
2013 Paraphrasing 4 Microblog Normalization
abstract
Compared to the edited genres that have played a central role in NLP research, microblog texts use a more informal register with nonstandard lexical items, abbreviations, and free orthographic variation.When confronted with such input, conventional text analysis tools often perform poorly.Normalization -replacing orthographically or lexically idiosyncratic forms with more standard variants -can improve performance.We propose a method for learning normalization rules from machine translations of a parallel corpus of microblog messages.To validate the utility of our approach, we evaluate extrinsically, showing that normalizing English tweets and then translating improves translation quality (compared to translating unnormalized text) using three standard web translation services as well as a phrase-based translation system trained on parallel microblog data.
Wang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso
EMNLP4
2013 Automated two-way entrainment to improve spoken dialog system performance
abstract
This paper proposes an approach to the use of lexical entrainment in Spoken Dialog Systems. This approach aims to increase the dialog success rate by adapting the lexical choices of the system to the user's lexical choices. If the system finds that the users lexical choice degrades the performance, it will try to establish a new conceptual pact, proposing other words that the user may adopt, in order to be more successful in task completion. The approach was implemented and tested in two different systems. Tests showed a relative dialog estimated error rate reduction of 10% and a relative reduction in the average number of turns per session of 6%.
José Lopes 0001, Maxine Eskénazi, Isabel Trancoso
ICASSP3
2013 Disfluency detection based on prosodic features for university lectures
abstract
This paper focuses on the identification of disfluent sequences and their distinct structural regions, based on acoustic and prosodic features. Reported experiments are based on a corpus of university lectures in European Portuguese, with roughly 32h, and a relatively high percentage of disfluencies (7.6%). The set of features automatically extracted from the corpus proved to be discriminant of the regions contained in the production of a disfluency. Several machine learning methods have been applied, but the best results were achieved using Classification and Regression Trees (CART). The set of features which was most informative for cross-region identification encompasses word duration ratios, word confidence score, silent ratios, and pitch and energy slopes. Features such as the number of phones and syllables per word proved to be more useful for the identification of the interregnum, whereas energy slopes were most suited for identifying the interruption point.
Henrique Medeiros, Helena Moniz, Fernando Batista, Isabel Trancoso, Luís Nunes
INTERSPEECH4
2013 A corpus-based study of elderly and young speakers of European Portuguese: acoustic correlates and their impact on speech recognition performance
abstract
This paper presents a study of European Portuguese elderly speech, in which the acoustic characteristics of two groups of elderly speakers (aged 60-75 and over 75) are compared with those of young adult speakers (aged 19-30). The correlation between age and a set of 14 acoustic features was investigated, and decision trees were used to establish the relative importance of the features. A greater use of pauses characterized speakers aged 60 and over. For female speakers, speech rate also appeared to correlate with age. For male speakers, jitter distinguished between speakers aged 60-75 and older. The correlation between the features and speech recognition performance was also investigated. Word error rate correlated mostly with the use of pauses, speech rate, and the ratio of long phone realizations. Finally, by comparing the phone sequences used by the recognizer on the most frequent words, we observed that the young adult speakers reduced schwas more than the elderly speakers. This result seems to confirm the common idea that young speakers reduce articulation more than older speakers. Further investigation is needed to confirm this result by determining whether this is due to ageing or to the generation gap.
Thomas Pellegrini, Annika Hämäläinen, Philippe Boula de Mareüil, Michael Tjalve, Isabel Trancoso, Sara Candeias, José Miguel Salles Dias, Daniela Braga
INTERSPEECH5
2013 Secure binary embeddings of front-end factor analysis for privacy preserving speaker verification
abstract
Remote speaker verification services typically rely on the system to have access to the users recordings, or features derived from them, and also a model of the users voice. This conventional scheme raises several privacy concerns. In this work, we address this privacy problem in the context of a speaker verification system using a factor analysis based front-end extractor, the so-called i-vectors. Speaker verification without exposing speaker data is achieved by transforming speaker i-vectors to bit strings in a way that allows the computation of approximate distances, instead of exact ones. The key to the transformation uses a hashing scheme known as Secure Binary Embeddings. Then, a modified SVM kernel permits operating on the i-vector hashes. Experiments on sub-sets of NIST SRE 2008 showed that the secure system yielded similar results as its non-private counterpart.
José Portelo, Alberto Abad, Bhiksha Raj, Isabel Trancoso
INTERSPEECH4
2013 euTV: a system for media monitoring and publishing
abstract
In this paper, we describe the euTV system, which provides a flexible approach to collect, manage, annotate and publish collections of images, videos and textual documents. The system is based on a Service Oriented Architecture that allows to combine and orchestrate a large set of web services for automatic and manual annotation, retrieval, browsing, ingestion and authoring of multimedia sources. euTV tools have been used to create several publicly available vertical applications, addressing different use cases. Positive results of user evaluations have shown that the system can be effectively used to create different types of applications.
Marco Bertini 0001, Alberto Del Bimbo, George Ioannidis, Emile Bijk, Isabel Trancoso, Hugo Meinedo
ACM Multimedia5
2013 Automatic word naming recognition for an on-line aphasia treatment system
Alberto Abad, Anna Pompili, Ângela Costa, Isabel Trancoso, José G. Fonseca, Gabriela Leal, Luisa Farrajota, Isabel P. Martins
Comput. Speech Lang.4
2013 ASR-based exercises for listening comprehension practice in European Portuguese
Thomas Pellegrini, Rui Correia, Isabel Trancoso, Jorge Baptista, Nuno J. Mamede, Maxine Eskénazi
Comput. Speech Lang.3
2012 Overview of Computer-assisted Language Learning for European Portuguese at L2f
Thomas Pellegrini, Wang Ling, Rui Correia, Isabel Trancoso, Jorge Baptista, Nuno J. Mamede
CSEDU (2)5
2012 Entropy-based Pruning for Phrase-based Machine Translation
Wang Ling, João Graça, Isabel Trancoso, Alan W. Black
EMNLP-CoNLL3
2012 Attacking a privacy preserving music matching algorithm
abstract
Secure multi-party computation based techniques are often used to perform audio database search tasks, such as music matching, with privacy. However, in spite of the security of individual components of the matching schemes, the overall scheme may still not be secure. This paper explains how such flaws may occur, using a privacy preserving music matching problem as a template, and provides a solution, and analyzes the resulting tradeoff between privacy and computational complexity. Although the paper focus on a music matching application, the principles can be easily adapted to perform other tasks, such as speaker verification and keyword spotting.
José Portelo, Bhiksha Raj, Isabel Trancoso
ICASSP3
2012 Automatic word naming recognition for treatment and assessment of aphasia
abstract
VITHEA is an on-line platform designed to act as a “virtual therapist” for the treatment of Portuguese speaking aphasic patients. Concretely, the system integrates automatic speech recognition technology to provide word naming exercises to individuals with lost or reduced word naming ability. In this paper, we present the solution adopted for the word naming task, which is based on a keyword spotting approach with hybrid HMM/MLP speech recognizer. Furthermore, we explore a simple cross-validation method that makes use of the patients measured word naming ability to automatically adapt to their speech particularities. A corpus with word naming therapy sessions of aphasic Portuguese native speakers has been collected to test the utility of the approach for both global evaluation and treatment. In spite of the different patient characteristics and speech quality conditions of the collected data, encouraging results have been obtained.
Alberto Abad, Anna Pompili, Ângela Costa, Isabel Trancoso
INTERSPEECH4
2012 Text-dependent pathological voice detection
abstract
While global characteristics of the speaker’s source and spectral features have been successfully employed in pathological voice detection, the underlying text has largely been ignored. In this work, we focus on experiments that exploit the text stimulus that is read by the subject. Features derived from text include the mean cepstral distortion of the subject from an average intelligible speaker, and prosodic features include the speaking rate, statistics of phoneme durations, etc. The phonetic labeling information is also exploited to ignore all the unvoiced regions of the speech samples to improve the discriminability between intelligible and pathological voices. We also designed features that capture the speaker’s overall closeness to intelligible instances of the same text stimulus from other speakers. Our experiments show that the proposed text-derived features improve the detection of pathological voices by 20%. Index Terms: Pathological voices, example based detection, text-driven features, fusion of classification methods.
Gopala Krishna Anumanchipalli, Hugo Meinedo, Miguel M. F. Bugalho, Isabel Trancoso, Luís C. Oliveira, Alan W. Black
INTERSPEECH4
2012 Prosodic contex-based analysis of disfluencies
abstract
This work explores prosodic cues of disfluencies in a corpus of university lectures. Results show three significant (p < 0.001) trends: pitch and energy slopes are significantly different between the disfluency and the onset of fluency; those features are also relevant to disfluency type differentiation; and they do not seem to be a speakereffect. The best combination of linguistic features one can use to better predict the onset of fluency are pitch and energy resets as well as the presence of a silent pause immediately before a repair. Our results, thus, point out to a strategy of prosodic contrast rather than of parallelism. With this work we hope to contribute to the analysis of the prosodic behaviors in the production of the so called disfluencies and in the fluency repair in European Portuguese.
Helena Moniz, Fernando Batista, Isabel Trancoso, Ana Isabel Mata
INTERSPEECH3
2012 Less errors with TTS? A dictation experiment with foreign language learners
Thomas Pellegrini, Ângela Costa, Isabel Trancoso
INTERSPEECH3
2012 Privacy-Preserving Speaker Authentication
Manas A. Pathak, José Portelo, Bhiksha Raj, Isabel Trancoso
ISC4
2012 Dealing with unknown words in statistical machine translation
João Pedro Carlos Gomes da Silva, Luísa Coheur, Ângela Costa, Isabel Trancoso
LREC4
2012 Bilingual Experiments on Automatic Recovery of Capitalization and Punctuation of Automatic Speech Transcripts
abstract
This paper focuses on the tasks of recovering capitalization and punctuation marks from texts without that information, such as spoken transcripts, produced by automatic speech recognition systems. These two practical rich transcription tasks were performed using the same discriminative approach, based on maximum entropy, suitable for on-the-fly usage. Reported experiments were conducted both over Portuguese and English broadcast news data. Both force aligned and automatic transcripts were used, allowing to measure the impact of the speech recognition errors. Capitalized words and named entities are intrinsically related, and are influenced by time variation effects. For that reason, the so-called language dynamics have been addressed for the capitalization task. Language adaptation results indicate, for both languages, that the capitalization performance is affected by the temporal distance between the training and testing data. In what regards the punctuation task, this paper covers the three most frequent punctuation marks: full stop, comma, and question marks. Different methods were explored for improving the baseline results for full stop and comma. The first uses punctuation information extracted from large written corpora. The second applies different levels of linguistic structure, including lexical, prosodic, and speaker related features. The comma detection improved significantly in the first method, thus indicating that it depends more on lexical features. The second method provided even better results, for both languages and both punctuation marks, best results being achieved mainly for full stop. As for question marks, there is a small gain, but differences are not very significant, due to the relatively small number of question marks in the corpora.
Fernando Batista, Helena Moniz, Isabel Trancoso, Nuno J. Mamede
IEEE Trans. Speech Audio Process.3
2011 Towards choosing better primes for spoken dialog systems
abstract
When humans and computers use the same terms (primes, when they entrain to one another), spoken dialogs proceed more smoothly. The goal of this paper is to describe initial steps we have found that will enable us to eventually automatically choose better primes in spoken dialog system prompts. Two different sets of prompts were used to understand what makes one prime more suitable than another. The impact of the primes chosen in speech recognition was evaluated. In addition, results reveal that users did adopt the new vocabulary introduced in the new system prompts. As a result of this, performance of the system improved, providing clues for the trade off needed when choosing between adequate primes in prompts and speech recognition performance.
José Lopes 0001, Maxine Eskénazi, Isabel Trancoso
ASRU3
2011 Multi-site heterogeneous system fusions for the Albayzin 2010 Language Recognition Evaluation
abstract
Best language recognition performance is commonly obtained by fusing the scores of several heterogeneous systems. Regardless the fusion approach, it is assumed that different systems may contribute complementary information, either because they are developed on different datasets, or because they use different features or different modeling approaches. Most authors apply fusion as a final resource for improving performance based on an existing set of systems. Though relative performance gains decrease as larger sets of systems are considered, best performance is usually attained by fusing all the available systems, which may lead to high computational costs. In this paper, we aim to discover which technologies combine the best through fusion and to analyse the factors (data, features, modeling methodologies, etc.) that may explain such a good performance. Results are presented and discussed for a number of systems provided by the participating sites and the organizing team of the Albayzin 2010 Language Recognition Evaluation. We hope the conclusions of this work help research groups make better decisions in developing language recognition technology.
Luis Javier Rodríguez-Fuentes, Mikel Peñagarikano, Amparo Varona, Mireia Díez, Germán Bordel, David Martínez González, Jesús Villalba 0001, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida, Alberto Abad, Oscar Koller, Isabel Trancoso, Paula Lopez-Otero, Laura Docío Fernández, Carmen García-Mateo, Rahim Saeidi, Mehdi Soufifar, Tomi Kinnunen, Torbjørn Svendsen, Pasi Fränti
ASRU13
2011 BP2EP - Adaptation of Brazilian Portuguese texts to European Portuguese
Luís Marujo, Nuno Grazina, Tiago Luís, Wang Ling, Luísa Coheur, Isabel Trancoso
EAMT6
2011 Parallel Transformation Network features for speaker recognition
abstract
The use of speaker adaptation transforms as features for speaker recognition is an appealing alternative to conventional short-term cepstral features. In general, this kind of methods are language dependent and limited by the need of speech recognition in the client speakers language. In this paper, we generalize a recently pro posed method -named Transformation Network features with SVM modeling- in order to become language independent and overcome the need for accurate speech recognition. This is accomplished by using a set of parallel acoustic models in several different languages to obtain a high-dimensional Parallel Transformation Network feature vector for speaker characterization.
Alberto Abad, Jordi Luque, Isabel Trancoso
ICASSP3
2011 A nativeness classifier for TED Talks
abstract
This paper presents a nativeness classifier for English. The detector was developed and tested with TED Talks collected from the web, where the major non-native cues are in terms of segmental aspects and prosody. The first experiments were made using only acoustic features, with Gaussian supervectors for training a classifier based on support vector machines. These experiments resulted in an equal error rate of 13.11%. The following experiments based on prosodic features alone did not yield good results. However, a fused system, combining acoustic and prosodic cues, achieved an equal error rate of 10.58%. A small human benchmark was conducted, showing an inter-rater agreement of 0.88. This value is also very close to the agreement value between humans and the best fused system.
José Lopes 0001, Isabel Trancoso, Alberto Abad
ICASSP2
2011 Discriminative Phrase-based Lexicalized Reordering Models using Weighted Reordering Graphs
Wang Ling, João Graça, David Martins de Matos, Isabel Trancoso, Alan W. Black
IJCNLP4
2011 Automatic Generation of Listening Comprehension Learning Material in European Portuguese
abstract
The goal of this work is the automatic selection of materials for a listening comprehension game. We would like to select automatically transcribed sentences from recent broadcast news corpora, in order to gather material for the games with little human effort. The recognized words are used as the ground solution of the exercises, thus sentences with misrecognitions need to be filtered out. Our experiments confirmed the feasibility of the filter chain that automatically selects sentences, although harder confidence thresholds may be needed. Together with the correct words, wrong candidates, namely distractors, are also needed to build the exercises. Two techniques of distractor generation are presented, either based on the confusion networks produced by the recognizer, or on phonetic distances. The experiments confirmed the complementarity of both approaches. Index Terms: CALL, Listening Comprehension, European Portuguese, ASR, distractors
Thomas Pellegrini, Rui Correia, Isabel Trancoso, Jorge Baptista, Nuno J. Mamede
INTERSPEECH3
2011 Temporal Video Segmentation to Scenes Using High-Level Audiovisual Features
abstract
In this paper, a novel approach to video temporal decomposition into semantic units, termed scenes, is presented. In contrast to previous temporal segmentation approaches that employ mostly low-level visual or audiovisual features, we introduce a technique that jointly exploits low-level and high-level features automatically extracted from the visual and the auditory channel. This technique is built upon the well-known method of the scene transition graph (STG), first by introducing a new STG approximation that features reduced computational cost, and then by extending the unimodal STG-based temporal segmentation technique to a method for multimodal scene segmentation. The latter exploits, among others, the results of a large number of TRECVID-type trained visual concept detectors and audio event detectors, and is based on a probabilistic merging process that combines multiple individual STGs while at the same time diminishing the need for selecting and fine-tuning several STG construction parameters. The proposed approach is evaluated on three test datasets, comprising TRECVID documentary films, movies, and news-related videos, respectively. The experimental results demonstrate the improved performance of the proposed approach in comparison to other unimodal and multimodal techniques of the relevant literature and highlight the contribution of high-level audiovisual features toward improved video segmentation to scenes.
Panagiotis Sidiropoulos, Vasileios Mezaris, Ioannis Kompatsiaris, Hugo Meinedo, Miguel M. F. Bugalho, Isabel Trancoso
IEEE Trans. Circuits Syst. Video Technol.6
2010 Context dependent modelling approaches for hybrid speech recognizers
abstract
Speech recognition based on connectionist approaches is one of the most successful alternatives to widespread Gaussian systems. One of the main claims against hybrid recognizers is the increased complexity for context-dependent phone modeling, which is a key aspect in medium to large size vocabulary tasks. In this paper, we investigate the use of context-dependent triphone models in a connectionist speech recognizer. Thus, most common triphone state clustering procedures for Gaussian models are compared and applied to our hybrid recognizer. The developed systems with clustered context-dependent triphones show above 20% relative word error rate reduction compared to a baseline hybrid system in two selected WSJ evaluation test sets. Additionally, the recent porting efforts of the proposed context modelling approaches to a LVCSR system for English Broadcast News transcription are reported. Index Terms: speech recognition, context modeling, connectionist system
Alberto Abad, Thomas Pellegrini, Isabel Trancoso, João Paulo da Silva Neto
INTERSPEECH3
2010 Speaker recognition experiments using connectionist transformation network features
abstract
The use of adaptation transforms common in speech recognition systems as features for speaker recognition is an appealing alternative approach to conventional short-term cepstral modelling of speaker characteristics. Recently, we have shown that it is possible to use transformation weights derived from adaptation techniques applied to the Multi Layer Perceptrons that form a connectionist speech recognizer. The proposed method – named Transformation Network features with SVM modelling (TN-SVM) – showed promising results on a sub-set of NIST SRE 2008 and allowed further improvements when it was combined with baseline systems. In this paper, we summarize the recently proposed TN-SVM approach and present new results. First, we explore two alternative approaches that may be used in the absence of high quality speech transcriptions. Second, we present results of the proposed approach with Nuisance Attribute Projection for session variability compensation.
Alberto Abad, Isabel Trancoso
INTERSPEECH2
2010 Extending the punctuation module for european portuguese
abstract
This paper describes our recent work on extending the punctuation module of automatic subtitles for Portuguese Broadcast News. The main improvement was achieved by the use of prosodic information. This enabled the extension of the previous module which covered only full stops and commas, to cover question marks as well. The approach uses lexical, acoustic and prosodic information. Our results show that the latter is relevant for all types of punctuation. An analysis of the results also shows what type of interrogative is better dealt with by our method, taking into account the specificities of Portuguese. This may lead to different results for different types of corpora, depending on the types of interrogatives that are more frequent.
Fernando Batista, Helena Moniz, Isabel Trancoso, Hugo Meinedo, Ana Isabel Mata, Nuno J. Mamede
INTERSPEECH3
2010 Exploiting variety-dependent phones in portuguese variety identification applied to broadcast news transcription
abstract
This paper presents a Variety IDentification (VID) approach and its application to broadcast news transcription for Portuguese. The phonotactic VID system, based on Phone Recognition and Language Modelling, focuses on a single tokenizer that combines distinctive knowledge about differences between the target varieties. This knowledge is introduced into a Multi-Layer Perceptron phone recognizer by training mono-phone models for two varieties as contrasting phone-like classes. Significant improvements in terms of identification rate were achieved compared to conventional single and fused phonotactic and acoustic systems. The VID system is used to select data to automatically train variety-specific acoustic models for broadcast news transcription. The impact of the selection is analyzed and variety-specific recognition is shown to improve results by up to 13% compared to a standard variety baseline. Index Terms: accent identification, recognition of accented speech.
Oscar Koller, Alberto Abad, Isabel Trancoso, Céu Viana
INTERSPEECH3
2010 Age and gender classification using fusion of acoustic and prosodic features
abstract
This paper presents a description of the INESC-ID Spoken Language Systems Laboratory (L2F) Age and Gender classification system submitted to the INTERSPEECH 2010 Paralinguistic Challenge. The L2F Age classification system and the Gender classification system are composed respectively by the fusion of four and six individual sub-systems trained with short and long term acoustic and prosodic features, different classification strategies (GMM-UBM, MLP and SVM) and using four different speech corpora. The best results obtained by the calibration and linear logistic regression fusion back-end show an absolute improvement of 4.1% on the unweighted accuracy value for the Age and 5.8% for the Gender when compared to the competition baseline systems in the development set.
Hugo Meinedo, Isabel Trancoso
INTERSPEECH2
2010 Improving ASR error detection with non-decoder based features
abstract
This study reports error detection experiments in large vocabulary automatic speech recognition (ASR) systems, by using statistical classifiers. We explored new features gathered from other knowledge sources than the decoder itself: a binary feature that compares outputs from two different ASR systems (word by word), a feature based on the number of hits of the hypothesized bigrams, obtained by queries entered into a very popular Web search engine, and finally a feature related to automatically infered topics at sentence and word levels. Experiments were conducted on a European Portuguese broadcast news corpus. The combination of baseline decoder-based features and two of these additional features led to significant improvements, from 13.87% to 12.16% classification error rate (CER) with a maximum entropy model, and from 14.01% to 12.39% CER with linear-chain conditional random fields, comparing to a baseline using only decoder-based features.
Thomas Pellegrini, Isabel Trancoso
INTERSPEECH2
2010 Multimedia learning materials
abstract
This paper describes the integration of multimedia documents in the Portuguese version of REAP, a tutoring system for vocabulary learning. The documents result from the pipeline processing of Broadcast News videos that automatically segments the audio files, transcribes them, adds punctuation and capitalization, and breaks them into stories classified by topics. The integration of these materials in REAP was done in a way that tries to decrease the impact of potential errors of the automatic chain in the learning process.
José Lopes 0001, Isabel Trancoso, Rui Correia, Thomas Pellegrini, Hugo Meinedo, Nuno J. Mamede, Maxine Eskénazi
SLT2
2009 Comparing automatic rich transcription for Portuguese, Spanish and English Broadcast News
abstract
This paper describes and evaluates a language independent approach for automatically enriching the speech recognition output with punctuation marks and capitalization information. The two tasks are treated as two classification problems, using a maximum entropy modeling approach, which achieves results within state-of-the-art. The language independence of the approach is attested with experiments conducted on Portuguese, Spanish and English broadcast news corpora. This paper provides the first comparative study between the three languages, concerning these tasks.
Fernando Batista, Isabel Trancoso, Nuno J. Mamede
ASRU2
2009 Non-speech audio event detection
abstract
Audio event detection is one of the tasks of the European project VIDIVIDEO. This paper focuses on the detection of non-speech events, and as such only searches for events in audio segments that have been previously classified as non-speech. Preliminary experiments with a small corpus of sound effects have shown the potential of this type of corpus for training purposes. This paper describes our experiments with SVM and HMM-based classifiers, using a 290-hour corpus of sound effects. Although we have only built detectors for 15 semantic concepts so far, the method seems easily portable to other concepts. The paper reports experiments with multiple features, different kernels and several analysis windows. Preliminary experiments on documentaries and films yielded promising results, despite the difficulties posed by the mixtures of audio events that characterize real sounds.
José Portelo, Miguel M. F. Bugalho, Isabel Trancoso, João Paulo da Silva Neto, Alberto Abad, António Joaquim Serralheiro
ICASSP3
2009 Audio contributions to semantic video search
abstract
This paper summarizes the contributions to semantic video search that can be derived from the audio signal. Because of space restrictions, the emphasis will be on non-linguistic cues. The paper thus covers what is generally known as audio segmentation, as well as audio event detection. Using machine learning approaches, we have built detectors for over 50 semantic audio concepts.
Isabel Trancoso, Thomas Pellegrini, José Portelo, Hugo Meinedo, Miguel M. F. Bugalho, Alberto Abad, João Paulo da Silva Neto
ICME1
2009 Porting an european portuguese broadcast news recognition system to brazilian portuguese
abstract
This paper reports on recent work in the context of the activities of the PoSTPort project aimed at porting a Broadcast News recognition system originally developed for European Portuguese to other varieties. Concretely, in this paper we have focused on porting to Brazilian Portuguese. The impact of some of the main sources of variability has been assessed, besides proposing solutions at the lexical, acoustic and syntactic levels. The ported Brazilian Portuguese Broadcast News system allowed a drastic performance improvement from 56.6% WER (obtained with the European Portuguese system) to 25.5%.
Alberto Abad, Isabel Trancoso, Nelson Neto 0001, Céu Viana
INTERSPEECH2
2009 Cross-variety rhythm typology in portuguese
abstract
This paper aims at proposing a measure of speech rhythm based on the inference of the coupling strength between the syllable oscillator and the stress group oscillator of an underlying coupled oscillators model. This coupling is inferred from the linear regression between the stress group duration and the number of syllables within the group, as well as from the multiple linear regression between the same parameters and an estimate of phrase stress prominence. This technique is applied to compare the rhythmic differences between European and Brazilian Portuguese in two speaking styles and three speakers per variety. Compared with a syllable-sized normalised PVI, the findings suggest that the coupling strength captures better the perceptual effects of the speakers’ renditions. Furthermore, it shows that stress group duration is much better predicted by adding phrase stress prominence to the regression. Index Terms: rhythmic variability, rhythm typology, coupledoscillators
Plínio Almeida Barbosa, Céu Viana, Isabel Trancoso
INTERSPEECH3
2009 Detecting audio events for semantic video search
abstract
This paper describes our work on audio event detection, one of our tasks in the European project VIDIVIDEO. Preliminary experiments with a small corpus of sound effects have shown the potential of this type of corpus for training purposes. This paper describes our experiments with SVM classifiers, and different features, using a 290-hour corpus of sound effects, which allowed us to build detectors for almost 50 semantic concepts. Although the performance of these detectors on the development set is quite good (achieving an average F-measure of 0.87), preliminary experiments on documentaries and films showed that the task is much harder in real-life videos, which so often include overlapping audio events. Index Terms: event detection, audio segmentation 1.
Miguel M. F. Bugalho, José Portelo, Isabel Trancoso, Thomas Pellegrini, Alberto Abad
INTERSPEECH3
2009 Classification of disfluent phenomena as fluent communicative devices in specific prosodic contexts
abstract
This work explores prosodic cues of disfluent phenomena. In our previous work, we conducted a perceptual experiment regarding (dis)fluency ratings. Results suggested that some disfluencies may be considered felicitous by listeners, namely filled pauses and prolongations. In an attempt to discriminate which linguistic features are more salient in the classification of disfluencies as either fluent or disfluent phenomena, we used CART techniques on a corpus of 3.5 hours of spontaneous and prepared non-scripted speech. CART results pointed out 2 splits: break indices and contour shape. The first split indicates that events uttered at breaks 3 and 4 are considered felicitous. The second shows that these events must have flat or ascending contours to be considered as such; otherwise they are strongly penalized. Our preliminary results suggest that there are regular trends in the production of these events, namely, prosodic phrasing and contour shape. Index Terms: prosody, disfluency, fluency rating
Helena Moniz, Isabel Trancoso, Ana Isabel Mata
INTERSPEECH2
2009 Multi-modal scene segmentation using scene transition graphs
abstract
In this work the problem of automatic decomposition of video into elementary semantic units, known in the literature as scenes, is addressed. Two multi-modal automatic scene segmentation techniques are proposed, both building upon the Scene Transition Graph (STG). In the first of the proposed approaches, speaker diarization results are used for introducing a post-processing step to the STG construction algorithm, with the objective of discarding scene boundaries erroneously identified according to visual-only dissimilarity. In the second approach, speaker diarization and additional audio analysis results are employed and a separate audio-based STG is constructed, in parallel to the original STG based on visual information. The two STGs are subsequently combined. Preliminary results from the application of the proposed techniques to broadcast videos reveal their improved performance over previous approaches.
Panagiotis Sidiropoulos, Vasileios Mezaris, Ioannis Kompatsiaris, Hugo Meinedo, Isabel Trancoso
ACM Multimedia5
2008 Portuguese variety identification on broadcast news
abstract
This paper describes an accent identification system for Portuguese, that explores different type of properties: acoustic, phono tactic and prosodic. The system is designed to be used as a pre-processing module for the Portuguese automatic speech recognition system developed at INESC-ID. In terms of variety identification, the overall rate of correct identification is 69.0% if all 7 varieties are considered, and the best results are obtained for Brazilian Portuguese, also the variety that proved easiest to identify in perceptual experiments. When distinguishing between European, Brazilian and African Portuguese, the identification rate goes up to 94.7%. The fact that the prosodic system alone can achieve an identification rate of 77% is also worth investigating.
Jean-Luc Rouas, Isabel Trancoso, Céu Viana, Mónica Abreu
ICASSP2
2008 Topic segmentation and indexation in a media watch system
abstract
The goal of this paper is the description of our current work in terms of topic segmentation and indexation and the comparison of their performance with the story boundaries and topics manually chosen by a professional media watch company. The segmentation module explores the typical structure of a broadcast news show, namely by cues provided by the audio pre-processing module, but an improved performance could be achieved by taking its contents into account. The topic indexation module was retrained for the media watch topics. The comparison showed how different criteria, together with manual labeling inconsistencies, may affect the performance.
Rui Amaral, Isabel Trancoso
INTERSPEECH2
2008 The impact of language dynamics on the capitalization of broadcast news
abstract
This paper investigates the impact of language dynamics on the capitalization of transcriptions of broadcast news. Most of the capitalization information is provided by a large newspaper corpus. Three different speech corpora subsets, from different time periods, are used for evaluation, assessing the importance of available training data in nearby time periods. Results are provided both for manual and automatic transcriptions, showing also the impact of the recognition errors in the capitalization task. Our approach is based on maximum entropy models, uses unlimited vocabulary, and is suitable for language adaptation. The language model for a given language period is produced by retraining a previous language model with data from that time period. The language model produced with this approach can be sorted and then pruned, in order to reduce computational resources, without much impact in the final results.
Fernando Batista, Nuno J. Mamede, Isabel Trancoso
INTERSPEECH3
2008 How can you use disfluencies and still sound as a good speaker?
Helena Moniz, Ana Isabel Mata, Isabel Trancoso, Céu Viana
INTERSPEECH3
2008 Training audio events detectors with a sound effects corpus
abstract
This paper describes the work done in the framework of the VIDIVIDEO European project in terms of audio event detection. Our first experiments concerned the detection of nonvoice sounds, such as birds, machines, traffic, water and steps. Given the unavailability of a corpus labelled in terms of audio events, we used a relatively small sound effect corpus for training. Our initial experiments with one-against-all SVM classifiers for these 5 classes showed us the feasibility of using this type of data for training, thus avoiding the extremely morose task of manual labelling of a very high number of audio events. Preliminary integration experiments are quite promising. Index Terms: audio segmentation, event detection 1.
Isabel Trancoso, José Portelo, Miguel M. F. Bugalho, João Paulo da Silva Neto, António Joaquim Serralheiro
INTERSPEECH1
2008 The LECTRA Corpus - Classroom Lecture Transcriptions in European Portuguese
Isabel Trancoso, Rui Martins, Helena Moniz, Ana Isabel Mata, Céu Viana
LREC1
2008 Impact of dynamic model adaptation beyond speech recognition
abstract
The application of speech recognition to live subtitling of Broadcast News has motivated the adaptation of the lexical and language models of the recognizer on a daily basis with text material retrieved from online newspapers. This paper studies the impact of this adaptation on two of the blocks following the speech recognition module: capitalization and topic indexation. We describe and evaluate different adaptation approaches that try to explore the language dynamics.
Fernando Batista, Rui Amaral, Isabel Trancoso, Nuno J. Mamede
SLT3
2008 Recovering capitalization and punctuation marks for automatic speech recognition: Case study for Portuguese broadcast news
Fernando Batista, Diamantino Caseiro, Nuno J. Mamede, Isabel Trancoso
Speech Commun.4
2008 Language and variety verification on broadcast news for Portuguese
Jean-Luc Rouas, Isabel Trancoso, Céu Viana, Mónica Abreu
Speech Commun.2
2008 Special Issue on Iberian Languages
Isabel Trancoso, Néstor Becerra Yoma, Plínio Almeida Barbosa, Rubén San-Segundo-Hernández, Kuldip K. Paliwal
Speech Commun.1
2007 Recovering punctuation marks for automatic speech recognition
abstract
This paper shows results of recovering punctuation over speech transcriptions for a Portuguese broadcast news corpus. The approach is based on maximum entropy models and uses word, part-of-speech, time and speaker information. The contribution of each type of feature is analyzed individually. Separate results for each focus condition are given, making it possible to analyze the differences of performance between planned and spontaneous speech. Index Terms: rich transcription, punctuation recovery, sentence boundary detection, maximum entropy.
Fernando Batista, Diamantino Caseiro, Nuno J. Mamede, Isabel Trancoso
INTERSPEECH4
2006 Spoken language technologies applied to digital talking books
abstract
Digital Talking Books (DTBs) offer to visually impaired users an evolution of analogue talking books that mimics the interaction possibilities of print books. This paper describes a new DTB player which tries to improve the usability and accessibility of current players, through the combination of the possibilities offered by multimodal interaction and interface adaptability, and the integration of several language processing components. Besides the potential for a greater enjoyment of the reader in general, these modifications also pave the way to the use of DTBs in different domains, from e-inclusion to e-learning applications. Index Terms: digital talking books, Portuguese. 1.
Isabel Trancoso, Carlos Duarte, António Joaquim Serralheiro, Diamantino Caseiro, Luís Carriço, Céu Viana
INTERSPEECH1
2006 Recognition of classroom lectures in european portuguese
abstract
Classroom lectures may be very challenging for automatic speech recognizers, because the vocabulary may be very specific and the speaking style very spontaneous. Our first experiments using a recognizer trained for Broadcast News resulted in word error rates near 60%, clearly confirming the need for adaptation to the specific topic of the lectures, on one hand, and for better strategies for handling spontaneous speech. This paper describes our efforts in these two directions: the different domain adaptation steps that lowered the error rate to 45%, with very little transcribed adaptation material, and the exploratory study of spontaneous speech phenomena in European Portuguese, namely concerning filled pauses. Index Terms: spontaneous speech recognition, Portuguese. 1.
Isabel Trancoso, Ricardo Nunes, Luís Neves, Céu Viana, Helena Moniz, Diamantino Caseiro, Ana Isabel Mata
INTERSPEECH1
2006 A specialized on-the-fly algorithm for lexicon and language model composition
abstract
This paper presents an algorithm for the composition of weighted finite-state transducers which is specially tailored to speech recognition applications: it composes the lexicon with the language model while simultaneously optimizing the resulting transducer. Furthermore, it performs these computations "on-the-fly" to allow easier management of the tradeoff between offline and online computation and memory. The algorithm is exact for local knowledge integration and optimization operations such as composition and determinization. Minimization and pushing operations are approximated. Our results have confirmed the efficiency of these approximations
Diamantino Caseiro, Isabel Trancoso
IEEE Trans. Speech Audio Process.2
2005 Finite-state transducer inference for a speech-input Portuguese-to-English machine translation system
abstract
Statistical techniques and grammatical inference have been used for dealing with automatic speech recognition with success, and can also be used for speech-to-speech machine translation. In this paper, new advances on a method for finite-state transducer inference are presented. This method has been tested experimentally in a speech-input translation task using a recognizer that allows a flexible use of models by means of efficient algorithms for on-thefly transducer composition. These are the first reported results of a speech-to-speech translation task involving European Portuguese input that we know of.
David Picó, Francisco Casacuberta, Diamantino Caseiro, Isabel Trancoso
INTERSPEECH5
2005 Aligning and recognizing spoken books in different varieties of Portuguese
abstract
This paper tries to present digital spoken books as a useful diagnostic tool for detecting alignment and recognition problems and for studying the porting of these technologies to different varieties of the same language- Portuguese, in our case. We summarize the main differences between European and Brazilian Portuguese (EP/BP) and describe how they affect the GtoP system. Despite the small size of our parallel spoken book corpus in the two varieties, our preliminary experiments confirmed our expectations in terms of the effectiveness of an EP-trained aligner used on BP spoken books. They also confirmed the inadequacy of an EP Broadcast News recognizer tested over literary contents, and the expected degradation in recognition scores caused by using that recognizer on a BP spoken book. Pronunciation adaptation was tested by adding variants derived by the BP GtoP system to our EP lexicon, resulting in a very small improvement in terms of recognition scores.
Isabel Trancoso, António Joaquim Serralheiro, Céu Viana, Diamantino Caseiro
INTERSPEECH1
2005 Editorial
Isabel Trancoso
IEEE Trans. Speech Audio Process.1
2004 Improving the topic indexation and segmentation modules of a media watch system
abstract
This paper describes on-going work related to the topic segmentation and indexation module of an alert system for selective dissemination of multimedia information. This system was submitted in the past year to a field trial which exposed a number of issues that should be dealt with in order to improve its performance. Some of our efforts involved the use of multiple topics, confidence measures and named entity extraction. This paper discusses these approaches and the corresponding results which, unfortunately, are still affected by the limited amount of topic-annotated training data.
Rui Amaral, Isabel Trancoso
INTERSPEECH2
2004 Poetry assistant
Isabel Trancoso, Paul Araújo, Céu Viana, Nuno J. Mamede
INTERSPEECH1
2004 From the Editor-in-Chief
Isabel Trancoso
IEEE Trans. Speech Audio Process.1
2003 A tail-sharing WFST composition algorithm for large vocabulary speech recognition
abstract
This paper presents an algorithm for approximating minimization in the context of the weighted finite-state transducers approach to large vocabulary speech recognition. The algorithm is designed for the integration of the lexicon with the language model and performs composition, determinization and pushing in one step. Furthermore, it uses tail-sharing in order to approximate minimization. Our results show that it is a good approximation to explicit minimization, with the added advantage that it can be used "on-the-fly" in a dynamic decoder.
Diamantino Caseiro, Isabel Trancoso
ICASSP (1)2
2003 Towards a repository of digital talking books
abstract
Considerable effort has been devoted at# # # to increase and broaden our speech and text data resources. Digital Talking Books (DTB), comprising both speech and text data are, as such, an invaluable asset as multimedia resources. Furthermore, those DTB have been under a speech-to-text alignment procedure, either word or phone-based, to increase their potential in research activities. This paper thus describes the motivation and the method that we used to accomplish this goal for aligning DTBs. This alignment allows specific access interfaces for persons with special needs, and also tools for easily detecting and indexing units (words, sentences, topics) in the spoken books. The alignment tool was implemented in a Weighted Finite State Transducer framework, which provides an efficient way to combine different types of knowledge sources, such as alternative pronunciation rules. With this tool, a 2-hour long spoken book was aligned in a single step in much less than real time. Last but not least, new browsing interfaces, allowing improved access and data retrieval to and from the DTBs, are described in this paper.
António Joaquim Serralheiro, Isabel Trancoso, Diamantino Caseiro, Teresa Chambel, Luís Carriço, Nuno Guimarães
INTERSPEECH2
2003 Evaluation of an alert system for selective dissemination of broadcast news
abstract
This paper describes the evaluation of the system for selective dissemination of Broadcast News that we developed in the context of the European project ALERT.Each component of the main processing block of our system was evaluated separately, using the ALERT corpus.Likewise, the user interface was also evaluated separately.Besides this modular evaluation which will be briefly mentioned here, as a reference, the system can also be evaluated as a whole, in a field trial from the point of view of a potential user.This is the main topic of this paper.The analysis of the main sources of problems hinted at a large number of issues that must be dealt with in order to improve the performance.In spite of these pending problems, we believe that having a fully operational system is a must for being able to address user needs in the future in this type of service. 10.
Isabel Trancoso, João Paulo da Silva Neto, Hugo Meinedo, Rui Amaral
INTERSPEECH1
2003 From the editor-in-chief
Isabel Trancoso
IEEE Trans. Speech Audio Process.1
2002 Using dynamic WFST composition for recognizing broadcast news
abstract
Our first application of weighted finite state transducers to the recognition of broadcast news provided us with an interesting framework to study several problems related to the optimization of the search space. The paper starts by describing how the use of our lexicon and language model "on-the-fly" composition algorithm is crucial in extending the transducer approach to large systems. We present an efficient representation for WFSTs, that allowed us to reduce runtime memory requirements, and discuss several types of language model optimizations, including a context-sharing algorithm. Experimental results obtained with the broadcast news corpus collected for European Portuguese illustrate the impact of the various possible optimizations of the components on the performance of the system.
Diamantino Caseiro, Isabel Trancoso
INTERSPEECH2
2002 Morphosyntactic Disambiguation for TTS Systems
Ricardo Ribeiro 0001, Luís C. Oliveira, Isabel Trancoso
LREC3
2002 Editorial
Harry Printz, Isabel Trancoso
IEEE Trans. Speech Audio Process.2
2001 The development of a portuguese version of a media watch system
abstract
This paper summarizes the work that has been done concerning the Portuguese language in the scope of the ALERT project during its first year.The media watch system that is the goal of this project comprises many different modules, some of them common among the three languages of the project.This paper concentrates on the definition and collection of the necessary linguistic resources for Portuguese, and the development of the speech recognition, topic and jingle detection modules.The first version of the ALERT demo for European Portuguese is also described.
Rui Amaral, Thibault Langlois, Hugo Meinedo, João Paulo da Silva Neto, Nuno Souto, Isabel Trancoso
INTERSPEECH6
2001 On integrating the lexicon with the language model
abstract
The goal of this work was to develop an algorithm for the integration of the lexicon with the language model which would be computationally efficient in terms of memory requirements, even in the case of large trigram models. Two specialized versions of the algorithm for transducer composition were implemented. The first one is basically a composition algorithm that uses the precomputed set of the output labels that can be reached from a particular epsilon edge of the lexicon
Diamantino Caseiro, Isabel Trancoso
INTERSPEECH2
2000 Phonetic vocoder assessment
abstract
The efficiency of phonetic vocoders stems from the fact that the only transmitted information is the index of the recognised units and the corresponding prosodic parameters. Hence, speaker recognisability is one of the main issues in this class of coders. Our approach to minimise this drawback was to include some speaker adaptation capability. The purpose of this paper is two-folded: on one hand, to describe the recognisability and intelligibility tests that were performed with our phonetic vocoder with and without speaker adaptation; on the other hand, to present our recent developments of this coder, using the SpeechDat corpus for Portuguese, that includes telephone calls from 5000 speakers. This allowed us to generate improved HMM models, codebooks, and quantization tables, and to investigate the performance of the coder in non-clean environments and with a much wider speaker population.
Carlos M. Ribeiro, Isabel Trancoso, Diamantino Caseiro
INTERSPEECH2
2000 Book review
Isabel Trancoso
Speech Commun.1
1999 On improving the decision algorithm for articulatory codebook search
abstract
This paper describes our progress on articulatory voice mimic. The objective is to achieve an articulatory voice mimic system as a basis for low bit-rate speech coding using articulatory codebooks. The articulatory codebook uses a suitable vocal tract model for generating shapes for all possible speech sounds. When building a codebook, unrealistic vocal tract shapes may be generated. In this paper, we describe the used filter based on physiological assumptions to avoid populating the codebook with such shapes. We also propose a new method for the codebook search. This method weights the prediction error for each area function parameter as a function of its contribution to match the input acoustic signal parameters. An interaction between the variations of the area function parameter is created to constraint the articulatory trajectory, which improves the search and increases the output speech quality of the voice mimic system.
Carlos Silva 0001, Samir Chennoukh, Isabel Trancoso
EUROSPEECH3
1999 On deriving rules for nativised pronunciation in navigation queries
abstract
Navigation queries are typical examples of contexts in which a recognizer may have to deal with non-native names. In order to build a pronunciation lexicon with these names, special GtoP rules may be derived. The paper addresses this problem in the context of navigation queries in French including German names and viceversa. The special GtoP rules were mostly based on statistics derived from cross-lingual spoken corpora.
Isabel Trancoso, Céu Viana, Isabel Mascarenhas
EUROSPEECH1
1998 Spoken language identification using the speechdat corpus
abstract
Current language identification systems vary significantly in their complexity. The systems that use higher level linguistic information have the best performance. Nevertheless, that information is hard to collect for each new language. The system presented in this paper is easily extendable to new languages because it uses very little linguistic information. In fact, the presented system needs only one language specific phone recogniser (in our case the Portuguese one), and is trained with speech from each of the other languages. With the SpeechDat-M corpus, with 6 European languages (English, French, German, Italian, Portuguese and Spanish) our system achieved an identification rate of 83.4% on 5-second utterances, this result shows an improvement of 5% over our previous version, mainly through the use of a neural network classifier. Both the baseline and the full system were implemented in realtime. 1. INTRODUCTION When designing a language identification system, we face the prob...
Diamantino Caseiro, Isabel Trancoso
ICSLP2
1998 Improving speaker recognisability in phonetic vocoders
abstract
Phonetic vocoding is one of the methods for coding speech below 1000 bit/s. The transmitter stage includes a phone recogniser whose index is transmitted together with prosodic information such as duration, energy and pitch variation. This type of coder does not transmit spectral speaker characteristics and speaker recognisability thus becomes a major problem. In our previous work, we adapted a speaker modification strategy to minimise this problem, modifying a codebook to match the spectral characteristics of the input speaker. This is done at the cost of transmitting the LSP averages computed for vowel and glide phones. This paper presents new codebook generation strategies, with gender dependence and interpolation frames, that lead to better speaker recognisability and speech quality. Relatively to our previous work, some effort was also devoted to deriving more efficient quantization methods for the speakerspecific information, that considerably reduced the average bit rate, without quality degradation. For the CD-ROM version, a set of audio files is also included.
Carlos M. Ribeiro, Isabel Trancoso
ICSLP2
1997 Phonetic vocoding with speaker adaptation
abstract
This paper describes a phonetic vocoding scheme which relies on speaker adaptation to capture important speaker characteristics. These are typically lost in phonetic vocoders which transmit only information about the phones which are recognized, together with some prosodic information. In our scheme, however, additional speaker characteristics are transmitted in vowel regions (average values of LSP coefficients for each phone). This additional information yielded potentially good speaker recognizability results, in informal listening tests, while still achieving a rather low average bit rate, suitable for many transmission and storage applications. This work extends our previous phonetic vocoding scheme described in [5]. The vocoder is now fully quantized and the number of transmitted parameters had been significantly reduced. 1. INTRODUCTION Segment or Phonetic vocoders are one of the most frequently proposed methods to code speech at rates below 1000 bit/s [4, 7]. This type of code...
Carlos M. Ribeiro, Isabel Trancoso
EUROSPEECH2
1997 Recognition of non-native accents
abstract
This paper deals with the problem of non-native accents in speech recognition. Reference tests were performed using whole-word and sub-word models trained either with a native accent or a pool of native and non-native accents. The results seem to indicate that the use of phonetic transcriptions for each specific accent may improve recognition scores with sub-word models. A data-driven process is used to derive transcription lattices. The recognition scores thus obtained were encouraging. 1. INTRODUCTION Our main concern in this study is to reduce the negative effects caused by non-native accents in automatic speech recognition. These effects account for a significant drop of word recognition scores when using a recogniser trained with material spoken with a native accent. Collecting large enough corpora for each non-native accent is generally not feasible. However, this problem turns out to be more complex, since even a recogniser trained with speech material from a specific non-nati...
Isabel Trancoso, António Joaquim Serralheiro
EUROSPEECH2
1997 On the pronunciation mode of acronyms in several European languages
Isabel Trancoso, Céu Viana
EUROSPEECH1
1996 Application of speaker modification techniques to phonetic vocoding
Carlos M. Ribeiro, Isabel Trancoso
ICSLP2
1996 Accent identification
Isabel Trancoso, António Joaquim Serralheiro
ICSLP2
1995 EUROM - a spoken language resource for the EU - the SAM projects
Dominic S. F. Chan, Adrian Fourcin, Dafydd Gibbon, Björn Granström, Mark A. Huckvale, George K. Kokkinakis, Knut Kvale, Lori Lamel, Børge Lindberg, Asunción Moreno, Jiannis Mouropoulos, Francesco Senia, Isabel Trancoso, Corin 't Veld, Jerome Zeiliger
EUROSPEECH13
1994 E-mail to voice-mail conversion using a portuguese text-to-speech system
Pedro M. Carvalho, Isabel Trancoso, Luís C. Oliveira
ICSLP3
1994 Rule-based vs neural network-based approaches to letter-to-phone conversion for portuguese common and proper names
Isabel Trancoso, Céu Viana, Fernando M. Silva, Goncalo C. Marques, Luís C. Oliveira
ICSLP1
1993 A software tool for speech collection, recognition and reproduction
Carlos M. Ribeiro, Isabel Trancoso, António Joaquim Serralheiro
EUROSPEECH2
1993 The relationship between spelled and spoken portuguese: implications for speech synthesis and recognition
Céu Viana, Isabel Trancoso, Carlos M. Ribeiro, Amalia Andrade, Ernesto d'Andrade
EUROSPEECH2
1992 A rule-based text-to-speech system for Portuguese
abstract
The latest progress in the development of a text-to-speech system for Portuguese is described. The system comprises four major modules: text normalization, linguistic and phonetic processing, generation of the synthesizer parameters and synthesis. The present rule-based version, based on the Klatt80 formant synthesizer, has achieved promising results, namely in what concerns the performance of stress assignment, phonetic transcription, and prosodic parsing. Each of the major modules is described, and some language-dependent issues are discussed.>
Luís C. Oliveira, Céu Viana, Isabel Trancoso
ICASSP3
1992 Word rejection using multiple sink models
Isabel Trancoso
ICSLP2
1991 Hybrid sinusoidal modeling of speech without voicing decision
Arnaldo J. Abrantes, Jorge S. Marques, Isabel Trancoso
EUROSPEECH3
1991 A recognition / synthesis system applied to database access through the telephone network
Filipe N. Carlos, Jose P. Carmona, Pedro M. Chagas, Luís C. Oliveira, António Joaquim Serralheiro, Isabel Trancoso
EUROSPEECH6
1991 Harmonic coding of speech: an experimental study
Jorge S. Marques, Isabel Trancoso, Arnaldo J. Abrantes
EUROSPEECH2
1991 DIXI - portuguese text-to-speech system
abstract
This paper describes the software architecture of the Portuguese text-to-speech system DIXI 1 . The system has three major modules. The first one contains the text normalizer and searches eachword in the lexicon. The second one is a multi-level rule based module for lexical stress assignment, orthographic to phonetic transcription, metrically based prosodic patterning and for generating the evolution of the synthesizer parameters. The final module is the Klatt 80 formant synthesizer. The paper describes each of these main modules, emphasizing the particularities of text-to-speech synthesis in the Portuguese language.
Luís C. Oliveira, Céu Viana, Isabel Trancoso
EUROSPEECH3
1991 A 4.8 kbps celp coder with post-processing
Carlos M. Ribeiro, Isabel Trancoso
EUROSPEECH2
1991 Spectral subtraction for front-end noise reduction in a speech recognizer
Isabel Trancoso
EUROSPEECH2
1990 Improved pitch prediction with fractional delays in CELP coding
abstract
A scheme is discussed for long-term prediction in CELP (code-excited linear predictive) coding using fractional delay prediction. This technique permits a more accurate representation of voiced speech and achieves an improvement of synthetic quality for female speakers. The higher complexity of this type of predictor relative to the classical one is its major disadvantage. Suboptimal schemes in which the search for the functional pitch delay is restricted to a neighborhood of an integer pitch estimate can be envisaged to decrease the computational load.>
Jorge S. Marques, Isabel Trancoso, José M. Tribolet, Luís B. Almeida
ICASSP2
1990 CELP: a candidate for GSM half-rate coding
abstract
A systematic study of code excited linear predictive (CELP) coders is presented. The design of an optimized version complying with the specifications for a GSM half-rate coder is described. This study is divided into three parts: election of an unquantized configuration, assuming unquantized scale factors and filter coefficients, fine tuning of the parameter quantization, and adaptation to the GSM specifications. A strong emphasis is placed on efficient codebook search procedures, and two approaches are briefly described: the truncated autocorrelation approach and the unity-magnitude approach. The final coder version achieves a very good combination of speech quality and implementation complexity.>
Isabel Trancoso, Carlos M. Ribeiro, Luís B. Almeida, Luís C. Oliveira, Jorge S. Marques
ICASSP1
1990 CELP and sinusoidal coders: Two solutions for speech coding at 4.8-9.6 kbps
Isabel Trancoso, Jorge S. Marques, Carlos M. Ribeiro
Speech Commun.1
1989 Pitch prediction with fractional delays in CELP coding
Jorge S. Marques, José M. Tribolet, Isabel Trancoso, Luís B. Almeida
EUROSPEECH3
1989 Adaptive and stochastic search procedures in CELP based coders
Isabel Trancoso, Carlos M. Ribeiro
EUROSPEECH1
1988 Quantization issues in harmonic coders [speech coding]
abstract
The introduction of vector quantization in the harmonic coder framework can affect several sets of parameters, namely the amplitudes and the phases of the harmonic coefficients. A systematic experimental study is reported in which this type of quantization is compared with traditional scalar techniques for quantizing the harmonic parameters. The experiments showed that it is possible to quantize voice speech below 8 kb/s and still obtain high-quality synthetic signals.>
Isabel Trancoso, Joaquim S. Rodrigues, Luís B. Almeida, Jorge S. Marques, António Joaquim Serralheiro, Diana Santos, José M. Tribolet
ICASSP1
1988 Harmonic coding - state of the art and future trends
Isabel Trancoso, Luís B. Almeida, Joaquim S. Rodrigues, Jorge S. Marques, José M. Tribolet
Speech Commun.1
1986 Efficient procedures for finding the optimum innovation in stochastic coders
abstract
Stochastic linear predictive coders have the potential for producing high quality synthetic speech at bit rates as low as 4.8 kbits/sec. However, these coders are computationally very complex requiring more than 500 million multiply-add operations per second - a capability much beyond the reach of present digital hardware. Most of the complexity in stochastic coders comes from the search procedure used to determine the optimum innovation sequence from a random code book of white Gaussian sequences. This paper will describe several procedures to simplify the search in stochastic coders. The computational efficiency is achieved by moving the search from time domain to another domain where the convolution (filtering) operation in the search procedure is reduced to a single multiplication operation. The new procedures have reduced the computational load by a factor of approximately 20 without producing any additional degradation in synthetic speech.
Isabel Trancoso, Bishnu S. Atal
ICASSP1
1986 A study on the relationships between stochastic and harmonic coding
abstract
Two of the coding methods which have appeared recently, attempting to solve the problem of high quality speech coding at bit rates below 9.6 kbits/sec are harmonic and stochastic coding. The two represent distinct approaches to this problem and yield synthetic speech of a very different kind. In this paper, we discuss the advantages and disadvantages of both methods for processing voiced and unvoiced speech segments, and examine the potential benefits of combining the two methods in coding applications.
Isabel Trancoso, Luís B. Almeida, José M. Tribolet
ICASSP1
1985 Pole-zero multipulse speech representation using harmonic modelling in the frequency domain
abstract
When trying to increase the efficiency of the multipulse speech representation through a reduction in the number of pulses per frame, the use of a pole-zero filter instead of the classical LPC filter appears to be an appropriate choice, for voiced segments. In this paper, the possibility and advantages of using this kind of filter are investigated. The filter design criterion is the minimization of the mean squared error (m.s.e.) between its output and the original voiced signal, when the excitation is a suitable periodic impulse train. This is, however, approximately equivalent to finding the filter that minimizes the m.s.e. between its transfer function and the signal harmonics, at the harmonic frequencies, taking into account both magnitude and phase. Some preliminary results are shown, and the conclusion is drawn that pole-zero filters actually are good candidates for a more efficient multipulse representation of voiced speech. The design of an algorithm for filter computation, and the incorporation of these ideas into the general multipulse framework are also discussed.
Isabel Trancoso, Luís B. Almeida, José M. Tribolet
ICASSP1
1984 A study on short-time phase and multipulse LPC
abstract
Multipulse LPC, as is often designated the model proposed by Atal and Remae, has been directed towards 9.6 kb/s speech coding. However, at such bit rate, the speech quality is not yet generally acceptable. This paper has a double purpose - one is to investigate the role of short-time phase in multipulse LPC. The other is to look for different modelling structures to be used with this method. This study was prompted by direct observation of the structure of the excitation, where patterns of pulses may be found which are associated with phase-correcting mechanisms of the LPC impulse response. Consequently, a new multipulse technique was developed, based on a special ARMA model formed by cascading an all-pole with an all-pass network. This new model will be referred to as the MAPAP (Multipulse All-Pole All-Pass) method. Another one was tried in which the synthetic speech is formed by combination of several MAPAP signals. We therefore denoted it "multichannel multipulse" method. The potential advantages of both single and multichannel models seem rather promising.
Isabel Trancoso, Ramón García-Gómez, José M. Tribolet
ICASSP1