Alberto Abad

dblp:71/1198 · DBLP profile ↗
← Back
83ranked-venue papers
17as first author
26since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 68 · 14 first-author · 22 since 2021Artificial intelligence and machine learning · 58 · 13 first-author · 19 since 2021Theory of computation · 3 · 2 first-author
YearPublicationVenuePosition
2026 FalAR: A Large-scale Speaker-Annotated European Portuguese Speech Corpus of Parliamentary Sessions
Francisco Teixeira, Carlos Carvalho 0003, Mariana Julião, Catarina Botelho, Rubén Solera-Ureña, Sérgio Paulo, Thomas Rolland, Ben Peters, Isabel Trancoso, Alberto Abad
LREC10
2026 Exploring features for membership inference in ASR model auditing
Francisco Teixeira, Karla Pizzi, Raphaël Olivier, Alberto Abad, Bhiksha Raj, Isabel Trancoso
Comput. Speech Lang.4
2025 CAMÕES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese
abstract
Existing resources for Automatic Speech Recognition in Portuguese are mostly focused on Brazilian Portuguese, leaving European Portuguese (EP) and other varieties underexplored. To bridge this gap, we introduce CAMÕES, the first open framework for EP and other Portuguese varieties. It consists of (1) a comprehensive evaluation benchmark, including 46 h of EP test data spanning multiple domains; and (2) a collection of state-of-the-art models. For the latter, we consider multiple foundation models, evaluating their zero-shot and fine-tuned performances, as well as E-Branchformer models trained from scratch. A curated set of $\mathbf{4 2 5 h}$ of EP was used for both fine-tuning and training. Our results show comparable performance for EP between fine-tuned foundation models and the E-Branchformer. Furthermore, the best-performing models achieve relative improvements above 35% WER, compared to the strongest zero-shot foundation model, establishing a new state-of-the-art for EP and other varieties.
Carlos Carvalho 0003, Francisco Teixeira, Catarina Botelho, Anna Pompili, Rubén Solera-Ureña, Sérgio Paulo, Mariana Julião, Thomas Rolland, John Mendonça, Diogo A. P. Nunes, Isabel Trancoso, Alberto Abad
ASRU12
2025 AC-Mix: Self-Supervised Adaptation for Low-Resource Automatic Speech Recognition using Agnostic Contrastive Mixup
abstract
Self-supervised learning (SSL) leverages large amounts of unlabelled data to learn rich speech representations, fostering improvements in automatic speech recognition (ASR), even when only a small amount of labelled data is available for fine-tuning. Despite the advances in SSL, a significant challenge remains when the data used for pre-training (source domain) mismatches the fine-tuning data (target domain). To tackle this domain mismatch challenge, we propose a new domain adaptation method for low-resource ASR focused on contrastive mixup for joint-embedding architectures named AC-Mix (agnostic contrastive mixup). In this approach, the SSL model is adapted through additional pre-training using mixed data views created by interpolating samples from the source and the target domains. Our proposed adaptation method consistently outperforms the baseline system, using approximately 11 hours of adaptation data and requiring only 1 hour of adaptation time on a single GPU with WavLM-Large.
Carlos Carvalho 0003, Alberto Abad
ICASSP2
2025 Acoustic and Linguistic Biomarkers for Cognitive Impairment Detection from Speech
Catarina Botelho, David Gimeno-Gómez, Francisco Teixeira, John Mendonça, Patrícia Pereira, Diogo A. P. Nunes, Thomas Rolland, Anna Pompili, Rubén Solera-Ureña, Maria Ponte, David Martins de Matos, Carlos D. Martínez-Hinarejos, Isabel Trancoso, Alberto Abad
INTERSPEECH14
2025 Exploring Linear Variant Transformers and k-NN Memory Inference for Long-Form ASR
Carlos Carvalho 0003, Jinchuan Tian, Yifan Peng 0003, Alberto Abad, Shinji Watanabe 0001
INTERSPEECH5
2025 On the Relevance of Clinical Assessment Tasks for the Automatic Detection of Parkinson's Disease Medication State from Speech
David Gimeno-Gómez, Rubén Solera-Ureña, Anna Pompili, Carlos D. Martínez-Hinarejos, Rita Cardoso, Isabel Guimarães, Joaquim Ferreira 0002, Alberto Abad
INTERSPEECH8
2025 Exploring Shared-Weight Mechanisms in Transformer and Conformer Architectures for Automatic Speech Recognition
Thomas Rolland, Alberto Abad
INTERSPEECH2
2025 Speech Reference Intervals: An Assessment of Feasibility in Depression Symptom Severity Prediction
abstract
Major Depressive Disorder (MDD) is a prevalent mental disorder. Combining speech features and machine learning has promise for predicting MDD, but interpretability is crucial for clinical applications. Reference intervals (RIs) represent a typical range for a speech feature in a population. RIs could increase interpretability and help clinicians identify deviations from norms. They could also replace conventional speech features in machine learning models. However, no work has yet assessed the feasibility of speech RIs in MDD. We generated and compared RIs from three reference datasets varying in size, elicitation prompt, and health information. We then calculated deviations from each RI set for people with MDD to compare performance on a depression symptom severity prediction task. Our RI-based models trained with demographic data performed similarly to each other and equivalent models using conventional features or demographics only, demonstrating the value of RI-derived features.
Lauren L. White, Ewan Carr, Judith Dineley, Catarina Botelho, Pauline Conde, Faith Matcham, Carolin Oetzmann, Amos Folarin, George Fairs, Agnes Norbury, Stefano Goria, Srinivasan Vairavan, Til Wykes, Richard J. B. Dobson, Vaibhav A. Narayan, Matthew Hotopf, Alberto Abad, Isabel Trancoso, Nicholas Cummins
INTERSPEECH17
2024 Exploring Adapters with Conformers for Children's Automatic Speech Recognition
abstract
The high variability in acoustic, pronunciation, and linguistic characteristics of children’s speech makes of children’s automatic speech recognition (ASR) a complex task. Training a dedicated ASR model from scratch for children remains challenging, mainly due to the limited availability of children’s data. To tackle this limitation, a common strategy involves fine-tuning a pre-trained ASR model. However, this approach faces challenges due to the diversity of speakers and data scarcity, especially when dealing with large ASR models like the Conformer. In this study, we explore an alternative approach known as Adapter transfer. Adapter transfer requires training fewer parameters and can be more effective in adapting large ASR models for children’s speech. In this paper, we assess various Adapter configurations in the literature and introduce a novel configuration called Two Serial Adapter (TSA). The experimental results indicate that Adapter transfer consistently outperforms traditional fine-tuning across various configurations for the Conformer model.
Thomas Rolland, Alberto Abad
ICASSP2
2024 Improved Children's Automatic Speech Recognition Combining Adapters and Synthetic Data Augmentation
abstract
Children’s automatic speech recognition (ASR) poses a significant challenge due to the high variability nature of children’s speech. The limited availability of training datasets hampers the effective modelling of this variability, which can be partially addressed using a text-to-speech (TTS) system for data augmentation. However, generated data may contain imperfections, potentially impacting performance. In this work, we use Adapters to handle the domain mismatch when fine-tuning with TTS data. This involves a two-step training process: training adapter layers with a frozen pre-trained model using synthetic data, then fine-tuning both adapters and the entire model with a mix of synthetic and real data, where only synthetic data passes through the adapters. Experimental results demonstrate up to 6% relative reduction in WER compared to the straightforward use of synthetic data, indicating the effectiveness of adapter-based architectures in learning from imperfect synthetic data.
Thomas Rolland, Alberto Abad
ICASSP2
2024 Macro-descriptors for Alzheimer's disease detection using large language models
Catarina Botelho, John Mendonça, Anna Pompili, Tanja Schultz, Alberto Abad, Isabel Trancoso
INTERSPEECH5
2024 Shared-Adapters: A Novel Transformer-based Parameter Efficient Transfer Learning Approach For Children's Automatic Speech Recognition
Thomas Rolland, Alberto Abad
INTERSPEECH2
2024 Introduction To Partial Fine-tuning: A Comprehensive Evaluation Of End-to-end Children's Automatic Speech Recognition Adaptation
abstract
Automatic Speech Recognition (ASR) encounters unique challenges when dealing with children's speech, mainly due to the scarcity of available data.Training large ASR models with constrained data presents a significant challenge.To address this, fine-tuning strategy is frequently employed.However, fine-tuning an entire large pre-trained model with limited children's speech data may overfit leading to decreased performance.This study offers a granular evaluation of children's ASR fine-tuning, departing from conventional whole-network tunning.We present a partial fine-tuning approach spotlighting the importance of the Encoder and Feedforward Neural Network modules in Transformer-based models.Remarkably, this method surpasses the efficacy of whole-model fine-tuning, with a relative word error rate improvement of 9% when dealing with limited data.Our findings highlight the critical role of partial fine-tuning in advancing children's ASR model development.
Thomas Rolland, Alberto Abad
INTERSPEECH2
2023 Towards Reducing Patient Effort for the Automatic Prediction of Speech Intelligibility in Head and Neck Cancers
abstract
The automatic prediction of speech intelligibility can be seen as a growing and relevant alternative to the perceptual evaluations used clinically, which are known to be biased, variant and subjective. We propose an automatic way to regress an intelligibility score based on a recurrent model with a self-attention mechanism. This approach not only presented a high correlation of 0.87 when applied to a pseudo-word task designed for head and neck cancers, but also a significant decrease in error of more than 50%, when compared to previous approaches. Moreover, we have also studied the reliability of the same system when operating with smaller amounts of data at inference time. The results suggest that we can reduce the linguistic sample size to only 30% of the full sample, without losing performance. This aspect validates the reliability of using a smaller subset of data when predicting intelligibility, which can be extremely useful to prevent patient’s fatigue by creating smaller batteries of clinical exams.
Sebastião Quintas, Alberto Abad, Julie Mauclair, Virginie Woisard, Julien Pinquier
ICASSP2
2023 Privacy-Preserving Automatic Speaker Diarization
abstract
Automatic Speaker Diarization (ASD) is an enabling technology with numerous applications, which deals with recordings of multiple speakers, raising special concerns in terms of privacy. In fact, in remote settings, where recordings are shared with a server, clients relinquish not only the privacy of their conversation, but also of all the information that can be inferred from their voices. However, to the best of our knowledge, the development of privacy-preserving ASD systems has been overlooked thus far. In this work, we tackle this problem using a combination of two cryptographic techniques, Secure Multiparty Computation (SMC) and Secure Modular Hashing, and apply them to the two main steps of a cascaded ASD system: speaker embedding extraction and agglomerative hierarchical clustering. Our system is able to achieve a reasonable trade-off between performance and efficiency, presenting real-time factors of 1.1 and 1.6, for two different SMC security settings.
Francisco Teixeira, Alberto Abad, Bhiksha Raj, Isabel Trancoso
ICASSP2
2023 Towards Reference Speech Characterization for Health Applications
Catarina Botelho, Alberto Abad, Tanja Schultz, Isabel Trancoso
INTERSPEECH2
2023 Memory-augmented conformer for improved end-to-end long-form ASR
abstract
Conformers have recently been proposed as a promising modelling approach for automatic speech recognition (ASR), outperforming recurrent neural network-based approaches and transformers. Nevertheless, in general, the performance of these end-to-end models, especially attention-based models, is particularly degraded in the case of long utterances. To address this limitation, we propose adding a fully-differentiable memory-augmented neural network between the encoder and decoder of a conformer. This external memory can enrich the generalization for longer utterances since it allows the system to store and retrieve more information recurrently. Notably, we explore the neural Turing machine (NTM) that results in our proposed Conformer-NTM model architecture for ASR. Experimental results using Librispeech train-clean-100 and train-960 sets show that the proposed system outperforms the baseline conformer without memory for long utterances.
Carlos Carvalho 0003, Alberto Abad
INTERSPEECH2
2022 Exploring Dementia Detection from Speech: Cross Corpus Analysis
abstract
In this work, we present a qualitative and quantitative analysis of speech and language features derived from two different corpora with the aim to predict early signs of dementia. One corpus consists of the Interdisciplinary Longitudinal Study on Adult Development and Aging (ILSE) designed to investigate satisfying and healthy aging. It consists of more than 6500 hours of biographic interviews from 1000 participants recorded over the course of 20 years. The other corpus is a cross sectional data set created for the ADReSS challenge 2020. In an experimental study we describe a large variety of acoustic and linguistic features that are automatically extracted from speech and corresponding transcriptions. We compare different traditional classifiers, i.e. Gaussian Mixture Models, Linear Discriminant Analysis, and Support Vector Machines. Our final performance results surpass the ADReSS benchmarks.
Ayimnisagul Ablimit, Catarina Botelho, Alberto Abad, Tanja Schultz, Isabel Trancoso
ICASSP3
2022 Challenges of using longitudinal and cross-domain corpora on studies of pathological speech
Catarina Botelho, Tanja Schultz, Alberto Abad, Isabel Trancoso
INTERSPEECH3
2022 Towards End-to-End Private Automatic Speaker Recognition
abstract
The development of privacy-preserving automatic speaker verification systems has been the focus of a number of studies with the intent of allowing users to authenticate themselves without risking the privacy of their voice. However, current privacy-preserving methods assume that the template voice representations (or speaker embeddings) used for authentication are extracted locally by the user. This poses two important issues: first, knowledge of the speaker embedding extraction model may create security and robustness liabilities for the authentication system, as this knowledge might help attackers in crafting adversarial examples able to mislead the system; second, from the point of view of a service provider the speaker embedding extraction model is arguably one of the most valuable components in the system and, as such, disclosing it would be highly undesirable. In this work, we show how speaker embeddings can be extracted while keeping both the speaker's voice and the service provider's model private, using Secure Multiparty Computation. Further, we show that it is possible to obtain reasonable trade-offs between security and computational cost. This work is complementary to those showing how authentication may be performed privately, and thus can be considered as another step towards fully private automatic speaker recognition.
Francisco Teixeira, Alberto Abad, Bhiksha Raj, Isabel Trancoso
INTERSPEECH2
2022 A new European Portuguese corpus for the study of Psychosis through speech analysis
abstract
Psychosis is a clinical syndrome characterized by the presence of symptoms such as hallucinations, thought disorder and disorganized speech. Several studies have used machine learning, combined with speech and natural language processing methods to aid in the diagnosis process of this disease. This paper describes the creation of the first European Portuguese corpus for the identification of the presence of speech characteristics of psychosis, which contains samples of 92 participants, 56 controls and 36 individuals diagnosed with psychosis and medicated. The corpus was used in a set of experiments that allowed identifying the most promising feature set to perform the classification: the combination of acoustic and speech metric features. Several classifiers were implemented to study which ones entailed the best performance depending on the task and feature set. The most promising results obtained for the entire corpus were achieved when identifying individuals with a Multi-Layer Perceptron classifier and reached an 87.5% accuracy. Focusing on the gender dependent results, the overall best results were 90.9% and 82.9% accuracy, for female and male subjects respectively. Lastly, the experiments performed lead us to conjecture that spontaneous speech presents more identifiable characteristics than read speech to differentiate healthy and patients diagnosed with psychosis.
Maria Forjó, Daniel Neto, Alberto Abad, H. Sofia Pinto, Joaquim Gago
LREC3
2022 Multilingual Transfer Learning for Children Automatic Speech Recognition
abstract
Despite recent advances in automatic speech recognition (ASR), the recognition of children’s speech still remains a significant challenge. This is mainly due to the high acoustic variability and the limited amount of available training data. The latter problem is particularly evident in languages other than English, which are usually less-resourced. In the current paper, we address children ASR in a number of less-resourced languages by combining several small-sized children speech corpora from these languages. In particular, we address the following research question: Does a novel two-step training strategy in which multilingual learning is followed by language-specific transfer learning outperform conventional single language/task training for children speech, as well as multilingual and transfer learning alone? Based on previous experimental results with English, we hypothesize that multilingual learning provides a better generalization of the underlying characteristics of children’s speech. Our results provide a positive answer to our research question, by showing that using transfer learning on top of a multilingual model for an unseen language outperforms conventional single language-specific learning.
Thomas Rolland, Alberto Abad, Catia Cucchiarini, Helmer Strik
LREC2
2021 FoolHD: Fooling Speaker Identification by Highly Imperceptible Adversarial Disturbances
abstract
Speaker identification models are vulnerable to carefully designed adversarial perturbations of their input signals that induce misclassification. In this work, we propose a white-box steganography-inspired adversarial attack that generates imperceptible adversarial perturbations against a speaker identification model. Our approach, FoolHD, uses a Gated Convolutional Autoencoder that operates in the DCT domain and is trained with a multi-objective loss function, to generate and conceal the adversarial perturbation within the original audio files. In addition to hindering speaker identification performance, this multi-objective loss accounts for human perception through a frame-wise cosine similarity between MFCC feature vectors extracted from the original and adversarial audio files. We validate the effectiveness of FoolHD with a 250-speaker identification x-vector network, trained using VoxCeleb, in terms of accuracy, success rate, and imperceptibility. Our results show that FoolHD generates highly imperceptible adversarial audio files (average PESQ scores above 4.30), while achieving a success rate of 99.6% and 99.2% in misleading the speaker identification model, for untargeted and targeted settings, respectively.
Ali Shahin Shamsabadi, Francisco Teixeira, Alberto Abad, Bhiksha Raj, Andrea Cavallaro, Isabel Trancoso
ICASSP3
2021 Visual Speech for Obstructive Sleep Apnea Detection
Catarina Botelho, Alberto Abad, Tanja Schultz, Isabel Trancoso
Interspeech2
2021 Transfer Learning-Based Cough Representations for Automatic Detection of COVID-19
abstract
In the last months, there has been an increasing interest in developing reliable, cost-effective, immediate and easy to use machine learning based tools that can help health care operators, institutions, companies, etc. to optimize their screening campaigns.In this line, several initiatives emerged aimed at the automatic detection of COVID-19 from speech, breathing and coughs, with inconclusive preliminary results.The ComParE 2021 COVID-19 Cough Sub-challenge provides researchers from all over the world a suitable test-bed for the evaluation and comparison of their work.In this paper, we present the INESC-ID contribution to the ComParE 2021 COVID-19 Cough Sub-challenge.We leverage transfer learning to develop a set of three expert classifiers based on deep cough representation extractors.A calibrated decision-level fusion system provides the final classification of coughs recordings as either COVID-19 positive or negative.Results show unweighted average recalls of 72.3% and 69.3% in the development and test sets, respectively.Overall, the experimental assessment shows the potential of this approach although much more research on extended respiratory sounds datasets is needed.
Rubén Solera-Ureña, Catarina Botelho, Francisco Teixeira, Thomas Rolland, Alberto Abad, Isabel Trancoso
Interspeech5
2020 Cross Lingual Transfer Learning for Zero-Resource Domain Adaptation
abstract
We propose a method for zero-resource domain adaptation of DNN acoustic models, for use in low-resource situations where the only in-language training data available may be poorly matched to the intended target domain. Our method uses a multi-lingual model in which several DNN layers are shared between languages. This architecture enables domain adaptation transforms learned for one well-resourced language to be applied to an entirely different low- resource language. First, to develop the technique we use English as a well-resourced language and take Spanish to mimic a low-resource language. Experiments in domain adaptation between the conversational telephone speech (CTS) domain and broadcast news (BN) domain demonstrate a 29% relative WER improvement on Spanish BN test data by using only English adaptation data. Second, we demonstrate the effectiveness of the method for low-resource languages with a poor match to the well-resourced language. Even in this scenario, the proposed method achieves relative WER improvements of 18-27% by using solely English data for domain adaptation. Compared to other related approaches based on multi-task and multi-condition training, the proposed method is able to better exploit well-resource language data for improved acoustic modelling of the low-resource target domain.
Alberto Abad, Peter Bell 0001, Andrea Carmantini, Steve Renals
ICASSP1
2020 Toward Silent Paralinguistics: Speech-to-EMG - Retrieving Articulatory Muscle Activity from Speech
abstract
Electromyographic (EMG) signals recorded during speech production encode information on articulatory muscle activity and also on the facial expression of emotion, thus representing a speech-related biosignal with strong potential for paralinguistic applications.In this work, we estimate the electrical activity of the muscles responsible for speech articulation directly from the speech signal.To this end, we first perform a neural conversion of speech features into electromyographic time domain features, and then attempt to retrieve the original EMG signal from the time domain features.We propose a feed forward neural network to address the first step of the problem (speech features to EMG features) and a neural network composed of a convolutional block and a bidirectional long short-term memory block to address the second problem (true EMG features to EMG signal).We observe that four out of the five originally proposed time domain features can be estimated reasonably well from the speech signal.Further, the five time domain features are able to predict the original speech-related EMG signal with a concordance correlation coefficient of 0.663.We further compare our results with the ones achieved on the inverse problem of generating acoustic speech features from EMG features.
Catarina Botelho, Lorenz Diener, Dennis Küster, Kevin Scheck, Shahin Amiriparian, Björn W. Schuller, Tanja Schultz, Alberto Abad, Isabel Trancoso
INTERSPEECH8
2020 Exploring Text and Audio Embeddings for Multi-Dimension Elderly Emotion Recognition
Mariana Julião, Alberto Abad, Helena Moniz
INTERSPEECH2
2020 Analyzing Breath Signals for the Interspeech 2020 ComParE Challenge
John Mendonça, Francisco Teixeira, Isabel Trancoso, Alberto Abad
INTERSPEECH4
2020 The INESC-ID Multi-Modal System for the ADReSS 2020 Challenge
abstract
This paper describes a multi-modal approach for the automatic detection of Alzheimer's disease proposed in the context of the INESC-ID Human Language Technology Laboratory participation in the ADReSS 2020 challenge.Our classification framework takes advantage of both acoustic and textual feature embeddings, which are extracted independently and later combined.Speech signals are encoded into acoustic features using DNN speaker embeddings extracted from pre-trained models.For textual input, contextual embedding vectors are first extracted using an English Bert model and then used either to directly compute sentence embeddings or to feed a bidirectional LSTM-RNNs with attention.Finally, an SVM classifier with linear kernel is used for the individual evaluation of the three systems.Our best system, based on the combination of linguistic and acoustic information, attained a classification accuracy of 81.25%.Results have shown the importance of linguistic features in the classification of Alzheimer's Disease, which outperforms the acoustic ones in terms of accuracy.Early stage features fusion did not provide additional improvements, confirming that the discriminant ability conveyed by speech in this case is smooth out by linguistic data.
Anna Pompili, Thomas Rolland, Alberto Abad
INTERSPEECH3
2020 Assessment of Parkinson's Disease Medication State Through Automatic Speech Analysis
abstract
Parkinson's disease (PD) is a progressive degenerative disorder of the central nervous system characterized by motor and non-motor symptoms. As the disease progresses, patients alternate periods in which motor symptoms are mitigated due to medication intake (ON state) and periods with motor complications (OFF state). The time that patients spend in the OFF condition is currently the main parameter employed to assess pharmacological interventions and to evaluate the efficacy of different active principles. In this work, we present a system that combines automatic speech processing and deep learning techniques to classify the medication state of PD patients by leveraging personal speech-based bio-markers. We devise a speaker-dependent approach and investigate the relevance of different acoustic-prosodic feature sets. Results show an accuracy of 90.54% in a test task with mixed speech and an accuracy of 95.27% in a semi-spontaneous speech task. Overall, the experimental assessment shows the potentials of this approach towards the development of reliable, remote daily monitoring and scheduling of medication intake of PD patients.
Anna Pompili, Rubén Solera-Ureña, Alberto Abad, Rita Cardoso, Isabel Guimarães, Margherita Fabbri, Isabel P. Martins, Joaquim Ferreira 0002
INTERSPEECH3
2019 Speech as a Biomarker for Obstructive Sleep Apnea Detection
abstract
Obstructive sleep apnea (OSA) is a prevalent sleep disorder, responsible for a decrease of people's quality of life, and significant morbidity and mortality associated with hypertension and cardiovascular diseases. OSA is caused by anatomical and functional alterations in the upper airways, thus we hypothesize that the speech properties of OSA patients are altered, making it possible to detect OSA through voice analysis. To address this hypothesis, we collected speech recordings from 25 OSA subjects and 20 controls, designed a feature set, and compared different machine learning algorithms for binary classification. We achieved a True-Positive-Rate of 88% and a True-Negative-Rate of 80% with a majority vote ensemble of SVM, LDA and kNN classifiers. These results were validated with in-the-wild data acquired from Youtube. Moreover, the negative impact of sleep disorders on working memory was also shown by the results obtained in one of the recorded verbal tasks.
Catarina Botelho, Isabel Trancoso, Alberto Abad, Teresa Paiva
ICASSP3
2019 Attentive Filtering Networks for Audio Replay Attack Detection
abstract
An attacker may use a variety of techniques to fool an automatic speaker verification system into accepting them as a genuine user. Anti-spoofing methods meanwhile aim to make the system robust against such attacks. The ASVspoof 2017 Challenge focused specifically on replay attacks, with the intention of measuring the limits of replay attack detection as well as developing countermeasures against them. In this work, we propose our replay attacks detection system - Attentive Filtering Network, which is composed of an attention-based filtering mechanism that enhances feature representations in both the frequency and time domains, and a ResNet-based classifier. We show that the network enables us to visualize the automatically acquired feature representations that are helpful for spoofing detection. Attentive Filtering Network attains an evaluation EER of 8.99% on the ASVspoof 2017 Version 2.0 dataset. With system fusion, our best system further obtains a 30% relative improvement over the ASVspoof 2017 enhanced baseline system.
Cheng-I Lai, Alberto Abad, Korin Richmond, Junichi Yamagishi, Najim Dehak, Simon King 0001
ICASSP2
2019 Privacy-preserving Paralinguistic Tasks
abstract
Speech is one of the primary means of communication for humans. It can be viewed as a carrier for information on several levels as it conveys not only the meaning and intention predetermined by a speaker, but also paralinguistic and extralinguistic information about the speaker's age, gender, personality, emotional state, health state and affect. This makes it a particularly sensitive biometric, that should be protected. In this work we intent to explore how Leveled Homomorphic Encryption can be combined with a Neural Network to create a privacy-preserving machine learning framework for speech-based health-related tasks. In particular, we will apply this framework to the detection and assessment of a Cold, Depression and Parkinson's Disease. Moreover, we will show how using a Quantized Neural Network, with discretized weights, allows us to apply a Leveled Homomorphic Encryption technique called batching that can be utilized to reduce the effective computational cost of this framework.
Francisco Teixeira, Alberto Abad, Isabel Trancoso
ICASSP2
2019 Recognition of Latin American Spanish Using Multi-Task Learning
Carlos Mendes, Alberto Abad, João Paulo da Silva Neto, Isabel Trancoso
INTERSPEECH2
2019 Preserving privacy in speaker and speech characterisation
abstract
Speech recordings are a rich source of personal, sensitive data that can be used to support a plethora of diverse applications, from health profiling to biometric recognition. It is therefore essential that speech recordings are adequately protected so that they cannot be misused. Such protection, in the form of privacy-preserving technologies, is required to ensure that: (i) the biometric profiles of a given individual (e.g., across different biometric service operators) are unlinkable; (ii) leaked, encrypted biometric information is irreversible, and that (iii) biometric references are renewable. Whereas many privacy-preserving technologies have been developed for other biometric characteristics, very few solutions have been proposed to protect privacy in the case of speech signals. Despite privacy preservation this is now being mandated by recent European and international data protection regulations. With the aim of fostering progress and collaboration between researchers in the speech, biometrics and applied cryptography communities, this survey article provides an introduction to the field, starting with a legal perspective on privacy preservation in the case of speech data. It then establishes the requirements for effective privacy preservation, reviews generic cryptography-based solutions, followed by specific techniques that are applicable to speaker characterisation (biometric applications) and speech characterisation (non-biometric applications). Glancing at non-biometrics, methods are presented to avoid function creep, preventing the exploitation of biometric information, e.g., to single out an identity in speech-assisted health care via speaker characterisation. In promoting harmonised research, the article also outlines common, empirical evaluation metrics for the assessment of privacy-preserving technologies for speech data.
Andreas Nautsch, Abelino Jiménez, Amos Treiber, Jascha Kolberg, Catherine Jasserand, Els Kindt, Héctor Delgado, Massimiliano Todisco, Mohamed Amine Hmani, Aymen Mtibaa, Mohammed Ahmed Abdelraheem, Alberto Abad, Francisco Teixeira, Driss Matrouf, Marta Gomez-Barrero, Dijana Petrovska-Delacrétaz, Gérard Chollet, Nicholas W. D. Evans, Christoph Busch 0001
Comput. Speech Lang.12
2018 Exploring Hashing and Cryptonet Based Approaches for Privacy-Preserving Speech Emotion Recognition
abstract
The outsourcing of machine learning classification and data mining tasks can be an effective solution for those parties that need machine learning services, but lack the appropriate resources, knowledge and/or tools to carry them out, in their own premises. This solution, however, raises major privacy concerns, in particular, when irrevocable biometric data such as speech is involved. In this work, we focus on the development of privacy-preserving schemes in a speech emotion recognition task, as a proof of concept that could be extended to other speech analytics tasks. Our aim is to prove that the implementation of privacy-preserving speech mining schemes in challenging tasks involving paralinguistic features is not only feasible, but also accurate. Using distance-preserving hashing techniques in a first approach, and homomorphic encryption in a second approach, we successfully protect sensitive data with little degradation costs regarding the accuracy of the predictive models.
José Miguel Salles Dias, Alberto Abad, Isabel Trancoso
ICASSP2
2018 Patient Privacy in Paralinguistic Tasks
Francisco Teixeira, Alberto Abad, Isabel Trancoso
INTERSPEECH2
2017 Depression Detection Using Automatic Transcriptions of De-Identified Speech
Paula Lopez-Otero, Laura Docío Fernández, Alberto Abad, Carmen García-Mateo
INTERSPEECH3
2016 Exploiting Phone Log-Likelihood Ratio Features for the Detection of the Native Language of Non-Native English Speakers
Alberto Abad, Eugénio Ribeiro, Fábio N. Kepler, Ramón Fernandez Astudillo, Isabel Trancoso
INTERSPEECH1
2016 SPA: Web-based Platform for easy Access to Speech Processing Modules
Fernando Batista, Pedro Curto, Isabel Trancoso, Alberto Abad, Jaime Ferreira, Eugénio Ribeiro, Helena Moniz, David Martins de Matos, Ricardo Ribeiro 0001
LREC4
2016 The SpeDial datasets: datasets for Spoken Dialogue Systems analytics
José Lopes 0001, Arodami Chorianopoulou, Elisavet Palogiannidi, Helena Moniz, Alberto Abad, Katerina Louka, Elias Iosif, Alexandros Potamianos
LREC5
2016 The DIRHA Portuguese Corpus: A Comparison of Home Automation Command Detection and Recognition in Simulated and Real Data
Miguel Matos, Alberto Abad, António Joaquim Serralheiro
LREC2
2015 Privacy-preserving Query-by-Example Speech Search
abstract
This paper investigates a new privacy-preserving paradigm for the task of Query-by-Example Speech Search using Secure Binary Embeddings, a hashing method that converts vector data to bit strings through a combination of random projections followed by banded quantization. The proposed method allows performing spoken query search in an encrypted domain, by analyzing ciphered information computed from the original recordings. Unlike other hashing techniques, the embeddings allow the computation of the distance between vectors that are close enough, but are not perfect matches. This paper shows how these hashes can be combined with Dynamic Time Warping based on posterior derived features to perform secure speech search. Experiments performed on a sub-set of the Speech-Dat Portuguese corpus showed that the proposed privacy-preserving system obtains similar results to its non-private counterpart.
José Portelo, Alberto Abad, Bhiksha Raj, Isabel Trancoso
ICASSP2
2015 Multi-channel speaker verification based on total variability modelling
Maria Joana Correia, Alessio Brutti, Alberto Abad
INTERSPEECH3
2015 Detecting repetitions in spoken dialogue systems using phonetic distances
abstract
This paper addresses the problem of automatic detection of re-peated turns in Spoken Dialogue Systems. Repetitions can be a symptom of problematic communication between users and systems. Such repetitions are often due to speech recognition errors, which in turn makes it hard to use speech recognition to detect repetitions. We present an approach to detect rep-etition using the phonetic distance to find the best alignment between turns in the same dialogue. The alignment score ob-tained is combined with different features to improve repeti-tion detection. To evaluate the method proposed we compare several alignment techniques from edit distance to DTW-based distance, previously used in Spoken-Term detection tasks. We also compare two different methods to compute the phonetic distance: the first one using the phoneme sequence, and the second one using the distance between the phone posterior vec-tors. Two different datasets were used in this evaluation: a bus-schedule information system (in English) and a call routing system (in Swedish). The results show that approaches using phoneme distances over-perform approaches using Levenshtein distances between ASR outputs for repetition detection. Index Terms: spoken dialogue systems, repetition detection, phonetic distance
José Lopes 0001, Giampiero Salvi, Gabriel Skantze, Alberto Abad, Joakim Gustafson, Fernando Batista, Raveesh Meena, Isabel Trancoso
INTERSPEECH4
2015 Combining multiple approaches to predict the degree of nativeness
abstract
Automatic speaker nativeness assessment has multiple applications, such as second language learning and IVR systems. In this paper we view this as a regression problem, since the available labels are on a continuous scale. Multiple approaches were applied, such as phonotactic models, i-vectors, and goodness of pronunciation, covering both segmental and suprasegmental features. Different phonotactic models were adopted, either trained with the challenge data, or using additional multilingual data from other domains. The obtained values were later combined in multiple ways and fed to a support vector machine regressor. Results on the test set surpass the provided baseline and are in line with the results obtained on the remaining sets. This suggests that our models generalize well to other datasets
Eugénio Ribeiro, Jaime Ferreira, Julia Olcoz, Alberto Abad, Helena Moniz, Fernando Batista, Isabel Trancoso
INTERSPEECH4
2014 Accounting for the residual uncertainty of multi-layer perceptron based features
abstract
Multi-Layer Perceptrons (MLPs) are often interpreted as modeling a posterior distribution over classes given input features using the mean field approximation. This approximation is fast but neglects the residual uncertainty of inference at each layer, making inference less robust. In this paper we introduce a new approximation of MLP inference that takes under consideration this residual uncertainty. The proposed algorithm propagates not only the mean, but also the variance of inference through the network. At the current stage, the proposed method can not be used with soft-max layers. Therefore, we illustrate the benefits of this algorithm in a tandem scheme. We use the residual uncertainty of inference of MLP-based features to compensate a GMM-HMM backend with uncertainty decoding. Experiments on the Aurora4 corpus show consistent improvement of performance against conventional MLPs for all scenarios, in particular for clean speech and multi-style training.
Ramón Fernandez Astudillo, Alberto Abad, Isabel Trancoso
ICASSP2
2014 The DIRHA simulated corpus
Luca Cristoforetti, Mirco Ravanelli, Maurizio Omologo, Alessandro Sosi, Alberto Abad, Martin Hagmüller, Petros Maragos
LREC5
2014 Exploiting magnitude and phase spectral information for converted speech detection
abstract
Speaker verification systems have been shown to be vulnerable in situations where voice conversion techniques are used to try to fool them, evidencing an important security breach in these applications. This work focuses on the development of a new converted speech detector able to robustly address this problem. The proposed detector uses four spectral features extracted from the magnitude and the phase spectrum of the speech signal. To evaluate the performance of the detector we use a subset of the core task of the NIST SRE2006 corpus as the natural data. The converted data was produced with two different voice conversion methods: Gaussian mixture model and unit selection, from other NIST SRE2006 conditions. The converted speech detector achieved a detection accuracy of 99.1% and 98.5% for natural and converted utterances, respectively.
Maria Joana Correia, Alberto Abad, Isabel Trancoso
SLT2
2013 On the calibration and fusion of heterogeneous spoken term detection systems
abstract
The combination of several heterogeneous systems is known to provide remarkable performance improvements in verification and detection tasks. In Spoken Term Detection (STD), two important issues arise: (1) how to define a com-mon set of detected candidates, and (2) how to combine system scores to produce a single score per candidate. In this paper, a discriminative calibration/fusion approach commonly applied in speaker and language recognition is adopted for STD. Under this approach, we first propose several heuristics to hypothesize scores for systems that do not detect a given candidate. In this way, the original problem of several unaligned detection can-didates is converted into a verification task. As for other ver-ification tasks, system weights and offsets are then estimated through linear logistic regression. As a result, the combined scores are well calibrated, and the detection threshold is au-tomatically given by application parameters (priors and costs). The proposed method not only offers an elegant solution for the problem of fusion and calibration of multiple detectors, but also provides consistent improvements over a baseline approach based on majority voting, according to experiments on the Me-diaEval 2012 Spoken Web Search (SWS) task involving 8 het-erogeneous systems developed at two different laboratories.
Alberto Abad, Luis Javier Rodríguez-Fuentes, Mikel Peñagarikano, Amparo Varona, Germán Bordel
INTERSPEECH1
2013 Secure binary embeddings of front-end factor analysis for privacy preserving speaker verification
abstract
Remote speaker verification services typically rely on the system to have access to the users recordings, or features derived from them, and also a model of the users voice. This conventional scheme raises several privacy concerns. In this work, we address this privacy problem in the context of a speaker verification system using a factor analysis based front-end extractor, the so-called i-vectors. Speaker verification without exposing speaker data is achieved by transforming speaker i-vectors to bit strings in a way that allows the computation of approximate distances, instead of exact ones. The key to the transformation uses a hashing scheme known as Secure Binary Embeddings. Then, a modified SVM kernel permits operating on the i-vector hashes. Experiments on sub-sets of NIST SRE 2008 showed that the secure system yielded similar results as its non-private counterpart.
José Portelo, Alberto Abad, Bhiksha Raj, Isabel Trancoso
INTERSPEECH2
2013 Automatic word naming recognition for an on-line aphasia treatment system
Alberto Abad, Anna Pompili, Ângela Costa, Isabel Trancoso, José G. Fonseca, Gabriela Leal, Luisa Farrajota, Isabel P. Martins
Comput. Speech Lang.1
2013 Integration of beamforming and uncertainty-of-observation techniques for robust ASR in multi-source environments
Ramón Fernandez Astudillo, Dorothea Kolossa, Alberto Abad, Steffen Zeiler, Rahim Saeidi, Pejman Mowlaee, João Paulo da Silva Neto, Rainer Martin 0001
Comput. Speech Lang.3
2012 Integration of beamforming and automatic speech recognition through propagation of the wiener posterior
abstract
This paper details one of the front-end components of the system used at the PASCAL-CHiME multi-source robust automatic speech recognition (ASR) challenge 2011. The presented approach uses uncertainty propagation techniques to integrate conventional beamforming with automatic speech recognition. The paper addresses the derivation of a complex Gaussian posterior for the multi-channel Wiener and the delay and sum beamformer and introduces a new approach based on the propagation of the Wiener posterior through the resynthesizing process. Results on the PASCAL-CHiME task for this algorithms show that they consistently outperform conventional beamfomers with a minimal increase in computational complexity.
Ramón Fernandez Astudillo, Alberto Abad, João Paulo da Silva Neto
ICASSP2
2012 Automatic word naming recognition for treatment and assessment of aphasia
abstract
VITHEA is an on-line platform designed to act as a “virtual therapist” for the treatment of Portuguese speaking aphasic patients. Concretely, the system integrates automatic speech recognition technology to provide word naming exercises to individuals with lost or reduced word naming ability. In this paper, we present the solution adopted for the word naming task, which is based on a keyword spotting approach with hybrid HMM/MLP speech recognizer. Furthermore, we explore a simple cross-validation method that makes use of the patients measured word naming ability to automatically adapt to their speech particularities. A corpus with word naming therapy sessions of aphasic Portuguese native speakers has been collected to test the utility of the approach for both global evaluation and treatment. In spite of the different patient characteristics and speech quality conditions of the collected data, encouraging results have been obtained.
Alberto Abad, Anna Pompili, Ângela Costa, Isabel Trancoso
INTERSPEECH1
2012 Uncertainty driven Compensation of Multi-Stream MLP Acoustic Models for Robust ASR
Ramón Fernandez Astudillo, Alberto Abad, João Paulo da Silva Neto
INTERSPEECH2
2012 The BLZ Submission to the NIST 2011 LRE: Data Collection, System Development and Performance
abstract
This paper describes the most relevant features of a collaborative multi-site submission to the NIST 2011 Language Recognition Evaluation (LRE), consisting of one primary and three contrastive systems, each fusing different combinations of 13 state-of-the-art (acoustic and phonotactic) language recognition subsystems.The collaboration focused on collecting and sharing training data for those target languages for which few development data were provided by NIST, and on defining a common development dataset to train backend and fusion parameters and select the best fusions.Official and post-key results are presented and compared, revealing that the greedy approach applied to select the best fusions provided suboptimal but very competitive performance.Several factors contributed to the high performance attained by BLZ systems, including the availability of training data for low resource target languages, the reliability of the development dataset (consisting only of data audited by NIST), the diversity of modeling approaches, features and datasets in the systems considered for fusion, and the effectiveness of the search for optimal fusions.
Luis Javier Rodríguez-Fuentes, Mikel Peñagarikano, Amparo Varona, Mireia Díez, Germán Bordel, Alberto Abad, David Martínez González, Jesús Villalba 0001, Alfonso Ortega Giménez, Eduardo Lleida
INTERSPEECH6
2012 Algorithm 924: TIDES, a Taylor Series Integrator for Differential EquationS
abstract
This article introduces the software package TIDES and revisits the use of the Taylor series method for the numerical integration of ODEs. The package TIDES provides an easy-to-use interface for standard double precision integrations, but also for quadruple precision and multiple precision integrations. The motivation for the development of this package is that more and more scientific disciplines need very high precision solution of ODEs, and a standard ODE method is not able to reach these precision levels. The TIDES package combines a preprocessor step in M athematica that generates Fortran or C programs with a library in C. Another capability of TIDES is the direct solution of sensitivities of the solution of ODE systems, which means that we can compute the solution of variational equations up to any order without formulating them explicitly. Different options of the software are discussed, and finally it is compared with other well-known available methods, as well as with different options of TIDES. From the numerical tests, TIDES is competitive, both in speed and accuracy, with standard methods, but it also provides new capabilities.
Alberto Abad, Roberto Barrio, Fernando Blesa, Marcos Rodríguez
ACM Trans. Math. Softw.1
2011 Multi-site heterogeneous system fusions for the Albayzin 2010 Language Recognition Evaluation
abstract
Best language recognition performance is commonly obtained by fusing the scores of several heterogeneous systems. Regardless the fusion approach, it is assumed that different systems may contribute complementary information, either because they are developed on different datasets, or because they use different features or different modeling approaches. Most authors apply fusion as a final resource for improving performance based on an existing set of systems. Though relative performance gains decrease as larger sets of systems are considered, best performance is usually attained by fusing all the available systems, which may lead to high computational costs. In this paper, we aim to discover which technologies combine the best through fusion and to analyse the factors (data, features, modeling methodologies, etc.) that may explain such a good performance. Results are presented and discussed for a number of systems provided by the participating sites and the organizing team of the Albayzin 2010 Language Recognition Evaluation. We hope the conclusions of this work help research groups make better decisions in developing language recognition technology.
Luis Javier Rodríguez-Fuentes, Mikel Peñagarikano, Amparo Varona, Mireia Díez, Germán Bordel, David Martínez González, Jesús Villalba 0001, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida, Alberto Abad, Oscar Koller, Isabel Trancoso, Paula Lopez-Otero, Laura Docío Fernández, Carmen García-Mateo, Rahim Saeidi, Mehdi Soufifar, Tomi Kinnunen, Torbjørn Svendsen, Pasi Fränti
ASRU11
2011 Parallel Transformation Network features for speaker recognition
abstract
The use of speaker adaptation transforms as features for speaker recognition is an appealing alternative to conventional short-term cepstral features. In general, this kind of methods are language dependent and limited by the need of speech recognition in the client speakers language. In this paper, we generalize a recently pro posed method -named Transformation Network features with SVM modeling- in order to become language independent and overcome the need for accurate speech recognition. This is accomplished by using a set of parallel acoustic models in several different languages to obtain a high-dimensional Parallel Transformation Network feature vector for speaker characterization.
Alberto Abad, Jordi Luque, Isabel Trancoso
ICASSP1
2011 A nativeness classifier for TED Talks
abstract
This paper presents a nativeness classifier for English. The detector was developed and tested with TED Talks collected from the web, where the major non-native cues are in terms of segmental aspects and prosody. The first experiments were made using only acoustic features, with Gaussian supervectors for training a classifier based on support vector machines. These experiments resulted in an equal error rate of 13.11%. The following experiments based on prosodic features alone did not yield good results. However, a fused system, combining acoustic and prosodic cues, achieved an equal error rate of 10.58%. A small human benchmark was conducted, showing an inter-rater agreement of 0.88. This value is also very close to the agreement value between humans and the best fused system.
José Lopes 0001, Isabel Trancoso, Alberto Abad
ICASSP3
2010 Context dependent modelling approaches for hybrid speech recognizers
abstract
Speech recognition based on connectionist approaches is one of the most successful alternatives to widespread Gaussian systems. One of the main claims against hybrid recognizers is the increased complexity for context-dependent phone modeling, which is a key aspect in medium to large size vocabulary tasks. In this paper, we investigate the use of context-dependent triphone models in a connectionist speech recognizer. Thus, most common triphone state clustering procedures for Gaussian models are compared and applied to our hybrid recognizer. The developed systems with clustered context-dependent triphones show above 20% relative word error rate reduction compared to a baseline hybrid system in two selected WSJ evaluation test sets. Additionally, the recent porting efforts of the proposed context modelling approaches to a LVCSR system for English Broadcast News transcription are reported. Index Terms: speech recognition, context modeling, connectionist system
Alberto Abad, Thomas Pellegrini, Isabel Trancoso, João Paulo da Silva Neto
INTERSPEECH1
2010 Speaker recognition experiments using connectionist transformation network features
abstract
The use of adaptation transforms common in speech recognition systems as features for speaker recognition is an appealing alternative approach to conventional short-term cepstral modelling of speaker characteristics. Recently, we have shown that it is possible to use transformation weights derived from adaptation techniques applied to the Multi Layer Perceptrons that form a connectionist speech recognizer. The proposed method – named Transformation Network features with SVM modelling (TN-SVM) – showed promising results on a sub-set of NIST SRE 2008 and allowed further improvements when it was combined with baseline systems. In this paper, we summarize the recently proposed TN-SVM approach and present new results. First, we explore two alternative approaches that may be used in the absence of high quality speech transcriptions. Second, we present results of the proposed approach with Nuisance Attribute Projection for session variability compensation.
Alberto Abad, Isabel Trancoso
INTERSPEECH1
2010 Exploiting variety-dependent phones in portuguese variety identification applied to broadcast news transcription
abstract
This paper presents a Variety IDentification (VID) approach and its application to broadcast news transcription for Portuguese. The phonotactic VID system, based on Phone Recognition and Language Modelling, focuses on a single tokenizer that combines distinctive knowledge about differences between the target varieties. This knowledge is introduced into a Multi-Layer Perceptron phone recognizer by training mono-phone models for two varieties as contrasting phone-like classes. Significant improvements in terms of identification rate were achieved compared to conventional single and fused phonotactic and acoustic systems. The VID system is used to select data to automatically train variety-specific acoustic models for broadcast news transcription. The impact of the selection is analyzed and variety-specific recognition is shown to improve results by up to 13% compared to a standard variety baseline. Index Terms: accent identification, recognition of accented speech.
Oscar Koller, Alberto Abad, Isabel Trancoso, Céu Viana
INTERSPEECH2
2009 Non-speech audio event detection
abstract
Audio event detection is one of the tasks of the European project VIDIVIDEO. This paper focuses on the detection of non-speech events, and as such only searches for events in audio segments that have been previously classified as non-speech. Preliminary experiments with a small corpus of sound effects have shown the potential of this type of corpus for training purposes. This paper describes our experiments with SVM and HMM-based classifiers, using a 290-hour corpus of sound effects. Although we have only built detectors for 15 semantic concepts so far, the method seems easily portable to other concepts. The paper reports experiments with multiple features, different kernels and several analysis windows. Preliminary experiments on documentaries and films yielded promising results, despite the difficulties posed by the mixtures of audio events that characterize real sounds.
José Portelo, Miguel M. F. Bugalho, Isabel Trancoso, João Paulo da Silva Neto, Alberto Abad, António Joaquim Serralheiro
ICASSP5
2009 Audio contributions to semantic video search
abstract
This paper summarizes the contributions to semantic video search that can be derived from the audio signal. Because of space restrictions, the emphasis will be on non-linguistic cues. The paper thus covers what is generally known as audio segmentation, as well as audio event detection. Using machine learning approaches, we have built detectors for over 50 semantic audio concepts.
Isabel Trancoso, Thomas Pellegrini, José Portelo, Hugo Meinedo, Miguel M. F. Bugalho, Alberto Abad, João Paulo da Silva Neto
ICME6
2009 Porting an european portuguese broadcast news recognition system to brazilian portuguese
abstract
This paper reports on recent work in the context of the activities of the PoSTPort project aimed at porting a Broadcast News recognition system originally developed for European Portuguese to other varieties. Concretely, in this paper we have focused on porting to Brazilian Portuguese. The impact of some of the main sources of variability has been assessed, besides proposing solutions at the lexical, acoustic and syntactic levels. The ported Brazilian Portuguese Broadcast News system allowed a drastic performance improvement from 56.6% WER (obtained with the European Portuguese system) to 25.5%.
Alberto Abad, Isabel Trancoso, Nelson Neto 0001, Céu Viana
INTERSPEECH1
2009 Detecting audio events for semantic video search
abstract
This paper describes our work on audio event detection, one of our tasks in the European project VIDIVIDEO. Preliminary experiments with a small corpus of sound effects have shown the potential of this type of corpus for training purposes. This paper describes our experiments with SVM classifiers, and different features, using a 290-hour corpus of sound effects, which allowed us to build detectors for almost 50 semantic concepts. Although the performance of these detectors on the development set is quite good (achieving an average F-measure of 0.87), preliminary experiments on documentaries and films showed that the task is much harder in real-life videos, which so often include overlapping audio events. Index Terms: event detection, audio segmentation 1.
Miguel M. F. Bugalho, José Portelo, Isabel Trancoso, Thomas Pellegrini, Alberto Abad
INTERSPEECH5
2008 Incorporating acoustical modelling of phone transitions in an hybrid ANN/HMM speech recognizer
abstract
Speech recognition based on connectionist approaches is one of the most successful alternatives to widespread Gaussian systems. One of the main claims against hybrid recognizers is the increased complexity for context-dependent phone modelling, which is a key aspect in medium to large size vocabulary tasks. In this paper, a baseline hybrid system based on monophone recognition units is improved by incorporating acoustical modelling of phone transitions. First, a single state monophone model is extended to multiple state sub-phoneme modelling. Then, a reduced set of diphone recognition units is incorporated to model phone transitions. The proposed approach shows a 26.8 % and 23.8 % relative word error rate reduction compared to baseline hybrid system in two selected WSJ evaluation test sets. Additionally, improved performance compared to a reference Gaussian system based on word-internal context-dependent triphones and comparable results to cross-word triphone system are reported. Index Terms: speech recognition, context modelling, connectionist system 1.
Alberto Abad, João Paulo da Silva Neto
INTERSPEECH1
2008 Speaker orientation estimation based on hybridation of GCC-PHAT and HLBR
abstract
This paper presents a novel approach to speaker orientation\nestimation in a SmartRoom environment equipped with\nmultiple microphones. The ratio between the high and low\nband energies (HLBR) received at each microphone has been\nshown in our previous work to be a potentially approach to estimate\nthe direction of the voice produced by a speaker. In this\nwork, for each microphone pair, a smoothed CPS phase is obtained\nby a proper windowing of the main peak of the crosscorrelation\nsequence estimated with the GCC-PHAT method,\nand a HLBR is computed from the processed CPS. The proposed\nmethod keeps the computational simplicity of the HLBR\nalgorithm while adding the robustness offered by the GCCPHAT\ntechnique. Experimental preliminary results were conducted\nover a database recorded purposely in the UPC Smart\nroom, and over the CLEAR head pose database. The proposed\nmethod performs consistently better than other state-of-the-art\ntechniques with both databases.
Carlos Segura, Alberto Abad, Javier Hernando, Climent Nadeu
INTERSPEECH2
2007 Multimodal Head Orientation Towards Attention Tracking in Smartrooms
abstract
This paper presents a multimodal approach to head pose estimation and 3D gaze orientation of individuals in a SmartRoom environment equipped with multiple cameras and microphones. We first introduce the two monomodal approaches as reference. In video, we estimate head orientation from color information by exploiting spatial redundancy among cameras. Audio information is processed to estimate the direction of the voice produced by a speaker making use of the directivity characteristics of the head radiation pattern. Two multimodal information fusion schemes working at data and decision levels are analyzed in terms of accuracy and robustness of the estimation. Experimental results conducted over the CLEAR evaluation database are reported and the comparison of the proposed multimodal head pose estimation algorithms with the reference monomodal approaches proves the effectiveness of the proposed approach.
Carlos Segura, Cristian Canton, Alberto Abad, Josep R. Casas, Javier Hernando
ICASSP (2)3
2007 Audio-based approaches to head orientation estimation in a smart-room
abstract
The head orientation of human speakers in a smart-room affects the quality of the signals recorded by far-field microphones, and consequently influences the performance of the technologies deployed based on those signals. Additionally, knowing the orientation in these environments can be useful for the development of several multimodal advanced services, for instance, in microphone network management. Consequently, head orientation estimation has recently become a growing interesting research topic. In this paper, we propose two different approaches to head orientation estimation on the basis of multimicrophone recordings: first, an approach based on the generalization of the well-known SRP-PHAT speaker localization algorithm, and second a new approach based on measurements of the ratio between the high and the low band speech energies. Promising results are obtained in both cases, with a generalized better performance of the algorithms based on speaker localization methods. Index Terms: head orientation estimation, microphone arrays, speaker tracking
Alberto Abad, Carlos Segura, Climent Nadeu, Javier Hernando
INTERSPEECH1
2006 Audio person tracking in a smart-room environment
abstract
Reliable measures of speaker positions are needed for computational perception of human activities taking place in a smartroom environment. In this work, it is described the development process and the experiments conducted in the design and implementation of an Audio Person Tracking system for smart-room environments. The proposed system is based on the SRP-PHAT algorithm, as it is known to perform robustly in most environmental conditions. Novelties proposed are aimed to enhance the accuracy of the system independently on the application scenario and to reduce the computational complexity. Index Terms: sound localization, source tracking, microphone arrays.
Alberto Abad, Carlos Segura, Dusan Macho, Javier Hernando, Climent Nadeu
INTERSPEECH1
2006 Architecture and dialogue design for a voice operated information system
Luis Villarejo, Javier Hernando, Núria Castell, Jaume Padrell, Alberto Abad
Appl. Intell.5
2005 Automatic Speech Activity Detection, Source Localization, and Speech Recognition on the Chil Seminar Corpus
abstract
To realize the long-term goal of ubiquitous computing, technological advances in multi-channel acoustic analysis are needed in order to solve several basic problems, including speaker localization and tracking, speech activity detection (SAD) and distant-talking automatic speech recognition (ASR). The European Commission integrated project CHIL, “ Computers in the Human Interaction Loop”, aims to make significant advances in these three technologies. In this work, we report the results of our initial automatic source localization, speech activity detection, and speech recognition experiments on the CHIL seminar corpus, which is comprised of spontaneous speech collected by both near- and far-field microphones. In addition to the audio sensors, the seminars were also recorded by calibrated video cameras. This simultaneous audio-visual data capture enables the realistic evaluation of component technologies as was never possible with earlier data bases.
Dusan Macho, Jaume Padrell, Alberto Abad, Climent Nadeu, Javier Hernando, John W. McDonough, Matthias Wölfel, Ulrich Klee, Maurizio Omologo, Alessio Brutti, Piergiorgio Svaizer, Gerasimos Potamianos, Stephen M. Chu
ICME3
2005 Effect of head orientation on the speaker localization performance in smart-room environment
abstract
Reliable measures of speaker positions are needed for computational perception of human activities taking place in a smart-room environment. In this work, we investigate the effect of talkers head orientation on the accuracy of acoustical source localization techniques and its relation with the talker directivity pattern and room reverberation. Two different representative speaker localization techniques are assessed, steered response power and a crossing lines based method, in both cases on the basis of the estimated delays between pairs of microphones with the GCC-PHAT algorithm. A small database has been collected at the UPC’s smart room for evaluation. The results show how the localization error heavily depends on the head orientation, and also the fact that the space exploration based technique is much more robust to head orientation changes than the crossing lines technique, due to the way the contributions from the various microphones are combined.
Alberto Abad, Dusan Macho, Carlos Segura, Javier Hernando, Climent Nadeu
INTERSPEECH1
2004 Speech enhancement and recognition by integrating adaptive beamforming and wiener filtering
abstract
A robust adaptive beamforming method is presented in this paper for speech enhancement and speech recognition with microphone arrays. The proposal is based on a modification of the Generalized Sidelobe Canceller with adaptive blocking matrix and the use of a Wiener filter. Alternatively to most of the previous reported works based on microphone arrays with postfiltering, the new technique integrates the Wiener filter in the structure of the adaptive beamformer in a single stage. Experimental results show that the proposed integrated adaptive Wiener-filtering (IAW) beamformer usually is more robust to directional and ambient noises than conventional postfiltering of the beamformer output with a lower level of degradation. Speech recognition experiments which show improvements with the proposed beamformer are also reported. 1.
Alberto Abad, Javier Hernando
INTERSPEECH1
2004 Jacobian adaptation with improved noise reference for speaker verification
Jan Anguita, Javier Hernando, Alberto Abad
INTERSPEECH3
2003 Jacobian adaptation based on the frequency-filtered spectral energies
abstract
Jacobian Adaptation (JA) of the acoustic models is an efficient adaptation technique for robust speech recognition. Several improvements for the JA have been proposed in the last years, either to generalize the Jacobian linear transformation for the case of large noise mismatch between training and testing or to extend the adaptation to other degrading factors, like channel distortion and vocal tract length. However, the JA technique has only been used so far with the conventional mel-frequency cepstral coefficients (MFCC). In this paper, the JA technique is applied to an alternative type of features, the Frequency-Filtered (FF) spectral energies, resulting in a more computationally efficient approach. Furthermore, in experimental tests with the database Aurora1, this new approach has shown an improved recognition performance with respect to the Jacobian adaptation with MFCCs.
Alberto Abad, Climent Nadeu, Javier Hernando, Jaume Padrell
INTERSPEECH1
2001 Algebraic and Symbolic Manipulation of Poisson Series
Juan Félix San-Juan, Alberto Abad
J. Symb. Comput.2
1997 PSPCLink: A Cooperation Between General Symbolic and Poisson Series Processors
Alberto Abad, Juan Félix San-Juan
J. Symb. Comput.1