VLDB 2026 Research / reviewers in the wild / expert
Laureano Moro-Velázquez
dblp:173/6483
· DBLP profile ↗
47ranked-venue papers
3as first author
39since 2021 · last 2026
0000-0002-3033-7005ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 2 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 2 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Transcripts: Iterative Peer-Editing with Audio Unlocks High-Quality Human Summaries of Conversational SpeechabstractThere are not enough established benchmarks for the task fo speech summarization. Creating new benchmarks demands human annotation, as LLMs could embed systemic errors and bias into datasets. We test ten annotation workflows varying input modality (audio, transcript, or both) and the inclusion of editing (self or peer-editing) to investigate potential quality tradeoffs from using human annotators to summarize audio. We compare human audio-based summaries to human transcript-based summaries to track the impact of the different information modalities on summary quality. We also compare the human outputs against four LLM benchmarks (three text, one audio) to examine whether human-written summaries are less informative than highly fluent automated outputs. We find that audio-based summaries are less informative and more compressed than transcript summaries. However, iterative peer-editing with audio mitigates this difference, enabling audio-based summaries to be as informative as their transcript counterparts and LLM summaries. These findings validate iterative peer-editing among human annotators for the creation of benchmarks informed by both lexical and prosodic information. This enables crucial dataset collection even in setting where transcripts are unavailable. Kaavya Chaparala, Thomas Thebaud, Jesús Villalba 0001, Laureano Moro-Velázquez, Peter Viechnicki, Najim Dehak |
LREC | 4 |
| 2026 | Towards Fair Speech Recognition: Mitigating Demographic Bias in End-to-End ASR Systems
Maliha Jahan, Thomas Thebaud, Zsuzsanna Fagyal, Jesús Villalba 0001, Mark Hasegawa-Johnson, Laureano Moro-Velázquez, Najim Dehak |
LREC | 6 |
| 2026 | Interpretable Features for the Assessment of Neurodegenerative Diseases Through Handwriting AnalysisabstractMotor dysfunction is a common sign of neurodegenerative diseases (NDs) such as Parkinson's disease (PD) and Alzheimer's disease (AD), but may be difficult to detect, especially in the early stages. In this work, we examine the behavior of a wide array of interpretable features extracted from the handwriting signals of 113 subjects performing multiple tasks on a digital tablet, as part of the Neurological Signals dataset. The aim is to measure their effectiveness in characterizing NDs, including AD and PD. To this end, task-agnostic and task-specific features are extracted from 14 distinct tasks. Subsequently, through statistical analysis and a series of classification experiments, we investigate which features provide greater discriminative power between NDs and healthy controls and amongst different NDs. Preliminary results indicate that the tasks at hand can all be effectively leveraged to distinguish between the considered set of NDs, specifically by measuring the stability, the speed of writing, the time spent not writing, and the pressure variations between groups from our handcrafted interpretable features, which shows a statistically significant difference between groups, across multiple tasks. Using various binary classification algorithms on the computed features, we obtain up to 87% accuracy for the discrimination between AD and healthy controls (CTL), and up to 69% for the discrimination between PD and CTL. Thomas Thebaud, Anna Favaro, Casey Chen, Gabriel Chávez, Laureano Moro-Velázquez, Emile Moukhebeir, Ankur A. Butala, Najim Dehak |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Paired by the Teacher: Turning Unpaired Data into High-Fidelity Pairs for Low-Resource Text GenerationabstractWe present Paired by the Teacher (PbT), a twostage teacher-student pipeline that synthesizes accurate input-output pairs without human labels or parallel data.In many low-resource natural language generation (NLG) scenarios, practitioners may have only raw outputs, like highlights, recaps, or questions, or only raw inputs, such as articles, dialogues, or paragraphs, but seldom both.This mismatch forces small models to learn from very few examples or rely on costly, broad-scope synthetic examples produced by large LLMs.PbT addresses this by asking a teacher LLM to compress each unpaired example into a concise intermediate representation (IR), and training a student to reconstruct inputs from IRs.This enables outputs to be paired with student-generated inputs, yielding high-quality synthetic data.We evaluate PbT on five benchmarks-document summarization (XSum, CNNDM), dialogue summarization (SAMSum, DialogSum), and question generation (SQuAD)-as well as an unpaired setting on SwitchBoard (paired with Dialog-Sum summaries).An 8B student trained only on PbT data outperforms models trained on 70 B teacher-generated corpora and other unsupervised baselines, coming within 1.2 ROUGE-L of human-annotated pairs and closing 82% of the oracle gap at one-third the annotation cost of direct synthesis.Human evaluation on SwitchBoard further confirms that only PbT produces concise, faithful summaries aligned with the target style, highlighting its advantage of generating in-domain sources that avoid the mismatch, limiting direct synthesis. Yen-Ju Lu, Thomas Thebaud, Laureano Moro-Velázquez, Najim Dehak, Jesús Villalba 0001 |
EMNLP | 3 |
| 2025 | Unveiling Performance Bias in ASR Systems: A Study on Gender, Age, Accent, and MoreabstractWith the recent advancements in speech recognition, it is crucial to ensure these systems are free from performance biases against any speaker subgroups. This study examined the performance of twenty variants of seven Automatic Speech Recognition models across four datasets in English language: L2 Arctic, Speech Accent Archive, CORAAL, and SBCSAE. We employed Poisson regression and drop-in-deviance tests to identify which attributes significantly contribute to the Word Error Rate. Our analysis revealed biases related to attributes such as native language, location, occupation, and birthplace. Most systems did not exhibit bias related to factors like gender and age. Additionally, we conducted an experiment to detect bias related to "variant" (accent and dialect) by combining the CORAAL (African American Vernacular English (AAVE)) and SBCSAE (General American English (GAE)) datasets, aiming to identify the sources of any observed bias. We found that both speaker variability and dialectal difference contribute to observed bias for variant. Maliha Jahan, Priyam Mazumdar, Thomas Thebaud, Mark Hasegawa-Johnson, Jesús Villalba 0001, Najim Dehak, Laureano Moro-Velázquez |
ICASSP | 7 |
| 2025 | Detecting Neurodegenerative Diseases using Frame-Level Handwriting EmbeddingsabstractIn this study, we explored the use of spectrograms to represent handwriting signals for assessing neurodegenerative diseases, including 42 healthy controls (CTL), 35 subjects with Parkinson’s Disease (PD), 21 with Alzheimer’s Disease (AD), and 15 with Parkinson’s Disease Mimics (PDM). We applied CNN and CNN-BLSTM models for binary classification using both multi-channel fixed-size and frame-based spectrograms. Our results showed that handwriting tasks and spectrogram channel combinations significantly impacted classification performance. The highest F1-score (89.8%) was achieved for AD vs. CTL, while PD vs. CTL reached 74.5%, and PD vs. PDM scored 77.97%. CNN consistently outperformed CNN-BLSTM. Different sliding window lengths were tested for constructing frame-based spectrograms. A 1-second window worked best for AD, longer windows improved PD classification, and window length had little effect on PD vs. PDM. Sarah Laouedj, Jesús Villalba 0001, Thomas Thebaud, Laureano Moro-Velázquez, Najim Dehak |
ICASSP | 5 |
| 2025 | Impact of Temporal Precision on Speech Synthesis Accuracy From Electrocorticographic Brain SignalsabstractAccurately time-aligned spectral targets are essential for training electrocorticographic (ECoG) brain-computer interfaces (BCIs) intended for real-time speech output. This alignment is particularly challenging with "silent speech," in which speech occurs without phonation, or in the extreme, without articulation. Crafting suitably precise targets for silent speech is complex and error-prone due to multiple sources of temporal imprecision. We investigated how these temporal inaccuracies impact deep neural network performance in synthesizing speech from a BCI clinical trial participant who retained some speech capability. By simulating silent speech conditions through distortions in known acoustic target timings, we observed significant performance degradations at both phonetic and syllabic timescales. Accuracy-based measures offered a more reliable assessment of intelligibility than traditional metrics like the short-time objective intelligibility index, which may overestimate performance in low-precision contexts. These results underscore the need for advanced alignment techniques with precise phonetic and syllabic guarantees for silent speech BCIs focused on providing immediate output. Qinwan Rabbani, Matthew S. Fifer, Nathan E. Crone, Laureano Moro-Velázquez |
ICASSP | 4 |
| 2025 | ADCeleb: A Longitudinal Speech Dataset from Public Figures for Early Detection of Alzheimer's Disease
Kunxiao Gao, Anna Favaro, Najim Dehak, Laureano Moro-Velázquez |
INTERSPEECH | 4 |
| 2025 | FaiST: A Benchmark Dataset for Fairness in Speech Technology
Maliha Jahan, Yinglun Sun, Priyam Mazumdar, Zsuzsanna Fagyal, Thomas Thebaud, Jesús Villalba 0001, Mark Hasegawa-Johnson, Najim Dehak, Laureano Moro-Velázquez |
INTERSPEECH | 9 |
| 2025 | The Interspeech 2025 Challenge on Speech Emotion Recognition in Naturalistic Conditions
Abinay Reddy Naini, Lucas Goncalves, Ali N. Salman, Pravin Mote, Ismail Rasim Ülgen, Thomas Thebaud, Laureano Moro-Velázquez, L. Paola García-Perera, Najim Dehak, Berrak Sisman, Carlos Busso |
INTERSPEECH | 7 |
| 2025 | Challenges and practical guidelines for atypical speech data collection, annotation, usage and sharing: A multi-project perspectiveabstractContains fulltext : 325867.pdf (Publisher’s version ) (Open Access) Zhengjun Yue, Mara Barberis, Tanvina Patel, Judith Dineley, Willemijn Doedens, Lottie Stipdonk, Elke De Witte, Erfan Loweimi, Hugo Van hamme, Djaina Satoer, Marina B. Ruiter, Laureano Moro-Velázquez, Nicholas Cummins, Odette Scharenborg |
INTERSPEECH | 13 |
| 2024 | Finding Spoken Identifications: Using GPT-4 Annotation for an Efficient and Fast Dataset Creation PipelineabstractThe growing emphasis on fairness in speech-processing tasks requires datasets with speakers from diverse subgroups that allow training and evaluating fair speech technology systems. However, creating such datasets through manual annotation can be costly. To address this challenge, we present a semi-automated dataset creation pipeline that leverages large language models. We use this pipeline to generate a dataset of speakers identifying themself or another speaker as belonging to a particular race, ethnicity, or national origin group. We use OpenaAI’s GPT-4 to perform two complex annotation tasks- separating files relevant to our intended dataset from the irrelevant ones (filtering) and finding and extracting information on identifications within a transcript (tagging). By evaluating GPT-4’s performance using human annotations as ground truths, we show that it can reduce resources required by dataset annotation while barely losing any important information. For the filtering task, GPT-4 had a very low miss rate of 6.93%. GPT-4’s tagging performance showed a trade-off between precision and recall, where the latter got as high as 97%, but precision never exceeded 45%. Our approach reduces the time required for the filtering and tagging tasks by 95% and 80%, respectively. We also present an in-depth error analysis of GPT-4’s performance. Maliha Jahan, Helin Wang, Thomas Thebaud, Yinglun Sun, Giang Ha Le, Zsuzsanna Fagyal, Odette Scharenborg, Mark Hasegawa-Johnson, Laureano Moro-Velázquez, Najim Dehak |
LREC/COLING | 9 |
| 2024 | Multimodal Emotion Recognition Harnessing the Complementarity of Speech, Language, and VisionabstractIn the realm of audiovisual emotion recognition, a significant challenge lies in developing neural network architectures capable of effectively harnessing and integrating multimodal information. This study introduces an advanced methodology for the Empathic Virtual Agent Challenge (EVAC), utilizing state-of-the-art speech, language, and image models. Specifically, we leverage cutting-edge pre-trained models, including multilingual variants fine-tuned in French for each modality, and integrate them using late fusion techniques. Through extensive experimentation and validation, we demonstrate the efficacy of our approach in achieving competitive results on the challenge dataset. Our findings highlight that multimodal approaches outperform unimodal methods across Core Affect Presence and Intensity and Appraisal Dimensions tasks, underscoring the effectiveness of integrating diverse modalities. This underscores the importance of leveraging multiple sources of information to capture nuanced emotional states more accurately and robustly in real-world applications. Thomas Thebaud, Anna Favaro, Yaohan Guan, Prabhav Singh, Jesús Villalba 0001, Laureano Moro-Velázquez, Najim Dehak |
ICMI | 7 |
| 2024 | Leveraging Universal Speech Representations for Detecting and Assessing the Severity of Mild Cognitive Impairment Across Languages
Anna Favaro, Tianyu Cao 0003, Najim Dehak, Laureano Moro-Velázquez |
INTERSPEECH | 4 |
| 2024 | Noise-robust Speech Separation with Fast Generative Correction
Helin Wang, Jesús Villalba 0001, Laureano Moro-Velázquez, Jiarui Hai, Thomas Thebaud, Najim Dehak |
INTERSPEECH | 3 |
| 2024 | Exploring the Complementary Nature of Speech and Eye Movements for Profiling Neurological Disorders
Anna Favaro, Thomas Thebaud, Jesús Villalba 0001, Najim Dehak, Laureano Moro-Velázquez |
INTERSPEECH | 6 |
| 2024 | CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech ProcessingabstractWe introduce Condition-Aware Self-Supervised Learning Representation (CA-SSLR), a generalist conditioning model broadly applicable to various speech-processing tasks. Compared to standard fine-tuning methods that optimize for downstream models, CA-SSLR integrates language and speaker embeddings from earlier layers, making the SSL model aware of the current language and speaker context.
This approach reduces the reliance on the input audio features while preserving the integrity of the base SSLR. CA-SSLR improves the model’s capabilities and demonstrates its generality on unseen tasks with minimal task-specific tuning. Our method employs linear modulation to dynamically adjust internal representations, enabling fine-grained adaptability without significantly altering the original model behavior. Experiments show that CA-SSLR reduces the number of trainable parameters, mitigates overfitting, and excels in under-resourced and unseen tasks. Specifically, CA-SSLR achieves a 10\% relative reduction in LID errors, a 37\% improvement in ASR CER on the ML-SUPERB benchmark, and a 27\% decrease in SV EER on VoxCeleb-1, demonstrating its effectiveness. Yen-Ju Lu, Thomas Thebaud, Laureano Moro-Velázquez, Ariya Rastrow, Najim Dehak, Jesús Villalba 0001 |
NeurIPS | 4 |
| 2024 | Slowness Regularized Contrastive Predictive Coding for Acoustic Unit DiscoveryabstractSelf-supervised methods such as Contrastive predictive Coding (CPC) have greatly improved the quality of the unsupervised representations. These representations significantly reduce the amount of labeled data needed for downstream task performance, such as automatic speech recognition. CPC learns representations by learning to predict future frames given current frames. Based on the observation that the acoustic information, e.g., phones, changes slower than the feature extraction rate in CPC, we propose regularization techniques that impose slowness constraints on the features. Here we propose two regularization techniques: Self-expressing constraint and Left-or-Right regularization. We evaluate the proposed model on ABX and linear phone classification tasks, acoustic unit discovery, and automatic speech recognition. The regularized CPC trained on 100 hours of unlabeled data matches the performance of the baseline CPC trained on 360 hours of unlabeled data. We also show that our regularization techniques are complementary to data augmentation and can further boost the system's performance. In monolingual, cross-lingual, or multilingual settings, with/without data augmentation, regardless of the amount of data used for training, our regularized models outperformed the baseline CPC models on the ABX task. Saurabhchand Bhati, Jesús Villalba 0001, Piotr Zelasko, Laureano Moro-Velázquez, Najim Dehak |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | Time-Domain Speech Super-Resolution With GAN Based Modeling for Telephony Speaker VerificationabstractAutomatic Speaker Verification(ASV) technology has become commonplace in virtual assistants. However, its performance suffers when there is a mismatch between the train and test domains. Mixed bandwidth training, i.e., pooling training data from both domains, is a preferred choice for developing a universal model that works for both narrowband and wideband domains. We propose complementing this technique by performing neural upsampling of narrowband signals, also known as bandwidth extension. We aim to discover and analyze high-performing time-domain Generative Adversarial Network (GAN) based models to improve our downstream state-of-the-art ASV system. We choose GANs since they 1) are powerful for learning conditional distribution and 2) allow flexibleplug-inusage as a pre-processor during the training of downstream tasks (ASV) with data augmentation. Prior work mainly focused on feature-domain bandwidth extension and limited experimental setups. We address these limitations by 1) using time-domain extension models, 2) reporting results on three real test sets, 3) extending training data, and 4) devising new test-time schemes. We compare supervised (conditional GAN) and unsupervised GANs (CycleGAN) and demonstrate an average relative improvement in the equal error rate of 8.6% and 7.7%, respectively. For further analysis, we study changes in the visual quality of the spectrogram, audio perceptual quality, t-SNE embeddings, and ASV score distributions. We show that our bandwidth extension leads to phenomena such as a shift of telephone (test) embeddings towards wideband (train) signals, a negative correlation of perceptual quality with downstream performance, and condition-independent score calibration. Saurabh Kataria 0001, Jesús Villalba 0001, Laureano Moro-Velázquez, Piotr Zelasko, Najim Dehak |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Model-Based Fairness Metric for Speaker VerificationabstractEnsuring that technological advancements benefit all groups of people equally is crucial. The first step towards fairness is identifying existing inequalities. The naive comparison of group error rates may lead to wrong conclusions. We introduce a new method to determine whether a speaker verification system is fair toward several population subgroups. We propose to model miss and false alarm probabilities as a function of multiple factors, including the population group effects, e.g., male and female, and a series of confounding variables, e.g., speaker effects, language, nationality, etc. This model can estimate error rates related to a group effect without the influence of confounding effects. We experiment with a synthetic dataset where we control group and confounding effects. Our metric achieves significantly lower false positive and false negative rates w.r.t. baseline. We also experiment with VoxCeleb and NIST SRE21 datasets on different ASV systems and present our conclusions. Maliha Jahan, Laureano Moro-Velázquez, Thomas Thebaud, Najim Dehak, Jesús Villalba 0001 |
ASRU | 2 |
| 2023 | Segmental SpeechCLIP: Utilizing Pretrained Image-text Models for Audio-Visual Learning
Saurabhchand Bhati, Jesús Villalba 0001, Laureano Moro-Velázquez, Thomas Thebaud, Najim Dehak |
INTERSPEECH | 3 |
| 2023 | Do Phonatory Features Display Robustness to Characterize Parkinsonian Speech Across Corpora?
Anna Favaro, Tianyu Cao 0003, Thomas Thebaud, Jesús Villalba 0001, Ankur A. Butala, Najim Dehak, Laureano Moro-Velázquez |
INTERSPEECH | 7 |
| 2023 | Self-FiLM: Conditioning GANs with self-supervised representations for bandwidth extension based speaker recognitionabstractSpeech super-resolution/Bandwidth Extension (BWE) can improve downstream tasks like Automatic Speaker Verification (ASV).We introduce a simple novel technique called Self-FiLM to inject self-supervision into existing BWE models via Feature-wise Linear Modulation.We hypothesize that such information captures domain/environment information, which can give zero-shot generalization.Self-FiLM Conditional GAN (CGAN) gives 18% relative improvement in Equal Error Rate and 8.5% in minimum Decision Cost Function using state-ofthe-art ASV system on SRE21 test.We further by 1) deep feature loss from time-domain models and 2) re-training of data2vec 2.0 models on naturalistic wideband (VoxCeleb) and telephone data (SRE Superset etc.).Lastly, we integrate selfsupervision with CycleGAN to present a completely unsupervised solution that matches the semi-supervised performance. Saurabh Kataria 0001, Jesús Villalba 0001, Laureano Moro-Velázquez, Thomas Thebaud, Najim Dehak |
INTERSPEECH | 3 |
| 2023 | DuTa-VC: A Duration-aware Typical-to-atypical Voice Conversion Approach with Diffusion Probabilistic Model
Helin Wang, Thomas Thebaud, Jesús Villalba 0001, Myra Sydnor, Becky Lammers, Najim Dehak, Laureano Moro-Velázquez |
INTERSPEECH | 7 |
| 2022 | Non-contrastive self-supervised learning of utterance-level speech representations
Raghavendra Pappagari, Piotr Zelasko, Laureano Moro-Velázquez, Jesús Villalba 0001, Najim Dehak |
INTERSPEECH | 4 |
| 2022 | Joint domain adaptation and speech bandwidth extension using time-domain GANs for speaker verificationabstractSpeech systems developed for a particular choice of acoustic domain and sampling frequency do not translate easily to others.The usual practice is to learn domain adaptation and bandwidth extension models independently.Contrary to this, we propose to learn both tasks together.Particularly, we learn to map narrowband conversational telephone speech to wideband microphone speech.We developed parallel and non-parallel learning solutions which utilize both paired and unpaired data.First, we first discuss joint and disjoint training of multiple generative models for our tasks.Then, we propose a two-stage learning solution where we use a pre-trained domain adaptation system for pre-processing in bandwidth extension training.We evaluated our schemes on a Speaker Verification downstream task.We used the JHU-MIT experimental setup for NIST SRE21, which comprises SRE16, SRE-CTS Superset and SRE21.Our results provide the first evidence that learning both tasks is better than learning just one.On SRE16, our best system achieves 22% relative improvement in Equal Error Rate w.r.t. a direct learning baseline and 8% w.r.t. a strong bandwidth expansion system. Saurabh Kataria 0001, Jesús Villalba 0001, Laureano Moro-Velázquez, Najim Dehak |
INTERSPEECH | 3 |
| 2022 | Vsameter: Evaluation of a New Open-Source Tool to Measure Vowel Space Area and Related MetricsabstractVowel space area (VSA) is an applicable metric for studying speech production deficits and intelligibility. Previous works suggest that the VSA accounts for almost 50% of the intelligibility variance, being an essential component of global intelligibility estimates. However, almost no study publishes a tool to estimate VSA automatically with publicly available codes. In this paper, we propose an open-source tool called VSAmeter to measure VSA and vowel articulation index (VAI) automatically and validate it with the VSA and VAI obtained from a dataset in which the formants and phone segments have been annotated manually. The results show that VSA and VAI values obtained by our proposed method strongly correlate with those generated by manually extracted F1 and F2 and alignments. Such a method can be utilized in speech applications, e.g., the automatic measurement of VAI for the evaluation of speakers with dysarthria. Tianyu Cao 0003, Laureano Moro-Velázquez, Piotr Zelasko, Jesús Villalba 0001, Najim Dehak |
SLT | 2 |
| 2022 | A Multi-Modal Array of Interpretable Features to Evaluate Language and Speech Patterns in Different Neurological DisordersabstractSpeech-based automatic approaches for evaluating neurological disorders (NDs) depend on feature extraction before the classification pipeline. It is preferable for these features to be interpretable to facilitate their development as diagnostic tools. This study focuses on the analysis of interpretable features obtained from the spoken responses of 88 subjects with NDs and controls (CN). Subjects with NDs have Alzheimer's disease (AD), Parkinson's disease (PD), or Parkinson's disease mimics (PDM). We configured three complementary sets of features related to cognition, speech, and language, and conducted a statistical analysis to examine which features differed between NDs and CN. Results suggested that features capturing response informativeness, reaction times, vocabulary richness, and syntactic complexity provided separability between AD and CN. Similarly, fundamental frequency variability helped differentiate PD from CN, while the number of salient informational units PDM from CN. Anna Favaro, Chelsie Motley, Tianyu Cao 0003, Miguel Iglesias, Ankur A. Butala, Esther S. Oh, Robert D. Stevens 0002, Jesús Villalba 0001, Najim Dehak, Laureano Moro-Velázquez |
SLT | 10 |
| 2022 | Discovering phonetic inventories with crosslingual automatic speech recognition
Piotr Zelasko, Siyuan Feng 0001, Laureano Moro-Velázquez, Ali Abavisani, Saurabhchand Bhati, Odette Scharenborg, Mark Hasegawa-Johnson, Najim Dehak |
Comput. Speech Lang. | 3 |
| 2022 | Unsupervised Speech Segmentation and Variable Rate Representation Learning Using Segmental Contrastive Predictive CodingabstractTypically, unsupervised segmentation of speech into the phone- and word-like units are treated as separate tasks and are often done via different methods which do not fully leverage the inter-dependence of the two tasks. Here, we unify them and propose a technique that can jointly perform both, showing that these two tasks indeed benefit from each other. Recent attempts employ self-supervised learning, such as contrastive predictive coding (CPC), where the next frame is predicted given past context. However, CPC only looks at the audio signal’s frame-level structure. We overcome this limitation with a segmental contrastive predictive coding (SCPC) framework to model the signal structure at a higher level, e.g., phone level. A convolutional neural network learns frame-level representation from the raw waveform via noise-contrastive estimation (NCE). A differentiable boundary detector finds variable-length segments, which are then used to optimize a segment encoder via NCE to learn segment representations. The differentiable boundary detector allows us to train frame-level and segment-level encoders jointly. Experiments show that our single model outperforms existing phone and word segmentation methods on TIMIT and Buckeye datasets. We analyze the impact of the threshold on boundary detector performance, and our results suggest that automatically learning the boundary threshold can be as effective as manually tuning that threshold. We discover that phone class impacts the boundary detection performance, and the boundaries between successive vowels or semivowels are the most difficult. Finally, we use SCPC to extract speech features at the segment level rather than at the uniformly spaced frame level (e.g., 10 ms) and produce variable rate representations that change according to the contents of the utterance. We can lower the feature extraction rate from the typical 100 Hz to as low as 14.5 Hz on average while still outperforming the hand-crafted features such as MFCC on the linear phone classification task. Saurabhchand Bhati, Jesús Villalba 0001, Piotr Zelasko, Laureano Moro-Velázquez, Najim Dehak |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Beyond Isolated Utterances: Conversational Emotion RecognitionabstractSpeech emotion recognition is the task of recognizing the speaker's emotional state given a recording of their utterance. While most of the current approaches focus on inferring emotion from isolated utterances, we argue that this is not sufficient to achieve conversational emotion recognition (CER) which deals with recognizing emotions in conversations. In this work, we propose several approaches for CER by treating it as a sequence labeling task. We investigated transformer architecture for CER and, compared it with ResNet-34 and BiLSTM architectures in both contextual and contextless scenarios using IEMOCAP corpus. Based on the inner workings of the self-attention mechanism, we proposed DiverseCatAugment (DCA), an augmentation scheme, which improved the transformer model performance by an absolute 3.3% micro-f1 on conversations and 3.6% on isolated utterances. We further enhanced the performance by introducing an interlocutor-aware transformer model where we learn a dictionary of interlocutor index embeddings to exploit diarized conversations. Raghavendra Pappagari, Piotr Zelasko, Jesús Villalba 0001, Laureano Moro-Velázquez, Najim Dehak |
ASRU | 4 |
| 2021 | How Phonotactics Affect Multilingual and Zero-Shot ASR PerformanceabstractThe idea of combining multiple languages’ recordings to train a single automatic speech recognition (ASR) model brings the promise of the emergence of universal speech representation. Recently, a Transformer encoder-decoder model has been shown to leverage multilingual data well in IPA transcriptions of languages presented during training. However, the representations it learned were not successful in zero-shot transfer to unseen languages. Because that model lacks an explicit factorization of the acoustic model (AM) and language model (LM), it is unclear to what degree the performance suffered from differences in pronunciation or the mismatch in phono-tactics. To gain more insight into the factors limiting zero-shot ASR transfer, we replace the encoder-decoder with a hybrid ASR system consisting of a separate AM and LM. Then, we perform an extensive evaluation of monolingual, multilingual, and crosslingual (zero-shot) acoustic and language models on a set of 13 phonetically diverse languages. We show that the gain from modeling crosslingual phonotactics is limited, and imposing a too strong model can hurt the zero-shot transfer. Furthermore, we find that a multilingual LM hurts a multilingual ASR system’s performance, and retaining only the target language’s phonotactic data in LM training is preferable. Siyuan Feng 0001, Piotr Zelasko, Laureano Moro-Velázquez, Ali Abavisani, Mark Hasegawa-Johnson, Odette Scharenborg, Najim Dehak |
ICASSP | 3 |
| 2021 | CopyPaste: An Augmentation Method for Speech Emotion RecognitionabstractData augmentation is a widely used strategy for training robust machine learning models. It partially alleviates the problem of limited data for tasks like speech emotion recognition (SER), where collecting data is expensive and challenging. This study proposes CopyPaste, a perceptually motivated novel augmentation procedure for SER. Assuming that the presence of emotions other than neutral dictates a speaker’s overall perceived emotion in a recording, concatenation of an emotional (emotion E) and a neutral utterance can still be labeled with emotion E. We hypothesize that SER performance can be improved using these concatenated utterances in model training. To verify this, three CopyPaste schemes are tested on two deep learning models: one trained independently and another using transfer learning from an x-vector model, a speaker recognition model. We observed that all three CopyPaste schemes improve SER performance on all the three datasets considered: MSP-Podcast, Crema-D, and IEMOCAP. Additionally, CopyPaste performs better than noise augmentation and, using them together improves the SER performance further. Our experiments on noisy test sets suggested that CopyPaste is effective even in noisy test conditions. Raghavendra Pappagari, Jesús Villalba 0001, Piotr Zelasko, Laureano Moro-Velázquez, Najim Dehak |
ICASSP | 4 |
| 2021 | Segmental Contrastive Predictive Coding for Unsupervised Word SegmentationabstractAutomatic detection of phoneme or word-like units is one of the core objectives in zero-resource speech processing. Recent attempts employ self-supervised training methods, such as contrastive predictive coding (CPC), where the next frame is predicted given past context. However, CPC only looks at the audio signal's frame-level structure. We overcome this limitation with a segmental contrastive predictive coding (SCPC) framework that can model the signal structure at a higher level e.g. at the phoneme level. In this framework, a convolutional neural network learns frame-level representation from the raw waveform via noise-contrastive estimation (NCE). A differentiable boundary detector finds variable-length segments, which are then used to optimize a segment encoder via NCE to learn segment representations. The differentiable boundary detector allows us to train frame-level and segment-level encoders jointly. Typically, phoneme and word segmentation are treated as separate tasks. We unify them and experimentally show that our single model outperforms existing phoneme and word segmentation methods on TIMIT and Buckeye datasets. We analyze the impact of boundary threshold and when is the right time to include the segmental loss in the learning process. Saurabhchand Bhati, Jesús Villalba 0001, Piotr Zelasko, Laureano Moro-Velázquez, Najim Dehak |
Interspeech | 4 |
| 2021 | Align-Denoise: Single-Pass Non-Autoregressive Speech Recognition
Nanxin Chen, Piotr Zelasko, Laureano Moro-Velázquez, Jesús Villalba 0001, Najim Dehak |
Interspeech | 3 |
| 2021 | Unsupervised Acoustic Unit Discovery by Leveraging a Language-Independent Subword Discriminative Feature RepresentationabstractThis paper tackles automatically discovering phone-like acoustic units (AUD) from unlabeled speech data. Past studies usually proposed single-step approaches. We propose a two-stage approach: the first stage learns a subword-discriminative feature representation and the second stage applies clustering to the learned representation and obtains phone-like clusters as the discovered acoustic units. In the first stage, a recently proposed method in the task of unsupervised subword modeling is improved by replacing a monolingual out-of-domain (OOD) ASR system with a multilingual one to create a subword-discriminative representation that is more language-independent. In the second stage, segment-level k-means is adopted, and two methods to represent the variable-length speech segments as fixed-dimension feature vectors are compared. Experiments on a very low-resource Mboshi language corpus show that our approach outperforms state-of-the-art AUD in both normalized mutual information (NMI) and F-score. The multilingual ASR improved upon the monolingual ASR in providing OOD phone labels and in estimating the phone boundaries. A comparison of our systems with and without knowing the ground-truth phone boundaries showed a 16% NMI performance gap, suggesting that the current approach can significantly benefit from improved phone boundary estimation. Siyuan Feng 0001, Piotr Zelasko, Laureano Moro-Velázquez, Odette Scharenborg |
Interspeech | 3 |
| 2021 | Deep Feature CycleGANs: Speaker Identity Preserving Non-Parallel Microphone-Telephone Domain Adaptation for Speaker VerificationabstractWith the increase in the availability of speech from varied domains, it is imperative to use such out-of-domain data to improve existing speech systems. Domain adaptation is a prominent pre-processing approach for this. We investigate it for adapt microphone speech to the telephone domain. Specifically, we explore CycleGAN-based unpaired translation of microphone data to improve the x-vector/speaker embedding network for Telephony Speaker Verification. We first demonstrate the efficacy of this on real challenging data and then, to improve further, we modify the CycleGAN formulation to make the adaptation task-specific. We modify CycleGAN's identity loss, cycle-consistency loss, and adversarial loss to operate in the deep feature space. Deep features of a signal are extracted from an auxiliary (speaker embedding) network and, hence, preserves speaker identity. Our 3D convolution-based Deep Feature Discriminators (DFD) show relative improvements of 5-10% in terms of equal error rate. To dive deeper, we study a challenging scenario of pooling (adapted) microphone and telephone data with data augmentations and telephone codecs. Finally, we highlight the sensitivity of CycleGAN hyper-parameters and introduce a parameter called probability of adaptation. Saurabh Kataria 0001, Jesús Villalba 0001, Piotr Zelasko, Laureano Moro-Velázquez, Najim Dehak |
Interspeech | 4 |
| 2021 | Automatic Detection and Assessment of Alzheimer Disease Using Speech and Language Technologies in Low-Resource Scenarios
Raghavendra Pappagari, Sonal Joshi, Laureano Moro-Velázquez, Piotr Zelasko, Jesús Villalba 0001, Najim Dehak |
Interspeech | 4 |
| 2021 | Study of Pre-Processing Defenses Against Adversarial Attacks on State-of-the-Art Speaker Recognition SystemsabstractAdversarial examples are designed to fool the speaker recognition (SR) system by adding a carefully crafted human-imperceptible noise to the speech signals. Posing a severe security threat to state-of-the-art SR systems, it becomes vital to deep-dive and study their vulnerabilities. Moreover, it is of greater importance to propose countermeasures that can protect the systems against these attacks. Addressing these concerns, we first investigated how state-of-the-art x-vector based SR systems are affected by white-box adversarial attacks, i.e., when the adversary has full knowledge of the system. x-Vector based SR systems are evaluated against white-box adversarial attacks common in the literature like fast gradient sign method (FGSM), basic iterative method (BIM)–a.k.a. iterative-FGSM–, projected gradient descent (PGD), and Carlini-Wagner (CW) attack. To mitigate against these attacks, we investigated four pre-processing defenses which do not need adversarial examples during training. The four pre-processing defenses–viz. randomized smoothing, DefenseGAN, variational autoencoder (VAE), and Parallel Wave-GAN vocoder (PWG) are compared against the baseline defense of adversarial training. Performing powerful adaptive white-box adversarial attack (i.e., when the adversary has full knowledge of the system, including the defense), our conclusions indicate that SR systems were extremely vulnerable under BIM, PGD, and CW attacks. Among the proposed pre-processing defenses, PWG combined with randomized smoothing offers the most protection against the attacks, with accuracy averaging 93% compared to 52% in the undefended system and an absolute improvement > 90% for BIM attacks with L∞ > 0.001 and CW attack. Sonal Joshi, Jesús Villalba 0001, Piotr Zelasko, Laureano Moro-Velázquez, Najim Dehak |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2020 | Using X-Vectors to Automatically Detect Parkinson's Disease from SpeechabstractThe promise of new neuroprotective treatments to stop or slow the advance of Parkinson's Disease (PD) urges for new biomarkers or detection schemes that can deliver a faster diagnosis. Given that speech is affected by PD, the combination of deep neural networks and speech processing can provide automatic detection schemes. Accordingly, in this study we analyze for the first time a new state-of-the-art speaker recognition technique, x-Vectors, in a different scenario: the automatic detection of PD from speech. The proposed approach is compared with another speaker recognition technique, i-Vectors, employed in previous works and used as baseline in this study. A corpus with 43 PD patients and 46 control speakers was used to evaluate the performance of these two techniques at two sampling frequencies: 8 and 16 kHz.The x-Vector approach provided the best results in terms of accuracy and AUC reaching values of 90% and 0.94, respectively. Consequently, results suggest that speaker embeddings obtained using deep neural networks are successful extracting acoustic information relative to patterns in articulation, prosody and/or phonation common in persons with PD. Laureano Moro-Velázquez, Jesús Villalba 0001, Najim Dehak |
ICASSP | 1 |
| 2020 | Using State of the Art Speaker Recognition and Natural Language Processing Technologies to Detect Alzheimer's Disease and Assess its Severity
Raghavendra Pappagari, Laureano Moro-Velázquez, Najim Dehak |
INTERSPEECH | 3 |
| 2020 | That Sounds Familiar: An Analysis of Phonetic Representations Transfer Across LanguagesabstractOnly a handful of the world's languages are abundant with the resources that enable practical applications of speech processing technologies. One of the methods to overcome this problem is to use the resources existing in other languages to train a multilingual automatic speech recognition (ASR) model, which, intuitively, should learn some universal phonetic representations. In this work, we focus on gaining a deeper understanding of how general these representations might be, and how individual phones are getting improved in a multilingual setting. To that end, we select a phonetically diverse set of languages, and perform a series of monolingual, multilingual and crosslingual (zero-shot) experiments. The ASR is trained to recognize the International Phonetic Alphabet (IPA) token sequences. We observe significant improvements across all languages in the multilingual setting, and stark degradation in the crosslingual setting, where the model, among other errors, considers Javanese as a tone language. Notably, as little as 10 hours of the target language training data tremendously reduces ASR error rates. Our analysis uncovered that even the phones that are unique to a single language can benefit greatly from adding training data from other languages - an encouraging result for the low-resource speech community. Piotr Zelasko, Laureano Moro-Velázquez, Mark Hasegawa-Johnson, Odette Scharenborg, Najim Dehak |
INTERSPEECH | 2 |
| 2020 | Analysis of the Effects of Supraglottal Tract Surgical Procedures in Automatic Speaker Recognition PerformanceabstractThis article evaluates the impact in the performance of state-of-the-art automatic speaker recognition schemes of three surgical procedures modifying the supraglottal tract structures of speakers. To do so, a new corpus (Cuco) was recorded, containing the speech of 107 speakers before and after surgery. Speakers were divided into four groups depending on the type of surgery: tonsillectomy, functional endoscopy sinus surgery (FESS), septoplasty, and controls. The analyzed speaker recognition schemes were i-vectors, i-vectors with supervised Universal Background Model, i-vectors employing Time-delay Deep Neural Networks and x-vectors. In all cases, probabilistic linear discriminant analysis was employed in the back-end. Results show changes in the speech of patients who underwent tonsillectomy or FESS after surgery in contrast to controls or patients who had a septoplasty, where not significant variations are observed. These changes increase the Equal Error Rate (EER) of the analyzed speaker recognition schemes for the septoplasty and FESS groups when employing enrollment data recorded before the surgery. Moreover, surgery has a similar influence in the speech of female and male speakers with respect to the analyzed schemes. In consequence, results suggest that it is advisable to update the speaker's enrollment speech after three months following supraglottal tract surgery to ensure that the effects of the operation and post-operative recovery period do not influence the performance of the automatic speaker recognition systems. Laureano Moro-Velázquez, Estefanía Hernández-García, Jorge Andrés Gómez García, Juan Ignacio Godino-Llorente, Najim Dehak |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Study of the Performance of Automatic Speech Recognition Systems in Speakers with Parkinson's DiseaseabstractParkinson’s Disease (PD) affects motor capabilities of patients, who in some cases need to use human-computer assistive technologies to regain independence. The objective of this work is to study in detail the differences in error patterns from state-of-the-art Automatic Speech Recognition (ASR) systems on speech from people with and without PD. Two different speech recognizers (attention-based end-to-end and Deep Neural Network - Hidden Markov Models hybrid systems) were trained on a Spanish language corpus and subsequently tested on speech from 43 speakers with PD and 46 without PD. The differences related to error rates, substitutions, insertions and deletions of characters and phonetic units between the two groups were analyzed, showing that the word error rate is 27% higher in speakers with PD than in control speakers, with a moderated correlation between that rate and the developmental stage of the disease. The errors were related to all manner classes, and were more pronounced in the vowel /u/. This study is the first to evaluate ASR systems’ responses to speech from patients at different stages of PD in Spanish. The analyses showed general trends but individual speech deficits must be studied in the future when designing new ASR systems for this population. Laureano Moro-Velázquez, Shinji Watanabe 0001, Mark Hasegawa-Johnson, Odette Scharenborg, Najim Dehak |
INTERSPEECH | 1 |
| 2019 | Emulating the perceptual capabilities of a human evaluator to map the GRB scale for the assessment of voice disorders
Jorge Andrés Gómez García, Laureano Moro-Velázquez, Janaina Mendes-Laureano, Germán Castellanos-Domínguez, Juan Ignacio Godino-Llorente |
Eng. Appl. Artif. Intell. | 2 |
| 2018 | ByoVoz Automatic Voice Condition Analysis System for the 2018 FEMH ChallengeabstractThis paper presents the methods and results used by the ByoVoz team for the design of an automatic voice condition analysis system, which was submitted to the 2018 Far East Memorial Hospital voice data challenge. The proposed methodology is based on a cascading scheme that firstly discriminates between pathological and normophonic voices, and then identifies the type of disorder. By using diverse feature selection techniques, a subset of complexity, spectral/cepstral and perturbation characteristics were identified for the proposed tasks. Then, several generative classification methodologies based on Gaussian Mixture Models and Gradient Boosting were employed to provide decisions about the input voices in the binary classification, and using one-vs-one classification systems based on Random Forests for the categorization according to the type of disorder. By using a 4-folds cross-validation approach on the training partition a sensitivity=0.93 and specificity=0.74 was obtained. Similarly, an unweighted average recall of 0.63 and an accuracy of 66% was obtained for the identification task. Using the scoring metric proposed in the challenge the final resulting score is of 0.77. When the testing partition was considered, a sensitivity of 0.92, specificity of 0.54, unweighted average recall of 0.61 and a score of 0.72 were obtained. This was the fifth best result in the competition. Julián D. Arias-Londoño, Jorge Andrés Gómez García, Laureano Moro-Velázquez, Juan Ignacio Godino-Llorente |
IEEE BigData | 3 |
| 2015 | Automatic age detection in normal and pathological voiceabstractSystems that automatically detect voice pathologies are usually trained with recordings belonging to population of all ages. However such an approach might be inadequate because of the acoustic variations in the voice caused by the natural aging process. In top of that, elder voices present some perturbations in quality similar to those related to voice disorders, which make the detection of pathologies more troublesome. With this in mind, the study of methodologies which automatically incorporate information about speakers’ age, aiming at a simplification in the detection of voice disorders is of interest. In this respect, the present paper introduces an age detector trained with normal and pathological voice, constituting a first step towards the study of age-dependent pathology detectors. The proposed system employs sustained vowels of the Saarbrucken database from which two age groups are examinated: adults and elders. Mel frequency cepstral coefficients for characterization, and Gaussian mixture models for classification are utilized. In addition, fusion of vowels at score level is considered to improve detection performance. Results suggest that age might be effectively recognized using normal and pathological voices when using sustained vowels as acoustical material, opening up possibilities for the design of automatic age-dependent voice pathology detection systems. Jorge Andrés Gómez García, Laureano Moro-Velázquez, Juan Ignacio Godino-Llorente, Germán Castellanos-Domínguez |
INTERSPEECH | 2 |