EDBT 2026 Demo / reviewers in the wild / expert
Ramón Fernandez Astudillo
dblp:56/7987
· DBLP profile ↗
45ranked-venue papers
14as first author
14since 2021 · last 2025
0000-0001-9443-2346ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 10 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 11 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Latent Principle Discovery for Language Model Self-ImprovementabstractWhen language model (LM) users aim to improve the quality of its generations, it is crucial to specify concrete behavioral attributes that the model should strive to reflect. However, curating such principles across many domains, even non-exhaustively, requires a labor-intensive annotation process. To automate this process, we propose eliciting these latent attributes that guide model reasoning toward human-preferred responses by explicitly modeling them in a self-correction setting. Our approach mines new principles from the LM itself and compresses the discovered elements to an interpretable set via clustering. Specifically, we employ a form of posterior-regularized Monte Carlo Expectation-Maximization to both identify a condensed set of the most effective latent principles and teach the LM to strategically invoke them in order to intrinsically refine its responses. We demonstrate that bootstrapping our algorithm over multiple iterations enables smaller language models (7-8B parameters) to self-improve, achieving +8-10\% in AlpacaEval win-rate, an average of +0.3 on MT-Bench, and +19-23\% in principle-following win-rate on IFEval. We also show that clustering the principles yields interpretable and diverse model-generated constitutions while retaining model performance. The gains that our method achieves highlight the potential of automated, principle-driven post-training recipes toward continual self-improvement. Keshav Ramji, Tahira Naseem, Ramón Fernandez Astudillo |
NeurIPS | 3 |
| 2024 | BRAIn: Bayesian Reward-conditioned Amortized Inference for natural language generation from feedbackabstractDistribution matching methods for language model alignment such as Generation with Distributional Control (GDC) and Distributional Policy Gradient (DPG) have not received the same level of attention in reinforcement learning from human feedback (RLHF) as contrastive methods such as Sequence Likelihood Calibration (SLiC), Direct Preference Optimization (DPO) and its variants. We identify high variance of the gradient estimate as the primary reason for the lack of success of these methods and propose a self-normalized baseline to reduce the variance. We further generalize the target distribution in DPG, GDC and DPO by using Bayes' rule to define the reward-conditioned posterior. The resulting approach, referred to as BRAIn - Bayesian Reward-conditioned Amortized Inference acts as a bridge between distribution matching methods and DPO and significantly outperforms prior art in summarization and Antropic HH tasks. Gaurav Pandey 0001, Yatin Nandwani, Tahira Naseem, Guangxuan Xu, Dinesh Raghu, Sachindra Joshi, Asim Munawar, Ramón Fernandez Astudillo |
ICML | 9 |
| 2023 | Laziness Is a Virtue When It Comes to Compositionality in Neural Semantic ParsingabstractMaxwell Crouse, Pavan Kapanipathi, Subhajit Chaudhury, Tahira Naseem, Ramon Fernandez Astudillo, Achille Fokoue, Tim Klinger. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Maxwell Crouse, Pavan Kapanipathi, Subhajit Chaudhury, Tahira Naseem, Ramón Fernandez Astudillo, Achille Fokoue, Tim Klinger |
ACL (1) | 5 |
| 2023 | Alignment via Mutual InformationabstractMany language learning tasks require learners to infer correspondences between data in two modalities.Often, these alignments are manyto-many and context-sensitive.For example, translating into morphologically rich languages requires learning not just how words, but morphemes, should be translated; words and morphemes may have different meanings (or groundings) depending on the context in which they are used.We describe an informationtheoretic approach to context-sensitive, manyto-many alignment.Our approach first trains a masked sequence model to place distributions over missing spans in (source, target) sequences.Next, it uses this model to compute pointwise mutual information between source and target spans conditional on context.Finally, it aligns spans with high mutual information.We apply this approach to two learning problems: character-based word translation (using alignments for joint morphological segmentation and lexicon learning) and visually grounded reference resolution (using alignments to jointly localize referents and learn word meanings).In both cases, our proposed approach outperforms both structured and neural baselines, showing that conditional mutual information offers an effective framework for formalizing alignment problems in general domains. Shinjini Ghosh, Ramón Fernandez Astudillo, Tahira Naseem, Jacob Andreas |
CoNLL | 3 |
| 2022 | X-FACTOR: A Cross-metric Evaluation of Factual Correctness in Abstractive SummarizationabstractSubhajit Chaudhury, Sarathkrishna Swaminathan, Chulaka Gunasekara, Maxwell Crouse, Srinivas Ravishankar, Daiki Kimura, Keerthiram Murugesan, Ramón Fernandez Astudillo, Tahira Naseem, Pavan Kapanipathi, Alexander Gray. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Subhajit Chaudhury, Sarathkrishna Swaminathan, R. Chulaka Gunasekara, Maxwell Crouse, Srinivas Ravishankar, Daiki Kimura, Keerthiram Murugesan, Ramón Fernandez Astudillo, Tahira Naseem, Pavan Kapanipathi, Alexander G. Gray |
EMNLP | 8 |
| 2022 | Inducing and Using Alignments for Transition-based AMR ParsingabstractAndrew Drozdov, Jiawei Zhou, Radu Florian, Andrew McCallum, Tahira Naseem, Yoon Kim, Ramón Astudillo. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Andrew Drozdov, Jiawei Zhou 0001, Radu Florian, Andrew McCallum, Tahira Naseem, Ramón Fernandez Astudillo |
NAACL-HLT | 7 |
| 2022 | Maximum Bayes Smatch Ensemble Distillation for AMR ParsingabstractYoung-Suk Lee, Ramón Astudillo, Hoang Thanh Lam, Tahira Naseem, Radu Florian, Salim Roukos. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Young-Suk Lee 0001, Ramón Fernandez Astudillo, Hoang Thanh Lam, Tahira Naseem, Radu Florian, Salim Roukos |
NAACL-HLT | 2 |
| 2022 | DocAMR: Multi-Sentence AMR Representation and EvaluationabstractTahira Naseem, Austin Blodgett, Sadhana Kumaravel, Tim O’Gorman, Young-Suk Lee, Jeffrey Flanigan, Ramón Astudillo, Radu Florian, Salim Roukos, Nathan Schneider. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Tahira Naseem, Austin Blodgett, Sadhana Kumaravel, Tim O'Gorman, Young-Suk Lee 0001, Jeffrey Flanigan, Ramón Fernandez Astudillo, Radu Florian, Salim Roukos, Nathan Schneider 0001 |
NAACL-HLT | 7 |
| 2021 | Structural Guidance for Transformer Language ModelsabstractPeng Qian, Tahira Naseem, Roger Levy, Ramón Fernandez Astudillo. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Tahira Naseem, Roger Levy, Ramón Fernandez Astudillo |
ACL/IJCNLP (1) | 4 |
| 2021 | Bootstrapping Multilingual AMR with Contextual Word AlignmentsabstractJanaki Sheth, Young-Suk Lee, Ramón Fernandez Astudillo, Tahira Naseem, Radu Florian, Salim Roukos, Todd Ward. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Janaki Sheth, Young-Suk Lee 0001, Ramón Fernandez Astudillo, Tahira Naseem, Radu Florian, Salim Roukos, Todd Ward |
EACL | 3 |
| 2021 | Structure-aware Fine-tuning of Sequence-to-sequence Transformers for Transition-based AMR ParsingabstractPredicting linearized Abstract Meaning Representation (AMR) graphs using pre-trained sequence-to-sequence Transformer models has recently led to large improvements on AMR parsing benchmarks.These parsers are simple and avoid explicit modeling of structure but lack desirable properties such as graph well-formedness guarantees or built-in graph-sentence alignments.In this work we explore the integration of general pre-trained sequence-to-sequence language models and a structure-aware transition-based approach.We depart from a pointer-based transition system and propose a simplified transition set, designed to better exploit pre-trained language models for structured fine-tuning.We also explore modeling the parser state within the pre-trained encoder-decoder architecture and different vocabulary strategies for the same purpose.We provide a detailed comparison with recent progress in AMR parsing and show that the proposed parser retains the desirable properties of previous transition-based approaches, while being simpler and reaching the new parsing state of the art for AMR 2.0, without the need for graph re-categorization. Jiawei Zhou 0001, Tahira Naseem, Ramón Fernandez Astudillo, Young-Suk Lee 0001, Radu Florian, Salim Roukos |
EMNLP (1) | 3 |
| 2021 | Eat: Enhanced ASR-TTS for Self-Supervised Speech RecognitionabstractSelf-supervised ASR-TTS models suffer in out-of-domain data conditions. Here we propose an enhanced ASR-TTS (EAT) model that incorporates two main features: 1) The ASR→TTS direction is equipped with a language model reward to penalize the ASR hypotheses before forwarding it to TTS. 2) In the TTS→ASR direction, a hyper-parameter is introduced to scale the attention context from synthesized speech before sending it to ASR to handle out-of-domain data. Training strategies and the effectiveness of the EAT model are explored under out-of-domain data conditions. The results show that EAT reduces the performance gap between supervised and self-supervised training significantly by absolute 2.6% and 2.7% on Librispeech and BABEL respectively. Murali Karthick Baskar, Lukás Burget, Shinji Watanabe 0001, Ramón Fernandez Astudillo, Jan Cernocký |
ICASSP | 4 |
| 2021 | AMR Parsing with Action-Pointer TransformerabstractJiawei Zhou, Tahira Naseem, Ramón Fernandez Astudillo, Radu Florian. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Jiawei Zhou 0001, Tahira Naseem, Ramón Fernandez Astudillo, Radu Florian |
NAACL-HLT | 3 |
| 2021 | Ensembling Graph Predictions for AMR ParsingabstractIn many machine learning tasks, models are trained to predict structure data such as graphs. For example, in natural language processing, it is very common to parse texts into dependency trees or abstract meaning representation (AMR) graphs. On the other hand, ensemble methods combine predictions from multiple models to create a new one that is more robust and accurate than individual predictions. In the literature, there are many ensembling techniques proposed for classification or regression problems, however, ensemble graph prediction has not been studied thoroughly. In this work, we formalize this problem as mining the largest graph that is the most supported by a collection of graph predictions. As the problem is NP-Hard, we propose an efficient heuristic algorithm to approximate the optimal solution. To validate our approach, we carried out experiments in AMR parsing problems. The experimental results demonstrate that the proposed approach can combine the strength of state-of-the-art AMR parsers to create new predictions that are more accurate than any individual models in five standard benchmark datasets. Hoang Thanh Lam, Gabriele Picco, Yufang Hou 0001, Young-Suk Lee 0001, Lam M. Nguyen, Dzung T. Phan, Vanessa López, Ramón Fernandez Astudillo |
NeurIPS | 8 |
| 2020 | GPT-too: A Language-Model-First Approach for AMR-to-Text GenerationabstractManuel Mager, Ramón Fernandez Astudillo, Tahira Naseem, Md Arafat Sultan, Young-Suk Lee, Radu Florian, Salim Roukos. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Manuel Mager, Ramón Fernandez Astudillo, Tahira Naseem, Md. Arafat Sultan, Young-Suk Lee 0001, Radu Florian, Salim Roukos |
ACL | 2 |
| 2020 | On the Importance of Diversity in Question Generation for QAabstractAutomatic question generation (QG) has shown promise as a source of synthetic training data for question answering (QA).In this paper we ask: Is textual diversity in QG beneficial for downstream QA?Using top-p nucleus sampling to derive samples from a transformer-based question generator, we show that diversity-promoting QG indeed provides better QA training than likelihood maximization approaches such as beam search.We also show that standard QG evaluation metrics such as BLEU, ROUGE and METEOR are inversely correlated with diversity, and propose a diversity-aware intrinsic measure of overall QG quality that correlates well with extrinsic evaluation on QA.1 Md. Arafat Sultan, Shubham Chandel, Ramón Fernandez Astudillo, Vittorio Castelli |
ACL | 3 |
| 2019 | Cycle-consistency Training for End-to-end Speech RecognitionabstractThis paper presents a method to train end-to-end automatic speech recognition (ASR) models using unpaired data. Although the end-to-end approach can eliminate the need for expert knowledge such as pronunciation dictionaries to build ASR systems, it still requires a large amount of paired data, i.e., speech utterances and their transcriptions. Cycle-consistency losses have been recently proposed as a way to mitigate the problem of limited paired data. These approaches compose a reverse operation with a given transformation, e.g., text-to-speech (TTS) with ASR, to build a loss that only requires unsupervised data, speech in this example. Applying cycle consistency to ASR models is not trivial since fundamental information, such as speaker traits, are lost in the intermediate text bottleneck. To solve this problem, this work presents a loss that is based on the speech encoder state sequence instead of the raw speech signal. This is achieved by training a Text-To-Encoder model and defining a loss based on the encoder reconstruction error. Experimental results on the LibriSpeech corpus show that the proposed cycle-consistency training reduced the word error rate by 14.7% from an initial model trained with 100-hour paired data, using an additional 360 hours of audio data without transcriptions. We also investigate the use of text-only data mainly for language modeling to further improve the performance in the unpaired data training scenario. Takaaki Hori, Ramón Fernandez Astudillo, Tomoki Hayashi, Yu Zhang 0033, Shinji Watanabe 0001, Jonathan Le Roux |
ICASSP | 2 |
| 2019 | Semi-Supervised Sequence-to-Sequence ASR Using Unpaired Speech and TextabstractSequence-to-sequence automatic speech recognition (ASR) models require large quantities of data to attain high performance. For this reason, there has been a recent surge in interest for unsupervised and semi-supervised training in such models. This work builds upon recent results showing notable improvements in semi-supervised training using cycle-consistency and related techniques. Such techniques derive training procedures and losses able to leverage unpaired speech and/or text data by combining ASR with Text-to-Speech (TTS) models. In particular, this work proposes a new semi-supervised loss combining an end-to-end differentiable ASR$\rightarrow$TTS loss with TTS$\rightarrow$ASR loss. The method is able to leverage both unpaired speech and text data to outperform recently proposed related techniques in terms of \%WER. We provide extensive results analyzing the impact of data quantity and speech and text modalities and show consistent gains across WSJ and Librispeech corpora. Our code is provided in ESPnet to reproduce the experiments. Murali Karthick Baskar, Shinji Watanabe 0001, Ramón Fernandez Astudillo, Takaaki Hori, Lukás Burget, Jan Cernocký |
INTERSPEECH | 3 |
| 2018 | Back-Translation-Style Data Augmentation for end-to-end ASRabstractIn this paper we propose a novel data augmentation method for attention-based end-to-end automatic speech recognition (E2E-ASR), utilizing a large amount of text which is not paired with speech signals. Inspired by the back-translation technique proposed in the field of machine translation, we build a neural text-to-encoder model which predicts a sequence of hidden states extracted by a pre-trained E2E-ASR encoder from a sequence of characters. By using hidden states as a target instead of acoustic features, it is possible to achieve faster attention learning and reduce computational cost, thanks to sub-sampling in E2E-ASR encoder, also the use of the hidden states can avoid to model speaker dependencies unlike acoustic features. After training, the text-to-encoder model generates the hidden states from a large amount of unpaired text, then E2E-ASR decoder is retrained using the generated hidden states as additional training data. Experimental evaluation using LibriSpeech dataset demonstrates that our proposed method achieves improvement of ASR performance and reduces the number of unknown words without the need for paired data. Tomoki Hayashi, Shinji Watanabe 0001, Yu Zhang 0033, Tomoki Toda, Takaaki Hori, Ramón Fernandez Astudillo, Kazuya Takeda |
SLT | 6 |
| 2017 | Segment Level Voice Conversion with Recurrent Neural Networks
Miguel Varela Ramos, Alan W. Black, Ramón Fernandez Astudillo, Isabel Trancoso, Nuno Fonseca |
INTERSPEECH | 3 |
| 2017 | A Semi-Supervised Learning Approach for Acoustic-Prosodic Personality Perception in Under-Resourced DomainsabstractAutomatic personality analysis has gained attention in the last years as a fundamental dimension in human-To-human and human-To-machine interaction. However, it still suffers from limited number and size of speech corpora for specific domains, such as the assessment of children's personality. This paper investigates a semi-supervised training approach to tackle this scenario. We devise an experimental setup with age and language mismatch and two training sets: A small labeled training set from the Interspeech 2012 Personality Sub-challenge, containing French adult speech labeled with personality OCEAN traits, and a large unlabeled training set of Portuguese children's speech. As test set, a corpus of Portuguese children's speech labeled with OCEAN traits is used. Based on this setting, we investigate a weak supervision approach that iteratively refines an initial model trained with the labeled data-set using the unlabeled data-set. We also investigate knowledge-based features, which leverage expert knowledge in acoustic-prosodic cues and thus need no extra data. Results show that, despite the large mismatch imposed by language and age differences, it is possible to attain improvements with these techniques, pointing both to the benefits of using a weak supervision and expert-based acoustic-prosodic features across age and language. Rubén Solera-Ureña, Helena Moniz, Fernando Batista, Vera Cabarrão, Anna Pompili, Ramón Fernandez Astudillo, Joana Campos 0001, Ana Paiva 0001, Isabel Trancoso |
INTERSPEECH | 6 |
| 2017 | Pushing the Limits of Translation Quality EstimationabstractTranslation quality estimation is a task of growing importance in NLP, due to its potential to reduce post-editing human effort in disruptive ways. However, this potential is currently limited by the relatively low accuracy of existing systems. In this paper, we achieve remarkable improvements by exploiting synergies between the related tasks of word-level quality estimation and automatic post-editing. First, we stack a new, carefully engineered, neural model into a rich feature-based word-level quality estimation system. Then, we use the output of an automatic post-editing system as an extra feature, obtaining striking results on WMT16: a word-level FMULT1 score of 57.47% (an absolute gain of +7.95% over the current state of the art), and a Pearson correlation score of 65.56% for sentence-level HTER prediction (an absolute gain of +13.36%). André F. T. Martins, Marcin Junczys-Dowmunt, Fábio N. Kepler, Ramón Fernandez Astudillo, Chris Hokamp, Roman Grundkiewicz |
Trans. Assoc. Comput. Linguistics | 4 |
| 2016 | A new uncertainty decoding scheme for DNN-HMM hybrid systems with multichannel speech enhancementabstractUncertainty decoding combines a probabilistic feature description with the acoustic model of a speech recognition system. For DNN-HMM hybrid systems, this can be realized by averaging the DNN outputs produced by a finite set of feature samples (drawn from an estimated probability distribution). In this article, we employ this sampling approach in combination with a multi-microphone speech enhancement system. We propose a new strategy for generating feature samples from multichannel signals, based on modeling the spatial coherence estimates between different microphone pairs as realizations of a latent random variable. From each coherence estimate, a spectral enhancement gain is computed and an enhanced feature vector is obtained, thus producing a finite set of feature samples, of which we average the respective DNN outputs. In the experimental part, this new uncertainty decoding strategy is shown to consistently improve the recognition accuracy of a DNN-HMM hybrid system for the 8-channel REVERB Challenge task. Christian Huemmer 0001, Andreas Schwarz, Roland Maas, Hendrik Barfuss, Ramón Fernandez Astudillo, Walter Kellermann |
ICASSP | 5 |
| 2016 | From Softmax to Sparsemax: A Sparse Model of Attention and Multi-Label ClassificationabstractWe propose sparsemax, a new activation function similar to the traditional softmax, but able to output sparse probabilities. After deriving its properties, we show how its Jacobian can be efficiently computed, enabling its use in a network trained with backpropagation. Then, we propose a new smooth and convex loss function which is the sparsemax analogue of the logistic loss. We reveal an unexpected connection between this new loss and the Huber classification loss. We obtain promising empirical results in multi-label classification problems and in attention-based neural networks for natural language inference. For the latter, we achieve a similar performance as the traditional softmax, but with a selective, more compact, attention focus. André F. T. Martins, Ramón Fernandez Astudillo |
ICML | 2 |
| 2016 | Exploiting Phone Log-Likelihood Ratio Features for the Detection of the Native Language of Non-Native English Speakers
Alberto Abad, Eugénio Ribeiro, Fábio N. Kepler, Ramón Fernandez Astudillo, Isabel Trancoso |
INTERSPEECH | 4 |
| 2016 | Uncertain LDA: Including Observation Uncertainties in Discriminative TransformsabstractLinear discriminant analysis (LDA) is a powerful technique in pattern recognition to reduce the dimensionality of data vectors. It maximizes discriminability by retaining only those directions that minimize the ratio of within-class and between-class variance. In this paper, using the same principles as for conventional LDA, we propose to employ uncertainties of the noisy or distorted input data in order to estimate maximally discriminant directions. We demonstrate the efficiency of the proposed uncertain LDA on two applications using state-of-the-art techniques. First, we experiment with an automatic speech recognition task, in which the uncertainty of observations is imposed by real-world additive noise. Next, we examine a full-scale speaker recognition system, considering the utterance duration as the source of uncertainty in authenticating a speaker. The experimental results show that when employing an appropriate uncertainty estimation algorithm, uncertain LDA outperforms its conventional LDA counterpart. Rahim Saeidi, Ramón Fernandez Astudillo, Dorothea Kolossa |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Learning Word Representations from Scarce and Noisy Data with Embedding SubspacesabstractRamon F. Astudillo, Silvio Amir, Wang Ling, Mário Silva, Isabel Trancoso. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Ramón Fernandez Astudillo, Silvio Amir, Wang Ling, Mário J. Silva, Isabel Trancoso |
ACL (1) | 1 |
| 2015 | Integration of DNN based speech enhancement and ASRabstractSpeech enhancement employing Deep Neural Networks (DNNs) is gaining strength as a data-driven alternative to classical Minimum Mean Square Error (MMSE) enhancement approaches. In the past, Observation Uncertainty approaches to integrate MMSE speech enhancement with Automatic Speech Recognition (ASR) have yielded good results as a lightweight alternative for robust ASR. In this paper we thus explore the integration of DNN-based speech enhancement with ASR by employing Observation Uncertainty techniques. For this purpose, we explore various techniques and approximations that allow propagating the uncertainty of inference of the DNN into feature domain. This uncertainty can then be used to dynamically compensate the ASR model utilizing techniques like uncertainty decoding. We test the proposed techniques on the AURORA4 corpus and show that notable improvements can be attained over the already effective DNN enhancement. Ramón Fernandez Astudillo, Maria Joana Correia, Isabel Trancoso |
INTERSPEECH | 1 |
| 2015 | Robust speech processing using observation uncertainty and uncertainty propagation: session and paper overview
Ramón Fernandez Astudillo, Shinji Watanabe 0001, Ahmed Hussen Abdelaziz, Dorothea Kolossa |
INTERSPEECH | 1 |
| 2015 | Uncertainty decoding for DNN-HMM hybrid systems based on numerical sampling
Christian Huemmer 0001, Roland Maas, Andreas Schwarz, Ramón Fernandez Astudillo, Walter Kellermann |
INTERSPEECH | 4 |
| 2014 | Accounting for the residual uncertainty of multi-layer perceptron based featuresabstractMulti-Layer Perceptrons (MLPs) are often interpreted as modeling a posterior distribution over classes given input features using the mean field approximation. This approximation is fast but neglects the residual uncertainty of inference at each layer, making inference less robust. In this paper we introduce a new approximation of MLP inference that takes under consideration this residual uncertainty. The proposed algorithm propagates not only the mean, but also the variance of inference through the network. At the current stage, the proposed method can not be used with soft-max layers. Therefore, we illustrate the benefits of this algorithm in a tandem scheme. We use the residual uncertainty of inference of MLP-based features to compensate a GMM-HMM backend with uncertainty decoding. Experiments on the Aurora4 corpus show consistent improvement of performance against conventional MLPs for all scenarios, in particular for clean speech and multi-style training. Ramón Fernandez Astudillo, Alberto Abad, Isabel Trancoso |
ICASSP | 1 |
| 2014 | The DIRHA-GRID corpus: baseline and tools for multi-room distant speech recognition using distributed microphonesabstractDistant speech recognition in real-world environments is still a challenging problem and a particularly interesting topic is the investigation of multi-channel processing in case of distributed microphones in home environments. This paper presents an initiative oriented to address the challenges of such a scenario; an experimental recognition framework comprising a multi-room, multi-channel corpus and the accompanying evaluation tools is made publicly available. The overall goal is to represent a common platform for comparing state-of-the-art algorithms, share ideas of different research communities and integrate several components in a realistic distant-talking recognition chain, e.g., voice activity detection, speech/feature enhancement, channel selection and fusion, model Marco Matassoni, Ramón Fernandez Astudillo, Athanasios Katsamanis, Mirco Ravanelli |
INTERSPEECH | 2 |
| 2013 | A propagation approach to modelling the joint distributions of clean and corrupted speech in the Mel-Cepstral domainabstractThis paper presents a closed form solution relating the joint distributions of corrupted and clean speech in the short-time Fourier Transform (STFT) and Mel-Frequency Cepstral Coefficient (MFCC) domains. This makes possible a tighter integration of STFT domain speech enhancement and feature and model-compensation techniques for robust automatic speech recognition. The approach directly utilizes the conventional speech distortion model for STFT speech enhancement, allowing for low cost, single pass, causal implementations. Compared to similar uncertainty propagation approaches, it provides the full joint distribution, rather than just the posterior distribution, which provides additional model compensation possibilities. The method is exemplified by deriving an MMSE-MFCC estimator from the propagated joint distribution. It is shown that similar performance to that of STFT uncertainty propagation (STFT-UP) can be obtained on the AURORA4, while deriving the full joint distribution. Ramón Fernandez Astudillo |
ASRU | 1 |
| 2013 | On the relation between speech corruption models in the spectral and the cepstral domainabstractThe Gaussian distortion model in the short-time Fourier transform (STFT) domain is the basis of many of the modern speech enhancement algorithms. One of the reasons is that additive sources and late reverberation can be analyzed and processed quite efficiently in this domain. The STFT domain is however not well related to acoustic quality and is also not well suited for learning models due to the high variability of speech in this domain. On the other hand, the cepstral domain has proved to be very well suited for these last two purposes, however, at the cost of loosing the simple linear relation between desired source and additive interferences. In this paper we explore the relation between the Gaussian distortion models in the STFT and the cepstral domain. We show how the assumption of a jointly Gaussian distortion model in the cepstrum domain is fulfilled for well-known distortion models in STFT domain. We provide closed-form solutions relating the joint distributions of corrupted and clean speech in the STFT and the cepstrum domain. We also propose various ways in which this model can be used to enhance speech. Ramón Fernandez Astudillo, Timo Gerkmann |
ICASSP | 1 |
| 2013 | Integration of beamforming and uncertainty-of-observation techniques for robust ASR in multi-source environments
Ramón Fernandez Astudillo, Dorothea Kolossa, Alberto Abad, Steffen Zeiler, Rahim Saeidi, Pejman Mowlaee, João Paulo da Silva Neto, Rainer Martin 0001 |
Comput. Speech Lang. | 1 |
| 2013 | An Extension of STFT Uncertainty Propagation for GMM-Based Super-Gaussian a Priori ModelsabstractFeature compensation is a low computational cost technique to achieve robust automatic speech recognition (ASR). Short-time Fourier Transform Uncertainty Propagation (STFT-UP) provides feature compensation in domains used for ASR as, e.g., Mel-Frequency Cepstra Coefficient (MFCC), while using STFT domain distortion models. However, STFT-UP is limited to Gaussian priors when modeling speech distortion, whereas super-Gaussian priors are known to provide improved performance. In this letter, an extension of STFT-UP is presented that uses approximate super-Gaussian priors. This is achieved by extending the conventional complex Gaussian priors to complex Gaussian mixture priors. The approach can be applied to any of the STFT-UP existing solutions, thus providing super-Gaussian uncertainty propagation. The method is exemplified by a Minimum Mean Square Error (MMSE) MFCC estimator with an approximate generalized Gamma speech prior. This estimator clearly outperforms the Gaussian-based MMSE-MFCC feature compensation on the AURORA4 corpus. Ramón Fernandez Astudillo |
IEEE Signal Process. Lett. | 1 |
| 2013 | Noise-Adaptive LDA: A New Approach for Speech Recognition Under Observation UncertaintyabstractAutomatic speech recognition (ASR) performance suffers severely from non-stationary noise, precluding widespread use of ASR in natural environments. Recently, so-termed uncertainty-of-observation techniques have helped to recover good performance. These consider the clean speech features as a hidden variable, of which the observable features are only an imperfect estimate. An estimated error variance of features is therefore used to further guide recognition. Based on the same idea, we introduce a new strategy: Reducing the speech feature dimensionality for optimal discriminance under observation uncertainty can yield significantly improved recognition performance, and is derived easily via Fisher's criterion of discriminant analysis. Dorothea Kolossa, Steffen Zeiler, Rahim Saeidi, Ramón Fernandez Astudillo |
IEEE Signal Process. Lett. | 4 |
| 2013 | Computing MMSE Estimates and Residual Uncertainty Directly in the Feature Domain of ASR using STFT Domain Speech Distortion ModelsabstractIn this paper we demonstrate how uncertainty propagation allows the computation of minimum mean square error (MMSE) estimates in the feature domain for various feature extraction methods using short-time Fourier transform (STFT) domain distortion models. In addition to this, a measure of estimate reliability is also attained which allows either feature re-estimation or the dynamic compensation of automatic speech recognition (ASR) models. The proposed method transforms the posterior distribution associated to a Wiener filter through the feature extraction using the STFT Uncertainty Propagation formulas. It is also shown that non-linear estimators in the STFT domain like the Ephraim-Malah filters can be seen as special cases of a propagation of the Wiener posterior. The method is illustrated by developing two MMSE-Mel-frequency Cepstral Coefficient (MFCC) estimators and combining them with observation uncertainty techniques. We discuss similarities with other MMSE-MFCC estimators and show how the proposed approach outperforms conventional MMSE estimators in the STFT domain on the AURORA4 robust ASR task. Ramón Fernandez Astudillo, Reinhold Orglmeister |
IEEE Trans. Speech Audio Process. | 1 |
| 2013 | Corpus-Based Speech Enhancement With Uncertainty Modeling and Cepstral SmoothingabstractWe present a new approach for corpus-based speech enhancement that significantly improves over a method published by Xiao and Nickel in 2010. Corpus-based enhancement systems do not merely filter an incoming noisy signal, but resynthesize its speech content via an inventory of pre-recorded clean signals. The goal of the procedure is to perceptually improve the sound of speech signals in background noise. The proposed new method modifies Xiao's method in four significant ways. Firstly, it employs a Gaussian mixture model (GMM) instead of a vector quantizer in the phoneme recognition front-end. Secondly, the state decoding of the recognition stage is supported with an uncertainty modeling technique. With the GMM and the uncertainty modeling it is possible to eliminate the need for noise dependent system training. Thirdly, the post-processing of the original method via sinusoidal modeling is replaced with a powerful cepstral smoothing operation. And lastly, due to the improvements of these modifications, it is possible to extend the operational bandwidth of the procedure from 4 kHz to 8 kHz. The performance of the proposed method was evaluated across different noise types and different signal-to-noise ratios. The new method was able to significantly outperform traditional methods, including the one by Xiao and Nickel, in terms of PESQ scores and other objective quality measures. Results of subjective CMOS tests over a smaller set of test samples support our claims. Robert M. Nickel, Ramón Fernandez Astudillo, Dorothea Kolossa, Rainer Martin 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Integration of beamforming and automatic speech recognition through propagation of the wiener posteriorabstractThis paper details one of the front-end components of the system used at the PASCAL-CHiME multi-source robust automatic speech recognition (ASR) challenge 2011. The presented approach uses uncertainty propagation techniques to integrate conventional beamforming with automatic speech recognition. The paper addresses the derivation of a complex Gaussian posterior for the multi-channel Wiener and the delay and sum beamformer and introduces a new approach based on the propagation of the Wiener posterior through the resynthesizing process. Results on the PASCAL-CHiME task for this algorithms show that they consistently outperform conventional beamfomers with a minimal increase in computational complexity. Ramón Fernandez Astudillo, Alberto Abad, João Paulo da Silva Neto |
ICASSP | 1 |
| 2012 | Inventory-style speech enhancement with uncertainty-of-observation techniquesabstractWe present a new method for inventory-style speech enhancement that significantly improves over earlier approaches [1]. Inventory-style enhancement attempts to resynthesize a clean speech signal from a noisy signal via corpus-based speech synthesis. The advantage of such an approach is that one is not bound to trade noise suppression against signal distortion in the same way that most traditional methods do. A significant improvement in perceptual quality is typically the result. Disadvantages of this new approach, however, include speaker dependency, increased processing delays, and the necessity of substantial system training. Earlier published methods relied on a-priori knowledge of the expected noise type during the training process [1]. In this paper we present a new method that exploits uncertainty-of-observation techniques to circumvent the need for noise specific training. Experimental results show that the new method is not only able to match, but outperform the earlier approaches in perceptual quality. Robert M. Nickel, Ramón Fernandez Astudillo, Dorothea Kolossa, Steffen Zeiler, Rainer Martin 0001 |
ICASSP | 2 |
| 2012 | Uncertainty driven Compensation of Multi-Stream MLP Acoustic Models for Robust ASR
Ramón Fernandez Astudillo, Alberto Abad, João Paulo da Silva Neto |
INTERSPEECH | 1 |
| 2011 | Propagation of Uncertainty Through Multilayer Perceptrons for Robust Automatic Speech Recognition
Ramón Fernandez Astudillo, João Paulo da Silva Neto |
INTERSPEECH | 1 |
| 2010 | A MMSE estimator in mel-cepstral domain for robust large vocabulary automatic speech recognition using uncertainty propagation
Ramón Fernandez Astudillo, Reinhold Orglmeister |
INTERSPEECH | 1 |
| 2009 | Accounting for the uncertainty of speech estimates in the complex domain for minimum mean square error speech enhancement
Ramón Fernandez Astudillo, Dorothea Kolossa, Reinhold Orglmeister |
INTERSPEECH | 1 |