VLDB 2026 Research / reviewers in the wild / expert
Mireia Díez
dblp:63/8158 · also Mireia Díez Sánchez
· DBLP profile ↗
44ranked-venue papers
12as first author
12since 2021 · last 2025
0000-0001-7894-8377ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 11 first-author · 10 since 2021Artificial intelligence and machine learning · 32 · 8 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Leveraging Self-Supervised Learning for Speaker DiarizationabstractEnd-to-end neural diarization has evolved considerably over the past few years, but data scarcity is still a major obstacle for further improvements. Self-supervised learning methods such as WavLM have shown promising performance on several downstream tasks, but their application on speaker diarization is somehow limited. In this work, we explore using WavLM to alleviate the problem of data scarcity for neural diarization training. We use the same pipeline as Pyannote and improve the local end-to-end neural diarization with WavLM and Conformer. Experiments on far-field AMI, AISHELL-4, and AliMeeting datasets show that our method substantially outperforms the Pyannote baseline and achieves new state-of-the-art results on AMI and AISHELL4, respectively. In addition, by analyzing the system performance under different data quantity scenarios, we show that WavLM representations are much more robust against data scarcity than filterbank features, enabling less data hungry training strategies. Furthermore, we found that simulated data, usually used to train end-to-end diarization models, does not help when using WavLM in our experiments. Additionally, we also evaluate our model on the recent CHiME8 NOTSOFAR-1 task where it achieves better performance than the Pyannote baseline. Our source code is publicly available at https://github.com/BUTSpeechFIT/DiariZen. Jiangyu Han, Federico Landini, Johan Rohdin, Anna Silnova, Mireia Díez, Lukás Burget |
ICASSP | 5 |
| 2025 | Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization
Jiangyu Han, Federico Landini, Johan Rohdin, Anna Silnova, Mireia Díez, Jan Cernocký, Lukás Burget |
INTERSPEECH | 5 |
| 2024 | Diacorrect: Error Correction Back-End for Speaker DiarizationabstractIn this work, we propose an error correction framework, named DiaCorrect, to refine the output of a diarization system in a simple yet effective way. This method is inspired by error correction techniques in automatic speech recognition. Our model consists of two parallel convolutional encoders and a transformer-based decoder. By exploiting the interactions between the input recording and the initial system’s outputs, DiaCorrect can automatically correct the initial speaker activities to minimize the diarization errors. Experiments on 2-speaker telephony data show that the proposed DiaCorrect can effectively improve the initial model’s results. Our source code is publicly available at https://github.com/BUTSpeechFIT/diacorrect. Jiangyu Han, Federico Landini, Johan Rohdin, Mireia Díez, Lukás Burget, Yuhang Cao, Jan Cernocký |
ICASSP | 4 |
| 2024 | Discriminative Training of VBx DiarizationabstractBayesian HMM clustering of x-vector sequences (VBx) has become a widely adopted diarization baseline model in publications and challenges. It uses an HMM to model speaker turns, a generatively trained probabilistic linear discriminant analysis (PLDA) for speaker distribution modeling, and Bayesian inference to estimate the assignment of x-vectors to speakers. This paper presents a new framework for updating the VBx parameters using discriminative training, which directly optimizes a predefined loss. We also propose a new loss that better correlates with the diarization error rate compared to binary cross-entropy — the default choice for diarization end-to-end systems. Proof-of-concept results across three datasets (AMI, CALLHOME, and DIHARD II) demonstrate the method’s capability of automatically finding hyperparameters, achieving comparable performance to those found by extensive grid search, which typically requires additional hyperparameter behavior knowledge. Moreover, we show that discriminative fine-tuning of PLDA can further improve the model’s performance. We release the source code with this publication. Dominik Klement, Mireia Díez, Federico Landini, Lukás Burget, Anna Silnova, Marc Delcroix, Naohiro Tawara |
ICASSP | 2 |
| 2024 | Spoof Diarization: "What Spoofed When" in Partially Spoofed Audio
Lin Zhang 0054, Xin Wang 0037, Erica Cooper, Mireia Díez, Federico Landini, Nicholas W. D. Evans, Junichi Yamagishi |
INTERSPEECH | 4 |
| 2024 | DiaPer: End-to-End Neural Diarization With Perceiver-Based AttractorsabstractUntil recently, the field of speaker diarization was dominated by cascaded systems. Due to their limitations, mainly regarding overlapped speech and cumbersome pipelines, end-to-end models have gained great popularity lately. One of the most successful models is end-to-end neural diarization with encoder-decoder based attractors (EEND-EDA). In this work, we replace the EDA module with a Perceiver-based one and show its advantages over EEND-EDA; namely obtaining better performance on the largely studied Callhome dataset, finding the quantity of speakers in a conversation more accurately, and faster inference time. Furthermore, when exhaustively compared with other methods, our model, DiaPer, reaches remarkable performance with a very lightweight design. Besides, we perform comparisons with other works and a cascaded baseline across more than ten public wide-band datasets. Together with this publication, we release the code of DiaPer as well as models trained on public and free data. Federico Landini, Mireia Díez, Themos Stafylakis, Lukás Burget |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Multi-Speaker and Wide-Band Simulated Conversations as Training Data for End-to-End Neural DiarizationabstractEnd-to-end diarization presents an attractive alternative to standard cascaded diarization systems because a single system can handle all aspects of the task at once. Many flavors of end-to-end models have been proposed but all of them require (so far non-existing) large amounts of annotated data for training. The compromise solution consists in generating synthetic data and the recently proposed simulated conversations (SC) have shown remarkable improvements over the original simulated mixtures (SM). In this work, we create SC with multiple speakers per conversation and show that they allow for substantially better performance than SM, also reducing the dependence on a fine-tuning stage. We also create SC with wide-band public audio sources and present an analysis on several evaluation sets. Together with this publication, we release the recipes for generating such data and models trained on public sets as well as the implementation to efficiently handle multiple speakers per conversation and an auxiliary voice activity detection loss. Federico Landini, Mireia Díez, Alicia Lozano-Diez, Lukás Burget |
ICASSP | 2 |
| 2023 | Multi-Stream Extension of Variational Bayesian HMM Clustering (MS-VBx) for Combined End-to-End and Vector Clustering-based Diarization
Marc Delcroix, Naohiro Tawara, Mireia Díez, Federico Landini, Anna Silnova, Atsunori Ogawa, Tomohiro Nakatani, Lukás Burget, Shoko Araki |
INTERSPEECH | 3 |
| 2022 | Speaker adaptation for Wav2vec2 based dysarthric ASR
Murali Karthick Baskar, Tim Herzig, Diana Nguyen, Mireia Díez, Tim Polzehl, Lukás Burget, Jan Cernocký |
INTERSPEECH | 4 |
| 2022 | From Simulated Mixtures to Simulated Conversations as Training Data for End-to-End Neural DiarizationabstractEnd-to-end neural diarization (EEND) is nowadays one of the most prominent research topics in speaker diarization. EEND presents an attractive alternative to standard cascaded diarization systems since a single system is trained at once to deal with the whole diarization problem. Several EEND variants and approaches are being proposed, however, all these models require large amounts of annotated data for training but available annotated data are scarce. Thus, EEND works have used mostly simulated mixtures for training. However, simulated mixtures do not resemble real conversations in many aspects. In this work we present an alternative method for creating synthetic conversations that resemble real ones by using statistics about distributions of pauses and overlaps estimated on genuine conversations. Furthermore, we analyze the effect of the source of the statistics, different augmentations and amounts of data. We demonstrate that our approach performs substantially better than the original one, while reducing the dependence on the fine-tuning stage. Experiments are carried out on 2-speaker telephone conversations of Callhome and DIHARD 3. Together with this publication, we release our implementations of EEND and the method for creating simulated conversations. Index Terms: speaker diarization, end-to-end neural diarization, simulated conversations Federico Landini, Alicia Lozano-Diez, Mireia Díez, Lukás Burget |
INTERSPEECH | 3 |
| 2022 | Bayesian HMM clustering of x-vector sequences (VBx) in speaker diarization: Theory, implementation and analysis on standard tasks
Federico Landini, Ján Profant, Mireia Díez, Lukás Burget |
Comput. Speech Lang. | 3 |
| 2021 | Analysis of the but Diarization System for Voxconverse ChallengeabstractThis paper describes the system developed by the BUT team for the fourth track of the VoxCeleb Speaker Recognition Challenge, focusing on diarization on the VoxConverse dataset. The system consists of signal pre-processing, voice activity detection, speaker embedding extraction, an initial agglomerative hierarchical clustering followed by diarization using a Bayesian hidden Markov model, a reclustering step based on per-speaker global embeddings and overlapped speech detection and handling. We provide comparisons for each of the steps and share the implementation of the most relevant modules of our system. Our system scored second in the challenge in terms of the primary metric (diarization error rate) and first according to the secondary metric (Jaccard error rate). Federico Landini, Ondrej Glembek, Pavel Matejka, Johan Rohdin, Lukás Burget, Mireia Díez, Anna Silnova |
ICASSP | 6 |
| 2020 | Optimizing Bayesian Hmm Based X-Vector Clustering for the Second Dihard Speech Diarization ChallengeabstractThis paper presents an analysis of our diarization system winning the second DIHARD speech diarization challenge, track 1. This system is based on clustering x-vector speaker embeddings extracted every 0.25s from short segments of the input recording. In this paper, we focus on the two x-vector clustering methods employed, namely Agglomerative Hierarchical Clustering followed by a clustering based on Bayesian Hidden Markov Model (BHMM). Even though the system submitted to the challenge had further post-processing steps, we will show that using this BHMM solely is enough to achieve the best performance in the challenge. The analysis will show improvements achieved by optimizing individual processing steps, including a simple procedure to effectively perform "domain adaptation" by Probabilistic Linear Discriminant Analysis model interpolation. All experiments are performed in the DIHARD II evaluation framework. Mireia Díez, Lukás Burget, Federico Landini, Shuai Wang 0016, Jan Cernocký |
ICASSP | 1 |
| 2020 | But System for the Second Dihard Speech Diarization ChallengeabstractThis paper describes the winning systems developed by the BUT team for the four tracks of the Second DIHARD Speech Diarization Challenge. For tracks 1 and 2 the systems were mainly based on performing agglomerative hierarchical clustering (AHC) of x-vectors, followed by another x-vector clustering based on Bayes hidden Markov model and variational Bayes inference. We provide a comparison of the improvement given by each step and share the implementation of the core of the system. For tracks 3 and 4 with recordings from the Fifth CHiME Challenge, we explored different approaches for doing multi-channel diarization and our best performance was obtained when applying AHC on the fusion of per channel probabilistic linear discriminant analysis scores. Federico Landini, Shuai Wang 0016, Mireia Díez, Lukás Burget, Pavel Matejka, Katerina Zmolíková, Ladislav Mosner, Anna Silnova, Oldrich Plchot, Ondrej Novotný, Hossein Zeinali, Johan Rohdin |
ICASSP | 3 |
| 2020 | 13 years of speaker recognition research at BUT, with longitudinal analysis of NIST SRE
Pavel Matejka, Oldrich Plchot, Ondrej Glembek, Lukás Burget, Johan Rohdin, Hossein Zeinali, Ladislav Mosner, Anna Silnova, Ondrej Novotný, Mireia Díez, Jan Cernocký |
Comput. Speech Lang. | 10 |
| 2020 | End-to-end DNN based text-independent speaker recognition for long and short utterances
Johan Rohdin, Anna Silnova, Mireia Díez, Oldrich Plchot, Pavel Matejka, Lukás Burget, Ondrej Glembek |
Comput. Speech Lang. | 3 |
| 2020 | Analysis of Speaker Diarization Based on Bayesian HMM With Eigenvoice PriorsabstractIn our previous work, we introduced our Bayesian Hidden Markov Model with eigenvoice priors, which has been recently recognized as the state-of-the-art model for Speaker Diarization. In this article we present a more complete analysis of the Diarization system. The inference of the model is fully described and derivations of all update formulas are provided for a complete understanding of the algorithm. An extensive analysis on the effect, sensitivity and interactions of all model parameters is provided, which might be used as a guide for their optimal setting. The newly introduced speaker regularization coefficient allows us to control the number of speakers inferred in an utterance. A naive speaker model merging strategy is also presented, which allows to drive the variational inference out of local optima. Experiments for the different diarization scenarios are presented on CALLHOME and DIHARD datasets. Mireia Díez, Lukás Burget, Federico Landini, Jan Cernocký |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Bayesian HMM Based x-Vector Clustering for Speaker Diarization
Mireia Díez, Lukás Burget, Shuai Wang 0016, Johan Rohdin, Jan Cernocký |
INTERSPEECH | 1 |
| 2018 | End-to-End DNN Based Speaker Recognition Inspired by I-Vector and PLDAabstractRecently, several end-to-end speaker verification systems based on deep neural networks (DNNs) have been proposed. These systems have been proven to be competitive for text-dependent tasks as well as for text-independent tasks with short utterances. However, for text-independent tasks with longer utterances, end-to-end systems are still outperformed by standard i-vector + PLDA systems. In this work, we develop an end-to-end speaker verification system that is initialized to mimic an i-vector + PLDA baseline. The system is then further trained in an end-to-end manner but regularized so that it does not deviate too far from the initial system. In this way we mitigate overfitting which normally limits the performance of end-to-end systems. The proposed system outperforms the i-vector + PLDA baseline on both long and short duration utterances. Johan Rohdin, Anna Silnova, Mireia Díez, Oldrich Plchot, Pavel Matejka, Lukás Burget |
ICASSP | 3 |
| 2018 | BUT System for DIHARD Speech Diarization Challenge 2018
Mireia Díez, Federico Landini, Lukás Burget, Johan Rohdin, Anna Silnova, Katerina Zmolíková, Ondrej Novotný, Karel Veselý, Ondrej Glembek, Oldrich Plchot, Ladislav Mosner, Pavel Matejka |
INTERSPEECH | 1 |
| 2017 | MGB-3 but system: Low-resource ASR on Egyptian YouTube dataabstractThis paper presents a series of experiments we performed during our work on the MGB-3 evaluations. We both describe the submitted system, as well as the post-evaluation analysis. Our initial BLSTM-HMM system was trained on 250 hours of MGB-2 data (Al-Jazeera), it was adapted with 5 hours of Egyptian data (YouTube). We included such techniques as diarization, n-gram language model adaptation, speed perturbation of the adaptation data, and the use of all 4 ‘correct’ references. The 4 references were either used for supervision with a ‘confusion network’, or we included each sentence 4x with the transcripts from all the annotators. Then, it was also helpful to blend the augmented MGB-3 adaptation data with 15 hours of MGB-2 data. Although we did not rank with our single system among the best teams in the evaluations, we believe that our analysis will be highly interesting not only for the other MGB-3 challenge participants. Karel Veselý, Murali Karthick Baskar, Mireia Díez, Karel Benes |
ASRU | 3 |
| 2017 | Analysis of Score Normalization in Multilingual Speaker Recognition
Pavel Matejka, Ondrej Novotný, Oldrich Plchot, Lukás Burget, Mireia Díez, Jan Cernocký |
INTERSPEECH | 5 |
| 2017 | Analysis and Description of ABC Submission to NIST SRE 2016
Oldrich Plchot, Pavel Matejka, Anna Silnova, Ondrej Novotný, Mireia Díez, Johan Rohdin, Ondrej Glembek, Niko Brümmer, Albert Swart, Jesús Jorrín-Prieto, L. Paola García-Perera, Luis Buera, Patrick Kenny, Jahangir Alam 0001, Gautam Bhattacharya |
INTERSPEECH | 5 |
| 2014 | High-performance Query-by-Example Spoken Term Detection on the SWS 2013 evaluationabstractIn the last years, the task of Query-by-Example Spoken Term Detection (QbE-STD), which aims to find occurrences of a spoken query in a set of audio documents, has gained the interest of the research community for its versatility in settings where untranscribed, multilingual and acoustically unconstrained spoken resources, or spoken resources in low-resource languages, must be searched. This paper describes and reports experimental results for a QbE-STD system that achieved the best performance in the recent Spoken Web Search (SWS) evaluation, held as part of MediaEval 2013. Though not optimized for speed, the system operates faster than real-time. The system exploits high-performance phone decoders to extract frame-level phone posteriors (a common representation in QbE-STD tasks). Then, given a query and a audio document, a distance matrix is computed between their phone posterior representations, followed by a newly introduced distance normalization technique and an iterative Dynamic Time Warping (DTW) matching procedure with some heuristic prunings. Results show that remarkable performance improvements can be achieved by using multiple examples per query and, specially, through the late (score-level) fusion of different subsystems, each based on a different set of phone posteriors. Luis Javier Rodríguez-Fuentes, Amparo Varona, Mikel Peñagarikano, Germán Bordel, Mireia Díez |
ICASSP | 5 |
| 2014 | Optimizing PLLR Features for Spoken Language RecognitionabstractPhone Log-Likelihood Ratios (PLLR) have been recently introduced as features for spoken language and speaker recognition systems. This representation has proven to be an effective way of retrieving acoustic-phonotactic information into frame-level vectors, which can be easily plugged into state-of-the-art systems. In a previous work, we began the search of reduced representations of PLLRs, as a mean of reducing computational costs. In this paper, we extend this search, by looking for the optimal compromise between feature vector size and system performance. Results achieved by Principal Component Analysis projection on the PLLR space are extensively analyzed. Also, to evaluate the effect of using larger temporal contexts, a Shifted Delta transformation is applied (and its optimal configuration explored) on highly reduced sets of PCA-projected PLLR features, leading to further performance improvements over the best PCA-projected PLLR set. Mireia Díez, Amparo Varona, Mikel Peñagarikano, Luis Javier Rodríguez-Fuentes, Germán Bordel |
ICPR | 1 |
| 2014 | On the complementarity of short-time fourier analysis windows of different lengths for improved language recognitionabstractPrevious works have shown that remarkable perfor-mance improvements can be attained in speaker and lan-guage recognition tasks by combining several heteroge-neous systems that provide complementary information. In this work, the complementarity of several i-vector lan-guage recognition systems, using Mel-Frequency Cep-stral-Coefficient (MFCC) features computed on Short-Time Fourier Analysis windows of different sizes, is stud-ied. Language recognition experiments carried out on the NIST 2007 and 2009 LRE datasets reveal relative per-formance gains of up to 33 % when fusing the systems, with regard to the best single system. Results suggest that combining acoustic systems based on analysis win-dows of different sizes may allow to get advantage from both the sharper characterization of short events provided by short windows and the better frequency resolution of stationary events provided by long windows. Mireia Díez, Mikel Peñagarikano, Germán Bordel, Amparo Varona, Luis Javier Rodríguez-Fuentes |
INTERSPEECH | 1 |
| 2014 | New insight into the use of phone log-likelihood ratios as features for language recognitionabstractPhone Log-Likelihood Ratio (PLLR) features have been recently introduced as an effective way of mak-ing use of frame-level phone posteriors in language and speaker recognition systems. In this paper, a deep insight into PLLR features is made and further evidence of the usefulness of these features in spoken language recognition tasks is provided, with a new set of experiments carried out on the NIST 2007 LRE dataset, combining the latest progresses made in optimiz-ing the features. PLLR features are projected into a subspace that enhances the information retrieved by the system. Then, di-mensionality reduction is performed on the projected subspace by means of Principal Component Analysis, and shifted deltas are computed on the reduced features to optimize performance. Figures attained are among the best reported so far on the NIST 2007 LRE dataset. Mireia Díez, Amparo Varona, Mikel Peñagarikano, Luis Javier Rodríguez-Fuentes, Germán Bordel |
INTERSPEECH | 1 |
| 2014 | PLLR features in language recognition system for RATSabstractIn this paper, we study the use of features based on frame-byframe phone posteriors (PLLRs) for language recognition. The results are reported on the datasets developed for the DARPA RATS (Robust Automatic Transcription of Speech) program, which seeks to advance state of the art detection capabilities on audio from highly degraded communication channels. We show that systems based on the PLLRs outperform the standard acoustic system based on PLP2 features. By experimenting with the system combinations, we also demonstrate that the PLLR-based systems contain complementary information with respect to the PLP2 system. Finally we make a comparison between the PLLR and phonotactic systems with the outcome favorable to the PLLR. Oldrich Plchot, Mireia Díez, Mehdi Soufifar, Lukás Burget |
INTERSPEECH | 2 |
| 2014 | KALAKA-3: a database for the recognition of spoken European languages on YouTube audios
Luis Javier Rodríguez-Fuentes, Mikel Peñagarikano, Amparo Varona, Mireia Díez, Germán Bordel |
LREC | 4 |
| 2014 | On the Complementarity of Phone Posterior Probabilities for Improved Speaker RecognitionabstractIn this letter, we apply Phone Log-Likelihood Ratio (PLLR) features to the task of speaker recognition. PLLRs, which are computed on the phone posterior probabilities provided by phone decoders, convey acoustic-phonetic information in a sequence of frame-level vectors, and therefore can be easily plugged into traditional acoustic systems, just by replacing the Mel-Frequency Cepstral Coefficients (MFCC) or an alternate representation. To study the performance of the proposed features, MFCC-based and PLLR-based systems are trained under an i-vector-PLDA approach. Results on the NIST 2010 and 2012 Speaker Recognition Evaluation databases show that, despite yielding lower performance than the acoustic system, the system based on PLLR features does provide significant gains when both systems are fused, which reveals a complementarity among features, and provides a suitable and effective way of using higher level phonetic information in speaker recognition systems. Mireia Díez, Amparo Varona, Mikel Peñagarikano, Luis Javier Rodríguez-Fuentes, Germán Bordel |
IEEE Signal Process. Lett. | 1 |
| 2014 | On the Projection of PLLRs for Unbounded Feature Distributions in Spoken Language RecognitionabstractThe so called Phone Log-Likelihood Ratio (PLLR) features have been recently introduced as a novel and effective way of retrieving acoustic-phonetic information in spoken language and speaker recognition systems. In this letter, an in-depth insight into the PLLR feature space is provided and the multidimensional distribution of these features is analyzed in a language recognition system. The study reveals that PLLR features are confined into a subspace that strongly bounds PLLR distributions. To enhance the information retrieved by the system, PLLR features are projected into a hyper-plane that provides a more suitable representation of the subspace where the features lie. After applying the projection method, PCA is used to decorrelate the features. Gains attained on each step of the proposed approach are outlined and compared to simple PCA projection. Experiments carried out on NIST 2007, 2009 and 2011 LRE datasets demonstrate the effectiveness of the proposed method, which yields up to a 27% relative improvement with regard to the system based on the original features. Mireia Díez, Amparo Varona, Mikel Peñagarikano, Luis Javier Rodríguez-Fuentes, Germán Bordel |
IEEE Signal Process. Lett. | 1 |
| 2013 | Dimensionality reduction of phone log-likelihood ratio features for spoken language recognitionabstractIn a previous work, we introduced the use of log-likelihood ratios of phone posterior probabilities, called Phone LogLikelihood Ratios (PLLR) as features for language recognition under an iVector-based approach, yielding high performance and promising results. However, the high dimensionality of the PLLR feature vectors (with regard to MFCC/SDC features) results in comparatively higher computational costs. In this work, several supervised and unsupervised dimensionality reduction techniques are studied, based on either fusions or selection of phone posteriors, finding that PLLR feature vectors can be reduced to almost a third of their original size attaining similar performance. Finally, Principal Component Analysis (PCA) is also applied to the original PLLR vector as a feature projection method for comparison purposes. Results show that PCA stands out among all the techniques studied, revealing that it does not only reduce computational costs, but also improves system performance significantly. Mireia Díez, Amparo Varona, Mikel Peñagarikano, Luis Javier Rodríguez-Fuentes, Germán Bordel |
INTERSPEECH | 1 |
| 2013 | Using phone log-likelihood ratios as features for speaker recognitionabstractThe so called Phone Log-Likelihood Ratio (PLLR) features, computed on phone posterior probabilities provided by phonetic decoders, convey acoustic-phonetic information in a sequence of frame-level vectors. Thus, PLLRs can be easily plugged into traditional acoustic systems just by replacing MFCCs, PLPs or whatever other representation. PLLR features were used under an iVector-PLDA approach in our submission to the NIST 2012 Speaker Recognition Evaluation (SRE). In this work, we present a report of the goodness of these features for speaker recognition. Results on the telephone clean speech condition of the NIST 2010 and 2012 SRE show that, although the system based on PLLR features does not reach state-ofthe-art performance, including it in a fusion with a traditional acoustic based system (trained on MFCC features) provides remarkable gains in performance (among the best reported in the NIST 2012 SRE telephone without added noise condition), revealing a fruitful way of using acoustic-phonetic information for speaker recognition. Mireia Díez, Amparo Varona, Mikel Peñagarikano, Luis Javier Rodríguez-Fuentes, Germán Bordel |
INTERSPEECH | 1 |
| 2013 | The albayzin 2012 language recognition evaluationabstractThe Albayzin 2008 Language Recognition Evaluation was held from May to October 2008, and their results presented and discussed among the participating teams at the 5th Biennial Workshop on Speech Technology [1], or-ganized by the Spanish Network on Speech Technologies [2] in November 2008. In this paper, we present (for the first time) a full description of the Albayzin 2008 LRE and analyze and discuss recognition results. The evalua-tion was designed according to the test procedures, pro-tocols and performance measures used in the NIST 2007 LRE. The KALAKA database [3], consisting of 16 kHz audio signals recorded from TV broadcasts, was created ad-hoc and used for the evaluation. The four official lan-guages spoken in Spain (Basque, Catalan, Galician and Luis Javier Rodríguez-Fuentes, Niko Brümmer, Mikel Peñagarikano, Amparo Varona, Germán Bordel, Mireia Díez |
INTERSPEECH | 6 |
| 2013 | Handling recordings acquired simultaneously over multiple channels with PLDAabstractIn some speaker recognition scenarios we find conversations recorded simultaneously over multiple channels.That is the case of the interviews in the NIST SRE dataset.To take advantage of that, we propose a modification of the PLDA model that considers two different inter-session variability terms.The first term is tied between all the recordings belonging to the same conversation whereas the second is not.Thus, the former mainly intends to capture the variability due to the phonetic content of the conversation while the latter tries to capture the channel variability.We test this approach on the NIST SRE12 core condition using multiple channels per interview to enroll the speakers.The proposed approach improves the minimum DCF by 26-29 % on telephone speech and by 1-8% on interviews compared to the standard PLDA (scored by the book). Jesús Villalba 0001, Mireia Díez, Amparo Varona, Eduardo Lleida |
INTERSPEECH | 2 |
| 2012 | Study of Different Backends in a State-Of-the-Art Language Recognition SystemabstractState of the art language recognition systems usually add a backend prior to the linear fusion of the subsystems scores. The backend plays a dual role. When the set of languages for which models have been trained does not match the set of target lan-guages, the backend maps the available scores to the space of target languages. On the other hand, the backend serves as a precalibration stage that adapts the space of scores. In this work, well known backends (Generative Gaussian Backend, Discrim-inative Gaussian Backend and Logistic Regression Backend) and newer proposals (Fully Bayesian Gaussian Backend and Gaussian Mixture Backend) are analyzed and compared. The effect of applying a T-Norm or a ZT-Norm is also analyzed. Fi-nally the effect of discarding development signals, those with the highest scores, is also studied. Experiments have been car-ried out on the NIST 2009 LRE database, using a state-of-the-art Language Recognition System consisting of the fusion of five subsystems: A Linearized Eigenchannel GMM (LE-GMM) subsystem, an iVector subsystem and three phone-lattice-SVM subsystems. Best performance was attained by Gaussian Mix-ture Backend (1.25 EER), yielding 23 % relative improvement with respect to the baseline (1.62 EER). Mikel Peñagarikano, Amparo Varona, Mireia Díez, Luis Javier Rodríguez-Fuentes, Germán Bordel |
INTERSPEECH | 3 |
| 2012 | The EHU Systems for the NIST 2011 Language Recognition EvaluationabstractThis paper describes the systems developed by the Software Technologies Working Group of the University of the Basque Country (EHU) for the NIST 2011 Language Recognition Eval-uation (LRE). One primary and three contrastive systems were submitted, all of them fusing five component subsystems: a Lin-earized Eigenchannel GMM (LE-GMM) subsystem, an iVector subsystem and three phone-lattice-SVM subsystems based on the publicly available BUT decoders for Czech, Hungarian an Russian. The four submitted systems were identical except for the backend approach and the development dataset used to esti-mate the backend and fusion parameters. Multiclass discrimina-tive fusion was performed separately for each nominal duration. A development set was defined, including the evaluation sets of LRE07 and LRE09 and the development data provided by NIST for 9 additional languages in the 2011 campaign. The official results, which were among the best submitted to the evaluation, are presented and briefly discussed. Post-key analyses are also addressed in the paper, including the performance attained by component subsystems and a study of their contribution to fu-sion performance by means of a greedy selection procedure. Mikel Peñagarikano, Amparo Varona, Luis Javier Rodríguez-Fuentes, Mireia Díez, Germán Bordel |
INTERSPEECH | 4 |
| 2012 | The BLZ Submission to the NIST 2011 LRE: Data Collection, System Development and PerformanceabstractThis paper describes the most relevant features of a collaborative multi-site submission to the NIST 2011 Language Recognition Evaluation (LRE), consisting of one primary and three contrastive systems, each fusing different combinations of 13 state-of-the-art (acoustic and phonotactic) language recognition subsystems.The collaboration focused on collecting and sharing training data for those target languages for which few development data were provided by NIST, and on defining a common development dataset to train backend and fusion parameters and select the best fusions.Official and post-key results are presented and compared, revealing that the greedy approach applied to select the best fusions provided suboptimal but very competitive performance.Several factors contributed to the high performance attained by BLZ systems, including the availability of training data for low resource target languages, the reliability of the development dataset (consisting only of data audited by NIST), the diversity of modeling approaches, features and datasets in the systems considered for fusion, and the effectiveness of the search for optimal fusions. Luis Javier Rodríguez-Fuentes, Mikel Peñagarikano, Amparo Varona, Mireia Díez, Germán Bordel, Alberto Abad, David Martínez González, Jesús Villalba 0001, Alfonso Ortega Giménez, Eduardo Lleida |
INTERSPEECH | 4 |
| 2012 | Using Time-Synchronous Phone Co-occurrences in a SVM-Phonotactic Dialect Recognition SystemabstractThis paper presents a simple approach to phonotactic dialect recognition which uses lattices of time-synchronous phone co-occurrences at the frame level. In previous works, we success-fully applied cross-decoder phone co-occurrences to improve performance in language recognition experiments on the 2007 NIST LRE database. We call phone co-occurrence to the si-multaneous (time-synchronous) presence of two phone units coming from two different phone decoders. In this work, the approach is ported to a Dialect Recognition task based on the assumption that co-occurrences can better represent the tiny differences among the dialects. Besides, a slightly different approach is presented, based on the simultaneous presence of two phone units in the lattice produced by a single decoder (intra-decoder phone co-occurrences). For evaluating the ap-proach, a choice of open software (Brno University of Technol-ogy phone decoders, HTK, SRILM, LIBLINEAR and FoCal) was used, and experiments were carried out on the Arabic di-alects of the NIST 2011 LRE database. The approach based on cross-decoder phone co-occurrences outperformed the base-line phonotactic system, yielding around 8 % relative improve-ment. The fusion of both systems yielded 7.31 % EER and CLLR = 0.497, meaning 19 % relative improvement. Amparo Varona, Mikel Peñagarikano, Luis Javier Rodríguez-Fuentes, Germán Bordel, Mireia Díez |
INTERSPEECH | 5 |
| 2012 | KALAKA-2: a TV Broadcast Speech Database for the Recognition of Iberian Languages in Clean and Noisy Environments
Luis Javier Rodríguez-Fuentes, Mikel Peñagarikano, Amparo Varona, Mireia Díez, Germán Bordel |
LREC | 4 |
| 2012 | On the use of phone log-likelihood ratios as features in spoken language recognitionabstractThis paper presents an alternative feature set to the traditional MFCC-SDC used in acoustic approaches to Spoken Language Recognition: the log-likelihood ratios of phone posterior probabilities, hereafter Phone Log-Likelihood Ratios (PLLR), produced by a phone recognizer. In this work, an iVector system trained on this set of features (plus dynamic coefficients) is evaluated and compared to (1) an acoustic iVector system (trained on the MFCC-SDC feature set) and (2) a phonotactic (Phone-lattice-SVM) system, using two different benchmarks: the NIST 2007 and 2009 LRE datasets. iVector systems trained on PLLR features proved to be competitive, reaching or even outperforming the MFCC-SDC-based iVector and the phonotactic systems. The fusion of the proposed approach with the acoustic and phonotactic systems provided even more significant improvements, outperforming state-of-the-art systems on both benchmarks. Mireia Díez, Amparo Varona, Mikel Peñagarikano, Luis Javier Rodríguez-Fuentes, Germán Bordel |
SLT | 1 |
| 2011 | Multi-site heterogeneous system fusions for the Albayzin 2010 Language Recognition EvaluationabstractBest language recognition performance is commonly obtained by fusing the scores of several heterogeneous systems. Regardless the fusion approach, it is assumed that different systems may contribute complementary information, either because they are developed on different datasets, or because they use different features or different modeling approaches. Most authors apply fusion as a final resource for improving performance based on an existing set of systems. Though relative performance gains decrease as larger sets of systems are considered, best performance is usually attained by fusing all the available systems, which may lead to high computational costs. In this paper, we aim to discover which technologies combine the best through fusion and to analyse the factors (data, features, modeling methodologies, etc.) that may explain such a good performance. Results are presented and discussed for a number of systems provided by the participating sites and the organizing team of the Albayzin 2010 Language Recognition Evaluation. We hope the conclusions of this work help research groups make better decisions in developing language recognition technology. Luis Javier Rodríguez-Fuentes, Mikel Peñagarikano, Amparo Varona, Mireia Díez, Germán Bordel, David Martínez González, Jesús Villalba 0001, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida, Alberto Abad, Oscar Koller, Isabel Trancoso, Paula Lopez-Otero, Laura Docío Fernández, Carmen García-Mateo, Rahim Saeidi, Mehdi Soufifar, Tomi Kinnunen, Torbjørn Svendsen, Pasi Fränti |
ASRU | 4 |
| 2011 | The Albayzin 2010 Language Recognition EvaluationabstractThe Albayzin 2008 Language Recognition Evaluation was held from May to October 2008, and their results presented and discussed among the participating teams at the 5th Biennial Workshop on Speech Technology [1], organized by the Spanish Network on Speech Technologies [2] in November 2008.In this paper, we present (for the first time) a full description of the Albayzin 2008 LRE and analyze and discuss recognition results.The evaluation was designed according to the test procedures, protocols and performance measures used in the NIST 2007 LRE.The KALAKA database [3], consisting of 16 kHz audio signals recorded from TV broadcasts, was created ad-hoc and used for the evaluation.The four official languages spoken in Spain (Basque, Catalan, Galician and Spanish) were taken as target languages, other (unknown) languages being also recorded to allow open-set verification tests.The best system, employing state-of-the-art technology, yielded C avg = 0, 0552 (around 5% EER) in closed-set verification tests on a set of 30-second segments.This reveals the difficulty of the task, despite using 16 kHz speech signals and having only four target languages.We plan to include also Portuguese and English as target languages for the next Albayzin 2010 LRE. Luis Javier Rodríguez-Fuentes, Mikel Peñagarikano, Amparo Varona, Mireia Díez, Germán Bordel |
INTERSPEECH | 4 |
| 2010 | KALAKA: A TV Broadcast Speech Database for the Evaluation of Language Recognition Systems
Luis Javier Rodríguez-Fuentes, Mikel Peñagarikano, Germán Bordel, Amparo Varona, Mireia Díez |
LREC | 5 |