Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Joaquín González-Rodríguez

dblp:44/4281 · also Joaquin Gonzalez-Rodriguez · DBLP profile ↗
← Back
67ranked-venue papers
12as first author
0since 2021 · last 2018
0000-0003-0910-2575ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 46 · 10 first-authorArtificial intelligence and machine learning · 41 · 8 first-authorSecurity and privacy · 3Human-computer interaction and ubiquitous computing · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
4 papers
Biometric security · 93% Digital forensics and information hiding · 7%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Biometric security › fingerprint recognition
fingerprint quality assessment
0.222008
Fingerprint Image-Quality Estimation and its Application to Multialgorithm Verification · IEEE Trans. Inf. Forensics Secur. 2008
A Comparative Study of Fingerprint Image-Quality Estimation Methods · IEEE Trans. Inf. Forensics Secur. 2007
Biometric security
fingerprint recognition
0.222008
Fingerprint Image-Quality Estimation and its Application to Multialgorithm Verification · IEEE Trans. Inf. Forensics Secur. 2008
A Comparative Study of Fingerprint Image-Quality Estimation Methods · IEEE Trans. Inf. Forensics Secur. 2007
Biometric security
biometric database
0.112010
The Multiscenario Multienvironment BioSecure Multimodal Database (BMDB) · IEEE Trans. Pattern Anal. Mach. Intell. 2010
Biometric security
biometric recognition
0.112010
The Multiscenario Multienvironment BioSecure Multimodal Database (BMDB) · IEEE Trans. Pattern Anal. Mach. Intell. 2010
Biometric security › multi-biometric systems
multimodal biometrics
0.112010
The Multiscenario Multienvironment BioSecure Multimodal Database (BMDB) · IEEE Trans. Pattern Anal. Mach. Intell. 2010
Biometric security › biometric fusion
multimodal biometric fusion
0.112008
Fingerprint Image-Quality Estimation and its Application to Multialgorithm Verification · IEEE Trans. Inf. Forensics Secur. 2008
Biometric security › biometric fusion
score-level fusion
0.112008
Fingerprint Image-Quality Estimation and its Application to Multialgorithm Verification · IEEE Trans. Inf. Forensics Secur. 2008
Biometric security
speaker recognition
0.112007
Emulating DNA: Rigorous Quantification of Evidential Weight in Transparent and Testable Forensic Speaker Recognition · IEEE Trans. Speech Audio Process. 2007
Natural language and speech › Speech recognition and synthesis
speaker recognition
0.012007
Emulating DNA: Rigorous Quantification of Evidential Weight in Transparent and Testable Forensic Speaker Recognition · IEEE Trans. Speech Audio Process. 2007
Biometric security › fingerprint recognition
fingerprint verification
0.012007
A Comparative Study of Fingerprint Image-Quality Estimation Methods · IEEE Trans. Inf. Forensics Secur. 2007

Methods — techniques the papers use, named apart from their topics

symmetry descriptors · 0.2cascaded fusion · 0.2bayes-based fusion · 0.2probabilistic similarity-typicality metric · 0.1likelihood ratio estimation · 0.1multimodal fusion · 0.1baseline evaluation · 0.1quality measure comparison · 0.1
YearPublicationVenuePosition
2018 DNN Based Embeddings for Language Recognition
abstract
In this work, we present a language identification (LID) system based on embeddings. In our case, an embedding is a fixed-length vector (similar to i-vector) that represents the whole utterance, but unlike i-vector it is designed to contain mostly information relevant to the target task (LID). In order to obtain these embeddings, we train a deep neural network (DNN) with sequence summarization layer to classify languages. In particular, we trained a DNN based on bidirectional long short-term memory (BLSTM) recurrent neural network (RNN) layers, whose frame-by-frame outputs are summarized into mean and standard deviation statistics. After this pooling layer, we add two fully connected layers whose outputs correspond to embeddings. Finally, we add a softmax output layer and train the whole network with multi-class cross-entropy objective to discriminate between languages. We report our results on NIST LRE 2015 and we compare the performance of embeddings and corresponding i-vectors both modeled by Gaussian Linear Classifier (GLC). Using only embeddings resulted in comparable performance to i-vectors and by performing score-level fusion we achieved 7.3% relative improvement over the baseline.
Alicia Lozano-Diez, Oldrich Plchot, Pavel Matejka, Joaquín González-Rodríguez
ICASSP4
2016 On the use of deep feedforward neural networks for automatic language identification
abstract
In this work, we present a comprehensive study on the use of deep neural networks (DNNs) for automatic language identification (LID). Motivated by the recent success of using DNNs in acoustic modeling for speech recognition, we adapt DNNs to the problem of identifying the language in a given utterance from its short-term acoustic features. We propose two different DNN-based approaches. In the first one, the DNN acts as an end-to-end LID classifier, receiving as input the speech features and providing as output the estimated probabilities of the target languages. In the second approach, the DNN is used to extract bottleneck features that are then used as inputs for a state-of-the-art i-vector system. Experiments are conducted in two different scenarios: the complete NIST Language Recognition Evaluation dataset 2009 (LRE'09) and a subset of the Voice of America (VOA) data from LRE'09, in which all languages have the same amount of training data. Results for both datasets demonstrate that the DNN-based systems significantly outperform a state-of-art i-vector system when dealing with short-duration utterances. Furthermore, the combination of the DNN-based and the classical i-vector system leads to additional performance improvements (up to 45% of relative improvement in both EER and Cavg on 3s and 10s conditions, respectively).
Ignacio López-Moreno, Javier Gonzalez-Dominguez, David Martinez, Oldrich Plchot, Joaquín González-Rodríguez, Pedro J. Moreno 0001
Comput. Speech Lang.5
2016 Linguistically-constrained formant-based i-vectors for automatic speaker recognition
Javier Franco-Pedroso, Joaquín González-Rodríguez
Speech Commun.2
2015 An end-to-end approach to language identification in short utterances using convolutional neural networks
abstract
In this work, we propose an end-to-end approach to the language identification (LID) problem based on Convolutional Deep Neural Networks (CDNNs). The use of CDNNs is mainly motivated by the ability they have shown when modeling speech signals, and their relatively low-cost with respect to other deep architectures in terms of number of free parameters. We evaluate different configurations in a subset of 8 languages within the NIST Language Recognition Evaluation 2009 Voice of America (VOA) dataset, for the task of short test durations (segments up to 3 seconds of speech). The proposed CDNN-based systems achieve comparable performances to our baseline i-vector system, while reducing drastically the number of parameters to tune (at least 100 times fewer parameters). Then, we combine these CDNN-based systems and the i-vector baseline with a simple fusion at score level. This combination outperforms our best standalone system (up to 11% of relative improvement in terms of EER).
Alicia Lozano-Diez, Rubén Zazo-Candil, Javier Gonzalez-Dominguez, Doroteo T. Toledano, Joaquín González-Rodríguez
INTERSPEECH5
2015 Frame-by-frame language identification in short utterances using deep neural networks
Javier Gonzalez-Dominguez, Ignacio López-Moreno, Pedro J. Moreno 0001, Joaquín González-Rodríguez
Neural Networks4
2014 Automatic language identification using deep neural networks
abstract
This work studies the use of deep neural networks (DNNs) to address automatic language identification (LID). Motivated by their recent success in acoustic modelling, we adapt DNNs to the problem of identifying the language of a given spoken utterance from short-term acoustic features. The proposed approach is compared to state-of-the-art i-vector based acoustic systems on two different datasets: Google 5M LID corpus and NIST LRE 2009. Results show how LID can largely benefit from using DNNs, especially when a large amount of training data is available. We found relative improvements up to 70%, in Cavg, over the baseline system.
Ignacio López-Moreno, Javier Gonzalez-Dominguez, Oldrich Plchot, David Martinez, Joaquín González-Rodríguez, Pedro J. Moreno 0001
ICASSP5
2014 Automatic language identification using long short-term memory recurrent neural networks
abstract
This work explores the use of Long Short-Term Memory (LSTM) recurrent neural networks (RNNs) for automatic lan-guage identification (LID). The use of RNNs is motivated by their better ability in modeling sequences with respect to feed forward networks used in previous works. We show that LSTM RNNs can effectively exploit temporal dependencies in acoustic data, learning relevant features for language discrimination pur-poses. The proposed approach is compared to baseline i-vector and feed forward Deep Neural Network (DNN) systems in the NIST Language Recognition Evaluation 2009 dataset. We show LSTM RNNs achieve better performance than our best DNN system with an order of magnitude fewer parameters. Further, the combination of the different systems leads to significant per-formance improvements (up to 28%). 1.
Javier Gonzalez-Dominguez, Ignacio López-Moreno, Hasim Sak, Joaquín González-Rodríguez, Pedro J. Moreno 0001
INTERSPEECH4
2014 Improving short utterance i-vector speaker verification using utterance variance modelling and compensation techniques
Ahilan Kanagasundaram, David Dean, Sridha Sridharan, Javier Gonzalez-Dominguez, Joaquín González-Rodríguez, Daniel Ramos-Castro
Speech Commun.5
2013 Security evaluation of i-vector based speaker verification systems against hill-climbing attacks
abstract
Proceedings of Interspeech 2013, Lyon (France)
Marta Gomez-Barrero, Javier Gonzalez-Dominguez, Javier Galbally, Joaquín González-Rodríguez
INTERSPEECH4
2013 Improving short utterance based i-vector speaker recognition using source and utterance-duration normalization techniques
abstract
A significant amount of speech is typically required for speaker verification system development and evaluation, especially in the presence of large intersession variability. This paper in-troduces a source and utterance-duration normalized linear dis-criminant analysis (SUN-LDA) approaches to compensate ses-sion variability in short-utterance i-vector speaker verification systems. Two variations of SUN-LDA are proposed where normalization techniques are used to capture source variation from both short and full-length development i-vectors, one based upon pooling (SUN-LDA-pooled) and the other on con-catenation (SUN-LDA-concat) across the duration and source-dependent session variation. Both the SUN-LDA-pooled and SUN-LDA-concat techniques are shown to provide improve-ment over traditional LDA on NIST 08 truncated 10sec-10sec evaluation conditions, with the highest improvement obtained with the SUN-LDA-concat technique achieving a relative im-provement of 8 % in EER for mis-matched conditions and over 3 % for matched conditions over traditional LDA approaches. Index Terms: speaker verification, i-vector, total-variability, LDA, WCCN
Ahilan Kanagasundaram, David Dean, Javier Gonzalez-Dominguez, Sridha Sridharan, Daniel Ramos-Castro, Joaquín González-Rodríguez
INTERSPEECH6
2013 Improving the PLDA based speaker verification in limited microphone data conditions
abstract
A significant amount of speech data is required to develop a robust speaker verification system, but it is difficult to find enough development speech to match all expected conditions.In this paper we introduce a new approach to Gaussian probabilistic linear discriminant analysis (GPLDA) to estimate reliable model parameters as a linearly weighted model taking more input from the large volume of available telephone data and smaller proportional input from limited microphone data.In comparison to a traditional pooled training approach, where the GPLDA model is trained over both telephone and microphone speech, this linear-weighted GPLDA approach is shown to provide better EER and DCF performance in microphone and mixed conditions in both the NIST 2008 and NIST 2010 evaluation corpora.Based upon these results, we believe that linear-weighted GPLDA will provide a better approach than pooled GPLDA, allowing for the further improvement of GPLDA speaker verification in conditions with limited development data.
Ahilan Kanagasundaram, David Dean, Javier Gonzalez-Dominguez, Sridha Sridharan, Daniel Ramos-Castro, Joaquín González-Rodríguez
INTERSPEECH6
2012 A linguistically-motivated speaker recognition front-end through session variability compensated cepstral trajectories in phone units
abstract
In this paper a new linguistically-motivated front-end is presented showing major performance improvements from the use of session variability compensated cepstral trajectories in phone units. Extending our recent work on temporal contours in linguistic units (TCLU), we have combined the potential of those unit-dependent trajectories with the ability of feature domain factor analysis techniques to compensate session variability effects, which has resulted in consistent and discriminant phone-dependent trajectories across different recording sessions. Evaluating with NIST SRE04 English-only 1s1s task, we report EERs as low as 5.40% from the trajectories in a single phone, with 29 different phones producing each of them EERs smaller than 10%, and additionally showing an excellent calibration performance per unit. The combination of different units shows significant complementarity reporting EERs as 1.63% (100×DCF=0.732) from a simple sum fusion of 23 best phones, or 0.68% (100×DCF=0.304) when fusing them through logistic regression.
Joaquín González-Rodríguez, Javier Gonzalez-Dominguez, Javier Franco-Pedroso, Daniel Ramos-Castro
ICASSP1
2011 Calibration and weight of the evidence by human listeners. The ATVS-UAM submission to NIST HUMAN-aided speaker recognition 2010
abstract
This work analyzes the performance of speaker recognition when carried out by human lay listeners. In forensics, judges and jurors usually manifest intuition that people is proficient to distinguish other people from their voices, and there fore opinions are easily elicited about speech evidence just by listening to it, or by means of panels of listeners. There is a danger, however, since little attention has been paid to scientifically measure the performance of human listeners, as well as to the strength with which they should elicit their opinions. In this work we perform such a rigorous analysis in the context of NIST Human-Aided Speaker Recognition 2010 (HASR). We have recruited a panel of listeners who have elicited opinions in the form of scores. Then, we have calibrated such scores using a development set, in order to generate calibrated likelihood ratios. Thus, the discriminating power and the strength with which human lay listeners should express their opinions about the speech evidence can be assessed, giving a measure of the amount of information given by human listeners to the speaker recognition process.
Daniel Ramos-Castro, Javier Franco-Pedroso, Joaquín González-Rodríguez
ICASSP3
2011 Speaker Recognition Using Temporal Contours in Linguistic Units: The Case of Formant and Formant-Bandwidth Trajectories
abstract
Proceedings of Interspeech 2011, Florence (Italy)
Joaquín González-Rodríguez
INTERSPEECH1
2011 Von Mises-Fisher Models in the Total Variability Subspace for Language Recognition
abstract
This letter proposes a new modeling approach for the Total Variability subspace within a Language Recognition task. Motivated by previous works in directional statistics, von Mises-Fisher distributions are used for assigning language-conditioned probabilities to language data, assumed to be spherically distributed in this subspace. The two proposed methods use Kernel Density Functions or Finite Mixture Models of such distributions. Experiments conducted on NIST LRE 2009 show that the proposed techniques significantly outperform the baseline cosine distance approach in most of the considered experimental conditions, including different speech conditions, durations and the presence of unseen languages.
Ignacio López-Moreno, Daniel Ramos-Castro, Javier Gonzalez-Dominguez, Joaquín González-Rodríguez
IEEE Signal Process. Lett.4
2010 Score-level compensation of extreme speech duration variability in speaker verification
Sergio Perez-Gomez, Daniel Ramos-Castro, Javier Gonzalez-Dominguez, Joaquín González-Rodríguez
INTERSPEECH4
2010 BiosecurID: a multimodal biometric database
Julian Fierrez, Javier Galbally, Javier Ortega-Garcia, Manuel R. Freire, Fernando Alonso-Fernandez, Daniel Ramos-Castro, Doroteo T. Toledano, Joaquín González-Rodríguez, Juan A. Sigüenza, Javier Garrido Salas
Pattern Anal. Appl.8
2010 The Multiscenario Multienvironment BioSecure Multimodal Database (BMDB)
abstract
A new multimodal biometric database designed and acquired within the framework of the European BioSecure Network of Excellence is presented. It is comprised of more than 600 individuals acquired simultaneously in three scenarios: 1) over the Internet, 2) in an office environment with desktop PC, and 3) in indoor/outdoor environments with mobile portable hardware. The three scenarios include a common part of audio/video data. Also, signature and fingerprint data have been acquired both with desktop PC and mobile portable hardware. Additionally, hand and iris data were acquired in the second scenario using desktop PC. Acquisition has been conducted by 11 European institutions. Additional features of the BioSecure Multimodal Database (BMDB) are: two acquisition sessions, several sensors in certain modalities, balanced gender and age distributions, multimodal realistic scenarios with simple and quick tasks per modality, cross-European diversity, availability of demographic data, and compatibility with other multimodal databases. The novel acquisition conditions of the BMDB allow us to perform new challenging research and evaluation of either monomodal or multimodal biometric systems, as in the recent BioSecure Multimodal Evaluation campaign. A description of this campaign including baseline results of individual modalities from the new database is also given. The database is expected to be available for research purposes through the BioSecure Association during 2008.
Javier Ortega-Garcia, Julian Fierrez, Fernando Alonso-Fernandez, Javier Galbally, Manuel R. Freire, Joaquín González-Rodríguez, Carmen García-Mateo, José Luis Alba-Castro, Elisardo González-Agulla, Enrique Otero Muras, Sonia Garcia-Salicetti, Lorène Allano, Van-Bao Ly, Bernadette Dorizzi, Josef Kittler, Thirimachos Bourlai, Norman Poh, Farzin Deravi, Ming W. R. Ng, Michael C. Fairhurst, Jean Hennebert, Andreas Humm, Massimo Tistarelli, Linda Brodo, Jonas Richiardi, Andrzej Drygajlo, Harald Ganster, Federico Sukno, Sri-Kaushik Pavani, Alejandro F. Frangi, Lale Akarun, Arman Savran
IEEE Trans. Pattern Anal. Mach. Intell.6
2010 Quality-Based Conditional Processing in Multi-Biometrics: Application to Sensor Interoperability
abstract
As biometric technology is increasingly deployed, it will be common to replace parts of operational systems with newer designs. The cost and inconvenience of reacquiring enrolled users when a new vendor solution is incorporated makes this approach difficult and many applications will require to deal with information from different sources regularly. These interoperability problems can dramatically affect the performance of biometric systems and thus, they need to be overcome. Here, we describe and evaluate the ATVS-UAM fusion approach submitted to the quality-based evaluation of the 2007 BioSecure Multimodal Evaluation Campaign, whose aim was to compare fusion algorithms when biometric signals were generated using several biometric devices in mismatched conditions. Quality measures from the raw biometric data are available to allow system adjustment to changing quality conditions due to device changes. This system adjustment is referred to as quality-based conditional processing. The proposed fusion approach is based on linear logistic regression, in which fused scores tend to be log-likelihood-ratios. This allows the easy and efficient combination of matching scores from different devices assuming low dependence among modalities. In our system, quality information is used to switch between different system modules depending on the data source (the sensor in our case) and to reject channels with low quality data during the fusion. We compare our fusion approach to a set of rule-based fusion schemes over normalized scores. Results show that the proposed approach outperforms all the rule-based fusion schemes. We also show that with the quality-based channel rejection scheme, an overall improvement of 25% in the equal error rate is obtained.
Fernando Alonso-Fernandez, Julian Fierrez, Daniel Ramos-Castro, Joaquín González-Rodríguez
IEEE Trans. Syst. Man Cybern. Part A4
2009 Forensic speaker recognition using traditional features comparing automatic and human-in-the-loop formant tracking
abstract
In this paper we compare forensic speaker recognition with traditional features using two different formant tracking strategies: one performed automatically and one semi-automatic performed by human experts.The main contribution of the work is the use of an automatic method for formant tracking, which allows a much faster recognition process and the use of a much higher amount of data for modelling background population, calibration, etc.This is especially important in likelihood-ratiobased forensic speaker recognition, where the variation of features among a population of speakers must be modelled in a statistically robust way.Experiments show that, although recognition using the human-in-the-loop approach is better than using the automatic scheme, the performance of the latter is also acceptable.Moreover, we present a novel feature selection method which allows the analysis of which feature of each formant has a greater contribution to the discriminating power of the whole recognition process, which can be used by the expert in order to decide which features in the available speech material are important.
Alberto de Castro, Daniel Ramos-Castro, Joaquín González-Rodríguez
INTERSPEECH3
2009 Speaker dependent emotion recognition using prosodic supervectors
abstract
Proceedings of Interspeech 2009, Brighton (United Kingdom)
Ignacio López-Moreno, Carlos Ortego-Resa, Joaquín González-Rodríguez, Daniel Ramos-Castro
INTERSPEECH3
2008 Forensic automatic speaker recognition: fiction or science?
Joaquín González-Rodríguez
INTERSPEECH1
2008 Anchor-model fusion for language recognition
abstract
Proceedings of Interspeech 2008, Brisbane (Australia)
Ignacio López-Moreno, Daniel Ramos-Castro, Joaquín González-Rodríguez, Doroteo T. Toledano
INTERSPEECH3
2008 Addressing database mismatch in forensic speaker recognition with Ahumada III: a public real-casework database in Spanish
abstract
Proceedings of Interspeech 2008, Brisbane (Australia)
Daniel Ramos-Castro, Joaquín González-Rodríguez, Javier Gonzalez-Dominguez, Jose Juan Lucena-Molina
INTERSPEECH2
2008 MAP and sub-word level t-norm for text-dependent speaker recognition
abstract
Proceedings of Interspeech 2008, Brisbane (Australia)
Doroteo T. Toledano, Daniel Hernández López, Cristina Esteve-Elizalde, Joaquín González-Rodríguez, Rubén Fernández Pozo, Luis A. Hernández Gómez
INTERSPEECH4
2008 BioSec Multimodal Biometric Database in Text-Dependent Speaker Recognition
Doroteo T. Toledano, Daniel Hernández López, Cristina Esteve-Elizalde, Julian Fierrez, Javier Ortega-Garcia, Daniel Ramos-Castro, Joaquín González-Rodríguez
LREC7
2008 Fingerprint Image-Quality Estimation and its Application to Multialgorithm Verification
abstract
Signal-quality awareness has been found to increase recognition rates and to support decisions in multisensor environments significantly. Nevertheless, automatic quality assessment is still an open issue. Here, we study the orientation tensor of fingerprint images to quantify signal impairments, such as noise, lack of structure, blur, with the help of symmetry descriptors. A strongly reduced reference is especially favorable in biometrics, but less information is not sufficient for the approach. This is also supported by numerous experiments involving a simpler quality estimator, a trained method (NFIQ), as well as the human perception of fingerprint quality on several public databases. Furthermore, quality measurements are extensively reused to adapt fusion parameters in a monomodal multialgorithm fingerprint recognition environment. In this study, several trained and nontrained score-level fusion schemes are investigated. A Bayes-based strategy for incorporating experts' past performances and current quality conditions, a novel cascaded scheme for computational efficiency, besides simple fusion rules, is presented. The quantitative results favor quality awareness under all aspects, boosting recognition rates and fusing differently skilled experts efficiently as well as effectively (by training).
Hartwig Fronthaler, Klaus Kollreider, Josef Bigün, Julian Fierrez, Fernando Alonso-Fernandez, Javier Ortega-Garcia, Joaquín González-Rodríguez
IEEE Trans. Inf. Forensics Secur.7
2007 Information-theoretical comparison of likelihood ratio methods of forensic evidence evaluation
abstract
Forensic evidence in the form of two-level hierarchical multivariate continuous data is modelled using a likelihood ratio approach. Data are available from fragments of glass and of paint. Cross-entropy is used to compare the results with a neutral method and a method using the correct answers.
Daniel Ramos-Castro, Joaquín González-Rodríguez, Grzegorz Zadora, Janina Zieba-Palus, Colin Aitken
IAS2
2007 Support vector regression for speaker verification
abstract
Proceedings of Interspeech 2007, Antwerp (Belgium)
Ignacio López-Moreno, Ismael Mateos-Garcia, Daniel Ramos-Castro, Joaquín González-Rodríguez
INTERSPEECH4
2007 Improved language recognition using better phonetic decoders and fusion with MFCC and SDC features
abstract
Proceedings of Interspeech 2007, Antwerp (Belgium)
Doroteo T. Toledano, Javier Gonzalez-Dominguez, Alejandro Abejón, Danilo Spada, Ismael Mateos-Garcia, Joaquín González-Rodríguez
INTERSPEECH6
2007 Biosec baseline corpus: A multimodal biometric database
Julian Fierrez, Javier Ortega-Garcia, Doroteo T. Toledano, Joaquín González-Rodríguez
Pattern Recognit.4
2007 HMM-based on-line signature verification: Feature extraction and signature modeling
Julian Fierrez, Javier Ortega-Garcia, Daniel Ramos-Castro, Joaquín González-Rodríguez
Pattern Recognit. Lett.4
2007 Speaker verification using speaker- and test-dependent fast score normalization
Daniel Ramos-Castro, Julian Fierrez, Joaquín González-Rodríguez, Javier Ortega-Garcia
Pattern Recognit. Lett.3
2007 Emulating DNA: Rigorous Quantification of Evidential Weight in Transparent and Testable Forensic Speaker Recognition
abstract
Forensic DNA profiling is acknowledged as the model for a scientifically defensible approach in forensic identification science, as it meets the most stringent court admissibility requirements demanding transparency in scientific evaluation of evidence and testability of systems and protocols. In this paper, we propose a unified approach to forensic speaker recognition (FSR) oriented to fulfil these admissibility requirements within a framework which is transparent, testable, and understandable, both for scientists and fact-finders. We show how the evaluation of DNA evidence, which is based on a probabilistic similarity-typicality metric in the form of likelihood ratios (LR), can also be generalized to continuous LR estimation, thus providing a common framework for phonetic-linguistic methods and automatic systems. We highlight the importance of calibration, and we exemplify with LRs from diphthongal F-pattern, and LRs in NIST-SRE06 tasks. The application of the proposed approach in daily casework remains a sensitive issue, and special caution is enjoined. Our objective is to show how traditional and automatic FSR methodologies can be transparent and testable, but simultaneously remain conscious of the present limitations. We conclude with a discussion on the combined use of traditional and automatic approaches and current challenges for the admissibility of speech evidence.
Joaquín González-Rodríguez, P. Rose, Daniel Ramos-Castro, Doroteo T. Toledano, Javier Ortega-Garcia
IEEE Trans. Speech Audio Process.1
2007 A Comparative Study of Fingerprint Image-Quality Estimation Methods
abstract
One of the open issues in fingerprint verification is the lack of robustness against image-quality degradation. Poor-quality images result in spurious and missing features, thus degrading the performance of the overall system. Therefore, it is important for a fingerprint recognition system to estimate the quality and validity of the captured fingerprint images. In this work, we review existing approaches for fingerprint image-quality estimation, including the rationale behind the published measures and visual examples showing their behavior under different quality conditions. We have also tested a selection of fingerprint image-quality estimation algorithms. For the experiments, we employ the BioSec multimodal baseline corpus, which includes 19 200 fingerprint images from 200 individuals acquired in two sessions with three different sensors. The behavior of the selected quality measures is compared, showing high correlation between them in most cases. The effect of low-quality samples in the verification performance is also studied for a widely available minutiae-based fingerprint matching system.
Fernando Alonso-Fernandez, Julian Fierrez, Javier Ortega-Garcia, Joaquín González-Rodríguez, Hartwig Fronthaler, Klaus Kollreider, Josef Bigün
IEEE Trans. Inf. Forensics Secur.4
2006 Using quality measures for multilevel speaker recognition
Daniel Garcia-Romero, Julian Fierrez, Joaquín González-Rodríguez, Javier Ortega-Garcia
Comput. Speech Lang.3
2006 Robust estimation, interpretation and assessment of likelihood ratios in forensic speaker recognition
Joaquín González-Rodríguez, Andrzej Drygajlo, Daniel Ramos-Castro, Marta Garcia-Gomar, Javier Ortega-Garcia
Comput. Speech Lang.1
2005 On the relationship between phonetic modeling precision and phonetic speaker recognition accuracy
abstract
Proceedings of Interspeech-Eurospeech 2005, Lisbon (Portugal)
Doroteo T. Toledano, Carlos Fombella, Joaquín González-Rodríguez, Luis A. Hernández Gómez
INTERSPEECH3
2005 Bayesian adaptation for user-dependent multimodal biometric authentication
Julian Fierrez, Daniel Garcia-Romero, Javier Ortega-Garcia, Joaquín González-Rodríguez
Pattern Recognit.4
2005 Discriminative multimodal biometric authentication based on quality measures
Julian Fierrez, Javier Ortega-Garcia, Joaquín González-Rodríguez, Josef Bigün
Pattern Recognit.3
2005 Adapted user-dependent multimodal biometric authentication exploiting general information
Julian Fierrez, Daniel Garcia-Romero, Javier Ortega-Garcia, Joaquín González-Rodríguez
Pattern Recognit. Lett.4
2005 Target dependent score normalization techniques and their application to signature verification
abstract
Score normalization methods in biometric verification, which encompass the more traditional user-dependent decision thresholding techniques, are reviewed from a test hypotheses point of view. These are classified into test dependent and target dependent methods. The focus of the paper is on target dependent score normalization techniques, which are further classified into impostor-centric, target-centric, and target-impostor methods. These are applied to an on-line signature verification system on signature data from the First International Signature Verification Competition (SVC 2004). In particular, a target-centric technique based on the cross-validation procedure provides the best relative performance improvement testing both with skilled (19%) and random forgeries (53%) as compared to the raw verification performance without score normalization (7.14% and 1.06% Equal Error Rate for skilled and random forgeries, respectively).
Julian Fierrez, Javier Ortega-Garcia, Joaquín González-Rodríguez
IEEE Trans. Syst. Man Cybern. Part C3
2004 Exploiting general knowledge in user-dependent fusion strategies for multimodal biometric verification
abstract
A novel strategy for combining general and user-dependent knowledge in a multimodal biometric verification system is presented. It is based on SVM classifiers and trade-off coefficients introduced in the standard SVM training problem. Experiments are reported on a bimodal biometric system based on fingerprint and on-line signature traits. A comparison between three fusion strategies, namely user-independent, user-dependent and the proposed adapted user-dependent, is carried out. As a result, the suggested approach outperforms the former ones. In particular, a highly remarkable relative improvement of 68% in the EER with respect to the user-independent approach is achieved. The severe and very common problem of training data scarcity in the user-dependent strategy is also relaxed by the proposed scheme, resulting in a relative improvement of 40% in the EER compared to the raw user-dependent strategy.
Julian Fierrez, Daniel Garcia-Romero, Javier Ortega-Garcia, Joaquín González-Rodríguez
ICASSP (5)4
2003 Support vector machine fusion of idiolectal and acoustic speaker information in Spanish conversational speech
abstract
This paper proposes a support vector machine (SVM) based combining scheme that incorporates ideolectal and acoustic characteristics for speaker recognition. Two statistical model paradigms, namely GMM for acoustic modeling and bigrams for language modeling, provide multilevel speaker information that affords a better classification performance when SVM-based fusion is accomplished. This combining approach is useful for all speaker recognition tasks where a considerable amount of data is available. Motivated by the absence of Spanish databases that made feasible our research experiments, more than nine hours of Spanish conversational speech was collected and manually transcribed from broadcasted radio talk shows.
Daniel Garcia-Romero, Julian Fierrez, Joaquín González-Rodríguez, Javier Ortega-Garcia
ICASSP (2)3
2003 Forensic identification reporting using automatic speaker recognition systems
abstract
We show how any speaker recognition system can be adapted to provide its results according to the Bayesian approach for evidence analysis and forensic reporting. This approach, firmly established in other forensic areas as fingerprint, DNA or fiber analysis, suits the needs of both the court and the forensic scientist. We show the inadequacy of the classical approach to forensic reporting because of the use of thresholds and the suppression of the prior probabilities related to the case. We also show how to assess the performance of those forensic systems through Tippet plots. Finally, an example is shown using NIST-Ahumada eval'2001 data, where the speaker recognition abilities of our system are assessed through DET plots, using then these raw scores as evidences into the forensic system, where relative to populations we will obtain the corresponding likelihood ratios values, which are assessed through Tippet (1968) plots.
Joaquín González-Rodríguez, Julian Fierrez, Javier Ortega-Garcia
ICASSP (2)1
2003 A real-time auditory-based microphone array assessed with E-RASTI evaluation proposal
abstract
In this paper, a real time nested microphone array based on the auditory properties of the human ear is presented. Three different stages in the development of the system are described. Firstly, the design of the new auditory-based microphone array is presented, obtaining better noise reduction using the masking properties of the human auditory system. Secondly, we show its validation through a new method called E-RASTI based in the well-known RASTI (Rapid STI - speech transmission index) intelligibility estimator. In addition to classical enhancement estimators as SNR, NMR (noise to masked ratio) or AI (articulation index), E-RASTI is proposed and validated for dereverberation assessment, used here with real speech signals and not with speech-like signals as in the original RASTI method. And finally as third stage, the real time implementation of this highly computing-demanding algorithm through the use of a DSP-based architecture based on the recent floating point TMS320C6701 is described.
José-Luis Sánchez-Bote, Joaquín González-Rodríguez, Javier Ortega-Garcia
ICASSP (5)2
2003 A comparative evaluation of global representation-based schemes for face verification
abstract
This paper is focused on algorithmic issues for biometric face verification (i.e., given an image of the face and an identity claim, decide whether they correspond to each other or not). Several alternatives for geometric normalization of images, photometric normalization, dimensionality reduction and similarity measures are proposed and compared using the XM2VTS database and the associated Lausanne protocol [K. Messer et al., 1999], [J. Luettin et al., 1998]. Experiments under this particular framework show that best verification results are obtained when holistic approaches for face recognition (such as eigenfaces or fisherfaces) are combined with techniques traditionally associated to local feature-based approaches, such as Gabor decompositions.
Julian Fierrez, S. Cruz-Llana, Javier Ortega-Garcia, Joaquín González-Rodríguez
ICIP (3)4
2003 Minutiae-based enhanced fingerprint verification assessment relaying on image quality factors
abstract
In this paper we evaluate authentication performance of the minutiae-based fingerprint automatic recognition system, previously proposed D. Simon-Zorita, et al. (2001) and recently completed, with the new large fingerprint image database, MCYT J. Ortega-Garcia, et al. (2002). The scheme includes: image enhancement, characteristic extraction and pattern recognition. The design of this database permits to analyse the influence of the variability factors appearing in the image acquisition phase. We focus in two factors: the finger position over the acquisition sensor, and the quality of the acquired fingerprint images. The analysis is accomplished in cases of supervised and nonsupervised databases. Score normalization is presented as an effective technique to improve the fingerprint verification system performance, yielding highly competitive EERs.
Danilo Simon-Zorita, Javier Ortega-Garcia, Marta Sanchez-Asenjo, Joaquín González-Rodríguez
ICIP (2)4
2003 Fusion strategies in multimodal biometric verification
abstract
The aim of this paper, regarding multimodal biometric verification, is twofold: on one hand, to compare experimentally a selection of them using as monomodal baseline systems as our template-based face, minutiae-based fingerprint and HMM-based on-line signature verification systems on the MCYT multimodal database. A new strategy is proposed and discussed in order to compute a multimodal combined score by means of support vector machine (SVM) classifiers.
Julian Fierrez, Javier Ortega-Garcia, Joaquín González-Rodríguez
ICME3
2003 Support vector machine fusion of idiolectal and acoustic speaker information in Spanish conversational speech
abstract
This paper proposed a support vector machine (SVM) based combining scheme that incorporates idiolectal and acoustic characteristics for speaker recognition. Two statistical model paradigms, namely GMM for acoustic modeling and bigrams for language modeling, provide multilevel speaker information that affords a better classification performance when SVM-based fusion is accomplished. This combining approach is useful for all speaker recognition tasks where a considerable amount of data is available. Motivated by the absence of Spanish databases that made feasible our research experiments, more than nine hours of Spanish conversational speech was collected and manually transcribed from broadcasted radio talk shows.
Daniel Garcia-Romero, Julian Fierrez, Joaquín González-Rodríguez, Javier Ortega-Garcia
ICME3
2003 Robust likelihood ratio estimation in Bayesian forensic speaker recognition
Joaquín González-Rodríguez, Daniel Garcia-Romero, Marta Garcia-Gomar, Daniel Ramos-Castro, Javier Ortega-Garcia
INTERSPEECH1
2002 A Multilingual Speaker Verification System: Architecture and Performance Evaluation
Francisco Javier Caminero Gil, Joaquín González-Rodríguez, Javier Ortega-Garcia, Daniel Tapias Merino, Pedro M. Ruz, Mercedes Solá
LREC2
2001 Minutiae extraction scheme for fingerprint recognition systems
abstract
A complete minutiae extraction scheme for automatic fingerprint recognition systems is presented. The proposed method uses improving alternatives for the image enhancement process, leading consequently to an increase in the reliability in the minutiae extraction task. In the first stages, image normalization and the orientation field of the fingerprint are calculated. The local orientation of the ridges serve as parameter for the next processing stages. Details of the adaptive morphological filtering used for ridge extraction and background noise elimination are described. Evaluation results are obtained from both inked and scanned fingerprints. Conclusions in terms of Goodness Index (GI), which compares the results obtained by automatic minutiae extraction with manually extracted ones, are provided in order to test the global performance of this approach.
Danilo Simon-Zorita, Javier Ortega-Garcia, Santiago Cruz-Llanas, Joaquín González-Rodríguez
ICIP (3)4
2001 A new auditory based microphone array and objective evaluation using e-RASTI
abstract
Two are the goals of the work presented in this paper. The first one is the implementation of a new method of speech enhancement using microphone arrays. This method gets noise reduction of speech signal using the masking properties of the human auditory system. The second goal of the paper is to use RASTI index (RApid Speech Transmission Index) for objective evaluation of speech signal quality through E-RASTI evaluation. What is new is that E-RASTI is applied to speech signals and no to RASTI-like signals. The E-RASTI index is specially suited to test reverberant speech and has been used here to evaluate the reverberation reduction produced by a microphone array based on all-pass and minimum-phase decomposition or multichannel liftering. Noise reduction evaluation has been performed with the E-RASTI index and also with more traditional methods, based on Signal to Noise Ratios (SNR). Results have demonstrated the good performance of the noise suppressor and the E-RASTI objective quality evaluator.
José-Luis Sánchez-Bote, Joaquín González-Rodríguez, Danilo Simon-Zorita
INTERSPEECH2
2000 Speech dereverberation and noise reduction with a combined microphone array approach
abstract
In this contribution we have addressed the problem of speech enhancement in noisy and reverberant rooms through the use of a new approach that combines the dereverberation abilities of a structure based in the separate processing of the minimum-phase and all-pass components of the input speech signals, and the noise rejection performance of a speech-activity-based Wiener filter able to cope both with coherent and diffuse noise. Experiments have been performed with the CMU real multichannel database, which includes a clean speech reference through a head-mounted microphone. This reference signal have been also used to perform simulation experiments in controlled conditions. Extensive results have been obtained, both with log area ratio and cepstral distances of input and processed signals to the reference, and with segSNR improvements, assessing the abilities of the new system to cope both with reverberation and coherent and diffuse noise in different acoustic environments.
Joaquín González-Rodríguez, José-Luis Sánchez-Bote, Javier Ortega-Garcia
ICASSP1
2000 Phonetic consistency in Spanish for pin-based speaker verification system
Javier Ortega-Garcia, Joaquín González-Rodríguez, Daniel Tapias Merino
INTERSPEECH2
2000 AHUMADA: A large speech corpus in Spanish for speaker characterization and identification
Javier Ortega-Garcia, Joaquín González-Rodríguez, Victoria Marrero-Aguiar
Speech Commun.2
1999 Concurrent speakers separation through binaural processing of stereo recordings
Joaquín González-Rodríguez, Santiago Cruz-Llanas, Javier Ortega-Garcia
EUROSPEECH1
1999 Facing severe channel variability in forensic speaker verification conditions
abstract
This paper proposes a distinction between existing multilingual synthesis systems and mixed-lingual or polyglot synthesis systems. The latter should be capable of synthesising with the same voice utterances which contain foreign language words or word groups. As a first step towards polyglot synthetic speech, the design and realisation of a 4-lingual single-speaker diphone inventory is detailed. The first results show that mixedlingual sentences can be synthesised using this inventory. Further work will focus on multilingual text analysis and prosodic modelling in order to create a complete polyglot TTS system.
Javier Ortega-Garcia, Santiago Cruz-Llanas, Joaquín González-Rodríguez
EUROSPEECH3
1998 AHUMADA: a large speech corpus in Spanish for speaker identification and verification
abstract
Speaker recognition is a major task when security applications through speech input are needed. Regarding speaker identity, several factors of variability must be considered: (a) factors concerning peculiar intra-speaker variability (manner of speaking, inter-session variability, dialectal variations, emotional condition, etc.) or forced intra-speaker variability (Lombard effect, cocktail-party effect), and (b) factors depending on external influences (kind of microphone, channel effects, noise, reverberation, etc). To cope with all these variability sources, a specific speech database called AHUMADA has been designed and collected for speaker recognition tasks in Castilian Spanish. AHUMADA incorporates six different recording sessions, including both in situ and telephone speech recordings. A total of 104 male speakers uttered isolated digits, digit strings, phonologically balanced short utterances, phonologically and syllabically balanced read text and more than one minute of spontaneous speech, so about 15 GB of speech material is available. Speaker verification results, concerning the available variability sources are also presented.
Javier Ortega-Garcia, Joaquín González-Rodríguez, Victoria Marrero-Aguiar, Juan J. Díaz-Gómez, Ramon Garcia-Jimenez, Jose Juan Lucena-Molina, José A. G. Sanchez-Molero
ICASSP2
1998 Coherence-based subband decomposition for robust speech and speaker recognition in noisy and reverberant rooms
abstract
In this paper, the acoustic characteristics of sound fields in enclosed rooms are studied in the joint presence of speech and noise, in order to design a broadband microphone array system capable of coping with both coherent and diffuse noises. Several state-of-the-art speech enhancement array structures are presented and compared to our new system in terms of correct word recognition rates in a simple command and control task. The proposed structure, based on a broadband subband-nested array, performs real-time estimations of the spatial coherence in order to determine the coherent/diffuse nature of the different subbands, using different filters in each case, improving also the classical Wiener post-filter, typically used for diffuse noise supression, for proper cancellation of coherent noises. The results obtained with a 15-channel simultaneous recording database in different reverberation and noise conditions show better performance than other structures previously proposed.
Joaquín González-Rodríguez, Santiago Cruz-Llanas, Javier Ortega-Garcia
ICSLP1
1998 Quantitative influence of speech variability factors for automatic speaker verification in forensic tasks
abstract
Regarding speaker identity in forensic conditions, several factors of variability must be taken into account, as peculiar intra-speaker variability, forced intra-speaker variability or channel-dependent external influences. Using ‘AHUMADA’ large speech database in Spanish, containing several recording sessions and channels, and including different tasks for 100 male speakers, automatic speaker verification experiments are accomplished. Due to the inherent non-cooperative nature of speakers in forensic applications, only text-independent recognizers are likely to be used. In this sense, a GMM-based verification system has been used in order to obtain quantitative results. Maximum likelihood estimation of the models is performed, and LPC-cepstra, deltaand delta-delta-LPCC, are used at the parameterization stage. With this baseline verification system, we intend to determine how some variability sources included in ‘AHUMADA’ affect speaker identification. Results including speaking rate influence, singleand multi-session training and cross-channel testing are presented when likelihood-domain normalization is applied.
Javier Ortega-Garcia, Santiago Cruz-Llanas, Joaquín González-Rodríguez
ICSLP3
1998 Speaker Recognition-Oriented 'AHUMADA' Large Speecb Corpus
Javier Ortega-Garcia, Victoria Marrero-Aguiar, Joaquín González-Rodríguez, José Javier Díaz Gómez, Ramon Garcia-Jimenez, Jose Juan Lucena-Molina, José A. G. Sanchez-Molero
LREC3
1997 Robust speaker recognition through acoustic array processing and spectral normalization
abstract
The development of a robust speaker recognition system obtained through the joint use of acoustic array processing and spectral normalization as input to a Gaussian mixture model speaker recognition system is described. Results obtained with these techniques have been reported previously by the authors, but operational problems appear if extensive testing with different configurations and testing conditions are intended. We describe an open system that has been developed to cope with this problem. The number and geometry of the microphones, the time delay estimation method, the array processing structure and the spectral normalization technique together with the room size, noise type and SNR are some of the options that can be easily changed. It will also allow testing with real multichannel databases and any new algorithm can easily be incorporated to the system.
Joaquín González-Rodríguez, Javier Ortega-Garcia
ICASSP1
1997 Providing single and multi-channel acoustical robustness to speaker identification systems
abstract
Acoustical mismatch between training and testing phases induces degradation of performance in automatic speaker recognition systems. Providing robustness to speaker recognizers has to be, therefore, a priority matter. Robustness in the acoustical stage can be accomplished through speech enhancement techniques as a prior stage to the recognizer. These techniques are oriented to the reduction of the impact that acoustical noise produces on the input signal. In this paper, several spectral subtraction-derived techniques are used to enhance single-channel noisy speech. Other perspectives, based in dual-channel (adaptive filtering) and multi-channel (microphone arrays) processing are also presented as optimal solutions to speech enhancement needs. A comparative analysis of the proposed techniques, with different types of noise at different SNRs, as a pre-processing stage to an ergodic HMM-based speaker recognizer, is presented.
Javier Ortega-Garcia, Joaquín González-Rodríguez
ICASSP2
1996 Increasing robustness in GMM speaker recognition systems for noisy and reverberant speech with low complexity microphone arrays
abstract
In this paper we describe the additive robustness obtained through the combined use of a first acoustic processing step based on a low complexity microphone array, followed by a spectral normalization step.Microphone arrays have shown to provide good results in reducing different sources of acoustic degradation.However, microphone arrays produce linear filtering effects that need to be compensated in order to obtain a minimal spectral distortion.In this contribution we will present the combination of a microphone array together with different well known spectral normalization techniques as preprocessing stages to a Gaussian Mixture Models (GMM) based text-independent speaker recognition system.We will show that the combination of these extensively used techniques in the fields of speech enhancement and robust speaker recognition respectively, greatly improves the results obtained when the system is tested in noisy reverberant environments with short utterances from unconstrained conversational speech.
Joaquín González-Rodríguez, Javier Ortega-Garcia, César Martin
ICSLP1
1996 Overview of speech enhancement techniques for automatic speaker recognition
abstract
Real world conditions differ from ideal or laboratory conditions, causing mismatch between training and testing phases, and consequently, inducing performance degradation in automatic speaker recognition systems [1].Many strategies have been adopted to cope with acoustical degradation; in some applications of speaker identification systems a clean sample of speech, prior to the recognition stage, is needed.This has justified the use of procedures that may reduce the impact of acoustical noise on the desired signal, giving rise to techniques involved in the enhancement of noisy speech [2,9].In this paper, a comparative performance analysis of singlechannel (based in classical spectral subtraction and some derived alternatives), dual-channel (based in adaptive noise cancelling) and multi-channel (using microphone arrays) speech enhancement techniques, with different types of noise at different SNRs, as a pre-processing stage to an ergodic HMMbased speaker recognizer, is presented.
Javier Ortega-Garcia, Joaquín González-Rodríguez
ICSLP2