EDBT 2026 Demo / reviewers in the wild / expert
M. Carmen Benítez
dblp:80/1091 · also M. Carmen Benítez Ortúzar, Maria C. Benitez
· DBLP profile ↗
57ranked-venue papers
5as first author
6since 2021 · last 2023
0000-0002-5407-8335ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 4 first-authorArtificial intelligence and machine learning · 20 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Classification & Clustering of Volcano-Seismic Events Using Supervised and Unsupervised MethodsabstractThis study looks into unsupervised and supervised methods for classifying and clustering volcano-seismic time series data. Due to the resource intensive labelling process required to understand volcano-seismic signals it is important to explore unsupervised analysis techniques in this domain. This study offers a comparison between the performance of supervised and unsupervised methods and offers insight into their potential for frame by frame classification of volcano-seismic signals. Joe Carthy, Manuel Titos, Carlos Martínez Clemente, M. Carmen Benítez |
IGARSS | 4 |
| 2023 | Generating a Mobility-Pattern Database for Urban Traffic Monitoring Using Distributed Acoustic SensingabstractThis work presents a methodology to generate a mobility-pattern database for urban traffic monitoring through distributed acoustic sensing. Registers from the continuous monitoring of a live experimental testbed provide three canonical types of events (buses, cars and pedestrians). Their separability for automatic detection and classification is inspected through PCA analysis of their approximate entropy and Hjorth parameters. An automatic detection and labeling procedure is presented and used to generate a database of labeled files accessible and usable for future machine learning approaches. Luz García 0001, Manuel Titos, Joe Carthy, José Camacho 0001, Sonia Mota, M. Carmen Benítez |
IGARSS | 7 |
| 2022 | Continuous Active Learning for Seismo-Volcanic MonitoringabstractDeep learning has advanced seismo-volcanic monitoring to unprecedented performance levels. Nevertheless, seismic data labeling still requires substantial annotation efforts, often delayed in time if the eruptive state alters the data conditions. The selective segmentation of which earthquake transients have to be reviewed by an expert can significantly reduce annotation time, speed up algorithmic training, and boost monitoring adaptability to unforeseen situations. In this work, we propose a Bayesian temporal convolutional neural network (B-TCN) to perform continuous detection and classification while extracting the most uncertain events from the continuous data stream. Formulated as an active learning (AL) procedure, our B-TCN outputs an uncertainty map over time, highlighting the class memberships that are needed to be reviewed. We attain a significant improvement in monitoring metrics, with only a fraction of the initial dataset to achieve a recognition performance of 83% for four seismo-volcanic events. Ángel Bueno, Manuel Titos, M. Carmen Benítez, Jesús Ibáñez 0003 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Recurrent Scattering Network Detects Metastable Behavior in Polyphonic Seismo-Volcanic Signals for Volcano Eruption ForecastingabstractWe introduce an end-to-end (E2E) deep neural network architecture designed to perform seismo-volcanic monitoring focused on detecting change. Due to the complexity of volcanic processes, this requires a polyphonic detection, segmentation, and classification approach. Through evolving epistemic uncertainty, invoking a Bayesian network strategy, we detect change and demonstrate its significance as an indicator for possible forecasting of eruptions using data from the Bezymianny and Etna volcanoes. Specifically, we propose morphing the scattering transform from previous work into a novel E2E hybrid and recurrent learnable deep scattering network to adapt to multi-scale temporal dependencies from streaming data. The time-dependent scattering is in some sense physics informed, namely, through time–frequency representation (TFR) of the data. At the same time, with a carefully designed deep convolutional LSTM (ConvLSTM) architecture, we learn intra-event, temporal dynamics from the scattering coefficients or features. We verify the effectiveness of transfer learning switching between volcanoes. Our experimental results set a new norm for semi-supervised seismo-volcanic monitoring. Ángel Bueno, Randall Balestriero, Silvio De Angelis, M. Carmen Benítez, Luciano Zuccarello, Richard G. Baraniuk, Jesús Ibáñez 0003, Maarten V. de Hoop |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Bayesian Monitoring of Seismo-Volcanic DynamicsabstractMethods for volcano monitoring that are based on analysis of geophysical data often rely on deterministic approaches without considering the complex and dynamic nature of volcanic systems. To detect subtle changes within seismic sequences associated with volcanic unrest, specialized workflows for data classification and analysis are required. Here, we present an inference framework based on Bayesian Deep Learning as a probabilistic proxy, which allows monitoring continuous changes in seismic activity at volcanoes. This architecture has been designed and trained to detect and to classify individual earthquake transients from continuous seismic data recorded in volcanic environments. We tested this new framework by analyzing seismic data associated with eruptions at Bezymianny Volcano (Russia) during 2007. Our results demonstrate efficient signal detection and classification accuracy, and effective detection of changes in the volcanic system in the hours preceding eruptive activity. This approach can be extended to other volcanoes and earthquake-prone areas, and demonstrates a new application of deep learning in the field of seismic monitoring. Ángel Bueno, M. Carmen Benítez, Luciano Zuccarello, Silvio De Angelis, Jesús Ibáñez 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | A Contribution to Deep Learning Approaches for Automatic Classification of Volcano-Seismic Events: Deep Gaussian ProcessesabstractThe automatic classification of volcano-seismic events is a key problem in volcanology. Due to its complexity, deep learning (DL) techniques have become the tool of choice for this problem, outperforming classical classifiers. The main drawback of this approach, when applied to the classification of volcano-seismic events, is its tendency to overfit because of the small-size available databases. In this work, we propose and analyze the use of the Gaussian processes (GPs) and Deep GPs (DGPs), and their hierarchical extension, for volcano-seismic event classification. We empirically prove the adequacy of the proposed modeling with an insightful and exhaustive comparison with state-of-the-art DL-based methods on a seismic database recorded at “Volcán de Fuego,” Colima, Mexico. The hierarchical structure of DGPs and the reduced number of parameters to be automatically estimated become essential to achieve excellent performance even on small databases, capturing well the complex patterns of seismic signals for all classes and, in particular, for those that have been hardly observed. Miguel López-Pérez, Luz García 0001, M. Carmen Benítez, Rafael Molina 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Automatic S-Phase Picking for Volcano-Tectonic Earthquakes Using Spectral Dissimilarity AnalysisabstractWe present an S-phase picking algorithm for volcano-tectonic earthquakes (VTs), based on the changes of frequency and amplitude expected in the plane transverse to the ray direction at S-phase arrival. A measure of these changes, called spectral dissimilarity, is proposed. Picking is performed in a particular waveform transformation that underlines such variations foreseen: horizontal instant power. Then, the algorithm provides a measure of its reliability, grounded on the low or high fluctuations of the picking instant obtained when applied to other horizontal components of the seismogram. Experiments are performed to test the algorithm with a challenging database of volcano seismic earthquakes from Mt. Etna, carefully picked and labeled by a human expert. The technique is compared to two well-known S-phase pickers: one based on the damped predominant period analysis and the other based on polarization and kurtosis rate analysis. The algorithm improves these techniques for the particular scenario of VTs, providing interesting results and possibilities of the application. Luz García 0001, Gerardo Alguacil, Manuel Titos, Ornella Cocina, Ángel de la Torre, M. Carmen Benítez |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2020 | Classification of Isolated Volcano-Seismic Events Based on Inductive Transfer LearningabstractDomain-specific problems where data collection is an expensive task are often represented by scarce or incomplete data. From a machine learning perspective, this type of problems has been addressed using models trained in different specific domains as the starting point for the final objective-model. The transfer of knowledge between domains, known as transfer learning (TL), helps to speed up training and improve the performance of the models in problems with limited amounts of data. In this letter, we introduce a TL approach to classify isolated volcano-seismic signals at “Volcán de Fuego”, Colima (Mexico). Using the well-known convolutional architecture (LeNet) as a feature extractor and a representative data set containing regional earthquakes, volcano-tectonic earthquakes, long-period events, volcanic tremors, explosions, and collapses, our proposal compares the generalization capabilities of the models when we only fine-tune the upper layers and fine-tune overall of them. Compared with the other state-of-the-art techniques, classification systems based on TL approaches provide good generalization capabilities (attaining nearly 94% of events correctly classified) and decreasing computational time resources. Manuel Titos, Ángel Bueno, Luz García 0001, M. Carmen Benítez, José C. Segura |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2020 | Volcano-Seismic Transfer Learning and Uncertainty Quantification With Bayesian Neural NetworksabstractOver the past few years, deep learning (DL) has emerged as an important tool in the fields of volcano and earthquake seismology. However, these methods have been applied without performing thorough analyses of the associated uncertainties. Here, we propose a solution to enhance volcano-seismic monitoring systems, through probabilistic Bayesian DL; we implement and demonstrate a workflow for waveform classification, rapid quantification of the associated uncertainty, and link these uncertainties to changes in volcanic unrest. Specifically, we introduce Bayesian neural networks (BNNs) to perform event identification, classification, and their estimated uncertainty on data gathered at two active volcanoes, Mount St. Helens, Washington, USA, and Bezymianny, Kamchatka, Russia. We demonstrate how BNNs achieve excellent performance (92.08%) in discriminating both the type of event and its origin when the two data sets are merged together, and no additional training information is provided. Finally, we demonstrate that the data representations learned by the BNNs are transferable across different eruptive periods. We also find that the estimated uncertainty is related to changes in the state of unrest at the volcanoes and propose that it could be used to gauge whether the learned models may be exported to other eruptive scenarios. Ángel Bueno, M. Carmen Benítez, Silvio De Angelis, Alejandro Diaz-Moreno, Jesús Ibáñez 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2019 | Detection and Classification of Continuous Volcano-Seismic Signals With Recurrent Neural NetworksabstractThis paper introduces recurrent neural networks (RNN), long short-term memory (LSTM), and gated recurrent unit (GRU) to detect and classify continuous sequences of volcano-seismic events at the Deception Island Volcano, Antarctica. A representative data set containing volcano-tectonic earthquakes, long-period events, volcanic tremors, and hybrid events was used to train these models. Experimental results show that RNN, LSTM, and GRU can exploit temporal and frequency information from continuous seismic data, attaining close to 90%, 94%, and 92% events correctly detected and classified. A second experiment is presented in this paper. The architectures described above, trained with data from campaigns of seismic records obtained in 1995-2002, have been tested with data from the recent seismic survey performed at the Deception Island Volcano in 2016-2017 by the Spanish Antarctic scientific campaign. Despite the variations in the geophysical properties of the seismic events within the volcano across eruptive periods, results provide good generalization accuracy. This result expands the possibilities of RNNs for real-time monitoring of volcanic activity, even if seismic sources change over time. Manuel Titos, Ángel Bueno, Luz García 0001, M. Carmen Benítez, Jesús Ibáñez 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2015 | Sub-band based histogram equalization in cepstral domain for speech recognition
Vikas Joshi, Raghavendra Bilgi, Srinivasan Umesh, Luz García 0001, M. Carmen Benítez |
Speech Commun. | 5 |
| 2013 | An Automatic P-Phase Picking Algorithm Based on Adaptive Multiband ProcessingabstractThis letter presents a novel picking algorithm which allows an automated determination of the P-phase onset time. The algorithm includes an adaptive multiband processing and noise-reduction techniques to allow a confident onset time estimation in signals strongly affected by background and/or nonstationary noise processes. Results using a set of 3780 computer-generated earthquake-like signals show that the accuracy is much better than that achieved by conventional STA/LTA algorithm. In addition, the accuracy of the proposed method is improved when it is combined with an autoregressive method. An application of the algorithm to a set of 400 natural earthquakes confirms that the combination of both algorithms provides a precise P-phase onset time estimation in real environments, overcoming the limitations associated with the autoregressive method. Isaac Álvarez, Luz García 0001, Sonia Mota, Guillermo Cortés, M. Carmen Benítez, Ángel de la Torre |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2012 | Robust speech recognition through selection of speaker and environment transformsabstractIn this paper, we address the problem of robustness to both noise and speaker-variability in automatic speech recognition (ASR). We propose the use of pre-computed Noise and Speaker transforms, and an optimal combination of these two transforms are chosen during test using maximum-likelihood (ML) criterion. These pre-computed transforms are obtained during training by using data obtained from different noise conditions that are usually encountered for that particular ASR task. The environment transforms are obtained during training using constrained-MLLR (CMLLR) framework, while for speaker-transforms we use the analytically determined linear-VTLN matrices. Even though the exact noise environment may not be encountered during test, the ML-based choice of the closest Environment transform provides “sufficient” cleaning and this is corroborated by experimental results with performance comparable to histogram equalization or Vector Taylor Series approaches on Aurora-2 task. The proposed method is simple since it involves only the choice of pre-computed environment and speaker transforms and therefore, can be applied with very little test data unlike many other speaker and noise-compensation methods. Raghavendra Bilgi, Vikas Joshi, Srinivasan Umesh, Luz García 0001, M. Carmen Benítez |
ICASSP | 5 |
| 2012 | Noise and speaker compensation in the Log filter bank domainabstractIn this paper, we propose a method to compensate for noise and speaker-variability directly in the Log filter-bank (FB) domain, so that MFCC features are robust to noise and speaker-variations. For noise-compensation, we use Vector Taylor Series (VTS) approach in the Log FB domain, and speaker-normalization is also done in the Log FB domain using Linear Vocal tract length (VTLN) matrices. For VTLN, optimal selection of warp-factor is done in Log FB domain using canonical GMM model, avoiding the two-pass approach needed by a HMM model. Further, this can be efficiently implemented using sufficient statistics obtained from the GMM and the FB-VTLN-matrices. The warp-factor selection using GMM can also be done in cepstral domain by applying DCT matrices without the usual approximations associated with conventional linear-VTLN. The elegance of the proposed approach is that given the speech data, we obtain directly MFCC features that are robust to noise and speaker-variations. The proposed approach, show a significant relative improvement of 31% over baseline on Aurora-4 task. Vikas Joshi, Raghavendra Bilgi, Srinivasan Umesh, Luz García 0001, M. Carmen Benítez |
ICASSP | 5 |
| 2012 | TELIAMADE ultrasonic indoor location system: Application as a teaching toolabstractThis paper proposes TELIAMADE (an indoor location system based on ultrasonic and radiofrequency signals) to be used as a teaching tool in the context of Telecommunication Engineering. Due to its simple design, the versatility of its configuration and the characteristics of the involved signals, TELIAMADE is an appropriate tool for teaching basic aspects in location systems, digital communication systems, encoded signalling, microcontroller programming, radio protocols or advanced signal processing techniques. The TELIAMADE design allows students to sample, store and analyze signals at different points of the circuits by using conventional oscilloscopes. Furthermore, some parameters can be configured, allowing students to assess the advantages and inconveniences of each specific configuration with respect to features such as bit-rate, range, robustness against noise or updating period. Our system presents advantages in the field of teaching for understanding commercial systems for location (like GPS) or communication (like wireless digital communication systems). Carlos Medina, Isaac Álvarez, José C. Segura, Ángel de la Torre, M. Carmen Benítez |
ICASSP | 5 |
| 2012 | Discriminative Feature Selection for Automatic Classification of Volcano-Seismic SignalsabstractFeature extraction is a critical element in automatic pattern classification. In this letter, we propose different sets of parameters for classification of volcano-seismic signals, and the discriminative feature selection (DFS) method is applied for selecting the minimum number of features containing most of the discriminative information. We have applied DFS to a conventional cepstral-based parameterization (with 39 features) and to an extended set of parameters (including 84 features). Classification experiments using seismograms recorded at Colima Volcano (Mexico) show that, for the most complex classifier and using the cepstral-based parameterization, DFS provided a reduction of the error rate from 24.3% (using 39 features) to 15.5% (ten components). When DFS is applied to the extended parameterization, the error rate decreased from 27.9% (84 features) to 13.8% (14 features). These results show the utility of DFS for identifying the best components from the original feature vector and for exploring new parameterizations for the classification of volcano-seismic signals. Isaac Álvarez, Luz García 0001, Guillermo Cortés, M. Carmen Benítez, Ángel de la Torre |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2012 | Class-Based Parametric Approximation to Histogram Equalization for ASRabstractThis letter assesses an improved equalization transformation for robust speech recognition in noisy environments. The proposal is an evolution of the parametric approximation to Histogram Equalization named PEQ into a two-step algorithm dealing separately with environmental and acoustic mismatch. A first parametric equalization is done to eliminate environmental mismatch. These equalized data are divided into classes, and parametrically re-equalized using class specific references to reduce the acoustic mismatch. Experiments have been conducted for Aurora 2 and Aurora 4 databases. A comparative analysis of the experimental results shows significant benefits for databases with high acoustic variability like Aurora 4. Luz García 0001, M. Carmen Benítez, Ángel de la Torre, José C. Segura |
IEEE Signal Process. Lett. | 2 |
| 2011 | Combining speaker and noise feature normalization techniques for Automatic Speech RecognitionabstractThis work deals with strategies to jointly reduce the speaker and environment mismatches in Automatic Speech Recognition. The consequences of environmental mismatch in the performance of conventional Vocal Tract Length Normalization algorithm are analyzed, observing the sensitivity of the warping factor distributions to the SNR fall. A new combined speaker-noise normalization strategy which reduces the effect of noise in VTLN by applying Histogram Equalization is proposed and experimented in AURORA2 and AU RORA4 databases. Solid results are obtained and discussed to analyze the effectiveness of the described technique. Luz García 0001, M. Carmen Benítez, José C. Segura |
ICASSP | 2 |
| 2011 | Efficient Speaker and Noise Normalization for Robust Speech Recognition
Vikas Joshi, Raghavendra Bilgi, Srinivasan Umesh, M. Carmen Benítez, Luz García 0001 |
INTERSPEECH | 4 |
| 2011 | Sub-Band Level Histogram Equalization for Robust Speech RecognitionabstractThis paper describes a novel modification of Histogram Equalization approach to robust speech recognition. We propose separate equalization of the high frequency and low frequency bands. We study different combinations of the sub-band equalization and obtain best results when we performs a twostage equalization. First, conventional Histogram Equalization (HEQ) is performed on the cepstral features, which does not completely equalize high frequency and low frequency bands, even though the overall histogram equalization is good. In the second stage, an equalization is done separately on the high frequency and the low frequency components of the above equalized cepstra. We refer to this approach as Sub-band Histogram Equalization (S-HEQ). The new set of features has better equalization of the sub-bands as well as the overall cepstral histogram. Recognition results show a relative improvement of 12% and 15% over conventional HEQ on Aurora-2 and Aurora4 databases respectively. Vikas Joshi, Raghavendra Bilgi, Srinivasan Umesh, Luz García 0001, M. Carmen Benítez |
INTERSPEECH | 5 |
| 2009 | Improving Feature Extraction in the Automatic Classification of Seismic Events. Application to Colima and Arenal VolcanoesabstractMonitoring of precursory seismicity in volcanoes is the most reliable and widely used technique in volcano monitoring. Since a visual inspection by human operators is a tedious task in a non-stop monitoring process, Hidden Markov Models have been previously proposed to automatically classify the different types of volcano-seismic events. Mel Frequency Cepstral Coefficients were successfully used as feature vector in this continuous classification system. In this paper seven novel features to be included in the MFCC feature vector are proposed. A very elementary GMM-based classifier has been implemented in order to assess the efficiency of the proposed parameters. Results using hundreds of events recorded from stations situated at Colima (Mexico) and Arenal (Costa Rica) volcanoes show that the proposed features improve the recognition accuracy and therefore they may be relevant in continuous volcano-seismic event automatic classification. Isaac Álvarez, Guillermo Cortés, Ángel de la Torre, M. Carmen Benítez, Luz García 0001, Philippe Lesage, Raúl Arámbula, Miguel González-Amezcua |
IGARSS (4) | 4 |
| 2009 | Evaluating Robustness of a HMM-based Classification System of Volcano-seismic Events at Colima and Popocatepetl VolcanoesabstractThis work presents a continuous volcano-seismic classification system based in the Hidden Markov Models as solution to recently strong needs for automatic event detection and recognition methods in early warning and monitoring scenarios. Furthermore, our system includes a reliable method to assign confidence measures to the recognized signals in order to evaluate the robustness of the results. Data from the two most active volcanoes have been used to probe the system reliability on a complex joint corpus achieving a recognition accuracy higher than 78% in blind recognition tests. Guillermo Cortés, Raúl Arámbula, Ligdamis A. Gutiérrez, M. Carmen Benítez, Jesús Ibáñez 0003, Philippe Lesage, Isaac Álvarez, Luz García 0001 |
IGARSS (2) | 4 |
| 2009 | Volcano-seismic Signal Detection and Classification Processing using Hidden Markov Models. Application to San Cristóbal Volcano, NicaraguaabstractWe present a method for automatic seismic event detection and classification, focusing on volcanic-seismic signals by means of the validity of the hidden Markov modeling (HMM) method in active volcanoes. Recordings of different seismic event types are studied at one active volcano; San Cristobal in Nicaragua. We use data from one field surveys carried out in February to March 2006. More than 600 hours of data in San Cristobal volcano were analyzed and 1098 seismic events were registered at short period stations. These events were manually labelled by a single expert technicians and identified three types classes of signals (S1, S2, S3) and tremor background seismic noise (NS). The method analyzes the seismograms comparing the characteristics of the data to a number of event classes defined beforehand. If a signal is present, the method detects its occurrence and produces a classification. From the application performed over our data set, we have demonstrated that in order to have a reliable result, a careful and adequate segmentation process is crucial. Also, each type of signals requires its own characterization. That is, each signal type must be represented by its own specific model, which would include the effects of source, path and sites. Once we have built this model, the success level of the system is high. Extensive performance evaluation is conducted to derive the optimal configuration of the different parameters Correct classification rates of up to 80% are achieved. The high success rates obtained imply that the method is fully able to detect, isolate, and identify seismic signals on raw seismic data. These results imply that, once an adequate training process has been used, the present method is particularly appropriate to work in real time, and in parallel to the data acquisition. Ligdamis A. Gutiérrez, Jesús Ibáñez 0003, Guillermo Cortés, Javier Ramírez 0001, M. Carmen Benítez, Virginia Tenorio, Isaac Álvarez |
IGARSS (4) | 5 |
| 2007 | Continuous HMM-Based Seismic-Event Classification at Deception Island, AntarcticaabstractThis paper shows a complete seismic-event classification and monitoring system that has been developed based on the seismicity observed during three summer Antarctic surveys at the Deception Island Volcano, Antarctica. The system is based on the state of the art in hidden Markov modeling (HMM) techniques successfully applied to other scenarios. A database that contains a representative set of different seismic events including volcano-tectonic earthquakes, long period (LP) events, volcanic tremor, and hybrid events that were recorded during the 1994–1995 and 1995–1996 seismic surveys was collected for training and testing. Simple left-to-right HMMs and multivariate Gaussian probability density functions with a diagonal covariance matrix were used. The feature vector consists of the log-energies of a filter bank that consists of 16 triangular weighting functions that were uniformly spaced between 0 and 20 Hz and the first- and second-order derivatives. The system is suitable to operate in real time, and its accuracy for this task is about 90%. On the other hand, when the system was tested with a different data set including mainly LP events that were registered during several seismic swarms during the 2001–2002 field survey, more than 95% of the recognized events were marked by the recognition system. M. Carmen Benítez, Javier Ramírez 0001, José C. Segura, Jesús Ibáñez 0003, Javier Almendros, Araceli García-Yeguas, Guillermo Cortés |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2006 | Continuous HMM-Based Volcano Monitoring at Deception Island, AntarcticaabstractThis paper shows a complete volcano monitoring system that has been developed on the basis of the seismicity observed during three summer Antarctic surveys at Deception Island Volcano (Antarctica). The system is based on the state of the art in hidden Markov modelling (HMM) techniques successfully applied to other scenarios. A database containing a representative set of different seismic events including volcano-tectonic earthquakes, long-period events, volcanic tremor and hybrid events recorded during the 1994-1995 and 1995-1996 seismic surveys was collected for training and testing. Simple left-to-right HMMs and multivariate Gaussian probability density functions (PDF) with diagonal covariance matrix were used. The feature vector consists of the log-energies of a filter-bank consisting of 16 triangular weighting functions uniformly spaced between 0 and 20 Hz plus the first and second order derivatives. The system is suitable to operate in real-time and its accuracy is close to 90%. When the system was tested with a different data set including mainly long-period events registered during several seismic swarms during the 2001-2002 field survey, more than 95% of the recognized events were correctly marked by the recognition system M. Carmen Benítez, Javier Ramírez 0001, José C. Segura, Antonio J. Rubio, Jesús Ibáñez 0003, Javier Almendros, Araceli García-Yeguas |
ICASSP (5) | 1 |
| 2006 | Parametric Nonlinear Feature Equalization for Robust Speech RecognitionabstractA new front-end normalization algorithm that uses a parametric nonlinear transformation is proposed in this paper. The method improves histogram equalization based nonlinear transformations by finding a simple and computationally inexpensive parametric expression of the nonlinear transformation. The new parametric approach relies on a two Gaussian model for the probability distribution of the features, and on a simple Gaussian classifier to label the input frames as belonging to the speech or non-speech classes. The result is a more robust equalization, less dependent on the percentage of speech and non-speech frames. Recognition experiments on the AURORA 4 database have been performed and the effectiveness of the algorithm is analyzed in comparison with other linear and nonlinear feature equalization techniques Luz García 0001, José C. Segura, Javier Ramírez 0001, Ángel de la Torre, M. Carmen Benítez |
ICASSP (1) | 5 |
| 2006 | Gaialab: A Weblab Project for Digital Communications Distributed LearningabstractThis paper presents GAIALAB, a Weblab project for digital communications distributed learning developed at the University of Granada (Spain). This initiative has been funded by a University program to improve teaching quality. The project addresses problems of teaching and learning innovatively and effectively by enabling the user interaction with the simulation engines through a Web browser. Our Weblab project provides an interface with JAVA applets that have been developed and integrated in the system. The Weblab project enables simulating digital communication systems and analyzing the results obtained through the Web. Moreover, all the applets that are included in the system have been designed to allow the modification of the simulation parameters, thus enabling a better understanding of the algorithms. The system also permits evaluating the experimental work developed by the students Javier Ramírez 0001, José C. Segura, Juan Manuel Górriz, M. Carmen Benítez, Antonio J. Rubio |
ICASSP (2) | 4 |
| 2006 | HMM-based continuous sign language recognition using a fast optical flow parameterization of visual informationabstractThis paper presents a preliminary study of an optical flow-based parameterization of visual information in a sign language recognition system using Hidden Markov Models (HMM). Current feature extraction processes need initialization, tracking and segmentation stages in order to describe signer gestures. Our aim is to develop a single and fast technique to reduce computational complexity which doesn't require these stages and is able to work in mobile devices with limited hardware resources. The Moving Block Distance (MBD) parameterization is an interesting first approach for this purpose, proved by two signers under a static background constraint. A lexicon of 33 basic word units (signemes) was used to build the data set containing phrases with a variable number of words. Continuous recognition results achieve more than 99 % accuracy in close test. Index Terms: multi-modal recognition, visual feature extraction, mobile devices, sign language, optical flow. 1. Guillermo Cortés, Luz García 0001, M. Carmen Benítez, José C. Segura |
INTERSPEECH | 3 |
| 2006 | Normalization of the inter-frame information using smoothing filteringabstractA filter that introduces inter-frame information into the voice features set is proposed in this paper. The filter adds the autocorrelations of the cepstral coefficients to the set of characteristics used for training and recognition. Those autocorrelations should not depend on the environment conditions. Because they should only depend on the information to recognize, a normalization of that inter-frame information is convenient. The filter defined implements this normalization by transforming the autocorrelations into a normalized domain defined with clean adaptation data. This temporal processing of the features is added to the Histogram Equalization of the cepstral coefficients (HEQ) used to normalize the MFCCs. An analysis is done about the most effective domain (original MFCCS or equalized MFCCs) on which the temporal processing should be executed. Performance results for the proposed algorithm are presented for AURORA2 and AURORA4 databases. Luz García 0001, José C. Segura, M. Carmen Benítez, Javier Ramírez 0001, Ángel de la Torre |
INTERSPEECH | 3 |
| 2006 | Noise robust model-based voice activity detectionabstractWe propose a model-based VAD derived from the Vector Taylor Series (VTS) approach. A Gaussian mixture (trained with clean speech) is used in order to provide an appropriate decision rule for speech/non-speech detection. Additionally, VTS approach adapts the Gaussian mixture to noise conditions, yielding a stable perfor- mance for a wide range of SNRs. We have evaluated its ability for speech/non-speech detection and also its application for robust speech recognition. When compared to other VAD methods, the proposed VAD shows the best trade-off in speech/non-speech de- tection. When applied for Wiener Filtering and for frame drop- ping, the proposed VAD also provides the best recognition results. Index Terms: voice activity detection (VAD), vector Taylor series approach (VTS), Gaussian mixture, Wiener filtering. Ángel de la Torre, Javier Ramírez 0001, M. Carmen Benítez, José C. Segura, Luz García 0001, Antonio J. Rubio |
INTERSPEECH | 3 |
| 2005 | Statistical voice activity detection using a multiple observation likelihood ratio testabstractCurrently, there are technology barriers inhibiting speech processing systems that work in extremely noisy conditions from meeting the demands of modern applications. This letter presents a new voice activity detector (VAD) for improving speech detection robustness in noisy environments and the performance of speech recognition systems. The algorithm defines an optimum likelihood ratio test (LRT) involving multiple and independent observations. The so-defined decision rule reports significant improvements in speech/nonspeech discrimination accuracy over existing VAD methods that are defined on a single observation and need empirically tuned hangover mechanisms. The algorithm has an inherent delay that, for several applications, including robust speech recognition, does not represent a serious implementation obstacle. An analysis of the overlap between the distributions of the decision variable shows the improved robustness of the proposed approach by means of a clear reduction of the classification error as the number of observations is increased. The proposed strategy is also compared to different VAD methods, including the G.729, AMR, and AFE standards, as well as recently reported algorithms showing a sustained advantage in speech/nonspeech detection accuracy and speech recognition performance. Javier Ramírez 0001, José C. Segura, M. Carmen Benítez, Luz García 0001, Antonio J. Rubio |
IEEE Signal Process. Lett. | 3 |
| 2005 | An effective subband OSF-based VAD with noise reduction for robust speech recognitionabstractAn effective voice activity detection (VAD) algorithm is proposed for improving speech recognition performance in noisy environments. The approach is based on the determination of the speech/nonspeech divergence by means of specialized order statistics filters (OSFs) working on the subband log-energies. This algorithm differs from many others in the way the decision rule is formulated. Instead of making the decision based on the current frame, it uses OSFs on the subband log-energies which significantly reduces the error probability when discriminating speech from nonspeech in a noisy signal. Clear improvements in speech/nonspeech discrimination accuracy demonstrate the effectiveness of the proposed VAD. It is shown that an increase of the OSF order leads to a better separation of the speech and noise distributions, thus allowing a more effective discrimination and a tradeoff between complexity and performance. The algorithm also incorporates a noise reduction block working in tandem with the VAD and showed to further improve its accuracy. A previous noise reduction block also improves the accuracy in detecting speech and nonspeech. The experimental analysis carried out on the AURORA databases and tasks provides an extensive performance evaluation together with an exhaustive comparison to the standard VADs such as ITU G.729, GSM AMR, and ETSI AFE for distributed speech recognition (DSR), and other recently reported VADs. Javier Ramírez 0001, José C. Segura, M. Carmen Benítez, Ángel de la Torre, Antonio J. Rubio |
IEEE Trans. Speech Audio Process. | 3 |
| 2005 | Histogram equalization of speech representation for robust speech recognitionabstractThis paper describes a method of compensating for nonlinear distortions in speech representation caused by noise. The method described here is based on the histogram equalization method often used in digital image processing. Histogram equalization is applied to each component of the feature vector in order to improve the robustness of speech recognition systems. The paper describes how the proposed method can be applied to robust speech recognition and it is compared with other compensation techniques. The recognition experiments, including results in the AURORA II framework, demonstrate the effectiveness of histogram equalization when it is applied either alone or in combination with other compensation techniques. Ángel de la Torre, Antonio M. Peinado, José C. Segura, José L. Pérez-Córdoba, M. Carmen Benítez, Antonio J. Rubio |
IEEE Trans. Speech Audio Process. | 5 |
| 2004 | A new voice activity detector using subband order-statistics filters for robust speech recognitionabstractCurrently, there are technology barriers inhibiting speech processing systems working under extreme noisy conditions. The emerging applications of speech technology, especially in the fields of wireless communications, digital hearing aids or speech recognition, are some examples of such systems often requiring a noise reduction technique in combination with a precise voice activity detector (VAD). This paper presents a new VAD for improving speech detection robustness in noisy environments and the performance of speech recognition systems. The algorithm uses long-term information about the speech signal to formulate the decision rule and estimates the subband SNR using specialized order statistics filters (OSF). The proposed algorithm is compared to the most commonly used VAD in the field, in terms of speech/nonspeech discrimination and also in terms of recognition performance when the VAD is used in an automatic speech recognition (ASR) system. Experimental results demonstrate a sustained advantage over different VAD methods including standard VAD such as G.729 and AMR which are used as a reference, the VAD of the Advanced Front-End (AFE) for distributed speech recognition (DSR), and recently reported algorithms. Javier Ramírez 0001, José C. Segura, M. Carmen Benítez, Ángel de la Torre, Antonio J. Rubio |
ICASSP (1) | 3 |
| 2004 | Voice activity detection with noise reduction and long-term spectral divergence estimationabstractThe paper mainly focusses on an improved voice activity detection algorithm employing long-term signal processing and maximum spectral component tracking. The benefits of this approach have been analyzed in a previous work (Ramirez, J. et al., Proc. EUROSPEECH 2003, p.3041-4, 2003) with clear improvements in speech/non-speech discriminability and speech recognition performance in noisy environments. Two clear aspects are now considered. The first one, which improves the performance of the VAD in low noise conditions, considers an adaptive length frame window to track the long-term spectral components. The second one reduces misclassification errors in highly noisy environments by using a noise reduction stage before the long-term spectral tracking. Experimental results show clear improvements over different VAD methods in speech/pause discrimination and speech recognition performance. Particularly, improvements in recognition rate were reported when the proposed VAD replaced the VADs of the ETSI advanced front-end (AFE) for distributed speech recognition (DSR). Javier Ramírez 0001, José C. Segura, M. Carmen Benítez, Ángel de la Torre, Antonio J. Rubio |
ICASSP (2) | 3 |
| 2004 | Improved voice activity detection combining noise reduction and subband divergence measuresabstractCurrently, new trends in wireless communications are demanding reliable human-machine interaction in real-life environments. However, there are obstacles inhibiting automatic speech recognition systems (ASR) working in noisy environments. The main difficulty is the degradation suffered by ASR systems due to a mismatch between training and test conditions. This paper shows an improved voice activity detector (VAD) combining noise reduction and subband divergence estimation for improving the reliability of speech recognizers operating in noisy environments. The algorithm formulates the decision rule by measuring the divergence between the subband spectral magnitude of speech and noise using the Kullback-Leibler (KL) distance on the denoised signal. Experiments demonstrate a sustained advantage over different VAD methods including standard VADs such as G.729 and AMR, which are used as a reference, recently reported algorithms, and the VADs of the advanced frontend (AFE) for distributed speech recognition (DSR). 1. Javier Ramírez 0001, José C. Segura, M. Carmen Benítez, Ángel de la Torre, Antonio J. Rubio |
INTERSPEECH | 3 |
| 2004 | Including uncertainty of speech observations in robust speech recognitionabstractNoise compensation methods for speech recognition provide a cleaned version of the speech representation. Usually this cleaned version is the expected value of the speech parameters given the observed noisy speech and the noise statistic. A more realistic representation should include the probability distribution of the cleaned speech instead of its expected value in order to represent the uncertainty associated to the compensation process due to the variability of the noise process. Recently, the inclusion of the uncertainty in the recognition process has been studied. Some approaches represent the uncertainty in the HMM parameters values. Other approaches represent it in the feature space. This second approach offers a much simpler system implementation and lower computational cost. In this paper we have developed a noise compensation technique that incorporates the variance of the cleaned speech into the speech representation. The variance is estimated using a Wiener filter during the speech feature enhancement process. This way of including the uncertainty implies the modification of the decoding rule. Experimental results using AURORA 2 database demonstrate a sustained improvement of the performance in the recognition system (about 21% word error rate reduction) when uncertainty is considered in the decoding rule. José C. Segura, Ángel de la Torre, Javier Ramírez 0001, Antonio J. Rubio, M. Carmen Benítez |
INTERSPEECH | 5 |
| 2004 | Efficient voice activity detection algorithms using long-term speech information
Javier Ramírez 0001, José C. Segura, M. Carmen Benítez, Ángel de la Torre, Antonio J. Rubio |
Speech Commun. | 3 |
| 2004 | A new Kullback-Leibler VAD for speech recognition in noiseabstractThis letter shows an innovative voice activity detector (VAD) based on the Kullback-Leibler (KL) divergence measure. The algorithm is evaluated in the context of the recently approved ETSI standard for distributed speech recognition (DSR). The VAD uses long-term information of the noisy speech signal in order to define a more robust decision rule yielding high accuracy. The mel-scaled filter bank log-energies (FBE) are modeled by means of Gaussian distributions, and a symmetric KL divergence is used for the estimation of the distance between speech and noise distributions. The decision rule is formulated in terms of the average subband KL divergence that is compared to a noise-adaptable threshold. An exhaustive analysis using the AURORA databases is conducted in order to assess the performance of the proposed method and to compare it to existing standard VAD methods. Javier Ramírez 0001, José C. Segura, M. Carmen Benítez, Ángel de la Torre, Antonio J. Rubio |
IEEE Signal Process. Lett. | 3 |
| 2004 | Cepstral domain segmental nonlinear feature transformations for robust speech recognitionabstractThis letter presents a new segmental nonlinear feature normalization algorithm to improve the robustness of speech recognition systems against variations of the acoustic environment. An experimental study of the best delay-performance tradeoff is conducted within the AURORA-2 framework, and a comparison with two commonly used normalization algorithms is presented. Computationally efficient algorithms based on order statistics are also presented. One of them is based on linear interpolation between sampling quantiles, and the other one is based on a point estimation of the probability distribution. The reduction in the computational cost does not degrade the performance significantly. José C. Segura, M. Carmen Benítez, Ángel de la Torre, Antonio J. Rubio, Javier Ramírez 0001 |
IEEE Signal Process. Lett. | 2 |
| 2003 | A new adaptive long-term spectral estimation voice activity detectorabstractThis paper shows an efficient voice activity detector (VAD) that is based on the estimation of the long-term spectral diver- gence (LTSD) between noise and speech periods. The proposed method decomposes the input signal into overlapped speech frames, uses a sliding window to compute the long-term spec- tral envelope and measures the speech/non-speech LTSD, thus yielding a high discriminating decision rule and minimizing the average number of decision errors. In order to increase non- speech detection accuracy, the decision threshold is adapted to the measured noise energy while a controlled hang-over is ac- tivated only when the observed signal-to-noise ratio (SNR) is low. An exhaustive analysis of the proposed VAD is carried out using the AURORA TIdigits and SpeechDat-Car (SDC) databases. The proposed VAD is compared to the most com- monly used ones in the field in terms of speech/non-speech detection and recognition performance. Experimental results demonstrate a sustained advantage over G.729, AMR and AFE VADs. Javier Ramírez 0001, José C. Segura, M. Carmen Benítez, Ángel de la Torre, Antonio J. Rubio |
INTERSPEECH | 3 |
| 2003 | Improved feature extraction based on spectral noise reduction and nonlinear feature normalizationabstractThis paper is mainly focused on showing experimental results of a feature extraction algorithm that combines spectral noise reduction and nonlinear feature normalization. The successfulness of this approach has been shown in a previous work, and in this one, we present several improvements that result in a performance comparable to that of the recently approved AFE for DSR. Noise reduction is now based on a Wiener filter instead of spectral subtraction. The voice activity detection based on the full-band energy has been replaced with a new one using spectral information. Relative improvements of 24.81% and 17.50% over our previous system are obtained for AURORA 2 and 3 respectively. Results for AURORA 2 are not as good as those for the AFE, but for AURORA 3 a relative improvement of 5.27% is obtained. José C. Segura, Javier Ramírez 0001, M. Carmen Benítez, Ángel de la Torre, Antonio J. Rubio |
INTERSPEECH | 3 |
| 2002 | VTS residual noise compensationabstractThe VTS approach for noise reduction is based on a statistical for mulation. It provides the expected value of the clean speech given the noisy observations and statistical models for the clean speech and the additive noise. The compensated signal is only an approximation of the clean one and retains a residual mismatch. The main objective of this work is to characterize this residual noise and to propose techniques to reduce its unwanted effects. Two different approaches to this problem are presented in this paper. The first one is based on linear filtering the time sequences of compensated acoustic parameters; for this purpose we use LDA-based RASTA-like FIR filters. The second approach is based on canceling the distortion introduced into the probability distribution of acoustic parameters and uses the well-known technique of histogram equalization. Results reported on AURORA database show that the proposed methods increase the recognition performance. José C. Segura, M. Carmen Benítez, Ángel de la Torre, Stéphane Dupont, Antonio J. Rubio |
ICASSP | 2 |
| 2002 | Non-linear transformations of the feature space for robust Speech RecognitionabstractThe noise usually produces a non-linear distortion of the feature space considered for Automatic Speech Recognition. This distortion causes a mismatch between the training and recognition conditions which significantly degrades the performance of speech recognizers. In this contribution we analyze the effect of the additive noise over cepstral based representations and we compare several approaches to compensate this effect. We discuss the importance of the non-linearities introduced by the noise and we propose a method (based on the histogram equalization technique) specifically oriented to the compensation of the non-linear transformation caused by the additive noise. The proposed method has been evaluated using the AURORA-2 database and task. The recognition results show significant improvements with respect to other compensation methods reported in the bibliography and reveals the importance of the non-linear effects of the noise and the utility of the proposed method. Ángel de la Torre, José C. Segura, M. Carmen Benítez, Antonio M. Peinado, Antonio J. Rubio |
ICASSP | 3 |
| 2002 | Feature extraction combining spectral noise reduction and cepstral histogram equalization for robust ASRabstractThis work is mainly focused on showing experimental results using a combination of two methods for noise compensation which are shown to be complementary: classical spectral subtraction algorithm and histogram equalization. While spectral subtraction is focused on the reduction of the additive noise in the spectral domain, histogram equalization is applied in the cepstral domain to compensate the remaining non-linear effects associated to channel distortion and additive noise. The estimation of the noise spectrum for the spectral subtraction method relies on a new algorithm for speech / non-speech detection (SND) based on order statistics. This SND classification is also used for dropping long speech pauses. Results on Aurora 2 and Aurora 3 are reported. José C. Segura, M. Carmen Benítez, Ángel de la Torre, Antonio J. Rubio |
INTERSPEECH | 2 |
| 2002 | Discriminative feature weighting for HMM-based continuous speech recognizers
Ángel de la Torre, Antonio M. Peinado, Antonio J. Rubio, José C. Segura, M. Carmen Benítez |
Speech Commun. | 5 |
| 2001 | Robust ASR front-end using spectral-based and discriminant features: experiments on the Aurora tasksabstractThis paper describes an automatic speech recognition frontend that combines low-level robust ASR feature extraction techniques, and higher-level linear and non-linear feature transformations. The low-level algorithms use data-derived filters, mean and variance normalization of the feature vectors, and dropping of noise frames. The feature vectors are then linearly transformed using Principal Components Analysis (PCA). An Artificial Neural Network (ANN) is also used to compute features that are useful for classification of speech sounds. It is trained for phoneme probability estimation on a large corpus of noisy speech. These transformations lead to two feature streams whose vectors are concatenated and then used for speech recognition. This method was tested on the set of speech corpora used for the “Aurora” evaluation. Using the feature stream generated without the ANN yields an overall 41% reduction of the error rate over Mel-Frequency Cepstral Coefficients (MFCC) reference features. Adding the ANN stream further reduces the error rate yielding a 46% reduction over the reference features. M. Carmen Benítez, Lukás Burget, Barry Y. Chen, Stéphane Dupont, Harinath Garudadri, Hynek Hermansky, Pratibha Jain, Sachin S. Kajarekar, Nelson Morgan, Sunil Sivadas |
INTERSPEECH | 1 |
| 2001 | Feature extraction from time-frequency matrices for robust speech recognitionabstractIn this paper we present a study about time-frequency distribution of acoustic-phonetic information for the Spanish language. This is based on a large Spanish database automatically labeled, and we conclude that results are similar to those obtained for hand-labeled english databases. We use bidimensional LDA [1] to extract discriminant features in time-frequency domain (TF) that are more robust in noise than the standard ones based on MFCC and time derivatives. We show that TF domain and its corresponding transformed domain (CTM) are equivalent from the point of view of LDA analysis and use this fact to reduce the dimensionality of the problem. Finally, cascade unidimensional LDA (CLDA) is applied first in frequency and then in time. This gives better estimates of projection vectors and better recognition performance. The proposed techniques are evaluated in a connected digit recognition task. Utterances have been artificially corrupted with additive real noises. José C. Segura, M. Carmen Benítez, Ángel de la Torre, Antonio J. Rubio |
INTERSPEECH | 2 |
| 2001 | Model-based compensation of the additive noise for continuous speech recognition. experiments using the Aurora II database and tasksabstractIn this paper we apply a model-based compensation method to cancel the effect of the additive noise in Automatic Speech Recognition systems. The method is formulated in a statistical framework in order to perform the optimal compensation of the noise effect given the observed noisy speech, a model describing the statistics of the speech recorded in a clean reference environment and the estimation of the noise in the noisy recognition environment. The noise is estimated using the first frames of the sentence to be recognized and a frame-by-frame noise compensation algorithm is performed, so that the compensation procedure does not constrain real-time speech recognition systems and is compatible with emerging technologies based on distributed speech recognition. José C. Segura, Ángel de la Torre, M. Carmen Benítez, Antonio M. Peinado |
INTERSPEECH | 3 |
| 2000 | Hard-Decision in COVQ over Waveform ChannelsabstractA channel optimized vector quantizer (COVQ) is studied for the case of transmission over waveform channels. In this work, a number of modulation schemes with multidimensional signal constellations are considered, specifically, results on the binary signalling. M-ary phase-shift keying (MPSK) and M-ary quadrature amplitude modulation (MQAM) performance using COVQ with hard-decision decoding, is optimized for additive white Gaussian noise (AWGN) and flat-fading Rayleigh channel. In addition, when a flat-fading Rayleigh channel is assumed, diversity techniques are used and evaluated to improve the performance of the system. José L. Pérez-Córdoba, Antonio J. Rubio, Juan M. López-Soler, M. Carmen Benítez |
Data Compression Conference | 4 |
| 2000 | Different confidence measures for word verification in speech recognition
M. Carmen Benítez, Antonio J. Rubio, Pedro García-Teodoro, Ángel de la Torre |
Speech Commun. | 1 |
| 1999 | A transcription-based approach to determine the difficulty of a speech recognition taskabstractA new parameter for estimating the difficulty of a continuous speech recognition task, called speech decoding difficulty, is presented. It is obtained from the language model defined for the recognition task and the phonetic similarity between the transcriptions of the words that make up the vocabulary used. Two variants of the proposed task difficulty measure are introduced: ideal speech decoding difficulty (ISDD), which is not influenced by practical considerations on the recognition system implemented, and a second, more realistic variant, called practical speech decoding difficulty (PSDD) to study the performance of a specific recognition system when confronting a given task. Pedro García-Teodoro, Antonio J. Rubio, Jesús Esteban Díaz Verdejo, M. Carmen Benítez, Juan M. López-Soler |
IEEE Trans. Speech Audio Process. | 4 |
| 1998 | Word verification using confidence measures in speech recognitionabstractIn this work we propose a novel way of discriminating the words that are recognized by a speech recognition system as correctly or incorrectly detected words. The procedure consists of the extraction of a set of characteristics for each word. Utilizing these characteristics, we have built two classifiers: the first one is a vector quantizer, while the second one, though also a vector quantizer, was trained using adaptative technique learning. The results obtained show an improvement in the performance of the recognizer achieved by reducing the number of insertions with no significant reduction in the correctly detected words. M. Carmen Benítez, Antonio J. Rubio, Pedro García-Teodoro, Jesús Esteban Díaz Verdejo |
ICSLP | 1 |
| 1998 | On the comparison of speech recognition tasks
Pedro García-Teodoro, Antonio J. Rubio, M. Carmen Benítez, Jesús Esteban Díaz Verdejo, Juan M. López-Soler |
LREC | 3 |
| 1998 | Speech recognitíon and the new technologies in conununication: a continuous speech recognition-based switchboard with an answering module for e-mailing voice messages
Pedro García-Teodoro, José C. Segura, Jesús Esteban Díaz Verdejo, Antonio José Rubio Ayuso, M. Carmen Benítez, Antonio M. Peinado, Juan M. López-Soler, José Luis Pérez Iglesias, Victoria E. Sánchez Calle, Ángel de la Torre, Ramón López-Cózar |
LREC | 5 |
| 1997 | STACC: an automatic service for information access using continuous speech recognition through telephone lineabstractThis work presents the STACC, Sistema Telef onico Autom atico de Consulta de Calificaciones (Automatic Telephone System for Consulting Marks). This system has been developed at our laboratory during 1996 and implements a service through telephone line that allows the students to consult by speech their marks after the exams by means of a simple phone call. This experience provided us an interesting point of view about the problems of real applications of speech technology. In this work we describe the system and some statistics about the use of STACC by the students are presented. Antonio J. Rubio, Pedro García-Teodoro, Ángel de la Torre, José C. Segura, Jesús Esteban Díaz Verdejo, M. Carmen Benítez, Victoria E. Sánchez, Antonio M. Peinado, Juan M. López-Soler, José L. Pérez-Córdoba |
EUROSPEECH | 6 |
| 1994 | Using multiple vector quantization and semicontinuous hidden Markov models for speech recognitionabstractAlthough the continuous HMM (CHMM) technique seems to be the most flexible and complete tool for speech modeling, it is not always used for the implementation of speech recognition systems due to several problems related to training and computational complexity. Besides, it is not clear the superiority of continuous models over other well-known types of HMMs, such as discrete (DHMM) or semicontinuous (SCHMM) models, or multiple vector quantization (MVQ) models, a new type of HMM modeling. The authors propose a new variant of HMM models, the SCMVQ, HMM models (semicontinuous multiple vector quantization HMM), that uses one VQ codebook per recognition unit and several quantization candidates, Formally, SCMVQ modeling is the closest one to CHMM, although requiring less computation than SCHMMs. Besides, the authors show that SCMVQs can obtain better recognition results than DHMMs, SCHMMs or MVQs.> Antonio M. Peinado, José C. Segura, Antonio J. Rubio, M. Carmen Benítez |
ICASSP (1) | 4 |