VLDB 2026 Research / reviewers in the wild / expert
José L. Pérez-Córdoba
dblp:11/5192 · also José Luis Pérez-Córdoba
· DBLP profile ↗
31ranked-venue papers
7as first author
2since 2021 · last 2026
0000-0002-1686-2401ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 7 first-author · 1 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 1 since 2021Computer networks · 2Databases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
2 papers |
Audio and music processing · 100% | |
| Artificial intelligence
2 papers |
Speech recognition and synthesis · 100% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Audio and music processing › speech coding
packet loss concealment |
0.2 | 1 | 2016 | An Error Mitigation Technique for Erasure Channels Based on a Wavelet Representation of the Speech Excitation Signal · IEEE Trans. Multim. 2016 |
Audio and music processing
speech coding |
0.2 | 1 | 2016 | An Error Mitigation Technique for Erasure Channels Based on a Wavelet Representation of the Speech Excitation Signal · IEEE Trans. Multim. 2016 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
robust speech recognition |
0.2 | 2 | 2010 | MMSE-Based Packet Loss Concealment for CELP-Coded Speech Recognition · IEEE Trans. Speech Audio Process. 2010 Histogram equalization of speech representation for robust speech recognition · IEEE Trans. Speech Audio Process. 2005 |
Natural language and speech › Speech recognition and synthesis › speech coding
packet loss concealment |
0.1 | 1 | 2010 | MMSE-Based Packet Loss Concealment for CELP-Coded Speech Recognition · IEEE Trans. Speech Audio Process. 2010 |
Natural language and speech › Speech recognition and synthesis
speech coding |
0.1 | 1 | 2010 | MMSE-Based Packet Loss Concealment for CELP-Coded Speech Recognition · IEEE Trans. Speech Audio Process. 2010 |
Audio and music processing › speech coding
linear predictive coding |
0.1 | 1 | 2016 | An Error Mitigation Technique for Erasure Channels Based on a Wavelet Representation of the Speech Excitation Signal · IEEE Trans. Multim. 2016 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › robust speech recognition
feature compensation |
0.1 | 1 | 2005 | Histogram equalization of speech representation for robust speech recognition · IEEE Trans. Speech Audio Process. 2005 |
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
distributed speech recognition |
0.0 | 1 | 2010 | MMSE-Based Packet Loss Concealment for CELP-Coded Speech Recognition · IEEE Trans. Speech Audio Process. 2010 |
Audio and music processing › speech recognition
acoustic modeling |
0.0 | 1 | 1996 | Discriminative codebook design using multiple vector quantization in HMM-based speech recognizers · IEEE Trans. Speech Audio Process. 1996 |
Audio and music processing
speech recognition |
0.0 | 1 | 1996 | Discriminative codebook design using multiple vector quantization in HMM-based speech recognizers · IEEE Trans. Speech Audio Process. 1996 |
Methods — techniques the papers use, named apart from their topics
minimum mean square error estimation · 0.4haar wavelet transform · 0.2codebook · 0.2hidden markov model · 0.1weighted viterbi algorithm · 0.1soft-data decoding · 0.1histogram equalization · 0.1maximum mutual information estimation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UGR-MINDVOICE: A multimodal EEG-audio dataset for overt and covert Iberian Spanish speech productionabstractWe present UGR-MINDVOICE, the University of Granada (UGR) multimodal electroencephalography (EEG) and audio dataset for overt and covert speech in Iberian Spanish intended for basic neuroscience and brain-computer interface (BCI) research. The dataset features EEG and audio recordings from 15 native Spanish speakers engaged in both overt and covert speech production tasks. This dataset is unique in its inclusion of all Spanish phonemes and a diverse set of words spanning various semantic categories and different usage frequencies. Validation of the dataset confirmed the presence of robust sensory event-related potentials, including the visual P100 and the auditory N1 (N100), indicating reliable early perceptual processing and sustained participant attention to both visual and auditory stimuli. Additionally, the EEG data were classified into rest, covert speech, and overt speech conditions with an accuracy of 81.40%, demonstrating active participant engagement in the tasks. By providing synchronised EEG and audio data for overt speech, along with EEG data for the same stimuli during covert speech, UGR-MINDVOICE constitutes a valuable resource for advancing research in basic neuroscience and brain-computer interfaces, particularly in the domain of silent speech communication. The full dataset is openly available on the Open Science Framework (OSF) ( https://osf.io/6sh5d ), and all accompanying code and analysis scripts are provided in a public GitHub repository ( https://github.com/owaismujtaba/mind-voice ). Ibon Vales Cortina, Owais Mujtaba Khanday, Marc Ouellet, José L. Pérez-Córdoba, Pablo Rodríguez San Esteban, Laura Miccoli, Alberto Galdón, Gonzalo Olivares Granados, José A. González 0001 |
Comput. Speech Lang. | 4 |
| 2025 | NeuroIncept Decoder for High-Fidelity Speech Reconstruction from Neural ActivityabstractThis paper introduces a novel algorithm designed for speech synthesis from neural activity recordings obtained using invasive electroencephalography (EEG) techniques. The proposed system offers a promising communication solution for individuals with severe speech impairments. Central to our approach is the integration of time-frequency features in the high-gamma band computed from EEG recordings with an advanced NeuroIncept Decoder architecture. This neural network architecture combines Convolutional Neural Networks (CNNs) and Gated Recurrent Units (GRUs) to reconstruct audio spectrograms from neural patterns. Our model demonstrates robust mean correlation coefficients between predicted and actual spectrograms, though inter-subject variability indicates distinct neural processing mechanisms among participants. Overall, our study highlights the potential of neural decoding techniques to restore communicative abilities in individuals with speech disorders and paves the way for future advancements in brain-computer interface technologies. Owais Mujtaba Khanday, José L. Pérez-Córdoba, Mohd Yaqub Mir, Ashfaq Ahmad Najar, José A. González 0001 |
ICASSP | 2 |
| 2018 | Speech excitation signal recovering based on a novel error mitigation scheme under erasure channel conditions
Domingo López-Oller, Nadir Benamirouche, Ángel M. Gómez, José L. Pérez-Córdoba |
Speech Commun. | 4 |
| 2016 | System-compatible robustness improvement for new generation dect decoders by G.722 soft-decision decodingabstractThe ITU-T Recommendation G.722 about subband adaptive differential pulse code modulation (SB-ADPCM) is the mandatory wideband speech codec in the new generation digital enhanced cordless telephony (NG-DECT). Although in ADPCM the difference signal instead of the original signal is quantized and adaptive prediction is employed, redundancy is yet observed within the quantized samples. In this paper we apply a soft-decision speech decoding technique which exploits this redundancy in terms of a priori knowledge and the channel reliability information to NG-DECT. In that way, we propose a novel scheme in a standard-compliant fashion which improves the robustness of the decoder. The performance of our proposal is evaluated in terms of speech quality and a noticeable improvement over the standard codec and its own packet loss concealment algorithm is observed. Domingo López-Oller, Sai Han, Ángel M. Gómez, José L. Pérez-Córdoba, Tim Fingscheidt |
ICASSP | 4 |
| 2016 | An Error Mitigation Technique for Erasure Channels Based on a Wavelet Representation of the Speech Excitation SignalabstractThe importance of packet-based speech transmissions has grown since it offers cheaper and efficient communications. However, frame erasures are a common hurdle in these networks and concealment techniques are necessary to ensure a minimum quality of service. In this paper, we propose a mitigation technique focused on the reconstruction of the linear prediction coding (LPC) coefficients and the excitation signal of the lost frame by using a replacement technique. These replacements are obtained by means of a minimum mean square error estimation based on a source model of the speech parameters (LPC coefficients and the excitation signal). As this approach critically relies on the quantization and representation of the excitation signal, we explore the Haar wavelet transform as a novel approach to represent the excitation signal for error mitigation. Thus, this paper describes how optimal codebook and estimates can be computed in a Haar transformed domain. As a result, the excitation signal of a frame can be decomposed in several partitions where each one is independently reconstructed. Objective and subjective tests are conducted in order to assess the quality of the concealed speech signal resulting from our proposal. Both evaluations confirm noticeable improvements over the default mitigation method included in the two tested standard codecs, adaptive multirate, and Internet low bitrate codec. Domingo López-Oller, Ángel M. Gómez, José L. Pérez-Córdoba, Victoria E. Sánchez |
IEEE Trans. Multim. | 3 |
| 2013 | Backwards-compatible error propagation recovery for the amr codec over erasure channelsabstractThis paper presents a recovery scheme for the error-propagation distortion which frequently appears after a frame erasure in CELP-based speech coders, in particular the AMR codec. The extensive use of predictive filters and parameter encoding allow a high-quality speech synthesis in these codecs, but makes them more vulnerable to frame erasures. Thus, when a frame is lost, an additional distortion appears in the subsequent frame, although that was correctly received, further degrading the speech quality. This degradation can also propagate over several frames, being even more damaging than the loss itself. This well known fact has motivated the development of techniques which prevent or mitigate the error propagation. Nevertheless, the previously proposed methods in some respect modify the transmission scheme (by including additional frames, FEC codes, etc.) making them incompatible with the original decoder. In this work, we apply a steganographic technique to embed recovery data to assist the decoder after a frame loss. This data mainly consist of resynchronization pulses and correction vectors for the excitation signal and the spectral envelope, respectively. PESQ results confirm that our proposal achieves a higher robustness against error propagation while the full backwards-compatibility with the AMR standard is retained. Ángel M. Gómez, José L. Pérez-Córdoba, Bernd Geiser |
ICASSP | 2 |
| 2010 | A multipulse FEC scheme based on amplitude estimation for CELP codecs over packet networks
José L. Carmona, Ángel M. Gómez, Antonio M. Peinado, José L. Pérez-Córdoba, José A. González 0001 |
INTERSPEECH | 4 |
| 2010 | MMSE-Based Packet Loss Concealment for CELP-Coded Speech RecognitionabstractIn this paper, we analyze the performance of network speech recognition (NSR) over IP networks, adapting and proposing new solutions to the packet loss problem for code excited linear prediction (CELP) codecs. NSR has a client-server architecture which places the recognizer at the server side using a standard speech codec for speech transmission. Its main advantage is that no changes are required for the existing client devices and networks. However, the use of speech codecs degrades its performance, mainly in the presence of packet losses. First, we study the degradations introduced by CELP codecs in lossy packet networks. Later, we propose a reconstruction technique based on minimum mean square error (MMSE) estimation using hidden Markov models. This approach also allows us to obtain reliability measures associated to each estimate. We show how to use this information to improve the recognition performance by means of soft-data decoding and weighted Viterbi algorithm. The experimental results are obtained for two well-known CELP codecs, G.729 and AMR 12.2 kbps, carrying out recognition from decoded speech. Finally, we analyze an efficient and improved implementation of the proposed techniques using an NSR system which extracts speech recognition features directly from the bit-stream parameters. The experimental results show that the different proposed NSR systems achieve a comparable performance to distributed speech recognition (DSR). José L. Carmona, Antonio M. Peinado, José L. Pérez-Córdoba, Ángel M. Gómez |
IEEE Trans. Speech Audio Process. | 3 |
| 2008 | A scalable coding scheme based on interframe dependency limitationabstractWhile VoIP (voice over IP) is gaining importance in comparison with other types of telephony, packet loss remains as the main source of degradation in VoIP systems. Traditional speech codecs, such as those based on the CELP (code excited linear prediction) paradigm, can achieve low bit-rates at the cost of introducing interframe dependencies. As a result, the effect of a packet loss burst is propagated to the frames correctly received after the burst. iLBC (internet low bit-rate codec) alleviates this problem by removing the interframe dependencies at the cost of a higher bit-rate. In this paper we propose a combination of iLBC with an ACELP (algebraic CELP) codec in which a variable number of ACELP-coded frames is inserted between every two iLBC-coded frames. The experimental results show that the combined codec can achieve a performance close to that of iLBC at different loss conditions but with a smaller bit-rate. Also, scalability is achieved by modifying the number of inserted ACELP-coded frames. José L. Carmona, José L. Pérez-Córdoba, Antonio M. Peinado, Ángel M. Gómez, José A. González 0001 |
ICASSP | 2 |
| 2007 | iLBC-Based Transparametrization: A Real Alternative to DSR for Speech Recognition Over Packet NetworksabstractThis paper proposes a method for the remote recognition of speech coded with the iLBC codec, which is employed by a number of VoIP systems. While the usual way of performing recognition of coded speech is to decode first the speech signal and use it as input to the recognition engine, our system directly converts the iLBC parameters into recognition features. The main advantage of this approach is to avoid any type of decoding post-processing which, although originally conceived to improve the speech perception, can be harmful for a recognition system. Our method ensures the compatibility between the speech spectra provided by the iLBC codec and those employed for cepstrum computation and introduces a robust and suitable packet loss concealment strategy. Our experimental results show that the proposed system achieves a performance better than that obtained from iLBC-decoded speech and similar to that of a distributed speech recognition system over a clean or degraded transmission channel. José L. Carmona, Antonio M. Peinado, José L. Pérez-Córdoba, Ángel M. Gómez, Victoria E. Sánchez |
ICASSP (4) | 3 |
| 2006 | An integrated solution for error concealment in DSR systems over wireless channels
Antonio M. Peinado, Ángel M. Gómez, Victoria E. Sánchez, José L. Pérez-Córdoba, Antonio J. Rubio |
INTERSPEECH | 4 |
| 2005 | Packet Loss Concealment Based on VQ Replicas and MMSE Estimation Applied to Distributed Speech RecognitionabstractThis paper proposes a new packet loss concealment technique based on the inclusion in each packet of a few FEC bits, representing data replicas, combined with a minimum mean square error estimation (MMSE). This technique is developed for an Aurora-2 distributed speech recognition system working over an IP network. In addition to the data representing the transmitted speech frames, each packet includes some FEC bits representing a strongly VQ-quantized version (replicas) of previous and subsequent frames. When a loss burst occurs, the lost frames can be reconstructed from the VQ replicas. In order to mitigate the degradation introduced by the coarse VQ quantization of the replicas, a model-based MMSE estimation is applied. The experimental results show that, under a strongly degraded channel, it is possible to obtain up to 83.31 % of word accuracy with only 4 FEC bits or 88.47 % with 8 FEC bits per packet, when the Aurora mitigation algorithm only obtains 76.98 %. Antonio M. Peinado, Ángel M. Gómez, Victoria E. Sánchez, José L. Pérez-Córdoba, Antonio J. Rubio |
ICASSP (1) | 4 |
| 2005 | Joint source-channel coding of LSP parameters for bursty channels
José L. Pérez-Córdoba, Antonio M. Peinado, Ángel M. Gómez, Antonio J. Rubio |
INTERSPEECH | 1 |
| 2005 | Histogram equalization of speech representation for robust speech recognitionabstractThis paper describes a method of compensating for nonlinear distortions in speech representation caused by noise. The method described here is based on the histogram equalization method often used in digital image processing. Histogram equalization is applied to each component of the feature vector in order to improve the robustness of speech recognition systems. The paper describes how the proposed method can be applied to robust speech recognition and it is compared with other compensation techniques. The recognition experiments, including results in the AURORA II framework, demonstrate the effectiveness of histogram equalization when it is applied either alone or in combination with other compensation techniques. Ángel de la Torre, Antonio M. Peinado, José C. Segura, José L. Pérez-Córdoba, M. Carmen Benítez, Antonio J. Rubio |
IEEE Trans. Speech Audio Process. | 4 |
| 2005 | Efficient MMSE-based channel error mitigation techniques. Application to distributed speech recognition over wireless channelsabstractThis work addresses the mitigation of channel errors by means of efficient minimum mean-square-error (MMSE) estimation. Although powerful model-based implementations have been recently proposed, the computational burden involved can make them impractical. We propose two new approaches that maintain a good level of performance with a low computational complexity. These approaches keep the simple structure and complexity of a raw MMSE estimation, although they enhance it with additional source a priori knowledge. The proposed techniques are built on a distributed speech recognition system. Different degrees of tradeoff between recognition performance and computational complexity are obtained. Antonio M. Peinado, Victoria E. Sánchez, José L. Pérez-Córdoba, Antonio J. Rubio |
IEEE Trans. Wirel. Commun. | 3 |
| 2004 | Mitigation of channel errors in EFR-based speech recognitionabstractNetwork-based speech recognition (NSR) using the conventional speech channel with the enhanced full rate (EFR) or the adaptive multi-rate (AMR) codec is a very attractive approach since no change to existing mobile phones is needed. However, NSR reveals a degrading performance due to both transmission channel errors and the speech encoding process in comparison with distributed speech recognition (DSR), where speech features are efficiently coded and transmitted on a data channel. We focus on the degradation of the speech features caused by channel errors in an NSR system and propose methods to improve the quality of these features. Applying these methods, it turns out that the performance of an NSR system based on EFR coding is comparable to that based on DSR. Ángel M. Gómez, Antonio M. Peinado, Victoria E. Sánchez, José L. Pérez-Córdoba, Antonio J. Rubio |
ICASSP (1) | 4 |
| 2003 | A study of joint source-channel coding of LSP parameters for wideband speech codingabstractA study of combined source and channel coding applied to LSP parameters in wideband speech coding is presented. The traditional approach to protect against channel errors is to increase the bit-rate for channel coding, decreasing the bit-rate of the source coding according the channel conditions. Joint source-channel coding is an alternative that provides a technique to mitigate channel errors without an increase of the bit-rate due to channel coding. This paper presents a study of channel optimized vector quantizer and channel optimized matrix quantizer applied to line spectral pairs (LSP) parameters in wideband speech coding. Gaussian and slow-fading Rayleigh channels are considered and GMSK (Gaussian minimum shift-keying) is used as the modulation technique. In addition, for comparison purposes, the performance of other schemes (split vector quantization, split matrix quantization and split multistage vector quantization) for quantizing the LSP parameters are evaluated. José L. Pérez-Córdoba, Antonio M. Peinado, Victoria E. Sánchez, Antonio J. Rubio |
ICASSP (2) | 1 |
| 2003 | Low complexity channel error mitigation for distributed speech recognition over wireless channelsabstractDistributed speech recognition (DSR) has been recently proposed as an efficient way of translating automatic speech recognition technologies to mobile and IP network application. In this paper we propose a channel error mitigation technique with a low computational complexity that improves the mitigation technique proposed in the ETSI standard for DSR (ETSI-ES-201-108 v1.12) for bad channel conditions. We also study the influence of the vector quantization index assignment on the proposed mitigation technique and design a new index assignment that gets some improvement on the proposed technique. Victoria E. Sánchez, Antonio M. Peinado, José L. Pérez-Córdoba |
ICC | 3 |
| 2003 | Entropy-optimized channel error mitigation with application to speech recognition over wireless
Victoria E. Sánchez, Antonio M. Peinado, Ángel M. Gómez, José L. Pérez-Córdoba |
INTERSPEECH | 4 |
| 2003 | HMM-based channel error mitigation and its application to distributed speech recognition
Antonio M. Peinado, Victoria E. Sánchez, José L. Pérez-Córdoba, Ángel de la Torre |
Speech Commun. | 3 |
| 2002 | Progressive image transmission over a noisy channel using wavelet transform and channel optimized vector quantizationabstractThis paper studies a progressive image transmission technique over waveform channels. The channel optimized vector quantization codec (COVQ) (Farvardin and Vaishampayan 1991) is applied to the image wavelet coefficients creating a robust progressive image transmission technique that mitigates the effects of a noisy channel on the reconstructed image. In order to evaluate the performance of our proposal, a Gaussian and slow-fading Rayleigh channel model, with several different values of channel signal to noise ratio (CSNR) were simulated in our experiments. Examples show a significant visual improvement of our application compared to other progressive image transmission techniques. José L. Pérez-Córdoba, Vicente González Ruiz, Inmaculada García |
ICIP (2) | 1 |
| 2002 | HMM-based methods for channel error mitigation in distributed speech recognitionabstractDistributed Speech Recognition involves the development of techniques to mitigate the degradations that the transmission channel introduces in the speech features. This work proposes an HMM framework from which different mitigation techniques oriented to bursty channels can be derived. In particular, two MMSE-based and a new Viterbi-based mitigation procedures are derived under this framework. Several implementation issues such as the channel SNR estimation or the application of hard decision on the received signal vectors are dealt with. Also, different boundary conditions suitable for the speech recognition application are studied for the different mitigation procedures. The experimental results show that the HMM-based techniques can effectively mitigate channel errors, even in very poor channel conditions. Antonio M. Peinado, Victoria E. Sánchez, José L. Pérez-Córdoba, José C. Segura, Antonio J. Rubio |
INTERSPEECH | 3 |
| 2001 | Channel optimized matrix quantization (COMQ) of LSP parameters over waveform channelsabstractCombined source and channel coding is a technique to mitigate channel errors without increasing the bit error rate. Channel optimized vector quantizer (COVQ) performs these objectives in the context of vector quantization. This paper presents a study of channel optimized matrix quantizer (COMQ) applied to quantize the line spectral pair (LSP) parameters as an extension of COVQ technique. Gaussian and slow-fading Rayleigh channels are considered and GMSK (Gaussian minimum shift-keying) is used as modulation technique. Several channel signal to noise ratio (CSNR) are considered to measure the performance of this system. In addition, for comparison purposes, the performance of other schemes for quantizing the LSP parameters are computed. José L. Pérez-Córdoba, Antonio J. Rubio, Juan M. López-Soler, Victoria E. Sánchez |
ICASSP | 1 |
| 2001 | MMSE-based channel error mitigation for distributed speech recognitionabstractRecently, the first version of an ETSI standard for Distributed Speech Recognition has been proposed. The main benefit of this approach is the possibility of maintaining a high recognition performance when accessing remote information systems. The use of a digital channel for transmission of the encoded speech parameters implies the introduction of several channel distortions. Our paper deals with the mitigation of such distortions. We study the application of MMSE estimation to this problem and propose a new MMSE procedure that obtains the probabilities needed for MMSE from a forward-backward algorithm. We show that MMSE estimation obtains better performance than the mitigation algorithm described in the ETSI standard under different channel conditions. Antonio M. Peinado, Victoria E. Sánchez, José C. Segura, José L. Pérez-Córdoba |
INTERSPEECH | 4 |
| 2001 | Joint source-channel coding for low bit-rate coding of LSP parametersabstractThis work presents a quantization technique for LSP parameters which results in a low bit-rate transmission while providing protection against channel errors. As a generalization of the so called Channel Optimized Vector Quantization (COVQ), Channel Optimized Matrix Quantization (COMQ) can remove intraframe and interframe LSP redundancy with the target of protecting the information sent through a channel in the presence of noise. Split COMQ is used in order to reduce storage requirements and complexity. Results show that Split COMQ gives better performance under certain error conditions and a lower bit rate transmission in all channel conditions compared to the reference quantization techniques. José L. Pérez-Córdoba, Antonio J. Rubio, Antonio M. Peinado, Ángel de la Torre |
INTERSPEECH | 1 |
| 2000 | Hard-Decision in COVQ over Waveform ChannelsabstractA channel optimized vector quantizer (COVQ) is studied for the case of transmission over waveform channels. In this work, a number of modulation schemes with multidimensional signal constellations are considered, specifically, results on the binary signalling. M-ary phase-shift keying (MPSK) and M-ary quadrature amplitude modulation (MQAM) performance using COVQ with hard-decision decoding, is optimized for additive white Gaussian noise (AWGN) and flat-fading Rayleigh channel. In addition, when a flat-fading Rayleigh channel is assumed, diversity techniques are used and evaluated to improve the performance of the system. José L. Pérez-Córdoba, Antonio J. Rubio, Juan M. López-Soler, M. Carmen Benítez |
Data Compression Conference | 1 |
| 2000 | Channel Optimized Matrix Quantizer (COMQ) in CELP CodingabstractWe present a study of a channel optimized matrix quantizer (COMQ) applied to quantize LSP (line spectral pair) parameters in a CELP (code excited linear prediction) coder. A modification of the DoD FS-1016 standard (Campbell et al., 1989) is used for this purpose. A Gaussian channel is considered as the channel through which information is sent and BPSK (binary phase shift-keying) is used as modulation technique. To measure the performance of this coder several channel signal to noise ratios are considered. Also, the performance of the FS-1016 standard coder is computed for comparison purposes. José L. Pérez-Córdoba, Antonio J. Rubio, Victoria E. Sánchez, Ángel de la Torre |
ICPR | 1 |
| 1997 | STACC: an automatic service for information access using continuous speech recognition through telephone lineabstractThis work presents the STACC, Sistema Telef onico Autom atico de Consulta de Calificaciones (Automatic Telephone System for Consulting Marks). This system has been developed at our laboratory during 1996 and implements a service through telephone line that allows the students to consult by speech their marks after the exams by means of a simple phone call. This experience provided us an interesting point of view about the problems of real applications of speech technology. In this work we describe the system and some statistics about the use of STACC by the students are presented. Antonio J. Rubio, Pedro García-Teodoro, Ángel de la Torre, José C. Segura, Jesús Esteban Díaz Verdejo, M. Carmen Benítez, Victoria E. Sánchez, Antonio M. Peinado, Juan M. López-Soler, José L. Pérez-Córdoba |
EUROSPEECH | 10 |
| 1996 | Discriminative codebook design using multiple vector quantization in HMM-based speech recognizersabstractResearch on multiple vector quantization (MVQ) has shown the suitability of such a technique for speech recognition. Basically, MVQ proposes the use of one separate VQ codebook for each recognition unit. Thus, a MVQHMM model is composed of a VQ codebook and a discrete HMM model. This technique allows the incorporation in the recognition dynamics of the input sequence information wasted by discrete HMM models in the VQ process. The use of distinct codebooks also allows one to train them in a discriminative manner. We propose a new VQ codebook design method for MVQ-based systems, obtained from a modified maximum mutual information estimation. This method provides meaningful error reductions and is performed independently from the estimation of the discrete HMM part of the MVQ model. The results show that the proposed discriminative design turns the MVQHMM technique into a powerful acoustic modeling tool in comparison with other classical methods such as discrete or semicontinuous HMMs. Antonio M. Peinado, José C. Segura, Antonio J. Rubio, Pedro García-Teodoro, José L. Pérez-Córdoba |
IEEE Trans. Speech Audio Process. | 5 |
| 1994 | Transform trellis coded quantization of speech using small frame sizesabstractProposes a transform trellis coded quantization (TTCQ) scheme suitable for low-delay speech coding. This scheme is based on a general transform domain formulation for small frame sizes and the trellis coded quantization technique previously proposed where Ungerboeck's amplitude modulation trellises and set partitioning ideas are used for source coding. Using the discrete cosine transform the authors apply this technique to the coding of 7 kHz wideband speech, a matter of interest due to ISDN based applications. A 32 kbps low-delay coder is developed, simulation results showing a very high speech quality comparable to that obtained by the G722 standard at 64 kbps.> Victoria E. Sánchez, José L. Pérez-Córdoba, Juan M. López-Soler, Antonio J. Rubio |
ICASSP (1) | 2 |
| 1993 | A new neuron model for an Alphanet-semicontinuous HMM
Jesús Esteban Díaz Verdejo, José C. Segura, Antonio J. Rubio, Antonio M. Peinado, José L. Pérez-Córdoba |
ICASSP (1) | 5 |