Javier Tejedor

dblp:97/4197 · DBLP profile ↗
← Back
30ranked-venue papers
9as first author
5since 2021 · last 2025
0000-0001-7699-5620ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 6 first-author · 4 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Speech recognition and synthesis · 48% Efficient and distributed learning · 30% Language models and text generation · 22%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
0.712023
Exploiting Low-Rank Tensor-Train Deep Neural Networks Based on Riemannian Gradient Descent With Illustrations of Speech Processing · IEEE ACM Trans. Audio Speech Lang. Process. 2023
Natural language and speech › Speech recognition and synthesis
speech enhancement
0.712023
Exploiting Low-Rank Tensor-Train Deep Neural Networks Based on Riemannian Gradient Descent With Illustrations of Speech Processing · IEEE ACM Trans. Audio Speech Lang. Process. 2023
Natural language and speech › Language models and text generation
language modeling
0.212016
Similar Word Model for Unfrequent Word Enhancement in Speech Recognition · IEEE ACM Trans. Audio Speech Lang. Process. 2016
Natural language and speech › Language models and text generation › language modeling
n-gram language model
0.212016
Similar Word Model for Unfrequent Word Enhancement in Speech Recognition · IEEE ACM Trans. Audio Speech Lang. Process. 2016
Natural language and speech › Speech recognition and synthesis › automatic speech recognition › keyword spotting
speech command recognition
0.212023
Exploiting Low-Rank Tensor-Train Deep Neural Networks Based on Riemannian Gradient Descent With Illustrations of Speech Processing · IEEE ACM Trans. Audio Speech Lang. Process. 2023
Natural language and speech › Speech recognition and synthesis
acoustic modeling
0.112012
Comparison of methods for language-dependent and language-independent query-by-example spoken term detection · ACM Trans. Inf. Syst. 2012
Information retrieval › document retrieval › spoken document retrieval
spoken term detection
0.112012
Comparison of methods for language-dependent and language-independent query-by-example spoken term detection · ACM Trans. Inf. Syst. 2012
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.112016
Similar Word Model for Unfrequent Word Enhancement in Speech Recognition · IEEE ACM Trans. Audio Speech Lang. Process. 2016
Information retrieval
indexing
0.012012
Comparison of methods for language-dependent and language-independent query-by-example spoken term detection · ACM Trans. Inf. Syst. 2012

Methods — techniques the papers use, named apart from their topics

tensor-train decomposition · 0.7riemannian gradient descent · 0.7convolutional neural network · 0.7weighted finite-state transducer · 0.3neural network · 0.3dynamic time warping · 0.3similar word model · 0.2class-based language models · 0.2gaussian mixture model/hidden markov model · 0.1gaussian mixture model-hidden markov model · 0.1
YearPublicationVenuePosition
2025 A hybrid cascade-parallel discriminative-generative model for pipeline integrity threat detection in a smart fiber optic surveillance system
Javier Tejedor, Javier Macías Guarasa, Hugo F. Martins, Sonia Martin-Lopez, Miguel González-Herráez
Multim. Tools Appl.1
2023 Optimizing Quantum Federated Learning Based on Federated Quantum Natural Gradient Descent
abstract
Quantum federated learning (QFL) is a quantum extension of the classical federated learning model across multiple local quantum devices. An efficient optimization algorithm is always expected to minimize the communication overhead among different quantum participants. In this work, we propose an efficient optimization algorithm, namely federated quantum natural gradient descent (FQNGD), and further, apply it to a QFL framework that is com-posed of a variational quantum circuit (VQC)-based quantum neural networks (QNN). Compared with stochastic gradient descent methods like Adam and Adagrad, the FQNGD algorithm admits much fewer training iterations for the QFL to get converged. Moreover, it can significantly reduce the total communication overhead among local quantum devices. Our experiments on a handwritten digit classification dataset justify the effectiveness of the FQNGD for the QFL framework in terms of a faster convergence rate on the training set and higher accuracy on the test set.
Jun Qi 0002, Xiao-Lei Zhang 0001, Javier Tejedor
ICASSP3
2023 Exploiting Low-Rank Tensor-Train Deep Neural Networks Based on Riemannian Gradient Descent With Illustrations of Speech Processing
abstract
This work focuses on designing low-complexity hybrid tensor networks by considering trade-offs between the model complexity and practical performance. Firstly, we exploit a low-rank tensor-train deep neural network (TT-DNN) to build an end-to-end deep learning pipeline, namely LR-TT-DNN. Secondly, a hybrid model combining LR-TT-DNN with a convolutional neural network (CNN), which is denoted as CNN+(LR-TT-DNN), is set up to boost the performance. Instead of randomly assigning large TT-ranks for TT-DNN, we leverage Riemannian gradient descent to determine a TT-DNN associated with small TT-ranks. Furthermore, CNN+(LR-TT-DNN) consists of convolutional layers at the bottom for feature extraction and several TT layers at the top to solve regression and classification problems. We separately assess the LR-TT-DNN and CNN+(LR-TT-DNN) models on speech enhancement and spoken command recognition tasks. Our empirical evidence demonstrates that the LR-TT-DNN and CNN+(LR-TT-DNN) models with fewer model parameters can outperform the TT-DNN and CNN+(TT-DNN) counterparts.
Jun Qi 0002, Chao-Han Huck Yang, Javier Tejedor
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 Exploiting Hybrid Models of Tensor-Train Networks For Spoken Command Recognition
abstract
This work aims to design a low complexity spoken command recognition (SCR) system by considering different trade-offs between the number of model parameters and classification accuracy. More specifically, we exploit a deep hybrid architecture of a tensor-train (TT) network to build an end-to-end SRC pipeline. Our command recognition system, namely CNN+(TT-DNN), is composed of convolutional layers at the bottom for spectral feature extraction and TT layers at the top for command classification. Compared with a traditional end-to-end CNN baseline for SCR, our proposed CNN+(TTDNN) model replaces fully connected (FC) layers with TT ones and it can substantially reduce the number of model parameters while maintaining the baseline performance of the CNN model. We initialize the CNN+(TT-DNN) model in a randomized manner or based on a well-trained CNN+DNN, and assess the CNN+(TT-DNN) models on the Google Speech Command Dataset. Our experimental results show that the proposed CNN+(TT-DNN) model attains a competitive accuracy of 96.31% with 4 times fewer model parameters than the CNN model. Furthermore, the CNN+(TT-DNN) model can obtain a 97.2% accuracy when the number of parameters is increased.
Jun Qi 0002, Javier Tejedor
ICASSP2
2022 Classical-To-Quantum Transfer Learning for Spoken Command Recognition Based on Quantum Neural Networks
abstract
This work investigates an extension of transfer learning applied in machine learning algorithms to the emerging hybrid end-to-end quantum neural network (QNN) for spoken command recognition (SCR). Our QNN-based SCR system is composed of classical and quantum components: (1) the classical part mainly relies on a 1D convolutional neural network (CNN) to extract speech features; (2) the quantum part is built upon the variational quantum circuit with a few learnable parameters. Since it is inefficient to train the hybrid end-to-end QNN from scratch on a noisy intermediate-scale quantum (NISQ) device, we put forth a hybrid transfer learning algorithm that allows a pre-trained classical network to be transferred to the classical part of the hybrid QNN model. The pre-trained classical network is further modified and augmented through jointly fine-tuning with a variational quantum circuit (VQC). The hybrid transfer learning methodology is particularly attractive for the task of QNN-based SCR because low-dimensional classical features are expected to be encoded into quantum states. We assess the hybrid transfer learning algorithm applied to the hybrid classical-quantum QNN for SCR on the Google speech command dataset, and our classical simulation results suggest that the hybrid transfer learning can boost our baseline performance on the SCR task.
Jun Qi 0002, Javier Tejedor
ICASSP2
2020 Submodular Rank Aggregation on Score-Based Permutations for Distributed Automatic Speech Recognition
abstract
Distributed automatic speech recognition (ASR) requires to aggregate outputs of distributed deep neural network (DNN)-based models. This work studies the use of submodular functions to design a rank aggregation on score-based permutations, which can be used for distributed ASR systems in both supervised and unsupervised modes. Specifically, we compose an aggregation rank function based on the Lovasz Bregman divergence for setting up linear structured convex and nested structured concave functions. The algorithm is based on stochastic gradient descent (SGD) and can obtain well-trained aggregation models. Our experiments on the distributed ASR system show that the submodular rank aggregation can obtain higher speech recognition accuracy than traditional aggregation methods like Adaboost. Code is available online1.
Jun Qi 0002, Chao-Han Huck Yang, Javier Tejedor
ICASSP3
2018 Distributed Submodular Maximization for Large Vocabulary Continuous Speech Recognition
abstract
Huge training datasets for automatic speech recognition (ASR) typically contain redundant information so that a subset of data is generally enough to obtain similar ASR performance to that obtained when the entire dataset is employed for training. Although the centralized submodular-based data selection methods have been successfully applied to obtain a representable subset involving the most significant information of the whole dataset, the submodular data selection conveys problems in adapting to an extremely massive dataset. This paper proposes to use distributed submodular maximization (DSM) for efficiently selecting a data subset that maintains the ASR performance, while reducing tremendously the computational overhead. There are two approaches for the distributed submodular maximization problem: one is based on an homogeneous submodular function, and the other relies on decomposable submodular functions in which heterogeneous submodular functions are applied. Our experiments show that the data subset output by the DSM algorithms can maintain the ASR performance, while significantly reducing the computational overhead.1
Jun Qi 0002, Xu Liu 0017, Shunsuke Kamijo, Javier Tejedor
ICASSP4
2016 Deep multi-view representation learning for multi-modal features of the schizophrenia and schizo-affective disorder
abstract
This work is originated from the MLSP 2014 Classification Challenge which tries to automatically detect subjects with schizophrenia and schizo-affective disorder by analyzing multi-modal features derived from magnetic resonance imaging (MRI) data. We employ Deep Neural Network (DNN)-based multi-view representation learning for combining multimodal features. The DNN-based multi-view models include deep canonical correlation analysis (DCCA) and deep canonically correlated auto-encoders (DCCAE). In addition, support vector machine with Gaussian kernel is used to conduct classification with the compact bottleneck features learned by the deep multi-view models. Our experiments on the dataset provided by the MLSP Classification Challenge show that bottleneck features learned via deep multi-view models obtain better results than the trimming features used in the baseline system in terms of the receiver operating characteristic (ROC) area under the curve (AUC).
Jun Qi 0002, Javier Tejedor
ICASSP2
2016 Robust submodular data partitioning for distributed speech recognition
abstract
Distributed deep neural networks are commonly employed for building automatic speech recognition (ASR) systems. In this work, we employ the robust submodular partitioning approach, which aims to split the training data into small disjoint data subsets and use each of these subsets to train a particular deep neural network. Two efficient algorithms are used as robust submodular functions [1], namely `Greedi-Max' and `Minorization-Maximization' [2], which are guaranteed to provide tight approximations to the submodular data partition problem. Experiments on TIMIT database show that each of the distributed neural networks trained by the submodular data subset obtains better results than that trained on any subset of data partitioned in a random way., In addition, multi-class adaboost is effectively used to fuse the outputs of the deep neural networks and provides competitive ASR results compared with the traditional ASR system. Besides, the time incurred by acoustic modeling is significantly reduced, which delivers us further benefits.
Jun Qi 0002, Javier Tejedor
ICASSP2
2016 Similar Word Model for Unfrequent Word Enhancement in Speech Recognition
abstract
The popular n-gram language model (LM) is weak for unfrequent words. Conventional approaches such as class-based LMs pre-define some sharing structures (e.g., word classes) to solve the problem. However, defining such structures requires prior knowledge, and the context sharing based on these structures is generally inaccurate. This paper presents a novel similar word model to enhance unfrequent words. In principle, we enrich the context of an unfrequent word by borrowing context information from some “similar words.” Compared to conventional class-based methods, this new approach offers a fine-grained context sharing by referring to words that best match the target word, and it is more flexible as no sharing structures need to be defined by hand. Experiments on a large-scale Chinese speech recognition task demonstrated that the similar word approach can improve performance on unfrequent words significantly, while keeping the performance on general tasks almost unchanged.
Xi Ma, Dong Wang 0013, Javier Tejedor
IEEE ACM Trans. Audio Speech Lang. Process.3
2014 A rule-based translation from written Spanish to Spanish Sign Language glosses
Jordi Porta, Fernando J. López-Colino, Javier Tejedor, José Colás Pasamontes
Comput. Speech Lang.3
2014 Feature analysis for discriminative confidence estimation in spoken term detection
Javier Tejedor, Doroteo T. Toledano, Dong Wang 0013, Simon King 0001, José Colás Pasamontes
Comput. Speech Lang.1
2013 Subspace models for bottleneck features
abstract
The bottleneck (BN) feature, particularly based on deep structures, has gained significant success in automatic speech recognition (ASR). However, applying the BN feature to small/medium-scale tasks is nontrivial. An obvious reason is that the limited training data prevent from training a complicated deep network; another reason, which is more subtle, is that the BN feature tends to possess high inter-dimensional correlation, thus being inappropriate to be modeled by the conventional diagonal Gaussian mixture model (GMM). This difficulty can be mitigated by increasing the number of Gaussian components and/or employing full covariance matrices. These approaches, however, are not applicable for small/medium-scale tasks for which only a limited amount of training data is available. In this paper, we study the subspace Gaussian mixture model (SGMM) for BN features. The SGMM assumes full but shared covariance matrices, and hence can address the interdimensional correlation in a parsimonious way. This is particularly attractive for the BN feature, especially on small/mediumscale tasks, where the inter-dimensional correlation is high but the full covariance modeling is not affordable due to the limited training data. Our preliminary experiments on the Resource Management (RM) database demonstrate that the SGMM can deliver significant performance improvement for ASR systems based on BN features.
Jun Qi 0002, Dong Wang 0013, Javier Tejedor
INTERSPEECH3
2013 Bottleneck features based on gammatone frequency cepstral coefficients
abstract
Recent work demonstrates impressive success of the bottleneck (BN) feature in speech recognition, particularly with deep networks plus appropriate pre-training. A widely admitted advantage associated with the BN feature is that the network structure can learn multiple environmental conditions with abundant training data. For tasks with limited training data, however, this multi-condition training is unavailable, and so the networks tend to be over-fitted and sensitive to acoustic condition changes. A possible solution is to base the BN features on a channel-robust primary feature. In this paper, we propose to derive the BN feature based on Gammatone frequency cepstral coefficients (GFCCs). The GFCC feature has shown nice robustness against acoustic change, due to its capability of simulating the auditory system of humans. The idea is to integrate the advantage of the GFCC feature in acoustic robustness and the advantage of the BN feature in signal representation, so that the BN feature can be improved in the condition of mismatched training/test channels. This is particularly useful for small-scale tasks for which the training data are often limited. The experiments are conducted on the WSJCAM0 database, where the test utterances are mixed with noises at various SNR levels to simulate the channel change. The results confirm that the GFCC-based BN feature is much more robust than the BN features based on the MFCC and the PLP. Furthermore, the primary GFCC feature and the GFCC-based BN feature can be concatenated, leading to a more robust combined feature which provides considerable performance gains in all the tested noise conditions.
Jun Qi 0002, Dong Wang 0013, Javier Tejedor
INTERSPEECH4
2013 Sequential model adaptation for speaker verification
abstract
GMM-UBM-based speaker verification heavily relies on well-trained UBMs.In practice, it is not often easy to obtain a UBM that fully matches the acoustic channel in operation.In a previous study, we proposed to address this problem by a novel sequential UBM adaptation approach based on MAP.This work extends the study by applying the sequential approach to speaker model adaptation.In addition, we investigate a new feature-space sequential adaptation approach based on feature MAP linear regression (fMAPLR) and compare it with the previously proposed model-space MAP approach.We find that these two approaches are complementary and can be combined to deliver additional performance gains.The experiments conducted on a time-varying speech database demonstrate that the proposed MAP-fMAPLR approach leads to significant EER reduction with two mismatched UBMs (25% and 39% respectively).
Jun Wang 0073, Dong Wang 0013, Thomas Fang Zheng, Javier Tejedor
INTERSPEECH5
2013 Evolutionary discriminative confidence estimation for spoken term detection
Javier Tejedor, Alejandro Echeverría, Dong Wang 0013, Ravichander Vipperla
Multim. Tools Appl.1
2012 The Spoken Web Search Task at MediaEval 2011
abstract
In this paper, we describe the “Spoken Web Search” Task, which was held as part of the 2011 MediaEval benchmark campaign. The purpose of this task was to perform audio search with audio input in four languages, with very few resources being available in each language. The data was taken from “spoken web” material collected over mobile phone connections by IBM India. We present results from several independent systems, developed by five teams and using different approaches, compare them, and provide analysis and directions for future research.
Florian Metze, Nitendra Rajput, Xavier Anguera Miró, Marelie H. Davel, Guillaume Gravier, Charl Johannes van Heerden, Gautam Varma Mantena, Armando Muscariello, Kishore Prahallad, Igor Szöke, Javier Tejedor
ICASSP11
2012 N-gram FST Indexing for Spoken Term Detection
abstract
An efficient indexing scheme is essentially important for spoken term detection (STD) on large databases, particularly for phone-based systems that have been widely adopted to achieve vocabulary-independent detection. While the finite state transducer (FST) composition provides a standard indexing approach, the n-gram reverse indexing is more flexible in connectivity representation and confidence measuring and therefore may result in better performance than searching within the original lattices or the equivalent FSTs. In this paper we present an n-gram FST indexing approach which combines the flexibility of n-gram indexing and the efficiency of FST indexing. Specifically, we employ the n-gram indexing to relax connectivity in original lattices and then formalize the indices into an FST for online search. We demonstrate this approach with a phone-based STD task where the lattice is sparse due to strong language models. The results show that n-gram FST indexing provides not only better detection performance than lattice search, but also a faster detection than both conventional n-gram and FST indexing. Index Terms: spoken term indexing, finite state transducer, spoken term detection, speech recognition
Chao Liu 0066, Dong Wang 0013, Javier Tejedor
INTERSPEECH3
2012 An On-Line, Cloud-Based Spanish-Spanish Sign Language Translation System
Javier Tejedor, Fernando J. López-Colino, Jordi Porta, José Colás Pasamontes
INTERSPEECH1
2012 Heterogeneous Convolutive Non-Negative Sparse Coding
abstract
Convolutive non-negative matrix factorization (CNMF) and its sparse version, convolutive non-negative sparse coding (CNSC), exhibit great success in speech processing. A particular limitation of the current CNMF/CNSC approaches is that the convolution ranges of the bases in learning are identical, resulting in patterns covering the same time span. This is obvious unideal as most of sequential signals, for example speech, involve patterns with a multitude of time spans. This paper extends the CMNF/CNSC algorithm and presents a heterogeneous learning approach which can learn bases with non-uniformed convolution ranges. The validity of this extension is demonstrated with a simple speech separation task
Dong Wang 0013, Javier Tejedor
INTERSPEECH2
2012 Term-Dependent Confidence Normalisation for Out-of-Vocabulary Spoken Term Detection
Dong Wang 0013, Javier Tejedor, Simon King 0001, Joe Frankel
J. Comput. Sci. Technol.2
2012 Comparison of methods for language-dependent and language-independent query-by-example spoken term detection
abstract
This article investigates query-by-example (QbE) spoken term detection (STD), in which the query is not entered as text, but selected in speech data or spoken. Two feature extractors based on neural networks (NN) are introduced: the first producing phone-state posteriors and the second making use of a compressive NN layer. They are combined with three different QbE detectors: while the Gaussian mixture model/hidden Markov model (GMM/HMM) and dynamic time warping (DTW) both work on continuous feature vectors, the third one, based on weighted finite-state transducers (WFST), processes phone lattices. QbE STD is compared to two standard STD systems with text queries: acoustic keyword spotting and WFST-based search of phone strings in phone lattices. The results are reported on four languages (Czech, English, Hungarian, and Levantine Arabic) using standard metrics: equal error rate (EER) and two versions of popular figure-of-merit (FOM). Language-dependent and language-independent cases are investigated; the latter being particularly interesting for scenarios lacking standard resources to train speech recognition systems. While the DTW and GMM/HMM approaches produce the best results for a language-dependent setup depending on the target language, the GMM/HMM approach performs the best dealing with a language-independent setup. As far as WFSTs are concerned, they are promising as they allow for indexing and fast search.
Javier Tejedor, Michal Fapso, Igor Szöke, Jan Cernocký, Frantisek Grézl
ACM Trans. Inf. Syst.1
2010 Augmented set of features for confidence estimation in spoken term detection
abstract
Discriminative confidence estimation along with confidence normalisation have been shown to construct robust decision maker modules in spoken term detection (STD) systems. Discriminative confidence estimation, making use of termdependent features, has been shown to improve the widely used lattice-based confidence estimation in STD. In this work, we augment the set of these term-dependent features and show a significant improvement in the STD performance both in terms of ATWV and DET curves in experiments conducted on a Spanish geographical corpus. This work also proposes a multiple linear regression analysis to carry out the feature selection. Next, the most informative features derived from it are used within the discriminative confidence on the STD system.
Javier Tejedor, Doroteo T. Toledano, Miguel Bautista, Simon King 0001, Dong Wang 0013, José Colás Pasamontes
INTERSPEECH1
2009 Posterior-based confidence measures for spoken term detection
abstract
Confidence measures play a key role in spoken term detection (STD) tasks. The confidence measure expresses the posterior probability of the search term appearing in the detection period, given the speech. Traditional approaches are based on the acoustic and language model scores for candidate detections found using automatic speech recognition, with Bayes' rule being used to compute the desired posterior probability. In this paper, we present a novel direct posterior-based confidence measure which, instead of resorting to the Bayesian formula, calculates posterior probabilities from a multi-layer perceptron (MLP) directly. Compared with traditional Bayesian-based methods, the direct-posterior approach is conceptually and mathematically simpler. Moreover, the MLP-based model does not require assumptions to be made about the acoustic features such as their statistical distribution and the independence of static and dynamic co-efficients. Our experimental results in both English and Spanish demonstrate that the proposed direct posterior-based confidence improves STD performance.
Dong Wang 0013, Javier Tejedor, Joe Frankel, Simon King 0001, José Colás Pasamontes
ICASSP2
2009 A posterior probability-based system hybridisation and combination for spoken term detection
abstract
Spoken term detection (STD) is a fundamental task for multimedia information retrieval. To improve the detection performance, we have presented a direct posterior-based confidence measure generated from a neural network. In this paper, we propose a detection-independent confidence estimation based on the direct posterior confidence measure, in which the decision making is totally separated from the term detection. Based on this idea, we first present a hybrid system which conducts the term detection and confidence estimation based on different sub-word units and then propose a combination method which merges detections from heterogeneous term detectors based on the direct posterior-based confidence. Experimental results demonstrated that the proposed methods improved system performance considerably for both English and Spanish. Index Terms: speech recognition, spoken term detection, confidence estimation, grapheme
Javier Tejedor, Dong Wang 0013, Simon King 0001, Joe Frankel, José Colás Pasamontes
INTERSPEECH1
2008 A comparison of phone and grapheme-based spoken term detection
abstract
We propose grapheme-based sub-word units for spoken term detection (STD). Compared to phones, graphemes have a number of potential advantages. For out-of-vocabulary search terms, phone- based approaches must generate a pronunciation using letter-to-sound rules. Using graphemes obviates this potentially error-prone hard decision, shifting pronunciation modelling into the statistical models describing the observation space. In addition, long-span grapheme language models can be trained directly from large text corpora. We present experiments on Spanish and English data, comparing phone and grapheme-based STD. For Spanish, where phone and grapheme-based systems give similar transcription word error rates (WERs), grapheme-based STD significantly outperforms a phone- based approach. The converse is found for English, where the phone- based system outperforms a grapheme approach. However, we present additional analysis which suggests that phone-based STD performance levels may be achieved by a grapheme-based approach despite lower transcription accuracy, and that the two approaches may usefully be combined. We propose a number of directions for future development of these ideas, and suggest that if grapheme-based STD can match phone-based performance, the inherent flexibility in dealing with out-of-vocabulary terms makes this a desirable approach.
Dong Wang 0013, Joe Frankel, Javier Tejedor, Simon King 0001
ICASSP3
2008 rre STC-TIMIT: Generation of a Single-channel Telephone Corpus
Nicolás Morales, Javier Tejedor, Javier Garrido Salas, José Colás Pasamontes, Doroteo T. Toledano
LREC2
2008 A comparison of grapheme and phoneme-based units for Spanish spoken term detection
Javier Tejedor, Dong Wang 0013, Joe Frankel, Simon King 0001, José Colás Pasamontes
Speech Commun.1
2004 Multimodal Control Centre for Handicapped People
Ana Granados, Víctor Tomico, Eduardo Campos, Javier Tejedor, Javier Garrido Salas
ICCHP4
2004 Multimedia Medicine Consultant for Visually Impaired People
Javier Tejedor, Daniel Bolaños, Nicolás Morales, Ana Granados, José Colás Pasamontes, Santiago Aguilera
ICCHP1