VLDB 2026 Research / reviewers in the wild / expert
Torbjørn Svendsen
dblp:15/7157 · also Torbjørn Karl Svendsen
· DBLP profile ↗
68ranked-venue papers
8as first author
14since 2021 · last 2025
0000-0003-0578-7941ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 54 · 8 first-author · 10 since 2021Artificial intelligence and machine learning · 52 · 3 first-author · 11 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Effects of Prosodic Information on Dialect Classification Using Whisper Features
Phoebe Parsons, Heming Strømholt Bremnes, Knut Kvale, Torbjørn Svendsen, Giampiero Salvi |
INTERSPEECH | 4 |
| 2024 | Collecting Linguistic Resources for Assessing Children's Pronunciation of Nordic LanguagesabstractThis paper reports on the experience collecting a number of corpora of Nordic languages spoken by children. The aim of the data collection is providing annotated data to develop and evaluate computer assisted pronunciation assessment systems both for non-native children learning a Nordic language (L2) and for L1 children with speech sound disorder (SSD). The paper presents the challenges encountered recording and annotating data for Finnish, Swedish and Norwegian, as well as the ethical considerations related with making this data publicly available. We hope that sharing this experience will encourage others to collect similar data for other languages. Of the different data collections, we were able to make the Norwegian corpus publicly available in the hope that it will serve as a reference in pronunciation assessment research. Anne Marte Haug Olstad, Anna-Riikka Smolander, Sofia Strömbergsson, Sari Ylinen, Minna Lehtonen, Mikko Kurimo, Yaroslav Getman, Tamás Grósz, Xinwei Cao, Torbjørn Svendsen, Giampiero Salvi |
LREC/COLING | 10 |
| 2024 | A Framework for Phoneme-Level Pronunciation Assessment Using CTC
Xinwei Cao, Zijian Fan, Torbjørn Svendsen, Giampiero Salvi |
INTERSPEECH | 3 |
| 2024 | Exploiting Foundation Models and Speech Enhancement for Parkinson's Disease Detection from Speech in Real-World Operative ConditionsabstractThis work is concerned with devising a robust Parkinson’s (PD) disease detector from speech in real-world operating conditions using (i) foundational models, and (ii) speech enhancement (SE) methods. To this end, we first fine-tune several foundational-based models on the standard PC-GITA (s-PCGITA) clean data. Our results demonstrate superior performance to previously proposed models. Second, we assess the generalization capability of the PD models on the extended PCGITA (e-PC-GITA) recordings, collected in real-world operative conditions, and observe a severe drop in performance moving from ideal to real-world conditions. Third, we align training and testing conditions applaying off-the-shelf SE techniques on e-PC-GITA, and a significant boost in performance is observed only for the foundational-based models. Finally, combining the two best foundational-based models trained on s-PCGITA, namely WavLM Base and Hubert Base, yielded top performance on the enhanced e-PC-GITA Moreno La Quatra, Maria Francesca Turco, Torbjørn Svendsen, Giampiero Salvi, Juan Rafael Orozco-Arroyave, Sabato Marco Siniscalchi |
INTERSPEECH | 3 |
| 2024 | On the Predictive Power of Objective Intelligibility Metrics for the Subjective Performance of Deep Complex Convolutional Recurrent Speech Enhancement NetworksabstractSpeech enhancement (SE) systems aim to improve the quality and intelligibility of degraded speech signals obtained from far-field microphones. Subjective evaluation of the intelligibility performance of these SE systems is uncommon. Instead, objective intelligibility measures (OIMs) are generally used to predict subjective performance increases. Many recent deep learning (DL) based SE systems, are expected to improve the intelligibility of degraded speech as measured by OIMs. However, validation of the ability of these OIMs to predict subjective intelligibility when enhancing a speech signal using DL-based systems is lacking. Therefore, in this study, we evaluate the predictive performance of five popular OIMs. We compare the metrics' predictions with subjective results. For this purpose, we recruited 50 human listeners, and subjectively tested both single channel and multi-channel Deep Complex Convolutional Recurrent Network (DCCRN) based speech enhancement systems. We found that none of the OIMs gave reliable predictions, and that all OIMs overestimated the intelligibility of ‘enhanced’ speech signals. Femke B. Gelderblom, Tron V. Tronstad, Torbjørn Svendsen, Tor André Myrvoll |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Using Modified Adult Speech as Data Augmentation for Child Speech RecognitionabstractData augmentation is a technique which enhances the size and quality of training data such that deep learning or machine learning models can achieve better performance. This paper proposes a novel way of applying data augmentation for child speech recognition in the low data resource scenario. Data augmentation is achieved by modifying existing adult speech signals. The procedure consists of two main parts, resampling, and time scaling. The experiment involves both speech from children aged from kindergarten to grade 10, and adults’ speech. We test the proposed method using both a TDNN-HMM and a GMM-HMM acoustic model. The results show that the proposed data augmentation scheme achieves a relative 7.95% reduction of WERs compared with 4.56% relative reduction when using a traditional bilinear frequency warping approach. Zijian Fan, Xinwei Cao, Giampiero Salvi, Torbjørn Svendsen |
ICASSP | 4 |
| 2023 | An Analysis of Goodness of Pronunciation for Child Speech
Xinwei Cao, Zijian Fan, Torbjørn Svendsen, Giampiero Salvi |
INTERSPEECH | 3 |
| 2023 | Perceptual and Task-Oriented Assessment of a Semantic Metric for ASR Evaluation
Janine Rugayan, Giampiero Salvi, Torbjørn Svendsen |
INTERSPEECH | 3 |
| 2022 | wav2vec2-based Speech Rating System for Children with Speech Sound DisorderabstractThe computational resources were provided by Aalto ScienceIT. This work was supported by NordForsk through the funding to Technology-enhanced foreign and second-language learning of Nordic languages, project number 103893. Yaroslav Getman, Ragheb Al-Ghezi, Katja Voskoboinik, Tamás Grósz, Mikko Kurimo, Giampiero Salvi, Torbjørn Svendsen, Sofia Strömbergsson |
INTERSPEECH | 7 |
| 2022 | Semantically Meaningful Metrics for Norwegian ASR SystemsabstractEvaluation metrics are important for quanitfying the performance of Automatic Speech Recognition (ASR) systems. However, the widely used word error rate (WER) captures errors at the word-level only and weighs each error equally, which makes it insufficient to discern ASR system performance for downstream tasks such as Natural Language Understanding (NLU) or information retrieval. We explore in this paper a more robust and discriminative evaluation metric for Norwegian ASR systems through the use of semantic information modeled by a transformer-based language model. We propose Aligned Semantic Distance (ASD) which employs dynamic programming to quantify the similarity between the reference and hypothesis text. First, embedding vectors are generated using the NorBERT model. Afterwards, the minimum global distance of the optimal alignment between these vectors is obtained and normalized by the sequence length of the reference embedding vector. In addition, we present results using Semantic Distance (SemDist), and compare them with ASD. Results show that for the same WER, ASD and SemDist values can vary significantly, thus, exemplifying that not all recognition errors can be considered equally important. We investigate the resulting data, and present examples which demonstrate the nuances of both metrics in evaluating various transcription errors. Janine Rugayan, Torbjørn Svendsen, Giampiero Salvi |
INTERSPEECH | 2 |
| 2022 | Acoustic-to-Articulatory Mapping With Joint Optimization of Deep Speech Enhancement and Articulatory Inversion ModelsabstractWe investigate the problem of speaker independent acoustic-to-articulatory inversion (AAI) in noisy conditions within the deep neural network (DNN) framework. In contrast with recent results in the literature, we argue that a DNN vector-to-vector regression front-end for speech enhancement (DNN-SE) can play a key role in AAI when used to enhance spectral features prior to AAI back-end processing. We experimented with single- and multi-task training strategies for the DNN-SE block finding the latter to be beneficial to AAI. Furthermore, we show that coupling DNN-SE producing enhanced speech features with an AAI trained on clean speech outperforms a multi-condition AAI (AAI-MC) when tested on noisy speech. We observe a 15% relative improvement in the Pearson's correlation coefficient (PCC) between our system and AAI-MC at 0dB signal-to-noise ratio on the Haskins corpus. Our approach also compares favourably against using a conventional DSP approach to speech enhancement (MMSE with IMCRA) in the front-end. Finally, we demonstrate the utility of articulatory inversion in a downstream speech application. We report significant WER improvements on an automatic speech recognition task in mismatched conditions based on the Wall Street Journal corpus (WSJ) when leveraging articulatory information estimated by AAI-MC system over spectral-alone speech features. Abdolreza Sabzi Shahrebabaki, Giampiero Salvi, Torbjørn Svendsen, Sabato Marco Siniscalchi |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | A Two-Stage Deep Modeling Approach to Articulatory InversionabstractThis paper proposes a two-stage deep feed-forward neural network (DNN) to tackle the acoustic-to-articulatory inversion (AAI) problem. DNNs are a viable solution for the AAI task, but the temporal continuity of the estimated articulatory values has not been exploited properly when a DNN is employed. In this work, we propose to address the lack of any temporal constraints while enforcing a parameter-parsimonious solution by deploying a two-stage solution based only on DNNs: (i) Articulatory trajectories are estimated in a first stage using DNN, and (ii) a temporal window of the estimated trajectories is used in a follow-up DNN stage as a refinement. The first stage estimation could be thought of as an auxiliary additional information that poses some constraints on the inversion process. Experimental evidence demonstrates an average error reduction of 7.51% in terms of RMSE compared to the baseline, and an improvement of 2.39% with respect to Pearson correlation is also attained. Finally, we should point out that AAI is still a highly challenging problem, mainly due to the non-linearity of the acoustic-to-articulatory and one-to-many mapping. It is thus promising that a significant improvement was attained with our simple yet elegant solution. Abdolreza Sabzi Shahrebabaki, Negar Olfati, Ali Shariq Imran, Magne Hallstein Johnsen, Sabato Marco Siniscalchi, Torbjørn Svendsen |
ICASSP | 6 |
| 2021 | Raw Speech-to-Articulatory Inversion by Temporal Filtering and DecimationabstractWe propose a novel sequence-to-sequence acoustic-to-articulatory inversion (AAI) neural architecture in the temporal waveform domain. In contrast to traditional AAI approaches that leverage hand-crafted short-time spectral features obtained from the windowed signal, such as LSFs, or MFCCs, our solution directly process the input speech signal in the time domain, avoiding any intermediate signal transformation, using a cascade of 1D convolutional filters in a deep model. The time-rate synchronization between raw speech signal and the articulatory signal is obtained through a decimation process that acts upon each convolution step. Decimation in time thus avoids degradation phenomena observed in the conventional AAI procedure, caused by the need of framing the speech signal to produce a feature sequence that perfectly matches the articulatory data rate. Experimental evidence on the “Haskins Production Rate Comparison” corpus demonstrates the effectiveness of the proposed solution, which outperforms a conventional state-of-the-art AAI system leveraging MFCCs with an 20% relative improvement in terms of Pearson correlation coefficient (PCC) in mismatched speaking rate conditions. Finally, the proposed approach attains the same accuracy as the conventional AAI solution in the typical matched speaking rate condition Abdolreza Sabzi Shahrebabaki, Sabato Marco Siniscalchi, Torbjørn Svendsen |
Interspeech | 3 |
| 2021 | A DNN Based Speech Enhancement Approach to Noise Robust Acoustic-to-Articulatory InversionabstractIn this work, we investigate the problem of speaker independent acoustic-to-articulatory inversion (AAI) in noisy condition within the deep neural network (DNN) framework. We claim that DNN vector-to-vector regression for speech enhancement (DNN-SE) can play a key role in AAI when used in a front-end stage to enhance speech features before AAI backend processing. Our claim contrasts recent literature reporting a drop in AAI accuracy on MMSE enhanced data and thereby sheds some light on the opportunities offered by DNN-SE in robust speech applications. We have also tested single- and multitask training strategies of the DNN-SE block and experimentally found the latter to be beneficial to AAI. Moreover, DNN-SE coupled with an AAI deep system tested on enhanced speech can outperform a multi-condition AAI deep system tested on noisy speech. We assess our approach on the Haskins corpus using the Pearson's correlation coefficient (PCC). A 15% relative PCC improvement is observed over a multi-condition AAI system at 0dB signal-to-noise ratio (SNR). Our approach also compares favorably against using a conventional DSP approach, namely MMSE with IMCRA, in the front-end stage. Abdolreza Sabzi Shahrebabaki, Sabato Marco Siniscalchi, Giampiero Salvi, Torbjørn Svendsen |
ISCAS | 4 |
| 2020 | Transfer Learning of Articulatory Information Through Phone InformationabstractArticulatory information has been argued to be useful for several speech tasks.However, in most practical scenarios this information is not readily available.We propose a novel transfer learning framework to obtain reliable articulatory information in such cases.We demonstrate its reliability both in terms of estimating parameters of speech production and its ability to enhance the accuracy of an end-to-end phone recognizer.Articulatory information is estimated from speaker independent phonemic features, using a small speech corpus, with electromagnetic articulography (EMA) measurements.Next, we employ a teacher-student model to learn estimation of articulatory features from acoustic features for the targeted phone recognition task.Phone recognition experiments, demonstrate that the proposed transfer learning approach outperforms the baseline transfer learning system acquired directly from an acousticto-articulatory (AAI) model.The articulatory features estimated by the proposed method, in conjunction with acoustic features, improved the phone error rate (PER) by 6.7% and 6% on the TIMIT core test and development sets, respectively, compared to standalone static acoustic features.Interestingly, this improvement is slightly higher than what is obtained by static+dynamic acoustic features, but with a significantly less.Adding articulatory features on top of static+dynamic acoustic features yields a small but positive PER improvement. Abdolreza Sabzi Shahrebabaki, Negar Olfati, Sabato Marco Siniscalchi, Giampiero Salvi, Torbjørn Svendsen |
INTERSPEECH | 5 |
| 2020 | Sequence-to-Sequence Articulatory Inversion Through Time Convolution of Sub-Band Frequency SignalsabstractWe propose a new acoustic-to-articulatory inversion (AAI) sequence-to-sequence neural architecture, where spectral sub-bands are independently processed in time by 1-dimensional (1-D) convolutional filters of different sizes. The learned features maps are then combined and processed by a recurrent block with bi-directional long short-term memory (BLSTM) gates for preserving the smoothly varying nature of the articulatory trajectories. Our experimental evidence shows that, on a speaker dependent AAI task, in spite of the reduced number of parameters, our model demonstrates better root mean squared error (RMSE) and Pearson's correlation coefficient (PCC) than a both a BLSTM model and an FC-BLSTM model where the first stages are fully connected layers. In particular, the average RMSE goes from 1.401 when feeding the filterbank features directly into the BLSTM, to 1.328 with the FC-BLSTM model, and to 1.216 with the proposed method. Similarly, the average PCC increases from 0.859 to 0.877, and 0.895, respectively. On a speaker independent AAI task, we show that our convolutional features outperform the original filterbank features, and can be combined with phonetic features bringing independent information to the solution of the problem. To the best of the authors' knowledge, we report the best results on the given task and data. Abdolreza Sabzi Shahrebabaki, Sabato Marco Siniscalchi, Giampiero Salvi, Torbjørn Svendsen |
INTERSPEECH | 4 |
| 2019 | Text-Independent Speaker ID Employing 2D-CNN for Automatic Video Lecture Categorization in a MOOC SettingabstractA new form of distance and blended education has hit the market in recent years with the advent of massive open online courses (MOOCs) which have brought many opportunities to the educational sector. Consequently, the availability of learning content to vast demographics of people and across locations has opened up a plethora of possibilities for everyone to gain new knowledge through MOOCs. This poses an immense issue to the content providers as the amount of manual effort required to structure properly and to organize the content automatically for millions of video lectures daily become incredibly challenging. This paper, therefore, addresses this issue as a small part of our proposed personalized content management system by exploiting the voice pattern of the lecturer for identification and for classifying video lectures to the right speaker category. The use of Mel frequency Cepstral coefficients (MFCC) as 2D input features maps to 2D-CNN has shown promising results in contrast to machine learning and deep learning classifiers - making text-independent speaker identification plausible in MOOC setting for automatic video lecture categorization. It will not only help categorize educational videos efficiently for easy search and retrieval but will also promote effective utilization of micro-lectures and multimedia video learning objects (MLO). Ali Shariq Imran, Zenun Kastrati, Torbjørn Svendsen, Arianit Kurti |
ICTAI | 3 |
| 2019 | A Phonetic-Level Analysis of Different Input Features for Articulatory InversionabstractThe challenge of articulatory inversion is to determine the tem- poral movement of the articulators from the speech waveform, or from acoustic-phonetic knowledge, e.g. derived from infor- mation about the linguistic content of the utterance. The actual position of the articulators is typically obtained from measured data, in our case position measurements obtained using EMA (Electromagnetic articulography). In this paper, we investigate the impact on articulatory inversion problem by using features derived from the acoustic waveform relative to using linguis- tic features related to the time aligned phone sequence of the utterance. Filterbank energies (FBE) are used as acoustic fea- tures, while phoneme identities and (binary) phonetic attributes are used as linguistic features. Experiments are performed on a speech corpus with synchronously recorded EMA measure- ments and employing a bidirectional long short-term memory (BLSTM) that estimates the articulators’ position. Acoustic FBE features performed better for vowel sounds. Phonetic fea- tures attained better results for nasal and fricative sounds except for /h/. Further improvements were obtained by combining FBE and linguistic features, which led to an average relative RMSE reduction of 9.8%, and a 3% relative improvement of the Pearson correlation coefficient. Abdolreza Sabzi Shahrebabaki, Negar Olfati, Ali Shariq Imran, Sabato Marco Siniscalchi, Torbjørn Svendsen |
INTERSPEECH | 5 |
| 2014 | An artificial neural network approach to automatic speech processing
Sabato Marco Siniscalchi, Torbjørn Svendsen, Chin-Hui Lee 0001 |
Neurocomputing | 2 |
| 2013 | Synthetic speaker models using VTLN to improve the performance of children in mismatched speaker conditions for ASR
Rama Sanand Doddipatla, Torbjørn Svendsen |
INTERSPEECH | 2 |
| 2013 | Universal attribute characterization of spoken languages for automatic spoken language recognition
Sabato Marco Siniscalchi, Jeremy Reed, Torbjørn Svendsen, Chin-Hui Lee 0001 |
Comput. Speech Lang. | 3 |
| 2013 | A Bottom-Up Modular Search Approach to Large Vocabulary Continuous Speech RecognitionabstractA novel bottom-up decoding framework for large vocabulary continuous speech recognition (LVCSR) with a modular search strategy is presented. Weighted finite state machines (WFSMs) are utilized to accomplish stage-by-stage acoustic-to-linguistic mappings from low-level speech attributes to high-level linguistic units in a bottom-up manner. Probabilistic attribute and phone lattices are used as intermediate vehicles to facilitate knowledge integration at different levels of the speech knowledge hierarchy. The final decoded sentence is obtained by performing lexical access and applying syntactical constraints. Two key factors are critical to warrant a high recognition accuracy, namely: (i) generation of high-precision sets of competing hypotheses at every intermediate stage; and (ii) low-error pruning of unlikely theories to reduce input lattice sizes while maintaining high-quality hypotheses for the next layers of knowledge integration. The decoupled nature of the proposed techniques allows us to obtain recognition results at all stages, including attribute, phone and word levels, and enables an integration of various knowledge sources not easily done in the state-of-the-art hidden Markov model (HMM) systems based on top-down knowledge integration. Evaluation on the Nov92 test set of the 5000-word, Wall Street Journal task demonstrates that high-accuracy attribute and phone classification can be attained. As for word recognition, the proposed WFSM-based framework achieves encouraging word error rates. Finally, by combining attribute scores with the conventional HMM likelihood scores and re-ordering the N-best lists obtained from the word lattices generated with the proposed WFSM system, the word error rate (WER) can be further reduced. Sabato Marco Siniscalchi, Torbjørn Svendsen, Chin-Hui Lee 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Experiments on Cross-Language Attribute Detection and Phone Recognition With Minimal Target-Specific Training DataabstractA state-of-the-art automatic speech recognition (ASR) system can often achieve high accuracy for most spoken languages of interest if a large amount of speech material can be collected and used to train a set of language-specific acoustic phone models. However, designing good ASR systems with little or no language-specific speech data for resource-limited languages is still a challenging research topic. As a consequence, there has been an increasing interest in exploring knowledge sharing among a large number of languages so that a universal set of acoustic phone units can be defined to work for multiple or even for all languages. This work aims at demonstrating that a recently proposed automatic speech attribute transcription framework can play a key role in designing language-universal acoustic models by sharing speech units among all target languages at the acoustic phonetic attribute level. The language-universal acoustic models are evaluated through phone recognition. It will be shown that good cross-language attribute detection and continuous phone recognition performance can be accomplished for “unseen” languages using minimal training data from the target languages to be recognized. Furthermore, a phone-based background model (PBM) approach will be presented to improve attribute detection accuracies. Sabato Marco Siniscalchi, Dau-Cheng Lyu, Torbjørn Svendsen, Chin-Hui Lee 0001 |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | Multi-site heterogeneous system fusions for the Albayzin 2010 Language Recognition EvaluationabstractBest language recognition performance is commonly obtained by fusing the scores of several heterogeneous systems. Regardless the fusion approach, it is assumed that different systems may contribute complementary information, either because they are developed on different datasets, or because they use different features or different modeling approaches. Most authors apply fusion as a final resource for improving performance based on an existing set of systems. Though relative performance gains decrease as larger sets of systems are considered, best performance is usually attained by fusing all the available systems, which may lead to high computational costs. In this paper, we aim to discover which technologies combine the best through fusion and to analyse the factors (data, features, modeling methodologies, etc.) that may explain such a good performance. Results are presented and discussed for a number of systems provided by the participating sites and the organizing team of the Albayzin 2010 Language Recognition Evaluation. We hope the conclusions of this work help research groups make better decisions in developing language recognition technology. Luis Javier Rodríguez-Fuentes, Mikel Peñagarikano, Amparo Varona, Mireia Díez, Germán Bordel, David Martínez González, Jesús Villalba 0001, Antonio Miguel, Alfonso Ortega Giménez, Eduardo Lleida, Alberto Abad, Oscar Koller, Isabel Trancoso, Paula Lopez-Otero, Laura Docío Fernández, Carmen García-Mateo, Rahim Saeidi, Mehdi Soufifar, Tomi Kinnunen, Torbjørn Svendsen, Pasi Fränti |
ASRU | 20 |
| 2011 | Pronunciation variation modeling of non-native proper names by discriminative tree searchabstractIn this paper, the task of selecting the optimal subset of pronunciation variants from a set of automatically generated candidates is recast as a tree search problem. In this approach, the optimal recognition lexicon corresponds with the optimal path through a search tree. We define a discriminative evaluation function to guide the search algorithm, which is based on estimates of the number of recognition errors before and after a lexicon change. The error rate for a given lexicon is estimated using the Minimum Classification Error framework. Selecting pronunciation candidates by means of this search algorithm clearly outperforms a baseline selection method, resulting in a reduction of both the error rate and the required number of variants in the recognition lexicon. Line Adde, Torbjørn Svendsen |
ICASSP | 2 |
| 2011 | A Bottom-Up Stepwise Knowledge-Integration Approach to Large Vocabulary Continuous Speech Recognition Using Weighted Finite State MachinesabstractA bottom-up, stepwise, knowledge integration framework is proposed to realize detection-based, large vocabulary continuous speech recognition (LVCSR) with a weighted finite state machine (WFSM). The WFSM framework offers a flexible architecture for different types of knowledge network compositions, each of them can be built and optimized independently. Speech attribute detectors are used as an intermediate block to obtain phoneme posterior probabilities over which a phoneme recognition network is designed. Lexical access and syntax knowledge integration over this phoneme network are then performed to deliver the decoded sentences. Experimental evidence illustrates that the proposed system outperforms several hybrid HMM/ANN systems with different configurations on the Wall Street Journal task while it is competitive with conventional LVCSR technology. Sabato Marco Siniscalchi, Torbjørn Svendsen, Chin-Hui Lee 0001 |
INTERSPEECH | 2 |
| 2011 | Frequency-Warped and Stabilized Time-Varying Cepstral Coefficients
Trond Skogstad, Torbjørn Svendsen |
INTERSPEECH | 2 |
| 2011 | iVector Approach to Phonotactic Language RecognitionabstractThis paper addresses a novel technique for representation and processing of n-gram counts in phonotactic language recognition (LRE): subspace multinomial modelling represents the vectors of n-gram counts by low dimensional vectors of coordinates in total variability subspace, called iVector. Two techniques for iVector scoring are tested: support vector machines (SVM), and logistic regression (LR). Using standard NIST LRE 2009 task as our evaluation set, the latter scoring approach was shown to outperform phonotactic LRE system based on direct SVM classification of n-gram count vectors. The proposed iVector paradigm also shows comparable results to previously proposed PCA-based phonotactic feature extraction. Index Terms: language recognition, subspace modeling, multinomial distribution. Mehdi Soufifar, Marcel Kockmann, Lukás Burget, Oldrich Plchot, Ondrej Glembek, Torbjørn Svendsen |
INTERSPEECH | 6 |
| 2010 | Experimental studies on continuous speech recognition using neural architectures with "adaptive" hidden activation functionsabstractThe choice of hidden non-linearity in a feed-forward multi-layer perceptron (MLP) architecture is crucial to obtain good generalization capability and better performance. Nonetheless, little attention has been paid to this aspect in the ASR field. In this work, we present some initial, yet promising, studies toward improving ASR performance by adopting hidden activation functions that can be automatically learned from the data and change shape during training. This adaptive capability is achieved through the use of orthonormal Hermite polynomials. The “adaptive” MLP is used in two neural architectures that generate phone posterior estimates, namely, a standalone configuration and a hierarchical structure. The posteriors are input to a hybrid phone recognition system with good results on the TIMIT corpus. A scheme for optimizing the contributions of high-accuracy neural architectures is also investigated, resulting in a relative improvement of ~9.0% over a non-optimized combination. Finally, initial experiments on the WSJ Nov92 task show that the proposed technique scales well up to large vocabulary continuous speech recognition (LVCSR) tasks. Sabato Marco Siniscalchi, Torbjørn Svendsen, Filippo Sorbello, Chin-Hui Lee 0001 |
ICASSP | 2 |
| 2010 | A minimum classification error approach to pronunciation variation modeling of non-native proper namesabstractIn automatic recognition of non-native proper names, it is critical to be able to handle a variety of different pronunciations. Traditionally, this has been solved by including alternative pronunciation variants in the recognition lexicon at the risk of introducing unwanted confusion between different name entries. In this paper we propose a pronunciation variant selection criterion that aims to avoid this risk by basing its decisions on scores which are calculated according to the minimum classification error (MCE) framework. By comparing the error rate before and after a lexicon change, the selection criterion chooses only the candidates that actually decrease the error rate. Selecting pronunciation candidates in this manner substantially reduces both the error rate and the required number of variants per name compared to a probability-based baseline selection method. Line Adde, Bert Réveil, Jean-Pierre Martens, Torbjørn Svendsen |
INTERSPEECH | 4 |
| 2010 | Exploiting context-dependency and acoustic resolution of universal speech attribute models in spoken language recognitionabstractThis paper expands a previously proposed universal acoustic characterization approach to spoken language identification (LID) by studying different ways of modeling attributes to improve language recognition. The motivation is to describe any spoken language with a common set of fundamental units. Thus, a spoken utterance is first tokenized into a sequence of universal attributes. Then a vector space modeling approach delivers the final LID decision. Context-dependent attribute models are now used to better capture spectral and temporal characteristics. Also, an approach to expand the set of attributes to increase the acoustic resolution is studied. Our experiments show that the tokenization accuracy positively affects LID results by producing a 2.8% absolute improvement over our previous 30-second NIST 2003 performance. This result also compares favorably with the best results on the same task known by the authors when the tokenizers are trained on language-dependent OGI-TS data. Sabato Marco Siniscalchi, Jeremy Reed, Torbjørn Svendsen, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2010 | Intra-frame variability as a predictor of frame classifiability
Trond Skogstad, Torbjørn Svendsen |
INTERSPEECH | 2 |
| 2010 | NameDat: A Database of English Proper Names Spoken by Native Norwegians
Line Adde, Torbjørn Svendsen |
LREC | 2 |
| 2010 | Spontal-N: A Corpus of Interactional Spoken Norwegian
Rein Ove Sikveland, Anton Öttl, Ingunn Amdal, Mirjam Ernestus, Torbjørn Svendsen, Jens Edlund |
LREC | 5 |
| 2010 | On the use of discriminative and non-discriminative pronunciation priors in pronunciation variation modeling of non-native proper namesabstractThe large amount of variation present in native speakers' pronunciation of non-native proper names is a big challenge for most automatic speech recognition systems today. The recognizer's ability to handle a variety of different pronunciations is therefore critical to achieve an acceptable recognition performance for this task. This problem has traditionally been solved by including alternative pronunciation variants in the recognition lexicon. To reduce the unwanted confusion that this might introduce between different lexicon entries, several studies have tried incorporating pronunciation prior probabilities into the recognition process. In this paper, we compare three different approaches for training these pronunciation prior probabilities based on: probabilities learned during the pronunciation variant generation and the discriminative frameworks Maximum Entropy and Minimum Classification Error. The recognition results obtained using these prior probabilities as a variant selection criterion are evaluated and a comparative error analysis is performed. Line Adde, Torbjørn Svendsen |
SLT | 2 |
| 2009 | Lexicon adaptation for subword speech recognitionabstractIn this paper we present two approaches to adapt a syllable-based recognition lexicon in an automatic speech recognition (ASR) setting. The motivation is to evaluate whether adaptation techniques commonly used on a word level can also be employed on a subword level. The first method predicts syllable variations, taking into account sub-syllabic phone cluster variations, and subsequently adapts the syllable lexicon. The second approach adds syllable bigrams to the lexicon to cope with acoustic confusability of subword units and syllable-inherent phone attachment ambiguities. We evaluate the methods on two German data sets, one consisting of planned and the other of spontaneous speech. Although the first method did not yield any improvement in the syllable error rate (SER), we could observe that the predicted confusions correlate with those observed in the test data. Bigram adaptation improved the SER by 1.3% and 0.8% absolute on the planned and spontaneous data sets, respectively. Timo Mertens, Daniel Schneider 0003, Arild Brandrud Næss, Torbjørn Svendsen |
ASRU | 4 |
| 2009 | A phonetic feature based lattice rescoring approach to LVCSRabstractLarge Vocabulary Continuous Speech Recognition (LVCSR) systems decode the input speech using diverse information sources, such as acoustic, lexical, and linguistic. Although most of the unreliable hypotheses are pruned during the recognition process, current state-of-the-art systems often make errors that are ldquounreasonablerdquo for human listeners. Several studies have shown that a proper integration of acoustic-phonetic information can be beneficial to reducing such errors. We have previously shown that high-accuracy phone recognition can be achieved if a bank of speech attribute detectors is used to compute a confidence score describing attribute activation levels that the current frame exhibits. In those experiments, the phone recognition system did not rely on the language model to follow their word sequence constraints, and the vocabulary was small. In this work, we extend our approach to LVCSR by introducing a second recognition step during which additional information not directly used during conventional log-likelihood based decoding is introduced. Experimental results show promising performance. Sabato Marco Siniscalchi, Torbjørn Svendsen, Chin-Hui Lee 0001 |
ICASSP | 2 |
| 2009 | Exploring universal attribute characterization of spoken languages for spoken language recognitionabstractWe propose a novel universal acoustic characterization approach to spoken language identification (LID), in which any spoken language is described with a common set of fundamental units defined “universally.” Specifically, manner and place of articulation form this unit inventory and are used to build a set of universal attribute models with data-driven techniques. Using the vector space modeling approaches to LID a spoken utterance is first decoded into a sequence of attributes. Then, a feature vector consisting of co-occurrence statistics of attribute units is created, and the final LID decision is implemented with a set of vector space language classifiers. Although the present study is just in its preliminary stage, promising results comparable to acoustically rich phone-based LID systems have already been obtained on the NIST 2003 LID task. The results provide clear insight for further performance improvements and encourage a continuing exploration of the proposed framework. Sabato Marco Siniscalchi, Jeremy Reed, Torbjørn Svendsen, Chin-Hui Lee 0001 |
INTERSPEECH | 3 |
| 2008 | Toward a detector-based universal phone recognizerabstractIn recent research, we have proposed a high-accuracy bottom-up detection-based paradigm for continuous phone speech recognition. The key component of our system was a bank of articulatory detectors each of which computes a score describing an activation level of the specified speech phonetic features that the current frame exhibits. In this work, we present our first attempt at designing a universal phone recognizer using the detection-based approach. We show that our technique is intrinsically language independent since reliable articulatory detectors can be designed for diverse languages, and robust detection can be performed across languages. Moreover, a universal set of detectors is designed by sharing the training material available for several diverse languages. We further demonstrate that our approach makes it possible to decode new target languages by neither retraining nor applying acoustic adaptation techniques. We report phone recognition performance that compares favorably with the best results known by the authors on the OGI Multi-language Telephone Speech corpus. Sabato Marco Siniscalchi, Torbjørn Svendsen, Chin-Hui Lee 0001 |
ICASSP | 2 |
| 2008 | A penalized logistic regression approach to detection based phone classificationabstractRecently, we have proposed a detection-based speech recognizer which has two main components: a bank of phonetic feature detectors implemented with hidden Markov models (HMMs), and an event merger. Each detector generates a score that pertains to some phonetic features, e.g. voicing. The merger combines all these scores to generate phone labels. The parameters of the detectors and the merger can be optimized either separately or jointly, and we showed that penalized logistic regression machine (PLRM) is a convenient tool for joint optimization. We validated our approach on a rescoring scheme. In this work, we tackle the phone classification problem and show that high level phone accuracy can be achieved without a direct modeling of the phones when PLRM is used. We also show that better results can be obtained by increasing the number of phonetic features, and that our method outperforms phone classifiers trained either by maximum likelihood estimation, or maximum mutual information Sabato Marco Siniscalchi, Torbjørn Svendsen, Chin-Hui Lee 0001 |
INTERSPEECH | 2 |
| 2008 | RUNDKAST: an Annotated Norwegian Broadcast News Speech Corpus
Ingunn Amdal, Ole Morten Strand, Jørn Almberg, Torbjørn Svendsen |
LREC | 4 |
| 2007 | Towards bottom-up continuous phone recognitionabstractWe present a novel approach to designing bottom-up automatic speech recognition (ASR) systems. The key component of the proposed approach is a bank of articulatory attribute detectors implemented using a set of feed-forward artificial neural networks (ANNs). Each detector computes a score describing an activation level of the specified speech attributes that the current frame exhibits. These cues are first combined by an event merger that provides some evidence about the presence of a higher level feature which is then verified by an evidence verifier to produce hypotheses at the phone or word level. We evaluate several configurations of our proposed system on a continuous phone recognition task. Experimental results on the TIMIT database show that the system achieves a phone error rate of 25% which is superior to results obtained with either hidden Markov model (HMM) or conditional random field (CRF) based recognizers. We believe the system's inherent flexibility and the ease of adding new detectors may provide further improvements. Sabato Marco Siniscalchi, Torbjørn Svendsen, Chin-Hui Lee 0001 |
ASRU | 2 |
| 2006 | FonDat1: A Speech Synthesis Corpus for Norwegian
Ingunn Amdal, Torbjørn Svendsen |
LREC | 2 |
| 2005 | Unit selection synthesis database development using utterance verificationabstractAccurate annotation of the unit inventory database is of vi-tal importance to the quality of unit selection text-to-speech synthesis. The time consuming manual work involved in database development limits the ability to produce new voices quickly and at low cost. Automatic annotation is therefore more and more in use. Misalignments due to mismatch between the predicted and pronounced unit sequence require manual correc-tion to achieve natural sounding synthesis. This paper proposes a new annotation assessment method using log likelihood ratio based utterance verification on the recorded database. The ut-terance verification is applied to detect utterances where there is a likely mismatch between the predicted pronunciation and what is actually spoken, or where an automated procedure for phonemic labelling misaligns the phone labels and the acoustic content. In a fully automated procedure, utterances failing the ver-ification test can be discarded. In semi-automatic procedures, the utterance verification can be applied to select utterances that need to be manually inspected, thereby reducing the manual ef-fort. Preliminary experiments are presented that show promis-ing figures for correct rejections. 1. Ingunn Amdal, Torbjørn Svendsen |
INTERSPEECH | 2 |
| 2005 | Comparing spectral distance measures for join cost optimization in concatenative speech synthesisabstractAbstract In concatenative synthesis the join cost function can be relatedtotheprobabilityofaperceiveddiscontinuityatthejoin. There-fore it is important that the distance measures in the cost func-tion correlate highly with human perceived discontinuities. Inthis paper the results of a listening test on joins in two Norwe-gian long vowels: /A:/ and /e:/, is presented. Five spectral dis-tancemeasuresandtheF0differencearecomparedaspredictorsof the human perceived discontinuities using Receiver Operat-ing Characteristic (ROC) curves. In addition, a linear join costfunction is optimized by means of stepwise linear regression. 1. Introduction Unit selection systems are considered state of the art in text tospeech synthesis, with capability of producing highly natural-sounding speech. Unit selection synthesis is based on concate-natingsmallunitsofspeech,selectedfromalargedatabasecon-taining multiple candidatesfor each unit. The search for the op-timal unit sequence is normally based on a combination of twocost functions: target cost and concatenation cost [1]. Targetcost measures how well a unit matches prosodic and phoneticfeatures of the target, while the concatenation cost is a measureof how well two neighboring units can be concatenated. Theoptimal unit sequence with respect to these two cost functionscan then be found by a Viterbi search for the lowest cost paththrough a lattice, where the database units are the nodes with anassociate target cost, and the concatenation cost is the cost ofthe path between two nodes.One problem with unit selection synthesis system is thelarge variability in quality, varying from almost perfect speechto very poor quality speech with many disturbing discontinu-ities. To improve quality we want the concatenation cost tohave a high correlation with human perception of discontinu-ities at concatenation points. The concatenation cost has to bemeasured with distance measures on physical measurable prop-erties such as spectral parameters, F0 (pitch) and power. Theprobability of perceived discontinuity at a concatenation pointcan be estimated as a composite function of the distance mea-sures. Defining the concatenation cost as a function propor-tional to the probability of perceived discontinuity then gives acost function Ingmund Bjrkan, Torbjørn Svendsen, Snorre Farner |
INTERSPEECH | 2 |
| 2005 | Distributed ASR using speech coder data for efficient feature vector representationabstractThis paper proposes an alternative approach to distributed speech recognition in scenarios where both reliable feature vectors and reconstruction of the speech signal are required. By transmitting the difference between speech coded information and the desired feature vectors, this system achieves both excellent quality speech reconstruction and ASR recognition performance. Experiments show that a transparent recognition rate is achieved with as little as 0.6 kbps of additional information supplementing the AMR speech coder operating at 4.75 kbps. The total rate is comparable to the the ETSI 202 211 extended front-end standard. Trond Skogstad, Torbjørn Svendsen |
INTERSPEECH | 2 |
| 2003 | Cross-lingual pronunciation modelling for indonesian speech recognitionabstractThe resources necessary to produce Automatic Speech Recognition systems for a new language are considerable, and for many languages these resources are not available. This emphasizes the need for the development of generic techniques which overcome this data shortage. Indonesian is one language which suffers from this problem and whose population and importance suggest it could benefit from speech enabled technology. Accordingly, we investigate using English acoustic models to recognize Indonesian speech. The mapping process, where the symbolic representation of the Source language acoustic models is equated to the Target language phonetic units, has typically been achieved using one to one mapping techniques. This mapping method does not allow for the incorporation of predictable allophonic variation in the lexicon. Accordingly, in this paper we present the use of cross-lingual pronunciation modelling to extract context dependant mapping rules, which are subsequently used to produce a more accurate cross lingual lexicon. Terrence Martin, Torbjørn Svendsen, Sridha Sridharan |
INTERSPEECH | 2 |
| 2003 | Multilingual phone clustering for recognition of spontaneous indonesian speech utilising pronunciation modelling techniquesabstractIn this paper, a multilingual acoustic model set derived from English, Hindi, and Spanish is utilised to recognise speech in Indonesian. In order to achieve this task we incorporate a two tiered approach to perform the cross-lingual porting of the multilingual models to a new language. In the first stage, we use an entropy based decision tree to merge similar phones from different languages intoclustersto forma newmultilingual model set. In the second stage, we propose the use of a cross-lingual pronunciation modelling technique to perform the mapping from the multilingual models to the Indonesian phone set. A set of mapping rules are derived from this process and are employed to convert the original Indonesian lexicon into a pronunciation lexicon in terms of the multilingual model set. Preliminary experimental results show that, compared to the common knowledge based approach, both of these techniques reduce the word error rate in a spontaneous speech recognition task. Eddie Wong, Terrence Martin, Torbjørn Svendsen, Sridha Sridharan |
INTERSPEECH | 3 |
| 2002 | Evaluation of Pronunciation Variants in the ASR Lexicon for Different Speaking Styles
Ingunn Amdal, Torbjørn Svendsen |
LREC | 2 |
| 2001 | Fast adaptation using constrained affine transformations with hierarchical priorsabstractIn this paper we present an approach to transformation based model adaptation that combines a fast, closed form solution to the MAP estimation of our transforms with robust priors. The robust priors are found using the technique of hierarchical priors, and a closed form solution is achieved by choosing diagonally constrained affine transformations and a suitable family of prior distributions for these transformations. We show that the method gives results comparable to other algorithms, but with significantly reduced computational complexity and memory demands. Experiments are conducted on the SI Recognition Outlier task from the Wall Street Journal corpus, where speaker independent models have to be adapted to handle speech from non-native speakers. Tor André Myrvoll, Kuldip K. Paliwal, Torbjørn Svendsen |
INTERSPEECH | 3 |
| 2000 | ASR-based subtitling of live TV-programs for the hearing impairedabstractA system for on-line generation of closed captions (subtitles) for broadcast of live TV-programs is described. During broadcast, a commentator formulates a possibly condensed, but semantically correct version of the original speech. These compressed phrases are recognized by a continuous speech recognizer, and the resulting captions are fed into the teletext system. This application will provide the hearing impaired with an option to read captions for live broadcast programs, i.e., when off-line captioning is not feasible. The main advantage in using a speech recognizer rather than a stenography-based system (e.g., Velotype) is the relaxed requirements for commentator training. Also, the amount of text generated by a system based on stenography tends to be large, thus making it harder to read. 1. Trym Holter, Erik Harborg, Magne Hallstein Johnsen, Torbjørn Svendsen |
INTERSPEECH | 4 |
| 2000 | Stochastic modeling of semantic content for use IN a spoken dialogue system
Magne Hallstein Johnsen, Trym Holter, Torbjørn Svendsen, Erik Harborg |
INTERSPEECH | 3 |
| 2000 | TABOR - a norwegian spoken dialogue system for bus travel information
Magne Hallstein Johnsen, Torbjørn Svendsen, Tore Amble, Trym Holter, Erik Harborg |
INTERSPEECH | 2 |
| 1999 | On-line captioning of TV-programs for the hearing impaired
Erik Harborg, Trym Holter, Magne Hallstein Johnsen, Torbjørn Svendsen |
EUROSPEECH | 4 |
| 1999 | Maximum likelihood modelling of pronunciation variation
Trym Holter, Torbjørn Svendsen |
Speech Commun. | 2 |
| 1997 | Incorporating linguistic knowledge and automatic baseform generation in acoustic subword unit based speech recognitionabstractA major challenge in speech recognition based on acoustic subword units is creating a lexicon which is robust to inter- and intra-speaker variations. In this paper we present two different approaches for incorporating simple word-level linguistic knowledge into the labelling step of the training procedure. The proposed systems also utilise a scheme for combined optimisation of baseforms and subword models. For the TI46 database, these methods are shown to greatly improve the performance compared to an acoustic subword based speech recogniser employing unsupervised labelling, and they are found to perform as well as systems utilising whole-word models and context independent phoneme models. 1. INTRODUCTION Traditionally, automatic speech recognisers employ phone-like units based upon a linguistic description of the language. On the other hand, the analysis of the actual speech signal is acoustically based. The resulting system is neither phonetically nor acoustically consistent, but is... Trym Holter, Torbjørn Svendsen |
EUROSPEECH | 2 |
| 1995 | Optimizing baseforms for HMM-based speech recognition
Torbjørn Svendsen, Frank K. Soong, Heiko Purnhagen |
EUROSPEECH | 1 |
| 1994 | Segmental quantization of speech spectral informationabstractThe majority of current speech coding algorithms for medium-to-low bit rates transmit two information components, a short-time spectrum estimate and an excitation signal. Even though advanced intraframe quantization schemes have been proposed, the spectral information still consumes large proportion of the available bit rate. For many speech sounds, the speech spectrum is relatively smooth for time intervals much longer than the sampling rate of the spectrum estimates. Thus, compression can be obtained by identifying smoothly varying segments of the speech spectrum and only transmitting the spectral information once for each segment. The segment spectral information is then an approximation to the true spectrum, but if the segmentation criterion is properly chosen, the induced distortion can be controlled to be within the acceptable 1 dB mean spectral distortion limit. In the present paper the author shows that segment quantization can be applied to reduce the required bit rate for the spectral information by a factor of approximately two without compromising the total spectral distortion.> Torbjørn Svendsen |
ICASSP (1) | 1 |
| 1993 | A time-frequency segmental neural network for phoneme recognition
Anjan Basu, Torbjørn Svendsen |
ICASSP (1) | 2 |
| 1993 | Cost232: speech recognition over the telephone line
Andrea Paoloni, Torbjørn Svendsen, Bernhard Kaspar, Denis Johnston, Gunnar Hult |
EUROSPEECH | 2 |
| 1993 | Efficient quantization of speech spectral information
Torbjørn Svendsen |
EUROSPEECH | 1 |
| 1991 | ANN-based speech recognition using a preprocessor for non-linear time compression
P. O. Husoy, Torbjørn Svendsen |
EUROSPEECH | 2 |
| 1990 | Automatic alignment of phonemic labels with continuous speech
Torbjørn Svendsen, Knut Kvale |
ICSLP | 1 |
| 1989 | An improved sub-word based speech recognizerabstractThe authors describe a system for speaker-dependent speech recognition based on acoustic subword units. Several strategies for automatic generation of an acoustic lexicon are outlined. Preliminary tests have been performed on a small vocabulary. In these tests, the proposed system showed results comparable to those of whole-word-based systems.> Torbjørn Svendsen, Kuldip K. Paliwal, Erik Harborg, P. O. Husoy |
ICASSP | 1 |
| 1987 | On the automatic segmentation of speech signalsabstractFor large vocabulary and continuous speech recognition, the sub-word-unit-based approach is a viable alternative to the whole-word-unit-based approach. For preparing a large inventory of subword units, an automatic segmentation is preferrable to manual segmentation as it substantially reduces the work associated with the generation of templates and gives more consistent results. In this paper we discuss some methods for automatically segmenting speech into phonetic units. Three different approaches are described, one based on template matching, one based on detecting the spectral changes that occur at the boundaries between phonetic units and one based on a constrained-clustering vector quantization approach. An evaluation of the performance of the automatic segmentation methods is given. Torbjørn Svendsen, Frank K. Soong |
ICASSP | 1 |
| 1986 | Multi-dimensional quantization applied to predictive coding of speechabstractThe properties of several multi-dimensional quantizers (VQ, tree and trellis coders) have been investigated. The multi-dimensional quantizers yield a superior distortion performance for direct quantization of stationary sources. Incorporated in a predictive speech coding scheme, NFC, the use of multi-dimensional quantizers improves speech quality noticably. A subjective test indicate that NFC with trellis coding is a promising candidate for speech coding at 16 kbit/s. Torbjørn Svendsen |
ICASSP | 1 |
| 1985 | A study of three coders (sub-band, RELP and MPE) for speech with additive white noiseabstractThe following three speech coders are implemented for a bitrate of 9.6 kbits/s 1) Sub-band coder, 2) Residual Excited Linear Predictive (RELP) coder, and 3) Multi-Pulse Excited linear predictive (MPE) coder. Performance of these coders is evaluated for speech corrupted by additive white noise. Evaluation of speech coders is done both subjectively and objectively. The MPE coder is found to give the best performance among the three coders. It is also shown that the MPE coder can be used for noisy speech with signal-to-noise ratio as low as -10 dB giving reasonably good quality speech provided 1) one does not use the error weighting filter and 2) one can use a better LP analysis algorithm which can estimate LP coefficients correctly from noisy speech. Kuldip K. Paliwal, Torbjørn Svendsen |
ICASSP | 2 |
| 1984 | Tree encoding of the LPC residualabstractIn speech coding quantization has traditionally been done on a sample by sample basis. According to rate distortion theory there is much to be gained by applying multidimensional schemes to the quantization at low bit rates. This paper presents the use of a tree-encoder for quantization of the LPC-residual. The perceptual speech quality for the straightforward encoding is however not satisfactory and a frequency weighted error criterion which greatly improves the perceptual speech quality is suggested. Torbjørn Svendsen |
ICASSP | 1 |