VLDB 2026 Research / reviewers in the wild / expert
Tom Bäckström
dblp:43/3827
· DBLP profile ↗
59ranked-venue papers
24as first author
8since 2021 · last 2025
0000-0002-5590-2349ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 50 · 20 first-author · 5 since 2021Artificial intelligence and machine learning · 37 · 11 first-author · 5 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Privacy in Speech TechnologyabstractSpeech technology for communication, accessing information, and services has rapidly improved in quality. It is convenient and appealing because speech is the primary mode of communication for humans. Such technology, however, also presents proven threats to privacy. Speech is a tool for communication, and thus, it will inherently contain private information. Importantly, it also contains a wealth of side information, including details about health, emotions, affiliations, and relationships, all of which are private. Exposing such private information can lead to serious threats such as price gouging, harassment, extortion, and stalking. This article is a tutorial on privacy issues related to speech technology, modeling their threats, approaches for protecting users’ privacy, measuring the performance of privacy-protecting methods, perception of privacy, as well as societal and legal consequences. In addition to a tutorial overview, it also presents lines for further development where improvements are most urgently needed. Tom Bäckström |
Proc. IEEE | 1 |
| 2024 | Privacy PORCUPINE: Anonymization of Speaker Attributes Using Occurrence Normalization for Space-Filling Vector QuantizationabstractSpeech signals contain a vast range of private information such as its text, speaker identity, emotions, and state of health. Privacy-preserving speech processing seeks to filter out any private information that is not needed for downstream tasks, for example with an information bottleneck, sufficiently tight that only the desired information can pass through. We however demonstrate that the occurrence frequency of codebook elements in bottlenecks using vector quantization have an uneven information rate, threatening privacy. We thus propose to use space-filling vector quantization (SFVQ) together with occurrence normalization, balancing the information rate and thus protecting privacy. Our experiments with speaker identification validate the proposed method. This approach thus provides a generic tool for quantizing information bottlenecks in any speech applications such that their privacy disclosure is predictable and quantifiable. Mohammad Hassan Vali, Tom Bäckström |
INTERSPEECH | 2 |
| 2023 | Stochastic Optimization of Vector Quantization Methods in Application to Speech and Image ProcessingabstractVector quantization (VQ) methods have been used in a wide range of applications for speech, image, and video data. While classic VQ methods often use expectation maximization, in this paper, we investigate the use of stochastic optimization employing our recently proposed noise substitution in vector quantization technique. We consider three variants of VQ including additive VQ, residual VQ, and product VQ, and evaluate their quality, complexity and bitrate in speech coding, image compression, approximate nearest neighbor search, and a selection of toy examples. Our experimental results demonstrate the trade-offs in accuracy, complexity, and bitrate such that using our open source implementations and complexity calculator, the best vector quantization method can be chosen for a particular problem. Mohammad Hassan Vali, Tom Bäckström |
ICASSP | 2 |
| 2023 | Optimizing the Performance of Text Classification Models by Improving the Isotropy of the Embeddings Using a Joint Loss Function
Joseph Attieh, Abraham Woubie, Vladimir Vlassov, Adrian Flanagan, Tom Bäckström |
ICDAR (5) | 5 |
| 2023 | Interpretable Latent Space Using Space-Filling Curves for Phonetic Analysis in Voice ConversionabstractVector quantized variational autoencoders (VQ-VAE) are well-known deep generative models, which map input data to a latent space that is used for data generation. Such latent spaces are unstructured and can thus be difficult to interpret. Some earlier approaches have introduced a structure to the latent space through supervised learning by defining data labels as latent variables. In contrast, we propose an unsupervised technique incorporating space-filling curves into vector quantization (VQ), which yields an arranged form of latent vectors such that adjacent elements in the VQ codebook refer to similar content. We applied this technique to the latent codebook vectors of a VQ-VAE, which encode the phonetic information of a speech signal in a voice conversion task. Our experiments show there is a clear arrangement in latent vectors representing speech phones, which clarifies what phone each latent vector corresponds to and facilitates other detailed interpretations of latent vectors. Mohammad Hassan Vali, Tom Bäckström |
INTERSPEECH | 2 |
| 2023 | The Internet of Sounds: Convergent Trends, Insights, and Future DirectionsabstractCurrent sound-based practices and systems developed in both academia and industry point to convergent research trends that bring together the field of Sound and Music Computing with that of the Internet of Things. This paper proposes a vision for the emerging field of the Internet of Sounds (IoS), which stems from such disciplines. The IoS relates to the network of Sound Things, i.e., devices capable of sensing, acquiring, processing, actuating, and exchanging data serving the purpose of communicating sound-related information. In the IoS paradigm, which merges under a unique umbrella the emerging fields of the Internet of Musical Things and the Internet of Audio Things, heterogeneous devices dedicated to musical and non-musical tasks can interact and cooperate with one another and with other things connected to the Internet to facilitate sound-based services and applications that are globally available to the users. We survey the state of the art in this space, discuss the technological and non-technological challenges ahead of us and propose a comprehensive research agenda for the field. Luca Turchet, Mathieu Lagrange, Cristina Rottondi, György Fazekas, Nils Peters, Jan Østergaard, Frederic Font, Tom Bäckström, Carlo Fischione |
IEEE Internet Things J. | 8 |
| 2021 | End-to-End Optimized Multi-Stage Vector Quantization of Spectral Envelopes for Speech and Audio CodingabstractSpectral envelope modeling is an instrumental part of speech and audio codecs, which can be used to enable efficient entropy coding of spectral components. Overall optimization of codecs, including envelope models, has however been difficult due to the complicated interactions between different modules of the codec. In this paper, we study an end-to-end optimization methodology to optimize all modules in a codec integrally with respect to each other while capturing all these complex interactions with a global loss function. For the quantization of the spectral envelope parameters with a fixed bitrate, we use multistage vector quantization which gives high quality, but yet has a computational complexity which can be realistically applied in embedded devices. The obtained results demonstrate benefits in terms of PESQ and PSNR in comparison to the 3GPP EVS, as well as our recently proposed PyAWNeS codecs. Mohammad Hassan Vali, Tom Bäckström |
Interspeech | 2 |
| 2021 | Cancellation of Local Competing Speaker with Near-Field Localization for Distributed ad-hoc Sensor NetworkabstractIn scenarios such as remote work, open offices and call centers, multiple people may simultaneously have independent spoken interactions with their devices in the same room. The speech of competing speakers will however be picked up by all microphones, both reducing the quality of audio and exposing speakers to breaches in privacy. We propose a cooperative cross-talk cancellation solution breaking the single active speaker assumption employed by most telecommunication systems. The proposed method applies source separation on the microphone signals of independent devices, to extract the dominant speaker in each device. It is realized using a localization estimator based on a deep neural network, followed by a time-frequency mask to separate the target speech from the interfering one at each time-frequency unit referring to its orientation. By experimental evaluation, we confirm that the proposed method effectively reduces crosstalk and exceeds the baseline expectation maximization method by 10 dB in terms of interference rejection. This performance makes the proposed method a viable solution for cross-talk cancellation in near-field conditions, thus protecting the privacy of external speakers in the same acoustic space. Pablo Pérez Zarazaga, Mariem Bouafif Mansali, Tom Bäckström, Zied Lachiri |
Interspeech | 3 |
| 2020 | Fundamental Frequency Model for Postfiltering at Low Bitrates in a Transform-Domain Speech and Audio CodecabstractS.2837-2841 Sneha Das, Tom Bäckström, Guillaume Fuchs |
INTERSPEECH | 2 |
| 2020 | Perception of Privacy Measured in the Crowd - Paired Comparison on the Effect of Background NoisesabstractDefence is held on 26.11.2021 12:00 – 15:00 Zoom: https://aalto.zoom.us/j/61255513284 Anna Leschanowsky, Sneha Das, Tom Bäckström, Pablo Pérez Zarazaga |
INTERSPEECH | 3 |
| 2019 | Overlap-add Windows with Maximum Energy Concentration for Speech and Audio ProcessingabstractProcessing of speech and audio signals with time-frequency representations require windowing methods which allow perfect reconstruction of the original signal and where processing artifacts have a predictable behavior. The most common approach for this purpose is overlap-add windowing, where signal segments are windowed before and after processing. Commonly used windows include the half-sine and a Kaiser-Bessel derived window. The latter is an approximation of the discrete prolate spherical sequence, and thus a maximum energy concentration window, adapted for overlap-add. We demonstrate that performance can be improved by including the overlap-add structure as a constraint in optimization of the maximum energy concentration criteria. The same approach can be used to find further special cases such as optimal low-overlap windows. Our experiments demonstrate that the proposed windows provide notable improvements in terms of reduction in side-lobe magnitude. Tom Bäckström |
ICASSP | 1 |
| 2019 | End-to-End Optimization of Source Models for Speech and Audio Coding Using a Machine Learning FrameworkabstractSpeech coding is the most commonly used application of speech processing. Accumulated layers of improvements have however made codecs so complex that optimization of individual modules becomes increasingly difficult. This work introduces machine learning methodology to speech and audio coding, such that we can optimize quality in terms of overall entropy. We can then use conventional quantization, coding and perceptual models without modification such that the codec adheres to conventional requirements on algorithmic complexity, latency and robustness to packet loss. Experiments demonstrate that end-to-end optimization of quantization accuracy of the spectral envelope can be used for a lossless reduction in bitrate of 0.4 kbits/s. Tom Bäckström |
INTERSPEECH | 1 |
| 2019 | Super-Wideband Spectral Envelope Modeling for Speech CodingabstractSignificant improvements in the quality of speech coders have been achieved by widening the coded frequency range from narrowband to wideband. However, existing speech coders still employ a limited band source-filter model extended by parametric coding of the higher band. In the present work, a superwideband source-filter model running at 32 kHz is considered and especially its spectral magnitude envelope modeling. To match super-wideband operating mode, we adapted and compared two methods; Linear Predictive Coding (LPC) and Distribution Quantization (DQ). LPC uses autoregressive modeling, while DQ quantifies the energy ratios between different parts of the spectrum. Parameters of both methods were quantized with a multi-stage vector quantization. Objective and subjective evaluations indicate that both methods used in a super-wideband source-filter coding scheme offer the same quality range, making them an attractive alternative to conventional speech coders that require additional bandwidth extension. Guillaume Fuchs, Chamran Ashour, Tom Bäckström |
INTERSPEECH | 3 |
| 2019 | Sound Privacy: A Conversational Speech Corpus for Quantifying the Experience of PrivacyabstractWith the growing popularity of social networks, cloud services and online applications, people are becoming concerned about the way companies store their data and the ways in which the data can be applied. Privacy with devices and services operated by the voice are of particular interest. To enable studies in privacy, this paper presents a database which quantifies the experience of privacy users have in spoken communication. We focus on the effect of the acoustic environment on that perception of privacy. Speech signals are recorded in scenarios simulating real-life situations, where the acoustic environment has an effect on the experience of privacy. The acoustic data is complemented with measures of the speakers’ experience of privacy, recorded using a questionnaire. The presented corpus enables studies in how acoustic environments affect peoples’ experience of privacy, which in turn, can be used to develop speech operated applications which are respectful of their right to privacy. Pablo Pérez Zarazaga, Sneha Das, Tom Bäckström, Vishnu Vidyadhara Raju Vegesna, Anil Kumar Vuppala |
INTERSPEECH | 3 |
| 2018 | GMM-Based Iterative Entropy Coding for Spectral Envelopes of Speech and AudioabstractSpectral envelope modelling is a central part of speech and audio codecs and is traditionally based on either vector quantization or scalar quantization followed by entropy coding. To bridge the coding performance of vector quantization with the low complexity of the scalar case, we propose an iterative approach for entropy coding the spectral envelope parameters. For each parameter, a univariate probability distribution is derived from a Gaussian mixture model of the joint distribution and the previously quantized parameters used as a-priori information. Parameters are then iteratively and individually scalar quantized and entropy coded. Unlike vector quantization, the complexity of proposed method does not increase exponentially with dimension and bitrate. Moreover, the coding resolution and dimension can be adaptively modified without retraining the model. Experimental results show that these important advantages do not impair coding efficiency compared to a state-of-art vector quantization scheme. Srikanth Korse, Guillaume Fuchs, Tom Bäckström |
ICASSP | 3 |
| 2018 | Dithered Quantization for Frequency-Domain Speech and Audio CodingabstractA common issue in coding speech and audio in the frequency domain, which appears with decreasing bitrate, is that quantization levels become increasingly sparse. With low accuracy, high-frequency components are typically quantized to zero, which leads to a muffled output signal and musical noise. Band-width extension and noise-filling methods attempt to treat the problem by inserting noise of similar energy as the original signal, at the cost of low signal to noise ratio. Dithering methods however provide an alternative approach, where both accuracy and energy are retained. We propose a hybrid coding approach where low-energy samples are quantized using dithering, instead of the conventional uniform quantizer. For dithering, we apply 1 bit quantization in a randomized sub-space. We further show that the output energy can be adjusted to the desired level using a scaling parameter. Objective measurements and listening tests demonstrate the advantages of the proposed methods. Tom Bäckström, Johannes Fischer 0002, Sneha Das |
INTERSPEECH | 1 |
| 2018 | Postfiltering with Complex Spectral Correlations for Speech and Audio CodingabstractState-of-the-art speech codecs achieve a good compromise between quality, bitrate and complexity. However, retaining performance outside the target bitrate range remains challenging. To improve performance, many codecs use pre- and post-filtering techniques to reduce the perceptual effect of quantization-noise. In this paper, we propose a postfiltering method to attenuate quantization noise which uses the complex spectral correlations of speech signals. Since conventional speech codecs cannot transmit information with temporal dependencies as transmission errors could result in severe error propagation, we model the correlation offline and employ them at the decoder, hence removing the need to transmit any side information. Objective evaluation indicates an average 4 dB improvement in the perceptual SNR of signals using the context-based post-filter, with respect to the noisy signal and an average 2 dB improvement relative to the conventional Wiener filter. These results are confirmed by an improvement of up to 30 MUSHRA points in a subjective listening test. Sneha Das, Tom Bäckström |
INTERSPEECH | 2 |
| 2018 | Postfiltering Using Log-Magnitude Spectrum for Speech and Audio CodingabstractAdvanced coding algorithms yield high quality signals with good coding efficiency within their target bit-rate ranges, but their performance suffer outside the target range. At lower bitrates, the degradation in performance is because the decoded signals are sparse, which gives a perceptually muffled and distorted characteristic to the signal. Standard codecs reduce such distortions by applying noise filling and post-filtering methods. In this paper, we propose a post-processing method based on modeling the inherent time-frequency correlation in the log-magnitude spectrum. The goal is to improve the perceptual SNR of the decoded signals and, to reduce the distortions caused by signal sparsity. Objective measures show an average improvement of 1.5 dB for input perceptual SNR in range 4 to 18 dB. The improvement is especially prominent in components which had been quantized to zero. Sneha Das, Tom Bäckström |
INTERSPEECH | 2 |
| 2018 | Fast Randomization for Distributed Low-Bitrate Coding of Speech and AudioabstractEfficient coding of speech and audio in a distributed system requires that quantization errors across nodes are uncorrelated. Yet, with conventional methods at low bitrates, quantization levels become increasingly sparse, which does not correspond to the distribution of the input signal and, importantly, also reduces coding efficiency in a distributed system. We have recently proposed a distributed speech and audio codec design, which applies quantization in a randomized domain such that quantization errors are randomly rotated in the output domain. Similar to dithering, this ensures that quantization errors across nodes are uncorrelated and coding efficiency is retained. In this paper, we improve this approach by proposing faster randomization methods, with a computational complexity of O(N log N). The presented experiments demonstrate that the proposed randomizations yield uncorrelated signals, that perceptual quality is competitive, and that the complexity of the proposed methods is feasible for practical applications. Tom Bäckström, Johannes Fischer 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | Estimation of the Probability Distribution of Spectral Fine Structure in the Speech SourceabstractThe efficiency of many speech processing methods rely on accurate modeling of the distribution of the signal spectrum and a majority of prior works suggest that the spectral components follow the Laplace distribution. To improve the probability distribution models based on our knowledge of speech source modeling, we argue that the model should in fact be a multiplicative mixture model, including terms for voiced and unvoiced utterances. While prior works have applied Gaussian mixture models, we demonstrate that a mixture of generalized Gaussian models more accurately follows the observations. The proposed estimation method is based on measuring the ratio of $L_p$-norms between spectral bands. Such ratios follow the Beta-distribution when the input signal is generalized Gaussian, whereby the estimated parameters can be used to determine the underlying parameters of the mixture of generalized Gaussian distributions. Tom Bäckström |
INTERSPEECH | 1 |
| 2017 | Quadratic Programming Approach to Glottal Inverse Filtering by Joint Norm-1 and Norm-2 OptimizationabstractThis study proposes an approach for glottal inverse filtering of acoustic speech signals using quadratic programming (QPR). The method aims to jointly model the effect of vocal tract and lip radiation with a single filter whose coefficients are optimized using QPR. This optimization is based on the principles of closed phase analysis, where the contribution of the glottal source is attenuated in optimizing the inverse model of the vocal tract. By expressing the optimization problem in terms of the output of a filter, we can apply physically motivated optimization such as flatness of the closed phase. The proposed method was objectively evaluated using a synthetic Liljencrants-Fant model based test set of sustained vowels, as well as a real speech test set where the glottal flow estimates' closed phases were compared in terms of their flatness. The results based on synthetic speech indicate that the proposed method is robust to changes in f0, and state-of-the-art quality results were obtained for high-pitched voices, when f0is in the range 330-450 Hz. The results based on real speech indicate that the proposed method produces glottal flow estimates that have flatter closed phases with less formant ripple in comparison to estimates computed with known reference methods. Manu Airaksinen, Tom Bäckström, Paavo Alku |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Automatic Glottal Inverse Filtering with Non-Negative Matrix Factorization
Manu Airaksinen, Lauri Juvela, Tom Bäckström, Paavo Alku |
INTERSPEECH | 3 |
| 2016 | Blind Recovery of Perceptual Models in Distributed Speech and Audio Coding
Tom Bäckström, Florin Ghido, Johannes Fischer 0002 |
INTERSPEECH | 1 |
| 2016 | Joint Enhancement and Coding of Speech by Incorporating Wiener Filtering in a CELP CodecabstractS.1730-1734 Johannes Fischer 0002, Tom Bäckström |
INTERSPEECH | 2 |
| 2016 | Entropy Coding of Spectral Envelopes for Speech and Audio Coding Using Distribution QuantizationabstractS.2543-2547 Srikanth Korse, Tobias Jähnel, Tom Bäckström |
INTERSPEECH | 3 |
| 2016 | Feature Extraction Using Power-Law Adjusted Linear Prediction With Application to Speaker Recognition Under Severe Vocal Effort MismatchabstractLinear prediction is one of the most established techniques in signal estimation, and it is widely utilized in speech signal processing. It has been long understood that the nerve firing rate of human auditory system can be approximated by power law non-linearity, and this has been the motivation behind using perceptual linear prediction in extracting acoustic features in a variety of speech processing applications. In this paper, we revisit the application of power law non-linearity in speech spectrum estimation by compressing/expanding power spectrum in autocorrelation-based linear prediction. The development of so-called LP- α is motivated by a desire to obtain spectral features that present less mismatch than conventionally used spectrum estimation methods when speech of normal loudness is compared to speech under vocal effort. The effectiveness of the proposed approach is demonstrated in a speaker recognition task conducted under severe vocal effort mismatch comparing shouted versus normal speech mode. Rahim Saeidi, Paavo Alku, Tom Bäckström |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Arithmetic coding of speech and audio spectra using tcx based on linear predictive spectral envelopesabstractUnified speech and audio codecs often use a frequency domain coding technique of the transform coded excitation (TCX) type. It is based on modeling the speech source with a linear predictor, spectral weighting by a perceptual model and entropy coding of the frequency components. While previous approaches have used neighbouring frequency components to form a probability model for the entropy coder of spectral components, we propose to use the magnitude of the linear predictor to estimate the variance of spectral components. Since the linear predictor is transmitted in any case, this method does not require any additional side info. Subjective measurements show that the proposed methods give a statistically significant improvement in perceptual quality when the bit-rate is held constant. Consequently, the proposed method has been adopted to the 3GPP Enhanced Voice Services speech coding standard. Tom Bäckström, Christian R. Helmrich |
ICASSP | 1 |
| 2015 | Finding line spectral frequencies using the fast fourier transformabstractMain-stream speech codecs are based on modelling the speech source by a linear predictor. An efficient domain for quantization and coding of this linear predictor is the line spectral frequency representation, where the predictor is encoded into an ordered set of frequencies that correspond to the roots of the corresponding line spectral polynomials. While this representation is robust in terms of quantization, methods available for finding the line spectral frequencies are computationally complex. In this work, we present a method for finding these frequencies using the FFT, including methods for limiting numerical range in fixed-point implementations. Our experiments show that, in comparison to a zero-crossing search in the Chebyshev domain, the proposed method reduces complexity and improves robustness, while retaining accuracy. Tom Bäckström, Christian Fischer Pedersen, Johannes Fischer 0002, Grzegorz Pietrzyk |
ICASSP | 1 |
| 2015 | Intelligibility evaluation of speech coding standards in severe background noise and packet loss conditionsabstractSpeech intelligibility is an important aspect of speech transmission but often when speech coding standards are compared only the quality is evaluated using perceptual tests. In this study, the performance of three wideband speech coding standards, adaptive multi-rate wideband (AMR-WB), G.718, and enhanced voice services (EVS), is evaluated in a subjective intelligibility test. The test covers different packet loss conditions as well as a near-end background noise condition. Additionally, an objective quality evaluation in different packet loss conditions is conducted. All of the test conditions extend beyond the specification range to evaluate the attainable performance of the codecs in extreme conditions. The results of the subjective tests show that both EVS and G.718 are better in terms of intelligibility than AMR-WB. EVS attains the same performance as G.718 with lower algorithmic delay. Emma Jokinen, Jérémie Lecomte, Nadja Schinkel-Bielefeld, Tom Bäckström |
ICASSP | 4 |
| 2015 | Glottal inverse filtering based on quadratic programming
Manu Airaksinen, Tom Bäckström, Paavo Alku |
INTERSPEECH | 2 |
| 2015 | Decorrelating MVDR Filterbanks Using the Non-Uniform Discrete Fourier TransformabstractMinimum variance distortionless response (MVDR) is a classic design criteria in signal adaptive spectral analysis as well as filter design. We extend this approach to filterbanks with the constraint that transform domain signal components must be uncorrelated. Our analysis shows that filterbanks based on Vandermonde decomposition of the autocorrelation matrix correspond to the non-uniform discrete Fourier transform and satisfies the MVDR criteria. Namely, the columns of the inverse Vandermonde matrix corresponds to filters with unit response at the pass-band while leakage is minimized with the constraint that components remain uncorrelated. In the special case that the autocorrelation matrix is rank deficient, the proposed filterbank coincides with Pisarenko's harmonic decomposition, which thus also satisfies the MVDR criteria. Tom Bäckström |
IEEE Signal Process. Lett. | 1 |
| 2014 | Automatic estimation of the lip radiation effect in glottal inverse filtering
Manu Airaksinen, Tom Bäckström, Paavo Alku |
INTERSPEECH | 2 |
| 2014 | Decorrelated innovative codebooks for ACELP using factorization of autocorrelation matrix
Tom Bäckström, Christian R. Helmrich |
INTERSPEECH | 1 |
| 2014 | Sparse time-frequency representation of speech by the vandermonde transform
Christian Fischer Pedersen, Tom Bäckström |
INTERSPEECH | 2 |
| 2013 | Computationally efficient objective function for algebraic codebook optimization in ACELPabstractS.3434-3438 Tom Bäckström |
INTERSPEECH | 1 |
| 2012 | Enumerative Algebraic Coding for ACELPabstractS.930-933 Tom Bäckström |
INTERSPEECH | 1 |
| 2009 | Pitch variation estimationabstractA method for estimating the normalised pitch variation is described. While pitch tracking is a classical problem, in applications where the pitch magnitude is not required but only the change in pitch, all the main problems of pitch tracking can be avoided, such as octave jumps and intricate peak-finding heuristics. The presented approach is efficient, accurate and unbiased. It was developed for use in speech and audio coding for pitch variation compensation, but can also be used as additional information for pitch tracking. Tom Bäckström, Stefan Bayer, Sascha Disch |
INTERSPEECH | 1 |
| 2009 | Stabilised weighted linear prediction
Carlo Magi, Jouni Pohjalainen, Tom Bäckström, Paavo Alku |
Speech Commun. | 3 |
| 2008 | DC-constrained linear prediction for glottal inverse filtering
Paavo Alku, Carlo Magi, Tom Bäckström |
INTERSPEECH | 3 |
| 2008 | Simple proofs of root locations of two symmetric linear prediction models
Carlo Magi, Tom Bäckström, Paavo Alku |
Signal Process. | 2 |
| 2007 | Stabilised weighted linear prediction - a robust all-pole method for speech processing
Carlo Magi, Tom Bäckström, Paavo Alku |
INTERSPEECH | 2 |
| 2007 | Effect of White-Noise Correction on Linear Predictive CodingabstractWhite-noise correction is a technique used in speech coders using linear predictive coding (LPC). This technique generates an artificial noise-floor in order to avoid stability problems caused by numerical round-off errors. In this letter, we study the effect of white-noise correction on the roots of the LPC model. The results demonstrate in analytic form the relation between the noise floor level and the stability radius of the LPC model Tom Bäckström, Carlo Magi |
IEEE Signal Process. Lett. | 1 |
| 2007 | Minimum Separation of Line Spectral FrequenciesabstractWe provide a theoretical lower limit on the distance of line spectral frequencies for both the line spectrum pair decomposition and the immittance spectrum pair decomposition. The result applies to line spectral frequencies computed from linear predictive polynomials with all roots within a zero-centered circle of radius r<1 Tom Bäckström, Carlo Magi, Paavo Alku |
IEEE Signal Process. Lett. | 1 |
| 2006 | Properties of line spectrum pair polynomials - A review
Tom Bäckström, Carlo Magi |
Signal Process. | 1 |
| 2005 | Objective Quality Measures for Glottal Inverse Filtering of Speech Pressure SignalsabstractGlottal inverse filtering is a process where the effects of the vocal tract are cancelled from the speech signal in order to estimate the voice source. Traditionally, inverse filtering methods have involved a high level of manual tuning of parameters, such as the vocal tract model order. We present objective heuristics for the measurement of the quality of the resulting glottal flow estimate. In addition, we propose an automatic method for determining the order of the vocal tract all-pole model in inverse filtering based on phase-plane analysis and estimation of the glottal flow kurtosis. Tom Bäckström, Matti Airas, Laura Lehto, Paavo Alku |
ICASSP (1) | 1 |
| 2005 | A toolkit for voice inverse filtering and parametrisation
Matti Airas, Hannu Pulakka, Tom Bäckström, Paavo Alku |
INTERSPEECH | 3 |
| 2005 | Group delay function as a means to assess quality of glottal inverse filtering
Paavo Alku, Matti Airas, Tom Bäckström, Hannu Pulakka |
INTERSPEECH | 3 |
| 2004 | Linear predictive method for improved spectral modeling of lower frequencies of speech with small prediction ordersabstractAn all-pole modeling technique, Linear Prediction with Low-frequency Emphasis (LPLE), which emphasizes the lower frequency range of the input signal, is presented. The method is based on first interpreting conventional linear predictive (LP) analyses of successive prediction orders with parallel structures using the concept of symmetric linear prediction. In these implementations, symmetric linear prediction is preceded by simple pre-filters, which are of either low or high frequency characteristics. Combining those symmetric linear predictors that are not preceded by high-frequency pre-filters yields the proposed LPLE predictor. It is proved that the all-pole filters computed by LPLE are always stable. The results achieved with vowels show that the proposed method is well-suited for those applications, where low-order all-pole models with improved modeling of the lowest formants, are needed. Paavo Alku, Tom Bäckström |
IEEE Trans. Speech Audio Process. | 2 |
| 2004 | A time-domain interpretation for the LSP decompositionabstractThe line spectrum pair (LSP) decomposition is a widely used method in speech coding. In this article, we will show that the LSP polynomials, whose trivial zeros have been removed, are equivalent to two optimal (in the mean square sense) predictors in which a sample is predicted from linear combinations of its previous averaged and differentiated values. Tom Bäckström, Paavo Alku, Tuomas Paatero, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 1 |
| 2003 | All-pole modeling of wide-band speech with symmetric linear predictionabstractA new linear predictive technique, all-pole modeling with symmetric linear prediction (ASLP), is presented. The starting point of the method is an implementation of conventional linear prediction (LP) with a parallel structure, where two symmetric linear predictors are combined to prefilters represented by first order FIRs. Modification of these prefilters yields the ASLP predictor, which is always minimum phase. Experiments indicate that the new method models the formant structure of wide-band speech more accurately than conventional LP, when the prediction order is smaller than the one required by the sampling frequency. Paavo Alku, Tom Bäckström |
ICASSP (1) | 2 |
| 2003 | On the stability of constrained linear predictive modelsabstractStability of the all-pole model in conventional, unconstrained linear prediction with the autocorrelation criterion is well known. By exerting constraints to the optimisation problem it is possible to define models of order m + l with m parameters. However, traditionally constraints have led to models whose stability is not guaranteed. In this paper, we discuss constrained linear predictive models where the constraint is one-dimensional (l = 1) and derive stability criteria for these models. Tom Bäckström, Paavo Alku |
ICASSP (6) | 1 |
| 2003 | Linear predictive method with low-frequency emphasis
Paavo Alku, Tom Bäckström |
INTERSPEECH | 2 |
| 2003 | A constrained linear predictive model with the minimum-phase property
Tom Bäckström, Paavo Alku |
Signal Process. | 1 |
| 2003 | All-pole modeling technique based on weighted sum of LSP polynomialsabstractThis study presents a new technique called weighted-sum line spectrum pair (WLSP) where an all-pole filter is defined by using a sum of weighted line spectrum pair polynomials. The WLSP yields a stable all-pole filter of order m, whose autocorrelation function coincides with that of the input signal between indices 0 and m-1. By sacrificing the exact matching at index m, the WLSP models the autocorrelation of the input signal at the indices above m more accurately than conventional linear prediction (LP). Experiments with vowels show that, in comparison to the conventional LP, WLSP yields all-pole spectra that model formants with an increased dynamic range between formant peaks and spectral valleys. Tom Bäckström, Paavo Alku |
IEEE Signal Process. Lett. | 1 |
| 2003 | On line spectral frequenciesabstractThe commonly used line spectral frequencies form the roots of symmetric and antisymmetric polynomials constructed from a linear predictor. We provide a new, simpler proof that the symmetric and antisymmetric polynomials can be regarded as optimal constrained predictors that correspond to predicting from the low-pass and high-pass filtered signal, respectively. W. Bastiaan Kleijn, Tom Bäckström, Paavo Alku |
IEEE Signal Process. Lett. | 2 |
| 2002 | All-pole modeling technique based on the Weighted Sum of the LSP polynomialsabstractWith the order of prediction equal to m, conventional linear prediction (LP) yields an all-pole filter that matches exactly the autocorrelation function of the input signal between indices 0 and m. This study presents a new technique, Weighted-Sum Line Spectrum Pair (WLSP), where an all-pole filter is defined by using a sum of weighted LSP polynomials. WLSP yields a stable all-pole filter of order m, whose autocorrelation function coincides to that of the input signal between indices 0 and m −1. By sacrificing the exact matching of the autocorrelations at index m, WLSP models the autocorrelation of the input signal at the indices above m more accurately than conventional LP. The current paper presents mathematical properties of WLSP together with preliminary results on modeling of vowels. The results indicate that WLSP, in comparison to conventional LP is able to yield all-pole filters that model especially the upper formants with a larger dynamic range between formant peaks and spectral valleys. Paavo Alku, Tom Bäckström |
ICASSP | 2 |
| 2002 | A time domain reformulation of linear prediction equivalent to the LSP decompositionabstractThe Line spectrum pair (LSP) decomposition is a widely used method in speech coding. In this paper, we will present a reformulation of conventional linear prediction which is equivalent to the LSP decomposition. The paper shows that the symmetric and antisymmetric polynomials of the LSP decomposition are equivalent to two filters, that are determined by predicting a signal sample using its averaged and differentiated previous values. Tom Bäckström, Paavo Alku, W. Bastiaan Kleijn |
ICASSP | 1 |
| 2002 | All-pole modeling of wide-band speech using weighted sum of the LSP polynomials
Paavo Alku, Tom Bäckström |
INTERSPEECH | 2 |
| 2002 | Time-domain parameterization of the closing phase of glottal airflow waveform from voices over a large intensity rangeabstractThe aim of this paper is to analyze and compare two time-domain parameterization methods of the glottal flow waveform on a large intensity range. The first parameter is the classical closing quotient which indicates the portion of a period where the glottis is closing. The second parameter is the normalized amplitude quotient which is defined using the ratio between the maximum flow amplitude and the negative peak amplitude of the differentiated glottal flow. The parameters are shown to be strongly correlated, and the normalized amplitude quotient to be a more accurate, consistent and robust measure than the closing quotient. The subjects, five female and six male, produced sustained phonations on a large intensity range. On this material, the normalized amplitude quotient is shown to vary systematically with sound pressure level, and it reveals information that for the closing quotient is hidden in the local variance. Tom Bäckström, Paavo Alku, Erkki Vilkman |
IEEE Trans. Speech Audio Process. | 1 |