Jan Skoglund

dblp:59/7156 · DBLP profile ↗
← Back
41ranked-venue papers
6as first author
11since 2021 · last 2025
0009-0008-0167-4628ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 35 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3
YearPublicationVenuePosition
2025 Perceptual Audio Coding: A 40-Year Historical Perspective
abstract
In the history of audio and acoustic signal processing, perceptual audio coding has certainly excelled as a bright success story by its ubiquitous deployment in virtually all digital media devices, such as computers, tablets, mobile phones, set-top-boxes, and digital radios. From a technology perspective, perceptual audio coding has undergone tremendous development from the first very basic perceptually driven coders (including the popular mp3 format) to today’s full-blown integrated coding/rendering systems. This paper provides a historical overview of this research journey by pinpointing the pivotal development steps in the evolution of perceptual audio coding. Finally, it provides thoughts about future directions in this area.
Jürgen Herre, Schuyler R. Quackenbush, Minje Kim 0001, Jan Skoglund
ICASSP4
2024 NOMAD: Unsupervised Learning of Perceptual Embeddings For Speech Enhancement and Non-Matching Reference Audio Quality Assessment
abstract
This paper presents NOMAD (Non-Matching Audio Distance), a differentiable perceptual similarity metric that measures the distance of a degraded signal against non-matching references. The proposed method is based on learning deep feature embeddings via a triplet loss guided by the Neurogram Similarity Index Measure (NSIM) to capture degradation intensity. During inference, the similarity score between any two audio samples is computed through Euclidean distance of their embeddings. NOMAD is fully unsupervised and can be used in general perceptual audio tasks for audio analysis e.g. quality assessment and generative tasks such as speech enhancement and speech synthesis. The proposed method is evaluated with 3 tasks. Ranking degradation intensity, predicting speech quality, and as a loss function for speech enhancement. Results indicate NOMAD outperforms other non-matching reference approaches in both ranking degradation intensity and quality assessment, exhibiting competitive performance with full-reference audio metrics. NOMAD demonstrates a promising technique that mimics human capabilities in assessing audio quality with non-matching references to learn perceptual embeddings without the need for human-generated labels.
Alessandro Ragano, Jan Skoglund, Andrew Hines
ICASSP2
2024 SCOREQ: Speech Quality Assessment with Contrastive Regression
abstract
In this paper, we present SCOREQ, a novel approach for speech quality prediction. SCOREQ is a triplet loss function for contrastive regression that addresses the domain generalisation shortcoming exhibited by state of the art no-reference speech quality metrics. In the paper we: (i) illustrate the problem of L2 loss training failing at capturing the continuous nature of the mean opinion score (MOS) labels; (ii) demonstrate the lack of generalisation through a benchmarking evaluation across several speech domains; (iii) outline our approach and explore the impact of the architectural design decisions through incremental evaluation; (iv) evaluate the final model against state of the art models for a wide variety of data and domains. The results show that the lack of generalisation observed in state of the art speech quality metrics is addressed by SCOREQ. We conclude that using a triplet loss function for contrastive regression improves generalisation for speech quality prediction models but also has potential utility across a wide range of applications using regression-based predictive models.
Alessandro Ragano, Jan Skoglund, Andrew Hines
NeurIPS2
2023 LMCodec: A Low Bitrate Speech Codec with Causal Transformer Models
abstract
We introduce LMCodec, a causal neural speech codec that provides high quality audio at very low bitrates. The backbone of the system is a causal convolutional codec that encodes audio into a hierarchy of coarse-to-fine tokens using residual vector quantization. LMCodec trains a Transformer language model to predict the fine tokens from the coarse ones in a generative fashion, allowing for the transmission of fewer codes. A second Transformer predicts the uncertainty of the next codes given the past transmitted codes, and is used to perform conditional entropy coding. A MUSHRA subjective test was conducted and shows that the quality is comparable to reference codecs at higher bitrates. Example audio is available at https://mjenrungrot.github.io/chrome-media-audio-papers/publications/lmcodec.
Teerapat Jenrungrot, Michael Chinen, W. Bastiaan Kleijn, Jan Skoglund, Zalan Borsos, Neil Zeghidour, Marco Tagliasacchi
ICASSP4
2023 Multi-Channel Audio Signal Generation
abstract
We present a multi-channel audio signal generation scheme based on machine-learning and probabilistic modeling. We start from modeling a multi-channel single-source signal. Such signals are naturally modeled as a single-channel reference signal and a spatial-arrangement (SA) model specified by an SA parameter sequence. We focus on the SA model and assume that the reference signal is described by some parameter sequence. The SA model parameters are described with a learned probability distribution that is conditioned by the reference-signal parameter sequence and, optionally, an SA conditioning sequence. If present, the SA conditioning sequence specifies a signal class or a specific signal. The single-source method can be used for multi-source signals by applying source separation or by using an SA model that operates on nonoverlapping frequency bands. Our GAN-based stereo coding implementation of the latter approach shows that our paradigm facilitates plausible high-quality rendering at a low bit rate for the SA conditioning.
W. Bastiaan Kleijn, Michael Chinen, Felicia Lim, Jan Skoglund
ICASSP4
2022 Using Rater and System Metadata to Explain Variance in the VoiceMOS Challenge 2022 Dataset
Michael Chinen, Jan Skoglund, Chandan K. A. Reddy, Alessandro Ragano, Andrew Hines
INTERSPEECH2
2022 Ultra-Low-Bitrate Speech Coding with Pretrained Transformers
abstract
Speech coding facilitates the transmission of speech over lowbandwidth networks with minimal distortion.Neural-network based speech codecs have recently demonstrated significant improvements in quality over traditional approaches.While this new generation of codecs is capable of synthesizing highfidelity speech, their use of recurrent or convolutional layers often restricts their effective receptive fields, which prevents them from compressing speech efficiently.We propose to further reduce the bitrate of neural speech codecs through the use of pretrained Transformers, capable of exploiting long-range dependencies in the input signal due to their inductive bias.As such, we use a pretrained Transformer in tandem with a convolutional encoder, which is trained end-to-end with a quantizer and a generative adversarial net decoder.Our numerical experiments show that supplementing the convolutional encoder of a neural speech codec with Transformer speech embeddings yields a speech codec with a bitrate of 600 bps that outperforms the original neural speech codec in synthesized speech quality when trained at the same bitrate.Subjective human evaluations suggest that the quality of the resulting codec is comparable or better than that of conventional codecs operating at three to four times the rate.
Ali Siahkoohi, Michael Chinen, Tom Denton, W. Bastiaan Kleijn, Jan Skoglund
INTERSPEECH5
2022 Speech quality assessment with WARP-Q: From similarity to subsequence dynamic time warp cost
abstract
Abstract Speech coding has been shown to achieve good speech quality using either waveform matching or parametric reconstruction. For very low bit rate streams, recently developed generative speech models can reconstruct high‐quality wideband speech from the bit streams of standard parametric encoders at less than 3 kb/s. Generative codecs produce high‐quality speech based on synthesising speech from a DNN and the parametric input. Existing objective speech quality models (e.g., ViSQOL and POLQA) cannot be used to accurately evaluate the quality of coded speech from generative models as they penalise based on signal differences not apparent in subjective listening test results. This paper presents WARP‐Q, a full‐reference objective speech quality metric that uses a dynamic time warping cost for MFCC representations of the signals. It is robust to low perceptual signal changes introduced by low bit rate neural vocoders. An evaluation using waveform matching, parametric, and generative neural vocoder‐based codecs as well as channel and environmental noise shows that WARP‐Q has better correlation and codec quality ranking for novel codecs compared to traditional metrics as well as the versatility of capturing other types of degradations, such as additive noise and transmission channel degradations.
Wissam A. Jassim, Jan Skoglund, Michael Chinen, Andrew Hines
IET Signal Process.2
2022 SoundStream: An End-to-End Neural Audio Codec
abstract
We presentSoundStream, a novel neural audio codec that can efficiently compress speech, music and general audio at bitrates normally targeted by speech-tailored codecs.SoundStreamrelies on a model architecture composed by a fully convolutional encoder/decoder network and a residual vector quantizer, which are trained jointly end-to-end. Training leverages recent advances in text-to-speech and speech enhancement, which combine adversarial and reconstruction losses to allow the generation of high-quality audio content from quantized embeddings. By training with structured dropout applied to quantizer layers, a single model can operate across variable bitrates from 3 kbps to 18 kbps, with a negligible quality loss when compared with models trained at fixed bitrates. In addition, the model is amenable to a low latency implementation, which supports streamable inference and runs in real time on a smartphone CPU. In subjective evaluations using audio at 24 kHz sampling rate,SoundStreamat 3 kbps outperforms Opus at 12 kbps and approaches EVS at 9.6 kbps. Moreover, we are able to perform joint compression and enhancement either at the encoder or at the decoder side with no additional latency, which we demonstrate through background noise suppression for speech.
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, Marco Tagliasacchi
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 Warp-Q: Quality Prediction for Generative Neural Speech Codecs
abstract
Good speech quality has been achieved using waveform matching and parametric reconstruction coders. Recently developed very low bit rate generative codecs can reconstruct high quality wideband speech with bit streams less than 3 kb/s. These codecs use a DNN with parametric input to synthesise high quality speech outputs. Existing objective speech quality models (e.g., POLQA, ViSQOL) do not accurately predict the quality of coded speech from these generative models underestimating quality due to signal differences not highlighted in subjective listening tests. We present WARP-Q, a full-reference objective speech quality metric that uses dynamic time warping cost for MFCC speech representations. It is robust to small perceptual signal changes. Evaluation using waveform matching, parametric and generative neural vocoder based codecs as well as channel and environmental noise shows that WARP-Q has better correlation and codec quality ranking for novel codecs compared to traditional metrics in addition to versatility for general quality assessment scenarios.
Wissam A. Jassim, Jan Skoglund, Michael Chinen, Andrew Hines
ICASSP2
2021 Generative Speech Coding with Predictive Variance Regularization
abstract
The recent emergence of machine-learning based generative models for speech suggests a significant reduction in bit rate for speech codecs is possible. However, the performance of generative models deteriorates significantly with the distortions present in real-world input signals. We argue that this deterioration is due to the sensitivity of the maximum likelihood criterion to outliers and the ineffectiveness of modeling a sum of independent signals with a single autoregressive model. We introduce predictive-variance regularization to reduce the sensitivity to outliers, resulting in a significant increase in performance. We show that noise reduction to remove unwanted signals can significantly increase performance. We provide extensive subjective performance evaluations that show that our system based on generative modeling provides state-of-the-art coding performance at 3 kb/s for real-world speech signals at reasonable computational complexity.
W. Bastiaan Kleijn, Andrew Storus, Michael Chinen, Tom Denton, Felicia Lim, Alejandro Luebs, Jan Skoglund, Hengchin Yeh
ICASSP7
2020 Robust Low Rate Speech Coding Based on Cloned Networks and Wavenet
abstract
Rapid advances in machine-learning based generative modeling of speech make its use in speech coding attractive. However, the current performance of such models drops rapidly with noise contamination of the input, preventing use in practical applications. We present a new speech-coding scheme that is based on features that are robust to the distortions occurring in speech-coder input signals. To this purpose, we encourage the feature encoder to provide the same independent features for each of a set of linguistically equivalent signals, obtained by adding various noises to a common clean signal. The independent features, subjected to scalar quantization, are used as a conditioning vector sequence for WaveNet. Our experiments show that a 1.8 kb/s implementation of the resulting coder provides state-of-the-art performance for clean signals, and is additionally robust to noisy input.
Felicia Lim, W. Bastiaan Kleijn, Michael Chinen, Jan Skoglund
ICASSP4
2020 Improving Opus Low Bit Rate Quality with Neural Speech Synthesis
abstract
The voice mode of the Opus audio coder can compress wideband speech at bit rates ranging from 6 kb/s to 40 kb/s. However, Opus is at its core a waveform matching coder, and as the rate drops below 10 kb/s, quality degrades quickly. As the rate reduces even further, parametric coders tend to perform better than waveform coders. In this paper we propose a backward-compatible way of improving low bit rate Opus quality by re-synthesizing speech from the decoded parameters. We compare two different neural generative models, WaveNet and LPCNet. WaveNet is a powerful, high-complexity, and high-latency architecture that is not feasible for a practical system, yet provides a best known achievable quality with generative models. LPCNet is a low-complexity, low-latency RNN-based generative model, and practically implementable on mobile phones. We apply these systems with parameters from Opus coded at 6 kb/s as conditioning features for the generative models. A listening test shows that for the same 6 kb/s Opus bit stream, synthesized speech using LPCNet clearly outperforms the output of the standard Opus decoder. This opens up ways to improve the decoding quality of existing speech and audio waveform coders without breaking compatibility.
Jan Skoglund, Jean-Marc Valin
INTERSPEECH1
2020 ViSQOL v3: An Open Source Production Ready Objective Speech and Audio Metric
abstract
Estimation of perceptual quality in audio and speech is possible using a variety of methods. The combined v3 release of ViSQOL and ViSQOLAudio (for speech and audio, respectively,) provides improvements upon previous versions, in terms of both design and usage. As an open source C++ library or binary with permissive licensing, ViSQOL can now be deployed beyond the research context into production usage. The feedback from internal production teams at Google has helped to improve this new release, and serves to show cases where it is most applicable, as well as to highlight limitations. The new model is benchmarked against real-world data for evaluation purposes. The trends and direction of future work is discussed.
Michael Chinen, Felicia Lim, Jan Skoglund, Nikita Gureev, Feargus O'Gorman, Andrew Hines
QoMEX3
2020 Speech Quality Factors for Traditional and Neural-Based Low Bit Rate Vocoders
abstract
This study compares the performances of different algorithms for coding speech at low bit rates. In addition to widely deployed traditional vocoders, a selection of recently developed generative-model-based coders at different bit rates are contrasted. Performance analysis of the coded speech is evaluated for different quality aspects: accuracy of pitch periods estimation, the word error rates for automatic speech recognition, and the influence of speaker gender and coding delays. A number of performance metrics of speech samples taken from a publicly available database were compared with subjective scores. Results from subjective quality assessment do not correlate well with existing full reference speech quality metrics. The results provide valuable insights into aspects of the speech signal that will be used to develop a novel metric to accurately predict speech quality from generative-model-based coders.
Wissam A. Jassim, Jan Skoglund, Michael Chinen, Andrew Hines
QoMEX2
2019 LPCNET: Improving Neural Speech Synthesis through Linear Prediction
abstract
Neural speech synthesis models have recently demonstrated the ability to synthesize high quality speech for text-to-speech and compression applications. These new models often require powerful GPUs to achieve real-time operation, so being able to reduce their complexity would open the way for many new applications. We propose LPCNet, a WaveRNN variant that combines linear prediction with recurrent neural networks to significantly improve the efficiency of speech synthesis. We demonstrate that LPCNet can achieve significantly higher quality than WaveRNN for the same network size and that high quality LPCNet speech synthesis is achievable with a complexity under 3 GFLOPS. This makes it easier to deploy neural synthesis applications on lower-power devices, such as embedded systems and mobile phones.
Jean-Marc Valin, Jan Skoglund
ICASSP2
2019 Salient Speech Representations Based on Cloned Networks
abstract
We define salient features as features that are shared by signals that are defined as being equivalent by a system designer. The definition allows the designer to contribute qualitative information. We aim to find salient features that are useful as conditioning for generative networks. We extract salient features by jointly training a set of clones of an encoder network. Each network clone receives as input a different signal from a set of equivalent signals. The objective function encourages the network clones to map their input into a set of features that is identical across the clones. It additionally encourages feature independence and, optionally, reconstruction of a desired target signal by a decoder. As an application, we train a system that extracts a time-sequence of feature vectors of speech and uses it as a conditioning of a WaveNet generative system, facilitating both coding and enhancement.
W. Bastiaan Kleijn, Felicia Lim, Michael Chinen, Jan Skoglund
INTERSPEECH4
2019 A Real-Time Wideband Neural Vocoder at 1.6kb/s Using LPCNet
abstract
Neural speech synthesis algorithms are a promising new approach for coding speech at very low bitrate. They have so far demonstrated quality that far exceeds traditional vocoders, at the cost of very high complexity. In this work, we present a low-bitrate neural vocoder based on the LPCNet model. The use of linear prediction and sparse recurrent networks makes it possible to achieve real-time operation on general-purpose hardware. We demonstrate that LPCNet operating at 1.6 kb/s achieves significantly higher quality than MELP and that uncompressed LPCNet can exceed the quality of a waveform codec operating at low bitrate. This opens the way for new codec designs based on neural synthesis models.
Jean-Marc Valin, Jan Skoglund
INTERSPEECH2
2018 Wavenet Based Low Rate Speech Coding
abstract
Traditional parametric coding of speech facilitates low rate but provides poor reconstruction quality because of the inadequacy of the model used. We describe how a WaveNet generative speech model can be used to generate high quality speech from the bit stream of a standard parametric coder operating at 2.4 kb/s. We compare this parametric coder with a waveform coder based on the same generative model and show that approximating the signal waveform incurs a large rate penalty. Our experiments confirm the high performance of the WaveNet based coder and show that the speech produced by the system is able to additionally perform implicit bandwidth extension and does not significantly impair recognition of the original speaker for the human listener, even when that speaker has not been used during the training of the generative model.
W. Bastiaan Kleijn, Felicia Lim, Alejandro Luebs, Jan Skoglund, Florian Stimberg, Thomas C. Walters
ICASSP4
2018 Beamforming with Partial Knowledge of the Acoustic Scenario
abstract
We address the problem of acoustic beamforming with a small number of microphones given only limited knowledge of the spatial scenario. We first identify a set of plausible target (desired source) scenarios and a set of interferer scenarios that have desirable suppression characteristics. We then design soft masks (postprocessors) for all target-interferer scenario pairs and select the composite scenario that maximizes the output variance over target scenarios and minimizes it over interferer scenarios. This corresponds to an approximate concatenation of soft masks. The result is a nonlinear beamformer with an adjustable beamwidth and an adjustable region where point interferers are strongly suppressed, even for two microphones. For the individual masks we use a new postprocessor formulation that is robust to scenario mismatch. The resulting system provides excellent performance with few microphones even when the acoustic scenario is ill-defined.
W. Bastiaan Kleijn, Christopher Laguna, Alejandro Luebs, Andrew MacDonald, Jan Skoglund
MMSP5
2018 AMBIQUAL - a full reference objective quality metric for ambisonic spatial audio
abstract
Streaming spatial audio over networks requires efficient encoding techniques that compress the raw audio content without compromising quality of experience. Streaming service providers such as YouTube need a perceptually relevant objective audio quality metric to monitor users' perceived quality and spatial localization accuracy. In this paper we introduce a full reference objective spatial audio quality metric, AMBIQUAL, which assesses both Listening Quality and Localization Accuracy. In our solution both metrics are derived directly from the B-format Ambisonic audio. The metric extends and adapts the algorithm used in ViSQOLAudio, a full reference objective metric designed for assessing speech and audio quality. In particular, Listening Quality is derived from the omnidirectional channel and Localization Accuracy is derived from a weighted sum of similarity from B-format directional channels. This paper evaluates whether the proposed AMBIQUAL objective spatial audio quality metric can predict two factors: Listening Quality and Localization Accuracy by comparing its predictions with results from MUSHRA subjective listening tests. In particular, we evaluated the Listening Quality and Localization Accuracy of First and Third-Order Ambisonic audio compressed with the OPUS 1.2 codec at various bitrates (i.e. 32, 128 and 256, 512kbps respectively). The sample set for the tests comprised both recorded and synthetic audio clips with a wide range of time-frequency characteristics. To evaluate Localization Accuracy of compressed audio a number of fixed and dynamic (moving vertically and horizontally) source positions were selected for the test samples. Results showed a strong correlation (PCC=0.919; Spearman=0.882 regarding Listening Quality and PCC=0.854; Spearman=0.842 regarding Localization Accuracy) between objective quality scores derived from the B-format Ambisonic audio using AMBIQUAL and subjective scores obtained during listening MUSHRA tests. AMBIQUAL displays very promising quality assessment predictions for spatial audio. Future work will optimise the algorithm to generalise and validate it for any Higher Order Ambisonic formats.
Miroslaw Narbutt, Andrew Allen, Jan Skoglund, Michael Chinen, Andrew Hines
QoMEX3
2018 Phase-Sensitive Joint Learning Algorithms for Deep Learning-Based Speech Enhancement
abstract
This letter presents a phase-sensitive joint learning algorithm for single-channel speech enhancement. Although a deep learning framework that estimates the time-frequency (T-F) domain ideal ratio masks demonstrates a strong performance, it is limited in the sense that the enhancement process is performed only in the magnitude domain, while the phase spectra remain unchanged. Thus, recent studies have been conducted to involve phase spectra in speech enhancement systems. A phase-sensitive mask (PSM) is a T-F mask that implicitly represents phase-related information. However, since the PSM has an unbounded value, the networks are trained to target its truncated values rather than directly estimating it. To effectively train the PSM, we first approximate it to have a bounded dynamic range under the assumption that speech and noise are uncorrelated. We then propose a joint learning algorithm that trains the approximated value through its parameterized variables in order to minimize the inevitable error caused by the truncation process. Specifically, we design a network that explicitly targets three parameterized variables: 1) speech magnitude spectra; 2) noise magnitude spectra; and 3) phase difference of clean to noisy spectra. To further improve the performance, we also investigate how the dynamic range of magnitude spectra controlled by a warping function affects the final performance in joint learning algorithms. Finally, we examined how the proposed additional constraint that preserves the sum of the estimated speech and noise power spectra affects the overall system performance. The experimental results show that the proposed learning algorithm outperforms the conventional learning algorithm with the truncated phase-sensitive approximation.
Jinkyu Lee 0002, Jan Skoglund, Turaj Zakizadeh Shabestary, Hong-Goo Kang
IEEE Signal Process. Lett.2
2017 Practically efficient nonlinear acoustic echo cancellers using cascaded block RLS and FLMS adaptive filters
abstract
This paper presents a practically efficient implementation for non-linear acoustic echo cancellation (NAEC). The echo path is modeled by a novel hybrid Taylor-Volterra pre-processor followed by a linear FIR filter. A cascaded block RLS and unconstrained FLMS adaptive algorithm is developed to jointly identify the pre-processor and the FIR filter. This implementation is validated via simulations.
Yiteng Huang, Jan Skoglund, Alejandro Luebs
ICASSP2
2016 An acoustic keystroke transient canceler for speech communication terminals using a semi-blind adaptive filter model
abstract
In many teleconferencing applications using modern laptop and net-book devices it is common to encounter annoying keyboard typing noise. In this paper we propose an acoustic keystroke transient canceler for speech communication terminals as a novel broadband adaptive filter application in such a hands-free scenario. We present this approach in the context of the Google Chromebook Pixel device which is equipped with a special audio reference channel providing various new signal processing possibilities. Our novel semi-blind/semi-supervised approach exploiting this new degree of freedom, combined with the system-based broadband estimation and a novel adaptation control yields a high-quality speech enhancement even under challenging acoustic conditions.
Herbert Buchner, Jan Skoglund, Simon J. Godsill
ICASSP2
2016 Globally optimized least-squares post-filtering for microphone array speech enhancement
abstract
Existing post-filtering techniques for microphone array speech enhancement have two common deficiencies. First, they assume that the noise is either white or diffuse and cannot deal with point inter-ferers. Second, they estimate the post-filter coefficients using only two microphones at a time and then perform averaging over all microphone pairs, yielding a suboptimal solution at best. In this paper, we present a novel post-filtering algorithm that alleviates the first limitation by using a more generalized signal model including not only white and diffuse but also point interferers, and overcomes the second deficiency by offering a globally optimized least-squares solution over all microphones. It is shown by simulations that the proposed method outperforms the existing algorithms in many different acoustic scenarios.
Yiteng Huang, Alejandro Luebs, Jan Skoglund, W. Bastiaan Kleijn
ICASSP3
2015 Direct-to-Reverberant Ratio estimation using a null-steered beamformer
abstract
Reverberation affects the quality and intelligibility of distant speech recorded in a room. Direct-to-Reverberant Ratio (DRR) is a useful measure for assessing the acoustic configuration and can be used to inform dereverberation algorithms. We describe a novel DRR estimation algorithm applicable where the signal was recorded with two or more microphones, such as mobile communications devices and laptops. The method uses a null-steered beamformer. In simulations the proposed method yields accurate DRR estimates to within ±4 dB across a wide variety of room sizes, reverberation times and source-receiver distances. It is also shown that the proposed method is more robust to background noise than a baseline approach. The best estimation accuracy is obtained in the region from -5 to 5 dB which is a relevant range for portable devices.
James Eaton, Alastair H. Moore, Patrick A. Naylor, Jan Skoglund
ICASSP4
2015 Detection and suppression of keyboard transient noise in audio streams with auxiliary keybed microphone
abstract
In this paper a problem in transient noise suppression for audio streams in laptop and netbook devices is addressed. One or more microphones record voice signals which are corrupted with ambient noise and also transient noise from keyboard and mouse clicks. In the current work, a synchronous reference microphone is embedded in the keyboard which allows for measurement of the key click noise, substantially unaffected by the voice signal and ambient noise. An algorithm is here presented for incorporation of the keybed microphone as a reference signal in a signal restoration process for the voice part. The problem is substantially complicated by the presence of nonlinear vibarations (we postulate) in the hinge and casework of the laptop, which renders a simple linear suppressor ineffective in some cases. Moreover, the transfer functions between key clicks and voice microphone depend strongly upon which key is being clicked. A very low-latency solution is proposed in which short-time transform data is processed sequentially in short frames and a robust statistical model is formulated and estimated using Bayesian inference procedures. Results with real recordings show a significant reduction of typing artefacts at the expense of small amounts of voice distortion.
Simon J. Godsill, Herbert Buchner, Jan Skoglund
ICASSP3
2014 Perceived Audio Quality for Streaming Stereo Music
abstract
Users of audio-visual streaming services expect an ever increasing quality of experience. Channel bandwidth remains a bottleneck commonly addressed with lossy compression schemes for both the video and audio streams. Anecdotal evidence suggests a strongly perceived link between bit rate and quality. This paper presents three audio quality listening experiments using the ITU MUSHRA methodology to assess a number of audio codecs typically used by streaming services. They were assessed for a range of bit rates using three presentation modes: consumer and studio quality headphones and loudspeakers. Our results indicate that with consumer quality headphones, listeners were not differentiating between codecs with bit rates greater than 48 kb/s (p>=0.228). For studio quality headphones and loudspeakers aac-lc at 128 kb/s and higher was differentiated over other codecs (p<=0.001). The results provide insights into quality of experience that will guide future development of objective audio quality metrics.
Andrew Hines, Eoin Gillen, Damien Kelly, Jan Skoglund, Anil C. Kokaram, Naomi Harte
ACM Multimedia4
2013 Robustness of speech quality metrics to background noise and network degradations: Comparing ViSQOL, PESQ and POLQA
abstract
The Virtual Speech Quality Objective Listener (ViSQOL) is a new objective speech quality model. It is a signal based full reference metric that uses a spectro-temporal measure of similarity between a reference and a test speech signal. ViSQOL aims to predict the overall quality of experience for the end listener whether the cause of speech quality degradation is due to ambient noise, or transmission channel degradations. This paper describes the algorithm and tests the model using two speech corpora: NOIZEUS and E4. The NOIZEUS corpus contains speech under a variety of background noise types, speech enhancement methods, and SNR levels. The E4 corpus contains voice over IP degradations including packet loss, jitter and clock drift. The results are compared with the ITU-T objective models for speech quality: PESQ and POLQA. The behaviour of the metrics are also evaluated under simulated time warp conditions. The results show that for both datasets ViSQOL performed comparably with PESQ. POLQA was shown to have lower correlation with subjective scores than the other metrics for the NOIZEUS database.
Andrew Hines, Jan Skoglund, Anil C. Kokaram, Naomi Harte
ICASSP2
2013 Monitoring the effects of temporal clipping on voIP speech quality
abstract
This paper presents work on a real-time temporal clipping monitoring tool for VoIP. Temporal clipping can occur as a result of voice activity detection (VAD) or echo cancellation where comfort noise in used in place of clipped speech segments. The algorithm presented will form part of a no-reference objective model for quantifying perceived speech quality in VoIP. The overall approach uses a modular design that will help pinpoint the reason for degradations in addition to quantifying their impact on speech quality. The new algorithm was tested for VAD compared over a range of thresholds and varied speech frame sizes. The results are compared to objective Mean Opinion Scores (MOS-LQO) from POLQA. The results show that the proposed algorithm can efficiently predict temporal clipping in speech and correlates well with the full reference quality predictions from POLQA. The model shows good potential for use in a real-time monitoring tool. Index Terms: temporal clipping, VAD, VoIP, POLQA 1.
Andrew Hines, Jan Skoglund, Anil C. Kokaram, Naomi Harte
INTERSPEECH2
2000 A combined WI and MELP coder at 5.2 kbps
abstract
This paper presents a low bit rate speech coder that encodes speech using two parallel coders into one embedded bit frame structure of 5.2 kbps. The two coders, a mixed excitation linear predictive coder (MELP) and a waveform interpolation (WI) coder, share the same pitch and linear prediction analysis and quantization. The combined coder is structured to provide a higher performance option of WI while maintaining compatibility with standard 2.4 kbps MELP via an embedded MELP mode.
Jan Skoglund, Richard V. Cox, John S. Collura
ICASSP1
2000 Vector quantization based on Gaussian mixture models
abstract
We model the underlying probability density function of vectors in a database as a Gaussian mixture (GM) model. The model is employed for high rate vector quantization analysis and for design of vector quantizers. It is shown that the high rate formulas accurately predict the performance of model-based quantizers. We propose a novel method for optimizing GM model parameters for high rate performance, and an extension to the EM algorithm for densities having bounded support is also presented. The methods are applied to quantization of LPC parameters in speech coding and we present new high rate analysis results for band-limited spectral distortion and outlier statistics. In practical terms, we find that an optimal single-stage VQ can operate at approximately 3 bits less than a state-of-the-art LSF-based 2-split VQ.
Per Hedelin, Jan Skoglund
IEEE Trans. Speech Audio Process.2
2000 On time-frequency masking in voiced speech
abstract
This paper addresses the issue of masking of noise in voiced speech. First, we examine the audibility of cyclostationary narrow-band noise bursts added to voiced speech generated by synthetic excitation. Varying the temporal location of noise within a pitch cycle corresponds to varying its phase spectrum. Using this fact, we found that a change of phase of the noise in the high frequency region is more perceptible for a low-pitched sound than for a high-pitched sound. We then propose a pitch-dependent temporal weighting function which can be employed in quantization of pitch cycle waveforms. In a second experiment, we found that the audibility of high-frequency noise added to natural speech can be significantly reduced using this weighting function.
Jan Skoglund, W. Bastiaan Kleijn
IEEE Trans. Speech Audio Process.1
1999 Performance bounds for LPC spectrum quantization
abstract
This paper presents a method for obtaining numerical estimates of high rate vector quantization (VQ) performance suitable for sources for which the PDF is not analytically available. In the proposed method, the VQ point density is described from a Gaussian mixture model optimized for the data. Employing this method for LPC spectrum quantization, we obtain high rate expressions for both the average spectral distortion (SD) and the distribution function of the SD. We estimate the minimum bits required for a quantizer to obtain an average SD of 1 dB and the outlier statistics for that quantizer. We find that approximately 3 bits can be saved as compared to a 2-split LSF-based vector quantizer.
Per Hedelin, Jan Skoglund, Jonas Samuelsson
ICASSP2
1999 Interframe LSF quantization for noisy channels
abstract
In linear predictive speech coding algorithms, transmission of linear predictive coding (LPC) parameters-often transformed to the line spectrum frequencies (LSF) representation-consumes a large part of the total bit rate of the coder. Typically, the LSF parameters are highly correlated from one frame to the next, and a considerable reduction in bit rate can be achieved by exploiting this interframe correlation. However, interframe coding leads to error propagation if the channel is noisy, which possibly cancels the achievable gain. In this paper, several algorithms for exploiting interframe correlation of LSF parameters are compared. Especially, performance for transmission over noisy channels is examined, and methods to improve noisy channel performance are proposed. By combining an interframe quantizer and a memoryless "safety-net" quantizer, we demonstrate that the advantages of both quantization strategies can be utilized, and the performance for both noiseless and noisy channels improves. The results indicate that the best interframe method performs as good as a memoryless quantizing scheme, with 4 bits less per frame. Subjective listening tests have been employed that verify the results from the objective measurements.
Thomas Eriksson, Jan Linden, Jan Skoglund
IEEE Trans. Speech Audio Process.3
1998 On nonlinear utilization of intervector dependency in vector quantization
abstract
This paper presents an approach to speech vector quantization of sources exhibiting intervector dependency. We present the optimal decoder based on a collection of received indices. We also present the optimal encoder for such decoding. The optimal decoder can be implemented as a table look-up decoder, however the size of the decoder codebook grows very fast with the size of the collection of utilized indices. This leads us to introduce a method for storing an approximation to the set of optimal decoder vectors, based on linear mapping of a block code vector quantization. In this approach a heavily reduced set of parameters is employed to represent the codebook. Furthermore, we illustrate that the proposed scheme has an interpretation as nonlinear predictive quantization. Numerical results indicate high gain over memoryless coding and memory quantization based on linear predictive coding. The results also show that the sub-optimal approach performs close to the optimal.
Mikael Skoglund, Jan Skoglund
ICASSP2
1998 On the significance of temporal masking in speech coding
Jan Skoglund, W. Bastiaan Kleijn
ICSLP1
1998 Analysis and quantization of glottal pulse shapes
Jan Skoglund
Speech Commun.1
1997 Predictive VQ for noisy channel spectrum coding: AR or MA?
abstract
In this paper, the performance of different predictive vector quantization (PVQ) structures is studied and compared for different degrees of channel noise. Predictive quantization schemes with an auto-regressive (AR) decoder structure are compared with schemes that employ a moving average (MA) decoder. For noisy channels MA prediction performs better than AR. It is shown here that a combination of a PVQ scheme (AR or MA) and a memoryless VQ outperforms both types of traditional predictive quantizer schemes in noiseless as well as noisy channels.
Jan Skoglund, Jan Linden
ICASSP1
1996 Exploiting interframe correlation in spectral quantization: a study of different memory VQ schemes
abstract
This paper addresses the problem of efficient transmission of the LSF parameters in speech coding using vector quantization (VQ). By performing a comparison of several memory VQ methods on the same database, we investigate what gains can be achieved by exploiting interframe correlation. The memory VQ methods studied are finite-state VQ and linear predictive VQ. By combining the memory VQ with a fixed memoryless VQ, called the safety-net, further improvements in performance can be obtained. It is found that memory VQ can improve the performance with 3-5 bits compared to memoryless VQ for error-free transmission. The best method in this study is a safety-net extended predictive VQ. For noisy channels, most memory methods perform worse than memoryless VQ, but the safety-net predictive VQ outperforms memoryless VQ for all tested channel error rates, with 4 bits less.
Thomas Eriksson, Jan Linden, Jan Skoglund
ICASSP3
1995 Vector quantization of glottal pulses
abstract
An efficient codebook driven voiced excitation coding method producing natural sounding speech is proposed. It can be incorporated as an essential part in a complete speech coder working at low bit rates. The inter-pulse correlation of such a coding scheme is investigated and exploited using linear predictive vector quantization and finite state vector quantization (FSVQ). A new and more robust FSVQ method that is able to more efficiently exploit this correlation is proposed. The new method dynamically combines two memoryless vector quantizers. This dynamic combination method is not restricted to pulse codebooks but can be employed for any coding scheme with inter vector correlation. 1. INTRODUCTION Most modern speech coders are based on the source-filter concept which tries to separate the excitation source waveform from the spectral characteristics of the vocal tract. For voiced speech, the excitation waveform essentially describes the airflow through the vocal folds. For coders op...
Thomas Eriksson, Jan Linden, Jan Skoglund
EUROSPEECH3