EDBT 2026 Demo / reviewers in the wild / expert
W. Bastiaan Kleijn
dblp:30/797 · also Willem B. Kleijn
· DBLP profile ↗
222ranked-venue papers
26as first author
20since 2021 · last 2025
0000-0002-1973-3920ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 160 · 23 first-author · 13 since 2021Artificial intelligence and machine learning · 88 · 5 first-author · 12 since 2021Computer networks · 7Databases, data management, data science and information retrieval · 4Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Kolmogorov-Arnold Networks Still Catastrophically Forget but Differently from MLPabstractCatastrophic forgetting is when a neural network loses previously learnt information after learning a new task sequentially. Avoiding catastrophic forgetting could reduce the resources necessary to update neural networks. Recently, Kolmogorov–Arnold Networks (KAN) gained the community's attention as preliminary experiments suggest KAN avoid catastrophic forgetting. KAN replace neural network edges with learnable B-splines and sum incoming edges in nodes. Proponents of KAN argue they avoid forgetting, are more accurate, are interpretable, and use fewer parameters. Our work investigates the claims that KAN avoid catastrophic forgetting, finding that they fail to do so on more complex datasets containing features that overlap between tasks. We give a simple explanation as to why and how KAN catastrophically forget. Motivated by evidence suggesting KAN are superior for symbolic regression, we augment KAN in the same ways as multilayer perceptron (MLP) to perform continual learning tasks, making special accommodations to support KAN. Our experiments found that unmodified KAN often forget more than MLP, but KAN can be better than MLP when combined with continual learning strategies. We aim to highlight some of the current shortcomings and strengths associated with KAN for continual learning. Anton Lee, Heitor Murilo Gomes, Yaqian Zhang 0004, W. Bastiaan Kleijn |
AAAI | 4 |
| 2025 | On Exact Bit-level Reversible Transformers Without Changing ArchitectureabstractIn this work we present the BDIA-transformer, which is an exact bit-level reversible transformer that uses an unchanged standard architecture for inference. The basic idea is to first treat each transformer block as the Euler integration approximation for solving an ordinary differential equation (ODE) and then incorporate the technique of bidirectional integration approximation (BDIA) (originally designed for diffusion inversion) into the neural architecture, together with activation quantization to make it exactly bit-level reversible. In the training process, we let a hyper-parameter $\gamma$ in BDIA-transformer randomly take one of the two values $\{0.5, -0.5\}$ per training sample per transformer block for averaging every two consecutive integration approximations. As a result, BDIA-transformer can be viewed as training an ensemble of ODE solvers parameterized by a set of binary random variables, which regularizes the model and results in improved validation accuracy. Lightweight side information is required to be stored in the forward process to account for binary quantization loss to enable exact bit-level reversibility. In the inference procedure, the expectation $\mathbb{E}(\gamma)=0$ is taken to make the resulting architecture identical to transformer up to activation quantization. Our experiments in natural language generation, image classification, and language translation show that BDIA-transformers outperform their conventional counterparts significantly in terms of validation performance while also requiring considerably less training memory. Thanks to the regularizing effect of the ensemble, the BDIA-transformer is particularly suitable for fine-tuning with limited data. Source-code can be found via https://github.com/guoqiang-zhang-x/BDIA-Transformer. Guoqiang Zhang 0003, John P. Lewis, W. Bastiaan Kleijn |
ICML | 3 |
| 2025 | Zero-Shot Mono-to-Binaural Speech Synthesis
Alon Levkovitch, Julian Salazar, Soroosh Mariooryad, R. J. Skerry-Ryan, Nadav Bar, W. Bastiaan Kleijn, Eliya Nachmani |
INTERSPEECH | 6 |
| 2025 | Revisiting 1-peer exponential graph for enhancing decentralized learning efficiencyabstractFor communication-efficient decentralized learning, it is essential to employ dynamic graphs designed to improve the expected spectral gap by reducing deviations from global averaging. The $1$-peer exponential graph demonstrates its finite-time convergence property--achieved by maximizing the expected spectral gap--but only when the number of nodes $n$ is a power of two. However, its efficiency across any $n$ and the commutativity of mixing matrices remain unexplored. We delve into the principles underlying the $1$-peer exponential graph to explain its efficiency across any $n$ and leverage them to develop new dynamic graphs. We propose two new dynamic graphs: the $k$-peer exponential graph and the null-cascade graph. Notably, the null-cascade graph achieves finite-time convergence for any $n$ while ensuring commutativity. Our experiments confirm the effectiveness of these new graphs, particularly the null-cascade graph, in most test settings. Kenta Niwa, Yuki Takezawa, Guoqiang Zhang 0003, W. Bastiaan Kleijn |
NeurIPS | 4 |
| 2025 | Deep Green's Function Tsunami InversionabstractAn important source of uncertainty in tsunami forecasting arises from uncertainty in the event’s initial conditions. In this work, we propose a dictionary-based inversion method that uses off-shore sensor data to recover the initial ocean condition, allowing inversion of any tsunami event, and leading to reduced uncertainty in forecasts. We show that deep learning models can be used to address the computational requirements that arise from dictionary-based inversion. We validate our method using simulations of historic events. Amr Morssy, Paul D. Teal, W. Bastiaan Kleijn |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Directed Diffusion: Direct Control of Object Placement through Attention GuidanceabstractText-guided diffusion models such as DALLE-2, Imagen, and Stable Diffusion are able to generate an effectively endless variety of images given only a short text prompt describing the desired image content. In many cases the images are of very high quality. However, these models often struggle to compose scenes containing several key objects such as characters in specified positional relationships. The missing capability to ``direct'' the placement of characters and objects both within and across images is crucial in storytelling, as recognized in the literature on film and animation theory. In this work, we take a particularly straightforward approach to providing the needed direction. Drawing on the observation that the cross-attention maps for prompt words reflect the spatial layout of objects denoted by those words, we introduce an optimization objective that produces ``activation'' at desired positions in these cross-attention maps. The resulting approach is a step toward generalizing the applicability of text-guided diffusion models beyond single images to collections of related images, as in storybooks. Directed Diffusion provides easy high-level positional control over multiple objects, while making use of an existing pre-trained model and maintaining a coherent blend between the positioned objects and the background. Moreover, it requires only a few lines to implement. Wan-Duo Kurt Ma, Avisek Lahiri, John P. Lewis, Thomas K. Leung, W. Bastiaan Kleijn |
AAAI | 5 |
| 2024 | Exact Diffusion Inversion via Bidirectional Integration Approximation
Guoqiang Zhang 0003, John P. Lewis, W. Bastiaan Kleijn |
ECCV (57) | 3 |
| 2024 | A Practical Online Multichannel Dereverberation Approach with Data-Reuse TechniqueabstractOne of the most effective online dereverberation algorithms is the weighted prediction error (WPE) method and its improved version, switching WPE (SwWPE). This paper proposes a practical online dereverberation approach to improve SwWPE by introducing a datareuse technique, DR-SwWPE; we then show analytically that DR-SwWPE is equivalent to a SwWPE followed by a post-filter to suppress the residual reverberation. Secondly, we compare this suppression technique and two explicit post-filtering schemes to SwWPE using both simulated data and real-world recordings. Experimental results show that DR-SwWPE out-performs SwWPE with little additional computational cost. Weilong Huang, Jinwei Feng, W. Bastiaan Kleijn |
ICASSP | 4 |
| 2024 | On Accelerating Diffusion-Based Sampling Processes via Improved Integration ApproximationabstractA popular approach to sample a diffusion-based generative model is to solve an ordinary differential equation (ODE). In existing samplers, the coefficients of the ODE solvers are pre-determined by the ODE formulation, the reverse discrete timesteps, and the employed ODE methods. In this paper, we consider accelerating several popular ODE-based sampling processes (including EDM, DDIM, and DPM-Solver) by optimizing certain coefficients via improved integration approximation (IIA). We propose to minimize, for each time step, a mean squared error (MSE) function with respect to the selected coefficients. The MSE is constructed by applying the original ODE solver for a set of fine-grained timesteps, which in principle provides a more accurate integration approximation in predicting the next diffusion state. The proposed IIA technique does not require any change of a pre-trained model, and only introduces a very small computational overhead for solving a number of quadratic optimization problems. Extensive experiments show that considerably better FID scores can be achieved by using IIA-EDM, IIA-DDIM, and IIA-DPM-Solver than the original counterparts when the neural function evaluation (NFE) is small (i.e., less than 25). Guoqiang Zhang 0003, Kenta Niwa, W. Bastiaan Kleijn |
ICLR | 3 |
| 2024 | TrailBlazer: Trajectory Control for Diffusion-Based Video Generation
Wan-Duo Kurt Ma, John P. Lewis, W. Bastiaan Kleijn |
SIGGRAPH Asia | 3 |
| 2023 | Lookahead Diffusion Probabilistic Models for Refining Mean EstimationabstractWe propose lookahead diffusion probabilistic models (LA-DPMs) to exploit the correlation in the outputs of the deep neural networks (DNNs) over subsequent timesteps in diffusion probabilistic models (DPMs) to refine the mean estimation of the conditional Gaussian distributions in the backward process. A typical DPM first obtains an estimate of the original data sample$x$by feeding the most recent state$z_{i}$and index$i$into the DNN model and then computes the mean vector of the conditional Gaussian distribution for$z_{i-1}$. We propose to calculate a more accurate estimate for$x$by performing extrapolation on the two estimates of$x$that are obtained by feeding ($z_{i+1},i+1$) and ($z_{i},i$) into the DNN model. The extrapolation can be easily integrated into the backward process of existing DPMs by introducing an additional connection over two consecutive timesteps, and fine-tuning is not required. Extensive experiments showed that plugging in the additional connection into DDPM, DDIM, DEIS, S-PNDM, and high-order DPM-Solvers leads to a significant performance gain in terms of Fréchet inception distance (FID) score. Our implementation is available at https://github.com/guoqiang-zhang-x/LA-DPM. Guoqiang Zhang 0003, Kenta Niwa, W. Bastiaan Kleijn |
CVPR | 3 |
| 2023 | LMCodec: A Low Bitrate Speech Codec with Causal Transformer ModelsabstractWe introduce LMCodec, a causal neural speech codec that provides high quality audio at very low bitrates. The backbone of the system is a causal convolutional codec that encodes audio into a hierarchy of coarse-to-fine tokens using residual vector quantization. LMCodec trains a Transformer language model to predict the fine tokens from the coarse ones in a generative fashion, allowing for the transmission of fewer codes. A second Transformer predicts the uncertainty of the next codes given the past transmitted codes, and is used to perform conditional entropy coding. A MUSHRA subjective test was conducted and shows that the quality is comparable to reference codecs at higher bitrates. Example audio is available at https://mjenrungrot.github.io/chrome-media-audio-papers/publications/lmcodec. Teerapat Jenrungrot, Michael Chinen, W. Bastiaan Kleijn, Jan Skoglund, Zalan Borsos, Neil Zeghidour, Marco Tagliasacchi |
ICASSP | 3 |
| 2023 | Multi-Channel Audio Signal GenerationabstractWe present a multi-channel audio signal generation scheme based on machine-learning and probabilistic modeling. We start from modeling a multi-channel single-source signal. Such signals are naturally modeled as a single-channel reference signal and a spatial-arrangement (SA) model specified by an SA parameter sequence. We focus on the SA model and assume that the reference signal is described by some parameter sequence. The SA model parameters are described with a learned probability distribution that is conditioned by the reference-signal parameter sequence and, optionally, an SA conditioning sequence. If present, the SA conditioning sequence specifies a signal class or a specific signal. The single-source method can be used for multi-source signals by applying source separation or by using an SA model that operates on nonoverlapping frequency bands. Our GAN-based stereo coding implementation of the latter approach shows that our paradigm facilitates plausible high-quality rendering at a low bit rate for the SA conditioning. W. Bastiaan Kleijn, Michael Chinen, Felicia Lim, Jan Skoglund |
ICASSP | 1 |
| 2023 | Neural Optimization Of Geometry And Fixed Beamformer For Linear Microphone ArraysabstractFixed beamforming based on uniform linear microphone arrays often suffers from non-optimal performance for broadband signals. This paper addresses the issue by jointly optimizing the array geometry and spatial filters through a neural network based model. The model, composed of two feed forward neural networks, is optimized in an end-to-end manner. It satisfies the distortionless constraint in the look direction. Experimental results show that the proposed model outperforms the previous state-of-the-art fixed beamformer with overall better scores. Moreover, the proposed model can control the tradeoff between Directivity Factor (DF) and White Noise Gain (WNG) in a flexible way. Longfei Yan 0001, Weilong Huang, W. Bastiaan Kleijn, Thushara D. Abhayapala |
ICASSP | 3 |
| 2022 | Wave-Domain Approach for Cancelling Noise Entering Open WindowsabstractActive control of noise propagating through apertures is commonly realized with closed-loop LMS algorithms. However, these algorithms require a large number of error microphones and provide only local attenuation. Slow convergence and high computational effort are additional disadvantages. We propose a wave-domain approach that converges instantaneously, operates with low computational effort and does not require error microphones. It inherently controls sound in all directions in the far-field. The soundfield from the aperture is matched in a least squares sense with the generated soundfield from the loudspeaker array using orthonormal basis functions. Compensation for algorithmic delay, induced by blockwise processing, can be based on microphone placement or signal prediction, at the cost of a loss in attenuation performance. Our simulation results indicate that wave-domain processing has the potential to outperform LMS-based methods in practical active noise control for apertures. Daan Ratering, W. Bastiaan Kleijn, Jean Gonzalez Silva, Riccardo M. G. Ferrari |
ICASSP | 2 |
| 2022 | Ultra-Low-Bitrate Speech Coding with Pretrained TransformersabstractSpeech coding facilitates the transmission of speech over lowbandwidth networks with minimal distortion.Neural-network based speech codecs have recently demonstrated significant improvements in quality over traditional approaches.While this new generation of codecs is capable of synthesizing highfidelity speech, their use of recurrent or convolutional layers often restricts their effective receptive fields, which prevents them from compressing speech efficiently.We propose to further reduce the bitrate of neural speech codecs through the use of pretrained Transformers, capable of exploiting long-range dependencies in the input signal due to their inductive bias.As such, we use a pretrained Transformer in tandem with a convolutional encoder, which is trained end-to-end with a quantizer and a generative adversarial net decoder.Our numerical experiments show that supplementing the convolutional encoder of a neural speech codec with Transformer speech embeddings yields a speech codec with a bitrate of 600 bps that outperforms the original neural speech codec in synthesized speech quality when trained at the same bitrate.Subjective human evaluations suggest that the quality of the resulting codec is comparable or better than that of conventional codecs operating at three to four times the rate. Ali Siahkoohi, Michael Chinen, Tom Denton, W. Bastiaan Kleijn, Jan Skoglund |
INTERSPEECH | 4 |
| 2022 | Dirichlet Process Mixture of Generalized Inverted Dirichlet Distributions for Positive Vector Data With Extended Variational InferenceabstractA Bayesian nonparametric approach for estimation of a Dirichlet process (DP) mixture of generalized inverted Dirichlet distributions [i.e., an infinite generalized inverted Dirichlet mixture model (InGIDMM)] has been proposed. The generalized inverted Dirichlet distribution has been proven to be efficient in modeling the vectors that contain only positive elements. Under the classical variational inference (VI) framework, the key challenge in the Bayesian estimation of InGIDMM is that the expectation of the joint distribution of data and variables cannot be explicitly calculated. Therefore, numerical methods are usually applied to simulate the optimal posterior distributions. With the recently proposed extended VI (EVI) framework, we introduce lower bound approximations to the original variational objective function in the VI framework such that an analytically tractable solution can be derived. Hence, the problem in numerical simulation has been overcome. By applying the DP mixture technique, an InGIDMM can automatically determine the number of mixture components from the observed data. Moreover, the DP mixture model with an infinite number of mixture components also avoids the problems of underfitting and overfitting. The performance of the proposed approach is demonstrated with both synthesized data and real-life data applications. Zhanyu Ma, Yuping Lai, Jiyang Xie 0001, Deyu Meng, W. Bastiaan Kleijn, Jun Guo 0002, Jingyi Yu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2021 | Generative Speech Coding with Predictive Variance RegularizationabstractThe recent emergence of machine-learning based generative models for speech suggests a significant reduction in bit rate for speech codecs is possible. However, the performance of generative models deteriorates significantly with the distortions present in real-world input signals. We argue that this deterioration is due to the sensitivity of the maximum likelihood criterion to outliers and the ineffectiveness of modeling a sum of independent signals with a single autoregressive model. We introduce predictive-variance regularization to reduce the sensitivity to outliers, resulting in a significant increase in performance. We show that noise reduction to remove unwanted signals can significantly increase performance. We provide extensive subjective performance evaluations that show that our system based on generative modeling provides state-of-the-art coding performance at 3 kb/s for real-world speech signals at reasonable computational complexity. W. Bastiaan Kleijn, Andrew Storus, Michael Chinen, Tom Denton, Felicia Lim, Alejandro Luebs, Jan Skoglund, Hengchin Yeh |
ICASSP | 1 |
| 2021 | Asynchronous Decentralized Optimization With Implicit Stochastic Variance ReductionabstractA novel asynchronous decentralized optimization method that follows Stochastic Variance Reduction (SVR) is proposed. Average consensus algorithms, such as Decentralized Stochastic Gradient Descent (DSGD), facilitate distributed training of machine learning models. However, the gradient will drift within the local nodes due to statistical heterogeneity of the subsets of data residing on the nodes and long communication intervals. To overcome the drift problem, (i) Gradient Tracking-SVR (GT-SVR) integrates SVR into DSGD and (ii) Edge-Consensus Learning (ECL) solves a model constrained minimization problem using a primal-dual formalism. In this paper, we reformulate the update procedure of ECL such that it implicitly includes the gradient modification of SVR by optimally selecting a constraint-strength control parameter. Through convergence analysis and experiments, we confirmed that the proposed ECL with Implicit SVR (ECL-ISVR) is stable and approximately reaches the reference performance obtained with computation on a single-node using full data set. Kenta Niwa, Guoqiang Zhang 0003, W. Bastiaan Kleijn, Noboru Harada, Hiroshi Sawada, Akinori Fujino |
ICML | 3 |
| 2021 | Room Acoustical Parameter Estimation From Room Impulse Responses Using Deep Neural NetworksabstractWe describe a new method to estimate the geometry of a room and reflection coefficients given room impulse responses. The method utilizes convolutional neural networks to estimate the room geometry and multilayer perceptrons to estimate the reflection coefficients. The mean square error is used as the loss function. In contrast to existing methods, we do not require the knowledge of the relative positions of sources and receivers in the room. The method can be used with only a single RIR between one source and one receiver. For simulated environments, the proposed estimation method can achieve an average of 0.04 m accuracy for each dimension in room geometry estimation and 0.09 accuracy in reflection coefficients. For real-world environments, the room geometry estimation method achieves an accuracy of an average of 0.065 m for each dimension. Wangyang Yu 0002, W. Bastiaan Kleijn |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | The HSIC Bottleneck: Deep Learning without Back-PropagationabstractWe introduce the HSIC (Hilbert-Schmidt independence criterion) bottleneck for training deep neural networks. The HSIC bottleneck is an alternative to the conventional cross-entropy loss and backpropagation that has a number of distinct advantages. It mitigates exploding and vanishing gradients, resulting in the ability to learn very deep networks without skip connections. There is no requirement for symmetric feedback or update locking. We find that the HSIC bottleneck provides performance on MNIST/FashionMNIST/CIFAR10 classification comparable to backpropagation with a cross-entropy target, even when the system is not encouraged to make the output resemble the classification labels. Appending a single layer trained with SGD (without backpropagation) to reformat the information further improves performance. Kurt Wan-Duo Ma, John P. Lewis, W. Bastiaan Kleijn |
AAAI | 3 |
| 2020 | A Linear-time Independence Criterion Based on a Finite Basis ApproximationabstractDetection of statistical dependence between random variables is an essential component in many machine learning algorithms. We propose a novel independence criterion for two random variables with linear-time complexity. We establish that our independence criterion is an upper bound of the Hirschfeld-Gebelein-Rényi maximum correlation coefficient between tested variables. A finite set of basis functions is employed to approximate the mapping functions that can achieve the maximal correlation. Using classic benchmark experiments based on independent component analysis, we demonstrate that our independence criterion performs comparably with the state-of-the-art quadratic-time kernel dependence measures like the Hilbert-Schmidt Independence Criterion, while being more efficient in computation. The experimental results also show that our independence criterion outperforms another contemporary linear-time kernel dependence measure, the Finite Set Independence Criterion. The potential application of our criterion in deep neural networks is validated experimentally. Longfei Yan 0001, W. Bastiaan Kleijn, Thushara D. Abhayapala |
AISTATS | 2 |
| 2020 | Projected Weight Regularization to Improve Neural Network GeneralizationabstractGeneralization of a deep neural network (DNN) is one major concern when employing the deep learning approach for solving practical problems. In this paper we propose a new technique, named projected weight regularization (PWR), to improve the generalization capacity of a DNN model. Consider a weight matrix W from a particular neural layer in the model. Our objective is to make the eigenvalues of the matrix product WWThave comparable or roughly the same magnitudes while allowing the DNN model to fit the training data sufficiently accurate. Intuitively speaking, by doing so, it would prevent the W matrix from matching the training data too well. Specifically, at each iteration, we first project the W matrix to a number of vectors along randomly generated directions. After that, we build an objective function of the projected vectors to regularize their behaviours towards comparable eigenvalue magnitudes of WWT. Experimental results on training VGG16 for CIFAR10 show that PWR combined with centered weight normalization (CWN) yields promising validation performance compared to orthonormal regularisation combined with CWN. Guoqiang Zhang 0003, Kenta Niwa, W. Bastiaan Kleijn |
ICASSP | 3 |
| 2020 | Robust Low Rate Speech Coding Based on Cloned Networks and WavenetabstractRapid advances in machine-learning based generative modeling of speech make its use in speech coding attractive. However, the current performance of such models drops rapidly with noise contamination of the input, preventing use in practical applications. We present a new speech-coding scheme that is based on features that are robust to the distortions occurring in speech-coder input signals. To this purpose, we encourage the feature encoder to provide the same independent features for each of a set of linguistically equivalent signals, obtained by adding various noises to a common clean signal. The independent features, subjected to scalar quantization, are used as a conditioning vector sequence for WaveNet. Our experiments show that a 1.8 kb/s implementation of the resulting coder provides state-of-the-art performance for clean signals, and is additionally robust to noisy input. Felicia Lim, W. Bastiaan Kleijn, Michael Chinen, Jan Skoglund |
ICASSP | 2 |
| 2020 | Distributed Summation Privacy for Speech Enhancement
Matthew O'Connor, W. Bastiaan Kleijn |
INTERSPEECH | 2 |
| 2020 | Edge-consensus Learning: Deep Learning on P2P Networks with Nonhomogeneous DataabstractAn effective Deep Neural Network (DNN) optimization algorithm that can use decentralized data sets over a peer-to-peer (P2P) network is proposed. In applications such as medical data analysis, the aggregation of data in one location may not be possible due to privacy issues. Hence, we formulate an algorithm to reach a global DNN model that does not require transmission of data among nodes. An existing solution for this issue is gossip stochastic gradient descend (SGD), which updates by averaging node models over a P2P network. However, in practical situations where the data are statistically heterogeneous across the nodes and/or where communication is asynchronous, gossip SGD often gets trapped in local minimum since the model gradients are noticeably different. To overcome this issue, we solve a linearly constrained DNN cost minimization problem, which results in variable update rules that restrict differences among all node models. Our approach can be based on the Primal-Dual Method of Multipliers (PDMM) or the Alternating Direction Method of Multiplier (ADMM), but the cost function is linearized to be suitable for deep learning. It facilitates asynchronous communication. The results of our numerical experiments using CIFAR-10 indicate that the proposed algorithms converge to a global recognition model even though statistically heterogeneous data sets are placed on the nodes. Kenta Niwa, Noboru Harada, Guoqiang Zhang 0003, W. Bastiaan Kleijn |
KDD | 4 |
| 2020 | Microphone Array Wiener Post Filtering Using Monotone Operator SplittingabstractFor array-based acoustic source enhancement, variants of multi-channel Wiener filters are commonly used. The approach includes a Wiener post-filter that requires the simultaneous estimation of the power spectral density (PSD) of the target source and of noise sources for each time-frame. Conventional methods generally do not exploit prior knowledge, such as sparsity of the source, in solving this simultaneous estimation problem. We show that, for common scenarios, the simultaneous PSD estimation with consideration of prior knowledge can be formulated as a convex optimization problem with linear constraints. We use monotone operator splitting (MOS) to solve the constrained optimization problem. Our experiments confirm that the proposed method improves the accuracy of the noise PSD estimation, and that the resulting enhanced target signal is of higher quality. Kenta Niwa, Hironobu Chiba, Noboru Harada, Guoqiang Zhang 0003, W. Bastiaan Kleijn |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2019 | Fast Edge-consensus Computing Based on Bregman Monotone Operator SplittingabstractEdge-consensus computing is a framework to optimize a global cost function when distributed nodes observe distinct data sets. The distributed primal-dual method of multipliers (PDMM) and distributed alternating direction method of multipliers (ADMM) find network-global optima for edge-consensus algorithms by exchanging variables rather than data sets among the nodes. Since the distributed PDMM follows traditional Peaceman-Rachford splitting, it has a faster convergence rate than the distributed ADMM. To further speed up the convergence rate, we propose a new edge-consensus computing algorithm based on Bregman Peaceman-Rachford splitting. In traditional Peaceman-Rachford splitting, the variable update is defined based on a Euclidean metric and the convergence rate and is a form of first-order gradient descent. By generalizing the metric to a Bregman divergence and designing the divergence adaptively, our fast edge-consensus computing algorithm corresponds to the Newton or an accelerated gradient descent method. The results of our experiments confirm that the proposed algorithm can significantly improve the convergence rate of edge-consensus computing over state-of-the-art algorithms. Kenta Niwa, Guoqiang Zhang 0003, W. Bastiaan Kleijn |
ICASSP | 3 |
| 2019 | Speech Enhancement with Variance Constrained Autoencoders
Daniel T. Braithwaite, W. Bastiaan Kleijn |
INTERSPEECH | 2 |
| 2019 | Salient Speech Representations Based on Cloned NetworksabstractWe define salient features as features that are shared by signals that are defined as being equivalent by a system designer. The definition allows the designer to contribute qualitative information. We aim to find salient features that are useful as conditioning for generative networks. We extract salient features by jointly training a set of clones of an encoder network. Each network clone receives as input a different signal from a set of equivalent signals. The objective function encourages the network clones to map their input into a set of features that is identical across the clones. It additionally encourages feature independence and, optionally, reconstruction of a desired target signal by a decoder. As an application, we train a system that extracts a time-sequence of feature vectors of speech and uses it as a conditioning of a WaveNet generative system, facilitating both coding and enhancement. W. Bastiaan Kleijn, Felicia Lim, Michael Chinen, Jan Skoglund |
INTERSPEECH | 1 |
| 2019 | Maximum a posteriori Speech Enhancement Based on Double Spectrum
Pejman Mowlaee, Daniel Scheran, Johannes Stahl 0003, Sean U. N. Wood, W. Bastiaan Kleijn |
INTERSPEECH | 5 |
| 2019 | Finite Approximate Consensus for Privacy in Distributed Sensor NetworksabstractWith concepts such as the Internet of Things becoming more commonplace, greater emphasis must be placed on data privacy in large-scale public networks for these to be used securely without the threat of data theft. Most current distributed processing research deals with improving the flexibility and convergence speed of algorithms for networks of finite size with no constraints on information sharing and no concept for expected levels of signal privacy. In this work we investigate the concept of data privacy in unbounded public networks, where processing approximation is seen as a means to restrict information travel. We describe a practical method to use during processing aggregation stages that may be implemented in hardware to restrict the distance that data is shared. This method is efficient to implement, and requires very few update iterations to perform. We simulate the method and demonstrate its performance for the task of distributed acoustic beamforming in microphone sensor networks. Matthew O'Connor, W. Bastiaan Kleijn |
PDCAT | 2 |
| 2019 | A Low Latency Approach for Blind Source SeparationabstractWe present a low latency approach for blind source separation (BSS). BSS algorithms generally require a long window to estimate the demixing parameters. In traditional approaches, the long analysis window leads to a long algorithmic delay. Hence, traditional BSS approaches cannot be used in real-time systems. In contrast, our approach reduces the algorithmic delay independently of the window length used for estimation, while retaining separation performance. The new method exploits that the information about the sources provided by additional microphones can be traded against algorithmic delay. The method can be integrated with existing BSS algorithms and can be implemented in the time domain or in the time-frequency domain. Our experimental results confirm the effectiveness of our approach. Jiawen Chua, W. Bastiaan Kleijn |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Variational Bayesian Learning for Dirichlet Process Mixture of Inverted Dirichlet Distributions in Non-Gaussian Image Feature ModelingabstractIn this paper, we develop a novel variational Bayesian learning method for the Dirichlet process (DP) mixture of the inverted Dirichlet distributions, which has been shown to be very flexible for modeling vectors with positive elements. The recently proposed extended variational inference (EVI) framework is adopted to derive an analytically tractable solution. The convergency of the proposed algorithm is theoretically guaranteed by introducing single lower bound approximation to the original objective function in the EVI framework. In principle, the proposed model can be viewed as an infinite inverted Dirichlet mixture model that allows the automatic determination of the number of mixture components from data. Therefore, the problem of predetermining the optimal number of mixing components has been overcome. Moreover, the problems of overfitting and underfitting are avoided by the Bayesian estimation approach. Compared with several recently proposed DP-related methods and conventional applied methods, the good performance and effectiveness of the proposed method have been demonstrated with both synthesized data and real data evaluations. Zhanyu Ma, Yuping Lai, W. Bastiaan Kleijn, Yi-Zhe Song, Liang Wang 0001, Jun Guo 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | Training Deep Neural Networks via Optimization Over GraphsabstractIn this work, we propose to train a deep neural network by distributed optimization over a graph. Two nonlinear functions are considered: the rectified linear unit (ReLU) and a linear unit with both lower and upper cutoffs (DCutLU). The problem reformulation over a graph is realized by explicitly representing ReLU or DCutLU using a set of slack variables. We then apply the alternating direction method of multipliers (ADMM) to update the weights of the network layer-wise by solving subproblems of the reformulated problem. Empirical results suggest that the ADMM-based method is less sensitive to overfitting than the stochastic gradient descent (SGD) and Adam methods. Guoqiang Zhang 0003, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2018 | On the Comparison of Two Room Compensation / Dereverberation Methods Employing Active Acoustic Boundary AbsorptionabstractIn this paper, we compare the performance of two active dereverberation techniques using a planar array of microphones and loudspeakers. The two techniques are based on a solution to the Kirchhoff-Helmholtz Integral Equation (KHIE). We adapt a Wave Field Synthesis (WFS) based method to the application of real-time 3D dereverberation by using a low-latency pre-filter design. The use of First-Order Differential (FOD) models is also proposed as an alternative method to the use of monopoles with WFS and which does not assume knowledge of the room geometry or primary sources. The two methods are compared by observing the suppression of reflections off a single active wall over the volume of a room in the time and (temporal) frequency domain. The FOD method provides better suppression of reflections than the WFS based method but at the expense of using higher order models. The equivalent absorption coefficients are comparable to passive fibre panel absorbers. Jacob Donley, Christian H. Ritz, W. Bastiaan Kleijn |
ICASSP | 3 |
| 2018 | Wavenet Based Low Rate Speech CodingabstractTraditional parametric coding of speech facilitates low rate but provides poor reconstruction quality because of the inadequacy of the model used. We describe how a WaveNet generative speech model can be used to generate high quality speech from the bit stream of a standard parametric coder operating at 2.4 kb/s. We compare this parametric coder with a waveform coder based on the same generative model and show that approximating the signal waveform incurs a large rate penalty. Our experiments confirm the high performance of the WaveNet based coder and show that the speech produced by the system is able to additionally perform implicit bandwidth extension and does not significantly impair recognition of the original speaker for the human listener, even when that speaker has not been used during the training of the generative model. W. Bastiaan Kleijn, Felicia Lim, Alejandro Luebs, Jan Skoglund, Florian Stimberg, Thomas C. Walters |
ICASSP | 1 |
| 2018 | Beamforming with Partial Knowledge of the Acoustic ScenarioabstractWe address the problem of acoustic beamforming with a small number of microphones given only limited knowledge of the spatial scenario. We first identify a set of plausible target (desired source) scenarios and a set of interferer scenarios that have desirable suppression characteristics. We then design soft masks (postprocessors) for all target-interferer scenario pairs and select the composite scenario that maximizes the output variance over target scenarios and minimizes it over interferer scenarios. This corresponds to an approximate concatenation of soft masks. The result is a nonlinear beamformer with an adjustable beamwidth and an adjustable region where point interferers are strongly suppressed, even for two microphones. For the individual masks we use a new postprocessor formulation that is robust to scenario mismatch. The resulting system provides excellent performance with few microphones even when the acoustic scenario is ill-defined. W. Bastiaan Kleijn, Christopher Laguna, Alejandro Luebs, Andrew MacDonald, Jan Skoglund |
MMSP | 1 |
| 2018 | Cross-modal subspace learning for fine-grained sketch-based image retrieval
Peng Xu 0005, Qiyue Yin, Yongye Huang, Yi-Zhe Song, Zhanyu Ma, Liang Wang 0001, Tao Xiang 0002, W. Bastiaan Kleijn, Jun Guo 0002 |
Neurocomputing | 8 |
| 2018 | Directional Emphasis in AmbisonicsabstractWe describe an ambisonics enhancement method that increases the signal strength in specified directions at a low computational cost. The strengthening method is referred to as a utilization of an enhancement operator. The operator can be used in a static setup to emphasize the signal arriving from a particular direction or set of directions, much like a spotlight amplifies the visibility of objects in one direction. It can also be used in an adaptive arrangement where it sharpens the directionality and reduces the distortions in timbre associated with low-degree ambisonics representations. The enhancement operator can be applied directly to time-domain ambisonics representations but also to time-frequency ambisonics representations. W. Bastiaan Kleijn |
IEEE Signal Process. Lett. | 1 |
| 2018 | An Instrumental Intelligibility Metric Based on Information TheoryabstractWe propose a monaural intrusive instrumental intelligibility metric called speech intelligibility in bits (SIIB). SIIB is an estimate of the amount of information shared between a talker and a listener in bits per second. Unlike existing information theoretic intelligibility metrics, SIIB accounts for talker variability and statistical dependencies between time-frequency units. Our evaluation shows that relative to state-of-the-art intelligibility metrics, SIIB is highly correlated with the intelligibility of speech that has been degraded by noise and processed by speech enhancement algorithms. Steven Van Kuyk, W. Bastiaan Kleijn, Richard C. Hendriks |
IEEE Signal Process. Lett. | 2 |
| 2018 | Multizone Soundfield Reproduction With Privacy- and Quality-Based Speech Masking FiltersabstractReproducing zones of personal sound is a challenging signal processing problem that has garnered considerable research interest in recent years. We introduce in this work an extended method to multizone soundfield reproduction that overcomes issues with speech privacy and quality. Measures of speech intelligibility contrast (SIC) and speech quality are used as cost functions in an optimization of speech privacy and quality. Novel spatial and (temporal) frequency domain speech masker filter designs are proposed to accompany the optimization process. Spatial masking filters are designed using multizone soundfield algorithms that are dependent on the target speech multizone reproduction. Combinations of estimates of acoustic contrast and long term average speech spectra are proposed to provide equal masking influence on speech privacy and quality. Spatial aliasing specific to multizone soundfield reproduction geometry is further considered in analytically derived low-pass filters. Simulated and real-world experiments are conducted to verify the performance of the proposed method using semi-circular and linear loudspeaker arrays. Simulated implementations of the proposed method show that significant SIC and speech quality is achievable between zones. A range of perceptual evaluation of speech quality mean opinion scores that indicate good quality are obtained while at the same time providing confidential privacy as indicated by SIC. The simulations also show that the method is robust to variations in the speech, virtual source location, array geometry, and number of loudspeakers. Real-world experiments confirm the practicality of the proposed methods by showing that good quality and confidential privacy are achievable. Jacob Donley, Christian H. Ritz, W. Bastiaan Kleijn |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | An Evaluation of Intrusive Instrumental Intelligibility MetricsabstractInstrumental intelligibility metrics are commonly used as an alternative to listening tests. This paper evaluates 12 monaural intrusive intelligibility metrics: SII, HEGP, CSII, HASPI, NCM, QSTI, STOI, ESTOI, MIKNN, SIMI, SIIB, and sEPSMcorr. In addition, this paper investigates the ability of intelligibility metrics to generalize to new types of distortions and analyzes why the top performing metrics have high performance. The intelligibility data were obtained from 11 listening tests described in the literature. The stimuli included Dutch, Danish, and English speech that was distorted by additive noise, reverberation, competing talkers, preprocessing enhancement, and postprocessing enhancement. SIIB and HASPI had the highest performance achieving a correlation with listening test scores on average of p = 0.92 and p = 0.89, respectively. The high performance of SIIB may, in part, be the result of SIIBs developers having access to all the intelligibility data considered in the evaluation. The results show that intelligibility metrics tend to perform poorly on datasets that were not used during their development. By modifying the original implementations of SIIB and STOI, the advantage of reducing statistical dependencies between input features is demonstrated. Additionally, this paper presents a new version of SIIB called SIIBGauss, which has similar performance to SIIB and HASPI, but takes less time to compute by two orders of magnitude. Steven Van Kuyk, W. Bastiaan Kleijn, Richard C. Hendriks |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Non-iterative impulse response shortening method for system latency reductionabstractIn this paper we present a non-iterative impulse response shortening method aiming to reduce the latency of a system. Our method exploits that smoothing the frequency-domain response generally leads to a shorter time-domain response. The method is simple to implement and has a computational complexity that is significantly lower than that of competing methods. Yet it achieves good performance. It can be used for applications involving system identification such as blind source separation (BSS), cross-talk cancellation and channel equalization. Our experimental results confirm the effectiveness of the method, demonstrating the benefit of the approach in the BSS and cross-talk cancelling applications. Jiawen Chua, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2017 | Active speech control using wave-domain processing with a linear wall of dipole secondary sourcesabstractIn this paper, we investigate the effects of compensating for wave-domain filtering delay in an active speech control system. An active control system utilising wave-domain processed basis functions is evaluated for a linear array of dipole secondary sources. The target control soundfield is matched in a least squares sense using orthogonal wavefields to a predicted future target soundfield. Filtering is implemented using a block-based short-time signal processing approach which induces an inherent delay. We present an autoregressive method for predictively compensating for the filter delay. An approach to block-length choice that maximises the soundfield control is proposed for a trade-off between soundfield reproduction accuracy and prediction accuracy. Results show that block-length choice has a significant effect on the active suppression of speech. Jacob Donley, Christian H. Ritz, W. Bastiaan Kleijn |
ICASSP | 3 |
| 2017 | Machine learning based non-intrusive quality estimation with an augmented feature setabstractWe present a method that improves the objective quality estimation of a speech utterance. We show that including raw features that are presumably redundant reduces the effect of input noise and improves the performance of linear regressors. To exploit this effect we propose the novel idea to augment the feature set with redundant features. The proposed augmented feature set and the neural network that consists of an auto-encoder and a linear regressor leads to improved prediction accuracy of the single-ended quality assessment approach. Evaluating the system on the ITU-T Supplement 23 database illustrates that the proposed approach outperforms the current state-of-the-art. Mona Hakami, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2017 | On the information rate of speech communicationabstractThe key to the success of speech-based technology is an understanding of human speech communication. While significant advances have been made, a unified theory of speech communication that is both comprehensive and quantitative is yet to emerge. In this paper we approach speech communication from an information theoretical perspective. Without relying on prior knowledge of speech production, language, or auditory processing, we develop a new methodology for measuring the information rate of speech. Instead we rely on having recordings of multiple talkers saying the same utterance. In general, our results are consistent with a linguistic understanding of speech communication. Steven Van Kuyk, W. Bastiaan Kleijn, Richard C. Hendriks |
ICASSP | 2 |
| 2017 | Distributed TV-L1 image fusion using PDMMabstractDistributed image fusion over networks has had little coverage in the literature, particularly considering the recent emergence of large networks of imaging sensors such as radio telescope arrays and wireless self-contained node networks. We present a fully asynchronous and distributed approach for image fusion in a general network with partially overlapping node fields of view. We use the example of an aerial surveillance drone network to present the advantages of our system and show that the communication power required for performing image fusion in-network is orders of magnitude lower than transmitting all raw images back to a distant central processor, while still achieving the same fusion performance. Matthew O'Connor, W. Bastiaan Kleijn, Thushara D. Abhayapala |
ICASSP | 2 |
| 2017 | Scatter Component Analysis: A Unified Framework for Domain Adaptation and Domain GeneralizationabstractThis paper addresses classification tasks on a particular target domain in which labeled training data are only available from source domains different from (but related to) the target. Two closely related frameworks, domain adaptation and domain generalization, are concerned with such tasks, where the only difference between those frameworks is the availability of the unlabeled target data: domain adaptation can leverage unlabeled target information, while domain generalization cannot. We propose Scatter Component Analyis (SCA), a fast representation learning algorithm that can be applied to both domain adaptation and domain generalization. SCA is based on a simple geometrical measure, i.e., scatter, which operates on reproducing kernel Hilbert space. SCA finds a representation that trades between maximizing the separability of classes, minimizing the mismatch between domains, and maximizing the separability of data; each of which is quantified through scatter. The optimization problem of SCA can be reduced to a generalized eigenvalue problem, which results in a fast and exact solution. Comprehensive experiments on benchmark cross-domain object recognition datasets verify that SCA performs much faster than several state-of-the-art algorithms and also provides state-of-the-art classification accuracy in both domain adaptation and domain generalization. We also show that scatter can be used to establish a theoretical generalization bound in the case of domain adaptation. Muhammad Ghifary, David Balduzzi, W. Bastiaan Kleijn, Mengjie Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Intelligibility Enhancement Based on Mutual InformationabstractSpeech intelligibility enhancement is considered for multiple-microphone acquisition and single loudspeaker rendering. This is based on the mutual information measured between the message spoken at far-end environment and the message perceived by a listener at near-end. We prove that the joint optimal processing can be decomposed into far-end and near-end processing. The former is a minimum variance distortionless response beamformer that reduces the noise in the talker environment and the latter is a post-filter that redistributes the power over the frequency bands. Disjoint processing is optimal provided that the post-filtering operation is aware of the residual noise from the beamforming operation. Our results show that both processing steps are necessary for the effective conveyance of a message and, importantly, that the second step must be aware of the remaining noise from the beamforming operation in the first step. In addition, we study the use of the mutual information applied on the perceptually more relevant powers per critical band. Seyran Khademi, Richard C. Hendriks, W. Bastiaan Kleijn |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | New Results in Modulation-Domain Single-Channel Speech EnhancementabstractWe investigate single-channel speech enhancement using the double spectrum (DS) consisting of pitch-synchronous and modulation transforms. We first explore the fundamentals of the proposed DS domain and its advantageous properties for pitch estimation and speech presence probability estimation. We then propose speech enhancement methods based on adaptive weighting and Wiener filtering in the DS domain. We demonstrate the effectiveness of the proposed DS-based methods compared to the conventional benchmarks in the modulation or short-time Fourier transform domains. Our results show a good tradeoff between improved perceived quality and slight degradation in speech intelligibility is achieved by the proposed method across different signal-to-noise ratios and noise types. Pejman Mowlaee, Martin Blass, W. Bastiaan Kleijn |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Deep Reconstruction-Classification Networks for Unsupervised Domain Adaptation
Muhammad Ghifary, W. Bastiaan Kleijn, Mengjie Zhang 0001, David Balduzzi |
ECCV (4) | 2 |
| 2016 | Distributed linear blind source separation over wireless sensor networks with arbitrary connectivity patternsabstractBroad areal coverage and low cost make wireless sensor networks natural platforms for blind source separation (BSS). In this context, distributed processing is attractive because of low power requirements and scalability. However, existing distributed BSS algorithms either require a fully connected pattern of connectivity or require a high computational load at each sensor node. We introduce a distributed robust BSS algorithm that uses a fully shared computation and can be applied over any connected graph. This enables us to facilitate a low computational load at each node as well as low data transmission rates. Comparative experimental results confirm the effectiveness of the new method. Seyed Reza Mir Alavi, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2016 | Improving speech privacy in personal sound zonesabstractThis paper proposes two methods for providing speech privacy between spatial zones in anechoic and reverberant environments. The methods are based on masking the content leaked between regions. The masking is optimised to maximise the speech intelligibility contrast (SIC) between the zones. The first method uses a uniform masker signal that is combined with desired multizone loudspeaker signals and requires acoustic contrast between zones. The second method computes a space-time domain masker signal in parallel with the loudspeaker signals so that the combination of the two emphasises the spectral masking in the targeted quiet zone. Simulations show that it is possible to achieve a significant SIC in anechoic environments whilst maintaining speech quality in the bright zone. Jacob Donley, Christian H. Ritz, W. Bastiaan Kleijn |
ICASSP | 3 |
| 2016 | Globally optimized least-squares post-filtering for microphone array speech enhancementabstractExisting post-filtering techniques for microphone array speech enhancement have two common deficiencies. First, they assume that the noise is either white or diffuse and cannot deal with point inter-ferers. Second, they estimate the post-filter coefficients using only two microphones at a time and then perform averaging over all microphone pairs, yielding a suboptimal solution at best. In this paper, we present a novel post-filtering algorithm that alleviates the first limitation by using a more generalized signal model including not only white and diffuse but also point interferers, and overcomes the second deficiency by offering a globally optimized least-squares solution over all microphones. It is shown by simulations that the proposed method outperforms the existing algorithms in many different acoustic scenarios. Yiteng Huang, Alejandro Luebs, Jan Skoglund, W. Bastiaan Kleijn |
ICASSP | 4 |
| 2016 | Jointly optimal near-end and far-end multi-microphone speech intelligibility enhancement based on mutual informationabstractThe processing required for the global maximization of the intelligibility of speech acquired by multiple microphones and rendered by a single loudspeaker, is considered in this paper. The intelligibility is quantized, based on the mutual information rate between the message spoken by the talker and the message as interpreted by the listener. We prove that then, in each of a set of narrow-band channels, the processing can be decomposed into a minimum variance distortionless response (MVDR) beamforming operation that reduces the noise in the talker environment, followed by a gain operation that, given the far-end noise and beamforming operation, accounts for the noise at the listener end. Our experiments confirm that both processing steps are necessary for the effective conveyance ofa message and, importantly, that the second step must be aware of the first step. Seyran Khademi, Richard C. Hendriks, W. Bastiaan Kleijn |
ICASSP | 3 |
| 2016 | Distributed sparse MVDR beamforming using the bi-alternating direction method of multipliersabstractUntil now, distributed acoustic beamforming has focused on optimizing for a beamformer over an entire network, with each node contributing to the beamformer output. We present a novel approach that introduces sparsity to this beamformer computation, where we attempt to optimize for a subset of nodes within the network that produce SNR gains roughly equivalent to that of the optimal MVDR case. Due to the physical nature of sound, this approach trades a small loss in SNR for a large reduction in communication power and iterations required to produce a beamformer output by reducing the active node set of our network. Our approach operates in a fully distributed and asynchronous manner and does not require a high update iteration rate to produce an output at each sample. Matthew O'Connor, W. Bastiaan Kleijn, Thushara D. Abhayapala |
ICASSP | 2 |
| 2016 | A distributed algorithm for robust LCMV beamformingabstractIn this paper we propose a distributed reformulation of the linearly constrained minimum variance (LCMV) beamformer for use in acoustic wireless sensor networks. The proposed distributed minimum variance (DMV) algorithm, for which we demonstrate implementations for both cyclic and acyclic networks, allows the optimal beamformer output to be computed at each node without the need for sharing raw data within the network. By exploiting the low rank structure of estimated covariance matrices in time-varying noise fields, the algorithm can also provide a reduction in the total amount of data transmitted during computation when compared to centralised solutions. This is particularly true when multiple microphones are used per node. We also compare the performance of DMV with state of the art distributed beamformers and demonstrate that it achieves greater improvements in SNR in dynamic noise fields with similar transmission costs. Thomas Sherson, W. Bastiaan Kleijn, Richard Heusdens |
ICASSP | 2 |
| 2016 | Single-Channel Speech Enhancement Using Double Spectrum
Martin Blass, Pejman Mowlaee, W. Bastiaan Kleijn |
INTERSPEECH | 3 |
| 2016 | Minimum Entropy Rate Simplification of Stochastic ProcessesabstractWe propose minimum entropy rate simplification (MERS), an information-theoretic, parameterization-independent framework for simplifying generative models of stochastic processes. Applications include improving model quality for sampling tasks by concentrating the probability mass on the most characteristic and accurately described behaviors while de-emphasizing the tails, and obtaining clean models from corrupted data (nonparametric denoising). This is the opposite of the smoothing step commonly applied to classification models. Drawing on rate-distortion theory, MERS seeks the minimum entropy-rate process under a constraint on the dissimilarity between the original and simplified processes. We particularly investigate the Kullback-Leibler divergence rate as a dissimilarity measure, where, compatible with our assumption that the starting model is disturbed or inaccurate, the simplification rather than the starting model is used for the reference distribution of the divergence. This leads to analytic solutions for stationary and ergodic Gaussian processes and Markov chains. The same formulas are also valid for maximum-entropy smoothing under the same divergence constraint. In experiments, MERS successfully simplifies and denoises models from audio, text, speech, and meteorology. Gustav Eje Henter, W. Bastiaan Kleijn |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Sparse HMM-based speech enhancement method for stationary and non-stationary noise environmentsabstractWe propose a sparse hidden Markov model (HMM)-based single-channel speech enhancement method that models the speech and noise gains accurately in both stationary and nonstationary environments. The objective function is augmented with an lp regularization term resulting in a sparse autoregressive HMM (SARHMM). The method encourages sparsity in the speech- and noise- modeling, which eliminates the ambiguity between noise and speech spectra and, as a consequence, provides improved tracking of the changes of both spectral shapes and power levels of non-stationary noise. Using the modeled speech and noise SARHMMs, we first construct an estimator to estimate the noise spectrum. Then a Bayesian speech estimator is used to obtain the enhanced speech. The test results indicate that the proposed speech enhancement scheme performs much better than the reference methods in non-stationary environments, while providing state-of-the-art performance for stationary conditions. Feng Deng, Changchun Bao, W. Bastiaan Kleijn |
ICASSP | 3 |
| 2015 | A robust region-based near-field beamformerabstractIn this paper, a broadband region-based near-field beamforming algorithm is proposed and demonstrated for acoustic applications. We use an eigenfilter structure with a minimum-energy cost function based on desired and undesired near-field regions. Robustness is thus achieved by focusing on signals generated from desired zones in space while rejecting signals from undesired zones. This construction leads to a linear matrix pencil formulated in terms of the array gain to these desired and undesired zones. We include a far-field model as part of the rejection zones that further improves performance in reverberant environments. We demonstrate the robustness of the algorithm in simulated and real scenarios. Jorge Martínez 0002, Nikolay D. Gaubitch, W. Bastiaan Kleijn |
ICASSP | 3 |
| 2015 | Domain Generalization for Object Recognition with Multi-task AutoencodersabstractThe problem of domain generalization is to take knowledge acquired from a number of related domains, where training data is available, and to then successfully apply it to previously unseen domains. We propose a new feature learning algorithm, Multi-Task Autoencoder (MTAE), that provides good generalization performance for cross-domain object recognition. The algorithm extends the standard denoising autoencoder framework by substituting artificially induced corruption with naturally occurring inter-domain variability in the appearance of objects. Instead of reconstructing images from noisy versions, MTAE learns to transform the original image into analogs in multiple related domains. It thereby learns features that are robust to variations across domains. The learnt features are then used as inputs to a classifier. We evaluated the performance of the algorithm on benchmark image recognition datasets, where the task is to learn features from multiple datasets and to then predict the image label from unseen datasets. We found that (denoising) MTAE outperforms alternative autoencoder-based models as well as the current state-of-the-art algorithms for domain generalization. Muhammad Ghifary, W. Bastiaan Kleijn, Mengjie Zhang 0001, David Balduzzi |
ICCV | 2 |
| 2015 | Line spectral frequencies modeling by a mixture of von Mises-Fisher distributions
Zhanyu Ma, Jalil Taghia, W. Bastiaan Kleijn, Arne Leijon, Jun Guo 0002 |
Signal Process. | 3 |
| 2015 | A method of speech periodicity enhancement using transform-domain signal decomposition
Feng Huang 0002, Tan Lee, W. Bastiaan Kleijn, Ying-Yee Kong |
Speech Commun. | 3 |
| 2015 | A Simple Model of Speech Communication and its Application to Intelligibility EnhancementabstractWe introduce a model of communication that includes noise inherent in the message production process as well as noise inherent in the message interpretation process. The production and interpretation noise processes have a fixed signal-to-noise ratio. The resulting system is a simple but effective model of human communication. The model naturally leads to a method to enhance the intelligibility of speech rendered in a noisy environment. State-of-the-art experimental results confirm the practical value of the model. W. Bastiaan Kleijn, Richard C. Hendriks |
IEEE Signal Process. Lett. | 1 |
| 2015 | Sparse Hidden Markov Models for Speech Enhancement in Non-Stationary Noise EnvironmentsabstractWe propose a sparse hidden Markov model (HMM)-based single-channel speech enhancement method that models the speech and noise gains accurately in non-stationary noise environments. Autoregressive models are employed to describe the speech and noise in a unified framework and the speech and noise gains are modeled as random processes with memory. The likelihood criterion for finding the model parameters is augmented with an lp regularization term resulting in a sparse autoregressive HMM (SARHMM) system that encourages sparsity in the speech- and noise- modeling. In the SARHMM only a small number of HMM states contribute significantly to the model of each particular observed speech segment. As it eliminates ambiguity between noise and speech spectra, the sparsity of speech and noise modeling helps to improve the tracking of the changes of both spectral shapes and power levels of non-stationary noise. Using the modeled speech and noise SARHMMs, we first construct a noise estimator to estimate the noise power spectrum. Then, a Bayesian speech estimator is derived to obtain the enhanced speech signal. The subjective and objective test results indicate that the proposed speech enhancement scheme can achieve a larger segmental SNR improvement, a lower log-spectral distortion and a better speech quality in stationary noise conditions than state-of-the-art reference methods. The advantage of the new method is largest for non-stationary noise conditions . Feng Deng, Changchun Bao, W. Bastiaan Kleijn |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | Theory and Design of Multizone Soundfield Reproduction Using Sparse MethodsabstractMultizone soundfield reproduction over an extended spatial region is a challenging problem in acoustic signal processing. We introduce a method of reproducing a multizone soundfield within a desired region in reverberant environments. It is based on the identification of the acoustic transfer function (ATF) from the loudspeaker over the desired reproduction region using a limited number of microphone measurements. We assume that the soundfield is sparse in the domain of planewave decomposition and identify the ATF using sparse methods. The estimates of the ATFs are then used to derive the optimal least-squares solution for the loudspeaker filters that minimize the reproduction error over the entire reproduction region. Simulations confirm that the method leads to a significantly reduced number of required microphones for accurate multizone sound reproduction, while it also facilitates the reproduction over a wide frequency range. Practical experiments are used to verify the sparse planewave representation of the reverberant soundfield in a real-world listening environment. Wenyu Jin 0002, W. Bastiaan Kleijn |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Spectral Dynamics Recovery for Enhanced Speech Intelligibility in NoiseabstractSpeech intelligibility in noisy environments decreases with an increase in the noise power. We hypothesize that the differences of subsequent short-term spectra of speech, which we collectively refer to as the speech spectral dynamics, can be used to characterize speech intelligibility. We propose a distortion measure to characterize the deviation of the dynamics of the noisy modified speech from the dynamics of natural speech. Optimizing this distortion measure, we derive a parametric relationship between the signal band-power before and after modification. The parametric nature of the solution ensures adaptation to the noise level, the speech statistics and a penalty on the power gain. A multi-band speech modification system based on the single-band optimal solution is designed under a total signal power constraint and evaluated in selected noise conditions. The results indicate that the proposed approach compares favorably to a reference method based on optimizing a measure of the speech intelligibility index. Very low computational complexity and high intelligibility gain make this an attractive approach for speech modification in a wide range of application scenarios. Petko Nikolov Petkov, W. Bastiaan Kleijn |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Calibration of distributed sound acquisition systems using TOA measurements from a moving acoustic sourceabstractWe present a method for calibrating a distributed microphone array using time-of-arrival (TOA) measurements. The calibration encompasses localization and gain equalization of the microphones, which are both important in applications such as beamforming. The availability of accurate TOA measurements between the microphones and a set of spatially distributed acoustic events is pivotal to the calibration task. We propose to use a moving acoustic source emitting a calibration signal at known intervals. We then show that the TOAs and the observed signals can be used to estimate the gain differences between microphones in addition to the more established microphone localization. Finally, we provide experimental results with simulated and real measured data to demonstrate that our approach facilitates accurate TOA measurements and hence, accurate localization and gain equalization, even in reverberant and noisy conditions. Nikolay D. Gaubitch, W. Bastiaan Kleijn, Richard Heusdens |
ICASSP | 2 |
| 2014 | Deep hybrid networks with good out-of-sample object recognitionabstractWe introduce Deep Hybrid Networks that are robust to the recognition of out-of-sample objects, i.e., ones that are drawn from a different probability distribution from the training data distribution. The networks are based on a particular combination of an auto-encoder and stacked Restricted Boltzmann Machines (RBMs). The autoencoder is used to extract sparse features, which are expected to be noise invariant in the observations. The stacked RBMs then observe the sparse features as inputs to learn the top hierarchical features. The use of RBMs is motivated by the fact that the stacked RBMs typically provide good performance when dealing with in-sample observations, as proven in the previous works. To improve the robustness against local noise, we propose a variant of our hybrid network by the usage of a mixture of sparse features and sparse connections in the auto-encoder layer. The experiments show that our proposed deep networks provide good performance in both the in-sample and out-of-sample situations, particularly when the number of training examples is small. Muhammad Ghifary, W. Bastiaan Kleijn, Mengjie Zhang 0001 |
ICASSP | 2 |
| 2014 | Multizone soundfield reproduction in reverberant rooms using compressed sensing techniquesabstractWe introduce a method of reproducing a multizone soundfield within the desired region in a reverberant room. It is based on determining the acoustic transfer function (ATF) between the loudspeaker over the reproduction region using a limited number of microphones. We assume that the soundfield is sparse in the Helmholtz solution domain and find the ATF using a compressed-sensing approach. This sparseness assumption facilitates the finding of the optimal characterization the original sound over the reproduction region based on scarce sound pressure measurements. The outcome of the first stage is then used to derive the optimal least-squares solution for the loudspeaker filter that minimizes the reproduction error over the whole reproduction region. Simulations confirm that the method leads to a significant reduction in the number of required microphones for accurate multizone sound reproduction, while it also facilitates the reproduction over a wide frequency range. Wenyu Jin 0002, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2014 | On the EM algorithm for the estimation of speech AR parameters in noiseabstractIn this paper, the estimation of speech AR parameters under noisy conditions is revisited. The EM algorithm serving this purpose was first proposed by Gannot et al. We present an extensive experimental study along with a new approach to implement the E-step of the algorithm. The new realization of the E-step uses matrix computations instead of a Kalman filter. By appropriate rearrangement of the E-step, the complexity O(P(p +q)3)of the Kalman filter approach has been reduced to O(P log P), where P is the frame length, p is the speech order and q is the noise order. In practice, a speed up of the E-step of at least two orders of magnitude has been achieved. An extensive evaluation of the algorithm shows that EM algorithm in its base form is unable to improve over a recent speech enhancement method proposed by Heusdens et al. and over an established Spectral Subtraction with Minimum Statistics method, as measured by various quality measures. However, with some modification it was possible to improve over these methods in terms of spectral distortion. Marcin Kuropatwinski, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2014 | Pitch enhancement motivated by rate-distortion theoryabstractA pitch enhancement filter is designed with the objective to approach the optimal rate-distortion trade-off. The filter shows significant perceptual benefits, restating that information-theoretical and perceptual criteria are usually consistent. The filter is easy to implement and can be used as a complement to existing audio codecs. Our experiments show that it can improve the reconstruction quality of the AMR-WB standard. Obada Alhaj Moussa, Minyue Li, W. Bastiaan Kleijn |
ICASSP | 3 |
| 2014 | Diffusion-based distributed MVDR beamformerabstractAdvances in hardware and communication technology make distributed sound acquisition increasingly attractive. We describe a distributed beamforming method based on the diffusion adaptation paradigm. In contrast to existing distributed beamforming methods, the method does not impose conditions on the topology or the structure of the network nor does it require knowledge of the noise co-variance matrix. The algorithm can continuously track changes in the noise covariance matrix, making it suitable for a practical, dynamic environment. It will typically perform one iteration per signal sample, limiting communication requirements. Our experiments confirm the effectiveness of the method. Matthew O'Connor, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2014 | On the convergence rate of the bi-alternating direction method of multipliersabstractIn this paper, we analyze the convergence rate of the bi-alternating direction method of multipliers (BiADMM). Differently from ADMM that optimizes an augmented Lagrangian function, Bi-ADMM optimizes an augmented primal-dual Lagrangian function. The new function involves both the objective functions and their conjugates, thus incorporating more information of the objective functions than the augmented Lagrangian used in ADMM. We show that BiADMM has a convergence rate of O(K-1) (K denotes the number of iterations) for general convex functions. We consider the lasso problem as an example application. Our experimental results show that BiADMM outperforms not only ADMM, but fast-ADMM as well. Guoqiang Zhang 0003, Richard Heusdens, W. Bastiaan Kleijn |
ICASSP | 3 |
| 2014 | Domain Adaptive Neural Networks for Object Recognition
Muhammad Ghifary, W. Bastiaan Kleijn, Mengjie Zhang 0001 |
PRICAI | 2 |
| 2014 | Introduction to the Special Issue on The listening talker: context-dependent speech production and perception
Martin Cooke, Simon King 0001, W. Bastiaan Kleijn, Yannis Stylianou |
Comput. Speech Lang. | 3 |
| 2014 | Dirichlet mixture modeling to estimate an empirical lower bound for LSF quantization
Zhanyu Ma, Saikat Chatterjee, W. Bastiaan Kleijn, Jun Guo 0002 |
Signal Process. | 3 |
| 2013 | Auto-localization in ad-hoc microphone arraysabstractWe present a method for automatic microphone localization in adhoc microphone arrays. The localization is based on time-of-arrival (TOA) measurements obtained from spatially distributed acoustic events. In practice, measured TOAs are an incomplete representation of the true TOAs due to unknown onset times of the acoustic events and internal delays in the capturing devices and make the localization problem insoluble if not addressed appropriately. The main contribution of the proposed method is an algorithm that identifies and corrects for such internal delays and acoustic event onset times in the measured TOAs. Experimental results using both simulated and real-world data demonstrate the performance of the method and highlight the significance of correct estimation of the internal delays and onset times. Nikolay D. Gaubitch, W. Bastiaan Kleijn, Richard Heusdens |
ICASSP | 2 |
| 2013 | Multizone soundfield reproduction using orthogonal basis expansionabstractWe introduce a method for 2-D spatial multizone soundfield reproduction based on describing the desired multizone soundfield as an orthogonal expansion of basis functions over the desired reproduction region. This approach finds the solution to the Helmholtz equation that is closest to the desired soundfield in a weighted least squares sense. The basis orthogonal set is formed using QR factorization with as input a suitable set of solutions of the Helmholtz equation. The coefficients of the Helmholtz solution wavefields can then be calculated, reducing the multizone sound reproduction problem to the reconstruction of a set of basis wavefields over the desired region. The method facilitates its application with a more practical loudspeaker configuration. The approach is shown effective for both accurately reproducing sound in the selected bright zone and minimizing sound leakage into the predefined quiet zone. Wenyu Jin 0002, W. Bastiaan Kleijn, David Virette |
ICASSP | 2 |
| 2013 | Speech coding based on pitch synchrony and two-stage transformationabstractIn this paper, an effective speech coder that is based on a sparse representation of speech by exploiting the strong dependencies between adjacent pitch cycles is proposed. In the proposed coder, a pitch-synchronous processing that consists of pitch warping and a two-stage transformation is used to achieve a compact representation of the voiced speech. Power spectral density preserving quantization (PSD-PQ) is adopted for quantizing the transform coefficients. The result is a coder that is efficient over a wide range of bit rates: it approaches perfect reconstruction with increasing rate, and has a parametric signal representation at low rates. Both objective PESQ results and subjective A/B listening tests show that the proposed coder outperforms the ITU-T G.722.1 codec. Changchun Bao, W. Bastiaan Kleijn |
ICASSP | 3 |
| 2013 | Reliable estimation of quality scores by a small calibrated listening panelabstractIn this work we propose the calibrated-mean-opinion-score (CMOS), which accounts for varying levels of precision and bias across subjects in a listening test. We adopt the Bayesian statistical framework, where the hyper-parameters of priors are learned via empirical Bayes, and the posterior is approximated by a computationally inexpensive variational technique. As our experimental results show, CMOS is more robust to noisy and biased subjects than MOS. As a result, CMOS can be used to improve the reliability of listening test results when a small test panel is used. To correct for the subjects in the test panel, calibration signals are required. Calibration signals are rated by a panel larger than the test panel. The key to saving human labor and cost is that only a few calibration signals are required, and that it is possible to share calibration signals across listening tests. Iman S. Mossavat, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2013 | Preservation of speech spectral dynamics enhances intelligibilityabstractSpeech is the most important communication modality for human interaction. Automatic speech recognition and speech synthesis have extended further the relevance of speech to man-machine interaction. Environment noise and various distortions, such as reverberation and speech processing artifacts, reduce the mutual information between the message modulated inthe clean speech and the message decoded from the observed signal. This degrades intelligibility and perceived quality, which are the two attributes associated with quality of service. An estimate of the state of these attributes provides important diagnostic information about the communication equipment and the environment. When the adverse effects occur at the presentation side, an objective measure of intelligibility facilitates speech signal modification for improved communication.The contributions of this thesis come from non-intrusive quality assessment and intelligibility-enhancing modification of speech. On the part of quality, the focus is on predictor design for limited training data. Paper A proposes a quality assessment model for bounded-support ratings that learns efficiently from a limited amount of training data, scales easily with the sampling frequency, and provides a platform for modeling variations in the individual subjective ratings. The predictive performance of the model for the mean of the subjective quality ratings compares favorably to the state-of-art in the field. Patterns in the spread of the individual ratings are captured in the feature space of the training data.Paper B focuses on enhancing predictive performance for the mean of the quality variable when the signal feature space is sparsely sampled by the training data. Using a Gaussian Processes framework, the deterministic signal-based feature set is augmented with a stochastic feature that is hypothesized to be jointly distributed with the target quality rating. An uncertainty propagation mechanism ensures that the variance of this feature is reflected in the prediction. The proposed architecture can take advantage of i) data that cannot be pooled due to subjective test protocol incompatibility and ii) models trained on data that are no longer available.With respect to intelligibility enhancement, a hierarchical perspective of the speech communication process, extended from foundational work in the field, is used in paper C to create a unified framework for method analysis and comparison. A high-level intelligibility measure related to the probability for correct recognition is derived using a hit-or-miss distortion criterion in the transcription domain. The measure is used to optimize two speech modifications at different levels of the message encoding hierarchy leading to significantly enhanced intelligibility in noise. The conceptual novelty of the method comes at the cost of higher complexity and the requirement for additional information including message transcription, sound segmentation, and a model of speech.Mapping the high-level measure to a lower level takes away the need for additional information and preserves asymptotically high-level optimality. Two methods are proposed to reduce degradation in the accuracy of the spectral dynamics due to additive noise. The focus of paper D is dynamics preservation in a range that is lower-bounded by an optimal band-power threshold. The performance of the method is competitive but allows for improvement in power efficiency. This issue is addressed in paper E which proposes and optimizes a distortion measure for spectral dynamics leading to a significant increase in intelligibility. Use of functional optimization techniques allows for families of solutions, among which are dynamic range compressors adaptive to the statistics of the speech and the noise. Petko Nikolov Petkov, W. Bastiaan Kleijn |
INTERSPEECH | 2 |
| 2013 | Rephrasing-based speech intelligibility enhancementabstractExisting algorithms for improving speech intelligibility in a noisy environment generally focus on modifying the acoustic features of live, recorded or synthesized speech while preserving the phonetic composition (the message). In this paper, we present an algorithm for text-to-speech systems that operates at a higher level of abstraction, the message-level. We use a paraphrasing system to adjust the linguistic content of the intended message such that the speech intelligibility improves under noisy conditions. To distinguish the intelligibility among paraphrases, we use the numerical integration of a normalized log-likelihood function over different signal-to-noise conditions. Objective evaluation results show that the developed measure is able to distinguish the intelligibility among paraphrases. Results from subjective evaluation confirm the effectiveness of our objective measure. Mengqiu Zhang, Petko Nikolov Petkov, W. Bastiaan Kleijn |
INTERSPEECH | 3 |
| 2013 | Perceptual Coding of High-Quality Digital AudioabstractThis paper introduces high-quality audio coding using psychoacoustic models. This technology is now abundant, with gadgets named after a standard (mp3 players) and the ability to play high-quality audio from literally billions of devices. The usual paradigm for these systems is based on filterbanks, followed by quantization and coding, controlled by a model of human hearing. The paper describes the basic technology, theoretical framework to apply to check for optimality, and the most prominent standards built on the basic ideas and newer work. Karlheinz Brandenburg, Christof Faller, Jürgen Herre, James D. Johnston, W. Bastiaan Kleijn |
Proc. IEEE | 5 |
| 2013 | Picking up the pieces: Causal states in noisy data, and how to recover them
Gustav Eje Henter, W. Bastiaan Kleijn |
Pattern Recognit. Lett. | 2 |
| 2013 | Vector quantization of LSF parameters with a mixture of dirichlet distributionsabstractQuantization of the linear predictive coding parameters is an important part in speech coding. Probability density function (PDF)-optimized vector quantization (VQ) has been previously shown to be more efficient than VQ based only on training data. For data with bounded support, some well-defined bounded-support distributions (e.g., the Dirichlet distribution) have been proven to outperform the conventional Gaussian mixture model (GMM), with the same number of free parameters required to describe the model. When exploiting both the boundary and the order properties of the line spectral frequency (LSF) parameters, the distribution of LSF differences LSF can be modelled with a Dirichlet mixture model (DMM). We propose a corresponding DMM based VQ. The elements in a Dirichlet vector variable are highly mutually correlated. Motivated by the Dirichlet vector variable's neutrality property, a practical non-linear transformation scheme for the Dirichlet vector variable can be obtained. Similar to the Karhunen-Loève transform for Gaussian variables, this non-linear transformation decomposes the Dirichlet vector variable into a set of independent beta-distributed variables. Using high rate quantization theory and by the entropy constraint, the optimal inter- and intra-component bit allocation strategies are proposed. In the implementation of scalar quantizers, we use the constrained-resolution coding to approximate the derived constrained-entropy coding. A practical coding scheme for DVQ is designed for the purpose of reducing the quantization error accumulation. The theoretical and practical quantization performance of DVQ is evaluated. Compared to the state-of-the-art GMM-based VQ and recently proposed beta mixture model (BMM) based VQ, DVQ performs better, with even fewer free parameters and lower computational cost Zhanyu Ma, Arne Leijon, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 3 |
| 2013 | Maximizing Phoneme Recognition Accuracy for Enhanced Speech Intelligibility in NoiseabstractAn effective measure of speech intelligibility is the probability of correct recognition of the transmitted message. We propose a speech pre-enhancement method based on matching the recognized text to the text of the original message. The selected criterion is accurately approximated by the probability of the correct transcription given an estimate of the noisy speech features. In the presence of environment noise, and with a decrease in the signal-to-noise ratio, speech intelligibility declines. We implement a speech pre-enhancement system that optimizes the proposed criterion for the parameters of two distinct speech modification strategies under an energy-preservation constraint. The proposed method requires prior knowledge in the form of a transcription of the transmitted message and acoustic speech models from an automatic speech recognition system. Performance results from an open-set subjective intelligibility test indicate a significant improvement over natural speech and a reference system that optimizes a perceptual-distortion-based objective intelligibility measure. The computational complexity of the approach permits use in on-line applications. Petko Nikolov Petkov, Gustav Eje Henter, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | Gaussian process dynamical models for nonparametric speech representation and synthesisabstractWe propose Gaussian process dynamical models (GPDMs) as a new, nonparametric paradigm in acoustic models of speech. These use multidimensional, continuous state-spaces to overcome familiar issues with discrete-state, HMM-based speech models. The added dimensions allow the state to represent and describe more than just temporal structure as systematic differences in mean, rather than as mere correlations in a residual (which dynamic features or AR-HMMs do). Being based on Gaussian processes, the models avoid restrictive parametric or linearity assumptions on signal structure. We outline GPDM theory, and describe model setup and initialization schemes relevant to speech applications. Experiments demonstrate subjectively better quality of synthesized speech than from comparable HMMs. In addition, there is evidence for unsupervised discovery of salient speech structure. Gustav Eje Henter, Marcus Frean, W. Bastiaan Kleijn |
ICASSP | 3 |
| 2012 | Transform-domain Wiener filter for speech periodicity enhancementabstractIn this paper, we present a transform-domain Wiener filtering approach for enhancing speech periodicity. The enhancement is performed on the linear prediction residual signal. Two sequential lapped frequency transforms are applied to the residual in a pitch-synchronous manner. The residual signal is effectively represented by two separate sets of transform coefficients that correspond to the periodic and aperiodic components, respectively. A Wiener filter operating on the transform coefficients is developed to restore periodicity and reduce noise. Different filter parameters are designed for the transform coefficients of the periodic and aperiodic components. A template-driven method is used to estimate the filter parameters for the periodic component. For the aperiodic components, the filter parameters are computed based on a local SNR for effective noise reduction. Experimental results confirm that the harmonic structure of the signal can be effectively restored with the proposed approach. Feng Huang 0002, Tan Lee, W. Bastiaan Kleijn |
ICASSP | 3 |
| 2012 | Audio codingwith power spectral density preserving quantizationabstractThe coding of audio-visual signals is generally based on different paradigms for high and low rates. At high rates the signal is approximated directly and at low rates only signal features are transmitted. The recently introduced distribution preserving quantization (DPQ) paradigm provides a seamless transition between these two regimes. In this paper we present a simplified scheme that preserves the power spectral density (PSD) rather than the probability distribution. In a practical system the PSD must be estimated. We show that both forward adaptive and backward adaptive PSD estimation are possible. Our experimental results confirm that preservation of PSD at finite precision leads to a unified coding paradigm that provides effective coding at both high and low rates. An audio coding application shows the perceptual benefits of PSD preserving quantization. Minyue Li, Janusz Klejsa, Alexey Ozerov, W. Bastiaan Kleijn |
ICASSP | 4 |
| 2012 | Enhancing Subjective Speech Intelligibility Using a Statistical Model of SpeechabstractThe intelligibility of speech in adverse noise conditions can be improved by modifying the characteristics of the clean speech prior to its presentation. An effective and flexible paradigm is to select the modification by optimizing a measure of objective intelligibility. Here we apply this paradigm at the text level and optimize a measure related to the classification error probability in an automatic speech recognition system. The proposed method was applied to a simple but powerful band-energy modification mechanism under an energy preservation constraint. Subjective evaluation results provide a clear indication of a significant gain in subjective intelligibility. In contrast to existing methods, the proposed approach is not restricted to a particular modification strategy and treats the notion of optimality at a level closer to that of subjective intelligibility. The computational complexity of the method is sufficiently low to enable its use in on-line applications. Petko Nikolov Petkov, W. Bastiaan Kleijn, Gustav Eje Henter |
INTERSPEECH | 2 |
| 2012 | A Hierarchical Bayesian Approach to Modeling Heterogeneity in Speech Quality AssessmentabstractThe development of objective speech quality measures generally involves fitting a model to subjective rating data. A typical data set comprises ratings generated by listening tests performed in different languages and across different laboratories. These factors as well as others, such as the sex and age of the talker, influence the subjective ratings and result in data heterogeneity. We use a linear hierarchical Bayes (HB) structure to account for heterogeneity. To make the structure effective, we develop a variational Bayesian inference for the linear HB structure that approximates not only the posterior over the model parameters, but also the model evidence. Using the approximate model evidence we are able to study and exploit the heterogeneity inducing factors in the Bayesian framework. The new approach yields a simple linear predictor with state-of-the-art predictive performance. Our experiments show that the new method compares favorably with systems based on more complex predictor structures such as ITU-T recommendation P.563, Bayesian MARS, and Gaussian processes. Iman S. Mossavat, Petko Nikolov Petkov, W. Bastiaan Kleijn, Oliver Amft |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | A Quantization Theoretic Perspective on Simulcast and Layered Multicast OptimizationabstractWe consider rate optimization in multicast systems that use several multicast trees on a communication network. The network is shared between different applications. For that reason, we model the available bandwidth for multicast as stochastic. For specific network topologies, we show that the multicast rate optimization problem is equivalent to the optimization of scalar quantization. We use results from rate-distortion theory to provide a bound on the achievable performance for the multicast rate optimization problem. A large number of receivers makes the possibility of adaptation to changing network conditions desirable in a practical system. To this end, we derive an analytical solution to the problem that is asymptotically optimal in the number of multicast trees. We derive local optimality conditions, which we use to describe a general class of iterative algorithms that give locally optimal solutions to the problem. Simulation results are provided for the multicast of an i.i.d. Gaussian process, an i.i.d. Laplacian process, and a video source. Ermin Kozica, W. Bastiaan Kleijn |
IEEE/ACM Trans. Netw. | 2 |
| 2011 | Bounding the Rate Region of the Two-Terminal Vector Gaussian CEO ProblemabstractThe rate region of the two-terminal vector Gaussian CEO problem is studied. A lower bound on the rate region is derived. It is obtained by lower-bounding a weighted sum rate for each supporting hyper plane of the rate region. The bound is in the form of a closed-form expression rather than the form of an optimization problem. Guoqiang Zhang 0003, W. Bastiaan Kleijn |
DCC | 2 |
| 2011 | Distributed blind source separation with an application to audio signalsabstractA scalable blind source separation paradigm aimed at sensor networks is described. The approach facilitates an unlimited number of sensors and sources and does not require a fusion centre. It is based on a so-called ownership principle, where each network node aims to extract (own) a source signal that is not already extracted (owned) by another network node. Nodes that own a source signal broadcast that signal to user nodes outside the network. Nodes that do not currently own a source signal do not transmit information and can be active intermittently. A natural application of the method is a distributed microphone network in a multi-talker environment, with as user nodes hearing aids or telephone interface devices. Such a network can stretch across buildings or neighbourhoods. Simulations using independent component analysis (ICA) indicate the validity of the principles of the method. Yusuke Hioka, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2011 | Quantization with an adjustable codeword length penaltyabstractQuantizers are generally designed either for fixed-rate coding or variable-rate coding. Under common conditions, variable-rate quantization leads to the same mean distortion for all quantization cells whereas fixed-rate is associated with a fixed bit allocation for each quantization operation. We show that both are special cases of the solution of a more general constraint that facilitates a practical trade-off between mean rate and rate outliers and mean distortion and outliers in distortion. The proposed constraint weights the code-word lengths exponentially. The resulting quantizers are suitable for commonly occurring scenarios such as those where a limited variability in rate can be tolerated by the network, those where the signal model may not be accurate, and those where source redundancy is used to counter channel errors. W. Bastiaan Kleijn, Moo Young Kim 0001 |
ICASSP | 1 |
| 2011 | Intermediate-State HMMs to Capture Continuously-Changing Signal FeaturesabstractTraditional discrete-state HMMs are not well suited for describing steadily evolving, path-following natural processes like motion capture data or speech. HMMs cannot represent incremental progress between behaviors, and sequences sampled from the models have unnatural segment durations, unsmooth transitions, and excessive rapid variation. We propose to address these problems by permitting the state variable to occupy positions between the discrete states, and present a concrete left-right model incorporating this idea. We call this intermediate-state HMMs. The state evolution remains Markovian. We describe training using the generalized EM-algorithm and present associated update formulas. An experiment shows that the intermediate-state model is capable of gradual transitions, with more natural durations and less noise in sampled sequences compared to a conventional HMM. Gustav Eje Henter, W. Bastiaan Kleijn |
INTERSPEECH | 2 |
| 2011 | Discrete Choice Models for Non-Intrusive Quality AssessmentabstractNon-intrusive signal quality assessment in general, and its application to speech signal processing, in particular, builds extensively upon statistical regression models. Commonly, the raw preference scores used for fitting these models belong to a categorical scale. Averaging the scores over a number of test subjects results in smooth, close-to-continuous ratings, thus justifying the use of regression as opposed to classification models. A form of marginalization, averaging subjective ratings takes away useful information about the reliability of individual test points. Using a model tailored to the raw data achieves highly competitive performance in terms of conventional performance measures while providing the additional advantage of identifying the usability of individual test points. In this paper, we consider the application of discrete choice models to non-intrusive quality assessment of speech. Petko Nikolov Petkov, W. Bastiaan Kleijn, Bert de Vries |
INTERSPEECH | 2 |
| 2011 | Auditory Model-Based Design and Optimization of Feature Vectors for Automatic Speech RecognitionabstractUsing spectral and spectro-temporal auditory models along with perturbation-based analysis, we develop a new framework to optimize a feature vector such that it emulates the behavior of the human auditory system. The optimization is carried out in an offline manner based on the conjecture that the local geometries of the feature vector domain and the perceptual auditory domain should be similar. Using this principle along with a static spectral auditory model, we modify and optimize the static spectral mel frequency cepstral coefficients (MFCCs) without considering any feedback from the speech recognition system. We then extend the work to include spectro-temporal auditory properties into designing a new dynamic spectro-temporal feature vector. Using a spectro-temporal auditory model, we design and optimize the dynamic feature vector to incorporate the behavior of human auditory response across time and frequency. We show that a significant improvement in automatic speech recognition (ASR) performance is obtained for any environmental condition, clean as well as noisy. Saikat Chatterjee, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Double-Ended Quality Assessment System for Super-Wideband SpeechabstractThis paper describes a double-ended quality assessment system for speech with a bandwidth of up to 14 kHz (so-called super-wideband speech). The quality assessment system is based on a combination of local and global features, where the local features are dependent on a time alignment procedure and the global features are not. The system is evaluated over a large set of subjectively scored narrowband, wideband and super-wideband speech databases. The system performs similarly to PESQ for narrowband speech and significantly better for wideband speech. L. Anders Ekman, Volodya Grancharov, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | Asymptotically Optimal Model Estimation for QuantizationabstractUsing high-rate theory approximations we introduce flexible practical quantizers based on possibly non-Gaussian models in both the constrained resolution (CR) and the constrained entropy cases. We derive model estimation criteria optimizing asymptotic (with increasing rate) quantizer performance. We show that in the CR case the optimal criterion is different from the maximum likelihood criterion commonly used for that purpose and introduce a new criterion that we call constrained resolution minimum description length (CR-MDL). We apply these principles to the generalized Gaussian scaled mixture model, which is accurate for many real-world signals. We provide an explanation of the reason why the CR-MDL improves quantization performance in the CR case and show that CR-MDL can compensate for a possible mismatch between model and data distribution. Thus, this criterion is of a great interest for practical applications. Our experiments apply the new quantization method to controllable artificial data and to the commonly used modulated lapped transform representation of audio signals. We show that both the CR-MDL criterion and a non-Gaussian modeling have significant advantages. Alexey Ozerov, W. Bastiaan Kleijn |
IEEE Trans. Commun. | 2 |
| 2011 | High-Rate Analysis of Symmetric L-Channel Multiple Description CodingabstractThis paper studies the tight rate-distortion bound for L-channel symmetric multiple-description coding of a scalar Gaussian source with two levels of receivers. Each of the first-level receivers obtains κ of the L descriptions (κ < L). The second-level receiver obtains all L descriptions. We find that if the central distortion (corresponding to the second-level receiver) is much smaller than the side distortion (corresponding to the first-level receivers), the product of a function of the side distortions and the central distortion is asymptotically independent of the redundancy between the descriptions. Using this property, we analyze the asymptotic behavior of a practical multiple-description lattice vector quantizer (MDLVQ). Our analysis includes the treatment of the MDLVQ system from a new geometric viewpoint, which results in an expression for the side distortions using the normalized second moment of a sphere of higher dimensionality than the quantization space. The expression of the distortion product derived from the lower bound is then applied as a criterion to assess the performance loss of the considered MDLVQ system. In principle, the efficiency of other practical MD systems can also be evaluated using the derived distortion product. Guoqiang Zhang 0003, Jan Østergaard, Janusz Klejsa, W. Bastiaan Kleijn |
IEEE Trans. Commun. | 4 |
| 2011 | Graph-Preserving Sparse Nonnegative Matrix Factorization With Application to Facial Expression RecognitionabstractIn this paper, a novel graph-preserving sparse nonnegative matrix factorization (GSNMF) algorithm is proposed for facial expression recognition. The GSNMF algorithm is derived from the original NMF algorithm by exploiting both sparse and graph-preserving properties. The latter may contain the class information of the samples. Therefore, GSNMF can be conducted as an unsupervised or a supervised dimension reduction method. A sparse representation of the facial images is obtained by minimizing the l(1)-norm of the basis images. Furthermore, according to the graph embedding theory, the neighborhood of the samples is preserved by retaining the graph structure in the mapped space. The GSNMF decomposition transforms the high-dimensional facial expression images into a locality-preserving subspace with sparse representation. To guarantee convergence, we use the projected gradient method to calculate the nonnegative solution of GSNMF. Experiments are conducted on the JAFFE database and the Cohn-Kanade database with unoccluded and partially occluded facial images. The results show that the GSNMF algorithm provides better facial representations and achieves higher recognition rates than nonnegative matrix factorization. Moreover, GSNMF is also more robust to partial occlusions than other tested methods. Ruicong Zhi, Markus Flierl, Qiuqi Ruan, W. Bastiaan Kleijn |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2010 | Bounding the Rate Region of Vector Gaussian Multiple Descriptions with Individual and Central ReceiversabstractThe problem of the rate region of the vector Gaussian multiple description with individual and central quadratic distortion constraints is studied. We have two main contributions. First, a lower bound on the rate region is derived. The bound is obtained by lower-bounding a weighted sum rate for each supporting hyperplane of the rate region. Second, the rate region for the scenario of the scalar Gaussian source is fully characterized by showing that the lower bound is tight. The optimal weighted sum rate for each supporting hyperplane is obtained by solving a single maximization problem. This is contrary to existing results, which require solving a min-max optimization problem. Guoqiang Zhang 0003, W. Bastiaan Kleijn, Jan Østergaard |
DCC | 2 |
| 2010 | Auditory model based modified MFCC featuresabstractUsing spectral and spectro-temporal auditory models, we develop a computationally simple feature vector based on the design architecture of existing mel frequency cepstral coefficients (MFCCs). Along with the use of an optimized static function to compress a set of filter bank energies, we propose to use a memory-based adaptive compression function to incorporate the behavior of human auditory response across time and frequency. We show that a significant improvement in automatic speech recognition (ASR) performance is obtained for any environmental condition, clean as well as noisy. Saikat Chatterjee, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2010 | Flexcode - flexible audio codingabstractModern networks are highly variable and, as a result, source coders are commonly used under conditions that they were not designed for. We address this problem with a source-coding philosophy that aims at the instantaneous re-optimization of a source coder to match a wide range of constraints on rate or quality and a wide range of packet-loss rates. We present a number of technologies that can be reconfigured by solving analytic relations that use the current conditions and a statistical description of the source as input. The technologies include distribution-preserving quantizers, flexible multiple-description quantizers, and a rate distribution scheme. Based on the generic technologies, we created a complete audio coder. Formal listening tests show that the resulting audio coding scheme with full flexibility provides a quality that is on-par with the best standardized codecs for any particular rate. Janusz Klejsa, Minyue Li, W. Bastiaan Kleijn |
ICASSP | 3 |
| 2010 | Selecting static and dynamic features using an advanced auditory model for speech recognitionabstractWe describe a method to select features for speech recognition that is based on a quantitative model of the human auditory periphery. The method maximizes the similarity of the geometry of the space spanned by the subset of features and the geometry of the space spanned by the auditory model output. The selection method uses a spectro-temporal auditory model that captures both frequency- and time-domain masking. The selection method is blind to the meaning of speech and does not require annotated speech data. We apply the method to the selection of a subset of features from a conventional set consisting of mel cepstra and their first-order and second-order time derivatives. Although our method uses only knowledge of the human auditory periphery, the experimental results show that it performs significantly better than feature-reduction algorithms based on linear and heteroscedastic discriminant analysis that require training with annotated speech data. Christos Koniaris, Saikat Chatterjee, W. Bastiaan Kleijn |
ICASSP | 3 |
| 2010 | Objective quality estimation of wide-band speech using a narrow-band priorabstractA fundamental challenge in the design of objective models for estimation of speech signal quality lies in the shortage of subjectively labelled databases. This problem is particularly relevant when developing quality assessment models for wide-band (16 kHz sampling rate) signals where databases are scarce. We explore the possibility for seamlessly integrating a quality prior in the form of a narrow-band quality estimate into the framework of a non-intrusive wide-band quality assessment algorithm. Experimental results confirm that the proposed approach can be used to improve performance over a baseline wide-band system without a narrow-band prior. Petko Nikolov Petkov, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2010 | Complexity-outsourced low-latency video encoding through feedback under a sum-rate constraintabstractIn live video communication, the quality of the video that is encoded on a mobile device is more often than not constrained by the available computational resources. We address this problem by allowing the encoder to “outsource” the encoding complexity to the decoder through the use of feedback, under a sum-rate constraint comprising the weighted sum of the forward and feedback rates. Analysis of such a complexity-outsourced framework using an analytically tractable video model reveals that the feedback rate should be optimally adapted to the video motion characteristics. Application of our framework to real-world video sequences reveals the efficacy of our proposed architecture, with experimental results validating that substantial gains in PSNR are achievable over state-of-the-art fast-search algorithms at a comparable level of complexity. These gains are significant even for mobile live video communication applications over the Internet. Ermin Kozica, Kannan Ramchandran, W. Bastiaan Kleijn |
ICIP | 3 |
| 2010 | Learning from images and speech with Non-negative Matrix Factorization enhanced by input space scalingabstractComputional learning from multimodal data is often done with matrix factorization techniques such as NMF (Non-negative Matrix Factorization), pLSA (Probabilistic Latent Semantic Analysis) or LDA (Latent Dirichlet Allocation). The different modalities of the input are to this end converted into features that are easily placed in a vectorized format. An inherent weakness of such a data representation is that only a subset of these data features actually aids the learning. In this paper, we first describe a simple NMF-based recognition framework operating on speech and image data. We then propose and demonstrate a novel algorithm that scales the inputs of this framework in order to optimize its recognition performance. Joris Driesen, Hugo Van hamme, W. Bastiaan Kleijn |
SLT | 3 |
| 2010 | The synergy between bounded-distance HMM and spectral subtraction for robust speech recognition
Jesús Vicente-Peña, Fernando Díaz-de-María, W. Bastiaan Kleijn |
Speech Commun. | 3 |
| 2010 | Distribution Preserving Quantization With Dithering and TransformationabstractA new quantization scheme that preserves the probability distribution of the source signal is presented. The distribution preserving quantization (DPQ) achieves the optimal trade-off between mean square error and bit rate asymptotically. It provides a continuum ranging from rate-distortion optimal signal quantization to parametric coding. The method can be used as a core component for scalable coding. Its efficacy is illustrated by applying the scheme to audio coding. Minyue Li, Janusz Klejsa, W. Bastiaan Kleijn |
IEEE Signal Process. Lett. | 3 |
| 2010 | Reduction of the Impact of Distortion Outliers and Source Mismatch in Resolution-Constrained QuantizationabstractThe rate-distortion performance of conventional resolution-constrained quantization (RCQ) based on the mean-squared error criterion (MSE-RCQ) is generally compromised by the impact of distortion outliers and source mismatch. Not only the mean distortion, but also the number of distortion outliers should be considered in quantizer design. Thus, we propose the use of a design criterion that gives more importance to the tail of the source distribution, which leads to RCQ based on the second moment of distortion (SMD-RCQ). A continuous range of alternatives between MSE-RCQ and SMD-RCQ is also defined and implemented based on the weighted arithmetic-mean measure (WAM-RCQ). It can be used to control the centroid density in the tail of the source distribution. Experimental results with a Gaussian source and line spectral frequencies (LSFs) show that the proposed WAM-RCQ not only produces a similar mean distortion as conventional MSE-RCQ, but has a lower percentage of distortion outliers and a significantly reduced sensitivity to source mismatch. Moo Young Kim 0001, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Analysis of K-Channel Multiple Description QuantizationabstractThis paper studies the tight rate-distortion bound for K-channel symmetric multiple-description coding for a memory less Gaussian source. We find that the product of a function of the individual side distortions (for single received descriptions) and the central distortion (for K received descriptions) is asymptotically independent of the redundancy among the descriptions. Using this property, we analyze the asymptotic behaviors of two different practical multiple-description lattice vector quantizers (MDLVQ). Our analysis includes the treatment of a MDLVQ system from a new geometric viewpoint, which results in an expression for the side distortions using the normalized second moment of a sphere of higher dimensionality than the quantization space. The expression of the distortion product derived from the lower bound is then applied as a criterion to assess the performance losses of the considered MDLVQ systems. Guoqiang Zhang 0003, Janusz Klejsa, W. Bastiaan Kleijn |
DCC | 3 |
| 2009 | Rate distribution between model and signal for multiple descriptionsabstractWe consider the rate allocation problem for multiple-description quantization of the signal described by an adaptive model with a fixed structure. The source modeling in coding generally results in a two-stage description of the data, where one of the stages describes the model parameters, and the other describes the signal. Such a setup implies the existence of a trade-off between the rate spent on the parameters and the rate spent on the signal. We optimize this trade-off analytically for the multiple-description case using a method inspired by Minimum Description Length principle. We also provide an algorithm for optimizing the rate allocation between the components of the model-based multiple description coder. Finally we experimentally confirm our results. Our method facilitates the rate-adaptive multiple-description coding. Janusz Klejsa, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2009 | Joint optimization of the redundancy of multiple-description coders for multicastabstractWe consider the optimization of multicast over packet-switched communication networks with a non-zero packet-loss probability. For the system setup consisting of a number of multiple-description coders, we jointly optimize these coders. We propose an analytic so-lution, asymptotically optimal in the number of multiple-description coders. The analytic solution allows for fast system adaptation to changing network conditions. A locally optimal optimization algo-rithm that is useful when the number of multicast groups is small is derived. Simulations show that the utilization of the analytic solution incurs a low overhead on the performance when compared to the locally optimal solution, even for a small number of multiple-description coders. Index Terms — Multicast, multiple-description coding, opti-mization, packet-loss Ermin Kozica, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2009 | Optimal parameter estimation for model-based quantizationabstractWe address optimal model estimation for model-based vector quantization for both the constrained resolution (CR) and constrained entropy (CE) cases. To this purpose we derive under high-rate (HR) theory assumptions the rate-distortion (RD) relations for these two quantization scenarios assuming a Gaussian model. Based on the RD relations we show that the maximum likelihood (ML) criterion leads to optimal performance for CE quantization, but not for CR quantization. We introduce a new model estimation criterion for CR quantization that is optimal (under HR theory assumptions) in terms of the RD relation. Our experiments confirm that the proposed criterion for model identification outperforms the ML criterion for a range of conditions. Alexey Ozerov, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2009 | Compressive sensing for sparsely excited speech signalsabstractCompressive sensing (CS) has been proposed for signals with sparsity in a linear transform domain. We explore a signal dependent unknown linear transform, namely the impulse response matrix operating on a sparse excitation, as in the linear model of speech production, for recovering compressive sensed speech. Since the linear transform is signal dependent and unknown, unlike the standard CS formulation, a codebook of transfer functions is proposed in a matching pursuit (MP) framework for CS recovery. It is found that MP is efficient and effective to recover CS encoded speech as well as jointly estimate the linear model. Moderate number of CS measurements and low order sparsity estimate will result in MP converge to the same linear transform as direct VQ of the LP vector derived from the original signal. There is also high positive correlation between signal domain approximation and CS measurement domain approximation for a large variety of speech spectra. Thippur V. Sreenivas, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2009 | Facial expression recognition based on graph-preserving sparse non-negative matrix factorizationabstractIn this paper, we present a novel algorithm for representing facial expressions. The algorithm is based on the non-negative matrix factorization (NMF) algorithm, which decomposes the original facial image matrix into two non-negative matrices, namely the coefficient matrix and the basis image matrix. We call the novel algorithm graph-preserving sparse non-negative matrix factorization (GSNMF). GSNMF utilizes both sparse and graph-preserving constraints to achieve a non-negative factorization. The graph-preserving criterion preserves the structure of the original facial images in the embedded subspace while considering the class information of the facial images. Therefore, GSNMF has more discriminant power than NMF. GSNMF is applied to facial images for the recognition of six basic facial expressions. Our experiments show that GSNMF achieves on average a recognition rate of 93.5% compared to that of discriminant NMF with 91.6%. Ruicong Zhi, Markus Flierl, Qiuqi Ruan, W. Bastiaan Kleijn |
ICIP | 4 |
| 2009 | Auditory model based optimization of MFCCs improves automatic speech recognition performanceabstractUsing a spectral auditory model along with perturbation based analysis, we develop a new framework to optimize a set of fea-tures such that it emulates the behavior of the human auditory sys-tem. The optimization is carried out in an off-line manner based on the conjecture that the local geometries of the feature domain and the perceptual auditory domain should be similar. Using this principle, we modify and optimize the static mel frequency cep-stral coefficients (MFCCs) without considering any feedback from the speech recognition system. We show that improved recognition performance is obtained for any environmental condition, clean as well as noisy. Index Terms: MFCC, auditory model, ASR. 1. Saikat Chatterjee, Christos Koniaris, W. Bastiaan Kleijn |
INTERSPEECH | 3 |
| 2009 | A Bayesian approach to non-intrusive quality assessment of speechabstractA Bayesian approach to non-intrusive quality assessment of narrow-band speech is presented. The speech features used to assess quality are the sample mean and variance of band-powers evaluated from the temporal envelope in the channels of an auditory filter-bank. Bayesian multivariate adaptive regression splines (BMARS) is used to map features into quality ratings. The proposed combination of features and regression method leads to a high performance quality assessment algorithm that learns efficiently from a small amount of training data and avoids overfitting. Use of the Bayesian approach also allows the derivation of credible intervals on the model predictions, which provide a quantitative measure of model confidence and can be used to identify the need for complementing the training databases. Petko Nikolov Petkov, Iman S. Mossavat, W. Bastiaan Kleijn |
INTERSPEECH | 3 |
| 2009 | Speech Watermarking for Analog Flat-Fading Bandpass ChannelsabstractWe present a blind speech watermarking algorithm that embeds the watermark data in the phase of non-voiced speech by replacing the excitation signal of an autoregressive speech signal representation. The watermark signal is embedded in a frequency subband, which facilitates robustness against bandpass filtering channels. We derive several sets of pulse shapes that prevent intersymbol interference and that allow the passband watermark signal to be created by simple filtering. A marker-based synchronization scheme robustly detects the location of the embedded watermark data without the occurrence of insertions or deletions. In light of a potential application to analog aeronautical voice radio communication, we present experimental results for embedding a watermark in narrowband speech at a bit-rate of 450 bit/s. The recursive least-squares (RLS) equalization-based watermark detector not only compensates for the vocal tract filtering, but also recovers the watermark data in the presence of nonlinear phase and bandpass filtering, amplitude modulation, and additive white Gaussian noise (AWGN), making the watermarking scheme highly robust. Konrad Hofbauer, Gernot Kubin, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 3 |
| 2009 | Scalable coding with side information for packet loss recoveryabstractThis paper presents a packet loss recovery method that uses an incomplete secondary encoding based on scalar quantization as redundancy. The method is redundancy bit rate scalable and allows an adaptation to varying loss scenarios and a varying packeting strategy. The recovery is performed by minimum mean squared error estimation incorporating a statistical model for the quantizers to facilitate real-time adaptation. A bit allocation algorithm is proposed that extends 'reverse water filling' to the problem of scalar encoding dependent variables for a decoder with a final estimation stage and available side information. We apply the method to the encoding of line-spectral frequencies (LSFs), which are commonly used in speech coding, illustrating the good performance of the method. Christian Feldbauer, W. Bastiaan Kleijn |
IEEE Trans. Commun. | 2 |
| 2009 | Feature Selection Under a Complexity ConstraintabstractClassification on mobile devices is often done in an uninterrupted fashion. This requires algorithms with gentle demands on the computational complexity. The performance of a classifier depends heavily on the set of features used as input variables. Existing feature selection strategies for classification aim at finding a ldquobestrdquo set of features that performs well in terms of classification accuracy, but are not designed to handle constraints on the computational complexity. We demonstrate that an extension of the performance measures used in state-of-the-art feature selection algorithms with a penalty on the feature extraction complexity leads to superior feature sets if the allowed computational complexity is limited. Our solution is independent of a particular classification algorithm. Jan H. Plasberg, W. Bastiaan Kleijn |
IEEE Trans. Multim. | 2 |
| 2008 | Adaptive resolution-constrained scalar multiple-description codingabstractWe consider adaptive two-channel multiple-description coding. We provide an analytical method for designing a resolution-constrained symmetrical multiple-description coder that uses an index assignment matrix. We use existing index assignment algorithms that are known for their good properties within our adaptive multiple-description coding architecture. These existing index assignment algorithms are parameterized and the coefficient of quantization of the side coders is described by a rational function. This leads to an analytical solution for the design problem, facilitating real-time adaptation based on information provided by feedback channels or rate conditions imposed by network management. Our experimental results show that the practical performance closely approximates theoretically obtained optimal behavior. Janusz Klejsa, Marcin Kuropatwinski, W. Bastiaan Kleijn |
ICASSP | 3 |
| 2008 | Autoregressive model-based speech packet-loss concealmentabstractWe study packet-loss concealment for speech based on autoregressive modeling using a rigorous minimum mean square error (MMSE) approach. The effect of the model estimation error on predicting the missing segment is studied and an upper bound on the mean square error is derived. Our experiments show that the upper bound is tight when the estimation error is less than the signal variance. We also consider the usage of perceptual weighting on prediction to improve speech quality. A rigorous argument is presented to show that perceptual weighting is not useful in this context. We create simple and practical MMSE-based systems using two signal models: a basic model capturing the short-term correlation and a more sophisticated model that also captures the long-term correlation. Subjective quality comparison tests show that the proposed MMSE-based system provides state-of-the-art performance. Guoqiang Zhang 0003, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2008 | Regularized Linear Prediction of SpeechabstractAll-pole spectral envelope estimates based on linear prediction (LP) for speech signals often exhibit unnaturally sharp peaks, especially for high-pitch speakers. In this paper, regularization is used to penalize rapid changes in the spectral envelope, which improves the spectral envelope estimate. Based on extensive experimental evidence, we conclude that regularized linear prediction outperforms bandwidth-expanded linear prediction. The regularization approach gives lower spectral distortion on average, and fewer outliers, while maintaining a very low computational complexity. L. Anders Ekman, W. Bastiaan Kleijn, Manohar N. Murthi |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | Generalized Postfilter for Speech Quality EnhancementabstractPostfilters are commonly used in speech coding for the attenuation of quantization noise. In the presence of acoustic background noise or distortion due to tandeming operations, the postfilter parameters are not adjusted and the performance is, therefore, not optimal. We propose a modification that consists of replacing the nonadaptive postfilter parameters with parameters that adapt to variations in spectral flatness, obtained from the noisy speech. This generalization of the postfiltering concept can handle a larger range of noise conditions, but has the same computational complexity and memory requirements as the conventional postfilter. Test results indicate that the presented algorithm improves on the standard postfilter, as well as on the combination of a noise attenuation preprocessor and the conventional postfilter. Volodya Grancharov, Jan H. Plasberg, Jonas Samuelsson, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 4 |
| 2008 | Online Noise Estimation Using Stochastic-Gain HMM for Speech EnhancementabstractWe propose a noise estimation algorithm for single-channel noise suppression in dynamic noisy environments. A stochastic-gain hidden Markov model (SG-HMM) is used to model the statistics of nonstationary noise with time-varying energy. The noise model is adaptive and the model parameters are estimated online from noisy observations using a recursive estimation algorithm. The parameter estimation is derived for the maximum-likelihood criterion and the algorithm is based on the recursive expectation maximization (EM) framework. The proposed method facilitates continuous adaptation to changes of both noise spectral shapes and noise energy levels, e.g., due to movement of the noise source. Using the estimated noise model, we also develop an estimator of the noise power spectral density (PSD) based on recursive averaging of estimated noise sample spectra. We demonstrate that the proposed scheme achieves more accurate estimates of the noise model and noise PSD, and as part of a speech enhancement system facilitates a lower level of residual noise. David Yuheng Zhao, W. Bastiaan Kleijn, Alexander Ypma, Bert de Vries |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | An Adaptive, Scalable Packet Loss Recovery MethodabstractWe propose a packet loss recovery method that uses an incomplete secondary encoding as redundancy. The recovery is performed by minimum mean squared error estimation. The method adapts to the loss scenario and is rate scalable. It incorporates a statistical model for the quantizers to facilitate real-time adaptation. We apply the method to the encoding of line-spectral frequencies, which are commonly used in speech coding, illustrating the good performance of the method. Christian Feldbauer, W. Bastiaan Kleijn |
ICASSP (4) | 2 |
| 2007 | A Canonical Representation of SpeechabstractIt is well known that usage of an appropriate representation of the speech signal improves the performance of speech coders, recognizers, and synthesizers. In this paper we present a representation of speech that has the efficiency, in terms of being compact, similar to that of parametric modeling, but additionally has the completeness property of signal expansions. The resulting canonical representation of speech is suited for a wide range of speech processing applications and we demonstrate this through experiments related to coding and prosodic modification. Mattias Nilsson 0002, Barbara Resch, Moo Young Kim 0001, W. Bastiaan Kleijn |
ICASSP (4) | 4 |
| 2007 | Interlacing Intraframes in Multiple-Description Video CodingabstractWe introduce a method to improve performance of multiple-description coding based on legacy video coders with pre-and postprocessing. The pre-and post-processing setup is general, making the method applicable to most legacy coders. For the case of two coders, a relative displacement of the intra-coding mode between the coders is shown to give improved robustness to packet loss. The optimal displacement of the intra-coding mode is found analytically, using a distortion minimization formulation where two independent Gilbert channels are assumed. The analytical results are confirmed by simulations. Tests with an H.263 coder show significant improvement in YPSNR over equivalent systems with no relative displacement of the intra-coding operation. Ermin Kozica, Dave Zachariah, W. Bastiaan Kleijn |
ICIP (4) | 3 |
| 2007 | Noise suppression based on extending a speech-dominated modulation bandabstractPrevious work on bandpass modulation filtering for noise suppression has resulted in unwanted perceptual artifacts and decreased speech clarity. Artifacts are introduced mainly due to half-wave rectification, which is employed to correct for negative power spectral values resultant from the filtering process. In this paper, modulation frequency estimation (i.e., bandwidth extension) is used to improve perceptual quality. Experiments demonstrate that speech-component lowpass modulation content can be reliably estimated from bandpass modulation content of speech-plus-noise components. Subjective listening tests corroborate that improved quality is attained when the removed speech lowpass modulation content is compensated for by the estimate. Tiago H. Falk, Svante Stadler, W. Bastiaan Kleijn, Wai-Yip Chan |
INTERSPEECH | 3 |
| 2007 | Mutual information and the speech signalabstractMutual information is commonly used in speech processing in the context of statistical mapping. Examples are the optimization of speech or speaker recognition algorithms, the computation of performance bounds on such algorithms, and bandwidth extension of narrow-band speech signals. It is generally ignored that speech-signal derived data usually have an intrinsic dimensionality that is lower than the dimensionality of the observation vectors (the dimensionality of the embedding space). In this paper, we show that such reduced dimensionality can affect the accuracy of the mutual information estimate significantly. We introduce a new method that removes the effects of singular probability density functions. The method does not require prior knowledge of the intrinsic dimensionality of the data. It is shown that the method is appropriate for speech-derived data. Mattias Nilsson 0002, W. Bastiaan Kleijn |
INTERSPEECH | 2 |
| 2007 | Improving the phase vocoder approach to pitch-shiftingabstractA class of methods known as phase vocoders allows for implementing pitch shifting in the spectral domain. We extend the approach of shifting the isolated harmonies of the spectrum by introducing a ... Petko Nikolov Petkov, W. Bastiaan Kleijn |
INTERSPEECH | 2 |
| 2007 | The Sensitivity Matrix: Using Advanced Auditory Models in Speech and Audio ProcessingabstractPerceptually optimal processing of speech and audio signals demands a rigorous approach using a distortion measure that resembles human perception. This requires distortion measures based on sophisticated, complex auditory models. Under the assumption of small distortions these models can be simplified by means of a sensitivity matrix. In this paper, we show the power of this approach. We present a method to derive the sensitivity matrix for distortion measures based on spectro-temporal auditory models. This method is applied to an example auditory model and the region of validity of the approximation and the application of linear algebra to analyze the characteristics of the given model are discussed. Furthermore, we show how to build a coder minimizing a sensitivity matrix distortion measure given the typically long support of a perceptual distortion measure Jan H. Plasberg, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Estimation of the Instantaneous Pitch of SpeechabstractAn accurate estimation of the pitch is essential for many speech processing applications, such as speech synthesis, speech coding, and speech enhancement. A widely used assumption in most common pitch estimation methods is that pitch is constant over a segment of short duration. This assumption does not apply in reality and leads to inaccurate pitch estimates. In this paper, we present a method for continuous pitch estimation that is able to track fast changes. In the presented framework, the pitch is modeled by a B-spline expansion and optimized in a multistage procedure for increased robustness. The performance of the continuous optimization procedure is compared to state-of-the-art pitch estimation methods and is evaluated both for artificial speech-like signals with known pitch, and for real speech signals. The results of the experiments show that our method leads to a higher accuracy of the estimate of the pitch than state-of-the-art methods Barbara Resch, Mattias Nilsson 0002, L. Anders Ekman, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 4 |
| 2007 | Codebook-Based Bayesian Speech Enhancement for Nonstationary EnvironmentsabstractIn this paper, we propose a Bayesian minimum mean squared error approach for the joint estimation of the short-term predictor parameters of speech and noise, from the noisy observation. We use trained codebooks of speech and noise linear predictive coefficients to model the a priori information required by the Bayesian scheme. In contrast to current Bayesian estimation approaches that consider the excitation variances as part of the a priori information, in the proposed method they are computed online for each short-time segment, based on the observation at hand. Consequently, the method performs well in nonstationary noise conditions. The resulting estimates of the speech and noise spectra can be used in a Wiener filter or any state-of-the-art speech enhancement system. We develop both memoryless (using information from the current frame alone) and memory-based (using information from the current and previous frames) estimators. Estimation of functions of the short-term predictor parameters is also addressed, in particular one that leads to the minimum mean squared error estimate of the clean speech signal. Experiments indicate that the scheme proposed in this paper performs significantly better than competing methods Sriram Srinivasan 0003, Jonas Samuelsson, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 3 |
| 2007 | HMM-Based Gain Modeling for Enhancement of Speech in NoiseabstractAccurate modeling and estimation of speech and noise gains facilitate good performance of speech enhancement methods using data-driven prior models. In this paper, we propose a hidden Markov model (HMM)-based speech enhancement method using explicit gain modeling. Through the introduction of stochastic gain variables, energy variation in both speech and noise is explicitly modeled in a unified framework. The speech gain models the energy variations of the speech phones, typically due to differences in pronunciation and/or different vocalizations of individual speakers. The noise gain helps to improve the tracking of the time-varying energy of nonstationary noise. The expectation-maximization (EM) algorithm is used to perform offline estimation of the time-invariant model parameters. The time-varying model parameters are estimated online using the recursive EM algorithm. The proposed gain modeling techniques are applied to a novel Bayesian speech estimator, and the performance of the proposed enhancement method is evaluated through objective and subjective tests. The experimental results confirm the advantage of explicit gain modeling, particularly for nonstationary noise sources David Yuheng Zhao, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | On the Estimation of Differential Entropy From Data Located on Embedded ManifoldsabstractEstimation of the differential entropy from observations of a random variable is of great importance for a wide range of signal processing applications such as source coding, pattern recognition, hypothesis testing, and blind source separation. In this paper, we present a method for estimation of the Shannon differential entropy that accounts for embedded manifolds. The method is based on high-rate quantization theory and forms an extension of the classical nearest-neighbor entropy estimator. The estimator is consistent in the mean square sense and an upper bound on the rate of convergence of the estimator is given. Because of the close connection between compression and Shannon entropy, the proposed method has an advantage over methods estimating the Renyi entropy. Through experiments on uniformly distributed data on known manifolds and real-world speech data we show the accuracy and usefulness of our proposed method. Mattias Nilsson 0002, W. Bastiaan Kleijn |
IEEE Trans. Inf. Theory | 2 |
| 2006 | Spectral Envelope Estimation and RegularizationabstractA well-known problem with linear prediction is that its estimate of the spectral envelope often has sharp peaks for high-pitch speakers. These peaks are anomalies resulting from contamination of the spectral envelope by the spectral fine structure. We investigate the method of regularized linear prediction to find a better estimate of the spectral envelope and compare the method to the commonly used approach of bandwidth expansion. We present simulations over voiced frames of female speakers from the TIMIT database, where the envelope modeling accuracy is measured using a log spectral distortion measure. We also investigate the coding properties of the methods. The results indicate that the new regularized LP method is superior to bandwidth expansion, with an insignificant increase in computational complexity L. Anders Ekman, W. Bastiaan Kleijn, Manohar N. Murthi |
ICASSP (1) | 2 |
| 2006 | Sub-Pixel Registration of Noisy ImagesabstractThe accurate registration of images observed in additive noise is a challenging task. The noise increases the number of misregistered regions, and decreases the accuracy of subpixel registration. To address this problem, we propose an intensity-based algorithm that performs registration based only on regions that are least affected by noise. We select these regions with a signal-to-noise ratio estimate that is obtained from an initial, less-accurate registration. Our simulations demonstrate that the proposed noise-adaptive scheme significantly outperforms the conventional registration approach. Volodya Grancharov, W. Bastiaan Kleijn, Alexander Georgiev |
ICASSP (2) | 2 |
| 2006 | HMM-Based Speech Enhancement using Explicit Gain ModelingabstractWe propose a hidden Markov model (HMM) based speech enhancement method using explicit modeling of speech and noise gains. The gains are considered to be stochastic variables in an HMM framework. The speech gain models the energy variations of speech phones, typically due to differences in pronunciation and/or different vocalizations of individual speakers. The noise gain helps to improve the tracking of the time-varying energy of non-stationary noise. The time-varying parameters of the gain models are estimated on-line using the recursive expectation maximization (EM) algorithm. The performance of the proposed enhancement system is evaluated through both objective and subjective tests. The experimental results confirm the advantage of explicit gain modeling, particularly for non-stationary noise sources David Yuheng Zhao, W. Bastiaan Kleijn |
ICASSP (1) | 2 |
| 2006 | Non-intrusive speech quality assessment with low computational complexityabstractWe describe an algorithm for monitoring subjective speech quality without access to the original signal that has very low computational and memory requirements. The features used in the proposed algorithm can be computed from commonly used speechcoding parameters. Reconstruction and perceptual transformation of the signal are not performed. The algorithm generates quality assessment ratings without explicit distortion modeling. The simulation results indicate that the proposed non-intrusive objective quality measure performs better than the ITU-T P.563 standard despite its very low computational complexity. Index Terms: non-intrusive quality assessment, quality of service. 1. Volodya Grancharov, David Yuheng Zhao, Jonas Lindblom, W. Bastiaan Kleijn |
INTERSPEECH | 4 |
| 2006 | Individual on-line variance adaptation of frequency filtered parameters for robust ASRabstractIn this paper we address the problem of robust speech recognition. We propose a new method based on the individual variance adaptation of frequency filtered parameters to reduce the deleterious effects of additive narrow-band noise. The method can be interpreted as a spectral weighting that assigns increased importance to the most reliable spectral components, typically the spectral peaks. The experiments confirm that the suggested method results in significantly improved recognition rates for additive narrow-band noise. Jesús Vicente-Peña, Fernando Díaz-de-María, W. Bastiaan Kleijn |
INTERSPEECH | 3 |
| 2006 | Resolution-Constrained Quantization With JND-Based Perceptual-Distortion MeasuresabstractWhen the squared error of observable signal parameters is below the just noticeable difference (JND), it is not registered by human perception. We modify commonly used distortion criteria to account for this phenomenon and study the implications for quantizer design and performance. Taking the JND into account in the design of the quantizer generally leads to improved performance in terms of mean distortion and the number of outliers. Moreover, the resulting quantizer exhibits better robustness against source mismatch Moo Young Kim 0001, W. Bastiaan Kleijn |
IEEE Signal Process. Lett. | 2 |
| 2006 | Multichannel parametric speech enhancementabstractWe present a parametric model-based multichannel approach for speech enhancement. By employing an autoregressive model for the speech signal and using a trained codebook of speech linear predictive coefficients, minimum mean square error estimation of the speech signal is performed. By explicitly accounting for steering errors in the signal model, robust estimates are obtained. Experiments show that the proposed method results in significant performance gains. Sriram Srinivasan 0003, Robert Aichner, W. Bastiaan Kleijn, Walter Kellermann |
IEEE Signal Process. Lett. | 3 |
| 2006 | On causal algorithms for speech enhancementabstractKalman filtering is a powerful technique for the estimation of a signal observed in noise that can be used to enhance speech observed in the presence of acoustic background noise. In a speech communication system, the speech signal is typically buffered for a period of 10-40 ms and, therefore, the use of either a causal or a noncausal filter is possible. We show that the causal Kalman algorithm is in conflict with the basic properties of human perception and address the problem of improving its perceptual quality. We discuss two approaches to improve perceptual performance. The first is based on a new method that combines the causal Kalman algorithm with pre- and postfiltering to introduce perceptual shaping of the residual noise. The second is based on the conventional Kalman smoother. We show that a short lag removes the conflict resulting from the causality constraint and we quantify the minimum lag required for this purpose. The results of our objective and subjective evaluations confirm that both approaches significantly outperform the conventional causal implementation. Of the two approaches, the Kalman smoother performs better if the signal statistics are precisely known, if this is not the case the perceptually weighted Kalman filter performs better. Volodya Grancharov, Jonas Samuelsson, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 3 |
| 2006 | Low-Complexity, Nonintrusive Speech Quality AssessmentabstractMonitoring of speech quality in emerging heterogeneous networks is of great interest to network operators. The most efficient way to satisfy such a need is through nonintrusive, objective speech quality assessment. In this paper, we describe a low-complexity algorithm for monitoring the speech quality over a network. The features used in the proposed algorithm can be computed from commonly used speech-coding parameters. Reconstruction and perceptual transformation of the signal is not performed. The critical advantage of the approach lies in generating quality assessment ratings without explicit distortion modeling. The results from the performed experiments indicate that the proposed nonintrusive objective quality measure performs better than the ITU-T P.563 standard Volodya Grancharov, David Yuheng Zhao, Jonas Lindblom, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 4 |
| 2006 | Estimation of the short-term predictor parameters of speech under noisy conditionsabstractSpeech coding algorithms that have been developed for clean speech are often used in a noisy environment. We describe maximum a posteriori (MAP) and minimum mean square error (MMSE) techniques to estimate the clean-speech short-term predictor (STP) parameters from noisy speech. The MAP and MMSE estimates are obtained using a likelihood function computed by means of the DFT or Kalman filtering and empirical probability distributions based on multidimensional histograms. The method is assessed in terms of the resulting root mean spectral distortion between the "clean" speech STP parameters and the STP parameters computed with the proposed method from noisy speech. The estimated parameters are also applied to obtain clean speech estimates by means of a Kalman filter. The quality of the estimated speech as compared to the "clean" speech is assessed by means of subjective tests, signal-to-noise ratio improvement, and the perceptual speech quality measurement method Marcin Kuropatwinski, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Codebook driven short-term predictor parameter estimation for speech enhancementabstractIn this paper, we present a new technique for the estimation of short-term linear predictive parameters of speech and noise from noisy data and their subsequent use in waveform enhancement schemes. The method exploits a priori information about speech and noise spectral shapes stored in trained codebooks, parameterized as linear predictive coefficients. The method also uses information about noise statistics estimated from the noisy observation. Maximum-likelihood estimates of the speech and noise short-term predictor parameters are obtained by searching for the combination of codebook entries that optimizes the likelihood. The estimation involves the computation of the excitation variances of the speech and noise auto-regressive models on a frame-by-frame basis, using the a priori information and the noisy observation. The high computational complexity resulting from a full search of the joint speech and noise codebooks is avoided through an iterative optimization procedure. We introduce a classified noise codebook scheme that uses different noise codebooks for different noise types. Experimental results show that the use of a priori information and the calculation of the instantaneous speech and noise excitation variances on a frame-by-frame basis result in good performance in both stationary and nonstationary noise conditions. Sriram Srinivasan 0003, Jonas Samuelsson, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 3 |
| 2006 | Rate-distortion optimized quantization in multistage audio codingabstractIn this work, we develop a new method for quantization in multistage audio coding. Given a (perceptual) distortion measure and a bit-rate constraint, we analytically derive the optimal rate distribution between subcoders (stages) and the corresponding optimal quantizers using high-rate theory. The analytical solutions for optimal quantizers allow a coder to easily adapt to changes in bit-rate requirements. As an illustration of the new method, we consider quantization in a two-stage sinusoidal/waveform coder that is a widely used combination in audio coding. We show that at low total rates most of the rate should be assigned to the sinusoidal (model-based, subspace) subcoder, while at high total rates most of the rate should be assigned to the waveform (full-space) subcoder. We compare the new method to a reference quantization method that does not use rate-distortion optimization. A significantly higher performance of the new method is shown by means of a listening test. Renat Vafin, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Comparative rate-distortion performance of multiple description coding for real-time audiovisual communication over the InternetabstractTo facilitate real-time audiovisual communication through the Internet, forward error correction (FEC) and multiple description coding (MDC) can be used as low-delay packet-loss recovery techniques. We use both a Gilbert channel model and data obtained from real IP connections to compare the rate-distortion performance of different variants of FEC and MDC. Using identical overall rates with stringent delay constraints, we find that side-distortion optimized MDC generally performs better than Reed-Solomon-based FEC. If the channel condition is known from feedback, then channel-optimized MDC can be used to exploit this information, resulting in significantly improved performance. Our results confirm that two-independent-channel transmission is preferred to single-channel transmission, both for FEC and MDC. Moo Young Kim 0001, W. Bastiaan Kleijn |
IEEE Trans. Commun. | 2 |
| 2005 | Improved Kalman Filtering for Speech EnhancementabstractThe Kalman recursion is a powerful technique for reconstruction of the speech signal observed in additive background noise. In contrast to Wiener filtering and spectral subtraction schemes, the Kalman algorithm can be easily implemented in both causal and noncausal form. After studying the perceptual differences between these two implementations we propose a novel algorithm that combines the low complexity and the robustness of the Kalman filter and the proper noise shaping of the Kalman smoother. Volodya Grancharov, Jonas Samuelsson, W. Bastiaan Kleijn |
ICASSP (1) | 3 |
| 2005 | Stochastic Integration and Long Term Predictor Estimation under Noisy Conditions for Speech EnhancementabstractWe propose a method to estimate the short term predictor (STP) and the long-term predictor (LTP) under noisy conditions. We assume the speech signal to be a single, dual or triple frame asymptotic mean stationary process. The a priori STP parameter distribution is represented as databases sampled from the speech training data. Stochastic integration is used to obtain the minimum mean square error estimates of the STP parameters. After computing the STP parameters, the LTP parameters from a database of pairs of taps and excitation variances are matched, together with the lag, using a likelihood criterion, to the noisy speech. The estimated STP and LTP parameters are also applied to obtain clean speech estimates by means of a Wiener or a Kalman filter. For car noise with an SNR of -5dB, the proposed enhancement method gives a mean opinion score of 3.3 as measured using the perceptual speech quality measure software. Marcin Kuropatwinski, W. Bastiaan Kleijn |
ICASSP (1) | 2 |
| 2005 | Codebook-Based Bayesian Speech EnhancementabstractIn this paper, we propose a Bayesian approach for the estimation of the short-term predictor parameters of speech and noise, from the noisy observation. The resulting estimates of the speech and noise spectra can be used in a Wiener filter or any state-of-the-art speech enhancement system. We utilize a-priori information about both speech and noise in the form of trained codebooks of linear predictive coefficients. In contrast to current Bayesian estimation approaches that consider the excitation variances as part of the a-priori information, in the proposed method they are computed analytically based on the observation at hand. Consequently, the method performs well in nonstationary noise conditions. Experimental results confirm the superior performance of the proposed method compared to existing Bayesian approaches, such as those based on hidden Markov models. Sriram Srinivasan 0003, Jonas Samuelsson, W. Bastiaan Kleijn |
ICASSP (1) | 3 |
| 2005 | Distortion measures for vector quantization of noisy spectrumabstractIn this paper we address the problem of vector quantization of speech in a noisy environment. We show that the performance of a vector quantization system can be improved by adapting the distortion measure to the changing environmental conditions. The proposed method emphasizes the distortion in spectral regions where the speech signal dominates. The method functions well even when conventional pre-processor methods fail because the noise statistics cannot be estimated reliably from speech pauses (as, e.g., in tandeming operations). Objective tests confirm that the use of environmentally adaptive measures significantly improves estimation accuracy in noisy speech, while preserving the quality in the case of clean input. 1. Volodya Grancharov, Jonas Samuelsson, W. Bastiaan Kleijn |
INTERSPEECH | 3 |
| 2005 | Denoising through source separation and minimum trackingabstractIn this paper, we develop a multi-channel noise reduction algorithm based on blind source separation (BSS). In contrast to general BSS algorithms that attempt to recover all the signals, we explicitly estimate only the speech signal. By tracking the minimum of the spectral density of the microphone signals, noise-only segments are identified. The coefficients of the unmixing matrix that are necessary to separate the speech are identified from these segments through the optimization of an appropriate energy criterion. Since the proposed method explicitly estimates the speech signal from the noisy mixture, it does not suffer from the permutation problem that is typical to conventional BSS techniques. The method is applicable to both instantaneous and convolutive mixtures and achieves the separation in a single step, without the need for iterations. Experimental results show superior performance compared to a general BSS algorithm. Sriram Srinivasan 0003, Mattias Nilsson 0002, W. Bastiaan Kleijn |
INTERSPEECH | 3 |
| 2005 | On noise gain estimation for HMM-based speech enhancementabstractTo address the variation of noise level in non-stationary noise signals, we study the noise gain estimation for speech enhancement using hidden Markov models (HMM). We consider the noise gain as a stochastic process and we approximate the probability density function (PDF) to be log-normal distributed. The PDF parameters are estimated for every signal block using the past noisy signal blocks. The approximated PDF is then used in a Bayesian speech estimator minimizing the Bayes risk for a novel cost function, that allows for an adjustable level of residual noise. As a more computationally efficient alternative, we also derive the maximum likelihood (ML) estimator, assuming the noise gain to be a deterministic parameter. The performance of the proposed gain-adaptive methods are evaluated and compared to two reference methods. The experimental results show significant improvement under noise conditions with time-varying noise energy. David Yuheng Zhao, W. Bastiaan Kleijn |
INTERSPEECH | 2 |
| 2005 | On frequency quantization in sinusoidal audio codingabstractIn this work, we develop a new method for jointly optimal quantization of sinusoidal frequencies, amplitudes, and phases and apply the method to sinusoidal audio coding. This is an extension of an earlier work on quantization of sinusoidal amplitudes and phases to frequencies. The optimization is performed for a set of sinusoids that models a short segment of an audio signal. For a given bit-rate constraint, the optimal quantizers minimize a single-letter weighted distortion measure that accounts for perceptual importance of sinusoids. The quantizers are derived analytically using high-rate theory. The method yields high performance and has a number of practical advantages over conventional sinusoidal quantization methods. Renat Vafin, Deep Prakash, W. Bastiaan Kleijn |
IEEE Signal Process. Lett. | 3 |
| 2005 | Entropy-constrained polar quantization and its application to audio codingabstractIn this work, we present a new method for quantization of sinusoidal amplitudes and phases, and apply the method to sinusoidal coding of speech and audio signals. The method is based on unrestricted polar quantization, where phase quantization accuracy depends on amplitude. Amplitude and phase quantizers are derived under an entropy (average rate) constraint using high-rate assumptions. First, we derive optimal quantizers for one sinusoid and a mean-squared error distortion measure. We provide a detailed analysis of entropy-constrained unrestricted polar quantization, showing its high performance and practicality even at low rates. Second, we find optimal quantizers for a set of sinusoids that model a short segment of an audio signal. The optimization is performed using a weighted error measure that can account for the masking effect in the human auditory system. We find the optimal rate distribution between sinusoids, as well as the corresponding optimal amplitude and phase quantizers, based on the perceptual importance of sinusoids defined by masking. The new method is used in an audio-coding application and is shown to significantly outperform a conventional sinusoidal quantization method where phase quantization accuracy is identical for all sinusoids. Renat Vafin, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | Multivariate block polar quantizationabstractWe introduce multivariate block polar quantization (MBPQ). MBPQ minimizes the weighted squared-error distortion for a set of complex variables representing one block of a signal under a resolution constraint for the entire block. MBPQ performs below the lower bound for classical bivariate quantization, both for Gaussian complex variables and for complex variables found from sinusoidal analysis of audio data. Still, it is of similar complexity as traditional polar quantizers. In the case of audio data, we found a performance gain of about 2.5 dB over the best performing conventional resolution-constrained polar quantization (an extension of unrestricted polar quantization). Harald Pobloth, Renat Vafin, W. Bastiaan Kleijn |
IEEE Trans. Commun. | 3 |
| 2004 | Multiple-description vector quantization using translated lattices with local optimization [joint source/channel coding]abstractMultiple-description coding is a joint source- and channel coding technique suitable for real-time multimedia transmission over erasure channels. This work improves the previous methods of multiple-description vector quantization using lattice structured codebooks by introducing translated lattices in the single-description codebooks. The quantizer can easily adapt to the current channel conditions, using the locally optimized combined-description codebooks, assuming that channel statistics are available at the encoder. Compared to previous methods, the central distortion is greatly reduced for noisy channels, without a significant effect on complexity. David Yuheng Zhao, W. Bastiaan Kleijn |
GLOBECOM | 2 |
| 2004 | Noise-dependent postfilteringabstractThe paper introduces a modification of the commonly used postfilter that improves performance when acoustic background noise is present. The modification consists of replacing the nonadaptive postfilter parameters that govern the degree of spectral emphasis (commonly denoted as /spl gamma//sub 1/ and /spl gamma//sub 2/) with parameters that adapt to the noise statistics. We describe an effective mapping from the noise statistics to the emphasis parameters and provide a low complexity noise estimation algorithm that is sufficient for this application. The resulting noise-adaptive postfilter successfully attenuates the background noise and naturally converges to the conventional postfilter at high SNR conditions. Thus, the speech enhancement problem is solved with minimal modification of legacy codecs, since the existing structure of the speech codec is used. Test results indicate that the presented algorithm significantly outperforms the standard postfilter with non-adaptive parameters. Volodya Grancharov, Jonas Samuelsson, W. Bastiaan Kleijn |
ICASSP (1) | 3 |
| 2004 | Multi-variate block polar quantization and an application to audioabstractWe introduce multi-variate block polar quantization (MBPQ). MBPQ minimizes a weighted distortion for a set of complex variables representing one block of a signal under a resolution constraint for the entire block. MBPQ allows for different probability distributions in different dimensions of the set of complex variables. It outperforms a block polar quantizer introduced earlier (Pobloth, H. et al., Proc. Eurospeech, p.1097-100, 2003), and unrestricted polar quantization (UPQ), for both Gaussian complex variables and sinusoids found from audio data. In the case of audio data, we found a performance gain of about 2.5 dB over the best performing conventional resolution-constrained polar quantization, UPQ. Harald Pobloth, Renat Vafin, W. Bastiaan Kleijn |
ICASSP (4) | 3 |
| 2004 | Estimation of short-term predictor parameters for coding and enhancement of noisy speechabstractWe describe a technique for obtaining estimates of the short-term predictor parameters of speech under noisy conditions. We use a-priori information about speech in the form of a trained codebook of speech linear predictive coefficients. Our contribution is two-fold. First, we provide a framework where the standard vector quantization search to obtain the quantized linear predictive coefficients can be replaced by a maximum likelihood search, given the noisy observation, the speech codebook and an estimate of the noise. This results in an enhancement method that is integrated with parametric coders such as linear predictive analysis-by-synthesis coders. Second, we provide a scheme where the chosen vector is not restricted to be an element of the codebook. An interpolative search between the maximum likelihood estimate and its nearest neighbors in the codebook is used to improve the precision of the estimated parameters. Such a scheme is relevant when enhancement is considered separately from coding. Experimental results show improved performance for the proposed methods. Sriram Srinivasan 0003, Jonas Samuelsson, W. Bastiaan Kleijn |
ICASSP (1) | 3 |
| 2004 | Towards optimal quantization in multistage audio codingabstractWe develop a new method for quantization in multistage audio coding. We consider the case of a two-stage sinusoidal/waveform coder. Given a distortion measure and a bit-rate constraint, we analytically derive the optimal rate distribution between subcoders (stages) and the corresponding optimal quantizers, which allows the coder to adapt easily to changes in bit-rate requirements. We verify that the performance, both in terms of signal-to-noise ratio (SNR) and perceptual quality, is higher if the input to the second stage is obtained by subtracting the quantized first-stage reconstruction from the original signal, as opposed to subtracting the unquantized reconstruction. Renat Vafin, W. Bastiaan Kleijn |
ICASSP (4) | 2 |
| 2004 | Comparison of transmitter - based packet-loss recovery techniques for voice transmissionabstractTo facilitate real-time voice communication through the Internet, forward error correction (FEC) and multiple description coding (MDC) can be used as low-delay packet-loss recovery techniques. We u ... Moo Young Kim 0001, W. Bastiaan Kleijn |
INTERSPEECH | 2 |
| 2004 | Speech enhancement using adaptive time-domain segmentationabstractIn this paper, we investigate the benets of using an adaptive seg-mentation of the speech signal in speech enhancement. The adap-tive segmentation scheme divides the signal into the longest seg-ments within which stationarity is preserved, thus providing a good time-frequency resolution. The segmentation is performed with the help of an orthogonal library of local cosine bases using a compu-tationally efcient tree-structured best-basis search. We show that such an adaptive segmentation results in improved speech enhance-ment compared to a xed segmentation. The resulting enhanced speech is free from musical noise, without any additional smooth-ing. 1. Sriram Srinivasan 0003, W. Bastiaan Kleijn |
INTERSPEECH | 2 |
| 2004 | A time-domain interpretation for the LSP decompositionabstractThe line spectrum pair (LSP) decomposition is a widely used method in speech coding. In this article, we will show that the LSP polynomials, whose trivial zeros have been removed, are equivalent to two optimal (in the mean square sense) predictors in which a sample is predicted from linear combinations of its previous averaged and differentiated values. Tom Bäckström, Paavo Alku, Tuomas Paatero, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 4 |
| 2004 | KLT-based adaptive classified VQ of the speech signalabstractCompared to scalar quantization (SQ), vector quantization (VQ) has memory, space-filling, and shape advantages. If the signal statistics are known, direct vector quantization (DVQ) according to these statistics provides the highest coding efficiency, but requires unmanageable storage requirements if the statistics are time varying. In code-excited linear predictive (CELP) coding, a single "compromise" codebook is trained in the excitation-domain and the space-filling and shape advantages of VQ are utilized in a nonoptimal, average sense. In this paper, we propose Karhunen-Loe/spl grave/ve transform (KLT)-based adaptive classified VQ (CVQ), where the space-filling advantage can be utilized since the Voronoi-region shape is not affected by the KLT. The memory and shape advantages can be also used, since each codebook is designed based on a narrow class of KLT-domain statistics. We further improve basic KLT-CVQ with companding. The companding utilizes the shape advantage of VQ more efficiently. Our experiments show that KLT-CVQ provides a higher SNR than basic CELP coding, and has a computational complexity similar to DVQ and much lower than CELP. With companding, even single-class KLT-CVQ outperforms CELP, both in terms of SNR and codebook search complexity. Moo Young Kim 0001, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 2 |
| 2003 | Minimum mean square error estimation of speech short-term predictor parameters under noisy conditionsabstractMinimum mean square error (MMSE) estimation of the speech short-term predictor (STP) parameters in the line spectral frequency (LSF) representation is considered. We exploit that the square error between LSF parameter vectors is a subjectively meaningful distortion criterion. As speech coding algorithms are often used in a noisy environment, it is relevant to estimate the STP parameters used in these algorithms under the inclusion of noise statistics. In our experiments, car noise is used as an example of an autoregressive (AR) noise process. The MMSE estimates are obtained using a likelihood function computed by means of Kalman filtering and empirical probability distributions. The method is assessed in terms of the resulting root mean spectral distortion between the 'clean' speech STP parameters and the STP parameters computed using the proposed method from noisy speech. Marcin Kuropatwinski, W. Bastiaan Kleijn |
ICASSP (1) | 2 |
| 2003 | Polar quantization of sinusoids from speech signal blocks
Harald Pobloth, Renat Vafin, W. Bastiaan Kleijn |
INTERSPEECH | 3 |
| 2003 | Speech enhancement using a-priori information
Sriram Srinivasan 0003, Jonas Samuelsson, W. Bastiaan Kleijn |
INTERSPEECH | 3 |
| 2003 | On line spectral frequenciesabstractThe commonly used line spectral frequencies form the roots of symmetric and antisymmetric polynomials constructed from a linear predictor. We provide a new, simpler proof that the symmetric and antisymmetric polynomials can be regarded as optimal constrained predictors that correspond to predicting from the low-pass and high-pass filtered signal, respectively. W. Bastiaan Kleijn, Tom Bäckström, Paavo Alku |
IEEE Signal Process. Lett. | 1 |
| 2002 | A time domain reformulation of linear prediction equivalent to the LSP decompositionabstractThe Line spectrum pair (LSP) decomposition is a widely used method in speech coding. In this paper, we will present a reformulation of conventional linear prediction which is equivalent to the LSP decomposition. The paper shows that the symmetric and antisymmetric polynomials of the LSP decomposition are equivalent to two filters, that are determined by predicting a signal sample using its averaged and differentiated previous values. Tom Bäckström, Paavo Alku, W. Bastiaan Kleijn |
ICASSP | 3 |
| 2002 | Spline-based continuous-time pitch estimationabstractPitch-synchronous speech coding algorithms can achieve low bit rates without compromising the quality. However, the effectiveness of pitch-synchronous coding depends strongly on the ability to estimate precisely and reliably the fundamental period of the speech signal. We present a novel pitch postprocessing method that significantly improves the accuracy and reliability of pitch estimation. In contrast to the classical schemes, the pitch is treated as a continuous function in time and amplitude. B-Spline signal processing, half wave rectification, and multi-stage, multi-resolution optimization are essential parts of the procedure. The performance of the method is evaluated objectively and subjectively using the Waveform Interpolation coder. The objective results show that, for voiced segments, the method significantly (60% on average) decreases the energy of the unvoiced component estimate compared to using an unprocessed pitch. Listening tests show a 90% preference of speech generated using our postprocessor over speech generated using a conventional method. Andrei Jefremov, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2002 | KLT-based classified VQ for the speech signalabstractIf the signal statistics are given, direct vector quantization (DVQ) according to these statistics provides the highest coding efficiency, but requires unmanageable storage requirements. In. code-excited linear predictive (CELP) coding. a single “compromise” codebook is trained in the prediction residual-domain and the space-filling and shape advantages of vector quantization (VQ) are utilized in a non-optimal, average sense. In this paper. we propose a Karhunen-Loève Transform (KLT)-based classified VQ (CVQ), where the space-filling advantage can be utilized since the Voronoi-region shape is not affected by the KLT. The memory and shape advantages can be also used, since each codebook is designed based on a narrow class of KL T -domain statistics. Our experiments show that the KLT-CVQ provides a higher SNR than CELP and (single-codebook) DVQ, and has a computational complexity similar to DVQ and much lower than CELP. Storage requirements are modest because of the energy concentration property of the KL T. Moo Young Kim 0001, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2002 | Gaussian mixture model based mutual information estimation between frequency bands in speechabstractIn this paper, we investigate the dependency between the spectral envelopes of speech in disjoint frequency bands, one covering the telephone bandwidth from 0.3 kHz to 3.4 kHz and one covering the frequencies from 3.7 kHz to 8 kHz. The spectral envelopes are jointly modeled with a Gaussian mixture model based on mel-frequency cepstral coefficients and the log-energy-ratio of the disjoint frequency bands. Using this model, we quantify the dependency between bands through their mutual information and the perceived entropy of the high frequency band. Our results indicate that the mutual information is only a small fraction of the perceived entropy of the high band. This suggests that speech bandwidth extension should not rely only on mutual information between narrow- and high-band spectra. Rather, such methods need to make use' of perceptual properties to ensure that the extended signal sounds pleasant. Mattias Nilsson 0002, Harald Gustaftson, Søren Vang Andersen, W. Bastiaan Kleijn |
ICASSP | 4 |
| 2002 | Entropy-constrained polar quantization: theory and an application to audio codingabstractIn this work, we develop entropy-constrained unrestricted polar quantizers, where phase quantization depends on the input amplitude. Formulas for amplitude and phase quantization point densities are derived under high-rate assumptions. It is shown that the mean-squared distortion is decreased considerably as compared to strictly polar quantization and approaches that of scalar rectangular quantization asymptotically with increasing rate. The unrestricted polar quantization is generalized to include a weighted error measure, such that it accounts for masking effects of the human auditory system. Both amplitude and phase quantization depend on the perceptual importance of sinusoids. The new method is applied to a sinusoidal audio coder, and is shown to outperform a conventional sinusoidal quantization method where the number of phase quantization bits is the same for all audible sinusoids. Renat Vafin, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2002 | On the relevance of bandwidth extension for speaker verificationabstractIn this paper, we consider the effect of a bandwidth extension of narrow-band speech signals (0.3-3.4 kHz) to 0.3-8 kHz on speaker verification.Using covariance matrix based verification systems together with detection error trade-off curves, we compare the performance between systems operating on narrowband, wide-band (0-8 kHz), and bandwidth-extended speech.The experiments were conducted using different short-time spectral parameterizations derived from microphone and ISDN speech databases.The studied bandwidth-extension algorithm did not introduce artifacts that affected the speaker verification task, and we achieved improvements between 1 and 10 percent (depending on the model order) over the verification system designed for narrow-band speech when mel-frequency cepstral coefficients for the short-time spectral parameterization were used. Marcos Faúndez-Zanuy, Mattias Nilsson 0002, W. Bastiaan Kleijn |
INTERSPEECH | 3 |
| 2002 | Sinusoidal modeling using psychoacoustic-adaptive matching pursuitsabstractWe propose a segment-based matching-pursuit algorithm where the psychoacoustical properties of the human auditory system are taken into account. Rather than scaling the dictionary elements according to auditory perception, we define a psychoacoustic-adaptive norm on the signal space that can be used for assigning the dictionary elements to the individual segments in a rate-distortion optimal way. The new algorithm is asymptotically equal to signal-to-mask-ratio-based algorithms in the limit of infinite-analysis window length. However, the new algorithm provides a significantly improved selection of the dictionary elements for finite window length. Richard Heusdens, Renat Vafin, W. Bastiaan Kleijn |
IEEE Signal Process. Lett. | 3 |
| 2001 | Sinusoidal modeling of audio and speech using psychoacoustic-adaptive matching pursuitsabstractWe propose a segment-based matching pursuit algorithm where the psychoacoustical properties of the human auditory system are taken into account. Rather than scaling the dictionary elements according to auditory perception, we define a psychoacoustic-adaptive norm on the signal space which can be used for assigning the dictionary elements to the individual segments in a rate-distortion optimal manner. The new algorithm is asymptotically equal to signal-to-mask ratio based algorithms in the limit of infinite analysis window length. However, the new algorithm provides a significantly improved selection of the dictionary elements for finite window length. Richard Heusdens, Renat Vafin, W. Bastiaan Kleijn |
ICASSP | 3 |
| 2001 | All-pole modelling of mixed excitation signalsabstractConventional linear prediction (LP) techniques can fail to adequately model speech spectra when the model order is too low and/or when the input is periodic (voiced speech). We view the LP modelling problem as a correlation matching problem. We introduce a correlation matching criterion which models the signal as a filtered mixture of a noise-like excitation and a periodic excitation. As such it is an extension of the discrete all-pole (DAP) modelling approach. The new technique provides a means to generate LP spectra that evolve more smoothly from frame to frame even when the excitation signal has a periodic component with changing period. Peter Kabal, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2001 | Estimation of the excitation variances of speech and noise AR-models for enhanced speech codingabstractIn this paper, we consider the estimation of short-term predictor (STP) parameters under noisy conditions. The possible autoregressive spectral shapes of the speech and additive noise are stored in AR-coefficient codebooks. The product codebook is then searched to maximize the likelihood function of the observed noisy speech signal frame. The maximum likelihood (ML) estimates of the variances of the driving term are computed for each pair of the speech and noise AR spectra. For further processing (e.g., Kalman filtering or speech coding using enhanced STP parameters), the spectra and variances that yield the maximum of the likelihood function are selected. To evaluate the proposed method, the estimates of the spectral shapes and variances are compared with those computed from clean speech signal using a common spectral distortion measure. Globally maximizing the likelihood function over some restricted region of the parameter space, the presented approach provides robust estimates. Marcin Kuropatwinski, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2001 | Avoiding over-estimation in bandwidth extension of telephony speechabstractWe present a new way of treating the problem of extending a narrow-band signal to a wide-band signal. For many cases of bandwidth extension, the high-band energy is overestimated, leading to undesirable audible artifacts. To overcome these problems we introduce an asymmetric cost-function in the estimation process of the high-band that penalizes over-estimates more than under-estimates of the energy in the high-band. We show that the resulting attenuation of the estimated high-band energy depends on the broadness of the a-posteriori distribution of the energy given the extracted information about the narrow-band. Thus, the uncertainty about how to extend the signal at the high-band influences the level of extension. Results from a listening test show that the proposed algorithm produces less artifacts. Mattias Nilsson 0002, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2001 | Modifying transients for efficient coding of audioabstractWe propose a method for efficient representation of transients in audio signals. We estimate the transient component of an original audio signal and modify the locations of the transients in such a way that the transients can occur only at locations defined by a relatively coarse time grid. This procedure allows an efficient representation of transients with damped sinusoids. We also verify that the introduced modifications do not result in a perceptual difference between the original and the modified audio signals. Renat Vafin, Richard Heusdens, W. Bastiaan Kleijn |
ICASSP | 3 |
| 2001 | Squared error as a measure of phase distortionabstractIn this article, we investigate how accurately the squared error captures perceptual errors introduced by Fourier phase spectrum changes. We measure the perceptual error using the Auditory Image Model by Patterson et al.. The squared error is found to represent the perceptual error well for low squared errors but it saturates. Thus, a further increase in squared error does on average not lead to any further increase in perceptual error. This suggests that encoding phase using squared-error trained codebooks only improves perceived quality when operating at high bit rates. To verify this, phase was encoded with codebooks of different sizes. As expected, increasing the codebook size has very little influence on the average perceptual error for low rates, which is confirmed by listening tests. Our results suggest that a direct phase codebook is an inefficient representation of the relevant information contained in phase. Harald Pobloth, W. Bastiaan Kleijn |
INTERSPEECH | 2 |
| 2000 | A frame interpretation of sinusoidal coding and waveform interpolationabstractConventional sinusoidal and waveform interpolation coders have a modeling error that limits performance at high rates. In addition, their time-frequency localization of the unvoiced speech component is often insufficient to characterize the speech signal in a perceptually accurate manner. Both problems can be addressed with frame expansions. The use of frame representations, that are very similar to conventional implementations of the fore-mentioned coders, eliminates modeling errors. Furthermore, frame representations can be selected so as to preserve the time-frequency location of the unvoiced component even when it is characterized with statistical parameters only. W. Bastiaan Kleijn |
ICASSP | 1 |
| 2000 | On the mutual information between frequency bands in speechabstractIn this paper we investigate the mutual information in speech between the spectral envelope of the high frequency band and low frequency bands of various widths. Direct methods on the computation of the mutual information often result in an excessive amount of data required even for modest situations. We reduce the required amount of data by quantizing the low band leading to a lower bound expression on the mutual information. We indicate by simulation that this lower bound is in the same order of magnitude as the true mutual information. Simulations on speech show that we have no less than 0.1 bit of shared information between the slope of the high band and the low frequency band from 0-4 kHz. Performing the analogous simulation with the gain of the high band we obtained no less than 0.45 bit of mutual information. Mattias Nilsson 0002, Søren Vang Andersen, W. Bastiaan Kleijn |
ICASSP | 3 |
| 2000 | Exploiting time and frequency masking in consistent sinusoidal analysis-synthesisabstractIn this paper, we elaborate on the issue of analysis-synthesis consistency in sinusoidal coding. Our analysis is based on windowed sinusoids, and uses the same amplitude-complementary window as is used in the overlap-add synthesis. Reconstructions of the neighboring segments are taken into account when forming a particular analysis segment. Sinusoidal estimation is based on a perceptual criterion. In our new procedure, when analyzing the current segment we take advantage of the forward masking effect due to estimated sinusoids in the previous segments (possibly overlapping with the current segment). Experimental results verify that the number of sinusoids can be reduced significantly with our time masking model, without introducing perceptual artifacts in the reconstructed signal. Renat Vafin, Søren Vang Andersen, W. Bastiaan Kleijn |
ICASSP | 3 |
| 2000 | On time-frequency masking in voiced speechabstractThis paper addresses the issue of masking of noise in voiced speech. First, we examine the audibility of cyclostationary narrow-band noise bursts added to voiced speech generated by synthetic excitation. Varying the temporal location of noise within a pitch cycle corresponds to varying its phase spectrum. Using this fact, we found that a change of phase of the noise in the high frequency region is more perceptible for a low-pitched sound than for a high-pitched sound. We then propose a pitch-dependent temporal weighting function which can be employed in quantization of pitch cycle waveforms. In a second experiment, we found that the audibility of high-frequency noise added to natural speech can be significantly reduced using this weighting function. Jan Skoglund, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 2 |
| 1999 | On speech coding in a perceptual domainabstractFor speech coders which fall within the class of waveform coders, the reconstructed signal approaches the original with increasing bit rate. In such coders, the distortion criterion generally operates on the speech signal or a signal obtained by adaptive linear filtering of the speech signal. To satisfy computational and delay constraints, the distortion criterion must be reduced to a very simple approximation of the auditory system. This drawback of conventional approaches motivates a new speech coding paradigm in which the coding is performed in a domain where the single-letter squared-error criterion forms an accurate representation of perception. The new paradigm requires a model of the auditory periphery which is accurate, can be be inverted with relatively low computational effort, and which represents the signal with relatively few parameters. We develop such a model of the auditory periphery and discuss its suitability for speech coding. The results indicate that the new paradigm in general and our auditory model in particular form a promising basis for the coding of both speech and audio at low bit rates. Gernot Kubin, W. Bastiaan Kleijn |
ICASSP | 2 |
| 1999 | On phase perception in speechabstractIn this paper we define perceptual phase capacity as the size of a codebook of phase spectra necessary to represent all possible phase spectra in a perceptually accurate manner. We determine the perceptual phase capacity for voiced speech. To this purpose, we use an auditory model which indicates if phase spectrum changes are audible or not. The correct performance of the model was adjusted and verified by listening tests. The perceptual phase capacity in low pitched speech is found to be much higher than it is for high pitched speech. Our results are consistent with the well known fact that speech coding schemes which preserve the phase accurately work better for male voices, while coders which put more weight on the amplitude spectrum of the speech signal result in better quality for female speech. Harald Pobloth, W. Bastiaan Kleijn |
ICASSP | 2 |
| 1998 | Removal of sparse-excitation artifacts in CELPabstractIn CELP, the use of codebooks with entries with only a few non-zero samples provides high speech quality and facilitates fast computation. With decreasing bit-rate, the intervals between the pulses increase, and the quality of the reconstructed signal begins to suffer from a particular type of artifact, which is strongest for noise-like segments. In this paper we describe experiments which show that the perceived artifacts are mainly concentrated at frequencies above 3 kHz, and this is consistent with our understanding of auditory theory. Our analysis leads to simple strategies to eliminate the artifacts, even at lower bit rates. We describe both a non-adaptive and an adaptive post-processing method to remove the artifacts. The methods are demonstrated to be efficient when used in the ACELP algorithm. A closed-loop method for ACELP is also described. Roar Hagen, Erik Ekudden, W. Bastiaan Kleijn |
ICASSP | 4 |
| 1998 | Waveform interpolation coding with pitch-spaced subbands
W. Bastiaan Kleijn, Huimin Yang, Ed F. Deprettere |
ICSLP | 1 |
| 1998 | On the significance of temporal masking in speech coding
Jan Skoglund, W. Bastiaan Kleijn |
ICSLP | 2 |
| 1997 | On optimal and minimum-entropy decodingabstractIn the quantization of a signal in speech coding, dependencies between its samples are often neglected. Generally, these dependencies are then also neglected at the decoder. However, usually a priori information about these dependencies is available, making it possible to improve decoder performance by means of enhanced decoding. An attractive feature of enhanced decoding is that it can be applied to existing coding standards. This paper describes several enhanced decoding methods, including a vector decoding method and a method which aims at reducing the differential entropy rate of the decoded signal. Experimental results are used to confirm that both these decoding procedures can provide better performance than conventional decoding for common signal/encoder combinations. W. Bastiaan Kleijn |
ICASSP | 1 |
| 1997 | Perceptual entropy rate estimates for the phonemes of American EnglishabstractWe estimated the perceptual entropy rate of the phonemes of American English and found that the upper limit of the perceptual entropy of voiced phonemes is approximately 1.4 bit/sample, whereas the perceptual entropy of unvoiced phonemes is approximately 0.9 bit/sample. Results indicate that a simple voiced/unvoiced classification is suboptimal when trying to minimize bit rate. We used two different methods for the entropy estimation, and the results of both methods show that short segments of unvoiced speech are approximately Gaussian. Vincent van de Laar, W. Bastiaan Kleijn, Ed F. Deprettere |
ICASSP | 2 |
| 1997 | Using a perception-based frequency scale in waveform interpolationabstractIn speech coding it is important to focus the coding effort on the perceptually important features of the speech signal. This paper describes new quantization techniques which take advantage of current knowledge of human perception in speech coders. The new procedures exploit the frequency-dependent frequency resolution of the human auditory system. The methods are applied to the waveform interpolation (WI) coder, and their effectiveness is confirmed with experimental results. The principles described in the paper are not restricted to the WI coder, but are also applicable to many other speech coding algorithms. Jes Thyssen, W. Bastiaan Kleijn, Roar Hagen |
ICASSP | 2 |
| 1997 | Quantization using wavelet based temporal decomposition of the LSF
Aweke N. Lemma, W. Bastiaan Kleijn, Ed F. Deprettere |
EUROSPEECH | 2 |
| 1996 | A low-complexity waveform interpolation coderabstractA recent independent survey found a 2.4 kbit/s waveform-interpolation (WI) algorithm to perform better than other state-of-the-art speech coders. However, this coder had a very high level of computational complexity. The introduction of various techniques, including a time-varying waveform sampling rate and a cubic B-spline waveform representation, has reduced the computational complexity by an order of magnitude. The new implementation allows full-duplex real-time operation on a single DSP device and on an average workstation (the latter using non-optimized compiled C source code). The new coder also contains a number of new features which improve the quality of the reconstructed speech signal. W. Bastiaan Kleijn, Yair Shoham, Deep Sen, Roar Hagen |
ICASSP | 1 |
| 1996 | On memoryless quantization in speech codingabstractIn memoryless quantization, neither the encoder nor the decoder has memory, and quantization noise shaping is not used. We show that, by constraining the parameter dynamics during quantization at the encoder, the performance of speech coders can be enhanced significantly without adding to the delay. The proposed method retains the advantages of memoryless quantization, including channel-error robustness. W. Bastiaan Kleijn, Roar Hagen |
IEEE Signal Process. Lett. | 1 |
| 1995 | A speech coder based on decomposition of characteristic waveformsabstractFor low-rate speech coding it is advantageous to represent the speech signal as an evolving characteristic waveform (CW). The CW evolves slowly when the speech signal is clearly voiced and rapidly when the speech signal is clearly unvoiced. The voiced (periodic) and unvoiced (nonperiodic) components of the speech signal can be separated by a simple nonadaptive filter in the CW domain. Because of perceptual effects, a significant increase in coding efficiency is obtained by coding these two components separately. A 2.4 kb/s coder using these principles was developed. In an independent evaluation, the performance of the 2.4 kb/s waveform interpolation (WI) coder was found to be at least equivalent to the 4.8 kb/s FS1016 standard for all of the many tests. W. Bastiaan Kleijn, Jesper Haagen |
ICASSP | 1 |
| 1995 | Spectral dynamics is more important than spectral distortionabstractLinear prediction coefficients are used to describe the power-spectrum envelope in the majority of low-bit-rate coders. The performance of quantizers for the linear-prediction coefficients is generally evaluated in terms of spectral distortion. This paper shows that the audible distortion in low-bit-rate coders is often more a function of the dynamics of the power-spectrum envelope than of the spectral distortion as usually evaluated. Smoothing the evolution of the power-spectrum envelope over time increases the reconstructed speech quality. A reasonable objective is to find the smoothest path that keeps the quantized parameters within the Voronoi regions associated with the transmitted quantization index. We demonstrate increased quantizer performance by such smoothing of the line-spectral frequencies. H. Petter Knagenhjelm, W. Bastiaan Kleijn |
ICASSP | 2 |
| 1994 | Time-scale modification of speech based on a nonlinear oscillator modelabstractIntroduces a new signal model for speech: a nonlinear oscillator with a slowly time-varying state-space representation produces the full speech waveform at its output. Short-time stationary intervals of the waveform are interpreted as (transient) attractors of the underlying dynamical system. The model is adapted by building a 'state-transition codebook' (a table of all possible state transitions) for each speech frame. As a first test for the model the authors have developed a novel time-scale modification method for speech. This method produces high quality output at moderate computational cost.> Gernot Kubin, W. Bastiaan Kleijn |
ICASSP (1) | 2 |
| 1994 | Transformation and decomposition of the speech signal for codingabstractThe speech signal is represented by an evolving characteristic waveform. The characteristic waveform is decomposed into a slowly evolving waveform and a rapidly evolving waveform, representing the quasi-periodic and other components of speech, respectively. These two evolving waveforms have fundamentally different quantization requirements. The decomposition allows efficient coding of voiced and unvoiced speech at bit rates between 2 and 8 kb/s.> W. Bastiaan Kleijn, Jesper Haagen |
IEEE Signal Process. Lett. | 1 |
| 1994 | On the periodicity of speech coded with linear-prediction based analysis by synthesis codersabstractThe closed-loop pitch predictor (CLPP) is an essential part of most linear-prediction based analysis-by-synthesis (LPAS) coders. The author discusses the relation between the periodicity of the original and reconstructed signals for LPAS coders with a CLPP. It is shown that the periodicity is unchanged if the periodicity of the original signal is generated with an autoregressive model and increases or decreases for several other signal classes. The theoretical findings are confirmed by experiments.> W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 1 |
| 1994 | Interpolation of the pitch-predictor parameters in analysis-by-synthesis speech codersabstractThe pitch-predictor contributes greatly to the efficiency of current analysis-by-synthesis speech coders by mapping the past reconstructed signal into the present. However, for good performance, it is required that its parameters are updated often (one every 2.5-7.5 ms). A slower update rate of the pitch-predictor delay results in time misalignment between the original signal and the pitch-predictor contribution to the reconstructed signal and the pitch-predictor contribution to the reconstructed signal. The authors introduce a new procedure, that allows a slow update rate of the pitch-predictor parameters without this problem. In this method the original signal is modified in a closed-loop fashion such that the parameter values obtained by interpolation of open-loop estimates form the optimal encoding of the modified signal. This new paradigm is a generalization of the familiar analysis-by-synthesis principle. The generalized analysis-by-synthesis principle can be used for interpolation of both the pitch-predictor delay and gain. The authors compare, by means of a subjective test, speech signals encoded with different versions of the code-excited linear predictor delay and gain. They compare, by means of a subjective test, speech signals encoded with different versions of the code-excited linear predictor (CELP) coder. The comparison shows that a pitch predictor exploiting the present interpolation strategy, with an update rate of 50 Hz, provides a subjective speed quality similar to a conventional pitch predictor where the parameters are updated for every pitch cycle. W. Bastiaan Kleijn, Ravi Prakash Ramachandran, Peter Kroon |
IEEE Trans. Speech Audio Process. | 1 |
| 1993 | A 5.85 kbits CELP algorithm for cellular applications
W. Bastiaan Kleijn, Peter Kroon, Luca Cellario, Daniele Sereno |
ICASSP (2) | 1 |
| 1993 | Encoding speech using prototype waveformsabstractVoiced speech is interpreted as a concentration of slowly evolving pitch-cycle waveforms. This signal can be reconstructed by interpolation from a downsampled sequence of pitch-cycle waveforms with a rate of one prototype waveform per 20-30 ms interval. The prototype waveform is described by a set of linear-prediction (LP) filter coefficients describing the formant structure and a prototype excitation waveform, quantized with analysis-by-synthesis procedures. The speech signal is reconstructed by filtering an excitation signal consisting of the concatenation of (infinitesimal) sections of the instantaneous excitation waveforms. To obtain the correct level of periodicity, the short-term and the long-term correlations between the instantaneous excitation waveforms can be controlled explicitly. Thus, distortions such as noise, reverberation, and buzziness can be prevented. The coding method is easily combined with existing LP-based speech coders, such as CELP, for unvoiced signals. Excellent voiced speech quality is obtained at rates between 3.0 and 4.0 kb/s.> W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 1 |
| 1992 | Generalized analysis-by-synthesis coding and its application to pitch predictionabstractMany modifications can be applied to a speech signal without changing its perceptual quality. For a particular speech coder, the coding efficiency will differ for distinct modifications. To exploit this, the authors introduced a generalized analysis-by-synthesis procedure. In this procedure, a search is performed over a multitude of modified original signals (on a blockwise basis), and the signal which can be encoded with the least distortion is selected for transmission. At the receiver, a quantized version of this modified original signal is constructed. The authors discuss the application of generalized analysis-by-synthesis coding to the pitch predictor of a code excited linear predictor (CELP) coder. The use of this technique makes it possible to transmit the pitch predictor parameters at a much lower rate than conventional approaches, without compromising speech quality.> W. Bastiaan Kleijn, Ravi Prakash Ramachandran, Peter Kroon |
ICASSP | 1 |
| 1992 | Efficient channel coding for CELP using source information
W. Bastiaan Kleijn, Rafid A. Sukkar |
Speech Commun. | 1 |
| 1991 | Continuous representations in linear predictive codingabstractA major source of audible distortion in current low-bit-rate speech coding algorithms is an inaccurate degree of periodicity of the voiced speech signal. If the correlations between neighboring pitch cycles are accurately reproduced, these audible distortions can be reduced significantly. To this purpose, a novel method of coding voiced speech is introduced, which transmits an encoded prototype waveform at 20-30 ms intervals. The prototype waveform describes a pitch cycle representative for the interval, and is quantized using analysis-by-synthesis methods. The speech signal is reconstructed by concatenation of interpolated prototype waveforms. The short-term and the long-term correlations between pitch cycles can be controlled explicitly. Unquantized reconstructed speech is virtually indistinguishable from the original signal. The method results in excellent speech quality at rates between 3.0 and 4.0 kb/s.> W. Bastiaan Kleijn |
ICASSP | 1 |
| 1991 | Acoustic to articulatory parameter mapping using an assembly of neural networksabstractThe authors describe an efficient procedure for acoustic-to-articulatory parameter mapping using neural networks. An assembly of multilayer perceptrons, each designated to a specific region in the articulatory space, is used to map acoustic parameters of the speech into tract areas. The training of this model is executed in two stages; in the first stage a codebook of suitably normalized articulatory parameters is used and in the second stage real speech data are used to further improve the mapping. In general, acoustic-to-articulatory parameter mapping is nonunique; several vocal tract shapes can result in identical spectral envelopes. The model accommodates this ambiguity. During synthesis, neural networks are selected by dynamic programming using a criterion that ensures smoothly varying vocal tract shapes while maintaining a good spectral match.> Mazin G. Rahim, W. Bastiaan Kleijn, Juergen Schroeter, Colin C. Goodyear |
ICASSP | 2 |
| 1990 | Source-dependent channel coding for CELPabstractSource-dependent channel-error codes for speech compression algorithms are obtained by minimizing an appropriate speech distortion criterion under channel-error conditions. A simulated annealing algorithm for optimizing such a criterion is described. The resulting channel codes are efficient since they provide error correction of nonuniform accuracy (highly probable quantization levels receive more accurate correction) and/or nonuniform error detection (serious errors are more likely to be detected). An optimal tradeoff between error correction and error detection (associated with a recovery procedure) is obtained. Any desired fraction of the codewords can be allocated for error protection, including noninteger bit allocations. The performance of the source-dependent channel-coding technique is reported for the code-excited linear prediction (CELP) algorithm.> W. Bastiaan Kleijn |
ICASSP | 1 |
| 1989 | Robust CELP coders for noisy backgrounds and noisy channelsabstractThe authors examine the robustness of the code-excited linear predictive (CELP) coder, such as its ability to cope with nonspeech and corrupted speech inputs or to survive errors in the transmission of the coder parameters. They describe how they determined the error sensitivity of each coder parameter and identified the error propagation mechanisms. They find that the coder is robust to many kinds of input signals but is sensitive to channel errors. They show that the effect of bit errors can be reduced significantly with relatively simple measures such as repetition of the coder parameters from the most recent error-free frame. Incorporating these techniques into CELP results in a robust system with an acceptable performance for both burst and random bit errors rates of up to 1%. The clear channel performance degrades slightly (0.2 dB) as a result of these techniques.> Richard V. Cox, W. Bastiaan Kleijn, Peter Kroon |
ICASSP | 2 |
| 1988 | Improved speech quality and efficient vector quantization in SELPabstractA SELP (stochastically excited linear prediction) algorithm consisting of a two-stage vector quantization using an adaptive codebook and a stochastic codebook is described. The adaptive codebook quantization is similar to a closed-loop long-term prediction filter if the predictor delay is more than one frame length. The performance of the adaptive codebook procedure is improved by extending its codebook to include candidate vectors constructed from past synthetic excitation displaying a high level of periodicity. An algorithm is introduced which through increased symmetry of the error criterion significantly reduces the computational effort required for the search through the adaptive codebook. The algorithm can also be used for stochastic codebooks consisting of overlapping candidate vectors. It is shown that for stochastic codebooks in which neighboring candidates overlap for all but two samples, the quantization performance is as high as for codebooks containing fully independent candidate vectors.> W. Bastiaan Kleijn, Daniel J. Krasinski, Richard H. Ketchum |
ICASSP | 1 |
| 1988 | An efficient stochastically excited linear predictive coding algorithm for high quality low bit rate transmission of speech
W. Bastiaan Kleijn, Daniel J. Krasinski, Richard H. Ketchum |
Speech Commun. | 1 |
| 1987 | Harmonic coding of speech at 4.8 kb/sabstractThis paper describes a new speech coding technique which yields improved speech quality over existing 2.4 kb/s LPC vocoders. The method is computationally efficient and operates at a data rate of 4.8 kb/s. Each speech frame is initially classified as voiced or unvoiced. Unvoiced frames are synthesized using a linear predictive coding filter with noise or multipulse excitation. Voiced frames are synthesized using a sum of sinusoids. The frequency of each sinusoid is defined by peaks in the frequency spectrum. A new interpolation technique provides a computationally efficient method of locating the spectral peaks. A real-time, fully quantized version has been implemented in hardware. Edward C. Bronson, Douglas A. Carlone, W. Bastiaan Kleijn, Kevin M. O'Dell, Joseph Picone, David L. Thomson |
ICASSP | 3 |