Zoran Cvetkovic

dblp:76/1907 · DBLP profile ↗
← Back
62ranked-venue papers
18as first author
19since 2021 · last 2026
0000-0002-5128-5099ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 39 · 12 first-author · 11 since 2021Artificial intelligence and machine learning · 23 · 1 first-author · 12 since 2021Databases, data management, data science and information retrieval · 4 · 4 first-authorTheory of computation · 4 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 2Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2026 Raw acoustic-articulatory multimodal dysarthric speech recognition
abstract
Automatic speech recognition (ASR) for dysarthric speech is challenging. The acoustic characteristics of dysarthric speech are highly variable and there are often fewer distinguishing cues between phonetic tokens. Multimodal ASR utilises the data from other modalities to facilitate the task when a single acoustic modality proves insufficient. Articulatory information, which encapsulates knowledge about the speech production process, may constitute such a complementary modality. Although multimodal acoustic-articulatory ASR has received increasing attention recently, incorporating real articulatory data is under-explored for dysarthric speech recognition. This paper investigates the effectiveness of multimodal acoustic modelling using real dysarthric speech articulatory information in combination with acoustic features, especially raw signal representations which are more informative than classic features, leading to learning representations tailored to dysarthric ASR. In particular, various raw acoustic-articulatory multimodal dysarthric speech recognition systems are developed and compared with similar systems with hand-crafted features. Furthermore, the difference between dysarthric and typical speech in terms of articulatory information is systematically analysed by using a statistical space distribution indicator called Maximum Articulator Motion Range (MAMR). Additionally, we used mutual information analysis to investigate the robustness and phonetic information content of the articulatory features, offering insights that support feature selection and the ASR results. Experimental results on the widely used TORGO dysarthric speech dataset show that combining the articulatory and raw acoustic features at the empirically found optimal fusion level achieves a notable performance gain, leading to up to 7.6% and 12.8% relative word error rate (WER) reduction for dysarthric and typical speech, respectively.
Zhengjun Yue, Erfan Loweimi, Zoran Cvetkovic, Jon Barker, Heidi Christensen
Comput. Speech Lang.3
2025 INFR-GC: Interpretable Feature Representations for Granger Causality in Cortico-muscular Interactions
abstract
Understanding the interactions between the central nervous system and muscular responses is essential for developing effective strategies to diagnose and manage movement disorders such as dystonia. This study addresses these complex interactions by introducing a novel non-linear forecasting method for time series data. We propose that mutual information, by detecting complex dependencies between time series, can uncover hidden relationships suggestive of Granger causality, thereby enhancing the scope and precision of causality analysis. Our approach emphasizes the selection of the most informative features for predicting the target variable through iterative extraction and evaluation. We employ an optimized gradient-boosted random forest algorithm, prioritizing features with the highest mutual information relative to the target variable. Additionally, a Granger causality metric, tailored for non-linear models, is developed to quantify the strength of the discovered interactions. Experimental validation on real physiological data demonstrates the effectiveness of our method in uncovering causal relationships and assessing feature importance, contributing to a deeper understanding of movement control mechanisms.
Farwa Abbas, Verity M. McClelland, Wei Dai 0001, Zoran Cvetkovic
ICASSP4
2025 On the Design of a Robust Superdirective Beamformer and Topology Parameter Optimization with Frustum-Shaped Microphone Arrays Featuring Multiple Rings
Kunlong Zhao, Gongping Huang, Jingdong Chen, Jacob Benesty, Zoran Cvetkovic
INTERSPEECH6
2025 White-box Differentiable Model of Perceived Localisation
abstract
Auditory models are useful tools for estimating perceptual attributes of a sound field. Integrating such auditory models in the optimisation of immersive sound systems is a promising strategy when listeners’ perception is central to the application. To that end, differentiability is key to allowing the perceptual model to be included in gradient-based optimisation loops. Existing differentiable models, however, are black-box deep-learning based, which limits their interpretability. In this paper, we propose an analytical white-box differentiable model of auditory localisation based on an existing non-differential model. Our evaluations show that the model produces outputs that are highly correlated with the outputs of the non-differential model and data collected in subjective listening tests. The proposed model also enables optimisation of amplitude panning laws in a stereophonic spatial sound field rendering through gradient descent. This study therefore demonstrates, more generally, the feasibility of designing and optimising immersive sound systems using white-box differentiable models of auditory perception.
Antoine R. Souchaud, Pedro Lladó, Annika Neidhardt, Zoran Cvetkovic, Enzo De Sena
MMSP4
2023 SS-ADMM: Stationary and Sparse Granger Causal Discovery for Cortico-Muscular Coupling
abstract
Cortico-muscular communication patterns reveal important information about motor control. However, inferring significant causal relationships between motor cortex electroencephalogram (EEG) and surface electromyogram (sEMG) of concurrently active muscles is challenging since relevant processes involved in muscle control are relatively weak compared to additive noise and background activities. In this paper, a framework for identification of cortico-muscular linear time invariant communication is proposed that simultaneously estimates model order and its parameters by enforcing sparsity and stationarity conditions in a convex optimization program. The experimental results demonstrate that our proposed algorithm outperforms existing techniques for autoregressive model estimation, in terms of computational speed and model identification for causality estimation.
Farwa Abbas, Verity M. McClelland, Zoran Cvetkovic, Wei Dai 0001
ICASSP3
2023 Structured Errors-in-Variables Modelling for Cortico-Muscular Coherence Enhancement
abstract
Functional coupling between the cortex and muscle is commonly quantified by cortico-muscular coherence (CMC) between electroencephalogram (EEG) and electromyogram (EMG) signals. However, the presence of noise in EEG and EMG often degrades CMC, making it challenging to detect: some healthy subjects with good motor skills show no significant CMC. This study proposes an approach based on structured errors-in-variables (EIV) modelling to estimate components of the cortex and muscle signals involved in movement control from noisy EEG and EMG signals for the purpose of coherence estimation. We describe three algorithms to identify the underlying EIV system: one based on total least squares; the other two on structured total least squares, in which the Toeplitz data matrix structure is preserved. The effectiveness of the proposed method is assessed using simulated and neurophysiological data, where it achieved considerable improvements in coherence levels.
Zhenghao Guo, Verity M. McClelland, Wei Dai 0001, Zoran Cvetkovic
ICASSP4
2023 Dysarthric Speech Recognition, Detection and Classification using Raw Phase and Magnitude Spectra
abstract
In this paper, we explore the effectiveness of deploying the raw phase and magnitude spectra for dysarthric speech recognition, detection and classification. In particular, we scrutinise the usefulness of various raw phase-based representations along with their combinations with the raw magnitude spectrum and filterbank features. We employed single and multi-stream architectures consisting of a cascade of convolutional, recurrent and fully-connected layers for acoustic modelling. Furthermore, we investigate various configurations and fusion schemes as well as their training dynamics. In addition, the accuracies of the raw phase and magnitude based systems in the detection and classification tasks are studied and discussed. We report the performance on the UASpeech and TORGO dysarthric speech databases and for different severity levels. Our best system achieved WERs of 31.2% and 9.1% for dysarthric and typical speech on TORGO and 30.2% on UASpeech, respectively.
Zhengjun Yue, Erfan Loweimi, Zoran Cvetkovic
INTERSPEECH3
2023 3D Perceptual Soundfield Reconstruction via Virtual Microphone Synthesis
abstract
Perceptual soundfield reconstruction (PSR) is a multichannel audio recording and reproduction framework based on time-intensity panning in the horizontal plane. A practical limitation of PSR is that the optimal directivity patterns required by the system cannot be trivially and precisely obtained in practice, and it is limited to the horizontal plane. This paper extends the horizontal PSR to three dimensions and proposes a virtual microphone synthesis approach to obtain the PSR directivity pattern via sound field extrapolation. The proposed 3D extension and virtual microphone synthesis are evaluated using numerical simulations and a subjective localisation test. Comparisons with second-order Ambisonics rendering indicate that subjects localise sources rendered using 3D PSR more accurately and also with a higher certainty, particularly at an off-centre listening position for the low-channel count reproduction system employed.
Ege Erdem, Zoran Cvetkovic, Hüseyin Hacihabiboglu
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Phonetic Error Analysis Beyond Phone Error Rate
abstract
In this paper, we analyse the performance of the TIMIT-based phone recognition systems beyond the overall phone error rate (PER) metric. We consider three broad phonetic classes (BPCs): {affricate, diphthong, fricative, nasal, plosive, semi-vowel, vowel, silence}, {consonant, vowel, silence} and {voiced, unvoiced, silence} and, calculate the contribution of each phonetic class in terms of the substitution, deletion, insertion and PER. Furthermore, for each BPC we investigate the following: evolution of PER during training, effect of noise (NTIMIT), importance of different spectral subbands (1, 2, 4, and 8 kHz), usefulness of bidirectional vs unidirectional sequential modelling, transfer learning from WSJ and regularisation via monophones. In addition, we construct a confusion matrix for each BPC and analyse the confusions via dimensionality reduction to 2D at the input (acoustic features) and output (logits) levels of the acoustic model. We also compare the performance and confusion matrices of the BLSTM-based hybrid baseline system with those of the GMM-HMM based hybrid, Conformer and wav2vec 2.0 based end-to-end phone recognisers. Finally, the relationship of the unweighted and weighted PERs with the broad phonetic class priors is studied for both the hybrid and end-to-end systems.
Erfan Loweimi, Andrea Carmantini, Peter Bell 0001, Steve Renals, Zoran Cvetkovic
IEEE ACM Trans. Audio Speech Lang. Process.5
2023 Multi-Stream Acoustic Modelling Using Raw Real and Imaginary Parts of the Fourier Transform
abstract
In this paper, we investigate multi-stream acoustic modelling using the raw real and imaginary parts of the Fourier transform of speech signals. Using the raw magnitude spectrum, or features derived from it, as a proxy for the real and imaginary parts leads to irreversible information loss and suboptimal information fusion. We discuss and quantify the importance of such information in terms of speech quality and intelligibility. In the proposed framework, the real and imaginary parts are treated as two streams of information, pre-processed via separate convolutional networks, and then combined at an optimal level of abstraction, followed by further post-processing via recurrent and fully-connected layers. The optimal level of information fusion in various architectures, training dynamics in terms of cross-entropy loss, frame classification accuracy and WER as well as the shape and properties of the filters learned in the first convolutional layer of single- and multi-stream models are analysed. We investigated the effectiveness of the proposed systems in various tasks: TIMIT/NTIMIT (phone recognition), Aurora-4 (noise robustness), WSJ (read speech), AMI (meeting) and TORGO (dysarthric speech). Across all tasks we achieved competitive performance: in Aurora-4, down to 4.6% WER on average, in WSJ down to 4.6% and 6.2% WERs for Eval-92 and Eval-93, for Dev/Eval sets of the AMI-IHM down to 23.3%/23.8% WERs and in the AMI-SDM down to 43.7%/47.6% WERs have been achieved. In TORGO, for dysarthric and typical speech we achieved down to 31.7% and 10.2% WERs, respectively.
Erfan Loweimi, Zhengjun Yue, Peter Bell 0001, Steve Renals, Zoran Cvetkovic
IEEE ACM Trans. Audio Speech Lang. Process.5
2022 Raw Source and Filter Modelling for Dysarthric Speech Recognition
abstract
Acoustic modelling for automatic dysarthric speech recognition (ADSR) is a challenging task. Data deficiency is a major problem and substantial differences between the typical and dysarthric speech complicates transfer learning. In this paper, we build acoustic models using the raw magnitude spectra of the source and filter components. The proposed multi-stream model consists of convolutional and recurrent layers. It allows for fusing the vocal tract and excitation components at different levels of abstraction and after per-stream pre-processing. We show that such a multi-stream processing leverages these two information streams and helps s model towards normalising the speaker attributes and speaking style. This potentially leads to better handling of the dysarthric speech with a large inter-speaker and intra-speaker variability. We compare the proposed system with various features, study the training dynamics, explore usefulness of the data augmentation and provide interpretation for the learned convolutional filters. On the widely used TORGO dysarthric speech corpus, the proposed approach results in up to 1.7% absolute WER reduction for dysarthric speech compared with the MFCC base-line. Our best model reaches up to 40.6% and 11.8% WER for dysarthric and typical speech, respectively.
Zhengjun Yue, Erfan Loweimi, Zoran Cvetkovic
ICASSP3
2022 Multi-Modal Acoustic-Articulatory Feature Fusion For Dysarthric Speech Recognition
abstract
Building automatic speech recognition (ASR) systems for speakers with dysarthria is a very challenging task. Although multi-modal ASR has received increasing attention recently, incorporating real articulatory data with acoustic features has not been widely explored in the dysarthric speech community. This paper investigates the effectiveness of multi-modal acoustic modelling for dysarthric speech recognition using acoustic features along with articulatory information. The proposed multi-stream architectures consist of convolutional, recurrent and fully-connected layers allowing for bespoke per-stream pre-processing, fusion at the optimal level of abstraction and post-processing. We study the optimal fusion level/scheme as well as training dynamics in terms of cross-entropy and WER using the popular TORGO dysarthric speech database. Experimental results show that fusing the acoustic and articulatory features at the empirically found optimal level of abstraction achieves a remarkable performance gain, leading to up to 4.6% absolute (9.6% relative) WER reduction for speakers with dysarthria.
Zhengjun Yue, Erfan Loweimi, Zoran Cvetkovic, Heidi Christensen, Jon Barker
ICASSP3
2022 Dysarthric Speech Recognition From Raw Waveform with Parametric CNNs
abstract
Raw waveform acoustic modelling has recently received increasing attention. Compared with the task-blind hand-crafted features which may discard useful information, representations directly learned from the raw waveform are task-specific and potentially include all task-relevant information. In the context of automatic dysarthric speech recognition (ADSR), raw waveform acoustic modelling is under-explored owing to data scarcity. Parametric convolutional neural networks (CNNs) can compensate for this problem due to having notably fewer parameters and requiring less training data in comparison with conventional non-parametric CNNs. In this paper, we explore the usefulness of raw waveform acoustic modelling using various parametric CNNs for ADSR. We investigate the properties of the learned filters and monitor the training dynamics of various models. Furthermore, we study the effectiveness of data augmentation and multi-stream acoustic modelling through combining the non-parametric and parametric CNNs fed by hand-crafted and raw waveform features. Experimental results on the TORGO dysarthric database show that the parametric CNNs significantly outperform the non-parametric CNNs, reaching up to 36.2% and 12.6% WERs (up to 3.4% and 1.1% absolute error reduction) for dysarthric and typical speech, respectively. Multi-stream acoustic modelling further improves the performance resulting in up to 33.2% and 10.3% WERs for dysarthric and typical speech, respectively.
Zhengjun Yue, Erfan Loweimi, Heidi Christensen, Jon Barker, Zoran Cvetkovic
INTERSPEECH5
2022 Scattering Delay Network Simulator of Coupled Volume Acoustics
abstract
Artificial reverberators provide a computationally viable alternative to full-scale room acoustics simulation methods for deployment in interactive, immersive systems. Scattering delay network (SDN) is an artificial reverberator that allows direct parametric control over the geometry of a simulated cuboid enclosure, as well as the directional characteristics of the simulated sound sources and microphones. This paper extends the concept of SDN reverberators to multiple enclosures coupled via an aperture. The extension allows independent control of the acoustical properties of the coupled enclosures and the size of the connecting aperture. Transfer functions of the coupled-volume SDN are derived. The effectiveness of the proposed method is evaluated in terms of rendered energy decay curves in comparison to full-scale ray-tracing models and scale model measurements.
Timuçin Berk Atalay, Zühre Sü Gül, Enzo De Sena, Zoran Cvetkovic, Hüseyin Hacihabiboglu
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 Towards Robust Waveform-Based Acoustic Models
abstract
We study the problem of learning robust acoustic models in adverse environments, characterized by a significant mismatch between training and test conditions. This problem is of paramount importance for the deployment of speech recognition systems that need to perform well in unseen environments. First, we characterize data augmentation theoretically as an instance of vicinal risk minimization, which aims at improving risk estimates during training by replacing the delta functions that define the empirical density over the input space with an approximation of the marginal population density in the vicinity of the training samples. More specifically, we assume that local neighborhoods centered at training samples can be approximated using a mixture of Gaussians, and demonstrate theoretically that this can incorporate robust inductive bias into the learning process. We then specify the individual mixture components implicitly via data augmentation schemes, designed to address common sources of spurious correlations in acoustic models. To avoid potential confounding effects on robustness due to information loss, which has been associated with standard feature extraction techniques (e.g.,fbankandmfccfeatures), we focus on the waveform-based setting. Our empirical results show that the approach can generalize to unseen noise conditions, with 150% relative improvement in out-of-distribution generalization compared to training using the standard risk minimization principle. Moreover, the results demonstrate competitive performance relative to models learned using a training sample designed to match the acoustic conditions characteristic of test utterances.
Dino Oglic, Zoran Cvetkovic, Peter Sollich, Steve Renals, Bin Yu 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Acoustic Modelling From Raw Source and Filter Components for Dysarthric Speech Recognition
abstract
Acoustic modelling for automatic dysarthric speech recognition (ADSR) is a challenging task. Data deficiency is a major problem and substantial differences between typical and dysarthric speech complicate the transfer learning. In this paper, we aim at building acoustic models using the raw magnitude spectra of the source and filter components for ADSR. The proposed multi-stream models consist of convolutional, recurrent and fully-connected layers allowing for pre-processing various information streams and fusing them at an optimal level of abstraction. We demonstrate that such a multi-stream processing leverages information encoded in the vocal tract and excitation components and leads to normalising nuisance factors such as speaker attributes and speaking style. This leads to a better handling of dysarthric speech that exhibits large inter- and intra-speaker variabilities and results in a notable performance gain. Furthermore, we analyse the learned convolutional filters and visualise the outputs of different layers after dimensionality reduction to demonstrate how the speaker-related attributes are normalised along the pipeline. We also compare the proposed multi-stream model with various systems based on MFCC, FBank, raw waveform and i-vector, and, study the training dynamics as well as usefulness of the feature normalisation and data augmentation via speed perturbation. On the widely used TORGO and UASpeech dysarthric speech corpora, the proposed approach leads to a competitive performance of up to 35.3% and 30.3% WERs for dysarthric speech, respectively.
Zhengjun Yue, Erfan Loweimi, Heidi Christensen, Jon Barker, Zoran Cvetkovic
IEEE ACM Trans. Audio Speech Lang. Process.5
2021 Speech Acoustic Modelling from Raw Phase Spectrum
abstract
Magnitude spectrum-based features are the most widely employed front-ends for acoustic modelling in automatic speech recognition (ASR) systems. In this paper, we investigate the possibility and efficacy of acoustic modelling using the raw short-time phase spectrum. In particular, we study the usefulness of the raw wrapped, unwrapped and minimum-phase phase spectra as well as the phase of the source and filter components for acoustic modelling. Furthermore, we explore the effectiveness of simultaneous deployment of the vocal tract and excitation components of the raw phase spectrum using multi-head CNNs and investigate multiple information fusion schemes. This paves the way for developing an effective phase-based multi-stream information processing systems for speech recognition. The performance, even for wrapped phase with a noise-like shape, is comparable to or better than the magnitude-based classic features, and up to 4.8% WER has been achieved in the WSJ (Eval-92) task.
Erfan Loweimi, Zoran Cvetkovic, Peter Bell 0001, Steve Renals
ICASSP2
2021 Speech Acoustic Modelling Using Raw Source and Filter Components
abstract
Source-filter modelling is among the fundamental techniques in speech processing with a wide range of applications. In acoustic modelling, features such as MFCC and PLP which parametrise the filter component are widely employed. In this paper, we investigate the efficacy of building acoustic models from the raw filter and source components. The raw magnitude spectrum, as the primary information stream, is decomposed into the excitation and vocal tract information streams via cepstral liftering. Then, acoustic models are built via multi-head CNNs which, among others, allow for processing each individual stream via a sequence of bespoke transforms and fusing them at an optimal level of abstraction. We discuss the possible advantages of such information factorisation and recombination, investigate the dynamics of these models and explore the optimal fusion level. Furthermore, we illustrate the CNN’s learned filters and provide some interpretation for the captured patterns. The proposed approach with optimal fusion scheme results in up to 14% and 7% relative WER reduction in WSJ and Aurora-4 tasks.
Erfan Loweimi, Zoran Cvetkovic, Peter Bell 0001, Steve Renals
Interspeech2
2021 Learning Waveform-Based Acoustic Models Using Deep Variational Convolutional Neural Networks
abstract
We investigate the potential of stochastic neural networks for learning effective waveform-based acoustic models. The waveform-based setting, inherent to fully end-to-end speech recognition systems, is motivated by several comparative studies of automatic and human speech recognition that associate standard non-adaptive feature extraction techniques with information loss, which can adversely affect robustness. Stochastic neural networks, on the other hand, are a class of models capable of incorporating rich regularization mechanisms into the learning process. We consider a deep convolutional neural network that first decomposes speech into frequency sub-bands via an adaptive parametric convolutional block where filters are specified by cosine modulations of compactly supported windows. The network then employs standard non-parametric 1D convolutions to extract relevant spectro-temporal patterns while gradually compressing the structured high dimensional representation generated by the parametric block. We rely on a probabilistic parametrization of the proposed neural architecture and learn the model using stochastic variational inference. This requires evaluation of an analytically intractable integral defining the Kullback-Leibler divergence term responsible for regularization, for which we propose an effective approximation based on the Gauss-Hermite quadrature. Our empirical results demonstrate a superior performance of the proposed approach over comparable waveform-based baselines and indicate that it could lead to robustness. Moreover, the approach outperforms a recently proposed deep convolutional neural network for learning of robust acoustic models with standard FBANK features.
Dino Oglic, Zoran Cvetkovic, Peter Sollich
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Deep Scattering Power Spectrum Features for Robust Speech Recognition
abstract
Deep scattering spectrum consists of a cascade of wavelet transforms and modulus non-linearity. It generates features of different orders, with the first order coefficients approximately equal to the Mel-frequency cepstrum, and higher order coefficients recovering information lost at lower levels. We investigate the effect of including the information recovered by higher order coefficients on the robustness of speech recognition. To that end, we also propose a modification to the original scattering transform tailored for noisy speech. In particular, instead of the modulus non-linearity we opt to work with power coefficients and, therefore, use the squared modulus non-linearity. We quantify the robustness of scattering features using the word error rates of acoustic models trained on clean speech and evaluated using sets of utterances corrupted with different noise types. Our empirical results show that the second order scattering power spectrum coefficients capture invariants relevant for noise robustness and that this additional information improves generalization to unseen noise conditions (almost 20% relative error reduction on aurora 4). This finding can have important consequences on speech recognition systems that typically discard the second order information and keep only the first order features (known for emulating mfcc and fbank values) when representing speech.
Neethu M. Joy, Dino Oglic, Zoran Cvetkovic, Peter Bell 0001, Steve Renals
INTERSPEECH3
2020 A Deep 2D Convolutional Network for Waveform-Based Speech Recognition
abstract
Due to limited computational resources, acoustic models of early automatic speech recognition ( asr) systems were built in low-dimensional feature spaces that incur considerable information loss at the outset of the process. Several comparative studies of automatic and human speech recognition suggest that this information loss can adversely affect the robustness of asr systems. To mitigate that and allow for learning of robust models, we propose a deep 2 d convolutional network in the waveform domain. The first layer of the network decomposes waveforms into frequency sub-bands, thereby representing them in a structured high-dimensional space. This is achieved by means of a parametric convolutional block defined via cosine modulations of compactly supported windows. The next layer embeds the waveform in an even higher-dimensional space of high-resolution spectro-temporal patterns, implemented via a 2 d convolutional block. This is followed by a gradual compression phase that selects most relevant spectro-temporal patterns using wide-pass 2 d filtering. Our results show that the approach significantly outperforms alternative waveform-based models on both noisy and spontaneous conversational speech (24% and 11% relative error reduction, respectively). Moreover, this study provides empirical evidence that learning directly from the waveform domain could be more effective than learning using hand-crafted features.
Dino Oglic, Zoran Cvetkovic, Peter Bell 0001, Steve Renals
INTERSPEECH2
2020 Localization Uncertainty in Time-Amplitude Stereophonic Reproduction
abstract
This article studies the effects of inter-channel time and level differences in stereophonic reproduction on perceived localization uncertainty, which is defined as how difficult it is for a listener to tell where a sound source is located. Towards this end, a computational model of localization uncertainty is proposed first. The model calculates inter-aural time and level difference cues, and compares them to those associated to free-field point-like sources. The comparison is carried out using a particular distance functional that replicates the increased uncertainty observed experimentally with inconsistent inter-aural time and level difference cues. The model is validated by formal listening tests, achieving a Pearson correlation of 0.99. The model is then used to predict localization uncertainty for stereophonic setups and a listener in central and off-central positions. Results show that amplitude methods achieve a slightly lower localization uncertainty for a listener positioned exactly in the center of the sweet spot. As soon as the listener moves away from that position, the situation reverses, with time-amplitude methods achieving a lower localization uncertainty.
Enzo De Sena, Zoran Cvetkovic, Hüseyin Hacihabiboglu, Marc Moonen, Toon van Waterschoot
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 Perceptual Soundfield Reconstruction in Three Dimensions via Sound Field Extrapolation
abstract
Perceptual sound field reconstruction (PSR) is a spatial audio recording and reproduction method based on the application of stereophonic panning laws in microphone array design. PSR allows rendering a perceptually veridical and stable auditory perspective in the horizontal plane of the listener, and involves recording using near-coincident microphone arrays. This paper extends the PSR concept to three dimensions using sound field extrapolation carried out in the spherical-harmonic domain. Sound field rendering is performed using a two-level loudspeaker rig. An active-intensity-based analysis of the rendered sound field shows that the proposed approach can render direction of monochromatic plane waves accurately.
Ege Erdem, Enzo De Sena, Hüseyin Hacihabiboglu, Zoran Cvetkovic
ICASSP4
2019 Bilinear Dictionary Update via Linear Least Squares
abstract
Algorithms for dictionary learning aim to learn a dictionary under which training data have sparse representations. This paper addresses the dictionary update sub-problem, the goal of which is to update the dictionary and the corresponding sparse coefficients given a fixed sparsity pattern. It is a non-convex bilinear inverse problem, and hence challenging to solve. Inspired by a recent work by Ling and Strohmer, we re-formulate the dictionary update problem as a linear least squares problem, which is convex and easy to solve. Necessary bounds on the number of training samples required for a unique solution are derived when exact sparsity pattern is known. Further, for dictionary update with unknown sparsity patterns, an efficient iterative algorithm based on total least squares is developed. Embedding the new dictionary update procedure into an overall dictionary learning algorithm achieves better numerical performance compared to state of the art algorithms.
Wei Dai 0001, Zoran Cvetkovic, Jubo Zhu
ICASSP3
2018 Cortico-Muscular Coherence Enhancement Via Sparse Signal Representation
abstract
Identifiction of specific cortico-muscular interactions is essential for understanding sensorimotor control. These interactions are commonly studied by analyzing cortico-muscular coherence (CMC) between electroencephalogram (EEG) and surface electromyogram (sEMG) recorded synchronously under a motor control task. However, the presence of noise and components irrelevant to the monitored task weakens CMC so that it is often very difficult to detect. This study proposes an approach based on dictionary learning and sparse signal representation combined with a component selection algorithm to extract versions of EEG and sEMG signals which contain higher relative levels of coherent components. Evaluations using neurophysiological data show that the method achieves substantial increase in CMC levels.
Yuhang Xu 0001, Wei Dai 0001, Zoran Cvetkovic, Verity M. McClelland
ICASSP4
2016 Delay estimation between EEG and EMG via coherence with time lag
abstract
The traditional way to estimate the time delay between the motor cortex and the periphery is based on the estimation of the slope of the phase of the cross spectral density between motor cortex electroencephalogram (EEG) and electromyography (EMG) signals recorded synchronously during a motor control task. There are several issues that could make the delay estimation using this method subject to errors, leading frequently to estimates which are in disagreement with underlying physiology. This study introduces cortico-muscular coherence with time lag (CMCTL) function and proposes a method for estimating the delay based on finding its local maxima. We further address the issue of the interpretation of such time delay in multi-path propagation systems. Delay estimates obtained using the proposed method are more consistent compared with results obtained using the phase method and in a better agreement with physiological facts.
Yuhang Xu 0001, Verity M. McClelland, Zoran Cvetkovic, Kerry R. Mills
ICASSP3
2015 Efficient Synthesis of Room Acoustics via Scattering Delay Networks
abstract
An acoustic reverberator consisting of a network of delay lines connected via scattering junctions is proposed. All parameters of the reverberator are derived from physical properties of the enclosure it simulates. It allows for simulation of unequal and frequency-dependent wall absorption, as well as directional sources and microphones. The reverberator renders the first-order reflections exactly, while making progressively coarser approximations of higher-order reflections. The rate of energy decay is close to that obtained with the image method (IM) and consistent with the predictions of Sabine and Eyring equations. The time evolution of the normalized echo density, which was previously shown to be correlated with the perceived texture of reverberation, is also close to that of the IM. However, its computational complexity is one to two orders of magnitude lower, comparable to the computational complexity of a feedback delay network and its memory requirements are negligible.
Enzo De Sena, Hüseyin Hacihabiboglu, Zoran Cvetkovic, Julius O. Smith III
IEEE ACM Trans. Audio Speech Lang. Process.3
2013 A computational model for the estimation of localisation uncertainty
abstract
A computational model for prediction of localisation uncertainty of phantom auditory sources is proposed. The interaural level and time difference pairs due to point sources in free field are used as a reference. The mismatch between these “natural” pairs and interaural time and level difference pairs elicited by phantom sources is quantified by means of the 0.5-norm distance, which is justified on psychoacoustic grounds. The model is validated by results of subjective listening tests, achieving a high level of correlation with experimental data.
Enzo De Sena, Zoran Cvetkovic
ICASSP2
2013 Analysis and Design of Multichannel Systems for Perceptual Sound Field Reconstruction
abstract
This paper presents a systematic framework for the analysis and design of circular multichannel surround sound systems. Objective analysis based on the concept of active intensity fields shows that for stable rendition of monochromatic plane waves it is beneficial to render each such wave by no more than two channels. Based on that finding, we propose a methodology for the design of circular microphone arrays, in the same configuration as the corresponding loudspeaker system, which aims to capture inter-channel time and intensity differences that ensure accurate rendition of the auditory perspective. The methodology is applicable to regular and irregular microphone/speaker layouts, and a wide range of microphone array radii, including the special case of coincident arrays which corresponds to intensity-based systems. Several design examples, involving first and higher-order microphones are presented. Results of formal listening tests suggest that the proposed design methodology achieves a performance comparable to prior art in the center of the loudspeaker array and a more graceful degradation away from the center.
Enzo De Sena, Hüseyin Hacihabiboglu, Zoran Cvetkovic
IEEE Trans. Speech Audio Process.3
2013 On Frequency Offset Estimation for OFDM
abstract
This paper presents a comparative study of Schmidl-Cox (SC) and Morelli-Mengali (MM) algorithms for frequency offset estimation in OFDM, along with a new least squares (LS) and a new modified SC algorithm. All algorithms have comparable accuracy approaching asymptotically the Cramer-Rao bound. The complexity of the LS algorithm is between O(N) and O(N log N) operations, where N is the length of the training sequence, while the complexity of the SC algorithm is between O(N log N) and O(N2) operations, and the complexity of the MM algorithm is O(N2) operations. The modified version of the SC algorithm requires only one training sequence as opposed to two required by the original SC algorithm, and significantly reduced O(N log N) complexity. The sensitivity of the three algorithms to quantization of the arg function (the argument of a complex number) is analyzed and quantified. The analysis and simulation results demonstrate that while all considered algorithms can be used with coarse quantization of the arg function, the LS algorithm is least affected and the SC algorithm is most affected by this quantization error.
Zoran Cvetkovic, Vahid Tarokh, Seokho Yoon
IEEE Trans. Wirel. Commun.1
2012 Multichannel Dereverberation Theorems and Robustness Issues
abstract
Multichannel dereverberation amounts to the inversion of a multiple-input/multiple-output linear time-invariant system. In this paper, necessary and sufficient conditions for perfect dereverberation using stable and finite impulse response (FIR) filters are established. It is then shown that the inverse system given by the pseudoinverse of the original transfer function matrix exhibits a noise reduction property. A necessary and sufficient condition under which this pseudoinverse system is FIR is also given. Further, an FIR approximation to the pseudoinverse system is considered and the effects of the length of this approximation on the dereverberation accuracy are investigated. Finally, an analytical and numerical assessment of the dependence of the dereverberation accuracy on the accuracy of the acquisition of room impulse responses is provided.
Hüseyin Hacihabiboglu, Zoran Cvetkovic
IEEE Trans. Speech Audio Process.2
2012 On the Design and Implementation of Higher Order Differential Microphones
abstract
A novel systematic approach to the design of directivity patterns of higher order differential microphones is proposed. The directivity patterns are obtained by optimizing a cost function which is a convex combination of a front-back energy ratio and uniformity within a frontal sector of interest. Most of the standard directivity patterns—omnidirectional, cardioid, subcardioid, hypercardioid, supercardioid—are particular solutions of this optimization problem with specific values of two free parameters: the angular width of the frontal sector and the convex combination factor. More general solutions of practical use are obtained by varying these two parameters. Many of these optimal directivity patterns are trigonometric polynomials with complex roots. A new differential array structure that enables the implementation of general higher order directivity patterns, with complex or real roots, is then proposed. The effectiveness of the proposed design framework and the implementation structure are illustrated by design examples, simulations, and measurements.
Enzo De Sena, Hüseyin Hacihabiboglu, Zoran Cvetkovic
IEEE Trans. Speech Audio Process.3
2011 A generalized design method for directivity patterns of spherical microphone arrays
abstract
Spherical microphone arrays provide a flexible solution to obtaining higher-order directivity patterns, which are useful in audio recording and reproduction. A general systematic approach to the design of directivity patterns for spherical microphone arrays is introduced in this paper. The directivity patterns are obtained by optimizing a cost function which is a convex combination of a front-back energy ratio and a smoothness term. Most of the standard directivity patterns - i.e. omnidirectional, cardioid, subcardioid, hypercardioid and supercardioid - are particular solutions of this optimization problem with specific values of two free parameters: the angle of the frontal sector, and the convex combination factor. By varying these two parameters, more general solutions of practical use are obtained.
Enzo De Sena, Hüseyin Hacihabiboglu, Zoran Cvetkovic
ICASSP3
2011 Combined waveform-cepstral representation for robust speech recognition
abstract
High-dimensional acoustic waveform representations are studied as a front-end for noise robust automatic speech recognition using generative methods, in particular Gaussian mixture models and hidden Markov models. The proposed representations are compared with standard cepstral features on phoneme classification and recognition tasks. While lower error rates are achieved using cepstral features at very low noise levels, the acoustic waveform representations are much more robust to noise. A convex combination of acoustic waveforms and cepstral features is then considered and it achieves higher accuracy than either of the individual representations across all noise levels.
Matthew Ager, Zoran Cvetkovic, Peter Sollich
ISIT2
2011 Combined Features and Kernel Design for Noise Robust Phoneme Classification Using Support Vector Machines
abstract
This paper proposes methods for combining cepstral and acoustic waveform representations for a front-end of support vector machine (SVM)-based speech recognition systems that are robust to additive noise. The key issue of kernel design and noise adaptation for the acoustic waveform representation is addressed first. Cepstral and acoustic waveform representations are then compared on a phoneme classification task. Experiments show that the cepstral features achieve very good performance in low noise conditions, but suffer severe performance degradation already at moderate noise levels. Classification in the acoustic waveform domain, on the other hand, is less accurate in low noise but exhibits a more robust behavior in high noise conditions. A combination of the cepstral and acoustic waveform representations achieves better classification performance than either of the individual representations over the entire range of noise levels tested, down to$-$18-dB SNR.
Jibran Yousafzai, Peter Sollich, Zoran Cvetkovic, Bin Yu 0001
IEEE Trans. Speech Audio Process.3
2010 ARMA regularization of cardiac perfusion modeling
abstract
Cardiac perfusion modelling using ARMA systems is studied. ARMA is a generalization of a recently proposed exponential approximation technique, which was shown to exhibit better performance than the widely used truncated singular value decomposition method. Experiments demonstrate that ARMA achieves results as accurate as the those obtained using the exponential approximation, but it its at the same time less sensitive to additive noise and model order selection.
Philipp G. Batchelor, Amedeo Chiribiri, Niloufar Zarinabad Nooralipour, Zoran Cvetkovic
ICASSP4
2010 Towards robust phoneme classification with hybrid features
abstract
In this paper, we investigate the robustness of phoneme classification to additive noise with hybrid features using support vector machines (SVMs). In particular, the cepstral features are combined with short term energy features of acoustic waveform segments to form a hybrid representation. The energy features are then taken into account separately in the SVM kernel, and a simple subtraction method allows them to be adapted effectively in noise. This hybrid representation contributes significantly to the robustness of phoneme classification and narrows the performance gap to the ideal baseline of classifiers trained under matched noise conditions.
Jibran Yousafzai, Zoran Cvetkovic, Peter Sollich
ISIT2
2010 Subband acoustic waveform front-end for robust speech recognition using support vector machines
abstract
A subband acoustic waveform front-end for robust speech recognition using support vector machines (SVMs) is developed. The primary issues of kernel design for subband components of acoustic waveforms and combination of the individual subband classifiers using stacked generalization are addressed. Experiments performed on the TIMIT phoneme classification task demonstrate the benefits of classification in frequency subbands: the subband classifier outperforms the cepstral classifiers in the presence of noise for signal-to-noise ratio (SNR) below 12dB.
Jibran Yousafzai, Zoran Cvetkovic, Peter Sollich
SLT2
2010 Simulation of Directional Microphones in Digital Waveguide Mesh-Based Models of Room Acoustics
abstract
Digital waveguide mesh (DWM) models are time-domain numerical methods providing computationally simple solutions for wave propagation problems. They have been used in various acoustical modeling and audio synthesis applications including synthesis of musical instrument sounds and speech, and modeling of room acoustics. A successful model of room acoustics should be able to account for source and receiver directivity. Methods for the simulation of directional sources in DWM models were previously proposed. This paper presents a method for the simulation of directional microphones in DWM-based models of room acoustics. The method is based on the directional weighting of the microphone response according to the instantaneous direction of incidence at a given point. The direction of incidence is obtained from instantaneous intensity that is calculated from local pressure values in the DWM model. The calculation of instantaneous intensity in DWM meshes and the directional accuracies of different mesh topologies are discussed. An intensity-based formulation for the response of a directional microphone is given. Simulation results for an actual microphone with frequency-dependent, non-ideal directivity function are presented.
Hüseyin Hacihabiboglu, Banu Gunel, Zoran Cvetkovic
IEEE Trans. Speech Audio Process.3
2009 Tuning support vector machines for robust phoneme classification with acoustic waveforms
abstract
This work focuses on the robustness of phoneme classification to additive noise in the acoustic waveform domain using support vector machines (SVMs). We address the issue of designing kernels for acoustic waveforms which imitate the state-of-the-art representations such as PLP and MFCC and are tuned to the physical properties of speech. For comparison, classification results in the PLP representation domain with cepstral mean-and-variance normalization (CMVN) using standard kernels are also reported. It is shown that our custom-designed kernels achieve better classification performance at high noise. Finally, we combine the PLP and acoustic waveform representations to attain better classification than either of the individual representations over the entire range of noise levels tested, from quiet condition up to - 18dB SNR.
Jibran Yousafzai, Zoran Cvetkovic, Peter Sollich
INTERSPEECH2
2007 An Efficient Multichannel Equalization Algorithm for Audio Applications
abstract
The challenge of multichannel equalization for audio applications lies in the physical properties of the underlying multi-input/multi-output (MIMO) linear time-invariant systems which are generally non-minimum phase and exhibit extremely long impulse responses, thereby imposing a considerable computational burden on the equalization task particularly when iterative solutions are sought. In this paper we propose a computationally efficient non-iterative multi-channel equalization algorithm. The proposed algorithm is based on the fast Fourier transform (FFT) and allows for faster and considerably more accurate inversion of MIMO systems compared to traditional deconvolution algorithms and adaptive solutions. We address the accuracy and limitations of the proposed algorithm and present simulation results illustrating its performance.
Jibran Yousafzai, Zoran Cvetkovic
ICASSP (1)2
2007 Single-Bit Oversampled A/D Conversion With Exponential Accuracy in the Bit Rate
abstract
A scheme for simple oversampled analog-to-digital (A/D) conversion using single-bit quantization is presented. The scheme is based on recording positions of zero-crossings of the input signal added to a deterministic dither function. This information can be represented in a manner such that the bit rate increases only logarithmically with the oversampling factor$r$. The input band-limited signal can be reconstructed from this information locally with$O(1/r)$pointwise error, resulting in an exponentially decaying distortion-rate characteristic.
Zoran Cvetkovic, Ingrid Daubechies, Benjamin F. Logan
IEEE Trans. Inf. Theory1
2006 Separation of Audio Signals Into Direct and Diffuse Soundfields for Surround Sound
abstract
This paper presents a new method for reproduction of music in multichannel audio systems. The proposed method separates signals from individual channels into their direct and diffuse components which are then sent to different speaker elements. The direct components are sent directly to the listener, while the diffuse components are additionally scattered. The purpose of this scattering of diffuse components is twofold: first, it eliminates spurious localization cues which may be created by reproducing the sound using a small number of speakers (three to five), and second, it provides additional diffusion which improves the envelopment experience. We investigate two methods to separate direct and diffuse soundfield components, both of which assume the knowledge of the impulse response of the performance auditorium to the corresponding microphones. One method is based on techniques for multichannel equalization, while the other uses techniques for signal reconstruction after oversampled filter-bank processing. The latter method turns out to be computationally more manageable. In listening tests, subjects preferred music reproduction which separates direct and diffuse soundfields to reproduction in which both soundfields are sent to the same speaker elements.
Benjamin S. Olswang, Zoran Cvetkovic
ICASSP (5)2
2006 Locally adaptive wavelet-based image interpolation
abstract
We describe a spatially adaptive algorithm for image interpolation. The algorithm uses a wavelet transform to extract information about sharp variations in the low-resolution image and then implicitly applies interpolation which adapts to the image local smoothness/singularity characteristics. The proposed algorithm yields images that are sharper compared to several other methods that we have considered in this paper. Better performance comes at the expense of higher complexity.
S. Grace Chang, Zoran Cvetkovic, Martin Vetterli
IEEE Trans. Image Process.2
2004 Frequency synchronization in OFDM
abstract
We present an OFDM frequency synchronization scheme. The scheme uses periodic OFDM symbols, similar to the algorithms proposed previously by Morelli and Mengali (1999) and Schmidl and Cox (1997). The proposed scheme attains considerably higher accuracy than the scheme by Schmidl and Cox requiring a similar computational load. Compared to the scheme by Morelli and Mengali, the proposed algorithm attains a somewhat inferior accuracy but at a significantly reduced computational complexity, i.e O(N) versus O(N/sup 2/) operations for N-tone OFDM. In addition to that, the scheme proposed here is considerably less sensitive to the accuracy of the involved computations than the other two schemes.
Zoran Cvetkovic, Vahid Tarokh, Seokho Yoon
ICASSP (4)1
2003 Nonuniform oversampled filter banks for audio signal processing
abstract
In emerging audio technology applications, there is a need for decompositions of audio signals into oversampled subband components with time-frequency resolution which mimics that of the cochlear filter bank and with high aliasing attenuation in each of the subbands independently, rather than aliasing cancellation properties. We present a design of nearly perfect reconstruction nonuniform oversampled filter banks which implement signal decompositions of this kind.
Zoran Cvetkovic, James D. Johnston
IEEE Trans. Speech Audio Process.1
2003 Resilience properties of redundant expansions under additive noise and quantization
abstract
Representing signals using coarsely quantized coefficients of redundant expansions is an interesting source coding paradigm, the most important practical case of which is oversampled analog-to-digital (A/D) conversion. Signal reconstruction from quantized redundant expansions and the accuracy of such representations are problems which are not well understood and we study them in this paper for uniform scalar quantization in finite-dimensional spaces. To give a more global perspective, we first present an analysis of the resilience of redundant expansions to degradation by additive noise in general, and then focus on the effects of uniform scalar quantization. The accuracy of signal representations obtained by applying uniform scalar quantization to coefficients of redundant expansions, measured as the mean-squared Euclidean norm of the reconstruction error, has been previously shown to be lower-bounded by an 1/r/sup 2/ expression. We establish some general conditions under which the 1/r/sup 2/ accuracy can actually be attained, and under those conditions prove a 1/r/sup 2/ upper error bound. For a particular kind of structured expansions, which includes many popular frame classes, we propose reconstruction algorithms which attain the 1/r/sup 2/ accuracy at low numerical complexity. These structured expansions, moreover, facilitate efficient encoding of quantized coefficients in a manner which requires only a logarithmic bit-rate increase in redundancy, resulting in an exponential error decay in the bit rate. Results presented in this paper are immediately applicable to oversampled A/D conversion of periodic bandlimited signals.
Zoran Cvetkovic
IEEE Trans. Inf. Theory1
2002 Interpolation of Bandlimited Functions from Quantized Irregular Samples
abstract
The problem of reconstructing a /spl pi/-bandlimited signal f from its quantized samples taken at an irregular sequence of points (t/sub k/)/sub k/spl isin//spl Zopf// arises in oversampled analog-to-digital conversion. The input signal can be reconstructed from the quantized samples (f(t/sub k/))/sub k/spl isin//spl Zopf// by estimating samples (f(n//spl lambda/))/sub n/spl isin//spl Zopf//, where /spl lambda/ is the average uniform density of the sequence (tk)/sub k/spl isin//spl Zopf//, assumed here to be greater than one, followed by linear low-pass filtering. We study three techniques for estimating samples (f(n//spl lambda/))/sub n/spl isin//spl Zopf// from quantized irregular samples (f(t/sub k/))/sub k/spl isin//spl Zopf//, including Lagrangian interpolation, and two other techniques which result in a better overall accuracy of oversampled A/D conversion.
Zoran Cvetkovic, Benjamin F. Logan, Ingrid Daubechies
DCC1
2002 Robust phoneme discrimination using acoustic waveforms
abstract
We present a study of separability of acoustic waveforms of speech at phoneme level. The analyzed data consist of 64ms segments of acoustic waveforms of individual phonemes from TIMIT data base, sampled at 16kHz. For each phoneme, by means of principal component analysis, we identify subspaces which contain a given proportion of the total energy of the available waveforms in time-domain, and also in spectral-magnitude domain. In order to assess separation between phonemes in the two domains, we perform pairwise classification of phonemes on clean data and on data immersed in white additive Gaussian noise up to 0dB signal to noise ratio. While the classification based on spectral magnitudes exhibits high sensitivity to additive noise, the time-domain classification proves to be very robust.
Zoran Cvetkovic, Baltasar Beferull-Lozano, Andreas Buja
ICASSP1
2001 On simple oversampled A/D conversion in L2(IR)
abstract
The accuracy of oversampled analog-to-digital (A/D) conversion, the dependence of accuracy on the sampling interval /spl tau/ and on the bit rate R are characteristics fundamental to A/D conversion but not completely understood. These characteristics are studied for oversampled A/D conversion of band-limited signals in L/sup 2/ (R). We show that the digital sequence obtained in the process of oversampled A/D conversion describes the corresponding analog signal with an error which tends to zero as /spl tau//sup 2/ in energy, provided that the quantization threshold crossings of the signal constitute a sequence of stable sampling in the respective space of band-limited functions. Further, we show that the sequence of quantized samples can be represented in a manner which requires only a logarithmic increase in the bit rate with the sampling frequency, R=O(|log/spl tau/|), and hence that the error of oversampled A/D conversion actually exhibits an exponential decay in the bit rate as the sampling interval tends to zero.
Zoran Cvetkovic, Martin Vetterli
IEEE Trans. Inf. Theory1
2000 Single-Bit Oversampled A/D Conversion with Exponential Accuracy in the Bit-Rate
abstract
We present a scheme for simple oversampled analog-to-digital conversion with single bit quantization and exponential error decay in the bit rate. The scheme is based on recording positions of zero-crossings of the input signal added to a deterministic dither function. This information can be represented in a manner which requires only logarithmic increase of the bit rate with the oversampling factor, r. The input-bandlimited signal can be reconstructed from this information locally, and with a mean squared error which is inversely proportional to the square of the oversampling factor, MSE=O(1/r/sup 2/). Consequently the mean squared error of this scheme exhibits exponential decay in the bit rate.
Zoran Cvetkovic, Ingrid Daubechies
Data Compression Conference1
2000 OFDM with biorthogonal demultiplexing
abstract
In this paper we study biorthogonal schemes for frequency division multiplexing. These are schemes which use one given set of orthogonal waveforms for frequency division multiplexing at the transmitter, and a different set of waveforms, which is orthogonal to the multiplexing waveforms, for demultiplexing at the receiver. The well known cyclic prefix scheme is in this category. This biorthogonal demultiplexing allows for more design freedom, and we investigate its impact on designing robust OFDM transmultiplexers, and focus on waveforms which extend over several OFDM symbol intervals.
Zoran Cvetkovic
ICASSP1
1999 Source Coding with Quantized Redundant Expansions: Accuracy and Reconstruction
abstract
Signal representations based on low-resolution quantization of redundant expansions is an interesting source coding paradigm, the most important practical case of which is oversampled A/D conversion. Signal reconstruction from quantized coefficients of a redundant expansion and accuracy of representations of this kind are problems which are still not well understood and these are studied in this paper in finite dimensional spaces. It has been previously proven that accuracy of signal representations based on quantized redundant expansions, measured as the squared Euclidean norm of the reconstruction error, cannot be better than O(1/(r/sup 2/)), where r is the expansion redundancy. We give some general conditions under which 1/(r/sup 2/) accuracy can be attained. We also suggest a form of structure for overcomplete families which facilitates reconstruction, and which enables efficient encoding of quantized coefficients with a logarithmic increase of the bit-rate in redundancy.
Zoran Cvetkovic
Data Compression Conference1
1999 Modulating waveforms for OFDM
abstract
Orthogonal frequency division multiplexing (OFDM) is a popular transmission technique that is employed in applications such as digital audio broadcasting, asymmetric digital subscriber line and wireless LAN. We consider design of modulating waveforms for OFDM in the presence of the delay spread and system impairments such as frequency offset and timing mismatch. We give a complete parameterization of OFDM modulating waveforms. Increasing robustness of OFDM to frequency offsets requires using long modulating waveforms. To make the implementation of OFDM systems with long modulating waveforms feasible we propose fast implementation algorithms. Some preliminary modulating waveform design examples are presented. The presented waveforms demonstrate that the robustness of OFDM systems to impairments can be improved by allowing certain degradation of unnecessarily good performances of the state of the art OFDM systems in ideal operating conditions.
Zoran Cvetkovic
ICASSP1
1998 Accurate Subband Coding with Low Resolution Quantization
abstract
Redundant signal expansions exhibit robustness to degradation by additive noise, at a degree which is proportional to expansion redundancy. This property can be exploited for obtaining accurate signal representations with low resolution quantization, and it has been used in oversampled A/D conversion. A problem with this approach to source coding is that the bit-rate increases linearly with expansion redundancy giving poor rate-distortion performances. Entropy coding schemes that could remedy the problem of fast increase of the bit-rate have not yet been studied in general. An example of a redundant expansion which is amenable to efficient encoding of quantized samples is discussed. A quantitative characterization of noise reduction properties of redundant expansions is given first. Then, a subband coding scheme based on overcomplete short-time Fourier expansions and low resolution quantization is discussed. This scheme allows for efficient lossless encoding of quantized expansion coefficients, with a logarithmic increase of the bit-rate with redundancy. Thus, it attains an exponential error decay in the bit-rate.
Zoran Cvetkovic
Data Compression Conference1
1998 Short-time Fourier analysis-a novel window design procedure
abstract
Weyl-Heisenberg frames are the tool for short-time Fourier analysis. These are generated from a prototype window function using translation on a rectangular grid in the time-frequency plane. Particularly appealing Weyl-Heisenberg frames are those which are tight as they allow for signal representations analogous to orthonormal expansions and have good numerical stability properties. Designing the window of a tight Weyl-Heisenberg frame requires optimization of the frequency characteristics of the window, usually some form of frequency selectivity, under a set of nonlinear constraints. For long windows this can be a formidable task, if not infeasible. We propose a new filter design method based on expansions with respect to prolate spheroidal sequences. The advantages of this new method are more and more pronounced as the redundancy of the frame increases. These advantages pertain to a reduction in computational complexity and the ability to describe good and long windows with a few parameters.
Zoran Cvetkovic
ICASSP1
1998 Error-Rate Characteristics of Oversampled Analog-to-Digital Conversion
abstract
Accuracy of simple analog-to-digital conversion depends on both resolution of discretization in amplitude and resolution of discretization in time. For implementation convenience, high conversion accuracy is attained by refining the discretization in time using oversampling. It is commonly believed that oversampling adversely impacts rate-distortion properties of the conversion, since the bit rate, B, increases linearly with oversampling, resulting in a slow error decay in the bit rate, on the order of O(1/B). We demonstrate that the information obtained in the process of oversampled analog-to-digital conversion can easily be encoded in a manner which requires only a logarithmic increase of the bit rate with redundancy, achieving an exponential error decay in the bit rate.
Zoran Cvetkovic, Martin Vetterli
IEEE Trans. Inf. Theory1
1996 Analysis of errors in quantization of Weyl-Heisenberg frame expansions and oversampled A/D conversion
abstract
Manifestations of robustness of overcomplete expansions, acquired as a result of the redundancy, have been observed for two primary sources of degradation: white additive noise and quantization. These two cases require different treatments since quantization error has a certain structure which is obscured by statistical analysis. The central issue of this paper is a deterministic approach to quantization error analysis, with a focus on redundant Weyl-Heisenberg expansions in L/sup 2/(R). This analysis demonstrates that overcomplete expansions exhibit a higher degree of robustness to quantization error than to a white additive noise, and indicates how this additional robustness can be exploited for improvement of representation accuracy. Further, it shows that under reasonable assumptions quantization error in Weyl-Heisenberg frame expansions can be reduced as /spl par/e/spl par//sup 2/=O(1/r/sup 2/), where r is the frame redundancy factor.
Zoran Cvetkovic
ICASSP1
1996 Oversampled FIR filter banks and frames in l2(Z)
abstract
Perfect reconstruction FIR filter banks are equivalent to a particular class of frames in l/sup 2/(Z). These frames are the subject of this paper. Necessary and sufficient conditions on a filter bank to implement a frame or a tight frame decomposition are given, as well as the necessary and sufficient condition for perfect reconstruction using FIR filters. Complete parameterizations of FIR filter banks satisfying these conditions are also given. Further, we study the condition under which the frame dual to the frame associated to an FIR filter bank is also FIR, and give a parameterization of a class of filter banks having this property.
Martin Vetterli, Zoran Cvetkovic
ICASSP2
1995 Resolution enhancement of images using wavelet transform extrema extrapolation
abstract
One problem of image interpolation refers to magnifying a small image without loss in image clarity. We propose a wavelet based method which estimates the higher resolution information needed to sharpen the image. This method extrapolates the wavelet transform of the higher resolution based on the evolution of the wavelet transform extrema across the scales. By identifying three constraints that the higher resolution information needs to obey, we enhance the reconstructed image through alternating projections onto the sets defined by these constraints.
S. Grace Chang, Zoran Cvetkovic, Martin Vetterli
ICASSP2
1995 Oversampled modulated filter banks and tight Gabor frames in l2(Z)
abstract
The subject of this study is paraunitary modulated filter banks. A factorization of the polyphase matrices of these filter banks, which is described, gives complete characterization of tight Gabor frames in l/sup 2/(Z), with arbitrary rational oversampling ratios. Tight Gabor frames, being less constrained than orthogonal bases, allow for filter bank designs with good localization in both time and frequency.
Zoran Cvetkovic
ICASSP1
1994 Wavelet extrema and zero-crossings representations: properties and consistent reconstruction
abstract
Properties of nonsubsampled FIR filter banks, used for wavelet extrema and zero-crossings representations, are investigated. Conditions under which iterated nonsubsampled filter banks implement frame or tight frame operators are given and relations with continuous time domain are established. Algorithms for consistent reconstruction of signals from wavelet extrema or zero crossings representations are proposed. These algorithms are characterized by a low computational complexity and a simple implementation. It is also shown that the wavelet extrema representation of a signal and the wavelet zero-crossings representation of its difference provide equivalent information on the signal.>
Zoran Cvetkovic, Martin Vetterli
ICASSP (3)1