EDBT 2026 Demo / reviewers in the wild / expert
Saikat Chatterjee
dblp:86/5806
· DBLP profile ↗
81ranked-venue papers
16as first author
19since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 63 · 14 first-author · 13 since 2021Artificial intelligence and machine learning · 19 · 5 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021Theory of computation · 4Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Selection of Layers from Self-supervised Learning Models for Predicting Mean-Opinion-Score of SpeechabstractSelf-supervised learning (SSL) models like Wav2Vec2, HuBERT, and WavLM have been widely used in speech processing. These transformer-based models consist of multiple layers, each capturing different levels of representation. While prior studies explored their layer-wise representations for efficiency and performance, speech quality assessment (SQA) models predominantly rely on last-layer features, leaving intermediate layers underexamined. In this work, we systematically evaluate different layers of multiple SSL models for predicting mean-opinion-score (MOS). Features from each layer are fed into a lightweight regression network to assess effectiveness. Our experiments consistently show early-layers features outperform or match those from the last layer, leading to significant improvements over conventional approaches and state-of-the-art MOS prediction models. These findings highlight the advantages of early-layer selection, offering enhanced performance and reduced system complexity. Fredrik Cumlin, Victor Ungureanu, Chandan K. A. Reddy, Christian Schüldt, Saikat Chatterjee |
ASRU | 6 |
| 2025 | Impairments are Clustered in Latents of Deep Neural Network-based Speech Quality ModelsabstractIn this article, we provide an experimental observation: Deep neural network (DNN) based speech quality assessment (SQA) models have inherent latent representations where many types of impairments are clustered. While DNN-based SQA models are not trained for impairment classification, our experiments show good impairment classification results in an appropriate SQA latent representation. We investigate the clustering of impairments using various kinds of audio degradations that include different types of noises, waveform clipping, gain transition, pitch shift, compression, reverberation, etc. To visualize the clusters we perform classification of impairments in the SQA-latent representation domain using a standard k-nearest neighbor (kNN) classifier. We also develop a new DNN-based SQA model, named DNSMOS+, to examine whether an improvement in SQA leads to an improvement in impairment classification. The classification accuracy is 94% for LibriAugmented dataset with 16 types of impairments and 54% for ESC-50 dataset with 50 types of real noises. Fredrik Cumlin, Victor Ungureanu, Chandan K. A. Reddy, Christian Schüldt, Saikat Chatterjee |
ICASSP | 6 |
| 2025 | Near-Field ISAC in 6G: Addressing Phase Nonlinearity via Lifted Super-ResolutionabstractIntegrated sensing and communications (ISAC) is a promising component of 6G networks, fusing communication and radar technologies to facilitate new services. Additionally, the use of extremely large-scale antenna arrays (ELAA) at the ISAC common receiver not only facilitates terahertz-rate communication links but also significantly enhances the accuracy of target detection in radar applications. In practical scenarios, communication scatterers and radar targets often reside in close proximity to the ISAC receiver. This, combined with the use of ELAA, fundamentally alters the electromagnetic characteristics of wireless and radar channels, shifting from far-field planar-wave propagation to near-field spherical wave propagation. Under the far-field planar-wave model, the phase of the array response vector varies linearly with the antenna index. In contrast, in the near-field spherical wave model, this phase relationship becomes nonlinear. This shift presents a fundamental challenge: the widely-used Fourier analysis can no longer be directly applied for target detection and communication channel estimation at the ISAC common receiver. In this work, we propose a feasible solution to address this fundamental issue. Specifically, we demonstrate that there exists a high-dimensional space in which the phase nonlinearity can be expressed as linear. Leveraging this insight, we develop a lifted super-resolution framework that simultaneously performs communication channel estimation and extracts target parameters with high precision. Sajad Daei, Amirreza Zamani, Saikat Chatterjee, Mikael Skoglund, Gábor Fodor 0001 |
ICASSP | 3 |
| 2025 | Particle-based Data-driven Nonlinear State Estimation of Model-free Process from Nonlinear MeasurementsabstractWe consider the problem of causal filtering of a model-free process from (noisy) nonlinear measurements. The ‘model-free process’ means that we do not have a state-space model (SSM) of the process dynamics, limiting the use of traditional model-driven filters, such as unscented Kalman filter (UKF) and particle filter (PF). To address the problem we propose a particle-based data-driven nonlinear state estimation (pDANSE) method. In pDANSE, a recurrent neural network (RNN) provides the statistical parameters of a Gaussian prior of the underlying state, and particles are then drawn from the prior to compute the posterior moments. pDANSE is typically trained in a semi-supervised fashion. For our experiments we study the use of half-wave rectification as a nonlinear transformation of measurements. We first show that an unsupervised learning-based method under-performs, and subsequently the semi-supervised learning-based pDANSE performs satisfactorily. Using Lorenz-63 system as benchmark, pDANSE is found to be competitive against a model-driven PF that knows the exact SSM. Anubhab Ghosh, Yonina C. Eldar, Saikat Chatterjee |
ICASSP | 3 |
| 2025 | iDANSE: Iterative Data-driven Nonlinear State Estimation of Model-free Hidden SequencesabstractWe introduce a model-free hidden sequence (MHS) estimation problem where the task is to estimate a long sequence of ‘model-free’ process hidden under additive Gaussian noise. To estimate the posterior of the hidden sequence from the noisy observation sequence, we have three main challenges: (a) the process is ‘model-free’, that means there is no underlying state-space-model (SSM), limiting the use of traditional SSM-informed model-based methods like Kalman filter (KF) and particle filter (PF); (b) only an unlabelled dataset comprised of noisy observation sequences is available as training data, and hence no supervised learning possible; and, (c) the use of sequential ancestral sampling based methods, like dynamical variational autoencoders (DVAEs), results in prohibitive complexity. To address the challenges we adopt a recently established recurrent neural network (RNN)-based method called DANSE (data-driven nonlinear state estimation). We develop an iterative DANSE (iDANSE) in unsupervised learning setup where a set of neural networks (NNs) are used iteratively. The set of NNs refines MHS estimation over the iterations. Using simulations, the performance of iDANSE is demonstrated for a benchmark Lorenz process (a chaotic attractor) and compared with SSM-informed extended KF (EKF) and unscented KF (UKF). Hang Qin, Anubhab Ghosh, Saikat Chatterjee |
ICASSP | 3 |
| 2025 | Enhancing Network Calibration for Low-Cost Gas Sensor Networks Through Adaptive Similarity SearchabstractIoT-based low-cost gas sensors networks are important for environmental monitoring, but their regular calibrations are needed to achieve acceptable sensing performance. A critical step in network calibration is identifying when sensors within the network are sensing the same phenomenon, which is essential for accurate calibration. In this paper, we propose an adaptive similarity-search-based method for detecting these periods of similarity under the assumption of linear sensor drift. Our method leverages the relationships between neighboring sensors’ measurements to enhance calibration accuracy, outperforming the commonly used Pearson correlation approach. We validate the effectiveness of our method through experiments with both synthetic data and real-world CO2sensor networks, demonstrating improved calibration accuracy and reliability. Saikat Chatterjee, Tobias J. Oechtering |
ICASSP | 2 |
| 2025 | Multivariate Probabilistic Assessment of Speech Quality
Fredrik Cumlin, Victor Ungureanu, Chandan K. A. Reddy, Christian Schüldt, Saikat Chatterjee |
INTERSPEECH | 6 |
| 2025 | BiGSM: Bayesian inference of gene regulatory network via sparse modellingabstractMOTIVATION: Inference of gene regulatory network (GRN) is challenging due to the inherent sparsity of the GRN matrix and noisy expression data, often leading to a high possibility of false positive or negative predictions. To address this, it is essential to leverage the sparsity of the GRN matrix and develop a robust method capable of handling varying levels of noise in the data. Moreover, most existing GRN inference methods produce only fixed point estimates, which lack the flexibility and informativeness for comprehensive network analysis. In contrast, a Bayesian approach that yields closed-form posterior distributions allows probabilistic link selection, offering insights into the statistical confidence of each possible link. Consequently, it is important to engineer a Bayesian GRN inference method and rigorously execute a benchmark evaluation compared to state-of-the-art methods. RESULTS: We propose a method-Bayesian inference of GRN via Sparse Modelling (BiGSM). BiGSM effectively exploits the sparsity of the GRN matrix and infers the posterior distributions of GRN links from noisy expression data by using the maximum likelihood based learning. We thoroughly benchmarked BiGSM using biological and simulated datasets including GeneNetWeaver, GeneSPIDER, and GRNbenchmark. The benchmark test evaluates its accuracy and robustness across varying noise levels and data models. Using point-estimate based performance measures, BiGSM provides an overall best performance in comparison with several state-of-the-art methods including GENIE3, LASSO, LSCON, and Zscore. Additionally, BiGSM is the only method in the set of competing methods that provides posteriors for the GRN weights, helping to decipher confidence across predictions. AVAILABILITY AND IMPLEMENTATION: Code implemented via MATLAB and Python are available at Github: https://github.com/SachLab/BiGSM and archived at zenodo. Hang Qin, Mateusz Garbulowski, Erik L. L. Sonnhammer, Saikat Chatterjee |
Bioinform. | 4 |
| 2024 | DNSMOS Pro: A Reduced-Size DNN for Probabilistic MOS of Speech
Fredrik Cumlin, Victor Ungureanu, Chandan K. A. Reddy, Christian Schüldt, Saikat Chatterjee |
INTERSPEECH | 6 |
| 2024 | IMU-based Online Multi-lidar CalibrationabstractModern autonomous systems typically use several sensors for perception. For best performance, accurate and reliable extrinsic calibration is necessary. In this research, we propose a reliable technique for the extrinsic calibration of several lidars on a vehicle without the need for odometry estimation or fiducial markers. First, our method generates an initial guess of the extrinsics by matching the raw signals of IMUs co-located with each lidar. This initial guess is then used in ICP and point cloud feature matching which refines and verifies this estimate. Furthermore, we can use observability criteria to choose a subset of the IMU measurements that have the highest mutual information — rather than comparing all the readings. We have successfully validated our methodology using data gathered from Scania test vehicles. Sandipan Das, Bengt Boberg, Maurice Fallon, Saikat Chatterjee |
IV | 4 |
| 2023 | DeePMOS: Deep Posterior Mean-Opinion-Score of Speech
Fredrik Cumlin, Christian Schüldt, Saikat Chatterjee |
INTERSPEECH | 4 |
| 2023 | M-LIO: Multi-lidar, multi-IMU odometry with sensor dropout toleranceabstractWe present a robust system for state estimation that fuses measurements from multiple lidars and inertial sensors with GNSS data. To initiate the method, we use the prior GNSS pose information. We then perform motion estimation in real-time, which produces robust motion estimates in a global frame by fusing lidar and IMU signals with GNSS translation components using a factor graph framework. We also propose methods to account for signal loss with a novel synchronization and fusion mechanism. To validate our approach extensive tests were carried out on data collected using Scania test vehicles (5 sequences for a total of ≈ 7 Km). From our evaluations, we show an average improvement of 61% in relative translation and 42% rotational error compared to a state-of-the-art estimator fusing a single lidar/inertial sensor pair, in sensor dropout scenarios. Sandipan Das, Navid Mahabadi, Maurice Fallon, Saikat Chatterjee |
IV | 4 |
| 2022 | Deterministic Transform Based Weight Matrices for Neural NetworksabstractWe propose to use deterministic transforms as weight matrices for several feedforward neural networks. The use of deterministic transforms helps to reduce the computational complexity in two ways: (1) matrix-vector product complexity in forward pass, helping real time complexity, and (2) fully avoiding backpropagation in the training stage. For each layer of a feedforward network, we pro-pose two unsupervised methods to choose the most appropriate deterministic transform from a set of transforms (a bag of well-known transforms). Experimental results show that the use of deterministic transforms is as good as traditional random matrices in the sense of providing similar classification performance. Pol Grau Jurado, Saikat Chatterjee |
ICASSP | 3 |
| 2022 | Extrinsic Calibration and Verification of Multiple Non-overlapping Field of View Lidar SensorsabstractWe demonstrate a multi-lidar calibration frame-work for large mobile platforms that jointly calibrate the extrinsic parameters of non-overlapping Field-of-View (FoV) lidar sensors, without the need for any external calibration aid. The method starts by estimating the pose of each lidar in its corresponding sensor frame in between subsequent timestamps. Since the pose estimates from the lidars are not necessarily synchronous, we first align the poses using a Dual Quaternion (DQ) based Screw Linear Interpolation. Afterward, a Hand-Eye based calibration problem is solved using the DQ-based formulation to recover the extrinsics. Furthermore, we verify the extrinsics by matching chosen lidar semantic features, obtained by projecting the lidar data into the camera perspective after time alignment using vehicle kinematics. Experimental results on the data collected from a Scania vehicle [~ 1 Km sequence] demonstrate the ability of our approach to obtain better calibration parameters than the provided vehicle CAD model calibration parameters. This setup can also be scaled to any combination of multiple lidars. Sandipan Das, Navid Mahabadi, Addi Djikic, Cesar Nassir, Saikat Chatterjee, Maurice Fallon |
ICRA | 5 |
| 2022 | Neural Greedy Pursuit for Feature SelectionabstractWe propose a greedy algorithm to select$N$important features among$P$input features for a non-linear prediction problem. The features are selected one by one sequentially, in an iterative loss minimization procedure. We use neural networks as predictors in the algorithm to compute the loss and hence, we refer to our method as neural greedy pursuit (NGP). NGP is efficient in selecting$N$features when$N\ll P$, and it provides a notion of feature importance in a descending order following the sequential selection procedure. We experimentally show that NGP provides better performance than several feature selection methods such as DeepLIFT and Drop-one-out loss. In addition, we experimentally show a phase transition behavior in which perfect selection of all$N$features without false positives is possible when the training data size exceeds a threshold. Sandipan Das, Alireza M. Javid, Prakash B. Gohain, Yonina C. Eldar, Saikat Chatterjee |
IJCNN | 5 |
| 2021 | A ReLU Dense Layer to Improve the Performance of Neural NetworksabstractWe propose ReDense as a simple and low complexity way to improve the performance of trained neural networks. We use a combination of random weights and rectified linear unit (ReLU) activation function to add a ReLU dense (ReDense) layer to the trained neural network such that it can achieve a lower training loss. The lossless flow property (LFP) of ReLU is the key to achieve the lower training loss while keeping the generalization error small. ReDense does not suffer from vanishing gradient problem in the training due to having a shallow structure. We experimentally show that ReDense can improve the training and testing performance of various neural network architectures with different optimization loss and activation functions. Finally, we test ReDense on some of the state-of-the-art architectures and show the performance improvement on benchmark datasets. Alireza M. Javid, Sandipan Das, Mikael Skoglund, Saikat Chatterjee |
ICASSP | 4 |
| 2021 | Feature Reuse for a Randomization Based Neural NetworkabstractWe propose a feature reuse approach for an existing multi-layer randomization based feedforward neural network. The feature representation is directly linked among all the necessary hidden layers. For the feature reuse at a particular layer, we concatenate features from the previous layers to construct a large-dimensional feature for the layer. The large-dimensional concatenated feature is then efficiently used to learn a limited number of parameters by solving a convex optimization problem. Experiments show that the proposed model improves the performance in comparison with the original neural network without a significant increase in computational complexity. Mikael Skoglund, Saikat Chatterjee |
ICASSP | 3 |
| 2021 | Detecting Signal Corruptions in Voice Recordings For Speech TherapyabstractIn this article we design an experimental setup to detect disturbances in voice recordings, such as additive noise, clipping, infrasound and random muting. The datasets are generated by introducing degradations into clean recordings. We test five different classification algorithms in both single- and multi-label settings: kernel substitution based support vector machine, convolutional neural network, long short-term memory (LSTM), and a hidden Markov model using either Gaussian mixture models or generative models in its state distribution. The LSTM achieved good results in both tests, most notably in the multi-label case where the average balanced accuracy was 82.7% on one dataset. Helmer Nylén, Saikat Chatterjee, Sten Ternström |
ICASSP | 2 |
| 2021 | Asynchronous Decentralized Learning of Randomization-Based Neural NetworksabstractIn a communication network, decentralized learning refers to the knowledge collaboration between the different local agents (processing nodes) to improve the local estimation performance without sharing private data. The ideal case is that the decentralized solution approximates the centralized solution, as if all the data are available at a single node, and requires low computational power and communication overhead. In this work, we propose a decentralized learning of randomization-based neural networks with asynchronous communication and achieve centralized equivalent performance. We propose an ARock-based alternating-direction-method-of-multipliers (ADMM) algorithm that enables individual node activation and one-sided communication in an undirected connected network, characterized by a doubly-stochastic network policy matrix. Besides, the proposed algorithm reduces the computational cost and communication overhead due to its asynchronous nature. We study the proposed algorithm on different randomization-based neural networks, including ELM, SSFN, RVFL, and its variants, to achieve the centralized equivalent performance under efficient computation and communication costs. We also show that the proposed asynchronous decentralized learning algorithm can outperform a synchronous learning algorithm regarding computational complexity” especially when the network connections are sparse. Alireza M. Javid, Mikael Skoglund, Saikat Chatterjee |
IJCNN | 4 |
| 2020 | Dual sentence representation model integrating prior knowledge for bio-text-miningabstractData mining, especially the extraction of the relationship between genes and proteins, plays an important role in the biomedical field. Several related models have been proposed for data mining in the biomedical domain. Furthermore, manually curated biomedical knowledge bases, which could assist the task, have been used to enhance the data-mining model. However, due to the limitation of methods, much prior knowledge information is not be fully exploited. In this work, we propose a novel method that reasonably applied the curated prior knowledge for biomedical text mining by dual sentence representation models; one model is for the experimental data and the other one is for the prior knowledge information sentence. We evaluated our method on two community-supported datasets; BioNLP and BioCreative corpora. The experimental results demonstrate that the dual sentence representation model can successfully utilize external prior knowledge information to extract relationship from biomedical text. Our method can achieve state-of-art results and it could be an application of biomedical relation extraction in the future. Zhijing Li 0005, YangYang Lan, Saikat Chatterjee, Pargorn Puttapirat, Xiangrong Zhang, Chen Li 0011 |
BIBM | 3 |
| 2020 | Powering Hidden Markov Model by Neural Network Based Generative ModelsabstractHidden Markov model (HMM) has been successfully used for sequential data modeling problems. In this work, we propose to power the modeling capacity of HMM by bringing in neural network based generative models. The proposed model is termed as GenHMM. In the proposed GenHMM, each HMM hidden state is associated with a neural network based generative model that has tractability of exact likelihood and provides efficient likelihood computation. A generative model in GenHMM consists of mixture of generators that are realized by flow models. A learning algorithm for GenHMM is proposed in expectation-maximization framework. The convergence of the learning GenHMM is analyzed. We demonstrate the efficiency of GenHMM by classification tasks on practical sequential data. Code available at this https URL. Dong Liu 0009, Antoine Honoré, Saikat Chatterjee, Lars K. Rasmussen |
ECAI | 3 |
| 2020 | Hidden Markov Models for Sepsis Detection in Preterm InfantsabstractWe explore the use of traditional and contemporary hidden Markov models (HMMs) for sequential physiological data analysis and sepsis prediction in preterm infants. We investigate the use of classical Gaussian mixture model based HMM, and a recently proposed neural network based HMM. To improve the neural network based HMM, we propose a discriminative training approach. Experimental results show the potential of HMMs over logistic regression, support vector machine and extreme learning machine. Antoine Honoré, Dong Liu 0009, David Forsberg, Karen Coste, Eric Herlenius, Saikat Chatterjee, Mikael Skoglund |
ICASSP | 6 |
| 2020 | High-Dimensional Neural Feature Using Rectified Linear Unit And Random Matrix InstanceabstractWe design a ReLU-based multilayer neural network to generate a rich high-dimensional feature vector. The feature guarantees a monotonically decreasing training cost as the number of layers increases. We design the weight matrix in each layer to extend the feature vectors to a higher dimensional space while providing a richer representation in the sense of training cost. Linear projection to the target in the higher dimensional space leads to a lower training cost if a convex cost is minimized. An ℓ2-norm convex constraint is used in the minimization to improve the generalization error and avoid overfitting. The regularization hyperparameters of the network are derived analytically to guarantee a monotonic decrement of the training cost and therefore, it eliminates the need for cross-validation to find the regularization hyperparameter in each layer. Alireza M. Javid, Arun Venkitaraman, Mikael Skoglund, Saikat Chatterjee |
ICASSP | 4 |
| 2020 | Asynchrounous Decentralized Learning of a Neural NetworkabstractIn this work, we exploit an asynchronous computing framework namely ARock to learn a deep neural network called self-size estimating feedforward neural network (SSFN) in a decentralized scenario. Using this algorithm namely asynchronous decentralized SSFN (dSSFN), we provide the centralized equivalent solution under certain technical assumptions. Asynchronous dSSFN relaxes the communication bottleneck by allowing one node activation and one side communication, which reduces the communication overhead significantly, consequently increasing the learning speed. We compare asynchronous dSSFN with traditional synchronous dSSFN in the experimental results, which shows the competitive performance of asynchronous dSSFN, especially when the communication network is sparse. Alireza M. Javid, Mikael Skoglund, Saikat Chatterjee |
ICASSP | 4 |
| 2020 | Gaussian Processes Over GraphsabstractKernel Regression over Graphs (KRG) was recently proposed for predicting graph signals in a supervised learning setting, where the inputs are agnostic to the graph. KRG model predicts targets that are smooth graph signals as over the given graph, given the input when all the signals are deterministic. In this work, we consider the development of a stochastic or Bayesian variant of KRG. Using priors and likelihood functions, our goal is to systematically derive a predictive distribution for the smooth graph signal target given the training data and a new input. We show that this naturally results in a Gaussian process formulation which we call Gaussian Processes over Graphs (GPG). Experiments with real-world datasets show that the performance of GPG is superior to a conventional Gaussian Process (without the graph-structure) for small training data sizes and under noisy training. Arun Venkitaraman, Saikat Chatterjee, Peter Händel |
ICASSP | 2 |
| 2020 | Recursive Prediction of Graph Signals With Incoming NodesabstractKernel and linear regression have been recently explored in the prediction of graph signals as the output, given arbitrary input signals that are agnostic to the graph. In many real-world problems, the graph expands over time as new nodes get introduced. Keeping this premise in mind, we propose a method to recursively obtain the optimal prediction or regression coefficients for the recently proposed Linear Regression over Graphs (LRG), as the graph expands with incoming nodes. This comes as a natural consequence of the structure of the regression problem, and obviates the need to solve a new regression problem each time a new node is added. Experiments with real-world graph signals show that our approach results in a good prediction performance which tends to be close to that obtained from knowing the entire graph apriori. Arun Venkitaraman, Saikat Chatterjee, Bo Wahlberg |
ICASSP | 2 |
| 2020 | Neural Network based Explicit Mixture Models and Expectation-maximization based LearningabstractWe propose two neural network based mixture models in this work. The proposed mixture models are explicit. The explicit models have analytical forms with the advantages of computing likelihood and efficiency of generating samples. Expectation-maximization based algorithms are developed for learning parameters of the proposed models. We provide sufficient conditions to realize the expectation-maximization based learning. The main requirements are invertibility of neural networks that are used as generators and Jacobian computation of functional form of the neural networks. The requirements are practically realized using a flow-based neural network. In our first mixture model, we use multiple flow-based neural networks as generators. Naturally the model is complex. A single latent variable is used as the common input to all the neural networks. The second mixture model uses a single flow-based neural network as a generator to reduce complexity. The single generator has a latent variable input that follows a Gaussian mixture distribution. The proposed models are verified via training with expectation-maximization based algorithms on practical datasets. We demonstrate efficiency of proposed mixture models through extensive experiments for generating samples and maximum likelihood based classification. Dong Liu 0009, Minh Thành Vu, Saikat Chatterjee, Lars K. Rasmussen |
IJCNN | 3 |
| 2020 | A Low Complexity Decentralized Neural Net with Centralized Equivalence using Layer-wise LearningabstractWe design a low complexity decentralized learning algorithm to train a recently proposed large neural network in distributed processing nodes (workers). We assume the communication network between the workers is synchronized and can be modeled as a doubly-stochastic mixing matrix without having any master node. In our setup, the training data is distributed among the workers but is not shared in the training process due to privacy and security concerns. Using alternating-direction-method-of-multipliers (ADMM) along with a layer-wise convex optimization approach, we propose a decentralized learning algorithm which enjoys low computational complexity and communication cost among the workers. We show that it is possible to achieve equivalent learning performance as if the data is available in a single place. Finally, we experimentally illustrate the time complexity and convergence behavior of the algorithm. Alireza M. Javid, Mikael Skoglund, Saikat Chatterjee |
IJCNN | 4 |
| 2020 | Online Spatiotemporal Popularity Learning via Variational Bayes for Cooperative CachingabstractHerein, we focus on an end-to-end design of a proactive cooperative caching strategy for a multi-cell network. The design is challenging as it involves two interrelated problems: the ability to predict future content popularity and to meet network operation characteristics. To this end, we first formulate a cooperative content caching in order to optimize the aggregated network cost for delivering contents to users. An efficient proactive caching policy requires an accurate prediction of time-varying content popularity. Content popularity has temporal and spatial dependencies and therefore, we develop a probabilistic dynamical model for content popularity prediction by exploiting its spatiotemporal correlations. To achieve an accurate tracking and prediction of content popularity evolution, the proposed dynamical model is non-linear and incorporates non-Gaussian distributions. We use Variational Bayes (VB) approach for estimating the model parameters. The VB provides mathematical tractability. We then develop an online VB method that works with streaming data where content request arrives sequentially. Using extensive simulations study on a real-world dataset, we show that our online VB based dynamical model provides improved performance compared to conventional content caching policies. Sajad Mehrizi, Saikat Chatterjee, Symeon Chatzinotas, Björn Ottersten 0001 |
IEEE Trans. Commun. | 2 |
| 2019 | Entropy-regularized Optimal Transport Generative ModelsabstractWe investigate the use of entropy-regularized optimal transport (EOT) cost in developing generative models to learn implicit distributions. Two generative models are proposed. One uses EOT cost directly in an one-shot optimization problem and the other uses EOT cost iteratively in an adversarial game. The proposed generative models show improved performance over contemporary models on scores of sample based test. Dong Liu 0009, Minh Thành Vu, Saikat Chatterjee, Lars K. Rasmussen |
ICASSP | 3 |
| 2019 | Kernel Regression for Graph Signal Prediction in Presence of Sparse NoiseabstractIn presence of sparse noise we propose kernel regression for predicting output vectors which are smooth over a given graph. Sparse noise models the training outputs being corrupted either with missing samples or large perturbations. The presence of sparse noise is handled using appropriate use of ℓ1-norm along-with use of ℓ2-norm in a convex cost function. For optimization of the cost function, we propose an iteratively reweighted least-squares (IRLS) approach that is suitable for kernel substitution or kernel trick due to availability of a closed form solution. Simulations using real-world temperature data show efficacy of our proposed method, mainly for limited-size training datasets. Arun Venkitaraman, Pascal Frossard, Saikat Chatterjee |
ICASSP | 3 |
| 2019 | On Hilbert transform, analytic signal, and modulation analysis for signals over graphs
Arun Venkitaraman, Saikat Chatterjee, Peter Händel |
Signal Process. | 2 |
| 2019 | Estimate exchange over network is good for distributed hard thresholding pursuit
Ahmed Zaki, Partha P. Mitra, Lars K. Rasmussen, Saikat Chatterjee |
Signal Process. | 4 |
| 2018 | Distributed Large Neural Network with Centralized EquivalenceabstractIn this article, we develop a distributed algorithm for learning a large neural network that is deep and wide. We consider a scenario where the training dataset is not available in a single processing node, but distributed among several nodes. We show that a recently proposed large neural network architecture called progressive learning network (PLN) can be trained in a distributed setup with centralized equivalence. That means we would get the same result if the data be available in a single node. Using a distributed convex optimization method called alternating-direction-method-of-multipliers (ADMM), we perform training of PLN in the distributed setup. Alireza M. Javid, Mikael Skoglund, Saikat Chatterjee |
ICASSP | 4 |
| 2018 | Multi-Kernel Regression for Graph Signal ProcessingabstractWe develop a multi-kernel based regression method for graph signal processing where the target signal is assumed to be smooth over a graph. In multi-kernel regression, an effective kernel function is expressed as a linear combination of many basis kernel functions. We estimate the linear weights to learn the effective kernel function by appropriate regularization based on graph smoothness. We show that the resulting optimization problem is shown to be convex and propose an accelerated projected gradient descent based solution. Simulation results using real-world graph signals show efficiency of the multi-kernel based approach over a standard kernel based approach. Arun Venkitaraman, Saikat Chatterjee, Peter Händel |
ICASSP | 2 |
| 2017 | Generalized fusion algorithm for compressive sampling reconstruction and RIP-based analysis
Ahmed Zaki, Saikat Chatterjee, Lars K. Rasmussen |
Signal Process. | 2 |
| 2016 | Automatic Recognition of Social Roles Using Long Term Role Transitions in Small Group InteractionsabstractRecognition of social roles in small group interactions is challenging because of the presence of disfluency in speech, frequent overlaps between speakers, short speaker turns and the need for reliable data annotation. In this work, we consider the problem of recognizing four roles, namely Gatekeeper, Protagonist, Neutral, and Supporter in small group interactions in AMI corpus. In general, Gatekeeper and Protagonist roles occur less frequently compared to Neutral, and Supporter. In this work, we exploit role transitions across segments in a meeting by incorporating role transition probabilities and formulating the role recognition as a decoding problem over the sequence of segments in an interaction. Experiments are performed in a five fold cross validation setup using acoustic, lexical and structural features with precision, recall and F-score as the performance metrics. The results reveal that precision averaged across all folds and different feature combinations improves in the case of Gatekeeper and Protagonist by 13.64% and 12.75% when the role transition information is used which in turn improves the F-score for Gatekeeper by 6.58% while the F-scores for the rest of the roles do not change significantly. Gaurav Fotedar, Aditya Gaonkar P., Saikat Chatterjee, Prasanta Kumar Ghosh |
INTERSPEECH | 3 |
| 2016 | Alternating strategies with internal ADMM for low-rank matrix reconstruction
Kezhi Li, Martin Sundin, Cristian R. Rojas, Saikat Chatterjee, Magnus Jansson |
Signal Process. | 4 |
| 2016 | Analysis of Regularized LS Reconstruction and Random Matrix Ensembles in Compressed SensingabstractThe performance of regularized least-squares estimation in noisy compressed sensing is analyzed in the limit when the dimensions of the measurement matrix grow large. The sensing matrix is considered to be from a class of random ensembles that encloses as special cases standard Gaussian, row-orthogonal, geometric, and so-called T-orthogonal constructions. Source vectors that have non-uniform sparsity are included in the system model. Regularization based on ℓ-norm and leading to LASSO estimation, or basis pursuit denoising, is given the main emphasis in the analysis. Extensions to ℓ-norm and zero-norm regularization are also briefly discussed. The analysis is carried out using the replica method in conjunction with some novel matrix integration results. Numerical experiments for LASSO are provided to verify the accuracy of the analytical results. The numerical experiments show that for noisy compressed sensing, the standard Gaussian ensemble is a suboptimal choice for the measurement matrix. Orthogonal constructions provide a superior performance in all considered scenarios and are easier to implement in practical applications. It is also discovered that for non-uniform sparsity patterns, the T-orthogonal matrices can further improve the mean square error behavior of the reconstruction when the noise level is not too high. However, as the additive noise becomes more prominent in the system, the simple row-orthogonal measurement matrix appears to be the best choice out of the considered ensembles. Mikko Vehkaperä, Yoshiyuki Kabashima, Saikat Chatterjee |
IEEE Trans. Inf. Theory | 3 |
| 2015 | Greedy minimization of l1-norm with high empirical successabstractWe develop a greedy algorithm for the basis-pursuit problem. The algorithm is empirically found to provide the same solution as convex optimization based solvers. The method uses only a subset of the optimization variables in each iteration and iterates until an optimality condition is satisfied. In simulations, the algorithm converges faster than standard methods when the number of measurements is small and the number of variables large. Martin Sundin, Saikat Chatterjee, Magnus Jansson |
ICASSP | 2 |
| 2014 | Distributed quantization for compressed sensingabstractWe study distributed coding of compressed sensing (CS) measurements using vector quantizer (VQ). We develop a distributed framework for realizing optimized quantizer that enables encoding CS measurements of correlated sparse sources followed by joint decoding at a fusion center. The optimality of VQ encoder-decoder pairs is addressed by minimizing the sum of mean-square errors between the sparse sources and their reconstruction vectors at the fusion center. We derive a lower-bound on the end-to-end performance of the studied distributed system, and propose a practical encoder-decoder design through an iterative algorithm. Amirpasha Shirazinia, Saikat Chatterjee, Mikael Skoglund |
ICASSP | 2 |
| 2014 | Analysis of regularized LS reconstruction and random matrix ensembles in compressed sensingabstractPerformance of regularized least-squares estimation in noisy compressed sensing is studied in the limit when the problem dimensions grow large. The sensing matrix is sampled from the rotationally invariant ensemble that encloses as special cases the standard IID and row-orthogonal constructions. The analysis is carried out using the replica method in conjunction with some novel matrix integration results. The numerical experiments show that for noisy compressed sensing, the standard IID ensemble is a suboptimal choice for the measurement matrix. Orthogonal constructions provide a superior performance in all considered scenarios and are easier to implement in practice. Mikko Vehkaperä, Yoshiyuki Kabashima, Saikat Chatterjee |
ISIT | 3 |
| 2014 | A sparsity based preprocessing for noise robust speech recognitionabstractWe show a method to sparsify the speech input that improves the robustness of an automatic speech recognizer. The proposed scheme is added to the system as a preprocessing module prior to the acoustic feature extraction. The preprocessing module passes the input speech signal through a linear predictive (LP) analysis filter and enforces sparsity in the LP residue domain. The sparsified prediction residue finally is filtered to generate the speech signal for computing a sequence of conventional feature vectors used in automatic speech recognition (ASR). Using standard feature vectors, our experiments show that sparsification in LP residue domain improves robustness in ASR performance. Christos Koniaris, Saikat Chatterjee |
SLT | 2 |
| 2014 | SEK: sparsity exploiting k-mer-based estimation of bacterial community compositionabstractMOTIVATION: Estimation of bacterial community composition from a high-throughput sequenced sample is an important task in metagenomics applications. As the sample sequence data typically harbors reads of variable lengths and different levels of biological and technical noise, accurate statistical analysis of such data is challenging. Currently popular estimation methods are typically time-consuming in a desktop computing environment. RESULTS: Using sparsity enforcing methods from the general sparse signal processing field (such as compressed sensing), we derive a solution to the community composition estimation problem by a simultaneous assignment of all sample reads to a pre-processed reference database. A general statistical model based on kernel density estimation techniques is introduced for the assignment task, and the model solution is obtained using convex optimization tools. Further, we design a greedy algorithm solution for a fast solution. Our approach offers a reasonably fast community composition estimation method, which is shown to be more robust to input data variation than a recently introduced related method. AVAILABILITY AND IMPLEMENTATION: A platform-independent Matlab implementation of the method is freely available at http://www.ee.kth.se/ctsoftware; source code that does not require access to Matlab is currently being tested and will be made available later through the above Web site. Saikat Chatterjee, David Koslicki, Siyuan Dong, Nicolas Innocenti, Lu Cheng 0004, Yueheng Lan, Mikko Vehkaperä, Mikael Skoglund, Lars K. Rasmussen, Erik Aurell, Jukka Corander |
Bioinform. | 1 |
| 2014 | Progressive fusion of reconstruction algorithms for low latency applications in compressed sensing
Sooraj K. Ambat, Saikat Chatterjee, K. V. S. Hari |
Signal Process. | 2 |
| 2014 | Dirichlet mixture modeling to estimate an empirical lower bound for LSF quantization
Zhanyu Ma, Saikat Chatterjee, W. Bastiaan Kleijn, Jun Guo 0002 |
Signal Process. | 2 |
| 2014 | Distributed greedy pursuit algorithms
Dennis Sundman, Saikat Chatterjee, Mikael Skoglund |
Signal Process. | 2 |
| 2013 | Fusion of algorithms for Compressed SensingabstractNumerous algorithms have been proposed recently for sparse signal recovery in Compressed Sensing (CS). In practice, the number of measurements can be very limited due to the nature of the problem and/or the underlying statistical distribution of the non-zero elements of the sparse signal may not be known a priori. It has been observed that the performance of any sparse signal recovery algorithm depends on these factors, which makes the selection of a suitable sparse recovery algorithm difficult. To take advantage in such situations, we propose to use a fusion framework using which we employ multiple sparse signal recovery algorithms and fuse their estimates to get a better estimate. Theoretical results justifying the performance improvement are shown. The efficacy of the proposed scheme is demonstrated by Monte Carlo simulations using synthetic sparse signals and ECG signals selected from MIT-BIH database. Sooraj K. Ambat, Saikat Chatterjee, K. V. S. Hari |
ICASSP | 2 |
| 2013 | Pilot design for MIMO channel estimation: An alternative to the Kronecker structure assumptionabstractThis work seeks to design a pilot signal, under a power constraint, such that the channel can be estimated with minimum mean square error. The procedure we derive does not assume Kronecker structure on the underlying covariance matrices, and the pilot signal is obtained in three main steps. Firstly, we solve a relaxed convex version of the original minimization problem. Secondly, its solution is projected onto the feasible set. Thirdly we use the projected solution as starting point for an augmented Lagrangian method. Numerical experiments indicate that this procedure may produce pilot signals that are far better than those obtained under the Kronecker structure assumption. John T. Flåm, Emil Björnson, Saikat Chatterjee |
ICASSP | 3 |
| 2013 | Channel-optimized vector quantizer design for compressed sensing measurementsabstractWe consider vector-quantized (VQ) transmission of compressed sensing (CS) measurements over noisy channels. Adopting mean-square error (MSE) criterion to measure the distortion between a sparse vector and its reconstruction, we derive channel-optimized quantization principles for encoding CS measurement vector and reconstructing sparse source vector. The resulting necessary optimal conditions are used to develop an algorithm for training channel-optimized vector quantization (COVQ) of CS measurements by taking the end-to-end distortion measure into account. Amirpasha Shirazinia, Saikat Chatterjee, Mikael Skoglund |
ICASSP | 2 |
| 2013 | Analysis-by-synthesis-based quantization of compressed sensing measurementsabstractWe consider a resource-constrained scenario where a compressed sensing- (CS) based sensor has a low number of measurements which are quantized at a low rate followed by transmission or storage. Applying this scenario, we develop a new quantizer design which aims to attain a high-quality reconstruction performance of a sparse source signal based on analysis-by-synthesis framework. Through simulations, we compare the performance of the proposed quantization algorithm vis-a-vis existing quantization methods. Amirpasha Shirazinia, Saikat Chatterjee, Mikael Skoglund |
ICASSP | 2 |
| 2013 | Distributed predictive subspace pursuitabstractIn a compressed sensing setup with jointly sparse, correlated data, we develop a distributed greedy algorithm called distributed predictive subspace pursuit. Based on estimates from neighboring sensor nodes, this algorithm operates iteratively in two steps: first forming a prediction of the signal and then solving the compressed sensing problem with an iterative linear minimum mean squared estimator. Through simulations we show that the algorithm provides better performance than current state-of-the-art algorithms. Dennis Sundman, Dave Zachariah, Saikat Chatterjee, Mikael Skoglund |
ICASSP | 3 |
| 2013 | Iteratively reweighted least squares for reconstruction of low-rank matrices with linear structureabstractThis paper considers the problem of reconstructing low-rank matrices from undersampled measurements, when the matrix has a known linear structure. Based on the iterative reweighted least-squares approach, we develop an algorithm that exploits the linear structure in an efficient way that allows for reconstruction in highly undersampled scenarios. The method also enables inferring an appropriate regularization parameter value from the observations. The performance of the method is tested in a missing data recovery problem. Dave Zachariah, Saikat Chatterjee, Magnus Jansson |
ICASSP | 2 |
| 2013 | Line spectrum estimation with probabilistic priors
Dave Zachariah, Petter Wirfält, Magnus Jansson, Saikat Chatterjee |
Signal Process. | 4 |
| 2012 | Detection of sparse random signals using compressive measurementsabstractWe consider the problem of detecting a sparse random signal from the compressive measurements without reconstructing the signal. Using a subspace model for the sparse signal where the signal parameters are drawn according to Gaussian law, we obtain the detector based on Neyman-Pearson criterion and analytically determine its operating characteristics when the signal covariance is known. These results are extended to situations where the covariance cannot be estimated. The results can be used to determine the number of measurements needed for a particular detector performance and also illustrate the presence of an optimal support for a given number of measurements. Bhavani Shankar, Saikat Chatterjee, Björn Ottersten 0001 |
ICASSP | 2 |
| 2012 | A greedy pursuit algorithm for distributed compressed sensingabstractWe develop a greedy pursuit algorithm for solving the distributed compressed sensing problem in a connected network. This algorithm is based on subspace pursuit and uses the mixed support-set signal model. Through experimental evaluation, we show that the distributed algorithm performs significantly better than the standalone (disconnected) solution and close to a centralized (fully connected to a central point) solution. Dennis Sundman, Saikat Chatterjee, Mikael Skoglund |
ICASSP | 2 |
| 2012 | Dynamic subspace pursuitabstractFor compressive sensing of dynamic sparse signals, we develop an iterative greedy search algorithm based on subspace pursuit (SP) that can incorporate sequential predictions, thereby taking advantage of its low complexity while improving recovery performance by exploiting correlations described by a state space model. The algorithm, which we call dynamic subspace pursuit (DSP), is presented and experimentally validated. It exhibits a graceful degradation at deteriorating signal conditions while capable of yielding substantial performance gains as conditions improve. Dave Zachariah, Saikat Chatterjee, Magnus Jansson |
ICASSP | 2 |
| 2012 | Performance bounds for vector quantized compressive sensing
Amirpasha Shirazinia, Saikat Chatterjee, Mikael Skoglund |
ISITA | 2 |
| 2012 | Analysis of sparse representations using bi-orthogonal dictionariesabstractThe sparse representation problem of recovering an N dimensional sparse vector x from M1-norm of x under the constraint y = Dx. In this paper, the performance of l1-reconstruction is analyzed, when the dictionary is bi-orthogonal D = [O1O2], where O1, O2are independent and drawn uniformly according to the Haar measure on the group of orthogonal M × M matrices. By an application of the replica method, we obtain the critical conditions under which perfect l1-recovery is possible with bi-orthogonal dictionaries. Mikko Vehkaperä, Yoshiyuki Kabashima, Saikat Chatterjee, Erik Aurell, Mikael Skoglund, Lars K. Rasmussen |
ITW | 3 |
| 2012 | Alternating Least-Squares for Low-Rank Matrix ReconstructionabstractFor reconstruction of low-rank matrices from undersampled measurements, we develop an iterative algorithm based on least-squares estimation. While the algorithm can be used for any low-rank matrix, it is also capable of exploiting a-priori knowledge of matrix structure. In particular, we consider linearly structured matrices, such as Hankel and Toeplitz, as well as positive semidefinite matrices. The performance of the algorithm, referred to as alternating least-squares (ALS), is evaluated by simulations and compared to the Cramér-Rao bounds. Dave Zachariah, Martin Sundin, Magnus Jansson, Saikat Chatterjee |
IEEE Signal Process. Lett. | 4 |
| 2011 | Look ahead orthogonal matching pursuitabstractFor compressive sensing, we endeavor to improve the recovery performance of the existing orthogonal matching pursuit (OMP) algorithm. To achieve a better estimate of the underlying support set progressively through iterations, we use a look ahead strategy. The choice of an atom in the current iteration is performed by checking its effect on the future iterations (look ahead strategy). Through experimental evaluations, the effect of look ahead strategy is shown to provide a significant improvement in performance. Saikat Chatterjee, Dennis Sundman, Mikael Skoglund |
ICASSP | 1 |
| 2011 | Gaussian mixture modeling for source localizationabstractExploiting prior knowledge, we use Bayesian estimation to localize a source heard by a fixed sensor network. The method has two main aspects: Firstly, the probability density function (PDF) of a function of the source location is approximated by a Gaussian mixture model (GMM). This approximation can theoretically be made arbitrarily accurate, and allows a closed form minimum mean square error (MMSE) estimator for that function. Secondly, the source location is retrieved by minimizing the Euclidean distance between the function and its MMSE estimate using a gradient method. Our method avoids the issues of a numerical MMSE estimator but shows comparable accuracy. John T. Flåm, Joakim Jaldén, Saikat Chatterjee |
ICASSP | 3 |
| 2011 | Analysis of MMSE estimation for compressive sensing of block sparse signalsabstractMinimum mean square error (MMSE) estimation of block sparse signals from noisy linear measurements is considered. Unlike in the standard compressive sensing setup where the non-zero entries of the signal are independently and uniformly distributed across the vector of interest, the information bearing components appear here in large mutually dependent clusters. Using the replica method from statistical physics, we derive a simple closed-form solution for the MMSE obtained by the optimum estimator. We show that the MMSE is a version of the Tse-Hanly formula with system load and MSE scaled by a parameter that depends on the sparsity pattern of the source. It turns out that this is equal to the MSE obtained by a genie-aided MMSE estimator which is informed in advance about the exact locations of the non-zero blocks. The asymptotic results obtained by the non-rigorous replica method are found to have an excellent agreement with finite sized numerical simulations. Mikko Vehkaperä, Saikat Chatterjee, Mikael Skoglund |
ITW | 2 |
| 2011 | Auditory Model-Based Design and Optimization of Feature Vectors for Automatic Speech RecognitionabstractUsing spectral and spectro-temporal auditory models along with perturbation-based analysis, we develop a new framework to optimize a feature vector such that it emulates the behavior of the human auditory system. The optimization is carried out in an offline manner based on the conjecture that the local geometries of the feature vector domain and the perceptual auditory domain should be similar. Using this principle along with a static spectral auditory model, we modify and optimize the static spectral mel frequency cepstral coefficients (MFCCs) without considering any feedback from the speech recognition system. We then extend the work to include spectro-temporal auditory properties into designing a new dynamic spectro-temporal feature vector. Using a spectro-temporal auditory model, we design and optimize the dynamic feature vector to incorporate the behavior of human auditory response across time and frequency. We show that a significant improvement in automatic speech recognition (ASR) performance is obtained for any environmental condition, clean as well as noisy. Saikat Chatterjee, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 1 |
| 2010 | Auditory model based modified MFCC featuresabstractUsing spectral and spectro-temporal auditory models, we develop a computationally simple feature vector based on the design architecture of existing mel frequency cepstral coefficients (MFCCs). Along with the use of an optimized static function to compress a set of filter bank energies, we propose to use a memory-based adaptive compression function to incorporate the behavior of human auditory response across time and frequency. We show that a significant improvement in automatic speech recognition (ASR) performance is obtained for any environmental condition, clean as well as noisy. Saikat Chatterjee, W. Bastiaan Kleijn |
ICASSP | 1 |
| 2010 | Selecting static and dynamic features using an advanced auditory model for speech recognitionabstractWe describe a method to select features for speech recognition that is based on a quantitative model of the human auditory periphery. The method maximizes the similarity of the geometry of the space spanned by the subset of features and the geometry of the space spanned by the auditory model output. The selection method uses a spectro-temporal auditory model that captures both frequency- and time-domain masking. The selection method is blind to the meaning of speech and does not require annotated speech data. We apply the method to the selection of a subset of features from a conventional set consisting of mel cepstra and their first-order and second-order time derivatives. Although our method uses only knowledge of the human auditory periphery, the experimental results show that it performs significantly better than feature-reduction algorithms based on linear and heteroscedastic discriminant analysis that require training with annotated speech data. Christos Koniaris, Saikat Chatterjee, W. Bastiaan Kleijn |
ICASSP | 2 |
| 2010 | A ratification of means: international law and assistive technology in the developing worldabstractSeveral nations around the world have ratified the UN Convention on the Rights of Persons with Disabilities (CRPD) since 2008. Ratifying states commit that national law will guarantee rights enumerated in the CRPD. The use of Assistive Technology (AT) in ensuring the social inclusion of people with disabilities is specifically mentioned in the convention. Although AT is increasingly seen as necessary in facilitating functional equity in social and economic participation, most AT and accessibility tools are not just expensive but are also and typically designed for use by people in industrialized nations. The practical implication of the CRPD's impact on AT for the developing world is a vast subject, in this paper we examine the cost of AT for people with vision impairments at current day costs, and find that the functional fulfillment of the CRPD for a lot of the signatory countries would be extremely difficult without significant technological innovation and market expansion in this space. Joyojeet Pal, Anjali Vartak, Vrutti Vyas, Saikat Chatterjee, Nektarios Paisios, Rahul Cherian |
ICTD | 4 |
| 2009 | Analysis-by-synthesis based switched transform domain split VQ using Gaussian mixture modelabstractUsing analysis-by-synthesis (AbS) approach, we develop a soft decision based switched vector quantization (VQ) method for high quality and low complexity coding of wideband speech line spectral frequency (LSF) parameters. For each switching region, a low complexity transform domain split VQ (TrSVQ) is designed. The overall rate-distortion (R/D) performance optimality of new switched quantizer is addressed in the Gaussian mixture model (GMM) based parametric framework. In the AbS approach, the reduction of quantization complexity is achieved through the use of nearest neighbor (NN) TrSVQs and splitting the transform domain vector into higher number of subvectors. Compared to the current LSF quantization methods, the new method is shown to provide competitive or better trade-off between R/D performance and complexity. Saikat Chatterjee, Thippur V. Sreenivas |
ICASSP | 1 |
| 2009 | Auditory model based optimization of MFCCs improves automatic speech recognition performanceabstractUsing a spectral auditory model along with perturbation based analysis, we develop a new framework to optimize a set of fea-tures such that it emulates the behavior of the human auditory sys-tem. The optimization is carried out in an off-line manner based on the conjecture that the local geometries of the feature domain and the perceptual auditory domain should be similar. Using this principle, we modify and optimize the static mel frequency cep-stral coefficients (MFCCs) without considering any feedback from the speech recognition system. We show that improved recognition performance is obtained for any environmental condition, clean as well as noisy. Index Terms: MFCC, auditory model, ASR. 1. Saikat Chatterjee, Christos Koniaris, W. Bastiaan Kleijn |
INTERSPEECH | 1 |
| 2008 | GMM based Bayesian approach to speech enhancement in signal / transform domainabstractConsidering a general linear model of signal degradation, by modeling the probability density function (PDF) of the clean signal using a Gaussian mixture model (GMM) and additive noise by a Gaussian PDF, we derive the minimum mean square error (MMSE) estimator. The derived MMSE estimator is non-linear and the linear MMSE estimator is shown to be a special case. For speech signal corrupted by independent additive noise, by modeling the joint PDF of time-domain speech samples of a speech frame using a GMM, we propose a speech enhancement method based on the derived MMSE estimator. We also show that the same estimator can be used for transform-domain speech enhancement. Achintya Kundu, Saikat Chatterjee, A. Sreenivasa Murthy, Thippur V. Sreenivas |
ICASSP | 2 |
| 2008 | Subspace based speech enhancement using Gaussian mixture modelabstractTraditional subspace based speech enhancement (SSE)methods \nuse linear minimum mean square error (LMMSE) estimation \nthat is optimal if the Karhunen Loeve transform (KLT) coefficients of speech and noise are Gaussian distributed. In this paper, we investigate the use of Gaussian mixture (GM) density for modeling the non-Gaussian statistics of the clean speech KLT coefficients. Using Gaussian mixture model (GMM), the optimum minimum mean square error (MMSE) estimator is found to be nonlinear and the traditional LMMSE estimator is shown to be a special case. Experimental results show that the proposed method provides better enhancement performance than the traditional subspace based methods.Index Terms: Subspace based speech enhancement, Gaussian mixture density, MMSE estimation. Achintya Kundu, Saikat Chatterjee, Thippur V. Sreenivas |
INTERSPEECH | 2 |
| 2008 | Optimum switched split vector quantization of LSF parameters
Saikat Chatterjee, Thippur V. Sreenivas |
Signal Process. | 1 |
| 2008 | Switched Conditional PDF-Based Split VQ Using Gaussian Mixture ModelabstractIn this letter, we develop switched conditional PDF-based split vector quantization (SCSVQ) method using the recently proposed conditional PDF-based split vector quantizer (CSVQ). The use of CSVQ allows us to alleviate the coding loss by exploiting the correlation between subvectors, in each switching region. Using the Gaussian mixture model (GMM)-based parametric framework, we also address the rate-distortion (R/D) performance optimality of the proposed SCSVQ method by allocating the bits optimally among the switching regions. For the wideband speech line spectrum frequency (LSF) parameter quantization, it is shown that the optimum parametric SCSVQ method provides nearly 2 bits/vector advantage over the recently proposed nonparametric switched split vector quantization (SSVQ) method. Saikat Chatterjee, Thippur V. Sreenivas |
IEEE Signal Process. Lett. | 1 |
| 2008 | Predicting VQ Performance Bound for LSF CodingabstractFor vector quantization (VQ) of speech line spectrum frequency (LSF) parameters, we experimentally determine a mapping function between the mean square error (MSE) measure and the perceptually motivated average spectral distortion (SD) measure. Using the mapping function, we estimate the minimum bits/vector required for transparent quantization of telephone-band and wide-band speech LSF parameters, respectively, as 22 bits/vector and 36 bits/vector, where the distribution of LSF vector is modeled as a Gaussian mixture model (GMM). Saikat Chatterjee, Thippur V. Sreenivas |
IEEE Signal Process. Lett. | 1 |
| 2008 | Optimum Transform Domain Split VQabstractIn this letter, we develop an optimum transform domain split vector quantization (TrSVQ) method. We address both the issues of achieving best rate-distortion (R/D) performance and less complexity. For quantizing a multivariate Gaussian source, we derive the mean-square error (MSE) performance expression for the TrSVQ method using high rate theory and optimum bit allocation. Also, to reduce the complexity, we develop a binary split-based iterative algorithm and use the algorithm in a tree structured manner to find the optimum subvectors' dimensions (i.e., optimum splits). Saikat Chatterjee, Thippur V. Sreenivas |
IEEE Signal Process. Lett. | 1 |
| 2007 | Sequential Split Vector Quantization of LSF Parameters using Conditional PdfabstractA better performing product code vector quantization (VQ) method is proposed for coding the line spectrum frequency (LSF) parameters; the method is referred to as sequential split vector quantization (SeSVQ). The split sub-vectors of the full LSF vector are quantized in sequence and thus uses conditional distribution derived from the previous quantized sub-vectors. Unlike the traditional split vector quantization (SVQ) method, SeSVQ exploits the inter sub-vector correlation and thus provides improved rate-distortion performance, but at the expense of higher memory. We investigate the quantization performance of SeSVQ over traditional SVQ and transform domain split VQ (TrSVQ) methods. Compared to SVQ, SeSVQ saves 1 bit and nearly 3 bits, for telephone-band and wide-band speech coding applications respectively. Saikat Chatterjee, Thippur V. Sreenivas |
ICASSP (4) | 1 |
| 2007 | Normalized two stage SVQ for minimum complexity wide-band LSF quantizationabstractWe develop a two stage split vector quantization method with optimum bit allocation, for achieving minimum computational complexity. This also results in much lower memory requirement than the recently proposed switched split vector quantization method. To improve the rate-distortion performance further, a region specific normalization is introduced, which results in 1 bit/vector improvement over the typical two stage split vector quantizer, for wide-band LSF quantization. Saikat Chatterjee, Thippur V. Sreenivas |
INTERSPEECH | 1 |
| 2007 | Conditional PDF-Based Split Vector Quantization of Wideband LSF ParametersabstractThe commonly used split vector quantization (SVQ) method is inferior to unconstrained quantization due to independent coding of the split subvectors, resulting in a coding loss. In this paper, we propose a conditional pdf-based split vector quantization (CSVQ) method to recover the coding loss. The CSVQ method is developed assuming the line spectrum frequency source distribution as a multivariate Gaussian and the subvectors are quantized sequentially to exploit the correlation between the subvectors. The new CSVQ method is evaluated for wideband speech LSF quantization; CSVQ is shown to outperform traditional SVQ and provide comparable performance to the recently proposed switched split vector quantization (SSVQ) method. In addition, the transform domain SVQ method is also realized to show that its performance is limited by the distance measure used in the transform domain. Saikat Chatterjee, Thippur V. Sreenivas |
IEEE Signal Process. Lett. | 1 |
| 2007 | Analysis of Conditional PDF-Based Split VQabstractThe split vector quantization (SVQ) method results in a “coding loss” due to the independent quantization of split subvectors. To recover the coding loss, conditional pdf-based split vector quantization methods were proposed recently. Saikat Chatterjee, Thippur V. Sreenivas |
IEEE Signal Process. Lett. | 1 |
| 2006 | Two stage transform vector quantization of LSFs for wideband speech codingabstractWe investigate the use of a two stage transform vector quantizer (TSTVQ) for coding of line spectral frequency (LSF) parameters in wideband speech coding. The first stage quantizer of TSTVQ, provides better matching of source distribution and the second stage quantizer provides additional coding gain through using an individual cluster specific decorrelating transform and variance normalization. Further coding gain is shown to be achieved by exploiting the slow time-varying nature of speech spectra and thus using inter-frame cluster continuity (ICC) property in the first stage of TSTVQ method. The proposed method saves 3-4 bits and reduces the computational complexity by 58-66%, compared to the traditional split vector quantizer (SVQ), but at the expense of 1.5-2.5 times of memory. Saikat Chatterjee, Thippur V. Sreenivas |
INTERSPEECH | 1 |
| 2006 | Comparison of prediction based LSF quantization methods using split VQabstractFurther improvement in performance, to achieve near transparent quality LSF quantization, is shown to be possible by using a higher order two dimensional (2-D) prediction in the coefficient domain. The prediction is performed in a closed-loop manner so that the LSF reconstruction error is the same as the quantization error of the prediction residual. We show that an optimum 2-D predictor, exploiting both inter-frame and intra-frame correlations, performs better than existing predictive methods. Computationally efficient split vector quantization technique is used to implement the proposed 2-D prediction based method. We show further improvement in performance by using weighted Euclidean distance. Saikat Chatterjee, Thippur V. Sreenivas |
INTERSPEECH | 1 |