VLDB 2026 Research / reviewers in the wild / expert
Chandra Sekhar Seelamantula
dblp:15/5637 · also S. Chandra Sekhar, Seelamantula Chandrasekhar
· DBLP profile ↗
119ranked-venue papers
13as first author
21since 2021 · last 2025
0000-0001-9049-1912ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 112 · 13 first-author · 19 since 2021Artificial intelligence and machine learning · 20 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | On the Design of Weakly-Convex Regularizers for Solving Linear Inverse ProblemsabstractLinear inverse problems are ubiquitous in signal processing and computational imaging. The prototypical problem is to recover a signal from noisy linear measurements. A typical optimization-based approach is to minimize the sum of a data-fidelity loss and a regularization function. The data-fidelity function ensures consistency with the measurements and the regularization function imparts desired properties to the solution. Convex regularization functions are typically preferred as one can provide theoretical guarantees. However, the convex-nonconvex (CNC) framework, which employs a convex data-fidelity loss and a nonconvex regularization function has been shown to be superior in terms of the quality of signal recovery. In this paper, we consider model-based and data-driven nonconvex regularization objectives to solve linear inverse problems. We consider the denoising problem and propose a constructive approach to design weakly convex regularization functions by minimizing a measure of maximum concavity. Our design approach captures known model-based regularization functions, including those that promote sparsity, and also includes learnable convolutional neural networks. Minimization of the objective follows first-order gradient-based methods. Our approach ensures that the overall reconstruction technique is provably convergent. We show that it outperforms state-of-the-art model-based techniques and is comparable to the benchmark learningbased methods. Crucially, our technique results in reconstructions with fewer artifacts compared to the state-of-the-art learning-based methods. Our reconstruction approach reduces the number of network parameters to be learnt for similar neural network architectures making it easier/faster to train. Abijith Jagannath Kamath, Abhishek Shreekant Bhandiwad, Chandra Sekhar Seelamantula |
ICASSP | 3 |
| 2025 | Neuromorphic Unlimited Sampling for High-Dynamic-Range Video AcquisitionabstractThe unlimited sampling framework (USF) is a computational sensing paradigm that addresses the practical bottleneck pertaining to finite dynamic range and quantization resolution of standard analog-to-digital converters (ADCs). The essence of unlimited sampling is to capture high-dynamic range (HDR) signals using practical sensors by folding the signal within the dynamic range of the sensor, followed by leveraging computational techniques for reconstruction. In this paper, we consider unlimited sampling using neuromorphic/event-driven encoders that enable acquisition of HDR signals as they simultaneously fold the signal and keep track of the folding instants using events. At ICASSP 2024, we proposed neuromorphic unlimited sampling for bandlimited signals. Herein, we extend the theory to include signals in principal shift-invariant spaces, and show that perfect reconstruction is possible without an oversampling requirement on the ADC. Samples of the folded signal along with the events are sufficient for perfect reconstruction with sampling rates that are independent of the dynamic-range of the ADC. Within the proposed technique, a large class of smooth signals that lie outside any shift-invariant space can be accommodated with the reconstruction error decreasing as the sampling interval is reduced. On the experimental front, we consider acquisition of HDR video signals using an array of neuromorphic unlimited samplers and demonstrate accurate reconstruction. The proposed neuromorphic design for HDR video acquisition does not require oversampling of the ADC because the HDR information is indirectly captured in a compressed form using binary events. Abijith Jagannath Kamath, Chandra Sekhar Seelamantula |
ICASSP | 2 |
| 2025 | Some Intriguing Observations on the Learnt Matrices in Deep Unfolded NetworksabstractDeep-unfolded networks (DUNs) have set new performance benchmarks in fields such as compressed sensing, image restoration, and wireless communications. DUNs are built from conventional iterative algorithms, where an iteration is transformed into a layer/block of a network with learnable parameters. Despite their huge success, the reasons behind their superior performance over their iterative counterparts are not fully understood. This paper focuses on enhancing the explainability of DUNs by investigating potential reasons behind their superior performance over traditional iterative methods. We concentrate on the Learnt Iterative Shrinkage-Thresholding Algorithm (LISTA), a foundational contribution that achieves sparse recovery with significantly fewer layers than its iterative counterpart, ISTA. Our findings reveal that the learnt matrices in LISTA always have Gaussian distributed entries regardless of whether the sensing matrix is random Gaussian, Bernoulli, exponential, or uniform. The findings also show that the singular values of the learnt matrices exceed unity, despite which, the reconstruction scheme is stable. We conjecture that the activation function may have a role to play in ensuring stability. We also present an unbiasing technique that substantially improves the sparse recovery performance by reestimating the amplitudes based on the converged support. Kartheek Kumar Reddy Nareddy, Inbasekaran Perumal, Chandra Sekhar Seelamantula |
ICASSP | 3 |
| 2025 | Monte Carlo Score Matching for Image GenerationabstractScore-based models are state-of-the-art generative models for image generation. We propose a novel loss namely the Monte Carlo Score Matching (MCSM) loss as an approximation of the original score matching loss. MCSM leverages a Taylor-series expansion of the score function to approximate the expensive calculation involved in computing the trace of the Jacobian of the score function. MCSM is competitive with models trained using the Sliced-Score Matching (SSM) loss. We validate the efficacy of the proposed technique in terms of negative log-likelihood and Fréchet Inception distance (FID) on MNIST and CelebA datasets, respectively. In particular, we show that FID of images generated with models trained using MCSM loss is on par with, and in some cases, better than, sliced score-matching for image generation. Nishanth Shetty, Chandra Sekhar Seelamantula |
ICASSP | 2 |
| 2025 | Diffusion Model Based Image Reconstruction in Lensless ImagingabstractLensless imaging systems eliminate the need for lenses by employing an encoding element to multiplex incident light signals, which are then captured directly onto a bare camera sensor. They present a promising alternative to traditional lens-based imaging systems by offering significant advantages in terms of compactness, versatility, and cost. Due to the multiplexed nature of measurements, image reconstruction takes place computationally. However, existing techniques for image reconstruction in lensless imaging fall short of the image quality offered by traditional lens-based imaging. In this work, we consider the application of diffusion models, a class of deep generative models, for image reconstruction in a lensless imaging modality. These models currently achieve state-of-the-art performance in image generation. Specifically, we focus on the PhlatCam lensless system, which consists of a coded phase mask as the encoding element placed close to the camera sensor. We use a ControlNet based diffusion model to improve the perceptual quality of image reconstruction. The performance is measured in terms of peak signal-to-noise ratio (PSNR), structural similarity index measure (SSIM), and learned perceptual image patch similarity (LPIPS). The proposed method improves the performance in these metrics for synthetic measurements. For real measurements, the improvement in image quality comes at the expense of a small bias in color, which is attributed to the generative nature of the diffusion prior itself. Vivek Boominathan, Ashok Veeraraghavan, Chandra Sekhar Seelamantula |
ICASSP | 4 |
| 2025 | Deep Unsupervised Despeckling With Unbiased Risk EstimationabstractDespeckling of Synthetic Aperture Radar (SAR) images has seen significant progress in recent years, largely driven by advancements in deep learning techniques. However, many of these approaches face challenges when applied to new SAR datasets, primarily due to their dependence on ground truth images, which are often unavailable for real-world sensors. In this paper, we address this limitation by extending the concept of unbiased risk estimation in the presence of Gamma-distributed multiplicative speckle. Specifically, we demonstrate that it is possible to train deep denoising networks without relying on ground truth data using our estimator. We introduce a new formulation of the Multiplicative Unbiased Risk Estimator (MURE) and present a computationally efficient Monte Carlo-based method that enables accurate estimation of the modified MURE cost, facilitating effective unsupervised training of deep neural networks from large datasets consisting solely of noisy SAR images. Experimental results on both synthetic datasets and real Sentinel-1 SAR images validate the suitability of our method for real-world applications. Even without ground truth, our method achieves performance that closely matches the Oracle-based denoiser and proves superior to the out-of-domain performance of popular supervised SAR despeckling methods. Ashutosh Gupta 0008, Chandra Sekhar Seelamantula, Thierry Blu, Nitant Dube, Shanmuganathan Raman |
ICIP | 2 |
| 2025 | LURE: An Unsupervised Denoising Framework for Multiplicative Lognormal NoiseabstractAbstract. Most image denoising problems focus on additive white Gaussian noise. In real-world imaging scenarios such as ultrasound, synthetic aperture radar, and optical-coherence tomography, the noise is multiplicative. Multiplicative noise has an extreme degradation effect compared to additive noise of the same variance. Further, in a practical imaging setting, one does not have access to the ground-truth clean images to train a deep neural network in a supervised fashion for image denoising. In this paper, we propose an unsupervised image denoising method for multiplicative noise. Specifically, we consider lognormal noise and develop an unbiased risk estimator of the mean-square error (MSE). We show that the resulting lognormal noise unbiased risk estimate, which we abbreviate as LURE, is an accurate estimator of the Oracle MSE. Computation of LURE involves the weighted trace of the Jacobian, which we estimate using a stochastic/Monte Carlo approximation method that is not only fast but also results in an accurate estimate of the MSE. The framework is flexible enough to accommodate a wide spectrum of denoisers—from wavelet denoising techniques to state-of-the-art deep learning techniques, subject to certain smoothness conditions on the denoiser. We deploy modern deep learning models such as U-Net, dilated-residual U-Net (DRUNet), and gradient step denoiser with DRUNet (GS-DRUNet) to establish the reliability of LURE. The performance measures used are peak signal-to-noise ratio (PSNR) and structural similarity index metric (SSIM). We show that minimizing Monte Carlo LURE in an unsupervised setting gives results that are on par with and sometimes even better than those obtained using the Oracle MSE loss in the supervised setting. We also provide comparisons with unsupervised despeckling techniques such as SAR2SAR, SAR-CNN, and Speckle2Void on real-world noisy images. Monalisa Bakshi, Gayathri Venkat, Nikhil Bisen, Chandra Sekhar Seelamantula, Thierry Blu |
SIAM J. Imaging Sci. | 4 |
| 2024 | Variational Analysis of Adversarial Regularization for Solving Inverse ProblemsabstractInverse problems form the backbone of modern signal/image processing and computational imaging, where signal reconstruction from corrupted measurements follows an optimization problem. The objective function is the sum of a data-fidelity term and a regularization functional that enforces desired properties in the reconstruction. The adversarial regularization (AR) framework is an unsupervised, data-driven approach for solving inverse problems, where the regularization function is learnt adversarially as a critique between the ground-truth distribution and the distribution of unregularized reconstructions. Thereafter, the solution to the regularized inverse problem follows an iterative technique. In this paper, we analyze the AR framework from a variational perspective, and, using Euler-Lagrange conditions, obtain the optimal regularization function in closed-form. The overall objective function is smooth and readily amenable to gradient descent minimization. We introduce momentum into the iterates as a natural extension to accelerate convergence. Since the optimal solutions are obtained in closed-form, our approach to solving inverse problems does not require prior training whilst being data-driven. We demonstrate the proposed technique on image deconvolution and show that the reconstruction performance of the proposed techniques measured in terms of peak signal-to-noise ratio (PSNR) and structural similarity index metric (SSIM) are identical to the learnt counterparts. Abhishek Shreekant Bhandiwad, Abijith Jagannath Kamath, Siddarth Asokan, Chandra Sekhar Seelamantula |
ICASSP | 4 |
| 2024 | Neuromorphic Sensing Meets Unlimited SamplingabstractUnlimited sampling is a computational sensing paradigm for high-dynamic range (HDR) acquisition of continuous-time signals. In standard analog-to-digital converters (ADC), a fixed input dynamic range is accommodated by clipping or saturation of the signal outside the dynamic range. In unlimited sampling, signal that lies outside the fixed dynamic range is folded back, thereby preserving the signal dynamic range. The folding is achieved using a self-reset ADC (SR-ADC), which performs a continuous-time modulo operation. In this paper, we use a neuromorphic encoder, which is an opportunistic and event-driven sampling device, and propose a new technique for unlimited sampling. We show that the neuromorphic encoder folds the signal that lies outside the dynamic range and simultaneously records a compressed representation of the error signal. Unlimited sampling is achieved by measuring the folded signal. We analyze sampling and reconstruction of finite-energy, bandlimited signals and show that perfect reconstruction is possible using uniform samples acquired at the Nyquist rate of the signal. The reconstruction technique operates in real-time, and can be readily extended to larger classes of continuous-time signals. We demonstrate the performance of our technique and report comparisons with state-of-the-art techniques using simulations. Abijith Jagannath Kamath, Chandra Sekhar Seelamantula |
ICASSP | 2 |
| 2024 | Image Restoration with Generalized L2 Loss and Convergent Plug-and-Play PriorsabstractImage restoration involves solving an optimization problem where the objective function is the sum of a data-fidelity term and a regularization functional that incorporates a desired image prior. Solving the optimization problem using proximal methods results in iterative algorithms that require computing a gradient step corresponding to the data-fidelity loss and a proximal update corresponding to enforcing the image prior. In this paper, we develop a novel formulation for image restoration considering a generalized data-fidelity loss and a convex regularization function that enforces a desired image prior, and we solve the problem using proximal gradient method. The choice of the data-fidelity loss is such that the adjoint operator is reminiscent of Wiener filtering when the forward operator is a convolutional operator (for instance, a shift-invariant blur kernel). The proposed gradient update ensures that the iterates remain in the solution-space of the linear measurement constraints. We further propose the plug-and-play counterpart of the restoration technique, which allows one to leverage off-the-shelf data-driven denoisers in place of the proximal operator. Experimental validations carried out on BSD500, Brodatz, Urban100, and DIV2K datasets show that the proposed technique gives rise to superior image reconstruction quality compared with the state-of-the-art techniques, with the performance measured in terms of peak signal-to-noise ratio (PSNR) and structural similarity index metric (SSIM), with comparable computational complexity. Kartheek Kumar Reddy Nareddy, Abijith Jagannath Kamath, Chandra Sekhar Seelamantula |
ICASSP | 3 |
| 2024 | Momentum-Imbued Langevin Dynamics (MILD) for Faster SamplingabstractScore-based generative models have emerged as the state-of-the-art in generative modeling. In this paper, we introduce a novel sampling scheme that can be combined with pre-trained score-based diffusion models to speed up sampling by a factor of two to five in terms of the number of function evaluations (NFEs) with a superior Fréchet Inception distance (FID), compared to Annealed Langevin dynamics in noise-conditional score network (NCSN) and improved noise-conditional score network (NCSN++). The proposed sampling algorithm is inspired by momentum-based accelerated gradient descent used in convex optimization techniques. We validate the sampling efficiency of the proposed algorithm in terms of FID on CIFAR-10 and CelebA datasets. Nishanth Shetty, Manikanta Bandla, Nishit Neema, Siddarth Asokan, Chandra Sekhar Seelamantula |
ICASSP | 5 |
| 2024 | Tight-Frame-Like Analysis-Sparse Recovery Using Nontight Sensing MatricesabstractAbstract. The choice of the sensing matrix is crucial in compressed sensing. Random Gaussian sensing matrices satisfy the restricted isometry property, which is crucial for solving the sparse recovery problem using convex optimization techniques. However, tight-frame sensing matrices result in minimum mean-squared-error recovery given oracle knowledge of the support of the sparse vector. If the sensing matrix is not tight, could one achieve the recovery performance assured by a tight frame by suitably designing the recovery strategy? This is the key question addressed in this paper. We consider the analysis-sparse [Formula: see text]-minimization problem with a generalized [Formula: see text]-norm-based data-fidelity and show that it effectively corresponds to using a tight-frame sensing matrix. The new formulation offers improved performance bounds when the number of nonzeros is large. One could develop a tight-frame variant of a known sparse recovery algorithm using the proposed formalism. We solve the analysis-sparse recovery problem in an unconstrained setting using proximal methods. Within the tight-frame sensing framework, we rescale the gradients of the data-fidelity loss in the iterative updates to further improve the accuracy of analysis-sparse recovery. Experimental results show that the proposed algorithms offer superior analysis-sparse recovery performance. Proceeding further, we also develop deep-unfolded variants, with a convolutional neural network as the sparsifying operator. On the application front, we consider compressed sensing image recovery. Experimental results on Set11, BSD68, Urban100, and DIV2K datasets show that the proposed techniques outperform the state-of-the-art techniques, with performance measured in terms of peak signal-to-noise ratio and structural similarity index metric. Kartheek Kumar Reddy Nareddy, Abijith Jagannath Kamath, Chandra Sekhar Seelamantula |
SIAM J. Imaging Sci. | 3 |
| 2023 | Spider GAN: Leveraging Friendly Neighbors to Accelerate GAN TrainingabstractTraining Generative adversarial networks (GANs) stably is a challenging task. The generator in GANs transform noise vectors, typically Gaussian distributed, into realistic data such as images. In this paper, we propose a novel approach for training GANs with images as inputs, but without enforcing any pairwise constraints. The intuition is that images are more structured than noise, which the generator can leverage to learn a more robust transformation. The process can be made efficient by identifying closely related datasets, or a “friendly neighborhood” of the target distribution, inspiring the moniker, Spider GAN. To define friendly neighborhoods leveraging proximity between datasets, we propose a new measure called the signed inception distance (SID), inspired by the polyharmonic kernel. We show that the Spider GAN formulation results in faster convergence, as the generator can discover correspondence even between seemingly unrelated datasets, for instance, between TinyImageNet and CelebA faces. Further, we demonstrate cascading Spider GAN, where the output distribution from a pre-trained GAN generator is used as the input to the subsequent network. Effectively, transporting one distribution to another in a cascaded fashion until the target is learnt – a new flavor of transfer learning. We demonstrate the efficacy of the Spider approach on DCGAN, conditional GAN, PGGAN, StyleGAN2 and StyleGAN3. The proposed approach achieves state-of-the-art Fréchet inception distance (FID) values, with one-fifth of the training iterations, in comparison to their baseline counterparts on high-resolution small datasets such as MetFaces, Ukiyo-E Faces and AFHQ-Cats. Siddarth Asokan, Chandra Sekhar Seelamantula |
CVPR | 2 |
| 2023 | A Game of Snakes and GansabstractGenerative adversarial networks (GANs) comprise generator and discriminator networks trained adversarially to learn the underlying distribution of a dataset. Recently, we have shown that the optimal GAN discriminator can be obtained in closed-form as the solution to the Poisson partial differential equation (PDE). While existing approaches either train a network or solve the PDE in closed-form, we propose training the generator through the gradient field of the optimal discriminator. In this paper, we establish a connection between active contour models (snakes) and GANs. We evolve a set of snake points over the gradient field of radial basis function (RBF) Coulomb GAN. The generator is then trained to follow the trajectory of the snake. The proposed approach benefits from both the sample diversity seen in flow-based approaches and the fast sampling capability of GANs. Experimental validation on 2-D synthetic data shows that the proposed approach leads to accelerated convergence, compared against the baseline approaches that either employ network or kernel-based discriminators. Siddarth Asokan, Fatwir Sheikh Mohammed, Chandra Sekhar Seelamantula |
ICASSP | 3 |
| 2023 | Multichannel Time-Encoding of Finite-Rate-of-Innovation SignalsabstractTime-encoding of continuous-time signals is an alternative sampling paradigm to Shannon sampling. In time-encoding or event-driven sampling, the signal is encoded using a sequence of time instants corresponding to an event. In this paper, we propose multichannel time-encoding of signals with a finite-rate-of-innovation (FRI) in single-input-multi-output (SIMO) and multi-input-multi-output (MIMO) configurations using the integrate-and-fire model. We demonstrate perfect reconstruction of FRI signals with common support from MIMO time-encoded measurements using a joint estimation technique, and perfect reconstruction of FRI signals from SIMO time-encoded measurements with reduced sampling requirement as compared to the single channel case. We provide sufficient conditions for perfect reconstruction with sampling requirement of the order of the rate of innovation of the signal. We substantiate our claims using simulations on noise-free and noisy measurements. Abijith Jagannath Kamath, Chandra Sekhar Seelamantula |
ICASSP | 2 |
| 2023 | Euler-Lagrange Analysis of Generative Adversarial NetworksabstractWe consider Generative Adversarial Networks (GANs) and address the underlying functional optimization problem ab initio within a variational setting. Strictly speaking, the optimization of the generator and discriminator functions must be carried out in accordance with the Euler-Lagrange conditions, which become particularly relevant in scenarios where the optimization cost involves regularizers comprising the derivatives of these functions. Considering Wasserstein GANs (WGANs) with a gradient-norm penalty, we show that the optimal discriminator is the solution to a Poisson differential equation. In principle, the optimal discriminator can be obtained in closed form without having to train a neural network. We illustrate this by employing a Fourier-series approximation to solve the Poisson differential equation. Experimental results based on synthesized Gaussian data demonstrate superior convergence behavior of the proposed approach in comparison with the baseline WGAN variants that employ weight-clipping, gradient or Lipschitz penalties on the discriminator on low-dimensional data. We also analyze the truncation error of the Fourier-series approximation and the estimation error of the Fourier coefficients in a high-dimensional setting. We demonstrate applications to real-world images considering latent-space prior matching in Wasserstein autoencoders and present performance comparisons on benchmark datasets such as MNIST, SVHN, CelebA, CIFAR-10, and Ukiyo-E. We demonstrate that the proposed approach achieves comparable reconstruction error and Frechet inception distance with faster convergence and up to two-fold improvement in image sharpness. Siddarth Asokan, Chandra Sekhar Seelamantula |
J. Mach. Learn. Res. | 2 |
| 2022 | Differentiate-and-Fire Time-Encoding of Finite-Rate-of-Innovation SignalsabstractTime-encoding or event-driven sampling of continuous-time signals is an alternative paradigm to uniform sampling. In this sampling scheme, the signal is encoded by a sequence of time-instants as opposed to a sequence of amplitudes in uniform sampling. Time-encoding is opportunistic by design – measurements are taken only when the signal exhibits significant variability. Consequently, the measurements are sparse, noise-robust, and require low power. However, standard processing and reconstruction methods do not apply. In this paper, we introduce a new time-encoding machine, namely, differentiate-and-fire time-encoding machine (DIF-TEM) inspired by the functioning of the human visual system. A DIF-TEM can be tuned to provide sampling sets with variable densities – sparse sets that mimic dynamic vision sensors (neuromorphic cameras) or dense sets that mimic classical time-encoding machines. We propose kernel-based time-encoding of finite-rate-of-innovation (FRI) signals using DIF-TEM via Fourier-domain analysis. We show that DIF-TEM measurements are sufficient for perfect signal reconstruction under certain conditions. We provide simulation results to substantiate our claims. Abijith Jagannath Kamath, Chandra Sekhar Seelamantula |
ICASSP | 2 |
| 2022 | Compressive Phase Retrieval Based On Sparse Latent Generative PriorsabstractWe address the problem of compressive phase retrieval (CPR) based on generative prior. The problem is ill-posed and requires structural assumptions. CPR techniques impose sparsity prior on the signal to perform reconstruction from compressive phaseless measurements. Recent developments in data-driven signal models in the form of generative priors have been shown to outperform sparsity priors with significantly fewer measurements. However, it is possible to improve upon the performance of generative prior based methods by introducing structure in the latent-space. We propose to introduce structure on the signal by enforcing sparsity in the latent-space via proximal method while training the generator. The optimization is called as proximal meta-learning (PML). Enforcing sparsity in the latent space naturally leads to a union-of-submanifolds model in the solution space. The overall framework of imposing sparsity along with PML is called as sparsity-driven latent space sampling (SDLSS). We demonstrate the efficacy of the proposed framework over the state-of-the-art deep phase retrieval (DPR) technique on MNIST and CelebA datasets. We evaluate the performance as a function of the number of measurements and sparsity factor using standard objective measures. The results show that SDLSS performs better at higher compression ratio and has faster recovery compared with DPR. Vinayak Killedar, Chandra Sekhar Seelamantula |
ICASSP | 2 |
| 2022 | An Ensemble of Proximal Networks for Sparse CodingabstractSparse coding methods are iterative and typically rely on proximal gradient methods. While the commonly used sparsity promoting penalty is the ℓ1norm, alternatives such as the minimax concave penalty (MCP) and smoothly clipped absolute deviation (SCAD) penalty have also been employed to obtain superior results. Combining various penalties to achieve robust sparse recovery is possible, but the challenge lies in parameter tuning. Given the connection between deep networks and unrolling of iterative algorithms, it is possible to unify the unfolded networks arising from different formulations. We propose an ensemble of proximal networks for sparse recovery, where the ensemble weights are learnt in a data-driven fashion. We found that the proposed network performs superior to or on par with the individual networks in the ensemble for synthetic data under various noise levels and sparsity conditions. We demonstrate an application to image denoising based on the convolutional sparse coding formulation. Kartheek Kumar Reddy Nareddy, Swapnil Mache, Praveen Kumar Pokala, Chandra Sekhar Seelamantula |
ICIP | 4 |
| 2022 | Iteratively Reweighted Minimax-Concave Penalty Minimization for Accurate Low-rank Plus Sparse Matrix Decompositionabstract-norm, respectively. Convex approximations are known to result in biased estimates, to overcome which, nonconvex regularizers such as weighted nuclear-norm minimization and weighted Schatten p-norm minimization have been proposed. However, works employing these regularizers have used heuristic weight-selection strategies. We propose weighted minimax-concave penalty (WMCP) as the nonconvex regularizer and show that it admits an equivalent representation that enables weight adaptation. Similarly, an equivalent representation to the weighted matrix gamma norm (WMGN) enables weight adaptation for the low-rank part. The optimization algorithms are based on the alternating direction method of multipliers technique. We show that the optimization frameworks relying on the two penalties, WMCP and WMGN, coupled with a novel iterative weight update strategy, result in accurate low-rank plus sparse matrix decomposition. The algorithms are also shown to satisfy descent properties and convergence guarantees. On the applications front, we consider the problem of foreground-background separation in video sequences. Simulation experiments and validations on standard datasets, namely, I2R, CDnet 2012, and BMC 2012 show that the proposed techniques outperform the benchmark techniques. Praveen Kumar Pokala, Raghu Vamshi Hemadri, Chandra Sekhar Seelamantula |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Sparsity Driven Latent Space Sampling for Generative Prior Based Compressive SensingabstractWe address the problem of recovering signals from compressed measurements based on generative priors. Recently, generative-model based compressive sensing (GMCS) methods have shown superior performance over traditional compressive sensing (CS) techniques in recovering signals from fewer measurements. However, it is possible to further improve the performance of GMCS by introducing controlled sparsity in the latent-space. We propose a proximal meta-learning (PML) algorithm to enforce sparsity in the latent-space while training the generator. Enforcing sparsity naturally leads to a union-of-submanifolds model in the solution space. The overall framework is named as sparsity driven latent space sampling (SDLSS). In addition, we derive the sample complexity bounds for the proposed model. Furthermore, we demonstrate the efficacy of the proposed framework over the state-of-the-art techniques with application to CS on standard datasets such as MNIST and CIFAR-10. In particular, we evaluate the performance of the proposed method as a function of the number of measurements and sparsity factor in the latent space using standard objective measures. Our findings show that the sparsity driven latent space sampling approach improves the accuracy and aids in faster recovery of the signal in GMCS. Vinayak Killedar, Praveen Kumar Pokala, Chandra Sekhar Seelamantula |
ICASSP | 3 |
| 2020 | Wirtinger Flow Algorithms for Phase Retrieval from Binary MeasurementsabstractWe consider the problem of Binary Phase Retrieval, wherein we attempt to recover signals from their quadratic measurements, which are further encoded as +1 or -1 depending on whether they exceed a threshold or not. Binary encoding is the extreme case of quantization in the phase retrieval setting and has been introduced only recently. We formulate a consistency based non-convex cost function, which requires the signal estimate to agree with the binary measurements. Since lifting the measurements does not scale well with respect to the signal dimension, we solve the problem using Wirtinger Flow and the recently proposed accelerated Wirtinger Flow algorithms. We also propose a spectral initialization for the binary measurement model. The proposed algorithms have low computational complexity and also scale well with respect to the signal dimension. Simulation results show that the Wirtinger flow solution to the Binary Phase Retrieval problem is nearly on par with that obtained using the principle of lifting - the loss in signal-to-reconstruction error ratio (SRER) is about 1 dB, but the consistency with the binary observations is 100% in the absence of noise. Further, simulation results show that the accelerated Binary Wirtinger Flow (BWF) gives a 2 dB improvement in SRER over BWF. Vinith Kishore, Chandra Sekhar Seelamantula |
ICASSP | 2 |
| 2020 | Epoch Estimation from a Speech Signal Using Gammatone Wavelets in a Scattering NetworkabstractIn speech production, epochs are glottal closure instants where significant energy is released from the lungs. Extracting an epoch accurately is important in speech synthesis, analysis, and pitch oriented studies. The time-varying characteristics of the source and the system, and channel attenuation of low-frequency components by telephone channels make estimation of epoch from a speech signal a challenging task. In this paper, we propose a new technique that employs a Gammatone wavelet filterbank and compute a scattering sequence whose local maxima define the candidate epochs in the speech signal. Results are presented for both normal and telephone channel speech by considering the differential electroglottograph from CMU-Arctic database as the ground-truth. The proposed method gives significant improvements with respect to multiple performance metrics when compared with state-of-the-art techniques for epoch estimation. Pavan Kulkarni, Jishnu Sadasivan, Aniruddha Adiga, Chandra Sekhar Seelamantula |
ICASSP | 4 |
| 2020 | Confirmnet: Convolutional Firmnet and Application to Image Denoising and InpaintingabstractWe address the problem of efficient convolutional sparse coding (CSC) and develop a non-convex-penalty-regularized CSC formulation, namely, minimax-concave CSC (MC2SC). MC2SC leads to an optimal sparse representation than the standard ℓ1-penalty based approach. In addition, suitable convergence guarantees can also be provided for MC2SC. We propose a convolutional iterative firm-thresholding algorithm (CIFTA) building on our previously proposed IFTA, and its deep-unfolded version, namely, convolutional-FirmNet (ConFirmNet). As an application, we develop the ConFirmNet based sparse autoencoder (ConFirmNet-SAE) for learning an application-specific convolutional dictionary, the applications being image denoising and inpainting. Further, we also show that training ConFirmNet-SAE with the Huber loss imparts robustness to outliers. It also turns out that ConFirmNet-SAE is robust to mismatch between training and test noise conditions than convolutional learned iterative soft-thresholding algorithm (LISTA). Praveen Kumar Pokala, Prakash Kumar Uttam, Chandra Sekhar Seelamantula |
ICASSP | 3 |
| 2020 | A Time-Based Sampling Framework for Finite-Rate-of-Innovation SignalsabstractTime-based sampling of continuous-time signals is an alternative to Shannon's sampling paradigm in which the signal is encoded using a sequence of nonuniform time instants. The standard methods for reconstructing signals in bandlimited and shift-invariant spaces from their nonuniform measurements employ alternating projections algorithms. In this paper, we consider the problem of sampling and perfect reconstruction of periodic finite-rate-of-innovation (FRI) signals using crossing-time-encoding machine (C-TEM) and integrate-and-fire TEM (IF-TEM). We formulate the reconstruction problem in the frequency domain and develop techniques to compute the Fourier coefficients, which contain the unknown parameters of the signal in the form of a sum of weighted complex exponentials. The parameters are then estimated using high-resolution spectral estimation techniques. Unlike state-of-the-art methods, the proposed method is generalized to incorporate reconstruction of periodic FRI signals consisting of weighted and shifted versions of an arbitrary pulse with arbitrarily close delays, and is compatible with a large class of sampling kernels. We provide sufficient conditions for sampling and perfect reconstruction using C-TEM and IF-TEM. We present simulation results to support our claims. We also discuss an extension to the sampling of aperiodic FRI signals. Sunil Rudresh, Abijith Jagannath Kamath, Chandra Sekhar Seelamantula |
ICASSP | 3 |
| 2020 | Non-Convex Optimization For Sparse Interferometric Phase EstimationabstractWe present a new sparsity based technique for interferometric phase estimation. We consider complex extensions of non-convex regularizers such as the minimax concave penalty (MCP) and smoothly clipped absolute deviation penalty (SCAD) for sparse recovery. We solve the problem of interferometric phase estimation based on complex-domain dictionary learning. We develop an algorithm, namely, improved sparse interferometric phase estimation (iSpInPhase) based on alternating direction method of multipliers (ADMM) and Wirtinger calculus for solving the optimization problem. Wiritinger calculus is employed because the cost functions are nonholomorphic. We evaluate the performance of iSpInPhase on synthetic data, namely, truncated Gaussian elevation and also on mountain terrain data, namely, Long's peak, for different noise levels. Performance comparisons show that iSpInPhase outperforms the state-of-the-art techniques in terms of standard performance assessment measures. Satvik Chemudupati, Praveen Kumar Pokala, Chandra Sekhar Seelamantula |
ICIP | 3 |
| 2020 | Cornet: Composite-Regularized Neural Network For Convolutional Sparse CodingabstractSparse recovery via composite regularization is an interesting approach proposed recently in the literature. One could design nonconvex regularizers through a convex combination of sparsity-promoting penalties with known proximal operators. We develop a new algorithm, namely, convolutional proximal-averaged thresholding algorithm (C-PATA) for composite-regularized convolutional sparse coding (CR-CSC) based on the recently proposed idea of proximal averaging. We develop an autoencoder structure based on the deep-unfolding of C-PATA iterations into neural network layers, which results in the composite-regularized neural network (CoRNet) architecture. The convolutional learned iterative soft-thresholding algorithm becomes a special case of CoRNet. We demonstrate the efficacy of CoRNet considering applications to image denoising and inpainting, and compare the performance with state-of-the-art techniques such as BM3D, convolutional LISTA, and fast and flexible convolutional sparse coding (FFCSC). Dhruv Jawali, Praveen Kumar Pokala, Chandra Sekhar Seelamantula |
ICIP | 3 |
| 2020 | Generalized Fast Iteratively Reweighted Soft-Thresholding Algorithm for Sparse Coding Under Tight Frames in the Complex-DomainabstractWe present a new method for fast magnetic resonance image (MRI) reconstruction in the complex-domain under tight frames. We propose a generalized problem formulation that allows for different weight-update strategies for iteratively reweighted ℓ1-minimization under tight frames. Further, we impose sufficient conditions on the function of the weights that leads to the reweighting strategy, which follows the interpretation originally given by Candès et al, but is more efficient than theirs. Since the objective function in complex-domain compressive sensing MRI (CS-MRI) reconstruction problem is nonholomorphic, we resort to Wirtinger calculus for deriving the update strategies. We develop an algorithm called generalized iteratively reweighted soft-thresholding algorithm (GIRSTA) and its fast variant, namely, generalized fast iteratively reweighted soft-thresholding algorithm (GFIRSTA). We provide convergence guarantees for GIRSTA and empirical convergence results for GFIRSTA. Our experiments show a remarkable performance of the proposed algorithms for complex-domain CS-MRI reconstruction considering both random sampling and radial sampling strategies. GFIRSTA outperforms state-of-the-art techniques in terms of peak signal-to-noise ratio (PSNR) and structural similarity index metric (SSIM). Praveen Kumar Pokala, Satvik Chemudupati, Chandra Sekhar Seelamantula |
ICIP | 3 |
| 2020 | Projected Improved Fista And Application To Image DeblurringabstractThe analysis-sparse model has been shown to be efficient as compared to the synthesis-sparse model when the sparsifying transform is redundant particularly for image restoration applications. We pose the image deblurring problem as an optimization problem based on the analysis-sparse model considering the sparsifying basis to be a tight frame, more specifically, the shift-invariant discrete wavelet transform (SIDWT). We propose two algorithms, namely, projected improved fast iterative soft-thresholding algorithm (piFISTA) and projected improved fast iterative soft-thresholding algorithm beyond Nesterov's momentum (piFISTA-BN). The proposed algorithms are the analysis counterparts of the improved fast iterative soft-thresholding algorithm (iFISTA) and improved fast iterative soft-thresholding algorithm beyond Nesterov's momentum (iFISTA-BN), respectively, both of which consider the synthesis-sparse model. We demonstrate that piFISTA and piFISTA-BN significantly outperform FISTA, pFISTA, iFISTA, and iFISTA-BN considering standard objective metrics such as peak signal-to-noise ratio (PSNR) and structural similarity index metric (SSIM). Further, we demonstrate empirically that the proposed algorithms converge faster than the state-of-the-art techniques. Praveen Kumar Pokala, Chandra Sekhar Seelamantula |
ICIP | 2 |
| 2020 | Nyquist Pulses for Sub-Nyquist Sampling - Application to Underwater ImagingabstractThe introduction of finite-rate-of-innovation (FRI) sampling has made it possible to sample and perfectly reconstruct certain classes of non-bandlimited signals. The design of sampling kernels for FRI framework relies on frequency-domain alias-cancellation and Strang-Fix conditions. We establish an equivalence between the alias-cancellation conditions and zero intersymbol interference (ISI) conditions required for distortionless transmission in the field of digital communication. Consequently, Nyquist pulses employed for ISI-free communication could also be used as FRI sampling kernels. As an example, a Nyquist pulse, namely, the raised-cosine pulse is employed as the FRI sampling kernel and its performance is analyzed in the presence of noise. In essence, we show that any admissible FRI sampling kernel follows the Nyquist pulse criterion and can be used interchangeably. As an application, we demonstrate super-resolution underwater imaging by employing the FRI signal model on experimental active sonar measurements. Compared with standard matched filtering, the FRI methodology results in superior-quality image reconstruction. Suhas Srinath, Sunil Rudresh, Chandra Sekhar Seelamantula, Hareesh G, Murali Krishna P. |
ICIP | 3 |
| 2020 | Teaching a GAN What Not to LearnabstractGenerative adversarial networks (GANs) were originally envisioned as unsupervised generative models that learn to follow a target distribution. Variants such as conditional GANs, auxiliary-classifier GANs (ACGANs) project GANs on to supervised and semi-supervised learning frameworks by providing labelled data and using multi-class discriminators. In this paper, we approach the supervised GAN problem from a different perspective, one that is motivated by the philosophy of the famous Persian poet Rumi who said, "The art of knowing is knowing what to ignore." In the GAN framework, we not only provide the GAN positive data that it must learn to model, but also present it with so-called negative samples that it must learn to avoid — we call this "The Rumi Framework." This formulation allows the discriminator to represent the underlying target distribution better by learning to penalize generated samples that are undesirable — we show that this capability accelerates the learning process of the generator. We present a reformulation of the standard GAN (SGAN) and least-squares GAN (LSGAN) within the Rumi setting. The advantage of the reformulation is demonstrated by means of experiments conducted on MNIST, Fashion MNIST, CelebA, and CIFAR-10 datasets. Finally, we consider an application of the proposed formulation to address the important problem of learning an under-represented class in an unbalanced dataset. The Rumi approach results in substantially lower FID scores than the standard GAN frameworks while possessing better generalization capability. Siddarth Asokan, Chandra Sekhar Seelamantula |
NeurIPS | 2 |
| 2020 | Musical noise suppression using a low-rank and sparse matrix decomposition approach
Jishnu Sadasivan, Jitendra Kumar Dhiman, Chandra Sekhar Seelamantula |
Speech Commun. | 3 |
| 2020 | Speech Enhancement Using a Risk Estimation Approach
Jishnu Sadasivan, Chandra Sekhar Seelamantula, Nagarjuna Reddy Muraka |
Speech Commun. | 2 |
| 2020 | Minkowski-Algebra-Based Super-Sparse Array Design for Super-Resolution Ultrasound ImagingabstractSparse arrays give rise to the possibility of miniaturization of sensor arrays for ultrasound imaging. The design of sparse arrays that preserves the full-array beampattern has received much attention in the recent past. Based on the transmit/receive pair array synthesis model proposed by Hoctor and Kassam, issues related to producing effective apertures twice their corresponding physical arrays have been addressed. We consider two sparse array designs -SCOBA and SCOBAR - proposed recently and investigate the sparsity of their acquisition apertures and sub-apertures. We constructively formulate the problem of designing sparse arrays using fundamental properties from Minkowski algebra. Our investigation leads to two novel super-sparse array designs, which are sparser than SCOBA and SCOBAR and computationally more efficient for performing convolutional beamforming. They also have identical beampatterns and consequently similar lateral resolution as SCOBA and SCOBAR. We substantiate our findings using Field-II simulations. Amol G. Mahurkar, Chandra Sekhar Seelamantula |
IEEE Signal Process. Lett. | 2 |
| 2020 | Neuromorphic Fringe Projection ProfilometryabstractWe address the problem of 3-D reconstruction using neuromorphic cameras (also known as event-driven cameras), which are a new class of vision-inspired imaging devices. Neuromorphic cameras are becoming increasingly popular for solving image processing and computer vision problems as they have significantly lower data rates than conventional frame-based cameras. We develop a neuromorphic-camera-based Fringe Projection Profilometry (FPP) system. We use the Dynamic Vision Sensor (DVS) in the DAVIS346 neuromorphic camera for acquiring measurements. Neuromorphic FPP is faster than a single-line-scanning method. Also, unlike frame-based FPP, the efficacy of the proposed method is not limited by the background while acquiring measurements. The working principle of the DVS also allows one to efficiently handle shadows thereby preventing ambiguities during 2-D phase unwrapping. Ashish Rao Mangalore, Chandra Sekhar Seelamantula, Chetan Singh Thakur |
IEEE Signal Process. Lett. | 2 |
| 2019 | Automatic Segmentation of Optic Disc Using Affine Snakes in Gradient Vector FieldabstractThe optic disc is one of the prominent features of a retinal fundus image, and its segmentation is a critical component in automated retinal screening systems for ophthalmic anomalies, such as diabetic retinopathy and glaucoma. In this paper, we propose a novel method for optic disc segmentation using affine snakes, where the snake evolves using an affine transformation and requires a priori knowledge of the desired object shape. We determine the affine transformation parameters by first computing a force field on the image and then deforming the snake till the net force on the snake is zero. The affine snakes technique excels in its speed of convergence. This is attributed to the fact that only six parameters require optimization, the six parameters being the horizontal and vertical scaling, shearing and translation components of an affine transformation. Localization of the optic disc is done using normalized cross-correlation and segmentation is done using the affine snakes technique. This technique is tested on publicly available fundus image datasets, such as IDRiD, Drishti-GS, RIM-ONE, DRIONS-DB, and Messidor, with Dice In-dices of 0.943, 0.958, 0.933, 0.913, and 0.912, respectively. Sidhartha Dey, Kapil Tahiliani, J. R. Harish Kumar, Adithya Kumar Pediredla, Chandra Sekhar Seelamantula |
ICASSP | 5 |
| 2019 | A Spectro-temporal Technique for Estimating Aperiodicity and Voiced/unvoiced Decision Boundaries of Speech SignalsabstractIn contrast to a 1-D short-time analysis of speech, 2-D approaches aim at characterizing the speech signal attributes jointly in time and frequency. In this paper, we focus on the quasi-periodicity of a voiced spectro-temporal patch and quantify it by proposing an aperiodicity measure defined using the underlying frequency modulations in the patch. We further propose a time-frequency aperiodicity map obtained by overlapping and adding the aperiodicity measures across patches. The proposed aperiodicity map is utilized to obtain band-wise aperiodicity parameters, which are essential for high-quality speech synthesis. The aperiodicity in unvoiced patches is addressed by identifying them using the coherence of the patch. In addition, the proposed technique also provides voiced/unvoiced decisions boundaries of a speech signal. The effectiveness of the proposed band-wise aperiodicity parameters and voiced/unvoiced decisions is verified by incorporating them in an existing state-of-the-art vocoder for speech synthesis. Subjective listening tests show that the quality of the reconstructed speech is on par with that of the state-of-the-art WORLD vocoder in terms of mean opinion score, indicating that spectrotemporal approaches are highly promising for speech analysis and synthesis applications. Jitendra Kumar Dhiman, Chandra Sekhar Seelamantula |
ICASSP | 2 |
| 2019 | A Learning Approach for Wavelet DesignabstractWavelet analysis and perfect reconstruction filterbanks (PRFBs) are closely related. Desired properties on the wavelet could be translated to equivalent properties on a PRFB. We propose a new learning-based approach towards designing compactly supported orthonormal wavelets with a specified number of vanishing moments. We view PRFBs as a special class of convolutional autoencoders, which places the problem of wavelet/PRFB design within a learning framework. One could then deploy several state-of-the-art deep learning tools to solve the design problem. The PRFBs are learned by minimizing a squared-error loss function using gradient-descent optimization. The model is trained using a dataset containing random samples drawn from the standard normal distribution. We demonstrate that imposing orthonormality and vanishing moment constraints in the learning framework gives rise to filters that generate an orthonormal wavelet basis. We present results for learning PRFBs with filter lengths 2 and 8. As an illustration, we show that the proposed framework is able to learn the Daubechies wavelet with four vanishing moments, as well as wavelets with an arbitrary number of vanishing moments. For all our results, the signal-to-reconstruction error ratio is greater than 200 dB, implying that perfect reconstruction is indeed achieved accurately up to machine precision. Dhruv Jawali, Chandra Sekhar Seelamantula |
ICASSP | 3 |
| 2019 | FRI Modelling of Fourier DescriptorsabstractFourier descriptors are used to parametrically represent closed contours. In practice, a finite set of Fourier descriptors can model a large class of smooth contours. In this paper, we propose a method for estimating the Fourier descriptors of a given contour from its partial samples. We take a sampling-theoretic approach to model the x and y coordinate functions of the shape and express them as a sum of weighted complex exponentials, which belong to the class of finite-rate-of-innovation (FRI) signals. The weights represent the Fourier descriptors of the shape. We use the FRI framework to estimate the shape parameters reliably from noisy and partial measurements. We model non-uniformities in sampling using the sampling jitter model and employ a prefiltering process to reduce the effect of measurement noise and jitter. The average sampling interval is estimated by a block annihilating filter, which is then followed by the estimation of Fourier descriptors using least-squares fitting. We demonstrate the robustness of the proposed algorithm to noise and sampling jitter. Monte Carlo performance analysis shows that the variances of the estimators are close to the Cramér-Rao lower bounds. We present results for outlining shapes in synthetic as well as real images. Abijith Jagannath Kamath, Sunil Rudresh, Chandra Sekhar Seelamantula |
ICASSP | 3 |
| 2019 | Automatic Segmentation of Common Carotid Artery in Longitudinal Mode Ultrasound Images Using Active OblongsabstractWe propose a fully automated algorithm for the segmentation of common carotid artery in longitudinal mode ultrasound images using active oblongs. The problem of segmentation and subsequent delineation of lumen-intima layer is solved as an optimization of a locally defined contrast function with respect to five degrees-of-freedom that characterize the active oblong. The detection of the common carotid artery and subsequent initialization of the active oblong inside the common carotid artery region has been done using a combination of binary thresholding, Hough transform, and pixel-offset operations. The algorithm has been validated on the Brno university signal processing lab B-mode ultrasound image database, which contains 84 longitudinal mode ultrasound images of the common carotid artery. The segmentation results are validated against the ground truth provided by two practising radiologists using Jaccard and Dice similarity measures. We have achieved a detection and segmentation accuracy of 95.2% and 97.5%, respectively. J. R. Harish Kumar, Kartik Teotia, P. Kevin Raj, Jasbon Andrade, Rajagopal Kadavigere, Chandra Sekhar Seelamantula |
ICASSP | 6 |
| 2019 | SAMIR: Sparsity Amplified Iteratively-reweighted Beamforming for High-rsolution Ultrasound ImagingabstractIn ultrasound imaging, one typically employs delay-and-sum (DAS) beamformers for image reconstruction. An apodization window is used to suppress the side-lobes of an array beam pattern. The application of an apodization window to suppress the side-lobes widens the main-lobe width. We consider a statistical beamformer and present two variants. The signal of interest is modeled as a Laplacian-distributed random variable and additive interference components as Gaussian distributed. The resultant LASSO formulation is known to suffer from underestimation of large signal amplitudes due to the ℓ1-norm regularization. In the first variant, we reformulate the LASSO problem with a minimax-concave penalty (called Sparsity AMplified (SAM)) to contain the bias, thereby enhancing the beamformed image. A closed-form pointwise estimator is obtained for the optimization problem. In the second variant, we propose Sparsity AMplified Iteratively-Reweighted (SAMIR) beamforming algorithm, which leverages the properties of an apodization function. In SAMIR beamforming, we jointly optimize the cost over the signal of interest and the extrinsic apodization weights. This beamformer results in high-resolution ultrasound images, especially in the lateral direction. The proposed methods are compared with the standard DAS and a recently proposed statistically-modeled beamformer, iMAP, for a different number of plane-wave insonifications. Amol G. Mahurkar, Praveen Kumar Pokala, Chetan Singh Thakur, Chandra Sekhar Seelamantula |
ICASSP | 4 |
| 2019 | FirmNet: A Sparsity Amplified Deep Network for Solving Linear Inverse ProblemsabstractRecovering a sparse signal from a noisy linear measurement is an important problem in signal processing. Typically, one employs greedy pursuit techniques such as OMP, CoSaMP to solve an ℓ0regularization problem. For large-scale problems, iterative shrinkage techniques such as ISTA, FISTA, AMP-ℓ1have been introduced. The underlying formulation in the iterative algorithms is a LASSO problem with an ℓ1-penalty. It is known in the literature that an ℓ1-penalty in LASSO suffers from underestimation of large signal amplitudes. Also, the iterative shrinkage-based approaches such as ISTA typically have only one free parameter to trade-off between noise variance and sparsity. We consider a minimax-concave penalty-based formulation, which offers an unbiased estimate of the sparse signal. The resulting iterative firm-thresholding algorithm is restructured as a DNN architecture called FirmNet. The proposed network, FirmNet, has two interpretable shrinkage function parameters - one that controls the noise variance, and the other that allows for explicit sparsity control. We compare the network with a broader network architecture of Learned-ISTA (LISTA), and show that it outperforms in terms of the probability-of-error-in-support (PES) - a strong support recovery metric, by at least three-fold. We also observe an improvement of 2 to 4 dB in reconstruction SNR compared with LISTA. Praveen Kumar Pokala, Amol G. Mahurkar, Chandra Sekhar Seelamantula |
ICASSP | 3 |
| 2019 | Optic Disc Segmentation Using Cascaded Multiresolution Convolutional Neural NetworksabstractOptic disc segmentation is a crucial step in the development of automated tools for the detection and diagnosis of optical pathologies such as glaucoma. In this paper, we build upon our previous work, where we introduced the Fine-Net [1] - a Convolutional Neural Network (CNN) for optic disc segmentation. In this work, we introduce a prior CNN called the P-Net, which is arranged in cascade with the Fine-Net, to generate a more accurate optic disc segmentation map. The P-Net generates a low-resolution (256 × 256) segmentation map which is then further upscaled along with the input image and is fed to the Fine-Net, which yields a high-resolution segmentation map (1024 × 1024). Both CNNs are separately trained on publicly available datasets: DRISHTI-GS, MESSIDOR, and DRIONS-DB. We demonstrate the advantage of providing a prior segmentation map via the P-Net and further improve on our previous predictions. We obtain state-of-the-art results with an average Dice coefficient of 0.966 and Jaccard coefficient of 0.934. Dhruv Mohan, J. R. Harish Kumar, Chandra Sekhar Seelamantula |
ICIP | 3 |
| 2019 | A Structure Tensor Based Voronoi Decomposition Technique for Optic Cup SegmentationabstractWe present a technique for segmentation of optic cup based on the structural features found in blood vessels surrounding the optic cup region. The advantage of using such features is that they are robust to variations in the properties of the fundus image such as brightness, contrast, etc. The main features used in the technique are vessel bends (also called as landmark points or kinks), which are identified by applying the Harris corner detection algorithm on the optic disc region, followed by a Voronoi image decomposition. Pratt's circle fitting algorithm is employed on the extracted landmark points to segment the optic cup region. The proposed technique is validated on a total of 191 images taken from publicly available fundus image datasets, namely, Drishti-GS and MESSIDOR. Performance metrics such as sensitivity, specificity, accuracy, Jaccarďs index, and Dice coefficient are computed to be 85%, 97%, 96%, 69.5%, and 81%, respectively, which indicates that the proposed technique for optic cup segmentation is competitive with the state-of-the-art methods. P. Kevin Raj, J. R. Harish Kumar, S. P. Subramanya Jois, S. Harsha, Chandra Sekhar Seelamantula |
ICIP | 5 |
| 2019 | On the Suitability of the Riesz Spectro-Temporal Envelope for WaveNet Based Speech SynthesisabstractWe address the problem of estimating the time-varying spectral envelope of a speech signal using a spectro-temporal demodulation technique. Unlike the conventional spectrogram, we consider a pitch-adaptive spectrogram and model a spectro-temporal patch using an amplitude- and frequency-modulated two-dimensional (2-D) cosine signal. We employ a demodulation technique based on the Riesz transform that we proposed recently to estimate the amplitude and frequency modulations. The amplitude modulation (AM) corresponds to the vocal-tract filter magnitude response (or envelope) and the frequency modulation (FM) corresponds to the excitation. We consider the AM and demonstrate its effectiveness by incorporating it as an acoustic feature for local conditioning in the statistical WaveNet vocoder for the task of speech synthesis. The quality of the synthesized speech obtained with the Riesz envelope is compared with that obtained using the envelope estimated by the WORLD vocoder. Objective measures and subjective listening tests on the CMU-Arctic database show that the quality of synthesis is superior to that obtained using the WORLD envelope. This study thus establishes the Riesz envelope as an efficient alternative to the WORLD envelope. Jitendra Kumar Dhiman, Nagaraj Adiga, Chandra Sekhar Seelamantula |
INTERSPEECH | 3 |
| 2018 | Phasesplit: A Variable Splitting Framework for Phase RetrievalabstractWe develop two techniques based on alternating minimization and alternating directions method of multipliers for phase retrieval (PR) by employing a variable-splitting approach in a maximum likelihood estimation framework. This leads to an additional equality constraint, which is incorporated in the optimization framework using a quadratic penalty. Both algorithms are iterative, wherein the updates are computed in closed-form. Experimental results show that: (i) the proposed techniques converge faster than the state-of-the-art PR algorithms; (ii) the complexity is comparable to the state of the art; and (iii) the performance does not depend critically on the choice of the penalty parameter. We also show how sparsity can be incorporated within the variable splitting framework and demonstrate concrete applications to image reconstruction in frequency-domain optical-coherence tomography. Subhadip Mukherjee, Suprosanna Shit, Chandra Sekhar Seelamantula |
ICASSP | 3 |
| 2018 | Wavelet-Based Reconstruction for Unlimited SamplingabstractSelf-reset analog-to-digital converters (ADCs) allow for digitization of a signal with a high dynamic range. The reset action is equivalent to a modulo operation performed on the signal. We consider the problem of recovering the original signal from the measured modulo-operated signal. In our formulation, we assume that the underlying signal is Lipschitz continuous. The modulo-operated signal can be expressed as the sum of the original signal and a piecewise-constant signal that captures the transitions. The reconstruction requires estimating the piecewise-constant signal. We rely on local smoothness of the modulo-operated signal and employ wavelets with sufficient vanishing moments to suppress the polynomial component. We employ Daubechies wavelets, which are most compact for a given number of vanishing moments. The wavelet filtering provides a sequence consisting of a sum of scaled and shifted versions of a kernel derived from the wavelet filter. The transition locations are estimated from the sequence using a sparse recovery technique. We derive a sufficient condition on the sampling frequency for ensuring perfect reconstruction of the smooth signal. We validate our reconstruction technique on a signal consisting of sinusoids in both clean and noisy conditions and compare the reconstruction quality with the recently developed repeated finite-difference method. Sunil Rudresh, Aniruddha Adiga, Basty Ajay Shenoy, Chandra Sekhar Seelamantula |
ICASSP | 4 |
| 2018 | Design of Sampling Kernels and Sampling Rates for Two-Dimensional Finite Rate of Innovation SignalsabstractThis paper focuses on the design of sampling kernels and critical sampling rates for reconstruction of two-dimensional finite-rate-of-innovation (2D-FRI) signals from their samples. A class of aperiodic 2D sum-of-modulated-splines (2D-SMS) sampling kernels and sampling rates lower than the other state-of-the-art techniques are proposed. The 2D kernels proposed do not require a periodization to handle aperiodic signals. The proposed design also allows for reduced computation complexity. The reconstruction proceeds using the standard FRI reconstruction machinery. We demonstrate successful recovery using signals that are a sum of 2D Gaussians. The reconstruction performance of various members of the proposed family of kernels is evaluated through Monte Carlo simulations and compared against the Cramér-Rae lower bound (CRLB) for various noise levels. Anindita De, Chandra Sekhar Seelamantula |
ICIP | 2 |
| 2018 | Automatic Segmentation of Lumen Intima Layer in Transverse Mode Ultrasound ImagesabstractWe propose an elliptical active disc technique for the segmentation of common carotid artery lumen intima layer from transverse mode ultrasound images. The segmentation and subsequent outlining problem is posed as one of optimization of a local energy function with respect to the five degrees-of-freedom that characterize the elliptical active disc. Gradient descent technique is used to find the minimum of the energy function with respect to the five parameters that describe the disc. In addition, we use Green's theorem to optimize the computation of the partial derivatives. For automatic initialization of the active disc, we use the normalized crosscorrelation technique. We report results of experimental validation on SPLab, Brno university database, which contains 971 transverse mode ultrasound images of the carotid artery. We achieve accurate carotid artery lumen intima detection in 97.63% of cases. In addition, for lumen intima layer segmentation we achieve an average Dice index of 94.83%. J. R. Harish Kumar, Chandra Sekhar Seelamantula, Jasbon Andrade, Rajagopal Kadavigere |
ICIP | 2 |
| 2018 | High-Performance Optic Disc Segmentation Using Convolutional Neural NetworksabstractWe present a framework for robust optic disc segmentation using convolutional neural networks. Optic disc is an important anatomical landmark in the fundus image used for the diagnosis of ophthalmological pathologies. Our objective is to develop a system for unsupervised, early and robust detection of diseases such as glaucoma. We introduce the Fine-Net, which generates a high-resolution optic disc segmentation map (1024 × 1024) from retinal fundus images. The network is trained on three publicly available datasets, MESSI-DOR, DRIONS-DB, and DRISHTI-GS. The proposed framework generalizes well as it performs reliably even on test images that have a significant variability. For experimental evaluation, we perform a five-fold cross-validation and achieve accurate optic disc localization in 99.4% of cases. Moreover, for optic disc segmentation we achieve an average Dice coefficient and Jaccard coefficient of 0.958 and 0.921, respectively. Dhruv Mohan, J. R. Harish Kumar, Chandra Sekhar Seelamantula |
ICIP | 3 |
| 2018 | Multicomponent 2-D AM-FM Modeling of Speech SpectrogramsabstractIn contrast to 1-D short-time analysis of speech, 2-D modeling of spectrograms provides a characterization of speech attributes directly in the joint time-frequency plane. Building on existing 2-D models to analyze a spectrogram patch, we propose a multicomponent 2-D AM-FM representation for spectrogram decomposition. The components of the proposed representation comprise a DC, a fundamental frequency carrier and its harmonics, and a spectrotemporal envelope, all in 2-D. The number of harmonics required is patch-dependent. The estimation of the AM and FM is done using the Riesz transform, and the component weights are estimated using a least-squares approach. The proposed representation provides an improvement over existing state-of-the-art approaches, for both male and female speakers. This is quantified using reconstruction SNR and perceptual evaluation of speech quality (PESQ) metric. Further, we perform an overlap-add on the DC component, pooling all the patches and obtain a time-frequency (t-f) a periodicity map for the speech signal. We verify its effectiveness in improving speech synthesis quality by using it in an existing state-of-the art vocoder. Jitendra Kumar Dhiman, Chandra Sekhar Seelamantula |
INTERSPEECH | 3 |
| 2018 | Speech Enhancement Using the Minimum-probability-of-error CriterionabstractWe propose a novel speech denoising framework by minimizing the probability of error (PE), which measures the deviation probability of the estimate from its true value. To develop the minimum PE (MPE) criterion, one requires the knowledge of the noise probability density function (p.d.f.), which may not be available in a parametric form in speech denoising applications. Therefore, we adopt two approaches for modeling the noise p.d.f.: (i) Gaussian modeling based on adaptive variance estimation; and (ii) a Gaussian mixture model (GMM) in view of its approximation capabilities. We consider discrete cosine transform (DCT) domain shrinkage, where the optimum shrinkage parameter is obtained by minimizing an estimate of the PE. A performance assessment for real-world noise types shows that for input signal-to-noise ratios (SNR) greater than 5 dB, the proposed MPE-based point-wise shrinkage estimators outperform three benchmark techniques in terms of segmental SNR and short-time objective intelligibility (STOI) scores. Jishnu Sadasivan, Subhadip Mukherjee, Chandra Sekhar Seelamantula |
INTERSPEECH | 3 |
| 2018 | An Optimization Framework for Recovery of Speech from Phase-Encoded SpectrogramsabstractIn general, reconstruction of a speech signal from the spectrogram is non-unique because of the unavailability of the phase spectrum. Considering zero phase would result in a minimum phase reconstruction. This limitation is overcome by computing the recently introduced phase-encoded spectrogram. In this approach, one modifies each frame of a speech signal to possess the causal, delta-dominant (CDD) property prior to computing the spectrogram. In an earlier publication, we showed that finite-length CDD sequences can be retrieved exactly from their magnitude spectra using a scepstrum technique. Although exactness is guaranteed in principle, practical implementations result in a limited, but high, reconstruction accuracy. In this paper, we focus on increasing the reconstruction accuracy. We formulate the reconstruction problem within an optimization framework and deploy a recently proposed iterative, alternating direction method of multipliers (ADMM) algorithm called autocorrelation retrieval Kolmogorov factorization (CoRK). Experimental validations show that the CoRK algorithm results in a reconstruction accurate up to machine precision. We also show that both CoRK and cepstrum techniques are robust and invariant to the choice of the window duration, the amount of overlap between consecutive speech frames, the strength of the delta used to impart the CDD property, and the presence of noise. Abhilash Sainathan, Sunil Rudresh, Chandra Sekhar Seelamantula |
INTERSPEECH | 3 |
| 2018 | Training-Free, Single-Image Super-Resolution Using a Dynamic Convolutional NetworkabstractThe typical approach for solving the problem of single-image super-resolution (SR) is to learn a nonlinear mapping between the low-resolution (LR) and high-resolution (HR) representations of images in a training set. Training-based approaches can be tuned to give high accuracy on a given class of images, but they call for retraining if the HR → LR generative model deviates or if the test images belong to a different class, which limits their applicability. On the other hand, we propose a solution that does not require a training dataset. Our method relies on constructing a dynamic convolutional network (DCN) to learn the relation between the consecutive scales of Gaussian and Laplacian pyramids. The relation is in turn used to predict the detail at a finer scale, thus leading to SR. Comparisons with state-of-the-art techniques on standard datasets show that the proposed DCN approach results in about 0.8 and 0.3 dB gain in peak signal-to-noise ratio for 2× and 3× SR, respectively. The structural similarity index is on par with the competing techniques. Aritra Bhowmik, Suprosanna Shit, Chandra Sekhar Seelamantula |
IEEE Signal Process. Lett. | 3 |
| 2018 | Phase Retrieval From Binary MeasurementsabstractWe consider the problem of signal reconstruction from quadratic measurements that are encoded as +1 or -1 depending on whether they exceed a predetermined positive threshold or not. Binary measurements are fast to acquire and inexpensive in terms of hardware. We formulate the problem of signal reconstruction using a consistency criterion, wherein one seeks to find a signal that is in agreement with the acquired measurements. To enforce consistency, we construct a convex cost using a one-sided quadratic penalty and minimize it using an iterative accelerated projected gradient-descent technique. The projected gradient-descent (PGD) scheme reduces the cost function in each iteration, whereas incorporating momentum into PGD, notwithstanding the lack of such a descent property, exhibits faster convergence than PGD empirically. We refer to the resulting algorithm as binary phase retrieval (BPR). Considering additive white noise contamination prior to quantization, we also derive the Cramér-Rao Bound (CRB) for the binary encoding model. Experimental results demonstrate that the BPR algorithm yields a signal-to-reconstruction error ratio (SRER) of approximately 25 dB in the absence of noise. In the presence of noise prior to quantization, the SRER is within 2 to 3 dB of the CRB. Subhadip Mukherjee, Chandra Sekhar Seelamantula |
IEEE Signal Process. Lett. | 2 |
| 2018 | TDOA-Based Multiple Acoustic Source Localization Without Association AmbiguityabstractMultiple source localization using time-differences of arrival (TDOAs) is challenging because of the ambiguity involved in associating the TDOAs computed across microphone pairs to the sources. We show that the association ambiguity of the TDOAs can be effectively resolved using the concept of an inverse delay interval region (IDIR), which we introduce in this paper. By examining the association between a spatial domain and the TDOAs, we define IDIR as an interhyperboloidal spatial region corresponding to an interval of delays for a given pair of microphones. The proposed scheme for localizing multiple sources involves two stages. In the first stage, the given enclosure is partitioned into nonoverlapping elemental regions and the ones that contain a source are detected using a measure based on the generalized cross-correlation with phase transform and the IDIRs. In the second stage, the sources are finely localized within each of the detected elemental regions by identifying the IDIRs containing a single source and a novel region-constrained localization approach. We evaluate the performance of the proposed approach on real recordings from the AV16.3 corpus and in a simulated reverberation setting with a reverberation time RT60 of up to 500 ms, and show that the DOA estimation error with two active speakers is within 2° and the spatial localization error is less than 30 cm for each speaker. Sundar Harshavardhan, Thippur V. Sreenivas, Chandra Sekhar Seelamantula |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | Automatic delineation of macular regions based on a locally defined contrast functionabstractWe consider the problem of fovea segmentation and develop a technique for delineation of macular regions based on the active-disc formalism that we recently introduced. The outlining problem is posed as one of the optimization of a locally defined contrast function using gradient-ascent maximization with respect to the affine transformation parameters that characterize the active disc. For automatic localization of the fovea and initialization of the active disc, we use the directional-derivative-based matched filter. We report validation results on three publicly available fundus image databases, amounting to a total of 1370 fundus images for automatic fovea localization and 370 fundus images for fovea segmentation and macular regions delineation. The proposed method results in a fovea localization accuracy of 100%, 92%, and 99.4%, and an average Dice similarity index of 77.78%, 67.46%, and 76.56% on DRIVE, DIARETDB0, and MESSIDOR fundus image databases, respectively. We have also developed an ImageJ plugin and an iOS App based on the proposed method. J. R. Harish Kumar, Rittwik Adhikari, Yogish S. Kamath, Rajani Jampala, Chandra Sekhar Seelamantula |
ICIP | 5 |
| 2017 | A Spectro-Temporal Demodulation Technique for Pitch EstimationabstractWe consider a two-dimensional demodulation framework for spectro-temporal analysis of the speech signal. We construct narrowband (NB) speech spectrograms, and demodulate them using the Riesz transform, which is a two-dimensional extension of the Hilbert transform. The demodulation results in timefrequency envelope (amplitude modulation or AM) and timefrequency carrier (frequency modulation or FM). The AM corresponds to the vocal tract and is referred to as the vocal tract spectrogram. The FM corresponds to the underlying excitation and is referred to as the carrier spectrogram. The carrier spectrogram exhibits a high degree of time-frequency consistency for voiced sounds. For unvoiced sounds, such a structure is lacking. In addition, the carrier spectrogram reflects the fundamental frequency (F0) variation of the speech signal. We develop a technique to determine the F0 from the carrier spectrogram. The time-frequency consistency is used to determine which time-frequency regions correspond to voiced segments. Comparisons with the state-of-the-art F0 estimation algorithms show that the proposed F0 estimator has high accuracy for telephone channel speech and is robust to noise. Jitendra Kumar Dhiman, Nagaraj Adiga, Chandra Sekhar Seelamantula |
INTERSPEECH | 3 |
| 2017 | Time-Frequency Coherence for Periodic-Aperiodic Decomposition of Speech SignalsabstractDecomposing speech signals into periodic and aperiodic components is an important task, finding applications in speech synthesis, coding, denoising, etc. In this paper, we construct a time-frequency coherence function to analyze spectro-temporal signatures of speech signals for distinguishing between deterministic and stochastic components of speech. The narrowband speech spectrogram is segmented into patches, which are represented as 2-D cosine carriers modulated in amplitude and frequency. Separation of carrier and amplitude/frequency modulations is achieved by 2-D demodulation using Riesz transform, which is the 2-D extension of Hilbert transform. The demodulated AM component reflects contributions of the vocal tract to spectrogram. The frequency modulated carrier (FM-carrier) signal exhibits properties of the excitation. The time-frequency coherence is defined with respect to FM-carrier and a coherence map is constructed, in which highly coherent regions represent nearly periodic and deterministic components of speech, whereas the incoherent regions correspond to unstructured components. The coherence map shows a clear distinction between deterministic and stochastic components in speech characterized by jitter, shimmer, lip radiation, type of excitation, etc. Binary masks prepared from the time-frequency coherence function are used for periodic-aperiodic decomposition of speech. Experimental results are presented to validate the efficiency of the proposed method. Karthika Vijayan, Jitendra Kumar Dhiman, Chandra Sekhar Seelamantula |
INTERSPEECH | 3 |
| 2016 | A divide-and-conquer dictionary learning algorithm and its performance analysisabstractWe address the problem of learning a sparsifying synthesis dictionary over large datasets that occur in numerous signal and image processing applications, such as inpainting, super-resolution, etc. We develop a dictionary learning algorithm that exploits the similarity of the training examples to reduce the training time. Training datasets containing correlated examples typically occur in image processing applications, as the datasets contain the patches extracted from natural images as training vectors. Our algorithm employs a divide- and-conquer approach, where one leverages the correlation within the training examples to segment the dataset into clusters containing similar examples, and learn local dictionaries for each of them. This constitutes the divide step of the algorithm. In the conquer step, a global dictionary is trained using the atoms of the local dictionaries as the training examples. We analyze the run-time complexity and the representation error of the proposed divide-and-conquer dictionary learning algorithm, and compare the performance with the batch and online dictionary learning algorithms, both on synthesized dataset and natural images. The analysis reveals that the proposed algorithm has an asymptotic complexity that is linear and logarithmic in the number of training examples, corresponding to sequential and parallel implementations, respectively. Subhadip Mukherjee, Chandra Sekhar Seelamantula |
ICASSP | 2 |
| 2016 | Joint dictionary training for bandwidth extension of speech signalsabstractWe address the problem of extending the bandwidth of speech signals, which is of importance to enhance the quality and intelligibility of the telephone speech. The low-pass filtering effect of the telephone communication channels eliminate the high-frequency components of the speech signal, and it is necessary to retrieve those to maintain the speech quality. We adopt a joint-dictionary training approach to recover the missing spectral information. By exploiting the sparsity of the spectrogram frames, the dictionaries for the wide-band (WB) and the corresponding narrow-band (NB) spectrogram frames are trained in a coupled manner in order to learn the mapping from NB to WB frames. We refer to this approach as the joint dictionary training for bandwidth extension (JDTBE). To ensure that the reconstructed bandwidth-extended speech is consistent with the measurement, we propose to apply a suitable affine transformation that depends on the properties of the telephone channel. We study the effect of the choice of sparsity on the quality of the reconstructed speech, for both male and female speakers. A comparison of the proposed JDTBE algorithm with a bandwidth extension technique based on stochastic modeling reveals the superiority of the JDTBE approach in terms of subjective listening test scores. Jishnu Sadasivan, Subhadip Mukherjee, Chandra Sekhar Seelamantula |
ICASSP | 3 |
| 2016 | An unbiased risk estimator for Gaussian mixture noise distributions - Application to speech denoisingabstractWe develop an unbiased estimate of mean-squared error (MSE), where the observations are assumed to be drawn from a Gaussian mixture (GM) distribution. Stein's unbiased risk estimate (SURE) is an unbiased estimate of the MSE, and was originally proposed for independent and identically distributed (i.i.d.) multivariate Gaussian observations. Subsequently, it was extended to the exponential family of distributions. In this paper, we extend the idea of SURE to observations drawn from a Gaussian mixture distribution (GMD). Since Gaussian mixture models (GMM) can model any given distribution sufficiently accurately, this generalized framework allows us to apply the SURE technique to the observations drawn from an arbitrary distribution. As an application, we consider the problem of denoising speech corrupted by a GM distributed noise. It is observed that the denoising performance of the algorithm developed using SURE based on GMD is superior in terms of the signal-to-noise ratio (SNR) and average segmental SNR (ASSNR), compared with that obtained using SURE based on the single Gaussian assumption. Jishnu Sadasivan, Chandra Sekhar Seelamantula |
ICASSP | 2 |
| 2016 | Active-disc-based Kalman filter technique for tracking of blood cells in microfluidic channelsabstractIdentification and tracking of red blood cells and leukocytes in micro-circulation is very important for drug validation and inflammation. Manual tracking is usually used for this purpose. However, due to its excessive time consumption and inaccuracy in the case of overlapped cells, there is a need for automated methods. In this paper, we propose a Kalman filter and active-disc-based method for automatic and accurate detection of cells. In addition to tracking both slow and fast moving blood cells, the proposed method is also able to track overlapped cells successfully with an overall tracking accuracy of 95.7% and at a processing rate of 25 ms per frame. Comparisons with five state-of-the-art methods show that the proposed method is better in terms of the detection rate and accuracy of instantaneous speed estimation. Badarish Colathur Arvind, Sujith Kumar Nagaraj, Chandra Sekhar Seelamantula, Gorthi Sai Siva |
ICIP | 3 |
| 2016 | Automatic segmentation of common carotid artery in transverse mode ultrasound imagesabstractWe consider the problem of carotid artery segmentation and develop an automated outlining technique based on the active disc formalism that we recently introduced. The outlining problem is posed as one of optimization of a locally defined contrast function with respect to the affine transformation parameters that characterize the active disc. It turns out that standard techniques based on gradient-descent minimization can be used to carry out the optimization, although more sophisticated optimizers could also be deployed. For the initialization, we use a matched filter with a template size chosen based on an estimate of the average size of the carotid artery. We report results of experimental validation on Brno university's signal processing (SP) lab database, which contains 971 transverse mode ultrasound images of the carotid artery. The images in the database are manually annotated using a circle with center and radius explicitly specified in pixels, which serves as the reference. The circular annotation is also a good match with the active disc template considered in this paper. The proposed method results in an average detection accuracy of 95.5% and an average Dice similarity measure of 87.36% and takes only a few seconds of processing time per image. Comparisons with other state-of-the-art techniques are also reported. J. R. Harish Kumar, Chandra Sekhar Seelamantula, Nikhil S. Narayan, Pina Marziliano |
ICIP | 2 |
| 2016 | Reverberation-Robust One-Bit TDOA Based Moving Source Localization for Automatic Camera SteeringabstractWe address the problem of moving acoustic source localization and automatic camera steering using one-bit measurement of the time-difference of arrival (TDOA) between two microphones in a given array. Given that the camera has a finite field of view (FoV), an algorithm with a coarse estimate of the source location would suffice for the purpose. We use a microphone array and develop an algorithm to obtain a coarse estimate of the source using only one-bit information of the TDOA, the sign of it, to be precise. One advantage of the one-bit approach is that the computational complexity is lower, which aids in real-time adaptation and localization of the moving source. We carried out experiments in a reverberant enclosure with a 60 dB reverberation time of 600 ms (RT60 = 600 ms). We analyzed the performance of the proposed approach using a circular microphone array. We report comparisons with a point source localization based automatic camera steering algorithm proposed in the literature. The proposed algorithm turned out to be more accurate in terms of always having the moving speaker within the field of view. Sundar Harshavardhan, Gokul Deepak Manavalan, Thippur V. Sreenivas, Chandra Sekhar Seelamantula |
INTERSPEECH | 4 |
| 2016 | A Novel Risk-Estimation-Theoretic Framework for Speech Enhancement in Nonstationary and Non-Gaussian Noise ConditionsabstractWe address the problem of suppressing background noise from noisy speech within a risk estimation framework, where the clean signal is estimated from the noisy observations by minimizing an unbiased estimate of a chosen risk function. For Gaussian noise, such a risk estimate was derived by Stein, which eventually went on to be called Stein's unbiased risk estimate (SURE). Stein's formalism is restricted to Gaussian noise and exclusive risk estimators have been developed for each noise type. On the other hand, we consider linear denoising functions and derive an unbiased risk estimate without making any assumption about the noise distribution. The proposed unbiased estimate depends only on the second-order statistics of noise and makes the proposed framework applicable to many practical denoising problems where the noise distribution is not known a priori, but one has access only to the samples of noise. We demonstrate the usefulness of the proposed methodology for speech enhancement using subband shrinkage, where the shrinkage parameters are obtained by minimizing the newly developed risk estimator. The proposed methodology is also applicable to nonstationary noise conditions. We show that the proposed denoising algorithm outperforms the state-of-the art algorithms in terms of standard speech-quality evaluation metrics. Jishnu Sadasivan, Chandra Sekhar Seelamantula |
INTERSPEECH | 2 |
| 2016 | Phase-Encoded Speech SpectrogramsabstractSpectrograms of speech and audio signals are time-frequency densities, and by construction, they are non-negative and do not have phase associated with them. Under certain conditions on the amount of overlap between consecutive frames and frequency sampling, it is possible to reconstruct the signal from the spectrogram. Deviating from this requirement, we develop a new technique to incorporate the phase of the signal in the spectrogram by satisfying what we call as the delta dominance condition, which in general is different from the well known minimum-phase condition. In fact, there are signals that are delta dominant but not minimum-phase and vice versa. The delta dominance condition can be satisfied in multiple ways, for example by placing a Kronecker impulse of the right amplitude or by choosing a suitable window function. A direct consequence of this novel way of constructing the spectrograms is that the phase of the signal is directly encoded or embedded in the spectrogram. We also develop a reconstruction methodology that takes such phase-encoded spectrograms and obtains the signal using the discrete Fourier transform (DFT). It is envisaged that the new class of phase-encoded spectrogram representations would find applications in various speech processing tasks such as analysis, synthesis, enhancement, and recognition. Chandra Sekhar Seelamantula |
INTERSPEECH | 1 |
| 2016 | ℓ1-K-SVD: A robust dictionary learning algorithm with simultaneous update
Subhadip Mukherjee, Rupam Basu, Chandra Sekhar Seelamantula |
Signal Process. | 3 |
| 2016 | A risk minimization framework for channel estimation in OFDM systems
Karthik Upadhya, Chandra Sekhar Seelamantula, K. V. S. Hari |
Signal Process. | 2 |
| 2016 | Ellipse Fitting Using the Finite Rate of Innovation Sampling PrincipleabstractStandard approaches for ellipse fitting are based on the minimization of algebraic or geometric distance between the given data and a template ellipse. When the data are noisy and come from a partial ellipse, the state-of-the-art methods tend to produce biased ellipses. We rely on the sampling structure of the underlying signal and show that the x - and y -coordinate functions of an ellipse are finite-rate-of-innovation (FRI) signals, and that their parameters are estimable from partial data. We consider both uniform and nonuniform sampling scenarios in the presence of noise and show that the data can be modeled as a sum of random amplitude-modulated complex exponentials. A low-pass filter is used to suppress noise and approximate the data as a sum of weighted complex exponentials. The annihilating filter used in FRI approaches is applied to estimate the sampling interval in the closed form. We perform experiments on simulated and real data, and assess both objective and subjective performances in comparison with the state-of-the-art ellipse fitting methods. The proposed method produces ellipses with lesser bias. Furthermore, the mean-squared error is lesser by about 2 to 10 dB. We show the applications of ellipse fitting in iris images starting from partial edge contours, and to free-hand ellipses drawn on a touch-screen tablet. Satish Mulleti, Chandra Sekhar Seelamantula |
IEEE Trans. Image Process. | 2 |
| 2015 | Periodic non-uniform sampling for FRI signalsabstractA typical finite-rate-of-innovation (FRI) signal reconstruction scheme is based on the measurement of uniform samples in time/frequency domain, and the application of the annihilating filter on the measured samples. We propose a continuous-time annihilation framework for a class of FRI signals. In particular, we show that FRI signals of sum-of-weighted exponential form can be annihilated by a composition of translation operators and show that the parameters of the signal can be estimated in the periodic non-uniform sampling (PNU) scenario. We discuss the advantages of PNU sampling over uniform sampling and extend it for general FRI signal sampling and reconstruction. Simulations are performed and the results are compared with state-of-the-art methods for signal-to-noise ratios ranging from -20 to 100 dB. An improvement in the estimation accuracy of 15-55 dB in terms of bias and mean-square error is achieved over conventional methods by a rearrangement of uniform samples to follow PNU sampling. Satish Mulleti, Chandra Sekhar Seelamantula |
ICASSP | 2 |
| 2015 | FRI sampling and reconstruction of asymmetric pulsesabstractWe consider the problem of modelling asymmetric pulse trains as finite-rate-of-innovation (FRI) signals. In particular, we show that the sum of amplitude-scaled and time-shifted pulses with different asymmetry factors is an FRI signal. Such signals frequently arise in applications such as ultrasound and radio detection and ranging (RADAR) where the received signal has skewed pulses. In this paper, we model the asymmetric component of a pulse using its derivative. A sampling kernel with a sum-of-sincs frequency response is used to measure the samples, and a modified annihilating filter method is applied on the samples to estimate the parameters of the FRI signal. We show accurate reconstruction for signals containing asymmetric Gaussian, Cauchy-Lorentz, and sinc pulses. Analysis of the proposed scheme in the presence of noise shows that the error in the estimated parameters decreases by oversampling the signal. Sudarshan Nagesh, Chandra Sekhar Seelamantula |
ICASSP | 2 |
| 2015 | Directional bilateral filtersabstractWe propose a bilateral filter with a locally controlled domain kernel for directional edge-preserving smoothing. Traditional bilateral filters use a range kernel, which is responsible for edge preservation, and a fixed domain kernel that performs smoothing. Our intuition is that orientation and anisotropy of image structures should be incorporated into the domain kernel while smoothing. For this purpose, we employ an oriented Gaussian domain kernel locally controlled by a structure tensor. The oriented domain kernel combined with a range kernel forms the directional bilateral filter. The two kernels assist each other in effectively suppressing the influence of the outliers while smoothing. To find the optimal parameters of the directional bilateral filter, we propose the use of Stein's unbiased risk estimate (SURE). We test the capabilities of the kernels separately as well as together, first on synthetic images, and then on real endoscopic images. The directional bilateral filter has better denoising performance than the Gaussian bilateral filter at various noise levels in terms of peak signal-to-noise ratio (PSNR). Manasij Venkatesh, Chandra Sekhar Seelamantula |
ICASSP | 2 |
| 2015 | Image denoising in multiplicative noiseabstractWe address the problem of denoising images corrupted by multiplicative noise. The noise is assumed to follow a Gamma distribution. Compared with additive noise distortion, the effect of multiplicative noise on the visual quality of images is quite severe. We consider the mean-square error (MSE) cost function and derive an expression for an unbiased estimate of the MSE. The resulting multiplicative noise unbiased risk estimator is referred to as MURE. The denoising operation is performed in the wavelet domain by considering the image-domain MURE. The parameters of the denoising function (typically, a shrinkage of wavelet coefficients) are optimized for by minimizing MURE. We show that MURE is accurate and close to the oracle MSE. This makes MURE-based image denoising reliable and on par with oracle-MSE-based estimates. Analogous to the other popular risk estimation approaches developed for additive, Poisson, and chi-squared noise degradations, the proposed approach does not assume any prior on the underlying noise-free image. We report denoising results for various noise levels and show that the quality of denoising obtained is on par with the oracle result and better than that obtained using some state-of-the-art denoisers. Chandra Sekhar Seelamantula, Thierry Blu |
ICIP | 1 |
| 2015 | A Zero-Crossing Rate Property of Power Complementary Analysis Filterbank OutputsabstractWe establish zero-crossing rate (ZCR) relations between the input and the subbands of a maximally decimated M-channel power complementary analysis filterbank when the input is a stationary Gaussian process. The ZCR at lag l is defined as the number of sign changes between the samples of a sequence and its l-sample shifted version, normalized by the sequence length. We derive the relationship between the ZCR of the Gaussian process at lags that are integer multiples of M and the subband ZCRs. Based on this result, we propose a robust iterative autocorrelation estimator for a signal consisting of a sum of sinusoids of fixed amplitudes and uniformly distributed random phases. Simulation results show that the performance of the proposed estimator is better than the sample autocorrelation over the SNR range of - 6 to 15 dB. Validation on a segment of a trumpet signal showed similar performance gains. Ravi R. Shenoy, Chandra Sekhar Seelamantula |
IEEE Signal Process. Lett. | 2 |
| 2015 | Demodulation of Narrowband Speech Spectrograms Using the Riesz TransformabstractWe propose a two-dimensional (2-D) multicomponent amplitude-modulation, frequency-modulation (AM-FM) model for a spectrogram patch corresponding to voiced speech, and develop a new demodulation algorithm to effectively separate the AM, which is related to the vocal tract response, and the carrier, which is related to the excitation. The demodulation algorithm is based on the Riesz transform and is developed along the lines of Hilbert-transform-based demodulation for 1-D AM-FM signals. We compare the performance of the Riesz transform technique with that of the sinusoidal demodulation technique on real speech data. Experimental results show that the Riesz-transform-based demodulation technique represents spectrogram patches accurately. The spectrograms reconstructed from the demodulated AM and carrier are inverted and the corresponding speech signal is synthesized. The signal-to-noise ratio (SNR) of the reconstructed speech signal, with respect to clean speech, was found to be 2 to 4 dB higher in case of the Riesz transform technique than the sinusoidal demodulation technique. Haricharan Aragonda, Chandra Sekhar Seelamantula |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Ellipse fitting using finite rate of innovation principlesabstractWe address the problem of parameter estimation of an ellipse from a limited number of samples. We develop a new approach for solving the ellipse fitting problem by showing that the x and y coordinate functions of an ellipse are finite-rate-of-innovation (FRI) signals. Uniform samples of x and y coordinate functions of the ellipse are modeled as a sum of weighted complex exponentials, for which we propose an efficient annihilating filter technique to estimate the ellipse parameters from the samples. The FRI framework allows for estimating the ellipse parameters reliably from partial or incomplete measurements even in the presence of noise. The efficiency and robustness of the proposed method is compared with state-of-art direct method. The experimental results show that the estimated parameters have lesser bias compared with the direct method and the estimation error is reduced by 5-10 dB relative to the direct method. Satish Mulleti, Chandra Sekhar Seelamantula |
ICASSP | 2 |
| 2014 | On the role of the Hilbert transform in boosting the performance of the annihilating filterabstractWe consider the problem of parameter estimation from real-valued multi-tone signals. Such problems arise frequently in spectral estimation. More recently, they have gained new importance in finite-rate-of-innovation signal sampling and reconstruction. The annihilating filter is a key tool for parameter estimation in these problems. The standard annihilating filter design has to be modified to result in accurate estimation when dealing with real sinusoids, particularly because the real-valued nature of the sinusoids must be factored into the annihilating filter design. We show that the constraint on the annihilating filter can be relaxed by making use of the Hilbert transform. We refer to this approach as the Hilbert annihilating filter approach. We show that accurate parameter estimation is possible by this approach. In the single-tone case, the mean-square error performance increases by 6 dB for signal-to-noise ratio (SNR) greater than 0 dB. We also present experimental results in the multi-tone case, which show that a significant improvement (about 6dB) is obtained when the parameters are close to 0 or π. In the mid-frequency range, the improvement is about 2 to 3dB. Sudarshan Nagesh, Satish Mulleti, Chandra Sekhar Seelamantula |
ICASSP | 3 |
| 2014 | An optimum shrinkage estimator based on minimum-probability-of-error criterion and application to signal denoisingabstractWe address the problem of designing an optimal pointwise shrinkage estimator in the transform domain, based on the minimum probability of error (MPE) criterion. We assume an additive model for the noise corrupting the clean signal. The proposed formulation is general in the sense that it can handle various noise distributions. We consider various noise distributions (Gaussian, Student's-t, and Laplacian) and compare the denoising performance of the estimator obtained with the mean-squared error (MSE)-based estimators. The MSE optimization is carried out using an unbiased estimator of the MSE, namely Stein's Unbiased Risk Estimate (SURE). Experimental results show that the MPE estimator outperforms the SURE estimator in terms of SNR of the denoised output, for low (0-10 dB) and medium values (10-20 dB) of the input SNR. Jishnu Sadasivan, Subhadip Mukherjee, Chandra Sekhar Seelamantula |
ICASSP | 3 |
| 2014 | Frequency domain linear prediction based on temporal analysisabstractFrequency-domain linear prediction (FDLP) is widely used in speech coding for modeling envelopes of transients signals, such as voiced and unvoiced stops, plosives, etc. FDLP fits an auto regressive model to the discrete cosine transform (DCT) coefficients of a sequence. The spectral prediction coefficients provide a parametric model of the temporal envelope. The prediction coefficients are obtained by solving the set of Yule-Walker equations expressing the relationship between lagged spectral autocorrelation values. A limitation of the direct approach of computing the spectral autocorrelation values is that the sequence has to be padded with a large number of zeros for the autocorrelation estimates to be reasonably accurate. This comes at the cost of increased computational complexity. We present an efficient and accurate method for computing the spectral autocorrelation samples. We show that the spectral autocorrelation can be computed as cosine-weighted temporal centroids, where the weighting function is dependent on time-index of the samples. Ravi R. Shenoy, Chandra Sekhar Seelamantula |
ICASSP | 2 |
| 2014 | Ultrasound image reconstruction using the finite-rate-of-innovation principleabstractRecently, a method of finding the spectral samples of non-periodic-finite-rate-of-innovation (NP-FRI) signals using a sum-of-sincs (SoS) sampling kernel was proposed in the literature. In the SoS approach, the kernel is repeated at a rate dependent on the delays of the FRI signal. The number of repetitions depends on both the duration and the delays of pulses constituting the FRI signal. In this paper, we show that the kernel repetition can be avoided and perfect reconstruction can be obtained by working with the SoS kernel directly provided that certain sampling criteria are satisfied. We place a lower bound on the sampling rate to ensure that exact signal reconstruction is achieved using filtered samples. To suppress the effect of noise, we use Cadzow denoising technique. Reconstruction is achieved using the annihilating filter method. We report results on data simulated using Field II software as well as real cardiac ultrasound data. The experimental results show that, with nearly 10 times less data than that required by the standard technique, the proposed method gives a comparable quality of reconstruction. The reconstruction accuracy can be controlled by choosing the model order of the NP-FRI signal appropriately. Satish Mulleti, Sudarshan Nagesh, Rajesh Langoju, Abhijit Patil, Chandra Sekhar Seelamantula |
ICIP | 5 |
| 2014 | Exact reconstruction in Quantitative Phase MicroscopyabstractWe address the problem of phase reconstruction in Quantitative Phase Microscopy (QPM). We develop sufficient conditions on the interfering reference wave for exact phase reconstruction in the absence of noise and propose a noniterative phase reconstruction technique. We show that the zero-order artifact, a commonly encountered problem in QPM, can be completely suppressed by the proposed reconstruction technique. The theoretical guarantees pertain to continuous-domain object functions with the additional assumption of being bandlimited. We present results on synthesized phase objects that are effectively bandlimited to show that the proposed method is capable of recovering the phase acurately. In cases of images where the phase exhibits large excursions, unwrapping becomes necessary, for which we employ Goldstein's phase unwrapping algorithm. Chandra Sekhar Seelamantula, Basty Ajay Shenoy, Severine Coquoz, Theo Lasser |
ICIP | 1 |
| 2014 | Controlled blurring for improving image reconstruction quality in flutter-shutter acquisitionabstractThe blurred images obtained using conventional cameras usually lack high-frequency details due to the impulse response associated with long exposure. This leads to imperfect reconstruction of the underlying scene. It has been shown that flutter-shutter (FS) cameras retain the high-frequency details by physically controlling the characteristics of exposure. This technique has been explored for conditions of fixed camera with objects moving in 1-D. In this paper, we propose a technique that is applicable to 2-D motion. First, we propose the use of fast iterative shrinkage thresholding algorithm (FISTA) for reconstruction of the underlying image convolved with slightly complex blurs. Second, we propose a new acquisition technique, called controlled blurring, which preserves high-frequency image content better than a flutter-shutter camera. Suraj Srinivas, Aniruddha Adiga, Chandra Sekhar Seelamantula |
ICIP | 3 |
| 2014 | Auditory-motivated Gammatone wavelet transform
Arun Venkitaraman, Aniruddha Adiga, Chandra Sekhar Seelamantula |
Signal Process. | 3 |
| 2014 | Fractional Hilbert transform extensions and associated analytic signal construction
Arun Venkitaraman, Chandra Sekhar Seelamantula |
Signal Process. | 2 |
| 2014 | Spatially Adaptive Kernel Regression Using Risk EstimationabstractAn important question in kernel regression is one of estimating the order and bandwidth parameters from available noisy data. We propose to solve the problem within a risk estimation framework. Considering an independent and identically distributed (i.i.d.) Gaussian observations model, we use Stein's unbiased risk estimator (SURE) to estimate a weighted mean-square error (MSE) risk, and optimize it with respect to the order and bandwidth parameters. The two parameters are thus spatially adapted in such a manner that noise smoothing and fine structure preservation are simultaneously achieved. On the application side, we consider the problem of image restoration from uniform/non-uniform data, and show that the SURE approach to spatially adaptive kernel regression results in better quality estimation compared with its spatially non-adaptive counterparts. The denoising results obtained are comparable to those obtained using other state-of-the-art techniques, and in some scenarios, superior. Sunder Ram Krishnan, Chandra Sekhar Seelamantula, Purvasha Chakravarti |
IEEE Signal Process. Lett. | 2 |
| 2014 | A Contraction Mapping Approach for Robust Estimation of Lagged AutocorrelationabstractWe consider the zero-crossing rate (ZCR) of a Gaussian process and establish a property relating the lagged ZCR (LZCR) to the corresponding normalized autocorrelation function. This is a generalization of Kedem's result for the lag-one case. For the specific case of a sinusoid in white Gaussian noise, we use the higher-order property between lagged ZCR and higher-lag autocorrelation to develop an iterative higher-order autoregressive filtering scheme, which stabilizes the ZCR and consequently provide robust estimates of the lagged autocorrelation. Simulation results show that the autocorrelation estimates converge in about 20 to 40 iterations even for low signal-to-noise ratio. Chandra Sekhar Seelamantula, Ravi R. Shenoy |
IEEE Signal Process. Lett. | 1 |
| 2014 | Binaural Signal Processing Motivated Generalized Analytic Signal Construction and AM-FM DemodulationabstractBinaural hearing studies show that the auditory system uses the phase-difference information in the auditory stimuli for localization of a sound source. Motivated by this finding, we present a method for demodulation of amplitude-modulated-frequency-modulated (AM-FM) signals using a signal and its arbitrary phase-shifted version. The demodulation is achieved using two allpass filters, whose impulse responses are related through the fractional Hilbert transform (FrHT). The allpass filters are obtained by cosine-modulation of a zero-phase flat-top prototype halfband lowpass filter. The outputs of the filters are combined to construct an analytic signal (AS) from which the AM and FM are estimated. We show that, under certain assumptions on the signal and the filter structures, the AM and FM can be obtained exactly. The AM-FM calculations are based on the quasi-eigenfunction approximation. We then extend the concept to the demodulation of multicomponent signals using uniform and non-uniform cosine-modulated filterbank (FB) structures consisting of flat bandpass filters, including the uniform cosine-modulated, equivalent rectangular bandwidth (ERB), and constant-Q filterbanks. We validate the theoretical calculations by considering application on synthesized AM-FM signals and compare the performance in presence of noise with three other multiband demodulation techniques, namely, the Teager-energy-based approach, the Gabor's AS approach, and the linear transduction filter approach. We also show demodulation results for real signals. Arun Venkitaraman, Chandra Sekhar Seelamantula |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | Riesz-transform-based demodulation of narrowband spectrograms of voiced speechabstractNarrowband spectrograms of voiced speech can be modeled as an outcome of two-dimensional (2-D) modulation process. In this paper, we develop a demodulation algorithm to estimate the 2-D amplitude modulation (AM) and carrier of a given spectrogram patch. The demodulation algorithm is based on the Riesz transform, which is a unitary, shift-invariant operator and is obtained as a 2-D extension of the well known 1-D Hilbert transform operator. Existing methods for spectrogram demodulation rely on extension of sinusoidal demodulation method from the communications literature and require precise estimate of the 2-D carrier. On the other hand, the proposed method based on Riesz transform does not require a carrier estimate. The proposed method and the sinusoidal demodulation scheme are tested on real speech data. Experimental results show that the demodulated AM and carrier from Riesz demodulation represent the spectrogram patch more accurately compared with those obtained using the sinusoidal demodulation. The signal-to-reconstruction error ratio was found to be about 2 to 6 dB higher in case of the proposed demodulation approach. Haricharan Aragonda, Chandra Sekhar Seelamantula |
ICASSP | 2 |
| 2013 | Bilateral edge detectorsabstractWe propose to employ bilateral filters to solve the problem of edge detection. The proposed methodology presents an efficient and noise robust method for detecting edges. Classical bilateral filters smooth images without distorting edges. In this paper, we modify the bilateral filter to perform edge detection, which is the opposite of bilateral smoothing. The Gaussian domain kernel of the bilateral filter is replaced with an edge detection mask, and Gaussian range kernel is replaced with an inverted Gaussian kernel. The modified range kernel serves to emphasize dissimilar regions. The resulting approach effectively adapts the detection mask according as the pixel intensity differences. The results of the proposed algorithm are compared with those of standard edge detection masks. Comparisons of the bilateral edge detector with Canny edge detection algorithm, both after non-maximal suppression, are also provided. The results of our technique are observed to be better and noise-robust than those offered by methods employing masks alone, and are also comparable to the results from Canny edge detector, outperforming it in certain cases. Abin Jose, Chandra Sekhar Seelamantula |
ICASSP | 2 |
| 2013 | Ridge detection using Savitzky-Golay filtering and steerable second-order Gaussian derivativesabstractWe propose a method for ridge detection at different widths using second-order Gaussian derivative masks. The width of the ridge extracted varies depending on the mask size and its parameter, σ. In the proposed method, the ridge orientations are estimated as an initial step by finding the zerocrossings of the first derivative of the second directional derivative. In order to compute the orientations from discrete samples of the image, we make use of the recently popularized Savitzky-Golay (S-G) filter. Once the directions are estimated, ridge detection is accomplished by steering a second-order Gaussian kernel, which closely approximates the ideal ridge template, in the computed directions. The method is computationally effective on two accounts: (1) The ridge orientations are determined efficiently using S-G filtering; and (2) Once the orientations are estimated, the steerability property is used to detect ridges. The output of the ridge detector is then improved using non-maximal suppression and hysteresis thresholding. The results obtained are compared with an efficient benchmark method for ridge extraction. Abin Jose, Sunder Ram Krishnan, Chandra Sekhar Seelamantula |
ICIP | 3 |
| 2013 | Sure-optimal two-dimensional Savitzky-Golay filters for image denoisingabstractSavitzky-Golay (SG) filters are linear, shift-invariant lowpass filters employed for data smoothing. In their pathbreaking paper published in Analytical Chemistry, Savitzky and Golay mathematically established that polynomial regression of data over local intervals and evaluation of their values at the center of the approximation window is equivalent to convolution with a finite impulse response filter. In this paper, we expound SURE (Stein's unbiased risk estimate) based adaptive SG filters for image denoising. Our goal is to optimally choose SG filter parameters, namely, order and window length, the optimality defined in terms of the mean squared error (MSE). In practical scenarios, only a single realization of the noisy image is available and the ground truth is inaccessible. Hence, we propose SURE, which is an unbiased estimate of MSE, to solve the parameter selection problem. It is observed that bandwidth of the minimum MSE (MMSE)-optimum SG filter is small at relatively slowly varying portions of the underlying image, and vice versa at abrupt transitions, thereby enabling us to trade off bias and variance to obtain near-optimal performance. The denoising results obtained exhibit considerable peak signal-to-noise-ratio (PSNR) improvement. At low SNRs, the filter performance is further enhanced by using a regularized cost function. Sreeram V. Menon, Chandra Sekhar Seelamantula |
ICIP | 2 |
| 2013 | A shape-template based two-stage corpus callosum segmentation technique for sagittal plane T1-weighted brain magnetic resonance imagesabstractWe propose a semi-automatic technique to segment corpus callosum (CC) using a two-stage snake formulation: A restricted affine transform (RAT) constrained snake followed by an unconstrained snake in an iterative fashion. A statistical model is developed to capture the shape variations of CC from a training set, which restrict the unconstrained snake to lie in the shape-space of CC. The geometry of the constrained snake is optimized using a local contrast-based energy over RAT space (which allows for five degrees of freedom). On the other hand, the unconstrained snake is optimized using a unified energy (region, gradient, and curvature energy) formulation. Joint optimization resulted in increased robustness to initialization as well as fast and accurate segmentation. The technique was validated on 243 images taken from the OASIS database and performance was quantified using Jaccard's distance, sensitivity, and specificity as the metrics. Jayanth Krishna Mogali, Naren Nallapareddy, Chandra Sekhar Seelamantula, Michael Unser |
ICIP | 3 |
| 2013 | Phase retrieval for a class of 2-D signals characterized by first-order difference equationsabstractWe address the problem of signal reconstruction from the Fourier transform magnitude of a certain class of two-dimensional (2-D) signals that are characterized by first-order difference equations. We show that when such a signal has a Z-transform that includes the unit sphere in the region of convergence, it can be reconstructed uniquely from the Fourier magnitude. We employ the annihilating filter approach to find the parameters of the rational transfer function. Basty Ajay Shenoy, Subhadip Mukherjee, Chandra Sekhar Seelamantula |
ICIP | 3 |
| 2013 | A Savitzky-Golay Filtering Perspective of Dynamic Feature ComputationabstractWe address the classical problem of delta feature computation, and interpret the operation involved in terms of Savitzky-Golay (SG) filtering. Features such as the mel-frequency cepstral coefficients (MFCCs), obtained based on short-time spectra of the speech signal, are commonly used in speech recognition tasks. In order to incorporate the dynamics of speech, auxiliary delta and delta-delta features, which are computed as temporal derivatives of the original features, are used. Typically, the delta features are computed in a smooth fashion using local least-squares (LS) polynomial fitting on each feature vector component trajectory. In the light of the original work of Savitzky and Golay, and a recent article by Schafer in IEEE Signal Processing Magazine, we interpret the dynamic feature vector computation for arbitrary derivative orders as SG filtering with a fixed impulse response. This filtering equivalence brings in significantly lower latency with no loss in accuracy, as validated by results on a TIMIT phoneme recognition task. The SG filters involved in dynamic parameter computation can be viewed as modulation filters, proposed by Hermansky. Sunder Ram Krishnan, Mathew Magimai-Doss, Chandra Sekhar Seelamantula |
IEEE Signal Process. Lett. | 3 |
| 2013 | On Computing Amplitude, Phase, and Frequency Modulations Using a Vector Interpretation of the Analytic SignalabstractThe amplitude-modulation (AM) and phase-modulation (PM) of an amplitude-modulated frequency-modulated (AM-FM) signal are defined as the modulus and phase angle, respectively, of the analytic signal (AS). The FM is defined as the derivative of the PM. However, this standard definition results in a PM with jump discontinuities in cases when the AM index exceeds unity, resulting in an FM that contains impulses. We propose a new approach to define smooth AM, PM, and FM for the AS, where the PM is computed as the solution to an optimization problem based on a vector interpretation of the AS. Our approach is directly linked to the fractional Hilbert transform (FrHT) and leads to an eigenvalue problem. The resulting PM and AM are shown to be smooth, and in particular, the AM turns out to be bipolar. We show an equivalence of the eigenvalue formulation to the square of the AS, and arrive at a simple method to compute the smooth PM. Some examples on synthesized and real signals are provided to validate the theoretical calculations. Arun Venkitaraman, Chandra Sekhar Seelamantula |
IEEE Signal Process. Lett. | 2 |
| 2013 | Temporal Envelope Fit of Transient Audio SignalsabstractWe address the problem of temporal envelope modeling for transient audio signals. We propose the Gamma distribution function (GDF) as a suitable candidate for modeling the envelope keeping in view some of its interesting properties such as asymmetry, causality, near-optimal time-bandwidth product, controllability of rise and decay, etc. The problem of finding the parameters of the GDF becomes a nonlinear regression problem. We overcome the hurdle by using a logarithmic envelope fit, which reduces the problem to one of linear regression. The logarithmic transformation also has the feature of dynamic range compression. Since temporal envelopes of audio signals are not uniformly distributed, in order to compute the amplitude, we investigate the importance of various loss functions for regression. Based on synthesized data experiments, wherein we have a ground truth, and real-world signals, we observe that the least-squares technique gives reasonably accurate amplitude estimates compared with other loss functions. Arun Venkitaraman, Chandra Sekhar Seelamantula |
IEEE Signal Process. Lett. | 2 |
| 2012 | Sure-fast bilateral filtersabstractEdge-preserving smoothing is widely used in image processing and bilateral filtering is one way to achieve it. Bilateral filter is a nonlinear combination of domain and range filters. Implementing the classical bilateral filter is computationally intensive, owing to the nonlinearity of the range filter. In the standard form, the domain and range filters are Gaussian functions and the performance depends on the choice of the filter parameters. Recently, a constant time implementation of the bilateral filter has been proposed based on raised-cosine approximation to the Gaussian to facilitate fast implementation of the bilateral filter. We address the problem of determining the optimal parameters for raised-cosine-based constant time implementation of the bilateral filter. To determine the optimal parameters, we propose the use of Stein's unbiased risk estimator (SURE). The fast bilateral filter accelerates the search for optimal parameters by faster optimization of the SURE cost. Experimental results show that the SURE-optimal raised-cosine-based bilateral filter has nearly the same performance as the SURE-optimal standard Gaussian bilateral filter and the Oracle mean squared error (MSE)-based optimal bilateral filter. Harini Kishan, Chandra Sekhar Seelamantula |
ICASSP | 2 |
| 2012 | An iterative algorithm for phase retrieval with sparsity constraints: application to frequency domain optical coherence tomographyabstractWe address the problem of phase retrieval, which is frequently encountered in optical imaging. The measured quantity is the magnitude of the Fourier spectrum of a function (in optics, the function is also referred to as an object). The goal is to recover the object based on the magnitude measurements. In doing so, the standard assumptions are that the object is compactly supported and positive. In this paper, we consider objects that admit a sparse representation in some orthonormal basis. We develop a variant of the Fienup algorithm to incorporate the condition of sparsity and to successively estimate and refine the phase starting from the magnitude measurements. We show that the proposed iterative algorithm possesses Cauchy convergence properties. As far as the modality is concerned, we work with measurements obtained using a frequency-domain optical-coherence tomography experimental setup. The experimental results on real measured data show that the proposed technique exhibits good reconstruction performance even with fewer coefficients taken into account for reconstruction. It also suppresses the autocorrelation artifacts to a significant extent since it estimates the phase accurately. Subhadip Mukherjee, Chandra Sekhar Seelamantula |
ICASSP | 2 |
| 2012 | A unified approach for optimization of Snakuscules and OvusculesabstractAutomated image segmentation techniques are useful tools in biological image analysis and are an essential step in tracking applications. Typically, snakes or active contours are used for segmentation and they evolve under the influence of certain internal and external forces. Recently, a new class of shape-specific active contours have been introduced, which are known as Snakuscules and Ovuscules. These contours are based on a pair of concentric circles and ellipses as the shape templates, and the optimization is carried out by maximizing a contrast function between the outer and inner templates. In this paper, we present a unified approach to the formulation and optimization of Snakuscules and Ovuscules by considering a specific form of affine transformations acting on a pair of concentric circles. We show how the parameters of the affine transformation may be optimized for, to generate either Snakuscules or Ovuscules. Our approach allows for a unified formulation and relies only on generic regularization terms and not shape-specific regularization functions. We show how the calculations of the partial derivatives may be made efficient thanks to the Green's theorem. Results on synthesized as well as real data are presented. Adithya Kumar Pediredla, Chandra Sekhar Seelamantula |
ICASSP | 2 |
| 2012 | Bilateral smoothing of gradient vector field and application to image segmentationabstractMedical image segmentation finds application in computer-aided diagnosis, computer-guided surgery, measuring tissue volumes, locating tumors, and pathologies. One approach to segmentation is to use active contours or snakes. Active contours start from an initialization (often manually specified) and are guided by image-dependent forces to the object boundary. Snakes may also be guided by gradient vector fields associated with an image. The first main result in this direction is that of Xu and Prince, who proposed the notion of gradient vector flow (GVF), which is computed iteratively. We propose a new formalism to compute the vector flow based on the notion of bilateral filtering of the gradient field associated with the edge map - we refer to it as the bilateral vector flow (BVF). The range kernel definition that we employ is different from the one employed in the standard Gaussian bilateral filter. The advantage of the BVF formalism is that smooth gradient vector flow fields with enhanced edge information can be computed noniteratively. The quality of image segmentation turned out to be on par with that obtained using the GVF and in some cases better than the GVF. Ravindra S. Hegadi, Adithya Kumar Pediredla, Chandra Sekhar Seelamantula |
ICIP | 3 |
| 2012 | Optimal parameter selection for bilateral filters using Poisson Unbiased Risk EstimateabstractBilateral filters perform edge-preserving smoothing and are widely used for image denoising. The denoising performance is sensitive to the choice of the bilateral filter parameters. We propose an optimal parameter selection for bilateral filtering of images corrupted with Poisson noise. We employ the Poisson's Unbiased Risk Estimate (PURE), which is an unbiased estimate of the Mean Squared Error (MSE). It does not require a priori knowledge of the ground truth and is useful in practical scenarios where there is no access to the original image. Experimental results show that quality of denoising obtained with PURE-optimal bilateral filters is almost indistinguishable with that of the Oracle-MSE-optimal bilateral filters. Harini Kishan, Chandra Sekhar Seelamantula |
ICIP | 2 |
| 2012 | A Technique to Compute Smooth Amplitude, Phase, and Frequency Modulations From the Analytic SignalabstractGabor's analytic signal (AS) is a unique complex signal corresponding to a real signal, but in general, it admits infinitely-many combinations of amplitude and frequency modulations (AM and FM, respectively). The standard approach is to enforce a non-negativity constraint on the AM, but this results in discontinuities in the corresponding phase modulation (PM), and hence, an FM with discontinuities particularly when the underlying AM-FM signal is over-modulated. In this letter, we analyze the phase discontinuities and propose a technique to compute smooth AM and FM from the AS, by relaxing the non-negativity constraint on the AM. The proposed technique is effective at handling over-modulated signals. We present simulation results to support the theoretical calculations. Arun Venkitaraman, Chandra Sekhar Seelamantula |
IEEE Signal Process. Lett. | 2 |
| 2012 | A Mixture Model Approach for Formant Tracking and the Robustness of Student's-t DistributionabstractWe address the problem of robust formant tracking in continuous speech in the presence of additive noise. We propose a new approach based on mixture modeling of the formant contours. Our approach consists of two main steps: (i) Computation of a pyknogram based on multiband amplitude-modulation/frequency-modulation (AM/FM) decomposition of the input speech; and (ii) Statistical modeling of the pyknogram using mixture models. We experiment with both Gaussian mixture model (GMM) and Student's-t mixture model (tMM) and show that the latter is robust with respect to handling outliers in the pyknogram data, parameter selection, accuracy, and smoothness of the estimated formant contours. Experimental results on simulated data as well as noisy speech data show that the proposed tMM-based approach is also robust to additive noise. We present performance comparisons with a recently developed adaptive filterbank technique proposed in the literature and the classical Burg's spectral estimator technique, which show that the proposed technique is more robust to noise. Sundar Harshavardhan, Chandra Sekhar Seelamantula, Thippur V. Sreenivas |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Snakes With an Ellipse-Reproducing PropertyabstractWe present a new class of continuously defined parametric snakes using a special kind of exponential splines as basis functions. We have enforced our bases to have the shortest possible support subject to some design constraints to maximize efficiency. While the resulting snakes are versatile enough to provide a good approximation of any closed curve in the plane, their most important feature is the fact that they admit ellipses within their span. Thus, they can perfectly generate circular and elliptical shapes. These features are appropriate to delineate cross sections of cylindrical-like conduits and to outline bloblike objects. We address the implementation details and illustrate the capabilities of our snake with synthetic and real data. Ricard Delgado-Gonzalo, Philippe Thévenaz, Chandra Sekhar Seelamantula, Michael Unser |
IEEE Trans. Image Process. | 3 |
| 2011 | Quadrature approximation properties of the spiral-phase quadrature transformabstractThe notion of the 1-D analytic signal is well understood and has found many applications. At the heart of the analytic signal concept is the Hilbert transform. The problem in extending the concept of analytic signal to higher dimensions is that there is no unique multidimensional definition of the Hilbert transform. Also, the notion of analyticity is not so well under stood in higher dimensions. Of the several 2-D extensions of the Hilbert transform, the spiral-phase quadrature transform or the Riesz transform seems to be the natural extension and has attracted a lot of attention mainly due to its isotropic properties. From the Riesz transform, Larkin et al. constructed a vortex operator, which approximates the quadratures based on asymptotic stationary-phase analysis. In this paper, we show an alternative proof for the quadrature approximation property by invoking the quasi-eigenfunction property of linear, shift-invariant systems. We show that the vortex operator comes up as a natural consequence of applying this property. We also characterize the quadrature approximation error in terms of its energy as well as the peak spatial-domain error. Such results are available for 1-D signals, but their counter part for 2-D signals have not been provided. We also provide simulation results to supplement the analytical calculations. Haricharan Aragonda, Chandra Sekhar Seelamantula |
ICASSP | 2 |
| 2011 | Spectral-envelope and group-delay models for transient signals - Applications to castanets and stop consonantsabstractWe present a novel approach to represent transients using spectral-domain amplitude-modulated/frequency -modulated (AM-FM) functions. The model is applied to the real and imaginary parts of the Fourier transform (FT) of the transient. The suitability of the model lies in the observation that since transients are well-localized in time, the real and imaginary parts of the Fourier spectrum have a modulation structure. The spectral AM is the envelope and the spectral FM is the group delay function. The group delay is estimated using spectral zero-crossings and the spectral envelope is estimated using a coherent demodulator. We show that the proposed technique is robust to additive noise. We present applications of the proposed technique to castanets and stop-consonants in speech. Ravi R. Shenoy, Chandra Sekhar Seelamantula |
ICASSP | 2 |
| 2011 | A Risk-Estimation-Based Comparison of Mean Square Error and Itakura-Saito Distortion Measures for Speech Enhancement
Nagarjuna Reddy Muraka, Chandra Sekhar Seelamantula |
INTERSPEECH | 2 |
| 2010 | A multimodal density function estimation approach to formant tracking
Sundar Harshavardhan, Chandra Sekhar Seelamantula, Thippur V. Sreenivas |
INTERSPEECH | 2 |
| 2009 | Blocking artifacts in speech/audio: Dynamic auditory model-based characterization and optimal time-frequency smoothing
Chandra Sekhar Seelamantula, Thippur V. Sreenivas |
Signal Process. | 1 |
| 2008 | Performance analysis of the cepstral technique for frequency-domain optical-coherence tomographyabstractRecently, we proposed a noniterative cepstral technique for exact signal recovery in frequency-domain optical-coherence tomography. In this paper, we address the influence of measurement noise on the performance of the method. We derive analytical expressions for the bias and variance of the tomogram under a small noise approximation, and show that our technique yields unbiased and consistent estimators, which have a variance that is proportional to that of the noise and inversely proportional to the data size. We present simulation results to confirm the theoretical derivations. We also derive approximate Cramer-Rao bounds (CRBs) on the achievable accuracy of reconstruction. Chandra Sekhar Seelamantula, Michael Unser |
ICASSP | 1 |
| 2008 | A Generalized Sampling Method for Finite-Rate-of-Innovation-Signal ReconstructionabstractThe problem of sampling signals that are not admissible within the classical Shannon framework has received much attention in the recent past. Typically, these signals have a parametric representation with a finite number of degrees of freedom per time unit. It was shown that, by choosing suitable sampling kernels, the parameters can be computed by employing high-resolution spectral estimation techniques. In this letter, we propose a simple acquisition and reconstruction method within the framework of multichannel sampling. In the proposed approach, an infinite stream of nonuniformly-spaced Dirac impulses can be sampled and accurately reconstructed provided that there is at most one Dirac impulse per sampling period. The reconstruction algorithm has a low computational complexity, and the parameters are computed on the fly. The processing delay is minimal just the sampling period. We propose sampling circuits using inexpensive passive devices such as resistors and capacitors. We also show how the approach can be extended to sample piecewise-constant signals with a minimal change in the system configuration. We provide some simulation results to confirm the theoretical findings. Chandra Sekhar Seelamantula, Michael Unser |
IEEE Signal Process. Lett. | 1 |
| 2007 | A New Technique for High-Resolution Frequency Domain Optical Coherence TomographyabstractFrequency domain optical coherence tomography (FDOCT) is a new technique that is well-suited for fast imaging of biological specimens, as well as non-biological objects. The measurements are in the frequency domain, and the objective is to retrieve an artifact-free spatial domain description of the specimen. In this paper, we develop a new technique for model-based retrieval of spatial domain data from the frequency domain data. We use a piecewise-constant model for the refractive index profile that is suitable for multi-layered specimens. We show that the estimation of the layered structure parameters can be mapped into a harmonic retrieval problem, which enables us to use high-resolution spectrum estimation techniques. The new technique that we propose is efficient and requires few measurements. We also analyze the effect of additive measurement noise on the algorithm performance. The experimental results show that the technique gives highly accurate parameter estimates. For example, at 25 dB signal-to-noise ratio, the mean square error in the position estimate is about 0.01 % of the actual value. Chandra Sekhar Seelamantula, Himanshu Nazkani, Thierry Blu, Michael Unser |
ICASSP (1) | 1 |
| 2007 | Robust and high-resolution voiced/unvoiced classification in noisy speech using a signal smoothness criterionabstractWe propose a novel technique for robust voiced/unvoiced segment detection in noisy speech, based on local polynomial regression. The local polynomial model is well-suited for voiced segments in speech. The unvoiced segments are noise-like and do not exhibit any smooth structure. This property of smoothness is used for devising a new metric called the variance ratio metric, which, after thresholding, indicates the voiced/unvoiced boundaries with 75% accuracy for 0dB global signal-to-noise ratio (SNR). A novelty of our algorithm is that it processes the signal continuously, sample-by-sample rather than frame-by-frame. Simulation results on TIMIT speech database (downsampled to 8kHz) for various SNRs are presented to illustrate the performance of the new algorithm. Results indicate that the algorithm is robust even in high noise levels. A. Sreenivasa Murthy, Chandra Sekhar Seelamantula, Thippur V. Sreenivas |
INTERSPEECH | 2 |
| 2006 | Signal-to-noise ratio estimation using higher-order moments
Chandra Sekhar Seelamantula, Thippur V. Sreenivas |
Signal Process. | 1 |
| 2004 | Novel approach to AM-FM decomposition with applications to speech and music analysisabstractWe present a new zero-crossing based algorithm for decomposing a bandpass signal into the amplitude modulation (AM) and frequency modulation (FM) components. In this sequential algorithm, the FM component is first estimated using zero-crossing instant information in a k-nearest neighbour (k-NN) framework. The AM component is estimated by coherent demodulation using a time-varying lowpass filter that uses the estimated instantaneous frequency. Simulation results show that the proposed algorithm gives more accurate envelope and frequency estimates compared to the discrete-energy separation algorithm (DESA) which uses the Teager energy operator. Using the proposed approach on bandpass filtered speech and music, we can extract the fine-structured modulations that occur on a micro-time scale, within an analysis frame. Chandra Sekhar Seelamantula, Thippur V. Sreenivas |
ICASSP (2) | 1 |
| 2004 | Effect of interpolation on PWVD computation and instantaneous frequency estimation
Chandra Sekhar Seelamantula, Thippur V. Sreenivas |
Signal Process. | 1 |
| 2003 | Instantaneous frequency estimation using level-crossing informationabstractWe discuss the problem of instantaneous frequency (IF) estimation of phase signals using their level-crossing (LC) instant information. We cast the problem to that of interpolating the instantaneous phase (IP), and hence finding the IF, from samples obtained at the level-crossing instants of the phase signal. These are inherently irregularly spaced and the problem essentially reduces to reconstructing a signal from the samples taken at irregularly sampled points for which we propose a 'line plus sum of sines' model. In the presence of noise, the temporal structure of the level-crossings can get distorted. To reduce the effects of noise, we use a short-time Fourier transform (STFT) based enhancement scheme. The performance of the proposed method is studied through Monte-Carlo simulations for a phase signal with composite IF for various SNRs. Different level-crossing based estimates are combined to obtain a new IF estimate. Simulation studies show that the estimates obtained using zero-crossing (ZC) and other very low level values perform better than those obtained with higher level values. Chandra Sekhar Seelamantula, Thippur V. Sreenivas |
ICASSP (6) | 1 |
| 2003 | Adaptive spectrogram vs. adaptive pseudo-Wigner-Ville distribution for instantaneous frequency estimation
Chandra Sekhar Seelamantula, Thippur V. Sreenivas |
Signal Process. | 1 |