Ritwik Giri

dblp:49/8285 · DBLP profile ↗
← Back
26ranked-venue papers
9as first author
6since 2021 · last 2024
0000-0002-8599-3229ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2024 Real-Time Stereo Speech Enhancement with Spatial-Cue Preservation Based on Dual-Path Structure
abstract
We introduce a real-time, multichannel speech enhancement algorithm which maintains the spatial cues of stereo recordings including two speech sources. Recognizing that each source has unique spatial information, our method utilizes a dual-path structure, ensuring the spatial cues remain unaffected during enhancement by applying source-specific common-band gain. This method also seamlessly integrates pretrained monaural speech enhancement, eliminating the need for retraining on stereo inputs. Source separation from stereo mixtures is achieved via spatial beamforming, with the steering vector for each source being adaptively updated using post-enhancement output signal. This ensures accurate tracking of the spatial information. The final stereo output is derived by merging the spatial images of the enhanced sources, with its efficacy not heavily reliant on the separation performance of the beamforming. The algorithm runs in real-time on 10-ms frames with a 40 ms of look-ahead. Evaluations reveal its effectiveness in enhancing speech and preserving spatial cues in both fully and sparsely overlapped mixtures.
Masahito Togami, Jean-Marc Valin, Karim Helwani, Ritwik Giri, Umut Isik, Michael M. Goodwin
ICASSP4
2023 A Framework for Unified Real-Time Personalized and Non-Personalized Speech Enhancement
abstract
In this study, we present an approach to train a single speech enhancement network that can perform both personalized and non-personalized speech enhancement. This is achieved by incorporating a frame-wise conditioning input that specifies the type of enhancement output. To improve the quality of the enhanced output and mitigate oversuppression, we experiment with re-weighting frames by the presence or absence of speech activity and applying augmentations to speaker embeddings. By training under a multi-task learning setting, we empirically show that the proposed unified model obtains promising results on both personalized and non-personalized speech enhancement benchmarks and reaches similar performance to models that are trained specialized for either task. The strong performance of the proposed method demonstrates that the unified model is a more economical alternative compared to keeping separate task-specific models during inference.
Zhepei Wang, Ritwik Giri, Devansh Shah, Jean-Marc Valin, Michael M. Goodwin, Paris Smaragdis
ICASSP2
2022 Improved Singing Voice Separation with Chromagram-Based Pitch-Aware Remixing
abstract
Singing voice separation aims to separate music into vocals and accompaniment components. One of the major constraints for the task is the limited amount of training data with separated vocals. Data augmentation techniques such as random source mixing have been shown to make better use of existing data and mildly improve model performance. We propose a novel data augmentation technique, chromagram-based pitch-aware remixing, where music segments with high pitch alignment are mixed. By performing controlled experiments in both supervised and semi-supervised settings, we demonstrate that training models with pitch-aware remixing significantly improves the test signal-to-distortion ratio (SDR).
Siyuan Yuan, Zhepei Wang, Umut Isik, Ritwik Giri, Jean-Marc Valin, Michael M. Goodwin, Arvindh Krishnaswamy
ICASSP4
2021 Enhancing into the Codec: Noise Robust Speech Coding with Vector-Quantized Autoencoders
abstract
Audio codecs based on discretized neural autoencoders have recently been developed and shown to provide significantly higher compression levels for comparable quality speech out-put. However, these models are tightly coupled with speech content, and produce unintended outputs in noisy conditions. Based on VQ-VAE autoencoders with WaveRNN decoders, we develop compressor-enhancer encoders and accompanying decoders, and show that they operate well in noisy conditions. We also observe that a compressor-enhancer model performs better on clean speech inputs than a compressor model trained only on clean speech.
Jonah Casebeer, Vinjai Vale, Umut Isik, Jean-Marc Valin, Ritwik Giri, Arvindh Krishnaswamy
ICASSP5
2021 Semi-Supervised Singing Voice Separation With Noisy Self-Training
abstract
Recent progress in singing voice separation has primarily focused on supervised deep learning methods. However, the scarcity of ground-truth data with clean musical sources has been a problem for long. Given a limited set of labeled data, we present a method to leverage a large volume of unlabeled data to improve the model’s performance. Following the noisy self-training framework, we first train a teacher network on the small labeled dataset and infer pseudo-labels from the large corpus of unlabeled mixtures. Then, a larger student network is trained on combined ground-truth and self-labeled datasets. Empirical results show that the proposed self-training scheme, along with data augmentation methods, effectively leverage the large unlabeled corpus and obtain superior performance compared to supervised methods.
Zhepei Wang, Ritwik Giri, Umut Isik, Jean-Marc Valin, Arvindh Krishnaswamy
ICASSP2
2021 Personalized PercepNet: Real-Time, Low-Complexity Target Voice Separation and Enhancement
abstract
The presence of multiple talkers in the surrounding environment poses a difficult challenge for real-time speech communication systems considering the constraints on network size and complexity. In this paper, we present Personalized PercepNet, a real-time speech enhancement model that separates a target speaker from a noisy multi-talker mixture without compromising on complexity of the recently proposed PercepNet. To enable speaker-dependent speech enhancement, we first show how we can train a perceptually motivated speaker embedder network to produce a representative embedding vector for the given speaker. Personalized PercepNet uses the target speaker embedding as additional information to pick out and enhance only the target speaker while suppressing all other competing sounds. Our experiments show that the proposed model significantly outperforms PercepNet and other baselines, both in terms of objective speech enhancement metrics and human opinion scores.
Ritwik Giri, Shrikant Venkataramani, Jean-Marc Valin, Umut Isik, Arvindh Krishnaswamy
Interspeech1
2020 Channel-Attention Dense U-Net for Multichannel Speech Enhancement
abstract
Supervised deep learning has gained significant attention for speech enhancement recently. The state-of-the-art deep learning methods perform the task by learning a ratio/binary mask that is applied to the mixture in the time-frequency domain to produce the clean speech. Despite the great performance in the single-channel setting, these frameworks lag in performance in the multichannel setting as the majority of these methods a) fail to exploit the available spatial information fully, and b) still treat the deep architecture as a black box which may not be well-suited for multichannel audio processing. This paper addresses these drawbacks, a) by utilizing complex ratio masking instead of masking on the magnitude of the spectrogram, and more importantly, b) by introducing a channel-attention mechanism inside the deep architecture to mimic beamforming. We propose Channel-Attention Dense U-Net, in which we apply the channel-attention unit recursively on feature maps at every layer of the network, enabling the network to perform non-linear beamforming. We demonstrate the superior performance of the network against the state-of-the-art approaches on the CHiME-3 dataset.
Bahareh Tolooshams, Ritwik Giri, Andrew H. Song, Umut Isik, Arvindh Krishnaswamy
ICASSP2
2020 PoCoNet: Better Speech Enhancement with Frequency-Positional Embeddings, Semi-Supervised Conversational Data, and Biased Loss
abstract
Neural network applications generally benefit from larger-sized models, but for current speech enhancement models, larger scale networks often suffer from decreased robustness to the variety of real-world use cases beyond what is encountered in training data. We introduce several innovations that lead to better large neural networks for speech enhancement. The novel PoCoNet architecture is a convolutional neural network that, with the use of frequency-positional embeddings, is able to more efficiently build frequency-dependent features in the early layers. A semi-supervised method helps increase the amount of conversational training data by pre-enhancing noisy datasets, improving performance on real recordings. A new loss function biased towards preserving speech quality helps the optimization better match human perceptual opinions on speech quality. Ablation experiments and objective and human opinion metrics show the benefits of the proposed improvements.
Umut Isik, Ritwik Giri, Neerad Phansalkar, Jean-Marc Valin, Karim Helwani, Arvindh Krishnaswamy
INTERSPEECH2
2020 A Perceptually-Motivated Approach for Low-Complexity, Real-Time Enhancement of Fullband Speech
abstract
Over the past few years, speech enhancement methods based on deep learning have greatly surpassed traditional methods based on spectral subtraction and spectral estimation. Many of these new techniques operate directly in the the short-time Fourier transform (STFT) domain, resulting in a high computational complexity. In this work, we propose PercepNet, an efficient approach that relies on human perception of speech by focusing on the spectral envelope and on the periodicity of the speech. We demonstrate high-quality, real-time enhancement of fullband (48 kHz) speech with less than 5% of a CPU core.
Jean-Marc Valin, Umut Isik, Neerad Phansalkar, Ritwik Giri, Karim Helwani, Arvindh Krishnaswamy
INTERSPEECH4
2018 Improved Noise Characterization for Relative Impulse Response Estimation
abstract
Relative Impulse Responses (ReIRs) have several applications in speech enhancement, noise suppression and source localization for multi-channel speech processing in reverberant environments. Noise is usually assumed to be white Gaussian during the estimation of the ReIR between two microphones. We show that the noise in this system identification problem is instead dependent upon the microphone measurements and the ReIR itself. We then present modifications that incorporate this new noise model into three prevalent methods: Least Squares, Non-Stationary Frequency Domain and Sparse Bayesian Learning based approaches. We demonstrated improvements with an experimental study using real-world measurements in various noise environments.
Tharun Adithya Srikrishnan, Bhaskar D. Rao, Ritwik Giri, Tao Zhang 0024
ICASSP3
2018 Perceptually Guided Speech Enhancement Using Deep Neural Networks
abstract
Human listeners often have difficulties understanding speech in the presence of background noise in the real world. Recently, supervised learning based speech enhancement approaches have achieved substantial success, and show significant improvements over the conventional approaches. However, existing supervised learning based approaches often try to minimize the mean squared error between the enhanced output and the pre-defined training target (e.g., the log power spectrum of clean speech), even though the purpose of such speech enhancement is to improve speech understanding in noise. In this paper, we propose a new deep neural networks based enhancement approach by incorporating a speech perception model into the loss function. Specifically, we use the short-time objective intelligibility metric in the loss in addition to the mean squared error. Optimizing the proposed perceptually guided loss is expected to improve speech intelligibility further. Systematic evaluations show that our proposed approach is able to improve speech intelligibility in a wide range of signal-to-noise ratios and noise types while maintaining speech quality.
Yan Zhao 0010, Buye Xu, Ritwik Giri, Tao Zhang 0024
ICASSP3
2018 A unified framework for sparse non-negative least squares using multiplicative updates and the non-negative matrix factorization problem
Igor Fedorov, Alican Nalci, Ritwik Giri, Bhaskar D. Rao, Truong Q. Nguyen, Harinath Garudadri
Signal Process.3
2017 Multivariate Scale mixtures for joint sparse regularization in multi-task learning
abstract
In this paper we address the problem of learning shared sparse representation across several tasks. Assuming that the tasks share a common set of relevant features across all tasks is highly restrictive. This acts as a motivation to look for a generalized model which will be able to learn any correlation structure present between the tasks. We propose a generalized scale mixture distribution, the Multivariate Power Exponential Scale Mixture (M-PESM), as a joint sparsity promoting prior and derive a unified framework which consists of many of the popular Multitask Learning algorithms. Our proposed unified model also has the ability to learn any present correlation structure between tasks which leads to a more robust framework.
Ritwik Giri, Bhaskar D. Rao
ICASSP1
2017 Bayesian Blind Deconvolution with application to acoustic Feedback Path modeling
abstract
Acoustic Feedback Path in a digital hearing aid is not only affected by the user's head and ear, but also by different acoustic environments. But some of these effects are common for a specific style of hearing aid and individual ear, i.e., this part will be invariant to the different acoustic environments and can be interpreted as the effects associated with that specific hearing aid and ear characteristics. In this article we propose a novel Bayesian Blind Deconvolution approach with exponentially decaying kernel and show its application in extracting the invariant part of the feedback path measurements of a digital hearing aid. Efficacy of our proposed approach in extracting the invariant part has been measured by using the extracted invariant part to model unseen test Feedback Path (FBP) measured from the same hearing aid but in a different acoustic environment, over existing methods.
Ritwik Giri, Tao Zhang 0024
ICASSP1
2017 Reweighted Algorithms for Independent Vector Analysis
abstract
In this letter, we consider the problem of joint blind source separation of multiple datasets simultaneously using an Independent vector analysis (IVA) framework. In particular we propose a new paradigm of reweighted algorithms for IVA by employing a source prior from a multivariate generalized scale mixture distribution family. In addition, our proposed reweighted algorithms can also exploit second-order statistical information across datasets by learning intrasource correlation matrix of each source component vector (SCV) along with higher order statistics. Experimental results are provided to show the efficacy of our proposed algorithms in achieving reliable source separation for both the cases, i.e., when there is correlation present within an SCV, and also when the sources are uncorrelated, i.e., no second order dependencies across datasets.
Ritwik Giri, Bhaskar D. Rao, Harinath Garudadri
IEEE Signal Process. Lett.1
2016 Dynamic relative impulse response estimation using structured sparse Bayesian learning
abstract
In this paper we present a novel Hierarchical Bayesian approach to estimate Relative Impulse Response (ReIR) using short, noisy and reverberant microphone recordings. The information contained in ReIRs between two microphones is useful for a wide range of multichannel speech processing applications such as speaker localization, speech enhancement, etc. It has been shown in several previous works that the Relative Transfer Function (RTF) corresponding to a given ReIR is dynamic and depends on the environment, microphone positions and target position. This acts as the main motivation of this work, as we develop a structured sparse Bayesian learning algorithm to estimate ReIR using very short recordings, which will be robust to changes in the environment. An extensive experimental study with real-world recordings has also been conducted to show the efficacy of our proposed approach over other competing approaches.
Ritwik Giri, Bhaskar D. Rao, Frédéric Mustière, Tao Zhang 0024
ICASSP1
2016 Robust Bayesian method for simultaneous block sparse signal recovery with applications to face recognition
abstract
In this paper, we present a novel Bayesian approach to recover simultaneously block sparse signals in the presence of outliers. The key advantage of our proposed method is the ability to handle non-stationary outliers, i.e. outliers which have time varying support. We validate our approach with empirical results showing the superiority of the proposed method over competing approaches in synthetic data experiments as well as the multiple measurement face recognition problem.
Igor Fedorov, Ritwik Giri, Bhaskar D. Rao, Truong Q. Nguyen
ICIP2
2015 Joint Clustering and Classification for Multiple Instance Learning
Karan Sikka, Ritwik Giri, Marian Stewart Bartlett
BMVC2
2015 Improving speech recognition in reverberation using a room-aware deep neural network and multi-task learning
abstract
In this paper, we propose two approaches to improve deep neural network (DNN) acoustic models for speech recognition in reverberant environments. Both methods utilize auxiliary information in training the DNN but differ in the type of information and the manner in which it is used. The first method uses parallel training data for multi-task learning, in which the network is trained to perform both a primary senone classification task and a secondary feature enhancement task using a shared representation. The second method uses a parameterization of the reverberant environment extracted from the observed signal to train a room-aware DNN. Experiments were performed on the single microphone task of the REVERB Challenge corpus. The proposed approach obtained a word error rate of 7.8% on the SimData test set, which is lower than all reported systems using the same training data and evaluation conditions, and 27.5% on the mismatched RealData test set, which is lower than all but two systems.
Ritwik Giri, Michael L. Seltzer, Jasha Droppo, Dong Yu 0001
ICASSP1
2014 Block sparse excitation based all-pole modeling of speech
abstract
In this paper, it is shown that an appropriate model for voiced speech is an all-pole filter excited by a block sparse excitation sequence. The modeling approach is generalized in a novel manner to deal with a wide spectrum of speech signal; voiced speech, unvoiced speech and mixed excitation speech. In this context, the input sequence to the all-pole model is modeled as a suitable weighted linear combination of a block sparse signal and white noise. We develop the corresponding estimation procedure to reconstruct the generalized input sequence and model parameters via sparse Bayesian learning methods employing the Expectation-Maximization based procedure. Rigorous experiments have been performed to show the efficacy of our proposed model for the speech modeling task. By imposing a block sparse structure on the input sequence, the problems associated with the commonly used Linear Prediction approach is alleviated leading to a more robust modeling scheme.
Ritwik Giri, Bhaskar D. Rao
ICASSP1
2014 User Behavior Modeling in a Cellular Network Using Latent Dirichlet Allocation
Ritwik Giri, Heesook Choi, Kevin Soo Hoo, Bhaskar D. Rao
IDEAL1
2011 An ecologically inspired direct search method for solving optimal control problems with Bézier parameterization
Arnob Ghosh, Swagatam Das, Aritra Chowdhury, Ritwik Giri
Eng. Appl. Artif. Intell.4
2011 An improved differential evolution algorithm with fitness-based adaptation of the control parameters
Arnob Ghosh, Swagatam Das, Aritra Chowdhury, Ritwik Giri
Inf. Sci.4
2010 Linear antenna array synthesis using fitness-adaptive differential evolution algorithm
abstract
Design of non-uniform linear antenna arrays is one of the most important electromagnetic optimization problems of current interest. In this article, an adaptive Differential Evolution (DE) algorithm has been used to optimize the spacing between the elements of the linear array to produce a radiation pattern with minimum side lobe level and null placement control. DE is arguably one of the best real parameter optimizers of current interest takes very few control parameters and is easy to implement in any programming language. In this study two very simple adaptation schemes are used to regulate the control parameters F and Cr, upon which the performance of DE is critically dependent. The adaptation schemes are based on the objective function values of the target vectors and donor vectors. The adaptive DE-variant has been used to solve three difficult instances of the design problem and the optimization goal in each example is easily achieved. The results of the proposed algorithm have been shown to meet or beat the recently published results obtained using other state-of-the-art metaheuristics like the Genetic Algorithm (GA), Particle Swarm Optimization (PSO), Memetic Algorithms (MA), and Tabu Search (TS) in a statistically meaningful way.
Aritra Chowdhury, Ritwik Giri, Arnob Ghosh, Swagatam Das, Ajith Abraham, Václav Snásel
IEEE Congress on Evolutionary Computation2
2010 A hybrid evolutionary direct search technique for solving Optimal Control problems
abstract
An Optimal Control is a set of differential equations describing the path of the control variables that minimize the cost functional (function of both state and control variables). Direct solution methods for optimal control problems treat them from the perspective of global optimization: perform a global search for the control function that optimizes the required objective. Invasive Weed Optimization (IWO) technique is used here for optimal control. However, the direct solution method operates on discrete n-dimensional vectors, not on continuous functions, and becomes computationally unmanageable for large values of n. Thus, a parameterization technique is required, which can represent control functions using a small number of real-valued parameters. Typically, direct methods using evolutionary techniques parameterize control functions with a piecewise constant approximation. This has obvious limitations, both for accuracy in representing arbitrary functions, and for optimization efficiency. In this paper a new parameterization is introduced, using Bézier curves, which can accurately represent continuous control functions with only a few parameters. It is combined with Invasive Weed Optimization into a new evolutionary direct method for optimal control. The effectiveness of the new method is demonstrated by solving a wide range of optimal control problems.
Arnob Ghosh, Aritra Chowdhury, Ritwik Giri, Swagatam Das, Ajith Abraham
HIS3
2010 A Modified Invasive Weed Optimization Algorithm for training of feed- forward Neural Networks
abstract
Invasive Weed Optimization Algorithm IWO) is an ecologically inspired metaheuristic that mimics the process of weeds colonization and distribution and is capable of solving multi-dimensional, linear and nonlinear optimization problems with appreciable efficiency. In this article a modified version of IWO has been used for training the feed-forward Artificial Neural Networks (ANNs) by adjusting the weights and biases of the neural network. It has been found that modified IWO performs better than another very competitive real parameter optimizer called Differential Evolution (DE) and a few classical gradient-based optimization algorithms in context to the weight training of feed-forward ANNs in terms of learning rate and solution quality. Moreover, IWO can also be used in validation of reached optima and in the development of regularization terms and non-conventional transfer functions that do not necessarily provide gradient information
Ritwik Giri, Aritra Chowdhury, Arnob Ghosh, Swagatam Das, Ajith Abraham, Václav Snásel
SMC1