EDBT 2026 Demo / reviewers in the wild / expert
Arvindh Krishnaswamy
dblp:52/6492
· DBLP profile ↗
17ranked-venue papers
3as first author
10since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Neural Speech Synthesis on a Shoestring: Improving the Efficiency of LpcnetabstractNeural speech synthesis models can synthesize high quality speech but typically require a high computational complexity to do so. In previous work, we introduced LPCNet, which uses linear prediction to significantly reduce the complexity of neural synthesis. In this work, we further improve the efficiency of LPCNet – targeting both algorithmic and computational improvements – to make it usable on a wide variety of devices. We demonstrate an improvement in synthesis quality while operating 2.5x faster. The resulting open-source1LPCNet algorithm can perform real-time neural synthesis on most existing phones and is even usable in some embedded devices. Jean-Marc Valin, Umut Isik, Paris Smaragdis, Arvindh Krishnaswamy |
ICASSP | 4 |
| 2022 | Improved Singing Voice Separation with Chromagram-Based Pitch-Aware RemixingabstractSinging voice separation aims to separate music into vocals and accompaniment components. One of the major constraints for the task is the limited amount of training data with separated vocals. Data augmentation techniques such as random source mixing have been shown to make better use of existing data and mildly improve model performance. We propose a novel data augmentation technique, chromagram-based pitch-aware remixing, where music segments with high pitch alignment are mixed. By performing controlled experiments in both supervised and semi-supervised settings, we demonstrate that training models with pitch-aware remixing significantly improves the test signal-to-distortion ratio (SDR). Siyuan Yuan, Zhepei Wang, Umut Isik, Ritwik Giri, Jean-Marc Valin, Michael M. Goodwin, Arvindh Krishnaswamy |
ICASSP | 7 |
| 2022 | Clock Skew Robust Acoustic Echo Cancellation
Karim Helwani, Erfan Soltanmohammadi, Michael M. Goodwin, Arvindh Krishnaswamy |
INTERSPEECH | 4 |
| 2022 | End-to-end LPCNet: A Neural Vocoder With Fully-Differentiable LPC EstimationabstractNeural vocoders have recently demonstrated high quality speech synthesis, but typically require a high computational complexity.LPCNet was proposed as a way to reduce the complexity of neural synthesis by using linear prediction (LP) to assist an autoregressive model.At inference time, LPCNet relies on the LP coefficients being explicitly computed from the input acoustic features.That makes the design of LPCNet-based systems more complicated, while adding the constraint that the input features must represent a clean speech spectrum.We propose an end-to-end version of LPCNet that lifts these limitations by learning to infer the LP coefficients from the input features in the frame rate network .Results show that the proposed end-toend approach equals or exceeds the quality of the original LPC-Net model, but without explicit LP analysis.Our open-source 1 end-to-end model still benefits from LPCNet's low complexity, while allowing for any type of conditioning features. Krishna Subramani, Jean-Marc Valin, Umut Isik, Paris Smaragdis, Arvindh Krishnaswamy |
INTERSPEECH | 5 |
| 2022 | Real-Time Packet Loss Concealment With Mixed Generative and Predictive Model
Jean-Marc Valin, Ahmed Mustafa, Christopher Montgomery, Timothy B. Terriberry, Michael Klingbeil, Paris Smaragdis, Arvindh Krishnaswamy |
INTERSPEECH | 7 |
| 2021 | Enhancing into the Codec: Noise Robust Speech Coding with Vector-Quantized AutoencodersabstractAudio codecs based on discretized neural autoencoders have recently been developed and shown to provide significantly higher compression levels for comparable quality speech out-put. However, these models are tightly coupled with speech content, and produce unintended outputs in noisy conditions. Based on VQ-VAE autoencoders with WaveRNN decoders, we develop compressor-enhancer encoders and accompanying decoders, and show that they operate well in noisy conditions. We also observe that a compressor-enhancer model performs better on clean speech inputs than a compressor model trained only on clean speech. Jonah Casebeer, Vinjai Vale, Umut Isik, Jean-Marc Valin, Ritwik Giri, Arvindh Krishnaswamy |
ICASSP | 6 |
| 2021 | Enhancing Audio Augmentation Methods with Consistency LearningabstractData augmentation is an inexpensive way to increase training data diversity, and is commonly achieved via transformations of existing data. For tasks such as classification, there is a good case for learning representations of the data that are invariant to such transformations, yet this is not explicitly enforced by classification losses such as the cross-entropy loss. This paper investigates the use of training objectives that explicitly impose this consistency constraint, and how it can impact downstream audio classification tasks. In the context of deep convolutional neural networks in the supervised setting, we show empirically that certain measures of consistency are not implicitly captured by the cross-entropy loss, and that incorporating such measures into the loss function can improve the performance of tasks such as audio tagging. Put another way, we demonstrate how existing augmentation methods can further improve learning by enforcing consistency. Turab Iqbal, Karim Helwani, Arvindh Krishnaswamy, Wenwu Wang 0001 |
ICASSP | 3 |
| 2021 | Low-Complexity, Real-Time Joint Neural Echo Control and Speech Enhancement Based On PercepnetabstractSpeech enhancement algorithms based on deep learning have greatly surpassed their traditional counterparts and are now being considered for the task of removing acoustic echo from hands-free communication systems. This is a challenging problem due to both real-world constraints like loudspeaker non-linearities, and to limited compute capabilities in some communication systems. In this work, we propose a system combining a traditional acoustic echo canceller, and a low-complexity joint residual echo and noise suppressor based on a hybrid signal processing/deep neural network (DSP/DNN) approach. We show that the proposed system outperforms both traditional and other neural approaches, while requiring only 5.5% CPU for real-time operation. We further show that the system can scale to even lower complexity levels. Jean-Marc Valin, Srikanth V. Tenneti, Karim Helwani, Umut Isik, Arvindh Krishnaswamy |
ICASSP | 5 |
| 2021 | Semi-Supervised Singing Voice Separation With Noisy Self-TrainingabstractRecent progress in singing voice separation has primarily focused on supervised deep learning methods. However, the scarcity of ground-truth data with clean musical sources has been a problem for long. Given a limited set of labeled data, we present a method to leverage a large volume of unlabeled data to improve the model’s performance. Following the noisy self-training framework, we first train a teacher network on the small labeled dataset and infer pseudo-labels from the large corpus of unlabeled mixtures. Then, a larger student network is trained on combined ground-truth and self-labeled datasets. Empirical results show that the proposed self-training scheme, along with data augmentation methods, effectively leverage the large unlabeled corpus and obtain superior performance compared to supervised methods. Zhepei Wang, Ritwik Giri, Umut Isik, Jean-Marc Valin, Arvindh Krishnaswamy |
ICASSP | 5 |
| 2021 | Personalized PercepNet: Real-Time, Low-Complexity Target Voice Separation and EnhancementabstractThe presence of multiple talkers in the surrounding environment poses a difficult challenge for real-time speech communication systems considering the constraints on network size and complexity. In this paper, we present Personalized PercepNet, a real-time speech enhancement model that separates a target speaker from a noisy multi-talker mixture without compromising on complexity of the recently proposed PercepNet. To enable speaker-dependent speech enhancement, we first show how we can train a perceptually motivated speaker embedder network to produce a representative embedding vector for the given speaker. Personalized PercepNet uses the target speaker embedding as additional information to pick out and enhance only the target speaker while suppressing all other competing sounds. Our experiments show that the proposed model significantly outperforms PercepNet and other baselines, both in terms of objective speech enhancement metrics and human opinion scores. Ritwik Giri, Shrikant Venkataramani, Jean-Marc Valin, Umut Isik, Arvindh Krishnaswamy |
Interspeech | 5 |
| 2020 | Efficient Trainable Front-Ends for Neural Speech EnhancementabstractMany neural speech enhancement and source separation systems operate in the time-frequency domain. Such models often benefit from making their Short-Time Fourier Transform (STFT) front-ends trainable. In current literature, these are implemented as large Discrete Fourier Transform matrices; which are prohibitively inefficient for low-compute systems. We present an efficient, trainable front-end based on the butterfly mechanism to compute the Fast Fourier Transform, and show its accuracy and efficiency benefits for low-compute neural speech enhancement models. We also explore the effects of making the STFT window trainable. Jonah Casebeer, Umut Isik, Shrikant Venkataramani, Arvindh Krishnaswamy |
ICASSP | 4 |
| 2020 | Channel-Attention Dense U-Net for Multichannel Speech EnhancementabstractSupervised deep learning has gained significant attention for speech enhancement recently. The state-of-the-art deep learning methods perform the task by learning a ratio/binary mask that is applied to the mixture in the time-frequency domain to produce the clean speech. Despite the great performance in the single-channel setting, these frameworks lag in performance in the multichannel setting as the majority of these methods a) fail to exploit the available spatial information fully, and b) still treat the deep architecture as a black box which may not be well-suited for multichannel audio processing. This paper addresses these drawbacks, a) by utilizing complex ratio masking instead of masking on the magnitude of the spectrogram, and more importantly, b) by introducing a channel-attention mechanism inside the deep architecture to mimic beamforming. We propose Channel-Attention Dense U-Net, in which we apply the channel-attention unit recursively on feature maps at every layer of the network, enabling the network to perform non-linear beamforming. We demonstrate the superior performance of the network against the state-of-the-art approaches on the CHiME-3 dataset. Bahareh Tolooshams, Ritwik Giri, Andrew H. Song, Umut Isik, Arvindh Krishnaswamy |
ICASSP | 5 |
| 2020 | PoCoNet: Better Speech Enhancement with Frequency-Positional Embeddings, Semi-Supervised Conversational Data, and Biased LossabstractNeural network applications generally benefit from larger-sized models, but for current speech enhancement models, larger scale networks often suffer from decreased robustness to the variety of real-world use cases beyond what is encountered in training data. We introduce several innovations that lead to better large neural networks for speech enhancement. The novel PoCoNet architecture is a convolutional neural network that, with the use of frequency-positional embeddings, is able to more efficiently build frequency-dependent features in the early layers. A semi-supervised method helps increase the amount of conversational training data by pre-enhancing noisy datasets, improving performance on real recordings. A new loss function biased towards preserving speech quality helps the optimization better match human perceptual opinions on speech quality. Ablation experiments and objective and human opinion metrics show the benefits of the proposed improvements. Umut Isik, Ritwik Giri, Neerad Phansalkar, Jean-Marc Valin, Karim Helwani, Arvindh Krishnaswamy |
INTERSPEECH | 6 |
| 2020 | A Perceptually-Motivated Approach for Low-Complexity, Real-Time Enhancement of Fullband SpeechabstractOver the past few years, speech enhancement methods based on deep learning have greatly surpassed traditional methods based on spectral subtraction and spectral estimation. Many of these new techniques operate directly in the the short-time Fourier transform (STFT) domain, resulting in a high computational complexity. In this work, we propose PercepNet, an efficient approach that relies on human perception of speech by focusing on the spectral envelope and on the periodicity of the speech. We demonstrate high-quality, real-time enhancement of fullband (48 kHz) speech with less than 5% of a CPU core. Jean-Marc Valin, Umut Isik, Neerad Phansalkar, Ritwik Giri, Karim Helwani, Arvindh Krishnaswamy |
INTERSPEECH | 6 |
| 2003 | Application of pitch tracking to South Indian classical musicabstractWe present results of applying pitch trackers to samples of South Indian classical (Carnatic) music. In particular, we investigate the various musical notes used and their intonation. We try different pitch tracking methods and observe their performance in Carnatic music analysis. Examining our data, we find only 12 distinct intervals per octave among the notes that are played with constant pitch. However, there are pitch inflexions used sometimes that are not mere ornamentations - they are essential to the correct rendition of certain notes. Though these inflexions can be viewed as different versions of a particular note, they are certainly not equivalent to constant-pitch intervals like just intonation intervals, semitones or quartertones. Arvindh Krishnaswamy |
ICASSP (5) | 1 |
| 2003 | Application of pitch tracking to South Indian classical musicabstractWe present results of applying pitch trackers to samples of South Indian classical (Carnatic) music. In particular, we investigate the various musical notes used and their intonation. We try different pitch tracking methods and observe their performance in Carnatic music analysis. Examining our data, we find only 12 distinct intervals per octave among the notes that are played with constant pitch. However, there are pitch inflexions used sometimes that are not mere ornamentations - they are essential to the correct rendition of certain notes. Though these inflexions can be viewed as different versions of a particular note, they are certainly not equivalent to constant-pitch intervals like just intonation intervals, semitones or quartertones. Arvindh Krishnaswamy |
ICME | 1 |
| 2003 | Inferring control inputs to an acoustic violin from audio spectraabstractWe present a method to analyze the streaming sound output of a real acoustic violin and infer the control inputs used by the musician. In particular, we extract the following parameters: which note was played, which string it was played on, whether the instrument was bowed or plucked and the location of the bowing or plucking point from a discrete set. The approach we use, pattern classification of the short-time Fourier transform (STFT) frames of the sound signal, is general enough to enable application to almost any musical instrument. Arvindh Krishnaswamy, Julius O. Smith III |
ICME | 1 |