EDBT 2026 Demo / reviewers in the wild / expert
Karan Nathwani
dblp:122/2654
· DBLP profile ↗
23ranked-venue papers
7as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 12 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Deep Multiscale Wavelet Scattering-Attention Model for Real-Time Underwater DoA EstimationabstractDirection-of-arrival (DoA) estimation is vital for underwater target localisation but is challenged by multipath propagation and noise in dynamic ocean environments. To tackle such issues, we introduce the multiscale attentive wavelet scattering fusion (MAWSF) for DoA estimation in the challenging underwater environment. This MAWSF is a deep learning-based network, trained on simulated underwater data, capable of generalising to real oceanic scenarios. MAWSF performs multi-scale feature extraction using directional convolutions and wavelet scattering transforms to capture spatial patterns and enhance robustness to noise. It also employs attention mechanisms to model global dependencies and enhance context awareness. To emulate real oceanic conditions, we simulate multipath propagation along with varying source-receiver depths, ranges, and ambient noise levels. These simulations are used to generate multi-channel spatial covariance matrices labelled with a one-hot encoded true DoAs vector for training. MAWSF is evaluated on the real SWellEx-96 dataset, where it demonstrated high angular resolution in offgrid DoA estimation. It achieves an RMSE of$14.76^{\circ }$, surpassing the performance of state-of-the-art networks. Unlike traditional methods, it requires no prior knowledge of noise or multipath counts, making it well-suited for real-time applications. Murtiza Ali, Karan Nathwani |
IEEE Signal Process. Lett. | 2 |
| 2025 | Exploiting Wavelet Scattering Transform & Squeeze-Excitation Blocks with Cross-Modal Attention for Multi-modal Emotion RecognitionabstractMulti-modal emotion recognition (MER) is crucial for improving human-computer interaction. Convolutional neural networks (CNNs) are the mainstream for MER tasks, but they require large databases, extensive memory, and significant energy, limiting their practical use. This paper proposes a novel MER system that leverages wavelet scattering transform (WST) to address these challenges, achieving improved performance with lower computational consumption. Moreover, the system benefits from the noise robustness provided by WST. By integrating WST as a non-trainable initial layer in a CNN model and employing an encoder module, our system effectively captures time-frequency, local and high-level features from both speech and video. We enhance feature integration and representation with cross-modal attention (CMA) and a squeeze-and-excitation (SE) block. The results demonstrate that our system performs consistently across varying noise levels and duration thresholds. Ablation studies reveal that the combination of MFCC, Mel spectrogram, and raw waveform features yields the highest accuracy, with Mel spectrogram being the most influential. Experimental results on the IEMOCAP and RAVDESS databases achieve emotion recognition accuracy of 83.2% and 97.8%, respectively, showcasing improved performance and robustness compared to state-of-the-art models, while using fewer trainable parameters. Jesin James, Karan Nathwani |
ICASSP | 3 |
| 2024 | Exploiting Wavelet Scattering Transform for an Unsupervised Speaker Diarization in Deep Neural Network Framework
Arunav Arya, Murtiza Ali, Karan Nathwani |
INTERSPEECH | 3 |
| 2024 | Exploiting Wavelet Scattering Transform and 1D-CNN for Unmanned Aerial Vehicle DetectionabstractRecent advancements in Unmanned Aerial Vehicles (UAVs) have prompted concerns regarding their potential misuse. Deep-learning techniques offer superior detection capabilities compared to traditional rule-based approaches, provided the training dataset is diverse and sufficiently large. We present a UAV acoustic dataset featuring a range of UAVs, from toys to high-speed models, recorded in an open field simulating airport environments. To ensure robust multi-conditional training, the dataset is augmented with noise at specific signal-to-noise ratios (SNRs). Further, we introduce a network called WST-CNN, which has a non-trainable Wavelet Scattering Transform (WST) layer with fixed initializations as the first layer in a Convolutional Neural Network (CNN). With raw audio as an input, exempting any pre-processing, WST provides a multi-resolution time-frequency representation resilient to noise and signal deformation. As a secondary result, we introduce a 1D-F-CNN network utilizing distinctive acoustic features for UAV detection. Murtiza Ali, Karan Nathwani |
IEEE Signal Process. Lett. | 2 |
| 2024 | Statistically Guided Near-End Speech Intelligibility Improvement Through Voice Transformation and Transfer LearningabstractIn recent developments, speech intelligibility has been improved through an optimal trapezoidal transformation function, which performed normal to Lombard speech conversion via formant shifting. Despite performing well, the optimization took very long to converge and led to artifacts in the modified signal due to aggressive formant shifts in unvoiced frames. Therefore, transfer learning was used to rapidly modify the optimized parameters for a target language to bypass re-optimization for a new language. However, such transfer across noises was left unaddressed. This work proposes a Gaussian transformation function to perform statistically guided normal to Lombard speech conversion. Optimizing fewer parameters ensures faster convergence than before. The new transformation function generates fewer artifacts during voice modification while performing at par with the earlier function. This work enhances transfer learning performance by mitigating the directional nature in case of language mismatch. We also propose the transfer learning across noises using the comparative estimations of noise magnitude spectra, which was not feasible earlier. The simultaneous transfer of parameters across languages and noises is now feasible via the proposed Gaussian transformation function. We also explore the statistical difference between formant shifts produced by the Gaussian transformation function and its predecessor and their effect on intelligibility improvement. All experiments were conducted on exhaustive combinations of three languages, four noise types, and three SNR levels. Ritujoy Biswas, Karan Nathwani, Vinayak Abrol |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Improved Multi-Modal Emotion Recognition Using Squeeze-and-Excitation Block in Cross-Modal AttentionabstractThe field of multi-modal emotion recognition has gained significant attention due to its applications in human-computer interaction. This paper explores recent advancements and challenges in multi-modal emotion recognition. We propose a neural network-based multi-modal emotion recognition model that applies attention mechanisms to fuse features extracted by the transformer model and convolutional neural networks. Specifically, we use pre-trained models, including BERT, ResNet, and DenseNet, to extract features from text, spectrograms, and face images, combining convolutional and recurrent layers to capture temporal information. We integrate all extracted features using cross-modal attention. In addition, a novel squeeze-and-excitation that enhances the representation ability of individual features is proposed to find the fine distinction between emotions, which improves the model’s recognition ability for similar emotional states. Our model is evaluated on the IEMOCAP and RAVDESS datasets, outperforming the state-of-the-art results. Jesin James, Karan Nathwani |
ASRU | 3 |
| 2023 | Exploiting Sparse Recovery Algorithms for Semi-Supervised Training of Deep Neural Networks for Direction-of-Arrival EstimationabstractThis paper proposes a semi-supervised training approach for a direction-of-arrival (DoA) estimation based on a convolutional neural network (CNN). We apply a sparse recovery algorithm called optMGD-ℓ1-SVD on the training dataset consisting of only unlabeled observed data to obtain binarized pseudo-spectra regarded as the CNN training targets (labels). The estimated DoAs are obtained at test time by performing peak picking on the CNN outputs. optMGD-ℓ1-SVD has been shown to perform well with a few sensors under low signal-to-noise ratio (SNR) conditions (up to −6 dB) by optimally reweighting the pseudo-spectra of ℓ1-SVD based on the application of group delay function on the pseudo-spectra of MUSIC. Since its hyperparameters are noise-sensitive, we assume that the SNR levels of the training dataset are known such that we can use the optimal ones. We also consider multi-condition training using data of multiple SNR levels to improve the robustness towards different noisy environments. We evaluated the trained networks, named optMGD-ℓ1-SVD-CNN and MGD-ℓ1-SVD-CNN, in terms of the average root-mean-square error and the resolution probability under low SNR conditions (up to −20 dB). We demonstrated that it performed well with a few sensors and snapshots, including at SNR levels unseen in the training data. Murtiza Ali, Aditya Arie Nugraha, Karan Nathwani |
ICASSP | 3 |
| 2022 | Group Delay based Methods for Detection and Recognition of Whispered SpeechabstractThe present study demonstrates the effectiveness of the group delay function for detection and recognition of whispered speech. The group delay function in its spectral form is able to differentiate phonated from the whispered speech. In particular, the mean height-bandwidth product of the formants in the lower and mid frequency regions of the short term spectrum of speech is used herein for whisper speech detection. The mean height-bandwidth vector is obtained herein across all frames and is further smoothed using a moving average filter. The smoothed temporal version on this vector is able to detect the phonated to whisper change points. Towards this end, the current study investigates the cepstral, linear prediction (LP), minimum variance distortionless response (MVDR) and numerator of the group delay based smoothing techniques for whisper change point detection. Experiments on whispered speech detection are performed on the CHAINS database. Experimental results are compared to various methods for whispered speech detection available in literature. Kishore Vedvyasan, Karan Nathwani, Rajesh M. Hegde |
ICPR | 2 |
| 2022 | Transformer-based quality assessment model for generalized user-generated multimedia audio content
Deebha Mumtaz, Ajit Jena, Vinit Jakhetiya, Karan Nathwani, Sharath Chandra Guntuku |
INTERSPEECH | 4 |
| 2022 | Nonintrusive Perceptual Audio Quality Assessment for User-Generated Content Using Deep LearningabstractWith the boom of social media communication, teleconferencing, and online classes, audiovisual communication over bandwidth strained networks has become an integral part of our lives. Consequently, the growing demand for the quality of experience necessitates developing algorithms to measure and enrich user experience. Prior studies have mainly focused on assessing speech quality and intelligibility with reference to audio quality assessment, while other categories in user-generated multimedia (UGM) are less explored. Moreover, frequency-domain properties of speech and UGM audio are significantly different from each other. Furthermore, there is a lack of a standard dataset for the quality assessment of UGM. Considering these limitations, in this article, we first develop the IIT-JMU-UGM audio dataset consisting of 1150 audio clips, with diverse context, content, and types of degradation commonly observed in real-world scenarios and annotated with the subjective quality scores. Finally, we propose a non-intrusive audio quality assessment metric using a stacked gated-recurrent-unit-based deep learning framework. The proposed model outperforms several baseline methods, including state-of-the-art non-intrusive and intrusive approaches. The resulting Pearson’s correlation coefficient of 0.834 indicates that the proposed method efficiently mirrors human auditory perception. Deebha Mumtaz, Vinit Jakhetiya, Karan Nathwani, Badri N. Subudhi, Sharath Chandra Guntuku |
IEEE Trans. Ind. Informatics | 3 |
| 2021 | Group Delay Based Re-Weighted Sparse Recovery Algorithms for Robust and High-Resolution Source Separation in DOA Framework
Murtiza Ali, Ashwani Koul, Karan Nathwani |
Interspeech | 3 |
| 2021 | Transfer Learning for Speech Intelligibility Improvement in Noisy Environments
Ritujoy Biswas, Karan Nathwani, Vinayak Abrol |
Interspeech | 2 |
| 2020 | Effect of Microphone Position Measurement Error on RIR and its Impact on Speech Intelligibility and Quality
Aditya Raikar, Karan Nathwani, Ashish Panda, Sunil Kumar Kopparapu |
INTERSPEECH | 2 |
| 2019 | Excitation Source and Vocal Tract System Based Acoustic Features for Detection of Nasals in Continuous Speech
Bhanu Teja Nellore, Sri Harsha Dumpala, Karan Nathwani, Suryakanth V. Gangashetty |
INTERSPEECH | 3 |
| 2018 | LSTM Based Attentive Fusion of Spectral and Prosodic Information for Keyword Spotting in Hindi Language
Laxmi Pandey, Karan Nathwani |
INTERSPEECH | 2 |
| 2018 | DNN Uncertainty Propagation Using GMM-Derived Uncertainty Features for Noise Robust ASRabstractThe uncertainty decoding framework is known to improve the deep neural network (DNN)-based automatic speech recognition (ASR) performance in noisy environments. It operates by estimating the statistical uncertainty about the input features and propagating it to the output senone posteriors by sampling. Unfortunately, this approximate propagation scheme limits the performance improvement. In this letter, we exploit the fact that uncertainty propagation can be achieved in closed form for Gaussian mixture acoustic models (GMMs). We introduce new GMM-derived (GMMD) uncertainty features for the robust DNN-based acoustic model training and decoding. The GMMD features are computed as the difference between the GMM log-likelihoods obtained with versus without uncertainty. They are concatenated with conventional acoustic features and used as inputs to the DNN. We evaluate the resulting ASR performance on the CHiME-2 and CHiME-3 datasets. The proposed features are shown to improve the performance on both datasets, both for the conventional decoding and for the uncertainty decoding with different uncertainty estimation/propagation techniques. Karan Nathwani, Emmanuel Vincent 0001, Irina Illina |
IEEE Signal Process. Lett. | 1 |
| 2017 | Consistent DNN uncertainty training and decoding for robust ASRabstractWe consider the problem of robust automatic speech recognition (ASR) in noisy conditions. The performance improvement brought by speech enhancement is often limited by residual distortions of the enhanced features, which can be seen as a form of statistical uncertainty. Uncertainty estimation and propagation methods have recently been proposed to improve the ASR performance with deep neural network (DNN) acoustic models. However, the performance is still limited due to the use of uncertainty only during decoding. In this paper, we propose a consistent approach to account for uncertainty in the enhanced features during both training and decoding. We estimate the variance of the distortions using a DNN uncertainty estimator that operates directly in the feature maximum likelihood linear regression (fMLLR) domain and we then sample the uncertain features using the unscented transform (UT). We report the resulting ASR performance on the CHiME-2 and CHiME-3 datasets for different uncertainty estimation/propagation techniques. The proposed DNN uncertainty training method brings 4% and 8% relative improvement on these two datasets, respectively, compared to a competitive fMLLR-domain DNN acoustic modeling baseline. Karan Nathwani, Emmanuel Vincent 0001, Irina Illina |
ASRU | 1 |
| 2017 | Speech intelligibility improvement in car noise environment by voice transformation
Karan Nathwani, Gaël Richard, Bertrand David 0002, Pierre Prablanc, Vincent Roussarie |
Speech Commun. | 1 |
| 2016 | Formant shifting for speech intelligibility improvement in car noise environmentabstractIn this paper, we propose a novel approach aiming at improving the intelligibility of speech in the context of in-car applications. Speech produced in noisy environments is subject to the Lombard effect which gathers a number of voice transformation effects compared to the speech produced in calm environments. To improve intelligibility of in car speech (radio, message alerts, ...), we propose to modify the original speech signal by incorporating one of the important Lombard effect, namely the shift of the lower formant center frequencies away from the competing noise regions. The proposed approach exploits traditional Linear Prediction analysis and overlap and add synthesis. We explore several modification strategies and the merit of each modification is evaluated using both objective and subjective tests. It is in particular shown that the improvement of speech intelligibility in car noise is significantly improved for a majority of listeners. Karan Nathwani, Morgane Daniel, Gaël Richard, Bertrand David 0002, Vincent Roussarie |
ICASSP | 1 |
| 2015 | Joint source separation and dereverberation using constrained spectral divergence optimization
Karan Nathwani, Rajesh M. Hegde |
Signal Process. | 1 |
| 2015 | Robust acoustic echo cancellation using Kalman filter in double talk scenario
Sanchit Goel, Karan Nathwani, Rajesh M. Hegde |
Speech Commun. | 3 |
| 2013 | Joint noise cancellation and dereverberation using multi-channel linearly constrained minimum variance filter
Karan Nathwani, Rajesh M. Hegde |
INTERSPEECH | 1 |
| 2013 | Group Delay Based Methods for Speaker Segregation and its Application in Multimedia Information RetrievalabstractA novel method of single channel speaker segregation using the group delay cross correlation function is proposed in this paper. The group delay function, which is the negative derivative of the phase spectrum, yields robust spectral estimates. Hence the group delay spectral estimates are first computed over frequency sub-bands after passing the speech signal through a bank of filters. The filter bank spacing is based on a multi-pitch algorithm that computes the pitch estimates of the competing speakers. An affinity matrix is then computed from the group delay spectral estimates of each frequency sub-band. This affinity matrix represents the correlations of the different sub-bands in the mixed broadband speech signal. The grouping of correlated harmonics present in the mixed speech signal is then carried out by using a new iterative graph cut method. The signals are reconstructed from the respective harmonic groups which represent individual speakers in the mixed speech signal. Spectrographic masks are then applied on the reconstructed signals to refine their perceptual quality. The quality of separated speech is evaluated using several objective and subjective criteria. Experiments on multi-speaker automatic speech recognition are conducted using mixed speech data from the GRID corpus. A cell phone based multimedia information retrieval system (MIRS) for multi-source meeting environments are also developed. Karan Nathwani, Pranav Pandit, Rajesh M. Hegde |
IEEE Trans. Multim. | 1 |