Vinayak Abrol

dblp:136/4698 · DBLP profile ↗
← Back
38ranked-venue papers
10as first author
21since 2021 · last 2026
0000-0001-8149-8151ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 7 first-author · 13 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 NADIR: Differential Attention Flow for Non-Autoregressive Transliteration in Indic Languages
abstract
In this work, we argue that not all sequence-to-sequence tasks require the strong inductive biases of autoregressive (AR) models. Tasks like multilingual transliteration, code refactoring, grammatical correction or text normalization often rely on local dependencies where the full modeling capacity of AR models can be overkill, creating a trade-off between their high accuracy and high inference latency. While non-autoregressive (NAR) models offer speed, they typically suffer from hallucinations and poor length control. To explore this trade-off, we focus on the multilingual transliteration task in Indic languages and introduce NADIR, a novel NAR architecture designed to strike a balance between speed and accuracy. NADIR integrates a Differential Transformer and a Mixture-of-Experts mechanism, enabling it to robustly model complex character mappings without sequential dependencies. NADIR achieves over a 13× speed-up compared to the state-of-the-art AR baseline. It maintains a competitive mean Character Error Rate of 15.78%, compared to 14.44% for the AR model and 21.88% for a standard NAR equivalent. Importantly, NADIR reduces Repetition errors by 49.53%, Substitution errors by 24.45%, Omission errors by 32.92%, and Insertion errors by 16.87%. This work provides a practical blueprint for building fast and reliable NAR systems, effectively bridging the gap between AR accuracy and the demands of real-time, large-scale deployment.
Lakshya Tomar, Vinayak Abrol, Puneet Agarwal
AAAI2
2026 BLEND: Balanced and Leaf-Enhanced Dual Fine-Tuning for Taxonomy Completion
abstract
Taxonomy completion is the task of integrating new concepts into an existing taxonomy by determining the appropriate hypernym--hyponym relations. Existing approaches often struggle with the inherent imbalance between leaf and non-leaf edges, which induces bias in representation learning. In this paper, we propose BLEND: B alanced and L eaf- En hanced D ual Fine-Tuning for Taxonomy Completion, a novel framework designed to mitigate this inductive bias. Our method employs independent fine-tuning of two lightweight large language models (LLMs): one optimized with a leaf-focused objective and the other trained with a balanced focused strategy. To further enhance structural understanding, we apply contrastive learning over structure-encoded paths and introduce a combined loss function, enabling more robust representation of hierarchical relations. Extensive experiments on three real-world benchmark datasets demonstrate that BLEND achieves up to 9.32% improvement in recall or hit metrics compared to state-of-the-art approaches. Moreover, BLEND delivers efficient inference while outperforming the latest baseline COMI, highlighting its effectiveness for taxonomy completion tasks.
Pankaj, Dhruv Kumar 0001, Vinayak Abrol, Vikram Goyal
WWW3
2026 AdaSpeech: Multiresolution modality adaptation for ASR in Hindi
Puneet Singh Bhooi, Vinayak Abrol
Speech Commun.2
2025 ArticulateX: End-to-End Monolingual Speech Translation in Articulator Space
Vinayak Abrol
INTERSPEECH2
2025 From Pretraining to Performance: Benchmarking Self-Supervised Speech Models for Interspeech-25 SER Challenge
Drishya Uniyal, Vinayak Abrol
INTERSPEECH2
2025 Evaluating Generative Models via Cubical Homology based Persistent Entropy
abstract
Topological tools have become popular in improving and evaluating the performance of generative models by exploring the connection between their representation power and topological properties.This has led to the development of various measures that can assess the diversity and quality of generated data.However, existing methods are impractical in higher dimensions and large-scale datasets/models.To address this, we propose a scalable framework based on persistent entropy.We first establish a theoretical relation between the homological complexity of the underlying topology and the persistent entropy.We then empirically study the topological transformation during training of the generated data manifold using cubical homology.The proposed method is domain & modelagnostic and scales well for various neural architectures at different depths.
Suryaka Suresh, Vinayak Abrol
KDD (2)2
2024 Sampling Rate Adaptive Speaker Verification from Raw Waveforms
Vinayak Abrol, Anshul Thakur, Akshat Gupta, Xiaomo Liu, Sameena Shah
ICPR (28)1
2024 QGAN: Low Footprint Quaternion Neural Vocoder for Speech Synthesis
Aryan Chaudhary, Vinayak Abrol
INTERSPEECH2
2024 Efficient CNNs with Quaternion Transformations and Pruning for Audio Tagging
Aryan Chaudhary, Arshdeep Singh, Vinayak Abrol, Mark D. Plumbley
INTERSPEECH3
2024 On characterizing the evolution of embedding space of neural networks using algebraic topology
Suryaka Suresh, Bishshoy Das, Vinayak Abrol, Sumantra Dutta Roy
Pattern Recognit. Lett.3
2024 Statistically Guided Near-End Speech Intelligibility Improvement Through Voice Transformation and Transfer Learning
abstract
In recent developments, speech intelligibility has been improved through an optimal trapezoidal transformation function, which performed normal to Lombard speech conversion via formant shifting. Despite performing well, the optimization took very long to converge and led to artifacts in the modified signal due to aggressive formant shifts in unvoiced frames. Therefore, transfer learning was used to rapidly modify the optimized parameters for a target language to bypass re-optimization for a new language. However, such transfer across noises was left unaddressed. This work proposes a Gaussian transformation function to perform statistically guided normal to Lombard speech conversion. Optimizing fewer parameters ensures faster convergence than before. The new transformation function generates fewer artifacts during voice modification while performing at par with the earlier function. This work enhances transfer learning performance by mitigating the directional nature in case of language mismatch. We also propose the transfer learning across noises using the comparative estimations of noise magnitude spectra, which was not feasible earlier. The simultaneous transfer of parameters across languages and noises is now feasible via the proposed Gaussian transformation function. We also explore the statistical difference between formant shifts produced by the Gaussian transformation function and its predecessor and their effect on intelligibility improvement. All experiments were conducted on exhaustive combinations of three languages, four noise types, and three SNR levels.
Ritujoy Biswas, Karan Nathwani, Vinayak Abrol
IEEE ACM Trans. Audio Speech Lang. Process.3
2024 On the Quantization of Neural Models for Speaker Verification
abstract
This paper addresses the sub-optimality of current post-training quantization (PTQ) and quantization-aware training (QAT) methods for state-of-the-art speaker verification (SV) models featuring intricate architectural elements such as channel aggregation and squeeze excitation modules. To address these limitations, we propose 1) a data-independent PTQ technique employing iterative low-precision calibration on pre-trained models; and 2) a data-dependent QAT method designed to reduce the performance gap between full-precision and integer models. Our QAT involves two progressive stages where FP-32 weights are initially transformed into FP-8, adapting precision based on the gradient norm, followed by the learning of quantizer parameters (scale and zero-point) for INT8 conversion. Experimental validation underscores the ingenuity of our method in model quantization, demonstrating reduced floating-point operations and INT8 inference time, all while maintaining performance on par with full-precision models.
Vinayak Abrol, Mathew Magimai-Doss
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 Incremental Trainable Parameter Selection in Deep Neural Networks
abstract
This article explores the utilization of the effective degree-of-freedom (DoF) of a deep learning model to regularize its stochastic gradient descent (SGD)-based training. The effective DoF of a deep learning model is defined only by a subset of its total parameters. This subset is highly responsive or sensitive toward the training loss, and its cardinality can be used to govern the effective DoF of a model during training. To this aim, the incremental trainable parameter selection (ITPS) algorithm is introduced in this article. The proposed ITPS algorithm acts as a wrapper over SGD and incrementally selects the parameters for updation that exhibit the maximum sensitivity toward the training loss. Hence, it gradually increases the DoF of the model during training. In ideal cases, the proposed algorithm arrives at a model configuration (i.e., DoF) optimum for the task at hand. This whole process results in a regularization-like behavior induced by a gradual increment of the DoF. Since the selection and updation of parameters is a function of the training loss, the proposed algorithm can be seen as a task and data-dependent regularization mechanism. This article exhibits the general utility of ITPS by evaluating it on various prominent neural network architectures such as CNNs, transformers, recurrent neural networks (RNNs), and multilayer perceptrons. These models are trained for image classification and healthcare tasks using the publicly available CIFAR-10, SLT-10, and MIMIC-III datasets.
Anshul Thakur, Vinayak Abrol, Pulkit Sharma, Tingting Zhu 0001, David A. Clifton
IEEE Trans. Neural Networks Learn. Syst.2
2023 Coordinate descent on the Stiefel manifold for deep neural network training
abstract
To alleviate the cost incurred by orthogonality constraints in optimization and model training, we propose a stochastic coordinate descent algorithm on the Stiefel manifold.We compute expressions for geodesics on the Stiefel manifold with initial velocity aligned with coordinates of the tangent space and show that, analogously to the orthogonal group, iterate updates of coordinate descent methods can be efficiently implemented in terms of multiplications by Givens matrices.We illustrate our proposed algorithm on deep neural network training.
Estelle M. Massart, Vinayak Abrol
ESANN2
2023 Investigating Acoustic Cues for Multilingual Abuse Detection
Yash Thakran, Vinayak Abrol
INTERSPEECH2
2023 Multi-Layer Acoustic & Linguistic Feature Fusion for ComParE-23 Emotion and Requests Challenge
abstract
The ACM Multimedia 2023 ComParE challenge focuses on classification/regression tasks for spoken customer-agent and emotionally rated conversations. The challenge baseline systems build upon the recent advancement in large-scale supervised/unsupervised foundational acoustic models that demonstrate consistently good performance across tasks. In this work, with the aim of improving the performance further, we present a novel multi-layer feature fusion method. In particular, the proposed approach leverages the hierarchical information from acoustic models using multi-layer statistics pooling, where we compute the weighted sum of layer-wise (mean and standard deviation) features. We further experimented with linguistic features and their late fusion with acoustic features, especially for subtasks involving complex conversations. Exploring various combinations of methods and features, we present four different systems tailored for each subchallenge, demonstrating significant performance gains over the baseline on the development and test set.
Siddhant R. Viksit, Vinayak Abrol
ACM Multimedia2
2022 Coordinate Descent on the Orthogonal Group for Recurrent Neural Network Training
abstract
We address the poor scalability of learning algorithms for orthogonal recurrent neural networks via the use of stochastic coordinate descent on the orthogonal group, leading to a cost per iteration that increases linearly with the number of recurrent states. This contrasts with the cubic dependency of typical feasible algorithms such as stochastic Riemannian gradient descent, which prohibits the use of big network architectures. Coordinate descent rotates successively two columns of the recurrent matrix. When the coordinate (i.e., indices of rotated columns) is selected uniformly at random at each iteration, we prove convergence of the algorithm under standard assumptions on the loss function, stepsize and minibatch noise. In addition, we numerically show that the Riemannian gradient has an approximately sparse structure. Leveraging this observation, we propose a variant of our proposed algorithm that relies on the Gauss-Southwell coordinate selection rule. Experiments on a benchmark recurrent neural network training problem show that the proposed approach is a very promising step towards the training of orthogonal recurrent neural networks with big architectures.
Estelle M. Massart, Vinayak Abrol
AAAI2
2022 Time-Frequency and Geometric Analysis of Task-Dependent Learning in Raw Waveform Based Acoustic Models
abstract
End-to-end raw-waveform modelling with learnable feature extraction front-ends has shown promising results in various speech/audio tasks. Despite its varied success, there have not been many attempts to understand how spectral/temporal feature integration from raw inputs helps recognize task-dependent information. Towards this aim, this work presents data-dependent and data-independent methods for understanding the modelling behavior of acoustic models. The first method employs time-frequency analysis to visualize input-specific response spectra as a function of short-time front-end block processing. The second method employs geometric properties of layer-wise weights to quantify the impact of architectural choices on signal propagation and trainability of the model. We demonstrate potential of the proposed methods with help of case studies on speech classification, speaker identification, and spoofing classification tasks.
Devansh Gupta, Vinayak Abrol
ICASSP2
2022 Data Pre-Processing Using Neural Processes for Modeling Personalized Vital-Sign Time-Series Data
abstract
Clinical time-series data retrieved from electronic medical records are widely used to build predictive models of adverse events to support resource management. Such data is often sparse and irregularly-sampled, which makes it challenging to use many common machine learning methods. Missing values may be interpolated by carrying the last value forward, or through linear regression. Gaussian process (GP) regression is also used for performing imputation, and often re-sampling of time-series at regular intervals. The use of GPs can require extensive, and likely adhoc, investigation to determine model structure, such as an appropriate covariance function. This can be challenging for multivariate real-world clinical data, in which time-series variables exhibit different dynamics to one another. In this work, we construct generative models to estimate missing values in clinical time-series data using a neural latent variable model, known as a Neural Process (NP). The NP model employs a conditional prior distribution in the latent space to learn global uncertainty in the data by modelling variations at a local level. In contrast to conventional generative modelling, this prior is not fixed and is itself learned during the training process. Thus, NP model provides the flexibility to adapt to the dynamics of the available clinical data. We propose a variant of the NP framework for efficient modelling of the mutual information between the latent and input spaces, ensuring meaningful learned priors. Experiments using the MIMIC III dataset demonstrate the effectiveness of the proposed approach as compared to conventional methods.
Pulkit Sharma, Farah Shamout, Vinayak Abrol, David A. Clifton
IEEE J. Biomed. Health Informatics3
2021 Transfer Learning for Speech Intelligibility Improvement in Noisy Environments
Ritujoy Biswas, Karan Nathwani, Vinayak Abrol
Interspeech3
2021 Improving Generative Modelling in VAEs Using Multimodal Prior
abstract
In this paper we propose a conditional generative modelling (CGM) approach for unsupervised disentangled representation learning using variational autoencoder (VAE). CGM employs a multimodal/categorical conditional prior distribution in the latent space to learn global uncertainty in data by modelling the variations at local level. Thus, the proposed framework enforces the model to independently estimate the inherent patterns within each category, which improves the interpretability of the latent representations learned by the VAE model. The evidence lower bound objective for training the generative model is maximized using a mutual information criterion between the global latent categorical variable and the encoded inputs. Further, the approach has a built-in mechanism for bounding the information flow between the encoder and the decoder which addresses the problems of posterior collapse in conventional VAE models. Experiments on a variety of datasets demonstrate that our objective can learn disentangled representations and the proposed approach achieves competitive results on various task such as generative modelling, image classification and image denoising.
Vinayak Abrol, Pulkit Sharma, Arijit Patra
IEEE Trans. Multim.1
2020 A Geometric Approach to Archetypal Analysis via Sparse Projections
abstract
Archetypal analysis (AA) aims to extract patterns using self-expressive decomposition of data as convex combinations of extremal points (on the convex hull) of the data. This work presents a computationally efficient greedy AA (GAA) algorithm. GAA leverages the underlying geometry of AA, is scalable to larger datasets, and has significantly faster convergence rate. To achieve this, archetypes are learned via sparse projection of data. In the transformed space, GAA employs an iterative subset selection approach to identify archetypes based on the sparsity of convex representations. The work further presents the use of GAA algorithm for extended AA models such as robust and kernel AA. Experimental results show that GAA is considerably faster while performing comparable to existing methods for tasks such as classification, data visualization/categorization.
Vinayak Abrol, Pulkit Sharma
ICML1
2020 Learning Hierarchy Aware Embedding From Raw Audio for Acoustic Scene Classification
abstract
Recent advancements in modeling speech and audio signals using deep neural networks have shown that systems learning both features and the classifier can be built directly from raw signal. However, the performance of such end-to-end systems for acoustic scene classification (ASC) task is still not at par with conventional systems built using spectral features. In this work, we propose a raw waveform based end-to-end ASC system using convolutional neural network. In contrast to the existing studies using a non-hierarchical model, our framework leverages the hierarchical relations between acoustic categories to improve the classification performance. To this aim, our multi-task model is trained with coarse and fine labels that correspond to different levels of abstraction. In order to ensure consistency in the encoded information via label hierarchy, the proposed framework uses a prototypical model. Such a model ensures that the learned representations at least match to one of the global categorical learned prototypes. We also employed a statistical pooling layer to aggregate hidden representations over multiple frames of the input audio signal. The statistics (mean and standard deviation) are concatenated together to form a fixed-length audio embedding. This aggregation is done via an attention module so as to guide the model's attention even in presence of relatively short or transient acoustic events. Further, the proposed framework incorporate two parallel feature processing pipelines to achieve different resolutions for extracting important acoustic cues. Various experiments on publicly available datasets are performed to demonstrate the effectiveness of the proposed framework for ASC task. Additional transfer learning experiments showed the proposed model's adaptation capability to unseen data. Network analysis and visualizations demonstrate the importance of individual modules and their impact on overall representation learning for ASC task.
Vinayak Abrol, Pulkit Sharma
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 CONV-codes: Audio Hashing for Bird Species Classification
abstract
We propose a supervised, convex representation based audio hashing framework for bird species classification. The proposed framework utilizes archetypal analysis, a matrix factorization technique, to obtain convex-sparse representations of a bird vocalization. These convex representations are hashed using Bloom filters with non-cryptographic hash functions to obtain compact binary codes, designated as conv-codes. The conv-codes extracted from the training examples are clustered using class-specific k-medoids clustering with Jaccard coefficient as the similarity metric. A hash table is populated using the cluster centers as keys while hash values/slots are pointers to the species identification information. During testing, the hash table is searched to find the species information corresponding to a cluster center that exhibits maximum similarity with the test conv-code. Hence, the proposed framework classifies a bird vocalization in the conv-code space and requires no explicit classifier or reconstruction error calculations. Apart from that, based on min-hash and direct addressing, we also propose a variant of the proposed framework that provides faster and effective classification. The performances of both these frameworks are compared with existing bird species classification frameworks on the audio recordings of 50 different bird species.
Anshul Thakur, Pulkit Sharma, Vinayak Abrol, Padmanabhan Rajan
ICASSP3
2019 Understanding and Visualizing Raw Waveform-Based CNNs
abstract
Modeling directly raw waveforms through neural networks for speech processing is gaining more and more attention. Despite its varied success, a question that remains is: what kind of information are such neural networks capturing or learning for different tasks from the speech signal? Such an insight is not only interesting for advancing those techniques but also for understanding better speech signal characteristics. This paper takes a step in that direction, where we develop a gradient based approach to estimate the relevance of each speech sample input on the output score. We show that analysis of the resulting ``relevance signal" through conventional speech signal processing techniques can reveal the information modeled by the whole network. We demonstrate the potential of the proposed approach by analyzing raw waveform CNN-based phone recognition and speaker identification systems.
Hannah Muckenhirn, Vinayak Abrol, Mathew Magimai-Doss, Sébastien Marcel
INTERSPEECH2
2018 Compressed Convex Spectral Embedding for Bird Species Classification
abstract
This paper focuses on the problem of bird species identification using audio recordings. Following recent developments in deep learning, we propose a multi-layer alternating sparse-dense framework for bird species identification. Temporal and frequency modulations in bird vocalizations are captured by concatenating frames of spectrograms, resulting in a high dimensional super-frame based representation. These super-frame representations are highly sparse. Hence, we propose to use random projections to compress these super-frames. This is followed by class-specific archetypal analysis, employed on these compressed super-frames for acoustic modeling, to obtain a convex-sparse representation. These convex-sparse representations are referred as compressed convex spectral embeddings (CCSE). It is observed that these representations efficiently capture species-specific discriminative information. Experimental results show compelling evidence that the proposed approach shows performance comparable to existing methods such as deep neural networks (DNN) and dynamic kernel based SVMs.
Anshul Thakur, Vinayak Abrol, Pulkit Sharma, Padmanabhan Rajan
ICASSP2
2018 ASe: Acoustic Scene Embedding Using Deep Archetypal Analysis and GMM
Pulkit Sharma, Vinayak Abrol, Anshul Thakur
INTERSPEECH2
2018 Deep Convex Representations: Feature Representations for Bioacoustics Classification
Anshul Thakur, Vinayak Abrol, Pulkit Sharma, Padmanabhan Rajan
INTERSPEECH2
2018 Sparse coding based features for speech units classification
Pulkit Sharma, Vinayak Abrol, Aroor Dinesh Dileep, Anil Kumar Sao
Comput. Speech Lang.2
2018 Reducing footprint of unit selection based text-to-speech system using compressed sensing and sparse representation
Pulkit Sharma, Vinayak Abrol, Nivedita, Anil Kumar Sao
Comput. Speech Lang.2
2017 Fast exemplar selection algorithm for matrix approximation and representation: A variant oASIS algorithm
abstract
Extracting inherent patterns from large data using decompositions of data matrix by a sampled subset of exemplars has found many applications in machine learning. We propose a computationally efficient algorithm for adaptive exemplar sampling, called fast exemplar selection (FES). The proposed algorithm can be seen as an efficient variant of the oASIS algorithm [1]. FES iteratively selects incoherent exemplars based on the exemplars that are already sampled. This is done by ensuring that the selected exemplars forms a positive definite Gram matrix which is checked by exploiting its Cholesky factorization in an incremental manner. FES is a deterministic rank revealing algorithm delivering a tighter matrix approximation bound. Further, FES can also be used to exactly represent low rank matrices and signals sampled from a unions of independent subspaces. Experimental results show that FES performs comparable to existing methods for tasks such as matrix approximation, feature selection, outlier detection, and clustering.
Vinayak Abrol, Pulkit Sharma, Anil Kumar Sao
ICASSP1
2017 Deep-Sparse-Representation-Based Features for Speech Recognition
abstract
Features derived using sparse representation (SR)-based approaches have been shown to yield promising results for speech recognition tasks. In most of the approaches, the SR corresponding to speech signal is estimated using a dictionary, which could be either exemplar based or learned. However, a single-level decomposition may not be suitable for the speech signal, as it contains complex hierarchical information about various hidden attributes. In this paper, we propose to use a multilevel decomposition (having multiple layers), also known as the deep sparse representation (DSR), to derive a feature representation for speech recognition. Instead of having a series of sparse layers, the proposed framework employs a dense layer between two sparse layers, which helps in efficient implementation. Our studies reveal that the representations obtained at different sparse layers of the proposed DSR model have complimentary information. Thus, the final feature representation is derived after concatenating the representations obtained at the sparse layers. This results in a more discriminative representation, and improves the speech recognition performance. Since the concatenation results in a high-dimensional feature, principal component analysis is used to reduce the dimension of the obtained feature. Experimental studies demonstrate that the proposed feature outperforms existing features for various speech recognition tasks.
Pulkit Sharma, Vinayak Abrol, Anil Kumar Sao
IEEE ACM Trans. Audio Speech Lang. Process.2
2016 Fast and robust FMRI unmixing using hierarchical dictionary learning
abstract
We propose a novel computationally efficient hierarchical dictionary learning (HDL) approach for data-driven unmixing and functional connectivity analysis of functional magnetic resonance imaging (fMRI) data. It is shown that by simultaneously exploiting the sparsity of the spatial brain maps and the incoherence among their evolution in time or task functions, one can achieve better performance while overcoming the drawbacks of existing approaches. The task functions constituting the dictionary, are learned using a hierarchical subset selection approach. Here, to enforce incoherence among atoms, any new atom is selected from suitable training candidates if it does not lie in the column span of past selected atoms. Also, since the sparsity of spatial maps is generally unknown and affected due to acquisition artifacts, HDL doesn't make use of an implicit sparse coding stage while dictionary update. This makes HDL a very fast and efficient data-driven approach for fMRI analysis. Experimental results on synthetic and real fMRI datasets provide compelling evidences that HDL performs better than existing state-of-the-art methods.
Vinayak Abrol, Pulkit Sharma, Shahrooz Faghih Roohi, Anil Kumar Sao, Ashraf A. Kassim
ICIP1
2016 Greedy dictionary learning for kernel sparse representation based classifier
Vinayak Abrol, Pulkit Sharma, Anil Kumar Sao
Pattern Recognit. Lett.1
2016 Greedy double sparse dictionary learning for sparse representation of speech signals
Vinayak Abrol, Pulkit Sharma, Anil Kumar Sao
Speech Commun.1
2015 Sparse coding based features for speech units classification
abstract
Abstract In this work, we propose sparse representation based features for speech units classification tasks. In order to effectively capture the variations in a speech unit, the proposed method employs multiple class specific dictionaries. Here, the training data belonging to each class is clustered into multiple clusters, and a principal component analysis (PCA) based dictionary is learnt for each cluster. It has been observed that coefficients corresponding to middle principal components can effectively discriminate among different speech units. Exploiting this observation, we propose to use a transformation function known as weighted decomposition (WD) of principal components, which is used to emphasize the discriminative information present in the PCA-based dictionary. In this paper, both raw speech samples and mel frequency cepstral coefficients (MFCC) are used as an initial representation for feature extraction. For comparison, various popular dictionary learning techniques such as K-singular value decomposition (KSVD), simultaneous codeword optimization (SimCO) and greedy adaptive dictionary (GAD) are also employed in the proposed framework. The effectiveness of the proposed features is demonstrated using continuous density hidden Markov model (CDHMM) based classifiers for (i) classification of isolated utterances of E-set of English alphabet, (ii) classification of consonant-vowel (CV) segments in Hindi language and (iii) classification of phoneme from TIMIT phonetic corpus.
Pulkit Sharma, Vinayak Abrol, Aroor Dinesh Dileep, Anil Kumar Sao
INTERSPEECH2
2015 Voiced/nonvoiced detection in compressively sensed speech signals
Vinayak Abrol, Pulkit Sharma, Anil Kumar Sao
Speech Commun.1
2013 Speech enhancement using compressed sensing
Vinayak Abrol, Pulkit Sharma, Anil Kumar Sao
INTERSPEECH1