EDBT 2026 Demo / reviewers in the wild / expert
Vinayak Abrol
dblp:136/4698
· DBLP profile ↗
38ranked-venue papers
10as first author
21since 2021 · last 2026
0000-0001-8149-8151ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 7 first-author · 13 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NADIR: Differential Attention Flow for Non-Autoregressive Transliteration in Indic LanguagesabstractIn this work, we argue that not all sequence-to-sequence tasks require the strong inductive biases of autoregressive (AR) models. Tasks like multilingual transliteration, code refactoring, grammatical correction or text normalization often rely on local dependencies where the full modeling capacity of AR models can be overkill, creating a trade-off between their high accuracy and high inference latency. While non-autoregressive (NAR) models offer speed, they typically suffer from hallucinations and poor length control. To explore this trade-off, we focus on the multilingual transliteration task in Indic languages and introduce NADIR, a novel NAR architecture designed to strike a balance between speed and accuracy. NADIR integrates a Differential Transformer and a Mixture-of-Experts mechanism, enabling it to robustly model complex character mappings without sequential dependencies. NADIR achieves over a 13× speed-up compared to the state-of-the-art AR baseline. It maintains a competitive mean Character Error Rate of 15.78%, compared to 14.44% for the AR model and 21.88% for a standard NAR equivalent. Importantly, NADIR reduces Repetition errors by 49.53%, Substitution errors by 24.45%, Omission errors by 32.92%, and Insertion errors by 16.87%. This work provides a practical blueprint for building fast and reliable NAR systems, effectively bridging the gap between AR accuracy and the demands of real-time, large-scale deployment. Lakshya Tomar, Vinayak Abrol, Puneet Agarwal |
AAAI | 2 |
| 2026 | BLEND: Balanced and Leaf-Enhanced Dual Fine-Tuning for Taxonomy CompletionabstractTaxonomy completion is the task of integrating new concepts into an existing taxonomy by determining the appropriate hypernym--hyponym relations. Existing approaches often struggle with the inherent imbalance between leaf and non-leaf edges, which induces bias in representation learning. In this paper, we propose BLEND: B alanced and L eaf- En hanced D ual Fine-Tuning for Taxonomy Completion, a novel framework designed to mitigate this inductive bias. Our method employs independent fine-tuning of two lightweight large language models (LLMs): one optimized with a leaf-focused objective and the other trained with a balanced focused strategy. To further enhance structural understanding, we apply contrastive learning over structure-encoded paths and introduce a combined loss function, enabling more robust representation of hierarchical relations. Extensive experiments on three real-world benchmark datasets demonstrate that BLEND achieves up to 9.32% improvement in recall or hit metrics compared to state-of-the-art approaches. Moreover, BLEND delivers efficient inference while outperforming the latest baseline COMI, highlighting its effectiveness for taxonomy completion tasks. Pankaj, Dhruv Kumar 0001, Vinayak Abrol, Vikram Goyal |
WWW | 3 |
| 2026 | AdaSpeech: Multiresolution modality adaptation for ASR in Hindi
Puneet Singh Bhooi, Vinayak Abrol |
Speech Commun. | 2 |
| 2025 | ArticulateX: End-to-End Monolingual Speech Translation in Articulator Space
Vinayak Abrol |
INTERSPEECH | 2 |
| 2025 | From Pretraining to Performance: Benchmarking Self-Supervised Speech Models for Interspeech-25 SER Challenge
Drishya Uniyal, Vinayak Abrol |
INTERSPEECH | 2 |
| 2025 | Evaluating Generative Models via Cubical Homology based Persistent EntropyabstractTopological tools have become popular in improving and evaluating the performance of generative models by exploring the connection between their representation power and topological properties.This has led to the development of various measures that can assess the diversity and quality of generated data.However, existing methods are impractical in higher dimensions and large-scale datasets/models.To address this, we propose a scalable framework based on persistent entropy.We first establish a theoretical relation between the homological complexity of the underlying topology and the persistent entropy.We then empirically study the topological transformation during training of the generated data manifold using cubical homology.The proposed method is domain & modelagnostic and scales well for various neural architectures at different depths. Suryaka Suresh, Vinayak Abrol |
KDD (2) | 2 |
| 2024 | Sampling Rate Adaptive Speaker Verification from Raw Waveforms
Vinayak Abrol, Anshul Thakur, Akshat Gupta, Xiaomo Liu, Sameena Shah |
ICPR (28) | 1 |
| 2024 | QGAN: Low Footprint Quaternion Neural Vocoder for Speech Synthesis
Aryan Chaudhary, Vinayak Abrol |
INTERSPEECH | 2 |
| 2024 | Efficient CNNs with Quaternion Transformations and Pruning for Audio Tagging
Aryan Chaudhary, Arshdeep Singh, Vinayak Abrol, Mark D. Plumbley |
INTERSPEECH | 3 |
| 2024 | On characterizing the evolution of embedding space of neural networks using algebraic topology
Suryaka Suresh, Bishshoy Das, Vinayak Abrol, Sumantra Dutta Roy |
Pattern Recognit. Lett. | 3 |
| 2024 | Statistically Guided Near-End Speech Intelligibility Improvement Through Voice Transformation and Transfer LearningabstractIn recent developments, speech intelligibility has been improved through an optimal trapezoidal transformation function, which performed normal to Lombard speech conversion via formant shifting. Despite performing well, the optimization took very long to converge and led to artifacts in the modified signal due to aggressive formant shifts in unvoiced frames. Therefore, transfer learning was used to rapidly modify the optimized parameters for a target language to bypass re-optimization for a new language. However, such transfer across noises was left unaddressed. This work proposes a Gaussian transformation function to perform statistically guided normal to Lombard speech conversion. Optimizing fewer parameters ensures faster convergence than before. The new transformation function generates fewer artifacts during voice modification while performing at par with the earlier function. This work enhances transfer learning performance by mitigating the directional nature in case of language mismatch. We also propose the transfer learning across noises using the comparative estimations of noise magnitude spectra, which was not feasible earlier. The simultaneous transfer of parameters across languages and noises is now feasible via the proposed Gaussian transformation function. We also explore the statistical difference between formant shifts produced by the Gaussian transformation function and its predecessor and their effect on intelligibility improvement. All experiments were conducted on exhaustive combinations of three languages, four noise types, and three SNR levels. Ritujoy Biswas, Karan Nathwani, Vinayak Abrol |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | On the Quantization of Neural Models for Speaker VerificationabstractThis paper addresses the sub-optimality of current post-training quantization (PTQ) and quantization-aware training (QAT) methods for state-of-the-art speaker verification (SV) models featuring intricate architectural elements such as channel aggregation and squeeze excitation modules. To address these limitations, we propose 1) a data-independent PTQ technique employing iterative low-precision calibration on pre-trained models; and 2) a data-dependent QAT method designed to reduce the performance gap between full-precision and integer models. Our QAT involves two progressive stages where FP-32 weights are initially transformed into FP-8, adapting precision based on the gradient norm, followed by the learning of quantizer parameters (scale and zero-point) for INT8 conversion. Experimental validation underscores the ingenuity of our method in model quantization, demonstrating reduced floating-point operations and INT8 inference time, all while maintaining performance on par with full-precision models. Vinayak Abrol, Mathew Magimai-Doss |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | Incremental Trainable Parameter Selection in Deep Neural NetworksabstractThis article explores the utilization of the effective degree-of-freedom (DoF) of a deep learning model to regularize its stochastic gradient descent (SGD)-based training. The effective DoF of a deep learning model is defined only by a subset of its total parameters. This subset is highly responsive or sensitive toward the training loss, and its cardinality can be used to govern the effective DoF of a model during training. To this aim, the incremental trainable parameter selection (ITPS) algorithm is introduced in this article. The proposed ITPS algorithm acts as a wrapper over SGD and incrementally selects the parameters for updation that exhibit the maximum sensitivity toward the training loss. Hence, it gradually increases the DoF of the model during training. In ideal cases, the proposed algorithm arrives at a model configuration (i.e., DoF) optimum for the task at hand. This whole process results in a regularization-like behavior induced by a gradual increment of the DoF. Since the selection and updation of parameters is a function of the training loss, the proposed algorithm can be seen as a task and data-dependent regularization mechanism. This article exhibits the general utility of ITPS by evaluating it on various prominent neural network architectures such as CNNs, transformers, recurrent neural networks (RNNs), and multilayer perceptrons. These models are trained for image classification and healthcare tasks using the publicly available CIFAR-10, SLT-10, and MIMIC-III datasets. Anshul Thakur, Vinayak Abrol, Pulkit Sharma, Tingting Zhu 0001, David A. Clifton |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Coordinate descent on the Stiefel manifold for deep neural network trainingabstractTo alleviate the cost incurred by orthogonality constraints in optimization and model training, we propose a stochastic coordinate descent algorithm on the Stiefel manifold.We compute expressions for geodesics on the Stiefel manifold with initial velocity aligned with coordinates of the tangent space and show that, analogously to the orthogonal group, iterate updates of coordinate descent methods can be efficiently implemented in terms of multiplications by Givens matrices.We illustrate our proposed algorithm on deep neural network training. Estelle M. Massart, Vinayak Abrol |
ESANN | 2 |
| 2023 | Investigating Acoustic Cues for Multilingual Abuse Detection
Yash Thakran, Vinayak Abrol |
INTERSPEECH | 2 |
| 2023 | Multi-Layer Acoustic & Linguistic Feature Fusion for ComParE-23 Emotion and Requests ChallengeabstractThe ACM Multimedia 2023 ComParE challenge focuses on classification/regression tasks for spoken customer-agent and emotionally rated conversations. The challenge baseline systems build upon the recent advancement in large-scale supervised/unsupervised foundational acoustic models that demonstrate consistently good performance across tasks. In this work, with the aim of improving the performance further, we present a novel multi-layer feature fusion method. In particular, the proposed approach leverages the hierarchical information from acoustic models using multi-layer statistics pooling, where we compute the weighted sum of layer-wise (mean and standard deviation) features. We further experimented with linguistic features and their late fusion with acoustic features, especially for subtasks involving complex conversations. Exploring various combinations of methods and features, we present four different systems tailored for each subchallenge, demonstrating significant performance gains over the baseline on the development and test set. Siddhant R. Viksit, Vinayak Abrol |
ACM Multimedia | 2 |
| 2022 | Coordinate Descent on the Orthogonal Group for Recurrent Neural Network TrainingabstractWe address the poor scalability of learning algorithms for orthogonal recurrent neural networks via the use of stochastic coordinate descent on the orthogonal group, leading to a cost per iteration that increases linearly with the number of recurrent states. This contrasts with the cubic dependency of typical feasible algorithms such as stochastic Riemannian gradient descent, which prohibits the use of big network architectures. Coordinate descent rotates successively two columns of the recurrent matrix. When the coordinate (i.e., indices of rotated columns) is selected uniformly at random at each iteration, we prove convergence of the algorithm under standard assumptions on the loss function, stepsize and minibatch noise. In addition, we numerically show that the Riemannian gradient has an approximately sparse structure. Leveraging this observation, we propose a variant of our proposed algorithm that relies on the Gauss-Southwell coordinate selection rule. Experiments on a benchmark recurrent neural network training problem show that the proposed approach is a very promising step towards the training of orthogonal recurrent neural networks with big architectures. Estelle M. Massart, Vinayak Abrol |
AAAI | 2 |
| 2022 | Time-Frequency and Geometric Analysis of Task-Dependent Learning in Raw Waveform Based Acoustic ModelsabstractEnd-to-end raw-waveform modelling with learnable feature extraction front-ends has shown promising results in various speech/audio tasks. Despite its varied success, there have not been many attempts to understand how spectral/temporal feature integration from raw inputs helps recognize task-dependent information. Towards this aim, this work presents data-dependent and data-independent methods for understanding the modelling behavior of acoustic models. The first method employs time-frequency analysis to visualize input-specific response spectra as a function of short-time front-end block processing. The second method employs geometric properties of layer-wise weights to quantify the impact of architectural choices on signal propagation and trainability of the model. We demonstrate potential of the proposed methods with help of case studies on speech classification, speaker identification, and spoofing classification tasks. Devansh Gupta, Vinayak Abrol |
ICASSP | 2 |
| 2022 | Data Pre-Processing Using Neural Processes for Modeling Personalized Vital-Sign Time-Series DataabstractClinical time-series data retrieved from electronic medical records are widely used to build predictive models of adverse events to support resource management. Such data is often sparse and irregularly-sampled, which makes it challenging to use many common machine learning methods. Missing values may be interpolated by carrying the last value forward, or through linear regression. Gaussian process (GP) regression is also used for performing imputation, and often re-sampling of time-series at regular intervals. The use of GPs can require extensive, and likely adhoc, investigation to determine model structure, such as an appropriate covariance function. This can be challenging for multivariate real-world clinical data, in which time-series variables exhibit different dynamics to one another. In this work, we construct generative models to estimate missing values in clinical time-series data using a neural latent variable model, known as a Neural Process (NP). The NP model employs a conditional prior distribution in the latent space to learn global uncertainty in the data by modelling variations at a local level. In contrast to conventional generative modelling, this prior is not fixed and is itself learned during the training process. Thus, NP model provides the flexibility to adapt to the dynamics of the available clinical data. We propose a variant of the NP framework for efficient modelling of the mutual information between the latent and input spaces, ensuring meaningful learned priors. Experiments using the MIMIC III dataset demonstrate the effectiveness of the proposed approach as compared to conventional methods. Pulkit Sharma, Farah Shamout, Vinayak Abrol, David A. Clifton |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | Transfer Learning for Speech Intelligibility Improvement in Noisy Environments
Ritujoy Biswas, Karan Nathwani, Vinayak Abrol |
Interspeech | 3 |
| 2021 | Improving Generative Modelling in VAEs Using Multimodal PriorabstractIn this paper we propose a conditional generative modelling (CGM) approach for unsupervised disentangled representation learning using variational autoencoder (VAE). CGM employs a multimodal/categorical conditional prior distribution in the latent space to learn global uncertainty in data by modelling the variations at local level. Thus, the proposed framework enforces the model to independently estimate the inherent patterns within each category, which improves the interpretability of the latent representations learned by the VAE model. The evidence lower bound objective for training the generative model is maximized using a mutual information criterion between the global latent categorical variable and the encoded inputs. Further, the approach has a built-in mechanism for bounding the information flow between the encoder and the decoder which addresses the problems of posterior collapse in conventional VAE models. Experiments on a variety of datasets demonstrate that our objective can learn disentangled representations and the proposed approach achieves competitive results on various task such as generative modelling, image classification and image denoising. Vinayak Abrol, Pulkit Sharma, Arijit Patra |
IEEE Trans. Multim. | 1 |
| 2020 | A Geometric Approach to Archetypal Analysis via Sparse ProjectionsabstractArchetypal analysis (AA) aims to extract patterns using self-expressive decomposition of data as convex combinations of extremal points (on the convex hull) of the data. This work presents a computationally efficient greedy AA (GAA) algorithm. GAA leverages the underlying geometry of AA, is scalable to larger datasets, and has significantly faster convergence rate. To achieve this, archetypes are learned via sparse projection of data. In the transformed space, GAA employs an iterative subset selection approach to identify archetypes based on the sparsity of convex representations. The work further presents the use of GAA algorithm for extended AA models such as robust and kernel AA. Experimental results show that GAA is considerably faster while performing comparable to existing methods for tasks such as classification, data visualization/categorization. Vinayak Abrol, Pulkit Sharma |
ICML | 1 |
| 2020 | Learning Hierarchy Aware Embedding From Raw Audio for Acoustic Scene ClassificationabstractRecent advancements in modeling speech and audio signals using deep neural networks have shown that systems learning both features and the classifier can be built directly from raw signal. However, the performance of such end-to-end systems for acoustic scene classification (ASC) task is still not at par with conventional systems built using spectral features. In this work, we propose a raw waveform based end-to-end ASC system using convolutional neural network. In contrast to the existing studies using a non-hierarchical model, our framework leverages the hierarchical relations between acoustic categories to improve the classification performance. To this aim, our multi-task model is trained with coarse and fine labels that correspond to different levels of abstraction. In order to ensure consistency in the encoded information via label hierarchy, the proposed framework uses a prototypical model. Such a model ensures that the learned representations at least match to one of the global categorical learned prototypes. We also employed a statistical pooling layer to aggregate hidden representations over multiple frames of the input audio signal. The statistics (mean and standard deviation) are concatenated together to form a fixed-length audio embedding. This aggregation is done via an attention module so as to guide the model's attention even in presence of relatively short or transient acoustic events. Further, the proposed framework incorporate two parallel feature processing pipelines to achieve different resolutions for extracting important acoustic cues. Various experiments on publicly available datasets are performed to demonstrate the effectiveness of the proposed framework for ASC task. Additional transfer learning experiments showed the proposed model's adaptation capability to unseen data. Network analysis and visualizations demonstrate the importance of individual modules and their impact on overall representation learning for ASC task. Vinayak Abrol, Pulkit Sharma |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | CONV-codes: Audio Hashing for Bird Species ClassificationabstractWe propose a supervised, convex representation based audio hashing framework for bird species classification. The proposed framework utilizes archetypal analysis, a matrix factorization technique, to obtain convex-sparse representations of a bird vocalization. These convex representations are hashed using Bloom filters with non-cryptographic hash functions to obtain compact binary codes, designated as conv-codes. The conv-codes extracted from the training examples are clustered using class-specific k-medoids clustering with Jaccard coefficient as the similarity metric. A hash table is populated using the cluster centers as keys while hash values/slots are pointers to the species identification information. During testing, the hash table is searched to find the species information corresponding to a cluster center that exhibits maximum similarity with the test conv-code. Hence, the proposed framework classifies a bird vocalization in the conv-code space and requires no explicit classifier or reconstruction error calculations. Apart from that, based on min-hash and direct addressing, we also propose a variant of the proposed framework that provides faster and effective classification. The performances of both these frameworks are compared with existing bird species classification frameworks on the audio recordings of 50 different bird species. Anshul Thakur, Pulkit Sharma, Vinayak Abrol, Padmanabhan Rajan |
ICASSP | 3 |
| 2019 | Understanding and Visualizing Raw Waveform-Based CNNsabstractModeling directly raw waveforms through neural networks for speech processing is gaining more and more attention. Despite its varied success, a question that remains is: what kind of information are such neural networks capturing or learning for different tasks from the speech signal? Such an insight is not only interesting for advancing those techniques but also for understanding better speech signal characteristics. This paper takes a step in that direction, where we develop a gradient based approach to estimate the relevance of each speech sample input on the output score. We show that analysis of the resulting ``relevance signal" through conventional speech signal processing techniques can reveal the information modeled by the whole network. We demonstrate the potential of the proposed approach by analyzing raw waveform CNN-based phone recognition and speaker identification systems. Hannah Muckenhirn, Vinayak Abrol, Mathew Magimai-Doss, Sébastien Marcel |
INTERSPEECH | 2 |
| 2018 | Compressed Convex Spectral Embedding for Bird Species ClassificationabstractThis paper focuses on the problem of bird species identification using audio recordings. Following recent developments in deep learning, we propose a multi-layer alternating sparse-dense framework for bird species identification. Temporal and frequency modulations in bird vocalizations are captured by concatenating frames of spectrograms, resulting in a high dimensional super-frame based representation. These super-frame representations are highly sparse. Hence, we propose to use random projections to compress these super-frames. This is followed by class-specific archetypal analysis, employed on these compressed super-frames for acoustic modeling, to obtain a convex-sparse representation. These convex-sparse representations are referred as compressed convex spectral embeddings (CCSE). It is observed that these representations efficiently capture species-specific discriminative information. Experimental results show compelling evidence that the proposed approach shows performance comparable to existing methods such as deep neural networks (DNN) and dynamic kernel based SVMs. Anshul Thakur, Vinayak Abrol, Pulkit Sharma, Padmanabhan Rajan |
ICASSP | 2 |
| 2018 | ASe: Acoustic Scene Embedding Using Deep Archetypal Analysis and GMM
Pulkit Sharma, Vinayak Abrol, Anshul Thakur |
INTERSPEECH | 2 |
| 2018 | Deep Convex Representations: Feature Representations for Bioacoustics Classification
Anshul Thakur, Vinayak Abrol, Pulkit Sharma, Padmanabhan Rajan |
INTERSPEECH | 2 |
| 2018 | Sparse coding based features for speech units classification
Pulkit Sharma, Vinayak Abrol, Aroor Dinesh Dileep, Anil Kumar Sao |
Comput. Speech Lang. | 2 |
| 2018 | Reducing footprint of unit selection based text-to-speech system using compressed sensing and sparse representation
Pulkit Sharma, Vinayak Abrol, Nivedita, Anil Kumar Sao |
Comput. Speech Lang. | 2 |
| 2017 | Fast exemplar selection algorithm for matrix approximation and representation: A variant oASIS algorithmabstractExtracting inherent patterns from large data using decompositions of data matrix by a sampled subset of exemplars has found many applications in machine learning. We propose a computationally efficient algorithm for adaptive exemplar sampling, called fast exemplar selection (FES). The proposed algorithm can be seen as an efficient variant of the oASIS algorithm [1]. FES iteratively selects incoherent exemplars based on the exemplars that are already sampled. This is done by ensuring that the selected exemplars forms a positive definite Gram matrix which is checked by exploiting its Cholesky factorization in an incremental manner. FES is a deterministic rank revealing algorithm delivering a tighter matrix approximation bound. Further, FES can also be used to exactly represent low rank matrices and signals sampled from a unions of independent subspaces. Experimental results show that FES performs comparable to existing methods for tasks such as matrix approximation, feature selection, outlier detection, and clustering. Vinayak Abrol, Pulkit Sharma, Anil Kumar Sao |
ICASSP | 1 |
| 2017 | Deep-Sparse-Representation-Based Features for Speech RecognitionabstractFeatures derived using sparse representation (SR)-based approaches have been shown to yield promising results for speech recognition tasks. In most of the approaches, the SR corresponding to speech signal is estimated using a dictionary, which could be either exemplar based or learned. However, a single-level decomposition may not be suitable for the speech signal, as it contains complex hierarchical information about various hidden attributes. In this paper, we propose to use a multilevel decomposition (having multiple layers), also known as the deep sparse representation (DSR), to derive a feature representation for speech recognition. Instead of having a series of sparse layers, the proposed framework employs a dense layer between two sparse layers, which helps in efficient implementation. Our studies reveal that the representations obtained at different sparse layers of the proposed DSR model have complimentary information. Thus, the final feature representation is derived after concatenating the representations obtained at the sparse layers. This results in a more discriminative representation, and improves the speech recognition performance. Since the concatenation results in a high-dimensional feature, principal component analysis is used to reduce the dimension of the obtained feature. Experimental studies demonstrate that the proposed feature outperforms existing features for various speech recognition tasks. Pulkit Sharma, Vinayak Abrol, Anil Kumar Sao |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Fast and robust FMRI unmixing using hierarchical dictionary learningabstractWe propose a novel computationally efficient hierarchical dictionary learning (HDL) approach for data-driven unmixing and functional connectivity analysis of functional magnetic resonance imaging (fMRI) data. It is shown that by simultaneously exploiting the sparsity of the spatial brain maps and the incoherence among their evolution in time or task functions, one can achieve better performance while overcoming the drawbacks of existing approaches. The task functions constituting the dictionary, are learned using a hierarchical subset selection approach. Here, to enforce incoherence among atoms, any new atom is selected from suitable training candidates if it does not lie in the column span of past selected atoms. Also, since the sparsity of spatial maps is generally unknown and affected due to acquisition artifacts, HDL doesn't make use of an implicit sparse coding stage while dictionary update. This makes HDL a very fast and efficient data-driven approach for fMRI analysis. Experimental results on synthetic and real fMRI datasets provide compelling evidences that HDL performs better than existing state-of-the-art methods. Vinayak Abrol, Pulkit Sharma, Shahrooz Faghih Roohi, Anil Kumar Sao, Ashraf A. Kassim |
ICIP | 1 |
| 2016 | Greedy dictionary learning for kernel sparse representation based classifier
Vinayak Abrol, Pulkit Sharma, Anil Kumar Sao |
Pattern Recognit. Lett. | 1 |
| 2016 | Greedy double sparse dictionary learning for sparse representation of speech signals
Vinayak Abrol, Pulkit Sharma, Anil Kumar Sao |
Speech Commun. | 1 |
| 2015 | Sparse coding based features for speech units classificationabstractAbstract In this work, we propose sparse representation based features for speech units classification tasks. In order to effectively capture the variations in a speech unit, the proposed method employs multiple class specific dictionaries. Here, the training data belonging to each class is clustered into multiple clusters, and a principal component analysis (PCA) based dictionary is learnt for each cluster. It has been observed that coefficients corresponding to middle principal components can effectively discriminate among different speech units. Exploiting this observation, we propose to use a transformation function known as weighted decomposition (WD) of principal components, which is used to emphasize the discriminative information present in the PCA-based dictionary. In this paper, both raw speech samples and mel frequency cepstral coefficients (MFCC) are used as an initial representation for feature extraction. For comparison, various popular dictionary learning techniques such as K-singular value decomposition (KSVD), simultaneous codeword optimization (SimCO) and greedy adaptive dictionary (GAD) are also employed in the proposed framework. The effectiveness of the proposed features is demonstrated using continuous density hidden Markov model (CDHMM) based classifiers for (i) classification of isolated utterances of E-set of English alphabet, (ii) classification of consonant-vowel (CV) segments in Hindi language and (iii) classification of phoneme from TIMIT phonetic corpus. Pulkit Sharma, Vinayak Abrol, Aroor Dinesh Dileep, Anil Kumar Sao |
INTERSPEECH | 2 |
| 2015 | Voiced/nonvoiced detection in compressively sensed speech signals
Vinayak Abrol, Pulkit Sharma, Anil Kumar Sao |
Speech Commun. | 1 |
| 2013 | Speech enhancement using compressed sensing
Vinayak Abrol, Pulkit Sharma, Anil Kumar Sao |
INTERSPEECH | 1 |