Vipul Arora 0001

dblp:43/521 · DBLP profile ↗
← Back
29ranked-venue papers
8as first author
17since 2021 · last 2025
0000-0002-1207-1258ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 BEST-STD: Bidirectional Mamba-Enhanced Speech Tokenization for Spoken Term Detection
abstract
Query-by-example spoken term detection (QbE-STD) is often hindered by reliance on frame-level features and the computationally intensive DTW-based template matching, limiting its practicality. To address these challenges, we propose a novel approach that encodes speech into discrete, speaker-agnostic semantic tokens. This facilitates fast retrieval using text-based search algorithms and effectively handles out-of-vocabulary terms. Our approach focuses on generating consistent token sequences across varying utterances of the same term. We also propose a bidirectional state space modeling within the Mamba encoder, trained in a self-supervised learning framework, to learn contextual frame-level features that are further encoded into discrete tokens. Our analysis shows that our speech tokens exhibit greater speaker invariance than those from existing tokenizers, making them more suitable for QbESTD tasks. Empirical evaluation on LibriSpeech and TIMIT databases indicates that our method outperforms existing baselines while being more efficient.
Anup Singh, Kris Demuynck, Vipul Arora 0001
ICASSP3
2025 ASR Confidence Estimation using True Class Lexical Similarity Score
abstract
sponsorship: We thank Prasar Bharati for funding this project. We acknowledge the assistance of the project's stakeholders for their consistent support and Param Siddhi for providing compute infrastructure. We also recognise the data annotation team members - Divyanshu Tripathi, Anika Kumari, Ankita Bhattacharya, Shruthi Mishra, Sachindananda Prajapathi and Shivnarayan Pandey for their meticulous effort in preparing the data. (Prasar Bharati)
Nagarathna Ravi, Thishyan Raj T, Ravi Teja Chaganti, Vipul Arora 0001
INTERSPEECH4
2025 H-QuEST: Accelerating Query-by-Example Spoken Term Detection with Hierarchical Indexing
abstract
sponsorship: This work was supported by a research grant from MeitY, Govt. of India, under the project BHASHINI. (MeitY, Govt. of India, under the project BHASHINI)
Yi-Ping Phoebe Chen, Vipul Arora 0001
INTERSPEECH3
2025 Language-Agnostic Speech Tokenizer for Spoken Term Detection with Efficient Retrieval
abstract
The surge in multilingual and code-switched spoken content demands efficient Query-by-Example Spoken Term Detection (STD) systems capable of handling diverse languages. Existing STD systems are monolingual; they typically require large labeled datasets for training and use costly DTW-based matching during inference, limiting their practicality. This paper proposes a novel speech tokenizer that converts speech into language-agnostic tokens. Furthermore, a multi-stage search algorithm enables fast and efficient retrieval from large datasets. In experimental evaluations, the tokens from the proposed tokenizer demonstrate strong speaker invariance, consistent performance across languages, and a capability to generalize effectively to unseen languages, outperforming the baselines significantly.
Anup Singh, Kris Demuynck, Vipul Arora 0001
INTERSPEECH3
2025 Leveraging Unsupervised Data and Domain Adaptation for Deep Regression in Low-Cost Sensor Calibration
abstract
Air quality monitoring is becoming an essential task with rising awareness about air quality. Low-cost air quality sensors are easy to deploy but are not as reliable as the costly and bulky reference monitors. The low-quality sensors can be calibrated against the reference monitors with the help of deep learning. In this article, we translate the task of sensor calibration into a semi-supervised domain adaptation problem and propose a novel solution for the same. The problem is challenging, because it is a regression problem with a covariate shift and label gap. We use histogram loss instead of mean-squared or mean absolute error (MAE), which is commonly used for regression, and find it useful against covariate shift. To handle the label gap, we propose the weighting of samples for adversarial entropy optimization. In experimental evaluations, the proposed scheme outperforms many competitive baselines, which are based on semi-supervised and supervised domain adaptation, in terms of $R^{2}$ score and MAE. Ablation studies show the relevance of each proposed component in the entire scheme.
Swapnil Dey, Vipul Arora 0001, Sachchida N. Tripathi
IEEE Trans. Neural Networks Learn. Syst.2
2024 Learning Ontology Informed Representations with Constraints for Acoustic Event Detection
abstract
Acoustic Event Detection (AED) has been of great interest for nearly a decade for diverse applications. Most open datasets contain meta information on the hierarchy of labels, which can be utilized for building robust AED systems. Our study aims at injecting this domain knowledge by enforcing ontology-informed constraints upon the output space. We show that constrained optimization allows a network to confuse less among the child classes and can back off to parent classes when not confident enough. We perform several experiments on different datasets signifying the robustness of the method. The experiments substantiate that the state of the art baselines do not follow ontology constraints, and perform poorer than the proposed method.
Akshay Raina, Sayeedul Islam Sheikh, Vipul Arora 0001
ICASSP3
2024 Adaptive Refiner Based Few-Shot Font Generation
Pratikhya Ranjit, Jon von Gillern, Venkat Yetrintala, Vipul Arora 0001
ICPR (6)5
2024 AudioNet: Supervised Deep Hashing for Retrieval of Similar Audio Events
abstract
This work presents a supervised deep hashing method for retrieving similar audio events. The proposed method, named AudioNet, is a deep-learning-based system for efficient hashing and retrieval of similar audio events using an audio example as a query. AudioNet achieves high retrieval performance on multiple standard datasets by generating binary hash codes for similar audio events, setting new benchmarks in the field, and highlighting its efficacy and effectiveness compare to other hashing methods. Through comprehensive experiments on standard datasets, our research represents a pioneering effort in evaluating the retrieval performance of similar audio events. A novel loss function is proposed which incorporates weighted contrastive and weighted pairwise loss along with hashcode balancing to improve the efficiency of audio event retrieval. The method adopts discrete gradient propagation, which allows gradients to be propagated through discrete variables during backpropagation. This enables the network to optimize the discrete hash codes using standard gradient-based optimization algorithms, which are typically used for continuous variables. The proposed method showcases promising retrieval performance, as evidenced by the experimental results, even when dealing with imbalanced datasets. The systematic analysis conducted in this study further supports the significant benefits of the proposed method in retrieval performance across multiple datasets. The findings presented in this work establish a baseline for future studies on the efficient retrieval of similar audio events using deep audio embeddings.
Sagar Dutta, Vipul Arora 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 TeLeS: Temporal Lexeme Similarity Score to Estimate Confidence in End-to-End ASR
abstract
Confidence estimation of predictions from an End-to-End (E2E) Automatic Speech Recognition (ASR) model benefitsASR’s downstream and upstream tasks. Class-probability-based confidence scores do not accurately represent the quality of overconfidentASRpredictions. An ancillary Confidence Estimation Model (CEM) calibrates the predictions. State-of-the-art (SOTA) solutions use binary target scores forCEMtraining. However, the binary labels do not reveal the granular information of predicted words, such as temporal alignment between reference and hypothesis and whether the predicted word is entirely incorrect or contains spelling errors. Addressing this issue, we propose a novelTemporal-LexemeSimilarity (TeLeS) confidence score to trainCEM. To address the data imbalance of target scores while trainingCEM, we use shrinkage loss to focus on hard-to-learn data points and minimise the impact of easily learned data points. We conduct experiments withASRmodels trained in three languages, namely Hindi, Tamil, and Kannada, with varying training data sizes. Experiments show thatTeLeSgeneralises well across domains. To demonstrate the applicability of the proposed method, we formulate a TeLeS-based Acquisition (TeLeS-A) function for sampling uncertainty in active learning.TeLeS-Aachieves a significant reduction in the Word Error Rate (WER) compared toSOTAmethods.
Nagarathna Ravi, Thishyan Raj T, Vipul Arora 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2024 Interactive Singing Melody Extraction Based on Active Adaptation
abstract
Extraction of predominant pitch from polyphonic audio is one of the fundamental tasks in the field of music information retrieval and computational musicology. To accomplish this task using machine learning, a large amount of labeled audio data is required to train the model. Moreover, a classical model pre-trained on data from one domain (source), e.g., songs of a particular singer or genre, may not perform comparatively well in extracting melody from songs of a different singer or genre in other domains (target). The performance of such models can be boosted by adapting the model using very little annotated data from the target domain. In this work, we propose an efficient interactive melody adaptation method. Our method selects the regions in the target audio that require human annotation using a confidence criterion based on normalized true class probability. The annotations are used by the model to adapt itself to the target domain using meta-learning. Our method also provides a novel meta-learning approach that handles class imbalance, i.e., a few representative samples from a few classes are available for adaptation in the target domain. Experimental results show that the proposed method outperforms other adaptive melody extraction baselines. The proposed method is model-agnostic and hence can be applied to other non-adaptive melody extraction models to boost their performance. Also, we released a Hindustani Alankaar and Raga (HAR) dataset containing 523 audio files of about 6.86 hours of duration intended for singing melody extraction tasks.
Kavya Ranjan Saxena, Vipul Arora 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 FlowHash: Accelerating Audio Search With Balanced Hashing via Normalizing Flow
abstract
Nearest neighbor search on context representation vectors is a formidable task due to challenges posed by high dimensionality, scalability issues, and potential noise within query vectors. Our novel approach leverages normalizing flow within a self-supervised learning framework to effectively tackle these challenges, specifically in the context of audio fingerprinting tasks. Audio fingerprinting systems incorporate two key components: audio encoding and indexing. The existing systems consider these components independently, resulting in suboptimal performance. Our approach optimizes the interplay between these components, facilitating the adaptation of vectors to the indexing structure. Additionally, we distribute vectors in the latent$\mathbb {R}^{K}$space using normalizing flow, resulting in balanced$K$-bit hash codes. This allows indexing vectors using a balanced hash table, where vectors are uniformly distributed across all possible$2^{K}$hash buckets. This significantly accelerates retrieval, achieving speedups of up to 2× and 1.4× compared to the Locality-Sensitive Hashing (LSH) and Product Quantization (PQ), respectively. We empirically demonstrate that our system is scalable, highly effective, and efficient in identifying short audio queries ($\leq$2 s), particularly at high noise and reverberation levels.
Anup Singh, Kris Demuynck, Vipul Arora 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Balanced Deep CCA for Bird Vocalization Detection
abstract
Event detection improves when events are captured by two different modalities rather than just one. But to train detection systems on multiple modalities is challenging, in particular when there is abundance of unlabelled data but limited amounts of labeled data. We develop a novel self-supervised learning technique for multi- modal data that learns (hidden) correlations between simultaneously recorded microphone (sound) signals and accelerometer (body vibration) signals. The key objective of this work is to learn useful embeddings associated with high performance in downstream event detection tasks when labeled data is scarce and the audio events of interest — songbird vocalizations — are sparse. We base our approach on deep canonical correlation analysis (DCCA) that suffers from event sparseness. We overcome the sparseness of positive labels by first learning a data sampling model from the labelled data and by applying DCCA on the output it produces. This method that we term balanced DCCA (b-DCCA) improves the performance of the unsupervised embeddings on the down-stream supervised audio detection task compared to classsical DCCA. Because data labels are frequently imbalanced, our method might be of broad utility in low-resource scenarios.
B. Anshuman, Linus Rüttimann, Richard H. R. Hahnloser, Vipul Arora 0001
ICASSP5
2023 SyncNet: Correlating Objective for Time Delay Estimation in Audio Signals
abstract
This study addresses the task of performing robust and reliable time-delay estimation in signals in noisy and reverberating environments. In contrast to the popular signal processing based methods, this paper proposes to transform the input signals using a deep neural network into another pair of sequences which show high cross correlation at the actual time delay. This is achieved with the help of a novel correlation function based objective function for training the network. The proposed approach is also intrinsically interpretable as it does not lose temporal information. Experimental evaluations are performed for estimating mutual time delays for different types of audio signals such as pulse, speech and musical beats. SyncNet out-performs other classical approaches, such as GCC-PHAT, and some other learning based approaches.
Akshay Raina, Vipul Arora 0001
ICASSP2
2023 Simultaneously Learning Robust Audio Embeddings and Balanced Hash Codes for Query-by-Example
abstract
Audio fingerprinting systems must efficiently and robustly identify query snippets in an extensive database. To this end, state-of-the-art systems use deep learning to generate compact audio fingerprints. These systems deploy indexing methods, which quantize fingerprints to hash codes in an unsupervised manner to expedite the search. However, these methods generate imbalanced hash codes, leading to their suboptimal performance. Therefore, we propose a self-supervised learning framework to compute fingerprints and balanced hash codes in an end-to-end manner to achieve both fast and accurate retrieval performance. We model hash codes as a balanced clustering process, which we regard as an instance of the optimal transport problem. Experimental results indicate that the proposed approach improves retrieval efficiency while preserving high accuracy, particularly at high distortion levels, compared to the competing methods. Moreover, our system is efficient and scalable in computational load and memory storage.
Anup Singh, Kris Demuynck, Vipul Arora 0001
ICASSP3
2023 wav2tok: Deep Sequence Tokenizer for Audio Retrieval
Adhiraj Banerjee, Vipul Arora 0001
ICLR2
2023 Enc-Dec RNN Acoustic Word Embeddings learned via Pairwise Prediction
Adhiraj Banerjee, Vipul Arora 0001
INTERSPEECH2
2022 EEG-Based Drowsiness Detection With Fuzzy Independent Phase-Locking Value Representations Using Lagrangian-Based Deep Neural Networks
abstract
Passive electroencephalogram (EEG) brain–computer interfaces (BCI) have common usage in the area of Driver Drowsiness Detection. The approach presented herein identifies the cognitive state of the user while no mental action is required. Data recorded in EEG-based BCI experiments are generally noisy, nonstationary, and contaminated with artifacts that can deteriorate any analyzer’s performance. Recently, common spatial patterns (CSPs) have been adapted with EEG state-space incorporating spatiospectral optimization using fuzzy time delay (FTD-CSSP). Temporal phase disparity sequence (TPDS) is used to measure synchrony between EEG signals. The output of Linear transforms operating on the TPDS constitute useful features for EEG regression problems. On similar lines, this article proposes spatiospectral optimized fuzzy-independent phase-locking value (SSO-FIPLV) representations (exploiting the spatiospectral information from TPDS) for EEG signals to monitor a user’s cognitive states. Specifically, we analyze changes in EEG synchronization for a car driver as she/he drifts between alert and drowsy states. We use neural networks (NNs) for prediction. This article also proposes a cutting-edge method for training NN using the Euler–Lagrangian formulation. A stability proof is provided for the intended training approach alongside, and the performance is corroborated on the EEG reaction time prediction task, both within and across subjects, using a publicly available dataset. The NN trained by the proposed approach performs better than other competitive approaches in terms of minimizing root-mean-squared error and maximizing correlation coefficient. Channelwise feature importance in terms of average relevance values calculated from NN feature representations is visualized in the form of Topoplots using layerwise relevance propagation for regression.
Tharun Kumar Reddy, Vipul Arora 0001, Vinay Gupta, Rupam Biswas, Laxmidhar Behera
IEEE Trans. Syst. Man Cybern. Syst.2
2020 Fuzzy Divergence Based Analysis for Eeg Drowsiness Detection Brain Computer Interfaces
abstract
EEG signals can be processed and classified into commands for brain-computer interface (BCI). Stable deciphering of EEG is one of the leading challenges in BCI design owing to low signal to noise ratio and non-stationarities. Presence of non-stationarities in the EEG signals significantly perturb the feature distribution thus deteriorating the performance of Brain Computer Interface. Stationary Subspace methods discover subspaces in which data distribution remains steady over time. In this paper, we develop novel spatial filtering based feature extraction methods for dealing with nonstationarity in EEG signals from a drowsiness detection problem (a machine learning regression problem). The proposed method: DivOVR-FuzzyCSP-WS based features clearly outperformed fuzzy CSP based baseline features in terms of both RMSE and CC performance metrics. It is hoped that the proposed feature extraction method based on DivOVR-FuzzyCSP-WS will bring in a lot of interest in researchers working in developing algorithms for signal processing, in general, for BCI regression problems.
Tharun Kumar Reddy, Vipul Arora 0001, Laxmidhar Behera, Yu-Kai Wang, Chin-Teng Lin
FUZZ-IEEE2
2020 Formulating Divergence Framework for Multiclass Motor Imagery EEG Brain Computer Interface
abstract
The ubiquitous presence of non-stationarities in the EEG signals significantly perturb the feature distribution thus deteriorating the performance of Brain Computer Interface. In this work, a novel method is proposed based on Joint Approximate Diagonalization (JAD) to optimize stationarity for multiclass motor imagery Brain Computer Interface (BCI) in an information theoretic framework. Specifically, in the proposed method, we estimate the subspace which optimizes the discriminability between the classes and simultaneously preserve stationarity within the motor imagery classes. We determine the subspace for the proposed approach through optimization using gradient descent on an orthogonal manifold. The performance of the proposed stationarity enforcing algorithm is compared to that of baseline One-Versus-Rest (OVR)-CSP and JAD on publicly available BCI competition IV dataset IIa. Results show that an improvement in average classification accuracies across the subjects over the baseline algorithms and thus essence of alleviating within session non-stationarities.
Satyam Kumar 0001, Tharun Kumar Reddy, Vipul Arora 0001, Laxmidhar Behera
ICASSP3
2019 Deep Embeddings for Rare Audio Event Detection with Imbalanced Data
abstract
In this paper, we present a method to handle data imbalance for classification with neural networks, and apply it to acoustic event detection (AED) problem. The common approach to tackle data imbalance is to use class-weights in the objective function while training. An existing more sophisticated approach is to map the input to clusters in an embedding space, so that learning is locally balanced by incorporating inter-cluster and inter-class margins. On these lines, we propose a method to learn the embedding using a novel objective function, called triple-header cross entropy. Our scheme integrates in a simple way with back-propagation based training, and is computationally more efficient than general hinge-loss based embedding learning schemes. The empirical evaluation results demonstrate the effectiveness of the proposed method for AED with imbalanced training data.
Vipul Arora 0001, Ming Sun 0007, Chao Wang 0018
ICASSP1
2019 Multiclass Fuzzy Time-Delay Common Spatio-Spectral Patterns With Fuzzy Information Theoretic Optimization for EEG-Based Regression Problems in Brain-Computer Interface (BCI)
abstract
Electroencephalogram (EEG) signals are one of the most widely used noninvasive signals in brain-computer interfaces. Large dimensional EEG recordings suffer from poor signal-tonoise ratio. These signals are very much prone to artifacts and noise, so sufficient preprocessing is done on raw EEG signals before using them for classification or regression. Properly selected spatial filters enhance the signal quality and subsequently improve the rate and accuracy of classifiers, but their applicability to solve regression problems is quite an unexplored objective. This paper extends common spatial patterns (CSP) to EEG state space using fuzzy time delay and thereby proposes a novel approach for spatial filtering. The approach also employs a novel fuzzy information theoretic framework for filter selection. Experimental performance on EEG-based reaction time (RT) prediction from a lane-keeping task data from 12 subjects demonstrated that the proposed spatial filters can significantly increase the EEG signal quality. A comparison based on root-mean-squared error (RMSE), mean absolute percentage error (MAPE), and correlation to true responses is made for all the subjects. In comparison to the baseline fuzzy CSP regression one versus rest, the proposed Fuzzy Time-delay Common Spatio-Spectral filters reduced the RMSE on an average by 9.94%, increased the correlation to true RT on an average by 7.38%, and reduced the MAPE by 7.09%.
Tharun Kumar Reddy, Vipul Arora 0001, Laxmidhar Behera, Yu-Kai Wang, Chin-Teng Lin
IEEE Trans. Fuzzy Syst.2
2017 HJB equation based learning scheme for neural networks
abstract
A control theoretic approach is presented in this paper for both batch and instantaneous updates of weights in feed-forward neural networks. The popular Hamilton-Jacobi-Bellman (HJB) equation has been used to generate an optimal weight update law. The main contribution in this paper is that a closed form solutions for both optimal cost and weight update can be achieved for any feed-forward network using HJB equation. The proposed approach has been compared with some of the existing best performing learning algorithms. It is found as expected that the proposed approach is faster in convergence in terms of computational time. Some of the benchmark test data such as 8-bit parity, breast cancer and credit approval, as well as 2D Gabor function have been used to validate our claims.
Vipul Arora 0001, Laxmidhar Behera, Tharun Kumar Reddy, Ajay Pratap Yadav
IJCNN1
2017 Phonological Feature Based Mispronunciation Detection and Diagnosis Using Multi-Task DNNs and Active Learning
abstract
This paper presents a phonological feature based computer aided pronunciation training system for the learners of a new language (L2).Phonological features allow analysing the learners' mispronunciations systematically and rendering the feedback more effectively.The proposed acoustic model consists of a multi-task deep neural network, which uses a shared representation for estimating the phonological features and HMM state probabilities.Moreover, an active learning based scheme is proposed to efficiently deal with the cost of annotation, which is done by expert teachers, by selecting the most informative samples for annotation.Experimental evaluations are carried out for German and Italian native-speakers speaking English.For mispronunciation detection, the proposed feature-based system outperforms conventional GOP measure and classifier based methods, while providing more detailed diagnosis.Evaluations also demonstrate the advantage of active learning based sampling over random sampling.
Vipul Arora 0001, Aditi Lahiri, Henning Reetz
INTERSPEECH1
2016 Attribute based shared hidden layers for cross-language knowledge transfer
abstract
Deep neural network (DNN) acoustic models can be adapted to under-resourced languages by transferring the hidden layers. An analogous transfer problem is popular as few-shot learning to recognise scantily seen objects based on their meaningful attributes. In similar way, this paper proposes a principled way to represent the hidden layers of DNN in terms of attributes shared across languages. The diverse phoneme sets of different languages can be represented in terms of phonological features that are shared by them. The DNN layers estimating these features could then be transferred in a meaningful and reliable way. Here, we evaluate model transfer from English to German, by comparing the proposed method with other popular methods on the task of phoneme recognition. Experimental results support that apart from providing interpretability to the DNN acoustic models, the proposed framework provides efficient means for their speedy adaptation to different languages, even in the face of scanty adaptation data.
Vipul Arora 0001, Aditi Lahiri, Henning Reetz
SLT1
2015 Multiple F0 Estimation and Source Clustering of Polyphonic Music Audio Using PLCA and HMRFs
abstract
Source transcription of pitched polyphonic music entails providing the pitch (F0) values corresponding to each source in a separate channel. This problem is an important step towards many important problems in music and speech processing. It involves 1) estimating the multiple F0 values in each short time frame, and 2) clustering the F0 values into streams corresponding to different sources. We address the problem in an unsupervised way, with only the total number of sources given beforehand. The framework of probabilistic latent component analysis (PLCA) is used to decompose the polyphonic short-time magnitude spectra for multiple F0 estimation and source-specific feature extraction. It is further embedded into the structure of hidden Markov random fields (HMRF) for clustering the F0s into different sources. This clustering is constrained by the cognitive grouping of continuous F0 contours as well as segregation of simultaneous F0s into different source streams. Such constraints are effectively and elegantly modeled by the HMRF's. Simulated annealing varies the degree of constraints for better clustering. The paper also proposes a novel strategy using the trade-off between precision and recall of multiple F0 estimation for better clustering. Evaluations over a variety of datasets show the efficacy of the proposed algorithm and its robustness to the presence of spurious F0s while clustering. It also outperforms a state-of-the-art unsupervised source streaming algorithm in a set of comparative experiments.
Vipul Arora 0001, Laxmidhar Behera
IEEE ACM Trans. Audio Speech Lang. Process.1
2014 Musical Source Clustering and Identification in Polyphonic Audio
abstract
For music transcription or musical source separation, apart from knowing the multi-F0 contours, it is also important to know which F0 has been played by which instrument. This paper focuses on this aspect, i.e. given the polyphonic audio along with its multiple F0 contours, the proposed system clusters them so as to decide `which instrument played when.' For the task of identifying the instrument or singers in the polyphonic audio, there are many supervised methods available. But many times individual source audio is not available for training. To address this problem, this paper proposes novel schemes using semi-supervised as well as unsupervised approach to source clustering. The proposed theoretical framework is based on auditory perception theory and is implemented using various tools like probabilistic latent component analysis and graph clustering, while taking into account various perceptual cues for characterizing a source. Experiments have been carried out over a wide variety of datasets - ranging from vocal to instrumental as well as from synthetic to real world music. The proposed scheme significantly outperforms a state of the art unsupervised scheme, which does not make use of the given F0 contours. The proposed semi-supervised approach also performs better than another semi-supervised scheme, which makes use of the given F0 information, in terms of computations as well as accuracy.
Vipul Arora 0001, Laxmidhar Behera
IEEE ACM Trans. Audio Speech Lang. Process.1
2013 On-Line Melody Extraction From Polyphonic Audio Using Harmonic Cluster Tracking
abstract
Extraction of predominant melody from the musical performances containing various instruments is one of the most challenging task in the field of music information retrieval and computational musicology. This paper presents a novel framework which estimates predominant vocal melody in real-time by tracking various sources with the help of harmonic clusters (combs) and then determining the predominant vocal source by using the harmonic strength of the source. The novel on-line harmonic comb tracking approach complies with both structural as well as temporal constraints simultaneously. It relies upon the strong higher harmonics for robustness against distortion of the first harmonic due to low frequency accompaniments, in contrast to the existing methods which track the pitch values. The predominant vocal source identification depends upon the novel idea of source dependant filtering of recognition score, which allows the algorithm to be implemented on-line. The proposed method, although on-line, is shown to significantly outperform our implementation of a state-of-the-art offline method for vocal melody extraction. Evaluations also show the reduction in octave error and the effectiveness of novel score filtering technique in enhancing the performance.
Vipul Arora 0001, Laxmidhar Behera
IEEE Trans. Speech Audio Process.1
2012 Design of Distribution Independent Noise Filters with Online PDF Estimation
Vipul Arora 0001, Laxmidhar Behera
ICONIP (1)1
2011 EEG denoising with a recurrent quantum neural network for a brain-computer interface
abstract
Brain-computer interface (BCI) technology is a means of communication that allows individuals with severe movement disability to communicate with external assistive devices using the electroencephalogram (EEG) or other brain signals. This paper presents an alternative neural information processing architecture using the Schrödinger wave equation (SWE) for enhancement of the raw EEG signal. The raw EEG signal obtained during the motor imagery (MI) of a BCI user is intrinsically embedded with non-Gaussian noise while the actual signal is still a mystery. The proposed work in the field of recurrent quantum neural network (RQNN) is designed to filter such non-Gaussian noise using an unsupervised learning scheme without making any assumption about the signal type. The proposed learning architecture has been modified to do away with the Hebbian learning associated with the existing RQNN architecture as this learning scheme was found to be unstable for complex signals such as EEG. Besides, this the soliton behaviour of the non-linear SWE was not properly preserved in the existing scheme. The unsupervised learning algorithm proposed in this paper is able to efficiently capture the statistical behaviour of the input signal while making the algorithm robust to parametric sensitivity. This denoised EEG signal is then fed as an input to the feature extractor to obtain the Hjorth features. These features are then used to train a Linear Discriminant Analysis (LDA) classifier. It is shown that the accuracy of the classifier output over the training and the evaluation datasets using the filtered EEG is much higher compared to that using the raw EEG signal. The improvement in classification accuracy computed over nine subjects is found to be statistically significant.
Vaibhav Gandhi, Vipul Arora 0001, Laxmidhar Behera, Girijesh Prasad, Damien Coyle, T. Martin McGinnity
IJCNN2