VLDB 2026 Research / reviewers in the wild / expert
Aroor Dinesh Dileep
dblp:152/1278 · also Dileep Aroor Dinesh
· DBLP profile ↗
32ranked-venue papers
6as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 6 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hierarchical Loss for Bi-Level Classification of Speech into Language and DialectsabstractSpoken dialect identification (DID) is a challenging task due to high similarities between the classes. This becomes further complicated when DID needs to be performed in a multi-lingual environment, especially when languages are from same language-family, as the scope for confusion is significantly high. In this paper, we propose multiple techniques to address this issue. These techniques motivate the DID system to respect the parent-child relationship between the languages and their dialects by performing bi-level classification of speech. Specifically, in our first approach, we train the language identification (LID) and DID models independently and then combine them together such that the output of LID model selects the DID model, which in turn decides the dialect. However, such independent training does not allow the model to learn the relations/similarities between dialects of different languages. To overcome this limitation, we propose to train an end-to-end DID model which is trained using dialects of all languages in the dataset. While such end-to-end training allows the model to learn inter-dialect similarities in a better way, it does not explicitly prevent a dialect from being misclassified into a dialect of a different language. To address this limitation, we propose a novel hierarchical loss, which motivates the model to maintain the parent-child relationship between language and dialects. Specifically, with the help of an auxiliary language classifier and a primary dialect classifier, the hierarchical loss penalizes the model heavily whenever parent-child relationship is not maintained, i.e., when predicted dialect does not belong to the predicted language. Experiments conducted on a set of closely related Indian languages shows that hierarchical loss based training leads to improvement in the performance. Ananya Angra, Muralikrishna H, Aroor Dinesh Dileep, Veena Thenkanidiyoor |
ICASSP | 3 |
| 2025 | A parallel computing approach to CNN-based QbE-STD using kernel-based matching
Manisha Naik Gaonkar, Veena Thenkanidiyoor, Aroor Dinesh Dileep |
J. Supercomput. | 3 |
| 2022 | Sieving Camera Trap Sequences in the Wild
Anoushka Banerjee, Aroor Dinesh Dileep, Arnav Bhavsar |
ICPRAM | 2 |
| 2022 | Spoken language identification in unseen channel conditions using modified within-sample similarity loss
Muralikrishna H, Aroor Dinesh Dileep |
Pattern Recognit. Lett. | 2 |
| 2021 | Spoken Language Identification in Unseen Target Domain Using Within-Sample Similarity LossabstractState-of-the-art spoken language identification (LID) networks are vulnerable to channel-mismatch that occurs due to the differences in the channels used to obtain the training and testing samples. The effect of channel-mismatch is severe when the training dataset contains very limited channel diversity. One way to address channel-mismatch is by learning a channel-invariant representation of the speech using adversarial multi-task learning (AMTL). But, AMTL approach cannot be used when the training samples do not contain the corresponding channel labels. To address this, we propose an auxiliary within-sample similarity loss (WSSL) which encourages the network to suppress the channel-specific contents in the speech. This does not require any channel labels. Specifically, WSSL gives the similarity between a pair of embeddings of same sample obtained by two separate embedding extractors. These embedding extractors are designed to capture similar information about the channel, but dissimilar LID-specific information in the speech. Furthermore, the proposed WSSL improves the noise-robustness of the LID-network by suppressing the background noise in the speech to some extent. We demonstrate the effectiveness of the proposed approach in both seen and unseen channel conditions using a set of datasets having significant channel-mismatch. Muralikrishna H, Shantanu Kapoor, Aroor Dinesh Dileep, Padmanabhan Rajan |
ICASSP | 3 |
| 2021 | An Empirical Study on Machine Learning Models for Potato Leaf Disease Classification using RGB Images
Soma Ghosh, Renu M. Rameshan, Aroor Dinesh Dileep |
ICPRAM | 3 |
| 2021 | Noise-Robust Spoken Language Identification Using Language Relevance Factor Based EmbeddingabstractState-of-the-art systems for spoken language identification (LID) use i-vector or embedding extracted using a deep neural network (DNN) to represent the utterance. These fixed-length representations are obtained without explicitly considering the relevance of individual frame-level feature vectors in deciding the class label. In this paper, we propose a new method to represent the utterance that considers the relevance of the individual frame-level features. The proposed representation can also preserve the locally available LID-specific information in the input features to some extent. To better utilize the local-level information in the new representation, we propose a novel segment-level matching kernel based support vector machine (SVM) classifier. The proposed representation of the utterance based on the relevance of frame-level features improves the robustness of the LID system to different background noise conditions in the speech. The experiments conducted on speech with different background conditions show that the proposed approach performs better than state-of-the-art approaches in noisy speech and performs similarly to the state-of-the-art systems in clean speech condition. Muralikrishna H, Aroor Dinesh Dileep, Padmanabhan Rajan |
SLT | 3 |
| 2021 | Recognition of varying size scene images using semantic analysis of deep activation maps
Aroor Dinesh Dileep, Veena Thenkanidiyoor |
Mach. Vis. Appl. | 2 |
| 2021 | Visual Semantic-Based Representation Learning Using Deep CNNs for Scene RecognitionabstractIn this work, we address the task of scene recognition from image data. A scene is a spatially correlated arrangement of various visual semantic contents also known as concepts, e.g., “chair,” “car,” “sky,” etc. Representation learning using visual semantic content can be regarded as one of the most trivial ideas as it mimics the human behavior of perceiving visual information. Semantic multinomial (SMN) representation is one such representation that captures semantic information using posterior probabilities of concepts. The core part of obtaining SMN representation is the building of concept models. Therefore, it is necessary to have ground-truth (true) concept labels for every concept present in an image. Moreover, manual labeling of concepts is practically not feasible due to the large number of images in the dataset. To address this issue, we propose an approach for generating pseudo-concepts in the absence of true concept labels. We utilize the pre-trained deep CNN-based architectures where activation maps (filter responses) from convolutional layers are considered as initial cues to the pseudo-concepts. The non-significant activation maps are removed using the proposed filter-specific threshold-based approach that leads to the removal of non-prominent concepts from data. Further, we propose a grouping mechanism to group the same pseudo-concepts using subspace modeling of filter responses to achieve a non-redundant representation. Experimental studies show that generated SMN representation using pseudo-concepts achieves comparable results for scene recognition tasks on standard datasets like MIT-67 and SUN-397 even in the absence of true concept labels. Krishan Sharma, Aroor Dinesh Dileep, Veena Thenkanidiyoor |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2020 | Attention-Driven Projections for Soundscape Classification
Dhanunjaya Varma Devalraju, Muralikrishna H, Padmanabhan Rajan, Aroor Dinesh Dileep |
INTERSPEECH | 4 |
| 2020 | Relevance feedback based online learning model for resource bottleneck prediction in cloud servers
Shaifu Gupta, Aroor Dinesh Dileep |
Neurocomputing | 2 |
| 2020 | Online Sparse BLSTM Models for Resource Usage Prediction in Cloud DatacentresabstractReal time resource usage prediction is an important part of resource provisioning in a cloud data centre. As cloud workloads vary dynamically, effective resource provisioning requires prediction of future resource usage trends. The problem is highly complicated because of highly time varying nature of cloud resource workloads. Training the future resource usage prediction models once, using a fixed set of observations is not sufficient to capture the variability in cloud workloads. In this work, we propose to use gradient descent (GD) and Levenberg-Marquardt (LM) adaptation algorithms for dynamic adaptation of resource utilization prediction models. We also propose a novel sparse framework for fast online adaptation of resource usage prediction models. We propose to analyze different algorithms such as ℓ1regularization, ℓ2regularization, optimal brain damage (OBD), optimal brain surgeon (OBS) for introducing sparsity. The proposed sparse framework for online adaptation of multivariate resource usage prediction models is validated for CPU usage prediction in the Google cluster trace and PlanetLab workload trace. A comparative analysis of different sparse frameworks shows that OBD-based LM adaptation algorithm performs better than other frameworks for online multivariate resource usage prediction in a cloud. Shaifu Gupta, Aroor Dinesh Dileep, Timothy A. Gonsalves |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2019 | Spoken Language Identification Using Bidirectional LSTM Based LID Sequential SenonesabstractThe effectiveness of features used to represent speech utterances influences the performance of spoken language identification (LID) systems. Recent LID systems use bottleneck features (BNFs) obtained from deep neural networks (DNNs) to represent the utterances. These BNFs do not encode language-specific features. The recent advances in DNNs have led to the usage of effective language-sensitive features such as LID-senones, obtained using convolutional neural network (CNN) based architecture. In this work, we propose a novel approach to obtain LID-senones. The proposed approach combines BNF with bidirectional long short-term memory (BLSTM) networks to generate LID-senones. Since each LID-senones preserve sequence information, we term it as LID-sequential-senones (LID-seq-senones). The proposed LID-seq-senones are then used for LID in two ways. In the first approach, we propose to build an end-to-end structure with BLSTM as front end LID-seq-senones extractor followed by a fully connected classification layer. In the second approach, we consider each utterance as a sequence of LID-seq-senones and propose to use support vector machine (SVM) with sequence kernel (GMM-based segment level pyramid match kernel) to classify the utterance. The effectiveness of proposed representation is evaluated on Oregon graduate institute multi-language telephone speech corpus (OGI-TS) and IIT Madras Indian language corpus (IITM-IL). Muralikrishna H, Pulkit Sapra, Anuksha Jain, Aroor Dinesh Dileep |
ASRU | 4 |
| 2018 | Scene Image Classification Using Reduced Virtual Feature Representation in Sparse FrameworkabstractIn this paper, we address the task of scene image classification in sparse framework. Recent scene image datasets consist of thousands of different size images with size of the order of 106pixels. Motivated by the fact that every image has a different size, we propose a dynamic kernel1which works over set of feature maps obtained for an image from last convolutional pooling layer of a pre-trained CNN. The size of feature maps depends on the input image size leading to the requirement of a dynamic kernel to compute similarity score between feature maps of different images. The kernel matrix obtained by using a dynamic kernel is large in size owing to the large number of training examples. To handle this we propose to use the concept of reduced virtual features (RVFs) obtained by diagonalizing the kernel matrix. RVF is a fixed length representation of a scene image irrespective of its true size. Classification is done in sparse framework by applying block sparsity constraint over sparse coefficients using dictionary built from RVFs. The proposed approach tested over standard datasets like Vogel-Schiele, MIT-8, MIT-67 and SUN-397 yields good results. Krishan Sharma, Aroor Dinesh Dileep, Renu M. Rameshan |
ICASSP | 3 |
| 2018 | Deep Spatial Pyramid Match Kernel for Scene Classification
Deepak Kumar Pradhan, Aroor Dinesh Dileep, Veena Thenkanidiyoor |
ICPRAM | 3 |
| 2018 | Modified Time Flexible Kernel for Video Activity Recognition using Support Vector Machines
Apurv Kumar, Sony Allappa, Veena Thenkanidiyoor, Aroor Dinesh Dileep |
ICPRAM | 5 |
| 2018 | Association Learning based Hybrid Model for Cloud Workload PredictionabstractCloud environment provides on-demand access to a shared pool of computing resources over the Internet. Failures are unavoidable in such a distributed and complex environment. In this work, we simulate such a scenario in a docker based virtual environment to aid a proactive approach for anomaly identification in a cloud environment. Proactive approach involves resource prediction first and then anomaly detection. This paper focuses only on resource prediction. We also propose a hybrid model of LSTM and BLSTM using association learning that captures the relationship between the related resource metrics to predict future resource workload in cloud. We use a mix of different types of workloads for simulating the workloads in a cloud environment. The proposed approach is validated on the collected trace of data in a docker based virtual environment as well as the Google cluster trace. It is observed that the proposed model works better as compared to the other state-of-the-art models for resource workload prediction. Siddhant Kumar, Neha Muthiyan, Shaifu Gupta, Aroor Dinesh Dileep, Aditya Nigam |
IJCNN | 4 |
| 2018 | A Context-aware Convolutional Natural Language Generation model for Dialogue SystemsabstractNatural language generation (NLG) is an important component in spoken dialog systems (SDSs).A model for NLG involves sequence to sequence learning.State-of-the-art NLG models are built using recurrent neural network (RNN) based sequence to sequence models (Dušek and Jurcicek, 2016a).Convolutional sequence to sequence based models have been used in the domain of machine translation but their application as natural language generators in dialogue systems is still unexplored.In this work, we propose a novel approach to NLG using convolutional neural network (CNN) based sequence to sequence learning.CNN-based approach allows to build a hierarchical model which encapsulates dependencies between words via shorter path unlike RNNs.In contrast to recurrent models, convolutional approach allows for efficient utilization of computational resources by parallelizing computations over all elements, and eases the learning process by applying constant number of nonlinearities.We also propose to use CNN-based reranker for obtaining responses having semantic correspondence with input dialogue acts.The proposed model is capable of entrainment.Studies using a standard dataset shows the effectiveness of the proposed CNN-based approach to NLG. Sourab Mangrulkar, Suhani Shrivastava, Veena Thenkanidiyoor, Aroor Dinesh Dileep |
SIGDIAL Conference | 4 |
| 2018 | Sparse coding based features for speech units classification
Pulkit Sharma, Vinayak Abrol, Aroor Dinesh Dileep, Anil Kumar Sao |
Comput. Speech Lang. | 3 |
| 2018 | A joint feature selection framework for multivariate resource usage prediction in cloud servers using stability and prediction performance
Shaifu Gupta, Aroor Dinesh Dileep, Timothy A. Gonsalves |
J. Supercomput. | 2 |
| 2017 | Text Classification Using Hierarchical Sparse Representation ClassifiersabstractIn this paper, we propose to use sparse representation classifier (SRC) for text classification. The sparse representation of an example is obtained by using an overcomplete dictionary made up of term frequency (TF) vectors corresponding to all the training documents. We propose to seed the dictionary using principal components of TF vector representation corresponding to training text documents. In this work, we also propose 2-level hierarchical SRC (HSRC) by exploiting the similarity among the classes. We propose to use weighted decomposition principal component analysis (WDPCA) in the second level of HSRC to seed the dictionary to discriminate the similar classes. The effectiveness of the proposed approach to build HSRC for text classification is demonstrated on 20 Newsgroup Corpus. Aroor Dinesh Dileep, Veena Thenkanidiyoor |
ICMLA | 2 |
| 2016 | Bird Call Identification Using Dynamic Kernel Based Support Vector Machines and Deep Neural NetworksabstractIn this paper, we apply speech and audio processing techniques to bird vocalizations and for the classification of birds found in the lower Himalayan regions. Mel frequency cepstral coefficients (MFCC) are extracted from each recording. As a result, the recordings are now represented as varying length sets of feature vectors. Dynamic kernel based support vector machines (SVMs) and deep neural networks (DNNs) are popularly used for the classification of such varying length patterns obtained from speech signals. In this work, we propose to use dynamic kernel based SVMs and DNNs for classification of bird calls represented as sets of feature vectors. Results of our studies show that both approaches give comparable performance. Deep Chakraborty, Paawan Mukker, Padmanabhan Rajan, Aroor Dinesh Dileep |
ICMLA | 4 |
| 2016 | Segment-Level Probabilistic Sequence Kernel Based Support Vector Machines for Classification of Varying Length Patterns of Speech
Veena Thenkanidiyoor, Aroor Dinesh Dileep |
ICONIP (4) | 3 |
| 2015 | Example-Specific Density Based Matching Kernels for Scene Classification Using Support Vector MachinesabstractIn this paper, we propose the example-specific density based matching kernel (ESDMK) for classification of scene images represented as sets of local feature vectors. The proposed kernel is computed between the pair of examples, represented as sets of local feature vectors, by matching the estimates of example-specific densities computed at every local feature vector in those two examples. In this work, the number of local feature vectors of an example among the K nearest neighbors of a local feature vector is considered as an estimate of the example-specific density. The minimum of the two example-specific densities, one for each example, at a local feature vector is considered as the matching score. The ESDMK is then computed as the sum of the matching score computed at every local feature vector in a pair of examples. We also propose the spatial ESDMK (SESDMK) to include spatial information present in the scene images while matching the pair of scene images. Each of the scene images is divided spatially into a fixed number of regions. Then the SESDMK is computed as a combination of region specific ESDMKs that match the corresponding regions. We study the performance of the support vector machine (SVM) based classifiers using the proposed ESDMKs for scene classification and compare with that of the SVM-based classifiers using the state-of-the-art kernels for sets of local feature vectors. Abhijeet Sachdev, Veena Thenkanidiyoor, Aroor Dinesh Dileep, Chellu Chandra Sekhar |
ICMLA | 3 |
| 2015 | Example-Specific Density Based Matching Kernel for Classification of Varying Length Patterns of Speech Using Support Vector Machines
Abhijeet Sachdev, Aroor Dinesh Dileep, Veena Thenkanidiyoor |
ICONIP (1) | 2 |
| 2015 | Sparse coding based features for speech units classificationabstractAbstract In this work, we propose sparse representation based features for speech units classification tasks. In order to effectively capture the variations in a speech unit, the proposed method employs multiple class specific dictionaries. Here, the training data belonging to each class is clustered into multiple clusters, and a principal component analysis (PCA) based dictionary is learnt for each cluster. It has been observed that coefficients corresponding to middle principal components can effectively discriminate among different speech units. Exploiting this observation, we propose to use a transformation function known as weighted decomposition (WD) of principal components, which is used to emphasize the discriminative information present in the PCA-based dictionary. In this paper, both raw speech samples and mel frequency cepstral coefficients (MFCC) are used as an initial representation for feature extraction. For comparison, various popular dictionary learning techniques such as K-singular value decomposition (KSVD), simultaneous codeword optimization (SimCO) and greedy adaptive dictionary (GAD) are also employed in the proposed framework. The effectiveness of the proposed features is demonstrated using continuous density hidden Markov model (CDHMM) based classifiers for (i) classification of isolated utterances of E-set of English alphabet, (ii) classification of consonant-vowel (CV) segments in Hindi language and (iii) classification of phoneme from TIMIT phonetic corpus. Pulkit Sharma, Vinayak Abrol, Aroor Dinesh Dileep, Anil Kumar Sao |
INTERSPEECH | 3 |
| 2014 | Class-specific GMM based intermediate matching kernel for classification of varying length patterns of long duration speech using support vector machines
Aroor Dinesh Dileep, Chellu Chandra Sekhar |
Speech Commun. | 1 |
| 2014 | GMM-Based Intermediate Matching Kernel for Classification of Varying Length Patterns of Long Duration Speech Using Support Vector MachinesabstractDynamic kernel (DK)-based support vector machines are used for the classification of varying length patterns. This paper explores the use of intermediate matching kernel (IMK) as a DK for classification of varying length patterns of long duration speech represented as sets of feature vectors. The main issue in construction of IMK is the choice for the set of virtual feature vectors used to select the local feature vectors for matching. This paper proposes to use components of class-independent Gaussian mixture model (CIGMM) as a representation for the set of virtual feature vectors. For every component of CIGMM, a local feature vector each from the two sets of local feature vectors that has the highest probability of belonging to that component is selected and a base kernel is computed between the selected local feature vectors. The IMK is computed as the sum of all the base kernels corresponding to different components of CIGMM. It is proposed to use the responsibility term weighted base kernels in computation of IMK to improve its discrimination ability. This paper also proposes the posterior probability weighted DKs (including the proposed IMKs) to improve their classification performance and reduce the number of support vectors. The performance of the support vector machine (SVM)-based classifiers using the proposed IMKs is studied for speech emotion recognition and speaker identification tasks and compared with that of the SVM-based classifiers using the state-of-the-art DKs. Aroor Dinesh Dileep, Chellu Chandra Sekhar |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2013 | HMM based pyramid match kernel for classification of sequential patterns of speech using support vector machinesabstractClassification of varying length sequences using support vector machine (SVM) requires a suitable kernel that measures the similarity between a pair of sequences. In this paper we propose a novel approach to design a pyramid match kernel (PMK) using hidden Markov model. We study the performance of the SVM-based classifiers using the proposed PMK for recognition of isolated utterances of E-set in English alphabet and recognition of consonant-vowel segments of speech in Hindi and compare with that of the SVM-based classifiers using score-space kernels and alignments kernels. Aroor Dinesh Dileep, Chellu Chandra Sekhar |
ICASSP | 1 |
| 2013 | HMM Based Intermediate Matching Kernel for Classification of Sequential Patterns of Speech Using Support Vector MachinesabstractIn this paper, we address the issues in the design of an intermediate matching kernel (IMK) for classification of sequential patterns using support vector machine (SVM) based classifier for tasks such as speech recognition. Specifically, we address the issues in constructing a kernel for matching sequences of feature vectors extracted from the speech signal data of utterances. The codebook based IMK and Gaussian mixture model (GMM) based IMK have been proposed earlier for matching the varying length patterns represented as sets of features vectors for tasks such as image classification and speaker recognition. These methods consider the centers of clusters and the components of GMM as the virtual feature vectors used in the design of IMK. As these methods do not use sequence information in matching the patterns, these methods are not suitable for matching sequential patterns. We propose the hidden Markov model (HMM) based IMK for matching sequential patterns of varying length. We consider two approaches to design the HMM-based IMK. In the first approach, each of the two sequences to be matched is segmented into subsequences with each subsequence aligned to a state of the HMM. Then the HMM-based IMK is constructed as a combination of state-specific GMM-based IMKs that match the subsequences aligned with the particular states of the HMM. In the second approach, the HMM-based IMK is constructed without segmenting sequences, and by matching the local feature vectors selected using the responsibility terms that account for being in a state and generating the feature vectors by a component of the GMM of that state. We study the performance of the SVM based classifiers using the proposed HMM-based IMK for recognition of isolated utterances of E-set in English alphabet and recognition of consonent–vowel segments in Hindi language. Aroor Dinesh Dileep, Chellu Chandra Sekhar |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2009 | Representation and feature selection using multiple kernel learningabstractMultiple kernel learning (MKL) approach for selecting and combining different representations of a data is presented. Selection of features from a representation of data using the MKL approach is also addressed. A base kernel function is used for each representation as well as for each feature from a representation. A new kernel is obtained as a linear combination of base kernels, weighted according to the relevance of representation or feature. The MKL approach helps to select and combine the representations as well as to select features from a representation. Issues in the MKL algorithm are addressed in the framework of support vector machines (SVM). Studies on the representation and feature selection are presented for an image categorization task. Aroor Dinesh Dileep, Chellu Chandra Sekhar |
IJCNN | 1 |
| 2006 | Identification of Block Ciphers using Support Vector MachinesabstractIn this paper, we propose an approach for identification of encryption method for block ciphers using support vector machines. The task of identification of encryption method from cipher text only is considered as a document categorization task. We address the issues in representing a cipher text by a document vector. We consider the common dictionary based method and the class specific dictionary based method for generating a document vector from a cipher text. As the dimension of document vector is large, support vector machines based classifiers are considered for identification of encryption method. We present the performance of the proposed approach for cipher texts generated using five block ciphers. Aroor Dinesh Dileep, Chellu Chandra Sekhar |
IJCNN | 1 |