VLDB 2026 Research / reviewers in the wild / expert
Chellu Chandra Sekhar
dblp:15/1757 · also C. Chandra Sekhar
· DBLP profile ↗
56ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0001-5551-4883ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DenseCapVCR: Utilizing Dense Captions for Visual Commonsense ReasoningabstractVisual Commonsense Reasoning (VCR) is a challenging task in the domain of visual cognition. Extending the Visual Question Answering (VQA) task, where models select only correct answers for a given image and a question, VCR involves not only answer selection but also identification of an appropriate rationale supporting the chosen answer. Existing VCR models rely mainly on visual cues, which can be insufficient for complete image understanding. To address this, we propose using dense captions as additional semantic information to enhance the VCR model’s reasoning capabilities. We introduce two approaches: an attention-based integration method to efficiently incorporate dense captions into existing VCR models, and a contrastive learning-based method that optimizes both cross-entropy and contrastive losses. The contrastive loss ensures robust representations from the generated dense captions, while the cross-entropy loss maintains discriminative power for response prediction. Experiments on the benchmark VCR dataset demonstrate that both methods improve reasoning performance compared to state-of-the-art models, confirming the effectiveness of integrating dense captions. Subham Das, Chellu Chandra Sekhar |
IJCNN | 2 |
| 2025 | Uncertainty-Guided Metric Learning Without LabelsabstractUnsupervised metric learning aims to learn the discriminative representations by grouping similar examples in the absence of labels. Many unsupervised metric learning algorithms combine clustering-based pseudo-label generation with embedding fine-tuning. However, pseudo-labels can be unreliable and noisy. This could affect metric learning and degrade the quality of the learned representations. In this work, we propose an approach to reduce the negative effect of label noise on learning discriminative embeddings by using context and prediction uncertainty. In particular, we refine the pseudo-labels by aggregating information from neighbors. We propose a function to weigh the pairs, leveraging their prediction confidence and uncertainty. We modify the metric learning loss function to incorporate this weight. Experimental results demonstrate the effectiveness of our proposed method on standard datasets for metric learning. Dhanunjaya Varma Devalraju, Chellu Chandra Sekhar |
WACV | 2 |
| 2024 | Leveraging Generated Image Captions for Visual Commonsense ReasoningabstractVisual Commonsense Reasoning (VCR) involves cognition-level visual understanding by drawing accurate conclusions based on thorough visual understanding. Unlike Visual Question Answering (VQA), where the model merely chooses a correct answer, VCR requires models to not only pick an answer but also identify an appropriate rationale. Traditionally, VCR models have predominantly relied on visual data for their reasoning processes. However, achieving a comprehensive understanding of an image remains a challenging task, as it often requires reasoning beyond visual cues alone. We propose to use the generated image captions to enhance the VCR model’s reasoning capabilities. We propose fusion strategies to integrate the image caption into the VCR model, enabling a better understanding of the image. To evaluate the effectiveness of our proposed approach, we conduct experiments on the benchmark VCR dataset. The results demonstrate that the late fusion strategy enhances the performance of baseline VCR models, yielding an improved accuracy and reasoning capability. Subham Das, Chellu Chandra Sekhar |
ICIP | 2 |
| 2024 | MetaFix: Semi-supervised Model Agnostic Meta-learning Using Consistency Regularization
Solarica Palit, Chellu Chandra Sekhar |
ICONIP (3) | 2 |
| 2023 | Descriptive and Coherent Paragraph Generation for Image Paragraph Captioning Using Vision Transformer and Post-processing
Naveen Vakada, Chellu Chandra Sekhar |
ACIVS | 2 |
| 2023 | Multi-Modal Hierarchical Attention-Based Dense Video CaptioningabstractMost of the existing dense video captioning models use a single modality of features for captioning. A video has a wide variety of information like spatial features, temporal features, audio features, and semantic features. In this paper, we propose a dense video captioning model that captures crossmodal attention between different types of features using an audio-visual attention block in the encoder and a hierarchical attention block in the decoder. The audio-visual attention block applies cross-modal attention between the RGB, flow, and audio features. The hierarchical attention block performs two-level attention between the semantic features and the features from the encoder for generating descriptions. The results show that the proposed approach performs better than the state-of-the-art approaches. Hemalatha Munusamy, Chellu Chandra Sekhar |
ICIP | 2 |
| 2023 | Multimodal attention-based transformer for video captioning
Hemalatha Munusamy, Chellu Chandra Sekhar |
Appl. Intell. | 2 |
| 2022 | Video captioning using Semantically Contextual Generative Adversarial Network
Hemalatha Munusamy, Chellu Chandra Sekhar |
Comput. Vis. Image Underst. | 2 |
| 2021 | Semi-Supervised Metric Learning: A Deep ResurrectionabstractDistance Metric Learning (DML) seeks to learn a discriminative embedding where similar examples are closer, and dissimilar examples are apart. In this paper, we address the problem of Semi-Supervised DML (SSDML) that tries to learn a metric using a few labeled examples, and abundantly available unlabeled examples. SSDML is important because it is infeasible to manually annotate all the examples present in a large dataset. Surprisingly, with the exception of a few classical approaches that learn a linear Mahalanobis metric, SSDML has not been studied in the recent years, and lacks approaches in the deep SSDML scenario. In this paper, we address this challenging problem, and revamp SSDML with respect to deep learning. In particular, we propose a stochastic, graph-based approach that first propagates the affinities between the pairs of examples from labeled data, to that of the unlabeled pairs. The propagated affinities are used to mine triplet based constraints for metric learning. We impose orthogonality constraint on the metric parameters, as it leads to a better performance by avoiding a model collapse. Ujjal Kr Dutta, Mehrtash Harandi, Chellu Chandra Sekhar |
AAAI | 3 |
| 2021 | Novel Architectures for Unsupervised Information Bottleneck Based Speaker Diarization of MeetingsabstractSpeaker diarization is an important problem that is topical, and is especially useful as a preprocessor for conversational speech related applications. The objective of this article is two-fold: (i) segment initialization by uniformly distributing speaker information across the initial segments, and (ii) incorporating speaker discriminative features within the unsupervised diarization framework. In the first part of the work, a varying length segment initialization technique for Information Bottleneck (IB) based speaker diarization system using phoneme rate as the side information is proposed. This initialization distributes speaker information uniformly across the segments and provides a better starting point for IB based clustering. In the second part of the work, we present a Two-Pass Information Bottleneck (TPIB) based speaker diarization system that incorporates speaker discriminative features during the process of diarization. The TPIB based speaker diarization system has shown improvement over the baseline IB based system. During the first pass of the TPIB system, a coarse segmentation is performed using IB based clustering. The alignments obtained are used to generate speaker discriminative features using a shallow feed-forward neural network and linear discriminant analysis. The discriminative features obtained are used in the second pass to obtain the final speaker boundaries. In the final part of the paper, variable segment initialization is combined with the TPIB framework. This leverages the advantages of better segment initialization and speaker discriminative features that results in an additional improvement in performance. An evaluation on standard meeting datasets shows that a significant absolute improvement of 3.9% and 4.7% is obtained on the NIST and AMI datasets, respectively. Nauman Dawalatabad, Srikanth R. Madikeri, Chellu Chandra Sekhar, Hema A. Murthy |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Unsupervised Metric Learning with Synthetic ExamplesabstractDistance Metric Learning (DML) involves learning an embedding that brings similar examples closer while moving away dissimilar ones. Existing DML approaches make use of class labels to generate constraints for metric learning. In this paper, we address the less-studied problem of learning a metric in an unsupervised manner. We do not make use of class labels, but use unlabeled data to generate adversarial, synthetic constraints for learning a metric inducing embedding. Being a measure of uncertainty, we minimize the entropy of a conditional probability to learn the metric. Our stochastic formulation scales well to large datasets, and performs competitive to existing metric learning methods. Ujjal Kr Dutta, Mehrtash Harandi, Chellu Chandra Sekhar |
AAAI | 3 |
| 2020 | A Geometric Approach for Unsupervised Similarity LearningabstractMetric learning groups similar examples together, while moving away dissimilar ones. This is a crucial task in image processing and computer vision. However, existing metric learning approaches require huge number of labeled examples for their success. In this paper, we propose a novel, unsupervised metric learning approach, that learns a similarity metric without making use of class labels. Using a graph-based clustering approach, we form a set of tuples, to provide constraints for metric learning. To efficiently handle high-dimensional data, we learn the metric in a lower dimensional latent space. A confidence function is devised to aid the convergence by appropriately weighting the loss functions. The parameters of our approach are jointly learned using Riemannian optimization. Ujjal Kr Dutta, Chellu Chandra Sekhar |
ICASSP | 2 |
| 2020 | Domain-Specific Semantics Guided Approach to Video CaptioningabstractIn video captioning, the description of a video usually relies on the domain to which the video belongs. Typically, the videos belong to wide range domains such as sports, music, news, cooking, etc. In many cases, a video can be associated with more than one domain. In this paper, we propose an approach to video captioning that uses domain-specific decoders. We build a domain classifier to obtain the estimates of probabilities of a video belonging to different domains. For each video, we identify the top - k domains based on the estimated probabilities. Each video in the training data set is shared in training the domain-specific decoders of top-k labels obtained from the domain classifier. The domain-specific decoders use the domain-specific semantic tags for generating captions. The proposed approach uses the Temporal VLAD for preprocessing the features extracted from 2D-CNN and 3D-CNN features. The preprocessed features provide better feature representation of the videos. The effectiveness of the proposed approach is demonstrated through the results of experimental studies on Microsoft Video Description (MSVD) corpus and MSR-VTT dataset. Hemalatha Munusamy, Chellu Chandra Sekhar |
WACV | 2 |
| 2019 | Incremental Transfer Learning in Two-pass Information Bottleneck Based Speaker Diarization System for MeetingsabstractThe two-pass information bottleneck (TPIB) based speaker diarization system operates independently on different conversational recordings. TPIB system does not consider previously learned speaker discriminative information while di-arizing new conversations. Hence, the real time factor (RTF) of TPIB system is high owing to the training time required for the artificial neural network (ANN). This paper attempts to improve the RTF of the TPIB system using an incremental transfer learning approach where the parameters learned by the ANN from other conversations are updated using current conversation rather than learning parameters from scratch. This reduces the RTF significantly. The effectiveness of the proposed approach compared to the baseline IB and the TPIB systems is demonstrated on standard NIST and AMI conversational meeting datasets. With a minor degradation in performance, the proposed system shows a significant improvement of 33.07% and 24.45% in RTF with respect to TPIB system on the NIST RT-04Eval and AMI-1 datasets, respectively. Nauman Dawalatabad, Srikanth R. Madikeri, Chellu Chandra Sekhar, Hema A. Murthy |
ICASSP | 3 |
| 2019 | Multi-label Classification Models for Detection of Phonetic Features in building Acoustic ModelsabstractAcoustic modeling in large vocabulary continuous speech recognition systems is commonly done by building the models for subword units such as phonemes, syllables or senones. In recent years, various end-to-end systems using acoustic models built at grapheme or phoneme level have also been explored. These systems either require a lot of data and/or heavily rely on the use of language models or pronunciation dictionary for good recognition performance. With the intention of reducing the dependence on data or external models, we have explored the usage of phonetic features in building acoustic models for speech recognition. The phonetic features describe a sound based on the speech production mechanism in humans. Multi-label classification models are built for detection of phonetic features in a given speech signal. The detected phonetic features are used along with the acoustic features as input to models for phoneme identification. The effectiveness of the proposed approach is demonstrated on TIMIT and Wall Street Journal corpora. Performance improvement over other phoneme recognition studies using the phonetic features is obtained. Rupam Ojha, Chellu Chandra Sekhar |
IJCNN | 2 |
| 2018 | Affinity Propagation Based Closed-Form Semi-supervised Metric Learning Framework
Ujjal Kr Dutta, Chellu Chandra Sekhar |
ICANN (1) | 2 |
| 2018 | Subspace Segmentation Based Metric LearningabstractDistance Metric Learning (DML) has been successfully applied in a variety of computer vision and image processing tasks. Laplacian Regularized Metric Learning (LRML) computes a distance metric by satisfying given sets of pairwise similarity and dissimilarity constraints while preserving the topological structure of the given data via a Laplacian regularizer which is dependent on an affinity matrix. This paper addresses the problem of semi-supervised DML using LRML for image data sampled from a union of low-dimensional subspaces by computing the affinity matrix using a self-representation based graph instead of traditional graph used in LRML, resulting in two variants of LRML called as L-NNLRS and L-NLSP. Ujjal Kr Dutta, Chellu Chandra Sekhar |
ICIP | 2 |
| 2018 | Information Bottleneck Based Percussion Instrument Diarization System for Taniavartanam Segments of Carnatic Music Concerts
Nauman Dawalatabad, Jom Kuriakose, Chellu Chandra Sekhar, Hema A. Murthy |
INTERSPEECH | 3 |
| 2018 | Distance metric learning-based kernel gram matrix learning for pattern analysis tasks in kernel feature space
B. S. Shajee Mohan, Chellu Chandra Sekhar |
Pattern Anal. Appl. | 2 |
| 2016 | Two-Pass IB Based Speaker Diarization System Using Meeting-Specific ANN Based FeaturesabstractIn this paper, we present a two-pass Information Bottleneck (IB) based system for speaker diarization which uses meetingspecific artificial neural network (ANN) based features.We first use IB based speaker diarization system to get the labelled speaker segments.These segments are re-segmented using Kullback-Leibler Hidden Markov Model (KL-HMM) based re-segmentation.The multi-layer ANN is then trained to discriminate these speakers using the re-segmented output labels and the spectral features.We then extract the bottleneck features from the trained ANN and perform principal component analysis (PCA) on these features.After performing PCA, these bottleneck features are used along with the different spectral features in the second pass using the same IB based system with KL-HMM re-segmentation.Our experiments on NIST RT and AMI datasets show that the proposed system performs better than the baseline IB system in terms of speaker error rate (SER) with a best case relative improvement of 28.6% amongst AMI datasets and 27.1% on NIST RT04eval dataset. Nauman Dawalatabad, Srikanth R. Madikeri, Chellu Chandra Sekhar, Hema A. Murthy |
INTERSPEECH | 3 |
| 2015 | Example-Specific Density Based Matching Kernels for Scene Classification Using Support Vector MachinesabstractIn this paper, we propose the example-specific density based matching kernel (ESDMK) for classification of scene images represented as sets of local feature vectors. The proposed kernel is computed between the pair of examples, represented as sets of local feature vectors, by matching the estimates of example-specific densities computed at every local feature vector in those two examples. In this work, the number of local feature vectors of an example among the K nearest neighbors of a local feature vector is considered as an estimate of the example-specific density. The minimum of the two example-specific densities, one for each example, at a local feature vector is considered as the matching score. The ESDMK is then computed as the sum of the matching score computed at every local feature vector in a pair of examples. We also propose the spatial ESDMK (SESDMK) to include spatial information present in the scene images while matching the pair of scene images. Each of the scene images is divided spatially into a fixed number of regions. Then the SESDMK is computed as a combination of region specific ESDMKs that match the corresponding regions. We study the performance of the support vector machine (SVM) based classifiers using the proposed ESDMKs for scene classification and compare with that of the SVM-based classifiers using the state-of-the-art kernels for sets of local feature vectors. Abhijeet Sachdev, Veena Thenkanidiyoor, Aroor Dinesh Dileep, Chellu Chandra Sekhar |
ICMLA | 4 |
| 2015 | Automatic Image Annotation Using Convex Deep Learning Models
Niharjyoti Sarangi, Chellu Chandra Sekhar |
ICPRAM (2) | 2 |
| 2014 | A Descriptor based on Intensity Binning for Image MatchingabstractThis paper proposes a method for extracting image
descriptors using intensity binning. It is based on the fact that,
when the intensities of the interest regions are quantized, the
pixels retain their bin labels under common image deformations, up
to a certain degree of perturbation. Consequently, the spatial
configuration and the shape of the connected regions of pixels
belonging to each bin become resilient to noise, which, as a whole,
capture the topography of the intensity map pertaining to that
region. We examine the effect of classical image deformations on this
representation and seek to find a compact yet robust representation
which remains unperturbed in the presence of noise and image
deformations. We use Oxford dataset in our experiments and the
results show that the proposed descriptor gives a better performance
than the existing methods for matching two images under common image
deformations. B. Balasanjeevi, Chellu Chandra Sekhar |
ICPRAM | 2 |
| 2014 | Class-specific GMM based intermediate matching kernel for classification of varying length patterns of long duration speech using support vector machines
Aroor Dinesh Dileep, Chellu Chandra Sekhar |
Speech Commun. | 2 |
| 2014 | GMM-Based Intermediate Matching Kernel for Classification of Varying Length Patterns of Long Duration Speech Using Support Vector MachinesabstractDynamic kernel (DK)-based support vector machines are used for the classification of varying length patterns. This paper explores the use of intermediate matching kernel (IMK) as a DK for classification of varying length patterns of long duration speech represented as sets of feature vectors. The main issue in construction of IMK is the choice for the set of virtual feature vectors used to select the local feature vectors for matching. This paper proposes to use components of class-independent Gaussian mixture model (CIGMM) as a representation for the set of virtual feature vectors. For every component of CIGMM, a local feature vector each from the two sets of local feature vectors that has the highest probability of belonging to that component is selected and a base kernel is computed between the selected local feature vectors. The IMK is computed as the sum of all the base kernels corresponding to different components of CIGMM. It is proposed to use the responsibility term weighted base kernels in computation of IMK to improve its discrimination ability. This paper also proposes the posterior probability weighted DKs (including the proposed IMKs) to improve their classification performance and reduce the number of support vectors. The performance of the support vector machine (SVM)-based classifiers using the proposed IMKs is studied for speech emotion recognition and speaker identification tasks and compared with that of the SVM-based classifiers using the state-of-the-art DKs. Aroor Dinesh Dileep, Chellu Chandra Sekhar |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2013 | HMM based pyramid match kernel for classification of sequential patterns of speech using support vector machinesabstractClassification of varying length sequences using support vector machine (SVM) requires a suitable kernel that measures the similarity between a pair of sequences. In this paper we propose a novel approach to design a pyramid match kernel (PMK) using hidden Markov model. We study the performance of the SVM-based classifiers using the proposed PMK for recognition of isolated utterances of E-set in English alphabet and recognition of consonant-vowel segments of speech in Hindi and compare with that of the SVM-based classifiers using score-space kernels and alignments kernels. Aroor Dinesh Dileep, Chellu Chandra Sekhar |
ICASSP | 2 |
| 2013 | Bayesian mixture of AR models for time series clustering
Venkataramana B. Kini, Chellu Chandra Sekhar |
Pattern Anal. Appl. | 2 |
| 2013 | HMM Based Intermediate Matching Kernel for Classification of Sequential Patterns of Speech Using Support Vector MachinesabstractIn this paper, we address the issues in the design of an intermediate matching kernel (IMK) for classification of sequential patterns using support vector machine (SVM) based classifier for tasks such as speech recognition. Specifically, we address the issues in constructing a kernel for matching sequences of feature vectors extracted from the speech signal data of utterances. The codebook based IMK and Gaussian mixture model (GMM) based IMK have been proposed earlier for matching the varying length patterns represented as sets of features vectors for tasks such as image classification and speaker recognition. These methods consider the centers of clusters and the components of GMM as the virtual feature vectors used in the design of IMK. As these methods do not use sequence information in matching the patterns, these methods are not suitable for matching sequential patterns. We propose the hidden Markov model (HMM) based IMK for matching sequential patterns of varying length. We consider two approaches to design the HMM-based IMK. In the first approach, each of the two sequences to be matched is segmented into subsequences with each subsequence aligned to a state of the HMM. Then the HMM-based IMK is constructed as a combination of state-specific GMM-based IMKs that match the subsequences aligned with the particular states of the HMM. In the second approach, the HMM-based IMK is constructed without segmenting sequences, and by matching the local feature vectors selected using the responsibility terms that account for being in a state and generating the feature vectors by a component of the GMM of that state. We study the performance of the SVM based classifiers using the proposed HMM-based IMK for recognition of isolated utterances of E-set in English alphabet and recognition of consonent–vowel segments in Hindi language. Aroor Dinesh Dileep, Chellu Chandra Sekhar |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2009 | Combination of generative models and SVM based classifier for speech emotion recognitionabstractModeling time series data of varying length is important in different domains. There are two paradigms for modeling the varying length sequential data. Tasks such as speech recognition need modeling the temporal dynamics and the correlations among the features. Hidden Markov models (HMM) are used for these tasks. In tasks such as speaker recognition, audio classification and speech emotion recognition, modeling the temporal dynamics is not critical. Gaussian mixture models (GMM) are commonly used for these tasks. Generative models such as HMMs and GMMs focus on estimating the density of the data and are not suitable for classifying the data of confusable classes. Discriminative classifiers such as support vector machines (SVM) are suitable for the fixed dimensional patterns. In this paper, we propose a hybrid framework where a generative front end is used for representing the varying length time series data and then a discriminative model is used for classification. A score based approach and a segment modeling based approach are proposed in this framework. Both the approaches are applied for speech emotion recognition. The performance is compared with that of an SVM classifier that uses different statistical features and also with that of the GMM classifiers that use maximum likelihood method and the variational Bayes method for parameter estimation. Both the proposed approaches outperform the methods used for comparison. S. Chandrakala 0001, Chellu Chandra Sekhar |
IJCNN | 2 |
| 2009 | Representation and feature selection using multiple kernel learningabstractMultiple kernel learning (MKL) approach for selecting and combining different representations of a data is presented. Selection of features from a representation of data using the MKL approach is also addressed. A base kernel function is used for each representation as well as for each feature from a representation. A new kernel is obtained as a linear combination of base kernels, weighted according to the relevance of representation or feature. The MKL approach helps to select and combine the representations as well as to select features from a representation. Issues in the MKL algorithm are addressed in the framework of support vector machines (SVM). Studies on the representation and feature selection are presented for an image categorization task. Aroor Dinesh Dileep, Chellu Chandra Sekhar |
IJCNN | 2 |
| 2008 | An SVM Based Approach to Cross-Language Adaptation for Indian Languages
A. Vijaya Rama Raju, Chellu Chandra Sekhar |
ICONIP (2) | 2 |
| 2008 | Large margin AR model for time series classificationabstractIn this paper we propose a new method for time series pattern classification. It is based on the generative modeling using Autoregressive(AR) model and optimizing the boundaries between these models using the large margin concepts. The developed model captures the correlations in the time series data. Multi-class classification can be performed directly without performing binary classification. The optimization is performed using genetic algorithm for obtaining global optimal parameters. The developed method is applied on simulated and ECG data and found to perform better than the methods which utilize the AR coefficients as the features for the classification. Venkataramana B. Kini, Chellu Chandra Sekhar |
ICPR | 2 |
| 2008 | A density based method for multivariate time series clustering in kernel feature spaceabstractTime series clustering finds applications in diverse fields of science and technology. Kernel based clustering methods like kernel K-means method need number of clusters as input and cannot handle outliers or noise. In this paper, we propose a density based clustering method in kernel feature space for clustering multivariate time series data of varying length. This method can also be used for clustering any type of structured data, provided a kernel which can handle that kind of data is used. We present heuristic methods to find the initial values of the parameters used in our proposed algorithm. To show the effectiveness of this method, this method is applied to two different online handwritten character data sets which are multivariate time series data of varying length, as a real world application. The performance of the proposed method is compared with the spectral clustering and kernel k-means clustering methods. Besides handling outliers, the proposed method performs as well as the spectral clustering method and outperforms the kernel k-means clustering method. S. Chandrakala 0001, Chellu Chandra Sekhar |
IJCNN | 2 |
| 2008 | Hyperparameters of Gaussian process as features for trajectory classificationabstractIn this paper, we address the trajectory classification problem in Gaussian process framework without using Gaussian process based classification directly. Properties of the function corresponding to a trajectory are captured into the hyperparameters of a Gaussian process. As different trajectories have different properties, hyperparameters are different for these trajectories. In the hyperparametric space, different clusters are formed for noisy, shifted versions of the trajectories. The hyperparameters are used as features representing a trajectory and the classification task is performed in the hyperparametric space. Classification performance of the proposed method is evaluated on simulated data and also on realworld time series data. G. Haranadh, Chellu Chandra Sekhar |
IJCNN | 2 |
| 2007 | Clustering of Nonlinearly Separable Data Using Spiking Neural Networks
Lakshmi Narayana Panuku, Chellu Chandra Sekhar |
ICANN (1) | 2 |
| 2007 | Spatiostructural Features for Recognition of Online Handwritten Characters in Devanagari and Tamil Scripts
H. Swethalakshmi, Chellu Chandra Sekhar, V. Srinivasa Chakravarthy |
ICANN (2) | 2 |
| 2007 | Multi-Scale Kernel Latent Variable Models for Nonlinear Time Series Pattern Matching
Venkataramana B. Kini, Chellu Chandra Sekhar |
ICONIP (2) | 2 |
| 2007 | Region-Based Encoding Method Using Multi-dimensional Gaussians for Networks of Spiking Neurons
Lakshmi Narayana Panuku, Chellu Chandra Sekhar |
ICONIP (1) | 2 |
| 2007 | Acoustic Modeling using Vector Quantization in Kernel Feature Space and Classification using String Kernel based Support Vector MachinesabstractIn this paper, we propose an approach to acoustic modeling using vector quantization in a Mercer kernel feature space to obtain a sequence of codebook indices, and then use a support vector machine based classifier to classify the sequence of codebook indices. Clustering and vector quantization in the kernel feature space induced by a nonlinear innerproduct kernel is helpful in proper separation of nonlinearly separable clusters in the input acoustic feature space. Effectiveness of the proposed approach to acoustic modeling is demonstrated for recognition of spoken letters in E-set of English alphabet, and for recognition of a large number of consonant-vowel type subword units in continuous speech of three Indian languages. Performance of the proposed approach to acoustic modeling is compared with that of a continuous density hidden Markov model based classifier in the input acoustic feature space. Though there is a significant loss of information due to discretization involved in vector quantization, the proposed approach gives a performance better than that of classifiers using the continuous valued acoustic feature representation. R. Anitha 0003, Chellu Chandra Sekhar |
IJCNN | 2 |
| 2007 | Local Density Estimation based ClusteringabstractIn this paper we propose a density based clustering approach. A kernel based density estimation technique is used to estimate the density of the given data set using a Gaussian kernel. Generally, a fixed width parameter is used for all the Gaussians in such methods. Here, a method to automatically determine the widths of Gaussians by considering the information available locally at a data point has been proposed. Cluster boundary information is subsequently extracted from the estimated density of the data. The performance of the proposed method is demonstrated on several data sets. Studies comparing the performance of the proposed method with that of DBSCAN and SVC are also presented. Sheetal Reddy Pamudurthy, S. Chandrakala 0001, Chellu Chandra Sekhar |
IJCNN | 3 |
| 2007 | Acoustic Modeling Using Continuous Density Hidden Markov Models in the Mercer Kernel Feature Space
R. Anitha 0003, Chellu Chandra Sekhar |
ISNN (1) | 2 |
| 2006 | Identification of Block Ciphers using Support Vector MachinesabstractIn this paper, we propose an approach for identification of encryption method for block ciphers using support vector machines. The task of identification of encryption method from cipher text only is considered as a document categorization task. We address the issues in representing a cipher text by a document vector. We consider the common dictionary based method and the class specific dictionary based method for generating a document vector from a cipher text. As the dimension of document vector is large, support vector machines based classifiers are considered for identification of encryption method. We present the performance of the proposed approach for cipher texts generated using five block ciphers. Aroor Dinesh Dileep, Chellu Chandra Sekhar |
IJCNN | 2 |
| 2006 | Kernel based Clustering and Vector Quantization for Speech SegmentationabstractIn this paper, we propose an approach to segmentation of continuous speech into syllable-like units where each unit has one or more consonants followed by a vowel. The proposed approach uses the clustering and vector quantization methods to identify the consonant, transition and vowel regions in continuous speech. We consider methods based on clustering and vector quantization in the Mercer kernel feature space for separation of nonlinearly separable clusters of data belonging to the different regions. Results of experimental studies demonstrate the effectiveness of the kernel based methods in improving the performance of the speech segmentation system. D. Srikrishna Satish, Chellu Chandra Sekhar |
IJCNN | 2 |
| 2004 | Kernel Based Clustering for Multiclass Data
D. Srikrishna Satish, Chellu Chandra Sekhar |
ICONIP | 2 |
| 2004 | A Fast and Efficient Face Detection Technique Using Support Vector Machine
R. Suguna, N. Sudha, Chellu Chandra Sekhar |
ICONIP | 3 |
| 2004 | Acoustic model combination for recognition of speech in multiple languages using support vector machinesabstractWe study the performance of support vector machine based classifiers in acoustic model combination for recognition of context dependent sub word units of speech in multiple languages. In acoustic model combination, the data for similar sub word units across languages are shared to train acoustic models for multilingual speech. Sharing of data across languages leads to an increase in the number of training examples for a subword unit common to the languages. It may also lead to increase in the variability of the data for a subword unit. In This work, we study the effect of data sharing on the classification accuracy and complexity of acoustic models built using support vector machines. We compare the performance of multilingual acoustic models with that of monolingual acoustic models in the recognition of a large number of consonant-vowel units in the broadcast news corpus of three Indian languages. Suryakanth V. Gangashetty, Chellu Chandra Sekhar, Bayya Yegnanarayana |
IJCNN | 2 |
| 2004 | Detection of vowel on set points in continuous speech using autoassociative neural network modelsabstractDetection of vowel onset points (VOPs) is important for spotting subword units in continuous speech. For consonant-vowel (CV) utterances, VOP is the instant at which the consonant part ends and the vowel part begins. Accurate detection of VOPs is important for recognition of CV units in continuous speech. In this paper, we propose an approach for detection of VOPs using autoassociative neural network (AANN) models. A pair of AANN models are trained for each CV class to capture the characteristics of speech signal in the consonant and vowel regions of that class. The trained AANN models are then used to detect VOPs in continuous speech. The results of studies show that the proposed approach leads to significantly less number of spurious hypotheses. Suryakanth V. Gangashetty, Chellu Chandra Sekhar, Bayya Yegnanarayana |
INTERSPEECH | 2 |
| 2003 | Constraint satisfaction model for enhancement of evidence in recognition of consonant-vowel utterancesabstractWe address the issues in recognition of a large number of subword units of speech with high confusability among several units. Evidence available from the classification models trained with a limited number of training examples may not be strong to correctly recognize the subword units. We present a constraint satisfaction neural network model that can be used to enhance the evidence for a particular unit with the supporting evidence available for a subset of units confusable with that unit. We demonstrate the enhancement of evidence by the proposed model in recognition of utterances of 145 consonant-vowel units. Suryakanth V. Gangashetty, Chellu Chandra Sekhar, Bayya Yegnanarayana |
ICASSP (2) | 2 |
| 2003 | Constraint satisfaction model for enhancement of evidence in recognition of consonant-vowel utterancesabstractIn this paper, we address the issues in recognition of a large number of subword units of speech with high confusability among several units. Evidence available from the classification models trained with a limited number of training examples may not be strong to correctly recognize the subword units. We present a constraint satisfaction neural network model that can be used to enhance the evidence for a particular unit with the supporting evidence available for a subset of units confusable with the unit. We demonstrate the enhancement of evidence by the proposed model in recognition of utterances of 145 consonant-vowel units. Suryakanth V. Gangashetty, Chellu Chandra Sekhar, Bayya Yegnanarayana |
ICME | 2 |
| 2003 | Combining evidence from multiple modular networks for recognition of consonant-vowel units of speechabstractIn this paper, we present a method to combine evidence from multiple classifiers to recognize a large number of subword units of speech using small size training data sets. Grouping criteria based on phonetic description are considered, to build multiple modular networks for recognition of the large number of units. Nonlinear compression of feature vectors is carried out to obtain reduced dimensional patterns, and multiple classifiers are trained separately using the uncompressed feature vectors and compressed feature vectors. Evidence from multiple classifiers at different stages in the recognition system is combined using the sum rule. Effectiveness of the proposed method is demonstrated for recognition of isolated utterances of 145 consonant-vowel units of speech. Suryakanth V. Gangashetty, K. Sreenivasa Rao, A. Nayeemulla Khan, Chellu Chandra Sekhar, Bayya Yegnanarayana |
IJCNN | 4 |
| 2002 | Recognition of continuous speech segments of monophone units using support vector machines
Weifeng Lee, Chellu Chandra Sekhar, Kazuya Takeda, Fumitada Itakura |
INTERSPEECH | 2 |
| 2002 | A constraint satisfaction model for recognition of stop consonant-vowel (SCV) utterancesabstractWe propose a model for recognition of utterances of consonant-vowel (CV) units. The acoustic-phonetic knowledge of the CV classes is incorporated in the form of constraints of a constraint satisfaction model. The model combines evidence from multiple classifiers. The significant feature of this model is that discrimination of the CV units could be enhanced by a combination of even weak evidence derived from the features. The evidence is obtained from multilayer feedforward neural networks trained for subgroups of CV classes. The evidence is enhanced using a set of feedback subnetworks in the constraint satisfaction model. The weights for the connections in the feedback subnetworks are derived using acoustic-phonetic knowledge and the performance statistics of the trained networks. The performance of the proposed model is demonstrated for recognition of utterances of a large number (80) of stop consonant-vowel units for the Indian language Hindi. Chellu Chandra Sekhar, Bayya Yegnanarayana |
IEEE Trans. Speech Audio Process. | 1 |
| 2001 | Recognition of consonant-vowel utterances using Support Vector Machines
Chellu Chandra Sekhar, Kazuya Takeda, Fumitada Itakura |
ESANN | 1 |
| 2001 | Close-Class-Set Discrimination Method for Recognition of Stop_Consonant-Vowel Utterances Using Support Vector Machines
Chellu Chandra Sekhar, Kazuya Takeda, Fumitada Itakura |
ICANN | 1 |
| 1991 | Synthesizing intonation for speech in hindi
A. S. Madhukumar, Chellu Chandra Sekhar, Bayya Yegnanarayana |
EUROSPEECH | 3 |
| 1989 | Parsing spoken utterances in an inflectional language
M. Prakash 0002, Venkata Ramana Rao Gadde, Chellu Chandra Sekhar, Bayya Yegnanarayana |
EUROSPEECH | 3 |