Jithendra Vepa

dblp:57/3658 · DBLP profile ↗
← Back
38ranked-venue papers
6as first author
14since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 5 first-author · 14 since 2021Artificial intelligence and machine learning · 32 · 5 first-author · 14 since 2021
YearPublicationVenuePosition
2024 Leveraging large language models for post-transcription correction in contact centers
Bramhendra Koilakuntla, Prajesh Rana, Paras Ahuja, Srikanth Konjeti, Jithendra Vepa
INTERSPEECH5
2024 Detection of background agents speech in contact centers
Srikanth Konjeti, Jithendra Vepa
INTERSPEECH3
2023 Cross-lingual/Cross-channel Intent Detection in Contact-Center Conversations
Suraj Agrawal, Aashraya Sachdeva, Soumya Jain, Cijo George, Jithendra Vepa
INTERSPEECH5
2023 Listening To Silences In Contact Center Conversations Using Textual Cues
Digvijay Ingle, Jithendra Vepa
INTERSPEECH3
2023 What questions are my customers asking?: Towards Actionable Insights from Customer Questions in Contact Center Calls
Varun Nathan, Devashish Deshpande, Cijo George, Jithendra Vepa
INTERSPEECH5
2023 CauSE: Causal Search Engine for Understanding Contact-Center Conversations
Anup Pattnaik, Tanay Narshana, Aashraya Sachdeva, Cijo George, Jithendra Vepa
INTERSPEECH5
2023 Tailored Real-Time Call Summarization System for Contact Centers
Aashraya Sachdeva, Sai Nishanth Padala, Anup Pattnaik, Varun Nathan, Cijo George, Jithendra Vepa
INTERSPEECH7
2023 COnVoy: A Contact Center Operated Pipeline for Voice of Customer Discovery
Rishabh Kumar Tripathi, Digvijay Ingle, Cijo George, Jithendra Vepa
INTERSPEECH5
2022 Interpretabilty of Speech Emotion Recognition modelled using Self-Supervised Speech and Text Pre-Trained Embeddings
K. V. Vijay Girish, Srikanth Konjeti, Jithendra Vepa
INTERSPEECH3
2022 Real-Time Monitoring of Silences in Contact Center Conversations
Digvijay Ingle, Krishnachaitanya Gogineni, Jithendra Vepa
INTERSPEECH4
2022 Does Utterance entails Intent?: Evaluating Natural Language Inference Based Setup for Few-Shot Intent Detection
Vijit Malik, Jithendra Vepa
INTERSPEECH3
2021 Audio Segmentation Based Conversational Silence Detection for Contact Center Calls
Krishnachaitanya Gogineni, Tarun Reddy Yadama, Jithendra Vepa
Interspeech3
2021 Addressing Compliance in Call Centers with Entity Extraction
Sai Guruju, Jithendra Vepa
Interspeech2
2021 PhonemeBERT: Joint Language Modelling of Phoneme Sequence and ASR Transcript
abstract
Recent years have witnessed significant improvement in ASR systems to recognize spoken utterances. However, it is still a challenging task for noisy and out-of-domain data, where substitution and deletion errors are prevalent in the transcribed text. These errors significantly degrade the performance of downstream tasks. In this work, we propose a BERT-style language model, referred to as PhonemeBERT, that learns a joint language model with phoneme sequence and ASR transcript to learn phonetic-aware representations that are robust to ASR errors. We show that PhonemeBERT can be used on downstream tasks using phoneme sequences as additional features, and also in low-resource setup where we only have ASR-transcripts for the downstream tasks with no phoneme information available. We evaluate our approach extensively by generating noisy data for three benchmark datasets - Stanford Sentiment Treebank, TREC and ATIS for sentiment, question and intent classification tasks respectively. The results of the proposed approach beats the state-of-the-art baselines comprehensively on each dataset.
Mukuntha N. S. 0001, Jithendra Vepa
Interspeech3
2020 Gated Mechanism for Attention Based Multi Modal Sentiment Analysis
abstract
Multimodal sentiment analysis has recently gained popularity because of its relevance to social media posts, customer service calls and video blogs. In this paper, we address three aspects of multimodal sentiment analysis; 1. Cross modal interaction learning, i.e. how multiple modalities contribute to the sentiment, 2. Learning long-term dependencies in multimodal interactions and 3. Fusion of unimodal and cross modal cues. Out of these three, we find that learning cross modal interactions is beneficial for this problem. We perform experiments on two benchmark datasets, CMU Multimodal Opinion level Sentiment Intensity (CMU-MOSI) and CMU Multimodal Opinion Sentiment and Emotion Intensity (CMU-MOSEI) corpus. Our approach on both these tasks yields accuracies of 83.9% and 81.1% respectively, which is 1.6% and 1.34% absolute improvement over current state-of-the-art.
Jithendra Vepa
ICASSP2
2018 CACTAS - Collaborative Audio Categorization and Transcription for ASR Systems
Mithul Mathivanan, Kinnera Saranu, Jithendra Vepa
INTERSPEECH4
2018 Leveraging Second-Order Log-Linear Model for Improved Deep Learning Based ASR Performance
Ankit Raj, Shakti P. Rath, Jithendra Vepa
INTERSPEECH3
2018 Hierarchical Accent Determination and Application in a Large Scale ASR System
Ramya Viswanathan, Periyasamy Paramasivam, Jithendra Vepa
INTERSPEECH3
2018 Speech Emotion Recognition Using Spectrogram & Phoneme Embedding
Promod Yenigalla, Abhay Kumar 0001, Suraj Tripathi, Chirag Singh, Sibsambhu Kar, Jithendra Vepa
INTERSPEECH6
2008 Segmentation of heart sounds using simplicity features and timing information
abstract
Automatic analysis of heart sounds aid physicians in the diagnosis of abnormal heart valve conditions. Segmentation, i.e. identification of first (S1) and second (S2) heart sounds, is the first step in the automatic analysis. In this work, we have proposed a segmentation method which uses energy-based and simplicity-based features computed from multi-level wavelet decomposition coefficients. This method utilizes timing information of S1 and S2 based on biomedical domain knowledge. Proposed method has been evaluated on several normal and abnormal heart sounds for identification of S1 and S2 and compared with windowed energy-based and simplicity-based approaches individually. The proposed method is an efficient and robust technique for identification and gating of S1 and S2, and yields better results than the above approaches.
Jithendra Vepa, Paresh Tolay
ICASSP1
2007 An Acoustic Model Based on Kullback-Leibler Divergence for Posterior Features
abstract
This paper investigates the use of features based on posterior probabilities of subword units such as phonemes. These features are typically transformed when used as inputs for a hidden Markov model with mixture of Gaussians as emission distribution (HMM/GMM). In this work, we introduce a novel acoustic model that avoids the Gaussian assumption and directly uses posterior features without any transformation. This model is described by a finite state machine where each state is characterized by a target distribution and the cost function associated to each state is given by the Kullback-Leibler (KL) divergence between its target distribution and the posterior features. Furthermore, hybrid HMM/ANN system can be seen as a particular case of this KL-based model where state target distributions are predefined. A recursive training algorithm to estimate the state target distributions is also presented.
Guillermo Aradilla, Jithendra Vepa, Hervé Bourlard
ICASSP (4)2
2007 The AMI System for the Transcription of Speech in Meetings
abstract
This paper describes the AMI transcription system for speech in meetings developed in collaboration by five research groups. The system includes generic techniques such as discriminative and speaker adaptive training, vocal tract length normalisation, heteroscedastic linear discriminant analysis, maximum likelihood linear regression, and phone posterior based features, as well as techniques specifically designed for meeting data. These include segmentation and cross-talk suppression, beam-forming, domain adaptation, Web-data collection, and channel adaptive training. The system was improved by more than 20% relative in word error rate compared to our previous system and was used in the NIST RT106 evaluations where it was found to yield competitive performance.
Thomas Hain, Vincent Wan, Lukás Burget, Martin Karafiát, John Dines, Jithendra Vepa, Giulia Garau, Mike Lincoln
ICASSP (4)6
2007 Direct optimisation of a multilayer perceptron for the estimation of cepstral mean and variance statistics
abstract
We propose an alternative means of training a multilayer perceptron for the task of speech activity detection based on a criterion to minimise the error in the estimation of mean and variance statistics for speech cepstrum based features using the Kullback-Leibler divergence. We present our baseline and proposed speech activity detection approaches for multi-channel meeting room recordings and demonstrate the effectiveness of the new criterion by comparing the two approaches when used to carry out cepstrum mean and variance normalisation of features used in our meeting ASR system.
John Dines, Jithendra Vepa
INTERSPEECH2
2007 Multi-stream features combination based on dempster-shafer rule for LVCSR system
abstract
This paper investigates the combination of two streams of acoustic features. Extending our previous work on small vocabulary task, we show that combination based on Dempster-Shafer rule outperforms several classical rules like sum, product and inverse entropy weighting even in LVCSR systems. We analyze results in terms of Frame Error Rate and Cross Entropy measures. Experimental framework uses meeting transcription task and results are provided on RT05 evaluation data. Results are consistent with what has been previously observed on smaller databases.
Fabio Valente, Jithendra Vepa, Hynek Hermansky
INTERSPEECH2
2007 Hierarchical neural networks feature extraction for LVCSR system
abstract
This paper investigates the use of a hierarchy of Neural Networks for performing data driven feature extraction.Two different hierarchical structures based on long and short temporal context are considered.Features are tested on two different LVCSR systems for Meetings data (RT05 evaluation data) and for Arabic Broadcast News (BNAT05 evaluation data).The hierarchical NNs outperforms the single NN features consistently on different type of data and tasks and provides significant improvements w.r.t.respective baselines systems.Best results are obtained when different time resolutions are used at different level of the hierarchy.
Fabio Valente, Jithendra Vepa, Christian Plahl, Christian Gollan, Hynek Hermansky, Ralf Schlüter
INTERSPEECH2
2006 Using Pitch as Prior Knowledge in Template-Based Speech Recognition
abstract
In a previous paper on speech recognition, we showed that templates can better capture the dynamics of speech signal compared to parametric models such as hidden Markov models. The key point in template matching approaches is finding the most similar templates to the test utterance. Traditionally, this selection is given by a distortion measure on the acoustic features. In this work, we propose to improve this template selection with the use of meta-linguistic information as prior knowledge. In this way, similarity is not only based on acoustic features but also on other sources of information that are present in the speech signal. Results on a continuous digit recognition task confirm the statement that similarity between words does not only depend on acoustic features since we obtained 24% relative improvement over the baseline. Interestingly, results are better even when compared to a system with no prior information but a larger number of templates
Guillermo Aradilla, Jithendra Vepa, Hervé Bourlard
ICASSP (1)2
2006 Using More Informative Posterior Probabilities for Speech Recognition
abstract
In this paper, we present initial investigations towards boosting posterior probability based speech recognition systems by estimating more informative posteriors taking into account acoustic context (e.g., the whole utterance), as well as possible prior information (such as phonetic and lexical knowledge). These posteriors are estimated based on HMM state posterior probability definition (typically used in standard HMMs training). This approach provides a new, principled, theoretical framework for hierarchical estimation/use of more informative posteriors integrating appropriate context and prior knowledge. In the present work, we used the resulting posteriors as local scores for decoding. On the OGI numbers database, this resulted in significant performance improvement, compared to using MLP estimated posteriors for decoding (hybrid HMM/ANN approach) for clean and more specially for noisy speech. The system is also shown to be much less sensitive to tuning factors (such as phone deletion penalty, language model scaling) compared to the standard HMM/ANN and HMM/GMM systems, thus practically it does not need to be tuned to achieve the best possible performance
Hamed Ketabdar, Jithendra Vepa, Samy Bengio, Hervé Bourlard
ICASSP (1)2
2006 Using posterior-based features in template matching for speech recognition
abstract
Given the availability of large speech corpora, as well as the increasing of memory and computational resources, the use of template matching approaches for automatic speech recognition (ASR) have recently attracted new attention. In such template-based approaches, speech is typically represented in terms of acoustic vector sequences, using spectral-based features such as MFCC of PLP, and local distances are usually based on Euclidean or Mahalanobis distances. In the present paper, we further investigate template-based ASR and show (on a continuous digit recognition task) that the use of posterior-based features significantly improves the standard template-based approaches, yielding to systems that are very competitive to state-of-the-art HMMs, even when using a very limited number (e.g., 10) of reference templates. Since those posteriors-based features can also be interpreted as a probability distribution, we also show that using Kullback-Leibler (KL) divergence as a local distance further improves the performance of the template-based approach, now beating state-of-the-art of more complex posterior-based HMMs systems (usually referred to as "Tandem").
Guillermo Aradilla, Jithendra Vepa, Hervé Bourlard
INTERSPEECH2
2006 The segmentation of multi-channel meeting recordings for automatic speech recognition
abstract
One major research challenge in the domain of the analysis of meeting room data is the automatic transcription of what is spoken during meetings, a task which has gained considerable attention within the ASR research community through the NIST rich transcription evaluations conducted over the last three years. One of the major difficulties in carrying out automatic speech recognition (ASR) on this data is dealing with the challenging recording environment, which has instigated the development of novel audio pre-processing approaches. In this paper we present a system for the automatic segmentation of multiple-channel individual headset microphone (IHM) meeting recordings for automatic speech recognition. The system relies on an MLP classifier trained from several meeting room corpora to identify speech/non-speech segments of the recordings. We give a detailed analysis of the segmentation performance for a number of system configurations, with our best system achieving ASR performance on automatically generated segments within 1.3\% (3.7\% relative) of a manual segmentation of the data.
John Dines, Jithendra Vepa, Thomas Hain
INTERSPEECH2
2006 Posterior based keyword spotting with a priori thresholds
abstract
In this paper, we propose a new posterior based scoring approach for keyword and non keyword (garbage) elements. The estimation of these scores is based on HMM state posterior probability definition, taking into account long contextual information and the prior knowledge (e.g. keyword model topology). The state posteriors are then integrated into keyword and garbage posteriors for every frame. These posteriors are used to make a decision on detection of the keyword at each frame. The frame level decisions are then accumulated (in this case, by counting) to make a global decision on having the keyword in the utterance. In this way, the contribution of possible outliers are minimized, as opposed to the conventional Viterbi decoding approach which accumulates likelihoods. Experiments on keywords from the Conversational Telephone Speech (CTS) and Numbers'95 databases are reported. Results show that the new scoring approach leads to better trade off between true and false alarms compared to the Viterbi decoding approach, while also providing the possibility to precalculate keyword specific spotting thresholds related to the length of the keywords.
Hamed Ketabdar, Jithendra Vepa, Samy Bengio, Hervé Bourlard
INTERSPEECH2
2006 Multi-stream ASR: an oracle perspective
abstract
Multi-stream based automatic speech recognition (ASR) systems are usually shown to outperform single stream systems, specially in noisy test conditions. And, indeed, there is a trend today in ASR towards using more and more acoustic features combined at the input (early integration, possibly preceded by some linear or nonlinear transformation) or later in the recognition process (e.g., at the level of likelihoods, then referred to as late integration). However, to guarantee optimal exploitation of such multi-stream systems, we need to use features that are as much complementary as possible, while also using the best combination method for those streams. In practice, it is never clear whether we fully exploit the potential of the available streams. This present paper investigates an ‘oracle ’ test to provide some insight in these issues. Although not providing us with an absolute performance upper bound, oracle is shown to indicate the complimentary of the feature streams used, and to provide a reasonable reference target to evaluate combination strategies. The oracle analysis is supported by results obtained on Numbers95 database using different feature streams and entropy based combination method.
Hemant Misra, Jithendra Vepa, Hervé Bourlard
INTERSPEECH2
2006 Subjective evaluation of join cost and smoothing methods for unit selection speech synthesis
abstract
In unit selection-based concatenative speech synthesis, join cost (also known as concatenation cost), which measures how well two units can be joined together, is one of the main criteria for selecting appropriate units from the inventory. Usually, some form of local parameter smoothing is also needed to disguise the remaining discontinuities. This paper presents a subjective evaluation of three join cost functions and three smoothing methods. We also describe the design and performance of a listening test. The three join cost functions were taken from our previous study, where we proposed join cost functions derived from spectral distances, which have good correlations with perceptual scores obtained for a range of concatenation discontinuities. This evaluation allows us to further validate their ability to predict concatenation discontinuities. The units for synthesis stimuli are obtained from a state-of-the-art unit selection text-to-speech system: rVoice from Rhetorical Systems Ltd. In this paper, we report listeners' preferences for each join cost in combination with each smoothing method
Jithendra Vepa, Simon King 0001
IEEE Trans. Speech Audio Process.1
2005 Improving speech recognition using a data-driven approach
abstract
In this paper, we investigate the possibility of enhancing state-of-the-art HMM-based speech recognition systems using data-driven techniques, where whole set of training utterances is used as reference models and recognition is then performed through the well-known template matching technique, DTW. This approach allows us to better capture the temporal dynamics of the speech signal while avoiding some of the HMM assumptions such as the piecewise stationarity. Potentially, such data-driven techniques also allow us to better exploit meta-data and environmental information, such as speaker, gender, accent and noise conditions. However, we cannot entirely abandon HMMs, which are very powerful and scalable models. Thus, we investigate one way to combine and take advantage of both the approaches, combining scores of HMMs and reference templates. Experiments on the Numbers95 database showed that this combination yields 22\% relative improvement in word error rate over the baseline HMM performance. Applying K-means clustering to the acoustic vectors speeds up the decoding, while still retaining a significant improvement in the recognition accuracy.
Guillermo Aradilla, Jithendra Vepa, Hervé Bourlard
INTERSPEECH2
2005 Developing and enhancing posterior based speech recognition systems
abstract
Local state or phone posterior probabilities are often investigated as local scores (e.g., hybrid HMM/ANN systems) or as transformed acoustic features (e.g., “Tandem”) to improve speech recognition systems. In this paper, we present initial results towards boosting these approaches by improving posterior estimates, using acoustic context (e.g., as available in the whole utterance), as well as possible prior information (such as topological constraints). In the present work, the enhanced posterior distribution is associated with the “gamma ” distribution typically used in standard HMMs training, and estimated from local likelihoods (GMM) or local posteriors (ANN). This approach results in a family of new HMM based systems, where only posterior probabilities are used, while also providing a new, principled, approach towards a hierarchical use/integration of these posteriors, from the frame level up to the phone and word levels, and integrating the appropriate context and prior knowledge in each level. In the present work, we used the resulting posteriors as local scores in a Viterbi decoder. On the OGI Numbers’95 database, this resulted in improved recognition performance, compared to a state-of-the-art hybrid HMM/ANN system. 1.
Hamed Ketabdar, Jithendra Vepa, Samy Bengio, Hervé Bourlard
INTERSPEECH2
2004 Subjective evaluation of join cost functions used in unit selection speech synthesis
abstract
In our previous papers, we have proposed join cost functions derived from spectral distances, which have good correlations with perceptual scores obtained for a range of concatenation discontinuities. To further validate their ability to predict concatenation discontinuities, we have chosen the best three spectral distances and evaluated them subjectively in a listening test. The unit sequences for synthesis stimuli are obtained from a state-of-the-art unit selection text-to-speech system: rVoice from Rhetorical Systems Ltd. In this paper, we report listeners' preferences for each of the three join cost functions.
Jithendra Vepa, Simon King 0001
INTERSPEECH1
2003 Kalman-filter based join cost for unit-selection speech synthesis
abstract
We introduce a new method for computing join cost in unit-selection speech synthesis which uses a linear dynamical model (also known as a Kalman filter) to model line spectral frequency trajectories. The model uses an underlying subspace in which it makes smooth, continuous trajectories. This subspace can be seen as an analogy for underlying articulator movement. Once trained, the model can be used to measure how well concatenated speech segments join together. The objective join cost is based on the error between model predictions and actual observations. We report correlations between this measure and mean listener scores obtained from a perceptual listening experiment. Our experiments use a state-of-the art unit-selection text-to-speech system: `rVoice' from Rhetorical Systems Ltd.
Jithendra Vepa, Simon King 0001
INTERSPEECH1
2002 A text-to-speech synthesis system for telugu
Jithendra Vepa, Jahnavi Ayachitam, K. V. K. Kalpana Reddy
INTERSPEECH1
2002 Objective distance measures for spectral discontinuities in concatenative speech synthesis
abstract
In unit selection based concatenative speech systems, join cost, which measures how well two units can be joined together, is one of the main criteria for selecting appropriate units from the inventory. The ideal join cost will measure perceived discontinuity, based on easily measurable spectral properties of the units being joined, in order to ensure smooth and natural-sounding synthetic speech. In this paper we report a perceptual experiment conducted to measure the correlation between subjective human perception and various objective spectrally-based measures proposed in the literature. Our experiments used a state-of-the art unit-selection text-to-speech system: rVoice from Rhetorical Systems Ltd. 1.
Jithendra Vepa, Simon King 0001, Paul Taylor 0001
INTERSPEECH1