VLDB 2026 Research / reviewers in the wild / expert
Ashish Verma 0001
dblp:41/442-1
· DBLP profile ↗
54ranked-venue papers
8as first author
6since 2021 · last 2024
0009-0003-5222-1702ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 7 first-author · 1 since 2021Artificial intelligence and machine learning · 19 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 10 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 since 2021Systems, architecture and hardware · 5 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Network and information security
1 paper |
Privacy and data protection · 87% Hardware security and side channels · 13% | |
| Artificial intelligence
2 papers |
Efficient and distributed learning · 35% Language models and text generation · 35% Question answering and dialogue systems · 15% | |
| Databases, data mining, and information retrieval
4 papers |
Data mining · 71% Web and social media mining · 21% Information retrieval · 7% | |
| Computer graphics and multimedia
1 paper |
Computer animation and physical simulation · 75% Audio and music processing · 25% |
Topics — the 15 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Privacy and data protection › privacy-preserving machine learning
federated learning privacy |
0.8 | 1 | 2024 | DeTA: Minimizing Data Leaks in Federated Learning via Decentralized and Trustworthy Aggregation · EuroSys 2024 |
Machine learning › Efficient and distributed learning › inference efficiency
efficient transformer inference |
0.4 | 1 | 2020 | PoWER-BERT: Accelerating BERT Inference via Progressive Word-vector Elimination · ICML 2020 |
Hardware security and side channels › trusted execution environments
confidential computing |
0.2 | 1 | 2024 | DeTA: Minimizing Data Leaks in Federated Learning via Decentralized and Trustworthy Aggregation · EuroSys 2024 |
Natural language and speech › Question answering and dialogue systems › question generation
question-answer pair generation |
0.2 | 1 | 2014 | Automatic generation of question answer pairs from noisy case logs · ICDE 2014 |
Natural language and speech › Information extraction and text analysis
text segmentation |
0.2 | 1 | 2014 | Automatic generation of question answer pairs from noisy case logs · ICDE 2014 |
Data mining
customer relationship management |
0.2 | 1 | 2013 | A CRM system for social media: challenges and experiences · WWW 2013 |
Web and social media mining
social media analysis |
0.2 | 1 | 2013 | A CRM system for social media: challenges and experiences · WWW 2013 |
Data mining
text mining |
0.2 | 1 | 2013 | A CRM system for social media: challenges and experiences · WWW 2013 |
Data mining
clustering |
0.1 | 1 | 2009 | Cross-Guided Clustering: Transfer of Relevant Supervision across Domains for Improved Clustering · ICDM 2009 |
Services computing and microservices
enterprise systems |
0.0 | 1 | 2013 | A CRM system for social media: challenges and experiences · WWW 2013 |
Computer animation and physical simulation
facial animation |
0.0 | 1 | 2004 | Animating expressive faces across languages · IEEE Trans. Multim. 2004 |
Computer animation and physical simulation › facial animation
lip synchronization |
0.0 | 1 | 2004 | Animating expressive faces across languages · IEEE Trans. Multim. 2004 |
Computer animation and physical simulation › facial animation
speech-driven facial animation |
0.0 | 1 | 2004 | Animating expressive faces across languages · IEEE Trans. Multim. 2004 |
Audio and music processing
speech processing |
0.0 | 1 | 2004 | Animating expressive faces across languages · IEEE Trans. Multim. 2004 |
Data mining › clustering
k-means clustering |
0.0 | 1 | 2009 | Cross-Guided Clustering: Transfer of Relevant Supervision across Domains for Improved Clustering · ICDM 2009 |
Methods — techniques the papers use, named apart from their topics
decentralized aggregation · 0.8confidential computing · 0.8self-attention significance measurement · 0.4progressive elimination · 0.4latent dirichlet allocation · 0.4hidden markov model · 0.4conditional random field · 0.4conversation mining · 0.3k-means · 0.1cross-domain similarity measure · 0.1case study · 0.1optical flow · 0.0morphing · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | DeTA: Minimizing Data Leaks in Federated Learning via Decentralized and Trustworthy AggregationabstractFederated learning (FL) relies on a central authority to oversee and aggregate model updates contributed by multiple participating parties in the training process. This centralization of sensitive model updates naturally raises concerns about the trustworthiness of the central aggregation server, as well as the potential risks associated with server failures or breaches, which could result in loss and leaks of model updates. Moreover, recent attacks have demonstrated that, by obtaining the leaked model updates, malicious actors can even reconstruct substantial amounts of private data belonging to training participants. This underscores the critical necessity to rethink the existing FL system architecture to mitigate emerging attacks in the evolving threat landscape. One straightforward approach is to fortify the central aggregator with confidential computing (CC), which offers hardware-assisted protection for runtime computation and can be remotely verified for execution integrity. However, a growing number of security vulnerabilities have surfaced in tandem with the adoption of CC, indicating that depending solely on this singular defense may not provide the requisite resilience to thwart data leaks. Pau-Chen Cheng, Kevin Eykholt, Zhongshu Gu, Hani Jamjoom, K. R. Jayaram, Enriquillo Valdez, Ashish Verma 0001 |
EuroSys | 7 |
| 2023 | Runtime Prediction of Machine Learning Algorithms in Automl SystemsabstractIn this paper we introduce a metalearning-based methodology for predicting the training runtime of various machine learning algorithms. This prediction is important for automated machine learning (AutoML) systems because they search by training and evaluating a large number of machine learning models in order to identify the best model for a given dataset. Our approach identifies the main factors that impact the runtime performance of state of the art algorithms used in AutoML systems and can be used to enhance their performance in resource-constrained settings. Parijat Dube, Theodoros Salonidis, Parikshit Ram, Ashish Verma 0001 |
ICASSP | 4 |
| 2023 | FLIPS: Federated Learning using Intelligent Participant SelectionabstractThis paper presents the design and implementation of FLIPS, a middleware system to manage data and participant heterogeneity in federated learning (FL) training workloads. In particular, we examine the benefits of label distribution clustering on participant selection in federated learning. FLIPS clusters parties involved in an FL training job based on the label distribution of their data apriori, and during FL training, ensures that each cluster is equitably represented in the participants selected. FLIPS can support the most common FL algorithms, including FedAvg, FedProx, FedDyn, FedOpt and FedYogi. To manage platform heterogeneity and dynamic resource availability, FLIPS incorporates a straggler management mechanism to handle changing capacities in distributed, smart community applications. Privacy of label distributions, clustering and participant selection is ensured through a trusted execution environment (TEE). Our comprehensive empirical evaluation compares FLIPS with random participant selection, as well as three other "smart" selection mechanisms -- Oort [51], TiFL [15] and gradient clustering [27] using four real-world datasets, two different non-IID distributions and three common FL algorithms (FedYogi, FedProx and FedAvg). We demonstrate that FLIPS significantly improves convergence, achieving higher accuracy by 17-20 percentage points with 20-60% lower communication costs, and these benefits endure in the presence of straggler participants. Rahul Atul Bhope, K. R. Jayaram, Nalini Venkatasubramanian, Ashish Verma 0001, Gegi Thomas |
Middleware | 4 |
| 2022 | Adaptive Aggregation For Federated LearningabstractIn this paper, we present a new scalable and adaptive architecture for FL aggregation. First, we demonstrate how traditional tree overlay based aggregation techniques (from P2P, publish-subscribe and stream processing research) can help FL aggregation scale, but are ineffective from a resource utilization and cost standpoint. Next, we present the design and implementation of AdaFed, which uses serverless/cloud functions to adaptively scale aggregation in a resource efficient and fault tolerant manner. We describe how AdaFed enables FL aggregation to be dynamically deployed only when necessary, elastically scaled to handle participant joins/leaves and is fault tolerant with minimal effort required on the (aggregation) programmer side. We also demonstrate that our prototype based on Ray [1] scales to thousands of participants, and is able to achieve a > 90% reduction in resource requirements and cost, with minimal impact on aggregation latency. K. R. Jayaram, Vinod Muthusamy, Gegi Thomas, Ashish Verma 0001, Mark Purcell |
IEEE Big Data | 4 |
| 2022 | Just-in-Time Aggregation for Federated LearningabstractThe increasing number and scale of federated learning (FL) jobs necessitates resource efficient scheduling and management of aggregation to make the economics of cloud-hosted aggregation work. Existing FL research has focused on the design of FL algorithms and optimization, and less on aggregation efficacy. In this paper, we propose a new FL aggregation paradigm - “just-in-time” (JIT) aggregation that leverages unique properties of FL jobs, especially the periodicity of model updates, to defer aggregation as much as possible and free compute resources for other FL jobs or other datacenter workloads. We describe a novel way to prioritize FL jobs for aggregation, and demonstrate using multiple datasets, models and FL aggregation algorithms that our techniques can reduce resource usage by 60+% when compared to eager aggregation used in existing FL platforms. We demonstrate that using JIT aggregation has negligible overhead and impact on the latency of the FL job. K. R. Jayaram, Ashish Verma 0001, Gegi Thomas, Vinod Muthusamy |
MASCOTS | 2 |
| 2021 | HyperASPO: Fusion of Model and Hyper Parameter Optimization for Multi-objective Machine LearningabstractCurrent state of the art methods for generating Pareto-optimal solutions for multi-objective optimization problems mostly rely on optimizing the hyper-parameters of the models (HPO - hyper-parameter Optimization). Few recent, less studied methods focus on optimizing over the space of model parameters, leveraging the problem specific knowledge. We present a generic first-of-a-kind method, referred to as HyperASPO, that combines optimization over the spaces of both hyper-parameters and model parameters for multi-objective optimization of learning problems. HyperASPO consists of two stages. First, we perform a coarse HPO to determine a set of favorable hyper-parameter configurations. In the second step, for each of these configurations, we solve a sequence of weighted single objective optimization problems for estimating Pareto-optimal solutions. We generate the weights in the second step using an adaptive mesh constructed iteratively based on the metrics of interest, resulting in further refinement of Pareto frontier efficiently. We consider the widely used XGBoost (Gradient Boosted Trees) model and validate our method on multiple classification datasets. Our proposed method shows up to 20% improvement over the hypervolumes of Pareto fronts obtained through state of the art HPO based methods with up to 2× reduction in computational time. Aswin Kannan, Anamitra R. Choudhury, Vaibhav Saxena, Saurabh Raje, Parikshit Ram, Ashish Verma 0001, Yogish Sabharwal |
IEEE BigData | 6 |
| 2020 | Variable batch size across layers for efficient prediction on CNNsabstractCNNs are used extensively for computer vision tasks like activity recognition, image classification, segmentation etc. The large compute memory required in these applications restricts the use of high batch size during inference, thereby increasing the overall prediction time. Prior work addresses this issue through various model compression mechanisms like weight/filter pruning, quantizing the parameters/intermediate outputs, etc. We propose a complementary technique where we improve inference time by using variable batch sizes (VBS) across the layers of a CNN. This optimises the memory-time trade-off for each layer and leads to better network throughput. Our approach does not make any modifications to the existing network (unlike pruning or quantization techniques) and thus there is no impact on the model accuracy. We develop a dynamic program (DP) based algorithm that takes inference time and memory required by different layers of the network as input, and computes the optimal batch sizes for each layer depending on the available resources (RAM, storage space etc.). We demonstrate our findings in two different settings: video inference on K80 GPUs and image inference on Edge devices. On video networks like C3D, our VBS algorithm gives up to 61% higher throughput compared to a fixed batch size baseline. On image networks like GoogleNet, ResNet50 etc., we achieve up to 60% higher throughput compared to a fixed batch size baseline. Anamitra R. Choudhury, Saurabh Goyal, Yogish Sabharwal, Ashish Verma 0001 |
CLOUD | 4 |
| 2020 | MYSTIKO: Cloud-Mediated, Private, Federated Gradient DescentabstractFederated learning enables multiple, distributed participants (potentially on different clouds) to collaborate and train machine/deep learning models by sharing parameters/gradients. However, sharing gradients, instead of centralizing data, may not be as private as one would expect. Reverse engineering attacks [1], [2] on plaintext gradients have been demonstrated to be practically feasible. Existing solutions for differentially private federated learning, while promising, lead to less accurate models and require nontrivial hyperparameter tuning. In this paper, we examine the use of additive homomorphic encryption (specifically the Paillier cipher) to design secure federated gradient descent techniques that (i) do not require addition of statistical noise or hyperparameter tuning, (ii) does not alter the accuracy or utility of the final model, (iii) ensure that the plaintext model parameters/gradients of a participant are never revealed to any other participant or third party coordinator involved in the federated learning job, (iv) minimize the trust placed in any third party coordinator and (v) are efficient, with minimal overhead, and cost effective. K. R. Jayaram, Archit Verma, Ashish Verma 0001, Gegi Thomas, Colin Sutcher-Shepard |
CLOUD | 3 |
| 2020 | PoWER-BERT: Accelerating BERT Inference via Progressive Word-vector EliminationabstractWe develop a novel method, called PoWER-BERT, for improving the inference time of the popular BERT model, while maintaining the accuracy. It works by: a) exploiting redundancy pertaining to word-vectors (intermediate transformer block outputs) and eliminating the redundant vectors. b) determining which word-vectors to eliminate by developing a strategy for measuring their significance, based on the self-attention mechanism. c) learning how many word-vectors to eliminate by augmenting the BERT model and the loss function. Experiments on the standard GLUE benchmark shows that PoWER-BERT achieves up to 4.5x reduction in inference time over BERT with < 1% loss in accuracy. We show that PoWER-BERT offers significantly better trade-off between accuracy and inference time compared to prior methods. We demonstrate that our method attains up to 6.8x reduction in inference time with < 1% loss in accuracy when applied over ALBERT, a highly compressed version of BERT. The code for PoWER-BERT is publicly available at https://github.com/IBM/PoWER-BERT. Saurabh Goyal, Anamitra R. Choudhury, Saurabh Raje, Venkatesan T. Chakaravarthy, Yogish Sabharwal, Ashish Verma 0001 |
ICML | 6 |
| 2020 | Effective Elastic Scaling of Deep Learning WorkloadsabstractWe examine the elastic scaling of Deep Learning (DL) jobs and propose a novel resource allocation strategy for DL training jobs, resulting in improved job run time performance as well as increased cluster utilization. We begin by analyzing DL workloads and exploit the fact that DL jobs can be run with a range of batch sizes without affecting their final accuracy. We formulate an optimization problem that explores a dynamic batch size allocation to individual DL jobs based on their scaling efficiency, when running on multiple nodes. We design a fast dynamic programming based optimizer to solve this problem in real-time to determine jobs that can be scaled up/down, and use this optimizer in an autoscaler to dynamically change the allocated resources and batch sizes of individual DL jobs. We demonstrate empirically that our elastic scaling algorithm can complete up to as many jobs as compared to a strong baseline algorithm that also scales the number of GPUs but does not change the batch size, with average completion times up to faster. Vaibhav Saxena, K. R. Jayaram, Saurav Basu, Yogish Sabharwal, Ashish Verma 0001 |
MASCOTS | 5 |
| 2018 | Efficient Training of Convolutional Neural Nets on Large Distributed SystemsabstractDeep Neural Networks (DNNs) have achieved impressive accuracy in many application domains including im-age classification. Training of DNNs is an extremely compute-intensive process and is solved using variants of the stochastic gradient descent (SGD) algorithm. A lot of recent research has focused on improving the performance of DNN training. In this paper, we present optimization techniques to improve the performance of the data parallel synchronous SGD algorithm using the Torch framework: (i) we maintain data in-memory to avoid file I/O overheads, (ii) we propose optimizations to the Torch data parallel table framework that handles multi-threading, and (iii) we present MPI optimization to minimize communication overheads. We evaluate the performance of our optimizations on a Power 8 Minsky cluster with 64 nodes and 256 NVidia Pascal P100 GPUs. With our optimizations, we are able to train 90 epochs of the ResNet-50 model on the Imagenet-1k dataset using 256 GPUs in just 48 minutes. This significantly improves on the previously best known performance of training 90 epochs of the ResNet-50 model on the same dataset using the same number of GPUs in 65 minutes. To the best of our knowledge, this is the best known training performance demonstrated for the Imagenet-1k dataset using 256 GPUs. Dheeraj Sreedhar, Vaibhav Saxena, Yogish Sabharwal, Ashish Verma 0001, Sameer Kumar 0001 |
CLUSTER | 4 |
| 2018 | Balancing Stragglers Against Staleness in Distributed Deep LearningabstractSynchronous SGD is frequently the algorithm of choice for training deep learning models on compute clusters within reasonable time frames. However, even if a large number of workers (CPUs or GPUs) are at disposal for training, hetero-geneity of compute nodes and unreliability of the interconnecting network frequently pose a bottleneck to the training speed. Since the workers have to wait for each other at every model update step, even a single straggler/slow worker can derail the whole training performance. In this paper, we propose a novel approach to mitigate the straggler problem in large compute clusters. We cluster the compute nodes into multiple groups where each group updates the model synchronously stored in its own parameter server. The parameter servers of the different groups update the model in a central parameter server in an asynchronous manner. Few stragglers in the same group (or even separate groups) have little effect on the computational performance. The staleness of the asynchronous updates can be controlled by limiting the number of groups. Our method, in essence, provides a mechanism to move seamlessly between a pure synchronous and a pure asynchronous setting, thereby balancing between the computational overhead of synchronous SGD and the accuracy degradation of a pure asynchronous SGD. We empirically show that with increasing delay from straggler nodes (more than 300% delay in a node), progressive grouping of available workers still finishes the training within 20% of the no-delay case, with the limit to the number of groups governed by the permissible degradation in accuracy (≤ 2.5% compared to the no-delay case). Saurav Basu, Vaibhav Saxena, Rintu Panja, Ashish Verma 0001 |
HiPC | 4 |
| 2014 | Automatic generation of question answer pairs from noisy case logsabstractIn a customer support scenario, a lot of valuable information is recorded in the form of `case logs'. Case logs are primarily written for future references or manual inspections and therefore are written in a hasty manner and are very noisy. In this paper, we propose techniques that exploit these case logs to mine real customer concerns or problems and then map them to well written knowledge articles for that enterprise. This mapping results into generation of question-answer (QA) pairs. These QA pairs can be used for a variety of applications such as dynamically updating the frequently-asked-questions (FAQs), updating the knowledge repository etc. In this paper we show the utility of these discovered QA pairs as training data for a question-answering system. Our approach for mining the case logs is based on a composite model consisting of two generative models, viz, hidden Markov model (HMM) and latent Dirichlet allocation (LDA) model. The LDA model explains the long-range dependencies across words due to their semantic similarity and HMM models the sequential patterns present in these case logs. Such processing results in crisp `problem statement' segments which are indicative of the real customer concerns. Our experiments show that this approach finds crisp problem-statements in 56% of the cases and outperforms other alternate methods for segmentation such as HMM, LDA and conditional random field (CRF). After finding these crisp problem-statements, appropriate answers are looked up from an existing knowledge repository index forming candidate QA pairs. We show that considering only the problemstatement segments for which the answers can be found further improves the segmentation performance to 82%. Finally, we show that when these QA pairs are used as training data, the performance of a question-answering system can be improved significantly. Jitendra Ajmera, Sachindra Joshi, Ashish Verma 0001, Amol Mittal |
ICDE | 3 |
| 2013 | Sparse hidden Markov models for purer clustersabstractThe hidden Markov model (HMM) is widely popular as the de facto tool for representing temporal data; in this paper, we add to its utility in the sequence clustering domain - we describe a novel approach that allows us to directly control purity in HMM-based clustering algorithms. We show that encouraging sparsity in the observation probabilities increases cluster purity and derive an algorithm based on lpregularization; as a corollary, we also provide a different and useful interpretation of the value of p in Renyi p-entropy. We test our method on the problem of clustering non-speech audio events from the BBC sound effects corpus. Experimental results confirm that our approach does learn purer clusters, with (unweighted) average purity as high as 0.88 - a considerable improvement over both the baseline HMM (0.72) and k-means clustering (0.69). Sujeeth Bharadwaj, Mark Hasegawa-Johnson, Jitendra Ajmera, Om Deshmukh, Ashish Verma 0001 |
ICASSP | 5 |
| 2013 | Intent focused summarization of caller-agent conversationsabstractIn this paper, we propose a conditional random field (CRF) based to identify segments within call center conversations that convey caller intent. A distinguishing aspect of our approach is the use of context information of the intent bearing segments to predict the presence or absence of intents within various segments. The context is represented through a set of phrase features that are frequently present in and around the intent bearing segments. These phrases, identified in a data-driven manner, are used along with conventional word features in a CRF based sequence labeling framework to assign intent/non-intent labels to each utterance in a conversation. Another distinguishing aspect of our approach is that instead of using 1-best label alignment, we extract N-best label alignments at the output of CRF and combine evidences from them to rank the utterances according to their intent bearing potential, so that top ranked utterances can be chosen as the intent summary. To demonstrate the effectiveness of our approach and to evaluate the influence of automatic speech recognition (ASR) errors we evaluated our approach using manually transcribed and ASR transcribed conversations. Experimental results show improved summarization accuracy using our approach. Specifically, in 92% of the manually transcribed conversations accurate summaries of just one utterance length can be extracted using the proposed approach. Shajith Ikbal, Ashish Verma 0001, Prasanta Ghosh, Kenneth Church 0001, Jeffrey Marcus |
ICASSP | 2 |
| 2013 | Content Analytics System for Social Customer Relationship Management
Meena Nagarajan, Danish Contractor, Stephen Dill, Jitendra Ajmera, Hyung-Il Ahn, Ashish Verma 0001, Matthew Denesuk |
ICWSM | 6 |
| 2013 | A CRM system for social media: challenges and experiencesabstractThe social Customer Relationship Management (CRM) landscape is attracting significant attention from customers and enterprises alike as a sustainable channel for tracking, managing and improving customer relations. Enterprises are taking a hard look at this open, unmediated platform because the community effect generated on this channel can have a telling effect on their brand image, potential market opportunity and customer loyalty. In this work we present our experiences in building a system that mines conversations on social platforms to identify and prioritize those posts and messages that are relevant to enterprises. The system presented in this work aims to empower an agent or a representative in an enterprise to monitor, track and respond to customer communication while also encouraging community participation. Jitendra Ajmera, Hyung-Il Ahn, Meena Nagarajan, Ashish Verma 0001, Danish Contractor, Stephen Dill, Matthew Denesuk |
WWW | 4 |
| 2012 | Towards a domain-independent ASR-confidence classifierabstractThis work addresses the problem of developing a domain-independent binary classifier for a test domain given labeled data from several training domains where the test domain is not necessarily present in training data. The classifier accepts or rejects the ASR hypothesis based on the confidence generated by the ASR system. In the proposed approach, training data is grouped into across-domain clusters and separate cluster-specific classifiers are trained. One of the main findings is that the cluster purity and the normalized mutual information of the clusters are not very high which suggests that the domains might not necessarily be natural clusters. The performance of these cluster-specific classifiers is better than that of: (a) a single classifier trained on data from all the domains, and (b) a set of classifiers trained separately for each of the training domains. At an operating point corresponding to low False Accept, the Correct Accept of the proposed technique is on an average 2.3% higher than that obtained by the single-classifier or the individual train-domain classifiers. Om Deshmukh, Ashish Verma 0001, Etienne Marcheret |
ICASSP | 2 |
| 2012 | Finding Influential Authors in Brand-Page Communities
Hemant Purohit, Jitendra Ajmera, Sachindra Joshi, Ashish Verma 0001, Amit P. Sheth |
ICWSM | 4 |
| 2012 | Spoken Document Clustering Using Word Confusion Networks
Shajith Ikbal, Sachindra Joshi, Ashish Verma 0001, Om Deshmukh |
INTERSPEECH | 3 |
| 2012 | Cross-Guided Clustering: Transfer of Relevant Supervision across TasksabstractLack of supervision in clustering algorithms often leads to clusters that are not useful or interesting to human reviewers. We investigate if supervision can be automatically transferred for clustering a target task, by providing a relevant supervised partitioning of a dataset from a different source task. The target clustering is made more meaningful for the human user by trading-off intrinsic clustering goodness on the target task for alignment with relevant supervised partitions in the source task, wherever possible. We propose a cross-guided clustering algorithm that builds on traditional k-means by aligning the target clusters with source partitions. The alignment process makes use of a cross-task similarity measure that discovers hidden relationships across tasks. When the source and target tasks correspond to different domains with potentially different vocabularies, we propose a projection approach using pivot vocabularies for the cross-domain similarity measure. Using multiple real-world and synthetic datasets, we show that our approach improves clustering accuracy significantly over traditional k-means and state-of-the-art semi-supervised clustering baselines, over a wide range of data characteristics and parameter settings. Indrajit Bhattacharya, Shantanu Godbole, Sachindra Joshi, Ashish Verma 0001 |
ACM Trans. Knowl. Discov. Data | 4 |
| 2011 | Role of nucleus based context in word-independent syllable stress classificationabstractAn acoustic-phonetics based word-independent technique which uses syllable context for classifying the lexical syllable stress of spoken English words is presented. Nucleus based clustering is re markably successful in moving from word-dependent syllable stress classification which is intrinsically not scalable to word-independent classification. This however is not possible without an inherent drop in accuracy due to the loss of important contextual information of the syllables. An approach based on incorporating the left and the right context-ID of the syllable nucleus is proposed which results in a 10% improvement in word-level accuracy for word-independent syllable stress classification. The proposed approach exhibits performances comparable to that of the best performing word-dependent classifiers without suffering from the latter's scalability issues. A 7% improvement in the syllable level accuracy is also reported. Harish Doddala, Om Deshmukh, Ashish Verma 0001 |
ICASSP | 3 |
| 2011 | A Cross-Lingual Spoken Content Search System
Jitendra Ajmera, Ashish Verma 0001 |
INTERSPEECH | 2 |
| 2011 | Acoustic-Similarity Based Technique to Improve Concept Recognition
Om Deshmukh, Shajith Ikbal, Ashish Verma 0001, Etienne Marcheret |
INTERSPEECH | 3 |
| 2011 | A Language Independent Approach to Audio Search
Vikram Gupta, Jitendra Ajmera, Arun Kumar 0007, Ashish Verma 0001 |
INTERSPEECH | 4 |
| 2010 | Building re-usable dictionary repositories for real-world text miningabstractText mining, though still a nascent industry, has been growing quickly along with the awareness of the importance of unstructured data in business analytics, customer retention and extension, social media, and legal applications. There has been a recent increase in the number of commercial text mining product and service offerings, but successful or wide-spread deployments are rare, mainly due to a dependence on the expertise and skill of practitioners. Accordingly, there is a growing need for re-usable repositories for text mining. In this paper, we focus on dictionary-based text mining and its role in enabling practitioners in understanding and analyzing large text datasets. We motivate and define the problem of exploratory dictionary construction for capturing concepts of interest, and propose a framework for efficient construction, tuning, and re-use of these dictionaries across datasets. The construction framework offers a range of interaction modes to the user to quickly build concept dictionaries over large datasets. We also show how to adapt one or more dictionaries across domains and tasks, thereby enabling reuse of knowledge and effort in industrial practice. We present results and case studies on real-life CRM analytics datasets, where such repositories and tooling significantly cut down practitioner time and effort for dictionary-based text mining. Shantanu Godbole, Indrajit Bhattacharya, Ashish Verma 0001 |
CIKM | 4 |
| 2010 | Role of language models in spoken fluency evaluation
Om Deshmukh, Harish Doddala, Ashish Verma 0001, Karthik Visweswariah |
INTERSPEECH | 3 |
| 2010 | Utilizing relationships between named entities to improve speech recognition in dialog systemsabstractIn this paper, we address the problem of improving recognition accuracy of spoken named entities in the context of dialog systems for transactional applications. We propose utilizing the knowledge of relationships, that typically exist in many applications, between named entities spoken across different dialog states. For example, in a bank customer database each customer name is associated with one or a few account numbers, addresses and vice versa. We utilize these relationships to build long-term dependency constraints in grammars (and thus in decoding graphs) representing these entities. This enforces the recognizer to use collective evidences from instances of all the entities to improve the recognition accuracy of each individual entity. Experiments conducted to evaluate our approach show significant accuracy improvements on a task of recognizing a person via a name and a location. Shajith Ikbal, Om Deshmukh, Karthik Visweswariah, Ashish Verma 0001 |
SLT | 4 |
| 2009 | Formant-based technique for automatic filled-pause detection in spontaneous spoken englishabstractDetection of filled pauses is a challenging research problem which has several practical applications. It can be used to evaluate the spoken fluency skills of the speaker, to improve the performance of automatic speech recognition systems or to predict the mental state of the speaker. This paper presents an algorithm for filled pause detection that is based on the premise that the vocal tract characteristics, and hence the formants, are stable during the production of a filled pause. The performance of the proposed algorithm is evaluated on real-life recordings of call center agents where the locations of the filled pauses are hand labeled. The proposed algorithm outperforms a standard cepstral stability based filled pause detection algorithm and a standard pitch-based detection technique. Kartik Audhkhasi, Kundan Kandhway, Om Deshmukh, Ashish Verma 0001 |
ICASSP | 4 |
| 2009 | Automatic evaluation of spoken english fluencyabstractThis paper presents a method to automatically quantify the spoken English fluency skills of speakers. The focus of this work is to automatically compute a numeric score of spoken fluency that is correlated with the numerical score the human assessors would assign. The proposed method combines several novel prosodic and lexical features to compute the fluency score. It is shown that the prosodic and the lexical features provide complementary information for fluency evaluation. Extensive evaluation on human-labeled utterances shows that the proposed technique exhibits similar trends in performance and confusions as shown by human assessors. The proposed technique leads to 84.2% classification accuracy when the two extreme classes of fluency are considered. Om Deshmukh, Kundan Kandhway, Ashish Verma 0001, Kartik Audhkhasi |
ICASSP | 3 |
| 2009 | Cross-Guided Clustering: Transfer of Relevant Supervision across Domains for Improved ClusteringabstractLack of supervision in clustering algorithms often leads to clusters that are not useful or interesting to human reviewers. We investigate if supervision can be automatically transferred to a clustering task in a target domain, by providing a relevant supervised partitioning of a dataset from a different source domain. The target clustering is made more meaningful for the human user by trading off intrinsic clustering goodness on the target dataset for alignment with relevant supervised partitions in the source dataset, wherever possible. We propose a cross-guided clustering algorithm that builds on traditional k-means by aligning the target clusters with source partitions. The alignment process makes use of a cross-domain similarity measure that discovers hidden relationships across domains with potentially different vocabularies. Using multiple real-world datasets, we show that our approach improves clustering accuracy significantly over traditional k-means. Indrajit Bhattacharya, Shantanu Godbole, Sachindra Joshi, Ashish Verma 0001 |
ICDM | 4 |
| 2009 | Enabling Scaleable, Efficient, Non-visual Web Browsing ServicesabstractOver the last few decades, the discipline of Web accessibility has been focused on building more efficient and more effective speech generators for Web browsers. The visual browser interface is central to the current paradigm. However, in many cases, visual interaction is not required or desired, e.g. it is not relevant for blind people. More generally, when the input and output points are WAV files, SMS messages or natural queries, it becomes very clear that going through the visual user interface is overkill. In this paper we introduce a solution to this problem - a scalable, efficient, non-visual Web browser that works with a Web whose central assumption is that visual interaction is an integral part of the user experience. Ashish Verma 0001, Tyrone Grandison, Himanshu Chauhan |
ICWS | 1 |
| 2009 | Enabling analysts in managed services for CRM analyticsabstractData analytics tools and frameworks abound, yet rapid deployment of analytics solutions that deliver actionable insights from business data remains a challenge. The primary reason is that on-field practitioners are required to be both technically proficient and knowledgeable about the business. The recent abundance of unstructured business data has thrown up new opportunities for analytics, but has also multiplied the deployment challenge, since interpretation of concepts derived from textual sources require a deep understanding of the business. In such a scenario, a managed service for analytics comes up as the best alternative. A managed analytics service is centered around a business analyst who acts as a liaison between the business and the technology. This calls for new tools that assist the analyst to be efficient in the tasks that she needs to execute. Also, the analytics needs to be repeatable, in that the delivered insights should not depend heavily on the expertise of specific analysts. These factors lead us to identify new areas that open up for KDD research in terms of 'time-to-insight' and repeatability for these analysts. We present our analytics framework in the form of a managed service offering for CRM analytics. We describe different analyst-centric tools using a case study from real-life engagements and demonstrate their effectiveness. Indrajit Bhattacharya, Shantanu Godbole, Ashish Verma 0001, Jeff Achtermann, Kevin English |
KDD | 4 |
| 2009 | Nucleus-level clustering for word-independent syllable stress classification
Om Deshmukh, Ashish Verma 0001 |
Speech Commun. | 2 |
| 2008 | Automatic pronunciation evaluation and classification
Om Deshmukh, Sachindra Joshi, Ashish Verma 0001 |
INTERSPEECH | 3 |
| 2008 | Acoustic-phonetic approach for automatic evaluation of spoken grammar
Om Deshmukh, Ashish Verma 0001 |
INTERSPEECH | 2 |
| 2007 | Sensei: Spoken language assessment for call center agentsabstractIn this paper, we present a system, called Sensei, for assessment of spoken English skills of call center agents. Sensei evaluates multiple parameters of spoken English skills, i.e., articulation of sounds, correctness of lexical stress in words and spoken grammar proficiency. Sensei provides an assessment test to be taken by a call center agent (or candidate) and generates score on each of the spoken English parameters as well as a combined score. It is implemented in the form of a web application so that it can be accessed through a web browser and doesn’t require any software to be installed at the client side. We describe how the individual parameters are assessed in Sensei using various speech processing techniques and the experiments conducted to evaluate these techniques. The performance is compared with assessment performed by human assessors. A correlation of 0.8 is obtained between overall score generated by Sensei and human assessors on a real life test dataset of 243 candidates which compares well with the corresponding human-to-human correlation of 0.91. Abhishek Chandel, Abhinav Parate, Maymon Madathingal, Himanshu Pant, Nitendra Rajput, Shajith Ikbal, Om Deshmukh, Ashish Verma 0001 |
ASRU | 8 |
| 2007 | Keyword Search using Modified Minimum Edit Distance MeasureabstractA popular approach for keyword search in speech files is the phone lattice search. Recently minimum edit distance (MED) has been used as a measure of similarity between strings rather than using simple string matching while searching the phone lattice for the keyword. In this paper, we propose a variation of the MED, where the substitution penalties are automatically derived from the phone confusion matrix of the recognizer, as compared to heuristic or class based penalties used earlier. The results show that the substitution penalties derived from the phone confusion matrix lead to a considerable improvement in the accuracy of the keyword search algorithm. Kartik Audhkhasi, Ashish Verma 0001 |
ICASSP (4) | 2 |
| 2007 | Evaluation of syllable stress using single class classifier
Abhinav Parate, Ashish Verma 0001, Jayanta Basak |
INTERSPEECH | 2 |
| 2007 | Language identification of person names using CF-IOF based weighing functionabstractInformation about the language of origin helps in generating pronunciation for foreign words, specially person names, in a text-to-speech synthesis system. It can be used to apply language specific letter-to-sound (LTS) rules to these words during synthesis. In this paper, we propose a novel approach for using substrings of a person name (called letter N-grams) to identify the language of its origin. We use a weight for the letter N-grams that is motivated by the techniques used in text document classification, different from the usual N-gram probabilities used in earlier approaches. We also propose a tree based approach to select the letter N-grams of different lengths for language identification. Several experiments have been conducted to evaluate the performance of the proposed approach and compare it with those of the earlier proposed approaches based on N-gram probabilities. We show an improvement in classification results over the earlier approaches without using any language specific rules. Samuel Thomas 0001, Ashish Verma 0001 |
INTERSPEECH | 2 |
| 2006 | Word Independent Model for Syllable Stress EvaluationabstractAnalyzing syllable stress in spoken English has been an area of research for a long time. In this paper, we analyze the performance of a novel method for evaluating syllable stress in spoken English. Specifically, we study the problem of determining if a word is spoken with the correct syllable stress pattern. The proposed method uses generalized models for stressed and unstressed syllables to analyze the constituent syllables of a word and determines if the word is spoken correctly. The performance of the proposed method is reported in terms of classification results on human labeled word utterances and it is compared with that of the word-dependent models using various classifiers. Ashish Verma 0001, Kunal Lal, Yuen Yee Lo, Jayanta Basak |
ICASSP (1) | 1 |
| 2005 | Introducing Roughness in Individuality Transformation through Jitter Modeling and ModificationabstractIndividuality transformation is a process to modify the speech signal in a person's voice so that it sounds as if it is spoken by another person. In most individuality transformation methods, pitch transformation is performed through a simple scaling considering the global pitch characteristics of the source and target speakers without considering the short-term pitch variation or jitter. We present a novel method to model and modify jitter in the speech signal to introduce a handle on roughness in the process of individuality transformation. The proposed method is based upon computing the average intensity in a band around the fundamental frequency in the spectrum of a speaker's mean normalized pitch contour. The validity of the proposed method to model jitter has been established by subjective tests for perceived roughness in the speaker's voice. It is also shown that modification of jitter by the proposed method results in an improved subjective rating for individuality transformation. Ashish Verma 0001, Arun Kumar 0007 |
ICASSP (1) | 1 |
| 2004 | Articulatory class based spectral envelope representation for voice fontsabstractVoice fonts are used to represent and transform the personality of speech. The spectral envelope is one of the main descriptors of voice fonts. We present a new technique to represent the spectral envelope in voice fonts which makes it possible to create a universal voice font for a given speaker which can be used for multiple languages. It enables voice fonts to be used for translingual personality transformation of speech. We evaluate the performance of the proposed approach through monolingual and translingual personality transformation experiments. We also compare the performance with the phoneme class based spectral envelope representation in voice fonts through various objective and subjective tests. Ashish Verma 0001, Arun Kumar 0007 |
ICME | 1 |
| 2004 | Animating expressive faces across languagesabstractThis paper describes a morphing-based audio driven facial animation system. Based on an incoming audio stream, a face image is animated with full lip synchronization and synthesized expressions. A novel scheme to implement a language independent system for audio-driven facial animation given a speech recognition system for just one language, in our case, English, is presented. The method presented here can also be used for text to audio-visual speech synthesis. Visemes in new expressions are synthesized to be able to generate animations with different facial expressions. An animation sequence using optical flow between visemes is constructed, given an incoming audio stream and still pictures of a face representing different visemes. The presented techniques give improved lip synchronization and naturalness to the animated video. Ashish Verma 0001, L. Venkata Subramaniam, Nitendra Rajput, Chalapathy Neti, Tanveer A. Faruquie |
IEEE Trans. Multim. | 1 |
| 2003 | Using phone and diphone based acoustic models for voice conversion: a step towards creating voice fontsabstractVoice conversion techniques attempt to modify the speech signal so that it is perceived as if spoken by another speaker, different from the original speaker. In this paper, we present a novel approach to perform voice conversion. Our approach uses acoustic models based on units of speech, like phones and diphones, for voice conversion. These models can be computed and used independently for a given speaker without being concerned about the source or target speaker. It avoids the use of a parallel speech corpus in the voices of source and target speakers. It is shown that by using the proposed approach, voice fonts can be created and stored which represent individual characteristics of a particular speaker, to be used for customization of synthetic speech. We also show through objective and subjective tests, that voice conversion quality is comparable to other approaches that require a parallel speech corpus. Arun Kumar 0007, Ashish Verma 0001 |
ICASSP (1) | 2 |
| 2003 | Using viseme based acoustic models for speech driven lip synthesisabstractSpeech driven lip synthesis is an interesting and important step toward human-computer interaction. An incoming speech signal is time aligned using a speech recognizer to generate a phonetic sequence which is then converted to the corresponding viseme sequence to be animated. We present a novel method for generation of the viseme sequence, which uses viseme based acoustic models, instead of the usual phone based acoustic models, to align the input speech signal. This results in higher accuracy and speed of the alignment procedure and allows a much simpler implementation of the speech driven lip synthesis system as it completely obviates the requirement of an acoustic unit to visual unit conversion. We show, through various experiments, that the proposed method results in about 53% relative improvement in classification accuracy and about 52% reduction in the time required to compute alignments. Ashish Verma 0001, Nitendra Rajput, L. Venkata Subramaniam |
ICASSP (5) | 1 |
| 2003 | Using phone and diphone based acoustic models for voice conversion: a step towards creating voice fontsabstractVoice conversion techniques attempt to modify speech signal so that it is perceived as if spoken by another speaker, different from the original speaker. In this paper, we present a novel approach to perform voice conversion. Our approach uses acoustic models based on units of speech, like phones and diphones, for voice conversion. These models can be computed and used independently for a given speaker without being concerned about the source or target speaker. It avoids the use of a parallel speech corpus in the voices of source and target speakers. It is shown that by using the proposed approach, voice fonts can be created and stored which will represent individual characteristics of a particular speaker, to be used for customization of synthetic speech. We also show through objective and subjective tests, that voice conversion quality is comparable to other approaches that require a parallel speech corpus. Arun Kumar 0007, Ashish Verma 0001 |
ICME | 2 |
| 2003 | Using viseme based acoustic models for speech driven lip synthesisabstractSpeech drive lip synthesis is an interesting and important step toward human-computer interaction. An incoming speech signal is time aligned using a speech recognizer to generate phonetic sequence, which is then converted to corresponding viseme sequence to be animated. In this paper, we present a novel method for generation of the viseme sequence, which uses viseme based acoustic models, instead of usual phone based acoustic models, to align the input speech signal. This results in higher accuracy and speed of the alignment procedure and allows a much simpler implementation of the speech driven lip synthesis system as it completely obviates the requirement of acoustic unit to visual unit conversion. We show through various experiments that the proposed method results in about 53% relative improvement in classification accuracy and about 52% reduction in time, required to compute alignments. Ashish Verma 0001, Nitendra Rajput, L. Venkata Subramaniam |
ICME | 1 |
| 2003 | Modeling speaking rate for voice fonts
Ashish Verma 0001, Arun Kumar 0007 |
INTERSPEECH | 1 |
| 2001 | Using likelihood L-statistics to measure confidence in audio-visual speech recognitionabstractThis paper describes previous work on decision fusion in audio-visual speech recognition. A novel approach is proposed to combine audio and video channel information in audio-visual speech recognition scenario. We have considered frame-level phonetic classification problem using two single-stream Gaussian mixture models. Audio and video streams are adaptively weighted using a cumulative mean of the sample confidence values over past frames in addition to the present sample confidence value. The confidence values for audio and video decisions are computed using an L-statistics (linear combination of order-statistics) of log-likelihoods against phone models. It is shown through various experiments, on a database of about 15000 sentences from large vocabulary continuous speech, that the proposed approach results in better classification accuracy as compared to other approaches. Arpita Ghosh, Ashish Verma 0001, Abhinanda Sarkar |
MMSP | 2 |
| 2001 | Robust detection of visual ROI for automatic speechreadingabstractWe present our work on visual pruning in an audio-visual (AV) speech recognition scenario. Visual speech information has been successfully used in circumstances where audio-only recognition suffers (e.g. noisy environments). Tracking and extraction of region-of-interest (ROI) (e.g., speaker's mouth region) from video is an essential component of such systems. It is important for the visual front-end to handle tracking errors that result in noisy visual data and hamper performance. We present our robust visual front-end, investigate methods to prune visual noise and its effect on the performance of the AV speech recognition systems. Specifically, we estimate the "goodness of ROI" using Gaussian mixture models and our experiments indicate that significant performance gains are achieved with good quality visual data. Giridharan Iyengar, Gerasimos Potamianos, Chalapathy Neti, Tanveer A. Faruquie, Ashish Verma 0001 |
MMSP | 5 |
| 2000 | On deriving a phoneme model for a new languageabstractWe present a method for building an initial phoneme model for training an HMM in a new language using an already trained recognition system in a base language. HMM based phoneme recognition systems are used to model the phonemes in most large vocabulary speech recognition tasks. Mappings between the phonetic spaces of the two languages are generated and are used to populate the phonetic space of the new language. The best possible alignment of the new language data is obtained and initial phone models are built on this labeled data. A classification experiment is performed in the new language to illustrate the goodness of initial phone models. Experiments are carried out with Hindi as the new language using an English language recognition system to derive the initial phone models for Hindi language. 1. Niloy Mukherjee, Nitendra Rajput, L. Venkata Subramaniam, Ashish Verma 0001 |
INTERSPEECH | 4 |
| 2000 | Adapting phonetic decision trees between languages for continuous speech recognitionabstractIn a continuous speech recognition system it is important to model the context dependent variations in the pronunciations of phones. In this work we have attempted to build decision trees for modeling phonetic context-dependency in Hindi. The approach followed is to modify a decision tree built to model context-dependency in American English. The reason the decision trees turn out to be different are that the English and Hindi phoneme sets are not identical. Then even for identical phonemes, the context-dependency is different for the two languages. Linguistic-Phonetic knowledge of Hindi is used to modify the English phone set. Since the Hindi phone set being used is derived from the English phone set, the adaptation of the English tree to Hindi follows naturally. Though here the adaptation is from English to Hindi, the method may be applicable for adapting between any two languages. The decision tree is built using either Hindi data or English data labeled with the correct Hindi contexts. This procedure is discussed and the limitations of both the methods are described. 1. Nitendra Rajput, L. Venkata Subramaniam, Ashish Verma 0001 |
INTERSPEECH | 3 |
| 1999 | Audio-visual large vocabulary continuous speech recognition in the broadcast domainabstractConsiders the problem of combining visual cues with audio signals for the purpose of improved automatic machine recognition of speech. Although significant progress has been made in the machine transcription of large-vocabulary continuous speech (LVCSR) over the last few years, the technology to date is most effective only under controlled conditions, such as low noise, speaker-dependent recognition, read speech (as opposed to conversational speech), etc. On the other hand, while augmenting the recognition of speech utterances with visual cues has attracted the attention of researchers over the last couple of years, most efforts in this domain can be considered to be only preliminary in the sense that, unlike LVCSR efforts, tasks have been limited to small vocabularies (e.g. commands, digits) and often to speaker-dependent training or isolated word speech, where word boundaries are artificially well-defined. Sankar Basu, Chalapathy Neti, Nitendra Rajput, Andrew W. Senior, L. Venkata Subramaniam, Ashish Verma 0001 |
MMSP | 6 |