EDBT 2026 Demo / reviewers in the wild / expert
Karthik Visweswariah
dblp:32/698
· DBLP profile ↗
63ranked-venue papers
15as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 9 first-authorGraphics, computer vision, multimedia, augmented reality and games · 36 · 8 first-authorDatabases, data management, data science and information retrieval · 7Theory of computation · 5 · 4 first-authorSoftware engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Speech recognition and synthesis · 57% Machine translation · 41% Probabilistic and Bayesian machine learning · 2% | |
| Theoretical computer science
5 papers |
Coding theory · 90% Information theory · 10% | |
| Software engineering, system software, and programming languages
1 paper |
Empirical software engineering · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 77% Data mining · 23% | |
| Computer graphics and multimedia
1 paper |
Audio and music processing · 100% |
Topics — the 25 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
question answering |
0.2 | 1 | 2014 | Unsupervised Solution Post Identification from Discussion Forums · ACL (1) 2014 |
Coding theory
source coding |
0.1 | 5 | 2002 | Universal lossless source coding with the Burrows Wheeler Transform · IEEE Trans. Inf. Theory 2002 Universal variable-to-fixed length source codes · IEEE Trans. Inf. Theory 2001 Separation of random number generation and resolvability · IEEE Trans. Inf. Theory 2000 |
Natural language and speech › Speech recognition and synthesis
acoustic modeling |
0.1 | 2 | 2007 | Discriminative Estimation of Subspace Constrained Gaussian Mixture Models for Speech Recognition · IEEE Trans. Speech Audio Process. 2007 Subspace constrained Gaussian mixture models for speech recognition · IEEE Trans. Speech Audio Process. 2005 |
Natural language and speech › Speech recognition and synthesis
automatic speech recognition |
0.1 | 2 | 2007 | Discriminative Estimation of Subspace Constrained Gaussian Mixture Models for Speech Recognition · IEEE Trans. Speech Audio Process. 2007 Subspace constrained Gaussian mixture models for speech recognition · IEEE Trans. Speech Audio Process. 2005 |
Natural language and speech › Speech recognition and synthesis › acoustic modeling
subspace constrained gaussian mixture models |
0.1 | 2 | 2007 | Discriminative Estimation of Subspace Constrained Gaussian Mixture Models for Speech Recognition · IEEE Trans. Speech Audio Process. 2007 Subspace constrained Gaussian mixture models for speech recognition · IEEE Trans. Speech Audio Process. 2005 |
Natural language and speech › Machine translation
word reordering |
0.1 | 1 | 2011 | A Word Reordering Model for Improved Machine Translation · EMNLP 2011 |
Empirical software engineering
developer studies |
0.1 | 1 | 2010 | Timesheet assistant: mining and reporting developer effort · ASE 2010 |
Empirical software engineering
mining software repositories |
0.1 | 1 | 2010 | Timesheet assistant: mining and reporting developer effort · ASE 2010 |
Coding theory › source coding
universal coding |
0.1 | 3 | 2002 | Universal lossless source coding with the Burrows Wheeler Transform · IEEE Trans. Inf. Theory 2002 Universal variable-to-fixed length source codes · IEEE Trans. Inf. Theory 2001 Universal coding of nonstationary sources · IEEE Trans. Inf. Theory 2000 |
Natural language and speech › Speech recognition and synthesis › acoustic modeling
discriminative acoustic model training |
0.1 | 1 | 2007 | Discriminative Estimation of Subspace Constrained Gaussian Mixture Models for Speech Recognition · IEEE Trans. Speech Audio Process. 2007 |
Audio and music processing › speech recognition
acoustic modeling |
0.1 | 1 | 2006 | Gaussian mixture models with covariances or precisions in shared multiple subspaces · IEEE Trans. Speech Audio Process. 2006 |
Audio and music processing
speech recognition |
0.1 | 1 | 2006 | Gaussian mixture models with covariances or precisions in shared multiple subspaces · IEEE Trans. Speech Audio Process. 2006 |
Data mining
clustering |
0.1 | 1 | 2014 | Unsupervised Solution Post Identification from Discussion Forums · ACL (1) 2014 |
Information theory
random number generation |
0.0 | 2 | 2000 | Separation of random number generation and resolvability · IEEE Trans. Inf. Theory 2000 Source Codes as Random Number Generators · IEEE Trans. Inf. Theory 1998 |
Natural language and speech › Machine translation
neural machine translation |
0.0 | 1 | 2011 | A Word Reordering Model for Improved Machine Translation · EMNLP 2011 |
Coding theory › source coding
burrows-wheeler transform |
0.0 | 1 | 2002 | Universal lossless source coding with the Burrows Wheeler Transform · IEEE Trans. Inf. Theory 2002 |
Coding theory › source coding
lossless compression |
0.0 | 1 | 2002 | Universal lossless source coding with the Burrows Wheeler Transform · IEEE Trans. Inf. Theory 2002 |
Empirical software engineering
software project management |
0.0 | 1 | 2010 | Timesheet assistant: mining and reporting developer effort · ASE 2010 |
Coding theory › source coding › source modeling
markov sources |
0.0 | 1 | 2001 | Universal variable-to-fixed length source codes · IEEE Trans. Inf. Theory 2001 |
Coding theory › source coding
variable-to-fixed length codes |
0.0 | 1 | 2001 | Universal variable-to-fixed length source codes · IEEE Trans. Inf. Theory 2001 |
Coding theory › source coding
lempel-ziv compression |
0.0 | 1 | 2000 | Universal coding of nonstationary sources · IEEE Trans. Inf. Theory 2000 |
Coding theory › source coding › source modeling
nonstationary source |
0.0 | 1 | 2000 | Universal coding of nonstationary sources · IEEE Trans. Inf. Theory 2000 |
Coding theory › source coding
resolvability |
0.0 | 1 | 2000 | Separation of random number generation and resolvability · IEEE Trans. Inf. Theory 2000 |
Audio and music processing › speech recognition
hidden markov model |
0.0 | 1 | 2006 | Gaussian mixture models with covariances or precisions in shared multiple subspaces · IEEE Trans. Speech Audio Process. 2006 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
hidden markov model |
0.0 | 1 | 2005 | Subspace constrained Gaussian mixture models for speech recognition · IEEE Trans. Speech Audio Process. 2005 |
Methods — techniques the papers use, named apart from their topics
expectation-maximization · 0.2translation model · 0.2language model · 0.2word reordering model · 0.1maximum likelihood estimation · 0.1statistical analysis · 0.1maximum mutual information · 0.1error-weighted training · 0.1subspace sharing · 0.1factor analysis · 0.1redundancy analysis · 0.0dictionary-free coding · 0.0lempel-ziv incremental parsing · 0.0achievability and converse bounds · 0.0lempel-ziv algorithm · 0.0entropy rate analysis · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | Unsupervised Solution Post Identification from Discussion ForumsabstractDiscussion forums have evolved into a dependable source of knowledge to solve common problems.However, only a minority of the posts in discussion forums are solution posts.Identifying solution posts from discussion forums, hence, is an important research problem.In this paper, we present a technique for unsupervised solution post identification leveraging a so far unexplored textual feature, that of lexical correlations between problems and solutions.We use translation models and language models to exploit lexical correlations and solution post character respectively.Our technique is designed to not rely much on structural features such as post metadata since such features are often not uniformly available across forums.Our clustering-based iterative solution identification approach based on the EM-formulation performs favorably in an empirical evaluation, beating the only unsupervised solution identification technique from literature by a very large margin.We also show that our unsupervised technique is competitive against methods that require supervision, outperforming one such technique comfortably. Deepak P 0001, Karthik Visweswariah |
ACL (1) | 2 |
| 2014 | When Transliteration Met Crowdsourcing : An Empirical Study of Transliteration via Crowdsourcing using Efficient, Non-redundant and Fair Quality Control
Mitesh M. Khapra, Ananthakrishnan Ramanathan, Anoop Kunchukuttan, Karthik Visweswariah, Pushpak Bhattacharyya |
LREC | 4 |
| 2013 | Cut the noise: Mutually reinforcing reordering and alignments for improved machine translation
Karthik Visweswariah, Mitesh M. Khapra, Ananthakrishnan Ramanathan |
ACL (1) | 1 |
| 2013 | Efficient multifaceted screening of job applicantsabstractBuilt on top of human resources management databases within the enterprise, we present a decision support system for managing and optimizing screening activities during the hiring process in a large organization. The basic idea is to prioritize the efforts of human resource practitioners to focus on candidates that are likely of high quality, that are likely to accept a job offer if made one, and that are likely to remain with the organization for the long term. To do so, the system first individually ranks candidates along several dimensions using a keyword matching algorithm and several bipartite ranking algorithms with univariate loss trained on historical actions. Next, individual rankings are aggregated to derive a single list that is presented to the recruitment team through an interactive portal. The portal supports multiple filters that facilitate effective identification of candidates. We demonstrate the usefulness of our system on data collected from a large organization over several years with business value metrics showing greater hiring yield with less interviews. Similarly, using historical pre-hire data we demonstrate accurate identification of candidates that will have quickly left the organization. The system has been deployed as described in a large globally integrated enterprise. Sameep Mehta, Rakesh Pimplikar, Amit Singh 0003, Lav R. Varshney, Karthik Visweswariah |
EDBT | 5 |
| 2013 | Semi-Supervised Answer Extraction from Discussion Forums
Rose Catherine, Rashmi Gangadharaiah, Karthik Visweswariah, Dinesh Raghu |
IJCNLP | 3 |
| 2013 | Improving reordering performance using higher order and structural features
Mitesh M. Khapra, Ananthakrishnan Ramanathan, Karthik Visweswariah |
HLT-NAACL | 3 |
| 2012 | Two-part segmentation of text documentsabstractWe consider the problem of segmenting text documents that have a two-part structure such as a problem part and a solution part. Documents of this genre include incident reports that typically involve description of events relating to a problem followed by those pertaining to the solution that was tried. Segmenting such documents into the component two parts would render them usable in knowledge reuse frameworks such as Case-Based Reasoning. This segmentation problem presents a hard case for traditional text segmentation due to the lexical inter-relatedness of the segments. We develop a two-part segmentation technique that can harness a corpus of similar documents to model the behavior of the two segments and their inter-relatedness using language models and translation models respectively. In particular, we use separate language models for the problem and solution segment types, whereas the inter-relatedness between segment types is modeled using an IBM Model 1 translation model. We model documents as being generated starting from the problem part that comprises of words sampled from the problem language model, followed by the solution part whose words are sampled either from the solution language model or from a translation model conditioned on the words already chosen in the problem part. We show, through an extensive set of experiments on real-world data, that our approach outperforms the state-of-the-art text segmentation algorithms in the accuracy of segmentation, and that such improved accuracy translates well to improved usability in Case-based Reasoning systems. We also analyze the robustness of our technique to varying amounts and types of noise and empirically illustrate that our technique is quite noise tolerant, and degrades gracefully with increasing amounts of noise. Deepak P 0001, Karthik Visweswariah, Nirmalie Wiratunga, Sadiq Sani |
CIKM | 2 |
| 2012 | A Comparison of Syntactic Reordering Methods for English-German Machine Translation
Jirí Navrátil 0001, Karthik Visweswariah, Ananthakrishnan Ramanathan |
COLING | 2 |
| 2012 | A Study of Word-Classing for MT Reordering
Ananthakrishnan Ramanathan, Karthik Visweswariah |
LREC | 2 |
| 2011 | Privacy protected knowledge management in services with emphasis on quality dataabstractImproving productivity of practitioners through effective knowledge management and delivering high quality service in Application Management Services (AMS) domain, are key focus areas for all IT services organizations. One source of historical knowledge in AMS is the large amount of resolved problem ticket data which are often confidential, immensely valuable, but majority of it is of very bad quality. In this paper, we present a knowledge management tool that detects the quality of information present in problem tickets and enables effective knowledge search in tickets by prioritizing quality data in the search ranking. The tool facilitates leveraging of knowledge across different AMS accounts, while preserving data privacy, by masking client confidential information. It also extracts several relevant entities contained in the noisy unstructured text entered in the tickets and presents them to the users. We present several experimental evaluations and a pilot study conducted with an AMS account which show that our tool is effective and leads to substantial improvement in productivity of the practitioners. Debapriyo Majumdar, Rose Catherine, Shajith Ikbal, Karthik Visweswariah |
CIKM | 4 |
| 2011 | CQC: classifying questions in CQA websitesabstractCommunity Question Answering portals like Yahoo! Answers have recently become a popular method for seeking information online. Users express their information need as questions for which other users generate potential answers. These questions are organized into pre-defined hierarchical categories to facilitate effective answering, hence Question Classification is an important aspect of these systems. In this paper we propose a novel system, CQC, for automatically classifying new questions into one of the hierarchical categories. Experiments conducted on large scale real data from Yahoo Answers! show that the proposed techniques are effective and outperform existing methods significantly. Amit Singh 0003, Karthik Visweswariah |
CIKM | 2 |
| 2011 | A Word Reordering Model for Improved Machine Translation
Karthik Visweswariah, Rajakrishnan Rajkumar, Ankur Gandhe, Ananthakrishnan Ramanathan, Jirí Navrátil 0001 |
EMNLP | 1 |
| 2011 | Front-end feature transforms with context filtering for speaker adaptationabstractFeature-space transforms such as feature-space maximum likelihood linear regression (FMLLR) are very effective speaker adaptation technique, especially on mismatched test data. In this study, we extend the full-rank square matrix of FMLLR to a non-square matrix that uses neighboring feature vectors in estimating the adapted central feature vector. Through optimizing an appropriate objective function we aim to filter out and transform features through the correlation of the feature context. We compare to FMLLR that just con sider the current feature vector only. Our experiments are conducted on the automobile data with different speed conditions. Results show that context filtering improves 23% on word error rate over conventional FMLLR on noisy 60mph data with adapted ML model, and 7%/9% improvement over the discriminatively trained FMMI/BMMI models. Jing Huang 0019, Karthik Visweswariah, Peder A. Olsen, Vaibhava Goel |
ICASSP | 2 |
| 2011 | Handling verb phrase morphology in highly inflected Indian languages for Machine Translation
Ankur Gandhe, Rashmi Gangadharaiah, Karthik Visweswariah, Ananthakrishnan Ramanathan |
IJCNLP | 3 |
| 2011 | Clause-Based Reordering Constraints to Improve Statistical Machine Translation
Ananthakrishnan Ramanathan, Pushpak Bhattacharyya, Karthik Visweswariah, Kushal Ladha, Ankur Gandhe |
IJCNLP | 3 |
| 2010 | PROSPECT: a system for screening candidates for recruitmentabstractCompanies often receive thousands of resumes for each job posting and employ dedicated screeners to short list qualified applicants. In this paper, we present PROSPECT, a decision support tool to help these screeners shortlist resumes efficiently. Prospect mines resumes to extract salient aspects of candidate profiles like skills, experience in each skill, education details and past experience. Extracted information is presented in the form of facets to aid recruiters in the task of screening. We also employ Information Retrieval techniques to rank all applicants for a given job opening. In our experiments we show that extracted information improves our ranking by 30% there by making screening task simpler and more efficient. Amit Singh 0003, Rose Catherine, Karthik Visweswariah, Vijil Chenthamarakshan, Nanda Kambhatla |
CIKM | 3 |
| 2010 | Syntax Based Reordering with Automatically Derived Rules for Improved Statistical Machine Translation
Karthik Visweswariah, Jirí Navrátil 0001, Jeffrey S. Sorensen, Vijil Chenthamarakshan, Nanda Kambhatla |
COLING | 1 |
| 2010 | Role of language models in spoken fluency evaluation
Om Deshmukh, Harish Doddala, Ashish Verma 0001, Karthik Visweswariah |
INTERSPEECH | 4 |
| 2010 | Timesheet assistant: mining and reporting developer effortabstractTimesheets are an important instrument used to track time spent by team members in a software project on the tasks assigned to them. In a typical project, developers fill timesheets manually on a periodic basis. This is often tedious, time consuming and error prone. Over or under reporting of time spent on tasks causes errors in billing development costs to customers and wrong estimation baselines for future work, which can have serious business consequences. In order to assist developers in filling their timesheets accurately, we present a tool called Timesheet Assistant (TA) that non-intrusively mines developer activities and uses statistical analysis on historical data to estimate the actual effort the developer may have spent on individual assigned tasks. TA further helps the developer or project manager by presenting the details of the activities along with effort data so that the effort may be seen in the context of the actual work performed. We report on an empirical study of TA in a software maintenance project at IBM that provides preliminary validation of its feasibility and usefulness. Some of the limitations of the TA approach and possible ways to address those are also discussed. Renuka Sindhgatta, Nanjangud C. Narendra, Bikram Sengupta, Karthik Visweswariah, Arthur G. Ryman |
ASE | 4 |
| 2010 | Utilizing relationships between named entities to improve speech recognition in dialog systemsabstractIn this paper, we address the problem of improving recognition accuracy of spoken named entities in the context of dialog systems for transactional applications. We propose utilizing the knowledge of relationships, that typically exist in many applications, between named entities spoken across different dialog states. For example, in a bank customer database each customer name is associated with one or a few account numbers, addresses and vice versa. We utilize these relationships to build long-term dependency constraints in grammars (and thus in decoding graphs) representing these entities. This enforces the recognizer to use collective evidences from instances of all the entities to improve the recognition accuracy of each individual entity. Experiments conducted to evaluate our approach show significant accuracy improvements on a task of recognizing a person via a name and a location. Shajith Ikbal, Om Deshmukh, Karthik Visweswariah, Ashish Verma 0001 |
SLT | 3 |
| 2010 | Call transcript segmentation using word cooccurrence modelabstractIn this paper, we propose a word cooccurrence model to perform topic segmentation of call center conversational speech. This model is estimated from training data to discriminatively represent how likely various pairs of words are to cooccur within homogeneous topic segments. We show that such model provide an effective measure of lexical cohesion and hence provide useful evidence of topical coherence or lack thereof between various parts of the call transcripts. We propose two approaches of utilizing such evidence for segmentation: 1) An efficient dynamic programming algorithm to perform segmentation simply utilizing the word cooccurrence model. 2) Extracting features based on word cooccurrence model to utilize them as additional features in conditional random field (CRF) based segmentation. Experimental evaluation of these approaches against state-of-the-art approaches show the effectiveness of word cooccurrence model for the topic segmentation task. Shajith Ikbal, Karthik Visweswariah |
SLT | 2 |
| 2009 | Improved decision trees for multi-stream HMM-based audio-visual continuous speech recognitionabstractHMM-based audio-visual speech recognition (AVSR) systems have shown success in continuous speech recognition by combining visual and audio information, especially in noisy environments. In this paper we study how to improve decision trees used to create context classes in HMM-based AVSR systems. Traditionally, visual models have been trained with the same context classes as the audio only models. In this paper we investigate the use of separate decision trees to model the context classes for the audio and visual streams independently. Additionally we investigate the use of viseme classes in the decision tree building for the visual stream. On experiments with a 37-speaker 1.5 hours test set (about 12000 words) of continuous digits in noise, we obtain about a 3% absolute (20% relative) gain on AVSR performance by using separate decision trees for the audio and visual streams when using viseme classes in decision tree building for the visual stream. Jing Huang 0019, Karthik Visweswariah |
ASRU | 2 |
| 2009 | Combined discriminative training for multi-stream HMM-based audio-visual speech recognition
Jing Huang 0019, Karthik Visweswariah |
INTERSPEECH | 2 |
| 2009 | CAESAR: A Context-Aware, Social Recommender System for Low-End Mobile DevicesabstractMobile-enabled social networks applications are becoming increasingly popular. Most of the current social network applications have been designed for high-end mobile devices, and they rely upon features such as GPS, capabilities of the world wide web, and rich media support. However, a significant fraction of mobile user base, especially in the developing world, own low-end devices that are only capable of voice and short text messages (SMS). In this context, a natural question is whether one can design meaningful social network-based applications that can work well with these simple devices, and if so, what the real challenges are. Towards answering these questions, this paper presents a social network-based recommender system that has been explicitly designed to work even with devices that just support phone calls and SMS. Our design of the social network based recommender system incorporates three features that complement each other to derive highly targeted ads. First, we analyze information such as customer's address books to estimate the level of social affinity among various users. This social affinity information is used to identify the recommendations to be sent to an individual user. Second, we combine the social affinity information with the spatio-temporal context of users and historical responses of the user to further refine the set of recommendations and to decide when a recommendation would be sent. Third, social affinity computation and spatio-temporal contextual association are continuously tuned through user feedback. We outline the challenges in building such a system, and outline approaches to deal with such challenges. Lakshmish Ramaswamy, Deepak P 0001, Ramana Polavarapu, Kutila Gunasekera, Dinesh Garg, Karthik Visweswariah, Shivkumar Kalyanaraman |
Mobile Data Management | 6 |
| 2008 | Semi-automated logging of contact center telephone callsabstractModern businesses use contact centers as a communication channel with users of their products and services. The largest factor in the expense of running a telephone contact center is the labor cost of its agents. IBM Research has built a new system, Contact-Center Agent Buddies (CAB), which is designed to help reduce the average handle time (AHT) for customer calls, thereby also reducing their cost. In this paper, we focus on the call logging subsystem, which helps agents reduce the time they spend documenting those calls. We built a Template CAB and a Call Logging CAB, using a pipeline consisting of audio capture of a telephone conversation, automatic speech recognition, text analysis, and log generation. We developed techniques for ASR text cleansing, including normalization of expressions and acronyms, domain terms, capitalization, and boundaries for sentences, paragraphs, and call segments. We found that simple heuristics suffice to generate high-quality logs from the normalized sentences. The pipeline yields a candidate call log which the agents can edit in less time than it takes them to generate call logs manually. Evaluation of the Call Logging CAB in an industrial contact center environment shows that it reduces the amount of time agents spend logging calls by at least 50% without compromising the quality of the resulting call documentation. Roy J. Byrd, Mary S. Neff, Wilfried Teiken, Youngja Park, Keh-Shin F. Cheng, Stephen C. Gates, Karthik Visweswariah |
CIKM | 7 |
| 2008 | Boosted MMI for model and feature-space discriminative trainingabstractWe present a modified form of the maximum mutual information (MMI) objective function which gives improved results for discriminative training. The modification consists of boosting the likelihoods of paths in the denominator lattice that have a higher phone error relative to the correct transcript, by using the same phone accuracy function that is used in Minimum Phone Error (MPE) training. We combine this with another improvement to our implementation of the Extended Baum-Welch update equations for MMI, namely the canceling of any shared part of the numerator and denominator statistics on each frame (a procedure that is already done in MPE). This change affects the Gaussian-specific learning rate. We also investigate another modification whereby we replace I-smoothing to the ML estimate with I-smoothing to the previous iteration's value. Boosted MMI gives better results than MPE in both model and feature-space discriminative training, although not consistently. Daniel Povey, Dimitri Kanevsky, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, Karthik Visweswariah |
ICASSP | 6 |
| 2008 | Learning essential speaker sub-space using hetero-associative neural networks for speaker clustering
Shajith Ikbal, Karthik Visweswariah |
INTERSPEECH | 2 |
| 2008 | An empirical analysis of word error rate and keyword error rateabstractThis paper studies the relationship between word error rate (WER) and keyword error rate (KER) in speech transcripts and their effect on the performance of speech analytics applications. Automatic speech recognition (ASR) systems are increasingly used as input for speech analytics, which raises the question of whether WER or KER is the more suitable performance metric for calibrating the ASR system. ASR systems are typically evaluated in terms of WER. Many speech analytics applications, Youngja Park, Siddharth Patwardhan, Karthik Visweswariah, Stephen C. Gates |
INTERSPEECH | 3 |
| 2007 | Efficient, Low Latency Adaptation for Speech RecognitionabstractConstrained or feature space maximum likelihood linear regression (FMLLR) is known to be an effective algorithm for adaptation to a new speaker or environment. It employs a single transformation matrix and bias vector to linearly transform the test speaker's features. FMLLR makes no assumption on the underlying noise, environment or speaker and estimates parameters to maximize likelihood of the test data. The standard implementation needs considerable computational power, requires significant amounts of storage, and requires a first pass decoding before adaptation can begin. In this paper, we propose a simplified implementation of FMLLR for embedded applications to address these problems. Here, we employ a simple speech/silence segmentation to estimate parameters. We operate in the 13 dimensional cepstral space, hence resource requirements are low. The algorithm does not require a first pass decoding (parameter estimation is accomplished entirely in the front end) and can be applied with low latency as compared to FMLLR. The algorithms we describe here provide an attractive tradeoff between the power of FMLLR and the computational simplicity of Cepstral Mean Subtraction. With minimal cost, we achieve nearly 15% relative gains on an embedded speech recognition task. Suleyman Serdar Kozat, Karthik Visweswariah, Ramesh A. Gopinath |
ICASSP (4) | 2 |
| 2007 | Improving speaker diarization for CHIL lecture meetingsabstractSpeaker diarization is often performed before automatic speech recognition (ASR) to label speaker segments. In this paper we present two simple schemes to improve the speaker diarization performance. The first is to iteratively refine GMM speaker models by frame level re-labelingand smoothingof the decision likelihood. The second is to use word level alignment information from the ASR process. We focus on the CHIL lecture meeting data. Our experiments on the NIST RT06 evaluation data show that these simple methods are quite effective in improving our baseline diarization system, with alignment information providing 1% absolute reduction in diarization error rate (DER) and the re-label smoothing providing an additional 3.51% absolute reduction in DER. The overall system generates a DER that is 6.8% relative better than the top performing system from the RT06 evaluation. Jing Huang 0019, Etienne Marcheret, Karthik Visweswariah |
INTERSPEECH | 3 |
| 2007 | Detection, diarization, and transcription of far-field lecture speechabstractSpeech processing of lectures recorded inside smart rooms has recently attracted much interest. In particular, the topic has been central to the Rich Transcription (RT) Meeting Recognition Evaluation campaign series, sponsored by NIST, with emphasis placed on benchmarking speech activity detection (SAD), speaker diarization (SPKR), speech-to-text (STT), and speakerattributed STT (SASTT) technologies. In this paper, we present the IBM systems developed to address these tasks in preparation for the RT 2007 evaluation, focusing on the far-field condition of lecture data collected as part of European project CHIL. For their development, the systems are benchmarked on a subset of the RT Spring 2006 (RT06s) evaluation test set, where they yield significant improvements for all SAD, SPKR, and STT tasks over RT06s results; for example, a 16% relative reduction in word error rate is reported in STT, attributed to a number of system advances discussed here. Initial results are also presented on SASTT, a task newly introduced in 2007 in place of the discontinued SAD. Index Terms: speech processing, speech recognition, speaker diarization, speech activity detection, lectures, smart rooms. Jing Huang 0019, Etienne Marcheret, Karthik Visweswariah, Vit Libal, Gerasimos Potamianos |
INTERSPEECH | 3 |
| 2007 | Discriminative Estimation of Subspace Constrained Gaussian Mixture Models for Speech RecognitionabstractIn this paper, we study discriminative training of acoustic models for speech recognition under two criteria: maximum mutual information (MMI) and a novel "error-weighted" training technique. We present a proof that the standard MMI training technique is valid for a very general class of acoustic models with any kind of parameter tying. We report experimental results for subspace constrained Gaussian mixture models (SCGMMs), where the exponential model weights of all Gaussians are required to belong to a common "tied" subspace, as well as for subspace precision and mean (SPAM) models which impose separate subspace constraints on the precision matrices (i.e., inverse covariance matrices) and means. It has been shown previously that SCGMMs and SPAM models generalize and yield significant error rate improvements over previously considered model classes such as diagonal models, models with semitied covariances, and extended maximum likelihood linear transformation (EMLLT) models. We show here that MMI and error-weighted training each individually result in over 20% relative reduction in word error rate on a digit task over maximum-likelihood (ML) training. We also show that a gain of as much as 28% relative can be achieved by combining these two discriminative estimation techniques Scott Axelrod, Vaibhava Goel, Ramesh A. Gopinath, Peder A. Olsen, Karthik Visweswariah |
IEEE Trans. Speech Audio Process. | 5 |
| 2006 | Feature Adaptation Based on Gaussian PosteriorsabstractIn this paper we consider the use of non-linear methods for feature adaptation to reduce the mismatch between test and training conditions. The non-linearity is introduced by using the posteriors of a set of Gaussians to (softly) partition the observation space for feature adaptation. The modeling framework used is based on the fMPE models (D. Povey et al., 2005) applied to FMLLR matrices directly. However, the parameters are estimated to maximize the likelihood of the test data. We observe a relative gain of 14% on top of FMLLR, which was a 42% relative gain over the baseline Suleyman Serdar Kozat, Karthik Visweswariah, Ramesh A. Gopinath |
ICASSP (1) | 2 |
| 2006 | Gaussian mixture models with covariances or precisions in shared multiple subspacesabstractWe introduce a class of Gaussian mixture models (GMMs) in which the covariances or the precisions (inverse covariances) are restricted to lie in subspaces spanned by rank-one symmetric matrices. The rank-one basis are shared between the Gaussians according to a sharing structure. We describe an algorithm for estimating the parameters of the GMM in a maximum likelihood framework given a sharing structure. We employ these models for modeling the observations in the hidden-states of a hidden Markov model based speech recognition system. We show that this class of models provide improvement in accuracy and computational efficiency over well-known covariance modeling techniques such as classical factor analysis, shared factor analysis and maximum likelihood linear transformation based models which are special instances of this class of models. We also investigate different sharing mechanisms. We show that for the same number of parameters, modeling precisions leads to better performance when compared to modeling covariances. Modeling precisions also gives a distinct advantage in computational and memory requirements Satya Dharanipragada, Karthik Visweswariah |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | Initializing Subspace Constrained Gaussian Mixture ModelsabstractA recent series of papers [1, 2, 3, 4] introduced subspace constrained Gaussian mixture models (SCGMM) and showed that SCGMM can very efficiently approximate full covariance Gaussian mixture models (FCGMM); a significant reduction in the number of parameters is achieved with little loss in the accuracy of the model. SCGMM were arrived at as a sequence of generalizations of diagonal covariance GMM. As an artifact of this process the initialization of SCGMM parameters in that work is complex, i.e., relies on best parameter settings of less general models. This paper overcomes this problem by showing how an FCGMM can be used to give a simple and direct initialization of an SCGMM. The initialization scheme is powerful enough that as the number of parameters in an SCGMM approaches that of an FCGMM (i.e., large SCGMM) further training of the SCGMM is unnecessary. Peder A. Olsen, Karthik Visweswariah, Ramesh A. Gopinath |
ICASSP (1) | 2 |
| 2005 | Rapid Feature Space Speaker Adaptation for Multi-Stream HMM-Based Audio-Visual Speech RecognitionabstractMulti-stream hidden Markov models (HMMs) have recently been very successful in audio-visual speech recognition, where the audio and visual streams are fused at the final decision level. In this paper we investigate fast feature space speaker adaptation using multi-stream HMMs for audio-visual speech recognition. In particular, we focus on studying the performance of feature-space maximum likelihood linear regression (fMLLR), a fast and effective method for estimating feature space transforms. Unlike the common speaker adaptation techniques of MAP or MLLR, fMLLR does not change the audio or visual HMM parameters, but simply applies a single transform to the testing features. We also address the problem of fast and robust on-line fMLLR adaptation using feature space maximum a posterior linear regression (fMAPLR). Adaptation experiments are reported on the IBM infrared headset audio-visual database. On average for a 20-speaker 1 hour independent test set, the multi-stream fMLLR achieves 31% relative gain on the clean audio condition, and 59% relative gain on the noisy audio condition (approximately 7 dB) as compared to the baseline multi-stream system Jing Huang 0019, Etienne Marcheret, Karthik Visweswariah |
ICME | 3 |
| 2005 | Improving lip-reading with feature space transforms for multi-stream audio-visual speech recognition
Jing Huang 0019, Karthik Visweswariah |
INTERSPEECH | 2 |
| 2005 | Speech activity detection fusing acoustic phonetic and energy features
Etienne Marcheret, Karthik Visweswariah, Gerasimos Potamianos |
INTERSPEECH | 2 |
| 2005 | Feature adaptation using projection of Gaussian posteriors
Karthik Visweswariah, Peder A. Olsen |
INTERSPEECH | 1 |
| 2005 | Subspace constrained Gaussian mixture models for speech recognitionabstractA standard approach to automatic speech recognition uses hidden Markov models whose state dependent distributions are Gaussian mixture models. Each Gaussian can be viewed as an exponential model whose features are linear and quadratic monomials in the acoustic vector. We consider here models in which the weight vectors of these exponential models are constrained to lie in an affine subspace shared by all the Gaussians. This class of models includes Gaussian models with linear constraints placed on the precision (inverse covariance) matrices (such as diagonal covariance, maximum likelihood linear transformation, or extended maximum likelihood linear transformation), as well as the LDA/HLDA models used for feature selection which tie the part of the Gaussians in the directions not used for discrimination. In this paper, we present algorithms for training these models using a maximum likelihood criterion. We present experiments on both small vocabulary, resource constrained, grammar-based tasks, as well as large vocabulary, unconstrained resource tasks to explore the rather large parameter space of models that fit within our framework. In particular, we demonstrate significant improvements can be obtained in both word error rate and computational complexity. Scott Axelrod, Vaibhava Goel, Ramesh A. Gopinath, Peder A. Olsen, Karthik Visweswariah |
IEEE Trans. Speech Audio Process. | 5 |
| 2004 | Stochastic gradient adaptation of front-end parametersabstractThis paper examines how any parameter in the typical front end of a speech recognizer, can be rapidly and inexpensively adapted with usage. It focusses on firstly demonstrating that effective adaptation can be accomplished using low CPU/Memory cost stochastic gradient descent methods, secondly showing that adaptation can be done at time scales small enough to make it effective with just a single utterance, and lastly showing that using a prior on the parameter significantly improves adaptation performance on small amounts of data. It extends previous work on stochastic gradient descent implementation of fMLLR [1] and work on adapting any parameter in the front-end chain using general 2nd order opimization techniques [2]. The framework for general stochastic gradient descent of any front-end parameter with a prior is presented, along with practical techniques to improve convergence. In addition the methods for obtaining the alignment at small time intervals before the end of the utterance are presented. Finally it shown that experimentally online causal adaptation can result in a 5-15% WER reduction across a variety of problems sets and noise conditions, even with just 1 or 2 utterances of adaptation data. Sreeram Balakrishnan, Karthik Visweswariah, Vaibhava Goel |
INTERSPEECH | 2 |
| 2004 | Fast clustering of Gaussians and the virtue of representing Gaussians in exponential model format
Peder A. Olsen, Karthik Visweswariah |
INTERSPEECH | 2 |
| 2004 | Adaptation of front end parameters in a speech recognizerabstractIn this paper we consider the problem of adapting parameters of the algorithm used for extraction of features. Typical speech recognition systems use a sequence of modules to extract features which are then used for recognition. We present a method to adapt the parameters in these modules under a variety of criteria, e.g maximum likelihood, maximum mutual information. This method works under the assumption that the functions that the modules implement are differentiable with respect to their inputs and parameters. We use this framework to optimize a linear transform preceding the linear discriminant analysis (LDA) matrix and show that it gives significantly better performance than a linear transform after the LDA matrix with small amounts of data. We show that linear transforms can be estimated by directly optimizing likelihood or the MMI objective without using auxiliary functions. We also apply the method to optimize the Mel bins, and the compression power in a system that uses power law compression. 1. Karthik Visweswariah, Ramesh A. Gopinath |
INTERSPEECH | 1 |
| 2004 | Task adaptation of acoustic and language models based on large quantities of dataabstractWe investigate use of large amounts, over 1500 hours, of untranscribed data recorded from a deployed conversational system to improve the acoustic and language models. The system that we considered allows users to perform transactions on their retirement accounts. Using all the untranscribed data we get over 19 % relative improvement in word error rate over a baseline system. In contrast, a system built using 70 hours of transcribed data results in over 31 % relative improvement. 1. Karthik Visweswariah, Ramesh A. Gopinath, Vaibhava Goel |
INTERSPEECH | 1 |
| 2003 | Dimensional reduction, covariance modeling, and computational complexity in ASR systemsabstractWe study acoustic modeling for speech recognition using mixtures of exponential models with linear and quadratic features tied across all context dependent states. These models are one version of the SPAM models introduced by Axelrod, Gopinath and Olsen (see Proc. ICSLP, 2002). They generalize diagonal covariance, MLLT, EMLLT, and full covariance models. Reduction of the dimension of the acoustic vectors using LDA/HDA projections corresponds to a special case of reducing the exponential model feature space. We see, in one speech recognition task, that SPAM models on an LDA projected space of varying dimensions achieve a significant fraction of the WER improvement in going from MLLT to full covariance modeling, while maintaining the low computational cost of the MLLT models. Further, the feature precomputation cost can be minimized using the hybrid feature technique of Visweswariah, Olsen, Gopinath and Axelrod (see ICASSP 2003); and the number of Gaussians one needs to compute can be greatly reducing using hierarchical clustering of the Gaussians (with fixed feature space). Finally, we show that reducing the quadratic and linear feature spaces separately produces models with better accuracy, but comparable computational complexity, to LDA/HDA based models. Scott Axelrod, Ramesh A. Gopinath, Peder A. Olsen, Karthik Visweswariah |
ICASSP (1) | 4 |
| 2003 | Covariance and precision modeling in shared multiple subspacesabstractWe introduce a class of Gaussian mixture models for HMM states in continuous speech recognition. In these models, the covariances or the precisions (inverse covariances) are restricted to lie in subspaces spanned by rank-one symmetric matrices. In both cases, the rank-one matrices are shared across classes of Gaussians. We show that, for the same number of parameters, modeling precisions leads to better performance when compared to modeling covariances. Modeling precisions however gives a distinct advantage in computational and memory requirements. We also show that this class of models provides improvement in accuracy (for the same number of parameters) over classical factor analysed models and the recently proposed EMLLT (extended maximum likelihood linear transform) models which are special instances of this class of models. Satya Dharanipragada, Karthik Visweswariah |
ICASSP (1) | 2 |
| 2003 | Maximum likelihood training of subspaces for inverse covariance modelingabstractSpeech recognition systems typically use mixtures of diagonal Gaussians to model the acoustics. Using Gaussians with a more general covariance structure can give improved performance; EM-LLT and SPAM models give improvements by restricting the inverse covariance to a linear/affine subspace spanned by rank one and full rank matrices respectively. We consider training these subspaces to maximize likelihood. For EMLLT ML training the subspace results in significant gains over the scheme proposed by Olsen and Gopinath (see Proceedings of ICASSP, 2002). For SPAM ML training of the subspace slightly improves performance over the method reported by Axelrod, Gopinath and Olsen (see Proceedings of ICSLP, 2002). For the same subspace size an EMLLT model is more efficient computationally than a SPAM model, while the SPAM model is more accurate. This paper proposes a hybrid method of structuring the inverse covariances that both has good accuracy and is computationally efficient. Karthik Visweswariah, Peder A. Olsen, Ramesh A. Gopinath, Scott Axelrod |
ICASSP (1) | 1 |
| 2003 | Large vocabulary conversational speech recognition with a subspace constraint on inverse covariance matricesabstractThis paper applies the recently proposed SPAM models for acoustic modeling in a Speaker Adaptive Training (SAT) context on large vocabulary conversational speech databases, including the Switchboard database. SPAM models are Gaussian mixture models in which a subspace constraint is placed on the precision and mean matrices (although this paper focuses on the case of unconstrained means). They include diagonal covariance, full covariance, MLLT, and EMLLT models as special cases. Adaptation is carried out with maximum likelihood estimation of the means and feature-space under the SPAM model. This paper shows the first experimental evidence that the SPAM models can achieve significant word-error-rate improvements over state-of-the-art diagonal covariance models, even when those diagonal models are given the benefit of choosing the optimal number of Gaussians (according to the Bayesian Information Criterion). This paper also is the first to apply SPAM models in a SAT context. All experiments are performed on the IBM "Superhuman" speech corpus, which is a challenging and diverse conversational speech test set that includes the Switchboard portion of the 1998 Hub5 evaluation data set. Scott Axelrod, Vaibhava Goel, Brian Kingsbury, Karthik Visweswariah, Ramesh A. Gopinath |
INTERSPEECH | 4 |
| 2003 | Discriminative estimation of subspace precision and mean (SPAM) modelsabstractThe SPAM model was recently proposed as a very general method for modeling Gaussians with constrained means and covariances. It has been shown to yield significant error rate improvements over other methods of constraining covariances such as diagonal, semi-tied covariances, and extended maximum likelihood linear transformations. In this paper we address the problem of discriminative estimation of SPAM model parameters, in an attempt to further improve its performance. We present discriminative estimation under two criteria: maximum mutual information (MMI) and an "error-weighted" training. We show that both these methods individually result in over 20% relative reduction in word error rate on a digit task over maximum likelihood (ML) estimated SPAM model parameters. We also show that a gain of as much as 28% relative can be achieved by combining these two discriminative estimation techniques. The techniques developed in this paper also apply directly to an extension of SPAM called subspace constrained exponential models. Vaibhava Goel, Scott Axelrod, Ramesh A. Gopinath, Peder A. Olsen, Karthik Visweswariah |
INTERSPEECH | 5 |
| 2003 | Toward domain-independent conversational speech recognitionabstractWe describe a multi-domain, conversational test set developed for IBM’s Superhuman speech recognition project and our 2002 benchmark system for this task. Through the use of multipass decoding, unsupervised adaptation and combination of hypotheses from systems using diverse feature sets and acoustic models, we achieve a word error rate of 32.0 % on data drawn from voicemail messages, two-person conversations and multiple-person meetings. 1. Brian Kingsbury, Lidia Mangu, George Saon, Geoffrey Zweig, Scott Axelrod, Vaibhava Goel, Karthik Visweswariah, Michael Picheny |
INTERSPEECH | 7 |
| 2003 | Acoustic modeling with mixtures of subspace constrained exponential modelsabstractGaussian distributions are usually parameterized with their natural parameters: the mean and the covariance #. They can also be re-parameterized as exponential models with canonical parameters P = # -1 and # = P. In this paper we consider modeling acoustics with mixtures of Gaussians parameterized with canonical parameters where the parameters are constrained to lie in a shared affine subspace. This class of models includes Gaussian models with various constraints on its parameters: diagonal covariances, MLLT models, and the recently proposed EMLLT and SPAM models. We describe how to perform maximum likelihood estimation of the subspace and parameters within a fixed subspace. In speech recognition experiments, we show that this model improves upon all of the above classes of models with roughly the same number of parameters and with little computational overhead. In particular we get 30-40% relative improvement over LDA+MLLT models when using roughly the same number of parameters. Karthik Visweswariah, Scott Axelrod, Ramesh A. Gopinath |
INTERSPEECH | 1 |
| 2002 | Rapid adaptation with linear combinations of rank-one matricesabstractLinear transforms are often used to adapt the acoustic models in speech recognition systems. When there is very little (5–10 sees.) acoustic data adaptation suffers from unreliable parameter estimation. Typically this problem is handled by imposing a diagonal or block diagonal structure on the transform. This paper proposes using transforms that are linear combinations of rank-one matrices. This approach is applied to the adaptation of the Gaussian means, Gaussian covariances and the acoustic features. Experimental results with varying amounts of adaptation data indicate that for the same number of parameters, our new parameterization performs significantly better than simpler transform parameterizations (diagonal and/or block-diagonal). Vaibhava Goel, Karthik Visweswariah, Ramesh A. Gopinath |
ICASSP | 2 |
| 2002 | Adaptation experiments on the SPINE database with the Extended Maximum Likelihood Linear Transformation (EMLLT) modelabstractThis paper applies the recently proposed Extended Maximum Likelihood Linear Transformation (EMLLT) model for inverse covariances in a Speaker Adaptive Training (SAT) context. The paper adapts standard algorithms for maximum likelihood estimation of linear transforms for mean, variance and feature space adaptation respectively, to the EMLLT model. Experimental results showing word-error-rate improvements are reported on the SPINE2 database. The system described here is the best-performing system submitted by IBM in the SPINE2 evaluation conducted by NIST in October 2001. Ramesh A. Gopinath, Vaibhava Goel, Karthik Visweswariah, Peder A. Olsen |
ICASSP | 3 |
| 2002 | Structuring linear transforms for adaptation using training time informationabstractLinear transforms are often used for adaptation to test data in speech recognition systems. However, when used with small amounts of test data, these techniques provide limited improvements if any. This paper proposes a two-step Bayesian approach where a) the transforms lie in a subspace obtained at training time and b) the expansion coefficients of the transform are obtained using MAP. Estimation algorithms are given for adaptation transforms for means, covariances, and feature spaces. Experimental results indicate that our method gives a significant improvement in performance over other methods. Karthik Visweswariah, Vaibhava Goel, Ramesh A. Gopinath |
ICASSP | 1 |
| 2002 | Large vocabulary conversational speech recognition with the extended maximum likelihood linear transformation (EMLLT) modelabstractThis paper applies the recently proposed Extended Maximum Likelihood Linear Transformation (EMLLT) model in a Speaker Adaptive Training (SAT) context on the Switchboard database. Adaptation is carried out with maximum likelihood estimation of linear transforms for the means, precisions (inverse covariances) and the feature-space under the EMLLT model. This paper shows the first experimental evidence that significant word-error-rate improvements can be achieved with the EMLLT model (in both VTL and VTL+SAT training contexts) over a state-of-the-art diagonal covariance model in a difficult large-vocabulary conversational speech recognition task. The improvements were of the order of 1% absolute in multiple scenarios. Jing Huang 0019, Vaibhava Goel, Ramesh A. Gopinath, Brian Kingsbury, Peder A. Olsen, Karthik Visweswariah |
INTERSPEECH | 6 |
| 2002 | Universal lossless source coding with the Burrows Wheeler TransformabstractThe Burrows Wheeler transform (1994) is a reversible sequence transformation used in a variety of practical lossless source-coding algorithms. In each, the BWT is followed by a lossless source code that attempts to exploit the natural ordering of the BWT coefficients. BWT-based compression schemes are widely touted as low-complexity algorithms giving lossless coding rates better than those of the Ziv-Lempel codes (commonly known as LZ'77 and LZ'78) and almost as good as those achieved by prediction by partial matching (PPM) algorithms. To date, the coding performance claims have been made primarily on the basis of experimental results. This work gives a theoretical evaluation of BWT-based coding. The main results of this theoretical evaluation include: (1) statistical characterizations of the BWT output on both finite strings and sequences of length n /spl rarr/ /spl infin/, (2) a variety of very simple new techniques for BWT-based lossless source coding, and (3) proofs of the universality and bounds on the rates of convergence of both new and existing BWT-based codes for finite-memory and stationary ergodic sources. The end result is a theoretical justification and validation of the experimentally derived conclusions: BWT-based lossless source codes achieve universal lossless coding performance that converges to the optimal coding performance more quickly than the rate of convergence observed in Ziv-Lempel style codes and, for some BWT-based codes, within a constant factor of the optimal rate of convergence for finite-memory sources. Michelle Effros, Karthik Visweswariah, Sanjeev R. Kulkarni, Sergio Verdú |
IEEE Trans. Inf. Theory | 2 |
| 2001 | Speech recognition for DARPA CommunicatorabstractWe report the results of investigations in acoustic modeling, language modeling and decoding techniques, for the DARPA Communicator, a speaker-independent, telephone-based dialog system. By a combination of methods, including enlarging the acoustic model, augmenting the recognizer vocabulary, conditioning the language model upon the dialog state, and applying a post-processing decoding method, we lowered the overall word error rate from 21.9% to 15.0%, a gain of 6.9% absolute and 31.5% relative. Andrew Aaron, Scott Saobing Chen, Paul S. Cohen, Satya Dharanipragada, Ellen Eide, Martin Franz, Jean-Michel LeRoux, X. Luo, Benoît Maison, Lidia Mangu, T. Mathes, Miroslav Novak, Peder A. Olsen, Michael Picheny, Harry Printz, Bhuvana Ramabhadran, Andrej Sakrajda, George Saon, Borivoj Tydlitát, Karthik Visweswariah, D. Yuk |
ICASSP | 20 |
| 2001 | Language models conditioned on dialog stateabstractWe consider various techniques for using the state of the dialog in language modeling. The language models we built were for use in an automated airline travel reservation system. The techniques that we explored include (1) linear interpolation with state specific models and (2) incorporating state information using maximum entropy techniques. We also consider using the system prompt as part of the language model history. We show that using state results in about a relative gain in perplexity and about a percent relative gain in word error rate over a system using a language model with no information of the state. Karthik Visweswariah, Harry Printz |
INTERSPEECH | 1 |
| 2001 | Universal variable-to-fixed length source codesabstractA universal variable-to-fixed length algorithm for binary memoryless sources which converges to the entropy of the source at the optimal rate is known. We study the problem of universal variable-to-fixed length coding for the class of Markov sources with finite alphabets. We give an upper bound on the performance of the code for large dictionary sizes and show that the code is optimal in the sense that no codes exist that have better asymptotic performance. The optimal redundancy is shown to be H log log M/log M where H is the entropy rate of the source and M is the code size. This result is analogous to Rissanen's (1984) result for fixed-to-variable length codes. We investigate the performance of a variable-to-fixed coding method which does not need to store the dictionaries, either at the coder or the decoder. We also consider the performance of both these source codes on individual sequences. For individual sequences we bound the performance in terms of the best code length achievable by a class of coders. All the codes that we consider are prefix-free and complete. Karthik Visweswariah, Sanjeev R. Kulkarni, Sergio Verdú |
IEEE Trans. Inf. Theory | 1 |
| 2000 | Impact of bucketing on performance of linearly interpolated language models
Karthik Visweswariah, Harry Printz, Michael Picheny |
INTERSPEECH | 1 |
| 2000 | Universal coding of nonstationary sourcesabstractWe investigate the performance of the Lempel-Ziv (1978) incremental parsing scheme on nonstationary sources. We show that it achieves the best rate achievable by a finite-state block coder for the nonstationary source. We also show a similar result for a lossy coding scheme given by Yang and Kieffer (see ibid., vol.42, p.239-45, 1996) which uses a Lempel-Ziv scheme to perform lossy coding. Karthik Visweswariah, Sanjeev R. Kulkarni, Sergio Verdú |
IEEE Trans. Inf. Theory | 1 |
| 2000 | Separation of random number generation and resolvabilityabstractWe consider the problem of determining when a given source can be used to approximate the output due to any input to a given channel. We provide achievability and converse results for a general source and channel. For the special case of a full-rank discrete memoryless channel we give a stronger converse result than we can give for a general channel. Karthik Visweswariah, Sanjeev R. Kulkarni, Sergio Verdú |
IEEE Trans. Inf. Theory | 1 |
| 1998 | Source Codes as Random Number GeneratorsabstractA random number generator generates fair coin flips by processing deterministically an arbitrary source of nonideal randomness. An optimal random number generator generates asymptotically fair coin flips from a stationary ergodic source at a rate of bits per source symbol equal to the entropy rate of the source. Since optimal noiseless data compression codes produce incompressible outputs, it is natural to investigate their capabilities as optimal random number generators. We show under general conditions that optimal variable-length source codes asymptotically achieve optimal variable-length random bit generation in a rather strong sense. In particular, we show in what sense the Lempel-Ziv (1978) algorithm can be considered an optimal universal random bit generator from arbitrary stationary ergodic random sources with unknown distributions. Karthik Visweswariah, Sanjeev R. Kulkarni, Sergio Verdú |
IEEE Trans. Inf. Theory | 1 |