Rohit Prasad

dblp:71/1397 · DBLP profile ↗
← Back
94ranked-venue papers
12as first author
4since 2021 · last 2026
0000-0003-1202-837XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 60 · 8 first-authorArtificial intelligence and machine learning · 57 · 7 first-authorDatabases, data management, data science and information retrieval · 18 · 1 first-authorSystems, architecture and hardware · 5 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-Partner Project: dAIEDGE - A Network of Excellence for Distributed, Trustworthy, Efficient and Scalable AI at the Edge
abstract
The dAIEDGE Network of Excellence (NoE) seeks to strengthen and support the development of a dynamic European cutting-edge Artificial intelligence (AI) ecosystem under the umbrella of the European Lighthouse for AI, and to sustain the development of advanced AI. dAIEDGE fosters the exchange of ideas, concepts, and trends on cutting-edge next generation AI, creating links between ecosystem actors to help both the European Commission (EC) and the European Union (EU) and the peripheral AI constituency identify strategies for future developments in Europe. Our main objective is to advance Europe’s innovation and technology base by developing a comprehensive policy and governance approach to AI in order for the EU to become a world leader in innovation in the data economy and its applications.
Alain Pagani, Haralampos-G. D. Stratigopoulos, Aysajan Abidin, Mhd Rashed Al Koutayni, Luca Benini, Angelos Bilas, Alessandro Capotondi, Roberto Cavicchioli, Brian Clerkin, Oscar Déniz-Suárez, Margaux Divernois, Baptiste Dupertuis, Dorvan Favre, Giulio Gambardella, Ander García Gangoiti, Carlo Augusto Grazia, Dominik Günzel, Jude Haris, Klodjan K. Hidri, Maïck Huguenin-Vuillemin, Manal Jammal, Paul Kling, Christos Kozanitis, Xavier Lessage, Srikanth Mandapati, Philippe Massonet, Alfio Di Mauro, Varesh Mishra, Juan Odriozola, Javier Parra 0001, Nuria Pazos, Viviane Potocnik, Miguel de Prado, Rohit Prasad, Spyridon Raptis, Gregoire Rebstein, Ignacio Sanudo Olmedo, Mohamed Selim, Chinmay Satish Shrivastav, Noelia Vállez, Giorgos Vasiliadis, Micaela Verrucchi, Enrico Vincenzi, Damian Vizár, Devendra Vyas, Stefan Wiehle
DATE35
2026 NX-CGRA: A Programmable Hardware Accelerator for Core Transformer Algorithms on Edge Devices
abstract
The increasing diversity and complexity of transformer workloads at the edge present significant challenges in balancing performance, energy efficiency, and architectural flexibility. This paper introduces NX-CGRA, a programmable hardware accelerator designed to support a range of transformer inference algorithms, including both linear and non-linear functions. Unlike fixed-function accelerators optimized for narrow use cases, NX-CGRA employs a coarse-grained reconfigurable array (CGRA) architecture with software-driven programmability, enabling efficient execution across varied kernel patterns. The architecture is evaluated using representative benchmarks derived from real-world transformer models, demonstrating high overall efficiency and favorable energy-area tradeoffs across different classes of operations. These results indicate the potential of NX-CGRA as a scalable and adaptable hardware solution for edge transformer deployment under constrained power and silicon budgets.
Rohit Prasad
DATE1
2026 SPEC CPU: The Next Generotion
abstract
The march toward developing relevant and robust CPU benchmarks continues with the introduction of SPEC CPU 2026, the next generation suite for measuring processor performance. This paper details the methodology behind its creation, showcasing a process centered on community collaboration and principled development. The suite is built upon a foundation of modern, open-source applications, selected and hardened through a process that emphasizes workload diversity, portability, and software longevity. A key contribution is Rolling-Round-Robin Rate, a novel and standardized approach to running heterogeneous, multiprogrammed workloads that addresses a long-standing gap in benchmarking practice. Additionally, the suite features an expanded set of multithreaded benchmarks and introduces workloads with distinct microarchitectural profiles, reflecting the demands of contemporary software. By detailing our principled approach to benchmark selection, adaptation, and validation, we demonstrate how the SPEC CPU 2026 suite sets the standard for performance evaluation in the next era of computer architecture research and development.
Mahesh Madhav, Allen Lee, Andres Mejia, Branden Moore, Charan Soppadandi, Chris Cambly, Christoph Müllner, Daniel Bowers, David Reiner, Denis Bakhvalov, Duane Voth, Frédérique Silber-Chaussumier, James Bucek, James Southern, Jiangning Liu, Jim Himer, John Henning, Kristen Yang, Kunal Kashyap, Mason Guy, Mat Colgrove, Michael Berg, Prasad Battini, Prasad Joshi, Rohit Prasad, Shayantika Bhattacharya, Sriyash Caculo, Stefan Reimbold, Sundar Iyengar, Van Smith, Zarko Todorovski
ISCA28
2025 J3DAI: A tiny DNN-Based Edge AI Accelerator for 3D-Stacked CMOS Image Sensor
abstract
This paper presents J3DAI, a tiny deep neural network-based hardware accelerator for a 3-layer 3D-stacked CMOS image sensor featuring an artificial intelligence (AI) chip integrating a Deep Neural Network (DNN)-based accelerator. The DNN accelerator is designed to efficiently perform neural network tasks such as image classification and segmentation. This paper focuses on the digital system of J3DAI, highlighting its Performance-Power-Area (PPA) characteristics and showcasing advanced edge AI capabilities on a CMOS image sensor.To support hardware, we utilized the Aidge comprehensive software framework, which enables the programming of both the host processor and the DNN accelerator. Aidge supports post-training quantization, significantly reducing memory footprint and computational complexity, making it crucial for deploying models on resource-constrained hardware like J3DAI.Our experimental results demonstrate the versatility and efficiency of this innovative design in the field of edge AI, showcasing its potential to handle both simple and computationally intensive tasks.
Benoît Tain, Raphael Millet, Romain Lemaire, Michal Szczepanski, Laurent Alacoque, Emmanuel Pluchart, Sylvain Choisnet, Rohit Prasad, Jérôme Chossat, Pascal Pierunek, Pascal Vivet, Sébastien Thuries
ISLPED8
2020 TRANSPIRE: An energy-efficient TRANSprecision floating-point Programmable archItectuRE
abstract
In recent years, Coarse Grain Reconfigurable Architecture (CGRA) accelerators have been increasingly deployed in Internet-of-Things (IoT) end nodes. A modern CGRA has to support and efficiently accelerate both integer and floating-point (FP) operations. In this paper, we propose an ultra-low-power tunable-precision CGRA architectural template, called TRANSprecision floating-point Programmable archItectuRE (TRANSPIRE), and its associated compilation flow supporting both integer and FP operations. TRANSPIRE employs transprecision computing and multiple Single Instruction Multiple Data (SIMD) to accelerate FP operations while boosting energy efficiency as well. Experimental results show that TRANSPIRE achieves a maximum of 10.06× performance gain and consumes 12.91× less energy w.r.t. a RISC-V based CPU with an enhanced ISA supporting SIMD-style vectorization and FP data-types, while executing applications for near-sensor computing and embedded machine learning, with an area overhead of 1.25× only.
Rohit Prasad, Satyajit Das, Kevin J. M. Martin, Giuseppe Tagliavini, Philippe Coussy, Luca Benini, Davide Rossi 0001
DATE1
2020 Energy Efficient Acceleration Of Floating Point Applications Onto CGRA
abstract
In this paper, we propose a novel CGRA architecture and associated compilation flow supporting both integer and floating-point computations for energy efficient acceleration of DSP applications. Experimental results show that the proposed accelerator achieves a maximum of 4.61 × speedup compared to a DSP optimized, ultra low power RISC-V based CPU while executing seizure detection, a representative of wide range of EEG signal processing applications with an area overhead of 1.9×. The proposed CGRA achieves a maximum of 6.5× energy efficiency compared to the CPU.
Satyajit Das, Rohit Prasad, Kevin J. M. Martin, Philippe Coussy
ICASSP2
2019 Alexa Everywhere: AI for Daily Convenience
abstract
The computing industry has been on an inexorable march toward simplifying human-computer interaction, and earlier this decade Amazon bet big on combining voice technology and artificial intelligence. In 2014, with the introduction of Echo and Alexa, Amazon created an entirely new technology category with an AI-first strategy and vision. Since then, Alexa has captured the imagination of customers across the globe, and the company has accelerated the pace of AI research and innovation in support of its promise to improve Alexa every day. In this presentation Rohit Prasad, Vice President and Head Scientist of Amazon Alexa, shares his insights into how recent scientific innovations are advancing Alexa.
Rohit Prasad
WSDM1
2013 ASR error detection in a conversational spoken language translation system
abstract
Detection of automatic speech recognition (ASR) errors is crucial to preventing their further propagation through statistical machine translation (SMT) in conversational spoken language translation (CSLT) systems. In this paper, we venture beyond traditional features obtained from the ASR decoder and hypothesized word sequence, and explore additional information streams provided by an error-robust CSLT system, including SMT confidence estimates and posteriors from named entity detection (NED). Another significant novelty of this work is the use of an automated word boundary detector based on acoustic-prosodic features to verify the existence of ASR-hypothesized word boundaries, which further improves ASR error detection. Offline evaluation on a test set designed to invoke ASR errors showed that at 10% false alarm rate, the proposed features provide 2.8% absolute (4.2% relative) improvement in detection rate over a state-of-the-art baseline error detector that uses a rich set of features traditionally employed in the existing literature.
Sankaranarayanan Ananthakrishnan, Rohit Kumar 0001, Rohit Prasad, Premkumar Natarajan
ICASSP4
2013 Robust EEG emotion classification using segment level decision fusion
abstract
In this paper we address single-trial binary classification of emotion dimensions (arousal, valence, dominance and liking) using electroencephalogram (EEG) signals that represent responses to audio-visual stimuli. We propose an innovative three step solution to this problem: (1) in contrast to the typical feature extraction on the response-level, we represent the EEG signal as a sequence of overlapping segments and extract feature vectors on the segment level; (2) transform segment level features to the response level features using projections based on a novel non-parametric nearest neighbor model; and (3) perform classification on the obtained response-level features. We demonstrate the efficacy of our approach by performing binary classification of emotion dimensions on DEAP (Dataset for Emotion Analysis using electroencephalogram, Physiological and Video Signals) and report state-of-the-art classification accuracies for all emotional dimensions.
Viktor Rozgic, Shiv Vitaladevuni, Rohit Prasad
ICASSP3
2013 Graph based multimodal word clustering for video event detection
abstract
Combining diverse low-level features from multiple modalities has consistently improved performance over a range of video processing tasks, including event detection. In our work, we study graph based clustering techniques for integrating information from multiple modalities by identifying word clusters spread across the different modalities. We present different methods to identify word clusters including word similarity graph partitioning, word-video co-clustering and Latent Semantic Indexing and the impact of different metrics to quantify the co-occurrence of words. We present experimental results on a ≈45000 video dataset used in the TRECVID MED 11 evaluations. Our experiments show that multimodal features have consistent performance gains over the use of individual features. Further, word similarity graph construction using a complete graph representation consistently improves over partite graphs and early fusion based multimodal systems. Finally, we see additional performance gains by fusing multimodal features with individual features.
Aravind Namandi Vembu, Pradeep Natarajan, Shuang Wu 0003, Rohit Prasad, Premkumar Natarajan
ICASSP4
2013 Detecting OOV Names in Arabic Handwritten Data
abstract
This paper presents a novel approach to detect Arabic OOV names from OCR'ed handwritten documents. In our approach, OOV names are searched for using approximate string match on character consensus networks (cnets). The retrieved regions are re-ranked using novel features representing the quality of the match and the likelihood of the detected region to be an OOV name. Our features that encode word boundary information into the approximate match algorithm significantly improve mean average precision (MAP) by 12.2% (absolute gains) for rank cut-off 100 (48.2% vs. 36.0%) and 11.9% for cut-off 1000 (47.0% vs. 35.1%) over the baseline system. Discriminative reranking based on maximum entropy classification using novel features, such as the probability of a retrieved region being an OOV name (called OOV name probability) from a conditional random field model, further improve MAP by 2.3% (absolute gains) for cut-off 100 and 3.0% for cut-off 1000. The improvements are consistent in DET (Detection Error Tradeoff) curves. Our results show that character cnet based OOV name search benefits clearly from the approximate match using word boundary information and the reranking algorithm. Our experiments also show that OOV name probability is very useful for reranking.
Jinying Chen, Rohit Prasad, Huaigu Cao, Premkumar Natarajan
ICDAR2
2013 Exploiting Stroke Orientation for CRF Based Binarization of Historical Documents
abstract
We present a novel binarization method that is especially effective on historical documents with the following characteristics: (a) the documents contain free-form cursive handwritten text with significant but consistent slant, (b) scanning artifacts resulting in the text and background pixels not having uniform intensity even within the same page, and (c) pages containing significant amount of bleeds from the other side of the page. In order to tackle the problem of non-uniform text and background intensity, we use a thresholding algorithm that works equally well for regions of the page containing text and regions of the page containing no text. We then combine this algorithm with a CRF-based framework which handles bleeds using a novel approach to further improve the quality of binarization. We compare the proposed binarization algorithm against other popular binarization algorithms both qualitatively using examples and quantitatively using the word error rate (WER) metric from performing optical character recognition (OCR) on binarized text using the BBN Byblos Offline Handwritten text recognition (OHR) system.
Xujun Peng, Huaigu Cao, Krishna Subramanian 0001, Rohit Prasad, Premkumar Natarajan
ICDAR4
2013 Variable-Span out-of-vocabulary named entity detection
Sankaranarayanan Ananthakrishnan, Rohit Prasad, Premkumar Natarajan
INTERSPEECH3
2013 Audio self organized units for high-level event detection
Xiaodan Zhuang, Shuang Wu 0003, Pradeep Natarajan, Rohit Prasad, Premkumar Natarajan
INTERSPEECH4
2013 Probabilistic trainable segmenter for call center audio using multiple features
Nina Zinovieva, Xiaodan Zhuang, Pat Peterson, Joe Alwan, Rohit Prasad
INTERSPEECH5
2013 Semi-Supervised Word Sense Disambiguation for Mixed-Initiative Conversational Spoken Language Translation
Sankaranarayanan Ananthakrishnan, Sanjika Hewavitharana, Rohit Kumar 0001, Enoch Kan, Rohit Prasad, Premkumar Natarajan
MTSummit5
2013 Ridge Regression based classifiers for large scale class imbalanced datasets
abstract
Large scale, class imbalanced data classification is a challenging task that occurs frequently in several computer vision tasks such as web video retrieval. A number of algorithms have been proposed in literature that approach this problem from different perspectives (e.g. Sampling, Cost-sensitive learning, Active learning). The challenge is two fold in this task - first the data imbalance causes many classification algorithms to learn trivial classifiers that declare all test examples to be from the majority class. Second, many algorithms do not scale to large dataset sizes. We address these two issues by using two different cost-sensitive versions of Ridge Regression as our binary classifiers. We demonstrate our approach for retrieving unstructured web videos from 10 events on the benchmark TRECVID MED 12 dataset containing ≈47000 videos. We empirically show that they perform at par with state-of-the-art support vector machine based classifiers using χ2kernels while being 30 to 60 times faster.
Devansh Arpit, Shuang Wu 0003, Pradeep Natarajan, Rohit Prasad, Premkumar Natarajan
WACV4
2013 Scene image categorization and video event detection using Naive Bayes Nearest Neighbor
abstract
We present a detailed study of Naive Bayes Nearest Neighbor (NBNN) proposed by Boiman et al., with application to scene categorization and video event detection. Our study indicates that using Dense-SIFT along with dimensionality reduction using PCA enables NBNN to obtain state-of-the-art results. We demonstrate this on two tasks: (1) scene image categorization on the UIUC 8 Sports Events Image Dataset (obtaining 84.67%) and the MIT 67 Indoor Scene Image Dataset (obtaining 48.84%); and (2) detecting videos depicting certain events of interest on the challenging MED'11 video dataset with only 15 positive training videos per event. We present an extension referred to as sparse-NBNN that constrains the number of training images that can used to match with a given test image for the image-to-class distance computation. Experiments indicate that this improves upon NBNN for handling of imbalanced training data.
Shiv Vitaladevuni, Pradeep Natarajan, Shuang Wu 0003, Xiaodan Zhuang, Rohit Prasad, Premkumar Natarajan
WACV5
2013 Batch-mode semi-supervised active learning for statistical machine translation
Sankaranarayanan Ananthakrishnan, Rohit Prasad, David Stallard, Premkumar Natarajan
Comput. Speech Lang.2
2013 BBN TransTalk: Robust multilingual two-way speech-to-speech translation for mobile platforms
Rohit Prasad, Premkumar Natarajan, David Stallard, Shirin Saleem, Shankar Ananthakrishnan, Stavros Tsakalidis, Chia-Lin Kao, Fred Choi, Ralf Meermeier, Mark Rawls, Jacob Devlin, Kriste Krstovski, Aaron Challenner
Comput. Speech Lang.1
2012 Automatic Detection of Psychological Distress Indicators and Severity Assessment from Online Forum Posts
Shirin Saleem, Rohit Prasad, Shiv Vitaladevuni, Maciej Pacula, Michael Crystal, Brian Marx, Denise Sloan, Jennifer Vasterling, Theodore Speroff
COLING2
2012 Multimodal feature fusion for robust event detection in web videos
abstract
Combining multiple low-level visual features is a proven and effective strategy for a range of computer vision tasks. However, limited attention has been paid to combining such features with information from other modalities, such as audio and videotext, for large scale analysis of web videos. In our work, we rigorously analyze and combine a large set of low-level features that capture appearance, color, motion, audio and audio-visual co-occurrence patterns in videos. We also evaluate the utility of high-level (i.e., semantic) visual information obtained from detecting scene, object, and action concepts. Further, we exploit multimodal information by analyzing available spoken and videotext content using state-of-the-art automatic speech recognition (ASR) and videotext recognition systems. We combine these diverse features using a two-step strategy employing multiple kernel learning (MKL) and late score level fusion methods. Based on the TRECVID MED 2011 evaluations for detecting 10 events in a large benchmark set of ~45000 videos, our system showed the best performance among the 19 international teams.
Pradeep Natarajan, Shuang Wu 0003, Shiv Vitaladevuni, Xiaodan Zhuang, Stavros Tsakalidis, Unsang Park, Rohit Prasad, Premkumar Natarajan
CVPR7
2012 Local Segmentation of Touching Characters Using Contour Based Shape Decomposition
abstract
We propose a contour based shape decomposition approach that provides local segmentation of touching characters. The shape contour is linearized into edge lets and edge lets are merged into boundary fragments. The connection cost between boundary fragments is obtained by considering local smoothness, connection length and a stroke-level property called the Same Stroke Rate. Samples of connections among boundary fragments are randomly generated and the one with the minimum global cost is selected to produce the final segmentation of the shape. To obtain a bipartite segmentation using this approach, we perform an iterative search for the parameters that finally yields two components on a shape. Experimental results on synthetic shape images and the LTP dataset show that this contour based shape decomposition technique is promising and it is effective for providing local segmentation of touching characters.
David S. Doermann, Huaigu Cao, Rohit Prasad, Premkumar Natarajan
Document Analysis Systems4
2012 Automatic Tune Set Generation for Machine Translation with Limited Indomain Data
Jinying Chen, Jacob Devlin, Huaigu Cao, Rohit Prasad, Premkumar Natarajan
EAMT4
2012 Multi-channel Shape-Flow Kernel Descriptors for Robust Video Event Detection and Retrieval
Pradeep Natarajan, Shuang Wu 0003, Shiv Vitaladevuni, Xiaodan Zhuang, Unsang Park, Rohit Prasad, Premkumar Natarajan
ECCV (2)6
2012 Automatic pronunciation prediction for text-to-speech synthesis of dialectal arabic in a speech-to-speech translation system
abstract
Text-to-speech synthesis (TTS) is the final stage in the speech-tospeech (S2S) translation pipeline, producing an audible rendition of translated text in the target language. TTS systems typically rely on a lexicon to look up pronunciations for each word in the input text. This is problematic when the target language is dialectal Arabic, because the statistical machine translation (SMT) system usually produces undiacritized text output. Many words in the latter possess multiple pronunciations; the correct choice must be inferred from context. In this paper, we present a weakly supervised pronunciation prediction approach for undiacritized dialectal Arabic in S2S systems that leverages automatic speech recognition (ASR) to obtain parallel training data for pronunciation prediction. Additionally, we show that incorporating source language features derived from SMT-generated automatic word alignment further improves automatic pronunciation prediction accuracy.
Sankaranarayanan Ananthakrishnan, Stavros Tsakalidis, Rohit Prasad, Premkumar Natarajan, Aravind Namandi Vembu
ICASSP3
2012 Applying Discriminatively Optimized Feature Transform for HMM-based Off-Line Handwriting Recognition
abstract
Feature extraction is an important step in off-line handwriting recognition systems to represent raw handwriting in a low-dimensional, tractable feature space. Traditionally, linear feature transforms such as Principle Component Analysis (PCA), Linear Discriminative Analysis (LDA) are commonly used. The assumptions they make, however, usually cannot be satisfied in practice and thus the best performance is not obtained. In this paper, we apply the Region-Dependent non-linear feature Transform (RDT) to handwriting recognition. RDT is one type of non-linear feature transforms which captures the discriminating power much better than traditional linear ones. We justify the effectiveness of RDT on handwriting features using an HMM-based handwriting recognition system on an Arabic handwriting dataset, which consists of 38K pages of handwriting, over 3M handwritten words. Experimental results show that RDT is able to decrease the word error rates (WERs) relatively by 4% to 7% with statistical significance, comparing to two LDA-based baseline systems.
Huaigu Cao, Rohit Prasad, Premkumar Natarajan
ICFHR4
2012 Statistical Machine Translation as a Language Model for Handwriting Recognition
abstract
When performing handwriting recognition on natural language text, the use of a word-level language model (LM) is known to significantly improve recognition accuracy. The most common type of language model, the n-gram model, decomposes sentences into short, overlapping chunks. In this paper, we propose a new type of language model which we use in addition to the standard n-gram LM. Our new model uses the likelihood score from a statistical machine translation system as a reranking feature. In general terms, we automatically translate each OCR hypothesis into another language, and then create a feature score based on how "difficult" it was to perform the translation. Intuitively, the difficulty of translation correlates with how well-formed the input sentence is. In an Arabic handwriting recognition task, we were able to obtain an 0.4% absolute improvement to word error rate (WER) on top of a powerful 5-gram LM.
Jacob Devlin, Matin Kamali, Krishna Subramanian 0001, Rohit Prasad, Premkumar Natarajan
ICFHR4
2012 Document recognition and translation system for unconstrained Arabic documents
Huaigu Cao, Jinying Chen, Jacob Devlin, Rohit Prasad, Premkumar Natarajan
ICPR4
2012 Extracting information from handwritten content in census forms
Huaigu Cao, Krishna Subramanian 0001, Xujun Peng, Jinying Chen, Rohit Prasad, Premkumar Natarajan
ICPR5
2012 Detecting near-duplicate document images using interest point matching
Shiv Vitaladevuni, Fred Choi, Rohit Prasad, Premkumar Natarajan
ICPR3
2012 Detecting OOV Named-Entities in Conversational Speech
Rohit Kumar 0001, Rohit Prasad, Sankaranarayanan Ananthakrishnan, Aravind Namandi Vembu, David Stallard, Stavros Tsakalidis, Premkumar Natarajan
INTERSPEECH2
2012 Emotion Recognition using Acoustic and Lexical Features
Viktor Rozgic, Sankaranarayanan Ananthakrishnan, Shirin Saleem, Rohit Kumar 0001, Aravind Namandi Vembu, Rohit Prasad
INTERSPEECH6
2012 Robust Event Detection From Spoken Content In Consumer Domain Videos
Stavros Tsakalidis, Xiaodan Zhuang, Roger Hsiao, Shuang Wu 0003, Pradeep Natarajan, Rohit Prasad, Premkumar Natarajan
INTERSPEECH6
2012 Compact Audio Representation for Event Detection in Consumer Media
Xiaodan Zhuang, Stavros Tsakalidis, Shuang Wu 0003, Pradeep Natarajan, Rohit Prasad, Premkumar Natarajan
INTERSPEECH5
2011 Model-based parametric features for emotion recognition from speech
abstract
Automatic emotion recognition from speech is desirable in many applications relying on spoken language processing. Telephone-based customer service systems, psychological healthcare initiatives, and virtual training modules are examples of real-world applications that would significantly benefit from such capability. Traditional utterance-level emotion recognition relies on a global feature set obtained by computing various statistics from raw segmental and supra-segmental measurements, including fundamental frequency (F0), energy, and MFCCs. In this paper, we propose a novel, model-based parametric feature set that better discriminates between the competing emotion classes. Our approach relaxes modeling assumptions associated with using global statistics (e.g. mean, standard deviation, etc.) of traditional segment-level features for classification, and results in significant improvements over the state-of-the-art in 7-way emotion classification accuracy on the standard, freely-available Berlin Emotional Speech Corpus. These improvements are consistent even in a reduced feature space obtained by Fisher's Multiple Linear Discriminant Analysis, demonstrating the signficantly higher discriminative power of the proposed feature set.
Sankaranarayanan Ananthakrishnan, Aravind Namandi Vembu, Rohit Prasad
ASRU3
2011 Wavelet Band-pass Filters for Matching Multiple Templates in Real-time
abstract
Many applications in image processing and computer vision require finding a particular template in an image or a video, that is, template matching. Given a template and an input, the matching algorithm finds the region of interest (ROI) that most closely matches the template in terms of some similarity measurement. According to the way similarity measurements are performed, the template matching methods can roughly be classified into two groups: 1) patch matching schemes, such as the sum of absolute difference (SAD) , the sum of squared difference (SSD) [1], or cross correlation (XCORR), where the similarity measurement directly relies on pixel information from the patch of interest; and 2) feature matching schemes, such as invariant features [4] and bags of features [5], where similarity measurement relies on features describing the template and the frame. Patch matching methods are not robust, especially when noise, skew, or errors occur. Further, they consume a large amount of time, because of expensive sliding window search for calculating the similarity score over all possible locations. Several techniques have been explored for accelerating such matching methods, including early rejections and correlation techniques [1]. However, the computation cost could still be unaffordable when the frame size is large. Typically other techniques, like frame difference, are used to reduce the search space in applications. Feature matching methods process the template and describe it with features, which are ideally invariant to rotation, skew, noise etc. However in many cases, the use of a more complicated model for similarity measurement results in higher computational cost. Further, sliding window search is also a costly stage for such methods. While there exist known algorithms for fast search of object instances in an image using branch-and-bound techniques, in our particular problem, methods of this type have two crucial limitations. First, they require a large number of training samples for each class to learn robust classifiers. Second, interest point detectors like SIFT [5] typically do not generate sufficient number of feature points, because of the small size of the provided logo, large homogenous regions and degradations. Wavelets based approaches have been extensively used in object detection and recognition. In [3], wavelet coefficients based image histogram are collected in bins and are used for classifying logos. In [6], wavelet coefficients are directly used and trained for pedestrian detection. In [8], wavelet coefficients are selected to form rotation-invariant features by using the angular-radial transform. However, matching logos within frames using [3, 6, 8] still requires expensive window searching and thus are not appropriate for real-time processing. In this paper, we propose a new matching method using the wavelet based band-pass filters (WBPFs). Instead of using direct distance measurement requiring expensive window search, the similarity is measured in the indirect way involving two stages. In the stage of offline template processing (see Figure 1), a template is automatically described by a set of three directional WBPFs, where only salient wavelet frequency components of the template are allowed to pass. In the stage of online frame processing (see Figure 2), a frame is transformed to the wavelet domain and its sub-bands are filtered with respect to the corresponding template WBFPs. Finally, the detection is made with respect to the region of the densest responses under spatial constraints [2, 4]. We show that the proposed template matching system has a very low computational cost, which is 50 times faster than the correlation based SSD [1] and 10 times faster than the orthogonal Haar transform (OHT) based SSD [7]. Further, the proposed method does not trade-off accuracy, since the use of subtemplate information makes it robust to skew and camera view change. Experimental results demonstrate our method for real-time logo detection in broadcast videos.
Yue Wu 0001, Pradeep Natarajan, Joseph P. Noonan, Rohit Prasad, Premkumar Natarajan
BMVC4
2011 Efficient Orthogonal Matching Pursuit using sparse random projections for scene and video classification
abstract
Sparse projection has been shown to be highly effective in several domains, including image denoising and scene / object classification. However, practical application to large scale problems such as video analysis requires efficient versions of sparse projection algorithms such as Orthogonal Matching Pursuit (OMP). In particular, random projection based locality sensitive hashing (LSH) has been proposed for OMP. In this paper, we propose a novel technique called Comparison Hadamard random projection (CHRP) for further improving the efficiency of LSH within OMP. CHRP combines two techniques:(1) The Fast Johnson-Lindenstrauss Transform (FJLT) which uses a randomized Hadamard transform and sparse projection matrix for LSH, and (2) Achlioptas' random projection that uses only addition and comparison operations. Our approach provides the robustness of FJLT while completely avoiding multiplications. We empirically validate CHRP's efficacy by performing a suite of experiments for image denoising, scene classification, and video categorization. Our experiments indicate that CHRP significantly speeds-up OMP with negligible loss in classification accuracy.
Shiv Vitaladevuni, Pradeep Natarajan, Rohit Prasad, Premkumar Natarajan
ICCV3
2011 OCR-Driven Writer Identification and Adaptation in an HMM Handwriting Recognition System
abstract
We present an OCR-driven writer identification algorithm in this paper. Our algorithm learns writer-specific characteristics more precisely from explicit character alignment using the Viterbi algorithm and shows significant reduction of close-set writer identification error rates, compared with the GMM-based method. With writers' identities retrieved, we improve the performance of handwriting recognition using the HMM trained adapted on the training data of that writer. In our system, writer identification and OCR are highly interactive. They improve the performance of each other and thus show close approximation of supervised text-dependent writer identification and writer-dependent HMM handwriting.
Huaigu Cao, Rohit Prasad, Premkumar Natarajan
ICDAR2
2011 Handwritten and Typewritten Text Identification and Recognition Using Hidden Markov Models
abstract
We present a system for identification and recognition of handwritten and typewritten text from document images using hidden Markov models (HMMs) in this paper. Our text type identification uses OCR decoding to generate word boundaries followed by word-level handwritten/typewritten identification using HMMs. We show that the contextual constraints from the HMM significantly improves the identification performance over the conventional Gaussian mixture model (GMM)-based method. Type identification is then used to estimate the frame sample rates and frame width of feature sequences for HMM OCR system for each type independently. This type-dependent approach to computing the frame sample rate and frame width shows significant improvement in OCR accuracy over type-independent approaches.
Huaigu Cao, Rohit Prasad, Premkumar Natarajan
ICDAR2
2011 Graph Clustering-Based Ensemble Method for Handwritten Text Line Segmentation
abstract
Handwritten text line segmentation on real-world data presents significant challenges that cannot be overcome by any single technique. Given the diversity of approaches and the recent advances in ensemble-based combination for pattern recognition problems, it is possible to improve the segmentation performance by combining the outputs from different line finding methods. In this paper, we propose a novel graph clustering-based approach to combine the output of an ensemble of text line segmentation algorithms. A weighted undirected graph is constructed with nodes corresponding to connected components and edge connecting pairs of connected components. Text line segmentation is then posed as the problem of minimum cost partitioning of the nodes in the graph such that each cluster corresponds to a unique line in the document image. Experimental results on a challenging Arabic field dataset using the ensemble method shows a relative gain of 18% in the F1 score over the best individual method within the ensemble.
Vasant Manohar, Shiv Vitaladevuni, Huaigu Cao, Rohit Prasad, Premkumar Natarajan
ICDAR4
2011 Baseline Dependent Percentile Features for Offline Arabic Handwriting Recognition
abstract
Handwritten text in Arabic and other languages exhibit significant variations in the slant and baseline of characters across words and also within a single word. Since the concept of baseline does not have a precise mathematical definition, existing approaches use heuristic methods to first identify a set of baseline relevant pixels and then fit lines/curves through them. However, for statistical features like percentiles that we use in our system, we only need an approximate curve that is close to the baseline to normalize the features. Hence we propose a two stage approach to estimate the approximate baseline. First we segment the text line into a set of components, and then estimate the baseline in each component using two methods max projection and smoothed centroid line. We incorpate the computed baseline into percentile feature computation in the BBN Byblos OCR system for an Arabic offline handwriting recognition task. Our new features, result in a 1% absolute gain and 3.1% relative gain in the word error rate on a large test set with 15K handwritten Arabic words, which is statistically significant with p-value<;0.001 using the matched pair comparison test. Further, our results show that computing fine-grained baselines from small line segments is significantly better than estimating a single baseline over the entire text line.
Pradeep Natarajan, David Belanger 0001, Rohit Prasad, Matin Kamali, Krishna Subramanian 0001, Premkumar Natarajan
ICDAR3
2011 Text Extraction from Video Using Conditional Random Fields
abstract
In this paper, we describe an approach to extract text from broadcast videos. Candidate blocks are detected based on edge extraction results. Corners and geometrical features are used for the purpose of initial classification which is carried out by using a support vector machine (SVM). Considering the spatial inter-dependencies of different regions in the image, we propose a novel conditional random field (CRF) based framework which integrates the outputs of SVM into the system to improve the accuracy of labeling for blocks. The experimental results show that the proposed system achieves reliable performance for text detection/extraction from videos.
Xujun Peng, Huaigu Cao, Rohit Prasad, Premkumar Natarajan
ICDAR3
2011 Automated image quality assessment for camera-captured OCR
abstract
Camera-captured optical character recognition (OCR) is a challenging area because of artifacts introduced during image acquisition with consumer-domain hand-held and Smart phone cameras. Critical information is lost if the user does not get immediate feedback on whether the acquired image meets the quality requirements for OCR. To avoid such information loss, we propose a novel automated image quality assessment method that predicts the degree of degradation on OCR. Unlike other image quality assessment algorithms which only deal with blurring, the proposed method quantifies image quality degradation across several artifacts and accurately predicts the impact on OCR error rate. We present evaluation results on a set of machine-printed document images which have been captured using digital cameras with different degradations.
Xujun Peng, Huaigu Cao, Krishna Subramanian 0001, Rohit Prasad, Premkumar Natarajan
ICIP4
2011 Large-scale, real-time logo recognition in broadcast videos
abstract
Robust, real-time, multi-class logo detection in high resolution broadcast videos presents several difficult challenges. For most logos we only have a few training samples, which makes training robust classifiers hard. Also, logos could potentially occur anywhere in the image, and traditional sliding window approaches for logo/object detection are computationally intensive. We present a system that addresses these issues by first identifying a small set of possible logo locations in a frame, based on temporal continuity and multi-resolution search, and then successively pruning these locations for each logo template, using a cascade of color and edge based features. We present experimental results that demonstrate our system for detecting a total of 270 different logo classes in broadcast video from 5 different languages (English, Indonesian, Malay, Simplified and Traditional Chinese).
Pradeep Natarajan, Yue Wu 0001, Shirin Saleem, Ehry MacRostie, Fred Bernardin, Rohit Prasad, Premkumar Natarajan
ICME6
2011 Source Error-Projection for Sample Selection in Phrase-Based SMT for Resource-Poor Languages
Sankaranarayanan Ananthakrishnan, Shiv Vitaladevuni, Rohit Prasad, Premkumar Natarajan
IJCNLP3
2011 On-Line Language Model Biasing for Multi-Pass Automatic Speech Recognition
Sankaranarayanan Ananthakrishnan, Stavros Tsakalidis, Rohit Prasad, Premkumar Natarajan
INTERSPEECH3
2011 Unsupervised Audio Analysis for Categorizing Heterogeneous Consumer Domain Videos
Pradeep Natarajan, Stavros Tsakalidis, Vasant Manohar, Rohit Prasad, Premkumar Natarajan
INTERSPEECH4
2011 Audio-visual fusion using bayesian model combination for web video retrieval
abstract
Combining features from multiple, heterogeneous, audio visual sources can significantly improve retrieval performance in consumer domain videos. However, such videos often contain unrelated overlaid audio content, or have significant camera motion to reliably extract visual features. We present an approach, which overcomes errors in individual feature streams by combining classifiers trained on multiple, heterogeneous feature streams using Bayesian model combination (BAYCOM). We demonstrate our method, by combining low-level audio and visual features, for classification of a large 200 hour web video corpus. The combined models outperform any of the individual features by 10%. Further, BAYCOM consistently outperforms traditional early and late fusion methods.
Vasant Manohar, Stavros Tsakalidis, Pradeep Natarajan, Rohit Prasad, Premkumar Natarajan
ACM Multimedia4
2011 Robust named entity detection from optical character recognition output
Krishna Subramanian 0001, Rohit Prasad, Premkumar Natarajan
Int. J. Document Anal. Recognit.2
2010 A Semi-Supervised Batch-Mode Active Learning Strategy for Improved Statistical Machine Translation
Sankaranarayanan Ananthakrishnan, Rohit Prasad, David Stallard, Premkumar Natarajan
CoNLL2
2010 Gabor features for offline Arabic handwriting recognition
abstract
Many feature extraction approaches for off-line handwriting recognition (OHR) rely on accurate binarization of gray-level images. However, high-quality binarization of most real-world documents is extremely difficult due to varying characteristics of noises artifacts common in such documents. Unlike most of these features, Gabor features do not require binarization of the document images, and thus are likely to be more robust to noises in document images. To demonstrate the efficacy of our proposed Gabor features, we perform subword recognition for off-line Arabic handwritten images using Support Vector Machines (SVM). We also compare the recognition performance with other binarization based features which have been proven to be effective in capturing shape characteristics of handwritten Arabic subwords, such as GSC (a set of gradient, structure, and concavity features) and skeleton based Graph features. Our preliminary experimental results show that Gabor features outperform Graph features and are slightly better than GSC features for Arabic subword recognition. In addition, by combining Gabor and GSC features, we obtain a significant reduction in classification error rate over using GSC or Gabor features alone.
Huaigu Cao, Rohit Prasad, Anurag Bhardwaj, Premkumar Natarajan
Document Analysis Systems3
2010 The BBN document analysis service: a platform for multilingual document translation
abstract
In this paper, we introduce a new operational platform for end-to-end document image analysis, recognition, and machine translation. The Raytheon BBN Document Analysis Service (BBN DAS) performs the following operations on scanned machine-print document images: (1) image pre-processing and segmentation to identify homogenous zones of text, (2) text recognition to convert the text zones into electronic text, (3) machine translation for converting the text from the native language of the document into English, and (4) document archiving and indexing for effective content-based search. BBN DAS uses a service-oriented architecture (SOA), which offers modularity and scalability for operation on hardware configurations ranging from a laptop to distributed multi-node server environments. This paper describes the platform architecture, the process of configuring it for Arabic newsprint documents and resulting performance results of the Arabic system.
Ehry MacRostie, Rohit Prasad, Stephen Rawls, Matin Kamali, Huaigu Cao, Krishna Subramanian 0001, Premkumar Natarajan
Document Analysis Systems2
2010 Discriminative Sample Selection for Statistical Machine Translation
Sankaranarayanan Ananthakrishnan, Rohit Prasad, David Stallard, Premkumar Natarajan
EMNLP2
2010 Pashto speech recognition with limited pronunciation lexicon
abstract
Automatic speech recognition (ASR) for low resource languages continues to be a difficult problem. In particular, colloquial dialects of Arabic, Farsi, and Pashto pose significant challenges in pronunciation dictionary creation. Therefore, most state-of-the-art ASR engines rely on the grapheme-as-phoneme approach for creating pronunciation dictionaries in these languages. While the grapheme approach simplifies ASR training, it performs significantly worse than a system trained with a high-quality phonetic dictionary. In this paper, we explore two techniques for bridging the performance gap between the grapheme and the phonetic approaches, without requiring manual pronunciations for all the words in the training data. The first approach is based on learning letter-to-sound rules from a small set of manual pronunciations in Pashto, and the second approach uses a hybrid phoneme/grapheme representation for recognition. Through experimental results on colloquial Pashto, we demonstrate that both techniques perform as well as a full phonetic system while requiring manual pronunciations for only a small fraction of the words in the acoustic training data.
Rohit Prasad, Stavros Tsakalidis, Ivan Bulyko, Chia-Lin Kao, Premkumar Natarajan
ICASSP1
2010 Evaluating different confirmation strategies for speech-to-speech translation systems
abstract
Speech-to-speech translation systems have made a great deal of progress in recent years. But users of such systems still face the problem of not knowing whether the system has translated their utterance correctly. Various confirmation strategies can be used to address this problem. Some of these generate a confirmation utterance for the user to approve, such as reading back the ASR result, or performing “back-translation” to translate the system's translation output back into the source language. Other strategies use automated methods such as confidence measures to eliminate likely mistranslations. We propose a methodology for quantitatively evaluating the effectiveness of these different strategies, and present results of experiments using this methodology.
David Stallard, Rohit Prasad, Shankar Ananthakrishnan, Fred Choi, Shirin Saleem, Premkumar Natarajan
ICASSP2
2010 Improvements in HMM Adaptation for Handwriting Recognition Using Writer Identification and Duration Adaptation
abstract
This paper presents two techniques for improving adaptation of hidden Markov models (HMMs) for offline handwriting recognition. The first technique uses a novel writer identification algorithm to select training data for adapting writer-dependent models. This helps us get enough annotated samples for adaptation when the writers of test samples are known to have written some manuscripts in the training set. The second technique adapts the transition probabilities of the HMM using estimated mean of model durations from the initial decoding. Experimental results show significant improvements over the standard unsupervised parameter adaptation in our handwriting recognition system.
Huaigu Cao, Rohit Prasad, Premkumar Natarajan
ICFHR2
2010 Stochastic Segment Model Adaptation for Offline Handwriting Recognition
abstract
In this paper, we present techniques for unsupervised adaptation of stochastic segment models to improve accuracy on large vocabulary offline handwriting recognition (OHR) tasks. We build upon our previous work on stochastic segment modeling for Arabic OHR. In our previous work, stochastic character segments for each n-best hypothesis were generated by a hidden Markov model (HMM) recognizer, and then a segmental model was used as an additional knowledge source for re-ranking the n-best list. Here, we describe a novel framework for unsupervised adaptation. It integrates both HMM and segment model adaptation to achieve significant gains over un-adapted recognition. Experimental results demonstrate the efficacy of our proposed method on a large corpus of handwritten Arabic documents.
Rohit Prasad, Anurag Bhardwaj, Krishna Subramanian 0001, Huaigu Cao, Premkumar Natarajan
ICPR1
2010 Consensus Network Based Hypotheses Combination for Arabic Offline Handwriting Recognition
abstract
Offline handwriting recognition (OHR) is an extremely challenging task because of many factors including variations in writing style, writing device and material, and noise in the scanning and collection process. Due to the diverse nature of the above challenges, it is highly unlikely that a single recognition technique can address all the characteristics of real-world handwritten documents. Therefore, one must consider designing different systems, each addressing specific challenges in the handwritten corpus, and then combining the hypotheses from these diverse systems. To that end, we present an innovative approach for combining hypotheses from multiple handwriting recognition systems. Our approach is based on generating a consensus network using hypotheses from a diverse set of handwriting recognition systems. Next, we decode the consensus network for producing the best possible hypothesis given an error criterion. Experimental results on an Arabic OHR task show that our combination algorithm outperforms the NIST ROVER technique and results in a 7% relative reduction in the word error rate over the single best OHR system.
Rohit Prasad, Matin Kamali, David Belanger 0001, Antti-Veikko I. Rosti, Spyridon Matsoukas, Premkumar Natarajan
ICPR1
2010 Phrase alignment confidence for statistical machine translation
Sankaranarayanan Ananthakrishnan, Rohit Prasad, Premkumar Natarajan
INTERSPEECH2
2010 Multi resolution discriminative models for subvocalic speech recognition
Mark Raugas, Vivek Kumar Rangarajan Sridhar, Rohit Prasad, Premkumar Natarajan
INTERSPEECH3
2010 Mutual information analysis for feature and sensor subset selection in surface electromyography based speech recognition
Vivek Kumar Rangarajan Sridhar, Rohit Prasad, Premkumar Natarajan
INTERSPEECH2
2010 An unsupervised boosting technique for refiningword alignment
abstract
Translation rules extracted from automatic word alignment form the basis of statistical machine translation (SMT) systems. An unsupervised expectation-maximization (EM) algorithm is typically used to obtain a word alignment from parallel corpora. Being statistically-driven, the alignments produced by this technique are often erroneous. In this paper, we propose an unsupervised boosting strategy for refining automatic word alignment with the goal of improving SMT performance. The proposed approach results in fewer unaligned words, a significant reduction in the number of extracted translation phrase pairs, a corresponding improvement in SMT decoding speed, and a consistent improvement in translation accuracy, as measured by BLEU, across multiple language pairs and test sets. The reduction in storage and processing requirements coupled with improved accuracy make the proposed technique ideally suited for interactive translation services, facilitating applications such as mobile speech-to-speech translation.
Sankaranarayanan Ananthakrishnan, Rohit Prasad, Premkumar Natarajan
SLT2
2009 Context-dependent pronunciation modeling for Iraqi ASR
abstract
In this paper, we introduce a novel pronunciation modeling technique that in contrast to existing techniques uses word context information. This context-dependent pronunciation modeling is designed to overcome the challenges posed by absence of diacritics in transcripts for training acoustic models for Arabic dialects. To demonstrate the efficacy of the proposed pronunciation modeling, we present experimental results with both manually created and automatically generated vowelized lexicons on the DARPA TRANSTAC colloquial Iraqi corpus.
Stavros Tsakalidis, Rohit Prasad, Premkumar Natarajan
ICASSP2
2009 Unsupervised HMM Adaptation Using Page Style Clustering
abstract
In this paper we present an innovative two-stage adaptation approach for handwriting recognition that is based on clustering of similar pages in the training data. In our approach, we first perform page clustering on training data using features such as contour slope, pen pressure, writing velocity, and stroke sparseness. Next, we adapt the writer-independent hidden Markov models (HMMs) to each cluster in the training data. While decoding a test page, we first determine the cluster the test page belongs to and then decode the page with the model associated with that cluster. Experimental results with the two-stage adaptation show significant gains on a held-out validation set.
Huaigu Cao, Rohit Prasad, Shirin Saleem, Premkumar Natarajan
ICDAR2
2009 Stochastic Segment Modeling for Offline Handwriting Recognition
abstract
In this paper, we present a novel approach for incorporating structural information into the hidden Markov modeling (HMM) framework for offline handwriting recognition. Traditionally, structural features have been used in recognition approaches that rely on accurate segmentation of words into smaller units (sub-words or characters). However, such segmentation based approaches do not perform well on real-world handwritten images, because breaks and merges in glyphs typically create new connected components that are not observed in the training data. To mitigate the problem of having to derive accurate segmentation from connected components, we present a novel framework where the HMM based recognition system trained on shorter-span features is used to generate the 2D character images (the ldquostochastic segmentsrdquo), and then another classifier that uses structural features extracted from the stochastic character segments generates a new set of scores. Finally, the scores from the HMM system and from structural matching are used in combination to generate a hypothesis that is better than the results from either the HMM or from structural matching alone. We demonstrate the efficacy of our approach by reporting experimental results on a large corpus of handwritten Arabic documents.
Premkumar Natarajan, Krishna Subramanian 0001, Anurag Bhardwaj, Rohit Prasad
ICDAR4
2009 Improvements in BBN's HMM-Based Offline Arabic Handwriting Recognition System
abstract
Offline handwriting recognition of free-flowing Arabic text is a challenging task due to the plethora of factors that contribute to the variability in the data. In this paper, we address some of these sources of variability, and present experimental results on a large corpus of handwritten documents. Specific techniques such as the application of context-dependent Hidden Markov Models (HMMs) for the cursive Arabic script, unsupervised adaptation to account for the stylistic variations across scribes, and image pre-processing to remove ruled-lines are explored. In particular, we proposed a novel integration of structural features in the HMM framework which exclusively results in a 9% relative improvement in performance. Overall, we demonstrate a relative reduction of 17% in word error rate over our baseline Arabic handwriting recognition system.
Shirin Saleem, Huaigu Cao, Krishna Subramanian 0001, Matin Kamali, Rohit Prasad, Premkumar Natarajan
ICDAR5
2009 Nested state indexing in pairwise Markov networks for fast handwritten document image rule-line removal
abstract
The Markov random field (MRF) has been applied to modeling the connectivity constraints of the text in document images for tasks like binarization and rule-line removal. One challenge of applying the MRF is its high computational cost. This paper presents a method using two nested set of states trained to reduce the computational cost of patch-based MRF. The two sets of states are trained at different levels in coarse-to-fine order. We show effective reduction of run time but very little loss of quality using rule-line removal experiments.
Huaigu Cao, Rohit Prasad, Premkumar Natarajan, Venu Govindaraju
ICIP2
2008 End-to-End Trainable Thai OCR System Using Hidden Markov Models
abstract
In this paper we present an end-to-end trainable optical character recognition (OCR) system for recognizing machine-printed text in Thai documents. The end-to-end OCR system is based on a script-independent methodology using hidden Markov models. Our system provides an integrated workflow beginning with annotation and transcription of training images to performing OCR on new images with models trained on transcribed training images. The efficacy of our end-to-end OCR system is demonstrated by rapidly configuring our OCR engine for the Thai script. We present experimental results on Thai documents to highlight the specific challenges posed by the Thai script and analyze the recognition performance as a function of amount of training data.
Kriste Krstovski, Ehry MacRostie, Rohit Prasad, Premkumar Natarajan
Document Analysis Systems3
2008 Semi-supervised topic classification for low resource languages
abstract
In this paper, we present a novel methodology for rapidly developing a topic-based document classification system for a language that has limited resources. Our approach, a hybrid one, combines supervised and unsupervised topic classification techniques. Given that access to native speakers is fairly limited for low resource languages, our approach requires annotating only a few broad “root” topics in the corpus. Next, unsupervised topic discovery (UTD) technique is used to automatically determine finer topics within the root topics. Lastly, we use the recently developed unsupervised topic clustering technique to organize the corpus into a hierarchical structure that enables browsing documents at multiple levels of granularity. Recognizing the need for reducing false alarms during runtime, we describe rejection techniques for discarding off-topic documents.
Daben Liu, Sam McVeety, Rohit Prasad, Premkumar Natarajan
ICASSP3
2008 Multi-frame combination for robust videotext recognition
abstract
Optical character recognition (OCR) of overlaid text in video streams is a challenging problem due to various factors including the presence of dynamic backgrounds, color, and low resolution. In video feeds such as Broadcast News, a particular overlaid text region usually persists for multiple frames during which the background may or may not vary. In this paper we explore two innovative techniques that exploit such multi-frame persistence of videotext. The first technique uses multiple instances to generate a single enhanced image for recognition. The second technique uses the NIST ROVER algorithm developed for speech recognition to combine 1-best hypotheses from different frames of a text region. Significant improvement in the word error rate (WER) is obtained by using ROVER when compared to recognizing a single instance. The WER is further reduced by combining hypotheses from frame instances, which were generated using character models trained with different binarization thresholds. A 20% relative reduction in the WER was achieved for multi-frame combination over decoding a single frame instance.
Rohit Prasad, Shirin Saleem, Ehry MacRostie, Premkumar Natarajan, Michael Decerbo
ICASSP1
2008 Recent improvements and performance analysis of ASR and MT in a speech-to-speech translation system
abstract
We report on recent ASR and MT work on our English/Iraqi Arabic speech-to-speech translation system. We present detailed results for both objective and subjective evaluations of translation quality, along with a detailed analysis and categorization of translation errors. We also present novel ideas for quantifying the relative importance of different subjective error categories, and for assigning the blame for an error to a particular phrase pair in the translation model.
David Stallard, Chia-Lin Kao, Kriste Krstovski, Daben Liu, Premkumar Natarajan, Rohit Prasad, Shirin Saleem, Krishna Subramanian 0001
ICASSP6
2008 Robust named entity detection in videotext using character lattices
abstract
Text in video sequences can provide key indexing information. In particular, videotext is rich in named entities (NEs) and detection of such entities is critical for search applications. Traditional approaches for detecting NEs in OCR output look for these NEs in the single-best recognition results. Due to inevitable presence of recognition errors in the single-best output, such approaches usually result in low recall. Given that a lattice is more likely to contain the correct answer, we explore NE detection from character lattices produced by our videotext OCR system. Furthermore, we use an approximate match criterion that allows insertion of punctuations during lookup. Experimental results show a 50% relative improvement in NE recall using lattices over exact lookup in the 1-best hypothesis. Since the improvement in recall is accompanied by a large number of false positives, we present techniques for reducing false alarms. In addition, we describe efficient techniques for reducing the time for detecting NEs.
Krishna Subramanian 0001, Rohit Prasad, Ehry MacRostie, Premkumar Natarajan
ICASSP2
2008 Improvements in hidden Markov model based Arabic OCR
abstract
This paper describes recent advances in hidden Markov model (HMM) based OCR for machine-printed arabic documents. A combination of script-independent and script-specific techniques are applied to glyph models and language models (LM). Script-independent techniques we applied are higher order n-gram LMs for N-best rescoring and discriminative estimation of glyph HMMs. Arabic specific techniques include the use of context-dependent HMMs for glyph modeling and Parts-of-Arabic-Words in language modeling. We present experimental results that demonstrate a 40% relative reduction in word error rate over the baseline configuration on a corpus of machine-printed Arabic documents.
Rohit Prasad, Shirin Saleem, Matin Kamali, Ralf Meermeier, Premkumar Natarajan
ICPR1
2008 Recent improvements in BBN's English/Iraqi speech-to-speech translation system
abstract
We report on recent improvements in our English/Iraqi Arabic speech-to-speech translation system. User interface improvements include a novel parallel approach to user confirmation which makes confirmation cost-free in terms of dialog duration. Automatic speech recognition improvements include the incorporation of state-of-the-art techniques in feature transformation and discriminative training. Machine translation improvements include a novel combination of multiple alignments derived from various pre-processing techniques, such as Arabic segmentation and English word compounding, higher order N-grams for target language model, and use of context in form of semantic classes and part-of-speech tags.
Fred Choi, Stavros Tsakalidis, Shirin Saleem, Chia-Lin Kao, Ralf Meermeier, Kriste Krstovski, Christine Moran, Krishna Subramanian 0001, Rohit Prasad, Premkumar Natarajan
SLT9
2008 Name aware speech-to-speech translation for English/Iraqi
abstract
In this paper, we describe a novel approach that exploits intra-sentence and dialog-level context for improving translation performance on spoken Iraqi utterances that contain named entities (NEs). Dialog-level context is used to predict whether the Iraqi response is likely to contain names and the intra-sentence context is used to determine words that are named entities. While we do not address the problem of translating out-of-vocabulary (OOV) NEs in spoken utterances, we show that our approach is capable of translating OOV names in text input. To demonstrate efficacy of our approach, we present results on internal test set as well as the 2008 June DARPA TRANSTAC name evaluation set.
Rohit Prasad, Christine Moran, Fred Choi, Ralf Meermeier, Shirin Saleem, Chia-Lin Kao, David Stallard, Premkumar Natarajan
SLT1
2007 Semantic translation error rate for evaluating translation systems
abstract
In this paper, we introduce a new metric which we call the semantic translation error rate, or STER, for evaluating the performance of machine translation systems. STER is based on the previously published translation error rate (TER) (Snover et al., 2006) and METEOR (Banerjee and Lavie, 2005) metrics. Specifically, STER extends TER in two ways: first, by incorporating word equivalence measures (WordNet and Porter stemming) standardly used by METEOR, and second, by disallowing alignments of concept words to non-concept words (aka stop words). We show how these features make STER alignments better suited for human-driven analysis than standard TER. We also present experimental results that show that STER is better correlated to human judgments than TER. Finally, we compare STER to METEOR, and illustrate that METEOR scores computed using the STER alignments have similar statistical properties to METEOR scores computed using METEOR alignments.
Krishna Subramanian 0001, David Stallard, Rohit Prasad, Shirin Saleem, Premkumar Natarajan
ASRU3
2007 Optimal Estimation of Rejection Thresholds for Topic Spotting
abstract
In many applications of topic spotting technology, especially those that require a human review of in-topic documents, a low false alarm rate is a key requirement. Topic spotting techniques typically include a rejection scheme to filter out off-topic documents. In this paper we present a robust methodology for rejecting off-topic messages that, in addition to modeling the topics of interest, uses a so-called alternate model for topics that are not included in the set of topics of interest. Specifically, we introduce two novel techniques for estimating topic-specific rejection thresholds - a parametric technique that can be viewed as transformation of topic-independent thresholds, and a nonparametric technique based on constrained optimization of false rejections subject to a pre-specified number of false acceptances. Our experiments on newsgroup messages demonstrate that when adequate training data is available topic-specific threshold estimation techniques can outperform topic-independent thresholds in terms of the ROC curve.
Krishna Subramanian 0001, Rohit Prasad, Premkumar Natarajan, Richard M. Schwartz
ICASSP (4)2
2007 Robust Page Segmentation Based on Smearing and Error Correction Unifying Top-down and Bottom-up Approaches
abstract
In this paper we present a robust multi-pass page segmentation algorithm. The first pass uses a modified smearing algorithm and the second pass performs a hybrid of bottom-up and top-down segmentation on the output of the first pass. Unlike traditional approaches, the bottom-up and top-down steps are based on primitive results of a smearing based page segmentation algorithm. Therefore, "split" and "merge" processes start with text blocks that are mostly true text blocks but a few of them are either touching or broken. We present experimental results on newspaper and journal documents from different languages to demonstrate the robustness and language independence of our approach.
Huaigu Cao, Rohit Prasad, Premkumar Natarajan, Ehry MacRostie
ICDAR2
2007 Improvements in machine translation for English/iraqi speech translation
Shirin Saleem, Krishna Subramanian 0001, Rohit Prasad, David Stallard, Chia-Lin Kao, Premkumar Natarajan, Raid Suleiman
INTERSPEECH3
2007 The BBN 2007 displayless English/iraqi speech-to-speech translation system
David Stallard, Fred Choi, Chia-Lin Kao, Kriste Krstovski, Premkumar Natarajan, Rohit Prasad, Shirin Saleem, Krishna Subramanian 0001
INTERSPEECH6
2007 Finding structure in noisy text: topic classification and unsupervised clustering
Premkumar Natarajan, Rohit Prasad, Krishna Subramanian 0001, Shirin Saleem, Fred Choi, Richard M. Schwartz
Int. J. Document Anal. Recognit.2
2006 Colloquial Iraqi ASR for speech translation
Shirin Saleem, Rohit Prasad, Premkumar Natarajan
INTERSPEECH2
2006 A hybrid phrase-based/statistical speech translation system
David Stallard, Fred Choi, Kriste Krstovski, Premkumar Natarajan, Rohit Prasad, Shirin Saleem
INTERSPEECH5
2006 Design and Evaluation of the 2006 BBN English/Iraqi Two-Way speech Translation System
abstract
In this paper, we present a 2-way speech-to-speech translation system for English and Iraqi colloquial Arabic, the dialect of Arabic spoken by ordinary people in Iraq. The application domain of the system is military force protection, including municipal services surveys, detainee screening, and descriptions of people, houses, vehicles, etc. The system uses statistical speech recognition, and a combination of prerecorded questions and statistical machine translation with speech synthesis to translate the speech recognition output. We present evaluation results, along with an analysis of the gap between Iraqi-to-English and English-to-Iraqi translation performance.
David Stallard, Fred Choi, Kriste Krstovski, Premkumar Natarajan, Rohit Prasad, Shirin Saleem, Raid Suleiman
SLT5
2006 Advances in transcription of broadcast news and conversational telephone speech within the combined EARS BBN/LIMSI system
abstract
This paper describes the progress made in the transcription of broadcast news (BN) and conversational telephone speech (CTS) within the combined BBN/LIMSI system from May 2002 to September 2004. During that period, BBN and LIMSI collaborated in an effort to produce significant reductions in the word error rate (WER), as directed by the aggressive goals of the Effective, Affordable, Reusable, Speech-to-text [Defense Advanced Research Projects Agency (DARPA) EARS] program. The paper focuses on general modeling techniques that led to recognition accuracy improvements, as well as engineering approaches that enabled efficient use of large amounts of training data and fast decoding architectures. Special attention is given on efforts to integrate components of the BBN and LIMSI systems, discussing the tradeoff between speed and accuracy for various system combination strategies. Results on the EARS progress test sets show that the combined BBN/LIMSI system achieved relative reductions of 47% and 51% on the BN and CTS domains, respectively.
Spyridon Matsoukas, Jean-Luc Gauvain, Gilles Adda, Thomas Colthurst, Chia-Lin Kao, Owen Kimball, Lori Lamel, Fabrice Lefèvre, Jeff Z. Ma, John Makhoul, Long Nguyen 0001, Rohit Prasad, Richard M. Schwartz, Holger Schwenk, Bing Xiang
IEEE Trans. Speech Audio Process.12
2005 Performance Improvements to the BBN Byblos OCR System
abstract
In this paper, we describe four recent enhancements to the BBN Byblos OCR system, a multilingual HMM-based character recognition system which has been demonstrated on a variety of languages, including English, Arabic, Chinese, and Japanese. These enhancements are implemented as optional extensions to the system and provide improved performance for certain scripts or domains. Projection-based re-estimation of line boundaries reduces instability in the presence of some types of noise. An alternate modeling strategy used in the first of two recognition search passes substantially increases speed on languages with a large number of characters. Another speed improvement comes from automatic discovery and modeling of sub-characters. The use of heteroschedastic linear discriminant analysis (HLDA) makes modeling more tractable by reducing feature-space dimensionality.
Michael Decerbo, Premkumar Natarajan, Rohit Prasad, Ehry MacRostie
ICDAR3
2005 Character Duration Modeling for Speed Improvements in the BBN Byblos OCR System
abstract
In this paper, we describe a recent enhancement to our HMM-based OCR system that results in a significant increase in the speed of the system without any impact on recognition accuracy. Recognition speed is, in part, a function of the number of distinct HMMs that constitute the model set. As a result, the recognition speed is much slower for ideographic scripts, such as Chinese and Japanese which contain thousands of glyphs, than for alphabetic scripts such as Latin and Arabic. In our current OCR system, methods like sub-character modeling and Gaussian shortlists are used to reduce the processing time. In this paper, we describe a simple character-based duration modeling technique that puts a duration constraint on the number of frames for which a character can stay active. Character durations were obtained from automatically labeled training data and a probability mass function (histogram) was used to model character durations. The use of a duration model yielded a 37% improvement in speed with no loss in accuracy.
Premkumar Natarajan, Ram Sundaram, Rohit Prasad, Ehry MacRostie
ICDAR3
2005 The 2004 BBN 1xRT recognition systems for English broadcast news and conversational telephone speech
abstract
This paper describes the BBN real-time recognition systems used in the 2004 Rich Transcription (RT) benchmark test for the English Conversational Telephone Speech (CTS) and Broadcast News (BN) tasks. We describe the system architecture, along withthe algorithms weused inorder to reduce computation with minimal impact on recognition accuracy. Particular choices in the design of thefinal system are analyzed toshow the trade-offs between speed and accuracy. We also present recently developed new architecture for the real-time systems, which outperforms the systems we submitted for the RT04 benchmark tests for both domains.
Spyridon Matsoukas, Rohit Prasad, Srinivas Laxminarayan, Bing Xiang, Long Nguyen 0001, Richard M. Schwartz
INTERSPEECH2
2005 The 2004 BBN/LIMSI 20xRT English conversational telephone speech recognition system
abstract
In this paper we describe the English Conversational Telephone Speech (CTS) recognition system jointly developed by BBN and LIMSI under the DARPA EARS program for the 2004 evalua-tion conducted by NIST. The 2004 BBN/LIMSI system achieved a word error rate (WER) of 13.5 % at 18.3xRT (real-time as mea-sured on Pentium 4 Xeon 3.4 GHz Processor) on the EARS progress test set. This translates into a 22.8 % relative improvement in WER over the 2003 BBN/LIMSI EARS evaluation system, which was run without any time constraints. In addition to reporting on the system architecture and the evaluation results, we also highlight the significant improvements made at both sites. 1.
Rohit Prasad, Spyridon Matsoukas, Chia-Lin Kao, Jeff Z. Ma, Dongxin Xu, Thomas Colthurst, Owen Kimball, Richard M. Schwartz, Jean-Luc Gauvain, Lori Lamel, Holger Schwenk, Gilles Adda, Fabrice Lefèvre
INTERSPEECH1
2004 Speech recognition in multiple languages and domains: the 2003 BBN/LIMSI EARS system
abstract
We report on the results of the first evaluations for the BBN/LIMSI system under the new DARPA EARS program. The evaluations were carried out for conversational telephone speech (CTS) and broadcast news (BN) for three languages: English, Mandarin, and Arabic. In addition to providing system descriptions and evaluation results, the paper highlights methods that worked well across the two domains and those few that worked well on one domain but not the other. For the BN evaluations, which had to be run under 10 times real-time, we demonstrated that a joint BBN/LIMSI system with a time constraint achieved better results than either system alone.
Richard M. Schwartz, Thomas Colthurst, Nicolae Duta, Herbert Gish, Rukmini Iyer, Chia-Lin Kao, Daben Liu, Owen Kimball, Jeff Z. Ma, John Makhoul, Spyridon Matsoukas, Long Nguyen 0001, Mohammed Noamany, Rohit Prasad, Bing Xiang, Dongxin Xu, Jean-Luc Gauvain, Lori Lamel, Holger Schwenk, Gilles Adda, Langzhou Chen
ICASSP (3)14
2002 A scalable architecture for Directory Assistance automation
abstract
We present a novel architecture for providing automated telephone Directory Assistance (DA). The architecture couples a large-vocabulary, statistical n-gram, speech recognition engine with a statistical retrieval system. The use of a statistical n-gram allows for the recognition of unconstrained spoken queries while the statistical retrieval engine allows for an inexact match between a particular spoken query and the training data. Allowing for unconstrained recognition and an inexact match provides the framework for high levels of automation. Once the retrieval engine returns a ranked set of frequently requested telephone numbers (FRN), the rejection module uses a classifier to compute a confidence-like score that is used to make the automation decision. With actual customer calls into an operational, automated DA call center and an FRN set size of 25000 numbers, the new architecture is capable of delivering more than 17% correct automation at a false accept rate of 0.76%.
Premkumar Natarajan, Rohit Prasad, Richard M. Schwartz, John Makhoul
ICASSP2
2002 Speech-enabled natural language call routing: BBN call director
Premkumar Natarajan, Rohit Prasad, Bernhard Suhm, Daniel McCarthy
INTERSPEECH2
2002 Automatic transcription of courtroom speech
Rohit Prasad, Long Nguyen 0001, Richard M. Schwartz, John Makhoul
INTERSPEECH1