VLDB 2026 Research / reviewers in the wild / expert
Premkumar Natarajan
dblp:40/2362 · also Prem Natarajan, Premkumar S. Natarajan
· DBLP profile ↗
157ranked-venue papers
8as first author
23since 2021 · last 2026
0000-0002-4386-6651ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 107 · 7 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 94 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 32 · 4 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rescind: Countering Image Misconduct in Biomedical Publications with Vision-Language and State-Space ModelingabstractScientific image manipulation in biomedical publications poses a growing threat to research integrity and reproducibility. Unlike natural image forensics, biomedical forgery detection is uniquely challenging due to domain-specific artifacts, complex textures, and unstructured figure layouts. We present the first vision-language guided framework for both generating and detecting biomedical image forgeries. By combining diffusion-based synthesis with vision-language prompting, our method enables realistic and semantically controlled manipulations—including duplication, splicing, and region removal—across diverse biomedical modalities. We introduce Rescind, a large-scale benchmark featuring fine-grained annotations and modality-specific splits, and propose Integscan, a structured state-space modeling framework that integrates attention-enhanced visual encoding with prompt-conditioned semantic alignment for precise forgery localization. To ensure semantic fidelity, we incorporate a VLM-based verification loop that filters generated forgeries based on consistency with intended prompts. Extensive experiments on Rescind and existing benchmarks demonstrate that Integscan achieves state-of-the-art performance in both detection and localization, establishing a strong foundation for automated scientific integrity analysis. Soumyaroop Nandi, Premkumar Natarajan |
AAAI | 2 |
| 2026 | Can Image Splicing and Copy-Move Forgery Be Detected by the Same Model? Forensim: An Attention-Based State-Space ApproachabstractWe introduce Forensim, an attention-based state-space framework for image forgery detection that jointly local-izes both manipulated (target) and source regions. Unlike traditional approaches that rely solely on artifact cues to detect spliced or forged areas, Forensim is designed to capture duplication patterns crucial for understanding context. In scenarios like protest imagery, detecting only the forged region—e.g., a duplicated act of violence inserted into a peaceful crowd—can mislead interpretation, highlighting the need for joint source-target localization. Forensim outputs three-class masks (pristine, source, target) and supports detection of both splicing and copy-move forgeries within a unified architecture. We propose a visual state-space model that leverages normalized attention maps to identify internal similarities, paired with a region-based block-attention module to distinguish manipulated regions. This design enables end-to-end training and precise localization. Forensim achieves state-of-the-art performance on standard benchmarks. We also release CMFD Anything, a new dataset addressing limitations of existing copy-move forgery datasets. Project page and code. Soumyaroop Nandi, Premkumar Natarajan |
WACV | 2 |
| 2024 | Agenda-Driven Question Generation: A Case Study in the Courtroom DomainabstractThis paper introduces a novel problem of automated question generation for courtroom examinations, CourtQG. While question generation has been studied in domains such as educational testing and product description, CourtQG poses several unique challenges owing to its non-cooperative and agenda-driven nature. Specifically, not only the generated questions need to be relevant to the case and underlying context, they also have to achieve certain objectives such as challenging the opponent’s arguments and/or revealing potential inconsistencies in their answers. We propose to leverage large language models (LLM) for CourtQG by fine-tuning them on two auxiliary tasks, agenda explanation (i.e., uncovering the underlying intents) and question type prediction. We additionally propose cold-start generation of questions from background documents without relying on examination history. We construct a dataset to evaluate our proposed method and show that it generates better questions according to standard metrics when compared to several baselines. Yi R. Fung 0001, Aram Galstyan, Heng Ji 0001, Premkumar Natarajan |
LREC/COLING | 5 |
| 2023 | User-Controllable Arbitrary Style Transfer via Entropy RegularizationabstractEnsuring the overall end-user experience is a challenging task in arbitrary style transfer (AST) due to the subjective nature of style transfer quality. A good practice is to provide users many instead of one AST result. However, existing approaches require to run multiple AST models or inference a diversified AST (DAST) solution multiple times, and thus they are either slow in speed or limited in diversity. In this paper, we propose a novel solution ensuring both efficiency and diversity for generating multiple user-controllable AST results by systematically modulating AST behavior at run-time. We begin with reformulating three prominent AST methods into a unified assign-and-mix problem and discover that the entropies of their assignment matrices exhibit a large variance. We then solve the unified problem in an optimal transport framework using the Sinkhorn-Knopp algorithm with a user input ε to control the said entropy and thus modulate stylization. Empirical results demonstrate the superiority of the proposed solution, with speed and stylization quality comparable to or better than existing AST and significantly more diverse than previous DAST works. Code is available at https://github.com/cplusx/eps-Assign-and-Mix. Jiaxin Cheng, Yue Wu 0001, Ayush Jaiswal, Xu Zhang 0022, Pradeep Natarajan, Premkumar Natarajan |
AAAI | 6 |
| 2023 | MASSIVE: A 1M-Example Multilingual Natural Language Understanding Dataset with 51 Typologically-Diverse LanguagesabstractJack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie, Kay Rottmann, Ana Sanchez, Aaron Nash, Liam Urbach, Vishesh Kakarala, Richa Singh, Swetha Ranganath, Laurie Crist, Misha Britan, Wouter Leeuwis, Gokhan Tur, Prem Natarajan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Jack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie, Kay Rottmann, Ana Sanchez, Aaron Nash, Liam Urbach, Vishesh Kakarala, Swetha Ranganath, Laurie Crist, Misha Britan, Wouter Leeuwis, Gökhan Tür, Premkumar Natarajan |
ACL (1) | 16 |
| 2023 | TAGPRIME: A Unified Framework for Relational Structure ExtractionabstractI-Hung Hsu, Kuan-Hao Huang, Shuning Zhang, Wenxin Cheng, Prem Natarajan, Kai-Wei Chang, Nanyun Peng. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. I-Hung Hsu, Kuan-Hao Huang, Wenxin Cheng, Premkumar Natarajan, Kai-Wei Chang 0001, Nanyun Peng 0001 |
ACL (1) | 5 |
| 2023 | AMPERE: AMR-Aware Prefix for Generation-Based Event Argument Extraction ModelabstractEvent argument extraction (EAE) identifies event arguments and their specific roles for a given event.Recent advancement in generationbased EAE models has shown great performance and generalizability over classificationbased models.However, existing generationbased EAE models mostly focus on problem reformulation and prompt design, without incorporating additional information that has been shown to be effective for classification-based models, such as the abstract meaning representation (AMR) of the input passages.Incorporating such information into generation-based models is challenging due to the heterogeneous nature of the natural language form prevalently used in generation-based models and the structured form of AMRs.In this work, we study strategies to incorporate AMR into generationbased EAE models.We propose AMPERE, which generates AMR-aware prefixes for every layer of the generation model.Thus, the prefix introduces AMR information to the generationbased EAE model and then improves the generation.We also introduce an adjusted copy mechanism to AMPERE to help overcome potential noises brought by the AMR graph.Comprehensive experiments and analyses on ACE2005 and ERE datasets show that AMPERE can get 4% -10% absolute F1 score improvements with reduced training data and it is in general powerful across different training sizes. I-Hung Hsu, Zhiyu Xie 0001, Kuan-Hao Huang, Premkumar Natarajan, Nanyun Peng 0001 |
ACL (1) | 4 |
| 2023 | DeepMaven: Deep Question Answering on Long-Distance Movie/TV Show Videos with Multimedia Knowledge Extraction and SynthesisabstractYi Fung, Han Wang, Tong Wang, Ali Kebarighotbi, Mohit Bansal, Heng Ji, Prem Natarajan. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Yi R. Fung 0001, Ali Kebarighotbi, Mohit Bansal, Heng Ji 0001, Premkumar Natarajan |
EACL | 7 |
| 2023 | Alexa Arena: A User-Centric Interactive Platform for Embodied AIabstractWe introduce Alexa Arena, a user-centric simulation platform to facilitate research in building assistive conversational embodied agents. Alexa Arena features multi-room layouts and an abundance of interactable objects. With user-friendly graphics and control mechanisms, the platform supports the development of gamified robotic tasks readily accessible to general human users, allowing high-efficiency data collection and EAI system evaluation. Along with the platform, we introduce a dialog-enabled task completion benchmark with online human evaluations. Qiaozi Gao, Govind Thattai, Suhaila M. Shakiah, Xiaofeng Gao 0002, Shreyas Pansare, Vasu Sharma, Gaurav S. Sukhatme, Hangjie Shi, Bofei Yang, Lucy Hu, Karthika Arumugam, Shui Hu, Matthew Wen, Dinakar Guthy, Shunan Chung, Rohan Khanna, Osman Ipek, Leslie Ball, Kate Bland, Heather Rocker, Michael Johnston, Reza Ghanadan, Dilek Hakkani-Tür, Premkumar Natarajan |
NeurIPS | 25 |
| 2023 | Incorporating Fairness in Large Scale NLU SystemsabstractNLU models power several user facing experiences such as conversations agents and chat bots. Building NLU models typically consist of 3 stages: a) building or finetuning a pre-trained model b) distilling or fine-tuning the pre-trained model to build task specific models and, c) deploying the task-specific model to production. In this presentation, we will identify fairness considerations that can be incorporated in the aforementioned three stages in the life-cycle of NLU model building: (i) selection/building of a large scale language model, (ii) distillation/fine-tuning the large model into task specific model and, (iii) deployment of the task specific model. We will present select metrics that can be used to quantify fairness in NLU models and fairness enhancement techniques that can be deployed in each of these stages. Finally, we will share some recommendations to successfully implement fairness considerations when building an industrial scale NLU system. Rahul Gupta 0001, Lisa Bauer, Kai-Wei Chang 0001, Jwala Dhamala, Aram Galstyan, Palash Goyal, Avni Khatri, Rohit Parimi, Charith Peris, Apurv Verma, Richard S. Zemel, Premkumar Natarajan |
WSDM | 13 |
| 2022 | Multilingual Generative Language Models for Zero-Shot Cross-Lingual Event Argument ExtractionabstractWe present a study on leveraging multilingual pre-trained generative language models for zero-shot cross-lingual event argument extraction (EAE).By formulating EAE as a language generation task, our method effectively encodes event structures and captures the dependencies between arguments.We design language-agnostic templates to represent the event argument structures, which are compatible with any language, hence facilitating the cross-lingual transfer.Our proposed model finetunes multilingual pre-trained generative language models to generate sentences that fill in the language-agnostic template with arguments extracted from the input passage.The model is trained on source languages and is then directly applied to target languages for event argument extraction.Experiments demonstrate that the proposed model outperforms the current state-of-the-art models on zero-shot cross-lingual EAE.Comprehensive studies and error analyses are presented to better understand the advantages and the current limitations of using generative language models for zero-shot cross-lingual transfer EAE. *The authors contribute equally.Attacker Place Target Attacker Target 接近高级军官的消息灵通人士 说,南斯拉夫 军队 不会离 开军营去干涉 反对派 起义。 Australian commandos , who have been operating deep in Iraq , destroyed a command and control post and killed a number of soldiers. Kuan-Hao Huang, I-Hung Hsu, Premkumar Natarajan, Kai-Wei Chang 0001, Nanyun Peng 0001 |
ACL (1) | 3 |
| 2022 | Transform-Retrieve-Generate: Natural Language-Centric Outside-Knowledge Visual Question AnsweringabstractOutside-knowledge visual question answering (OK-VQA) requires the agent to comprehend the image, make use of relevant knowledge from the entire web, and digest all the information to answer the question. Most previous works address the problem by first fusing the image and question in the multi-modal space, which is inflexible for further fusion with a vast amount of external knowledge. In this paper, we call for an alternative paradigm for the OK-VQA task, which transforms the image into plain text, so that we can enable knowledge passage retrieval, and generative question-answering in the natural language space. This paradigm takes advantage of the sheer volume of gigantic knowledge bases and the richness of pretrained language models. A Transform-Retrieve-Generate framework (TRiG) framework is proposed11The code of this work will be made public., which can be plug-and-played with alternative image-to-text models and textual knowledge bases. Experimental results show that our TRiG framework outperforms all state-of-the-art supervised methods by at least 11.1 % absolute margin. Feng Gao 0013, Qing Ping, Govind Thattai, Aishwarya N. Reganti, Ying Nian Wu, Premkumar Natarajan |
CVPR | 6 |
| 2022 | MONet: Multi-Scale Overlap Network for Duplication Detection in Biomedical ImagesabstractManipulation of biomedical images to misrepresent experimental results has plagued the biomedical community for a while. Recent interest in the problem led to the curation of a dataset and associated tasks to promote the development of biomedical forensic methods. Of these, the largest manipulation detection task focuses on the detection of duplicated regions between images. Traditional computer-vision based forensic models trained on natural images are not designed to overcome the challenges presented by biomedical images. We propose a multi-scale overlap detection model to detect duplicated image regions. Our model is structured to find duplication hierarchically, so as to reduce the number of patch operations. It achieves state-of-the-art performance overall and on multiple biomedical image categories. Ekraam Sabir, Soumyaroop Nandi, Wael Abd-Almageed, Premkumar Natarajan |
ICIP | 4 |
| 2022 | Alexa Teacher Model: Pretraining and Distilling Multi-Billion-Parameter Encoders for Natural Language Understanding SystemsabstractWe present results from a large-scale experiment on pretraining encoders with non-embedding parameter counts ranging from 700M to 9.3B, their subsequent distillation into smaller models ranging from 17M-170M parameters, and their application to the Natural Language Understanding (NLU) component of a virtual assistant system. Though we train using 70% spoken-form data, our teacher models perform comparably to XLM-R and mT5 when evaluated on the written-form Cross-lingual Natural Language Inference (XNLI) corpus. We perform a second stage of pretraining on our teacher models using in-domain data from our system, improving error rates by 3.86% relative for intent classification and 7.01% relative for slot filling. We find that even a 170M-parameter model distilled from our Stage 2 teacher model has 2.88% better intent classification and 7.69% better slot filling error rates when compared to the 2.3B-parameter teacher trained only on public data (Stage 1), emphasizing the importance of in-domain data for pretraining. When evaluated offline using labeled NLU data, our 17M-parameter Stage 2 distilled model outperforms both XLM-R Base (85M params) and DistillBERT (42M params) by 4.23% to 6.14%, respectively. Finally, we present results from a full virtual assistant experimentation platform, where we find that models trained using our pretraining and distillation pipeline outperform models distilled from 85M-parameter teachers by 3.74%-4.91% on an automatic measurement of full-system user dissatisfaction. Jack FitzGerald, Shankar Ananthakrishnan, Konstantine Arkoudas, Davide Bernardi, Abhishek Bhagia, Claudio Delli Bovi, Jin Cao 0003, Rakesh Chada, Amit Chauhan, Luoxin Chen, Anurag Dwarakanath, Satyam Dwivedi, Turan Gojayev, Karthik Gopalakrishnan 0001, Thomas Gueudré, Dilek Hakkani-Tür, Wael Hamza, Jonathan J. Hüser, Kevin Martin Jose, Haidar Khan, Beiye Liu, Jianhua Lu, Alessandro Manzotti, Pradeep Natarajan, Karolina Owczarzak, Gokmen Oz, Enrico Palumbo, Charith Peris, Chandana Satya Prakash, Stephen Rawls, Andy Rosenbaum, Anjali Shenoy, Saleh Soltan, Mukund Sridhar, Lizhen Tan, Fabian Triefenbach, Pan Wei, Shuai Zheng 0004, Gökhan Tür, Premkumar Natarajan |
KDD | 41 |
| 2022 | DEGREE: A Data-Efficient Generation-Based Event Extraction ModelabstractI-Hung Hsu, Kuan-Hao Huang, Elizabeth Boschee, Scott Miller, Prem Natarajan, Kai-Wei Chang, Nanyun Peng. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. I-Hung Hsu, Kuan-Hao Huang, Elizabeth Boschee, Premkumar Natarajan, Kai-Wei Chang 0001, Nanyun Peng 0001 |
NAACL-HLT | 5 |
| 2021 | Societal Biases in Language Generation: Progress and ChallengesabstractEmily Sheng, Kai-Wei Chang, Prem Natarajan, Nanyun Peng. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Emily Sheng, Kai-Wei Chang 0001, Premkumar Natarajan, Nanyun Peng 0001 |
ACL/IJCNLP (1) | 3 |
| 2021 | Style-Aware Normalized Loss for Improving Arbitrary Style TransferabstractNeural Style Transfer (NST) has quickly evolved from single-style to infinite-style models, also known as Arbitrary Style Transfer (AST). Although appealing results have been widely reported in literature, our empirical studies on four well-known AST approaches (GoogleMagenta [14], AdaIN [19], LinearTransfer [29], and SANet [37]) show that more than 50% of the time, AST stylized images are not acceptable to human users, typically due to under- or over-stylization. We systematically study the cause of this imbalanced style transferability (IST ) and propose a simple yet effective solution to mitigate this issue. Our studies show that the IST issue is related to the conventional AST style loss, and reveal that the root cause is the equal weightage of training samples irrespective of the properties of their corresponding style images, which biases the model towards certain styles. Through investigation of the theoretical bounds of the AST style loss, we propose a new loss that largely overcomes IST . Theoretical analysis and experimental results validate the effectiveness of our loss, with over 80% relative improvement in style deception rate and 98% relatively higher preference in human evaluation. Jiaxin Cheng, Ayush Jaiswal, Yue Wu 0001, Pradeep Natarajan, Premkumar Natarajan |
CVPR | 5 |
| 2021 | Learning Better Visual Dialog Agents With Pretrained Visual-Linguistic RepresentationabstractGuessWhat?! is a visual dialog guessing game which incorporates a Questioner agent that generates a sequence of questions, while an Oracle agent answers the respective questions about a target object in an image. Based on this dialog history between the Questioner and the Oracle, a Guesser agent makes a final guess of the target object. While previous work has focused on dialogue policy optimization and visual-linguistic information fusion, most work learns the vision-linguistic encoding for the three agents solely on the GuessWhat?! dataset without shared and prior knowledge of vision-linguistic representation. To bridge these gaps, this paper proposes new Oracle, Guesser and Questioner models that take advantage of a pretrained vision-linguistic model, VilBERT. For Oracle model, we introduce a two-way background/target fusion mechanism to understand both intra and inter-object questions. For Guesser model, we introduce a state-estimator that best utilizes VilBERT’s strength in single-turn referring expression comprehension. For the Questioner, we share the state-estimator from pretrained Guesser with Questioner to guide the question generator. Experimental results show that our proposed models outperform state-of-the-art models significantly by 7%, 10%, 12% for Oracle, Guesser and End-to-End Questioner respectively. Tao Tu 0002, Qing Ping, Govindarajan Thattai, Gökhan Tür, Premkumar Natarajan |
CVPR | 5 |
| 2021 | Lifelong Event Detection with Knowledge TransferabstractTraditional supervised Information Extraction (IE) methods can extract structured knowledge elements from unstructured data, but they are limited to a pre-defined target ontology.In reality, the ontology of interest may change over time, adding emergent new types or more fine-grained subtypes.We propose a new lifelong learning framework to address this challenge.We focus on lifelong event detection as an exemplar case and propose a new problem formulation that is also generalizable to other IE tasks.In event detection and more general IE tasks, rich correlations or semantic relatedness exist among hierarchical knowledge element types.In our proposed framework, knowledge is being transferred between learned old event types and new event types.Specifically, we update old knowledge with the mentions of new event types using a selftraining loss.In addition, we aggregate the representations of old event types based on their similarities with new event types to initialize the representations of new event types.Experimental results show that our framework outperforms competitive baselines with a 5.1% absolute gain in the F1 score.Moreover, our proposed framework can boost the F1 score for over 30% absolute gain on some new long-tail rare event types with few training instances.Our knowledge transfer module improves the performance on both learned event types and new event types under the lifelong learning setting, showing that it helps consolidate old knowledge and improve novel knowledge acquisition.1 Pengfei Yu 0001, Heng Ji 0001, Premkumar Natarajan |
EMNLP (1) | 3 |
| 2021 | SIGN: Spatial-information Incorporated Generative Network for Generalized Zero-shot Semantic SegmentationabstractUnlike conventional zero-shot classification, zero-shot semantic segmentation predicts a class label at the pixel level instead of the image level. When solving zero-shot semantic segmentation problems, the need for pixel-level prediction with surrounding context motivates us to incorporate spatial information using positional encoding. We improve standard positional encoding by introducing the concept of Relative Positional Encoding, which integrates spatial information at the feature level and can handle arbitrary image sizes. Furthermore, while self-training is widely used in zero-shot semantic segmentation to generate pseudo-labels, we propose a new knowledge-distillation-inspired self-training strategy, namely Annealed Self-Training, which can automatically assign different importance to pseudo-labels to improve performance. We systematically study the proposed Relative Positional Encoding and Annealed Self-Training in a comprehensive experimental evaluation, and our empirical results confirm the effectiveness of our method on three benchmark datasets. Jiaxin Cheng, Soumyaroop Nandi, Premkumar Natarajan, Wael Abd-Almageed |
ICCV | 3 |
| 2021 | BioFors: A Large Biomedical Image Forensics DatasetabstractResearch in media forensics has gained traction to combat the spread of misinformation. However, most of this research has been directed towards content generated on social media. Biomedical image forensics is a related problem, where manipulation or misuse of images reported in biomedical research documents is of serious concern. The problem has failed to gain momentum beyond an academic discussion due to an absence of benchmark datasets and standardized tasks. In this paper we present BioFors1– the first dataset for benchmarking common biomedical image manipulations. BioFors comprises 47,805 images extracted from 1,031 open-source research papers. Images in BioFors are divided into four categories – Microscopy, Blot/Gel, FACS and Macroscopy. We also propose three tasks for forensic analysis – external duplication detection, internal duplication detection and cut/sharp-transition detection. We benchmark BioFors on all tasks with suitable state-of-the-art algorithms. Our results and analysis show that existing algorithms developed on common computer vision datasets are not robust when applied to biomedical images, validating that more research is required to address the unique challenges of biomedical image forensics. Ekraam Sabir, Soumyaroop Nandi, Wael Abd-Almageed, Premkumar Natarajan |
ICCV | 4 |
| 2021 | "Nice Try, Kiddo": Investigating Ad Hominems in Dialogue ResponsesabstractEmily Sheng, Kai-Wei Chang, Prem Natarajan, Nanyun Peng. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Emily Sheng, Kai-Wei Chang 0001, Premkumar Natarajan, Nanyun Peng 0001 |
NAACL-HLT | 3 |
| 2021 | Class-agnostic Object DetectionabstractObject detection models perform well at localizing and classifying objects that they are shown during training. However, due to the difficulty and cost associated with creating and annotating detection datasets, trained models detect a limited number of object types with unknown objects treated as background content. This hinders the adoption of conventional detectors in real-world applications like large-scale object matching, visual grounding, visual relation prediction, obstacle detection (where it is more important to determine the presence and location of objects than to find specific types), etc. We propose class-agnostic object detection as a new problem that focuses on detecting objects irrespective of their object-classes. Specifically, the goal is to predict bounding boxes for all objects in an image but not their object-classes. The predicted boxes can then be consumed by another system to perform application-specific classification, retrieval, etc. We propose training and eval uation protocols for benchmarking class-agnostic detectors to advance future research in this domain. Finally, we propose (1) baseline methods and (2) a new adversarial learning framework for class-agnostic detection that forces the model to exclude class-specific information from features used for predictions. Experimental results show that adversarial learning improves class-agnostic detection efficacy. Ayush Jaiswal, Yue Wu 0001, Pradeep Natarajan, Premkumar Natarajan |
WACV | 4 |
| 2020 | Invariant Representations through Adversarial ForgettingabstractWe propose a novel approach to achieving invariance for deep neural networks in the form of inducing amnesia to unwanted factors of data through a new adversarial forgetting mechanism. We show that the forgetting mechanism serves as an information-bottleneck, which is manipulated by the adversarial training to learn invariance to unwanted factors. Empirical results show that the proposed framework achieves state-of-the-art performance at learning invariance in both nuisance and bias settings on a diverse collection of datasets and tasks. Ayush Jaiswal, Daniel Moyer, Greg Ver Steeg, Wael Abd-Almageed, Premkumar Natarajan |
AAAI | 5 |
| 2020 | MEG: Multi-Evidence GNN for Multimodal Semantic ForensicsabstractFake news often involves semantic manipulations across modalities such as image, text, location etc and requires the development of multimodal semantic forensics for its detection. Recent research has centered the problem around images, calling it image repurposing - where a digitally unmanipulated image is semantically misrepresented by means of its accompanying multimodal metadata such as captions, location, etc. The image and metadata together comprise a multimedia package. The problem setup requires algorithms to perform multimodal semantic forensics to authenticate a query multimedia package using a reference dataset of potentially related packages as evidences. Existing methods are limited to using a single evidence (retrieved package), which ignores potential performance improvement from the use of multiple evidences. In this work, we introduce a novel graph neural network based model for multimodal semantic forensics, which effectively utilizes multiple retrieved packages as evidences and is scalable with the number of evidences. We compare the scalability and performance of our model against existing methods. Experimental results show that the proposed model outperforms existing state-of-the-art algorithms with an error reduction of up to 25 %. Ekraam Sabir, Ayush Jaiswal, Wael Abd-Almageed, Premkumar Natarajan |
ICPR | 4 |
| 2020 | Improving face verification using facial marks and deep CNN: IARPA Janus benchmark-A
Sidra Riaz, Unsang Park, Premkumar Natarajan |
Image Vis. Comput. | 3 |
| 2019 | ManTra-Net: Manipulation Tracing Network for Detection and Localization of Image Forgeries With Anomalous FeaturesabstractTo fight against real-life image forgery, which commonly involves different types and combined manipulations, we propose a unified deep neural architecture called ManTraNet. Unlike many existing solutions, ManTra-Net is an end-to-end network that performs both detection and localization without extra preprocessing and postprocessing. ManTra-Net is a fully convolutional network and handles images of arbitrary sizes and many known forgery types such splicing, copy-move, removal, enhancement, and even unknown types. This paper has three salient contributions. We design a simple yet effective self-supervised learning task to learn robust image manipulation traces from classifying 385 image manipulation types. Further, we formulate the forgery localization problem as a local anomaly detection problem, design a Z-score feature to capture local anomaly, and propose a novel long short-term memory solution to assess local anomalies. Finally, we carefully conduct ablation experiments to systematically optimize the proposed network design. Our extensive experimental results demonstrate the generalizability, robustness and superiority of ManTra-Net, not only in single types of manipulations/forgeries, but also in their complicated combinations. Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan |
CVPR | 3 |
| 2019 | QATM: Quality-Aware Template Matching for Deep LearningabstractFinding a template in a search image is one of the core problems in many computer vision applications, such as template matching, image semantic alignment, image-to-GPS verification \etc. In this paper, we propose a novel quality-aware template matching method, which is not only used as a standalone template matching algorithm, but also a trainable layer that can be easily plugged in any deep neural network. Specifically, we assess the quality of a matching pair as its soft-ranking among all matching pairs, and thus different matching scenarios like 1-to-1, 1-to-many, and many-to-many will be all reflected to different values. Our extensive studies in the classic template matching problem and deep learning tasks demonstrate the effectiveness of QATM: it not only outperforms SOTA template matching methods when used alone, but also largely improves existing DNN solutions when used in DNN. Jiaxin Cheng, Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan |
CVPR | 4 |
| 2019 | AIRD: Adversarial Learning Framework for Image Repurposing DetectionabstractImage repurposing is a commonly used method for spreading misinformation on social media and online forums, which involves publishing untampered images with modified metadata to create rumors and further propaganda. While manual verification is possible, given vast amounts of verified knowledge available on the internet, the increasing prevalence and ease of this form of semantic manipulation call for the development of robust automatic ways of assessing the semantic integrity of multimedia data. In this paper, we present a novel method for image repurposing detection that is based on the real-world adversarial interplay between a bad actor who repurposes images with counterfeit metadata and a watchdog who verifies the semantic consistency between images and their accompanying metadata, where both players have access to a reference dataset of verified content, which they can use to achieve their goals. The proposed method exhibits state-of-the-art performance on location-identity, subject-identity and painting-artist verification, showing its efficacy across a diverse set of scenarios. Ayush Jaiswal, Yue Wu 0001, Wael Abd-Almageed, Iacopo Masi, Premkumar Natarajan |
CVPR | 5 |
| 2019 | The Woman Worked as a Babysitter: On Biases in Language GenerationabstractEmily Sheng, Kai-Wei Chang, Premkumar Natarajan, Nanyun Peng. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Emily Sheng, Kai-Wei Chang 0001, Premkumar Natarajan, Nanyun Peng 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | Layout-aware Subfigure Decomposition for Complex Figures in the Biomedical LiteratureabstractPublished scientific figure is a valuable information resource, but often occur as composite images. The ImageCLEF meeting presented a shared evaluation in 2016 to use machine learning to split these composite figures into components automatically. We adapted an existing high-performance object detection method to analyze the substructure of published biomedical figures by developing a novel multi-branch output convolution neural network to predict irregular panel layouts and provide augmented training data to drive learning. Our system has an accuracy of 86.8% on the 2016 ImageCLEF Medical dataset and 83.1% on a new dataset derived from open access papers from the INTACT database of molecular interactions. Xiangyang Shi, Yue Wu 0001, Huaigu Cao, Gully A. P. C. Burns, Premkumar Natarajan |
ICASSP | 5 |
| 2019 | A Study of Script Language Effects in Deep Neural-Network-Based Scene Text DetectionabstractThis study is different from most of the recent text detection work which focuses on creating a robust text detector system. In this work we studied how script languages affect a text detector's performance by using a multi-language synthetic dataset-namely, the Synthetic Octa-Language (SOL) dataset. The effect of script languages continues to be largely unexplored. Previously, this kind of experiment was infeasible because too many factors influence the performance of a text detector. We really cannot tell what role the factor X plays, neither positive nor negative. To overcome these difficulties, we used controlled synthesized data, which allows us to explicitly control factors such as base image, script language, text content, text color, font face, and font size. With the SOL dataset, we were able to investigate the effect that script languages have on on deep neural-network (DNN)-based methods under different scenarios. Moreover, this dataset can be used in other script-language-related text detection research as well. Jiaxin Cheng, Achin Gupta, Yue Wu 0001, Premkumar Natarajan |
ICDAR | 4 |
| 2019 | NIESR: Nuisance Invariant End-to-End Speech RecognitionabstractDeep neural network models for speech recognition have achieved great success recently, but they can learn incorrect associations between the target and nuisance factors of speech (e.g., speaker identities, background noise, etc.), which can lead to overfitting. While several methods have been proposed to tackle this problem, existing methods incorporate additional information about nuisance factors during training to develop invariant models. However, enumeration of all possible nuisance factors in speech data and the collection of their annotations is difficult and expensive. We present a robust training scheme for end-to-end speech recognition that adopts an unsupervised adversarial invariance induction framework to separate out essential factors for speech-recognition from nuisances without using any supplementary labels besides the transcriptions. Experiments show that the speech recognition model trained with the proposed training scheme achieves relative improvements of 5.48% on WSJ0, 6.16% on CHiME3, and 6.61% on TIMIT dataset over the base model. Additionally, the proposed method achieves a relative improvement of 14.44% on the combined WSJ0+CHiME3 dataset. I-Hung Hsu, Ayush Jaiswal, Premkumar Natarajan |
INTERSPEECH | 3 |
| 2019 | Age-invariant face recognition using gender specific 3D aging modeling
Sidra Riaz, Zahid Ali 0001, Unsang Park, Jongmoo Choi, Iacopo Masi, Premkumar Natarajan |
Multim. Tools Appl. | 6 |
| 2019 | Age progression by gender-specific 3D aging model
Sidra Riaz, Unsang Park, Jongmoo Choi, Premkumar Natarajan |
Mach. Vis. Appl. | 4 |
| 2019 | Learning Pose-Aware Models for Pose-Invariant Face Recognition in the WildabstractWe propose a method designed to push the frontiers of unconstrained face recognition in the wild with an emphasis on extreme out-of-plane pose variations. Existing methods either expect a single model to learn pose invariance by training on massive amounts of data or else normalize images by aligning faces to a single frontal pose. Contrary to these, our method is designed to explicitly tackle pose variations. Our proposed Pose-Aware Models (PAM) process a face image using several pose-specific, deep convolutional neural networks (CNN). 3D rendering is used to synthesize multiple face poses from input images to both train these models and to provide additional robustness to pose variations at test time. Our paper presents an extensive analysis of the IARPA Janus Benchmark A (IJB-A), evaluating the effects that landmark detection accuracy, CNN layer selection, and pose model selection all have on the performance of the recognition pipeline. It further provides comparative evaluations on IJB-A and the PIPA dataset. These tests show that our approach outperforms existing methods, even surprisingly matching the accuracy of methods that were specifically fine-tuned to the target dataset. Parts of this work previously appeared in [1] and [2]. Iacopo Masi, Feng-Ju Chang, Jongmoo Choi, Shai Harel, Jungyeon Kim, KangGeon Kim, Jatuporn Toy Leksut, Stephen Rawls, Yue Wu 0001, Tal Hassner, Wael Abd-Almageed, Gérard G. Medioni, Louis-Philippe Morency, Premkumar Natarajan, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 14 |
| 2018 | A Batch Learning Framework for Scalable Personalized RankingabstractIn designing personalized ranking algorithms, it is desirable to encourage a high precision at the top of the ranked list. Existing methods either seek a smooth convex surrogate for a non-smooth ranking metric or directly modify updating procedures to encourage top accuracy. In this work we point out that these methods do not scale well in a large-scale setting, and this is partly due to the inaccurate pointwise or pairwise rank estimation. We propose a new framework for personalized ranking. It uses batch-based rank estimators and smooth rank-sensitive loss functions. This new batch learning framework leads to more stable and accurate rank approximations compared to previous work. Moreover, it enables explicit use of parallel computation to speed up training. We conduct empirical evaluations on three item recommendation tasks, and our method shows a consistent accuracy improvement over current state-of-the-art methods. Additionally, we observe time efficiency advantages when data scale increases. Premkumar Natarajan |
AAAI | 2 |
| 2018 | Image-to-GPS Verification Through a Bottom-Up Pattern Matching Network
Jiaxin Cheng, Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan |
ACCV (5) | 4 |
| 2018 | Bidirectional Conditional Generative Adversarial Networks
Ayush Jaiswal, Wael Abd-Almageed, Yue Wu 0001, Premkumar Natarajan |
ACCV (3) | 4 |
| 2018 | Weighted Feature Pooling Network in Template-Based Recognition
Zekun Li 0007, Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan |
ACCV (5) | 4 |
| 2018 | BusterNet: Detecting Copy-Move Image Forgery with Source/Target Localization
Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan |
ECCV (6) | 3 |
| 2018 | Deep Multimodal Image-Repurposing DetectionabstractNefarious actors on social media and other platforms often spread rumors and falsehoods through images whose metadata (e.g., captions) have been modified to provide visual substantiation of the rumor/falsehood. This type of modification is referred to as image repurposing, in which often an unmanipulated image is published along with incorrect or manipulated metadata to serve the actor's ulterior motives. We present the Multimodal Entity Image Repurposing (MEIR) dataset, a substantially challenging dataset over that which has been previously available to support research into image repurposing detection. The new dataset includes location, person, and organization manipulations on real-world data sourced from Flickr. We also present a novel, end-to-end, deep multimodal learning model for assessing the integrity of an image by combining information extracted from the image with related information from a knowledge base. The proposed method is compared against state-of-the-art techniques on existing datasets as well as MEIR, where it outperforms existing methods across the board, with AUC improvement up to 0.23. Ekraam Sabir, Wael Abd-Almageed, Yue Wu 0001, Premkumar Natarajan |
ACM Multimedia | 4 |
| 2018 | Unsupervised Adversarial InvarianceabstractData representations that contain all the information about target variables but are invariant to nuisance factors benefit supervised learning algorithms by preventing them from learning associations between these factors and the targets, thus reducing overfitting. We present a novel unsupervised invariance induction framework for neural networks that learns a split representation of data through competitive training between the prediction task and a reconstruction task coupled with disentanglement, without needing any labeled information about nuisance factors or domain knowledge. We describe an adversarial instantiation of this framework and provide analysis of its working. Our unsupervised model outperforms state-of-the-art methods, which are supervised, at inducing invariance to inherent nuisance factors, effectively using synthetic data augmentation to learn invariance, and domain adaptation. Our method can be applied to any prediction task, eg., binary/multi-class classification or regression, without loss of generality. Ayush Jaiswal, Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan |
NeurIPS | 4 |
| 2018 | Image Copy-Move Forgery Detection via an End-to-End Deep Neural NetworkabstractIn this paper, for the first time, we introduce a new end-to-end deep neural network predicting forgery masks to the image copy-move forgery detection problem. Specifically, we use a convolutional neural network to extract block-like features from an image, compute self-correlations between different blocks, use a pointwise feature extractor to locate matching points, and reconstruct a forgery mask through a deconvolutional network. Unlike classic solutions requiring multiple stages of training and parameter tuning, ranging from feature extraction to postprocessing, the proposed solution is fully trainable and can be jointly optimized for the forgery mask reconstruction loss. Our experimental results demonstrate that the proposed method achieves better forgery detection performance than classic approaches relying on different features and matching schemes, and it is more robust against various known attacks like affine transformation, JPEG compression, blurring, etc. Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan |
WACV | 3 |
| 2018 | Facial Landmark Detection with Tweaked Convolutional Neural NetworksabstractThis paper concerns the problem of facial landmark detection. We provide a unique new analysis of the features produced at intermediate layers of a convolutional neural network (CNN) trained to regress facial landmark coordinates. This analysis shows that while being processed by the CNN, face images can be partitioned in an unsupervised manner into subsets containing faces in similar poses (i.e., 3D views) and facial properties (e.g., presence or absence of eye-wear). Based on this finding, we describe a novel CNN architecture, specialized to regress the facial landmark coordinates of faces in specific poses and appearances. To address the shortage of training data, particularly in extreme profile poses, we additionally present data augmentation techniques designed to provide sufficient training examples for each of these specialized sub-networks. The proposed Tweaked CNN (TCNN) architecture is shown to outperform existing landmark detection methods in an extensive battery of tests on the AFW, ALFW, and 300W benchmarks. Finally, to promote reproducibility of our results, we make code and trained models publicly available through our project webpage. Yue Wu 0001, Tal Hassner, KangGeon Kim, Gérard G. Medioni, Premkumar Natarajan |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2017 | EPAT: Euclidean Perturbation Analysis and Transform - An Agnostic Data Adaptation Framework for Improving Facial Landmark DetectorsabstractWe propose EPAT, (Euclidean Perturbation Analysis and Transform) a novel unsupervised adaptation approach for improving the accuracy of any facial landmark detector by characterizing the stability of landmark prediction on test images. In EPAT, a test image is transformed several times using a set of Euclidean transforms, producing several perturbed images. The black box landmark detector is used to find facial landmarks on each perturbed version of the test image. Subsequently, inverse transforms are applied to the corresponding landmarks in order to map them back to the original image. Mean and variance are calculated for all inversely transformed detection. Mean and variance represent the new ensemble prediction and the sensitivity of the underlying landmark detector, respectively. We also introduce affine variance (AV) of facial landmarks. AV is used as a measure of the stability of the predicted landmarks and a criterion for selecting a good data adaptation model which effectively addresses potential mismatches between test and training data of the underlying landmark detector. EPAT is evaluated using four state-of-the-art landmark detectors on the standard 300W dataset and also incorporated into a face recognition pipeline to show improved recognition accuracy on the challenging IJB-A dataset. Yue Wu 0001, Wael Abd-Almageed, Stephen Rawls, Premkumar Natarajan |
FG | 4 |
| 2017 | Self-Organized Text Detection with Minimal Post-processing via Border LearningabstractIn this paper we propose a new solution to the text detection problem via border learning. Specifically, we make four major contributions: 1) We analyze the insufficiencies of the classic non-text and text settings for text detection. 2) We introduce the border class to the text detection problem for the first time, and validate that the decoding process is largely simplified with the help of text border. 3) We collect and release a new text detection PPT dataset containing 10,692 images with non-text, border, and text annotations. 4) We develop a lightweight (only 0.28M parameters), fully convolutional network (FCN) to effectively learn borders in text images. The results of our extensive experiments show that the proposed solution achieves comparable performance, and often outperforms state-of-theart approaches on standard benchmarks-even though our solution only requires minimal post-processing to parse a bounding box from a detected text map, while others often require heavy post-processing. Yue Wu 0001, Premkumar Natarajan |
ICCV | 2 |
| 2017 | 1990 US Census Form Recognition Using CTC Network, WFST Language Model, and Surname CorrectionabstractThis paper presents a system for transcribing 1990 US census forms. Extraction of information from census forms is useful for creating a genealogy database and better archiving census forms. We trained CTC/LSTM-RNN networks as our OCR engine. We solved the major challenge in language modeling by defining syntactical constraints with WFST language models. We made two major technical contributions in this paper. Firstly, 1990 US census forms were automatically transcribed with compelling accuracy for the first time using our system, which can be useful in downstream study in information extracted from census forms. Secondly, we designed a novel post-processing algorithm that improved the recognition accuracy of surnames significantly. Huaigu Cao, Stephen Rawls, Premkumar Natarajan |
ICDAR | 3 |
| 2017 | Using Convolutional Encoder-Decoder for Document Image BinarizationabstractDocument image binarization is one of the critical initial steps for document analysis and understanding. Previous work mostly focused on exploiting hand-crafted features to build statistical models for distinguishing text from background. However, these approaches only achieved limited success because: (a) the effectiveness of hand-crafted features is limited by the researcher's domain knowledge and understanding on the documents, and (b) a universal model cannot always capture the complexity of different document degradations. In order to address these challenges, we propose a convolutional encoder-decoder model with deep learning for document image binarization in this paper. In the proposed method, mid-level document image representations are learnt by a stack of convolutional layers, which compose the encoder in this architecture. Then the binarization image is obtained by mapping low resolution representations to the original size through the decoder, which is composed by a series of transposed convolutional layers. We compare the proposed binarization method with other binarization algorithms both qualitatively and quantitatively on the public dataset. The experimental results show that the proposed method has comparable performance to the other hand-crafted binarization approaches and has more generalization capabilities with limited in-domain training data. Xujun Peng, Huaigu Cao, Premkumar Natarajan |
ICDAR | 3 |
| 2017 | Combining Convolutional Neural Networks and LSTMs for Segmentation-Free OCRabstractWe present a novel end-to-end trainable OCR system combining a CNN for feature extraction with 1-D LSTMs for sequence modeling. We present results on English and Arabic handwriting data, and on English machine print data, showing state-of-the-art performance. We believe that our method is simpler than existing 2D LSTM models, and will make it easier to use techniques borrowed from CNN research in computer vision to improve OCR performance. Stephen Rawls, Huaigu Cao, Senthil Kumar, Premkumar Natarajan |
ICDAR | 4 |
| 2017 | Multimedia Semantic Integrity Assessment Using Joint Embedding Of Images And TextabstractReal-world multimedia data is often composed of multiple modalities such as an image or a video with associated text (e.g., captions, user comments, etc.) and metadata. Such multimodal data packages are prone to manipulations, where a subset of these modalities can be altered to misrepresent or repurpose data packages, with possible malicious intent. It is therefore important to develop methods to assess or verify the integrity of these multimedia packages. Using computer vision and natural language processing methods to directly compare the image (or video) and the associated caption to verify the integrity of a media package is only possible for a limited set of objects and scenes. In this paper we present a novel deep-learning-based approach that uses a reference set of multimedia packages to assess the semantic integrity of multimedia packages containing images and captions. We construct a joint embedding of images and captions with deep multimodal representation learning on the reference dataset in a framework that also provides image-caption consistency scores (ICCSs). The integrity of query media packages is assessed as the inlierness of the query ICCSs with respect to the reference dataset. We present the MultimodAl Information Manipulation dataset (MAIM), a new dataset of media packages from Flickr, which we are making available to the research community. We use both the newly created dataset as well as Flickr30K and MS COCO datasets to quantitatively evaluate our proposed approach. The reference dataset does not contain unmanipulated versions of tampered query packages. Our method is able to achieve F-1 scores of 0.75, 0.89 and 0.94 on MAIM, Flickr30K and MS COCO, respectively, for detecting semantically incoherent media packages. Ayush Jaiswal, Ekraam Sabir, Wael Abd-Almageed, Premkumar Natarajan |
ACM Multimedia | 4 |
| 2017 | Deep Matching and Validation Network: An End-to-End Solution to Constrained Image Splicing Localization and DetectionabstractImage splicing is a very common image manipulation technique that is sometimes used for malicious purposes. A splicing detection and localization algorithm usually takes an input image and produces a binary decision indicating whether the input image has been manipulated, and also a segmentation mask that corresponds to the spliced region. Most existing splicing detection and localization pipelines suffer from two main shortcomings: 1) they use handcrafted features that are not robust against subsequent processing (e.g., compression), and 2) each stage of the pipeline is usually optimized independently. In this paper we extend the formulation of the underlying splicing problem to consider two input images, a query image and a potential donor image. Here the task is to estimate the probability that the donor image has been used to splice the query image, and obtain the splicing masks for both the query and donor images. We introduce a novel deep convolutional neural network architecture, called Deep Matching and Validation Network (DMVN), which simultaneously localizes and detects image splicing. The proposed approach does not depend on handcrafted features and uses raw input images to create deep learned representations. Furthermore, the DMVN is end-to-end optimized to produce the probability estimates and the segmentation masks. Our extensive experiments demonstrate that this approach outperforms state-of-the-art splicing detection methods by a large margin in terms of both AUC score and speed. Yue Wu 0001, Wael Abd-Almageed, Premkumar Natarajan |
ACM Multimedia | 3 |
| 2016 | Modeling Concept Dependencies in a Scientific CorpusabstractOur goal is to generate reading lists for students that help them optimally learn technical material.Existing retrieval algorithms return items directly relevant to a query but do not return results to help users read about the concepts supporting their query.This is because the dependency structure of concepts that must be understood before reading material pertaining to a given query is never considered.Here we formulate an information-theoretic view of concept dependency and present methods to construct a "concept graph" automatically from a text corpus.We perform the first human evaluation of concept dependency edges (to be published as open data), and the results verify the feasibility of automatic approaches for inferring concepts and their dependency relations.This result can support search capabilities that may be tuned to help users learn a subject rather than retrieve documents based on a single query. Jonathan Gordon 0001, Linhong Zhu, Aram Galstyan, Premkumar Natarajan, Gully A. P. C. Burns |
ACL (1) | 4 |
| 2016 | Pose-Aware Face Recognition in the WildabstractWe propose a method to push the frontiers of unconstrained face recognition in the wild, focusing on the problem of extreme pose variations. As opposed to current techniques which either expect a single model to learn pose invariance through massive amounts of training data, or which normalize images to a single frontal pose, our method explicitly tackles pose variation by using multiple pose specific models and rendered face images. We leverage deep Convolutional Neural Networks (CNNs) to learn discriminative representations we call Pose-Aware Models (PAMs) using 500K images from the CASIA WebFace dataset. We present a comparative evaluation on the new IARPA Janus Benchmark A (IJB-A) and PIPA datasets. On these datasets PAMs achieve remarkably better performance than commercial products and surprisingly also outperform methods that are specifically fine-tuned on the target dataset. Iacopo Masi, Stephen Rawls, Gérard G. Medioni, Premkumar Natarajan |
CVPR | 4 |
| 2016 | Document Image Quality Assessment Using Discriminative Sparse RepresentationabstractThe goal of document image quality assessment (DIQA) is to build a computational model which can predict the degree of degradation for document images. Based on the estimated quality scores, the immediate feedback can be provided by document processing and analysis systems, which helps to maintain, organize, recognize and retrieve the information from document images. Recently, the bag-of-visual-words (BoV) based approaches have gained increasing attention from researchers to fulfill the task of quality assessment, but how to use BoV to represent images more accurately is still a challenging problem. In this paper, we propose to utilize a sparse representation based method to estimate document image's quality with respect to the OCR capability. Unlike the conventional sparse representation approaches, we introduce the target quality scores into the training phase of sparse representation. The proposed method improves the discriminability of the system and ensures the obtained codebook is more suitable for our assessment task. The experimental results on a public dataset show that the proposed method outperforms other hand-crafted and BoV based DIQA approaches. Xujun Peng, Huaigu Cao, Premkumar Natarajan |
DAS | 3 |
| 2016 | Learning document image binarization from dataabstractWe present a fully trainable solution for binarization of degraded document images using extremely randomized trees. Unlike previous attempts that often use simple features, our method encodes all heuristics about whether or not a pixel is foreground text into a high-dimensional feature vector and learns a more complicated decision function. We introduce two novel features, the Logarithm Intensity Percentile (LIP) and the Relative Darkness Index (RDI), and combine them with low level features, and reformulated features from existing binarization methods. Experimental results show that using small sample size (about 1.5% of all available training data), we can achieve a binarization performance comparable to manually-tuned, state-of-the-art methods. Additionally, the trained document binarization classifier shows good generalization capabilities on out-of-domain data. Yue Wu 0001, Premkumar Natarajan, Stephen Rawls, Wael Abd-Almageed |
ICIP | 2 |
| 2016 | Computationally efficient template-based face recognitionabstractClassically, face recognition depends on computing the similarity (or distance) between a pair of face images and/or their respective representations, where each subject is represented by one image. Template-based face recognition was introduced by the release of IARPA's Janus Benchmark-A (IJB-A) dataset, in which each enrolled subject is represented by a group of one or more images, called a template. The group of images comprising a template might have been acquired using different head poses, illuminations, ages and facial expressions. Template images could come from still images or video frames. Therefore, measuring the similarity between templates representing two subjects significantly increases the number of pairwise image comparisons (i.e., O(NM), where N and M are the number of image templates being compared). As the number of enrolled subjects, K, increases, both computational and space requirements become computationally prohibitive. To address this challenge, we present a novel approximate nearest-neighbor (ANN) search-based solution. Given a query template, ANN methods are used to find similar face images. Retrieved images are used to construct a template pool that is used to find the correct identity of the query subject. The proposed approach largely reduces the number of imposter template-pair comparisons. Experimental results on the IJB-A dataset show that the proposed approach achieves significant speed-up and storage savings, without sacrificing accuracy. Yue Wu 0001, Wael Abd-Almageed, Stephen Rawls, Premkumar Natarajan |
ICPR | 4 |
| 2016 | Face recognition using deep multi-pose representationsabstractWe introduce our method and system for face recognition using multiple pose-aware deep learning models. In our representation, a face image is processed by several pose-specific deep convolutional neural network (CNN) models to generate multiple pose-specific features. 3D rendering is used to generate multiple face poses from the input image. Sensitivity of the recognition system to pose variations is reduced since we use an ensemble of pose-specific CNN features. The paper presents extensive experimental results on the effect of landmark detection, CNN layer selection and pose model selection on the performance of the recognition pipeline. Our novel representation achieves better results than the state-of-the-art on IARPA's CS2 and NIST's IJB-A in both verification and identification (i.e. search) tasks. Wael Abd-Almageed, Yue Wu 0001, Stephen Rawls, Shai Harel, Tal Hassner, Iacopo Masi, Jongmoo Choi, Jatuporn Toy Leksut, Jungyeon Kim, Premkumar Natarajan, Ramakant Nevatia, Gérard G. Medioni |
WACV | 10 |
| 2015 | Document image OCR accuracy prediction via latent Dirichlet allocationabstractOptical character recognition (OCR) accuracy of document images is an important factor for the success of many document processing and analysis tasks, especially for unconstraint captured document images. Although several document image OCR capability assessment methods are proposed, they mostly model the problem based on the empirically defined rules of image degradation, which cause the existing approaches infeasible for predicting the OCR scores. In this paper, a computational model is presented to automatically predict document image quality towards facilitating the OCR accuracy without references. Unlike conventional methods that use heuristically designed features, in our work the raw features are learned from training images and a generative quality model is built based on latent Dirichlet allocation, which is used to assess the document's OCR capability. We present evaluation results on a public dataset which have been captured using digital cameras with different level of blur degradation. The experimental results show that the proposed method outperforms traditional document image quality assessment approaches. Xujun Peng, Huaigu Cao, Premkumar Natarajan |
ICDAR | 3 |
| 2015 | Building and Using a Knowledge Graph to Combat Human Trafficking
Pedro A. Szekely, Craig A. Knoblock, Jason Slepicka, Andrew Philpot, Chengye Yin, Dipsy Kapoor, Premkumar Natarajan, Daniel Marcu, Kevin Knight, David Stallard, Subessware S. Karunamoorthy, Rajagopal Bojanapalli, Steven Minton, Brian Amanatullah, Todd Hughes, Mike Tamayo, David Flynt, Rachel Artiss, Shih-Fu Chang, Tao Chen 0015, Gerald Hiebel, Lidia Silva Ferreira |
ISWC (2) | 8 |
| 2015 | Integrating natural language processing with image document analysis: what we learned from two real-world applications
Jinying Chen, Huaigu Cao, Premkumar Natarajan |
Int. J. Document Anal. Recognit. | 3 |
| 2014 | Text Classification via iVector Based Feature RepresentationabstractIn this paper, we address the problem of text classification: classifying modern machine-printed text, handwritten text and historical typewritten text from degraded noisy documents. We propose a novel text classification approach based on iVector, a newly developed concept in speaker verification. To a given text line, the iVector is a fixed-length feature vector representation, transformed from a high-dimensional super vector based on means of Gaussian mixture model (GMM), where the text dependent component is separated from a universal background model (UBM) and can be represented by a low dimensional set of factors. We classify the text lines with a discriminative classifier - support vector machine (SVM) in iVector space. A baseline approach of text classification using GMM in feature space is also presented for evaluation purpose. Experimental results on an Arabic document database show accuracy of 92.04% for text line classification using the proposed method. Furthermore, the relative word error rate (WER) of 9.6% is decreased in optical character recognition (OCR) when coupled with the proposed iVector-SVM classifier. The proposed iVector-SVM approach is language independent, thus, can be applied to other scripts as well. Shengxin Zha, Xujun Peng, Huaigu Cao, Xiaodan Zhuang, Pradeep Natarajan, Premkumar Natarajan |
Document Analysis Systems | 6 |
| 2014 | Progress in the Raytheon BBN Arabic Offline Handwriting Recognition SystemabstractThis paper presents the most recent progress and state of the art result obtained from BBN's Arabic offline handwriting recognition research. Our system is based a left-to-right hidden Markov model and integrates discriminative learning methods including discriminative MPE and n-best rescoring using the scores of glyph classifiers (SVM, DNN) and the RNNLM. Arabic-related features for n-best rescoring are also investigated in this paper. Multi-stage MAP/MLLR and writer verification are applied to adapt the recognizer in all training situations. Consensus network is extensively researched for system combination and improving challenging preprocessing problems. Huaigu Cao, Premkumar Natarajan, Xujun Peng, Krishna Subramanian 0001, David Belanger 0001 |
ICFHR | 2 |
| 2014 | Applications of Recurrent Neural Network Language Model in Offline Handwriting Recognition and Word SpottingabstractThe recurrent neural network language model (RNNLM) is a discriminative, non-Markovian model that can capture long-span word history in natural language. It has been proved to be successful in automatic speech recognition and machine translation. In this work, we applied RNNLM to the n-best rescoring stage of the state-of-the-art BBN Byblos OCR (optical character recognition) system for handwriting recognition.1 With RNNLM scores as additional features, our system achieved significant improvement (p <; 0.001), a 3.5% relative reduction on OCR word error rate, compared with a high baseline that uses n-gram language model for rescoring. We have also developed a novel method to integrate the OCR n-best RNNLM scores into the word posterior probabilities in OCR confusion networks, which resulted in consistent observable improvements in word spotting for OCR'ed handwritten documents, as measured by both mean average precision (MAP) and detection-error tradeoff (DET) curves. Jinying Chen, Huaigu Cao, Premkumar Natarajan |
ICFHR | 5 |
| 2014 | Confusion Network Based Recurrent Neural Network Language Modeling for Chinese OCR Error DetectionabstractThis paper presents a new framework for OCR error detection, which uses a conditional random field model to combine rich features from multiple sources, including confusion networks (c-nets), lexical local context and recurrent neural network language model (RNNLM)1. We propose a novel, efficient method for computing character-level c-net based RNNLM scores by using dynamic programming and c-net partial unfolding. Our experiments show that our error detection model has consistent observable improvements over a high baseline employed by our current OCR demo system, as measured by average precision and detection error trade-off curve on two test sets of Chinese image documents. Both linguistic and recognition features contribute to the high performance, with the former especially informative. In addition, we show that the new feature we proposed, the c-net RNNLM feature, plays a remarkable beneficial role in improving error detection rate. These results suggest that applications on top of image text recognition can benefit substantially from a hybrid strategy that combines techniques from optical character recognition and natural language processing. Jinying Chen, Yue Wu 0001, Huaigu Cao, Premkumar Natarajan |
ICPR | 4 |
| 2014 | Fast blockwise SURE shrinkage for image denoising
Yue Wu 0001, Brian Tracey, Premkumar Natarajan, Joseph P. Noonan |
Signal Process. | 3 |
| 2013 | Leveraging Arabic-English Bilingual Corpora with Crowd Sourcing-Based Annotation for Arabic-Hebrew SMT
Manish Gaurav, Guruprasad Saikumar, Amit Srivastava, Premkumar Natarajan, Shankar Ananthakrishnan, Spyridon Matsoukas |
CICLing (2) | 4 |
| 2013 | ASR error detection in a conversational spoken language translation systemabstractDetection of automatic speech recognition (ASR) errors is crucial to preventing their further propagation through statistical machine translation (SMT) in conversational spoken language translation (CSLT) systems. In this paper, we venture beyond traditional features obtained from the ASR decoder and hypothesized word sequence, and explore additional information streams provided by an error-robust CSLT system, including SMT confidence estimates and posteriors from named entity detection (NED). Another significant novelty of this work is the use of an automated word boundary detector based on acoustic-prosodic features to verify the existence of ASR-hypothesized word boundaries, which further improves ASR error detection. Offline evaluation on a test set designed to invoke ASR errors showed that at 10% false alarm rate, the proposed features provide 2.8% absolute (4.2% relative) improvement in detection rate over a state-of-the-art baseline error detector that uses a rich set of features traditionally employed in the existing literature. Sankaranarayanan Ananthakrishnan, Rohit Kumar 0001, Rohit Prasad, Premkumar Natarajan |
ICASSP | 5 |
| 2013 | Graph based multimodal word clustering for video event detectionabstractCombining diverse low-level features from multiple modalities has consistently improved performance over a range of video processing tasks, including event detection. In our work, we study graph based clustering techniques for integrating information from multiple modalities by identifying word clusters spread across the different modalities. We present different methods to identify word clusters including word similarity graph partitioning, word-video co-clustering and Latent Semantic Indexing and the impact of different metrics to quantify the co-occurrence of words. We present experimental results on a ≈45000 video dataset used in the TRECVID MED 11 evaluations. Our experiments show that multimodal features have consistent performance gains over the use of individual features. Further, word similarity graph construction using a complete graph representation consistently improves over partite graphs and early fusion based multimodal systems. Finally, we see additional performance gains by fusing multimodal features with individual features. Aravind Namandi Vembu, Pradeep Natarajan, Shuang Wu 0003, Rohit Prasad, Premkumar Natarajan |
ICASSP | 5 |
| 2013 | Detecting OOV Names in Arabic Handwritten DataabstractThis paper presents a novel approach to detect Arabic OOV names from OCR'ed handwritten documents. In our approach, OOV names are searched for using approximate string match on character consensus networks (cnets). The retrieved regions are re-ranked using novel features representing the quality of the match and the likelihood of the detected region to be an OOV name. Our features that encode word boundary information into the approximate match algorithm significantly improve mean average precision (MAP) by 12.2% (absolute gains) for rank cut-off 100 (48.2% vs. 36.0%) and 11.9% for cut-off 1000 (47.0% vs. 35.1%) over the baseline system. Discriminative reranking based on maximum entropy classification using novel features, such as the probability of a retrieved region being an OOV name (called OOV name probability) from a conditional random field model, further improve MAP by 2.3% (absolute gains) for cut-off 100 and 3.0% for cut-off 1000. The improvements are consistent in DET (Detection Error Tradeoff) curves. Our results show that character cnet based OOV name search benefits clearly from the approximate match using word boundary information and the reranking algorithm. Our experiments also show that OOV name probability is very useful for reranking. Jinying Chen, Rohit Prasad, Huaigu Cao, Premkumar Natarajan |
ICDAR | 4 |
| 2013 | Exploiting Stroke Orientation for CRF Based Binarization of Historical DocumentsabstractWe present a novel binarization method that is especially effective on historical documents with the following characteristics: (a) the documents contain free-form cursive handwritten text with significant but consistent slant, (b) scanning artifacts resulting in the text and background pixels not having uniform intensity even within the same page, and (c) pages containing significant amount of bleeds from the other side of the page. In order to tackle the problem of non-uniform text and background intensity, we use a thresholding algorithm that works equally well for regions of the page containing text and regions of the page containing no text. We then combine this algorithm with a CRF-based framework which handles bleeds using a novel approach to further improve the quality of binarization. We compare the proposed binarization algorithm against other popular binarization algorithms both qualitatively using examples and quantitatively using the word error rate (WER) metric from performing optical character recognition (OCR) on binarized text using the BBN Byblos Offline Handwritten text recognition (OHR) system. Xujun Peng, Huaigu Cao, Krishna Subramanian 0001, Rohit Prasad, Premkumar Natarajan |
ICDAR | 5 |
| 2013 | Variable-Span out-of-vocabulary named entity detection
Sankaranarayanan Ananthakrishnan, Rohit Prasad, Premkumar Natarajan |
INTERSPEECH | 4 |
| 2013 | Audio self organized units for high-level event detection
Xiaodan Zhuang, Shuang Wu 0003, Pradeep Natarajan, Rohit Prasad, Premkumar Natarajan |
INTERSPEECH | 5 |
| 2013 | Semi-Supervised Word Sense Disambiguation for Mixed-Initiative Conversational Spoken Language Translation
Sankaranarayanan Ananthakrishnan, Sanjika Hewavitharana, Rohit Kumar 0001, Enoch Kan, Rohit Prasad, Premkumar Natarajan |
MTSummit | 6 |
| 2013 | Ridge Regression based classifiers for large scale class imbalanced datasetsabstractLarge scale, class imbalanced data classification is a challenging task that occurs frequently in several computer vision tasks such as web video retrieval. A number of algorithms have been proposed in literature that approach this problem from different perspectives (e.g. Sampling, Cost-sensitive learning, Active learning). The challenge is two fold in this task - first the data imbalance causes many classification algorithms to learn trivial classifiers that declare all test examples to be from the majority class. Second, many algorithms do not scale to large dataset sizes. We address these two issues by using two different cost-sensitive versions of Ridge Regression as our binary classifiers. We demonstrate our approach for retrieving unstructured web videos from 10 events on the benchmark TRECVID MED 12 dataset containing ≈47000 videos. We empirically show that they perform at par with state-of-the-art support vector machine based classifiers using χ2kernels while being 30 to 60 times faster. Devansh Arpit, Shuang Wu 0003, Pradeep Natarajan, Rohit Prasad, Premkumar Natarajan |
WACV | 5 |
| 2013 | Scene image categorization and video event detection using Naive Bayes Nearest NeighborabstractWe present a detailed study of Naive Bayes Nearest Neighbor (NBNN) proposed by Boiman et al., with application to scene categorization and video event detection. Our study indicates that using Dense-SIFT along with dimensionality reduction using PCA enables NBNN to obtain state-of-the-art results. We demonstrate this on two tasks: (1) scene image categorization on the UIUC 8 Sports Events Image Dataset (obtaining 84.67%) and the MIT 67 Indoor Scene Image Dataset (obtaining 48.84%); and (2) detecting videos depicting certain events of interest on the challenging MED'11 video dataset with only 15 positive training videos per event. We present an extension referred to as sparse-NBNN that constrains the number of training images that can used to match with a given test image for the image-to-class distance computation. Experiments indicate that this improves upon NBNN for handling of imbalanced training data. Shiv Vitaladevuni, Pradeep Natarajan, Shuang Wu 0003, Xiaodan Zhuang, Rohit Prasad, Premkumar Natarajan |
WACV | 6 |
| 2013 | Batch-mode semi-supervised active learning for statistical machine translation
Sankaranarayanan Ananthakrishnan, Rohit Prasad, David Stallard, Premkumar Natarajan |
Comput. Speech Lang. | 4 |
| 2013 | BBN TransTalk: Robust multilingual two-way speech-to-speech translation for mobile platforms
Rohit Prasad, Premkumar Natarajan, David Stallard, Shirin Saleem, Shankar Ananthakrishnan, Stavros Tsakalidis, Chia-Lin Kao, Fred Choi, Ralf Meermeier, Mark Rawls, Jacob Devlin, Kriste Krstovski, Aaron Challenner |
Comput. Speech Lang. | 2 |
| 2013 | Local Shannon entropy measure with statistical tests for image randomness
Yue Wu 0001, Yicong Zhou, George Saveriades, Sos S. Agaian, Joseph P. Noonan, Premkumar Natarajan |
Inf. Sci. | 6 |
| 2013 | James-Stein Type Center Pixel Weights for Non-Local Means Image DenoisingabstractNon-Local Means (NLM) and its variants have proven to be effective and robust in many image denoising tasks. In this letter, we study approaches to selecting center pixel weights (CPW) in NLM. Our key contributions are 1) we give a novel formulation of the CPW problem from a statistical shrinkage perspective; 2) we construct the James–Stein shrinkage estimator in the CPW context; and 3) we propose a new local James–Stein type CPW (LJSCPW) that is locally tuned for each image pixel. Our experimental results showed that compared to existing CPW solutions, the LJSCPW is more robust and effective under various noise levels. In particular, the NLM with the LJSCPW attains higher means with smaller variances in terms of the peak signal and noise ratio (PSNR) and structural similarity (SSIM), implying it improves the NLM denoising performance and makes the denoising less sensitive to parameter changes. Yue Wu 0001, Brian Tracey, Premkumar Natarajan, Joseph P. Noonan |
IEEE Signal Process. Lett. | 3 |
| 2013 | Probabilistic Non-Local MeansabstractIn this letter, we propose a so-called probabilistic non-local means (PNLM) method for image denoising. Our main contributions are: 1) we point out defects of the weight function used in the classic NLM; 2) we successfully derive all theoretical statistics of patch-wise differences for Gaussian noise; and 3) we employ this prior information and formulate the probabilistic weights truly reflecting the similarity between two noisy patches. Our simulation results indicate the PNLM outperforms the classic NLM and many NLM recent variants in terms of the peak signal noise ratio (PSNR) and the structural similarity (SSIM) index. Encouraging improvements are also found when we replace the NLM weights with the PNLM weights in tested NLM variants. Yue Wu 0001, Brian Tracey, Premkumar Natarajan, Joseph P. Noonan |
IEEE Signal Process. Lett. | 3 |
| 2012 | Multimodal feature fusion for robust event detection in web videosabstractCombining multiple low-level visual features is a proven and effective strategy for a range of computer vision tasks. However, limited attention has been paid to combining such features with information from other modalities, such as audio and videotext, for large scale analysis of web videos. In our work, we rigorously analyze and combine a large set of low-level features that capture appearance, color, motion, audio and audio-visual co-occurrence patterns in videos. We also evaluate the utility of high-level (i.e., semantic) visual information obtained from detecting scene, object, and action concepts. Further, we exploit multimodal information by analyzing available spoken and videotext content using state-of-the-art automatic speech recognition (ASR) and videotext recognition systems. We combine these diverse features using a two-step strategy employing multiple kernel learning (MKL) and late score level fusion methods. Based on the TRECVID MED 2011 evaluations for detecting 10 events in a large benchmark set of ~45000 videos, our system showed the best performance among the 19 international teams. Pradeep Natarajan, Shuang Wu 0003, Shiv Vitaladevuni, Xiaodan Zhuang, Stavros Tsakalidis, Unsang Park, Rohit Prasad, Premkumar Natarajan |
CVPR | 8 |
| 2012 | Local Segmentation of Touching Characters Using Contour Based Shape DecompositionabstractWe propose a contour based shape decomposition approach that provides local segmentation of touching characters. The shape contour is linearized into edge lets and edge lets are merged into boundary fragments. The connection cost between boundary fragments is obtained by considering local smoothness, connection length and a stroke-level property called the Same Stroke Rate. Samples of connections among boundary fragments are randomly generated and the one with the minimum global cost is selected to produce the final segmentation of the shape. To obtain a bipartite segmentation using this approach, we perform an iterative search for the parameters that finally yields two components on a shape. Experimental results on synthetic shape images and the LTP dataset show that this contour based shape decomposition technique is promising and it is effective for providing local segmentation of touching characters. David S. Doermann, Huaigu Cao, Rohit Prasad, Premkumar Natarajan |
Document Analysis Systems | 5 |
| 2012 | Automatic Tune Set Generation for Machine Translation with Limited Indomain Data
Jinying Chen, Jacob Devlin, Huaigu Cao, Rohit Prasad, Premkumar Natarajan |
EAMT | 5 |
| 2012 | Multi-channel Shape-Flow Kernel Descriptors for Robust Video Event Detection and Retrieval
Pradeep Natarajan, Shuang Wu 0003, Shiv Vitaladevuni, Xiaodan Zhuang, Unsang Park, Rohit Prasad, Premkumar Natarajan |
ECCV (2) | 7 |
| 2012 | Automatic pronunciation prediction for text-to-speech synthesis of dialectal arabic in a speech-to-speech translation systemabstractText-to-speech synthesis (TTS) is the final stage in the speech-tospeech (S2S) translation pipeline, producing an audible rendition of translated text in the target language. TTS systems typically rely on a lexicon to look up pronunciations for each word in the input text. This is problematic when the target language is dialectal Arabic, because the statistical machine translation (SMT) system usually produces undiacritized text output. Many words in the latter possess multiple pronunciations; the correct choice must be inferred from context. In this paper, we present a weakly supervised pronunciation prediction approach for undiacritized dialectal Arabic in S2S systems that leverages automatic speech recognition (ASR) to obtain parallel training data for pronunciation prediction. Additionally, we show that incorporating source language features derived from SMT-generated automatic word alignment further improves automatic pronunciation prediction accuracy. Sankaranarayanan Ananthakrishnan, Stavros Tsakalidis, Rohit Prasad, Premkumar Natarajan, Aravind Namandi Vembu |
ICASSP | 4 |
| 2012 | Applying Discriminatively Optimized Feature Transform for HMM-based Off-Line Handwriting RecognitionabstractFeature extraction is an important step in off-line handwriting recognition systems to represent raw handwriting in a low-dimensional, tractable feature space. Traditionally, linear feature transforms such as Principle Component Analysis (PCA), Linear Discriminative Analysis (LDA) are commonly used. The assumptions they make, however, usually cannot be satisfied in practice and thus the best performance is not obtained. In this paper, we apply the Region-Dependent non-linear feature Transform (RDT) to handwriting recognition. RDT is one type of non-linear feature transforms which captures the discriminating power much better than traditional linear ones. We justify the effectiveness of RDT on handwriting features using an HMM-based handwriting recognition system on an Arabic handwriting dataset, which consists of 38K pages of handwriting, over 3M handwritten words. Experimental results show that RDT is able to decrease the word error rates (WERs) relatively by 4% to 7% with statistical significance, comparing to two LDA-based baseline systems. Huaigu Cao, Rohit Prasad, Premkumar Natarajan |
ICFHR | 5 |
| 2012 | Statistical Machine Translation as a Language Model for Handwriting RecognitionabstractWhen performing handwriting recognition on natural language text, the use of a word-level language model (LM) is known to significantly improve recognition accuracy. The most common type of language model, the n-gram model, decomposes sentences into short, overlapping chunks. In this paper, we propose a new type of language model which we use in addition to the standard n-gram LM. Our new model uses the likelihood score from a statistical machine translation system as a reranking feature. In general terms, we automatically translate each OCR hypothesis into another language, and then create a feature score based on how "difficult" it was to perform the translation. Intuitively, the difficulty of translation correlates with how well-formed the input sentence is. In an Arabic handwriting recognition task, we were able to obtain an 0.4% absolute improvement to word error rate (WER) on top of a powerful 5-gram LM. Jacob Devlin, Matin Kamali, Krishna Subramanian 0001, Rohit Prasad, Premkumar Natarajan |
ICFHR | 5 |
| 2012 | Document recognition and translation system for unconstrained Arabic documents
Huaigu Cao, Jinying Chen, Jacob Devlin, Rohit Prasad, Premkumar Natarajan |
ICPR | 5 |
| 2012 | Extracting information from handwritten content in census forms
Huaigu Cao, Krishna Subramanian 0001, Xujun Peng, Jinying Chen, Rohit Prasad, Premkumar Natarajan |
ICPR | 6 |
| 2012 | Detecting near-duplicate document images using interest point matching
Shiv Vitaladevuni, Fred Choi, Rohit Prasad, Premkumar Natarajan |
ICPR | 4 |
| 2012 | Detecting OOV Named-Entities in Conversational Speech
Rohit Kumar 0001, Rohit Prasad, Sankaranarayanan Ananthakrishnan, Aravind Namandi Vembu, David Stallard, Stavros Tsakalidis, Premkumar Natarajan |
INTERSPEECH | 7 |
| 2012 | Robust Event Detection From Spoken Content In Consumer Domain Videos
Stavros Tsakalidis, Xiaodan Zhuang, Roger Hsiao, Shuang Wu 0003, Pradeep Natarajan, Rohit Prasad, Premkumar Natarajan |
INTERSPEECH | 7 |
| 2012 | Compact Audio Representation for Event Detection in Consumer Media
Xiaodan Zhuang, Stavros Tsakalidis, Shuang Wu 0003, Pradeep Natarajan, Rohit Prasad, Premkumar Natarajan |
INTERSPEECH | 6 |
| 2011 | Wavelet Band-pass Filters for Matching Multiple Templates in Real-timeabstractMany applications in image processing and computer vision require finding a particular template in an image or a video, that is, template matching. Given a template and an input, the matching algorithm finds the region of interest (ROI) that most closely matches the template in terms of some similarity measurement. According to the way similarity measurements are performed, the template matching methods can roughly be classified into two groups: 1) patch matching schemes, such as the sum of absolute difference (SAD) , the sum of squared difference (SSD) [1], or cross correlation (XCORR), where the similarity measurement directly relies on pixel information from the patch of interest; and 2) feature matching schemes, such as invariant features [4] and bags of features [5], where similarity measurement relies on features describing the template and the frame. Patch matching methods are not robust, especially when noise, skew, or errors occur. Further, they consume a large amount of time, because of expensive sliding window search for calculating the similarity score over all possible locations. Several techniques have been explored for accelerating such matching methods, including early rejections and correlation techniques [1]. However, the computation cost could still be unaffordable when the frame size is large. Typically other techniques, like frame difference, are used to reduce the search space in applications. Feature matching methods process the template and describe it with features, which are ideally invariant to rotation, skew, noise etc. However in many cases, the use of a more complicated model for similarity measurement results in higher computational cost. Further, sliding window search is also a costly stage for such methods. While there exist known algorithms for fast search of object instances in an image using branch-and-bound techniques, in our particular problem, methods of this type have two crucial limitations. First, they require a large number of training samples for each class to learn robust classifiers. Second, interest point detectors like SIFT [5] typically do not generate sufficient number of feature points, because of the small size of the provided logo, large homogenous regions and degradations. Wavelets based approaches have been extensively used in object detection and recognition. In [3], wavelet coefficients based image histogram are collected in bins and are used for classifying logos. In [6], wavelet coefficients are directly used and trained for pedestrian detection. In [8], wavelet coefficients are selected to form rotation-invariant features by using the angular-radial transform. However, matching logos within frames using [3, 6, 8] still requires expensive window searching and thus are not appropriate for real-time processing. In this paper, we propose a new matching method using the wavelet based band-pass filters (WBPFs). Instead of using direct distance measurement requiring expensive window search, the similarity is measured in the indirect way involving two stages. In the stage of offline template processing (see Figure 1), a template is automatically described by a set of three directional WBPFs, where only salient wavelet frequency components of the template are allowed to pass. In the stage of online frame processing (see Figure 2), a frame is transformed to the wavelet domain and its sub-bands are filtered with respect to the corresponding template WBFPs. Finally, the detection is made with respect to the region of the densest responses under spatial constraints [2, 4]. We show that the proposed template matching system has a very low computational cost, which is 50 times faster than the correlation based SSD [1] and 10 times faster than the orthogonal Haar transform (OHT) based SSD [7]. Further, the proposed method does not trade-off accuracy, since the use of subtemplate information makes it robust to skew and camera view change. Experimental results demonstrate our method for real-time logo detection in broadcast videos. Yue Wu 0001, Pradeep Natarajan, Joseph P. Noonan, Rohit Prasad, Premkumar Natarajan |
BMVC | 5 |
| 2011 | Open-set speaker identification in broadcast newsabstractIn this paper, we examine the problem of text-independent open-set speaker identification (OS-SI) in broadcast news. Particularly, the impact of the population of registered speakers to OS-SI performance is investigated, which is the central issue for designing practical OS-SI system. We amend the maximum mutual information (MMI)-based discriminative training scheme to facilitate its incorporation in OS-SI systems. We also improve the implementation to allow the application of MMI based approach with 2048-component Gaussian mixture models. All systems are evaluated using NIST RT-03, RT-04 and FBIS corpora, with a maximum of 82 registered speakers. Our study shows that notable performance improvement can be obtained with MMI-based discriminative training, which reduces the equal error rate (EER) by 15.9% relatively, in comparison to the GMM-MAP scheme. Guruprasad Saikumar, Amit Srivastava, Premkumar Natarajan |
ICASSP | 4 |
| 2011 | Efficient Orthogonal Matching Pursuit using sparse random projections for scene and video classificationabstractSparse projection has been shown to be highly effective in several domains, including image denoising and scene / object classification. However, practical application to large scale problems such as video analysis requires efficient versions of sparse projection algorithms such as Orthogonal Matching Pursuit (OMP). In particular, random projection based locality sensitive hashing (LSH) has been proposed for OMP. In this paper, we propose a novel technique called Comparison Hadamard random projection (CHRP) for further improving the efficiency of LSH within OMP. CHRP combines two techniques:(1) The Fast Johnson-Lindenstrauss Transform (FJLT) which uses a randomized Hadamard transform and sparse projection matrix for LSH, and (2) Achlioptas' random projection that uses only addition and comparison operations. Our approach provides the robustness of FJLT while completely avoiding multiplications. We empirically validate CHRP's efficacy by performing a suite of experiments for image denoising, scene classification, and video categorization. Our experiments indicate that CHRP significantly speeds-up OMP with negligible loss in classification accuracy. Shiv Vitaladevuni, Pradeep Natarajan, Rohit Prasad, Premkumar Natarajan |
ICCV | 4 |
| 2011 | OCR-Driven Writer Identification and Adaptation in an HMM Handwriting Recognition SystemabstractWe present an OCR-driven writer identification algorithm in this paper. Our algorithm learns writer-specific characteristics more precisely from explicit character alignment using the Viterbi algorithm and shows significant reduction of close-set writer identification error rates, compared with the GMM-based method. With writers' identities retrieved, we improve the performance of handwriting recognition using the HMM trained adapted on the training data of that writer. In our system, writer identification and OCR are highly interactive. They improve the performance of each other and thus show close approximation of supervised text-dependent writer identification and writer-dependent HMM handwriting. Huaigu Cao, Rohit Prasad, Premkumar Natarajan |
ICDAR | 3 |
| 2011 | Handwritten and Typewritten Text Identification and Recognition Using Hidden Markov ModelsabstractWe present a system for identification and recognition of handwritten and typewritten text from document images using hidden Markov models (HMMs) in this paper. Our text type identification uses OCR decoding to generate word boundaries followed by word-level handwritten/typewritten identification using HMMs. We show that the contextual constraints from the HMM significantly improves the identification performance over the conventional Gaussian mixture model (GMM)-based method. Type identification is then used to estimate the frame sample rates and frame width of feature sequences for HMM OCR system for each type independently. This type-dependent approach to computing the frame sample rate and frame width shows significant improvement in OCR accuracy over type-independent approaches. Huaigu Cao, Rohit Prasad, Premkumar Natarajan |
ICDAR | 3 |
| 2011 | Graph Clustering-Based Ensemble Method for Handwritten Text Line SegmentationabstractHandwritten text line segmentation on real-world data presents significant challenges that cannot be overcome by any single technique. Given the diversity of approaches and the recent advances in ensemble-based combination for pattern recognition problems, it is possible to improve the segmentation performance by combining the outputs from different line finding methods. In this paper, we propose a novel graph clustering-based approach to combine the output of an ensemble of text line segmentation algorithms. A weighted undirected graph is constructed with nodes corresponding to connected components and edge connecting pairs of connected components. Text line segmentation is then posed as the problem of minimum cost partitioning of the nodes in the graph such that each cluster corresponds to a unique line in the document image. Experimental results on a challenging Arabic field dataset using the ensemble method shows a relative gain of 18% in the F1 score over the best individual method within the ensemble. Vasant Manohar, Shiv Vitaladevuni, Huaigu Cao, Rohit Prasad, Premkumar Natarajan |
ICDAR | 5 |
| 2011 | Baseline Dependent Percentile Features for Offline Arabic Handwriting RecognitionabstractHandwritten text in Arabic and other languages exhibit significant variations in the slant and baseline of characters across words and also within a single word. Since the concept of baseline does not have a precise mathematical definition, existing approaches use heuristic methods to first identify a set of baseline relevant pixels and then fit lines/curves through them. However, for statistical features like percentiles that we use in our system, we only need an approximate curve that is close to the baseline to normalize the features. Hence we propose a two stage approach to estimate the approximate baseline. First we segment the text line into a set of components, and then estimate the baseline in each component using two methods max projection and smoothed centroid line. We incorpate the computed baseline into percentile feature computation in the BBN Byblos OCR system for an Arabic offline handwriting recognition task. Our new features, result in a 1% absolute gain and 3.1% relative gain in the word error rate on a large test set with 15K handwritten Arabic words, which is statistically significant with p-value<;0.001 using the matched pair comparison test. Further, our results show that computing fine-grained baselines from small line segments is significantly better than estimating a single baseline over the entire text line. Pradeep Natarajan, David Belanger 0001, Rohit Prasad, Matin Kamali, Krishna Subramanian 0001, Premkumar Natarajan |
ICDAR | 6 |
| 2011 | Text Extraction from Video Using Conditional Random FieldsabstractIn this paper, we describe an approach to extract text from broadcast videos. Candidate blocks are detected based on edge extraction results. Corners and geometrical features are used for the purpose of initial classification which is carried out by using a support vector machine (SVM). Considering the spatial inter-dependencies of different regions in the image, we propose a novel conditional random field (CRF) based framework which integrates the outputs of SVM into the system to improve the accuracy of labeling for blocks. The experimental results show that the proposed system achieves reliable performance for text detection/extraction from videos. Xujun Peng, Huaigu Cao, Rohit Prasad, Premkumar Natarajan |
ICDAR | 4 |
| 2011 | Automated image quality assessment for camera-captured OCRabstractCamera-captured optical character recognition (OCR) is a challenging area because of artifacts introduced during image acquisition with consumer-domain hand-held and Smart phone cameras. Critical information is lost if the user does not get immediate feedback on whether the acquired image meets the quality requirements for OCR. To avoid such information loss, we propose a novel automated image quality assessment method that predicts the degree of degradation on OCR. Unlike other image quality assessment algorithms which only deal with blurring, the proposed method quantifies image quality degradation across several artifacts and accurately predicts the impact on OCR error rate. We present evaluation results on a set of machine-printed document images which have been captured using digital cameras with different degradations. Xujun Peng, Huaigu Cao, Krishna Subramanian 0001, Rohit Prasad, Premkumar Natarajan |
ICIP | 5 |
| 2011 | Large-scale, real-time logo recognition in broadcast videosabstractRobust, real-time, multi-class logo detection in high resolution broadcast videos presents several difficult challenges. For most logos we only have a few training samples, which makes training robust classifiers hard. Also, logos could potentially occur anywhere in the image, and traditional sliding window approaches for logo/object detection are computationally intensive. We present a system that addresses these issues by first identifying a small set of possible logo locations in a frame, based on temporal continuity and multi-resolution search, and then successively pruning these locations for each logo template, using a cascade of color and edge based features. We present experimental results that demonstrate our system for detecting a total of 270 different logo classes in broadcast video from 5 different languages (English, Indonesian, Malay, Simplified and Traditional Chinese). Pradeep Natarajan, Yue Wu 0001, Shirin Saleem, Ehry MacRostie, Fred Bernardin, Rohit Prasad, Premkumar Natarajan |
ICME | 7 |
| 2011 | Source Error-Projection for Sample Selection in Phrase-Based SMT for Resource-Poor Languages
Sankaranarayanan Ananthakrishnan, Shiv Vitaladevuni, Rohit Prasad, Premkumar Natarajan |
IJCNLP | 4 |
| 2011 | On-Line Language Model Biasing for Multi-Pass Automatic Speech Recognition
Sankaranarayanan Ananthakrishnan, Stavros Tsakalidis, Rohit Prasad, Premkumar Natarajan |
INTERSPEECH | 4 |
| 2011 | Online Speech Activity Detection in Broadcast News
Guruprasad Saikumar, Saurabh Khanwalkar, Avi Herscovici, Amit Srivastava, Premkumar Natarajan |
INTERSPEECH | 7 |
| 2011 | Unsupervised Audio Analysis for Categorizing Heterogeneous Consumer Domain Videos
Pradeep Natarajan, Stavros Tsakalidis, Vasant Manohar, Rohit Prasad, Premkumar Natarajan |
INTERSPEECH | 5 |
| 2011 | Audio-visual fusion using bayesian model combination for web video retrievalabstractCombining features from multiple, heterogeneous, audio visual sources can significantly improve retrieval performance in consumer domain videos. However, such videos often contain unrelated overlaid audio content, or have significant camera motion to reliably extract visual features. We present an approach, which overcomes errors in individual feature streams by combining classifiers trained on multiple, heterogeneous feature streams using Bayesian model combination (BAYCOM). We demonstrate our method, by combining low-level audio and visual features, for classification of a large 200 hour web video corpus. The combined models outperform any of the individual features by 10%. Further, BAYCOM consistently outperforms traditional early and late fusion methods. Vasant Manohar, Stavros Tsakalidis, Pradeep Natarajan, Rohit Prasad, Premkumar Natarajan |
ACM Multimedia | 5 |
| 2011 | Robust named entity detection from optical character recognition output
Krishna Subramanian 0001, Rohit Prasad, Premkumar Natarajan |
Int. J. Document Anal. Recognit. | 3 |
| 2010 | A Semi-Supervised Batch-Mode Active Learning Strategy for Improved Statistical Machine Translation
Sankaranarayanan Ananthakrishnan, Rohit Prasad, David Stallard, Premkumar Natarajan |
CoNLL | 4 |
| 2010 | Gabor features for offline Arabic handwriting recognitionabstractMany feature extraction approaches for off-line handwriting recognition (OHR) rely on accurate binarization of gray-level images. However, high-quality binarization of most real-world documents is extremely difficult due to varying characteristics of noises artifacts common in such documents. Unlike most of these features, Gabor features do not require binarization of the document images, and thus are likely to be more robust to noises in document images. To demonstrate the efficacy of our proposed Gabor features, we perform subword recognition for off-line Arabic handwritten images using Support Vector Machines (SVM). We also compare the recognition performance with other binarization based features which have been proven to be effective in capturing shape characteristics of handwritten Arabic subwords, such as GSC (a set of gradient, structure, and concavity features) and skeleton based Graph features. Our preliminary experimental results show that Gabor features outperform Graph features and are slightly better than GSC features for Arabic subword recognition. In addition, by combining Gabor and GSC features, we obtain a significant reduction in classification error rate over using GSC or Gabor features alone. Huaigu Cao, Rohit Prasad, Anurag Bhardwaj, Premkumar Natarajan |
Document Analysis Systems | 5 |
| 2010 | The BBN document analysis service: a platform for multilingual document translationabstractIn this paper, we introduce a new operational platform for end-to-end document image analysis, recognition, and machine translation. The Raytheon BBN Document Analysis Service (BBN DAS) performs the following operations on scanned machine-print document images: (1) image pre-processing and segmentation to identify homogenous zones of text, (2) text recognition to convert the text zones into electronic text, (3) machine translation for converting the text from the native language of the document into English, and (4) document archiving and indexing for effective content-based search. BBN DAS uses a service-oriented architecture (SOA), which offers modularity and scalability for operation on hardware configurations ranging from a laptop to distributed multi-node server environments. This paper describes the platform architecture, the process of configuring it for Arabic newsprint documents and resulting performance results of the Arabic system. Ehry MacRostie, Rohit Prasad, Stephen Rawls, Matin Kamali, Huaigu Cao, Krishna Subramanian 0001, Premkumar Natarajan |
Document Analysis Systems | 7 |
| 2010 | Discriminative Sample Selection for Statistical Machine Translation
Sankaranarayanan Ananthakrishnan, Rohit Prasad, David Stallard, Premkumar Natarajan |
EMNLP | 4 |
| 2010 | Pashto speech recognition with limited pronunciation lexiconabstractAutomatic speech recognition (ASR) for low resource languages continues to be a difficult problem. In particular, colloquial dialects of Arabic, Farsi, and Pashto pose significant challenges in pronunciation dictionary creation. Therefore, most state-of-the-art ASR engines rely on the grapheme-as-phoneme approach for creating pronunciation dictionaries in these languages. While the grapheme approach simplifies ASR training, it performs significantly worse than a system trained with a high-quality phonetic dictionary. In this paper, we explore two techniques for bridging the performance gap between the grapheme and the phonetic approaches, without requiring manual pronunciations for all the words in the training data. The first approach is based on learning letter-to-sound rules from a small set of manual pronunciations in Pashto, and the second approach uses a hybrid phoneme/grapheme representation for recognition. Through experimental results on colloquial Pashto, we demonstrate that both techniques perform as well as a full phonetic system while requiring manual pronunciations for only a small fraction of the words in the acoustic training data. Rohit Prasad, Stavros Tsakalidis, Ivan Bulyko, Chia-Lin Kao, Premkumar Natarajan |
ICASSP | 5 |
| 2010 | Evaluating different confirmation strategies for speech-to-speech translation systemsabstractSpeech-to-speech translation systems have made a great deal of progress in recent years. But users of such systems still face the problem of not knowing whether the system has translated their utterance correctly. Various confirmation strategies can be used to address this problem. Some of these generate a confirmation utterance for the user to approve, such as reading back the ASR result, or performing “back-translation” to translate the system's translation output back into the source language. Other strategies use automated methods such as confidence measures to eliminate likely mistranslations. We propose a methodology for quantitatively evaluating the effectiveness of these different strategies, and present results of experiments using this methodology. David Stallard, Rohit Prasad, Shankar Ananthakrishnan, Fred Choi, Shirin Saleem, Premkumar Natarajan |
ICASSP | 6 |
| 2010 | Improvements in HMM Adaptation for Handwriting Recognition Using Writer Identification and Duration AdaptationabstractThis paper presents two techniques for improving adaptation of hidden Markov models (HMMs) for offline handwriting recognition. The first technique uses a novel writer identification algorithm to select training data for adapting writer-dependent models. This helps us get enough annotated samples for adaptation when the writers of test samples are known to have written some manuscripts in the training set. The second technique adapts the transition probabilities of the HMM using estimated mean of model durations from the initial decoding. Experimental results show significant improvements over the standard unsupervised parameter adaptation in our handwriting recognition system. Huaigu Cao, Rohit Prasad, Premkumar Natarajan |
ICFHR | 3 |
| 2010 | Stochastic Segment Model Adaptation for Offline Handwriting RecognitionabstractIn this paper, we present techniques for unsupervised adaptation of stochastic segment models to improve accuracy on large vocabulary offline handwriting recognition (OHR) tasks. We build upon our previous work on stochastic segment modeling for Arabic OHR. In our previous work, stochastic character segments for each n-best hypothesis were generated by a hidden Markov model (HMM) recognizer, and then a segmental model was used as an additional knowledge source for re-ranking the n-best list. Here, we describe a novel framework for unsupervised adaptation. It integrates both HMM and segment model adaptation to achieve significant gains over un-adapted recognition. Experimental results demonstrate the efficacy of our proposed method on a large corpus of handwritten Arabic documents. Rohit Prasad, Anurag Bhardwaj, Krishna Subramanian 0001, Huaigu Cao, Premkumar Natarajan |
ICPR | 5 |
| 2010 | Consensus Network Based Hypotheses Combination for Arabic Offline Handwriting RecognitionabstractOffline handwriting recognition (OHR) is an extremely challenging task because of many factors including variations in writing style, writing device and material, and noise in the scanning and collection process. Due to the diverse nature of the above challenges, it is highly unlikely that a single recognition technique can address all the characteristics of real-world handwritten documents. Therefore, one must consider designing different systems, each addressing specific challenges in the handwritten corpus, and then combining the hypotheses from these diverse systems. To that end, we present an innovative approach for combining hypotheses from multiple handwriting recognition systems. Our approach is based on generating a consensus network using hypotheses from a diverse set of handwriting recognition systems. Next, we decode the consensus network for producing the best possible hypothesis given an error criterion. Experimental results on an Arabic OHR task show that our combination algorithm outperforms the NIST ROVER technique and results in a 7% relative reduction in the word error rate over the single best OHR system. Rohit Prasad, Matin Kamali, David Belanger 0001, Antti-Veikko I. Rosti, Spyridon Matsoukas, Premkumar Natarajan |
ICPR | 6 |
| 2010 | Phrase alignment confidence for statistical machine translation
Sankaranarayanan Ananthakrishnan, Rohit Prasad, Premkumar Natarajan |
INTERSPEECH | 3 |
| 2010 | Multi resolution discriminative models for subvocalic speech recognition
Mark Raugas, Vivek Kumar Rangarajan Sridhar, Rohit Prasad, Premkumar Natarajan |
INTERSPEECH | 4 |
| 2010 | Mutual information analysis for feature and sensor subset selection in surface electromyography based speech recognition
Vivek Kumar Rangarajan Sridhar, Rohit Prasad, Premkumar Natarajan |
INTERSPEECH | 3 |
| 2010 | An unsupervised boosting technique for refiningword alignmentabstractTranslation rules extracted from automatic word alignment form the basis of statistical machine translation (SMT) systems. An unsupervised expectation-maximization (EM) algorithm is typically used to obtain a word alignment from parallel corpora. Being statistically-driven, the alignments produced by this technique are often erroneous. In this paper, we propose an unsupervised boosting strategy for refining automatic word alignment with the goal of improving SMT performance. The proposed approach results in fewer unaligned words, a significant reduction in the number of extracted translation phrase pairs, a corresponding improvement in SMT decoding speed, and a consistent improvement in translation accuracy, as measured by BLEU, across multiple language pairs and test sets. The reduction in storage and processing requirements coupled with improved accuracy make the proposed technique ideally suited for interactive translation services, facilitating applications such as mobile speech-to-speech translation. Sankaranarayanan Ananthakrishnan, Rohit Prasad, Premkumar Natarajan |
SLT | 3 |
| 2009 | Context-dependent pronunciation modeling for Iraqi ASRabstractIn this paper, we introduce a novel pronunciation modeling technique that in contrast to existing techniques uses word context information. This context-dependent pronunciation modeling is designed to overcome the challenges posed by absence of diacritics in transcripts for training acoustic models for Arabic dialects. To demonstrate the efficacy of the proposed pronunciation modeling, we present experimental results with both manually created and automatically generated vowelized lexicons on the DARPA TRANSTAC colloquial Iraqi corpus. Stavros Tsakalidis, Rohit Prasad, Premkumar Natarajan |
ICASSP | 3 |
| 2009 | Unsupervised HMM Adaptation Using Page Style ClusteringabstractIn this paper we present an innovative two-stage adaptation approach for handwriting recognition that is based on clustering of similar pages in the training data. In our approach, we first perform page clustering on training data using features such as contour slope, pen pressure, writing velocity, and stroke sparseness. Next, we adapt the writer-independent hidden Markov models (HMMs) to each cluster in the training data. While decoding a test page, we first determine the cluster the test page belongs to and then decode the page with the model associated with that cluster. Experimental results with the two-stage adaptation show significant gains on a held-out validation set. Huaigu Cao, Rohit Prasad, Shirin Saleem, Premkumar Natarajan |
ICDAR | 4 |
| 2009 | Stochastic Segment Modeling for Offline Handwriting RecognitionabstractIn this paper, we present a novel approach for incorporating structural information into the hidden Markov modeling (HMM) framework for offline handwriting recognition. Traditionally, structural features have been used in recognition approaches that rely on accurate segmentation of words into smaller units (sub-words or characters). However, such segmentation based approaches do not perform well on real-world handwritten images, because breaks and merges in glyphs typically create new connected components that are not observed in the training data. To mitigate the problem of having to derive accurate segmentation from connected components, we present a novel framework where the HMM based recognition system trained on shorter-span features is used to generate the 2D character images (the ldquostochastic segmentsrdquo), and then another classifier that uses structural features extracted from the stochastic character segments generates a new set of scores. Finally, the scores from the HMM system and from structural matching are used in combination to generate a hypothesis that is better than the results from either the HMM or from structural matching alone. We demonstrate the efficacy of our approach by reporting experimental results on a large corpus of handwritten Arabic documents. Premkumar Natarajan, Krishna Subramanian 0001, Anurag Bhardwaj, Rohit Prasad |
ICDAR | 1 |
| 2009 | Improvements in BBN's HMM-Based Offline Arabic Handwriting Recognition SystemabstractOffline handwriting recognition of free-flowing Arabic text is a challenging task due to the plethora of factors that contribute to the variability in the data. In this paper, we address some of these sources of variability, and present experimental results on a large corpus of handwritten documents. Specific techniques such as the application of context-dependent Hidden Markov Models (HMMs) for the cursive Arabic script, unsupervised adaptation to account for the stylistic variations across scribes, and image pre-processing to remove ruled-lines are explored. In particular, we proposed a novel integration of structural features in the HMM framework which exclusively results in a 9% relative improvement in performance. Overall, we demonstrate a relative reduction of 17% in word error rate over our baseline Arabic handwriting recognition system. Shirin Saleem, Huaigu Cao, Krishna Subramanian 0001, Matin Kamali, Rohit Prasad, Premkumar Natarajan |
ICDAR | 6 |
| 2009 | Nested state indexing in pairwise Markov networks for fast handwritten document image rule-line removalabstractThe Markov random field (MRF) has been applied to modeling the connectivity constraints of the text in document images for tasks like binarization and rule-line removal. One challenge of applying the MRF is its high computational cost. This paper presents a method using two nested set of states trained to reduce the computational cost of patch-based MRF. The two sets of states are trained at different levels in coarse-to-fine order. We show effective reduction of run time but very little loss of quality using rule-line removal experiments. Huaigu Cao, Rohit Prasad, Premkumar Natarajan, Venu Govindaraju |
ICIP | 3 |
| 2008 | End-to-End Trainable Thai OCR System Using Hidden Markov ModelsabstractIn this paper we present an end-to-end trainable optical character recognition (OCR) system for recognizing machine-printed text in Thai documents. The end-to-end OCR system is based on a script-independent methodology using hidden Markov models. Our system provides an integrated workflow beginning with annotation and transcription of training images to performing OCR on new images with models trained on transcribed training images. The efficacy of our end-to-end OCR system is demonstrated by rapidly configuring our OCR engine for the Thai script. We present experimental results on Thai documents to highlight the specific challenges posed by the Thai script and analyze the recognition performance as a function of amount of training data. Kriste Krstovski, Ehry MacRostie, Rohit Prasad, Premkumar Natarajan |
Document Analysis Systems | 4 |
| 2008 | Semi-supervised topic classification for low resource languagesabstractIn this paper, we present a novel methodology for rapidly developing a topic-based document classification system for a language that has limited resources. Our approach, a hybrid one, combines supervised and unsupervised topic classification techniques. Given that access to native speakers is fairly limited for low resource languages, our approach requires annotating only a few broad “root” topics in the corpus. Next, unsupervised topic discovery (UTD) technique is used to automatically determine finer topics within the root topics. Lastly, we use the recently developed unsupervised topic clustering technique to organize the corpus into a hierarchical structure that enables browsing documents at multiple levels of granularity. Recognizing the need for reducing false alarms during runtime, we describe rejection techniques for discarding off-topic documents. Daben Liu, Sam McVeety, Rohit Prasad, Premkumar Natarajan |
ICASSP | 4 |
| 2008 | Multi-frame combination for robust videotext recognitionabstractOptical character recognition (OCR) of overlaid text in video streams is a challenging problem due to various factors including the presence of dynamic backgrounds, color, and low resolution. In video feeds such as Broadcast News, a particular overlaid text region usually persists for multiple frames during which the background may or may not vary. In this paper we explore two innovative techniques that exploit such multi-frame persistence of videotext. The first technique uses multiple instances to generate a single enhanced image for recognition. The second technique uses the NIST ROVER algorithm developed for speech recognition to combine 1-best hypotheses from different frames of a text region. Significant improvement in the word error rate (WER) is obtained by using ROVER when compared to recognizing a single instance. The WER is further reduced by combining hypotheses from frame instances, which were generated using character models trained with different binarization thresholds. A 20% relative reduction in the WER was achieved for multi-frame combination over decoding a single frame instance. Rohit Prasad, Shirin Saleem, Ehry MacRostie, Premkumar Natarajan, Michael Decerbo |
ICASSP | 4 |
| 2008 | Recent improvements and performance analysis of ASR and MT in a speech-to-speech translation systemabstractWe report on recent ASR and MT work on our English/Iraqi Arabic speech-to-speech translation system. We present detailed results for both objective and subjective evaluations of translation quality, along with a detailed analysis and categorization of translation errors. We also present novel ideas for quantifying the relative importance of different subjective error categories, and for assigning the blame for an error to a particular phrase pair in the translation model. David Stallard, Chia-Lin Kao, Kriste Krstovski, Daben Liu, Premkumar Natarajan, Rohit Prasad, Shirin Saleem, Krishna Subramanian 0001 |
ICASSP | 5 |
| 2008 | Robust named entity detection in videotext using character latticesabstractText in video sequences can provide key indexing information. In particular, videotext is rich in named entities (NEs) and detection of such entities is critical for search applications. Traditional approaches for detecting NEs in OCR output look for these NEs in the single-best recognition results. Due to inevitable presence of recognition errors in the single-best output, such approaches usually result in low recall. Given that a lattice is more likely to contain the correct answer, we explore NE detection from character lattices produced by our videotext OCR system. Furthermore, we use an approximate match criterion that allows insertion of punctuations during lookup. Experimental results show a 50% relative improvement in NE recall using lattices over exact lookup in the 1-best hypothesis. Since the improvement in recall is accompanied by a large number of false positives, we present techniques for reducing false alarms. In addition, we describe efficient techniques for reducing the time for detecting NEs. Krishna Subramanian 0001, Rohit Prasad, Ehry MacRostie, Premkumar Natarajan |
ICASSP | 4 |
| 2008 | Improvements in hidden Markov model based Arabic OCRabstractThis paper describes recent advances in hidden Markov model (HMM) based OCR for machine-printed arabic documents. A combination of script-independent and script-specific techniques are applied to glyph models and language models (LM). Script-independent techniques we applied are higher order n-gram LMs for N-best rescoring and discriminative estimation of glyph HMMs. Arabic specific techniques include the use of context-dependent HMMs for glyph modeling and Parts-of-Arabic-Words in language modeling. We present experimental results that demonstrate a 40% relative reduction in word error rate over the baseline configuration on a corpus of machine-printed Arabic documents. Rohit Prasad, Shirin Saleem, Matin Kamali, Ralf Meermeier, Premkumar Natarajan |
ICPR | 5 |
| 2008 | Recent improvements in BBN's English/Iraqi speech-to-speech translation systemabstractWe report on recent improvements in our English/Iraqi Arabic speech-to-speech translation system. User interface improvements include a novel parallel approach to user confirmation which makes confirmation cost-free in terms of dialog duration. Automatic speech recognition improvements include the incorporation of state-of-the-art techniques in feature transformation and discriminative training. Machine translation improvements include a novel combination of multiple alignments derived from various pre-processing techniques, such as Arabic segmentation and English word compounding, higher order N-grams for target language model, and use of context in form of semantic classes and part-of-speech tags. Fred Choi, Stavros Tsakalidis, Shirin Saleem, Chia-Lin Kao, Ralf Meermeier, Kriste Krstovski, Christine Moran, Krishna Subramanian 0001, Rohit Prasad, Premkumar Natarajan |
SLT | 10 |
| 2008 | Name aware speech-to-speech translation for English/IraqiabstractIn this paper, we describe a novel approach that exploits intra-sentence and dialog-level context for improving translation performance on spoken Iraqi utterances that contain named entities (NEs). Dialog-level context is used to predict whether the Iraqi response is likely to contain names and the intra-sentence context is used to determine words that are named entities. While we do not address the problem of translating out-of-vocabulary (OOV) NEs in spoken utterances, we show that our approach is capable of translating OOV names in text input. To demonstrate efficacy of our approach, we present results on internal test set as well as the 2008 June DARPA TRANSTAC name evaluation set. Rohit Prasad, Christine Moran, Fred Choi, Ralf Meermeier, Shirin Saleem, Chia-Lin Kao, David Stallard, Premkumar Natarajan |
SLT | 8 |
| 2007 | Semantic translation error rate for evaluating translation systemsabstractIn this paper, we introduce a new metric which we call the semantic translation error rate, or STER, for evaluating the performance of machine translation systems. STER is based on the previously published translation error rate (TER) (Snover et al., 2006) and METEOR (Banerjee and Lavie, 2005) metrics. Specifically, STER extends TER in two ways: first, by incorporating word equivalence measures (WordNet and Porter stemming) standardly used by METEOR, and second, by disallowing alignments of concept words to non-concept words (aka stop words). We show how these features make STER alignments better suited for human-driven analysis than standard TER. We also present experimental results that show that STER is better correlated to human judgments than TER. Finally, we compare STER to METEOR, and illustrate that METEOR scores computed using the STER alignments have similar statistical properties to METEOR scores computed using METEOR alignments. Krishna Subramanian 0001, David Stallard, Rohit Prasad, Shirin Saleem, Premkumar Natarajan |
ASRU | 5 |
| 2007 | Optimal Estimation of Rejection Thresholds for Topic SpottingabstractIn many applications of topic spotting technology, especially those that require a human review of in-topic documents, a low false alarm rate is a key requirement. Topic spotting techniques typically include a rejection scheme to filter out off-topic documents. In this paper we present a robust methodology for rejecting off-topic messages that, in addition to modeling the topics of interest, uses a so-called alternate model for topics that are not included in the set of topics of interest. Specifically, we introduce two novel techniques for estimating topic-specific rejection thresholds - a parametric technique that can be viewed as transformation of topic-independent thresholds, and a nonparametric technique based on constrained optimization of false rejections subject to a pre-specified number of false acceptances. Our experiments on newsgroup messages demonstrate that when adequate training data is available topic-specific threshold estimation techniques can outperform topic-independent thresholds in terms of the ROC curve. Krishna Subramanian 0001, Rohit Prasad, Premkumar Natarajan, Richard M. Schwartz |
ICASSP (4) | 3 |
| 2007 | Robust Page Segmentation Based on Smearing and Error Correction Unifying Top-down and Bottom-up ApproachesabstractIn this paper we present a robust multi-pass page segmentation algorithm. The first pass uses a modified smearing algorithm and the second pass performs a hybrid of bottom-up and top-down segmentation on the output of the first pass. Unlike traditional approaches, the bottom-up and top-down steps are based on primitive results of a smearing based page segmentation algorithm. Therefore, "split" and "merge" processes start with text blocks that are mostly true text blocks but a few of them are either touching or broken. We present experimental results on newspaper and journal documents from different languages to demonstrate the robustness and language independence of our approach. Huaigu Cao, Rohit Prasad, Premkumar Natarajan, Ehry MacRostie |
ICDAR | 3 |
| 2007 | Character-Stroke Detection for Text-Localization and ExtractionabstractIn this paper, we present a new approach for analysis of images for text-localization and extraction. Our approach puts very few constraints on the font, size and color of text and is capable of handling both scene text and artificial text well. In this paper, we exploit two well-known features of text: approximately constant stroke width and local contrast, and develop a fast, simple, and effective algorithm to detect character strokes. We also show how these can be used for accurate extraction and motivate some advantages of using this approach for text localization over other color-space segmentation based approaches. We analyze the performance of our stroke detection algorithm on images collected for the robust-reading competitions at ICDAR 2003. Krishna Subramanian 0001, Premkumar Natarajan, Michael Decerbo, David A. Castañón |
ICDAR | 2 |
| 2007 | Improvements in machine translation for English/iraqi speech translation
Shirin Saleem, Krishna Subramanian 0001, Rohit Prasad, David Stallard, Chia-Lin Kao, Premkumar Natarajan, Raid Suleiman |
INTERSPEECH | 6 |
| 2007 | The BBN 2007 displayless English/iraqi speech-to-speech translation system
David Stallard, Fred Choi, Chia-Lin Kao, Kriste Krstovski, Premkumar Natarajan, Rohit Prasad, Shirin Saleem, Krishna Subramanian 0001 |
INTERSPEECH | 5 |
| 2007 | Finding structure in noisy text: topic classification and unsupervised clustering
Premkumar Natarajan, Rohit Prasad, Krishna Subramanian 0001, Shirin Saleem, Fred Choi, Richard M. Schwartz |
Int. J. Document Anal. Recognit. | 1 |
| 2006 | Colloquial Iraqi ASR for speech translation
Shirin Saleem, Rohit Prasad, Premkumar Natarajan |
INTERSPEECH | 3 |
| 2006 | A hybrid phrase-based/statistical speech translation system
David Stallard, Fred Choi, Kriste Krstovski, Premkumar Natarajan, Rohit Prasad, Shirin Saleem |
INTERSPEECH | 4 |
| 2006 | Design and Evaluation of the 2006 BBN English/Iraqi Two-Way speech Translation SystemabstractIn this paper, we present a 2-way speech-to-speech translation system for English and Iraqi colloquial Arabic, the dialect of Arabic spoken by ordinary people in Iraq. The application domain of the system is military force protection, including municipal services surveys, detainee screening, and descriptions of people, houses, vehicles, etc. The system uses statistical speech recognition, and a combination of prerecorded questions and statistical machine translation with speech synthesis to translate the speech recognition output. We present evaluation results, along with an analysis of the gap between Iraqi-to-English and English-to-Iraqi translation performance. David Stallard, Fred Choi, Kriste Krstovski, Premkumar Natarajan, Rohit Prasad, Shirin Saleem, Raid Suleiman |
SLT | 4 |
| 2005 | Performance Improvements to the BBN Byblos OCR SystemabstractIn this paper, we describe four recent enhancements to the BBN Byblos OCR system, a multilingual HMM-based character recognition system which has been demonstrated on a variety of languages, including English, Arabic, Chinese, and Japanese. These enhancements are implemented as optional extensions to the system and provide improved performance for certain scripts or domains. Projection-based re-estimation of line boundaries reduces instability in the presence of some types of noise. An alternate modeling strategy used in the first of two recognition search passes substantially increases speed on languages with a large number of characters. Another speed improvement comes from automatic discovery and modeling of sub-characters. The use of heteroschedastic linear discriminant analysis (HLDA) makes modeling more tractable by reducing feature-space dimensionality. Michael Decerbo, Premkumar Natarajan, Rohit Prasad, Ehry MacRostie |
ICDAR | 2 |
| 2005 | Character Duration Modeling for Speed Improvements in the BBN Byblos OCR SystemabstractIn this paper, we describe a recent enhancement to our HMM-based OCR system that results in a significant increase in the speed of the system without any impact on recognition accuracy. Recognition speed is, in part, a function of the number of distinct HMMs that constitute the model set. As a result, the recognition speed is much slower for ideographic scripts, such as Chinese and Japanese which contain thousands of glyphs, than for alphabetic scripts such as Latin and Arabic. In our current OCR system, methods like sub-character modeling and Gaussian shortlists are used to reduce the processing time. In this paper, we describe a simple character-based duration modeling technique that puts a duration constraint on the number of frames for which a character can stay active. Character durations were obtained from automatically labeled training data and a probability mass function (histogram) was used to model character durations. The use of a duration model yielded a 37% improvement in speed with no loss in accuracy. Premkumar Natarajan, Ram Sundaram, Rohit Prasad, Ehry MacRostie |
ICDAR | 1 |
| 2003 | Design and evaluation of a limited two-way speech translator
David Stallard, John Makhoul, Fred Choi, Ehry MacRostie, Premkumar Natarajan, Richard M. Schwartz, Bushra Zawaydeh |
INTERSPEECH | 5 |
| 2003 | Surprise! What's in a Cebuano or Hindi Name?abstractEmpirical results are presented for creating training data and training a statistical name learning algorithm on Cebuano and Hindi in roughly three weeks time. The empirical study compares performance in a compressed time frame against performance of the same statistical language model in English (where there was no compressed time frame). Rapid development of several co-reference heuristics in Hindi are also described, and co-reference performance in Hindi is compared to previously developed English techniques. Jonathan May, Ada Brunstein, Premkumar Natarajan, Ralph M. Weischedel |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2002 | A scalable architecture for Directory Assistance automationabstractWe present a novel architecture for providing automated telephone Directory Assistance (DA). The architecture couples a large-vocabulary, statistical n-gram, speech recognition engine with a statistical retrieval system. The use of a statistical n-gram allows for the recognition of unconstrained spoken queries while the statistical retrieval engine allows for an inexact match between a particular spoken query and the training data. Allowing for unconstrained recognition and an inexact match provides the framework for high levels of automation. Once the retrieval engine returns a ranked set of frequently requested telephone numbers (FRN), the rejection module uses a classifier to compute a confidence-like score that is used to make the automation decision. With actual customer calls into an operational, automated DA call center and an FRN set size of 25000 numbers, the new architecture is capable of delivering more than 17% correct automation at a false accept rate of 0.76%. Premkumar Natarajan, Rohit Prasad, Richard M. Schwartz, John Makhoul |
ICASSP | 1 |
| 2002 | Speech-enabled natural language call routing: BBN call director
Premkumar Natarajan, Rohit Prasad, Bernhard Suhm, Daniel McCarthy |
INTERSPEECH | 1 |
| 2002 | Design for a speech-to-speech translator for field use
David Stallard, Premkumar Natarajan, Mohammed Noamany, Richard M. Schwartz, John Makhoul |
INTERSPEECH | 2 |
| 2001 | Videotext OCR Using Hidden Markov ModelsabstractWe present a method for performing optical character recognition (OCR) of text in video images. Recognition of videotext is a challenging problem due to various factors such as the presence of rich, dynamic backgrounds, low resolution, color, etc. Our strategy is to process the video images to produce high-resolution binarized text images that resemble printed text. We describe a novel clustering and relaxation procedure that combines stroke and color information to separate the text from the background. The binarized text image is then recognized with our Byblos OCR engine (Natarajan et al., 2001; Schwartz et al., 1996) using hidden Markov models trained on similar data. We present experimental results on a video-data corpus collected from broadcast news programs. Currently the system delivers a character error rate of 8.3% on independent multi-font test data from this corpus. Premkumar Natarajan, Baback Elmieh, Richard M. Schwartz, John Makhoul |
ICDAR | 1 |
| 2001 | Multilingual Machine Printed OCRabstractThis paper presents a script-independent methodology for optical character recognition (OCR) based on the use of hidden Markov models (HMM). The feature extraction, training and recognition components of the system are all designed to be script independent. The training and recognition components were taken without modification from a continuous speech recognition system; the only component that is specific to OCR is the feature extraction component. To port the system to a new language, all that is needed is text image training data from the new language, along with ground truth which gives the identity of the sequences of characters along each line of each text image, without specifying the location of the characters on the image. The parameters of the character HMMs are estimated automatically from the training data, without the need for laborious handwritten rules. The system does not require presegmentation of the data, neither at the word level nor at the character level. Thus, the system is able to handle languages with connected characters in a straightforward manner. The script independence of the system is demonstrated in three languages with different types of script: Arabic, English, and Chinese. The robustness of the system is further demonstrated by testing the system on fax data. An unsupervised adaptation method is then described to improve performance under degraded conditions. Premkumar Natarajan, Zhidong Lu, Richard M. Schwartz, Issam Bazzi, John Makhoul |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1999 | Advances in the BBN BYBLOS OCR SystemabstractWe present some recent advances in the BBN BYBLOS OCR system. This OCR system can be used to recognize Arabic, Chinese, and English with high accuracy. A major change in the system is the use of continuous-density HMMs, which allow us to take advantage of a large amount of training data and to use unsupervised adaptation methods to improve accuracy in many cases, e.g., on degraded data. Another advance is the substantial increase in recognition speed. With this increased speed, the system is fast enough for practical use on Arabic and English data. The extension of the system to Chinese further demonstrated the language independence of this system and showed that this system can be used on languages with large character sets and complicated character structures. The Chinese OCR system yielded high accuracy on newspaper data. Zhidong Lu, Richard M. Schwartz, Premkumar Natarajan, Issam Bazzi, John Makhoul |
ICDAR | 3 |
| 1999 | Robust OCR of Degraded DocumentsabstractThis paper is concerned with techniques for performing robust OCR of degraded documents, such us faxed text, using a hidden Markov model (HMM) based OCR system. We present two strategies for dealing with degraded documents. The first strategy is to train the system on degraded documents that have been subjected to the same, or similar, degradation process as the documents to be recognized. The second, more sophisticated, strategy is to use adaptation to adjust the parameters of the trained model in order to improve recognition accuracy on a specific document. This adjustment of model parameters is typically posed as a constrained optimization problem wherein a certain prespecified objective function is to be optimized. We present a comparative study of two objective functions. The likelihood function and the posterior probability. A variation of the basic posterior probability method is also discussed. Using adaptation with a model trained on fax-degraded data we have reduced, by a factor of three, the character error rate on fax-degraded text images generated from the University of Washington English Image Database I. Premkumar Natarajan, Issam Bazzi, Zhidong Lu, John Makhoul, Richard M. Schwartz |
ICDAR | 1 |