Rajiv Jain

dblp:95/123 · DBLP profile ↗
← Back
64ranked-venue papers
14as first author
25since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 5 first-author · 19 since 2021Systems, architecture and hardware · 27 · 7 first-authorDatabases, data management, data science and information retrieval · 10 · 6 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 FormA11y - Research and Development of a Tool for Remediating PDF Forms for Accessibility
abstract
PDF documents are usually not born-accessible, and so document authors need to put in additional work (remediation) to make them accessible for people with disabilities. Unfortunately, this step is often overlooked and hard to execute, resulting in a large number of inaccessible PDF documents on the internet. Previously, there have been research efforts to investigate potential solutions for remediating PDF documents for accessibility. However, most of the existing research focuses on accessibility of long or scientific PDF documents meant for passive reading. PDF documents come in different types, and this research project focuses on a distinct type of PDF document-forms-where the user is required to interact with the PDF document and enter data. Through our research work we identified that the PDF form remediation process is non-intuitive, repetitive, and overwhelming due to the high-information density of PDF forms, and existing research and tools do not yet address the challenges. Our research work culminated in the creation of a tool - FormA11y - that addresses these challenges by making the repetitive and painstaking process of form remediation easier. To evaluate the effectiveness and efficiency of FormA11y against the industry standard tool - Adobe Acrobat - for PDF form remediation, we performed a within-subject user study with 20 participants. With FormA11y, users remediated forms 2.8 times faster while creating more accurately accessible PDF forms.
Sparsh Paliwal, Joshua Hoeflich, J. Bern Jordan, Rajiv Jain, Vlad I. Morariu, Alexa F. Siu, Jonathan Lazar
ACM Trans. Comput. Hum. Interact.4
2024 DocScript: Document-level Script Event Prediction
abstract
We present a novel task of document-level script event prediction, which aims to predict the next event given a candidate list of narrative events in long-form documents. To enable this, we introduce DocSEP, a challenging dataset in two new domains - contractual documents and Wikipedia articles, where timeline events may be paragraphs apart and may require multi-hop temporal and causal reasoning. We benchmark existing baselines and present a novel architecture called DocScript to learn sequential ordering between events at the document scale. Our experimental results on the DocSEP dataset demonstrate that learning longer-range dependencies between events is a key challenge and show that contemporary LLMs such as ChatGPT and FlanT5 struggle to solve this task, indicating their lack of reasoning abilities for understanding causal relationships and temporal sequences within long texts.
Puneet Mathur, Vlad I. Morariu, Aparna Garimella, Franck Dernoncourt, Jiuxiang Gu, Ramit Sawhney, Preslav Nakov, Dinesh Manocha, Rajiv Jain
LREC/COLING9
2024 DocEdit-v2: Document Structure Editing Via Multimodal LLM Grounding
abstract
Manan Suri, Puneet Mathur, Franck Dernoncourt, Rajiv Jain, Vlad I Morariu, Ramit Sawhney, Preslav Nakov, Dinesh Manocha. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Manan Suri, Puneet Mathur, Franck Dernoncourt, Rajiv Jain, Vlad I. Morariu, Ramit Sawhney, Preslav Nakov, Dinesh Manocha
EMNLP4
2024 FlexDoc: Flexible Document Adaptation through Optimizing both Content and Layout
abstract
Designing adaptive documents that are visually appealing across various devices and for diverse viewers is a challenging task. This is due to the wide variety of devices and different viewer requirements and preferences. Alterations to a document’s content, style, or layout often necessitate numerous adjustments, potentially leading to a complete layout redesign. We introduce FlexDoc, a framework for creating and consuming documents that seamlessly adapt to different devices, author, and viewer preferences and interactions. It eliminates the need to manually create multiple document layouts, as FlexDoc enables authors to define desired document properties using templates and employs both discrete and continuous optimization in a novel comprehensive optimization process, which leverages automatic text summarization and image carving techniques to adapt both layout and content during consumption dynamically. Further, we demonstrate FlexDoc in real-world scenarios.
Yue Jiang 0002, Christof Lutteroth, Rajiv Jain, Chris Tensmeyer, Varun Manjunatha, Wolfgang Stuerzlinger, Vlad I. Morariu
VL/HCC3
2023 DocEdit: Language-Guided Document Editing
abstract
Professional document editing tools require a certain level of expertise to perform complex edit operations. To make editing tools accessible to increasingly novice users, we investigate intelligent document assistant systems that can make or suggest edits based on a user's natural language request. Such a system should be able to understand the user's ambiguous requests and contextualize them to the visual cues and textual content found in a document image to edit localized unstructured text and structured layouts. To this end, we propose a new task of language-guided localized document editing, where the user provides a document and an open vocabulary editing request, and the intelligent system produces a command that can be used to automate edits in real-world document editing software. In support of this task, we curate the DocEdit dataset, a collection of approximately 28K instances of user edit requests over PDF and design templates along with their corresponding ground truth software executable commands. To our knowledge, this is the first dataset that provides a diverse mix of edit operations with direct and indirect references to the embedded text and visual objects such as paragraphs, lists, tables, etc. We also propose DocEditor, a Transformer-based localization-aware multimodal (textual, spatial, and visual) model that performs the new task. The model attends to both document objects and related text contents which may be referred to in a user edit request, generating a multimodal embedding that is used to predict an edit command and associated bounding box localizing it. Our proposed model empirically outperforms other baseline deep learning approaches by 15-18%, providing a strong starting point for future work.
Puneet Mathur, Rajiv Jain, Jiuxiang Gu, Franck Dernoncourt, Dinesh Manocha, Vlad I. Morariu
AAAI2
2023 Characteristics of Deep and Skim Reading on Smartphones vs. Desktop: A Comparative Study
abstract
Deep reading fosters text comprehension, memory, and critical thinking. The growing prevalance of digital reading on mobile interfaces raises concerns that deep reading is being replaced by skimming and sifting through information, but this is currently unmeasured. Traditionally, reading quality is assessed using comprehension tests, which require readers to explicitly answer a set of carefully composed questions. To quantify and understand reading behaviour in natural settings and at scale, however, implicit measures are needed of deep versus skim reading across desktop and mobile devices, the most prominent digital reading platforms. In this paper, we present an approach to systematically induce deep and skim reading and subsequently train classifiers to discriminate these two reading styles based on eye movement patterns and interaction data. Based on a user study with 29 participants, we created models that detect deep reading on both devices with up to 0.82 AUC. We present the characteristics of deep reading and discuss how our models can be used to measure the effect of reading UI design and monitor long-term changes in reading behaviours.
Xiuge Chen, Namrata Srivastava, Rajiv Jain, Jennifer A. Healey, Tilman Dingler
CHI3
2023 Exposing Model Theft: A Robust and Transferable Watermark for Thwarting Model Extraction Attacks
abstract
The increasing prevalence of Deep Neural Networks (DNNs) in cloud-based services has led to their widespread use through various APIs. However, recent studies reveal the susceptibility of these public APIs to model extraction attacks, where adversaries attempt to create a local duplicate of the private model using data and API-generated predictions. Existing defense methods often involve perturbing prediction distributions to hinder an attacker's training goals, inadvertently affecting API utility. In this study, we extend the concept of digital watermarking to protect DNNs' APIs. We suggest embedding a watermark into the safeguarded APIs; thus, any model attempting to copy will inherently carry the watermark, allowing the defender to verify any suspicious models. We propose a simple yet effective framework to increase watermark transferability. By requiring the model to memorize the preset watermarks in the final decision layers, we significantly enhance the transferability of watermarks. Comprehensive experiments show that our proposed framework not only successfully watermarks APIs but also maintains their utility.
Ruixiang Tang, Hongye Jin, Mengnan Du, Curtis Wigington, Rajiv Jain, Xia Ben Hu
CIKM5
2023 Computable Contracts by Extracting Obligation Logic Graphs
abstract
The emergence of contract specific programming languages has struggled to translate into widespread adoption of computable contracts due largely to high conversion costs. In this work, we present the first system for converting natural language contracts into code through the extraction of key entities, relationships, and formulas into a graph representation called the Obligation Logic Graph (OLG). This approach allows the semantic meaning of contract obligations, including dependencies between obligations, to be captured through the OLG and mapped to code downstream. We also introduce OLG extraction as a new joint entity and relation prediction task for legal contracts, and present the Contract-OLG dataset, consisting of 1,876 contract provisions, 18,597 entities and 18,170 relationships. We perform detailed experiments to understand the capabilities of state-of-the-art Transformer and graph-based models at completing these tasks, and identify where there is currently a significant gap between human expert and machine performance, particularly for relation extraction.
Sergio Servantez, Nedim Lipka, Alexa F. Siu, Milan Aggarwal, Balaji Krishnamurthy, Aparna Garimella, Kristian J. Hammond, Rajiv Jain
ICAIL8
2023 LayerDoc: Layer-wise Extraction of Spatial Hierarchical Structure in Visually-Rich Documents
abstract
Digital documents often contain images and scanned text. Parsing such visually-rich documents is a core task for work-flow automation, but it remains challenging since most documents do not encode explicit layout information, e.g., how characters and words are grouped into boxes and ordered into larger semantic entities. Current state-of-the-art layout extraction methods are challenged by such documents as they rely on word sequences to have correct reading order and do not exploit their hierarchical structure. We propose LayerDoc, an approach that uses visual features, textual semantics, and spatial coordinates along with constraint inference to extract the hierarchical layout structure of documents in a bottom-up layer-wise fashion. LayerDoc recursively groups smaller regions into larger semantic elements in 2D to infer complex nested hierarchies. Experiments show that our approach outperforms competitive baselines by 10-15% on three diverse datasets of forms and mobile app screen layouts for the tasks of spatial region classification, higher-order group identification, layout hierarchy extraction, reading order detection, and word grouping.
Puneet Mathur, Rajiv Jain, Ashutosh Mehra 0002, Jiuxiang Gu, Franck Dernoncourt, Anandhavelu Natarajan, Quan Hung Tran, Verena Kaynig, Ani Nenkova, Dinesh Manocha, Vlad I. Morariu
WACV2
2023 Editorial for special issue on "advanced topics in document analysis and recognition"
Koichi Kise, Richard Zanibbi, Rajiv Jain, Gernot A. Fink
Int. J. Document Anal. Recognit.3
2022 User-Entity Differential Privacy in Learning Natural Language Models
abstract
In this paper, we introduce a novel concept of user-entity differential privacy (UeDP) to provide formal privacy protection simultaneously to both sensitive entities in textual data and data owners in learning natural language models (NLMs). To preserve UeDP, we developed a novel algorithm, called UeDP-Alg, optimizing the trade-off between privacy loss and model utility with a tight sensitivity bound derived from seamlessly combining user and sensitive entity sampling processes. An extensive theoretical analysis and evaluation show that our UeDP-Alg outperforms baseline approaches in model utility under the same privacy budget consumption on several NLM tasks, using benchmark datasets.
Phung Lai, NhatHai Phan, Tong Sun 0005, Rajiv Jain, Franck Dernoncourt, Jiuxiang Gu, Nikolaos Barmpalios
IEEE Big Data4
2022 Secure and Efficient Agreement Signing Atop Blockchain and Decentralized Identity
Songlin He, Tong Sun 0005, Qiang Tang 0005, Chase Qishi Wu, Nedim Lipka, Curtis Wigington, Rajiv Jain
BlockSys7
2022 MACRONYM: A Large-Scale Dataset for Multilingual and Multi-Domain Acronym Extraction
abstract
Acronym extraction is the task of identifying acronyms and their expanded forms in texts that is necessary for various NLP applications. Despite major progress for this task in recent years, one limitation of existing AE research is that they are limited to the English language and certain domains (i.e., scientific and biomedical). Challenges of AE in other languages and domains are mainly unexplored. As such, lacking annotated datasets in multiple languages and domains has been a major issue to prevent research in this direction. To address this limitation, we propose a new dataset for multilingual and multi-domain AE. Specifically, 27,200 sentences in 6 different languages and 2 new domains, i.e., legal and scientific, are manually annotated for AE. Our experiments on the dataset show that AE in different languages and learning settings has unique challenges, emphasizing the necessity of further research on multilingual and multi-domain AE.
Amir Pouran Ben Veyseh, Nicole Meister, Seunghyun Yoon 0002, Rajiv Jain, Franck Dernoncourt, Thien Huu Nguyen
COLING4
2022 Keyphrase Prediction from Video Transcripts: New Dataset and Directions
abstract
Keyphrase Prediction (KP) is an established NLP task, aiming to yield representative phrases to summarize the main content of a given document. Despite major progress in recent years, existing works on KP have mainly focused on formal texts such as scientific papers or weblogs. The challenges of KP in informal-text domains are not yet fully studied. To this end, this work studies new challenges of KP in transcripts of videos, an understudied domain for KP that involves informal texts and non-cohesive presentation styles. A bottleneck for KP research in this domain involves the lack of high-quality and large-scale annotated data that hinders the development of advanced KP models. To address this issue, we introduce a large-scale manually-annotated KP dataset in the domain of live-stream video transcripts obtained by automatic speech recognition tools. Concretely, transcripts of 500+ hours of videos streamed on the behance.net platform are manually labeled with important keyphrases. Our analysis of the dataset reveals the challenging nature of KP in transcripts. Moreover, for the first time in KP, we demonstrate the idea of improving KP for long documents (i.e., transcripts) by feeding models with paragraph-level keyphrases, i.e., hierarchical extraction. To foster future research, we will publicly release the dataset and code.
Amir Pouran Ben Veyseh, Quan Hung Tran, Seunghyun Yoon 0002, Varun Manjunatha, Hanieh Deilamsalehy, Rajiv Jain, Trung Bui, Walter Chang, Franck Dernoncourt, Thien Huu Nguyen
COLING6
2022 Certified Neural Network Watermarks with Randomized Smoothing
abstract
Watermarking is a commonly used strategy to protect creators’ rights to digital images, videos and audio. Recently, watermarking methods have been extended to deep learning models – in principle, the watermark should be preserved when an adversary tries to copy the model. However, in practice, watermarks can often be removed by an intelligent adversary. Several papers have proposed watermarking methods that claim to be empirically resistant to different types of removal attacks, but these new techniques often fail in the face of new or better-tuned adversaries. In this paper, we propose the first certifiable watermarking method. Using the randomized smoothing technique, we show that our watermark is guaranteed to be unremovable unless the model parameters are changed by more than a certain $\ell_2$ threshold. In addition to being certifiable, our watermark is also empirically more robust compared to previous watermarking methods.
Arpit Bansal, Ping-Yeh Chiang, Michael J. Curry, Rajiv Jain, Curtis Wigington, Varun Manjunatha, John Dickerson 0001, Tom Goldstein
ICML4
2022 DocLayoutTTS: Dataset and Baselines for Layout-informed Document-level Neural Speech Synthesis
Puneet Mathur, Franck Dernoncourt, Quan Hung Tran, Jiuxiang Gu, Ani Nenkova, Vlad I. Morariu, Rajiv Jain, Dinesh Manocha
INTERSPEECH7
2022 DocTime: A Document-level Temporal Dependency Graph Parser
abstract
Puneet Mathur, Vlad Morariu, Verena Kaynig-Fittkau, Jiuxiang Gu, Franck Dernoncourt, Quan Tran, Ani Nenkova, Dinesh Manocha, Rajiv Jain. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Puneet Mathur, Vlad I. Morariu, Verena Kaynig, Jiuxiang Gu, Franck Dernoncourt, Quan Hung Tran, Ani Nenkova, Dinesh Manocha, Rajiv Jain
NAACL-HLT9
2021 Syntopical Graphs for Computational Argumentation Tasks
abstract
Joe Barrow, Rajiv Jain, Nedim Lipka, Franck Dernoncourt, Vlad Morariu, Varun Manjunatha, Douglas Oard, Philip Resnik, Henning Wachsmuth. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Joe Barrow, Rajiv Jain, Nedim Lipka, Franck Dernoncourt, Vlad I. Morariu, Varun Manjunatha, Douglas W. Oard, Philip Resnik, Henning Wachsmuth
ACL/IJCNLP (1)2
2021 SelfDoc: Self-Supervised Document Representation Learning
abstract
We propose SelfDoc, a task-agnostic pre-training framework for document image understanding. Because documents are multimodal and are intended for sequential reading, our framework exploits the positional, textual, and visual information of every semantically meaningful component in a document, and it models the contextualization between each block of content. Unlike existing document pre-training models, our model is coarse-grained instead of treating individual words as input, therefore avoiding an overly fine-grained with excessive contextualization. Beyond that, we introduce cross-modal learning in the model pre-training phase to fully leverage multimodal information from unlabeled documents. For downstream usage, we propose a novel modality-adaptive attention mechanism for multimodal feature fusion by adaptively emphasizing language and vision signals. Our framework benefits from self-supervised pre-training on documents without requiring annotations by a feature masking training strategy. It achieves superior performance on multiple downstream tasks with significantly fewer document images used in the pre-training stage compared to previous works.
Peizhao Li, Jiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao, Rajiv Jain, Varun Manjunatha, Hongfu Liu 0001
CVPR6
2021 Black-Box Explanation of Object Detectors via Saliency Maps
abstract
We propose D-RISE, a method for generating visual explanations for the predictions of object detectors. Utilizing the proposed similarity metric that accounts for both localization and categorization aspects of object detection allows our method to produce saliency maps that show image areas that most affect the prediction. D-RISE can be considered "black-box" in the software testing sense, as it only needs access to the inputs and outputs of an object detector. Compared to gradient-based methods, D-RISE is more general and agnostic to the particular type of object detector being tested, and does not need knowledge of the inner workings of the model. We show that D-RISE can be easily applied to different object detectors including one-stage detectors such as YOLOv3 and two-stage detectors such as Faster-RCNN. We present a detailed analysis of the generated visual explanations to highlight the utilization of context and possible biases learned by object detectors.
Vitali Petsiuk, Rajiv Jain, Varun Manjunatha, Vlad I. Morariu, Ashutosh Mehra 0002, Vicente Ordonez, Kate Saenko
CVPR2
2021 ClauseRec: A Clause Recommendation Framework for AI-aided Contract Authoring
abstract
Contracts are a common type of legal document that frequent in several day-to-day business workflows.However, there has been very limited NLP research in processing such documents, and even lesser in generating them.These contracts are made up of clauses, and the unique nature of these clauses calls for specific methods to understand and generate such documents.In this paper, we introduce the task of clause recommendation, as a first step to aid and accelerate the authoring of contract documents.We propose a twostaged pipeline to first predict if a specific clause type is relevant to be added in a contract, and then recommend the top clauses for the given type based on the contract context.We pretrain BERT on an existing library of clauses with two additional tasks and use it for our prediction and recommendation.We experiment with classification methods and similarity-based heuristics for clause relevance prediction, and generation-based methods for clause recommendation, and evaluate the results from various methods on several clause types.We provide analyses on the results, and further outline the advantages and limitations of the various methods for this line of research.
Vinay Aggarwal, Aparna Garimella, Balaji Vasan Srinivasan, Anandhavelu Natarajan, Rajiv Jain
EMNLP (1)5
2021 IGA: An Intent-Guided Authoring Assistant
abstract
Simeng Sun, Wenlong Zhao, Varun Manjunatha, Rajiv Jain, Vlad Morariu, Franck Dernoncourt, Balaji Vasan Srinivasan, Mohit Iyyer. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Simeng Sun, Wenlong Zhao 0001, Varun Manjunatha, Rajiv Jain, Vlad I. Morariu, Franck Dernoncourt, Balaji Vasan Srinivasan, Mohit Iyyer
EMNLP (1)4
2021 Towards Interpreting and Mitigating Shortcut Learning Behavior of NLU models
abstract
Mengnan Du, Varun Manjunatha, Rajiv Jain, Ruchi Deshpande, Franck Dernoncourt, Jiuxiang Gu, Tong Sun, Xia Hu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Mengnan Du, Varun Manjunatha, Rajiv Jain, Ruchi Deshpande, Franck Dernoncourt, Jiuxiang Gu, Tong Sun 0005, Xia Ben Hu
NAACL-HLT3
2021 UniDoc: Unified Pretraining Framework for Document Understanding
abstract
Document intelligence automates the extraction of information from documents and supports many business applications. Recent self-supervised learning methods on large-scale unlabeled document datasets have opened up promising directions towards reducing annotation efforts by training models with self-supervised objectives. However, most of the existing document pretraining methods are still language-dominated. We present UDoc, a new unified pretraining framework for document understanding. UDoc is designed to support most document understanding tasks, extending the Transformer to take multimodal embeddings as input. Each input element is composed of words and visual features from a semantic region of the input document image. An important feature of UDoc is that it learns a generic representation by making use of three self-supervised losses, encouraging the representation to model sentences, learn similarities, and align modalities. Extensive empirical analysis demonstrates that the pretraining procedure learns better joint representations and leads to improvements in downstream tasks.
Jiuxiang Gu, Jason Kuen, Vlad I. Morariu, Handong Zhao, Rajiv Jain, Nikolaos Barmpalios, Ani Nenkova, Tong Sun 0005
NeurIPS5
2021 Inducing Rich Interaction Structures Between Words for Document-Level Event Argument Extraction
Amir Pouran Ben Veyseh, Franck Dernoncourt, Quan Hung Tran, Varun Manjunatha, Rajiv Jain, Doo Soon Kim, Walter Chang, Thien Huu Nguyen
PAKDD (2)6
2020 A Joint Model for Document Segmentation and Segment Labeling
abstract
Text segmentation aims to uncover latent structure by dividing text from a document into coherent sections.Where previous work on text segmentation considers the tasks of document segmentation and segment labeling separately, we show that the tasks contain complementary information and are best addressed jointly.We introduce the Segment Pooling LSTM (S-LSTM) model, which is capable of jointly segmenting a document and labeling segments.In support of joint training, we develop a method for teaching the model to recover from errors by aligning the predicted and ground truth segments.We show that S-LSTM reduces segmentation error by 30% on average, while also improving segment labeling.
Joe Barrow, Rajiv Jain, Vlad I. Morariu, Varun Manjunatha, Douglas W. Oard, Philip Resnik
ACL2
2020 Text and Style Conditioned GAN for the Generation of Offline-Handwriting Lines
Brian L. Davis, Bryan S. Morse, Brian L. Price, Chris Tensmeyer, Curtis Wigington, Rajiv Jain
BMVC6
2020 Generative-Discriminative Feature Representations for Open-Set Recognition
abstract
We address the problem of open-set recognition, where the goal is to determine if a given sample belongs to one of the classes used for training a model (known classes). The main challenge in open-set recognition is to disentangle open-set samples that produce high class activations from known-set samples. We propose two techniques to force class activations of open-set samples to be low. First, we train a generative model for all known classes and then augment the input with the representation obtained from the generative model to learn a classifier. This network learns to associate high classification probabilities both when image content is from the correct class as well as when the input and the reconstructed image are consistent with each other. Second, we use self-supervision to force the network to learn more informative featues when assigning class scores to improve separation of classes from each other and from open-set samples. We evaluate the performance of the proposed method with recent open-set recognition works across three datasets, where we obtain state-of-the-art results.
Pramuditha Perera, Vlad I. Morariu, Rajiv Jain, Varun Manjunatha, Curtis Wigington, Vicente Ordonez, Vishal M. Patel
CVPR3
2019 Multimodal Document Image Classification
abstract
State-of-the-art methods for document image classification rely on visual features extracted by deep convolutional neural networks (CNNs). These methods do not utilize rich semantic information present in the text of the document, which can be extracted using Optical Character Recognition (OCR). We first study the performance of state-of-the-art text classification approaches when applied to noisy text obtained from OCR. We then show that fusing this textual information with visual CNN methods produces state-of-the-art results on the RVL-CDIP classification dataset.
Rajiv Jain, Curtis Wigington
ICDAR1
2015 Novel line verification for multiple instance focused retrieval in document collections
abstract
Spatial verification is typically employed to check the spatial consistency among matched local features and to remove outliers. However, when looking for multiple instances of the query within a target image, RANSAC algorithms which are widely applied in many one-to-one matching applications might fail due to the large proportion of “outliers” - correct matches corresponding to other instances. On the other hand, geometrical verification methods are more robust to outliers but usually suffer from high computational costs. In this paper, we introduce a novel two-step line verification method which is more flexible than existing methods and leads to lower computational complexity especially when multiple instances of a query are sought. We study this approach within an information extraction scenario, where the objective is to locate document structures indicative of certain type of information (e.g. different records on invoices).
Hongxing Gao, Marçal Rusiñol, Dimosthenis Karatzas, Josep Lladós 0001, Rajiv Jain, David S. Doermann
ICDAR5
2015 Localized document image change detection
abstract
Given two versions of a document image, the goal of document image change detection is to automatically determine exactly what content was added, deleted or modified. Typically, one would accomplish this by first performing Optical Character Recognition (OCR) on the two documents and then performing a “diff” to identify the changes. However, this approach can fail due to OCR errors, poor segmentation, or the inability to handle graphical content. We compare the OCR baseline with two techniques based on SIFT features that detect changes in the image at the word level. The first approach performs the “diff” on SIFT features extracted from the center line of the text image. The second approach performs a segmentation free alignment of text blocks using dense SIFT to address the more general cases where segmentation fails or graphical objects are modified. Results on two experimental datasets show the improvement of the segmentation free approach over the baseline approach.
Rajiv Jain, David S. Doermann
ICDAR1
2014 Combining Local Features for Offline Writer Identification
abstract
Several powerful approaches have recently been proposed for writer identification, which rely on local descriptors that capture the texture, shape and curvature properties of the handwriting. In this paper we use combinations of three of these features (K-Adjacent Segments, SURF, and Contour Gradient Descriptors), to address the writer identification problem. Experiments demonstrate that feature combinations outperform individual features, resulting in state-of-the-art performance on three datasets.
Rajiv Jain, David S. Doermann
ICFHR1
2013 VisualDiff: Document Image Verification and Change Detection
abstract
This paper explores the related problems of verification and change detection in document images. The goal is to determine if two document images differ, and if so, to determine precisely what content may have been added, deleted, or otherwise modified. This problem has many potential applications, especially for important legal documents such as contractual agreements. These agreements are often edited, shared and stored as scanned or hardcopy documents, where small, undetected changes between edits could create major differences in the contractual language and thus have severe repercussions. One can view the problem of change detection as tracing the revision history of a set of documents. Thus, in order to validate the performance of this approach, we created the "Enron Revisions" dataset. This dataset contains realistic revisions obtained from attachments in the Enron Corpus, and a series of before and after snapshots of the revisions in images with varying levels of noise from resolution, binarization, and blur. The approach taken in this paper utilizes the SIFT descriptor to align two document images without the benefit of OCR and once aligned, to compare dense descriptors to determine changes that have occurred within the image. As a baseline, this "VisualDiff" is compared to a UNIX diff-like approach on text extracted through OCR and results demonstrate the effectiveness of this approach.
Rajiv Jain, David S. Doermann
ICDAR1
2013 Writer Identification Using an Alphabet of Contour Gradient Descriptors
abstract
This paper presents a new method for writer identification, which emulates the approach taken by forensic document examiners. It combines a novel feature, which uses contour gradients to capture local shape and curvature, with character segmentation to create a pseudo-alphabet for a given handwriting sample. A distance metric is then defined between elements of these alphabets that captures character similarity between two handwriting samples. This approach achieves a Top-1 identification rate of 96.5% on the benchmark IAM dataset, reducing the error rate of previous approaches by 50%.
Rajiv Jain, David S. Doermann
ICDAR1
2012 Logo Retrieval in Document Images
abstract
This paper presents a scalable algorithm for segmentation free logo retrieval in document images. The contributions include the use of the SURF feature for logo retrieval, a novel indexing algorithm for efficient retrieval and a method to filter results using the orientation of local features and geometric constraints. Results demonstrate that logo retrieval can be performed with high accuracy and efficiently scaled to a large datasets.
Rajiv Jain, David S. Doermann
Document Analysis Systems1
2011 Offline Writer Identification Using K-Adjacent Segments
abstract
This paper presents a method for performing offline writer identification by using K-adjacent segment (KAS) features in a bag-of-features framework to model a user's handwriting. This approach achieves a top 1 recognition rate of 93% on the benchmark IAM English handwriting dataset, which outperforms current state of the art features. Results further demonstrate that identification performance improves as the number of training samples increase, and additionally, that the performance of the KAS features extend to Arabic handwriting found in the MADCAT dataset.
Rajiv Jain, David S. Doermann
ICDAR1
1996 Incorporating performance and testability constraints during binding in high-level synthesis
abstract
Module and register binding during high-level synthesis is one of the most important steps in generating an RTL design from a behavioral description. The binding phase determines the structure of the final design, and hence issues related to area, performance and testability of the RTL design have to be addressed in this step. In this paper, we present algorithms for module and register binding which generate RTL designs having high performance and/or high testability. The binding problem is decomposed into a sequence of subproblems, each of which is modeled as a minimum-cost network flow problem. The relative impact of the possible bindings is expressed in terms of the costs associated with the edges of the network. The model is simple and can be solved quickly to obtain a low cost flow solution. Putting together the solutions to the subproblems gives low cost bindings. We also propose cost functions that can be used with varying emphasis on delay and testing. The results demonstrate the effectiveness of our algorithm; the final designs produced by the algorithms require a smaller clock cycle or are easier to test as compared to designs generated without the performance or testability constraints.
Ashutosh Mujumdar, Rajiv Jain, Kewal K. Saluja
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1996 Valid Transformations: A New Class of Loop Transformations for High-Level Synthesis and Pipelined Scheduling Applications
abstract
In this paper we present a new class of loop optimizing transformations called valid transformations, which are suitable for fine-grain parallelization applications such as high-level synthesis of VLSI designs or compilers for super-scalar or VLIW machines. This class of transformations are different from existing ones in that valid transformations can be illegal. Nevertheless, if a transformation is valid, the transformed loop has a feasible pipeline schedule. We present an example valid transformation called loop expansion which can help produce cost-performance efficient designs and explore a larger design space for a satisfactory design. Several examples are used to demonstrate the efficacy of the proposed technique.
Minjoong Rim, Rajiv Jain
IEEE Trans. Parallel Distributed Syst.2
1995 Test application time reduction for scan based sequential circuits
abstract
This paper addresses the issue of reducing test application time in sequential circuits with partial scan using a single clock configuration without freezing the state of the non-scan flip-flops. Experimental results show that this technique significantly reduces test application time. Further, we study the effect of ordering the scan flip-flops on the test vector length and also present a non-atomic two-clock scan method which can be easily incorporated in conventional test generation environment.
Kewal K. Saluja, Rajiv Jain
Great Lakes Symposium on VLSI3
1995 PARAS: system-level concurrent partitioning and scheduling
abstract
Partitioning for the ASIC designs is examined and the interaction between high-level synthesis and partitioning is studied and incorporated in the solution. Four algorithms (called PARAS) which can exploit this interaction by solving the scheduling and partitioning problems concurrently are presented. PARAS maximizes the overall performance of the final design and considers different chip configurations and communication structures. Experiments, conducted with specifications ranging in size from few to hundreds of operations, demonstrate the success of this approach.
Wing Hang Wong, Rajiv Jain
ICCAD2
1995 Global scheduling with code-motions for high-level synthesis applications
abstract
In this paper, we present a global scheduling technique for synthesis applications. The algorithm accepts a specification containing conditional branches and while-loop constructs and schedules it for a given set of resources. The algorithm performs several types of code motions across different basic blocks and trades off cost with performance. Several real-life examples taken from Numerical Recipes in C are used to demonstrate the efficacy of the approach. The results indicate that code-motions are very important for achieving significant speed-ups for synthesis applications.>
Minjoong Rim, Yaw Fann, Rajiv Jain
IEEE Trans. Very Large Scale Integr. Syst.3
1994 Global Scheduling for High-Level Synthesis Applications
abstract
In this paper, we present a resource-constrained global scheduling technique for synthesis applications. The algorithm accepts speci cations containing conditional branches and while loops and schedules them for a given set of resources. The algorithm performs several types of code motions across di erent basic blocks and trades o cost with performance. Several real-life examples are used to demonstrate the e cacy of the approach. 1
Yaw Fann, Minjoong Rim, Rajiv Jain
DAC3
1994 RECALS II: a new list scheduling algorithm
abstract
Presents a new scheduling heuristic RECALS II for high-level synthesis applications. RECALS II accepts a directed acyclic graph and a set of resources and schedules the graph while minimizing the number of clock cycles required to execute it. Experiments show that schedules produced by RECALS II are close to the optimal solutions.>
Minjoong Rim, Rajiv Jain
ICASSP (2)2
1994 Register Estimation from Behavioral Specifications
abstract
Provides answers to the following problems: (1) Given a data flow graph and a performance constraint, determine a lower-bound on the storage area required for executing the data flow graph while satisfying the performance constraint. (2) Determine a lower-bound on performance for executing a data flow graph under fixed storage area constraints. The results demonstrate that our approach produces solutions which are very close to the optimal.>
Alok Sharma, Rajiv Jain
ICCD2
1994 Valid Transformations: A New Class of Loop Transformations
abstract
In this paper we present a new class of loop optimizing transformations called valid transformations. This class of transformations are different from existing ones in that valid transformations can be illegal and can result in incorrect non-pipelined designs. Nevertheless, valid transformations have feasible pipeline schedules which is important for scheduling loops. We present an example valid transformation called loop expansion which can help explore a larger design space and helps in producing cost-performance efficient designs. Several examples are used to demonstrate the efficacy of the proposed transformations.
Minjoong Rim, Rajiv Jain
ICPP (2)2
1994 Estimating Performance Characteristics of Loop Transformations
abstract
In this paper we present estimation techniques for loop transformations. These estimation techniques can be used for design space exploration. They can be used to quickly evaluate the cost and performance characteristics of the transformations, thus helping the designer make the decision of applying the transformation or not. The estimates are verified for several transformations and are very close to the actual results.>
Minjoong Rim, Rajiv Jain
ISCAS2
1994 Incorporating testability considerations in high-level synthesis
Ashutosh Mujumdar, Rajiv Jain, Kewal K. Saluja
J. Electron. Test.2
1994 Lower-bound performance estimation for the high-level synthesis scheduling problem
abstract
A given behavioral specification can be implemented on a large number of register-transfer level designs. Instead of producing several designs and selecting the best one, synthesis systems may use estimation to reduce the design space. In this paper, we present a new technique for computing a lower-bound completion time for non-pipelined resource-constrained scheduling problem. Given a data flow graph, a set of resources, resource delays and a clock cycle, we derive a lower-bound on the completion time of a schedule. Our technique can handle chaining, multi-cycle operations and pipelined modules. The technique is very fast and experimental results show that it is also very tight.>
Minjoong Rim, Rajiv Jain
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
1994 Optimal and heuristic algorithms for solving the binding problem
abstract
In this paper we present an optimal and a heuristic approach to solve the binding problem which occurs in high-level synthesis of digital systems. The optimal approach is based on an integer linear programming formulation. Given that such an approach is not practical for large problems, we then derive a heuristic from the ILP formulation which produces very good solutions in order of seconds. The heuristic is based on a network flow model and also considers floorplanning during the design process to minimize the interconnection area.>
Minjoong Rim, Ashutosh Mujumdar, Rajiv Jain, Renato De Leone
IEEE Trans. Very Large Scale Integr. Syst.3
1993 InSyn: Integrated Scheduling for DSP Applications
abstract
In this paper, we present the l~SV~, an integrated allocation and schedulingapproachfor high-levelsynthesisapplications.The scheduler considers functional units, busses and registers while performing time-step assignment.The results show that incorporating all these features during scheduling can produce very good designs.
Alok Sharma, Rajiv Jain
DAC2
1993 Estimating Architectural Resources and Performance for High-Level Synthesis Applications
abstract
In this paper we present a solution to the following problems related to architectural synthesis. Given an input specification and a perfomance constraint, determine a lower bound number of resources (active and interconnect) required to execute the data flow graph while satisfying the performance constraint. Conversely, determine a lower bound performance for executing an input specification for a given number of resources (active and interconnect). The generated bounds are close to the actual designs synthesized by several existing systems.
Alok Sharma, Rajiv Jain
DAC2
1993 Estimating architectural resources and performance for high-level synthesis applications
abstract
The authors present a solution to the following problems related to architectural synthesis. (1) Given an input specification and a performance constraint, determine a lower bound number of resources (active and interconnect) required to execute the data flow graph while satisfying the performance constraint. (2) Determine a lower bound performance for executing an input specification for a given number of resources (active and interconnect). These bounds are close to the actual designs synthesized by several existing systems.>
Alok Sharma, Rajiv Jain
IEEE Trans. Very Large Scale Integr. Syst.2
1992 Representing Conditional Branches for High-Level Synthesis Applications
Minjoong Rim, Rajiv Jain
DAC2
1992 Optimal Allocation and Binding in High-Level Synthesis
Minjoong Rim, Rajiv Jain, Renato De Leone
DAC2
1992 Estimating Lower-Bound Performance of Schedules Using a Relaxation Technique
abstract
A technique for computing a lower bound for nonpipelined resource-constrained scheduling performance is presented. Given a data-flow graph, a set of resources, resource delays, and clock cycle, a lower bound on the performance of a schedule is derived. The technique is fast (typical runtime of 50 ms), and the experimental lower-bounds are within two time steps of the actual schedules produced by a scheduling heuristic. The technique is also constructive and can be incorporated in a branch-and-bound method for solving the scheduling problem. The lower bound can be used to reduce a design search space. The method is applicable to lower-bound estimation in high-level synthesis.>
Minjoong Rim, Rajiv Jain
ICCD2
1992 Predicting system-level area and delay for pipelined and nonpipelined designs
abstract
The ability to predict area-delay characteristics of designs without actually implementing them is important in producing quality designs in a reasonable time. A mathematical model for predicting the area-delay tradeoff curve for pipelined and nonpipelined data paths, given a data flow graph and a choice of module styles, is proposed. The model has been validated against designs generated by pipelined and nonpipelined data-path synthesis programs.>
Rajiv Jain, Alice C. Parker, Nohbyung Park
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
1991 Empirical Evaluation of Some High-Level Synthesis Scheduling Heuristics
abstract
Over the past few years a large number of heuristics for performing scheduling
Rajiv Jain, Ashutosh Mujumdar, Alok Sharma, Hueymin Wang
DAC1
1990 MOSP: Module Selection for Pipelined Designs with Multi-Cycle Operations
abstract
Selection of appropriate module types from a design library which will be used in the final implementation is called module selection. A solution is presented to the module selection problem for pipelined designs with multicycle operations. The proposed solution technique is based on an area-delay analysis of an RTL design and produces optimal results in milliseconds.>
Rajiv Jain
ICCAD1
1989 Experience with ADAM Synthesis System
abstract
The ADAM synthesis system consists of two major subsystems: the program tools which synthesize RTL designs from behavioral descriptions and the prediction tools which guide the designer in exploring the design space for a good design. In this paper, we demonstrate the necessity for predictions in narrowing the search space. With the aid of an example, we describe the interaction of a designer with the two subsystems in designing an RTL implementation which maximizes performance while meeting a given area constraint.
Rajiv Jain, Kayhan Küçükçakar, Mitch J. Mlinar, Alice C. Parker
DAC1
1988 Module Selection for Pipelined Synthesis
Rajiv Jain, Alice C. Parker, Nohbyung Park
DAC1
1988 Area-time model for synthesis of non-pipelined designs
abstract
A mathematical model is presented for predicting the area-time tradeoff curve for nonpipelined data paths given a data-flow graph and a module set. Specifically, it examines operator cost and delay to predict the lower bound noninferior area-time curve. The model has been validated against designs generated by a program which synthesizes nonpipelined data paths.>
Rajiv Jain, Mitch J. Mlinar, Alice C. Parker
ICCAD1
1988 The POTATO chip architecture: a study in tradeoffs for signal processing chip design
abstract
The authors describe an example signal-processing design which illustrates partitioning, performance, cost, and fault-tolerance tradeoffs. They focus on high-performance multiplication using the power-of-two number representation as implemented in the POTATO (power of two arithmetic time-optimized) chip architecture. The implementation is compared to more conventional designs, and performance estimates are given. It is concluded that the design compares favourably to more conventional implementations.>
B. Sharma, Rajiv Jain, Melvin A. Breuer, Alice C. Parker, Cauligi S. Raghavendra, C. Y. Tseng
ICCD2
1988 On Array Storage for Conflict-Free Memory Access for Parallel Processors
Meera Balakrishnan, Rajiv Jain, Cauligi S. Raghavendra
ICPP (1)2
1987 Predicting Area-Time Tradeoffs for Pipelined Design
abstract
In this paper we give a model for predicting the shape of cost-speed tradeoff curves for pipelined designs. The model includes prediction of the number of operators, registers and multiplexers from a behavioral specification. It has been verified with the designs generated by an automated pipeline synthesis program, Sehwa. This model was developed as a part of the ADAM Advanced Design Automation System of the University of Southern California.
Rajiv Jain, Alice C. Parker, Nohbyung Park
DAC1