VLDB 2026 Research / reviewers in the wild / expert
Srirangaraj Setlur
dblp:80/1388
· DBLP profile ↗
68ranked-venue papers
3as first author
28since 2021 · last 2026
0000-0002-7118-9280ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 3 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 19 since 2021Databases, data management, data science and information retrieval · 23 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 6 · 5 since 2021Security and privacy · 5 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Teaching Spell Checkers to Teach: Pedagogical Program Synthesis for Interactive LearningabstractSpelling taught through memorization often fails many learners, particularly children with language-based learning disorders who struggle with the phonological skills necessary to spell words accurately. Educators such as speech-language pathologists (SLPs) address this instructional gap by using an inquiry-based approach to teach spelling that targets the phonology, morphology, meaning, and etymology of words. Yet, these strategies rarely appear in everyday writing tools, which simply detect and autocorrect errors. We introduce SPIRE (Spelling Inquiry Engine), a spell check system that brings this inquiry-based pedagogy into the act of composition. SPIRE implements Pedagogical Program Synthesis, a novel approach for operationalizing the inherently dynamic pedagogy of spelling instruction. SPIRE represents SLP instructional moves in a domain-specific language, synthesizes tailored programs in real-time from learner errors, and renders them as interactive interfaces for inquiry-based interventions. With SPIRE, spelling errors become opportunities to explore word meanings, word structures, morphological families, word origins, and grapheme-phoneme correspondences, supporting metalinguistic reasoning alongside correction. Evaluation with SLPs and learners shows alignment with professional practice and potential for integration into writing workflows. Momin Naushad Siddiqui, Vincent Cavez, Sahana Rangasrinivasan, Abbie Olszewski, Srirangaraj Setlur, Maneesh Agrawala, Hari Subramonyam |
IUI | 5 |
| 2026 | AF-CANet: Augmentation-free, curriculum-guided attention UNet framework for few-shot document layout analysis
Hadia Showkat Kawoosa, Sahana Rangasrinivasan, Srirangaraj Setlur, Puneet Goyal, Venu Govindaraju |
Pattern Recognit. Lett. | 3 |
| 2025 | From Scribbles to Text: A Novel Transformer-Based Recognition Model for Child Handwriting
Sahana Rangasrinivasan, M. S. Sumi Suresh, Srirangaraj Setlur, Bharat Jayaraman, Venu Govindaraju |
ICDAR (1) | 3 |
| 2025 | Ridgeformer: Mutli-Stage Contrastive Training for Fine-Grained Cross-Domain Fingerprint RecognitionabstractThe increasing demand for hygienic and portable biometric systems has underscored the critical need for advancements in contactless fingerprint recognition. Despite its potential, this technology faces notable challenges, including out-of-focus image acquisition, reduced contrast between fingerprint ridges and valleys, variations in finger positioning, and perspective distortion. These factors significantly hinder the accuracy and reliability of contactless fingerprint matching. To address these issues, we propose a novel multi-stage transformer-based contactless fingerprint matching approach that first captures global spatial features and subsequently refines localized feature alignment across fingerprint samples. By employing a hierarchical feature extraction and matching pipeline, our method ensures fine-grained, cross-sample alignment while maintaining the robustness of global feature representation. We perform extensive evaluations on publicly available datasets such as HKPolyU and RidgeBase under different evaluation protocols, such as contactless-to-contact matching and contactless-to-contactless matching and demonstrate that our proposed approach outperforms existing methods, including COTS solutions. Our codebase is available at https://github.com/KNITPhoenix/Ridgeformer Shubham Pandey, Bhavin Jawade, Srirangaraj Setlur |
ICIP | 3 |
| 2025 | AutoMisty: A Multi-Agent LLM Framework for Automated Code Generation in the Misty Social RobotabstractThe social robot’s open API allows users to customize open-domain interactions. However, it remains inaccessible to those without programming experience. We introduce AutoMisty, the first LLM-powered multi-agent framework that converts natural-language commands into executable Misty robot code by decomposing high-level instructions, generating sub-task code, and integrating everything into a deployable program. Each agent employs a two-layer optimization mechanism: first, a self-reflective loop that instantly validates and automatically executes the generated code, regenerating whenever errors emerge; second, human review for refinement and final approval, ensuring alignment with user preferences and preventing error propagation. To evaluate AutoMisty’s effectiveness, we designed a benchmark task set spanning four levels of complexity and conducted experiments in a real Misty robot environment. Extensive evaluations demonstrate that AutoMisty not only consistently generates high-quality code but also enables precise code control, significantly outperforming direct reasoning with ChatGPT-4o and ChatGPT-o1. All code, optimized APIs, and experimental videos will be publicly released through the webpage: AutoMisty. Lu Dong 0004, Sahana Rangasrinivasan, Ifeoma Nwogu, Srirangaraj Setlur, Venu Govindaraju |
IROS | 5 |
| 2025 | SCOT: Self-Supervised Contrastive Pretraining for Zero-Shot Compositional RetrievalabstractCompositional image retrieval (CIR) is a multimodal learning task where a model combines a query image with a user-provided text modification to retrieve a target image. CIR finds applications in a variety of domains including product retrieval (e-commerce) and web search. Existing methods primarily focus on fully-supervised learning, wherein models are trained on datasets of labeled triplets such as FashionIQ and CIRR. This poses two significant challenges: (i) curating such triplet datasets is labor intensive; and (ii) models lack generalization to un-seen objects and domains. In this work, we propose SCOT (Self-supervised COmpositional Training), a novel zero-shot compositional pretraining strategy that combines existing large image-text pair datasets with the generative capabilities of large language models to contrastively train an embedding composition network. Specifically, we show that the text embedding from a large-scale contrastively-pretrained vision-language model can be utilized as proxy target supervision during compositional pretraining, replacing the target image embedding. In zero-shot settings, this strategy surpasses SOTA zero-shot compositional re-trieval methods as well as many fully-supervised methods on standard benchmarks such as FashionIQ and CIRR. Our code and models are available at https://github.com/yahoo/SCOT. Bhavin Jawade, João V. B. Soares, Kapil Thadani, Deen Dayal Mohan, Amir Erfan Eshratifar, Benjamin Culpepper, Paloma de Juan, Srirangaraj Setlur, Venu Govindaraju |
WACV | 8 |
| 2024 | GestSpoof: Gesture Based Spatio-Temporal Representation Learning for Robust Fingerprint Presentation Attack DetectionabstractFingerprint spoof attacks represent one of the most prevalent forms of biometric presentation attacks. While significant progress has been made in framing fingerprint spoof detection as a general image classification problem, limited attention has been given to treating it as a temporal learning problem. The distinctions in the elastic properties between authentic and synthetically created counterfeit fingerprints can be more accurately captured under motion-induced gestures during acquisition. In this study, we introduce a novel method for detecting fake fingerprints by deliberately introducing distortions through sliding and twisting motions during acquisition. As widely used spoof datasets such as those from LivDet 2009 to 2021 or MSU FPAD lack the temporal information essential for this investigation, we assembled a new dataset focused on distortion-based fake and real fingerprints, encompassing various types of spoof materials and diverse distortions. This gesture-equipped dataset comprises more than 3680 videos gathered from 184 unique fingers. Additionally, we present a novel spatial-temporal multi-modal network for detecting fingerprint spoofs using intentional-distortion. Our proposed approach yields significantly improved results compared to traditional static classification-based methods for spoof detection, across various metrics and for both known and unknown (generalization) scenarios, thereby highlighting the substantial impact that introducing gestures can have on enhancing fingerprint spoof detection. The dataset can be downloaded from here: https://www.buffalo.edu/cubs/research/datasets/gestspoof-dataset.html Bhavin Jawade, Shreeram Subramanya, Atharv Dabhade, Srirangaraj Setlur, Venu Govindaraju |
FG | 4 |
| 2024 | A Comparative Study of Video-Based Human Representations for American Sign Language Alphabet GenerationabstractSign language is a complex visual language, and automatic interpretations of sign language can facilitate communication involving deaf individuals. As one of the essential components of sign language, fingerspelling connects the natural spoken languages to the sign language and expands the scale of sign language vocabulary. In practice, it is challenging to analyze fingerspelling alphabets due to their signing speed and small motion range. The usage of synthetic data has the potential of further improving fingerspelling alphabets analysis at scale. In this paper, we evaluate how different video-based human representations perform in a framework for Alphabet Generation for American Sign Language (ASL). We tested three mainstream video-based human representations: two-stream inflated 3D ConvNet, 3D landmarks of body joints, and rotation matrices of body joints. We also evaluated the effect of different skeleton graphs and selected body joints. The generation process of ASL fingerspelling used a transformer-based Conditional Variational Autoencoder. To train the model, we collected ASL alphabet signing videos from 17 signers with dynamic alphabet signing. The generated alphabets were evaluated using automatic metrics of quality such as FID, and we also considered supervised metrics by recognizing the generated entries using Spatio-Temporal Graph Convolutional Networks. Our experiments show that using the rotation matrices of the upper body joints and the signing hand give the best results for the generation of ASL alphabet signing. Going forward, our goal is to produce articulated fingerspelling words by combining individual alphabets learned in this work. Lipisha Chaudhary, Lu Dong 0004, Srirangaraj Setlur, Venu Govindaraju, Ifeoma Nwogu |
FG | 4 |
| 2024 | Fine-Grained Engine Fault Sound Event Detection Using Multimodal SignalsabstractSound event detection (SED) is an active area of audio research that aims to detect the temporal occurrence of sounds. In this paper, we apply SED to engine fault detection by introducing a multimodal SED framework that detects fine-grained engine faults of automobile engines using audio and accelerometer-recorded vibration. We first introduce the problem of engine fault SED on a dataset collected from a large variety of vehicles with expertly-labeled engine fault sound events. Next, we propose a SED model to temporally detect ten fine-grained engine faults that occur within vehicle engines and further explore a pretraining strategy using a large-scale weakly-labeled engine fault dataset. Through multiple evaluations, we show our proposed framework is able to effectively detect engine fault sound events. Finally, we investigate the interaction and characteristics of each modality and show that fusing features from audio and vibration improves overall engine fault SED capabilities. Dennis Fedorishin, Livio Forte, Philip Schneider, Srirangaraj Setlur, Venu Govindaraju |
ICASSP | 4 |
| 2024 | Audio Match Cutting: Finding and Creating Matching Audio Transitions in Movies and VideosabstractA "match cut" is a common video editing technique where a pair of shots that have a similar composition transition fluidly from one to another. Although match cuts are often visual, certain match cuts involve the fluid transition of audio, where sounds from different sources merge into one indistinguishable transition between two shots. In this paper, we explore the ability to automatically find and create "audio match cuts" within videos and movies. We create a self-supervised audio representation for audio match cutting and develop a coarse-to-fine audio match pipeline that recommends matching shots and creates the blended audio. We further annotate a dataset for the proposed audio match cut task and compare the ability of multiple audio representations to find audio match cut candidates. Finally, we evaluate multiple methods to blend two matching audio candidates with the goal of creating a smooth transition. Project page and examples are available at: https://denfed.github.io/audiomatchcut/ Dennis Fedorishin, Lie Lu, Srirangaraj Setlur, Venu Govindaraju |
ICASSP | 3 |
| 2024 | DIOR: Dataset for Indoor-Outdoor Reidentification Long Range 3D/2D Skeleton Gait Collection Pipeline, Semi-Automated Gait Keypoint Labeling and Baseline Evaluation MethodsabstractRecently, there has been growing interest in the identification and re-identification of individuals from long distances using rooftop cameras, UAV cameras, street cams, and similar devices. This type of recognition extends beyond facial recognition, utilizing whole-body markers such as gait. However, datasets to train and test such recognition algorithms are scarce and often lack labeling. This paper introduces DIOR—a comprehensive framework for data collection, semi-automated annotation, and a dataset comprising 1.649 million RGB frames labeled with 3D/2D skeleton gait markers across 14 subjects. This dataset includes 200,000 RGB frames captured from long-range cam-eras. Our approach employs advanced 3D computer vision techniques to achieve pixel-level accuracy in indoor environments using motion capture systems. For outdoor, long-range environments, we eliminate the reliance on motion capture systems and implement a cost-effective, hybrid 3D computer vision and learning pipeline using only four inexpensive RGB cameras. This method successfully achieves precise skeleton labeling of distant subjects, even when their visual size is as small as 20-25 pixels within an RGB frame. We benchmark models trained on existing datasets such as CASIA-B, on our proposed dataset for the task of Gait recognition. Our pipeline and the accompanying dataset will be made publicly available following acceptance. Praveen Raj Masilamani, Bhavin Jawade, Srirangaraj Setlur, Karthik Dantu |
IJCB | 4 |
| 2024 | CHART-Info 2024: A Dataset for Chart Analysis and Recognition
Kenny Davila, Rupak Lazarus, Nicole Rodríguez Alcántara, Srirangaraj Setlur, Venu Govindaraju, Ajoy Mondal, C. V. Jawahar |
ICPR (19) | 5 |
| 2024 | Maximizing Coverage over a Surveillance Region Using a Specific Number of Cameras
M. S. Sumi Suresh, Vivek Menon, Srirangaraj Setlur, Venu Govindaraju |
ICPR (22) | 3 |
| 2024 | ProxyFusion: Face Feature Aggregation Through Sparse ExpertsabstractFace feature fusion is indispensable for robust face recognition, particularly in scenarios involving long-range, low-resolution media (unconstrained environments) where not all frames or features are equally informative. Existing methods often rely on large intermediate feature maps or face metadata information, making them incompatible with legacy biometric template databases that store pre-computed features. Additionally, real-time inference and generalization to large probe sets remains challenging.
To address these limitations, we introduce a linear time O(N) proxy based sparse expert selection and pooling approach for context driven feature-set attention. Our approach is order invariant on the feature-set, generalizes to large sets, is compatible with legacy template stores, and utilizes significantly less parameters making it suitable real-time inference and edge use-cases. Through qualitative experiments, we demonstrate that ProxyFusion learns discriminative information for importance weighting of face features without relying on intermediate features. Quantitative evaluations on challenging low-resolution face verification datasets such as IARPA BTS3.1 and DroneSURF show the superiority of ProxyFusion in unconstrained long-range face recognition setting.
Our code and pretrained models are available at: https://github.com/bhavinjawade/ProxyFusion Bhavin Jawade, Alexander Stone, Deen Dayal Mohan, Srirangaraj Setlur, Venu Govindaraju |
NeurIPS | 5 |
| 2023 | CoNAN: Conditional Neural Aggregation Network For Unconstrained Face Feature FusionabstractFace recognition from image sets acquired under unregulated and uncontrolled settings, such as at large distances, low resolutions, varying viewpoints, illumination, pose, and atmospheric conditions, is challenging. Face feature aggregation, which involves aggregating a set of N feature representations present in a template into a single global representation, plays a pivotal role in such recognition systems. Existing works in traditional face feature aggregation either utilize metadata or high-dimensional intermediate feature representations to estimate feature quality for aggregation. However, generating high-quality metadata or style information is not feasible for extremely low-resolution faces captured in long-range and high altitude settings. To overcome these limitations, we propose a feature distribution conditioning approach called CoNAN for template aggregation. Specifically, our method aims to learn a context vector conditioned over the distribution information of the incoming feature set, which is utilized to weigh the features based on their estimated informativeness. The proposed method produces state-of-the-art results on long-range unconstrained face recognition datasets such as BTS, and DroneSURF, validating the advantages of such an aggregation strategy. Bhavin Jawade, Deen Dayal Mohan, Dennis Fedorishin, Srirangaraj Setlur, Venu Govindaraju |
IJCB | 4 |
| 2023 | Liveness Detection Competition - Noncontact-based Fingerprint Algorithms and Systems (LivDet-2023 Noncontact Fingerprint)abstractLiveness Detection (LivDet) is an international competition series open to academia and industry with the objective to assess and report state-of-the-art in Presentation Attack Detection (PAD). LivDet-2023 Noncontact Fingerprint is the first edition of the noncontact fingerprint-based PAD competition for algorithms and systems. The competition serves as an important benchmark in noncontact-based fingerprint PAD, offering (a) independent assessment of the state-of-the-art in noncontact-based fingerprint PAD for algorithms and systems, and (b) common evaluation protocol, which includes finger photos of a variety of Presentation Attack Instruments (PAIs) and live fingers to the biometric research community (c) provides standard algorithm and system evaluation protocols, along with the comparative analysis of state-of-the-art algorithms from academia and industry with both old and new android smartphones. The winning algorithm achieved an APCER of 11.35% averaged over all PAIs and a BPCER of 0.62%. The winning system achieved an APCER of 13.0.4%, averaged over all PAIs tested over all the smartphones, and a BPCER of 1.68% over all smartphones tested. Four-finger systems that make individual finger-based PAD decisions were also tested. The dataset used for competition will be available1, to all researchers as per data share protocol.1https://noncontactfingerprint2023.1ivdet.org/index.php Sandip Purnapatra, Humaira Rezaie, Bhavin Jawade, Yu Liu 0069, Luke Brosell, Mst Rumana Sumi, Lambert Igene, Alden Dimarco, Srirangaraj Setlur, Soumyabrata Dey, Stephanie Schuckers, Marco Huber, Jan Niklas Kolf, Meiling Fang, Naser Damer, Banafsheh Adami, Raul Chitic, Karsten Seelert, Vishesh Mistry, Rahul Parthe, Umit Kacar |
IJCB | 10 |
| 2023 | RealCQA: Scientific Chart Question Answering as a Test-Bed for First-Order Logic
Saleem Ahmed, Bhavin Jawade, Shubham Pandey, Srirangaraj Setlur, Venu Govindaraju |
ICDAR (3) | 4 |
| 2023 | SpaDen: Sparse and Dense Keypoint Estimation for Real-World Chart Understanding
Saleem Ahmed, Pengyu Yan, David S. Doermann, Srirangaraj Setlur, Venu Govindaraju |
ICDAR (2) | 4 |
| 2023 | Hear The Flow: Optical Flow-Based Self-Supervised Visual Sound Source LocalizationabstractLearning to localize the sound source in videos without explicit annotations is a novel area of audio-visual research. Existing work in this area focuses on creating attention maps to capture the correlation between the two modalities to localize the source of the sound. In a video, oftentimes, the objects exhibiting movement are the ones generating the sound. In this work, we capture this characteristic by modeling the optical flow in a video as a prior to better aid in localizing the sound source. We further demonstrate that the addition of flow-based attention substantially improves visual sound source localization. Finally, we benchmark our method on standard sound source localization datasets and achieve state-of-the-art performance on the Soundnet Flickr and VGG Sound Source datasets. Code: https://github.com/denfed/heartheflow. Dennis Fedorishin, Deen Dayal Mohan, Bhavin Jawade, Srirangaraj Setlur, Venu Govindaraju |
WACV | 4 |
| 2023 | NAPReg: Nouns As Proxies Regularization for Semantically Aware Cross-Modal EmbeddingsabstractCross-modal retrieval is a fundamental vision-language task with a broad range of practical applications. Text-to-image matching is the most common form of cross-modal retrieval where, given a large database of images and a textual query, the task is to retrieve the most relevant set of images. Existing methods utilize dual encoders with an attention mechanism and a ranking loss for learning embeddings that can be used for retrieval based on cosine similarity. Despite the fact that these methods attempt to perform semantic alignment across visual regions and textual words using tailored attention mechanisms, there is no explicit supervision from the training objective to enforce such alignment. To address this, we propose NAPReg, a novel regularization formulation that projects high-level semantic entities i.e Nouns into the embedding space as shared learnable proxies. We show that using such a formulation allows the attention mechanism to learn better word-region alignment while also utilizing region information from other samples to build a more generalized latent representation for semantic concepts. Experiments on three benchmark datasets i.e. MS-COCO, Flickr30k and Flickr8k demonstrate that our method achieves state-of-the-art results in cross-modal metric learning for text-image and image-text retrieval tasks. Code: https://github.com/bhavinjawade/NAPReq Bhavin Jawade, Deen Dayal Mohan, Naji Mohamed Ali, Srirangaraj Setlur, Venu Govindaraju |
WACV | 4 |
| 2022 | Attribute De-biased Vision Transformer (AD-ViT) for Long-Term Person Re-identificationabstractPerson re-identification (re-ID) aims to retrieve images of the same identity from a gallery of person images across cameras and viewpoints. However, most works in person re-ID assume a short-term setting characterized by invariance in appearance. In contrast, a high visual variance can be frequently seen in a long-term setting due to changes in apparel and accessories, which makes the task more challenging. Therefore, learning identity-specific features agnostic of temporally variant features is crucial for robust long-term person Re-ID. To this end, we propose an Attribute De-biased Vision Transformer (AD-ViT) to provide direct supervision to learn identity-specific features. Specifically, we produce attribute labels for person instances and utilize them to guide our model to focus on identity features through gradient reversal. Our experiments on two long-term re-ID datasets - LTCC and NKUP show that the proposed work consistently outperforms current state-of-the-art methods. Kyung Won Lee, Bhavin Jawade, Deen Dayal Mohan, Srirangaraj Setlur, Venu Govindaraju |
AVSS | 4 |
| 2022 | RidgeBase: A Cross-Sensor Multi-Finger Contactless Fingerprint DatasetabstractContactless fingerprint matching using smartphone cameras can alleviate major challenges of traditional fingerprint systems including hygienic acquisition, portability and presentation attacks. However, development of practical and robust contactless fingerprint matching techniques is constrained by the limited availablity of large scale real-world datasets. To motivate further advances in contactless fingerprint matching across sensors, we introduce the RidgeBase benchmark dataset. RidgeBase consists of more than 15,000 contactless and contact-based fingerprint image pairs acquired from 88 individuals under different background and lighting conditions using two smartphone cameras and one flatbed contact sensor. Unlike existing datasets, RidgeBase is designed to promote research under different matching scenarios that include Single Finger Matching and Multi-Finger Matching for both contactless-to-contactless (CL2CL) and contact-to-contactless (C2CL) verification and identification. Furthermore, due to the high intra-sample variance in contactless fingerprints belonging to the same finger, we propose a set-based matching protocol inspired by the advances in facial recognition datasets. This protocol is specifically designed for pragmatic contactless fingerprint matching that can account for variances in focus, polarity and finger-angles. We report qualitative and quantitative baseline results for different protocols using a COTS fingerprint matcher (Verifinger) and a Deep CNN based approach on the RidgeBase dataset. The dataset can be downloaded here: https://www.buffalo.edu/cubs/research/datasets/ridgebase-benchmark-dataset.html Bhavin Jawade, Deen Dayal Mohan, Srirangaraj Setlur, Nalini K. Ratha, Venu Govindaraju |
IJCB | 3 |
| 2022 | Synthetic Data Generation for Semantic Segmentation of Lecture Videos
Kenny Davila, James Molina, Srirangaraj Setlur, Venu Govindaraju |
ICFHR | 4 |
| 2022 | ICPR 2022: Challenge on Harvesting Raw Tables from Infographics (CHART-Infographics)abstractThe outcomes of the third Challenge on HArvesting Raw Tables from Infographics (ICPR 2022 CHART-Infographics) are presented in this work. Recognizing charts is a difficult process which we divided into the following task: Chart Image Classification (Task 1), Text Detection and Recognition (Task 2), Text Role Classification (Task 3), Axis Analysis (Task 4), Legend Analysis (Task 5), Plot Element Detection and Classification (Task 6.a), Data Extraction (Task 6.b), and End-to-End Data Extraction (Task 7). We have provided a novel dataset for training reusing all available data from previous challenges, and we also provide a brand new testing dataset for the evaluation of submissions. Both datasets were constructed by manually annotating charts extracted from the Open Access section of the PubMed Central. A total of 9 teams registered out of which 5 submitted results for different tasks of the challenge. Many submissions are based on state-of-the-art methods from computer vision, but the final scores imply that more work will be required to solve the chart recognition problem. The data, annotation tools, and evaluation scripts have been publicly released for academic use. Kenny Davila, Saleem Ahmed, David A. Mendoza, Srirangaraj Setlur, Venu Govindaraju |
ICPR | 5 |
| 2022 | Large-Scale Acoustic Automobile Fault Detection: Diagnosing Engines Through SoundabstractIn this paper we present AMPNet, an acoustic abnormality detection model deployed at ACV Auctions to automatically identify engine faults of vehicles listed on the ACV Auctions platform. We investigate the problem of engine fault detection and discuss our approach of deep-learning based audio classification on a large-scale automobile dataset collected at ACV Auctions. Specifically, we discuss our data collection pipeline and its challenges, dataset preprocessing and training procedures, and deployment of our trained models into a production setting. We perform empirical evaluations of AMPNet and demonstrate that our framework is able to successfully capture various engine anomalies agnostic of vehicle type. Finally we demonstrate the effectiveness and impact of AMPNet in the real world, specifically showing a 20.85% reduction in vehicle arbitrations on ACV Auctions' live auction platform. Dennis Fedorishin, Justas Birgiolas, Deen Dayal Mohan, Livio Forte, Philip Schneider, Srirangaraj Setlur, Venu Govindaraju |
KDD | 6 |
| 2021 | Bayesian Personalized-Wardrobe Model (BP-WM) for Long-Term Person Re-IdentificationabstractLong-term surveillance applications often involve having to re-identify individuals over several days. The task is made even more challenging due to changes in appearance features such as clothing over a longitudinal time-span of days or longer. In this paper, we propose a novel approach called Bayesian Personalized-Wardrobe Model (BPWM) for long-term person re-identification (re-ID) by employing a Bayesian Personalized Ranking (BPR) for clothing features extracted from video sequences. In contrast to previous long-term person re-ID works, we exploit the fact that people typically choose their attire based on their personal preferences and that knowing a person’s chosen wardrobe can be used as a soft-biometric to distinguish identities in the long-term. We evaluate the performance of our proposed BP-WM on the extended Indoor Long-term Re-identification Wardrobe (ILRW) dataset. Experimental results show that our method achieves state-of-the-art performance and that BP-WM can be used as a reliable soft-biometric for person re-identification. Kyung Won Lee, Nishant Sankaran, Deen Dayal Mohan, Kenny Davila, Dennis Fedorishin, Srirangaraj Setlur, Venu Govindaraju |
AVSS | 6 |
| 2021 | TADPool: Target Adaptive Pooling for Set Based Face RecognitionabstractA majority of the modern methods used for template aggregation of set-based face recognition systems rely on learning to quantify the quality of images present in a template. While focusing on weighting the feature embedding based on this quality factor, they have overlooked aggregation strategies that can adapt the template's features to the paired template involved in matching. In this paper, we explore the potential of such adaptive methods for feature aggregation. We propose a template feature aggregation strategy that tailors a template's image set to mirror the properties exhibited by the target template. The proposed method produces state-of-the-art results on standard unconstrained face recognition datasets such as IJB-A, IJB-C and YouTubeFaces, validating the advantages of such an aggregation strategy. Nishant Sankaran, Deen Dayal Mohan, Sergey Tulyakov, Srirangaraj Setlur, Venu Govindaraju |
FG | 4 |
| 2021 | Chart Mining: A Survey of Methods for Automated Chart AnalysisabstractCharts are useful communication tools for the presentation of data in a visually appealing format that facilitates comprehension. There have been many studies dedicated to chart mining, which refers to the process of automatic detection, extraction and analysis of charts to reproduce the tabular data that was originally used to create them. By allowing access to data which might not be available in other formats, chart mining facilitates the creation of many downstream applications. This paper presents a comprehensive survey of approaches across all components of the automated chart mining pipeline, such as (i) automated extraction of charts from documents; (ii) processing of multi-panel charts; (iii) automatic image classifiers to collect chart images at scale; (iv) automated extraction of data from each chart image, for popular chart types as well as selected specialized classes; (v) applications of chart mining; and (vi) datasets for training and evaluation, and the methods that were used to build them. Finally, we summarize the main trends found in the literature and provide pointers to areas for further research in chart mining. Kenny Davila, Srirangaraj Setlur, David S. Doermann, Bhargava Urala Kota, Venu Govindaraju |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Moving in the Right Direction: A Regularization for Deep Metric LearningabstractDeep metric learning leverages carefully designed sampling strategies and loss functions that aid in optimizing the generation of a discriminable embedding space. While effective sampling of pairs is critical for shaping the metric space during training, the relative interactions between pairs, and consequently the forces exerted on these pairs that direct their displacement in the embedding space can significantly impact the formation of well separated clusters. In this work, we identify a shortcoming of existing loss formulations which fail to consider more optimal directions of pair displacements as another criterion for optimization. We propose a novel direction regularization to explicitly account for the layout of sampled pairs and attempt to introduce orthogonality in the representations. The proposed regularization is easily integrated into existing loss functions providing considerable performance improvements. We experimentally validate our hypothesis on the Cars-196, CUB-200 and InShop datasets and outperform existing methods to yield state-of-the-art results on these datasets. Deen Dayal Mohan, Nishant Sankaran, Dennis Fedorishin, Srirangaraj Setlur, Venu Govindaraju |
CVPR | 4 |
| 2020 | Learning Guided Attention Masks for Facial Action Unit RecognitionabstractHumans have the innate ability to rapidly spot and react to a person's emotional response. For computers to be able to understand expressions in a similar way, the gap between the perception of expressions between the humans and computers needs to be minimized. Inspired by human visual fixations, we propose a guided attention mechanism that facilitates the network to `look' at the most important features of a face. Rather than imposing hard attention, we learn the attention maps from the intermediate representation for Action Units (AUs). We propose a joint attention learning and AU classification module with minimal increase in the network parameters. We demonstrate the efficiency of our approach on three standard datasets: BP4D, MMSE and DISFA and obtain state of the art and near state of the art results respectively. Nagashri N. Lakshminarayana, Srirangaraj Setlur, Venu Govindaraju |
FG | 2 |
| 2020 | Equation Attention Relationship Network (EARN) : A Geometric Deep Metric Framework for Learning Similar Math Expression EmbeddingabstractRepresentational Learning in the form of high dimensional embeddings have been used for multiple pattern recognition applications. There has been a significant interest in building embedding based systems for learning representations in the mathematical domain. At the same time, retrieval of structured information such as mathematical expressions is an important need for modern IR systems. In this work, our motivation is to introduce a robust framework for learning representations for similarity based retrieval of mathematical expressions. Given a query by example, the embedding can find the closest matching expression as a function of euclidean distance between them. We leverage recent advancements in image-based and graph-based deep learning algorithms to learn our similarity embeddings. We do this first, by using unimodal encoders in graph space and image space and then, a multi-modal combination of the same. To overcome the lack of training data, we force the networks to learn a deep metric using triplets generated with a heuristic scoring function. We also adopt a custom strategy for mining hard samples to train our neural networks. Our system produces rankings similar to those generated by the original scoring function, but using only a fraction of the time. Our results establish the viability of using such a multi-modal embedding for this task. Saleem Ahmed, Kenny Davila, Srirangaraj Setlur, Venu Govindaraju |
ICPR | 3 |
| 2020 | Automated Whiteboard Lecture Video Summarization by Content Region Detection and RepresentationabstractLecture videos are rapidly becoming an invaluable source of information for students across the globe. Given the large number of online courses currently available, it is important to condense the information within these videos into a compact yet representative summary that can be used for search-based applications. We propose a framework to summarize whiteboard lecture videos by finding feature representations of detected handwritten content regions to determine unique content. We investigate multi-scale histogram of gradients and embeddings from deep metric learning for feature representation. We explicitly handle occluded, growing and disappearing handwritten content. Our method is capable of producing two kinds of lecture video summaries - the unique regions themselves or so-called key content and keyframes (which contain all unique content in a video segment). We use weighted spatio-temporal conflict minimization to segment the lecture and produce keyframes from detected regions and features. We evaluate both types of summaries and find that we obtain state-of-the-art peformance in terms of number of summary keyframes while our unique content recall and precision are comparable to state-of-the-art. Bhargava Urala Kota, Alexander Stone, Kenny Davila, Srirangaraj Setlur, Venu Govindaraju |
ICPR | 4 |
| 2020 | Domain adaptive representation learning for facial action unit recognition
Nishant Sankaran, Deen Dayal Mohan, Nagashri N. Lakshminarayana, Srirangaraj Setlur, Venu Govindaraju |
Pattern Recognit. | 4 |
| 2019 | Tangent-V: Math Formula Image Search Using Line-of-Sight Graphs
Kenny Davila, Ritvik Joshi, Srirangaraj Setlur, Venu Govindaraju, Richard Zanibbi |
ECIR (1) | 3 |
| 2019 | Multimodal Deep Feature Aggregation for Facial Action Unit Recognition using Visible Images and Physiological SignalsabstractIn this paper we present a feature aggregation method to combine the information from the visible light domain and the physiological signals for predicting the 12 facial action units in the MMSE dataset. Although multimodal affect analysis has gained lot of attention, the utility of physiological signals in recognizing facial action units is relatively unexplored. In this paper we investigate if physiological signals such as Electro Dermal Activity (EDA), Respiration Rate and Pulse Rate can be used as metadata for action unit recognition. We exploit the effectiveness of deep learning methods to learn an optimal combined representation that is derived from the individual modalities. We obtained an improved performance on MMSE dataset further validating our claim. To the best of our knowledge this is the first study on facial action unit recognition using physiological signals. Nagashri N. Lakshminarayana, Nishant Sankaran, Srirangaraj Setlur, Venu Govindaraju |
FG | 3 |
| 2019 | Representation Learning Through Cross-Modality SupervisionabstractLearning robust representations for applications with multiple modalities of input can have a significant impact on its performance. Traditional representation learning methods rely on projecting the input modalities to a common subspace to maximize agreement amongst the modalities for a particular task. We propose a novel approach to representation learning that uses a latent representation decoder to reconstruct the target modality and thereby employs the target modality purely as a supervision signal for discovering correlations between the modalities. Through cross modality supervision, we demonstrate that the learnt representation is able to improve the performance of the task of facial action unit (AU) recognition when compared with the modality specific representations and even their fused counterparts. Our experiments on three AU recognition datasets - MMSE, BP4D and DISFA, show strong performance gains producing state-of-the-art results in spite of the absence of a modality. Nishant Sankaran, Deen Dayal Mohan, Srirangaraj Setlur, Venu Govindaraju, Dennis Fedorishin |
FG | 3 |
| 2019 | ICDAR 2019 Competition on Harvesting Raw Tables from Infographics (CHART-Infographics)abstractThis work summarizes the results of the first Competition on Harvesting Raw Tables from Infographics (ICDAR 2019 CHART-Infographics). The complex process of automatic chart recognition is divided into multiple tasks for the purpose of this competition, including Chart Image Classification (Task 1), Text Detection and Recognition (Task 2), Text Role Classification (Task 3), Axis Analysis (Task 4), Legend Analysis (Task 5), Plot Element Detection and Classification (Task 6.a), Data Extraction (Task 6.b), and End-to-End Data Extraction (Task 7). We provided a large synthetic training set and evaluated submitted systems using newly proposed metrics on both synthetic charts and manually-annotated real charts taken from scientific literature. A total of 8 groups registered for the competition out of which 5 submitted results for tasks 1-5. The results show that some tasks can be performed highly accurately on synthetic data, but all systems did not perform as well on real world charts. The data, annotation tools, and evaluation scripts have been publicly released for academic use. Kenny Davila, Bhargava Urala Kota, Srirangaraj Setlur, Venu Govindaraju, Chris Tensmeyer, Ritwick Chaudhry |
ICDAR | 3 |
| 2019 | Content Extraction from Lecture Video via Speaker Action Classification Based on Pose InformationabstractOnline lecture videos are increasingly important e-learning materials for students. Automated content extraction from lecture videos facilitates information retrieval applications that improve access to the lecture material. A significant number of lecture videos include the speaker in the image. Speakers perform various semantically meaningful actions during the process of teaching. Among all the movements of the speaker, key actions such as writing or erasing potentially indicate important features directly related to the lecture content. In this paper, we present a methodology for lecture video content extraction using the speaker actions. Each lecture video is divided into small temporal units called action segments. Using a pose estimator, body and hands skeleton data are extracted and used to compute motion-based features describing each action segment. Then, the dominant speaker action of each of these segments is classified using Random forests and the motion-based features. With the temporal and spatial range of these actions, we implement an alternative way to draw key-frames of handwritten content from the video. In addition, for our fixed camera videos, we also use the skeleton data to compute a mask of the speaker writing locations for the subtraction of the background noise from the binarized key-frames. Our method has been tested on a publicly available lecture video dataset, and it shows reasonable recall and precision results, with a very good compression ratio which is better than previous methods based on content analysis. Kenny Davila, Srirangaraj Setlur, Venu Govindaraju |
ICDAR | 3 |
| 2019 | Generalized framework for summarization of fixed-camera lecture videos by detecting and binarizing handwritten content
Bhargava Urala Kota, Kenny Davila, Alexander Stone, Srirangaraj Setlur, Venu Govindaraju |
Int. J. Document Anal. Recognit. | 4 |
| 2019 | Learning deep features for online person tracking using non-overlapping cameras: A survey
Neeti Narayan, Nishant Sankaran, Srirangaraj Setlur, Venu Govindaraju |
Image Vis. Comput. | 3 |
| 2018 | Wardrobe Model for Long Term Re-identification and Appearance PredictionabstractLong-term surveillance applications often involve having to re-identify individuals over several days or weeks. The task is made even more challenging with the lack of sufficient visibility of the subjects faces. We address this problem by modeling the wardrobe of individuals using discriminative features and labels extracted from their clothing information from video sequences. In contrast to previous person re-id works, we exploit that people typically own a limited amount of clothing and that knowing a person's wardrobe can be used as a soft-biometric to distinguish identities. We a) present a new dataset consisting of more than 70,000 images recorded over 30 days of 25 identities; b) model clothing features using CNNs that minimize intra-garments variations while maximizing inter-garments differences; and c) build a reference wardrobe model that captures each persons set of clothes that can be used for re-id. We show that these models open new perspectives to long-term person re-id problem using clothing information. Kyung Won Lee, Nishant Sankaran, Srirangaraj Setlur, Nils Napp, Venu Govindaraju |
AVSS | 3 |
| 2018 | Knowledge Transfer Using Neural Network Based Approach for Handwritten Text RecognitionabstractThe goal of a writer adaptive handwriting recognition system is to build a model that improves the recognition of a generic recognition model for a specific author. In this work, we show how structural representation learned from a generic writer-independent handwriting recognition model can be customized to individual authors. Convolutional Neural Network has shown outstanding performance in learning image-based representation that was used for classification. Additionally, they have been used along with Recurrent Neural Network (RNN) or its variations like, LSTM and GRU layers to analyze and understand sequences in handwriting recognition, sentence analysis, voice recognition etc. In most cases, the CNNs serve as a feature extractor instead of low-level hand-designed features that were used previously for the above-mentioned classification tasks. We design a method to reuse weights from layers trained on the IAM offline handwritten dataset to compute mid-level image representation for text in the Washington and Moore dataset. We show that despite differences in the writing style, fonts across these datasets, the transferred representation is able to capture a spatio-temporal representation leading to significantly improved recognition results. We hypothesize that the performance is solely not dependent on the number of samples and the model is evaluated with varying amount of fine-tuning samples showing promising results backing the hypothesis. Rathin Radhakrishnan Nair, Nishant Sankaran, Bhargava Urala Kota, Sergey Tulyakov, Srirangaraj Setlur, Venu Govindaraju |
DAS | 5 |
| 2018 | Automated Detection of Handwritten Whiteboard Content in Lecture Videos for SummarizationabstractOnline lecture videos are a valuable resource for students across the world. The ability to find videos based on their content could make them even more useful. Methods for automatic extraction of this content reduce the amount of manual effort required to make indexing and retrieval of such videos possible. We adapt a deep learning based method for scene text detection, for the purpose of detection of handwritten text, math expressions and sketches in lecture videos. We detect handwritten elements on the whiteboard to generate a summary of all content over time in the lecture, while also dealing with occluded content due to motion of the lecturer. We train, test on the publicly available AccessMath lecture video dataset and evaluate our framework on the basis of number of summary frames, as well as recall and precision of all whiteboard content in the set of test lecture videos. We found that our method increases the precision of the state-of-the-art while there is potential to increase recall as well. We have added to the existing ground truth in the AccessMath dataset by providing timestamp-based, semantically meaningful bounding box annotations for the handwritten whiteboard content, which has been released. Bhargava Urala Kota, Kenny Davila, Alexander Stone, Srirangaraj Setlur, Venu Govindaraju |
ICFHR | 4 |
| 2017 | Score normalization in stratified biometric systemsabstractStratified biometric system can be defined as a system in which the subjects, their templates or matching scores can be separated into two or more categories, or strata, and the matching decisions can be made separately for each stratum. In this paper we investigate the properties of the strat-ifiedbiometric system and, in particular, possible strata creation strategies, score normalization and acceptance decisions, expected performance improvements due to stratification. We perform our experiments on face recognition matching scores from IARPA Janus CS2 dataset. Sergey Tulyakov, Nishant Sankaran, Srirangaraj Setlur, Venu Govindaraju |
IJCB | 3 |
| 2013 | A Model Based Framework for Table Processing in Degraded Document ImagesabstractThis paper describes a model based framework for detection and extraction of the contents of table cells from degraded handwritten document images that contain tables. Given the very poor quality of the target documents, the table cell detection problem is formulated conceptually as a two-step process. The first step is to identify the location of the table and extract the content of table cells given a model of the structure of the table present in the image. The second step is to identify the model of the table present in a document image from a list of given table models. A model-based representation for tables is introduced and is used for matching table candidates with the given model to identify and extract the contents of table cells. The approach for detecting potential table candidates is based on the detection of horizontal and vertical table line candidates. The table representation is a matrix of horizontal and vertical table line crossings, and the matching algorithm is formulated as a minimization problem where the optimal table candidate is obtained using the minimal distance between the candidate and model table matrices which is then used for extraction of the table cell contents. A similar approach is used to solve the model selection problem where the best fitting location in the document page for each of the candidate models is identified using the distance minimization approach along with a confidence score and the model with the highest confidence score is selected as the correct model. The approach was tested on document page images containing tables from the challenge set of the DARPA MADCAT handwritten document image data. Results indicate that the method is effective for both model selection as well as table cell content extraction. Zhixin Shi, Srirangaraj Setlur, Venu Govindaraju |
ICDAR | 2 |
| 2013 | IBM_UB_1: A Dual Mode Unconstrained English Handwriting DatasetabstractIn this paper we present a new dual mode, twin-folio structured English handwriting dataset IBM_UB_1. IBM_UB_1 is our first major release from a large multilingual handwriting corpus. Containing over 6000 pages of handwritten matter, this dataset can not only be used for unconstrained handwriting recognition, more importantly, the dataset's unique twin-folio structure presents a natural fit for research on writer identification, keyword spotting, indexing and various forms of handwritten document search and retrieval. We first describe two central characteristics of the dataset - the twin-folio structure and dual modality (online/offline) - and their relevance to current research problems. Secondly, we describe the dataset, its collection and construction, and provide key descriptive statistics. Finally, we evaluate the dataset on two different research domains - handwriting recognition and writer identification - and present related experimental results. Arti Shivram, Chetan Ramaiah, Srirangaraj Setlur, Venu Govindaraju |
ICDAR | 3 |
| 2013 | Segmentation Based Online Word Recognition: A Conditional Random Field Driven Beam Search StrategyabstractWe propose a segmentation based online word recognition approach which uses a Conditional Random Field (CRF) driven beam search strategy. An efficient trie-lexicon directed, breadth-first beam search algorithm is employed in a combined segmentation-and-recognition framework to accomplish real-time recognition of online handwritten cursive English words. This framework is developed by building a candidate lattice of primitive segments obtained through over segmentation of the word pattern. The search space for the lattice is expanded by synchronously matching the lattice nodes to likely character patterns from a trie-dictionary constructed out of the target lexicon. The probable paths are evaluated by integrating character recognition scores with physical and spatial characteristics of the handwritten segments in a CRF (conditional random field) model and a beam search strategy is used to prune the set of likely paths. This approach has been benchmarked on the new IBM_UB_1 dataset as well as on the UNIPEN dataset for comparison. Arti Shivram, Bilan Zhu, Srirangaraj Setlur, Masaki Nakagawa, Venu Govindaraju |
ICDAR | 3 |
| 2013 | Online Handwritten Cursive Word Recognition Using Segmentation-Free MRF in Combination with P2DBMN-MQDFabstractThis paper describes an online handwritten English cursive word recognition method using a segmentation-free Markov random field (MRF) model in combination with an offline recognition method which uses pseudo 2D bi-moment normalization (P2DBMN) and modified quadratic discriminant function (MQDF). It extracts feature points along the pen-tip trace from pen-down to pen-up and uses the feature point coordinates as unary features and the differences in coordinates between the neighboring feature points as binary features. Each character is modeled as a MRF and word MRFs are constructed by concatenating character MRFs according to a trie lexicon of words during recognition. Our method expands the search space using a character-synchronous beam search strategy to search the segmentation and recognition paths. This method restricts the search paths from the trie lexicon of words and preceding paths, as well as the lengths of feature points during path search. We also combine it with a P2DBMN-MQDF recognizer that is widely used for Chinese and Japanese character recognition. Bilan Zhu, Arti Shivram, Srirangaraj Setlur, Venu Govindaraju, Masaki Nakagawa |
ICDAR | 3 |
| 2013 | Handwritten text separation from annotated machine printed documents using Markov Random Fields
Xujun Peng, Srirangaraj Setlur, Venu Govindaraju, Ramachandrula Sitaram |
Int. J. Document Anal. Recognit. | 2 |
| 2012 | Keyword Spotting Framework Using Dynamic Background ModelabstractAn important task in Keyword Spotting in handwritten documents is to separate Keywords from Non Keywords. Very often this is achieved by learning a filler or background model. A common method of building a background model is to allow all possible sequences or transitions of characters. However, due to large variation in handwriting styles, allowing all possible sequences of characters as background might result in an increased false reject. A weak background model could result in high false accept. We propose a novel way of learning the background model dynamically. The approach first used in word spotting in speech uses a feature vector of top K local scores per character and top N global scores of matching hypotheses. A two class classifier is learned on these features to classify between Keyword and Non Keyword. Zhixin Shi, Srirangaraj Setlur, Venu Govindaraju, Ramachandrula Sitaram |
ICFHR | 3 |
| 2012 | Using a boosted tree classifier for text segmentation in hand-annotated documents
Xujun Peng, Srirangaraj Setlur, Venu Govindaraju, Ramachandrula Sitaram |
Pattern Recognit. Lett. | 2 |
| 2011 | Image Enhancement for Degraded Binary Document ImagesabstractThis paper presents a novel set of image enhancement algorithms for binary images of poorly scanned real world page documents. Problems that are targeted by the methods described include large blobs or clutter noise, salt-and-pepper noise and detection and removal of non-text objects such as form lines or rule-lines. The algorithms described are shown to be very effective in removing clutter noise and pepper noise as well as form lines and rule-lines. A region growing algorithm is also described to enhance the quality of the text and to fix the problems arising from the salt noise which leaves holes in the text and creates broken strokes. The methods were tested on 204 images from the challenge set of the DARPA MADCAT Arabic handwritten document image data. The results indicate that the methods described are robust and are capable of significantly improving the image quality for downstream OCR systems. Zhixin Shi, Srirangaraj Setlur, Venu Govindaraju |
ICDAR | 2 |
| 2010 | Latent Dirichlet allocation based writer identification in offline handwritingabstractIn this paper, we describe a novel approach to Writer Identification in Offline handwriting using Latent Dirichlet Allocation. State-of-the-art methods for writer identification employ the traditional feature-classification paradigm which does not provide enough information about the handwriting attributes such as writing style which are key components in any forensic analysis of handwriting. This problem is also compounded due to lack of efficient rules for defining a particular writing style that can capture writer specific characteristics over a large dataset. We propose to address this issue by using a generative model in form of Latent Dirichlet Allocation(LDA) that automatically infers writing styles from handwritten document collection without any pre-defined set of rules. This information is then used to represent each writer as a distribution over multiple writing style for classifying any unknown writer sample. We describe our approach on two different feature sets consisting of contour angle features as well as structural and concavity features. Our experimental results show comparable performance with baseline systems and also demonstrate the efficacy of LDA for learning multiple handwriting styles. Anurag Bhardwaj, Manavender R. Malgireddy, Srirangaraj Setlur, Venu Govindaraju, Ramachandrula Sitaram |
Document Analysis Systems | 3 |
| 2010 | Overlapped text segmentation using Markov random field and aggregationabstractSeparating machine printed text and handwriting from overlapping text is a challenging problem in the document analysis field and no reliable algorithms have been developed thus far. In this paper, we propose a novel approach for separating handwriting from binary image of overlapped text. Instead of using fixed size training patches, we describe an aggregation method which uses shape context features to extract training samples automatically. We use a Markov Random Field (MRF) to model the overlapped text. The neighbor system is inherited from a coarsening procedure and the prior and likelihood of the MRF is learned based on a distance metric. Experimental results show that the proposed method can achieve 87.97% recall for handwriting and 91.44% recall for machine printed text. Xujun Peng, Srirangaraj Setlur, Venu Govindaraju, Ramachandrula Sitaram |
Document Analysis Systems | 2 |
| 2010 | A Framework for Hand Gesture Recognition and Spotting Using Sub-gesture ModelingabstractHand gesture interpretation is an open research problem in Human Computer Interaction (HCI), which involves locating gesture boundaries (Gesture Spotting) in a continuous video sequence and recognizing the gesture. Existing techniques model each gesture as a temporal sequence of visual features extracted from individual frames which is not efficient due to the large variability of frames at different timestamps. In this paper, we propose a new sub-gesture modeling approach which represents each gesture as a sequence of fixed sub-gestures (a group of consecutive frames with locally coherent context) and provides a robust modeling of the visual features. We further extend this approach to the task of gesture spotting where the gesture boundaries are identified using a filler model and gesture completion model. Experimental results show that the proposed method outperforms state-of-the-art Hidden Conditional Random Fields (HCRF) based methods and baseline gesture spotting techniques. Manavender R. Malgireddy, Jason J. Corso, Srirangaraj Setlur, Venu Govindaraju, Dinesh Mandalapu |
ICPR | 3 |
| 2010 | Text Separation from Mixed Documents Using a Tree-Structured ClassifierabstractIn this paper, we propose a tree-structured multi-class classifier to identify annotations and overlapping text from machine printed documents. Each node of the tree-structured classifier is a binary weak learner. Unlike normal decision tree(DT) which only considers a subset of training data at each node and is susceptible to over-fitting, we boost the tree using all training data at each node with different weights. The evaluation of the proposed method is presented on a set of machine printed documents which have been annotated by multiple writers in an office/collaborative environment. Xujun Peng, Srirangaraj Setlur, Venu Govindaraju, Ramachandrula Sitaram |
ICPR | 2 |
| 2010 | Removing Rule-Lines from Binary Handwritten Arabic Document Images Using Directional Local ProfileabstractIn this paper, we present a novel approach for detecting and removing pre-printed rule-lines from binary handwritten Arabic document images. The proposed technique is based on a directional local profiling approach for the detection of the rule-line locations. Then a refined adaptive vertical run-length search is designed for removing the rule-line pixels without much damaging to the text. They are also tolerate to the variations in the rule-lines such as broken lines, orientation changes and variation in the thickness of the rule-lines. Analysis of experimental results on the DARPA MADCAT Arabic handwritten document data indicates that the method is robust and is capable of correctly removing rule-lines. Zhixin Shi, Srirangaraj Setlur, Venu Govindaraju |
ICPR | 2 |
| 2009 | Markov Random Field Based Text Identification from Annotated Machine Printed DocumentsabstractIn this paper, we describe an approach to segment handwritten text, machine printed text and noise from annotated machine printed documents. Three categories of word level features are extracted. We use a modified K-Means clustering algorithm for classification followed by a relabeling procedure using Markov Random Field(MRF) based on a concept of neighboring patches and Belief Propagation(BP) rules. Experimental results on an imbalanced data set show that our approach achieves an overall recall of 96.33%. Xujun Peng, Srirangaraj Setlur, Venu Govindaraju, Ramachandrula Sitaram, Kiran Bhuvanagiri |
ICDAR | 2 |
| 2009 | A Steerable Directional Local Profile Technique for Extraction of Handwritten Arabic Text LinesabstractIn this paper, we present a new text line extraction method for handwritten Arabic documents. The proposed technique is based on a generalized adaptive local connectivity map (ALCM) using a steerable directional filter. The algorithm is designed to solve the particularly complex problems seen in handwritten documents such as fluctuating, touching or crossing text lines. The proposed algorithm consists of three steps. Firstly, a steerable filter is used to probe and determine foreground intensity along multiple directions at each pixel while generating the ALCM. The ALCM is then binarized using an adaptive thresholding algorithm to get a rough estimate of the location of the text lines. In the second step, connected component analysis is used to classify text and non text patterns in the generated ALCM to refine the location of the text lines. Finally, the text lines are separated by superimposing the text line patterns in the ALCM on the original document image and extracting the connected components covered by the pattern mask. Analysis of experimental results on the DARPA MADCAT Arabic handwritten document data indicate that the method is robust and is capable of correctly isolating handwritten text lines even on challenging document images. Zhixin Shi, Srirangaraj Setlur, Venu Govindaraju |
ICDAR | 2 |
| 2009 | Devanagari OCR using a recognition driven segmentation framework and stochastic language models
Suryaprakash Kompalli, Srirangaraj Setlur, Venu Govindaraju |
Int. J. Document Anal. Recognit. | 2 |
| 2005 | Challenges in OCR of Dev anagari DocumentsabstractOCR of Devanagari script presents a wide range of challenges that are not seen in Latin based scripts. This paper outlines the implementation of a neural network based Devanagari OCR. Experimental results on a standard data set are reported and analyzed. Suryaprakash Kompalli, Sankalp Nayak, Srirangaraj Setlur, Venu Govindaraju |
ICDAR | 3 |
| 2005 | A Lexicon Reduction Strategy in the Context of Handwritten Medical FormsabstractTraditional handwriting recognition algorithms rely heavily on small lexicons and clean word images. Unfortunately, emergency medical documents do not satisfy either of these conditions. This is a significant road-block that is hampering efforts to rapidly convert valuable offline healthcare handwriting data into digital content that can be efficiently mined for information. This paper describes a strategy whereby given an image representing a noisy handwritten word from a medical document, and a large lexicon consisting of English, medical and pharmacological words, symbols, abbreviations and acronyms, significantly reduces the size of the lexicon while keeping the unknown desired entry within the lexicon. The approach combines geometric interpretations of the word image along with contextual inference of concepts to reduce lexicons for word recognition. The data extracted can then be efficiently and securely disseminated for epidemiological and outbreak detection/analysis. Experimental results on NY State PCR forms are reported. Robert Milewski, Srirangaraj Setlur, Venu Govindaraju |
ICDAR | 2 |
| 2005 | Text Extraction from Gray Scale Historical Document Images Using Adaptive Local Connectivity MapabstractThis paper presents an algorithm using adaptive local connectivity map for retrieving text lines from the complex handwritten documents such as handwritten historical manuscripts. The algorithm is designed for solving the particularly complex problems seen in handwritten documents. These problems include fluctuating text lines, touching or crossing text lines and low quality image that do not lend themselves easily to binarizations. The algorithm is based on connectivity features similar to local projection profiles, which can be directly extracted from gray scale images. The proposed technique is robust and has been tested on a set of complex historical handwritten documents such as Newton's and Galileo's manuscripts. A preliminary testing shows a successful location rate of above 95% for the test set. Zhixin Shi, Srirangaraj Setlur, Venu Govindaraju |
ICDAR | 2 |
| 2004 | DL Architecture for Indic Scripts
Suryaprakash Kompalli, Srirangaraj Setlur, Venu Govindaraju |
Document Analysis Systems | 2 |
| 2003 | Text - Image Separation in Devanagari DocumentsabstractIn this paper we present a top-down, projection-profile based algorithm to separate text blocks from image blocks in a Devanagari document. We use a distinctive feature of Devanagari text, called Shirorekha (Header Line) to analyze the pattern produced by Devanagari text in the horizontal profile. The horizontal profile corresponding to a text block possesses certain regularity in frequency, orientation and shows spatial cohesion. The algorithm uses these features to identify text blocks in a document image containing both text and graphics. Swapnil Khedekar, Vemulapati Ramanaprasad, Srirangaraj Setlur, Venu Govindaraju |
ICDAR | 3 |
| 2002 | Large scale address recognition systems Truthing, testing, tools, and other evaluation issues
Srirangaraj Setlur, Alfred Lawson, Venu Govindaraju, Sargur N. Srihari |
Int. J. Document Anal. Recognit. | 1 |
| 2001 | Truthing, Testing and Evaluation Issues in Complex SystemsabstractThis paper describes the issues involved in the design of a system for evaluating improvements in the performance of a real-time address recognition system being used by the United States Postal Service for processing mail-piece images. Evaluation of the performance of recognition systems is normally carried out by measuring the performance of the system on a representative sample of images. Designing a comprehensive and valid testing scenario is a complex task that requires careful attention. Sampling live mail-stream to generate a deck of images representative of the general mail-stream for testing, truthing (generating reference data on a significant number of images), grading and evaluation, and designing tools to facilitate these functions are important topics that need to be addressed. This paper describes the efforts of the United States Postal Service and CEDAR towards developing an infrastructure for sampling, truthing and testing of mail-stream images. Srirangaraj Setlur, Venu Govindaraju, Sargur N. Srihari, Alfred Lawson |
ICDAR | 1 |
| 1994 | Generating manifold samples from a handwritten word
Srirangaraj Setlur, Venu Govindaraju |
Pattern Recognit. Lett. | 1 |