VLDB 2026 Research / reviewers in the wild / expert
R. Manmatha
dblp:m/RManmatha · also Raghavan Manmatha
· DBLP profile ↗
76ranked-venue papers
8as first author
16since 2021 · last 2025
0000-0003-2315-8583ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 49 · 6 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 39 · 5 first-author · 13 since 2021Databases, data management, data science and information retrieval · 26 · 2 first-author · 1 since 2021Computer networks · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scaling up Image Segmentation across Data and TasksabstractTraditional segmentation models, while effective in isolated tasks, often fail to generalize to more complex and open-ended segmentation problems, such as free-form, open-vocabulary, and in-the-wild scenarios. To bridge this gap, we propose to scale up image segmentation across diverse datasets and tasks such that the knowledge across different tasks and datasets can be integrated while improving the generalization ability. Mixed-Query Transformer (MQ-Former), a novel segmentation framework, is introduced and designed to scale seamlessly across both data size and task diversity. It is built upon a dynamic object query mechanism called mixed query, which fuses different types of queries using cross-attention. This hybrid approach enables the model to balance between instance- and stuff-level segmentation, providing enhanced scalability for handling diverse object types. We further enhance scalability by leveraging synthetic data-generating segmentation masks and captions for pixel-level and open-vocabulary tasks-drastically reducing the need for costly human annotations. By training on multiple datasets and tasks at scale, MQ-Former continuously improves performance as the volume and diversity of data and tasks increase. It exhibits strong generalization capabilities, boosting performance in open-set segmentation tasks SeginW by 7 points. These advancements mark a key step toward universal, scalable segmentation models capable of addressing the demands of real-world applications. Zhaowei Cai, Hao Yang 0043, Ashwin Swaminathan, R. Manmatha, Stefano Soatto |
CVPR | 5 |
| 2024 | DocFormerv2: Local Features for Document UnderstandingabstractWe propose DocFormerv2, a multi-modal transformer for Visual Document Understanding (VDU). The VDU domain entails understanding documents (beyond mere OCR predictions) e.g., extracting information from a form, VQA for documents and other tasks. VDU is challenging as it needs a model to make sense of multiple modalities (visual, language and spatial) to make a prediction. Our approach, termed DocFormerv2 is an encoder-decoder transformer which takes as input - vision, language and spatial features. DocFormerv2 is pre-trained with unsupervised tasks employed asymmetrically i.e., two novel document tasks on encoder and one on the auto-regressive decoder. The unsupervised tasks have been carefully designed to ensure that the pre-training encourages local-feature alignment between multiple modalities. DocFormerv2 when evaluated on nine challenging datasets shows state-of-the-art performance on all over strong baselines - On TabFact (+4.3%), InfoVQA (+1.4%), FUNSD (+1.0%). Furthermore, to show generalization capabilities, on three VQA tasks involving scene-text, DocFormerv2 outperforms previous comparably-sized models and even does better than much larger models (such as GIT2, PaLI and Flamingo) on these tasks. Extensive ablations show that due to its novel pre-training tasks, DocFormerv2 understands multiple modalities better than prior-art in VDU. Srikar Appalaraju, Peng Tang 0005, Nishant Sankaran, Yichu Zhou, R. Manmatha |
AAAI | 6 |
| 2024 | No Head Left Behind - Multi-Head Alignment Distillation for TransformersabstractKnowledge distillation aims at reducing model size without compromising much performance. Recent work has applied it to large vision-language (VL) Transformers, and has shown that attention maps in the multi-head attention modules of vision-language Transformers contain extensive intra-modal and cross-modal co-reference relations to be distilled. The standard approach is to apply a one-to-one attention map distillation loss, i.e. the Teacher's first attention head instructs the Student's first head, the second teaches the second, and so forth, but this only works when the numbers of attention heads in the Teacher and Student are the same. To remove this constraint, we propose a new Attention Map Alignment Distillation (AMAD) method for Transformers with multi-head attention, which works for a Teacher and a Student with different numbers of attention heads. Specifically, we soft-align different heads in Teacher and Student attention maps using a cosine similarity weighting. The Teacher head contributes more to the Student heads for which it has a higher similarity weight. Each Teacher head contributes to all the Student heads by minimizing the divergence between the attention activation distributions for the soft-aligned heads. No head is left behind. This distillation approach operates like cross-attention. We experiment on distilling VL-T5 and BLIP, and apply AMAD loss on their T5, BERT, and ViT sub-modules. We show, under vision-language setting, that AMAD outperforms conventional distillation methods on VQA-2.0, COCO captioning, and Multi30K translation datasets. We further show that even without VL pre-training, the distilled VL-T5 models outperform corresponding VL pre-trained VL-T5 models that are further fine-tuned by ground-truth signals, and that fine-tuning distillation can also compensate to some degree for the absence of VL pre-training for BLIP models. Tianyang Zhao 0004, Kunwar Yashraj Singh, Srikar Appalaraju, Peng Tang 0005, Vijay Mahadevan, R. Manmatha, Ying Nian Wu |
AAAI | 6 |
| 2024 | On the Scalability of Diffusion-based Text-to-Image GenerationabstractScaling up model and data size has been quite successful for the evolution of LLMs. However, the scaling law for the diffusion based text-to-image (T2I) models is not fully explored. It is also unclear how to efficiently scale the model for better performance at reduced cost. The different training settings and expensive training cost make a fair model comparison extremely difficult. In this work, we empirically study the scaling properties of diffusion based T2I models by performing extensive and rigours ablations on scaling both denoising backbones and training set, including training scaled UNet and Transformer variants ranging from 0.4B to 4B parameters on datasets upto 600M images. For model scaling, we find the location and amount of cross attention distinguishes the performance of existing UNet designs. And increasing the transformer blocks is more parameter-efficient for improving text-image alignment than increasing channel numbers. We then identify an efficient UNet variant, which is 45% smaller and 28% faster than SDXL's UNet. On the data scaling side, we show the quality and diversity of the training set matters more than simply dataset size. Increasing caption density and diversity improves text-image alignment performance and the learning efficiency. Finally, we provide scaling functions to predict the text-image alignment performance as functions of the scale of model size, compute and dataset size. Orchid Majumder, Yusheng Xie, R. Manmatha, Ashwin Swaminathan, Zhuowen Tu, Stefano Ermon, Stefano Soatto |
CVPR | 6 |
| 2024 | VisFocus: Prompt-Guided Vision Encoders for OCR-Free Dense Document Understanding
Ofir Abramovich, Niv Nayman, Sharon Fogel, Inbal Lavi, Ron Litman, Shahar Tsiper, Royee Tichauer, Srikar Appalaraju, Shai Mazor, R. Manmatha |
ECCV (8) | 10 |
| 2024 | DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding ModelsabstractSungnyun Kim, Haofu Liao, Srikar Appalaraju, Peng Tang, Zhuowen Tu, Ravi Kumar Satzoda, R. Manmatha, Vijay Mahadevan, Stefano Soatto. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Sungnyun Kim, Haofu Liao, Srikar Appalaraju, Peng Tang 0005, Zhuowen Tu, Ravi Kumar Satzoda, R. Manmatha, Vijay Mahadevan, Stefano Soatto |
EMNLP | 7 |
| 2024 | ICDAR 2024 Competition on Recognition and VQA on Handwritten Documents
Ajoy Mondal, Vijay Mahadevan, R. Manmatha, C. V. Jawahar |
ICDAR (6) | 3 |
| 2024 | Improving Semantic Segmentation via Efficient Self-TrainingabstractStarting from the seminal work of Fully Convolutional Networks (FCN), there has been significant progress on semantic segmentation. However, deep learning models often require large amounts of pixelwise annotations to train accurate and robust models. Given the prohibitively expensive annotation cost of segmentation masks, we introduce a self-training framework in this paper to leverage pseudo labels generated from unlabeled data. In order to handle the data imbalance problem of semantic segmentation, we propose a centroid sampling strategy to uniformly select training samples from every class within each epoch. We also introduce a fast training schedule to alleviate the computational burden. This enables us to explore the usage of large amounts of pseudo labels. Our Centroid Sampling based Self-Training framework (CSST) achieves state-of-the-art results on Cityscapes and CamVid datasets. On PASCAL VOC 2012 test set, our models trained with the original train set even outperform the same models trained on the much bigger augmented train set. This indicates the effectiveness of CSST when there are fewer annotations. We also demonstrate promising few-shot generalization capability from Cityscapes to BDD100K and from Cityscapes to Mapillary datasets. Yi Zhu 0001, Chongruo Wu, Zhi Zhang 0005, Tong He 0002, Hang Zhang 0005, R. Manmatha, Mu Li 0003, Alexander J. Smola |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | PolyFormer: Referring Image Segmentation as Sequential Polygon GenerationabstractIn this work, instead of directly predicting the pixel-level segmentation masks, the problem of referring image seg-mentation is formulated as sequential polygon generation, and the predicted polygons can be later converted into segmentation masks. This is enabled by a new sequence-to-sequence framework, Polygon Transformer (PolyFormer), which takes a sequence of image patches and text query to-kens as input, and outputs a sequence of polygon vertices autoregressively. For more accurate geometric localization, we propose a regression-based decoder, which predicts the precise floating-point coordinates directly, without any co-ordinate quantization error. In the experiments, PolyFormer outperforms the prior art by a clear margin, e.g., 5.40% and 4.52% absolute improvements on the challenging Re-fCOCO+ and RefCOCOg datasets. It also shows strong generalization ability when evaluated on the referring video segmentation task without fine-tuning, e.g., achieving competitive 61.5% J&F on the Ref-DAVIS17 dataset. Zhaowei Cai, Ravi Kumar Satzoda, Vijay Mahadevan, R. Manmatha |
CVPR | 7 |
| 2023 | DocTr: Document Transformer for Structured Information Extraction in DocumentsabstractWe present a new formulation for structured information extraction (SIE) from visually rich documents. We address the limitations of existing IOB tagging and graph-based formulations, which are either overly reliant on the correct ordering of input text or struggle with decoding a complex graph. Instead, motivated by anchor-based object detectors in computer vision, we represent an entity as an anchor word and a bounding box, and represent entity linking as the association between anchor words. This is more robust to text ordering, and maintains a compact graph for entity linking. The formulation motivates us to introduce 1) a Document Transformer (DocTr) that aims at detecting and associating entity bounding boxes in visually rich documents, and 2) a simple pre-training strategy that helps learn entity detection in the context of language. Evaluations on three SIE benchmarks show the effectiveness of the proposed formulation, and the overall approach outperforms existing solutions. Haofu Liao, Aruni Roy Chowdhury, Ankan Bansal, Zhuowen Tu, Ravi Kumar Satzoda, R. Manmatha, Vijay Mahadevan |
ICCV | 8 |
| 2022 | LaTr: Layout-Aware Transformer for Scene-Text VQAabstractWe propose a novel multimodal architecture for Scene Text Visual Question Answering (STVQA), named Layout-Aware Transformer (LaTr). The task of STVQA requires models to reason over different modalities. Thus, we first investigate the impact of each modality, and reveal the importance of the language module, especially when enriched with layout information. Accounting for this, we propose a single objective pre-training scheme that requires only text and spatial cues. We show that applying this pre-training scheme on scanned documents has certain advantages over using natural images, despite the domain gap. Scanned documents are easy to procure, text-dense and have a variety of layouts, helping the model learn various spatial cues (e.g. left-of, below etc.) by tying together language and layout information. Compared to existing approaches, our method performs vocabulary-free decoding and, as shown, generalizes well beyond the training vocabulary. We further demonstrate that LaTr improves robustness towards OCR errors, a common reason for failure cases in STVQA. In addition, by leveraging a vision transformer, we eliminate the need for an external object detector. LaTr outperforms state-of-the-art STVQA methods on multiple datasets. In particular, +7.6% on TextVQA, +10.8% on ST-VQA and +4.0% on OCR-VQA (all absolute accuracy numbers). Ali Furkan Biten, Ron Litman, Yusheng Xie, Srikar Appalaraju, R. Manmatha |
CVPR | 5 |
| 2022 | Towards Weakly-Supervised Text Spotting using a Multi-Task TransformerabstractText spotting end-to-end methods have recently gained attention in the literature due to the benefits of jointly optimizing the text detection and recognition components. Existing methods usually have a distinct separation between the detection and recognition branches, requiring exact annotations for the two tasks. We introduce TextTranSpotter (TTS), a transformer-based approach for text spotting and the first text spotting framework which may be trained with both fully- and weakly-supervised settings. By learning a single latent representation per word detection, and using a novel loss function based on the Hungarian loss, our method alleviates the need for expensive localization annotations. Trained with only text transcription annotations on real data, our weakly-supervised method achieves competitive performance with previous state-of-the-art fully-supervised methods. When trained in a fully-supervised manner, TextTranSpotter shows state-of-the-art results on multiple benchmarks. Yair Kittenplon, Inbal Lavi, Sharon Fogel, Yarin Bar, R. Manmatha, Pietro Perona |
CVPR | 5 |
| 2022 | GLASS: Global to Local Attention for Scene-Text Spotting
Roi Ronen, Shahar Tsiper, Oron Anschel, Inbal Lavi, Amir Markovitz, R. Manmatha |
ECCV (28) | 6 |
| 2021 | Sequence-to-Sequence Contrastive Learning for Text RecognitionabstractWe propose a framework for sequence-to-sequence contrastive learning (SeqCLR) of visual representations, which we apply to text recognition. To account for the sequence-to-sequence structure, each feature map is divided into different instances over which the contrastive loss is computed. This operation enables us to contrast in a sub-word level, where from each image we extract several positive pairs and multiple negative examples. To yield effective visual representations for text recognition, we further suggest novel augmentation heuristics, different encoder architectures and custom projection heads. Experiments on hand-written text and on scene text show that when a text decoder is trained on the learned representations, our method out-performs non-sequential contrastive methods. In addition, when the amount of supervision is reduced, SeqCLR significantly improves performance compared with supervised training, and when fine-tuned with 100% of the labels, our method achieves state-of-the-art results on standard hand-written text recognition benchmarks. Aviad Aberdam, Ron Litman, Shahar Tsiper, Oron Anschel, Ron Slossberg, Shai Mazor, R. Manmatha, Pietro Perona |
CVPR | 7 |
| 2021 | DocFormer: End-to-End Transformer for Document UnderstandingabstractWe present DocFormer - a multi-modal transformer based architecture for the task of Visual Document Understanding (VDU). VDU is a challenging problem which aims to understand documents in their varied formats (forms, receipts etc.) and layouts. In addition, DocFormer is pre-trained in an unsupervised fashion using carefully designed tasks which encourage multi-modal interaction. DocFormer uses text, vision and spatial features and combines them using a novel multi-modal self-attention layer. DocFormer also shares learned spatial embeddings across modalities which makes it easy for the model to correlate text to visual tokens and vice versa. DocFormer is evaluated on 4 different datasets each with strong baselines. DocFormer achieves state-of-the-art results on all of them, sometimes beating models 4x its size (in no. of parameters). Srikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota, Yusheng Xie, R. Manmatha |
ICCV | 5 |
| 2021 | Saliency Driven Perceptual Image CompressionabstractThis paper proposes a new end-to-end trainable model for lossy image compression, which includes several novel components. The method incorporates 1) an adequate perceptual similarity metric; 2) saliency in the images; 3) a hierarchical auto-regressive model. This paper demonstrates that the popularly used evaluations metrics such as MS-SSIM and PSNR are inadequate for judging the performance of image compression techniques as they do not align with the human perception of similarity. Alternatively, a new metric is proposed, which is learned on perceptual similarity data specific to image compression. The proposed compression model incorporates the salient regions and optimizes on the proposed perceptual similarity metric. The model not only generates images which are visually better but also gives superior performance for subsequent computer vision tasks such as object detection and segmentation when compared to existing engineered or learned compression techniques. Srikar Appalaraju, R. Manmatha |
WACV | 3 |
| 2020 | SCATTER: Selective Context Attentional Scene Text RecognizerabstractScene Text Recognition (STR), the task of recognizing text against complex image backgrounds, is an active area of research. Current state-of-the-art (SOTA) methods still struggle to recognize text written in arbitrary shapes. In this paper, we introduce a novel architecture for STR, named Selective Context ATtentional Text Recognizer (SCATTER). SCATTER utilizes a stacked block architecture with intermediate supervision during training, that paves the way to successfully train a deep BiLSTM encoder, thus improving the encoding of contextual dependencies. Decoding is done using a two-step 1D attention mechanism. The first attention step re-weights visual features from a CNN backbone together with contextual features computed by a BiLSTM layer. The second attention step, similar to previous papers, treats the features as a sequence and attends to the intra-sequence relationships. Experiments show that the proposed approach surpasses SOTA performance on irregular text recognition benchmarks by 3.7% on average. Ron Litman, Oron Anschel, Shahar Tsiper, Roee Litman, Shai Mazor, R. Manmatha |
CVPR | 6 |
| 2019 | Dependence Models for Searching Text in Document ImagesabstractThe main goal of existing word spotting approaches for searching document images has been the identification of visually similar word images in the absence of high quality text recognition output. Searching for a piece of arbitrary text is not possible unless the user identifies a sample word image from the document collection or generates the query word image synthetically. To address this problem, a Markov Random Field (MRF) framework is proposed for searching document images and shown to be effective for searching arbitrary text in real time for books printed in English (Latin script), Telugu and Ottoman scripts. The English experiments demonstrate that the dependencies between the visual terms and letter bigrams can be automatically learned using noisy OCR output. It is also shown that OCR text search accuracy can be significantly improved if it is combined with the proposed approach. No commercial OCR engine is available for Telugu or Ottoman script. In these cases the dependencies are trained using manually annotated document images. It is demonstrated that the trained model can be directly used to resolve arbitrary text queries across books despite font type and size differences. The proposed approach outperforms a state-of-the-art BLSTM baseline in these contexts. Ismet Zeki Yalniz, R. Manmatha |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Compressed Video Action RecognitionabstractTraining robust deep video representations has proven to be much more challenging than learning deep image representations. This is in part due to the enormous size of raw video streams and the high temporal redundancy; the true and interesting signal is often drowned in too much irrelevant data. Motivated by that the superfluous information can be reduced by up to two orders of magnitude by video compression (using H.264, HEVC, etc.), we propose to train a deep network directly on the compressed video. This representation has a higher information density, and we found the training to be easier. In addition, the signals in a compressed video provide free, albeit noisy, motion information. We propose novel techniques to use them effectively. Our approach is about 4.6 times faster than Res3D and 2.7 times faster than ResNet-152. On the task of action recognition, our approach outperforms all the other methods on the UCF-101, HMDB-51, and Charades dataset. Chao-Yuan Wu, Manzil Zaheer, Hexiang Hu, R. Manmatha, Alexander J. Smola, Philipp Krähenbühl |
CVPR | 4 |
| 2017 | Sampling Matters in Deep Embedding LearningabstractDeep embeddings answer one simple question: How similar are two images? Learning these embeddings is the bedrock of verification, zero-shot learning, and visual search. The most prominent approaches optimize a deep convolutional network with a suitable loss function, such as contrastive loss or triplet loss. While a rich line of work focuses solely on the loss functions, we show in this paper that selecting training examples plays an equally important role. We propose distance weighted sampling, which selects more informative and stable examples than traditional approaches. In addition, we show that a simple margin based loss is sufficient to outperform all other loss functions. We evaluate our approach on the Stanford Online Products, CAR196, and the CUB200-2011 datasets for image retrieval and clustering, and on the LFW dataset for face verification. Our method achieves state-of-the-art performance on all of them. R. Manmatha, Chao-Yuan Wu, Alexander J. Smola, Philipp Krähenbühl |
ICCV | 1 |
| 2016 | Deep Decision Network for Multi-class Image ClassificationabstractIn this paper, we present a novel Deep Decision Network (DDN) that provides an alternative approach towards building an efficient deep learning network. During the learning phase, starting from the root network node, DDN automatically builds a network that splits the data into disjoint clusters of classes which would be handled by the subsequent expert networks. This results in a tree-like structured network driven by the data. The proposed method provides an insight into the data by identifying the group of classes that are hard to classify and require more attention when compared to others. DDN also has the ability to make early decisions thus making it suitable for timesensitive applications. We validate DDN on two publicly available benchmark datasets: CIFAR-10 and CIFAR-100 and it yields state-of-the-art classification performance on both the datasets. The proposed algorithm has no limitations to be applied to any generic classification problems. Venkatesh N. Murthy, Vivek K. Singh 0002, Terrence Chen, R. Manmatha, Dorin Comaniciu |
CVPR | 4 |
| 2016 | Image Annotation using Multi-scale Hypergraph Heat Diffusion FrameworkabstractThe task of automatic image annotation involves assigning relevant multiple labels/tags to query images based on their visual content. One of the key challenge in multi-label image annotation task is the class imbalance problem where frequently occurring labels suppress the participation of rarely occurring labels. In this paper, we propose to exploit the multi-scale behavior in hypergraph heat diffusion framework for the automatic image annotation task. The proposed novel technique enables to model the higher order relationship among images in the feature space and provides a multi-scale label diffusion mechanism to address the class imbalance problem in the data. Venkatesh N. Murthy, Avinash Sharma 0001, Visesh Chari, R. Manmatha |
ICMR | 4 |
| 2015 | Automatic Image Annotation using Deep Learning RepresentationsabstractWe propose simple and effective models for the image annotation that make use of Convolutional Neural Network (CNN) features extracted from an image and word embedding vectors to represent their associated tags. Our first set of models is based on the Canonical Correlation Analysis (CCA) framework that helps in modeling both views - visual features (CNN feature) and textual features (word embedding vectors) of the data. Results on all three variants of the CCA models, namely linear CCA, kernel CCA and CCA with k-nearest neighbor (CCA-KNN) clustering, are reported. The best results are obtained using CCA-KNN which outperforms previous results on the Corel-5k and the ESP-Game datasets and achieves comparable results on the IAPRTC-12 dataset. In our experiments we evaluate CNN features in the existing models which bring out the advantages of it over dozens of handcrafted features. We also demonstrate that word embedding vectors perform better than binary vectors as a representation of the tags associated with an image. In addition we compare the CCA model to a simple CNN based linear regression model, which allows the CNN layers to be trained using back-propagation. Venkatesh N. Murthy, Subhransu Maji, R. Manmatha |
ICMR | 3 |
| 2014 | Sequential Word Spotting in Historical Handwritten DocumentsabstractIn this work we present a handwritten word spotting approach that takes advantage of the a priori known order of appearance of the query words. Given an ordered sequence of query word instances, the proposed approach performs a sequence alignment with the words in the target collection. Although the alignment is quite sparse, i.e. the number of words in the database is higher than the query set, the improvement in the overall performance is sensitively higher than isolated word spotting. As application dataset, we use a collection of handwritten marriage licenses taking advantage of the ordered index pages of family names. David Fernández Mota, R. Manmatha, Alicia Fornés, Josep Lladós 0001 |
Document Analysis Systems | 2 |
| 2014 | Modeling Concept Dependencies for Event DetectionabstractEvent detection is a recent and challenging task. The aim is to retrieve the relevant videos given an event description. A set of training examples associated with the events are generally provided as well, since retrieving relevant videos from textual queries solely is not feasible. Early attempts of event detection are based on low-level features. High level features such as concepts for event detection have been introduced as an alternative to low-level features since high-level features provide semantically richer information. In this work, we focus on object-based concepts and exploit their dependencies using a Markov Random Field (MRF) based model for event detection. This enables us to model likelihood of concepts, either pairwise or individually, present in the videos. Here, we propose a method incorporating the strengths of concepts and MRF based model for event detection task. We evaluate our models on an Multimedia Event Detection (MED) dataset from NIST's 2011 TRECVID Multimedia, which consists of approximately 45,000 unconstrained videos. This type of work is beneficial from several respects. First, we focus on the task of concept-based event detection using a very large number of unconstrained Youtube videos. Second, we introduce the application of MRF's for the event detection purpose, which can further be enhanced incorporating other features or temporal information. At last but not means least, we exploit the occurrence and co-occurrence of object based concepts for event detection that enables us to reveal interactions of such concepts in the video level. Experimental results show that revealing these interactions provide promising event detection results. Ethem F. Can, R. Manmatha |
ICMR | 2 |
| 2014 | A Hybrid Model for Automatic Image AnnotationabstractIn this work, we present a hybrid model (SVM-DMBRM) combining a generative and a discriminative model for the image annotation task. A support vector machine (SVM) is used as the discriminative model and a Discrete Multiple Bernoulli Relevance Model (DMBRM) is used as the generative model. The idea of combining both the models is to take advantage of the distinct capabilities of each model. The SVM tries to address the problem of poor annotation (images are not annotated with all relevant keywords), while the DMBRM model tries to address the problem of data imbalance (large variations in number of positive samples). Since DMBRM does not work well with high-dimensional data, a Latent Dirichlet Allocation (LDA) model is used to reduce the dimensionality of vector quantized features before using it. The hybrid model's results are comparable to or better than the state-of-the-art results on three standard datasets: Corel-5k, ESP-Game and IAPRTC-12. Venkatesh N. Murthy, Ethem F. Can, R. Manmatha |
ICMR | 3 |
| 2014 | Incorporating query-specific feedback into learning-to-rank modelsabstractRelevance feedback has been shown to improve retrieval for a broad range of retrieval models. It is the most common way of adapting a retrieval model for a specific query. In this work, we expand this common way by focusing on an approach that enables us to do query-specific modification of a retrieval model for learning-to-rank problems. Our approach is based on using feedback documents in two ways: 1) to improve the retrieval model directly and 2) to identify a subset of training queries that are more predictive than others. Experiments with the Gov2 collection show that this approach can obtain statistically significant improvements over two baselines; learning-to-rank (SVM-rank) with no feedback and learning-to-rank with standard relevance feedback. Ethem F. Can, W. Bruce Croft, R. Manmatha |
SIGIR | 3 |
| 2014 | Large scale document image retrieval by automatic word annotation
K. Pramod Sankar, R. Manmatha, C. V. Jawahar |
Int. J. Document Anal. Recognit. | 2 |
| 2014 | Special issue on Multimedia Event Detection
Thomas B. Moeslund, Omar Javed, Yu-Gang Jiang 0001, R. Manmatha |
Mach. Vis. Appl. | 4 |
| 2013 | Predicting retweet count using visual cuesabstractSocial media platforms allow rapid information diffusion, and serve as a source of information to many of the users. Particularly, in Twitter information provided by tweets diffuses over the users through retweets. Hence, being able to predict the retweet count of a given tweet is important for understanding and controlling information diffusion on Twitter. Since the length of a tweet is limited to 140 characters, extracting relevant features to predict the retweet count is a challenging task. However, visual features of images linked in tweets may provide predictive features. In this study, we focus on predicting the expected retweet count of a tweet by using visual cues of an image linked in that tweet in addition to content and structure-based features. Ethem F. Can, Hüseyin Oktay, R. Manmatha |
CIKM | 3 |
| 2013 | Creating an Improved Version Using Noisy OCR from Multiple EditionsabstractThis paper evaluates an automated scheme for aligning and combining optical character recognition (OCR) output from three scans of a book to generate a composite version with fewer OCR errors. While there has been some previous work on aligning multiple OCR versions of the same scan, the scheme introduced in this paper does not require that scans be from the same copy of the book, or even the same edition. The three OCR outputs are combined using an algorithm which builds upon an technique which aligns two sequences at a time. In the algorithm a multiple sequence alignment of the scans is generated by stitching together pair wise alignments and is used in turn to construct a corrected text. The approach is able to correct OCR errors so long as they do not occur in multiple scans. The proposed approach is shown to be effective even if some of the books contain additional content such as introductions or commentary. This scheme is used to generate improved versions from OCR texts taken from the Internet Archive. The accuracy of the original scans and the composite text are evaluated by comparing them to the version available from Project Gutenberg. David Wemhoener, Ismet Zeki Yalniz, R. Manmatha |
ICDAR | 3 |
| 2012 | An Efficient Framework for Searching Text in Noisy Document ImagesabstractAn efficient word spotting framework is proposed to search text in scanned books. The proposed method allows one to search for words when optical character recognition (OCR) fails due to noise or for languages where there is no OCR. Given a query word image, the aim is to retrieve matching words in the book sorted by the similarity. In the offline stage, SIFT descriptors are extracted over the corner points of each word image. Those features are quantized into visual terms (visterms) using hierarchical K-Means algorithm and indexed using an inverted file. In the query resolution stage, the candidate matches are efficiently identified using the inverted index. These word images are then forwarded to the next stage where the configuration of visterms on the image plane are tested. Configuration matching is efficiently performed by projecting the visterms on the horizontal axis and searching for the Longest Common Subsequence (LCS) between the sequences of visterms. The proposed framework is tested on one English and two Telugu books. It is shown that the proposed method resolves a typical user query under 10 milliseconds providing very high retrieval accuracy (Mean Average Precision 0.93). The search accuracy for the English book is comparable to searching text in the high accuracy output of a commercial OCR engine. Ismet Zeki Yalniz, R. Manmatha |
Document Analysis Systems | 2 |
| 2012 | On Influence of Line Segmentation in Efficient Word Segmentation in Old ManuscriptsabstractThe objective of this work is to show the importance of a good line segmentation to obtain better results in the segmentation of words of historical documents. We have used the approach developed by Manmatha and Rothfeder to segment words in old handwritten documents. In their work the lines of the documents are extracted using projections. In this work, we have developed an approach to segment lines more efficiently. The new line segmentation algorithm tackles with skewed, touching and noisy lines, so it is significantly improves word segmentation. Experiments using Spanish documents from the Marriages Database of the Barcelona Cathedral show that this approach reduces the error rate by more than 20%. Josep Lladós 0001, Alicia Fornés, R. Manmatha |
ICFHR | 4 |
| 2012 | A framework for manipulating and searching multiple retrieval typesabstractConventional retrieval systems view documents as a unit and look at different retrieval types within a document. We introduce Proteus, a frame-work for seamlessly navigating books as dynamic collections which are defined on the fly. Proteus allows us to search various retrieval types. Navigable types include pages, books, named persons, locations, and pictures in a collection of books taken from the Internet Archive. The demonstration shows the value of multi-type browsing in dynamic collections to peruse new data. Marc-Allen Cartright, Ethem F. Can, William Dabney, Jeff Dalton 0001, Logan Giorda, Kriste Krstovski, Xiaoye Wu, Ismet Zeki Yalniz, James Allan 0001, R. Manmatha, David A. Smith |
SIGIR | 10 |
| 2012 | Finding translations in scanned book collectionsabstractThis paper describes an approach for identifying translations of books in large scanned book collections with OCR errors. The method is based on the idea that although individual sentences do not necessarily preserve the word order when translated, a book must preserve the linear progression of ideas for it to be a valid translation. Consider two books in two different languages, say English and German. The English book in the collection is represented by the sequence of words (in the order they appear in the text) which appear only once in the book. Similarly, the book in German is represented by its sequence of words which appear only once. An English-German dictionary is used to transform the word sequence of the English book into German by translating individual words in place. It is not necessary to translate all the words and this method works even with small dictionaries. Both sequences are now in German and can, therefore, be aligned using a Longest Common Subsequence (LCS) algorithm. We describe two scoring functions TRANS-cs and TRANS-its which account for both the LCS length and the lengths of the original word sequences. Experiments demonstrate that TRANS-its is particularly successful in finding translations of books and outperforms several baselines including metadata search based on matching titles and authors. Experiments performed on a Europarl parallel corpus for four language pairs, English-Finnish, English-French, English-German, English-Spanish, and a scanned book collection of 50K English-German books show that the proposed method retrieves translations of books with an average MAP score of 1.0 and a speed of 10K book pair comparisons per second on a single core. Ismet Zeki Yalniz, R. Manmatha |
SIGIR | 2 |
| 2012 | A Novel Word Spotting Method Based on Recurrent Neural NetworksabstractKeyword spotting refers to the process of retrieving all instances of a given keyword from a document. In the present paper, a novel keyword spotting method for handwritten documents is described. It is derived from a neural network-based system for unconstrained handwriting recognition. As such it performs template-free spotting, i.e., it is not necessary for a keyword to appear in the training set. The keyword spotting is done using a modification of the CTC Token Passing algorithm in conjunction with a recurrent neural network. We demonstrate that the proposed systems outperform not only a classical dynamic time warping-based approach but also a modern keyword spotting system, based on hidden Markov models. Furthermore, we analyze the performance of the underlying neural networks when using them in a recognition task followed by keyword spotting on the produced transcription. We point out the advantages of keyword spotting when compared to classic text line recognition. Volkmar Frinken, Andreas Fischer 0002, R. Manmatha, Horst Bunke |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | Partial duplicate detection for large book collectionsabstractA framework is presented for discovering partial duplicates in large collections of scanned books with optical character recognition (OCR) errors. Each book in the collection is represented by the sequence of words (in the order they appear in the text) which appear only once in the book. These words are referred to as "unique words" and they constitute a small percentage of all the words in a typical book. Along with the order information the set of unique words provides a compact representation which is highly descriptive of the content and the flow of ideas in the book. By aligning the sequence of unique words from two books using the longest common subsequence (LCS) one can discover whether two books are duplicates. Experiments on several datasets show that DUPNIQ is more accurate than traditional methods for duplicate detection such as shingling and is fast. On a collection of 100K scanned English books DUPNIQ detects partial duplicates in 30 min using 350 cores and has precision 0.996 and recall 0.833 compared to shingling with precision 0.992 and recall 0.720. The technique works on other languages as well and is demonstrated for a French dataset. Ismet Zeki Yalniz, Ethem F. Can, R. Manmatha |
CIKM | 3 |
| 2011 | BLSTM Neural Network Based Word Retrieval for Hindi DocumentsabstractRetrieval from Hindi document image collections is a challenging task. This is partly due to the complexity of the script, which has more than 800 unique ligatures. In addition, segmentation and recognition of individual characters often becomes difficult due to the writing style as well as degradations in the print. For these reasons, robust OCRs are non existent for Hindi. Therefore, Hindi document repositories are not amenable to indexing and retrieval. In this paper, we propose a scheme for retrieving relevant Hindi documents in response to a query word. This approach uses BLSTM neural networks. Designed to take contextual information into account, these networks can handle word images that can not be robustly segmented into individual characters. By zoning the Hindi words, we simplify the problem and obtain high retrieval rates. Our simplification suits the retrieval problem, while it does not apply to recognition. Our scalable retrieval scheme avoids explicit recognition of characters. An experimental evaluation on a dataset of word images gathered from two complete books demonstrates good accuracy even in the presence of printing variations and degradations. The performance is compared with baseline methods. Raman Jain, Volkmar Frinken, C. V. Jawahar, R. Manmatha |
ICDAR | 4 |
| 2011 | A Fast Alignment Scheme for Automatic OCR Evaluation of BooksabstractThis paper aims to evaluate the accuracy of optical character recognition (OCR) systems on real scanned books. The ground truth e-texts are obtained from the Project Gutenberg website and aligned with their corresponding OCR output using a fast recursive text alignment scheme (RETAS). First, unique words in the vocabulary of the book are aligned with unique words in the OCR output. This process is recursively applied to each text segment in between matching unique words until the text segments become very small. In the final stage, an edit distance based alignment algorithm is used to align these short chunks of texts to generate the final alignment. The proposed approach effectively segments the alignment problem into small sub problems which in turn yields dramatic time savings even when there are large pieces of inserted or deleted text and the OCR accuracy is poor. This approach is used to evaluate the OCR accuracy of real scanned books in English, French, German and Spanish. Ismet Zeki Yalniz, R. Manmatha |
ICDAR | 2 |
| 2010 | Nearest neighbor based collection OCRabstractConventional optical character recognition (OCR) systems operate on individual characters and words, and do not normally exploit document or collection context. We describe a Collection OCR which takes advantage of the fact that multiple examples of the same word (often in the same font) may occur in a document or collection. The idea here is that an OCR or a reCAPTCHA like process generates a partial set of recognized words. In the second stage, a nearest neighbor algorithm compares the remaining word-images to those already recognized and propagates labels from the nearest neighbors. It is shown that by using an approximate fast nearest neighbor algorithm based on Hierarchical K-Means (HKM), we can do this accurately and efficiently. It is also shown that profile based features perform much better than SIFT and Pyramid Histogram of Gradient (PHOG) features. We believe that this is because profile features are more robust to word degradations (common in our documents). This approach is applied to a collection of Telugu books - a language for which no commercial OCR exists. We show from a selection of 33 Telugu books that starting with OCR labels for only 30% of the collection we can recognize the remaining 70% of the words in the collection with 70% accuracy using this approach. Since the approach makes no language specific assumptions, it should be applicable to a large number of languages. In particular we are interested in its applicability to Indic languages and scripts. K. Pramod Sankar, C. V. Jawahar, R. Manmatha |
Document Analysis Systems | 3 |
| 2010 | Adapting BLSTM Neural Network Based Keyword Spotting Trained on Modern Data to Historical DocumentsabstractBeing able to search for words or phrases in historic handwritten documents is of paramount importance when preserving cultural heritage. Storing scanned pages of written text can save the information from degradation, but it does not make the textual information readily available. Automatic keyword spotting systems for handwritten historic documents can fill this gap. However, most such systems have trouble with the great variety of writing styles. It is not uncommon for handwriting processing systems to be built for just a single book. In this paper we show that neural network based keyword spotting systems are flexible enough to be used successfully on historic data, even when they are trained on a modern handwriting database. We demonstrate that with little transcribed historic text, added to the training set, the performance can further be enhanced. Volkmar Frinken, Andreas Fischer 0002, Horst Bunke, R. Manmatha |
ICFHR | 4 |
| 2009 | Robust Recognition of Documents by Fusing Results of Word ClustersabstractThe word error rate of any optical character recognition system (OCR) is usually substantially below its component or character error rate. This is especially true of Indic languages in which a word consists of many components. Current OCRs recognize each character or word separately and do not take advantage of document level constraints. We propose a document level OCR which incorporates information from the entire document to reduce word error rates. Word images are first clustered using a locality sensitive hashing technique. Individual words are then recognized using a (regular) OCR. The OCR outputs of word images in a cluster are then corrected probabilistically by comparing with the OCR outputs of other members of the same cluster. The approach may be applied to improve the accuracy of any OCR run on documents in any language. In particular, we demonstrate it for Telugu, where the use of language models for post-processing is not promising. We show a relative improvement of 28% for long words and 12% for all words which appear at least twice in the corpus. Venkat Rasagna, Anand Kumar 0001, C. V. Jawahar, R. Manmatha |
ICDAR | 4 |
| 2009 | Finding words in alphabet soup: Inference on freeform character recognition for historical scripts
Nicholas R. Howe, Shaolei Feng 0001, R. Manmatha |
Pattern Recognit. | 3 |
| 2008 | Distributed image search in camera sensor networksabstractRecent advances in sensor networks permit the use of a large number of relatively inexpensive distributed computational nodes with camera sensors linked in a network and possibly linked to one or more central servers. We argue that the full potential of such a distributed system can be realized if it is designed as a distributed search engine where images from different sensors can be captured, stored, searched and queried. However, unlike traditional image search engines that are focused on resource-rich situations, the resource limitations of camera sensor networks in terms of energy, bandwidth, computational power, and memory capacity present significant challenges. In this paper, we describe the design and implementation of a distributed search system over a camera sensor network where each node is a search engine that senses, stores and searches information. Our work involves innovation at many levels including local storage, local search, and distributed search, all of which are designed to be efficient under the resource constraints of sensor networks. We present an implementation of the search engine on a network of iMote2 sensor nodes equipped with low-power cameras and extended flash storage. We evaluate our system for a dataset comprising book images, and demonstrate more than two orders of magnitude reduction in the amount of data communicated and up to 5x reduction in overall energy consumption over alternate techniques. Tingxin Yan, Deepak Ganesan, R. Manmatha |
SenSys | 3 |
| 2007 | Efficient Search in Document Image Collections
Anand Kumar 0001, C. V. Jawahar, R. Manmatha |
ACCV (1) | 3 |
| 2007 | Further explorations in text alignment with handwritten documents
E. Micah Kornfield, R. Manmatha, James Allan 0001 |
Int. J. Document Anal. Recognit. | 2 |
| 2007 | Word spotting for historical documents
Toni M. Rath, R. Manmatha |
Int. J. Document Anal. Recognit. | 2 |
| 2007 | Word spotting for historical documents
Toni M. Rath, R. Manmatha |
Int. J. Document Anal. Recognit. | 2 |
| 2006 | Aligning Transcripts to Automatically Segmented Handwritten Manuscripts
Jamie L. Rothfeder, R. Manmatha, Toni M. Rath |
Document Analysis Systems | 2 |
| 2005 | Combining text and audio-visual features in video indexingabstractWe discuss the opportunities, state of the art, and open research issues in using multi-modal features in video indexing. Specifically, we focus on how imperfect text data obtained by automatic speech recognition (ASR) may be used to help solve challenging problems, such as story segmentation, concept detection, retrieval, and topic clustering. We review the frameworks and machine learning techniques that are used to fuse the text features with audio-visual features. Case studies showing promising performance are described, primarily in the broadcast news video domain. Shih-Fu Chang, R. Manmatha, Tat-Seng Chua |
ICASSP (5) | 2 |
| 2005 | Classification Models for Historical Manuscript RecognitionabstractThis paper investigates different machine learning models to solve the historical handwritten manuscript recognition problem. In particular, we test and compare support vector machines, conditional maximum entropy models and Naive Bayes with kernel density estimates and explore their behaviors and properties when solving this problem. We focus on a whole word problem to avoid having to do character segmentation which is difficult with degraded handwritten documents. Our results on a publicly available standard dataset of 20 pages of George Washington's manuscripts show that Naive Bayes with Gaussian kernel density estimates significantly outperforms the other models and prior work using hidden Markov models on this heavily unbalanced dataset. Shaolei Feng 0001, R. Manmatha |
ICDAR | 2 |
| 2005 | Joint visual-text modeling for automatic retrieval of multimedia documentsabstractIn this paper we describe a novel approach for jointly modeling the text and the visual components of multimedia documents for the purpose of information retrieval(IR). We propose a novel framework where individual components are developed to model different relationships between documents and queries and then combined into a joint retrieval framework. In the state-of-the-art systems, a late combination between two independent systems, one analyzing just the text part of such documents, and the other analyzing the visual part without leveraging any knowledge acquired in the text processing, is the norm. Such systems rarely exceed the performance of any single modality (i.e. text or video) in information retrieval tasks. Our experiments indicate that allowing a rich interaction between the modalities results in significant improvement in performance over any single modality. We demonstrate these results using the TRECVID03 corpus, which comprises 120 hours of broadcast news videos. Our results demonstrate over 14 % improvement in IR performance over the best reported text-only baseline and ranks amongst the best results reported on this corpus. Giridharan Iyengar, Pinar Duygulu, Shaolei Feng 0001, Pavel Ircing, Sanjeev Khudanpur, Dietrich Klakow, M. R. Krause, R. Manmatha, Harriet J. Nock, D. Petkova, Brock Pytlik, Paola Virga |
ACM Multimedia | 8 |
| 2005 | Boosted decision trees for word recognition in handwritten document retrievalabstractRecognition and retrieval of historical handwritten material is an unsolved problem. We propose a novel approach to recognizing and retrieving handwritten manuscripts, based upon word image classification as a key step. Decision trees with normalized pixels as features form the basis of a highly accurate AdaBoost classifier, trained on a corpus of word images that have been resized and sampled at a pyramid of resolutions. To stem problems from the highly skewed distribution of class frequencies, word classes with very few training samples are augmented with stochastically altered versions of the originals. This increases recognition performance substantially. On a standard corpus of 20 pages of handwritten material from the George Washington collection the recognition performance shows a substantial improvement in performance over previous published results (75% vs 65%). Following word recognition, retrieval is done using a language model over the recognized words. Retrieval performance also shows substantially improved results over previously published results on this database. Recognition/retrieval results on a more challenging database of 100 pages from the George Washington collection are also presented. Nicholas R. Howe, Toni M. Rath, R. Manmatha |
SIGIR | 3 |
| 2005 | A Scale Space Approach for Automatically Segmenting Words from Historical Handwritten DocumentsabstractMany libraries, museums, and other organizations contain large collections of handwritten historical documents, for example, the papers of early presidents like George Washington at the Library of Congress. The first step in providing recognition/ retrieval tools is to automatically segment handwritten pages into words. State of the art segmentation techniques like the gap metrics algorithm have been mostly developed and tested on highly constrained documents like bank checks and postal addresses. There has been little work on full handwritten pages and this work has usually involved testing on clean artificial documents created for the purpose of research. Historical manuscript images, on the other hand, contain a great deal of noise and are much more challenging. Here, a novel scale space algorithm for automatically segmenting handwritten (historical) documents into words is described. First, the page is cleaned to remove margins. This is followed by a gray-level projection profile algorithm for finding lines in images. Each line image is then filtered with an anisotropic Laplacian at several scales. This procedure produces blobs which correspond to portions of characters at small scales and to words at larger scales. Crucial to the algorithm is scale selection, that is, finding the optimum scale at which blobs correspond to words. This is done by finding the maximum over scale of the extent or area of the blobs. This scale maximum is estimated using three different approaches. The blobs recovered at the optimum scale are then bounded with a rectangular box to recover the words. A postprocessing filtering step is performed to eliminate boxes of unusual size which are unlikely to correspond to words. The approach is tested on a number of different data sets and it is shown that, on 100 sampled documents from the George Washington corpus of handwritten document images, a total error rate of 17 percent is observed. The technique outperforms a state-of-the-art gap metrics word-segmentation algorithm on this collection. R. Manmatha, Jamie L. Rothfeder |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | Multiple Bernoulli Relevance Models for Image and Video Annotation
Shaolei Feng 0001, R. Manmatha, Victor Lavrenko |
CVPR (2) | 2 |
| 2004 | Statistical models for automatic video annotation and retrievalabstractWe apply a continuous relevance model (CRM) to the problem of directly retrieving the visual content of videos using text queries. The model computes a joint probability model for image features and words using a training set of annotated images. The model may then be used to annotate unseen test images. The probabilistic annotations are used for retrieval using text queries. We also propose a modified model - the normalized CRM - which substantially improves performance on a subset of the TREC video dataset. Victor Lavrenko, Shaolei Feng 0001, R. Manmatha |
ICASSP (3) | 3 |
| 2004 | A search engine for historical manuscript imagesabstractMany museum and library archives are digitizing their large collections of handwritten historical manuscripts to enable public access to them. These collections are only available in image formats and require expensive manual annotation work for access to them. Current handwriting recognizers have word error rates in excess of 50% and therefore cannot be used for such material. We describe two statistical models for retrieval in large collections of handwritten manuscripts given a text query. Both use a set of transcribed page images to learn a joint probability distribution between features computed from word images and their transcriptions. The models can then be used to retrieve unlabeled images of handwritten documents given a text query. We show experiments with a training set of 100 transcribed pages and a test set of 987 handwritten page images from the George Washington collection. Experiments show that the precision at 20 documents is about 0.4 to 0.5 depending on the model. To the best of our knowledge, this is the first automatic retrieval system for historical manuscripts using text queries, without manual transcription of the original corpus. Toni M. Rath, R. Manmatha, Victor Lavrenko |
SIGIR | 2 |
| 2003 | Word Image Matching Using Dynamic Time WarpingabstractLibraries and other institutions are interested in providing access to scanned versions of their large collections of handwritten historical manuscripts on electronic media. Convenient access to a collection requires an index, which is manually created at great labor and expense. Since current handwriting recognizers do not perform well on historical documents, a technique called word spotting has been developed: clusters with occurrences of the same word in a collection are established using image matching. By annotating "interesting" clusters, an index can be built automatically. We present an algorithm for matching handwritten words in noisy historical documents. The segmented word images are preprocessed to create sets of 1-dimensional features, which are then compared using dynamic time warping. We present experimental results on two different data sets from the George Washington collection. Our experiments show that this algorithm performs better and is faster than competing matching techniques. Toni M. Rath, R. Manmatha |
CVPR (2) | 2 |
| 2003 | Features for Word Spotting in Historical ManuscriptsabstractFor the transition from traditional to digital libraries, the large number of handwritten manuscripts that exist pose a great challenge. Easy access to such collections requires an index, which is currently created manually at great cost. Because automatic handwriting recognizers fail on historical manuscripts, the word spotting technique has been developed: the words in a collection are matched as images and grouped into clusters which contain all instances of the same word. By annotating "interesting" clusters, an index that links words to the locations where they occur can be built automatically. Due to the noise in historical documents, selecting the right features for matching words is crucial. We analyzed a range of features suitable for matching words using dynamic time warping (DTW), which aligns and compares sets of features extracted from two images. Each feature's individual performance was measured on a test set. With an average precision of 72%, a combination of features outperforms competing techniques in speed and precision. Toni M. Rath, R. Manmatha |
ICDAR | 2 |
| 2003 | Mobile Distributed Information Retrieval for Highly-Partitioned NetworksabstractWe propose and evaluate a mobile, peer-to-peer information retrieval system. Such a system can, for example, support medical care in a disaster by allowing access to a large collections of medical literature. In our system, documents in a collection are replicated in an overlapping manner at mobile peers. This provides resilience in the face of node failures, malicious attacks, and network partitions. We show that our design manages the randomness of node mobility. Although nodes contact only direct neighbors (who change frequently) and do not use any ad hoc routing, the system maintains good IR performance. This makes our design applicable to mobility situations where routing partitions are common. Our evaluation shows that our scheme provides significant savings in network costs, and increased access to information over ad-hoc routing-based approaches; nodes in our system require only a modest amount of additional storage on average. Katrina M. Hanna, Brian Neil Levine, R. Manmatha |
ICNP | 3 |
| 2003 | A Model for Learning the Semantics of PicturesabstractWe propose an approach to learning the semantics of images which al- lows us to automatically annotate an image with keywords and to retrieve images based on text queries. We do this using a formalism that models the generation of annotated images. We assume that every image is di- vided into regions, each described by a continuous-valued feature vector. Given a training set of images with annotations, we compute a joint prob- abilistic model of image features and words which allow us to predict the probability of generating a word given the image regions. This may be used to automatically annotate and retrieve images given a word as a query. Experiments show that our model significantly outperforms the best of the previously reported results on the tasks of automatic image annotation and retrieval. Victor Lavrenko, R. Manmatha, Jiwoon Jeon |
NIPS | 2 |
| 2003 | Automatic image annotation and retrieval using cross-media relevance modelsabstractLibraries have traditionally used manual image annotation for indexing and then later retrieving their image collections. However, manual image annotation is an expensive and labor intensive procedure and hence there has been great interest in coming up with automatic ways to retrieve images based on content. Here, we propose an automatic approach to annotating and retrieving images based on a training set of images. We assume that regions in an image can be described using a small vocabulary of blobs. Blobs are generated from image features using clustering. Given a training set of images with annotations, we show that probabilistic models allow us to predict the probability of generating a word given the blobs in an image. This may be used to automatically annotate and retrieve images given a word as a query. We show that relevance models allow us to derive these probabilities in a natural way. Experiments show that the annotation performance of this cross-media relevance model is almost six times as good (in terms of mean precision) than a model based on word-blob co-occurrence model and twice as good as a state of the art model derived from machine translation. Our approach shows the usefulness of using formal information retrieval models for the task of image annotation and retrieval. Jiwoon Jeon, Victor Lavrenko, R. Manmatha |
SIGIR | 3 |
| 2002 | A critical examination of TDT's cost functionabstractTopic Detection and Tracking (TDT) tasks are evaluated using a cost function. The standard TDT cost function assumes a constant probability of relevance P(rel) across all topics. In practice, P(rel) varies widely across topics. We argue using both theoretical and experimental evidence that the cost function should be modified to account for the varying P(rel). R. Manmatha, Ao Feng, James Allan 0001 |
SIGIR | 1 |
| 2001 | Automatic Segmentation and Indexing in a Database of Bird ImagesabstractThe aim of this work is to index images in domain specific databases using colors computed from the object of interest only, instead of the whole image. The main problem in this task is the segmentation of the region of interest from the background. Viewing segmentation as a figure/ground segregation problem leads to a new approach-eliminating the background leaves the figure or object of interest. To find possible object colors, we first find background colors and eliminate them. We then use an edge image at an appropriate scale to eliminate those parts of the image which are not in focus and do not contain significant structures. The edge information is combined with the color-based background elimination to produce object (figure) regions. We test our approach on a database of bird images. We show that in 87% or 450 bird images tested, the segmentation is sufficient to determine the colors of the bird correctly for retrieval purposes. We also show that our approach provides improved retrieval performance. Madirakshi Das, R. Manmatha |
ICCV | 2 |
| 2001 | Modeling Score Distributions for Combining the Outputs of Search EnginesabstractIn this paper the score distributions of a number of text search engines are modeled. It is shown empirically that the score distributions on a per query basis may be fitted using an exponential distribution for the set of non-relevant documents and a normal distribution for the set of relevant documents. Experiments show that this model fits TREC-3 and TREC-4 data for not only probabilistic search engines like INQUERY but also vector space search engines like SMART for English. We have also used this model to fit the output of other search engines like LSI search engines and search engines indexing other languages like Chinese. R. Manmatha, Toni M. Rath, Fangfang Feng |
SIGIR | 1 |
| 1999 | TextFinder: An Automatic System to Detect and Recognize Text In ImagesabstractA robust system is proposed to automatically detect and extract text in images from different sources, including video, newspapers, advertisements, stock certificates, photographs, and checks. Text is first detected using multiscale texture segmentation and spatial cohesion constraints, then cleaned up and extracted using a histogram-based binarization algorithm. An automatic performance evaluation scheme is also proposed. Victor Wu, R. Manmatha, Edward M. Riseman |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1998 | Retrieving Images by AppearanceabstractA system to retrieve images using a description of the image intensity surface is presented. Gaussian derivative filters at several scales are applied to the image and low order 2D differential invariants are computed. The resulting multi-scale representation is indexed for rapid retrieval. Queries are designed by the users from an example image by selecting appropriate regions. The invariant vectors corresponding to these regions are matched with the database counterparts both in feature and coordinate space. This yields a match score per image. Images are sorted by the match score and displayed. Experiments conducted with over 1500 images of objects embedded in arbitrary backgrounds are described. It is observed that images similar in appearance and whose viewpoint is within small view variations of the query can be retrieved with an average precision of 50%. S. Chandu Ravela, R. Manmatha |
ICCV | 2 |
| 1998 | Indexing flowers by color names using domain knowledge-driven segmentationabstractWe describe a solution to the problem of indexing images of flowers for searching a flower patents database by color. We use a natural language color classification derived from the ISCC-NBS color system and the X Window color names to effectively use the domain knowledge available, provide perceptually correct retrieval and allow natural language queries. We have developed an automatic iterative segmentation algorithm with knowledge-driven feedback to isolate a flower region from the background. The color of the flower is defined by the color names present in the flower region and their relative proportions. The database can be queried by example and by color names. We demonstrate the effectiveness of the strategy on a test database. Madirakshi Das, R. Manmatha, Edward M. Riseman |
WACV | 2 |
| 1998 | On computing global similarity in imagesabstractThe retrieval of images based on their visual similarity to an example image is an important and fascinating area of research. Here, a method to characterize visual appearance for determining global similarity in images is described. Images are filtered with Gaussian derivatives and geometric features are computed from the filtered images. The geometric features used here are curvature and phase. Two images may be said to be similar if they have similar distributions of such features. Global similarity may, therefore, be deduced by comparing histograms of these features. This allows for rapid retrieval and examples from collection of gray-level and trademark images are shown. S. Chandu Ravela, R. Manmatha |
WACV | 2 |
| 1997 | Image Retrieval by AppearanceabstractA system to retrieve images using a syntactic description of appearance is presented.A multi-scale invariant wxtor representation is obtained by first filtering images in the database with Gaussian derivative filters at several acalea and then computing low order differential invariants.The multi-scale representation is indexed for rapid retrieval.Queries are designed by the users horn an example image by eelecting appropriate regions.The invariant xnxtors corrmponding to these regions are matched with those in the database both in feature space as well M in coordinate space and a match score is obtained for each image.The results are then displayed to the user sorted by the match score.I?kom experiments conducted with ovsr 1500 images it is shown that imagee similar in appearance and whose viewpoint is within 25 degrees of the query image can be retrieved with an average precision of 57% S. Chandu Ravela, R. Manmatha |
SIGIR | 2 |
| 1996 | Word Spotting: A New Approach to Indexing HandwritingabstractThere are many historical manuscripts written in a single hand which it would be useful to index. Examples include the W.B. DuBois collection at the University of Massachusetts and the early Presidential libraries at the Library of Congress. Since Optical Character Recognition (OCR) does not work well on handwriting, an alternative scheme based on matching the images of the words is proposed for indexing such texts. The current paper deals with the matching aspects of this process. Two different techniques for matching words are discussed. The first method matches words assuming that the transformation between the words may be modelled by a translation (shift). The second method matches words assuming that the transformation between the words may be modelled by an affine transform. Experiments are shown demonstrating the feasibility of the approach for indexing handwriting. The method should also be applicable to retrieving previously stored material from personal digital assistants (PDAs). R. Manmatha, Chengfeng Han, Edward M. Riseman |
CVPR | 1 |
| 1996 | Image Retrieval Using Scale-Space Matching
S. Chandu Ravela, R. Manmatha, Edward M. Riseman |
ECCV (1) | 2 |
| 1994 | A framework for recovering affine transforms using points, lines or image brightnessesabstractImage deformations due to the relative motion between an observer and an object may be used to infer 3D structure. Up to first order, these deformations can be written in terms of an affine transform. A new framework for measuring affine transforms, which correctly handles the problem of corresponding deformed patches, is presented. In this framework, points, lines or image brightnesses may be used to derive the affine transform between image patches. No correspondence is required. The patches are filtered using Gaussians and derivatives of Gaussians, and the filters are deformed according to the affine transform. The problem of finding the affine transform is therefore reduced to that of finding the appropriate deformed filter to use. The method is local and can handle large affine deformations. Experiments demonstrate that this technique can find scale changes and optical flow in situations where other methods fail.> R. Manmatha |
CVPR | 1 |
| 1994 | Measuring the Affine Transform Using Gaussian Filters
R. Manmatha |
ECCV (2) | 1 |
| 1993 | Extracting affine deformations from image patches. I. Finding scale and rotationabstractImage deformations due to relative motion between an observer and an object may be used to infer 3-D structure. Up to the first order, these deformations can be written in terms of an affine transform. A novel approach is adopted to measuring affine transforms which correctly handles the problem of corresponding deformed patches. The patches are filtered using Gaussians and derivatives of Gaussians and the filters deformed according to the affine transform. The problem of finding the affine transform is therefore reduced to that of finding the appropriate deformed filter to use. In the special case where the affine transform can be written as a scale change and an in-plane rotation, the Gaussian and first derivative equations are solved for the scale. The robustness of the method is demonstrated experimentally.> R. Manmatha, John Oliensis |
CVPR | 1 |
| 1989 | A data set for quantitative motion analysisabstractThe collection of image sequences for quantitative experiments in motion analysis is discussed. To assess the effectiveness of a motion algorithm it is necessary to obtain motion data with ground truth of known accuracy. The authors describe motion data collected using the imaging facilities provided by the autonomous land vehicle (ALV). A total of eight sequences of about 30 frames each were collected at five different outdoor sites using move-and-shoot and stop-and-shoot methods. For all the sequences accurate ground truth of environmental objects were determined using theodolites and a laser range finder. The camera egomotion parameters (this included both translation and rotation) were determined using a land navigation system on the ALV. Example images of the data set collected and an analysis of one of the sequences are presented.> Rabindranath Dutta, R. Manmatha, Lance R. Williams, Edward M. Riseman |
CVPR | 2 |