EDBT 2026 Demo / reviewers in the wild / expert
Gaurav Harit
dblp:58/4372
· DBLP profile ↗
10ranked-venue papers in the field
5as first author
2since 2021 · last 2025
0000-0001-7943-0123ORCID · corroborated
Domains — venue-derived; a paper can count in several
Other / Interdisciplinary · 10 (5 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Table Detection with Active Learning
Somraj Gautam, Nachiketa Purohit, Gaurav Harit |
ICDAR (3) | 3 |
| 2023 | TransDocAnalyser: A Framework for Semi-structured Offline Handwritten Documents Analysis with an Application to Legal Domain
Sagar Chakraborty, Gaurav Harit |
ICDAR (1) | 2 |
| 2017 | Core Region Detection for Off-Line Unconstrained Handwritten Latin Words Using Word EnvelopsabstractZone extraction is acclaimed as a significant pre-processing step in handwriting analysis. This paper presents a new method for separating ascenders and descenders from an unconstrained handwritten word and identifying its core-region. The method estimates correct core-region for complexities like long horizontal strokes, skewed words, first letter capital, hill and dale writing, jumping baselines and words with long descender curves, cursive handwriting, calligraphic words, title case words, very short words as shown in Fig. 1. It extracts two envelops from the word image and selects sample points that constitute the core region envelop. The method is tested on CVL, ICDAR-2013, ICFHR-2012, and IAM benchmark datasets of handwritten words written by multiple writers. We also created our own dataset of 100 words authored by 2 writers comprising all the above mentioned handwriting complexities. Due to non-availability of the Ground Truth for core-region extraction we created it manually for all the datasets. Our work reports an accuracy of 90.16% for correctly identifying all the three zones on 17,100 Latin words written by 802 individuals. Promising results are obtained by our core-region detection method when compared with the current state of the art methods. Shilpa Pandey, Gaurav Harit |
ICDAR | 2 |
| 2007 | Word image based latent semantic indexing for conceptual querying in document image databasesabstractIn this paper we present an application of latent semantic analysis (LSA) for indexing and retrieval of document images with text. The query is specified as a set of word images and the documents which best match with the query representation in the the latent semantic space are retrieved. We show through extensive experiments on a large database that use of LSA for document images provides improvements in retrieval precision as is the case with electronic text documents. Sameek Banerjee, Gaurav Harit, Santanu Chaudhury |
ICDAR | 2 |
| 2007 | Pàtrà: A Novel Document Architecture for Integrating Handwriting with Audio-Visual InformationabstractIn this paper we present P`atr`a - an integrated docu- ment architecture which incorporates handwritten illustra- tions captured and rendered in a temporal fashion synchro- nized with audio, video, text, and image data. The architec- ture of P`atr`a permits non-linear growth in the form of mul- tiple hierarchically organized play streams. Semantic meta- data is also an integral part of P`atr`a which serves a useful purpose of organizing such documents in a collection. We have developed an email application in which the users are provided with an authoring and rendering environment to compose, view, and reply to messages in the form of P`atr`a. Gaurav Harit, V. Mankar, Santanu Chaudhury |
ICDAR | 1 |
| 2006 | Using Multimedia Ontology for Generating Conceptual Annotations and Hyperlinks in Video CollectionsabstractTo enable seamless integration of video information on the semantic Web, we require that the knowledge of a video domain be formally specified in ontology. We present a novel approach for defining video domain concepts in an ontology using properties that can be observed from the media. We use the ontology specified knowledge for recognizing concepts relevant to a video scene by making observations for the media properties of concepts as well as making inferences from other ontological concept definitions and relations. For this purpose we introduce new language constructs to OWL (Web Ontology Language), which are used to specify the inherently uncertain nature of media observations. The new constructs also allow additional semantics concerned with the association of media properties with concepts. We propose the use of Bayesian network as the reasoning mechanism for doing inferencing tasks in the presence of uncertainty. The video is annotated with the relevant concepts defined in the ontology. These conceptual annotations are used to create hyperlinks in the video collection Gaurav Harit, Santanu Chaudhury, Hiranmay Ghosh |
Web Intelligence | 1 |
| 2005 | Ontology Guided Access to Document ImagesabstractIn this paper, we propose a scheme for accessing document images using ontology. We make use of an extension of OWL (ontology language for Web) to allow encoding of ontologies for document images. We experimentally demonstrate that reasoning with the concepts defined in ontology and their observation models provide a mechanism to support conceptual querying and automated hyperlinking of document images. Gaurav Harit, Santanu Chaudhury, Jagrati Paranjpe |
ICDAR | 1 |
| 2005 | Improved Geometric Feature Graph: A Script Independent Representation of Word Images for Compression, and RetrievalabstractIn this paper, we discuss a new representation scheme for word images which exploits the structural features. The word image features are represented in the form of a graph called as the geometric feature graph (GFG). The GFG is encoded in the form of a string which serves as a compressed representation of the word image skeleton. We demonstrate reconstruction, and retrieval of word images for 3 different scripts using the GFG string. Gaurav Harit, Richa Jain, Santanu Chaudhury |
ICDAR | 1 |
| 2003 | Devising Interactive Access Techniques for Indian Language Document Imagesabstractexist only in paper form. Web based interactive access techniques for images of these documents can ensure wider dissemination and easy availability. In this paper, we have proposed an access mechanism based on word based indexing and personalized annotation. The word based indexing scheme exploits typical structural characteristics of Indian scripts. We have combined this word indexing technique with personalized annotation based hyperlinking and query scheme for providing an interactive access interface to a collection of Indian language documents. Santanu Chaudhury, Geetika Sethi, Anand Vyas, Gaurav Harit |
ICDAR | 4 |
| 2001 | A Model Guided Document Image Analysis SchemeabstractThis paper presents a new model-based document image segmentation scheme that uses XML-DTDs (eXtensible Markup Language Document Type Definitions). Given a document image, the algorithm has the ability to select the appropriate model. A new wavelet-based tool has been designed for distinguishing text from non-text regions and characterization of font sizes. Our model-based analysis scheme makes use of this tool for identifying the logical components of a document image. Gaurav Harit, Santanu Chaudhury, Neeti Vohra, Shiv Dutt Joshi |
ICDAR | 1 |