Ajoy Mondal

dblp:138/1190 · DBLP profile ↗
← Back
43ranked-venue papers
20as first author
29since 2021 · last 2026
0000-0002-4808-8860ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 13 first-author · 21 since 2021Databases, data management, data science and information retrieval · 17 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Can VLMs Understand Handwritten Mathematical Documents?
Shree Mitra, Ajoy Mondal, C. V. Jawahar
ICDAR (3)2
2026 Learning Beyond Labels: Self-Supervised Handwritten Text Recognition
abstract
This paper addresses a key challenge in Handwritten Text Recognition (HTR): the dependence on large volumes of labeled data. To overcome this, we propose a self-supervised learning (SSL) framework, LoGo-HTR, that minimizes labeling requirements while achieving strong recognition performance. We introduce a large-scale dataset, SSL-HWD of 10 million word-level handwritten images from diverse scanned documents, partitioned into a small labeled subset and a much larger unlabeled subset.LoGo-HTR combines a local contrastive loss for spatial consistency and a global decorrelation loss to enhance feature diversity. This dual objective enables robust, invariant, and spatially discriminative feature learning. After self-supervised pretraining, we fine-tune a transformer-based decoder using limited labeled data. Extensive experiments on standard HTR benchmarks, which include multilingual and historical data, demonstrate that, after SSL pretraining on our unlabeled dataset, our method consistently outperforms state-of-the-art approaches, even when fine-tuned using only 80% and 20% of the available labeled training data from the respective benchmarks. Ablation studies highlight the effectiveness of our dual loss design and demonstrate the potential of scalable, label-efficient handwritten text recognition. The SSL-HWD dataset and LoGoHTR model with code are publicly available at https://logo-ssl.github.io/.
Shree Mitra, Ajoy Mondal, C. V. Jawahar
WACV2
2026 UniTabBank: A Large Scale Multi-Lingual, Multi-Layout, Multi-Type, Multi-Format Dataset for Table Detection
abstract
Tables play a key role in conveying structured data across documents. Accurate table detection is crucial for downstream tasks like structure recognition and information extraction. However, current datasets lack diversity in format, language, and layout, limiting real-world generalization. This underscores the need for well-annotated datasets that are multi-lingual, layout-diverse, document-agnostic, and format-richTo address these limitations, we introduce UniTabBank, a large scale, diverse table detection dataset designed to reflect realistic use cases. UniTabBank is characterized by five key attributes: (i) Multi-Lingual — supporting 28 languages (including Arabic, English, Hindi, etc.); (ii) Multi-Layout — encompassing both single-column and multi-column documents; (iii) Multi-Type — covering a wide range of document genres such as annual reports, books, newspapers, and magazines; (iv) Multi-Format — comprising scanned documents, photographed pages, and PDFs; and finally (v) Scale and Annotation Quality — consists of 55,443 document page images with 82,114 accurately annotated table instances, offering scale and annotation precisionAdditionally, we introduce UniTabDet, a YOLO-based model for table detection, which outperforms state-of-the-arts on eight out of nine table detection benchmarks. Cross-benchmark evaluation highlights the strong generalization capability of UniTabBank compared to existing benchmarks. The dataset and models are available here.
Ajoy Mondal, Saumya Mundra, Avijit Dasgupta, C. V. Jawahar
WACV1
2026 MIST: Multilingual Incidental Dataset for Scene Text Detection
abstract
Scene text detection has progressed rapidly, largely driven by curated datasets and benchmarks. However, many of these have reached evaluation saturation and are heavily biased toward focused scenes, limiting their effectiveness in real-world environments where detection is hindered by environmental factors. To address this, we introduce MIST – a Multilingual Incidental Scene Text dataset featuring diverse text instances across 11 languages. MIST provides language, legibility, and fine-grained polygon-shaped annotations across 12K scene images and 600K word-level text instances. Images are captured along roads using a GoPro mounted on a moving car to capture real-world complexities, ensuring the scenes are incidental rather than deliberately framed. MIST establishes a new challenging benchmark to enable robust evaluation of scene text detection methods in real-world scenarios. The datasets and code are available at https://saumya-svm.github.io/mist/.
Saumya Mundra, Ajoy Mondal, C. V. Jawahar
WACV2
2026 HW-MLVQA: a novel handwritten multilingual dataset for visual question answering and evaluation
Aniket Pal, Ajoy Mondal, C. V. Jawahar
Int. J. Document Anal. Recognit.2
2026 From pixels to tables: reconstructing complex tables from document images
Sachin Raja, Ajoy Mondal, C. V. Jawahar
Int. J. Document Anal. Recognit.2
2025 Adapting Vision-Language Models for Hindi OCR
Shaon Bhattacharyya, Prantik Deb, Ajoy Mondal, C. V. Jawahar
ICDAR (3)4
2025 AI-Generated Lecture Slides for Improving Slide Element Detection and Retrieval
Suyash Maniyar, Vishvesh Trivedi, Ajoy Mondal, Anand Mishra 0001, C. V. Jawahar
ICDAR (1)3
2025 ICDAR 2025 Handwritten Notes Understanding Challenge
Aniket Pal, Sanket Biswas, Alloy Das, Ayush Lodh, Priyanka Banerjee, Soumitri Chattopadhyay, Ajoy Mondal, Dimosthenis Karatzas, Josep Lladós 0001, C. V. Jawahar
ICDAR (5)7
2025 EviFiVQA: A Benchmark for Evidence-Grounded Multi-hop Reasoning in Financial VQA
Sachin Raja, Ajoy Mondal, C. V. Jawahar
ICDAR (4)2
2025 UniLayDet: Simple Multi-dataset Document Layout Analysis
Prasidh Srikumar, Ajoy Mondal, C. V. Jawahar
ICDAR (1)2
2024 ICDAR 2024 Competition on Reading Documents Through Aria Glasses
Soumya Jahagirdar, Ajoy Mondal, Yuheng (Carl) Ren, Omkar M. Parkhi, C. V. Jawahar
ICDAR (6)2
2024 ICDAR 2024 Competition on Recognition and VQA on Handwritten Documents
Ajoy Mondal, Vijay Mahadevan, R. Manmatha, C. V. Jawahar
ICDAR (6)1
2024 Bridging the Gap in Resource for Offline English Handwritten Text Recognition
Ajoy Mondal, Krishna Tulsyan, C. V. Jawahar
ICDAR (2)1
2024 Indic Scene Text on the Roadside
Ajoy Mondal, Krishna Tulsyan, C. V. Jawahar
ICDAR (5)1
2024 CHART-Info 2024: A Dataset for Chart Analysis and Recognition
Kenny Davila, Rupak Lazarus, Nicole Rodríguez Alcántara, Srirangaraj Setlur, Venu Govindaraju, Ajoy Mondal, C. V. Jawahar
ICPR (19)7
2024 ICPR 2024 Competition on Word Image Recognition from Indic Scene Images
Harsh Lunia, Ajoy Mondal, C. V. Jawahar
ICPR (34)2
2024 Towards Deployable OCR Models for Indic Languages
Minesh Mathew, Ajoy Mondal, C. V. Jawahar
ICPR (19)2
2024 Unconstrained Camera Captured Indic Offline Handwritten Dataset
Ajoy Mondal, C. V. Jawahar
ICPR (19)1
2023 ICDAR 2023 Competition on Indic Handwriting Text Recognition
Ajoy Mondal, C. V. Jawahar
ICDAR (2)1
2023 ICDAR 2023 Competition on Visual Question Answering on Business Document Images
Sachin Raja, Ajoy Mondal, C. V. Jawahar
ICDAR (2)2
2023 Deep semantic binarization for document images
Ajoy Mondal, Chetan Reddy 0002, C. V. Jawahar
Multim. Tools Appl.1
2023 Dataset agnostic document object detection
Ajoy Mondal, Madhav Agarwal, C. V. Jawahar
Pattern Recognit.1
2022 Enhancing Indic Handwritten Text Recognition Using Global Semantic Information
Ajoy Mondal, C. V. Jawahar
ICFHR1
2022 Visual Understanding of Complex Table Structures from Document Images
abstract
Table structure recognition is necessary for a comprehensive understanding of documents. Tables in unstructured business documents are tough to parse due to the high diversity of layouts, varying alignments of contents, and the presence of empty cells. The problem is particularly difficult because of challenges in identifying individual cells using visual or linguistic contexts or both. Accurate detection of table cells (including empty cells) simplifies structure extraction and hence, it becomes the prime focus of our work. We propose a novel object-detection-based deep model that captures the inherent alignments of cells within tables and is fine-tuned for fast optimization. Despite accurate detection of cells, recognizing structures for dense tables may still be challenging because of difficulties in capturing long-range row/column dependencies in presence of multi-row/column spanning cells. Therefore, we also aim to improve structure recognition by deducing a novel rectilinear graph-based formulation. From a semantics perspective, we highlight the significance of empty cells in a table. To take these cells into account, we suggest an enhancement to a popular evaluation criterion. Finally, we introduce a modestly sized evaluation dataset with an annotation style inspired by human cognition to encourage new approaches to the problem. Our framework improves the previous state-of-the-art performance by a 2.7% average F1-score on benchmark datasets.
Sachin Raja, Ajoy Mondal, C. V. Jawahar
WACV2
2022 Deep neural networks for automatic grain-matrix segmentation in plane and cross-polarized sandstone photomicrographs
Rajdeep Das, Ajoy Mondal, Tapan Chakraborty, Kuntal Ghosh
Appl. Intell.2
2022 Camouflage design, assessment and breaking techniques: a survey
Ajoy Mondal
Multim. Syst.1
2021 Occluded object tracking using object-background prototypes and particle filter
Ajoy Mondal
Appl. Intell.1
2021 New performance measures for object tracking under complex environments
Ajoy Mondal
Multim. Syst.1
2020 Graph Representation Ensemble Learning
abstract
Representation learning on graphs has been gaining attention due to its wide applicability in predicting missing links and classifying and recommending nodes. Most embedding methods aim to preserve specific properties of the original graph in the low dimensional space. However, real-world graphs have a combination of several features that are difficult to characterize and capture by a single approach. In this work, we introduce the problem of graph representation ensemble learning and provide a first of its kind framework to aggregate multiple graph embedding methods efficiently. We provide analysis of our framework and analyze - theoretically and empirically - the dependence between state-of-the-art embedding methods. We test our models on the node classification task on four realworld graphs and show that proposed ensemble approaches can outperform the state-of-the-art methods by up to 20% on macro-F1. We further show that the strategy is even more beneficial for underrepresented classes with an improvement of up to 40%.
Palash Goyal, Sachin Raja, Sujit Rokka Chhetri, Arquimedes Canedo, Ajoy Mondal, Jaya Shree, C. V. Jawahar
ASONAM6
2020 IIIT-AR-13K: A New Dataset for Graphical Object Detection in Documents
Ajoy Mondal, Peter Lipps, C. V. Jawahar
DAS1
2020 A Benchmark System for Indian Language Text Recognition
Krishna Tulsyan, Nimisha Srivastava, Ajoy Mondal, C. V. Jawahar
DAS3
2020 Table Structure Recognition Using Top-Down and Bottom-Up Cues
Sachin Raja, Ajoy Mondal, C. V. Jawahar
ECCV (28)2
2020 CDeC-Net: Composite Deformable Cascade Network for Table Detection in Document Images
abstract
Localizing page elements/objects such as tables, figures, equations, etc. is the primary step in extracting information from document images. We propose a novel end-to-end trainable deep network, (cnec-xet) for detecting tables present in the documents. The proposed network consists of a multistage extension of Mask R-CNN with a dual backbone having deformable convolution for detecting tables varying in scale with high detection accuracy at higher IoU threshold. We empirically evaluate CDeC-Net on the publicly available benchmark datasets with extensive experiments. Our solution has three important properties: (i) a single trained model CDeC-Net‡that performs well across all the popular benchmark datasets; (ii) we report excellent performances across multiple, including higher, thresholds of IoU; (iii) by following the same protocol of the recent papers for each of the benchmarks, we consistently demonstrate the superior quantitative performance. Our code and models are publicly available at https://github.com/mdv3101/CDeCNet for enabling reproducibility of the results.
Madhav Agarwal, Ajoy Mondal, C. V. Jawahar
ICPR2
2020 Fuzzy energy based active contour model for multi-region image segmentation
Ajoy Mondal
Multim. Tools Appl.1
2020 State-of-the-art fuzzy active contour models for image segmentation
Ajoy Mondal, Kuntal Ghosh
Soft Comput.1
2019 Textual Description for Mathematical Equations
abstract
Reading of mathematical expression or equation in the document images is very challenging due to the large variability of mathematical symbols and expressions. In this paper, we pose reading of mathematical equation as a task of the generation of the textual description which interprets the internal meaning of this equation. Inspired by the natural image captioning problem in computer vision, we present a mathematical equation description ( MED ) model, a novel end-to-end trainable deep neural network based approach that learns to generate a textual description for reading mathematical equation images. Our MED model consists of a convolution neural network as an encoder that extracts features of input mathematical equation images and a recurrent neural network with attention mechanism which generates description related to the input mathematical equation images. Due to the unavailability of mathematical equation image data sets with their textual descriptions, we generate two data sets for experimental purpose. To validate the effectiveness of our MED model, we conduct a real-world experiment to see whether the students are able to write equations by only reading or listening their textual descriptions or not. Experiments conclude that the students are able to write most of the equations correctly by reading their textual descriptions only.
Ajoy Mondal, C. V. Jawahar
ICDAR1
2019 Graphical Object Detection in Document Images
abstract
Graphical elements: particularly tables and figures contain a visual summary of the most valuable information contained in a document. Therefore, localization of such graphical objects in the document images is the initial step to understand the content of such graphical objects or document images. In this paper, we present a novel end-to-end trainable deep learning based framework to localize graphical objects in the document images called as Graphical Object Detection ( GOD ). Our framework is data-driven and does not require any heuristics or meta-data to locate graphical objects in the document images. The GOD explores the concept of transfer learning and domain adaptation to handle scarcity of labeled training images for graphical object detection task in the document images. Performance analysis carried out on the various public benchmark data sets: ICDAR -2013, ICDAR - POD2017 and UNLV shows that our model yields promising results as compared to state-of-the-art techniques.
Ranajit Saha, Ajoy Mondal, C. V. Jawahar
ICDAR2
2019 Neuro-probabilistic model for object tracking
Ajoy Mondal
Pattern Anal. Appl.1
2017 Partially Camouflaged Object Tracking using Modified Probabilistic Neural Network and Fuzzy Energy based Active Contour
Ajoy Mondal, Susmita Ghosh, Ashish Ghosh
Int. J. Comput. Vis.1
2016 Robust image segmentation using global and local fuzzy energy based active contour
abstract
Though various image segmentation techniques have been developed, it is still a very challenging task to design a robust and efficient algorithm to segment (noisy, blurred or even discontinuous edged) images having high intensity inhomogeneity or non-homogeneity. In this article, a robust fuzzy energy based active contour, using both global and local information, is proposed to detect objects in a given image based on curve evolution. The local energy is generated by considering both local spatial and gray level/color information. The proposed model can better deal with images having high intensity inhomogeneity or non-homogeneity, noise and blurred boundary or discontinuous edges by incorporating local energy term in the proposed active contour energy function. The global energy term is used to avoid unsatisfactory results due to bad initialization. We show a realization of the proposed method and demonstrate its performance (both qualitatively and quantitatively) with respect to state-of-the-art techniques on several images having such kind of artifacts. Analysis of results concludes that the proposed method can detect objects from given images in a better way than the existing ones.
Ajoy Mondal, K. Ramachandra Murthy, Ashish Ghosh, Susmita Ghosh
FUZZ-IEEE1
2016 Maximum Class Boundary Criterion for supervised dimensionality reduction
abstract
Participation of class-wise noisy patterns may mislead the selection process of relevant patterns for subspace projection. And modelling between-class scatter for each class using the patterns that are nearer to the corresponding class decision boundary may improve the quality of feature generation. In this manuscript, a novel dimensionality reduction method, named Maximum Class Boundary Criterion (MCBC) is proposed. MCBC increases class separability by realizing the significant class-boundary and class-non-boundary patterns after the elimination of noisy patterns. The objective of MCBC is modeled such that the class-boundary patterns are pushed away from the corresponding class means and class-non-boundary patterns are forced towards their class means. As a result, the classification performance of the extracted MCBC features is improved. Experimental study is performed on UCI machine learning and face recognition data to highlight the performance of MCBC. The results conclude that MCBC can generate better discriminative features compared to the state-of-the-art dimensionality reduction methods.
K. Ramachandra Murthy, Ajoy Mondal, Ashish Ghosh
IJCNN2
2016 Efficient silhouette-based contour tracking using local information
Ajoy Mondal, Susmita Ghosh, Ashish Ghosh
Soft Comput.1