Sanket Biswas

dblp:228/8453 · DBLP profile ↗
← Back
28ranked-venue papers
2as first author
28since 2021 · last 2026
0000-0001-6648-8270ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 2 first-author · 23 since 2021Databases, data management, data science and information retrieval · 13 · 1 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 10 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From Chunks to Graphs: Training-Free Multimodal Late Interaction for Document Understanding
Ayush Lodh, Souparni Mazumder, Sanket Biswas, Josep Lladós 0001, Nisha Singh Chauhan
ICDAR (3)3
2026 Hybrid Classical-Quantum Architecture for Vectorised Image Classification of Hand-Written Sketches
Yeray Cordero, Sanket Biswas, Fernando Vilariño, Matías Bilkis
ICPR (6)2
2025 Where Layout Meets Language: Lightweight Spatial Enhancement to Large Language Models for Document Understanding
Nil Biescas, Sanket Biswas, Josep Lladós 0001, Jordy Van Landeghem
ICDAR (4)2
2025 Doc2GraphFormer: Bridging Structured Graph Learning with Transformer Attention for Efficient Document Understanding
Souparni Mazumder, Sanket Biswas, Aniket Pal, Alloy Das, Umapada Pal 0001, Josep Lladós 0001
ICDAR (4)2
2025 ICDAR 2025 Handwritten Notes Understanding Challenge
Aniket Pal, Sanket Biswas, Alloy Das, Ayush Lodh, Priyanka Banerjee, Soumitri Chattopadhyay, Ajoy Mondal, Dimosthenis Karatzas, Josep Lladós 0001, C. V. Jawahar
ICDAR (5)2
2025 BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks
abstract
Multimodal AI has the potential to significantly enhance document-understanding tasks, such as processing receipts, understanding workflows, extracting data from documents, and summarizing reports. Code generation tasks that require long-structured outputs can also be enhanced by multimodality. Despite this, their use in commercial applications is often limited due to limited access to relevant training data and restrictive licensing, which hinders open access. To address these limitations, we introduce BigDocs-7.5M, a high-quality, open-access dataset comprising 7.5 million multimodal documents across 30 tasks. We use an efficient data curation process to ensure that our data is high quality and license-permissive. Our process emphasizes accountability, responsibility, and transparency through filtering rules, traceable metadata, and careful content analysis. Additionally, we introduce BigDocs-Bench,, a benchmark suite with 10 novel tasks where we carefully create datasets that reflect real-world use cases involving reasoning over Graphical User Interfaces (GUI) and code generation from images. Our experiments show that training with BigDocs-Bench, improves average performance up to 25.8% over closed-source GPT-4o in document reasoning and structured output tasks such as Screenshot2HTML or Image2Latex generation. Finally, human evaluations revealed that participants preferred the outputs from models trained with BigDocs over those from GPT-4o. This suggests that BigDocs can help both academics and the open-source community utilize and improve AI tools to enhance multimodal capabilities and document reasoning.
Juan A. Rodríguez, Xiangru Jian, Siba Smarak Panigrahi, Aarash Feizi, Abhay Puri, Akshay Kalkunte Suresh, François Savard, Ahmed Masry, Shravan Nayak, Rabiul Awal, Mahsa Massoud, Amirhossein Abaskohi, Suyuchen Wang, Pierre-André Noël, Mats Leon Richter, Saverio Vadacchino, Sanket Biswas
ICLR20
2025 GlobalDoc: A Cross-Modal Vision-Language Framework for Real-World Document Image Retrieval and Classification
abstract
Visual document understanding (VDU) has rapidly advanced with the development of powerful multi-modal language models. However, these models typically require extensive document pre-training data to learn intermediate representations and often suffer a significant performance drop in real-world online industrial settings. A primary issue is their heavy reliance on OCR engines to extract local positional information within document pages, which limits the models' ability to capture global information and hinders their generalizability, flexibility, and robustness. In this paper, we introduce GlobalDoc, a cross modal transformer-based architecture pre-trained in a self supervised manner using three novel pretext objective tasks. GlobalDoc improves the learning of richer semantic concepts by unifying language and visual representations, resulting in more transferable models. For proper evaluation, we also propose two novel document-level downstream VDU tasks, Few-Shot Document Image Classification (DIC) and Content-based Document Image Retrieval (DIR), designed to simulate industrial scenarios more closely. Extensive experimentation has been conducted to demonstrate GlobalDoc's effectiveness in practical settings.
Souhail Bakkali, Sanket Biswas, Zuheng Ming, Mickaël Coustaty, Marçal Rusiñol, Oriol Ramos Terrades, Josep Lladós 0001
WACV2
2025 FASTER: A Font-Agnostic Scene Text Editing and Rendering Framework
abstract
Scene Text Editing (STE) is a challenging research prob-lem, that primarily aims towards modifying existing texts in an image while preserving the background and the font style of the original text. Despite its utility in numerous real-world applications, existing style-transfer-based approaches have shown sub-par editing performance due to (1) complex image backgrounds, (2) diverse font attributes, and (3) varying word lengths within the text. To address such limitations, in this paper, we propose a novel font-agnostic scene text editing and rendering framework, named FASTER, for simultaneously generating text in arbitrary styles and locations while preserving a natural and realistic appearance and structure. A combined fusion of target mask generation and style transfer units, with a cascaded self-attention mech-anism has been proposed to focus on multi-level text region edits to handle varying word lengths. Extensive evaluation on a real-world database withfurther subjective human eval-uation study indicates the superiority of FASTER in both scene text editing and rendering tasks, in terms of model per-formance and efficiency. The code and pre-trained models have been released in our Gi thub repo.
Alloy Das, Sanket Biswas, Prasun Roy, Subhankar Ghosh, Umapada Pal 0001, Michael Blumenstein, Josep Lladós 0001, Saumik Bhattacharya
WACV2
2024 Towards Generative Class Prompt Learning for Fine-grained Visual Recognition
Soumitri Chattopadhyay, Sanket Biswas, Emanuele Vivoli, Josep Lladós 0001
BMVC2
2024 GraphKD: Exploring Knowledge Distillation Towards Document Object Detection with Structured Graph Creation
Ayan Banerjee 0002, Sanket Biswas, Josep Lladós 0001, Umapada Pal 0001
ICDAR (3)2
2024 GeoContrastNet: Contrastive Key-Value Edge Learning for Language-Agnostic Document Understanding
Nil Biescas, Carlos Boned, Josep Lladós 0001, Sanket Biswas
ICDAR (1)4
2024 DistilDoc: Knowledge Distillation for Visually-Rich Document Applications
Jordy Van Landeghem, Subhajit Maity, Ayan Banerjee 0002, Matthew B. Blaschko, Marie-Francine Moens, Josep Lladós 0001, Sanket Biswas
ICDAR (4)7
2024 Recurrent Few-Shot Model for Document Verification
Maxime Talarmain, Carlos Boned Riera, Sanket Biswas, Oriol Ramos Terrades
ICDAR (1)3
2024 SketchGPT: Autoregressive Modeling for Sketch Generation and Recognition
Adarsh Tiwari, Sanket Biswas, Josep Lladós 0001
ICDAR (5)2
2024 FastTextSpotter: A High-Efficiency Transformer for Multilingual Scene Text Spotting
Alloy Das, Sanket Biswas, Umapada Pal 0001, Josep Lladós 0001, Saumik Bhattacharya
ICPR (20)2
2024 Diving into the Depths of Spotting Text in Multi-Domain Noisy Scenes
abstract
When used in a real-world noisy environment, the capacity to generalize to multiple domains is essential for any autonomous scene text spotting system. However, existing state-of-the-art methods employ pretraining and fine-tuning strategies on natural scene datasets, which do not exploit the feature interaction across other complex domains. In this work, we explore and investigate the problem of domain-agnostic scene text spotting, i.e., training a model on multi-domain source data such that it can directly generalize to target domains rather than being specialized for a specific domain or scenario. In this regard, we present the community a text spotting validation benchmark called Under-Water Text (UWT) for noisy underwater scenes to establish an important case study. Moreover, we also design an efficient super-resolution based end-to-end transformer baseline called DA-TextSpotter which achieves comparable or superior performance over existing text spotting architectures for both regular and arbitrary-shaped scene text spotting benchmarks in terms of both accuracy and model efficiency. The dataset, code and pre-trained models have been released in our Github.
Alloy Das, Sanket Biswas, Umapada Pal 0001, Josep Lladós 0001
ICRA2
2024 Harnessing the Power of Multi-Lingual Datasets for Pre-training: Towards Enhancing Text Spotting Performance
abstract
The adaptation capability to a wide range of domains is crucial for scene text spotting models when deployed to real-world conditions. However, existing SOTA approaches usually incorporate scene text detection and recognition simply by pretraining on natural scene text datasets, which do not directly exploit the intermediate feature representations between multiple domains. Here, we investigate the problem of domain-adaptive scene text spotting, i.e., training a model on multi-domain source data such that it can directly adapt to target domains rather than being specialized for a specific domain or scenario. Further, we investigate a transformer baseline called Swin-TESTR to focus on solving scene-text spotting for both regular and arbitraryshaped text along with an exhaustive evaluation. The results demonstrate the potential of intermediate representations to gain significant performance on text spotting benchmarks across multiple domains (e.g. language, synth-to-real, and documents). both in terms of accuracy and efficiency.
Alloy Das, Sanket Biswas, Ayan Banerjee 0002, Josep Lladós 0001, Umapada Pal 0001, Saumik Bhattacharya
WACV2
2024 Beyond Document Page Classification: Design, Datasets, and Challenges
abstract
This paper highlights the need to bring document classification benchmarking closer to real-world applications, both in the nature of data tested (X: multi-channel, multi-paged, multi-industry; Y : class distributions and label set variety) and in classification tasks considered (f: multi-page document, page stream, and document bundle classification, …). We identify the lack of public multi-page document classification datasets, formalize different classification tasks arising in application scenarios, and motivate the value of targeting efficient multi-page document representations. An experimental study on proposed multi-page document classification datasets demonstrates that current benchmarks have become irrelevant and need to be updated to evaluate complete documents, as they naturally occur in practice. This reality check also calls for more mature evaluation methodologies, covering calibration evaluation, inference complexity (time-memory), and a range of realistic distribution shifts (e.g., born-digital vs. scanning noise, shifting page order). Our study ends on a hopeful note by recommending concrete avenues for future improvements.
Jordy Van Landeghem, Sanket Biswas, Matthew B. Blaschko, Marie-Francine Moens
WACV2
2024 SemiDocSeg: harnessing semi-supervised learning for document layout analysis
Ayan Banerjee 0002, Sanket Biswas, Josep Lladós 0001, Umapada Pal 0001
Int. J. Document Anal. Recognit.2
2024 A unified representation framework for the evaluation of Optical Music Recognition systems
abstract
Abstract Modern-day Optical Music Recognition (OMR) is a fairly fragmented field. Most OMR approaches use datasets that are independent and incompatible between each other, making it difficult to both combine them and compare recognition systems built upon them. In this paper we identify the need of a common music representation language and propose the Music Tree Notation format, with the idea to construct a common endpoint for OMR research that allows coordination, reuse of technology and fair evaluation of community efforts. This format represents music as a set of primitives that group together into higher-abstraction nodes, a compromise between the expression of fully graph-based and sequential notation formats. We have also developed a specific set of OMR metrics and a typeset score dataset as a proof of concept of this idea.
Pau Torras, Sanket Biswas, Alicia Fornés
Int. J. Document Anal. Recognit.2
2023 Text-DIAE: A Self-Supervised Degradation Invariant Autoencoder for Text Recognition and Document Enhancement
abstract
In this paper, we propose a Text-Degradation Invariant Auto Encoder (Text-DIAE), a self-supervised model designed to tackle two tasks, text recognition (handwritten or scene-text) and document image enhancement. We start by employing a transformer-based architecture that incorporates three pretext tasks as learning objectives to be optimized during pre-training without the usage of labelled data. Each of the pretext objectives is specifically tailored for the final downstream tasks. We conduct several ablation experiments that confirm the design choice of the selected pretext tasks. Importantly, the proposed model does not exhibit limitations of previous state-of-the-art methods based on contrastive losses, while at the same time requiring substantially fewer data samples to converge. Finally, we demonstrate that our method surpasses the state-of-the-art in existing supervised and self-supervised settings in handwritten and scene text recognition and document image enhancement. Our code and trained models will be made publicly available at https://github.com/dali92002/SSL-OCR
Mohamed Ali Souibgui, Sanket Biswas, Andrés Mafla, Ali Furkan Biten, Alicia Fornés, Yousri Kessentini, Josep Lladós 0001, Lluís Gómez i Bigorda, Dimosthenis Karatzas
AAAI2
2023 SwinDocSegmenter: An End-to-End Unified Domain Adaptive Transformer for Document Instance Segmentation
Ayan Banerjee 0002, Sanket Biswas, Josep Lladós 0001, Umapada Pal 0001
ICDAR (1)2
2023 ICDAR 2023 Competition on Document UnderstanDing of Everything (DUDE)
Jordy Van Landeghem, Rubèn Tito, Lukasz Borchmann, Michal Pietruszka, Dawid Jurkiewicz, Rafal Powalski, Pawel Józiak, Sanket Biswas, Mickaël Coustaty, Tomasz Stanislawek
ICDAR (2)8
2023 SelfDocSeg: A Self-supervised Vision-Based Approach Towards Document Segmentation
Subhajit Maity, Sanket Biswas, Siladittya Manna, Ayan Banerjee 0002, Josep Lladós 0001, Saumik Bhattacharya, Umapada Pal 0001
ICDAR (1)2
2022 A Few Shot Multi-representation Approach for N-Gram Spotting in Historical Manuscripts
Giuseppe De Gregorio, Sanket Biswas, Mohamed Ali Souibgui, Asma Bensalah, Josep Lladós 0001, Alicia Fornés, Angelo Marcelli
ICFHR2
2022 DocEnTr: An End-to-End Document Image Enhancement Transformer
abstract
Document images can be affected by many degradation scenarios, which cause recognition and processing difficulties. In this age of digitization, it is important to denoise them for proper usage. To address this challenge, we present a new encoder-decoder architecture based on vision transformers to enhance both machine-printed and handwritten document images, in an end-to-end fashion. The encoder operates directly on the pixel patches with their positional information without the use of any convolutional layers, while the decoder reconstructs a clean image from the encoded patches. Conducted experiments show a superiority of the proposed model compared to the state-of-the-art methods on several DIBCO benchmarks. Code and models will be publicly available at: https://github.com/dali92002/DocEnTR.
Mohamed Ali Souibgui, Sanket Biswas, Sana Khamekhem Jemni, Yousri Kessentini, Alicia Fornés, Josep Lladós 0001, Umapada Pal 0001
ICPR2
2021 DocSynth: A Layout Guided Approach for Controllable Document Image Synthesis
Sanket Biswas, Pau Riba, Josep Lladós 0001, Umapada Pal 0001
ICDAR (3)1
2021 Beyond document object detection: instance-level segmentation of complex layouts
Sanket Biswas, Pau Riba, Josep Lladós 0001, Umapada Pal 0001
Int. J. Document Anal. Recognit.1