VLDB 2026 Research / reviewers in the wild / expert
Ayan Banerjee 0002
dblp:86/2304-2
· DBLP profile ↗
16ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0002-0269-2202ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CraftSVG: Multi-Object Text-to-SVG Synthesis via Layout Guided DiffusionabstractGenerating SVGs from text is a challenging vision task, requiring diverse yet realistic depictions of the seen as well as unseen entities. Existing research has been mostly limited to generating single-object rather than comprehensive scenes comprising multiple elements. In response, CraftSVG introduces an end-to-end framework for creating SVGs depicting entire scenes from a textual description. Utilizing a pre-trained LLM for layout generation from text via iterative in-context learning, CraftSVG introduces a per-box mask latent mechanism for accurate object placement. A fusion mechanism is developed to integrate the attention maps, employing a diffusion U-Net for coherent composition, which accelerates stroke initialization. Recognizing the importance of abstract SVGs in communication, we incorporated an MLP-based mechanism to simplify SVGs, with alignment and perceptual loss via differential rendering and opacity modulation to improve aesthetics. CraftSVG outperforms previous methods in abstraction, recognizability, and detail, as depicted by its CLIP-T: 0.5013, Aesthetic: 7.0779, score. The code is available at github.com/CraftSVG. Ayan Banerjee 0002, Nityanand Mathur, Josep Lladós 0001, Umapada Pal 0001, Anjan Dutta 0001 |
WACV | 1 |
| 2024 | GraphKD: Exploring Knowledge Distillation Towards Document Object Detection with Structured Graph Creation
Ayan Banerjee 0002, Sanket Biswas, Josep Lladós 0001, Umapada Pal 0001 |
ICDAR (3) | 1 |
| 2024 | DistilDoc: Knowledge Distillation for Visually-Rich Document Applications
Jordy Van Landeghem, Subhajit Maity, Ayan Banerjee 0002, Matthew B. Blaschko, Marie-Francine Moens, Josep Lladós 0001, Sanket Biswas |
ICDAR (4) | 3 |
| 2024 | Harnessing the Power of Multi-Lingual Datasets for Pre-training: Towards Enhancing Text Spotting PerformanceabstractThe adaptation capability to a wide range of domains is crucial for scene text spotting models when deployed to real-world conditions. However, existing SOTA approaches usually incorporate scene text detection and recognition simply by pretraining on natural scene text datasets, which do not directly exploit the intermediate feature representations between multiple domains. Here, we investigate the problem of domain-adaptive scene text spotting, i.e., training a model on multi-domain source data such that it can directly adapt to target domains rather than being specialized for a specific domain or scenario. Further, we investigate a transformer baseline called Swin-TESTR to focus on solving scene-text spotting for both regular and arbitraryshaped text along with an exhaustive evaluation. The results demonstrate the potential of intermediate representations to gain significant performance on text spotting benchmarks across multiple domains (e.g. language, synth-to-real, and documents). both in terms of accuracy and efficiency. Alloy Das, Sanket Biswas, Ayan Banerjee 0002, Josep Lladós 0001, Umapada Pal 0001, Saumik Bhattacharya |
WACV | 3 |
| 2024 | A new deep CNN for 3D text localization in the wild through shadow removal
Palaiahnakote Shivakumara, Ayan Banerjee 0002, Lokesh Nandanwar, Umapada Pal 0001, Apostolos Antonacopoulos, Tong Lu 0002, Michael Blumenstein |
Comput. Vis. Image Underst. | 2 |
| 2024 | SemiDocSeg: harnessing semi-supervised learning for document layout analysis
Ayan Banerjee 0002, Sanket Biswas, Josep Lladós 0001, Umapada Pal 0001 |
Int. J. Document Anal. Recognit. | 1 |
| 2024 | Soft set-based MSER end-to-end system for occluded scene text detection, recognition and predictionabstractThe presence of unpredictable occlusions on natural scene text is a significant challenge, exacerbating the difficulties already posed on text detection and recognition by the variability of such images. Addressing the need for a robust, consistently performing approach that can effectively address the above challenges, this paper presents a new Soft Set-based end-to-end system for text detection, recognition and prediction in occluded natural scene images. This is the first approach to integrate text detection, recognition and prediction , unlike existing systems developed for end-to-end text spotting (text detection and recognition) only. For candidate text components detection, the proposed combination of Soft Sets with Maximally Stable Extremal Regions (SS-MSER) improves text detection and spotting in natural scene images, irrespectively of the presence of arbitrarily orientated and shaped text, complex backgrounds and occlusion. Furthermore, a Graph Recurrent Neural Network is proposed for grouping candidate text components into text lines and for fitting accurate bounding boxes to each word. Finally, a Convolutional Recurrent Neural Network (CRNN) is proposed for the recognition of text and for predicting missing characters due to occlusion. Experimental results on a new occluded scene text dataset (OSTD) and on the most relevant benchmark natural scene text datasets demonstrate that the proposed system outperforms the state-of-the-art in text detection, recognition and prediction. The code and dataset are available at https://github.com/alloydas/Softset-MSER-Based-Occluded-Scene-Text-Spotting/blob/master/Soft_set_MSER.ipynb Alloy Das, Palaiahnakote Shivakumara, Ayan Banerjee 0002, Apostolos Antonacopoulos, Umapada Pal 0001 |
Knowl. Based Syst. | 3 |
| 2024 | An end-to-end model for multi-view scene text recognition
Ayan Banerjee 0002, Palaiahnakote Shivakumara, Saumik Bhattacharya, Umapada Pal 0001, Cheng-Lin Liu 0001 |
Pattern Recognit. | 1 |
| 2024 | TTS: Hilbert Transform-Based Generative Adversarial Network for Tattoo and Scene Text SpottingabstractText spotting in natural scenes is of increasing interest and significance due to its critical role in several applications, such as visual question answering, named entity recognition and event rumor detection on social media. One of the newly emerging challenging problems is Tattoo Text Spotting (TTS) in images for assisting forensic teams and for person identification. Unlike the generally simpler scene text addressed by current state-of-the-art methods, tattoo text is typically characterized by the presence of decorative backgrounds, calligraphic handwriting and several distortions due to the deformable nature of the skin. This paper describes the first approach to address TTS in a real-world application context by designing an end-to-end text spotting method employing a Hilbert transform-based Generative Adversarial Network (GAN). To reduce the complexity of the TTS task, the proposed approach first detects fine details in the image using the Hilbert transform and the Optimum Phase Congruency (OPC). To overcome the challenges of only having a relatively small number of training samples, a GAN is then used for generating suitable text samples and descriptors for text spotting (i.e., both detection and recognition). The superior performance of the proposed TTS approach, for both tattoo and general scene text, over the state-of-the-art methods is demonstrated on a new TTS-specific dataset (publicly available) as well as on the existing benchmark natural scene text datasets: Total-Text, CTW1500 and ICDAR 2015. Ayan Banerjee 0002, Palaiahnakote Shivakumara, Umapada Pal 0001, Apostolos Antonacopoulos, Tong Lu 0002, Josep Lladós 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | SwinDocSegmenter: An End-to-End Unified Domain Adaptive Transformer for Document Instance Segmentation
Ayan Banerjee 0002, Sanket Biswas, Josep Lladós 0001, Umapada Pal 0001 |
ICDAR (1) | 1 |
| 2023 | SelfDocSeg: A Self-supervised Vision-Based Approach Towards Document Segmentation
Subhajit Maity, Sanket Biswas, Siladittya Manna, Ayan Banerjee 0002, Josep Lladós 0001, Saumik Bhattacharya, Umapada Pal 0001 |
ICDAR (1) | 4 |
| 2023 | A New Language-Independent Deep CNN for Scene Text Detection and Style Transfer in Social Media ImagesabstractDue to the adverse effect of quality caused by different social media and arbitrary languages in natural scenes, detecting text from social media images and transferring its style is challenging. This paper presents a novel end-to-end model for text detection and text style transfer in social media images. The key notion of the proposed work is to find dominant information, such as fine details in the degraded images (social media images), and then restore the structure of character information. Therefore, we first introduce a novel idea of extracting gradients from the frequency domain of the input image to reduce the adverse effect of different social media, which outputs text candidate points. The text candidates are further connected into components and used for text detection via a UNet++ like network with an EfficientNet backbone (EffiUNet++). Then, to deal with the style transfer issue, we devise a generative model, which comprises a target encoder and style parameter networks (TESP-Net) to generate the target characters by leveraging the recognition results from the first stage. Specifically, a series of residual mapping and a position attention module are devised to improve the shape and structure of generated characters. The whole model is trained end-to-end so as to optimize the performance. Experiments on our social media dataset, benchmark datasets of natural scene text detection and text style transfer show that the proposed model outperforms the existing text detection and style transfer methods in multilingual and cross-language scenario. Palaiahnakote Shivakumara, Ayan Banerjee 0002, Umapada Pal 0001, Lokesh Nandanwar, Tong Lu 0002, Cheng-Lin Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | TWD: A New Deep E2E Model for Text Watermark/Caption and Scene Text Detection in VideoabstractText watermark detection in video images is challenging because text watermark characteristics are different from caption and scene texts in the video images. Developing a successful model for detecting text watermark, caption, and scene texts is an open challenge. This study aims at developing a new Deep End-to-End model for Text Watermark Detection (TWD), caption and scene text in video images. To standardize non-uniform contrast, quality, and resolution, we explore the U-Net3+ model for enhancing poor quality text without affecting high-quality text. Similarly, to address the challenges of arbitrary orientation, text shapes and complex background, we explore Stacked Hourglass Encoded Fourier Contour Embedding Network (SFCENet) by feeding the output of the U-Net3+ model as input. Furthermore, the proposed work integrates enhancement and detection models as an end-to-end model for detecting multi-type text in video images. To validate the proposed model, we create our own dataset (named TW-866), which provides video images containing text watermark, caption (subtitles), as well as scene text. The proposed model is also evaluated on standard natural scene text detection datasets, namely, ICDAR 2019 MLT, CTW1500, Total-Text, and DAST1500. The results show that the proposed method outperforms the existing methods. This is the first work on text watermark detection in video images to the best of our knowledge. Ayan Banerjee 0002, Palaiahnakote Shivakumara, Parikshit Acharya, Umapada Pal 0001, Josep Lladós 0001 |
ICPR | 1 |
| 2022 | Missing value estimation of microarray data using Sim-GAN
Soumen Kumar Pati, Manan Kumar Gupta, Rinita Shai, Ayan Banerjee 0002 |
Knowl. Inf. Syst. | 4 |
| 2022 | SHUBHCHINTAK
Ayan Banerjee 0002, Dibyendu Maji, Rajdeep Datta, Subhas Barman, Debasis Samanta, Samiran Chattopadhyay |
Multim. Tools Appl. | 1 |
| 2022 | A comprehensive scheme for tattoo text detectionabstractTattoo text detection provides a vital clue for person and crime identification. Due to the freestyle and unconstrained nature of handwritten tattoo text over skin regions, accurate tattoo text detection is very challenging. This paper proposes a comprehensive scheme for tattoo text detection which comprises (a) adaptive Deformable Convolutional Neural Network (DCNN) for skin region detection to reduce text detection complexity (b) a Decoupled Gradient Text Detector (DGTD) for tattoo text detection from skin region (c) a Deep Q-Network (DQN) to refine the bounding boxes detected by DGTD, and (d) a Term-Frequency-Inverse-Document-Frequency (TF-IDF) model to group the words into text lines based on semantic information to fix the bounding box for the line. To test the effectiveness, the proposed method is evaluated on different datasets, namely, (i) a newly developed tattoo text dataset, (ii) benchmark bib number dataset of the marathon, and (iii) person re-identification dataset. The proposed method achieves 91.2, 87.5, and 88.8 F-scores from these three respective datasets. To demonstrate its superior performance, the text detection module (without skin detection) is also compared with state-of-the-art scene text detection methods on benchmark datasets, namely, ICDAR 2019 ArT, Total-Text, and DAST1500 and the proposed method achieves 90.3, 88.5 and 89.8 F-score from these respective datasets. Ayan Banerjee 0002, Palaiahnakote Shivakumara, Umapada Pal 0001, Ramachandra Raghavendra, Cheng-Lin Liu 0001 |
Pattern Recognit. Lett. | 1 |