VLDB 2026 Research / reviewers in the wild / expert
Prasun Roy
dblp:224/0012
· DBLP profile ↗
10ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0002-1733-5670ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 6 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FASTER: A Font-Agnostic Scene Text Editing and Rendering FrameworkabstractScene Text Editing (STE) is a challenging research prob-lem, that primarily aims towards modifying existing texts in an image while preserving the background and the font style of the original text. Despite its utility in numerous real-world applications, existing style-transfer-based approaches have shown sub-par editing performance due to (1) complex image backgrounds, (2) diverse font attributes, and (3) varying word lengths within the text. To address such limitations, in this paper, we propose a novel font-agnostic scene text editing and rendering framework, named FASTER, for simultaneously generating text in arbitrary styles and locations while preserving a natural and realistic appearance and structure. A combined fusion of target mask generation and style transfer units, with a cascaded self-attention mech-anism has been proposed to focus on multi-level text region edits to handle varying word lengths. Extensive evaluation on a real-world database withfurther subjective human eval-uation study indicates the superiority of FASTER in both scene text editing and rendering tasks, in terms of model per-formance and efficiency. The code and pre-trained models have been released in our Gi thub repo. Alloy Das, Sanket Biswas, Prasun Roy, Subhankar Ghosh, Umapada Pal 0001, Michael Blumenstein, Josep Lladós 0001, Saumik Bhattacharya |
WACV | 3 |
| 2024 | λ-Color: Amplifying Long-Range Dependencies for Image Colorization
Subhankar Ghosh, Saumik Bhattacharya, Prasun Roy, Umapada Pal 0001, Michael Blumenstein |
ICPR (22) | 3 |
| 2024 | d-Sketch: Improving Visual Fidelity of Sketch-to-Image Translation with Pretrained Latent Diffusion Models without Retraining
Prasun Roy, Saumik Bhattacharya, Subhankar Ghosh, Umapada Pal 0001, Michael Blumenstein |
ICPR (25) | 1 |
| 2024 | Semantically Consistent Person Image Generation
Prasun Roy, Saumik Bhattacharya, Subhankar Ghosh, Umapada Pal 0001, Michael Blumenstein |
ICPR (25) | 1 |
| 2024 | TIC: text-guided image colorization using conditional generative modelabstractAbstract Image colorization is a well-known problem in computer vision. However, due to the ill-posed nature of the task, image colorization is inherently challenging. Though several attempts have been made by researchers to make the colorization pipeline automatic, these processes often produce unrealistic results due to a lack of conditioning. In this work, we attempt to integrate textual descriptions as an auxiliary condition, along with the grayscale image that is to be colorized, to improve the fidelity of the colorization process. To the best of our knowledge, this is one of the first attempts to incorporate textual conditioning in the colorization pipeline. To do so, a novel deep network has been proposed that takes two inputs (the grayscale image and the respective encoded text description) and tries to predict the relevant color gamut. As the respective textual descriptions contain color information of the objects present in the scene, the text encoding helps to improve the overall quality of the predicted colors. The proposed model has been evaluated using different metrics like SSIM, PSNR, LPISPS and achieved scores of 0.917, 23.27,0.223, respectively. These quantitative metrics have shown that the proposed method outperforms the SOTA techniques in most of the cases. Subhankar Ghosh, Prasun Roy, Saumik Bhattacharya, Umapada Pal 0001, Michael Blumenstein |
Multim. Tools Appl. | 2 |
| 2023 | Multi-scale attention guided pose transfer
Prasun Roy, Saumik Bhattacharya, Subhankar Ghosh, Umapada Pal 0001 |
Pattern Recognit. | 1 |
| 2022 | TIPS: Text-Induced Pose Synthesis
Prasun Roy, Subhankar Ghosh, Saumik Bhattacharya, Umapada Pal 0001, Michael Blumenstein |
ECCV (38) | 1 |
| 2022 | Scene Aware Person Image Generation through Global Contextual ConditioningabstractPerson image generation is an intriguing yet challenging problem. However, this task becomes even more difficult under constrained situations. In this work, we propose a novel pipeline to generate and insert contextually relevant person images into an existing scene while preserving the global semantics. More specifically, we aim to insert a person such that the location, pose, and scale of the person being inserted blends in with the existing persons in the scene. Our method uses three individual networks in a sequential pipeline. At first, we predict the potential location and the skeletal structure of the new person by conditioning a Wasserstein Generative Adversarial Network (WGAN) on the existing human skeletons present in the scene. Next, the predicted skeleton is refined through a shallow linear network to achieve higher structural accuracy in the generated image. Finally, the target image is generated from the refined skeleton using another generative network conditioned on a given image of the target person. In our experiments, we achieve high-resolution photo-realistic generation results while preserving the general context of the scene. We conclude our paper with multiple qualitative and quantitative benchmarks on the results. Prasun Roy, Subhankar Ghosh, Saumik Bhattacharya, Umapada Pal 0001, Michael Blumenstein |
ICPR | 1 |
| 2020 | STEFANN: Scene Text Editor Using Font Adaptive Neural NetworkabstractTextual information in a captured scene plays an important role in scene interpretation and decision making. Though there exist methods that can successfully detect and interpret complex text regions present in a scene, to the best of our knowledge, there is no significant prior work that aims to modify the textual information in an image. The ability to edit text directly on images has several advantages including error correction, text restoration and image reusability. In this paper, we propose a method to modify text in an image at character-level. We approach the problem in two stages. At first, the unobserved character (target) is generated from an observed character (source) being modified. We propose two different neural network architectures - (a) FANnet to achieve structural consistency with source font and (b) Colornet to preserve source color. Next, we replace the source character with the generated character maintaining both geometric and visual consistency with neighboring characters. Our method works as a unified platform for modifying text in images. We present the effectiveness of our method on COCO-Text and ICDAR datasets both qualitatively and quantitatively. Prasun Roy, Saumik Bhattacharya, Subhankar Ghosh, Umapada Pal 0001 |
CVPR | 1 |
| 2018 | A CNN Based Framework for Unistroke Numeral Recognition in Air-WritingabstractAir-writing refers to virtually writing linguistic characters through hand gestures in three dimensional space with six degrees of freedom. In this paper a generic video camera dependent convolutional neural network (CNN) based air-writing framework has been proposed. Gestures are performed using a marker of fixed color in front of a generic video camera followed by color based segmentation to identify the marker and track the trajectory of marker tip. A pre-trained CNN is then used to classify the gesture. The recognition accuracy is further improved using transfer learning with the newly acquired data. The performance of the system varies greatly on the illumination condition due to color based segmentation. In a less fluctuating illumination condition the system is able to recognize isolated unistroke numerals of multiple languages. The proposed framework achieved 97.7%, 95.4% and 93.7% recognition rate in person independent evaluation over English, Bengali and Devanagari numerals, respectively. Prasun Roy, Subhankar Ghosh, Umapada Pal 0001 |
ICFHR | 1 |