Prasun Roy

dblp:224/0012 · DBLP profile ↗
← Back
10ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0002-1733-5670ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 6 first-author · 6 since 2021
YearPublicationVenuePosition
2025 FASTER: A Font-Agnostic Scene Text Editing and Rendering Framework
abstract
Scene Text Editing (STE) is a challenging research prob-lem, that primarily aims towards modifying existing texts in an image while preserving the background and the font style of the original text. Despite its utility in numerous real-world applications, existing style-transfer-based approaches have shown sub-par editing performance due to (1) complex image backgrounds, (2) diverse font attributes, and (3) varying word lengths within the text. To address such limitations, in this paper, we propose a novel font-agnostic scene text editing and rendering framework, named FASTER, for simultaneously generating text in arbitrary styles and locations while preserving a natural and realistic appearance and structure. A combined fusion of target mask generation and style transfer units, with a cascaded self-attention mech-anism has been proposed to focus on multi-level text region edits to handle varying word lengths. Extensive evaluation on a real-world database withfurther subjective human eval-uation study indicates the superiority of FASTER in both scene text editing and rendering tasks, in terms of model per-formance and efficiency. The code and pre-trained models have been released in our Gi thub repo.
Alloy Das, Sanket Biswas, Prasun Roy, Subhankar Ghosh, Umapada Pal 0001, Michael Blumenstein, Josep Lladós 0001, Saumik Bhattacharya
WACV3
2024 λ-Color: Amplifying Long-Range Dependencies for Image Colorization
Subhankar Ghosh, Saumik Bhattacharya, Prasun Roy, Umapada Pal 0001, Michael Blumenstein
ICPR (22)3
2024 d-Sketch: Improving Visual Fidelity of Sketch-to-Image Translation with Pretrained Latent Diffusion Models without Retraining
Prasun Roy, Saumik Bhattacharya, Subhankar Ghosh, Umapada Pal 0001, Michael Blumenstein
ICPR (25)1
2024 Semantically Consistent Person Image Generation
Prasun Roy, Saumik Bhattacharya, Subhankar Ghosh, Umapada Pal 0001, Michael Blumenstein
ICPR (25)1
2024 TIC: text-guided image colorization using conditional generative model
abstract
Abstract Image colorization is a well-known problem in computer vision. However, due to the ill-posed nature of the task, image colorization is inherently challenging. Though several attempts have been made by researchers to make the colorization pipeline automatic, these processes often produce unrealistic results due to a lack of conditioning. In this work, we attempt to integrate textual descriptions as an auxiliary condition, along with the grayscale image that is to be colorized, to improve the fidelity of the colorization process. To the best of our knowledge, this is one of the first attempts to incorporate textual conditioning in the colorization pipeline. To do so, a novel deep network has been proposed that takes two inputs (the grayscale image and the respective encoded text description) and tries to predict the relevant color gamut. As the respective textual descriptions contain color information of the objects present in the scene, the text encoding helps to improve the overall quality of the predicted colors. The proposed model has been evaluated using different metrics like SSIM, PSNR, LPISPS and achieved scores of 0.917, 23.27,0.223, respectively. These quantitative metrics have shown that the proposed method outperforms the SOTA techniques in most of the cases.
Subhankar Ghosh, Prasun Roy, Saumik Bhattacharya, Umapada Pal 0001, Michael Blumenstein
Multim. Tools Appl.2
2023 Multi-scale attention guided pose transfer
Prasun Roy, Saumik Bhattacharya, Subhankar Ghosh, Umapada Pal 0001
Pattern Recognit.1
2022 TIPS: Text-Induced Pose Synthesis
Prasun Roy, Subhankar Ghosh, Saumik Bhattacharya, Umapada Pal 0001, Michael Blumenstein
ECCV (38)1
2022 Scene Aware Person Image Generation through Global Contextual Conditioning
abstract
Person image generation is an intriguing yet challenging problem. However, this task becomes even more difficult under constrained situations. In this work, we propose a novel pipeline to generate and insert contextually relevant person images into an existing scene while preserving the global semantics. More specifically, we aim to insert a person such that the location, pose, and scale of the person being inserted blends in with the existing persons in the scene. Our method uses three individual networks in a sequential pipeline. At first, we predict the potential location and the skeletal structure of the new person by conditioning a Wasserstein Generative Adversarial Network (WGAN) on the existing human skeletons present in the scene. Next, the predicted skeleton is refined through a shallow linear network to achieve higher structural accuracy in the generated image. Finally, the target image is generated from the refined skeleton using another generative network conditioned on a given image of the target person. In our experiments, we achieve high-resolution photo-realistic generation results while preserving the general context of the scene. We conclude our paper with multiple qualitative and quantitative benchmarks on the results.
Prasun Roy, Subhankar Ghosh, Saumik Bhattacharya, Umapada Pal 0001, Michael Blumenstein
ICPR1
2020 STEFANN: Scene Text Editor Using Font Adaptive Neural Network
abstract
Textual information in a captured scene plays an important role in scene interpretation and decision making. Though there exist methods that can successfully detect and interpret complex text regions present in a scene, to the best of our knowledge, there is no significant prior work that aims to modify the textual information in an image. The ability to edit text directly on images has several advantages including error correction, text restoration and image reusability. In this paper, we propose a method to modify text in an image at character-level. We approach the problem in two stages. At first, the unobserved character (target) is generated from an observed character (source) being modified. We propose two different neural network architectures - (a) FANnet to achieve structural consistency with source font and (b) Colornet to preserve source color. Next, we replace the source character with the generated character maintaining both geometric and visual consistency with neighboring characters. Our method works as a unified platform for modifying text in images. We present the effectiveness of our method on COCO-Text and ICDAR datasets both qualitatively and quantitatively.
Prasun Roy, Saumik Bhattacharya, Subhankar Ghosh, Umapada Pal 0001
CVPR1
2018 A CNN Based Framework for Unistroke Numeral Recognition in Air-Writing
abstract
Air-writing refers to virtually writing linguistic characters through hand gestures in three dimensional space with six degrees of freedom. In this paper a generic video camera dependent convolutional neural network (CNN) based air-writing framework has been proposed. Gestures are performed using a marker of fixed color in front of a generic video camera followed by color based segmentation to identify the marker and track the trajectory of marker tip. A pre-trained CNN is then used to classify the gesture. The recognition accuracy is further improved using transfer learning with the newly acquired data. The performance of the system varies greatly on the illumination condition due to color based segmentation. In a less fluctuating illumination condition the system is able to recognize isolated unistroke numerals of multiple languages. The proposed framework achieved 97.7%, 95.4% and 93.7% recognition rate in person independent evaluation over English, Bengali and Devanagari numerals, respectively.
Prasun Roy, Subhankar Ghosh, Umapada Pal 0001
ICFHR1