Prateek Keserwani

dblp:232/2520 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
7since 2021 · last 2024
0000-0001-7611-6462ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2024 DCDM: Diffusion-Conditioned-Diffusion Model for Scene Text Image Super-Resolution
Shrey Singh, Prateek Keserwani, Masakazu Iwamura, Partha Pratim Roy 0001
ECCV (15)2
2024 Efficient Adapter on Pre-trained Visual Feature Reliance in Medical Visual Question Answering
Aakansha Mishra, Prateek Keserwani, Vikram Nelvoy Rajendiran, Ashok K. Senapati
ICPR (28)2
2023 Receptive Field Reliant Zero-Cost Proxies for Neural Architecture Search
abstract
Neural Architecture Search (NAS) is a fast growing technology for automatic design of deep-learning architectures. NAS includes three stages: search space design, search strategy, and evaluation criterion. Among these, the evaluation of various architectures is very cost-intensive task. In this work, we have proposed a set of receptive field reliant zero-cost proxies which need only one iteration of training and thereby reduce the computational time associated with evaluation criterion during the NAS. The proposed zero-cost proxies are based on layer-wise binding of the prune-at-initialization score with its receptive field for more effective measure as compared to the vanilla counterparts to achieve generalizability. The proposed zero-cost proxies are validated on the set of PyTorchCV models, and NAS-Bench-201 benchmarking datasets. The proposed zero-cost proxies have performed better for set of PyTorchCV models and competitively with vanilla counterparts for NAS-Bench-201. The efficiency of the proposed method is also demonstrated in NAS on NAS-Bench-201 using Aging Evolution as controller.
Prateek Keserwani, Srinivas Soumitri Miriyala, Vikram Nelvoy Rajendiran, Pradeep N. Shivamurthappa
ICASSP1
2022 Cross-Session Motor Imagery EEG Classification using Self-Supervised Contrastive Learning
abstract
Among various Brain-Computer Interfaces (BCI) categories, Electroencephalography (EEG) signals have advantages over their other counterparts, such as fNIRS and fMRI, for their ease of acquisition and affordable recording devices. Despite the ease of data recording, the labeling of the data needs expert knowledge, which is expensive. Also, the data collected from the same subjects in different sessions are prone to varying distribution, making the cross-session data classification more challenging. Supervised learning has been widely used for EEG signal analysis, but its performance is highly dependent on a vast amount of annotated data. Self-supervised learning (SSL) has been proven to be a highly effective technique that allows deep learning architectures to learn general features from unlabelled data without needing an extensive amount of labeled data. This paper proposes a self-supervised contrastive learning method of cross-session based EEG data for motor imagery classification. The pretext task of signal transformations has been used to train the network during the training phase to reduce the distance between the similar transformation pairs and maximize the distance between the dissimilar transformation pairs generated from the EEG signals. It helps the network learn generalized features of cross-session EEG signals with the help of representation learning and in the absence of labeled data. The diverse nature of cross-session EEG data leads the network toward learning more generalized features. The performance of the proposed model upon pretraining on the BCI-IV 2a dataset using self-supervised contrastive learning and the fine-tuning on the same dataset using supervised learning gives the average accuracy of 50.81% and performed significantly better than the compared supervised learning-based methods.
Taveena Lotey, Prateek Keserwani, Gaurav Wasnik, Partha Pratim Roy 0001
ICPR2
2022 Text Region Conditional Generative Adversarial Network for Text Concealment in the Wild
abstract
Textual information appearing on the captured image may contain personal information. In various circumstances, publishing such images in the public domain may create a threat of privacy leak. To avoid these situations, we propose a text concealment method. To accomplish this task, we have used a conditional generator, which is a text region conditioned concealment network. The text regions predicted by the detector network are used as a conditioning criterion for the concealment process. The text region prediction in the form of a word-level bounding box may contain stroke pixels as well as the background pixels. However, to reduce the number of background pixels for proper conditioning of text concealment network, a character level annotation is used for the generator in place of word-level bounding box annotation. It helps to focus more on strokes as compared to the background pixels. A character-level symmetric line representation of text has been proposed to obtain finer level text region prediction as compared to the character-level bounding box. The proposed model is trainable from end-to-end. The text region conditioned generator is trained from the loss of global and local discriminators. The proposed method is validated on public scene text image datasets such as ICDAR 2015, COCO-Text and Synthesis dataset. The proposed architecture shows competitive results as compared to other state-of-the-art approaches.
Prateek Keserwani, Partha Pratim Roy 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 Robust Scene Text Detection for Partially Annotated Training Data
abstract
This article analyzed the impact of training data containing un-annotated text instances, i.e., partial annotation in scene text detection, and proposed a text region refinement approach to address it. Scene text detection is a problem that has attracted the attention of the research community for decades. Impressive results have been obtained for fully supervised scene text detection with recent deep learning approaches. These approaches, however, need a vast amount of completely labeled datasets, and the creation of such datasets is a challenging and time-consuming task. Research literature lacks the analysis of the partial annotation of training data for scene text detection. We have found that the performance of the generic scene text detection method drops significantly due to the partial annotation of training data. We have proposed a text region refinement method that provides robustness against the partially annotated training data in scene text detection. The proposed method works as a two-tier scheme. Text-probable regions are obtained in the first tier by applying hybrid loss that generates pseudo-labels to refine text regions in the second-tier during training. Extensive experiments have been conducted on a dataset generated from ICDAR 2015 by dropping the annotations with various drop rates and on a publicly available SVT dataset. The proposed method exhibits a significant improvement over the baseline and existing approaches for the partially annotated training data.
Prateek Keserwani, Rajkumar Saini, Marcus Liwicki, Partha Pratim Roy 0001
IEEE Trans. Circuits Syst. Video Technol.1
2021 Logo detection using weakly supervised saliency map
Prateek Keserwani, Partha Pratim Roy 0001, Debi Prosad Dogra
Multim. Tools Appl.2
2019 Zero Shot Learning Based Script Identification in the Wild
abstract
The text recognition system for natural images or video frames containing multilingual text needs a method to first identify the written script and then recognize the word in the identified script. However, the occurrence of some scripts is rare as compared to others. Due to the availability of a few samples of the rare script, the supervised learning of the deep neural networks is difficult. To overcome this problem, we have proposed a zero-shot learning based method for script identification. We have also proposed architecture for script identification which fuses the global feature vector and the semantic embedding vector. The semantic embedding of the script is obtained by using the spatial dependency of the stroke's sequence via the recurrent neural network. The proposed architecture shows superior results as compared to the baseline approaches.
Prateek Keserwani, Kanjar De, Partha Pratim Roy 0001, Umapada Pal 0001
ICDAR1