Rajkumar Saini

dblp:181/4611 · DBLP profile ↗
← Back
29ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0001-8532-0895ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Modified TSception for analyzing driver drowsiness and mental workload from EEG
Gourav Siddhad, Rajkumar Saini, Partha Pratim Roy 0001
Neural Comput. Appl.3
2025 ASTrA: Adversarial Self-supervised Training with Adaptive-Attacks
abstract
Existing self-supervised adversarial training (self-AT) methods rely on hand-crafted adversarial attack strategies for PGD attacks, which fail to adapt to the evolving learning dynamics of the model and do not account for instance-specific characteristics of images. This results in sub-optimal adversarial robustness and limits the alignment between clean and adversarial data distributions. To address this, we propose $\textit{ASTrA}$ ($\textbf{A}$dversarial $\textbf{S}$elf-supervised $\textbf{Tr}$aining with $\textbf{A}$daptive-Attacks), a novel framework introducing a learnable, self-supervised attack strategy network that autonomously discovers optimal attack parameters through exploration-exploitation in a single training episode. ASTrA leverages a reward mechanism based on contrastive loss, optimized with REINFORCE, enabling adaptive attack strategies without labeled data or additional hyperparameters. We further introduce a mixed contrastive objective to align the distribution of clean and adversarial examples in representation space. ASTrA achieves state-of-the-art results on CIFAR10, CIFAR100, and STL10 while integrating seamlessly as a plug-and-play module for other self-AT methods. ASTrA shows scalability to larger datasets, demonstrates strong semi-supervised performance, and is resilient to robust overfitting, backed by explainability analysis on optimal attack strategies. Project page for source code and other details at https://prakashchhipa.github.io/projects/ASTrA.
Prakash Chandra Chhipa, Gautam Vashishtha, Settur Jithamanyu, Rajkumar Saini, Mubarak Shah, Marcus Liwicki
ICLR4
2024 LCM: Log Conformal Maps for Robust Representation Learning to Mitigate Perspective Distortion
Meenakshi Subhash Chippa, Prakash Chandra Chhipa, Kanjar De, Marcus Liwicki, Rajkumar Saini
ACCV (8)5
2024 Möbius Transform for Mitigating Perspective Distortions in Representation Learning
Prakash Chandra Chhipa, Meenakshi Subhash Chippa, Kanjar De, Rajkumar Saini, Marcus Liwicki, Mubarak Shah
ECCV (73)4
2024 Attention Dynamics: Estimating Attention Levels of ADHD using Swin Transformer
Debashis Das Chakladar, Anand Shankar, Foteini Liwicki, Shovan Barma, Rajkumar Saini
ICPR (11)5
2024 Vehicle Detection Performance in Nordic Region
Hamam Mokayed, Rajkumar Saini, Oluwatosin Adewumi, Lama Alkhaled, Björn Backe, Palaiahnakote Shivakumara, Olle Hagner, Yan Chai Hum
ICPR (22)2
2024 SimBrainNet: Evaluating Brain Network Similarity for Attention Disorders
Debashis Das Chakladar, Foteini Liwicki, Rajkumar Saini
MICCAI (2)3
2023 ICDAR 2023 CROHME: Competition on Recognition of Handwritten Mathematical Expressions
Yejing Xie, Harold Mouchère, Foteini Liwicki, Sumit Rakesh, Rajkumar Saini, Masaki Nakagawa, Cuong Tuan Nguyen, Thanh-Nghia Truong
ICDAR (2)5
2023 Functional Knowledge Transfer with Self-supervised Representation Learning
abstract
This work investigates the unexplored usability of self-supervised representation learning in the direction of functional knowledge transfer. In this work, functional knowledge transfer is achieved by joint optimization of self-supervised learning pseudo task and supervised learning task, improving supervised learning task performance. Recent progress in self-supervised learning uses a large volume of data, which becomes a constraint for its applications on small-scale datasets. This work shares a simple yet effective joint training framework that reinforces human-supervised task learning by learning self-supervised representations just-in-time and vice versa. Experiments on three public datasets from different visual domains, Intel Image, CIFAR, and APTOS, reveal a consistent track of performance improvements on classification tasks during joint optimization. Qualitative analysis also supports the robustness of learnt representations. Source code and trained models are available on GitHub1.
Prakash Chandra Chhipa, Muskaan Chopra, Gopal Mengi, Varun Gupta 0005, Richa Upadhyay, Meenakshi Subhash Chippa, Kanjar De, Rajkumar Saini, Seiichi Uchida, Marcus Liwicki
ICIP8
2023 Multi-Task Meta Learning: learn how to adapt to unseen tasks
abstract
This work proposes Multi-task Meta Learning (MTML), integrating two learning paradigms Multi-Task Learning (MTL) and meta learning, to bring together the best of both worlds. In particular, it focuses simultaneous learning of multiple tasks, an element of MTL and promptly adapting to new tasks, a quality of meta learning. It is important to highlight that we focus on heterogeneous tasks, which are of distinct kind, in contrast to typically considered homogeneous tasks (e.g., if all tasks are classification or if all tasks are regression tasks). The fundamental idea is to train a multi-task model, such that when an unseen task is introduced, it can learn in fewer steps whilst offering a performance at least as good as conventional single task learning on the new task or inclusion within the MTL. By conducting various experiments, we demonstrate this paradigm on two datasets and four tasks: NYU-v2 and the taskonomy dataset for which we perform semantic segmentation, depth estimation, surface normal estimation, and edge detection. MTML achieves state-of-the-art results for three out of four tasks for the NYU-v2 dataset and two out of four for the taskonomy dataset. In the taskonomy dataset, it was discovered that many pseudo-labeled segmentation masks lacked classes that were expected to be present in the ground truth; however, our MTML approach was found to be effective in detecting these missing classes, delivering good qualitative results. While, quantitatively its performance was affected due to the presence of incorrect ground truth labels. The the source code for reproducibility can be found at https://github.com/ricupa/MTML-learn-how-to-adapt-to-unseen-tasks.
Richa Upadhyay, Prakash Chandra Chhipa, Ronald Phlypo, Rajkumar Saini, Marcus Liwicki
IJCNN4
2023 Magnification Prior: A Self-Supervised Method for Learning Representations on Breast Cancer Histopathological Images
abstract
This work presents a novel self-supervised pre-training method to learn efficient representations without labels on histopathology medical images utilizing magnification factors. Other state-of-the-art works mainly focus on fully supervised learning approaches that rely heavily on human annotations. However, the scarcity of labeled and unlabeled data is a long-standing challenge in histopathology. Currently, representation learning without labels remains unexplored in the histopathology domain. The proposed method, Magnification Prior Contrastive Similarity (MPCS), enables self-supervised learning of representations without labels on small-scale breast cancer dataset BreakHis by exploiting magnification factor, inductive transfer, and reducing human prior. The proposed method matches fully supervised learning state-of-the-art performance in malignancy classification when only 20% of labels are used in fine-tuning and outperform previous works in fully supervised learning settings for three public breast cancer datasets, including BreakHis. Further, It provides initial support for a hypothesis that reducing human-prior leads to efficient representation learning in self-supervision, which will need further investigation. The implementation of this work is available online on GitHub1.
Prakash Chandra Chhipa, Richa Upadhyay, Gustav Pihlgren, Rajkumar Saini, Seiichi Uchida, Marcus Liwicki
WACV4
2022 Robust Scene Text Detection for Partially Annotated Training Data
abstract
This article analyzed the impact of training data containing un-annotated text instances, i.e., partial annotation in scene text detection, and proposed a text region refinement approach to address it. Scene text detection is a problem that has attracted the attention of the research community for decades. Impressive results have been obtained for fully supervised scene text detection with recent deep learning approaches. These approaches, however, need a vast amount of completely labeled datasets, and the creation of such datasets is a challenging and time-consuming task. Research literature lacks the analysis of the partial annotation of training data for scene text detection. We have found that the performance of the generic scene text detection method drops significantly due to the partial annotation of training data. We have proposed a text region refinement method that provides robustness against the partially annotated training data in scene text detection. The proposed method works as a two-tier scheme. Text-probable regions are obtained in the first tier by applying hybrid loss that generates pseudo-labels to refine text regions in the second-tier during training. Extensive experiments have been conducted on a dataset generated from ICDAR 2015 by dropping the annotations with various drop rates and on a publicly available SVT dataset. The proposed method exhibits a significant improvement over the baseline and existing approaches for the partially annotated training data.
Prateek Keserwani, Rajkumar Saini, Marcus Liwicki, Partha Pratim Roy 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 3D word spotting using leap motion sensor
Partha Pratim Roy 0001, Pradeep Kumar 0002, Shweta Patidar, Rajkumar Saini
Multim. Tools Appl.4
2021 Modeling local and global behavior for trajectory classification using graph based algorithm
Rajkumar Saini, Pradeep Kumar 0002, Partha Pratim Roy 0001, Umapada Pal 0001
Pattern Recognit. Lett.1
2019 ICDAR 2019 Historical Document Reading Challenge on Large Structured Chinese Family Records
abstract
In this paper, we present a large historical database of Chinese family records with the aim to develop robust systems for historical document analysis. In this direction, we propose a Historical Document Reading Challenge on Large Chinese Structured Family Records (ICDAR 2019 HDRC-CHINESE). The objective of the competition is to recognize and analyze the layout, and finally detect and recognize the textlines and characters of the large historical document image dataset containing more than 100000 pages. Cascade R-CNN, CRNN, and U-Net based architectures were trained to evaluate the performances in these tasks. Error rate of 0.01 has been recorded for textline recognition (Task1) whereas a Jaccard Index of 99:54% has been recorded for layout analysis (Task2). The graph edit distance based total error ratio of 1:5% has been recorded for complete integrated textline detection and recognition (Task3).
Rajkumar Saini, Derek Dobson, Jon Morrey, Marcus Liwicki, Foteini Liwicki
ICDAR1
2019 Recognizing gender from human facial regions using genetic algorithm
Avirup Bhattacharyya, Rajkumar Saini, Partha Pratim Roy 0001, Debi Prosad Dogra, Samarjit Kar
Soft Comput.2
2019 Multimodal Gait Recognition With Inertial Sensor Data and Video Using Evolutionary Algorithm
abstract
Evolutionary decision fusion has applications in biometric authentication and verification. Gray wolf optimizer (GWO) is one such evolutionary decision fusion approach that can be used to tune the fusion parameters in a multimodal data acquisition system. Human gait is a proven biometric trait with applications in security and authentication. However, acquiring human-gait data can be erroneous due to various factors and multimodal fusion of such erroneous gait data can be challenging. In this paper, we propose a new decision fusion-based approach to solve the above problem. Gait data is recorded simultaneously using motion sensors and visible-light camera. The signals of the motion sensors are modeled using a long short-term memory neural network and corresponding video recordings are processed using a three-dimensional convolutional neural network. GWO has been used to optimize the parameters during fusion. It has been chosen based on the underlying hunting strategy that leads to better approximation of the solution. Interestingly, in our case it converges quicker than other optimization techniques such as genetic algorithm or particle swarm optimization. To test the model, a dataset involving 23 males and females has been recorded while they perform four different types of walks, including, normal walk, fast walk, walking while listening to music, and walking while watching multimedia content on a mobile. An overall accuracy of 91.3% has been recorded across all test scenarios. Results reveal that the proposed study can further be explored to design robust gait biometric systems.
Pradeep Kumar 0002, Subham Mukherjee, Rajkumar Saini, Pallavi Kaushik, Partha Pratim Roy 0001, Debi Prosad Dogra
IEEE Trans. Fuzzy Syst.3
2019 A novel point-line duality feature for trajectory classification
Rajkumar Saini, Partha Pratim Roy 0001, Debi Prosad Dogra
Vis. Comput.1
2018 A segmental HMM based trajectory classification using genetic algorithm
Rajkumar Saini, Partha Pratim Roy 0001, Debi Prosad Dogra
Expert Syst. Appl.1
2018 A novel framework of continuous human-activity recognition using Kinect
Rajkumar Saini, Pradeep Kumar 0002, Partha Pratim Roy 0001, Debi Prosad Dogra
Neurocomputing1
2018 Don't just sign use brain too: A novel multimodal approach for user identification and verification
Rajkumar Saini, Barjinder Kaur, Priyanka Singh 0001, Pradeep Kumar 0002, Partha Pratim Roy 0001, Balasubramanian Raman
Inf. Sci.1
2018 A position and rotation invariant framework for sign language recognition (SLR) using Kinect
Pradeep Kumar 0002, Rajkumar Saini, Partha Pratim Roy 0001, Debi Prosad Dogra
Multim. Tools Appl.2
2018 Frame selection for OCR from video stream of book flipping
Dibyayan Chakraborty, Partha Pratim Roy 0001, Rajkumar Saini, José M. Álvarez 0004, Umapada Pal 0001
Multim. Tools Appl.3
2018 A lexicon-free approach for 3D handwriting recognition using classifier combination
Pradeep Kumar 0002, Rajkumar Saini, Partha Pratim Roy 0001, Umapada Pal 0001
Pattern Recognit. Lett.2
2018 Envisioned speech recognition using EEG sensors
Pradeep Kumar 0002, Rajkumar Saini, Partha Pratim Roy 0001, Pawan Kumar Sahu, Debi Prosad Dogra
Pers. Ubiquitous Comput.2
2017 A bio-signal based framework to secure mobile devices
Pradeep Kumar 0002, Rajkumar Saini, Partha Pratim Roy 0001, Debi Prosad Dogra
J. Netw. Comput. Appl.2
2017 3D text segmentation and recognition using leap motion
Pradeep Kumar 0002, Rajkumar Saini, Partha Pratim Roy 0001, Debi Prosad Dogra
Multim. Tools Appl.2
2017 Analysis of EEG signals and its application to neuromarketing
Mahendra Yadava, Pradeep Kumar 0002, Rajkumar Saini, Partha Pratim Roy 0001, Debi Prosad Dogra
Multim. Tools Appl.3
2016 On the applicability of diploid genetic algorithms in dynamic environments
Harsh Bhasin, Gitanshu Behal, Nimish Aggarwal, Rajkumar Saini, Shivani Choudhary
Soft Comput.4