Kuldeep Singh Yadav

dblp:266/0153 · DBLP profile ↗
← Back
17ranked-venue papers
10as first author
16since 2021 · last 2026
0000-0002-9761-9023ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Towards generalized real-world image super-resolution: an adaptive zero-shot and efficient generative approach for handling unknown degradations
Masuma Aktar, Kuldeep Singh Yadav, Rabul Hussain Laskar
Neural Comput. Appl.2
2026 FER20E: An Extended Facial Expression Recognition Dataset With 20 Discrete Emotions
abstract
Facial emotion recognition (FER) has traditionally focused on a limited set of basic expressions, often failing to capture the complexity, subtlety, and cultural variability of real-world human emotions. To address these limitations, this paper introduces FER20E, a large-scale facial expression dataset comprising 20 emotion categories, including both basic and fine-grained affective states. The proposed emotion taxonomy is systematically derived by integrating facial action coding system (FACS)-based action units (AUs) with the valence-arousal circumplex model, ensuring both interpretability and psychological validity. To enable scalable and reliable annotation, we develop a data annotation tool (DL-DAT) that follows a semi-automated, human-in-the-loop pipeline. To validate the effectiveness and relevance of FER20E, we conduct extensive experiments using recent state-of-the-art models, including convolutional neural networks and transformer-based models. Results demonstrate that lightweight models such as MobileNetV2 and SqueezeNet achieve competitive performance while incurring significantly lower computational cost, enabling real-time deployment. Furthermore, transformer-based models with large-scale pretraining (ViT21k) achieve superior recognition accuracy, highlighting the importance of representation learning. Additional analysis reveals challenges related to emotion ambiguity, overlapping AUs, and cross-cultural variations, underscoring the need for fine-grained, robust FER systems. The FER20E dataset provides a comprehensive benchmark for advancing emotion recognition in unconstrained and real-world scenarios. The dataset and implementation details will be publicly available at https://github.com/akstheme/FER20E.
Kuldeep Singh Yadav, Lalan Kumar
IEEE Trans. Image Process.1
2025 EDSGAN: An Edge-Informed Generative Adversarial Network for Enhanced Perceptual Quality Single Image Super-Resolution
abstract
Single Image Super-Resolution (SISR) faces a persistent challenge in reconstructing high-frequency edge details, which are paramount for human perceptual quality. While Generative Adversarial Networks (GANs) have significantly advanced SISR, they often struggle to generate truly sharp and realistic edges, often due to their inherent loss functions. To address this critical limitation, we propose an Edge-Informed Super-Resolution GAN (EDSGAN). EDSGAN employs a dual-path edge-informed discriminator that simultaneously analyzes image content and edge maps, enabling more effective discrimination of realistic versus artifact-prone edge structures. Concurrently, an innovative edge-aware loss guides the generator towards reconstructing perceptually sharper and more accurate edges. Extensive experiments on benchmark datasets demonstrate EDSGAN's superior perceptual quality. Notably, on the Set14 dataset, EDSGAN achieves a remarkable improvement of 9% in LPIPS (from 0.1329 to 0.121) and 8% in PI (from 2.9261 to 2.689) over the ESRGAN baseline. Our method effectively strikes an optimal balance between visual realism and pixel-level accuracy.
Masuma Aktar, Kuldeep Singh Yadav, Rabul Hussain Laskar
TENCON2
2024 A cascaded deep learning framework for iris centre localization in facial image
abstract
Abstract Accurate iris centre localization is crucial in many computer vision and facial biometric applications such as gaze estimation, human–computer interaction, iris recognition, and liveness detection. However, it is challenging in an uncontrolled environment due to variations like pose, scale, rotation, specular reflection, and image quality. Therefore, a cascaded deep learning framework for iris centre localization in facial images is proposed that is robust to the abovementioned variations. The proposed approach consists of (i) YOLOv3 for eye detection, (ii) UNet for iris segmentation, and (iii) statistical modelling for iris centre localization. The eyes are first detected using the YOLOv3, and subsequently, iris segmentation is performed within the detected eyes using the UNet. Following iris segmentation, statistical modelling is employed to enhance the localization accuracy of the iris centre. Experiments were performed on benchmark databases, resulting in a standardized error measure SED of 3.405 pixels for BioID and 3.259 pixels for GI4E databases. In addition, the robustness of the proposed eye detection model was further evaluated on the Yale B for illumination variations and the CAS‐PEAL for pose variations.
Naseem Ahmad, Muhammad Ghulam, Kuldeep Singh Yadav, Rabul Hussain Laskar, Ashraf Hossain, Zulfiqar Ali 0001
Expert Syst. J. Knowl. Eng.3
2024 Design and development of an integrated approach towards detection and tracking of iris using deep learning
Naseem Ahmad, Kuldeep Singh Yadav, Anish Monsley K., Saharul Alom Barlaskar, Rabul Hussain Laskar, Ashraf Hossain
Multim. Tools Appl.2
2024 Design of a two-stage ASCII recognizer for the case-sensitive inputs in handwritten and gesticulation mode of the text-entry interface
Anish Monsley K., Kuldeep Singh Yadav, Naragoni Saidulu, Saharul Alom Barlaskar, Rabul Hussain Laskar
Multim. Tools Appl.2
2024 End-to-end bare-hand localization system for human-computer interaction: a comprehensive analysis and viable solution
Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar
Vis. Comput.1
2023 GCR-Net: A deep learning-based bare hand detection and gesticulated character recognition system for human-computer interaction
abstract
Summary Precisely detecting bare hands and recognizing the characters are two major stages in gesticulated character recognition systems. It is very challenging to implement them in an uncontrolled environment. Additional variations, particularly (i) background feature domination (BFD) effect and motion blur in detection, (ii) gesturing style, pattern, and case sensitivity in recognition, make the system more complex. To address these challenges, a gesticulated character recognition (GCR‐Net) model is designed. To detect the bare hand precisely, a pixel‐wise segmentation approach, HandSNet, is presented, which is able to overcome the BFD effect. To handle the motion blur in the frames, a tracking module comprised of a point‐tracker and Kalman filter is applied. To reduce the computational time, a mini‐SqueezeNet network is designed, which is used in HandSNet and recognition models as the backend network. It has 0.39 million parameters only. Four separate deep convolutional neural networks (DCNNs) are connected with the network section module at the recognition end. This network selection module activates one DCNN at a time to recognize the gesticulated characters accurately. The proposed GCR‐Net reduces the complexity between similar characters and provides a high precision rate compared to the existing approaches.
Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar
Concurr. Comput. Pract. Exp.1
2023 Exploration of deep learning models for localizing bare-hand in the practical environment
Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar, Naseem Ahmad
Eng. Appl. Artif. Intell.1
2023 Gesture objects detection and tracking for virtual text entry keyboard interface
Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar
Multim. Tools Appl.1
2023 Detection, tracking, and recognition of isolated multi-stroke gesticulated characters
Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar, Manas Kamal Bhuyan
Pattern Anal. Appl.1
2022 Design and development of a vision-based system for detection, tracking and recognition of isolated dynamic bare hand gesticulated characters
abstract
Abstract Detection and tracking are the vital stages to form the gesture trajectory in gesture recognition. It becomes more challenging when the variations in illumination, pose, position, occlusion, scale, speed, blurring effect and complex environment are introduced. Additionally, the background feature domination effect affects the existing deep learning models. A semantic segmentation model is implemented in this work to detect the bare hand to overcome these challenges. A pre‐trained network VGG‐16 is utilized by training with the proposed NITS S‐Net database. Evaluation of the SegNet model is done on EgoHands, Oxford and OUHands databases. To track the bare hand, a SegNet‐based detection and tracking approach is proposed using Kalman filter and point‐tracker. This model achieves 97.01% accuracy (a relative improvement of ~8% from the baseline models) at 0.068 s per frame computational time on NITS hand gesture database VIIIB. The gesticulated characters, that is, alphabets, numbers, operators, special characters, are gesticulated without any constraints on the pattern/strokes. To recognize these 95 multi‐stroke gestures, a deep convolutional neural network (DCNN) is presented using AlexNet. The DCNN model achieves 97.60% (a relative improvement of ~14% from the baseline models) accuracy on the NITS hand gesture database VIIIB merged. Evaluation of the handwritten EMNIST merged (balanced) database resulted in average recognition accuracy of 91.60%.
Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar, Manas Kamal Bhuyan
Expert Syst. J. Knowl. Eng.1
2022 Development of an intelligent recognition system for dynamic mid-air gesticulation of isolated alphanumeric keys
Anish Monsley K., Kuldeep Singh Yadav, Rabul Hussain Laskar, Manas Kamal Bhuyan
Expert Syst. Appl.2
2022 A selective region-based detection and tracking approach towards the recognition of dynamic bare hand gesture using deep neural network
Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar, Songhita Misra, Manas Kamal Bhuyan
Multim. Syst.1
2021 Recognition of isolated characters across different input interfaces using 2D DCNN
abstract
Recognition of the characters has gained much attention due to its potential applications like document analysis, license plate detection, house number detection, virtual text entry system, etc., in pattern recognition. However, it is very challenging to recognize the characters under the variations in pattern, style, translation, scale, rotation. This work develops a computationally efficient deep learning model to recognize handwritten, printable, and gesticulated characters. For gesture, the NITS gesticulated database having 60 characters (10 digits, 26 English uppercase alphabets, 4 operations, 18 special symbols) is proposed with the variation in pattern, style, scale in this work. To evaluate the ability and robustness of the proposed model, the handwritten characters (MNIST, EMNIST), printable characters (SVHN, Chars74) databases are considered. This network achieves 94.55%, 89.54%, 87.33, and 93.90% recognition accuracy on NITS gesticulated, EMNIST merge (balanced), SVHN, and Chars74 databases.
Kuldeep Singh Yadav, Anish Monsley K., Saharul Alom Barlaskar, Naseem Ahmad, Rabul Hussain Laskar, Manas Kamal Bhuyan
TENCON1
2021 Segregation of meaningful strokes, a pre-requisite for self co-articulation removal in isolated dynamic gestures
abstract
Abstract Gesture formation, a pre‐processing step, has its importance when variations in patterns, scale, and speed come into play. Self co‐articulations are intentional movements performed by an individual to complete a gesture, whose presence in the trajectory alters its original meaning. For recognition, most researchers have directly used the trajectory formed along with these self co‐articulated strokes, with a few removing it using visible trait‐like velocity. Usage of velocity has shortcomings as gesturing in air differs from gesturing over a solid surface; hence, we propose a gesture formation model, which incorporates global and local measures to remove these self co‐articulations. The global measure uses Euclidean distance, instantaneous velocity, and polarity calculated from the complete gesture, while the local measure segments the gesture into stroke‐level segments by using the minimum–maximum‐polarity algorithm and applies the selective bypass rules. The proposed model, when experimented on gestures patterns with premeditated speed variation, has a mean error rate of 0.0069 and 7.40% self co‐articulations;individuals’ natural gesticulation has a mean error rate of 0.0371 and 12.07% self co‐articulations. Experimentation on each gesture of NITS hand gesture databases showed a relative improvement of 40% (accuracy 97%) over the existing baseline models.
Anish Monsley K., Kuldeep Singh Yadav, Songhita Misra, Manas Kamal Bhuyan, Rabul Hussain Laskar
IET Image Process.2
2020 Facial expression recognition using modified Viola-John's algorithm and KNN classifier
Kuldeep Singh Yadav, Joyeeta Singha
Multim. Tools Appl.1