VLDB 2026 Research / reviewers in the wild / expert
Kuldeep Singh Yadav
dblp:266/0153
· DBLP profile ↗
17ranked-venue papers
10as first author
16since 2021 · last 2026
0000-0002-9761-9023ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards generalized real-world image super-resolution: an adaptive zero-shot and efficient generative approach for handling unknown degradations
Masuma Aktar, Kuldeep Singh Yadav, Rabul Hussain Laskar |
Neural Comput. Appl. | 2 |
| 2026 | FER20E: An Extended Facial Expression Recognition Dataset With 20 Discrete EmotionsabstractFacial emotion recognition (FER) has traditionally focused on a limited set of basic expressions, often failing to capture the complexity, subtlety, and cultural variability of real-world human emotions. To address these limitations, this paper introduces FER20E, a large-scale facial expression dataset comprising 20 emotion categories, including both basic and fine-grained affective states. The proposed emotion taxonomy is systematically derived by integrating facial action coding system (FACS)-based action units (AUs) with the valence-arousal circumplex model, ensuring both interpretability and psychological validity. To enable scalable and reliable annotation, we develop a data annotation tool (DL-DAT) that follows a semi-automated, human-in-the-loop pipeline. To validate the effectiveness and relevance of FER20E, we conduct extensive experiments using recent state-of-the-art models, including convolutional neural networks and transformer-based models. Results demonstrate that lightweight models such as MobileNetV2 and SqueezeNet achieve competitive performance while incurring significantly lower computational cost, enabling real-time deployment. Furthermore, transformer-based models with large-scale pretraining (ViT21k) achieve superior recognition accuracy, highlighting the importance of representation learning. Additional analysis reveals challenges related to emotion ambiguity, overlapping AUs, and cross-cultural variations, underscoring the need for fine-grained, robust FER systems. The FER20E dataset provides a comprehensive benchmark for advancing emotion recognition in unconstrained and real-world scenarios. The dataset and implementation details will be publicly available at https://github.com/akstheme/FER20E. Kuldeep Singh Yadav, Lalan Kumar |
IEEE Trans. Image Process. | 1 |
| 2025 | EDSGAN: An Edge-Informed Generative Adversarial Network for Enhanced Perceptual Quality Single Image Super-ResolutionabstractSingle Image Super-Resolution (SISR) faces a persistent challenge in reconstructing high-frequency edge details, which are paramount for human perceptual quality. While Generative Adversarial Networks (GANs) have significantly advanced SISR, they often struggle to generate truly sharp and realistic edges, often due to their inherent loss functions. To address this critical limitation, we propose an Edge-Informed Super-Resolution GAN (EDSGAN). EDSGAN employs a dual-path edge-informed discriminator that simultaneously analyzes image content and edge maps, enabling more effective discrimination of realistic versus artifact-prone edge structures. Concurrently, an innovative edge-aware loss guides the generator towards reconstructing perceptually sharper and more accurate edges. Extensive experiments on benchmark datasets demonstrate EDSGAN's superior perceptual quality. Notably, on the Set14 dataset, EDSGAN achieves a remarkable improvement of 9% in LPIPS (from 0.1329 to 0.121) and 8% in PI (from 2.9261 to 2.689) over the ESRGAN baseline. Our method effectively strikes an optimal balance between visual realism and pixel-level accuracy. Masuma Aktar, Kuldeep Singh Yadav, Rabul Hussain Laskar |
TENCON | 2 |
| 2024 | A cascaded deep learning framework for iris centre localization in facial imageabstractAbstract Accurate iris centre localization is crucial in many computer vision and facial biometric applications such as gaze estimation, human–computer interaction, iris recognition, and liveness detection. However, it is challenging in an uncontrolled environment due to variations like pose, scale, rotation, specular reflection, and image quality. Therefore, a cascaded deep learning framework for iris centre localization in facial images is proposed that is robust to the abovementioned variations. The proposed approach consists of (i) YOLOv3 for eye detection, (ii) UNet for iris segmentation, and (iii) statistical modelling for iris centre localization. The eyes are first detected using the YOLOv3, and subsequently, iris segmentation is performed within the detected eyes using the UNet. Following iris segmentation, statistical modelling is employed to enhance the localization accuracy of the iris centre. Experiments were performed on benchmark databases, resulting in a standardized error measure SED of 3.405 pixels for BioID and 3.259 pixels for GI4E databases. In addition, the robustness of the proposed eye detection model was further evaluated on the Yale B for illumination variations and the CAS‐PEAL for pose variations. Naseem Ahmad, Muhammad Ghulam, Kuldeep Singh Yadav, Rabul Hussain Laskar, Ashraf Hossain, Zulfiqar Ali 0001 |
Expert Syst. J. Knowl. Eng. | 3 |
| 2024 | Design and development of an integrated approach towards detection and tracking of iris using deep learning
Naseem Ahmad, Kuldeep Singh Yadav, Anish Monsley K., Saharul Alom Barlaskar, Rabul Hussain Laskar, Ashraf Hossain |
Multim. Tools Appl. | 2 |
| 2024 | Design of a two-stage ASCII recognizer for the case-sensitive inputs in handwritten and gesticulation mode of the text-entry interface
Anish Monsley K., Kuldeep Singh Yadav, Naragoni Saidulu, Saharul Alom Barlaskar, Rabul Hussain Laskar |
Multim. Tools Appl. | 2 |
| 2024 | End-to-end bare-hand localization system for human-computer interaction: a comprehensive analysis and viable solution
Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar |
Vis. Comput. | 1 |
| 2023 | GCR-Net: A deep learning-based bare hand detection and gesticulated character recognition system for human-computer interactionabstractSummary Precisely detecting bare hands and recognizing the characters are two major stages in gesticulated character recognition systems. It is very challenging to implement them in an uncontrolled environment. Additional variations, particularly (i) background feature domination (BFD) effect and motion blur in detection, (ii) gesturing style, pattern, and case sensitivity in recognition, make the system more complex. To address these challenges, a gesticulated character recognition (GCR‐Net) model is designed. To detect the bare hand precisely, a pixel‐wise segmentation approach, HandSNet, is presented, which is able to overcome the BFD effect. To handle the motion blur in the frames, a tracking module comprised of a point‐tracker and Kalman filter is applied. To reduce the computational time, a mini‐SqueezeNet network is designed, which is used in HandSNet and recognition models as the backend network. It has 0.39 million parameters only. Four separate deep convolutional neural networks (DCNNs) are connected with the network section module at the recognition end. This network selection module activates one DCNN at a time to recognize the gesticulated characters accurately. The proposed GCR‐Net reduces the complexity between similar characters and provides a high precision rate compared to the existing approaches. Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar |
Concurr. Comput. Pract. Exp. | 1 |
| 2023 | Exploration of deep learning models for localizing bare-hand in the practical environment
Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar, Naseem Ahmad |
Eng. Appl. Artif. Intell. | 1 |
| 2023 | Gesture objects detection and tracking for virtual text entry keyboard interface
Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar |
Multim. Tools Appl. | 1 |
| 2023 | Detection, tracking, and recognition of isolated multi-stroke gesticulated characters
Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar, Manas Kamal Bhuyan |
Pattern Anal. Appl. | 1 |
| 2022 | Design and development of a vision-based system for detection, tracking and recognition of isolated dynamic bare hand gesticulated charactersabstractAbstract Detection and tracking are the vital stages to form the gesture trajectory in gesture recognition. It becomes more challenging when the variations in illumination, pose, position, occlusion, scale, speed, blurring effect and complex environment are introduced. Additionally, the background feature domination effect affects the existing deep learning models. A semantic segmentation model is implemented in this work to detect the bare hand to overcome these challenges. A pre‐trained network VGG‐16 is utilized by training with the proposed NITS S‐Net database. Evaluation of the SegNet model is done on EgoHands, Oxford and OUHands databases. To track the bare hand, a SegNet‐based detection and tracking approach is proposed using Kalman filter and point‐tracker. This model achieves 97.01% accuracy (a relative improvement of ~8% from the baseline models) at 0.068 s per frame computational time on NITS hand gesture database VIIIB. The gesticulated characters, that is, alphabets, numbers, operators, special characters, are gesticulated without any constraints on the pattern/strokes. To recognize these 95 multi‐stroke gestures, a deep convolutional neural network (DCNN) is presented using AlexNet. The DCNN model achieves 97.60% (a relative improvement of ~14% from the baseline models) accuracy on the NITS hand gesture database VIIIB merged. Evaluation of the handwritten EMNIST merged (balanced) database resulted in average recognition accuracy of 91.60%. Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar, Manas Kamal Bhuyan |
Expert Syst. J. Knowl. Eng. | 1 |
| 2022 | Development of an intelligent recognition system for dynamic mid-air gesticulation of isolated alphanumeric keys
Anish Monsley K., Kuldeep Singh Yadav, Rabul Hussain Laskar, Manas Kamal Bhuyan |
Expert Syst. Appl. | 2 |
| 2022 | A selective region-based detection and tracking approach towards the recognition of dynamic bare hand gesture using deep neural network
Kuldeep Singh Yadav, Anish Monsley K., Rabul Hussain Laskar, Songhita Misra, Manas Kamal Bhuyan |
Multim. Syst. | 1 |
| 2021 | Recognition of isolated characters across different input interfaces using 2D DCNNabstractRecognition of the characters has gained much attention due to its potential applications like document analysis, license plate detection, house number detection, virtual text entry system, etc., in pattern recognition. However, it is very challenging to recognize the characters under the variations in pattern, style, translation, scale, rotation. This work develops a computationally efficient deep learning model to recognize handwritten, printable, and gesticulated characters. For gesture, the NITS gesticulated database having 60 characters (10 digits, 26 English uppercase alphabets, 4 operations, 18 special symbols) is proposed with the variation in pattern, style, scale in this work. To evaluate the ability and robustness of the proposed model, the handwritten characters (MNIST, EMNIST), printable characters (SVHN, Chars74) databases are considered. This network achieves 94.55%, 89.54%, 87.33, and 93.90% recognition accuracy on NITS gesticulated, EMNIST merge (balanced), SVHN, and Chars74 databases. Kuldeep Singh Yadav, Anish Monsley K., Saharul Alom Barlaskar, Naseem Ahmad, Rabul Hussain Laskar, Manas Kamal Bhuyan |
TENCON | 1 |
| 2021 | Segregation of meaningful strokes, a pre-requisite for self co-articulation removal in isolated dynamic gesturesabstractAbstract Gesture formation, a pre‐processing step, has its importance when variations in patterns, scale, and speed come into play. Self co‐articulations are intentional movements performed by an individual to complete a gesture, whose presence in the trajectory alters its original meaning. For recognition, most researchers have directly used the trajectory formed along with these self co‐articulated strokes, with a few removing it using visible trait‐like velocity. Usage of velocity has shortcomings as gesturing in air differs from gesturing over a solid surface; hence, we propose a gesture formation model, which incorporates global and local measures to remove these self co‐articulations. The global measure uses Euclidean distance, instantaneous velocity, and polarity calculated from the complete gesture, while the local measure segments the gesture into stroke‐level segments by using the minimum–maximum‐polarity algorithm and applies the selective bypass rules. The proposed model, when experimented on gestures patterns with premeditated speed variation, has a mean error rate of 0.0069 and 7.40% self co‐articulations;individuals’ natural gesticulation has a mean error rate of 0.0371 and 12.07% self co‐articulations. Experimentation on each gesture of NITS hand gesture databases showed a relative improvement of 40% (accuracy 97%) over the existing baseline models. Anish Monsley K., Kuldeep Singh Yadav, Songhita Misra, Manas Kamal Bhuyan, Rabul Hussain Laskar |
IET Image Process. | 2 |
| 2020 | Facial expression recognition using modified Viola-John's algorithm and KNN classifier
Kuldeep Singh Yadav, Joyeeta Singha |
Multim. Tools Appl. | 1 |