Sanjay Singh 0001

dblp:27/6533-1 · DBLP profile ↗
← Back
33ranked-venue papers
0as first author
31since 2021 · last 2026
0000-0002-2249-799XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 The Action-Engagement-Collaboration Triad: A Multimodal Analytical Framework for Human-Robot Collaboration
abstract
Industrial human-robot-collaboration (HRC) depends on under-standing how people act, stay engaged, and work together with robots during shared tasks. In most prior work, these aspects are studied in isolation, which makes it hard to see the full picture of real collaboration. This study tackles that problem with an improved multimodal framework that jointly captures fine-grained human actions, engagement levels, and collaboration outcomes in an industrial assembly setting. It introduces a three-layer analytical model called MICRO-MESO-MACRO (M3), which links detailed action patterns to engagement dynamics and overall system efficiency. All data streams are aligned in time using a custom synchronization process so that visual, motion, and behavioral signals can be analyzed together in a consistent way. The results show clear relationships between varied action patterns, stable engagement, and better collaboration efficiency, providing a solid foundation for designing adaptive, human-aware robotic partners. The annotated dataset and code are available at: https://github.com/arvindsihag/m3_analyzer.
Arvind 0002, Naval Kishore Mehta, Sumeet Saurav, Sanjay Singh 0001
HRI5
2025 A Multimodal Dataset for Enhancing Industrial Task Monitoring and Engagement Prediction
abstract
Detecting and interpreting operator actions, engagement, and object interactions in dynamic industrial workflows remains a significant challenge in human-robot collaboration research, especially within complex, real-world environments. Traditional unimodal methods often fall short of capturing the intricacies of these unstructured industrial settings. To address this gap, we present a novel Multimodal Industrial Activity Monitoring (MIAM) dataset that captures realistic assembly and disassembly tasks, facilitating the evaluation of key meta-tasks such as action localization, object interaction, and engagement prediction. The dataset comprises multi-view RGB, depth, and Inertial Measurement Unit (IMU) data collected from 22 sessions, amounting to 290 minutes of untrimmed video, annotated in detail for task performance and operator behavior. Its distinctiveness lies in the integration of multiple data modalities and its emphasis on real-world, untrimmed industrial workflows-key for advancing research in human-robot collaboration and operator monitoring. Additionally, we propose a multimodal network that fuses RGB frames, IMU data, and skeleton sequences to predict engagement levels during industrial tasks. Our approach improves the accuracy of recognizing engagement states, providing a robust solution for monitoring operator performance in dynamic industrial environments. The dataset and code can be accessed from https://github.com/navalkishoremehta95/MIAM/.
Naval Kishore Mehta, Arvind 0002, Abeer Banerjee, Sumeet Saurav, Sanjay Singh 0001
HRI6
2025 Enhanced YOLO Object Detector for Insulator Defect Detection in Power Line Infrastructure
abstract
Deep learning has shown remarkable capabilities in automatic defect detection in power line infrastructure, but the scarcity of defect-specific labeled datasets often limits its effectiveness. This work addresses the critical challenge of detecting missing disc insulator defects under data-limited conditions by proposing a data augmentation-enhanced YOLOv12 framework. Starting with only 128 original defect images, we systematically applied geometric augmentations, including multi-angle rotations$({10}^{\circ}, 20}^{\circ}, {30}^{\circ}$) and spatial shifts (horizontal/vertical shifts of 0.1−0.3) to generate 27 distinct variations per image. This strategy expanded the dataset to 3, 456 synthetic samples, enriching defect diversity while preserving realistic defect characteristics. The YOLOv12-based framework was evaluated using 5-fold cross-validation, with parallel GPU training used to fully utilize computational resources and reduce training time. Experimental results demonstrate that the diversity of synthetic data, combined with the advanced detection capabilities of YOLOv12, significantly improves model robustness, achieving a 30-34% increase in mAP over non-augmented training and surpassing existing augmentation-based and improved fault detection methods. This study provides a practical approach to overcome data scarcity and advance reliable defect detection in power line inspection applications.
Seema Choudhary, Sumeet Saurav, Ravi Saini, Sanjay Singh 0001
TENCON4
2025 Mobile Application for Real-Time Fabric Pattern Classification to Assist Visually Impaired and Blind: A Proof-of-Concept Implementation
abstract
ABSTRACT Visual impairment has a drastic impact on the psychological and cognitive well‐being of individuals. Recent progress in advanced assistive technologies (AATs) has emerged as an essential tool to mitigate the adverse impact of blindness and enhance the quality of life of visually impaired persons (VIPs). Like generic object identification, the VIPs face difficulties in identifying their garments. Such a limitation severely impacts their identity as they cannot select dresses according to their preferences for different contexts and occasions. To this end, in this paper, we present a proof‐of‐concept (POC) implementation of a mobile application for real‐time fabric pattern classification to assist VIPs in selecting the cloth with fabric patterns of their choice. The proposed framework uses a robust and compute‐efficient convolutional neural network (CNN) named FabricNet to classify four types of fabric patterns (lattice, printed, solid, and stripe). The designed FabricNet model uses efficient feature enhancement (EFE), efficient feature refinement (EFR), and enhanced feature fusion (EFF) blocks to extract discriminative texture features from the fabric pattern images. We evaluated the performance of the proposed FabricNet on a recent open‐source fabric image dataset. Compared to the existing baseline, the FabricNet model, with 1.64 M parameters, 6.25 MB memory storage size, and 4.09 GFLOPs, attained state‐of‐the‐art classification accuracy with reduced model storage size and competitive computation time. Finally, we optimized the trained model and integrated it with a mobile application developed using the Android Studio software development kit (SDK) for real‐time classification of fabric images. The designed mobile application on an Android mobile phone uses its back camera to capture garment images and provides predicted fabric type using audio feedback. Thus, it can assist VIPs in selecting clothing with fabric patterns of their choice in real time.
Sumeet Saurav, Seema Choudhary, Sanjay Singh 0001
Comput. Intell.3
2025 Towards lensless image deblurring with prior-embedded implicit neural representations in the low-data regime
Abeer Banerjee, Sanjay Singh 0001
Eng. Appl. Artif. Intell.2
2025 An integrated attention-guided deep convolutional neural network for facial expression recognition in the wild
Sumeet Saurav, Ravi Saini, Sanjay Singh 0001
Multim. Tools Appl.3
2025 Optimizing Multitask Industrial Processes With Predictive Action Guidance
Naval Kishore Mehta, Arvind 0002, Shyam Sunder Prasad, Sumeet Saurav, Sanjay Singh 0001
IEEE Trans Autom. Sci. Eng.5
2024 DF Sampler: A Self-Supervised Method for Adaptive Keyframe Sampling
abstract
Video understanding has become crucial in computer vision research due to the vast amounts of video data generated from various sources. Efficient video analysis, especially the extraction of keyframes from long video sequences, is crucial when combined with state-of-the-art deep learning algorithms for video understanding. This has numerous applications, including smart surveillance, autonomous driving, video summarization, and human activity monitoring in human-robot collaboration (HRC) settings. In the context of Industry 5.0, such analysis aids in enhancing productivity, safety, and ergonomics for industrial workers. However, there is a notable lack of research on adaptive frame selection for video-based Human Activity Recognition (HAR) systems in challenging, uncontrolled industrial environments, leading to high memory and computational demands. To address this, we introduce Dynamic Frame Sampler (DF Sampler), a novel adaptive keyframe sampling method designed to tackle the challenges of processing variable-duration action sequences in complex industrial environments. DF Sampler uses a self-supervised learning approach to adaptively select keyframes, prioritizing segments with significant motion while optimizing the computational efficiency of HAR systems. The method is evaluated on the publicly available HRI30 dataset.
Naval Kishore Mehta, Shyam Sunder Prasad, Sumeet Saurav, Sanjay Singh 0001
ETFA4
2024 PROACT: Anticipatory Action Modeling and Anomaly Prevention in Industrial Workflows
abstract
Human-Robot Collaboration (HRC) requires seamless interaction between humans and robots, necessitating an accurate understanding and response to human intentions, despite their unpredictability. This challenge is particularly acute in industrial settings, where the complexity of tasks and the variability of human actions demand robust solutions. Effective Human Activity Recognition (HAR) is essential, particularly where real-time anomaly detection is critical for maintaining productivity and efficiency. To address these challenges, we introduce PROACT, a novel framework that predicts operator actions and detects anomalies by constructing a reference graph from action sequences. By anticipating behavior and identifying deviations early, PROACT enhances safety and operational efficiency. The framework has been validated using the Meccano dataset, which simulates an industrial-like environment, demonstrating its capability in monitoring, proactive guidance, and anomaly detection.
Naval Kishore Mehta, Arvind 0002, Shyam Sunder Prasad, Sumeet Saurav, Sanjay Singh 0001
IECON5
2024 Generalized Gaze-Vector Estimation in Low-light with Encoded Event-driven Neural Network
abstract
In this paper, we address the intricate challenge of gaze vector prediction, a pivotal task with applications ranging from human-computer interaction to driver monitoring systems. Our innovative approach is designed for the demanding setting of extremely low-light conditions, leveraging a novel temporal event encoding scheme, and a dedicated neural network architecture. The temporal encoding method seamlessly integrates Dynamic Vision Sensor (DVS) events with grayscale guide frames, generating consecutively encoded images for input into our neural network. This unique solution not only captures diverse gaze responses from participants within the active age group but also introduces a curated dataset tailored for low-light conditions. The encoded temporal frames paired with our network showcase impressive spatial localization and reliable gaze direction in their predictions. Achieving a remarkable 100-pixel accuracy of 100%, our research underscores the potency of our neural network to work with temporally consecutive encoded images for precise gaze vector predictions in challenging low-light videos, contributing to the advancement of gaze prediction technologies.
Abeer Banerjee, Naval Kishore Mehta, Shyam Sunder Prasad, Sumeet Saurav, Sanjay Singh 0001
IJCNN6
2024 Object detection in power line infrastructure: A review of the challenges and solutions
Pratibha Sharma, Sumeet Saurav, Sanjay Singh 0001
Eng. Appl. Artif. Intell.3
2024 Attention-guided generator with dual discriminator GAN for real-time video anomaly detection
Rituraj Singh, Anikeit Sethi, Krishanu Saini, Sumeet Saurav, Aruna Tiwari, Sanjay Singh 0001
Eng. Appl. Artif. Intell.6
2024 CVAD-GAN: Constrained video anomaly detection via generative adversarial network
Rituraj Singh, Anikeit Sethi, Krishanu Saini, Sumeet Saurav, Aruna Tiwari, Sanjay Singh 0001
Image Vis. Comput.6
2024 Exploration of deep learning architectures for real-time yoga pose recognition
Sumeet Saurav, Prashant Sadashiv Gidde, Sanjay Singh 0001
Multim. Tools Appl.3
2024 ParaColorizer-Realistic image colorization using parallel generative networks
Abeer Banerjee, Sumeet Saurav, Sanjay Singh 0001
Vis. Comput.4
2023 Reconstructing Synthetic Lensless Images in the Low-Data Regime
Abeer Banerjee, Sumeet Saurav, Sanjay Singh 0001
BMVC4
2023 Physics-Informed Deep Deblurring: Over-parameterized vs. Under-parameterized
abstract
Image deblurring is a classic problem in inverse computational imaging, where the blur is characterized by a kernel, known as the point spread function (PSF). For a complicated PSF, the resulting blurred image is usually incomprehensible to the human eye, such as in the case of lensless images. In this paper, we design and compare the reconstruction performance of under-parameterized and over-parameterized networks for the inverse imaging problem of lensless image reconstruction. We perform an untrained iterative reconstruction against a physics-informed loss function that requires knowledge of the forward imaging process. We conduct an extensive performance evaluation using multiple evaluation metrics and explore the obtained results contrastively by identifying the strengths and weaknesses of under-parameterization. Also, we provide reconstructions obtained using our custom PSF captured with a random diffuser.
Abeer Banerjee, Sumeet Saurav, Sanjay Singh 0001
ICIP3
2023 Capsule networks for computer vision applications: a comprehensive review
Seema Choudhary, Sumeet Saurav, Ravi Saini, Sanjay Singh 0001
Appl. Intell.4
2023 STemGAN: spatio-temporal generative adversarial network for video anomaly detection
Rituraj Singh, Krishanu Saini, Anikeit Sethi, Aruna Tiwari, Sumeet Saurav, Sanjay Singh 0001
Appl. Intell.6
2023 A dual-channel ensembled deep convolutional neural network for facial expression recognition in the wild
abstract
Abstract Facial expression recognition (FER) in the wild is an active and challenging field of research. A system for automatic FER finds use in a wide range of applications related to advanced human–computer interaction (HCI), human–robot interaction (HRI), human behavioral analysis, gaming and entertainment, etc. Since their inception, convolutional neural networks (CNNs) have attained state‐of‐the‐art accuracy in the facial analysis task. However, recognizing facial expressions in the wild with high confidence running on a low‐cost embedded device remains challenging. To this end, this study presents an efficient dual‐channel ensembled deep CNN (DCE‐DCNN) for FER in the wild. Initially, two DCNNs, namely the and , are trained separately on the grayscale and Scharr‐convolved vertical gradient facial images, respectively. The proposed network later integrates the two pre‐trained DCNNs to obtain the dual‐channel integrated DCNN (DCI‐DCNN). Finally, all three neural networks, namely the , , and DCI‐DCNN, are jointly fine‐tuned to get a single dual‐channel‐multi‐output model. The multi‐output model produces three prediction scores for the given input facial image. The prediction scores are thus fused using the max‐voting ensemble scheme to obtain the DCE‐DCNN with the final classification label. On the FER2013, RAF‐DB, NCAER‐S, AffectNet, and CKPlus benchmark FER datasets, the proposed DCE‐DCNN consistently outperforms the two individual DCNNs and numerous state‐of‐the‐art CNNs. Moreover, the network achieves competitive recognition accuracy on all four FER in the wild datasets with reduced memory storage size and parameters. The proposed DCE‐DCNN model with high throughput on resource‐limited embedded devices is suitable for applications that seek real‐time classification of facial expressions in the wild with high confidence.
Sumeet Saurav, Ravi Saini, Sanjay Singh 0001
Comput. Intell.3
2023 JS-SpoofNet: A jointly supervised parallel branched neural network for spoof detection
Shyam Sunder Prasad, Naval Kishore Mehta, Abeer Banerjee, Sumeet Saurav, Sanjay Singh 0001
Neurocomputing5
2023 An attention-guided convolutional neural network for automated classification of brain tumor from MRI
Sumeet Saurav, Ayush Sharma, Ravi Saini, Sanjay Singh 0001
Neural Comput. Appl.4
2023 Fast facial expression recognition using Boosted Histogram of Oriented Gradient (BHOG) features
Sumeet Saurav, Ravi Saini, Sanjay Singh 0001
Pattern Anal. Appl.3
2022 Three-dimensional DenseNet self-attention neural network for automatic detection of student's engagement
Naval Kishore Mehta, Shyam Sunder Prasad, Sumeet Saurav, Ravi Saini, Sanjay Singh 0001
Appl. Intell.5
2022 Vision-based techniques for fall detection in 360∘ videos using deep learning: Dataset and baseline results
Sumeet Saurav, Ravi Saini, Sanjay Singh 0001
Multim. Tools Appl.3
2022 Deep learning inspired intelligent embedded system for haptic rendering of facial emotions to the blind
Sumeet Saurav, Anil Kumar Saini, Ravi Saini, Sanjay Singh 0001
Neural Comput. Appl.4
2022 A dual-stream fused neural network for fall detection in multi-camera and $360^{\circ }$ videos
Sumeet Saurav, Ravi Saini, Sanjay Singh 0001
Neural Comput. Appl.3
2022 Dual integrated convolutional neural network for real-time facial expression recognition in the wild
Sumeet Saurav, Prashant Sadashiv Gidde, Ravi Saini, Sanjay Singh 0001
Vis. Comput.4
2021 Video Classification using SlowFast Network via Fuzzy rule
abstract
Anomalous events occur rarely and are challenging to model. Therefore, automatic recognition of abnormal activities in surveillance videos is a non-trivial task. Though with the availability of video datasets of abnormal activities, there has been some progress, recognition of abnormal activities in real-time with high confidence remains unsolved. Existing video-based anomaly detection techniques using traditional machine learning and deep-learning are compute-intensive and give low recognition accuracy. This paper presents a robust and computationally efficient deep learning-based framework to recognize different real-world anomalies from the video. The proposed scheme uses a Fuzzy rule to summarize the video to scale the problem into fewer frames and the slow-fast neural network for classification. Intuitively, the designed pipeline aims to solve two significant problems that arise with video classification; one is to reduce the redundant frames and avoid the computation of optical flow for a video that has a substantial computational requirement. The proposed scheme tested on the UCF-crime dataset and has achieved recognition accuracy of 53%.
Rituraj, Aruna Tiwari, Santanu Chaudhury, Sanjay Singh 0001, Sumeet Saurav
FUZZ-IEEE4
2021 EmNet: a deep integrated convolutional neural network for facial emotion recognition in the wild
Sumeet Saurav, Ravi Saini, Sanjay Singh 0001
Appl. Intell.3
2021 Three-dimensional CNN-inspired deep learning architecture for Yoga pose recognition in the real-world environment
Shrajal Jain, Aditya Rustagi, Sumeet Saurav, Ravi Saini, Sanjay Singh 0001
Neural Comput. Appl.5
2019 Facial Expression Recognition Using Histogram of Oriented Gradients with SVM-RFE Selected Features
Sumeet Saurav, Sanjay Singh 0001, Ravi Saini
HIS2
2018 An FPGA Based Hardware Accelerator for Classification of Handwritten Digits
R. Gautham Sundar Ram, Nitin Chaturvedi, Sumeet Saurav, Sanjay Singh 0001
ISDA (1)4