Praveen Kumar 0005

dblp:95/448-5 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
12since 2021 · last 2025
0000-0003-4820-3088ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2025 An integrative survey on Indian sign language recognition and translation
abstract
Abstract Hard of hearing (HoH) people commonly use sign languages (SLs) to communicate. They face major impediments in communicating with hearing individuals, mostly because hearing people are unaware of SLs. Therefore, it is important to promote tools that enable communication between users of sign language and users of spoken languages. The study of sign language recognition and translation (SLRT) is a step forward in this direction, as it tries to create a spoken‐language translation of a sign‐language video or vice versa. This study aims to survey the Indian sign language (ISL) interpretation literature and gives pertinent information about ISL recognition and translation (ISLRT). It provides an overview of recent advances in ISLRT, including the use of machine learning based, deep learning based, and gesture‐based techniques. This work also summarizes the development of ISL datasets and dictionaries. It highlights the gaps in the literature and provides recommendations for future research opportunities for ISLRT development.
Rina Damdoo, Praveen Kumar 0005
IET Image Process.2
2025 SignEdgeLVM transformer model for enhanced sign language translation on edge devices
abstract
Transformer architectures have accelerated the research in Continuous Sign Language Recognition and Translation (CSLRT), which involves predicting sign gloss patterns from video and converting them into spoken language. This process is challenging due to the lack of direct alignment between sign glosses and spoken words. While Transformers are effective due to their ability to process inputs in parallel, their high memory consumption makes the transformers less suitable for edge devices. To address this issue, we propose the SignEdgeLVM, a model that uses a Global Relative Attention Matrix (GRAM) and Dynamic Point Frame Sampling (DPFS) module. The SignEdgeLVM significantly lowers attention mechanism memory consumption by 78.22 MB (99.93 percent) per head in the attention layer for our implementation. Evaluated on the PHOENIX14T dataset, this optimization makes SignEdgeLVM suitable for edge devices.
Rina Damdoo, Praveen Kumar 0005
Discov. Comput.2
2024 A visual intelligent system for students' behavior classification using body pose and facial features in a smart classroom
Chakradhar Pabba, Vishal Bhardwaj, Praveen Kumar 0005
Multim. Tools Appl.3
2024 A vision-based multi-cues approach for individual students' and overall class engagement monitoring in smart classroom environments
Chakradhar Pabba, Praveen Kumar 0005
Multim. Tools Appl.2
2023 A novel multi-modal depression detection approach based on mobile crowd sensing and task-based mechanisms
Ravi Prasad Thati, Abhishek Singh Dhadwal, Praveen Kumar 0005, Sainaba P
Multim. Tools Appl.3
2022 A survey on parallel computing for traditional computer vision
abstract
Summary The applications of computer vision (CV) are continuously increasing along with the enormous demand for real‐time data processing. This visual data processing is done with various compute‐intensive image/video processing algorithms that may belong to traditional approaches or deep learning approaches. This article aims to provide a survey of state‐of‐the‐art hardware platforms and software frameworks for parallel implementation of traditional CV applications. The article discusses various options for hardware platforms for centralized‐computing architecture and edge‐computing architecture, and various software frameworks that can be used to leverage the hardware. This discussion is based on a systematic survey of studies/works that show the use of various hardware platforms and software frameworks in order to achieve real‐time processing for CV algorithms. Based on the survey, some possible future directions are also discussed.
Deepak Jaiswal, Praveen Kumar 0005
Concurr. Comput. Pract. Exp.2
2022 A comparative study on SoC embedded low power GPUs for real-time edge-based automated traffic surveillance
abstract
Summary Achieving real‐time processing for automated traffic surveillance is a major challenge due to the huge amount of data generated by a large number of surveillance cameras. For this, centralized computing has been a standard architecture for many years. But due to the increasing need for real‐time processing and limitations of centralized computing architecture (such as network congestion due to limited bandwidth and the need for costly high‐end servers), the paradigm is shifting toward edge processing. This article compares the suitability of two popular but different types of SoCs (System on Chip) (from NVIDIA and Qualcomm) embedded with low‐power GPUs as processing units for edge devices. These GPUs can be programmed as general‐purpose GPUs (GPGPU) to speed up compute‐intensive tasks, so as to achieve real‐time processing at the edge of the network. The article also discusses the architectural features and differences between these GPUs and recommends optimization techniques to leverage them. For quantitative comparison, we implement a wrong‐way vehicle detection algorithm (using background‐subtraction‐based moving object detection and Kalman‐filter‐based trajectory tracking), and optimized it for these SoCs. The experimental results show that real‐time processing can be achieved for HD videos with implementations optimized for these SoCs.
Deepak Jaiswal, Praveen Kumar 0005
Concurr. Comput. Pract. Exp.2
2022 An intelligent system for monitoring students' engagement in large classroom teaching through facial expression recognition
abstract
Abstract Students' disengagement problem has become critical in the modern scenario due to various distractions and lack of student‐teacher interactions. This problem is exacerbated with large offline classrooms, where it becomes challenging for teachers to monitor students' engagement and maintain the right‐level of interactions. Traditional ways of monitoring students' engagement rely on self‐reporting or using physical devices, which have limitations for offline classroom use. Student's academic affective states (e.g., moods and emotions) analysis has potential for creating intelligent classrooms, which can autonomously monitor and analyse students' engagement and behaviours in real‐time. In recent literature, a few computer vision based methods have been proposed, but they either work only in the e‐learning domain or have limitations in real‐time processing and scalability for large offline classes. This paper presents a real‐time system for student group engagement monitoring by analysing their facial expressions and recognizing academic affective states: ‘boredom,’ ‘confuse,’ ‘focus,’ ‘frustrated,’ ‘yawning,’ and ‘sleepy,’ which are pertinent in the learning environment. The methodology includes certain pre‐processing steps like face detection, a convolutional neural network (CNN) based facial expression recognition model, and post‐processing steps like frame‐wise group engagement estimation. For training the CNN model, we created a dataset of the aforementioned facial expressions from classroom lecture videos and added related samples from three publicly available datasets, BAUM‐1, DAiSEE, and YawDD, to generalize the model predictions. The trained model has achieved train and test accuracy of 78.70% and 76.90%, respectively. The proposed methodology gave promising results when compared with self‐reported engagement levels by students.
Chakradhar Pabba, Praveen Kumar 0005
Expert Syst. J. Knowl. Eng.2
2022 Deep learning and RGB-D based human action, human-human and human-object interaction recognition: A survey
Pushpajit Khaire, Praveen Kumar 0005
J. Vis. Commun. Image Represent.2
2022 Fall detection approach based on combined displacement of spatial features for intelligent indoor surveillance
Anurag De, Ashim Saha, Praveen Kumar 0005
Multim. Tools Appl.3
2022 Fall detection method based on Spatio-temporal feature fusion using combined two-channel classification
Anurag De, Ashim Saha, Praveen Kumar 0005, Gautam Pal
Multim. Tools Appl.3
2022 Online suspicious event detection in a constrained environment with RGB+D camera using multi-stream CNNs and SVM
Pushpajit Khaire, Praveen Kumar 0005
Multim. Tools Appl.2
2018 User specific context construction for personalized multimedia retrieval
Shirish Singh, Praveen Kumar 0005
Multim. Tools Appl.2
2018 Combining CNN streams of RGB-D and skeletal data for human activity recognition
Pushpajit Khaire, Praveen Kumar 0005, Javed Imran
Pattern Recognit. Lett.2
2012 Efficient compression and network adaptive video coding for distributed video surveillance
Praveen Kumar 0005, Amit Pande, Ankush Mittal
Multim. Tools Appl.1
2012 OS-Guard: on-site signature based framework for multimedia surveillance data management
Praveen Kumar 0005, Sujoy Roy, Ankush Mittal
Multim. Tools Appl.1
2011 Parallel flux tensor analysis for efficient moving object detection
Kannappan Palaniappan, Ilker Ersoy, Guna Seetharaman, Shelby R. Davis, Praveen Kumar 0005, Raghuveer M. Rao, Richard W. Linderman
FUSION5
2010 Efficient feature extraction and likelihood fusion for vehicle tracking in low frame rate airborne video
Kannappan Palaniappan, Filiz Bunyak, Praveen Kumar 0005, Ilker Ersoy, Stefan Jäger 0001, Koyeli Ganguli, Anoop Haridas, Joshua Fraser, Raghuveer M. Rao, Guna Seetharaman
FUSION3
2009 Parallel Blob Extraction Using the Multi-core Cell Processor
Praveen Kumar 0005, Kannappan Palaniappan, Ankush Mittal, Guna Seetharaman
ACIVS1