VLDB 2026 Research / reviewers in the wild / expert
Santosh Kumar Yadav
dblp:92/6302
· DBLP profile ↗
17ranked-venue papers
16as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 13 first-author · 11 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scalable energy optimization of resources for mobile cloud computing using sensor enabled cluster based system
Santosh Kumar Yadav |
Wirel. Networks | 1 |
| 2024 | ASME-SKYR framework: a comprehensive task scheduling framework for mobile cloud computing
Santosh Kumar Yadav |
Wirel. Networks | 1 |
| 2023 | SLBERT: A Novel Pre-Training Framework for Joint Speech and Language ModelingabstractWe propose SLBERT (Speech and Language pre-training framework for BERT), an end-to-end trainable framework for learning joint representations of speech and language modalities. We enhance the well-known BERT architecture to provide a dual-stream multimodal architecture that processes both speech and language input. To enable effective information exchange between the two modalities, we introduce a novel attention fusion mechanism via AF-Blocks. To acquire robust contrastive representations for speech and language processing applications, we pre-train SLBERT on three auxiliary tasks: Masked Language Modeling, Masked Speech Modeling, and Speech-Language Matching. We evaluate our proposed model on two well-known multimodal tasks: intent classification and sentiment analysis. Our model achieves state-of-the-art results on both benchmarks while surpassing even larger baselines. Onkar Susladkar, Prajwal Gatti, Santosh Kumar Yadav |
ICASSP | 3 |
| 2023 | SWTA: Sparse Weighted Temporal Attention for Drone-Based Activity RecognitionabstractDrone-camera based human activity recognition (HAR) has received significant attention from the computer vision research community in the past few years. A robust and efficient HAR system has a pivotal role in fields like video surveillance, crowd behavior analysis, sports analysis, and human-computer interaction. What makes it challenging are the complex poses, understanding different viewpoints, and the environmental scenarios where the action is taking place. To address such complexities, in this paper, we propose a novel Sparse Weighted Temporal Attention (SWTA) module to utilize sparsely sampled video frames for obtaining global weighted temporal attention. The proposed SWTA is divided into two components. First, temporal segment network that sparsely samples a given set of frames. Second, weighted temporal attention, which incorporates a fusion of attention maps derived from optical flow, with raw RGB images. This is followed by a basenet network, which comprises a convolutional neural network (CNN) module along with fully connected layers that provide us with activity recognition. The SWTA network can be used as a plug-in module to the existing deep CNN architectures, for optimizing them to learn temporal information by eliminating the need for a separate temporal stream. It has been evaluated on three publicly available benchmark datasets, namely Okutama, MOD20, and Drone-Action. The proposed model has received an accuracy of 72.76%, 92.56%, and 78.86% on the respective datasets thereby surpassing the previous state-of-the-art performances by a margin of 25.26%, 18.56%, and 2.94%, respectively. Santosh Kumar Yadav, Esha Pahwa, Achleshwar Luthra, Kamlesh Tiwari, Hari Mohan Pandey |
IJCNN | 1 |
| 2023 | DroneAttention: Sparse weighted temporal attention for drone-camera based activity recognition
Santosh Kumar Yadav, Achleshwar Luthra, Esha Pahwa, Kamlesh Tiwari, Heena Rathore, Hari Mohan Pandey, Peter Corcoran 0001 |
Neural Networks | 1 |
| 2022 | WTM: Weighted Temporal Attention Module for Group Activity RecognitionabstractGroup Activity Recognition requires spatiotemporal modeling of an exponential number of semantic and geometric relations among various individuals in a scene. Previous attempts model these relations by aggregating independently derived spatial and temporal features. This increases the modeling complexity and results in sparse information due to lack of feature correlation. In this paper, we propose Weighted Temporal Attention Mechanism (WTM), a representational mechanism that combines spatial and temporal features of a local subset of a visual sequence into a single 2D image representation, highlighting areas of a frame where actor motion is significant. Pairwise dense optical flow maps representing the temporal characteristic of individuals over a sequence are used as attention masks over raw RGB images through a multi-layer weighted aggregation. We demonstrate a strong correlation between spatial and temporal features, which helps localize actions effectively in a multi-person scenario. The simplicity of the input representation allows the model to be trained by 2D image classification architectures in a plug-and-play fashion, which outperforms its multi-stream and multi-dimensional counterparts. The proposed method achieves the lowest computational complexity in comparison to other works. We demonstrate the performance of WTM on two widely used public benchmark datasets, namely the Collective Activity Dataset (CAD) and the Volleyball Dataset. and achieve state-of-the-art accuracies of 95.1% and 94.6% respectively. We also discuss the application of this method to other datasets and general scenarios. The code is being made publicly available. Santosh Kumar Yadav, Palaash Agrawal, Kamlesh Tiwari, Ehsan Adeli-Mosabbeb, Hari Mohan Pandey, Ali Akbar Shaikh |
IJCNN | 1 |
| 2022 | MS-KARD: A Benchmark for Multimodal Karate Action RecognitionabstractClassifying complex human motion sequences is a major research challenge in the domain of human activity recognition. Currently, most popular datasets lack a specialized set of classes pertaining to similar action sequences (in terms of spatial trajectories). To recognize such complex action sequences with high inter-class similarity, such as those in karate, multiple streams are required. To fulfill this need, we propose MS-KARD, a Multi-Stream Karate Action Recognition Dataset that uses multiple vision perspectives, as well as sensor data - accelerometer and gyroscope. It includes 1518 video clips along with their corresponding sensor data. Each video was shot at 30fps and lasts around one minute, equating to a total of 2,814,930 frames and 5,623,734 sensor data samples. The dataset has been collected for 23 classes like Jodan Zuki, Oi Zuki, etc. The data acquisition setting involves the combination of 2 orthogonal web cameras and 3 wearable inertial sensors recording both vision and inertial data respectively. The aim of this dataset is to aid research that deals with recognizing human actions that have similar spatial trajectories. The paper describes statistics of the dataset, acquisition setting, and provides baseline performance figures using popular action recognizers. We propose an ensemble-based method, KarateNet, that performs decision-level fusion on the two input modalities (vision and sensor data) to classify actions. For the first stream, the RGB frames are extracted from the videos and passed into action recognition networks like Temporal Segment Network (TSN) and Temporal Shift Module (TSM). For the second stream, the sensor data is converted into a 2-D image and fed into a Convolutional Neural Network (CNN). The results reported were obtained on performing a fusion of the 2 streams. We also report results on ablations that use fusion with various input settings. The dataset and code will be made publicly available. Santosh Kumar Yadav, Aditya Deshmukh, Raghurama Varma Gonela, Shreyas Bhat Kera, Kamlesh Tiwari, Hari Mohan Pandey, Ali Akbar Shaikh |
IJCNN | 1 |
| 2022 | TBAC: Transformers Based Attention Consensus for Human Activity RecognitionabstractHuman Activity Recognition is an important task in Computer Vision that involves the utilization of spatio-temporal features of videos to classify human actions. The temporal portion of videos contains vital information needed for accurate classification. However, common Deep Learning methods simply average the temporal features, thereby giving all frames equal importance irrespective of their relevance, which negatively impacts the accuracy of the model. To combat this adverse effect, this paper proposes a novel Transformer Based Attention Consensus (TBAC) module. The TBAC module can be used in a plug-and-play manner as an alternate to the conventional consensus meth-ods of any existing video action recognition network. The TBAC module contains four components: (i) Query Sampling Unit, (ii) Attention Extraction Unit, (iii) Softening Unit, and (iv) Attention Consensus Unit. Our experiments demonstrate that the use of the TBAC module in place of classical consensus can improve the performance of the CNN-based action recognition models, such as Channel Separated Convolutional Network (CSN), Temporal Shift Module (TSM), and Temporal Segment Network (TSN). We also propose the Decision Consensus (DC) algorithm that utilizes multiple independent but related action recognizer models in order to improve upon the performance of most of these constituent models, using a novel fusion algorithm. Results have been obtained on two benchmark human action recognition datasets, HMDB51 and HAA500. The use of the proposed TBAC module along with Decision Consensus achieves state-of-the-art performances, with 85.23% and 83.73% classification accuracies on the two databases HMDB51 and HAA500, respectively. The code will be made publicly available. Santosh Kumar Yadav, Shreyas Bhat Kera, Raghurama Varma Gonela, Kamlesh Tiwari, Hari Mohan Pandey, Ali Akbar Shaikh |
IJCNN | 1 |
| 2022 | YogaTube: A Video Benchmark for Yoga Action RecognitionabstractYoga can be seen as a set of fitness exercises involving various body postures. Most of the available pose and action recognition datasets are comprised of easy-to-moderate body pose orientations and do not offer much challenge to the learning algorithms in terms of the complexity of pose. In order to observe action recognition from a different perspective, we introduce YogaTube, a new large-scale video benchmark dataset for yoga action recognition. YogaTube aims at covering a wide range of complex yoga postures, which consist of 5484 videos belonging to a taxonomy of 82 classes of yoga asanas. Also, a three-stream architecture has been designed for yoga asanas pose recognition using two modules, feature extraction, and classification. Feature extraction comprises three parallel components. First, pose is estimated using the part affinity fields model to extract meaningful cues from the practitioner. Second, optical flow is used to extract temporal features. Third, raw RGB videos are used for extracting the spatiotemporal features. Finally in the classification module, pose, optical flow, and RGB streams are fused to get the final results of the yoga asanas. To the best of our knowledge, this is the first attempt to establish a video benchmark yoga recognition dataset. The code and dataset will be released soon. Santosh Kumar Yadav, Guntaas Singh, Manisha Verma, Kamlesh Tiwari, Hari Mohan Pandey, Ali Akbar Shaikh, Peter Corcoran 0001 |
IJCNN | 1 |
| 2022 | YogNet: A two-stream network for realtime multiperson yoga action recognition and posture correction
Santosh Kumar Yadav, Aayush Agarwal, Kamlesh Tiwari, Hari Mohan Pandey, Ali Akbar Shaikh |
Knowl. Based Syst. | 1 |
| 2022 | ARFDNet: An efficient activity recognition & fall detection system using latent feature pooling
Santosh Kumar Yadav, Achleshwar Luthra, Kamlesh Tiwari, Hari Mohan Pandey, Ali Akbar Shaikh |
Knowl. Based Syst. | 1 |
| 2022 | CSITime: Privacy-preserving human activity recognition using WiFi channel state information
Santosh Kumar Yadav, Siva Sai, Akshay Gundewar, Heena Rathore, Kamlesh Tiwari, Hari Mohan Pandey, Mohit Mathur |
Neural Networks | 1 |
| 2022 | Skeleton-based human activity recognition using ConvLSTM and guided feature learningabstractAbstract Human activity recognition aims to determine actions performed by a human in an image or video. Examples of human activity include standing, running, sitting, sleeping,etc. These activities may involve intricate motion patterns and undesired events such as falling. This paper proposes a novel deep convolutional long short-term memory (ConvLSTM) network for skeletal-based activity recognition and fall detection. The proposed ConvLSTM network is a sequential fusion of convolutional neural networks (CNNs), long short-term memory (LSTM) networks, and fully connected layers. The acquisition system applies human detection and pose estimation to pre-calculate skeleton coordinates from the image/video sequence. The ConvLSTM model uses the raw skeleton coordinates along with their characteristic geometrical and kinematic features to construct the novel guided features. The geometrical and kinematic features are built upon raw skeleton coordinates using relative joint position values, differences between joints, spherical joint angles between selected joints, and their angular velocities. The novel spatiotemporal-guided features are obtained using a trained multi-player CNN-LSTM combination. Classification head including fully connected layers is subsequently applied. The proposed model has been evaluated on the KinectHAR dataset having 130,000 samples with 81 attribute values, collected with the help of a Kinect (v2) sensor. Experimental results are compared against the performance of isolated CNNs and LSTM networks. Proposed ConvLSTM have achieved an accuracy of 98.89% that is better than CNNs and LSTMs having an accuracy of 93.89 and 92.75%, respectively. The proposed system has been tested in realtime and is found to be independent of the pose, facing of the camera, individuals, clothing,etc. The code and dataset will be made publicly available. Santosh Kumar Yadav, Kamlesh Tiwari, Hari Mohan Pandey, Ali Akbar Shaikh |
Soft Comput. | 1 |
| 2021 | A review of multimodal human activity recognition with special emphasis on classification, applications, challenges and future directions
Santosh Kumar Yadav, Kamlesh Tiwari, Hari Mohan Pandey, Ali Akbar Shaikh |
Knowl. Based Syst. | 1 |
| 2019 | Real-time Yoga recognition using deep learning
Santosh Kumar Yadav, Amitojdeep Singh, Jagdish Lal Raheja |
Neural Comput. Appl. | 1 |
| 2015 | Electrocardiogram signal denoising using non-local wavelet transform domain filteringabstractElectrocardiogram (ECG) signals are usually corrupted by baseline wander, power‐line interference, muscle noise etc. Numerous methods have been proposed to remove these noises. However, in case of wireless recording of the ECG signal it gets corrupted by the additive white Gaussian noise (AWGN). For the correct diagnosis, removal of AWGN from ECG signals becomes necessary as it affects the diagnostic features. The natural signals exhibit correlation among their samples and this property has been exploited in various signal restoration tasks. Motivated by that, in this study we propose a non‐local wavelet transform domain ECG signal denoising method which exploits the correlations among both local and non‐local samples of the signal. In the proposed method, the similar blocks of the samples are grouped in a matrix and then denoising is achieved by the shrinkage of its two‐dimensional discrete wavelet transform coefficients. The experiments performed on a number of ECG signals show significant quantitative and qualitative improvements in denoising performance over the existing ECG signal denoising methods. Santosh Kumar Yadav, Rohit Sinha 0003, Prabin Kumar Bora |
IET Signal Process. | 1 |
| 2015 | An Efficient SVD Shrinkage for Rank EstimationabstractMatrix rank estimation is a classical problem with many applications in statistical signal processing. In this letter, a logistic function based thresholding of the singular values is proposed for the rank estimation purpose. Parameters of the proposed shrinkage function are tuned using Stein's unbiased risk estimator. The proposed method is shown to outperform the state-of-the-art methods in terms of rank estimation accuracy. Further, it is also noted to result in a better denoising performance. Santosh Kumar Yadav, Rohit Sinha 0003, Prabin Kumar Bora |
IEEE Signal Process. Lett. | 1 |