Dinesh Singh 0001

dblp:127/1101-1 · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
9since 2021 · last 2026
0000-0001-8889-9847ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 7 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 SPOT-Face: Forensic Face Identification Using Attention Guided Optimal Transport
Ravi Shankar Prasad, Dinesh Singh 0001
ICPR (9)2
2026 Cross-domain identity representation for skull-to-face matching with benchmark dataset
Ravi Shankar Prasad, Dinesh Singh 0001
J. Vis. Commun. Image Represent.2
2025 Towards Scene Text Recognition in Rainy Weather Conditions
Anandita Jamwal, Lalithya Koneti, Manikandan Ravikiran, Dinesh Singh 0001, Rohit Saluja
ICDAR (5)4
2025 Drone-UP: Drone-based Unauthorized Parking Detection Method for Edge-devices
abstract
In this paper, we present Drone-UP, a framework for the automatic detection of unauthorized parked vehicles using drone-based surveillance videos in real-time for the resource-constrained on-board edge devices. The proposed framework consists of two stages, namely, unauthorized region identification and vehicle speed estimation. The goal of the first stage is to identify the potential unauthorized region in the live video stream from the drone camera using the YOLO-seg-based method fine-tuned for the segmentation of roads with side walks. In the second stage, it tracks the vehicles in the unauthorized area and estimates their speed to determine the violators. To meet real-time performance on the edge devices, we used post-training quantization and low bit representation techniques to speed up the inference of the segmentation and object detection models. To evaluate our method, we collected a benchmark dataset of drone video with varying traffic conditions, height, and angles, captured with different cameras. Our experimental results demonstrate that the integration of our optimized segmentation, detection, and speed estimation modules achieves robust performance in identifying unauthorized parking violations while maintaining real-time efficiency on edge devices.
Tarun Saini, Madhumita Guha, Sachin Kumar Bairwa, Vandita Dutt, Kajal Singh, Dinesh Singh 0001
IJCNN6
2025 Para-X: Graph-based Facial Paralysis Detection using Structural Deformations of Facial Expression
abstract
Facial paralysis, marked by the inability to move specific facial muscles, often results from nerve damage, strokes, or neurological disorders. Prompt and accurate detection is essential for effective diagnosis and treatment, potentially improving recovery outcomes. Recent studies identified facial paralysis by analyzing alterations in facial expressions of affected individuals relative to those of unaffected individuals, utilizing facial attributes and landmark data; nevertheless, they overlooked the structural information among the diverse face features. This work utilizes structural information to offer a graphical depiction of facial features. In the graph representation of facial attributes, key-points serve as the vertices. At the same time, the edges are established based on the closeness of the key-points and the similarity of the local appearance of the facial attributes conveyed via the vision transformer. Advanced graph convolutional networks integrate structural information into face attribute encoding to enhance the identification of facial expressions. Consequently, Para-X acquires highly expressive semantic representations from facial attribute graphs. In contrast, the vision transformer and graph convolutional blocks enable the framework to leverage local and global dependencies among facial attributes, which are crucial for recognizing facial paralysis. Comprehensive studies demonstrate the robustness and generalizability of the proposed methodology to identify facial paralysis in various facial paralysis datasets such as AFLFP, YFPD, FDPDI, and our FPD dataset.
Nandani Sharma, Kajal Singh, Dinesh Singh 0001
IJCNN3
2024 Tackling Data Heterogeneity in Federated Learning through Global Density Estimation
abstract
Federated learning is gaining popularity for its largely accepted paradigm of privacy protection. However, data heterogeneity among clients is often the primary challenge in federated learning, hindering the convergence of deep neural networks. The non-IID nature of data across clients escalates the computational cost and communication overhead for models trained locally on-device and shared for global averaging. To mitigate this issue, we try to preserve the statistical parameters of local clients by estimating the global density using Gaussian mixture model. Our study primarily focuses on preserving client data privacy while addressing the statistical heterogeneity in data distribution across all the clients. A federated implementation of the distribution-preserving sampling algorithm, FedDpS is put forward to mitigate the high heterogeneity of data among clients thereby facilitating the training of DNN models for faster convergence. Local models are built at the client level using our on-device algorithm designed to tackle data heterogeneity among clients. Significant improvements in test accuracy, F1-score, and other evaluation metrics have been observed when trained using state-of-the-art optimization models like FedAvg, FedProx, FedAdam, FedAwS, and MOON. Our proposed method achieves the target performance in fewer communication rounds thereby reducing the overall communication cost. The code for our implementation can be found at https://github.com/sagnik04g/FedDpS.
Sagnik Ghosh, Avinash Kushwaha, Dinesh Singh 0001
IEEE Big Data3
2023 Graph Representation for Weakly-Supervised Spatio-Temporal Action Detection
abstract
Spatio-temporal action recognition and localization are crucial in several computer vision applications including video surveillance, video captioning to name a few. However, most of the existing action recognition and localization approaches are for offline use, perform well only on trimmed action clips. Also, they need precise annotations at the clip, frame, and pixel levels which is labor-intensive and thus undermines their usage for real-world large-scale scenarios. In this paper, we propose a weakly-supervised spatio-temporal action recognition and localization based on graph representation in untrimmed videos. More specifically, we propose an efficient graph representation of videos using only the clip level annotations, while existing approaches are either supervised or unsupervised learning approach. For graph construction, the local actions are determined based on the key interesting demeanor in an action clip and assigned the class label the same as that of the clip. This weak annotation impacts both action recognition and localization significantly because the local actions have considerable intra-class variability and inter-class similarity. To handle the intra-class variability and inter-class similarity, we use a weakly-supervised deep multiple instance ranking framework on the local action descriptors. To classify a graph of local actions into one of the action classes, we use a support vector machine along with a graph kernel and then localize the recognized action as a non-cubic shaped-portion of the video based on local actions in the graph. The experimental results show that the proposed approach outperforms the state-of-the-art methods on the three benchmark datasets, namely, THUMOS14, UCF-Sports, and JHMDB-21.
Dinesh Singh 0001
IJCNN1
2023 FsNet: Feature Selection Network on High-dimensional Biological Data
abstract
Biological data, including gene expression data, are generally high-dimensional and require efficient, generalizable, and scalable machine-learning methods to discover complex nonlinear patterns. Recent advances in machine learning can be attributed to deep neural networks (DNNs), which perform various tasks in terms of computer vision and natural language processing. However, standard DNNs are inappropriate for high-dimensional datasets generated in biology because they consider numerous parameters, which in turn require numerous samples. In this paper, we propose a DNN-based, nonlinear feature selection method, called the feature selection network (FsNet), for high-dimensional and small sample data. Specifically, FsNet comprises a selection layer that selects features and a reconstruction layer that stabilizes the training. Because a large number of parameters in the selection and reconstruction layers can easily result in overfitting under a limited number of samples, we utilized two tiny networks to predict the large virtual weight matrices of the selection and reconstruction layers. Experimental results on several real-world high-dimensional biological datasets demonstrate the efficacy of the proposed method.
Dinesh Singh 0001, Héctor Climente-González, Mathis Petrovich, Eiryo Kawakami, Makoto Yamada
IJCNN1
2023 GraphLIME: Local Interpretable Model Explanations for Graph Neural Networks
abstract
Graph structured data has wide applicability in various domains such as physics, chemistry, biology, computer vision, and social networks, to name a few. Recently, graph neural networks (GNN) were shown to be successful in effectively representing graph structured data because of their good performance and generalization ability. However, explaining the effectiveness of GNN models is a challenging task because of the complex nonlinear transformations made over the iterations. In this paper, we propose GraphLIME, a local interpretable model explanation for graphs using the Hilbert-Schmidt Independence Criterion (HSIC) Lasso, which is a nonlinear feature selection method. GraphLIME is a generic GNN-model explanation framework that learns a nonlinear interpretable model locally in the subgraph of the node being explained. Through experiments on two real-world datasets, the explanations of GraphLIME are found to be of extraordinary degree and more descriptive in comparison to the existing explanation methods.
Makoto Yamada, Yuan Tian 0016, Dinesh Singh 0001, Yi Chang 0001
IEEE Trans. Knowl. Data Eng.4
2019 Deep Spatio-Temporal Representation for Detection of Road Accidents Using Stacked Autoencoder
abstract
Vision-based detection of road accidents using traffic surveillance video is a highly desirable but challenging task. In this paper, we propose a novel framework for automatic detection of road accidents in surveillance videos. The proposed framework automatically learns feature representation from the spatiotemporal volumes of raw pixel intensity instead of traditional hand-crafted features. We consider the accident of the vehicles as an unusual incident. The proposed framework extracts deep representation using denoising autoencoders trained over the normal traffic videos. The possibility of an accident is determined based on the reconstruction error and the likelihood of the deep representation. For the likelihood of the deep representation, an unsupervised model is trained using one class support vector machine. Also, the intersection points of the vehicle's trajectories are used to reduce the false alarm rate and increase the reliability of the overall system. We evaluated out proposed approach on real accident videos collected from the CCTV surveillance network of Hyderabad City in India. The experiments on these real accident videos demonstrate the efficacy of the proposed approach.
Dinesh Singh 0001, C. Krishna Mohan
IEEE Trans. Intell. Transp. Syst.1
2018 Projection-SVM: Distributed Kernel Support Vector Machine for Big Data using Subspace Partitioning
abstract
The training of kernel support vector machine (SVM) is a computationally complex task for large datasets where the number of samples ranges in millions. This is because kernel matrix (in general not sparse) is both computation expensive and memory intensive. Existing methods hardly achieve a linear scale and suffer from high approximation loss. We propose Projection-SVM, a distributed implementation of kernel support vector machine for large datasets using subspace partitioning. In subspace partitioning, a decision tree is constructed on projection of data along the direction of maximum variance (i.e., dominant eigenvector) to obtain smaller partitions (i.e., subspaces) of the dataset. On each of these partitions, a kernel SVM is trained independently over a cluster thereby reducing the overall training time. Also, it results in reducing the prediction time significantly. We demonstrate the efficacy of the proposed approach on eight standard large datasets from various application domains, namely, mnist8m, kddcup99, webspam, etc. where Projection-SVM is on an average 150 times faster than sequential SVM while maintaining the classification accuracy. The experimental results also show the superiority of the Projection-SVM over the state-of-the-art approaches for distributed kernel SVMs, such as DCSVM, CASVM, and DTSVM.
Dinesh Singh 0001, C. Krishna Mohan
IEEE BigData1
2018 Fast-BoW: Scaling Bag-of-Visual-Words Generation
Dinesh Singh 0001, Abhijeet Bhure, Sumit Mamtani, C. Krishna Mohan
BMVC1
2017 Detection of motorcyclists without helmet in videos using convolutional neural network
abstract
In order to ensure the safety measures, the detection of traffic rule violators is a highly desirable but challenging task due to various difficulties such as occlusion, illumination, poor quality of surveillance video, varying whether conditions, etc. In this paper, we present a framework for automatic detection of motorcyclists driving without helmets in surveillance videos. In the proposed approach, first we use adaptive background subtraction on video frames to get moving objects. Later convolutional neural network (CNN) is used to select motorcyclists among the moving objects. Again, we apply CNN on upper one fourth part for further recognition of motorcyclists driving without a helmet. The performance of the proposed approach is evaluated on two datasets, IITH_Helmet_1 contains sparse traffic and IITH_Helmet_2 contains dense traffic, respectively. The experiments on real videos successfully detect 92.87% violators with a low false alarm rate of 0.5% on an average and thus shows the efficacy of the proposed approach.
Chalavadi Vishnu, Dinesh Singh 0001, C. Krishna Mohan, Sobhan Babu Chintapalli
IJCNN2
2017 Graph formulation of video activities for abnormal activity recognition
Dinesh Singh 0001, C. Krishna Mohan
Pattern Recognit.1
2017 DiP-SVM : Distribution Preserving Kernel Support Vector Machine for Big Data
abstract
In literature, the task of learning a support vector machine for large datasets has been performed by splitting the dataset into manageable sized “partitions” and training a sequential support vector machine on each of these partitions separately to obtain local support vectors. However, this process invariably leads to the loss in classification accuracy as global support vectors may not have been chosen as local support vectors in their respective partitions. We hypothesize that retaining the original distribution of the dataset in each of the partitions can help solve this issue. Hence, we present DiP-SVM, a distribution preserving kernel support vector machine where the first and second order statistics of the entire dataset are retained in each of the partitions. This helps in obtaining local decision boundaries which are in agreement with the global decision boundary, thereby reducing the chance of missing important global support vectors. We show that DiP-SVM achieves a minimal loss in classification accuracy among other distributed support vector machine techniques on several benchmark datasets. We further demonstrate that our approach reduces communication overhead between partitions leading to faster execution on large datasets and making it suitable for implementation in cloud environments.
Dinesh Singh 0001, Debaditya Roy, C. Krishna Mohan
IEEE Trans. Big Data1
2016 Distributed quadratic programming solver for kernel SVM using genetic algorithm
abstract
Support vector machine (SVM) is a powerful tool for classification and regression problems, however, its time and space complexities make it unsuitable for large datasets. In this paper, we present GeneticSVM, an evolutionary computing based distributed approach to find optimal solution of quadratic programming (QP) for kernel support vector machine. In Ge-neticSVM, novel encoding method and crossover operation help in obtaining the better solution. In order to train a SVM from large datasets, we distribute the training task over the graphics processing units (GPUs) enabled cluster. It leverages the benefit of the GPUs for large matrix multiplication. The experiments show better performance in terms of classification accuracy as well as computational time on standard datasets like GISETTE, ADULT, etc.
Dinesh Singh 0001, C. Krishna Mohan
CEC1
2016 Spontaneous Facial Expression Recognition: A Part Based Approach
abstract
A part-based approach for spontaneous expression recognition using audio-visual feature and deep convolution neural network (DCNN) is proposed. The ability of convolution neural network to handle variations in translation and scale is exploited for extracting visual features. The sub-regions, namely, eye and mouth parts extracted from the video faces are given as an input to the deep CNN (DCNN) inorder to extract convnet features. The audio features, namely, voice-report, voice intensity, and other prosodic features are used to obtain complementary information useful for classification. The confidence scores of the classifier trained on different facial parts and audio information are combined using different fusion rules for recognizing expressions. The effectiveness of the proposed approach is demonstrated on acted facial expression in wild (AFEW) dataset.
Nazil Perveen, Dinesh Singh 0001, C. Krishna Mohan
ICMLA2
2016 Visual Big Data Analytics for Traffic Monitoring in Smart City
abstract
The application such as video surveillance for traffic control in smart cities needs to analyze the large amount (hours/days) of video footage in order to locate the people who are violating the traffic rules. The traditional computer vision techniques are unable to analyze such a huge amount of visual data generated in real-time. So, there is a need for visual big data analytics which involves processing and analyzing large scale visual data such as images or videos to find semantic patterns that are useful for interpretation. In this paper, we propose a framework for visual big data analytics for automatic detection of bike-riders without helmet in city traffic. We also discuss challenges involved in visual big data analytics for traffic control in a city scale surveillance data and explore opportunities for future research.
Dinesh Singh 0001, Chalavadi Vishnu, C. Krishna Mohan
ICMLA1
2016 Automatic detection of bike-riders without helmet using surveillance videos in real-time
abstract
In this paper, we propose an approach for automatic detection of bike-riders without helmet using surveillance videos in real time. The proposed approach first detects bike riders from surveillance video using background subtraction and object segmentation. Then it determines whether bike-rider is using a helmet or not using visual features and binary classifier. Also, we present a consolidation approach for violation reporting which helps in improving reliability of the proposed approach. In order to evaluate our approach, we have provided a performance comparison of three widely used feature representations namely histogram of oriented gradients (HOG), scale-invariant feature transform (SIFT), and local binary patterns (LBP) for classification. The experimental results show detection accuracy of 93.80% on the real world surveillance data. It has also been shown that proposed approach is computationally less expensive and performs in real-time with a processing time of 11.58 ms per frame.
Kunal Dahiya, Dinesh Singh 0001, C. Krishna Mohan
IJCNN2