VLDB 2026 Research / reviewers in the wild / expert
Debi Prosad Dogra
dblp:97/7630
· DBLP profile ↗
65ranked-venue papers
2as first author
24since 2021 · last 2026
0000-0002-3904-732XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 2 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Illumination-Invariant Representations for Vehicle Re-identification
Arabinda Panda, Debi Prosad Dogra, Partha Pratim Dey |
ICPR (12) | 2 |
| 2026 | VAST-ReID: A Low-Light Benchmark Dataset for Person Re-Identification with Visual and Attribute-Rich Semantic TrackingabstractPerson Re-Identification (ReID) task is important for designing intelligent surveillance systems. ReID can be highly challenging in low-light and low resolution scenarios. Existing ReID datasets predominantly feature cropped pedestrian images captured in well-lit environments, often lacking semantic richness, frame-level temporal continuity, and robustness to adverse conditions. To address these limitations, we introduce VAST-ReID, a new benchmark dataset specifically designed for the low-light person ReID task in real-world surveillance contexts. VAST-ReID consists of 1,441 surveillance videos collected at 24 different locations, capturing 256 distinct pedestrians of various age groups. The dataset emphasizes naturally low-light and visually degraded scenarios. Each identity is annotated with dense bounding boxes and enriched with auxiliary semantic labels, including pedestrian attributes and LLM-generated descriptions. While these annotations are not used during supervised training, they provide valuable semantic context for advancing research in language-guided retrieval and attribute-aware modeling. Additionally, we release identity-aligned image crops under the BoxTrack-ReID subset, which has over 18.7K frames sampled at 1fps from the raw videos, with standard training, gallery, and query splits compatible with the Market-1501 evaluation protocol, enabling straightforward benchmarking. The dataset has been benchmarked against SOTA methods, and experiments reveal that there is huge scope for improvement in ReID research. VAST-ReID is available at: https://github.com/Byte0wl/VAST-ReID Hammad Khan, Rakesh Kumar Giri, Thakare Kamalakar Vijay, Heeseung Choi, Hyungjoo Jung, Debi Prosad Dogra, Ig-Jae Kim |
WACV | 6 |
| 2026 | IMPACT: Interpretable Most Important Person Analysis and Classification using Transformer-based ModelsabstractIdentifying the Most Important Person (MIP) in complex social and sports events remains a challenging problem due to the dynamic nature of group interactions, subtle visual cues, and context-dependent semantics. Traditional methods often struggle to accurately capture the interplay between individuals and the overarching activity, especially in unstructured real-world environments. In addition, the lack of strong supervision and the need for a deeper contextual understanding further complicate the task. In this work, we propose IMPACT, a novel multi-modal framework that leverages recent advances in vision language models to bridge the gap between visual perception and semantic reasoning. Our approach integrates structured scene understanding, natural language generation, and cross-modal learning to jointly model activity recognition and MIP localization. The method integrates language, vision, and spatial reasoning to improve scene interpretability as well as accuracy in group activity recognition tasks. By incorporating language-based representations, the proposed method enables interpretable and robust performance in sports-centric group activity scenarios. Comprehensive experiments on C-Sports and NCAA datasets demonstrate that the framework significantly enhances the localization of key individuals as well as the accuracy of activity prediction, laying the groundwork for a holistic scene understanding in human-centric video and image analysis. Our proposed method achieves an accuracy of 81.6% when compared with human annotator markings and an increase in mAP scores by ∼ 5% for MIP identification. Akshat Rampuria, Kamakshya Prasad Nayak, Thakare Kamalakar Vijay, Tushar Joshi, Aditya Dhananjay Singh, Haesol Park, Heeseung Choi, Hyungjoo Jung, Debi Prosad Dogra, Ig-Jae Kim |
WACV | 9 |
| 2026 | Physics-constrained dynamic graph and multi-resolution temporal modeling for traffic forecasting
Arabinda Panda, Debi Prosad Dogra, Partha Pratim Dey |
Inf. Sci. | 2 |
| 2025 | Can Person-Level Attributes Improve Group Re-Identification?abstractGroup re-identification (G-ReID) attempts to recognize human groups across multiple camera perspectives. It is a challenging task due to occlusion, perspective variation, and illumination change. While Person Attribute Recognition (PAR) methods have shown robustness under similar challenges, yet their potential in G-ReID remains unexplored. Though existing G-ReID datasets are well-crafted, however, they lack person-level attribute annotations. This restricts G-ReID methods to explore attribute-based matching, essentially limiting their capability of multi-modal analysis. In this work, we bridge this gap by utilizing person-level attributes for group-level re-identification. We introduce PAG-ReID (Person Attribute based Group Re-identification), a large-scale dataset constructed by combining three popular G-ReID datasets: CM-Group, Road Group, and CUHK-SYSU-Group. PAG-ReID includes 19K group images encompassing 2,148 groups and 5,504 unique person IDs. At the person level, it provides 25K individual images, each annotated with 19 diverse human attributes, resulting in nearly 475K fine-grained annotations. Next, we propose an effective baseline (CLIP-based) that transforms attribute information into natural language descriptions, enabling joint multi-modal (visual-textual) reasoning for PAR as well as G-ReID tasks. Experiments demonstrate the effectiveness of our approach, setting a new direction for person attribute-centric group re-identification. To our knowledge, this is the first work to unify G-ReID and PAR in a single multi-modal framework. PAG-ReID can be found at https://github.com/draxler1/PAG-ReID. Kamakshya Prasad Nayak, Thakare Kamalakar Vijay, Ashesh Xalxo, Lalit Lohani, Debi Prosad Dogra |
ACM Multimedia | 5 |
| 2025 | CLIPping Imbalances: A Novel Evaluation Baseline and PEARL Dataset for Pedestrian Attribute RecognitionabstractPedestrian Attribute Recognition (PAR) serves as a fun-damental task in computer vision and is crucial for upgradign security systems. It helps in precisely identifying and characterizing various attributes of pedestrians. However, current PAR datasets have certain issues in representing a wide range of attributes correctly, which makes the ex-isting PAR methods less effective in real-world scenarios. Addressing this limitation, this paper introduces PEARL, a comprehensive dataset comprising of diverse pedestrian images annotated with 146 attributes. These samples have been sourced from surveillance videos across twelve coun-tries. This paper also formulates an image-based PAR using language-image fusion strategy and utilizes CLIP as a new evaluation baseline. Specifically, we leverage textual infor-mation by transforming sets of attributes into meaningful sentences. Addressing the inherent data imbalance in PAR, we provide three types of prompt settings to optimize the training of the CLIP model. Our evaluation encompasses a thorough assessment of the proposed baseline model across various datasets, including PEARL dataset as well as estab-lished PAR benchmarks such as PA100K, RAP, and PETA. Thakare Kamalakar Vijay, Lalit Lohani, Kamakshya Prasad Nayak, Debi Prosad Dogra, Heeseung Choi, Hyungjoo Jung, Ig-Jae Kim |
WACV | 4 |
| 2024 | Pedestrian Attribute Recognition Using Hierarchical Transformers
Lalit Lohani, Thakare Kamalakar Vijay, Kamakshya Prasad Nayak, Debi Prosad Dogra, Heeseung Choi, Hyungjoo Jung, Ig-Jae Kim |
ICPR (16) | 4 |
| 2024 | Extracting Vitals from ICU Monitor Images: An Insight from Analysis of 10K Patient Data
Akshat Rampuria, Kushagra Khare, Ayush Soni, Debi Prosad Dogra |
ICPR (12) | 4 |
| 2024 | Let's Observe Them Over Time: An Improved Pedestrian Attribute Recognition ApproachabstractDespite poor image quality, occlusions, and small training datasets, recent pedestrian attribute recognition (PAR) methods have achieved considerable performance. However, leveraging only spatial information of different attributes limits their reliability and generalizability. This paper introduces a multi-perspective approach to reduce over-dependence on spatial clues of a single perspective and exploits other aspects available in multiple perspectives. In order to tackle image quality and occlusions, we exploit different spatial clues present across images and handpick the best attribute-specific features to classify. Precisely, we extract the class-activation energy of each attribute and correlate it with the corresponding energy present across other images using the proposed Self-Attentive Cross Relation Module. In the next stage, we fuse this correlation information with similar clues accumulated from the other images. Lastly, we train a classification neural network using combined correlation information with two different losses. We have validated our method on four widely used PAR datasets, namely Market1501, PETA, PA-100k, and Duke. Our method achieves superior performance over most existing methods, demonstrating the effectiveness of a multi-perspective approach in PAR. Thakare Kamalakar Vijay, Debi Prosad Dogra, Heeseung Choi, Haksub Kim, Ig-Jae Kim |
WACV | 2 |
| 2024 | PlutoAR: a scalable marker-based augmented reality application for interactive and inclusive education
Ajaya Kumar Dash, Santosh Kumar Behera, Debi Prosad Dogra |
Multim. Tools Appl. | 3 |
| 2024 | Interactions with 3D virtual objects in augmented reality using natural gestures
Ajaya Kumar Dash, Koniki Venkata Balaji, Debi Prosad Dogra, Byung-Gyu Kim |
Vis. Comput. | 3 |
| 2023 | DyAnNet: A Scene Dynamicity Guided Self-Trained Video Anomaly Detection NetworkabstractUnsupervised approaches for video anomaly detection may not perform as good as supervised approaches. However, learning unknown types of anomalies using an unsupervised approach is more practical than a supervised approach as annotation is an extra burden. In this paper, we use isolation tree-based unsupervised clustering to partition the deep feature space of the video segments. The RGB-stream generates a pseudo anomaly score and the flow stream generates a pseudo dynamicity score of a video segment. These scores are then fused using a majority voting scheme to generate preliminary bags of positive and negative segments. However, these bags may not be accurate as the scores are generated only using the current segment which does not represent the global behavior of a typical anomalous event. We then use a refinement strategy based on a cross-branch feed-forward network designed using a popular I3D network to refine both scores. The bags are then refined through a segment re-mapping strategy. The intuition of adding the dynamicity score of a segment with the anomaly score is to enhance the quality of the evidence. The method has been evaluated on three popular video anomaly datasets, i.e., UCF-Crime, CCTV-Fights, and UBI-Fights. Experimental results reveal that the proposed framework achieves competitive accuracy as compared to the state-of-the-art video anomaly detection methods. Thakare Kamalakar Vijay, Yash Raghuwanshi, Debi Prosad Dogra, Heeseung Choi, Ig-Jae Kim |
WACV | 3 |
| 2023 | RareAnom: A Benchmark Video Dataset for Rare Type AnomaliesabstractExisting video anomaly detection methods and datasets suffer from restricted anomaly categories containing single-source (CCTV) videos recorded in controlled environment, inadequate annotations, and lack of adequate supervision. To mitigate these problems, we introduce a new dataset ( RareAnom ) containing 17 rare types of real-world anomalies (2200 videos) recorded using multiple sources (e.g., CCTV , handheld cameras, dash-cams, and mobile phones) with rich temporal annotations. A new fully unsupervised anomaly detection and classification method has been proposed. It has three stages: training of a 3D Convolution Autoencoder using pseudo-labelled video segments, anomaly detection using latent features, and classification. Unlike the existing datasets, we have benchmarked RareAnom using three levels of supervision: fully, weakly, and unsupervised. It has been compared with UCF-Crime and XD-Violence datasets. The proposed anomaly detection and classification method beats the latest unsupervised methods by 4.49%, 8.66%, and 6.77% on RareAnom, UCF-Crime, and XD-violence datasets, respectively. Thakare Kamalakar Vijay, Debi Prosad Dogra, Heeseung Choi, Haksub Kim, Ig-Jae Kim |
Pattern Recognit. | 2 |
| 2023 | Crowd Characterization in Surveillance Videos Using Deep-Graph Convolutional Neural NetworkabstractCrowd behavior is a natural phenomenon that can provide valuable insight into the crowd characterization process. Modeling the visual appearance of a large crowd gathering can reveal meaningful information about its dynamics. Parametric modeling can be used to develop efficient and robust crowd monitoring systems. A crowd can be structured or unstructured based on the organization. In this article, crowd characterization has been mapped to a graph classification problem to classify movements based on order parameter ( ϕ ), active force components, and steadiness (Reynolds number). The graphs are constructed from the motion groups obtained using an active Langevin framework. These graphs are processed using a deep graph convolutional neural network for crowd characterization. For experimentation, we have prepared a dataset comprising of videos from popular publicly available datasets and our own recorded videos. The proposed framework has been compared with the latest deep learning-based frameworks in terms of accuracy and area under the curve (AUC). We have obtained a 4%-5% improvement in accuracy and AUC values over the existing frameworks. The insights obtained from the proposed framework can be used for better crowd monitoring and management. Shreetam Behera, Debi Prosad Dogra, Malay Kumar Bandyopadhyay, Partha Pratim Roy 0001 |
IEEE Trans. Cybern. | 2 |
| 2023 | Detection of Road Accidents Using Synthetically Generated Multi-Perspective Accident VideosabstractRoad accidents are often caused by short abnormal events, including traffic violations, abrupt change in vehicular motion, driver fatigue, etc. Observing an accident event from the right camera perspective plays a crucial role while detecting accidents. However, it may not be possible to capture such abnormal events from a limited camera perspective. We present a deep learning framework to analyze the accident events recorded from multiple perspectives. First, we estimate feature similarity in videos recorded from multiple perspectives. We then divided the video samples into high and low feature similarity groups. Next, we extract spatio-temporal features from each group using two-branch DCNNs and fuse them using a rank-based weighted average pooling strategy followed by classification. We present a new road accident video dataset (MP-RAD), where each accident event is synthetically generated and captured from five independent camera perspectives using a computer gaming platform. Most of the existing road accident datasets use egocentric views or they are captured in fixed camera setups. However, our dataset is large and multi-perspective that can be used to validate ITS-related tasks such as accident detection, accident localization, traffic monitoring, etc. The dataset contains 400 accident events with a total of 2000 videos. We provide temporal annotations of all videos. The proposed framework and the dataset have been cross-validated with latest accident detection baselines trained on real-world road accident videos and vice-versa. The sub-optimal detection accuracy obtained using the baselines indicates that the proposed framework and the dataset can be useful for ITS related research. Code and dataset is available at: https://github.com/draxler1/MP-RAD-Dataset-ITS- Thakare Kamalakar Vijay, Debi Prosad Dogra, Heeseung Choi, Gi Pyo Nam, Ig-Jae Kim |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Person re-identification in indoor videos by information fusion using Graph Convolutional Networks
Komal Soni, Debi Prosad Dogra, Arif Ahmed 0002, Samarjit Kar, Heeseung Choi, Ig-Jae Kim |
Expert Syst. Appl. | 2 |
| 2022 | A multi-stream deep neural network with late fuzzy fusion for real-world anomaly detection
Thakare Kamalakar Vijay, Nitin Sharma 0004, Debi Prosad Dogra, Heeseung Choi, Ig-Jae Kim |
Expert Syst. Appl. | 3 |
| 2022 | Video based exercise recognition and correct pose detection
Tushar Rangari, Sudhanshu Kumar 0002, Partha Pratim Roy 0001, Debi Prosad Dogra, Byung-Gyu Kim |
Multim. Tools Appl. | 4 |
| 2022 | Vehicular Trajectory Classification and Traffic Anomaly Detection in Videos Using a Hybrid CNN-VAE ArchitectureabstractVisual surveillance has become indispensable in the evolution of Intelligent Transportation Systems (ITS). Video object trajectories are key to many of the visual surveillance applications. Classifying varying length time series data such as video object trajectories using conventional neural networks, can be challenging. In this paper, we propose trajectory classification and anomaly detection using a hybrid Convolutional Neural Network (CNN) and Variational Autoencoder (VAE) architecture. First, we introduce a high level features for varying length object trajectories using color gradient representation. In the next stage, a semi-supervised way to annotate moving object trajectories extracted using Temporally Incremental Gravitational Model (TIGM) is used for class labeling. For training, anomalous trajectories are identified using t-Distributed Stochastic Neighbor Embedding (t-SNE). Finally, a hybrid CNN-VAE architecture has been proposed for trajectory classification and anomaly detection. The results obtained using publicly available surveillance video datasets reveal that the proposed method can successfully identify traffic anomalies such as violations in lane driving, sudden speed variations, abrupt termination of vehicle during movement, and vehicles moving in wrong directions. The accuracy of trajectory classification improves by a margin of 1-6% against popular neural networks-based classifiers across various datasets using the proposed high-level features. The gradient representation also improves the anomaly detection accuracy significantly (30-35%). Code and dataset can be found athttps://github.com/santhoshkelathodi/CNN-VAE. Santhosh Kelathodi Kumaran, Debi Prosad Dogra, Partha Pratim Roy 0001, Adway Mitra |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Object Interaction-Based Localization and Description of Road Accident Events Using Deep LearningabstractDetection and localization of road accidents in real-time is an integral part of the Intelligent Transportation System (ITS). Even though the existing road accident detection methods show promising results, the process suffers from some drawbacks. For example, existing methods require a large number of sample videos for feature learning. Moreover, features such as temporal gradients or flow fields are time-consuming. To address these issues, we introduce a new method that uses objects and their positions to detect accidents in real-time. Apart from localization of the accident events in videos, we perform a high-level post processing to describe the severity and context of an accident. Firstly, we divide an input video into pre-accident, accident and post-accident stages to extract object interactions. These interaction proposals are then filtered using a refinement algorithm. We then adopt an iterative training procedure to classify normal and accident interactions. We also highlight the damaged zone using heat maps. Finally, we generate high-level textual descriptions to quantify the context and severity of an accident. We have trained the proposed model using offline setups. However, it can be deployed online to detect road accident events in real-time by taking the video inputs directly from the CCTV camera. Moreover, with a minimal supervision, the model can be retrained for online surveillance. Extensive experiments carried out on UCF Crime and CADP datasets reveal that the proposed framework achieves state-of-the-art performance when compared with the recently proposed accident event detection methods in terms of AUC (UCF Crime: 69.70% and CADP: 72.59%) and FAR (UCF Crime: 0.8 and CADP: 2.2). The high-level description of the accident is an added advantage that will certainly help the traffic police to react in a timely manner. Thakare Kamalakar Vijay, Debi Prosad Dogra, Heeseung Choi, Haksub Kim, Ig-Jae Kim |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | PIDLNet: A Physics-Induced Deep Learning Network for Characterization of Crowd VideosabstractHuman visual perception regarding crowd gatherings can provide valuable information about behavioral movements. Empirical analysis on visual perception about orderly moving crowds has revealed that such movements are often structured in nature with relatively higher order parameter and lower entropy as compared to unstructured crowd, and vice-versa. This paper proposes a Physics-Induced Deep Learning Network (PIDLNet), a deep learning framework trained on conventional 3D convolutional features combined with physics-based features. We have computed frame-level entropy and order parameter from the motion flows extracted from the crowd videos. These features are then integrated with the 3D convolutional features at a later stage in the feature extraction pipeline to aid in the crowd characterization process. Experiments reveal that the proposed network can characterize video segments depicting crowd movements with accuracy as high as 91.63%. We have obtained overall AUC of 0.9913 on highly challenging publicly available video dataset. The method outperforms existing deep-learning frameworks and conventional crowd characterization frameworks by a notable margin. Shreetam Behera, Thakare Kamalakar Vijay, H. Manish Kausik, Debi Prosad Dogra |
AVSS | 4 |
| 2021 | Logo detection using weakly supervised saliency map
Prateek Keserwani, Partha Pratim Roy 0001, Debi Prosad Dogra |
Multim. Tools Appl. | 4 |
| 2021 | Understanding crowd flow patterns using active-Langevin model
Shreetam Behera, Debi Prosad Dogra, Malay Kumar Bandyopadhyay, Partha Pratim Roy 0001 |
Pattern Recognit. | 2 |
| 2021 | Trajectory-Based Scene Understanding Using Dirichlet Process Mixture ModelabstractAppropriate modeling of a surveillance scene is essential for the detection of anomalies in road traffic. Learning usual paths can provide valuable insight into road traffic conditions and thus can help in identifying unusual routes taken by commuters/vehicles. If usual traffic paths are learned in a nonparametric way, manual interventions in road marking can be avoided. In this paper, we propose an unsupervised and nonparametric method to learn the frequently used paths from the tracks of moving objects in Θ(kn) time, where k denotes the number of paths and n represents the number of tracks. In the proposed method, temporal dependencies of the moving objects are considered to make the clustering meaningful using temporally incremental gravity model (TIGM). In addition, the distance-based scene learning makes it intuitive to estimate the model parameters. Further, we have extended the TIGM hierarchically as a dynamically evolving model (DEM) to represent notable traffic dynamics of a scene. The experimental validation reveals that the proposed method can learn a scene quickly without prior knowledge about the number of paths ( k ). We have compared the results with various state-of-the-art methods. We have also highlighted the advantages of the proposed method over the existing techniques popularly used for designing traffic monitoring applications. It can be used for administrative decision making to control traffic at junctions or crowded places and generate alarm signals, if necessary. Santhosh Kelathodi Kumaran, Debi Prosad Dogra, Partha Pratim Roy 0001, Bidyut B. Chaudhuri |
IEEE Trans. Cybern. | 2 |
| 2020 | Estimation of linear motion in dense crowd videos using Langevin model
Shreetam Behera, Debi Prosad Dogra, Malay Kumar Bandyopadhyay, Partha Pratim Roy 0001 |
Expert Syst. Appl. | 2 |
| 2020 | Person Re-identification in Videos by Analyzing Spatio-temporal TubesabstractAbstract Typical person re-identification frameworks search for k best matches in a gallery of images that are often collected in varying conditions. The gallery usually contains image sequences for video re-identification applications. However, such a process is time consuming as video re-identification involves carrying out the matching process multiple times. In this paper, we propose a new method that extracts spatio-temporal frame sequences or tubes of moving persons and performs the re-identification in quick time. Initially, we apply a binary classifier to remove noisy images from the input query tube. In the next step, we use a key-pose detection-based query minimization technique. Finally, a hierarchical re-identification framework is proposed and used to rank the output tubes. Experiments with publicly available video re-identification datasets reveal that our framework is better than existing methods. It ranks the tubes with an average increase in the CMC accuracy of 6-8% across multiple datasets. Also, our method significantly reduces the number of false positives. A new video re-identification dataset, named Tube-based Re-identification Video Dataset (TRiViD), has been prepared with an aim to help the re-identification research community. Arif Ahmed 0002, Debi Prosad Dogra, Heeseung Choi, Seungho Chae, Ig-Jae Kim |
Multim. Tools Appl. | 2 |
| 2020 | Retrieval of colour and texture images using local directional peak valley binary pattern
Partha Pratim Roy 0001, Debi Prosad Dogra, Byung-Gyu Kim |
Pattern Anal. Appl. | 3 |
| 2020 | Can we automate diagrammatic reasoning?abstractDiagrammatic reasoning (DR) problems are well known. However, solving DR problems represented in 4 × 1 Raven’s Progressive Matrix (RPM) form using computer vision and pattern recognition has not yet been tried. Emergence of deep learning techniques aided by advanced computing can be exploited to solve such DR problems. In this paper, we propose a new learning framework by combining LSTM and Convolutional LSTM to solve 4 × 1 DR problems. Initially, the elementary geometrical shapes in such problems are detected using a typical CNN-based detector. Next, relations of various shapes are analyzed and a high-level feature set is produced and processed in the LSTM framework. A new 4 × 1 DR dataset has been prepared and made available to the research community. We believe, it will be helpful in advancing this research further. We have compared our method with some of the existing frameworks that can be used for solving RPM-guided DR problems. We have recorded 18–20% increase in the average prediction accuracy as compared to the prior frameworks when applied to RPM-guided DR problems. We believe the CV research community will be interested to carry out similar research, particularly to investigate the feasibility of solving other types of known DR problems. Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Partha Pratim Roy 0001, Dilip K. Prasad |
Pattern Recognit. | 2 |
| 2020 | Video trajectory analysis using unsupervised clustering and multi-criteria rankingabstractAbstract Surveillance camera usage has increased significantly for visual surveillance. Manual analysis of large video data recorded by cameras may not be feasible on a larger scale. In various applications, deep learning-guided supervised systems are used to track and identify unusual patterns. However, such systems depend on learning which may not be possible. Unsupervised methods relay on suitable features and demand cluster analysis by experts. In this paper, we propose an unsupervised trajectory clustering method referred to as t-Cluster. Our proposed method prepares indexes of object trajectories by fusing high-level interpretable features such as origin, destination, path, and deviation. Next, the clusters are fused using multi-criteria decision making and trajectories are ranked accordingly. The method is able to place abnormal patterns on the top of the list. We have evaluated our algorithm and compared it against competent baseline trajectory clustering methods applied to videos taken from publicly available benchmark datasets. We have obtained higher clustering accuracies on public datasets with significantly lesser computation overhead. Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Partha Pratim Roy 0001 |
Soft Comput. | 2 |
| 2020 | Query-Based Video Synopsis for Intelligent Traffic Monitoring ApplicationsabstractSynopsis of a long-duration video has many applications in intelligent transportation systems. It can help to monitor traffic with lesser manpower. However, generating meaningful synopsis of a long-duration video recording can be challenging. Often summarized outputs include redundant contents or activities that may not be helpful to the observer. Moving object trajectories are possible sources of information that can be used to generate the synopsis of long-duration videos. The synopsis generation faces challenges due to object tracking, grouping of the trajectories with respect to activity type, object category, and contextual information, and generating smooth synopsis according to a query. In this paper, we propose a method to generate meaningful and smooth synopsis of long-duration videos according to the users' query. We have tracked moving objects and adopted deep learning to classify the objects into known categories (e.g., car, bike, and pedestrians). We then identify regions in the surveillance scene with the help of unsupervised clustering. Each tube (spatiotemporal object trajectory) is represented by the source and the destination. In the final stage, we take a query from the user and generate the synopsis video by smoothly blending the appropriate tubes over the background frame through energy minimization. The proposed method has been evaluated on two publicly available datasets and our own surveillance datasets. We have compared the method with popular state-of-the-art techniques. The experiments reveal that the proposed method is superior to the existing techniques and it produces visually seamless video synopsis. Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Renuka Patnaik, Seung-Cheol Lee, Heeseung Choi, Gi Pyo Nam, Ig-Jae Kim |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2019 | Natural Gestures to Interact with 3D Virtual Objects using Deep Learning FrameworkabstractThis paper presents a system for freehand interaction with 3D objects using gestures. A secondary contribution of this work is to present a system to recognize a set of suitable gestures which are used for manipulation of the 3D objects interactively, with the help of a deep learning framework using the 3D raw images captured via Leap motion interface. For this study, we have used our own dataset, a collection of images acquired using Leap motion controller comprises with six naturally occurring gestures. A deep learning framework built with CNN has been trained on this dataset and the validation accuracy has been found to be as high as 99%. This model is then used to predict the user's gesture to interact with the 3D objects with bare hands. The application has been tried and assessed by ten subjects. The subjects had no prior experience on how to interact with Leap motion sensor, making it an interesting study to explore the possibility of its usage by mass. Suraj Tripathy, Rohan Sahoo, Ajaya Kumar Dash, Debi Prosad Dogra |
TENCON | 4 |
| 2019 | Queuing theory guided intelligent traffic scheduling through video analysis using Dirichlet process mixture model
Santhosh Kelathodi Kumaran, Debi Prosad Dogra, Partha Pratim Roy 0001 |
Expert Syst. Appl. | 2 |
| 2019 | Computer vision-guided intelligent traffic signaling for isolated intersections
Santhosh Kelathodi Kumaran, Shrohan Mohapatra, Debi Prosad Dogra, Partha Pratim Roy 0001, Byung-Gyu Kim |
Expert Syst. Appl. | 3 |
| 2019 | Fingertip detection and tracking for recognition of air-writing in videos
Sohom Mukherjee, Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Partha Pratim Roy 0001 |
Expert Syst. Appl. | 3 |
| 2019 | Likelihood learning in modified Dirichlet Process Mixture Model for video analysis
Santhosh Kelathodi Kumaran, Adyasha Chakravarty, Debi Prosad Dogra, Partha Pratim Roy 0001 |
Pattern Recognit. Lett. | 3 |
| 2019 | Recognizing gender from human facial regions using genetic algorithm
Avirup Bhattacharyya, Rajkumar Saini, Partha Pratim Roy 0001, Debi Prosad Dogra, Samarjit Kar |
Soft Comput. | 4 |
| 2019 | Trajectory-Based Surveillance Analysis: A SurveyabstractDue to the advancement of camera hardware and machine learning techniques, video object tracking for surveillance has received noticeable attention from the computer vision research community. Object tracking and trajectory modeling have important applications in surveillance video analysis. For example, trajectory clustering, summarization or synopsis generation, and detection of anomalous or abnormal events in videos are mainly being exploited by the research community. However, barring one research work (which is almost a decade old), there is no recent review that emphasizes the use of video object trajectories, particularly in the perspective of visual surveillance. This paper presents a survey of trajectory-based surveillance applications with a focus on clustering, anomaly detection, summarization, and synopsis generation. The methods reviewed in this paper broadly summarize the abovementioned applications. The main purpose of this survey is to summarize the state-of-the-art video object trajectory analysis techniques used in the indoor and outdoor surveillance. Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Partha Pratim Roy 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Multimodal Gait Recognition With Inertial Sensor Data and Video Using Evolutionary AlgorithmabstractEvolutionary decision fusion has applications in biometric authentication and verification. Gray wolf optimizer (GWO) is one such evolutionary decision fusion approach that can be used to tune the fusion parameters in a multimodal data acquisition system. Human gait is a proven biometric trait with applications in security and authentication. However, acquiring human-gait data can be erroneous due to various factors and multimodal fusion of such erroneous gait data can be challenging. In this paper, we propose a new decision fusion-based approach to solve the above problem. Gait data is recorded simultaneously using motion sensors and visible-light camera. The signals of the motion sensors are modeled using a long short-term memory neural network and corresponding video recordings are processed using a three-dimensional convolutional neural network. GWO has been used to optimize the parameters during fusion. It has been chosen based on the underlying hunting strategy that leads to better approximation of the solution. Interestingly, in our case it converges quicker than other optimization techniques such as genetic algorithm or particle swarm optimization. To test the model, a dataset involving 23 males and females has been recorded while they perform four different types of walks, including, normal walk, fast walk, walking while listening to music, and walking while watching multimedia content on a mobile. An overall accuracy of 91.3% has been recorded across all test scenarios. Results reveal that the proposed study can further be explored to design robust gait biometric systems. Pradeep Kumar 0002, Subham Mukherjee, Rajkumar Saini, Pallavi Kaushik, Partha Pratim Roy 0001, Debi Prosad Dogra |
IEEE Trans. Fuzzy Syst. | 6 |
| 2019 | Temporal Unknown Incremental Clustering Model for Analysis of Traffic Surveillance VideosabstractOptimized scene representation is an important characteristic of a framework for detecting abnormalities on live videos. One of the challenges for detecting abnormalities in live videos is real-time detection of objects in a non-parametric way. Another challenge is to efficiently represent the state of objects temporally across frames. In this paper, a Gibbs sampling-based heuristic model referred to as temporal unknown incremental clustering has been proposed to cluster pixels with motion. Pixel motion is first detected using optical flow and a Bayesian algorithm has been applied to associate pixels belonging to a similar cluster in subsequent frames. The algorithm is fast and produces accurate results in Θ(kn) time, where k is the number of clusters and n the number of pixels. Our experimental validation with publicly available data sets reveals that the proposed framework has good potential to open up new opportunities for real-time traffic analysis. Santhosh Kelathodi Kumaran, Debi Prosad Dogra, Partha Pratim Roy 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2019 | A novel point-line duality feature for trajectory classification
Rajkumar Saini, Partha Pratim Roy 0001, Debi Prosad Dogra |
Vis. Comput. | 3 |
| 2018 | Air Signature Recognition Using Deep Convolutional Neural Network-Based Sequential ModelabstractDeep convolutional neural networks are becoming extremely popular in classification, especially when the inputs are non-sequential in nature. Though it seems unrealistic to adopt such networks as sequential classifiers, however, researchers have started to use them for applications that primarily deal with sequential data. It is possible, if the sequential data can be represented in the conventional way the inputs are provided in CNNs. Signature recognition is one of the important tasks for biometric applications. Signatures represent the signer's identity. Air signatures can make traditional biometric systems more secure and robust than conventional pen-paper or stylus guided interfaces. In this paper, we propose a new set of geometrical features to represent 3D air signatures captured using Leap motion sensor. The features are then arranged such that they can be fed to a deep convolutional neural network architecture with application specific tuning of the model parameters. It has been observed that the proposed features in combination with the CNN architecture can act as a good sequential classifier when tested on a moderate size air signature dataset. Experimental results reveal that the proposed biometric system performs better as compared to the state-of-the-art geometrical features with average accuracy improvement of 4%. Santosh Kumar Behera, Ajaya Kumar Dash, Debi Prosad Dogra, Partha Pratim Roy 0001 |
ICPR | 3 |
| 2018 | Surveillance scene representation and trajectory abnormality detection using aggregation of multiple concepts
Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Partha Pratim Roy 0001 |
Expert Syst. Appl. | 2 |
| 2018 | Fast recognition and verification of 3D air signatures using convex hulls
Santosh Kumar Behera, Debi Prosad Dogra, Partha Pratim Roy 0001 |
Expert Syst. Appl. | 2 |
| 2018 | A segmental HMM based trajectory classification using genetic algorithm
Rajkumar Saini, Partha Pratim Roy 0001, Debi Prosad Dogra |
Expert Syst. Appl. | 3 |
| 2018 | A novel framework of continuous human-activity recognition using Kinect
Rajkumar Saini, Pradeep Kumar 0002, Partha Pratim Roy 0001, Debi Prosad Dogra |
Neurocomputing | 4 |
| 2018 | Independent Bayesian classifier combination based sign language recognition using facial expression
Pradeep Kumar 0002, Partha Pratim Roy 0001, Debi Prosad Dogra |
Inf. Sci. | 3 |
| 2018 | A position and rotation invariant framework for sign language recognition (SLR) using Kinect
Pradeep Kumar 0002, Rajkumar Saini, Partha Pratim Roy 0001, Debi Prosad Dogra |
Multim. Tools Appl. | 4 |
| 2018 | Analysis of 3D signatures recorded using leap motion sensor
Santosh Kumar Behera, Debi Prosad Dogra, Partha Pratim Roy 0001 |
Multim. Tools Appl. | 2 |
| 2018 | Motion anomaly detection and trajectory analysis in visual surveillance
Manaswi Chebiyyam, Rohit Desam Reddy, Debi Prosad Dogra, Harish Bhaskar, Lyudmila Mihaylova |
Multim. Tools Appl. | 3 |
| 2018 | Exercise classification and event segmentation in Hammersmith Infant Neurological Examination videos
Abdul Fatir Ansari, Partha Pratim Roy 0001, Debi Prosad Dogra |
Mach. Vis. Appl. | 3 |
| 2018 | Envisioned speech recognition using EEG sensors
Pradeep Kumar 0002, Rajkumar Saini, Partha Pratim Roy 0001, Pawan Kumar Sahu, Debi Prosad Dogra |
Pers. Ubiquitous Comput. | 5 |
| 2018 | Unsupervised classification of erroneous video object trajectories
Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Partha Pratim Roy 0001 |
Soft Comput. | 2 |
| 2017 | A multimodal framework for sensor based sign language recognition
Pradeep Kumar 0002, Himaanshu Gauba, Partha Pratim Roy 0001, Debi Prosad Dogra |
Neurocomputing | 4 |
| 2017 | A bio-signal based framework to secure mobile devices
Pradeep Kumar 0002, Rajkumar Saini, Partha Pratim Roy 0001, Debi Prosad Dogra |
J. Netw. Comput. Appl. | 4 |
| 2017 | Localization of region of interest in surveillance scene
Arif Ahmed 0002, Debi Prosad Dogra, Samarjit Kar, Byung-Gyu Kim, Paul R. Hill, Harish Bhaskar |
Multim. Tools Appl. | 2 |
| 2017 | 3D text segmentation and recognition using leap motion
Pradeep Kumar 0002, Rajkumar Saini, Partha Pratim Roy 0001, Debi Prosad Dogra |
Multim. Tools Appl. | 4 |
| 2017 | Analysis of EEG signals and its application to neuromarketing
Mahendra Yadava, Pradeep Kumar 0002, Rajkumar Saini, Partha Pratim Roy 0001, Debi Prosad Dogra |
Multim. Tools Appl. | 5 |
| 2017 | Prediction of advertisement preference by fusing EEG response and sentiment analysis
Himaanshu Gauba, Pradeep Kumar 0002, Partha Pratim Roy 0001, Priyanka Singh 0001, Debi Prosad Dogra, Balasubramanian Raman |
Neural Networks | 5 |
| 2017 | Coupled HMM-based multi-sensor data fusion for sign language recognition
Pradeep Kumar 0002, Himaanshu Gauba, Partha Pratim Roy 0001, Debi Prosad Dogra |
Pattern Recognit. Lett. | 4 |
| 2017 | Moving object detection using modified temporal differencing and local fuzzy thresholding
Nihal Paul, Abhishek Midya, Partha Pratim Roy 0001, Debi Prosad Dogra |
J. Supercomput. | 5 |
| 2016 | Classification of head movement patterns to aid patients undergoing home-based cervical spine rehabilitationabstractPhysical rehabilitation under close supervision of experts is often recommended to patients suffering from cervical pain and join related injuries. However, complete supervised rehabilitation may not be possible due to various socio-economical parameters. On the other hand, unsupervised rehabilitation may lead to post-injury complications. In this paper, we propose a pattern classification based method to understand head movement that is necessary to recognize unusual patterns during cervical spine rehabilitation. The proposed system takes the help of Kalman filter to fuse data acquired through camera and motion sensors. Patterns of the fused signals represent head displacement during left-right and front-back movements with respect to the central axes parallel to coronal and sagittal planes. Normal and abnormal patterns of movements are thereby represented using time series data that describes temporal change in the angle of head on above planes. Experimental validation has been done with data collected from several users. It has been observed that, our proposed methodology can precisely detect and represent abnormalities in head movements during cervical spine rehabilitation. K. M. Vamsikrishna, Debi Prosad Dogra, Harish Bhaskar |
ICASSP | 2 |
| 2016 | Smart video summarization using mealy machine-based trajectory modelling for surveillance applications
Debi Prosad Dogra, Arif Ahmed 0002, Harish Bhaskar |
Multim. Tools Appl. | 1 |
| 2015 | Autonomous detection and tracking under illumination changes, occlusions and moving camera
Harish Bhaskar, Kartik Dwivedi, Debi Prosad Dogra, Mohammed E. Al-Mualla, Lyudmila Mihaylova |
Signal Process. | 3 |
| 2012 | Evaluation of segmentation techniques using region area and boundary matching information
Debi Prosad Dogra, Arun K. Majumdar, Shamik Sural |
J. Vis. Commun. Image Represent. | 1 |
| 2011 | A Web Enabled Health Information System for the Neonatal Intensive Care Unit (NICU)abstractInformation Systems are needed for modernization of ICUs to deliver better health care services. EHR systems can improve the work flow management in health care delivery. This work proposes a secure web enabled system based on a multi-tier architecture for carrying out routine and special operations of Neonatal Intensive Care Unit (NICU). The system adopts a service oriented approach for execution of various tasks that are performed for managing NICU activities. It also facilitates decision support systems for a number of critical tasks of NICU. A prototype of the system has been installed in the neonatology department of SSKM Hospital, Kolkata, India and the staff of the hospital including doctors, nurses, laboratory personals and technicians are using it in a regular manner. Soumendranath Ray, Debi Prosad Dogra, Bhaskar Saha, Arunava Biswas, Arun K. Majumdar, Jayanta Mukhopadhyay, Bandana Majumdar, A. Paria, Suchandra Mukherjee, S. Das Bhattacharya |
SERVICES | 2 |