Thakare Kamalakar Vijay

dblp:310/7408 · also Kamalakar Vijay Thakare · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0003-4587-4126ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 VAST-ReID: A Low-Light Benchmark Dataset for Person Re-Identification with Visual and Attribute-Rich Semantic Tracking
abstract
Person Re-Identification (ReID) task is important for designing intelligent surveillance systems. ReID can be highly challenging in low-light and low resolution scenarios. Existing ReID datasets predominantly feature cropped pedestrian images captured in well-lit environments, often lacking semantic richness, frame-level temporal continuity, and robustness to adverse conditions. To address these limitations, we introduce VAST-ReID, a new benchmark dataset specifically designed for the low-light person ReID task in real-world surveillance contexts. VAST-ReID consists of 1,441 surveillance videos collected at 24 different locations, capturing 256 distinct pedestrians of various age groups. The dataset emphasizes naturally low-light and visually degraded scenarios. Each identity is annotated with dense bounding boxes and enriched with auxiliary semantic labels, including pedestrian attributes and LLM-generated descriptions. While these annotations are not used during supervised training, they provide valuable semantic context for advancing research in language-guided retrieval and attribute-aware modeling. Additionally, we release identity-aligned image crops under the BoxTrack-ReID subset, which has over 18.7K frames sampled at 1fps from the raw videos, with standard training, gallery, and query splits compatible with the Market-1501 evaluation protocol, enabling straightforward benchmarking. The dataset has been benchmarked against SOTA methods, and experiments reveal that there is huge scope for improvement in ReID research. VAST-ReID is available at: https://github.com/Byte0wl/VAST-ReID
Hammad Khan, Rakesh Kumar Giri, Thakare Kamalakar Vijay, Heeseung Choi, Hyungjoo Jung, Debi Prosad Dogra, Ig-Jae Kim
WACV3
2026 IMPACT: Interpretable Most Important Person Analysis and Classification using Transformer-based Models
abstract
Identifying the Most Important Person (MIP) in complex social and sports events remains a challenging problem due to the dynamic nature of group interactions, subtle visual cues, and context-dependent semantics. Traditional methods often struggle to accurately capture the interplay between individuals and the overarching activity, especially in unstructured real-world environments. In addition, the lack of strong supervision and the need for a deeper contextual understanding further complicate the task. In this work, we propose IMPACT, a novel multi-modal framework that leverages recent advances in vision language models to bridge the gap between visual perception and semantic reasoning. Our approach integrates structured scene understanding, natural language generation, and cross-modal learning to jointly model activity recognition and MIP localization. The method integrates language, vision, and spatial reasoning to improve scene interpretability as well as accuracy in group activity recognition tasks. By incorporating language-based representations, the proposed method enables interpretable and robust performance in sports-centric group activity scenarios. Comprehensive experiments on C-Sports and NCAA datasets demonstrate that the framework significantly enhances the localization of key individuals as well as the accuracy of activity prediction, laying the groundwork for a holistic scene understanding in human-centric video and image analysis. Our proposed method achieves an accuracy of 81.6% when compared with human annotator markings and an increase in mAP scores by ∼ 5% for MIP identification.
Akshat Rampuria, Kamakshya Prasad Nayak, Thakare Kamalakar Vijay, Tushar Joshi, Aditya Dhananjay Singh, Haesol Park, Heeseung Choi, Hyungjoo Jung, Debi Prosad Dogra, Ig-Jae Kim
WACV3
2025 Can Person-Level Attributes Improve Group Re-Identification?
abstract
Group re-identification (G-ReID) attempts to recognize human groups across multiple camera perspectives. It is a challenging task due to occlusion, perspective variation, and illumination change. While Person Attribute Recognition (PAR) methods have shown robustness under similar challenges, yet their potential in G-ReID remains unexplored. Though existing G-ReID datasets are well-crafted, however, they lack person-level attribute annotations. This restricts G-ReID methods to explore attribute-based matching, essentially limiting their capability of multi-modal analysis. In this work, we bridge this gap by utilizing person-level attributes for group-level re-identification. We introduce PAG-ReID (Person Attribute based Group Re-identification), a large-scale dataset constructed by combining three popular G-ReID datasets: CM-Group, Road Group, and CUHK-SYSU-Group. PAG-ReID includes 19K group images encompassing 2,148 groups and 5,504 unique person IDs. At the person level, it provides 25K individual images, each annotated with 19 diverse human attributes, resulting in nearly 475K fine-grained annotations. Next, we propose an effective baseline (CLIP-based) that transforms attribute information into natural language descriptions, enabling joint multi-modal (visual-textual) reasoning for PAR as well as G-ReID tasks. Experiments demonstrate the effectiveness of our approach, setting a new direction for person attribute-centric group re-identification. To our knowledge, this is the first work to unify G-ReID and PAR in a single multi-modal framework. PAG-ReID can be found at https://github.com/draxler1/PAG-ReID.
Kamakshya Prasad Nayak, Thakare Kamalakar Vijay, Ashesh Xalxo, Lalit Lohani, Debi Prosad Dogra
ACM Multimedia2
2025 CLIPping Imbalances: A Novel Evaluation Baseline and PEARL Dataset for Pedestrian Attribute Recognition
abstract
Pedestrian Attribute Recognition (PAR) serves as a fun-damental task in computer vision and is crucial for upgradign security systems. It helps in precisely identifying and characterizing various attributes of pedestrians. However, current PAR datasets have certain issues in representing a wide range of attributes correctly, which makes the ex-isting PAR methods less effective in real-world scenarios. Addressing this limitation, this paper introduces PEARL, a comprehensive dataset comprising of diverse pedestrian images annotated with 146 attributes. These samples have been sourced from surveillance videos across twelve coun-tries. This paper also formulates an image-based PAR using language-image fusion strategy and utilizes CLIP as a new evaluation baseline. Specifically, we leverage textual infor-mation by transforming sets of attributes into meaningful sentences. Addressing the inherent data imbalance in PAR, we provide three types of prompt settings to optimize the training of the CLIP model. Our evaluation encompasses a thorough assessment of the proposed baseline model across various datasets, including PEARL dataset as well as estab-lished PAR benchmarks such as PA100K, RAP, and PETA.
Thakare Kamalakar Vijay, Lalit Lohani, Kamakshya Prasad Nayak, Debi Prosad Dogra, Heeseung Choi, Hyungjoo Jung, Ig-Jae Kim
WACV1
2024 Pedestrian Attribute Recognition Using Hierarchical Transformers
Lalit Lohani, Thakare Kamalakar Vijay, Kamakshya Prasad Nayak, Debi Prosad Dogra, Heeseung Choi, Hyungjoo Jung, Ig-Jae Kim
ICPR (16)2
2024 Let's Observe Them Over Time: An Improved Pedestrian Attribute Recognition Approach
abstract
Despite poor image quality, occlusions, and small training datasets, recent pedestrian attribute recognition (PAR) methods have achieved considerable performance. However, leveraging only spatial information of different attributes limits their reliability and generalizability. This paper introduces a multi-perspective approach to reduce over-dependence on spatial clues of a single perspective and exploits other aspects available in multiple perspectives. In order to tackle image quality and occlusions, we exploit different spatial clues present across images and handpick the best attribute-specific features to classify. Precisely, we extract the class-activation energy of each attribute and correlate it with the corresponding energy present across other images using the proposed Self-Attentive Cross Relation Module. In the next stage, we fuse this correlation information with similar clues accumulated from the other images. Lastly, we train a classification neural network using combined correlation information with two different losses. We have validated our method on four widely used PAR datasets, namely Market1501, PETA, PA-100k, and Duke. Our method achieves superior performance over most existing methods, demonstrating the effectiveness of a multi-perspective approach in PAR.
Thakare Kamalakar Vijay, Debi Prosad Dogra, Heeseung Choi, Haksub Kim, Ig-Jae Kim
WACV1
2023 DyAnNet: A Scene Dynamicity Guided Self-Trained Video Anomaly Detection Network
abstract
Unsupervised approaches for video anomaly detection may not perform as good as supervised approaches. However, learning unknown types of anomalies using an unsupervised approach is more practical than a supervised approach as annotation is an extra burden. In this paper, we use isolation tree-based unsupervised clustering to partition the deep feature space of the video segments. The RGB-stream generates a pseudo anomaly score and the flow stream generates a pseudo dynamicity score of a video segment. These scores are then fused using a majority voting scheme to generate preliminary bags of positive and negative segments. However, these bags may not be accurate as the scores are generated only using the current segment which does not represent the global behavior of a typical anomalous event. We then use a refinement strategy based on a cross-branch feed-forward network designed using a popular I3D network to refine both scores. The bags are then refined through a segment re-mapping strategy. The intuition of adding the dynamicity score of a segment with the anomaly score is to enhance the quality of the evidence. The method has been evaluated on three popular video anomaly datasets, i.e., UCF-Crime, CCTV-Fights, and UBI-Fights. Experimental results reveal that the proposed framework achieves competitive accuracy as compared to the state-of-the-art video anomaly detection methods.
Thakare Kamalakar Vijay, Yash Raghuwanshi, Debi Prosad Dogra, Heeseung Choi, Ig-Jae Kim
WACV1
2023 RareAnom: A Benchmark Video Dataset for Rare Type Anomalies
abstract
Existing video anomaly detection methods and datasets suffer from restricted anomaly categories containing single-source (CCTV) videos recorded in controlled environment, inadequate annotations, and lack of adequate supervision. To mitigate these problems, we introduce a new dataset ( RareAnom ) containing 17 rare types of real-world anomalies (2200 videos) recorded using multiple sources (e.g., CCTV , handheld cameras, dash-cams, and mobile phones) with rich temporal annotations. A new fully unsupervised anomaly detection and classification method has been proposed. It has three stages: training of a 3D Convolution Autoencoder using pseudo-labelled video segments, anomaly detection using latent features, and classification. Unlike the existing datasets, we have benchmarked RareAnom using three levels of supervision: fully, weakly, and unsupervised. It has been compared with UCF-Crime and XD-Violence datasets. The proposed anomaly detection and classification method beats the latest unsupervised methods by 4.49%, 8.66%, and 6.77% on RareAnom, UCF-Crime, and XD-violence datasets, respectively.
Thakare Kamalakar Vijay, Debi Prosad Dogra, Heeseung Choi, Haksub Kim, Ig-Jae Kim
Pattern Recognit.1
2023 Detection of Road Accidents Using Synthetically Generated Multi-Perspective Accident Videos
abstract
Road accidents are often caused by short abnormal events, including traffic violations, abrupt change in vehicular motion, driver fatigue, etc. Observing an accident event from the right camera perspective plays a crucial role while detecting accidents. However, it may not be possible to capture such abnormal events from a limited camera perspective. We present a deep learning framework to analyze the accident events recorded from multiple perspectives. First, we estimate feature similarity in videos recorded from multiple perspectives. We then divided the video samples into high and low feature similarity groups. Next, we extract spatio-temporal features from each group using two-branch DCNNs and fuse them using a rank-based weighted average pooling strategy followed by classification. We present a new road accident video dataset (MP-RAD), where each accident event is synthetically generated and captured from five independent camera perspectives using a computer gaming platform. Most of the existing road accident datasets use egocentric views or they are captured in fixed camera setups. However, our dataset is large and multi-perspective that can be used to validate ITS-related tasks such as accident detection, accident localization, traffic monitoring, etc. The dataset contains 400 accident events with a total of 2000 videos. We provide temporal annotations of all videos. The proposed framework and the dataset have been cross-validated with latest accident detection baselines trained on real-world road accident videos and vice-versa. The sub-optimal detection accuracy obtained using the baselines indicates that the proposed framework and the dataset can be useful for ITS related research. Code and dataset is available at: https://github.com/draxler1/MP-RAD-Dataset-ITS-
Thakare Kamalakar Vijay, Debi Prosad Dogra, Heeseung Choi, Gi Pyo Nam, Ig-Jae Kim
IEEE Trans. Intell. Transp. Syst.1
2022 A multi-stream deep neural network with late fuzzy fusion for real-world anomaly detection
Thakare Kamalakar Vijay, Nitin Sharma 0004, Debi Prosad Dogra, Heeseung Choi, Ig-Jae Kim
Expert Syst. Appl.1
2022 Object Interaction-Based Localization and Description of Road Accident Events Using Deep Learning
abstract
Detection and localization of road accidents in real-time is an integral part of the Intelligent Transportation System (ITS). Even though the existing road accident detection methods show promising results, the process suffers from some drawbacks. For example, existing methods require a large number of sample videos for feature learning. Moreover, features such as temporal gradients or flow fields are time-consuming. To address these issues, we introduce a new method that uses objects and their positions to detect accidents in real-time. Apart from localization of the accident events in videos, we perform a high-level post processing to describe the severity and context of an accident. Firstly, we divide an input video into pre-accident, accident and post-accident stages to extract object interactions. These interaction proposals are then filtered using a refinement algorithm. We then adopt an iterative training procedure to classify normal and accident interactions. We also highlight the damaged zone using heat maps. Finally, we generate high-level textual descriptions to quantify the context and severity of an accident. We have trained the proposed model using offline setups. However, it can be deployed online to detect road accident events in real-time by taking the video inputs directly from the CCTV camera. Moreover, with a minimal supervision, the model can be retrained for online surveillance. Extensive experiments carried out on UCF Crime and CADP datasets reveal that the proposed framework achieves state-of-the-art performance when compared with the recently proposed accident event detection methods in terms of AUC (UCF Crime: 69.70% and CADP: 72.59%) and FAR (UCF Crime: 0.8 and CADP: 2.2). The high-level description of the accident is an added advantage that will certainly help the traffic police to react in a timely manner.
Thakare Kamalakar Vijay, Debi Prosad Dogra, Heeseung Choi, Haksub Kim, Ig-Jae Kim
IEEE Trans. Intell. Transp. Syst.1
2021 PIDLNet: A Physics-Induced Deep Learning Network for Characterization of Crowd Videos
abstract
Human visual perception regarding crowd gatherings can provide valuable information about behavioral movements. Empirical analysis on visual perception about orderly moving crowds has revealed that such movements are often structured in nature with relatively higher order parameter and lower entropy as compared to unstructured crowd, and vice-versa. This paper proposes a Physics-Induced Deep Learning Network (PIDLNet), a deep learning framework trained on conventional 3D convolutional features combined with physics-based features. We have computed frame-level entropy and order parameter from the motion flows extracted from the crowd videos. These features are then integrated with the 3D convolutional features at a later stage in the feature extraction pipeline to aid in the crowd characterization process. Experiments reveal that the proposed network can characterize video segments depicting crowd movements with accuracy as high as 91.63%. We have obtained overall AUC of 0.9913 on highly challenging publicly available video dataset. The method outperforms existing deep-learning frameworks and conventional crowd characterization frameworks by a notable margin.
Shreetam Behera, Thakare Kamalakar Vijay, H. Manish Kausik, Debi Prosad Dogra
AVSS2