Chhavi Dhiman

dblp:232/7900 · DBLP profile ↗
← Back
22ranked-venue papers
5as first author
19since 2021 · last 2026
0000-0002-3401-596XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 8 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing memes analysis with multi-task multimodal M2CLIP-SEmoNet architecture
Chhavi Dhiman, Saihajpreet Singh, Satyam Mohan
Signal Process. Image Commun.1
2026 A Unified Approach for Multimodal Emotion Recognition Using Counterfactual Learning
Armaan Singh, Chhavi Dhiman
Signal Process. Image Commun.2
2025 GExSent: Gated Experts for Robust Sentiment Analysis Across Modalities
abstract
The field of multimodal sentiment analysis has gained huge attention through its broad applications in different forms of communication, from personal interactions to social media, rapid identification and understanding of public opinions and sentiments related to various social issues. Initially, traditional sentiment analysis relied on rule-based systems and predefined lexicons, which struggled to accurately interpret complex emotional cues. Recent advancements have improved the robustness and adaptability of sentiment analysis across diverse datasets. In this study, a novel architecture is proposed that combines the powerful encoding capabilities of CLIP and Modern BERT. Furthermore, it introduces two feature enhancement modules—Hierarchical Gated Mixture of Experts (H-GMoE) and Gated Attention Mechanism (GAM)—to support task-specific learning and enhance the relevance of multimodal features. Experimental results on benchmark datasets, including MOSI, MOSEI, and Memotion, demonstrate superior performance when compared to current state-of-the-art methods. The code for this work is available at https://github.com/Subhanshusethi/GExSENT-MOE.git
Subhanshu Sethi, Divij Saini, Sarthak Singh, Chhavi Dhiman
IJCNN4
2025 Unmasking Deception: A Comprehensive Survey on the Evolution of Face Anti-spoofing Methods
Aashania Antil, Chhavi Dhiman
Neurocomputing2
2025 Decoding fake news fabrications and trends: A comprehensive survey
Chhavi Dhiman
Neurocomputing2
2025 UnMA-CapSumT: Unified and Multi-Head attention-driven caption summarization transformer
Dhruv Sharma, Chhavi Dhiman, Dinesh Kumar 0001
J. Vis. Commun. Image Represent.2
2025 Predicting pedestrian intentions with multimodal IntentFormer: A Co-learning approach
Chhavi Dhiman, Indu Sreedevi
Pattern Recognit.2
2025 Cross-Modal Pedestrian Behavior Prediction: A Dual-Task Approach With Progressive Denoising Attention and CVAE
abstract
Pedestrian intention and trajectory prediction are crucial for advancing intelligent transportation systems and autonomous vehicles, significantly enhancing urban mobility’s safety and efficiency. Traditional approaches have evolved from capturing pedestrian dynamics through image features and bounding box coordinates to leveraging multiple modalities and attention mechanisms. However, challenges in robust cross-modal feature integration and adaptation to complex scenarios persist. This paper introduces a dual-task approach that simultaneously predicts short-term pedestrian crossing intentions and long-term trajectories by integrating features from pedestrian regions of interest (ROIs), scene attributes, and past trajectories. For crossing intention prediction, Progressive Denoising Attention (PDA) is developed, which iteratively refines cross-modal features to augment inter-class variations. Additionally, a three-phase counterfactual training approach is employed that manipulates pedestrian ROIs and segmentation maps to further enhance model robustness in complex scenarios. For trajectory prediction, a Conditional Variational Autoencoder (CVAE) is implemented, guided by contextual embeddings from the novel Context-Aware Feature Fusion Module (CAFFM) to significantly reduce mean squared error by integrating rich spatiotemporal ROI and context information. Experimental results on benchmark datasets JAAD and PIE demonstrate the superior performance of the proposed approach in understanding and predicting pedestrian intent. The code is available at: https://github.com/neha013/DPITRA
Chhavi Dhiman, Sreedevi Indu
IEEE Trans. Intell. Transp. Syst.2
2024 Securing Faces: A GAN-Powered Defense Against Spoofing with MSRCR and CBAM
Aashania Antil, Chhavi Dhiman
ICPR (13)2
2024 Fight detection with spatial and channel wise attention-based ConvLSTM model
abstract
Abstract An automated detection of aggressive and violent behaviour in videos has immense potential. It enables efficient online content filtering by restricting access to extreme content and also, when integrated with security systems, helps to monitor violence in surveillance videos. In this work, a convolutional neural network is combined with the proposed Spatial and Channel wise Attention‐based ConvLSTM encoder (SCan‐ConvLSTM). The proposed architecture performs an efficient spatiotemporal fusion of the features extracted from the video sequences containing fight scenes. In order to focus selectively on regions of utmost importance, this blended attention mechanism adjusts the weights of outputs in different locations and across different channels. This recurrent attention mechanism enhances the sequential refinement of activation maps and boosts the model performance. Finally, the experimental results have been presented that show the proposed architecture achieves superior results on the benchmark datasets (RWF‐2000, Violent‐flow, Hockey‐fights, and Movies).
Kunal Chaturvedi, Chhavi Dhiman, Dinesh Kumar Vishwakarma
Expert Syst. J. Knowl. Eng.2
2024 XGL-T transformer model for intelligent image captioning
Dhruv Sharma, Chhavi Dhiman, Dinesh Kumar 0001
Multim. Tools Appl.2
2024 AP-TransNet: a polarized transformer based aerial human action recognition framework
Chhavi Dhiman, Anunay Varshney, Ved Vyapak
Mach. Vis. Appl.1
2024 FDT - Dr2T: a unified Dense Radiology Report Generation Transformer framework for X-ray images
Dhruv Sharma, Chhavi Dhiman, Dinesh Kumar 0001
Mach. Vis. Appl.2
2024 MF2ShrT: Multimodal Feature Fusion Using Shared Layered Transformer for Face Anti-spoofing
abstract
In recent times, Face Anti-spoofing (FAS) has gained significant attention in both academic and industrial domains. Although various convolutional neural network (CNN)-based solutions have emerged, multimodal approaches incorporating RGB, depth, and information retrieval (IR) have exhibited better performance than unimodal classifiers. The increasing veracity of modern presentation attack instruments results in a persistent need to enhance the performance of such models. Recently, self-attention-based vision transformers (ViT) have become a popular choice in this field. Their fundamental aspects for multimodal FAS have not been thoroughly explored yet. Therefore, we propose a novel framework for FAS called MF 2 ShrT, which is based on a pretrained vision transformer. The proposed framework uses overlap patches and parameter sharing in the ViT network, allowing it to utilize multiple modalities in a computationally efficient manner. Furthermore, to effectively fuse intermediate features from different encoders of each ViT, we explore a T-encoder-based hybrid feature block enabling the system to identify correlations and dependencies across different modalities. MF 2 ShrT outperforms conventional vision transformers and achieves state-of-the-art performance on benchmarks CASIA-SURF and WMCA, demonstrating the efficiency of transformer-based models for presentation attack detection PAD).
Aashania Antil, Chhavi Dhiman
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Evolution of visual data captioning Methods, Datasets, and evaluation Metrics: A comprehensive survey
Dhruv Sharma, Chhavi Dhiman, Dinesh Kumar 0001
Expert Syst. Appl.2
2023 A two stream face anti-spoofing framework using multi-level deep features and ELBP features
Aashania Antil, Chhavi Dhiman
Multim. Syst.2
2022 A sparse coded composite descriptor for human activity recognition
abstract
Abstract This paper proposes a novel algorithm for computing discriminative descriptors named as a sparse coded composite descriptor (SCCD) for robust human activity recognition. The proposed method blends the state‐of‐the‐art handcrafted features and the discriminative nature of the sparse representation of visual information. The human activity is firstly modelled using any handcrafted feature, and then the sparse codes computed on a discriminative sparse dictionary of these features are embedded to provide discrimination in the feature set. Finally, a support vector machine (SVM) is trained using the proposed SCCDs to perform classification of different human activities. A new feature named as differential motion descriptor (DMD) is also proposed to extract the motion as well as spatial information from an activity video. The simulation results reveal that in comparison with the handcrafted feature, the corresponding SCCD improves the recognition accuracy significantly. The proposed method is compared with state‐of‐the‐art methods on KTH, Ballet, UCF50, and HMDB51 datasets and the proposed methodology of composite features outperforms these methods in terms of recognition accuracy.
Kuldeep Singh 0002, Chhavi Dhiman, Dinesh Kumar Vishwakarma, Himanshu Makhija, Gurjit Singh Walia
Expert Syst. J. Knowl. Eng.2
2022 Pedestrian Intention Prediction for Autonomous Vehicles: A Comprehensive Survey
Chhavi Dhiman, Indu Sreedevi
Neurocomputing2
2021 Part-wise Spatio-temporal Attention Driven CNN-based 3D Human Action Recognition
abstract
Recently, human activity recognition using skeleton data is increasing due to its ease of acquisition and finer shape details. Still, it suffers from a wide range of intra-class variation, inter-class similarity among the actions and view variation due to which extraction of discriminative spatial and temporal features is still a challenging problem. In this regard, we present a novel Residual Inception Attention Driven CNN (RIAC-Net) Network, which visualizes the dynamics of the action in a part-wise manner. The complete skeletonis partitioned into five key parts: Head to Spine, Left Leg, Right Leg, Left Hand, Right Hand. For each part, a Compact Action Skeleton Sequence (CASS) is defined. Part-wise skeleton-based motion dynamics highlights discriminative local features of the skeleton that helps to overcome the challenges of inter-class similarity and intra-class variation with improved recognition performance. The RIAC-Net architecture is inspired by the concept of inception-residual representation that unifies the Attention Driven Residues (ADR) with inception-based Spatio-Temporal Convolution Features (STCF) to learn efficient salient action features. An ablation study is also carried out to analyze the effect of ADR over simple residue-based action representation. The robustness of the proposed framework is evaluated by performing an extensive experiment on four challenging datasets: UT Kinect Action 3D, Florence 3D action, MSR Daily Action3D, and NTU RGB-D datasets, which consistently demonstrate the superiority of the proposed method over other state-of-the-art methods.
Chhavi Dhiman, Dinesh Kumar Vishwakarma, Paras Agarwal
ACM Trans. Multim. Comput. Commun. Appl.1
2020 View-Invariant Deep Architecture for Human Action Recognition Using Two-Stream Motion and Shape Temporal Dynamics
abstract
Human action Recognition for unknown views, is a challenging task. We propose a deep view-invariant human action recognition framework, which is a novel integration of two important action cues: motion and shape temporal dynamics (STD). The motion stream encapsulates the motion content of action as RGB Dynamic Images (RGB-DIs), which are generated by Approximate Rank Pooling (ARP) and processed by using finetuned InceptionV3 model. The STD stream learns long-term view-invariant shape dynamics of action using a sequence of LSTM and Bi-LSTM learning models. Human Pose Model (HPM) generates view-invariant features of structural similarity index matrix (SSIM) based key depth human pose frames. The final prediction of the action is made on the basis of three types of late fusion techniques i.e. maximum (max), average (avg) and multiply (mul), applied on individual stream scores. To validate the performance of the proposed novel framework, the experiments are performed using both cross-subject and cross-view validation schemes on three publically available benchmarks- NUCLA multi-view dataset, UWA3D-II Activity dataset and NTU RGB-D Activity dataset. Our algorithm outperforms existing state-of-the-arts significantly, which is measured in terms of recognition accuracy, receiver operating characteristic (ROC) curve and area under the curve (AUC).
Chhavi Dhiman, Dinesh Kumar Vishwakarma
IEEE Trans. Image Process.1
2019 A review of state-of-the-art techniques for abnormal human activity recognition
Chhavi Dhiman, Dinesh Kumar Vishwakarma
Eng. Appl. Artif. Intell.1
2019 A unified model for human activity recognition using spatial distribution of gradients and difference of Gaussian kernel
Dinesh Kumar Vishwakarma, Chhavi Dhiman
Vis. Comput.2