VLDB 2026 Research / reviewers in the wild / expert
Gnana Praveen Rajasekhar
dblp:158/9777 · also Gnana Praveen R., Praveen R. Gnana, R. Gnana Praveen
· DBLP profile ↗
12ranked-venue papers
9as first author
9since 2021 · last 2026
0000-0002-4698-9198ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 6 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Weakly Supervised Learning for Facial Affective Behavior Analysis: A ReviewabstractRecent advances in deep learning (DL) and computational capacity have enabled facial affective behavior analysis (FABA) to progress from static images captured in controlled settings to fine-grained analysis of facial expressions in real world video data. However, training accurate DL models for FABA typically requires large-scale, expert-annotated datasets, which are costly to obtain and inherently noisy due to the ambiguity of labeling subtle facial expressions and action units (AUs). To mitigate these challenges, weakly supervised learning (WSL) has emerged as a promising paradigm for training models with weak annotations. In this paper, we present a structured taxonomy of WSL scenarios for FABA, organized according to the type of weak annotation and the specific affective task. Building on this taxonomy, we provide a critical synthesis of representative WSL methods for both classification (expression and AU recognition) and regression (expression and AU intensity estimation) tasks, focusing on their core methodological ideas, strengths, and limitations. Furthermore, we systematically summarize the comparative performance of WSL approaches along with widely adopted experimental setups and evaluation proto cols. Our critical assessment identifies key challenges and future research directions, including the need for efficient adaptation of foundation models and for the development of robust, scalable FABA systems suitable for real-world applications. Gnana Praveen Rajasekhar, Patrick Cardinal, Eric Granger |
IEEE Trans. Affect. Comput. | 1 |
| 2025 | LAVViT: Latent Audio-Visual Vision Transformers for Speaker VerificationabstractRecently, Vision Transformers (ViTs) have shown remarkable success in various computer vision applications. In this work, we have explored the potential of ViTs, pre-trained on visual data, for audio-visual speaker verification. To cope with the challenges of large-scale training, we introduce the Latent Audio-Visual Vision Transformer (LAVViT) adapters, where we exploit the existing pre-trained models on visual data without fine-tuning their parameters and train only the parameters of LAVViT adapters. The LAVViT adapters are injected into every layer of the ViT architecture to effectively fuse the audio and visual modalities using a small set of latent tokens, forming an attention bottleneck, thereby reducing the quadratic computational cost of cross-attention across the modalities. The proposed approach has been evaluated on the Voxceleb1 dataset and shows promising performance using only a few trainable parameters. Code is available at https://github.com/praveena2j/LAVViT Gnana Praveen Rajasekhar, Jahangir Alam 0001 |
ICASSP | 1 |
| 2024 | Dynamic Cross Attention for Audio-Visual Person VerificationabstractAlthough person or identity verification has been predominantly explored using individual modalities such as face and voice, audio-visual fusion has recently shown immense potential to outperform unimodal approaches. Audio and visual modalities are often expected to pose strong complementary relationships, which plays a crucial role for effective audio-visual fusion. However, they may not always strongly complement each other, they may also exhibit weak complementary relationships, resulting in poor audio-visual feature representations. In this paper, we propose a Dynamic Cross Attention (DCA) model that can dynamically select the cross-attended or unattended features on the fly based on the strong or weak complementary relationships, respectively, across audio and visual modalities. In particular, a conditional gating layer is designed to evaluate the contribution of the cross-attention mechanism and choose cross-attended features only when they exhibit strong complementary relationships, otherwise unattended features. Extensive experiments are conducted on the Voxceleb1 dataset to demonstrate the robustness of the proposed model. Results indicate that the proposed model consistently improves the performance on multiple variants of cross-attention while outperforming the state-of-the-art methods. Code is available at https://github.com/praveena2j/DCAforPersonVerification Gnana Praveen Rajasekhar, Jahangir Alam 0001 |
FG | 1 |
| 2024 | Audio-Visual Person Verification Based on Recursive Fusion of Joint Cross-AttentionabstractPerson or identity verification has been recently gaining a lot of attention using audio-visual fusion as faces and voices share close associations with each other. Conventional approaches based on audio-visual fusion rely on score-level or early feature-level fusion techniques. Though existing approaches showed improvement over unimodal systems, the potential of audio-visual fusion for person verification is not fully exploited. In this paper, we have investigated the prospect of effectively capturing both intra- and inter-modal relationships across audio and visual modalities, which can play a crucial role in significantly improving the fusion performance over unimodal systems. In particular, we introduce a recursive fusion of a joint cross-attentional model, where a joint audio-visual feature representation is employed in the cross-attention framework in a recursive fashion to progressively refine the feature representations that can efficiently capture the intra- and inter-modal relationships. To further enhance the audio-visual feature representations, we have also explored BLSTMs to improve the temporal modeling of audio-visual feature representations. Extensive experiments are conducted on the Voxceleb1 dataset to evaluate the proposed model. Results indicate that the proposed model shows promising improvement in fusion performance by adeptly capturing the intra- and inter-modal relationships across audio and visual modalities. Code is available at https://github.com/praveena2j/RJCAforSpeakerVerification Gnana Praveen Rajasekhar, Jahangir Alam 0001 |
FG | 1 |
| 2024 | Cross-Attention is not always needed: Dynamic Cross-Attention for Audio-Visual Dimensional Emotion RecognitionabstractIn video-based emotion recognition, audio and visual modalities are often expected to have a complementary relationship, which is widely explored using cross-attention. However, they may also exhibit weak complementary relationships, resulting in poor representations of audio-visual features, thus degrading the performance of the system. To address this issue, we propose Dynamic Cross-Attention (DCA) that can dynamically select cross-attended or unattended features on the fly based on their strong or weak complementary relationships respectively. Specifically, a simple yet efficient gating layer is designed to evaluate the contribution of the cross-attention mechanism and choose cross-attended features only when they exhibit a strong complementary relationship, otherwise unattended features. We evaluate the performance of the proposed approach on the challenging RECOLA and Aff-Wild2 datasets. We also compare the proposed approach with other variants of cross-attention and show that the proposed model consistently improves the performance on both datasets. Gnana Praveen Rajasekhar, Jahangir Alam 0001 |
ICME | 1 |
| 2023 | Recursive Joint Attention for Audio-Visual Fusion in Regression Based Emotion RecognitionabstractIn video-based emotion recognition (ER), it is important to effectively leverage the complementary relationship among audio (A) and visual (V) modalities, while retaining the intramodal characteristics of individual modalities. In this paper, a recursive joint attention model is proposed along with long short-term memory (LSTM) modules for the fusion of vocal and facial expressions in regression-based ER. Specifically, we investigated the possibility of exploiting the complementary nature of A and V modalities using a joint cross-attention model in a recursive fashion with LSTMs to capture the intramodal temporal dependencies within the same modalities as well as among the A-V feature representations. By integrating LSTMs with recursive joint cross-attention, our model can efficiently leverage both intra- and inter-modal relationships for the fusion of A and V modalities. The results of extensive experiments1performed on the challenging Affwild2 and Fatigue (private) datasets indicate that the proposed A-V fusion model can significantly outperform state-of-art-methods. Gnana Praveen Rajasekhar, Eric Granger, Patrick Cardinal |
ICASSP | 1 |
| 2021 | Holistic Guidance for Occluded Person Re-Identification
Madhu Kiran, Gnana Praveen Rajasekhar, Le Thanh Nguyen-Meidine, Soufiane Belharbi, Louis-Antoine Blais-Morin, Eric Granger |
BMVC | 2 |
| 2021 | Cross Attentional Audio-Visual Fusion for Dimensional Emotion RecognitionabstractMultimodal analysis has recently drawn much interest in affective computing, since it can improve the overall accuracy of emotion recognition over isolated uni-modal approaches. The most effective techniques for multimodal emotion recognition efficiently leverage diverse and complimentary sources of information, such as facial, vocal, and physiological modalities, to provide comprehensive feature representations. In this paper, we focus on dimensional emotion recognition based on the fusion of facial and vocal modalities extracted from videos, where complex spatiotemporal relationships may be captured. Most of the existing fusion techniques rely on recurrent networks or conventional attention mechanisms that do not effectively leverage the complimentary nature of audiovisual (A-V) modalities. We introduce a cross-attentional fusion approach to extract the salient features across A - V modalities, allowing for accurate prediction of continuous values of valence and arousal. Our new cross-attentional A - V fusion model efficiently leverages the inter-modal relationships. In particular, it computes cross-attention weights to focus on the more contributive features across individual modalities, and thereby combine contributive feature representations, which are then fed to fully connected layers for the prediction of valence and arousal. The effectiveness of the proposed approach is validated experimentally on videos from the RECOLA and Fatigue (private) data-sets. Results indicate that our cross-attentional A - V fusion model is a cost-effective approach that outperforms state-of-the-art fusion approaches. Code is available: https://github.com/praveena2j/Cross-Attentional-AV-Fusion. Gnana Praveen Rajasekhar, Eric Granger, Patrick Cardinal |
FG | 1 |
| 2021 | Deep domain adaptation with ordinal regression for pain assessment using weakly-labeled videos
Gnana Praveen Rajasekhar, Eric Granger, Patrick Cardinal |
Image Vis. Comput. | 1 |
| 2020 | Deep Weakly Supervised Domain Adaptation for Pain Localization in VideosabstractAutomatic pain assessment has an important potential diagnostic value for populations that are incapable of articulating their pain experiences. As one of the dominating nonverbal channels for eliciting pain expression events, facial expressions has been widely investigated for estimating the pain intensity of individual. However, using state-of-the-art deep learning (DL) models in real-world pain estimation applications poses several challenges related to the subjective variations of facial expressions, operational capture conditions, and lack of representative training videos with labels. Given the cost of annotating intensity levels for every video frame, we propose a weakly-supervised domain adaptation (WSDA) technique that allows for training 3D CNNs for spatiotemporal pain intensity estimation using weakly labeled videos, where labels are provided on a periodic basis. In particular, WSDA integrates multiple instance learning into an adversarial deep domain adaptation framework to train an Inflated 3D-CNN (I3D) model such that it can accurately estimate pain intensities in the target operational domain. The training process relies on weak target loss, along with domain loss and source loss for domain adaptation of the I3D model. Experimental results obtained using labeled source domain RECOLA videos and weakly-labeled target domain UNBC-McMaster videos indicate that the proposed deep WSDA approach can achieve significantly higher level of sequence (bag)-level and frame (instance)-level pain localization accuracy than related state-of-the-art approaches. Gnana Praveen Rajasekhar, Eric Granger, Patrick Cardinal |
FG | 1 |
| 2015 | Compressed domain human action recognition in H.264/AVC video streams
Manu Tom, Venkatesh Babu Radhakrishnan, Gnana Praveen Rajasekhar |
Multim. Tools Appl. | 3 |
| 2014 | Super-pixel based crowd flow segmentation in H.264 compressed videosabstractIn this paper, we have proposed a simple yet robust novel approach for segmentation of high density crowd flows based on super-pixels in H.264 compressed videos. The collective representation of the motion vectors of the compressed video sequence is transformed to color map and super-pixel segmentation is performed at various scales for clustering the coherent motion vectors. The number of dynamically meaningful flow segments is determined by measuring the confidence score of the accumulated multi-scale super-pixel boundaries. The final crowd flow segmentation is obtained from the edges that are consistent across all the super-pixel resolutions. Hence, our major contribution involves obtaining the flow segmentation by clustering the motion vectors and determination of number of flow segments using only motion super-pixels without any prior assumption of the number of flow segments. The proposed approach was bench-marked on standard crowd flow dataset. Experiments demonstrated better accuracy and speedup for the proposed approach compared to the state-of-the-art methods. Sovan Biswas, Gnana Praveen Rajasekhar, Venkatesh Babu Radhakrishnan |
ICIP | 2 |