Abhinav Dhall

dblp:05/7591 · DBLP profile ↗
← Back
95ranked-venue papers
21as first author
46since 2021 · last 2026
0000-0002-2230-1440ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 57 · 7 first-author · 31 since 2021Artificial intelligence and machine learning · 43 · 6 first-author · 21 since 2021Human-computer interaction and ubiquitous computing · 22 · 12 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Security and privacy · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 Hierarchical Cross-Attention Transformer for Non-contact Multimodal Pain Classification using Remote Physiological Signals and Visual Features
Anup Kumar Gupta 0001, Puneet Gupta 0002, Abhinav Dhall
FG3
2026 VFace: A Training-Free Approach for Diffusion-Based Video Face Swapping
abstract
We present a training-free, plug-and-play method, namely VFace, for high-quality face swapping in videos. It can be seamlessly integrated with image-based face swapping approaches built on diffusion models. First, we introduce a Frequency Spectrum Attention Interpolation technique to facilitate generation and intact key identity characteristics. Second, we achieve Target Structure Guidance via plug-and-play attention injection to better align the structural features from the target frame to the generation. Third, we present a Flow-Guided Attention Temporal Smoothening mechanism that enforces spatiotemporal coherence without modifying the underlying diffusion model to reduce temporal inconsistencies typically encountered in frame-wise generation. Our method requires no additional training or video-specific fine-tuning. Extensive experiments show that our method significantly enhances temporal consistency and visual fidelity, offering a practical and modular solution for video-based face swapping. Our code is available at VFace.
Sanoojan Baliah, Yohan Abeysinghe, Rusiru Thushara, Khan Muhammad 0001, Abhinav Dhall, Karthik Nandakumar, Muhammad Haris Khan
WACV5
2026 BrandFusion: Aligning Image Generation with Brand Styles
abstract
While recent text-to-image models excel at generating realistic content, they struggle to capture the nuanced visual characteristics that define a brand’s distinctive style—such as lighting preferences, photography genres, color palettes, and compositional choices. This work introduces BrandFusion, a novel framework that automatically generates brand-aligned promotional images by decoupling brand style learning from image generation. Our approach consists of two components: a Brand-aware Vision-Language Model (BrandVLM) that predicts brand-relevant style characteristics and corresponding visual embeddings from marketer-provided contextual information, and a Brand-aware Diffusion Model (BrandDM) that generates images conditioned on these learned style representations. Unlike existing personalization methods that require separate finetuning for each brand, BrandFusion maintains scalability while preserving interpretability through textual style characteristics. Our method generalizes effectively to unseen brands by leveraging common industry sector-level visual patterns. Extensive evaluation demonstrates consistent improvements over existing approaches across multiple brand alignment metrics, with a 66.11% preference rate in human evaluation study. This work paves the way for AI-assisted on-brand content creation in marketing workflows.
Varun Khurana, Yaman Singla, Balaji Krishnamurthy, Abhinav Dhall
WACV5
2026 SynchroRaMa : Lip-Synchronized and Emotion-Aware Talking Face Generation via Multi-Modal Emotion Embedding
abstract
Audio-driven talking face generation has received growing interest, particularly for applications requiring expressive and natural human-avatar interaction. However, most existing emotion-aware methods rely on a single modality (either audio or image) for emotion embedding, limiting their ability to capture nuanced affective cues. Additionally, most methods condition on a single reference image, restricting the model’s ability to represent dynamic changes in actions or attributes across time. To address these issues, we introduce SynchroRaMa, a novel framework that integrates a multi-modal emotion embedding by combining emotional signals from text (via sentiment analysis) and audio (via speech-based emotion recognition and audio-derived valence-arousal features), enabling the generation of talking face videos with richer and more authentic emotional expressiveness and fidelity. To ensure natural head motion and accurate lip synchronization, SynchroRaMa includes an audio-to-motion (A2M) module that generates motion frames aligned with the input audio. Finally, SynchroRaMa incorporates scene descriptions generated by Large Language Model (LLM) as additional textual input, enabling it to capture dynamic actions and high-level semantic attributes. Conditioning the model on both visual and textual cues enhances temporal consistency and visual realism. Quantitative and qualitative experiments on benchmark datasets demonstrate that SynchroRaMa outperforms the state-of-the-art, achieving improvements in image quality, expression preservation, and motion realism. A user study further confirms that SynchroRaMa achieves higher subjective ratings than competing methods in overall naturalness, motion diversity, and video smoothness. Our project page is available at https://novicemm.github.io/synchrorama.
Phyo Thet Yee, Dimitris Kollias, Sudeepta Mishra, Abhinav Dhall
WACV4
2026 A Survey on Deep Learning for Group-Level Emotion Recognition
abstract
With the rapid advancement of artificial intelligence, group-level emotion recognition (GER) has emerged as an important domain in human behavior analysis. Early GER methods primarily relied on handcrafted features. However, the recent success of deep learning has shifted the focus toward neural network-based solution, enabling more effective exploitation of the rich visual and contextual cues in group images and videos. Unlike individual-level emotion recognition, GER must account for the diversity and dynamics of multiple individuals within varied social contexts. Over the past decade, numerous deep learning-based methods have been proposed, achieving substantial performance gains. This survey provides a comprehensive review of deep learning-centric review of GER, introducing a new taxonomy that spans representation learning, graph-based modeling, attention and transformer architectures, and multimodal fusion strategies. We summarize benchmark datasets, outline prevailing GER pipelines, and consolidate performance trends from recent state-of-the-art approaches. In addition, we discuss the integration of foundation models and large language model-guided multimodal reasoning into GER. Key challenges are identified, and potential research directions are proposed to support the development of robust, real-world GER systems. This work aims to serve as a pivotal reference for future research in this evolving field.
Xiaohua Huang 0003, Xiaopeng Hong, Qirong Mao, Wenming Zheng, Abhinav Dhall
IEEE Trans. Comput. Soc. Syst.5
2026 Generation and Detection of Sign Language Deepfakes: A Linguistic and Visual Analysis
abstract
This research explores the positive application of deepfake technology for upper body generation, specifically sign language for the D(d)eaf and hard of hearing (DHoH) community. Given the complexity of sign language and the scarcity of experts, the generated videos are vetted by a sign language expert for accuracy. We construct a reliable deepfake dataset, evaluating its technical and visual credibility using computer vision and natural language processing models. The dataset, consisting of over 1200 videos featuring both seen and unseen individuals to the generation model, is also used to detect deepfake videos targeting vulnerable individuals. Expert annotations confirm that the generated videos are comparable to real sign language content. Linguistic analysis, using textual similarity scores and interpreter evaluations, shows that the interpretation of generated videos is at least 90% similar to authentic sign language. Visual analysis demonstrates that convincingly realistic deepfakes can be produced, even for new subjects. Using a pose/style transfer model, we pay close attention to detail, ensuring hand movements are accurate and align with the driving video. We also apply machine learning algorithms to establish a baseline for deepfake detection on this dataset, contributing to the detection of fraudulent sign language videos.
Shahzeb Naeem, Muhammad Riyyan Khan, Usman Tariq, Abhinav Dhall, Carlos Ivan Colon, Hasan Al-Nashash
IEEE Trans. Comput. Soc. Syst.4
2025 Investigating the Viability of Employing Multi-modal Large Language Models in the Context of Audio Deepfake Detection
abstract
While Vision-Language Models (VLMs) and Multimodal Large Language Models (MLLMs) have shown strong generalisation in detecting image and video deepfakes, their use for audio deepfake detection remains largely unexplored. In this work, we aim to explore the potential of MLLMs for audio deepfake detection. Combining audio inputs with a range of text prompts as queries to find out the viability of MLLMs to learn robust representations across modalities for audio deepfake detection. Therefore, we attempt to explore text-aware and context-rich, question-answer based prompts with binary decisions. We hypothesise that such a feature-guided reasoning will help in facilitating deeper multimodal understanding and enable robust feature learning for audio deepfake detection. We evaluate the performance of two MLLMs, Qwen2-Audio-7B-Instruct and SALMONN, in two evaluation modes: (a) zero-shot and (b) fine-tuned. Our experiments demonstrate that combining audio with a multi-prompt approach could be a viable way forward for audio deepfake detection. Our experiments show that the models perform poorly without task-specific training and struggle to generalise to out-of-domain data. However, they achieve good performance on in-domain data with minimal supervision, indicating promising potential for audio deepfake detection.
Akanksha Chuchra, Shukesh Reddy, Sudeepta Mishra, Abhijit Das 0001, Abhinav Dhall
IJCB5
2025 LayLens: Improving Deepfake Understanding through Simplified Explanations
Abhijeet Narang, Liuyijia Su, Abhinav Dhall
ICMI4
2025 MRAC 2025: 3rd International Workshop on Multimodal, Generative and Responsible Affective Computing
abstract
Multimodal, generative, and responsible affective computing aims to enhance people's lives. In recent years, the AI revolution has already begun to impact daily life, with virtual assistants being deployed across various sectors such as healthcare, banking, transportation, and education. It is clear that, in the near future, humans may interact with AI-powered systems as much or maybe even more than direct human-to-human interactions. Affective computing has numerous applications, including innovative approaches to forecasting and preventing anxiety, stress, and mental health issues; enhancing robotic empathy; assisting individuals with communication, behavior, and emotion regulation challenges; and promoting awareness of health and well-being. Many of these applications require enhanced control and protection of sensitive, private, and personal data. Therefore, it is crucial to further develop the creation, evaluation, and deployment of emotionally intelligent systems that are both responsive and responsible. Additionally, improving the accuracy and interpretability of emotion prediction results can significantly enhance the application of this technology in the downstream tasks mentioned above. MRAC'25 is the continuation of MRAC'23 and MRAC'24. Through this workshop, we aim to bring together researchers to discuss the potential and development of affective computing.
Zheng Lian 0004, Shreya Ghosh 0001, Erik Cambria, Zhixi Cai, Guoying Zhao 0001, Abhinav Dhall, Björn W. Schuller, Roland Göcke, Jianhua Tao 0001, Tom Gedeon
ACM Multimedia6
2025 AV-Deepfake1M++: A Large-Scale Audio-Visual Deepfake Benchmark with Real-World Perturbations
abstract
The rapid surge of text-to-speech and face-voice reenactment models makes video fabrication easier and highly realistic. To encounter this problem, we require datasets that rich in type of generation methods and perturbation strategy which is usually common for online videos. To this end, we propose AV-Deepfake1M++, an extension of the AV-Deepfake1M having 2 million video clips with diversified manipulation strategy and audio-visual perturbation. This paper includes the description of data generation strategies along with benchmarking of AV-Deepfake1M++ using state-of-the-art methods. We believe that this dataset will play a pivotal role in facilitating research in Deepfake domain. Based on this dataset, we host the 2025 1M-Deepfakes Detection Challenge. The challenge details, dataset and evaluation scripts are available online under a research-only license at https://deepfakes1m.github.io/2025.
Zhixi Cai, Kartik Kuckreja, Shreya Ghosh 0001, Akanksha Chuchra, Muhammad Haris Khan, Usman Tariq, Tom Gedeon, Abhinav Dhall
ACM Multimedia8
2025 Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations
Shreya Ghosh 0001, Tom Gedeon, Thanh-Toan Do, Abhinav Dhall
ACM Multimedia5
2025 GTA-HDR: A Large-Scale Synthetic Dataset for HDR Image Reconstruction
abstract
High Dynamic Range (HDR) content (i.e., images and videos) has a broad range of applications. However, capturing HDR content from real-world scenes is expensive and time-consuming. Therefore, the challenging task of reconstructing visually accurate HDR images from their Low Dynamic Range (LDR) counterparts is gaining attention in the vision research community. A major challenge is the lack of datasets, which capture diverse scene conditions (e.g., lighting, weather, locations) and various image features (e.g., color, contrast, saturation). To address this gap, we introduce GTA-HDR, a large-scale synthetic dataset of photo-realistic HDR images sampled from the GTA-V video game. We perform thorough evaluation of the proposed dataset, which enables significant qualitative and quantitative improvements of the state-of-the-art HDR image reconstruction methods. Furthermore, we demonstrate the effectiveness of the proposed dataset and its impact on additional computer vision tasks including 3D human pose estimation, human body part segmentation, and holistic scene segmentation. The dataset, data collection pipeline, and evaluation code are available at: https://github.com/HrishavBakulBarua/GTA-HDR.
Hrishav Bakul Barua, Kalin Stefanov, Koksheik Wong, Abhinav Dhall, Ganesh Krishnasamy
WACV4
2025 MIP-GAF: A MLLM-Annotated Benchmark for Most Important Person Localization and Group Context Understanding
abstract
Estimating the Most Important Person (MIP) in any social event setup is a challenging problem mainly due to contextual complexity and scarcity of labeled data. Moreover, the causality aspects of MIP estimation are quite subjective and diverse. To this end, we aim to address the problem by annotating a large-scale ‘in-the-wild’ dataset for iden-tifying human perceptions about the ‘Most Important Person (MIP)‘ in an image. The paper provides a thorough description of our proposed Multimodal Large Language Model (MLLM) based data annotation strategy, and a thor-ough data quality analysis. Further, we perform a comprehensive benchmarking of the proposed dataset utilizing state-of-the-art MIP localization methods, indicating a significant drop in performance compared to existing datasets. The performance drop shows that the existing MIP localization algorithms must be more robust with respect to ‘in-the-wild’ situations. We believe the proposed dataset will play a vital role in building the next-generation social situation understanding methods. The dataset and associated code will be made available for research purposes.
Surbhi Madan, Shreya Ghosh 0001, Lownish Rai Sookha, M. A. Ganaie 0001, Subramanian Ramanathan, Abhinav Dhall, Tom Gedeon
WACV6
2025 Dual stage semantic information based generative adversarial network for image super-resolution
Shailza Sharma, Abhinav Dhall, Shikhar Johri, Vivek Singh Bawa
Comput. Vis. Image Underst.2
2025 Speech-aided facial video super resolution with accurate lip motion and enhanced frequency details
Shailza Sharma, Vivek Singh Bawa, Abhinav Dhall
Mach. Vis. Appl.3
2025 Multiview Attention Fusion for Explainable Body Language Behavior Recognition
abstract
Body language behavior, including gestures and fine-grained movements not only reflects human emotions, but also serves as a versatile cue for enhancing emotional intelligence and creating responsive technologies. In this work, we explore the efficacy ofmultiview-multimodal cuesforexplainable predictionof bodily behavior. This paper proposes an attention fusion method that combines features extracted from (1) multiview videos termed “RGB”, (2) their multiview Discrete Cosine Transform representations termed “DCT” and (3) three stream skeleton features termed “Skeleton”, via a transformer-based approach. We evaluate our approach on the diverse BBSI (Balazia et al., 2022) and Drive&Act (Martin et al., 2019) datasets. Empirical results confirm that the RGB, DCT and Skeleton features enable discovery of multiple class-specific behaviors resulting in explainable predictions. Our key findings are: (a) Multimodal approaches outperform unimodal counterparts in categorizing bodily behavioral classes; (b) Efficient class predictions and plausible explanations are achieved with both unimodal and multimodal approaches; and (c) Empirical results confirm the superiority of our approach compared to state-of-the-art methods on both datasets.
Surbhi Madan, Subramanian Ramanathan, Abhinav Dhall
IEEE Trans. Affect. Comput.4
2024 Conditional Distribution Modelling for Few-Shot Image Synthesis with Diffusion Models
Munawar Hayat, Abhinav Dhall, Thanh-Toan Do
ACCV (5)3
2024 INDIFACE: Illuminating India's Deepfake Landscape with a Comprehensive Synthetic Dataset
abstract
Due to the recent progress in Deepfake generation, several datasets and manipulation techniques have been proposed in the recent literature with various effective face-swap and face-reenactment methods. Deepfake is an emerging threat to society and government as it can jeopardize law enforcement and cause personal loss. Investigations in the literature established that demographic variation had impacted the performance of Deepfake detection. To date, Deepfake detection has not been studied in the Indian context; hence, in this work, we proposed a Deepfake dataset INDIFACE entirely with Indian subjects. We have collected 101 original videos and used two different manipulation techniques for Deepfake generation. We provide detailed benchmarking with state-of-the-art methods on Deepfake datasets, showcasing that the existing model is insufficient to detect Deepfake detection for the Indian scenario. Hence, more attention is required to this area of research. The proposed dataset INDIFACE is publicly available at.
Kartik Kuckreja, Ximi Hoque, Nishit Poddar, Shukesh Reddy, Abhinav Dhall, Abhijit Das 0001
FG5
2024 Real, Fake and Synthetic Faces - Does the Coin Have Three Sides?
abstract
With the ever-growing power of generative artificial intelligence, deepfake and artificially generated (synthetic) media have continued to spread online, which creates various ethical and moral concerns regarding their usage. To tackle this, we thus present a novel exploration of the trends and patterns observed in real, deepfake and synthetic facial images. The proposed analysis is done in two parts: firstly, we incorporate eight deep learning models and analyze their performances in distinguishing between the three classes of images. Next, we look to further delve into the similarities and differences between these three sets of images by investigating their image properties both in the context of the entire image as well as in the context of specific regions within the image. ANOVA test was also performed and provided further clarity amongst the patterns associated between the images of the three classes. From our findings, we observe that the investigated deep-learning models found it easier to detect synthetic facial images, with the ViT Patch-16 model performing best on this task with a class-averaged sensitivity, specificity, precision, and accuracy of 97.37%, 98.69%, 97.48%, and 98.25%, respectively. This observation was supported by further analysis of various image properties. We saw noticeable differences across the three category of images. This analysis can help us build better algorithms for facial image generation, and also shows that synthetic, deepfake and real face images are indeed three different classes.
Shahzeb Naeem, Ramzi Al-Sharawi, Muhammad Riyyan Khan, Usman Tariq, Abhinav Dhall, Hasan Al-Nashash
FG5
2024 A Spectro-Statistical Approach for Emotion Identification from EEG Signals
abstract
Automatic identification of emotions is important in human-centered computing. It allows machines to better understand user emotions. Identifying emotions via neural sensing techniques such as electroencephalogram (EEG) is a promising approach. In this paper, we aim to identify the emotions class from EEG signals. We frame emotion identification as a classification task and apply spectral and statistical encoders to extract the relevant features. We validate our approach on EmoNeuroDB dataset. Our method outperforms the EmoNeuroDB baseline, achieving a 42.10% increase in class prediction accuracy.
Lownish Rai Sookha, Gulshan Sharma, M. A. Ganaie 0001, Abhinav Dhall
FG4
2024 ClipSwap: Towards High Fidelity Face Swapping via Attributes and CLIP-Informed Loss
abstract
This paper introduces ClipSwap, a new frame-work designed for high-fidelity face swapping. Earlier methods for face swapping often struggle in identity transfer due to the mismatches in attributes between the target and source images. To handle this issue, an attributes-aware face swapping approach is proposed in our work. We use a conditional Generative Adversarial Network and a CLIP-based encoder, which extracts rich semantic knowledge to achieve attributes-aware face swapping. Our framework uses CLIP embedding in the face swapping process for improving the transmission of source image's identity details to the swapped image by refining the high-level semantic attributes obtained from the source image. And source image serves as the input reference image for CLIP and ensures a more accurate and detailed identity representation in the final result. Additionally, we apply Contrastive Loss to guide the transformation of source facial attributes onto the swapped image from various viewpoints. We also introduce Attributes Preservation Loss, which penalizes the network to keep the facial attributes of the target image. Thorough quantitative and qualitative evaluations on multiple datasets illustrate the high-quality swapping results. Our proposed ClipSwap outperforms prior state-of-the-art (SOTA) methods in face swapping, particularly in terms of identity transfer and facial attribute features.
Phyo Thet Yee, Sudeepta Mishra, Abhinav Dhall
FG3
2024 Histohdr-Net: Histogram Equalization for Single LDR to HDR Image Translation
abstract
High Dynamic Range (HDR) imaging aims to replicate the high visual quality and clarity of real-world scenes. Due to the high costs associated with HDR imaging, the literature offers various data-driven methods for HDR image reconstruction from Low Dynamic Range (LDR) counterparts. A common limitation of these approaches is missing details in regions of the reconstructed HDR images, which are overor under-exposed in the input LDR images. To this end, we propose a simple and effective method, HistoHDR-Net, to recover the fine details (e.g., color, contrast, saturation, and brightness) of HDR images via a fusion-based approach utilizing histogram-equalized LDR images along with self-attention guidance. Our experiments demonstrate the efficacy of the proposed approach over the state-of-art methods.
Hrishav Bakul Barua, Ganesh Krishnasamy, Koksheik Wong, Abhinav Dhall, Kalin Stefanov
ICIP4
2024 DREAMS: Diverse Reactions of Engagement and Attention Mind States Dataset
Monisha Singh, Gulshan Sharma, Ximi Hoque, Abhinav Dhall
ICPR (14)4
2024 AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake Dataset
abstract
The detection and localization of highly realistic deepfake audio-visual content are challenging even for the most advanced state-of-the-art methods. While most of the research efforts in this domain are focused on detecting high-quality deepfake images and videos, only a few works address the problem of the localization of small segments of audio-visual manipulations embedded in real videos. In this research, we emulate the process of such content generation and propose the AV-Deepfake1M dataset. The dataset contains content-driven (i) video manipulations, (ii) audio manipulations, and (iii) audio-visual manipulations for more than 2K subjects resulting in a total of more than 1M videos. The paper provides a thorough description of the proposed data generation pipeline accompanied by a rigorous analysis of the quality of the generated data. The comprehensive benchmark of the proposed dataset utilizing state-of-the-art deepfake detection and localization methods indicates a significant drop in performance compared to previous datasets. The proposed dataset will play a vital role in building the next-generation deepfake localization methods. The dataset and associated code are available at https://github.com/ControlNet/AV-Deepfake1M.
Zhixi Cai, Shreya Ghosh 0001, Aman Pankaj Adatia, Munawar Hayat, Abhinav Dhall, Tom Gedeon, Kalin Stefanov
ACM Multimedia5
2024 1M-Deepfakes Detection Challenge
abstract
The detection and localization of deepfake content, particularly when small fake segments are seamlessly mixed with real videos, remains a significant challenge in the field of digital media security. Based on the recently released AV-Deepfake1M dataset, which contains more than 1 million manipulated videos across more than 2,000 subjects, we introduce the 1M-Deepfakes Detection Challenge. This challenge is designed to engage the research community in developing advanced methods for detecting and localizing deepfake manipulations within the large-scale high-realistic audio-visual dataset. The participants can access the AV-Deepfake1M dataset and are required to submit their inference results for evaluation across the metrics for detection or localization tasks. The methodologies developed through the challenge will contribute to the development of next-generation deepfake detection and localization systems. Evaluation scripts, baseline models, and accompanying code will be available on https://github.com/ControlNet/AV-Deepfake1M.
Zhixi Cai, Abhinav Dhall, Shreya Ghosh 0001, Munawar Hayat, Dimitris Kollias, Kalin Stefanov, Usman Tariq
ACM Multimedia2
2024 Towards Engagement Prediction: A Cross-Modality Dual-Pipeline Approach using Visual and Audio Features
abstract
Engagement estimation is crucial for advancing natural human-computer interaction, allowing artificial agents to dynamically adjust their responses based on user engagement levels and creating more intuitive and immersive experiences. Despite advancements in automating real-time engagement estimation, challenges persist in real-world scenarios due to the complex nature of multi-modal human social signals. This paper proposes a novel cross-modality fusion-based methodology to address these challenges by leveraging multi-modal data. Our approach integrates visual and audio features, such as facial motion, acoustic characteristics, Contrastive Language-Image Pretraining (CLIP), and semantic embeddings. These features first pass through a transformer encoder, are then combined and processed through a cross-modal fusion mechanism, ensuring robust integration. The final integrated features are then used to predict engagement scores. This hierarchical and self-normalizing approach enhances the accuracy of engagement estimation by effectively capturing dependencies within and between modalities. The experiments are conducted on multimediate's NoXI and MPIIGroupInteraction datasets and the results demonstrates competitive performance in estimating engagement levels, addressing the complex, context-dependent nature of human engagement. Specifically, our approach achieves a Global Concordance Correlation Coefficient (CCC) score approximately (56.1%) higher than the baseline. This work contributes to developing more intelligent and responsive artificial systems, enhancing user experiences across various interactive applications.
Surbhi Madan, Abhinav Dhall, Balasubramanian Raman
ACM Multimedia4
2024 Automatic Gaze Analysis: A Survey of Deep Learning Based Approaches
abstract
Eye gaze analysis is an important research problem in the field of Computer Vision and Human-Computer Interaction. Even with notable progress in the last 10 years, automatic gaze analysis still remains challenging due to the uniqueness of eye appearance, eye-head interplay, occlusion, image quality, and illumination conditions. There are several open questions, including what are the important cues to interpret gaze direction in an unconstrained environment without prior knowledge and how to encode them in real-time. We review the progress across a range of gaze analysis tasks and applications to elucidate these fundamental questions, identify effective methods in gaze analysis, and provide possible future directions. We analyze recent gaze estimation and segmentation methods, especially in the unsupervised and weakly supervised domain, based on their advantages and reported evaluation metrics. Our analysis shows that the development of a robust and generic gaze analysis method still needs to address real-world challenges such as unconstrained setup and learning with less supervision. We conclude by discussing future research directions for designing a real-world gaze analysis system that can propagate to other domains including Computer Vision, Augmented Reality (AR), Virtual Reality (VR), and Human Computer Interaction (HCI).
Shreya Ghosh 0001, Abhinav Dhall, Munawar Hayat, Jarrod Knibbe
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 MARLIN: Masked Autoencoder for facial video Representation LearnINg
abstract
This paper proposes a self-supervised approach to learn universal facial representations from videos, that can transfer across a variety of facial analysis tasks such as Facial Attribute Recognition (FAR), Facial Expression Recognition (FER), DeepFake Detection (DFD), and Lip Synchronization (LS). Our proposed framework, named MARLIN, is a facial video masked autoencoder, that learns highly robust and generic facial embeddings from abundantly available non-annotated web crawled facial videos. As a challenging auxiliary task, MARLIN reconstructs the spatio-temporal details of the face from the densely masked facial regions which mainly include eyes, nose, mouth, lips, and skin to capture local and global aspects that in turn help in encoding generic and transferable features. Through a variety of experiments on diverse downstream tasks, we demonstrate MARLIN to be an excellent facial video encoder as well as feature extractor, that performs consistently well across a variety of downstream tasks including FAR (1.13% gain over supervised benchmark), FER (2.64% gain over unsupervised benchmark), DFD (1.86% gain over unsupervised benchmark), LS (29.36% gain for Frechet Inception Distance), and even in low data regime. Our code and models are available at https://github.com/ControlNet/MARLIN.
Zhixi Cai, Shreya Ghosh 0001, Kalin Stefanov, Abhinav Dhall, Jianfei Cai 0001, Seyed Hamid Rezatofighi, Gholamreza Haffari, Munawar Hayat
CVPR4
2023 BEAMER: Behavioral Encoder to Generate Multiple Appropriate Facial Reactions
abstract
This paper presents a framework for generating appropriate facial expressions for a listener engaged in a dyadic conversation. The ability to produce contextually suitable facial gestures in response to user interactions may enhance the user experience for avatars and social robots interaction. We propose a Transformer and Siamese architecture-based approach for generating appropriate facial expressions. Positive and negative Speaker-Listener pairs are created, applying a contrastive loss to facilitate learning. Furthermore, an ensemble of reconstruction quality sensitive loss functions is added to the network for learning discriminative features. The listener's facial reactions are represented with a combination of the 3D Morphable Model's coefficients and affect-related attributes (facial action units). The inputs to the network are pre-trained Transformer-based feature MARLIN and affect-related features. Experimental analysis demonstrate the effectiveness of the proposed method across various metrics in the form of an increase in performance compared to a variational auto-encoder-based baseline.
Ximi Hoque, Adamay Mann, Gulshan Sharma, Abhinav Dhall
ACM Multimedia4
2023 MAGIC-TBR: Multiview Attention Fusion for Transformer-based Bodily Behavior Recognition in Group Settings
abstract
Bodily behavioral language is an important social cue, and its automated analysis helps in enhancing the understanding of artificial intelligence systems. Furthermore, behavioral language cues are essential for active engagement in social agent-based user interactions. Despite the progress made in computer vision for tasks like head and body pose estimation, there is still a need to explore the detection of finer behaviors such as gesturing, grooming, or fumbling. This paper proposes a multiview attention fusion method named MAGIC-TBR that combines features extracted from videos and their corresponding Discrete Cosine Transform coefficients via a transformer-based approach. The experiments are conducted on the BBSI dataset and the results demonstrate the effectiveness of the proposed feature fusion with multiview attention. The code is available at: https://github.com/surbhimadan92/MAGIC-TBR
Surbhi Madan, Gulshan Sharma, Subramanian Ramanathan, Abhinav Dhall
ACM Multimedia5
2023 Glitch in the matrix: A large scale benchmark for content driven audio-visual forgery detection and localization
abstract
Most deepfake detection methods focus on detecting spatial and/or spatio-temporal changes in facial attributes and are centered around the binary classification task of detecting whether a video is real or fake. This is because available benchmark datasets contain mostly visual-only modifications present in the entirety of the video. However, a sophisticated deepfake may include small segments of audio or audio-visual manipulations that can completely change the meaning of the video content. To addresses this gap, we propose and benchmark a new dataset, Localized Audio Visual DeepFake (LAV-DF), consisting of strategic content-driven audio, visual and audio-visual manipulations. The proposed baseline method, Boundary Aware Temporal Forgery Detection (BA-TFD), is a 3D Convolutional Neural Network-based architecture which effectively captures multimodal manipulations. We further improve (i.e. BA-TFD+) the baseline method by replacing the backbone with a Multiscale Vision Transformer and guide the training process with contrastive, frame classification, boundary matching and multimodal boundary matching loss functions. The quantitative analysis demonstrates the superiority of BA-TFD+ on temporal forgery localization and deepfake detection tasks using several benchmark datasets including our newly proposed dataset. The dataset, models and code are available at https://github.com/ControlNet/LAV-DF.
Zhixi Cai, Shreya Ghosh 0001, Abhinav Dhall, Tom Gedeon, Kalin Stefanov, Munawar Hayat
Comput. Vis. Image Underst.3
2023 Audio-Visual Automatic Group Affect Analysis
abstract
Affective computing has progressed well due to methods, which can identify a person’s posed and spontaneous perceived affect with high accuracy. This paper focuses on group-level affect analysis on videos, which is one of the first few multimodal group-level affect analysis studies. There are many challenges on video-based group-level affect analysis as most of the work is focused on either a single person's affect recognition or image-based group affect analysis. To address this, first, we present an audio-visual perceived group affect dataset - ‘Video-level Group AFfect (VGAF)’. VGAF is a large-scale dataset consisting of 4,183 group videos. The videos are collected from YouTube with large variations in the keywords for collecting data across different genders, group settings, group sizes, illuminations and poses. The variety within the dataset will help the study of perception of group affect in a real environment. The data is manually annotated for three group affect classes - positive, neutral, and negative. Further, a fusion based audio-visual method is proposed to set a benchmark performance on the proposed dataset. The experimental results show the effectiveness of facial, holistic and speech features for group-level affect analysis. The baseline code, dataset, and pre-trained models are available at [LINK].
Abhinav Dhall, Jianfei Cai 0001
IEEE Trans. Affect. Comput.2
2022 'Labelling the Gaps': A Weakly Supervised Automatic Eye Gaze Estimation
Shreya Ghosh 0001, Abhinav Dhall, Munawar Hayat, Jarrod Knibbe
ACCV (4)2
2022 AV-GAZE: A Study on the Effectiveness of Audio Guided Visual Attention Estimation for Non-profilic Faces
abstract
In challenging real-life conditions such as extreme head-pose, occlusions, and low-resolution images where the visual information fails to estimate visual attention/gaze direction, audio signals could provide important and complementary information. In this paper, we explore if audio-guided coarse head-pose can further enhance visual attention estimation performance for non-prolific faces. Since it is difficult to annotate audio signals for estimating the head-pose of the speaker, we use off-the-shelf state-of-the-art models to facilitate cross-modal weak-supervision. During the training phase, the framework learns complementary information from synchronized audio-visual modality. Our model can utilize any of the available modalities i.e. audio, visual or audio-visual for task-specific inference. It is interesting to note that, when AV-Gaze is tested on benchmark datasets with these specific modalities, it achieves competitive results on multiple datasets, while being highly adaptive toward challenging scenarios.
Shreya Ghosh 0001, Abhinav Dhall, Munawar Hayat, Jarrod Knibbe
ICIP2
2022 Neural Encoding of Songs is Modulated by Their Enjoyment
abstract
We examine user and song identification from neural (EEG) signals. Owing to perceptual subjectivity in human-media interaction, music identification from brain signals is a challenging task. We demonstrate that subjective differences in music perception aid user identification, but hinder song identification. In an attempt to address intrinsic complexities in music identification, we provide empirical evidence on the role of enjoyment in song recognition. Our findings reveal that considering song enjoyment as an additional factor can improve EEG-based song recognition.
Gulshan Sharma, Pankaj Pandey, Subramanian Ramanathan, Krishna P. Miyapuram, Abhinav Dhall
ICMI5
2022 Sentiment-aware Classifier for Out-of-Context Caption Detection
abstract
In this work we propose additions to the COSMOS and COSMOS on Steroids pipelines for the detection of Cheapfakes for Task 1 of the ACM Grand Challenge for Detecting Cheapfakes. We compute sentiment features, namely polarity and subjectivity, using the news image captions. Multiple logistic regression results show that these sentiment features are significant in prediction of the outcome. We then combine the sentiment features with the four image-text features obtained in the aforementioned previous works to train an MLP. This classifies sets of inputs into being out-of-context (OOC) or not-out-of-context (NOOC). On a test set of 400 samples, the MLP with all features achieved a score of 87.25%, and that with only the image-text features a score of 88%. In addition to the challenge requirements, we also propose a separate pipeline to automatically construct caption pairs and annotations using the images and captions provided in the large, un-annotated training dataset. We hope that this endeavor will open the door for improvements, since hand-annotating cheapfake labels is time-consuming. To evaluate the performance on the test set, the Docker image with the models is available at: https://hub.docker.com/repository/docker/malkaddour/mmsys22cheapfakes. The open-source code for the project is accessible at: https://github.com/malkaddour/ACMM-22-Cheapfake-Detection-Sentiment-aware-Classifier-for-Out-of-Context-Caption-Detection.
Muhannad Alkaddour, Abhinav Dhall, Usman Tariq, Hasan Al-Nashash, Fares Al-Shargie
ACM Multimedia2
2022 A Transformer Based Approach for Activity Detection
abstract
Non-invasive physiological sensors allow for the collection of user-specific data in realistic environments. In this paper, using physiological data, we investigate the effectiveness of Convolutional Neural Network (CNN) based feature embeddings and Transformer architecture for the human activity recognition task. 1D-CNN representation is used for the heart rate, and 2D-CNN is used for short-term Fourier transformation of the accelerometer data. Post fusion, the feature is input into a transformer. The experiments are performed on the harAGE dataset. The findings indicate the discriminative ability of the feature-fusion on transformer-based architecture, and the method outperforms the harAGE baseline by an absolute 3.7%.
Gulshan Sharma, Abhinav Dhall, Subramanian Ramanathan
ACM Multimedia2
2022 Graph-based Group Modelling for Backchannel Detection
abstract
The brief responses given by listeners in group conversations are known as backchannels rendering the task of backchannel detection an essential facet of group interaction analysis. Most of the current backchannel detection studies explore various audio-visual cues for individuals. However, analysing all group members is of utmost importance for backchannel detection, like any group interaction. This study uses a graph neural network to model group interaction through all members' implicit and explicit behaviours. The proposed method achieves the best and second best performance on agreement estimation and backchannel detection tasks, respectively, of the 2022 MultiMediate: Multi-modal Group Behaviour Analysis for Artificial Mediation challenge.
Kalin Stefanov, Abhinav Dhall, Jianfei Cai 0001
ACM Multimedia3
2022 MTGLS: Multi-Task Gaze Estimation with Limited Supervision
abstract
Robust gaze estimation is a challenging task, even for deep CNNs, due to the non-availability of large-scale labeled data. Moreover, gaze annotation is a time-consuming process and requires specialized hardware setups. We propose MTGLS: a Multi-Task Gaze estimation framework with Limited Supervision, which leverages abundantly available non-annotated facial image data. MTGLS distills knowledge from off-the-shelf facial image analysis models, and learns strong feature representations of human eyes, guided by three complementary auxiliary signals: (a) the line of sight of the pupil (i.e. pseudo-gaze) defined by the localized facial landmarks, (b) the head-pose given by Euler angles, and (c) the orientation of the eye patch (left/right eye). To overcome inherent noise in the supervisory signals, MT-GLS further incorporates a noise distribution modelling approach. Our experimental results show that MTGLS learns highly generalized representations which consistently perform well on a range of datasets. Our proposed framework outperforms the unsupervised state-of-the-art on CAVE (by ∼ 6.43%) and even supervised state-of-the-art methods on Gaze360 (by ∼ 6.59%) datasets.
Shreya Ghosh 0001, Munawar Hayat, Abhinav Dhall, Jarrod Knibbe
WACV3
2022 Frequency aware face hallucination generative adversarial network with semantic structural constraint
Shailza Sharma, Abhinav Dhall
Comput. Vis. Image Underst.2
2022 Self-Supervised Approach for Facial Movement Based Optical Flow
abstract
Computing optical flow is a fundamental problem in computer vision. However, deep learning-based optical flow techniques do not perform well for non-rigid movements such as those found in faces, primarily due to lack of the training data representing the fine facial motion. We hypothesize that learning optical flow on face motion data will improve the quality of predicted flow on faces. This work aims to: (1) exploring self-supervised techniques to generate optical flow ground truth for face images; (2) computing baseline results on the effects of using face data to train Convolutional Neural Networks (CNN) for predicting optical flow; and (3) using the learned optical flow in micro-expression recognition to demonstrate its effectiveness. We generate optical flow ground truth using facial key-points in the BP4D-Spontaneous dataset. This optical flow is used to train the FlowNetS architecture to test its performance on the Extended Cohn-Kanade dataset and a portion of the generated dataset. The performance of FlowNetS trained on face images surpassed that of other optical flow CNN architectures. Our optical flow features are further compared with other methods using the STSTNet micro-expression classifier, and the results indicate that the optical flow obtained using this work has promising applications in facial expression analysis.
Muhannad Alkaddour, Usman Tariq, Abhinav Dhall
IEEE Trans. Affect. Comput.3
2022 Automatic Prediction of Group Cohesiveness in Images
abstract
This article discusses the prediction of cohesiveness of a group of people in images. The cohesiveness of a group is an essential indicator of the emotional state, structure, and success of the group. We study the factors that influence the perception of group-level cohesion and propose methods for estimating the human-perceived cohesion on the group cohesiveness scale. To identify the visual cues (attributes) for cohesion, we conducted a user survey. Image analysis is performed at a group-level via a multi-task convolutional neural network. A capsule network is explored for analyzing the contribution of facial expressions of the group members on predicting the Group Cohesion Score (GCS). We add GCS to the Group Affect database and propose the ‘GAF-Cohesion database’. The proposed model performs well on the database and achieves near human-level performance in predicting a group's cohesion score. It is interesting to note that group cohesion as an attribute, when jointly trained for group-level emotion prediction, helps in increasing the performance for the later task. This suggests that group-level emotion and cohesion are correlated. Further, we investigate the effect of face-level similarity, body pose and subset of a group on the task of automatic cohesion perception.
Shreya Ghosh 0001, Abhinav Dhall, Nicu Sebe, Tom Gedeon
IEEE Trans. Affect. Comput.2
2022 Analyzing Group-Level Emotion with Global Alignment Kernel based Approach
abstract
From the perspective of social science, understanding group emotion has become increasingly important for teams to considerably accomplish organizational work. Currently, automatically analyzing the perceived affect of a group of people has been received increasingly interest in affective computing community. The variability in group size makes difficulty for group-level emotion recognition to straightforwardly measure the feature distance of two group-level images. Recent works attempted to resolve the preceding problem by using feature encoding. However, the early works lack of efficiency. To alleviate this problem, this article aims to design a new method to effectively analyze the group behavior from a group-level image. Motivated by time-series kernel approaches explored in dynamic facial expression classification, this article mainly concentrates on global alignment kernel and design support vector machine with the combined global alignment kernels (SVM-CGAK) to better recognize group-level emotion. Specifically, we first propose to use global alignment kernel to explicitly measure the distance of two group-level images. For improving the performance of global alignment kernel, we use the global weight sort scheme based on their spatial relation information to sort the faces from group-level image, making an efficient data structure to the global alignment kernel. With this new global alignment kernel, we construct the backbone of SVM-CGAK, namely, support vector machine with global alignment kernel. Furthermore, considering the challenging environment, we construct two global alignment kernels based on Reisz-based Volume Local Binary Pattern and deep convolutional neural network features, respectively. Lastly, to make the robustness of group-level emotion recognition, we propose SVM-CGAK combining both global alignment kernels with multiple kernel learning approach. It can enhance the discriminative ability of each global alignment kernel. Intensive experiments are conducted on three challenging group-level emotion databases. The experimental results demonstrate that the proposed approach achieves promising performance for group-level emotion recognition compared with the recent state-of-the-art methods.
Xiaohua Huang 0003, Abhinav Dhall, Roland Göcke, Matti Pietikäinen, Guoying Zhao 0001
IEEE Trans. Affect. Comput.2
2021 ADGD'21: 1st Workshop on Synthetic Multimedia - Audiovisual Deepfake Generation and Detection
abstract
Deepfakes, i.e.synthetic or "fake" media content generated using deep learning, are a double-edged sword. On one hand, they pose new threats and risks in the form of scams, fraud, disinformation, social manipulation, or celebrity porn. On the other hand, deepfakes have just as many meaningful and beneficial applications - they allow us to create and experience things that no longer exist, or that have never existed, enabling numerous exciting applications in entertainment, education, and even privacy.
Stefan Winkler 0001, Abhinav Dhall, Pavel Korshunov
ACM Multimedia3
2021 Hyperrealistic Image Inpainting with Hypergraphs
abstract
Image inpainting is a non-trivial task in computer vision due to multiple possibilities for filling the missing data, which may be dependent on the global information of the image. Most of the existing approaches use the attention mechanism to learn the global context of the image. This attention mechanism produces semantically plausible but blurry results because of incapability to capture the global context. In this paper, we introduce hypergraph convolution on spatial features to learn the complex relationship among the data. We introduce a trainable mechanism to connect nodes using hyperedges for hypergraph convolution. To the best of our knowledge, hypergraph convolution have never been used on spatial features for any image-to-image tasks in computer vision. Further, we introduce gated convolution in the discriminator to enforce local consistency in the predicted image. The experiments on Places2, CelebA-HQ, Paris Street View, and Facades datasets, show that our approach achieves state-of-the-art results.
Gourav Wadhwa, Abhinav Dhall, M. Subrahmanyam 0001, Usman Tariq
WACV2
2021 Editorial for the special issue of IMAVIS on automatic face analytics for human behavior understanding
Xiaohua Huang 0003, Abhinav Dhall, Guoying Zhao 0001, Wenming Zheng, Matti Pietikäinen
Image Vis. Comput.2
2020 SSBC 2020: Sclera Segmentation Benchmarking Competition in the Mobile Environment
abstract
The paper presents a summary of the 2020 Sclera Segmentation Benchmarking Competition (SSBC), the 7th in the series of group benchmarking efforts centred around the problem of sclera segmentation. Different from previous editions, the goal of SSBC 2020 was to evaluate the performance of sclera-segmentation models on images captured with mobile devices. The competition was used as a platform to assess the sensitivity of existing models to i) differences in mobile devices used for image capture and ii) changes in the ambient acquisition conditions. 26 research groups registered for SSBC 2020, out of which 13 took part in the final round and submitted a total of 16 segmentation models for scoring. These included a wide variety of deep-learning solutions as well as one approach based on standard image processing techniques. Experiments were conducted with three recent datasets. Most of the segmentation models achieved relatively consistent performance across images captured with different mobile devices (with slight differences across devices), but struggled most with low-quality images captured in challenging ambient conditions, i.e., in an indoor environment and with poor lighting.
Matej Vitek, Abhijit Das 0001, Yann Pourcenoux, Alexandre Missler, C. Paumier, Sumanta Das, Ishita De Ghosh, Diego Rafael Lucio, Luiz Antonio Zanlorensi, David Menotti, Fadi Boutros, Naser Damer, Jonas Henry Grebe, Arjan Kuijper, Junxing Hu, Yong He 0009, Caiyong Wang, Yunlong Wang 0003, Zhenan Sun, Dailé Osorio Roig, Christian Rathgeb, Christoph Busch 0001, Juan E. Tapia, Andres Valenzuela, Georgios Zampoukis, Lazaros T. Tsochatzidis, Ioannis Pratikakis, Sabari Nathan, R. Suganya 0001, Vineet Mehta, Abhinav Dhall, Kiran B. Raja, Gourav Gupta, Jalil Nourmohammadi-Khiarak, Mohsen Akbari-Shahper, Farhang Jaryani, Meysam Asgari-Chenaghlu, Ritesh Vyas, Sristi Dakshit, Peter Peer, Umapada Pal 0001, Vitomir Struc
IJCB32
2020 Depth Estimation From Single Image And Semantic Prior
abstract
The multi-modality sensor fusion technique is an active research area in scene understating. In this work, we explore the RGB image and semantic-map fusion methods for depth estimation. The LiDARs, Kinect, and TOF depth sensors are unable to predict the depth-map at illuminate and monotonous pattern surface. In this paper, we propose a semantic-to-depth generative adversarial network (S2D-GAN) for depth estimation from RGB image and its semantic-map. In the first stage, the proposed S2D-GAN estimates the coarse level depthmap using a semantic-to-coarse-depth generative adversarial network (S2CD-GAN) while the second stage estimates the fine-level depth-map using a cascaded multi-scale spatial pooling network. The experimental analysis of the proposed S2D-GAN performed on NYU-Depth-V2 dataset shows that the proposed S2D-GAN gives outstanding result over existing single image depth estimation and RGB with sparse samples methods. The proposed S2D-GAN also gives efficient results on the real-world indoor and outdoor image depth estimation.
Praful Hambarde, Akshay Dudhane, Prashant W. Patil, M. Subrahmanyam 0001, Abhinav Dhall
ICIP5
2020 EmotiW 2020: Driver Gaze, Group Emotion, Student Engagement and Physiological Signal based Challenges
abstract
This paper introduces the Eighth Emotion Recognition in the Wild (EmotiW) challenge. EmotiW is a benchmarking effort run as a grand challenge of the 22nd ACM International Conference on Multimodal Interaction 2020. It comprises of four tasks related to automatic human behavior analysis: a) driver gaze prediction; b) audio-visual group-level emotion recognition; c) engagement prediction in the wild; and d) physiological signal based emotion recognition. The motivation of EmotiW is to bring researchers in affective computing, computer vision, speech processing and machine learning to a common platform for evaluating techniques on a test data. We discuss the challenge protocols, databases and their associated baselines.
Abhinav Dhall, Roland Göcke, Tom Gedeon
ICMI1
2020 The eyes know it: FakeET- An Eye-tracking Database to Understand Deepfake Perception
abstract
We present FakeET -- an eye-tracking database to understand human visual perception of deepfake videos. Given that the principal purpose of deepfakes is to deceive human observers, FakeET is designed to understand and evaluate the ability of viewers to detect synthetic video artifacts. FakeET contains viewing patterns compiled from 40 users via the Tobii desktop eye-tracker for 811 videos from the Google Deepfake dataset, with a minimum of two viewings per video. Additionally, EEG responses acquired via the Emotiv sensor are also available. The compiled data confirms (a) distinct eye movement characteristics for real vs fake videos; (b) utility of the eye-track saliency maps for spatial forgery localization and detection, and (c) Error Related Negativity (ERN) triggers in the EEG responses, and the ability of the raw EEG signal to distinguish between real and fake videos.
Komal Chugh, Abhinav Dhall, Subramanian Ramanathan
ICMI3
2020 Motion and Region Aware Adversarial Learning for Fall Detection with Thermal Imaging
abstract
Automatic fall detection is a vital technology for ensuring the health and safety of people. Home-based camera systems for fall detection often put people's privacy at risk. Thermal cameras can partially or fully obfuscate facial features, thus preserving the privacy of a person. Another challenge is the less occurrence of falls in comparison to the normal activities of daily living. As fall occurs rarely, it is non-trivial to learn algorithms due to class imbalance. To handle these problems, we formulate fall detection as an anomaly detection within an adversarial framework using thermal imaging. We present a novel adversarial network that comprises of two-channel 3D convolutional autoencoders which reconstructs the thermal data and the optical flow input sequences respectively. We introduce a technique to track the region of interest, a region-based difference constraint, and a joint discriminator to compute the reconstruction error. A larger reconstruction error indicates the occurrence of a fall. The experiments on a publicly available thermal fall dataset show the superior results obtained compared to the standard baseline.
Vineet Mehta, Abhinav Dhall, Sujata Pal, Shehroz S. Khan
ICPR2
2020 Not made for each other- Audio-Visual Dissonance-based Deepfake Detection and Localization
abstract
We propose detection of deepfake videos based on the dissimilarity between the audio and visual modalities, termed as the Modality Dissonance Score (MDS). We hypothesize that manipulation of either modality will lead to dis-harmony between the two modalities, e.g., loss of lip-sync, unnatural facial and lip movements, etc. MDS is computed as the mean aggregate of dissimilarity scores between audio and visual segments in a video. Discriminative features are learnt for the audio and visual channels in a chunk-wise manner, employing the cross-entropy loss for individual modalities, and a contrastive loss that models inter-modality similarity. Extensive experiments on the DFDC and DeepFake-TIMIT Datasets show that our approach outperforms the state-of-the-art by up to 7%. We also demonstrate temporal forgery localization, and show how our technique identifies the manipulated video segments.
Komal Chugh, Abhinav Dhall, Subramanian Ramanathan
ACM Multimedia3
2020 Large Scale Hierarchical Anomaly Detection and Temporal Localization
abstract
Abnormal event detection is a non-trivial task in machine learning. The primary reason behind this is that the abnormal class occurs sparsely, and its temporal location may not be available. In this paper, we propose a multiple feature-based approach for CitySCENE challenge-based anomaly detection. For motion and context information, Res3D and Res101 architectures are used. Object-level information is extracted by object detection feature-based pooling. Fusion of three channels above gives relatively high performance on the challenge Test set for the general anomaly task. We also show how our method can be used for temporal localisation of the abnormal activity event in a video.
Soumil Kanwal, Vineet Mehta, Abhinav Dhall
ACM Multimedia3
2019 Expression Empowered ResiDen Network for Facial Action Unit Detection
abstract
The paper explores the topic of Facial Action Unit (FAU) detection in the wild. In particular, we are interested in answering the following questions: (1) How useful are residual connections across dense blocks for face analysis? (2) How useful is the information from a network trained for categorical Facial Expression Recognition (FER) for the task of FAU detection? The proposed network (ResiDen) exploits dense blocks along with residual connections and uses auxiliary information from a FER network. The experiments are performed on the EmotionNet and DISFA datasets. The experiments show the usefulness of facial expression information for AU detection. The proposed network achieves state-of-the-art results on the two datasets. Analysis of the results for cross dataset protocol shows the effectiveness of the network.
Shreyank Jyoti, Abhinav Dhall
FG3
2019 Domain Adaptation based Topic Modeling Techniques for Engagement Estimation in the Wild
abstract
In recent years, student engagement estimation has gained focus in the affective computing community. The absence of student monitoring during online MOOC courses makes it challenging to estimate behavioural student engagement during online classes. The non availability of consistent engagement datasets makes it difficult to build cross data automatic behavioural engagement estimation technique. In this paper, we propose an unsupervised topic modeling technique for engagement detection as it captures multiple behavioral cues which are indicators of engagement level such as eye gaze, head movement, facial expression and body posture. We have addressed the various challenges such as less volume of our datasets, large decision unit (annotated for 5 minutes duration) and uneven distribution of different engagement categories with domain adaptation based solution for cross data implementation. We present results on engagement prediction using different clustering techniques such as K-Means and Latent Dirichlet Allocation (LDA) along with different regressors and neural network based attention mechanisms.
Amanjot Kaur, Bishal Ghosh, Naman D. Singh, Abhinav Dhall
FG4
2019 EmotiW 2019: Automatic Emotion, Engagement and Cohesion Prediction Tasks
abstract
This paper describes the Seventh Emotion Recognition in the Wild (EmotiW) Challenge. The EmotiW benchmarking platform provides researchers with an opportunity to evaluate their methods on affect labelled data. This year EmotiW 2019 encompasses three sub-challenges: a) Group-level cohesion prediction; b) Audio-Video emotion recognition; and c) Student engagement prediction. We discuss the databases used, the experimental protocols and the baselines.
Abhinav Dhall
ICMI1
2019 DIF : Dataset of Perceived Intoxicated Faces for Drunk Person Identification
abstract
Traffic accidents cause over a million deaths every year, of which a large fraction is attributed to drunk driving. An automated intoxicated driver detection system in vehicles will be useful in reducing accidents and related financial costs. Existing solutions require special equipment such as electrocardiogram, infrared cameras or breathalyzers. In this work, we propose a new dataset called DIF (Dataset of perceived Intoxicated Faces) which contains audio-visual data of intoxicated and sober people obtained from online sources. To the best of our knowledge, this is the first work for automatic bimodal non-invasive intoxication detection. Convolutional Neural Networks (CNN) and Deep Neural Networks (DNN) are trained for computing the video and audio baselines, respectively. 3D CNN is used to exploit the Spatio-temporal changes in the video. A simple variation of the traditional 3D convolution block is proposed based on inducing non-linearity between the spatial and temporal channels. Extensive experiments are performed to validate the approach and baselines.
Vineet Mehta, Sai Srinadhu Katta, Devendra Pratap Yadav, Abhinav Dhall
ICMI4
2019 Unsupervised Learning of Eye Gaze Representation from the Web
abstract
Automatic eye gaze estimation has interested researchers for a while now. In this paper, we propose an unsupervised learning based method for estimating the eye gaze region. To train the proposed network "Ize-Net" in self-supervised manner, we collect a large `in the wild' dataset containing 1,54,251 images from the web. For the images in the database, we divide the gaze into three regions based on an automatic technique based on pupil-centers localization and then use a feature-based technique to determine the gaze region. The performance is evaluated on the Tablet Gaze and CAVE datasets by fine-tuning results of Ize-Net for the task of eye gaze estimation. The feature representation learned is also used to train traditional machine learning algorithms for eye gaze estimation. The results demonstrate that the proposed method learns a rich data representation, which can be efficiently finetuned for any eye gaze estimation dataset.
Neeru Dubey, Shreya Ghosh 0001, Abhinav Dhall
IJCNN3
2019 Predicting Group Cohesiveness in Images
abstract
The cohesiveness of a group is an essential indicator of the emotional state, structure and success of a group of people. We study the factors that influence the perception of group-level cohesion and propose methods for estimating the human-perceived cohesion on the group cohesiveness scale. In order to identify the visual cues (attributes) for cohesion, we conducted a user survey. Image analysis is performed at a group-level via a multi-task convolutional neural network. For analyzing the contribution of facial expressions of the group members for predicting the Group Cohesion Score (GCS), a capsule network is explored. We add GCS to the Group Affect database and propose the `GAF-Cohesion database'. The proposed model performs well on the database and is able to achieve near human-level performance in predicting a group's cohesion score. It is interesting to note that group cohesion as an attribute, when jointly trained for group-level emotion prediction, helps in increasing the performance for the later task. This suggests that group-level emotion and cohesion are correlated.
Shreya Ghosh 0001, Abhinav Dhall, Nicu Sebe, Tom Gedeon
IJCNN2
2019 Automatic Speech-Gesture Mapping and Engagement Evaluation in Human Robot Interaction
abstract
In this paper, we present an end-to-end system for enhancing the effectiveness of non-verbal gestures in human robot interaction. We identify prominently used gestures in performances by TED talk speakers and map them to their corresponding speech context and modulated speech based upon the attention of the listener. Gestures are localised with convolution neural networks based approach. Dominant gestures of TED speakers are used for learning the gesture-to-speech mapping. We evaluated the engagement of the robot with people by conducting a social survey. The effectiveness of the performance was monitored by the robot and it self-improvised its speech pattern on the basis of the attention level of the audience, which was calculated using visual feedback from the camera. The effectiveness of interaction as well as the decisions made during improvisation was further evaluated based on the head-pose detection and an interaction survey.
Bishal Ghosh, Abhinav Dhall, Ekta Singla
RO-MAN2
2018 Multi-level Dense Capsule Networks
Sai Samarth R. Phaye, Apoorva Sikka, Abhinav Dhall, Deepti R. Bathula
ACCV (5)3
2018 Fast Face and Saliency Aware Collage Creation for Mobile Phones
abstract
This demonstration is of an automated method for creating image collage in a mobile phone. The app creates semantically meaningful collages from images based on faces, image saliency and hybrid blending. The algorithm is designed for computational efficiency inorder for it run on a mobile device. Due to the increase in the use of social networks, users, are uploading images and videos from social events. Collage presents a useful crisper summary of the event. It is popular among users as evident from Layout from Instagram: Collage app, which has over 100 million downloads. A limitation of these apps is that they are dependent on the user for selecting the grid layout and optimal cropping of the images. Our smart collage over comes these limitations and to the best of our knowledge this is one of the first mobile based solutions, which considers the faces and the region importance and later merges the salient images automatically in a non-rigid grid layout. Other interesting works either have a semi-rigid [1] or a fully rigid grid based layout [2] or are desktop based [3].
Love Mehta, Abhinav Dhall
FG2
2018 Hybrid Neural Networks Based Approach for Holoscopic Micro-Gesture Recognition in Images and Videos
abstract
This paper presents an approach for hand based micro-gesture recognition in images and videos as part of the Holoscopic Micro-Gesture Recognition (HoMGR) challenge. The database consists of Holoscopic 3D Micro-Gesture images and videos. The proposed framework is an ensemble of convolutional neural network and deep neural network. The framework performs feature fusion technique on both handcrafted (local phase quantization) and deep features extracted from the neural network, to leverage on complimentary information. The powerful discriminative nature of the fused features has proved beneficial on the given HoMGR challenge data. The experiments show that the proposed approach is effective and outperforms the baseline on the Test set by an absolute margin of 26.67% for images and 2.47% for videos, respectively.
Shreyank Jyoti, Abhinav Dhall
FG3
2018 Automatic Group Affect Analysis in Images via Visual Attribute and Feature Networks
abstract
This paper proposes a pipeline for automatic group-level affect analysis. A deep neural network-based approach, which leverages on the facial-expression information, scene information and a high-level facial visual attribute information is proposed. A capsule network-based architecture is used to predict the facial expression. Transfer learning is used on Inception-V3 to extract global image-based features which contain scene information. Another network is trained for inferring the facial attributes of the group members. Further, these attributes are pooled at a group-level to train a network for inferring the group-level affect. The facial attribute prediction network, although is simple yet, is effective and generates result comparable to the state-of-the-art methods. Later, model integration is performed from the three channels. The experiments show the effectiveness of the proposed techniques on three `in the wild' databases: Group Affect Database, HAPPEI and UCLA-Protest database.
Shreya Ghosh 0001, Abhinav Dhall, Nicu Sebe
ICIP2
2018 EmotiW 2018: Audio-Video, Student Engagement and Group-Level Affect Prediction
abstract
This paper details the sixth Emotion Recognition in the Wild (EmotiW) challenge. EmotiW 2018 is a grand challenge in the ACM International Conference on Multimodal Interaction 2018, Colarado, USA. The challenge aims at providing a common platform to researchers working in the affective computing community to benchmark their algorithms on 'in the wild' data. This year EmotiW contains three sub-challenges: a) Audio-video based emotion recognition; b) Student engagement prediction; and c) Group-level emotion recognition. The databases, protocols and baselines are discussed in detail.
Abhinav Dhall, Amanjot Kaur, Roland Göcke, Tom Gedeon
ICMI1
2018 Automatic Eye Gaze Estimation using Geometric & Texture-based Networks
abstract
Eye gaze estimation is an important problem in automatic human behavior understanding. This paper proposes a deep learning based method for inferring the eye gaze direction. The method is based on the use of ensemble of networks, which capture both the geometric and texture information. Firstly, a Deep Neural Network (DNN) is trained using the geometric features that are extracted from the facial landmark locations. Secondly, for the texture based features, three Convolutional Neural Networks (CNN) are trained i.e. for the patch around the left eye, right eye, and the combined eyes, respectively. Finally, the information from the four channels is fused with concatenation and dense layers are trained to predict the final eye gaze. The experiments are performed on the two publicly available datasets: Columbia eye gaze and TabletGaze. The extensive evaluation shows the superior performance of the proposed framework. We also evaluate the performance of the recently proposed swish activation function as compared to Rectified Linear Unit (ReLU) for eye gaze estimation.
Shreyank Jyoti, Abhinav Dhall
ICPR2
2018 MsEDNet: Multi-Scale Deep Saliency Learning for Moving Object Detection
abstract
Moving object detection (foreground and background) is an important problem in computer vision. Most of the works in this problem are based on background subtraction. However, these approaches are not able to handle scenarios with infrequent motion of object, illumination changes, shadow, camouflage etc. To overcome these, here a two stage robust and compact method for moving object detection (MOD) is proposed. In first stage, to generate the saliency map, background image is estimated using a temporal histogram technique with the help of several input frames. In the second stage, multiscale encoder-decoder network is used to learn multiscale semantic feature of estimated saliency for foreground extraction. The encoder is used to extract multi-scale features from multi-scale saliency map. The decoder part is designed to learn the mapping of low resolution multi-scale features into high resolution output frame. To observe the efficacy of proposed MsEDNet, experiments are conducted on two benchmark datasets (change detection (CDnet-2014) [1] and Wallflower [2]) for MOD. The precision, recall and F-measure are used as performance parameter for comparison with the existing state-of-the-art methods. Experimental results show a significant improvement in detection accuracy and decrement in execution time as compared to the state-of-the-art methods for MOD.
Prashant W. Patil, M. Subrahmanyam 0001, Abhinav Dhall, Sachin Chaudhary
SMC3
2018 Multimodal Framework for Analyzing the Affect of a Group of People
abstract
With the advances in multimedia and the world wide web, users upload millions of images and videos everyone on social networking platforms on the Internet. From the perspective of automatic human behavior understanding, it is of interest to analyze and model the affects that are exhibited by groups of people who are participating in social events in these images. However, the analysis of the affect that is expressed by multiple people is challenging due to the varied indoor and outdoor settings. Recently, a few interesting works have investigated face-based group-level emotion recognition (GER). In this paper, we propose a multimodal framework for enhancing the affective analysis ability of GER in challenging environments. Specifically, for encoding a person's information in a group-level image, we first propose an information aggregation method for generating feature descriptions of face, upper body, and scene. Later, we revisit localized multiple kernel learning for fusing face, upper body, and scene information for GER against challenging environments. Intensive experiments are performed on two challenging group-level emotion databases (HAPPEI and GAFF) to investigate the roles of the face, upper body, scene information, and the multimodal framework. Experimental results demonstrate that the multimodal framework achieves promising performance for GER.
Xiaohua Huang 0003, Abhinav Dhall, Roland Göcke, Matti Pietikäinen, Guoying Zhao 0001
IEEE Trans. Multim.2
2017 From individual to group-level emotion recognition: EmotiW 5.0
abstract
Research in automatic affect recognition has come a long way. This paper describes the fifth Emotion Recognition in the Wild (EmotiW) challenge 2017. EmotiW aims at providing a common benchmarking platform for researchers working on different aspects of affective computing. This year there are two sub-challenges: a) Audio-video emotion recognition and b) group-level emotion recognition. These challenges are based on the acted facial expressions in the wild and group affect databases, respectively. The particular focus of the challenge is to evaluate method in `in the wild' settings. `In the wild' here is used to describe the various environments represented in the images and videos, which represent real-world (not lab like) scenarios. The baseline, data, protocol of the two challenges and the challenge participation are discussed in detail in this paper.
Abhinav Dhall, Roland Göcke, Shreya Ghosh 0001, Jyoti Joshi, Jesse Hoey, Tom Gedeon
ICMI1
2016 Emotion recognition in the wild challenge 2016
abstract
The fourth Emotion Recognition in the Wild (EmotiW) challenge is a grand challenge in the ACM International Conference on Multimodal Interaction 2016, Tokyo. EmotiW is a series of benchmarking and competition effort for researchers working in the area of automatic emotion recognition in the wild. The fourth EmotiW has two sub-challenges: Video based emotion recognition (VReco) and Group-level emotion recognition (GReco). The VReco sub-challenge is being run for the fourth time and GReco is a new sub-challenge this year.
Abhinav Dhall, Roland Göcke, Jyoti Joshi, Tom Gedeon
ICMI1
2016 EmotiW 2016: video and group-level emotion recognition challenges
abstract
This paper discusses the baseline for the Emotion Recognition in the Wild (EmotiW) 2016 challenge. Continuing on the theme of automatic affect recognition `in the wild', the EmotiW challenge 2016 consists of two sub-challenges: an audio-video based emotion and a new group-based emotion recognition sub-challenges. The audio-video based sub-challenge is based on the Acted Facial Expressions in the Wild (AFEW) database. The group-based emotion recognition sub-challenge is based on the Happy People Images (HAPPEI) database. We describe the data, baseline method, challenge protocols and the challenge results. A total of 22 and 7 teams participated in the audio-video based emotion and group-based emotion sub-challenges, respectively.
Abhinav Dhall, Roland Göcke, Jyoti Joshi, Jesse Hoey, Tom Gedeon
ICMI1
2015 A temporally piece-wise fisher vector approach for depression analysis
abstract
Depression and other mood disorders are common, disabling disorders with a profound impact on individuals and families. Inspite of its high prevalence, it is easily missed during the early stages. Automatic depression analysis has become a very active field of research in the affective computing community in the past few years. This paper presents a framework for depression analysis based on unimodal visual cues. Temporally piece-wise Fisher Vectors (FV) are computed on temporal segments. As a low-level feature, block-wise Local Binary Pattern-Three Orthogonal Planes descriptors are computed. Statistical aggregation techniques are analysed and compared for creating a discriminative representative for a video sample. The paper explores the strength of FV in representing temporal segments in a spontaneous clinical data. This creates a meaningful representation of the facial dynamics in a temporal segment. The experiments are conducted on the Audio Video Emotion Challenge (AVEC) 2014 German speaking depression database. The superior results of the proposed framework show the effectiveness of the technique as compared to the current state-of-art.
Abhinav Dhall, Roland Göcke
ACII1
2015 Riesz-based Volume Local Binary Pattern and A Novel Group Expression Model for Group Happiness Intensity Analysis
abstract
Automatic emotion analysis and understanding has received much attention over the years in affective computing. Recently, there are increasing interests in inferring the emotional intensity of a group of people. For group emotional intensity analysis, feature extraction and group expression model are two critical issues. In this paper, we propose a new method to estimate the happiness intensity of a group of people in an image. Firstly, we combine the Riesz transform and the local binary pattern descriptor, named Riesz-based volume local binary pattern, which considers neighbouring changes not only in the spatial domain of a face but also along the different Riesz faces. Secondly, we exploit the continuous conditional random fields for constructing a new group expression model, which considers global and local attributes. Intensive experiments are performed on three challenging facial expression databases to evaluate the novel feature. Furthermore, experiments are conducted on the HAPPEI database to evaluate the new group expression model with the new feature. Our experimental results demonstrate the promising performance for group happiness intensity analysis.
Xiaohua Huang 0003, Abhinav Dhall, Guoying Zhao 0001, Roland Göcke, Matti Pietikäinen
BMVC2
2015 Video and Image based Emotion Recognition Challenges in the Wild: EmotiW 2015
abstract
The third Emotion Recognition in the Wild (EmotiW) challenge 2015 consists of an audio-video based emotion and static image based facial expression classification sub-challenges, which mimics real-world conditions. The two sub-challenges are based on the Acted Facial Expression in the Wild (AFEW) 5.0 and the Static Facial Expression in the Wild (SFEW) 2.0 databases, respectively. The paper describes the data, baseline method, challenge protocol and the challenge results. A total of 12 and 17 teams participated in the video based emotion and image based expression sub-challenges, respectively.
Abhinav Dhall, O. V. Ramana Murthy, Roland Göcke, Jyoti Joshi, Tom Gedeon
ICMI1
2015 Automatic Group Happiness Intensity Analysis
abstract
The recent advancement of social media has given users a platform to socially engage and interact with a larger population. Millions of images and videos are being uploaded everyday by users on the web from different events and social gatherings. There is an increasing interest in designing systems capable of understanding human manifestations of emotional attributes and affective displays. As images and videos from social events generally contain multiple subjects, it is an essential step to study these groups of people. In this paper, we study the problem of happiness intensity analysis of a group of people in an image using facial expression analysis. A user perception study is conducted to understand various attributes, which affect a person's perception of the happiness intensity of a group. We identify the challenges in developing an automatic mood analysis system and propose three models based on the attributes in the study. An `in the wild' image-based database is collected. To validate the methods, both quantitative and qualitative experiments are performed and applied to the problem of shot selection, event summarisation and album creation. The experiments show that the global and local attributes defined in the paper provide useful information for theme expression analysis, with results close to human perception results.
Abhinav Dhall, Roland Göcke, Tom Gedeon
IEEE Trans. Affect. Comput.1
2014 Emotion Recognition In The Wild Challenge 2014: Baseline, Data and Protocol
abstract
The Second Emotion Recognition In The Wild Challenge (EmotiW) 2014 consists of an audio-video based emotion classification challenge, which mimics the real-world conditions. Traditionally, emotion recognition has been performed on data captured in constrained lab-controlled like environment. While this data was a good starting point, such lab controlled data poorly represents the environment and conditions faced in real-world situations. With the exponential increase in the number of video clips being uploaded online, it is worthwhile to explore the performance of emotion recognition methods that work `in the wild'. The goal of this Grand Challenge is to carry forward the common platform defined during EmotiW 2013, for evaluation of emotion recognition methods in real-world conditions. The database in the 2014 challenge is the Acted Facial Expression In Wild (AFEW) 4.0, which has been collected from movies showing close-to-real-world conditions. The paper describes the data partitions, the baseline method and the experimental protocol.
Abhinav Dhall, Roland Göcke, Jyoti Joshi, Karan Sikka, Tom Gedeon
ICMI1
2014 A discriminative parts based model approach for fiducial points free and shape constrained head pose normalisation in the wild
abstract
Continuous Confidence Map Based Normalisation: While continuous head pose normalisation is not the goal of this paper, we demonstrate as a proof of concept that it is possible to extend the current method for continuous head pose normalisation. For dealing with faces in videos [1], continuous head pose normalisation is required. [2] argue that the appearance of a part does not changes with a subtle pose change, therefore a detector for part i in pose angle p can be shared for the same part i for a pose angle p + δ. Further experiments in [2] showed that sharing based models and independent model have comparable performance. However, sharing based models are faster upto ten times as compared to the independent models [2]. The confidence maps based methods (CM-HPNPSand CM-HPNPI) can be extended from discrete to continuous by sharing part-specific regression models R, which are shared among neighboring pose angles.
Abhinav Dhall, Karan Sikka, Gwen Littlewort, Roland Göcke, Marian Stewart Bartlett
WACV1
2014 A discriminative parts based model approach for fiducial points free and shape constrained head pose normalisation in the wild
abstract
This paper proposes a method for parts-based view-invariant head pose normalisation, which works well even in difficult real-world conditions. Handling pose is a classical problem in facial analysis. Recently, parts-based models have shown promising performance for facial landmark points detection `in the wild'. Leveraging on the success of these models, the proposed data-driven regression framework computes a constrained normalised virtual frontal head pose. The response maps of a discriminatively trained part detector are used as texture information. These sparse texture maps are projected from non-frontal to frontal pose using block-wise structured regression. Finally, a facial kinematic shape constraint is achieved by applying a shape model. The advantages of the proposed approach are: a) no explicit dependence on the outputs of a facial parts detector and, thus, avoiding any error propagation owing to their failure; (b) the application of a shape prior on the reconstructed frontal maps provides an anatomically constrained facial shape; and c) modelling head pose as a mixture-of-parts model allows the framework to work without any prior pose information. Experiments are performed on the Multi-PIE and the `in the wild' SFEW databases. The results demonstrate the effectiveness of the proposed method.
Abhinav Dhall, Karan Sikka, Gwen Littlewort, Roland Göcke, Marian Stewart Bartlett
WACV1
2014 Classification and weakly supervised pain localization using multiple segment representation
Karan Sikka, Abhinav Dhall, Marian Stewart Bartlett
Image Vis. Comput.2
2013 Context Based Facial Expression Analysis in the Wild
abstract
With the advances in the computer vision in the past few years, analysis of human facial expressions has gained attention. Facial expression analysis is now an active field of research for over two decades now. However, still there are a lot of questions unanswered. This project will explore and devise algorithms and techniques for facial expression analysis in practical environments. Methods will also be developed for inferring the emotion of a group of people. The central hypothesis of the project is that close to real-world data can be extracted from movies and facial expression analysis on movies is a stepping stone for moving to analysis in the real-world. The data extracted from movies carries with it rich meta-data information which will be useful for exploring the role of context in emotion recognition in the wild. For the analysis of groups of people various attributes effect the perception of mood. Study will be conducted and thorough literature survey will be performed on what are these contextual attributes. A system which can classify the mood of a group of people in videos will be developed and will be used to solve the problem of efficient image browsing and retrieval based on emotion.
Abhinav Dhall
ACII1
2013 Relative Body Parts Movement for Automatic Depression Analysis
abstract
In this paper, a human body part motion analysis based approach is proposed for depression analysis. Depression is a serious psychological disorder. The absence of an (automated) objective diagnostic aid for depression leads to a range of subjective biases in initial diagnosis and ongoing monitoring. Researchers in the affective computing community have approached the depression detection problem using facial dynamics and vocal prosody. Recent works in affective computing have shown the significance of body pose and motion in analysing the psychological state of a person. Inspired by these works, we explore a body parts motion based approach. Relative orientation and radius are computed for the body parts detected using the pictorial structures framework. A histogram of relative parts motion is computed. To analyse the motion on a holistic level, space-time interest points are computed and a bag of words framework is learnt. The two histograms are fused and a support vector machine classifier is trained. The experiments conducted on a clinical database, prove the effectiveness of the proposed method.
Jyoti Joshi, Abhinav Dhall, Roland Göcke, Jeffrey F. Cohn
ACII2
2013 Modeling Stress Using Thermal Facial Patterns: A Spatio-temporal Approach
abstract
Stress is a serious concern facing our world today, motivating the development of better objective understanding using non-intrusive means for stress recognition. The aim for the work was to use thermal imaging of facial regions to detect stress automatically. The work uses facial regions captured in videos in thermal (TS) and visible (VS) spectrums and introduces our database ANU StressDB. It describes the experiment conducted for acquiring TS and VS videos of observers of stressed and not-stressed films for the ANU StressDB. Further, it presents an application of local binary patterns on three orthogonal planes (LBP-TOP) on VS and TS videos for stress recognition. It proposes a novel method to capture dynamic thermal patterns in histograms (HDTP) to utilize thermal and spatio-temporal characteristics associated in TS videos. Individual-independent support vector machine classifiers were developed for stress recognition. Results show that a fusion of facial patterns from VS and TS videos produced significantly better stress recognition rates than patterns from only VS or TS videos with p <; 0.01. The best stress recognition rate was 72% and it was obtained from HDTP features fused with LBP-TOP features for TS and VS videos respectively.
Nandita Sharma, Abhinav Dhall, Tom Gedeon, Roland Göcke
ACII2
2013 Monocular Image 3D Human Pose Estimation under Self-Occlusion
abstract
In this paper, an automatic approach for 3D pose reconstruction from a single image is proposed. The presence of human body articulation, hallucinated parts and cluttered background leads to ambiguity during the pose inference, which makes the problem non-trivial. Researchers have explored various methods based on motion and shading in order to reduce the ambiguity and reconstruct the 3D pose. The key idea of our algorithm is to impose both kinematic and orientation constraints. The former is imposed by projecting a 3D model onto the input image and pruning the parts, which are incompatible with the anthropomorphism. The latter is applied by creating synthetic views via regressing the input view to multiple oriented views. After applying the constraints, the 3D model is projected onto the initial and synthetic views, which further reduces the ambiguity. Finally, we borrow the direction of the unambiguous parts from the synthetic views to the initial one, which results in the 3D pose. Quantitative experiments are performed on the Human Eva-I dataset and qualitatively on unconstrained images from the Image Parse dataset. The results show the robustness of the proposed approach to accurately reconstruct the 3D pose form a single image.
Ibrahim Radwan, Abhinav Dhall, Roland Göcke
ICCV2
2013 Emotion recognition in the wild challenge (EmotiW) challenge and workshop summary
abstract
The Emotion Recognition In The Wild Challenge and Workshop (EmotiW) 2013 Grand Challenge consists of an audio-video based emotion classification challenge, which mimics real-world conditions. In total, 27 teams participated in the challenge. The database in the 2013 challenge is the Acted Facial Expression in the Wild (AFEW), which has been collected from movies showing close-to-real-world conditions.
Abhinav Dhall, Roland Göcke, Jyoti Joshi, Michael Wagner 0004, Tom Gedeon
ICMI1
2013 Emotion recognition in the wild challenge 2013
abstract
Emotion recognition is a very active field of research. The Emotion Recognition In The Wild Challenge and Workshop (EmotiW) 2013 Grand Challenge consists of an audio-video based emotion classification challenges, which mimics real-world conditions. Traditionally, emotion recognition has been performed on laboratory controlled data. While undoubtedly worthwhile at the time, such laboratory controlled data poorly represents the environment and conditions faced in real-world situations. The goal of this Grand Challenge is to define a common platform for evaluation of emotion recognition methods in real-world conditions. The database in the 2013 challenge is the Acted Facial Expression in the Wild (AFEW), which has been collected from movies showing close-to-real-world conditions.
Abhinav Dhall, Roland Göcke, Jyoti Joshi, Michael Wagner 0004, Tom Gedeon
ICMI1
2013 Expression analysis in the wild: from individual to groups
abstract
With the advances in the computer vision in the past few years, analysis of human facial expressions has gained attention. Facial expression analysis is now an active field of research for over two decades now. However, still there are a lot of questions unanswered. This project will explore and devise algorithms and techniques for facial expression analysis in practical environments. Methods will also be developed for inferring the emotion of a group of people. The central hypothesis of the project is that close to real-world data can be extracted from movies and facial expression analysis on movies is a stepping stone for moving to analysis in the real-world. For the analysis of groups of people various attributes effect the perception of mood. A system which can classify the mood of a group of people in videos will be developed and will be used to solve the problem of efficient image browsing and retrieval based on emotion.
Abhinav Dhall
ICMR1
2012 Finding Happiest Moments in a Social Context
Abhinav Dhall, Jyoti Joshi, Ibrahim Radwan, Roland Göcke
ACCV (2)1
2012 Regression Based Pose Estimation with Automatic Occlusion Detection and Rectification
abstract
Human pose estimation is a classic problem in computer vision. Statistical models based on part-based modelling and the pictorial structure framework have been widely used recently for articulated human pose estimation. However, the performance of these models has been limited due to the presence of self-occlusion. This paper presents a learning-based framework to automatically detect and recover self-occluded body parts. We learn two different models: one for detecting occluded parts in the upper body and another one for the lower body. To solve the key problem of knowing which parts are occluded, we construct Gaussian Process Regression (GPR) models to learn the parameters of the occluded body parts from their corresponding ground truth parameters. Using these models, the pictorial structure of the occluded parts in unseen images is automatically rectified. The proposed framework outperforms a state-of-the-art pictorial structure approach for human pose estimation on 3 different datasets.
Ibrahim Radwan, Abhinav Dhall, Jyoti Joshi, Roland Göcke
ICME2
2012 Group expression intensity estimation in videos via Gaussian Processes
Abhinav Dhall, Roland Göcke
ICPR1
2012 Neural-net classification for spatio-temporal descriptor based depression analysis
Jyoti Joshi, Abhinav Dhall, Roland Göcke, Michael Breakspear, Gordon Parker
ICPR2
2012 Correcting pose estimation with implicit occlusion detection and rectification
Ibrahim Radwan, Abhinav Dhall, Roland Göcke
ICPR2
2012 Facial Performance Transfer via Deformable Models and Parametric Correspondence
abstract
The issue of transferring facial performance from one person's face to another's has been an area of interest for the movie industry and the computer graphics community for quite some time. In recent years, deformable face models, such as the Active Appearance Model (AAM), have made it possible to track and synthesize faces in real time. Not surprisingly, deformable face model-based approaches for facial performance transfer have gained tremendous interest in the computer vision and graphics community. In this paper, we focus on the problem of real-time facial performance transfer using the AAM framework. We propose a novel approach of learning the mapping between the parameters of two completely independent AAMs, using them to facilitate the facial performance transfer in a more realistic manner than previous approaches. The main advantage of modeling this parametric correspondence is that it allows a "meaningful" transfer of both the nonrigid shape and texture across faces irrespective of the speakers' gender, shape, and size of the faces, and illumination conditions. We explore linear and nonlinear methods for modeling the parametric correspondence between the AAMs and show that the sparse linear regression method performs the best. Moreover, we show the utility of the proposed framework for a cross-language facial performance transfer that is an area of interest for the movie dubbing industry.
Akshay Asthana, Miles de la Hunty, Abhinav Dhall, Roland Göcke
IEEE Trans. Vis. Comput. Graph.3
2011 A SSIM-based approach for finding similar facial expressions
abstract
There are various scenarios where finding the most similar expression is the requirement rather than classifying one into discrete, pre-defined classes, for example, for facial expression transfer and facial expression based automatic album generation. This paper proposes a novel method for finding the most similar facial expression. Instead of the regular L2 norm distance, we investigate the use of the Structural SIMilarity (SSIM) metric for similarity comparison as a distance metric in a nearest neighbour unsupervised algorithm. The feature vectors are generated using Active Appearance Models (AAM). We also demonstrate how this technique can be extended and used for finding corresponding facial expression images across two or more subjects, which is useful in applications such as facial animation and automatic expression transfer. Person-independent facial expression performance results are shown on the Multi-PIE, FEEDTUM and AVOZES databases. We also compare the performance of the SSIM metric versus other distance metrics in a nearest neighbour search for finding the most similar facial expression to a given image.
Abhinav Dhall, Akshay Asthana, Roland Göcke
FG1
2011 Emotion recognition using PHOG and LPQ features
abstract
We propose a method for automatic emotion recognition as part of the FERA 2011 competition. The system extracts pyramid of histogram of gradients (PHOG) and local phase quantisation (LPQ) features for encoding the shape and appearance information. For selecting the key frames, K-means clustering is applied to the normalised shape vectors derived from constraint local model (CLM) based face tracking on the image sequences. Shape vectors closest to the cluster centers are then used to extract the shape and appearance features. We demonstrate the results on the SSPNET GEMEP-FERA dataset. It comprises of both person specific and person independent partitions. For emotion classification we use support vector machine (SVM) and largest margin nearest neighbour (LMNN) and compare our results to the pre-computed FERA 2011 emotion challenge baseline.
Abhinav Dhall, Akshay Asthana, Roland Göcke, Tom Gedeon
FG1
2010 Facial Expression Based Automatic Album Creation
Abhinav Dhall, Akshay Asthana, Roland Göcke
ICONIP (2)1