Shreya Ghosh 0001

dblp:133/4864-1 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
15since 2021 · last 2025
0000-0002-2639-8374ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 MRAC 2025: 3rd International Workshop on Multimodal, Generative and Responsible Affective Computing
abstract
Multimodal, generative, and responsible affective computing aims to enhance people's lives. In recent years, the AI revolution has already begun to impact daily life, with virtual assistants being deployed across various sectors such as healthcare, banking, transportation, and education. It is clear that, in the near future, humans may interact with AI-powered systems as much or maybe even more than direct human-to-human interactions. Affective computing has numerous applications, including innovative approaches to forecasting and preventing anxiety, stress, and mental health issues; enhancing robotic empathy; assisting individuals with communication, behavior, and emotion regulation challenges; and promoting awareness of health and well-being. Many of these applications require enhanced control and protection of sensitive, private, and personal data. Therefore, it is crucial to further develop the creation, evaluation, and deployment of emotionally intelligent systems that are both responsive and responsible. Additionally, improving the accuracy and interpretability of emotion prediction results can significantly enhance the application of this technology in the downstream tasks mentioned above. MRAC'25 is the continuation of MRAC'23 and MRAC'24. Through this workshop, we aim to bring together researchers to discuss the potential and development of affective computing.
Zheng Lian 0004, Shreya Ghosh 0001, Erik Cambria, Zhixi Cai, Guoying Zhao 0001, Abhinav Dhall, Björn W. Schuller, Roland Göcke, Jianhua Tao 0001, Tom Gedeon
ACM Multimedia2
2025 AV-Deepfake1M++: A Large-Scale Audio-Visual Deepfake Benchmark with Real-World Perturbations
abstract
The rapid surge of text-to-speech and face-voice reenactment models makes video fabrication easier and highly realistic. To encounter this problem, we require datasets that rich in type of generation methods and perturbation strategy which is usually common for online videos. To this end, we propose AV-Deepfake1M++, an extension of the AV-Deepfake1M having 2 million video clips with diversified manipulation strategy and audio-visual perturbation. This paper includes the description of data generation strategies along with benchmarking of AV-Deepfake1M++ using state-of-the-art methods. We believe that this dataset will play a pivotal role in facilitating research in Deepfake domain. Based on this dataset, we host the 2025 1M-Deepfakes Detection Challenge. The challenge details, dataset and evaluation scripts are available online under a research-only license at https://deepfakes1m.github.io/2025.
Zhixi Cai, Kartik Kuckreja, Shreya Ghosh 0001, Akanksha Chuchra, Muhammad Haris Khan, Usman Tariq, Tom Gedeon, Abhinav Dhall
ACM Multimedia3
2025 Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations
Shreya Ghosh 0001, Tom Gedeon, Thanh-Toan Do, Abhinav Dhall
ACM Multimedia2
2025 MIP-GAF: A MLLM-Annotated Benchmark for Most Important Person Localization and Group Context Understanding
abstract
Estimating the Most Important Person (MIP) in any social event setup is a challenging problem mainly due to contextual complexity and scarcity of labeled data. Moreover, the causality aspects of MIP estimation are quite subjective and diverse. To this end, we aim to address the problem by annotating a large-scale ‘in-the-wild’ dataset for iden-tifying human perceptions about the ‘Most Important Person (MIP)‘ in an image. The paper provides a thorough description of our proposed Multimodal Large Language Model (MLLM) based data annotation strategy, and a thor-ough data quality analysis. Further, we perform a comprehensive benchmarking of the proposed dataset utilizing state-of-the-art MIP localization methods, indicating a significant drop in performance compared to existing datasets. The performance drop shows that the existing MIP localization algorithms must be more robust with respect to ‘in-the-wild’ situations. We believe the proposed dataset will play a vital role in building the next-generation social situation understanding methods. The dataset and associated code will be made available for research purposes.
Surbhi Madan, Shreya Ghosh 0001, Lownish Rai Sookha, M. A. Ganaie 0001, Subramanian Ramanathan, Abhinav Dhall, Tom Gedeon
WACV2
2025 Empathy Detection From Text, Audiovisual, Audio or Physiological Signals: A Systematic Review of Task Formulations and Machine Learning Methods
abstract
Empathy indicates an individual's ability to understand others. Over the past few years, empathy has drawn attention from various disciplines, including but not limited to Affective Computing, Cognitive Science, and Psychology. Detecting empathy has potential applications in society, healthcare and education. Despite being a broad and overlapping topic, the avenue of empathy detection leveraging Machine Learning remains underexplored from a systematic literature review perspective. We collected 849 papers from 10 well-known academic databases, systematically screened them and analysed the final 82 papers. Our analyses reveal several prominent task formulations – including empathy on localised utterances or overall expressions, unidirectional or parallel empathy, and emotional contagion – in monadic, dyadic and group interactions. Empathy detection methods are summarised based on four input modalities – text, audiovisual, audio and physiological signals – thereby presenting modality-specific network architecture design protocols. We discuss challenges, research gaps and potential applications in the Affective Computing-basedempathydomain, which can facilitate new avenues of exploration. We further enlist the public availability of datasets and codes. This paper, therefore, provides a structured overview of recent advancements and remaining challenges towards developing a robust empathy detection system that could meaningfully contribute to enhancing human well-being.
Md. Rakibul Hasan 0001, Shreya Ghosh 0001, Aneesh Krishna, Tom Gedeon
IEEE Trans. Affect. Comput.3
2024 AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake Dataset
abstract
The detection and localization of highly realistic deepfake audio-visual content are challenging even for the most advanced state-of-the-art methods. While most of the research efforts in this domain are focused on detecting high-quality deepfake images and videos, only a few works address the problem of the localization of small segments of audio-visual manipulations embedded in real videos. In this research, we emulate the process of such content generation and propose the AV-Deepfake1M dataset. The dataset contains content-driven (i) video manipulations, (ii) audio manipulations, and (iii) audio-visual manipulations for more than 2K subjects resulting in a total of more than 1M videos. The paper provides a thorough description of the proposed data generation pipeline accompanied by a rigorous analysis of the quality of the generated data. The comprehensive benchmark of the proposed dataset utilizing state-of-the-art deepfake detection and localization methods indicates a significant drop in performance compared to previous datasets. The proposed dataset will play a vital role in building the next-generation deepfake localization methods. The dataset and associated code are available at https://github.com/ControlNet/AV-Deepfake1M.
Zhixi Cai, Shreya Ghosh 0001, Aman Pankaj Adatia, Munawar Hayat, Abhinav Dhall, Tom Gedeon, Kalin Stefanov
ACM Multimedia2
2024 1M-Deepfakes Detection Challenge
abstract
The detection and localization of deepfake content, particularly when small fake segments are seamlessly mixed with real videos, remains a significant challenge in the field of digital media security. Based on the recently released AV-Deepfake1M dataset, which contains more than 1 million manipulated videos across more than 2,000 subjects, we introduce the 1M-Deepfakes Detection Challenge. This challenge is designed to engage the research community in developing advanced methods for detecting and localizing deepfake manipulations within the large-scale high-realistic audio-visual dataset. The participants can access the AV-Deepfake1M dataset and are required to submit their inference results for evaluation across the metrics for detection or localization tasks. The methodologies developed through the challenge will contribute to the development of next-generation deepfake detection and localization systems. Evaluation scripts, baseline models, and accompanying code will be available on https://github.com/ControlNet/AV-Deepfake1M.
Zhixi Cai, Abhinav Dhall, Shreya Ghosh 0001, Munawar Hayat, Dimitris Kollias, Kalin Stefanov, Usman Tariq
ACM Multimedia3
2024 Automatic Gaze Analysis: A Survey of Deep Learning Based Approaches
abstract
Eye gaze analysis is an important research problem in the field of Computer Vision and Human-Computer Interaction. Even with notable progress in the last 10 years, automatic gaze analysis still remains challenging due to the uniqueness of eye appearance, eye-head interplay, occlusion, image quality, and illumination conditions. There are several open questions, including what are the important cues to interpret gaze direction in an unconstrained environment without prior knowledge and how to encode them in real-time. We review the progress across a range of gaze analysis tasks and applications to elucidate these fundamental questions, identify effective methods in gaze analysis, and provide possible future directions. We analyze recent gaze estimation and segmentation methods, especially in the unsupervised and weakly supervised domain, based on their advantages and reported evaluation metrics. Our analysis shows that the development of a robust and generic gaze analysis method still needs to address real-world challenges such as unconstrained setup and learning with less supervision. We conclude by discussing future research directions for designing a real-world gaze analysis system that can propagate to other domains including Computer Vision, Augmented Reality (AR), Virtual Reality (VR), and Human Computer Interaction (HCI).
Shreya Ghosh 0001, Abhinav Dhall, Munawar Hayat, Jarrod Knibbe
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 MARLIN: Masked Autoencoder for facial video Representation LearnINg
abstract
This paper proposes a self-supervised approach to learn universal facial representations from videos, that can transfer across a variety of facial analysis tasks such as Facial Attribute Recognition (FAR), Facial Expression Recognition (FER), DeepFake Detection (DFD), and Lip Synchronization (LS). Our proposed framework, named MARLIN, is a facial video masked autoencoder, that learns highly robust and generic facial embeddings from abundantly available non-annotated web crawled facial videos. As a challenging auxiliary task, MARLIN reconstructs the spatio-temporal details of the face from the densely masked facial regions which mainly include eyes, nose, mouth, lips, and skin to capture local and global aspects that in turn help in encoding generic and transferable features. Through a variety of experiments on diverse downstream tasks, we demonstrate MARLIN to be an excellent facial video encoder as well as feature extractor, that performs consistently well across a variety of downstream tasks including FAR (1.13% gain over supervised benchmark), FER (2.64% gain over unsupervised benchmark), DFD (1.86% gain over unsupervised benchmark), LS (29.36% gain for Frechet Inception Distance), and even in low data regime. Our code and models are available at https://github.com/ControlNet/MARLIN.
Zhixi Cai, Shreya Ghosh 0001, Kalin Stefanov, Abhinav Dhall, Jianfei Cai 0001, Seyed Hamid Rezatofighi, Gholamreza Haffari, Munawar Hayat
CVPR2
2023 Glitch in the matrix: A large scale benchmark for content driven audio-visual forgery detection and localization
abstract
Most deepfake detection methods focus on detecting spatial and/or spatio-temporal changes in facial attributes and are centered around the binary classification task of detecting whether a video is real or fake. This is because available benchmark datasets contain mostly visual-only modifications present in the entirety of the video. However, a sophisticated deepfake may include small segments of audio or audio-visual manipulations that can completely change the meaning of the video content. To addresses this gap, we propose and benchmark a new dataset, Localized Audio Visual DeepFake (LAV-DF), consisting of strategic content-driven audio, visual and audio-visual manipulations. The proposed baseline method, Boundary Aware Temporal Forgery Detection (BA-TFD), is a 3D Convolutional Neural Network-based architecture which effectively captures multimodal manipulations. We further improve (i.e. BA-TFD+) the baseline method by replacing the backbone with a Multiscale Vision Transformer and guide the training process with contrastive, frame classification, boundary matching and multimodal boundary matching loss functions. The quantitative analysis demonstrates the superiority of BA-TFD+ on temporal forgery localization and deepfake detection tasks using several benchmark datasets including our newly proposed dataset. The dataset, models and code are available at https://github.com/ControlNet/LAV-DF.
Zhixi Cai, Shreya Ghosh 0001, Abhinav Dhall, Tom Gedeon, Kalin Stefanov, Munawar Hayat
Comput. Vis. Image Underst.2
2022 'Labelling the Gaps': A Weakly Supervised Automatic Eye Gaze Estimation
Shreya Ghosh 0001, Abhinav Dhall, Munawar Hayat, Jarrod Knibbe
ACCV (4)1
2022 AV-GAZE: A Study on the Effectiveness of Audio Guided Visual Attention Estimation for Non-profilic Faces
abstract
In challenging real-life conditions such as extreme head-pose, occlusions, and low-resolution images where the visual information fails to estimate visual attention/gaze direction, audio signals could provide important and complementary information. In this paper, we explore if audio-guided coarse head-pose can further enhance visual attention estimation performance for non-prolific faces. Since it is difficult to annotate audio signals for estimating the head-pose of the speaker, we use off-the-shelf state-of-the-art models to facilitate cross-modal weak-supervision. During the training phase, the framework learns complementary information from synchronized audio-visual modality. Our model can utilize any of the available modalities i.e. audio, visual or audio-visual for task-specific inference. It is interesting to note that, when AV-Gaze is tested on benchmark datasets with these specific modalities, it achieves competitive results on multiple datasets, while being highly adaptive toward challenging scenarios.
Shreya Ghosh 0001, Abhinav Dhall, Munawar Hayat, Jarrod Knibbe
ICIP1
2022 MTGLS: Multi-Task Gaze Estimation with Limited Supervision
abstract
Robust gaze estimation is a challenging task, even for deep CNNs, due to the non-availability of large-scale labeled data. Moreover, gaze annotation is a time-consuming process and requires specialized hardware setups. We propose MTGLS: a Multi-Task Gaze estimation framework with Limited Supervision, which leverages abundantly available non-annotated facial image data. MTGLS distills knowledge from off-the-shelf facial image analysis models, and learns strong feature representations of human eyes, guided by three complementary auxiliary signals: (a) the line of sight of the pupil (i.e. pseudo-gaze) defined by the localized facial landmarks, (b) the head-pose given by Euler angles, and (c) the orientation of the eye patch (left/right eye). To overcome inherent noise in the supervisory signals, MT-GLS further incorporates a noise distribution modelling approach. Our experimental results show that MTGLS learns highly generalized representations which consistently perform well on a range of datasets. Our proposed framework outperforms the unsupervised state-of-the-art on CAVE (by ∼ 6.43%) and even supervised state-of-the-art methods on Gaze360 (by ∼ 6.59%) datasets.
Shreya Ghosh 0001, Munawar Hayat, Abhinav Dhall, Jarrod Knibbe
WACV1
2022 Automatic Prediction of Group Cohesiveness in Images
abstract
This article discusses the prediction of cohesiveness of a group of people in images. The cohesiveness of a group is an essential indicator of the emotional state, structure, and success of the group. We study the factors that influence the perception of group-level cohesion and propose methods for estimating the human-perceived cohesion on the group cohesiveness scale. To identify the visual cues (attributes) for cohesion, we conducted a user survey. Image analysis is performed at a group-level via a multi-task convolutional neural network. A capsule network is explored for analyzing the contribution of facial expressions of the group members on predicting the Group Cohesion Score (GCS). We add GCS to the Group Affect database and propose the ‘GAF-Cohesion database’. The proposed model performs well on the database and achieves near human-level performance in predicting a group's cohesion score. It is interesting to note that group cohesion as an attribute, when jointly trained for group-level emotion prediction, helps in increasing the performance for the later task. This suggests that group-level emotion and cohesion are correlated. Further, we investigate the effect of face-level similarity, body pose and subset of a group on the task of automatic cohesion perception.
Shreya Ghosh 0001, Abhinav Dhall, Nicu Sebe, Tom Gedeon
IEEE Trans. Affect. Comput.1
2021 Depression Intensity Estimation via Social Media: A Deep Learning Approach
abstract
Depression has become a big problem in our society today. It is also a major reason for suicide, especially among teenagers. In the current outbreak of coronavirus disease (COVID-19), the affected countries have recommended social distancing and lockdown measures. Resulting in interpersonal isolation, these measures have raised serious concerns for mental health and depression. Generally, clinical psychologists diagnose depressed people via face-to-face interviews following the clinical depression criteria. However, often patients tend to not consult doctors in their early stages of depression. Nowadays, people are increasingly using social media to express their moods. In this article, we aim to predict depressed users as well as estimate their depression intensity via leveraging social media (Twitter) data, in order to aid in raising an alarm. We model this problem as a supervised learning task. We start with weakly labeling the Twitter data in a self-supervised manner. A rich set of features, including emotional, topical, behavioral, user level, and depression-related$n$-gram features, are extracted to represent each user. Using these features, we train a small long short-term memory (LSTM) network using Swish as an activation function, to predict the depression intensities. We perform extensive experiments to demonstrate the efficacy of our method. We outperform the baseline models for depression intensity estimation by achieving the lowest mean squared error of 1.42 and also outperform the existing state-of-the-art binary classification method by more than 2% of accuracy. We found that the depressed users frequently use negative words such as stress and sad, mostly post during late nights, highly use personal pronouns and sometimes also share personal events.
Shreya Ghosh 0001, Tarique Anwar
IEEE Trans. Comput. Soc. Syst.1
2020 LSTM-DNN based Approach for Pain Intensity and Protective Behaviour Prediction
abstract
This paper proposes an approach for pain intensity recognition and protective behaviour prediction task from body movements as a part of the EmoPain challenge. The given dataset consists of body part based sensor data for both the tasks. The proposed network is a lightweight LSTM-DNN model, which takes the angle, angle energy and sEMG features as input and predicts pain intensity level and protective behaviour as output. The performance of LSTM, Bi-LSTM, attention-LSTM and LSTM-DNN models are compared for this problem on the same dataset. In order to enhance the model’s discriminating power, joint training of all the models are performed, combining respective task labels with exercise type as an additional label. The experiments show that the proposed approach is effective and outperforms the baseline on the validation set by a margin of 35.00% for pain intensity prediction and 47.72% for protective behaviour prediction, respectively.
Shreya Ghosh 0001, Jyoti Joshi, Sharon L. Oviatt
FG2
2019 Unsupervised Learning of Eye Gaze Representation from the Web
abstract
Automatic eye gaze estimation has interested researchers for a while now. In this paper, we propose an unsupervised learning based method for estimating the eye gaze region. To train the proposed network "Ize-Net" in self-supervised manner, we collect a large `in the wild' dataset containing 1,54,251 images from the web. For the images in the database, we divide the gaze into three regions based on an automatic technique based on pupil-centers localization and then use a feature-based technique to determine the gaze region. The performance is evaluated on the Tablet Gaze and CAVE datasets by fine-tuning results of Ize-Net for the task of eye gaze estimation. The feature representation learned is also used to train traditional machine learning algorithms for eye gaze estimation. The results demonstrate that the proposed method learns a rich data representation, which can be efficiently finetuned for any eye gaze estimation dataset.
Neeru Dubey, Shreya Ghosh 0001, Abhinav Dhall
IJCNN2
2019 Predicting Group Cohesiveness in Images
abstract
The cohesiveness of a group is an essential indicator of the emotional state, structure and success of a group of people. We study the factors that influence the perception of group-level cohesion and propose methods for estimating the human-perceived cohesion on the group cohesiveness scale. In order to identify the visual cues (attributes) for cohesion, we conducted a user survey. Image analysis is performed at a group-level via a multi-task convolutional neural network. For analyzing the contribution of facial expressions of the group members for predicting the Group Cohesion Score (GCS), a capsule network is explored. We add GCS to the Group Affect database and propose the `GAF-Cohesion database'. The proposed model performs well on the database and is able to achieve near human-level performance in predicting a group's cohesion score. It is interesting to note that group cohesion as an attribute, when jointly trained for group-level emotion prediction, helps in increasing the performance for the later task. This suggests that group-level emotion and cohesion are correlated.
Shreya Ghosh 0001, Abhinav Dhall, Nicu Sebe, Tom Gedeon
IJCNN1
2018 Automatic Group Affect Analysis in Images via Visual Attribute and Feature Networks
abstract
This paper proposes a pipeline for automatic group-level affect analysis. A deep neural network-based approach, which leverages on the facial-expression information, scene information and a high-level facial visual attribute information is proposed. A capsule network-based architecture is used to predict the facial expression. Transfer learning is used on Inception-V3 to extract global image-based features which contain scene information. Another network is trained for inferring the facial attributes of the group members. Further, these attributes are pooled at a group-level to train a network for inferring the group-level affect. The facial attribute prediction network, although is simple yet, is effective and generates result comparable to the state-of-the-art methods. Later, model integration is performed from the three channels. The experiments show the effectiveness of the proposed techniques on three `in the wild' databases: Group Affect Database, HAPPEI and UCLA-Protest database.
Shreya Ghosh 0001, Abhinav Dhall, Nicu Sebe
ICIP1
2017 From individual to group-level emotion recognition: EmotiW 5.0
abstract
Research in automatic affect recognition has come a long way. This paper describes the fifth Emotion Recognition in the Wild (EmotiW) challenge 2017. EmotiW aims at providing a common benchmarking platform for researchers working on different aspects of affective computing. This year there are two sub-challenges: a) Audio-video emotion recognition and b) group-level emotion recognition. These challenges are based on the acted facial expressions in the wild and group affect databases, respectively. The particular focus of the challenge is to evaluate method in `in the wild' settings. `In the wild' here is used to describe the various environments represented in the images and videos, which represent real-world (not lab like) scenarios. The baseline, data, protocol of the two challenges and the challenge participation are discussed in detail in this paper.
Abhinav Dhall, Roland Göcke, Shreya Ghosh 0001, Jyoti Joshi, Jesse Hoey, Tom Gedeon
ICMI3