Roland Göcke

dblp:g/RolandGocke · also Roland Goecke · DBLP profile ↗
← Back
96ranked-venue papers
1as first author
15since 2021 · last 2026
0000-0003-2279-7041ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 59 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 55 · 7 since 2021Human-computer interaction and ubiquitous computing · 16 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MDNet: A Lightweight Multidomain 1-D CNN for Embedded Pain Assessment Using EDA Signals
abstract
Accurate pain assessment is critical for effective clinical intervention and patient management. Traditional pain assessment methods, such as self-reported scales, are subjective and often unsuitable for patients with cognitive impairments or limited communication abilities. This study proposes MDNet, a multi-domain one-dimensional convolutional neural network designed for real-time pain assessment using electrodermal activity signals. The proposed model extracts information from time, spectral, and cepstral domains to enhance pain classification accuracy. We evaluated MDNet using leave-one-subject-out cross-validation on the AI4Pain and BioVid datasets. The model achieved 92.08% accuracy for binary pain classification (No Pain vs. High Pain) and 71.91% accuracy for multiclass classification (No Pain vs. Low Pain vs. High Pain). Additionally, we deployed MDNet on a Raspberry Pi 5 (RPi-5), demonstrating its feasibility for real-time embedded healthcare applications with low end-to-end latency (102–119 ms), low CPU load (27.5%—30.6%), and efficient power consumption (4.6–4.8 W). These results highlight the potential of integrating deep learning-based pain assessment into wearable and embedded systems, offering a promising step toward automated, real-time pain monitoring in clinical and non-clinical settings.
Sumair Aziz, Girija Chetty, Roland Göcke, Raul Fernandez Rojas
IEEE Internet Things J.3
2026 Exploring Prefrontal Cortex Involvement in Postural Control Across Degraded Sensory Conditions Using fNIRS and Classification
abstract
The prefrontal cortex (PFC) of the brain is involved in processing visual, vestibular, and somatosensory inputs to stabilise postural balance. However, the PFC's activation map for a standing person and different sensory inputs remains unclear. This study aimed to explore the PFC activity map and distinct haemodynamic responses during postural control when sensory inputs change. To this end, functional near-infrared spectroscopy (fNIRS) was employed to capture the haemodynamic responses throughout the PFC from a group of young adults standing in four sensory conditions. The results revealed distinct PFC activation patterns supporting sensory processing, motor planning, and cognitive control to maintain balance under different degraded sensory conditions. Furthermore, by applying machine learning classifiers and multivariate feature selection, the PFC locations and haemodynamic responses indicative of different sensory conditions were identified. The findings of this study offer valuable insights for optimising rehabilitation approaches, enhancing the design of fNIRS studies, and advancing brain-computer interface technologies for balance assessment and training.
Yasaman Baradaran, Raul Fernandez Rojas, Roland Göcke, Maryam Ghahramani
IEEE J. Biomed. Health Informatics3
2025 MRAC 2025: 3rd International Workshop on Multimodal, Generative and Responsible Affective Computing
abstract
Multimodal, generative, and responsible affective computing aims to enhance people's lives. In recent years, the AI revolution has already begun to impact daily life, with virtual assistants being deployed across various sectors such as healthcare, banking, transportation, and education. It is clear that, in the near future, humans may interact with AI-powered systems as much or maybe even more than direct human-to-human interactions. Affective computing has numerous applications, including innovative approaches to forecasting and preventing anxiety, stress, and mental health issues; enhancing robotic empathy; assisting individuals with communication, behavior, and emotion regulation challenges; and promoting awareness of health and well-being. Many of these applications require enhanced control and protection of sensitive, private, and personal data. Therefore, it is crucial to further develop the creation, evaluation, and deployment of emotionally intelligent systems that are both responsive and responsible. Additionally, improving the accuracy and interpretability of emotion prediction results can significantly enhance the application of this technology in the downstream tasks mentioned above. MRAC'25 is the continuation of MRAC'23 and MRAC'24. Through this workshop, we aim to bring together researchers to discuss the potential and development of affective computing.
Zheng Lian 0004, Shreya Ghosh 0001, Erik Cambria, Zhixi Cai, Guoying Zhao 0001, Abhinav Dhall, Björn W. Schuller, Roland Göcke, Jianhua Tao 0001, Tom Gedeon
ACM Multimedia8
2025 Pain Assessment Using Multi-Kernel-FCN-LSTM and Haemoglobin Difference in fNIRS
abstract
This study investigates the effectiveness of various machine learning and deep learning models for automated pain detection using Functional Near-Infrared Spectroscopy (fNIRS) data from the AI4Pain Grand Challenge dataset. Four different near-infrared spectroscopy metrics—Oxygenated Haemoglobin (HbO2), Deoxygenated Haemoglobin (HHb), Total Haemoglobin (HT) and Haemoglobin Difference (HbDiff)—were investigated to determine their contributions to pain assessment and identify which metric offers the most reliable performance. Across all models, both traditional and deep learning, HbDiff consistently outperformed the other metrics in terms of classification accuracy. The Multi-Kernel Fully Convolutional Network Hybrid with Long Short-Term Memory (MK-FCN-LSTM) model, particularly when utilising the HbDiff metric, achieved superior performance with a binary classification accuracy of 64.73%. These findings suggest that haemoglobin difference may provide more sensitive and reliable features for pain assessment, highlighting its potential as a key biomarker in fNIRS-based pain detection systems.
Ghazal Bargshady, Sumair Aziz, Stefanos Gkikas, Manolis Tsiknakis, Roland Göcke, Raul Fernandez Rojas
ACM Trans. Comput. Heal.5
2025 Velocity control of a Stephenson III six-bar linkage-based gait rehabilitation robot using deep reinforcement learning
Akim Kapsalyamov, Nicholas A. T. Brown, Roland Göcke, Prashant Kumar Jamwal, Shahid Hussain 0003
Neural Comput. Appl.3
2025 Inverse kinematics solution for a six-degree-of-freedom upper limb rehabilitation robot using deep learning models
abstract
Abstract The inverse kinematics problem in serially manipulated upper limb rehabilitation robots implies the usage of the end-effector position to obtain the joint rotation angles. In contrast to the forward kinematics, there are no systematic approaches for solving the inverse kinematics problem. Furthermore, for some morphology of the upper limb rehabilitation robots, the inverse kinematics problem is particularly challenging to solve. Conventional methods to solve the inverse kinematics problem reported in the literature are computationally expensive. In the present work, we propose a deep learning-based model to acquire the joint angles for a given end-effector position. The proposed approach exhibits high efficacy in determining the joint angles for various target positions and can accurately predict the end-effector positions once trained, improving the ability of the upper limb rehabilitation robot to adapt to varying patient needs. Due to its improved capability and effectiveness to track positions, the proposed algorithm lays the foundation for the development of efficient controllers in future.
Muhammad Faizan Shah, Naveed Ahmad Khan, Prashant Kumar Jamwal, Girija Chetty, Roland Göcke, Shahid Hussain 0003
Neural Comput. Appl.5
2023 A Weakly Supervised Approach to Emotion-change Prediction and Improved Mood Inference
abstract
Whilst a majority of affective computing research focuses on inferring emotions, examining mood or understanding the mood-emotion interplay has received significantly less attention. Building on prior work, we (a) deduce and incorporate emotion-change ($\Delta$) information for inferring mood, without resorting to annotated labels, and (b) attempt mood prediction for long duration video clips, in alignment with the characterisation of mood. We generate the emotion-change ($\Delta$) labels via metric learning from a pre-trained Siamese Network, and use these in addition to mood labels for mood classification. Experiments evaluating unimodal (training only using mood labels) vs muttimodat (training using mood plus $\Delta$ labels) models show that mood prediction benefits from the incorporation of emotion-change information, emphasising the importance of modelling the moodemotion interplay for effective mood inference.
Soujanya Narayana, Ibrahim Radwan, Ravikiran Parameshwara, Iman Abbasnejad, Akshay Asthana, Subramanian Ramanathan, Roland Göcke
ACII7
2023 Examining Subject-Dependent and Subject-Independent Human Affect Inference from Limited Video Data
abstract
Continuous human affect estimation from video data entails modelling the dynamic emotional state from a sequence of facial images. Though multiple affective video databases exist, they are limited in terms of data and dy-namic annotations, as assigning continuous affective labels to video data is subjective, onerous and tedious. While studies have established the existence of signature facial expressions corresponding to the basic categorical emotions, individual differences in emoting facial expressions nevertheless exist; factoring out these idiosyncrasies is critical for effective emotion inference. This work explores continuous human affect recognition using AFEW-VA, an ‘in-the-wild’ video dataset with limited data, employing subject-independent (SI) and subject-dependent (SD) settings. The SI setting involves the use of training and test sets with mutually exclusive subjects, while training and test samples corresponding to the same subject can occur in the SD setting. A novel, dynamically-weighted loss function is employed with a Convolutional Neural Network (CNN)-Long Short- Term Memory (LSTM) architecture to optimise dynamic affect prediction. Superior prediction is achieved in the SD setting, as compared to the SI counterpart.
Ravikiran Parameshwara, Ibrahim Radwan, Subramanian Ramanathan, Roland Göcke
FG4
2023 Efficient Labelling of Affective Video Datasets via Few-Shot & Multi-Task Contrastive Learning
abstract
Whilst deep learning techniques have achieved excellent emotion prediction, they still require large amounts of labelled training data, which are (a) onerous and tedious to compile, and (b) prone to errors and biases. We propose Multi-Task Contrastive Learning for Affect Representation (MT-CLAR) for few-shot affect inference. MT-CLAR combines multi-task learning with a Siamese network trained via contrastive learning to infer from a pair of expressive facial images (a) the (dis)similarity between the facial expressions, and (b) the difference in valence and arousal levels of the two faces. We further extend the image-based MT-CLAR framework for automated video labelling where, given one or a few labelled video frames (termed support-set), MT-CLAR labels the remainder of the video for valence and arousal. Experiments are performed on the AFEW-VA dataset with multiple support-set configurations; moreover, supervised learning on representations learnt via MT-CLAR are used for valence, arousal and categorical emotion prediction on the AffectNet and AFEW-VA datasets. The results show that valence and arousal predictions via MT-CLAR are very comparable to the state-of-the-art (SOTA), and we significantly outperform SOTA with a support-set ≈6% the size of the video dataset.
Ravikiran Parameshwara, Ibrahim Radwan, Akshay Asthana, Iman Abbasnejad, Subramanian Ramanathan, Roland Göcke
ACM Multimedia6
2023 An Investigation of Video Vision Transformers for Depression Severity Estimation from Facial Video Data
Ghazal Bargshady, Roland Göcke
PSIVT2
2023 Synthesis of a six-bar mechanism for generating knee and ankle motion trajectories using deep generative neural network
Akim Kapsalyamov, Shahid Hussain 0003, Nicholas A. T. Brown, Roland Göcke, Munawar Hayat, Prashant Kumar Jamwal
Eng. Appl. Artif. Intell.4
2023 Interpretation of Depression Detection Models via Feature Selection Methods
abstract
Given the prevalence of depression worldwide and its major impact on society, several studies employed artificial intelligence modelling to automatically detect and assess depression. However, interpretation of these models and cues are rarely discussed in detail in the AI community, but have received increased attention lately. In this study, we aim to analyse the commonly selected features using a proposed framework of several feature selection methods and their effect on the classification results, which will provide an interpretation of the depression detection model. The developed framework aggregates and selects the most promising features for modelling depression detection from 38 feature selection algorithms of different categories. Using three real-world depression datasets, 902 behavioural cues were extracted from speech behaviour, speech prosody, eye movement and head pose. To verify the generalisability of the proposed framework, we applied the entire process to depression datasets individually and when combined. The results from the proposed framework showed that speech behaviour features (e.g. pauses) are the most distinctive features of the depression detection model. From the speech prosody modality, the strongest feature groups were F0, HNR, formants, and MFCC, while for the eye activity modality they were left-right eye movement and gaze direction, and for the head modality it was yaw head movement. Modelling depression detection using the selected features (even though there are only 9 features) outperformed using all features in all the individual and combined datasets. Our feature selection framework did not only provide an interpretation of the model, but was also able to produce a higher accuracy of depression detection with a small number of features in varied datasets. This could help to reduce the processing time needed to extract features and creating the model.
Sharifa Alghowinem, Tom Gedeon, Roland Göcke, Jeffrey F. Cohn, Gordon Parker
IEEE Trans. Affect. Comput.3
2022 Analyzing Group-Level Emotion with Global Alignment Kernel based Approach
abstract
From the perspective of social science, understanding group emotion has become increasingly important for teams to considerably accomplish organizational work. Currently, automatically analyzing the perceived affect of a group of people has been received increasingly interest in affective computing community. The variability in group size makes difficulty for group-level emotion recognition to straightforwardly measure the feature distance of two group-level images. Recent works attempted to resolve the preceding problem by using feature encoding. However, the early works lack of efficiency. To alleviate this problem, this article aims to design a new method to effectively analyze the group behavior from a group-level image. Motivated by time-series kernel approaches explored in dynamic facial expression classification, this article mainly concentrates on global alignment kernel and design support vector machine with the combined global alignment kernels (SVM-CGAK) to better recognize group-level emotion. Specifically, we first propose to use global alignment kernel to explicitly measure the distance of two group-level images. For improving the performance of global alignment kernel, we use the global weight sort scheme based on their spatial relation information to sort the faces from group-level image, making an efficient data structure to the global alignment kernel. With this new global alignment kernel, we construct the backbone of SVM-CGAK, namely, support vector machine with global alignment kernel. Furthermore, considering the challenging environment, we construct two global alignment kernels based on Reisz-based Volume Local Binary Pattern and deep convolutional neural network features, respectively. Lastly, to make the robustness of group-level emotion recognition, we propose SVM-CGAK combining both global alignment kernels with multiple kernel learning approach. It can enhance the discriminative ability of each global alignment kernel. Intensive experiments are conducted on three challenging group-level emotion databases. The experimental results demonstrate that the proposed approach achieves promising performance for group-level emotion recognition compared with the recent state-of-the-art methods.
Xiaohua Huang 0003, Abhinav Dhall, Roland Göcke, Matti Pietikäinen, Guoying Zhao 0001
IEEE Trans. Affect. Comput.3
2021 Micro-Expression Recognition Based On Video Motion Magnification And Pre-Trained Neural Network
abstract
This paper investigates the effects of using video motion magnification methods based on amplitude and phase, respectively, to amplify small facial movements. We hypothesise that this approach will assist in the micro-expression recognition task. To this end, we apply the pre-trained VGGFace2 model with its excellent facial feature capturing ability to transfer learn the magnified micro-expression movement, then encode the spatial information and decode the spatial and temporal information by Bi-LSTM model. Moreover, Grad-CAM is utilised to map the model and visually explain the operating mechanism of the spatio-temporal network. Experiments on the SMIC database confirm that the proposed framework significantly improves the micro-expression recognition rate compared to without video magnification (baseline) and other state-of-the-art methods.
Mengjiong Bai, Roland Göcke, Damith Chandana Herath
ICIP2
2021 Deeply Supervised Discriminative Learning for Adversarial Defense
abstract
Deep neural networks can easily be fooled by an adversary with minuscule perturbations added to an input image. The existing defense techniques suffer greatly under white-box attack settings, where an adversary has full knowledge of the network and can iterate several times to find strong perturbations. We observe that the main reason for the existence of such vulnerabilities is the close proximity of different class samples in the learned feature space of deep models. This allows the model decisions to be completely changed by adding an imperceptible perturbation to the inputs. To counter this, we propose to class-wise disentangle the intermediate feature representations of deep networks, specifically forcing the features for each class to lie inside a convex polytope that is maximally separated from the polytopes of other classes. In this manner, the network is forced to learn distinct and distant decision regions for each class. We observe that this simple constraint on the features greatly enhances the robustness of learned models, even against the strongest white-box attacks, without degrading the classification performance on clean images. We report extensive evaluations in both black-box and white-box attack scenarios and show significant gains in comparison to state-of-the-art defenses.
Aamir Mustafa, Salman Khan 0001, Munawar Hayat, Roland Göcke, Jianbing Shen, Ling Shao 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 Feature Map Augmentation to Improve Rotation Invariance in Convolutional Neural Networks
Dharmendra Sharma 0001, Roland Göcke
ACIVS3
2020 EmotiW 2020: Driver Gaze, Group Emotion, Student Engagement and Physiological Signal based Challenges
abstract
This paper introduces the Eighth Emotion Recognition in the Wild (EmotiW) challenge. EmotiW is a benchmarking effort run as a grand challenge of the 22nd ACM International Conference on Multimodal Interaction 2020. It comprises of four tasks related to automatic human behavior analysis: a) driver gaze prediction; b) audio-visual group-level emotion recognition; c) engagement prediction in the wild; and d) physiological signal based emotion recognition. The motivation of EmotiW is to bring researchers in affective computing, computer vision, speech processing and machine learning to a common platform for evaluating techniques on a test data. We discuss the challenge protocols, databases and their associated baselines.
Abhinav Dhall, Roland Göcke, Tom Gedeon
ICMI3
2019 Automated Measurement of Head Movement Synchrony during Dyadic Depression Severity Interviews
abstract
With few exceptions, most research in automated assessment of depression has considered only the patient's behavior to the exclusion of the therapist's behavior. We investigated the interpersonal coordination (synchrony) of head movement during patient-therapist clinical interviews. Our sample consisted of patients diagnosed with major depressive disorder. They were recorded in clinical interviews (Hamilton Rating Scale for Depression, HRSD) at 7-week intervals over a period of 21 weeks. For each session, patient and therapist 3D head movement was tracked from 2D videos. Head angles in the horizontal (pitch) and vertical (yaw) axes were used to measure head movement. Interpersonal coordination of head movement between patients and therapists was measured using windowed cross-correlation. Patterns of coordination in head movement were investigated using the peak picking algorithm. Changes in head movement coordination over the course of treatment were measured using a hierarchical linear model (HLM). The results indicated a strong effect for patient-therapist head movement synchrony. Within-dyad variability in head movement coordination was found to be higher than between-dyad variability, meaning that differences over time in a dyad were higher as compared to the differences between dyads. Head movement synchrony did not change over the course of treatment with change in depression severity. To the best of our knowledge, this study is the first attempt to analyze the mutual influence of patient-therapist head movement in relation to depression severity.
Shalini Bhatia, Roland Göcke, Zakia Hammal, Jeffrey F. Cohn
FG2
2019 Estimation of Missing Human Body Parts Via Bidirectional LSTM
abstract
In this paper, a bi-directional long-short term memory (LSTM) based approach is proposed for the estimation of missing body parts in a human pose estimation context. Accurate human pose estimation is often a key component for accurate human action and activity recognition. The key idea of our algorithm is to learn the temporal consistencies of the human body poses between previous and subsequent frames. This helps in estimating missing body parts and improves the general smoothness of the pose detection results. The approach acts as a post-processing step after the application of any off-the-shelf body part detector and has been evaluated on the PoseTrack dataset for both validation and testing sequences. The results show consistent improvement in the detection across all body parts.
Ibrahim Radwan, Akshay Asthana, Hafsa Ismail, Byron W. Keating, Roland Göcke
FG5
2019 Adversarial Defense by Restricting the Hidden Space of Deep Neural Networks
abstract
Deep neural networks are vulnerable to adversarial attacks which can fool them by adding minuscule perturbations to the input images. The robustness of existing defenses suffers greatly under white-box attack settings, where an adversary has full knowledge about the network and can iterate several times to find strong perturbations. We observe that the main reason for the existence of such perturbations is the close proximity of different class samples in the learned feature space. This allows model decisions to be totally changed by adding an imperceptible perturbation in the inputs. To counter this, we propose to class-wise disentangle the intermediate feature representations of deep networks. Specifically, we force the features for each class to lie inside a convex polytope that is maximally separated from the polytopes of other classes. In this manner, the network is forced to learn distinct and distant decision regions for each class. We observe that this simple constraint on the features greatly enhances the robustness of learned models, even against the strongest white-box attacks, without degrading the classification performance on clean images. We report extensive evaluations in both black-box and white-box attack scenarios and show significant gains in comparison to state-of-the art defenses.
Aamir Mustafa, Salman Khan 0001, Munawar Hayat, Roland Göcke, Jianbing Shen, Ling Shao 0001
ICCV4
2019 An investigation of linguistic stress and articulatory vowel characteristics for automatic depression classification
Brian Stasak, Julien Epps, Roland Göcke
Comput. Speech Lang.3
2019 Automatic depression classification based on affective read sentences: Opportunities for text-dependent analysis
Brian Stasak, Julien Epps, Roland Göcke
Speech Commun.3
2018 EmotiW 2018: Audio-Video, Student Engagement and Group-Level Affect Prediction
abstract
This paper details the sixth Emotion Recognition in the Wild (EmotiW) challenge. EmotiW 2018 is a grand challenge in the ACM International Conference on Multimodal Interaction 2018, Colarado, USA. The challenge aims at providing a common platform to researchers working in the affective computing community to benchmark their algorithms on 'in the wild' data. This year EmotiW contains three sub-challenges: a) Audio-video based emotion recognition; b) Student engagement prediction; and c) Group-level emotion recognition. The databases, protocols and baselines are discussed in detail.
Abhinav Dhall, Amanjot Kaur, Roland Göcke, Tom Gedeon
ICMI3
2018 Multimodal Depression Detection: Fusion Analysis of Paralinguistic, Head Pose and Eye Gaze Behaviors
abstract
An estimated 350 million people worldwide are affected by depression. Using affective sensing technology, our long-term goal is to develop an objective multimodal system that augments clinical opinion during the diagnosis and monitoring of clinical depression. This paper steps towards developing a classification system-oriented approach, where feature selection, classification and fusion-based experiments are conducted to infer which types of behaviour (verbal and nonverbal) and behaviour combinations can best discriminate between depression and non-depression. Using statistical features extracted from speaking behaviour, eye activity, and head pose, we characterise the behaviour associated with major depression and examine the performance of the classification of individual modalities and when fused. Using a real-world, clinically validated dataset of 30 severely depressed patients and 30 healthy control subjects, a Support Vector Machine is used for classification with several feature selection techniques. Given the statistical nature of the extracted features, feature selection based on T-tests performed better than other methods. Individual modality classification results were considerably higher than chance level (83 percent for speech, 73 percent for eye, and 63 percent for head). Fusing all modalities shows a remarkable improvement compared to unimodal systems, which demonstrates the complementary nature of the modalities. Among the different fusion approaches used here, feature fusion performed best with up to 88 percent average accuracy. We believe that is due to the compatible nature of the extracted statistical features.
Sharifa Alghowinem, Roland Göcke, Michael Wagner 0004, Julien Epps, Matthew Hyett, Gordon Parker, Michael Breakspear
IEEE Trans. Affect. Comput.2
2018 MSMCT: Multi-State Multi-Camera Tracker
abstract
Visual tracking of multiple persons simultaneously is an important tool for group behaviour analysis. In this paper, we demonstrate that multi-target tracking in a network of non-overlapping cameras can be formulated in a framework, where the association among all given target hypotheses both within and between cameras is performed simultaneously. Our approach helps to overcome the fragility of multi-camera-based tracking, where the performance relies on the single-camera tracking results obtained at input level. In particular, we formulate an estimation of the target states as a multi-state graph optimization problem, in which the likelihood of each target hypothesis belonging to different identities is modeled. In addition, we learn the target-specific model to improve the similarity measure among targets based on the appearance cues. We also handle the occluded targets when there is no reliable evidence for the target's presence and each target trajectory is expected to be fragmented into multiple tracks. An iterative procedure is proposed to solve the optimization problem, resulting in final trajectories that reveal the true states of the targets. The performance of the proposed approach has been extensively evaluated on challenging multi-camera non-overlapping tracking data sets, in which many difficulties, such as occlusion, viewpoint, and illumination variation, are present. The results of systematic experiments conducted on a large set of sequences show that the proposed approach outperforms several state-of-the-art trackers.
Behzad Bozorgtabar, Roland Göcke
IEEE Trans. Circuits Syst. Video Technol.2
2018 Multimodal Framework for Analyzing the Affect of a Group of People
abstract
With the advances in multimedia and the world wide web, users upload millions of images and videos everyone on social networking platforms on the Internet. From the perspective of automatic human behavior understanding, it is of interest to analyze and model the affects that are exhibited by groups of people who are participating in social events in these images. However, the analysis of the affect that is expressed by multiple people is challenging due to the varied indoor and outdoor settings. Recently, a few interesting works have investigated face-based group-level emotion recognition (GER). In this paper, we propose a multimodal framework for enhancing the affective analysis ability of GER in challenging environments. Specifically, for encoding a person's information in a group-level image, we first propose an information aggregation method for generating feature descriptions of face, upper body, and scene. Later, we revisit localized multiple kernel learning for fusing face, upper body, and scene information for GER against challenging environments. Intensive experiments are performed on two challenging group-level emotion databases (HAPPEI and GAFF) to investigate the roles of the face, upper body, scene information, and the multimodal framework. Experimental results demonstrate that the multimodal framework achieves promising performance for GER.
Xiaohua Huang 0003, Abhinav Dhall, Roland Göcke, Matti Pietikäinen, Guoying Zhao 0001
IEEE Trans. Multim.3
2017 Heart rate estimation from facial videos for depression analysis
abstract
Automated facial video analysis is useful in numerous health care applications. For example, spatio-temporal analysis of such videos has been previously done for assisting clinicians in the diagnosis of depression. Physiological measures, such as an individual's heart rate, provide very important cues to understand a person's mental health. Unobtrusively estimated heart rate has not been previously used to analyse individuals' mental health. In this paper, we automatically estimate heart rate activity from facial videos. We then study the association of the estimated heart rate activity with the person's mental health, as diagnosed by clinicians. Specifically, from the heart rate activity in response to watching different movies, we classify individuals as either depressed or healthy. The efficacy of the proposed scheme is demonstrated by experimental evaluations on a clinically validated dataset. Our results suggest unobtrusively estimated heart rate to be very effective for depression analysis.
Aamir Mustafa, Shalini Bhatia, Munawar Hayat, Roland Göcke
ACII4
2017 Joint Registration and Representation Learning for Unconstrained Face Identification
abstract
Recent advances in deep learning have resulted in human-level performances on popular unconstrained face datasets including Labeled Faces in the Wild and YouTube Faces. To further advance research, IJB-A benchmark was recently introduced with more challenges especially in the form of extreme head poses. Registration of such faces is quite demanding and often requires laborious procedures like facial landmark localization. In this paper, we propose a Convolutional Neural Networks based data-driven approach which learns to simultaneously register and represent faces. We validate the proposed scheme on template based unconstrained face identification. Here, a template contains multiple media in the form of images and video frames. Unlike existing methods which synthesize all template media information at feature level, we propose to keep the template media intact. Instead, we represent gallery templates by their trained one-vs-rest discriminative models and then employ a Bayesian strategy which optimally fuses decisions of all medias in a query template. We demonstrate the efficacy of the proposed scheme on IJB-A, YouTube Celebrities and COX datasets where our approach achieves significant relative performance boosts of 3.6%, 21.6% and 12.8% respectively.
Munawar Hayat, Salman Khan 0001, Naoufel Werghi, Roland Göcke
CVPR4
2017 A Video-Based Facial Behaviour Analysis Approach to Melancholia
abstract
Recent years have seen a lot of activity in affective computing for automated analysis of depression. However, no research has so far directly evaluated the performance of facial behavioural analysis methods in classifying different subtypes of depression such as melancholia. The mental state assessment of a mood disorder depends largely on appearance, behaviour, speech, thought, perception, mood and facial affect. Mood and facial affect mainly contribute to distinguishing melancholia from nonmelancholia. These are assessed by clinicians, and hence vulnerable to subjective judgement. As a result, clinical assessment alone may not accurately capture the presence or absence of specific disorders such as melancholia, a distressing condition whose presence has important treatment implications. Melancholia is characterised by severe anhedonia and psychomotor disturbance, which can be a mix of motor retardation with periods of superimposed agitation. To the best of our knowledge, this study is the first attempt to perform facial behavioural analysis to disambiguate melancholia from non-melancholia and healthy controls on the basis of facial behavioural characteristics. We report the sensitivity and specificity of classification in depressive subtypes. These results serve as a baseline for more fine-grained depression classification and analysis.
Shalini Bhatia, Munawar Hayat, Michael Breakspear, Gordon Parker, Roland Göcke
FG5
2017 Human Postural Sway Estimation from Noisy Observations
abstract
Postural sway is a reflection of brain signals that are generated to control a person’s balance. During the process of ageing, the postural sway changes, which increases the likelihood of a fall. Thus far, expensive specialist equipment is required, such as a force plate, in order to detect such changes over time, which makes the process costly and impractical. Our long-term goal is to investigate the use of inexpensive, everyday video technology as an alternative. This paper describes a study that establishes a 3-way correlation between the clinical gold standard (force plate), a highly accurate multi-camera 3D video tracking system (Vicon) and a standard RGB video camera. To this end, a dataset of 18 subjects performing the BESS balance test on the force plate was recorded, while simultaneously recording the 3D Vicon data, and the RGB video camera data. Then, using Gaussian process regression and a recurrent neural network, models were built to predict the lateral postural sway in the force plate data from the RGB video data. The predicted results show high correlation with the actual force plate signals, which supports the hypothesis that lateral postural sway can be accurately predicted from video data alone. Detecting changes to a person’s postural sway can be used to improve elderly people’s life by monitoring the likelihood of a fall and detecting its increase well before a fall occurs, so that countermeasures (e.g. exercises) can be put in place to prevent falls occurring.
Hafsa Ismail, Ibrahim Radwan, Hanna Suominen, Gordon Waddington, Roland Göcke
FG5
2017 A multimodal system to characterise melancholia: cascaded bag of words approach
abstract
Recent years have seen a lot of activity in affective computing for automated analysis of depression. However, no research has so far proposed a multimodal system for classifying different subtypes of depression such as melancholia. The mental state assessment of a mood disorder depends primarily on appearance, behaviour, speech, thought, perception, mood and facial affect. Mood and facial affect mainly contribute to distinguishing melancholia from non-melancholia. These are assessed by clinicians, and hence vulnerable to subjective judgement. As a result, clinical assessment alone may not accurately capture the presence or absence of specific disorders such as melancholia, a distressing condition whose presence has important treatment implications. Melancholia is characterised by severe anhedonia and psychomotor disturbance, which can be a combination of motor retardation with periods of superimposed agitation. Psychomotor disturbance can be sensed in both face and voice. To the best of our knowledge, this study is the first attempt to propose a multimodal system to differentiate melancholia from non-melancholia and healthy controls. We report the sensitivity and specificity of classification in depressive subtypes.
Shalini Bhatia, Munawar Hayat, Roland Göcke
ICMI3
2017 From individual to group-level emotion recognition: EmotiW 5.0
abstract
Research in automatic affect recognition has come a long way. This paper describes the fifth Emotion Recognition in the Wild (EmotiW) challenge 2017. EmotiW aims at providing a common benchmarking platform for researchers working on different aspects of affective computing. This year there are two sub-challenges: a) Audio-video emotion recognition and b) group-level emotion recognition. These challenges are based on the acted facial expressions in the wild and group affect databases, respectively. The particular focus of the challenge is to evaluate method in `in the wild' settings. `In the wild' here is used to describe the various environments represented in the images and videos, which represent real-world (not lab like) scenarios. The baseline, data, protocol of the two challenges and the challenge participation are discussed in detail in this paper.
Abhinav Dhall, Roland Göcke, Shreya Ghosh 0001, Jyoti Joshi, Jesse Hoey, Tom Gedeon
ICMI2
2017 Elicitation Design for Acoustic Depression Classification: An Investigation of Articulation Effort, Linguistic Complexity, and Word Affect
Brian Stasak, Julien Epps, Roland Göcke
INTERSPEECH3
2016 Extending Long Short-Term Memory for Multi-View Structured Learning
Shyam Sundar Rajagopalan, Louis-Philippe Morency, Tadas Baltrusaitis, Roland Göcke
ECCV (7)4
2016 Emotion recognition in the wild challenge 2016
abstract
The fourth Emotion Recognition in the Wild (EmotiW) challenge is a grand challenge in the ACM International Conference on Multimodal Interaction 2016, Tokyo. EmotiW is a series of benchmarking and competition effort for researchers working in the area of automatic emotion recognition in the wild. The fourth EmotiW has two sub-challenges: Video based emotion recognition (VReco) and Group-level emotion recognition (GReco). The VReco sub-challenge is being run for the fourth time and GReco is a new sub-challenge this year.
Abhinav Dhall, Roland Göcke, Jyoti Joshi, Tom Gedeon
ICMI2
2016 EmotiW 2016: video and group-level emotion recognition challenges
abstract
This paper discusses the baseline for the Emotion Recognition in the Wild (EmotiW) 2016 challenge. Continuing on the theme of automatic affect recognition `in the wild', the EmotiW challenge 2016 consists of two sub-challenges: an audio-video based emotion and a new group-based emotion recognition sub-challenges. The audio-video based sub-challenge is based on the Acted Facial Expressions in the Wild (AFEW) database. The group-based emotion recognition sub-challenge is based on the Happy People Images (HAPPEI) database. We describe the data, baseline method, challenge protocols and the challenge results. A total of 22 and 7 teams participated in the audio-video based emotion and group-based emotion sub-challenges, respectively.
Abhinav Dhall, Roland Göcke, Jyoti Joshi, Jesse Hoey, Tom Gedeon
ICMI2
2016 Cross-Cultural Depression Recognition from Vocal Biomarkers
abstract
No studies have investigated cross-cultural and cross-language characteristics of depressed speech. We investigated the generalisability of a vocal biomarker-based approach to depression detection in clinical interviews recorded in three countries (Australia, the USA and Germany), two languages (German and English) and different accents (Australian and American). Several approaches to training and testing within and between datasets were evaluated. Using the same experimental protocol separately within each dataset, (cross-classification) accuracy was high.combining datasets, high accuracy was high again and consistent across language, recording environment, and culture. Training and testing between datasets, however, attenuated accuracy. These finding emphasize the importance of heterogeneous training sets for robust depression detection.
Sharifa Alghowinem, Roland Göcke, Julien Epps, Michael Wagner 0004, Jeffrey F. Cohn
INTERSPEECH2
2016 An Investigation of Emotional Speech in Depression Classification
Brian Stasak, Julien Epps, Nicholas Cummins, Roland Göcke
INTERSPEECH4
2016 Efficient multi-target tracking via discovering dense subgraphs
Behzad Bozorgtabar, Roland Göcke
Comput. Vis. Image Underst.2
2016 Dimensionality reduction of Fisher vectors for human action recognition
abstract
Automatic analysis of human behaviour in large collections of videos is rapidly gaining interest, even more so with the advent of file sharing sites such as YouTube. From one perspective, it can be observed that the size of feature vectors used for human action recognition from videos has been increasing enormously in the last five years, in the order of ∼100–500K. One possible reason might be the growing number of action classes/videos and hence the requirement of discriminating features (that usually end up to be higher‐dimensional for larger databases). In this study, the authors review and investigate feature projection as a means to reduce the dimensions of the high‐dimensional feature vectors and show their effectiveness in terms of performance. They hypothesise that dimensionality reduction techniques often unearth latent structures in the feature space and are effective in applications such as the fusion of high‐dimensional features of different types; and action recognition in untrimmed videos. They conduct all the authors’ experiments using a Bag‐of‐Words framework for consistency and results are presented on large class benchmark databases such as the HMDB51 and UCF101 datasets.
O. V. Ramana Murthy, Roland Göcke
IET Comput. Vis.2
2015 A temporally piece-wise fisher vector approach for depression analysis
abstract
Depression and other mood disorders are common, disabling disorders with a profound impact on individuals and families. Inspite of its high prevalence, it is easily missed during the early stages. Automatic depression analysis has become a very active field of research in the affective computing community in the past few years. This paper presents a framework for depression analysis based on unimodal visual cues. Temporally piece-wise Fisher Vectors (FV) are computed on temporal segments. As a low-level feature, block-wise Local Binary Pattern-Three Orthogonal Planes descriptors are computed. Statistical aggregation techniques are analysed and compared for creating a discriminative representative for a video sample. The paper explores the strength of FV in representing temporal segments in a spontaneous clinical data. This creates a meaningful representation of the facial dynamics in a temporal segment. The experiments are conducted on the Audio Video Emotion Challenge (AVEC) 2014 German speaking depression database. The superior results of the proposed framework show the effectiveness of the technique as compared to the current state-of-art.
Abhinav Dhall, Roland Göcke
ACII2
2015 Riesz-based Volume Local Binary Pattern and A Novel Group Expression Model for Group Happiness Intensity Analysis
abstract
Automatic emotion analysis and understanding has received much attention over the years in affective computing. Recently, there are increasing interests in inferring the emotional intensity of a group of people. For group emotional intensity analysis, feature extraction and group expression model are two critical issues. In this paper, we propose a new method to estimate the happiness intensity of a group of people in an image. Firstly, we combine the Riesz transform and the local binary pattern descriptor, named Riesz-based volume local binary pattern, which considers neighbouring changes not only in the spatial domain of a face but also along the different Riesz faces. Secondly, we exploit the continuous conditional random fields for constructing a new group expression model, which considers global and local attributes. Intensive experiments are performed on three challenging facial expression databases to evaluate the novel feature. Furthermore, experiments are conducted on the HAPPEI database to evaluate the new group expression model with the new feature. Our experimental results demonstrate the promising performance for group happiness intensity analysis.
Xiaohua Huang 0003, Abhinav Dhall, Guoying Zhao 0001, Roland Göcke, Matti Pietikäinen
BMVC4
2015 Multi-level action detection via learning latent structure
abstract
Detecting actions in videos is still a demanding task due to large intra-class variation caused by varying pose, motion and scales. Conventional approaches use a Bag-of-Words model in the form of space-time motion feature pooling followed by learning a classifier. However, since the informative body parts motion only appear in specific regions of the body, these methods have limited capability. In this paper, we seek to learn a model of the interaction among regions of interest via a graph structure. We first discover several space-time video segments representing persistent moving body parts observed sparsely in video. Then, via learning the hidden graph structure (a subset of the graph), we identify both spatial and temporal relations between the subsets of these segments. In order to seize the more discriminative motion patterns and handle different interactions between body parts from simple to composite action, we present a multi-level action model representation. Consequently, for action classification, the classifier learned through each action model labels the test video based on the action model that gives the highest probability score. Experiments on challenging datasets, such as MSR II and UCF-Sports including complex motions and dynamic backgrounds, demonstrate the effectiveness of the proposed approach that outperforms state-of-the-art methods in this context.
Behzad Bozorgtabar, Roland Göcke
ICIP2
2015 Video and Image based Emotion Recognition Challenges in the Wild: EmotiW 2015
abstract
The third Emotion Recognition in the Wild (EmotiW) challenge 2015 consists of an audio-video based emotion and static image based facial expression classification sub-challenges, which mimics real-world conditions. The two sub-challenges are based on the Acted Facial Expression in the Wild (AFEW) 5.0 and the Static Facial Expression in the Wild (SFEW) 2.0 databases, respectively. The paper describes the data, baseline method, challenge protocol and the challenge results. A total of 12 and 17 teams participated in the video based emotion and image based expression sub-challenges, respectively.
Abhinav Dhall, O. V. Ramana Murthy, Roland Göcke, Jyoti Joshi, Tom Gedeon
ICMI3
2015 Ordered trajectories for human action recognition with large number of classes
O. V. Ramana Murthy, Roland Göcke
Image Vis. Comput.2
2015 Automatic Group Happiness Intensity Analysis
abstract
The recent advancement of social media has given users a platform to socially engage and interact with a larger population. Millions of images and videos are being uploaded everyday by users on the web from different events and social gatherings. There is an increasing interest in designing systems capable of understanding human manifestations of emotional attributes and affective displays. As images and videos from social events generally contain multiple subjects, it is an essential step to study these groups of people. In this paper, we study the problem of happiness intensity analysis of a group of people in an image using facial expression analysis. A user perception study is conducted to understand various attributes, which affect a person's perception of the happiness intensity of a group. We identify the challenges in developing an automatic mood analysis system and propose three models based on the attributes in the study. An `in the wild' image-based database is collected. To validate the methods, both quantitative and qualitative experiments are performed and applied to the problem of shot selection, event summarisation and album creation. The experiments show that the global and local attributes defined in the paper provide useful information for theme expression analysis, with results close to human perception results.
Abhinav Dhall, Roland Göcke, Tom Gedeon
IEEE Trans. Affect. Comput.2
2014 Enhanced Laplacian Group Sparse Learning with Lifespan Outlier Rejection for Visual Tracking
Behzad Bozorgtabar, Roland Göcke
ACCV (5)2
2014 Joint sparsity-based robust visual tracking
abstract
In this paper, we propose a new object tracking in a particle filter framework utilising a joint sparsity-based model. Based on the observation that a target can be reconstructed from several templates that are updated dynamically, we jointly analyse the representation of the particles under a single regression framework and with the shared underlying structure. Two convex regularisations are combined and used in our model to enable sparsity as well as facilitate coupling information between particles. Unlike the previous methods that consider a model commonality between particles or regard them as independent tasks, we simultaneously take into account a structure inducing norm and an outlier detecting norm. Such a formulation is shown to be more flexible in terms of handling various types of challenges including occlusion and cluttered background. To derive the optimal solution efficiently, we propose to use a Preconditioned Conjugate Gradient method, which is computationally affordable for high-dimensional data. Furthermore, an online updating procedure scheme is included in the dictionary learning, which makes the proposed tracker less vulnerable to outliers. Experiments on challenging video sequences demonstrate the robustness of the proposed approach to handling occlusion, pose and illumination variation and outperform state-of-the-art trackers in tracking accuracy.
Behzad Bozorgtabar, Roland Göcke
ICIP2
2014 Dense body part trajectories for human action recognition
abstract
Several techniques have been proposed for human action recognition from videos. It has been observed that incorporating mid-level viz. human body and/or high-level information viz. pose estimation in the computation of low-level features viz. trajectories yields the best performance in action recognition where full body is presumed. However, in datasets with a large number of classes, where the full body may not be visible at all times, incorporating such mid- and high-level information is unexplored. Moreover, changes and developments in any stage will require a recompute of all low-level features. We decouple mid-level and low-level feature computation and study on benchmark action recognition datasets such as UCF50, UCF101 and HMDB51, containing the largest number of action classes to date. Further, we employ a part-based model for human body part detection in frames statically, thus also investigating classes where the full body is not present. We also track dense regions around the detected human body parts by Hungarian particle linking, thus minimising most of the wrongly detected body parts and enriching the mid-level information.
O. V. Ramana Murthy, Ibrahim Radwan, Roland Göcke
ICIP3
2014 Detecting self-stimulatory behaviours for autism diagnosis
abstract
Autism Spectrum Disorders (ASD), often referred to as autism, are neurological disorders characterised by deficits in cognitive skills, social and communicative behaviours. A common way of diagnosing ASD is by studying behavioural cues expressed by the children. An algorithm for detecting three types of self-stimulatory behaviours from publicly available unconstrained videos is proposed here. The child's body is tracked in the video by a careful selection of poselet bounding box predictions using a nearest neighbour algorithm. A global motion descriptor - Histogram of Dominant Motions (HDM) - is computed using the dominant motion flow in the detected body regions. The motion model built using this descriptor is used for detecting the self-stimulatory behaviours. Experiments conducted on the recently released unconstrained SSBD video dataset show significant improvement in detection accuracy over the baseline approach. The robustness of the method is validated using benchmark action recognition datasets. The proposed poselet bounding box selection algorithm is validated against the ground truth annotation data provided with the UCF101 dataset.
Shyam Sundar Rajagopalan, Roland Göcke
ICIP2
2014 Emotion Recognition In The Wild Challenge 2014: Baseline, Data and Protocol
abstract
The Second Emotion Recognition In The Wild Challenge (EmotiW) 2014 consists of an audio-video based emotion classification challenge, which mimics the real-world conditions. Traditionally, emotion recognition has been performed on data captured in constrained lab-controlled like environment. While this data was a good starting point, such lab controlled data poorly represents the environment and conditions faced in real-world situations. With the exponential increase in the number of video clips being uploaded online, it is worthwhile to explore the performance of emotion recognition methods that work `in the wild'. The goal of this Grand Challenge is to carry forward the common platform defined during EmotiW 2013, for evaluation of emotion recognition methods in real-world conditions. The database in the 2014 challenge is the Acted Facial Expression In Wild (AFEW) 4.0, which has been collected from movies showing close-to-real-world conditions. The paper describes the data partitions, the baseline method and the experimental protocol.
Abhinav Dhall, Roland Göcke, Jyoti Joshi, Karan Sikka, Tom Gedeon
ICMI2
2014 Automatic Prediction of Perceived Traits Using Visual Cues under Varied Situational Context
abstract
Automatic assessment of human personality traits is a non-trivial problem, especially when perception is marked over a fairly short duration of time. In this study, thin slices of behavioral data are analyzed. Perceived physical and behavioral traits are assessed by external observers (raters). Along with the big-five personality trait model, four new traits are introduced and assessed in this work. The relationship between various traits is investigated to obtain a better understanding of observer perception and assessment. Perception change is also considered when participants interact with several virtual characters each with a distinct emotional style. Encapsulating these observations and analysis, an automated system is proposed by firstly computing low level visual features. Using these features a separate model is trained for each trait and performance is evaluated. Further, a weighted model based on rater credibility is proposed to address observer biases. Experimental results indicate that a weighted model show major improvement for automatic prediction of perceived physical and behavioral traits.
Jyoti Joshi, Hatice Gunes, Roland Göcke
ICPR3
2014 A discriminative parts based model approach for fiducial points free and shape constrained head pose normalisation in the wild
abstract
Continuous Confidence Map Based Normalisation: While continuous head pose normalisation is not the goal of this paper, we demonstrate as a proof of concept that it is possible to extend the current method for continuous head pose normalisation. For dealing with faces in videos [1], continuous head pose normalisation is required. [2] argue that the appearance of a part does not changes with a subtle pose change, therefore a detector for part i in pose angle p can be shared for the same part i for a pose angle p + δ. Further experiments in [2] showed that sharing based models and independent model have comparable performance. However, sharing based models are faster upto ten times as compared to the independent models [2]. The confidence maps based methods (CM-HPNPSand CM-HPNPI) can be extended from discrete to continuous by sharing part-specific regression models R, which are shared among neighboring pose angles.
Abhinav Dhall, Karan Sikka, Gwen Littlewort, Roland Göcke, Marian Stewart Bartlett
WACV4
2014 A discriminative parts based model approach for fiducial points free and shape constrained head pose normalisation in the wild
abstract
This paper proposes a method for parts-based view-invariant head pose normalisation, which works well even in difficult real-world conditions. Handling pose is a classical problem in facial analysis. Recently, parts-based models have shown promising performance for facial landmark points detection `in the wild'. Leveraging on the success of these models, the proposed data-driven regression framework computes a constrained normalised virtual frontal head pose. The response maps of a discriminatively trained part detector are used as texture information. These sparse texture maps are projected from non-frontal to frontal pose using block-wise structured regression. Finally, a facial kinematic shape constraint is achieved by applying a shape model. The advantages of the proposed approach are: a) no explicit dependence on the outputs of a facial parts detector and, thus, avoiding any error propagation owing to their failure; (b) the application of a shape prior on the reconstructed frontal maps provides an anatomically constrained facial shape; and c) modelling head pose as a mixture-of-parts model allows the framework to work without any prior pose information. Experiments are performed on the Multi-PIE and the `in the wild' SFEW databases. The results demonstrate the effectiveness of the proposed method.
Abhinav Dhall, Karan Sikka, Gwen Littlewort, Roland Göcke, Marian Stewart Bartlett
WACV4
2013 Head Pose and Movement Analysis as an Indicator of Depression
abstract
Depression is a common and disabling mental health disorder, which impacts not only on the sufferer but also their families, friends and the economy overall. Our ultimate aim is to develop an automatic objective affective sensing system that supports clinicians in their diagnosis and monitoring of clinical depression. Here, we analyse the performance of head pose and movement features extracted from face videos using a 3D face model projected on a 2D Active Appearance Model (AAM). In a binary classification task (depressed vs. non-depressed), we modelled low-level and statistical functional features for an SVM classifier using real-world clinically validated data. Although the head pose and movement would be used as a complementary cue in detecting depression in practice, their recognition rate was impressive on its own, giving 71.2% on average, which illustrates that head pose and movement hold effective cues in diagnosing depression. When expressing positive and negative emotions, recognising depression using positive emotions was more accurate than using negative emotions. We conclude that positive emotions are expressed less in depressed subjects at all times, and that negative emotions have less discriminatory power than positive emotions in detecting depression. Analysing the functional features statistically illustrates several behaviour patterns for depressed subjects: (1) slower head movements, (2) less change of head position, (3) longer duration of looking to the right, (4) longer duration of looking down, which may indicate fatigue and eye contact avoidance. We conclude that head movements are significantly different between depressed patients and healthy subjects, and could be used as a complementary cue.
Sharifa Alghowinem, Roland Göcke, Michael Wagner 0004, Gordon Parker, Michael Breakspear
ACII2
2013 Relative Body Parts Movement for Automatic Depression Analysis
abstract
In this paper, a human body part motion analysis based approach is proposed for depression analysis. Depression is a serious psychological disorder. The absence of an (automated) objective diagnostic aid for depression leads to a range of subjective biases in initial diagnosis and ongoing monitoring. Researchers in the affective computing community have approached the depression detection problem using facial dynamics and vocal prosody. Recent works in affective computing have shown the significance of body pose and motion in analysing the psychological state of a person. Inspired by these works, we explore a body parts motion based approach. Relative orientation and radius are computed for the body parts detected using the pictorial structures framework. A histogram of relative parts motion is computed. To analyse the motion on a holistic level, space-time interest points are computed and a bag of words framework is learnt. The two histograms are fused and a support vector machine classifier is trained. The experiments conducted on a clinical database, prove the effectiveness of the proposed method.
Jyoti Joshi, Abhinav Dhall, Roland Göcke, Jeffrey F. Cohn
ACII3
2013 Modeling Stress Using Thermal Facial Patterns: A Spatio-temporal Approach
abstract
Stress is a serious concern facing our world today, motivating the development of better objective understanding using non-intrusive means for stress recognition. The aim for the work was to use thermal imaging of facial regions to detect stress automatically. The work uses facial regions captured in videos in thermal (TS) and visible (VS) spectrums and introduces our database ANU StressDB. It describes the experiment conducted for acquiring TS and VS videos of observers of stressed and not-stressed films for the ANU StressDB. Further, it presents an application of local binary patterns on three orthogonal planes (LBP-TOP) on VS and TS videos for stress recognition. It proposes a novel method to capture dynamic thermal patterns in histograms (HDTP) to utilize thermal and spatio-temporal characteristics associated in TS videos. Individual-independent support vector machine classifiers were developed for stress recognition. Results show that a fusion of facial patterns from VS and TS videos produced significantly better stress recognition rates than patterns from only VS or TS videos with p <; 0.01. The best stress recognition rate was 72% and it was obtained from HDTP features fused with LBP-TOP features for TS and VS videos respectively.
Nandita Sharma, Abhinav Dhall, Tom Gedeon, Roland Göcke
ACII4
2013 Detecting depression: A comparison between spontaneous and read speech
abstract
Major depressive disorders are mental disorders of high prevalence, leading to a high impact on individuals, their families, society and the economy. In order to assist clinicians to better diagnose depression, we investigate an objective diagnostic aid using affective sensing technology with a focus on acoustic features. In this paper, we hypothesise that (1) classifying the general characteristics of clinical depression using spontaneous speech will give better results than using read speech, (2) that there are some acoustic features that are robust and would give good classification results in both spontaneous and read, and (3) that a `thin-slicing' approach using smaller parts of the speech data will perform similarly if not better than using the whole speech data. By examining and comparing recognition results for acoustic features on a real-world clinical dataset of 30 depressed and 30 control subjects using SVM for classification and a leave-one-out cross-validation scheme, we found that spontaneous speech has more variability, which increases the recognition rate of depression. We also found that jitter, shimmer, energy and loudness feature groups are robust in characterising both read and spontaneous depressive speech. Remarkably, thin-slicing the read speech, using either the beginning of each sentence or the first few sentences performs better than using all reading task data.
Sharifa Alghowinem, Roland Göcke, Michael Wagner 0004, Julien Epps, Michael Breakspear, Gordon Parker
ICASSP2
2013 A comparative study of different classifiers for detecting depression from spontaneous speech
abstract
Accurate detection of depression from spontaneous speech could lead to an objective diagnostic aid to assist clinicians to better diagnose depression. Little thought has been given so far to which classifier performs best for this task. In this study, using a 60-subject real-world clinically validated dataset, we compare three popular classifiers from the affective computing literature - Gaussian Mixture Models (GMM), Support Vector Machines (SVM) and Multilayer Perceptron neural networks (MLP) - as well as the recently proposed Hierarchical Fuzzy Signature (HFS) classifier. Among these, a hybrid classifier using GMM models and SVM gave the best overall classification results. Comparing feature, score, and decision fusion, score fusion performed better for GMM, HFS and MLP, while decision fusion worked best for SVM (both for raw data and GMM models). Feature fusion performed worse than other fusion methods in this study. We found that loudness, root mean square, and intensity were the voice features that performed best to detect depression in this dataset.
Sharifa Alghowinem, Roland Göcke, Michael Wagner 0004, Julien Epps, Tom Gedeon, Michael Breakspear, Gordon Parker
ICASSP2
2013 Monocular Image 3D Human Pose Estimation under Self-Occlusion
abstract
In this paper, an automatic approach for 3D pose reconstruction from a single image is proposed. The presence of human body articulation, hallucinated parts and cluttered background leads to ambiguity during the pose inference, which makes the problem non-trivial. Researchers have explored various methods based on motion and shading in order to reduce the ambiguity and reconstruct the 3D pose. The key idea of our algorithm is to impose both kinematic and orientation constraints. The former is imposed by projecting a 3D model onto the input image and pruning the parts, which are incompatible with the anthropomorphism. The latter is applied by creating synthetic views via regressing the input view to multiple oriented views. After applying the constraints, the 3D model is projected onto the initial and synthetic views, which further reduces the ambiguity. Finally, we borrow the direction of the unambiguous parts from the synthetic views to the initial one, which results in the 3D pose. Quantitative experiments are performed on the Human Eva-I dataset and qualitatively on unconstrained images from the Image Parse dataset. The results show the robustness of the proposed approach to accurately reconstruct the 3D pose form a single image.
Ibrahim Radwan, Abhinav Dhall, Roland Göcke
ICCV3
2013 Eye movement analysis for depression detection
abstract
Depression is a common and disabling mental health disorder, which impacts not only on the sufferer but also on their families, friends and the economy overall. Despite its high prevalence, current diagnosis relies almost exclusively on patient self-report and clinical opinion, leading to a number of subjective biases. Our aim is to develop an objective affective sensing system that supports clinicians in their diagnosis and monitoring of clinical depression. In this paper, we analyse the performance of eye movement features extracted from face videos using Active Appearance Models for a binary classification task (depressed vs. non-depressed). We find that eye movement low-level features gave 70% accuracy using a hybrid classifier of Gaussian Mixture Models and Support Vector Machines, and 75% accuracy when using statistical measures with SVM classifiers over the entire interview. We also investigate differences while expressing positive and negative emotions, as well as the classification performance in gender-dependent versus gender-independent modes. Interestingly, even though the blinking rate was not significantly different between depressed and healthy controls, we find that the average distance between the eyelids (`eye opening') was significantly smaller and the average duration of blinks significantly longer in depressed subjects, which might be an indication of fatigue or eye contact avoidance.
Sharifa Alghowinem, Roland Göcke, Michael Wagner 0004, Gordon Parker, Michael Breakspear
ICIP2
2013 Emotion recognition in the wild challenge (EmotiW) challenge and workshop summary
abstract
The Emotion Recognition In The Wild Challenge and Workshop (EmotiW) 2013 Grand Challenge consists of an audio-video based emotion classification challenge, which mimics real-world conditions. In total, 27 teams participated in the challenge. The database in the 2013 challenge is the Acted Facial Expression in the Wild (AFEW), which has been collected from movies showing close-to-real-world conditions.
Abhinav Dhall, Roland Göcke, Jyoti Joshi, Michael Wagner 0004, Tom Gedeon
ICMI2
2013 Emotion recognition in the wild challenge 2013
abstract
Emotion recognition is a very active field of research. The Emotion Recognition In The Wild Challenge and Workshop (EmotiW) 2013 Grand Challenge consists of an audio-video based emotion classification challenges, which mimics real-world conditions. Traditionally, emotion recognition has been performed on laboratory controlled data. While undoubtedly worthwhile at the time, such laboratory controlled data poorly represents the environment and conditions faced in real-world situations. The goal of this Grand Challenge is to define a common platform for evaluation of emotion recognition methods in real-world conditions. The database in the 2013 challenge is the Acted Facial Expression in the Wild (AFEW), which has been collected from movies showing close-to-real-world conditions.
Abhinav Dhall, Roland Göcke, Jyoti Joshi, Michael Wagner 0004, Tom Gedeon
ICMI2
2013 Adaptive Multiple Component Metric Learning for Robust Visual Tracking
Behzad Bozorgtabar, Roland Göcke
ICONIP (3)2
2013 Characterising depressed speech for classification
abstract
Depression is a serious psychiatric disorder that affects mood, thoughts, and the ability to function in everyday life. This pa-per investigates the characteristics of depressed speech for the purpose of automatic classification by analysing the effect of different speech features on the classification results. We anal-ysed voiced, unvoiced and mixed speech in order to gain a better understanding of depressed speech and to bridge the gap be-tween physiological and affective computing studies. This un-derstanding may ultimately lead to an objective affective sens-ing system that supports clinicians in their diagnosis and mon-itoring of clinical depression. The characteristics of depressed speech were statistically analysed using ANOVA and linked to their classification results using GMM and SVM. Features were extracted and classified over speech utterances of 30 clinically depressed patients against 30 controls (both gender-matched) in a speaker-independent manner. Most feature classification re-sults were consistent with their statistical characteristics, pro-viding a link between physiological and affective computing studies. The classification results from low-level features were slightly better than the statistical functional features, which in-dicates a loss of information in the latter. We found that both mixed and unvoiced speech were as useful in detecting depres-sion as voiced speech, if not better. Index Terms: depression, speech characteristics, mood classi-fication
Sharifa Alghowinem, Roland Göcke, Michael Wagner 0004, Julien Epps, Gordon Parker, Michael Breakspear
INTERSPEECH2
2013 Modeling spectral variability for the classification of depressed speech
abstract
Quantifying how the spectral content of speech relates to changes in mental state may be crucial in building an objective speech-based depression classification system with clinical utility. This paper investigates the hypothesis that important depression based information can be captured within the covariance structure of a Gaussian Mixture Model (GMM) of recorded speech. Significant negative correlations found between a speaker’s average weighted variance- a GMM-based indicator of speaker variability- and their level of depression support this hypothesis. Further evidence is provided by the comparison of classification accuracies from seven different GMM-UBM systems, each formed by varying different parameter combinations during MAP adaption. This analysis shows that variance-only adaptation either outperforms or matches the de facto standard mean-only adaptation when classifying both the presence and severity of depression. This result is perhaps the first of its kind seen in GMM-UBM speech classification.
Nicholas Cummins, Julien Epps, Vidhyasaharan Sethu, Michael Breakspear, Roland Göcke
INTERSPEECH5
2013 R-norm: improving inter-speaker variability modelling at the score level via regression score normalisation
abstract
This paper presents a new method of score post-processing which utilises previously hidden relationships among client models and test probes that are found within the scores produced by an automatic speaker recognition system. We suggest the name r-Norm (for Regression Normalisation) for the method, which can be viewed as both a score normalisation process and as a novel and improved modelling technique of inter-speaker variability. The key component of the method lies in learning a regression model between development data scores and an 'ideal' score matrix, which can either be derived from clean data or created synthetically. To generate scores for experimental validation of the proposed idea we perform a classic GMM-UBM experiment employing mel-cepstral features on the 1sp-female task of the NIST 2003 SRE corpus. Comparisons of the r-Norm results are made with standard score postprocessing/ normalisation methods t-Norm and z-Norm. The r- Norm method is shown to perform very strongly, improving the EER from 18.5% to 7.01%, significantly outperforming both z-Norm and t-Norm in this case. The baseline system performance was deemed acceptable for the aims of this experiment, which were focused on evaluating and comparing the performance of the proposed r-Norm idea.
David Vandyke, Michael Wagner 0004, Roland Göcke
INTERSPEECH3
2012 Finding Happiest Moments in a Social Context
Abhinav Dhall, Jyoti Joshi, Ibrahim Radwan, Roland Göcke
ACCV (2)4
2012 Regression Based Pose Estimation with Automatic Occlusion Detection and Rectification
abstract
Human pose estimation is a classic problem in computer vision. Statistical models based on part-based modelling and the pictorial structure framework have been widely used recently for articulated human pose estimation. However, the performance of these models has been limited due to the presence of self-occlusion. This paper presents a learning-based framework to automatically detect and recover self-occluded body parts. We learn two different models: one for detecting occluded parts in the upper body and another one for the lower body. To solve the key problem of knowing which parts are occluded, we construct Gaussian Process Regression (GPR) models to learn the parameters of the occluded body parts from their corresponding ground truth parameters. Using these models, the pictorial structure of the occluded parts in unseen images is automatically rectified. The proposed framework outperforms a state-of-the-art pictorial structure approach for human pose estimation on 3 different datasets.
Ibrahim Radwan, Abhinav Dhall, Jyoti Joshi, Roland Göcke
ICME4
2012 An Improved NN Training Scheme Using Two-Stage LDA Features for Face Recognition
Behzad Bozorgtabar, Roland Göcke
ICONIP (5)2
2012 Group expression intensity estimation in videos via Gaussian Processes
Abhinav Dhall, Roland Göcke
ICPR2
2012 Neural-net classification for spatio-temporal descriptor based depression analysis
Jyoti Joshi, Abhinav Dhall, Roland Göcke, Michael Breakspear, Gordon Parker
ICPR3
2012 Correcting pose estimation with implicit occlusion detection and rectification
Ibrahim Radwan, Abhinav Dhall, Roland Göcke
ICPR3
2012 Facial Performance Transfer via Deformable Models and Parametric Correspondence
abstract
The issue of transferring facial performance from one person's face to another's has been an area of interest for the movie industry and the computer graphics community for quite some time. In recent years, deformable face models, such as the Active Appearance Model (AAM), have made it possible to track and synthesize faces in real time. Not surprisingly, deformable face model-based approaches for facial performance transfer have gained tremendous interest in the computer vision and graphics community. In this paper, we focus on the problem of real-time facial performance transfer using the AAM framework. We propose a novel approach of learning the mapping between the parameters of two completely independent AAMs, using them to facilitate the facial performance transfer in a more realistic manner than previous approaches. The main advantage of modeling this parametric correspondence is that it allows a "meaningful" transfer of both the nonrigid shape and texture across faces irrespective of the speakers' gender, shape, and size of the faces, and illumination conditions. We explore linear and nonlinear methods for modeling the parametric correspondence between the AAMs and show that the sparse linear regression method performs the best. Moreover, we show the utility of the proposed framework for a cross-language facial performance transfer that is an area of interest for the movie dubbing industry.
Akshay Asthana, Miles de la Hunty, Abhinav Dhall, Roland Göcke
IEEE Trans. Vis. Comput. Graph.4
2011 Pose Normalization via Learned 2D Warping for Fully Automatic Face Recognition
abstract
We present a novel approach to pose-invariant face recognition that handles continuous pose variations, is not database-specific, and achieves high accuracy without any manual intervention. Our method uses multidimensional Gaussian process regression to learn a nonlinear mapping function from the 2D shapes of faces at any non-frontal pose to the corresponding 2D frontal face shapes. We use this mapping to take an input image of a new face at an arbitrary pose and pose-normalize it, generating a synthetic frontal image of the face that is then used for recognition. Our fully automatic system for face recognition includes automatic methods for extracting 2D facial feature points and accurately estimating 3D head pose, and this information is used as input to the 2D pose-normalization algorithm. The current system can handle pose variation up to 45 degrees to the left or right (yaw angle) and up to 30 degrees up or down (pitch angle). The system demonstrates high accuracy in recognition experiments on the CMU-PIE, USF 3D, and Multi-PIE databases, showing excellent generalization across databases and convincingly outperforming other automatic methods.
Akshay Asthana, Michael J. Jones 0001, Tim K. Marks, Kinh H. Tieu, Roland Göcke
BMVC5
2011 A SSIM-based approach for finding similar facial expressions
abstract
There are various scenarios where finding the most similar expression is the requirement rather than classifying one into discrete, pre-defined classes, for example, for facial expression transfer and facial expression based automatic album generation. This paper proposes a novel method for finding the most similar facial expression. Instead of the regular L2 norm distance, we investigate the use of the Structural SIMilarity (SSIM) metric for similarity comparison as a distance metric in a nearest neighbour unsupervised algorithm. The feature vectors are generated using Active Appearance Models (AAM). We also demonstrate how this technique can be extended and used for finding corresponding facial expression images across two or more subjects, which is useful in applications such as facial animation and automatic expression transfer. Person-independent facial expression performance results are shown on the Multi-PIE, FEEDTUM and AVOZES databases. We also compare the performance of the SSIM metric versus other distance metrics in a nearest neighbour search for finding the most similar facial expression to a given image.
Abhinav Dhall, Akshay Asthana, Roland Göcke
FG3
2011 Emotion recognition using PHOG and LPQ features
abstract
We propose a method for automatic emotion recognition as part of the FERA 2011 competition. The system extracts pyramid of histogram of gradients (PHOG) and local phase quantisation (LPQ) features for encoding the shape and appearance information. For selecting the key frames, K-means clustering is applied to the normalised shape vectors derived from constraint local model (CLM) based face tracking on the image sequences. Shape vectors closest to the cluster centers are then used to extract the shape and appearance features. We demonstrate the results on the SSPNET GEMEP-FERA dataset. It comprises of both person specific and person independent partitions. For emotion classification we use support vector machine (SVM) and largest margin nearest neighbour (LMNN) and compare our results to the pre-computed FERA 2011 emotion challenge baseline.
Abhinav Dhall, Akshay Asthana, Roland Göcke, Tom Gedeon
FG3
2011 Building an Audio-Visual Corpus of Australian English: Large Corpus Collection with an Economical Portable and Replicable Black Box
abstract
The Big Australian Speech Corpus project incorporates the strategic goals of 30 Chief Investigators from various speech science areas. Speech from 1000 geographically and socially diverse speakers is being recorded using a uniform and automated protocol plus standardized hardware and software to produce a widely applicable and extensible database – AusTalk. Here we describe the project’s major components and organization; share the lessons learnt from difficulties and challenges; and present the results achieved so far.
Denis Burnham, Dominique Estival, Steven Fazio, Jette Viethen, Felicity Cox, Robert Dale, Steve Cassidy, Julien Epps, Roberto Togneri, Michael Wagner 0004, Yuko Kinoshita, Roland Göcke, Joanne Arciuli, Mark Onslow, Trent W. Lewis, Andrew Butcher, John Hajek
INTERSPEECH12
2011 An Investigation of Depressed Speech Detection: Features and Normalization
abstract
In recent years, the problem of automatic detection of mental illness from the speech signal has gained some initial interest, however questions remaining include how speech segments should be selected, what features provide good discrimination, and what benefits feature normalization might bring given the speaker-specific nature of mental disorders. In this paper, these questions are addressed empirically using classifier configurations employed in emotion recognition from speech, evaluated on a 47-speaker depressed/neutral read sentence speech database. Results demonstrate that (1) detailed spectral features are well suited to the task, (2) speaker normalization provides benefits mainly for less detailed features, and (3) dynamic information appears to provide little benefit. Classification accuracy using a combination of MFCC and formant based features approached 80 % for this database. Index Terms: mental state recognition, depressed speech, feature comparison, MFCC, Gaussian mixture models
Nicholas Cummins, Julien Epps, Michael Breakspear, Roland Göcke
INTERSPEECH4
2011 Regression based automatic face annotation for deformable model building
Akshay Asthana, Simon Lucey, Roland Göcke
Pattern Recognit.3
2010 Facial Expression Based Automatic Album Creation
Abhinav Dhall, Akshay Asthana, Roland Göcke
ICONIP (2)3
2010 Linear Facial Expression Transfer with Active Appearance Models
abstract
The issue of transferring facial expressions from one person's face to another's has been an area of interest for the movie industry and the computer graphics community for quite some time. In recent years, with the proliferation of online image and video collections and web applications, such as Google Street View, the question of preserving privacy through face de-identification has gained interest in the computer vision community. In this paper, we focus on the problem of real-time dynamic facial expression transfer using an Active Appearance Model framework. We provide a theoretical foundation for a generalisation of two well-known expression transfer methods and demonstrate the improved visual quality of the proposed linear extrapolation transfer method on examples of face swapping and expression transfer using the AVOZES data corpus. Realistic talking faces can be generated in real-time at low computational cost.
Miles de la Hunty, Akshay Asthana, Roland Göcke
ICPR3
2010 Illumination and Expression Invariant Recognition Using SSIM Based Sparse Representation
abstract
The sparse representation technique has provided a new way of looking at object recognition. As we demonstrate in this paper, however, the mean-squared error (MSE) measure, which is at the heart of this technique, is not a very robust measure when it comes to comparing facial images, which differ significantly in luminance values, as it only performs pixel-by-pixel comparisons. This requires a significantly large training set with enough variations in it to offset the drawback of the MSE measure. A large training set, however, is often not available. We propose the replacement of the MSE measure by the structural similarity (SSIM) measure in the sparse representation algorithm, which performs a more robust comparison using only one training sample per subject. In addition, since the off-the-shelf sparsifiers are also written using the MSE measure, we developed our own sparsifier using genetic algorithms that use the SSIM measure. We applied the modified algorithm to the Extended Yale Face B database as well as to the Multi-PIE database with expression and illumination variations. The improved performance demonstrates the effectiveness of the proposed modifications.
Asim A. Khwaja, Akshay Asthana, Roland Göcke
ICPR3
2009 Learning-based Face Synthesis for Pose-Robust Recognition from Single Image
abstract
Face recognition in real-world conditions requires the ability to deal with a number of conditions, such as variations in pose, illumination and expression. In this paper, we focus on variations in head pose and use a computationally efficient regression-based approach for synthesising face images in different poses, which are used to extend the face recognition training set. In this data-driven approach, the correspondences between facial landmark points in frontal and non-frontal views are learnt offline from manually annotated training data via Gaussian Process Regression. We then use this learner to synthesise non-frontal face images from any unseen frontal image. To demonstrate the utility of this approach, two frontal face recognition systems (the commonly used PCA and the recent Multi-Region Histograms) are augmented with synthesised non-frontal views for each person. This synthesis and augmentation approach is experimentally validated on the FERET dataset, showing a considerable improvement in recognition rates for ±40° and ±60° views, while maintaining high recognition rates for ±15° and ±25° views.
Akshay Asthana, Conrad Sanderson, Tom Gedeon, Roland Göcke
BMVC4
2009 Learning based automatic face annotation for arbitrary poses and expressions from frontal images only
abstract
Statistical approaches for building non-rigid deformable models, such as the active appearance model (AAM), have enjoyed great popularity in recent years, but typically require tedious manual annotation of training images. In this paper, a learning based approach for the automatic annotation of visually deformable objects from a single annotated frontal image is presented and demonstrated on the example of automatically annotating face images that can be used for building AAMs for fitting and tracking. This approach employs the idea of initially learning the correspondences between landmarks in a frontal image and a set of training images with a face in arbitrary poses. Using this learner, virtual images of unseen faces at any arbitrary pose for which the learner was trained can be reconstructed by predicting the new landmark locations and warping the texture from the frontal image. View-based AAMs are then built from the virtual images and used for automatically annotating unseen images, including images of different facial expressions, at any random pose within the maximum range spanned by the virtually reconstructed images. The approach is experimentally validated by automatically annotating face images from three different databases.
Akshay Asthana, Roland Göcke, Novi Quadrianto, Tom Gedeon
CVPR2
2009 Automatic frontal face annotation and AAM building for arbitrary expressions from a single frontal image only
abstract
Statistically motivated approaches for the registration and tracking of non-rigid objects, such as the active appearance model (AAM), have become very popular. A major drawback of these approaches is that they require manual annotation of all training images which can be tedious and error prone. In this paper, a MPEG-4 based approach for the automatic annotation of frontal face images, having any arbitrary facial expression, from a single annotated frontal image is presented. This approach utilises the MPEG-4 based facial animation system to generate virtual images having different expressions and uses the existing AAM framework to automatically annotate unseen images. The approach demonstrates an excellent generalisability by automatically annotating face images from two different databases.
Akshay Asthana, Asim A. Khwaja, Roland Göcke
ICIP3
2009 Learning AAM fitting through simulation
Jason M. Saragih, Roland Göcke
Pattern Recognit.2
2008 Optical flow estimation using Fourier Mellin Transform
abstract
In this paper, we propose a novel method of computing the optical flow using the Fourier Mellin Transform (FMT). Each image in a sequence is divided into a regular grid of patches and the optical flow is estimated by calculating the phase correlation of each pair of co-sited patches using the FMT. By applying the FMT in calculating the phase correlation, we are able to estimate not only the pure translation, as limited in the case of the basic phase correlation techniques, but also the scale and rotation motion of image patches, i.e. full similarity transforms. Moreover, the motion parameters of each patch can be estimated to sub-pixel accuracy based on a recently proposed algorithm that uses a 2D esinc function in fitting the data from the phase correlation output. We also improve the estimation of the optical flow by presenting a method of smoothing the field by using a vector weighted average filter. Finally, experimental results, using publicly available data sets are presented, demonstrating the accuracy and improvements of our method over previous optical flow methods.
Huy Tho Ho, Roland Göcke
CVPR2
2008 A Hybrid Fuzzy Approach for Human Eye Gaze Pattern Recognition
Dingyun Zhu, B. Sumudu U. Mendis, Tom Gedeon, Akshay Asthana, Roland Göcke
ICONIP (2)5
2008 A composite framework for affective sensing
abstract
A system capable of interpreting affect from a speaking face must recognise and fuse signals from multiple cues. Building such a system requires the integration of software components to perform tasks such as image registration, video segmentation, speech recognition and classification. Such software components tend to be idiosyncratic, purpose-built, and driven by scripts and textual configuration files. Integrating components to achieve the necessary degree of flexibility to perform full multimodal affective recognition is challenging. We discuss the key requirements and describe a system to perform multimodal affect sensing which integrates such software components and meets these requirements. Index Terms:emotion recognition, affective sensing 1.
Gordon McIntyre, Roland Göcke
INTERSPEECH2
2007 Monocular and Stereo Methods for AAM Learning from Video
abstract
The active appearance model (AAM) is a powerful method for modeling deformable visual objects. One of the major drawbacks of the AAM is that it requires a training set of pseudo-dense correspondences over the whole database. In this work, we investigate the utility of stereo constraints for automatic model building from video. First, we propose a new method for automatic correspondence finding in monocular images which is based on an adaptive template tracking paradigm. We then extend this method to take the scene geometry into account, proposing three approaches, each accounting for the availability of the fundamental matrix and calibration parameters or the lack thereof. The performance of the monocular method was first evaluated on a pre-annotated database of a talking face. We then compared the monocular method against its three stereo extensions using a stereo database.
Jason M. Saragih, Roland Göcke
CVPR2
2007 A Nonlinear Discriminative Approach to AAM Fitting
abstract
The Active Appearance Model (AAM) is a powerful generative method for modeling and registering deformable visual objects. Most methods for AAM fitting utilize a linear parameter update model in an iterative framework. Despite its popularity, the scope of this approach is severely restricted, both in fitting accuracy and capture range, due to the simplicity of the linear update models used. In this paper, we present an new AAM fitting formulation, which utilizes a nonlinear update model. To motivate our approach, we compare its performance against two popular fitting methods on two publicly available face databases, in which this formulation boasts significant performance improvements.
Jason M. Saragih, Roland Göcke
ICCV2
2007 Automatic Parametrisation for an Image Completion Method Based on Markov Random Fields
abstract
Recently, a new exemplar-based method for image completion, texture synthesis and image inpainting was proposed which uses a discrete global optimization strategy based on Markov random fields. Its main advantage lies in the use of priority belief propagation and dynamic label pruning to reduce the computational cost of standard belief propagation while producing high quality results. However, one of the drawbacks of the method is its use of a heuristically chosen parameter set. In this paper, a method for automatically determining the parameters for the belief propagation and dynamic label pruning steps is presented. The method is based on an information theoretic approach making use of the entropy of the image patches and the distribution of pairwise node potentials. A number of image completion results are shown demonstrating the effectiveness of our method.
Huy Tho Ho, Roland Göcke
ICIP (3)2
2004 The audio-video australian English speech data corpus AVOZES
abstract
This paper presents the Audio-Video Australian English Speech data corpus AVOZES. It contains recordings of 20 speakers ut-tering a variety of phrases. The corpus was designed for re-search on the statistical relationship of audio and video speech parameters with an audio-video (AV) automatic speech recog-nition (ASR) task in mind, but may be useful for other research tasks. AVOZES is the first published AV speaking-face data corpus for Australian English and is novel in its use of a stereo camera system for the video recordings and its modular design. 1.
J. Bruce Millar, Roland Göcke
INTERSPEECH2
2004 Aspects of speaking-face data corpus design methodology
abstract
S.1157-1160
J. Bruce Millar, Michael Wagner 0004, Roland Göcke
INTERSPEECH3
2002 Noisy audio feature enhancement using audio-visual speech data
abstract
We investigate improving automatic speech recognition (ASR) in noisy conditions by enhancing noisy audio features using visual speech captured from the speaker's face. The enhancement is achieved by applying a linear filter to the concatenated vector of noisy audio and visual features, obtained by mean square error estimation of the clean audio features in a training stage. The performance of the enhanced audio features is evaluated on two ASR tasks: A connected digits task and speaker-independent, large-vocabulary, continuous speech recognition. In both cases and at sufficiently low signal-to-noise ratios (SNRs), ASR trained on the enhanced audio features significantly outperforms ASR trained on the noisy audio, achieving for example a 46% relative reduction in word error rate on the digits task at −3.5 dB SNR. However, the method fails to capture the full visual modality benefit to ASR, as demonstrated by its comparison to discriminant audio-visual feature fusion introduced in previous work.
Roland Göcke, Gerasimos Potamianos, Chalapathy Neti
ICASSP1