Mohammad H. Mahoor

dblp:36/3051 · DBLP profile ↗
← Back
69ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0001-8923-4660ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 45 · 3 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 5 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 since 2021Systems, architecture and hardware · 6 · 1 since 2021Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2026 Contrastive learning-based video quality assessment-jointed video vision transformer for video recognition
Mohammad H. Mahoor
Neural Comput. Appl.2
2026 AffectNet+: A Database for Enhancing Facial Expression Recognition With Soft-Labels
abstract
Automated Facial Expression Recognition (FER) is challenging due to intra-class variations and inter-class similarities. FER can be especially difficult when facial expressions reflect a mixture of various emotions (aka compound expressions). Existing FER datasets, such as AffectNet, provide discrete emotion labels (hard-labels), where a single category of emotion is assigned to an expression. To alleviate inter- and intra-class challenges, as well as provide a better facial expression descriptor, we propose a new approach to create FER datasets through a labeling method in which an image is labeled with more than one emotion (calledsoft-labels), each with a different confidence. Specifically, we introduce the notion ofsoft-labelsfor facial expression datasets, a new approach to affective computing for more realistic recognition of facial expressions. To achieve this goal, we propose a novel methodology to accurately calculatesoft-labels:a vector representing the extent to which multiple categories of emotion are simultaneously present within a single facial expression. Finding smoother decision boundaries, enabling multi-labeling, and mitigating bias and imbalanced data are some of the advantages of our proposed method. Building upon AffectNet, we introduce AffectNet+, the next-generation facial expression dataset. This dataset containssoft-labels, three categories of data complexity subsets, and additional metadata such as age, gender, race, head pose, facial landmarks, valence, and arousal. AffectNet+ will be made publicly accessible to researchers.
Ali Pourramezan Fard, Timothy D. Sweeny, Mohammad H. Mahoor
IEEE Trans. Affect. Comput.4
2025 A Reinforcement Learning-Based Social Robot for Personalized Learning in Children with Autism
abstract
This work hypothesizes that a social robot that uses reinforcement learning can effectively adapt to individual differences in teaching imitation skills (e.g., facial expressions) to children with autism spectrum disorder. We developed an active learning method based on reinforcement learning to personalize human-robot interaction sessions based on each child's imitation performance and preference. We evaluated this method with five children with autism spectrum disorder, and the results demonstrated varying responses to different methods of presenting facial expressions to teach imitation skills. We found that the robot consistently promoted increased shared attention, including visual contact and physical proximity during imitation tasks. This suggests that adaptive human-robot interactions can cater to the unique needs of children with autism, offering a promising avenue for personalized intervention. Additionally, we discuss observed qualitative insights from our study and considerations for robot behavior mitigation strategies to sustain engagement.
Farzaneh Askari, Hojjat Abdollahi, Kerstin S. Haring, Mohammad H. Mahoor
ICRA4
2025 Social Robot-led Yoga for Older Adults: A Feasibility Study
abstract
Yoga and other forms of exercise have demonstrated protective health benefits for older adults, including enhanced body flexibility, balance, joint mobility, and cognitive function. Social robots are increasingly used as caregivers for older adults, including instruction to improve mental and physical well-being. The primary objective of this study was to measure the impact of social robot-led yoga with six female older adult participants across twenty-four sessions of social robot-led yoga. This study used a pre-peri-post design to compare participants’ biopsychosocial measurements before, during, and after yoga participation. We analyzed impact across measures of fall resilience, overhead flexion, mindfulness skills, and task engagement. Participants demonstrated statistically significant improvement in fall resilience, statistically non-significant improvements in overhead flexion and mindfulness skills, and increased task engagement throughout the study. As one of the first studies to investigate the impact of social robotled yoga with older adults, the findings from this study provide statistically significant evidence that older adults may see improved nondominant single-leg balance from social robot-led yoga. When considered in conjunction with other research on the benefits of yoga with older adults, these findings suggest that robot-led social yoga instruction warrants additional investigation.
Sean Mapoles, Stephanie Melgar-Donis, Jason Ghiglieri, Jarid Siewierski, Hojjat Abdollahi, Kim A. Gorgens, Mohammad H. Mahoor
RO-MAN7
2024 Mild cognitive impairment detection from facial video interviews by applying spatial-to-temporal attention module
Muath Alsuhaibani, Hiroko H. Dodge, Mohammad H. Mahoor
Expert Syst. Appl.3
2024 MC-ViViT: Multi-branch Classifier-ViViT to detect Mild Cognitive Impairment in older adults using facial videos
Hiroko H. Dodge, Mohammad H. Mahoor
Expert Syst. Appl.3
2023 Topical language generation using transformers
abstract
Abstract Large-scale transformer-based language models (LMs) demonstrate impressive capabilities in open-text generation. However, controlling the generated text’s properties such as the topic, style, and sentiment is challenging and often requires significant changes to the model architecture or retraining and fine-tuning the model on new supervised data. This paper presents a novel approach for topical language generation (TLG) by combining a pre-trained LM with topic modeling information. We cast the problem using Bayesian probability formulation with topic probabilities as a prior, LM probabilities as the likelihood, and TLG probability as the posterior. In learning the model, we derive the topic probability distribution from the user-provided document’s natural structure. Furthermore, we extend our model by introducing new parameters and functions to influence the quantity of the topical features presented in the generated text. This feature would allow us to easily control the topical properties of the generated text. Our experimental results demonstrate that our model outperforms the state-of-the-art results on coherency, diversity, and fluency while being faster in decoding.
Rohola Zandie, Mohammad H. Mahoor
Nat. Lang. Eng.2
2023 Artificial Emotional Intelligence in Socially Assistive Robots for Older Adults: A Pilot Study
abstract
This paper presents our recent research on integrating artificial emotional intelligence in a social robot (Ryan) and studies the robot's effectiveness in engaging older adults. Ryan is a socially assistive robot designed to provide companionship for older adults with depression and dementia through conversation. We used two versions of Ryan for our study, empathic and non-empathic. The empathic Ryan utilizes a multimodal emotion recognition algorithm and a multimodal emotion expression system. Using different input modalities for emotion, i.e. facial expression and speech sentiment, the empathic Ryan detects users emotional state and utilizes an affective dialogue manager to generate a response. On the other hand, the non-empathic Ryan lacks facial expression and uses scripted dialogues that do not factor in the users emotional state. We studied these two versions of Ryan with 10 older adults living in a senior care facility. The statistically significant improvement in the users' reported face-scale mood measurement indicates an overall positive effect from the interaction with both the empathic and non-empathic versions of Ryan. However, the number of spoken words measurement and the exit survey analysis suggest that the users perceive the empathic Ryan as more engaging and likable.
Hojjat Abdollahi, Mohammad H. Mahoor, Rohola Zandie, Jarid Siewierski, Sara H. Qualls
IEEE Trans. Affect. Comput.2
2023 Deep Siamese Neural Networks for Facial Expression Recognition in the Wild
abstract
This article introduces an algorithm for facial expression recognition (FER) using deep Siamese Neural Networks (SNNs) that preserve the local structure of images in the embedding similarity space. We designed the network to reveal the input pairs similarity by comparing features through a designed metric. Furthermore, we developed a novel image pairing (i.e., positive and negative pairs) strategy technique to train our Siamese model. Our Siamese model comprises of a verification framework and an identification framework to learn a joint embedding space. The verification path reduces the intra-class variations by minimizing the distance between the extracted features from the same class, while the identification path increases the inter-class variations by maximizing the distance between the features extracted from different classes. We apply transfer learning to only use the identification model for facial expression classification. We evaluated our algorithm using AffectNet, FER2013, and Compound Facial Expressions of Emotion (CFEE) datasets, where better results are achieved compared to other deep learning-based approaches.
Wassan Hayale, Pooran Singh Negi, Mohammad H. Mahoor
IEEE Trans. Affect. Comput.3
2022 ACR Loss: Adaptive Coordinate-based Regression Loss for Face Alignment
abstract
Although deep neural networks have achieved reasonable accuracy in solving face alignment, it is still a challenging task, specifically when dealing with facial images, under occlusion, or extreme head poses. Heatmap-based Regression (HBR) and Coordinate-based Regression (CBR) are among the two mainly used methods for face alignment. CBR methods require less computer memory, though their performance is less than HBR methods. In this paper, we propose an Adaptive Coordinate-based Regression (ACR) loss to improve the accuracy of CBR for face alignment. Inspired by the Active Shape Model (ASM), we generate Smooth-Face objects, a set of facial landmark points with fewer variations compared to the ground truth landmark points. We then introduce a method to estimate the level of difficulty in predicting each landmark point for the network by comparing the distribution of the ground truth landmark points and the corresponding Smooth-Face objects. Our proposed ACR Loss can adaptively modify its curvature and the influence of the loss based on the difficulty level of predicting each landmark point in a face. Accordingly, the ACR Loss guides the network toward more challenging points than easier points, which improves the accuracy of the face alignment task. Our extensive evaluation shows the capabilities of the proposed ACR Loss in predicting facial landmark points in various facial images.
Ali Pourramezan Fard, Mohammad H. Mahoor
ICPR2
2022 Facial landmark points detection using knowledge distillation-based neural networks
Ali Pourramezan Fard, Mohammad H. Mahoor
Comput. Vis. Image Underst.2
2022 BReG-NeXt: Facial Affect Computing Using Adaptive Residual Networks With Bounded Gradient
abstract
This article introducesBReG-NeXt, a residual-based network architecture using a function wtih bounded derivative instead of a simple shortcut path (a.k.a.identity mapping) in the residual units for automatic recognition of facial expressions based on the categorical and dimensional models of affect. Compared to ResNet, our proposed adaptive complex mapping results in a shallower network with less numbers of training parameters and floating point operations per second (FLOPs). Adding trainable parameters to the bypass function further improves fitting and training the network and hence recognizing subtle facial expressions such as contempt with a higher accuracy. We conducted comprehensive experiments on the categorical and dimensional models of affect on the challenging in-the-wild databases of AffectNet, FER2013, and Affect-in-Wild. Our experimental results show that our adaptive complex mapping approach outperforms the original ResNet consisting of a simple identity mapping as well as other state-of-the-art methods for Facial Expression Recognition (FER). Various metrics are reported in both affect models to provide a comprehensive evaluation of our method. In the categorical model, BReG-NeXt-50 with only 3.1M training parameters and 15 MFLOPs, achieves 68.50 and 71.53 percent accuracy on AffectNet and FER2013 databases, respectively. In the dimensional model, BReG-NeXt achieves 0.2577 and 0.2882 RMSE value on AffectNet and Affect-in-Wild databases, respectively.
Behzad Hassani, Pooran Singh Negi, Mohammad H. Mahoor
IEEE Trans. Affect. Comput.3
2021 RyanSpeech: A Corpus for Conversational Text-to-Speech Synthesis
abstract
This paper introduces RyanSpeech, a new speech corpus for research on automated text-to-speech (TTS) systems. Publicly available TTS corpora are often noisy, recorded with multiple speakers, or lack quality male speech data. In order to meet the need for a high quality, publicly available male speech corpus within the field of speech recognition, we have designed and created RyanSpeech which contains textual materials from real-world conversational settings. These materials contain over 10 hours of a professional male voice actor's speech recorded at 44.1 kHz. This corpus's design and pipeline make RyanSpeech ideal for developing TTS systems in real-world applications. To provide a baseline for future research, protocols, and benchmarks, we trained 4 state-of-the-art speech models and a vocoder on RyanSpeech. The results show 3.36 in mean opinion scores (MOS) in our best model. We have made both the corpus and trained models for public use.
Rohola Zandie, Mohammad H. Mahoor, Julia Madsen, Eshrat S. Emamian
Interspeech2
2020 Deep Learning-based Symbolic Indoor Positioning using the Serving eNodeB
abstract
This paper presents a novel indoor positioning method designed for residential apartments. The proposed method makes use of cellular signals emitting from a serving eNodeB which eliminates the need for specialized positioning infrastructure. Additionally, it utilizes Denoising Autoencoders to mitigate the effects of cellular signal loss. We evaluated the pro-posed method using real-world data collected from two different smartphones inside a representative apartment of eight symbolic spaces. Experimental results verify that the proposed method outperforms conventional symbolic indoor positioning techniques in various performance metrics. To promote reproducibility and foster new research efforts, we made all the data and codes associated with this work publicly available.
Fahad Alhomayani, Mohammad H. Mahoor
ICMLA2
2019 Bounded Residual Gradient Networks (BReG-Net) for Facial Affect Computing
abstract
Residual-based neural networks have shown remarkable results in various visual recognition tasks including Facial Expression Recognition (FER). Despite the tremendous efforts have been made to improve the performance of FER systems using DNNs, existing methods are not generalizable enough for practical applications. This paper introduces Bounded Residual Gradient Networks (BReG-Net) for facial expression recognition, in which the shortcut connection between the input and the output of the ResNet module is replaced with a differentiable function with a bounded gradient. This configuration prevents the network from facing the vanishing or exploding gradient problem. We show that utilizing such non-linear units will result in shallower networks with better performance. Further, by using a weighted loss function which gives a higher priority to less represented categories, we can achieve an overall better recognition rate. The results of our experiments show that BReG-Nets outperform state-of-the-art methods on three publicly available facial databases in the wild, on both the categorical and dimensional models of affect.
Behzad Hassani, Pooran Singh Negi, Mohammad H. Mahoor
FG3
2019 Facial Expression Recognition Using Deep Siamese Neural Networks with a Supervised Loss function
abstract
This paper presents a novel algorithm for an end-to-end facial expression recognition(FER) system based on deep Siamese neural networks equipped with a supervised loss function. Our method learns a powerful FER system by dynamically modulating verification signal over identification/classification signal. The identification signal increases the inter-class variations by maximizing the distance between the features for different classes, while the verification signal reduces the intra-class variations by minimizing the distance between features for the same class. We have evaluated our method on the AffectNet dataset [10] and achieved promising results compared to other deep learning models.
Wassan Hayale, Pooran Singh Negi, Mohammad H. Mahoor
FG3
2019 Delivering Cognitive Behavioral Therapy Using A Conversational Social Robot
abstract
Social robots are becoming an integrated part of our daily life due to their ability to provide companionship and entertainment. A subfield of robotics, Socially Assistive Robotics (SAR), is particularly suitable for expanding these benefits into the healthcare setting because of its unique ability to provide cognitive, social, and emotional support. This paper presents our recent research on developing SAR by evaluating the ability of a life-like conversational social robot, called Ryan, to administer internet-delivered cognitive behavioral therapy (iCBT) to older adults with depression. For Ryan to administer the therapy, we developed a dialogue-management system, called Program-R. Using an accredited CBT manual for the treatment of depression, we created seven hour-long iCBT dialogues and integrated them into Program-R using Artificial Intelligence Markup Language (AIML). To assess the effectiveness of Robot-based iCBT and users' likability of our approach, we conducted an HRI study with a cohort of elderly people with mild-to-moderate depression over a period of four weeks. Quantitative analyses of participant's spoken responses (e.g. word count and sentiment analysis), face-scale mood scores, and exit surveys, strongly support the notion robot-based iCBT is a viable alternative to traditional human-delivered therapy.
Francesca Dino, Rohola Zandie, Hojjat Abdollahi, Sarah Schoeder, Mohammad H. Mahoor
IROS5
2019 An Adaptive Bayesian Source Separation Method for Intensity Estimation of Facial AUs
abstract
Automated measurement of the intensity of spontaneous facial Action Units (AU) defined by the Facial Action Coding System (FACS) in video sequences is a challenging problem. This paper proposes a person-adaptive methodology for the intensity estimation of spontaneous AUs. We formulate this problem as a source separation problem where we consider the observed AUs as the source signals to be separated from each other and other information given by a sequence of facial images. We first compute an initial estimation of the sources, called observations, using sparse linear regression functions. We then develop and apply a Bayesian source separation method that recruits the prior information of the sources to iteratively improve the initial estimations/observations in an adaptive fashion. Furthermore, our approach adaptively uses some testing information (but not the ground-truth labels) to improve the performance of the approach (i.e., Person-Adaptive model). Our experimental results on DISFA, UNBC-McMaster and FERA2015 databases show that this approach is very promising for automated measurement of the intensity of spontaneous facial AUs.
Mohammad Reza Mohammadi, Emad Fatemizadeh, Mohammad H. Mahoor
IEEE Trans. Affect. Comput.3
2019 AffectNet: A Database for Facial Expression, Valence, and Arousal Computing in the Wild
abstract
Automated affective computing in the wild setting is a challenging problem in computer vision. Existing annotated databases of facial expressions in the wild are small and mostly cover discrete emotions (aka the categorical model). There are very limited annotated facial databases for affective computing in the continuous dimensional model (e.g., valence and arousal). To meet this need, we collected, annotated, and prepared for public distribution a new database of facial emotions in the wild (called AffectNet). AffectNet contains more than 1,000,000 facial images from the Internet by querying three major search engines using 1,250 emotion related keywords in six different languages. About half of the retrieved images were manually annotated for the presence of seven discrete facial expressions and the intensity of valence and arousal. AffectNet is by far the largest database of facial expression, valence, and arousal in the wild enabling research in automated facial expression recognition in two different emotion models. Two baseline deep neural networks are used to classify images in the categorical model and predict the intensity of valence and arousal. Various evaluation metrics show that our deep neural network baselines can perform better than conventional machine learning methods and off-the-shelf facial expression recognition systems.
Ali Mollahosseini, Behzad Hassani, Mohammad H. Mahoor
IEEE Trans. Affect. Comput.3
2018 Improved indoor geomagnetic field fingerprinting for smartwatch localization using deep learning
abstract
Geomagnetic field fingerprinting has attracted researchers in recent years as a promising alternative to WiFi and Bluetooth fingerprinting. It is omnipresent, stable, and does not require the deployment of specialized infrastructure to be realized. While several studies have utilized these characteristics in developing indoor positioning systems, the positioning accuracy still can be improved. This paper presents a Convolutional Neural Network (CNN) based method for designing and developing a novel smartwatch-based indoor geomagnetic field positioning system. We tested the proposed system on real world data in an indoor environment composed of three corridors of different lengths and three rooms of different sizes. Experimental results show a promising location classification accuracy of 97.77% with a mean localization error of 0.136 meters. We also demonstrate how the Softmax (SM) layer of the network can be exploited to further improve the localization accuracy in user tracking scenarios.
Fahad Alhomayani, Mohammad H. Mahoor
IPIN2
2018 A Pilot Study on Facial Expression Recognition Ability of Autistic Children Using Ryan, A Rear-Projected Humanoid Robot
abstract
Rear-projected robots use computer graphics technology to create facial animations and project them on a mask to show the robot's facial cues and expressions. These types of robots are becoming commercially available, though more research is required to understand how they can be effectively used as a socially assistive robotic agent. This paper presents the results of a pilot study on comparing the facial expression recognition abilities of children with Autism Spectrum Disorder (ASD) with typically developing (TD) children using a rear-projected humanoid robot called Ryan. Six children with ASD and six TD children participated in this research, where Ryan showed them six basic expressions (i.e. anger, disgust, fear, happiness, sadness, and surprise) with different intensity levels. Participants were asked to identify the expressions portrayed by Ryan. The results of our study show that there is not any general impairment in expression recognition ability of the ASD group comparing to the TD control group; however, both groups showed deficiencies in identifying disgust and fear. Increasing the intensity of Ryan's facial expressions significantly improved the expression recognition accuracy. Both groups were successful to recognize the expressions demonstrated by Ryan with high average accuracy.
Farzaneh Askari, Huanghao Feng, Timothy D. Sweeny, Mohammad H. Mahoor
RO-MAN4
2018 Studying Effects of Incorporating Automated Affect Perception with Spoken Dialog in Social Robots
abstract
Social robots are becoming an integrated part of our daily lives with the goal of understanding humans' social intentions and feelings, a capability which is often referred to as empathy. Despite significant progress towards the development of empathic social agents, current social robots have yet to reach the full emotional and social capabilities. This paper presents our recent effort on incorporating an automated Facial Expression Recognition (FER) system based on deep neural networks into the spoken dialog of a social robot (Ryan) to extend and enrich its capabilities beyond spoken dialog and integrate the user's affect state into the robot's responses. In order to evaluate whether this incorporation can improve social capabilities of Ryan, we conducted a series of Human-Robot-Interaction (HRI) experiments. In these experiments the subjects watched some videos and Ryan engaged them in a conversation driven by user's facial expressions perceived by the robot. We measured the accuracy of the automated FER system on the robot when interacting with different human subjects as well as three social/interactive aspects, namely task engagement, empathy, and likability of the robot. The results of our HRI study indicate that the subjects rated empathy and likability of the affect-aware Ryan significantly higher than non-empathic (the control condition) Ryan. Interestingly, we found that the accuracy of the FER system is not a limiting factor, as subjects rated the affect-aware agent equipped with a low accuracy FER system as empathic and likable as when facial expression was recognized by a human observer.
Ali Mollahosseini, Hojjat Abdollahi, Mohammad H. Mahoor
RO-MAN3
2018 A wavelet-based approach to emotion classification using EDA signals
Huanghao Feng, Hosein M. Golshan, Mohammad H. Mahoor
Expert Syst. Appl.3
2018 Role of embodiment and presence in human perception of robots' facial cues
Ali Mollahosseini, Hojjat Abdollahi, Timothy D. Sweeny, Ronald A. Cole, Mohammad H. Mahoor
Int. J. Hum. Comput. Stud.5
2017 Spatio-Temporal Facial Expression Recognition Using Convolutional Neural Networks and Conditional Random Fields
abstract
Automated Facial Expression Recognition (FER) has been a challenging task for decades. Many of the existing works use hand-crafted features such as LBP, HOG, LPQ, and Histogram of Optical Flow (HOF) combined with classifiers such as Support Vector Machines for expression recognition. These methods often require rigorous hyperparameter tuning to achieve good results. Recently Deep Neural Networks (DNN) have shown to outperform traditional methods in visual object recognition. In this paper, we propose a two-part network consisting of a DNN-based architecture followed by a Conditional Random Field (CRF) module for facial expression recognition in videos. The first part captures the spatial relation within facial images using convolutional layers followed by three Inception- ResNet modules and two fully-connected layers. To capture the temporal relation between the image frames, we use linear chain CRF in the second part of our network. We evaluate our proposed network on three publicly available databases, viz. CK+, MMI, and FERA. Experiments are performed in subjectindependent and cross-database manners. Our experimental results show that cascading the deep network architecture with the CRF module considerably increases the recognition of facial expressions in videos and in particular it outperforms the stateof- the-art methods in the cross-database experiments and yields comparable results in the subject-independent experiments.
Behzad Hassani, Mohammad H. Mahoor
FG2
2017 An FFT-based synchronization approach to recognize human behaviors using STN-LFP signal
abstract
Classification of human behavior is a key step to developing closed-loop Deep Brain Stimulation (DBS) systems, which may decrease the power consumption and side effects of the existing systems. Recent studies have shown that the Local Field Potential (LFP) signals from both Subthalamic Nuclei (STN) of the brain can be used to recognize human behavior. Since the DBS leads implanted in each STN can collect three bipolar signals, the selection of a suitable pair of LFPs that achieves optimal recognition performance is still an open problem to address. Considering the presence of synchronized aggregate activity in the basal ganglia, this paper presents an FFT-based synchronization approach to automatically select a relevant pair of LFPs and use the pair together with an SVM-based MKL classifier for behavior recognition purposes. Our experiments on five subjects show the superiority of the proposed approach compared to other methods used for behavior classification.
Hosein M. Golshan, Adam O. Hebb, Sara J. Hanrahan, Joshua Nedrud, Mohammad H. Mahoor
ICASSP5
2017 A joint dictionary learning and regression model for intensity estimation of facial AUs
Mohammad Reza Mohammadi, Emad Fatemizadeh, Mohammad H. Mahoor
J. Vis. Commun. Image Represent.3
2016 Robot-based therapeutic protocol for training children with Autism
abstract
Robots are commonly used artificial agents with powerful capabilities in navigation, perception and execution in the physical world. One interesting question is how well robots can assist and engage individuals with social and behavioral deficits (such as autism) to acquire new skills? Preliminary studies in autism research demonstrate that in many cases individuals with Autism Spectrum Disorder (ASD) interact more actively and engagingly with robots than humans. As there are limited investigations for utilizing robots in social and behavioral treatments of individuals with ASD, we designed and evaluated a robot-based intervention protocol using a social robot (NAO) to deliver behavioral training mechanism for children with ASD. Results of our pilot study on seven verbal children with high functioning autism show behavioral response improvement, including pointing and facial expression recognition in the majority of the participants as a consequence of the behavioral intervention delivered directly through the robot. Results also show that individuals were able to engage in these learned skills during human-human follow-up sessions.
Seyed Mohammad Mavadati, Huanghao Feng, Michelle J. Salvador, Sophia Silver, Anibal Gutierrez, Mohammad H. Mahoor
RO-MAN6
2016 Going deeper in facial expression recognition using deep neural networks
abstract
Automated Facial Expression Recognition (FER) has remained a challenging and interesting problem in computer vision. Despite efforts made in developing various methods for FER, existing approaches lack generalizability when applied to unseen images or those that are captured in wild setting (i.e. the results are not significant). Most of the existing approaches are based on engineered features (e.g. HOG, LBPH, and Gabor) where the classifier's hyper-parameters are tuned to give best recognition accuracies across a single database, or a small collection of similar databases. This paper proposes a deep neural network architecture to address the FER problem across multiple well-known standard face datasets. Specifically, our network consists of two convolutional layers each followed by max pooling and then four Inception layers. The network is a single component architecture that takes registered facial images as the input and classifies them into either of the six basic or the neutral expressions. We conducted comprehensive experiments on seven publicly available facial expression databases, viz. MultiPIE, MMI, CK+, DISFA, FERA, SFEW, and FER2013. The results of our proposed architecture are comparable to or better than the state-of-the-art methods and better than traditional convolutional neural networks in both accuracy and training time.
Ali Mollahosseini, Mohammad H. Mahoor
WACV3
2016 Task-dependent multi-task multiple kernel learning for facial action unit detection
Xiao Zhang 0003, Mohammad H. Mahoor
Pattern Recognit.2
2016 Intensity Estimation of Spontaneous Facial Action Units Based on Their Sparsity Properties
abstract
Automatic measurement of spontaneous facial action units (AUs) defined by the facial action coding system (FACS) is a challenging problem. The recent FACS user manual defines 33 AUs to describe different facial activities and expressions. In spontaneous facial expressions, a subset of AUs are often occurred or activated at a time. Given this fact that AUs occurred sparsely over time, we propose a novel method to detect the absence and presence of AUs and estimate their intensity levels via sparse representation (SR). We use the robust principal component analysis to decompose expression from facial identity and then estimate the intensity of multiple AUs jointly using a regression model formulated based on dictionary learning and SR. Our experiments on Denver intensity of spontaneous facial action and UNBC-McMaster shoulder pain expression archive databases show that our method is a promising approach for measurement of spontaneous facial AUs.
Mohammad Reza Mohammadi, Emad Fatemizadeh, Mohammad H. Mahoor
IEEE Trans. Cybern.3
2015 An emotion recognition comparative study of autistic and typically-developing children using the zeno robot
abstract
In this paper we present the results of our recent study on comparing the emotion expression recognition abilities of children diagnosed with high functioning Autism (ASD) with those of typically developing (TD) children through use of a humanoid robot, Zeno. In our study we investigated the effect of incorporating gestures to the emotion expression prediction accuracy of both child groups. Although the idea that ASD individuals suffer from general emotion recognition deficits is widely assumed [1], we found no significant impairment in the general emotion prediction. However, a specific deficit in correctly identifying Fear was found for children with Autism when compared to the TD children. Furthermore, we found that gestures can significantly impact the prediction accuracy of both ASD and TD children in a negative or positive manner depending on the specific expression. Thus, the use of gestures for conveying emotional expressions by a humanoid robot in a social skill therapy setting is relevant. The methodology and experimental protocol are presented and additional discussion of the Zeno R-50 robot used is given.
Michelle J. Salvador, Sophia Silver, Mohammad H. Mahoor
ICRA3
2015 Facial expression recognition using lp-norm MKL multiclass-SVM
Xiao Zhang 0003, Mohammad H. Mahoor, Seyed Mohammad Mavadati
Mach. Vis. Appl.2
2015 Measuring the intensity of spontaneous facial action units with dynamic Bayesian network
Seyed Mohammad Mavadati, Mohammad H. Mahoor, Yongping Zhao
Pattern Recognit.3
2014 Jointly detecting infants' multiple facial action units expressed during spontaneous face-to-face communication
abstract
Automatic detection of spontaneous facial Action Units (AUs) in video has many applications including understanding infants' emotion-mediated interactions and development. The target AUs for detection are those essential to positive and negative emotion (i.e., AU 6, AU 12, and AU 20). Tracking and extraction of facial features is especially challenging in infants. Face shape and texture markedly differ from that in adults, jaw contour often is occult, sudden changes in pose and expression are common, and AU often occur in complex combinations. We investigate the association among AUs central to positive and negative emotion and propose a methodology for jointly detecting positively correlated facial AUs of infants during spontaneous interactions with their parents. We apply a subject-independent structural output model to (1) recognize combinations of AUs simultaneously, and (2) model the dependencies between AUs. Using this approach, we improved the reliability of automatic detection of AU 12 and AU 20 in a total 90-minute video of infant-parent interaction of 12 infants.
Nazanin Zaker, Mohammad H. Mahoor, Daniel S. Messinger, Jeffrey F. Cohn
ICIP2
2014 Temporal Facial Expression Modeling for Automated Action Unit Intensity Measurement
abstract
Spontaneous facial expression recognition using temporal patterns is a relatively unexplored area in facial image analysis. Several factors such as head orientation, co-occurrence and presence of subtle facial action units (AUs), and time variability of AUs make the problem more challenging. This paper presents a methodology to model and automatically recognize the intensity of spontaneous AUs in videos. Our method exploits localized Gabor features and Hidden Markov Model (HMM) to represent and model the dependencies of AU dynamics in both subject-dependent (SD) and subject-independent (SI) settings. Our experimental results show that temporal information can improve the recognition of AUs and their intensity levels compared to static methods.
Seyed Mohammad Mavadati, Mohammad H. Mahoor
ICPR2
2014 Simultaneous Detection of Multiple Facial Action Units via Hierarchical Task Structure Learning
abstract
Automatic facial action unit (AU) detection is a challenging research topic in computer vision and pattern recognition. Most of the existing approaches design classifiers to detect AUs individually without considering their intrinsic relations. This paper proposes a novel framework to jointly learn the classifiers for detecting the presence and absence of multiple AUs. In our method, hierarchical structure is defined to model the relations among a set of AU detection tasks, where each leaf denotes a specific AU. The relatedness among AUs is captured by introducing a latent layer whose nodes represent the common properties across several subsets of AUs. Multi-task multiple kernel learning (MTMKL) approach is utilized to simultaneously learn the similarities between AUs within our hierarchical model and the SVM discriminative hyper plane for detecting each AU. Extensive experiments on the CK+ and DISFA databases show that by exploiting the AU inter-relations, our proposed method has achieved encouraging performance on AU detection compared to several state-of-the-art methods.
Xiao Zhang 0003, Mohammad H. Mahoor
ICPR2
2014 ReFrESH: A self-adaptation framework to support fault tolerance in field mobile robots
abstract
Mobile robots are being employed far more often in extreme environments, such as urban search and rescue, with greater levels of autonomy; yet recent studies on field robotics show that numerous failure modes affect the reliability of the robot in meeting mission objectives. Therefore, fault tolerance is increasingly important for field robots operating in unpredictable environments to ensure safety and effectiveness of the system. This paper demonstrates a self-adaptation framework, ReFrESH, that contains mechanisms for fault detection and fault mitigation. The goal of ReFrESH is to provide diagnosable and maintainable infrastructure support, built into a real-time operating system, to manage task performance in the presence of unexpected uncertainties. ReFrESH augments the port-based object framework by attaching evaluation and estimation mechanisms to each functional component so that the robot can easily detect and locate faults. In conjunction, a task level decision mechanism interacts with the fault detection elements in order to generate and choose an optimal approach to mitigating faults. Moreover, to increase flexibility of the fault tolerance, ReFrESH provides self-adaptation support for both software and hardware functionality. To our knowledge, this is the first framework to support both software and hardware self-adaptation. A demonstrative application of ReFrESH illustrates its applicability through a target tracking task deployed on a mobile robot system.
Yanzhe Cui, Richard M. Voyles, Joshua T. Lane, Mohammad H. Mahoor
IROS4
2014 eBear: An expressive Bear-Like robot
abstract
This paper presents an anthropomorphic robotic bear for the exploration of human-robot interaction including verbal and non-verbal communications. This robot is implemented with a hybrid face composed of a mechanical faceplate with 10 DOFs and an LCD-display-equipped mouth. The facial emotions of the bear are designed based on the description of the Facial Action Coding System as well as some animal-like gestures described by Darwin. The mouth movements are realized by synthesizing emotions with speech. User acceptance investigations have been conducted to evaluate the likability of these facial behaviors exhibited by the eBear. Multiple Kernel Learning is proposed to fuse different features for recognizing user's facial expressions. Our experimental results show that the developed Bear-Like robot can perceive basic facial expressions and provide emotive conveyance towards human beings.
Xiao Zhang 0003, Ali Mollahosseini, Amir H. Kargar B., Evan Boucher, Richard M. Voyles, Rodney D. Nielsen, Mohammad H. Mahoor
RO-MAN7
2014 Simultaneous recognition of facial expression and identity via sparse representation
abstract
Automatic recognition of facial expression and facial identity from visual data are two challenging problems that are tied together. In the past decade, researchers have mostly tried to solve these two problems separately to come up with face identification systems that are expression-independent and facial expressions recognition systems that are person-independent. This paper presents a new framework using sparse representation for simultaneous recognition of facial expression and identity. Our framework is based on the assumption that any facial appearance is a sparse combination of identities and expressions (i.e., one identity and one expression). Our experimental results using the CK+ and MMI face datasets show that the proposed approach outperforms methods that conduct face identification and face recognition individually.
Mohammad Reza Mohammadi, Emad Fatemizadeh, Mohammad H. Mahoor
WACV3
2014 A lp-norm MTMKL framework for simultaneous detection of multiple facial action units
abstract
Facial action unit (AU) detection is a challenging topic in computer vision and pattern recognition. Most existing approaches design classifiers to detect AUs individually or AU combinations without considering the intrinsic relations among AUs. This paper presents a novel method, lp-norm multi-task multiple kernel learning (MTMKL), that jointly learns the classifiers for detecting the absence and presence of multiple AUs. lp-norm MTMKL is an extension of the regularized multi-task learning, which learns shared kernels from a given set of base kernels among all the tasks within Support Vector Machines (SVM). Our approach has several advantages over existing methods: (1) AU detection work is transformed to a MTL problem, where given a specific frame, multiple AUs are detected simultaneously by exploiting their inter-relations; (2) lp-norm multiple kernel learning is applied to increase the discriminant power of classifiers. Our experimental results on the CK+ and DISFA databases show that the proposed method outperforms the state-of-the-art methods for AU detection.
Xiao Zhang 0003, Mohammad H. Mahoor, Seyed Mohammad Mavadati, Jeffrey F. Cohn
WACV2
2014 Nonverbal social withdrawal in depression: Evidence from manual and automatic analyses
Jeffrey M. Girard, Jeffrey F. Cohn, Mohammad H. Mahoor, Seyed Mohammad Mavadati, Zakia Hammal, Dean P. Rosenwald
Image Vis. Comput.3
2014 PCA-based dictionary building for accurate facial expression recognition via sparse representation
Mohammad Reza Mohammadi, Emad Fatemizadeh, Mohammad H. Mahoor
J. Vis. Commun. Image Represent.3
2014 Human activity recognition using multi-features and multiple kernel learning
Salah Althloothi, Mohammad H. Mahoor, Xiao Zhang 0003, Richard M. Voyles
Pattern Recognit.2
2014 Non-negative sparse decomposition based on constrained smoothed ℓ0 norm
Mohammad Reza Mohammadi, Emad Fatemizadeh, Mohammad H. Mahoor
Signal Process.3
2013 Head Movement Dynamics during Normal and Perturbed Parent-Infant Interaction
abstract
We investigated the dynamics of head motion in parents and infants during an age-appropriate, well-validated emotion induction, the Face-to-Face/Still-Face procedure. Participants were 12 ethnically diverse 6-month-old infants and their mother or father. During infant gaze toward the parent, infant angular amplitude and velocity of pitch and yaw decreased from face-to-face (FF) to still-face (SF) episodes and remained lower in the following Reunion (RE). During infant gaze away from the parent, angular velocity of pitch decreased from FF to SF and remained lower in the RE. Windowed cross-correlation suggested strong bidirectional effects with frequent shifts in the direction of influence. The number of significant positive and negative peaks was higher during FF than RE. Gaze toward and away from the parent was modestly predicted by head orientation. Together, these findings suggest that head motion is strongly related to age-appropriate emotion challenge, are consistent with the hypothesis that perturbations of normal responsiveness carry-over even after the parent resumes normal responsiveness in the reunion, and that there are frequent changes in direction of influence in the postural domain.
Zakia Hammal, Jeffrey F. Cohn, Daniel S. Messinger, Whitney I. Mattson, Mohammad H. Mahoor
ACII5
2013 Mobile robot connectivity maintenance based on RF mapping
abstract
This paper presents a method for proactive robot communication connectivity maintenance based on electromagnetic field (EMF) recognition and signal strength (SS) gradient estimation for mobile robots. To achieve these goals in an efficient manner, we combine EMF recognition method and gradient descent of SS measurements into a proactive robot motion control algorithm in a way that maintains connectivity among mobile robots in the presence of a radio frequency (RF) obstacle. The EMF recognition method utilizes hidden Markov models (HMMs) for learning EMF environments based on SS measurements. The proposed motion control algorithm uses the EMF recognition and gradient method results to drive the robots towards favorable locations in which robots can communicate. The numerical simulation demonstrates promising EMF recognition, robot motion control results and confirms their abilities in proactive robot motion control for connectivity maintenance.
Mustafa A. Ayad, Jun Jason Zhang, Richard M. Voyles, Mohammad H. Mahoor
IROS4
2013 DISFA: A Spontaneous Facial Action Intensity Database
abstract
Access to well-labeled recordings of facial expression is critical to progress in automated facial expression recognition. With few exceptions, publicly available databases are limited to posed facial behavior that can differ markedly in conformation, intensity, and timing from what occurs spontaneously. To meet the need for publicly available corpora of well-labeled video, we collected, ground-truthed, and prepared for distribution the Denver intensity of spontaneous facial action database. Twenty-seven young adults were video recorded by a stereo camera while they viewed video clips intended to elicit spontaneous emotion expression. Each video frame was manually coded for presence, absence, and intensity of facial action units according to the facial action unit coding system. Action units are the smallest visibly discriminable changes in facial action; they may occur individually and in combinations to comprise more molar facial expressions. To provide a baseline for use in future research, protocols and benchmarks for automated action unit intensity measurement are reported. Details are given for accessing the database for research in computer vision, machine learning, and affective and behavioral science.
Seyed Mohammad Mavadati, Mohammad H. Mahoor, Kevin Bartlett, Philip Trinh, Jeffrey F. Cohn
IEEE Trans. Affect. Comput.2
2013 A Robust Method for Rotation Estimation Using Spherical Harmonics Representation
abstract
This paper presents a robust method for 3D object rotation estimation using spherical harmonics representation and the unit quaternion vector. The proposed method provides a closed-form solution for rotation estimation without recurrence relations or searching for point correspondences between two objects. The rotation estimation problem is casted as a minimization problem, which finds the optimum rotation angles between two objects of interest in the frequency domain. The optimum rotation angles are obtained by calculating the unit quaternion vector from a symmetric matrix, which is constructed from the two sets of spherical harmonics coefficients using eigendecomposition technique. Our experimental results on hundreds of 3D objects show that our proposed method is very accurate in rotation estimation, robust to noisy data, missing surface points, and can handle intra-class variability between 3D objects.
Salah Althloothi, Mohammad H. Mahoor, Richard M. Voyles
IEEE Trans. Image Process.2
2012 Automatic detection of non-posed facial action units
abstract
Automatic facial expression recognition has received great attention in the past two decades due to many applications, such as developmental psychology and human-computer interface design. In most of the current studies, less attention has been paid to the recognition of non-posed facial expressions or measuring their intensity levels. In this paper, we first introduce a novel spontaneous facial expression database called DISFA, which contains videos of 27 young adults, expressing non-posed facial expressions. In this database, the absence and presence of 12 action units (AUs) as well as their intensity levels (i.e., 0-5 scale) were coded by a facial action coding system (FACS) coder. Then, we present an automatic system which can detect facial AUs described by FACS. We compared different facial representation techniques and classifiers for automatic AU detection. As a result of our experiments, we achieved 97% average detection rate using localized Gabor features with SVM classifiers.
Seyed Mohammad Mavadati, Mohammad H. Mahoor, Kevin Bartlett, Philip Trinh
ICIP2
2011 Development of a Wearable Sensor System for Measuring Body Joint Flexion
abstract
This paper presents a novel approach for measuring and monitoring human body joint angles using wearable sensors. This type of monitoring is beneficial for therapists and physicians as it allows them to assess patients' activities remotely. In our approach multiple flex-sensors are mounted on supportive cloth to measure the flexion of a joint. The changes in the resistivity of the flex-sensors are measured using an electronic board. We utilize an Extended Kalman Filter (EKF) to predict the joint angle based on the dynamic model of the joint movement and the measurements obtained from the flex-sensors. Due to variations in the measured angle by each sensor, the outputs are fussed to reduce the error and estimate the best value for the actual body joint angle. We evaluated the effectiveness and performance of our approach for measuring knee joint angle by comparing with the measured angles using goniometer. The result shows that the average of error is 6.92Ë with correlation of 0.98.
Saba Bakhshi, Mohammad H. Mahoor
BSN2
2011 Facial action unit recognition with sparse representation
abstract
This paper presents a novel framework for recognition of facial action unit (AU) combinations by viewing the classification as a sparse representation problem. Based on this framework, we represent a facial image exhibiting the combination of AUs as a sparse linear combination of basis constituting an overcomplete dictionary. We build an overcomplete dictionary whose main elements are mean Gabor features of AU combinations under examination. The other elements of the dictionary are randomly sampled from a distribution (e.g., Gaussian distribution) that guarantees sparse signal recovery. Afterwards, by solving L1-norm minimization, a facial image is represented as a sparse vector which is used to distinguish various AU patterns. After calculating the sparse representation, the classification problem is simply viewed as a rank maximal problem. The index of the maximal value of the sparse vector is regarded as the class label of the facial image under test. Extensive experiments on the Cohn-Kanade facial expressions database demonstrate that this sparse learning framework is promising for recognition of AU combinations.
Mohammad H. Mahoor, Mu Zhou, Kevin L. Veon, Seyed Mohammad Mavadati, Jeffrey F. Cohn
FG1
2011 Video stabilization using SIFT-ME features and fuzzy clustering
abstract
We propose a digital video stabilization process using information that the scale-invariant feature transform (SIFT) provides for each frame. We use a fuzzy clustering scheme to separate the SIFT features representing global motion from those representing local motion. We then calculate the global orientation change and translation between the current frame and the previous frame. Each frame's translation and orientation is added to an accumulated total, and a Kalman filter is applied to estimate the desired motion. We provide experimental results from five video sequences using peak signal-to-noise ratio (PSNR) and qualitative analysis.
Kevin L. Veon, Mohammad H. Mahoor, Richard M. Voyles
IROS2
2011 Analysis of eye gaze pattern of infants at risk of autism spectrum disorder using Markov models
abstract
This paper presents the possibility of using pattern recognition algorithms of infant gaze patterns at six months of age among children at high risk for an autism spectrum disorder (ASD). ASDs, which must be diagnosed by 3 years of age, are characterized by communication and interaction impairments which frequently involve disturbances of visual attention and gaze patterning. We used video cameras to record the face-to-face interactions of 32 infant subjects with their parents. The video was manually coded to determine the eye gaze pattern of infants by marking where the infant was looking in each frame (either at their parent's face or away from their parent's face). In order to identify infants ASD diagnosis at three years, we analyzed infant eye gaze patterns at six months. Variable-order Markov Models (VMM) were used to create models for typically developing comparison children as well as children with an ASD. The models correctly classified infants who did and did not develop an ASD diagnosis with an accuracy rate of 93.75 percent. Employing an assessment tool at a very young age offers the hope of early intervention, potentially mitigating the effects of the disorder throughout the rest of the child's life.
David Alie, Mohammad H. Mahoor, Whitney I. Mattson, Daniel R. Anderson, Daniel S. Messinger
WACV2
2011 Localized support vector machines using Parzen window for incomplete sets of categories
abstract
This paper describes a novel approach to pattern classification that combines Parzen window and support vector machines. Pattern classification is usually performed in universes where all possible categories are defined. Most of the current supervised learning classification techniques do not account for undefined categories. In a universe that is only partially defined, there may be objects that do not fall into the known set of categories. It would be a mistake to always classify these objects as a known category. We propose a Parzen window-based approach which is capable of classifying an object as not belonging to a known class. In our approach we use a Parzen window to identify local neighbors of a test point and train a localized support vector machine on the identified neighbors. Visual category recognition experiments are performed to compare the results of our approach, localized support vector machines using a k-nearest neighbors approach, and global support vector machines. Our experiments show that our Parzen window approach has superior results when testing with incomplete sets, and comparable results when testing with complete sets.
Kevin L. Veon, Mohammad H. Mahoor
WACV2
2009 Detecting Local Audio-visual Synchrony in Monologues Utilizing Vocal Pitch and Facial Landmark Trajectories
abstract
We describe a novel approach for determining the audio-visual synchrony of a monologue video sequence utilizing vocal pitch and facial landmark trajectories as descriptors of the audio and visual modalities, respectively. The visual component is represented by the horizontal and vertical displacement of corresponding facial landmarks between subsequent frames. These facial landmarks are acquired using the statistical modeling technique, known as the Active Shape Model (ASM). The audio component is represented by the fundamental frequency, or pitch, obtained using the subharmonic-to-harmonic ratio (SHR). The synchrony between the audio and visual feature vectors is computed using Gaussian mutual information. The raw synchrony estimates obtained using this method may contain spurious synchrony values due to over-sensitivity. A filtering method is employed for discarding synchrony values that occur during non-associated audio/visual events. The human visual system is capable of distinguishing rigid and non-rigid motion of an articulator during speech. In an attempt to emulate this process, we separate rigid and non-rigid motion and compute the synchrony attributed to each. Experiments are conducted on a dataset of monologue video clip pairs. Each pair is composed of an asynchronous and synchronous version of the video clip. For the asynchronous video clips, the audio signal is displaced with respect to the visual signal. Experimental results indicate that the proposed approach is successful in detecting facial regions that demonstrate synchrony, and in distinguishing between synchronous and asynchronous sequences. © 2009. The copyright of this document resides with its authors.
Steven Cadavid, Mohamed Abdel-Mottaleb, Daniel S. Messinger, Mohammad H. Mahoor, Lorraine E. Bahrick
BMVC4
2009 Multi-modal ear and face modeling and recognition
abstract
In this paper we describe a multi-modal ear and face biometric system. The system is comprised of two components: a 3D ear recognition component and a 2D face recognition component. For the 3D ear recognition, a series of frames is extracted from a video clip and the region of interest (i.e., ear) in each frame is independently reconstructed in 3D using Shape From Shading. The resulting 3D models are then registered using the iterative closest point algorithm. We iteratively consider each model in the series as a reference model and calculate the similarity between the reference model and every model in the series using a similarity cost function. Cross validation is performed to assess the relative fidelity of each 3D model. The model that demonstrates the greatest overall similarity is determined to be the most stable 3D model and is subsequently enrolled in the database. For the 2D face recognition, a set of facial landmarks is extracted from frontal facial images using the Active Shape Model. Then, the response of facial images to a series of Gabor filters at the locations of facial landmarks are calculated. The Gabor features (attributes) are stored in the database as the face model for recognition. The similarity between the Gabor features of a probe facial image and the reference models are utilized to determine the best match. The match scores of the ear recognition and face recognition modalities are fused to boost the overall recognition rate of the system. Experiments are conducted using a gallery set of 402 video clips and a probe of 60 video clips (images). As a result, a rank-one identification rate of 100% was achieved using the weighted sum technique for fusion.
Mohammad H. Mahoor, Steven Cadavid, Mohamed Abdel-Mottaleb
ICIP1
2009 Fast image blending using watersheds and graph cuts
Nuno Gracias, Mohammad H. Mahoor, Shahriar Negahdaripour, Arthur Gleason
Image Vis. Comput.2
2009 A multimodal approach for 3D face modeling and recognition using 3D deformable facial mask
A-Nasser Ansari, Mohamed Abdel-Mottaleb, Mohammad H. Mahoor
Mach. Vis. Appl.3
2009 Face recognition based on 3D ridge images obtained from range data
Mohammad H. Mahoor, Mohamed Abdel-Mottaleb
Pattern Recognit.1
2008 Multi-modal (2-D and 3-D) face modeling and recognition using Attributed Relational Graph
abstract
In this paper we present a unified graph model, called Attributed Relational Graph (ARG), for multi-modal face modeling and recognition. Based on the ARG model, the 2-D and 3-D data are included in a single model. The developed ARG model consists of nodes, edges, and mutual relations. The nodes of the graph correspond to the landmark points that are extracted by an improved Active Shape Model (ASM) technique. Then, at each node of the graph, the responses of a set of log-Gabor filters to the facial image texture and shape information (depth values) are calculated; the filter responses are used to model the local structure of the face at each node of the graph. The edges of the graph are defined based on Delaunay triangulation and a set of mutual relations between the sides of the triangles are defined. The mutual relations boost the final performance of the system. The results of face matching using the 2-D and 3-D attributes and the mutual relations are fused at the score level. A rank-one identification rate of 99% is achieved by experimenting on the University of Miami face database.
Mohammad H. Mahoor, A-Nasser Ansari, Mohamed Abdel-Mottaleb
ICIP1
2008 A Multimodal Approach for Face Modeling and Recognition
abstract
In this paper, we present a fully automated multi- modal (3-D and 2-D) face recognition system. For the 3-D modality, we model the facial image as a 3-D binary ridge image that contains the ridge lines on the face. We use the principal curvature to extract the locations of the ridge lines around the important facial regions on the range image (i.e., the eyes, the nose, and the mouth.) For matching, we utilize a fast variant of the iterative closest point to match the ridge image of a given probe image to the archived ridge images in the database. The main advantage of this approach is reducing the computational complexity by two orders of magnitude by relying on the ridge lines. For the 2-D modality, we model the face by an attributed relational graph (ARG), where each node of the graph corresponds to a facial feature point. At each facial feature point, a set of attributes is extracted by applying Gabor wavelets to the 2-D image and assigned to the node of the graph. The edges of the graph are defined based on Delaunay triangulation and a set of geometrical features that defines the mutual relations between the edges is extracted from the Delaunay triangles and stored in the ARG model. The similarity measure between the ARG models that represent the probe and gallery images is used for 2-D face recognition. Finally, we fuse the matching results of the 3-D and the 2-D modalities at the score level to improve the overall performance of the system. Different techniques for fusion, such as the Dempster-Shafer theory of evidence and weighted sum of scores are employed and tested using the facial images in the third experiment dataset of the Face Recognition Grand Challenge version 2.0.
Mohammad H. Mahoor, Mohamed Abdel-Mottaleb
IEEE Trans. Inf. Forensics Secur.1
2007 3D Face Mesh Modeling from Range Images for 3D Face Recognition
abstract
We present an algorithm for 3D face deformation and modeling using range data captured by a 3D scanner. Using only three facial feature points extracted from the range images and a 3D generic face model, the algorithm first aligns the 3D model to the entire range data of a given subject's face. Then each aligned triangle of the mesh model, with three vertices, is treated as a surface plane which is then fitted to the corresponding interior 3D range data, using least squares plane fitting. Via triangular vertices subdivisions, a higher resolution model is generated from the coordinates of the aligned and fitted model. Finally the model and its triangular surfaces are fitted once again resulting in a smoother mesh model that resembles and captures the surface characteristic of the face. Application of the final deformed model in 3D face recognition, using a publicly available database, shows promising results.
A-Nasser Ansari, Mohamed Abdel-Mottaleb, Mohammad H. Mahoor
ICIP (4)3
2007 3D Face Recognition Based on 3D Ridge Lines in Range Data
abstract
In this paper we present an approach for 3D face recognition from range data based on the principal curvature, kmax, and Hausdorff distance. We use the principal curvature, kmax, to represent the face image as a 3D binary image called ridge image. The ridge image shows the locations of the ridge lines around the important facial regions on the face (i.e. the eyes, the nose, and the mouth). We utilize Hausdorff distance to match the ridge image of a given probe to the created ridge images of the subjects in the gallery. For pose alignment, we extract the locations of three feature points, the inner corners of the two eyes and the tip of the nose using Gaussian curvature. These three feature points plus an auxiliary point in the center of the triangle, made by averaging the coordinates of the three feature points, are used for initial 3D face alignment. In the face recognition stage, we find the optimum pose alignment between the probe image and the gallery, which gives the minimum Hasusdorff distance between the two sets of features. This approach is used for identification of both neutral faces and faces with smile expression. Experiments on a public face database of 61 subjects resulted in 93.5% ranked one recognition rate for neutral expression and 82.0% for the faces with smile expression.
Mohammad H. Mahoor, Mohamed Abdel-Mottaleb
ICIP (1)1
2006 Fast Image Blending using Watersheds and Graph Cuts
abstract
This paper presents a novel approach for combining a set of registered images into a composite mosaic with no visible seams and minimal texture distortion. To promote execution speed in building large area mosaics, the mosaic space is divided into disjoint regions of image intersection based on a geometric criterion. Pair-wise image blending is performed independently in each region by means of watershed segmentation and graph cut optimization. A contribution of this work – use of watershed segmentation to find possible cuts over areas of low photometric difference – allows for searching over a much smaller set of watershed segments, instead of over the entire set of pixels in the intersection zone. The proposed method presents several advantages. The use of graph cuts over image pairs guarantees the globally optimal solution for each intersection region. The independence of such regions makes the algorithm suitable for parallel implementation. The separated use of the geometric and photometric criteria frees the need for a weighting parameter. Finally, it allows the efficient creation of large mosaics, without user intervention. We illustrate the performance of the approach on image sequences with prominent 3D content and moving objects. 1
Nuno Gracias, Arthur Gleason, Shahriar Negahdaripour, Mohammad H. Mahoor
BMVC4
2006 Disparity-Based 3D Face Modeling for 3D Face Recognition
abstract
We present an automatic disparity-based approach for 3D face modeling, from two frontal and one profile view stereo images, for 3D face recognition applications. Once the images are captured, the algorithm starts by extracting selected 2D facial features from one of the frontal views and computes a dense disparity map from the two frontal images. We then align a low resolution 2D mesh model to the selected features, adjust some of its vertices along the profile line using the profile view, increase its triangular vertices to a higher resolution, and re-project them back on the frontal image. Using the coordinates of the re-projected vertices and their corresponding disparities, we capture and compute the 3D facial shape variations using stereo vision. The final result is a deformed 3D model specific to a given subject's face. Application of the model in 3D face recognition validates the algorithm and shows a promising 98 % recognition rate.
A-Nasser Ansari, Mohamed Abdel-Mottaleb, Mohammad H. Mahoor
ICIP3
2006 Disparity-Based 3D Face Modeling using 3D Deformable Facial Mask for 3D Face Recognition
abstract
We present an automatic disparity-based approach for 3D face modeling, from two frontal and one profile view stereo images, for 3D face recognition applications. Once the images are captured, the algorithm starts by extracting selected 2D facial features from one of the frontal views and computes a dense disparity map from the two frontal images. Using the extracted 2D features plus their corresponding disparities in the disparity map, we compute their 3D coordinates. We next align a low resolution 3D mesh model to the 3D features, re-project its vertices on the frontal 2D image and adjust its profile line vertices using the profile view. We increase the resolutions of the resulting 2D model only at its center region to obtain a facial mask model covering distinctive features of the face. The computation of the 2D vertices coordinates with their disparities results in a deformed 3D model mask specific to a give subject face. Application of the model in 3D face recognition validates the algorithm and shows a high recognition rate
A-Nasser Ansari, Mohamed Abdel-Mottaleb, Mohammad H. Mahoor
ICME3
2005 Classification and numbering of teeth in dental bitewing images
Mohammad H. Mahoor, Mohamed Abdel-Mottaleb
Pattern Recognit.1
2004 Automatic classification of teeth in bitewing dental images
abstract
We present an automated algorithm to classify teeth in bitewing dental images, using Bayesian classification, and assign an absolute number to each tooth based on common numbering system used in dentistry. Fourier descriptors of the contours of the molar and the premolar teeth in bitewing images are used in the Bayesian classification of these two types of the teeth. Then, the spatial relation between the two types of the teeth is considered to number each tooth and correct the misclassification of some teeth in order to obtain high precision results. Experiments with 50 bitewing images containing more than 400 teeth show that our method is capable of classifying and assigning absolute index number to the teeth with high accuracy.
Mohammad H. Mahoor, Mohamed Abdel-Mottaleb
ICIP1