Ayoub Al-Hamadi

dblp:51/6942 · DBLP profile ↗
← Back
91ranked-venue papers
2as first author
15since 2021 · last 2026
0000-0002-3632-2402ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 55 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 38 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 3 since 2021Human-computer interaction and ubiquitous computing · 14 · 3 since 2021Databases, data management, data science and information retrieval · 4
YearPublicationVenuePosition
2026 Constructing a multimodal feature set for pain intensity classification
abstract
Abstract Previous approaches to pain intensity classification have typically relied on small sets of top-performing features to maximize accuracy. While effective in constrained scenarios, such strategies neglect the diverse range of modalities available in modern pain databases. In this work, we introduce a multimodal feature set (MMFS) that integrates heterogeneous features from each biosignal modality in the BioVid and X-ITE databases. Our approach captures a broad spectrum of complementary information, maintaining robustness even when individual modalities are unavailable. Experimental results show consistent performance improvements, with classification accuracy increasing by up to 8% overall and by 6% when evaluating all pain levels. Through detailed analyses of individual modalities, confusion matrices, and a modality ablation study, we demonstrate that the combined effect of multimodality and balanced information distribution drives these gains. Furthermore, feature importance analysis reveals which inputs contribute most to final predictions and which features are most beneficial for the more challenging pain classes.
Sören Nienaber, Laslo Dinges, Ayoub Al-Hamadi
Neural Comput. Appl.4
2025 Gaze Matters: Eye Contact Detection in Unscripted Human-Robot Interaction Scenarios
abstract
In human-robot interaction (HRI), eye contact is a crucial mechanism in nonverbal communication, yet its detection from the robot’s perspective remains largely underexplored. This paper presents a user study in a realistic collaborative HRI scenario with unscripted tasks to evaluate the performance of current eye contact detection models under realistic conditions. As part of this study, a new dataset of 5,888 manually annotated visual engagement instances was created to reflect the complexity of real-world interactions and to validate our previous work on NITEC, a large-scale dataset for eye contact detection. The results indicate that models trained on the NITEC dataset perform best with an AUC of 0.74 compared to other models, highlighting the importance of training on diverse and unconstrained data for robust eye contact recognition in HRI. These findings highlight both the technical challenges and the potential for more human-centric robotic systems.
Thorsten Hempel, Magnus Jung, Dominykas Strazdas, Ayoub Al-Hamadi
SMC4
2025 A Real-Time Digital Twin Framework for the TIAGo Service Robot
abstract
The rapid development of humanoid robots to perform human-like tasks has introduced new possibilities in automation. In the context of advanced autonomous robotic operations, the potential damage due to unanticipated actions is a significant concern. To address this challenge, this paper presents a digital twin designed to replicate and monitor robotic operations in a virtual environment. Key features include the integration of manufacturer-provided kinematic models, real-time sensor fusion, and a user-friendly graphical interface to support reinforcement learning applications. The digital twin successfully mirrors the movements of a physical robot and records visual data for downstream machine learning tasks. Further, we implemented integration of procedurally generated virtual environments that allows for dynamic scenario creation, enabling robots to train and adapt to diverse and unpredictable conditions without the risks associated with real-world testing. This paves the way toward transitioning our entire laboratory into a digital ecosystem, providing a safe and controlled environment where robots can operate, learn, and improve autonomously.
Malte Herrmann, Thorsten Hempel, Ayoub Al-Hamadi
SMC3
2025 Multi-Modal AI-Based Pain Detection in Intermediate Care Patients in the Postoperative Phase
abstract
"Multi-modal AI-Based Pain Detection in Intermediate Care Patients in the Postoperative Phase" is an interdisciplinary research work that operates in the domain of automated pain detection. It aims to improve previous work, based on pain databases like BioVid and UNBC shoulder pain, as well as AI-based approaches using computer vision and signal processing to analyze available modalities. Thus, we present our basic research idea on how to improve automatic pain detection in three major steps. The first step focuses on collecting pain data from postoperative patients in intermediate care stations (IMC). In addition, patients who are not fully oriented should be included in a separate data collection as a second focus group. Then, improvements on the state-of-the-art models should not only advance general pain detection, but also help bridge the gap to the real-world setting of the IMC data. Improvements include transferability analysis, feature selection evaluation, and balancing of data distribution to deliver better classification performance. In a last step, we aim to test, verify and evaluate the classification performance on the IMC data with the support of medical practitioners.
Sören Nienaber, Thorsten Hempel, Steffen Walter 0001, Eberhard Barth, Ayoub Al-Hamadi
SMC6
2025 Towards efficient and robust face recognition through attention-integrated multi-level CNN
abstract
Abstract The rapid advancement of deep Convolutional Neural Networks (CNNs) has led to remarkable progress in computer vision, contributing to the development of numerous face verification architectures. However, the inherent complexity of these architectures, often characterized by millions of parameters and substantial computational demands, presents significant challenges for deployment on resource-constrained devices. To address these challenges, we introduce RobFaceNet, a robust and efficient CNN designed explicitly for face recognition (FR). The proposed RobFaceNet optimizes accuracy while preserving computational efficiency, a balance achieved by incorporating multiple features and attention mechanisms. These features include both low-level and high-level attributes extracted from input face images and aggregated from multiple levels. Additionally, the model incorporates a newly developed bottleneck that integrates both channel and spatial attention mechanisms. The combination of multiple features and attention mechanisms enables the network to capture more significant facial features from the images, thereby enhancing its robustness and the quality of facial feature extraction. Experimental results across state-of-the-art FR datasets demonstrate that our RobFaceNet achieves higher recognition performance. For instance, RobFaceNet achieves 95.95% and 92.23% on the CA-LFW and CP-LFW datasets, respectively, compared to 95.45% and 92.08% for very deep ArcFace model. Meanwhile, RobFaceNet exhibits a more lightweight model complexity. In terms of computation cost, RobFaceNet has 337M Floating Point Operations Per Second (FLOPs) compared to ArcFace’s 24211M, with only 3% of the parameters. Consequently, RobFaceNet is well-suited for deployment across various platforms, including robots, embedded systems, and mobile devices.
Aly Khalifa, Ahmed A. Abdelrahman, Thorsten Hempel, Ayoub Al-Hamadi
Multim. Tools Appl.4
2025 Mobgazenet: robust gaze estimation mobile network based on progressive attention mechanisms
abstract
Abstract Gaze estimation is a fundamental task in computer vision with a wide range of applications. Recently, convolution neural network approaches have made notable progress in inferring gaze from facial images. However, these methods often struggle to capture fine-grained gaze features and reflect spatial contextual relationships, as the most crucial gaze information exists in the eye area, which constitutes only a small portion of the face images. In this paper, we introduce MobGazeNet, an efficient and lightweight network that leverages a progressive combination of attention mechanisms, including squeeze-and-excitation, convolutional block attention module, and coordinate attention. The combination of attention mechanisms helps to emphasize crucial eye features and allows the model to consider both local and global spatial relationships without increasing computational overhead. Furthermore, we introduce the rotation matrix formalism for gaze ground truth to avoid discontinuity and ambiguity in spherical angle representation. Building upon this, we propose a continuous 6D rotation matrix representation to enable efficient and reliable direct regression which we further enhance with a geodesic-based loss. To evaluate our model, we conduct experiments on three popular datasets collected in unconstrained settings. Our proposed model surpasses current SOTA methods in both performance and efficiency, showcasing its superior capability in gaze estimation. Our code is available at: https://github.com/Ahmednull/MobGazeNet .
Ahmed A. Abdelrahman, Thorsten Hempel, Aly Khalifa, Dominykas Strazdas, Ayoub Al-Hamadi
Mach. Vis. Appl.5
2024 NITEC: Versatile Hand-Annotated Eye Contact Dataset for Ego-Vision Interaction
abstract
Eye contact is a crucial non-verbal interaction modality and plays an important role in our everyday social life. While humans are very sensitive to eye contact, the capabilities of machines to capture a person’s gaze are still mediocre. We tackle this challenge and present NITEC, a hand-annotated eye contact dataset for ego-vision interaction. NITEC exceeds existing datasets for ego-vision eye contact in size and variety of demographics, social contexts, and lighting conditions, making it a valuable resource for advancing ego-vision-based eye contact research. Our extensive evaluations on NITEC demonstrate strong cross-dataset performance, emphasizing its effectiveness and adaptability in various scenarios, that allows seamless utilization to the fields of computer vision, human-computer interaction, and social robotics. We make our NITEC dataset publicly available to foster reproducibility and further exploration in the field of ego-vision interaction1.
Thorsten Hempel, Magnus Jung, Ahmed A. Abdelrahman, Ayoub Al-Hamadi
WACV4
2024 Fine-grained gaze estimation based on the combination of regression and classification losses
abstract
Abstract Human gaze is a crucial cue used in various applications such as human-robot interaction, autonomous driving, and virtual reality. Recently, convolution neural network (CNN) approaches have made notable progress in predicting gaze angels. However, estimating accurate gaze direction in-the-wild is still a challenging problem due to the difficulty of obtaining the most crucial gaze information that exists in the eye area which constitutes a small part of the face images. In this paper, we introduce a novel two-branch CNN architecture with a multi-loss approach to estimate gaze angles (pitch and yaw) from face images. Our approach utilizes separate fully connected layers for each gaze angle prediction, allowing explicit learning of discriminative features and emphasizing the distinct information associated with each gaze angle. Moreover, we adopt a multi-loss approach, incorporating both classification and regression losses. This allows for joint optimization of the combined loss for each gaze angle, resulting in improved overall gaze performance. To evaluate our model, we conduct experiments on three popular datasets collected under unconstrained settings: MPIIFaceGaze, Gaze360, and RT-GENE. Our proposed model surpasses current state-of-the-art methods and achieves state-of-the-art performance on all three datasets, showcasing its superior capability in gaze estimation.
Ahmed A. Abdelrahman, Thorsten Hempel, Aly Khalifa, Ayoub Al-Hamadi
Appl. Intell.4
2024 Exploring facial cues: automated deception detection using artificial intelligence
abstract
Abstract Deception detection is an interdisciplinary field attracting researchers from psychology, criminology, computer science, and economics. Automated deception detection presents unique challenges compared to traditional polygraph tests, but also offers novel economic applications. In this spirit, we propose an approach combining deep learning with discriminative models for deception detection. Therefore, we train CNNs for the facial modalities of gaze, head pose, and facial expressions, allowing us to compute facial cues. Due to the very limited availability of training data for deception, we utilize early fusion on the CNN outputs to perform deception classification. We evaluate our approach on five datasets, including four well-known publicly available datasets and a new economically motivated rolling dice experiment. Results reveal performance differences among modalities, with facial expressions outperforming gaze and head pose overall. Combining multiple modalities and feature selection consistently enhances detection performance. The observed variations in expressed features across datasets with different contexts affirm the importance of scenario-specific training data for effective deception detection, further indicating the influence of context on deceptive behavior. Cross-dataset experiments reinforce these findings. Notably, low-stake datasets, including the rolling dice Experiment, present more challenges for deception detection compared to the high-stake Real-Life trials dataset. Nevertheless, various evaluation measures show deception detection performance surpassing chance levels. Our proposed approach and comprehensive evaluation highlight the challenges and potential of automating deception detection from facial cues, offering promise for future research.
Laslo Dinges, Marc-André Fiedler, Ayoub Al-Hamadi, Thorsten Hempel, Ahmed A. Abdelrahman, Joachim Weimann, Dmitri Bershadskyy, Johann Steiner
Neural Comput. Appl.3
2024 Toward Robust and Unconstrained Full Range of Rotation Head Pose Estimation
abstract
Estimating the head pose of a person is a crucial problem for numerous applications that is yet mainly addressed as a subtask of frontal pose prediction. We present a novel method for unconstrained end-to-end head pose estimation to tackle the challenging task of full range of orientation head pose prediction. We address the issue of ambiguous rotation labels by introducing the rotation matrix formalism for our ground truth data and propose a continuous 6D rotation matrix representation for efficient and robust direct regression. This allows to efficiently learn full rotation appearance and to overcome the limitations of the current state-of-the-art. Together with new accumulated training data that provides full head pose rotation data and a geodesic loss approach for stable learning, we design an advanced model that is able to predict an extended range of head orientations. An extensive evaluation on public datasets demonstrates that our method significantly outperforms other state-of-the-art methods in an efficient and robust manner, while its advanced prediction range allows the expansion of the application area. We open-source our training and testing code along with our trained models: https://github.com/thohemp/6DRepNet360.
Thorsten Hempel, Ahmed A. Abdelrahman, Ayoub Al-Hamadi
IEEE Trans. Image Process.3
2023 Classification networks for continuous automatic pain intensity monitoring in video using facial expression on the X-ITE Pain Database
abstract
So far, the current methods in the clinical application do not facilitate continuous monitoring for pain and are unreliable, especially for vulnerable patients. In contrast, several automated methods have been proposed for this task by using facial features that were extracted independently from every frame of a given sequence. However, the obtained results were poor due to the failure to represent movement dynamics. To solve this problem, this work introduces three distinct methods regarding classification to monitor continuous pain intensity: (1) A Random Forest classifier (RFc) baseline method, (2) Long-Short Term Memory (LSTM) method, and (3) LSTM using sample weighting method (LSTM-SW). In this study, we conducted experiments with 11 datasets regarding classification, then compared results to regression results in Othman et al. (2021). Experimental results showed that the LSTM & LSTM-SW methods for continuous automatic pain intensity recognition performed better than guessing and RFc except with small datasets such as the reduced tonic datasets.
Ehsan Othman, Philipp Werner, Frerk Saxen, Ayoub Al-Hamadi, Sascha Gruss, Steffen Walter 0003
J. Vis. Commun. Image Represent.4
2023 JAMsFace: joint adaptive margins loss for deep face recognition
abstract
Abstract Deep feature learning has become crucial in large-scale face recognition, and margin-based loss functions have demonstrated impressive success in this field. These methods aim to enhance the discriminative power of the softmax loss by increasing the feature margin between different classes. These methods assume class balance, where a fixed margin is sufficient to squeeze intra-class variation equally. However, real-face datasets often exhibit imbalanced classes, where the fixed margin is suboptimal, limiting the discriminative power and generalizability of the face recognition model. Furthermore, margin-based approaches typically focus on enhancing discrimination either in the angle or cosine space, emphasizing one boundary while disregarding the other. To overcome these limitations, we propose a joint adaptive margins loss function (JAMsFace) that learns class-related margins for both angular and cosine spaces. This approach allows adaptive margin penalties to adjust adaptively for different classes. We explain and analyze the proposed JAMsFace geometrically and present comprehensive experiments on multiple face recognition benchmarks. The results show that JAMsFace outperforms existing face recognition losses in mainstream face recognition tasks. Specifically, JAMsFace advances the state-of-the-art face recognition performance on LFW, CPLFW, and CFP-FP and achieves comparable results on CALFW and AgeDB-30. Furthermore, for the challenging IJB-B and IJB-C benchmarks, JAMsFace achieves impressive true acceptance rates (TARs) of 89.09% and 91.81% at a false acceptance rate (FAR) of 1e-4, respectively.
Aly Khalifa, Ayoub Al-Hamadi
Neural Comput. Appl.2
2022 6d Rotation Representation For Unconstrained Head Pose Estimation
abstract
In this paper, we present a method for unconstrained end-to-end head pose estimation. We address the problem of ambiguous rotation labels by introducing the rotation matrix formalism for our ground truth data and propose a continuous 6D rotation matrix representation for efficient and robust direct regression. This way, our method can learn the full rotation appearance which exceeds the capabilities of previous approaches that restrict the pose prediction to a narrow-angle for satisfactory results. In addition, we propose a geodesic distance-based loss to penalize our network with respect to the SO(3) manifold geometry. Experiments on the public AFLW2000 and BIWI datasets demonstrate that our proposed method significantly outperforms other state-of-the-art methods by up to 20%. We open-source our training and testing code along with our trained models: https://github.com/thohemp/6DRepNet.
Thorsten Hempel, Ahmed A. Abdelrahman, Ayoub Al-Hamadi
ICIP3
2022 An online semantic mapping system for extending and enhancing visual SLAM
Thorsten Hempel, Ayoub Al-Hamadi
Eng. Appl. Artif. Intell.2
2022 Automatic Recognition Methods Supporting Pain Assessment: A Survey
abstract
Pain is a complex phenomenon, involving sensory and emotional experience, that is often poorly understood, especially in infants, anesthetized patients, and others who cannot speak. Technology supporting pain assessment has the potential to help reduce suffering; however, advances are needed before it can be adopted clinically. This survey paper assesses the state of the art and provides guidance for researchers to help make such advances. First, we overview pain’s biological mechanisms, physiological and behavioral responses, emotional components, as well as assessment methods commonly used in the clinic. Next, we discuss the challenges hampering the development and validation of pain recognition technology, and we survey existing datasets together with evaluation methods. We then present an overview of all automated pain recognition publications indexed in the Web of Science as well as from the proceedings of the major conferences on biomedical informatics and artificial intelligence, to provide understanding of the current advances that have been made. We highlight progress in both non-contact and contact-based approaches, tools using face, voice, physiology, and multi-modal information, the importance of context, and discuss challenges that exist, including identification of ground truth. Finally, we identify underexplored areas such as chronic pain and connections to treatments, and describe promising opportunities for continued advances.
Philipp Werner, Daniel Lopez Martinez, Steffen Walter 0001, Ayoub Al-Hamadi, Sascha Gruss, Rosalind W. Picard
IEEE Trans. Affect. Comput.4
2019 Generalizing to Unseen Head Poses in Facial Expression Recognition and Action Unit Intensity Estimation
abstract
Facial expression analysis is challenged by the numerous degrees of freedom regarding head pose, identity, illumination, occlusions, and the expressions itself. It currently seems hardly possible to densely cover this enormous space with data for training a universal well-performing expression recognition system. In this paper we address the sub-challenge of generalizing to head poses that were not seen in the training data, aiming at getting along with sparse coverage of the pose subspace. For this purpose we (1) propose a novel face normalization method called FaNC that massively reduces pose-induced image variance; (2) we compare the impact of the proposed and other normalization methods on (a) action unit intensity estimation with the FERA 2017 challenge data (achieving new state of the art) and (b) facial expression recognition with the Multi-PIE dataset; and (3) we discuss the head pose distribution needed to train a pose-invariant CNN-based recognition system. The proposed FaNC method normalizes pose and facial proportions while retaining expression information and runs in less than 2 ms. When comparing results achieved by training a CNN on the output images of FaNC and other normalization methods, FaNC generalizes significantly better than others to unseen poses if they deviate more than 20° from the poses available during training. Code and data are available.
Philipp Werner, Frerk Saxen, Ayoub Al-Hamadi, Hui Yu 0001
FG3
2019 Detecting Arbitrarily Rotated Faces for Face Analysis
abstract
Current face detection concentrates on detecting tiny faces and severely occluded faces. Face analysis methods, however, require a good localization and would benefit greatly from some rotation information. We propose to predict a face direction vector (FDV), which provides the face size and orientation and can be learned by a common object detection architecture better than the traditional bounding box. It provides a more consistent definition of face location and size. Using the FDV is promising for all succeeding face analysis methods. As an example, we show that facial landmark detection can highly benefit from pre-aligned faces.
Frerk Saxen, Sebastian Handrich, Philipp Werner, Ehsan Othman, Ayoub Al-Hamadi
ICIP5
2018 3D Human Pose Estimation Using Stochastic Optimization in Real Time
abstract
Random Tree Walkers (RTW) are a well-established method for human pose estimation, because they deliver state-of-the-art performance at low computational cost. As the forests capabilities for generalization are limited, the algorithm fails to estimate unlearned poses very quickly. The proposed method pushes this limitation by combining the RTW with optimization methods such as iterative closest point (ICP) and a stochastic search. The RTW is being used to initialize various hypotheses in different ways which are then passed to the optimization stage of the proposed method. The quality of each hypothesis is assessed by a cost function measuring the discrepancy between the data and a human body model generated for each hypothesis. Experimental results show a greater number of correctly estimated poses over a single RTW result.
Sebastian Handrich, Philipp Waxweiler, Philipp Werner, Ayoub Al-Hamadi
ICIP4
2018 Pattern Optimization for 3D Surface Reconstruction with an Active Line Scan Camera System
abstract
The 3d surface reconstruction with active line-scan camera systems generally require a special approach for generating powerful structured light. For our system, we use a structured light approach based on distributed projection of a LED line light. The pattern generated by this kind of structured light are extremely height dependent and provide an inhomogeneous measurement accuracy. Thus, it is necessary to determine pattern sequences providing precise 3d measurements over the whole measuring range. For this we first define average-free pattern sequences, which can be normalized and evaluated in terms of their local correlation properties. Based on that definition we propose a heuristic algorithm to choose optimal pattern sequences and verify its basic functionality.
Erik Lilienblum, Ayoub Al-Hamadi
ICIP2
2018 How the Region of Interest Impacts Contact Free Heart Rate Estimation Algorithms
abstract
The contact free camera-based estimation of human vital signs is more comfortable than the classical contact-based methods. Current methods suffer in realistic environments from e.g. occlusions by hair or glasses. Several approaches use complex methods to extract the pulse signals from the skin, but use basic geometric defined regions of the face, which ignore possible occlusions. In this paper we compare for the first time the influence of several region based methods and a newly proposed parameter-free skin segmentation approach on the performance of the heart rate estimation for nine different algorithms from the literature on two datasets with strong facial and head movement.
Michal Rapczynski, Philipp Werner, Frerk Saxen, Ayoub Al-Hamadi
ICIP4
2018 A Multi-Spectral Database for NIR Heart Rate Estimation
abstract
This paper presents a new dataset for contact-free heart rate estimation in the near-infrared spectrum. Most published works are either using an approach based on RGB cameras or near-infrared light of a specific wavelength. The data from these experiments is usually not available for other researchers so that no public dataset for multi-spectral heart rate estimation exists. Our new contributed dataset consists of 38 subjects, covering a spectral range of 675nm up to 950nm in 25 narrow spectral bands and allows a thorough comparison of heart rate estimation algorithms and the assessment of the chosen wavelength on the measurement accuracy.
Michal Rapczynski, Ayoub Al-Hamadi, Gunther Notni
ICIP3
2018 Intention-Based Anticipatory Interactive Systems
abstract
Intention-based, anticipatory, interactive systems (IAIS) represent a new class of user-centered assistance systems. IAIS uses actions and system intentions derived from signal data, and the affective state of the user. By anticipating the further action of the user, solutions are interactively negotiated. The active roles of humans and systems change strategically, which requires behavioral models, which in turn can be specified by life sciences' results. Deployed human-machine-systems and lab tests have the goal of understanding of the situated interaction. This supports integration of assistance systems for Industry 4.0 and in the context of demographic change. We provide a definition of IAIS, discuss the underlying requirements, goals and challenges and provide a brief review of the state-of-the-art in the involved research areas.
Andreas Wendemuth, Ronald Böck, Andreas Nürnberger, Ayoub Al-Hamadi, André Brechmann, Frank W. Ohl
SMC4
2018 Facial point localization via neural networks in a cascade regression framework
Anwar Saeed, Ayoub Al-Hamadi, Heiko Neumann
Multim. Tools Appl.2
2017 Facial action unit intensity estimation and feature relevance visualization with random regression forests
abstract
Automatic facial action unit intensity estimation can be useful for various applications in affective computing. In this paper, we apply random regression forests for this task and propose modifications that improve predictive performance compared to the original random forest. Further, we introduce a way to estimate and visualize the relevance of the features for an individual prediction and the forest in general. We conduct experiments on the FERA 2017 challenge dataset (which outperform the FERA baseline results), show the performance gain by the modifications, and illustrate feature relevance.
Philipp Werner, Sebastian Handrich, Ayoub Al-Hamadi
ACII3
2017 Automatic recognition of common Arabic handwritten words based on OCR and N-GRAMS
abstract
Comprehensive databases are vital for training and validation of word recognition systems. To overcome the lack of offline databases of Arabic handwritten words, especially regarding the generality of the underlying vocabulary, we used a synthesis system to generate a database of common Arabic handwritings. Subsequently, we validate a new word recognition system on these synthetic handwritings, to analyze the performance of its segmentation, character recognition, and error correction module. We found, that a dynamic character classifier, that is capable to adapted to the variations that are caused by the segmentation, clearly improves word recognition accuracy. For error detection and correction, n-grams as well as the Levenstein distance to a vocabulary of up to 50,000 valid words have been used.
Laslo Dinges, Ayoub Al-Hamadi, Moftah Elzobi, Andreas Nürnberger
ICIP2
2017 Localizing body joints from single depth images using geodetic distances and random tree walk
abstract
We address the problem of human pose estimation from single depth images. We extend a previously proposed approach made by Jung and learn the direction toward skeleton joints from both depth and geodetic features. Experimental evaluation shows that our approach achieves a higher precision or a similar precision, but with much smaller regression trees.
Sebastian Handrich, Ayoub Al-Hamadi
ICIP2
2017 Landmark based head pose estimation benchmark and method
abstract
Head pose estimation can help in understanding human behavior or to improve head pose invariance in various face analysis applications. Ready-to-use pose estimators are available with several facial landmark trackers, but their accuracy is commonly unknown. Following the goal to find the best landmark based pose estimator, we introduce a new database (called SyLaHP), propose a new benchmark protocol, and describe and implement a method to learn a pose estimator on top of any landmark detector (called HPFL). The experiments (including cross database) reveal that OpenFace comes with the best pose estimator. Further, HPFL models trained on top of landmark trackers outperform the respective built-in pose estimators. The SyLaHP database, source code, and trained models are publicly available for research.
Philipp Werner, Frerk Saxen, Ayoub Al-Hamadi
ICIP3
2017 Automatic Pain Assessment with Facial Activity Descriptors
abstract
Pain is a primary symptom in medicine, and accurate assessment is needed for proper treatment. However, today's pain assessment methods are not sufficiently valid and reliable in many cases. Automatic recognition systems may contribute to overcome this problem by facilitating objective and continuous assessment. In this article we propose a novel feature set for describing facial actions and their dynamics, which we call facial activity descriptors. We apply them to detect pain and estimate the pain intensity. The proposed method outperforms previous state-of-the-art approaches in sequence-level pain classification on both, the BioVid Heat Pain and the UNBC-McMaster Shoulder Pain Expression database. We further discuss major challenges of pain recognition research, benefits of temporal integration, and shortcomings of widely used frame-based pain intensity ground truth.
Philipp Werner, Ayoub Al-Hamadi, Kerstin Limbrecht, Steffen Walter 0001, Sascha Gruss, Harald C. Traue
IEEE Trans. Affect. Comput.2
2016 Continuous low latency heart rate estimation from painful faces in real time
abstract
Video based heart rate estimation has several advantages compared to the classical method. Current approaches use long time windows (30sec) to calculate heart rates, which results in high latency and is a big disadvantage for a practical use. To overcome this constraint, we propose a low latency approach for continuous frame based heart rate estimation. It is based on combination of face tracking and skin detection using short time windows (10sec) to filter and analyze the extracted PPG signals in real time. In experiments the presented approach performs with high accuracy (85,2%, with error <;3 BPM) under stable illumination conditions using a pain recognition data set including facial expressions and head movement for validation.
Michal Rapczynski, Philipp Werner, Ayoub Al-Hamadi
ICPR3
2016 A framework for joint facial expression recognition and point localization
abstract
Unlike many approaches that use detected facial points to infer facial expressions, in this work, we propose an approach in which we jointly tackle the two tasks on a frame basis. After ensuring the consistent face cropping, our framework makes use of geometric- and appearance-based methods for the facial expression recognition, and of cascade regression and local-based methods for the facial point detection. For the data fusion, we adapted the Viterbi algorithm. The training and testing were carried out on two public benchmark databases. With the proposed framework, we improved the recognition rate of the facial expression and the accuracy of the facial point localization by at least 5.4% and 8.9%, respectively, in comparison to the conventional sequential methods.
Anwar Saeed, Ayoub Al-Hamadi
ICPR2
2016 A Multiresolution Approach to Model-Based 3-D Surface Quality Inspection
abstract
We propose a novel model-based surface approximation method for three-dimensional (3-D) surface quality inspection that combines a machine learning approach with multiresolution paradigms. Acceptable surface deviations are modeled by learning a number of 3-D measurements of tolerance samples. At the same time, areas with high surface details are hierarchically refined, allowing an improved spatial localization of the surface model. The method is based on a dual eigenvalue decomposition, which leads to fast computation for large datasets of ordered 3-D point clouds. The proposed algorithm is easy to configure and requires few parameters by automatically determining the areas for local refinement. It yields a better defect detection on deformable parts and parts with high tolerance ranges, especially on critical areas with high surface curvature. Experimental results show the effectiveness compared to model-based approaches without multiresolution as well as nonmodel-based methods. Examples are given for successful defect detection where previous methods have failed.
Sebastian von Enzberg, Ayoub Al-Hamadi
IEEE Trans. Ind. Informatics2
2015 Optical Sensor Tracking and 3D-Reconstruction of Hydrogen-Induced Cracking
Christian Freye, Christian Bendicks, Erik Lilienblum, Ayoub Al-Hamadi
ACIVS4
2015 Full-Body Human Pose Estimation by Combining Geodesic Distances and 3D-Point Cloud Registration
Sebastian Handrich, Ayoub Al-Hamadi
ACIVS2
2015 Handling Data Imbalance in Automatic Facial Action Intensity Estimation
abstract
Automatic Action Unit (AU) intensity estimation is a key problem in facial expression analysis. But limited research attention has been paid to the inherent class imbalance, which usually leads to suboptimal performance. To handle the imbalance, we propose (1) a novel multiclass under-sampling method and (2) its use in an ensemble. We compare our approach with state of the art sampling methods used for AU intensity estimation. Multiple datasets and widely varying performance measures are used in the literature, making direct comparison difficult. To address these shortcomings, we compare different performance measures for AU intensity estimation and evaluate our proposed approach on three publicly available datasets, with a comparison to state of the art methods along with a cross dataset evaluation.
Philipp Werner, Frerk Saxen, Ayoub Al-Hamadi
BMVC3
2015 Utilizing the Bezier descriptors for hand gesture recognition
abstract
In this paper, a novel approach is proposed for hand gesture recognition by modelling the Bezier curves. We have adapted a three-step approach which begins with the skin-based hand segmentation method using the normal Gaussian distribution. It is followed by the feature extraction module where the hand centroid points are computed which are then fitted with Bezier curves. These fitted Bezier curve points are quantized and concatenated to build the Bezier descriptors. The extracted Bezier descriptors are finally classified by Hidden Markov Models (HMM) using Left-Right Banded (LRB) topology for hand gesture recognition. We have tested our proposed approach with different HMM models on the same hand centroid points (i.e., control points) fitted with Bezier curves and compare results. The experimental results show that our proposed approach is capable to detect hands, model the Bezier curves to build descriptors and classify the descriptors for hand gesture recognition in real situations which proves its applicability in the domain of Human Computer Interaction.
Omer Rashid Rashid, Ayoub Al-Hamadi
ICIP2
2015 Boosted human head pose estimation using kinect camera
abstract
Head pose estimation is essential for several computer vision applications. For example, it has been employed in facial expression recognition, head gesture detection, and driver monitoring systems. In this work, we present a boosted method to estimate the head pose using Kinect camera. This estimation is cooperatively performed with the help of RGB and depth images. The human face is located in the RGB image using frontal and profile Viola-Jones (VJ) face detector, where the depth information is used to confine the size and location of the search window. Appearance features, extracted from the detected face patch in the RGB image and its corresponding in the depth image, are passed to Support Vector Machine (SVM) regressors to infer the head pose. Evaluation on two public benchmark databases demonstrates that our proposed approach compares favorably to state-of-the-art approaches.
Anwar Saeed, Ayoub Al-Hamadi
ICIP2
2015 Quantitative Analysis of Surface Reconstruction Accuracy Achievable with the TSDF Representation
Diana Werner, Philipp Werner, Ayoub Al-Hamadi
ICVS3
2015 Towards the Separation of Rigid and Non-rigid Motions for Facial Expression Analysis
abstract
In intelligent environments, computer systems not solely serve as passive input devices waiting for user interaction but actively analyze their environment and adapt their behaviour according to changes in environmental parameters. One essential ability to achieve this goal is to analyze the mood, emotions and dispositions a user experiences while interacting with such intelligent systems. Features allowing to infer such parameters can be extracted from auditive, as well as visual sensory input streams. For the visual feature domain, in particular facial expressions are known to contain rich information about a user's emotional state and can be detected by using either static and/or dynamic image features. During interaction facial expressions are rarely performed in isolation, but most of the time co-occur with movements of the head. Thus, optical flow based facial features are often compromised by additional motions. Parts of the optical flow may be caused by rigid head motions, while other parts reflect deformations resulting from facial expressivity (non-rigid motions). In this work, we propose the first steps towards an optical flow based separation of rigid head motions from non-rigid motions caused by facial expressions. We suggest that after their separation, both, head movements and facial expressions can be used as a basis for the recognition of a user's emotions and dispositions and thus allow a technical system to effectively adapt to the user's state.
Georg Layher, Stephan Tschechne, Robert Niese, Ayoub Al-Hamadi, Heiko Neumann
Intelligent Environments4
2014 Dynamical evaluation Of academic performance in e-learning systems using neural networks modeling (time response approach)
abstract
This paper explores a relatively new methodological approach for the field integrating learning and education, with other research areas, such as neurobiological, cognitive, and computational sciences. Specifically, presented work is an interdisciplinary piece of research aiming to simulate appropriately a challenging and critical issue concerned with academic performance in e-learning systems. Namely, considering face to face tutoring phenomenon observed while an interactive e-learning process is performed. Referring to strong interest announced by educationalists to know how neurons' synapses inside the brain are interconnected. Together to perform communication processing among brain regions. Herein, a special attention has been developed towards dynamical academic evaluation of timely based brain learning via face to face (FTF) interactive tutoring. In other words, this piece of research presents an interdisciplinary realistic dynamic investigation. For academic performance phenomenon associated with e-learners' contribution as time response performed human's brain neuronal function. Accordingly, Artificial Neural Networks (ANNS) have been adopted for realistic modeling of academic performance evaluation based on timely dependant student's response till attaining learning convergence (desired output). After running of designed realistic simulation program, some interesting results have been presented. Interestingly, individual differences' phenomenon observed via after statistical analysis of obtained simulation results.
Hassan M. H. Mustafa, Ayoub Al-Hamadi, Nosipho Dladlu, Nada M. Al-Shenawy, Eyas El-Qwasmeh
EDUCON2
2014 Color-based skin segmentation: An evaluation of the state of the art
abstract
Skin segmentation is widely used, e.g. in face detection and gesture recognition. In the last years, the number of skin segmentation approaches has grown. However, multiple datasets and varying performance measurements make direct comparison difficult. We address these shortcomings and evaluate 5 threshold-based methods, 5 model-based methods, and 2 region-based state-of-the-art skin segmentation methods. We discuss each algorithm and provide the segmentation performance along with the processing time. All methods are evaluated on the ECU dataset which provides a great amount of training data besides other important attributes.
Frerk Saxen, Ayoub Al-Hamadi
ICIP2
2014 Automatic heart rate estimation from painful faces
abstract
Non-contact measurement of the heart rate is more comfortable than classical methods and can facilitate new applications. However, current approaches are very susceptible to motion. Aiming at overcoming this limitation, we propose a new, more robust approach to estimate the heart rate from a videotaped face. It features non-planar motion compensation, fusion of multiple ROI signals, and a RANSAC-like time-domain heart rate estimation algorithm. In experiments with a comprehensive pain recognition dataset we show that our approach outperforms previous methods in the presence of spontaneous head movement and facial expression.
Philipp Werner, Ayoub Al-Hamadi, Steffen Walter 0001, Sascha Gruss, Harald C. Traue
ICIP2
2014 A Defect Recognition System for Automated Inspection of Non-rigid Surfaces
abstract
The goal of this work is the automated recognition of 3D surface defects for quality inspection in industrial production. For complexly shaped work pieces that are non-rigid and have non-uniform tolerance ranges, it is hard to distinguish acceptable surface deviations from defects. We propose a 3-stage defect recognition system based on 3D measurement of the defective part. First, a variable B-Spline surface model is used to adapt to acceptable tolerance ranges. The remaining model deviations are then used for segmentation of possible defects. Finally, a SVM-based classifier separates true defects from pseudo defects. On a real world data set of a series of measurements for a car front hood, the effectiveness of the approach is proven.
Sebastian von Enzberg, Ayoub Al-Hamadi
ICPR2
2014 Automatic Pain Recognition from Video and Biomedical Signals
abstract
How much does it hurt? Accurate assessment of pain is very important for selecting the right treatment, however current methods are not sufficiently valid and reliable in many cases. Automatic pain monitoring may help by providing an objective and continuous assessment. In this paper we propose an automatic pain recognition system combining information from video and biomedical signals, namely facial expression, head movement, galvanic skin response, electromyography and electrocardiogram. Using the BioVid Heat Pain Database, the system is evaluated in the task of pain detection showing significant improvement over the current state of the art. Further, we discuss the relevance of the modalities and compare person-specific and generic classification models.
Philipp Werner, Ayoub Al-Hamadi, Robert Niese, Steffen Walter 0001, Sascha Gruss, Harald C. Traue
ICPR2
2014 Detection and tracking approach using an automotive radar to increase active pedestrian safety
abstract
Vulnerable road users, e.g. pedestrians, have a high impact on fatal accident numbers. To reduce these statistics, car manufactures are intensively developing suitable safety systems. Hereby, fast and reliable environment recognition is a major challenge. In this paper we describe a tracking approach that is only based on a 24 GHz radar sensor. While common radar signal processing loses much information, we make use of a track-before-detect filter to incorporate raw measurements. It is explained how the Range-Doppler spectrum of pedestrian can help to initiate and stabilize tracking even in occultation scenarios compared to sensors in series.
Michael Heuer, Ayoub Al-Hamadi, A. Rain, Marc-Michael Meinecke
Intelligent Vehicles Symposium2
2014 Comparative Learning Applied to Intensity Rating of Facial Expressions of pain
abstract
Together with classification of facial expressions, the rating of their intensities is of major interest. Classical supervised learning techniques require labeling of the intensities, which is labor intensive and requires expert knowledge, but nevertheless is not guaranteed to be objective. We propose a new approach to learn an intensity rating function which does not require expert knowledge, because it simplifies the labeling task by avoiding the difficulty of selecting an absolute intensity value and to keep the labeling consistent for the whole dataset. It is based on a novel kind of ground truth which we call Comparative Labeling. It specifies sample pairs for which the first element is desired to have a lower intensity than the second. We introduce a learning scheme to find an optimal intensity function in respect of the Comparative Labeling and propose performance measures to assess the quality of the learned function. The technique is applied to rate the intensity of facial expressions of posed pain. The evaluation results show that the learned function is well suited for determining dynamic intensity variation over time. We also assess the suitability of the rating as an inter-individual intensity measure by comparing it to the intensity ratings given by human observers.
Philipp Werner, Ayoub Al-Hamadi, Robert Niese
Int. J. Pattern Recognit. Artif. Intell.2
2013 A New Approach for Hand Augmentation Based on Patch Modelling
Omer Rashid Ahmed, Ayoub Al-Hamadi
ACIVS2
2013 Automatic User-Specific Avatar Parametrisation and Emotion Mapping
Stephanie Behrens, Ayoub Al-Hamadi, Robert Niese, Eicke Redweik
ACIVS2
2013 Tracking of a Handheld Ultrasonic Sensor for Corrosion Control on Pipe Segment Surfaces
Christian Bendicks, Erik Lilienblum, Christian Freye, Ayoub Al-Hamadi
ACIVS4
2013 Upper-Body Pose Estimation Using Geodesic Distances and Skin-Color
Sebastian Handrich, Ayoub Al-Hamadi
ACIVS2
2013 Towards Pain Monitoring: Facial Expression, Head Pose, a new Database, an Automatic System and Remaining
abstract
Pain is what the patient says it is. But what about these who cannot utter? Automatic pain monitoring opens up prospects for better treatment, but accurate assessment of pain is challenging due to the subjective nature of pain. To facilitate advances, we contribute a new dataset, the BioVid Heat Pain Database which contains videos and physiological data of 90 persons subjected to well-defined pain stimuli of 4 intensities. We propose a fully automatic recognition system utilizing facial expression, head pose information and their dynamics. The approach is evaluated with the task of pain detection on the new dataset, also outlining open challenges for pain monitoring in general. Additionally, we analyze the relevance of head pose information for pain recognition and compare person-specific and general classification models.
Philipp Werner, Ayoub Al-Hamadi, Robert Niese, Steffen Walter 0001, Sascha Gruss, Harald C. Traue
BMVC2
2013 On optimality of teaching quality for a mathematical topic using Neural Networks (with a case study)
abstract
This paper addresses an interdisciplinary approach integrating evaluation of an educational issue with Artificial Neural Network (ANN) modeling. Specifically, it is concerned with ANN modeling of two Computer Assisted Learning (CAL) packages/modules using various learning rate values. Both packages are considered for teaching a mathematical topic: “How to solve long division problem?”. They have been submitted at the fifth grade classroom level in elementary schools (as a case study with or without associated teacher's voice). Furthermore, after the application of the suggested CAL packages, practical findings have been compared with classical learning obtained results. Interestingly, all findings are shown to be in agreement with simulation results after running the introduced realistic ANN model. Finally, this work investigated well how measured mathematical teaching quality could be fairly improved via assessment of two learning parameters' performance (achievement level & response time).
Saeed A. Al-Ghamdi, Hassan M. H. Mustafa, Ayoub Al-Hamadi, Mohammed H. Kortam, Abdel Aziz M. Al-Bassiouni
EDUCON3
2013 A Locale Group Based Line Segmentation Approach for Non Uniform Skewed and Curved Arabic Handwritings
abstract
In this paper we present a novel local group based method for extracting skewed and curved handwritten text lines in Arabic document images. We first detect all connected components and use a Support Vector Machine (SVM) to classify them either as Piece of Arabic Word (PAW) or diacritic. We then, use novel distance measures like sigmoid function based shapes to calculate the nearest neighbors for all PAWs. A subsequently graph based grouping algorithms, which follows the text lines from right to left, generates multiple candidate lines. After assessing the quality of all line candidates the final line representation is chosen. In a final step all PAWs which are not already part of a final line are inserted into the one that is closest. Experimental results show a successfully line segmentation for documents of different writers and styles.
Laslo Dinges, Ayoub Al-Hamadi, Moftah Elzobi
ICDAR2
2013 An Approach for Arabic Handwriting Synthesis Based on Active Shape Models
abstract
Comprehensive handwriting databases are crucial to train and test script recognition systems. However their generation is expensive in sense of manpower and time. As a result there is a lack of such databases which impedes research and development. This is especially true in case of holistic word recognition, since various samples must be available for each entry of the underlying vocabulary. To bypass this problem for Arabic, we present an efficient system that automatically generates images of synthetic handwritten words or text lines from unicode. A total of 28046 online samples of multiple writers are created to compute Active Shape Models (ASM) for over hundred letter classes. ASMs are used to generate unique letter representations for each synthesis. Subsequently these representations are modified by affine transformations, smoothed by B-Spline interpolation and composed to text. Finally the text is rendered and saved. In this way our system produces off-line pseudo handwritten samples with variations in shape and texture. We compare samples of the IFN/ENIT database with corresponding syntheses to show that these can be used to surrogate real samples.
Laslo Dinges, Ayoub Al-Hamadi, Moftah Elzobi
ICDAR2
2013 A Hidden Markov Model-Based Approach with an Adaptive Threshold Model for Off-Line Arabic Handwriting Recognition
abstract
In contrast to the mainstream HMM-based approaches dedicated for the recognition of offline handwritten Arabic, this paper proposes an HMM-based approach that built upon an explicit segmentation module. And shape representative based rather than sliding window based features, are extracted and used to build a reference as well as a confirmation model for each letter in each handwritten form. Additionally, we constructed an HMM-based threshold model by ergodically connecting all letter models, in order to detect false segmentation as well as nonletter segments. IESK-arDB and IFN/ENIT databases are used for testing and evaluation of the proposed approach respectively, and satisfactory results are achieved.
Moftah Elzobi, Ayoub Al-Hamadi, Laslo Dinges, Mahmoud Elmezain, Anwar Saeed
ICDAR2
2013 Automatic Realtime User Performance-Driven Avatar Animation
abstract
In this paper an approach for automatic user-specific 3D model generation and expression classification is proposed. User performance-driven avatar animation is recently in the focus of research due to the increasing amount of low-cost acquisition devices with integrated depth map computation. Thereby challenging is the user-specific emotion classification without a complex manual initialisation. Correct classification and emotion intensity identification can only be done with known expression specific facial feature displacement which differs from user to user. The use of facial feature tracking on predefined 3D model expression animations is presented here as solution statement for automatic emotion classification and intensity calculation. Consequently with this approach partial occlusions of a presented face do not hamper expression identification due to the symmetrical structure of human faces. Thus, a marker less, automatic and easy to use performance-driven avatar animation approach is presented.
Stephanie Behrens, Ayoub Al-Hamadi, Eicke Redweik, Robert Niese
SMC2
2013 A Robust Method for Human Pose Estimation Based on Geodesic Distance Features
abstract
In this work, we propose a real-time capable and robust method for human pose estimation based on geodesic distance features from depth images. Although a lot of work has been done in the field of the human pose estimation, it remains a challenging task - especially because of the high variability of human poses and self occlusions. The pose estimation focuses on the upper body, as it is the relevant part for a subsequent gesture and posture recognition and therefore the basis for a real human-machine-interaction. A graph-based representation of the 3D point cloud data is determined which allows for the measurement of pose-independent geodesic distances on the surface of the body. Based on these distances we determine feature points that are used for the adaptation of a kinematic skeleton model of the human upper body. The method does not need any pre-trained pose classifiers and can therefore track arbitrary poses as long as the user is not turned away from the camera.
Sebastian Handrich, Ayoub Al-Hamadi
SMC2
2013 Accurate, Fast and Robust Realtime Face Pose Estimation Using Kinect Camera
abstract
Since its release in late 2010 the Microsoft Kinect depth sensor has boosted real time gesture recognition and new man-machine interaction endeavors in the computer vision community. Based on depth image data, in this paper we propose an accurate, fast and robust face pose estimation approach, which for example can be of interest for user behavior analysis, or be of use as a means of man machine interaction modality. In our method we apply the depth sensor to create a user specific model which is fitted with an Iterative Closest Point algorithm. This model consists of point vertices and surface normals. In the fitting procedure we employ the normal vectors for the minimization of distances between the model and the measured point cloud. As the experimental results show, our method is precise, fast and robust in case of strong head rotation, even during facial expression and partial face occlusion.
Robert Niese, Philipp Werner, Ayoub Al-Hamadi
SMC3
2013 Gestic-Based Human Machine Interface for Robot Control
abstract
Mobile robots can assist humans in disaster management or environmental perception by building a multi robot team, working as a distributed sensor actor system. In order to coordinate the operations of a multi robot team the human machine interface is required to decode orders which will potentially be performed by a different robot, e.g. pointing to an area to be scanned. The human operator sets his statement in such cases of via hand action as interaction modality, registered by camera, which is the base for feature extraction. In this paper we address a gesture and hand posture based HMI-system. For segmentation of hand regions we combined color and depths information. As feature vectors a varying combination of: Fourier descriptors, cosine descriptors, Hu-moments and geometric features are extracted from the image and depth data. For classification of hand postures the feature vector is processed by an artificial neural network. A maximum overall classification rate of 93% is achieved for single image processing. Stabilizing the hand shape classification for online-sequences using a time histogram enables a robust robot control. The HMI serves hereby as communication basis for a multi-robot based enviroment perception and disaster management.
Michael Tornow, Ayoub Al-Hamadi, Vinzenz Borrmann
SMC2
2013 IESK-ArDB: a database for handwritten Arabic and an optimized topological segmentation approach
Moftah Elzobi, Ayoub Al-Hamadi, Zaher Al Aghbari, Laslo Dinges
Int. J. Document Anal. Recognit.2
2012 Multi hypotheses based object tracking in HCI environments
abstract
Gesture recogntion plays an important role in Human Computer Interaction (HCI). However, in most HCI systems the user is limited to use only one hand or two hands under optimal conditions. Challenges are for instance non-homogeneous backgrounds, hand-hand or hand-face overlapping or brightness modifications which will be met in real HCI scenarios. In this work, we propose a novel method that solves the ambiguities due to the hand overlap robustly based on multi-hypotheses object association. The results of tracking build the basis for further feature extraction and gesture recognition.
Sebastian Handrich, Ayoub Al-Hamadi
ICIP2
2012 An SVM approach for activity recognition based on chord-length-function shape features
abstract
Despite their high stability and compactness, chord-length features have received little attention in activity recognition literature. In this paper, we present an SVM approach for activity recognition, based on chord-length shape features. The main contribution of the paper is two-fold. We first show how a compact computationally-efficient shape descriptor is constructed using 1-D chord-length functions. Secondly, we unfold how to use fuzzy membership functions to partition action snippets into a number of temporal states. When tested on KTH benchmark dataset, the approach achieves promising results that compare very favorably with those reported in the literature, while maintaining real-time performance.
Samy Sadek, Ayoub Al-Hamadi, Bernd Michaelis, Usama Sayed
ICIP2
2012 Pain recognition and intensity rating based on Comparative Learning
abstract
Automatic pain recognition can improve medical treatment, especially when the patient is not able to utter on his pain experience. Facial expressions with their intensities and dynamics contain valuable information for recognising pain. We propose a concept for distinguishing facial expressions of pain from others and assessing the pain expression intensity. It is based on a Support Vector Machine (SVM) classifier and a function model for intensity rating. The intensity model is trained using Comparative Learning, a new technique that simplifies labelling of the data. Using a database of 3D posed pain sequences we show the suitability of the concept to recognise pain expressions, distinguish different intensities and spot even slight intensity changes in its temporal context.
Philipp Werner, Ayoub Al-Hamadi, Robert Niese
ICIP2
2012 Improving of Gesture Recognition Using Multi-hypotheses Object Association
Sebastian Handrich, Ayoub Al-Hamadi, Omer Rashid Ahmed
ICISP2
2012 Speaker Tracking Using Multi-modal Fusion Framework
Anwar Saeed, Ayoub Al-Hamadi, Michael Heuer
ICISP2
2012 Flow Modeling and skin-based Gaussian pruning to recognize gestural actions using HMM
Omer Rashid Ahmed, Ayoub Al-Hamadi
ICPR2
2012 Human action recognition via affine moment invariants
Samy Sadek, Ayoub Al-Hamadi, Bernd Michaelis, Usama Sayed
ICPR2
2012 Neutral-independent geometric features for facial expression recognition
abstract
Improving Human-Computer Interaction (HCI) necessitates building an efficient human emotion recognition approach that involves various modalities such as facial expressions, hand gestures, acoustic data, and biophysiological data. In this paper, we address the perception of the universal human emotions (happy, surprise, anger, disgust, fear, and sadness) from facial expressions. In our companion-based assistant system, facial expression is considered as complementary aspect to the hand gestures. Unlike many other approaches, we do not rely on prior knowledge of the neutral state to infer the emotion because annotating the neutral state usually involves human intervention. We use features extracted from just eight fiducial facial points. Our results are in a good agreement with those of a state-of-the-art approach that exploits features derived from 68 facial points and requires prior knowledge of the neutral state. Then, we evaluate our approach on two databases. Finally, we investigate the influence of the facial points detection error on our emotion recognition approach.
Anwar Saeed, Ayoub Al-Hamadi, Robert Niese
ISDA2
2012 Image-based gesture recognition for user interaction with mobile companion-based assistance systems
abstract
In this paper, we present image-based methods for robust recognition of static and dynamic hand gestures in real-time. These methods are used for an intuitive interaction with an assistance-system in which the skin-tones are used to segment the hands. The segmentation builds the basis of feature extraction for the static and dynamic gestures. In the static gestures, the activation of particular region leads us to associated actions whereas HMM classifier is used to extract the dynamic gestures dependent upon the flow. The assistance-system supports the workers in manual working tasks in the context of assembling complex products. This paper is focused on the interaction of the user with this system and describes the work in progress with the initial results from an application scenario.
Frerk Saxen, Omer Rashid Ahmed, Ayoub Al-Hamadi, Simon Adler, Alexa Kernchen, Rüdiger Mecke
ISDA3
2012 Recognizing gestural actions
abstract
In visual interaction environments, hands and arm provide a natural communication medium to interact with computers for recognizing meaningful actions. So, in this context, a novel approach is proposed to recognize the gestural actions by visual modality which comprises of four main modules. First, the dynamic contents in the scene are captured by optical flow and marginalized using Gaussian Mixture Model (GMM). Second, the resulting mixture of Gaussians are pruned by applying skin-based criterion to obtain the Gaussians containing both the skin and flow information. Third, the intra-level merging step is performed to obtain a concrete Gaussian representation spatially and then Kullback-Leibler (KL) divergence is used to obtain the inter-level linking among these Gaussians temporally. Fourth, the temporal features are computed from these linked Gaussians which are classified with Hidden Markov Model (HMM) to recognize the gestural actions. The experimental results show that our proposed approach is capable to recognize gestural actions in real situations which proves its applicability and usability in the domain of Human Computer Interactions (HCI).
Omer Rashid Ahmed, Ayoub Al-Hamadi
SMC2
2012 LDCRFs-based hand gesture recognition
abstract
This paper proposes a system to recognize isolated American Sign Language and numbers in real-time from Bumblebee stereo camera using Latent-Dynamic Conditional Random Fields (LDCRFs). Our system is based on three main stages: preprocessing, feature extraction and classification. In preprocessing stage, color and 3D depth map are used to detect and track the hand. The second stage, combining features of location, orientation and velocity with respected to Polar systems are used. The depth information is to identify the region of interest and consequently reduces the cost of searching and increases the processing speed. In the final stage, the hand gesture path is recognized using LDCRFs, which are more restricted to the number of hidden states owned by each class label to make training and inferencing processes tractable. Experimental results demonstrate that, our system can successfully recognize gestures with 96.14% recognition rate. Such results have the potential to compare very favorably to those of other investigators published in the literature.
Mahmoud Elmezain, Ayoub Al-Hamadi
SMC2
2012 Facial feature point detection using simplified gabor wavelets and confidence-based grouping
abstract
One of the first steps in most facial expression and facial analysis systems is the localization of prominent facial feature points. In this paper we present a novel approach for facial feature point detection using Simplified Gabor Wavelets (SGW). The classifier is built in cascades, where each stage of the cascade is a Gentle-AdaBoost trained classifier. In addition, we suggest a confidence based weighted grouping of multi-detected feature points to enhance accuracy. We have trained and tested our algorithm with a shuffled mix of four available labeled databases with more than 700 individuals. Our experimental results achieve approximately 82% detection rate in average, which is a considerable result, since the databases contain not only frontal faces.
Axel Panning, Ayoub Al-Hamadi, Bernd Michaelis
SMC2
2012 Stereo-Camera-Based Urban Environment Perception Using Occupancy Grid and Object Tracking
abstract
This paper deals with environment perception for automobile applications. Environment perception comprises measuring the surrounding field with onboard sensors such as cameras, radar, lidars, etc., and signal processing to extract relevant information for the planned safety or assistance function. Relevant information is primarily supplied using two well-known methods, namely, object based and grid based. In the introduction, we discuss the advantages and disadvantages of the two methods and subsequently present an approach that combines the two methods to achieve better results. The first part outlines how measurements from stereo sensors can be mapped onto an occupancy grid using an appropriate inverse sensor model. We employ the Dempster-Shafer theory to describe the occupancy grid, which has certain advantages over Bayes' theorem. Furthermore, we generate clusters of grid cells that potentially belong to separate obstacles in the field. These clusters serve as input for an object-tracking framework implemented with an interacting multiple-model estimator. Thereby, moving objects in the field can be identified, and this, in turn, helps update the occupancy grid more effectively. The first experimental results are illustrated, and the next possible research intentions are also discussed.
Thien-Nghia Nguyen, Bernd Michaelis, Ayoub Al-Hamadi, Michael Tornow, Marc-Michael Meinecke
IEEE Trans. Intell. Transp. Syst.3
2011 Efficient KNN search by linear projection of image clusters
abstract
K-nearest neighbors (KNN) search in a high-dimensional vector space is an important paradigm for a variety of applications. Despite the continuous efforts in the past years, algorithms to find the exact KNN answer set at high dimensions are outperformed by a linear scan method. In this paper, we propose a technique to find the exact KNN image objects to a given query object. First, the proposed technique clusters the images using a self-organizing map algorithm and then it projects the found clusters into points in a linear space based on the distances between each cluster and a selected reference point. These projected points are then organized in a simple, compact, and yet fast index structure called array-index. Unlike most indexes that support KNN search, the array-index requires a storage space that is linear in the number of projected points. The experiments show that the proposed technique is more efficient and robust to dimensionality as compared to other well-known techniques because of its simplicity and compactness. © 2011 Wiley Periodicals, Inc.
Zaher Al Aghbari, Ayoub Al-Hamadi
Int. J. Intell. Syst.2
2010 Toward Robust Action Retrieval in Video
abstract
Retrieving human actions from video databases is a paramount but challenging task in computer vision. In this work, we develop such a framework for robustly recognizing human actions in video sequences. The contribution of the paper is twofold. First a reliable neural model, the Multi-level Sigmoidal Neural Network (MSNN) as a classifier for the task of action recognition is presented. Second we unfold how the temporal shape variations can be accurately captured based on both temporal self-similarities and fuzzy log-polar histograms. When the method is evaluated on the popular KTH dataset, an average recognition rate of 94.3% is obtained. Such results have the potential to compare very favorably to those of other investigators published in the literature. Further the approach is amenable for real-time applications due to its low computational requirements.
Samy Bakheet, Ayoub Al-Hamadi, Bernd Michaelis, Usama Sayed
BMVC2
2010 Variable block-size image authentication with localization and self-recovery
abstract
In this paper, a self-recovery image authentication technique is proposed using randomly-sized blocks. To dived an image into randomly-sized blocks, it undergoes recursive arbitrarily-asymmetric quad tree partitioning. Multiple description coding (MDC) that enhances reliability of altered block recovery, is utilized to generate two block descriptions. We propose the use of a chain with multiple links for embedding several signature copies and two descriptions per block in arbitrarily distant blocks. The experimental results deposit that the proposed technique successfully both localizes and compensates the alterations. Furthermore, it is robust against the vector quantization (VQ) attack.
Ammar M. Hassan, Ayoub Al-Hamadi, Yassin M. Y. Hasan, Mohamed A. A. Wahab, Axel Panning, Bernd Michaelis
ICIP2
2010 Towards completely rotated Simplified Gabor Wavelets for fast facial feature point detection
abstract
Facial Feature point detection, and object detection in general, is not only a problem of accuracy but also a problem of computation time. Integral images provide fast computation time. Several classes of features working on these integral images have been introduced. So far in most cases only integral images of 0° and 45° have been used. We will present our first results of exploring the performance of additional 26.5° and 63.5° rotated integral images, which have been introduced several years ago but widely been ignored so far. As features we will use the so called Simplified Gabor Wavelets. The extension of the available feature set for pattern recognition tasks should reduce the number of used features and can save computation time in the later search process.
Axel Panning, Ayoub Al-Hamadi, Bernd Michaelis
ICIP2
2010 A Robust Method for Hand Gesture Segmentation and Recognition Using Forward Spotting Scheme in Conditional Random Fields
abstract
This paper proposes a forward spotting method that handles hand gesture segmentation and recognition simultaneously without time delay. To spot meaningful gestures of numbers (0-9) accurately, a stochastic method for designing a non-gesture model using Conditional Random Fields (CRFs) is proposed without training data. The non-gesture model provides a confidence measures that are used as an adaptive threshold to find the start and the end point of meaningful gestures. Experimental results show that the proposed method can successfully recognize isolated gestures with 96.51% and meaningful gestures with 90.49% reliability.
Mahmoud Elmezain, Ayoub Al-Hamadi, Bernd Michaelis
ICPR2
2010 Secure Self-Recovery Image Authentication Using Randomly-Sized Blocks
abstract
In this paper, a secure variable-size block-based image authentication technique is proposed that can not only localize the alteration detection but also recover the missing data. An image undergoes recursive arbitrarily-asymmetric binary tree partitioning to obtain randomly-sized blocks spanning the entire image. To enhance reliability of altered block recovery, multiple description coding (MDC) is utilized to generate two block descriptions. Block signature copies and the two block descriptions are embedded into two relatively-distant blocks making a doubly linked chain. The experimental results deposit that the proposed technique successfully both localizes and compensates the alterations. Furthermore, it is robust against the vector quantization (VQ) attack.
Ammar M. Hassan, Ayoub Al-Hamadi, Bernd Michaelis, Yassin M. Y. Hasan, Mohamed A. A. Wahab
ICPR2
2010 Real-Time Automatic Traffic Accident Recognition Using HFG
abstract
Recently, the problem of automatic traffic accident recognition has appealed to the machine vision community due to its implications on the development of autonomous Intelligent Transportation Systems (ITS). In this paper, a new framework for real-time automated traffic accidents recognition using Histogram of Flow Gradient (HFG) is proposed. This framework performs two major steps. First, HFG-based features are extracted from video shots. Second, logistic regression is employed to develop a model for the probability of occurrence of an accident by fitting data to a logistic curve. In case of occurrence of an accident, the trajectory of vehicle by which the accident was occasioned is determined. Preliminary results on real video sequences confirm the effectiveness and the applicability of the proposed approach, and it can offer delay guarantees for real-time surveillance and monitoring scenarios.
Samy Sadek, Ayoub Al-Hamadi, Bernd Michaelis, Usama Sayed
ICPR2
2010 Audio-Visual Data Fusion Using a Particle Filter in the Application of Face Recognition
abstract
This paper describes a methodology by which audio and visual data about a scene can be fused in a meaningful manner in order to locate a speaker in a scene. This fusion is implemented within a Particle Filter such that a single speaker can be identified in the presence of multiple visual observations. The advantages of this fusion are that weak sensory data from either modality can be reinforced and the presence of noise can be reduced.
Michael Alan Steer, Ayoub Al-Hamadi, Bernd Michaelis
ICPR2
2009 OIF - An Online Inferential Framework for Multi-object Tracking with Kalman Filter
Saira Saleem Pathan, Ayoub Al-Hamadi, Bernd Michaelis
CAIP2
2009 Hand trajectory-based gesture spotting and recognition using HMM
abstract
In this paper, we propose an automatic system that executes hand gesture spotting and recognition simultaneously without any time delay based on Hidden Markov Models (HMM). Our system is based on three main stages; preprocessing, feature extraction and classification. In preprocessing stage, color and 3D depth map are used to detect hands. The hand trajectory will take place in further steps using Mean-shift algorithm and Kalman filter. The second stage, Orientation dynamic features are obtained from spatio-temporal trajectories and then are quantized to generate its codewords. In the final stage, the gestures are segmented by finding the start and the end points of meaningful gestures that are embedded in the input stream and then are recognized by Viterbi algorithm. Experimental results demonstrate that, our system can successfully recognize spotted hand gestures with a 95.87% recognition rate for Arabic numbers from 0 to 9.
Mahmoud Elmezain, Ayoub Al-Hamadi, Bernd Michaelis
ICIP2
2009 A novel public key self-embedding fragile watermarking technique for image authentication
abstract
Digital images are vulnerable to various malicious manipulations. So, effective techniques are demanded for assuring the integrity of image contents. In this paper, a public key self-embedding image authentication technique is proposed which can detect and localize alterations of the image contents and can restore the distorted regions. A doubly linked chain is used to embed block signature copies and two block codes into relatively distant blocks. Each block code can not only approximately rebuild a distorted block but also can, combined with another block code, rebuild the block with high quality. Moreover, the proposed technique uses a public key in the verification process that enables verification of the image authenticity without knowing the private key used in the embedding process. The experimental results show that the proposed technique is sensitive to pixel changes and provides secure embedding of the blockwise image signatures. Furthermore, it thwarts many attacks, specially the collage (also called vector quantization, blind copy, or pattern matching) attacks and can successively recover the image attacked areas.
Ammar M. Hassan, Yassin M. Y. Hasan, Ayoub Al-Hamadi, Mohamed A. A. Wahab, Bernd Michaelis
ICIP3
2009 Cubic-splines neural network- based system for Image Retrieval
abstract
Research in content-based image retrieval (CBIR) shows that high-level semantic concepts in image cannot be constantly depicted using low-level image features. So the process of designing a CBIR system should take into account diminishing the existing gap between low-level visual image features and the high-level semantic concepts. In this paper, we propose a new architecture for a CBIR system named SNNIR (splines neural network-based image retrieval). SNNIR system makes use of a rapid and precise neural model. This model employs a cubic-splines activation function. By using the spline neural model, the gap between the low-level visual features and the high-level concepts is minimized. Experimental results show that the proposed system achieves high accuracy and effectiveness in terms of precision and recall compared with other CBIR systems.
Samy Sadek, Ayoub Al-Hamadi, Bernd Michaelis, Usama Sayed
ICIP2
2008 Robust facial expression recognition based on 3-d supported feature extraction and SVM classification
abstract
Facial expression recognition is an imperative task in human computer interaction systems. In this work we propose a new system for automatic expression recognition in video sequences. Our system uses color information to extract the facial features. Additionally, it includes a camera model and a registration step, in which we automatically build a person specific face model from stereo. Photogrammetric techniques are applied to determine real world geometric measures and to build the feature vector. Feature normalization is carried out and support vector machine is trained over the normalized feature vector. Using SVM classification we reach a minimal mixing between different expression classes. Our framework achieves robust and superior expression classification results across a variety of head poses with resulting perspective foreshortening and changing face size. Moreover, the method has shown robustness across a variety of skin colors while reaching high performance.
Robert Niese, Ayoub Al-Hamadi, Faisal Aziz, Bernd Michaelis
FG2
2008 A Hidden Markov Model-based continuous gesture recognition system for hand motion trajectory
abstract
In this paper, we propose an automatic system that recognizes both isolated and continuous gestures for Arabic numbers (0-9) in real-time based on hidden Markov model (HMM). To handle isolated gestures, HMM using ergodic, left-right (LR) and left-right banded (LRB) topologies with different number of states ranging from 3 to 10 is applied. Orientation dynamic features are obtained from spatio-temporal trajectories and then quantized to generate its codewords. The continuous gestures are recognized by our novel idea of zero-codeword detection with static velocity motion. Therefore, the LRB topology in conjunction with forward algorithm presents the best performance and achieves average rate recognition 98.94% and 95.7% for isolated and continuous gestures, respectively.
Mahmoud Elmezain, Ayoub Al-Hamadi, Jörg Appenrodt, Bernd Michaelis
ICPR2
2007 Real-Time Capable Method for Facial Expression Recognition in Color and Stereo Vision
Robert Niese, Ayoub Al-Hamadi, Axel Panning, Bernd Michaelis
ICCSA (1)2
2005 Spike-timing-dependent plasticity in 'small world' networks
Karsten Kube, Andreas Herzog, Bernd Michaelis, Ayoub Al-Hamadi, Ana D. de Lima, Thomas B. Voigt
ESANN4
2003 Improvement of the Fail-Safe Characteristics in Motion Analysis Using Adaptive Technique
Ayoub Al-Hamadi, Rüdiger Mecke, Bernd Michaelis
CIARP1
2003 Another Paradigm for the Solution of the Correspondence Problem in Motion Analysis
Ayoub Al-Hamadi, Robert Niese, Bernd Michaelis
CIARP1
1998 A robust method for block-based motion estimation in RGB-image sequences
abstract
In this paper the block-based motion analysis in monocular RGB-image sequences is considered. The motion analysis is realized using a combined matching criterion, that represents a robust similarity measure for image blocks in successive images. This criterion is calculated as a weighted combination of the 3 color-specific matching criteria. The weight of each of these channel-specific criteria is determined based on their reliability. The combined criterion is used to measure translational and rotational motion between 2 successive images. Assuming these measuring values, a model based recursive estimation technique (Kalman filter) is applied to estimate the position and the velocity of the respective region. In order to cope with problematic image situations (e.g. partial occlusion) an extension of the conventional filter by a recognition system is proposed. This system recognizes typical distorted matching criteria and controls the adaptation of the filter.
Rüdiger Mecke, Ayoub Al-Hamadi, Bernd Michaelis
ICPR2