Mohamed Abouelenien

dblp:123/9584 · DBLP profile ↗
← Back
24ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0001-5351-5778ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 7 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards Personalized Multimodal Efficient Detection of Human Circadian States
Kapotaksha Das, Mihai Burzo, Mohamed Abouelenien
ICPR (7)3
2026 More Than Meets the Ear: Multimodal Driver Alertness Detection Leveraging LLMs and Synthetic Speech
Salem Sharak, Kapotaksha Das, Mihai Burzo, Mohamed Abouelenien
ICPR (13)4
2025 PyraSegNet: A Novel Framework for Thermal Facial Image Segmentation
abstract
Thermal facial region segmentation is critical for applications such as physiological monitoring, stress detection, and human-computer interaction. Despite its importance, progress in this domain has been hindered by challenges such as occlusions, the lack of thermal annotated datasets and the absence of tailored, high-performing models on thermal images. In this paper, we address these challenges through four key contributions. First, we introduce the TFR dataset, a novel thermal face dataset that includes 5050 annotated thermal images extracted from 166 video sequences using 163 different subjects, which is, to our knowledge, the largest thermal dataset with comprehensive annotated regions in terms of the number of subjects, captured under diverse conditions, including variations in lighting, occlusions, and demographics, ensuring robustness and adaptability. Second, we propose PyraSegNet, an advanced deep learning architecture combining a customized encoder-decoder structure with a Feature Pyramid Network (FPN) to enhance multi-scale feature learning. Third, we compare PyraSegNet to other state-of-the-art methods. Finally, we perform an in-depth comparative analysis of loss functions, including Dice Loss, Tversky Loss, and Jaccard Loss, providing actionable guidance on optimizing thermal facial segmentation performance. Our findings underscore the significance of the TFR dataset, the effectiveness of PyraSegNet, and provide insights into optimal loss function selection for thermal facial region segmentation. This research paves the way for enhanced accuracy and reliability in applications leveraging thermal facial analysis.
Kais Riani, Mohamed Abouelenien, Oumaima Jouiri
FG2
2024 PyraMoT: A Novel Framework for Enhanced Facial Thermal Landmarks Detection
abstract
Facial analysis and recognition are vital components in many real-world applications, such as driver safety monitoring, security systems, and healthcare. Despite the need for thermal images in many of these applications, there has been limited research conducted on facial landmark recognition for thermal images. This scarcity can be partly attributed to a lack of thermal datasets for comprehensive analysis. In this paper, we make three major contributions to the field of thermal facial landmark detection. First, we present D5050, a novel thermal face dataset that includes 166 video sequences and 5050 annotated thermal images from 163 different subjects, which is, to our knowledge, the largest dataset with comprehensive landmark annotation in terms of the number of subjects. Second, we propose PyraMoT, an innovative thermal facial landmark detection framework that combines a customized encoder-decoder structure and a Feature Pyramid Network (FPN) with the efficiency of MobileNetV2and the incorporation of the anisotropic loss function. Third, we conduct a thorough comparative analysis with existing approaches, including six methods, three datasets, and four color palettes for thermal facial landmark detection. This study employs measures such as Normalized Mean Error (NME) and Failure Rate (FR) to determine and comparatively evaluate the accuracy and reliability of the various detection methods. Overall, our proposed approach outperforms other methods on different datasets and contributes to the field of thermal image processing, proposing three main advancements that address current limitations in the field.
Kais Riani, Salem Sharak, Mohamed Abouelenien
FG3
2023 Towards Modeling Human Biorhythms Using Non-Invasive Approaches
abstract
Biorhythms in humans hold vital importance that is directly related to a person’s health and well-being. This paper introduces a novel dataset for modeling biorhythms while focusing primarily on using an ingestible pill to monitor core body temperature and four physiological signals. We use two different settings to label our data to develop different scenarios of energization and enervation. We propose an approach that could potentially act as a reliable alternative to the usage of core temperature pills and genome sequencing while avoiding their invasiveness and complexities. We also explore the effect of the data size on the overall performance. Our findings indicate that our proposed approach is reliable and feasible, and presents future promise in using less invasive approaches to monitor and predict the circadian state of humans.
Kapotaksha Das, John Elson, Mohamed Abouelenien, Ali Hassani 0007, Mihai Burzo, Clay Maranville
BIBM3
2023 Non-Contact Based Modeling of Enervation
abstract
Significant research is currently carried out with a focus on autonomous vehicles; research is starting to focus on areas such as the modeling of occupant states and behavioral elements. This paper contributes to this line of research by developing a pipeline that extracts physiological signals from thermal imagery and modeling occupant enervation using a fully non-contact based approach. These signals are obtained via a multimodal dataset of 36 subjects across multiple channels, including the thermal and physiological modalities. Moreover, we provide a comparative analysis of non-contact and contact based channels to model the enervation state of individuals. Our analysis indicates that non-contact physiological signals extracted from thermal imagery can reach and exceed the performance of contact-based physiological signals. In addition, modeling of enervation is possible using said non-contact physiological signals and thermal features, with an accuracy of up to 70% in identifying energized and enervated occupant states. Our findings provide a novel approach for future research and opens the possibility for integration of unrestrictive sensors in future automobiles.
Kais Riani, Salem Sharak, Mohamed Abouelenien, Mihai Burzo, Rada Mihalcea, John Elson, Clay Maranville, Kwaku O. Prakah-Asante, Waqas Manzoor
FG3
2022 Contact Versus Noncontact Detection of Driver's Drowsiness
abstract
With an estimated number of injuries in the millions, accidents caused due to drowsy driving remain a significant source of financial costs and loss of life. Accurate detection of driver’s drowsiness could provide a clear avenue towards eliminating a great majority of the associated accidents and losses. Existing research on the subject could be defined as either contact-based or noncontact-based alertness detection. This paper utilizes a novel multimodal driver’s alertness dataset consisting of 45 subjects via seven recorded channels, including four contact-based and three noncontact-based channels, to investigate the performance of said modalities in detecting driver’s drowsiness as well as provide a novel comparison between the results of multiple contact and noncontact methods. Our results highlight the viability of noncontact methods to detect driver’s drowsiness as an implementable technology in automobiles.
Salem Sharak, Kapotaksha Das, Kais Riani, Mohamed Abouelenien, Mihai Burzo, Rada Mihalcea
ICPR4
2022 A Topological Approach for Facial Region Segmentation in Thermal Images
abstract
The use of thermal images as a non-contact modality has expanded drastically in recent years, including applications such as biometrics, detection of specific human behaviors, and extraction of physiological signals. For instance, using thermal images to monitor physiological signals was found to provide better information compared to RGB images, especially in situations where using physical contact sensors is not a possibility. In this paper, we present a novel topological approach that can segment the cheeks and forehead of a human face in a thermal image, with the key benefit that explicit prior information about the cheeks and forehead is unneeded. Our approach leverages topological properties present in a thermal image to provide a non-rectangular bounding curve for the cheeks and forehead, which we compare against a dataset of 1000 thermal images that were manually annotated for this task. By generating a Vietoris-Rips complex on a thermal image filtered by the Canny edge detector, our approach can segment the cheeks and forehead with recall scores as high as 90.4% and 78.4%, respectively.
Michael Lilley, Kapotaksha Das, Kais Riani, Mohamed Abouelenien
ISM4
2022 Multimodal Deception Detection Using Real-Life Trial Data
abstract
Hearings of witnesses and defendants play a crucial role when reaching court trial decisions. Given the high-stakes nature of trial outcomes, developing computational models that assist the decision-making process is an important research venue. In this article, we address the identification of deception in real-life trial data. We use a dataset consisting of videos collected from public court trials. We explore the use of verbal and non-verbal modalities to build a multimodal deception detection system that aims to discriminate between truthful and deceptive statements provided by defendants and witnesses. In particular, three complementary modalities (visual, acoustic and linguistic) are evaluated for the classification of deception at the subject level. The final classifier is obtained by combining the three modalities via score-level classification, achieving 83.05 percent accuracy in subject-level deceit detection. To place our results in perspective, we present a human deception detection study where we evaluate the human capability of detecting deception using different modalities and compare the results to the developed system. The results show that our system outperforms the average non-expert human capability of identifying deceit.
Mehmet Umut Sen, Verónica Pérez-Rosas, Berrin A. Yanikoglu, Mohamed Abouelenien, Mihai Burzo, Rada Mihalcea
IEEE Trans. Affect. Comput.4
2022 Detection and Recognition of Driver Distraction Using Multimodal Signals
abstract
Distracted driving is a leading cause of accidents worldwide. The tasks of distraction detection and recognition have been traditionally addressed as computer vision problems. However, distracted behaviors are not always expressed in a visually observable way. In this work, we introduce a novel multimodal dataset of distracted driver behaviors, consisting of data collected using twelve information channels coming from visual, acoustic, near-infrared, thermal, physiological and linguistic modalities. The data were collected from 45 subjects while being exposed to four different distractions (three cognitive and one physical). For the purposes of this paper, we performed experiments with visual, physiological, and thermal information to explore potential of multimodal modeling for distraction recognition. In addition, we analyze the value of different modalities by identifying specific visual, physiological, and thermal groups of features that contribute the most to distraction characterization. Our results highlight the advantage of multimodal representations and reveal valuable insights for the role played by the three modalities on identifying different types of driving distractions.
Kapotaksha Das, Michalis Papakostas, Kais Riani, Andrew Brian Gasiorowski, Mohamed Abouelenien, Mihai Burzo, Rada Mihalcea
ACM Trans. Interact. Intell. Syst.5
2021 Towards Classifying Human Circadian Rhythm Using Multiple Modalities
abstract
Autonomous vehicles represent one of the most active technologies currently being developed, with research areas addressing, among others, the modeling of the states and behavioral elements of the occupants. This paper contributes to this line of research by studying the circadian rhythm of individuals using a novel multimodal dataset of 36 subjects consisting of five information channels. These channels include visual, thermal, physiological, linguistic, and background data. Moreover, we propose a framework to explore whether the circadian rhythm can be modeled without continuous monitoring and investigate the hypothesis that multimodal features have a greater propensity for improved performance using data points specific to certain times during the day. Our analysis shows that multimodal fusion can lead to an accuracy of up to 77% on identifying energized and enervated states of the participants. Our findings highlight the validity of our hypothesis and present a novel approach for future research.
Kais Riani, Salem Sharak, Kapotaksha Das, Mohamed Abouelenien, Mihai Burzo, Rada Mihalcea, John Elson, Clay Maranville, Kwaku O. Prakah-Asante, Waqas Manzoor
ACII4
2021 Multimodal Detection of Drivers Drowsiness and Distraction
abstract
Considering the ever-growing presence of automobiles around the world, ensuring the safety of those on and near roadways is of great importance. From the causes of accidents, drowsiness and distractedness are among the most consequential. In this paper, we use a multimodal dataset consisting of 11 recorded channels over 45 subjects to model driver’s drowsiness and distraction. Our work puts forward the application of this dataset by using segmented windows as features, resulting in four main contributions. We explore the performance of each individual modality and specify which signals and features have a better capability of detecting drowsiness and different kinds of distractions. In addition, we analyze the effects of early fusion on the classification of the driver’s state using multiple physiological and thermal channels. Finally, we use cascaded late fusion and test three voting strategies to evaluate the performance of our proposed approach. Our results confirm the effectiveness of utilizing a multimodal approach in detecting both drowsiness and distraction as two separate factors influencing the driver and provide guidelines on which signals are appropriate for detecting different driver’s states.
Kapotaksha Das, Salem Sharak, Kais Riani, Mohamed Abouelenien, Mihai Burzo, Michalis Papakostas
ICMI4
2021 Understanding Driving Distractions: A Multimodal Analysis on Distraction Characterization
abstract
Distracted driving is a leading cause of accidents worldwide. The tasks of distraction detection and recognition have been traditionally addressed as computer vision problems. However, distracted behaviors are not always expressed in a visually observable way. In this work, we introduce a novel multimodal dataset of distracted driver behaviors, consisting of data collected using twelve information channels coming from visual, acoustic, near-infrared, thermal, physiological and linguistic modalities. The data were collected from 45 subjects while being exposed to four different distractions (three cognitive and one physical). For the purposes of this paper, we experiment with visual and physiological information and explore the potential of multimodal modeling for distraction recognition. In addition, we analyze the value of different modalities by identifying specific visual and physiological groups of features that contribute the most to distraction characterization. Our results highlight the advantage of multimodal representations and reveal valuable insights for the role played by the two modalities on identifying different types of driving distractions.
Michalis Papakostas, Kais Riani, Andrew Brian Gasiorowski, Mohamed Abouelenien, Rada Mihalcea, Mihai Burzo
IUI5
2021 MUSER: MUltimodal Stress detection using Emotion Recognition as an Auxiliary Task
abstract
Yiqun Yao, Michalis Papakostas, Mihai Burzo, Mohamed Abouelenien, Rada Mihalcea. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Yiqun Yao, Michalis Papakostas, Mihai Burzo, Mohamed Abouelenien, Rada Mihalcea
NAACL-HLT4
2020 MORSE: MultimOdal sentiment analysis for Real-life SEttings
abstract
Multimodal sentiment analysis aims to detect and classify sentiment expressed in multimodal data. Research to date has focused on datasets with a large number of training samples, manual transcriptions, and nearly-balanced sentiment labels. However, data collection in real settings often leads to small datasets with noisy transcriptions and imbalanced label distributions, which are therefore significantly more challenging than in controlled settings. In this work, we introduce MORSE, a domain-specific dataset for MultimOdal sentiment analysis in Real-life SEttings. The dataset consists of 2,787 video clips extracted from 49 interviews with panelists in a product usage study, with each clip annotated for positive, negative, or neutral sentiment. The characteristics of MORSE include noisy transcriptions from raw videos, naturally imbalanced label distribution, and scarcity of minority labels. To address the challenging real-life settings in MORSE, we propose a novel two-step fine-tuning method for multimodal sentiment classification using transfer learning and the Transformer model architecture; our method starts with a pre-trained language model and one step of fine-tuning on the language modality, followed by the second step of joint fine-tuning that incorporates the visual and audio modalities. Experimental results show that while MORSE is challenging for various baseline models such as SVM and Transformer, our two-step fine-tuning method is able to capture the dataset characteristics and effectively address the challenges. Our method outperforms related work that uses both single and multiple modalities in the same transfer learning settings.
Yiqun Yao, Verónica Pérez-Rosas, Mohamed Abouelenien, Mihai Burzo
ICMI3
2018 A regularized ensemble framework of deep learning for cancer detection from multi-class, imbalanced training data
Xiaohui Yuan 0001, Lijun Xie, Mohamed Abouelenien
Pattern Recognit.3
2017 Multimodal gender detection
abstract
Automatic gender classification is receiving increasing attention in the computer interaction community as the need for personalized, reliable, and ethical systems arises. To date, most gender classification systems have been evaluated on textual and audiovisual sources. This work explores the possibility of enhancing such systems with physiological cues obtained from thermography and physiological sensor readings. Using a multimodal dataset consisting of audiovisual, thermal, and physiological recordings of males and females, we extract features from five different modalities, namely acoustic, linguistic, visual, thermal, and physiological. We then conduct a set of experiments where we explore the gender prediction task using single and combined modalities. Experimental results suggest that physiological and thermal information can be used to recognize gender at reasonable accuracy levels, which are comparable to the accuracy of current gender prediction systems. Furthermore, we show that the use of non-contact physiological measurements, such as thermography readings, can enhance current systems that are based on audio or visual input. This can be particularly useful for scenarios where non-contact approaches are preferred, i.e., when data is captured under noisy audiovisual conditions or when video or speech data are not available due to ethical considerations.
Mohamed Abouelenien, Verónica Pérez-Rosas, Rada Mihalcea, Mihai Burzo
ICMI1
2017 Identity Deception Detection
abstract
This paper addresses the task of detecting identity deception in language. Using a novel identity deception dataset, consisting of real and portrayed identities from 600 individuals, we show that we can build accurate identity detectors targeting both age and gender, with accuracies of up to 88. We also perform an analysis of the linguistic patterns used in identity deception, which lead to interesting insights into identity portrayers.
Verónica Pérez-Rosas, Quincy Davenport, Anna Mengdan Dai, Mohamed Abouelenien, Rada Mihalcea
IJCNLP(1)4
2017 Detecting Deceptive Behavior via Integration of Discriminative Features From Multiple Modalities
abstract
Deception detection has received an increasing amount of attention in recent years, due to the significant growth of digital media, as well as increased ethical and security concerns. Earlier approaches to deception detection were mainly focused on law enforcement applications and relied on polygraph tests, which had proved to falsely accuse the innocent and free the guilty in multiple cases. In this paper, we explore a multimodal deception detection approach that relies on a novel data set of 149 multimodal recordings, and integrates multiple physiological, linguistic, and thermal features. We test the system on different domains, to measure its effectiveness and determine its limitations. We also perform feature analysis using a decision tree model, to gain insights into the features that are most effective in detecting deceit. Our experimental results indicate that our multimodal approach is a promising step toward creating a feasible, non-invasive, and fully automated deception detection system.
Mohamed Abouelenien, Verónica Pérez-Rosas, Rada Mihalcea, Mihai Burzo
IEEE Trans. Inf. Forensics Secur.1
2015 Verbal and Nonverbal Clues for Real-life Deception Detection
abstract
Deception detection has been receiving an increasing amount of attention from the computational linguistics, speech, and multimodal processing communities. One of the major challenges encountered in this task is the availability of data, and most of the research work to date has been conducted on acted or artificially collected data. The generated deception models are thus lacking real-world evidence. In this paper, we explore the use of multimodal real-life data for the task of deception detection. We develop a new deception dataset consisting of videos from reallife scenarios, and build deception tools relying on verbal and nonverbal features. We achieve classification accuracies in the range of 77-82% when using a model that extracts and fuses features from the linguistic and visual modalities. We show that these results outperform the human capability of identifying deceit.
Verónica Pérez-Rosas, Mohamed Abouelenien, Rada Mihalcea, C. J. Linton, Mihai Burzo
EMNLP2
2015 Deception Detection using Real-life Trial Data
abstract
Hearings of witnesses and defendants play a crucial role when reaching court trial decisions. Given the high-stake nature of trial outcomes, implementing accurate and effective computational methods to evaluate the honesty of court testimonies can offer valuable support during the decision making process. In this paper, we address the identification of deception in real-life trial data. We introduce a novel dataset consisting of videos collected from public court trials. We explore the use of verbal and non-verbal modalities to build a multimodal deception detection system that aims to discriminate between truthful and deceptive statements provided by defendants and witnesses. We achieve classification accuracies in the range of 60-75% when using a model that extracts and fuses features from the linguistic and gesture modalities. In addition, we present a human deception detection study where we evaluate the human capability of detecting deception in trial hearings. The results show that our system outperforms the human capability of identifying deceit.
Verónica Pérez-Rosas, Mohamed Abouelenien, Rada Mihalcea, Mihai Burzo
ICMI2
2014 Deception detection using a multimodal approach
abstract
In this paper we address the automatic identification of deceit by using a multimodal approach. We collect deceptive and truthful responses using a multimodal setting where we acquire data using a microphone, a thermal camera, as well as physiological sensors. Among all available modalities, we focus on three modalities namely, language use, physiological response, and thermal sensing. To our knowledge, this is the first work to integrate these specific modalities to detect deceit. Several experiments are carried out in which we first select representative features for each modality, and then we analyze joint models that integrate several modalities. The experimental results show that the combination of features from different modalities significantly improves the detection of deceptive behaviors as compared to the use of one modality at a time. Moreover, the use of non-contact modalities proved to be comparable with and sometimes better than existing contact-based methods. The proposed method increases the efficiency of detecting deceit by avoiding human involvement in an attempt to move towards a completely automated non-invasive deception detection process.
Mohamed Abouelenien, Verónica Pérez-Rosas, Rada Mihalcea, Mihai Burzo
ICMI1
2012 A Boosting Method for Learning from Uneven Data for Improved Face Recognition
abstract
In this paper, we propose a multi-class boosting method (multiBoost.imb) to address difficulties of learning from imbalanced data set as well as employment of stable base learners. A random resampling strategy is incorporated to diversify the training data set and to recover balance among all classes. Extending AdaBoost by adding an error adjustment parameter, early termination in the training phase is avoided in multi-class scenarios. Experiments were conducted using three public face databases and two synthetic data sets. It is demonstrated that stable learners can be used in our ensemble method. In the multi-class problems, the ensemble overcomes the early termination even when stable learner is employed. It was evident that our method improves learning performance in all cases, especially when imbalance ratio is high. Comparison to the SMOTEboost and RUSboost also reveals the advantage of our method in handling multi-class, imbalanced face recognition problems.
Xiaohui Yuan 0001, Mohamed Abouelenien
ICMLA (2)2
2012 SampleBoost: Improving boosting performance by destabilizing weak learners based on weighted error analysis
Mohamed Abouelenien, Xiaohui Yuan 0001
ICPR1