VLDB 2026 Research / reviewers in the wild / expert
Fady Shibata-Alnajjar
dblp:208/1541 · also Fady Alnajjar, Fady S. Alnajjar
· DBLP profile ↗
31ranked-venue papers
8as first author
18since 2021 · last 2027
0000-0001-6102-3765ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 7 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 9 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 5 since 2021Systems, architecture and hardware · 1Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | TCubes: Non-privacy invasive dataset for activities of daily living as thermal cubesabstractThis work introduces the first open multi–infrared (IR) thermal array sensor dataset for recognizing 19 activities of daily living (ADLs) performed by 74 subjects. Each activity sample forms a 32×32×32 thermal tensor cube-hence the name TCubes. The dataset surpasses existing resources in scale, featuring the highest number of activities, twice as many sensors, and three times as many participants as comparable datasets. Thermal patterns were captured using nine low-resolution (8×8) IR thermal sensors, providing a non-invasive and privacy-preserving means of activity recognition. A comprehensive benchmarking study evaluates both convolutional and transformer-based architectures—including C3D, R(2+1)D-18, MViTv2, and Swin-T—to assess their ability to learn spatiotemporal representations from coarse thermal imagery. Results are highly promising: R(2+1)D-18 achieves the most consistent performance with an F1-score of 0.900, while transformer models such as MViTv2 and Swin-T effectively capture subtle gestures and generalize well across 18 and 19 activity classes. The lightweight 3DCNN-Mixed model further demonstrates strong efficiency for resource-constrained applications, highlighting the trade-off between accuracy and computational cost. Analyses leveraging entropy, mutual confusion, and frequency-domain representations reveal how factors such as activity location, posture, temporal dynamics, and motion periodicity influence recognition accuracy. Overall, this dataset and benchmarking suite establish a robust foundation for future research in low-resolution, low-compute, non-invasive, and privacy-preserving human activity recognition, with broad implications for eldercare, healthcare monitoring, and smart environments. Luubaatar Badarch, Byambaa Dorj, Hansaem Park, Gantumur Tsogtgerel, Fady Shibata-Alnajjar, Munkhjargal Gochoo |
Expert Syst. Appl. | 5 |
| 2026 | Removal notice to "Deep learning with image-based autism spectrum disorder analysis: a systematic review" [Eng. Appl. Artif. Intell. 127 (2024) 107185]
Md. Zasim Uddin, Md. Arif Shahriar, Md. Nadim Mahamood, Fady Shibata-Alnajjar, Md. Ileas Pramanik, Md. Atiqur Rahman Ahad |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Comparing Emotion Detection Methods in Online Classrooms: YOLO Models, Multimodal LLM, and Human BaselineabstractThe COVID-19 pandemic has transformed learning environments, challenging educators to understand students' behaviors during the online mode of learning, in particular emotions associated with students' attention during virtual classrooms. As learning transitions between physical and virtual spaces, the ability to interpret student attention and engagement has become complex. In response to this challenge, our research investigates the use of GPT-4o, a multimodal large language model, for identifying student emotions by analyzing images in diverse learning settings. The study involved analyzing online classroom images featuring 149 faces, utilizing three distinct approaches: a computer vision model (YOLO), the multimodal LLM (GPT-4o), and a human-annotated baseline. The analysis systematically categorized facial expressions into eight emotional categories: Happy, Sad, Angry, Neutral, Contempt, Disgust, Fear, and Surprise. The findings indicate that multimodal LLMs can effectively detect student emotions, achieving an average accuracy of 93.8%, which aligns with the human baseline accuracy of 97.0%. In contrast, YOLO models maintained an average accuracy of 81.9%, performing well for basic emotions but struggling with subtle expressions. This research contributes to enhancing educational practices by providing valuable insights regarding the application of multimodal LLMs to assist educators in comprehending student emotions within both physical and digital classroom settings. Medha Mohan Ambali Parambil, Salah Bouktif, Munkhjargal Gochoo, Fady Shibata-Alnajjar |
EDUCON | 4 |
| 2025 | Driving with Voices: How Virtual Agent Affects Driver Stress and PerformanceabstractAuditory features like voice pitch and speech style significantly influence human-computer interaction, especially in driving where voice assistants provide real-time guidance. This study examined how two voice pitches (Alto and Bass) and two speech styles (Casual and Frozen) affected drivers’ cognitive and emotional responses. Eighty-four participants in a driving simulator received voice reminders from an assistant with varied auditory features while stress, heart rate variability, emotions, and task performance were recorded. Cognitive load was assessed using the NASA-TLX questionnaire. Results showed that the Frozen speech style combined with Alto pitch triggered higher stress and cognitive load, while speech style had a stronger effect than pitch on emotions. These insights aid in designing user-friendly voice assistants that reduce driver stress and improve safety. Future studies should investigate long-term effects and individual auditory differences. Zhao Zou, Fady Shibata-Alnajjar, Michael Lwin, Aila Khan, Abdullah Al Mahmud 0001, Omar Mubin |
HAI | 2 |
| 2025 | Enhancing road safety with DL vision: Do driver distraction alerts hold the key?
Luqman Ali, Muhammed Swavaf, Fady Shibata-Alnajjar, Zhao Zou, Medha Mohan Ambali Parambil, Omar Mubin, Hamad Aljassmi |
Neural Comput. Appl. | 3 |
| 2024 | SWIN and Vision Transformer-Driven Crack Detection in the Al Qattara Oasis, UAE: Towards Sustainable Infrastructure ManagementabstractHeritage sites are central to the cultural identity and historical narrative of a community. The Al Qattara Oasis, located in the United Arab Emirates (UAE), illustrates this function well. This study addresses the pressing need to preserve and maintain built heritage by employing modern technological solutions for infrastructure upkeep. Specifically, it focuses on crack detection, a critical aspect of ensuring the structural integrity of heritage buildings. Utilizing data collected from various structures across the UAE, obtained through handheld cameras, the effectiveness of Vision Transformers (ViTs) and Swin Transformers in identifying cracks within Al Qattara is assessed. The research thoroughly evaluates models, with particular attention to different input patch sizes. Through systematic experimentation and analysis over 100 epochs, ViT models with a patch size 16 exhibit significant promise in crack detection. Notably, the ViT-16 model achieves good performance metrics, including training and validation accuracies of 85% and (82%) respectively, and correctly classifies 791 out of 957 cracks and 772 out of 903 non-cracks. In contrast, Swin Transformers show lower validation accuracies (70-74%) and higher misclassification rates. The outcomes underscore the potential of ViT models in enhancing infrastructure maintenance efforts within heritage sites like Qattara Oasis. By combining advanced technology with a profound respect for historical preservation, this study aims to contribute to the sustainable conservation and protection of cultural heritage, ensuring that future generations can continue to appreciate the enduring legacy of sites such as Qattara Oasis. Luqman Ali, Medha Mohan Ambali Parambil, Muhammed Swavaf, Fady Shibata-Alnajjar, Hamad Aljassmi, Adriaan De Man |
BDCAT | 4 |
| 2024 | Empowering Helpers: Reversing Roles in Paediatric Rehab with Humanoid Robots and Sensory GamesabstractAbstract-In this study a new conceptual framework is presented for a rehabilitation game system designed specifically for children with cerebral palsy (CP), particularly those with high-level disabilities. Humanoid robot Pepper integrates advanced sensory technologies into this framework, leveraging empathetic interactions to foster children's motivation and engagement to assist. This innovative approach seeks to reverse the conventional dynamic in therapeutic care by empowering children to aid the robot, thereby potentially boosting their motivation and engagement. An observational study in a pediatric rehabilitation center, where therapists and children interacted during rehabilitation games/exercises, formed the basis of this study's methodological framework. Following a thematic analysis of these observations and previous studies findings, three empathic game scenarios were conceptualized. Through tablet interfaces enabled by IMU sensors, children assist the robot, creating a sense of accomplishment and agency among them. Despite the lack of direct experimental data with children-robot interaction, the paper describes the potential technical architecture of the system, including features such as real-time monitoring. Future empirical research on empathetic, interactive robotics in pediatric rehabilitation settings can build upon this framework to achieve further development and validation. Leila Mouzehkesh Pirborj, Fady Shibata-Alnajjar, Saman Shafigh |
IE | 2 |
| 2024 | Deep learning with image-based autism spectrum disorder analysis: A systematic review
Md. Zasim Uddin, Md. Arif Shahriar, Md. Nadim Mahamood, Fady Shibata-Alnajjar, Md. Ileas Pramanik, Md. Atiqur Rahman Ahad |
Eng. Appl. Artif. Intell. | 4 |
| 2023 | Fine-Tuning Vision Transformer for Arabic Sign Language Video Recognition on Augmented Small-Scale DatasetabstractWith the rise of AI, the recognition of Sign Language (SL) through sign-to-text has gained significance in the field of computer vision and deep machine learning. However, there are only a few medium to large open datasets available for this task, as it requires a vast dataset of thousands of signs for words/phrases in different environments, which is a time-consuming and tedious process. Furthermore, there has been very little effort towards Arabic Sign Language Recognition (ArSLR). This research paper presents the results of fine-tuning the Vision Transformer (ViT) model on a small-scale in-house dataset of ArSL. The main goal is to attain satisfactory results by utilizing minimal computing power and a small dataset involving less than 10 individuals, with only one recording made for each sign in every environment. The dataset comprises 49 classes/signs, all of which were made with two hands and belong to the Level I category in terms of popularity. To enhance the dataset, three types of augmentations - translation, shear, and rotation were employed. The ViT model, pre-trained on the Kinetics dataset, was trained on the variation of augmented datasets with 2 to 40 times samples for each original video, where the training set includes original and augmented videos of 8 volunteers and the test set includes only original videos of one particular volunteer. Experimental results reveal that the combination of rotation and shear outperformed the others, achieving an accuracy of 93% on the 20 times augmented samples per class per signer dataset. We believe this study sheds light on small-scale dataset-based SLR tasks and video/action recognition in general. Munkhjargal Gochoo, Ganzorig Batnasan, Ahmed Abdelhadi Ahmed, Munkh-Erdene Otgonbold, Fady Shibata-Alnajjar, Timothy K. Shih, Tan-Hsu Tan, Khin Wee Lai |
SMC | 5 |
| 2023 | Autism Spectrum Disorder Classification via Local and Global Feature Representation of Facial ImageabstractAutism Spectrum Disorder (ASD) is a neurodevelopmental disorder that affects social communication and interaction. Early diagnosis of ASD can mitigate the severity and help with ideal treatment direction. Computer vision-based methods with traditional machine learning and deep learning are employed in the literature for automatic diagnosis. Recently, deep learning with a facial image-based ASD classification has gained interest due to its ease of collection and non-invasiveness. We observed that the existing approaches utilized either local or global features of facial images to diagnose ASD. However, its important to consider both local and global features to obtain fine-grained details and larger contextual information for accurate detection and classification. This paper proposes a sequencer-based patch-wise Local Feature Extractor along with a Global Feature Extractor. Finally, the features from these modules are aggregated to obtain the final feature for the classification of ASD. Experiments on a publicly available Autism Facial Image Dataset demonstrate that our proposed framework achieves state-of-the-art performance. We achieved accuracy, precision, recall, and F1-score of 94.7%, 94.0%, 95.3%, and 94.6%, respectively. Md. Nadim Mahamood, Md. Zasim Uddin, Md. Arif Shahriar, Fady Shibata-Alnajjar, Md. Atiqur Rahman Ahad |
SMC | 4 |
| 2023 | A simulated measurement for COVID-19 pandemic using the effective reproductive number on an empirical portion of population: epidemiological models
Belal Alsinglawi, Omar Mubin, Fady Shibata-Alnajjar, Khalid Kheirallah, Mahmoud Elkhodr, Mohammed Al-Zobbi, Mauricio Novoa, Mudassar Arsalan, Tahmina Nasrin Poly, Munkhjargal Gochoo, Gulfaraz Khan, Kapal Dev |
Neural Comput. Appl. | 3 |
| 2022 | ArSL21L: Arabic Sign Language Letter Dataset Benchmarking and an Educational Avatar for Metaverse ApplicationsabstractIt is complicated for the PwHL (people with hearing loss) to make a relationship with social majority, which naturally demands an interactive auto computer systems that have ability to understand sign language. With a trending Metaverse applications using augmented reality (AR) and virtual reality (VR), it is easier and interesting to teach sign language remotely using an avatar that mimics the gesture of a person using AI (Artificial Intelligence)-based system. There are various proposed methods and datasets for English SL (sign language); however, it is limited for Arabic sign language. Therefore, we present our collected and annotated Arabic Sign Language Letters Dataset (ArSL21L) consisting of 14202 images of 32 letter signs with various backgrounds collected from 50 people. We benchmarked our ArSL21L dataset on state-of-the-art object detection models, i.e., 4 versions of YOLOv5. Among the models, YOLOv5l achieved the best result with COCOmAP of 0.83. Moreover, we provide comparison results of classification task between ArSL2018 dataset, the only Arabic sign language letter dataset for classification task, and our dataset by running classification task on in-house short video. The results revealed that the model trained on our dataset has a superior performance over the model trained on ArSL2018. Moreover, we have created our prototype avatar which can mimic the ArSL (Arabic Sign Language) gestures for Metaverse applications. Finally, we believe, ArSL21L and the ArSL avatar will offer an opportunity to enhance the research and educational applications for not only the PwHL, but also in general real and virtual world applications. Our ArSL21L benchmark dataset is publicly available for research use on the Mendeley. Ganzorig Batnasan, Munkhjargal Gochoo, Munkh-Erdene Otgonbold, Fady Shibata-Alnajjar, Timothy K. Shih |
EDUCON | 4 |
| 2022 | What doesn't kill you makes you stronger: Conceptualising social robot's pain and Consumer's Empathic response through touchabstractTo be welcomed as assistant robots in our daily lives, robots must be liberated from their rigid, programmed logic and made more emotive and empathic to engage with people on their terms [1]. One of the key factors in designing and developing more human-like robots is to understand human emotions and behaviours regarding their pain empathy. This research focuses on the role of emotional touch (painful touch) in humanoid robots in the field of empathy. Leila Mouzehkesh Pirborj, Omar Mubin, Michael Lwin, Aila Khan, Fady Shibata-Alnajjar |
HAI | 5 |
| 2022 | Emotion and memory model for social robots: a reinforcement learning based behaviour selectionabstractIn this paper, we propose a reinforcement learning (RL) mechanism for social robots to select an action based on users’ learning performance and social engagement. We applied this behavior selection mechanism to extend the emotion and memory model, which allows a robot to create a memory account of the user’s emotional events and adapt its behavior based on the developed memory. We evaluated the model in a vocabulary-learning task at a school during a children’s game involving robot interaction to see if the model results in maintaining engagement and improving vocabulary learning across the four different interaction sessions. Generally, we observed positive findings based on child vocabulary learning and sustaining social engagement during all sessions. Compared to the trends of a previous study, we observed a higher level of social engagement across sessions in terms of the duration of the user gaze toward the robot. For vocabulary retention, we saw similar trends in general but also showing high vocabulary retention across some sessions. The findings indicate the benefits of applying RL techniques that have a reward system based on multi-modal user signals or cues. Muneeb Imtiaz Ahmad, Fady Shibata-Alnajjar, Suleman Shahid, Omar Mubin |
Behav. Inf. Technol. | 3 |
| 2022 | HCI Research in the Middle East and North Africa: A Bibliometric and Socioeconomic OverviewabstractThe field of human–computer interaction (HCI) is a growing area of research. However, its ascendancy and evolution is not understood in the Middle East and North Africa (MENA) region, which is knocking on the doors of innovation and technological advancements. In our analysis, we present a bibliometric and socioeconomic overview of the progress of HCI research in the MENA region. Using a pool of 10 premier venues in the discipline, we extracted 549 papers. We show that authors from high-income MENA countries are the most prolific in terms of both research output and first authorship and that authors from upper-middle-income countries are more likely to publish their work in journals. The topical focus of conducted research was also observed to vary as per the socioeconomic category to which a country belonged to. The USA emerged as the standout collaborative partner of the MENA region, and there was an observed lack of intra-MENA associations. We conclude with the assertion that additional efforts are warranted to promote collaborations amongst the region. Our work is one of the first attempts to quantitatively chart the development of HCI in the MENA region using the socioeconomic level of a country as an important predictor. Omar Mubin, Fady Shibata-Alnajjar, Mudassar Arsalan |
Int. J. Hum. Comput. Interact. | 2 |
| 2022 | Bidirectional parallel echo state network for speech emotion recognitionabstractSpeech is an effective way for communicating and exchanging complex information between humans. Speech signal has involved a great attention in human-computer interaction. Therefore, emotion recognition from speech has become a hot research topic in the field of interacting machines with humans. In this paper, we proposed a novel speech emotion recognition system by adopting multivariate time series handcrafted feature representation from speech signals. Bidirectional echo state network with two parallel reservoir layers has been applied to capture additional independent information. The parallel reservoirs produce multiple representations for each direction from the bidirectional data with two stages of concatenation. The sparse random projection approach has been adopted to reduce the high-dimensional sparse output for each direction separately from both reservoirs. Random over-sampling and random under-sampling methods are used to overcome the imbalanced nature of the used speech emotion datasets. The performance of the proposed parallel ESN model is evaluated from the speaker-independent experiments on EMO-DB, SAVEE, RAVDESS, and FAU Aibo datasets. The results show that the proposed SER model is superior to the single reservoir and the state-of-the-art studies. Hemin Ibrahim, Chu Kiong Loo, Fady Shibata-Alnajjar |
Neural Comput. Appl. | 3 |
| 2021 | Grouped Echo State Network with Late Fusion for Speech Emotion Recognition
Hemin Ibrahim, Chu Kiong Loo, Fady Shibata-Alnajjar |
ICONIP (3) | 3 |
| 2021 | Ultra-Low Resolution Infrared Sensor-Based Wireless Sensor Network for Privacy-Preserved Recognition of Daily Activities of LivingabstractIn the last few decades, the number of elderly people living alone surge worldwide due to an increase in human life expectancy. Researchers suggest different types of methods to monitor elderly residents living alone to prevent them from being unable to get help on time after some incidents such as accidental falling or heart attack. High-resolution data generating methods in monitoring elderly people with RGB cameras or wearable devices are inconvenient for elderly residents. While the former raises a privacy concern, the latter is not practical to wear on and off or charge its battery frequently. One of the solutions that solve both aforementioned problems is ultra-low resolution infrared (IR) sensor arrays. We offer a dataset for Activities of Daily Living (ADL) collected with 8×8 IR sensor arrays from 74 volunteers. For ADL recognition, four types of deep learning models, Convolutional Neural Network (CNN), two types of Recurrent Neural Networks (RNN), and Transformer models are employed. Among them, CNN and Transformer models showed promising results. We believe the dataset is a good contribution to versatile data sources for researchers to accelerate their work on the development of privacy-preserved ADL recognition systems. Luubaatar Badarch, Munkhjargal Gochoo, Ganzorig Batnasan, Fady Shibata-Alnajjar, Tan-Hsu Tan |
NCA | 4 |
| 2020 | Benchmarking Predictive Models in Electronic Health Records: Sepsis Length of Stay Prediction
Belal Alsinglawi, Fady Shibata-Alnajjar, Omar Mubin, Mauricio Novoa, Ola Karajeh, Omar A. Darwish |
AINA | 2 |
| 2020 | Lownet: Privacy Preserved Ultra-Low Resolution Posture Image ClassificationabstractIndoor posture recognition is vital for monitoring/detecting exercises, activities of daily living, accidental falls, unusual behavior, etc. However, high-resolution image based systems have a high accuracy, they are considered as intrusive and most of the current state-of-the-art image classifiers (VGG, ImageNet, ResNext) are not applicable for ultra-low resolution (<; 32 pixels in extent) image classification due to their downsizing feature extraction architecture. Thus, we propose a shallow LowNet model for classifying privacy preserved 16x16 posture images with its feature preserving architecture, variable ReLU slopes, and a custom loss function. LowNet outperformed, with an Accuracy of 98.94% and F1-score of 79.86%, the existing models (LeNet, ResNet1, ResNet-2) which can run on our Ultra lowresolution Thermal Posture Image (UTPI38) dataset (offered here) with 38 classes (4374 samples) collected from 23 volunteers. More experimental results are discussed on the custom loss, and variable ReLU slopes which gave 8.2% performance increase. Thus, we conclude that LowNet is useful in a multiclass ultra-low-resolution thermal posture image classification task. Munkhjargal Gochoo, Tan-Hsu Tan, Fady Shibata-Alnajjar, Jun-Wei Hsieh, Ping-Yang Chen |
ICIP | 3 |
| 2019 | A Low-Cost Autonomous Attention Assessment System for Robot Intervention with Autistic ChildrenabstractAttention is an essential mental process that is important to achieve learning progress. We cannot get better in our academic learning unless we concentrate our attention on the person giving the educational material, such as the teacher or trainer. Children with Autism Spectrum Disorder (ASD) may have attention difficulties that can directly influence their academic skills. In recent years, robot intervention in autism therapy and assessment is becoming a popular research topic due to its role in enhancing children's attention more than a regular human therapist, as well as, the increasing number of autism children compared to the availability of professional therapists. Robot intervention helps in reducing therapy time and makes early therapeutics sessions easier and much promising. Many researches have been conducted to develop robot intervention techniques for ASD children, and some methods have already been used to assess autistic individuals' attention during the robot intervention sessions. Yet, the existing attention assessment methods are either very complex or simple with one measured interaction cue only. This paper presents a practical and low-cost automatic approach to assess autistic individuals' attention during robot intervention; addressing multiple interaction cues. Experimental results show that the proposed attention assessment system could accurately measure the child attention and enhance therapy progress. This automatic attention system can open a new era for utilizing technologies to monitor students' attentions in the class to enhance educational systems. Fady Shibata-Alnajjar, Abdulrahman Majed Renawi, Massimiliano Lorenzo Cappuccio, Omar Mubin |
EDUCON | 1 |
| 2019 | Novel IoT-Based Privacy-Preserving Yoga Posture Recognition System Using Low-Resolution Infrared Sensors and Deep LearningabstractIn recent years, the number of yoga practitioners has been drastically increased and there are more men and older people practice yoga than ever before. Internet of Things (IoT)-based yoga training system is needed for those who want to practice yoga at home. Some studies have proposed RGB/Kinect camera-based or wearable device-based yoga posture recognition methods with a high accuracy; however, the former has a privacy issue and the latter is impractical in the long-term application. Thus, this paper proposes an IoT-based privacy-preserving yoga posture recognition system employing a deep convolutional neural network (DCNN) and a low-resolution infrared sensor-based wireless sensor network (WSN). The WSN has three nodes (x, y, and z-axes) where each integrates 8 × 8 pixels' thermal sensor module and a Wi-Fi module for connecting the deep learning server. We invited 18 volunteers to perform 26 yoga postures for two sessions each lasted for 20 s. First, recorded sessions are saved as .csv files, then preprocessed and converted to grayscale posture images. Totally, 93200 posture images are employed for the validation of the proposed DCNN models. The tenfold cross-validation results revealed that F1-scores of the models trained with xyz (all 3-axes) and y (only y-axis) posture images were 0.9989 and 0.9854, respectively. An average latency for a single posture image classification on the server was 107 ms. Thus, we conclude that the proposed IoT-based yoga posture recognition system has a great potential in the privacy-preserving yoga training system. Munkhjargal Gochoo, Tan-Hsu Tan, Shih-Chia Huang, Tsedevdorj Batjargal, Jun-Wei Hsieh, Fady Shibata-Alnajjar, Yung-fu Chen |
IEEE Internet Things J. | 6 |
| 2019 | Unobtrusive Activity Recognition of Elderly People Living Alone Using Anonymous Binary Sensors and DCNNabstractElderly population (over the age of 60) is predicted to be 1.2 billion by 2025. Most of the elderly people would like to stay alone in their own house due to the high eldercare cost and privacy invasion. Unobtrusive activity recognition is the most preferred solution for monitoring daily activities of the elderly people living alone rather than the camera and wearable devices based systems. Thus, we propose an unobtrusive activity recognition classifier using deep convolutional neural network (DCNN) and anonymous binary sensors that are passive infrared motion sensors and door sensors. We employed Aruba annotated open data set that was acquired from a smart home where a voluntary single elderly woman was living inside for eight months. First, ten basic daily activities, namely, Eating, Bed_to_Toilet, Relax, Meal_Preparation, Sleeping, Work, Housekeeping, Wash_Dishes, Enter_Home, and Leave_Home are segmented with different sliding window sizes, and then converted into binary activity images. Next, the activity images are employed as the ground truth for the proposed DCNN model. The 10-fold cross-validation evaluation results indicated that our proposed DCNN model outperforms the existing models with F1-score of 0.79 and 0.951 for all ten activities and eight activities (excluding Leave_Home and Wash_Dishes), respectively. Munkhjargal Gochoo, Tan-Hsu Tan, Shing-Hong Liu, Fu-Rong Jean, Fady Shibata-Alnajjar, Shih-Chia Huang |
IEEE J. Biomed. Health Informatics | 5 |
| 2018 | Grasp-training Robot to Activate Neural Control Loop for Reflex and Experimental VerificationabstractUsing a rehabilitation robot to activate motion intention and reflex response simultaneously is an effective approach to aiding recovery from paralysis caused by neurological disorders. Mechanical motions supported by conventional robots are, however, not enough to activate reflex. In this paper, we propose a grasp-training robot that can stimulate the grasp reflex of a paralyzed hand by pushing the hand onto an elastic bar while supporting the grasping movements. In addition to this feature, we discuss the robot design in relation to its usability and wearability for ease of use in clinical practice. Experimental results obtained from healthy subjects show that the proposed robot can support grasping in a way similar to the traditional range-of-motion exercise used by therapists for grasp rehabilitation. Combining this appropriate grasping-motion support and the mechanism for pushing the hand onto an elastic bar succeeds in activating the grasp reflex of a completely paralyzed patient in a clinical test that involves monitoring electromyography signals from the paralyzed hand. Shotaro Okajima, Fady Shibata-Alnajjar, Hiroshi Yamasaki, Matti Itkonen, Álvaro Costa-García 0001, Yasuhisa Hasegawa, Shingo Shimoda |
ICRA | 2 |
| 2012 | Static and dynamic memory to simulate higher-order cognitive tasksabstractThe foremost objective of our research series is to construct a neurocomputational model that aims to achieve a Large-Scale Brain Network, and to suggest a possible insight of how the macro-level anatomical structures, such as the connectivity between the frontal lobe regions and their dynamic properties, can be self-organized to obtain the higher-order cognitive mechanisms, such as: planning, reasoning, task switching, cognitive branching, etc. For addressing these issues, this paper, in particular, focuses in proposing a model that intends to clarify the neural structure and mechanisms underlying the task switching and the cognitive branching condition. Although both tasks requiring varying degree of a working memory, in contrast to the switching task, where the primary ongoing task is entirely replaced by a new task, in the branching task, a delaying to the execution of an original task occurs until the completion of a subordinate task. The proposed model is constructed by a hierarchical Multiple Timescale Recurrent Neural Network (MTRNN) and conducted on a humanoid robot in a physical environment. Experimental results suggest essential factors related to the neural activities and network's structure necessary to form a suitable working memory for accomplishing such tasks. Fady Shibata-Alnajjar, Yuichi Yamashita, Jun Tani |
IJCNN | 1 |
| 2009 | A Novel Hierarchical Constructive BackPropagation with Memory for Teaching a Robot the Names of Things
Fady Shibata-Alnajjar, Abdul Rahman Hafiz, Kazuyuki Murase |
ICONIP (1) | 1 |
| 2009 | Vision-Motor Abstraction toward Robot Cognition
Fady Shibata-Alnajjar, Abdul Rahman Hafiz, Indra Bin Mohd Zin, Kazuyuki Murase |
ICONIP (2) | 1 |
| 2008 | Vision-sensorimotor abstraction and imagination towards exploring robot's inner worldabstractBased on indications from the neuroscience and psychology, both perception and action can be internally simulated by activating sensor and motor areas in the brain without external sensory input or without any resulting overt behavior. This hypothesis, however, can be highly useful in the real robot applications. The robot, for instance, can cover some of the corrupted sensory inputs by replacing them with its internal simulation. The accuracy of this hypothesis is strongly based on the agentpsilas experiences. As much as the agent knows about the environment, as much as it can build a strong internal representation about it. Although many works have been presented regarding to this hypothesis with various levels of success. At the sensorimotor abstraction level, where extracting data from the environment occur, however, none of them have so far used the robotpsilas vision as a sensory input. In this study, vision-sensorimotor abstraction is presented through memory-based learning in a real mobile robot ldquoHemissonrdquo to investigate the possibilities of explaining its inner world based on internal simulation of perception and action at the abstract level. The analysis of the experiments illustrate that our robot with vision sensory input has developed some kind of simple associations or anticipation mechanism through interacting with the environment, which enables, based on its history and the present situation, to guide its behavior in the absence of any external interaction. Fady Shibata-Alnajjar, Abdul Rahman Hafiz, Indra Bin Mohd Zin, Kazuyuki Murase |
IJCNN | 1 |
| 2008 | Sensor-fusion in spiking neural network that generates autonomous behavior in real mobile robotabstractWe here introduce a novel adaptive controller for autonomous mobile robot that binds N types of sensory information. For each sensory modality, sensory-motor connection is made by a three-layered spiking neural network (SNN). The synaptic weights in the model have the property of spike timing-dependent plasticity (STDP) and regulated by presynaptic modulation signal from the sensory neurons. Each synaptic weight is incrementally adapted depending upon the firing rate of the presynaptic modulation signal and that of the hidden-layer neurons). Information from different types of sensors are bound at the motor neurons. A real mobile robot Khepera with the SNN controller quickly adapted into an open environment and performed the desired task successfully. This approach could be applicable to a robot with inputs of various sensory modalities and various types of motor outputs. Fady Shibata-Alnajjar, Kazuyuki Murase |
IJCNN | 1 |
| 2008 | A Spiking Neural Network with dynamic memory for a real autonomous mobile robot in dynamic environmentabstractThis work concerns practical issues surrounding the application of learning and memory in a real mobile robot towards optimal navigation in dynamic environments. A novel control system that contains two-level units (low-level and high-level) is developed and trained in a physical mobile robot ldquoe-Puckrdquo. In the low-level unit, the robotpsilas task is to navigate in a various local environments, by training N numbers of spiking neural networks (SNN) that have the property of spike time-dependent plasticity. All the trained SNNs are stored in a tree-type memory structure, which is located in the high-level unit. These stored networks are used as experiences for the robot to enhance its navigation ability in new and previously trained environments. The memory is designed to hold memories of various lengths and has a simple searching mechanism. For controlling the memory size, forgetting and on-line dynamic clustering techniques are used. Experimental results have proved that the proposed model can provide a robot with learning and memorizing capabilities enable it to survive in complex and dynamic environments. Fady Shibata-Alnajjar, Indra Bin Mohd Zin, Kazuyuki Murase |
IJCNN | 1 |
| 2006 | Self-organization of Spiking Neural Network that Generates Autonomous Behavior in a Real Mobile RobotabstractIn this paper, we propose self-organization algorithm of spiking neural network (SNN) applicable to autonomous robot for generation of adoptive and goal-directed behavior. First, we formulated a SNN model whose inputs and outputs were analog and the hidden unites are interconnected each other. Next, we implemented it into a miniature mobile robot Khepera. In order to see whether or not a solution(s) for the given task(s) exists with the SNN, the robot was evolved with the genetic algorithm in the environment. The robot acquired the obstacle avoidance and navigation task successfully, exhibiting the presence of the solution. After that, a self-organization algorithm based on a use-dependent synaptic potentiation and depotentiation at synapses of input layer to hidden layer and of hidden layer to output layer was formulated and implemented into the robot. In the environment, the robot incrementally organized the network and the given tasks were successfully performed. The time needed to acquire the desired adoptive and goal-directed behavior using the proposed self-organization method was much less than that with the genetic evolution, approximately one fifth. Fady Shibata-Alnajjar, Kazuyuki Murase |
Int. J. Neural Syst. | 1 |