Amr Gomaa

dblp:186/7372 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0003-0955-3181ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 7 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Do You (Dis)agree With Me? Modelling Implicit User Disagreement in Human-AI Interaction Using Gaze Data
abstract
The widespread use of generative AI has led to increased focus on human–AI interaction. However, AI systems can generate unexpected outputs, leading to disagreement or human–AI conflict. This paper focuses on modelling user disagreement using machine learning (ML) by observing users’ implicit viewing behaviour. We conducted a controlled study with 30 participants evaluating captions from a simulated ML image-captioning system. Participants indicated agreement or disagreement with each caption while we recorded their gaze and facial-expression data, which we used to predict (dis)agreement. We show that unimodal gaze-based personalised modelling (0.684 average balanced accuracy) outperforms generalised modelling (0.570), whereas multimodal approaches did not improve performance. Our exploratory post hoc gaze-based analysis highlights the importance of feature selection and temporal dynamics, which help guide system design and future work. We release the dataset to support reproducibility and further work. Due to the nature of this research, we also discuss the potential ethical and privacy implications of continuous passive gaze and facial monitoring.
Abdulrahman Mohamed Selim, Omair Shahzad Bhatti, Amr Gomaa, Michael Barz, Daniel Sonntag
CHI3
2026 Looking for an Even Better Fit? A Personalized Incremental Reinforcement Learning Multimodal Object Referencing Framework
abstract
The rapid advancement of the automotive industry toward automated and semi-automated vehicles has rendered traditional methods of vehicle interaction, such as touch-based and voice command systems, inadequate for a widening range of non-driving related tasks, such as referencing objects outside of the vehicle. Consequently, research has shifted toward gestural input (e.g., hand, gaze, and head pose gestures) as a more suitable mode of interaction during driving. However, due to the dynamic nature of driving and individual variation, there are significant differences in drivers’ gestural input performance. While, in theory, this inherent variability could be moderated by substantial data-driven machine learning models, prevalent methodologies lean toward constrained, single-instance trained models for object referencing. These models show a limited capacity to continuously adapt to the divergent behaviors of individual drivers and the variety of driving scenarios. To address this, we proposed, in our previous work, an incremental learning approach, adapting to changing user behavior. Although this method proved superior over state-of-the-art approaches, it still lacked in performance and had some computational resource limitations, mainly the small-sized memory needed to store a subset of old data for incremental learning. Therefore, in this work, we propose an alternative solution using reinforcement learning that can be used to overcome memory storage limitations, achieving a zero data storage approach for incremental learning without performance loss. Our newly enhanced solution has been added to our previous framework at https://github.com/amrgomaaelhady/IcRegress .
Amr Gomaa, Kiran Gani, Michael Feld, Antonio Krüger
ACM Trans. Interact. Intell. Syst.1
2024 Towards a Surgeon-in-the-Loop Ophthalmic Robotic Apprentice using Reinforcement and Imitation Learning
abstract
Robot-assisted surgical systems have demonstrated significant potential in enhancing surgical precision and minimizing human errors. However, existing systems cannot accommodate individual surgeons’ unique preferences and requirements. Additionally, they primarily focus on general surgeries (e.g., laparoscopy) and are unsuitable for highly precise microsurgeries, such as ophthalmic procedures. Thus, we propose an image-guided approach for surgeon-centered autonomous agents that can adapt to the individual surgeon’s skill level and preferred surgical techniques during ophthalmic cataract surgery. Our approach trains reinforcement and imitation learning agents simultaneously using curriculum learning approaches guided by image data to perform all tasks of the incision phase of cataract surgery. By integrating the surgeon’s actions and preferences into the training process, our approach enables the robot to implicitly learn and adapt to the individual surgeon’s unique techniques through surgeon-in-the-loop demonstrations. This results in a more intuitive and personalized surgical experience for the surgeon while ensuring consistent performance for the autonomous robotic apprentice. We define and evaluate the effectiveness of our approach in a simulated environment using our proposed metrics and highlight the trade-off between a generic agent and a surgeon-centered adapted agent. Finally, our approach has the potential to extend to other ophthalmic and microsurgical procedures, opening the door to a new generation of surgeon-in-the-loop autonomous surgical robots. We provide an open-source simulation framework for future development and reproducibility at https://github.com/amrgomaaelhady/CataractAdaptSurgRobot.
Amr Gomaa, Bilal Mahdy, Niko Kleer, Antonio Krüger
IROS1
2024 Bridging the Gap to Natural Language-based Grasp Predictions through Semantic Information Extraction
abstract
Enabling multi-fingered robots to choose an appropriate grasp on an object from natural language instructions poses great difficulties for such systems. The diversity, imprecision, and limited information contained in the language make this task particularly challenging. However, speech serves humans as a natural communication interface that can aid robots in adapting to the environment more easily. Therefore, providing robots with relevant data about the objects they interact with is essential for them to understand how to carry out object manipulation tasks. By leveraging Named Entity Recognition (NER) to automatically extract semantic data, our work introduces a novel approach to text-based grasp predictions. Our methodology involves a multistage learning approach using a semantic information extractor that provides significant features to a grasp prediction model. To assess the effectiveness of our approach, we conducted experiments on an existing corpus and two corpora generated by ChatGPT. Our results demonstrate superior performance compared to similar grasp prediction models while overcoming limitations in the literature. Additionally, we open-source our training data for reproducibility and future research advancement.
Niko Kleer, Martin Feick, Amr Gomaa, Michael Feld, Antonio Krüger
IROS3
2024 Looking for a better fit? An Incremental Learning Multimodal Object Referencing Framework adapting to Individual Drivers
abstract
The rapid advancement of the automotive industry towards automated and semi-automated vehicles has rendered traditional methods of vehicle interaction, such as touch-based and voice command systems, inadequate for a widening range of non-driving related tasks, such as referencing objects outside of the vehicle. Consequently, research has shifted toward gestural input (e.g., hand, gaze, and head pose gestures) as a more suitable mode of interaction during driving. However, due to the dynamic nature of driving and individual variation, there are significant differences in drivers’ gestural input performance. While, in theory, this inherent variability could be moderated by substantial data-driven machine learning models, prevalent methodologies lean towards constrained, single-instance trained models for object referencing. These models show a limited capacity to continuously adapt to the divergent behaviors of individual drivers and the variety of driving scenarios. To address this, we propose IcRegress, a novel regression-based incremental learning approach that adapts to changing behavior and the unique characteristics of drivers engaged in the dual task of driving and referencing objects. We suggest a more personalized and adaptable solution for multimodal gestural interfaces, employing continuous lifelong learning to enhance driver experience, safety, and convenience. Our approach was evaluated using an outside-the-vehicle object referencing use case, highlighting the superiority of the incremental learning models adapted over a single trained model across various driver traits such as handedness, driving experience, and numerous driving conditions. Finally, to facilitate reproducibility, ease deployment, and promote further research, we offer our approach as an open-source framework at https://github.com/amrgomaaelhady/IcRegress.
Amr Gomaa, Guillermo Reyes, Michael Feld, Antonio Krüger
IUI1
2024 SynthoGestures: A Multi-Camera Framework for Generating Synthetic Dynamic Hand Gestures for Enhanced Vehicle Interaction
abstract
Dynamic hand gesture recognition is crucial for human-machine interfaces in the automotive domain. However, creating a diverse and comprehensive dataset of hand gestures can be challenging and time-consuming, especially in dynamic dual-task situations like driving. To address these challenges, we propose using synthetic gesture datasets generated by virtual 3D models as an alternative. Our framework synthesizes realistic hand gestures using a combination of 3D models and animation software, particularly utilizing Unreal Engine. This approach enables the creation of diverse and customizable gesture datasets, reducing the risk of overfitting and improving the model’s generalizability. Specifically, our framework generates natural-looking dynamic hand gestures with multiple variants, including gesture speed, performance, and hand shape. Moreover, we simulate various camera locations, such as above the driver and behind the wheel, and different camera types, such as RGB, infrared, and depth cameras, without incurring additional time and cost to obtain these cameras. Our experiments demonstrate that our proposed framework, SynthoGestures (available at https://github.com/amrgomaaelhady/SynthoGestures), can augment or replace existing real-hand datasets with additional enhancement in gesture recognition accuracy. Our tool for generating synthetic static and dynamic hand gestures saves time and effort in creating large datasets, facilitating the faster development of gesture recognition systems for automotive applications.
Amr Gomaa, Robin Zitt, Guillermo Reyes, Antonio Krüger
IV1
2024 Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation
abstract
There is a growing interest in using Large Language Models (LLMs) in multi-agent systems to tackle interactive real-world tasks that require effective collaboration and assessing complex situations. Yet, we have a limited understanding of LLMs' communication and decision-making abilities in multi-agent setups. The fundamental task of negotiation spans many key features of communication, such as cooperation, competition, and manipulation potentials. Thus, we propose using scorable negotiation to evaluate LLMs. We create a testbed of complex multi-agent, multi-issue, and semantically rich negotiation games. To reach an agreement, agents must have strong arithmetic, inference, exploration, and planning capabilities while integrating them in a dynamic and multi-turn setup. We propose metrics to rigorously quantify agents' performance and alignment with the assigned role. We provide procedures to create new games and increase games' difficulty to have an evolving benchmark. Importantly, we evaluate critical safety aspects such as the interaction dynamics between agents influenced by greedy and adversarial players. Our benchmark is highly challenging; GPT-3.5 and small models mostly fail, and GPT-4 and SoTA large models (e.g., Llama-3 70b) still underperform in reaching agreement in non-cooperative and more difficult games.
Sahar Abdelnabi, Amr Gomaa, Sarath Sivaprasad, Lea Schönherr, Mario Fritz
NeurIPS2
2024 Incorporation of the Intended Task into a Vision-based Grasp Type Predictor for Multi-fingered Robotic Grasping
abstract
Robots that make use of multi-fingered or fully anthropomorphic end-effectors can engage in highly complex manipulation tasks. However, the choice of a suitable grasp for manipulating an object is strongly influenced by factors such as the physical properties of an object and the intended task. This makes predicting an appropriate grasping pose for carrying out a concrete task notably challenging. At the same time, current grasp type predictors rarely consider the task as a part of the prediction process. This work proposes a learning model that considers the task in addition to an object’s visual features for predicting a suitable grasp type. Furthermore, we generate a synthetic dataset by simulating robotic grasps on 3D object models based on the BarrettHand end-effector. With an angular similarity of 0.9 and above, our model achieves competitive prediction results compared to grasp type predictors that do not consider the intended task for learning grasps. Finally, to foster research in the field, we make our synthesized dataset available to the research community.
Niko Kleer, Ole Keil, Martin Feick, Amr Gomaa, Tim Schwartz, Michael Feld
RO-MAN4
2023 It's all about you: Personalized in-Vehicle Gesture Recognition with a Time-of-Flight Camera
abstract
Despite significant advances in gesture recognition technology, recognizing gestures in a driving environment remains challenging due to limited and costly data and its dynamic, ever-changing nature. In this work, we propose a model-adaptation approach to personalize the training of a CNNLSTM model and improve recognition accuracy while reducing data requirements. Our approach contributes to the field of dynamic hand gesture recognition while driving by providing a more efficient and accurate method that can be customized for individual users, ultimately enhancing the safety and convenience of in-vehicle interactions, as well as driver’s experience and system trust. We incorporate hardware enhancement using a time-of-flight camera and algorithmic enhancement through data augmentation, personalized adaptation, and incremental learning techniques. We evaluate the performance of our approach in terms of recognition accuracy, achieving up to 90%, and show the effectiveness of personalized adaptation and incremental learning for a user-centered design.
Guillermo Reyes, Amr Gomaa, Michael Feld
AutomotiveUI2
2022 Adaptive User-Centered Multimodal Interaction towards Reliable and Trusted Automotive Interfaces
abstract
With the recently increasing capabilities of modern vehicles, novel approaches for interaction emerged that go beyond traditional touch-based and voice command approaches. Therefore, hand gestures, head pose, eye gaze, and speech have been extensively investigated in automotive applications for object selection and referencing. Despite these significant advances, existing approaches mostly employ a one-model-fits-all approach unsuitable for varying user behavior and individual differences. Moreover, current referencing approaches either consider these modalities separately or focus on a stationary situation, whereas the situation in a moving vehicle is highly dynamic and subject to safety-critical constraints. In this paper, I propose a research plan for a user-centered adaptive multimodal fusion approach for referencing external objects from a moving vehicle. The proposed plan aims to provide an open-source framework for user-centered adaptation and personalization using user observations and heuristics, multimodal fusion, clustering, transfer-of-learning for model adaptation, and continuous learning, moving towards trusted human-centered artificial intelligence.
Amr Gomaa
ICMI1
2021 ML-PersRef: A Machine Learning-based Personalized Multimodal Fusion Approach for Referencing Outside Objects From a Moving Vehicle
abstract
Over the past decades, the addition of hundreds of sensors to modern vehicles has led to an exponential increase in their capabilities. This allows for novel approaches to interaction with the vehicle that go beyond traditional touch-based and voice command approaches, such as emotion recognition, head rotation, eye gaze, and pointing gestures. Although gaze and pointing gestures have been used before for referencing objects inside and outside vehicles, the multimodal interaction and fusion of these gestures have so far not been extensively studied. We propose a novel learning-based multimodal fusion approach for referencing outside-the-vehicle objects while maintaining a long driving route in a simulated environment. The proposed multimodal approaches outperform single-modality approaches in multiple aspects and conditions. Moreover, we also demonstrate possible ways to exploit behavioral differences between users when completing the referencing task to realize an adaptable personalized system for each driver. We propose a personalization technique based on the transfer-of-learning concept for exceedingly small data sizes to enhance prediction and adapt to individualistic referencing behavior. Our code is publicly available at https://github.com/amr-gomaa/ML-PersRef.
Amr Gomaa, Guillermo Reyes, Michael Feld
ICMI1
2020 Studying Person-Specific Pointing and Gaze Behavior for Multimodal Referencing of Outside Objects from a Moving Vehicle
abstract
Hand pointing and eye gaze have been extensively investigated in automotive applications for object selection and referencing. Despite significant advances, existing outside-the-vehicle referencing methods consider these modalities separately. Moreover, existing multimodal referencing methods focus on a static situation, whereas the situation in a moving vehicle is highly dynamic and subject to safety-critical constraints. In this paper, we investigate the specific characteristics of each modality and the interaction between them when used in the task of referencing outside objects (e.g. buildings) from the vehicle. We furthermore explore person-specific differences in this interaction by analyzing individuals' performance for pointing and gaze patterns, along with their effect on the driving task. Our statistical analysis shows significant differences in individual behaviour based on object's location (i.e. driver's right side vs. left side), object's surroundings, driving mode (i.e. autonomous vs. normal driving) as well as pointing and gaze duration, laying the foundation for a user-adaptive approach.
Amr Gomaa, Guillermo Reyes, Alexandra Alles, Lydia Rupp, Michael Feld
ICMI1