VLDB 2026 Research / reviewers in the wild / expert
Youssef Mohamed
dblp:282/0389
· DBLP profile ↗
11ranked-venue papers
9as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 8 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | XProvence: Zero-Cost Multilingual Context Pruning for Retrieval-Augmented Generation
Youssef Mohamed, Mohamed Elhoseiny 0001, Thibault Formal, Nadezhda Chirkova |
ECIR (2) | 1 |
| 2025 | DeepChest: Dynamic Gradient-Free Task Weighting for Effective Multi-Task Learning in Chest X-Ray Classification
Youssef Mohamed, Noran Mohamed, Khaled Abouhashad, Sara Atito Ali Ahmed, Shoaib Jameel, Muhammad Imran Razzak, Ahmed B. Zaky |
IEEE Big Data | 1 |
| 2025 | Towards AI-Assisted Psychotherapy: Emotion-Guided Generative InterventionsabstractLarge language models (LLMs) hold promise for therapeutic interventions, yet most existing datasets rely solely on text, overlooking nonverbal emotional cues essential to real-world therapy.To address this, we introduce a multimodal dataset of 1,441 publicly sourced therapy session videos containing both dialogue and non-verbal signals such as facial expressions and vocal tone.Inspired by Hochschild's concept of emotional labor, we propose a computational formulation of emotional dissonance-the mismatch between facial and vocal emotion-and use it to guide emotionally aware prompting.Our experiments show that integrating multimodal cues, especially dissonance, improves the quality of generated interventions.We also find that LLM-based evaluators misalign with expert assessments in this domain, highlighting the need for humancentered evaluation.Project page link: https: //kilichbek.github.io/webpage/mental/ Kilichbek Haydarov, Youssef Mohamed, Emilio Goldenhersch, Paul OCallaghan, Li-jia Li |
EMNLP | 2 |
| 2025 | Fusion in Context: A Multimodal Approach to Affective State RecognitionabstractAccurate recognition of human emotions is a crucial challenge in affective computing and human-robot interaction (HRI). Emotional states play a vital role in shaping behaviors, decisions, and social interactions. However, emotional expressions can be influenced by contextual factors, leading to misinterpretations if context is not considered. Multimodal fusion, combining modalities like facial expressions, speech, and physiological signals, has shown promise in improving affect recognition. This paper proposes a transformer-based multimodal fusion approach that leverages facial thermal data, facial action units, and textual context information for context-aware emotion recognition. We explore modality-specific encoders to learn tailored representations, which are then fused and processed by a shared transformer encoder to capture temporal dependencies and interactions. The proposed method is evaluated on a dataset collected from participants engaged in a tangible tabletop Pacman game designed to induce various affective states. Our results demonstrate improvements from incorporating contextual information and multimodal fusion, achieving 89% F1 score with our full model compared to 65% for action units alone and 30% for thermal data alone. Youssef Mohamed, Séverin Lemaignan, Arzu Güneysu, Patric Jensfelt, Christian Smith |
RO-MAN | 1 |
| 2025 | Are You an Expert? Instruction Adaptation Using Multi-Modal Affect Detections with Thermal Imaging and ContextabstractHuman-robot interactions increasingly require adaptive instruction delivery, yet robots struggle to calibrate instruction detail levels without explicit user input. We present a system that automatically modulates instruction granularity using real-time affect detection through multi-modal fusion of thermal imaging, facial expressions, and contextual information. Our transformer-based architecture integrates these signals to enable decisions about instruction delivery based on detected user states. In a between-subjects study (N=40), participants completed assembly tasks under either manual adjustment or automatic adaptation conditions. Results showed significantly fewer manual adjustments in the adaptive condition (0.7 vs 2.0 per session), with comparable user satisfaction across conditions. This work shows the effectiveness of affect-driven adaptive instruction in human-robot interaction, contributing to more responsive robotic interfaces while providing guidelines for balancing automation with user control. Youssef Mohamed, Séverin Lemaignan, Arzu Güneysu, Patric Jensfelt, Christian Smith |
RO-MAN | 1 |
| 2024 | No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 LanguagesabstractYoussef Mohamed, Runjia Li, Ibrahim Said Ahmad, Kilichbek Haydarov, Philip Torr, Kenneth Church, Mohamed Elhoseiny. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Youssef Mohamed, Runjia Li, Ibrahim Said Ahmad, Kilichbek Haydarov, Philip Torr 0001, Kenneth Church 0001, Mohamed Elhoseiny 0001 |
EMNLP | 1 |
| 2024 | Continual Learning on a Diet: Learning from Sparsely Labeled Streams Under Constrained ComputationabstractWe propose and study a realistic Continual Learning (CL) setting where learning algorithms are granted a restricted computational budget per time step while training. We apply this setting to large-scale semi-supervised Continual Learning scenarios with sparse label rate. Previous proficient CL methods perform very poorly in this challenging setting. Overfitting to the sparse labeled data and insufficient computational budget are the two main culprits for such a poor performance. Our new setting encourages learning methods to effectively and efficiently utilize the unlabeled data during training. To that end, we propose a simple but highly effective baseline, DietCL, which utilizes both unlabeled and labeled data jointly. DietCL meticulously allocates computational budget for both types of data. We validate our baseline, at scale, on several datasets, e.g., CLOC, ImageNet10K, and CGLM, under constraint budget setup. DietCL outperforms, by a large margin, all existing supervised CL algorithms as well as more recent continual semi-supervised methods. Our extensive analysis and ablations demonstrate that DietCL is stable under a full spectrum of label sparsity, computational budget and various other ablations. Wenxuan Zhang 0003, Youssef Mohamed, Bernard Ghanem, Philip Torr 0001, Adel Bibi, Mohamed Elhoseiny 0001 |
ICLR | 2 |
| 2022 | It is Okay to Not Be Okay: Overcoming Emotional Bias in Affective Image Captioning by Contrastive Data CollectionabstractDatasets that capture the connection between vision, language, and affection are limited, causing a lack of understanding of the emotional aspect of human intelligence. As a step in this direction, the ArtEmis dataset was recently introduced as a large-scale dataset of emotional reactions to images along with language explanations of these chosen emotions. We observed a significant emotional bias towards instance-rich emotions, making trained neural speakers less accurate in describing under-represented emotions. We show that collecting new data, in the same way, is not effective in mitigating this emotional bias. To remedy this problem, we propose a contrastive data collection approach to balance ArtEmis with a new complementary dataset such that a pair of similar images have contrasting emotions (one positive and one negative). We collected 260,533 instances using the proposed method, we combine them with ArtEmis, creating a second iteration of the dataset. The new combined dataset, dubbed ArtEmis v2.0, has a balanced distribution of emotions with explanations revealing more fine details in the associated painting. Our experiments show that neural speakers trained on the new dataset improve CIDEr and METEOR evaluation metrics by 20% and 7%, respectively, compared to the biased dataset. Finally, we also show that the performance per emotion of neural speakers is improved across all the emotion categories, significantly on under-represented emotions. The collected dataset and code are available at https://artemisdataset-v2.org. Youssef Mohamed, Faizan Farooq Khan, Kilichbek Haydarov, Mohamed Elhoseiny 0001 |
CVPR | 1 |
| 2022 | ArtELingo: A Million Emotion Annotations of WikiArt with Emphasis on Diversity over Language and CultureabstractYoussef Mohamed, Mohamed Abdelfattah, Shyma Alhuwaider, Feifan Li, Xiangliang Zhang, Kenneth Church, Mohamed Elhoseiny. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Youssef Mohamed, Mohamed Abdelfattah, Shyma Alhuwaider, Xiangliang Zhang 0001, Kenneth Church 0001, Mohamed Elhoseiny 0001 |
EMNLP | 1 |
| 2022 | Automatic Frustration Detection Using Thermal ImagingabstractTo achieve seamless interactions, robots have to be capable of reliably detecting affective states in real time. One of the possible states that humans go through while interacting with robots is frustration. Detecting frustration from RGB images can be challenging in some real-world situations; thus, we investigate in this work whether thermal imaging can be used to create a model that is capable of detecting frustration induced by cognitive load and failure. To train our model, we collected a data set from 18 participants experiencing both types of frustration induced by a robot. The model was tested using features from several modalities: thermal, RGB, Electrodermal Activity (EDA), and all three combined. When data from both frustration cases were combined and used as training input, the model reached an accuracy of 89% with just RGB features, 87% using only thermal features, 84% using EDA, and 86% when using all modalities. Furthermore, the highest accuracy for the thermal data was reached using three facial regions of interest: nose, forehead and lower lip. Youssef Mohamed, Giulia Ballardini, Maria Teresa Parreira, Séverin Lemaignan, Iolanda Leite |
HRI | 1 |
| 2021 | ROS for Human-Robot InteractionabstractIntegrating real-time, complex social signal processing into robotic systems – especially in real-world, multi-party interaction situations – is a challenge faced by many in the Human-Robot Interaction (HRI) community. The difficulty is compounded by the lack of any standard model for human representation that would facilitate the development and interoperability of social perception components and pipelines. We introduce in this paper a set of conventions and standard interfaces for HRI scenarios, designed to be used with the Robot Operating System (ROS). It directly aims at promoting interoperability and re-usability of core functionality between the many HRI-related software tools, from skeleton tracking, to face recognition, to natural language processing. Importantly, these interfaces are designed to be relevant to a broad range of HRI applications, from high-level crowd simulation, to group-level social interaction modelling, to detailed modelling of human kinematics. We demonstrate these interfaces by providing a reference pipeline implementation, packaged to be easily downloaded and evaluated by the community. Youssef Mohamed, Séverin Lemaignan |
IROS | 1 |