Hanan Salam

dblp:96/9605 · DBLP profile ↗
← Back
21ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0001-6971-5264ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 1 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 9 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 CCD-Bench: Probing Cultural Conflict in Large Language Model Decision-Making
abstract
Large language models (LLMs) increasingly shape interpersonal and societal decision-making, yet their ability to navigate explicit conflicts between legitimate cultural values remains underexplored. Existing benchmarks focus on cultural knowledge (CulturalBench), value inference (WorldValuesBench), or single-axis bias (CDEval), but none assess how LLMs adjudicate when multiple cultural frameworks directly clash. We introduce CCD-Bench (Culture-Conflict Decision Benchmark), a benchmark for evaluating LLM decision-making under cross-cultural value conflict. CCD-Bench contains 2,182 open-ended dilemmas across seven domains, each with ten anonymized response options aligned with the ten GLOBE cultural clusters spanning 62 societies. Using a Stratified Latin Square design, we evaluate 17 leading LLMs and find clear biases: models favor Nordic Europe (20.2%) and Germanic Europe (12.4%), while Eastern Europe and Middle East & North Africa responses are least preferred (≈5–6%). Although 87.9% of model rationales reference multiple cultural dimensions, this pluralism is shallow, dominated by Future and Performance Orientation, with limited attention to Assertiveness or Gender Egalitarianism (
Hasibur Rahman, Hanan Salam
AAAI2
2025 Decompose-ToM: Enhancing Theory of Mind Reasoning in Large Language Models through Simulation and Task Decomposition
abstract
Theory of Mind (ToM) is the ability to understand and reflect on the mental states of others. Although this capability is crucial for human interaction, testing on Large Language Models (LLMs) reveals that they possess only a rudimentary understanding of it. Although the most capable closed-source LLMs have come close to human performance on some ToM tasks, they still perform poorly on complex variations of the task that involve more structured reasoning. In this work, we utilize the concept of “pretend-play”, or “Simulation Theory” from cognitive psychology to propose “Decompose-ToM”: an LLM-based inference algorithm that improves model performance on complex ToM tasks. We recursively simulate user perspectives and decompose the ToM task into a simpler set of tasks: subject identification, question-reframing, world model updation, and knowledge availability. We test the algorithm on higher-order ToM tasks and a task testing for ToM capabilities in a conversational setting, demonstrating that our approach shows significant improvement across models compared to baseline methods while requiring minimal prompt tuning across tasks and no additional model training. Our code is publicly available.
Sneheel Sarangi, Maha Elgarf, Hanan Salam
COLING3
2025 DynamicNER: A Dynamic, Multilingual, and Fine-Grained Dataset for LLM-based Named Entity Recognition
abstract
Hanjun Luo, Yingbin Jin, Yiran Wang, Xinfeng Li, Tong Shang, Xuecheng Liu, Ruizhe Chen, Kun Wang, Hanan Salam, Qingsong Wen, Zuozhu Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Hanjun Luo, Yingbin Jin, Xinfeng Li, Tong Shang, Xuecheng Liu, Ruizhe Chen, Kun Wang 0056, Hanan Salam, Qingsong Wen, Zuozhu Liu
EMNLP9
2025 A Study Companion for Productivity: Exploring the Role of a Social Robot for College Students with ADHD
abstract
Attention-Deficit Hyperactivity Disorder (ADHD) significantly impacts the academic, social, and mental well-being of young adults. While Socially Assistive Robots have shown promise in supporting individuals with ADHD, most existing systems focus on children and lack productivity tools for college students. To address this gap, we developed Alex, a study companion that assists with task prioritization, scheduling, and maintaining focus during work sessions. In a study involving 15 university students self-reporting ADHD symptoms, Alex's support alleviated overwhelm and improved focus. Notably, 12 participants expressed interest in using it again, demonstrating its potential to enhance productivity.
Himanshi Lalwani, Mira Saleh, Hanan Salam
HRI3
2025 AgentAuditor: Human-level Safety and Security Evaluation for LLM Agents
abstract
Despite the rapid advancement of LLM-based agents, the reliable evaluation of their safety and security remains a significant challenge. Existing rule-based or LLM-based evaluators often miss dangers in agents' step-by-step actions, overlook subtle meanings, fail to see how small issues compound, and get confused by unclear safety or security rules. To overcome this evaluation crisis, we introduce AgentAuditor, a universal, training-free, memory-augmented reasoning framework that empowers LLM evaluators to emulate human expert evaluators. AgentAuditor constructs an experiential memory by having an LLM adaptively extract structured semantic features (e.g., scenario, risk, behavior) and generate associated chain-of-thought reasoning traces for past interactions. A multi-stage, context-aware retrieval-augmented generation process then dynamically retrieves the most relevant reasoning experiences to guide the LLM evaluator's assessment of new cases. Moreover, we developed ASSEBench, the first benchmark designed to check how well LLM-based evaluators can spot both safety risks and security threats. ASSEBench comprises 2293 meticulously annotated interaction records, covering 15 risk types across 29 application scenarios. A key feature of ASSEBench is its nuanced approach to ambiguous risk situations, employing "Strict" and "Lenient" judgment standards. Experiments demonstrate that AgentAuditor not only consistently improves the evaluation performance of LLMs across all benchmarks but also sets a new state-of-the-art in LLM-as-a-judge for agent safety and security, achieving human-level accuracy. Our work is openly accessible at https://github.com/Astarojth/AgentAuditor-ASSEBench.
Hanjun Luo, Shenyu Dai, Chiming Ni, Xinfeng Li, Guibin Zhang, Kun Wang 0056, Tongliang Liu, Hanan Salam
NeurIPS8
2025 Supporting Productivity Skill Development in College Students through Social Robot Coaching: A Proof-of-Concept
abstract
College students often face academic challenges that hamper their productivity and well-being. Although self-help books and productivity apps are popular, they often fall short. Books provide generalized, non-interactive guidance, and apps are not inherently educational and can hinder the development of key organizational skills. Traditional productivity coaching offers personalized support, but is resource-intensive and difficult to scale. In this study, we present a proof-of-concept for a socially assistive robot (SAR) as an educational coach and a potential solution to the limitations of existing productivity tools and coaching approaches. The SAR delivers six different lessons on time management and task prioritization. Users interact via a chat interface, while the SAR responds through speech (with a toggle option). An integrated dashboard monitors progress, mood, engagement, confidence per lesson, and time spent per lesson. It also offers personalized productivity insights to foster reflection and self-awareness. We evaluated the system with 15 college students, achieving a System Usability Score of 79.2 and high ratings for overall experience and engagement. Our findings suggest that SAR-based productivity coaching can offer an effective and scalable solution to improve productivity among college students.
Himanshi Lalwani, Hanan Salam
RO-MAN2
2025 Ethically-Aware Participatory Design of a Productivity Social Robot for College Students
abstract
College students often face academic and life stressors affecting productivity, especially students with Attention Deficit Hyperactivity Disorder (ADHD) who experience executive functioning challenges. Conventional productivity tools typically demand sustained self-discipline and consistent use, which many students struggle with, leading to disruptive app-switching behaviors. Socially Assistive Robots (SARs), known for their intuitive and interactive nature, offer promising potential to support productivity in academic environments, having been successfully utilized in domains like education, cognitive development, and mental health. To leverage SARs effectively in addressing student productivity, this study employed a Participatory Design (PD) approach, directly involving college students and a Student Success and Well-Being Coach in the design process. Through interviews and a collaborative workshop, we gathered detailed insights on productivity challenges and identified desirable features for a productivity-focused SAR. Importantly, ethical considerations were integrated from the onset, facilitating responsible and user-aligned design choices. Our contributions include comprehensive insights into student productivity challenges, SAR design preferences, and actionable recommendations for effective robot characteristics. Additionally, we present stakeholder-derived ethical guidelines to inform responsible future implementations of productivity-focused SARs in higher education.
Himanshi Lalwani, Hanan Salam
RO-MAN2
2024 Towards Generalised and Incremental Bias Mitigation in Personality Computing
abstract
Building systems for predicting human socio-emotional states has promising applications; however, if trained on biased data, such systems could inadvertently yield biased decisions. Bias mitigation remains an open problem, which tackles the correction of a model's disparate performance over different groups defined by particular sensitive attributes (e.g., gender, age, and race). In this work, we design a novel fairness loss function named Multi-Group Parity (MGP) to provide a generalised approach for bias mitigation in personality computing. In contrast to existing works in the literature, MGP is generalised as it features four ‘multiple’ properties (4Mul): multiple tasks, multiple modalities, multiple sensitive attributes, and multi-valued attributes. Moreover, we explore how to incrementally mitigate the biases when more sensitive attributes are taken into consideration sequentially. Towards this problem, we introduce a novel algorithm that utilises an incremental learning framework to mitigate bias against one attribute data at a time without compromising past fairness. Extensive experiments on two large-scale multi-modal personality recognition datasets validate the effectiveness of our approach in achieving superior bias mitigation under the proposed four properties and incremental debiasing settings.
Jian Jiang 0001, Viswonathan Manoranjan, Hanan Salam, Oya Çeliktutan
IEEE Trans. Affect. Comput.3
2024 Automatic Context-Aware Inference of Engagement in HMI: A Survey
abstract
Engagement is the process by which participants establish, maintain, and end their perceived connection. Automatic engagement inference is one of the tasks required to develop successful human-centered HMI applications. Engagement is a multi-faceted multimodal construct requiring high accuracy in interpretating contextual, verbal and non-verbal cues, making the development of an intelligent automated engagement inference system challenging. Existing surveys concentrate on specific application settings, and a comprehensive survey covering the different engagement facets, definition and inference across various contexts is lacking. Moreover, despite the importance of context-aware modeling, the literature lacks a systematic context-aware overview on the topic. This paper presents a comprehensive survey on previous work in engagement for HMI, entailing interdisciplinary definition, engagement components, publicly available datasets, ground truth assessment, and commonly used features and methods, serving as a guide for the development of future HMI interfaces with reliable context-aware engagement inference capability. An in-depth review across embodied and disembodied interaction modes, and an emphasis on the interaction context of which engagement is studied sets apart this survey from existing ones. Our findings suggest four important directions for future research: (1) context-aware computational modeling, (2) temporal dynamics, (3) personalised computing, and (4) bias and fairness of engagement inference systems.
Hanan Salam, Oya Çeliktutan, Hatice Gunes, Mohamed Chetouani
IEEE Trans. Affect. Comput.1
2023 A holistic AI-based approach for pharmacovigilance optimization from patients behavior on social media
Valentin Roche, Jean-Philippe Robert, Hanan Salam
Artif. Intell. Medicine3
2022 Personalized Productive Engagement Recognition in Robot-Mediated Collaborative Learning
abstract
In this paper, we propose and compare personalized models for Productive Engagement (PE) recognition. PE is defined as the level of engagement that maximizes learning. Previously, in the context of robot-mediated collaborative learning, a framework of productive engagement was developed by utilizing multimodal data of 32 dyads and learning profiles, namely, Expressive Explorers (EE), Calm Tinkerers (CT), and Silent Wanderers (SW) were identified which categorize learners according to their learning gain. Within the same framework, a PE score was constructed in a non-supervised manner for real-time evaluation. Here, we use these profiles and the PE score within an AutoML deep learning framework to personalize PE models. We investigate two approaches for this purpose: (1) Single-task Deep Neural Architecture Search (ST-NAS), and (2) Multitask NAS (MT-NAS). In the former approach, personalized models for each learner profile are learned from multimodal features and compared to non-personalized models. In the MT-NAS approach, we investigate whether jointly classifying the learners’ profiles with the engagement score through multi-task learning would serve as an implicit personalization of PE. Moreover, we compare the predictive power of two types of features: incremental and non-incremental features. Non-incremental features correspond to features computed from the participant’s behaviours in fixed time windows. Incremental features are computed by accounting to the behaviour from the beginning of the learning activity till the time window where productive engagement is observed. Our experimental results show that (1) personalized models improve the recognition performance with respect to non-personalized models when training models for the gainer vs. non-gainer groups, (2) multitask NAS (implicit personalization) also outperforms non-personalized models, (3) the speech modality has high contribution towards prediction, and (4) non-incremental features outperform the incremental ones overall.
Vetha Vikashini Chithrra Raghuram, Hanan Salam, Jauwairia Nasir, Barbara Bruno, Oya Çeliktutan
ICMI2
2021 Identification of Signs of Depression Relapse using Audio-visual Cues: A Preliminary Study
abstract
Depression is a serious mental disorder that affects many individuals across the globe. Depression (unipolar or bipolar) is characterized by a high rate of relapse or recurrence where a person might experience depressive episodes after non-depressive ones. The symptom patterns for recurrent depressive episodes have not been properly analyzed. Thus, there is a pressing need for systems which can monitor the mental health of individuals at risk to detect initial signs of relapse and recurrence. This points towards an automated system which identifies such signs and facilitates in timely treatment. In this paper, we introduce for the first time a deep learning based prospective monitoring system for the identification of relapse signs using audio-visual cues. The proposed model approximates relapse as the similarity between non-depression and depression samples. Experiments were performed on the DAIC-WOZ dataset and a highest accuracy of 73.21% was obtained using a Siamese network-based approach with one-shot learning regime.
Muhammad Muzammel, Alice Othmani, Himadri Mukherjee, Hanan Salam
CBMS4
2021 Towards Automatic Narrative Coherence Prediction
abstract
Research in Psychology has shown that stories people tell about themselves, and how they recall their experiences, reveal a lot about their individual characteristics and mental well-being. The Narrative Coherence Coding Scheme (NaCCS) is a set of guidelines established in psychology research for annotating the “coherence” of a narrative along three dimensions: context, chronology and theme. A significant correlation was found between a narrative’s coherence score and independently collected mental health markers of the narrator. Currently, all coherence annotations are done manually; a time consuming task which drains vital resources. In this paper, we propose an Artificial Intelligence based approach involving Natural Language Processing (NLP) to predict a narrative’s coherence score (4-class classification problem). We explore a number of techniques, ranging from traditional machine learning models such as Support Vector Machines (SVM) to pre-trained language models such as BERT (Bidirectional Encoder Representations from Transformers). BERT produced the best results for all dimensions in terms of accuracy: 53.7% (context), 71.8% (chronology), and 69.6% (theme). The location of information in the narratives (beginning, end, throughout) was helpful in improving predictions.
Filip Bendevski, Jumana Ibrahim, Tina Krulec, Theodore Waters, Nizar Habash, Hanan Salam, Himadri Mukherjee, Christin Camia
ICMI6
2021 Lung Health Analysis: Adventitious Respiratory Sound Classification Using Filterbank Energies
abstract
Audio-based healthcare technologies are among the most significant applications of pattern recognition and Artificial Intelligence. Lately, a major chunk of the World population has been infected with serious respiratory diseases such as COVID-19. Early recognition of lung health abnormalities can facilitate early intervention, and decrease the mortality rate of the infected population. Research has shown that it is possible to automatically monitor lung health abnormalities through respiratory sounds. In this paper, we propose an approach that employs filter bank energy-based features and Random Forests to classify lung problem types from respiratory sounds. The adventitious sounds, crackles and wheezes appear distinct to the human ear. Moreover, different sounds are characterized by different frequency ranges that are dominant. The proposed approach attempts to distinguish the adventitious sounds (crackles and wheezes) by modeling the human auditory perception of these sounds. Specifically, we propose a respiratory sounds representation technique capable of modeling the dominant frequency range present in such sounds. On a publicly available dataset (ICBHI) of size 6898 cycles spanning over 5[Formula: see text]h, our results can be compared with the state-of-the-art results, in distinguishing two different types of adventitious sounds: crackles and wheezes.
Himadri Mukherjee, Hanan Salam, KC Santosh
Int. J. Pattern Recognit. Artif. Intell.2
2021 Multi-facial patches aggregation network for facial expression recognition and facial regions contributions to emotion display
Ahmed Rachid Hazourli, Amine Djeghri, Hanan Salam, Alice Othmani
Multim. Tools Appl.3
2018 A survey on face modeling: building a bridge between face analysis and synthesis
Hanan Salam, Renaud Séguier
Vis. Comput.1
2015 Engagement detection based on mutli-party cues for human robot interaction
abstract
In this paper, we address the problematic of automatic detection of engagement in multi-party Human-Robot Interaction scenarios. The aim is to investigate to what extent are we able to infer the engagement of one of the entities of a group based solely on the cues of the other entities present in the interaction. In a scenario featuring 3 entities: 2 participants and a robot, we extract behavioural cues that concern each of the entities, we then build models based solely on each of these entities' cues and on combinations of them to predict the engagement level of each of the participants. Person-level cross validation shows that we are capable of detecting the engagement of the participant in question using solely the behavioural cues of the robot with a high accuracy compared to using the participant's cues himself (75.91% vs. 74.32%). Moreover using the behavioural cues of the other participant is also informative where it permits the detection of the engagement of the participant in question at an accuracy of 62.15% on average. The correlation between the features of the other participant with the engagement labels of the participant in question suggests a high cohesion between the two participants. In addition, the similarity of the most significantly correlated features among the two participants suggests a high synchrony between the two parties.
Hanan Salam, Mohamed Chetouani
ACII1
2012 A multi-texture approach for estimating iris positions in the eye using 2.5D Active Appearance Models
abstract
This paper describes a new approach for the detection of the iris center. Starting from a learning base that only contains people in frontal view and looking in front of them, our model (based on 2.5D Active Appearance Models (AAM)) is capable of capturing the iris movements for both people in frontal view and with different head poses. We merge an iris model and a local eye model where holes are put in the place of the white-iris region. The iris texture slides under the eye hole permitting to synthesize and thus analyze any gaze direction. We propose a multi-objective optimization technique to deal with large head poses. We compared our method to a 2.5D AAM trained on faces with different gaze directions and showed that our proposition outperforms it in robustness and accuracy of detection specifically when head pose varies and with subjects wearing eyeglasses.
Hanan Salam, Nicolas Stoiber, Renaud Séguier
ICIP1
2012 A multimodal fuzzy inference system using a continuous facial expression representation for emotion detection
abstract
This paper presents a multimodal fuzzy inference system for emotion detection. The system extracts and merges visual, acoustic and context relevant features. The experiments have been performed as part of the AVEC 2012 challenge. Facial expressions play an important role in emotion detection. However, having an automatic system to detect facial emotional expressions on unknown subjects is still a challenging problem. Here, we propose a method that adapts to the morphology of the subject and that is based on an invariant representation of facial expressions. Our method relies on 8 key expressions of emotions of the subject. In our system, each image of a video sequence is defined by its relative position to these 8 expressions. These 8 expressions are synthesized for each subject from plausible distortions learnt on other subjects and transferred on the neutral face of the subject. Expression recognition in a video sequence is performed in this space with a basic intensity-area detector. The emotion is described in the 4 dimensions: valence, arousal, power and expectancy. The results show that the duration of high intensity smile is an expression that is meaningful for continuous valence detection and can also be used to improve arousal detection. The main variations in power and expectancy are given by context data.
Catherine Soladié, Hanan Salam, Catherine Pelachaud, Nicolas Stoiber, Renaud Séguier
ICMI2
2012 Facial Action Recognition Combining Heterogeneous Features via Multikernel Learning
abstract
This paper presents our response to the first international challenge on facial emotion recognition and analysis. We propose to combine different types of features to automatically detect action units (AUs) in facial images. We use one multikernel support vector machine (SVM) for each AU we want to detect. The first kernel matrix is computed using local Gabor binary pattern histograms and a histogram intersection kernel. The second kernel matrix is computed from active appearance model coefficients and a radial basis function kernel. During the training step, we combine these two types of features using the recently proposed SimpleMKL algorithm. SVM outputs are then averaged to exploit temporal information in the sequence. To evaluate our system, we perform deep experimentation on several key issues: influence of features and kernel function in histogram-based SVM approaches, influence of spatially independent information versus geometric local appearance information and benefits of combining both, sensitivity to training data, and interest of temporal context adaptation. We also compare our results with those of the other participants and try to explain why our method had the best performance during the facial expression recognition and analysis challenge.
Thibaud Senechal, Vincent Rapp, Hanan Salam, Renaud Séguier, Kevin Bailly, Lionel Prevost
IEEE Trans. Syst. Man Cybern. Part B3
2011 Combining AAM coefficients with LGBP histograms in the multi-kernel SVM framework to detect facial action units
abstract
This study presents a combination of geometric and appearance features used to automatically detect Action Units in face images. We use one multi-kernel SVM for each Action Unit we want to detect. The first kernel matrix is computed using Local Gabor Binary Pattern (LGBP) histograms and a histogram intersection kernel. The second kernel matrix is computed from AAM coefficients and a RBF kernel. During the training step, we combine these two type s of features using the recent SimpleMKL algorithm. SVM outputs are then filtered to exploit dynamic relationships between Action Units.
Thibaud Senechal, Vincent Rapp, Hanan Salam, Renaud Séguier, Kevin Bailly, Lionel Prevost
FG3