VLDB 2026 Research / reviewers in the wild / expert
Mohamed Chetouani
dblp:70/2508
· DBLP profile ↗
73ranked-venue papers
8as first author
19since 2021 · last 2026
0000-0002-2920-4539ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 4 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 31 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 6 since 2021Systems, architecture and hardware · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Inferring Implicit Goals Across Differing Task ModelsabstractOne of the significant challenges to generating value-aligned behavior is to not only account for the specified user objectives but also any implicit or unspecified user requirements. The existence of such implicit requirements could be particularly common in settings where the user's understanding of the task model may differ from the agent's estimate of the model. Under this scenario, the user may incorrectly expect some agent behavior to be inevitable or guaranteed. This paper addresses such expectation mismatch in the presence of differing models by capturing the possibility of unspecified user subgoal in the context of a task captured as a Markov Decision Process (MDP) and querying for it as required. Our method identifies bottleneck states and uses them as candidates for potential implicit subgoals. We then introduce a querying strategy that will generate the minimal number of queries required to identify a policy guaranteed to achieve the underlying goal. Our empirical evaluations demonstrate the effectiveness of our approach in inferring and achieving unstated goals across various tasks. Silvia Tulli, Stylianos Loukas Vasileiou, Mohamed Chetouani, Sarath Sreedharan |
AAAI | 3 |
| 2026 | Human-Interactive Robot Learning: Definition, Challenges, and RecommendationsabstractRobot learning from humans has been proposed and researched for several decades as a means to enable robots to learn new skills or adapt existing ones to new situations. Recent advances in AI, including learning approaches like reinforcement learning and architectures like transformers and foundation models, combined with access to massive datasets, have created attractive opportunities to apply those data-hungry techniques to this problem. We argue that the focus on massive amounts of pre-collected data, and the resulting learning paradigm, where humans demonstrate and robots learn in isolation, is overshadowing a specialized area of work we term Human-Interactive Robot Learning (HIRL). This paradigm, wherein robots and humans interact during the learning process , is at the intersection of multiple fields (AI, robotics, human–computer interaction, design and others) and holds unique promise. Using HIRL, robots can achieve greater sample efficiency (as humans can provide task knowledge through interaction), align with human preferences (as humans can guide the robot behavior toward their expectations), and explore more meaningfully and safely (as humans can utilize domain knowledge to guide learning and prevent catastrophic failures). This can result in robotic systems that can more quickly and easily adapt to new tasks in human environments. The objective of this article is to provide a broad and consistent overview of HIRL research and to guide researchers toward understanding the scope of HIRL, and current open or underexplored challenges related to four themes—namely, human, robot learning, interaction, and broader context. The article includes concrete use cases to illustrate the interaction between these challenges and inspire further research according to broad recommendations and a call for action for the growing HIRL community. Kim Baraka, Ifrah Idrees, Taylor Kessler Faulkner, Erdem Biyik, Serena Booth, Mohamed Chetouani, Daniel H. Grollman, Akanksha Saran, Emmanuel Senft, Silvia Tulli, Anna-Lisa Vollmer, Antonio Andriella, Helen Beierling, Tiffany Horter, Jens Kober, Isaac S. Sheidlower, Matthew E. Taylor, Sanne van Waveren, Xuesu Xiao |
ACM Trans. Hum. Robot Interact. | 6 |
| 2025 | Demographic User Modeling for Social Robotics with Multimodal Pre-trained ModelsabstractInternational audience Hamed Rahimi, Mouad Abrini, Jeanne Malecot, Ying Lai, Adrien Jacquet Crétides, Mahdi Khoramshahi, Mohamed Chetouani |
ICMI | 7 |
| 2025 | USER-VLM 360: Personalized Vision Language Models with User-aware Tuning for Social Human-Robot InteractionsabstractInternational audience Hamed Rahimi, Adil Bahaj, Mouad Abrini, Mahdi Khoramshahi, Mounir Ghogho, Mohamed Chetouani |
ICMI | 6 |
| 2025 | Task-Aware Robotic Grasping by evaluating Quality Diversity Solutions through Foundation ModelsabstractTask-aware robotic grasping is a challenging problem that requires the integration of semantic understanding and geometric reasoning. This paper proposes a novel framework that leverages Large Language Models (LLMs) and Quality Diversity (QD) algorithms to enable zero-shot task-conditioned grasp synthesis. The framework segments objects into meaningful subparts and labels each subpart semantically, creating structured representations that can be used to prompt an LLM. By coupling semantic and geometric representations of an object’s structure, the LLM’s knowledge about tasks and which parts to grasp can be applied in the physical world. The QD-generated grasp archive provides a diverse set of grasps, allowing us to select the most suitable grasp based on the task. We evaluated the proposed method on a subset of the YCB dataset with a Franka Emika robot. A consolidated ground truth for task-specific grasp regions is established through a survey. Our work achieves a weighted intersection over union (IoU) of 73.6% in predicting task-conditioned grasp regions in 65 task-object combinations. An end-to-end validation study on a smaller subset further confirms the effectiveness of our approach, with 88% of responses favoring the task-aware grasp over the control group. A binomial test shows that participants significantly prefer the task-aware grasp. Aurel Appius, Émiland Garrabé, François Hélénon, Mahdi Khoramshahi, Mohamed Chetouani, Stéphane Doncieux |
IROS | 5 |
| 2025 | Reasoning LLMs for User-Aware Multimodal Conversational AgentsabstractPersonalization in social robotics is critical for fostering effective human-robot interactions, yet systems often face the cold start problem, where initial user preferences or characteristics are unavailable. This paper proposes a novel framework called USER-LLM R1 for a user-aware conversational agent that addresses this challenge through dynamic user profiling and model initiation. Our approach integrates chain-of-thought (CoT) reasoning models to iteratively infer user preferences and vision-language models (VLMs) to initialize user profiles from multimodal inputs, enabling personalized interactions from the first encounter. Leveraging a Retrieval-Augmented Generation (RAG) architecture, the system dynamically refines user representations within an inherent CoT process, ensuring contextually relevant and adaptive responses. Evaluations on the ElderlyTech-Vqa Bench demonstrate significant improvements in ROUGE-1 (+23.2%) ROUGE-2 (+0.6%) and ROUGE-L (+8%) F1 scores over state-of-the-art baselines, with ablation studies underscoring the impact of reasoning model size on performance. Human evaluations further validate the framework’s efficacy, particularly for elderly users, where tailored responses enhance engagement and trust. Ethical considerations, including privacy preservation and bias mitigation, are rigorously discussed and addressed to ensure responsible deployment. Hamed Rahimi, Jeanne Cattoni, Meriem Beghili, Mouad Abrini, Mahdi Khoramshahi, Maribel Pino, Mohamed Chetouani |
RO-MAN | 7 |
| 2024 | Legibot: Generating Legible Motions for Service Robots Using Cost-Based Local Plannersabstracthyperref With the increasing presence of social robots in various environments and applications, there is an increasing need for these robots to exhibit socially-compliant behaviors. Legible motion, characterized by the ability of a robot to clearly and quickly convey intentions and goals to the individuals in its vicinity, through its motion, holds significant importance in this context. This will improve the overall user experience and acceptance of robots in human environments. In this paper, we introduce a novel approach to incorporate legibility into local motion planning for mobile robots. This can enable robots to generate legible motions in real-time and dynamic environments. To demonstrate the effectiveness of our proposed methodology, we also provide a robotic stack designed for deploying legibility-aware motion planning in a social robot, by integrating perception and localization components.The code and the data used in this work are available at https://legibot.github.io. Javad Amirian, Mouad Abrini, Mohamed Chetouani |
RO-MAN | 3 |
| 2024 | Modeling the Interplay Between Cohesion Dimensions: A Challenge for Group Affective Emergent StatesabstractEmergent states are temporal group phenomena that arise from collective affective, behavioral, and cognitive processes shared among the group's members during their interactions. Cohesion is one such state, mainly conceptualized by scholars as affective in nature, and frequently distinguished into the two dimensions social and task cohesion. Whereas social cohesion is related to the need of belonging to a group, task cohesion is related to the group's goals and tasks. In this article, we emphasize the importance of behavioral interaction dynamics to predict cohesion's dynamics. Drawing from Social Science insights, we investigate the interplay between social and task cohesion to predict their dynamics across group tasks from nonverbal behavioral features. Three computational architectures exploiting transfer learning are presented. Transfer learning capitalizes on information learnt by a model for a specific dimension to predict the dynamics of the other dimension. Results show that integrating the influence of social cohesion to predict dynamics of task cohesion outperforms state-of-the-art. To predict dynamics of social cohesion, a model integrating the reciprocal impact of social and task cohesion significantly improves performance with respect to the state-of-the-art and a model only integrating the impact of task cohesion on dynamics of social cohesion. Lucien Maman, Nale Lehmann-Willenbrock, Mohamed Chetouani, Laurence Likforman-Sulem, Giovanna Varni |
IEEE Trans. Affect. Comput. | 3 |
| 2024 | Automatic Context-Aware Inference of Engagement in HMI: A SurveyabstractEngagement is the process by which participants establish, maintain, and end their perceived connection. Automatic engagement inference is one of the tasks required to develop successful human-centered HMI applications. Engagement is a multi-faceted multimodal construct requiring high accuracy in interpretating contextual, verbal and non-verbal cues, making the development of an intelligent automated engagement inference system challenging. Existing surveys concentrate on specific application settings, and a comprehensive survey covering the different engagement facets, definition and inference across various contexts is lacking. Moreover, despite the importance of context-aware modeling, the literature lacks a systematic context-aware overview on the topic. This paper presents a comprehensive survey on previous work in engagement for HMI, entailing interdisciplinary definition, engagement components, publicly available datasets, ground truth assessment, and commonly used features and methods, serving as a guide for the development of future HMI interfaces with reliable context-aware engagement inference capability. An in-depth review across embodied and disembodied interaction modes, and an emphasis on the interaction context of which engagement is studied sets apart this survey from existing ones. Our findings suggest four important directions for future research: (1) context-aware computational modeling, (2) temporal dynamics, (3) personalised computing, and (4) bias and fairness of engagement inference systems. Hanan Salam, Oya Çeliktutan, Hatice Gunes, Mohamed Chetouani |
IEEE Trans. Affect. Comput. | 4 |
| 2023 | Humans' Spatial Perspective-Taking When Interacting with a Robotic ArmabstractPerceiving the environment from another person’s perspective, in other words, being in someone else’s shoes spatially, is not always an easy task. Perspective-taking can be even more challenging when working with a robot as a collaborator. The study reported here aims at investigating humans’ level 2 spatial perspective-taking performance when interacting with a collaborative robotic arm through a novel inperson experiment. First, a robotic arm drew ambiguous shapes on a whiteboard and participants had to answer questions that require performing spatial perspective-taking. A metric was used to compute a score based on their responses. Second, participants completed the PTSOT, a test measuring spatial orientation and perspective-taking ability. The results revealed a correlation between the scores computed using our metric and those obtained in the PTSOT. This suggests the efficiency of our new setup and associated evaluation metric in assessing spatial perspective-taking skills in a human-robot interaction context, as well as the validity of our findings, in line with prior studies on perspective-taking Mouad Abrini, Malika Auvray, Mohamed Chetouani |
RO-MAN | 3 |
| 2022 | Pragmatically Learning from Pedagogical Demonstrations in Multi-Goal EnvironmentsabstractLearning from demonstration methods usually leverage close to optimal demonstrations to accelerate training. By contrast, when demonstrating a task, human teachers deviate from optimal demonstrations and pedagogically modify their behavior by giving demonstrations that best disambiguate the goal they want to demonstrate. Analogously, human learners excel at pragmatically inferring the intent of the teacher, facilitating communication between the two agents. These mechanisms are critical in the few demonstrations regime, where inferring the goal is more difficult. In this paper, we implement pedagogy and pragmatism mechanisms by leveraging a Bayesian model of Goal Inference from demonstrations. We highlight the benefits of this model in multi-goal teacher-learner setups with two artificial agents that learn with goal-conditioned Reinforcement Learning. We show that combining BGI-agents (a pedagogical teacher and a pragmatic learner) results in faster learning and reduced goal ambiguity over standard learning from demonstrations, especially in the few demonstrations regime. Hugo Caselles-Dupré, Olivier Sigaud, Mohamed Chetouani |
NeurIPS | 3 |
| 2022 | Does what users say match what they do? Comparing self-reported attitudes and behaviours towards a social robotabstractConstructs intended to capture social attitudes and behaviour towards social robots are incredibly varied, with little overlap or consistency in how they may be related. In this study we conduct exploratory analyses between participants’ self-reported attitudes and behaviour towards a social robot. We designed an autonomous interaction where 102 participants interacted with a social robot (Pepper) in a hypothetical travel planning scenario, during which the robot displayed various multi-modal social behaviours. Several behavioural measures were embedded throughout the interaction, followed by a self-report questionnaire targeting participant’s social attitudes towards the robot (social trust, liking, rapport, competency trust, technology acceptance, mind perception, social presence, and social information processing). Several relationships were identified between participant’s behaviour and self-reported attitudes towards the robot. Implications for how to conceptualise and measure interactions with social robots are discussed. Rebecca Stower, Karen Tatarian, Damien Rudaz, Marine Chamoux, Mohamed Chetouani, Arvid Kappas |
RO-MAN | 5 |
| 2022 | SLOT-V: Supervised Learning of Observer Models for Legible Robot Motion Planning in ManipulationabstractWe present SLOT-V, a novel supervised learning framework that learns observer models (human preferences) from robot motion trajectories in a legibility context. Legibility measures how easily a (human) observer can infer the robot’s goal from a robot motion trajectory. When generating such trajectories, existing planners often rely on an observer model that estimates the quality of trajectory candidates. These observer models are frequently hand-crafted or, occasionally, learned from demonstrations. Here, we propose to learn them in a supervised manner using the same data format that is frequently used during the evaluation of aforementioned approaches. We then demonstrate the generality of SLOT-V using a Franka Emika in a simulated manipulation environment. For this, we show that it can learn to closely predict various hand-crafted observer models, i.e., that SLOT-V’s hypothesis space encompasses existing handcrafted models. Next, we showcase SLOT-V’s ability to generalize by showing that a trained model continues to perform well in environments with unseen goal configurations and/or goal counts. Finally, we benchmark SLOT-V’s sample efficiency (and performance) against an existing IRL approach and show that SLOT-V learns better observer models with less data. Combined, these results suggest that SLOT-V can learn viable observer models. Better observer models imply more legible trajectories, which may -in turn - lead to better and more transparent human-robot interaction. Sebastian Wallkötter, Mohamed Chetouani, Ginevra Castellano |
RO-MAN | 2 |
| 2021 | Using Valence Emotion to Predict Group Cohesion's Dynamics: Top-down and Bottom-up ApproachesabstractCohesion is an affective group phenomenon. It has received a lot of attention from scholars both in Social Sciences and in Affective Computing that showed that cohesion and emotion influence each other, highlighting the need to jointly analyze them. This study presents 2 deep neural network architectures grounded on multitask learning to jointly predict cohesion and emotion. Inspired by 2 major Social Sciences approaches on group emotion (i.e., Top-down and Bottom-up), these architectures exploit cohesion and emotion interdependencies intending to improve the prediction of the dynamics (i.e. changes over time) of the Social and Task dimensions of cohesion. Emotion, here, is addressed in terms of its valence. Both architectures are evaluated against the performances of a similar model that only predicts the dynamics of both the Social and Task dimensions of cohesion, without integrating valence. Statistical analysis shows that only the deep model implementing the Bottom-up approach significantly improved the predictions of the Task cohesion’s dynamics. This result confirms the theoretical and practical benefits of multitasking as it takes full advantage of the inherent relationships between group emotion and cohesion to improve Task cohesion’s predictions. Lucien Maman, Mohamed Chetouani, Laurence Likforman-Sulem, Giovanna Varni |
ACII | 2 |
| 2021 | Grounding Language to Autonomously-Acquired Skills via Goal Generation
Ahmed Akakzia, Cédric Colas, Pierre-Yves Oudeyer, Mohamed Chetouani, Olivier Sigaud |
ICLR | 4 |
| 2021 | Exploiting the Interplay between Social and Task Dimensions of Cohesion to Predict its Dynamics Leveraging Social SciencesabstractEmergent states are behavioral, cognitive and affective processes appearing among the members of a group when they interact together. In the last decade, the development of computational approaches received a growing interest in building Human-Centered systems. Such a development is particularly difficult because some of these states have several dimensions interplaying somehow and somewhere over time. In this paper, we focus on cohesion, its dimensions and their interplay. Several definitions of cohesion exist, it can be simply defined as the tendency of a group to stick together to pursue goals and/or affective needs. This plethora of definitions resulted in many different cohesion dimensions. Social and Task dimensions are the most investigated both in Social Sciences and Computer Science since they both play an important role in a wide range of contexts and groups. To the best of our knowledge, however, no previous work on the prediction of cohesion dynamics focused on how these 2 dimensions interplay. We leverage Social Sciences to address this issue. In particular, we take advantage of the importance of Social cohesion for creating flexible and constructive relationships to reinforce Task cohesion. We describe a Deep Neural Network architecture (DNN) for predicting the dynamics of Task cohesion by applying transfer learning from a pre-trained model dedicated to the prediction of Social cohesion dynamics. Our architecture is evaluated against several baselines. Results show that it significantly improves the predictions of the Task cohesion dynamics, confirming the benefits of integrating Social Sciences insights into models architectures. Lucien Maman, Laurence Likforman-Sulem, Mohamed Chetouani, Giovanna Varni |
ICMI | 3 |
| 2021 | Robot Gaze Behavior and Proxemics to Coordinate Conversational Roles in Group InteractionsabstractWith more social robots entering different industries such as educational systems, health-care facilities, and even airports, it is important to tackle problems that may hinder high quality interactions in a wild setting including group conversations. This paper presents an autonomous group conversational role coordinator system based on the proxemics of participants in the group with the robot including their distances and orientations. The system accordingly assigns to the group participants around the robot three different statuses: active, bystander, and overhearer. Once the statuses are estimated, the robot autonomously adjusts its gaze pattern in order to adapt to the group dynamics and attributes its attention in relation to the role the member in the group is playing. This system was evaluated through a pilot study (N=16), in which two participants at a time played a trivia game with the robot and had different roles to play within the interaction. The primary results imply that the participants interacting with a robot having this adaptive gaze behavior based on conversational role coordination are more likely to stand closer to the robot. In addition, the robot was perceived as more adaptable, sociable, and socially present as well as more likely to make the participants feel more attended to. Karen Tatarian, Marine Chamoux, Amit Kumar Pandey, Mohamed Chetouani |
RO-MAN | 4 |
| 2021 | Towards Transparent Robot Learning Through TDRL-Based Emotional ExpressionsabstractRobots and virtual agents need to adapt existing and learn novel behavior to function autonomously in our society. Robot learning is often in interaction with or in the vicinity of humans. As a result the learning process needs to be transparent to humans. Reinforcement Learning (RL) has been used successfully for robot task learning. However, this learning process is often not transparent to the users. This results in a lack of understanding of what the robot is trying to do and why. The lack of transparency will directly impact robot learning. The expression of emotion is used by humans and other animals to signal information about the internal state of the individual in a language-independent, and even species-independent way, also during learning and exploration. In this article we argue that simulation and subsequent expression of emotion should be used to make the learning process of robots more transparent. We propose that the TDRL Theory of Emotion gives sufficient structure on how to develop such an emotionally expressive learning robot. Finally, we argue that next to such a generic model of RL-based emotion simulation we need personalized emotion interpretation for robots to better cope with individual expressive differences of users. Joost Broekens, Mohamed Chetouani |
IEEE Trans. Affect. Comput. | 2 |
| 2021 | Explainable Embodied Agents Through Social Cues: A ReviewabstractThe issue of how to make embodied agents explainable has experienced a surge of interest over the past 3 years, and there are many terms that refer to this concept, such as transparency and legibility. One reason for this high variance in terminology is the unique array of social cues that embodied agents can access in contrast to that accessed by non-embodied agents. Another reason is that different authors use these terms in different ways. Hence, we review the existing literature on explainability and organize it by (1) providing an overview of existing definitions, (2) showing how explainability is implemented and how it exploits different social cues, and (3) showing how the impact of explainability is measured. Additionally, we present a list of open questions and challenges that highlight areas that require further investigation by the community. This provides the interested reader with an overview of the current state of the art. Sebastian Wallkötter, Silvia Tulli, Ginevra Castellano, Ana Paiva 0001, Mohamed Chetouani |
ACM Trans. Hum. Robot Interact. | 5 |
| 2020 | Touch Recognition with Attentive End-to-End ModelabstractTouch is the earliest sense to develop and the first mean of contact with the external world. Touch also plays a key role in our socio-emotional communication: we use it to communicate our feelings, elicit strong emotions in others and modulate behavior (e.g compliance). Although its relevance, touch is an understudied modality in Human-Machine-Interaction compared to audition and vision. Most of the social touch recognition systems require a feature engineering step making them difficult to compare and to generalize to other databases. In this paper, we propose an end-to-end approach. We present an attention-based end-to-end model for touch gesture recognition evaluated on two public datasets (CoST and HAART) in the context of the ICMI 15 Social Touch Challenge. Our model gave a similar level of accuracy: 61% for CoST and 68% for HAART and uses self-attention as an alternative to feature engineering and Recurrent Neural Networks. Wail El Bani, Mohamed Chetouani |
ICMI | 2 |
| 2020 | Integrating an Observer in Interactive Reinforcement Learning to Learn Legible TrajectoriesabstractAn important aspect of Human-Robot-cooperation is that the robot is capable of clearly communicating its intentions to its human collaborator. This communication of intentions often requires the generation of legible motion trajectories. The concept of legible motion is usually not studied together with machine learning. Studying these fields together is an important step towards better Human-Robot cooperation. In this paper, we investigate interactive robot learning approaches with the aim of developing models that are able to generate legible motions by taking observer feedback into account. We explore how to integrate the observer feedback into a Reinforcement Learning (RL) framework. We do this by proposing three different observer algorithms as observer strategies in an interactive RL scheme and compare with one non-interactive RL algorithm as baseline. For the observer strategies we vary the method how the observer estimates how likely the agent is going for the target goal. We evaluate our approach on five environments and calculate the legibility of the learned trajectories. The results show that the legibility of the learned trajectories is significantly higher while integrating the feedback from the observer compared with a standard Q-Learning algorithm not using the observer feedback. Manuel Bied, Mohamed Chetouani |
RO-MAN | 2 |
| 2020 | TIRL: Enriching Actor-Critic RL with non-expert human teachers and a Trust ModelabstractReinforcement learning (RL) algorithms have been demonstrated to be very attractive tools to train agents to achieve sequential tasks. However, these algorithms require too many training data to converge to be efficiently applied to physical robots. By using a human teacher, the learning process can be made faster and more robust, but the overall performance heavily depends on the quality and availability of teacher demonstrations or instructions. In particular, when these teaching signals are inadequate, the agent may fail to learn an optimal policy. In this paper, we introduce a trust-based interactive task learning approach. We propose an RL architecture able to learn both from environment rewards and from various sparse teaching signals provided by non-expert teachers, using an actor-critic agent, a human model and a trust model. We evaluate the performance of this architecture on 4 different setups using a maze environment with different simulated teachers and show that the benefits of the trust model. Félix Rutard, Olivier Sigaud, Mohamed Chetouani |
RO-MAN | 3 |
| 2020 | Interactively shaping robot behaviour with unlabeled human instructions
Anis Najar, Olivier Sigaud, Mohamed Chetouani |
Auton. Agents Multi Agent Syst. | 3 |
| 2020 | Computational Study of Primitive Emotional Contagion in Dyadic InteractionsabstractInterpersonal human-human interaction is a dynamical exchange and coordination of social signals, feelings and emotions usually performed through and across multiple modalities such as facial expressions, gestures, and language. Developing machines able to engage humans in rich and natural interpersonal interactions requires capturing such dynamics. This paper addresses primitive emotional contagion during dyadic interactions in which roles are prefixed. Primitive emotional contagion was defined as the tendency people have to automatically mimic and synchronize their multimodal behavior during interactions and, consequently, to emotionally converge. To capture emotional contagion, a cross-recurrence based methodology that explicitly integrates short and long-term temporal dynamics through the analysis of both facial expressions and sentiment was developed. This approach is employed to assess emotional contagion at unimodal, multimodal and cross-modal levels and is evaluated on the Solid SAL-SEMAINE corpus. Interestingly, the approach is able to show the importance of the adoption of cross-modal strategies for addressing emotional contagion. Giovanna Varni, Isabelle Hupont, Chloé Clavel, Mohamed Chetouani |
IEEE Trans. Affect. Comput. | 4 |
| 2019 | How unitizing affects annotation of cohesionabstractThis paper investigates how unitizing affects external observers' annotation of group cohesion. We compared unitizing techniques belonging to these categories: interval coding, continuous coding, and a technique inspired by a cognitive theory on event perception. We applied such techniques for sampling coding units from a set of recordings of social interactions rich in behaviors related to cohesion. Then, we compared the cohesion scores the observers assigned to each coding unit. Results show that the three techniques can lead to suitable ratings and that the technique inspired to cognitive theories leads to scores reflecting variability in cohesion better than the other ones. Eleonora Ceccaldi, Nale Lehmann-Willenbrock, Erica Volta, Mohamed Chetouani, Gualtiero Volpe, Giovanna Varni |
ACII | 4 |
| 2019 | CURIOUS: Intrinsically Motivated Modular Multi-Goal Reinforcement LearningabstractIn open-ended environments, autonomous learning agents must set their own goals and build their own curriculum through an intrinsically motivated exploration. They may consider a large diversity of goals, aiming to discover what is controllable in their environments, and what is not. Because some goals might prove easy and some impossible, agents must actively select which goal to practice at any moment, to maximize their overall mastery on the set of learnable goals. This paper proposes CURIOUS , an algorithm that leverages 1) a modular Universal Value Function Approximator with hindsight learning to achieve a diversity of goals of different kinds within a unique policy and 2) an automated curriculum learning mechanism that biases the attention of the agent towards goals maximizing the absolute learning progress. Agents focus sequentially on goals of increasing complexity, and focus back on goals that are being forgotten. Experiments conducted in a new modular-goal robotic environment show the resulting developmental self-organization of a learning curriculum, and demonstrate properties of robustness to distracting goals, forgetting and changes in body properties. Cédric Colas, Pierre-Yves Oudeyer, Olivier Sigaud, Pierre Fournier, Mohamed Chetouani |
ICML | 5 |
| 2019 | Affective and behavioural computing: Lessons learnt from the First Computational Paralinguistics Challenge
Björn W. Schuller, Felix Weninger, Yue Zhang 0014, Fabien Ringeval, Anton Batliner, Stefan Steidl, Florian Eyben, Erik Marchi, Alessandro Vinciarelli, Klaus R. Scherer, Mohamed Chetouani, Marcello Mortillaro |
Comput. Speech Lang. | 11 |
| 2019 | Region-based facial representation for real-time Action Units intensity detection across datasets
Isabelle Hupont, Mohamed Chetouani |
Pattern Anal. Appl. | 2 |
| 2019 | Quantifying patterns of joint attention during human-robot interactions: An application for autism spectrum disorder assessmentabstractHAL is a multi-disciplinary open access archive for the deposit and dissemination of scientific research documents, whether they are published or not.The documents may come from teaching and research institutions in France or abroad, or from public Salvatore Maria Anzalone, Jean Xavier, Sofiane Boucenna, Lucia Billeci, Antonio Narzisi, Filippo Muratori, Mohamed Chetouani |
Pattern Recognit. Lett. | 8 |
| 2018 | The Attribution of Emotional State - How Embodiment Features and Social Traits Affect the Perception of an Artificial AgentabstractUnderstanding emotional states is a challenging task which frequently leads to misinterpretation even in human observers. While the perception of emotions has been studied extensively in human psychology, little is known about what factors influence the human perception of emotions in robots and virtual characters. In this paper, we build on the Brunswik lens model to investigate the influence of (a) the agent's embodiment using a 2D virtual character, a 3D blended embodiment, a recording of the 3D platform and a recording of a human, as well as (b) the level of human-likeness on people's ability to interpret emotional facial expressions in an agent. In addition, we measure social traits of the human observers and analyze how they correlate to the success in recognizing emotional expressions. We find that interpersonal differences play a minor role in the perception of emotional states. However, both embodiment and human-likeness as well as related perceptual dimensions such as perceived social presence and uncanniness have an effect on the attribution of emotional states. Maike Paetzel-Prüsmann, Ginevra Castellano, Giovanna Varni, Isabelle Hupont, Mohamed Chetouani, Christopher Peters 0001 |
RO-MAN | 5 |
| 2018 | Multimodal Stress Detection from Multiple AssessmentsabstractStress is a complex phenomenon that impacts the body and the mind at several levels. It has been studied for more than a century from different perspectives, which result in different definitions and different ways to assess the presence of stress. This paper introduces a methodology for analyzing multimodal stress detection results by taking into account the variety of stress assessments. As a first step, we have collected video, depth and physiological data from 25 subjects in a stressful situation: a socially evaluated mental arithmetic test. As a second step, we have acquired three different assessments of stress: self-assessment, assessments from external observers and assessment from a physiology expert. Finally, we extract 101 behavioural and physiological features and evaluate their predictive power for the three collected assessments using a classification task. Using multimodal features, we obtain average F1 scores up to 0.85. By investigating the composition of the best selected feature subsets and the individual feature classification performances, we show that several features provide valuable information for the classification of the three assessments: features related to body movement, blood volume pulse and heart rate. From a methodological point of view, we argue that a multiple assessment approach provide more robust results. Jonathan Aigrain, Michel Spodenkiewicz, Séverine Dubuisson, Marcin Detyniecki, Mohamed Chetouani |
IEEE Trans. Affect. Comput. | 6 |
| 2017 | The influence of individual social traits on robot learning in a human-robot interactionabstractInteractive Machine Learning considers that a robot is learning with and/or from a human. In this paper, we investigate the impact of human social traits on the robot learning. We explore social traits such as age (children vs. adult) and pathology (typical developing children vs. children with autistic spectrum disorders). In particular, we consider learning to recognize both postures and identity of a human partner. A human-robot posture imitation learning, based on a neural network architecture, is used to develop a multi-task learning framework. This architecture exploits three learning levels : 1) visual feature representation, 2) posture classification and 3) human partner identification. During the experiment the robot interacts with children with autism spectrum disorders (ASD), typical developing children (TD) and healthy adults. Previous works assessed the impact on learning of these social traits at the group level. In this paper, we focus on the analysis of individuals separately. The results show that the robot is impacted by the social traits of these different groups' individuals. First, the architecture needs to learn more visual features when interacting with a child with ASD (compared to a TD child) or with a TD child (compared to an adult). However, this surplus in the number of neurons helped the robot to improve the TD children's posture recognition but not that of children with ASD. Second, preliminary results show that this need of a neurons surplus while interacting with children with ASD is also generalizable to the identity recognition task. Hakim Guedjou, Sofiane Boucenna, Jean Xavier, Mohamed Chetouani |
RO-MAN | 5 |
| 2017 | Investigating the influence of embodiment on facial mimicry in HRI using computer vision-based measuresabstractMimicry plays an important role in social interaction. In human communication, it is used to establish rapport and bonding both with other humans, as well as robots and virtual characters. However, little is known about the underlying factors that elicit mimicry in humans when interacting with a robot. In this work, we study the influence of embodiment on participants' ability to mimic a social character. Participants were asked to intentionally mimic the laughing behavior of the Furhat mixed embodied robotic head and a 2D virtual version of the same character. To explore the effect of embodiment, we present two novel approaches to automatically assess people's ability to mimic based solely on videos of their facial expressions. In contrast to participants' self-assessment, the analysis of video recordings suggests a better ability to mimic when people interact with the 2D embodiment. Maike Paetzel-Prüsmann, Giovanna Varni, Isabelle Hupont, Mohamed Chetouani, Christopher Peters 0001, Ginevra Castellano |
RO-MAN | 4 |
| 2017 | Semantic-based interaction for teaching robot behavior compositionsabstractAllowing humans to teach robot behaviors will facilitate acceptability as well as long-term interactions. Humans would mainly use speech to transfer knowledge or to teach highlevel behaviors. In this paper, we propose a proof-of-concept application allowing a Pepper robot to learn behaviors from their natural-language-based description, provided by naive human users. In our model, natural language input is provided by grammar-free speech recognition, and is then processed to produce semantic knowledge, grounded in language and primitive behaviors. The same semantic knowledge is used to represent any kind of perceived input as well as actions the robot can perform. The experiment shows that the system can work independently from the domain of application, but also that it has limitations. Progress in semantic extraction, behavior planning and interaction scenario could stretch these limits. Victor Paleologue, Jocelyn Martin, Amit Kumar Pandey, Miranda Coninx, Mohamed Chetouani |
RO-MAN | 5 |
| 2016 | On leveraging crowdsourced data for automatic perceived stress detectionabstractResorting to crowdsourcing platforms is a popular way to obtain annotations. Multiple potentially noisy answers can thus be aggregated to retrieve an underlying ground truth. However, it may be irrelevant to look for a unique ground truth when we ask crowd workers for opinions, notably when dealing with subjective phenomena such as stress. In this paper, we discuss how we can better use crowdsourced annotations with an application to automatic detection of perceived stress. Towards this aim, we first acquired video data from 44 subjects in a stressful situation and gathered answers to a binary question using a crowdsourcing platform. Then, we propose to integrate two measures derived from the set of gathered answers into the machine learning framework. First, we highlight that using the consensus level among crowd worker answers substantially increases classification accuracies. Then, we show that it is suitable to directly predict for each video the proportion of positive answers to the question from the different crowd workers. Hence, we propose a thorough study on how crowdsourced annotations can be used to enhance performance of classification and regression methods. Jonathan Aigrain, Arnaud Dapogny, Kevin Bailly, Séverine Dubuisson, Marcin Detyniecki, Mohamed Chetouani |
ICMI | 6 |
| 2016 | International workshop on social learning and multimodal interaction for designing artificial agents (workshop summary)abstractThe “social learning and multimodal interaction for designing artificial agents” workshop aims at presenting scientific and philosophical advances related to social learning and multimodal interaction for enhancing the design of artificial agents. Papers presented in the workshop include studies on human behavior modeling, on social robotics and on virtual agents. Our two invited speakers, Prof. Catherine Pelachaud and Prof. Louis-Philippe Morency will enrich and open the door to further discussion by bringing their widely acknowledged expertise in the field. Mohamed Chetouani, Salvatore Maria Anzalone, Giovanna Varni, Isabelle Hupont, Ginevra Castellano, Angelica Lim, Gentiane Venture |
ICMI | 1 |
| 2016 | ASSP4MI2016: 2nd international workshop on advancements in social signal processing for multimodal interaction (workshop summary)abstractThis paper gives a summary of the 2nd International Workshop on Advancements in Social Signal Processing for Multimodal Interaction (ASSP4MI). Following our successful 1st International Workshop on Advancements in Social Signal Processing for Multimodal Interaction, held during ICMI-2015, we proposed the 2nd ASSP4MI workshop during ICMI-2016. The topics addressed and discussions fostered during last year's workshop are considered very relevant and alive in the research community. In this year's workshop, we continued addressing important topics and fostering fruitful discussions among researchers from different disciplines working in the fields of Social Signal Processing (SSP) and multimodal interaction. Khiet P. Truong, Dirk Heylen, Toyoaki Nishida, Mohamed Chetouani |
ICMI | 4 |
| 2016 | Automatic Analysis of Typical and Atypical Encoding of Spontaneous Emotion in the Voice of ChildrenabstractInternational audience Fabien Ringeval, Erik Marchi, Charline Grossard, Jean Xavier, Mohamed Chetouani, Björn W. Schuller |
INTERSPEECH | 5 |
| 2016 | Seventh International Workshop on Human Behavior Understanding (HBU 2016)abstractWith advances in pattern recognition and multimedia computing, it becomes possible to analyze human behavior via multimodal sensors at varying time-scales, levels of analysis, and meaning. This ability opens up far-ranging possibilities for multimedia and multimodal interaction. Research has the, potential to endow computers with the capacity to detect and understand people's actions and activities and infer their attitudes, preferences, personality, and social relationships. This workshop brings together researchers in this rapidly emerging area and especially those concerned with behavior analysis and multimedia in children. Mohamed Chetouani, Jeffrey F. Cohn, Albert Ali Salah |
ACM Multimedia | 1 |
| 2016 | Training a robot with evaluative feedback and unlabeled guidance signalsabstractIn this paper, we present a new method for training a robot by natural interaction using evaluative feedback and unlabeled guidance signals. Feedback signals are directly mapped to reward values and used for learning both the task and the meaning of the guidance signals. The learned guidance signals are used in return to bootstrap task learning. We propose to use unlabeled guidance signals as an alternative solution to preprogrammed guidance. We evaluate our method both in simulation and on a real robot. Anis Najar, Olivier Sigaud, Mohamed Chetouani |
RO-MAN | 3 |
| 2016 | Robots learning how and where to approach peopleabstractRobot navigation in human environments has been in the eyes of researchers for the last few years. Robots operating under these circumstances have to take human awareness into consideration for safety and acceptance reasons. Nonetheless, navigation have been often treated as going towards a goal point or avoiding people, without considering the robot engaging a person or a group of people in order to interact with them. This paper presents two navigation approaches based on the use of inverse reinforcement learning (IRL) from exemplar situations. This allow us to implement two path planners that take into account social norms for navigation towards isolated people. For the first planner, we learn an appropriate way to approach a person in an open area without static obstacles, this information is used to generate robot's path plan. As for the second planner, we learn the weights of a linear combination of continuous functions that we use to generate a costmap for the approach-behavior. This costmap is then combined with others, e.g. a costmap with higher cost around obstacles, and finally a path is generated with Dijkstra's algorithm. Omar A. Islas Ramírez, Harmish Khambhaita, Raja Chatila 0001, Mohamed Chetouani, Rachid Alami 0001 |
RO-MAN | 4 |
| 2016 | Modeling the dynamics of individual behaviors for group detection in crowds using low-level featuresabstractThis paper introduces two novel algorithms for detecting groups of people standing or freely moving in a crowded environment. The proposed algorithms exploit low-level features extracted from videos. The first algorithm, the Link Method, uses a learning and forgetting strategy for modeling dynamics of proxemics between individuals. Two versions of this algorithm are proposed: they differ in the analysis of proxemics. The second one, called Interpersonal Synchrony Method, explicitly adopts interpersonal synchrony to refine clusters of persons detected by combining together proxemics and 2D field of view of individuals. The algorithms are evaluated on both simulated and real-world video sequences from state-of-the-art databases. Clustering metrics such as the Adjusted Mutual Information shows that our models outperform the approach based on F-formations. This work developed algorithms that can be readily applied in robotics, to allow robots to automatically detect groups in crowded environments. Omar A. Islas Ramírez, Giovanna Varni, Mihai Andries, Mohamed Chetouani, Raja Chatila 0001 |
RO-MAN | 4 |
| 2016 | Real-time facial action unit intensity prediction with regularized metric learning
Jérémie Nicolle, Kevin Bailly, Mohamed Chetouani |
Image Vis. Comput. | 3 |
| 2016 | Dynamics of Non-Verbal Vocalizations and Hormones during Father-Infant InteractionabstractAlthough researchers have established the roles of oxytocin (OT) in promoting affiliative bonds and cortisol (CT) in adapting to stress, the investigation of their interplay with non-verbal behaviors has only recently begun. In this study, we employed social signal-processing techniques to investigate relationships between non-verbal features: infant and father vocalizations, infant-directed speech, speech turn-taking (STT) and hormonal dynamics (OT and CT). Thirty-five fathers were asked to interact with their infants following the fathers self-administration of OT or placebo. We consider the three episodes of the Still Face (SF) paradigm: (1) a baseline normal interaction episode, (2) the SF episode, in which the father becomes unresponsive and maintains a neutral facial expression, and (3) a reunion in which parents and their infants re-engage in interaction. This paradigm elicited stress in the infant. Statistical relationships are assessed by correlation analysis and linear mixed models (LMMs). The results indicate that (i) infant vocalization and STT are key social cues regulating interactions during the stress-inducing and reunion episodes, with infant vocalization leading the interaction dynamics; (ii) father empty pause was the main adaptive behavior of fathers after SF; (iii) OT did not modulate infant STT or father STT/fatherese; (iv) CT appeared to modulate the interaction. Omri Weisman, Mohamed Chetouani, Catherine Saint-Georges, Nadege Bourvis, Emilie Delaherche, Orna Zagoory-Sharon, Ruth Feldman |
IEEE Trans. Affect. Comput. | 2 |
| 2015 | Engagement detection based on mutli-party cues for human robot interactionabstractIn this paper, we address the problematic of automatic detection of engagement in multi-party Human-Robot Interaction scenarios. The aim is to investigate to what extent are we able to infer the engagement of one of the entities of a group based solely on the cues of the other entities present in the interaction. In a scenario featuring 3 entities: 2 participants and a robot, we extract behavioural cues that concern each of the entities, we then build models based solely on each of these entities' cues and on combinations of them to predict the engagement level of each of the participants. Person-level cross validation shows that we are capable of detecting the engagement of the participant in question using solely the behavioural cues of the robot with a high accuracy compared to using the participant's cues himself (75.91% vs. 74.32%). Moreover using the behavioural cues of the other participant is also informative where it permits the detection of the engagement of the participant in question at an accuracy of 62.15% on average. The correlation between the features of the other participant with the engagement labels of the participant in question suggests a high cohesion between the two participants. In addition, the similarity of the most significantly correlated features among the two participants suggests a high synchrony between the two parties. Hanan Salam, Mohamed Chetouani |
ACII | 2 |
| 2015 | Automatic measure of imitation during social interaction: A behavioral and hyperscanning-EEG benchmark
Emilie Delaherche, Guillaume Dumas, Jacqueline Nadel, Mohamed Chetouani |
Pattern Recognit. Lett. | 4 |
| 2014 | Potential human reaction aware mobile robot motion planner: Potential cost minimization frameworkabstractRobots have been stepping steadily into our daily-life. As humans are likely to exist in daily-life scenarios, “human-aware” becomes an essential point for robot navigation in these scenarios. Existing research works on human aware robot navigation rarely take potential human reaction into account in the motion planner. In contrast, we propose in this paper the concept of potential human reaction aware motion planner (PHRAMP), i.e. the robot had better take potential human reaction into account in its motion planning. We propose a general methodology coined as the potential cost minimization (PCM) framework for realizing PHRAMPs. We present some theoretical characteristics of this framework and explain its relationship with existing methods. We describe an instantiation of the PCM framework which serves as a concrete example PHRAMP to demonstrate the utility and advantage of the PCM framework. Omar A. Islas Ramírez, Mohamed Chetouani |
RO-MAN | 3 |
| 2013 | Locating facial landmarks with binary map cross-correlationsabstractPrecise facial landmark localization in still images is a key step for many face analysis applications, such as biometrics or automatic emotion recognition. In this paper, we propose a framework for facial point detection in frontal and near-frontal images. We introduce a new appearance model based on binary map cross-correlations that efficiently uses LBP and LPQ in a localization context. Inclusion of shape-related constraints is performed by a nonparametric voting method using relational properties within triplets of points, designed to correct outliers without losing precision for accurately detected points. We tested our system's performance on the widely used as benchmark BioID database obtaining state-of-the-art results. We also discuss evaluation metrics used to compare facial landmarking systems and which have been mixed up in recent literature. Jérémie Nicolle, Kevin Bailly, Vincent Rapp, Mohamed Chetouani |
ICIP | 4 |
| 2013 | The INTERSPEECH 2013 computational paralinguistics challenge: social signals, conflict, emotion, autismabstractInternational audience Björn W. Schuller, Stefan Steidl, Anton Batliner, Alessandro Vinciarelli, Klaus R. Scherer, Fabien Ringeval, Mohamed Chetouani, Felix Weninger, Florian Eyben, Erik Marchi, Marcello Mortillaro, Hugues Salamin, Anna Polychroniou, Fabio Valente, Samuel Kim |
INTERSPEECH | 7 |
| 2012 | Robust continuous prediction of human emotions using multiscale dynamic cuesabstractDesigning systems able to interact with humans in a natural manner is a complex and far from solved problem. A key aspect of natural interaction is the ability to understand and appropriately respond to human emotions. This paper details our response to the Audio/Visual Emotion Challenge (AVEC'12) whose goal is to continuously predict four affective signals describing human emotions (namely valence, arousal, expectancy and power). The proposed method uses log-magnitude Fourier spectra to extract multiscale dynamic descriptions of signals characterizing global and local face appearance as well as head movements and voice. We perform a kernel regression with very few representative samples selected via a supervised weighted-distance-based clustering, that leads to a high generalization power. For selecting features, we introduce a new correlation-based measure that takes into account a possible delay between the labels and the data and significantly increases robustness. We also propose a particularly fast regressor-level fusion framework to merge systems based on different modalities. Experiments have proven the efficiency of each key point of the proposed method and we obtain very promising results. Jérémie Nicolle, Vincent Rapp, Kevin Bailly, Lionel Prevost, Mohamed Chetouani |
ICMI | 5 |
| 2012 | Novel Metrics of Speech Rhythm for the Assessment of EmotionabstractInternational audience Fabien Ringeval, Mohamed Chetouani, Björn W. Schuller |
INTERSPEECH | 2 |
| 2012 | Interpersonal Synchrony: A Survey of Evaluation Methods across DisciplinesabstractSynchrony refers to individuals' temporal coordination during social interactions. The analysis of this phenomenon is complex, requiring the perception and integration of multimodal communicative signals. The evaluation of synchrony has received multidisciplinary attention because of its role in early development, language learning, and social connection. Originally studied by developmental psychologists, synchrony has now captured the interest of researchers in such fields as social signal processing, robotics, and machine learning. This paper emphasizes the current questions asked by synchrony evaluation and the state-of-the-art related methods. First, we present definitions and functions of synchrony in youth and adulthood. Next, we review the noncomputational and computational approaches of annotating, evaluating, and modeling interactional synchrony. Finally, the current limitations and future research directions in the fields of developmental robotics, social robotics, and clinical studies are discussed. Emilie Delaherche, Mohamed Chetouani, Ammar Mahdhaoui, Catherine Saint-Georges, Sylvie Viaux |
IEEE Trans. Affect. Comput. | 2 |
| 2011 | Characterization of coordination in an imitation task: human evaluation and automatically computable cuesabstractUnderstanding the ability to coordinate with a partner constitutes a great challenge in social signal processing and social robotics. In this paper, we designed a child-adult imitation task to investigate how automatically computable cues on turn-taking and movements can give insight into high-level perception of coordination. First we collected a human questionnaire to evaluate the perceived coordination of the dyads. Then, we extracted automatically computable cues and information on dialog acts from the video clips. The automatic cues characterized speech and gestural turn-takings and coordinated movements of the dyad. We finally confronted human scores with automatic cues to search which cues could be informative on the perception of coordination during the task. We found that the adult adjusted his behavior according to the child need and that a disruption of the gestural turn-taking rhythm was badly perceived by the judges. We also found, that judges rated negatively the dyads that talked more as speech intervenes when the child had difficulties to imitate. Finally, coherence measures between the partners' movement features seemed more adequate than correlation to characterize their coordination. Emilie Delaherche, Mohamed Chetouani |
ICMI | 2 |
| 2011 | Supervised and semi-supervised infant-directed speech classification for parent-infant interaction analysis
Ammar Mahdhaoui, Mohamed Chetouani |
Speech Commun. | 2 |
| 2011 | Automatic Intonation Recognition for the Prosodic Assessment of Language-Impaired ChildrenabstractThis study presents a preliminary investigation into the automatic assessment of language-impaired children's (LIC) prosodic skills in one grammatical aspect: sentence modalities. Three types of language impairments were studied: autism disorder (AD), pervasive developmental disorder-not otherwise specified (PDD-NOS), and specific language impairment (SLI). A control group of typically developing (TD) children that was both age and gender matched with LIC was used for the analysis. All of the children were asked to imitate sentences that provided different types of intonation (e.g., descending and rising contours). An automatic system was then used to assess LIC's prosodic skills by comparing the intonation recognition scores with those obtained by the control group. The results showed that all LIC have difficulties in reproducing intonation contours because they achieved significantly lower recognition scores than TD children on almost all studied intonations (p <; 0.05). Regarding the “Rising” intonation, only SLI children had high recognition scores similar to TD children, which suggests a more pronounced pragmatic impairment in AD and PDD-NOS children. The automatic approach used in this study to assess LIC's prosodic skills confirms the clinical descriptions of the subjects' communication impairments. Fabien Ringeval, Jean Demouy, György Szaszák, Mohamed Chetouani, L. Robel, Jean Xavier, Monique Plaza |
IEEE Trans. Speech Audio Process. | 4 |
| 2010 | Automatic gait characterization for a mobility assistance systemabstractThis paper addresses gait analysis for a mobility assistance robot designed for the elderly people. Six patients and ten healthy peoples were invited to be part of our first pilot experiment. We designed two experiments so as to firstly detect gait parameters and secondly to identify a change of speed. For the first trial, we compared the temporal-distance parameters of the healthy people and of individuals suffering of mobility problems. The percentages of the gait cycle for duration of stance are higher for people with mobility impairment than for healthy people. In the second experiment, we detected the change in walking speed from the ten healthy peoples. Two different metrics derived from the Kullback-Leibler (KL) divergence and from the Generalized Likelihood Ratio (GLR) were employed for walking change detection. The Receiver Operating Characteristic (ROC) curves show a better performance for the signal obtained with the accelerometer sensor than that obtained with the infrared distance sensor. Nevertheless, the results of our experiments demonstrated that both methodologies (KL and GLR) can be used to detect the change points during walking at high or slow speed. Cong Zong, Mohamed Chetouani, Adriana Tapus |
ICARCV | 2 |
| 2010 | Emotional Speech Classification Based on Multi View CharacterizationabstractEmotional speech classification is a key problem in social interaction analysis. Traditional emotional speech classification methods are completely supervised and require large amounts of labeled data. In addition, various feature sets are usually used to characterize the emotional speech signals. Therefore, we propose a new co-training algorithm based on multi-view features. More specifically, we adopt different features for the characterization of speech signals to form different views for classification, so as to extract as much discriminative information as possible. We then use the co-training algorithm to classify emotional speech with only few annotations. In this article, a dynamic weighted co-training algorithm is developed to combine different features (views) to predict the common class variable. Experiments prove the validity and effectiveness of this method compared to self-training algorithm. Ammar Mahdhaoui, Mohamed Chetouani |
ICPR | 2 |
| 2010 | Voice and graphical -based interfaces for interaction with a robot dedicated to elderly and people with cognitive disordersabstractHuman-robot interaction (HRI) takes place especially through interfaces. The design of such interfaces is a very delicate and crucial phase because it influences the robot accessibility and usability by the user. In this paper, we describe and analyze the results of 2 tests conducted so as to understand some of the optimal features that should characterize the robot voice and graphical-based user interfaces. Our test platform is an assistive robot developed for the elderly with mild cognitive impairments. Therefore, the user interfaces must be clear and simple. The ambiguities must be eliminated so as to facilitate the use of the robot and hence not to discourage the elderly population to use new technologies. Consuelo Granata, Mohamed Chetouani, Adriana Tapus, Philippe Bidaud, Vincent Dupourqué |
RO-MAN | 2 |
| 2009 | Generating Robot/Agent backchannels during a storytelling experimentabstractThis work presents the development of a real-time framework for the research of multimodal feedback of robots/talking agents in the context of Human Robot Interaction (HRI) and Human Computer Interaction (HCI). For evaluating the framework, a Multimodal corpus is built (ENTERFACE_STEAD), and a study on the important multimodal features was done for building an active Robot/Agent listener of a storytelling experience with Humans. The experiments show that even when building the same reactive behavior models for Robot and Talking Agents, the interpretation and the realization of the behavior communicated is different due to the different communicative channels Robots/Agents offer be it physical but less human-like in Robots, and virtual but more expressive and human-like in Talking agents. Sames Al Moubayed, Malek Baklouti, Mohamed Chetouani, Thierry Dutoit, Ammar Mahdhaoui, Jean-Claude Martin, Stanislav Ondás, Catherine Pelachaud, Jérôme Urbain |
ICRA | 3 |
| 2009 | Investigation on LP-residual representations for speaker identification
Mohamed Chetouani, Marcos Faúndez-Zanuy, Bruno Gas, Jean-Luc Zarader |
Pattern Recognit. | 1 |
| 2009 | Optimizing feature complementarity by evolution strategy: Application to automatic speaker verification
Christophe Charbuillet, Bruno Gas, Mohamed Chetouani, Jean-Luc Zarader |
Speech Commun. | 3 |
| 2009 | Special issue on non-linear and non-conventional speech processing
Mohamed Chetouani, Marcos Faúndez-Zanuy, Amir Hussain 0001, Bruno Gas, Jean-Luc Zarader, Kuldip K. Paliwal |
Speech Commun. | 1 |
| 2009 | Maximum likelihood linear programming data fusion for speaker recognition
Enric Monte-Moreno, Mohamed Chetouani, Marcos Faúndez-Zanuy, Jordi Solé i Casals |
Speech Commun. | 2 |
| 2008 | Motherese detection based on segmental and supra-segmental featuresabstractIn this paper, we present an automatic motherese detection system for the study of parent-infant interaction analysis. Motherese is a speech register directed towards infants and it is characterized by higher pitch, slower tempo, and exaggerated intonation. The goal of this paper is to propose and evaluate different approaches for the detection of motherese from home movies. We investigated the characterization by supra-segmental features (prosody) but also by segmental ones namely the MFCC (Mel Frequency Cepstral Coefficients). Concerning the classification stage, we investigated two different methods: the k-nn (k-nearest neighbors) and the GMM (Gaussian Mixture Models). Experimental results show that segmental features play a major role on the detection. Ammar Mahdhaoui, Mohamed Chetouani, Cong Zong |
ICPR | 2 |
| 2008 | A vowel based approach for acted emotion recognitionabstractThis paper is devoted to the description of a new approach for emotion recognition. Our contribution is based on both the extraction and the characterization of phonemic units such as vowels and consonants, which are provided by a pseudophonetic speech segmentation phase combined with a vowel detector. Concerning the emotion recognition task, we explore acoustic and prosodic features from these pseudo-phonetic segments (vowels and consonants), and we compare this approach with traditional voiced and unvoiced segments. The classification is realized by the well-known k-nn classifier (k nearest neighbors) from two different emotional speech databases: Berlin (German) and Aholab (Basque). Fabien Ringeval, Mohamed Chetouani |
INTERSPEECH | 2 |
| 2007 | Complementary Features for Speaker Verification Based on Genetic AlgorithmsabstractSpeech recognition systems usually need a feature extraction stage aiming at obtaining the best signal representation. State of the art speaker verification systems are based on cepstrals features like MFCC, LFCC or LPCC. In this article, we propose to use a genetic algorithm to provide new features able to complete the LFCC's. We present an adaptation of the common LFCC feature extractor which consists in designing a filter bank, optimized for a high level fusion issue. Experiments are carried out using a state of the art speaker verification system. Results show that the proposed method improves the system performances on the 2006 Nist SRE Database. Christophe Charbuillet, Bruno Gas, Mohamed Chetouani, Jean-Luc Zarader |
ICASSP (4) | 3 |
| 2006 | Filter Bank Design for Speaker Diarization Based on Genetic AlgorithmsabstractSpeech recognition systems usually need a feature extraction stage aiming at obtaining the best signal representation. In this article we propose to use genetic algorithms to design a feature extraction method adapted to the speaker diarization task. We present an adaptation of the common MFCC feature extractor which consists in designing a filter bank, with optimized bandwidths. Experiments are carried out using a state-of-the-art speaker diarization system. The proposed method outperforms the original filter bank based on the Mel scale one. Furthermore, the obtained filter bank reveals the importance of some specific spectral information for speaker recognition Christophe Charbuillet, Bruno Gas, Mohamed Chetouani, Jean-Luc Zarader |
ICASSP (1) | 3 |
| 2005 | Non-linear Predictive Models for Speech Processing
Mohamed Chetouani, Amir Hussain 0001, Marcos Faúndez-Zanuy, Bruno Gas |
ICANN (2) | 1 |
| 2005 | Predictive Kohonen Map for Speech Features Extraction
Bruno Gas, Mohamed Chetouani, Jean-Luc Zarader, Christophe Charbuillet |
ICANN (2) | 2 |
| 2004 | A new nonlinear feature extraction algorithm for speaker verificationabstractIn this paper we propose a new feature extraction algorithm based on nonlinear prediction: the Neural Predictive Coding model which is an extension of the classical LPC one. This model is applied to speaker verification by the Arithmetic-Harmonic Sphericity (AHS) method. Two different initialization methods are proposed for the coding method based on the Neural Predictive Coding (NPC): classical neural networks initialization and linear initialization. The first model obtains smaller rates. For the linear initialization, we obtain significant improvement in comparison to the most used methods (LPCC, MFCC). This study opens a new way towards different feature extraction schemes that offers better accuracy on speaker recognition tasks. Mohamed Chetouani, Bruno Gas, Jean-Luc Zarader, Marcos Faúndez-Zanuy |
INTERSPEECH | 1 |
| 2004 | Discriminant neural predictive coding applied to phoneme recognition
Bruno Gas, Jean-Luc Zarader, Cyril Chavy, Mohamed Chetouani |
Neurocomputing | 4 |
| 2003 | Modular neural predictive coding for discriminative feature extractionabstractWe present an architecture called the modular neural predictive coding architecture (Modular NPC). The Modular NPC is used for discriminative feature extraction (DFE). It provides an architecture based on phonetics knowledge applied to phoneme recognition. The phonemes are extracted from the Darpa-Timit speech database. Comparisons with coding methods (LPC, MFCC, PLP) are presented: they put in obviousness an improvement of the recognition rates. Mohamed Chetouani, Bruno Gas, Jean-Luc Zarader |
ICASSP (2) | 1 |
| 2002 | Neural predictive coding for speech discriminant feature extraction: The DFE-NPC
Mohamed Chetouani, Bruno Gas, Jean-Luc Zarader, Cyril Chavy |
ESANN | 1 |