VLDB 2026 Research / reviewers in the wild / expert
Chloé Clavel
dblp:50/2768
· DBLP profile ↗
76ranked-venue papers
7as first author
39since 2021 · last 2026
0000-0003-4850-3398ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 56 · 3 first-author · 32 since 2021Human-computer interaction and ubiquitous computing · 16 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 2 since 2021Computer networks · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SPOT: An Annotated French Corpus and Benchmark for Detecting Critical Interventions in Online ConversationsabstractWe introduce SPOT (Stopping Points in Online Threads), the first annotated corpus translating the sociological concept of stopping point into a reproducible NLP task. Stopping points are ordinary critical interventions that pause or redirect online discussions through a range of forms (irony, subtle doubt or fragmentary arguments) that frameworks like counterspeech or social correction often overlook. We operationalize this concept as a binary classification task and provide reliable annotation guidelines. The corpus contains 43,305 manually annotated French Facebook comments linked to URLs flagged as false information by social media users, enriched with contextual metadata (article, post, parent comment, page or group, and source). We benchmark fine-tuned encoder models (CamemBERT) and instruction-tuned LLMs under various prompting strategies. Results show that fine-tuned encoders outperform prompted LLMs in F1 score by more than 10 percentage points, confirming the importance of supervised learning for emerging non-English social media tasks. Incorporating contextual metadata further improves encoder models F1 scores from 0.75 to 0.78. We release the anonymized dataset, along with the annotation guidelines and code in our code repository, to foster transparency and reproducible research. Manon Berriche, Célia Nouri, Chloé Clavel, Jean-Philippe Cointet |
LREC | 3 |
| 2026 | Human vs LLM in Conversational Repair Annotation: A New Resource and Comparative StudyabstractInternational audience Anh Ngo, Nicolas Rollet, Catherine Pelachaud, Chloé Clavel |
LREC | 4 |
| 2026 | Scare Quotes as Markers of "Questionable" Word Usages and Misalignment in Conversation: An Annotation Study
Aina Garí Soler, Juan Carlos Zevallos Huaco, Matthieu Labeau, Chloé Clavel |
LREC | 4 |
| 2026 | Automatic Analysis of Collaboration through Human Conversational Data Resources: A Review
Maria Boritchev, Chloé Clavel |
LREC | 3 |
| 2025 | Graphically Speaking: Unmasking Abuse in Social Media with Conversation InsightsabstractDetecting abusive language in social media conversations poses significant challenges, as identifying abusiveness often depends on the conversational context, characterized by the content and topology of preceding comments.Traditional Abusive Language Detection (ALD) models often overlook this context, which can lead to unreliable performance metrics.Recent Natural Language Processing (NLP) approaches that incorporate conversational context often rely on limited or overly simplified representations of this context, leading to inconsistent and sometimes inconclusive results.In this paper, we propose a novel approach that utilizes graph neural networks (GNNs) to model social media conversations as graphs, where nodes represent comments, and edges capture reply structures.We systematically investigate various graph representations and context windows to identify the optimal configurations for ALD.Our GNN model outperforms both context-agnostic baselines and linear context-aware methods, achieving significant improvements in F1 scores.These findings demonstrate the critical role of structured conversational context and establish GNNs as a robust framework for advancing context-aware ALD.Our code is available at this link. Célia Nouri, Chloé Clavel, Jean-Philippe Cointet |
ACL (1) | 2 |
| 2025 | "Mm, Wat?" Detecting Other-initiated Repair Requests in DialogueabstractMaintaining mutual understanding is a key component in human-human conversation to avoid conversation breakdowns, in which repair, particularly Other-Initiated Repair (OIR, when one speaker signals trouble and prompts the other to resolve), plays a vital role.However, Conversational Agents (CAs) still fail to recognize user repair initiation, leading to breakdowns or disengagement.This work proposes a multimodal model to automatically detect repair initiation in Dutch dialogues by integrating linguistic and prosodic features grounded in Conversation Analysis.The results show that prosodic cues complement linguistic features and significantly improve the results of pretrained text and audio embeddings, offering insights into how different features interact.Future directions include incorporating visual cues, exploring multilingual and cross-context corpora to assess the robustness and generalizability. Anh Ngo, Nicolas Rollet, Catherine Pelachaud, Chloé Clavel |
EMNLP | 4 |
| 2025 | Decoding Persuasiveness in Eloquence Competitions: An Investigation into the LLM's Ability to Assess Public SpeakingabstractInternational audience Alisa Barkar, Mathieu Chollet, Matthieu Labeau, Béatrice Biancardi, Chloé Clavel |
ICAART (3) | 5 |
| 2025 | EmoDynamiX: Emotional Support Dialogue Strategy Prediction by Modelling MiXed Emotions and Discourse DynamicsabstractChenwei Wan, Matthieu Labeau, Chloé Clavel. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Chenwei Wan, Matthieu Labeau, Chloé Clavel |
NAACL (Long Papers) | 3 |
| 2025 | Benchmarking Linguistic Diversity of Large Language ModelsabstractAbstract The development and evaluation of Large Language Models (LLMs) has primarily focused on their task-solving capabilities, with recent models even surpassing human performance in some areas. However, this focus often neglects whether machine-generated language matches the human level of diversity, in terms of vocabulary choice, syntactic construction, and expression of meaning, raising questions about whether the fundamentals of language generation have been fully addressed. This paper emphasizes the importance of examining the preservation of human linguistic richness by language models, given the concerning surge in online content produced or aided by LLMs. We adapt a comprehensive framework for evaluating LLMs from various linguistic diversity perspectives including lexical, syntactic, and semantic dimensions. Using this framework, we benchmark several state-of-the-art LLMs across all diversity dimensions, and conduct an in-depth analysis for syntactic diversity. Finally, we analyze how the design, development, and deployment choices of LLMs impact the linguistic diversity of their outputs, focusing on the creative task of story generation. Yanzhu Guo, Guokan Shang, Chloé Clavel |
Trans. Assoc. Comput. Linguistics | 3 |
| 2024 | MAFALDA: A Benchmark and Comprehensive Study of Fallacy Detection and ClassificationabstractChadi Helwe, Tom Calamai, Pierre-Henri Paris, Chloé Clavel, Fabian Suchanek. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Chadi Helwe, Tom Calamai, Pierre-Henri Paris, Chloé Clavel, Fabian M. Suchanek |
NAACL-HLT | 4 |
| 2024 | Exploration of Human Repair Initiation in Task-oriented Dialogue: A Linguistic Feature-based ApproachabstractIn daily conversations, people often encounter problems prompting conversational repair to enhance mutual understanding.By employing an automatic coreference solver, alongside examining repetition, we identify various linguistic features that distinguish turns when the addressee initiates repair from those when they do not.Our findings reveal distinct patterns that characterize the repair sequence and each type of other-repair initiation. Anh Ngo, Dirk Heylen, Nicolas Rollet, Catherine Pelachaud, Chloé Clavel |
SIGDIAL | 5 |
| 2024 | Do Language Models Enjoy Their Own Stories? Prompting Large Language Models for Automatic Story EvaluationabstractAbstract Storytelling is an integral part of human experience and plays a crucial role in social interactions. Thus, Automatic Story Evaluation (ASE) and Generation (ASG) could benefit society in multiple ways, but they are challenging tasks which require high-level human abilities such as creativity, reasoning, and deep understanding. Meanwhile, Large Language Models (LLMs) now achieve state-of-the-art performance on many NLP tasks. In this paper, we study whether LLMs can be used as substitutes for human annotators for ASE. We perform an extensive analysis of the correlations between LLM ratings, other automatic measures, and human annotations, and we explore the influence of prompting on the results and the explainability of LLM behaviour. Most notably, we find that LLMs outperform current automatic measures for system-level evaluation but still struggle at providing satisfactory explanations for their answers. Cyril Chhun, Fabian M. Suchanek, Chloé Clavel |
Trans. Assoc. Comput. Linguistics | 3 |
| 2024 | The Impact of Word Splitting on the Semantic Content of Contextualized Word RepresentationsabstractAbstract When deriving contextualized word representations from language models, a decision needs to be made on how to obtain one for out-of-vocabulary (OOV) words that are segmented into subwords. What is the best way to represent these words with a single vector, and are these representations of worse quality than those of in-vocabulary words? We carry out an intrinsic evaluation of embeddings from different models on semantic similarity tasks involving OOV words. Our analysis reveals, among other interesting findings, that the quality of representations of words that are split is often, but not always, worse than that of the embeddings of known words. Their similarity values, however, must be interpreted with caution. Aina Garí Soler, Matthieu Labeau, Chloé Clavel |
Trans. Assoc. Comput. Linguistics | 3 |
| 2023 | An Adaptive Layer to Leverage Both Domain and Task Specific Information from Scarce DataabstractMany companies make use of customer service chats to help the customer and try to solve their problem. However, customer service data is confidential and as such, cannot easily be shared in the research community. This also implies that these data are rarely labeled, making it difficult to take advantage of it with machine learning methods. In this paper we present the first work on a customer’s problem status prediction and identification of problematic conversations. Given very small subsets of labeled textual conversations and unlabeled ones, we propose a semi-supervised framework dedicated to customer service data leveraging speaker role information to adapt the model to the domain and the task using a two-step process. Our framework, Task-Adaptive Fine-tuning, goes from predicting customer satisfaction to identifying the status of the customer’s problem, with the latter being the main objective of the multi-task setting. It outperforms recent inductive semi-supervised approaches on this novel task while only considering a relatively low number of parameters to train on during the final target task. We believe it can not only serve models dedicated to customer service but also to any other application making use of confidential conversational data where labeled sets are rare. Source code is available at https://github.com/gguibon/taft Gaël Guibon, Matthieu Labeau, Luce Lefeuvre, Chloé Clavel |
AAAI | 4 |
| 2023 | A New Task for Predicting Emotions and Dialogue Strategies in Task-Oriented DialogueabstractWith a focus on task-oriented conversational agents, we introduce a new approach to make the generated response more relevant to the emotional context of the interlocutor. This approach aims to predict the labels of the agent’s next speaker turn, in order to condition the generated response and ensure its consistency to the user’s socio-emotional context. First, we propose a new formulation of this prediction task, based on the joint prediction of dialogue strategies and emotional labels associated with the next speaker turn. To handle this new task, we describe a new annotation protocol for task-oriented dialogue systems, that we implement on real customer-relationship interactions provided by a company. Lastly, we conduct an experiment using two approaches: a classification and a generation model. We evaluate them on the new prediction task using the annotated data, before discussing the results. As expected, we see that our classifier tends to predict safe, similar labels where the generator has a more diverse output in spite of its lower performance on traditional metrics. Lorraine Vanel, Alya Yacoubi, Chloé Clavel |
ACII | 3 |
| 2023 | How About Kind of Generating Hedges using End-to-End Neural Models?abstractHedging is a strategy for softening the impact of a statement in conversation.In reducing the strength of an expression, it may help to avoid embarrassment (more technically, "face threat") to one's listener.For this reason, it is often found in contexts of instruction, such as tutoring.In this work, we develop a model of hedge generation based on i) fine-tuning stateof-the-art language models trained on humanhuman tutoring data, followed by ii) reranking to select the candidate that best matches the expected hedging strategy within a candidate pool using a hedge classifier.We apply this method to a natural peer-tutoring corpus containing a significant number of disfluencies, repetitions, and repairs.The results show that generation in this noisy environment is feasible with reranking.By conducting an error analysis for both approaches, we reveal the challenges faced by systems attempting to accomplish both social and task-oriented goals in conversation. Alafate Abulimiti, Chloé Clavel, Justine Cassell |
ACL (1) | 2 |
| 2023 | Leveraging Interactional Sociology for Trust Analysis in Multiparty Human-Robot InteractionabstractBy leveraging Interactional Sociology theories, multimodal behavioral features and recurrent neural architectures, we incrementally build computational models for trust analysis in multiparty human-robot interactions (HRI). We show that the model’s performance improves when i) modeling group dynamics with different granularities (i.e. group member, dyadic, and group as a whole), and ii) modeling users-robot interactions as a question-answer sequence. Marc Hulcelle, Léo Hemamou, Giovanna Varni, Nicolas Rollet, Chloé Clavel |
HAI | 5 |
| 2023 | Comparing a Mentalist and an Interactionist Approach for Trust Analysis in Human-Robot InteractionabstractTrust is an important aspect of a human-robot interaction (HRI) as it mitigates the performance of many activities. Users’ trust may be impacted when robots make mistakes. To be able to properly time trust-reparation actions, robots should detect trust variations during the interaction. There are very few computational models of trust for such a task. The existing ones relied on either Psychological or Sociological theories that gave place to different definitions and analysis tools. We can distinguish two main approaches in the trust literature: the mentalist and the interactionist one. In this paper, we compare both approaches for trust detection, and explore how the adoption of two different assessment tools on an HRI dataset may lead to different results. We identify criteria that set them apart, and provide guidelines on the possibilities that each approach offers depending on the target computational model of trust. Marc Hulcelle, Giovanna Varni, Nicolas Rollet, Chloé Clavel |
HAI | 4 |
| 2023 | A Survey of Socio-Emotional Strategies for Generation-Based Conversational Agents
Lorraine Vanel, Alya Yacoubi, Chloé Clavel |
ICAART (3) | 3 |
| 2023 | When to generate hedges in peer-tutoring interactionsabstractThis paper explores the application of machine learning techniques to predict where hedging occurs in peer-tutoring interactions.The study uses a naturalistic face-to-face dataset annotated for natural language turns, conversational strategies, tutoring strategies, and nonverbal behaviors.These elements are processed into a vector representation of the previous turns, which serves as input to several machine learning models, including MLP and LSTM.The results show that embedding layers, capturing the semantic information of the previous turns, significantly improves the model's performance.Additionally, the study provides insights into the importance of various features, such as interpersonal rapport and nonverbal behaviors, in predicting hedges by using Shapley values (Hart, 1989) for feature explanation.We discover that the eye gaze of both the tutor and the tutee has a significant impact on hedge prediction.We further validate this observation through a follow-up ablation study. Alafate Abulimiti, Chloé Clavel, Justine Cassell |
SIGDIAL | 2 |
| 2023 | Multimodal Hierarchical Attention Neural Network: Looking for Candidates Behaviour Which Impact Recruiter's DecisionabstractAutomatic analysis of job interviews has gained in interest amongst academic and industrial research. The particular case of asynchronous video interviews allows to collect vast corpora of videos where candidates answer standardized questions in monologue videos, enabling the use of deep learning algorithms. On the other hand, state-of-the-art approaches still face some obstacles, among which the fusion of information from multiple modalities and the interpretability of the predictions. We study the task of predicting candidates performance in asynchronous video interviews using three modalities (verbal content, prosody and facial expressions) independently or simultaneously, using data from real interviews which take place in real conditions. We propose a sequential and multimodal deep neural network model, called Multimodal HireNet. We compare this model to state-of-the-art approaches and show a clear improvement of the performance. Moreover, the architecture we propose is based on attention mechanism, which provides interpretability about which questions, moments and modalities contribute the most to the output of the network. While other deep learning systems use attention mechanisms to offer a visualization of moments with attention values, the proposed methodology enables an in-depth interpretation of the predictions by an overall analysis of the features of social signals contained in these moments. Léo Hemamou, Arthur Guillon, Jean-Claude Martin, Chloé Clavel |
IEEE Trans. Affect. Comput. | 4 |
| 2022 | InfoLM: A New Metric to Evaluate Summarization & Data2Text GenerationabstractAssessing the quality of natural language generation (NLG) systems through human annotation is very expensive. Additionally, human annotation campaigns are time-consuming and include non-reusable human labour. In practice, researchers rely on automatic metrics as a proxy of quality. In the last decade, many string-based metrics (e.g., BLEU or ROUGE) have been introduced. However, such metrics usually rely on exact matches and thus, do not robustly handle synonyms. In this paper, we introduce InfoLM a family of untrained metrics that can be viewed as a string-based metric that addresses the aforementioned flaws thanks to a pre-trained masked language model. This family of metrics also makes use of information measures allowing the possibility to adapt InfoLM to different evaluation criteria. Using direct assessment, we demonstrate that InfoLM achieves statistically significant improvement and two figure correlation gains in many configurations compared to existing metrics on both summarization and data2text generation tasks. Pierre Colombo, Chloé Clavel, Pablo Piantanida |
AAAI | 2 |
| 2022 | Domain Adaptation for Stance Detection towards Unseen Target on Social MediaabstractStance detection aims at identifying people's stand-point towards a given target. New targets are constantly appearing on social media, making previous annotated data unusable by stance detection models relying on classical supervised machine learning. Thus, cross-target stance detection which uses labeled data from source targets to learn a model that can be adapted to the destination new target, has become a prevailing research direction. However, previous methods rely on manually chosen similar source-destination target pairs and lack generalization to unseen targets with no explicit relation to known ones. To this end, we investigate the problem from a domain adaptation perspective and further propose a novel Unified Target-aware Domain Adaptation method (UTDA) that leverages knowledge transfer capability of transformer-based language model. The proposed method can effectively extract critical target-shared features for detecting stance by feature disentanglement and automatically learn to identify target relations. UTDA can easily be applied to a new unseen target since it does not rely on any pre-defined target pairs. Experimental results on two benchmark stance datasets demonstrate that our method achieves better performance than strong baselines. Ruofan Deng, Li Pan 0002, Chloé Clavel |
ACII | 3 |
| 2022 | "You might think about slightly revising the title": Identifying Hedges in Peer-tutoring InteractionsabstractHedges play an important role in the management of conversational interaction.In peertutoring, they are notably used by tutors in dyads (pairs of interlocutors) experiencing low rapport to tone down the impact of instructions and negative feedback.Pursuing the objective of building a tutoring agent that manages rapport with students in order to improve learning, we used a multimodal peer-tutoring dataset to construct a computational framework for identifying hedges.We compared approaches relying on pre-trained resources with others that integrate insights from the social science literature.Our best performance involved a hybrid approach that outperforms the existing baseline while being easier to interpret.We employ a model explainability tool to explore the features that characterize hedges in peer-tutoring conversations, and we identify some novel features, and the benefits of such a hybrid model approach.Models RB MLP (KDF) MLP (PTE) MLP (K+P) CNN (PTE) LSTM (KDF) LSTM(PTE) LSTM (K+P) BERT (PTE) LGB (KDF) LGB (PTE) LGB (K+P) Rule-based No Yes No No No Yes No No Yes Yes Yann Raphalen, Chloé Clavel, Justine Cassell |
ACL (1) | 2 |
| 2022 | Of Human Criteria and Automatic Metrics: A Benchmark of the Evaluation of Story GenerationabstractResearch on Automatic Story Generation (ASG) relies heavily on human and automatic evaluation. However, there is no consensus on which human evaluation criteria to use, and no analysis of how well automatic criteria correlate with them. In this paper, we propose to re-evaluate ASG evaluation. We introduce a set of 6 orthogonal and comprehensive human criteria, carefully motivated by the social sciences literature. We also present HANNA, an annotated dataset of 1,056 stories produced by 10 different ASG systems. HANNA allows us to quantitatively evaluate the correlations of 72 automatic metrics with human criteria. Our analysis highlights the weaknesses of current metrics for ASG and allows us to formulate practical recommendations for ASG evaluation. Cyril Chhun, Pierre Colombo, Fabian M. Suchanek, Chloé Clavel |
COLING | 4 |
| 2022 | One Word, Two Sides: Traces of Stance in Contextualized Word RepresentationsabstractThe way we use words is influenced by our opinion. We investigate whether this is reflected in contextualized word embeddings. For example, is the representation of “animal” different between people who would abolish zoos and those who would not? We explore this question from a Lexical Semantic Change standpoint. Our experiments with BERT embeddings derived from datasets with stance annotations reveal small but significant differences in word representations between opposing stances. Aina Garí Soler, Matthieu Labeau, Chloé Clavel |
COLING | 3 |
| 2022 | Questioning the Validity of Summarization Datasets and Improving Their Factual ConsistencyabstractThe topic of summarization evaluation has recently attracted a surge of attention due to the rapid development of abstractive summarization systems.However, the formulation of the task is rather ambiguous, neither the linguistic nor the natural language processing community has succeeded in giving a mutually agreed-upon definition.Due to this lack of well-defined formulation, a large number of popular abstractive summarization datasets are constructed in a manner that neither guarantees validity nor meets one of the most essential criteria of summarization: factual consistency.In this paper, we address this issue by combining state-of-the-art factual consistency models to identify the problematic instances present in popular summarization datasets.We release SummFC, a filtered summarization dataset with improved factual consistency, and demonstrate that models trained on this dataset achieve improved performance in nearly all quality aspects.We argue that our dataset should become a valid benchmark for developing and evaluating summarization systems. Yanzhu Guo, Chloé Clavel, Moussa Kamal Eddine, Michalis Vazirgiannis |
EMNLP | 2 |
| 2022 | Opinions in Interactions : New Annotations of the SEMAINE DatabaseabstractIn this paper, we present the process we used in order to collect new annotations of opinions over the multimodal corpus SEMAINE composed of dyadic interactions. The dataset had already been annotated continuously in two affective dimensions related to the emotions: Valence and Arousal. We annotated the part of SEMAINE called Solid SAL composed of 79 interactions between a user and an operator playing the role of a virtual agent designed to engage a person in a sustained, emotionally colored conversation. We aligned the audio at the word level using the available high-quality manual transcriptions. The annotated dataset contains 5627 speech turns for a total of 73,944 words, corresponding to 6 hours 20 minutes of dyadic interactions. Each interaction has been labeled by three annotators at the speech turn level following a three-step process. This method allows us to obtain a precise annotation regarding the opinion of a speaker. We obtain thus a dataset dense in opinions, with more than 48% of the annotated speech turns containing at least one opinion. We then propose a new baseline for the detection of opinions in interactions improving slightly a state of the art model with RoBERTa embeddings. The obtained results on the database are promising with a F1-score at 0.72. Valentin Barrière, Slim Essid, Chloé Clavel |
LREC | 3 |
| 2022 | EZCAT: an Easy Conversation Annotation ToolabstractUsers generate content constantly, leading to new data requiring annotation. Among this data, textual conversations are created every day and come with some specificities: they are mostly private through instant messaging applications, requiring the conversational context to be labeled. These specificities led to several annotation tools dedicated to conversation, and mostly dedicated to dialogue tasks, requiring complex annotation schemata, not always customizable and not taking into account conversation-level labels. In this paper, we present EZCAT, an easy-to-use interface to annotate conversations in a two-level configurable schema, leveraging message-level labels and conversation-level labels. Our interface is characterized by the voluntary absence of a server and accounts management, enhancing its availability to anyone, and the control over data, which is crucial to confidential conversations. We also present our first usage of EZCAT along with our annotation schema we used to annotate confidential customer service conversations. EZCAT is freely available at https://gguibon.github.io/ezcat. Gaël Guibon, Luce Lefeuvre, Matthieu Labeau, Chloé Clavel |
LREC | 4 |
| 2022 | Polysemy in Spoken Conversations and Written TextsabstractOur discourses are full of potential lexical ambiguities, due in part to the pervasive use of words having multiple senses. Sometimes, one word may even be used in more than one sense throughout a text. But, to what extent is this true for different kinds of texts? Does the use of polysemous words change when a discourse involves two people, or when speakers have time to plan what to say? We investigate these questions by comparing the polysemy level of texts of different nature, with a focus on spontaneous spoken dialogs; unlike previous work which examines solely scripted, written, monolog-like data. We compare multiple metrics that presuppose different conceptualizations of text polysemy, i.e., they consider the observed or the potential number of senses of words, or their sense distribution in a discourse. We show that the polysemy level of texts varies greatly depending on the kind of text considered, with dialog and spoken discourses having generally a higher polysemy level than written monologs. Additionally, our results emphasize the need for relaxing the popular “one sense per discourse” hypothesis. Aina Garí Soler, Matthieu Labeau, Chloé Clavel |
LREC | 3 |
| 2021 | Don't Judge Me by My Face: An Indirect Adversarial Approach to Remove Sensitive Information From Multimodal Neural Representation in Asynchronous Job Video InterviewsabstractUse of machine learning for automatic analysis of job interview videos has recently seen increased interest. Despite claims of fair output regarding sensitive information such as gender or ethnicity of the candidates, the current approaches rarely provide proof of unbiased decision-making, or that sensitive information is not used. Recently, adversarial methods have been proved to effectively remove sensitive information from the latent representation of neural networks. However, these methods rely on the use of explicitly labeled protected variables (e.g. gender), which cannot be collected in the context of recruiting in some countries (e.g. France). In this article, we propose a new adversarial approach to remove sensitive information from the latent representation of neural networks without the need to collect any sensitive variable. Using only a few frames of the interview, we train our model to not be able to find the face of the candidate related to the job interview in the inner layers of the model. This, in turn, allows us to remove relevant private information from these layers. Comparing our approach to a standard baseline on a public dataset with gender and ethnicity annotations, we show that it effectively removes sensitive information from the main network. Moreover, to the best of our knowledge, this is the first application of adversarial techniques for obtaining a multimodal fair representation in the context of video job interviews. In summary, our contributions aim at improving fairness of the upcoming automatic systems processing videos of job interviews for equality in job selection. Léo Hemamou, Arthur Guillon, Jean-Claude Martin, Chloé Clavel |
ACII | 4 |
| 2021 | TURIN: A coding system for Trust in hUman Robot INteractionabstractA natural human-robot interaction (HRI) relies on the robot’s capacity to understand the users’ behaviors through psychological and sociological concepts. Users expect the robot to act in a realistic manner to create a more human-like relationship. In this context, trust is an essential concept as it determines the effectiveness of the system and its acceptance by users. The understanding of trust dynamics in HRI is still low and systematic studies of multimodal trust-related behaviors in HRI are relatively rare given the rising popularity of the topic. To bridge this gap, in this paper we present a novel coding system TURIN (Trust in hUman Robot INteraction) to study trust in HRI. A preliminary assessment of the coding system was carried out on the Vernissage dataset. Results show a significant agreement between expert annotators. Marc Hulcelle, Giovanna Varni, Nicolas Rollet, Chloé Clavel |
ACII | 4 |
| 2021 | A Novel Estimator of Mutual Information for Learning to Disentangle Textual RepresentationsabstractPierre Colombo, Pablo Piantanida, Chloé Clavel. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Pierre Colombo, Pablo Piantanida, Chloé Clavel |
ACL/IJCNLP (1) | 3 |
| 2021 | Improving Multimodal fusion via Mutual Dependency MaximisationabstractMultimodal sentiment analysis is a trending area of research, and the multimodal fusion is one of its most active topic.Acknowledging humans communicate through a variety of channels (i.e visual, acoustic, linguistic), multimodal systems aim at integrating different unimodal representations into a synthetic one.So far, a consequent effort has been made on developing complex architectures allowing the fusion of these modalities.However, such systems are mainly trained by minimising simple losses such as L 1 or cross-entropy.In this work, we investigate unexplored penalties and propose a set of new objectives that measure the dependency between modalities.We demonstrate that our new penalties lead to a consistent improvement (up to 4.3 on accuracy) across a large variety of state-of-the-art models on two well-known sentiment analysis datasets: CMU-MOSI and CMU-MOSEI.Our method not only achieves a new SOTA on both datasets but also produces representations that are more robust to modality drops.Finally, a by-product of our methods includes a statistical network which can be used to interpret the high dimensional representations learnt by the model. Pierre Colombo, Emile Chapuis, Matthieu Labeau, Chloé Clavel |
EMNLP (1) | 4 |
| 2021 | Code-switched inspired losses for spoken dialog representationsabstractSpoken dialog systems need to be able to handle both multiple languages and multilinguality inside a conversation (e.g in case of codeswitching).In this work, we introduce new pretraining losses tailored to learn multilingual spoken dialog representations.The goal of these losses is to expose the model to codeswitched language.To scale up training, we automatically build a pretraining corpus composed of multilingual conversations in five different languages (French, Italian, English, German and Spanish) from OpenSubtitles, a huge multilingual corpus composed of 24.3G tokens.We test the generic representations on MIAM, a new benchmark composed of five dialog act corpora on the same aforementioned languages as well as on two novel multilingual downstream tasks (i.e multilingual mask utterance retrieval and multilingual inconsistency identification).Our experiments show that our new code switched-inspired losses achieve a better performance in both monolingual and multilingual settings. Pierre Colombo, Emile Chapuis, Matthieu Labeau, Chloé Clavel |
EMNLP (1) | 4 |
| 2021 | Automatic Text Evaluation through the Lens of Wasserstein BarycentersabstractA new metric BaryScore to evaluate text generation based on deep contextualized embeddings (e.g., BERT, Roberta, ELMo) is introduced.This metric is motivated by a new framework relying on optimal transport tools, i.e., Wasserstein distance and barycenter.By modelling the layer output of deep contextualized embeddings as a probability distribution rather than by a vector embedding; this framework provides a natural way to aggregate the different outputs through the Wasserstein space topology.In addition, it provides theoretical grounds to our metric and offers an alternative to available solutions (e.g., Mover-Score and BertScore).Numerical evaluation is performed on four different tasks: machine translation, summarization, data2text generation and image captioning.Our results show that BaryScore outperforms other BERT based metrics and exhibits more consistent behaviour in particular for text summarization. Pierre Colombo, Guillaume Staerman, Chloé Clavel, Pablo Piantanida |
EMNLP (1) | 3 |
| 2021 | Few-Shot Emotion Recognition in Conversation with Sequential Prototypical NetworksabstractSeveral recent studies on dyadic humanhuman interactions have been done on conversations without specific business objectives.However, many companies might benefit from studies dedicated to more precise environments such as after sales services or customer satisfaction surveys.In this work, we place ourselves in the scope of a live chat customer service in which we want to detect emotions and their evolution in the conversation flow.This context leads to multiple challenges that range from exploiting restricted, small and mostly unlabeled datasets to finding and adapting methods for such context.We tackle these challenges by using Few-Shot Learning while making the hypothesis it can serve conversational emotion classification for different languages and sparse labels.We contribute by proposing a variation of Prototypical Networks for sequence labeling in conversation that we name ProtoSeq.We test this method on two datasets with different languages: daily conversations in English and customer service chat conversations in French.When applied to emotion classification in conversations, our method proved to be competitive even when compared to other ones.The code for Proto-Seq is available at https://github.com/ gguibon/ProtoSeq. Gaël Guibon, Matthieu Labeau, Hélène Flamein, Luce Lefeuvre, Chloé Clavel |
EMNLP (1) | 5 |
| 2021 | CATS2021: International Workshop on Corpora And Tools for Social skills annotationabstractThis Workshop aims at stimulating multi-disciplinary discussions about the challenges related to corpus creation and annotation for social skills behavior analysis. Contributions from computational, psychological and psychometrics perspectives, as well as applications including platforms to share corpora and annotations, are welcomed. The main challenges related to corpus creation include the choice of the best setup and sensors, finding a trade-off between eliciting natural interactions, limiting invasiveness and collecting precise information. The second issue in this context regards the process of annotation. The choice of the type of annotators (experts vs. nonexperts), the type of annotations (automatic vs. manual, continue vs. discrete), the temporal segmentation (windowed vs. holistic) is crucial for a correct measure of the phenomenon of interest and getting significant results. The topics of CATS2021 will have a strong impact on researchers and stakeholders across different disciplines, such as Computer Science, Social Signal Processing, Psychology, Statistics. Leveraging the opportunities offered by such a multidisciplinary environment, the participants could enrich their perspective, strengthen their practices and methodologies and draw together a research roadmap tackling the discussed challenges, which might be taken up in future collaborations. Béatrice Biancardi, Eleonora Ceccaldi, Chloé Clavel, Mathieu Chollet, Tanvi Dinkar |
ICMI | 3 |
| 2021 | Early Detection of User Engagement Breakdown in Spontaneous Human-Humanoid InteractionabstractThis paper presents a supervised classification system for forecasting a potential user engagement breakdown in human-robot interaction. We define engagement breakdown as a failure to successfully complete a predefined interaction scenario, where the user leaves before the expected end. The goal is thus to detect as early as possible such a potential engagement breakdown during the interaction between a human and a humanoid robot. To this end, we exploit a dataset that we have collected in real-world conditions where a set of participants were left to spontaneously engage in an interaction with the robot. The dataset is labeled according to the presence/absence of engagement breakdown. This study investigates the use of a multimodal approach to this problem, where a set of non-verbal features is considered to characterize the users' behavior. The use of combined multimodal features is found to effectively improve the performance of the system. The optimal set of data streams useful for this task is the combination of the distance to the robot, gaze and head motion, as well as facial expressions and speech. We study the time extent over which a user's departure can be anticipated. We find that this ability to anticipate the departure depends on the window during which we observe the user behavior. Atef Ben Youssef, Chloé Clavel, Slim Essid |
IEEE Trans. Affect. Comput. | 2 |
| 2020 | Guiding Attention in Sequence-to-Sequence Models for Dialogue Act PredictionabstractThe task of predicting dialog acts (DA) based on conversational dialog is a key component in the development of conversational agents. Accurately predicting DAs requires a precise modeling of both the conversation and the global tag dependencies. We leverage seq2seq approaches widely adopted in Neural Machine Translation (NMT) to improve the modelling of tag sequentiality. Seq2seq models are known to learn complex global dependencies while currently proposed approaches using linear conditional random fields (CRF) only model local tag dependencies. In this work, we introduce a seq2seq model tailored for DA classification using: a hierarchical encoder, a novel guided attention mechanism and beam search applied to both training and inference. Compared to the state of the art our model does not require handcrafted features and is trained end-to-end. Furthermore, the proposed approach achieves an unmatched accuracy score of 85% on SwDA, and state-of-the-art accuracy score of 91.6% on MRDA. Pierre Colombo, Emile Chapuis, Matteo Manica, Emmanuel Vignon, Giovanna Varni, Chloé Clavel |
AAAI | 6 |
| 2020 | The importance of fillers for text representations of speech transcriptsabstractWhile being an essential component of spoken language, fillers (e.g."um" or "uh") often remain overlooked in Spoken Language Understanding (SLU) tasks. We explore the possibility of representing them with deep contextualised embeddings, showing improvements on modelling spoken language and two downstream tasks - predicting a speaker's stance and expressed confidence. Tanvi Dinkar, Pierre Colombo, Matthieu Labeau, Chloé Clavel |
EMNLP (1) | 4 |
| 2020 | How confident are you? Exploring the role of fillers in the automatic prediction of a speaker's confidenceabstract"Fillers", example "um" in English, have been linked to the "Feeling of Another’s Knowing (FOAK)" or the listener’s perception of a speaker’s expressed confidence. Yet, in Spoken Language Processing (SLP) they remain unexplored, or overlooked as noise. We introduce a new and challenging task, that is the prediction of FOAK, which we think has widespread applicability, given the increasing popularity of automatic processing of educational and job interviews, reviews and speeches. We design a set of filler features based on linguistic literature, and investigate their potential in FOAK prediction. We show that the integration of information related to implicature meanings allows an improvement in the FOAK model and that the different functions of fillers are differently correlated with confidence. Tanvi Dinkar, Ioana Vasilescu, Catherine Pelachaud, Chloé Clavel |
ICASSP | 4 |
| 2020 | HRI-RNN: A User-Robot Dynamics-Oriented RNN for Engagement Decrease DetectionabstractNatural and fluid human-robot interaction (HRI) systems rely on the robot's ability to accurately assess the user's engagement in the interaction. Current HRI systems for engagement analysis , and more broadly emotion recognition, only consider user data while discarding robot data which, in many cases, affects the user state. We present a novel recurrent neural architecture for online detection of user engagement decrease in a spontaneous HRI setting that exploits the robot data. Our architecture models the user as a distinct party in the conversation and uses the robot data as contextual information to help assess engagement. We evaluate our approach on a real-world highly imbal-anced data set, where we observe up to 2.13% increase in F1 score compared to a standard gated recurrent unit (GRU). Asma Atamna, Chloé Clavel |
INTERSPEECH | 2 |
| 2020 | The POTUS Corpus, a Database of Weekly Addresses for the Study of Stance in Politics and Virtual AgentsabstractOne of the main challenges in the field of Embodied Conversational Agent (ECA) is to generate socially believable agents. The common strategy for agent behaviour synthesis is to rely on dedicated corpus analysis. Such a corpus is composed of multimedia files of socio-emotional behaviors which have been annotated by external observers. The underlying idea is to identify interaction information for the agent’s socio-emotional behavior by checking whether the intended socio-emotional behavior is actually perceived by humans. Then, the annotations can be used as learning classes for machine learning algorithms applied to the social signals. This paper introduces the POTUS Corpus composed of high-quality audio-video files of political addresses to the American people. Two protagonists are present in this database. First, it includes speeches of former president Barack Obama to the American people. Secondly, it provides videos of these same speeches given by a virtual agent named Rodrigue. The ECA reproduces the original address as closely as possible using social signals automatically extracted from the original one. Both are annotated for social attitudes, providing information about the stance observed in each file. It also provides the social signals automatically extracted from Obama’s addresses used to generate Rodrigue’s ones. Thomas Janssoone, Kevin Bailly, Gaël Richard, Chloé Clavel |
LREC | 4 |
| 2020 | Multimodal Analysis of Cohesion in Multi-party InteractionsabstractGroup cohesion is an emergent phenomenon that describes the tendency of the group members’ shared commitment to group tasks and the interpersonal attraction among them. This paper presents a multimodal analysis of group cohesion using a corpus of multi-party interactions. We utilize 16 two-minute segments annotated with cohesion from the AMI corpus. We define three layers of modalities: non-verbal social cues, dialogue acts and interruptions. The initial analysis is performed at the individual level and later, we combine the different modalities to observe their impact on perceived level of cohesion. Results indicate that occurrence of laughter and interruption are higher in high cohesive segments. We also observe that, dialogue acts and head nods did not have an impact on the level of cohesion by itself. However, when combined there was an impact on the perceived level of cohesion. Overall, the analysis shows that multimodal cues are crucial for accurate analysis of group cohesion. Reshmashree B. Kantharaju, Caroline Langlet, Mukesh Barange, Chloé Clavel, Catherine Pelachaud |
LREC | 4 |
| 2020 | Heavy-tailed Representations, Text Polarity Classification & Data AugmentationabstractThe dominant approaches to text representation in natural language rely on learning embeddings on massive corpora which have convenient properties such as compositionality and distance preservation. In this paper, we develop a novel method to learn a heavy-tailed embedding with desirable regularity properties regarding the distributional tails, which allows to analyze the points far away from the distribution bulk using the framework of multivariate extreme value theory. In particular, a classifier dedicated to the tails of the proposed embedding is obtained which exhibits a scale invariance property exploited in a novel text generation method for label preserving dataset augmentation. Experiments on synthetic and real text data show the relevance of the proposed framework and confirm that this method generates meaningful sentences with controllable attribute, e.g. positive or negative sentiments. Hamid Jalalzai, Pierre Colombo, Chloé Clavel, Éric Gaussier, Giovanna Varni, Emmanuel Vignon, Anne Sabourin |
NeurIPS | 3 |
| 2020 | Computational Study of Primitive Emotional Contagion in Dyadic InteractionsabstractInterpersonal human-human interaction is a dynamical exchange and coordination of social signals, feelings and emotions usually performed through and across multiple modalities such as facial expressions, gestures, and language. Developing machines able to engage humans in rich and natural interpersonal interactions requires capturing such dynamics. This paper addresses primitive emotional contagion during dyadic interactions in which roles are prefixed. Primitive emotional contagion was defined as the tendency people have to automatically mimic and synchronize their multimodal behavior during interactions and, consequently, to emotionally converge. To capture emotional contagion, a cross-recurrence based methodology that explicitly integrates short and long-term temporal dynamics through the analysis of both facial expressions and sentiment was developed. This approach is employed to assess emotional contagion at unimodal, multimodal and cross-modal levels and is evaluated on the Solid SAL-SEMAINE corpus. Interestingly, the approach is able to show the importance of the adoption of cross-modal strategies for addressing emotional contagion. Giovanna Varni, Isabelle Hupont, Chloé Clavel, Mohamed Chetouani |
IEEE Trans. Affect. Comput. | 3 |
| 2020 | Introduction to the Special Section on Computational Modeling and Understanding of Emotions in Conflictual Social InteractionsabstractInternational audience Rossana Damiano, Viviana Patti, Chloé Clavel, Paolo Rosso |
ACM Trans. Internet Techn. | 3 |
| 2019 | HireNet: A Hierarchical Attention Model for the Automatic Analysis of Asynchronous Video Job InterviewsabstractNew technologies drastically change recruitment techniques. Some research projects aim at designing interactive systems that help candidates practice job interviews. Other studies aim at the automatic detection of social signals (e.g. smile, turn of speech, etc...) in videos of job interviews. These studies are limited with respect to the number of interviews they process, but also by the fact that they only analyze simulated job interviews (e.g. students pretending to apply for a fake position). Asynchronous video interviewing tools have become mature products on the human resources market, and thus, a popular step in the recruitment process. As part of a project to help recruiters, we collected a corpus of more than 7000 candidates having asynchronous video job interviews for real positions and recording videos of themselves answering a set of questions. We propose a new hierarchical attention model called HireNet that aims at predicting the hirability of the candidates as evaluated by recruiters. In HireNet, an interview is considered as a sequence of questions and answers containing salient socials signals. Two contextual sources of information are modeled in HireNet: the words contained in the question and in the job position. Our model achieves better F1-scores than previous approaches for each modality (verbal content, audio and video). Results from early and late multimodal fusion suggest that more sophisticated fusion schemes are needed to improve on the monomodal results. Finally, some examples of moments captured by the attention mechanisms suggest our model could potentially be used to help finding key moments in an asynchronous job interview. Léo Hemamou, Ghazi Felhi, Vincent Vandenbussche, Jean-Claude Martin, Chloé Clavel |
AAAI | 5 |
| 2019 | Slices of Attention in Asynchronous Video Job InterviewsabstractThe impact of non verbal behaviour in a hiring decision remains an open question. Investigating this question is important, as it could provide a better understanding on how to train candidates for job interviews and make recruiters be aware of influential non verbal behaviour. This research has recently been accelerated due to the development of tools for the automatic analysis of social signals (facial expression detection, speech processing, etc), and the emergence of machine learning methods. However, these studies are still mainly based on hand engineered features, which imposes a limit to the discovery of influential social signals. On the other side, deep learning methods are a promising tool to discover complex patterns without the necessity of feature engineering. In this paper, we focus on studying influential non verbal social signals in asynchronous job video interviews that are discovered by deep learning methods. We use a previously published deep learning system that aims at inferring the hirability of a candidate with regard to a sequence of interview questions. One particularity of this system is the use of attention mechanisms, which aim at identifying the relevant parts of an answer. Thus, information at a fine-grained temporal level could be extracted using global (at the interview level) annotations on hirability. While most of the deep learning systems use attention mechanisms to offer a quick visualization of slices when a rise of attention occurs, we perform an in-depth analysis to understand what happens during these moments. First, we propose a methodology to automatically extract slices where there is a rise of attention (attention slices). Second, we study the content of attention slices by comparing them with randomly sampled slices. Finally, we show that they bear significantly more information for hirability than randomly sampled slices, and that such information is related to visual cues associated with anxiety and turn taking. Léo Hemamou, Ghazi Felhi, Jean-Claude Martin, Chloé Clavel |
ACII | 4 |
| 2019 | From the Token to the Review: A Hierarchical Multimodal approach to Opinion MiningabstractAlexandre Garcia, Pierre Colombo, Florence d’Alché-Buc, Slim Essid, Chloé Clavel. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Alexandre Garcia 0001, Pierre Colombo, Florence d'Alché-Buc, Slim Essid, Chloé Clavel |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Gesture Class Prediction by Recurrent Neural Network and Attention MechanismabstractOur objective is to develop a machine-learning model that allows a virtual agent to automatically perform appropriate communicative gestures. Our first step is to compute when a gesture should be performed. We express this as classification problem. We initially split the data into NoGesture class and HasGesture class. We develop a model based on recurrent neural network with attention mechanism to compute the class based on the speech prosody. We apply the model on a dialog corpus segmented into different gesture classes and gesture phases. We treat the prosody as the input sequence and the gesture classes as the output sequence. Fajrian Yunus, Chloé Clavel, Catherine Pelachaud |
IVA | 2 |
| 2018 | Attitude Classification in Adjacency Pairs of a Human-Agent Interaction with Hidden Conditional Random FieldsabstractIn this paper, the main goal is to classify, in a human-agent interaction, the attitude of the user using hidden conditional random fields. This model allows us to capture the dynamics of the interaction in the pairs of speech turns (adjacency pairs) analyzed by our system. High level linguistic features are computed at word level. The features include syntactic features, a statistical word embedding model and subjectivity lexicons. The proposed system is evaluated on the SEMAINE corpus. We obtain a Fl-score of 0.80, labeling using the most probable sequence of hidden states. Valentin Barrière, Chloé Clavel, Slim Essid |
ICASSP | 2 |
| 2018 | Detecting User's Likes and Dislikes for a Virtual Negotiating AgentabstractThis article tackles the issue of the detection of the user's likes and dislikes in a negotiation with a virtual agent for helping the creation of a model of user's preferences. We introduce a linguistic model of user's likes and dislikes as they are expressed in a negotiation context. The identification of syntactic and semantic features enables the design of formal grammars embedded in a bottom-up and rule-based system. It deals with conversational context by considering simple and collaborative likes and dislikes within adjacency pairs. We present the annotation campaign we conduct by recruiting annotators on CrowdFlower and using a dedicated annotation platform. Finally, we measure agreement between our system and the human reference. The obtained scores show substantial agreement. Caroline Langlet, Chloé Clavel |
ICMI | 2 |
| 2018 | Structured Output Learning with Abstention: Application to Accurate Opinion PredictionabstractMotivated by Supervised Opinion Analysis, we propose a novel framework devoted to Structured Output Learning with Abstention (SOLA). The structure prediction model is able to abstain from predicting some labels in the structured output at a cost chosen by the user in a flexible way. For that purpose, we decompose the problem into the learning of a pair of predictors, one devoted to structured abstention and the other, to structured output prediction. To compare fully labeled training data with predictions potentially containing abstentions, we define a wide class of asymmetric abstention-aware losses. Learning is achieved by surrogate regression in an appropriate feature space while prediction with abstention is performed by solving a new pre-image problem. Thus, SOLA extends recent ideas about Structured Output Prediction via surrogate problems and calibration theory and enjoys statistical guarantees on the resulting excess risk. Instantiated on a hierarchical abstention-aware loss, SOLA is shown to be relevant for fine-grained opinion mining and gives state-of-the-art results on this task. Moreover, the abstention-aware representations can be used to competitively predict user-review ratings based on a sentence-level opinion predictor. Alexandre Garcia 0001, Chloé Clavel, Slim Essid, Florence d'Alché-Buc |
ICML | 2 |
| 2017 | UE-HRI: a new dataset for the study of user engagement in spontaneous human-robot interactionsabstractIn this paper, we present a new dataset of spontaneous interactions between a robot and humans, of which 54 interactions (between 4 and 15-minute duration each) are freely available for download and use. Participants were recorded while holding spontaneous conversations with the robot Pepper. The conversations started automatically when the robot detected the presence of a participant and kept the recording if he/she accepted the agreement (i.e. to be recorded). Pepper was in a public space where the participants were free to start and end the interaction when they wished. The dataset provides rich streams of data that could be used by research and development groups in a variety of areas. Atef Ben Youssef, Chloé Clavel, Slim Essid, Miriam Bilac, Marine Chamoux, Angelica Lim |
ICMI | 2 |
| 2017 | Opinion Dynamics Modeling for Movie Review Transcripts Classification with Hidden Conditional Random FieldsabstractIn this paper, the main goal is to detect a movie reviewer's opinion using hidden conditional random fields. This model allows us to capture the dynamics of the reviewer's opinion in the transcripts of long unsegmented audio reviews that are analyzed by our system. High level linguistic features are computed at the level of inter-pausal segments. The features include syntactic features, a statistical word embedding model and subjectivity lexicons. The proposed system is evaluated on the ICT-MMMO corpus. We obtain a F1-score of 82\%, which is better than logistic regression and recurrent neural network approaches. We also offer a discussion that sheds some light on the capacity of our system to adapt the word embedding model learned from general written texts data to spoken movie reviews and thus model the dynamics of the opinion. Valentin Barrière, Chloé Clavel, Slim Essid |
INTERSPEECH | 2 |
| 2017 | A Web-Based Platform for Annotating Sentiment-Related Phenomena in Human-Agent Conversations
Caroline Langlet, Guillaume Dubuisson Duplessis, Chloé Clavel |
IVA | 3 |
| 2017 | Automatic Measures to Characterise Verbal Alignment in Human-Agent InteractionabstractThis work aims at characterising verbal alignment processes for improving virtual agent communicative capabilities.We propose computationally inexpensive measures of verbal alignment based on expression repetition in dyadic textual dialogues.Using these measures, we present a contrastive study between Human-Human and Human-Agent dialogues on a negotiation task.We exhibit quantitative differences in the strength and orientation of verbal alignment showing the ability of our approach to characterise important aspects of verbal alignment. Guillaume Dubuisson Duplessis, Chloé Clavel, Frédéric Landragin |
SIGDIAL Conference | 2 |
| 2017 | Affect and Interaction in Agent-Based Systems and Social Media: Guest Editors' Introductionabstract[EN] Today¿s Internet is evolving toward an open society of humans and computational\nentities, where intelligent agent systems increasingly support the interaction between\nusers and computational components. In this scenario, affect plays a key role, with functions\nthat span from creating and maintaining interpersonal relations, to establishing\ncooperation and trust. Artificial systems¿which can act as both actors and facilitators\nof these interactions¿are more and more requested to integrate affective components\nin order to achieve truly realistic social behaviors and foster the creation of bonds with\nthe users, with timely reactions to their affective input and appropriate expressions\nof affect. Achieving this integration requires to understand and reproduce the role of\naffect in human expressive capabilities, and to account for affect-related phenomena\n(e.g., sentiment, emotions, and mood) that engage social abilities, such as empathy,\nand expressive means, such as irony. The expectation for complex phenomena such as\nempathy and irony is that effective approaches require an interdisciplinary approach\nand, above all, the integration of representational models and data-oriented processing\ntechniques. The aim of this special section is to bring together leading research\non computational models of affect-related phenomena in interactions occurring either\nin social media or agent-based systems, by attaining cross-fertilization between two\nrelevant perspectives: on the one hand, research on agent architectures and cognitive models, mainly concerned with the integration of affective states into agents and open\nto the creation of virtual and embodied agents; on the other hand, research on techniques\nfor sentiment analysis and opinion mining, mainly focused on the processing\nof affective information in social media, typically (but not only) expressed through\ntext. The integration of models and methods between the two perspectives can open\nthe way to the development of a new generation of social and interactive applications\nthat leverage the affective dimension to promote improved, spontaneous technologymediated\ninteractions, including human¿computer and human¿human interactions,\non small and large scale.\nOur goal here is to provide an overview of the open research challenges for the\ncommunity of researchers interested in analyzing and modeling the interplay between\naffect and interaction in agent-based systems and social media, as well as an introduction\nto the special section. The rest of the article is structured as follows. Section 2\ndiscusses a set of research challenges relevant in the context of this special issue,\nSection 3 briefly introduces the articles included in this ACM TOIT special section.\nSection 4 concludes the article. Chloé Clavel, Rossana Damiano, Viviana Patti, Paolo Rosso |
ACM Trans. Internet Techn. | 1 |
| 2016 | Using Temporal Association Rules for the Synthesis of Embodied Conversational Agents with a Specific Stance
Thomas Janssoone, Chloé Clavel, Kevin Bailly, Gaël Richard |
IVA | 2 |
| 2016 | Grounding the detection of the user's likes and dislikes on the topic structure of human-agent interactions
Caroline Langlet, Chloé Clavel |
Knowl. Based Syst. | 2 |
| 2016 | Sentiment Analysis: From Opinion Mining to Human-Agent InteractionabstractThe opinion mining and human-agent interaction communities are currently addressing sentiment analysis from different perspectives that comprise, on the one hand, disparate sentiment-related phenomena and computational representations, and on the other hand, different detection and dialog management methods. In this paper we identify and discuss the growing opportunities for cross-disciplinary work that may increase individual advances. Sentiment/opinion detection methods used in human-agent interaction are indeed rare and, when they are employed, they are not different from the ones used in opinion mining and consequently not designed for socio-affective interactions (timing constraint of the interaction, sentiment analysis as an input and an output of interaction strategies). To support our claims, we present a comparative state of the art which analyzes the sentiment-related phenomena and the sentiment detection methods used in both communities and makes an overview of the goals of socio-affective human-agent strategies. We propose then different possibilities for mutual benefit, specifying several research tracks and discussing the open questions and prospects. To show the feasibility of the general guidelines proposed we also approach them from a specific perspective by applying them to the case of the Greta embodied conversational agents platform and discuss the way they can be used to make a more significative sentiment analysis for human-agent interactions in two different use cases: job interviews and dialogs with museum visitors. Chloé Clavel, Zoraida Callejas Carrión |
IEEE Trans. Affect. Comput. | 1 |
| 2015 | An ECA expressing appreciationsabstractIn this paper, we propose a computational model that provides an Embodied Conversational Agent (ECA) with the ability to generate verbal other-repetition (repetitions of some of the words uttered in the previous user speaker turn) when interacting with a user in a museum setting. We focus on the generation of other-repetitions expressing emotional stances in appreciation sentences. Emotional stances and their semantic features are selected according to the user's verbal input, and ECA's utterance is generated according to these features. We present an evaluation of this model through users' subjective reports. Results indicate that the expression of emotional stances by the ECA has a positive effect oIn this paper, we propose a computational model that provides an Embodied Conversational Agent (ECA) with the ability to generate verbal other-repetition (repetitions of some of the words uttered in the previous user speaker turn) when interacting with a user in a museum setting. We focus on the generation of other-repetitions expressing emotional stances in appreciation sentences. Emotional stances and their semantic features are selected according to the user's verbal input, and ECA's utterance is generated according to these features. We present an evaluation of this model through users' subjective reports. Results indicate that the expression of emotional stances by the ECA has a positive effect on user engagement, and that ECA's behaviours are rated as more believable by users when the ECA utters other-repetitions.n user engagement, and that ECA's behaviours are rated as more believable by users when the ECA utters other-repetitions. Sabrina Campano, Caroline Langlet, Nadine Glas, Chloé Clavel, Catherine Pelachaud |
ACII | 4 |
| 2015 | Adapting sentiment analysis to face-to-face human-agent interactions: From the detection to the evaluation issuesabstractThis paper introduces a sentiment analysis method suitable to the human-agent and face-to-face interactions. We present the positioning of our system and its evaluation protocol according to the existing sentiment analysis literature and detail how the proposed system integrates the human-agent interaction issues. Finally, we provide an in-depth analysis of the results obtained by the evaluation, opening the discussion on the different difficulties and the remaining challenges of sentiment analysis in human-agent interactions. Caroline Langlet, Chloé Clavel |
ACII | 2 |
| 2015 | Improving social relationships in face-to-face human-agent interactions: when the agent wants to know user's likes and dislikesabstractCaroline Langlet, Chloé Clavel. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Caroline Langlet, Chloé Clavel |
ACL (1) | 2 |
| 2014 | A CRF-based approach to automatic disfluency detection in a French call-centre corpusabstractInternational audience Camille Dutrey, Chloé Clavel, Sophie Rosset, Ioana Vasilescu, Martine Adda-Decker |
INTERSPEECH | 2 |
| 2014 | Comparative analysis of verbal alignment in human-human and human-agent interactions
Sabrina Campano, Jessica Durand, Chloé Clavel |
LREC | 3 |
| 2013 | A Multimodal Corpus Approach to the Design of Virtual RecruitersabstractThis paper presents the analysis of the multimodal behavior of experienced practitioners of job interview coaching, and describes a methodology to specify their behavior in Embodied Conversational Agents acting as virtual recruiters displaying different interpersonal stances. In a first stage, we collect a corpus of videos of job interview enactments, and we detail the coding scheme used to encode multimodal behaviors and contextual information. From the annotations of the practitioners' behaviors we observe specificities of behavior across different levels, namely monomodal behavior variations, inter-modalities behavior influences, and contextual influences on behavior. Finally we propose the adaptation of an existing agent architecture to model these specificities in a virtual recruiter's behavior. Mathieu Chollet, Magalie Ochs, Chloé Clavel, Catherine Pelachaud |
ACII | 3 |
| 2011 | Fiction support for realistic portrayals of fear-type emotional manifestations
Chloé Clavel, Ioana Vasilescu, Laurence Devillers |
Comput. Speech Lang. | 1 |
| 2010 | Combining text categorization and dialog modeling for speaker role identification on call center conversationsabstractIn this paper, we address the problem of speaker role identification on a corpus of manually transcribed call center conversations. We first tackle it as a text categorization task. Then, we combine these categorization results with a dialog modeling approach. We achieve 93% of correct role assignment with the least method. Our method also offers the possibility to extract text spans specific to each role. These strings slightly improve the role identification results and are an interesting element for conversation analysis. Rémi Lavalley, Chloé Clavel, Patrice Bellot, Marc El-Bèze |
INTERSPEECH | 2 |
| 2008 | Fear-type emotion recognition for future audio-based surveillance systems
Chloé Clavel, Ioana Vasilescu, Laurence Devillers, Gaël Richard, Thibaut Ehrette |
Speech Commun. | 1 |
| 2007 | Detection and Analysis of Abnormal Situations Through Fear-Type Acoustic ManifestationsabstractRecent work on emotional speech processing has demonstrated the interest to consider the information conveyed by the emotional component in speech to enhance the understanding of human behaviors. But to date, there has been little integration of emotion detection systems in effective applications. The present research focuses on the development of a fear-type emotions recognition system to detect and analyze abnormal situations for surveillance applications. The Fear vs. Neutral classification gets a mean accuracy rate at 70.3%. It corresponds to quite optimistic results given the diversity of fear manifestations illustrated in the data. More specific acoustic models are built inside the fear class by considering the context of emergence of the emotional manifestations, i.e. the type of the threat during which they occur, and which has a strong influence on fear acoustic manifestations. The potential use of these models for a threat type recognition system is also investigated. Such information about the situation can indeed be useful for surveillance systems. Chloé Clavel, Laurence Devillers, Gaël Richard, Ioana Vasilescu, Thibaut Ehrette |
ICASSP (4) | 1 |
| 2006 | Fear-type emotions of the SAFE Corpus: annotation issues
Chloé Clavel, Ioana Vasilescu, Laurence Devillers, Thibaut Ehrette, Gaël Richard |
LREC | 1 |
| 2005 | Events Detection for an Audio-Based Surveillance SystemabstractThe present research deals with audio events detection in noisy environments for a multimedia surveillance application. In surveillance or homeland security most of the systems aiming to automatically detect abnormal situations are only based on visual clues while, in some situations, it may be easier to detect a given event using the audio information. This is in particular the case for the class of sounds considered in this paper, sounds produced by gun shots. The automatic shot detection system presented is based on a novelty detection approach which offers a solution to detect abnormality (abnormal audio events) in continuous audio recordings of public places. We specifically focus on the robustness of the detection against variable and adverse conditions and the reduction of the false rejection rate which is particularly important in surveillance applications. In particular, we take advantage of potential similarity between the acoustic signatures of the different types of weapons by building a hierarchical classification system Chloé Clavel, Thibaut Ehrette, Gaël Richard |
ICME | 1 |
| 2004 | Fiction database for emotion detection in abnormal situationsabstractThe present research focuses on the acquisition and annotation of vocal resources for emotion detection. We are interested in detecting emotions occurring in abnormal situations and particularly in detecting ”fear”. The present study considers a preliminary database of audiovisual sequences extracted from movie fictions. The sequences selected provide various manifestations of target emotions and are described with a multimodal annotation tool. We focus on audio cues in the annotation strategy and we use the video as support for validating the audio labels. The present article deals with the description of the methodology of data acquisition and annotation. The validation of annotation is realized via two perceptual paradigms in which the +/-video condition in stimuli presentation varies. We show the perceptual significance of the audio cues and the presence of target emotions. Ioana Vasilescu, Laurence Devillers, Chloé Clavel, Thibaut Ehrette |
INTERSPEECH | 3 |