VLDB 2026 Research / reviewers in the wild / expert
Laurent Prévot 0001
dblp:16/3156
· DBLP profile ↗
40ranked-venue papers
8as first author
11since 2021 · last 2025
0000-0002-2463-2382ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 8 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Segmenting a Large French Meeting Corpus into Elementary Discourse UnitsabstractDespite growing interest in discourse-related tasks, the limited quantity and diversity of discourse-annotated data remain a major issue. Existing resources are largely based on written corpora, while spoken conversational genres are underrepresented. Although discourse segmentation into elementary discourse units (EDUs) is considered to be nearly solved for canonical written texts, conversational spontaneous speech transcripts present different challenges. In this paper, we introduce a large French corpus of segmented meeting dialogues, including 20 hours of manually transcribed and discourse-annotated conversations, and 80 hours of automatically transcribed and discourse-segmented data. We describe our annotation campaign, discuss inter-annotator agreement and segmentation guidelines, and present results from fine-tuning a model for EDU segmentation on this resource. Laurent Prévot 0001, Roxane Bertrand, Julie Hunter 0001 |
SIGDIAL | 1 |
| 2025 | Zero-Shot Evaluation of Conversational Language Competence in Data-Efficient LLMs Across English, Mandarin, and FrenchabstractLarge Language Models (LLMs) have achieved oustanding performance across various natural language processing tasks, including those from Discourse and Dialogue traditions. However, these achievements are typically obtained thanks to pretraining on huge datasets. In contrast, humans learn to speak and communicate through dialogue and spontaneous speech with only a fraction of the language exposure. This disparity has spurred interest in evaluating whether smaller, more carefully selected and curated pretraining datasets can support robust performance on specific tasks. Drawing inspiration from the BabyLM initiative, we construct small (10M-token) pretraining datasets from different sources, including conversational transcripts and Wikipedia-style text. To assess the impact of these datasets, we develop evaluation benchmarks focusing on discourse and interactional markers, extracted from high-quality spoken corpora in English, French, and Mandarin. Employing a zero-shot classification framework inspired by the BLiMP benchmark, we design tasks wherein the model must determine, between a genuine utterance extracted from a corpus and its minimally altered counterpart, which one is the authentic instance. Our findings reveal that the nature of pretraining data significantly influences model performance on discourse-related tasks. Models pretrained on conversational data exhibit a clear advantage in handling discourse and interactional markers compared to those trained on written or encyclopedic text. Furthermore, the models, trained on small amount spontaneous speech transcripts, perform comparably to standard LLMs. Sheng-Fu Wang, Ri-Sheng Huang, Shu-Kai Hsieh, Laurent Prévot 0001 |
SIGDIAL | 4 |
| 2024 | CHICA: A Developmental Corpus of Child-Caregiver's Face-to-face vs. Video Call Conversations in Middle ChildhoodabstractExisting studies of naturally occurring language-in-interaction have largely focused on the two ends of the developmental spectrum, i.e., early childhood and adulthood, leaving a gap in our knowledge about how development unfolds, especially across middle childhood. The current work contributes to filling this gap by introducing CHICA (for Child Interpersonal Communication Analysis), a developmental corpus of child-caregiver conversations at home, involving groups of French-speaking children aged 7, 9, and 11 years old. Each dyad was recorded twice: once in a face-to-face setting and once using computer-mediated video calls. For the face-to-face settings, we capitalized on recent advances in mobile, lightweight eye-tracking and head motion detection technology to optimize the naturalness of the recordings, allowing us to obtain both precise and ecologically valid data. Further, we mitigated the challenges of manual annotation by relying – to the extent possible – on automatic tools in speech processing and computer vision. Finally, to demonstrate the richness of this corpus for the study of child communicative development, we provide preliminary analyses comparing several measures of child-caregiver conversational dynamics across developmental age, modality, and communicative medium. We hope the current corpus will allow new discoveries into the properties and mechanisms of multimodal communicative development across middle childhood. Dhia-Elhak Goumri, Abhishek Agrawal, Mitja Nikolaus, Hong Duc Thang Vu, Kübra Bodur, Elias Emmar, Cassandre Armand, Chiara Mazzocconi, Shreejata Gupta, Laurent Prévot 0001, Benoît Favre, Leonor Becerra-Bonache, Abdellah Fourtassi |
LREC/COLING | 10 |
| 2024 | Conversational Feedback in Scripted versus Spontaneous Dialogues: A Comparative AnalysisabstractScripted dialogues such as movie and TV subtitles constitute a widespread source of training data for conversational NLP models.However, there are notable linguistic differences between these dialogues and spontaneous interactions, especially regarding the occurrence of communicative feedback such as backchannels, acknowledgments, or clarification requests.This paper presents a quantitative analysis of such feedback phenomena in both subtitles and spontaneous conversations.Based on conversational data spanning eight languages and multiple genres, we extract lexical statistics, classifications from a dialogue act tagger, expert annotations and labels derived from a fine-tuned Large Language Model (LLM).Our main empirical findings are that (1) communicative feedback is markedly less frequent in subtitles than in spontaneous dialogues and (2) subtitles contain a higher proportion of negative feedback.We also show that dialogues generated by standard LLMs lie much closer to scripted dialogues than spontaneous interactions in terms of communicative feedback. Ildikó Pilán, Laurent Prévot 0001, Hendrik Buschmeier, Pierre Lison |
SIGDIAL | 2 |
| 2023 | Transcribing and Aligning Conversational Speech: A Hybrid Pipeline Applied to French ConversationsabstractWith the advent of transformer based models, the use of fully automated ASR-pipelines in a real-world context has come close to a reality. However, for faithful transcription of conversational speech, there remain challenges both in terms of the content predicted by these models (hallucinations, unintended normalizations of disfluencies and transcriptions of background noises) and in terms of alignment accuracy. In this paper we present a hybrid ASR-pipeline which augments transformer models with other algorithms in order to transcribe conversational data. Through experiments on two French datasets, we show that: 1) VAD preprocessing can significantly improve transcription quality as well as word level temporal alignment, 2) prompting can reduce unintended normalizations of disfluencies, 3) heuristic-based detection of untranscribed sounds can further improve alignment quality. We conclude that our hybrid pipeline is an efficient way to improve and augment existing ASR-models. Hiroyoshi Yamasaki, Jérôme Louradour, Julie Hunter 0001, Laurent Prévot 0001 |
ASRU | 4 |
| 2023 | Communicative Feedback in Response to Children's Grammatical Errors
Mitja Nikolaus, Laurent Prévot 0001, Abdellah Fourtassi |
CogSci | 2 |
| 2023 | Bringing Together Ergonomic Concepts and Cognitive Mechanisms for Human - AI Agents CooperationabstractThe deployment of artificial intelligence from experimental settings to concrete applications implies to consider the social aspects of the environment and consequently to conceive the interaction between humans and computers endowed with the aim of being partners in action. This article proposes a review of the research initiatives regarding human-artificial agents interaction, including eXplainable Artificial Intelligence (XAI) and HRI/HCI. We argue that even if vocabulary and approaches are different, the concepts converge on the necessity for the artificial agents to provide an accurate mental model of their behavior to the humans they are interacting with. This has different implications depending on whether we consider a tool/user interaction or a cooperation interaction—which is far less documented despite being at the heart of the future concepts of autonomous vehicles. From this observation, the article uses the cognitive science corpus on joint-action to raise finer cognitive mechanisms proved to be essential for human joint-action which could be considered as cognitive requirements for future artificial agents, including shared task representation and mentalization. Finally, interactions content hypotheses are arisen to satisfy the identified mechanisms, including the ability for the artificial agent to elicit its intentions and to trigger mentalization toward them from the human cooperators. Marin Le Guillou, Laurent Prévot 0001, Bruno Berberian |
Int. J. Hum. Comput. Interact. | 2 |
| 2022 | Backchannel Behavior in Child-Caregiver Zoom-Mediated Conversations
Kübra Bodur, Mitja Nikolaus, Abdellah Fourtassi, Laurent Prévot 0001 |
CogSci | 4 |
| 2022 | Communicative Feedback as a Mechanism Supporting the Production of Intelligible Speech in Early Childhood
Mitja Nikolaus, Laurent Prévot 0001, Abdellah Fourtassi |
CogSci | 2 |
| 2022 | Listen and tell me who the user is talking to: Automatic detection of the interlocutor's type during a conversationabstractIn the well-known Turing test, humans have to judge whether they write to another human or a chatbot. In this article, we propose a reversed Turing test adapted to live conversations: based on the speech of the human, we have developed a model that automatically detects whether she/he speaks to an artificial agent or a human. We propose in this work a prediction methodology combining a step of specific features extraction from behaviour and a specific deep learning model based on recurrent neural networks. The prediction results show that our approach, and more particularly the considered features, improves significantly the predictions compared to the traditional approach in the field of automatic speech recognition systems, which is based on spectral features, such as Mel-frequency Cepstral Coefficients (MFCCs). Our approach allows evaluating automatically the type of conversational agent, human or artificial agent, solely based on the speech of the human interlocutor. Most importantly, this model provides a novel and very promising approach to weigh the importance of the behaviour cues used to make correctly recognize the nature of the interlocutor, in other words, what aspects of the human behaviour adapts to the nature of its interlocutor. Youssef Hmamouche, Magalie Ochs, Thierry Chaminade, Laurent Prévot 0001 |
RO-MAN | 4 |
| 2021 | Large-scale study of speech acts' development using automatic labelling
Mitja Nikolaus, Juliette Maes, Jérémy Auguste, Laurent Prévot 0001, Abdellah Fourtassi |
CogSci | 4 |
| 2020 | Exploring the Dependencies between Behavioral and Neuro-physiological Time-series Extracted from Conversations between Humans and Artificial AgentsabstractInternational audience Youssef Hmamouche, Magalie Ochs, Laurent Prévot 0001, Thierry Chaminade |
ICPRAM | 3 |
| 2020 | Neural Representations of Dialogical History for Improving Upcoming Turn Acoustic Parameters PredictionabstractInternational audience Simone Fuscone, Benoît Favre, Laurent Prévot 0001 |
INTERSPEECH | 3 |
| 2020 | Identifying Causal Relationships Between Behavior and Local Brain Activity During Natural ConversationabstractInternational audience Youssef Hmamouche, Laurent Prévot 0001, Magalie Ochs, Thierry Chaminade |
INTERSPEECH | 2 |
| 2020 | The ISO Standard for Dialogue Act Annotation, Second EditionabstractISO standard 24617-2 for dialogue act annotation, established in 2012, has in the past few years been used both in corpus annotation and in the design of components for spoken and multimodal dialogue systems. This has brought some inaccuracies and undesirbale limitations of the standard to light, which are addressed in a proposed second edition. This second edition allows a more accurate annotation of dependence relations and rhetorical relations in dialogue. Following the ISO 24617-4 principles of semantic annotation, and borrowing ideas from EmotionML, a triple-layered plug-in mechanism is introduced which allows dialogue act descriptions to be enriched with information about their semantic content, about accompanying emotions, and other information, and allows the annotation scheme to be customised by adding application-specific dialogue act types. Harry Bunt, Volha Petukhova, Emer Gilmartin, Catherine Pelachaud, Alex Chengyu Fang, Simon Keizer, Laurent Prévot 0001 |
LREC | 7 |
| 2020 | BrainPredict: a Tool for Predicting and Visualising Local Brain ActivityabstractIn this paper, we present a tool allowing dynamic prediction and visualization of an individual’s local brain activity during a conversation. The prediction module of this tool is based on classifiers trained using a corpus of human-human and human-robot conversations including fMRI recordings. More precisely, the module takes as input behavioral features computed from raw data, mainly the participant and the interlocutor speech but also the participant’s visual input and eye movements. The visualisation module shows in real-time the dynamics of brain active areas synchronised with the behavioral raw data. In addition, it shows which integrated behavioral features are used to predict the activity in individual brain areas. Youssef Hmamouche, Laurent Prévot 0001, Magalie Ochs, Thierry Chaminade |
LREC | 2 |
| 2020 | Multimodal Corpus of Bidirectional Conversation of Human-human and Human-robot Interaction during fMRI ScanningabstractIn this paper we present investigation of real-life, bi-directional conversations. We introduce the multimodal corpus derived from these natural conversations alternating between human-human and human-robot interactions. The human-robot interactions were used as a control condition for the social nature of the human-human conversations. The experimental set up consisted of conversations between the participant in a functional magnetic resonance imaging (fMRI) scanner and a human confederate or conversational robot outside the scanner room, connected via bidirectional audio and unidirectional videoconferencing (from the outside to inside the scanner). A cover story provided a framework for natural, real-life conversations about images of an advertisement campaign. During the conversations we collected a multimodal corpus for a comprehensive characterization of bi-directional conversations. In this paper we introduce this multimodal corpus which includes neural data from functional magnetic resonance imaging (fMRI), physiological data (blood flow pulse and respiration), transcribed conversational data, as well as face and eye-tracking recordings. Thus, we present a unique corpus to study human conversations including neural, physiological and behavioral data. Birgit Rauchbauer, Youssef Hmamouche, Brigitte Bigi, Laurent Prévot 0001, Magalie Ochs, Thierry Chaminade |
LREC | 4 |
| 2020 | Exploiting weak-supervision for classifying Non-Sentential Utterances in Mandarin Conversations
Xin-Yi Chen, Laurent Prévot 0001 |
PACLIC | 2 |
| 2020 | Filtering conversations through dialogue acts labels for improving corpus-based convergence studiesabstractDuring an interaction the tendency of speakers to change their speech production to make it more similar to their interlocutor's speech is called convergence.Convergence had been studied due to its relevance for cognitive models of communication as well as for dialogue system adaptation to the user.Convergence effects have been established on controlled data sets while tracking its dynamics on generic corpora has provided positive but more contrasted outcomes.We propose to enrich large conversational corpora with dialogue acts information and to use these acts as filters to create subsets of homogeneous conversational activity.Those subsets allow a more precise comparison between speakers' speech variables.We compare convergence on acoustic variables (Energy, Pitch and Speech Rate) measured on raw data sets, with human and automatically data sets labelled with dialog acts type.We found that such filtering helps in observing convergence suggesting that future studies should consider such high level dialogue activity types and the related NLP techniques as important tools for analyzing conversational interpersonal dynamics. Simone Fuscone, Benoît Favre, Laurent Prévot 0001 |
SIGdial | 3 |
| 2018 | Brain Neurophysiology to Objectify the Social Competence of Conversational AgentsabstractWe present an approach to objectify the social competence of artificial agents using human brain neurophysiology. Whole brain activity is recorded with functional Magnetic Resonance Imaging (fMRI) while participants discuss either with a human confederate or an artificial agent. This allows a direct comparison of local brain responses, including deep brain structures invisible to other neuroimaging techniques, as a function of the nature of the interlocutor. The present data (9 participants, artificial agent is the robotic conversational head Furhat controlled with a Wizard of Oz procedure) demonstrates the feasibility of this approach, and results confirm an increased activity in the hypothalamic region when interacting with a human compared to an artificial agent. Thierry Chaminade, Birgit Rauchbauer, Bruno Nazarian, Morgane Bourhis, Magalie Ochs, Laurent Prévot 0001 |
HAI | 6 |
| 2017 | Studying the Link Between Inter-Speaker Coordination and Speech Imitation Through Human-Machine InteractionsabstractInternational audience Leonardo Lancia, Thierry Chaminade, Noël Nguyen, Laurent Prévot 0001 |
INTERSPEECH | 4 |
| 2016 | 4Couv: A New Treebank for French
Philippe Blache, Grégoire de Montcheuil, Laurent Prévot 0001, Stéphane Rauzy |
LREC | 3 |
| 2016 | A CUP of CoFee: A large Collection of feedback Utterances Provided with communicative function annotations
Laurent Prévot 0001, Jan Gorisch, Roxane Bertrand |
LREC | 1 |
| 2016 | LexFr: Adapting the LexIt Framework to Build a Corpus-based French Subcategorization Lexicon
Giulia Rambelli, Gianluca Lebani, Laurent Prévot 0001, Alessandro Lenci |
LREC | 3 |
| 2015 | Audio synchronisation with a tunnel matrix for time series and dynamic programmingabstractPrecise multimodal studies require precise synchronisation between audio and video signals. However, raw audio and audio from video recordings can be out of sync for several reasons. In order to re-synchronise them, a dynamic programming (DP) approach is presented here. Traditionally, DP is performed on the rectangular distance matrix comparing each value in signal A with each value in signal B. Previous work limited the search space using for example the Sakoe Chiba Band (Sakoe and Chiba, 1978). However, the overall space of the distance matrix remains identical. Here, a tunnel matrix and its according DP-algorithm are presented. The matrix contains merely the computed distance of two signals to a pre-specified bandwidth and the computational cost is equally reduced. An example implementation demonstrates the functionality on artificial data and on data from real audio and video recordings. Jan Gorisch, Laurent Prévot 0001 |
ICASSP | 2 |
| 2015 | Annotation and Classification of French Feedback Communicative Functions
Laurent Prévot 0001, Jan Gorisch, Sankar Mukherjee |
PACLIC | 1 |
| 2015 | A SIP of CoFee : A Sample of Interesting Productions of Conversational FeedbackabstractFeedback utterances are among the most frequent in dialogue.Feedback is also a crucial aspect of linguistic theories that take social interaction, involving language, into account.This paper introduces the corpora and datasets of a project scrutinizing this kind of feedback utterances in French.We present the genesis of the corpora (for a total of about 16 hours of transcribed and phone force-aligned speech) involved in the project.We introduce the resulting datasets and discuss how they are being used in on-going work with focus on the form-function relationship of conversational feedback.All the corpora created and the datasets produced in the framework of this project will be made available for research purposes. Laurent Prévot 0001, Jan Gorisch, Roxane Bertrand, Emilien Gorene, Brigitte Bigi |
SIGDIAL Conference | 1 |
| 2014 | Representing Multimodal Linguistic Annotated data
Brigitte Bigi, Tatsuya Watanabe, Laurent Prévot 0001 |
LREC | 3 |
| 2014 | Aix Map Task corpus: The French multimodal corpus of task-oriented dialogue
Jan Gorisch, Corine Astésano, Ellen Gurman Bard, Brigitte Bigi, Laurent Prévot 0001 |
LREC | 5 |
| 2014 | Segmentation evaluation metrics, a comparison grounded on prosodic and discourse units
Klim Peshkov, Laurent Prévot 0001 |
LREC | 2 |
| 2013 | A Quantitative Comparative Study of Prosodic and Discourse Units, the Case of French and Taiwan Mandarin
Laurent Prévot 0001, Shu-Chuan Tseng, Alvin Cheng-Hsien Chen, Klim Peshkov |
PACLIC | 1 |
| 2013 | A quantitative view of feedback lexical markers in conversational French
Laurent Prévot 0001, Brigitte Bigi, Roxane Bertrand |
SIGDIAL Conference | 1 |
| 2012 | An empirical resource for discovering cognitive principles of discourse organisation: the ANNODIS corpus
Stergos D. Afantenos, Nicholas Asher, Farah Benamara, Myriam Bras, Cécile Fabre, Lydia-Mai Ho-Dac, Anne Le Draoulec, Philippe Muller, Marie-Paule Péry-Woodley, Laurent Prévot 0001, Josette Rebeyrolle, Ludovic Tanguy, Marianne Vergez-Couret, Laure Vieu |
LREC | 10 |
| 2010 | The OTIM Formal Annotation Model: A Preliminary Step before Annotation Scheme
Philippe Blache, Roxane Bertrand, Mathilde Guardiola, Marie-Laure Guénot, Christine Meunier, Irina Nesterenko, Berthille Pallaud, Laurent Prévot 0001, Béatrice Priego-Valverde, Stéphane Rauzy |
LREC | 8 |
| 2010 | Computational Modeling of Verb Acquisition, from a Monolingual to a Bilingual Study
Laurent Prévot 0001, Chun-Han Chang, Yann Desalle |
PACLIC | 1 |
| 2009 | Using Extra-Linguistic Material for Mandarin-French Verbal Constructions Comparison
Pierre Magistry, Laurent Prévot 0001, Hintat Cheung, Chien-yun Shiao, Yann Desalle, Bruno Gaume |
PACLIC | 2 |
| 2008 | Extracting Concrete Senses of Lexicon through Measurement of Conceptual Similarity in Ontologies
Siaw-Fong Chung, Laurent Prévot 0001, Kathleen Ahrens, Shu-Kai Hsieh, Chu-Ren Huang |
LREC | 2 |
| 2007 | Rethinking Chinese Word Segmentation: Tokenization, Character Classification, or Wordbreak Identification
Chu-Ren Huang, Petr Simon, Shu-Kai Hsieh, Laurent Prévot 0001 |
ACL | 4 |
| 2006 | Infrastructure for Standardization of Asian Language Resources
Takenobu Tokunaga, Virach Sornlertlamvanich, Thatsanee Charoenporn, Nicoletta Calzolari, Monica Monachini, Claudia Soria, Chu-Ren Huang, Yingju Xia, Hao Yu 0005, Laurent Prévot 0001, Kiyoaki Shirai |
ACL | 10 |
| 2006 | Using the Swadesh list for creating a simple common taxonomy
Laurent Prévot 0001, Chu-Ren Huang, I-Li Su |
PACLIC | 1 |