VLDB 2026 Research / reviewers in the wild / expert
Jan de Wit
dblp:179/5350 · also Jan M. S. de Wit
· DBLP profile ↗
16ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-7299-7992ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 13 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | How Well Can Large Language Models Reflect? A Human Evaluation of LLM-generated Reflections for Motivational Interviewing DialoguesabstractMotivational Interviewing (MI) is a counseling technique that promotes behavioral change through reflective responses to mirror or refine client statements. While advanced Large Language Models (LLMs) can generate engaging dialogues, challenges remain for applying them in a sensitive context such as MI. This work assesses the potential of LLMs to generate MI reflections via three LLMs: GPT-4, Llama-2, and BLOOM, and explores the effect of dialogue context size and integration of MI strategies for reflection generation by LLMs. We conduct evaluations using both automatic metrics and human judges on four criteria: appropriateness, relevance, engagement, and naturalness, to assess whether these LLMs can accurately generate the nuanced therapeutic communication required in MI. While we demonstrate LLMs’ potential in generating MI reflections comparable to human therapists, content analysis shows that significant challenges remain. By identifying the strengths and limitations of LLMs in generating empathetic and contextually appropriate reflections in MI, this work contributes to the ongoing dialogue in enhancing LLM’s role in therapeutic counseling. Mustafa Erkan Basar, Xin Sun 0016, Iris Hendrickx, Jan de Wit, Tibor Bosse, Gert-Jan de Bruijn, Jos A. Bosch, Emiel Krahmer |
COLING | 4 |
| 2025 | Rethinking the Alignment of Psychotherapy Dialogue Generation with Motivational Interviewing StrategiesabstractRecent advancements in large language models (LLMs) have shown promise in generating psychotherapeutic dialogues, particularly in the context of motivational interviewing (MI). However, the inherent lack of transparency in LLM outputs presents significant challenges given the sensitive nature of psychotherapy. Applying MI strategies, a set of MI skills, to generate more controllable therapeutic-adherent conversations with explainability provides a possible solution. In this work, we explore the alignment of LLMs with MI strategies by first prompting the LLMs to predict the appropriate strategies as reasoning and then utilizing these strategies to guide the subsequent dialogue generation. We seek to investigate whether such alignment leads to more controllable and explainable generations. Multiple experiments including automatic and human evaluations are conducted to validate the effectiveness of MI strategies in aligning psychotherapy dialogue generation. Our findings demonstrate the potential of LLMs in producing strategically aligned dialogues and suggest directions for practical applications in psychotherapeutic settings. Xin Sun 0016, Abdallah El Ali, Zhuying Li 0001, Pengjie Ren, Jan de Wit, Jiahuan Pei, Jos A. Bosch |
COLING | 6 |
| 2025 | Script-Strategy Aligned Generation: Aligning LLMs with Expert-Crafted Dialogue Scripts and Therapeutic Strategies for PsychotherapyabstractChatbots or conversational agents (CAs) are increasingly used to improve access to digital psychotherapy. Many current systems rely on rigid, rule-based designs, heavily dependent on expert-crafted dialogue scripts for guiding therapeutic conversations. Although advances in large language models (LLMs) offer potential for more flexible interactions, their lack of controllability and explanability poses challenges in high-stakes contexts like psychotherapy. To address this, we conducted two studies in this work to explore how aligning LLMs with expert-crafted scripts can enhance psychotherapeutic chatbot performance. In Study 1 (N=43), an online experiment with a within-subjects design, we compared rule-based, pure LLM, and LLMs aligned with expert-crafted scripts via fine-tuning and prompting. Results showed that aligned LLMs significantly outperformed the other types of chatbots in empathy, dialogue relevance, and adherence to therapeutic principles. Building on findings, we proposed ''Script-Strategy Aligned Generation (SSAG)'', a more flexible alignment approach that reduces reliance on fully scripted content while maintaining LLMs' therapeutic adherence and controllability. In a 10-day field Study 2 (N=21), SSAG achieved comparable therapeutic effectiveness to full-scripted LLMs while requiring less than 40% of expert-crafted dialogue content. Beyond these results, this work advances LLM applications in psychotherapy by providing a controllable and scalable solution, reducing reliance on expert effort. By enabling domain experts to align LLMs through high-level strategies rather than full scripts, SSAG supports more efficient co-development and expands access to a broader context of psychotherapy. Xin Sun 0016, Jan de Wit, Zhuying Li 0001, Jiahuan Pei, Abdallah El Ali, Jos A. Bosch |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2024 | Eliciting Motivational Interviewing Skill Codes in Psychotherapy with LLMs: A Bilingual Dataset and Analytical StudyabstractBehavioral coding (BC) in motivational interviewing (MI) holds great potential for enhancing the efficacy of MI counseling. However, manual coding is labor-intensive, and automation efforts are hindered by the lack of data due to the privacy of psychotherapy. To address these challenges, we introduce BiMISC, a bilingual dataset of MI conversations in English and Dutch, sourced from real counseling sessions. Expert annotations in BiMISC adhere strictly to the motivational interviewing skills code (MISC) scheme, offering a pivotal resource for MI research. Additionally, we present a novel approach to elicit the MISC expertise from Large language models (LLMs) for MI coding. Through the in-depth analysis of BiMISC and the evaluation of our proposed approach, we demonstrate that the LLM-based approach yields results closely aligned with expert annotations and maintains consistent performance across different languages. Our contributions not only furnish the MI community with a valuable bilingual dataset but also spotlight the potential of LLMs in MI coding, laying the foundation for future MI research. Xin Sun 0016, Jiahuan Pei, Jan de Wit, Mohammad Aliannejadi, Emiel Krahmer, Jos T. P. Dobber, Jos A. Bosch |
LREC/COLING | 3 |
| 2024 | Digital Confessions: The Willingness to Disclose Intimate Information to a Chatbot and its Impact on Emotional Well-BeingabstractAbstract Chatbots have several features that may stimulate self-disclosure, such as accessibility, anonymity, convenience and their perceived non-judgmental nature. The aim of this study is to investigate if people disclose (more) intimate information to a chatbot, compared to a human, and to what extent this enhances their emotional well-being through feelings of relief. An experiment with a 2 (human vs. chatbot) by 2 (low empathetic vs. high empathetic) design was conducted (N = 286). Results showed that there was no difference in the self-reported intimacy of self-disclosure between the human and chatbot conditions. Furthermore, people perceived less fear of judgment in the chatbot condition, but more trust in the human interactant compared to the chatbot interactant. Perceived anonymity was the only variable to directly impact self-disclosure intimacy. The finding that humans disclose equally intimate information to chatbots and humans is in line with the CASA paradigm, which states that people can react in a social manner to both computers and humans. Emmelyn A. J. Croes, Marjolijn L. Antheunis, Chris van der Lee, Jan de Wit |
Interact. Comput. | 4 |
| 2023 | Co-Designing with a Social Robot Facilitator: Effects of Robot Mood Expression on Human Group DynamicsabstractSocial robots can be designed to support the facilitation of co-design sessions. Facilitators regulate group dynamics to promote effective collaboration among stakeholders. Group dynamics are sensitive to mood expressions: Positive mood expressions by the facilitator promote cooperation in the group, whereas negative expressions promote conflict. However, whether mood expressions by a social robot facilitator also influence human group dynamics is an open scientific and practical question. To learn more, an experiment (N = 98) was conducted where small groups engaged in a co-design session led by a social robot facilitator. The robot displayed positive, neutral, or negative mood expressions throughout the session. The results showed that positive robot expressions, compared to neutral or negative expressions, increased perceived robot valence. Perceived robot valence increased cooperation and decreased conflict in the human groups. These findings contribute novel insight into how social robots can be used to innovate how co-design is facilitated. Alwin de Rooij, Simone van den Broek, Michelle Bouw, Jan de Wit |
HAI | 4 |
| 2023 | The Design and Observed Effects of Robot-performed Manual Gestures: A Systematic ReviewabstractCommunication using manual (hand) gestures is considered a defining property of social robots, and their physical embodiment and presence, therefore, we see a need for a comprehensive overview of the state-of-the-art in social robots that use gestures. This systematic literature review aims to address this need by (1) describing the gesture production process of a social robot, including the design and planning steps, and (2) providing a survey of the effects of robot-performed gestures on human-robot interactions in a multitude of domains. We identify patterns and themes from the existing body of literature, resulting in nine outstanding questions for research on robot-performed gestures regarding: developments in sensor technology and AI, structuring the gesture design and evaluation process, the relationship between physical appearance and gestures, the effects of planning on the overall interaction, standardizing measurements of gesture “quality,” individual differences, gesture mirroring, whether human-likeness is desirable, and universal accessibility of robots. We also reflect on current methodological practices in studies of robot-performed gestures and suggest improvements regarding replicability, external validity, measurement instruments used, and connections with other disciplines. These outstanding questions and methodological suggestions can guide future work in this field of research. Jan de Wit, Paul Vogt, Emiel Krahmer |
ACM Trans. Hum. Robot Interact. | 1 |
| 2021 | Designing and Evaluating Iconic Gestures for Child-Robot Second Language LearningabstractAbstract In this paper, we examine the process of designing robot-performed iconic hand gestures in the context of a long-term study into second language tutoring with children of approximately 5 years old. We explore four factors that may relate to their efficacy in supporting second language tutoring: the age of participating children; differences between gestures for various semantic categories, e.g. measurement words, such as small, versus counting words, such as five; the quality (comprehensibility) of the robot’s gestures; and spontaneous reenactment or imitation of the gestures. Age was found to relate to children’s learning outcomes, with older children benefiting more from the robot’s iconic gestures than younger children, particularly for measurement words. We found no conclusive evidence that the quality of the gestures or spontaneous reenactment of said gestures related to learning outcomes. We further propose several improvements to the process of designing and implementing a robot’s iconic gesture repertoire. Jan de Wit, Bram Willemsen, Mirjam de Haas, Rianne van den Berghe, Paul M. Leseman, Ora Oudgenoeg-Paz, Josje Verhagen, Paul Vogt, Emiel Krahmer |
Interact. Comput. | 1 |
| 2020 | Using Self-Determination Theory in Social Robots to Increase Motivation in L2 Word LearningabstractThis study presents a second language word learning experiment using a social robot with motivational strategies. These strategies were implemented in a social robot tutor to stimulate preschool children's intrinsic motivation. Subsequently, we investigated their effect on children's task engagement and word learning performance. The strategies were derived from the Self-Determination Theory, a well-known psychological theory that assumes that intrinsic motivation is strongly related to the fulfilment of three basic human needs, namely the need for autonomy, competence, and relatedness. We found an increase in the strength and duration of task engagement when all three psychological needs were supported by the robot. However, no significant results for learning gains were observed. Our intervention appears a promising method for improving child-robot interactions in educational settings, especially to sustain in long-term interactions. Peggy van Minkelen, Carmen Gruson, Pleun van Hees, Mirle Willems, Jan de Wit, Rian Aarts, Jaap Denissen, Paul Vogt |
HRI | 5 |
| 2020 | Varied Human-Like Gestures for Social Robots: Investigating the Effects on Children's Engagement and Language LearningabstractTo investigate whether a humanoid robot's use of gestures improves children's learning of second language vocabulary, and if variation in gestures strengthens this effect, we conducted a field study where a total of 94 children (aged 4-6 years old) played a language learning game with a NAO robot. The robot either used no gestures at all, repeated the same gesture every time a target word was presented, or produced a different gesture for each occurrence of a target word. We found that, contrary to what the majority of existing research suggests, the robot's use of gestures did not result in increased learning outcomes, compared to a robot that did not use gestures. However, engagement between child and robot was higher in both the repeated and varied gesture conditions, compared to the condition without gestures. An exploratory analysis showed that age played a role: the older children in the study learned more than the younger children when the robot used gestures. It is therefore important to carefully consider the design and application of robot gestures to support the learning process. The contribution of this work is twofold: it is a conceptual reproduction of a previous study, and we have taken first steps towards exploring the role of variation in gestures. The study was preregistered, and all materials are made publicly available. Jan de Wit, Arold Brandse, Emiel Krahmer, Paul Vogt |
HRI | 1 |
| 2019 | Robots for Learning - R4L: Adaptive LearningabstractThe Robots for Learning workshop series aims at advancing the research topics related to the use of social robots in educational contexts. This year's half-day workshop follows on previous events in Human-Robot Interaction conferences focusing on efforts to design, develop and test new robotics systems that help learners. This 5th edition of the workshop will be dealing in particular on the potential use of robots for adaptive learning. Since the past few years, inclusive education have been a key policy in a number of countries, aiming to provide equal changes and common ground to all. In this workshop, we aim to discuss strategies to design robotics system able to adapt to the learners' abilities, to provide assistance and to demonstrate long-term learning effects. Wafa Johal, Anara Sandygulova, Jan de Wit, Mirjam de Haas, Brian Scassellati |
HRI | 3 |
| 2019 | Second Language Tutoring Using Social Robots: L2TOR - The MovieabstractThis video illustrates the large-scale experiment of the L2TOR project that will be presented at the HRI 2019 conference. The experiment aimed to investigate how 192 Dutch 5-year-old children could learn 34 English words from a NAO robot in 7 lessons. The experiment compared 4 conditions: 1) robot using iconic gestures, 2) robot without iconic gestures, 3) tablet only, and 4) a control group. The results revealed that children could learn more English words in all experimental conditions compared to the control group. The three experimental conditions did not show any significant differences regarding the learning outcomes. Paul Vogt, Rianne van den Berghe, Mirjam de Haas, Laura Kunold, Junko Kanero, Ezgi Mamus, Jean-Marc Montanier, Cansu Oranç, Ora Oudgenoeg-Paz, Daniel Hernández García, Fotios Papadopoulos, Thorsten Schodde, Josje Verhagen, Christopher D. Wallbridge, Bram Willemsen, Jan de Wit, Tony Belpaeme, Tilbe Göksun, Stefan Kopp, Emiel Krahmer, Aylin C. Küntay, Paul M. Leseman, Amit Kumar Pandey |
HRI | 16 |
| 2019 | Second Language Tutoring Using Social Robots: A Large-Scale StudyabstractWe present a large-scale study of a series of seven lessons designed to help young children learn English vocabulary as a foreign language using a social robot. The experiment was designed to investigate 1) the effectiveness of a social robot teaching children new words over the course of multiple interactions (supported by a tablet), 2) the added benefit of a robot's iconic gestures on word learning and retention, and 3) the effect of learning from a robot tutor accompanied by a tablet versus learning from a tablet application alone. For reasons of transparency, the research questions, hypotheses and methods were preregistered. With a sample size of 194 children, our study was statistically well-powered. Our findings demonstrate that children are able to acquire and retain English vocabulary words taught by a robot tutor to a similar extent as when they are taught by a tablet application. In addition, we found no beneficial effect of a robot's iconic gestures on learning gains. Paul Vogt, Rianne van den Berghe, Mirjam de Haas, Laura Kunold, Junko Kanero, Ezgi Mamus, Jean-Marc Montanier, Cansu Oranç, Ora Oudgenoeg-Paz, Daniel Hernández García, Fotios Papadopoulos, Thorsten Schodde, Josje Verhagen, Christopher D. Wallbridge, Bram Willemsen, Jan de Wit, Tony Belpaeme, Tilbe Göksun, Stefan Kopp, Emiel Krahmer, Aylin C. Küntay, Paul M. Leseman, Amit Kumar Pandey |
HRI | 16 |
| 2019 | Playing Charades with a Robot: Collecting a Large Dataset of Human Gestures Through HRIabstractThis work documents a playful human-robot interaction, in the form of a game of charades, through which a humanoid robot is able to learn how to produce and recognize gestures by interacting with human participants. We describe an extensive dataset of gesture recordings, which can be used for future research into gestures, specifically for human-robot interaction applications. Jan de Wit, Bram Willemsen, Mirjam de Haas, Emiel Krahmer, Paul Vogt, Marije Merckens, Reinjet Oostdijk, Chani Savelberg, Sabine Verdult, Pieter Wolfert |
HRI | 1 |
| 2018 | The Effect of a Robot's Gestures and Adaptive Tutoring on Children's Acquisition of Second Language VocabulariesabstractThis paper presents a study in which children, four to six years old, were taught words in a second language by a robot tutor. The goal is to evaluate two ways for a robot to provide scaffolding for students: the use of iconic gestures, combined with adaptively choosing the next learning task based on the child»s past performance. The results show a positive effect on long-term memorization of novel words, and an overall higher level of engagement during the learning activities when gestures are used. The adaptive tutoring strategy reduces the extent to which the level of engagement is diminishing during the later part of the interaction. Jan de Wit, Thorsten Schodde, Bram Willemsen, Kirsten Bergmann, Mirjam de Haas, Stefan Kopp, Emiel Krahmer, Paul Vogt |
HRI | 1 |
| 2017 | Design and Evaluation of RaPIDO, A Platform for Rapid Prototyping of Interactive Outdoor GamesabstractOutdoor, multi-player games involving social interaction and physical activity are an emerging class of applications particularly interesting for children, for whom the attraction and the health and developmental benefits are clear cut. Implementing and prototyping such games present non-trivial technical challenges to interaction and game designers; this hampers iterative prototyping and testing cycles that are core to user-centred design and game development processes. This insight has motivated the development of RaPIDO (Rapid prototyping of Physical Interaction Design for Outdoor games), a prototyping platform for physical computing, targeting interaction designers with limited electronics or software skills. RaPIDO has been evaluated in a user test, evaluating RaPIDOs software library, and in a case study involving two designers who used it to develop outdoor games for children. We illustrate how RaPIDO enabled broader exploration of the design space and faster iterations than would otherwise be possible, allowing designers to focus on the core game concepts rather than complex and low-level engineering issues. Iris Soute, Tudor Vacaretu, Jan de Wit, Panos Markopoulos 0001 |
ACM Trans. Comput. Hum. Interact. | 3 |