Khiet P. Truong

dblp:26/5664 · DBLP profile ↗
← Back
68ranked-venue papers
18as first author
12since 2021 · last 2026
0000-0002-7243-0523ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 45 · 15 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 17 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 26 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Age Against the Machine: How Age Relates to Listeners' Ability to Recognize Emotions in Robots' Semantic-Free Utterances
abstract
Semantic-Free Utterances (SFUs, sounds conveying intention without using words) are being increasingly adopted for human-robot interaction (HRI) to communicate affect. In healthcare, where older adults are overrepresented, affective robotics are becoming more common to reduce healthcare professionals' workload. Hence, understanding how older adults perceive and communicate with robots is crucial. Although previous studies have demonstrated a decline in older adults' ability to categorize emotions, it remains unclear how this impacts their comprehension of SFUs used in HRI. This paper investigates the effect of age (and other factors) on listeners' ability to categorize emotions in SFUs designed for HRI. Additionally, we explore listeners' preferences of SFUs for a healthcare robot. Listeners indicated that SFUs' similarities to natural language, the need for a distinction between human and robot, and their expectations of how a hospital robot should sound like, influenced their preferences. Furthermore, we conducted an online emotion categorization task to investigate how age, emotion category, type of SFU (with varying degrees of robot-likeness), listeners' gender, and their experience with robots relate to listeners' ability to categorize emotions. Results confirm that as age increases, there is a decline in emotion categorization performance of SFUs varying by emotion category and type of SFU.
Hideki Garcia Goo, Laura Ermers, Esther Janse, Jan Kolkmeier, Bob Schadenberg, Vanessa Evers, Khiet P. Truong
IEEE Trans. Affect. Comput.7
2025 Enhancing Transcripts of Open-Source Automatic Speech Recognition Models Through Fine-Tuning with Laughter and Speech-Laugh
abstract
Non-lexical sounds such as laughter are considered important discourse markers in conversation as these sounds complement the lexical information in shaping interpersonal relations, managing the conversation, and in expressing attitudes and affect. However, general purpose open-source automatic speech recognition (ASR) systems are typically not focusing on these sounds. Laughter transcription is an understudied task in ASR: laughter is often not modelled as a token and is discarded in the evaluation of ASR. In our study, we investigate how current open-source ASR models are handling laughter and speech-laugh (speech interspersed with laughter). Using Switchboard and BuckEye as conversational speech corpora, we fine-tuned and evaluated Whisper and Wav2Vec2. Our results show that laughter can be integrated in ASR transcriptions without substantially degrading word error rate.
Phuoc Hoang Ho, Dragos Alexandru Balan, Dirk Heylen, Khiet P. Truong
INTERSPEECH4
2025 Do I Sound as Capable as I Look? Impact of Robot Communication Style and Appearance on User Perception
abstract
Semantic-free utterances are thought to lower expectations of robot capabilities.However, to our knowledge, it still needs to be investigated how the type of robot communication (semantic-free utterance vs natural language), appearance (humanlike vs robotlike), and perceived voice-appearance alignment influence perceived robot capabilities.In a 2x2 online study, participants viewed videos with robots varying in communication and appearance.Linear mixed modelling showed that higher perceived voice-appearance alignment ratings were associated with higher perceived ratings of robot competency and agency.Contrary to expectations, semantic-free utterances and robot-like appearance were not significantly associated with perceived competency and agency.
Nevio F. Kavuza, Hideki Garcia Goo, Khiet P. Truong
IVA3
2025 The role of voice and appearance in gender perception of speaking robots
abstract
To enhance the design of gender-ambiguous speaking robots, designers could benefit from more insights into how people integrate robot appearance and voice in the gender perception of robots. In an online survey, we presented audio-only, visual-only, and audiovisual displays of gendered (and gender-ambiguous) robots and asked participants to rate the perceived level of femininity, masculinity, and gender-ambiguity of these robots. We investigated how the addition of an embodiment feature (voice, robot appearance) or gender (feminine, masculine, ambiguous) affected the perceived gender of speaking robots. Results showed a complex interplay between the variables under study: the magnitude and direction of effect on perceived gender are dependent on the embodiment feature, the gender of the added feature, and the gender identity of the basis that was added to. Furthermore, we found that the addition of a voice to a gender-ambiguous identity has more impact than the addition of a robot appearance on the perceived gender-ambiguity by lowering perceived ambiguity. For the design of gender-ambiguous robots, efforts should focus on making the appearance more ambiguous to counterbalance the gendered effect of voice.
Sjoerd van der Veen, Cesco Willemse, Hideki Garcia Goo, Khiet P. Truong
RO-MAN4
2024 A Conversational Robot for Children's Access to a Cultural Heritage Multimedia Archive
Thomas Beelen, Roeland Ordelman, Khiet P. Truong, Vanessa Evers, Theo Huibers
ECIR (5)3
2024 'Uhm... Are you sure?' An Exploratory Study of Trust Indicators in Robot-Directed Child Speech
abstract
In order to calibrate children’s trust in robots toward appropriate levels in the interaction, reliable trust measures are necessary. Current trust measures are not suitable for measuring children’s trust in a real-time manner. While speech from adult speakers has proven to contain information on their trust, this paper presents a first exploration investigating whether these results hold up in the context of a child-robot interaction. Fifty-eight conversations between children and robots were recorded (N=29), evoking high and low trust moments in the interaction. Correlation tests showed no (strong) predictors of children’s trust in their speech. Limitations and possibilities of how to advance the investigations of an automatic trust measure for child-robot interaction are discussed.
Ella Velner, Thomas Beelen, Bob Schadenberg, Roeland Ordelman, Theo Huibers, Khiet P. Truong, Vanessa Evers
IVA6
2023 Effects of perceived gender on the perceived social function of laughter
Joop Arts, Khiet P. Truong
INTERSPEECH2
2023 Laughter in task-based settings: whom we talk to affects how, when, and how often we laugh
abstract
Map task corpora are not typically used to study laughter, but they allow an interesting analysis of multiple factors such as familiarity between the participants, their gender, and eye contact. We conducted linear/generalized mixed-effects analysis to study if co-laughter, laughter rate, and the percentage of voiced frames in laughs are influenced by such factors. Our results show that, in conversations without eye contact, the gender of the participant was statistically relevant regarding laughter rate and the percentage of voiced frames, and the difference in gender was relevant regarding co-laughter. On the other hand, with eye contact, familiarity was statistically relevant with respect to co-laughter, laughter rate, and the percentage of voiced frames. Most of our results align and extend what has been previously found, except for voiced laughs between friends. This study emphasizes the highly variable character of laughter and its dependence on interlocutors' characteristics.
Catarina Branco, Isabel Trancoso, Paulo Infante, Khiet P. Truong
INTERSPEECH4
2023 Acoustic characteristics of depression in older adults' speech: the role of covariates
abstract
Contains fulltext : 299421.pdf (Publisher’s version ) (Open Access)
Carmen Mijnders, Esther Janse, Paul Naarding, Khiet P. Truong
INTERSPEECH4
2022 The Robot That Showed Remorse: Repairing Trust with a Genuine Apology
abstract
In the current state-of-the-art, robots are bound to make errors in a human-robot interaction (HRI). Trust is one of the important concepts in HRI that is often lowered by these errors. Fortunately, research has shown there are strategies that can help rebuild trust. An apology made by the robot is one of those strategies. However, apologies can take different forms. We designed a study in which Nao first built trust with the users, then violated that trust by making a speech recognition error, and then tried to restore it by either an apology with display of remorse, without remorse, or no apology at all. The results showed expected trends; an apology with remorse in most cases rebuilt trust the strongest. Although the effect of the type of apology on the trusting beliefs were not significant, the effect on the trusting behaviours was found to be just significant. Suggestions for future research include repeating the study without its current limitations (small sample size, offline) and investigating the accuracy of the portrayed remorse by the robot.
Babiche L. Pompe, Ella Velner, Khiet P. Truong
RO-MAN3
2021 How Familiarity Influences the Frequency, Temporal Dynamics and Acoustics of Laughter
abstract
Laughter is an affective and social signal that serves many functions. Similar to other social affective signals, laughter production and perception is at least partially context dependent. Familiarity of conversation partners has been shown to be a contextual influence on laughter production. However, the literature is still scarce and divided on how this complex interaction between familiarity and laughter works. Our goal with this paper is to further study this interaction using a newly acquired and annotated corpus and contrast our findings with existing findings in a comprehensive overview of the literature. Using a series of Linear Mixed-Effect Models, we studied if familiarity with the conversational partner and the sex of same-sex conversation pairs affect the laughter frequency, co-laughter frequency or laughter acoustics produced by the subjects in the corpus. The model outputs show that the frequency of laughter and co-laughter is not influenced by familiarity or the sex of conversation pairs. Interestingly, the percentage of co-laughs is significantly influenced by familiarity. Laughter voicedness is influenced by both the familiarity and sex of conversation pairs, where duration of laughter is only influenced by familiarity of the conversation pair. We conclude that familiarity of conversation pairs play an important role in laughter production during interactions and should be systematically explored. Furthermore we make several suggestions for improving the methods in future work.
Michel-Pierre Jansen, Khiet P. Truong, Dirk Heylen
ACII2
2021 Uncanny, Sexy, and Threatening Robots: The Online Community's Attitude to and Perceptions of Robots Varying in Humanlikeness and Gender
abstract
To get a better understanding of people's natural responses to humanlike robots outside the lab, we analyzed commentary on online videos depicting robots of different humanlikeness and gender. We built on previous work, which compared online video commentary of moderately and highly humanlike robots with respect to valence, uncanny valley, threats, and objectification. Additionally, we took into account the robot's gender, its appearance, its societal impact, the attribution of mental states, and how people attribute human stereotypes to robots. The results are mostly in line with previous work. Overall, the findings indicate that moderately humanlike robot design may be preferable over highly humanlike robot design because it is less associated with negative attitudes and perceptions. Robot designers should therefore be cautious when designing highly humanlike and gendered robots.
Quirien R. M. Hover, Ella Velner, Thomas Beelen, Mieke Boon, Khiet P. Truong
HRI5
2020 Introducing MULAI: A Multimodal Database of Laughter during Dyadic Interactions
abstract
Although laughter has gained considerable interest from a diversity of research areas, there still is a need for laughter specific databases. We present the Multimodal Laughter during Interaction (MULAI) database to study the expressive patterns of conversational and humour related laughter. The MULAI database contains 2 hours and 14 minutes of recorded and annotated dyadic human-human interactions and includes 601 laughs, 168 speech-laughs and 538 on- or offset respirations. This database is unique in several ways; 1) it focuses on different types of social laughter including conversational- and humour related laughter, 2) it contains annotations from participants, who understand the social context, on how humourous they perceived themselves and their interlocutor during each task, and 3) it contains data rarely captured by other laughter databases including participant personality profiles and physiological responses. We use the MULAI database to explore the link between acoustic laughter properties and annotated humour ratings over two settings. The results reveal that the duration, pitch and intensity of laughs from participants do not correlate with their own perception of how humourous they are, however the acoustics of laughter do correlate with how humourous they are being perceived by their conversational partner.
Michel-Pierre Jansen, Khiet P. Truong, Dirk Heylen, Deniece S. Nazareth
LREC2
2019 Context in Human Emotion Perception for Automatic Affect Detection: A Survey of Audiovisual Databases
abstract
An important aspect of human emotion perception is the use of contextual information to understand others' feelings even in situations where their behavior is not very expressive or has an emotionally ambiguous meaning. For technology to successfully detect affect, it must mimic this human ability when analyzing audiovisual input. Databases upon which machine learning algorithms are trained should capture the context of social interactions as well as the behavior expressed in them. However, there is a lack of consensus about what constitutes relevant context in such databases. In this article, we make two contributions towards overcoming this challenge: (a) we identify two principal sources of context for emotion perceptions based on psychological theory, and (b) we provide an overview of how each of these has been considered in published databases covering social interactions. Our results show that a similar set of contextual features are present across the reviewed databases. Between all the different databases researchers seem to have taken into account a set of contextual features reflecting the sources of context seen in psychological theory. However, within individual databases, these features are not yet systematically varied. This is problematic because it prevents them from being used directly as resources for the modeling of context-sensitive affect detection. Based on our findings, we suggest improvements for the future development of affective databases.
Bernd Dudzik, Michel-Pierre Jansen, Franziska Burger, Frank Kaptein, Joost Broekens, Dirk Heylen, Hayley Hung, Mark A. Neerincx, Khiet P. Truong
ACII9
2019 MEMOA: Introducing the Multi-Modal Emotional Memories of Older Adults Database
abstract
In order to contribute to the need of spontaneous multi-modal affective databases for the automatic recognition of emotions in older adults, this paper presents a novel Dutch multi-modal database consisting of emotional memories of older adults. The data consists of positive and negative memories of older adults eliciting through two emotion reliving tasks: autobiographical memory recall in the first session and life story books to discuss these memories in depth in the second session. Data collection was carried out at the participants' home or at a place comfortable to them. Audio was recorded for the first session whereas audio, video and physiological data were recorded for the second session. As this database introduces a novel way of using autobiographical memories to study emotional expressions in older adults, a first step of the complex coding of emotions is presented in this paper. We reflect on the challenges encountered in the database and propose ways to address these issues.
Deniece S. Nazareth, Michel-Pierre Jansen, Khiet P. Truong, Gerben Westerhof, Dirk Heylen
ACII3
2019 Emotional prosthesis for animating awe through performative biofeedback
abstract
Awe is a heightened emotional state of fear and wonder that creates a physiological response resulting in a cascade of hairs standing on end, also known as piloerection or goose-bumps. This latent sense once served an animalian purpose of survival, but now lies dormant and is often not experienced consciously. In fact, 55 percent of the population reports to not feel this sensation that is noted to be healthy. The AWE Goosebumps artifact is an emotion prosthesis that animates the latent sensation of awe for embodiment and externalizes cues for communication. As the sensation is not experienced consciously, the techno fashion invites an opportunity to be a second skin for frisson biofeedback, behavior training, and expression to others as a tool to transform the doldrums of modern day to performative states of wonder.
Kristin Neidlinger, Lianne Toussaint, Edwin Dertien, Khiet P. Truong, Hermie Hermens, Vanessa Evers
UbiComp4
2019 An Acoustic and Lexical Analysis of Emotional Valence in Spontaneous Speech: Autobiographical Memory Recall in Older Adults
abstract
Analyzing emotional valence in spontaneous speech remains complex and challenging. We present an acoustic and lexical analysis of emotional valence in spontaneous speech of older adults. Data was collected by recalling autobiographical memories through a word association task. Due to the complex and personal nature of memories, we propose a novel coding scheme for emotional valence. We explore acoustic properties of speech as well as the use of affective words to predict emotional valence expressed in autobiographical memories. Using mixed-effect regression modelling, we compared predictive models based on acoustic information only, lexical information only, or a combination of both. Results show that the combined model accounts for the highest proportion of explained variance, with the acoustic features accounting for a smaller share of the total variance than the lexical features. Several acoustic and lexical features predicted valence. As a first attempt at analyzing spontaneous emotional speech in older adults autobiographical memories, the study provides more insight in which acoustic features can be used to predict valence (automatically) in a more ecologically valid setting.
Deniece S. Nazareth, Ellen Tournier, Sarah Leimkötter, Esther Janse, Dirk Heylen, Gerben Westerhof, Khiet P. Truong
INTERSPEECH7
2019 Towards an Annotation Scheme for Complex Laughter in Speech Corpora
abstract
Although laughter research has gained quite some interest over the past few years, a shared description of how to annotate laughter and its sub-units is still missing. We present a first attempt towards an annotation scheme that contributes to improving the homogeneity and transparency with which laughter is annotated. This includes the integration of respiratory noises as well as stretches of speech-laughs, and to a limited extend to smiled speech and short silent intervals. Inter-annotator agreement is assessed while applying the scheme to different corpora where laughter is evoked through different methods and varying settings. Annotating laughter becomes more complex when the situation in which laughter occurs becomes more spontaneous and social. There is a substantial disagreement among the annotators with respect to temporal alignment (when does a unit start and when does it end) and unit classification, particularly the determination of starts/ends of laughter episodes. In summary, this detailed laughter annotation study reflects the need for better investigations of the various components of laughter.
Khiet P. Truong, Jürgen Trouvain, Michel-Pierre Jansen
INTERSPEECH1
2018 A Dyadic Conversation Dataset on Moral Emotions
abstract
In this paper, we present a dyadic conversation dataset involving topics related to moral emotions which are ethically relevant. To the best of our knowledge, it is the first dataset where the main focus is moral emotions. This dataset also focuses on speaker-listener reactions during a dyadic conversation. Although some of the currently available datasets contain dyadic conversations, they were not conceived with the idea of focusing on the speaker-listener setup. Thus making it difficult to use them to study reactions related to speakers and listeners. Some preliminary analyses of the data are presented as well as our thoughts on future work related to this dataset.
Louise Heron, Jaebok Kim, Minha Lee, Kevin El Haddad, Stéphane Dupont, Thierry Dutoit, Khiet P. Truong
FG7
2018 Nanogami: the microbiome expanded. speak your truth. listen to your gut
abstract
Nanogami is a bioresponsive garment to visualize the importance of the microbiome on collective wellbeing. The microbiome is the group of bacteria, viruses, and cells that live within and on our bodies. This galaxy of particles makes up more than half of the human body and are noted to be responsible for overall health and mood.
Kristin Neidlinger, Colin Willson, Khiet P. Truong, Hermie Hermens, Vanessa Evers
UbiComp3
2018 Automatic temporal ranking of children's engagement levels using multi-modal cues
Jaebok Kim, Khiet P. Truong, Vanessa Evers
Comput. Speech Lang.2
2017 Learning spectro-temporal features with 3D CNNs for speech emotion recognition
abstract
In this paper, we propose to use deep 3-dimensional convolutional networks (3D CNNs) in order to address the challenge of modelling spectro-temporal dynamics for speech emotion recognition (SER). Compared to a hybrid of Convolutional Neural Network and Long-Short-Term-Memory (CNN-LSTM), our proposed 3D CNNs simultaneously extract short-term and long-term spectral features with a moderate number of parameters. We evaluated our proposed and other state-of-the-art methods in a speaker-independent manner using aggregated corpora that give a large and diverse set of speakers. We found that 1) shallow temporal and moderately deep spectral kernels of a homogeneous architecture are optimal for the task; and 2) our 3D CNNs are more effective for spectro-temporal feature learning compared to other methods. Finally, we visualised the feature space obtained with our proposed method using t-distributed stochastic neighbour embedding (T-SNE) and could observe distinct clusters of emotions.
Jaebok Kim, Khiet P. Truong, Gwenn Englebienne, Vanessa Evers
ACII2
2017 Exploring moral conflicts in speech: Multidisciplinary analysis of affect and stress
abstract
Our moral conscience as the “inner light” that guides us shines brighter during moments of ethical conflicts, when we notice a tension between our many oughts and/or wants. We present the first analyses on speech related stress and affect in accounts of moral conflicts. For our exploratory study, we started with interviews on moral and immoral events at work with entrepreneurs. Qualitative analysis revealed that interviewees do share personal moral conflicts without researchers probing for them. Quantitative analysis showed quiet and even toned voice features when discussing moral conflicts, and speech was laced with emotively positive and negative words, though more negative words were used. Moreover, we find promising results on our automatic classification experiment using speech features. How and what moral conflicts people deliberate on in real-life may be pertinent to future research in affective computing, as well as applications for decision-making support, ethical competences coaching, therapy, and healthy moral selfhood.
Minha Lee, Jaebok Kim, Khiet P. Truong, Yvonne de Kort, Femke Beute, Wijnand A. IJsselsteijn
ACII3
2017 A Simple Nod of the Head: The Effect of Minimal Robot Movements on Children's Perception of a Low-Anthropomorphic Robot
abstract
In this note, we present minimal robot movements for robotic technology for children. Two types of minimal gaze movements were designed: social-gaze movements to communicate social engagement and deictic-gaze movements to communicate task-related referential information. In a two (social-gaze movements vs. none) by two (deictic-gaze movements vs. none) video-based study (n=72), we found that social-gaze movements significantly increased children's perception of animacy and likeability of the robot. Deictic-gaze and social-gaze movements significantly increased children's perception of helpfulness. Our findings show the compelling communicative power of social-gaze movements, and to a lesser extent deictic-gaze movements, and have implications for designers who want to achieve animacy, likeability and helpfulness with simple and easily implementable minimal robot movements. Our work contributes to human-robot interaction research and design by providing a first indication of the potential of minimal robot movements to communicate social engagement and helpful referential information to children.
Cristina Zaga, Roelof Anne Jelle de Vries, Jamy Li, Khiet P. Truong, Vanessa Evers
CHI4
2017 Towards Speech Emotion Recognition "in the Wild" Using Aggregated Corpora and Deep Multi-Task Learning
abstract
One of the challenges in Speech Emotion Recognition (SER) "in the wild" is the large mismatch between training and test data (e.g.speakers and tasks).In order to improve the generalisation capabilities of the emotion models, we propose to use Multi-Task Learning (MTL) and use gender and naturalness as auxiliary tasks in deep neural networks.This method was evaluated in within-corpus and various cross-corpus classification experiments that simulate conditions "in the wild".In comparison to Single-Task Learning (STL) based state of the art methods, we found that our MTL method proposed improved performance significantly.Particularly, models using both gender and naturalness achieved more gains than those using either gender or naturalness separately.This benefit was also found in the high-level representations of the feature space, obtained from our method proposed, where discriminative emotional clusters could be observed.
Jaebok Kim, Gwenn Englebienne, Khiet P. Truong, Vanessa Evers
INTERSPEECH3
2017 Deep Temporal Models using Identity Skip-Connections for Speech Emotion Recognition
abstract
Deep architectures using identity skip-connections have demonstrated groundbreaking performance in the field of image classification. Recently, empirical studies suggested that identity skip-connections enable ensemble-like behaviour of shallow networks, and that depth is not a solo ingredient for their success. Therefore, we examine the potential of identity skip-connections for the task of Speech Emotion Recognition (SER) where moderately deep temporal architectures are often employed. To this end, we propose a novel architecture which regulates unimpeded feature flows and captures long-term dependencies via gate-based skip-connections and a memory mechanism. Our proposed architecture is compared to other state-of-the-art methods of SER and is evaluated on large aggregated corpora recorded in different contexts. Our proposed architecture outperforms the state-of-the-art methods by 9 - 15% and achieves an Unweighted Accuracy of 80.5% in an imbalanced class distribution. In addition, we examine a variant adopting simplified skip-connections of Residual Networks (ResNet) and show that gate-based skip-connections are more effective than simplified skip-connections.
Jaebok Kim, Gwenn Englebienne, Khiet P. Truong, Vanessa Evers
ACM Multimedia3
2017 AWElectric: That Gave Me Goosebumps, Did You Feel It Too?
abstract
Awe is a powerful, visceral sensation described as a sudden chill or shudder accompanied by goosebumps. People feel awe in the face of extraordinary experiences: the sublimity of nature, the beauty of art and music, the adrenaline rush of fear. Awe is healthy, both physically and mentally. It can be shared by people who are witnessing the same phenomenon, but traditionally it cannot be communicated remotely across time or distance: to feel awe involves real time experience, and explaining the experience that gave rise to it does not always induce the feeling of awe itself. We want to make this sensation something that can be transmitted, and therefore present AWElectric, a wearable interface that can detect awe, enhance it, and create it in another person. Our shared goosebump design embeds inflatable biometric displays in 3D print fabric.The AudioTactile fabric transmits an awe-inducing sound frequency to the partner that physically manifests the tingles, chills, and goosebumps that awe provokes.
Kristin Neidlinger, Khiet P. Truong, Caty Telfair, Loe M. G. Feijs, Edwin Dertien, Vanessa Evers
TEI2
2017 A word of advice: how to tailor motivational text messages based on behavior change theory to personality and gender
abstract
Developing systems that motivate people to change their behaviors, such as an exercise application for the smartphone, is challenging. One solution is to implement motivational strategies from existing behavior change theory and tailor these strategies to preferences based on personal characteristics, like personality and gender. We operationalized strategies by collecting representative motivational text messages and aligning the messages to ten theory-based behavior change strategies. We conducted an online survey with 350 participants, where the participants rated 50 of our text messages (each aligned to one of the ten strategies) on how motivating they found them. Results show that differences in personality and gender relate to significant differences in the evaluations of nine out of ten strategies. Eight out of ten strategies were perceived as either more or less motivating in relation to scores on the personality traits Openness, Extraversion, and Agreeableness. Four strategies were perceived as more motivating by men than by women. These findings show that personality and gender influence how motivational strategies are perceived. We conclude that our theory-based behavior change strategies can be more motivating by tailoring them to personality and gender of users of behavior change systems.
Roelof Anne Jelle de Vries, Khiet P. Truong, Cristina Zaga, Jamy Li, Vanessa Evers
Pers. Ubiquitous Comput.2
2016 Crowd-Designed Motivation: Motivational Messages for Exercise Adherence Based on Behavior Change Theory
abstract
Developing motivational technology to support long-term behavior change is a challenge. A solution is to incorporate insights from behavior change theory and design technology to tailor to individual users. We carried out two studies to investigate whether the processes of change, from the Transtheoretical Model, can be effectively represented by motivational text messages. We crowdsourced peer-designed text messages and coded them into categories based on the processes of change. We evaluated whether people perceived messages tailored to their stage of change as motivating. We found that crowdsourcing is an effective method to design motivational messages. Our results indicate that different messages are perceived as motivating depending on the stage of behavior change a person is in. However, while motivational messages related to later stages of change were perceived as motivational for those stages, the motivational messages related to earlier stages of change were not. This indicates that a person's stage of change may not be the (only) key factor that determines behavior change. More individual factors need to be considered to design effective motivational technology.
Roelof Anne Jelle de Vries, Khiet P. Truong, Sigrid Kwint, C. H. C. Drossaert, Vanessa Evers
CHI2
2016 Help-Giving Robot Behaviors in Child-Robot Games: Exploring Semantic Free Utterances
abstract
We present initial findings from an experiment where we used Semantic Free Utterances - vocalizations and sounds without semantic content - as an alternative to Natural Language in a child-robot collaborative game. We tested (i) if two types of Semantic Free Utterances could be accurately recognized by the children; (ii) what effect the type of Semantic Free Utterances had as part of help-giving behaviors with in situ child-robot interaction. We discuss the potential benefits and pitfalls of Semantic Free Utterances for child-robot interaction.
Cristina Zaga, Roelof Anne Jelle de Vries, Sem J. Spenkelink, Khiet P. Truong, Vanessa Evers
HRI4
2016 ERM4CT 2016: 2nd international workshop on emotion representations and modelling for companion systems (workshop summary)
abstract
In this paper the organisers present a brief overview of the 2nd International Workshop on Emotion Representations and Modelling for Companion Systems (ERM4CT). The ERM4CT 2016 Workshop is held in conjunction with the 18th ACM International Conference on Multimodal Interaction (ICMI 2016) taking place Tokyo, Japan. The ERM4CT is the follow-up of three previous workshops on emotion modelling for affective human-computer interaction and companion systems. Apart from its usual focus on emotion representations and models, this year's ERM4CT puts special emphasis on how to model adequate affective system behaviour. For the first time, this year's ERM4CT gave out a dataset, which all attendees could investigate to jointly discuss their findings.
Kim Hartmann, Ingo Siegert, Albert Ali Salah, Khiet P. Truong
ICMI4
2016 ASSP4MI2016: 2nd international workshop on advancements in social signal processing for multimodal interaction (workshop summary)
abstract
This paper gives a summary of the 2nd International Workshop on Advancements in Social Signal Processing for Multimodal Interaction (ASSP4MI). Following our successful 1st International Workshop on Advancements in Social Signal Processing for Multimodal Interaction, held during ICMI-2015, we proposed the 2nd ASSP4MI workshop during ICMI-2016. The topics addressed and discussions fostered during last year's workshop are considered very relevant and alive in the research community. In this year's workshop, we continued addressing important topics and fostering fruitful discussions among researchers from different disciplines working in the fields of Social Signal Processing (SSP) and multimodal interaction.
Khiet P. Truong, Dirk Heylen, Toyoaki Nishida, Mohamed Chetouani
ICMI1
2016 Crowd-Designed Motivation: Combining Personality and the Transtheoretical Model
Roelof Anne Jelle de Vries, Khiet P. Truong, Vanessa Evers
PERSUASIVE2
2016 The Geneva Minimalistic Acoustic Parameter Set (GeMAPS) for Voice Research and Affective Computing
abstract
Work on voice sciences over recent decades has led to a proliferation of acoustic parameters that are used quite selectively and are not always extracted in a similar fashion. With many independent teams working in different research areas, shared standards become an essential safeguard to ensure compliance with state-of-the-art methods allowing appropriate comparison of results across studies and potential integration and combination of extraction and recognition systems. In this paper we propose a basic standard acoustic parameter set for various areas of automatic voice analysis, such as paralinguistic or clinical speech analysis. In contrast to a large brute-force parameter set, we present a minimalistic set of voice parameters here. These were selected based on a) their potential to index affective physiological changes in voice production, b) their proven value in former studies as well as their automatic extractability, and c) their theoretical significance. The set is intended to provide a common baseline for evaluation of future research and eliminate differences caused by varying parameter sets or even different implementations of the same parameters. Our implementation is publicly available with the openSMILE toolkit. Comparative evaluations of the proposed feature set and large baseline feature sets of INTERSPEECH challenges show a high performance of the proposed set in relation to its size.
Florian Eyben, Klaus R. Scherer, Björn W. Schuller, Johan Sundberg, Elisabeth André, Carlos Busso, Laurence Devillers, Julien Epps, Petri Laukka, Shri Narayanan, Khiet P. Truong
IEEE Trans. Affect. Comput.11
2015 Vocal turn-taking patterns in groups of children performing collaborative tasks: an exploratory study
abstract
Since children (5-9 years old) are still developing their emotional and social skills, their social interactional behaviors in small groups might differ from adults' interactional behaviors. In order to develop a robot that is able to support children performing collaborative tasks in small groups, it is necessary to gain a better understanding of how children interact with each other. We were interested in investigating vocal turn-taking patterns as we expect these to reveal relations to collaborative and conflict behaviors, especially with children behaviors as previous literature suggests. To that end, we collected an audiovisual corpus of children performing collaborative tasks together in groups of three. Through automatic turn-taking analyses, our results showed that speaker changes with overlaps are more common than without overlaps and children seemed to show smoother turn-taking patterns, i.e., less frequent and longer lasting speaker changes, during collaborative than conflict behaviors.
Jaebok Kim, Khiet P. Truong, Vicky Charisi, Cristina Zaga, Manja Lohse, Dirk Heylen, Vanessa Evers
INTERSPEECH2
2015 Prosodic characteristics of read speech before and after treadmill running
abstract
Physical activity leads to a respiratory behaviour that is very different to a resting state and that influences speech production. How speech parameters are exactly affected by physical activity remains largely unknown. Hence, we investigated how several prosodic parameters change under influence of physical activity and focused on temporal and breathing characteristics which have not been addressed in detail before. Speech from subjects reading aloud a text before and after a treadmill running exercise was analysed for prosodic differences between before and after running. The most important findings include a higher articulation rate, longer averaged pause and breath durations, a higher in-breath intensity, a higher out-breath rate, and a higher mean F0 for speech recorded immediately after vigorous treadmill running. These findings provide fundamental insights into how speech characteristics are affected by physical effort, and may help advance automatic classification of physical stress in speech.
Jürgen Trouvain, Khiet P. Truong
INTERSPEECH2
2015 A database for analysis of speech under physical stress: detection of exercise intensity while running and talking
abstract
One of the ways to gauge your own exercise intensity while running, is to assess your capability of talking while running: if you can still speak comfortably, you are running within the recommended intensity guidelines. This subjective way of estimating one's exercise intensity by talking (i.e. the Talk Test) motivated us to investigate how speech characteristics are affected during running and whether it is possible to develop a more objective way of estimating exercise intensity levels while running through voice analysis. To this end, we developed the Talk & Run Speech database that contains speech recorded from people before, during, and after running. We present our database and show that it is possible to detect exercise intensity below or above the anaerobic threshold in speech during running with a performance of 73.5% and 60.0% (unweighted average recall) for female and male speakers respectively.
Khiet P. Truong, Arne Nieuwenhuys, Peter Beek, Vanessa Evers
INTERSPEECH1
2015 Selecting the right robot: Influence of user attitude, robot sociability and embodiment on user preferences
abstract
Selecting the suitable form of a robot, i.e. physical or virtual, for a task is not straightforward. The choice for a physical robot is not self-evident when the task is not physical but entirely social in nature. Results from previous studies comparing robots with different body types are found to be inconclusive. We performed a user study to provide a more sound comparison between a virtual and physical robot operating in a social setting. Besides body type, we manipulated the sociability of the robot. Our results show that 1) user preferences indicate that robot sociability is more important than body type for selecting a robot in a non-physical social setting, and 2) the user's attitude towards robots is an important moderating factor influencing robot preference.
Mike Ligthart, Khiet P. Truong
RO-MAN2
2015 Dynamics of social positioning patterns in group-robot interactions
abstract
When a mobile robot interacts with a group of people, it has to consider its position and orientation. We introduce a novel study aimed at generating hypotheses on suitable behavior for such social positioning, explicitly focusing on interaction with small groups of users and allowing for the temporal and social dynamics inherent in most interactions. In particular, the interactions we look at are approach, converse and retreat. In this study, groups of three participants and a telepresence robot (controlled remotely by a fourth participant) solved a task together while we collected quantitative and qualitative data, including tracking of positioning/orientation and ratings of the behaviors used. In the data we observed a variety of patterns that can be extrapolated to hypotheses using inductive reasoning. One such pattern/hypothesis is that a (telepresence) robot could pass through a group when retreating, without this affecting how comfortable that retreat is for the group members. Another is that a group will rate the position/orientation of a (telepresence) robot as more comfortable when it is aimed more at the center of that group.
Jered Vroon, Michiel Joosse, Manja Lohse, Jan Kolkmeier, Jaebok Kim, Khiet P. Truong, Gwenn Englebienne, Dirk Heylen, Vanessa Evers
RO-MAN6
2014 Investigating prosodic relations between initiating and responding laughs
abstract
In dialogue, it is not uncommon for people to laugh together. This joint laughter often results in overlapping laughter, consisting of an initiating laugh (the first one), and a responding laugh (the second one). In previous studies, we found that overlapping laughs are acoustically different from non-overlapping ones. So far, we have considered overlapping laughs as one category. Consequently, it is unknown whether there are also acoustic differences between initiating laughs and responding laughs. In this paper, we make a distinction between initiating, responding, and non-overlapping laughs and compare their acoustic characteristics. In particular, we will investigate the prosodic relations between initiating and responding laughs. Do these relations point to a form of accommodation and mimicry? To what extent are initiating and responding laughs paired to each other? The analyses were performed on two speech corpora containing spontaneous conversations between two speakers. Results show indications that initiating and responding laughs share several similar acoustic features that point towards accommodation and mimicry mechanisms.
Khiet P. Truong, Jürgen Trouvain
INTERSPEECH1
2014 An annotation scheme for sighs in spontaneous dialogue
abstract
Sighs are non-verbal vocalisations that can carry important information about a speaker’s emotional (and psychological) state. Although sighs are commonly associated with negative emotions (e.g. giving up on something, ‘a sigh of despair’, sad-ness), sighs can also be associated with positive emotions such as relief. In order to gain a better understanding of sighing as a social and affective signal in dialogue, and to advance towards an automatic classification and interpretation of the emotional content of sighs, it is necessary to learn more about the various phonetic characteristics of sighs. To that end, we developed an annotation scheme for sighs that takes the variation in phonetic form into account. Using this scheme, an oral history corpus containing emotionally-coloured dialogues was annotated for sighs. Results show that sighs can be annotated with a suffi-cient level of reliability (Cohen’s Kappa of 0.713), and that in-deed, various types of sighs can be identified as well (Cohen’s Kappa between 0.637 and 0.805). Through a preliminary anal-ysis of emotional content words, indications were found that certain types of sighs can be associated with specific emotional contexts. Index Terms: sigh, non-verbal vocalisation, dialogue, social signal processing, annotation, oral history
Khiet P. Truong, Gerben Westerhof, Franciska de Jong, Dirk Heylen
INTERSPEECH1
2013 Detection of nonverbal vocalizations using Gaussian mixture models: looking for fillers and laughter in conversational speech
abstract
In this paper, we analyze acoustic profiles of fillers (i.e. filled pauses, FPs) and laughter with the aim to automatically localize these nonverbal vocalizations in a stream of audio. Among other features, we use voice quality features to capture the distinctive production modes of laughter and spectral similarity measures to capture the stability of the oral tract that is characteristic for FPs. Classification experiments with Gaussian Mixture Models and various sets of features are performed. We find that Mel-Frequency Cepstrum Coefficients are performing relatively well in comparison to other features for both FPs and laughter. In order to address the large variation in the frame-wise decision scores (e.g., log-likelihood ratios) observed in sequences of frames we apply a median filter to these scores, which yields large performance improvements. Our analyses and results are presented within the framework of this year’s Interspeech Computational Paralinguistics sub-Challenge on Social Signals.
Teun F. Krikke, Khiet P. Truong
INTERSPEECH2
2013 Classification of cooperative and competitive overlaps in speech using cues from the context, overlapper, and overlappee
abstract
One of the major properties of overlapping speech is that it can be perceived as competitive or cooperative. For the development of real-time spoken dialog systems and the analysis of affective and social human behavior in conversations, it is important to (automatically) distinguish between these two types of overlap. We investigate acoustic characteristics of cooperative and competitive overlaps with the aim to develop automatic classifiers for the classification of overlaps. In addition to acoustic features, we also use information from gaze and head movement annotations. Contexts preceding and during the overlap are taken into account, as well as the behaviors of both the overlapper and the overlappee. We compare various feature sets in classification experiments that are performed on the AMI corpus. The best performances obtained lie around 27%–30% EER.
Khiet P. Truong
INTERSPEECH1
2013 Perceptual evaluation of backchannel strategies for artificial listeners
Ronald Poppe, Khiet P. Truong, Dirk Heylen
Auton. Agents Multi Agent Syst.2
2012 Measuring prosodic alignment in cooperative task-based conversations
abstract
In this paper, we investigate prosodic alignment in task-based conversations. We use the HCRC Map Task Corpus and investigate how familiarity affects prosodic alignment and how task success is related to prosodic alignment. A variety of existing alignment measures is used and applied to our data. In particular, a windowed cross-correlation procedure, that has been used previously in visual behavior research, is applied to prosodic features. In addition, we address the issue of how to separate genuine observed alignment from alignment that is a result from random coincidental behavior. Using these measures, we find some indications of prosodic convergence and synchrony in the map task conversations. Alignment tendencies are strongest for intensity, and familiarity seems to play a role in convergence. Finally, weak evidence was found for a correlation between prosodic alignment measures and task success.
Khiet P. Truong, Dirk Heylen
INTERSPEECH1
2012 On the acoustics of overlapping laughter in conversational speech
abstract
The social nature of laughter invites people to laugh together. This joint vocal action often results in overlapping laughter. In this paper, we show that the acoustics of overlapping laughs are different from non-overlapping laughs. We found that overlapping laughs are stronger prosodically marked than non-overlapping ones, in terms of higher values for duration, mean F0, mean and maximum intensity, and the amount of voicing. This effect is intensified by the number of people joining in the laughter event, which suggests that entrainment is at work. We also found that group size affects the number of overlapping laughs which illustrates the contagious nature of laughter. Finally, people appear to join laughter simultaneously at a delay of approximately 500 ms; a delay that must be considered when developing spoken dialogue systems that are able to respond to users’ laughs.
Khiet P. Truong, Jürgen Trouvain
INTERSPEECH1
2012 Speech-based recognition of self-reported and observed emotion in a dimensional space
Khiet P. Truong, David A. van Leeuwen, Franciska de Jong
Speech Commun.1
2011 Automatic Understanding of Affective and Social Signals by Multimodal Mimicry Recognition
Xiaofan Sun, Anton Nijholt, Khiet P. Truong, Maja Pantic
ACII (2)3
2011 Online detection of vocal Listener Responses with maximum latency constraints
abstract
When human listeners utter Listener Responses (e.g. back-channels or acknowledgments) such as 'yeah' and 'mmhmm', interlocutors commonly continue to speak or resume their speech even before the listener has finished his/her response. This type of speech interactivity results in frequent speech overlap which is common in human human conversation. To allow for this type of speech interactivity to occur between humans and spoken dialog systems, which will result in more human-like continuous and smoother human-machine inter action, we propose an on-line classifier which can classify incoming speech as Listener Responses. We show that it is possible to detect vocal Listener Responses using maximum latency thresholds of 100-500 ms, thereby obtaining equal error rates ranging from 34% to 28% by using an energy based voice activity detector.
Daniel Neiberg, Khiet P. Truong
ICASSP2
2011 A Multimodal Analysis of Vocal and Visual Backchannels in Spontaneous Dialogs
abstract
Backchannels (BCs) are short vocal and visual listener responses that signal attention, interest, and understanding to the speaker. Previous studies have investigated BC prediction in telephone-style dialogs from prosodic cues. In contrast, we consider spontaneous face-to-face dialogs. The additional visual modality allows speaker and listener to monitor each other's attention continuously, and we hypothesize that this affects the BC-inviting cues. In this study, we investigate how gaze, in addition to prosody, can cue BCs. Moreover, we focus on the type of BC performed, with the aim to find out whether vocal and visual BCs are invited by similar cues. In contrast to telephone-style dialogs, we do not find rising/falling pitch to be a BC-inviting cue. However, in a face-to-face setting, gaze appears to cue BCs. In addition, we find that mutual gaze occurs significantly more often during visual BCs. Moreover, vocal BCs are more likely to be timed during pauses in the speaker's speech.
Khiet P. Truong, Ronald Poppe, Iwan de Kok, Dirk Heylen
INTERSPEECH1
2011 Backchannels: Quantity, Type and Timing Matters
Ronald Poppe, Khiet P. Truong, Dirk Heylen
IVA2
2011 Towards visual and vocal mimicry recognition in human-human interactions
abstract
During face-to-face interpersonal interaction, people have a tendency to mimic each other. People not only mimic postures, mannerisms, moods or emotions, but they also mimic several speech-related behaviors. In this paper we describe how visual and vocal behavioral information expressed between two interlocutors can be used to detect and identify visual and vocal mimicry. We investigate expressions of mimicry and aim to learn more about in which situation and to what extent mimicry occurs. The observable effects of mimicry can be explored by representing and recognizing mimicry using visual and vocal features. In order to automatically analyze how to extract and integrate this behavioral information into a multimodal mimicry detection framework for improving affective computing, this paper addresses the main challenge: mimicry representation in terms of optimal behavioral feature extraction and automatic integration in both audio and video modalities.
Xiaofan Sun, Khiet P. Truong, Maja Pantic, Anton Nijholt
SMC2
2010 Towards affective state modeling in narrative and conversational settings
abstract
We carry out two studies on affective state modeling for communication settings that involve unilateral intent on the part of one participant (the evoker) to shift the affective state of another participant (the experiencer). The first investigates viewer response in a narrative setting using a corpus of docu-mentaries annotated with viewer-reported narrative peaks. The second investigates affective triggers in a conversational set-ting using a corpus of recorded interactions, annotated with continuous affective ratings, between a human interlocutor and an emotionally colored agent. In each case, we build a “one-sided ” model using indicators derived from the speech of one participant. Our classification experiments confirm the viabil-ity of our models and provide insight into useful features. Index Terms: affect, speech recognition, audio analysis, natural language communication
Bart Jochems, Martha A. Larson, Roeland Ordelman, Ronald Poppe, Khiet P. Truong
INTERSPEECH5
2010 Disambiguating the functions of conversational sounds with prosody: the case of 'yeah'
abstract
In this paper, we look at how prosody can be used to automatically distinguish between different dialogue act functions and how it determines degree of speaker incipiency. We focus on the different uses of 'yeah'. Firstly, we investigate ambiguous dialogue act functions of 'yeah': 'yeah' is most frequently used as a backchannel or an assessment. Secondly, we look at the degree of speakership incipiency of 'yeah': some 'yeah' items display a greater intent of the speaker to take the floor. Classi﬿cation experiments with decision trees were performed to assess the role of prosody: we found that prosody indeed plays a role in disambiguating dialogue act functions and in determining degree of speaker incipiency of 'yeah'.
Khiet P. Truong, Dirk Heylen
INTERSPEECH1
2010 A rule-based backchannel prediction model using pitch and pause information
abstract
We manually designed rules for a backchannel (BC) prediction model based on pitch and pause information. In short, the model predicts a BC when there is a pause of a certain length that is preceded by a falling or rising pitch. This model was validated against the Dutch IFADV Corpus in a corpus-based evaluation method. The results showed that our model performs slightly better than another well-known rule-based BC prediction model that uses only pitch information. We observed that the length of a pause preceding a BC is one of the important features in this model, next to the duration of the pitch slope at the end of an utterance. Further, we discuss implications of a corpus-based approach to BC prediction evaluation.
Khiet P. Truong, Ronald Poppe, Dirk Heylen
INTERSPEECH1
2010 How Turn-Taking Strategies Influence Users' Impressions of an Agent
Mark ter Maat, Khiet P. Truong, Dirk Heylen
IVA2
2010 Backchannel Strategies for Artificial Listeners
Ronald Poppe, Khiet P. Truong, Dennis Reidsma, Dirk Heylen
IVA2
2010 Automatic role recognition based on conversational and prosodic behaviour
abstract
This paper proposes an approach for the automatic recognition of roles in settings like news and talk-shows, where roles correspond to specific functions like Anchorman, Guest or Interview Participant. The approach is based on purely nonverbal vocal behavioral cues, including who talks when and how much (turn-taking behavior), and statistical properties of pitch, formants, energy and speaking rate (prosodic behavior). The experiments have been performed over a corpus of around 50 hours of broadcast material and the accuracy, percentage of time correctly labeled in terms of role, is up to 89%. Both turn-taking and prosodic behavior lead to satisfactory results. Furthermore, on one database, their combination leads to a statistically significant improvement.
Hugues Salamin, Alessandro Vinciarelli, Khiet P. Truong, Gelareh Mohammadi
ACM Multimedia3
2009 Arousal and valence prediction in spontaneous emotional speech: felt versus perceived emotion
abstract
Contains fulltext : 91351.pdf (author's version ) (Open Access)
Khiet P. Truong, David A. van Leeuwen, Mark A. Neerincx, Franciska de Jong
INTERSPEECH1
2009 Comparing different approaches for automatic pronunciation error detection
Helmer Strik, Khiet P. Truong, Febe de Wet, Catia Cucchiarini
Speech Commun.2
2008 Multimodal Subjectivity Analysis of Multiparty Conversation
Stephan Raaijmakers, Khiet P. Truong, Theresa Wilson
EMNLP2
2008 Assessing agreement of observer- and self-annotations in spontaneous multimodal emotion data
abstract
Contains fulltext : 91352.pdf (Publisher’s version ) (Open Access)
Khiet P. Truong, Mark A. Neerincx, David A. van Leeuwen
INTERSPEECH1
2007 An open-set detection evaluation methodology applied to language and emotion recognition
abstract
This paper introduces a detection methodology for recognition technologies in speech for which it is dif cult to obtain an abundance of non-target classes. An example is language recognition, where we would like to be able to measure the detection capability of a single target language without confounding with the modeling capability of non-target languages. The evaluation framework is based on a cross validation scheme leaving the non-target class out of the allowed training material for the detector. The framework allows us to use Detection Error Tradeoff curves properly. As another application example we apply the evaluation scheme to emotion recognition in order to obtain single-emotion detection performance assessment. Index Terms: detection methodology, open-set evaluation, language, emotion.
David A. van Leeuwen, Khiet P. Truong
INTERSPEECH2
2007 Comparing classifiers for pronunciation error detection
abstract
CITATION: Strik, H. et al. 2007. Comparing classifiers for pronunciation error detection. In Hamme, H. van; Son, R. van (ed.), Proceedings of Interspeech 2007, pp. 1837-1840.
Helmer Strik, Khiet P. Truong, Febe de Wet, Catia Cucchiarini
INTERSPEECH2
2007 Visualizing acoustic similarities between emotions in speech: an acoustic map of emotions
abstract
In this paper, we introduce a visual analysis method to assess the discriminability and confusiability between emotions according to automatic emotion classifiers. The degree of acoustic similarities between emotions can be defined in terms of distances that are based on pair-wise emotion discrimination experiments. By employing Multidimensional Scaling, the discriminability between emotions can then be visualized in a two-dimensional plot that is relatively easy to interpret. This ‘map of emotions’ is compared to the well-known ‘Feeltrace’ two-dimensional mapping of emotions. While there is correlation with the ‘arousal’ dimension of Feeltrace, it appears that the ‘valence’ dimension is difficult to relate to the acoustic map.
Khiet P. Truong, David A. van Leeuwen
INTERSPEECH1
2007 Automatic discrimination between laughter and speech
Khiet P. Truong, David A. van Leeuwen
Speech Commun.1
2005 Automatic detection of laughter
abstract
In the context of detecting ‘paralinguistic events’ with the aim to make classification of the speaker’s emotional state possible, a detector was developed for one of the most obvious ‘paralinguistic events’, namely laughter. Gaussian Mixture Models were trained with Perceptual Linear Prediction features, pitch&energy, pitch&voicing and modulation spectrum features to model laughter and speech. Data from the ICSI Meeting Corpus and the Dutch CGN corpus were used for our classification experiments. The results showed that Gaussian Mixture Models trained with Perceptual Linear Prediction features performed best with Equal Error Rates ranging from 7.1%-20.0%.
Khiet P. Truong, David A. van Leeuwen
INTERSPEECH1
2005 Automatic detection of frequent pronunciation errors made by L2-learners
abstract
Contains fulltext : 41035.pdf (Publisher’s version ) (Open Access)
Khiet P. Truong, Ambra Neri, Febe de Wet, Catia Cucchiarini, Helmer Strik
INTERSPEECH1